You can not select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
 
 
 
 

28 KiB

实施计划:多模态输入支持

Spec: SPEC-multimodal-input.md 方案: B — content 从 string 改为 ContentPart[] 统一多模态结构 范围: 仅图片(PNG/JPG/WebP),先上传再引用 URL,单条消息最多 4 张


Phase 0:类型定义与基础设施

Task 0.1 — 扩展 app/types/chat.ts 类型定义

文件:app/types/chat.ts(修改)

内容:

  1. 新增 ContentTextPart、ContentImagePart、ContentPart 类型:

    export interface ContentTextPart {
      type: "text";
      text: string;
    }
    
    export interface ContentImagePart {
      type: "image";
      image: string;       // 服务器返回的 URL
      mimeType: string;    // image/png | image/jpeg | image/webp
      name?: string;       // 原始文件名
    }
    
    export type ContentPart = ContentTextPart | ContentImagePart;
    
  2. MessagePartType 新增 "image":

    export type MessagePartType =
      | "text"
      | "reasoning"
      | "tool-call"
      | "tool-result"
      | "tool-approval"
      | "image";
    
  3. MessagePart 接口新增 image 字段:

    export interface MessagePart {
      id: string;
      type: MessagePartType;
      text?: string;
      toolName?: string;
      toolCallId?: string;
      args?: unknown;
      result?: unknown;
      state?: string;
      image?: {
        url: string;
        mimeType: string;
        name?: string;
      };
    }
    

验证:npx vue-tsc --noEmit 无类型错误

提交:feat(types): add ContentPart and image MessagePart types for multimodal


Task 0.2 — 修复 Guest models API 保留 type 字段

文件:server/api/agents/[agentSlug]/models/index.get.ts(修改)

内容:

  1. systemModels.map 中新增 type: m.type:

    const list = systemModels.map((m) => ({
      id: m.id,
      name: m.name,
      modelId: m.modelId,
      providerName: "—",
      supportsTools: m.supportsTools,
      type: m.type,               // ← 新增
    }));
    
  2. 若有 defaultModelId 补充逻辑,同样加上 type: row.model.type。

验证:Guest 用户访问 /api/agents/[slug]/models 返回的 JSON 中每个 model 含 type 字段

提交:fix(api): preserve model type field in guest models API


Phase 1:后端改造

Task 1.1 — chat-engine.ts 支持 ContentPart[]

文件:server/service/agent/chat-engine.ts(修改)

内容:

  1. 导入类型:

    import type { ContentPart } from "#shared/types/chat";
    

    若 #shared/types/chat 路径不可用,则在 server/service/agent/chat-engine.ts 顶部内联定义 ContentPart 类型,或从 app/types/chat.ts 导入。

  2. ChatEngineParams.body 从 content: string 改为 parts: ContentPart[]:

    interface ChatEngineParams {
      // ...
      body: {
        parts: ContentPart[];          // ← 替代 content
        editMessageId?: string;
        regenerate?: boolean;
        continueAfterApproval?: boolean;
        approvalToolCallId?: string;
        approved?: boolean;
        approvalReason?: string;
        enableThinking?: boolean;
        enableTools?: boolean;
      };
      // ...
    }
    
  3. executeChat 中提取纯文本 + 序列化 parts JSON:

    // 从 parts 提取纯文本 content
    const textContent = body.parts
      .filter((p): p is { type: "text"; text: string } => p.type === "text")
      .map((p) => p.text)
      .join("\n");
    
    // parts JSON 包含所有 part(text + image)
    const partsJson = JSON.stringify(body.parts);
    
    // saveMessage 调用改为
    await saveMessage(sessionId, "user", textContent, partsJson);
    

    注意:editMessageId 和 regenerate 场景下也需要从 body.parts 提取。

  4. buildModelMessages 中 user 消息处理改为解析 parts JSON:

    // 当前代码:
    // { type: "text", text: m.content }
    //
    // 改为:
    function parseUserContent(m: { content: string; parts: string | null }): string | ContentPart[] {
      if (m.parts) {
        try {
          const parsed = JSON.parse(m.parts) as ContentPart[];
          if (parsed.length > 0) {
            return parsed.map((p) => {
              if (p.type === "text") return { type: "text", text: p.text };
              if (p.type === "image") return { type: "image", image: p.image };
              return null;
            }).filter(Boolean);
          }
        } catch {
          // fallback
        }
      }
      return m.content;  // 兼容旧数据,返回 string
    }
    
    // 在 buildModelMessages 中:
    const userContent = parseUserContent(m);
    modelMessages.push({
      role: "user",
      content: userContent,
    });
    

    AI SDK streamText 的 messages 参数中,user 消息 content 可以是 string 或 Array<TextPart | ImagePart>。ImagePart 的 image 字段接受 URL 字符串。

验证:

  • 单元测试或手动测试:发送 { parts: [{type:"text",text:"hello"}] } → 后端正确存储 + 调用 LLM
  • 发送 { parts: [{type:"text",text:"看图"},{type:"image",image:"/static/upload/test.png",mimeType:"image/png"}] } → LLM 收到图片

提交:feat(chat-engine): support ContentPart[] for multimodal messages


Task 1.2 — Agent chat POST API 接收 parts

文件:server/api/agents/[agentSlug]/chat/index.post.ts(修改)

内容:

  1. body 解构从 content 改为 parts,兼容旧 content: string:

    const body = await readBody(event);
    const { sessionId, parts: rawParts, content: rawContent, editMessageId, regenerate, continueAfterApproval, approvalToolCallId, approved, approvalReason, enableThinking, enableTools } = body as {
      sessionId: string;
      parts?: ContentPart[];
      content?: string;  // 兼容旧客户端
      editMessageId?: string;
      regenerate?: boolean;
      continueAfterApproval?: boolean;
      approvalToolCallId?: string;
      approved?: boolean;
      approvalReason?: string;
      enableThinking?: boolean;
      enableTools?: boolean;
    };
    
    // 兼容处理:parts 优先,否则从 content 包装
    const parts: ContentPart[] = rawParts
      ?? (rawContent ? [{ type: "text", text: rawContent }] : []);
    
  2. 校验改为:

    const isApprovalContinue = continueAfterApproval === true && approvalToolCallId;
    if (!isApprovalContinue && parts.length === 0) {
      throw createError({ statusCode: 400, statusMessage: "参数无效" });
    }
    
  3. executeChat 调用中 body 改为传 parts:

    body: {
      parts,
      editMessageId,
      regenerate,
      continueAfterApproval,
      approvalToolCallId,
      approved,
      approvalReason,
      enableThinking,
      enableTools,
    },
    

验证:

  • 发送 { parts: [{type:"text",text:"hi"}], sessionId } → 正常
  • 发送 { content: "hi", sessionId }(旧格式)→ 正常兼容
  • 发送 { sessionId }(无内容)→ 400 错误

提交:feat(api): accept ContentPart[] in agent chat POST with string fallback


Task 1.3 — 独立 chat API 支持 parts

文件:server/api/llm/chat/index.post.ts(修改)

内容:

  1. body 新增 parts?: ContentPart[] 字段:

    const { modelId: llmModelId, messages, parts, enableThinking, enableTools } = body as {
      modelId: number;
      messages: any[];
      parts?: ContentPart[];
      enableThinking?: boolean;
      enableTools?: boolean;
    };
    
  2. 若 parts 存在,将其追加到 messages 末尾作为最新 user 消息:

    let allMessages = messages;
    if (parts && parts.length > 0) {
      allMessages = [
        ...messages,
        {
          role: "user",
          content: parts.map((p) => {
            if (p.type === "text") return { type: "text", text: p.text };
            if (p.type === "image") return { type: "image", image: p.image };
            return null;
          }).filter(Boolean),
        },
      ];
    }
    const modelMessages = await convertToModelMessages(allMessages);
    

    convertToModelMessages 接受 UIMessage[],但这里 messages 是前端传来的格式。若 messages 已是 UIMessage[] 格式,则 parts 追加的 user 消息需符合 UIMessage 结构。需根据实际前端 useChat 发送格式调整。

    简化方案:如果独立 chat API 前端使用 useChat(AI SDK),useChat 的 sendMessage 已支持 FileUIPart,则前端只需在 useChat 的消息 parts 中包含 { type: "file", mediaType: "image/...", url: "..." },convertToModelMessages 自动处理。此情况下后端无需修改。

    决策:先检查独立 chat API 前端是否使用 useChat。若是,则此 Task 仅需确认 convertToModelMessages 已支持 FileUIPart(已验证支持),后端无需改动。若前端是自定义 fetch,则按上述方式追加 parts。

验证:

  • 确认独立 chat API 前端使用方式
  • 若使用 useChat:前端在 sendMessage 中附加 FileUIPart 即可,后端无需改动
  • 若自定义 fetch:发送含 parts 的请求 → LLM 收到图片

提交:feat(api): support multimodal parts in standalone chat API(若需修改)


Phase 2:前端 Composable 改造

Task 2.1 — useAgentChat.ts 支持 ContentPart[]

文件:app/composables/useAgentChat.ts(修改)

内容:

  1. 导入类型:

    import type { ContentPart, MessagePart } from "~/types/chat";
    
  2. AgentMessage 接口新增 contentParts:

    export interface AgentMessage {
      id: string;
      role: "user" | "assistant";
      content: string;
      contentParts?: ContentPart[];       // ← 新增
      parts?: MessagePart[];
      modelId?: string;
      inputTokens?: number;
      outputTokens?: number;
      createdAt?: string;
      feedback?: "like" | "dislike" | null;
    }
    
  3. send 方法签名从 send(content: string, opts?) 改为 send(parts: ContentPart[], opts?):

    async function send(parts: ContentPart[], opts?: {
      editMessageId?: string;
      regenerate?: boolean;
    }) {
      // ...
      const body = {
        parts,
        sessionId: sessionId.value,
        modelId: options.modelId,
        enableThinking: options.enableThinking,
        enableTools: options.enableTools,
        ...opts,
      };
      // fetch POST ...
    }
    
  4. normalizeMessage — 加载历史消息时兼容映射:

    function normalizeMessage(m: RawMessage): AgentMessage {
      let contentParts: ContentPart[] | undefined;
    
      if (m.parts) {
        try {
          const parsed = JSON.parse(m.parts);
          if (Array.isArray(parsed)) {
            contentParts = parsed.filter(
              (p: any) => p.type === "text" || p.type === "image"
            );
          }
        } catch {
          // ignore
        }
      }
    
      if (!contentParts) {
        contentParts = [{ type: "text", text: m.content }];
      }
    
      return {
        id: m.id,
        role: m.role,
        content: m.content,
        contentParts,
        parts: m.role === "assistant" ? parseAssistantParts(m.parts) : undefined,
        // ...其他字段
      };
    }
    

    注意:assistant 消息的 parts JSON 含 reasoning/tool-call 等,需单独解析。user 消息的 parts JSON 含 text/image content parts。需区分处理。

  5. 乐观更新 — send 中本地插入用户消息时使用 contentParts:

    // 提取纯文本用于 content 字段
    const textContent = parts
      .filter((p): p is ContentTextPart => p.type === "text")
      .map((p) => p.text)
      .join("\n");
    
    messages.value.push({
      id: "temp-" + Date.now(),
      role: "user",
      content: textContent,
      contentParts: parts,
      createdAt: new Date().toISOString(),
    });
    
  6. handleRegenerate — 从最后一条 user 消息提取 contentParts:

    function handleRegenerate() {
      const lastUserMsg = [...messages.value].reverse().find((m) => m.role === "user");
      if (lastUserMsg && lastUserMsg.contentParts) {
        send(lastUserMsg.contentParts, { regenerate: true });
      }
    }
    

验证:

  • npx vue-tsc --noEmit 无类型错误
  • 手动测试:发送文本 → 消息列表正确显示
  • 手动测试:加载历史消息 → 旧消息(无 parts)正确映射为 [{type:"text",text:content}]

提交:feat(useAgentChat): support ContentPart[] send and normalize legacy messages


Phase 3:前端组件改造

Task 3.1 — AgentInput.vue 支持图片上传

文件:app/components/agent/AgentInput.vue(修改)

内容:

  1. 新增 props:

    const props = defineProps<{
      editing?: boolean;
      supportsVision?: boolean;           // ← 新增
      editParts?: ContentPart[];          // ← 新增,编辑模式填充
    }>();
    
  2. 新增 emits:

    const emit = defineEmits<{
      send: [parts: ContentPart[]];       // ← 从 [content: string] 改为
      cancelEdit: [];
    }>();
    
  3. 新增图片状态:

    const images = ref<{ url: string; mimeType: string; name: string }[]>([]);
    const isUploading = ref(false);
    const fileInputRef = ref<HTMLInputElement | null>(null);
    
  4. 上传逻辑:

    async function uploadFiles(files: File[]) {
      if (!props.supportsVision) return;
      const imageFiles = files.filter((f) => f.type.startsWith("image/"));
      const remaining = 4 - images.value.length;
      if (imageFiles.length > remaining) {
        toast.warning(`最多 4 张图片,已添加 ${images.value.length} 张`);
      }
      const toUpload = imageFiles.slice(0, remaining);
      if (toUpload.length === 0) return;
    
      isUploading.value = true;
      try {
        const formData = new FormData();
        for (const f of toUpload) formData.append("files", f);
        const res = await $fetch<{ name: string; url: string; mimeType: string }[]>(
          "/api/file/upload",
          { method: "POST", body: formData }
        );
        for (const r of res) {
          images.value.push({ url: r.url, mimeType: r.mimeType, name: r.name });
        }
      } catch (e) {
        toast.error("图片上传失败");
      } finally {
        isUploading.value = false;
      }
    }
    
  5. 粘贴图片:

    function handlePaste(e: ClipboardEvent) {
      if (!props.supportsVision) return;
      const items = e.clipboardData?.items;
      if (!items) return;
      const files: File[] = [];
      for (const item of items) {
        if (item.type.startsWith("image/")) {
          const file = item.getAsFile();
          if (file) files.push(file);
        }
      }
      if (files.length > 0) {
        e.preventDefault();
        uploadFiles(files);
      }
    }
    
  6. 拖拽图片:

    function handleDrop(e: DragEvent) {
      if (!props.supportsVision) return;
      e.preventDefault();
      const files = Array.from(e.dataTransfer?.files ?? []);
      uploadFiles(files);
    }
    function handleDragOver(e: DragEvent) {
      if (!props.supportsVision) return;
      e.preventDefault();
    }
    
  7. 发送逻辑:

    function handleSend() {
      const parts: ContentPart[] = [];
      const text = inputText.value.trim();
      if (text) parts.push({ type: "text", text });
      for (const img of images.value) {
        parts.push({ type: "image", image: img.url, mimeType: img.mimeType, name: img.name });
      }
      if (parts.length === 0) return;
      emit("send", parts);
      inputText.value = "";
      images.value = [];
    }
    
  8. 编辑模式填充:

    watch(() => props.editParts, (parts) => {
      if (!parts) return;
      const textParts = parts.filter((p): p is ContentTextPart => p.type === "text");
      inputText.value = textParts.map((p) => p.text).join("\n");
      images.value = parts
        .filter((p): p is ContentImagePart => p.type === "image")
        .map((p) => ({ url: p.image, mimeType: p.mimeType, name: p.name ?? "image" }));
    }, { immediate: true });
    
  9. 删除图片:

    function removeImage(index: number) {
      images.value.splice(index, 1);
    }
    
  10. 模板新增:

    • textarea 上方图片预览区(缩略图 + 删除按钮)
    • 工具栏图片上传按钮(<input type="file" accept="image/*" multiple>,supportsVision 为 false 时隐藏)
    • textarea 绑定 @paste、@drop、@dragover
    • 上传中状态显示
  11. defineExpose 扩展:

    defineExpose({
      setText: (text: string) => { inputText.value = text; },
      setParts: (parts: ContentPart[]) => { /* 同 editParts watch 逻辑 */ },
      focus: () => textareaRef.value?.focus(),
    });
    

验证:

  • npx vue-tsc --noEmit 无类型错误
  • 手动测试:选择图片上传 → 预览 → 发送 → parts 含 image
  • 手动测试:粘贴图片 → 自动上传 → 预览
  • 手动测试:拖拽图片 → 自动上传 → 预览
  • 手动测试:supportsVision=false → 上传按钮隐藏,粘贴/拖拽无效
  • 手动测试:超过 4 张 → toast 提示
  • 手动测试:编辑模式 → editParts 填充文本 + 图片

提交:feat(AgentInput): add image upload, paste, drag-drop with ContentPart[] emit


Task 3.2 — AgentChatArea.vue 传递 supportsVision + emit 签名变更

文件:app/components/agent/AgentChatArea.vue(修改)

内容:

  1. ModelOption 接口新增 type:

    interface ModelOption {
      id: number;
      name: string;
      modelId: string;
      providerName: string;
      supportsTools: number;
      type?: "text" | "vision" | "multimodal";  // ← 新增
    }
    
  2. send emit 签名变更:

    const emit = defineEmits<{
      send: [parts: ContentPart[]];       // ← 从 [content: string] 改为
      edit: [messageId: string, parts: ContentPart[]];  // ← 从 [messageId, content: string] 改为
      regenerate: [];
      // ...其他不变
    }>();
    
  3. supportsVision computed:

    const supportsVision = computed(() => {
      if (props.modelId === null) return false;
      const model = props.models.find((m) => m.id === props.modelId);
      return model?.type === "vision" || model?.type === "multimodal";
    });
    
  4. 传给 AgentInput:

    <AgentInput
      :editing="!!editingMessageId"
      :supports-vision="supportsVision"
      :edit-parts="editParts"
      @send="handleSend"
      @cancel-edit="handleCancelEdit"
    />
    
  5. handleSend 签名变更:

    function handleSend(parts: ContentPart[]) {
      emit("send", parts);
    }
    
  6. handleEdit 签名变更:

    function handleEdit(messageId: string, parts: ContentPart[]) {
      editingMessageId.value = messageId;
      editParts.value = parts;
      emit("edit", messageId, parts);
    }
    
  7. 新增 editParts ref:

    const editParts = ref<ContentPart[]>([]);
    

验证:npx vue-tsc --noEmit 无类型错误

提交:feat(AgentChatArea): pass supportsVision, update emit signatures for ContentPart[]


Task 3.3 — AgentMessageItem.vue 渲染 contentParts

文件:app/components/agent/AgentMessageItem.vue(修改)

内容:

  1. 导入类型:

    import type { ContentPart } from "~/types/chat";
    
  2. 用户消息渲染改为遍历 contentParts:

    <div class="user-bubble">
      <template v-for="(part, idx) in userContentParts" :key="idx">
        <span v-if="part.type === 'text'">{{ part.text }}</span>
        <img
          v-else-if="part.type === 'image'"
          :src="part.image"
          :alt="part.name ?? ''"
          class="user-image"
          @click="previewImage(part.image)"
        />
      </template>
    </div>
    
  3. userContentParts computed:

    const userContentParts = computed<ContentPart[]>(() => {
      if (props.message.contentParts) return props.message.contentParts;
      return [{ type: "text", text: props.message.content }];
    });
    
  4. 编辑按钮 emit 改为传 ContentPart[]:

    function handleEdit() {
      emit("edit", props.message.id, userContentParts.value);
    }
    
  5. 图片预览:

    const previewUrl = ref<string | null>(null);
    function previewImage(url: string) {
      previewUrl.value = url;
    }
    function closePreview() {
      previewUrl.value = null;
    }
    
    <Teleport to="body">
      <div v-if="previewUrl" class="image-preview-overlay" @click="closePreview">
        <img :src="previewUrl" class="preview-image" />
      </div>
    </Teleport>
    
  6. 样式:

    .user-image {
      max-width: 200px;
      max-height: 200px;
      border-radius: 8px;
      margin: 4px 0;
      cursor: pointer;
      display: block;
    }
    .image-preview-overlay {
      position: fixed;
      inset: 0;
      background: rgba(0,0,0,0.8);
      display: flex;
      align-items: center;
      justify-content: center;
      z-index: 9999;
    }
    .preview-image {
      max-width: 90vw;
      max-height: 90vh;
    }
    

验证:

  • 用户消息含图片 → 正确渲染缩略图
  • 点击缩略图 → 放大预览
  • 旧消息(无 contentParts)→ fallback 到 content 纯文本
  • 编辑按钮 → emit 传 ContentPart[]

提交:feat(AgentMessageItem): render contentParts with image preview for user messages


Task 3.4 — AgentMessageParts.vue 新增 image part 渲染

文件:app/components/agent/AgentMessageParts.vue(修改)

内容:

新增 image 类型处理:

<template v-if="part.type === 'image'">
  <img
    :src="part.image?.url"
    :alt="part.image?.name ?? ''"
    class="assistant-image"
  />
</template>
.assistant-image {
  max-width: 300px;
  max-height: 300px;
  border-radius: 8px;
  margin: 8px 0;
}

验证:npx vue-tsc --noEmit 无类型错误

提交:feat(AgentMessageParts): add image part rendering


Task 3.5 — AgentModelSelector.vue + AgentToolbar.vue 扩展 ModelOption

文件:

  • app/components/agent/AgentModelSelector.vue(修改)
  • app/components/agent/AgentToolbar.vue(修改)

内容:

两个文件的 ModelOption 接口均新增 type 字段:

interface ModelOption {
  id: number;
  name: string;
  modelId: string;
  providerName: string;
  supportsTools: number;
  type?: "text" | "vision" | "multimodal";  // ← 新增
}

验证:npx vue-tsc --noEmit 无类型错误

提交:feat(agent): add type field to ModelOption in selector and toolbar


Task 3.6 — app/pages/chat/[agentSlug]/index.vue 页面入口

文件:app/pages/chat/[agentSlug]/index.vue(修改)

内容:

  1. ModelOption 接口新增 type:

    interface ModelOption {
      id: number;
      name: string;
      modelId: string;
      providerName: string;
      supportsTools: number;
      type?: "text" | "vision" | "multimodal";  // ← 新增
    }
    
  2. loadModels() map 时保留 type:

    models.value = list.map((m) => ({
      id: m.id,
      name: m.name,
      modelId: m.modelId,
      providerName: m.providerName ?? "—",
      supportsTools: m.supportsTools ?? 1,
      type: m.type ?? "text",
    }));
    
  3. handleSend 签名变更:

    function handleSend(parts: ContentPart[]) {
      chat.send(parts);
    }
    
  4. handleEdit 签名变更:

    function handleEdit(messageId: string, parts: ContentPart[]) {
      chat.send(parts, { editMessageId: messageId });
    }
    
  5. handleRegenerate 改为从最后一条 user 消息提取 contentParts:

    function handleRegenerate() {
      const lastUserMsg = [...chat.messages.value].reverse().find((m) => m.role === "user");
      if (lastUserMsg && lastUserMsg.contentParts) {
        chat.send(lastUserMsg.contentParts, { regenerate: true });
      }
    }
    

验证:

  • npx vue-tsc --noEmit 无类型错误
  • 页面加载 → models 列表含 type 字段
  • 发送消息 → chat.send(parts) 调用正确

提交:feat(page): update chat page for ContentPart[] send/edit/regenerate


Phase 4:集成测试与验证

Task 4.1 — 类型检查与 Lint

命令:

npx vue-tsc --noEmit
bun run lint

验证:无类型错误,无 lint 错误

提交:(无新提交,仅验证)


Task 4.2 — 手动集成测试

测试场景:

  1. 纯文本发送(向后兼容)

    • 输入文本 → 发送 → 消息列表显示 → LLM 回复
  2. 文本 + 图片发送

    • 选择 vision 模型 → 上传图片 → 输入文本 → 发送 → 消息列表显示文本 + 缩略图 → LLM 回复
  3. 纯图片发送

    • 上传图片(无文本)→ 发送 → LLM 回复
  4. 粘贴图片

    • 粘贴剪贴板中的图片 → 自动上传 → 预览 → 发送
  5. 拖拽图片

    • 拖拽图片文件到输入框 → 自动上传 → 预览 → 发送
  6. 图片预览

    • 点击消息中的缩略图 → 放大显示
  7. 最多 4 张限制

    • 添加 4 张图片 → 上传入口禁用
    • 尝试添加第 5 张 → toast 提示
  8. 非 vision 模型拦截

    • 选择 text 模型 → 上传按钮隐藏 → 粘贴/拖拽无效
  9. 切换模型提示

    • vision 模型下添加图片 → 切换到 text 模型 → toast 提示
  10. 编辑消息

    • 点击编辑按钮 → 输入框填充文本 + 图片 → 修改 → 发送 → 重新生成
  11. 重生成

    • 点击重生成 → 沿用原 user 消息的 contentParts(含图片)→ 重新生成
  12. 加载历史消息

    • 刷新页面 → 旧消息(无 parts)正确显示为纯文本
    • 新消息(含 image parts)正确显示文本 + 缩略图
  13. 独立 chat API(若适用)

    • /api/llm/chat 发送含图片的消息 → LLM 回复

验证:所有场景通过


Task 4.3 — E2E 测试更新(可选)

文件:e2e/specs/agent-chat.spec.ts(修改,可选)

内容:新增多模态测试用例:

  • 上传图片 + 发送 → 验证消息含图片
  • 非 vision 模型 → 验证上传按钮隐藏
  • 最多 4 张限制

验证:bun run test:e2e -- --grep "multimodal"

提交:test(e2e): add multimodal input tests(可选)


执行顺序

Phase 0: 类型定义 + Guest models API 修复
  ↓
Phase 1: 后端改造(chat-engine → POST API → 独立 chat API)
  ↓
Phase 2: 前端 composable(useAgentChat)
  ↓
Phase 3: 前端组件(AgentInput → AgentChatArea → AgentMessageItem → AgentMessageParts → ModelSelector/Toolbar → 页面入口)
  ↓
Phase 4: 集成测试与验证

风险与注意事项

  1. AI SDK ImagePart 格式:ImagePart 已废弃但仍支持,image 字段接受 URL 字符串。若 AI SDK 未来版本移除 ImagePart,需迁移到 FilePart。验证方式:实际发送图片消息确认 LLM 收到。

  2. parts JSON 语义重叠:agentMessages.parts 字段对 user 消息存 ContentPart[](text/image),对 assistant 消息存 MessagePart[](reasoning/tool-call 等)。buildModelMessages 需根据 role 区分解析逻辑。

  3. 独立 chat API 前端:需确认前端使用 useChat 还是自定义 fetch。若 useChat,后端可能无需改动(FileUIPart 自动处理)。

  4. 编辑模式状态管理:AgentChatArea 需维护 editParts ref,传给 AgentInput 的 editParts prop。取消编辑时需清空。

  5. 图片 URL 持久性:图片存储在服务器 /static/upload/ 目录,URL 为相对路径。需确认 LLM provider 能访问该 URL(若 LLM 在云端,需公网可访问的 URL)。若 LLM 无法访问内网 URL,需改为 Base64 内嵌或使用公网 URL。这是潜在阻塞点,需在 Phase 1 验证。