file: ./content/toc.en.mdx
meta: {
"title": "FastGPT Toc",
"description": "FastGPT Toc"
}
* [/en/faq/chat](/en/faq/chat)
* [/en/guide/admin/sso](/en/guide/admin/sso)
* [/en/guide/admin/teamMode](/en/guide/admin/teamMode)
* [/en/guide/build/agentv2/debug](/en/guide/build/agentv2/debug)
* [/en/guide/build/agentv2/settings](/en/guide/build/agentv2/settings)
* [/en/guide/build/agentv2/vm](/en/guide/build/agentv2/vm)
* [/en/guide/build/evaluation](/en/guide/build/evaluation)
* [/en/guide/build/faq](/en/guide/build/faq)
* [/en/guide/build/general/ai\_settings](/en/guide/build/general/ai_settings)
* [/en/guide/build/general/chat\_input\_guide](/en/guide/build/general/chat_input_guide)
* [/en/guide/build/general/fileInput](/en/guide/build/general/fileInput)
* [/en/guide/build/general/voiceInput](/en/guide/build/general/voiceInput)
* [/en/guide/build/general/welcomeText](/en/guide/build/general/welcomeText)
* [/en/guide/build/publish/dingtalk](/en/guide/build/publish/dingtalk)
* [/en/guide/build/publish/feishu](/en/guide/build/publish/feishu)
* [/en/guide/build/publish/link](/en/guide/build/publish/link)
* [/en/guide/build/publish/mcp\_server](/en/guide/build/publish/mcp_server)
* [/en/guide/build/publish/official\_account](/en/guide/build/publish/official_account)
* [/en/guide/build/publish/openapi](/en/guide/build/publish/openapi)
* [/en/guide/build/publish/wechat](/en/guide/build/publish/wechat)
* [/en/guide/build/publish/wecom](/en/guide/build/publish/wecom)
* [/en/guide/build/skill/development](/en/guide/build/skill/development)
* [/en/guide/build/skill/initialization](/en/guide/build/skill/initialization)
* [/en/guide/build/skill/integration](/en/guide/build/skill/integration)
* [/en/guide/build/skill/intro](/en/guide/build/skill/intro)
* [/en/guide/build/skill/version](/en/guide/build/skill/version)
* [/en/guide/build/tools/mcp\_tools](/en/guide/build/tools/mcp_tools)
* [/en/guide/build/tools/system-plugins/upload\_system\_tool](/en/guide/build/tools/system-plugins/upload_system_tool)
* [/en/guide/build/workflow/intro](/en/guide/build/workflow/intro)
* [/en/guide/build/workflow/nodes/ai\_chat](/en/guide/build/workflow/nodes/ai_chat)
* [/en/guide/build/workflow/nodes/content\_extract](/en/guide/build/workflow/nodes/content_extract)
* [/en/guide/build/workflow/nodes/coreferenceResolution](/en/guide/build/workflow/nodes/coreferenceResolution)
* [/en/guide/build/workflow/nodes/custom\_feedback](/en/guide/build/workflow/nodes/custom_feedback)
* [/en/guide/build/workflow/nodes/dataset\_search](/en/guide/build/workflow/nodes/dataset_search)
* [/en/guide/build/workflow/nodes/document\_parsing](/en/guide/build/workflow/nodes/document_parsing)
* [/en/guide/build/workflow/nodes/form\_input](/en/guide/build/workflow/nodes/form_input)
* [/en/guide/build/workflow/nodes/http](/en/guide/build/workflow/nodes/http)
* [/en/guide/build/workflow/nodes/knowledge\_base\_search\_merge](/en/guide/build/workflow/nodes/knowledge_base_search_merge)
* [/en/guide/build/workflow/nodes/loop](/en/guide/build/workflow/nodes/loop)
* [/en/guide/build/workflow/nodes/loop\_run](/en/guide/build/workflow/nodes/loop_run)
* [/en/guide/build/workflow/nodes/parallel\_run](/en/guide/build/workflow/nodes/parallel_run)
* [/en/guide/build/workflow/nodes/question\_classify](/en/guide/build/workflow/nodes/question_classify)
* [/en/guide/build/workflow/nodes/reply](/en/guide/build/workflow/nodes/reply)
* [/en/guide/build/workflow/nodes/sandbox-v2](/en/guide/build/workflow/nodes/sandbox-v2)
* [/en/guide/build/workflow/nodes/text\_editor](/en/guide/build/workflow/nodes/text_editor)
* [/en/guide/build/workflow/nodes/tfswitch](/en/guide/build/workflow/nodes/tfswitch)
* [/en/guide/build/workflow/nodes/tool](/en/guide/build/workflow/nodes/tool)
* [/en/guide/build/workflow/nodes/user-selection](/en/guide/build/workflow/nodes/user-selection)
* [/en/guide/build/workflow/nodes/variable\_update](/en/guide/build/workflow/nodes/variable_update)
* [/en/guide/chat/htmlRendering](/en/guide/chat/htmlRendering)
* [/en/guide/chat/quoteList](/en/guide/chat/quoteList)
* [/en/guide/dataset/collection\_tags](/en/guide/dataset/collection_tags)
* [/en/guide/dataset/dataset\_engine](/en/guide/dataset/dataset_engine)
* [/en/guide/dataset/faq](/en/guide/dataset/faq)
* [/en/guide/dataset/rag](/en/guide/dataset/rag)
* [/en/guide/dataset/template](/en/guide/dataset/template)
* [/en/guide/dataset/third-party/api\_dataset](/en/guide/dataset/third-party/api_dataset)
* [/en/guide/dataset/third-party/dingtalk\_dataset](/en/guide/dataset/third-party/dingtalk_dataset)
* [/en/guide/dataset/third-party/lark\_dataset](/en/guide/dataset/third-party/lark_dataset)
* [/en/guide/dataset/third-party/third\_dataset](/en/guide/dataset/third-party/third_dataset)
* [/en/guide/dataset/third-party/yuque\_dataset](/en/guide/dataset/third-party/yuque_dataset)
* [/en/guide/dataset/websync](/en/guide/dataset/websync)
* [/en/guide/getting-started/index](/en/guide/getting-started/index)
* [/en/guide/getting-started/quick-start](/en/guide/getting-started/quick-start)
* [/en/guide/index](/en/guide/index)
* [/en/guide/version/cloud/faq](/en/guide/version/cloud/faq)
* [/en/guide/version/cloud/intro](/en/guide/version/cloud/intro)
* [/en/guide/version/cloud/privacy](/en/guide/version/cloud/privacy)
* [/en/guide/version/cloud/terms](/en/guide/version/cloud/terms)
* [/en/guide/version/commercial](/en/guide/version/commercial)
* [/en/guide/version/opensource/intro](/en/guide/version/opensource/intro)
* [/en/guide/version/opensource/license](/en/guide/version/opensource/license)
* [/en/guide/workspace/customDomain](/en/guide/workspace/customDomain)
* [/en/guide/workspace/team/invitation\_link](/en/guide/workspace/team/invitation_link)
* [/en/guide/workspace/team/team\_roles\_permissions](/en/guide/workspace/team/team_roles_permissions)
* [/en/openapi/app](/en/openapi/app)
* [/en/openapi/chat](/en/openapi/chat)
* [/en/openapi/dataset](/en/openapi/dataset)
* [/en/openapi/index](/en/openapi/index)
* [/en/openapi/intro](/en/openapi/intro)
* [/en/plugin/index](/en/plugin/index)
* [/en/plugin/intro](/en/plugin/intro)
* [/en/plugin/model-presets](/en/plugin/model-presets)
* [/en/plugin/system-tool-development](/en/plugin/system-tool-development)
* [/en/self-host/config/env](/en/self-host/config/env)
* [/en/self-host/config/model/intro](/en/self-host/config/model/intro)
* [/en/self-host/config/model/minimax](/en/self-host/config/model/minimax)
* [/en/self-host/config/model/siliconCloud](/en/self-host/config/model/siliconCloud)
* [/en/self-host/config/object-storage](/en/self-host/config/object-storage)
* [/en/self-host/config/remote-debug-suite](/en/self-host/config/remote-debug-suite)
* [/en/self-host/config/sandbox/opensandbox](/en/self-host/config/sandbox/opensandbox)
* [/en/self-host/config/sandbox/sealosdevbox](/en/self-host/config/sandbox/sealosdevbox)
* [/en/self-host/config/signoz](/en/self-host/config/signoz)
* [/en/self-host/custom-models/bge-rerank](/en/self-host/custom-models/bge-rerank)
* [/en/self-host/custom-models/chatglm2](/en/self-host/custom-models/chatglm2)
* [/en/self-host/custom-models/chatglm2-m3e](/en/self-host/custom-models/chatglm2-m3e)
* [/en/self-host/custom-models/m3e](/en/self-host/custom-models/m3e)
* [/en/self-host/custom-models/marker](/en/self-host/custom-models/marker)
* [/en/self-host/custom-models/mineru](/en/self-host/custom-models/mineru)
* [/en/self-host/custom-models/ollama](/en/self-host/custom-models/ollama)
* [/en/self-host/custom-models/xinference](/en/self-host/custom-models/xinference)
* [/en/self-host/deploy/docker](/en/self-host/deploy/docker)
* [/en/self-host/deploy/sealos](/en/self-host/deploy/sealos)
* [/en/self-host/design/dataset](/en/self-host/design/dataset)
* [/en/self-host/dev](/en/self-host/dev)
* [/en/self-host/index](/en/self-host/index)
* [/en/self-host/migration/docker\_db](/en/self-host/migration/docker_db)
* [/en/self-host/migration/docker\_mongo](/en/self-host/migration/docker_mongo)
* [/en/self-host/troubleshooting/attention](/en/self-host/troubleshooting/attention)
* [/en/self-host/troubleshooting/faq](/en/self-host/troubleshooting/faq)
* [/en/self-host/troubleshooting/methods](/en/self-host/troubleshooting/methods)
* [/en/self-host/troubleshooting/model-errors](/en/self-host/troubleshooting/model-errors)
* [/en/self-host/troubleshooting/s3-issues](/en/self-host/troubleshooting/s3-issues)
* [/en/self-host/upgrading/4-12/4120](/en/self-host/upgrading/4-12/4120)
* [/en/self-host/upgrading/4-12/4121](/en/self-host/upgrading/4-12/4121)
* [/en/self-host/upgrading/4-12/4122](/en/self-host/upgrading/4-12/4122)
* [/en/self-host/upgrading/4-12/4123](/en/self-host/upgrading/4-12/4123)
* [/en/self-host/upgrading/4-12/4124](/en/self-host/upgrading/4-12/4124)
* [/en/self-host/upgrading/4-13/4130](/en/self-host/upgrading/4-13/4130)
* [/en/self-host/upgrading/4-13/4131](/en/self-host/upgrading/4-13/4131)
* [/en/self-host/upgrading/4-13/4132](/en/self-host/upgrading/4-13/4132)
* [/en/self-host/upgrading/4-14/4140](/en/self-host/upgrading/4-14/4140)
* [/en/self-host/upgrading/4-14/4141](/en/self-host/upgrading/4-14/4141)
* [/en/self-host/upgrading/4-14/41410](/en/self-host/upgrading/4-14/41410)
* [/en/self-host/upgrading/4-14/41411](/en/self-host/upgrading/4-14/41411)
* [/en/self-host/upgrading/4-14/41412](/en/self-host/upgrading/4-14/41412)
* [/en/self-host/upgrading/4-14/41413](/en/self-host/upgrading/4-14/41413)
* [/en/self-host/upgrading/4-14/41414](/en/self-host/upgrading/4-14/41414)
* [/en/self-host/upgrading/4-14/41415](/en/self-host/upgrading/4-14/41415)
* [/en/self-host/upgrading/4-14/41416](/en/self-host/upgrading/4-14/41416)
* [/en/self-host/upgrading/4-14/41419](/en/self-host/upgrading/4-14/41419)
* [/en/self-host/upgrading/4-14/4142](/en/self-host/upgrading/4-14/4142)
* [/en/self-host/upgrading/4-14/41420](/en/self-host/upgrading/4-14/41420)
* [/en/self-host/upgrading/4-14/41421](/en/self-host/upgrading/4-14/41421)
* [/en/self-host/upgrading/4-14/41422](/en/self-host/upgrading/4-14/41422)
* [/en/self-host/upgrading/4-14/41424](/en/self-host/upgrading/4-14/41424)
* [/en/self-host/upgrading/4-14/41425](/en/self-host/upgrading/4-14/41425)
* [/en/self-host/upgrading/4-14/41426](/en/self-host/upgrading/4-14/41426)
* [/en/self-host/upgrading/4-14/41427](/en/self-host/upgrading/4-14/41427)
* [/en/self-host/upgrading/4-14/41428](/en/self-host/upgrading/4-14/41428)
* [/en/self-host/upgrading/4-14/41429](/en/self-host/upgrading/4-14/41429)
* [/en/self-host/upgrading/4-14/4143](/en/self-host/upgrading/4-14/4143)
* [/en/self-host/upgrading/4-14/4144](/en/self-host/upgrading/4-14/4144)
* [/en/self-host/upgrading/4-14/4145](/en/self-host/upgrading/4-14/4145)
* [/en/self-host/upgrading/4-14/41451](/en/self-host/upgrading/4-14/41451)
* [/en/self-host/upgrading/4-14/4146](/en/self-host/upgrading/4-14/4146)
* [/en/self-host/upgrading/4-14/4147](/en/self-host/upgrading/4-14/4147)
* [/en/self-host/upgrading/4-14/4148](/en/self-host/upgrading/4-14/4148)
* [/en/self-host/upgrading/4-14/41481](/en/self-host/upgrading/4-14/41481)
* [/en/self-host/upgrading/4-14/4149](/en/self-host/upgrading/4-14/4149)
* [/en/self-host/upgrading/4-14/41930](/en/self-host/upgrading/4-14/41930)
* [/en/self-host/upgrading/4-15/41500](/en/self-host/upgrading/4-15/41500)
* [/en/self-host/upgrading/4-15/41501](/en/self-host/upgrading/4-15/41501)
* [/en/self-host/upgrading/4-15/41502](/en/self-host/upgrading/4-15/41502)
* [/en/self-host/upgrading/4-15/41503](/en/self-host/upgrading/4-15/41503)
* [/en/self-host/upgrading/4-15/41504](/en/self-host/upgrading/4-15/41504)
* [/en/self-host/upgrading/4-15/41505](/en/self-host/upgrading/4-15/41505)
* [/en/self-host/upgrading/4-15/41506](/en/self-host/upgrading/4-15/41506)
* [/en/self-host/upgrading/4-15/41507](/en/self-host/upgrading/4-15/41507)
* [/en/self-host/upgrading/4-15/4151](/en/self-host/upgrading/4-15/4151)
* [/en/self-host/upgrading/4-15/4152](/en/self-host/upgrading/4-15/4152)
* [/en/self-host/upgrading/4-15/4153](/en/self-host/upgrading/4-15/4153)
* [/en/self-host/upgrading/4-15/4154](/en/self-host/upgrading/4-15/4154)
* [/en/self-host/upgrading/4-15/4155](/en/self-host/upgrading/4-15/4155)
* [/en/self-host/upgrading/4-15/4156](/en/self-host/upgrading/4-15/4156)
* [/en/self-host/upgrading/4-15/4157](/en/self-host/upgrading/4-15/4157)
* [/en/self-host/upgrading/4-16/41601](/en/self-host/upgrading/4-16/41601)
* [/en/self-host/upgrading/4-16/41602](/en/self-host/upgrading/4-16/41602)
* [/en/self-host/upgrading/outdated/40](/en/self-host/upgrading/outdated/40)
* [/en/self-host/upgrading/outdated/41](/en/self-host/upgrading/outdated/41)
* [/en/self-host/upgrading/outdated/4100](/en/self-host/upgrading/outdated/4100)
* [/en/self-host/upgrading/outdated/4101](/en/self-host/upgrading/outdated/4101)
* [/en/self-host/upgrading/outdated/4110](/en/self-host/upgrading/outdated/4110)
* [/en/self-host/upgrading/outdated/4111](/en/self-host/upgrading/outdated/4111)
* [/en/self-host/upgrading/outdated/42](/en/self-host/upgrading/outdated/42)
* [/en/self-host/upgrading/outdated/421](/en/self-host/upgrading/outdated/421)
* [/en/self-host/upgrading/outdated/43](/en/self-host/upgrading/outdated/43)
* [/en/self-host/upgrading/outdated/44](/en/self-host/upgrading/outdated/44)
* [/en/self-host/upgrading/outdated/441](/en/self-host/upgrading/outdated/441)
* [/en/self-host/upgrading/outdated/442](/en/self-host/upgrading/outdated/442)
* [/en/self-host/upgrading/outdated/445](/en/self-host/upgrading/outdated/445)
* [/en/self-host/upgrading/outdated/446](/en/self-host/upgrading/outdated/446)
* [/en/self-host/upgrading/outdated/447](/en/self-host/upgrading/outdated/447)
* [/en/self-host/upgrading/outdated/45](/en/self-host/upgrading/outdated/45)
* [/en/self-host/upgrading/outdated/451](/en/self-host/upgrading/outdated/451)
* [/en/self-host/upgrading/outdated/452](/en/self-host/upgrading/outdated/452)
* [/en/self-host/upgrading/outdated/46](/en/self-host/upgrading/outdated/46)
* [/en/self-host/upgrading/outdated/461](/en/self-host/upgrading/outdated/461)
* [/en/self-host/upgrading/outdated/462](/en/self-host/upgrading/outdated/462)
* [/en/self-host/upgrading/outdated/463](/en/self-host/upgrading/outdated/463)
* [/en/self-host/upgrading/outdated/464](/en/self-host/upgrading/outdated/464)
* [/en/self-host/upgrading/outdated/465](/en/self-host/upgrading/outdated/465)
* [/en/self-host/upgrading/outdated/466](/en/self-host/upgrading/outdated/466)
* [/en/self-host/upgrading/outdated/467](/en/self-host/upgrading/outdated/467)
* [/en/self-host/upgrading/outdated/468](/en/self-host/upgrading/outdated/468)
* [/en/self-host/upgrading/outdated/469](/en/self-host/upgrading/outdated/469)
* [/en/self-host/upgrading/outdated/47](/en/self-host/upgrading/outdated/47)
* [/en/self-host/upgrading/outdated/471](/en/self-host/upgrading/outdated/471)
* [/en/self-host/upgrading/outdated/48](/en/self-host/upgrading/outdated/48)
* [/en/self-host/upgrading/outdated/481](/en/self-host/upgrading/outdated/481)
* [/en/self-host/upgrading/outdated/4810](/en/self-host/upgrading/outdated/4810)
* [/en/self-host/upgrading/outdated/4811](/en/self-host/upgrading/outdated/4811)
* [/en/self-host/upgrading/outdated/4812](/en/self-host/upgrading/outdated/4812)
* [/en/self-host/upgrading/outdated/4813](/en/self-host/upgrading/outdated/4813)
* [/en/self-host/upgrading/outdated/4814](/en/self-host/upgrading/outdated/4814)
* [/en/self-host/upgrading/outdated/4815](/en/self-host/upgrading/outdated/4815)
* [/en/self-host/upgrading/outdated/4816](/en/self-host/upgrading/outdated/4816)
* [/en/self-host/upgrading/outdated/4817](/en/self-host/upgrading/outdated/4817)
* [/en/self-host/upgrading/outdated/4818](/en/self-host/upgrading/outdated/4818)
* [/en/self-host/upgrading/outdated/4819](/en/self-host/upgrading/outdated/4819)
* [/en/self-host/upgrading/outdated/482](/en/self-host/upgrading/outdated/482)
* [/en/self-host/upgrading/outdated/4820](/en/self-host/upgrading/outdated/4820)
* [/en/self-host/upgrading/outdated/4821](/en/self-host/upgrading/outdated/4821)
* [/en/self-host/upgrading/outdated/4822](/en/self-host/upgrading/outdated/4822)
* [/en/self-host/upgrading/outdated/4823](/en/self-host/upgrading/outdated/4823)
* [/en/self-host/upgrading/outdated/483](/en/self-host/upgrading/outdated/483)
* [/en/self-host/upgrading/outdated/484](/en/self-host/upgrading/outdated/484)
* [/en/self-host/upgrading/outdated/485](/en/self-host/upgrading/outdated/485)
* [/en/self-host/upgrading/outdated/486](/en/self-host/upgrading/outdated/486)
* [/en/self-host/upgrading/outdated/487](/en/self-host/upgrading/outdated/487)
* [/en/self-host/upgrading/outdated/488](/en/self-host/upgrading/outdated/488)
* [/en/self-host/upgrading/outdated/489](/en/self-host/upgrading/outdated/489)
* [/en/self-host/upgrading/outdated/490](/en/self-host/upgrading/outdated/490)
* [/en/self-host/upgrading/outdated/491](/en/self-host/upgrading/outdated/491)
* [/en/self-host/upgrading/outdated/4910](/en/self-host/upgrading/outdated/4910)
* [/en/self-host/upgrading/outdated/4911](/en/self-host/upgrading/outdated/4911)
* [/en/self-host/upgrading/outdated/4912](/en/self-host/upgrading/outdated/4912)
* [/en/self-host/upgrading/outdated/4913](/en/self-host/upgrading/outdated/4913)
* [/en/self-host/upgrading/outdated/4914](/en/self-host/upgrading/outdated/4914)
* [/en/self-host/upgrading/outdated/492](/en/self-host/upgrading/outdated/492)
* [/en/self-host/upgrading/outdated/493](/en/self-host/upgrading/outdated/493)
* [/en/self-host/upgrading/outdated/494](/en/self-host/upgrading/outdated/494)
* [/en/self-host/upgrading/outdated/495](/en/self-host/upgrading/outdated/495)
* [/en/self-host/upgrading/outdated/496](/en/self-host/upgrading/outdated/496)
* [/en/self-host/upgrading/outdated/497](/en/self-host/upgrading/outdated/497)
* [/en/self-host/upgrading/outdated/498](/en/self-host/upgrading/outdated/498)
* [/en/self-host/upgrading/outdated/499](/en/self-host/upgrading/outdated/499)
* [/en/self-host/upgrading/upgrade-intruction](/en/self-host/upgrading/upgrade-intruction)
file: ./content/toc.mdx
meta: {
"title": "FastGPT 文档目录",
"description": "FastGPT 文档目录"
}
* [/faq/chat](/faq/chat)
* [/guide/admin/sso](/guide/admin/sso)
* [/guide/admin/teamMode](/guide/admin/teamMode)
* [/guide/build/agentv2/debug](/guide/build/agentv2/debug)
* [/guide/build/agentv2/settings](/guide/build/agentv2/settings)
* [/guide/build/agentv2/vm](/guide/build/agentv2/vm)
* [/guide/build/evaluation](/guide/build/evaluation)
* [/guide/build/faq](/guide/build/faq)
* [/guide/build/general/ai\_settings](/guide/build/general/ai_settings)
* [/guide/build/general/chat\_input\_guide](/guide/build/general/chat_input_guide)
* [/guide/build/general/fileInput](/guide/build/general/fileInput)
* [/guide/build/general/voiceInput](/guide/build/general/voiceInput)
* [/guide/build/general/welcomeText](/guide/build/general/welcomeText)
* [/guide/build/publish/dingtalk](/guide/build/publish/dingtalk)
* [/guide/build/publish/feishu](/guide/build/publish/feishu)
* [/guide/build/publish/link](/guide/build/publish/link)
* [/guide/build/publish/mcp\_server](/guide/build/publish/mcp_server)
* [/guide/build/publish/official\_account](/guide/build/publish/official_account)
* [/guide/build/publish/openapi](/guide/build/publish/openapi)
* [/guide/build/publish/wechat](/guide/build/publish/wechat)
* [/guide/build/publish/wecom](/guide/build/publish/wecom)
* [/guide/build/skill/development](/guide/build/skill/development)
* [/guide/build/skill/initialization](/guide/build/skill/initialization)
* [/guide/build/skill/integration](/guide/build/skill/integration)
* [/guide/build/skill/intro](/guide/build/skill/intro)
* [/guide/build/skill/version](/guide/build/skill/version)
* [/guide/build/tools/mcp\_tools](/guide/build/tools/mcp_tools)
* [/guide/build/tools/system-plugins/upload\_system\_tool](/guide/build/tools/system-plugins/upload_system_tool)
* [/guide/build/workflow/intro](/guide/build/workflow/intro)
* [/guide/build/workflow/nodes/ai\_chat](/guide/build/workflow/nodes/ai_chat)
* [/guide/build/workflow/nodes/content\_extract](/guide/build/workflow/nodes/content_extract)
* [/guide/build/workflow/nodes/coreferenceResolution](/guide/build/workflow/nodes/coreferenceResolution)
* [/guide/build/workflow/nodes/custom\_feedback](/guide/build/workflow/nodes/custom_feedback)
* [/guide/build/workflow/nodes/dataset\_search](/guide/build/workflow/nodes/dataset_search)
* [/guide/build/workflow/nodes/document\_parsing](/guide/build/workflow/nodes/document_parsing)
* [/guide/build/workflow/nodes/form\_input](/guide/build/workflow/nodes/form_input)
* [/guide/build/workflow/nodes/http](/guide/build/workflow/nodes/http)
* [/guide/build/workflow/nodes/knowledge\_base\_search\_merge](/guide/build/workflow/nodes/knowledge_base_search_merge)
* [/guide/build/workflow/nodes/loop](/guide/build/workflow/nodes/loop)
* [/guide/build/workflow/nodes/loop\_run](/guide/build/workflow/nodes/loop_run)
* [/guide/build/workflow/nodes/parallel\_run](/guide/build/workflow/nodes/parallel_run)
* [/guide/build/workflow/nodes/question\_classify](/guide/build/workflow/nodes/question_classify)
* [/guide/build/workflow/nodes/reply](/guide/build/workflow/nodes/reply)
* [/guide/build/workflow/nodes/sandbox-v2](/guide/build/workflow/nodes/sandbox-v2)
* [/guide/build/workflow/nodes/text\_editor](/guide/build/workflow/nodes/text_editor)
* [/guide/build/workflow/nodes/tfswitch](/guide/build/workflow/nodes/tfswitch)
* [/guide/build/workflow/nodes/tool](/guide/build/workflow/nodes/tool)
* [/guide/build/workflow/nodes/user-selection](/guide/build/workflow/nodes/user-selection)
* [/guide/build/workflow/nodes/variable\_update](/guide/build/workflow/nodes/variable_update)
* [/guide/chat/htmlRendering](/guide/chat/htmlRendering)
* [/guide/chat/quoteList](/guide/chat/quoteList)
* [/guide/dataset/collection\_tags](/guide/dataset/collection_tags)
* [/guide/dataset/dataset\_engine](/guide/dataset/dataset_engine)
* [/guide/dataset/faq](/guide/dataset/faq)
* [/guide/dataset/rag](/guide/dataset/rag)
* [/guide/dataset/template](/guide/dataset/template)
* [/guide/dataset/third-party/api\_dataset](/guide/dataset/third-party/api_dataset)
* [/guide/dataset/third-party/dingtalk\_dataset](/guide/dataset/third-party/dingtalk_dataset)
* [/guide/dataset/third-party/lark\_dataset](/guide/dataset/third-party/lark_dataset)
* [/guide/dataset/third-party/third\_dataset](/guide/dataset/third-party/third_dataset)
* [/guide/dataset/third-party/yuque\_dataset](/guide/dataset/third-party/yuque_dataset)
* [/guide/dataset/websync](/guide/dataset/websync)
* [/guide/getting-started/index](/guide/getting-started/index)
* [/guide/getting-started/quick-start](/guide/getting-started/quick-start)
* [/guide/index](/guide/index)
* [/guide/version/cloud/faq](/guide/version/cloud/faq)
* [/guide/version/cloud/intro](/guide/version/cloud/intro)
* [/guide/version/cloud/privacy](/guide/version/cloud/privacy)
* [/guide/version/cloud/terms](/guide/version/cloud/terms)
* [/guide/version/commercial](/guide/version/commercial)
* [/guide/version/opensource/intro](/guide/version/opensource/intro)
* [/guide/version/opensource/license](/guide/version/opensource/license)
* [/guide/workspace/customDomain](/guide/workspace/customDomain)
* [/guide/workspace/team/invitation\_link](/guide/workspace/team/invitation_link)
* [/guide/workspace/team/team\_roles\_permissions](/guide/workspace/team/team_roles_permissions)
* [/openapi/app](/openapi/app)
* [/openapi/chat](/openapi/chat)
* [/openapi/dataset](/openapi/dataset)
* [/openapi/index](/openapi/index)
* [/openapi/intro](/openapi/intro)
* [/plugin/index](/plugin/index)
* [/plugin/intro](/plugin/intro)
* [/plugin/model-presets](/plugin/model-presets)
* [/plugin/system-tool-development](/plugin/system-tool-development)
* [/self-host/config/env](/self-host/config/env)
* [/self-host/config/model/intro](/self-host/config/model/intro)
* [/self-host/config/model/minimax](/self-host/config/model/minimax)
* [/self-host/config/model/siliconCloud](/self-host/config/model/siliconCloud)
* [/self-host/config/object-storage](/self-host/config/object-storage)
* [/self-host/config/remote-debug-suite](/self-host/config/remote-debug-suite)
* [/self-host/config/sandbox/opensandbox](/self-host/config/sandbox/opensandbox)
* [/self-host/config/sandbox/sealosdevbox](/self-host/config/sandbox/sealosdevbox)
* [/self-host/config/signoz](/self-host/config/signoz)
* [/self-host/custom-models/bge-rerank](/self-host/custom-models/bge-rerank)
* [/self-host/custom-models/chatglm2](/self-host/custom-models/chatglm2)
* [/self-host/custom-models/chatglm2-m3e](/self-host/custom-models/chatglm2-m3e)
* [/self-host/custom-models/m3e](/self-host/custom-models/m3e)
* [/self-host/custom-models/marker](/self-host/custom-models/marker)
* [/self-host/custom-models/mineru](/self-host/custom-models/mineru)
* [/self-host/custom-models/ollama](/self-host/custom-models/ollama)
* [/self-host/custom-models/xinference](/self-host/custom-models/xinference)
* [/self-host/deploy/docker](/self-host/deploy/docker)
* [/self-host/deploy/sealos](/self-host/deploy/sealos)
* [/self-host/design/dataset](/self-host/design/dataset)
* [/self-host/dev](/self-host/dev)
* [/self-host/index](/self-host/index)
* [/self-host/migration/docker\_db](/self-host/migration/docker_db)
* [/self-host/migration/docker\_mongo](/self-host/migration/docker_mongo)
* [/self-host/troubleshooting/attention](/self-host/troubleshooting/attention)
* [/self-host/troubleshooting/faq](/self-host/troubleshooting/faq)
* [/self-host/troubleshooting/methods](/self-host/troubleshooting/methods)
* [/self-host/troubleshooting/model-errors](/self-host/troubleshooting/model-errors)
* [/self-host/troubleshooting/s3-issues](/self-host/troubleshooting/s3-issues)
* [/self-host/upgrading/4-12/4120](/self-host/upgrading/4-12/4120)
* [/self-host/upgrading/4-12/4121](/self-host/upgrading/4-12/4121)
* [/self-host/upgrading/4-12/4122](/self-host/upgrading/4-12/4122)
* [/self-host/upgrading/4-12/4123](/self-host/upgrading/4-12/4123)
* [/self-host/upgrading/4-12/4124](/self-host/upgrading/4-12/4124)
* [/self-host/upgrading/4-13/4130](/self-host/upgrading/4-13/4130)
* [/self-host/upgrading/4-13/4131](/self-host/upgrading/4-13/4131)
* [/self-host/upgrading/4-13/4132](/self-host/upgrading/4-13/4132)
* [/self-host/upgrading/4-14/4140](/self-host/upgrading/4-14/4140)
* [/self-host/upgrading/4-14/4141](/self-host/upgrading/4-14/4141)
* [/self-host/upgrading/4-14/41410](/self-host/upgrading/4-14/41410)
* [/self-host/upgrading/4-14/41411](/self-host/upgrading/4-14/41411)
* [/self-host/upgrading/4-14/41412](/self-host/upgrading/4-14/41412)
* [/self-host/upgrading/4-14/41413](/self-host/upgrading/4-14/41413)
* [/self-host/upgrading/4-14/41414](/self-host/upgrading/4-14/41414)
* [/self-host/upgrading/4-14/41415](/self-host/upgrading/4-14/41415)
* [/self-host/upgrading/4-14/41416](/self-host/upgrading/4-14/41416)
* [/self-host/upgrading/4-14/41417](/self-host/upgrading/4-14/41417)
* [/self-host/upgrading/4-14/41418](/self-host/upgrading/4-14/41418)
* [/self-host/upgrading/4-14/41419](/self-host/upgrading/4-14/41419)
* [/self-host/upgrading/4-14/4142](/self-host/upgrading/4-14/4142)
* [/self-host/upgrading/4-14/41420](/self-host/upgrading/4-14/41420)
* [/self-host/upgrading/4-14/41421](/self-host/upgrading/4-14/41421)
* [/self-host/upgrading/4-14/41422](/self-host/upgrading/4-14/41422)
* [/self-host/upgrading/4-14/41424](/self-host/upgrading/4-14/41424)
* [/self-host/upgrading/4-14/41425](/self-host/upgrading/4-14/41425)
* [/self-host/upgrading/4-14/41426](/self-host/upgrading/4-14/41426)
* [/self-host/upgrading/4-14/41427](/self-host/upgrading/4-14/41427)
* [/self-host/upgrading/4-14/41428](/self-host/upgrading/4-14/41428)
* [/self-host/upgrading/4-14/41429](/self-host/upgrading/4-14/41429)
* [/self-host/upgrading/4-14/4143](/self-host/upgrading/4-14/4143)
* [/self-host/upgrading/4-14/4144](/self-host/upgrading/4-14/4144)
* [/self-host/upgrading/4-14/4145](/self-host/upgrading/4-14/4145)
* [/self-host/upgrading/4-14/41451](/self-host/upgrading/4-14/41451)
* [/self-host/upgrading/4-14/4146](/self-host/upgrading/4-14/4146)
* [/self-host/upgrading/4-14/4147](/self-host/upgrading/4-14/4147)
* [/self-host/upgrading/4-14/4148](/self-host/upgrading/4-14/4148)
* [/self-host/upgrading/4-14/41481](/self-host/upgrading/4-14/41481)
* [/self-host/upgrading/4-14/4149](/self-host/upgrading/4-14/4149)
* [/self-host/upgrading/4-14/41930](/self-host/upgrading/4-14/41930)
* [/self-host/upgrading/4-15/41500](/self-host/upgrading/4-15/41500)
* [/self-host/upgrading/4-15/41501](/self-host/upgrading/4-15/41501)
* [/self-host/upgrading/4-15/41502](/self-host/upgrading/4-15/41502)
* [/self-host/upgrading/4-15/41503](/self-host/upgrading/4-15/41503)
* [/self-host/upgrading/4-15/41504](/self-host/upgrading/4-15/41504)
* [/self-host/upgrading/4-15/41505](/self-host/upgrading/4-15/41505)
* [/self-host/upgrading/4-15/41506](/self-host/upgrading/4-15/41506)
* [/self-host/upgrading/4-15/41507](/self-host/upgrading/4-15/41507)
* [/self-host/upgrading/4-15/4151](/self-host/upgrading/4-15/4151)
* [/self-host/upgrading/4-15/4152](/self-host/upgrading/4-15/4152)
* [/self-host/upgrading/4-15/4153](/self-host/upgrading/4-15/4153)
* [/self-host/upgrading/4-15/4154](/self-host/upgrading/4-15/4154)
* [/self-host/upgrading/4-15/4155](/self-host/upgrading/4-15/4155)
* [/self-host/upgrading/4-15/4156](/self-host/upgrading/4-15/4156)
* [/self-host/upgrading/4-15/4157](/self-host/upgrading/4-15/4157)
* [/self-host/upgrading/4-16/41601](/self-host/upgrading/4-16/41601)
* [/self-host/upgrading/4-16/41602](/self-host/upgrading/4-16/41602)
* [/self-host/upgrading/outdated/40](/self-host/upgrading/outdated/40)
* [/self-host/upgrading/outdated/41](/self-host/upgrading/outdated/41)
* [/self-host/upgrading/outdated/4100](/self-host/upgrading/outdated/4100)
* [/self-host/upgrading/outdated/4101](/self-host/upgrading/outdated/4101)
* [/self-host/upgrading/outdated/4110](/self-host/upgrading/outdated/4110)
* [/self-host/upgrading/outdated/4111](/self-host/upgrading/outdated/4111)
* [/self-host/upgrading/outdated/42](/self-host/upgrading/outdated/42)
* [/self-host/upgrading/outdated/421](/self-host/upgrading/outdated/421)
* [/self-host/upgrading/outdated/43](/self-host/upgrading/outdated/43)
* [/self-host/upgrading/outdated/44](/self-host/upgrading/outdated/44)
* [/self-host/upgrading/outdated/441](/self-host/upgrading/outdated/441)
* [/self-host/upgrading/outdated/442](/self-host/upgrading/outdated/442)
* [/self-host/upgrading/outdated/445](/self-host/upgrading/outdated/445)
* [/self-host/upgrading/outdated/446](/self-host/upgrading/outdated/446)
* [/self-host/upgrading/outdated/447](/self-host/upgrading/outdated/447)
* [/self-host/upgrading/outdated/45](/self-host/upgrading/outdated/45)
* [/self-host/upgrading/outdated/451](/self-host/upgrading/outdated/451)
* [/self-host/upgrading/outdated/452](/self-host/upgrading/outdated/452)
* [/self-host/upgrading/outdated/46](/self-host/upgrading/outdated/46)
* [/self-host/upgrading/outdated/461](/self-host/upgrading/outdated/461)
* [/self-host/upgrading/outdated/462](/self-host/upgrading/outdated/462)
* [/self-host/upgrading/outdated/463](/self-host/upgrading/outdated/463)
* [/self-host/upgrading/outdated/464](/self-host/upgrading/outdated/464)
* [/self-host/upgrading/outdated/465](/self-host/upgrading/outdated/465)
* [/self-host/upgrading/outdated/466](/self-host/upgrading/outdated/466)
* [/self-host/upgrading/outdated/467](/self-host/upgrading/outdated/467)
* [/self-host/upgrading/outdated/468](/self-host/upgrading/outdated/468)
* [/self-host/upgrading/outdated/469](/self-host/upgrading/outdated/469)
* [/self-host/upgrading/outdated/47](/self-host/upgrading/outdated/47)
* [/self-host/upgrading/outdated/471](/self-host/upgrading/outdated/471)
* [/self-host/upgrading/outdated/48](/self-host/upgrading/outdated/48)
* [/self-host/upgrading/outdated/481](/self-host/upgrading/outdated/481)
* [/self-host/upgrading/outdated/4810](/self-host/upgrading/outdated/4810)
* [/self-host/upgrading/outdated/4811](/self-host/upgrading/outdated/4811)
* [/self-host/upgrading/outdated/4812](/self-host/upgrading/outdated/4812)
* [/self-host/upgrading/outdated/4813](/self-host/upgrading/outdated/4813)
* [/self-host/upgrading/outdated/4814](/self-host/upgrading/outdated/4814)
* [/self-host/upgrading/outdated/4815](/self-host/upgrading/outdated/4815)
* [/self-host/upgrading/outdated/4816](/self-host/upgrading/outdated/4816)
* [/self-host/upgrading/outdated/4817](/self-host/upgrading/outdated/4817)
* [/self-host/upgrading/outdated/4818](/self-host/upgrading/outdated/4818)
* [/self-host/upgrading/outdated/4819](/self-host/upgrading/outdated/4819)
* [/self-host/upgrading/outdated/482](/self-host/upgrading/outdated/482)
* [/self-host/upgrading/outdated/4820](/self-host/upgrading/outdated/4820)
* [/self-host/upgrading/outdated/4821](/self-host/upgrading/outdated/4821)
* [/self-host/upgrading/outdated/4822](/self-host/upgrading/outdated/4822)
* [/self-host/upgrading/outdated/4823](/self-host/upgrading/outdated/4823)
* [/self-host/upgrading/outdated/483](/self-host/upgrading/outdated/483)
* [/self-host/upgrading/outdated/484](/self-host/upgrading/outdated/484)
* [/self-host/upgrading/outdated/485](/self-host/upgrading/outdated/485)
* [/self-host/upgrading/outdated/486](/self-host/upgrading/outdated/486)
* [/self-host/upgrading/outdated/487](/self-host/upgrading/outdated/487)
* [/self-host/upgrading/outdated/488](/self-host/upgrading/outdated/488)
* [/self-host/upgrading/outdated/489](/self-host/upgrading/outdated/489)
* [/self-host/upgrading/outdated/490](/self-host/upgrading/outdated/490)
* [/self-host/upgrading/outdated/491](/self-host/upgrading/outdated/491)
* [/self-host/upgrading/outdated/4910](/self-host/upgrading/outdated/4910)
* [/self-host/upgrading/outdated/4911](/self-host/upgrading/outdated/4911)
* [/self-host/upgrading/outdated/4912](/self-host/upgrading/outdated/4912)
* [/self-host/upgrading/outdated/4913](/self-host/upgrading/outdated/4913)
* [/self-host/upgrading/outdated/4914](/self-host/upgrading/outdated/4914)
* [/self-host/upgrading/outdated/492](/self-host/upgrading/outdated/492)
* [/self-host/upgrading/outdated/493](/self-host/upgrading/outdated/493)
* [/self-host/upgrading/outdated/494](/self-host/upgrading/outdated/494)
* [/self-host/upgrading/outdated/495](/self-host/upgrading/outdated/495)
* [/self-host/upgrading/outdated/496](/self-host/upgrading/outdated/496)
* [/self-host/upgrading/outdated/497](/self-host/upgrading/outdated/497)
* [/self-host/upgrading/outdated/498](/self-host/upgrading/outdated/498)
* [/self-host/upgrading/outdated/499](/self-host/upgrading/outdated/499)
* [/self-host/upgrading/upgrade-intruction](/self-host/upgrading/upgrade-intruction)
file: ./content/faq/chat.en.mdx
meta: {
"title": "Chat Interface",
"description": "Common FastGPT chat interface questions"
}
## I updated my app in the workspace, but the chat isn't reflecting the changes?
You need to publish the app first. Chat only picks up changes after publishing.
## Browser doesn't support voice input
1. Make sure microphone permissions are enabled in both your browser and OS settings.
2. Confirm the browser has permission to use the microphone for this site, and that the correct microphone source is selected.
3. The site must have an SSL certificate for microphone access to work.
file: ./content/faq/chat.mdx
meta: {
"title": "聊天框问题",
"description": "FastGPT 常见聊天框问题"
}
## 我修改了工作台的应用,为什么在“聊天”时没有更新配置?
应用需要点击发布后,聊天才会更新应用。
## 浏览器不支持语音输入
1. 首先需要确保浏览器、电脑本身麦克风权限的开启。
2. 确认浏览器允许该站点使用麦克风,并且选择正确的麦克风来源。
3. 需有 SSL 证书的站点才可以使用麦克风。
file: ./content/faq/index.en.mdx
meta: {
"title": "FAQ",
"description": "FastGPT frequently asked questions"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/faq/index.mdx
meta: {
"title": "使用案例",
"description": "FastGPT 使用案例"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/guide/index.en.mdx
meta: {
"title": "User Guide",
"description": "FastGPT User Guide"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/guide/index.mdx
meta: {
"title": "使用指南",
"description": "FastGPT 使用指南"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/plugin/index.en.mdx
meta: {
"title": "Plugin System",
"description": "FastGPT plugin system documentation"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/plugin/index.mdx
meta: {
"title": "插件系统",
"description": "FastGPT 插件系统文档"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/plugin/intro.en.mdx
meta: {
"title": "Plugin System Overview",
"description": "FastGPT plugin system overview"
}
> This document applies to FastGPT Plugin v1.0.0 and later.
## Background
FastGPT capabilities were previously maintained inside the FastGPT main service and organized as a Monorepo. System plugins also existed as a sub-repository under `FastGPT/packages/plugin`.
As the number of system tools and community contributions grew, the old structure exposed several problems:
1. System plugins had to be released together with the FastGPT main service, which slowed plugin iteration.
2. Community contributors needed to run the full FastGPT application and submit PRs directly to the main repository.
3. Custom plugins required maintaining a FastGPT fork and manually handling upgrades and merges.
4. The Next.js/webpack build model was not suitable for mounting new plugins at runtime.
System plugins have therefore been split into a standalone repository:
[FastGPT Plugin](https://github.com/labring/fastgpt-plugin)
FastGPT Plugin v1.0.0 systematically refactors the plugin project so plugin installation, version management, runtime isolation, and operations configuration share one model.
## Design Goals
The main goals of FastGPT Plugin are:
1. Decoupling and modularization: system tools, model presets, app templates, and future capabilities such as RAG algorithms, Agent strategies, and third-party integrations can evolve independently.
2. Unified plugin package protocol: `.pkg` files manage plugin installation, updates, and distribution, with extension points reserved for future plugin types.
3. Runtime isolation: plugin execution is managed by a unified runtime. Each plugin version has its own process pool, queue, and runtime configuration.
4. Lower development complexity: contributors can develop, debug, check, and package system tools independently through the CLI and SDK.
5. Plugin Marketplace: official and community plugins can be displayed and distributed through Marketplace.
## Core Concepts
| Name | Description |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------- |
| Plugin | An independent, reusable capability module. Plugins can have different types, such as tools, model presets, and dataset sources. |
| Plugin package | The packaged `.pkg` file for a plugin. All plugin types are installed, updated, and managed through plugin packages. |
| Tool | A plugin type that usually wraps third-party services, internal APIs, or local computation and can be called by workflows and Agents. |
| Tool suite | A plugin that exposes multiple related child tools while sharing plugin metadata and secret configuration. |
| Plugin Marketplace | A centralized platform where users can search, download, and install plugins. |
| Runtime | The backend implementation responsible for executing plugin code. The current default runtime is `local-pool`. |
| Pod | A single plugin child process in the local process pool. One plugin service can own multiple Pods. |
## Repository Structure
`fastgpt-plugin` is a pnpm workspace Monorepo designed with Clean Architecture and DDD as references.
```text
fastgpt-plugin/
├── apps/
│ ├── cli/ # CLI for plugin development, build, check, pack, and debug
│ ├── server/ # FastGPT Plugin HTTP service
│ └── debug-runtime-monitor/ # Local runtime monitoring and debugging panel
├── packages/
│ ├── domain/ # Domain entities, value objects, and port definitions
│ ├── usecase/ # Application use cases for plugins, tools, models, runtime, and more
│ ├── interface-adapter/ # HTTP contracts, DTOs, and auth adapters
│ ├── infrastructure/ # Hono, Mongo, S3, Redis, runtime, logging, metrics, and other implementations
│ └── shared/ # Cross-layer pure utilities
├── sdk/
│ ├── client/ # Client SDK for calling the FastGPT Plugin service
│ └── factory/ # Plugin author SDK
├── test/ # Cross-package test utilities and fixtures
└── docs/ # Project documentation
```
Core dependency direction:
* `domain` defines business concepts and ports. It is the innermost layer and does not depend on application entrypoints or infrastructure.
* `usecase` orchestrates business flows and depends on `domain` entities, value objects, and ports.
* `interface-adapter` defines HTTP contracts, DTOs, and auth inputs/outputs. It converts external protocols into structures the application can understand.
* `infrastructure` implements ports and runtime capabilities, including the HTTP framework, database, object storage, Redis, plugin runtime, logging, and metrics.
* `apps/*` are composition roots that assemble dependencies, register routes, start processes, or provide development commands.
* `sdk/*` is published for external users and provides service calls and plugin development capabilities.
For system tool development, see [System Tool Development Guide](./system-tool-development.en.mdx). For model presets, see [Add Model Presets](./model-presets.en.mdx).
## Repository Responsibilities
The FastGPT Plugin ecosystem mainly involves these repositories:
| Repository | Purpose |
| --------------------------- | ---------------------------------------------------------------------- |
| `labring/fastgpt-plugin` | Plugin service, SDK, CLI, debug monitor, and infrastructure code. |
| `fastgpt-official-plugins` | Plugins maintained or reviewed by FastGPT officials. |
| `fastgpt-community-plugins` | Community third-party plugins. |
| `fastgpt-business-plugins` | Private plugins, customer-customized plugins, and commercial delivery. |
The `fastgpt-plugin` repository only provides development, build, check, packaging, and server runtime capabilities. Specific plugin source code is usually placed in the official, community, or business plugin repositories.
## Marketplace And Usage Boundaries
FastGPT Marketplace is the plugin distribution channel for centrally displaying and distributing official and community plugins. Current boundaries:
* Marketplace is a SaaS distribution service and does not provide a private deployment version.
* Community plugins must first be submitted to the Community Plugins repository, pass basic review, and then enter Marketplace.
* The FastGPT cloud service does not yet support direct custom plugin uploads by users.
* Third-party custom plugins are currently mainly used through self-deployment or administrator upload in the business edition.
## Plugin Installation And Management
The FastGPT Plugin service is responsible for plugin package management, runtime registration, plugin call forwarding, and system-level configuration. The FastGPT main service invokes plugins through the plugin runtime interface, and the plugin service dispatches each call to the corresponding runtime.
System plugins can be installed in two main ways:
1. System-level installation: the root user uploads a `.pkg` file on the plugin management page or installs a plugin from Marketplace. The installed plugin is visible to the whole system.
2. Team-level installation: reserved for team administrators or members with plugin management permission. The plugin is visible only within that team.
After a plugin is installed, the service saves the plugin package file, parses plugin metadata, and registers the plugin with the runtime when it is enabled. System administrators can manage plugin status, system secrets, and runtime parameters.
Plugin statuses include:
* Normal: the plugin is available for normal use.
* Pending offline: existing workflows continue to run, but the plugin can no longer be added to new workflows.
* Offline: the plugin cannot be used.
System-level plugins can configure system secrets for other users in the system to reuse when invoking the plugin. Secrets are hosted by the plugin service. Callers reference them through plugin configuration and never access plaintext secrets directly.
## `.pkg` Plugin Package Protocol
New system tools no longer depend on the legacy built-in source directory `modules/tool/packages`; they are delivered through unified `.pkg` files.
Build artifacts usually include:
* `dist/index.js`
* `dist/manifest.json`
* icon files
* optional `README.md`
* optional `assets/**`
`.pkg` files are used for upload, installation, listing, and version management. Plugin metadata, input/output schemas, secret schemas, and icon assets are included in the build output for FastGPT pages, workflows, and Agents.
## local-pool Runtime
The current default runtime is the local process pool, `local-pool`. It manages Pods and request queues per plugin service.
After a plugin call enters a service, scheduling proceeds as follows:
1. Prefer an existing available Pod and dispatch the request immediately.
2. If no Pod is available and `pods + pendingPods < maxPods`, create a new Pod first and dispatch the current request after startup succeeds.
3. If `maxPods` has been reached, startup backoff is active, or a Pod cannot be created temporarily, the request enters a bounded queue.
4. When a Pod is released, startup succeeds, configuration is updated, or a crash is recovered, the queue continues to drain.
5. When queue length reaches `maxQueueSize`, new requests are rejected. Requests also fail after waiting longer than `queueTimeout`.
Each tool plugin can configure four runtime parameters:
| Parameter | Default | Description |
| ------------------------------------ | ---------- | -------------------------------------------------------------------------------- |
| Minimum worker nodes | `0` | Values above `0` warm up Pods and try to keep at least this many Pods available. |
| Maximum worker nodes | `5` | The service can scale out to this limit when no Pod is available. |
| Node timeout | `120000ms` | Timeout for one plugin call inside a Pod. |
| Maximum concurrent requests per node | `10` | Maximum concurrent requests one Pod can process. |
Environment variables provide default runtime parameters and global limits:
| Environment variable | Description |
| ---------------------------------------------- | --------------------------------------------------------------------------------- |
| `POOL_HEALTH_CHECK_INTERVAL` | Health check interval in milliseconds. |
| `POOL_MAX_TOTAL_PODS` | Total limit for all plugin Pods in the current server process. |
| `POOL_SERVICE_MIN_PODS` | Default minimum worker nodes for one plugin. |
| `POOL_SERVICE_MAX_PODS` | Default maximum worker nodes for one plugin. |
| `POOL_SERVICE_IDLE_TIMEOUT` | Pod idle recycle time in milliseconds. |
| `POOL_SERVICE_POD_TIMEOUT` | Execution timeout for one plugin call in milliseconds. |
| `POOL_SERVICE_MAX_CONCURRENT_REQUESTS_PER_POD` | Default maximum concurrent requests for one Pod. |
| `POOL_SERVICE_MAX_REQUESTS_PER_POD` | Maximum requests one Pod can process before replacement. |
| `POOL_SERVICE_MAX_QUEUE_SIZE` | Maximum request queue capacity for one plugin service. |
| `POOL_SERVICE_QUEUE_TIMEOUT` | Maximum time a request can wait in queue for an available Pod, in milliseconds. |
| `POOL_SERVICE_STARTUP_RETRY_BASE_DELAY` | Base delay for exponential backoff after Pod startup timeout, in milliseconds. |
| `POOL_SERVICE_STARTUP_RETRY_MAX_DELAY` | Maximum delay for exponential backoff after Pod startup timeout, in milliseconds. |
Pod startup errors are recorded and classified. Consecutive non-timeout startup failures trigger startup circuit breaking after the threshold is reached, preventing more Pods from being created. Startup timeouts are usually treated as resource pressure, enter exponential backoff, and retry later.
## Development And Distribution
System tool plugins are developed with `@fastgpt-plugin/cli` and `@fastgpt-plugin/sdk-factory`.
Developers use the CLI to create single-tool or tool-suite skeletons, and use the SDK to declare `manifest`, `inputSchema`, `outputSchema`, `secretSchema`, and handler logic. After development, run tests, build, check, and pack to generate a `.pkg` file.
Continue with [System Tool Development Guide](./system-tool-development.en.mdx) to develop system tools.
## References
* [FastGPT Plugin](https://github.com/labring/fastgpt-plugin)
* [FastGPT Plugin System Design](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/design.md)
* [FastGPT Plugin Architecture](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/architecture.md)
* [System Plugin Development Guide](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/how-to-devlop-plugin.en.md)
file: ./content/plugin/intro.mdx
meta: {
"title": "插件系统说明",
"description": "FastGPT 插件系统说明"
}
> 本文档适用于 FastGPT Plugin v1.0.0 及以上版本的插件系统。
## 背景
原先 FastGPT 的各项能力均在 FastGPT 主服务内维护,并通过 Monorepo 方式组织。系统插件也曾作为一个子仓库存在于 `FastGPT/packages/plugin` 下。
随着系统工具数量和社区贡献增加,旧结构暴露出几个问题:
1. 系统插件必须伴随 FastGPT 主服务一起发版,限制了插件迭代速度。
2. 社区贡献插件需要运行完整 FastGPT 应用,并直接向主仓库提交 PR。
3. 使用自定义插件需要维护 FastGPT fork,手动处理升级和合并。
4. Next.js/webpack 构建模型不适合在运行时挂载新插件。
因此,系统插件被拆分到独立仓库:
[FastGPT Plugin](https://github.com/labring/fastgpt-plugin)
FastGPT Plugin v1.0.0 对插件项目进行了系统性重构,目标是让插件的安装、版本管理、运行隔离和运维配置形成统一模型。
## 设计目标
FastGPT Plugin 的核心目标:
1. 解耦和模块化:系统工具、模型预设、App 模板等能力可以独立迭代,后续也能扩展 RAG 算法、Agent 策略和第三方接入。
2. 插件包统一协议:使用 `.pkg` 文件管理插件安装、更新和分发,为后续插件类型预留扩展空间。
3. 运行隔离:通过运行时统一管理插件执行,每个插件版本拥有独立进程池、队列和运行配置。
4. 降低开发复杂度:贡献系统工具时可以通过 CLI 和 SDK 独立开发、调试、检查和打包。
5. 插件市场:通过 Marketplace 集中展示和分发官方及社区插件。
## 核心概念
| 名称 | 说明 |
| ---- | ------------------------------------------- |
| 插件 | 独立、可复用的功能模块,可以有不同类型,例如工具、模型预设、知识库来源等。 |
| 插件包 | 插件打包后的 `.pkg` 文件。不同类型插件都通过插件包完成安装、更新和管理。 |
| 工具 | 一类插件,通常封装第三方服务、内部接口或本地计算逻辑,可被工作流和 Agent 调用。 |
| 工具集 | 一个插件暴露多个相关子工具,共享插件元信息和密钥配置。 |
| 插件市场 | 集中管理插件的平台,用户可以在其中搜索、下载和安装插件。 |
| 运行时 | 负责执行插件代码的后端实现,当前默认运行时是 `local-pool`。 |
| Pod | 本地进程池中的单个插件子进程。一个插件 service 可以拥有多个 Pod。 |
## 仓库结构
`fastgpt-plugin` 使用 pnpm workspace 组织 Monorepo,参考 Clean Architecture 和 DDD 分层设计。
```text
fastgpt-plugin/
├── apps/
│ ├── cli/ # 插件开发、构建、检查、打包、调试命令行
│ ├── server/ # FastGPT Plugin HTTP 服务
│ └── debug-runtime-monitor/ # 本地运行时监控调试面板
├── packages/
│ ├── domain/ # 领域实体、值对象、端口定义
│ ├── usecase/ # 插件、工具、模型、runtime 等应用用例
│ ├── interface-adapter/ # HTTP contract、DTO、鉴权适配
│ ├── infrastructure/ # Hono、Mongo、S3、Redis、运行时、日志、指标等实现
│ └── shared/ # 跨层复用的纯工具函数
├── sdk/
│ ├── client/ # 调用 FastGPT Plugin 服务的客户端 SDK
│ └── factory/ # 插件作者侧 SDK
├── test/ # 跨包测试工具与 fixtures
└── docs/ # 项目文档
```
核心依赖方向:
* `domain` 定义业务概念和端口,是最内层,不依赖应用入口和基础设施。
* `usecase` 负责编排业务流程,依赖 `domain` 的实体、值对象和端口。
* `interface-adapter` 定义 HTTP 合约、DTO、鉴权输入输出,负责把外部协议转换为应用可理解的数据结构。
* `infrastructure` 实现端口和运行环境能力,包括 HTTP 框架、数据库、对象存储、Redis、插件运行时、日志与指标。
* `apps/*` 是组合根,负责装配依赖、注册路由、启动进程或提供开发命令。
* `sdk/*` 面向外部使用者发布,提供服务调用和插件开发能力。
系统工具开发结构可以参考 [系统工具开发指南](./system-tool-development.mdx)。模型预设维护可以参考 [增加模型预设](./model-presets.mdx)。
## 仓库分工
FastGPT Plugin 生态主要涉及以下仓库:
| 仓库 | 作用 |
| --------------------------- | -------------------------- |
| `labring/fastgpt-plugin` | 插件服务、SDK、CLI、调试监视器和基础设施代码。 |
| `fastgpt-official-plugins` | 官方维护或审核通过的插件。 |
| `fastgpt-community-plugins` | 社区第三方插件。 |
| `fastgpt-business-plugins` | 私有插件、客户定制插件和商业交付插件。 |
`fastgpt-plugin` 仓库只提供开发、构建、检查、打包和服务端运行能力。具体插件源码通常放在 official、community 或 business 插件仓库中。
## 插件市场与使用边界
FastGPT Marketplace 是插件分发渠道,用于集中展示和分发官方及社区插件。当前边界如下:
* Marketplace 是 SaaS 分发服务,不提供私有化部署版本。
* 社区插件需要先提交到 Community Plugins 仓库,经基础审核后再进入 Marketplace。
* 云服务版本 FastGPT 暂未支持用户直接上传自定义插件。
* 第三方自定义插件目前主要通过自部署或商业版的管理员上传方式使用。
## 插件安装与管理
FastGPT Plugin 服务负责插件包管理、运行时注册、插件调用转发和系统级配置管理。FastGPT 主服务通过插件运行时接口调用插件,插件服务负责把调用分发到对应运行时。
系统插件安装主要有两种方式:
1. 系统级安装:root 用户在插件管理页面上传 `.pkg` 文件,或从插件市场安装。安装后全系统可见。
2. 团队级安装:预留给团队管理员或有插件管理权限的成员,仅团队内可见。
插件安装后会保存插件包文件、解析插件元信息,并在插件启用时注册到运行时。系统管理员可以管理插件状态、系统密钥和运行时参数。
插件状态包括:
* 正常:插件正常使用。
* 即将下线:不影响已有工作流运行,但无法再被新增到工作流中。
* 已下线:插件无法正常使用。
系统级插件可以配置“系统密钥”,供系统内其他用户在调用插件时复用。密钥由插件服务托管,调用方通过插件配置引用,不直接接触明文密钥。
## `.pkg` 插件包协议
新版系统工具不再依赖旧的 `modules/tool/packages` 内置源码目录,而是使用统一 `.pkg` 文件交付。
构建产物通常包含:
* `dist/index.js`
* `dist/manifest.json`
* 图标文件
* 可选的 `README.md`
* 可选的 `assets/**`
`.pkg` 文件用于上传、安装、上架和版本管理。插件元信息、输入输出 schema、密钥 schema 和图标资源都会进入构建产物,供 FastGPT 页面、工作流和 Agent 调用使用。
## local-pool 运行时
当前默认运行时是本地进程池,即 `local-pool`。它按单插件 service 维度管理 Pod 和请求队列。
一次插件调用进入 service 后,调度顺序如下:
1. 优先选择已有可用 Pod,立即派发请求。
2. 没有可用 Pod 且 `pods + pendingPods < maxPods` 时,先创建新 Pod,启动成功后派发当前请求。
3. 达到 `maxPods`、处于启动退避期或暂时无法创建 Pod 时,请求进入有界队列等待。
4. Pod 释放、创建成功、配置更新或崩溃恢复时,队列继续被消费。
5. 队列长度达到 `maxQueueSize` 后,新请求会被拒绝;请求等待超过 `queueTimeout` 后会超时失败。
每个工具插件可以单独配置 4 个运行参数:
| 参数 | 默认值 | 说明 |
| -------- | ---------- | --------------------------------- |
| 最小工作节点数 | `0` | 大于 `0` 时会预热 Pod,并尽量维持不少于该数量的 Pod。 |
| 最大工作节点数 | `5` | 没有可用 Pod 时可扩容到该上限。 |
| 节点超时时间 | `120000ms` | 单次插件调用在 Pod 内执行的超时时间。 |
| 每节点最大并发数 | `10` | 单个 Pod 同时处理的最大并发请求数。 |
环境变量提供默认运行参数和全局限制:
| 环境变量 | 说明 |
| ---------------------------------------------- | --------------------------- |
| `POOL_HEALTH_CHECK_INTERVAL` | 健康检查间隔,单位毫秒。 |
| `POOL_MAX_TOTAL_PODS` | 当前 server 进程内所有插件 Pod 的总上限。 |
| `POOL_SERVICE_MIN_PODS` | 单插件默认最小工作节点数。 |
| `POOL_SERVICE_MAX_PODS` | 单插件默认最大工作节点数。 |
| `POOL_SERVICE_IDLE_TIMEOUT` | Pod 空闲回收时间,单位毫秒。 |
| `POOL_SERVICE_POD_TIMEOUT` | 单次插件调用执行超时时间,单位毫秒。 |
| `POOL_SERVICE_MAX_CONCURRENT_REQUESTS_PER_POD` | 单个 Pod 默认最大并发请求数。 |
| `POOL_SERVICE_MAX_REQUESTS_PER_POD` | 单个 Pod 最大处理请求数;超过后自动替换。 |
| `POOL_SERVICE_MAX_QUEUE_SIZE` | 单插件 service 请求队列最大容量。 |
| `POOL_SERVICE_QUEUE_TIMEOUT` | 请求在队列中等待可用 Pod 的最长时间,单位毫秒。 |
| `POOL_SERVICE_STARTUP_RETRY_BASE_DELAY` | Pod 启动超时后的指数退避基础延迟,单位毫秒。 |
| `POOL_SERVICE_STARTUP_RETRY_MAX_DELAY` | Pod 启动超时后的指数退避最大延迟,单位毫秒。 |
Pod 启动错误会被记录并分类。连续非超时启动失败达到阈值后会触发启动熔断,阻止继续创建 Pod;启动超时通常按资源繁忙处理,会进入指数退避后重试。
## 开发与分发
系统工具插件使用 `@fastgpt-plugin/cli` 和 `@fastgpt-plugin/sdk-factory` 开发。
开发者通过 CLI 创建单工具或工具集骨架,使用 SDK 声明 `manifest`、`inputSchema`、`outputSchema`、`secretSchema` 和 handler。插件开发完成后运行测试、构建、检查和打包,最终生成 `.pkg` 文件。
开发系统工具可以继续阅读 [系统工具开发指南](./system-tool-development.mdx)。
## 参考
* [FastGPT Plugin](https://github.com/labring/fastgpt-plugin)
* [FastGPT 插件系统设计文档](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/design.zh.md)
* [FastGPT Plugin 架构文档](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/architecture.zh.md)
* [系统插件开发指南](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/how-to-devlop-plugin.md)
file: ./content/plugin/model-presets.en.mdx
meta: {
"title": "Add Model Presets",
"description": "Model preset notes for the FastGPT plugin system"
}
Model presets are maintained in the `fastgpt-plugin` repository. They provide FastGPT with built-in model providers, model lists, model capabilities, and default request parameters. After FastGPT loads these static presets, users can select the corresponding models in model configuration, AIProxy channels, and plugin-related features.
This page follows the plugin system code structure for version 1.0 and later. The old `modules/model/*` paths are no longer the primary maintenance entry point.
## Related Directories
```text
packages/infrastructure/src/static-data/models/
├── index.ts
├── model.ts
├── type.ts
├── channel-avatar/
└── provider/
└── {Provider}/
├── index.ts
└── logo.svg
```
* `provider/{Provider}/index.ts`: model presets for one provider.
* `index.ts`: registers all providers and generates `staticModelList` and provider lists.
* `model.ts`: maintains provider display names in `ModelProviderMap` and AIProxy channels in `aiproxyChannels`.
* `type.ts`: defines schemas for provider configs and model presets.
* `provider/{Provider}/logo.svg`: provider logo.
* `channel-avatar/`: AIProxy channel avatars.
## Add a Model to an Existing Provider
### 1. Confirm the provider is registered
First check `packages/infrastructure/src/static-data/models/index.ts` and make sure the provider is imported and included in `staticModelProviderConfigs`:
```ts
import openai from './provider/OpenAI';
export const staticModelProviderConfigs = [openai];
```
If you are only adding a model to an existing provider, you do not need to change `index.ts`.
### 2. Update the provider model list
Open the provider file, for example:
```text
packages/infrastructure/src/static-data/models/provider/OpenAI/index.ts
```
Add the model to the `list` array. Prefer cloning the closest model from the same provider, type, and family, then adjust the fields based on official documentation.
Examples for all five model types:
```ts
import { ModelTypeEnum, type ProviderConfigType } from '../../type';
const ttsVoices = [
{
label: 'Default voice',
value: 'default'
}
];
const models: ProviderConfigType = {
provider: 'ExampleProvider',
list: [
{
type: ModelTypeEnum.llm,
model: 'example-chat',
maxContext: 128000,
maxTokens: 16384,
quoteMaxToken: 120000,
maxTemperature: 1,
responseFormatList: ['text', 'json_schema'],
vision: true,
reasoning: false,
reasoningEffort: false,
toolChoice: true
},
{
type: ModelTypeEnum.embedding,
model: 'example-embedding',
defaultToken: 512,
maxToken: 8192,
normalization: true
},
{
type: ModelTypeEnum.rerank,
model: 'example-rerank',
maxToken: 8192
},
{
type: ModelTypeEnum.tts,
model: 'example-tts',
voices: ttsVoices
},
{
type: ModelTypeEnum.stt,
model: 'example-stt'
}
]
};
export default models;
```
Common fields:
| Field | Description |
| -------------------- | ------------------------------------------------------------------------------ |
| `type` | Model type from `ModelTypeEnum`: `llm`, `embedding`, `rerank`, `tts`, or `stt` |
| `model` | Actual model ID used in requests |
| `name` | Optional display name; defaults to `model` when omitted |
| `maxContext` | Maximum LLM context length |
| `maxTokens` | Maximum LLM output length |
| `quoteMaxToken` | Maximum token budget FastGPT can use for quoted Knowledge Base content |
| `maxTemperature` | Maximum temperature; use `null` when the model does not support temperature |
| `responseFormatList` | Supported response formats, such as `text`, `json_object`, and `json_schema` |
| `vision` | Whether vision input is supported |
| `reasoning` | Whether this is a reasoning model |
| `reasoningEffort` | Whether reasoning effort can be configured |
| `toolChoice` | Whether tool choice is supported |
| `fieldMap` | Field-name mapping for non-standard OpenAI-compatible APIs |
| `defaultConfig` | Default request parameters sent with the model request |
| `defaultToken` | Default chunk token count for Embedding models |
| `maxToken` | Maximum input token count for Embedding/Rerank models |
| `normalization` | Whether Embedding vectors should be normalized |
| `voices` | Available voice list for TTS models |
`index.ts` automatically adds the following when building `staticModelList`:
* `provider`: from the current provider config.
* `name`: defaults to `model` when not explicitly set.
* Some default LLM capability switches, such as Knowledge Base processing, classification, extraction, tool calling, and evaluation.
### 3. Do not rely on model names alone
Before adding or changing a model, verify it against official model docs, official model-list APIs, or official pricing/model pages. Do not rely on search results, third-party blogs, or aggregator pages as proof that a model exists.
Recommended rules:
* Model presets support five model types: `llm`, `embedding`, `rerank`, `tts`, and `stt`. Choose the type based on the model's real capabilities and fill in the fields required by that type's schema.
* Do not remove preview, experimental, or dated models just because a stable-looking sibling exists. Remove them only when official docs mark them as deprecated, retired, unavailable, or no longer recommended.
* For open catalogs such as OpenRouter, Ollama, HuggingFace, and Other, avoid deleting local placeholders or models users may customize.
* Preserve the existing ordering style in each provider file. Newer or more capable models are usually placed first.
## Add a New Model Provider
Only add a provider directory when you need to support a completely new provider.
### 1. Create the provider directory
Create a directory under `provider/` using the provider identifier:
```text
packages/infrastructure/src/static-data/models/provider/NewProvider/
├── index.ts
└── logo.svg
```
`logo.svg` is the model provider avatar. When the plugin service initializes static model assets, it uploads `provider/{Provider}/logo.svg` as `models/{Provider}/logo`, and the `/models/get-providers` API returns that URL as the provider `avatar`.
Basic `index.ts` structure:
```ts
import { ModelTypeEnum, type ProviderConfigType } from '../../type';
const models: ProviderConfigType = {
provider: 'NewProvider',
list: [
{
type: ModelTypeEnum.llm,
model: 'new-provider-chat',
maxContext: 128000,
maxTokens: 8192,
quoteMaxToken: 120000,
maxTemperature: 1,
responseFormatList: ['text'],
vision: false,
reasoning: false,
reasoningEffort: false,
toolChoice: true
}
]
};
export default models;
```
### 2. Register the provider
Import it in `packages/infrastructure/src/static-data/models/index.ts` and add it to `staticModelProviderConfigs`:
```ts
import newProvider from './provider/NewProvider';
export const staticModelProviderConfigs: ProviderConfigType[] = [newProvider];
```
### 3. Add provider display names
Add multilingual display names to `ModelProviderMap` in `packages/infrastructure/src/static-data/models/model.ts`:
```ts
NewProvider: {
en: 'NewProvider',
'zh-CN': 'New Provider',
'zh-Hant': 'New Provider'
}
```
If you do not add the provider to `ModelProviderMap`, the system falls back to the raw `provider` string as its display name. Formal providers should include multilingual display names.
## Add an AIProxy Protocol
Adding an AIProxy protocol is not the same as adding a model provider:
* Model provider: decides which `provider` owns the model presets, maintains the model list and model capabilities, and uses `provider/{Provider}/logo.svg` as its avatar.
* AIProxy protocol: decides whether the protocol appears in the AIProxy channel list. AIProxy routes requests to the corresponding adaptor by `channelId`, and FastGPT uses `channel-avatar/{avatar}.svg` as the channel avatar.
If you are only adding model presets, you may not need to add an AIProxy protocol. Maintain `aiproxyChannels` only when FastGPT needs to display that protocol in the AIProxy channel list.
### 1. Check AIProxy protocols and get channelId
The `channelId` must match the `ChannelType` value defined in [`core/model/chtype.go`](https://github.com/labring/aiproxy/blob/main/core/model/chtype.go) in the AIProxy repository. Do not guess the `channelId` from the provider name.
Run this in the AIProxy repository:
```bash
rg -n "ChannelType.*=" core/model/chtype.go
```
Examples:
| AIProxy type | ID | FastGPT `channelId` |
| ------------------------- | ---- | ------------------- |
| `ChannelTypeOpenAI` | `1` | `1` |
| `ChannelTypeAnthropic` | `14` | `14` |
| `ChannelTypeAli` | `17` | `17` |
| `ChannelTypeGoogleGemini` | `24` | `24` |
| `ChannelTypeDeepseek` | `36` | `36` |
| `ChannelTypeDoubao` | `40` | `40` |
| `ChannelTypeSiliconflow` | `43` | `43` |
| `ChannelTypeAntLing` | `54` | `54` |
Use the current `core/model/chtype.go` file on the AIProxy main branch as the source of truth.
### 2. Add the protocol declaration in fastgpt-plugin
After confirming AIProxy supports the protocol, add an entry to `aiproxyChannels` in `packages/infrastructure/src/static-data/models/model.ts`:
```ts
export const aiproxyChannels: AIProxyChannelsType = [
{
channelId: 54,
name: {
en: 'Ant Ling',
'zh-CN': '蚂蚁百灵',
'zh-Hant': '螞蟻百靈'
},
avatar: 'antling'
}
];
```
Field reference:
| Field | Description |
| ----------- | ---------------------------------------------------------------------- |
| `channelId` | Numeric AIProxy `ChannelType` ID. It must match `core/model/chtype.go` |
| `name` | Multilingual display name in the FastGPT channel list |
| `avatar` | Channel avatar filename without the extension |
Also add the avatar file under `channel-avatar/`:
```text
packages/infrastructure/src/static-data/models/channel-avatar/antling.svg
```
The `avatar` value must match the filename under `channel-avatar/`. Supported avatar extensions are `svg`, `png`, `jpeg`, `webp`, and `jpg`.
If AIProxy does not support the protocol yet, add the `ChannelType` and adaptor in the AIProxy repository first, and confirm the adaptor is imported in [`core/relay/adaptors/register.go`](https://github.com/labring/aiproxy/blob/main/core/relay/adaptors/register.go). The FastGPT plugin side only declares channel display data; it does not implement AIProxy adaptor logic.
## Validation
After updating presets, run at least:
```bash
pnpm typecheck
```
If you changed many providers, model schemas, or static asset loading logic, also run:
```bash
pnpm test
```
Before submitting, review the diff under `packages/infrastructure/src/static-data/models/` and make sure no unrelated provider models were removed, model types are correct, and the new provider logo or `channel-avatar` file is included.
file: ./content/plugin/model-presets.mdx
meta: {
"title": "增加模型预设",
"description": "FastGPT 插件系统中的模型预设说明"
}
模型预设维护在 `fastgpt-plugin` 仓库中,用于向 FastGPT 提供内置模型供应商、模型列表、模型能力和默认参数。FastGPT 读取这些静态预设后,用户才能在模型配置、AIProxy 渠道和相关插件能力中选择对应模型。
本文基于 1.0 版本以上的插件系统代码结构,旧版 `modules/model/*` 路径已经不再作为主要维护入口。
## 相关目录
```text
packages/infrastructure/src/static-data/models/
├── index.ts
├── model.ts
├── type.ts
├── channel-avatar/
└── provider/
└── {Provider}/
├── index.ts
└── logo.svg
```
* `provider/{Provider}/index.ts`:单个模型供应商的模型预设列表。
* `index.ts`:注册所有供应商,生成 `staticModelList` 和供应商列表。
* `model.ts`:维护供应商显示名 `ModelProviderMap` 和 AIProxy 渠道 `aiproxyChannels`。
* `type.ts`:定义供应商配置和模型预设的输入 schema。
* `provider/{Provider}/logo.svg`:模型供应商 Logo。
* `channel-avatar/`:AIProxy 渠道头像。
## 给已有供应商增加模型
### 1. 确认供应商已经注册
先在 `packages/infrastructure/src/static-data/models/index.ts` 中确认供应商已经被引入,并存在于 `staticModelProviderConfigs`:
```ts
import openai from './provider/OpenAI';
export const staticModelProviderConfigs = [openai];
```
如果只是给已有供应商增加模型,不需要修改 `index.ts`。
### 2. 修改供应商模型列表
进入对应供应商目录,例如:
```text
packages/infrastructure/src/static-data/models/provider/OpenAI/index.ts
```
在 `list` 数组中增加模型。优先复制同供应商、同类型、同模型家族中最接近的一项,再根据官方文档调整字段。
五类模型示例:
```ts
import { ModelTypeEnum, type ProviderConfigType } from '../../type';
const ttsVoices = [
{
label: '默认音色',
value: 'default'
}
];
const models: ProviderConfigType = {
provider: 'ExampleProvider',
list: [
{
type: ModelTypeEnum.llm,
model: 'example-chat',
maxContext: 128000,
maxTokens: 16384,
quoteMaxToken: 120000,
maxTemperature: 1,
responseFormatList: ['text', 'json_schema'],
vision: true,
reasoning: false,
reasoningEffort: false,
toolChoice: true
},
{
type: ModelTypeEnum.embedding,
model: 'example-embedding',
defaultToken: 512,
maxToken: 8192,
normalization: true
},
{
type: ModelTypeEnum.rerank,
model: 'example-rerank',
maxToken: 8192
},
{
type: ModelTypeEnum.tts,
model: 'example-tts',
voices: ttsVoices
},
{
type: ModelTypeEnum.stt,
model: 'example-stt'
}
]
};
export default models;
```
常用字段说明:
| 字段 | 说明 |
| -------------------- | ----------------------------------------------------------------- |
| `type` | 模型类型,来自 `ModelTypeEnum`,可选 `llm`、`embedding`、`rerank`、`tts`、`stt` |
| `model` | 真实请求时使用的模型 ID |
| `name` | 可选显示名,不填时默认使用 `model` |
| `maxContext` | LLM 最大上下文长度 |
| `maxTokens` | LLM 最大输出长度 |
| `quoteMaxToken` | FastGPT 引用知识库内容时可使用的最大 token |
| `maxTemperature` | 最大温度;不支持温度时填 `null` |
| `responseFormatList` | 支持的返回格式,如 `text`、`json_object`、`json_schema` |
| `vision` | 是否支持视觉输入 |
| `reasoning` | 是否为推理模型 |
| `reasoningEffort` | 是否支持推理强度配置 |
| `toolChoice` | 是否支持工具调用选择 |
| `fieldMap` | 字段名映射,用于适配非标准 OpenAI 兼容接口 |
| `defaultConfig` | 请求默认参数,会随模型请求一起发送 |
| `defaultToken` | Embedding 默认分段 token 数 |
| `maxToken` | Embedding/Rerank 最大输入 token 数 |
| `normalization` | Embedding 是否做归一化处理 |
| `voices` | TTS 可选音色列表 |
`index.ts` 会在生成 `staticModelList` 时自动补充:
* `provider`:来自当前供应商配置的 `provider`。
* `name`:未显式填写时使用 `model`。
* LLM 的部分默认能力开关,例如知识库处理、分类、内容提取、工具调用和评测。
### 3. 不要只看模型名称
新增或修改模型前,需要以官方模型文档、官方模型列表 API 或官方价格/模型页为依据。不要只根据搜索结果、第三方博客或聚合站判断模型是否存在。
维护时建议遵守以下规则:
* 模型预设支持 `llm`、`embedding`、`rerank`、`tts`、`stt` 五类模型。按模型真实能力选择对应类型,并补齐该类型 schema 要求的字段。
* 不要仅因为存在稳定版名称就删除 preview、experimental 或 dated 模型;只有官方明确废弃、下线或不再推荐时再移除。
* 对 OpenRouter、Ollama、HuggingFace、Other 这类开放目录,避免删除本地占位或用户可能自定义的模型。
* 保持文件内原有排序风格,通常把更新或能力更强的模型放在前面。
## 新增模型供应商
只有在需要接入全新的模型供应商时才新增供应商目录。
### 1. 创建供应商目录
在 `provider/` 下新增目录,目录名使用供应商标识:
```text
packages/infrastructure/src/static-data/models/provider/NewProvider/
├── index.ts
└── logo.svg
```
`logo.svg` 是模型供应商头像。插件服务初始化静态模型资源时,会把 `provider/{Provider}/logo.svg` 上传为 `models/{Provider}/logo`,`/models/get-providers` 接口会把它作为该模型供应商的 `avatar` 返回。
`index.ts` 基本结构:
```ts
import { ModelTypeEnum, type ProviderConfigType } from '../../type';
const models: ProviderConfigType = {
provider: 'NewProvider',
list: [
{
type: ModelTypeEnum.llm,
model: 'new-provider-chat',
maxContext: 128000,
maxTokens: 8192,
quoteMaxToken: 120000,
maxTemperature: 1,
responseFormatList: ['text'],
vision: false,
reasoning: false,
reasoningEffort: false,
toolChoice: true
}
]
};
export default models;
```
### 2. 注册供应商
在 `packages/infrastructure/src/static-data/models/index.ts` 中引入并加入 `staticModelProviderConfigs`:
```ts
import newProvider from './provider/NewProvider';
export const staticModelProviderConfigs: ProviderConfigType[] = [newProvider];
```
### 3. 增加供应商显示名
在 `packages/infrastructure/src/static-data/models/model.ts` 的 `ModelProviderMap` 中增加多语言显示名:
```ts
NewProvider: {
en: 'NewProvider',
'zh-CN': '新供应商',
'zh-Hant': '新供應商'
}
```
如果不增加 `ModelProviderMap`,系统会使用 `provider` 字符串作为兜底显示名,但正式供应商应补齐多语言显示名。
## 增加 AIProxy 协议
增加 AIProxy 协议不等于增加模型供应商:
* 模型供应商:决定模型预设属于哪个 `provider`,维护模型列表和模型能力,使用 `provider/{Provider}/logo.svg` 作为头像。
* AIProxy 协议:决定 AIProxy 渠道列表中是否出现该协议,最终由 AIProxy 根据 `channelId` 路由到对应 adaptor,使用 `channel-avatar/{avatar}.svg` 作为头像。
如果只是新增模型预设,不一定要增加 AIProxy 协议。只有当 FastGPT 需要在 AIProxy 渠道列表中展示该协议时,才需要维护 `aiproxyChannels`。
### 1. 查看 AIProxy 支持的协议并获取 channelId
`channelId` 必须和 AIProxy 仓库中 [`core/model/chtype.go`](https://github.com/labring/aiproxy/blob/main/core/model/chtype.go) 定义的 `ChannelType` 数值一致。不要根据供应商名称猜测 `channelId`。
在 AIProxy 仓库中执行:
```bash
rg -n "ChannelType.*=" core/model/chtype.go
```
例如:
| AIProxy 类型 | ID | FastGPT `channelId` |
| ------------------------- | ---- | ------------------- |
| `ChannelTypeOpenAI` | `1` | `1` |
| `ChannelTypeAnthropic` | `14` | `14` |
| `ChannelTypeAli` | `17` | `17` |
| `ChannelTypeGoogleGemini` | `24` | `24` |
| `ChannelTypeDeepseek` | `36` | `36` |
| `ChannelTypeDoubao` | `40` | `40` |
| `ChannelTypeSiliconflow` | `43` | `43` |
| `ChannelTypeAntLing` | `54` | `54` |
完整列表以 AIProxy 主分支的 `core/model/chtype.go` 为准。
### 2. 在 fastgpt-plugin 中增加协议声明
确认 AIProxy 已经支持该协议后,在 `packages/infrastructure/src/static-data/models/model.ts` 的 `aiproxyChannels` 中增加声明:
```ts
export const aiproxyChannels: AIProxyChannelsType = [
{
channelId: 54,
name: {
en: 'Ant Ling',
'zh-CN': '蚂蚁百灵',
'zh-Hant': '螞蟻百靈'
},
avatar: 'antling'
}
];
```
字段说明:
| 字段 | 说明 |
| ----------- | ------------------------------------------------------------ |
| `channelId` | AIProxy `ChannelType` 对应的数字 ID,必须和 `core/model/chtype.go` 一致 |
| `name` | FastGPT 渠道列表中的多语言显示名 |
| `avatar` | 渠道头像文件名,不包含扩展名 |
同时在 `channel-avatar/` 下增加头像文件:
```text
packages/infrastructure/src/static-data/models/channel-avatar/antling.svg
```
`avatar` 字段必须和 `channel-avatar/` 下的文件名一致。支持的头像扩展名包括 `svg`、`png`、`jpeg`、`webp`、`jpg`。
如果 AIProxy 仓库还没有该协议,需要先在 AIProxy 中新增 `ChannelType` 和 adaptor,并确认 adaptor 已在 [`core/relay/adaptors/register.go`](https://github.com/labring/aiproxy/blob/main/core/relay/adaptors/register.go) 中被引入。FastGPT 插件侧只声明渠道展示信息,不负责实现 AIProxy 协议适配逻辑。
## 校验
修改完成后,至少运行:
```bash
pnpm typecheck
```
如果修改了较多供应商、模型 schema 或静态资源加载逻辑,再运行:
```bash
pnpm test
```
提交前检查 `packages/infrastructure/src/static-data/models/` 的 diff,确认没有误删其他供应商模型、没有填错模型类型,并且新增的 `provider` Logo 或 `channel-avatar` 头像文件已经提交。
file: ./content/plugin/system-tool-development.en.mdx
meta: {
"title": "System Tool Development Guide",
"description": "FastGPT system tool development guide"
}
## Introduction
This document targets system tool development after FastGPT v4.15.0. The new FastGPT Plugin service unifies system tools, model presets, and similar capabilities as installable, updatable, runtime-isolated plugin packages. A plugin is eventually delivered to the FastGPT Plugin service as a `.pkg` file.
The currently stable system tool plugin types are:
* Single tool: one plugin exposes one tool and is declared with `defineTool()`.
* Tool suite: one plugin exposes multiple related child tools and is declared with `defineToolSet()`.
System tool plugins run in the runtime provided by the FastGPT Plugin service. The FastGPT main service invokes tools through the plugin service, and plugin code uses `@fastgpt-plugin/sdk-factory` to describe input, output, secret configuration, and execution logic.
## Differences From The Legacy Mechanism
1. The deployment relationship between FastGPT and FastGPT Plugin remains an external extension model, and the overall architecture is still microservice-based.
2. The plugin package protocol upgrades from the old built-in system tool directory to a unified `.pkg` format, making installation, version management, hot updates, and future plugin type expansion easier.
3. The plugin runtime is managed by the server. The current default runtime is `local-pool`, where each plugin version has its own process pool, queue, and runtime configuration.
4. Plugin metadata, input/output schemas, secret schemas, and icon assets are included in build artifacts for use by FastGPT pages, workflows, and Agents.
5. Tool development uses `@fastgpt-plugin/cli` and `@fastgpt-plugin/sdk-factory`. The legacy `config.ts`, `versionList`, and `bun run build:pkg` flow is no longer the primary development model.
## Information To Collect Before Development
Clarify these items before coding:
| Information | Description |
| -------------------------------- | ------------------------------------------------------------------------------------------- |
| Plugin type | `tool` or `tool-suite`. |
| Plugin ID | `pluginId`, globally stable and unique. Keep it unchanged after release. |
| Child tool ID | Required for tool suites. `children[].id` stays unchanged after release. |
| Chinese and English names | `name.en` and `name.zh-CN`. |
| Chinese and English descriptions | `description.en` and `description.zh-CN`. |
| Inputs | Type, constraints, default value, UI title, and description for each field. |
| Outputs | Type, meaning, and downstream usage for each field. |
| Secrets | API Key, Base URL, username/password, and similar values, described through `secretSchema`. |
| External API | Request method, auth method, timeout, rate limit, error response, and test account. |
| File capability | Use `ctx.invoke.uploadFile()` when file upload is needed. |
| Streaming output | Use `ctx.streamResponse()` when intermediate progress should be shown to the user. |
| Test cases | Include at least success, invalid parameters, auth failure, and upstream failure. |
Missing information that affects plugin ID, auth method, billing, or listing security should be confirmed first. Other missing information can use reasonable defaults, with assumptions recorded in the submission notes.
## Developing With An Agent
When using Claude Code, Codex, or another agent tool, copy this prompt:
```plaintext
请根据以下 FastGPT 官方插件开发 Skill 开发插件:
https://raw.githubusercontent.com/labring/fastgpt-official-plugins/refs/heads/main/.agents/skills/develop-fastgpt-plugin/SKILL.md
执行要求:
1. 先读取并理解该 Skill 的完整内容,后续开发流程以该 Skill 为准。
2. 在开始编码前,收集插件名称、插件类型、中文/英文名称与描述、输入输出、密钥、外部 API、预期行为、错误处理和测试样例。
3. 如需求缺失,最多提出 3 个关键问题;如果可以合理默认,说明假设后继续推进。
4. 使用 `@fastgpt-plugin/cli` 创建插件骨架,并优先遵循仓库内已有插件的结构、命名、测试和构建方式。
5. 实现完成后运行必要验证,包括测试、构建、插件检查和打包;无法验证的项目需要说明原因。
6. 最终输出变更文件、验证结果、剩余假设和需要人工确认的外部 API 行为。
```
When developing or maintaining SDK/CLI in the `fastgpt-plugin` repository, also refer to local skills:
* `sdk/factory/skills/fastgpt-plugin-development/SKILL.md`
* `sdk/factory/skills/fastgpt-system-tool-development/SKILL.md`
* `sdk/factory/skills/fastgpt-sdk-factory/SKILL.md`
## 1. Prepare Environment
Recommended environment:
* Node.js version that satisfies the target plugin repository.
* `pnpm`; the `fastgpt-plugin` repository uses pnpm workspace.
* Git.
* GitHub CLI `gh`, used for forking, creating repositories, and submitting PRs.
When developing community plugins, first fork and clone the community repository:
```bash
gh repo fork labring/fastgpt-community-plugins --clone
cd fastgpt-community-plugins
pnpm install
```
When debugging the CLI or SDK in the `fastgpt-plugin` repository, install dependencies and build the CLI/SDK first:
```bash
pnpm install
pnpm build:sdk-factory
pnpm build:cli
```
## 2. Create Plugin Skeleton
Single-tool plugin:
```bash
pnpx @fastgpt-plugin/cli create my-tool --type tool --cwd packages/tools
```
Tool-suite plugin:
```bash
pnpx @fastgpt-plugin/cli create my-tool-suite --type tool-suite --cwd packages/tools
```
You can also enter the target directory and create interactively:
```bash
pnpx @fastgpt-plugin/cli create
```
The CLI creates the plugin directory and common files:
| File | Purpose |
| ------------------ | ------------------------------------------------------------------------- |
| `index.ts` | Plugin entry, default-exporting `defineTool()` or `defineToolSet()`. |
| `package.json` | Plugin dependencies and `build`, `build:dev`, `pack`, and `test` scripts. |
| `tsconfig.json` | TypeScript config. |
| `vitest.config.ts` | Test config. |
| `README.md` | Plugin description. |
| `logo.svg` | Main plugin icon. |
## 3. Implement Single Tool
The system tool entry must default-export an SDK factory instance:
```ts
import {
createToolHandler,
defineTool,
type InputSchemaMetaType,
type OutputSchemaMetaType,
type SecretSchemaMetaType
} from '@fastgpt-plugin/sdk-factory';
import z from 'zod';
const secretSchema = z.object({
apiKey: z
.string()
.min(1)
.meta({
title: 'API Key',
isSecret: true
} satisfies SecretSchemaMetaType)
});
const handler = createToolHandler({
inputSchema: z.object({
query: z
.string()
.min(1)
.meta({
title: 'Query',
description: 'Search keyword'
} satisfies InputSchemaMetaType)
}),
outputSchema: z.object({
result: z.string().meta({
title: 'Result'
} satisfies OutputSchemaMetaType)
}),
secretSchema,
handler: async (input, ctx) => {
return {
result: input.query
};
}
});
export default defineTool({
manifest: {
pluginId: 'example-search',
version: '1.0.0',
name: {
en: 'Example Search',
'zh-CN': '示例搜索'
},
description: {
en: 'Search example data',
'zh-CN': '搜索示例数据'
},
versionDescription: {
en: 'Initial version',
'zh-CN': '初始版本'
},
tags: ['tools']
},
handler
});
```
Core rules:
* Keep `pluginId`, child tool `id`, input field names, and output field names stable after publishing.
* Use `{ en, 'zh-CN' }` for `manifest.name`, `manifest.description`, and `versionDescription`.
* Describe inputs, outputs, and secrets with Zod schemas.
* Add `InputSchemaMetaType` to input fields and `OutputSchemaMetaType` to output fields.
* Add `SecretSchemaMetaType` to secret fields and set `isSecret: true` for sensitive fields.
* Handler return values must match `outputSchema`.
* Convert external API errors into actionable messages and avoid exposing secrets, tokens, or complete sensitive responses.
* Use `ctx.invoke.uploadFile()` when host file upload is needed, and prefer preserving the returned `err`.
* Use `ctx.streamResponse()` when progress should be shown to users.
## 4. Implement Tool Suite
Use `defineToolSet()` for tool suites. Put shared information in the top-level `manifest` and `secretSchema`, and declare each child tool's independent `id`, name, description, and handler in `children`.
```ts
import {
createToolHandler,
defineToolSet,
type InputSchemaMetaType,
type OutputSchemaMetaType,
type SecretSchemaMetaType
} from '@fastgpt-plugin/sdk-factory';
import z from 'zod';
const secretSchema = z.object({
apiKey: z.string().meta({
title: 'API Key',
isSecret: true
} satisfies SecretSchemaMetaType)
});
const searchHandler = createToolHandler({
inputSchema: z.object({
query: z.string().meta({
title: 'Query'
} satisfies InputSchemaMetaType)
}),
outputSchema: z.object({
items: z.array(z.string()).meta({
title: 'Items'
} satisfies OutputSchemaMetaType)
}),
secretSchema,
handler: async (input) => ({ items: [input.query] })
});
const summaryHandler = createToolHandler({
inputSchema: z.object({
content: z.string().meta({
title: 'Content'
} satisfies InputSchemaMetaType)
}),
outputSchema: z.object({
summary: z.string().meta({
title: 'Summary'
} satisfies OutputSchemaMetaType)
}),
secretSchema,
handler: async (input) => ({ summary: input.content.slice(0, 100) })
});
export default defineToolSet({
manifest: {
pluginId: 'text-tools',
version: '1.0.0',
name: {
en: 'Text Tools',
'zh-CN': '文本工具集'
},
description: {
en: 'Search and summarize text',
'zh-CN': '搜索和总结文本'
}
},
children: [
{
id: 'search',
name: { en: 'Search', 'zh-CN': '搜索' },
description: { en: 'Search text', 'zh-CN': '搜索文本' },
toolDescription: 'Search text by query',
handler: searchHandler
},
{
id: 'summary',
name: { en: 'Summary', 'zh-CN': '总结' },
description: { en: 'Summarize text', 'zh-CN': '总结文本' },
toolDescription: 'Summarize text content',
handler: summaryHandler
}
],
secretSchema
});
```
## 5. Icon Conventions
During build, the CLI scans icons in the plugin root and writes them into the built `manifest.json`.
| Scenario | File name |
| --------------------- | --------------------------------------------------------------------------- |
| Main plugin icon | `logo.svg`, `logo.png`, `logo.jpg`, `logo.jpeg`, `logo.webp`, or `logo.gif` |
| Tool-suite child icon | `.logo.svg`, `.logo.png`, and similar names |
Notes:
* Put icon files in the plugin root.
* The `` of a child icon must exactly match `children[].id`.
* Keep only one extension for the same icon to avoid ambiguous scan results.
* Child tools without their own icons reuse the main plugin icon by default.
* After build, check the `icon` field in `dist/manifest.json`.
## 6. Local Debugging
Install dependencies in the plugin directory first:
```bash
cd packages/tools/my-tool
pnpm install
```
View plugin and debuggable tool information:
```bash
pnpx @fastgpt-plugin/cli debug .
```
Run one single-tool debug invocation:
```bash
pnpx @fastgpt-plugin/cli debug . --run --input '{"query":"hello"}' --secrets '{"apiKey":"test"}'
```
Run a child tool in a tool suite:
```bash
pnpx @fastgpt-plugin/cli debug . --run --tool search --input '{"query":"hello"}' --secrets '{"apiKey":"test"}'
```
Use files when input, secrets, or system variables are large:
```bash
pnpx @fastgpt-plugin/cli debug . --run --input-file input.json --secrets-file secrets.json --system-var-file system-var.json
```
Local debug boundaries:
* `ctx.invoke.uploadFile()` uses a local mock implementation and defaults to `.fastgpt-plugin-debug/uploads`.
* Local debug quickly validates plugin logic and schemas.
* Local debug does not simulate the production child-process pool, real Node.js IPC, network environment, server timeout, or queue scheduling.
* Before listing official plugins, still manually install plugins in a test environment and complete end-to-end testing.
## 7. Remote Debugging
Remote debugging connects a locally developed plugin to a FastGPT test environment. The FastGPT page authenticates the user and generates a debug link, while the CLI uses that link to create a WSS debug channel. Debug plugins are visible only to the current debugger.
Before using it, confirm that the test environment has deployed the FastGPT Plugin service and Connection Gateway, and that your local machine can reach the Gateway WSS endpoint returned by the test environment.
### 7.1 Generate A Debug Link
1. Sign in to the FastGPT test environment.
2. Go to the System Tools page and click Local Debug.

3. In the modal, click Generate Link and copy the debug link.
4. If a debug session already exists, click Refresh Link to generate a new connection key; the old link becomes invalid.

The debug link is only for connecting your local CLI to the test environment. Do not commit it to code repositories, documentation examples, or chat logs.
### 7.2 Start A Local Remote-Debug Session
Run this command in a plugin directory or a workspace that contains multiple plugin directories:
```bash
fastgpt-plugin dev
```
After startup, paste the debug link copied from FastGPT into the TUI. The CLI exchanges the connection key from the link for a short-lived WSS connect token, then mounts local plugins to the FastGPT debug channel.
Scripts and Agents can use non-interactive mode:
```bash
fastgpt-plugin dev --no-interactive \
--connect "https://fastgpt.example.com/api/plugin/debug-channel/connection-key/exchange?connectionKey=fpg_dbg_..."
```
When passing only a raw connection key, tell the CLI where the exchange endpoint is:
```bash
FASTGPT_PLUGIN_DEBUG_CONNECT_URL=https://fastgpt.example.com/api/plugin/debug-channel/connection-key/exchange \
fastgpt-plugin dev --no-interactive --connect "fpg_dbg_..."
```
After `--connect` connects successfully, it saves the connection key so later `fastgpt-plugin dev` runs can reuse the local config. In the TUI, press `c` to enter and save a new debug link or connection key.
### 7.3 Specify Plugin Directories And Watch Changes
When no plugin directory is passed, `dev` auto-discovers plugins from the current directory. If the current directory contains `index.ts`, it is used as the plugin entry; otherwise, the CLI scans one level of child directories for `index.ts`.
You can also pass one or more plugin directories explicitly:
```bash
fastgpt-plugin dev ./plugins/getTime ./plugins/dbops --watch
```
`--watch` reloads local plugins and recreates the remote-debug session after local file changes. The CLI reconnects by default when disconnected; use `--no-reconnect` to disable automatic reconnect.
### 7.4 Verify In FastGPT
After the CLI reports that remote debugging is ready, return to the FastGPT test environment:
1. Check the debug plugin on the System Tools page.
2. Select the debug tool in an app, workflow, or Agent.
3. Fill in secrets and input parameters, then start a real invocation.
4. Check local handler logs and errors in the CLI terminal.
The debug tool `source` is bound to the currently signed-in member, so other members do not see that debug plugin by default.
### 7.5 End Debugging
Press `Ctrl+C` in the local terminal to close the current CLI debug session; press `Ctrl+C` again to force exit.
The End Debugging action in FastGPT revokes the current member's debug channel and removes the debug plugin entry from the page. If the debug link is exposed, the signed-in member changes, or authorization needs to be renewed, use Refresh Link to generate a new link.
## 8. Build, Check, And Pack
Inside a plugin directory, usually run:
```bash
pnpm run test
pnpm run build
pnpx @fastgpt-plugin/cli check --entry . --output ./dist
pnpm run pack
```
You can also pass directories explicitly:
```bash
pnpx @fastgpt-plugin/cli build --entry packages/tools/my-tool --output packages/tools/my-tool/dist --minify
pnpx @fastgpt-plugin/cli check --entry packages/tools/my-tool --output packages/tools/my-tool/dist
pnpx @fastgpt-plugin/cli pack --entry packages/tools/my-tool --dist ./dist --output packages/tools/my-tool/out
```
Build artifacts should include:
* `dist/index.js`
* `dist/manifest.json`
* icon files
* optional `README.md`
* optional `assets/**`
Packaging produces a `.pkg` file. Uploading, installation, and listing should all use that `.pkg` file.
## 9. Verification Checklist
Before submitting, confirm:
* `index.ts` default export is correct.
* `manifest.pluginId`, `manifest.version`, Chinese and English names, and descriptions are complete.
* Tool suite `children[].id` values are stable and unique.
* `inputSchema` covers all user inputs and includes required type and range constraints.
* `outputSchema` matches handler return values.
* `secretSchema` covers all secret configuration and sensitive fields set `isSecret: true`.
* External API success, failure, empty response, timeout, and auth failure are handled.
* Error messages help locate issues and do not leak secrets or sensitive responses.
* `pnpm run test` passes, or the reason it cannot be tested is documented.
* `build`, `check`, and `pack` pass.
* Icons and schemas in `dist/manifest.json` are as expected.
* Remote debugging completes a real invocation in the test environment, or the reason remote debugging is not needed for this change is documented.
* `.pkg` can be installed in a test environment and complete a real invocation.
## 10. Release Flow
### Community Plugins
Community plugins usually start by creating and pushing an independent GitHub repository from the plugin directory:
```bash
cd packages/tools/my-tool
git init
git add .
git commit -m "feat: add my-tool plugin"
gh repo create --public --source=. --remote=origin --push
```
Then return to the `fastgpt-community-plugins` repository, submit the submodule or reference update, and open a PR to `labring/fastgpt-community-plugins`.
### Official Plugins
Official plugins require:
1. Code review.
2. Build, check, test, and package.
3. Manual `.pkg` installation in a test environment.
4. Complete functional testing, including external APIs, secret configuration, error paths, and concurrent calls.
5. Pre-listing security checks, focusing on SSRF, secret leakage, arbitrary file access, command execution, and dependency risk.
### Business Plugins
Business plugins are released to private repositories. Manage versions, secrets, installation packages, and acceptance records according to the customer delivery process. Security boundaries for external APIs, customer private addresses, and account secrets should be recorded separately.
If you do not need official inclusion, see [Upload System Tool](../guide/build/tools/system-plugins/upload_system_tool.en.mdx) to use the plugin in your own FastGPT deployment.
## FAQ
### How should I choose between `tool` and `tool-suite`?
Use `tool` for a single capability. Use `tool-suite` for multiple capabilities that share authentication, the same upstream API, and strong business relevance, such as search, detail, and task creation in one plugin.
### How should plugin versions be managed?
Use semantic versioning for `manifest.version`. Upgrade patch for compatible fixes, minor for compatible new features, and major when changing input/output fields, child tool IDs, or user configuration. Evaluate existing workflow compatibility before major changes.
### Can I put API keys in code or environment variables?
Plugins should declare secrets through `secretSchema` and read them through `ctx.secrets`. Real secrets should not appear in code repositories, test snapshots, error logs, or README files.
### Is a test environment still needed after local debug passes?
Yes. Local debug quickly validates plugin logic and schemas. Test environment validation confirms real installation, runtime, host reverse invocation, network, and permission behavior.
## References
* [FastGPT Plugin Repository](https://github.com/labring/fastgpt-plugin)
* [System Plugin Development Guide](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/how-to-devlop-plugin.en.md)
* [SDK Factory Guide](https://github.com/labring/fastgpt-plugin/blob/main/sdk/factory/README.en.md)
* [CLI Guide](https://github.com/labring/fastgpt-plugin/blob/main/apps/cli/README.en.md)
file: ./content/plugin/system-tool-development.mdx
meta: {
"title": "系统工具开发指南",
"description": "FastGPT 系统工具开发指南"
}
## 介绍
本文面向 FastGPT v4.15.0 之后的系统工具开发。新版 FastGPT Plugin 服务把系统工具、模型预设等能力统一抽象为可安装、可更新、可运行隔离的插件包,插件最终以 `.pkg` 文件交付给 FastGPT Plugin 服务。
当前稳定支持的系统工具插件类型有两种:
* 单工具:一个插件只暴露一个工具,使用 `defineTool()` 声明。
* 工具集:一个插件暴露多个相关子工具,使用 `defineToolSet()` 声明。
系统工具插件运行在 FastGPT Plugin 服务提供的运行时中。FastGPT 主服务通过插件服务调用工具,插件代码通过 `@fastgpt-plugin/sdk-factory` 描述输入、输出、密钥配置和执行逻辑。
## 与旧版机制的区别
1. FastGPT 和 FastGPT Plugin 的部署关系保持外置扩展模式,整体仍然是微服务架构。
2. 插件包协议从旧的内置系统工具目录升级为统一 `.pkg` 格式,便于安装、版本管理、热更新和后续扩展其他插件类型。
3. 插件运行时由服务端统一管理,当前默认运行时是 `local-pool`,每个插件版本拥有独立进程池、队列和运行时配置。
4. 插件元信息、输入输出 schema、密钥 schema 和图标资源都会进入构建产物,供 FastGPT 页面、工作流和 Agent 调用使用。
5. 工具开发使用 `@fastgpt-plugin/cli` 和 `@fastgpt-plugin/sdk-factory`,不再以旧版 `config.ts`、`versionList` 和 `bun run build:pkg` 作为主要开发方式。
## 开发前准备
开始编码前先明确这些信息:
| 信息 | 说明 |
| ------ | -------------------------------------------- |
| 插件类型 | `tool` 或 `tool-suite`。 |
| 插件 ID | `pluginId`,全局稳定唯一,发布后保持不变。 |
| 子工具 ID | 工具集需要,`children[].id` 发布后保持不变。 |
| 中英文名称 | `name.en` 和 `name.zh-CN`。 |
| 中英文描述 | `description.en` 和 `description.zh-CN`。 |
| 输入 | 每个字段的类型、约束、默认值、UI 标题和说明。 |
| 输出 | 每个字段的类型、含义和下游使用方式。 |
| 密钥 | API Key、Base URL、账号密码等,通过 `secretSchema` 描述。 |
| 外部 API | 请求方式、鉴权方式、超时、限流、错误响应和测试账号。 |
| 文件能力 | 需要上传文件时使用 `ctx.invoke.uploadFile()`。 |
| 流式输出 | 需要展示中间进度时使用 `ctx.streamResponse()`。 |
| 测试样例 | 至少包含成功路径、参数错误、鉴权失败和上游失败。 |
影响插件 ID、鉴权方式、计费或上架安全性的信息需要先确认。其他信息可以使用合理默认值继续推进,并在提交说明中记录假设。
## 使用 Agent 开发
使用 Claude Code、Codex 或其他 Agent 工具时,可直接复制下面的提示词:
```plaintext
请根据以下 FastGPT 官方插件开发 Skill 开发插件:
https://raw.githubusercontent.com/labring/fastgpt-official-plugins/refs/heads/main/.agents/skills/develop-fastgpt-plugin/SKILL.md
执行要求:
1. 先读取并理解该 Skill 的完整内容,后续开发流程以该 Skill 为准。
2. 在开始编码前,收集插件名称、插件类型、中文/英文名称与描述、输入输出、密钥、外部 API、预期行为、错误处理和测试样例。
3. 如需求缺失,最多提出 3 个关键问题;如果可以合理默认,说明假设后继续推进。
4. 使用 `@fastgpt-plugin/cli` 创建插件骨架,并优先遵循仓库内已有插件的结构、命名、测试和构建方式。
5. 实现完成后运行必要验证,包括测试、构建、插件检查和打包;无法验证的项目需要说明原因。
6. 最终输出变更文件、验证结果、剩余假设和需要人工确认的外部 API 行为。
```
在 `fastgpt-plugin` 仓库内开发或维护 SDK/CLI 时,也可以参考本地 Skill:
* `sdk/factory/skills/fastgpt-plugin-development/SKILL.md`
* `sdk/factory/skills/fastgpt-system-tool-development/SKILL.md`
* `sdk/factory/skills/fastgpt-sdk-factory/SKILL.md`
## 1. 准备开发环境
推荐环境:
* Node.js 版本满足目标插件仓库要求。
* `pnpm`,当前 `fastgpt-plugin` 仓库使用 pnpm workspace。
* Git。
* GitHub CLI `gh`,用于 fork、创建仓库和提交 PR。
开发社区插件时,先 fork 并 clone 社区插件仓库:
```bash
gh repo fork labring/fastgpt-community-plugins --clone
cd fastgpt-community-plugins
pnpm install
```
在 `fastgpt-plugin` 仓库内调试 CLI 或 SDK 时,先安装依赖并构建 CLI/SDK:
```bash
pnpm install
pnpm build:sdk-factory
pnpm build:cli
```
## 2. 创建插件骨架
单工具插件:
```bash
pnpx @fastgpt-plugin/cli create my-tool --type tool --cwd packages/tools
```
工具集插件:
```bash
pnpx @fastgpt-plugin/cli create my-tool-suite --type tool-suite --cwd packages/tools
```
也可以进入目标目录后交互式创建:
```bash
pnpx @fastgpt-plugin/cli create
```
CLI 会创建插件目录,并生成常见文件:
| 文件 | 作用 |
| ------------------ | --------------------------------------------- |
| `index.ts` | 插件入口,默认导出 `defineTool()` 或 `defineToolSet()`。 |
| `package.json` | 插件依赖和 `build`、`build:dev`、`pack`、`test` 脚本。 |
| `tsconfig.json` | TypeScript 配置。 |
| `vitest.config.ts` | 测试配置。 |
| `README.md` | 插件说明。 |
| `logo.svg` | 插件主图标。 |
## 3. 实现单工具
系统工具入口必须默认导出 SDK factory 实例:
```ts
import {
createToolHandler,
defineTool,
type InputSchemaMetaType,
type OutputSchemaMetaType,
type SecretSchemaMetaType
} from '@fastgpt-plugin/sdk-factory';
import z from 'zod';
const secretSchema = z.object({
apiKey: z
.string()
.min(1)
.meta({
title: 'API Key',
isSecret: true
} satisfies SecretSchemaMetaType)
});
const handler = createToolHandler({
inputSchema: z.object({
query: z
.string()
.min(1)
.meta({
title: 'Query',
description: 'Search keyword'
} satisfies InputSchemaMetaType)
}),
outputSchema: z.object({
result: z.string().meta({
title: 'Result'
} satisfies OutputSchemaMetaType)
}),
secretSchema,
handler: async (input, ctx) => {
return {
result: input.query
};
}
});
export default defineTool({
manifest: {
pluginId: 'example-search',
version: '1.0.0',
name: {
en: 'Example Search',
'zh-CN': '示例搜索'
},
description: {
en: 'Search example data',
'zh-CN': '搜索示例数据'
},
versionDescription: {
en: 'Initial version',
'zh-CN': '初始版本'
},
tags: ['tools']
},
handler
});
```
核心规则:
* `pluginId`、子工具 `id`、输入字段名、输出字段名发布后保持稳定。
* `manifest.name`、`manifest.description` 和 `versionDescription` 使用 `{ en, 'zh-CN' }`。
* 输入、输出和密钥都用 Zod schema 描述。
* 输入字段补充 `InputSchemaMetaType`,输出字段补充 `OutputSchemaMetaType`。
* 密钥字段补充 `SecretSchemaMetaType`,敏感字段设置 `isSecret: true`。
* handler 返回值必须匹配 `outputSchema`。
* 外部 API 错误需要转成可定位的错误信息,并避免输出密钥、令牌和完整敏感响应。
* 调用宿主文件上传能力时,使用 `ctx.invoke.uploadFile()`,并优先保留返回的 `err`。
* 展示进度时,使用 `ctx.streamResponse()`。
## 4. 实现工具集
工具集使用 `defineToolSet()`,把共用信息放在顶层 `manifest` 和 `secretSchema`,每个子工具在 `children` 中声明独立 `id`、名称、描述和 handler。
```ts
import {
createToolHandler,
defineToolSet,
type InputSchemaMetaType,
type OutputSchemaMetaType,
type SecretSchemaMetaType
} from '@fastgpt-plugin/sdk-factory';
import z from 'zod';
const secretSchema = z.object({
apiKey: z.string().meta({
title: 'API Key',
isSecret: true
} satisfies SecretSchemaMetaType)
});
const searchHandler = createToolHandler({
inputSchema: z.object({
query: z.string().meta({
title: 'Query'
} satisfies InputSchemaMetaType)
}),
outputSchema: z.object({
items: z.array(z.string()).meta({
title: 'Items'
} satisfies OutputSchemaMetaType)
}),
secretSchema,
handler: async (input) => ({ items: [input.query] })
});
const summaryHandler = createToolHandler({
inputSchema: z.object({
content: z.string().meta({
title: 'Content'
} satisfies InputSchemaMetaType)
}),
outputSchema: z.object({
summary: z.string().meta({
title: 'Summary'
} satisfies OutputSchemaMetaType)
}),
secretSchema,
handler: async (input) => ({ summary: input.content.slice(0, 100) })
});
export default defineToolSet({
manifest: {
pluginId: 'text-tools',
version: '1.0.0',
name: {
en: 'Text Tools',
'zh-CN': '文本工具集'
},
description: {
en: 'Search and summarize text',
'zh-CN': '搜索和总结文本'
}
},
children: [
{
id: 'search',
name: { en: 'Search', 'zh-CN': '搜索' },
description: { en: 'Search text', 'zh-CN': '搜索文本' },
toolDescription: 'Search text by query',
handler: searchHandler
},
{
id: 'summary',
name: { en: 'Summary', 'zh-CN': '总结' },
description: { en: 'Summarize text', 'zh-CN': '总结文本' },
toolDescription: 'Summarize text content',
handler: summaryHandler
}
],
secretSchema
});
```
## 5. 图标规范
CLI 构建时会扫描插件根目录中的图标并写入构建后的 `manifest.json`。
| 场景 | 文件名 |
| -------- | --------------------------------------------------------------------- |
| 主插件图标 | `logo.svg`、`logo.png`、`logo.jpg`、`logo.jpeg`、`logo.webp` 或 `logo.gif` |
| 工具集子工具图标 | `.logo.svg`、`.logo.png` 等 |
注意事项:
* 图标文件放在插件根目录。
* 子工具图标的 `` 与 `children[].id` 完全一致。
* 同一个图标只保留一个扩展名,避免扫描结果不明确。
* 子工具没有独立图标时,默认复用主插件图标。
* 构建后检查 `dist/manifest.json` 中的 `icon` 字段。
## 6. 本地调试
先进入插件目录安装依赖:
```bash
cd packages/tools/my-tool
pnpm install
```
查看插件和可调试工具信息:
```bash
pnpx @fastgpt-plugin/cli debug .
```
执行一次单工具调试:
```bash
pnpx @fastgpt-plugin/cli debug . --run --input '{"query":"hello"}' --secrets '{"apiKey":"test"}'
```
执行工具集中的某个子工具:
```bash
pnpx @fastgpt-plugin/cli debug . --run --tool search --input '{"query":"hello"}' --secrets '{"apiKey":"test"}'
```
输入、密钥和系统变量较大时,使用文件传入:
```bash
pnpx @fastgpt-plugin/cli debug . --run --input-file input.json --secrets-file secrets.json --system-var-file system-var.json
```
本地 debug 的边界:
* `ctx.invoke.uploadFile()` 使用本地虚拟实现,默认输出到 `.fastgpt-plugin-debug/uploads`。
* 本地 debug 用于快速验证插件逻辑和 schema。
* 本地 debug 不模拟生产子进程池、真实 Node.js IPC、网络环境、服务端超时和队列调度。
* 上架官方插件前仍需在测试环境中手动安装插件并完成端到端测试。
## 7. 远程调试
远程调试用于把本地正在开发的插件接入 FastGPT 测试环境。FastGPT 页面负责鉴权并生成调试链接,CLI 通过该链接建立 WSS 调试通道;调试插件仅对当前调试者本人可见。
使用前确认测试环境已部署 FastGPT Plugin 服务和 Connection Gateway,并且本地开发机可以访问测试环境返回的 Gateway WSS 地址。
### 7.1 生成调试链接
1. 登录 FastGPT 测试环境。
2. 进入「系统工具」页面,点击「本地调试」。

3. 在弹窗中点击「生成链接」,复制生成的调试链接。
4. 已有调试会话时,可点击「刷新链接」生成新的 connection key;旧链接会失效。

调试链接只用于本地 CLI 连接测试环境,不应提交到代码仓库、文档示例或聊天记录中。
### 7.2 启动本地远程调试会话
在插件目录或包含多个插件目录的工作区中运行:
```bash
fastgpt-plugin dev
```
启动后,将 FastGPT 页面复制的调试链接粘贴到 TUI 中。CLI 会用链接中的 connection key 换取短期 WSS connect token,并把本地插件挂载到 FastGPT 的调试通道。
脚本或 Agent 场景可以使用非交互模式:
```bash
fastgpt-plugin dev --no-interactive \
--connect "https://fastgpt.example.com/api/plugin/debug-channel/connection-key/exchange?connectionKey=fpg_dbg_..."
```
如果只传入裸 connection key,需要让 CLI 知道 exchange 接口地址:
```bash
FASTGPT_PLUGIN_DEBUG_CONNECT_URL=https://fastgpt.example.com/api/plugin/debug-channel/connection-key/exchange \
fastgpt-plugin dev --no-interactive --connect "fpg_dbg_..."
```
`--connect` 成功连接后会保存 connection key,后续可直接运行 `fastgpt-plugin dev` 复用本地配置。TUI 中按 `c` 可重新输入并保存新的调试链接或 connection key。
### 7.3 指定插件目录和监听变化
`dev` 未传插件目录时会自动探测当前目录:当前目录存在 `index.ts` 时使用当前目录;否则扫描下一层子目录中的 `index.ts`。
也可以手动传入一个或多个插件目录:
```bash
fastgpt-plugin dev ./plugins/getTime ./plugins/dbops --watch
```
`--watch` 会在本地文件变化后重新加载插件并重建远程调试会话。CLI 默认开启断线重连;如需关闭自动重连,可加 `--no-reconnect`。
### 7.4 在 FastGPT 中验证
CLI 显示远程调试已就绪后,回到 FastGPT 测试环境:
1. 在「系统工具」页面查看调试插件。
2. 在应用、工作流或 Agent 中选择该调试工具。
3. 填写密钥和输入参数,发起真实调用。
4. 在 CLI 终端查看本地 handler 日志和错误信息。
调试工具的 `source` 会绑定到当前登录成员,其他成员默认看不到该调试插件。
### 7.5 结束调试
本地终端按 `Ctrl+C` 会关闭当前 CLI 调试会话;再次按 `Ctrl+C` 会强制退出。
FastGPT 页面中的「结束调试」会撤销当前成员的 debug channel,并清理页面上的调试插件入口。调试链接泄露、成员切换或需要重新授权时,优先使用「刷新链接」生成新链接。
## 8. 构建、检查和打包
插件目录中通常可以直接运行:
```bash
pnpm run test
pnpm run build
pnpx @fastgpt-plugin/cli check --entry . --output ./dist
pnpm run pack
```
也可以显式传入目录:
```bash
pnpx @fastgpt-plugin/cli build --entry packages/tools/my-tool --output packages/tools/my-tool/dist --minify
pnpx @fastgpt-plugin/cli check --entry packages/tools/my-tool --output packages/tools/my-tool/dist
pnpx @fastgpt-plugin/cli pack --entry packages/tools/my-tool --dist ./dist --output packages/tools/my-tool/out
```
构建产物应包含:
* `dist/index.js`
* `dist/manifest.json`
* 图标文件
* 可选的 `README.md`
* 可选的 `assets/**`
打包后会生成 `.pkg` 文件。上传、安装和上架都应使用该 `.pkg` 文件。
## 9. 验证清单
提交前至少确认:
* `index.ts` 默认导出正确。
* `manifest.pluginId`、`manifest.version`、中英文名称和描述完整。
* 工具集的 `children[].id` 稳定且没有重复。
* `inputSchema` 覆盖所有用户输入,并有必要的类型和范围约束。
* `outputSchema` 与 handler 返回值一致。
* `secretSchema` 覆盖全部密钥配置,敏感字段设置 `isSecret: true`。
* 外部 API 的成功、失败、空响应、超时和鉴权失败都有处理。
* 错误信息可定位问题,并且不会泄露密钥或敏感响应。
* `pnpm run test` 通过,或明确说明无法测试的原因。
* `build`、`check`、`pack` 通过。
* `dist/manifest.json` 中图标和 schema 符合预期。
* 使用远程调试完成测试环境真实调用,或明确说明本次无需远程调试的原因。
* `.pkg` 能在测试环境中安装并完成真实调用。
## 10. 发布流程
### 社区插件
社区插件通常先在插件目录创建并推送独立 GitHub 仓库:
```bash
cd packages/tools/my-tool
git init
git add .
git commit -m "feat: add my-tool plugin"
gh repo create --public --source=. --remote=origin --push
```
然后回到 `fastgpt-community-plugins` 仓库,提交 submodule 或引用更新,并向 `labring/fastgpt-community-plugins` 提 PR。
### 官方插件
官方插件需要完成:
1. 代码 review。
2. 构建、检查、测试和打包。
3. 在测试环境手动安装 `.pkg`。
4. 完整功能测试,包括外部 API、密钥配置、错误路径和并发调用。
5. 上架前安全检查,重点关注 SSRF、密钥泄露、任意文件访问、命令执行和依赖风险。
### 商业插件
商业插件发布到私有仓库,按客户交付流程管理版本、密钥、安装包和验收记录。对外部 API、客户私有地址和账号密钥的处理需要单独记录安全边界。
如无需官方收录,可参考 [上传系统工具](../guide/build/tools/system-plugins/upload_system_tool.mdx) 在自己部署的 FastGPT 中使用。
## 常见问题
### `tool` 和 `tool-suite` 如何选择?
单一能力使用 `tool`。多个共享鉴权、共享上游 API、业务上强相关的能力使用 `tool-suite`,例如搜索、详情、创建任务放在同一个插件中。
### 插件版本如何管理?
`manifest.version` 使用语义化版本。修复兼容性问题升级 patch,新增兼容功能升级 minor,修改输入输出字段、子工具 ID 或用户配置方式时升级 major,并提前评估已有工作流兼容性。
### 可以把 API Key 写在代码或环境变量里吗?
插件应通过 `secretSchema` 声明密钥,并通过 `ctx.secrets` 读取。代码仓库、测试快照、错误日志和 README 中都不应出现真实密钥。
### 本地 debug 通过后还需要测试环境验证吗?
需要。本地 debug 用于快速验证插件逻辑和 schema,测试环境验证用于确认真实安装、运行时、宿主反向调用、网络和权限行为。
## 参考
* [FastGPT Plugin 仓库](https://github.com/labring/fastgpt-plugin)
* [系统插件开发指南](https://github.com/labring/fastgpt-plugin/blob/main/docs/dev/how-to-devlop-plugin.md)
* [SDK Factory 使用指南](https://github.com/labring/fastgpt-plugin/blob/main/sdk/factory/README.md)
* [CLI 使用指南](https://github.com/labring/fastgpt-plugin/blob/main/apps/cli/README.md)
file: ./content/openapi/app.en.mdx
meta: {
"title": "Application API",
"description": "FastGPT OpenAPI Application Interface"
}
## Prerequisites
1. Prepare your API Key: You can use the global API Key directly
2. Get your application's AppId

## Log API
### Get Application Overall Statistics
```bash
curl --location --request GET 'https://cloud.fastgpt.cn/api/proApi/core/app/logs/getTotalData?appId=68c46a70d950e8850ae564ba' \
--header 'Authorization: Bearer apikey'
```
```bash
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"totalUsers": 0,
"totalChats": 0,
"totalPoints": 0
}
}
```
**Request Parameters:**
* appId: Application ID
**Response Parameters:**
* totalUsers: Total number of users
* totalChats: Total number of conversations
* totalPoints: Total points consumed
### Get Application Chart Data
```bash
curl --location --request POST 'https://cloud.fastgpt.cn/api/proApi/core/app/logs/getChartData' \
--header 'Authorization: Bearer apikey' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "68c46a70d950e8850ae564ba",
"dateStart": "2025-09-19T16:00:00.000Z",
"dateEnd": "2025-09-27T15:59:59.999Z",
"offset": 1,
"source": [
"test",
"online",
"share",
"api",
"cronJob",
"team",
"feishu",
"official_account",
"wecom",
"mcp"
],
"userTimespan": "day",
"chatTimespan": "day",
"appTimespan": "day"
}'
```
```bash
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"userData": [
{
"timestamp": 1758585600000,
"summary": {
"userCount": 1,
"newUserCount": 0,
"retentionUserCount": 0,
"points": 1.1132600000000001,
"sourceCountMap": {
"test": 1,
"online": 0,
"share": 0,
"api": 0,
"cronJob": 0,
"team": 0,
"feishu": 0,
"official_account": 0,
"wecom": 0,
"mcp": 0
}
}
}
],
"chatData": [
{
"timestamp": 1758585600000,
"summary": {
"chatItemCount": 1,
"chatCount": 1,
"errorCount": 0,
"points": 1.1132600000000001
}
}
],
"appData": [
{
"timestamp": 1758585600000,
"summary": {
"goodFeedBackCount": 0,
"badFeedBackCount": 0,
"chatCount": 1,
"totalResponseTime": 22.31
}
}
]
}
}
```
**Request Parameters:**
* appId: Application ID
* dateStart: Start time
* dateEnd: End time
* source: Log source
* offset: User retention offset. The unit follows userTimespan
* userTimespan: User data timespan //day|week|month|quarter
* chatTimespan: Chat data timespan //day|week|month|quarter
* appTimespan: Application data timespan //day|week|month|quarter
**Response Parameters:**
* userData: User data array
* timestamp: Timestamp
* summary: Summary data object
* userCount: Active user count
* newUserCount: New user count
* retentionUserCount: Retained user count
* points: Total points consumed
* sourceCountMap: User count by source
* chatData: Chat data array
* timestamp: Timestamp
* summary: Summary data object
* chatItemCount: Chat message count
* chatCount: Session count
* errorCount: Error count
* points: Total points consumed
* appData: Application data array
* timestamp: Timestamp
* summary: Summary data object
* goodFeedBackCount: Positive feedback count
* badFeedBackCount: Negative feedback count
* chatCount: Chat count
* totalResponseTime: Total response time
file: ./content/openapi/app.mdx
meta: {
"title": "应用接口",
"description": "FastGPT OpenAPI 应用接口"
}
## 前置准备
1. 准备 API key: 可用直接使用全局 apikey
2. 准备应用的 AppId

## 日志接口
### 获取应用总体数据统计
```bash
curl --location --request GET 'https://cloud.fastgpt.cn/api/proApi/core/app/logs/getTotalData?appId=68c46a70d950e8850ae564ba' \
--header 'Authorization: Bearer apikey'
```
```bash
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"totalUsers": 0,
"totalChats": 0,
"totalPoints": 0
}
}
```
**入参:**
* appId: 应用 ID
**出参:**
* totalUsers: 累积使用用户数量
* totalChats: 累积对话数量
* totalPoints: 累积积分消耗
### 获取应用图表数据
```bash
curl --location --request POST 'https://cloud.fastgpt.cn/api/proApi/core/app/logs/getChartData' \
--header 'Authorization: Bearer apikey' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "68c46a70d950e8850ae564ba",
"dateStart": "2025-09-19T16:00:00.000Z",
"dateEnd": "2025-09-27T15:59:59.999Z",
"offset": 1,
"source": [
"test",
"online",
"share",
"api",
"cronJob",
"team",
"feishu",
"official_account",
"wecom",
"mcp"
],
"userTimespan": "day",
"chatTimespan": "day",
"appTimespan": "day"
}'
```
```bash
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"userData": [
{
"timestamp": 1758585600000,
"summary": {
"userCount": 1,
"newUserCount": 0,
"retentionUserCount": 0,
"points": 1.1132600000000001,
"sourceCountMap": {
"test": 1,
"online": 0,
"share": 0,
"api": 0,
"cronJob": 0,
"team": 0,
"feishu": 0,
"official_account": 0,
"wecom": 0,
"mcp": 0
}
}
}
],
"chatData": [
{
"timestamp": 1758585600000,
"summary": {
"chatItemCount": 1,
"chatCount": 1,
"errorCount": 0,
"points": 1.1132600000000001
}
}
],
"appData": [
{
"timestamp": 1758585600000,
"summary": {
"goodFeedBackCount": 0,
"badFeedBackCount": 0,
"chatCount": 1,
"totalResponseTime": 22.31
}
}
]
}
}
```
**入参:**
* appId: 应用 ID
* dateStart: 开始时间
* dateEnd: 结束时间
* source: 日志来源
* offset: 用户留存偏移量,单位随 userTimespan 变化
* userTimespan: 用户数据时间跨度 //day|week|month|quarter
* chatTimespan: 对话数据时间跨度 //day|week|month|quarter
* appTimespan: 应用数据时间跨度 //day|week|month|quarter
**出参:**
* userData: 用户数据数组
* timestamp: 时间戳
* summary: 汇总数据对象
* userCount: 活跃用户数量
* newUserCount: 新用户数量
* retentionUserCount: 留存用户数量
* points: 总积分消耗
* sourceCountMap: 各来源用户数量
* chatData: 对话数据数组
* timestamp: 时间戳
* summary: 汇总数据对象
* chatItemCount: 对话次数
* chatCount - 会话次数
* errorCount - 错误对话次数
* points - 总积分消耗
* appData: 应用数据数组
* timestamp - 时间戳
* summary - 汇总数据对象
* goodFeedBackCount - 好评反馈数量
* badFeedBackCount - 差评反馈数量
* chatCount - 对话次数
* totalResponseTime - 总响应时间
file: ./content/openapi/chat.en.mdx
meta: {
"title": "Chat API",
"description": "FastGPT OpenAPI Chat Interface"
}
# How to Get AppId
You can find the AppId in your application details URL.

# Start a Conversation
* Authenticate with an API Key. When calling `chat/completions` , passing `appId` in the request body is recommended.
* For OpenAI SDK compatibility, `Authorization: Bearer -` is also supported. The suffix is only a transport compatibility format and is not stored.
* To proxy a team member identity through `authProxy` , the team owner must enable `authProxy` when creating or editing the key. The proxied member must still have permission to access the target app and chat.
* Some packages require adding `v1` to the `BaseUrl` . If you get a 404 error, try adding `v1` and retry.
{/* * 对话现在有`v1`和`v2`两个接口,可以按需使用,v2 自 4.9.4 版本新增,v1 接口同时不再维护 */}
## Start Chat
The `v1` chat API is compatible with the `GPT` interface! If you're using the standard `GPT` official API, you can access FastGPT by simply changing the `BaseUrl` and `Authorization` . However, note these rules:
* Parameters like `model` and `temperature` are ignored. These values are determined by your workflow configuration.
* Won't return actual `Token` consumed. If needed, set `detail=true` and manually calculate `tokens` from `responseData` .
### Request
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"chatId": "my_chatId",
"stream": false,
"detail": false,
"responseChatItemId": "my_responseChatItemId",
"variables": {
"uid": "asdfadsfasfd2323",
"name": "张三"
},
"messages": [
{
"role": "user",
"content": "导演是谁"
}
]
}'
```
* Only `messages` differs slightly; other parameters are the same.
* Direct file uploads are not supported. Upload files to your object storage and provide the URL.
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"chatId": "abcd",
"stream": false,
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "导演是谁"
},
{
"type": "image_url",
"image_url": {
"url": "图片链接"
}
},
{
"type": "file_url",
"name": "文件名",
"url": "文档链接,支持 txt md html word pdf ppt csv excel"
}
]
}
]
}'
```
* headers.Authorization: Bearer \[apikey]
* chatId: string | undefined.
* Empty or omitted: FastGPT context is not used, and context is built entirely from `messages` .
* Non-empty string: uses `chatId` for the chat, automatically reads messages from the FastGPT session, and uses only the last item in `messages` as the user question. Other messages are ignored. Make sure `chatId` is unique and shorter than 250 characters.
* messages: Same structure as [GPT chat messages](https://platform.openai.com/docs/api-reference/chat/object) .
* responseChatItemId: string | undefined. If provided, FastGPT uses it as the response message ID and stores it in the database. Make sure it is unique under the current `chatId` .
* detail: Whether to return intermediate values. In `stream` mode, they are separated by `event` ; in non-stream mode, they are stored in `responseData` .
* variables: Module variables. This object replaces `[key]` placeholders in input fields.
### Response
```json
{
"id": "adsfasf",
"model": "",
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 1
},
"choices": [
{
"message": {
"role": "assistant",
"content": "电影《铃芽之旅》的导演是新海诚。"
},
"finish_reason": "stop",
"index": 0
}
]
}
```
```bash
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":""},"index":0,"finish_reason":null}]}
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":"电"},"index":0,"finish_reason":null}]}
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":"影"},"index":0,"finish_reason":null}]}
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":"《"},"index":0,"finish_reason":null}]}
```
```json
{
"responseData": [
// 不同模块的响应值, 不同版本具体值可能有差异,可先 log 自行查看最新值。
{
"moduleName": "Dataset Search",
"price": 1.2000000000000002,
"model": "Embedding-2",
"tokens": 6,
"similarity": 0.61,
"limit": 3
},
{
"moduleName": "AI Chat",
"price": 454.5,
"model": "FastAI-4k",
"tokens": 303,
"question": "导演是谁",
"answer": "电影《铃芽之旅》的导演是新海诚。",
"maxToken": 2050,
"quoteList": [
{
"dataset_id": "646627f4f7b896cfd8910e38",
"id": "8099",
"q": "本作的主人公是谁?",
"a": "本作的主人公是名叫铃芽的少女。",
"source": "手动修改"
},
{
"dataset_id": "646627f4f7b896cfd8910e38",
"id": "8686",
"q": "电影《铃芽之旅》男主角是谁?",
"a": "电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。",
"source": ""
},
{
"dataset_id": "646627f4f7b896cfd8910e38",
"id": "19339",
"q": "电影《铃芽之旅》的导演是谁?22",
"a": "电影《铃芽之旅》的导演是新海诚。",
"source": "手动修改"
}
],
"completeMessages": [
{
"obj": "System",
"value": "下面是知识库内容:\n1. [本作的主人公是谁?\n本作的主人公是名叫铃芽的少女。]\n2. [电影《铃芽之旅》男主角是谁?\n电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。]\n3. [电影《铃芽之旅》的导演是谁?22\n电影《铃芽之旅》的导演是新海诚。]\n"
},
{
"obj": "System",
"value": "1.请记住,你的身份是百度的下一代知识增强语言模型,能够完全根据知识库提供的内容回答问题。\n\n2. 你忘记了关于电影《铃芽之旅》以外的内容。"
},
{
"obj": "System",
"value": "你仅回答关于电影《玲芽之旅》的问题,其余问题直接回复: 我不清楚。"
},
{
"obj": "Human",
"value": "导演是谁"
},
{
"obj": "AI",
"value": "电影《铃芽之旅》的导演是新海诚。"
}
]
}
],
"id": "",
"model": "",
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 1
},
"choices": [
{
"message": {
"role": "assistant",
"content": "电影《铃芽之旅》的导演是新海诚。"
},
"finish_reason": "stop",
"index": 0
}
]
}
```
```bash
event: flowNodeStatus
data: {"status":"running","name":"知识库搜索"}
event: flowNodeStatus
data: {"status":"running","name":"AI 对话"}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"电影"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"《铃"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"芽之旅》"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"的导演是新"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"海诚。"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{},"index":0,"finish_reason":"stop"}]}
event: answer
data: [DONE]
event: flowResponses
data: [{"moduleName":"知识库搜索","moduleType":"datasetSearchNode","runningTime":1.78},{"question":"导演是谁","quoteList":[{"id":"654f2e49b64caef1d9431e8b","q":"电影《铃芽之旅》的导演是谁?","a":"电影《铃芽之旅》的导演是新海诚!","indexes":[{"type":"qa","dataId":"3515487","text":"电影《铃芽之旅》的导演是谁?","_id":"654f2e49b64caef1d9431e8c","defaultIndex":true}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8935586214065552},{"id":"6552e14c50f4a2a8e632af11","q":"导演是谁?","a":"电影《铃芽之旅》的导演是新海诚。","indexes":[{"defaultIndex":true,"type":"qa","dataId":"3644565","text":"导演是谁?\n电影《铃芽之旅》的导演是新海诚。","_id":"6552e14dde5cc7ba3954e417"}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8890955448150635},{"id":"654f34a0b64caef1d946337e","q":"本作的主人公是谁?","a":"本作的主人公是名叫铃芽的少女。","indexes":[{"type":"qa","dataId":"3515541","text":"本作的主人公是谁?","_id":"654f34a0b64caef1d946337f","defaultIndex":true}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8738770484924316},{"id":"654f3002b64caef1d944207a","q":"电影《铃芽之旅》男主角是谁?","a":"电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。","indexes":[{"type":"qa","dataId":"3515538","text":"电影《铃芽之旅》男主角是谁?","_id":"654f3002b64caef1d944207b","defaultIndex":true}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8607980012893677},{"id":"654f2fc8b64caef1d943fd46","q":"电影《铃芽之旅》的编剧是谁?","a":"新海诚是本片的编剧。","indexes":[{"defaultIndex":true,"type":"qa","dataId":"3515550","text":"电影《铃芽之旅》的编剧是谁?22","_id":"654f2fc8b64caef1d943fd47"}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8468944430351257}],"moduleName":"AI 对话","moduleType":"chatNode","runningTime":1.86}]
```
Event values:
* answer: Text returned to the client (counts as the final answer)
* chatTitle: Chat title generated from the current user question
* fastAnswer: Preset reply text returned to the client (counts as the final answer)
* toolCall: Tool execution
* toolParams: Tool parameters
* toolResponse: Tool response
* flowNodeStatus: Current workflow step status
* flowResponses: Complete workflow step responses
* updateVariables: Updated variables
* interactive: Interactive config
* error: Error
### Response
If your workflow contains interactive nodes, still call this API with `detail=true` :
* `stream=true` : read the interactive config from `event=interactive` in `data.interactive` .
* `stream=false` : read the element that contains the `interactive` field from `choices[].message.content` .
The `interactive` payload returned to external callers is display config only. It contains only `type` and `params` ; internal runtime fields such as `entryNodeIds` , `memoryEdges` , `nodeOutputs` , and `nodeResponseId` are not returned. If the workflow internally hits a children / loop / tool wrapper interaction, the API returns the deepest user-facing interaction.
When calling a workflow with interactive steps, if an interaction is encountered, it returns immediately. The examples below show the element inside `choices[].message.content[]` when `stream=false` ; when `stream=true` , `event=interactive` returns `{ "interactive": ... }` as its `data` :
```json
{
"interactive": {
"type": "userSelect",
"params": {
"description": "测试",
"userSelectOptions": [
{
"value": "Confirm",
"key": "option1"
},
{
"value": "Cancel",
"key": "option2"
}
]
}
}
}
```
```json
{
"interactive": {
"type": "userInput",
"params": {
"description": "测试",
"inputForm": [
{
"type": "input",
"key": "测试 1",
"label": "测试 1",
"description": "",
"value": "",
"defaultValue": "",
"valueType": "string",
"required": false,
"list": [
{
"label": "",
"value": ""
}
]
},
{
"type": "numberInput",
"key": "测试 2",
"label": "测试 2",
"description": "",
"value": "",
"defaultValue": "",
"valueType": "number",
"required": false,
"list": [
{
"label": "",
"value": ""
}
]
}
]
}
}
}
```
### Continue Interaction
After receiving interactive info, render your UI to guide user input or selection. Then call this API again to continue the workflow. Use this format:
For user selection, simply pass the selected value to messages.
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"stream": true,
"detail": true,
"chatId":"22222231",
"messages": [
{
"role": "user",
"content": "Confirm"
}
]
}'
```
Form input is slightly more complex. Serialize the input as a JSON string for `messages` . Object keys match form keys, values are user inputs. Ensure `chatId` is consistent.
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"stream": true,
"detail": true,
"chatId":"22231",
"messages": [
{
"role": "user",
"content": "{\"测试 1\":\"这是输入框的内容\",\"测试 2\":666}"
}
]
}'
```
## Request Plugin
Plugin API is identical to chat API, with slight parameter differences:
* 调用插件 Type 的应用时,接口默认为 `detail` 模式。
* No need to pass `chatId` since plugins run only once.
* No need to pass `messages` .
* Pass `variables` to represent plugin inputs.
* Get plugin outputs from `pluginData` .
### Request
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer test-xxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"stream": false,
"chatId": "test",
"variables": {
"query":"你好" # 我的插件输入有一个参数,变量名叫 query
}
}'
```
### Response
* Find plugin output by locating `moduleType=pluginOutput` in `responseData` . Its `pluginOutput` contains the output.
* Stream output is still available via `choices` .
```json
{
"responseData": [
{
"nodeId": "fdDgXQ6SYn8v",
"moduleName": "AI 对话",
"moduleType": "chatNode",
"totalPoints": 0.685,
"model": "FastAI-3.5",
"tokens": 685,
"query": "你好",
"maxToken": 2000,
"historyPreview": [
{
"obj": "Human",
"value": "你好"
},
{
"obj": "AI",
"value": "你好!有什么可以帮助你的吗?欢迎向我提问。"
}
],
"contextTotalLen": 14,
"runningTime": 1.73
},
{
"nodeId": "pluginOutput",
"moduleName": "插件输出",
"moduleType": "pluginOutput",
"totalPoints": 0,
"pluginOutput": {
"result": "你好!有什么可以帮助你的吗?欢迎向我提问。"
},
"runningTime": 0
}
],
"newVariables": {
"query": "你好"
},
"id": "safsafsa",
"model": "",
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 1
},
"choices": [
{
"message": {
"role": "assistant",
"content": "你好!有什么可以帮助你的吗?欢迎向我提问。"
},
"finish_reason": "stop",
"index": 0
}
]
}
```
* Get plugin output by deserializing the `event=flowResponses` string into an array. Find `moduleType=pluginOutput` element; its `pluginOutput` contains the output.
* Stream output works the same as chat API.
```bash
event: flowNodeStatus
data: {"status":"running","name":"AI 对话"}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":""},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"你"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"好"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"!"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"有"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"什"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"么"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"可以"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"帮"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"助"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"你"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"的"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"吗"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"?"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":""},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{},"index":0,"finish_reason":"stop"}]}
event: answer
data: [DONE]
event: flowResponses
data: [{"nodeId":"fdDgXQ6SYn8v","moduleName":"AI 对话","moduleType":"chatNode","totalPoints":0.033,"model":"FastAI-3.5","tokens":33,"query":"你好","maxToken":2000,"historyPreview":[{"obj":"Human","value":"你好"},{"obj":"AI","value":"你好!有什么可以帮助你的吗?"}],"contextTotalLen":2,"runningTime":1.42},{"nodeId":"pluginOutput","moduleName":"插件输出","moduleType":"pluginOutput","totalPoints":0,"pluginOutput":{"result":"你好!有什么可以帮助你的吗?"},"runningTime":0}]
```
event 取值:
* answer: 返回给客户端的文本(最终会算作回答)
* fastAnswer: 指定回复返回给客户端的文本(最终会算作回答)
* toolCall: 执行工具
* toolParams: 工具参数
* toolResponse: 工具返回
* flowNodeStatus: 运行到的节点状态
* flowResponses: 节点完整响应
* updateVariables: 更新变量
* error: 报错
# Chat CRUD
* The following APIs can be called with any `API Key` .
* 4.8.12 and above
\***\*Important Fields\*\***
* chatId - The ID of a session under an application
* dataId - The ID of a message under a session
## Session Management
### Get Session List
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/history/getHistories' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"offset": 0,
"pageSize": 20,
"source": "api"
}'
```
* appId - Application ID
* offset - Offset (starting position)
* pageSize - Number of items
* source - Chat source. `source=api` means get API-created sessions only (excludes web UI sessions)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"chatId": "usdAP1GbzSGu",
"updateTime": "2024-10-13T03:29:05.779Z",
"appId": "66e29b870b24ce35330c0f08",
"customTitle": "",
"title": "你好",
"top": false
},
{
"chatId": "lC0uTAsyNBlZ",
"updateTime": "2024-10-13T03:22:19.950Z",
"appId": "66e29b870b24ce35330c0f08",
"customTitle": "",
"title": "测试",
"top": false
}
],
"total": 2
}
}
```
### Update Session Title
```bash
curl --location --request PUT 'http://localhost:3000/api/core/chat/history/updateHistory' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"customTitle": "自定义标题"
}'
```
* appId - Application ID
* chatId - Session ID
* customTitle - Custom session title
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### Update Session Pin Status
```bash
curl --location --request PUT 'http://localhost:3000/api/core/chat/history/updateHistory' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"top": true
}'
```
* appId - Application ID
* chatId - Session ID
* top - Whether to pin. true = pin, false = unpin
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### Delete a Session
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/chat/history/delHistory?chatId=[chatId]&appId=[appId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - Application ID
* chatId - Session ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### Clear App Sessions
Only clears sessions created via API Key. Does not clear sessions from web UI, share links, or other sources.
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/chat/history/clearHistories?appId=[appId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - Application ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## Message Management
Operations on messages under a specific session.
### Get Session Basic Info
```bash
curl --location --request GET 'http://localhost:3000/api/core/chat/init?appId=[appId]&chatId=[chatId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - Application ID
* chatId - Session ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"chatId": "sPVOuEohjo3w",
"appId": "66e29b870b24ce35330c0f08",
"variables": {},
"app": {
"chatConfig": {
"questionGuide": true,
"ttsConfig": {
"type": "web"
},
"whisperConfig": {
"open": false,
"autoSend": false,
"autoTTSResponse": false
},
"chatInputGuide": {
"open": false,
"textList": [],
"customUrl": ""
},
"instruction": "",
"variables": [],
"fileSelectConfig": {
"canSelectFile": true,
"canSelectImg": true,
"maxFiles": 10
},
"_id": "66f1139aaab9ddaf1b5c596d",
"welcomeText": ""
},
"chatModels": ["GPT-4o-mini"],
"name": "测试",
"avatar": "/imgs/app/avatar/workflow.svg",
"intro": "",
"type": "advanced",
"pluginInputs": []
}
}
}
```
### Get Message List
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/record/getPaginationRecords' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"offset": 0,
"pageSize": 10,
"loadCustomFeedbacks": true
}'
```
* appId - Application ID
* chatId - Session ID
* offset - Offset
* pageSize - Number of items
* loadCustomFeedbacks - Whether to load custom feedbacks (optional)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"_id": "670b84e6796057dda04b0fd2",
"dataId": "jzqdV4Ap1u004rhd2WW8yGLn",
"obj": "Human",
"value": [
{
"text": {
"content": "你好"
}
}
],
"customFeedbacks": []
},
{
"_id": "670b84e6796057dda04b0fd3",
"dataId": "x9KQWcK9MApGdDQH7z7bocw1",
"obj": "AI",
"value": [
{
"text": {
"content": "你好!有什么我可以帮助你的吗?"
}
}
],
"customFeedbacks": [],
"totalQuoteList": [],
"totalRunningTime": 2.42,
"useAgentSandbox": false
}
],
"total": 2
}
}
```
### Get Message Run Details
```bash
curl --location --request GET 'http://localhost:3000/api/core/chat/record/getResData?appId=[appId]&chatId=[chatId]&dataId=[dataId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - Application ID
* chatId - Session ID
* dataId - Message ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": [
{
"id": "mVlxkz8NfyfU",
"nodeId": "448745",
"moduleName": "common:core.module.template.work_start",
"moduleType": "workflowStart",
"runningTime": 0
},
{
"id": "b3FndAdHSobY",
"nodeId": "z04w8JXSYjl3",
"moduleName": "AI 对话",
"moduleType": "chatNode",
"runningTime": 1.22,
"totalPoints": 0.02475,
"model": "GPT-4o-mini",
"tokens": 75,
"query": "测试",
"maxToken": 2000,
"historyPreview": [
{
"obj": "Human",
"value": "你好"
},
{
"obj": "AI",
"value": "你好!有什么我可以帮助你的吗?"
},
{
"obj": "Human",
"value": "测试"
},
{
"obj": "AI",
"value": "测试成功!请问你有什么具体的问题或者需要讨论的话题吗?"
}
],
"contextTotalLen": 4
}
]
}
```
### Delete Message
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/chat/record/delete?contentId=[contentId]&chatId=[chatId]&appId=[appId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - Application ID
* chatId - Session ID
* contentId - Message ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### Update Feedback (Like / Dislike)
Like / unlike:
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/feedback/updateUserFeedback' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"dataId": "dataId",
"userGoodFeedback": "yes"
}'
```
Dislike / remove dislike:
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/feedback/updateUserFeedback' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"dataId": "dataId",
"userBadFeedback": "yes"
}'
```
* appId - Application ID
* chatId - Session ID
* dataId - Message ID
* userGoodFeedback - User feedback when liking (optional). Omit to unlike.
* userBadFeedback - User feedback when disliking (optional). Omit to remove dislike.
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## Question Suggestions
**4.8.16 New API (version**
The suggested questions feature requires both appId and chatId. It automatically fetches the last 6 message turns from the session as context.
```bash
curl --location --request POST 'http://localhost:3000/api/core/ai/agent/v2/createQuestionGuide' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"questionGuide": {
"open": true,
"model": "GPT-4o-mini",
"customPrompt": "你是一个智能助手,请根据用户的问题生成猜你想问。"
}
}'
```
| 参数名 | 类型 | 必填 | 说明 |
| ------------- | ------ | -- | -------------------------------- |
| appId | string | ✅ | 应用 ID |
| chatId | string | ✅ | Session ID |
| questionGuide | object | | 自定义配置,不传的话,则会根据 appId,取最新发布版本的配置 |
```ts
type CreateQuestionGuideParams = OutLinkChatAuthProps & {
appId: string;
chatId: string;
questionGuide?: {
open: boolean;
model?: string;
customPrompt?: string;
};
};
```
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": ["你对AI有什么看法?", "想了解AI的应用吗?", "你希望AI能做什么?"]
}
```
file: ./content/openapi/chat.mdx
meta: {
"title": "对话接口",
"description": "FastGPT OpenAPI 对话接口"
}
## 如何获取 AppId
可在应用详情的路径里获取 AppId。

## 发起会话
### 密钥使用规范
* 使用 APIKey 鉴权。调用 `chat/completions` 时,推荐在请求体传入 `body.appId`。
* 为兼容 OpenAI SDK,也支持 `Authorization: Bearer -`,此时不需要传递 `body.appId`。
* 有些 SDK 调用时,`BaseUrl` 需要添加 `v1` 路径,有些不需要,如果出现 404 情况,可补充 `v1` 重试。
* appId 的优先级:`body.appId` , `-` , `apikey 关联的 appId(旧版适配)`
### 注意事项
* 如需通过 `authProxy` 代理团队成员身份,需要团队所有者在创建或编辑该 key 时开启 `authProxy`;代理身份仍需要具备目标应用和会话权限。(仅适用于 FastGPT >= v4.15.0)
* 传入的 `model`,`temperature` 等参数字段均无效,这些字段由编排决定,不会根据 API 参数改变。
* 不会返回实际消耗 `Token` 值,如果需要,可以设置 `detail=true`,并手动计算 `responseData` 里的 `tokens` 值。
### 请求
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"chatId": "my_chatId",
"stream": false,
"detail": false,
"responseChatItemId": "my_responseChatItemId",
"variables": {
"uid": "asdfadsfasfd2323",
"name": "张三"
},
"messages": [
{
"role": "user",
"content": "导演是谁"
}
]
}'
```
* 仅 `messages` 有部分区别,其他参数一致。
* 目前不支持上传文件,需上传到自己的对象存储中,获取对应的文件链接。
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"chatId": "abcd",
"stream": false,
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "导演是谁"
},
{
"type": "image_url",
"image_url": {
"url": "图片链接"
}
},
{
"type": "file_url",
"name": "文件名",
"url": "文档链接,支持 txt md html word pdf ppt csv excel"
}
]
}
]
}'
```
* headers.Authorization: Bearer \[apikey]
* chatId: string | undefined。
* 为时(不传入),不使用 FastGpt 提供的上下文功能,完全通过传入的 messages 构建上下文。
* 为 `非空字符串` 时,意味着使用 chatId 进行对话,自动从 FastGpt 数据库取会话,并使用 messages 数组最后一个内容作为用户问题,其余 message 会被忽略。请自行确保 chatId 唯一,长度小于 250,通常可以是自己系统的对话框 ID。
* messages: 结构与 [GPT 接口](https://platform.openai.com/docs/api-reference/chat/object) chat 模式一致。
* responseChatItemId: string | undefined。如果传入,则会将该值作为本次对话的响应消息的 ID,FastGPT 会自动将该 ID 存入数据库。请确保,在当前 `chatId` 下,`responseChatItemId` 是唯一的。
* detail: 是否返回中间值(模块状态,响应的完整结果等),`stream 模式` 下会通过 `event` 进行区分,`非 stream 模式` 结果保存在 `responseData` 中。
* variables: 模块变量,一个对象,会替换模块中,输入框内容里的 `[key]`
### 响应
```json
{
"id": "adsfasf",
"model": "",
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 1
},
"choices": [
{
"message": {
"role": "assistant",
"content": "电影《铃芽之旅》的导演是新海诚。"
},
"finish_reason": "stop",
"index": 0
}
]
}
```
```bash
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":""},"index":0,"finish_reason":null}]}
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":"电"},"index":0,"finish_reason":null}]}
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":"影"},"index":0,"finish_reason":null}]}
data: {"id":"","object":"","created":0,"choices":[{"delta":{"content":"《"},"index":0,"finish_reason":null}]}
```
```json
{
"responseData": [
// 不同模块的响应值, 不同版本具体值可能有差异,可先 log 自行查看最新值。
{
"moduleName": "Dataset Search",
"price": 1.2000000000000002,
"model": "Embedding-2",
"tokens": 6,
"similarity": 0.61,
"limit": 3
},
{
"moduleName": "AI Chat",
"price": 454.5,
"model": "FastAI-4k",
"tokens": 303,
"question": "导演是谁",
"answer": "电影《铃芽之旅》的导演是新海诚。",
"maxToken": 2050,
"quoteList": [
{
"dataset_id": "646627f4f7b896cfd8910e38",
"id": "8099",
"q": "本作的主人公是谁?",
"a": "本作的主人公是名叫铃芽的少女。",
"source": "手动修改"
},
{
"dataset_id": "646627f4f7b896cfd8910e38",
"id": "8686",
"q": "电影《铃芽之旅》男主角是谁?",
"a": "电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。",
"source": ""
},
{
"dataset_id": "646627f4f7b896cfd8910e38",
"id": "19339",
"q": "电影《铃芽之旅》的导演是谁?22",
"a": "电影《铃芽之旅》的导演是新海诚。",
"source": "手动修改"
}
],
"completeMessages": [
{
"obj": "System",
"value": "下面是知识库内容:\n1. [本作的主人公是谁?\n本作的主人公是名叫铃芽的少女。]\n2. [电影《铃芽之旅》男主角是谁?\n电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。]\n3. [电影《铃芽之旅》的导演是谁?22\n电影《铃芽之旅》的导演是新海诚。]\n"
},
{
"obj": "System",
"value": "1.请记住,你的身份是百度的下一代知识增强语言模型,能够完全根据知识库提供的内容回答问题。\n\n2. 你忘记了关于电影《铃芽之旅》以外的内容。"
},
{
"obj": "System",
"value": "你仅回答关于电影《玲芽之旅》的问题,其余问题直接回复: 我不清楚。"
},
{
"obj": "Human",
"value": "导演是谁"
},
{
"obj": "AI",
"value": "电影《铃芽之旅》的导演是新海诚。"
}
]
}
],
"id": "",
"model": "",
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 1
},
"choices": [
{
"message": {
"role": "assistant",
"content": "电影《铃芽之旅》的导演是新海诚。"
},
"finish_reason": "stop",
"index": 0
}
]
}
```
```bash
event: flowNodeStatus
data: {"status":"running","name":"知识库搜索"}
event: flowNodeStatus
data: {"status":"running","name":"AI 对话"}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"电影"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"《铃"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"芽之旅》"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"的导演是新"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"content":"海诚。"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{},"index":0,"finish_reason":"stop"}]}
event: answer
data: [DONE]
event: flowResponses
data: [{"moduleName":"知识库搜索","moduleType":"datasetSearchNode","runningTime":1.78},{"question":"导演是谁","quoteList":[{"id":"654f2e49b64caef1d9431e8b","q":"电影《铃芽之旅》的导演是谁?","a":"电影《铃芽之旅》的导演是新海诚!","indexes":[{"type":"qa","dataId":"3515487","text":"电影《铃芽之旅》的导演是谁?","_id":"654f2e49b64caef1d9431e8c","defaultIndex":true}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8935586214065552},{"id":"6552e14c50f4a2a8e632af11","q":"导演是谁?","a":"电影《铃芽之旅》的导演是新海诚。","indexes":[{"defaultIndex":true,"type":"qa","dataId":"3644565","text":"导演是谁?\n电影《铃芽之旅》的导演是新海诚。","_id":"6552e14dde5cc7ba3954e417"}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8890955448150635},{"id":"654f34a0b64caef1d946337e","q":"本作的主人公是谁?","a":"本作的主人公是名叫铃芽的少女。","indexes":[{"type":"qa","dataId":"3515541","text":"本作的主人公是谁?","_id":"654f34a0b64caef1d946337f","defaultIndex":true}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8738770484924316},{"id":"654f3002b64caef1d944207a","q":"电影《铃芽之旅》男主角是谁?","a":"电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。","indexes":[{"type":"qa","dataId":"3515538","text":"电影《铃芽之旅》男主角是谁?","_id":"654f3002b64caef1d944207b","defaultIndex":true}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8607980012893677},{"id":"654f2fc8b64caef1d943fd46","q":"电影《铃芽之旅》的编剧是谁?","a":"新海诚是本片的编剧。","indexes":[{"defaultIndex":true,"type":"qa","dataId":"3515550","text":"电影《铃芽之旅》的编剧是谁?22","_id":"654f2fc8b64caef1d943fd47"}],"datasetId":"646627f4f7b896cfd8910e38","collectionId":"653279b16cd42ab509e766e8","sourceName":"data (81).csv","sourceId":"64fd3b6423aa1307b65896f6","score":0.8468944430351257}],"moduleName":"AI 对话","moduleType":"chatNode","runningTime":1.86}]
```
event 取值:
* answer: 返回给客户端的文本(最终会算作回答)
* chatTitle: 根据本轮用户问题生成的对话标题
* fastAnswer: 指定回复返回给客户端的文本(最终会算作回答)
* toolCall: 执行工具
* toolParams: 工具参数
* toolResponse: 工具返回
* flowNodeStatus: 运行到的节点状态
* flowResponses: 节点完整响应
* updateVariables: 更新变量
* interactive: 交互节点配置
* error: 报错
### 交互节点响应
如果工作流中包含交互节点,依然是调用该 API 接口,需要设置 `detail=true`:
* `stream=true`:可从 `event=interactive` 的 `data.interactive` 中获取交互节点配置。
* `stream=false`:可从 `choices[].message.content` 中获取包含 `interactive` 字段的元素。
返回给外部调用方的 `interactive` 是展示配置,只包含 `type` 和 `params`;`entryNodeIds` / `memoryEdges` / `nodeOutputs` / `nodeResponseId` 等内部运行态字段不会返回。若内部命中 children / loop / tool 包装交互,接口会返回最深层面向用户的交互节点。
当你调用一个带交互节点的工作流时,如果工作流遇到了交互节点,那么会直接返回。下面示例展示 `stream=false` 时 `choices[].message.content[]` 中的元素;`stream=true` 时 `event=interactive` 的 `data` 为 `{ "interactive": ... }`:
```json
{
"interactive": {
"type": "userSelect",
"params": {
"description": "测试",
"userSelectOptions": [
{
"value": "Confirm",
"key": "option1"
},
{
"value": "Cancel",
"key": "option2"
}
]
}
}
}
```
```json
{
"interactive": {
"type": "userInput",
"params": {
"description": "测试",
"inputForm": [
{
"type": "input",
"key": "测试 1",
"label": "测试 1",
"description": "",
"value": "",
"defaultValue": "",
"valueType": "string",
"required": false,
"list": [
{
"label": "",
"value": ""
}
]
},
{
"type": "numberInput",
"key": "测试 2",
"label": "测试 2",
"description": "",
"value": "",
"defaultValue": "",
"valueType": "number",
"required": false,
"list": [
{
"label": "",
"value": ""
}
]
}
]
}
}
}
```
### 交互节点继续运行
紧接着上一节,当你接收到交互节点信息后,可以根据这些数据进行 UI 渲染,引导用户输入或选择相关信息。然后需要再次发起会话,来继续工作流。调用的接口与仍是该接口,你需要按以下格式来发起请求:
对于用户选择,你只需要直接传递一个选择的结果给 messages 即可。
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"stream": true,
"detail": true,
"chatId":"22222231",
"messages": [
{
"role": "user",
"content": "Confirm"
}
]
}'
```
表单输入稍微麻烦一点,需要将输入的内容,以对象形式并序列化成字符串,作为 `messages` 的值。对象的 key 对应表单的 key,value 为用户输入的值。务必确保 `chatId` 是一致的。
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer fastgpt-xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"stream": true,
"detail": true,
"chatId":"22231",
"messages": [
{
"role": "user",
"content": "{\"测试 1\":\"这是输入框的内容\",\"测试 2\":666}"
}
]
}'
```
## 请求插件
插件的接口与对话接口一致,仅请求参数略有区别,有以下规定:
* 调用插件类型的应用时,接口默认为 `detail` 模式。
* 无需传入 `chatId`,因为插件只能运行一轮。
* 无需传入 `messages`。
* 通过传递 `variables` 来代表插件的输入。
* 通过获取 `pluginData` 来获取插件输出。
### 请求示例
```bash
curl --location --request POST 'http://localhost:3000/api/v1/chat/completions' \
--header 'Authorization: Bearer test-xxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "your_app_id",
"stream": false,
"chatId": "test",
"variables": {
"query":"你好" # 我的插件输入有一个参数,变量名叫 query
}
}'
```
### 响应示例
* 插件的输出可以通过查找 `responseData` 中, `moduleType=pluginOutput` 的元素,其 `pluginOutput` 是插件的输出。
* 流输出,仍可以通过 `choices` 进行获取。
```json
{
"responseData": [
{
"nodeId": "fdDgXQ6SYn8v",
"moduleName": "AI 对话",
"moduleType": "chatNode",
"totalPoints": 0.685,
"model": "FastAI-3.5",
"tokens": 685,
"query": "你好",
"maxToken": 2000,
"historyPreview": [
{
"obj": "Human",
"value": "你好"
},
{
"obj": "AI",
"value": "你好!有什么可以帮助你的吗?欢迎向我提问。"
}
],
"contextTotalLen": 14,
"runningTime": 1.73
},
{
"nodeId": "pluginOutput",
"moduleName": "插件输出",
"moduleType": "pluginOutput",
"totalPoints": 0,
"pluginOutput": {
"result": "你好!有什么可以帮助你的吗?欢迎向我提问。"
},
"runningTime": 0
}
],
"newVariables": {
"query": "你好"
},
"id": "safsafsa",
"model": "",
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 1
},
"choices": [
{
"message": {
"role": "assistant",
"content": "你好!有什么可以帮助你的吗?欢迎向我提问。"
},
"finish_reason": "stop",
"index": 0
}
]
}
```
* 插件的输出可以通过获取 `event=flowResponses` 中的字符串,并将其反序列化后得到一个数组。同样的,查找 `moduleType=pluginOutput` 的元素,其 `pluginOutput` 是插件的输出。
* 流输出,仍和对话接口一样获取。
```bash
event: flowNodeStatus
data: {"status":"running","name":"AI 对话"}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":""},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"你"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"好"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"!"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"有"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"什"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"么"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"可以"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"帮"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"助"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"你"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"的"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"吗"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":"?"},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{"role":"assistant","content":""},"index":0,"finish_reason":null}]}
event: answer
data: {"id":"","object":"","created":0,"model":"","choices":[{"delta":{},"index":0,"finish_reason":"stop"}]}
event: answer
data: [DONE]
event: flowResponses
data: [{"nodeId":"fdDgXQ6SYn8v","moduleName":"AI 对话","moduleType":"chatNode","totalPoints":0.033,"model":"FastAI-3.5","tokens":33,"query":"你好","maxToken":2000,"historyPreview":[{"obj":"Human","value":"你好"},{"obj":"AI","value":"你好!有什么可以帮助你的吗?"}],"contextTotalLen":2,"runningTime":1.42},{"nodeId":"pluginOutput","moduleName":"插件输出","moduleType":"pluginOutput","totalPoints":0,"pluginOutput":{"result":"你好!有什么可以帮助你的吗?"},"runningTime":0}]
```
event 取值:
* answer: 返回给客户端的文本(最终会算作回答)
* fastAnswer: 指定回复返回给客户端的文本(最终会算作回答)
* toolCall: 执行工具
* toolParams: 工具参数
* toolResponse: 工具返回
* flowNodeStatus: 运行到的节点状态
* flowResponses: 节点完整响应
* updateVariables: 更新变量
* error: 报错
# 对话 CRUD
**重要字段**
* appId - 应用 ID。
* chatId - 指一个应用下,某一个会话的 ID
* dataId - 指一个会话下,某一个对话的 ID
## 会话管理
### 获取会话列表
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/history/getHistories' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"offset": 0,
"pageSize": 20,
"source": "api"
}'
```
* appId - 应用 ID
* offset - 偏移量,即从第几条数据开始取
* pageSize - 记录数量
* source - 对话源。source=api,表示获取通过 API 创建的会话(不会获取页面上的会话)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"chatId": "usdAP1GbzSGu",
"updateTime": "2024-10-13T03:29:05.779Z",
"appId": "66e29b870b24ce35330c0f08",
"customTitle": "",
"title": "你好",
"top": false
},
{
"chatId": "lC0uTAsyNBlZ",
"updateTime": "2024-10-13T03:22:19.950Z",
"appId": "66e29b870b24ce35330c0f08",
"customTitle": "",
"title": "测试",
"top": false
}
],
"total": 2
}
}
```
### 修改会话标题
```bash
curl --location --request PUT 'http://localhost:3000/api/core/chat/history/updateHistory' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"customTitle": "自定义标题"
}'
```
* appId - 应用 ID
* chatId - 会话 ID
* customTitle - 自定义会话名
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### 修改会话置顶状态
```bash
curl --location --request PUT 'http://localhost:3000/api/core/chat/history/updateHistory' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"top": true
}'
```
* appId - 应用 ID
* chatId - 会话 ID
* top - 是否置顶,true 置顶,false 取消置顶
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### 删除单个会话
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/chat/history/delHistory?chatId=[chatId]&appId=[appId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - 应用 ID
* chatId - 会话 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### 清空应用会话
仅会清空通过 API Key 创建的会话,不会清空在线使用、分享链接等其他来源的会话。
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/chat/history/clearHistories?appId=[appId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - 应用 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## 对话管理
指的是某个会话下的会话操作。
### 获取会话基本信息
```bash
curl --location --request GET 'http://localhost:3000/api/core/chat/init?appId=[appId]&chatId=[chatId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - 应用 ID
* chatId - 会话 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"chatId": "sPVOuEohjo3w",
"appId": "66e29b870b24ce35330c0f08",
"variables": {},
"app": {
"chatConfig": {
"questionGuide": true,
"ttsConfig": {
"type": "web"
},
"whisperConfig": {
"open": false,
"autoSend": false,
"autoTTSResponse": false
},
"chatInputGuide": {
"open": false,
"textList": [],
"customUrl": ""
},
"instruction": "",
"variables": [],
"fileSelectConfig": {
"canSelectFile": true,
"canSelectImg": true,
"maxFiles": 10
},
"_id": "66f1139aaab9ddaf1b5c596d",
"welcomeText": ""
},
"chatModels": ["GPT-4o-mini"],
"name": "测试",
"avatar": "/imgs/app/avatar/workflow.svg",
"intro": "",
"type": "advanced",
"pluginInputs": []
}
}
}
```
### 获取对话列表
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/record/getPaginationRecords' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"offset": 0,
"pageSize": 10,
"loadCustomFeedbacks": true
}'
```
* appId - 应用 ID
* chatId - 会话 ID
* offset - 偏移量
* pageSize - 记录数量
* loadCustomFeedbacks - 是否读取自定义反馈(可选)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"_id": "670b84e6796057dda04b0fd2",
"dataId": "jzqdV4Ap1u004rhd2WW8yGLn",
"obj": "Human",
"value": [
{
"text": {
"content": "你好"
}
}
],
"customFeedbacks": []
},
{
"_id": "670b84e6796057dda04b0fd3",
"dataId": "x9KQWcK9MApGdDQH7z7bocw1",
"obj": "AI",
"value": [
{
"text": {
"content": "你好!有什么我可以帮助你的吗?"
}
}
],
"customFeedbacks": [],
"totalQuoteList": [],
"totalRunningTime": 2.42
}
],
"total": 2
}
}
```
### 获取单个对话运行详情
```bash
curl --location --request GET 'http://localhost:3000/api/core/chat/record/getResData?appId=[appId]&chatId=[chatId]&dataId=[dataId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - 应用 ID
* chatId - 会话 ID
* dataId - 对话 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": [
{
"id": "mVlxkz8NfyfU",
"nodeId": "448745",
"moduleName": "common:core.module.template.work_start",
"moduleType": "workflowStart",
"runningTime": 0
},
{
"id": "b3FndAdHSobY",
"nodeId": "z04w8JXSYjl3",
"moduleName": "AI 对话",
"moduleType": "chatNode",
"runningTime": 1.22,
"totalPoints": 0.02475,
"model": "GPT-4o-mini",
"tokens": 75,
"query": "测试",
"maxToken": 2000,
"historyPreview": [
{
"obj": "Human",
"value": "你好"
},
{
"obj": "AI",
"value": "你好!有什么我可以帮助你的吗?"
},
{
"obj": "Human",
"value": "测试"
},
{
"obj": "AI",
"value": "测试成功!请问你有什么具体的问题或者需要讨论的话题吗?"
}
],
"contextTotalLen": 4
}
]
}
```
### 删除对话
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/chat/record/delete?contentId=[contentId]&chatId=[chatId]&appId=[appId]' \
--header 'Authorization: Bearer [apikey]'
```
* appId - 应用 ID
* chatId - 会话 ID
* contentId - 对话 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### 更新反馈(点赞 / 点踩)
点赞 / 取消点赞:
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/feedback/updateUserFeedback' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"dataId": "dataId",
"userGoodFeedback": "yes"
}'
```
点踩 / 取消点踩:
```bash
curl --location --request POST 'http://localhost:3000/api/core/chat/feedback/updateUserFeedback' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"dataId": "dataId",
"userBadFeedback": "yes"
}'
```
* appId - 应用 ID
* chatId - 会话 ID
* dataId - 对话 ID
* userGoodFeedback - 用户点赞时的信息(可选),取消点赞时不填此参数即可
* userBadFeedback - 用户点踩时的信息(可选),取消点踩时不填此参数即可
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## 猜你想问
**4.8.16 后新版接口**
新版猜你想问必须包含 appId 和 chatId 参数。系统会根据 chatId 拉取最近 6 轮对话作为上下文来引导回答。
```bash
curl --location --request POST 'http://localhost:3000/api/core/ai/agent/v2/createQuestionGuide' \
--header 'Authorization: Bearer [apikey]' \
--header 'Content-Type: application/json' \
--data-raw '{
"appId": "appId",
"chatId": "chatId",
"questionGuide": {
"open": true,
"model": "GPT-4o-mini",
"customPrompt": "你是一个智能助手,请根据用户的问题生成猜你想问。"
}
}'
```
| 参数名 | 类型 | 必填 | 说明 |
| ------------- | ------ | -- | -------------------------------- |
| appId | string | ✅ | 应用 ID |
| chatId | string | ✅ | 会话 ID |
| questionGuide | object | | 自定义配置,不传的话,则会根据 appId,取最新发布版本的配置 |
```ts
type CreateQuestionGuideParams = OutLinkChatAuthProps & {
appId: string;
chatId: string;
questionGuide?: {
open: boolean;
model?: string;
customPrompt?: string;
};
};
```
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": ["你对AI有什么看法?", "想了解AI的应用吗?", "你希望AI能做什么?"]
}
```
file: ./content/openapi/dataset.en.mdx
meta: {
"title": "Dataset API",
"description": "FastGPT OpenAPI Dataset API"
}
| How to Get Dataset ID (datasetId) | How to Get Collection ID (collection\_id) |
| --------------------------------------- | ----------------------------------------- |
|  |  |
## Dataset
### Create Knowledge Base
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/create' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"parentId": null,
"type": "dataset",
"name":"测试",
"intro":"介绍",
"avatar": "",
"vectorModel": "text-embedding-ada-002",
"agentModel": "gpt-3.5-turbo-16k",
"vlmModel": "gpt-4.1"
}'
```
* parentId - Parent ID for building directory structure. Usually can be null or omitted.
* type - `dataset` or `folder`, represents regular dataset or folder. If not provided, creates a regular dataset.
* name - Dataset name (required)
* intro - Description (optional)
* avatar - Avatar URL (optional)
* vectorModel - Vector model (recommended to leave empty, use system default)
* agentModel - Text processing model (recommended to leave empty, use system default)
* vlmModel - Image understanding model (recommended to leave empty, use system default)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": "65abc9bd9d1448617cba5e6c"
}
```
### Get Dataset List
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/list?parentId=' \
--header 'Authorization: Bearer xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"parentId":""
}'
```
* parentId - Parent ID. Pass empty string or null to get datasets in the root directory
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": [
{
"_id": "65abc9bd9d1448617cba5e6c",
"parentId": null,
"avatar": "",
"name": "测试",
"intro": "",
"type": "dataset",
"permission": "private",
"canWrite": true,
"isOwner": true,
"vectorModel": {
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"charsPointsPrice": 0,
"defaultToken": 512,
"maxToken": 8000,
"weight": 100
}
}
]
}
```
### Get Knowledge Base Details
```bash
curl --location --request GET 'http://localhost:3000/api/core/dataset/detail?id=6593e137231a2be9c5603ba7' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: Dataset ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"_id": "6593e137231a2be9c5603ba7",
"parentId": null,
"teamId": "65422be6aa44b7da77729ec8",
"tmbId": "65422be6aa44b7da77729ec9",
"type": "dataset",
"status": "active",
"avatar": "/icon/logo.svg",
"name": "FastGPT test",
"vectorModel": {
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"charsPointsPrice": 0,
"defaultToken": 512,
"maxToken": 8000,
"weight": 100
},
"agentModel": {
"model": "gpt-3.5-turbo-16k",
"name": "FastAI-16k",
"maxContext": 16000,
"maxResponse": 16000,
"charsPointsPrice": 0
},
"intro": "",
"permission": "private",
"updateTime": "2024-01-02T10:11:03.084Z",
"canWrite": true,
"isOwner": true
}
}
```
### Delete Knowledge Base
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/dataset/delete?id=65abc8729d1448617cba5df6' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: Dataset ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## Collection
### Common Creation Parameters (Must Read)
**Request**
| Parameter | Description | Required |
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | -------- |
| datasetId | Dataset ID | ✅ |
| parentId: | Parent ID. Defaults to root directory if not provided | |
| trainingType | Data processing method. chunk: split by text length; qa: Q\&A extraction | ✅ |
| indexPrefixTitle | Whether to auto-generate title index | |
| customPdfParse | Whether to enable enhanced PDF parsing. Default false: disabled; true: enabled | |
| autoIndexes | Whether to auto-generate indexes (commercial version only) | |
| imageIndex | Whether to auto-generate image indexes (commercial version only) | |
| chunkSettingMode | Chunk parameter mode. auto: system default; custom: manual specification | |
| chunkSplitMode | Chunk split mode. size: split by length; char: split by character. Ineffective when chunkSettingMode=auto. | |
| chunkSize | Chunk size, default 1500. Ineffective when chunkSettingMode=auto. | |
| indexSize | Index size, default 512, must be less than index model max token. Ineffective when chunkSettingMode=auto. | |
| chunkSplitter | Custom highest priority split symbol. Won't split further unless exceeding file processing max context. Ineffective when chunkSettingMode=auto. | |
| qaPrompt | QA split prompt | |
| tags | Collection tags (string array) | |
| createTime | File creation time (Date / String) | |
**Response**
* collectionId - New collection ID
* insertLen:Number of inserted chunks
### Create Empty Collection/Folder
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"name":"测试",
"type":"virtual",
"metadata":{
"test":111
}
}'
```
* datasetId: Dataset ID (required)
* parentId: Parent ID. Defaults to root directory if not provided
* name: Collection name (required)
* type:
* folder: Folder
* virtual: Virtual collection (manual collection)
* metadata: Metadata (not currently used)
data is the collection ID.
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": "65abcd009d1448617cba5ee1"
}
```
### Create a Text Collection
Pass in text to create a collection. The text will be split accordingly.
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/text' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"text":"xxxxxxxx",
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"name":"测试训练",
"trainingType": "qa",
"chunkSettingMode": "auto",
"qaPrompt":"",
"metadata":{}
}'
```
* text: Original text
* datasetId: Dataset ID (required)
* parentId: Parent ID. Defaults to root directory if not provided
* name: Collection name (required)
* metadata: Metadata (not currently used)
data is the collection ID.
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abcfab9d1448617cba5f0d",
"results": {
"insertLen": 5, // Split into how many segments
"overToken": [],
"repeat": [],
"error": []
}
}
}
```
### Create a Link Collection
Pass in a web link to create a collection. Content will be fetched from the webpage first, then split.
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/link' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"link":"https://doc.fastgpt.io/guide/getting-started/quick-start",
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"trainingType": "chunk",
"chunkSettingMode": "auto",
"qaPrompt":"",
"metadata":{
"webPageSelector":".docs-content"
}
}'
```
* link: Web link
* datasetId: Dataset ID (required)
* parentId: Parent ID. Defaults to root directory if not provided
* metadata.webPageSelector: Web page selector to specify which element to use as text (optional)
data is the collection ID.
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abd0ad9d1448617cba6031",
"results": {
"insertLen": 1,
"overToken": [],
"repeat": [],
"error": []
}
}
}
```
### Create a File Collection
Pass in a file to create a collection. File content will be read and split. Currently supports: pdf, docx, md, txt, html, csv.
When uploading via code, note that Chinese filenames need to be encoded to avoid garbled text.
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/localFile' \
--header 'Authorization: Bearer {{authorization}}' \
--form 'file=@"C:\\Users\\user\\Desktop\\fastgpt测试File\\index.html"' \
--form 'data="{\"datasetId\":\"6593e137231a2be9c5603ba7\",\"parentId\":null,\"trainingType\":\"chunk\",\"chunkSize\":512,\"chunkSplitter\":\"\",\"qaPrompt\":\"\",\"metadata\":{}}"'
```
Use POST form-data format for upload. Contains file and data fields.
* file: File
* data: Dataset-related info (pass as serialized JSON). See "Common Creation Parameters" above
data is the collection ID.
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abc044e4704bac793fbd81",
"results": {
"insertLen": 1,
"overToken": [],
"repeat": [],
"error": []
}
}
}
```
### Create a Collection from an API Dataset (V1)
Pass in a file ID to create a collection. File content will be read and split. Currently supports: pdf, docx, md, txt, html, csv.
When uploading via code, note that Chinese filenames need to be encoded to avoid garbled text.
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/apiCollection' \
--header 'Authorization: Bearer fastgpt-xxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"name": "A Quick Guide to Building a Discord Bot.pdf",
"apiFileId":"A Quick Guide to Building a Discord Bot.pdf",
"datasetId": "674e9e479c3503c385495027",
"parentId": null,
"trainingType": "chunk",
"chunkSize":512,
"chunkSplitter":"",
"qaPrompt":""
}'
```
Use POST form-data format for upload. Contains file and data fields.
* name: Collection name, recommended to use filename, required.
* apiFileId: File ID, required.
* datasetId: Dataset ID (required)
* parentId: Parent ID. Defaults to root directory if not provided
* trainingType: Training mode (required)
* chunkSize: Length of each chunk (optional). chunk mode: 100~~3000; qa mode: 4000~~model max token (16k models usually recommended not to exceed 10000)
* chunkSplitter: Custom highest priority split symbol (optional)
* qaPrompt: QA split custom prompt (optional)
data is the collection ID.
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abc044e4704bac793fbd81",
"results": {
"insertLen": 1,
"overToken": [],
"repeat": [],
"error": []
}
}
}
```
### Create an External File Collection (Commercial)
```bash
curl --location --request POST 'http://localhost:3000/api/proApi/core/dataset/collection/create/externalFileUrl' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'User-Agent: Apifox/1.0.0 (https://apifox.com)' \
--header 'Content-Type: application/json' \
--data-raw '{
"externalFileUrl":"https://image.xxxxx.com/fastgpt-dev/%E6%91%82.pdf",
"externalFileId":"1111",
"createTime": "2024-05-01T00:00:00.000Z",
"filename":"自定义File名.pdf",
"datasetId":"6642d105a5e9d2b00255b27b",
"parentId": null,
"tags": ["tag1","tag2"],
"trainingType": "chunk",
"chunkSize":512,
"chunkSplitter":"",
"qaPrompt":""
}'
```
| Parameter | Description | Required |
| --------------- | ----------------------------------------------- | -------- |
| externalFileUrl | File access URL (can be temporary) | ✅ |
| externalFileId | External file ID | |
| filename | Custom filename with extension | |
| createTime | File creation time (Date or ISO string both ok) | |
data is the collection ID.
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "6646fcedfabd823cdc6de746",
"results": {
"insertLen": 1,
"overToken": [],
"repeat": [],
"error": []
}
}
}
```
### Get Collection List
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/listV2' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"offset":0,
"pageSize": 10,
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"searchText":""
}'
```
* offset: Offset
* pageSize: Items per page, max 30 (optional)
* datasetId: Dataset ID (required)
* parentId: Parent ID (optional)
* searchText: Fuzzy search text (optional)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"_id": "6593e137231a2be9c5603ba9",
"parentId": null,
"tmbId": "65422be6aa44b7da77729ec9",
"type": "virtual",
"name": "Manual entry",
"updateTime": "2099-01-01T00:00:00.000Z",
"dataAmount": 3,
"trainingAmount": 0,
"externalFileId": "1111",
"tags": ["11", "测试的"],
"forbid": false,
"trainingType": "chunk",
"permission": {
"value": 4294967295,
"isOwner": true,
"hasManagePer": true,
"hasWritePer": true,
"hasReadPer": true
}
},
{
"_id": "65abd0ad9d1448617cba6031",
"parentId": null,
"tmbId": "65422be6aa44b7da77729ec9",
"type": "link",
"name": "快速上手 | FastGPT",
"rawLink": "https://doc.fastgpt.io/guide/getting-started/quick-start",
"updateTime": "2024-01-20T13:54:53.031Z",
"dataAmount": 3,
"trainingAmount": 0,
"externalFileId": "222",
"tags": ["测试的"],
"forbid": false,
"trainingType": "chunk",
"permission": {
"value": 4294967295,
"isOwner": true,
"hasManagePer": true,
"hasWritePer": true,
"hasReadPer": true
}
}
],
"total": 93
}
}
```
### Get Collection Details
```bash
curl --location --request GET 'http://localhost:3000/api/core/dataset/collection/detail?id=65abcfab9d1448617cba5f0d' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: Collection ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"_id": "65abcfab9d1448617cba5f0d",
"parentId": null,
"teamId": "65422be6aa44b7da77729ec8",
"tmbId": "65422be6aa44b7da77729ec9",
"datasetId": {
"_id": "6593e137231a2be9c5603ba7",
"parentId": null,
"teamId": "65422be6aa44b7da77729ec8",
"tmbId": "65422be6aa44b7da77729ec9",
"type": "dataset",
"status": "active",
"avatar": "/icon/logo.svg",
"name": "FastGPT test",
"vectorModel": "text-embedding-ada-002",
"agentModel": "gpt-3.5-turbo-16k",
"intro": "",
"permission": "private",
"updateTime": "2024-01-02T10:11:03.084Z"
},
"type": "virtual",
"name": "测试训练",
"trainingType": "qa",
"chunkSize": 8000,
"chunkSplitter": "",
"qaPrompt": "11",
"rawTextLength": 40466,
"hashRawText": "47270840614c0cc122b29daaddc09c2a48f0ec6e77093611ab12b69cba7fee12",
"createTime": "2024-01-20T13:50:35.838Z",
"updateTime": "2024-01-20T13:50:35.838Z",
"canWrite": true,
"sourceName": "测试训练"
}
}
```
### Update Dataset Collection Info
**Update Collection Info by Collection ID**
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/update' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"id":"65abcfab9d1448617cba5f0d",
"parentId": null,
"name": "测2222试",
"tags": ["tag1", "tag2"],
"forbid": false,
"createTime": "2024-01-01T00:00:00.000Z"
}'
```
**Update Collection Info by External File ID**, Just replace id with datasetId and externalFileId.
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/update' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"datasetId":"6593e137231a2be9c5603ba7",
"externalFileId":"1111",
"parentId": null,
"name": "测2222试",
"tags": ["tag1", "tag2"],
"forbid": false,
"createTime": "2024-01-01T00:00:00.000Z"
}'
```
* id: Collection ID
* parentId: Update parent ID (optional)
* name: Update collection name (optional)
* tags: Update collection tags (optional)
* forbid: Update collection disabled status (optional)
* createTime: Update collection creation time (optional)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### Delete Collection
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/delete' \
--header 'Authorization: Bearer fastgpt-' \
--header 'Content-Type: application/json' \
--data-raw '{
"collectionIds": ["65a8cdcb0d70d3de0bf08d0a"]
}'
```
* collectionIds: Collection ID list
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## Data
### Data Structure
**Data Structure**
| Field | Type | Description | Required |
| ------------- | -------- | -------------- | -------- |
| teamId | String | Team ID | ✅ |
| tmbId | String | Member ID | ✅ |
| datasetId | String | Dataset ID | ✅ |
| collectionId | String | CollectionID | ✅ |
| q | String | Primary data | ✅ |
| a | String | Auxiliary data | ✖ |
| fullTextToken | String | Tokenization | ✖ |
| indexes | Index\[] | Vector indexes | ✅ |
| updateTime | Date | Update time | ✅ |
| chunkIndex | Number | Chunk index | ✖ |
**Index Structure**
Maximum 5 custom indexes per data group
| Field | Type | Description | Required |
| ------ | ------ | ----------------------------------------------------------------------------------------------------------------------------------- | -------- |
| type | String | Optional index types: default-default index; custom-custom index; summary-summary index; question-question index; image-image index | |
| dataId | String | Associated vector ID. Pass this ID when updating data for incremental updates instead of full updates | |
| text | String | Text content | ✅ |
`type` If not provided, defaults to `custom` index. A default index will also be created based on q/a. If a default index is provided, no additional one will be created.
### Push Data to Training Queue
Each request can push up to 200 data groups. FastGPT automatically creates the training usage record, so you do not need to provide `billId`.
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/data/pushData' \
--header 'Authorization: Bearer apikey' \
--header 'Content-Type: application/json' \
--data-raw '{
"collectionId": "64663f451ba1676dbdef0499",
"trainingType": "chunk",
"prompt": "Optional. QA split guide prompt, ignored in chunk mode",
"data": [
{
"q": "Who are you?",
"a": "I'm FastGPT Assistant"
},
{
"q": "What can you do?",
"a": "I can do anything",
"indexes": [
{
"text":"Custom index 1"
},
{
"text":"Custom index 2"
}
]
}
]
}'
```
* collectionId: Collection ID (required)
* trainingType: Training mode (required)
* prompt: Custom QA split prompt. Must follow template strictly. Recommended not to pass. (optional)
* data:(Specific data)
* q: Primary data(Required)
* a: Auxiliary data (optional)
* indexes: Custom indexes (optional). Can omit or pass empty array. By default, an index will be created from q and a.
```json
{
"code": 200,
"statusText": "",
"data": {
"insertLen": 1, // Final number of successful insertions
"overToken": [], // Exceeding token
"repeat": [], // Number of duplicates
"error": [] // Other errors
}
}
```
\[theme] content can be replaced with the data theme. Default: They may contain multiple theme contents
```
I'll give you a text, [theme], learn it, and organize the learning results, requirements:
1. Propose up to 25 questions.
2. Provide answers to each question.
3. Answers should be detailed and complete, and can include plain text, links, code, tables, formulas, media links, and other markdown elements.
4. Return multiple questions and answers in format:
Q1: Question.
A1: Answer.
Q2:
A2:
……
My text:"""{{text}}"""
```
### Get Collection Data List
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/data/v2/list' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"offset": 0,
"pageSize": 10,
"collectionId":"65abd4ac9d1448617cba6171",
"searchText":""
}'
```
* offset: Offset (optional)
* pageSize: Items per page, max 30 (optional)
* collectionId: Collection ID (required)
* searchText: Fuzzy search term (optional)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"_id": "65abd4b29d1448617cba61db",
"datasetId": "65abc9bd9d1448617cba5e6c",
"collectionId": "65abd4ac9d1448617cba6171",
"q": "N o . 2 0 2 2 1 2中 国 信 息 通 信 研 究 院京东探索研究院2022年 9月人工智能生成内容(AIGC)白皮书(2022 年)版权声明本白皮书版权属于中国信息通信研究院和京东探索研究院,并受法律保护。转载、摘编或利用其它方式使用本白皮书文字or观点的,应注明“来源:中国信息通信研究院和京东探索研究院”。违反上述声明者,编者将追究其相关法律责任。前 言习近平总书记曾指出,“数字技术正以新理念、新业态、新模式全面融入人类经济、政治、文化、社会、生态文明建设各领域和全过程”。在当前数字世界和物理世界加速融合的大背景下,人工智能生成内容(Artificial Intelligence Generated Content,简称 AIGC)正在悄然引导着一场深刻的变革,重塑甚至颠覆数字内容的生产方式和消费模式,将极大地丰富人们的数字生活,是未来全面迈向数字文明新时代不可或缺的支撑力量。",
"a": "",
"chunkIndex": 0
},
{
"_id": "65abd4b39d1448617cba624d",
"datasetId": "65abc9bd9d1448617cba5e6c",
"collectionId": "65abd4ac9d1448617cba6171",
"q": "本白皮书重点从 AIGC 技术、应用和治理等维度进行了阐述。在技术层面,梳理提出了 AIGC 技术体系,既涵盖了对现实世界各种内容的数字化呈现和增强,也包括了基于人工智能的自主内容创作。在应用层面,重点分析了 AIGC 在传媒、电商、影视等行业和场景的应用情况,探讨了以虚拟数字人、写作机器人等为代表的新业态和新应用。在治理层面,从政策监管、技术能力、企业应用等视角,分析了AIGC 所暴露出的版权纠纷、虚假信息传播等各种Question.最后,从政府、行业、企业、社会等层面,给出了 AIGC 发展和治理建议。由于人工智能仍处于飞速发展阶段,我们对 AIGC 的认识还有待进一步深化,白皮书中存在不足之处,敬请大家批评指正。目 录一、 人工智能生成内容的发展历程与概念.............................................................. 1(一)AIGC 历史沿革 .......................................................................................... 1(二)AIGC 的概念与内涵 .................................................................................. 4二、人工智能生成内容的技术体系及其演进方向.................................................... 7(一)AIGC 技术升级步入深化阶段 .................................................................. 7(二)AIGC 大模型架构潜力凸显 .................................................................... 10(三)AIGC 技术演化出三大前沿能力 ............................................................ 18三、人工智能生成内容的应用场景.......................................................................... 26(一)AIGC+传媒:人机协同生产,",
"a": "",
"chunkIndex": 1
}
],
"total": 63
}
}
```
### Get Single Data Details
```bash
curl --location --request GET 'http://localhost:3000/api/core/dataset/data/detail?id=65abd4b29d1448617cba61db' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: Data ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"id": "65abd4b29d1448617cba61db",
"q": "N o . 2 0 2 2 1 2中 国 信 息 通 信 研 究 院京东探索研究院2022年 9月人工智能生成内容(AIGC)白皮书(2022 年)版权声明本白皮书版权属于中国信息通信研究院和京东探索研究院,并受法律保护。转载、摘编或利用其它方式使用本白皮书文字or观点的,应注明“来源:中国信息通信研究院和京东探索研究院”。违反上述声明者,编者将追究其相关法律责任。前 言习近平总书记曾指出,“数字技术正以新理念、新业态、新模式全面融入人类经济、政治、文化、社会、生态文明建设各领域和全过程”。在当前数字世界和物理世界加速融合的大背景下,人工智能生成内容(Artificial Intelligence Generated Content,简称 AIGC)正在悄然引导着一场深刻的变革,重塑甚至颠覆数字内容的生产方式和消费模式,将极大地丰富人们的数字生活,是未来全面迈向数字文明新时代不可或缺的支撑力量。",
"a": "",
"chunkIndex": 0,
"indexes": [
{
"type": "default",
"dataId": "3720083",
"text": "N o . 2 0 2 2 1 2中 国 信 息 通 信 研 究 院京东探索研究院2022年 9月人工智能生成内容(AIGC)白皮书(2022 年)版权声明本白皮书版权属于中国信息通信研究院和京东探索研究院,并受法律保护。转载、摘编或利用其它方式使用本白皮书文字or观点的,应注明“来源:中国信息通信研究院和京东探索研究院”。违反上述声明者,编者将追究其相关法律责任。前 言习近平总书记曾指出,“数字技术正以新理念、新业态、新模式全面融入人类经济、政治、文化、社会、生态文明建设各领域和全过程”。在当前数字世界和物理世界加速融合的大背景下,人工智能生成内容(Artificial Intelligence Generated Content,简称 AIGC)正在悄然引导着一场深刻的变革,重塑甚至颠覆数字内容的生产方式和消费模式,将极大地丰富人们的数字生活,是未来全面迈向数字文明新时代不可或缺的支撑力量。",
"_id": "65abd4b29d1448617cba61dc"
}
],
"datasetId": "65abc9bd9d1448617cba5e6c",
"collectionId": "65abd4ac9d1448617cba6171",
"sourceName": "中文-AIGC白皮书2022.pdf",
"sourceId": "65abd4ac9d1448617cba6166",
"isOwner": true,
"canWrite": true
}
}
```
### Update Single Data
```bash
curl --location --request PUT 'http://localhost:3000/api/core/dataset/data/update' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"dataId":"65abd4b29d1448617cba61db",
"q":"Test 111",
"a":"sss",
"indexes":[
{
"dataId": "xxxx",
"type": "default",
"text": "Default index"
},
{
"dataId": "xxx",
"type": "custom",
"text": "旧的Custom index 1"
},
{
"type":"custom",
"text":"New custom index"
}
]
}'
```
* dataId: Data ID
* q: Primary data (optional)
* a: Auxiliary data (optional)
* indexes: Custom indexes (optional). See `Batch Add Data to Collection` for types. If custom indexes exist when created,
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### Delete Single Data
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/dataset/data/delete?id=65abd4b39d1448617cba624d' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: Data ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": "success"
}
```
## Search Test
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/searchTest' \
--header 'Authorization: Bearer fastgpt-xxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"datasetId": "Dataset ID",
"text": "Who is the director",
"limit": 5000,
"similarity": 0,
"searchMode": "embedding",
"usingReRank": false,
"datasetSearchUsingExtensionQuery": true,
"datasetSearchExtensionModel": "gpt-5",
"datasetSearchExtensionBg": ""
}'
```
* datasetId - Dataset ID
* text - Text to test
* limit - Maximum tokens
* similarity - Minimum similarity (0\~1, optional)
* searchMode - Search mode: embedding | fullTextRecall | mixedRecall
* usingReRank - Use rerank
* datasetSearchUsingExtensionQuery - Use query extension
* datasetSearchExtensionModel - Query extension model
* datasetSearchExtensionBg - Query extension background description
Returns top k results. limit is the maximum tokens, up to 20000 tokens.
```json
{
"code": 200,
"statusText": "",
"data": [
{
"id": "65599c54a5c814fb803363cb",
"q": "你是谁",
"a": "I'm FastGPT Assistant",
"datasetId": "6554684f7f9ed18a39a4d15c",
"collectionId": "6556cd795e4b663e770bb66d",
"sourceName": "GBT 15104-2021 装饰单板贴面人造板.pdf",
"sourceId": "6556cd775e4b663e770bb65c",
"score": 0.8050316572189331
},
......
]
}
```
file: ./content/openapi/dataset.mdx
meta: {
"title": "知识库接口",
"description": "FastGPT OpenAPI 知识库接口"
}
| 如何获取知识库 ID(datasetId) | 如何获取文件集合 ID(collection\_id) |
| --------------------------------------- | -------------------------------------- |
|  |  |
## 知识库
### 创建知识库
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/create' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"parentId": null,
"type": "dataset",
"name":"测试",
"intro":"介绍",
"avatar": "",
"vectorModel": "text-embedding-ada-002",
"agentModel": "gpt-3.5-turbo-16k",
"vlmModel": "gpt-4.1"
}'
```
* parentId - 父级 ID,用于构建目录结构。通常可以为 null 或者直接不传。
* type - `dataset` 或者 `folder`,代表普通知识库和文件夹。不传则代表创建普通知识库。
* name - 知识库名(必填)
* intro - 介绍(可选)
* avatar - 头像地址(可选)
* vectorModel - 向量模型(建议传空,用系统默认的)
* agentModel - 文本处理模型(建议传空,用系统默认的)
* vlmModel - 图片理解模型(建议传空,用系统默认的)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": "65abc9bd9d1448617cba5e6c"
}
```
### 获取知识库列表
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/list?parentId=' \
--header 'Authorization: Bearer xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"parentId":""
}'
```
* parentId - 父级 ID,传空字符串或者 null,代表获取根目录下的知识库
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": [
{
"_id": "65abc9bd9d1448617cba5e6c",
"parentId": null,
"avatar": "",
"name": "测试",
"intro": "",
"type": "dataset",
"permission": "private",
"canWrite": true,
"isOwner": true,
"vectorModel": {
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"charsPointsPrice": 0,
"defaultToken": 512,
"maxToken": 8000,
"weight": 100
}
}
]
}
```
### 获取知识库详情
```bash
curl --location --request GET 'http://localhost:3000/api/core/dataset/detail?id=6593e137231a2be9c5603ba7' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: 知识库的 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"_id": "6593e137231a2be9c5603ba7",
"parentId": null,
"teamId": "65422be6aa44b7da77729ec8",
"tmbId": "65422be6aa44b7da77729ec9",
"type": "dataset",
"status": "active",
"avatar": "/icon/logo.svg",
"name": "FastGPT test",
"vectorModel": {
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"charsPointsPrice": 0,
"defaultToken": 512,
"maxToken": 8000,
"weight": 100
},
"agentModel": {
"model": "gpt-3.5-turbo-16k",
"name": "FastAI-16k",
"maxContext": 16000,
"maxResponse": 16000,
"charsPointsPrice": 0
},
"intro": "",
"permission": "private",
"updateTime": "2024-01-02T10:11:03.084Z",
"canWrite": true,
"isOwner": true
}
}
```
### 删除知识库
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/dataset/delete?id=65abc8729d1448617cba5df6' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: 知识库的 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## 集合
### 通用创建参数说明(必看)
**入参**
| 参数 | 说明 | 必填 |
| ---------------- | ----------------------------------------------------------------- | -- |
| datasetId | 知识库 ID | ✅ |
| parentId | 父级 ID,不填则默认为根目录 | |
| trainingType | 数据处理方式。chunk: 按文本长度进行分割;qa: 问答对提取 | ✅ |
| indexPrefixTitle | 是否自动生成标题索引 | |
| customPdfParse | 是否开启 PDF 增强解析, 默认 false: 关闭;true: 开启; | |
| autoIndexes | 是否自动生成索引(仅商业版支持) | |
| imageIndex | 是否自动生成图片索引(仅商业版支持) | |
| chunkSettingMode | 分块参数模式。auto: 系统默认参数; custom: 手动指定参数 | |
| chunkSplitMode | 分块拆分模式。size: 按长度拆分; char: 按字符拆分。chunkSettingMode=auto 时不生效。 | |
| chunkSize | 分块大小,默认 1500。chunkSettingMode=auto 时不生效。 | |
| indexSize | 索引大小,默认 512,必须小于索引模型最大 token。chunkSettingMode=auto 时不生效。 | |
| chunkSplitter | 自定义最高优先分割符号,除非超出文件处理最大上下文,否则不会进行进一步拆分。chunkSettingMode=auto 时不生效。 | |
| qaPrompt | qa 拆分提示词 | |
| tags | 集合标签(字符串数组) | |
| createTime | 文件创建时间(Date / String) | |
**出参**
* collectionId - 新建的集合 ID
* insertLen:插入的块数量
### 创建空集合/目录
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"name":"测试",
"type":"virtual",
"metadata":{
"test":111
}
}'
```
* datasetId: 知识库的 ID(必填)
* parentId:父级 ID,不填则默认为根目录
* name: 集合名称(必填)
* type:
* folder:文件夹
* virtual:虚拟集合(手动集合)
* metadata:元数据(暂时没啥用)
data 为集合的 ID。
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": "65abcd009d1448617cba5ee1"
}
```
### 创建一个纯文本集合
传入一段文字,创建一个集合,会根据传入的文字进行分割。
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/text' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"text":"xxxxxxxx",
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"name":"测试训练",
"trainingType": "qa",
"chunkSettingMode": "auto",
"qaPrompt":"",
"metadata":{}
}'
```
* text: 原文本
* datasetId: 知识库的 ID(必填)
* parentId:父级 ID,不填则默认为根目录
* name: 集合名称(必填)
* metadata:元数据(暂时没啥用)
data 为集合的 ID。
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abcfab9d1448617cba5f0d",
"results": {
"insertLen": 5
}
}
}
```
### 创建一个链接集合
传入一个网络链接,创建一个集合,会先去对应网页抓取内容,再抓取的文字进行分割。
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/link' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"link":"https://doc.fastgpt.io/guide/getting-started/quick-start",
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"trainingType": "chunk",
"chunkSettingMode": "auto",
"qaPrompt":"",
"metadata":{
"webPageSelector":".docs-content"
}
}'
```
* link: 网络链接
* datasetId: 知识库的 ID(必填)
* parentId:父级 ID,不填则默认为根目录
* metadata.webPageSelector: 网页选择器,用于指定网页中的哪个元素作为文本(可选)
data 为集合的 ID。
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abd0ad9d1448617cba6031",
"results": {
"insertLen": 1
}
}
}
```
### 创建一个文件集合
传入一个文件,创建一个集合,会读取文件内容进行分割。目前支持:pdf, docx, md, txt, html, csv。
使用代码上传时,请注意中文 filename 需要进行 encode 处理,否则容易乱码。
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/localFile' \
--header 'Authorization: Bearer {{authorization}}' \
--form 'file=@"C:\\Users\\user\\Desktop\\fastgpt测试文件\\index.html"' \
--form 'data="{\"datasetId\":\"6593e137231a2be9c5603ba7\",\"parentId\":null,\"trainingType\":\"chunk\",\"chunkSize\":512,\"chunkSplitter\":\"\",\"qaPrompt\":\"\",\"metadata\":{}}"'
```
需要使用 POST form-data 的格式上传。包含 file 和 data 两个字段。
* file: 文件
* data: 知识库相关信息(json 序列化后传入),参数说明见上方"通用创建参数说明"
data 为集合的 ID。
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abc044e4704bac793fbd81",
"results": {
"insertLen": 1
}
}
}
```
### 通过 API 数据集创建集合(V1)
传入一个文件的 id,创建一个集合,会读取文件内容进行分割。目前支持:pdf, docx, md, txt, html, csv。
使用代码上传时,请注意中文 filename 需要进行 encode 处理,否则容易乱码。
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/create/apiCollection' \
--header 'Authorization: Bearer fastgpt-xxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"name": "A Quick Guide to Building a Discord Bot.pdf",
"apiFileId":"A Quick Guide to Building a Discord Bot.pdf",
"datasetId": "674e9e479c3503c385495027",
"parentId": null,
"trainingType": "chunk",
"chunkSize":512,
"chunkSplitter":"",
"qaPrompt":""
}'
```
需要使用 POST form-data 的格式上传。包含 file 和 data 两个字段。
* name: 集合名,建议就用文件名,必填。
* apiFileId: 文件的 ID,必填。
* datasetId: 知识库的 ID(必填)
* parentId:父级 ID,不填则默认为根目录
* trainingType:训练模式(必填)
* chunkSize: 每个 chunk 的长度(可选). chunk 模式:100~~3000; qa 模式: 4000~~ 模型最大 token(16k 模型通常建议不超过 10000)
* chunkSplitter: 自定义最高优先分割符号(可选)
* qaPrompt: qa 拆分自定义提示词(可选)
data 为集合的 ID。
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "65abc044e4704bac793fbd81",
"results": {
"insertLen": 1
}
}
}
```
### 创建一个外部文件库集合(商业版)
```bash
curl --location --request POST 'http://localhost:3000/api/proApi/core/dataset/collection/create/externalFileUrl' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'User-Agent: Apifox/1.0.0 (https://apifox.com)' \
--header 'Content-Type: application/json' \
--data-raw '{
"externalFileUrl":"https://image.xxxxx.com/fastgpt-dev/%E6%91%82.pdf",
"externalFileId":"1111",
"createTime": "2024-05-01T00:00:00.000Z",
"filename":"自定义文件名.pdf",
"datasetId":"6642d105a5e9d2b00255b27b",
"parentId": null,
"tags": ["tag1","tag2"],
"trainingType": "chunk",
"chunkSize":512,
"chunkSplitter":"",
"qaPrompt":""
}'
```
| 参数 | 说明 | 必填 |
| --------------- | ------------------------ | -- |
| externalFileUrl | 文件访问链接(可以是临时链接) | ✅ |
| externalFileId | 外部文件 ID | |
| filename | 自定义文件名,需要带后缀 | |
| createTime | 文件创建时间(Date ISO 字符串都 ok) | |
data 为集合的 ID。
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"collectionId": "6646fcedfabd823cdc6de746",
"results": {
"insertLen": 1
}
}
}
```
### 获取集合列表
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/listV2' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"offset":0,
"pageSize": 10,
"datasetId":"6593e137231a2be9c5603ba7",
"parentId": null,
"searchText":""
}'
```
* offset: 偏移量
* pageSize: 每页数量,最大 30(选填)
* datasetId: 知识库的 ID(必填)
* parentId: 父级 ID(选填)
* searchText: 模糊搜索文本(选填)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"_id": "6593e137231a2be9c5603ba9",
"parentId": null,
"tmbId": "65422be6aa44b7da77729ec9",
"type": "virtual",
"name": "手动录入",
"updateTime": "2099-01-01T00:00:00.000Z",
"dataAmount": 3,
"trainingAmount": 0,
"externalFileId": "1111",
"tags": ["11", "测试的"],
"forbid": false,
"trainingType": "chunk",
"permission": {
"value": 4294967295,
"isOwner": true,
"hasManagePer": true,
"hasWritePer": true,
"hasReadPer": true
}
},
{
"_id": "65abd0ad9d1448617cba6031",
"parentId": null,
"tmbId": "65422be6aa44b7da77729ec9",
"type": "link",
"name": "快速上手 | FastGPT",
"rawLink": "https://doc.fastgpt.io/guide/getting-started/quick-start",
"updateTime": "2024-01-20T13:54:53.031Z",
"dataAmount": 3,
"trainingAmount": 0,
"externalFileId": "222",
"tags": ["测试的"],
"forbid": false,
"trainingType": "chunk",
"permission": {
"value": 4294967295,
"isOwner": true,
"hasManagePer": true,
"hasWritePer": true,
"hasReadPer": true
}
}
],
"total": 93
}
}
```
### 获取集合详情
```bash
curl --location --request GET 'http://localhost:3000/api/core/dataset/collection/detail?id=65abcfab9d1448617cba5f0d' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: 集合的 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"_id": "65abcfab9d1448617cba5f0d",
"parentId": null,
"teamId": "65422be6aa44b7da77729ec8",
"tmbId": "65422be6aa44b7da77729ec9",
"datasetId": {
"_id": "6593e137231a2be9c5603ba7",
"parentId": null,
"teamId": "65422be6aa44b7da77729ec8",
"tmbId": "65422be6aa44b7da77729ec9",
"type": "dataset",
"status": "active",
"avatar": "/icon/logo.svg",
"name": "FastGPT test",
"vectorModel": "text-embedding-ada-002",
"agentModel": "gpt-3.5-turbo-16k",
"intro": "",
"permission": "private",
"updateTime": "2024-01-02T10:11:03.084Z"
},
"type": "virtual",
"name": "测试训练",
"trainingType": "qa",
"chunkSize": 8000,
"chunkSplitter": "",
"qaPrompt": "11",
"rawTextLength": 40466,
"hashRawText": "47270840614c0cc122b29daaddc09c2a48f0ec6e77093611ab12b69cba7fee12",
"createTime": "2024-01-20T13:50:35.838Z",
"updateTime": "2024-01-20T13:50:35.838Z",
"canWrite": true,
"sourceName": "测试训练"
}
}
```
### 更新数据集集合信息
**通过集合 ID 修改集合信息**
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/update' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"id":"65abcfab9d1448617cba5f0d",
"parentId": null,
"name": "测2222试",
"tags": ["tag1", "tag2"],
"forbid": false,
"createTime": "2024-01-01T00:00:00.000Z"
}'
```
**通过外部文件 ID 修改集合信息**,只需要把 ID 换成 datasetId 和 externalFileId。
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/update' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"datasetId":"6593e137231a2be9c5603ba7",
"externalFileId":"1111",
"parentId": null,
"name": "测2222试",
"tags": ["tag1", "tag2"],
"forbid": false,
"createTime": "2024-01-01T00:00:00.000Z"
}'
```
* id: 集合的 ID
* parentId: 修改父级 ID(可选)
* name: 修改集合名称(可选)
* tags: 修改集合标签(可选)
* forbid: 修改集合禁用状态(可选)
* createTime: 修改集合创建时间(可选)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### 删除集合
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/collection/delete' \
--header 'Authorization: Bearer fastgpt-' \
--header 'Content-Type: application/json' \
--data-raw '{
"collectionIds": ["65a8cdcb0d70d3de0bf08d0a"]
}'
```
* collectionIds: 集合的 ID 列表
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
## 数据
### 数据的结构
**Data 结构**
| 字段 | 类型 | 说明 | 必填 |
| ------------- | -------- | ------ | -- |
| teamId | String | 团队 ID | ✅ |
| tmbId | String | 成员 ID | ✅ |
| datasetId | String | 知识库 ID | ✅ |
| collectionId | String | 集合 ID | ✅ |
| q | String | 主要数据 | ✅ |
| a | String | 辅助数据 | ✖ |
| fullTextToken | String | 分词 | ✖ |
| indexes | Index\[] | 向量索引 | ✅ |
| updateTime | Date | 更新时间 | ✅ |
| chunkIndex | Number | 分块下表 | ✖ |
**Index 结构**
每组数据的自定义索引最多 5 个
| 字段 | 类型 | 说明 | 必填 |
| ------ | ------ | ------------------------------------------------------------------------------- | -- |
| type | String | 可选索引类型:default- 默认索引; custom- 自定义索引; summary- 总结索引; question- 问题索引; image- 图片索引 | |
| dataId | String | 关联的向量 ID,变更数据时候传入该 ID,会进行差量更新,而不是全量更新 | |
| text | String | 文本内容 | ✅ |
`type` 不填则默认为 `custom` 索引,还会基于 q/a 组成一个默认索引。如果传入了默认索引,则不会额外创建。
### 推送数据到训练队列
注意,每次最多推送 200 组数据。训练账单会在接口调用时自动创建,无需手动传入 `billId`。
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/data/pushData' \
--header 'Authorization: Bearer apikey' \
--header 'Content-Type: application/json' \
--data-raw '{
"collectionId": "64663f451ba1676dbdef0499",
"trainingType": "chunk",
"prompt": "可选。qa 拆分引导词,chunk 模式下忽略",
"data": [
{
"q": "你是谁?",
"a": "我是FastGPT助手"
},
{
"q": "你会什么?",
"a": "我什么都会",
"indexes": [
{
"text":"自定义索引1"
},
{
"text":"自定义索引2"
}
]
}
]
}'
```
* collectionId: 集合 ID(必填)
* trainingType:训练模式(必填)
* prompt: 自定义 QA 拆分提示词,需严格按照模板,建议不要传入。(选填)
* data:(具体数据)
* q: 主要数据(必填)
* a: 辅助数据(选填)
* indexes: 自定义索引(选填)。可以不传或者传空数组,默认都会使用 q 和 a 组成一个索引。
```json
{
"code": 200,
"statusText": "",
"data": {
"insertLen": 1 // 最终插入成功的数量
}
}
```
\[theme] 里的内容可以换成数据的主题。默认为:它们可能包含多个主题内容
```
我会给你一段文本,[theme],学习它们,并整理学习成果,要求为:
1. 提出最多 25 个问题。
2. 给出每个问题的答案。
3. 答案要详细完整,答案可以包含普通文字、链接、代码、表格、公示、媒体链接等 markdown 元素。
4. 按格式返回多个问题和答案:
Q1: 问题。
A1: 答案。
Q2:
A2:
……
我的文本:"""{{text}}"""
```
### 获取集合的数据列表
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/data/v2/list' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"offset": 0,
"pageSize": 10,
"collectionId":"65abd4ac9d1448617cba6171",
"searchText":""
}'
```
* offset: 偏移量(选填)
* pageSize: 每页数量,最大 30(选填)
* collectionId: 集合的 ID(必填)
* searchText: 模糊搜索词(选填)
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"list": [
{
"_id": "65abd4b29d1448617cba61db",
"datasetId": "65abc9bd9d1448617cba5e6c",
"collectionId": "65abd4ac9d1448617cba6171",
"q": "N o . 2 0 2 2 1 2中 国 信 息 通 信 研 究 院京东探索研究院2022年 9月人工智能生成内容(AIGC)白皮书(2022 年)版权声明本白皮书版权属于中国信息通信研究院和京东探索研究院,并受法律保护。转载、摘编或利用其它方式使用本白皮书文字或者观点的,应注明“来源:中国信息通信研究院和京东探索研究院”。违反上述声明者,编者将追究其相关法律责任。前 言习近平总书记曾指出,“数字技术正以新理念、新业态、新模式全面融入人类经济、政治、文化、社会、生态文明建设各领域和全过程”。在当前数字世界和物理世界加速融合的大背景下,人工智能生成内容(Artificial Intelligence Generated Content,简称 AIGC)正在悄然引导着一场深刻的变革,重塑甚至颠覆数字内容的生产方式和消费模式,将极大地丰富人们的数字生活,是未来全面迈向数字文明新时代不可或缺的支撑力量。",
"a": "",
"chunkIndex": 0
},
{
"_id": "65abd4b39d1448617cba624d",
"datasetId": "65abc9bd9d1448617cba5e6c",
"collectionId": "65abd4ac9d1448617cba6171",
"q": "本白皮书重点从 AIGC 技术、应用和治理等维度进行了阐述。在技术层面,梳理提出了 AIGC 技术体系,既涵盖了对现实世界各种内容的数字化呈现和增强,也包括了基于人工智能的自主内容创作。在应用层面,重点分析了 AIGC 在传媒、电商、影视等行业和场景的应用情况,探讨了以虚拟数字人、写作机器人等为代表的新业态和新应用。在治理层面,从政策监管、技术能力、企业应用等视角,分析了AIGC 所暴露出的版权纠纷、虚假信息传播等各种问题。最后,从政府、行业、企业、社会等层面,给出了 AIGC 发展和治理建议。由于人工智能仍处于飞速发展阶段,我们对 AIGC 的认识还有待进一步深化,白皮书中存在不足之处,敬请大家批评指正。目 录一、 人工智能生成内容的发展历程与概念.............................................................. 1(一)AIGC 历史沿革 .......................................................................................... 1(二)AIGC 的概念与内涵 .................................................................................. 4二、人工智能生成内容的技术体系及其演进方向.................................................... 7(一)AIGC 技术升级步入深化阶段 .................................................................. 7(二)AIGC 大模型架构潜力凸显 .................................................................... 10(三)AIGC 技术演化出三大前沿能力 ............................................................ 18三、人工智能生成内容的应用场景.......................................................................... 26(一)AIGC+传媒:人机协同生产,",
"a": "",
"chunkIndex": 1
}
],
"total": 63
}
}
```
### 获取单条数据详情
```bash
curl --location --request GET 'http://localhost:3000/api/core/dataset/data/detail?id=65abd4b29d1448617cba61db' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: 数据的 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": {
"id": "65abd4b29d1448617cba61db",
"q": "N o . 2 0 2 2 1 2中 国 信 息 通 信 研 究 院京东探索研究院2022年 9月人工智能生成内容(AIGC)白皮书(2022 年)版权声明本白皮书版权属于中国信息通信研究院和京东探索研究院,并受法律保护。转载、摘编或利用其它方式使用本白皮书文字或者观点的,应注明“来源:中国信息通信研究院和京东探索研究院”。违反上述声明者,编者将追究其相关法律责任。前 言习近平总书记曾指出,“数字技术正以新理念、新业态、新模式全面融入人类经济、政治、文化、社会、生态文明建设各领域和全过程”。在当前数字世界和物理世界加速融合的大背景下,人工智能生成内容(Artificial Intelligence Generated Content,简称 AIGC)正在悄然引导着一场深刻的变革,重塑甚至颠覆数字内容的生产方式和消费模式,将极大地丰富人们的数字生活,是未来全面迈向数字文明新时代不可或缺的支撑力量。",
"a": "",
"chunkIndex": 0,
"indexes": [
{
"type": "default",
"dataId": "3720083",
"text": "N o . 2 0 2 2 1 2中 国 信 息 通 信 研 究 院京东探索研究院2022年 9月人工智能生成内容(AIGC)白皮书(2022 年)版权声明本白皮书版权属于中国信息通信研究院和京东探索研究院,并受法律保护。转载、摘编或利用其它方式使用本白皮书文字或者观点的,应注明“来源:中国信息通信研究院和京东探索研究院”。违反上述声明者,编者将追究其相关法律责任。前 言习近平总书记曾指出,“数字技术正以新理念、新业态、新模式全面融入人类经济、政治、文化、社会、生态文明建设各领域和全过程”。在当前数字世界和物理世界加速融合的大背景下,人工智能生成内容(Artificial Intelligence Generated Content,简称 AIGC)正在悄然引导着一场深刻的变革,重塑甚至颠覆数字内容的生产方式和消费模式,将极大地丰富人们的数字生活,是未来全面迈向数字文明新时代不可或缺的支撑力量。",
"_id": "65abd4b29d1448617cba61dc"
}
],
"datasetId": "65abc9bd9d1448617cba5e6c",
"collectionId": "65abd4ac9d1448617cba6171",
"sourceName": "中文-AIGC白皮书2022.pdf",
"sourceId": "65abd4ac9d1448617cba6166",
"isOwner": true,
"canWrite": true
}
}
```
### 修改单条数据
```bash
curl --location --request PUT 'http://localhost:3000/api/core/dataset/data/update' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"dataId":"65abd4b29d1448617cba61db",
"q":"测试111",
"a":"sss",
"indexes":[
{
"dataId": "xxxx",
"type": "default",
"text": "默认索引"
},
{
"dataId": "xxx",
"type": "custom",
"text": "旧的自定义索引1"
},
{
"type":"custom",
"text":"新增的自定义索引"
}
]
}'
```
* dataId: 数据的 ID
* q: 主要数据(选填)
* a: 辅助数据(选填)
* indexes: 自定义索引(选填),类型参考 `为集合批量添加添加数据`。如果创建时候有自定义索引,
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": null
}
```
### 删除单条数据
```bash
curl --location --request DELETE 'http://localhost:3000/api/core/dataset/data/delete?id=65abd4b39d1448617cba624d' \
--header 'Authorization: Bearer {{authorization}}' \
```
* id: 数据的 ID
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": "success"
}
```
## 搜索测试
```bash
curl --location --request POST 'http://localhost:3000/api/core/dataset/searchTest' \
--header 'Authorization: Bearer fastgpt-xxxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"datasetId": "知识库的ID",
"text": "导演是谁",
"limit": 5000,
"similarity": 0,
"searchMode": "embedding",
"usingReRank": false,
"datasetSearchUsingExtensionQuery": true,
"datasetSearchExtensionModel": "gpt-5",
"datasetSearchExtensionBg": ""
}'
```
* datasetId - 知识库 ID
* text - 需要测试的文本
* limit - 最大 tokens 数量
* similarity - 最低相关度(0\~1,可选)
* searchMode - 搜索模式:embedding | fullTextRecall | mixedRecall
* usingReRank - 使用重排
* datasetSearchUsingExtensionQuery - 使用问题优化
* datasetSearchExtensionModel - 问题优化模型
* datasetSearchExtensionBg - 问题优化背景描述
返回 top k 结果,limit 为最大 Tokens 数量,最多 20000 tokens。
```json
{
"code": 200,
"statusText": "",
"data": [
{
"id": "65599c54a5c814fb803363cb",
"q": "你是谁",
"a": "我是FastGPT助手",
"datasetId": "6554684f7f9ed18a39a4d15c",
"collectionId": "6556cd795e4b663e770bb66d",
"sourceName": "GBT 15104-2021 装饰单板贴面人造板.pdf",
"sourceId": "6556cd775e4b663e770bb65c",
"score": 0.8050316572189331
},
......
]
}
```
file: ./content/openapi/index.en.mdx
meta: {
"title": "OpenAPI Documentation",
"description": "FastGPT OpenAPI Documentation"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/openapi/index.mdx
meta: {
"title": "OpenAPI 文档",
"description": "FastGPT OpenAPI 文档"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/openapi/intro.en.mdx
meta: {
"title": "API Documentation Introduction",
"description": "Introduction to FastGPT API Documentation"
}
Starting with `4.15.0`, FastGPT API documentation is generated automatically with `zod-openapi` (some legacy endpoints have not been migrated, so they are not shown). You can view the latest endpoint status by opening the API documentation URL. The manually edited endpoint descriptions in the left sidebar of this documentation are no longer updated.
FastGPT API documentation is split into two sets:
* Dev API: all development APIs. Not every endpoint can be called with an API Key.
* System OpenAPI: all system public endpoints, callable with a system API Key.
## API Documentation URL
`endpoint` is your FastGPT access URL. Append the corresponding path to open the documentation.
* Dev API: `{{endpoint}}/apidoc/devapi`
* System OpenAPI: `{{endpoint}}/apidoc/systemopenapi`
## Cloud API Documentation URL
**Dev API:**
* [China Mainland documentation](https://cloud.fastgpt.cn/apidoc/devapi)
* [International documentation](https://cloud.fastgpt.io/apidoc/devapi)
**System OpenAPI**
* [China Mainland documentation](https://cloud.fastgpt.cn/apidoc/systemopenapi)
* [International documentation](https://cloud.fastgpt.io/apidoc/systemopenapi)
## Usage Notes
FastGPT OpenAPI endpoints let you authenticate with an API Key to operate related FastGPT services and resources, such as calling app chat endpoints, uploading Knowledge Base data, and running search tests. For compatibility and security reasons, not all endpoints can be accessed with an API Key.
### How to Get an API Key
You can find API Keys in two places:
1. `Account` - `API Keys`
2. `App` - `Publish Channels` - `API Access`
### API Key Scope
An API Key acts as the current account's access credential within the current team. In other words, any resource the account can access in that team can also be operated through the API Key.
### How to Find the BaseURL
**Note: BaseURL is not an endpoint URL. It is the root URL for all endpoints, and requesting the BaseURL directly does nothing.**

### Basic Configuration
In OpenAPI, all endpoints authenticate through `Header.Authorization`.
```
baseUrl: "http://localhost:3000/api"
headers: {
Authorization: "Bearer {{apikey}}"
}
```
file: ./content/openapi/intro.mdx
meta: {
"title": "API 文档介绍",
"description": "FastGPT API 文档介绍"
}
从 `4.15.0` 开始,FastGPT API 文档均采用 `zod-openapi` 自动生成的方式(部分旧接口未改造,所以不显示)。可通过访问 API 文档地址查看最新的接口情况,该文档里左侧手动编辑的接口说明不再更新。
FastGPT API 文档一共分成两套:
* Dev API: 所有开发的 API,不一定能通过 ApiKey 调用。
* System OpenAPI: 系统所有开放的接口,可以通过系统 ApiKey 调用。
## API 文档地址
endpoint 是你的 FastGPT 访问地址,拼上对应 path 即可打开文档。
* Dev API: `{{endpoint}}/apidoc/devapi`
* System OpenAPI: `{{endpoint}}/apidoc/systemopenapi`
## 云服务 API 文档地址
**Dev API:**
* [中国大陆版文档](https://cloud.fastgpt.cn/apidoc/devapi)
* [国际版文档](https://cloud.fastgpt.io/apidoc/devapi)
**System OpenAPI**
* [中国大陆版文档](https://cloud.fastgpt.cn/apidoc/systemopenapi)
* [国际版文档](https://cloud.fastgpt.io/apidoc/systemopenapi)
## 使用说明
FastGPT OpenAPI 接口允许你使用 API Key 进行鉴权,从而操作 FastGPT 上的相关服务和资源,例如:调用应用对话接口、上传知识库数据、搜索测试等等。出于兼容性和安全考虑,并不是所有的接口都允许通过 API Key 访问。
### 如何获取 API Key
系统里有两个地方可看到 API 密钥
1. 在 `账号` - `Api 密钥` 中获取
2. 在 `应用` - `发布渠道` - `API 访问` 里查看。
### API 密钥可用范围
API 密钥相当于当前账号,在当前团队下的访问凭证。也就是,在该团队下有权限的资源,都可以通过 API 密钥进行操作。
### 如何查看 BaseURL
**注意:BaseURL 不是接口地址,而是所有接口的根地址,直接请求 BaseURL 是没有用的。**

### 基本配置
OpenAPI 中,所有的接口都通过 Header.Authorization 进行鉴权。
```
baseUrl: "http://localhost:3000/api"
headers: {
Authorization: "Bearer {{apikey}}"
}
```
file: ./content/self-host/dev.en.mdx
meta: {
"title": "Local Development Setup",
"description": "Develop and debug FastGPT locally"
}
import { Alert } from '@/components/docs/Alert';
import FastGPTLink from '@/components/docs/linkFastGPT';
This guide covers how to set up your development environment to build and test FastGPT.
## Prerequisites
Install and configure these dependencies on your machine to build FastGPT:
* [Git](https://git-scm.com/)
* [Docker](https://www.docker.com/)
* [Node.js v20.14.0](https://nodejs.org) (match this version closely; use [nvm](https://github.com/nvm-sh/nvm) to manage Node versions)
* [pnpm](https://pnpm.io/) recommended version 9.4.0 (current official dev environment)
We recommend developing on \*nix environments (Linux, macOS, Windows WSL).
## Local Development
### 1. Fork the FastGPT Repository
Fork the [FastGPT repository](https://github.com/labring/FastGPT).
### 2. Clone the Repository
Clone your forked repository from GitHub:
```
git clone git@github.com:/FastGPT.git
```
### 3. Start the Development Environment with Docker
If you're already running FastGPT locally via Docker, stop it first to avoid port conflicts.
Navigate to `FastGPT/deploy/dev` and run `docker compose up -d` to start FastGPT's dependencies:
```bash
cd FastGPT/deploy/dev
docker compose up -d
```
1. If you can't pull images, use the China mirror version: `docker compose -f
docker-compose.cn.yml up -d` 2. For MongoDB, add the `directConnection=true` parameter to your
connection string to connect to the replica set.
### 4. Initial Configuration
All files below are in the `projects/app` directory.
```bash
# Make sure you're in projects/app
pwd
# Should output /xxxx/xxxx/xxx/FastGPT/projects/app
```
**1. Environment Variables**
Copy `.env.template` to create `.env.local` in the same directory. Only changes in `.env.local` take effect.
See `.env.template` for variable descriptions.
If you haven't modified variables in docker-compose.yaml, the defaults in `.env.template` work as-is. Otherwise, match the values in your `yml` file.
```bash
cp .env.template .env.local
```
**2. config.json Configuration File**
Copy `data/config.json` to create `data/config.local.json`. For detailed parameters, see [Configuration Guide](./config/model/intro.en.mdx).
```bash
cp data/config.json data/config.local.json
```
This file usually doesn't need changes. Key `systemEnv` parameters:
* `vectorMaxProcess`: Max vector generation processes. Depends on database and key concurrency — for a 2c4g server, set to 10–15.
* `qaMaxProcess`: Max QA generation processes
* `vlmMaxProcess`: Max image understanding model processes
* `hnswEfSearch`: Vector search parameter (PG and OB only). Higher values = better accuracy but slower speed.
### 5. Run
See `dev.md` in the project root. The first compile may take a while — be patient.
```bash
# Run from the code root directory to install all dependencies
# If isolate-vm installation fails, see: https://github.com/laverdet/isolated-vm?tab=readme-ov-file#requirements
pwd # Should be in the code root directory
pnpm i
cd projects/app
pnpm dev
```
Next.js runs on port 3000 by default. Visit [http://localhost:3000](http://localhost:3000)
### 6. Build
We recommend using Docker for builds.
```bash
# Without proxy
docker build -f ./projects/app/Dockerfile -t fastgpt . --build-arg name=app
# With Taobao proxy
docker build -f ./projects/app/Dockerfile -t fastgpt. --build-arg name=app --build-arg proxy=taobao
```
Without Docker, you'd need to manually execute all the run-stage commands from the `Dockerfile` (not recommended).
## Contributing to the Open Source Repository
1. Make sure your code is forked from the [FastGPT](https://github.com/labring/FastGPT) repository.
2. Keep commits small and focused — each should address one issue.
3. Submit a PR to FastGPT's main branch. The FastGPT team and community will review it with you.
If you run into issues like merge conflicts, check GitHub's [pull request tutorial](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests). Once your PR is merged, you'll be listed in the [contributors table](https://github.com/labring/FastGPT/graphs/contributors).
## QA
### System Time Anomaly
If your default timezone is `Asia/Shanghai`, system time may be incorrect in non-Linux environments. For local development, change your timezone to UTC (+0).
### Can't Connect to Local Database
1. For remote databases, check if the port is open.
2. For local databases, try changing `host` to `localhost` or `127.0.0.1`.
3. For local connections to remote MongoDB, add `directConnection=true` to connect to replica sets.
4. Use `mongocompass` for MongoDB connection testing and visual management.
5. Use `navicat` for PostgreSQL connection and management.
### sh ./scripts/postinstall.sh Permission Denied
FastGPT runs a `postinstall` script after `pnpm i` to auto-generate ChakraUI types. If you get a permission error, run `chmod -R +x ./scripts/` first, then `pnpm i`.
If that doesn't work, manually execute the contents of `./scripts/postinstall.sh`.
*On Windows, use git bash to add execute permissions and run the script.*
### TypeError: Cannot read properties of null (reading 'useMemo')
Delete all `node_modules` and reinstall with Node 18 — newer Node versions may have issues. Local dev workflow:
1. Root directory: `pnpm i`
2. Copy `config.json` -> `config.local.json`
3. Copy `.env.template` -> `.env.local`
4. `cd projects/app`
5. `pnpm dev`
### Error response from daemon: error while creating mount source path 'XXX': mkdir XXX: file exists
This may be caused by leftover files from a previous container stop. Make sure all related containers are stopped, then manually delete the files or restart Docker.
## Join the Community
Having trouble? Join the Lark group to connect with developers and users.
## Code Structure
### Next.js
FastGPT uses Next.js page routing. To separate frontend and backend code, directories are split into global, service, and web subdirectories for shared, backend-only, and frontend-only code respectively.
### Monorepo
FastGPT uses pnpm workspace for its monorepo structure, with two main parts:
* projects/app - FastGPT main project
* packages/ - Submodules
* global - Shared code: functions, type declarations, and constants usable on both frontend and backend
* service - Server-side code
* web - Frontend code
* plugin - Custom workflow plugin code
### Domain-Driven Design (DDD)
FastGPT's code modules follow DDD principles, divided into these domains:
* core - Core features (knowledge base, workflow, app, conversation)
* support - Supporting features (user system, billing, authentication, etc.)
* common - Base features (log management, file I/O, etc.)
Code Structure Details
```
.
├── .github // GitHub config
├── .husky // Formatting config
├── document // Documentation
├── files // External files, e.g., docker-compose, helm
├── packages // Subpackages
│ ├── global // Frontend/backend shared subpackage
│ ├── plugins // Workflow plugins (for custom packages)
│ ├── service // Backend subpackage
│ └── web // Frontend subpackage
├── projects
│ └── app // FastGPT main project
├── python // Model code, unrelated to FastGPT itself
└── scripts // Automation scripts
├── icon // Icon scripts: pnpm initIcon (write SVG to code), pnpm previewIcon (preview icons)
└── postinstall.sh // ChakraUI custom theme TS type initialization
├── package.json // Top-level monorepo
├── pnpm-lock.yaml
├── pnpm-workspace.yaml // Monorepo declaration
├── Dockerfile
├── LICENSE
├── README.md
├── README_en.md
├── README_ja.md
├── dev.md
```
file: ./content/self-host/dev.mdx
meta: {
"title": "本地开发",
"description": "对 FastGPT 进行开发调试"
}
import { Alert } from '@/components/docs/Alert';
import FastGPTLink from '@/components/docs/linkFastGPT';
本文档介绍了如何设置开发环境以构建和测试 FastGPT。
## 前置开发环境
您需要在计算机上安装和配置以下依赖项才能构建 FastGPT:
* [Git](https://git-scm.com/)
* [Docker](https://www.docker.com/)
* [Node.js >=20](https://nodejs.org)(版本尽量一样,可以使用 [nvm](https://github.com/nvm-sh/nvm) 管理 Node.js 版本)
* [pnpm](https://pnpm.io/) 需要使用 10.x
建议在 \*nix 环境进行开发 (Linux, MacOS, Windows WSL)
## 开始本地开发
### 1. Fork FastGPT 存储库
您需要 Fork [FastGPT 存储库](https://github.com/labring/FastGPT)。
### 2. 克隆存储库
克隆您在 GitHub 上 Fork 的存储库:
```
git clone git@github.com:/FastGPT.git
```
### 3. 通过 docker 启动开发环境
若您本地已经通过 docker 启动了 FastGPT,则需要先关闭,否则会有端口冲突。
切换到 `FastGPT/deploy/dev` 目录,执行 `docker compose up -d` 运行 FastGPT 的各种依赖。
```bash
cd FastGPT/deploy/dev
docker compose up -d
```
1. 如果无法获取镜像,可以选择国内镜像版本的 docker-compose.yml 文件:`docker compose -f
docker-compose.cn.yml up -d` 2. Mongo 数据库需要注意,需要注意在连接地址中增加
`directConnection=true` 参数,才能连接上副本集的数据库。
### 4. 初始配置
以下文件均在 `projects/app` 路径下。
```bash
# 确保你现在在 projects/app 下
pwd
# 应当输出 /xxxx/xxxx/xxx/FastGPT/projects/app
```
**1. 环境变量**
复制 `.env.template` 文件,在同级目录下生成一个 `.env.local` 文件,修改 `.env.local` 里内容才是有效的变量。变量说明见 `.env.template`
如果没有修改 docker-compose.yaml 中的变量,`.env.template` 中的默认值就可以,不需要进行修改,否则需要和 `yml` 中的变量一致。
```bash
cp .env.template .env.local
```
**2. config.json 配置文件**
复制 `data/config.json` 文件,生成一个 `data/config.local.json` 配置文件,具体配置参数说明,可参考 [config 配置说明](./config/model/intro.mdx)
```bash
cp data/config.json data/config.local.json
```
这个文件大部分时候不需要修改。只需要关注 `systemEnv` 里的参数:
* `vectorMaxProcess` : 向量生成最大进程,根据数据库和 key 的并发数来决定,通常单个 120 号,2c4g 服务器设置 10\~15。
* `qaMaxProcess` : QA 生成最大进程
* `vlmMaxProcess` : 图片理解模型最大进程
* `hnswEfSearch` : 向量搜索参数,仅对 PG 和 OB 生效,越大搜索精度越高但是速度越慢。
### 5. 运行
可参考项目根目录下的 `dev.md`,第一次编译运行可能会有点慢,需要点耐心哦
```bash
# 代码根目录下执行,会安装根 package、projects 和 packages 内所有依赖
# 如果提示 isolate-vm 安装失败,可以参考:https://github.com/laverdet/isolated-vm?tab=readme-ov-file#requirements
pwd # 应该在代码的根目录
pnpm i
cd projects/app
pnpm dev
```
默认 next 将运行在 3000 端口,访问 [http://localhost:3000](http://localhost:3000)
### 6. 打包
建议直接使用 Docker 进行打包。
```bash
# 没有 Proxy
docker build -f ./projects/app/Dockerfile -t fastgpt . --build-arg name=app
# Taobao Proxy
docker build -f ./projects/app/Dockerfile -t fastgpt. --build-arg name=app --build-arg proxy=taobao
```
如果不使用 `docker` 打包,需要手动把 `Dockerfile` 里 run 阶段的内容全部手动执行一遍(非常不推荐)。
## 提交代码至开源仓库
1. 确保你的代码是 Fork [FastGPT](https://github.com/labring/FastGPT) 仓库
2. 尽可能少量的提交代码,每次提交仅解决一个问题。
3. 向 FastGPT 的 main 分支提交一个 PR,提交请求后,FastGPT 团队/社区的其他人将与您一起审查它。
如果遇到问题,比如合并冲突或不知道如何打开拉取请求,请查看 GitHub 的[拉取请求教程](https://docs.github.com/en/pull-requests/collaborating-with-pull-requests),了解如何解决合并冲突和其他问题。一旦您的 PR 被合并,您将自豪地被列为[贡献者表](https://github.com/labring/FastGPT/graphs/contributors)中的一员。
## QA
### 获取系统时间异常
如果用户默认的时区为 `Asia/Shanghai` , 非 linux 环境时,获取系统时间会异常,本地开发时,可以将用户的时区调整成 UTC(+0)。
### 本地数据库无法连接
1. 如果你是连接远程的数据库,先检查对应的端口是否开放。
2. 如果是本地运行的数据库,可尝试 `host` 改成 `localhost` 或 `127.0.0.1`
3. 本地连接远程的 Mongo,需要增加 `directConnection=true` 参数,才能连接上副本集的数据库。
4. mongo 使用 `mongocompass` 客户端进行连接测试和可视化管理。
5. pg 使用 `navicat` 进行连接和管理。
### sh ./scripts/postinstall.sh 没权限
FastGPT 在 `pnpm i` 后会执行 `postinstall` 脚本,用于自动生成 `ChakraUI` 的 `Type`。如果没有权限,可以先执行 `chmod -R +x ./scripts/`,再执行 `pnpm i`。
仍不可行的话,可以手动执行 `./scripts/postinstall.sh` 里的内容。*如果是 Windows 下的话,可以使用 git bash 给 `postinstall` 脚本添加执行权限并执行 sh 脚本*
### TypeError: Cannot read properties of null (reading 'useMemo' )
删除所有的 `node_modules`,用 Node18 重新 install 试试,可能最新的 Node.js 有问题。本地开发流程:
1. 根目录: `pnpm i`
2. 复制 `config.json` -> `config.local.json`
3. 复制 `.env.template` -> `.env.local`
4. `cd projects/app`
5. `pnpm dev`
### Error response from daemon: error while creating mount source path 'XXX': mkdir XXX: file exists
这个错误可能是之前停止容器时有文件残留导致的,首先需要确认相关镜像都全部关闭,然后手动删除相关文件或者重启 docker 即可
## 加入社区
遇到困难了吗?有任何问题吗? 加入飞书群与开发者和用户保持沟通。
## 代码结构说明
### nextjs
FastGPT 使用了 nextjs 的 page route 作为框架。为了区分好前后端代码,在目录分配上会分成 global, service, web 3 个自目录,分别对应着 `前后端共用`、`后端专用`、`前端专用` 的代码。
### monorepo
FastGPT 采用 pnpm workspace 方式构建 monorepo 项目,主要分为两个部分:
* projects/app - FastGPT 主项目
* packages/ - 子模块
* global - 共用代码,通常是放一些前后端都能执行的函数、类型声明、常量。
* service - 服务端代码
* web - 前端代码
* plugin - 工作流自定义插件的代码
### 领域驱动模式(DDD)
FastGPT 在代码模块划分时,按 DDD 的思想进行划分,主要分为以下几个领域:
* core - 核心功能(知识库,工作流,应用,对话)
* support - 支撑功能(用户体系,计费,鉴权等)
* common - 基础功能(日志管理,文件读写等)
代码结构说明
```
.
├── .github // github 相关配置
├── .husky // 格式化配置
├── document // 文档
├── files // 一些外部文件,例如 docker-compose, helm
├── packages // 子包
│ ├── global // 前后端通用子包
│ ├── plugins // 工作流插件(需要自定义包时候使用到)
│ ├── service // 后端子包
│ └── web // 前端子包
├── projects
│ └── app // FastGPT 主项目
├── python // 存放一些模型代码,和 FastGPT 本身无关
└── scripts // 一些自动化脚本
├── icon // icon预览脚本,可以在顶层 pnpm initIcon(把svg写入到代码中), pnpm previewIcon(预览icon)
└── postinstall.sh // chakraUI自定义theme初始化 ts 类型
├── package.json // 顶层monorepo
├── pnpm-lock.yaml
├── pnpm-workspace.yaml // monorepo 声明
├── Dockerfile
├── LICENSE
├── README.md
├── README_en.md
├── README_ja.md
├── dev.md
```
file: ./content/self-host/index.en.mdx
meta: {
"title": "Self-Host",
"description": "FastGPT Self-Host"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/self-host/index.mdx
meta: {
"title": "自部署",
"description": "FastGPT 自部署"
}
import { Redirect } from '@/components/docs/Redirect';
file: ./content/guide/admin/sso.en.mdx
meta: {
"title": "SSO & External Member Sync",
"description": "FastGPT External Member System Integration and Configuration"
}
import { Alert } from '@/components/docs/Alert';
If you don't need SSO or member sync, or only need quick login via GitHub, Google, Microsoft, or WeChat Official Account, you can skip this section. This guide is for users who need to integrate their own member systems or mainstream office IMs.
## Overview
To simplify integration with **external member systems**, FastGPT provides a set of **standard interfaces** for connecting to external systems, along with a FastGPT-SSO-Service image that serves as an **adapter**.
Through these standard interfaces, you can:
1. SSO login. After a callback from an external system, create a user in FastGPT.
2. Member and organizational structure sync (referred to as "member sync" below).
**How It Works**
FastGPT-pro includes a standard set of SSO and member sync interfaces. The system performs SSO and member sync operations based on these interfaces.
FastGPT-SSO-Service aggregates SSO and member sync interfaces from different sources and converts them into the format recognized by fastgpt-pro.

## System Configuration Tutorial
### 1. Deploy the SSO-Service Image
Deploy using docker-compose:
```yaml
fastgpt-sso:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-sso-service:v4.9.0 # This version must match the FastGPT image version
container_name: fastgpt-sso
restart: always
networks:
- fastgpt
environment:
- SSO_PROVIDER=example
- AUTH_TOKEN=xxxxx # Auth token, used by fastgpt-pro
# Provider-specific environment variables below
```
Depending on the provider, you'll need different environment variables. Below are the built-in protocols/IMs:
Protocol/Feature
SSO
Member Sync Support
Lark
Yes
Yes
WeCom
Yes
Yes
DingTalk
Yes
No
SAML 2.0
Yes
No
OAuth 2.0
Yes
No
### 2. Configure fastgpt-pro
#### 1. Configure Environment Variables
The `EXTERNAL_USER_SYSTEM_BASE_URL` environment variable should be set to the internal network address. For example, with the configuration above:
```yaml
env:
- EXTERNAL_USER_SYSTEM_BASE_URL=http://fastgpt-sso:3000
- EXTERNAL_USER_SYSTEM_AUTH_TOKEN=xxxxx
```
#### 2. Configure button text, icons, etc. in the commercial version admin panel.
WeCom
DingTalk
Lark



#### 3. Enable Member Sync (Optional)
If you need to sync members from an external system, you can enable member sync. For team mode details, see: [Team Mode Documentation](./teamMode.en.mdx)

#### 4. Optional Configuration
1. Automatic scheduled member sync
Set the fastgpt-pro environment variable to enable automatic member sync:
```yaml
env:
- "SYNC_MEMBER_CRON=0 0 * * *" # Cron expression, runs daily at 00:00. Note: uses UTC (timezone 0). For example, to sync at 12:00 Beijing time, set this to "0 4 * * *" (UTC 04:00)
```
## Built-in Protocol/IM Configuration Examples
### Lark
#### 1. Get Parameters
App ID and App Secret
Go to the developer console, click on your enterprise self-built app, and view the app credentials on the Credentials & Basic Info page.

#### 2. Permission Configuration
Go to the developer console, click on your enterprise self-built app, and enable permissions on the Permission Management page under Development Configuration.

You can use the **Batch Import/Export Permissions** feature to import the following permission configuration:
```json
{
"scopes": {
"tenant": [
"contact:user.phone:readonly",
"contact:contact.base:readonly",
"contact:department.base:readonly",
"contact:department.organize:readonly",
"contact:user.base:readonly",
"contact:user.department:readonly",
"contact:user.email:readonly",
"contact:user.employee_id:readonly"
],
"user": []
}
}
```
Note: The accessible data scope must be set to visible to all members.
#### 3. Redirect URL
Go to the developer console, click on your enterprise self-built app, and set the redirect URL in Security Settings under Development Configuration.
The redirect URL should follow the format `https://www.fastgpt.cn/login/provider` — replace the domain with your publicly accessible FastGPT domain.

#### 4. yml Configuration Example
```yaml
fastgpt-sso:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-sso-service:v4.9.0
container_name: fastgpt-sso
restart: always
networks:
- fastgpt
environment:
- SSO_PROVIDER=feishu
- AUTH_TOKEN=xxxxx
# OAuth endpoint (for private Lark deployments, replace with your private address; same below)
- SSO_TARGET_URL=https://accounts.feishu.cn/open-apis/authen/v1/authorize
# Token endpoint
- FEISHU_TOKEN_URL=https://open.feishu.cn/open-apis/authen/v2/oauth/token
# User info endpoint
- FEISHU_GET_USER_INFO_URL=https://open.feishu.cn/open-apis/authen/v1/user_info
# Redirect address — must match the URL from step 3 exactly
- FEISHU_REDIRECT_URI=https://fastgpt.cn/login/provider
# Lark App ID, usually starts with cli
- FEISHU_APP_ID=xxx
# Lark App Secret
- FEISHU_APP_SECRET=xxx
```
### DingTalk
#### 1. Get Parameters
CLIENT\_ID and CLIENT\_SECRET
Go to the DingTalk Open Platform, click App Development, select your app, and record the Client ID and Client Secret on the Credentials & Basic Info page.

#### 2. Permission Configuration
Go to the DingTalk Open Platform, click App Development, select your app, and manage permissions on the Permission Management page under Development Configuration. Required permissions:
1. ***Personal phone number information***
2. ***Contact personal information read permission***
3. ***Basic permission to obtain DingTalk open interface user access credentials***
#### 3. Redirect URL
Go to the DingTalk Open Platform, click App Development, select your app, and configure on the Security Settings page under Development Configuration.
Two items need to be filled in:
1. Server egress IP (list of server IPs calling DingTalk server-side APIs)
2. Redirect URL (callback domain)
#### 4. yml Configuration Example
```yaml
fastgpt-sso:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-sso-service:v4.9.0
container_name: fastgpt-sso
restart: always
networks:
- fastgpt
environment:
- SSO_PROVIDER=dingtalk
- AUTH_TOKEN=xxxxx
# OAuth endpoint
- SSO_TARGET_URL=https://login.dingtalk.com/oauth2/auth
# Token endpoint
- DINGTALK_TOKEN_URL=https://api.dingtalk.com/v1.0/oauth2/userAccessToken
# User info endpoint
- DINGTALK_GET_USER_INFO_URL=https://oapi.dingtalk.com/v1.0/contact/users/me
# DingTalk App ID
- DINGTALK_CLIENT_ID=xxx
# DingTalk App Secret
- DINGTALK_CLIENT_SECRET=xxx
```
### WeCom
#### 1. Get Parameters
1. Enterprise CorpID
a. Log in to the WeCom admin console with an admin account: `https://work.weixin.qq.com/wework_admin/loginpage_wx`
b. Go to the "My Enterprise" page and find the Enterprise ID

2. Create an internal app for FastGPT:
a. Get the app's AgentID and Secret
b. Ensure the app's visibility scope is set to all (i.e., root department)


3. A domain name with the following requirements:
a. Resolves to a publicly accessible server
b. Can serve static files at the root path (for domain ownership verification — follow the prompts, you only need to host one static file, which can be removed after verification)
c. Configure web authorization, JS-SDK, and WeCom authorization login
d. You can set "Hide app in workbench" at the bottom of the WeCom Authorization Login page



4. Get the "Contact Sync Assistant" secret
Retrieving contacts and organization member IDs requires the "Contact Sync Assistant" secret
Security & Management -- Management Tools -- Contact Sync

5. Enable interface sync
6. Get the Secret
7. Configure enterprise trusted IPs

#### 2. yml Configuration Example
```yaml
fastgpt-sso:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-sso-service:v4.9.0
container_name: fastgpt-sso
restart: always
networks:
- fastgpt
environment:
- AUTH_TOKEN=xxxxx
- SSO_PROVIDER=wecom
# OAuth endpoint, used in WeCom client
- WECOM_TARGET_URL_OAUTH=https://open.weixin.qq.com/connect/oauth2/authorize
# SSO endpoint, QR code scan
- WECOM_TARGET_URL_SSO=https://login.work.weixin.qq.com/wwlogin/sso/login
# Get user ID (returns ID only)
- WECOM_GET_USER_ID_URL=https://qyapi.weixin.qq.com/cgi-bin/auth/getuserinfo
# Get detailed user info (everything except name)
- WECOM_GET_USER_INFO_URL=https://qyapi.weixin.qq.com/cgi-bin/auth/getuserdetail
# Get user info (has name, no other info)
- WECOM_GET_USER_NAME_URL=https://qyapi.weixin.qq.com/cgi-bin/user/get
# Get department ID list
- WECOM_GET_DEPARTMENT_LIST_URL=https://qyapi.weixin.qq.com/cgi-bin/department/list
# Get user ID list
- WECOM_GET_USER_LIST_URL=https://qyapi.weixin.qq.com/cgi-bin/user/list_id
# WeCom CorpId
- WECOM_CORPID=
# WeCom App AgentId, usually 1000xxx
- WECOM_AGENTID=
# WeCom App Secret
- WECOM_APP_SECRET=
# Contact Sync Assistant Secret
- WECOM_SYNC_SECRET=
```
### Standard OAuth 2.0
We provide OAuth 2.0 integration support using the authorization code grant from RFC 6749.
References:
* [RFC 6749](https://datatracker.ietf.org/doc/html/rfc6749) documentation
* [Ruan Yifeng's blog post on OAuth 2.0](https://www.ruanyifeng.com/blog/2014/05/oauth_2_0.html)
#### Parameter Requirements
##### Three Endpoints
We provide a standard OAuth 2.0 integration flow requiring three endpoints:
1. Login authorization endpoint (users are redirected here with parameters after clicking the SSO button), e.g., `http://example.com/oauth/authorize`
```bash
curl -X GET\
"http://example.com/oauth/authorize?response_type=code&client_id=s6BhdRkqt3&state=xyz&redirect_uri=https%3A%2F%2Ffastgpt.cn%2Flogin%2Fprovider"
```
After entering credentials, users are redirected to redirect\_uri with a code parameter:
`https://fastgpt.cn/login/provider?code=4/P7qD2qAz4&state=xyz`
2. Access token endpoint. After obtaining the code, make a *server-side request* to this endpoint to get the access\_token, e.g., `http://example.com/oauth/access_token`
```bash
curl -X POST\
-H "Content-Type: application/x-www-form-urlencoded"\
"http://example.com/oauth/access_token?grant_type=authorization_code&client_id=s6BhdRkqt3&client_secret=xxx&code=4/P7qD2qAz4&redirect_uri=https%3A%2F%2Ffastgpt.cn%2Flogin%2Fprovider"
```
Note: Content-Type must be application/x-www-form-urlencoded, not application/json
3. User info endpoint, requires passing the access\_token, e.g., `http://example.com/oauth/user_info`
```bash
curl -X GET\
-H "Authorization: Bearer 4/P7qD2qAz4"\
"http://example.com/oauth/user_info"
```
Note: access\_token is passed as the Authorization header in the format: Bearer xxxx
##### Parameter Configuration
* CLIENT\_ID: Required
* CLIENT\_SECRET: Optional, skip if not needed
* SCOPE: Optional, skip if not needed
> The redirect\_uri parameter is auto-populated based on the runtime environment
>
> Other fixed parameters like grant\_type and response\_type are auto-populated
#### Configuration Example
```yaml
fastgpt-sso:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-sso-service:v4.9.0
container_name: fastgpt-sso
restart: always
networks:
- fastgpt
environment:
- SSO_PROVIDER=oauth2
- AUTH_TOKEN=xxxxx
# OAuth2.0
# === Request URLs ===
# 1. OAuth2 login authorization URL (required)
- OAUTH2_AUTHORIZE_URL=
# 2. OAuth2 access token URL (required)
- OAUTH2_TOKEN_URL=
# 3. OAuth2 user info URL (required)
- OAUTH2_USER_INFO_URL=
# === Parameters ===
# 1. client_id (required)
- OAUTH2_CLIENT_ID=
# 2. client_secret (optional)
- OAUTH2_CLIENT_SECRET=
# 3. scope (optional)
- OAUTH2_SCOPE=
# === Field Mapping ===
# OAuth2 username field mapping (required)
- OAUTH2_USERNAME_MAP=
# OAuth2 avatar field mapping (optional)
- OAUTH2_AVATAR_MAP=
# OAuth2 member name field mapping (optional)
- OAUTH2_MEMBER_NAME_MAP=
# OAuth2 contact field mapping (optional)
- OAUTH2_CONTACT_MAP=
```
## Standard Interface Documentation
Below is the standard interface documentation for SSO and member sync in FastGPT-pro. If you need to integrate with a non-standard system, refer to this section for development.

FastGPT provides the following standard interfaces:
1. [https://example.com/login/oauth/getAuthURL](https://example.com/login/oauth/getAuthURL) - Get the authorization redirect URL
2. [https://example.com/login/oauth/getUserInfo?code=xxxxx](https://example.com/login/oauth/getUserInfo?code=xxxxx) - Consume the code and exchange it for user info
3. [https://example.com/org/list](https://example.com/org/list) - Get the organization list
4. [https://example.com/user/list](https://example.com/user/list) - Get the member list
### Get SSO Login Redirect URL
Returns a redirect login URL. FastGPT will automatically redirect to this URL. The redirect\_uri is automatically appended to the URL query string.
```bash
curl -X GET "https://redict.example/login/oauth/getAuthURL?redirect_uri=xxx&state=xxxx" \
-H "Authorization: Bearer your_token_here" \
-H "Content-Type: application/json"
```
Success:
```json
{
"success": true,
"message": "",
"authURL": "https://example.com/somepath/login/oauth?redirect_uri=https%3A%2F%2Ffastgpt.cn%2Flogin%2Fprovider%0A"
}
```
Failure:
```json
{
"success": false,
"message": "Error message",
"authURL": ""
}
```
### SSO Get User Info
This interface accepts a code parameter for authentication, consumes the code, and returns user info.
```bash
curl -X GET "https://oauth.example/login/oauth/getUserInfo?code=xxxxxx" \
-H "Authorization: Bearer your_token_here" \
-H "Content-Type: application/json"
```
Success:
```json
{
"success": true,
"message": "",
"username": "fastgpt-123456789",
"avatar": "https://example.webp",
"contact": "+861234567890",
"memberName": "Member name (optional)",
}
```
Failure:
```json
{
"success": false,
"message": "Error message",
"username": "",
"avatar": "",
"contact": ""
}
```
### Get Organizations
```bash
curl -X GET "https://example.com/org/list" \
-H "Authorization: Bearer your_token_here" \
-H "Content-Type: application/json"
```
Warning: Only one root department can exist. If your system has multiple root departments, you need to add a virtual root department first. Return type:
```ts
type OrgListResponseType = {
message?: string; // Error message
success: boolean;
orgList: {
id: string; // Unique department ID
name: string; // Name
parentId: string; // parentId — empty string for root department
}[];
}
```
```json
{
"success": true,
"message": "",
"orgList": [
{
"id": "od-125151515",
"name": "Root Department",
"parentId": ""
},
{
"id": "od-51516152",
"name": "Sub Department",
"parentId": "od-125151515"
}
]
}
```
### Get Members
```bash
curl -X GET "https://example.com/user/list" \
-H "Authorization: Bearer your_token_here" \
-H "Content-Type: application/json"
```
Return type:
```typescript
type UserListResponseListType = {
message?: string; // Error message
success: boolean;
userList: {
username: string; // Unique ID. username must match the username returned by the SSO interface. Must include a prefix, e.g., sync-aaaaa, consistent with the SSO interface prefix
memberName?: string; // Name, used as tmbname
avatar?: string;
contact?: string; // email or phone number
orgs?: string[]; // IDs of organizations the member belongs to. Pass [] if no organization
}[];
}
```
curl example
```json
{
"success": true,
"message": "",
"userList": [
{
"username": "fastgpt-123456789",
"memberName": "John Doe",
"avatar": "https://example.webp",
"contact": "+861234567890",
"orgs": ["od-125151515", "od-51516152"]
},
{
"username": "fastgpt-12345678999",
"memberName": "Jane Smith",
"avatar": "",
"contact": "",
"orgs": ["od-125151515"]
}
]
}
```
## How to Integrate Non-Standard Systems
1. Self-development: Build according to the standard interfaces provided by FastGPT, then enter the deployed service address into fastgpt-pro.
You can use this template repository as a starting point: [fastgpt-sso-template](https://github.com/labring/fastgpt-sso-template)
2. Custom development by the FastGPT team:
a. Provide the system's SSO documentation, member and organization retrieval documentation, and an external test address.
b. In fastgpt-sso-service, add the corresponding provider and environment variables, and write the integration code.
file: ./content/guide/admin/sso.mdx
meta: {
"title": "SSO & 外部成员同步",
"description": "FastGPT 外部成员系统接入设计与配置"
}
import { Alert } from '@/components/docs/Alert';
如果你不需要用到 SSO/成员同步功能,或者是只需要用 Github、google、microsoft、公众号的快速登录,可以跳过本章节。本章适合需要接入自己的成员系统或主流 办公IM 的用户。
## 介绍
为了方便地接入**外部成员系统**,FastGPT 提供一套接入外部系统的**标准接口**,以及一个 FastGPT-SSO-Service 镜像作为**适配器**。
通过这套标准接口,你可以可以实现:
1. SSO 登录。从外部系统回调后,在 FastGPT 中创建一个用户。
2. 成员和组织架构同步(下面都简称成员同步)。
**原理**
FastGPT-pro 中,有一套标准的SSO 和成员同步接口,系统会根据这套接口进行 SSO 和成员同步操作。
FastGPT-SSO-Service 是为了聚合不同来源的 SSO 和成员同步接口,将他们转成 fastgpt-pro 可识别的接口。

## 系统配置教程
### 1. 部署 SSO-service 镜像
使用 docker-compose 部署:
```yaml
fastgpt-sso:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-sso-service:v4.14.16 # 目前sso最新版本,可直接使用当前版本
container_name: fastgpt-sso
restart: always
networks:
- fastgpt
environment:
- SSO_PROVIDER=example
- AUTH_TOKEN=xxxxx # 鉴权信息,fastgpt-pro 会用到。
# 具体对接提供商的环境变量。
```
根据不同的提供商,你需要配置不同的环境变量,下面是内置的通用协议/IM:
### Multi-team Mode (Default)
In multi-team mode, a default team owned by the user is automatically created when each user is created.
### Single-team Mode
Single-team mode is a new feature introduced in v4.9. To simplify personnel and resource management for enterprises, when single-team mode is enabled, new users no longer get their own default team — instead, they are added to the root user's team.
### Sync Mode
When system configuration is complete and sync mode is enabled, members from external member systems are automatically synced to FastGPT.
For specific sync methods and rules, see [SSO & External Member Sync](./sso.en.mdx).
## Configuration
In `fastgpt-pro`'s System Configuration - Member Configuration, you can configure the team mode.

file: ./content/guide/admin/teamMode.mdx
meta: {
"title": "团队模式说明文档",
"description": "FastGPT 团队模式说明文档"
}
## 介绍
目前支持的团队模式:
1. 多团队模式(默认模式)
2. 单团队模式(全局只有一个团队)
3. 成员同步模式(所有成员自外部同步)
团队模式
短信/邮箱 注册
管理员直接添加
SSO 注册
是否创建默认团队
是否加入 Root 团队
是否创建默认团队
是否加入 Root 团队
是否创建默认团队
是否加入 Root 团队
单团队模式
❌
✅
❌
✅
❌
✅
多团队模式
✅
❌
✅
❌
✅
❌
同步模式
❌
❌
❌
❌
❌
✅
### 多团队模式(默认模式)
多团队模式下,每个用户创建时默认创建以自己为所有者的默认团队。
### 单团队模式
单团队模式是 v4.9 推出的新功能。为了简化企业进行人员和资源的管理,开启单团队模式后,所有新增的用户都不再创建自己的默认团队,而是加入 root 用户所在的团队。
### 同步模式
在完成系统配置,开启同步模式的情况下,外部成员系统的成员会自动同步到 FastGPT 中。
具体的同步方式和规则请参考 [SSO & 外部成员同步](./sso.mdx)。
## 配置
在 `fastgpt-pro` 的`系统配置-成员配置`中,可以配置团队模式。

file: ./content/guide/chat/htmlRendering.en.mdx
meta: {
"title": "Dialog Boxes & HTML Rendering",
"description": "How to embed HTML code blocks in FastGPT via Markdown, with fullscreen, source code toggle, and other interactive features"
}
| Source Mode | Preview Mode | Fullscreen Mode |
| ----------------------------- | ----------------------------- | ----------------------------- |
|  |  |  |
### 1. Design Background
While Markdown natively supports embedded HTML tags, many platforms restrict HTML rendering for security reasons -- especially for dynamic content, interactive elements, and external resources. These restrictions limit flexibility when authoring complex documents that need embedded HTML. To address this, FastGPT uses `iframe` to embed and render HTML content, combined with the `sandbox` attribute to ensure safe rendering.
### 2. Feature Overview
This module extends FastGPT's Markdown rendering to support embedded HTML content. Since rendering uses an iframe, the content height cannot be determined automatically, so FastGPT sets a fixed height for the iframe. JavaScript execution within the HTML is not supported.
### 3. Technical Implementation
This module implements HTML rendering and interactivity through:
* **Component Design:** The module displays HTML content via `iframe`-type code blocks using a custom `IframeBlock` component. The `sandbox` attribute ensures embedded content security by restricting behaviors like script execution and form submissions. Helper functions integrate with the Markdown renderer to handle `iframe`-embedded HTML content.
* **Security Mechanism:** The `iframe`'s `sandbox` attribute and `referrerPolicy` prevent potential security risks. The `sandbox` attribute provides fine-grained control, allowing specific capabilities (scripts, forms, popups, etc.) to run in a restricted environment so rendered HTML cannot compromise the system.
* **Display & Interaction:** Users can switch between display modes (fullscreen, preview, source code) for flexible viewing and control of embedded HTML. The `iframe` adapts to the parent container's width while ensuring content displays properly.
### 4. How to Use
Simply use a Markdown code block with the language set to `html`. For example:
````md
```html
Welcome to FastGPT
````
```
```
file: ./content/guide/chat/htmlRendering.mdx
meta: {
"title": "对话框与HTML渲染",
"description": "如何在FastGPT中通过Markdown嵌入HTML代码块,并提供全屏、源代码切换等交互功能"
}
| 源码模式 | 预览模式 | 全屏模式 |
| ----------------------------- | ----------------------------- | ----------------------------- |
|  |  |  |
### 1. **设计背景**
尽管Markdown本身支持嵌入HTML标签,但由于安全问题,许多平台和环境对HTML的渲染进行了限制,特别是在渲染动态内容、交互式元素以及外部资源时。这些限制大大降低了用户在撰写和展示复杂文档时的灵活性,尤其是当需要嵌入外部HTML内容时。为了应对这一问题,我们通过使用 `iframe` 来嵌入和渲染HTML内容,并结合 `sandbox` 属性,保障了外部HTML的安全渲染。
### 2. 功能简介
该功能模块的主要目的是扩展FastGPT在Markdown渲染中的能力,支持嵌入和渲染HTML内容。由于是利用 Iframe 渲染,所以无法确认内容的高度,FastGPT 中会给 Iframe 设置一个固定高度来进行渲染。并且不支持 HTML 中执行 js 脚本。
### 3. 技术实现
本模块通过以下方式实现了HTML渲染和互动功能:
* **组件设计**:该模块通过渲染 `iframe` 类型的代码块展示HTML内容。使用自定义的 `IframeBlock` 组件,结合 `sandbox` 属性来保障嵌入内容的安全性。`sandbox` 限制了外部HTML中的行为,如禁用脚本执行、限制表单提交等,确保HTML内容的安全性。通过辅助函数与渲染Markdown内容的部分结合,处理 `iframe` 嵌入的HTML内容。
* **安全机制**:通过 `iframe` 的 `sandbox` 属性和 `referrerPolicy` 来防止潜在的安全风险。`sandbox` 属性提供了细粒度的控制,允许特定的功能(如脚本、表单、弹出窗口等)在受限的环境中执行,以确保渲染的HTML内容不会对系统造成威胁。
* **展示与互动功能**:用户可以通过不同的展示模式(如全屏、预览、源代码模式)自由切换,以便更灵活地查看和控制嵌入的HTML内容。嵌入的 `iframe` 自适应父容器的宽度,同时保证 `iframe`嵌入的内容能够适当显示。
### 4. 如何使用
你只需要通过 Markdown 代码块格式,并标记语言为 `html` 即可。例如:
````md
```html
欢迎使用FastGPT
````
file: ./content/guide/chat/quoteList.en.mdx
meta: {
"title": "Knowledge Base Chunk Reader",
"description": "FastGPT Chunk Reader feature overview"
}
In enterprise AI deployments, the accuracy and transparency of document citations have always been a key concern. The Knowledge Base Chunk Reader introduced in FastGPT 4.9.1 solves this pain point, making AI citations no longer a "black box."
# Why a Chunk Reader?
In traditional AI conversations, when a model cites content from an enterprise knowledge base, users typically only see the cited fragment without the full context. This makes content verification and deeper understanding difficult. The Chunk Reader lets users view the complete source document directly within the conversation and jump to the exact citation location, bringing true explainability to AI citations.
## Limitations of Traditional Citations
Previously, after uploading documents to the knowledge base, traditional citations only displayed the matched chunks with no way to see the surrounding context:
| Question | Citation |
| --------------------------- | --------------------------- |
|  |  |
## FastGPT Chunk Reader: Precise Positioning, Seamless Reading
With FastGPT's Chunk Reader, the same knowledge base content and questions are presented in a fundamentally better way:

When AI cites knowledge base content, click the citation link to open a popup showing the full original text with the cited passage clearly highlighted. This ensures traceability while providing a convenient reading experience.
# Core Features
## Full-Text Display & Positioning
The Chunk Reader lets users see exactly where AI responses draw from in the knowledge base.
In the conversation interface, when AI cites knowledge base content, source information appears below the reply. Click any citation link to open a popup with the complete original text and the cited passage highlighted.
This design ensures answer traceability and makes it easy to verify AI accuracy and review surrounding context.

## Citation Navigation
The top-right corner of the Chunk Reader provides simple navigation controls for switching between multiple citations. The navigation area displays the current citation index and total count (e.g., "7/10"), so you always know your browsing progress.

## Citation Quality Scoring
Each citation includes a relevance score label showing its ranking among all matched knowledge fragments. Hover over the label to see full scoring details, including why the citation was selected and how its relevance score breaks down.

## One-Click Document Export
The Chunk Reader includes a content export feature so valuable information is never lost. Users with read access to the knowledge base can save the full cited document to their local device with a single click.

# Advanced Features
## Flexible Visibility Control
FastGPT provides flexible citation visibility settings to balance openness and security. For example, with anonymous share links, administrators can precisely control what external visitors can see.
When set to "citation content only," external users clicking a citation link will only see the specific cited text fragments, not the full source document. The Chunk Reader automatically adjusts its display mode accordingly.
| | |
| --------------------------- | --------------------------- |
|  |  |
## Instant Annotation
While browsing, authorized users can annotate and correct citation content in real time. The system processes updates without interrupting the conversation. Modified content is clearly marked with an "Updated" label, maintaining both citation accuracy and conversation history integrity.
This seamless knowledge refinement workflow is ideal for team collaboration, allowing the knowledge base to evolve during actual use so AI responses always draw from the latest, most accurate sources.
## Smart Document Performance
For real-world scenarios with ultra-long documents containing thousands of chunks, FastGPT uses advanced performance optimization to keep the Chunk Reader responsive.
The system manages loading intelligently based on citation relevance ranking and database indexing, implementing on-demand rendering -- only content the user actually needs to view is loaded into memory. Whether jumping to a specific citation or scrolling through a document, the experience stays smooth regardless of document size.
This optimization lets FastGPT handle enterprise-scale knowledge bases efficiently, even for professional documents with massive amounts of content.
file: ./content/guide/chat/quoteList.mdx
meta: {
"title": "知识库引用分块阅读器",
"description": "FastGPT 分块阅读器功能介绍"
}
在企业 AI 应用落地过程中,文档知识引用的精确性和透明度一直是用户关注的焦点。FastGPT 4.9.1 版本带来的知识库分块阅读器,巧妙解决了这一痛点,让 AI 引用不再是"黑盒"。
# 为什么需要分块阅读器?
传统的 AI 对话中,当模型引用企业知识库内容时,用户往往只能看到被引用的片段,无法获取完整语境,这给内容验证和深入理解带来了挑战。分块阅读器的出现,让用户可以在对话中直接查看引用内容的完整文档,并精确定位到引用位置,实现了引用的"可解释性"。
## 传统引用体验的局限
以往在知识库中上传文稿后,当我们在工作流中输入问题时,传统的引用方式只会展示引用到的分块,无法确认分块在文章中的上下文:
| 问题 | 引用 |
| --------------------------- | --------------------------- |
|  |  |
## FastGPT 分块阅读器:精准定位,无缝阅读
而在 FastGPT 全新的分块式阅读器中,同样的知识库内容和问题,呈现方式发生了质的飞跃

当 AI 引用知识库内容时,用户只需点击引用链接,即可打开一个浮窗,呈现完整的原文内容,并通过醒目的高亮标记精确显示引用的文本片段。这既保证了回答的可溯源性,又提供了便捷的原文查阅体验。
# 核心功能
## 全文展示与定位
"分块阅读器" 让用户能直观查看AI回答引用的知识来源。
在对话界面中,当 AI 引用了知识库内容,系统会在回复下方展示出处信息。用户只需点击这些引用链接,即可打开一个优雅的浮窗,呈现完整的原文内容,并通过醒目的高亮标记精确显示 AI 引用的文本片段。
这一设计既保证了回答的可溯源性,又提供了便捷的原文查阅体验,让用户能轻松验证AI回答的准确性和相关上下文。

## 便捷引用导航
分块阅读器右上角设计了简洁实用的导航控制,用户可以通过这对按钮轻松在多个引用间切换浏览。导航区还直观显示当前查看的引用序号及总引用数量(如 "7/10"),帮助用户随时了解浏览进度和引用内容的整体规模。

## 引用质量评分
每条引用内容旁边都配有智能评分标签,直观展示该引用在所有知识片段中的相关性排名。用户只需将鼠标悬停在评分标签上,即可查看完整的评分详情,了解这段引用内容为何被AI选中以及其相关性的具体构成。

## 文档内容一键导出
分块阅读器贴心配备了内容导出功能,让有效信息不再流失。只要用户拥有相应知识库的阅读权限,便可通过简单点击将引用涉及的全文直接保存到本地设备。

# 进阶特性
## 灵活的可见度控制
FastGPT提供灵活的引用可见度设置,让知识共享既开放又安全。以免登录链接为例,管理员可精确控制外部访问者能看到的信息范围。
当设置为"仅引用内容可见"时,外部用户点击引用链接将只能查看 AI 引用的特定文本片段,而非完整原文档。如图所示,分块阅读器此时智能调整显示模式,仅呈现相关引用内容。
| | |
| --------------------------- | --------------------------- |
|  |  |
## 即时标注优化
在浏览过程中,授权用户可以直接对引用内容进行即时标注和修正,系统会智能处理这些更新而不打断当前的对话体验。所有修改过的内容会通过醒目的"已更新"标签清晰标识,既保证了引用的准确性,又维持了对话历史的完整性。
这一无缝的知识优化流程特别适合团队协作场景,让知识库能在实际使用过程中持续进化,确保AI回答始终基于最新、最准确的信息源。
## 智能文档性能优化
面对现实业务中可能包含成千上万分块的超长文档,FastGPT采用了先进的性能优化策略,确保分块阅读器始终保持流畅响应。
系统根据引用相关性排序和数据库索引进行智能加载管理,实现了"按需渲染"机制——根据索引排序和数据库 id,只有当用户实际需要查看的内容才会被加载到内存中。这意味着无论是快速跳转到特定引用,还是自然滚动浏览文档,都能获得丝滑的用户体验,不会因为文档体积庞大而出现卡顿或延迟。
这一技术优化使FastGPT能够轻松应对企业级的大规模知识库场景,让即使是包含海量信息的专业文档也能高效展示和查阅。
file: ./content/guide/build/evaluation.en.mdx
meta: {
"title": "App Evaluation (Beta)",
"description": "A quick overview of FastGPT app evaluation"
}
Starting from FastGPT v4.11.0, batch app evaluation is supported. By providing multiple QA pairs, the system automatically scores your app's responses, enabling quantitative assessment of app performance.
The system supports three evaluation metrics: answer accuracy, question relevance, and semantic accuracy. The current beta only includes answer accuracy — the remaining metrics will be added in future releases.
## Create an App Evaluation
### Go to the Evaluation Page

Navigate to the App Evaluation section under Workspace and click the "Create Task" button in the upper right corner.
### Fill in Evaluation Details

On the task creation page, provide the following:
* **Task Name**: A label to identify this evaluation
* **Evaluation Model**: The model used for scoring
* **Target App**: The app to be evaluated
### Prepare Evaluation Data

After selecting the target app, a button appears to download the CSV template. The template includes these fields:
* Global variables
* q (question)
* a (expected answer)
* Chat history
**Notes:**
* Maximum of 1,000 QA pairs
* Follow the template format when filling in data
Upload the completed file and click "Start Evaluation" to create the task.
## View Evaluation Results
### Evaluation List

The evaluation list shows all tasks with key information:
* **Progress**: Current execution status
* **Created By**: The user who created the task
* **Target App**: The app being evaluated
* **Start/End Time**: Execution time range
* **Overall Score**: The task's aggregate score
Use this to compare results across iterations as you improve your app.
### Evaluation Details

Click "View Details" to open the detail page:
**Task Overview**: The top section shows overall task information, including evaluation configuration and summary statistics.
**Detailed Results**: The bottom section lists each QA pair with its score, showing:
* User question
* Expected output
* App output
file: ./content/guide/build/evaluation.mdx
meta: {
"title": "应用评测(Beta)",
"description": "快速了解 FastGPT 应用评测功能"
}
FastGPT v4.11.0 版本开始支持应用批量评测功能。通过传入多组问答对,系统会对应用执行结果进行自动打分,实现应用运行效果的定量评估。
系统支持三种评估指标:回答准确性、问题相关性和语义准确性。当前测试版仅包含回答准确性这一个指标,其余指标将在后续版本中补充完善。
## 创建应用评测
### 进入评测页面

进入工作台下的应用评测目录,点击右上角的"创建任务"按钮。
### 填写评测信息

在创建任务页面中,需要填写以下信息:
* **评测任务名**:任务的标识名称
* **评测模型**:用于本次任务打分的模型
* **评测应用**:需要被打分的应用
### 准备评测数据

选择评测应用后,系统会弹出下载CSV模板的按钮。模板包含以下字段:
* 全局变量
* q(问题)
* a(标准答案)
* 历史记录
**注意事项:**
* 最多支持1000组问答对
* 请按照模板格式填写数据
填写完成后上传文件并点击"开始评测",即可创建一个应用评测任务
## 查看应用评测
### 评测列表

评测列表页面显示所有评测任务,包含以下关键信息:
* **进度**:当前评测任务的执行状态
* **执行人**:创建评测任务的用户
* **评测应用**:被评测的应用名称
* **开始时间/结束时间**:评测任务的执行时间范围
* **综合评分**:评测任务的整体得分
通过这些信息,可以清晰地比较每次应用改进后的效果。
### 评测详情

点击"查看详情"可进入评测任务的详情页面:
**任务概览**:页面顶部显示任务的整体信息,包括评测配置和统计结果。
**详细结果**:页面下方展示评测任务中的每一条问答对及其评分,可以查看:
* 用户问题
* 标准输出
* 应用输出
file: ./content/guide/build/faq.en.mdx
meta: {
"title": "App Building FAQ",
"description": "Common FastGPT app building questions, including simple apps, workflows, and plugins"
}
## Multi-Turn Classification
The Question Classification node has access to conversation context. When two consecutive questions are closely related, the model can usually classify them accurately based on their connection. For example, if a user asks "How do I use this feature?" followed by "What are the limitations?", the model leverages context to understand and respond correctly.
However, when consecutive questions have little relation to each other, classification accuracy may drop. To handle this, you can use a global variable to store the classification result. In subsequent classification steps, check the global variable first — if a result exists, reuse it; otherwise, let the model classify on its own.
Tip: Build batch test scripts to evaluate your question classification accuracy.
## Scheduled Execution Timing
If a user opens a shared link and stays on the page, scheduled execution still works as expected — it takes effect after the app is published and runs in the background.
## Changes Not Applied
After changing an app, click **Publish**. Chat and published channels only use the updated app configuration after publishing.
## Disable Markdown Formatting
Edit the Knowledge Base default prompt. The built-in standard template instructs the model to use Markdown. You can remove that requirement:
| | |
| ----------------------- | ----------------------- |
|  |  |
## Inconsistent App Results
Q: The app produces different results in debug mode vs. production, or when called via API.
A: This is usually caused by differences in context. Check the conversation logs, find the relevant entry, and compare the run details side by side.
| | | |
| ----------------------- | ----------------------- | ----------------------- |
|  |  |  |
The Knowledge Base response settings require a custom prompt. Without one, the default prompt (which includes Markdown formatting instructions) is used.
## Skip Classification for Follow-Ups
Scenario: A workflow starts with a Question Classification node that routes to different branches, each with its own Knowledge Base and AI Chat. After the first AI response, you want subsequent questions to skip classification and go straight to the Knowledge Base with chat history as context.
Solution: Add a condition check — if it's the first message (history count is 0), route through Question Classification. Otherwise, go directly to the Knowledge Base and AI Chat.
## Formula Rendering Issues
Add a prompt to guide the model to output formulas in LaTeX/Markdown format:
```bash
Latex inline: \(x^2\)
Latex block: $$e=mc^2$$
```
file: ./content/guide/build/faq.mdx
meta: {
"title": "常见问题",
"description": "FastGPT 应用构建常见问题,包括简易应用、工作流和插件"
}
## 多轮对话分类
问题分类节点具有获取上下文信息的能力,当处理两个关联性较大的问题时,模型的判断准确性往往依赖于这两个问题之间的联系和模型的能力。例如,当用户先问“我该如何使用这个功能?”接着又询问“这个功能有什么限制?”时,模型借助上下文信息,就能够更精准地理解并响应。
但是,当连续问题之间的关联性较小,模型判断的准确度可能会受到限制。在这种情况下,我们可以引入全局变量的概念来记录分类结果。在后续的问题分类阶段,首先检查全局变量是否存有分类结果。如果有,那么直接沿用该结果;若没有,则让模型自行判断。
建议:构建批量运行脚本进行测试,评估问题分类的准确性。
## 定时执行触发时机
系统编排配置中的定时执行,如果用户打开分享的连接,停留在那个页面,定时执行触发问题:
定时执行会在应用发布后生效,会在后台生效。
## 修改后未生效
应用变更后,需要点击发布后,聊天和发布渠道的使用才会更新应用。
## 取消 Markdown 输出
修改知识库默认提示词, 默认用的是标准模板提示词,会要求按 Markdown 输出,可以去除该要求:
| | |
| ----------------------- | ----------------------- |
|  |  |
## 不同来源效果不一致
Q: 应用在调试和正式发布后,效果不一致;在 API 调用时,效果不一致。
A: 通常是由于上下文不一致导致,可以在对话日志中,找到对应的记录,并查看运行详情来进行比对。
| | | |
| ----------------------- | ----------------------- | ----------------------- |
|  |  |  |
在针对知识库的回答要求里有, 要给它配置提示词,不然他就是默认的,默认的里面就有该语法。
## 后续问题跳过分类节点
做个判断器,如果是初次开始对话也就是历史记录为 0,就走问题分类;不为零直接走知识库和 ai。
## 公式无法正常显示
添加相关提示词,引导模型按 Markdown 输出公式
```bash
Latex inline: \(x^2\)
Latex block: $$e=mc^2$$
```
file: ./content/guide/dataset/collection_tags.en.mdx
meta: {
"title": "Knowledge Base Collection Tags",
"description": "How to use collection tags in FastGPT Knowledge Base"
}
Collection tags are a commercial-edition feature in FastGPT. They let you tag and categorize data collections within a knowledge base for more efficient data management.
You can also use tags as collection filters during knowledge base searches for more precise results.
| | | |
| -------------------------------- | -------------------------------- | -------------------------------- |
|  |  |  |
## Basic Tag Operations
On the knowledge base detail page, you can manage tags with the following operations:
* Create a tag
* Rename a tag
* Delete a tag
* Assign a tag to multiple collections
* Add multiple tags to a single collection
You can also filter collections by tags.
## Collection Filtering in Knowledge Base Search
Tags can be used to filter collections during knowledge base searches by filling in the "Collection Filter" field. Here's an example:
```json
{
"tags": {
"$and": ["Tag 1","Tag 2"],
"$or": ["When $and tags are present, $and takes effect and $or is ignored"]
},
"createTime": {
"$gte": "YYYY-MM-DD HH:mm format, matches collections created after this time",
"$lte": "YYYY-MM-DD HH:mm format, matches collections created before this time. Can be used together with $gte"
}
}
```
Two important notes:
* Tag values can be a `string` tag name or `null`, where `null` represents collections with no tags assigned
* There are two filter condition types: `$and` and `$or`. When both are set, only `$and` takes effect
file: ./content/guide/dataset/collection_tags.mdx
meta: {
"title": "知识库集合标签",
"description": "FastGPT 知识库集合标签使用说明"
}
知识库集合标签是 FastGPT 商业版特有功能。它允许你对知识库中的数据集合添加标签进行分类,更高效地管理知识库数据。
而进一步可以在问答中,搜索知识库时添加集合过滤,实现更精确的搜索。
| | | |
| -------------------------------- | -------------------------------- | -------------------------------- |
|  |  |  |
## 标签基础操作说明
在知识库详情页面,可以对标签进行管理,可执行的操作有
* 创建标签
* 修改标签名
* 删除标签
* 将一个标签赋给多个数据集合
* 给一个数据集合添加多个标签
也可以利用标签对数据集合进行筛选
## 知识库搜索-集合过滤说明
利用标签可以在知识库搜索时,通过填写「集合过滤」这一栏来实现更精确的搜索,具体的填写示例如下
```json
{
"tags": {
"$and": ["标签 1","标签 2"],
"$or": ["有 $and 标签时,and 生效,or 不生效"]
},
"createTime": {
"$gte": "YYYY-MM-DD HH:mm 格式即可,集合的创建时间大于该时间",
"$lte": "YYYY-MM-DD HH:mm 格式即可,集合的创建时间小于该时间,可和 $gte 共同使用"
}
}
```
在填写时有两个注意的点,
* 标签值可以为 `string` 类型的标签名,也可以为 `null`,而 `null` 代表着未设置标签的数据集合
* 标签过滤有 `$and` 和 `$or` 两种条件类型,在同时设置了 `$and` 和 `$or` 的情况下,只有 `$and` 会生效
file: ./content/guide/dataset/dataset_engine.en.mdx
meta: {
"title": "Knowledge Base Search Methods and Parameters",
"description": "This section covers FastGPT's knowledge base architecture, including its QA storage format and multi-vector mapping, to help you build better knowledge bases. It also explains each search parameter. This guide focuses on practical usage rather than in-depth theory."
}
## Understanding Vectors
FastGPT uses an Embedding-based RAG approach for its knowledge base. To use FastGPT effectively, you need a basic understanding of how `Embedding` vectors work and their characteristics.
Human text, images, and other media cannot be directly understood by computers. To determine whether two pieces of text are similar or related, they typically need to be converted into a computer-readable format — vectors are one such method.
A vector is essentially an array of numbers. The "distance" between two vectors can be calculated using mathematical formulas — the smaller the distance, the more similar the vectors. This maps back to text, images, and other media to measure similarity between them. Vector search leverages this principle.
Since text comes in many types with countless combinations, exact matching is hard to guarantee when converting to vectors for similarity comparison. In vector-based knowledge bases, a `top-k` recall approach is typically used — finding the top `k` most similar results and passing them to an LLM for further `semantic evaluation`, `logical reasoning`, and `summarization`, enabling knowledge base Q\&A. This makes vector search the most critical step in the process.
Many factors affect vector search accuracy, including: vector model quality, data quality (length, completeness, diversity), and retriever precision (the speed vs. accuracy tradeoff). Search query quality is equally important.
Retriever precision is relatively straightforward to address, and training vector models is more complex, so optimizing data and query quality becomes a key focus.
### Improving Vector Search Accuracy
1. Better tokenization and chunking: When a text segment has complete and singular structure and semantics, accuracy improves. Many systems optimize their tokenizers to preserve data completeness.
2. Streamline `index` content by reducing vector content length: Shorter, more precise `index` content improves search accuracy, though it may narrow the search scope. Best suited for scenarios requiring strict answers.
3. Increase `index` quantity: Add multiple `index` entries for the same `chunk` to improve recall.
4. Optimize search queries: In practice, user questions are often vague or incomplete. Refining the query (search term) can significantly improve accuracy.
5. Fine-tune vector models: Off-the-shelf vector models are general-purpose and may underperform in specific domains. Fine-tuning can greatly improve domain-specific search results.
## FastGPT Knowledge Base Architecture
### Data Storage Structure
In FastGPT, a knowledge base consists of three parts: libraries, collections, and data entries. A collection can be thought of as a "file." A library can contain multiple collections, and a collection can contain multiple data entries. The smallest searchable unit is the library — searches span the entire library. Collections are only for organizing and managing data and do not affect search results (at least for now).

### Vector Storage Structure
FastGPT uses `PostgreSQL`'s `PG Vector` extension as the vector retriever, with `HNSW` indexing. `PostgreSQL` is used solely for vector search (this engine can be swapped for other databases), while `MongoDB` handles all other data storage.
In `MongoDB`'s `dataset.datas` collection, vector source data is stored along with an `indexes` field that records corresponding vector IDs. This is an array, meaning a single data entry can map to multiple vectors. In addition to default text indexes, image content can also generate image description indexes or image vector indexes when the configured models support it.
In `PostgreSQL`, a `vector` field stores the vectors. During search, vectors are recalled first, then their IDs are used to look up the original data in `MongoDB`. If multiple vectors map to the same source data, they are merged and the highest vector score is used.

### Purpose and Usage of Multi-Vector Mapping
In a single vector, content length and semantic richness are often at odds. FastGPT uses multi-vector mapping to map a single data entry to multiple vectors, preserving both data completeness and semantic richness.
You can add multiple vectors to a longer text so that if any one vector is matched during search, the entire data entry is recalled.
This means you can continuously improve data chunk accuracy through annotation.
### Overall Search Strategy
A Knowledge Base search is not simply "user question -> vector database -> result." Depending on the input and search parameters, FastGPT combines text, images, semantic recall, full-text recall, query optimization, and reranking, then fuses multiple result paths into the final quoted content.
1. Use `Query Optimization` for coreference resolution and query expansion, improving multi-turn conversation search capability and semantic richness.
2. Use `Semantic Search`, `Full-Text Search`, or `Hybrid Search` to recall candidate content.
3. If the input contains images, use image description search or image vector search depending on model capability.
4. Use `RRF` (Reciprocal Rank Fusion) to merge results from multiple search channels.
5. Use `Rerank` for secondary sorting to improve text result relevance.
6. Apply similarity filtering and the reference limit to produce the final quoted content sent to the model.

### Image Search Method
In Knowledge Base search, images can participate in retrieval in addition to text questions. FastGPT handles images differently depending on the configured model capabilities.
Image search mainly works in two ways:
1. Image description search: If an available vision model is configured, the system can understand the image first, generate a text description, and use that description in regular text retrieval.
2. Image vector search: If the selected embedding model supports image input, the system can generate vectors for images directly and match them against image vectors in the Knowledge Base.
Image search is not a separate system outside the Knowledge Base. It adds an image-input path to the existing Knowledge Base search pipeline.
Common usage patterns include:
* Text-to-image search: enter text to find semantically related image content.
* Image-to-image search: enter an image to find visually or semantically similar image content.
* Text + image search: enter both text and an image, using the text question as an additional constraint on image search results.
Image search quality usually depends on image clarity, whether the image content is easy for the model to understand, whether a vision model is configured, and whether the embedding model supports image vectors.
Whether an image can be retrieved does not only depend on uploading an image at search time. It also depends on which indexes were created during ingestion:
| Knowledge Base capability | Text-only query | Image-only query | Text + image query |
| ------------------------------------------------- | -------------------------------------------- | --------------------------------------------------------------- | ------------------------------------------------- |
| Regular embedding model, no vision model | Normal text retrieval | Usually unavailable | Mainly uses the text part |
| Regular embedding model with a vision model | Normal text retrieval | Converts the image into a description, then uses text retrieval | Text + image description participate in retrieval |
| Image-capable embedding model, no vision model | Normal text retrieval | Image vector retrieval | Text retrieval + image vector retrieval |
| Image-capable embedding model with a vision model | Text retrieval, including image descriptions | Image description + image vector retrieval | Text + image description + image vector retrieval |
So when image-to-image search performs poorly, do not only adjust search parameters. Also check whether the Knowledge Base is configured with a vision model or an image-capable embedding model, and whether valid image indexes were generated during ingestion.
### Result Ranking and Fusion
FastGPT fuses results from different recall paths instead of using only one path. Common paths include text vector recall, full-text recall, image description recall, image vector recall, and reranked results.
Keep these points in mind:
1. `Semantic Search` relies more on vector similarity and is better for natural-language questions and semantically related content.
2. `Full-Text Search` relies more on keyword matches and is better for IDs, model numbers, proper nouns, error codes, and other exact queries.
3. `Hybrid Search` uses both semantic recall and full-text recall, then merges the results with `RRF`.
4. `Rerank` re-sorts candidate text results and works best when the question is clear and there are enough candidates.
5. Image search adds image description or image vector results, which are then fused with text-side results.
This means final quoted content may not be strictly sorted by a single vector similarity score. Content matched by multiple recall paths is usually more likely to rank higher.
## Search Parameters
| | | |
| ------------------------------------- | ------------------------------------- | ------------------------------------- |
|  |  |  |
### Search Modes
#### Semantic Search
Semantic search calculates the vector distance between the user's query and knowledge base content to determine "similarity" — mathematical similarity, not linguistic.
Pros:
* Understands similar semantics
* Cross-language understanding (e.g., Chinese query matching English content)
* Multimodal understanding (text, images, etc., depending on model capability)
Cons:
* Depends on model training quality
* Inconsistent accuracy
* Affected by keywords and sentence completeness
#### Full-Text Search
Uses traditional full-text search. Best for finding key subjects, predicates, and other specific terms.
#### Hybrid Search
Combines vector search and full-text search, merging results using the RRF formula. Generally produces richer and more accurate results.
Since hybrid search covers a large range and cannot directly filter by similarity, a rerank model is typically used to re-sort results and filter by rerank scores.
#### Result Reranking
Uses a `ReRank` model to re-sort search results. In most cases, this significantly improves accuracy. Rerank models work better with complete questions (with proper subjects and predicates), so query optimization is usually applied before search and reranking. Reranking produces a score between `0-1` representing the relevance between the search content and the query — this score is typically more accurate than vector similarity scores and can be used for filtering.
FastGPT uses `RRF` to merge rerank results, vector search results, and full-text search results into the final output.
### Search Filters
#### Reference Limit
The maximum number of `tokens` to reference per search.
Instead of using `top k`, we found that in mixed knowledge bases (Q\&A + document), different `chunk` lengths vary significantly, making `top k` results unstable. Using a `token` limit provides more consistent control.
#### Minimum Relevance
A value between `0-1` that filters out low-relevance search results.
This only takes effect when using `Semantic Search` or `Result Reranking`.
Note that minimum relevance is a filtering threshold, not the final sorting rule. After query optimization, hybrid search, image search, or result reranking is enabled, final results may be fused from multiple recall paths and may not be strictly sorted by a single vector similarity score.
### Query Optimization
#### Background
In RAG, we need to perform embedding searches against the database based on the input query to find similar content (i.e., knowledge base search).
During search — especially in multi-turn conversations — follow-up questions often fail to find relevant content because knowledge base search only uses the "current" question. Consider this example:

When the user asks "What's the second point?", the system searches for "What's the second point?" in the knowledge base, which returns nothing useful. The actual query should be "What is the QA structure?". This is why we need a Query Optimization module to complete the user's current question, enabling the knowledge base search to find relevant content. Here's the result after optimization:

#### How It Works
Before performing `data retrieval`, the model first performs `coreference resolution` and `query expansion`. This resolves ambiguous references and enriches the query's semantic content. You can view the optimized query in the conversation details after each interaction.
Query Optimization adds an extra model call before the actual search. It often improves retrieval in multi-turn conversations, but it also increases total latency. If the current question is already clear, or response speed is more important, decide whether to enable it based on actual results.
### Common Tuning Tips
If search results are not as expected, start from the symptom. Avoid changing every parameter at once.
| Symptom | What to check or adjust first |
| -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| No results found | Confirm the data has finished training; lower the minimum relevance; increase the reference limit; check whether the question is too short or missing a subject |
| Results are too broad or off-topic | Raise the minimum relevance; reduce the reference limit; improve chunking; check whether recalled content contains too many unrelated chunks |
| IDs, model numbers, or proper nouns are inaccurate | Use full-text or hybrid search; reduce semantic search weight; avoid overusing query optimization for exact ID queries |
| Natural-language questions do not retrieve well | Use semantic or hybrid search; enable query optimization; add more accurate indexes to the data |
| Query optimization makes search slower | Query optimization adds an extra model call. Use a faster optimization model, or enable it only for follow-up questions and short queries |
| Rerank still gives poor ordering | Check whether the user question is complete; make sure enough candidates are recalled; adjust minimum relevance and reference limit |
| Image-to-image search is weak | Confirm the embedding model supports image input; confirm image vector indexes were generated during ingestion; check whether the image is clear and has an obvious subject |
| Text + image search is unstable | Clarify whether text or image should be more important; if you only want visual similarity, reduce extra text constraints |
file: ./content/guide/dataset/dataset_engine.mdx
meta: {
"title": "知识库搜索方案和参数",
"description": "本节会详细介绍 FastGPT 知识库结构设计,理解其 QA 的存储格式和多向量映射,以便更好的构建知识库。同时会介绍每个搜索参数的功能。这篇介绍主要以使用为主,详细原理不多介绍。"
}
## 理解向量
FastGPT 采用了 RAG 中的 Embedding 方案构建知识库,要使用好 FastGPT 需要简单的理解 `Embedding` 向量是如何工作的及其特点。
人类的文字、图片等媒介是无法直接被计算机理解的,要想让计算机理解两段文字是否有相似性、相关性,通常需要将它们转成计算机可以理解的语言,向量是其中的一种方式。
向量可以简单理解为一个数字数组,两个向量之间可以通过数学公式得出一个 `距离`,距离越小代表两个向量的相似度越大。从而映射到文字、图片等媒介上,可以用来判断两个媒介之间的相似度。向量搜索便是利用了这个原理。
而由于文字是有多种类型,并且拥有成千上万种组合方式,因此在转成向量进行相似度匹配时,很难保障其精确性。在向量方案构建的知识库中,通常使用 `topk` 召回的方式,也就是查找前 `k` 个最相似的内容,丢给大模型去做更进一步的 `语义判断`、`逻辑推理` 和 `归纳总结`,从而实现知识库问答。因此,在知识库问答中,向量搜索的环节是最为重要的。
影响向量搜索精度的因素非常多,主要包括:向量模型的质量、数据的质量(长度,完整性,多样性)、检索器的精度(速度与精度之间的取舍)。与数据质量对应的就是检索词的质量。
检索器的精度比较容易解决,向量模型的训练略复杂,因此数据和检索词质量优化成了一个重要的环节。
### 提高向量搜索精度的方法
1. 更好分词分段:当一段话的结构和语义是完整的,并且是单一的,精度也会提高。因此,许多系统都会优化分词器,尽可能的保障每组数据的完整性。
2. 精简 `index` 的内容,减少向量内容的长度:当 `index` 的内容更少,更准确时,检索精度自然会提高。但与此同时,会牺牲一定的检索范围,适合答案较为严格的场景。
3. 丰富 `index` 的数量,可以为同一个 `chunk` 内容增加多组 `index`。
4. 优化检索词:在实际使用过程中,用户的问题通常是模糊的或是缺失的,并不一定是完整清晰的问题。因此优化用户的问题(检索词)很大程度上也可以提高精度。
5. 微调向量模型:由于市面上直接使用的向量模型都是通用型模型,在特定领域的检索精度并不高,因此微调向量模型可以很大程度上提高专业领域的检索效果。
## FastGPT 构建知识库方案
### 数据存储结构
在 FastGPT 中,整个知识库由库、集合和数据 3 部分组成。集合可以简单理解为一个 `文件`。一个 `库` 中可以包含多个 `集合`,一个 `集合` 中可以包含多组 `数据`。最小的搜索单位是 `库`,也就是说,知识库搜索时,是对整个 `库` 进行搜索,而集合仅是为了对数据进行分类管理,与搜索效果无关。(起码目前还是)

### 向量存储结构
FastGPT 采用了 `PostgresSQL` 的 `PG Vector` 插件作为向量检索器,索引为 `HNSW`。且 `PostgresSQL` 仅用于向量检索(该引擎可以替换成其它数据库),`MongoDB` 用于其他数据的存取。
在 `MongoDB` 的 `dataset.datas` 表中,会存储向量原数据的信息,同时有一个 `indexes` 字段,会记录其对应的向量 ID,这是一个数组,也就是说,一组数据可以对应多个向量。除默认文本索引外,如果模型能力支持,图片内容也可以生成图片描述索引或图片向量索引。
在 `PostgresSQL` 的表中,设置一个 `vector` 字段用于存储向量。在检索时,会先召回向量,再根据向量的 ID,去 `MongoDB` 中寻找原数据内容,如果对应了同一组原数据,则进行合并,向量得分取最高得分。

### 多向量的目的和使用方式
在一组向量中,内容的长度和语义的丰富度通常是矛盾的,无法兼得。因此,FastGPT 采用了多向量映射的方式,将一组数据映射到多组向量中,从而保障数据的完整性和语义的丰富度。
你可以为一组较长的文本,添加多组向量,从而在检索时,只要其中一组向量被检索到,该数据也将被召回。
意味着,你可以通过标注数据块的方式,不断提高数据块的精度。
### 整体检索方案
一次知识库检索不是简单的“用户问题 -> 向量库 -> 返回结果”。FastGPT 会根据输入内容和搜索参数,将文本、图片、语义召回、全文召回、问题优化和重排等能力组合起来,最后再把多路结果融合成引用内容。
1. 通过 `问题优化` 实现指代消除和问题扩展,从而增加连续对话的检索能力以及语义丰富度。
2. 通过 `语义检索`、`全文检索` 或 `混合检索` 召回候选内容。
3. 如果输入中包含图片,会根据模型能力额外进行图片描述检索或图片向量检索。
4. 通过 `RRF` 合并方式,综合多个渠道的检索效果。
5. 通过 `Rerank` 来二次排序,提高文本结果的相关性。
6. 最终经过相似度过滤和引用上限裁剪,得到返回给模型的引用内容。

### 图片检索方案
在知识库搜索中,除了文本问题外,也可以让图片参与检索。FastGPT 会根据当前模型能力,对图片进行不同处理。
图片检索主要有两种方式:
1. 图片描述检索:如果配置了可用的视觉模型,系统可以先理解图片内容,并生成一段文本描述,再使用这段描述参与普通文本检索。
2. 图片向量检索:如果当前向量模型支持图片输入,系统可以直接对图片生成向量,并与知识库中的图片向量进行相似度匹配。
因此,图片检索不是独立于知识库之外的一套能力,而是在原有知识库搜索链路上增加了图片输入的处理路径。
常见使用方式包括:
* 文搜图:输入文字,搜索语义相关的图片内容。
* 图搜图:输入图片,搜索视觉或语义相似的图片内容。
* 图文混合搜索:同时输入文字和图片,让文字问题对图片搜索结果进行补充约束。
图片检索效果通常取决于图片清晰度、图片内容是否容易被模型理解、是否配置了视觉模型,以及向量模型是否支持图片向量。
需要注意的是,图片能否被检索到,不只取决于搜索时是否上传了图片,也取决于入库时是否建立了对应索引:
| 知识库能力 | 纯文本查询 | 纯图片查询 | 图文混合查询 |
| --------------- | ------------- | --------------- | -------------------- |
| 普通向量模型,无视觉模型 | 正常文本检索 | 基本不可用 | 主要使用文字部分 |
| 普通向量模型,有视觉模型 | 正常文本检索 | 图片先转成描述,再参与文本检索 | 文字 + 图片描述共同参与检索 |
| 支持图片的向量模型,无视觉模型 | 正常文本检索 | 图片向量检索 | 文本检索 + 图片向量检索 |
| 支持图片的向量模型,有视觉模型 | 文本检索,也可命中图片描述 | 图片描述 + 图片向量双路检索 | 文本 + 图片描述 + 图片向量多路检索 |
所以,图搜图效果不理想时,除了调整搜索参数,也要确认当前知识库是否配置了视觉模型或支持图片的向量模型,以及图片入库时是否生成了有效的图片索引。
### 结果排序与融合
FastGPT 会把不同召回路径的结果进行融合,而不是简单采用某一路结果。常见路径包括文本向量召回、全文召回、图片描述召回、图片向量召回和重排结果。
因此,最终排序需要这样理解:
1. `语义检索` 更依赖向量相似度,适合自然语言问题和语义相近内容。
2. `全文检索` 更依赖关键词命中,适合编号、型号、专有名词、错误码等精确查询。
3. `混合检索` 会同时使用语义召回和全文召回,再通过 `RRF` 融合结果。
4. `Rerank` 会对候选文本进行二次排序,更适合文本问题明确、候选结果较多的场景。
5. 图片检索会额外引入图片描述或图片向量结果,最终和文本侧结果一起融合。
这意味着,最终引用内容不一定严格按照单一向量相似度排序。某条内容如果同时被多路召回命中,通常会更容易排在前面。
## 搜索参数
| | | |
| ------------------------------------- | ------------------------------------- | ------------------------------------- |
|  |  |  |
### 搜索模式
#### 语义检索
语义检索是通过向量距离,计算用户问题与知识库内容的距离,从而得出“相似度”,当然这并不是语文上的相似度,而是数学上的。
优点:
* 相近语义理解
* 跨多语言理解(例如输入中文问题匹配英文知识点)
* 多模态理解(文本、图片等,取决于模型能力)
缺点:
* 依赖模型训练效果
* 精度不稳定
* 受关键词和句子完整度影响
#### 全文检索
采用传统的全文检索方式。适合查找关键的主谓语等。
#### 混合检索
同时使用向量检索和全文检索,并通过 RRF 公式进行两个搜索结果合并,一般情况下搜索结果会更加丰富准确。
由于混合检索后的查找范围很大,并且无法直接进行相似度过滤,通常需要进行利用重排模型进行一次结果重新排序,并利用重排的得分进行过滤。
#### 结果重排
利用 `ReRank` 模型对搜索结果进行重排,绝大多数情况下,可以有效提高搜索结果的准确率。不过,重排模型与问题的完整度(主谓语齐全)有一些关系,通常会先走问题优化后再进行搜索 - 重排。重排后可以得到一个 `0-1` 的得分,代表着搜索内容与问题的相关度,该分数通常比向量的得分更加精确,可以根据得分进行过滤。
FastGPT 会使用 `RRF` 对重排结果、向量搜索结果、全文检索结果进行合并,得到最终的搜索结果。
### 搜索过滤
#### 引用上限
每次搜索最多引用 `n` 个 `tokens` 的内容。
之所以不采用 `top k`,是发现在混合知识库(问答库、文档库)时,不同 `chunk` 的长度差距很大,会导致 `top k` 的结果不稳定,因此采用了 `tokens` 的方式进行引用上限的控制。
#### 最低相关度
一个 `0-1` 的数值,会过滤掉一些低相关度的搜索结果。
该值仅在 `语义检索` 或使用 `结果重排` 时生效。
需要注意的是,最低相关度是过滤阈值,不是最终排序规则。开启问题优化、混合检索、图片检索或结果重排后,最终结果可能会经过多路召回融合,不一定严格按照单一向量相似度排序。
### 问题优化
#### 背景
在 RAG 中,我们需要根据输入的问题去数据库里执行 embedding 搜索,查找相关的内容,从而查找到相似的内容(简称知识库搜索)。
在搜索的过程中,尤其是连续对话的搜索,我们通常会发现后续的问题难以搜索到合适的内容,其中一个原因是知识库搜索只会使用“当前”的问题去执行。看下面的例子:

用户在提问“第二点是什么”的时候,只会去知识库里查找“第二点是什么”,压根查不到内容。实际上需要查询的是“QA 结构是什么”。因此我们需要引入一个【问题优化】模块,来对用户当前的问题进行补全,从而使得知识库搜索能够搜索到合适的内容。使用补全后效果如下:

#### 实现方式
在进行 `数据检索` 前,会先让模型进行 `指代消除` 与 `问题扩展`,一方面可以可以解决指代对象不明确问题,同时可以扩展问题的语义丰富度。你可以通过每次对话后的对话详情,查看补全的结果。
问题优化会在正式检索前增加一次模型调用,因此通常会提升连续对话检索效果,但也会增加整体耗时。如果当前问题本身已经非常明确,或对响应速度要求更高,可以根据实际效果决定是否开启。
### 常见调参建议
如果搜索结果不符合预期,可以先根据现象定位问题,不建议一次性调整所有参数。
| 现象 | 优先检查和调整 |
| -------------- | ---------------------------------------------- |
| 搜不到内容 | 确认数据是否已完成训练;适当降低最低相关度;提高引用上限;检查问题是否过短或缺少主体 |
| 结果太泛、答非所问 | 提高最低相关度;减少引用上限;优化数据分块;检查召回内容是否包含过多无关片段 |
| 编号、型号、专有名词搜不准 | 使用全文检索或混合检索;降低语义检索权重;避免对精确编号类问题过度使用问题优化 |
| 自然语言问法搜不准 | 使用语义检索或混合检索;开启问题优化;补充更准确的数据索引 |
| 开启问题优化后变慢 | 问题优化会额外调用模型,可以换更快的优化模型,或只在多轮追问、短问题场景中开启 |
| Rerank 后仍然排序不准 | 确认用户问题是否完整;检查召回候选是否足够;适当调整最低相关度和引用上限 |
| 图搜图效果弱 | 确认向量模型是否支持图片输入;确认入库时是否生成图片向量索引;检查图片是否清晰、主体是否明确 |
| 图文混合结果不稳定 | 明确文字和图片哪个更重要;如果只想找视觉相似图片,减少额外文字约束 |
file: ./content/guide/dataset/faq.en.mdx
meta: {
"title": "Knowledge Base Usage",
"description": "Common Knowledge Base usage questions"
}
## Garbled File Content
Re-save the file with UTF-8 encoding.
## Processing Model vs. Index Model
* **File Processing Model**: Used for **Enhanced Processing** and **Q\&A Splitting** during data ingestion. Enhanced Processing generates related questions and summaries; Q\&A Splitting generates question-answer pairs.
* **Index Model**: Used for vectorization — it processes and organizes text data into a structure optimized for fast retrieval.
## Excel File Import
Yes. You can upload xlsx and other spreadsheet formats, not just CSV.
## Token Calculation
All token counts use the GPT-3.5 tokenizer as the standard.
## Restore a Rerank Model

Add the rerank model configuration in your `config.json` file, then you'll be able to select it again.
## Data Retention After Expiration
On the free plan, Knowledge Base data is cleared after 30 days of inactivity (no login). Apps are not affected. Paid plans automatically downgrade to the free plan upon expiration.

## Too Many Results Interrupt Answers
FastGPT calculates the maximum response length as:
Max Response = min(Configured Max Response, Max Context Window - History)
For example, with an 18K context model, input + output share the same window. As output grows, available input shrinks.
To fix this:
1. Check your configured max response (response limit) setting.
2. Reduce input to free up space for output — specifically, reduce the number of chat history turns included in the workflow.
Where to find the max response setting:


For self-hosted deployments, you can reserve headroom when configuring model context limits. For example, set a 128K model to 120K — the remaining space will be allocated to output.
## Chat History Context Limits
FastGPT calculates the maximum response length as:
Max Response = min(Configured Max Response, Max Context Window - History)
For example, with an 18K context model, input + output share the same window. As output grows, available input shrinks.
To fix this:
1. Check your configured max response (response limit) setting.
2. Reduce input to free up space for output — specifically, reduce the number of chat history turns included in the workflow.
Where to find the max response setting:


For self-hosted deployments, you can reserve headroom when configuring model context limits. For example, set a 128K model to 120K — the remaining space will be allocated to output.
file: ./content/guide/dataset/faq.mdx
meta: {
"title": "常见问题",
"description": "知识库常见问题"
}
## 文件解析失败
未打开 PDF 增强解析。如果在上传文件设置参数时,没有打开【PDF 增强解析】设置时,需要在 Admin 后台正确配置 OCR 模块以支持增强解析。
## 文件中文乱码
将文件另存为 UTF-8 编码格式。
## 文件处理模型与索引模型
* **文件处理模型**:用于数据处理的【增强处理】和【问答拆分】。在【增强处理】中,生成相关问题和摘要,在【问答拆分】中执行问答对生成。
* **索引模型**:用于向量化,即通过对文本数据进行处理和组织,构建出一个能够快速查询的数据结构。
## Excel 文件导入
xlsx 等都可以上传的,不止支持 CSV。
## Tokens 计算方式
统一按 gpt3.5 标准。
## 恢复重排模型

config.json 文件里面配置后就可以勾选重排模型
## 套餐到期后的数据保留
免费版是三十天不登录后清空知识库,应用不会动。其他付费套餐到期后自动切免费版。
## 知识库结果过多导致回答中断
FastGPT 回复长度计算公式:
最大回复=min(配置的最大回复(内置的限制),最大上下文(输入和输出的总和)- 历史记录)
18K 模型 ->输入与输出的和
输出增多 ->输入减小
所以可以:
1. 检查配置的最大回复(回复上限)
2. 减小输入来增大输出,即减小历史记录,在工作流其实也就是“聊天记录”
配置的最大回复:


另外私有化部署的时候,后台配模型参数,可以在配置最大上文时,预留一些空间,比如 128000 的模型,可以只配置 120000, 剩余的空间后续会被安排给输出
## 聊天记录触发上下文限制
FastGPT 回复长度计算公式:
最大回复=min(配置的最大回复(内置的限制),最大上下文(输入和输出的总和)- 历史记录)
18K 模型 ->输入与输出的和
输出增多 ->输入减小
所以可以:
1. 检查配置的最大回复(回复上限)
2. 减小输入来增大输出,即减小历史记录,在工作流其实也就是“聊天记录”
配置的最大回复:


另外,私有化部署的时候,后台配模型参数,可以在配置最大上文时,预留一些空间,比如 128000 的模型,可以只配置 120000, 剩余的空间后续会被安排给输出。
## 知识库页面闪烁
未配置索引模型,补齐索引模型配置。
file: ./content/guide/dataset/rag.en.mdx
meta: {
"title": "Knowledge Base Fundamentals",
"description": "This section covers the core mechanisms, application scenarios, advantages, and limitations of the RAG model in generation tasks."
}
[RAG Documentation](https://huggingface.co/docs/transformers/model_doc/rag)
# 1. Introduction
As natural language processing (NLP) technology has advanced rapidly, generative language models (such as GPT and BART) have excelled at text generation tasks, particularly in language generation and context understanding. However, purely generative models have inherent limitations when handling factual tasks. Since these models rely on fixed pre-training data, they may "hallucinate" — fabricating information when answering questions that require up-to-date or real-time knowledge, leading to inaccurate or unfounded results. Additionally, generative models often struggle with long-tail questions and complex reasoning tasks due to a lack of domain-specific external knowledge.
Meanwhile, retrieval models (Retrievers) can quickly locate relevant information across massive document collections, addressing factual query needs. However, traditional retrieval models (such as BM25) often return isolated results when facing ambiguous queries or cross-domain questions, and cannot generate coherent natural language answers. Without contextual reasoning capabilities, the answers they produce tend to lack coherence and completeness.
To address the shortcomings of both approaches, Retrieval-Augmented Generation (RAG) was developed. RAG combines the strengths of generative and retrieval models by fetching relevant information from external knowledge bases in real time and incorporating it into the generation process. This ensures that generated text is both contextually coherent and factually grounded. This hybrid architecture performs particularly well in intelligent Q\&A, information retrieval and reasoning, and domain-specific content generation.
## 1.1 Definition of RAG
RAG is a hybrid architecture that combines information retrieval with generative models. First, the retriever fetches content fragments relevant to the user's query from an external knowledge base or document collection. Then, the generator produces natural language output based on these retrieved fragments, ensuring the output is information-rich, highly relevant, and accurate.
# 2. Core Mechanisms of RAG
RAG models consist of two main modules: the Retriever and the Generator. These modules work together to ensure generated text contains relevant external knowledge while maintaining natural, fluent language.
## 2.1 Retriever
The retriever's primary task is to fetch the most relevant content from an external knowledge base or document collection for a given input query. Common techniques in RAG include:
* Vector retrieval: Using models like BERT to convert documents and queries into vector space representations, then matching them via similarity calculations. Vector retrieval excels at capturing semantic similarity rather than relying solely on lexical matching.
* Traditional retrieval algorithms: Such as BM25, which uses term frequency and inverse document frequency (TF-IDF) weighted scoring to rank and retrieve documents. BM25 works well for straightforward keyword matching tasks.
The retriever in RAG provides contextual background for the generator, enabling it to produce more relevant answers based on the retrieved document fragments.
## 2.2 Generator
The generator is responsible for producing the final natural language output. Common generators in RAG systems include:
* BART: A sequence-to-sequence model focused on text generation, capable of improving output quality through various noise-handling techniques.
* GPT series: Pre-trained language models that excel at generating fluent, natural text, particularly strong in generation tasks thanks to large-scale training data.
After receiving document fragments from the retriever, the generator uses them as context alongside the input query to produce relevant, natural text answers. This ensures the output draws on both existing knowledge and the latest external information.
## 2.3 RAG Workflow
The RAG model workflow can be summarized as follows:
1. Input query: The user submits a question, which the system converts to a vector representation.
2. Document retrieval: The retriever extracts the most relevant document fragments from the knowledge base, typically using vector retrieval or traditional techniques like BM25.
3. Answer generation: The generator receives the retrieved fragments and produces a natural language answer based on both the original query and the retrieved context, providing richer, more contextually relevant responses.
4. Output: The generated answer is returned to the user, ensuring they receive an accurate response grounded in relevant, up-to-date information.
# 3. How RAG Works
## 3.1 Retrieval Phase
In RAG, the user's query is first converted to a vector representation, then vector search is performed against the knowledge base. The retriever typically uses pre-trained models like BERT to generate vector representations of both queries and document fragments, matching the most relevant fragments through similarity calculations (such as cosine similarity). RAG's retriever goes beyond simple keyword matching by using semantic-level vector representations, enabling more accurate results even for complex or ambiguous queries. This step is critical because retrieval quality directly determines the context available to the generator.
## 3.2 Generation Phase
The generation phase is the core of RAG. The generator produces coherent, natural text answers based on retrieved content. RAG generators like BART or GPT combine the user's query with retrieved document fragments to produce more precise and comprehensive answers. Unlike traditional generative models, RAG's generator can incorporate factual information from external knowledge bases, improving accuracy.
## 3.3 Multi-Turn Interaction and Feedback
RAG models effectively support multi-turn interactions in dialogue systems. Each round's query and generated results serve as input for the next round. Through this feedback loop, RAG progressively refines its retrieval and generation strategies, producing increasingly relevant answers across multiple conversation turns. This also enhances RAG's adaptability in complex dialogue scenarios involving cross-turn knowledge integration and reasoning.
# 4. Advantages and Limitations of RAG
## 4.1 Advantages
* Information completeness: RAG combines retrieval and generation, producing text that is both naturally fluent and grounded in real-time information from external knowledge bases. This significantly improves accuracy in knowledge-intensive scenarios like medical Q\&A or legal opinion generation, avoiding the risk of hallucinated information.
* Knowledge reasoning: RAG can efficiently retrieve from large-scale external knowledge bases and reason with real data to generate fact-based answers. Compared to traditional generative models, RAG handles more complex tasks, particularly cross-domain or cross-document reasoning — such as legal case analysis or financial report generation.
* Strong domain adaptability: RAG adapts well across domains, performing efficient retrieval and generation within specific fields. In domains like healthcare, law, and finance that require real-time updates and high accuracy, RAG outperforms models that rely solely on pre-training.
## 4.2 Limitations
Despite its strong potential and cross-domain adaptability, RAG faces several key limitations in practice that constrain large-scale deployment and optimization:
#### 4.2.1 Retriever Dependency and Quality Issues
RAG performance depends heavily on the quality of documents returned by the retriever. If retrieved fragments are irrelevant or inaccurate, the generated text may be biased or misleading — especially with ambiguous queries or cross-domain retrieval.
* Challenge: Improving retriever precision for complex queries across large, diverse knowledge bases remains difficult. Methods like BM25 have limitations, particularly with semantically ambiguous queries where keyword matching falls short.
* Solution: Adopt hybrid retrieval combining sparse retrieval (BM25) with dense retrieval (vector search). For example, Faiss enables BERT-based dense vector representations that significantly improve semantic matching, reducing the impact of irrelevant documents on generation.
#### 4.2.2 Generator Computational Complexity and Performance Bottlenecks
Combining retrieval and generation modules significantly increases computational complexity. When processing large datasets or long texts, the generator must integrate information from multiple document fragments, increasing generation time and reducing inference speed. This is a major bottleneck for real-time Q\&A systems.
* Challenge: As knowledge base scale grows, both retrieval computation and the generator's multi-fragment integration capabilities significantly impact system efficiency. GPU and memory consumption can multiply in multi-turn dialogues or complex generation tasks.
* Solution: Use model compression and knowledge distillation to reduce generator complexity and inference time. Distributed computing and model parallelization techniques like [DeepSpeed](https://www.deepspeed.ai/) can effectively handle high computational demands in large-scale scenarios.
#### 4.2.3 Knowledge Base Updates and Maintenance
RAG models typically rely on a pre-built external knowledge base containing documents, papers, legal provisions, and other information. The timeliness and accuracy of this content directly affects the credibility of generated results. Over time, knowledge base content may become outdated, producing answers that don't reflect current information — particularly problematic in fast-moving fields like healthcare and finance.
* Challenge: Knowledge bases need frequent updates, but manual updates are time-consuming and error-prone. Implementing continuous automated updates without impacting system performance is a significant challenge.
* Solution: Use automated crawlers and information extraction systems (such as Scrapy) to automatically fetch and update knowledge base content. Combined with [dynamic indexing techniques](https://arxiv.org/pdf/2102.03315), retrievers can update indexes in real time. Incremental learning allows the generator to gradually absorb new information, avoiding outdated answers.
#### 4.2.4 Generated Content Controllability and Transparency
RAG models have controllability and transparency challenges. In complex tasks or with ambiguous user input, the generator may produce incorrect reasoning based on inaccurate document fragments. Due to RAG's "black box" nature, users find it difficult to understand how the generator uses retrieved information — a significant concern in sensitive domains like law and healthcare, potentially eroding user trust.
* Challenge: Insufficient model transparency makes it hard for users to verify the source and credibility of generated answers. For tasks requiring high explainability (medical consultations, legal advice), inability to trace answer sources undermines trust.
* Solution: Introduce explainable AI (XAI) techniques like LIME or SHAP ([link](https://github.com/marcotcr/lime)) to provide detailed provenance for each generated answer, showing which knowledge fragments were referenced. Additionally, rule constraints and user feedback mechanisms can progressively optimize generator output for greater trustworthiness.
# 5. RAG Improvement Directions
RAG model performance depends on knowledge base accuracy and retrieval efficiency. Optimizing data collection, content chunking, retrieval precision, and answer generation are key to improving overall effectiveness.
## 5.1 Data Collection and Knowledge Base Construction
RAG's core dependency is knowledge base data quality and breadth — the knowledge base serves as "external memory." A high-quality knowledge base should include content from diverse, authoritative sources such as scientific literature databases (PubMed, IEEE Xplore), established news media, and industry standards and reports. It also needs automated update capabilities to stay current.
* Challenges:
* Single or limited data sources leading to insufficient coverage across domains
* Inconsistent data quality from non-authoritative or low-quality sources introducing bias
* Lack of regular update mechanisms, especially in fast-changing fields like law, finance, and technology
* Time-consuming and error-prone data processing workflows
* Data sensitivity and privacy concerns in domains like healthcare, law, and finance
* Improvements:
* Expand data source coverage across multiple domains, including specialized databases like PubMed, LexisNexis, and financial databases
* Build data quality review and filtering mechanisms using automated detection algorithms combined with manual review
* Implement automated knowledge base updates using web crawlers with change detection algorithms
* Adopt efficient data cleaning and classification using NLP techniques like BERT for entity recognition and text denoising
* Strengthen data security with de-identification, anonymization, and differential privacy protection
* Standardize data formats using JSON, XML, or knowledge graphs for structured storage
* Incorporate user feedback mechanisms to continuously optimize knowledge base content
## 5.2 Data Chunking and Content Management
Proper chunking strategies help models efficiently locate target information and provide clear context during answer generation. Chunking by paragraph, section, or topic improves retrieval efficiency and prevents redundant data from interfering with generation — especially important in complex, long-form text.
* Challenges:
* Unreasonable chunking breaking information chains and context
* Redundant data causing repetitive or overloaded generated content
* Inappropriate chunk granularity affecting retrieval precision
* Difficulty implementing topic-based or logic-based chunking for complex texts
* Improvements:
* Use NLP techniques (syntactic analysis, semantic segmentation) for automated, logic-based chunking
* Apply deduplication and information consolidation using similarity algorithms (TF-IDF, cosine similarity)
* Dynamically adjust chunk granularity based on task requirements
* Introduce topic-based chunking using topic models (LDA) or embedding-based text clustering
* Implement feedback mechanisms to continuously evaluate and optimize chunking strategies
## 5.3 Retrieval Optimization
The retrieval module determines the relevance and accuracy of generated answers. Hybrid retrieval strategies (combining BM25 and DPR) complement each other — BM25 handles keyword matching efficiently while DPR excels at deep semantic understanding.
* Challenges:
* Single retrieval strategies causing answer bias
* Tension between retrieval efficiency and resource consumption
* Redundant retrieval results leading to repetitive content
* Poor adaptability of fixed retrieval strategies across different task types
* Improvements:
* Combine BM25 and DPR in a hybrid retrieval strategy — BM25 for initial keyword filtering, then DPR for deep semantic matching
* Optimize retrieval efficiency using caching for frequent queries and distributed computing for parallel processing
* Apply deduplication and ranking optimization algorithms to retrieval results
* Dynamically adjust retrieval strategies based on task type — favoring semantic retrieval for medical Q\&A, keyword matching for news scenarios
* Integrate retrieval optimization frameworks like Haystack for enhanced extensibility
## 5.4 Answer Generation and Optimization
The generator produces natural language answers based on retrieved context. Accuracy and logical coherence directly impact user experience. Knowledge graphs and structured information help the generator better understand and connect context for more coherent, accurate answers.
* Challenges:
* Insufficient context leading to logically incoherent answers
* Inadequate accuracy in specialized domain answers
* Difficulty effectively integrating multi-turn user feedback
* Insufficient controllability and consistency in generated content
* Improvements:
* Integrate knowledge graphs and structured data to enhance context understanding
* Design domain-specific generation rules and terminology constraints
* Optimize user feedback mechanisms for dynamic generation logic adjustment
* Implement collaborative optimization between generator and retriever — allowing the generator to request additional context as needed
* Apply consistency detection and semantic correction to ensure uniform terminology and logical structure
## 5.5 RAG Pipeline

1. Data loading and query input:
1. The user submits a natural language query through the UI or API.
2. The input is passed to a vectorizer (such as BERT or Sentence Transformer) to convert the query into a vector representation.
2. Document retrieval:
1. The vectorized query is passed to the retriever, which finds the most relevant document fragments in the knowledge base.
2. Retrieval can use sparse techniques (BM25) or dense techniques (DPR) for improved matching efficiency and precision.
3. Generator processing and natural language generation:
1. Retrieved document fragments are fed to the generator (such as GPT, BART, or T5), which produces a natural language answer based on the query and document content.
2. The generator combines external retrieval results with pre-trained language knowledge for more precise, natural answers.
4. Result output:
1. The generated answer is returned to the user via API or UI, ensuring coherence and factual accuracy.
5. Feedback and optimization:
1. Users can provide feedback on generated answers, which the system uses to optimize retrieval and generation.
2. Through model fine-tuning or retrieval weight adjustments, the system progressively improves accuracy and efficiency.
# 6. RAG Case Studies
[RAG Across Various Domains](https://github.com/hymie122/RAG-Survey)
# 7. RAG Applications
RAG models have been widely adopted across multiple domains:
## 7.1 Intelligent Q\&A Systems
* RAG generates accurate, detailed answers by retrieving from external knowledge bases in real time, avoiding the hallucination issues of traditional generative models. For example, in medical Q\&A systems, RAG can incorporate the latest medical literature to generate answers with current treatment protocols, helping medical professionals quickly access the latest research and clinical recommendations.
* [Medical Q\&A System Case Study](https://www.apexon.com/blog/empowering-discovery-the-role-of-rag-architecture-generative-ai-in-healthcare-life-sciences/)
* 
* User submits a query through the web application:
1. The user enters a query in the web app, which enters the backend system and initiates the data processing pipeline.
* Authentication via Azure AD:
1. The system authenticates the user through Azure Active Directory (Azure AD), ensuring only authorized users can access the system and data.
* User permission check:
1. The system filters accessible content based on user group permissions managed by Azure AD.
* Azure AI Search Service:
1. The filtered query is passed to Azure AI Search, which finds relevant content in indexed databases or documents using semantic search.
* Document intelligence processing:
1. The system uses OCR and document extraction to convert unstructured data into structured, searchable data for Azure AI retrieval.
* Document sources:
1. Documents come from pre-stored collections that have been processed and indexed before user queries.
* Azure OpenAI generates response:
1. After retrieving relevant information, data is passed to Azure OpenAI, which uses natural language generation (NLG) to produce a coherent answer based on the query and retrieval results.
* Response returned to user:
1. The final answer is returned through the web application, completing the query-to-response flow.
* The entire pipeline demonstrates Azure AI technology integration, handling complex queries through document retrieval, intelligent processing, and natural language generation while ensuring data security and compliance.
## 7.2 Information Retrieval and Text Generation
* Text generation: RAG can not only retrieve relevant documents but also generate summaries, reports, or document abstracts, enhancing coherence and accuracy. For example, in the legal domain, RAG can integrate relevant statutes and case law to generate detailed legal opinions, ensuring comprehensiveness and rigor — particularly valuable for lawyers and legal practitioners to improve efficiency.
* [Legal Domain RAG Case Study](https://www.apexon.com/blog/empowering-discovery-the-role-of-rag-architecture-generative-ai-in-healthcare-life-sciences/)
* Summary:
* Background: Traditional LLMs perform well in generation tasks but have limitations with complex legal tasks. Legal documents have unique structures and terminology that standard retrieval benchmarks often fail to capture. LegalBench-RAG aims to provide a dedicated benchmark for evaluating legal document retrieval.
* LegalBench-RAG structure:
1. 
2. Workflow:
3. User inputs a question (Q: ?, A: ?): The user submits a query through the interface.
4. Embed + Retrieve module: Receives the query, embeds it as a vector, and performs similarity search in external knowledge bases or documents.
5. Answer generation (A): Based on the most relevant retrieved information, the generation model produces a coherent natural language answer.
6. Compare and return results: The generated answer is compared with previous related answers and returned to the user.
7. The benchmark is built on the LegalBench dataset with 6,858 query-answer pairs traced to exact locations in original legal documents.
8. LegalBench-RAG focuses on precisely retrieving small passages from legal texts rather than broad, contextually irrelevant fragments.
9. The dataset covers various legal document types including contracts and privacy policies, ensuring coverage across multiple legal scenarios.
* Significance: LegalBench-RAG is the first publicly available benchmark specifically for legal retrieval systems, providing a standardized framework for comparing retrieval algorithms in high-precision legal tasks such as citation lookup and clause interpretation.
* Key challenges:
1. RAG's generation component depends on retrieved information — incorrect retrieval can lead to incorrect generation.
2. The length and terminological complexity of legal documents increase retrieval and generation difficulty.
* Quality control: The dataset construction process ensures high-quality human annotations and textual precision, with multiple rounds of manual verification when mapping annotation categories and document IDs to specific text fragments.
## 7.3 Other Applications
RAG can also be applied to multimodal generation scenarios, including image, audio, and 3D content generation. Cross-modal applications like ReMoDiffuse and Make-An-Audio leverage RAG technology for generation across different data modalities. In enterprise decision support, RAG can rapidly retrieve external resources (industry reports, market data) to generate high-quality forward-looking reports, enhancing strategic decision-making capabilities.
## 8. Summary
This document systematically covers the core mechanisms, advantages, and applications of Retrieval-Augmented Generation (RAG). By combining generative and retrieval models, RAG addresses the hallucination problem of traditional generative models in factual tasks and the inability of retrieval models to produce coherent natural language output. RAG models retrieve information from external knowledge bases in real time, generating content that is both factually accurate and linguistically fluent — applicable to knowledge-intensive domains like healthcare, law, and intelligent Q\&A systems.
In practice, while RAG offers significant advantages in information completeness, reasoning capability, and cross-domain adaptability, it also faces challenges around data quality, computational resource consumption, and knowledge base maintenance. To further improve RAG performance, this document proposes comprehensive improvements across data collection, content chunking, retrieval strategy optimization, and answer generation — including knowledge graph integration, user feedback optimization, and efficient deduplication algorithms — to enhance model applicability and efficiency.
RAG has demonstrated strong potential in intelligent Q\&A, information retrieval, and text generation, and continues to expand into multimodal generation and enterprise decision support. Through hybrid retrieval techniques, knowledge graphs, and dynamic feedback mechanisms, RAG can flexibly address complex user needs, generating factually grounded and logically coherent answers. Going forward, RAG will further improve trustworthiness and practicality in specialized domains through enhanced model transparency and controllability, providing broader applications for intelligent information retrieval and content generation.
file: ./content/guide/dataset/rag.mdx
meta: {
"title": "知识库基础原理介绍",
"description": "本节详细介绍RAG模型的核心机制、应用场景及其在生成任务中的优势与局限性。"
}
[RAG文档](https://huggingface.co/docs/transformers/model_doc/rag)
# 1. 引言
随着自然语言处理(NLP)技术的迅猛发展,生成式语言模型(如GPT、BART等)在多种文本生成任务中表现卓越,尤其在语言生成和上下文理解方面。然而,纯生成模型在处理事实类任务时存在一些固有的局限性。例如,由于这些模型依赖于固定的预训练数据,它们在回答需要最新或实时信息的问题时,可能会出现“编造”信息的现象,导致生成结果不准确或缺乏事实依据。此外,生成模型在面对长尾问题和复杂推理任务时,常因缺乏特定领域的外部知识支持而表现不佳,难以提供足够的深度和准确性。
与此同时,检索模型(Retriever)能够通过在海量文档中快速找到相关信息,解决事实查询的问题。然而,传统检索模型(如BM25)在面对模糊查询或跨域问题时,往往只能返回孤立的结果,无法生成连贯的自然语言回答。由于缺乏上下文推理能力,检索模型生成的答案通常不够连贯和完整。
为了解决这两类模型的不足,检索增强生成模型(Retrieval-Augmented Generation,RAG)应运而生。RAG通过结合生成模型和检索模型的优势,实时从外部知识库中获取相关信息,并将其融入生成任务中,确保生成的文本既具备上下文连贯性,又包含准确的知识。这种混合架构在智能问答、信息检索与推理、以及领域特定的内容生成等场景中表现尤为出色。
## 1.1 RAG的定义
RAG是一种将信息检索与生成模型相结合的混合架构。首先,检索器从外部知识库或文档集中获取与用户查询相关的内容片段;然后,生成器基于这些检索到的内容生成自然语言输出,确保生成的内容既信息丰富,又具备高度的相关性和准确性。
# 2. RAG模型的核心机制
RAG 模型由两个主要模块构成:检索器(Retriever)与生成器(Generator)。这两个模块相互配合,确保生成的文本既包含外部的相关知识,又具备自然流畅的语言表达。
## 2.1 检索器(Retriever)
检索器的主要任务是从一个外部知识库或文档集中获取与输入查询最相关的内容。在RAG中,常用的技术包括:
* 向量检索:如BERT向量等,它通过将文档和查询转化为向量空间中的表示,并使用相似度计算来进行匹配。向量检索的优势在于能够更好地捕捉语义相似性,而不仅仅是依赖于词汇匹配。
* 传统检索算法:如BM25,主要基于词频和逆文档频率(TF-IDF)的加权搜索模型来对文档进行排序和检索。BM25适用于处理较为简单的匹配任务,尤其是当查询和文档中的关键词有直接匹配时。
RAG中检索器的作用是为生成器提供一个上下文背景,使生成器能够基于这些检索到的文档片段生成更为相关的答案。
## 2.2 生成器(Generator)
生成器负责生成最终的自然语言输出。在RAG系统中,常用的生成器包括:
* BART:BART是一种序列到序列的生成模型,专注于文本生成任务,可以通过不同层次的噪声处理来提升生成的质量 。
* GPT系列:GPT是一个典型的预训练语言模型,擅长生成流畅自然的文本。它通过大规模数据训练,能够生成相对准确的回答,尤其在任务-生成任务中表现尤为突出 。
生成器在接收来自检索器的文档片段后,会利用这些片段作为上下文,并结合输入的查询,生成相关且自然的文本回答。这确保了模型的生成结果不仅仅基于已有的知识,还能够结合外部最新的信息。
## 2.3 RAG的工作流程
RAG模型的工作流程可以总结为以下几个步骤:
1. 输入查询:用户输入问题,系统将其转化为向量表示。
2. 文档检索:检索器从知识库中提取与查询最相关的文档片段,通常使用向量检索技术或BM25等传统技术进行。
3. 生成答案:生成器接收检索器提供的片段,并基于这些片段生成自然语言答案。生成器不仅基于原始的用户查询,还会利用检索到的片段提供更加丰富、上下文相关的答案。
4. 输出结果:生成的答案反馈给用户,这个过程确保了用户能够获得基于最新和相关信息的准确回答。
# 3. RAG模型的工作原理
## 3.1 检索阶段
在RAG模型中,用户的查询首先被转化为向量表示,然后在知识库中执行向量检索。通常,检索器采用诸如BERT等预训练模型生成查询和文档片段的向量表示,并通过相似度计算(如余弦相似度)匹配最相关的文档片段。RAG的检索器不仅仅依赖简单的关键词匹配,而是采用语义级别的向量表示,从而在面对复杂问题或模糊查询时,能够更加准确地找到相关知识。这一步骤对于最终生成的回答至关重要,因为检索的效率和质量直接决定了生成器可利用的上下文信息 。
## 3.2 生成阶段
生成阶段是RAG模型的核心部分,生成器负责基于检索到的内容生成连贯且自然的文本回答。RAG中的生成器,如BART或GPT等模型,结合用户输入的查询和检索到的文档片段,生成更加精准且丰富的答案。与传统生成模型相比,RAG的生成器不仅能够生成语言流畅的回答,还可以根据外部知识库中的实际信息提供更具事实依据的内容,从而提高了生成的准确性 。
## 3.3 多轮交互与反馈机制
RAG模型在对话系统中能够有效支持多轮交互。每一轮的查询和生成结果会作为下一轮的输入,系统通过分析和学习用户的反馈,逐步优化后续查询的上下文。通过这种循环反馈机制,RAG能够更好地调整其检索和生成策略,使得在多轮对话中生成的答案越来越符合用户的期望。此外,多轮交互还增强了RAG在复杂对话场景中的适应性,使其能够处理跨多轮的知识整合和复杂推理 。
# 4. RAG的优势与局限
## 4.1 优势
* 信息完整性:RAG 模型结合了检索与生成技术,使得生成的文本不仅语言自然流畅,还能够准确利用外部知识库提供的实时信息。这种方法能够显著提升生成任务的准确性,特别是在知识密集型场景下,如医疗问答或法律意见生成。通过从知识库中检索相关文档,RAG 模型避免了生成模型“编造”信息的风险,确保输出更具真实性 。
* 知识推理能力:RAG 能够利用大规模的外部知识库进行高效检索,并结合这些真实数据进行推理,生成基于事实的答案。相比传统生成模型,RAG 能处理更为复杂的任务,特别是涉及跨领域或跨文档的推理任务。例如,法律领域的复杂判例推理或金融领域的分析报告生成都可以通过RAG的推理能力得到优化 。
* 领域适应性强:RAG 具有良好的跨领域适应性,能够根据不同领域的知识库进行特定领域内的高效检索和生成。例如,在医疗、法律、金融等需要实时更新和高度准确性的领域,RAG 模型的表现优于仅依赖预训练的生成模型 。
## 4.2 局限
RAG(检索增强生成)模型通过结合检索器和生成器,实现了在多种任务中知识密集型内容生成的突破性进展。然而,尽管其具有较强的应用潜力和跨领域适应能力,但在实际应用中仍然面临着一些关键局限,限制了其在大规模系统中的部署和优化。以下是RAG模型的几个主要局限性:
#### 4.2.1 检索器的依赖性与质量问题
RAG模型的性能很大程度上取决于检索器返回的文档质量。由于生成器主要依赖检索器提供的上下文信息,如果检索到的文档片段不相关、不准确,生成的文本可能出现偏差,甚至产生误导性的结果。尤其在多模糊查询或跨领域检索的情况下,检索器可能无法找到合适的片段,这将直接影响生成内容的连贯性和准确性。
* 挑战:当知识库庞大且内容多样时,如何提高检索器在复杂问题下的精确度是一大挑战。当前的方法如BM25等在特定任务上有局限,尤其是在面对语义模糊的查询时,传统的关键词匹配方式可能无法提供语义上相关的内容。
* 解决途径:引入混合检索技术,如结合稀疏检索(BM25)与密集检索(如向量检索)。例如,Faiss的底层实现允许通过BERT等模型生成密集向量表示,显著提升语义级别的匹配效果。通过这种方式,检索器可以捕捉深层次的语义相似性,减少无关文档对生成器的负面影响。
#### 4.2.2 生成器的计算复杂度与性能瓶颈
RAG模型将检索和生成模块结合,尽管生成结果更加准确,但也大大增加了模型的计算复杂度。尤其在处理大规模数据集或长文本时,生成器需要处理来自多个文档片段的信息,导致生成时间明显增加,推理速度下降。对于实时问答系统或其他需要快速响应的应用场景,这种高计算复杂度是一个主要瓶颈。
* 挑战:当知识库规模扩大时,检索过程中的计算开销以及生成器在多片段上的整合能力都会显著影响系统的效率。同时,生成器也面临着资源消耗的问题,尤其是在多轮对话或复杂生成任务中,GPU和内存的消耗会成倍增加。
* 解决途径:使用模型压缩技术和知识蒸馏来减少生成器的复杂度和推理时间。此外,分布式计算与模型并行化技术的引入,如[DeepSpeed](https://www.deepspeed.ai/)和模型压缩工具,可以有效应对生成任务的高计算复杂度,提升大规模应用场景中的推理效率。
#### 4.2.3 知识库的更新与维护
RAG模型通常依赖于一个预先建立的外部知识库,该知识库可能包含文档、论文、法律条款等各类信息。然而,知识库内容的时效性和准确性直接影响到RAG生成结果的可信度。随着时间推移,知识库中的内容可能过时,导致生成的回答不能反映最新的信息。这对于需要实时信息的场景(如医疗、金融)尤其明显。
* 挑战:知识库需要频繁更新,但手动更新知识库既耗时又容易出错。如何在不影响系统性能的情况下实现知识库的持续自动更新是当前的一大挑战。
* 解决途径:利用自动化爬虫和信息提取系统,可以实现对知识库的自动化更新,例如,Scrapy等爬虫框架可以自动抓取网页数据并更新知识库。结合[动态索引技术](https://arxiv.org/pdf/2102.03315),可以帮助检索器实时更新索引,确保知识库反映最新信息。同时,结合增量学习技术,生成器可以逐步吸收新增的信息,避免生成过时答案。此外,动态索引技术也可以帮助检索器实时更新索引,确保知识库检索到的文档反映最新的内容。
#### 4.2.4 生成内容的可控性与透明度
RAG模型结合了检索与生成模块,在生成内容的可控性和透明度上存在一定问题。特别是在复杂任务或多义性较强的用户输入情况下,生成器可能会基于不准确的文档片段生成错误的推理,导致生成的答案偏离实际问题。此外,由于RAG模型的“黑箱”特性,用户难以理解生成器如何利用检索到的文档信息,这在高敏感领域如法律或医疗中尤为突出,可能导致用户对生成内容产生不信任感。
* 挑战:模型透明度不足使得用户难以验证生成答案的来源和可信度。对于需要高可解释性的任务(如医疗问诊、法律咨询等),无法追溯生成答案的知识来源会导致用户不信任模型的决策。
* 解决途径:为提高透明度,可以引入可解释性AI(XAI)技术,如LIME或SHAP([链接](https://github.com/marcotcr/lime)),为每个生成答案提供详细的溯源信息,展示所引用的知识片段。这种方法能够帮助用户理解模型的推理过程,从而增强对模型输出的信任。此外,针对生成内容的控制,可以通过加入规则约束或用户反馈机制,逐步优化生成器的输出,确保生成内容更加可信。
# 5. RAG整体改进方向
RAG模型的整体性能依赖于知识库的准确性和检索的效率,因此在数据采集、内容分块、精准检索和回答生成等环节进行优化,是提升模型效果的关键。通过加强数据来源、改进内容管理、优化检索策略及提升回答生成的准确性,RAG模型能够更加适应复杂且动态的实际应用需求。
## 5.1 数据采集与知识库构建
RAG模型的核心依赖在于知识库的数据质量和广度,知识库在某种程度上充当着“外部记忆”的角色。因此,高质量的知识库不仅应包含广泛领域的内容,更要确保数据来源的权威性、可靠性以及时效性。知识库的数据源应涵盖多种可信的渠道,例如科学文献数据库(如PubMed、IEEE Xplore)、权威新闻媒体、行业标准和报告等,这样才能提供足够的背景信息支持RAG在不同任务中的应用。此外,为了确保RAG模型能够提供最新的回答,知识库需要具备自动化更新的能力,以避免数据内容老旧,导致回答失准或缺乏现实参考。
* 挑战:
* 尽管数据采集是构建知识库的基础,但在实际操作中仍存在以下几方面的不足:
* 数据采集来源单一或覆盖不全
1. RAG模型依赖多领域数据的支持,然而某些知识库过度依赖单一或有限的数据源,通常集中在某些领域,导致在多任务需求下覆盖不足。例如,依赖医学领域数据而缺乏法律和金融数据会使RAG模型在跨领域问答中表现不佳。这种局限性削弱了RAG模型在处理不同主题或多样化查询时的准确性,使得系统在应对复杂或跨领域任务时能力欠缺。
* 数据质量参差不齐
1. 数据源的质量差异直接影响知识库的可靠性。一些数据可能来源于非权威或低质量渠道,存在偏见、片面或不准确的内容。这些数据若未经筛选录入知识库,会导致RAG模型生成偏差或不准确的回答。例如,在医学领域中,如果引入未经验证的健康信息,可能导致模型给出误导性回答,产生负面影响。数据质量不一致的知识库会大大降低模型输出的可信度和适用性。
* 缺乏定期更新机制
1. 许多知识库缺乏自动化和频繁的更新机制,特别是在信息变动频繁的领域,如法律、金融和科技。若知识库长期未更新,则RAG模型无法提供最新信息,生成的回答可能过时或不具备实时参考价值。对于用户而言,特别是在需要实时信息的场景下,滞后的知识库会显著影响RAG模型的可信度和用户体验。
* 数据处理耗时且易出错
1. 数据的采集、清洗、分类和结构化处理是一项繁琐而复杂的任务,尤其是当数据量巨大且涉及多种格式时。通常,大量数据需要人工参与清洗和结构化,而自动化处理流程也存在缺陷,可能会产生错误或遗漏关键信息。低效和易出错的数据处理流程会导致知识库内容不准确、不完整,进而影响RAG模型生成的答案的准确性和连贯性。
* 数据敏感性和隐私问题
1. 一些特定领域的数据(如医疗、法律、金融)包含敏感信息,未经适当的隐私保护直接引入知识库可能带来隐私泄露的风险。此外,某些敏感数据需要严格的授权和安全存储,以确保在知识库使用中避免违规或隐私泄漏。若未能妥善处理数据隐私问题,不仅会影响系统的合规性,还可能对用户造成严重后果。
* 改进:
* 针对以上不足,可以从以下几个方面进行改进,以提高数据采集和知识库构建的有效性:
* 扩大数据源覆盖范围,增加数据的多样性
1. 具体实施:将知识库的数据源扩展至多个重要领域,确保包含医疗、法律、金融等关键领域的专业数据库,如PubMed、LexisNexis和金融数据库。使用具有开放许可的开源数据库和经过认证的数据,确保来源多样化且权威性强。
2. 目的与效果:通过跨领域数据覆盖,知识库的广度和深度得以增强,确保RAG模型能够在多任务场景下提供可靠回答。借助多领域合作机构的数据支持,在应对多样化需求时将更具优势。
* 构建数据质量审查与过滤机制
1. 具体实施:采用自动化数据质量检测算法,如文本相似度检查、情感偏差检测等工具,结合人工审查过滤不符合标准的数据。为数据打分并构建“数据可信度评分”,基于来源可信度、内容完整性等指标筛选数据。
2. 目的与效果:减少低质量、偏见数据的干扰,确保知识库内容的可靠性。此方法保障了RAG模型输出的权威性,特别在回答复杂或专业问题时,用户能够获得更加精准且中立的答案。
* 实现知识库的自动化更新
1. 具体实施:引入自动化数据更新系统,如网络爬虫,定期爬取可信站点、行业数据库的最新数据,并利用变化检测算法筛选出与已有知识库重复或已失效的数据。更新机制可以结合智能筛选算法,仅采纳与用户查询高相关性或时效性强的数据。
2. 目的与效果:知识库保持及时更新,确保模型在快速变化的领域(如金融、政策、科技)中提供最新信息。用户体验将因此大幅提升,特别是在需要动态或最新信息的领域,输出的内容将更具时效性。
* 采用高效的数据清洗与分类流程
1. 具体实施:使用自然语言处理技术,如BERT等模型进行数据分类、实体识别和文本去噪,结合去重算法清理重复内容。采用自动化的数据标注和分类算法,将不同数据类型分领域存储。
2. 目的与效果:数据清洗和分领域管理可以大幅提高数据处理的准确性,减少低质量数据的干扰。此改进确保RAG模型的回答生成更流畅、上下文更连贯,提升用户对生成内容的理解和信赖。
* 强化数据安全与隐私保护措施
1. 具体实施:针对医疗、法律等敏感数据,采用去标识化处理技术(如数据脱敏、匿名化等),并结合差分隐私保护。建立数据权限管理和加密存储机制,对敏感信息进行严格管控。
2. 目的与效果:在保护用户隐私的前提下,确保使用的数据合规、安全,适用于涉及个人或敏感数据的应用场景。此措施进一步保证了系统的法律合规性,并有效防止隐私泄露风险。
* 优化数据格式与结构的标准化
1. 具体实施:建立统一的数据格式与标准编码格式,例如使用JSON、XML或知识图谱形式组织结构化数据,以便于检索系统在查询时高效利用。同时,使用知识图谱等结构化工具,将复杂数据间的关系进行系统化存储。
2. 目的与效果:提高数据检索效率,确保模型在生成回答时能够高效使用数据的关键信息。标准化的数据结构支持高效的跨领域检索,并提高了RAG模型的内容准确性和知识关系的透明度。
* 用户反馈机制
1. 具体实施:通过用户反馈系统,记录用户对回答的满意度、反馈意见及改进建议。使用机器学习算法从反馈中识别知识库中的盲区与信息误差,反馈至数据管理流程中进行更新和优化。
2. 目的与效果:利用用户反馈作为数据质量的调整依据,帮助知识库持续优化内容。此方法不仅提升了RAG模型的实际效用,还使知识库更贴合用户需求,确保输出内容始终符合用户期望。
## 5.2 数据分块与内容管理
RAG模型的数据分块与内容管理是优化检索与生成流程的关键。合理的分块策略能够帮助模型高效定位目标信息,并在回答生成时提供清晰的上下文支持。通常情况下,将数据按段落、章节或主题进行分块,不仅有助于检索效率的提升,还能避免冗余数据对生成内容造成干扰。尤其在复杂、长文本中,适当的分块策略可保证模型生成的答案具备连贯性、精确性,避免出现内容跳跃或上下文断裂的问题。
* 挑战:
* 在实际操作中,数据分块与内容管理环节存在以下问题:
* 分块不合理导致的信息断裂
1. 部分文本过度切割或分块策略不合理,可能导致信息链条被打断,使得模型在回答生成时缺乏必要的上下文支持。这会使生成内容显得零散,不具备连贯性,影响用户对答案的理解。例如,将法律文本或技术文档随意切割成小段落会导致重要的上下文关系丢失,降低模型的回答质量。
* 冗余数据导致生成内容重复或信息过载
1. 数据集中往往包含重复信息,若不去重或优化整合,冗余数据可能导致生成内容的重复或信息过载。这不仅影响用户体验,还会浪费计算资源。例如,在新闻数据或社交媒体内容中,热点事件的描述可能重复出现,模型在生成回答时可能反复引用相同信息。
* 分块粒度选择不当影响检索精度
1. 如果分块粒度过细,模型可能因缺乏足够的上下文而生成不准确的回答;若分块过大,检索时将难以定位具体信息,导致回答内容冗长且含有无关信息。选择适当的分块粒度对生成答案的准确性和相关性至关重要,特别是在问答任务中需要精确定位答案的情况下,粗放的分块策略会明显影响用户的阅读体验和回答的可读性。
* 难以实现基于主题或内容逻辑的分块
1. 某些复杂文本难以直接按主题或逻辑结构进行分块,尤其是内容密集或领域专业性较强的数据。基于关键字或简单的规则切割往往难以识别不同主题和信息层次,导致模型在回答生成时信息杂乱。对内容逻辑或主题的错误判断,尤其是在医学、金融等场景下,会大大影响生成答案的准确度和专业性。
* 改进:
* 为提高数据分块和内容管理的有效性,可以从以下几方面进行优化:
* 引入NLP技术进行自动化分块和上下文分析
1. 具体实施:借助自然语言处理(NLP)技术,通过句法分析、语义分割等方式对文本进行逻辑切割,以确保分块的合理性。可以基于BERT等预训练模型实现主题识别和上下文分析,确保每个片段均具备完整的信息链,避免信息断裂。
2. 目的与效果:确保文本切割基于逻辑或语义关系,避免信息链条被打断,生成答案时能够更具连贯性,尤其适用于长文本和复杂结构的内容,使模型在回答时上下文更加完整、连贯。
* 去重与信息整合,优化内容简洁性
1. 具体实施:利用相似度算法(如TF-IDF、余弦相似度)识别冗余内容,并结合聚类算法自动合并重复信息。针对内容频繁重复的情况,可设置内容标记或索引,避免生成时多次引用相同片段。
2. 目的与效果:通过去重和信息整合,使数据更具简洁性,避免生成答案中出现重复信息。减少冗余信息的干扰,使用户获得简明扼要的回答,增强阅读体验,同时提升生成过程的计算效率。
* 根据任务需求动态调整分块粒度
1. 具体实施:根据模型任务的不同,设置动态分块策略。例如,在问答任务中对关键信息较短的内容可采用小粒度分块,而在长文本或背景性内容中采用较大粒度。分块策略可基于查询需求或内容复杂度自动调整。
2. 目的与效果:分块粒度的动态调整确保模型在检索和生成时既能准确定位关键内容,又能为回答提供足够的上下文支持,提升生成内容的精准性和相关性,确保用户获取的信息既准确又不冗长。
* 引入基于主题的分块方法以提升上下文完整性
1. 具体实施:使用主题模型(如LDA)或嵌入式文本聚类技术,对文本内容按主题进行自动分类与分块。基于相同主题内容的聚合分块,有助于模型识别不同内容层次,尤其适用于复杂的学术文章或多章节的长篇报告。
2. 目的与效果:基于主题的分块确保同一主题的内容保持在一个片段内,提升模型在回答生成时的上下文连贯性。适用于主题复杂、层次清晰的内容场景,提高回答的专业性和条理性,使用户更容易理解生成内容的逻辑关系。
* 实时评估分块策略与内容呈现效果的反馈机制
1. 具体实施:通过用户反馈机制和生成质量评估系统实时监测生成内容的连贯性和准确性。对用户反馈中涉及分块效果差的部分进行重新分块,通过用户使用数据优化分块策略。
2. 目的与效果:用户反馈帮助识别不合理的分块和内容呈现问题,实现分块策略的动态优化,持续提升生成内容的质量和用户满意度。
## 5.3 检索优化
在RAG模型中,检索模块决定了生成答案的相关性和准确性。有效的检索策略可确保模型获取到最适合的上下文片段,使生成的回答更加精准且贴合查询需求。常用的混合检索策略(如BM25和DPR结合)能够在关键词匹配和语义检索方面实现优势互补:BM25适合高效地处理关键字匹配任务,而DPR在理解深层语义上表现更为优异。因此,合理选用检索策略有助于在不同任务场景下达到计算资源和检索精度的平衡,以高效提供相关上下文供生成器使用。
* 挑战:
* 检索优化过程中,仍面临以下不足之处:
* 检索策略单一导致的回答偏差
1. 当仅依赖BM25或DPR等单一技术时,模型可能难以平衡关键词匹配与语义理解。BM25在处理具象关键字时表现良好,但在面对复杂、含义丰富的语义查询时效果欠佳;相反,DPR虽然具备深度语义匹配能力,但对高频关键词匹配的敏感度较弱。检索策略单一将导致模型难以适应复杂的用户查询,回答中出现片面性或不够精准的情况。
* 检索效率与资源消耗的矛盾
1. 检索模块需要在短时间内处理大量查询,而语义检索(如DPR)需要进行大量的计算和存储操作,计算资源消耗高,影响系统响应速度。特别是对于需要实时响应的应用场景,DPR的计算复杂度往往难以满足实际需求,因此在实时性和资源利用率上亟需优化。
* 检索结果的冗余性导致内容重复
1. 当检索策略未对结果进行去重或排序优化时,RAG模型可能从知识库中检索出相似度高但内容冗余的文档片段。这会导致生成的回答中包含重复信息,影响阅读体验,同时增加无效信息的比例,使用户难以迅速获取核心答案。
* 不同任务需求下检索策略的适配性差
1. RAG模型应用场景丰富,但不同任务对检索精度、速度和上下文长度的需求不尽相同。固定检索策略难以灵活应对多样化的任务需求,导致在应对不同任务时,模型检索效果受限。例如,面向精确性较高的医疗问答场景时,检索策略应偏向语义准确性,而在热点新闻场景中则应偏重检索速度。
* 改进:
* 针对上述不足,可以从以下几个方面优化检索模块:
* 结合BM25与DPR的混合检索策略
1. 具体实施:采用BM25进行关键词初筛,快速排除无关信息,然后使用DPR进行深度语义匹配筛选。这样可以有效提升检索精度,平衡关键词匹配和语义理解。
2. 目的与效果:通过多层筛选过程,确保检索结果在语义理解和关键词匹配方面互补,提升生成内容的准确性,特别适用于多意图查询或复杂的长文本检索。
* 优化检索效率,控制计算资源消耗
1. 具体实施:利用缓存机制存储近期高频查询结果,避免对相似查询的重复计算。同时,可基于分布式计算结构,将DPR的语义计算任务分散至多节点并行处理。
2. 目的与效果:缓存与分布式计算结合可显著减少检索计算压力,使系统能够在有限资源下提高响应速度,适用于高并发、实时性要求高的应用场景。
* 引入去重和排序优化算法
1. 具体实施:在检索结果中应用余弦相似度去重算法,筛除冗余内容,并基于用户偏好或时间戳对检索结果排序,以确保输出内容的丰富性和新鲜度。
2. 目的与效果:通过去重和优化排序,确保生成内容更加简洁、直接,减少重复信息的干扰,提高用户获取信息的效率和体验。
* 动态调整检索策略适应多任务需求
1. 具体实施:设置不同检索策略模板,根据任务类型自动调整检索权重、片段长度和策略组合。例如在医疗场景中偏向语义检索,而在金融新闻场景中更重视快速关键词匹配。
2. 目的与效果:动态调整检索策略使RAG模型更加灵活,能够适应不同任务需求,确保检索的精准性和生成答案的上下文适配性,显著提升多场景下的用户体验。
* 借助Haystack等检索优化框架
1. 具体实施:在RAG模型中集成Haystack框架,以实现更高效的检索效果,并利用框架中的插件生态系统来增强检索模块的可扩展性和可调节性。
2. 目的与效果:Haystack提供了检索和生成的整合接口,有助于快速优化检索模块,并适应复杂多样的用户需求,在多任务环境中提供更稳定的性能表现。
## 5.4 回答生成与优化
在RAG模型中,生成器负责基于检索模块提供的上下文,为用户查询生成自然语言答案。生成内容的准确性和逻辑性直接决定了用户的体验,因此优化生成器的表现至关重要。通过引入知识图谱等结构化信息,生成器能够更准确地理解和关联上下文,从而生成逻辑连贯、准确的回答。此外,生成器的生成逻辑可结合用户反馈持续优化,使回答风格和内容更加符合用户需求。
* 挑战:
* 在回答生成过程中,RAG模型仍面临以下不足:
* 上下文不充分导致的逻辑不连贯
1. 当生成器在上下文缺失或信息不完整的情况下生成回答时,生成内容往往不够连贯,特别是在处理复杂、跨领域任务时。这种缺乏上下文支持的问题,容易导致生成器误解或忽略关键信息,最终生成内容的逻辑性和完整性欠佳。如在医学场景中,若生成器缺少对病例或症状的全面理解,可能导致回答不准确或不符合逻辑,影响专业性和用户信任度。
* 专业领域回答的准确性欠佳
1. 在医学、法律等高专业领域中,生成器的回答需要高度的准确性。然而,生成器可能因缺乏特定知识而生成不符合领域要求的回答,出现内容偏差或理解错误,尤其在涉及专业术语和复杂概念时更为明显。如在法律咨询中,生成器可能未能正确引用相关法条或判例,导致生成的答案不够精确,甚至可能产生误导。
* 难以有效整合多轮用户反馈
1. 生成器缺乏有效机制来利用多轮用户反馈进行自我优化。用户反馈可能涉及回答内容的准确性、逻辑性以及风格适配等方面,但生成器在连续对话中缺乏充分的调节机制,难以持续调整生成策略和回答风格。如在客服场景中,生成器可能连续生成不符合用户需求的回答,降低了用户满意度。
* 生成内容的可控性和一致性不足
1. 在特定领域回答生成中,生成器的输出往往不具备足够的可控性和一致性。由于缺乏领域特定的生成规则和约束,生成内容的专业性和风格一致性欠佳,难以满足高要求的应用场景。如在金融报告生成中,生成内容需要确保一致的风格和术语使用,否则会影响输出的专业性和可信度。
* 改进:
* 针对以上不足,可以从以下方面优化回答生成模块:
* 引入知识图谱与结构化数据,增强上下文理解
1. 具体实施:结合知识图谱或知识库,将医学、法律等专业领域的信息整合到生成过程中。生成器在生成回答时,可以从知识图谱中提取关键信息和关联知识点,确保回答具备连贯的逻辑链条。
2. 目的与效果:知识图谱的引入提升了生成内容的连贯性和准确性,尤其在高专业性领域中,通过丰富的上下文理解,使生成器能够产生符合逻辑的回答。
* 设计专业领域特定的生成规则和约束
1. 具体实施:在生成模型中加入领域特定的生成规则和用语约束,特别针对医学、法律等领域的常见问答场景,设定回答模板、术语库等,以提高生成内容的准确性和一致性。
2. 目的与效果:生成内容更具领域特征,输出风格和内容的专业性增强,有效降低了生成器在专业领域中的回答偏差,满足用户对专业性和可信度的要求。
* 优化用户反馈机制,实现动态生成逻辑调整
1. 具体实施:利用机器学习算法对用户反馈进行分析,从反馈中提取生成错误或用户需求的调整信息,动态调节生成器的生成逻辑和策略。同时,在多轮对话中逐步适应用户的需求和风格偏好。
2. 目的与效果:用户反馈的高效利用能够帮助生成器优化生成内容,提高连续对话中的响应质量,提升用户体验,并使回答更贴合用户需求。
* 引入生成器与检索器的协同优化机制
1. 具体实施:通过协同优化机制,在生成器生成答案之前,允许生成器请求检索器补充缺失的上下文信息。生成器可基于回答需求自动向检索器发起上下文补充请求,从而获取完整的上下文。
2. 目的与效果:协同优化机制保障了生成器在回答时拥有足够的上下文支持,避免信息断层或缺失,提升回答的完整性和准确性。
* 实施生成内容的一致性检测和语义校正
1. 具体实施:通过一致性检测算法对生成内容进行术语、风格的统一管理,并结合语义校正模型检测生成内容是否符合用户需求的逻辑结构。在复杂回答生成中,使用语义校正对不符合逻辑的生成内容进行自动优化。
2. 目的与效果:生成内容具备高度一致性和逻辑性,特别是在多轮对话和专业领域生成中,保障了内容的稳定性和专业水准,提高了生成答案的可信度和用户满意度。
## 5.5 RAG流程

1. 数据加载与查询输入:
1. 用户通过界面或API提交自然语言查询,系统接收查询作为输入。
2. 输入被传递至向量化器,利用向量化技术(如BERT或Sentence Transformer)将自然语言查询转换为向量表示。
2. 文档检索:
1. 向量化后的查询会传递给检索器,检索器通过在知识库中查找最相关的文档片段。
2. 检索可以基于稀疏检索技术(如BM25)或密集检索技术(如DPR)来提高匹配效率和精度。
3. 生成器处理与自然语言生成:
1. 检索到的文档片段作为生成器的输入,生成器(如GPT、BART或T5)基于查询和文档内容生成自然语言回答。
2. 生成器结合了外部检索结果和预训练模型的语言知识,使回答更加精准、自然。
4. 结果输出:
1. 系统生成的答案通过API或界面返回给用户,确保答案连贯且知识准确。
5. 反馈与优化:
1. 用户可以对生成的答案进行反馈,系统根据反馈优化检索与生成过程。
2. 通过微调模型参数或调整检索权重,系统逐步改进其性能,确保未来查询时更高的准确性与效率。
# 6. RAG相关案例整合
[各种分类领域下的RAG](https://github.com/hymie122/RAG-Survey)
# 7. RAG模型的应用
RAG模型已在多个领域得到广泛应用,主要包括:
## 7.1 智能问答系统中的应用
* RAG通过实时检索外部知识库,生成包含准确且详细的答案,避免传统生成模型可能产生的错误信息。例如,在医疗问答系统中,RAG能够结合最新的医学文献,生成包含最新治疗方案的准确答案,避免生成模型提供过时或错误的建议。这种方法帮助医疗专家快速获得最新的研究成果和诊疗建议,提升医疗决策的质量。
* [医疗问答系统案例](https://www.apexon.com/blog/empowering-discovery-the-role-of-rag-architecture-generative-ai-in-healthcare-life-sciences/)
* 
* 用户通过Web应用程序发起查询:
1. 用户在一个Web应用上输入查询请求,这个请求进入后端系统,启动了整个数据处理流程。
* 使用Azure AD进行身份验证:
1. 系统通过Azure Active Directory (Azure AD) 对用户进行身份验证,确保只有经过授权的用户才能访问系统和数据。
* 用户权限检查:
1. 系统根据用户的组权限(由Azure AD管理)过滤用户能够访问的内容。这个步骤保证了用户只能看到他们有权限查看的信息。
* Azure AI搜索服务:
1. 过滤后的用户查询被传递给Azure AI搜索服务,该服务会在已索引的数据库或文档中查找与查询相关的内容。这个搜索引擎通过语义搜索技术检索最相关的信息。
* 文档智能处理:
1. 系统使用OCR(光学字符识别)和文档提取等技术处理输入的文档,将非结构化数据转换为结构化、可搜索的数据,便于Azure AI进行检索。
* 文档来源:
1. 这些文档来自预先存储的输入文档集合,这些文档在被用户查询之前已经通过文档智能处理进行了准备和索引。
* Azure Open AI生成响应:
1. 在检索到相关信息后,数据会被传递到Azure Open AI,该模块利用自然语言生成(NLG)技术,根据用户的查询和检索结果生成连贯的回答。
* 响应返回用户:
1. 最终生成的回答通过Web应用程序返回给用户,完成整个查询到响应的流程。
* 整个流程展示了Azure AI技术的集成,通过文档检索、智能处理以及自然语言生成来处理复杂的查询,并确保了数据的安全和合规性。
## 7.2 信息检索与文本生成
* 文本生成:RAG不仅可以检索相关文档,还能根据这些文档生成总结、报告或文档摘要,从而增强生成内容的连贯性和准确性。例如,法律领域中,RAG可以整合相关法条和判例,生成详细的法律意见书,确保内容的全面性和严谨性。这在法律咨询和文件生成过程中尤为重要,可以帮助律师和法律从业者提高工作效率。
* [法律领域检索增强生成案例](https://www.apexon.com/blog/empowering-discovery-the-role-of-rag-architecture-generative-ai-in-healthcare-life-sciences/)
* 内容总结:
* 背景: 传统的大语言模型 (LLMs) 在生成任务中表现优异,但在处理法律领域中的复杂任务时存在局限。法律文档具有独特的结构和术语,标准的检索评估基准往往无法充分捕捉这些领域特有的复杂性。为了弥补这一不足,LegalBench-RAG 旨在提供一个评估法律文档检索效果的专用基准。
* LegalBench-RAG 的结构:
1. 
2. 工作流程:
3. 用户输入问题(Q: ?,A: ?):用户通过界面输入查询问题,提出需要答案的具体问题。
4. 嵌入与检索模块(Embed + Retrieve):该模块接收到用户的查询后,会对问题进行嵌入(将其转化为向量),并在外部知识库或文档中执行相似度检索。通过检索算法,系统找到与查询相关的文档片段或信息。
5. 生成答案(A):基于检索到的最相关信息,生成模型(如GPT或类似的语言模型)根据检索的结果生成连贯的自然语言答案。
6. 对比和返回结果:生成的答案会与之前的相关问题答案进行对比,并最终将生成的答案返回给用户。
7. 该基准基于 LegalBench 的数据集,构建了 6858 个查询-答案对,并追溯到其原始法律文档的确切位置。
8. LegalBench-RAG 侧重于精确地检索法律文本中的小段落,而非宽泛的、上下文不相关的片段。
9. 数据集涵盖了合同、隐私政策等不同类型的法律文档,确保涵盖多个法律应用场景。
* 意义: LegalBench-RAG 是第一个专门针对法律检索系统的公开可用的基准。它为研究人员和公司提供了一个标准化的框架,用于比较不同的检索算法的效果,特别是在需要高精度的法律任务中,例如判决引用、条款解释等。
* 关键挑战:
1. RAG 系统的生成部分依赖检索到的信息,错误的检索结果可能导致错误的生成输出。
2. 法律文档的长度和术语复杂性增加了模型检索和生成的难度。
* 质量控制: 数据集的构建过程确保了高质量的人工注释和文本精确性,特别是在映射注释类别和文档ID到具体文本片段时进行了多次人工校验。
## 7.3 其它应用场景
RAG还可以应用于多模态生成场景,如图像、音频和3D内容生成。例如,跨模态应用如ReMoDiffuse和Make-An-Audio利用RAG技术实现不同数据形式的生成。此外,在企业决策支持中,RAG能够快速检索外部资源(如行业报告、市场数据),生成高质量的前瞻性报告,从而提升企业战略决策的能力。
## 8 总结
本文档系统阐述了检索增强生成(RAG)模型的核心机制、优势与应用场景。通过结合生成模型与检索模型,RAG解决了传统生成模型在面对事实性任务时的“编造”问题和检索模型难以生成连贯自然语言输出的不足。RAG模型能够实时从外部知识库获取信息,使生成内容既包含准确的知识,又具备流畅的语言表达,适用于医疗、法律、智能问答系统等多个知识密集型领域。
在应用实践中,RAG模型虽然有着信息完整性、推理能力和跨领域适应性等显著优势,但也面临着数据质量、计算资源消耗和知识库更新等挑战。为进一步提升RAG的性能,提出了针对数据采集、内容分块、检索策略优化以及回答生成的全面改进措施,如引入知识图谱、优化用户反馈机制、实施高效去重算法等,以增强模型的适用性和效率。
RAG在智能问答、信息检索与文本生成等领域展现了出色的应用潜力,并在不断发展的技术支持下进一步拓展至多模态生成和企业决策支持等场景。通过引入混合检索技术、知识图谱以及动态反馈机制,RAG能够更加灵活地应对复杂的用户需求,生成具有事实支撑和逻辑连贯性的回答。未来,RAG将通过增强模型透明性与可控性,进一步提升在专业领域中的可信度和实用性,为智能信息检索与内容生成提供更广泛的应用空间。
file: ./content/guide/dataset/template.en.mdx
meta: {
"title": "Template Import",
"description": "Batch-import knowledge base data from a CSV or Excel template"
}
Template import lets you add prepared content or question-answer pairs to a knowledge base in batches. FastGPT accepts `.csv` and `.xlsx` files and creates knowledge base entries from the questions, answers, indexes, and metadata in the template.
## File Structure
The first row must contain the headers. The following columns are supported:
| Header | Required | Count | Description |
| ---------- | -------- | ---------- | --------------------------------------------------------------- |
| `q` | Yes | Exactly 1 | Content or a question |
| `a` | Yes | Exactly 1 | The answer; it can be empty when importing standalone content |
| `index` | No | Repeatable | A custom index. A row can contain multiple indexes |
| `metadata` | No | At most 1 | A JSON object for custom information such as source or category |
Each row represents one knowledge base entry. `q` and `a` should not both be empty. Headers can appear in any order, but do not add unsupported headers.
### CSV Example
```csv
q,a,index,index,metadata
"What is FastGPT?","FastGPT is an AI agent development platform.","FastGPT overview","AI agent platform","{""source"":""product-doc"",""category"":""overview""}"
"How do I import knowledge base data?","Use a CSV or Excel template.","knowledge base import","template import","{""source"":""help-center""}"
```
Use UTF-8 encoding for CSV files. Cells that contain commas, line breaks, or double quotes must be escaped according to CSV rules.
### Excel Example
Excel files use the same headers and data structure as CSV files:
| q | a | index | index | metadata |
| ------------------------------------ | -------------------------------------------- | --------------------- | ----------------- | ------------------------------------------------ |
| What is FastGPT? | FastGPT is an AI agent development platform. | FastGPT overview | AI agent platform | `{"source":"product-doc","category":"overview"}` |
| How do I import knowledge base data? | Use a CSV or Excel template. | knowledge base import | template import | `{"source":"help-center"}` |
Excel files must meet these requirements:
* Use the `.xlsx` extension. `.xls` files are not supported.
* Include exactly one worksheet.
* Do not contain merged cells.
* Use the first row for the template headers.
## Import a Template
1. Open the target knowledge base and select **Template Import** from the import menu.

2. Select **Download CSV Template** for an example, or prepare an `.xlsx` file with the same structure.
3. Add your data and verify the headers, cell contents, and file format.
4. Select the file and confirm the import. You can import one file at a time.
5. After the import finishes, review the data and training status in the knowledge base collection.

## Metadata
Use `metadata` to attach structured information to each entry. The cell must contain a valid JSON object, for example:
```json
{ "source": "product-doc", "category": "overview", "version": 2 }
```
Do not use an array, plain text, or invalid JSON. In CSV files, escape the JSON according to CSV rules. In Excel files, enter the JSON string directly in the cell.
## Invalid File Format
If FastGPT reports an invalid file format, check the following:
* The file uses the `.csv` or `.xlsx` extension.
* The header row contains only supported columns.
* There is exactly one `q` column and one `a` column, with no more than one `metadata` column.
* Quotes, commas, and line breaks are correctly escaped in CSV files.
* The Excel file contains exactly one worksheet and no merged cells.
Start with a small test file. After confirming the format, import larger datasets in batches.
file: ./content/guide/dataset/template.mdx
meta: {
"title": "模板导入",
"description": "使用 CSV 或 Excel 模板批量导入知识库数据"
}
模板导入适合将已经整理好的内容或问答对批量写入知识库。FastGPT 支持导入 `.csv` 和 `.xlsx` 文件,并根据模板中的问题、答案、索引和元数据创建知识库数据。
## 文件结构
文件的第一行必须是表头。支持以下列:
| 表头 | 是否必需 | 数量 | 说明 |
| ---------- | ---- | ------ | ----------------------- |
| `q` | 是 | 1 列 | 内容或问题 |
| `a` | 是 | 1 列 | 答案;导入普通内容时可以留空 |
| `index` | 否 | 可重复 | 自定义索引,同一行可以填写多个索引 |
| `metadata` | 否 | 最多 1 列 | JSON 对象,用于保存来源、分类等自定义信息 |
每一行代表一条知识库数据,`q` 和 `a` 不应同时为空。表头顺序不受限制,但不要使用模板之外的表头。
### CSV 示例
```csv
q,a,index,index,metadata
"FastGPT 是什么?","FastGPT 是一个 AI Agent 构建平台。","FastGPT 简介","AI Agent 平台","{""source"":""product-doc"",""category"":""overview""}"
"如何导入知识库数据?","可以使用 CSV 或 Excel 模板导入。","知识库导入","模板导入","{""source"":""help-center""}"
```
CSV 文件建议使用 UTF-8 编码。如果单元格中包含逗号、换行或双引号,需要按照 CSV 规则正确转义。
### Excel 示例
Excel 文件使用与 CSV 相同的表头和数据结构:
| q | a | index | index | metadata |
| ------------ | -------------------------- | ---------- | ----------- | ------------------------------------------------ |
| FastGPT 是什么? | FastGPT 是一个 AI Agent 构建平台。 | FastGPT 简介 | AI Agent 平台 | `{"source":"product-doc","category":"overview"}` |
| 如何导入知识库数据? | 可以使用 CSV 或 Excel 模板导入。 | 知识库导入 | 模板导入 | `{"source":"help-center"}` |
Excel 文件需要满足以下要求:
* 文件扩展名为 `.xlsx`,不支持 `.xls`
* 只能包含一个工作表
* 不能包含合并单元格
* 第一行必须是模板表头
## 导入步骤
1. 打开目标知识库,在导入菜单中选择「模板导入」。

2. 点击「下载 CSV 模板」获取示例文件,或者按照相同结构准备 `.xlsx` 文件。
3. 填写数据并检查表头、单元格内容和文件格式。
4. 选择文件并确认导入。每次只能导入一个文件。
5. 导入完成后,在知识库集合中检查数据及训练状态。

## 元数据
`metadata` 用于为每条数据附加结构化信息。单元格内容应为有效的 JSON 对象,例如:
```json
{ "source": "product-doc", "category": "overview", "version": 2 }
```
不要填写数组、纯文本或包含语法错误的 JSON。CSV 中的 JSON 需要按照 CSV 规则转义;Excel 单元格中可以直接填写 JSON 字符串。
## 文件格式异常
出现「文件格式异常」提示时,请依次检查:
* 文件是否为 `.csv` 或 `.xlsx`
* 表头是否包含且仅包含支持的列
* `q`、`a` 是否各有一列,`metadata` 是否不超过一列
* CSV 的引号、逗号和换行是否正确转义
* Excel 是否只有一个工作表且没有合并单元格
建议先使用少量数据测试,确认格式正确后再分批导入大量数据。
file: ./content/guide/dataset/websync.en.mdx
meta: {
"title": "Web Site Sync",
"description": "Introduction and usage of the FastGPT Web Site Sync feature"
}

This feature is currently only available to commercial edition users.
## What is Web Site Sync
Web Site Sync uses crawler technology to automatically discover all pages under the `same domain` from an entry URL, supporting up to `200` sub-pages. For compliance and security reasons, FastGPT only supports crawling `static sites`, primarily intended for quickly building knowledge bases from documentation sites.
Tip: Most China-based media sites are not supported, including WeChat Official Accounts, CSDN, Zhihu, etc. You can verify whether a site is static by sending a `curl` request from the terminal:
```bash
curl https://doc.fastgpt.io/guide/getting-started
```
## How to Use
### 1. Create a New Knowledge Base and Select Web Site Sync


### 2. Click to Configure Site Information

### 3. Enter the URL and Selector


Click Start Sync and wait for the system to automatically crawl the site content.
## Create an App and Bind the Knowledge Base

## How to Use Selectors
Selectors are based on HTML/CSS/JS. You can use selectors to target specific content to crawl rather than the entire site. Here's how:
### Open the Browser DevTools (usually F12, or Right-click > Inspect)


### Enter the Element Selector
For a CSS selectors reference, see the [MDN CSS Selectors guide](https://developer.mozilla.org/en-US/docs/Web/CSS/CSS_selectors).
In the image above, we selected an area corresponding to a `div` tag with three attributes: `data-prismjs-copy`, `data-prismjs-copy-success`, and `data-prismjs-copy-error`. We only need one, so the selector is:
**`div[data-prismjs-copy]`**
Besides attribute selectors, class and ID selectors are also common. For example:

The `class` in the image contains class names (there may be multiple separated by spaces — just pick one). The selector would be: **`.docs-content`**
### Using Multiple Selectors
In the earlier demo, we used multiple selectors for the FastGPT documentation site, separated by commas.

We want to select content from the two tags shown above, which requires two selectors. The first is: `.docs-content .mb-0.d-flex`, meaning child elements under the `docs-content` class that have both the `mb-0` and `d-flex` classes.
The second is `.docs-content div[data-prismjs-copy]`, meaning `div` elements under the `docs-content` class that have the `data-prismjs-copy` attribute.
Separate the two selectors with a comma: `.docs-content .mb-0.d-flex, .docs-content div[data-prismjs-copy]`
file: ./content/guide/dataset/websync.mdx
meta: {
"title": "Web 站点同步",
"description": "FastGPT Web 站点同步功能介绍和使用方式"
}

该功能目前仅向商业版用户开放。
## 什么是 Web 站点同步
Web 站点同步利用爬虫的技术,可以通过一个入口网站,自动捕获`同域名`下的所有网站,目前最多支持`200`个子页面。出于合规与安全角度,FastGPT 仅支持`静态站点`的爬取,主要用于各个文档站点快速构建知识库。
Tips: 国内的媒体站点基本不可用,公众号、csdn、知乎等。可以通过终端发送`curl`请求检测是否为静态站点,例如:
```bash
curl https://doc.fastgpt.io/guide/getting-started
```
## 如何使用
### 1. 新建知识库,选择 Web 站点同步


### 2. 点击配置站点信息

### 3. 填写网址和选择器


好了, 现在点击开始同步,静等系统自动抓取网站信息即可。
## 创建应用,绑定知识库

## 选择器如何使用
选择器是 HTML CSS JS 的产物,你可以通过选择器来定位到你需要抓取的具体内容,而不是整个站点。使用方式为:
### 首先打开浏览器调试面板(通常是 F12,或者【右键 - 检查】)


### 输入对应元素的选择器
[菜鸟教程 css 选择器](https://www.runoob.com/cssref/css-selectors.html),具体选择器的使用方式可以参考菜鸟教程。
上图中,我们选中了一个区域,对应的是`div`标签,它有 `data-prismjs-copy`, `data-prismjs-copy-success`, `data-prismjs-copy-error` 三个属性,这里我们用到一个就够。所以选择器是:
**`div[data-prismjs-copy]`**
除了属性选择器,常见的还有类和ID选择器。例如:

上图 class 里的是类名(可能包含多个类名,都是空格隔开的,选择一个即可),选择器可以为:**`.docs-content`**
### 多选择器使用
在开头的演示中,我们对 FastGPT 文档是使用了多选择器的方式来选择,通过逗号隔开了两个选择器。

我们希望选中上图两个标签中的内容,此时就需要两组选择器。一组是:`.docs-content .mb-0.d-flex`,含义是 `docs-content` 类下同时包含 `mb-0`和`d-flex` 两个类的子元素;
另一组是`.docs-content div[data-prismjs-copy]`,含义是`docs-content` 类下包含`data-prismjs-copy`属性的`div`元素。
把两组选择器用逗号隔开即可:`.docs-content .mb-0.d-flex, .docs-content div[data-prismjs-copy]`
file: ./content/guide/getting-started/index.en.mdx
meta: {
"title": "Quick Overview of FastGPT",
"description": "FastGPT's capabilities and advantages"
}
import { Alert } from '@/components/docs/Alert';
import FastGPTLink from '@/components/docs/linkFastGPT';
FastGPT is an AI Agent application development platform built on large language models. It combines Knowledge Base Q\&A, visual Workflows, Agent orchestration, tool calling, and skill extensions so developers and business users can quickly build custom AI applications.
Try FastGPT now
* International: {'https://fastgpt.io'}
* China Mainland: {'https://fastgpt.cn'}
| | |
| ---------------------------------------------- | ---------------------------------------------- |
|  |  |
|  |  |
## Why FastGPT
### 1. Simple and Flexible, Like Building Blocks 🧱
Build AI applications as easily as snapping LEGO bricks together. FastGPT provides rich functional modules that let you create personalized AI apps through simple drag-and-drop — no coding required, even for complex business processes.
### 2. Make Your Data Smarter 🧠
FastGPT provides a complete data intelligence solution — from data import and preprocessing to knowledge matching and intelligent Q\&A — fully automated. Combined with visual workflow design, you can easily build professional-grade AI applications.
### 3. Open Source and Easy to Integrate 🔗
FastGPT supports custom development. Integrate quickly through standard APIs without modifying source code. It supports mainstream models including ChatGPT, Claude, DeepSeek, and ERNIE Bot, with continuous iteration to keep the product evolving.
***
## What Can FastGPT Do
### 1. Comprehensive Knowledge Base
Import documents and data with automatic knowledge structuring. Features intelligent Q\&A with multi-turn context understanding and a continuously improving knowledge base management experience.

### 2. Visual Workflow
FastGPT's intuitive drag-and-drop interface lets you build complex business processes with zero code. Rich functional node components handle diverse business needs with flexible process orchestration.

### 3. Intelligent Data Parsing
FastGPT's knowledge base system handles imported data with great flexibility — intelligently processing complex PDF structures while preserving images, tables, and LaTeX formulas. It automatically recognizes scanned files and structures content into clean Markdown format. It also supports automatic image annotation and indexing, making visual content searchable and ensuring knowledge is presented accurately in AI Q\&A.

### 4. Workflow Orchestration
Flow-based workflow orchestration lets you design complex Q\&A processes — such as querying databases, checking inventory, or booking lab resources.

### 5. Powerful API Integration
FastGPT is fully compatible with the OpenAI API interface, supporting one-click integration with WeCom, WeChat Official Account, Lark, DingTalk, and more — bringing AI capabilities into your business workflows.

***
## Core Features
* Out-of-the-box knowledge base system
* Visual low-code workflow orchestration
* Support for mainstream LLMs
* Simple and easy-to-use API interface
* Flexible data processing capabilities
***
## Knowledge Base Core Process Diagram

***
## Community
FastGPT is an open source project driven by users and contributors. If you have questions or suggestions, try the following support channels. Our team and community will do our best to help.
* 📱 Scan to join the Lark community group 👇
* 🐞 Submit any FastGPT bugs, issues, or feature requests to [GitHub Issues](https://github.com/labring/fastgpt/issues/new/choose).
file: ./content/guide/getting-started/index.mdx
meta: {
"title": "快速了解 FastGPT",
"description": "FastGPT 的能力与优势"
}
import { Alert } from '@/components/docs/Alert';
import FastGPTLink from '@/components/docs/linkFastGPT';
FastGPT 是一个基于大语言模型的 AI Agent 应用开发平台,集知识库问答、可视化工作流、Agent 编排、工具调用和技能扩展于一体,让开发者和业务人员都能快速构建专属 AI 应用。
快速开始体验
* 国际版:{'https://fastgpt.io'}
* 中国大陆版:{'https://fastgpt.cn'}
| | |
| ---------------------------------------------- | ---------------------------------------------- |
|  |  |
|  |  |
## FastGPT 的优势
### 1. 简单灵活,像搭积木一样简单 🧱
像搭乐高一样简单有趣,FastGPT 提供丰富的功能模块,通过简单拖拽就能搭建出个性化的 AI 应用,零代码也能实现复杂的业务流程。
### 2. 让数据更智能 🧠
FastGPT 提供完整的数据智能化解决方案,从数据导入、预处理到知识匹配,再到智能问答,全流程自动化。配合可视化的工作流设计,轻松打造专业级 AI 应用。
### 3. 开源开放,易于集成 🔗
FastGPT 支持二次开发。通过标准 API 即可快速接入,无需修改源码。支持 ChatGPT、Claude、DeepSeek 和文心一言等主流模型,持续迭代优化,始终保持产品活力。
***
## FastGPT 能做什么
### 1. 全能知识库
可轻松导入各式各样的文档及数据,能自动对其开展知识结构化处理工作。同时,具备支持多轮上下文理解的智能问答功能,还可为用户带来持续优化的知识库管理体验。

### 2. 可视化工作流
FastGPT 直观的拖拽式界面设计,可零代码搭建复杂业务流程。还拥有丰富的功能节点组件,能应对多种业务需求,有着灵活的流程编排能力,按需定制业务流程。

### 3. 数据智能解析
FastGPT 知识库系统对导入数据的处理极为灵活,可以智能处理 PDF 文档的复杂结构,保留图片、表格和 LaTeX 公式,自动识别扫描文件,并将内容结构化为清晰的 Markdown 格式。同时支持图片自动标注和索引,让视觉内容可被理解和检索,确保知识在 AI 问答中能被完整、准确地呈现和应用。

### 4. 工作流编排
基于 Flow 模块的工作流编排,可以帮助你设计更加复杂的问答流程。例如查询数据库、查询库存、预约实验室等。

### 5. 强大的 API 集成
FastGPT 完全对齐 OpenAI 官方接口,支持一键接入企业微信、公众号、飞书、钉钉等平台,让 AI 能力轻松融入您的业务场景。

***
## 核心特性
* 开箱即用的知识库系统
* 可视化的低代码工作流编排
* 支持主流大模型
* 简单易用的 API 接口
* 灵活的数据处理能力
***
## 知识库核心流程图

***
## 社区交流群
FastGPT 是一个由用户和贡献者参与推动的开源项目,如果您对产品使用存在疑问和建议,可尝试以下方式寻求支持。我们的团队与社区会竭尽所能为您提供帮助。
* 📱 扫码加入飞书交流群👇
* 🐞 请将任何 FastGPT 的 Bug、问题和需求提交到 [GitHub Issue](https://github.com/labring/fastgpt/issues/new/choose)。
file: ./content/guide/getting-started/quick-start.en.mdx
meta: {
"title": "Quick Start",
"description": "Quickly experience FastGPT through four use cases: Conversational Agent, Knowledge Base, Workflow, and Agent V2"
}
This article uses four complete use cases to help you quickly understand FastGPT's core application types and complete the basic setup from simple conversations to complex task orchestration.
This page is suitable for first-time FastGPT users, as well as pre-sales, delivery, operations, legal, and administrative roles who want to quickly experience the platform's capabilities. After completing this page, you will build the following in order:
1. Conversational Agent: Corporate email writing assistant.
2. Knowledge Base + Conversational Agent: Civil Code Q\&A assistant.
3. Workflow: Content review and automatic rewriting.
4. Agent V2: Intelligent data analysis Agent.
We recommend preparing the following in advance:
* An available AI model, such as GLM-5.1 or other configured models.
* A knowledge base test file, such as the Civil Code, company policies, product manuals, etc.
* If you want to test the Email tool, prepare an email SMTP authorization code.
* If you want to test Agent V2 data analysis, prepare a sample Excel or CSV file.
When reading, focus on three things: what kind of problems each application type is suitable for, why the key configurations are written this way, and what to observe during validation. The parameters and prompts in this article are reusable starting points; for production deployment, you can replace them with your own business materials, review rules, notification channels, and data files.
## Case 1: Conversational Agent — Corporate Email Writing Assistant
### 1.1 Use Cases
Conversational Agents are suitable for lightweight Q\&A, content generation, copy refinement, and standardized output. This case does not link a knowledge base or rely on complex workflows; it simply uses model configuration, prompts, and an Email tool to build a corporate email writing assistant.
Employees often need to handle emails for project updates, cross-departmental collaboration, customer replies, meeting minutes, and more. Using AI to standardize email formats and expression styles can improve writing efficiency and maintain the professionalism of external corporate communications.
The focus of this case is not to have the AI randomly generate an email, but to consolidate the stable requirements of corporate email writing into the prompt, such as subject lines, salutations, body structure, action items, and risk reminders. For beginners, it also serves as a minimal closed loop for understanding FastGPT application configuration: first define the role, then constrain the output, and finally extend execution actions through tools.
### 1.2 Configuration Steps
1. **Create a Conversational Agent**
Click "New" in the workspace, select "Conversational Agent", and fill in the application name as `Corporate Email Writing Assistant`.

2. **Enter the Application Configuration Page**
After creation, enter the application configuration page. The page is usually divided into left and right columns: the left is AI configuration, and the right is debug preview.

3. **Select a Model**
In the AI configuration, select the base model for the Conversational Agent. This case uses GLM-5.1, but you can replace it with other available models configured in your current environment.
When selecting a model, prioritize two aspects: first, whether the model is good at business writing, and second, whether it consistently follows the required format. Email writing is a low-risk, low-structure task, so you usually do not need the strongest model available, but you should still choose one with a natural tone and reliable instruction following.

4. **Write the Prompt**
The prompt needs to clearly describe the assistant's role, output format, and constraints. You can use the following example:
```md
You are a corporate email writing expert, helping employees write professional, clear, and appropriate work emails.
Output format:
- **Subject line**: Concise and clear
- **Salutation**: Choose "Dear Mr./Ms. XX" or "Hi XX" based on the relationship with the recipient
- **Body**: Three-part structure (Background → Core content → Action items)
- **Sign-off**: Name, Title, Department
Rules:
- Keep the body between 200-500 words
- List action items and to-dos with bullet points
- When involving sensitive content like salary, HR, or legal matters, remind the user to send with caution
- Use a neutral and polite tone when unsure of the relationship with the recipient
```
Write the prompt into the prompt module.
This prompt consists of three parts: the role definition stabilizes the assistant's identity, the output format constrains the email structure, and the rules control risk boundaries. In actual business, you can continue to add corporate tone requirements, brand terminology, banned words, signature formats, and more to make the output more aligned with internal standards.

5. **Add the Email Tool**
Tools encapsulate complex operations. This case uses the Email sending tool to give the AI assistant the ability to send emails after generating them. Click the plus sign on the right side of the tools and select "Send Email".
You can think of tools as the AI's external capabilities. Without tools, the assistant can only generate the email body; after adding the Email tool, the assistant can execute the sending action after user confirmation. In a production environment, it is recommended to have the assistant generate an email draft first, and then have the user confirm the sending, to avoid accidental sending or sending to the wrong recipient.

6. **Configure the Email Tool**
Enter the tool configuration page and fill in the parameters related to the email service.

7. **Activate the Tool**
Click "Settings" to enter the tool activation page, then click "Activate Tool".

8. **Fill in the Email SMTP Information**
This case uses QQ Mail as an example. When testing, you can fill it out as follows:
```text
SMTP Server Address: smtp.qq.com
SMTP Port: 465
Enable SSL
SMTP Username: Email address
SMTP Password: Authorization code
```
For the authorization code, please refer to the [Authorization Code Acquisition Tutorial](https://cloud.tencent.com/developer/article/2177098).

9. **Set the Conversation Opening**
The conversation opening is used to tell users what this AI assistant can do. You can fill in:
```text
Hello! I am the Email Writing Assistant 📧
Please tell me: Who is the recipient? What is the purpose of the email? What key information needs to be included?
I will help you generate a professional and appropriate email.
```
After filling it out, you can see the effect in the preview area on the right.

10. **Validate the Result**
After configuration, enter your email requirements in the debug preview to check whether the assistant can generate an email with a clear structure and appropriate tone.

If the user's information is insufficient, the assistant should proactively prompt to supplement the recipient, email purpose, key information, etc.

When verifying, it is recommended to test at least three types of inputs: an email requirement with complete information, a vague requirement missing the recipient or purpose, and an email requirement containing sensitive information. An email assistant ready for production not only needs to be able to "write", but also needs to be able to ask follow-up questions when information is insufficient, and remind users to send with caution in sensitive scenarios.
### 1.3 Business Value
* **Improve Writing Efficiency**: Transform high-frequency emails like project updates, customer follow-ups, and meeting minutes from "writing from scratch" to "generating after filling in key information", reducing repetitive labor.
* **Unify Communication Standards**: Consolidate standards for salutations, body structure, action items, and sign-offs into the Prompt, reducing the communication costs caused by differences in writing styles among employees.
* **Reduce Sending Risks**: Add sensitive information reminders, missing information follow-ups, and neutral tone constraints through the Prompt to reduce incomplete, inappropriate, or over-promising emails.
* **Expand Office Automation**: Combined with the Email tool, you can continue to integrate office workflows such as notifications, approvals, Lark, DingTalk, and WeCom, extending email writing from content generation to business action execution.
## Case 2: Knowledge Base + Conversational Agent — Civil Code Q\&A Assistant
### 2.1 Use Cases
A knowledge base is ideal for scenarios where answers must be based on provided materials, such as policy Q\&A, product manual Q\&A, legal article retrieval, and customer service knowledge support. Without a knowledge base, the AI primarily relies on its own model capabilities to answer questions. Once a knowledge base is connected, the AI first searches your materials and then organizes answers based on the search results.
In this case, we import the *Civil Code of the People's Republic of China* into the knowledge base and create a Civil Code Q\&A assistant. When users ask questions in natural language, the assistant should prioritize citing the original Civil Code text to reduce fabricated responses.
Think of the knowledge base as a "reference room" configured for the AI. The model itself has general knowledge but doesn't know your company policies, product details, internal processes, or specific regulatory versions. The knowledge base turns these materials into searchable content, allowing the AI to look up information before organizing an answer. For legal, policy, customer service, and after-sales scenarios, this is more controllable than relying solely on the model's memory.
### 2.2 Preparation
* Example file: Prepare a copy of the *Civil Code of the People's Republic of China* or another regulatory/documentation file.
* A usable conversation model and vector model.
* A legal question for testing, e.g., "My lease isn't up yet, but the landlord wants to sell the house. What should I do?"
### 2.3 Configuration Steps
1. **Create a Knowledge Base**
Click **Knowledge Base** on the left side of the homepage, then click **New** in the top right corner. Name the knowledge base `Civil Code Q&A Assistant`. Keep the other default settings for this case.

2. **Create a Text Dataset**
Click **New**, then select **Text Dataset** to import local documents.

3. **Upload a Local File**
Click **Upload Local File** and select the example file. In real use, you can upload multiple files at once; for this case, we upload only one file as a demonstration.

4. **Set Parsing Parameters**
After entering the parameter settings page, choose the parsing method based on the file type.

Recommended common settings:
* **File Parsing Settings**: Enable when uploading PDFs; for regular Word, Markdown, TXT files, start with the default configuration.
* **Processing Method**: For most scenarios, choose chunked storage — it's lower cost and faster for retrieval.
* **Chunking Conditions**: Controls how many tokens each chunk contains. Use the default values for quick testing.
* **Index Enhancement**: For plain text, usually check the first two options; if the document contains images, enable image-related enhancements.
These parameters directly affect the quality of subsequent Q\&A. If chunks are too large, search results may include too much irrelevant content, making answers verbose. If chunks are too small, key context may be split, causing answers to lack supporting evidence. For the quick-start phase, use the default configuration. When officially integrating enterprise policies, contracts, or product manuals, adjust these parameters gradually based on Q\&A performance.
5. **Preview Chunking Results**
After reaching the data preview step, check whether the chunks are complete and readable. If chunks are too long or too short, go back and adjust the parameters.
When previewing chunks, focus on three things: whether paragraphs are abnormally cut off, whether titles and body text remain within the same semantic range, and whether tables, clauses, or numbering are still readable. If the quality of these knowledge fragments is unstable, even a well-written Prompt in the application will struggle to consistently produce accurate answers.

6. **Wait for the Knowledge Base to Be Ready**
Once the data status changes to "Ready," the knowledge base can be referenced by applications.

7. **Create and Link a Conversational Agent**
Create a new Conversational Agent, also named `Civil Code Q&A Assistant`. After creation, link the knowledge base you just created in the application configuration.

After linking the knowledge base, the application's response flow changes from simple conversation to "user question → knowledge base search → model summarizes answer." This is the key difference between Case 1 and Case 2: Case 1 emphasizes content generation, while Case 2 emphasizes answering based on provided materials.
8. **Configure the Q\&A Prompt**
The Civil Code Q\&A assistant needs to emphasize "answer based on the knowledge base" and "cite the original text." You can use the following Prompt:
```md
You are a professional Civil Code Q&A assistant, answering legal questions based on the original text of the _Civil Code of the People's Republic of China_.
Rules:
- Strictly answer based on the Civil Code articles retrieved from the knowledge base; do not fabricate legal provisions.
- Every answer must cite the original Civil Code text (book, chapter, article).
- If there is no directly corresponding provision in the Civil Code, state this honestly and do not give legal advice.
- When applying the law to specific cases, remind the user: "This answer is for reference only; please consult a professional lawyer."
- Provide plain-language explanations of legal terms so that users without a legal background can understand.
Output format:
1. **Legal Conclusion** (1–3 sentence summary)
2. **Relevant Article Citation** (original excerpt + book/chapter/article number)
3. **Plain-Language Explanation** (explain the meaning of the article in everyday language)
4. **Practical Advice** (2–3 actionable suggestions)
5. **Disclaimer** ("This answer is based on the original Civil Code text and does not constitute legal advice. For specific cases, please consult a professional lawyer.")
```
For legal Q\&A scenarios, it's especially important to define boundaries: what can be answered is a general explanation based on the materials; the model's output should never be packaged as formal legal advice. Requiring article citations, stating uncertainty, and adding disclaimers in the Prompt all aim to make the output more traceable and compliant with high-risk knowledge Q\&A usage norms.
9. **Configure the Opening Message**
The opening greeting can include a few example questions to help users quickly understand how to use this assistant:
```text
Hello! I'm the Civil Code Q&A Assistant ⚖️
I answer legal questions based on the original text of the *Civil Code of the People's Republic of China*.
You can ask me:
["My lease isn't up yet, but the landlord wants to sell the house. What should I do?"]
["I received a product I bought online and it was broken, but the seller won't accept a return. What does the law say?"]
["The upstairs neighbor's leak soaked my ceiling. Can I claim compensation?"]
⚠️ Note: My answers are for reference only. For specific legal issues, please consult a professional lawyer.
```
If you add `[question content]` in the opening greeting, users can click the question to ask it directly — useful for demonstrations and guidance.

10. **Verify Q\&A Performance**
Ask a question related to the Civil Code and check whether the answer includes a legal conclusion, article citation, plain-language explanation, and disclaimer.

When verifying, don't just check whether the answer "looks like a legal answer" — also check whether it actually references the knowledge base content. It's recommended to test with questions inside the materials, outside the materials, and ambiguous questions: questions inside the materials should cite the original text; questions outside the materials should state that a direct confirmation is not possible; ambiguous questions should proactively prompt the user to provide more facts.
### 2.4 Business Value
* **Lower the barrier to material retrieval**: Turn lengthy regulations, policies, and manuals into a natural language Q\&A entry point, allowing business users to quickly find relevant content without needing to know keywords or directory locations first.
* **Improve answer credibility**: Through knowledge base retrieval and original text citations, answers have a source basis, reducing the risk of the model fabricating or giving vague responses based on experience.
* **Accumulate organizational knowledge**: Internal company policies, product FAQs, after-sales SOPs, contract templates, and other materials can be continuously added to the knowledge base, forming maintainable and reusable knowledge assets.
* **Adapt to high-frequency support scenarios**: Legal, HR, administrative, customer service, and delivery teams can all use a similar model to turn repetitive inquiries into self-service Q\&A, improving response efficiency.
## Case 3: Workflow — Content Review and Automatic Rewriting
### 3.1 Use Cases
Workflows are ideal for tasks with fixed steps, clear logic, and the need for branching decisions or human confirmation. This case breaks down content compliance review into several stages: knowledge base retrieval, AI classification, conditional branching, automatic rewriting, rejection explanation, and human confirmation, simulating the review process before enterprise content is published.
Its core value lies in two aspects:
1. **Automation**: Once triggered, the system automatically executes multiple steps according to the preset workflow.
2. **Standardization**: The same input goes through the same judgment and processing, reducing human variability.
To determine whether a task is suitable for a workflow, check if it has three characteristics: "stable steps, clear rules, and repeatable execution." Content review is a typical scenario: the input is content to be published, the rules come from a compliance knowledge base, and the output is usually pass, rewrite, or reject, with the option to add human confirmation in between. This improves processing efficiency while retaining risk control.
### 3.2 Preparation
First, create a "Content Compliance Rules" knowledge base and upload a simple text file. You can directly use the following rule template:
```text
Content Compliance Rules
Safe Content (can be published directly)
- Objective factual statements
- Normal event notifications, meeting arrangements
- Product feature descriptions (based on real data)
- Industry knowledge sharing
Sensitive Wording (needs rewriting)
- Absolute language: "best," "first," "100%," "absolute," "only"
- Exaggerated claims: "disrupting the industry," "unprecedented," "unmatched"
- Unverified data: conversion rates, satisfaction rates, growth rates without sources
- Comparative disparagement: directly naming competitors and belittling them
Prohibited Content (must not be published)
- Illegal information: involving pornography, gambling, drugs, fraud, pyramid schemes
- Personal attacks: insults, defamation against individuals or groups
- False information: fabricated data, forged qualifications, impersonating official sources
- Sensitive topics: political sensitivity, religious discrimination, regional attacks
```
The rule base does not need to be complex at the start. For quick validation, split the rules into three categories: "safe, needs rewriting, prohibited." For formal use, you can continue adding rules by industry, brand, channel, or region, such as advertising-sensitive terms, medical compliance requirements, prohibited financial marketing claims, brand tone guidelines, and so on.
### 3.3 Workflow Design
Before configuring nodes, it's recommended to confirm the complete workflow. This case can be designed as follows:
1. **User Input**: As the workflow starting point, receives the content to be reviewed.
2. **Knowledge Base Retrieval**: Recalls relevant rules from the content compliance rules base. It is recommended to set the citation limit to 1-2 entries.
3. **AI Content Compliance Classification**: Combines user input and retrieval results to classify the content as "Safe / Sensitive but Rewritable / Prohibited."
4. **Conditional Branching**: Enters different branches based on the classification result.
5. **Safe Branch**: Directly outputs the original text, indicating it can be published.
6. **Sensitive but Rewritable Branch**: Calls AI to rewrite the content, then submits it for user confirmation.
7. **Prohibited Branch**: Outputs a rejection explanation and provides revision suggestions.
8. **Final Output**: Returns the original text, rewritten version, or rejection explanation.
When designing a workflow, first determine the responsibility of each node to avoid having a single AI node simultaneously handle "retrieving rules, judging classification, rewriting content, explaining reasons," etc. Once responsibilities are clearly separated, each node's prompt will be shorter and more stable, making it easier to locate issues later: whether the knowledge base failed to recall rules, the review node misclassified, or the condition in the decision node didn't match.
### 3.4 Configuration Steps
1. **Create a Workflow**
Go to the workflow homepage, click "New Workflow," and name it `Content Review and Automatic Rewriting`.

After creation, enter the workflow editing page.

The complete workflow diagram for this case is as follows:

2. **Configure the Opening Message**
In the system configuration, fill in the opening statement to explain the assistant's review rules and usage.
```md
Hello! I am the Content Compliance Review Assistant 🛡️
Please send me the content you need reviewed, and I will automatically judge it according to the rules:
- **✅ Safe** — Content has no sensitive information and can be published directly
- **⚠️ Sensitive but Rewritable** — Contains correctable wording; I will rewrite it and send it back for your confirmation
- **🚫 Prohibited** — Contains red-line content and is rejected with an explanation
Supports single text review or batch submission (multiple items separated by line breaks).
**Let's get started.**
```
3. **Call the Knowledge Base**
Add a knowledge base retrieval node in the workflow, select the previously created content compliance rules base. The input for this node uses the user input, and the output is referenced by subsequent AI nodes.
The purpose of this node is not to have the model read the entire rule base, but to retrieve the rule fragments most relevant to the current content for the subsequent review node. You can initially set the citation limit to 1-2 entries to keep the prompt context focused; if the rule base is larger or the categories are more detailed, gradually increase the citation count.


4. **Configure the Content Review Node**
Add an AI dialogue node and rename it to "Content Review Node." This node is responsible for outputting the classification result based on the user input and knowledge base rules.
```md
You are a content compliance review expert. Based on the compliance rules retrieved from the knowledge base, determine the compliance level of the user's input content.
Classification criteria:
- Safe: Content has no sensitive information and can be published directly
- Sensitive but Rewritable: Contains correctable sensitive wording; can be published after rewriting
- Prohibited: Contains red-line content and cannot be published
Output format:
Output only one of three labels: Safe / Sensitive but Rewritable / Prohibited
```
Note: The knowledge base reference must select the previously created knowledge base, otherwise the AI cannot read the rule content.
The output of the review node should be as stable as possible, because it directly affects the decision node's branching. For a quick demo, you can output only the three labels "Safe / Sensitive but Rewritable / Prohibited"; if you need stricter automation integration later, you can switch to structured output, such as returning the classification, reason, and matched rules together, making it easier for downstream nodes to perform precise matching and logging.

5. **Test the Review Node**
Click "Run" in the top right corner, input a piece of content to be reviewed, and confirm that the review node returns a stable classification result.
Test this node individually instead of waiting until the entire workflow is built. Input a normal event announcement, marketing copy containing absolute language, and clearly prohibited text, and observe whether the classification meets expectations. If the classification is unstable, first adjust the rule base wording and the review prompt before continuing to configure subsequent branches.

6. **Hide Intermediate Output**
Click the settings button on the right side of the model to adjust the node's basic settings. Since the user only needs the final result, it is recommended to hide the intermediate output of the content review node.

7. **Configure the Decision Node**
Add a decision node to branch based on the "Safe / Sensitive but Rewritable / Prohibited" output from the content review node.
The decision node acts as a switch in the workflow. The more stable the output from the previous step, the easier it is to configure the decision node; if the review node's output contains explanatory text, the condition may fail to match. Therefore, in branching workflows, it is common to first constrain the output format of the upstream node before configuring the decision node conditions.

8. **Configure the Safe Branch**
When the review result is "Safe," directly output the original text. In a production scenario, you could also add spot-checking or human confirmation.

The output effect is as follows:

9. **Configure the Sensitive but Rewritable Branch**
When the review result is "Sensitive but Rewritable," add an AI dialogue node to automatically rewrite the content. You can use the following prompt:
```md
You are a content rewriting expert. Rewrite the user's input content into a compliant version.
Rewriting principles:
- Preserve the original meaning and information, do not change the core expression
- Replace absolute language with objective statements, e.g., "best" → "industry-leading"
- Replace unverified data statements with reasonable speculation, e.g., "100% effective" → "most users report it effective"
- Replace sensitive wording with neutral expressions
- Maintain the original style and tone
```
After rewriting, you can add a user choice node to let the user accept the rewrite, continue rewriting, or abandon it.
The human confirmation node is suitable for paths that are risky but correctable. AI can propose rewrite suggestions, but whether to publish is still left to the business personnel to confirm. This reduces the time spent on initial screening and repeated revisions while not handing over the final publishing decision entirely to the automated workflow.


If the user accepts the rewrite, output the rewritten result; if the user chooses to continue rewriting, it can be connected back to the content rewriting node; if the user abandons it, output the original text.

10. **Configure the Prohibited Branch**
When the review result is "Prohibited," add an AI dialogue node to output a rejection explanation. You can use the following prompt:
```md
The content submitted by the user cannot be published because it contains prohibited information. Please explain the reason in a polite and professional tone.
Output format:
1. One sentence stating that the content cannot be published
2. List specific violation points (1-3 items)
3. Provide alternative suggestions (e.g., suggest which aspects to modify before resubmitting)
```
11. **Verify the Complete Workflow**
Input safe content, sensitive content, and prohibited content separately, and confirm that all three branches return the expected results.

When verifying the complete workflow, it is recommended to record the input, review classification, branch entered, and final output for each test case. If the output does not meet expectations, troubleshoot node by node: first check whether the knowledge base recalled the correct rules, then whether the review node classified accurately, and finally whether the decision node conditions and branch output are configured correctly.
### 3.5 Business Value
* **Standardize review rules into a workflow**: Embed the judgment criteria for pre-publication content review into a knowledge base and node configuration, reducing reliance on personal experience passed down verbally.
* **Improve processing efficiency**: Safe content can pass quickly, sensitive content is automatically rewritten, and prohibited content receives a direct explanation, allowing reviewers to focus on content that requires judgment.
* **Retain human control points**: Add user confirmation for gray-area scenarios like "sensitive but rewritable," preventing the automated workflow from making publishing decisions directly for business personnel.
* **Easy to reuse and extend**: The same workflow can be applied to marketing copy, customer service scripts, announcements, event pages, short video scripts, and other content review scenarios by replacing the rule base and a few prompts.
* **Reduce compliance risk**: Through fixed branches and rejection explanations, high-risk content has a clear interception path, reducing the risk of accidental publication, exaggerated claims, or non-compliant expressions.
## Case 4: Agent V2 — Intelligent Data Analysis Agent
### 4.1 Use Cases
Agent V2 is ideal for open-ended, multi-step tasks that require dynamic planning. Unlike workflows, where every step must be predefined, Agent V2 is better suited for tasks with “unfixed steps,” such as data analysis, file processing, multi-tool collaboration, and complex problems that require follow-up clarification.
This case simulates the daily data analysis needs of operations, product, or sales teams: upload an Excel file, ask questions in natural language, and let the Agent autonomously read the file, formulate an analysis plan, and execute the analysis in a virtual machine.
In Case 3, the workflow required you to design a fixed path of “retrieval rules → classification → branching → rewriting or rejection.” Agent V2, on the other hand, acts more like an autonomous executor. You don’t need to define every step in advance; just provide the goal and the file. The Agent decides whether to read the file first, perform statistics, ask follow-up questions, or run code based on the data structure and problem complexity.
Data analysis is a great way to experience Agent V2 because it naturally has three characteristics: the analysis path is not fixed, it often requires multi-step reasoning, and the requirements may need clarification. With the same Excel file, different users might care about product sales, channel ROI, regional trends, or anomalous orders. A fixed workflow can hardly cover all paths in advance, but Agent V2 can dynamically plan based on the question.
### 4.2 Preparation
* Sample file: Prepare a sales data Excel or CSV table.
* A conversation model that supports Agent V2.
* Virtual machine capability is available in the current environment.
### 4.3 Configuration Steps
1. **Create an Agent V2 Application**
In the workspace, click “Create Application,” select `Conversational Agent V2`, and name it `Intelligent Data Analysis Agent`.

2. **Configure Model and Prompt**
Continue using the GLM-5.1 model for AI configuration. The system prompt can refer to:
```md
You are a senior data analyst Agent. You can read data files uploaded by users, run Python analysis code in the sandbox, and proactively ask users for clarification.
Tools:
- 📄 **Read File** — Read Excel/CSV data uploaded by the user
- 💻 **Sandbox Execution** — Run Python scripts (pandas/matplotlib/numpy)
- ❓ **Proactive Follow-up** — Confirm with the user when analysis requirements are unclear
Workflow:
1. After receiving data and a question, first read the file to understand the data structure and content
2. Formulate an analysis plan and present it to the user as a list of steps
3. Execute step by step according to the plan, showing key findings at each step
4. Proactively ask follow-up questions when encountering ambiguous requirements, such as unclear metric definitions or missing comparison baselines
5. Dynamically update the plan based on follow-up results
6. Output an analysis report containing data overview, core findings, visual charts, and business recommendations
Security Rules:
- Only data analysis is allowed in the sandbox; no network access, no writing files to the host machine
- Data is used only for this analysis; do not expose raw sensitive data in the report
- Mark confidence levels for uncertain conclusions
Output Format:
1. 📋 Analysis Plan (automatically generated based on data)
2. 📊 Data Overview (row count, column names, missing values, basic statistics)
3. 🔍 Core Findings (3-5 key insights, supported by charts)
4. 💡 Business Recommendations (actionable suggestions based on data)
```
The key point of this prompt is to have the Agent plan before executing, rather than jumping straight to conclusions. For data analysis tasks, first reading the data structure, confirming field meanings, and formulating an analysis plan can significantly reduce the risk of misunderstanding requirements or misusing metrics. In production use, you can continue to supplement internal metric definitions, such as GMV, ROI, conversion rate, active customers, repurchase rate, etc.
3. **Configure the Opening Message**
The opening remarks guide the user to upload a data file and ask analysis questions:
```text
Hello! I am the Intelligent Data Analysis Agent 📊
Just drag and drop your Excel or CSV file here and tell me what you want to analyze.
For example:
- "Analyze this sales data and find the best-selling products and trends"
- "Help me look at changes in user activity and find the reasons for the decline"
- "Compare the conversion rates of three channels, which one has the highest ROI?"
I will first understand your data, formulate an analysis plan, and then run the analysis code in the sandbox.
I will proactively ask you for clarification when needed.
```
4. **Enable Virtual Machine Capability**
Data analysis usually requires reading files and executing code, so the virtual machine configuration needs to be enabled.
The virtual machine capability is used to isolate the code execution environment. The Agent can run data analysis scripts, read uploaded files, and generate statistical results within it, without directly affecting the local host environment. For scenarios that require running Python, processing Excel, drawing charts, or performing batch calculations, this is an important capability that distinguishes Agent V2 from ordinary conversation applications.

5. **Upload File and Test**
Upload the sample sales data file and enter the question:
```text
Help me analyze this sales data to see which products sell well and which channel has the highest ROI
```
The Agent will first break down the task and then execute the analysis step by step.

During execution, you can see the task running in the virtual machine without affecting the local environment.

When testing, focus on whether the Agent has a complete analysis process: does it first identify the table fields, explain the analysis plan, run code when needed, and give business recommendations based on the results? If the problem description is unclear, the ideal behavior is not to force an analysis but to first ask the user for key definitions.
### 4.4 Verification of Results
The final result should include the analysis plan, data overview, core findings, and business recommendations.

When verifying results, it is recommended to focus on four dimensions: whether the conclusions come from actual data, whether the metric definitions are clear, whether the charts or statistics support the conclusions, and whether the business recommendations are actionable. The value of a data analysis Agent is not just to output a summary, but to connect the process of “reading data, calculating metrics, explaining results, and proposing recommendations.”
### 4.5 Business Value
* **Lower the barrier to data analysis**: Business users can directly upload Excel or CSV files and ask questions in natural language, without needing to write SQL, Python, or complex formulas first.
* **Support open-ended exploration**: The same data can be repeatedly queried around sales, channels, regions, customers, trends, outliers, etc., suitable for scenarios without a fixed analysis path.
* **Increase analysis transparency**: The Agent shows the analysis plan and key steps, so users can see how it understands the data, calculates metrics, and draws conclusions.
* **Isolate code execution risk**: By executing analysis scripts in a virtual machine, risks of local environment contamination, dependency conflicts, and permission misuse are reduced.
* **Consolidate business analysis capabilities**: Scenarios such as sales reviews, weekly operations reports, campaign attribution, and product metric diagnostics can all reuse this type of Agent, turning data analysis from an expert task into a daily workflow.
## Choosing Between the Four Types
After completing the four cases, you can understand the common application types of FastGPT as a progression from simple to complex capabilities: Conversational Agent solves "how to answer and generate content," Knowledge Base solves "what materials to base answers on," Workflow solves "what fixed process to follow," and Agent V2 solves "how to autonomously plan and execute open-ended tasks."
| Application Type | Suitable Scenarios | Core Capabilities |
| ------------------------------------- | --------------------------------------------------------------------- | ------------------------------------------------------------ |
| Conversational Agent | Lightweight Q\&A, copywriting, standardized output | Prompt, model configuration, tool calling |
| Knowledge Base + Conversational Agent | Q\&A based on documents, policies, regulations, product manuals | File import, knowledge base retrieval, citing sources |
| Workflow | Fixed steps, conditional branches, review flows, automated processing | Node orchestration, decision nodes, human confirmation |
| Agent V2 | Data analysis, complex tasks, multi-step reasoning, dynamic planning | Autonomous planning, tool calling, virtual machine execution |
When selecting, you can judge based on task complexity:
1. If it's just lightweight conversation or standardized copy generation, prioritize the Conversational Agent.
2. If answers must be based on existing materials, choose Knowledge Base + Conversational Agent.
3. If the process is fixed and requires conditional branches, human confirmation, or automated processing, choose Workflow.
4. If the task is open-ended with unfixed steps and requires autonomous analysis, tool calling, or code execution, choose Agent V2.
In real projects, you can also combine these capabilities. For example, a customer service assistant can use a Knowledge Base to answer product questions and then query orders via tools; content review can use a Workflow to fix the review path while maintaining rules with a Knowledge Base; data analysis scenarios can first use Agent V2 for exploration, then solidify stable analysis steps into a Workflow.
file: ./content/guide/getting-started/quick-start.mdx
meta: {
"title": "快速上手",
"description": "通过对话 Agent、知识库、工作流和 Agent V2 四个案例快速体验 FastGPT"
}
本文通过四个完整案例,帮助你快速理解 FastGPT 的核心应用类型,并完成从简单对话到复杂任务编排的基础搭建。
本页适合第一次接触 FastGPT 的用户,也适合售前、交付、运营、法务、行政等角色快速体验平台能力。完成本页后,你将依次搭建:
1. 对话 Agent:企业邮件撰写助手。
2. 知识库 + 对话 Agent:民法典问答助手。
3. 工作流:内容审核与自动改写。
4. Agent V2:智能数据分析 Agent。
建议提前准备以下内容:
* 可用的 AI 模型,例如 GLM-5.1 或其他已配置模型。
* 一个知识库测试文件,例如民法典、公司制度、产品手册等。
* 如果要测试 Email 工具,准备邮箱 SMTP 授权码。
* 如果要测试 Agent V2 数据分析,准备一个 Excel 或 CSV 示例文件。
阅读时建议重点关注三件事:每种应用类型适合解决什么问题、关键配置为什么这样写、验证时应该观察哪些效果。本文中的参数和 Prompt 都是可复用的起点,正式落地时可以替换为你自己的业务资料、审核规则、通知渠道和数据文件。
## 案例一:对话 Agent—企业邮件撰写助手
### 1.1 适用场景
对话 Agent 适合轻量问答、内容生成、文案润色、标准化输出等场景。本案例不关联知识库,也不依赖复杂流程,只通过模型配置、Prompt 和 Email 工具完成一个企业邮件撰写助手。
企业员工经常需要处理项目同步、跨部门协作、客户回复、会议纪要等邮件。通过 AI 统一邮件格式和表达风格,可以提升写作效率,也能保持企业对外沟通的专业性。
这个案例的重点不是让 AI 随意生成一封邮件,而是把企业邮件写作中的稳定要求沉淀到 Prompt 中,例如主题行、称呼、正文结构、行动项和风险提醒。对于新手来说,它也是理解 FastGPT 应用配置的最小闭环:先定义角色,再约束输出,最后通过工具扩展执行动作。
### 1.2 配置步骤
1. **创建对话 Agent**
在工作台点击新建,选择对话 Agent,应用名称填写为 `企业邮件撰写助手`。

2. **进入应用配置页**
创建完成后进入应用配置页。页面通常分为左右两栏:左侧是 AI 配置,右侧是调试预览。

3. **选择模型**
在 AI 配置中选择对话 Agent 使用的基础模型。本案例使用 GLM-5.1,你也可以替换为当前环境中已配置的其他可用模型。
选择模型时优先关注两点:一是模型是否擅长中文商务表达,二是输出是否稳定遵守格式。邮件撰写属于低风险、低结构复杂度任务,通常不需要过度追求最强模型,但要确保语气自然、指令遵循能力较好。

4. **编写 Prompt**
Prompt 需要清晰描述助手的角色、输出格式和约束。可以使用以下示例:
```md
你是一个企业邮件撰写专家,帮助员工撰写专业、清晰、得体的工作邮件。
输出格式:
- **主题行**:简洁明确
- **称呼**:根据收件人关系选择“尊敬的 XX 总”或“Hi XX”
- **正文**:三段式(背景 → 核心内容 → 行动项)
- **落款**:署名、职位、部门
规则:
- 正文控制在 200-500 字
- 行动项和待办用项目符号列出
- 涉及薪资、人事、法律等敏感内容时,提醒用户谨慎发送
- 不确定收件人关系时使用中性礼貌语气
```
将 Prompt 写入提示词模块。
这段 Prompt 由三部分组成:角色定义用于稳定助手身份,输出格式用于约束邮件结构,规则用于控制风险边界。实际业务中可以继续补充企业语气要求、品牌用语、禁用词、签名格式等内容,让输出更贴合内部规范。

5. **添加 Email 工具**
工具是对复杂操作的封装。本案例使用 Email 邮件发送工具,让 AI 助手在生成邮件后具备发送能力。点击工具右侧的加号,选择 Email 邮件发送。
可以把工具理解为 AI 的外部能力。没有工具时,助手只能生成邮件正文;添加 Email 工具后,助手可以在用户确认后执行发送动作。正式环境中建议先让助手生成邮件草稿,再由用户确认发送,避免误发或发送给错误收件人。

6. **配置 Email 工具**
进入工具配置页,填写邮件服务相关参数。

7. **激活工具**
点击设置后进入工具激活页面,再点击工具激活。

8. **填写邮箱 SMTP 信息**
本案例以 QQ 邮箱为例。测试时可以按以下方式填写:
```text
SMTP 服务器地址:smtp.qq.com
SMTP 端口:465
启用 SSL
SMTP 用户名:邮箱地址
SMTP 密码:授权码
```
授权码可参考 [授权码获取教程](https://cloud.tencent.com/developer/article/2177098)。

9. **设置对话开场白**
对话开场白用于告诉用户这个 AI 助手可以做什么。可以填写:
```text
你好!我是邮件撰写助手 📧
请告诉我:收件人是谁?邮件目的是什么?需要包含哪些关键信息?
我会帮你生成一封专业、得体的邮件。
```
填写后,可以在右侧预览区域看到效果。

10. **验证运行效果**
配置完成后,在调试预览中输入邮件需求,检查助手是否能生成结构清晰、语气得体的邮件。

如果用户信息不足,助手应主动提示补充收件人、邮件目的、关键信息等内容。

验证时建议至少测试三类输入:信息完整的邮件需求、缺少收件人或目的的模糊需求、包含敏感信息的邮件需求。一个可上线的邮件助手不只要“能写”,还要能在信息不足时追问,在敏感场景下提醒用户谨慎发送。
### 1.3 业务价值
* **提升写作效率**:将项目同步、客户跟进、会议纪要等高频邮件从“从零写”变成“填写关键信息后生成”,减少重复劳动。
* **统一沟通标准**:把称呼、正文结构、行动项、落款等规范固化到 Prompt 中,降低不同员工写作风格差异带来的沟通成本。
* **降低发送风险**:通过 Prompt 增加敏感信息提醒、信息缺失追问和中性语气约束,减少不完整、不恰当或过度承诺的邮件。
* **扩展办公自动化**:配合 Email 工具后,可以继续接入通知、审批、飞书、钉钉、企业微信等办公流程,将邮件撰写从内容生成延伸到业务动作执行。
## 案例二:知识库 + 对话 Agent—民法典问答助手
### 2.1 适用场景
知识库适合“回答必须基于资料”的场景,例如制度问答、产品手册问答、法律条文检索、客服知识支持等。没有知识库时,AI 主要依赖模型自身能力回答;接入知识库后,AI 会先检索你的资料,再基于检索结果组织答案。
本案例将《中华人民共和国民法典》导入知识库,并创建一个民法典问答助手。用户用自然语言提问时,助手需要优先引用民法典原文,减少凭空生成。
可以把知识库理解为给 AI 配置的“资料室”。模型本身具备通用知识,但不了解你的公司制度、产品细节、内部流程或指定法规版本;知识库把这些资料变成可检索内容,让 AI 在回答前先查资料,再组织答案。对于法律、制度、客服、售后等场景,这比单纯依赖模型记忆更可控。
### 2.2 准备内容
* 示例文件:准备一份《中华人民共和国民法典》或其他法规制度类文档。
* 一个可用的对话模型和向量模型。
* 一个用于测试的法律问题,例如“租房合同没到期,房东要卖房,我该怎么办?”。
### 2.3 配置步骤
1. **创建知识库**
在主页点击左侧的知识库,进入后点击右上角的新建。知识库名称填写为 `民法典问答助手`,案例阶段其他配置保持默认即可。

2. **创建文本数据集**
点击新建,再选择文本数据集,用于导入本地文档。

3. **上传本地文件**
点击上传本地文件,选择示例文件。实际使用时可以同时上传多个文件,本案例只上传一个文件作为演示。

4. **设置解析参数**
进入参数设置页后,根据文件类型选择解析方式。

常用配置建议:
* **文件解析设置**:上传 PDF 时建议开启;普通 Word、Markdown、TXT 等文件可先使用默认配置。
* **处理方式**:大多数场景选择分块存储,成本较低,检索速度也更快。
* **分块条件**:控制每个分块包含多少 Token,快速测试时使用默认值即可。
* **索引增强**:纯文本通常勾选前两个选项;如果文档包含图片,可勾选图片相关增强。
这些参数会直接影响后续问答质量。分块过大时,检索结果可能包含太多无关内容,回答容易变得冗长;分块过小时,关键上下文可能被拆散,回答容易缺少依据。快速上手阶段可以先使用默认配置,正式接入企业制度、合同、产品手册时,再根据问答效果逐步调整。
5. **预览分块效果**
到数据预览步骤后,检查分块是否完整、可读。如果分块过长或过短,再返回调整参数。
预览分块时重点看三点:段落是否被异常截断,标题与正文是否保留在同一语义范围内,表格、条款或编号是否仍然可读。只要这里的知识片段质量不稳定,后续应用即使 Prompt 写得很好,也很难稳定给出准确答案。

6. **等待知识库就绪**
当数据状态变为“已就绪”后,知识库即可被应用引用。

7. **创建并关联对话 Agent**
创建一个新的对话 Agent,名称同样填写为 `民法典问答助手`。创建完成后,在应用配置中关联刚刚创建的知识库。

关联知识库后,应用的回答链路会从单纯对话变成“用户问题 → 知识库检索 → 模型总结回答”。这也是案例一和案例二的关键差异:案例一强调内容生成,案例二强调基于资料回答。
8. **配置问答 Prompt**
民法典问答助手需要强调“基于知识库回答”和“引用原文”。可以使用以下 Prompt:
```md
你是一个专业的民法典问答助手,基于《中华人民共和国民法典》原文回答法律问题。
规则:
- 严格基于知识库检索到的民法典条文回答,不得臆造法条
- 每条回答必须引用民法典原文(编、章、条)
- 如果民法典中没有直接对应的规定,如实说明,不做法律建议
- 涉及具体案件的法律适用时,提醒用户“本回答仅供参考,建议咨询专业律师”
- 对法律术语做通俗解释,让非法学背景的用户也能理解
输出格式:
1. **法律结论**(1-3 句概述)
2. **相关法条引用**(原文摘录 + 编章节条号)
3. **通俗解读**(用日常语言解释法条含义)
4. **实务建议**(2-3 条可操作建议)
5. **免责声明**(“本回答基于民法典原文,不构成法律意见,具体案件请咨询专业律师”)
```
法律问答类场景尤其要强调边界:能回答的是基于资料的通用解释,不能把模型回答包装成正式法律意见。Prompt 中要求引用条文、说明不确定性、添加免责声明,目的都是让输出更可追溯,也更符合高风险知识问答的使用规范。
9. **配置开场白**
开场白可以内置几个示例问题,帮助用户快速理解这个助手的使用方式:
```text
你好!我是民法典问答助手 ⚖️
我基于《中华人民共和国民法典》原文,帮你解答法律问题。
你可以问我:
["租房合同没到期,房东要卖房,我该怎么办?"]
["网购商品收到后发现坏了,商家不给退,法律怎么规定?"]
["楼上漏水把我家的天花板泡了,能索赔吗?"]
⚠️ 提示:我的回答仅供参考,具体法律问题建议咨询专业律师。
```
如果在开场白里添加 `[问题内容]`,用户可以点击问题一键提问,适合用于演示和引导。

10. **验证问答效果**
提出一个民法典相关问题,检查回答是否包含法律结论、法条引用、通俗解读和免责声明。

验证时不要只看回答是否“像法律回答”,还要看它是否真的引用了知识库内容。建议同时测试资料内问题、资料外问题和模糊问题:资料内问题应能引用原文,资料外问题应说明无法直接确认,模糊问题应主动提示需要补充事实。
### 2.4 业务价值
* **降低资料检索门槛**:把长篇法规、制度、手册转成自然语言问答入口,让业务人员不用先知道关键词或目录位置,也能快速找到相关内容。
* **提升回答可信度**:通过知识库检索和原文引用,让答案具备来源依据,减少模型凭经验编造或泛泛而谈的风险。
* **沉淀组织知识**:企业内部制度、产品 FAQ、售后 SOP、合同模板等资料可以持续进入知识库,形成可维护、可复用的知识资产。
* **适配高频支持场景**:法务、HR、行政、客服、交付团队都可以使用类似模式,把重复咨询转成自助问答,提高响应效率。
## 案例三:工作流—内容审核与自动改写
### 3.1 适用场景
工作流适合步骤固定、逻辑清晰、需要分支判断或人工确认的任务。本案例将“内容合规审核”拆成知识库检索、AI 分类、条件判断、自动改写、拒绝说明和人工确认几个环节,用来模拟企业内容发布前的审核流程。
它的核心价值有两个:
1. **自动化**:一次触发后,系统可以按预设流程自动执行多个步骤。
2. **标准化**:同样的输入会经过同样的判断和处理,减少人工差异。
判断一个任务是否适合工作流,可以看它是否具备“步骤稳定、规则明确、可重复执行”三个特征。内容审核就是典型场景:输入是一段待发布内容,规则来自合规知识库,输出通常是通过、改写或拒绝,中间还可以加入人工确认,既提升处理效率,也保留风险控制。
### 3.2 准备内容
先创建一个“内容合规规则库”知识库,并上传一个简单的文本文件。可以直接使用以下规则模板:
```text
内容合规规则
安全内容(可直接发布)
- 客观事实陈述
- 正常的活动通知、会议安排
- 产品功能介绍(基于真实数据)
- 行业知识分享
敏感措辞(需改写)
- 绝对化用语:“最好的”“第一”“100%”“绝对”“唯一”
- 夸张宣传:“颠覆行业”“史无前例”“无人能比”
- 未证实数据:无来源的转化率、满意度、增长率
- 对比贬低:直接点名竞品并贬低
违规内容(禁止发布)
- 违法信息:涉及黄赌毒、诈骗、传销
- 人身攻击:针对个人或群体的侮辱、诽谤
- 虚假信息:伪造数据、虚构资质、冒充官方
- 敏感话题:政治敏感、宗教歧视、地域攻击
```
规则库不需要一开始就非常复杂。快速验证时,先把规则拆成“安全、需改写、禁止发布”三类即可;正式使用时,可以继续按行业、品牌、渠道、地区增加规则,例如广告法敏感词、医疗合规要求、金融营销禁用表达、品牌语气规范等。
### 3.3 流程设计
正式配置节点之前,建议先确认完整流程。这个案例可以按以下方式设计:
1. **用户输入**:作为工作流起点,接收待审核内容。
2. **知识库检索**:从内容合规规则库中召回相关规则,建议引用上限设置为 1-2 条。
3. **AI 内容合规分类**:结合用户输入和检索结果,将内容分为“安全 / 敏感可改 / 违规”。
4. **分支判断**:根据分类结果进入不同分支。
5. **安全分支**:直接输出原文,表示可发布。
6. **敏感可改分支**:调用 AI 改写内容,再交由用户确认。
7. **违规分支**:输出拒绝说明,并给出修改建议。
8. **最终输出**:返回原文、改写稿或拒绝说明。
工作流设计时要先确定每个节点的职责,避免让一个 AI 节点同时承担“检索规则、判断分类、改写内容、解释原因”等过多任务。职责拆清楚后,每个节点的 Prompt 会更短、更稳定,后续也更容易定位问题:是知识库没有召回规则,是审核节点分类不准,还是判断器条件没有匹配上。
### 3.4 配置步骤
1. **创建工作流**
进入工作流主页,点击新建工作流,名称填写为 `内容审核与自动改写`。

创建完成后进入工作流编辑页。

本案例的完整流程图如下:

2. **配置开场白**
在系统配置中填写开场白,说明助手的审核规则和使用方式。
```md
你好!我是内容合规审核助手 🛡️
请把需要审核的内容发给我,我会按规则自动判断:
- **✅ 安全** — 内容无敏感信息,可以直接发布
- **⚠️ 敏感可改** — 包含可修正的措辞,我会改写后发回你确认
- **🚫 违规** — 包含红线内容,直接拒绝并说明原因
支持单个文本审核,也可以批量发(多条用换行分隔)。
**开始吧。**
```
3. **调用知识库**
在工作流中添加知识库检索节点,选择前面创建的内容合规规则库。该节点的输入使用用户输入,输出供后续 AI 节点引用。
这个节点的目的不是让模型阅读完整规则库,而是把与当前内容最相关的规则片段召回给后续审核节点。引用上限可以先设置为 1-2 条,保证提示词上下文足够聚焦;如果规则库较大或分类较细,再逐步提高引用数量。


4. **配置内容审核节点**
添加 AI 对话节点,并重命名为“内容审核节点”。该节点负责根据用户输入和知识库规则输出分类结果。
```md
你是一个内容合规审核专家。结合知识库检索到的合规规则,判断用户输入内容的合规等级。
分类标准:
- 安全:内容无敏感信息,可以直接发布
- 敏感可改:包含可修正的敏感措辞,改写后可发布
- 违规:包含红线内容,不可发布
输出格式:
仅输出三个词之一:安全 / 敏感可改 / 违规
```
注意:知识库引用处必须选择前面创建的知识库,否则 AI 无法读取规则内容。
审核节点的输出要尽量稳定,因为它会直接影响判断器分支。快速演示时可以只输出“安全 / 敏感可改 / 违规”三个词;如果后续要做更严格的自动化集成,可以改成结构化输出,例如同时返回分类、原因和命中的规则,方便下游节点做精确匹配和日志记录。

5. **测试审核节点**
点击右上角运行,输入一段待审核内容,确认审核节点能够返回稳定的分类结果。
这一步建议单独测试节点,而不是等全部流程搭完再调试。可以分别输入正常活动通知、包含绝对化宣传的营销文案、明显违规的文本,观察分类是否符合预期。如果分类不稳定,优先调整规则库表达和审核 Prompt,再继续配置后续分支。

6. **隐藏中间输出**
点击模型右侧的设置按钮,调整节点基础设置。由于用户只需要最终结果,建议隐藏内容审核节点的中间输出。

7. **配置判断器**
添加判断器节点,根据内容审核节点输出的“安全 / 敏感可改 / 违规”进入不同分支。
判断器相当于流程中的分流开关。上一步输出越稳定,判断器越容易配置;如果审核节点输出包含解释性文字,判断条件就可能匹配失败。因此在分支类工作流中,通常要先约束上游节点的输出格式,再配置判断器条件。

8. **配置安全分支**
当审核结果为“安全”时,直接输出原文即可。生产场景中也可以增加抽样复核或人工确认。

输出效果如下:

9. **配置敏感可改分支**
当审核结果为“敏感可改”时,添加 AI 对话节点自动改写内容。可以使用以下 Prompt:
```md
你是一个内容改写专家。将用户输入的内容改写为合规版本。
改写原则:
- 保留原意和信息量,不改变核心表达
- 将绝对化用语替换为客观表述,例如“最好的”改为“行业领先的”
- 将未证实的数据表述替换为合理推测,例如“100% 有效”改为“多数用户反馈有效”
- 将敏感措辞替换为中性表达
- 保持原文风格和语气
```
改写完成后,可以添加用户选择节点,让用户选择接受改写、继续改写或放弃。
人工确认节点适合放在有风险但可修正的路径上。AI 可以负责提出改写建议,但是否发布仍交给业务人员确认。这样既能减少人工初筛和反复改稿的时间,也不会把最终发布权完全交给自动化流程。


如果用户接受改写,输出改写结果;如果用户选择继续改写,可连接回内容改写节点;如果用户放弃,则输出原文。

10. **配置违规分支**
当审核结果为“违规”时,添加 AI 对话节点输出拒绝说明。可以使用以下 Prompt:
```md
用户提交的内容因涉及违规信息无法发布。请用礼貌、专业的语气说明原因。
输出格式:
1. 一句话说明内容无法发布
2. 列举具体的违规点(1-3 条)
3. 提供替代建议(如:建议修改哪些方面后重新提交)
```
11. **验证完整流程**
分别输入安全内容、敏感内容和违规内容,确认三条分支都能返回预期结果。

完整流程验证时,建议记录每条测试内容的输入、审核分类、进入分支和最终输出。若输出不符合预期,按节点顺序排查:先看知识库是否召回正确规则,再看审核节点分类是否准确,最后看判断器条件和分支输出是否配置正确。
### 3.5 业务价值
* **把审核规则流程化**:将内容发布前的判断标准沉淀为知识库和节点配置,减少依赖个人经验口头传递。
* **提高处理效率**:安全内容可快速通过,敏感内容自动改写,违规内容直接说明原因,让审核人员把精力集中在需要判断的内容上。
* **保留人工控制点**:对“敏感可改”这类灰度场景加入用户确认,避免自动化流程直接替业务人员做发布决策。
* **便于复用和扩展**:同样的流程可以迁移到营销文案、客服话术、公告通知、活动页面、短视频脚本等内容审核场景,只需要替换规则库和部分 Prompt。
* **降低合规风险**:通过固定分支和拒绝说明,让高风险内容有明确拦截路径,减少误发、夸大宣传或不合规表达带来的风险。
## 案例四:Agent V2—智能数据分析 Agent
### 4.1 适用场景
Agent V2 适合开放式、多步骤、需要动态规划的任务。与工作流不同,工作流需要提前定义每一步;Agent V2 更适合“步骤不固定”的任务,例如数据分析、文件处理、多工具协作和需要追问澄清的复杂问题。
本案例模拟运营、产品或销售团队的日常数据分析需求:上传 Excel 文件后,直接用自然语言提出问题,让 Agent 自主读取文件、制定分析计划,并在虚拟机中执行分析。
案例三的工作流需要你先设计“检索规则 → 分类 → 分支 → 改写或拒绝”的固定路径;Agent V2 则更像一个可以自主规划的执行者。你不需要提前定义每一步,只需要提供目标和文件,Agent 会根据数据结构和问题复杂度决定先读取文件、再做统计、是否需要追问、是否需要运行代码。
数据分析很适合用来体验 Agent V2,因为它天然具备三个特点:分析路径不固定、经常需要多步推理、需求可能需要澄清。同一份 Excel,不同用户可能关心产品销量、渠道 ROI、区域趋势或异常订单,固定工作流很难提前覆盖所有路径,而 Agent V2 可以根据问题动态规划。
### 4.2 准备内容
* 示例文件:准备一份销售数据 Excel 或 CSV 表格。
* 一个支持 Agent V2 的对话模型。
* 虚拟机能力已在当前环境中可用。
### 4.3 配置步骤
1. **创建 Agent V2 应用**
工作台点击新建应用,选择 `对话 Agent V2`,名称填写为 `智能数据分析 Agent`。

2. **配置模型与 Prompt**
AI 配置继续使用 GLM-5.1 模型。系统提示词可以参考:
```md
你是一个资深数据分析师 Agent。你可以读取用户上传的数据文件、在沙箱中运行 Python 分析代码、并主动向用户追问澄清需求。
工具:
- 📄 **读取文件** — 读取用户上传的 Excel/CSV 数据
- 💻 **沙箱执行** — 运行 Python 脚本(pandas/matplotlib/numpy)
- ❓ **主动追问** — 分析需求不明确时向用户确认
工作方式:
1. 收到数据和问题后,先读取文件了解数据结构和内容
2. 制定分析计划,以步骤列表展示给用户
3. 按计划逐步执行,每步展示关键发现
4. 遇到模糊需求时主动追问,如指标定义不明确或缺少对比基准
5. 根据追问结果动态更新计划
6. 输出分析报告,包含数据概览、核心发现、可视化图表、业务建议
安全规则:
- 沙箱中只能做数据分析,禁止访问网络,禁止写文件到宿主机
- 数据仅用于本次分析,不在报告中暴露原始敏感数据
- 不确定的结论标注置信度
输出格式:
1. 📋 分析计划(根据数据自动生成)
2. 📊 数据概览(行数、列名、缺失值、基本统计)
3. 🔍 核心发现(3-5 个关键洞察,图表辅助)
4. 💡 业务建议(基于数据的可操作建议)
```
这段 Prompt 的重点是让 Agent 先规划再执行,而不是直接给结论。对于数据分析类任务,先读取数据结构、确认字段含义、制定分析计划,可以显著减少误解需求或误用指标的风险。正式使用时,可以继续补充企业内部指标定义,例如 GMV、ROI、转化率、有效客户、复购率等口径。
3. **配置开场白**
开场白用于引导用户上传数据文件并提出分析问题:
```text
你好!我是智能数据分析 Agent 📊
直接把 Excel 或 CSV 文件拖进来,告诉我你想分析什么。
例如:
- "分析这份销售数据,找出销量最好的产品和趋势"
- "帮我看看用户活跃度的变化,找出下降的原因"
- "对比三个渠道的转化率,哪个 ROI 最高?"
我会先了解你的数据,制定分析计划,然后在沙箱中运行分析代码。
需要确认的地方我会主动问你。
```
4. **开启虚拟机能力**
数据分析通常需要读取文件并执行代码,因此需要开启虚拟机配置。
虚拟机能力用于隔离代码执行环境。Agent 可以在其中运行数据分析脚本、读取上传文件、生成统计结果,但不会直接影响本地宿主环境。对于需要运行 Python、处理 Excel、绘制图表或做批量计算的场景,这是 Agent V2 区别于普通对话应用的重要能力。

5. **上传文件并测试**
上传销售数据示例文件,并输入问题:
```text
帮我分析这份销售数据,看看哪些产品卖得好,哪个渠道 ROI 最高
```
Agent 会先拆解任务,再逐步执行分析。

执行过程中可以看到任务在虚拟机中运行,不会影响本地环境。

测试时重点观察 Agent 是否具备完整的分析过程:是否先识别表格字段,是否解释分析计划,是否在需要时运行代码,是否根据结果给出业务建议。如果问题描述不清晰,理想行为不是强行分析,而是先向用户追问关键口径。
### 4.4 验证效果
最终结果应包含分析计划、数据概览、关键发现和业务建议。

结果验证时建议关注四个维度:结论是否来自实际数据,指标口径是否清楚,图表或统计是否能支撑结论,业务建议是否可执行。数据分析 Agent 的价值不只是输出一段总结,而是把“读数据、算指标、解释结果、提出建议”的过程串起来。
### 4.5 业务价值
* **降低数据分析门槛**:业务人员可以直接上传 Excel 或 CSV,用自然语言提出问题,不必先写 SQL、Python 或复杂公式。
* **支持开放式探索**:同一份数据可以围绕销量、渠道、区域、客户、趋势、异常值等方向反复追问,适合没有固定分析路径的场景。
* **提升分析透明度**:Agent 会展示分析计划和关键步骤,用户可以看到它如何理解数据、如何计算指标、如何得出结论。
* **隔离代码执行风险**:通过虚拟机执行分析脚本,降低本地环境污染、依赖冲突和权限误用风险。
* **沉淀业务分析能力**:销售复盘、运营周报、活动归因、产品指标诊断等场景都可以复用这一类 Agent,把数据分析从专家任务变成日常工作流。
## 四种类型选型回顾
完成四个案例后,可以把 FastGPT 的常见应用类型理解为从简单到复杂的能力递进:对话 Agent 解决“怎么回答和生成内容”,知识库解决“基于什么资料回答”,工作流解决“按什么固定流程处理”,Agent V2 解决“面对开放任务如何自主规划和执行”。
| 应用类型 | 适合场景 | 核心能力 |
| -------------- | ------------------- | ---------------- |
| 对话 Agent | 轻量问答、文案生成、标准化输出 | Prompt、模型配置、工具调用 |
| 知识库 + 对话 Agent | 基于资料、制度、法条、产品手册的问答 | 文件导入、知识库检索、引用原文 |
| 工作流 | 固定步骤、条件分支、审核流、自动化处理 | 节点编排、判断器、人工确认 |
| Agent V2 | 数据分析、复杂任务、多步推理、动态规划 | 自主规划、工具调用、虚拟机执行 |
选型时可以按任务复杂度判断:
1. 如果只是做轻量对话或标准化文案生成,优先选择对话 Agent。
2. 如果回答必须基于已有资料,选择知识库 + 对话 Agent。
3. 如果流程固定且需要条件分支、人工确认或自动化处理,选择工作流。
4. 如果任务开放、步骤不固定,并且需要自主分析、工具调用或代码执行,选择 Agent V2。
实际项目中也可以组合使用这些能力。例如客服助手可以使用知识库回答产品问题,再通过工具查询订单;内容审核可以用工作流固定审核路径,同时用知识库维护规则;数据分析场景可以先用 Agent V2 完成探索,再把稳定下来的分析步骤沉淀成工作流。
file: ./content/guide/workspace/customDomain.en.mdx
meta: {
"title": "Configure Custom Domain",
"description": "How to configure a custom domain in FastGPT"
}
FastGPT Cloud supports custom domain configuration starting from v4.14.4.
## How to Configure a Custom Domain
### 1. Open the "Custom Domain" Page
In the sidebar, go to "Account" -> "Custom Domain" to open the configuration page.
If your plan does not support this feature, follow the on-screen instructions to upgrade.

### 2. Add a Custom Domain
1. Have your domain ready. Your domain must have an ICP filing. Currently supported filing providers are Alibaba Cloud, Tencent Cloud, and Volcano Engine.
2. Click the "Edit" button to enter edit mode.
3. Enter your domain, e.g. [www.example.com](http://www.example.com)
4. In your domain provider's DNS management console, add the CNAME record shown on the screen.
5. After adding the DNS record, click "Save". The system will automatically verify the DNS configuration -- this usually takes less than a minute. If verification takes too long, try again.
6. Once the status shows "Active", click "Confirm" to finish.

You can now access FastGPT services and call FastGPT APIs using your own domain.
## DNS Resolution Failure
The system checks DNS resolution daily. If the DNS record becomes invalid, the custom domain will be disabled. You can re-verify it by clicking "Edit" on the "Custom Domain" management page.

To change your custom domain or switch providers, delete the existing configuration and set it up again.
## Use Cases
* [Integrate with WeCom Bot](../build/publish/wecom.en.mdx)
file: ./content/guide/workspace/customDomain.mdx
meta: {
"title": "配置自定义域名",
"description": "如何在 FastGPT 中配置自定义域名"
}
FastGPT 云服务版自 v4.14.4 后支持配置自定义域名。
## 如何配置自定义域名
### 1. 打开“自定义域名”页面
在侧边栏选择“账号” -> “自定义域名”,打开自定义域名配置页。
如果您的套餐等级不支持配置,请根据页面的指引升级套餐。

### 2. 添加自定义域名
1. 准备好您的域名。您的域名必须先经过备案,目前支持“阿里云”、“腾讯云”、“火山引擎”三家服务商的备案域名。
2. 点击“编辑”按钮,进入编辑状态。
3. 填入您的域名,例如 [www.example.com](http://www.example.com)
4. 在域名服务商的域名解析处,添加界面中提示的 DNS 记录,注意记录类型为 CNAME。
5. 添加解析记录后,点击"保存"按钮。系统将自动检查 DNS 解析情况,一般情况下,在一分钟内就可以获取到解析记录。如果长时间没有获取到记录,可以重试一次。
6. 待状态提示显示为“已生效”后,点击“确认”按钮即可。

现在您可以通过您自己的域名访问 fastgpt 服务、调用 fastgpt 的 API 了。
## 域名解析失效
系统会每天对 DNS 解析进行检查,如果发现 DNS 解析记录失效,则会停用该自定义域名,可以在"自定义域名"管理界面中点击"编辑"进行重新解析。

如果您需要修改自定义域名、或修改服务商,则需要删除自定义域名配置后进行重新配置。
## 使用案例
* [接入企业微信智能机器人](../build/publish/wecom.mdx)
file: ./content/self-host/config/env.en.mdx
meta: {
"title": "Environment Variables",
"description": "Environment variables for projects/app, projects/code-sandbox, and pro/admin"
}
This page describes the environment variables commonly used in a self-hosted FastGPT deployment. `projects/app` and `pro/admin` both reuse many settings from `packages/service/env.ts`, so database, secret, object storage, vector database, and service-level variables are documented together. Variables that are only read by `projects/app` or `pro/admin` are listed separately.
## Notes
* `projects/app`: the main Next.js application, including pages, API routes, workflows, Knowledge Bases, object storage, and vector storage.
* `pro/admin`: the commercial Admin service. Besides its own Admin variables, it also reuses App/Service settings such as database, secrets, object storage, models, and logging.
* `projects/code-sandbox`: the code execution sandbox service. It exposes the `/sandbox` endpoint and is called by App through `CODE_SANDBOX_URL`.
* `packages/service/env.ts` exports `serviceEnv`; `projects/app/src/env.ts` exports `appEnv`.
* Shared App/Admin boolean variables use `true`, `1`, `yes`, or `y` to enable a feature. Other values are treated as disabled.
* `FILE_TOKEN_KEY`, `AES256_SECRET_KEY`, and `INVOKE_TOKEN_SECRET` are required at runtime. Use strong random secrets and do not use the example values in production.
## Shared App/Admin Variables
These variables are mainly validated by `packages/service/env.ts` and apply to `projects/app` and to `pro/admin` when it imports `@fastgpt/service`. A few App-side switches are still defined in `packages/service/env.ts`; they are also called out in the App-specific section below.
### Basics and Secrets
| Variable | Default | Description |
| --------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `DB_MAX_LINK` | `5` | Maximum connection pool size for MongoDB, PG, OceanBase, openGauss, and other databases. |
| `SYNC_INDEX` | `true` | Whether to create missing MongoDB indexes and remove explicitly declared deprecated indexes at startup. Maintain indexes manually when disabled. |
| `FILE_TOKEN_KEY` | None, **required** | Secret for file read and file authorization flows. Must be at least 6 characters. |
| `AES256_SECRET_KEY` | None, **required** | Secret used by AES encryption and decryption. Must be at least 6 characters. |
| `INVOKE_TOKEN_SECRET` | None, **required** | JWT secret for Invoke reverse calls. Must be at least 32 characters. |
| `ROOT_KEY` | `fastgpt_root_key` | Admin API key for the current system. It can call `/api/admin/**` APIs and must be at least 6 characters. |
| `PRO_TOKEN` | Empty | Token for FastGPT app server calls to pro/admin internal APIs. It must match the pro/admin configuration and is required when App configures `PRO_URL`. |
| `PRO_URL` | Empty | Commercial service URL. When set, App can call Pro APIs, and the domain is allowed by file URL validation. |
### Service URLs and Integrations
| Variable | Default | Description |
| ------------------------ | ----------------------------------- | --------------------------------------------------------------------------------------------------------- |
| `PLUGIN_BASE_URL` | `http://localhost:3004` | FastGPT Plugin service URL. Deployment templates usually set this to the internal Plugin service URL. |
| `PLUGIN_TOKEN` | `token` | Authentication token for calling the Plugin service. It must match the Plugin service configuration. |
| `CODE_SANDBOX_URL` | `http://localhost:3002` | Code Sandbox service URL. Deployment templates usually set this to the internal Code Sandbox service URL. |
| `CODE_SANDBOX_TOKEN` | `codesandbox` | Token used by App when calling Code Sandbox. It must match the sandbox service `SANDBOX_TOKEN`. |
| `AIPROXY_API_ENDPOINT` | Empty | AI Proxy service URL. When configured, model requests prefer AI Proxy. |
| `AIPROXY_API_TOKEN` | Empty | Token for calling AI Proxy. |
| `OPENAI_BASE_URL` | `https://api.openai.com/v1` | Default OpenAI-compatible model endpoint when AI Proxy is not configured. |
| `CHAT_API_KEY` | Empty | Default OpenAI-compatible model API key when AI Proxy token is not configured. |
| `CRM_API_URL` | Empty | Lead attribution CRM API base URL (including `/api/v1`). Empty disables identity reporting. |
| `CRM_API_KEY` | Empty | CRM admin API key used to bind a FastGPT user to `visitor_id` after registration or login. |
| `MARKETPLACE_URL` | `https://v2.marketplace.fastgpt.cn` | Plugin marketplace API URL. |
| `FEISHU_BASE_URL` | `https://open.feishu.cn` | Lark Open Platform URL. Use your private Lark domain when self-hosting Lark. |
| `DINGTALK_BASE_URL` | `https://api.dingtalk.com` | DingTalk new API base URL. |
| `DINGTALK_OAPI_BASE_URL` | `https://oapi.dingtalk.com` | DingTalk OAPI base URL. |
| `YUQUE_DATASET_BASE_URL` | `https://www.yuque.com` | Yuque Knowledge Base URL. |
### Agent Sandbox
| Variable | Default | Description |
| ------------------------------------------------ | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `AGENT_SANDBOX_PROVIDER` | Empty | Agent Sandbox provider. Supported values are `sealosdevbox` and `opensandbox`. Empty disables Sandbox. Once set, the matching provider variables are required. `fastgpt-app` also requires all three proxy variables, while `fastgpt-pro` requires only the preview proxy URL. |
| `AGENT_SANDBOX_SEALOS_BASEURL` | Empty | Sealos Devbox service URL. |
| `AGENT_SANDBOX_SEALOS_TOKEN` | Empty | Sealos Devbox access token. |
| `AGENT_SANDBOX_SEALOS_WORK_DIRECTORY` | `/home/devbox/workspace` | Working directory inside the Sealos Devbox sandbox. |
| `AGENT_SANDBOX_SEALOS_IMAGE` | Empty | Runtime image used by Sealos Devbox. Required when `sealosdevbox` is enabled. |
| `AGENT_SANDBOX_OPENSANDBOX_BASEURL` | Empty | OpenSandbox service URL. |
| `AGENT_SANDBOX_OPENSANDBOX_API_KEY` | Empty | OpenSandbox API key. Required when OpenSandbox is enabled, and must match OpenSandbox server `[server].api_key`. |
| `AGENT_SANDBOX_OPENSANDBOX_RUNTIME` | `docker` | OpenSandbox runtime, either `docker` or `kubernetes`. |
| `AGENT_SANDBOX_OPENSANDBOX_IMAGE` | Empty | Full runtime image used by OpenSandbox. Required when `opensandbox` is enabled. |
| `AGENT_SANDBOX_OPENSANDBOX_USE_SERVER_PROXY` | `true` | Whether OpenSandbox access goes through the server proxy. |
| `AGENT_SANDBOX_OPENSANDBOX_VOLUME_MANAGER_URL` | Empty | Required in OpenSandbox mode. Volume Manager service URL. |
| `AGENT_SANDBOX_OPENSANDBOX_VOLUME_MANAGER_TOKEN` | Empty | Required in OpenSandbox mode. Volume Manager authentication token. |
| `AGENT_SANDBOX_OPENSANDBOX_VOLUME_NAME_PREFIX` | `fastgpt-session` | Prefix used by the FastGPT app when generating persistent OpenSandbox volume `claimName` values. When upgrading, reuse the previous `VM_VOLUME_NAME_PREFIX` value. |
| `AGENT_SANDBOX_PROXY_SECRET` | Empty | Shared HMAC secret for the app and agent-sandbox-proxy. Required by `fastgpt-app` when Agent Sandbox is enabled; must be at least 32 bytes. |
| `AGENT_SANDBOX_PROXY_URL` | Empty | Browser-accessible WebSocket URL for agent-sandbox-proxy. Required by `fastgpt-app` when Agent Sandbox is enabled; must start with `ws://` or `wss://`. |
| `AGENT_SANDBOX_PREVIEW_PROXY_URL` | Empty | Browser-accessible HTTP(S) URL for Sandbox file previews. You must add it to both `fastgpt-app` and `fastgpt-pro` when Agent Sandbox is enabled. Use an origin separate from the FastGPT application. |
| `AGENT_SANDBOX_FREE_TIP` | `false` | Whether the frontend shows the Agent Sandbox free-use hint. |
| `AGENT_SANDBOX_STORAGE_SIZE_GI` | `1` | Agent Sandbox storage size in Gi. FastGPT derives the archive, Skill, and single-file limits from this value. |
| `AGENT_SANDBOX_SUSPEND_MINUTES` | `60` | Number of inactive minutes before a running Agent Sandbox is automatically suspended. |
| `AGENT_SANDBOX_ARCHIVE_INACTIVE_DAYS` | `7` | Number of inactive days before a suspended Agent Sandbox is automatically archived. |
| `AGENT_SANDBOX_MAX_EDIT_DEBUG` | `100` | Limit for Agent edit/debug sandboxes. |
| `AGENT_SANDBOX_NPM_REGISTRY` | Empty | npm registry used by npm, yarn, pnpm, and bun inside Agent sandboxes. |
| `AGENT_SANDBOX_PYPI_INDEX_URL` | Empty | PyPI index URL used by pip, `python -m pip`, and uv inside Agent sandboxes. |
### Databases, Cache, and Vector Stores
| Variable | Default | Description |
| ---------------------------------------------- | ------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| `REDIS_URL` | `redis://default:mypassword@localhost:6379` | Redis connection URL. |
| `STREAM_RESUME_TTL_SECONDS` | `300` | TTL for an active stream resume mirror, in seconds. |
| `STREAM_RESUME_POST_COMPLETE_TTL_SECONDS` | `30` | Shortened TTL after a stream completes, in seconds. |
| `STREAM_RESUME_REDIS_MAXMEMORY_RATIO` | `0.5` | When Redis used memory divided by `maxmemory` reaches this ratio, new stream resume mirrors are skipped. |
| `STREAM_RESUME_REDIS_MEMORY_CHECK_INTERVAL_MS` | `5000` | Redis memory watermark cache duration, in milliseconds. |
| `MONGODB_URI` | Local MongoDB example URL | Main business MongoDB connection URL. |
| `MONGODB_LOG_URI` | Same example as `MONGODB_URI` | MongoDB connection URL for logs. If unset, it can reuse the main database. |
| `VECTOR_VQ_LEVEL` | `32` | Vector quantization level. Supported ranges depend on the vector store. |
| `PG_URL` | Empty | PostgreSQL/pgvector connection URL. |
| `OCEANBASE_URL` | Empty | OceanBase vector store connection URL. |
| `SEEKDB_URL` | Empty | SeekDB vector store connection URL. |
| `MILVUS_ADDRESS` | Empty | Milvus/Zilliz address. |
| `MILVUS_TOKEN` | Empty | Milvus/Zilliz access token. |
| `OPENGAUSS_URL` | Empty | openGauss vector store connection URL. |
### Object Storage
| Variable | Default | Description |
| --------------------------------------- | ----------------------- | ----------------------------------------------------------------------------------------------------- |
| `STORAGE_VENDOR` | `minio` | Object storage vendor. Supported values are `minio`, `aws-s3`, `r2`, `cos`, and `oss`. |
| `STORAGE_PUBLIC_BUCKET` | `fastgpt-public` | Public file bucket. |
| `STORAGE_PRIVATE_BUCKET` | `fastgpt-private` | Private file bucket. |
| `STORAGE_REGION` | `us-east-1` | Object storage region. |
| `STORAGE_EXTERNAL_ENDPOINT` | Empty | Externally reachable object storage endpoint for browsers or external services. |
| `STORAGE_R2_PUBLIC_ENDPOINT` | Empty | Public HTTPS domain for a Cloudflare R2 bucket; required when `STORAGE_VENDOR=r2`. |
| `STORAGE_S3_CDN_ENDPOINT` | Empty | CDN endpoint used for temporary `short-redirect` download URLs. Requires `STORAGE_EXTERNAL_ENDPOINT`. |
| `STORAGE_DOWNLOAD_URL_MODE` | `short-proxy` | Download mode: `short-proxy` or `short-redirect`. External URLs are always FastGPT short links. |
| `STORAGE_DOWNLOAD_REDIRECT_TTL_SECONDS` | `300` | Lifetime of the temporary object storage or CDN URL used by `short-redirect`, in seconds. |
| `STORAGE_S3_ENDPOINT` | `http://localhost:9000` | S3/MinIO-compatible API endpoint. |
| `STORAGE_PUBLIC_ACCESS_EXTRA_SUB_PATH` | Empty | Extra sub-path for public file access URLs. |
| `STORAGE_ACCESS_KEY_ID` | `minioadmin` | Object storage access key. |
| `STORAGE_SECRET_ACCESS_KEY` | `minioadmin` | Object storage secret key. |
| `STORAGE_S3_FORCE_PATH_STYLE` | `false` | Whether S3 path-style access is forced. MinIO usually requires this. |
| `STORAGE_S3_MAX_RETRIES` | `3` | Maximum S3 client retry count. |
| `STORAGE_COS_PROTOCOL` | `https:` | Tencent Cloud COS protocol, either `https:` or `http:`. |
| `STORAGE_COS_USE_ACCELERATE` | `false` | Whether Tencent Cloud COS acceleration domain is used. |
| `STORAGE_COS_CNAME_DOMAIN` | Empty | Tencent Cloud COS custom CNAME domain. |
| `STORAGE_COS_PROXY` | Empty | Tencent Cloud COS proxy URL. |
| `STORAGE_OSS_ENDPOINT` | Empty | Alibaba Cloud OSS endpoint. |
| `STORAGE_OSS_CNAME` | `false` | Whether Alibaba Cloud OSS uses CNAME. |
| `STORAGE_OSS_INTERNAL` | `false` | Whether Alibaba Cloud OSS uses an internal endpoint. |
| `STORAGE_OSS_SECURE` | `false` | Whether Alibaba Cloud OSS uses HTTPS. |
| `STORAGE_OSS_ENABLE_PROXY` | `true` | Whether Alibaba Cloud OSS proxy access is enabled. |
### Logging, Metrics, and Tracing
| Variable | Default | Description |
| --------------------------- | ---------------- | ----------------------------------------------------------------------------------------------------- |
| `LOG_ENABLE_CONSOLE` | `true` | Whether console logging is enabled. |
| `LOG_CONSOLE_LEVEL` | `debug` | Console log level. Supported values are `trace`, `debug`, `info`, `warning`, `error`, and `fatal`. |
| `LOG_DEPTH` | `3` | Legacy template variable for log object depth. New structured logging mainly uses log-level settings. |
| `LOG_ENABLE_OTEL` | `false` | Whether OpenTelemetry log export is enabled. |
| `LOG_OTEL_LEVEL` | `info` | OTEL log level. |
| `LOG_OTEL_SERVICE_NAME` | `fastgpt-client` | OTEL log service name. |
| `LOG_OTEL_URL` | Empty | OTEL log export URL. |
| `METRICS_ENABLE_OTEL` | `false` | Whether OpenTelemetry metrics export is enabled. |
| `METRICS_EXPORT_INTERVAL` | `30000` | Metrics export interval, in milliseconds. |
| `METRICS_OTEL_SERVICE_NAME` | `fastgpt-client` | OTEL metrics service name. |
| `METRICS_OTEL_URL` | Empty | OTEL metrics export URL. |
| `TRACING_ENABLE_OTEL` | `false` | Whether OpenTelemetry tracing is enabled. |
| `TRACING_OTEL_SERVICE_NAME` | `fastgpt-client` | OTEL tracing service name. |
| `TRACING_OTEL_URL` | Empty | OTEL tracing export URL. |
| `TRACING_OTEL_SAMPLE_RATIO` | Empty | Trace sampling ratio from `0` to `1`. |
| `CHAT_LOG_URL` | Empty | Chat log push service URL. Empty disables pushing. |
| `CHAT_LOG_INTERVAL` | Empty | Chat log batch push interval, in milliseconds. |
| `CHAT_LOG_SOURCE_ID_PREFIX` | `fastgpt-` | Prefix for chat log source IDs. |
| `TRACK_BATCH_UPDATE_TIME` | `10000` | Event counter batch write interval, in milliseconds. |
### Domains, Frontend, and Runtime
| Variable | Default | Description |
| ------------------------- | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `FE_DOMAIN` | Required | The origin clients use to access FastGPT, including the scheme, host, and optional port. It completes file and image URLs. Local development can use `http://localhost:3000`. |
| `FILE_DOMAIN` | Empty | File access domain. It usually points to FastGPT, but a separate domain can isolate file risk. |
| `NEXT_PUBLIC_BASE_URL` | Empty | Next.js sub-path deployment prefix, such as `/fastgpt`. It must be fixed when building the image. |
| `HOSTNAME` | `localhost` | Service host used for internal URLs and SSRF local-address detection. Containers commonly set it to `0.0.0.0`. |
| `PORT` | `3000` | Next.js listening port. Also used for local-address detection. |
| `NODE_ENV` | Empty | Standard Node/Next.js runtime environment. Production images set it to `production`. |
| `NEXT_TELEMETRY_DISABLED` | `1` | Disables Next.js Telemetry in production images. |
| `NODE_OPTIONS` | `--max-old-space-size=4096` | Node options used during production image builds to increase the build memory limit. |
### Security
| Variable | Default | Description |
| ----------------------------------- | ------- | --------------------------------------------------------------------------------------------------- |
| `USE_IP_LIMIT` | `false` | Whether IP rate limiting is enabled for selected APIs. |
| `CHECK_INTERNAL_IP` | `false` | Whether internal IP checks are enabled to reduce SSRF risk. |
| `AUTH_COOKIE_SECURE` | `false` | Whether login cookies use the `Secure` attribute. Enable only when the site is HTTPS-only. |
| `TRUSTED_PROXY_ENABLE` | `false` | Whether trusted reverse proxy client IP validation is enabled. Disabled keeps legacy behavior. |
| `TRUSTED_PROXY_IPS` | Empty | Trusted reverse proxy IP/CIDR list, separated by commas or whitespace. |
| `PASSWORD_LOGIN_MINUTE_LIMIT_COUNT` | `10` | Maximum password login requests per account per minute. |
| `MAX_LOGIN_SESSION` | `10` | Maximum login clients per account. |
| `ALLOWED_ORIGINS` | Empty | Allowed CORS origins. Use commas to separate multiple origins. Empty allows all origins by default. |
| `MULTIPLE_DATA_TO_BASE64` | `false` | Whether images are forced into base64 before being sent to models. |
| `DISABLE_CACHE` | `false` | Whether system cache hits are disabled, mainly for debugging. |
| `HTTP_PROXY` | Empty | Outbound HTTP proxy for Node and workers. |
| `HTTPS_PROXY` | Empty | Outbound HTTPS proxy for Node and workers. |
| `NO_PROXY` | Empty | Address list that bypasses proxies. |
| `ALL_PROXY` | Empty | General outbound proxy. |
### Feature Flags and Limits
| Variable | Default | Description |
| -------------------------------------- | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `AGENT_ENGINE` | `fastAgent` | Agent engine. Supported values are `fastAgent` and `piAgent`. |
| `SKIP_FILE_TYPE_CHECK` | `false` | Whether upload file type checks are skipped. |
| `WECHAT_CHANNEL_CONCURRENCY` | `1000` | WeChat channel poll worker concurrency. Minimum value is `10`. |
| `PARSE_FILE_WORKERS` | `5` | Resident file parsing worker count. |
| `HTML_TO_MARKDOWN_WORKERS` | `10` | Resident HTML-to-Markdown worker count. |
| `TEXT_TO_CHUNKS_WORKERS` | `10` | Resident text chunking worker count. |
| `PARSE_FILE_TIMEOUT_SECONDS` | `600` | Timeout for one file parsing task, in seconds. |
| `WORKFLOW_MAX_RUN_TIMES` | `500` | Maximum workflow run count to avoid extreme infinite loops. |
| `WORKFLOW_MAX_LOOP_TIMES` | `100` | Maximum input array length for loop and parallel nodes. |
| `WORKFLOW_PARALLEL_MAX_CONCURRENCY` | `10` | Parallel node concurrency limit. It must not exceed `WORKFLOW_MAX_LOOP_TIMES`. |
| `SYSTEM_MAX_STRING_LENGTH_M` | `100` | Maximum character length for synchronous system string operations such as variable replacement, in M characters. `1` means `1,000,000` characters. Valid range: `1` to `100`. |
| `CHAT_MAX_QPM` | `5000` | Chat QPM limit. User plan limits take precedence when configured. |
| `SERVICE_REQUEST_MAX_CONTENT_LENGTH` | `10` | Maximum request body size accepted by the service, in MB. |
| `MAX_FOLDER_DEPTH` | `4` | Maximum folder depth. The default allows up to 4 folder levels under the root. Valid range: `2` to `20`. |
| `APP_FOLDER_MAX_AMOUNT` | `1000` | Maximum number of App folders. |
| `DATASET_FOLDER_MAX_AMOUNT` | `1000` | Maximum number of dataset folders. |
| `UPLOAD_FILE_MAX_SIZE` | `1000` | Maximum upload file size, in MB. |
| `UPLOAD_FILE_MAX_AMOUNT` | `1000` | Maximum upload file count. |
| `LLM_REQUEST_TRACKING_RETENTION_HOURS` | `6` | LLM request tracking retention, in hours. |
| `MAX_HTML_TRANSFORM_CHARS` | `1000000` | Maximum number of characters for HTML-to-Markdown conversion. Larger content is not converted. |
## Additional App Variables
These variables are mainly read by `projects/app`. Some are currently defined in `packages/service/env.ts` for shared validation, but their actual consumers are still App-side code.
| Variable | Default | Description |
| ------------------------------- | -------- | --------------------------------------------------------------------------------------------------- |
| `DEFAULT_ROOT_PSW` | `123456` | Default password for initializing the root user. |
| `SYSTEM_NAME` | `AI` | Default system name for the page title. |
| `SYSTEM_DESCRIPTION` | Empty | Page meta description. If unset, the default i18n text is used. |
| `SYSTEM_FAVICON` | Empty | Page favicon URL. If unset, the favicon from system config is used. |
| `CHINESE_IP_REDIRECT_URL` | Empty | China IP redirect URL in frontend config. |
| `PAY_FORM_URL` | Empty | Payment form URL in frontend config. |
| `SHOW_COUPON` | `false` | Whether redemption codes are shown. |
| `SHOW_DISCOUNT_COUPON` | `false` | Whether discount coupons are shown. |
| `HIDE_CHAT_COPYRIGHT_SETTING` | `false` | Whether copyright settings are hidden. |
| `WECOM_LOGIN_AUTO_REDIRECT` | `false` | Whether WeCom terminals automatically redirect to login. |
| `APP_REGISTRATION_URL` | Empty | App registration application URL. Currently kept mostly for compatibility. |
| `PASSWORD_EXPIRED_MONTH` | Empty | Password expiration period in months. Empty means passwords do not expire. |
| `OPENAPI_KEY_MAX_COUNT` | `100` | Maximum number of system API Keys one team member can create. Minimum is 1. |
| `SSE_MCP_SERVER_PROXY_ENDPOINT` | Empty | MCP SSE server proxy URL. Do not include a trailing slash. Required when publishing an SSE MCP App. |
### Open-Source Only
Starting with v4.15.0, the open-source edition no longer reads `config.json`. When upgrading from an earlier version, remove the file's volume mount and migrate the old settings to environment variables using the table below. If you did not use these optional settings, you do not need to add them.
| Former `config.json` field | Current environment variable | Default | Description |
| ------------------------------------------- | ---------------------------- | -------- | ------------------------------------------------------------------------------ |
| `systemEnv.customPdfParse.url` | `CUSTOM_PDF_PARSE_URL` | Empty | Custom PDF parsing service URL. |
| `systemEnv.customPdfParse.key` | `CUSTOM_PDF_PARSE_KEY` | Empty | Custom PDF parsing service key. |
| `systemEnv.customPdfParse.doc2xKey` | `DOC2X_KEY` | Empty | Doc2x PDF parsing service key. |
| `systemEnv.customPdfParse.textinAppId` | `TEXTIN_APP_ID` | Empty | TextIn service App ID. |
| `systemEnv.customPdfParse.textinSecretCode` | `TEXTIN_SECRET_CODE` | Empty | TextIn service Secret Code. |
| `systemEnv.hnswEfSearch` | `HNSW_EF_SEARCH` | `100` | The `hnsw.ef_search` vector search parameter for PG, OceanBase, and openGauss. |
| `systemEnv.hnswMaxScanTuples` | `HNSW_MAX_SCAN_TUPLES` | `100000` | Maximum number of tuples scanned during vector search. Applies only to PG. |
| `systemEnv.datasetParseMaxProcess` | `DATASET_PARSE_MAX_PROCESS` | `10` | Maximum concurrency for the Knowledge Base file parsing queue. |
| `systemEnv.vectorMaxProcess` | `VECTOR_MAX_PROCESS` | `10` | Maximum concurrency for the vector training queue. |
| `systemEnv.qaMaxProcess` | `QA_MAX_PROCESS` | `10` | Maximum concurrency for the Q\&A splitting queue. |
| `systemEnv.vlmMaxProcess` | `VLM_MAX_PROCESS` | `10` | Maximum concurrency for the image understanding model queue. |
#### Enhanced PDF Parsing
The open-source edition supports custom PDF parsing services, SoMark, TextIn, and Doc2x. Configure only one service. If you configure more than one, FastGPT uses this priority order: custom PDF parsing service, SoMark, TextIn, then Doc2x.
##### Use the Sealos PDF Parsing Service
1. Open [Sealos AI Proxy](https://hzh.sealos.run/?uid=fnWRt09fZP\&openapp=system-aiproxy) and create an API key.
2. Add the API key to FastGPT:
```dotenv
CUSTOM_PDF_PARSE_URL=https://aiproxy.hzh.sealos.run/v1/parse/pdf?model=parse-pdf
CUSTOM_PDF_PARSE_KEY=your-sealos-api-key
```
##### Use SoMark
1. Open [SoMark Studio](https://somark.ai/Studio/apikey) and create an API key.
2. Add the API key to FastGPT:
```dotenv
SOMARK_API_KEY=sk-your-api-key
```
The SoMark synchronous parsing endpoint accepts files up to 200 MB and 300 pages. See the [SoMark API documentation](https://docs.somark.ai/en/api-reference) for the complete limits and error codes.
##### Use Another Custom PDF Parsing Service
```dotenv
CUSTOM_PDF_PARSE_URL=https://your-pdf-parser.example.com/v2/parse/file
CUSTOM_PDF_PARSE_KEY=your-service-key
```
`CUSTOM_PDF_PARSE_KEY` is optional. When set, FastGPT sends it to the parsing service as `Authorization: Bearer `. The service must accept a `multipart/form-data` POST request with a `file` field and return JSON in this format:
```json
{
"pages": 10,
"markdown": "Parsed Markdown content"
}
```
##### Use TextIn
```dotenv
TEXTIN_APP_ID=your-app-id
TEXTIN_SECRET_CODE=your-secret-code
```
##### Use Doc2x
```dotenv
DOC2X_KEY=your-api-key
```
Restart FastGPT after changing the environment variables. Then enable **Enhanced PDF Parsing** when importing files into a Knowledge Base or configuring App file uploads. PDFs use the configured enhanced parsing service only when this option is enabled; otherwise, FastGPT uses its built-in parser.
## Admin-Specific Variables
These variables are mainly read by `pro/admin`. Admin also uses the shared App/Admin variables above.
| Variable | Default | Description |
| ------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------ |
| `PRO_TOKEN` | None, **required** | Service-to-service token for FastGPT app calls to pro/admin internal APIs. Must be at least 32 characters and match App. |
| `EVAL_LINE_LIMIT` | `1000` | Maximum number of rows allowed when creating one evaluation task. Also sent to frontend config. |
| `BATCH_UPDATE_TIME` | `3000` | Wallet balance batch update interval, in milliseconds. |
| `INVOICE_FEISHU_WEBHOOK_URL` | Empty | Lark webhook URL for invoice application notifications. |
| `INVOICE_FEISHU_WEBHOOK_CALLBACK_URL` | Empty | Callback URL for buttons in invoice notifications. |
| `SMS_PROXY` | Empty | SMS sending proxy service URL. |
| `MAX_CRAWL_PAGE` | `2000` | Maximum number of pages to crawl during website sync. |
| `CRAWL_MAX_HTML_SIZE` | `10` | Estimated maximum HTML size for one static crawled page, in MB. |
| `CRAWL_EXCLUDE_LIST` | Empty | Crawler exclusion rules for domains or paths. Use commas to separate values. |
| `SHOW_GIT` | `false` | Whether Git information is shown in Admin. |
| `CLEAR_FREE_ACCOUNT` | `false` | Whether free account resource cleanup is enabled. |
| `SYNC_MEMBER_CRON` | Empty | Cron expression for automatic member sync. Empty disables the sync task. |
| `WORKORDER_BASE_URL` | Empty | Work order system URL. When set, the frontend shows work order entry points. |
| `WORKORDER_JWT_SECRET` | Empty | Secret used to sign JWTs when creating work orders. |
| `EXTERNAL_USER_SYSTEM_BASE_URL` | Empty | External user system URL. |
| `EXTERNAL_USER_SYSTEM_AUTH_TOKEN` | Empty | Authentication token for the external user system. |
| `BAIDU_CONVERSION_TOKEN` | Empty | Baidu conversion tracking token. |
| `BAIDU_CONVERSION_BASE_URL` | Empty | Baidu conversion tracking API URL. |
| `BING_ADS_DEVELOPER_TOKEN` | Empty | Bing Ads developer token. |
| `BING_ADS_CUSTOMER_ID` | Empty | Bing Ads customer ID. |
| `BING_ADS_CUSTOMER_ACCOUNT_ID` | Empty | Bing Ads customer account ID. |
| `BING_ADS_CONVERSION_NAME` | `fastgptcn` | Bing Ads conversion goal name. |
| `BING_OAUTH_CLIENT_ID` | Empty | Bing OAuth client ID. |
| `BING_OAUTH_CLIENT_SECRET` | Empty | Bing OAuth client secret. |
| `BING_OAUTH_REFRESH_TOKEN` | Empty | Bing OAuth refresh token. |
| `SHOW_WECOM_CONFIG` | `false` | Whether WeCom configuration is shown. |
| `WECOM_DEV` | `false` | Development mode switch for WeCom Pay. |
## Code Sandbox Variables
These variables are loaded and validated by `projects/code-sandbox/src/env.ts`. When App calls the sandbox, `CODE_SANDBOX_TOKEN` must match `SANDBOX_TOKEN`.
| Variable | Default | Description |
| --------------------------------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| `SANDBOX_PORT` | `3000` | Code Sandbox listening port. |
| `SANDBOX_TOKEN` | Empty | Bearer token for the `/sandbox` endpoint. Empty disables API authentication. It only allows printable ASCII characters and cannot contain spaces. |
| `SANDBOX_POOL_SIZE` | `20` | Number of pre-warmed JS/Python workers, from `1` to `100`. |
| `SANDBOX_QUEUE_ID_CONCURRENCY` | Empty | Number of requests with the same `queueId` that may enter execution concurrently. Empty disables `queueId` queueing. Range: `1` to `100`. |
| `SANDBOX_API_MAX_BODY_MB` | `8` | Maximum `/sandbox` API JSON body size, including `variables`, in MB. Range: `1` to `100`. |
| `SANDBOX_MAX_TIMEOUT` | `60000` | Timeout for one code execution, in milliseconds. Range: `1000` to `600000`. |
| `SANDBOX_MAX_MEMORY_MB` | `256` | Maximum memory for one sandbox, in MB. Range: `32` to `4096`. The runtime reserves an extra `50` MB for overhead. |
| `SANDBOX_MAX_OUTPUT_MB` | `10` | Maximum output JSON size for one code execution, including return values and logs, in MB. Range: `1` to `100`. |
| `CHECK_INTERNAL_IP` | `true` | Whether internal IP checks are enabled for sandbox network requests. |
| `SANDBOX_REQUEST_MAX_COUNT` | `30` | Maximum number of network requests allowed during one code execution. Range: `1` to `1000`. |
| `SANDBOX_REQUEST_TIMEOUT` | `60000` | Timeout for one network request from inside the sandbox, in milliseconds. Range: `1000` to `300000`. |
| `SANDBOX_REQUEST_MAX_RESPONSE_MB` | `10` | Maximum response body size for one sandbox network request, in MB. Range: `1` to `100`. |
| `SANDBOX_REQUEST_MAX_BODY_MB` | `5` | Maximum request body size for one sandbox network request, in MB. Range: `1` to `100`. |
| `SANDBOX_JS_ALLOWED_MODULES` | `lodash,dayjs,moment,uuid,crypto-js,qs,url,querystring` | Module allowlist for JavaScript code. Use commas to separate modules. |
| `SANDBOX_PYTHON_ALLOWED_MODULES` | Common standard libraries plus `numpy,pandas,matplotlib` | Module allowlist for Python code. Use commas to separate modules. |
| `NODE_ENV` | Empty | Standard Node environment variable. Internal address checks are relaxed in `development`. |
| `HOSTNAME` | `localhost` | Sandbox service host used for local-address detection. |
| `PORT` | `3000` | Sandbox local service port used for local-address detection. Actual listening uses `SANDBOX_PORT` first. |
## Volume Manager Variables
These variables are loaded and validated by `projects/volume-manager/src/env.ts`. The `AGENT_SANDBOX_OPENSANDBOX_VOLUME_MANAGER_TOKEN` used by FastGPT for persistent OpenSandbox volumes must match `VM_AUTH_TOKEN`.
| Variable | Default | Description |
| -------------------------- | ---------------------- | ---------------------------------------------------------------- |
| `PORT` | `3000` | Volume Manager listening port. |
| `VM_AUTH_TOKEN` | None, **required** | API authentication token for Volume Manager. |
| `VM_RUNTIME` | `kubernetes` | Runtime type, either `docker` or `kubernetes`. |
| `VM_DOCKER_SOCKET` | `/var/run/docker.sock` | Docker socket path. Required only in `docker` mode. |
| `VM_DOCKER_API_VERSION` | `v1.44` | Docker API version. Required only in `docker` mode. |
| `VM_K8S_NAMESPACE` | `opensandbox` | Kubernetes namespace. Required only in `kubernetes` mode. |
| `VM_K8S_PVC_STORAGE_CLASS` | `standard` | Kubernetes PVC StorageClass. Required only in `kubernetes` mode. |
| `VM_LOG_LEVEL` | `info` | Log level. Supported values are `debug`, `info`, and `none`. |
file: ./content/self-host/config/env.mdx
meta: {
"title": "环境变量说明",
"description": "projects/app、projects/code-sandbox 与 pro/admin 环境变量说明"
}
本文说明 FastGPT 自部署时常用服务的环境变量。`projects/app` 与 `pro/admin` 会大量复用 `packages/service/env.ts` 中的服务端配置,因此数据库、密钥、对象存储、向量库等变量合并说明;只有 `projects/app` 或 `pro/admin` 自己读取的变量单独列出。
## 说明
* `projects/app`:主应用服务,包含 Next.js 页面、API 路由、工作流、知识库、对象存储、向量库等能力。
* `pro/admin`:商业版 Admin 服务。除自己的后台功能变量外,也会复用 App/Service 的数据库、密钥、对象存储、模型、日志等变量。
* `projects/code-sandbox`:代码沙箱服务,对外暴露 `/sandbox` 执行接口,供 App 通过 `CODE_SANDBOX_URL` 调用。
* 代码中 `packages/service/env.ts` 导出名为 `serviceEnv`,`projects/app/src/env.ts` 导出名为 `appEnv`。
* App/Admin 共享布尔变量使用 `true`、`1`、`yes` 或 `y` 表示开启;其他值视为关闭。
* `FILE_TOKEN_KEY`、`AES256_SECRET_KEY` 与 `INVOKE_TOKEN_SECRET` 为运行期必填,建议使用随机强密钥,不要使用示例值。
## App/Admin 共享变量
这些变量主要由 `packages/service/env.ts` 校验,适用于 `projects/app`,也适用于会导入 `@fastgpt/service` 的 `pro/admin`。注意:`packages/service/env.ts` 当前也包含少量 App 侧开关;这类变量在下方 `projects/app` 额外变量中单独列出。
### 基础与密钥
| 变量 | 默认值 | 说明 |
| --------------------- | ------------------ | --------------------------------------------------------------------------- |
| `DB_MAX_LINK` | `5` | MongoDB、PG、OceanBase、openGauss 等数据库连接池最大连接数。 |
| `SYNC_INDEX` | `true` | 是否在启动时创建缺失的 MongoDB 索引并清理显式声明的废弃索引;关闭后需自行维护索引。 |
| `FILE_TOKEN_KEY` | 无,**必填** | 文件读取、文件鉴权相关密钥,长度至少 6 位。 |
| `AES256_SECRET_KEY` | 无,**必填** | AES 加解密密钥,长度至少 6 位。 |
| `INVOKE_TOKEN_SECRET` | 无,**必填** | Invoke 反向调用 JWT 密钥,长度至少 32 位。 |
| `ROOT_KEY` | `fastgpt_root_key` | 当前系统管理员 API 密钥,可用于调用 `/api/admin/**` 接口,长度至少 6 位。 |
| `PRO_TOKEN` | 空 | FastGPT app 服务端调用 pro/admin 内部接口的凭证,需与 pro/admin 配置一致;App 配置 `PRO_URL` 时必填。 |
| `PRO_URL` | 空 | 商业版服务地址,配置后 App 可调用 Pro API,也会作为文件 URL 安全校验允许域名。 |
### 服务地址与集成
| 变量 | 默认值 | 说明 |
| ------------------------ | ----------------------------------- | ----------------------------------------------------------- |
| `PLUGIN_BASE_URL` | `http://localhost:3004` | FastGPT Plugin 服务地址;部署模板通常会配置为内部 Plugin 服务地址。 |
| `PLUGIN_TOKEN` | `token` | 调用 Plugin 服务使用的认证 Token;需与 Plugin 服务配置一致。 |
| `CODE_SANDBOX_URL` | `http://localhost:3002` | Code Sandbox 服务地址;部署模板通常会配置为内部 Code Sandbox 服务地址。 |
| `CODE_SANDBOX_TOKEN` | `codesandbox` | App 调用 Code Sandbox 时使用的认证 Token,需与沙箱服务 `SANDBOX_TOKEN` 一致。 |
| `AIPROXY_API_ENDPOINT` | 空 | AI Proxy 服务地址;配置后模型请求会优先走 AI Proxy。 |
| `AIPROXY_API_TOKEN` | 空 | 调用 AI Proxy 使用的认证 Token。 |
| `OPENAI_BASE_URL` | `https://api.openai.com/v1` | 未配置 AI Proxy 时,兼容 OpenAI 协议的默认模型接口地址。 |
| `CHAT_API_KEY` | 空 | 未配置 AI Proxy Token 时,兼容 OpenAI 协议的默认模型 API Key。 |
| `CRM_API_URL` | 空 | 官网访客归因 CRM 的 API 基础地址(包含 `/api/v1`);为空时不进行身份上报。 |
| `CRM_API_KEY` | 空 | CRM 管理 API Key,用于注册或登录成功后按 `visitor_id` 绑定 FastGPT 用户。 |
| `MARKETPLACE_URL` | `https://v2.marketplace.fastgpt.cn` | 插件市场接口地址。 |
| `FEISHU_BASE_URL` | `https://open.feishu.cn` | 飞书开放平台地址,私有化飞书可改为对应域名。 |
| `DINGTALK_BASE_URL` | `https://api.dingtalk.com` | 钉钉新版 API 基础地址。 |
| `DINGTALK_OAPI_BASE_URL` | `https://oapi.dingtalk.com` | 钉钉 OAPI 基础地址。 |
| `YUQUE_DATASET_BASE_URL` | `https://www.yuque.com` | 语雀知识库地址。 |
### Agent Sandbox
| 变量 | 默认值 | 说明 |
| ------------------------------------------------ | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------- |
| `AGENT_SANDBOX_PROVIDER` | 空 | Agent 沙箱提供方,可选 `sealosdevbox`、`opensandbox`;为空时不启用沙箱。配置后必须同时配置对应 provider 的必填变量;`fastgpt-app` 还需要三项 Proxy 变量,`fastgpt-pro` 只需要预览 Proxy URL。 |
| `AGENT_SANDBOX_SEALOS_BASEURL` | 空 | Sealos Devbox 服务地址。 |
| `AGENT_SANDBOX_SEALOS_TOKEN` | 空 | Sealos Devbox 访问 Token。 |
| `AGENT_SANDBOX_SEALOS_WORK_DIRECTORY` | `/home/devbox/workspace` | Sealos Devbox 沙箱内工作目录。 |
| `AGENT_SANDBOX_SEALOS_IMAGE` | 空 | Sealos Devbox 使用的运行态镜像;启用 `sealosdevbox` 时必填。 |
| `AGENT_SANDBOX_OPENSANDBOX_BASEURL` | 空 | OpenSandbox 服务地址。 |
| `AGENT_SANDBOX_OPENSANDBOX_API_KEY` | 空 | OpenSandbox API Key;启用 OpenSandbox 时必填,并且必须与 OpenSandbox server 的 `[server].api_key` 一致。 |
| `AGENT_SANDBOX_OPENSANDBOX_RUNTIME` | `docker` | OpenSandbox 运行时,可选 `docker` 或 `kubernetes`。 |
| `AGENT_SANDBOX_OPENSANDBOX_IMAGE` | 空 | OpenSandbox 使用的完整运行态镜像;启用 `opensandbox` 时必填。 |
| `AGENT_SANDBOX_OPENSANDBOX_USE_SERVER_PROXY` | `true` | OpenSandbox 是否通过服务端代理访问。 |
| `AGENT_SANDBOX_OPENSANDBOX_VOLUME_MANAGER_URL` | 空 | OpenSandbox 模式下必填,Volume Manager 服务地址。 |
| `AGENT_SANDBOX_OPENSANDBOX_VOLUME_MANAGER_TOKEN` | 空 | OpenSandbox 模式下必填,Volume Manager 认证 Token。 |
| `AGENT_SANDBOX_OPENSANDBOX_VOLUME_NAME_PREFIX` | `fastgpt-session` | FastGPT app 生成 OpenSandbox 持久卷 `claimName` 时使用的前缀;升级时需沿用旧 `VM_VOLUME_NAME_PREFIX` 的值。 |
| `AGENT_SANDBOX_PROXY_SECRET` | 空 | agent-sandbox-proxy 与主站共用的 HMAC 密钥;`fastgpt-app` 启用 Agent Sandbox 时必填,至少 32 字节。 |
| `AGENT_SANDBOX_PROXY_URL` | 空 | 浏览器访问 agent-sandbox-proxy 的 WebSocket 地址;`fastgpt-app` 启用 Agent Sandbox 时必填,必须以 `ws://` 或 `wss://` 开头。 |
| `AGENT_SANDBOX_PREVIEW_PROXY_URL` | 空 | 浏览器访问 Sandbox 文件预览的 HTTP(S) 地址;`fastgpt-app` 和 `fastgpt-pro` 启用 Agent Sandbox 时都必须增加。建议使用与 FastGPT 主站不同的 origin。 |
| `AGENT_SANDBOX_FREE_TIP` | `false` | 前端是否展示 Agent Sandbox 免费提示。 |
| `AGENT_SANDBOX_STORAGE_SIZE_GI` | `1` | Agent Sandbox 存储容量,单位 Gi;FastGPT 根据该值计算归档、Skill 和单文件限制。 |
| `AGENT_SANDBOX_SUSPEND_MINUTES` | `60` | 运行中的 Agent 沙箱持续未活跃多少分钟后自动暂停。 |
| `AGENT_SANDBOX_ARCHIVE_INACTIVE_DAYS` | `7` | 已暂停的 Agent 沙箱持续未活跃多少天后自动归档。 |
| `AGENT_SANDBOX_MAX_EDIT_DEBUG` | `100` | Agent 编辑/调试沙箱数量限制。 |
| `AGENT_SANDBOX_NPM_REGISTRY` | 空 | Agent 沙箱内 npm、yarn、pnpm、bun 使用的 npm registry。 |
| `AGENT_SANDBOX_PYPI_INDEX_URL` | 空 | Agent 沙箱内 pip、`python -m pip`、uv 使用的 PyPI index URL。 |
### 数据库、缓存与向量库
| 变量 | 默认值 | 说明 |
| ---------------------------------------------- | ------------------------------------------- | -------------------------------------------- |
| `REDIS_URL` | `redis://default:mypassword@localhost:6379` | Redis 连接地址。 |
| `STREAM_RESUME_TTL_SECONDS` | `300` | 流式恢复镜像在生成中的 TTL,单位秒。 |
| `STREAM_RESUME_POST_COMPLETE_TTL_SECONDS` | `30` | 流结束后恢复镜像的缩短 TTL,单位秒。 |
| `STREAM_RESUME_REDIS_MAXMEMORY_RATIO` | `0.5` | Redis 已用内存与 `maxmemory` 比例达到该值后,不再创建新的流恢复镜像。 |
| `STREAM_RESUME_REDIS_MEMORY_CHECK_INTERVAL_MS` | `5000` | Redis 内存水位检测缓存时间,单位毫秒。 |
| `MONGODB_URI` | 本地 MongoDB 示例地址 | 主业务 MongoDB 连接地址。 |
| `MONGODB_LOG_URI` | 同 `MONGODB_URI` 示例地址 | 日志 MongoDB 连接地址;不配置时可复用主库。 |
| `VECTOR_VQ_LEVEL` | `32` | 向量量化等级;不同向量库支持范围不同。 |
| `PG_URL` | 空 | PostgreSQL/pgvector 向量库连接地址。 |
| `OCEANBASE_URL` | 空 | OceanBase 向量库连接地址。 |
| `SEEKDB_URL` | 空 | SeekDB 向量库连接地址。 |
| `MILVUS_ADDRESS` | 空 | Milvus/Zilliz 连接地址。 |
| `MILVUS_TOKEN` | 空 | Milvus/Zilliz 访问 Token。 |
| `OPENGAUSS_URL` | 空 | openGauss 向量库连接地址。 |
### 对象存储
| 变量 | 默认值 | 说明 |
| --------------------------------------- | ----------------------- | ------------------------------------------------------------------------ |
| `STORAGE_VENDOR` | `minio` | 对象存储类型,可选 `minio`、`aws-s3`、`r2`、`cos`、`oss`。 |
| `STORAGE_PUBLIC_BUCKET` | `fastgpt-public` | 公开文件 Bucket。 |
| `STORAGE_PRIVATE_BUCKET` | `fastgpt-private` | 私有文件 Bucket。 |
| `STORAGE_REGION` | `us-east-1` | 对象存储 Region。 |
| `STORAGE_EXTERNAL_ENDPOINT` | 空 | 外部可访问的对象存储地址,用于浏览器或外部服务访问。 |
| `STORAGE_R2_PUBLIC_ENDPOINT` | 空 | Cloudflare R2 公开 bucket 的 HTTPS 公网域名;`STORAGE_VENDOR=r2` 时必填。 |
| `STORAGE_S3_CDN_ENDPOINT` | 空 | `short-redirect` 临时下载地址使用的 CDN 地址;配置时必须同时配置 `STORAGE_EXTERNAL_ENDPOINT`。 |
| `STORAGE_DOWNLOAD_URL_MODE` | `short-proxy` | 下载模式,可选 `short-proxy` 或 `short-redirect`;对外始终返回 FastGPT 短链。 |
| `STORAGE_DOWNLOAD_REDIRECT_TTL_SECONDS` | `300` | `short-redirect` 模式下临时对象存储/CDN 下载地址的有效时间,单位秒。 |
| `STORAGE_S3_ENDPOINT` | `http://localhost:9000` | S3/MinIO 兼容 API 地址。 |
| `STORAGE_PUBLIC_ACCESS_EXTRA_SUB_PATH` | 空 | 公开文件访问路径的额外子路径。 |
| `STORAGE_ACCESS_KEY_ID` | `minioadmin` | 对象存储 Access Key。 |
| `STORAGE_SECRET_ACCESS_KEY` | `minioadmin` | 对象存储 Secret Key。 |
| `STORAGE_S3_FORCE_PATH_STYLE` | `false` | S3 是否强制 path-style 访问,MinIO 通常需要开启。 |
| `STORAGE_S3_MAX_RETRIES` | `3` | S3 客户端最大重试次数。 |
| `STORAGE_COS_PROTOCOL` | `https:` | 腾讯云 COS 访问协议,可选 `https:` 或 `http:`。 |
| `STORAGE_COS_USE_ACCELERATE` | `false` | 腾讯云 COS 是否使用全球加速域名。 |
| `STORAGE_COS_CNAME_DOMAIN` | 空 | 腾讯云 COS 自定义 CNAME 域名。 |
| `STORAGE_COS_PROXY` | 空 | 腾讯云 COS 代理地址。 |
| `STORAGE_OSS_ENDPOINT` | 空 | 阿里云 OSS Endpoint。 |
| `STORAGE_OSS_CNAME` | `false` | 阿里云 OSS 是否使用 CNAME。 |
| `STORAGE_OSS_INTERNAL` | `false` | 阿里云 OSS 是否使用内网 Endpoint。 |
| `STORAGE_OSS_SECURE` | `false` | 阿里云 OSS 是否使用 HTTPS。 |
| `STORAGE_OSS_ENABLE_PROXY` | `true` | 阿里云 OSS 是否启用代理访问。 |
### 日志、指标与追踪
| 变量 | 默认值 | 说明 |
| --------------------------- | ---------------- | ------------------------------------------------------------ |
| `LOG_ENABLE_CONSOLE` | `true` | 是否输出控制台日志。 |
| `LOG_CONSOLE_LEVEL` | `debug` | 控制台日志等级,可选 `trace`、`debug`、`info`、`warning`、`error`、`fatal`。 |
| `LOG_DEPTH` | `3` | 历史模板变量,用于日志对象展开深度;当前新版结构化日志主要使用日志等级配置。 |
| `LOG_ENABLE_OTEL` | `false` | 是否启用 OpenTelemetry 日志上报。 |
| `LOG_OTEL_LEVEL` | `info` | OTEL 日志等级。 |
| `LOG_OTEL_SERVICE_NAME` | `fastgpt-client` | OTEL 日志服务名。 |
| `LOG_OTEL_URL` | 空 | OTEL 日志上报地址。 |
| `METRICS_ENABLE_OTEL` | `false` | 是否启用 OpenTelemetry 指标上报。 |
| `METRICS_EXPORT_INTERVAL` | `30000` | 指标导出间隔,单位毫秒。 |
| `METRICS_OTEL_SERVICE_NAME` | `fastgpt-client` | OTEL 指标服务名。 |
| `METRICS_OTEL_URL` | 空 | OTEL 指标上报地址。 |
| `TRACING_ENABLE_OTEL` | `false` | 是否启用 OpenTelemetry 链路追踪。 |
| `TRACING_OTEL_SERVICE_NAME` | `fastgpt-client` | OTEL 追踪服务名。 |
| `TRACING_OTEL_URL` | 空 | OTEL 追踪上报地址。 |
| `TRACING_OTEL_SAMPLE_RATIO` | 空 | 追踪采样比例,范围 `0` 到 `1`。 |
| `CHAT_LOG_URL` | 空 | 对话日志推送服务地址;为空时不推送。 |
| `CHAT_LOG_INTERVAL` | 空 | 对话日志批量推送间隔,单位毫秒。 |
| `CHAT_LOG_SOURCE_ID_PREFIX` | `fastgpt-` | 对话日志来源 ID 前缀。 |
| `TRACK_BATCH_UPDATE_TIME` | `10000` | 事件计数批量写入间隔,单位毫秒。 |
### 域名、前端与运行时
| 变量 | 默认值 | 说明 |
| ------------------------- | --------------------------- | ----------------------------------------------------------------------------------- |
| `FE_DOMAIN` | 必填 | 客户端访问 FastGPT 时使用的地址(由协议、主机和可选端口组成),用于补全文件、图片等资源路径;本地开发可使用 `http://localhost:3000`。 |
| `FILE_DOMAIN` | 空 | 文件访问域名,通常也指向 FastGPT 服务;可独立域名隔离文件风险。 |
| `NEXT_PUBLIC_BASE_URL` | 空 | Next.js 子路径部署前缀,例如 `/fastgpt`;需要在构建镜像时确定。 |
| `HOSTNAME` | `localhost` | 服务本机 Host,用于内部 URL 与 SSRF 本地地址识别;容器中常设为 `0.0.0.0`。 |
| `PORT` | `3000` | Next.js 服务监听端口,也用于本地地址识别。 |
| `NODE_ENV` | 空 | 标准 Node/Next.js 运行环境变量,生产镜像中为 `production`。 |
| `NEXT_TELEMETRY_DISABLED` | `1` | 生产镜像中关闭 Next.js Telemetry。 |
| `NODE_OPTIONS` | `--max-old-space-size=4096` | 生产镜像构建阶段使用的 Node.js 启动参数,用于提高构建内存上限。 |
### 安全配置
| 变量 | 默认值 | 说明 |
| ----------------------------------- | ------- | ------------------------------------------- |
| `USE_IP_LIMIT` | `false` | 是否启用部分接口的 IP 限流。 |
| `CHECK_INTERNAL_IP` | `false` | 是否启用内网 IP 检查,用于降低 SSRF 风险。 |
| `AUTH_COOKIE_SECURE` | `false` | 是否为登录 Cookie 添加 `Secure` 属性;仅在全站 HTTPS 时启用。 |
| `TRUSTED_PROXY_ENABLE` | `false` | 是否启用可信反向代理客户端 IP 校验;关闭时兼容旧逻辑。 |
| `TRUSTED_PROXY_IPS` | 空 | 可信反向代理 IP/CIDR 列表,逗号或空白分隔。 |
| `PASSWORD_LOGIN_MINUTE_LIMIT_COUNT` | `10` | 单账号每分钟允许的密码登录请求次数。 |
| `MAX_LOGIN_SESSION` | `10` | 单账号最大登录客户端数量。 |
| `ALLOWED_ORIGINS` | 空 | 允许跨域来源,多个来源使用英文逗号分隔;为空默认允许所有跨域。 |
| `MULTIPLE_DATA_TO_BASE64` | `false` | 是否强制将图片转成 base64 传递给模型。 |
| `DISABLE_CACHE` | `false` | 是否关闭系统缓存命中,主要用于调试。 |
| `HTTP_PROXY` | 空 | Node/worker 出站 HTTP 代理。 |
| `HTTPS_PROXY` | 空 | Node/worker 出站 HTTPS 代理。 |
| `NO_PROXY` | 空 | 不走代理的地址列表。 |
| `ALL_PROXY` | 空 | 通用出站代理。 |
### 功能开关与限制
| 变量 | 默认值 | 说明 |
| -------------------------------------- | ----------- | -------------------------------------------------------------- |
| `AGENT_ENGINE` | `fastAgent` | Agent 引擎,可选 `fastAgent` 或 `piAgent`。 |
| `SKIP_FILE_TYPE_CHECK` | `false` | 是否跳过上传文件类型检查。 |
| `WECHAT_CHANNEL_CONCURRENCY` | `1000` | 微信渠道 poll worker 并发数,最小 `10`。 |
| `PARSE_FILE_WORKERS` | `5` | 文件解析 worker 常驻线程数。 |
| `HTML_TO_MARKDOWN_WORKERS` | `10` | HTML 转 Markdown worker 常驻线程数。 |
| `TEXT_TO_CHUNKS_WORKERS` | `10` | 文本切块 worker 常驻线程数。 |
| `PARSE_FILE_TIMEOUT_SECONDS` | `600` | 文件解析单任务超时时间,单位秒。 |
| `WORKFLOW_MAX_RUN_TIMES` | `500` | 工作流最大运行次数,避免极端死循环。 |
| `WORKFLOW_MAX_LOOP_TIMES` | `100` | 循环/并行节点最大输入数组长度。 |
| `WORKFLOW_PARALLEL_MAX_CONCURRENCY` | `10` | 并行节点并发上限,且不能超过 `WORKFLOW_MAX_LOOP_TIMES`。 |
| `SYSTEM_MAX_STRING_LENGTH_M` | `100` | 系统变量替换等同步字符串处理最大字符数,单位 M;`1` 表示 `1,000,000` 字符,范围 `1` 到 `100`。 |
| `CHAT_MAX_QPM` | `5000` | 聊天 QPM 限制;若用户套餐另有限制,以套餐限制为准。 |
| `SERVICE_REQUEST_MAX_CONTENT_LENGTH` | `10` | 服务端接收请求体最大大小,单位 MB。 |
| `MAX_FOLDER_DEPTH` | `4` | 允许的最深文件夹层级,根目录下最多 4 层文件夹;范围 `2` 到 `20`。 |
| `APP_FOLDER_MAX_AMOUNT` | `1000` | 应用文件夹最大数量。 |
| `DATASET_FOLDER_MAX_AMOUNT` | `1000` | 数据集文件夹最大数量。 |
| `UPLOAD_FILE_MAX_SIZE` | `1000` | 最大上传文件大小,单位 MB。 |
| `UPLOAD_FILE_MAX_AMOUNT` | `1000` | 最大上传文件数量。 |
| `LLM_REQUEST_TRACKING_RETENTION_HOURS` | `6` | LLM 请求追踪保留时长,单位小时。 |
| `MAX_HTML_TRANSFORM_CHARS` | `1000000` | HTML 转 Markdown 的最大字符数,超过后不转换。 |
## App 额外变量
以下变量主要由 `projects/app` 读取。其中部分变量当前定义在 `packages/service/env.ts` 中做统一校验,但实际消费点仍在 App 层。
| 变量 | 默认值 | 说明 |
| ------------------------------- | -------- | ------------------------------------------------- |
| `DEFAULT_ROOT_PSW` | `123456` | 初始化 root 用户默认密码。 |
| `SYSTEM_NAME` | `AI` | 页面标题默认系统名。 |
| `SYSTEM_DESCRIPTION` | 空 | 页面 Meta 描述,不配置时使用默认国际化文案。 |
| `SYSTEM_FAVICON` | 空 | 页面 favicon 地址,不配置时使用系统配置中的 favicon。 |
| `CHINESE_IP_REDIRECT_URL` | 空 | 前端配置中的中国 IP 跳转地址。 |
| `PAY_FORM_URL` | 空 | 前端配置中的付费表单地址。 |
| `SHOW_COUPON` | `false` | 是否展示兑换码功能。 |
| `SHOW_DISCOUNT_COUPON` | `false` | 是否展示优惠券功能。 |
| `HIDE_CHAT_COPYRIGHT_SETTING` | `false` | 是否隐藏版权信息配置项。 |
| `WECOM_LOGIN_AUTO_REDIRECT` | `false` | 是否允许企微终端自动跳转登录。 |
| `APP_REGISTRATION_URL` | 空 | 应用备案申请地址;当前主要作为兼容配置保留。 |
| `PASSWORD_EXPIRED_MONTH` | 空 | 密码过期月份数;为空表示不过期。 |
| `OPENAPI_KEY_MAX_COUNT` | `100` | 单个团队成员最多可创建的系统 API Key 数量,最小值为 1。 |
| `SSE_MCP_SERVER_PROXY_ENDPOINT` | 空 | MCP SSE Server 代理地址,末尾不要带 `/`。发布 SSE MCP 应用时需要配置。 |
### 开源版特有
从 4.15.0 起,开源版不再读取 `config.json`。如果从旧版本升级,请删除该文件的 volume 挂载,并按下表将原配置改为环境变量。没有使用过这些可选配置时,无需额外添加。
| 原 `config.json` 字段 | 当前环境变量 | 默认值 | 说明 |
| ------------------------------------------- | --------------------------- | -------- | ------------------------------------------------------- |
| `systemEnv.customPdfParse.url` | `CUSTOM_PDF_PARSE_URL` | 空 | 自定义 PDF 解析服务地址。 |
| `systemEnv.customPdfParse.key` | `CUSTOM_PDF_PARSE_KEY` | 空 | 自定义 PDF 解析服务密钥。 |
| `systemEnv.customPdfParse.doc2xKey` | `DOC2X_KEY` | 空 | Doc2x PDF 解析服务密钥。 |
| `systemEnv.customPdfParse.textinAppId` | `TEXTIN_APP_ID` | 空 | 合合信息 TextIn 服务 App ID。 |
| `systemEnv.customPdfParse.textinSecretCode` | `TEXTIN_SECRET_CODE` | 空 | 合合信息 TextIn 服务 Secret Code。 |
| `systemEnv.hnswEfSearch` | `HNSW_EF_SEARCH` | `100` | 向量检索的 `hnsw.ef_search` 参数,仅对 PG、OceanBase、openGauss 生效。 |
| `systemEnv.hnswMaxScanTuples` | `HNSW_MAX_SCAN_TUPLES` | `100000` | 向量检索最大扫描数据量,仅对 PG 生效。 |
| `systemEnv.datasetParseMaxProcess` | `DATASET_PARSE_MAX_PROCESS` | `10` | 知识库文件解析队列最大并发数。 |
| `systemEnv.vectorMaxProcess` | `VECTOR_MAX_PROCESS` | `10` | 向量训练队列最大并发数。 |
| `systemEnv.qaMaxProcess` | `QA_MAX_PROCESS` | `10` | 问答拆分队列最大并发数。 |
| `systemEnv.vlmMaxProcess` | `VLM_MAX_PROCESS` | `10` | 图片理解模型处理队列最大并发数。 |
#### PDF 增强解析配置
开源版支持接入自定义 PDF 解析服务、SoMark、TextIn 或 Doc2x。选择一种服务配置即可;如果同时配置多种服务,调用优先级为:自定义 PDF 解析服务、SoMark、TextIn、Doc2x。
##### 使用 Sealos PDF 解析服务
1. 打开 [Sealos AI Proxy](https://hzh.sealos.run/?uid=fnWRt09fZP\&openapp=system-aiproxy),申请 API Key。
2. 将 API Key 配置到 FastGPT:
```dotenv
CUSTOM_PDF_PARSE_URL=https://aiproxy.hzh.sealos.run/v1/parse/pdf?model=parse-pdf
CUSTOM_PDF_PARSE_KEY=your-sealos-api-key
```
##### 使用 SoMark
1. 打开 [SoMark Studio](https://somark.ai/Studio/apikey),创建 API Key。
2. 将 API Key 配置到 FastGPT:
```dotenv
SOMARK_API_KEY=sk-your-api-key
```
SoMark 同步解析接口单文件最大支持 200 MB、300 页。完整限制和错误码见 [SoMark API 文档](https://docs.somark.ai/en/api-reference)。
##### 使用其他自定义 PDF 解析服务
```dotenv
CUSTOM_PDF_PARSE_URL=https://your-pdf-parser.example.com/v2/parse/file
CUSTOM_PDF_PARSE_KEY=your-service-key
```
`CUSTOM_PDF_PARSE_KEY` 可选。配置后,FastGPT 会通过 `Authorization: Bearer ` 请求解析服务。解析服务需接收包含 `file` 字段的 `multipart/form-data` POST 请求,并返回以下 JSON:
```json
{
"pages": 10,
"markdown": "Parsed Markdown content"
}
```
##### 使用 TextIn
```dotenv
TEXTIN_APP_ID=your-app-id
TEXTIN_SECRET_CODE=your-secret-code
```
##### 使用 Doc2x
```dotenv
DOC2X_KEY=your-api-key
```
修改环境变量后需要重启 FastGPT。然后在知识库导入文件或应用文件上传配置中勾选“PDF 增强解析”,上传的 PDF 才会使用已配置的增强解析服务;未勾选时仍使用 FastGPT 内置解析器。
## Admin 额外变量
以下变量主要由 `pro/admin` 读取。Admin 同时也会使用上面的 App/Admin 共享变量。
| 变量 | 默认值 | 说明 |
| ------------------------------------- | ----------- | --------------------------------------------------------- |
| `PRO_TOKEN` | 无,**必填** | FastGPT app 服务端调用 pro/admin 内部接口的服务间凭证,至少 32 位,需与 App 一致。 |
| `EVAL_LINE_LIMIT` | `1000` | 单次创建评估任务允许的最大数据行数,也会下发给前端配置。 |
| `BATCH_UPDATE_TIME` | `3000` | 钱包余额批量更新间隔,单位毫秒。 |
| `INVOICE_FEISHU_WEBHOOK_URL` | 空 | 发票申请通知飞书 Webhook 地址。 |
| `INVOICE_FEISHU_WEBHOOK_CALLBACK_URL` | 空 | 发票通知中按钮回调地址。 |
| `SMS_PROXY` | 空 | 短信发送代理服务地址。 |
| `MAX_CRAWL_PAGE` | `2000` | 网站同步最大抓取页面数。 |
| `CRAWL_MAX_HTML_SIZE` | `10` | 静态网页爬虫单页 HTML 估算大小上限,单位 MB。 |
| `CRAWL_EXCLUDE_LIST` | 空 | 爬虫排除域名或路径规则,多个值使用英文逗号分隔。 |
| `SHOW_GIT` | `false` | 是否在后台展示 Git 信息。 |
| `CLEAR_FREE_ACCOUNT` | `false` | 是否启用免费账号资源清理任务。 |
| `SYNC_MEMBER_CRON` | 空 | 成员自动同步 Cron 表达式;为空则不启动同步任务。 |
| `WORKORDER_BASE_URL` | 空 | 工单系统地址;配置后前端展示工单入口。 |
| `WORKORDER_JWT_SECRET` | 空 | 创建工单时签发 JWT 使用的密钥。 |
| `EXTERNAL_USER_SYSTEM_BASE_URL` | 空 | 外部用户系统地址。 |
| `EXTERNAL_USER_SYSTEM_AUTH_TOKEN` | 空 | 外部用户系统认证 Token。 |
| `BAIDU_CONVERSION_TOKEN` | 空 | 百度转化跟踪 Token。 |
| `BAIDU_CONVERSION_BASE_URL` | 空 | 百度转化跟踪接口地址。 |
| `BING_ADS_DEVELOPER_TOKEN` | 空 | Bing Ads Developer Token。 |
| `BING_ADS_CUSTOMER_ID` | 空 | Bing Ads Customer ID。 |
| `BING_ADS_CUSTOMER_ACCOUNT_ID` | 空 | Bing Ads Customer Account ID。 |
| `BING_ADS_CONVERSION_NAME` | `fastgptcn` | Bing Ads 转化目标名称。 |
| `BING_OAUTH_CLIENT_ID` | 空 | Bing OAuth Client ID。 |
| `BING_OAUTH_CLIENT_SECRET` | 空 | Bing OAuth Client Secret。 |
| `BING_OAUTH_REFRESH_TOKEN` | 空 | Bing OAuth Refresh Token。 |
| `SHOW_WECOM_CONFIG` | `false` | 是否展示企业微信相关配置。 |
| `WECOM_DEV` | `false` | 企业微信支付相关开发模式开关。 |
## Code Sandbox 变量
这些变量由 `projects/code-sandbox/src/env.ts` 加载和校验。App 调用沙箱时,`CODE_SANDBOX_TOKEN` 需要与这里的 `SANDBOX_TOKEN` 保持一致。
| 变量 | 默认值 | 说明 |
| --------------------------------- | ------------------------------------------------------- | ----------------------------------------------------------------- |
| `SANDBOX_PORT` | `3000` | Code Sandbox 服务监听端口。 |
| `SANDBOX_TOKEN` | 空 | `/sandbox` 接口 Bearer Token;为空时不启用接口认证。仅允许 ASCII 可打印字符且不能包含空格。 |
| `SANDBOX_POOL_SIZE` | `20` | JS/Python 预热 worker 数量,范围 `1` 到 `100`。 |
| `SANDBOX_QUEUE_ID_CONCURRENCY` | 空 | 同一个 `queueId` 同时可进入执行流程的请求数;为空时不启用 `queueId` 排队,范围 `1` 到 `100`。 |
| `SANDBOX_API_MAX_BODY_MB` | `8` | `/sandbox` API JSON 请求体总大小上限,包含 `variables`,单位 MB,范围 `1` 到 `100`。 |
| `SANDBOX_MAX_TIMEOUT` | `60000` | 单次代码执行超时时间,单位毫秒,范围 `1000` 到 `600000`。 |
| `SANDBOX_MAX_MEMORY_MB` | `256` | 单个沙箱最大内存,单位 MB,范围 `32` 到 `4096`;运行时会额外预留 `50` MB 开销。 |
| `SANDBOX_MAX_OUTPUT_MB` | `10` | 单次代码执行输出 JSON 大小上限,包含返回值和日志,单位 MB,范围 `1` 到 `100`。 |
| `CHECK_INTERNAL_IP` | `true` | 是否在沙箱网络请求中启用内网 IP 检查。 |
| `SANDBOX_REQUEST_MAX_COUNT` | `30` | 单次代码执行允许发起的最大网络请求数,范围 `1` 到 `1000`。 |
| `SANDBOX_REQUEST_TIMEOUT` | `60000` | 沙箱内单次网络请求超时时间,单位毫秒,范围 `1000` 到 `300000`。 |
| `SANDBOX_REQUEST_MAX_RESPONSE_MB` | `10` | 沙箱内单次网络响应体最大大小,单位 MB,范围 `1` 到 `100`。 |
| `SANDBOX_REQUEST_MAX_BODY_MB` | `5` | 沙箱内单次网络请求体最大大小,单位 MB,范围 `1` 到 `100`。 |
| `SANDBOX_JS_ALLOWED_MODULES` | `lodash,dayjs,moment,uuid,crypto-js,qs,url,querystring` | JS 代码允许导入的模块白名单,使用英文逗号分隔。 |
| `SANDBOX_PYTHON_ALLOWED_MODULES` | 内置常用标准库与 `numpy,pandas,matplotlib` | Python 代码允许导入的模块白名单,使用英文逗号分隔。 |
| `NODE_ENV` | 空 | 标准 Node.js 环境变量;`development` 下内网地址检查会放宽。 |
| `HOSTNAME` | `localhost` | 沙箱服务本机 Host,用于本地地址识别。 |
| `PORT` | `3000` | 沙箱本地服务端口识别;实际监听优先使用 `SANDBOX_PORT`。 |
## Volume Manager 变量
这些变量由 `projects/volume-manager/src/env.ts` 加载和校验。FastGPT 侧 OpenSandbox 持久化 Volume 使用的 `AGENT_SANDBOX_OPENSANDBOX_VOLUME_MANAGER_TOKEN` 需要与这里的 `VM_AUTH_TOKEN` 保持一致。
| 变量 | 默认值 | 说明 |
| -------------------------- | ---------------------- | ------------------------------------------------ |
| `PORT` | `3000` | Volume Manager 服务监听端口。 |
| `VM_AUTH_TOKEN` | 无,**必填** | Volume Manager API 鉴权 Token。 |
| `VM_RUNTIME` | `kubernetes` | 运行时类型,可选 `docker` 或 `kubernetes`。 |
| `VM_DOCKER_SOCKET` | `/var/run/docker.sock` | Docker socket 路径,仅 `docker` 模式需要。 |
| `VM_DOCKER_API_VERSION` | `v1.44` | Docker API 版本,仅 `docker` 模式需要。 |
| `VM_K8S_NAMESPACE` | `opensandbox` | Kubernetes 命名空间,仅 `kubernetes` 模式需要。 |
| `VM_K8S_PVC_STORAGE_CLASS` | `standard` | Kubernetes PVC StorageClass,仅 `kubernetes` 模式需要。 |
| `VM_LOG_LEVEL` | `info` | 日志等级,可选 `debug`、`info` 或 `none`。 |
file: ./content/self-host/config/object-storage.en.mdx
meta: {
"title": "Object Storage Configuration",
"description": "How to configure and connect to various object storage providers via environment variables, and common configuration issues"
}
import { Alert } from '@/components/docs/Alert';
import FastGPTLink from '@/components/docs/linkFastGPT';
## Object Storage Configuration
This guide covers environment variable configuration for object storage providers supported by FastGPT, including self-hosted MinIO, AWS S3, Cloudflare R2, Alibaba Cloud OSS, and Tencent Cloud COS.
FastGPT supports MinIO, AWS S3, Alibaba Cloud OSS, Tencent Cloud COS, and Cloudflare R2. Except for local MinIO development, create `STORAGE_PUBLIC_BUCKET` and `STORAGE_PRIVATE_BUCKET` ahead of time and grant the FastGPT access key read/write permission on both buckets.
## Access Modes
* Uploads always go through the FastGPT backend proxy.
* External download URLs are always FastGPT short links. FastGPT no longer returns object storage presigned URLs directly.
* `STORAGE_DOWNLOAD_URL_MODE` supports two modes and defaults to `short-proxy`:
* `short-proxy`: FastGPT validates the short link and proxies the file stream. No public object storage endpoint is required.
* `short-redirect`: FastGPT validates the short link, then redirects to a short-lived object storage or CDN URL. File traffic bypasses FastGPT.
* Self-hosted MinIO requires `STORAGE_EXTERNAL_ENDPOINT` when using `short-redirect`.
## Provider Configuration
### MinIO
> MinIO has strong AWS S3 protocol support and is suitable for local development and self-hosted deployments. In theory, any object storage with S3 protocol support comparable to MinIO will work, such as SeaweedFS or RustFS.
* `STORAGE_S3_ENDPOINT` Internal connection address. Can be a container ID, e.g., `http://fastgpt-minio:9000`
* `STORAGE_EXTERNAL_ENDPOINT` An address accessible by both **server** and **client** to reach the bucket. Use a fixed host IP or domain name — don't use `127.0.0.1` or `localhost` (containers can't access loopback addresses). This variable does not change the download mode automatically.
* `STORAGE_S3_CDN_ENDPOINT` \[Optional] CDN endpoint used for temporary `short-redirect` download URLs. This variable does not change the default download mode and requires `STORAGE_EXTERNAL_ENDPOINT`. Uploads still go through the FastGPT backend proxy and do not use the CDN.
* `STORAGE_S3_FORCE_PATH_STYLE` \[Optional] Virtual-hosted-style or path-style routing. If vendor is `minio`, this is fixed to `true`.
* `STORAGE_S3_MAX_RETRIES` \[Optional] Maximum request retry attempts. Default: 3
**Complete Example**
> If using Sealos object storage, set `STORAGE_VENDOR` to `minio`
```dotenv
STORAGE_VENDOR=minio
STORAGE_REGION=us-east-1
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_S3_ENDPOINT=http://127.0.0.1:9000
STORAGE_S3_FORCE_PATH_STYLE=true
STORAGE_S3_MAX_RETRIES=3
```
### AWS S3
AWS S3 uses the same S3-compatible variables as MinIO. For production, create separate public and private buckets in advance and configure public-read or CloudFront/custom-domain access only for the public bucket.
```dotenv
STORAGE_VENDOR=aws-s3
STORAGE_REGION=ap-southeast-1
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_S3_ENDPOINT=https://s3.ap-southeast-1.amazonaws.com
STORAGE_S3_FORCE_PATH_STYLE=false
STORAGE_S3_MAX_RETRIES=3
```
### Alibaba Cloud OSS
> * [CORS Configuration](https://help.aliyun.com/zh/oss/user-guide/configure-cross-origin-resource-sharing/?spm=5176.8466032.console-base_help.dexternal.1bcd1450Wau6J6#b58400ec36rqf)
* `STORAGE_OSS_ENDPOINT` Alibaba Cloud OSS hostname. Default is usually `{region}.aliyuncs.com`, e.g., `oss-cn-hangzhou.aliyuncs.com`. If using a custom domain, enter it here, e.g., `your-domain.com`
* `STORAGE_OSS_CNAME` Whether custom domain is enabled
* `STORAGE_OSS_SECURE` Whether TLS is enabled. Disable if your domain doesn't have a certificate.
* `STORAGE_OSS_INTERNAL` \[Optional] Whether to use internal network access. Enable if your service is also on Alibaba Cloud to save bandwidth. Default: disabled
Set the OSS public bucket to public-read and keep the private bucket private. The same Access Key can be used for both buckets, but the bucket names must remain distinct.
**Complete Example**
```dotenv
STORAGE_VENDOR=oss
STORAGE_REGION=oss-cn-hangzhou
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_OSS_ENDPOINT=oss-cn-hangzhou.aliyuncs.com
STORAGE_OSS_CNAME=false
STORAGE_OSS_SECURE=false
STORAGE_OSS_INTERNAL=false
```
### Tencent Cloud COS
> * [CORS Configuration](https://cloud.tencent.com/document/product/436/13318)
* `STORAGE_COS_PROTOCOL` Options: `https:`, `http:` — don't forget the `:`. If your custom domain doesn't have a certificate, don't use `https:`
* `STORAGE_COS_USE_ACCELERATE` \[Optional] Enable global acceleration domain. Default: false. If true, the bucket must have global acceleration enabled.
* `STORAGE_COS_CNAME_DOMAIN` \[Optional] Custom domain, e.g., `your-domain.com`
* `STORAGE_COS_PROXY` \[Optional] Proxy server, e.g., `http://localhost:7897`
COS bucket names must include the account App ID suffix, for example `fastgpt-public-1250000000`. Configure anonymous read only for the public bucket and keep the private bucket private.
**Complete Example**
```dotenv
STORAGE_VENDOR=cos
STORAGE_REGION=ap-shanghai
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_COS_PROTOCOL=http:
STORAGE_COS_USE_ACCELERATE=false
STORAGE_COS_CNAME_DOMAIN=
STORAGE_COS_PROXY=
```
### Cloudflare R2
R2 uses the S3-compatible API. Set `STORAGE_REGION` to `auto` and use the account-level S3 endpoint from Cloudflare as `STORAGE_S3_ENDPOINT`. FastGPT does not rewrite R2 presigned URLs through `STORAGE_S3_CDN_ENDPOINT`; private objects should normally use the default `short-proxy` download mode.
`STORAGE_R2_PUBLIC_ENDPOINT` is required for public objects. Set it to the HTTPS custom domain (or another public HTTPS domain bound to the bucket). This is separate from the R2 S3 API endpoint and must not contain query parameters.
For production, use a custom domain instead of the rate-limited `r2.dev` development URL. Create both R2 buckets in advance; FastGPT checks that production buckets exist at startup and does not create them automatically.
```dotenv
STORAGE_VENDOR=r2
STORAGE_REGION=auto
STORAGE_S3_ENDPOINT=https://.r2.cloudflarestorage.com
STORAGE_R2_PUBLIC_ENDPOINT=https://assets.example.com
STORAGE_ACCESS_KEY_ID=
STORAGE_SECRET_ACCESS_KEY=
STORAGE_PUBLIC_BUCKET=
STORAGE_PRIVATE_BUCKET=
STORAGE_S3_FORCE_PATH_STYLE=false
```
file: ./content/self-host/config/object-storage.mdx
meta: {
"title": "对象存储配置",
"description": "如何通过环境变量配置并连接个各厂商的对象存储"
}
import { Alert } from '@/components/docs/Alert';
import FastGPTLink from '@/components/docs/linkFastGPT';
## 对象存储服务配置介绍
这里提供了 FastGPT 目前支持的对象存储厂商,包括自部署的 MinIO、AWS S3、Cloudflare R2、阿里云 OSS 和腾讯云 COS 的环境变量配置说明
FastGPT 支持 MinIO、AWS S3、Alibaba Cloud OSS、Tencent Cloud COS 和 Cloudflare R2。除 MinIO 本地开发外,建议提前创建 `STORAGE_PUBLIC_BUCKET` 和 `STORAGE_PRIVATE_BUCKET`,并确保 FastGPT 使用的 Access Key 对两个桶都有读写权限。
## 访问模式说明
* 上传固定走 FastGPT 后端代理。
* 对外下载地址固定为 FastGPT 短链,不再直接返回对象存储预签名长链接。
* `STORAGE_DOWNLOAD_URL_MODE` 支持两种模式,默认值为 `short-proxy`:
* `short-proxy`:FastGPT 校验短链并代理文件流,无需配置公网对象存储地址。
* `short-redirect`:FastGPT 校验短链后 302 到短时效对象存储/CDN 地址,文件流量不经过 FastGPT。
* 自部署 MinIO 使用 `short-redirect` 时必须配置 `STORAGE_EXTERNAL_ENDPOINT`。
## 提供商配置
### MinIO
> MinIO 对 AWS S3 协议支持比较完整,适合本地开发和自部署场景。理论上任何对 AWS S3 协议的支持程度至少和 MinIO 相当的对象存储服务也可以使用,比如 SeaweedFS、RustFS。
* `STORAGE_S3_ENDPOINT` 内网连接地址,可以是容器 ID 连接,比如 `http://fastgpt-minio:9000`
* `STORAGE_EXTERNAL_ENDPOINT` 一个**服务器**和**客户端**均可访问到存储桶的地址,可以是固定的宿主机 IP 或者域名,注意不要填写成 127.0.0.1 或者 localhost 等本地回环地址(因为容器里无法使用)。该变量不会自动改变下载模式。
* `STORAGE_S3_CDN_ENDPOINT`【可选】`short-redirect` 临时下载地址使用的 CDN 地址。该变量不会改变默认下载模式,且配置时必须同时配置 `STORAGE_EXTERNAL_ENDPOINT`。上传仍走 FastGPT 后端代理,不会使用 CDN。
* `STORAGE_S3_FORCE_PATH_STYLE`【可选】虚拟主机风格路由或路径路由风格,其中如果厂商填写了 `minio` 的话,该值被固定为 `true`
* `STORAGE_S3_MAX_RETRIES`【可选】请求最大尝试次数,默认为 3 次
**完整示例**
> 如果使用的是 Sealos 的对象存储服务请将 `STORAGE_VENDOR` 填写为 `minio`
```dotenv
STORAGE_VENDOR=minio
STORAGE_REGION=us-east-1
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_S3_ENDPOINT=http://127.0.0.1:9000
STORAGE_S3_FORCE_PATH_STYLE=true
STORAGE_S3_MAX_RETRIES=3
```
### AWS S3
AWS S3 与 MinIO 使用同一套 S3 兼容变量。生产环境建议提前创建 public/private 两个 bucket,并为 public bucket 配置公开读取策略或 CloudFront/自定义域名。
```dotenv
STORAGE_VENDOR=aws-s3
STORAGE_REGION=ap-southeast-1
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_S3_ENDPOINT=https://s3.ap-southeast-1.amazonaws.com
STORAGE_S3_FORCE_PATH_STYLE=false
STORAGE_S3_MAX_RETRIES=3
```
### 阿里云 OSS
> * [跨域配置](https://help.aliyun.com/zh/oss/user-guide/configure-cross-origin-resource-sharing/?spm=5176.8466032.console-base_help.dexternal.1bcd1450Wau6J6#b58400ec36rqf)
* `STORAGE_OSS_ENDPOINT` 阿里云对象存储连接主机名,厂商提供的默认值一般都是 `{地区}.aliyuncs.com`,如 `oss-cn-hangzhou.aliyuncs.com`;注意,如果配置了自定义域名的话也填在这里,比如 `your-domain.com`
* `STORAGE_OSS_CNAME` 是否开启自定义域名
* `STORAGE_OSS_SECURE` 是否开启了 TLS,如果域名没有认证证书的话,请关闭该选项
* `STORAGE_OSS_INTERNAL`【可选】是否开启内网访问,如果你的服务也在阿里云的话可以开启并节省流量,默认关闭
OSS 的 public bucket 需要设置为公开读,private bucket 保持私有。两个 bucket 可以使用同一组 Access Key,但不要把两个 bucket 配成同名。
**完整示例**
```dotenv
STORAGE_VENDOR=oss
STORAGE_REGION=oss-cn-hangzhou
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_OSS_ENDPOINT=oss-cn-hangzhou.aliyuncs.com
STORAGE_OSS_CNAME=false
STORAGE_OSS_SECURE=false
STORAGE_OSS_INTERNAL=false
```
### 腾讯云 COS
> * [跨域配置](https://cloud.tencent.com/document/product/436/13318)
* `STORAGE_COS_PROTOCOL` 枚举可选值 `https:`、`http:`,注意不要忘记 `:`;如果自定义域名没有上传证书的话,请不要设置为 `https:`
* `STORAGE_COS_USE_ACCELERATE`【可选】是否启用全球加速域名,默认为 false。若改为 true,需要存储桶开启全球加速功能
* `STORAGE_COS_CNAME_DOMAIN`【可选】自定义域名,如 `your-domain.com`
* `STORAGE_COS_PROXY`【可选】代理服务器,如 `http://localhost:7897`
COS bucket 名称必须包含账号 App ID 后缀,例如 `fastgpt-public-1250000000`。public bucket 需要配置匿名读,private bucket 保持私有。
**完整示例**
```dotenv
STORAGE_VENDOR=cos
STORAGE_REGION=ap-shanghai
STORAGE_ACCESS_KEY_ID=your_access_key
STORAGE_SECRET_ACCESS_KEY=your_secret_key
STORAGE_PUBLIC_BUCKET=fastgpt-public
STORAGE_PRIVATE_BUCKET=fastgpt-private
STORAGE_COS_PROTOCOL=http:
STORAGE_COS_USE_ACCELERATE=false
STORAGE_COS_CNAME_DOMAIN=
STORAGE_COS_PROXY=
```
### Cloudflare R2
R2 使用 S3 兼容 API。`STORAGE_REGION` 固定填写 `auto`,`STORAGE_S3_ENDPOINT` 填写 Cloudflare 控制台提供的账户级 S3 endpoint。R2 不支持通过 FastGPT 的 `STORAGE_S3_CDN_ENDPOINT` 重写预签名 URL;私有对象仍建议使用默认的 `short-proxy` 下载模式。
`STORAGE_R2_PUBLIC_ENDPOINT` 必须配置为公开 bucket 的自定义域名(或其他已绑定到该 bucket 的公开 HTTPS 域名),用于生成公开文件 URL。该地址不是 R2 S3 API endpoint,也不应包含查询参数。
R2 生产环境建议使用自定义域名,不建议使用受速率限制的 `r2.dev` 公共开发 URL。R2 public/private bucket 都应提前创建;FastGPT 启动时只检查 bucket 是否存在,不会自动创建生产 bucket。
```dotenv
STORAGE_VENDOR=r2
STORAGE_REGION=auto
STORAGE_S3_ENDPOINT=https://.r2.cloudflarestorage.com
STORAGE_R2_PUBLIC_ENDPOINT=https://assets.example.com
STORAGE_ACCESS_KEY_ID=
STORAGE_SECRET_ACCESS_KEY=
STORAGE_PUBLIC_BUCKET=
STORAGE_PRIVATE_BUCKET=
STORAGE_S3_FORCE_PATH_STYLE=false
```
file: ./content/self-host/config/remote-debug-suite.en.mdx
meta: {
"title": "System Plugin Remote Debugging Suite Configuration",
"description": "Configure the system plugin remote debugging suite for self-hosted FastGPT deployments"
}
import { Alert } from '@/components/docs/Alert';
## When to Use It
The system plugin remote debugging suite temporarily connects FastGPT system plugins running on a developer's local machine to a FastGPT test environment. It is intended for system plugin development, integration testing, and acceptance checks, not as a production plugin runtime.
The system plugin remote debugging suite is available only in the commercial edition.
We recommend using remote debugging in the FastGPT Cloud version first. Self-hosted deployments require you to operate Plugin Server, Connection Gateway, Redis, reverse proxy, TLS, and secret rotation yourself.
The default Docker Compose deployment only includes the FastGPT main service and the regular `fastgpt-plugin` runtime. It does not include the public WebSocket setup required by Connection Gateway. For self-hosted deployments, deploy the system plugin remote debugging suite separately.
## Components
The remote debug flow includes these components:
| Component | Purpose |
| -------------------- | --------------------------------------------------------------------------------------- |
| FastGPT main service | Provides the UI and APIs for enabling, refreshing, and revoking a debug channel. |
| Plugin Server | Manages `connectionKey`, debug source, and forwards debug invocations to Gateway. |
| Connection Gateway | Maintains CLI WebSocket connections, sessions, mailboxes, and debug invocation streams. |
| Redis | Stores Gateway sessions, source owner leases, and mailbox data. |
| `fastgpt-plugin dev` | Runs plugins locally and connects to Gateway through WebSocket. |
Main flow:
```mermaid
sequenceDiagram
participant User as Developer
participant FastGPT as FastGPT
participant Plugin as Plugin Server
participant Gateway as Connection Gateway
participant CLI as fastgpt-plugin dev
User->>FastGPT: Enable debug channel
FastGPT->>Plugin: Create debug channel
Plugin-->>FastGPT: connectionKey / connectionUrl / source
User->>CLI: fastgpt-plugin dev --connect
CLI->>FastGPT: Exchange connectionKey
FastGPT->>Plugin: Forward connectionKey exchange
Plugin-->>CLI: gatewayUrl / connectToken / source
CLI->>Gateway: WebSocket bind
FastGPT->>Plugin: Invoke plugin under debug source
Plugin->>Gateway: Send plugin-debug.run
Gateway->>CLI: Forward debug request
CLI-->>Gateway: Return execution result
Gateway-->>Plugin: Stream result
```
## Prerequisites
1. The FastGPT main service can access `fastgpt-plugin`, and `PLUGIN_TOKEN` / `AUTH_TOKEN` are the same on both sides.
2. Your `fastgpt-plugin` version includes remote debugging. Use the plugin version required by your current FastGPT release.
3. The Gateway WebSocket URL must be reachable from the developer's local machine. In production, expose it through HTTPS reverse proxy as `wss://`.
4. The Gateway internal HTTP API should only be reachable from the Plugin Server's private network.
5. The Redis used by Gateway must support Stream.
6. All production secrets must be at least 32 characters and must not use example values, defaults, or weak passwords.
## Deploy Connection Gateway
Connection Gateway is maintained in the `fastgpt-plugin` repository. Choose the China Mainland or global image based on your network environment:
```dotenv
# China Mainland
CONNECTION_GATEWAY_IMAGE=registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-plugin-connection-gateway:8a52896d1d5b866308778871526cfdff9d22c547
# Global
CONNECTION_GATEWAY_IMAGE=ghcr.io/labring/fastgpt-plugin-connection-gateway:8a52896d1d5b866308778871526cfdff9d22c547
```
A minimal setup looks like this:
```yaml
services:
connection-gateway:
image: ${CONNECTION_GATEWAY_IMAGE}
restart: unless-stopped
environment:
NODE_ENV: production
REDIS_URL: redis://default:mypassword@fastgpt-redis:6379
AUTH_TOKEN: ${CONNECTION_GATEWAY_AUTH_TOKEN}
CONNECTION_GATEWAY_AUTH_TOKEN: ${CONNECTION_GATEWAY_AUTH_TOKEN}
JWT_SECRET: ${CONNECTION_GATEWAY_JWT_SECRET}
CONNECTION_GATEWAY_PORT: 3000
CONNECTION_GATEWAY_WS_PORT: 3001
CONNECTION_GATEWAY_WS_PATH: /connection-gateway/v1
ports:
- '3010:3000'
- '3011:3001'
```
Port notes:
| Port | Purpose | Exposure requirement |
| ------ | ------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- |
| `3010` | Gateway HTTP API, mapped to container port `3000`, including `/health`, `/internal/*`, and `/metrics`. | Public exposure is not required. Plugin Server only needs private network access. |
| `3011` | Gateway WebSocket, mapped to container port `3001`, default path `/connection-gateway/v1`. | Must be reachable from the developer's local CLI, usually exposed as a public `wss://` URL through reverse proxy. |
| Redis | Stores Gateway sessions, source owner leases, and mailboxes. | Public exposure is not required. The Redis version must support Stream. |
## Configure Plugin Server
Add the Gateway-related environment variables to the `fastgpt-plugin` service:
```dotenv
# Private HTTP address used by Plugin Server to call Gateway internal APIs
CONNECTION_GATEWAY_BASE_URL=http://connection-gateway:3000
# WebSocket address returned to the local CLI; it must be reachable from developer machines
CONNECTION_GATEWAY_PUBLIC_URL=wss://debug-gateway.example.com/connection-gateway/v1
# Bearer token used by Plugin Server for Gateway /internal/* and /metrics APIs
CONNECTION_GATEWAY_AUTH_TOKEN=replace-with-a-random-token-at-least-32-chars
# HMAC secret for Gateway connect tokens; must exactly match Connection Gateway
JWT_SECRET=replace-with-a-random-jwt-secret-at-least-32-chars
```
Restart `fastgpt-plugin` after updating the configuration. When `CONNECTION_GATEWAY_BASE_URL` is unset, Plugin Server disables remote debugging.
## Configure FastGPT Main Service
The FastGPT main service keeps using the regular plugin configuration:
```dotenv
PLUGIN_BASE_URL=http://fastgpt-plugin:3000
PLUGIN_TOKEN=replace-with-the-same-value-as-plugin-auth-token
NEXT_PUBLIC_BASE_URL=https://fastgpt.example.com
```
`NEXT_PUBLIC_BASE_URL` affects the generated debug connection link. For public access, set it to the FastGPT URL reachable by the browser.
## Configure Reverse Proxy
Expose only the Gateway WebSocket endpoint. Keep the Gateway internal HTTP API private.
Nginx example:
```nginx
location /connection-gateway/v1 {
proxy_pass http://connection-gateway:3001;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_read_timeout 3600s;
}
```
Do not expose `/internal/*`, `/metrics`, or the Gateway HTTP port directly to the public internet.
## Developer Connection
1. Enable the debug channel from the FastGPT plugin debug entry and copy the generated connection link.
2. Run this command in the local plugin directory:
```bash
fastgpt-plugin dev --connect ''
```
After the connection succeeds, the local CLI reports plugin metadata through Gateway. The local plugins appear in FastGPT under the current debug source. The debug source format is:
```text
debug:tmbId:{tmbId}
```
## Verification
1. Check Gateway health:
```bash
curl http://connection-gateway:3000/health
```
2. Enable the debug channel in FastGPT and confirm the status changes from `enabled` to `connected`.
3. Run `fastgpt-plugin dev` locally and confirm the CLI reports an active WebSocket connection.
4. Select a tool under the debug source in FastGPT and invoke it once. The result should come from the local plugin.
## Security Notes
* `CONNECTION_GATEWAY_AUTH_TOKEN`, `JWT_SECRET`, `connectionKey`, and `connectToken` are sensitive. Do not write them to logs, screenshots, or public docs.
* `CONNECTION_GATEWAY_AUTH_TOKEN` is only for Plugin Server. The local CLI does not need it and should never receive it.
* `connectionKey` is a long-lived debug connection secret. It is returned in plaintext only when the debug channel is enabled or refreshed. Refresh or revoke the debug channel immediately if it leaks.
* Debug source invocations use the remote debug path. If the connection or session is missing, the invocation fails instead of falling back to the production plugin runtime.
* Multi-replica Gateway deployments must route session deletion requests to the node that owns the WebSocket, or accept that calls fail after the Redis session is deleted.
## FAQ
### The debug channel opens, but the CLI cannot connect
Check whether `CONNECTION_GATEWAY_PUBLIC_URL` is reachable from the developer's local machine. The browser and CLI run on the developer's computer, so Docker private hostnames will not work.
### The CLI is connected, but FastGPT shows disconnected
Check whether Plugin Server can access `CONNECTION_GATEWAY_BASE_URL`, and confirm that `CONNECTION_GATEWAY_AUTH_TOKEN` matches the Gateway configuration.
### Tool invocation times out after connection
Check Gateway Redis, WebSocket upgrade in the reverse proxy, `proxy_read_timeout`, and whether the local CLI is still online.
### connect token validation fails
Check whether `JWT_SECRET` is exactly the same in Plugin Server and Connection Gateway.
file: ./content/self-host/config/remote-debug-suite.mdx
meta: {
"title": "系统插件的远程调试功能套件配置",
"description": "FastGPT 接入系统插件的远程调试功能套件"
}
import { Alert } from '@/components/docs/Alert';
## 适用场景
系统插件的远程调试功能套件用于把开发者本地运行的 FastGPT 系统插件临时接入 FastGPT 测试环境。它适合系统插件开发、联调和验收,不适合作为生产插件运行时。
系统插件的远程调试功能套件仅商业版支持。
优先推荐在 FastGPT 云服务版本中使用远程调试能力。自部署需要额外维护 Plugin Server、Connection Gateway、Redis、反向代理、TLS 和密钥轮换,运维成本更高。
默认的 Docker Compose 部署脚本只包含 FastGPT 主服务和常规 `fastgpt-plugin` 运行环境,不包含 Connection Gateway 的公网 WebSocket 接入配置。自部署环境需要按本文额外部署系统插件的远程调试功能套件。
## 组件关系
远程调试链路包含以下组件:
| 组件 | 作用 |
| -------------------- | ----------------------------------------------- |
| FastGPT 主服务 | 提供开启、刷新、关闭调试通道的页面和 API。 |
| Plugin Server | 管理 `connectionKey`、调试 source,并把调试调用转发给 Gateway。 |
| Connection Gateway | 维护 CLI WebSocket 长连接、session、mailbox 和调试调用流转。 |
| Redis | 保存 Gateway session、source owner 和 mailbox 数据。 |
| `fastgpt-plugin dev` | 在开发者本地运行插件,并通过 WebSocket 连接 Gateway。 |
主路径如下:
```mermaid
sequenceDiagram
participant User as Developer
participant FastGPT as FastGPT
participant Plugin as Plugin Server
participant Gateway as Connection Gateway
participant CLI as fastgpt-plugin dev
User->>FastGPT: 开启调试通道
FastGPT->>Plugin: 创建 debug channel
Plugin-->>FastGPT: connectionKey / connectionUrl / source
User->>CLI: fastgpt-plugin dev --connect
CLI->>FastGPT: 兑换 connectionKey
FastGPT->>Plugin: 转发 connectionKey exchange
Plugin-->>CLI: gatewayUrl / connectToken / source
CLI->>Gateway: WebSocket bind
FastGPT->>Plugin: 调用 debug source 下的插件
Plugin->>Gateway: 发送 plugin-debug.run
Gateway->>CLI: 转发调试请求
CLI-->>Gateway: 返回执行结果
Gateway-->>Plugin: 流式返回结果
```
## 部署前提
1. FastGPT 主服务已能正常访问 `fastgpt-plugin`,并且两侧的 `PLUGIN_TOKEN` / `AUTH_TOKEN` 一致。
2. `fastgpt-plugin` 版本需要包含远程调试能力;建议与当前 FastGPT 版本要求的 plugin 版本保持一致。
3. Gateway WebSocket 地址需要从开发者本地可访问,生产建议使用 HTTPS 反向代理暴露为 `wss://`。
4. Gateway internal HTTP API 只允许 Plugin Server 所在内网访问。
5. Gateway 使用的 Redis 必须支持 Stream。
6. 所有生产密钥至少 32 位,且不要使用示例值、默认值或弱口令。
## 部署 Connection Gateway
Connection Gateway 由 `fastgpt-plugin` 仓库维护。部署时按网络环境选择国内版或海外版镜像:
```dotenv
# 国内版
CONNECTION_GATEWAY_IMAGE=registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-plugin-connection-gateway:8a52896d1d5b866308778871526cfdff9d22c547
# 海外版
CONNECTION_GATEWAY_IMAGE=ghcr.io/labring/fastgpt-plugin-connection-gateway:8a52896d1d5b866308778871526cfdff9d22c547
```
最小配置形态如下:
```yaml
services:
connection-gateway:
image: ${CONNECTION_GATEWAY_IMAGE}
restart: unless-stopped
environment:
NODE_ENV: production
REDIS_URL: redis://default:mypassword@fastgpt-redis:6379
AUTH_TOKEN: ${CONNECTION_GATEWAY_AUTH_TOKEN}
CONNECTION_GATEWAY_AUTH_TOKEN: ${CONNECTION_GATEWAY_AUTH_TOKEN}
JWT_SECRET: ${CONNECTION_GATEWAY_JWT_SECRET}
CONNECTION_GATEWAY_PORT: 3000
CONNECTION_GATEWAY_WS_PORT: 3001
CONNECTION_GATEWAY_WS_PATH: /connection-gateway/v1
ports:
- '3010:3000'
- '3011:3001'
```
端口说明:
| 端口 | 用途 | 暴露要求 |
| ------ | -------------------------------------------------------------------- | ------------------------------------------- |
| `3010` | Gateway HTTP API,对应容器内 `3000`,包含 `/health`、`/internal/*`、`/metrics`。 | 不需要公网暴露,Plugin Server 可通过内网访问即可。 |
| `3011` | Gateway WebSocket,对应容器内 `3001`,默认路径 `/connection-gateway/v1`。 | 需要让开发者本地 CLI 可访问,通常通过反向代理暴露为公网 `wss://` 地址。 |
| Redis | Gateway session、source owner 和 mailbox 存储。 | 不需要公网暴露;Redis 版本必须支持 Stream。 |
## 配置 Plugin Server
在 `fastgpt-plugin` 服务中增加 Gateway 相关环境变量:
```dotenv
# Plugin Server 调用 Gateway internal HTTP API 的内网地址
CONNECTION_GATEWAY_BASE_URL=http://connection-gateway:3000
# 返回给本地 CLI 的 WebSocket 地址,必须能从开发者本地访问
CONNECTION_GATEWAY_PUBLIC_URL=wss://debug-gateway.example.com/connection-gateway/v1
# Plugin Server 调用 Gateway /internal/* 和 /metrics 的 bearer token
CONNECTION_GATEWAY_AUTH_TOKEN=replace-with-a-random-token-at-least-32-chars
# Gateway connect token 的 HMAC secret,必须与 Connection Gateway 完全一致
JWT_SECRET=replace-with-a-random-jwt-secret-at-least-32-chars
```
配置后重启 `fastgpt-plugin`。`CONNECTION_GATEWAY_BASE_URL` 未配置时,Plugin Server 会关闭远程调试能力。
## 配置 FastGPT 主服务
FastGPT 主服务继续使用常规插件配置:
```dotenv
PLUGIN_BASE_URL=http://fastgpt-plugin:3000
PLUGIN_TOKEN=replace-with-the-same-value-as-plugin-auth-token
NEXT_PUBLIC_BASE_URL=https://fastgpt.example.com
```
`NEXT_PUBLIC_BASE_URL` 会影响调试连接链接的生成。公网用户访问 FastGPT 时,应配置为浏览器可访问的 FastGPT 地址。
## 配置反向代理
建议只暴露 Gateway WebSocket 入口,对 Gateway internal HTTP API 保持内网访问。
Nginx 示例:
```nginx
location /connection-gateway/v1 {
proxy_pass http://connection-gateway:3001;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_read_timeout 3600s;
}
```
`/internal/*`、`/metrics` 和 Gateway HTTP 端口不要直接暴露到公网。
## 开发者连接
1. 在 FastGPT 插件调试入口开启调试通道,复制页面返回的连接链接。
2. 在本地插件目录运行:
```bash
fastgpt-plugin dev --connect ''
```
连接成功后,本地 CLI 会通过 Gateway 上报插件 metadata。FastGPT 工具列表中会出现当前调试 source 下的本地插件。调试 source 的格式为:
```text
debug:tmbId:{tmbId}
```
## 验证
1. 访问 Gateway 健康检查:
```bash
curl http://connection-gateway:3000/health
```
2. 在 FastGPT 页面开启调试通道,确认状态从 `enabled` 变为 `connected`。
3. 运行本地 `fastgpt-plugin dev`,确认 CLI 显示 WebSocket 已连接。
4. 在 FastGPT 中选择调试 source 下的工具并触发一次调用,确认结果由本地插件返回。
## 安全注意事项
* `CONNECTION_GATEWAY_AUTH_TOKEN`、`JWT_SECRET`、`connectionKey` 和 `connectToken` 都属于敏感信息,禁止写入日志、截图或公开文档。
* `CONNECTION_GATEWAY_AUTH_TOKEN` 只给 Plugin Server 使用,本地 CLI 不需要也不应获取。
* `connectionKey` 是长期调试连接密钥,只在开启或刷新调试通道时明文返回;泄露后应立即刷新或关闭调试通道。
* 调试 source 命中后按远程调试路径处理,断连或 session 不存在时会失败,不会回退到生产插件运行时。
* 多副本 Gateway 部署需要保证 session 删除请求能路由到持有 WebSocket 的节点,或接受 Redis session 删除后后续调用失败。
## 常见问题
### 页面可以开启调试,但 CLI 连接失败
检查 `CONNECTION_GATEWAY_PUBLIC_URL` 是否为开发者本地可访问地址。浏览器和 CLI 在开发者电脑上运行,不能使用 Docker 内网域名。
### CLI 已连接,但 FastGPT 显示 disconnected
检查 Plugin Server 是否能访问 `CONNECTION_GATEWAY_BASE_URL`,并确认 `CONNECTION_GATEWAY_AUTH_TOKEN` 与 Gateway 配置一致。
### 连接后调用工具超时
检查 Gateway Redis 是否正常、反向代理是否保留 WebSocket upgrade、`proxy_read_timeout` 是否过短,以及本地 CLI 是否仍在线。
### connect token 校验失败
检查 Plugin Server 和 Connection Gateway 的 `JWT_SECRET` 是否完全一致。
file: ./content/self-host/config/signoz.en.mdx
meta: {
"title": "Integrate SigNoz Service Monitoring",
"description": "FastGPT integration with SigNoz service monitoring"
}
## Introduction
[SigNoz](https://signoz.io/) is an open-source Application Performance Monitoring (APM) and observability platform that provides comprehensive service monitoring for FastGPT. Built on the OpenTelemetry standard, it collects, processes, and visualizes telemetry data from distributed systems, including tracing, metrics, and logging.
**Key Features:**
* **Distributed Tracing**: Track the complete call chain of user requests across FastGPT services
* **Performance Monitoring**: Monitor key metrics like API response times and throughput
* **Error Tracking**: Automatically capture and record system exceptions for troubleshooting
* **Log Aggregation**: Centrally collect and manage application logs with structured query support
* **Real-time Alerts**: Set alert rules based on metric thresholds to detect anomalies early
## Deploy SigNoz
You can use [SigNoz](https://signoz.io/) cloud service or self-host it. Here's how to quickly deploy SigNoz on Sealos.
1. Click the card below to deploy SigNoz with one click.
[](https://hzh.sealos.run/?uid=fnWRt09fZP\&openapp=system-template%3FtemplateName%3Dsignoz)
2. Enable external access for SigNoz
After deployment, click **Details** in P1 to open the app details page, then click **Change** in the top right and enable the external address for port 4318 (skip this step if using internal network).
| P1 | P2 | P3 |
| ----------------------------------------------- | ----------------------------------------------- | ----------------------------------------------- |
|  |  |  |
3. Get the SigNoz access address
After the change completes, wait for the public address to be ready, copy it, and enter it in FastGPT. If using internal network, copy the internal address for port 4318 directly.

## Configure FastGPT
1. Update FastGPT environment variables
**Log level options**: `trace` | `debug` | `info` | `warning` | `error` | `fatal`
```dotenv
LOG_ENABLE_CONSOLE=true # Enable console logging
LOG_CONSOLE_LEVEL=debug # Minimum log level for console output
LOG_ENABLE_OTEL=true # Enable OTEL log collection
LOG_OTEL_LEVEL=info # Minimum log level for OTEL collection
LOG_OTEL_SERVICE_NAME=fastgpt-client # Service name passed to the OTLP collector
LOG_OTEL_URL=http://localhost:4318/v1/logs # Your OTLP collector address — don't omit /v1/logs
```
2. Restart FastGPT
## Verify the Setup
Go back to the Sealos app management list, open the SigNoz frontend project, and access its public address to open the dashboard.
| | |
| ----------------------------------------------- | ----------------------------------------------- |
|  |  |
First-time access requires creating an account (data is stored in the local database) — fill in anything.

After logging in, if `logs` and `traces` are lit up in the COMPLETED steps on the right side, the configuration is successful.


## Notes
1. Adjust log retention period
SigNoz monitoring is very disk-intensive. First, avoid storing FastGPT debug logs in SigNoz. Also consider setting the log retention period to 7 days. If SigNoz data stops growing while memory keeps increasing, the disk is full — expand capacity.

file: ./content/self-host/config/signoz.mdx
meta: {
"title": "Signoz 监控服务",
"description": "FastGPT 接入 Signoz 监控服务"
}
## 介绍
[SigNoz](https://signoz.io/) 是一款开源的应用性能监控(APM)和可观测性平台,为 FastGPT 提供全面的服务监控能力。它基于 OpenTelemetry 标准,能够收集、处理和可视化分布式系统的遥测数据,包括链路追踪(Tracing)、指标监控(Metrics)和日志分析(Logging)。
**主要功能:**
* **链路追踪**:跟踪用户请求在 FastGPT 各个服务间的完整调用链路
* **性能监控**:监控 API 响应时间、吞吐量等关键性能指标
* **错误追踪**:自动捕获和记录系统异常,便于问题排查
* **日志聚合**:集中收集和管理应用日志,支持结构化查询
* **实时告警**:基于指标阈值设置告警规则,及时发现系统异常
## 部署 Signoz
可以使用 [SigNoz](https://signoz.io/) 官方云服务,或者私有部署,下面介绍在 Sealos 上快速部署 Signoz。
1. 点击下方的卡片,即可一键部署 Signoz。
[](https://hzh.sealos.run/?uid=fnWRt09fZP\&openapp=system-template%3FtemplateName%3Dsignoz)
2. 开启 Signoz 外网访问
部署后,可点击 P1 中的详情,进入应用详情页, 然后点击右上角的变更,并开启 4318 端口的外网地址(如果走内网服务,可忽略该步骤)。
| P1 | P2 | P3 |
| ----------------------------------------------- | ----------------------------------------------- | ----------------------------------------------- |
|  |  |  |
3. 获取 Signoz 访问地址
变更完成后,等待公网地址就绪,复制该地址,将其填入 FastGPT 中。如果是走内网服务,可以直接复制 4318 端口的内网地址。

## 配置 FastGPT
1. 修改 FastGPT 环境变量
**日志等级枚举**: `trace` | `debug` | `info` | `warning` | `error` | `fatal`
```dotenv
LOG_ENABLE_CONSOLE=true # 是否开启控制台打印
LOG_CONSOLE_LEVEL=debug # 控制台打印最低日志等级
LOG_ENABLE_OTEL=true # 是否开启 OTEL 日志收集
LOG_OTEL_LEVEL=info # OTEL 日志收集的最低日志等级
LOG_OTEL_SERVICE_NAME=fastgpt-client # 传递给 OTLP 收集器的服务名称
LOG_OTEL_URL=http://localhost:4318/v1/logs # 你的 OTLP 收集器的地址,不要把 /v1/logs 遗漏了
```
2. 重启 FastGPT
## 查看效果
返回 Sealos 应用管理列表,点击进入 Signoz 前端项目,并访问其公网地址,进入管理台。
| | |
| ----------------------------------------------- | ----------------------------------------------- |
|  |  |
首次注册需要注册一个账号(数据是存储本地数据库),随便填写即可。

登录进去后,如果看到右侧 COMPLETED 的步骤条中,logs 和 traces 亮起,则说明配置成功。


## 注意事项
1. 调整日志存储时长
Signoz 监控是一个非常占用磁盘的服务,首先不要把 FastGPT debug 日志也存储进来,另外可以将日志存储时长调整为 7 天。如果突然发现 Signoz 数据不增加了,并且内存一直追加,则说明是磁盘满了,需要扩大容量。

file: ./content/guide/version/commercial.en.mdx
meta: {
"title": "FastGPT Commercial Edition",
"description": "FastGPT Commercial Edition overview"
}
import { Alert } from '@/components/docs/Alert';
## Overview
FastGPT Commercial Edition is an enhanced version built on top of the Community Edition with additional exclusive features. Simply install the commercial image and configure the internal network address on your existing Community Edition setup to get started.
## Feature Comparison
| | Community Edition | Commercial Edition | Cloud Service |
| ------------------------------------------------------ | ------------------------------------------------------- | ------------------ | ------------- |
| **App Building** | | | |
| Workflow orchestration | ✅ | ✅ | ✅ |
| Share links and API | ✅ | ✅ | ✅ |
| App publishing security config | ❌ | ✅ | ✅ |
| Third-party publishing (Lark, WeChat Official Account) | ❌ | ✅ | ✅ |
| Run log dashboard | ❌ | ✅ | ✅ |
| App evaluation | ❌ | ✅ | ✅ |
| Agent and Skill assisted generation | ❌ | ✅ | ✅ |
| System tool remote debugging | ❌ | ✅ | ✅ |
| **Knowledge Base** | | | |
| Knowledge base | ✅ | ✅ | ✅ |
| Third-party knowledge base scheduled sync | ❌ | ✅ | ✅ |
| Knowledge base index enhancement | ❌ | ✅ | ✅ |
| Website sync | ❌ | ✅ | ✅ |
| Image knowledge base | ❌ | ✅ | ✅ |
| **General Features** | | | |
| Multi-model configuration | ✅ | ✅ | ✅ |
| Model log dashboard | ✅ | ✅ | ✅ |
| Model content moderation | ❌ | ✅ | ✅ |
| **Enterprise Features** | | | |
| Custom branding | ❌ | ✅ | In design |
| Multi-tenancy & billing | ❌ | ✅ | ✅ |
| Team spaces & permissions | ❌ | ✅ | ✅ |
| Admin dashboard | ❌ | ✅ | Not needed |
| SSO login | ❌ | ✅ | In design |
| Commercial license | [View open source license](./opensource/license.en.mdx) | Full | Full |
## Pricing
FastGPT Commercial Edition offers 3 pricing models based on deployment type. Below are the common details for each. If you have further questions, [contact us](https://fael3z0zfze.feishu.cn/share/base/form/shrcnjJWtKqjOI9NbQTzhNyzljc?prefill_S=doc\&hide_S=1).
**Included with all plans**
1. SaaS commercial license — use for any commercial purpose during the license period.
2. Free initial deployment assistance.
3. Priority support ticket handling.
**Plan-specific features**
| Deployment Type | Included Features | Time to Launch | Starting Price |
| --------------------------------- | ------------------------------------------------------------------------------------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Sealos Fully Managed | 1. Free upgrades during license period. 2. No ops or database management needed. | Half day | Starting at ¥10,000/month (3-month minimum) or Starting at ¥120,000/year 8C32G resources; additional resources billed separately. |
| Sealos Fully Managed (Multi-node) | 1. Free upgrades during license period. 2. No ops or database management needed. | Half day | Starting at ¥22,000/month (3-month minimum) or Starting at ¥264,000/year 32C128G resources; additional resources billed separately. |
| Self-hosted | 1. Free upgrade support for 6 versions. | Within 14 days | [Contact us for pricing](https://fael3z0zfze.feishu.cn/share/base/form/shrcnjJWtKqjOI9NbQTzhNyzljc?prefill_S=doc\&hide_S=1) |
* "6 versions of upgrade support" means the FastGPT team assists with 6 upgrades — not that the
software stops working after 6 versions. Most upgrades are straightforward enough to handle
yourself. - Fully managed is ideal for teams without dedicated ops staff — just focus on your
business. - Self-hosted gives you full control with deployment on your own servers. - Single-node
is suitable for small to mid-sized teams providing internal services; you'll manage database
backups yourself. - High-availability is designed for public-facing services, including visual
monitoring, replicas, load balancing, and automated database backups.
## Contact Us
Fill out the [inquiry form](https://fael3z0zfze.feishu.cn/share/base/form/shrcnjJWtKqjOI9NbQTzhNyzljc?prefill_S=doc\&hide_S=1) and we'll get back to you shortly.
## Technical Support
### App Customization
We can build custom workflow orchestrations tailored to your needs, delivered as a complete app configuration. Pricing is negotiable based on scope.
### Technical Services (Custom Development, Maintenance, Migration, Third-party Integration)
¥2,000 – ¥3,000 per person per day
### Upgrade Fees
Most upgrades just require pulling the new image and running the initialization script — no extra steps needed.
For cross-version or complex upgrades, follow the documentation to upgrade yourself, or pay for support at the standard technical service rate.
## FAQ
### How is delivery handled?
Full application = Community Edition image + Commercial Edition image
We provide a Commercial Edition image that requires a License to start.
### How does custom development work?
You can modify the Community Edition source code, but the Commercial Edition image cannot be modified. Since the full version = Community Edition + Commercial Edition image, you can customize part of the codebase. However, if you fork the code, you'll need to handle code merges yourself during future upgrades.
### Sealos Usage Costs
Sealos cloud services use pay-as-you-go billing. Here's the pricing table:

## Admin Dashboard Screenshots
| | | |
| ------------------------------- | ------------------------------- | ------------------------------- |
|  |  |  |
file: ./content/guide/version/commercial.mdx
meta: {
"title": "FastGPT 商业版",
"description": "FastGPT 商业版相关说明"
}
import { Alert } from '@/components/docs/Alert';
## 简介
FastGPT 商业版是基于 FastGPT 社区版的增强版本,增加了一些独有的功能。只需安装一个商业版镜像,并在社区版基础上填写对应的内网地址,即可快速使用商业版。
## 功能差异
| | 社区版 | 商业版 | 云服务版 |
| ------------------ | ---------------------------------- | --- | ---- |
| **应用构建** | | | |
| 工作流编排 | ✅ | ✅ | ✅ |
| 分享链接和 API | ✅ | ✅ | ✅ |
| 应用发布安全配置 | ❌ | ✅ | ✅ |
| 第三方应用发布(飞书、公众号) | ❌ | ✅ | ✅ |
| 运行日志看板 | ❌ | ✅ | ✅ |
| 应用评测 | ❌ | ✅ | ✅ |
| Agent 与 Skill 辅助生成 | ❌ | ✅ | ✅ |
| 系统工具远程调试 | ❌ | ✅ | ✅ |
| **知识库** | | | |
| 知识库 | ✅ | ✅ | ✅ |
| 第三方知识库定时同步 | ❌ | ✅ | ✅ |
| 知识库索引增强 | ❌ | ✅ | ✅ |
| web 站点同步 | ❌ | ✅ | ✅ |
| 图片知识库 | ❌ | ✅ | ✅ |
| **通用功能** | | | |
| 多模型配置 | ✅ | ✅ | ✅ |
| 模型日志看板 | ✅ | ✅ | ✅ |
| 模型内容审核 | ❌ | ✅ | ✅ |
| **企业级功能** | | | |
| 自定义版权信息 | ❌ | ✅ | 设计中 |
| 多租户与支付 | ❌ | ✅ | ✅ |
| 团队空间 & 权限 | ❌ | ✅ | ✅ |
| 管理后台 | ❌ | ✅ | 不需要 |
| SSO 登录 | ❌ | ✅ | 设计中 |
| 商业授权 | [查看开源协议](./opensource/license.mdx) | 完整 | 完整 |
## 商业版软件价格
FastGPT 商业版软件根据不同的部署方式,分为 3 类收费模式。下面列举各种部署方式一些常规内容,如仍有问题,可[联系咨询](https://fael3z0zfze.feishu.cn/share/base/form/shrcnjJWtKqjOI9NbQTzhNyzljc?prefill_S=doc\&hide_S=1)
**共有服务**
1. SaaS 商业授权许可 - 在商业版有效期内,可提供任意形式的商业服务。
2. 首次免费帮助部署。
3. 优先问题工单处理。
**特有服务**
| 部署方式 | 特有服务 | 上线时长 | 标品价格 |
| --------------- | ------------------------------- | ----- | ----------------------------------------------------------------------------------------------------------------- |
| Sealos 全托管 | 1. 有效期内免费升级。 2. 免运维服务&数据库。 | 半天 | 10000 元起/月(3 个月起) 或 120000 元起/年 8C32G 资源,额外资源另外收费。 |
| Sealos 全托管(多节点) | 1. 有效期内免费升级。 2. 免运维服务&数据库。 | 半天 | 22000 元起/月(3 个月起) 或 264000 元起/年 32C128G 资源,额外资源另外收费。 |
| 自有服务器部署 | 1. 6 个版本免费升级支持。 | 14 天内 | 具体价格和优惠可[联系咨询](https://fael3z0zfze.feishu.cn/share/base/form/shrcnjJWtKqjOI9NbQTzhNyzljc?prefill_S=doc\&hide_S=1) |
* 6 个版本的升级服务不是指只能用 6 个版本,而是指依赖 FastGPT
团队提供的升级服务。大部分时候,建议自行升级,也不麻烦。-
全托管版本适合技术人员紧缺的团队,仅需关注业务推动,无需关心服务是否正常运行。-
自有服务器部署版可以完全部署在自己服务器中。-
单机版适合中小团队对内提供服务,需要自己维护数据库备份等。-
高可用版适合对外提供在线服务,包含可视化监控、多副本、负载均衡、数据库自动备份等生产环境的基础设施。
## 联系方式
请填写[咨询问卷](https://fael3z0zfze.feishu.cn/share/base/form/shrcnjJWtKqjOI9NbQTzhNyzljc?prefill_S=doc\&hide_S=1),我们会尽快与您联系。
## 技术支持
### 应用定制
根据需求,定制实现某个需求的编排功能,最终会交付一个应用编排。可根据实际情况商讨。
### 技术服务费(定开、维护、迁移、三方接入等)
2000 \~ 3000 元/人/天
### 更新升级费用
大部分更新升级,重新拉镜像,然后执行一下初始化脚本就可以了,不需要执行额外操作。
跨版本更新或复杂更新可参考文档自行更新;或付费支持,标准与技术服务费一致。
## QA
### 如何交付?
完整版应用 = 社区版镜像 + 商业版镜像
我们会提供一个商业版镜像给你使用,该镜像需要一个 License 启动。
### 二次开发如何操作?
可以修改社区版部分代码,不支持修改商业版镜像。完整版本=社区版+商业版镜像,所以是可以修改部分内容的。但是如果二开了,后续则需要自己进行代码合并升级。
### Sealos 运行费用
Sealos 云服务属于按量计费,下面是它的价格表:

## 管理后台部分截图
| | | |
| ------------------------------- | ------------------------------- | ------------------------------- |
|  |  |  |
file: ./content/self-host/custom-models/bge-rerank.en.mdx
meta: {
"title": "Integrating bge-rerank Reranking Model",
"description": "Integrating bge-rerank reranking model with FastGPT"
}
## Recommended Configuration by Model
| Model Name | RAM | VRAM | Disk Space | Start Command |
| ------------------ | ----- | ----- | ---------- | ------------- |
| bge-reranker-base | >=4GB | >=4GB | >=8GB | python app.py |
| bge-reranker-large | >=8GB | >=8GB | >=8GB | python app.py |
| bge-reranker-v2-m3 | >=8GB | >=8GB | >=8GB | python app.py |
## Source Code Deployment
### 1. Environment Setup
* Python 3.9 or 3.10
* CUDA 11.7
* Network access to download models
### 2. Download Code
Code repositories for the 3 models:
1. [https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-base](https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-base)
2. [https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-large](https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-large)
3. [https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-v2-m3](https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-v2-m3)
### 3. Install Dependencies
```sh
pip install -r requirements.txt
```
### 4. Download Models
HuggingFace repositories for the 3 models:
1. [https://huggingface.co/BAAI/bge-reranker-base](https://huggingface.co/BAAI/bge-reranker-base)
2. [https://huggingface.co/BAAI/bge-reranker-large](https://huggingface.co/BAAI/bge-reranker-large)
3. [https://huggingface.co/BAAI/bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3)
Clone the model into the corresponding code directory. Directory structure:
```
bge-reranker-base/
app.py
Dockerfile
requirements.txt
```
### 5. Run
```bash
python app.py
```
On successful startup, you should see an address like this:

> `http://0.0.0.0:6006` is the connection address.
## Docker Deployment
**Image names:**
1. registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-base:v0.1 (4 GB+)
2. registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-large:v0.1 (5 GB+)
3. registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-v2-m3:v0.1 (5 GB+)
**Port**
6006
**Environment Variables**
```
ACCESS_TOKEN=your_access_token (used in request header: Authorization: Bearer ${ACCESS_TOKEN})
```
**Run Command Example**
```sh
# auth token set to mytoken
docker run -d --name reranker -p 6006:6006 -e ACCESS_TOKEN=mytoken --gpus all registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-base:v0.1
```
**docker-compose.yml Example**
```
version: "3"
services:
reranker:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-base:v0.1
container_name: reranker
# GPU runtime. If the host doesn't have GPU drivers installed, comment out the deploy section.
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
ports:
- 6006:6006
environment:
- ACCESS_TOKEN=mytoken
```
## Integrate with FastGPT
1. Open the FastGPT model configuration and add a new reranking model.
2. Fill in the model configuration form: set the Model ID to `bge-reranker-base` and the address to `{{host}}/v1/rerank`, where host is your deployed domain or IP:Port.

## FAQ
### 403 Error
The custom request token in FastGPT does not match the ACCESS\_TOKEN environment variable.
### Docker reports `Bus error (core dumped)`
Try adding the `shm_size` option to your `docker-compose.yml` to increase the shared memory size in the container.
```
...
services:
reranker:
...
container_name: reranker
shm_size: '2gb'
...
```
file: ./content/self-host/custom-models/bge-rerank.mdx
meta: {
"title": "接入 bge-rerank 重排模型",
"description": "接入 bge-rerank 重排模型"
}
## 不同模型推荐配置
推荐配置如下:
| 模型名 | 内存 | 显存 | 硬盘空间 | 启动命令 |
| ------------------ | ----- | ----- | ----- | ------------- |
| bge-reranker-base | >=4GB | >=4GB | >=8GB | python app.py |
| bge-reranker-large | >=8GB | >=8GB | >=8GB | python app.py |
| bge-reranker-v2-m3 | >=8GB | >=8GB | >=8GB | python app.py |
## 源码部署
### 1. 安装环境
* Python 3.9, 3.10
* CUDA 11.7
* 科学上网环境
### 2. 下载代码
3 个模型代码分别为:
1. [https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-base](https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-base)
2. [https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-large](https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-large)
3. [https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-v2-m3](https://github.com/labring/FastGPT/tree/main/plugins/model/rerank-bge/bge-reranker-v2-m3)
### 3. 安装依赖
```sh
pip install -r requirements.txt
```
### 4. 下载模型
3个模型的 huggingface 仓库地址如下:
1. [https://huggingface.co/BAAI/bge-reranker-base](https://huggingface.co/BAAI/bge-reranker-base)
2. [https://huggingface.co/BAAI/bge-reranker-large](https://huggingface.co/BAAI/bge-reranker-large)
3. [https://huggingface.co/BAAI/bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3)
在对应代码目录下 clone 模型。目录结构:
```
bge-reranker-base/
app.py
Dockerfile
requirements.txt
```
### 5. 运行代码
```bash
python app.py
```
启动成功后应该会显示如下地址:

> 这里的 `http://0.0.0.0:6006` 就是连接地址。
## docker 部署
**镜像名分别为:**
1. registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-base:v0.1 (4 GB+)
2. registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-large:v0.1 (5 GB+)
3. registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-v2-m3:v0.1 (5 GB+)
**端口**
6006
**环境变量**
```
ACCESS_TOKEN=访问安全凭证,请求时,Authorization: Bearer ${ACCESS_TOKEN}
```
**运行命令示例**
```sh
# auth token 为mytoken
docker run -d --name reranker -p 6006:6006 -e ACCESS_TOKEN=mytoken --gpus all registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-base:v0.1
```
**docker-compose.yml示例**
```
version: "3"
services:
reranker:
image: registry.cn-hangzhou.aliyuncs.com/fastgpt/bge-rerank-base:v0.1
container_name: reranker
# GPU运行环境,如果宿主机未安装,将deploy配置隐藏即可
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
ports:
- 6006:6006
environment:
- ACCESS_TOKEN=mytoken
```
## 接入 FastGPT
1. 打开 FastGPT 模型配置,新增一个重排模型。
2. 填写模型配置表单:模型 ID 为`bge-reranker-base`,地址填写`{{host}}/v1/rerank`,host 为你部署的域名/IP:Port。

## QA
### 403报错
FastGPT中,自定义请求 Token 和环境变量的 ACCESS\_TOKEN 不一致。
### Docker 运行提示 `Bus error (core dumped)`
尝试增加 `docker-compose.yml` 配置项 `shm_size` ,以增加容器中的共享内存目录大小。
```
...
services:
reranker:
...
container_name: reranker
shm_size: '2gb'
...
```
file: ./content/self-host/custom-models/chatglm2-m3e.en.mdx
meta: {
"title": "Integrating ChatGLM2 and M3E Models",
"description": "Integrating private ChatGLM2 and m3e-large models with FastGPT"
}
## Introduction
FastGPT uses OpenAI's LLM and embedding models by default. For private deployment, you can use ChatGLM2 and m3e-large as replacements. The following method was contributed by community user @不做了睡大觉. This image bundles both M3E-Large and ChatGLM2-6B models, ready to use out of the box.
## Deploy the Image
* Image: `stawky/chatglm2-m3e:latest`
* China mirror: `registry.cn-hangzhou.aliyuncs.com/fastgpt_docker/chatglm2-m3e:latest`
* Port: 6006
```
# Set the security token (used as the channel key in OneAPI)
Default: sk-aaabbbcccdddeeefffggghhhiiijjjkkk
You can also set it via the environment variable: sk-key. Refer to Docker documentation for how to pass environment variables.
```
## Connect to OneAPI
Documentation: [One API](../config/model/intro.en.mdx)
Add a channel for chatglm2 and m3e-large respectively, with the following parameters:

Here, m3e is used as the embedding model and chatglm2 as the language model.
## Test
curl examples:
```bash
curl --location --request POST 'https://domain/v1/embeddings' \
--header 'Authorization: Bearer sk-aaabbbcccdddeeefffggghhhiiijjjkkk' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "m3e",
"input": ["What is FastGPT"]
}'
```
```bash
curl --location --request POST 'https://domain/v1/chat/completions' \
--header 'Authorization: Bearer sk-aaabbbcccdddeeefffggghhhiiijjjkkk' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "chatglm2",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
Set Authorization to sk-aaabbbcccdddeeefffggghhhiiijjjkkk. The model field should match the custom model name you entered in One API.
## Integrate with FastGPT
Edit the config.json file. Add chatglm2 to `llmModels` and M3E to `vectorModels`:
```json
"llmModels": [
// Other chat models
{
"model": "chatglm2",
"name": "chatglm2",
"maxToken": 8000,
"price": 0,
"quoteMaxToken": 4000,
"maxTemperature": 1.2,
"defaultSystemChatPrompt": ""
}
],
"vectorModels": [
{
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"price": 0.2,
"defaultToken": 500,
"maxToken": 3000
},
{
"model": "m3e",
"name": "M3E (for testing)",
"price": 0.1,
"defaultToken": 500,
"maxToken": 1800
}
],
```
## Usage
**M3E model:**
1. Select the M3E model when creating a Knowledge Base.
Note: once selected, the embedding model for the Knowledge Base cannot be changed.

2. Import data
3. Test search

4. Bind the Knowledge Base to an app
Note: an app can only bind Knowledge Bases that use the same embedding model -- cross-model binding is not supported. You may also need to adjust the similarity threshold, as different embedding models produce different similarity (distance) scores. Test and tune accordingly.

**ChatGLM2 model:**
Simply select chatglm2 as the model.
file: ./content/self-host/custom-models/chatglm2-m3e.mdx
meta: {
"title": "接入 ChatGLM2-m3e 模型",
"description": " 将 FastGPT 接入私有化模型 ChatGLM2和m3e-large"
}
## 前言
FastGPT 默认使用了 OpenAI 的 LLM 模型和向量模型,如果想要私有化部署的话,可以使用 ChatGLM2 和 m3e-large 模型。以下是由用户@不做了睡大觉 提供的接入方法。该镜像直接集成了 M3E-Large 和 ChatGLM2-6B 模型,可以直接使用。
## 部署镜像
* 镜像名: `stawky/chatglm2-m3e:latest`
* 国内镜像名: `registry.cn-hangzhou.aliyuncs.com/fastgpt_docker/chatglm2-m3e:latest`
* 端口号: 6006
```
# 设置安全凭证(即 AI Proxy 中的渠道密钥)
默认值:sk-aaabbbcccdddeeefffggghhhiiijjjkkk
也可以通过环境变量引入:sk-key。有关docker环境变量引入的方法请自寻教程,此处不再赘述。
```
## 接入 AI Proxy
文档链接:[AI Proxy](../config/model/intro.mdx)
为 chatglm2 和 m3e-large 各添加一个渠道,参数如下:

这里我填入 m3e 作为向量模型,chatglm2 作为语言模型
## 测试
curl 例子:
```bash
curl --location --request POST 'https://domain/v1/embeddings' \
--header 'Authorization: Bearer sk-aaabbbcccdddeeefffggghhhiiijjjkkk' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "m3e",
"input": ["laf是什么"]
}'
```
```bash
curl --location --request POST 'https://domain/v1/chat/completions' \
--header 'Authorization: Bearer sk-aaabbbcccdddeeefffggghhhiiijjjkkk' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "chatglm2",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
Authorization 为 sk-aaabbbcccdddeeefffggghhhiiijjjkkk。model 为刚刚在 One API 填写的自定义模型。
## 接入 FastGPT
修改 config.json 配置文件,在 llmModels 中加入 chatglm2, 在 vectorModels 中加入 M3E 模型:
```json
"llmModels": [
//其他对话模型
{
"model": "chatglm2",
"name": "chatglm2",
"maxToken": 8000,
"price": 0,
"quoteMaxToken": 4000,
"maxTemperature": 1.2,
"defaultSystemChatPrompt": ""
}
],
"vectorModels": [
{
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"price": 0.2,
"defaultToken": 500,
"maxToken": 3000
},
{
"model": "m3e",
"name": "M3E(测试使用)",
"price": 0.1,
"defaultToken": 500,
"maxToken": 1800
}
],
```
## 测试使用
M3E 模型的使用方法如下:
1. 创建知识库时候选择 M3E 模型。
注意,一旦选择后,知识库将无法修改向量模型。

2. 导入数据
3. 搜索测试

4. 应用绑定知识库
注意,应用只能绑定同一个向量模型的知识库,不能跨模型绑定。并且,需要注意调整相似度,不同向量模型的相似度(距离)会有所区别,需要自行测试实验。

chatglm2 模型的使用方法如下:
模型选择 chatglm2 即可
file: ./content/self-host/custom-models/chatglm2.en.mdx
meta: {
"title": "Integrating ChatGLM2-6B",
"description": "Integrating the private ChatGLM2-6B model with FastGPT"
}
import { Alert } from '@/components/docs/Alert';
## Introduction
FastGPT lets you use your own OpenAI API KEY to quickly call OpenAI APIs. It currently integrates GPT-3.5, GPT-4, and embedding models for building Knowledge Bases. However, for data security reasons, you may not want to send all data to cloud-based LLMs.
So how do you connect a private model to FastGPT? This guide walks through integrating Tsinghua's ChatGLM2 as an example.
## ChatGLM2-6B Overview
ChatGLM2-6B is the second-generation version of the open-source bilingual (Chinese-English) chat model ChatGLM-6B. For details, see the [ChatGLM2-6B project page](https://github.com/THUDM/ChatGLM2-6B).
Note: ChatGLM2-6B weights are fully open for academic research. Commercial use requires official written permission. This tutorial only demonstrates one integration method and does not grant any license.
## Recommended Configuration
According to official data, generating 8192 tokens requires 12.8GB VRAM at FP16, 8.1GB at int8, and 5.1GB at int4. Quantization slightly affects performance, but not significantly.
Recommended configurations:
| Type | RAM | VRAM | Disk Space | Start Command |
| ---- | ------ | ------ | ---------- | ------------------------ |
| fp16 | >=16GB | >=16GB | >=25GB | python openai\_api.py 16 |
| int8 | >=16GB | >=9GB | >=25GB | python openai\_api.py 8 |
| int4 | >=16GB | >=6GB | >=25GB | python openai\_api.py 4 |
## Deployment
### Environment Requirements
* Python 3.8.10
* CUDA 11.8
* Network access to download models
### Source Code Deployment
1. Set up the environment as described above;
2. Download the [Python file](https://github.com/labring/FastGPT/blob/main/plugins/model/llm-ChatGLM2/openai_api.py)
3. Run `pip install -r requirements.txt`;
4. Open the Python file and configure the token in the `verify_token` method -- this adds a layer of authentication to prevent unauthorized access;
5. Run `python openai_api.py --model_name 16`. Choose the number based on the configuration table above.
Wait for the model to download and load. If you encounter errors, try asking GPT for help.
On successful startup, you should see an address like this:

> `http://0.0.0.0:6006` is the connection address.
### Docker Deployment
**Image and Port**
* Image: `stawky/chatglm2:latest`
* China mirror: `registry.cn-hangzhou.aliyuncs.com/fastgpt_docker/chatglm2:latest`
* Port: 6006
```
# Set the security token (used as the channel key in OneAPI)
Default: sk-aaabbbcccdddeeefffggghhhiiijjjkkk
You can also set it via the environment variable: sk-key. Refer to Docker documentation for how to pass environment variables.
```
## Connect to One API
Add a channel for chatglm2 with the following parameters:

Here, chatglm2 is used as the language model.
## Test
curl example:
```bash
curl --location --request POST 'https://domain/v1/chat/completions' \
--header 'Authorization: Bearer sk-aaabbbcccdddeeefffggghhhiiijjjkkk' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "chatglm2",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
Set Authorization to sk-aaabbbcccdddeeefffggghhhiiijjjkkk. The model field should match the custom model name you entered in One API.
## Integrate with FastGPT
Edit the config.json file and add chatglm2 to `llmModels`:
```json
"llmModels": [
// Existing models
{
"model": "chatglm2",
"name": "chatglm2",
"maxContext": 4000,
"maxResponse": 4000,
"quoteMaxToken": 2000,
"maxTemperature": 1,
"vision": false,
"defaultSystemChatPrompt": ""
}
]
```
## Usage
Simply select chatglm2 as the model.
file: ./content/self-host/custom-models/chatglm2.mdx
meta: {
"title": "接入 ChatGLM2-6B",
"description": " 将 FastGPT 接入私有化模型 ChatGLM2-6B"
}
import { Alert } from '@/components/docs/Alert';
## 前言
FastGPT 允许你使用自己的 OpenAI API KEY 来快速调用 OpenAI 接口,目前集成了 GPT-3.5, GPT-4 和 embedding,可构建自己的知识库。但考虑到数据安全的问题,我们并不能将所有的数据都交付给云端大模型。
那么如何在 FastGPT 上接入私有化模型呢?本文就以清华的 ChatGLM2 为例,为各位讲解如何在 FastGPT 中接入私有化模型。
## ChatGLM2-6B 简介
ChatGLM2-6B 是开源中英双语对话模型 ChatGLM-6B 的第二代版本,具体介绍可参阅 [ChatGLM2-6B 项目主页](https://github.com/THUDM/ChatGLM2-6B)。
注意,ChatGLM2-6B 权重对学术研究完全开放,在获得官方的书面许可后,亦允许商业使用。本教程只是介绍了一种用法,无权给予任何授权!
## 推荐配置
依据官方数据,同样是生成 8192 长度,量化等级为 FP16 要占用 12.8GB 显存、int8 为 8.1GB 显存、int4 为 5.1GB 显存,量化后会稍微影响性能,但不多。
因此推荐配置如下:
| 类型 | 内存 | 显存 | 硬盘空间 | 启动命令 |
| ---- | ------ | ------ | ------ | ------------------------ |
| fp16 | >=16GB | >=16GB | >=25GB | python openai\_api.py 16 |
| int8 | >=16GB | >=9GB | >=25GB | python openai\_api.py 8 |
| int4 | >=16GB | >=6GB | >=25GB | python openai\_api.py 4 |
## 部署
### 环境要求
* Python 3.8.10
* CUDA 11.8
* 科学上网环境
### 源码部署
1. 根据上面的环境配置配置好环境,具体教程自行 GPT;
2. 下载 [python 文件](https://github.com/labring/FastGPT/blob/main/plugins/model/llm-ChatGLM2/openai_api.py)
3. 在命令行输入命令 `pip install -r requirements.txt`;
4. 打开你需要启动的 py 文件,在代码的 `verify_token` 方法中配置 token,这里的 token 只是加一层验证,防止接口被人盗用;
5. 执行命令 `python openai_api.py --model_name 16`。这里的数字根据上面的配置进行选择。
然后等待模型下载,直到模型加载完毕为止。如果出现报错先问 GPT。
启动成功后应该会显示如下地址:

> 这里的 `http://0.0.0.0:6006` 就是连接地址。
### docker 部署
**镜像和端口**
* 镜像名: `stawky/chatglm2:latest`
* 国内镜像名: `registry.cn-hangzhou.aliyuncs.com/fastgpt_docker/chatglm2:latest`
* 端口号: 6006
```
# 设置安全凭证(即oneapi中的渠道密钥)
默认值:sk-aaabbbcccdddeeefffggghhhiiijjjkkk
也可以通过环境变量引入:sk-key。有关docker环境变量引入的方法请自寻教程,此处不再赘述。
```
## 接入 One API
为 chatglm2 添加一个渠道,参数如下:

这里我填入 chatglm2 作为语言模型
## 测试
curl 例子:
```bash
curl --location --request POST 'https://domain/v1/chat/completions' \
--header 'Authorization: Bearer sk-aaabbbcccdddeeefffggghhhiiijjjkkk' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "chatglm2",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
Authorization 为 sk-aaabbbcccdddeeefffggghhhiiijjjkkk。model 为刚刚在 One API 填写的自定义模型。
## 接入 FastGPT
修改 config.json 配置文件,在 llmModels 中加入 chatglm2 模型:
```json
"llmModels": [
//已有模型
{
"model": "chatglm2",
"name": "chatglm2",
"maxContext": 4000,
"maxResponse": 4000,
"quoteMaxToken": 2000,
"maxTemperature": 1,
"vision": false,
"defaultSystemChatPrompt": ""
}
]
```
## 测试使用
chatglm2 模型的使用方法如下:
模型选择 chatglm2 即可
file: ./content/self-host/custom-models/m3e.en.mdx
meta: {
"title": "Integrating M3E Embedding Model",
"description": "Integrating the private M3E embedding model with FastGPT"
}
## Introduction
FastGPT uses OpenAI's embedding model by default. For private deployment, you can replace it with the M3E embedding model. M3E is a lightweight model with low resource requirements -- it can even run on CPU. The following tutorial is based on an image provided by community contributor "睡大觉".
## Deploy the Image
Image: `stawky/m3e-large-api:latest`
China mirror: `registry.cn-hangzhou.aliyuncs.com/fastgpt_docker/m3e-large-api:latest`
Port: 6008
Environment variables:
```
# Set the security token (used as the channel key in OneAPI)
Default: sk-aaabbbcccdddeeefffggghhhiiijjjkkk
You can also set it via the environment variable: sk-key. Refer to Docker documentation for how to pass environment variables.
```
## Connect to One API
Add a channel with the following parameters:

## Test
curl example:
```bash
curl --location --request POST 'https://domain/v1/embeddings' \
--header 'Authorization: Bearer xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "m3e",
"input": ["What is FastGPT"]
}'
```
Set Authorization to your sk-key. The model field should match the custom model name you entered in One API.
## Integrate with FastGPT
Edit the config.json file and add the M3E model to `vectorModels`:
```json
"vectorModels": [
{
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"price": 0.2,
"defaultToken": 500,
"maxToken": 3000
},
{
"model": "m3e",
"name": "M3E (for testing)",
"price": 0.1,
"defaultToken": 500,
"maxToken": 1800
}
]
```
## Usage
1. Select the M3E model when creating a Knowledge Base.
Note: once selected, the embedding model for the Knowledge Base cannot be changed.

2. Import data
3. Test search

4. Bind the Knowledge Base to an app
Note: an app can only bind Knowledge Bases that use the same embedding model -- cross-model binding is not supported. You may also need to adjust the similarity threshold, as different embedding models produce different similarity (distance) scores. Test and tune accordingly.

file: ./content/self-host/custom-models/m3e.mdx
meta: {
"title": "接入 M3E 向量模型",
"description": " 将 FastGPT 接入私有化模型 M3E"
}
## 前言
FastGPT 默认使用了 openai 的 embedding 向量模型,如果你想私有部署的话,可以使用 M3E 向量模型进行替换。M3E 向量模型属于小模型,资源使用不高,CPU 也可以运行。下面教程是基于 “睡大觉” 同学提供的一个的镜像。
## 部署镜像
镜像名: `stawky/m3e-large-api:latest`\
国内镜像: `registry.cn-hangzhou.aliyuncs.com/fastgpt_docker/m3e-large-api:latest`
端口号: 6008
环境变量:
```
# 设置安全凭证(即oneapi中的渠道密钥)
默认值:sk-aaabbbcccdddeeefffggghhhiiijjjkkk
也可以通过环境变量引入:sk-key。有关docker环境变量引入的方法请自寻教程,此处不再赘述。
```
## 接入 One API
添加一个渠道,参数如下:

## 测试
curl 例子:
```bash
curl --location --request POST 'https://domain/v1/embeddings' \
--header 'Authorization: Bearer xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "m3e",
"input": ["laf是什么"]
}'
```
Authorization 为 sk-key。model 为刚刚在 One API 填写的自定义模型。
## 接入 FastGPT
修改 config.json 配置文件,在 vectorModels 中加入 M3E 模型:
```json
"vectorModels": [
{
"model": "text-embedding-ada-002",
"name": "Embedding-2",
"price": 0.2,
"defaultToken": 500,
"maxToken": 3000
},
{
"model": "m3e",
"name": "M3E(测试使用)",
"price": 0.1,
"defaultToken": 500,
"maxToken": 1800
}
]
```
## 测试使用
1. 创建知识库时候选择 M3E 模型。
注意,一旦选择后,知识库将无法修改向量模型。

2. 导入数据
3. 搜索测试

4. 应用绑定知识库
注意,应用只能绑定同一个向量模型的知识库,不能跨模型绑定。并且,需要注意调整相似度,不同向量模型的相似度(距离)会有所区别,需要自行测试实验。

file: ./content/self-host/custom-models/marker.en.mdx
meta: {
"title": "Integrating Marker PDF Parsing",
"description": "Use Marker to parse PDF documents with image extraction and layout recognition"
}
## Background
PDF is a relatively complex file format. FastGPT's built-in PDF parser relies on the pdfjs library, which uses logical parsing and cannot effectively handle complex PDF files. When parsing PDFs containing images, tables, formulas, or other non-plain-text content, the results are often poor.
There are several PDF parsing solutions available. [Marker](https://github.com/VikParuchuri/marker) uses the Surya model for vision-based parsing, effectively extracting images, tables, formulas, and other complex content.
Starting from `FastGPT v4.9.0`, community edition users can add the `systemEnv.customPdfParse` configuration in `config.json` to use Marker for PDF parsing. Commercial edition users can configure this directly in the Admin panel via the form. You need to pull the latest Marker image, as the API format has changed.
## Tutorial
### 1. Install Marker
Refer to the [Marker installation guide](https://github.com/labring/FastGPT/tree/main/plugins/model/pdf-marker) to install the Marker model. The bundled API is already compatible with FastGPT's custom parsing service.
Quick Docker installation:
```dockerfile
docker pull crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.2
docker run --gpus all -itd -p 7231:7232 --name model_pdf_v2 -e PROCESSES_PER_GPU="2" crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.2
```
### 2. Add FastGPT Configuration
```json
{
xxx
"systemEnv": {
xxx
"customPdfParse": {
"url": "http://xxxx.com/v2/parse/file", // Custom PDF parsing service URL for Marker v0.2
"key": "", // Custom PDF parsing service key
"doc2xKey": "", // doc2x service key
"price": 0 // PDF parsing service price
}
}
}
```
Restart the service after making changes.
### 3. Test
Upload a PDF file through the Knowledge Base and enable the `Enhanced PDF Parsing` option.

After uploading, you should see the following logs (LOG\_LEVEL must be set to info or debug):
```
[Info] 2024-12-05 15:04:42 Parsing files from an external service
[Info] 2024-12-05 15:07:08 Custom file parsing is complete, time: 1316ms
```
You'll notice that PDFs parsed by Marker include image links:

Similarly, in apps you can enable `Enhanced PDF Parsing` in the file upload settings.

## Results
Using Tsinghua's [ChatDev Communicative Agents for Software Develop.pdf](https://arxiv.org/abs/2307.07924) as an example:
| | | |
| ---------------------------------------------- | ---------------------------------------------- | ---------------------------------------------- |
|  |  |  |
|  |  |  |
The top row shows chunked results; the bottom row shows the original PDF. Images, formulas, and tables are all extracted effectively.
Note that [Marker](https://github.com/VikParuchuri/marker) is licensed under `GPL-3.0 license`. Please ensure compliance with the license when using it.
## Legacy Marker Usage
For FastGPT versions before V4.9.0, you can use the following method for Marker parsing.
Install and run the Marker service:
```dockerfile
docker pull crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.1
docker run --gpus all -itd -p 7231:7231 --name model_pdf_v1 -e PROCESSES_PER_GPU="2" crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.1
```
Then modify the FastGPT environment variables:
```
CUSTOM_READ_FILE_URL=http://xxxx.com/v1/parse/file
CUSTOM_READ_FILE_EXTENSION=pdf
```
* CUSTOM\_READ\_FILE\_URL - The custom parsing service URL. Replace the host with your parsing service address; the path must remain unchanged.
* CUSTOM\_READ\_FILE\_EXTENSION - Supported file extensions. Use commas to separate multiple file types.
file: ./content/self-host/custom-models/marker.mdx
meta: {
"title": "接入 Marker PDF 文档解析",
"description": "使用 Marker 解析 PDF 文档,可实现图片提取和布局识别"
}
## 背景
PDF 是一个相对复杂的文件格式,在 FastGPT 内置的 pdf 解析器中,依赖的是 pdfjs 库解析,该库基于逻辑解析,无法有效的理解复杂的 pdf 文件。所以我们在解析 pdf 时候,如果遇到图片、表格、公式等非简单文本内容,会发现解析效果不佳。
市面上目前有多种解析 PDF 的方法,比如使用 [Marker](https://github.com/VikParuchuri/marker),该项目使用了 Surya 模型,基于视觉解析,可以有效提取图片、表格、公式等复杂内容。
在 `FastGPT v4.9.0` 版本中,社区版用户可以在`config.json`文件中添加`systemEnv.customPdfParse`配置,来使用 Marker 解析 PDF 文件。商业版用户直接在 Admin 后台根据表单指引填写即可。需重新拉取 Marker 镜像,接口格式已变动。
## 使用教程
### 1. 安装 Marker
参考文档 [Marker 安装教程](https://github.com/labring/FastGPT/tree/main/plugins/model/pdf-marker),安装 Marker 模型。封装的 API 已经适配了 FastGPT 自定义解析服务。
这里介绍快速 Docker 安装的方法:
```dockerfile
docker pull crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.2
docker run --gpus all -itd -p 7231:7232 --name model_pdf_v2 -e PROCESSES_PER_GPU="2" crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.2
```
### 2. 添加 FastGPT 文件配置
```json
{
xxx
"systemEnv": {
xxx
"customPdfParse": {
"url": "http://xxxx.com/v2/parse/file", // 自定义 PDF 解析服务地址 marker v0.2
"key": "", // 自定义 PDF 解析服务密钥
"doc2xKey": "", // doc2x 服务密钥
"price": 0 // PDF 解析服务价格
}
}
}
```
需要重启服务。
### 3. 测试效果
通过知识库上传一个 pdf 文件,并勾选上 `PDF 增强解析`。

确认上传后,可以在日志中看到 LOG (LOG\_LEVEL需要设置 info 或者 debug):
```
[Info] 2024-12-05 15:04:42 Parsing files from an external service
[Info] 2024-12-05 15:07:08 Custom file parsing is complete, time: 1316ms
```
然后你就可以发现,通过 Marker 解析出来的 pdf 会携带图片链接:

同样的,在应用中,你可以在文件上传配置里,勾选上 `PDF 增强解析`。

## 效果展示
以清华的 [ChatDev Communicative Agents for Software Develop.pdf](https://arxiv.org/abs/2307.07924) 为例,展示 Marker 解析的效果:
| | | |
| ---------------------------------------------- | ---------------------------------------------- | ---------------------------------------------- |
|  |  |  |
|  |  |  |
上图是分块后的结果,下图是 pdf 原文。整体图片、公式、表格都可以提取出来,效果还是杠杠的。
不过要注意的是,[Marker](https://github.com/VikParuchuri/marker) 的协议是`GPL-3.0 license`,请在遵守协议的前提下使用。
## 旧版 Marker 使用方法
FastGPT V4.9.0 版本之前,可以用以下方式,试用 Marker 解析服务。
安装和运行 Marker 服务:
```dockerfile
docker pull crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.1
docker run --gpus all -itd -p 7231:7231 --name model_pdf_v1 -e PROCESSES_PER_GPU="2" crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/marker11/marker_images:v0.1
```
并修改 FastGPT 环境变量:
```
CUSTOM_READ_FILE_URL=http://xxxx.com/v1/parse/file
CUSTOM_READ_FILE_EXTENSION=pdf
```
* CUSTOM\_READ\_FILE\_URL - 自定义解析服务的地址, host改成解析服务的访问地址,path 不能变动。
* CUSTOM\_READ\_FILE\_EXTENSION - 支持的文件后缀,多个文件类型,可用逗号隔开。
file: ./content/self-host/custom-models/mineru.en.mdx
meta: {
"title": "Integrating MinerU PDF Parsing",
"description": "Use MinerU to parse PDF documents with image extraction, layout recognition, table recognition, and formula recognition"
}
## Background
PDF is a relatively complex file format. FastGPT's built-in PDF parser relies on the pdfjs library, which uses logical parsing and cannot effectively handle complex PDF files. When parsing PDFs containing images, tables, formulas, or other non-plain-text content, the results are often poor.
There are several PDF parsing solutions available. [MinerU](https://github.com/opendatalab/MinerU) uses YOLO, PaddleOCR, and table recognition models for vision-based parsing, effectively extracting images, tables, formulas, and other complex content.
Community edition users can add the `systemEnv.customPdfParse` configuration in `config.json` to use MinerU for PDF parsing. Commercial edition users can configure this directly in the Admin panel via the form -- details are covered in the tutorial below.
## Tutorial
Hardware requirements: 16GB+ GPU VRAM, minimum 16GB+ RAM (32GB+ recommended). See the [official page](https://github.com/opendatalab/MinerU) for other requirements.
### 1. Install MinerU
Quick Docker installation:
Pull the fastgpt-mineru image --> Create and start the parsing service container --> Add the deployed URL to the FastGPT configuration file
```dockerfile
docker pull crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/fastgpt_ck/mineru:v1
docker run --gpus all -itd -p 7231:8001 --name mode_pdf_minerU crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/fastgpt_ck/mineru:v1
```
This MinerU integration uses pipeline mode with built-in parallelization inside the Docker container. It creates multiple processes based on the number of GPUs to handle uploaded PDFs concurrently.
### 2. Add FastGPT Configuration
```json
{
xxx
"systemEnv": {
xxx
"customPdfParse": {
"url": "http://xxxx.com/v2/parse/file", // Custom PDF parsing service URL for MinerU
"key": "", // Custom PDF parsing service key
"doc2xKey": "", // doc2x service key
"price": 0 // PDF parsing service price
}
}
}
```
For the commercial edition, configure as shown below:

**Note:** Services added via the configuration file require a restart to take effect.
### 3. Test
Upload a PDF file through the Knowledge Base and enable the `Enhanced PDF Parsing` option.

After uploading, you should see the following logs (LOG\_LEVEL must be set to info or debug):
```
[Info] 2024-12-05 15:04:42 Parsing files from an external service
[Info] 2024-12-05 15:07:08 Custom file parsing is complete, time: 1316ms
```
Similarly, in apps you can enable `Enhanced PDF Parsing` in the file upload settings.

## Results
Using Tsinghua's [ChatDev Communicative Agents for Software Develop.pdf](https://arxiv.org/abs/2307.07924) as an example:
| | | |
| ----------------------------------------------- | ----------------------------------------------- | ----------------------------------------------- |
|  |  |  |
|  |  |  |
The top row shows chunked results; the bottom row shows the original PDF. Images, formulas, and OCR handwriting are all extracted effectively.
Note that [MinerU](https://github.com/opendatalab/MinerU) is licensed under `GPL-3.0 license`. Please ensure compliance with the license when using it.
file: ./content/self-host/custom-models/mineru.mdx
meta: {
"title": "接入 MinerU PDF 文档解析",
"description": "使用 MinerU 解析 PDF 文档,可实现图片提取、布局识别、表格识别和公式识别"
}
## 背景
PDF 是一个相对复杂的文件格式,在 FastGPT 内置的 pdf 解析器中,依赖的是 pdfjs 库解析,该库基于逻辑解析,无法有效的理解复杂的 pdf 文件。所以我们在解析 pdf 时候,如果遇到图片、表格、公式等非简单文本内容,会发现解析效果不佳。
市面上目前有多种解析 PDF 的方法,比如使用 [MinerU](https://github.com/opendatalab/MinerU),该项目使用了 YOLO、PaddleOCR以及表格识别等模型,基于视觉解析,可以有效提取图片、表格、公式等复杂内容。
社区版用户可以在`config.json`文件中添加`systemEnv.customPdfParse`配置,来使用 MinerU 解析 PDF 文件。商业版用户直接在 Admin 后台根据表单指引填写即可,使用教程中会详细解释。
## 使用教程
硬件需求:16g+ 的gpu显存推理卡,最小 16GB+, 推荐 32GB+的内存,其他要求查看[官网](https://github.com/opendatalab/MinerU)
### 1. 安装 MinerU
这里介绍快速 Docker 安装的方法:
拉取fastgpt-mineru镜像 ---> 创建容器启动解析服务 ---> 把部署好的url地址接入到fastgpt配置文件中
```dockerfile
docker pull crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/fastgpt_ck/mineru:v1
docker run --gpus all -itd -p 7231:8001 --name mode_pdf_minerU crpi-h3snc261q1dosroc.cn-hangzhou.personal.cr.aliyuncs.com/fastgpt_ck/mineru:v1
```
这里的mineru接入的是pipeline模式,并且在docker内部进行了并行化,会根据gpu数量创建多个进程来同时处理上传的pdf数据
### 2. 添加 FastGPT 文件配置
```json
{
xxx
"systemEnv": {
xxx
"customPdfParse": {
"url": "http://xxxx.com/v2/parse/file", // 自定义 PDF 解析服务地址 MinerU
"key": "", // 自定义 PDF 解析服务密钥
"doc2xKey": "", // doc2x 服务密钥
"price": 0 // PDF 解析服务价格
}
}
}
```
商业版请按下图配置

**注意:** 通过配置文件添加的服务需要重启服务。
### 3. 测试效果
通过知识库上传一个 pdf 文件,并勾选上 `PDF 增强解析`。

确认上传后,可以在日志中看到 LOG (LOG\_LEVEL需要设置 info 或者 debug):
```
[Info] 2024-12-05 15:04:42 Parsing files from an external service
[Info] 2024-12-05 15:07:08 Custom file parsing is complete, time: 1316ms
```
同样的,在应用中,你可以在文件上传配置里,勾选上 `PDF 增强解析`。

## 效果展示
以清华的 [ChatDev Communicative Agents for Software Develop.pdf](https://arxiv.org/abs/2307.07924) 为例,展示 MinerU 解析的效果:
| | | |
| ----------------------------------------------- | ----------------------------------------------- | ----------------------------------------------- |
|  |  |  |
|  |  |  |
上图是分块后的结果,下图是 pdf 原文。整体图片、公式、ocr手写体都可以提取出来,效果还是可以的。
不过要注意的是,[MinerU](https://github.com/opendatalab/MinerU) 的协议是`GPL-3.0 license`,请在遵守协议的前提下使用。
file: ./content/self-host/custom-models/ollama.en.mdx
meta: {
"title": "Integrating Local Models with Ollama",
"description": "Deploy your own models using Ollama"
}
[Ollama](https://ollama.com/) is an open-source AI model deployment tool focused on simplifying the deployment and usage of large language models. It supports one-click download and running of various LLMs.
## Installing Ollama
Ollama supports multiple installation methods, but Docker is recommended. If you install Ollama directly on your host machine, you'll need to figure out how to let the FastGPT Docker container access Ollama on the host, which can be tricky.
### Docker Installation (Recommended)
Use Ollama's official Docker image for one-click installation and startup (make sure Docker is installed on your machine):
```bash
docker pull ollama/ollama
docker run --rm -d --name ollama -p 11434:11434 ollama/ollama
```
If your FastGPT is deployed in Docker, make sure the Ollama container is on the same network as FastGPT. Otherwise, FastGPT may not be able to access it:
```bash
docker run --rm -d --name ollama --network (your FastGPT container network) -p 11434:11434 ollama/ollama
```
### Host Installation
If you prefer not to use Docker, you can install directly on the host machine.
#### MacOS
If you're on macOS with Homebrew installed:
```bash
brew install ollama
ollama serve # Start the service after installation
```
#### Linux
On Linux, you can use a package manager. For Ubuntu:
```bash
curl https://ollama.com/install.sh | sh # Downloads and runs the official install script
ollama serve # Start the service after installation
```
#### Windows
On Windows, download the installer from the Ollama official website. Run the installer and follow the wizard. After installation, start the service in Command Prompt or PowerShell:
```bash
ollama serve # After installation, visit http://localhost:11434 in your browser to verify Ollama is running
```
#### Additional Notes
If you installed Ollama as a host application (not via Docker), make sure Ollama listens on 0.0.0.0.
##### 1. Linux
If Ollama runs as a systemd service, edit the service file with `sudo systemctl edit ollama.service`. Add `Environment="OLLAMA_HOST=0.0.0.0"` under the \[Service] section. Save and exit, then run `sudo systemctl daemon-reload` and `sudo systemctl restart ollama` to apply.
##### 2. MacOS
Open a terminal and run `launchctl setenv ollama_host "0.0.0.0"`, then restart the Ollama application.
##### 3. Windows
Open "Edit system environment variables" from the Start menu or search bar. In "System Properties", click "Environment Variables". Under "System variables", click "New" and create a variable named OLLAMA\_HOST with value 0.0.0.0. Click "OK" to save, then restart Ollama from the Start menu.
### Pull Model Images
After installing Ollama, no models are available locally -- you need to pull them:
```bash
# For Docker deployment, enter the container first: docker exec -it [Ollama container name] /bin/sh
ollama pull [model name]
```

### Test Communication
After installation, verify connectivity by entering the FastGPT container and trying to reach Ollama:
```bash
docker exec -it [FastGPT container name] /bin/sh
curl http://XXX.XXX.XXX.XXX:11434 # Container: "http://[container name]:[port]", Host: "http://[host IP]:[port]" (host IP cannot be localhost)
```
If you see that the Ollama service is running, communication is working.
## Integrating Ollama with FastGPT
### 1. Check Available Models
First, check which models Ollama has:
```bash
# For Docker-deployed Ollama: docker exec -it [Ollama container name] /bin/sh
ollama ls
```

### 2. AI Proxy Integration
If you're using FastGPT's default configuration from [here](../deploy/docker.en.mdx), AI Proxy is enabled by default.

Make sure your FastGPT can access the Ollama container. If not, refer to the [installation section](#installing-ollama) above -- check whether the host isn't listening on 0.0.0.0 or the containers aren't on the same network.

In FastGPT, go to Account -> Model Providers -> Model Configuration -> Add Model. Make sure the model ID matches the model name in OneAPI. See details [here](../config/model/intro.en.mdx).


Run FastGPT, then go to Account -> Model Providers -> Model Channels -> Add Channel. Select Ollama as the channel type, add your pulled model, and fill in the proxy address. For container-deployed Ollama, the address is [http://address:port](http://address:port). Note: container deployment uses "http\://\[container name]:\[port]", host installation uses "http\://\[host IP]:\[port]" (host IP cannot be localhost).

Create an app in the workspace and select the model you added. The model name shown is the alias you set. Note: the same model cannot be added multiple times -- the system uses the alias from the most recent addition.

### 3. OneAPI Integration
If you want to use OneAPI, pull the OneAPI image and run it on the same network as FastGPT:
```bash
# Pull the OneAPI image
docker pull intel/oneapi-hpckit
# Run the container on the FastGPT network
docker run -it --network [FastGPT network] --name container_name intel/oneapi-hpckit /bin/bash
```
In the OneAPI page, add a new channel with type Ollama. Enter your Ollama model name (must match exactly), then fill in the Ollama proxy address below -- default is [http://address:port](http://address:port), without /v1. Test the channel after adding. This example uses Docker-deployed Ollama; for host-installed Ollama, use http\://\[host IP]:\[port].

After adding the channel, click Token -> Add Token, fill in the name, and configure as needed.

Edit the FastGPT docker-compose.yml file: comment out AI Proxy, set OPENAI\_BASE\_URL to your OneAPI address (default [http://address:port/v1](http://address:port/v1) -- /v1 is required), and set KEY to your OneAPI token.

Then [jump to section 5](#5-model-addition-and-usage) to add and use models.
### 4. Direct Integration
If you don't want to use AI Proxy or OneAPI, you can connect directly. Edit the FastGPT docker-compose.yml: comment out AI Proxy code, set OPENAI\_BASE\_URL to your Ollama address (default [http://address:port/v1](http://address:port/v1) -- /v1 is required), and set KEY to any value (Ollama has no authentication by default; if you've enabled it, use the correct key). Everything else is the same as the OneAPI approach -- just add your model in FastGPT. This example uses Docker-deployed Ollama; for host-installed Ollama, use http\://\[host IP]:\[port].

After completing the setup, [click here](#5-model-addition-and-usage) to add and use models.
### 5. Model Addition and Usage
In FastGPT, go to Account -> Model Providers -> Model Configuration -> Add Model. Make sure the model ID matches the model name in OneAPI.


Create an app in the workspace and select the model you added. The model name shown is the alias you set. Note: the same model cannot be added multiple times -- the system uses the alias from the most recent addition.

### 6. Additional Notes
For the Ollama proxy addresses above: host-installed Ollama uses "http\://\[host IP]:\[port]", container-deployed Ollama uses "http\://\[container name]:\[port]".
file: ./content/self-host/custom-models/ollama.mdx
meta: {
"title": "使用 Ollama 接入本地模型 ",
"description": " 采用 Ollama 部署自己的模型"
}
[Ollama](https://ollama.com/) 是一个开源的AI大模型部署工具,专注于简化大语言模型的部署和使用,支持一键下载和运行各种大模型。
## 安装 Ollama
Ollama 本身支持多种安装方式,但是推荐使用 Docker 拉取镜像部署。如果是个人设备上安装了 Ollama 后续需要解决如何让 Docker 中 FastGPT 容器访问宿主机 Ollama的问题,较为麻烦。
### Docker 安装(推荐)
你可以使用 Ollama 官方的 Docker 镜像来一键安装和启动 Ollama 服务(确保你的机器上已经安装了 Docker),命令如下:
```bash
docker pull ollama/ollama
docker run --rm -d --name ollama -p 11434:11434 ollama/ollama
```
如果你的 FastGPT 是在 Docker 中进行部署的,建议在拉取 Ollama 镜像时保证和 FastGPT 镜像处于同一网络,否则可能出现 FastGPT 无法访问的问题,命令如下:
```bash
docker run --rm -d --name ollama --network (你的 Fastgpt 容器所在网络) -p 11434:11434 ollama/ollama
```
### 主机安装
如果你不想使用 Docker ,也可以采用主机安装,以下是主机安装的一些方式。
#### MacOS
如果你使用的是 macOS,且系统中已经安装了 Homebrew 包管理器,可通过以下命令来安装 Ollama:
```bash
brew install ollama
ollama serve #安装完成后,使用该命令启动服务
```
#### Linux
在 Linux 系统上,你可以借助包管理器来安装 Ollama。以 Ubuntu 为例,在终端执行以下命令:
```bash
curl https://ollama.com/install.sh | sh #此命令会从官方网站下载并执行安装脚本。
ollama serve #安装完成后,同样启动服务
```
#### Windows
在 Windows 系统中,你可以从 Ollama 官方网站 下载 Windows 版本的安装程序。下载完成后,运行安装程序,按照安装向导的提示完成安装。安装完成后,在命令提示符或 PowerShell 中启动服务:
```bash
ollama serve #安装完成并启动服务后,你可以在浏览器中访问 http://localhost:11434 来验证 Ollama 是否安装成功。
```
#### 补充说明
如果你是采用的主机应用 Ollama 而不是镜像,需要确保你的 Ollama 可以监听0.0.0.0。
##### 1. Linxu 系统
如果 Ollama 作为 systemd 服务运行,打开终端,编辑 Ollama 的 systemd 服务文件,使用命令sudo systemctl edit ollama.service,在\[Service]部分添加Environment="OLLAMA\_HOST=0.0.0.0"。保存并退出编辑器,然后执行sudo systemctl daemon - reload和sudo systemctl restart ollama使配置生效。
##### 2. MacOS 系统
打开终端,使用launchctl setenv ollama\_host "0.0.0.0"命令设置环境变量,然后重启 Ollama 应用程序以使更改生效。
##### 3. Windows 系统
通过 “开始” 菜单或搜索栏打开 “编辑系统环境变量”,在 “系统属性” 窗口中点击 “环境变量”,在 “系统变量” 部分点击 “新建”,创建一个名为OLLAMA\_HOST的变量,变量值设置为0.0.0.0,点击 “确定” 保存更改,最后从 “开始” 菜单重启 Ollama 应用程序。
### Ollama 拉取模型镜像
在安装 Ollama 后,本地是没有模型镜像的,需要自己去拉取 Ollama 中的模型镜像。命令如下:
```bash
# Docker 部署需要先进容器,命令为: docker exec -it [ Ollama 容器名 ] /bin/sh
ollama pull [模型名]
```

### 测试通信
在安装完成后,需要进行检测测试,首先进入 FastGPT 所在的容器,尝试访问自己的 Ollama ,命令如下:
```bash
docker exec -it [ FastGPT 所在的容器名 ] /bin/sh
curl http://XXX.XXX.XXX.XXX:11434 #容器部署地址为“http://[容器名]:[端口]”,主机安装地址为"http://[主机IP]:[端口]",主机IP不可为localhost
```
看到访问显示自己的 Ollama 服务以及启动,说明可以正常通信。
## 将 Ollama 接入 FastGPT
### 1. 查看 Ollama 所拥有的模型
首先采用下述命令查看 Ollama 中所拥有的模型,
```bash
# Docker 部署 Ollama,需要此命令 docker exec -it [ Ollama 容器名 ] /bin/sh
ollama ls
```

### 2. AI Proxy 接入
如果你采用的是 FastGPT 中的默认配置文件部署[这里](../deploy/docker.mdx),即默认采用 AI Proxy 进行启动。

以及在确保你的 FastGPT 可以直接访问 Ollama 容器的情况下,无法访问,参考上文[点此跳转](#安装-ollama)的安装过程,检测是不是主机不能监测0.0.0.0,或者容器不在同一个网络。

在 FastGPT 中点击账号->模型提供商->模型配置->新增模型,添加自己的模型即可,添加模型时需要保证模型ID和 OneAPI 中的模型名称一致。详细参考[这里](../config/model/intro.mdx)


运行 FastGPT ,在页面中选择账号->模型提供商->模型渠道->新增渠道。之后,在渠道选择中选择 Ollama ,然后加入自己拉取的模型,填入代理地址,如果是容器中安装 Ollama ,代理地址为[http://地址:端口,补充:容器部署地址为“http://\[容器名\]:\[端口\]”,主机安装地址为"http://\[主机IP\]:\[端口\]",主机IP不可为localhost](http://地址:端口,补充:容器部署地址为“http://\[容器名]:\[端口]”,主机安装地址为"http://\[主机IP]:\[端口]",主机IP不可为localhost)

在工作台中创建一个应用,选择自己之前添加的模型,此处模型名称为自己当时设置的别名。注:同一个模型无法多次添加,系统会采取最新添加时设置的别名。

### 3. OneAPI 接入
如果你想使用 OneAPI ,首先需要拉取 OneAPI 镜像,然后将其在 FastGPT 容器的网络中运行。具体命令如下:
```bash
# 拉取 oneAPI 镜像
docker pull intel/oneapi-hpckit
# 运行容器并指定自定义网络和容器名
docker run -it --network [ FastGPT 网络 ] --name 容器名 intel/oneapi-hpckit /bin/bash
```
进入 OneAPI 页面,添加新的渠道,类型选择 Ollama ,在模型中填入自己 Ollama 中的模型,需要保证添加的模型名称和 Ollama 中一致,再在下方填入自己的 Ollama 代理地址,默认[http://地址:端口,不需要填写/v1。添加成功后在](http://地址:端口,不需要填写/v1。添加成功后在) OneAPI 进行渠道测试,测试成功则说明添加成功。此处演示采用的是 Docker 部署 Ollama 的效果,主机 Ollama需要修改代理地址为http\://\[主机IP]:\[端口]

渠道添加成功后,点击令牌,点击添加令牌,填写名称,修改配置。

修改部署 FastGPT 的 docker-compose.yml 文件,在其中将 AI Proxy 的使用注释,在 OPENAI\_BASE\_URL 中加入自己的 OneAPI 开放地址,默认是[http://地址:端口/v1,v1必须填写。KEY](http://地址:端口/v1,v1必须填写。KEY) 中填写自己在 OneAPI 的令牌。

[直接跳转5](#5-模型添加和使用)添加模型,并使用。
### 4. 直接接入
如果你既不想使用 AI Proxy,也不想使用 OneAPI,也可以选择直接接入,修改部署 FastGPT 的 docker-compose.yml 文件,在其中将 AI Proxy 的使用注释,采用和 OneAPI 的类似配置。注释掉 AIProxy 相关代码,在OPENAI\_BASE\_URL中加入自己的 Ollama 开放地址,默认是[http://地址:端口/v1,强调:v1必须填写。在KEY中随便填入,因为](http://地址:端口/v1,强调:v1必须填写。在KEY中随便填入,因为) Ollama 默认没有鉴权,如果开启鉴权,请自行填写。其他操作和在 OneAPI 中加入 Ollama 一致,只需在 FastGPT 中加入自己的模型即可使用。此处演示采用的是 Docker 部署 Ollama 的效果,主机 Ollama需要修改代理地址为http\://\[主机IP]:\[端口]

完成后[点击这里](#5-模型添加和使用)进行模型添加并使用。
### 5. 模型添加和使用
在 FastGPT 中点击账号->模型提供商->模型配置->新增模型,添加自己的模型即可,添加模型时需要保证模型ID和 OneAPI 中的模型名称一致。


在工作台中创建一个应用,选择自己之前添加的模型,此处模型名称为自己当时设置的别名。注:同一个模型无法多次添加,系统会采取最新添加时设置的别名。

### 6. 补充
上述接入 Ollama 的代理地址中,主机安装 Ollama 的地址为“http\://\[主机IP]:\[端口]”,容器部署 Ollama 地址为“http\://\[容器名]:\[端口]”
file: ./content/self-host/custom-models/xinference.en.mdx
meta: {
"title": "Integrating Local Models with Xinference",
"description": "One-stop local LLM private deployment"
}
[Xinference](https://github.com/xorbitsai/inference) is an open-source model inference platform. Beyond LLMs, it can also deploy Embedding and ReRank models, which are critical for enterprise-grade RAG. Xinference also provides advanced features like Function Calling and supports distributed deployment, meaning it can scale horizontally as your application usage grows.
## Installing Xinference
Xinference supports multiple inference engines as backends for different deployment scenarios. Below we introduce these backends by use case.
### 1. Server
If you're deploying LLMs on a Linux or Windows server, you can choose Transformers or vLLM as Xinference's inference backend:
* [Transformers](https://huggingface.co/docs/transformers/index): By integrating Hugging Face's Transformers library, Xinference can quickly adopt the most cutting-edge NLP models, including LLMs.
* [vLLM](https://vllm.ai/): An open-source library developed by UC Berkeley for efficiently serving LLMs. It introduces the PagedAttention algorithm for improved memory management of attention keys and values. Throughput can reach 24x that of Transformers, making vLLM suitable for production environments with high-concurrency access.
If your server has an NVIDIA GPU, refer to [this article for CUDA installation instructions](https://xorbits.cn/blogs/langchain-streamlit-doc-chat) to maximize GPU acceleration with Xinference.
#### Docker Deployment
Use Xinference's official Docker image for one-click installation and startup (make sure Docker is installed):
```bash
docker run -p 9997:9997 --gpus all xprobe/xinference:latest xinference-local -H 0.0.0.0
```
#### Direct Deployment
First, prepare a Python 3.9+ environment. We recommend installing conda first, then creating a Python 3.11 environment:
```bash
conda create --name py311 python=3.11
conda activate py311
```
Install Xinference with Transformers and vLLM as inference backends:
```bash
pip install "xinference[transformers]"
pip install "xinference[vllm]"
pip install "xinference[transformers,vllm]" # Install both
```
PyPI automatically installs PyTorch with Transformers and vLLM, but the auto-installed CUDA version may not match your environment. If so, manually install per PyTorch's [installation guide](https://pytorch.org/get-started/locally/).
Start the Xinference service:
```bash
xinference-local -H 0.0.0.0
```
Xinference starts locally on port 9997 by default. With the `-H 0.0.0.0` parameter, non-local clients can access the service via the machine's IP address.
### 2. Personal Devices
To deploy LLMs on your MacBook or personal computer, we recommend CTransformers as Xinference's inference backend. CTransformers is a C++ implementation of Transformers using GGML.
[GGML](https://ggml.ai/) is a C++ library that enables LLMs to [run on consumer hardware](https://github.com/ggerganov/llama.cpp/discussions/205). Its key feature is model quantization -- reducing weight precision to lower resource requirements. For example, representing a high-precision float (like 0.0001) requires more space than a low-precision one (like 0.1). Since LLMs must be loaded into memory for inference, you need sufficient disk space for storage and enough RAM for execution. GGML supports many quantization strategies, each offering different efficiency-performance trade-offs.
Install CTransformers as Xinference's backend:
```bash
pip install xinference
pip install ctransformers
```
Since GGML is a C++ library, Xinference uses `llama-cpp-python` for language bindings. Different hardware platforms require different compilation parameters:
* Apple Metal (MPS): `CMAKE_ARGS="-DLLAMA_METAL=on" pip install llama-cpp-python`
* Nvidia GPU: `CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install llama-cpp-python`
* AMD GPU: `CMAKE_ARGS="-DLLAMA_HIPBLAS=on" pip install llama-cpp-python`
After installation, run `xinference-local` to start the Xinference service on your Mac.
## Creating and Deploying Models (Qwen-14B Example)
### 1. Launch via WebUI
After starting Xinference, open `http://127.0.0.1:9997` in your browser to access the Xinference Web UI.
Go to the "Launch Model" tab, search for qwen-chat, select the launch parameters, then click the rocket button in the lower left of the model card to deploy. The default Model UID is qwen-chat (used to access the model later).

On first launch, Xinference downloads model parameters from HuggingFace, which takes a few minutes. Model files are cached locally for subsequent launches. Xinference also supports downloading from other sources like [modelscope](https://inference.readthedocs.io/en/latest/models/sources/sources.html).
### 2. Launch via Command Line
You can also use Xinference's CLI to launch models. The default Model UID is qwen-chat.
```bash
xinference launch -n qwen-chat -s 14 -f pytorch
```
Beyond WebUI and CLI, Xinference also provides Python SDK and RESTful API. For more details, see the [Xinference documentation](https://inference.readthedocs.io/en/latest/getting_started/index.html).
## Integrate Local Models with One API
For One API deployment and setup, refer to [here](../config/model/intro.en.mdx).
Add a channel for qwen1.5-chat. Set the Base URL to the Xinference service endpoint and register qwen-chat (the model's UID).

Test with this command:
```bash
curl --location --request POST 'https://[oneapi_url]/v1/chat/completions' \
--header 'Authorization: Bearer [oneapi_token]' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "qwen-chat",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
Replace \[oneapi\_url] with your One API address and \[oneapi\_token] with your One API token. The model field should match the custom model name you entered in One API.
## Integrate Local Models with FastGPT
Add the qwen-chat model to the `llmModels` section of FastGPT's `config.json`:
```json
...
"llmModels": [
{
"model": "qwen-chat", // Model name (matches the channel model name in OneAPI)
"name": "Qwen", // Display name
"avatar": "/imgs/model/Qwen.svg", // Model logo
"maxContext": 125000, // Max context length
"maxResponse": 4000, // Max response length
"quoteMaxToken": 120000, // Max quote content tokens
"maxTemperature": 1.2, // Max temperature
"charsPointsPrice": 0, // n points/1k tokens (Commercial Edition)
"censor": false, // Enable content moderation (Commercial Edition)
"vision": true, // Supports image input
"toolChoice": true, // Supports tool choice (used in classification, extraction, tool calling)
"functionCall": false, // Supports function calling (used in classification, extraction, tool calling. toolChoice takes priority; if false, falls back to functionCall; if still false, uses prompt mode)
"customCQPrompt": "", // Custom classification prompt (for models without tool/function calling support)
"customExtractPrompt": "", // Custom content extraction prompt
"defaultSystemChatPrompt": "", // Default system prompt for conversations
"defaultConfig": {} // Default config sent with API requests (e.g., GLM4's top_p)
}
],
...
```
Restart FastGPT to select the Qwen model in app configuration:
## 
* Reference: [FastGPT + Xinference: One-Stop Local LLM Private Deployment and Application Development](https://xorbits.cn/blogs/fastgpt-weather-chat)
file: ./content/self-host/custom-models/xinference.mdx
meta: {
"title": "使用 Xinference 接入本地模型",
"description": "一站式本地 LLM 私有化部署"
}
[Xinference](https://github.com/xorbitsai/inference) 是一款开源模型推理平台,除了支持 LLM,它还可以部署 Embedding 和 ReRank 模型,这在企业级 RAG 构建中非常关键。同时,Xinference 还提供 Function Calling 等高级功能。还支持分布式部署,也就是说,随着未来应用调用量的增长,它可以进行水平扩展。
## 安装 Xinference
Xinference 支持多种推理引擎作为后端,以满足不同场景下部署大模型的需要,下面会分使用场景来介绍一下这三种推理后端,以及他们的使用方法。
### 1. 服务器
如果你的目标是在一台 Linux 或者 Window 服务器上部署大模型,可以选择 Transformers 或 vLLM 作为 Xinference 的推理后端:
* [Transformers](https://huggingface.co/docs/transformers/index):通过集成 Huggingface 的 Transformers 库作为后端,Xinference 可以最快地 集成当今自然语言处理(NLP)领域的最前沿模型(自然也包括 LLM)。
* [vLLM](https://vllm.ai/): vLLM 是由加州大学伯克利分校开发的一个开源库,专为高效服务大型语言模型(LLM)而设计。它引入了 PagedAttention 算法, 通过有效管理注意力键和值来改善内存管理,吞吐量能够达到 Transformers 的 24 倍,因此 vLLM 适合在生产环境中使用,应对高并发的用户访问。
假设你服务器配备 NVIDIA 显卡,可以参考[这篇文章中的指令来安装 CUDA](https://xorbits.cn/blogs/langchain-streamlit-doc-chat),从而让 Xinference 最大限度地利用显卡的加速功能。
#### Docker 部署
你可以使用 Xinference 官方的 Docker 镜像来一键安装和启动 Xinference 服务(确保你的机器上已经安装了 Docker),命令如下:
```bash
docker run -p 9997:9997 --gpus all xprobe/xinference:latest xinference-local -H 0.0.0.0
```
#### 直接部署
首先我们需要准备一个 3.9 以上的 Python 环境运行来 Xinference,建议先根据 conda 官网文档安装 conda。 然后使用以下命令来创建 3.11 的 Python 环境:
```bash
conda create --name py311 python=3.11
conda activate py311
```
以下两条命令在安装 Xinference 时,将安装 Transformers 和 vLLM 作为 Xinference 的推理引擎后端:
```bash
pip install "xinference[transformers]"
pip install "xinference[vllm]"
pip install "xinference[transformers,vllm]" # 同时安装
```
PyPi 在 安装 Transformers 和 vLLM 时会自动安装 PyTorch,但自动安装的 CUDA 版本可能与你的环境不匹配,此时你可以根据 PyTorch 官网中的[安装指南](https://pytorch.org/get-started/locally/)来手动安装。
只需要输入如下命令,就可以在服务上启动 Xinference 服务:
```bash
xinference-local -H 0.0.0.0
```
Xinference 默认会在本地启动服务,端口默认为 9997。因为这里配置了-H 0.0.0.0参数,非本地客户端也可以通过机器的 IP 地址来访问 Xinference 服务。
### 2. 个人设备
如果你想在自己的 Macbook 或者个人电脑上部署大模型,推荐安装 CTransformers 作为 Xinference 的推理后端。CTransformers 是用 GGML 实现的 C++ 版本 Transformers。
[GGML](https://ggml.ai/) 是一个能让大语言模型在[消费级硬件上运行](https://github.com/ggerganov/llama.cpp/discussions/205)的 C++ 库。 GGML 最大的特色在于模型量化。量化一个大语言模型其实就是降低权重表示精度的过程,从而减少使用模型所需的资源。 例如,表示一个高精度浮点数(例如 0.0001)比表示一个低精度浮点数(例如 0.1)需要更多空间。由于 LLM 在推理时需要加载到内存中的,因此你需要花费硬盘空间来存储它们,并且在执行期间有足够大的 RAM 来加载它们,GGML 支持许多不同的量化策略,每种策略在效率和性能之间提供不同的权衡。
通过以下命令来安装 CTransformers 作为 Xinference 的推理后端:
```bash
pip install xinference
pip install ctransformers
```
因为 GGML 是一个 C++ 库,Xinference 通过 `llama-cpp-python` 这个库来实现语言绑定。对于不同的硬件平台,我们需要使用不同的编译参数来安装:
* Apple Metal(MPS):`CMAKE_ARGS="-DLLAMA_METAL=on" pip install llama-cpp-python`
* Nvidia GPU:`CMAKE_ARGS="-DLLAMA_CUBLAS=on" pip install llama-cpp-python`
* AMD GPU:`CMAKE_ARGS="-DLLAMA_HIPBLAS=on" pip install llama-cpp-python`
安装后只需要输入 `xinference-local`,就可以在你的 Mac 上启动 Xinference 服务。
## 创建并部署模型(以 Qwen-14B 模型为例)
### 1. WebUI 方式启动模型
Xinference 启动之后,在浏览器中输入: `http://127.0.0.1:9997`,我们可以访问到本地 Xinference 的 Web UI。
打开“Launch Model”标签,搜索到 qwen-chat,选择模型启动的相关参数,然后点击模型卡片左下方的小火箭🚀按钮,就可以部署该模型到 Xinference。 默认 Model UID 是 qwen-chat(后续通过将通过这个 ID 来访问模型)。

当你第一次启动 Qwen 模型时,Xinference 会从 HuggingFace 下载模型参数,大概需要几分钟的时间。Xinference 将模型文件缓存在本地,这样之后启动时就不需要重新下载了。 Xinference 还支持从其他模型站点下载模型文件,例如 [modelscope](https://inference.readthedocs.io/en/latest/models/sources/sources.html)。
### 2. 命令行方式启动模型
我们也可以使用 Xinference 的命令行工具来启动模型,默认 Model UID 是 qwen-chat(后续通过将通过这个 ID 来访问模型)。
```bash
xinference launch -n qwen-chat -s 14 -f pytorch
```
除了 WebUI 和命令行工具, Xinference 还提供了 Python SDK 和 RESTful API 等多种交互方式, 更多用法可以参考 [Xinference 官方文档](https://inference.readthedocs.io/en/latest/getting_started/index.html)。
## 将本地模型接入 One API
One API 的部署和接入请参考[这里](../config/model/intro.mdx)。
为 qwen1.5-chat 添加一个渠道,这里的 Base URL 需要填 Xinference 服务的端点,并且注册 qwen-chat (模型的 UID) 。

可以使用以下命令进行测试:
```bash
curl --location --request POST 'https://[oneapi_url]/v1/chat/completions' \
--header 'Authorization: Bearer [oneapi_token]' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "qwen-chat",
"messages": [{"role": "user", "content": "Hello!"}]
}'
```
将 \[oneapi\_url] 替换为你的 One API 地址,\[oneapi\_token] 替换为你的 One API 令牌。model 为刚刚在 One API 填写的自定义模型。
## 将本地模型接入 FastGPT
修改 FastGPT 的 `config.json` 配置文件的 llmModels 部分加入 qwen-chat 模型:
```json
...
"llmModels": [
{
"model": "qwen-chat", // 模型名(对应OneAPI中渠道的模型名)
"name": "Qwen", // 模型别名
"avatar": "/imgs/model/Qwen.svg", // 模型的logo
"maxContext": 125000, // 最大上下文
"maxResponse": 4000, // 最大回复
"quoteMaxToken": 120000, // 最大引用内容
"maxTemperature": 1.2, // 最大温度
"charsPointsPrice": 0, // n积分/1k token(商业版)
"censor": false, // 是否开启敏感校验(商业版)
"vision": true, // 是否支持图片输入
"toolChoice": true, // 是否支持工具选择(分类,内容提取,工具调用会用到。)
"functionCall": false, // 是否支持函数调用(分类,内容提取,工具调用会用到。会优先使用 toolChoice,如果为false,则使用 functionCall,如果仍为 false,则使用提示词模式)
"customCQPrompt": "", // 自定义文本分类提示词(不支持工具和函数调用的模型
"customExtractPrompt": "", // 自定义内容提取提示词
"defaultSystemChatPrompt": "", // 对话默认携带的系统提示词
"defaultConfig": {} // 请求API时,挟带一些默认配置(比如 GLM4 的 top_p)
}
],
...
```
然后重启 FastGPT 就可以在应用配置中选择 Qwen 模型进行对话:
## 
* 参考:[FastGPT + Xinference:一站式本地 LLM 私有化部署和应用开发](https://xorbits.cn/blogs/fastgpt-weather-chat)
file: ./content/self-host/design/dataset.en.mdx
meta: {
"title": "Dataset Design",
"description": "FastGPT dataset file and data design"
}
## Relationship Between Files and Data
In FastGPT, files are stored using MongoDB's GridFS, while the actual data is stored in PostgreSQL. Each row in PG has a `file_id` column that references the corresponding file. For backward compatibility and to support manual input and annotated data, `file_id` has some special values:
* manual: Manually entered data
* mark: Manually annotated data
Note: `file_id` is only written at data insertion time and cannot be modified afterward.
## File Import Process
1. Upload the file to MongoDB GridFS and obtain a `file_id`. The file is marked as `unused` at this point.
2. The browser parses the file to extract text and chunks.
3. Each chunk is tagged with the `file_id`.
4. Click upload: the file status changes to `used`, and the data is pushed to the mongo `training` collection to await processing.
5. The training thread pulls data from mongo, generates vectors, and inserts them into PG.
file: ./content/self-host/design/dataset.mdx
meta: {
"title": "数据集",
"description": "FastGPT 数据集中文件与数据的设计方案"
}
## 文件与数据的关系
在 FastGPT 中,文件会通过 MongoDB 的 FS 存储,而具体的数据会通过 PostgreSQL 存储,PG 中的数据会有一列 file\_id,关联对应的文件。考虑到旧版本的兼容,以及手动输入、标注数据等,我们给 file\_id 增加了一些特殊的值,如下:
* manual: 手动输入
* mark: 手动标注的数据
注意,file\_id 仅在插入数据时会写入,变更时无法修改。
## 文件导入流程
1. 上传文件到 MongoDB 的 FS 中,获取 file\_id,此时文件标记为 `unused` 状态
2. 浏览器解析文件,获取对应的文本和 chunk
3. 给每个 chunk 打上 file\_id
4. 点击上传数据:将文件的状态改为 `used`,并将数据推送到 mongo `training` 表中等待训练
5. 由训练线程从 mongo 中取数据,并在获取向量后插入到 pg。
file: ./content/self-host/migration/docker_db.en.mdx
meta: {
"title": "Docker Database Migration (Simple Method)",
"description": "FastGPT Docker database backup and migration"
}
## 1. Stop Services
```bash
docker-compose down
```
## 2. Copy Directories
Docker-deployed databases mount local directories into containers via volumes. To migrate, simply copy these directories.
`PG data`: pg/data
`Mongo data`: mongo/data
Just copy the entire pg and mongo directories to the new location.
file: ./content/self-host/migration/docker_db.mdx
meta: {
"title": "Docker 数据库迁移(无脑操作)",
"description": "FastGPT Docker 数据库备份和迁移"
}
## 1. 停止服务
```bash
docker-compose down
```
## 2. Copy文件夹
Docker 部署数据库都会通过 volume 挂载本地的目录进入容器,如果要迁移,直接复制这些目录即可。
`PG 数据`: pg/data
`Mongo 数据`: mongo/data
直接把pg 和 mongo目录全部复制走即可。
file: ./content/self-host/migration/docker_mongo.en.mdx
meta: {
"title": "Docker MongoDB Migration (Dump Mode)",
"description": "FastGPT Docker MongoDB migration"
}
## Author
[https://github.com/samqin123](https://github.com/samqin123)
[Related PR -- open this to discuss with the author](https://github.com/labring/FastGPT/pull/1426)
## Overview
How to use mongodump to migrate FastGPT's MongoDB from Environment A to Environment B.
Prerequisites:
* Environment A: Your existing FastGPT deployment (e.g., on Alibaba Cloud) that needs to be migrated.
* Environment B: The new FastGPT deployment (e.g., on Tencent Cloud, or a NAS like Synology/QNAP). Note: NAS deployments may require MongoDB 4.2 or 4.4, while cloud deployments support the default FastGPT MongoDB version.
* Environment C: Your local machine, used as a staging area to hold files and coordinate the transfer.
## 1. Prepare: Access Docker MongoDB \[Environment A]
```
docker exec -it mongo sh
mongo -u 'username' -p 'password'
>> show dbs
```
Confirm you can see the fastgpt database and note the database name for export.
##### Preparation:
Create a temporary directory for import/export on both the container and the host, e.g., data/backup \[Environment A + Environment C].
#### Create the directory in \[Environment A] for the dump operation
Enter the FastGPT Docker container:
```
docker exec -it fastgpt sh
mkdir -p /data/backup
```
Once created, exported MongoDB data will appear in the `data/backup` directory under your local FastGPT installation folder (auto-synced via volume mount). If it doesn't sync automatically, you can manually create the directory and use `docker cp` to copy files out (this rarely happens).
#### Then set up the \[Environment C] host directory for syncing uploaded files into the container.
Navigate to the FastGPT directory, go into the mongo folder, and create a backup subdirectory:
```
mkdir -p /fastgpt/data/backup
```
Also create a directory in the new \[Environment B]:
```
mkdir -p /fastgpt/mongobackup
```
\###2. Export Data from \[Environment A]
Enter Environment A and use mongodump to export the MongoDB database.
#### 2.1 Export
Run mongodump to export data files to the temporary directory (data/backup).
\[The export path is set to /data/backup in the command. Since the FastGPT config already has data persistence set up, the exported files will sync to the host's fastgpt/mongo/data/backup directory.]
Single command to export (run on the host, no need to enter the container):
```
docker exec -it mongo bash -c "mongodump --db fastgpt -u 'username' -p 'password' --authenticationDatabase admin --out /data/backup"
```
You can also enter the container and combine directory creation with the export:
```
1.docker exec -it fastgpt sh
2.mkdir -p /data/backup
3. mongodump --host 127.0.0.1:27017 --db fastgpt -u "username" -p "password" --authenticationDatabase admin --out /data/backup
```
##### Fallback: if files don't auto-sync, manually copy them to the host \[Environment A]:
```
docker cp mongo:/data/backup [local-fastgpt-dir]:/fastgpt/data/backup>
```
2.2 For beginners, it's recommended to compress the directory and download it to your local staging environment \[A -> C] for verification. This ensures you have a backup and can check file counts. Experienced users can transfer directly to the new server \[A -> B].
2.2.1 Navigate to the \[Environment A] source system's local fastgpt/mongo/data directory:
```
cd /usr/fastgpt/mongo/data
```
Compress the files:
```
tar -czvf ../fastgpt-mongo-backup-$(date +%Y-%m-%d).tar.gz ./
```
Download the archive to your local machine \[A -> C] for verification. Experienced users can sync directly to Environment B's fastgpt data directory.
```
scp -i /Users/path/[your-pem-file] root@[cloud-server-ip]:/usr/fastgpt/mongo/fastgptbackup-2024-05-03.tar.gz /[local-path]/Downloads/fastgpt
```
Experienced users can transfer directly to the new environment:
```
scp -i /Users/path/[your-pem-file] root@[old-server-ip]:/usr/fastgpt/mongo/fastgptbackup-2024-05-03.tar.gz root@[new-server-ip]:/Downloads/fastgpt2
```
2.2 \[Environment C] Verify the archive is complete. If not, re-export. Cross-environment scp transfers can occasionally lose data.
After downloading the archive to Environment C, extract it to a custom directory, e.g., user/fastgpt/mongobackup/data:
```
tar -xvzf fastgptbackup-2024-05-03.tar.gz -C user/fastgpt/mongobackup/data
```
The extracted files should be .bson files. Verify the file count matches the source. If they don't match, the new FastGPT environment will have no data after import.
If everything looks good, upload the archive to Environment B's designated directory (e.g., /fastgpt/mongobackup). Do not place it in fastgpt/data/ -- that directory will be cleared later, and having extra files there will cause import errors.
```
scp -rfv [local-path]/Downloads/fastgpt/fastgptbackup-2024-05-03.tar.gz root@[new-server-ip]:/Downloads/fastgpt/backup
```
## 3. Import and Restore
### 3.1. Extract the archive on the new FastGPT environment
```
tar -xvzf fastgptbackup-2024-05-03.tar.gz -C user/fastgpt/mongobackup/data
```
Verify the file count again against your earlier check.
Experienced users can use tar to verify archive integrity. The above steps are for beginners to facilitate comparison.
### 3.2 Manually copy files into the new FastGPT Docker container \[Environment C]
Since the files aren't in the data/ directory, they won't auto-sync into the container. Also ensure the container's data directory is clean, or the import will fail.
```
docker cp user/fastgpt/mongobackup/data mongo:/tmp/backup
```
### 3.3 Initialize docker compose -- run it once to create the new mongo/data persistence directory
If the mongo/db directory isn't freshly initialized, mongorestore may fail. If you encounter errors, try initializing mongo.
Commands:
```
cd /fastgpt-install-dir/mongo/data
rm -rf *
```
4. Restore with mongorestore \[Environment C]
Run this from the host to import in one command (you can also run it inside the container):
```
docker exec -it mongo mongorestore -u "username" -p "password" --authenticationDatabase admin /tmp/backup/ --db fastgpt
```
Note: if the imported file count seems too low, the import likely failed. A failed import means you can log in to FastGPT but see no data.
5. Restart containers \[Environment C]
```
docker compose restart
docker logs -f mongo # Strongly recommended: check mongo logs before logging in. If mongo has errors, the web UI will also show errors.
```
If mongo starts normally, you should see output like this (not "mongo is restarting" -- that indicates an error):
Error state:
6. After starting the FastGPT container, log in to the web UI. If all your original data is displayed, the migration was successful.
file: ./content/self-host/migration/docker_mongo.mdx
meta: {
"title": "Docker Mongo迁移(dump模式)",
"description": "FastGPT Docker Mongo迁移"
}
## 作者
[https://github.com/samqin123](https://github.com/samqin123)
[相关PR。有问题可打开这里与作者交流](https://github.com/labring/FastGPT/pull/1426)
## 介绍
如何使用Mongodump来完成从A环境到B环境的Fastgpt的mongodb迁移
前提说明:
A环境:我在阿里云上部署的fastgpt,现在需要迁移到B环境。
B环境:是新环境比如腾讯云新部署的fastgpt,更特殊一点的是,NAS(群晖或者QNAP)部署了fastgpt,mongo必须改成4.2或者4.4版本(其实云端更方便,支持fastgpt mongo默认版本)
C环境:妥善考虑,用本地电脑作为C环境过渡,保存相关文件并分离操作
## 1. 环境准备:进入 docker mongo 【A环境】
```
docker exec -it mongo sh
mongo -u 'username' -p 'password'
>> show dbs
```
看到fastgpt数据库,以及其它几个,确定下导出数据库名称
准备:
检查数据库,容器和宿主机都创建一下 backup 目录 【A环境 + C环境】
##### 准备:
检查数据库,容器和宿主机都创建一下"数据导出导入"临时目录 ,比如data/backup 【A环境建目录 + C环境建目录用于同步到容器中】
#### 先在【A环境】创建文件目录,用于dump导出操作
容器:(先进入fastgpt docker容器)
```
docker exec -it fastgpt sh
mkdir -p /data/backup
```
建好后,未来导出mongo的数据,会在A环境本地fastgpt的安装目录/Data/下看到自动同步好的目录,数据会在data\backup中,然后可以衔接后续的压缩和下载转移动作。如果没有同步到本地,也可以手动建一下,配合docker cp 把文件拷到本地用(基本不会发生)
#### 然后,【C环境】宿主机目录类似操作,用于把上传的文件自动同步到C环境部署的fastgpt容器里。
到fastgpt目录,进入mongo目录,有data目录,下面建backup
```
mkdir -p /fastgpt/data/backup
```
准备好后,后续上传
```
### 新fastgpt环境【B】中也需要建一个,比如/fastgpt/mongobackup目录,注意不要在fastgpt/data目录下建立目录
```
mkdir -p /fastgpt/mongobackup
```
###2. 正题开始,从fastgpt老环境【A】中导出数据
进入A环境,使用mongodump 导出mongo数据库。
#### 2.1 导出
可以使用mongodump在源头容器中导出数据文件, 导出路径为上面指定临时目录,即"data\backup"
[导出的文件在代码中指定为/data/backup,因为fastgpt配置文件已经建立了data的持久化,所以会同步到容器所在环境本地fast/mongo/data应该就能看到这个导出的目录:backup,里面有文件]
一行指令导出代码,在服务器本地环境运行,不需要进入容器。
```
docker exec -it mongo bash -c "mongodump --db fastgpt -u 'username' -p 'password' --authenticationDatabase admin --out /data/backup"
```
也可以进入环境,熟手可以结合建目录,一次性完成建导出目录,以及使用mongodump导出数据到该目录
```
1.docker exec -it fastgpt sh
2.mkdir -p /data/backup
3. mongodump --host 127.0.0.1:27017 --db fastgpt -u "username" -p "password" --authenticationDatabase admin --out /data/backup
##### 补充:万一没自动同步,也可以将mongodump导出的文件,手工导出到宿主机【A环境】,备用指令如下:
````
docker cp mongo:/data/backup [A环境本地fastgpt目录]:/fastgpt/data/backup>
```
2.2 对新手,建议稳妥起见,压缩这个文件目录,并将压缩文件下载到本地过渡环境【A环境 -> C环境】;原因是因为留存一份,并且检查文件数量是否一致。
熟手可以直接复制到新部署服务器(腾讯云或者NAS)【A环境-> B环境】
2.2.1 先进入 【A环境】源头系统的本地环境 fastgpt/mongo/data 目录
````
cd /usr/fastgpt/mongo/data
```
#执行,压缩文件命令
```
tar -czvf ../fastgpt-mongo-backup-$(date +%Y-%m-%d).tar.gz ./ 【A环境】
```
#接下来,把压缩包下载到本地 【A环境-> C环境】,以便于检查和留存版本。熟手,直接将该压缩包同步到B环境中新fastgpt目录data目录下备用。
```
scp -i /Users/path/\[user.pem换成你自己的pem文件链接] root@\[fastgpt所在云服务器地址]:/usr/fastgpt/mongo/fastgptbackup-2024-05-03.tar.gz /\[本地电脑路径]/Downloads/fastgpt
```
熟手直接换成新环境地址
```
scp -i /Users/path/\[user.pem换成你自己的pem文件链接] root@\[老环境fastgpt服务器地址]:/usr/fastgpt/mongo/fastgptbackup-2024-05-03.tar.gz root@\[新环境fastgpt服务器地址]:/Downloads/fastgpt2
```
2.2 【C环境】检查压缩文件是否完整,如果不完整,重新导出。事实上,我也出现过问题,因为跨环境scp会出现丢数据的情况。
压缩数据包导入到C环境本地后,可以考虑在宿主机目录解压缩,放在一个自定义目录比如. [ user/fastgpt/mongobackup/data]
```
tar -xvzf fastgptbackup-2024-05-03.tar.gz -C user/fastgpt/mongobackup/data
```
解压缩后里面是bson文件,这里可以检查下,压缩文件数量是否一致。如果不一致,后续启动新环境的fastgpt容器,也不会有任何数据。
如果没问题,准备进入下一步,将压缩包文件上传到B环境,也就是新fastgpt环境里的指定目录,比如/fastgpt/mongobackup, 注意不要放到fastgpt/data目录下,因为下面会先清空一次这个目录,否则导入会报错。
```
scp -rfv \[本地电脑路径]/Downloads/fastgpt/fastgptbackup-2024-05-03.tar.gz root@\[新环境fastgpt服务器地址]:/Downloads/fastgpt/backup
```
## 3 导入恢复: 实际恢复和导入步骤
### 3.1. 进入新fastgpt本地环境的安装目录后,找到迁移的压缩文件包fastgptbackup-2024-05-03.tar.gz,解压缩到指定目录
```
tar -xvzf fastgptbackup-2024-05-03.tar.gz -C user/fastgpt/mongobackup/data
```
再次核对文件数量,和上面对比一下。
熟手可以用tar指令检查文件完整性,上面是给新手准备的,便于比对核查。
### 3.2 手动上传新fastgpt docker容器里备用 【C环境】
说明:因为没有放在data里,所以不会自动同步到容器里。而且要确保容器的data目录被清理干净,否则导入时会报错。
```
docker cp user/fastgpt/mongobackup/data mongo:/tmp/backup
\`\`\`
### 3.3 建议初始化一次docker compose ,运行后建立新的 mongo/data 持久化目录
如果不是初始化的 mongo/db 目录, mongorestore 导入可能会报错。如果报错,建议尝试初始化mongo。
操作指令
```
cd /fastgpt安装目录/mongo/data
rm -rf *
```
4.恢复: mongorestore 恢复 【C环境】
简单一点,退回到本地环境,用 docker 命令一键导入,当然你也可以在容器里操作
```
docker exec -it mongo mongorestore -u "username" -p "password" --authenticationDatabase admin /tmp/backup/ --db fastgpt
```
注意:导入文件数量量级太少,大概率是没导入成功的表现。如果导入不成功,新环境fastgpt可以登入,但是一片空白。
5.重启容器 【C环境】
```
docker compose restart
docker logs -f mongo **强烈建议先检查mongo运行情况,在去做登录动作,如果mongo报错,访问web也会报错"
```
如果mongo启动正常,显示的是类似这样的,而不是 "mongo is restarting",后者就是错误
报错情况
6. 启动fastgpt容器服务后,登录新fastgpt web,能看到原来的数据库内容完整显示,说明已经导入系统了。
file: ./content/self-host/deploy/docker.en.mdx
meta: {
"title": "Deploy with Docker Compose",
"description": "Quickly deploy FastGPT using Docker Compose"
}
import { Alert } from '@/components/docs/Alert';
import { CurrentOriginCodeBlockUpdater } from '@/components/docs/CurrentOriginCodeBlockUpdater';
## Prerequisites
1. Basic networking knowledge: ports, firewalls, etc.
2. Docker and Docker Compose basics
## Deployment Architecture

* MongoDB: Stores all data except vectors
* PostgreSQL/Milvus/Oceanbase/SeekDB: Stores vector data
* AIProxy: Aggregates various AI APIs with multi-model support (for any model issues, test with OneAPI first)
## Recommended Specs
### PgVector Version
Very lightweight, suitable for knowledge base indexes under 50 million.
| Environment | Minimum (Single Node) | Recommended |
| ---------------------------------- | --------------------- | ------------ |
| Testing (reduce compute processes) | 2c4g | 2c8g |
| 1M vector groups | 4c8g 50GB | 4c16g 50GB |
| 5M vector groups | 8c32g 200GB | 16c64g 200GB |
### Milvus Version
Better performance for 100M+ vectors.
[View Milvus official recommended specs](https://milvus.io/docs/prerequisite-docker.md)
| Environment | Minimum (Single Node) | Recommended |
| ---------------- | --------------------- | ----------- |
| Testing | 2c8g | 4c16g |
| 1M vector groups | Not tested | |
| 5M vector groups | | |
### Zilliz Cloud Version
Zilliz Cloud is built by the Milvus team — a fully managed SaaS vector database with better performance than Milvus and SLA guarantees. [Try Zilliz Cloud](https://zilliz.com.cn/).
Since the vector database runs in the cloud, no local resources are needed.
### SeekDB Version
SeekDB is a high-performance vector database based on MySQL protocol, fully compatible with OceanBase, supporting efficient vector retrieval.
| Environment | Minimum (Single Node) | Recommended |
| ---------------------------------- | --------------------- | ------------ |
| Testing (reduce compute processes) | 2c4g | 2c8g |
| 1M vector groups | 4c8g 50GB | 4c16g 50GB |
| 5M vector groups | 8c32g 200GB | 16c64g 200GB |
SeekDB uses MySQL protocol, fully compatible with OceanBase:
* Supports 1536-dimensional vector retrieval
* Built-in HNSW index algorithm
* Batch insert and query optimization
* Automatic retry and connection pool management
## Preparation
### Prepare Docker Environment
```bash
# Install Docker
curl -fsSL https://get.docker.com | bash -s docker --mirror Aliyun
systemctl enable --now docker
# Install docker-compose
curl -L https://github.com/docker/compose/releases/download/v2.20.3/docker-compose-`uname -s`-`uname -m` -o /usr/local/bin/docker-compose
chmod +x /usr/local/bin/docker-compose
# Verify installation
docker -v
docker compose -v
# If it fails, search online for solutions
```
We recommend [Orbstack](https://orbstack.dev/). Install via Homebrew:
```bash
brew install orbstack
```
Or [download the installer](https://orbstack.dev/download) directly.
We recommend storing source code and data in the Linux filesystem when binding to Linux containers, not the Windows filesystem.
You can [install Docker Desktop with WSL 2 backend on Windows](https://docs.docker.com/desktop/wsl/).
Or [install the command-line version of Docker directly in WSL 2](https://nickjanetakis.com/blog/install-docker-in-wsl-2-without-docker-desktop).
## Start Deployment
### 1. Get Configuration Files
#### Method 1: Deploy with an AI Agent
Copy the following content to your Coding Agent:
```text
Refer to https://doc.fastgpt.cn/deploy/SKILL.md and deploy FastGPT with Docker for me.
```
#### Method 2: Interactive Script Deployment
Run in Linux/MacOS/Windows WSL. The script guides you through selecting deployment environment, vector database version, IP address, etc.
```bash
FASTGPT_DEPLOY_BASE_URL=https://doc.fastgpt.cn bash <(curl -fsSL https://doc.fastgpt.cn/deploy/install.sh)
```
If the documentation site uses a custom domain, an internal domain, or a local address, set `FASTGPT_DEPLOY_BASE_URL` to choose the download source. You can provide either the site root or a URL ending in `/deploy`; the script downloads YAML and `config.json` from that source:
```bash
FASTGPT_DEPLOY_BASE_URL=https://doc.fastgpt.cn bash <(curl -fsSL https://doc.fastgpt.cn/deploy/install.sh)
```
Non-interactive mode also requires `FASTGPT_FE_DOMAIN`, the full URL users use to access FastGPT, such as `https://fastgpt.example.com`, and `FASTGPT_SANDBOX_PROXY_URL`, the Sandbox WebSocket URL, such as `wss://sandbox-proxy.example.com`. Version 4.16 also requires `FASTGPT_SANDBOX_PREVIEW_PROXY_URL` for the HTTP preview URL. In interactive mode, the script prompts for the addresses required by each version; 4.15 prompts only for the WebSocket URL.
The script automatically:
* Downloads `docker-compose.yml`.
* Guides you through selecting externally accessible S3 and MCP addresses, then writes them into the config files.
* Generates a random `root` login password, service tokens, app keys, and component passwords, then writes them into `docker-compose.yml`.
* Detects the host Docker socket path and updates the mount path in `docker-compose.yml` when needed.
After the script finishes, the terminal prints the generated `root` login password. Keep the generated `docker-compose.yml` safe. For future upgrades, start from this file so you do not lose the generated passwords and keys.
To use an existing local `docker-compose.yml` file, for example when testing a version that has not been published to the docs site yet, choose `本地 docker-compose.yml` (local docker-compose.yml) in the deployment version step and enter the local file path. You can also pass the path with an environment variable:
```bash
FASTGPT_LOCAL_COMPOSE_PATH=/path/to/docker-compose.yml bash <(curl -fsSL https://doc.fastgpt.cn/deploy/install.sh)
```
#### Method 3: Manual Download
If you need to pin deployment to a specific `docker-compose.yml` file, we recommend downloading both `docker-compose.yml` and `install.sh`, then using the script's local compose mode to generate the final config. This keeps the script's random credential generation, S3/MCP address updates, and Docker socket detection.
1. Download the required `docker-compose.yml` file to the server, for example:
```bash
curl -fsSL https://doc.fastgpt.cn/deploy/docker/v4.15/cn/docker-compose.pg.yml -o docker-compose.source.yml
```
Click to view docker-compose config file download links for different databases
* **Pgvector**
* China mirror (Alibaba Cloud): [docker-compose.pg.yml](/deploy/docker/v4.15/cn/docker-compose.pg.yml)
* Global mirror (dockerhub, ghcr): [docker-compose.pg.yml](/deploy/docker/v4.15/global/docker-compose.pg.yml)
* **Oceanbase**
* China mirror (Alibaba Cloud): [docker-compose.oceanbase.yml](/deploy/docker/v4.15/cn/docker-compose.oceanbase.yml)
* Global mirror (dockerhub, ghcr): [docker-compose.oceanbase.yml](/deploy/docker/v4.15/global/docker-compose.oceanbase.yml)
* **Milvus**
* China mirror (Alibaba Cloud): [docker-compose.milvus.yml](/deploy/docker/v4.15/cn/docker-compose.milvus.yml)
* Global mirror (dockerhub, ghcr): [docker-compose.milvus.yml](/deploy/docker/v4.15/global/docker-compose.milvus.yml)
* **Zilliz**
* China mirror (Alibaba Cloud): [docker-compose.zilliz.yml](/deploy/docker/v4.15/cn/docker-compose.zilliz.yml)
* Global mirror (dockerhub, ghcr): [docker-compose.zilliz.yml](/deploy/docker/v4.15/global/docker-compose.zilliz.yml)
* **SeekDB**
* China mirror (Alibaba Cloud): [docker-compose.seekdb.yml](/deploy/docker/v4.15/cn/docker-compose.seekdb.yml)
* Global mirror (dockerhub, ghcr): [docker-compose.seekdb.yml](/deploy/docker/v4.15/global/docker-compose.seekdb.yml)
2. Download `install.sh` to the server:
```bash
curl -fsSL https://doc.fastgpt.cn/deploy/install.sh -o install.sh
```
3. Run `install.sh` with the local compose file to generate the final deployment config:
```bash
FASTGPT_LOCAL_COMPOSE_PATH=./docker-compose.source.yml bash install.sh
```
The script copies this compose file to the final `docker-compose.yml`, then generates login passwords and credentials, and writes the S3/MCP addresses. After generation, log in with the root password printed in the terminal.
For a fully offline environment, prepare `docker-compose.yml` and `install.sh` in advance. If the script cannot run, manually update `DEFAULT_ROOT_PSW`, service tokens, database passwords, and S3/MCP addresses.
#### Deploy with a Custom Image Registry
If you use an internal Harbor, private registry, or image mirror, download `docker-compose.yml` first, replace all `image:` values with your own registry addresses, and then use local compose mode:
```bash
FASTGPT_LOCAL_COMPOSE_PATH=./docker-compose.yml bash install.sh
```
If Agent/Skill Sandbox is enabled, also replace the sandbox-related images in the Compose file and update `AGENT_SANDBOX_SEALOS_IMAGE` or `AGENT_SANDBOX_OPENSANDBOX_IMAGE` so the sandbox provider can pull the matching images. See [OpenSandbox Configuration](../config/sandbox/opensandbox) for details.
### 2. Modify Environment Variables
You must set `FE_DOMAIN` in `fastgpt-app` to the full URL users use to access FastGPT, such as `https://fastgpt.example.com`. It must include a scheme, host, and optional port; do not leave it empty or use an internal container address.
When Agent/Skill Sandbox is enabled, also configure:
* `AGENT_SANDBOX_PROXY_URL`: the browser-accessible Sandbox Proxy WebSocket URL using `ws://` or `wss://`, such as `wss://sandbox-proxy.example.com`, pointing to port 3006.
* Version 4.16 additionally requires `AGENT_SANDBOX_PREVIEW_PROXY_URL`: the browser-accessible HTTP(S) URL for sandbox file previews, such as `https://sandbox-proxy.example.com`, also pointing to port 3006.
The interactive install script prompts for these addresses before the final confirmation.
For `Zilliz version`, you also need credentials — see [Deploy Zilliz Version: Get Account and Credentials](#deploy-zilliz-version-get-account-and-credentials). Other versions can skip to the next step.
### 3. Open External Ports / Configure Domain
These ports must be accessible:
1. Port 3000 (FastGPT main service)
2. Port 9000 (S3 service)
3. Port 3003 (FastGPT SSE MCP server service)
4. Port 3006 (FastGPT Agent Sandbox Proxy service)
### 4. Start Containers
Run in the same directory as docker-compose.yml. Ensure `docker-compose` version is 2.17+, or automated commands may fail.
```bash
# Pre-pull all service and sandbox runtime images
docker compose --profile prepull pull
# Start containers
docker compose up -d
```
### 5. Access FastGPT
Access FastGPT via the port/domain opened in step 3.
Login username is `root`, password is the `DEFAULT_ROOT_PSW` set in `docker-compose.yml` environment variables.
If you deploy with the interactive script, it randomly generates `DEFAULT_ROOT_PSW` and prints the login password when it finishes. If you deploy manually, change the default password in `docker-compose.yml` before starting the service. Each container restart automatically updates the root user's password based on `DEFAULT_ROOT_PSW`.
### 6. Configure Models
* After first login, the system prompts that `Language Model` and `Index Model` are not configured and automatically redirects to the model configuration page. At least these two model types are required.
* If the redirect doesn't happen, go to `Account - Model Providers` to configure models. [View tutorial](../config/model/intro.en.mdx)
* Known issue: after first entering the system, the browser tab may become unresponsive. Close the tab and reopen it.
### 7. Install System Plugins as Needed
Starting from V4.14.0, the fastgpt-plugin image only provides the runtime environment without pre-installed system plugins. All FastGPT systems must manually install system plugins.
* Install via the plugin marketplace — by default it fetches from the public FastGPT Marketplace.
* If your FastGPT can't access the marketplace, visit [FastGPT Plugin Marketplace](https://marketplace.fastgpt.cn/), download .pkg files, and import them via file upload.
* You can also sort tools, set default installations, and manage tags.

## FAQ
### FastGPT and FastGPT-plugin Version Compatibility
| FastGPT-plugin Version | FastGPT Main Service |
| ---------------------- | --------------------- |
| 1.x | 4.15.x |
| 0.6.x | >= 4.14.11, \< 4.15.0 |
| 0.5.x | >= 4.14.6, \< 4.14.11 |
| \< 0.5.0 | \< 4.14.6 |
### S3 Connection Issues
Check the `STORAGE_EXTERNAL_ENDPOINT` variable — it must be accessible by both the client and FastGPT service.
**Important:**
> Don't use `127.0.0.1` or `localhost` or other loopback addresses. Use the host machine's local IP when deploying with Docker, but set it to a static IP; or use a fixed domain name. This prevents 403 errors caused by URL mismatches when signing object storage URLs.
>
> See [Object Storage Configuration & Common Issues](../config/object-storage.en.mdx)
### Browser Unresponsive After Login
Can't click anything, refresh doesn't help. Close the tab and reopen it.
### Mongo Replica Set Auto-Initialization Failed
The latest docker-compose examples have fully automated Mongo replica set initialization. Tested on Ubuntu 20/22, CentOS 7, WSL2, macOS, and Windows. If it still won't start, the CPU likely doesn't support AVX instructions — switch to Mongo 4.x.
To manually initialize the replica set:
1. Create a mongo key in the terminal:
```bash
openssl rand -base64 756 > ./mongodb.key
chmod 600 ./mongodb.key
# Change key permissions — some systems use admin, others use root
chown 999:root ./mongodb.key
```
2. Modify docker-compose.yml to mount the key:
```yml
mongo:
# image: mongo:5.0.18
# image: registry.cn-hangzhou.aliyuncs.com/fastgpt/mongo:5.0.18 # Alibaba Cloud
container_name: mongo
ports:
- 27017:27017
networks:
- fastgpt
command: mongod --keyFile /data/mongodb.key --replSet rs0
environment:
# Default username and password, only effective on first run
- MONGO_INITDB_ROOT_USERNAME=myusername
- MONGO_INITDB_ROOT_PASSWORD=mypassword
volumes:
- ./mongo/data:/data/db
- ./mongodb.key:/data/mongodb.key
```
3. Restart services:
```bash
docker compose down
docker compose up -d
```
4. Enter the container and initialize the replica set:
```bash
# Check if mongo container is running
docker ps
# Enter container
docker exec -it mongo bash
# Connect to database (use your Mongo username and password)
mongo -u myusername -p mypassword --authenticationDatabase admin
# Initialize replica set. For external access, add directConnection=true to the Mongo connection parameters
rs.initiate({
_id: "rs0",
members: [
{ _id: 0, host: "mongo:27017" }
]
})
# Check status — if it shows rs0 status, it's running successfully
rs.status()
```
### How to Change API Address and Key
By default, OneAPI connection address and key are configured. Modify the environment variables in the fastgpt container in `docker-compose.yml`:
`OPENAI_BASE_URL` (API endpoint, must include /v1)
`CHAT_API_KEY` (API credentials)
After modifying, restart:
```bash
docker compose down
docker compose up -d
```
### How to Update Versions?
1. Check the [update documentation](../upgrading/upgrade-intruction.en.mdx) to confirm the target version — avoid skipping versions.
2. Change the image tag to the target version
3. Run these commands to pull and restart:
```bash
docker compose up -d
```
4. Run initialization scripts (if any)
### How to Customize Environment Variables?
Edit the `environment` section of `fastgpt-app` in `docker-compose.yml`, then run `docker compose up -d` to restart the container. For details, see [Environment Variables](../config/env.en.mdx).
### How to Check if Environment Variables Loaded
1. `docker exec -it fastgpt sh` to enter the container.
2. Run `env` to view all environment variables.
### Why Can't I Connect to Local Model Images
`docker-compose.yml` uses bridge mode to create the `fastgpt` network. To access other images via 0.0.0.0 or image name, add those images to the same network.
### How to Resolve Port Conflicts?
Docker-compose port format: `mapped_port:running_port`.
In bridge mode, container running ports don't conflict, but mapped ports can. Change the mapped port to a different value.
If `container1` needs to connect to `container2`, use `container2:running_port`.
(Brush up on Docker basics as needed)
### relation "modeldata" does not exist
PG database not connected or initialization failed — check logs. FastGPT initializes tables on each PG connection. Errors will appear in the logs.
1. Check if the database container started normally
2. For non-Docker deployments, manually install the pg vector extension
3. Check fastgpt logs for related errors
### Illegal instruction
Possible causes:
1. ARM architecture — use the official Mongo image: mongo:5.0.18
2. CPU doesn't support AVX — switch to mongo4.x. Change the mongo image to: mongo:4.4.29
### Operation `auth_codes.findOne()` buffering timed out after 10000ms
Mongo connection failed — check mongo's running status and **logs**.
Possible causes:
1. Mongo service didn't start (some CPUs don't support AVX — switch to mongo4.x, find the latest 4.x on Docker Hub, update the image version, and rerun)
2. Database connection environment variables are wrong (username/password, check host and port — for non-container network connections, use public IP and add directConnection=true)
3. Replica set startup failed, causing the container to keep restarting
4. `Illegal instruction.... Waiting for MongoDB to start`: CPU doesn't support AVX — switch to mongo4.x
### First Deployment: Root User Shows Unregistered
Logs will show error messages. Most likely Mongo replica set mode wasn't started.
### Can't Export Knowledge Base / Can't Use Voice Input or Playback
SSL certificate not configured — some features require it.
### Login Shows Network Error
Caused by service initialization errors triggering a restart.
* 90% of cases: incorrect config file causing JSON parsing errors
* The rest: usually because the vector database can't connect
### How to Change Password
Modify `DEFAULT_ROOT_PSW` in `docker-compose.yml` and restart — the password auto-updates.
### Deploy Zilliz Version: Get Account and Credentials
Open [Zilliz Cloud](https://zilliz.com.cn/), create an instance, and get the credentials.

1. Set `MILVUS_ADDRESS` and `MILVUS_TOKEN` to match Zilliz's `Public Endpoint` and `Api key`. Remember to add your IP to the whitelist.
file: ./content/self-host/deploy/docker.mdx
meta: {
"title": "Docker Compose 部署",
"description": "使用 Docker Compose 快速部署 FastGPT"
}
import { Alert } from '@/components/docs/Alert';
import { CurrentOriginCodeBlockUpdater } from '@/components/docs/CurrentOriginCodeBlockUpdater';
## 前置知识
1. 基础的网络知识:端口,防火墙……
2. Docker 和 Docker Compose 基础知识
## 部署架构图

* MongoDB:用于存储除了向量外的各类数据
* PostgreSQL/Milvus/Oceanbase/SeekDB:存储向量数据
* AIProxy: 聚合各类 AI API,支持多模型调用(任何模型问题,先自行通过 OneAPI 测试校验)
## 推荐配置
### PgVector 版本
非常轻量,适合知识库索引量在 5000 万以下。
| 环境 | 最低配置(单节点) | 推荐配置 |
| ---------------- | ----------- | ------------ |
| 测试(可以把计算进程设置少一些) | 2c4g | 2c8g |
| 100w 组向量 | 4c8g 50GB | 4c16g 50GB |
| 500w 组向量 | 8c32g 200GB | 16c64g 200GB |
### Milvus 版本
对于亿级以上向量性能更优秀。
[点击查看 Milvus 官方推荐配置](https://milvus.io/docs/prerequisite-docker.md)
| 环境 | 最低配置(单节点) | 推荐配置 |
| -------- | --------- | ----- |
| 测试 | 2c8g | 4c16g |
| 100w 组向量 | 未测试 | |
| 500w 组向量 | | |
### zilliz cloud 版本
Zilliz Cloud 由 Milvus 原厂打造,是全托管的 SaaS 向量数据库服务,性能优于 Milvus 并提供 SLA,点击使用 [Zilliz Cloud](https://zilliz.com.cn/)。
由于向量库使用了 Cloud,无需占用本地资源,无需太关注。
### SeekDB 版本
SeekDB 是基于 MySQL 协议的高性能向量数据库,与 OceanBase 协议完全兼容,支持高效的向量检索。
| 环境 | 最低配置(单节点) | 推荐配置 |
| ---------------- | ----------- | ------------ |
| 测试(可以把计算进程设置少一些) | 2c4g | 2c8g |
| 100w 组向量 | 4c8g 50GB | 4c16g 50GB |
| 500w 组向量 | 8c32g 200GB | 16c64g 200GB |
SeekDB 使用 MySQL 协议,与 OceanBase 完全兼容:
* 支持 1536 维向量检索
* 内置 HNSW 索引算法
* 提供批量插入和查询优化
* 自动重试和连接池管理
## 前置工作
### 准备 Docker-compose 环境
```bash
# 安装 Docker
curl -fsSL https://get.docker.com | bash -s docker --mirror Aliyun
systemctl enable --now docker
# 安装 docker-compose
curl -L https://github.com/docker/compose/releases/download/v2.20.3/docker-compose-`uname -s`-`uname -m` -o /usr/local/bin/docker-compose
chmod +x /usr/local/bin/docker-compose
# 验证安装
docker -v
docker compose -v
# 如失效,自行百度~
```
推荐直接使用 [Orbstack](https://orbstack.dev/)。可直接通过 Homebrew 来安装:
```bash
brew install orbstack
```
或者直接[下载安装包](https://orbstack.dev/download)进行安装。
我们建议将源代码和其他数据绑定到 Linux 容器中时,将其存储在 Linux 文件系统中,而不是 Windows 文件系统中。
可以选择直接[使用 WSL 2 后端在 Windows 中安装 Docker Desktop](https://docs.docker.com/desktop/wsl/)。
也可以直接[在 WSL 2 中安装命令行版本的 Docker](https://nickjanetakis.com/blog/install-docker-in-wsl-2-without-docker-desktop)。
## 开始部署
### 1. 获取配置文件
#### 方法一:使用 AI Agent 代部署
将以下内容复制给你的 Coding Agent:
```text
参考 https://doc.fastgpt.cn/deploy/SKILL.md 帮我部署 FastGPT Docker 版本。
```
#### 方法二:使用交互式脚本部署
需要在 Linux/MacOS/Windows WSL 环境下执行,引导用户选择部署环境、向量库版本,IP 地址等。
```bash
FASTGPT_DEPLOY_BASE_URL=https://doc.fastgpt.cn bash <(curl -fsSL https://doc.fastgpt.cn/deploy/install.sh)
```
非交互模式还必须通过 `FASTGPT_FE_DOMAIN` 指定用户访问 FastGPT 的完整地址,例如 `https://fastgpt.example.com`,并通过 `FASTGPT_SANDBOX_PROXY_URL` 指定沙盒 WebSocket 地址,例如 `wss://sandbox-proxy.example.com`。4.16 还需要通过 `FASTGPT_SANDBOX_PREVIEW_PROXY_URL` 指定 HTTP 预览地址。交互模式下脚本会按版本询问这些地址;4.15 只询问 WebSocket 地址。
脚本会自动完成以下操作:
* 下载 `docker-compose.yml`。
* 引导选择 S3 与 MCP 的外部访问地址,并写入配置文件。
* 随机生成 `root` 登录密码、服务间 Token、应用密钥和组件密码,并写入 `docker-compose.yml`。
* 自动检测宿主机 Docker socket 路径,必要时替换 `docker-compose.yml` 中的挂载路径。
执行完成后,终端会输出本次生成的 `root` 登录密码,请妥善保存生成后的 `docker-compose.yml`。后续升级时建议基于该文件调整,不要直接丢失已生成的密码和密钥。
#### 方法三:手动下载部署
如果需要固定使用某个 `docker-compose.yml` 文件,推荐先手动下载 `docker-compose.yml` 和 `install.sh`,再通过 `install.sh` 的本地 compose 模式生成最终配置。这样仍然可以复用脚本里的随机密码、S3/MCP 地址写入、Docker socket 检测等能力。
1. 下载所需的 `docker-compose.yml` 文件到服务器,例如:
```bash
curl -fsSL https://doc.fastgpt.cn/deploy/docker/v4.15/cn/docker-compose.pg.yml -o docker-compose.source.yml
```
点击展开查看不同数据库的 docker-compose 配置文件下载地址
* **Pgvector**
* 中国大陆地区镜像源(阿里云):[docker-compose.pg.yml](/deploy/docker/v4.15/cn/docker-compose.pg.yml)
* 全球镜像源(dockerhub, ghcr):[docker-compose.pg.yml](/deploy/docker/v4.15/global/docker-compose.pg.yml)
* **Oceanbase**
* 中国大陆地区镜像源(阿里云):[docker-compose.oceanbase.yml](/deploy/docker/v4.15/cn/docker-compose.oceanbase.yml)
* 全球镜像源(dockerhub, ghcr):[docker-compose.oceanbase.yml](/deploy/docker/v4.15/global/docker-compose.oceanbase.yml)
* **Milvus**
* 中国大陆地区镜像源(阿里云):[docker-compose.milvus.yml](/deploy/docker/v4.15/cn/docker-compose.milvus.yml)
* 全球镜像源(dockerhub, ghcr):[docker-compose.milvus.yml](/deploy/docker/v4.15/global/docker-compose.milvus.yml)
* **Zilliz**
* 中国大陆地区镜像源(阿里云):[docker-compose.zilliz.yml](/deploy/docker/v4.15/cn/docker-compose.zilliz.yml)
* 全球镜像源(dockerhub, ghcr):[docker-compose.zilliz.yml](/deploy/docker/v4.15/global/docker-compose.zilliz.yml)
* **SeekDB**
* 中国大陆地区镜像源(阿里云):[docker-compose.seekdb.yml](/deploy/docker/v4.15/cn/docker-compose.seekdb.yml)
* 全球镜像源(dockerhub, ghcr):[docker-compose.seekdb.yml](/deploy/docker/v4.15/global/docker-compose.seekdb.yml)
2. 下载 `install.sh` 到服务器:
```bash
curl -fsSL https://doc.fastgpt.cn/deploy/install.sh -o install.sh
```
3. 使用 `install.sh` 读取本地 compose 文件并生成最终部署配置:
```bash
FASTGPT_LOCAL_COMPOSE_PATH=./docker-compose.source.yml bash install.sh
```
脚本会复制该 compose 文件为最终的 `docker-compose.yml`,并继续随机生成登录密码和各类凭证、写入 S3/MCP 地址。生成完成后,按终端输出的 root 密码登录。
完全离线环境下,需要同时准备 `docker-compose.yml` 和 `install.sh`。如果无法运行脚本,则需要手动修改 `DEFAULT_ROOT_PSW`、服务 Token、数据库密码、S3/MCP 地址等配置。
#### 自定义镜像源部署
如果需要使用企业内网 Harbor、私有 Registry 或自建镜像加速源,可以先下载 `docker-compose.yml`,把所有 `image:` 改成自己的镜像地址,再走本地 compose 模式:
```bash
FASTGPT_LOCAL_COMPOSE_PATH=./docker-compose.yml bash install.sh
```
如果启用 Agent/Skill 沙盒,还需要同步替换 Compose 文件中的沙盒相关镜像,以及 `AGENT_SANDBOX_SEALOS_IMAGE` 或 `AGENT_SANDBOX_OPENSANDBOX_IMAGE`,确保沙盒 provider 可以拉取对应镜像。具体配置见 [OpenSandbox 配置](../config/sandbox/opensandbox)。
### 2. 修改环境变量
必须填写 `fastgpt-app` 中的 `FE_DOMAIN`,设置为用户实际访问 FastGPT 的完整地址,例如 `https://fastgpt.example.com`。该地址由协议、主机和可选端口组成,不能留空,也不要填写容器内部地址。
启用 Agent/Skill 沙盒时还必须配置:
* `AGENT_SANDBOX_PROXY_URL`:浏览器访问 Sandbox Proxy 的 WebSocket 地址,使用 `ws://` 或 `wss://`,例如 `wss://sandbox-proxy.example.com`,需要指向 3006 端口。
* 4.16 版本额外配置 `AGENT_SANDBOX_PREVIEW_PROXY_URL`:浏览器访问沙盒文件预览的 HTTP(S) 地址,例如 `https://sandbox-proxy.example.com`,同样需要指向 3006 端口。
使用交互式安装脚本时,脚本会在确认部署前询问这些地址。
对于 `Zilliz 版本` 还需要获取密钥,参考 [部署 Zilliz 版本获取账号和密钥](#部署-zilliz-版本获取账号和密钥), 其他版本可直接下一步。
### 3. 开放外网端口/配置域名
以下端口必须被访问到:
1. 3000 端口(FastGPT 主服务)
2. 9000 端口(S3 服务)
3. 3003 端口(FastGPT SSE MCP server 服务)
4. 3006 端口(FastGPT Agent Sandbox Proxy 服务)
### 4. 启动容器
在 docker-compose.yml 同级目录下执行。请确保 `docker-compose` 版本最好在 2.17 以上,否则可能无法执行自动化命令。
```bash
# 预拉取所有服务及沙盒运行时镜像
docker compose --profile prepull pull
# 启动容器
docker compose up -d
```
### 5. 访问 FastGPT
可通过第三步开放的端口/域名访问 FastGPT。登录用户名为 `root`,密码为 `docker-compose.yml` 环境变量里设置的 `DEFAULT_ROOT_PSW`。
如果使用交互式脚本部署,脚本会随机生成 `DEFAULT_ROOT_PSW`,并在执行完成后输出本次登录密码;如果手动下载部署,请自行修改 `docker-compose.yml` 中的默认密码后再启动服务。每次重启容器,都会按 `DEFAULT_ROOT_PSW` 自动更新 root 用户密码。
### 6. 配置模型
* 首次登录 FastGPT 后,系统会提示未配置 `语言模型` 和 `索引模型`,并自动跳转模型配置页面。系统必须至少有这两类模型才能正常使用。
* 如果系统未正常跳转,可以在 `账号-模型提供商` 页面,进行模型配置。[点击查看相关教程](../config/model/intro.mdx)
* 目前已知可能问题:首次进入系统后,整个浏览器 tab 无法响应。此时需要删除该 tab,重新打开一次即可。
### 7. 按需安装系统插件
从 V4.14.0 版本开始,fastgpt-plugin 镜像仅提供运行环境,不再预装系统插件,所有 FastGPT 系统需手动安装系统插件。
* 通过插件市场安装,默认会向公开的 FastGPT Marketplace 获取数据进行安装。
* 如果你的 FastGPT 无法访问插件市场,则可以手动访问 [FastGPT 插件市场](https://marketplace.fastgpt.cn/),先下载 .pkg 文件,再通过文件导入的方式安装到系统里。
* 除了安装外,还可对工具进行排序、默认安装、标签管理等。

## FAQ
### FastGPT 和 FastGPT-plugin 版本对应
| FastGPT-plugin 版本 | FastGPT 主服务 |
| ----------------- | --------------------- |
| 1.x | 4.15.x |
| 0.6.x | >= 4.14.11, \< 4.15.0 |
| 0.5.x | >= 4.14.6, \< 4.14.11 |
| \< 0.5.0 | \< 4.14.6 |
### S3 无法正常连接
检查 `STORAGE_EXTERNAL_ENDPOINT` 变量,需设置成客户端和 FastGPT 服务均可访问的地址。
**重要:**
> 填入的地址不可为 `127.0.0.1` 或者 `localhost` 等本地回环地址,可填 Docker 部署时的宿主机本地 IP,但是需要把宿主机固定为静态 IP;或者统一为一个固定域名;目的是为了避免对象存储签名 URL 时,签发与上传的 URL 不一致导致的 403 错误。
>
> 具体查看 [对象存储配置及常见问题](../config/object-storage.mdx)
### 登录系统后,浏览器无法响应
无法点击任何内容,刷新也无效。此时需要删除该 tab,重新打开一次即可。
### Mongo 副本集自动初始化失败
最新的 docker-compose 示例优化 Mongo 副本集初始化,实现了全自动。目前在 unbuntu20,22 centos7, wsl2, mac, window 均通过测试。仍无法正常启动,大部分是因为 cpu 不支持 AVX 指令集,可以切换 Mongo4.x 版本。
如果是由于,无法自动初始化副本集合,可以手动初始化副本集:
1. 终端中执行下面命令,创建 mongo 密钥:
```bash
openssl rand -base64 756 > ./mongodb.key
chmod 600 ./mongodb.key
# 修改密钥权限,部分系统是admin,部分是root
chown 999:root ./mongodb.key
```
2. 修改 docker-compose.yml,挂载密钥
```yml
mongo:
# image: mongo:5.0.18
# image: registry.cn-hangzhou.aliyuncs.com/fastgpt/mongo:5.0.18 # 阿里云
container_name: mongo
ports:
- 27017:27017
networks:
- fastgpt
command: mongod --keyFile /data/mongodb.key --replSet rs0
environment:
# 默认的用户名和密码,只有首次允许有效
- MONGO_INITDB_ROOT_USERNAME=myusername
- MONGO_INITDB_ROOT_PASSWORD=mypassword
volumes:
- ./mongo/data:/data/db
- ./mongodb.key:/data/mongodb.key
```
3. 重启服务
```bash
docker compose down
docker compose up -d
```
4. 进入容器执行副本集合初始化
```bash
# 查看 mongo 容器是否正常运行
docker ps
# 进入容器
docker exec -it mongo bash
# 连接数据库(这里要填Mongo的用户名和密码)
mongo -u myusername -p mypassword --authenticationDatabase admin
# 初始化副本集。如果需要外网访问,mongo:27017 。如果需要外网访问,需要增加Mongo连接参数:directConnection=true
rs.initiate({
_id: "rs0",
members: [
{ _id: 0, host: "mongo:27017" }
]
})
# 检查状态。如果提示 rs0 状态,则代表运行成功
rs.status()
```
### 如何修改 API 地址和密钥
默认是写了 OneAPi 的连接地址和密钥,可以通过修改 `docker-compose.yml` 中,fastgpt 容器的环境变量实现。
`OPENAI_BASE_URL`(API 接口的地址,需要加/v1)`CHAT_API_KEY`(API 接口的凭证)。
修改完后重启:
```bash
docker compose down
docker compose up -d
```
### 如何更新版本?
1. 查看[更新文档](../upgrading/upgrade-intruction.mdx),确认要升级的版本,避免跨版本升级。
2. 修改镜像 tag 到指定版本
3. 执行下面命令会自动拉取镜像:
```bash
docker compose up -d
```
4. 执行初始化脚本(如果有)
### 如何自定义环境变量?
修改 `docker-compose.yml` 中 `fastgpt-app` 的 `environment` 配置,并执行 `docker compose up -d` 重启容器。具体配置参考[环境变量说明](../config/env.mdx)。
### 如何检查环境变量是否正常加载
1. `docker exec -it fastgpt sh` 进入 FastGPT 容器。
2. 直接输入 `env` 命令查看所有环境变量。
### 为什么无法连接 `本地模型` 镜像
`docker-compose.yml` 中使用了桥接的模式建立了 `fastgpt` 网络,如想通过 0.0.0.0 或镜像名访问其它镜像,需将其它镜像也加入到网络中。
### 端口冲突怎么解决?
docker-compose 端口定义为:`映射端口:运行端口`。
桥接模式下,容器运行端口不会有冲突,但是会有映射端口冲突,只需将映射端口修改成不同端口即可。
如果 `容器1` 需要连接 `容器2`,使用 `容器2:运行端口` 来进行连接即可。
(自行补习 docker 基本知识)
### relation "modeldata" does not exist
PG 数据库没有连接上/初始化失败,可以查看日志。FastGPT 会在每次连接上 PG 时进行表初始化,如果报错会有对应日志。
1. 检查数据库容器是否正常启动
2. 非 docker 部署的,需要手动安装 pg vector 插件
3. 查看 fastgpt 日志,有没有相关报错
### Illegal instruction
可能原因:
1. arm 架构。需要使用 Mongo 官方镜像:mongo:5.0.18
2. cpu 不支持 AVX,无法用 mongo5,需要换成 mongo4.x。把 mongo 的 image 换成: mongo:4.4.29
### Operation `auth_codes.findOne()` buffering timed out after 10000ms
mongo 连接失败,查看 mongo 的运行状态**对应日志**。
可能原因:
1. mongo 服务有没有起来(有些 cpu 不支持 AVX,无法用 mongo5,需要换成 mongo4.x,可以 docker hub 找个最新的 4.x,修改镜像版本,重新运行)
2. 连接数据库的环境变量填写错误(账号密码,注意 host 和 port,非容器网络连接,需要用公网 ip 并加上 directConnection=true)
3. 副本集启动失败。导致容器一直重启。
4. `Illegal instruction.... Waiting for MongoDB to start` : cpu 不支持 AVX,无法用 mongo5,需要换成 mongo4.x
### 首次部署,root 用户提示未注册
日志会有错误提示。大概率是没有启动 Mongo 副本集模式。
### 无法导出知识库、无法使用语音输入/播报
没配置 SSL 证书,无权使用部分功能。
### 登录提示 Network Error
由于服务初始化错误,系统重启导致。
* 90%是由于配置文件写不对,导致 JSON 解析报错
* 剩下的基本是因为向量数据库连不上
### 如何修改密码
修改 `docker-compose.yml` 文件中 `DEFAULT_ROOT_PSW` 并重启即可,密码会自动更新。
### 部署 Zilliz 版本,获取账号和密钥
打开 [Zilliz Cloud](https://zilliz.com.cn/) , 创建实例并获取相关秘钥。

1. 修改 `MILVUS_ADDRESS` 和 `MILVUS_TOKEN` 链接参数,分别对应 `zilliz` 的 `Public Endpoint` 和 `Api key`,记得把自己 ip 加入白名单。
file: ./content/self-host/deploy/sealos.en.mdx
meta: {
"title": "Deploy with Sealos",
"description": "One-click FastGPT deployment using Sealos"
}
import { Alert } from '@/components/docs/Alert';
## Deployment Architecture

## Multi-Model Support
FastGPT uses the one-api project to manage model pools, supporting OpenAI, Azure, mainstream domestic models, and local models.
See: [Quick OneAPI Deployment on Sealos](../config/model/intro.en.mdx)
## One-Click Deployment
With Sealos, you don't need to purchase servers or domains. It supports high concurrency and dynamic scaling, and databases use KubeBlocks with far better I/O performance than simple Docker container deployments. Choose a region below based on your needs.
### Singapore Region
Singapore servers are overseas with direct access to OpenAI, but users in mainland China need a VPN. International pricing is slightly higher. Click below to deploy 👇
### Beijing Region
The Beijing region is hosted by Volcano Engine. Users in mainland China get stable access, but it can't reach OpenAI or other overseas services. Pricing is about 1/4 of the Singapore region. Click below to deploy 👇
### 1. Start Deployment
Since databases need to be deployed, wait 2–4 minutes after deployment before accessing. The default uses minimal resources, so the first access may be slow.
Follow the prompts to enter `root_password` and the `openai`/`oneapi` address and key.

After clicking deploy, you'll be redirected to the app management page. Click the details button on the right side of the `fastgpt` main app (named fastgpt-xxxx), as shown below.

After clicking details, you'll see the FastGPT deployment management page. Click the link in the external access address to open the FastGPT service.
To bind a custom domain or modify deployment parameters, click **Change** in the top right and follow Sealos' instructions.

### 2. Log In
Username: `root`
Password: the `root_password` you set during one-click deployment
### 3. Configure Models
### 4. Configure Models
You must configure at least one model set, or the system won't work properly.
[View model configuration tutorial](../config/model/intro.en.mdx)
## Pricing
Sealos uses pay-as-you-go billing based on allocated CPU, memory, and disk. For specific pricing, open the **Cost Center** in the Sealos control panel.
## Using Sealos
### Overview
FastGPT Commercial Edition includes 2 apps (fastgpt, fastgpt-plus) and 2 databases. When using multiple API keys, install OneAPI (1 app and 1 database), totaling 3 apps and 3 databases.

Click details on the right to view each app's information.
### Modifying Config Files and Environment Variables
In Sealos, open **App Launchpad** to see deployed FastGPT apps, and open **Database** to see corresponding databases.
In **App Launchpad**, select FastGPT, click **Change**, and you'll see environment variables and config files.

On Sealos, FastGPT runs 1 service and 2 databases. When pausing or deleting, handle the databases
together. (You can start them during the day and pause at night to save costs.)
### How to Update/Upgrade FastGPT
[Upgrade script documentation](../upgrading/upgrade-intruction.en.mdx) — read the docs first to determine which version to upgrade to. Do not skip versions.
For example, if you're on version 4.5 and want to upgrade to 4.5.1: change the image version to v4.5.1, run the upgrade script, wait for completion, then continue upgrading. If the target version doesn't require initialization, skip it.
Upgrade steps:
1. Check the [update documentation](../upgrading/upgrade-intruction.en.mdx) to confirm the target version — avoid skipping versions.
2. Open Sealos app management
3. There are 2 apps: fastgpt, fastgpt-pro
4. Click the 3 dots on the right side of the app, then **Change**. Or click details, then **Change** in the top right.
5. Modify the image version number

6. Click **Change/Restart** to automatically pull the latest image and update
7. Run the initialization script for the corresponding version (if applicable)
### How to Get the FastGPT Access Link
Open the corresponding app and click the external access address.

### Configure a Custom Domain
Click **Change** on the app -> **Custom Domain** -> enter domain -> configure domain CNAME -> confirm -> confirm change.

### How to Modify Environment Variables
Open Sealos app management -> find the app -> **Change** -> edit environment variables -> click confirm change in the top right.

[Environment Variables](../config/env.en.mdx)
### Modify Site Name and Favicon
Add these environment variables to the app:
```
SYSTEM_NAME=FastGPT
SYSTEM_DESCRIPTION=
SYSTEM_FAVICON=/favicon.ico
HOME_URL=/dashboard/agent
```
SYSTEM\_FAVICON can be a URL.

### Mount a Logo
Currently, the browser logo can't be fully replaced — only SVG is supported. Full replacement will be available after visual customization is implemented.
Add a mounted file with path: `/app/projects/app/public/icon/logo.svg`, with the SVG content as the value.


### Commercial Edition Config File
```
{
"license": "",
"system": {
"title": "" // System name
}
}
```
### Using OneAPI
[See OneAPI usage guide](../config/model/intro.en.mdx)
file: ./content/self-host/deploy/sealos.mdx
meta: {
"title": "Sealos 部署",
"description": "使用 Sealos 一键部署 FastGPT"
}
import { Alert } from '@/components/docs/Alert';
## 部署架构图

## 多模型支持
FastGPT 使用了 one-api 项目来管理模型池,其可以兼容 OpenAI、Azure、国内主流模型和本地模型等。
可参考:[Sealos 快速部署 OneAPI](../config/model/intro.mdx)
## 一键部署
使用 Sealos 服务,无需采购服务器、无需域名,支持高并发 & 动态伸缩,并且数据库应用采用 kubeblocks 的数据库,在 IO 性能方面,远超于简单的 Docker 容器部署。可以根据需求,再下面两个区域选择部署。
### 新加坡区
新加披区的服务器在国外,可以直接访问 OpenAI,但国内用户需要梯子才可以正常访问新加坡区。国际区价格稍贵,点击下面按键即可部署👇
### 北京区
北京区服务提供商为火山云,国内用户可以稳定访问,但无法访问 OpenAI 等境外服务,价格约为新加坡区的 1/4。点击下面按键即可部署👇
### 1. 开始部署
由于需要部署数据库,部署完后需要等待 2\~4 分钟才能正常访问。默认用了最低配置,首次访问时会有些慢。
根据提示,输入 `root_password`,和 `openai` / `oneapi` 的地址和密钥。

点击部署后,会跳转到应用管理页面。可以点击 `fastgpt` 主应用右侧的详情按键(名字为 fastgpt-xxxx),如下图所示。

点击详情后,会跳转到 fastgpt 的部署管理页面,点击外网访问地址中的链接,即可打开 fastgpt 服务。
如需绑定自定义域名、修改部署参数,可以点击右上角变更,根据 sealos 的指引完成。

### 2. 登录
用户名:`root`
密码是刚刚一键部署时设置的 `root_password`
### 3. 配置模型
### 4. 配置模型
务必先配置至少一组模型,否则系统无法正常使用。
[点击查看模型配置教程](../config/model/intro.mdx)
## 收费
Sealos 采用按量计费的方式,也就是申请了多少 cpu、内存、磁盘,就按该申请量进行计费。具体的计费标准,可以打开 `sealos` 控制面板中的 `费用中心` 进行查看。
## Sealos 使用
### 简介
FastGPT 商业版共包含了 2 个应用(fastgpt, fastgpt-plus)和 2 个数据库,使用多 API Key 时候需要安装 OneAPI(一个应用和一个数据库),总计 3 个应用和 3 个数据库。

点击右侧的详情,可以查看对应应用的详细信息。
### 修改配置文件和环境变量
在 Sealos 中,你可以打开 `应用管理`(App Launchpad)看到部署的 FastGPT,可以打开 `数据库`(Database)看到对应的数据库。
在 `应用管理` 中,选中 FastGPT,点击变更,可以看到对应的环境变量和配置文件。

在 Sealos 上,FastGPT 一共运行了 1 个服务和 2
个数据库,如暂停和删除请注意数据库一同操作。(你可以白天启动,晚上暂停它们,省钱大法)
### 如何更新/升级 FastGPT
[升级脚本文档](../upgrading/upgrade-intruction.mdx)先看下文档,看下需要升级哪个版本。注意,不要跨版本升级。
例如,目前是 4.5 版本,要升级到 4.5.1,就先把镜像版本改成 v4.5.1,执行一下升级脚本,等待完成后再继续升级。如果目标版本不需要执行初始化,则可以跳过。
升级步骤:
1. 查看[更新文档](../upgrading/upgrade-intruction.mdx),确认要升级的版本,避免跨版本升级。
2. 打开 sealos 的应用管理
3. 有 2 个应用 fastgpt、fastgpt-pro
4. 点击对应应用右边 3 个点,变更。或者点详情后右上角的变更。
5. 修改镜像的版本号

6. 点击变更/重启,会自动拉取最新镜像进行更新
7. 执行对应版本的初始化脚本(如果有)
### 如何获取 FastGPT 访问链接
打开对应的应用,点击外网访问地址。

### 配置自定义域名
点击对应应用的变更 ->点击自定义域名 ->填写域名 -> 操作域名 Cname -> 确认 -> 确认变。

### 如何修改环境变量
打开 Sealos 的应用管理 -> 找到对应的应用 -> 变更 -> 修改环境变量 -> 点击右上角确认变。

[环境变量说明](../config/env.mdx)
### 修改站点名称以及 favicon
修改应用的环境变量,增加
```
SYSTEM_NAME=FastGPT
SYSTEM_DESCRIPTION=
SYSTEM_FAVICON=/favicon.ico
HOME_URL=/dashboard/agent
```
SYSTEM\_FAVICON 可以是一个网络地址

### 挂载 logo
目前暂时无法把浏览器上的 logo 替换。仅支持 svg,待后续可视化做了后可以全部替换。新增一个挂载文件,文件名为:/app/projects/app/public/icon/logo.svg,值为 svg 对应的值。


### 商业版镜像配置文件
```
{
"license": "",
"system": {
"title": "" // 系统名称
}
}
```
### One API 使用
[参考 OneAPI 使用步骤](../config/model/intro.mdx)
file: ./content/self-host/troubleshooting/attention.en.mdx
meta: {
"title": "Troubleshooting Notes",
"description": "FastGPT usage notes"
}
# Troubleshooting Notes
If you encounter issues while using FastGPT, follow the steps below to troubleshoot and resolve them.
## 1. Check Version and Upgrade
Many known issues are fixed in newer releases. Before reporting a problem, verify your version first:
* **Check version:** View the current running version on the FastGPT homepage or in the admin panel.
* **Upgrade recommendation:** If you are not on the latest version, follow the [Upgrade Guide](../upgrading/upgrade-intruction) to update to the latest stable release.
## 2. Troubleshooting Steps
If the issue still exists after upgrading, check in this order:
* **Check logs:** Review Docker container logs or server logs and locate the specific error stack.
* **Clear cache:** Clear browser cache or retry in incognito mode.
* **Environment check:** Ensure MongoDB and PostgreSQL/Milvus connections are healthy and API keys are valid.
## 3. Prevent Spoofed Client IPs Behind a Reverse Proxy
FastGPT reads the client IP for IP rate limiting, share-link IP allowlists, chat log IP records, and IP geolocation. If your self-hosted FastGPT is behind Nginx, a load balancer, an Ingress controller, or a CDN, make sure clients cannot spoof `X-Forwarded-For` or `X-Real-IP` headers.
Recommended setup:
* **Overwrite incoming IP headers in Nginx:** the last reverse proxy should not pass through a user-supplied `X-Forwarded-For` header. It should overwrite the header with the real connection source.
* **Enable trusted proxy validation in FastGPT:** trust forwarded IP headers only when they come from Nginx, the load balancer, or the Ingress controller.
* **Restrict direct access to FastGPT:** firewall or security group rules should allow only the reverse proxy to access the FastGPT service port.
FastGPT environment variable example:
```dotenv
TRUSTED_PROXY_ENABLE=true
TRUSTED_PROXY_IPS=172.18.0.0/16
```
`TRUSTED_PROXY_IPS` should contain the previous-hop proxy IP or CIDR that FastGPT sees directly, such as the Docker subnet for the Nginx container, the Ingress Controller private address, or the load balancer origin address. Do not use a trust-all CIDR such as `0.0.0.0/0` because it would trust every source, and do not add normal client networks to the trusted list.
For a single Nginx layer exposed directly to users, use:
```nginx
server {
listen 80;
server_name fastgpt.example.com;
location / {
proxy_pass http://fastgpt:3000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
}
```
If a CDN or load balancer is in front of Nginx, configure Nginx to trust only those upstream egress IPs first. Then forward the restored client IP to FastGPT:
```nginx
server {
listen 80;
server_name fastgpt.example.com;
# Add only your CDN or load balancer egress IP/CIDR ranges. Do not trust every source.
set_real_ip_from 10.0.0.0/8;
set_real_ip_from 172.16.0.0/12;
real_ip_header X-Forwarded-For;
real_ip_recursive on;
location / {
proxy_pass http://fastgpt:3000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
}
```
If your CDN uses a dedicated real-IP header, use that header in `real_ip_header` and set `set_real_ip_from` to the official egress IP ranges published by the CDN. Cloudflare uses `CF-Connecting-IP` as one example.
After updating Nginx, run:
```bash
nginx -t && nginx -s reload
```
You can verify the setup with spoofed headers:
```bash
curl -H 'X-Forwarded-For: 6.6.6.6' -H 'X-Real-IP: 6.6.6.6' https://fastgpt.example.com
```
If the configuration is correct, FastGPT should still record and validate the real client IP, not the spoofed value `6.6.6.6` from the request.
## 4. Contact Technical Support
If the issue still cannot be resolved, contact us through:
* **Community feedback:** Search for similar issues in GitHub Issues or community channels.
* **Provide details:** When contacting support, include:
* Full version number currently in use.
* Detailed issue description with reproduction steps.
* Related system error logs or screenshots.
file: ./content/self-host/troubleshooting/attention.mdx
meta: {
"title": "排查注意",
"description": "FastGPT注意事项"
}
# 注意事项
在使用 FastGPT 过程中遇到问题时,请参考以下步骤进行排查和解决。
## 1. 版本检查与升级
很多已知问题已在最新版本中得到修复。在反馈问题前,请务必确认您的版本情况:
* **检查版本**:在 FastGPT 首页或管理后台查看当前运行的版本号。
* **升级建议**:如果当前不是最新版本,建议先参考 [更新指南](../upgrading/upgrade-intruction) 升级至最新稳定版。
## 2. 问题排查步骤
若升级后问题依然存在,请按以下顺序排查:
* **查看日志**:检查 Docker 容器或服务器日志,寻找具体的错误报错信息(Error Stack)。
* **清理缓存**:尝试清理浏览器缓存或使用无痕模式重新访问。
* **环境检查**:确认数据库(MongoDB, PostgreSQL/Milvus)连接是否正常,以及 API 密钥是否有效。
## 3. 反向代理客户端 IP 防伪造
FastGPT 会在 IP 限流、分享链接 IP 白名单、对话日志 IP 记录、IP 属地展示等场景读取客户端 IP。自部署时如果 FastGPT 前面有 Nginx、负载均衡、Ingress 或 CDN,需要避免客户端伪造 `X-Forwarded-For` 或 `X-Real-IP` 请求头。
推荐同时完成以下配置:
* **Nginx 覆盖外部传入的 IP 请求头**:最后一层反向代理不要透传用户原始 `X-Forwarded-For`,而是用真实连接来源覆盖。
* **FastGPT 开启可信代理校验**:只信任来自 Nginx、负载均衡或 Ingress 的转发头,不信任普通客户端直连请求里的 IP 头。
* **限制 FastGPT 端口暴露范围**:防火墙或安全组只允许反向代理访问 FastGPT 服务端口,避免用户绕过 Nginx 直连 FastGPT。
FastGPT 环境变量示例:
```dotenv
TRUSTED_PROXY_ENABLE=true
TRUSTED_PROXY_IPS=172.18.0.0/16
```
`TRUSTED_PROXY_IPS` 需要填写 FastGPT 直接看到的上一跳代理 IP 或 CIDR,例如 Nginx 容器所在 Docker 网段、Ingress Controller 内网地址或负载均衡回源地址。不要填写 `0.0.0.0/0`,也不要把普通客户端网段加入可信列表。
单层 Nginx 直接对外时,可参考:
```nginx
server {
listen 80;
server_name fastgpt.example.com;
location / {
proxy_pass http://fastgpt:3000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
}
```
如果 Nginx 前面还有 CDN 或负载均衡,需要先让 Nginx 只信任这些上游的出口 IP,再把还原后的真实客户端 IP 转发给 FastGPT:
```nginx
server {
listen 80;
server_name fastgpt.example.com;
# 只填写你的 CDN 或负载均衡出口 IP/CIDR,不要信任所有来源。
set_real_ip_from 10.0.0.0/8;
set_real_ip_from 172.16.0.0/12;
real_ip_header X-Forwarded-For;
real_ip_recursive on;
location / {
proxy_pass http://fastgpt:3000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $remote_addr;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
}
```
如果 CDN 使用专用真实 IP 头,例如 `CF-Connecting-IP`,需要把 `real_ip_header` 改成对应头名,并把 `set_real_ip_from` 配置为该 CDN 官方公布的出口 IP 段。
修改完成后,执行:
```bash
nginx -t && nginx -s reload
```
可以用伪造头验证配置是否生效:
```bash
curl -H 'X-Forwarded-For: 6.6.6.6' -H 'X-Real-IP: 6.6.6.6' https://fastgpt.example.com
```
如果配置正确,FastGPT 记录和校验的仍应是真实客户端 IP,而不是 `6.6.6.6`。
## 4. 联系技术支持
若以上步骤均无法解决您的问题,请通过以下方式联系我们:
* **社区反馈**:在 GitHub Issues 或相关社群中搜索类似问题。
* **提供信息**:联系技术人员时,请务必提供:
* 当前使用的完整版本号。
* 问题的详细描述(包括复现步骤)。
* 相关的系统错误日志或截图。
file: ./content/self-host/troubleshooting/faq.en.mdx
meta: {
"title": "General Troubleshooting",
"description": "FastGPT Self-Hosting General Troubleshooting"
}
### Frontend Page Crash
1. 90% of cases are due to incorrect model configuration: ensure that at least one model is enabled for each category; check if some `object` parameters in the model are abnormal (arrays and objects). If empty, try giving an empty array or empty object.
2. A small part is due to browser compatibility issues. Since the project contains some high-level syntax, lower version browsers may not be compatible. You can provide specific operation steps and error information in the console to the issue.
3. Turn off the browser translation function. If the browser has translation enabled, it may cause the page to crash.
***
### If deployed via sealos, are there no limitations of local deployment?

This is the length limit of the indexing model. It is the same regardless of the deployment method, but the configuration of different indexing models is different, and parameters can be modified in the background.
***
### How to mount the Mini Program configuration file
Mount the verification file to the specified location: /app/projects/app/public/xxxx.txt
Then restart. For example:

***
### Database port 3306 is occupied, service startup failed

Change the port mapping to 3307 or similar, for example 3307:3306.
***
### Can it run purely locally?
Yes. You need to prepare the vector model and LLM model.
***
### Other models cannot perform question classification/content extraction
1. Check the logs. If it prompts JSON invalid, not support tool, etc., it means that the model does not support tool calling or function calling. You need to set `toolChoice=false` and `functionCall=false`, and it will default to the prompt mode. Currently, the built-in prompts are only tested for commercial model APIs. Question classification is basically usable, but content extraction is not very good.
2. If the configuration is normal and there are no error logs, it means that the prompt may not be suitable for the model. You can customize the prompt by modifying `customCQPrompt`.
***
### Page Crash
1. Turn off translation.
2. Check if the configuration file is loaded normally. If it is not loaded normally, system information will be missing, and it will cause a null pointer in some operations.
* 95% of cases are incorrect configuration files. It will prompt xxx undefined.
* Prompt `URI malformed`, please Issue feedback specific operations and pages, this is due to special string encoding parsing errors.
3. Some api incompatibility issues (rare).
***
### After enabling content completion, the response speed becomes slow
1. Question completion requires a round of AI generation.
2. 3\~5 rounds of queries will be performed. If the database performance is insufficient, there will be a significant impact.
***
### Normal reply in the page, API error
The page uses stream=true mode, so the API also needs to set stream=true for testing. Some model interfaces (mostly domestic) are a bit garbage in non-Stream compatibility.
Same as the previous question, curl test.
***
### Knowledge base indexing has no progress/indexing is very slow
First look at the log error information. There are several situations:
1. Can verify, but indexing has no progress: vector model (vectorModels) is not configured.
2. Cannot verify, nor index: API call failed. Maybe not connected to OneAPI or OpenAI.
3. Has progress, but very slow: api key is not good, OpenAI free account, only 3 times or 60 times a minute. 200 times a day limit.
***
### Connection error
Network exception. Domestic servers cannot request OpenAI, check whether the connection with the AI model is normal.
Or FastGPT cannot request OneAPI (not in the same network).
***
### How to change the root password
Modify the `DEFAULT_ROOT_PSW` environment variable, and then restart FastGPT.
***
file: ./content/self-host/troubleshooting/faq.mdx
meta: {
"title": "通用问题排查",
"description": "FastGPT 私有部署常见问题排查方式"
}
### 前端页面崩溃
1. 90% 情况是模型配置不正确:确保每类模型都至少有一个启用;检查模型中一些 `对象` 参数是否异常(数组和对象),如果为空,可以尝试给个空数组或空对象。
2. 少部分是由于浏览器兼容问题,由于项目中包含一些高阶语法,可能低版本浏览器不兼容,可以将具体操作步骤和控制台中错误信息提供 issue。
3. 关闭浏览器翻译功能,如果浏览器开启了翻译,可能会导致页面崩溃。
***
### 通过 sealos 部署的话,是否没有本地部署的一些限制?

这是索引模型的长度限制,通过任何方式部署都一样的,但不同索引模型的配置不一样,可以在后台修改参数。
***
### 怎么挂载小程序配置文件
将验证文件,挂载到指定位置:/app/projects/app/public/xxxx.txt
然后重启。例如:

***
### 数据库 3306 端口被占用了,启动服务失败

把端口映射改成 3307 之类的,例如 3307:3306。
***
### 能否纯本地运行
可以。需要准备好向量模型和 LLM 模型。
***
### 其他模型没法进行问题分类/内容提取
1. 看日志。如果提示 JSON invalid,not support tool 之类的,说明该模型不支持工具调用或函数调用,需要设置 `toolChoice=false` 和 `functionCall=false`,就会默认走提示词模式。目前内置提示词仅针对了商业模型 API 进行测试。问题分类基本可用,内容提取不太行。
2. 如果已经配置正常,并且没有错误日志,则说明可能提示词不太适合该模型,可以通过修改 `customCQPrompt` 来自定义提示词。
***
### 页面崩溃
1. 关闭翻译
2. 检查配置文件是否正常加载,如果没有正常加载会导致缺失系统信息,在某些操作下会导致空指针。
* 95%情况是配置文件不对。会提示 xxx undefined
* 提示 `URI malformed`,请 Issue 反馈具体操作和页面,这是由于特殊字符串编码解析报错。
3. 某些 API 不兼容问题(较少)
***
### 开启内容补全后,响应速度变慢
1. 问题补全需要经过一轮 AI 生成。
2. 会进行 3\~5 轮的查询,如果数据库性能不足,会有明显影响。
***
### 页面中可以正常回复,API 报错
页面中是用 stream=true 模式,所以 API 也需要设置 stream=true 来进行测试。部分模型接口(国产居多)非 Stream 的兼容有点垃圾。和上一个问题一样,curl 测试。
***
### 知识库索引没有进度/索引很慢
先看日志报错信息。有以下几种情况:
1. 可以对话,但是索引没有进度:没有配置向量模型(vectorModels)
2. 不能对话,也不能索引:API 调用失败。可能是没连上 OneAPI 或 OpenAI
3. 有进度,但是非常慢:API key 不行,OpenAI 的免费号,一分钟只有 3 次还是 60 次。一天上限 200 次。
***
### Connection error
网络异常。国内服务器无法请求 OpenAI,自行检查与 AI 模型的连接是否正常。
或者是 FastGPT 请求不到 OneAPI(没放同一个网络)
***
### 修改了 vectorModels 但是没有生效
1. 重启容器,确保模型配置已经加载(可以在日志或者新建知识库时候看到新模型)
2. 记得刷新一次浏览器。
3. 如果是已经创建的知识库,需要删除重建。向量模型是创建时候绑定的,不会动态更新。
***
### 如何修改 root 密码
修改环境变量中的 `DEFAULT_ROOT_PSW`,然后重启 FastGPT。
***
file: ./content/self-host/troubleshooting/methods.en.mdx
meta: {
"title": "Troubleshooting Methods",
"description": "FastGPT Self-Hosting Common Troubleshooting Methods"
}
## 1. Troubleshooting Methods
You can first look for [Issue](https://github.com/labring/FastGPT/issues), or raise a new Issue. For private deployment errors, be sure to provide detailed operation steps, logs, and screenshots, otherwise it is difficult to troubleshoot.
### (1) Get Backend Errors
1. `docker ps -a` View the running status of all containers, check if they are all running. If there is an abnormality, try `docker logs container_name` to view the corresponding log.
2. If the containers are running normally, `docker logs container_name` to view the error log.
***
### (2) Frontend Errors
When a frontend error occurs, the page will crash and prompt to check the console log. You can open the browser console and view the log in `console`. You can also click the hyperlink of the corresponding log, which will prompt to the specific error file. You can provide these detailed error information to facilitate troubleshooting.
***
file: ./content/self-host/troubleshooting/methods.mdx
meta: {
"title": "错误排查方式",
"description": "FastGPT 私有部署常见问题排查方式"
}
## 一、错误排查方式
可以先找找[Issue](https://github.com/labring/FastGPT/issues),或新提 Issue,私有部署错误,务必提供详细的操作步骤、日志、截图,否则很难排查。
### (1)获取后端错误
1. `docker ps -a` 查看所有容器运行状态,检查是否全部 running,如有异常,尝试`docker logs 容器名`查看对应日志。
2. 容器都运行正常的,`docker logs 容器名` 查看报错日志
***
### (2)前端错误
前端报错时,页面会出现崩溃,并提示检查控制台日志。可以打开浏览器控制台,并查看`console`中的 log 日志。还可以点击对应 log 的超链接,会提示到具体错误文件,可以把这些详细错误信息提供,方便排查。
***
file: ./content/self-host/troubleshooting/model-errors.en.mdx
meta: {
"title": "Model Troubleshooting",
"description": "FastGPT Self-Hosting Model Troubleshooting"
}
### (1) How to check model availability issues
1. For privately deployed models, first confirm whether the deployed model is normal.
2. Directly test whether the upstream model is running normally through CURL request (cloud model or private model are both tested).
3. Request OneAPI through CURL request to test whether the model is normal.
4. Use the model for testing in FastGPT.
Here are a few test CURL examples:
```bash
curl https://api.openai.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello!"
}
]
}'
```
```bash
curl https://api.openai.com/v1/embeddings \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": "The food was delicious and the waiter...",
"model": "text-embedding-ada-002",
"encoding_format": "float"
}'
```
```bash
curl --location --request POST 'https://xxxx.com/api/v1/rerank' \
--header 'Authorization: Bearer {{ACCESS_TOKEN}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "bge-rerank-m3",
"query": "Who is the director",
"documents": [
"Who are you?\nI am the assistant of the movie 'Suzume'"
]
}'
```
```bash
curl https://api.openai.com/v1/audio/speech \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "The quick brown fox jumped over the lazy dog.",
"voice": "alloy"
}' \
--output speech.mp3
```
```bash
curl https://api.openai.com/v1/audio/transcriptions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: multipart/form-data" \
-F file="@/path/to/file/audio.mp3" \
-F model="whisper-1"
```
***
### (2) Error - Model response is empty/Model error
This error is due to the fact that under stream mode, oneapi directly ended the stream request and did not return any content.
Version 4.8.10 added error logs. When an error occurs, the actual Body parameters sent will be printed in the log. You can copy the parameters and send a request test to oneapi through curl.
Since oneapi cannot correctly capture errors in stream mode, sometimes you can set `stream=false` to get the exact error.
Possible error issues:
1. Domestic models hit risk control.
2. Unsupported model parameters: only keep messages and necessary parameters for testing, delete other parameters for testing.
***
### (3) "Current group upstream load is saturated, please try again later"
If you encounter this error (e.g. `request id:xxx`) in the logs or response, this is typically an OneAPI channel issue. Try switching to a different model or a different relay provider.
***
### (4) "Connection Error" in logs when using the API
Most likely the API key is pointing to OpenAI's endpoint, but the server is deployed in mainland China and can't reach overseas endpoints. Use a relay service or reverse proxy to resolve the connectivity issue.
***
### (5) Enable image indexing reports 400
You need to correctly configure the OCR model in `Admin` -> `System Configuration`.
file: ./content/self-host/troubleshooting/model-errors.mdx
meta: {
"title": "模型问题排查",
"description": "FastGPT 私有部署模型问题排查"
}
### (1)如何检查模型可用性问题
1. 私有部署模型,先确认部署的模型是否正常。
2. 通过 CURL 请求,直接测试上游模型是否正常运行(云端模型或私有模型均进行测试)
3. 通过 CURL 请求,请求 OneAPI 去测试模型是否正常。
4. 在 FastGPT 中使用该模型进行测试。
下面是几个测试 CURL 示例:
```bash
curl https://api.openai.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"model": "gpt-4o",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello!"
}
]
}'
```
```bash
curl https://api.openai.com/v1/embeddings \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": "The food was delicious and the waiter...",
"model": "text-embedding-ada-002",
"encoding_format": "float"
}'
```
```bash
curl --location --request POST 'https://xxxx.com/api/v1/rerank' \
--header 'Authorization: Bearer {{ACCESS_TOKEN}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "bge-rerank-m3",
"query": "导演是谁",
"documents": [
"你是谁?\n我是电影《铃芽之旅》助手"
]
}'
```
```bash
curl https://api.openai.com/v1/audio/speech \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "The quick brown fox jumped over the lazy dog.",
"voice": "alloy"
}' \
--output speech.mp3
```
```bash
curl https://api.openai.com/v1/audio/transcriptions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: multipart/form-data" \
-F file="@/path/to/file/audio.mp3" \
-F model="whisper-1"
```
***
### (2)报错 - 模型响应为空/模型报错
该错误是由于 stream 模式下,oneapi 直接结束了流请求,并且未返回任何内容导致。
4.8.10 版本新增了错误日志,报错时,会在日志中打印出实际发送的 Body 参数,可以复制该参数后,通过 curl 向 oneapi 发起请求测试。
由于 oneapi 在 stream 模式下,无法正确捕获错误,有时候可以设置成 `stream=false` 来获取到精确的错误。
可能的报错问题:
1. 国内模型命中风控
2. 不支持的模型参数:只保留 messages 和必要参数来测试,删除其他参数测试。
3. 参数不符合模型要求:例如有的模型 temperature 不支持 0,有些不支持两位小数。max\_tokens 超出,上下文超长等。
4. 模型部署有问题,stream 模式不兼容。
测试示例如下,可复制报错日志中的请求体进行测试:
```bash
curl --location --request POST 'https://api.openai.com/v1/chat/completions' \
--header 'Authorization: Bearer sk-xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "xxx",
"temperature": 0.01,
"max_tokens": 1000,
"stream": true,
"messages": [
{
"role": "user",
"content": " 你是饿"
}
]
}'
```
***
### (3)如何测试模型是否支持工具调用
需要模型提供商和 oneapi 同时支持工具调用才可使用,测试方法如下:
##### 1. 通过 `curl` 向 `oneapi` 发起第一轮 stream 模式的 tool 测试。
```bash
curl --location --request POST 'https://oneapi.xxx/v1/chat/completions' \
--header 'Authorization: Bearer sk-xxxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "gpt-5",
"temperature": 0.01,
"max_tokens": 8000,
"stream": true,
"messages": [
{
"role": "user",
"content": "几点了"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "hCVbIY",
"description": "获取用户当前时区的时间。",
"parameters": {
"type": "object",
"properties": {},
"required": []
}
}
}
],
"tool_choice": "auto"
}'
```
##### 2. 检查响应参数
如果能正常调用工具,会返回对应 `tool_calls` 参数。
```json
{
"id": "chatcmpl-A7kwo1rZ3OHYSeIFgfWYxu8X2koN3",
"object": "chat.completion.chunk",
"created": 1726412126,
"model": "gpt-5",
"system_fingerprint": "fp_483d39d857",
"choices": [
{
"index": 0,
"id": "call_0n24eiFk8OUyIyrdEbLdirU7",
"type": "function",
"function": {
"name": "mEYIcFl84rYC",
"arguments": ""
}
}
],
"refusal": null
},
"logprobs": null,
"finish_reason": null
}
],
"usage": null
}
```
##### 3. 通过 `curl` 向 `oneapi` 发起第二轮 stream 模式的 tool 测试。
第二轮请求是把工具结果发送给模型。发起后会得到模型回答的结果。
```bash
curl --location --request POST 'https://oneapi.xxxx/v1/chat/completions' \
--header 'Authorization: Bearer sk-xxx' \
--header 'Content-Type: application/json' \
--data-raw '{
"model": "gpt-5",
"temperature": 0.01,
"max_tokens": 8000,
"stream": true,
"messages": [
{
"role": "user",
"content": "几点了"
},
{
"role": "assistant",
"tool_calls": [
{
"id": "kDia9S19c4RO",
"type": "function",
"function": {
"name": "hCVbIY",
"arguments": "{}"
}
}
]
},
{
"tool_call_id": "kDia9S19c4RO",
"role": "tool",
"name": "hCVbIY",
"content": "{\n \"time\": \"2024-09-14 22:59:21 Sunday\"\n}"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "hCVbIY",
"description": "获取用户当前时区的时间。",
"parameters": {
"type": "object",
"properties": {},
"required": []
}
}
}
],
"tool_choice": "auto"
}'
```
***
### (4)向量检索得分大于 1
由于模型没有归一化导致的。目前仅支持归一化的模型。
***
### (5) 当前分组上游负载已饱和,请稍后再试
如果在日志或请求中遇到此错误(如 `request id:xxx`),这通常是 OneAPI 渠道的问题,可以换个模型使用或者换一家中转站。
***
### (6) 使用API时在日志中报错 Connection Error
大概率是 API Key 填写了 OpenAI 的地址,但是部署的服务器在国内,不能访问海外的 API。可以使用中转或者反代的手段解决访问不到的问题。
***
### (7) 开启图片索引报 400
需在 `Admin` -> `系统配置` 中正确配置 OCR 模型。
file: ./content/self-host/troubleshooting/s3-issues.en.mdx
meta: {
"title": "S3 Issues Troubleshooting",
"description": "FastGPT Self-Hosting Common S3 Issues Troubleshooting"
}
## 1. Log shows ERR level "Failed to ensure external public/private bucket exists", resulting in inability to connect to object storage
### 1.1 Error Stack Display
* error: Error: getaddrinfo ENOTFOUND
Example
* 
### Possible Errors
* STORAGE\_S3\_FORCE\_PATH\_STYLE configuration error
### Solution
* Turn on the STORAGE\_S3\_FORCE\_PATH\_STYLE option to `true`, otherwise the client cannot find the target service
***
## 2. Upload conversation file / knowledge base file error
Example
* 
### 2.1 SignatureDoesNotMatched
* Signature inconsistency, mostly due to Nginx configuration error
### Possible Errors
* Necessary request headers (such as Headers, Host) were not passed during Nginx forwarding
### Solution
* Configure proxy\_set\_header Host $http\_host, do not set to $host, Nginx's $host built-in variable will remove the port, set to $http\_host
***
file: ./content/self-host/troubleshooting/s3-issues.mdx
meta: {
"title": "存储桶问题排查",
"description": "FastGPT 私有部署存储桶问题排查方式"
}
## 1. 日志出现 ERR 等级的 “Failed to ensure external public/private bucket exists”,导致无法连接上对象存储
### 1.1 错误栈显示
* error: Error: getaddrinfo ENOTFOUND
示例
* 
### 可能的错误
* STORAGE\_S3\_FORCE\_PATH\_STYLE 配置错误
### 解决
* 将 STORAGE\_S3\_FORCE\_PATH\_STYLE 选项开启为 `true`,否则客户端无法找到目标服务
***
## 2. 上传对话文件 / 知识库文件报错
示例
* 
### 2.1 SignatureDoesNotMatched
* 签名不一致,大部分情况是因为 Nginx 配置错误
### 可能的错误
* Nginx 转发时未透传必要的请求头(如 Headers、Host)
### 解决
* 配置 proxy\_set\_header Host $http\_host,不要设置成 $host,Nginx 的 $host内置变量会把端口去掉,要设置成 $http\_host
***
file: ./content/self-host/upgrading/upgrade-intruction.en.mdx
meta: {
"title": "Upgrade Guide",
"description": "FastGPT version upgrade guide"
}
Upgrading FastGPT involves two steps:
1. Update the image
2. Run the upgrade initialization script
## Image Names
**GitHub Container Registry**
* FastGPT main image: ghcr.io/labring/fastgpt:latest
* Plugin image: ghcr.io/labring/fastgpt-plugin
* Code sandbox image: ghcr.io/labring/fastgpt-sandbox
* MCP SSE server image: ghcr.io/labring/fastgpt-mcp\_server
* Commercial edition image: ghcr.io/c121914yu/fastgpt-pro:latest
**Alibaba Cloud**
* FastGPT main image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt
* Plugin image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-plugin
* Code sandbox image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-sandbox
* MCP SSE server image: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-mcp\_server
* Commercial edition image: ghcr:registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-pro
An image consists of the image name and a `Tag`. For example, registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt:v4.6.1 refers to the `4.6.3` version image. Check Docker Hub or the GitHub repository for details.
## Updating Images on Sealos
1. Open [Sealos Cloud](https://cloud.sealos.io?uid=fnWRt09fZP) and find App Management on the desktop.

2. Select the corresponding app - click the three dots on the right - Update.

3. Update the image - Confirm changes.
If you need to modify the configuration file, scroll down to the `Configuration File` section to make changes.

## Updating Images with Docker Compose
Simply modify the `image:` field in the `yml` file, then run:
```bash
docker-compose pull
docker-compose up -d
```
## Running the Upgrade Initialization Script
After updating the images, check the version notes in the documentation. Versions that require an upgrade script are typically labeled with "includes upgrade script". Open the corresponding documentation and follow the instructions to run the **upgrade script** -- in most cases, you just need to send a `POST` request.
## FAQ
### Why do I need to run an upgrade script?
When there are significant changes to the database schema that cannot be handled through default values, or when the migration logic is complex, an upgrade script is used to update certain database fields.
Following the initialization steps carefully will not cause any data loss. However, if the data volume is large, the initialization may take a while, during which the service may be temporarily unavailable.
### What is `{{host}}`?
`{{}}` denotes a variable. `{{host}}` refers to a variable named "host", which is your server's domain name or IP address.
On Sealos, you can find your domain name as shown below:

### How to get the rootkey
You can find it in the `environment` section of your `docker-compose.yml` file -- it's the value of `ROOT_KEY`.
On Sealos, you can find it in the environment variables panel shown in the image above.
### How to upgrade across multiple versions
Back up your data first!
You can upgrade to the latest version and then run all the upgrade scripts in order. However, for stability, we recommend upgrading one version at a time. For example, if your current version is 4.4.7 and you need to upgrade to 4.6:
1. Update the image to 4.5, run the upgrade script
2. Update the image to 4.5.1, run the upgrade script
3. Update the image to 4.5.2, run the upgrade script
4. Update the image to 4.6, run the upgrade script
5. .....
Upgrade one version at a time.
file: ./content/self-host/upgrading/upgrade-intruction.mdx
meta: {
"title": "版本&升级说明",
"description": "FastGPT 版本&升级说明"
}
## 版本说明
从 4.14.11 开始,为了区分稳定版和快速迭代版,对版本命名进行了调整,未来将按以下方式进行版本命名:
1. 维护 2 个稳定版本。例如当前迭代功能处于 4.16.x 版本,则会维护 4.14.x 和 4.15.x 两个文档版本。
2. 稳定版本命名不带后缀,例如:4.14.11, 4.14.12, 4.15.0, 4.15.1。如果 4.14.11 有问题,会修复后发布 4.14.12,并同步修复到 4.15.x 的稳定版,以确保修复问题同时不引入新的功能。
3. 快速迭代版本命名带后缀,例如:4.16.0-beta.1, 4.16.0-beta.2, 4.16.0-beta.3。
4. 迭代版本约 2 个月发布一次稳定版,并且会提供一个聚合的升级脚本,用户只需要执行一次请求,即可完成所有迭代版本的升级。
总结来说,后续用户可以直接升级不带 beta 后缀的稳定版本,以确保稳定性,官方会单独发布修复版本并确保不会引入新功能。
## 升级说明
FastGPT 升级通常包括两个步骤:
1. 修改镜像名
2. 执行升级初始化脚本
## 镜像名
**git版**
* FastGPT 主镜像: ghcr.io/labring/fastgpt:latest
* Plugin 镜像: ghcr.io/labring/fastgpt-plugin
* 代码沙箱镜像: ghcr.io/labring/fastgpt-code-sandbox
* MCP SSE setver 镜像: ghcr.io/labring/fastgpt-mcp\_server
* 商业版镜像: ghcr.io/c121914yu/fastgpt-pro:latest
**阿里云**
* FastGPT 主镜像: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt
* Plugin 镜像: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-plugin
* 代码沙箱镜像: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-code-sandbox
* MCP SSE setver 镜像: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-mcp\_server
* 商业版镜像: ghcr:registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt-pro
镜像由镜像名和`Tag`组成,例如: registry.cn-hangzhou.aliyuncs.com/fastgpt/fastgpt:v4.6.1 代表`4.6.3`版本镜像,具体可以看 docker hub, github 仓库。
## Sealos 修改镜像
1. 打开 [Sealos Cloud](https://cloud.sealos.io?uid=fnWRt09fZP), 找到桌面上的应用管理

2. 选择对应的应用 - 点击右边三个点 - 变更

3. 修改镜像 - 确认变更
如果要修改配置文件,可以拉到下面的`配置文件`进行修改。

## Docker-Compose 修改镜像
直接修改`yml`文件中的`image: `即可。随后执行:
```bash
docker-compose pull
docker-compose up -d
```
## 执行升级初始化脚本
镜像更新完后,可以查看文档中的`版本介绍`,通常需要执行升级脚本的版本都会标明`包含升级脚本`,打开对应的文档,参考说明执行**升级脚本**即可,大部分时候都是需要发送一个`POST`请求。
## QA
### 为什么需要执行升级脚本
数据表出现大幅度变更,无法通过设置默认值,或复杂度较高时,会通过升级脚本来更新部分数据表字段。
严格按初始化步骤进行操作,不会造成旧数据丢失。但在初始化过程中,如果数据量大,需要初始化的时间较长,这段时间可能会造成服务无法正常使用。
### `{{host}}` 是什么
`{{}}` 代表变量, `{{host}}`代表一个名为 host 的变量。指的是你服务器的域名或 IP。
Sealos 中,你可以在下图中找到你的域名:

### 如何获取 rootkey
从`docker-compose.yml`中的`environment`中获取,对应的是`ROOT_KEY`的值。
sealos 中可以从上图左侧的环境变量中获取。
### 如何跨版本升级!!
先进行数据备份!!!
可以升级至最新版本,然后将所有升级脚本的版本都执行一遍。不过为了稳定,建议逐一版本升级。例如,当前版本是4.4.7,需要升级到4.6。
1. 修改镜像到4.5,执行升级脚本
2. 修改镜像到4.5.1,执行升级脚本
3. 修改镜像到4.5.2,执行升级脚本
4. 修改镜像到4.6,执行升级脚本
5. .....
逐一升级
file: ./content/guide/build/agentv2/debug.en.mdx
meta: {
"title": "Assisted Generation and Debugging",
"description": "Detailed guide to generating Prompts via the AI Helper Bot and testing and debugging the Agent using the Chat Preview window and runtime details."
}
## Assisted generation
An AI Helper Bot optimized for Agent V2 is integrated under the "Assisted generation" tab.

### 1. Smart Analysis and Scheme Optimization
The AI Helper Bot automatically reads the current model, datasets, and candidate tools configured in your left panel. You only need to describe your expectations in natural language (e.g., "I want an assistant that analyzes sales data uploaded in Excel and generates visual charts"), and the Helper Bot will automatically:
* **Generate and Refine Structured System Prompts**, setting reasonable boundaries and execution rules.
* **Recommend the Most Suitable Tools** or virtual machine commands for the task.
* **Suggest Relevant Datasets** to supplement domain background knowledge.
### 2. Automatic Apply
Once the Helper Bot generates the optimized scheme, the system will automatically write back the recommended prompts, tool list, associated datasets, and sandbox toggles to the left configuration panel. The entire process is fully automated without requiring manual clicks, streamlining the setup workflow.
***
## Chat Preview
The "Chat Preview" tab provides a real-time conversational testing environment, allowing you to interact with the Agent as a real user and verify the application's effectiveness before publishing.
### 1. General Debugging
Regardless of whether the virtual machine is enabled, you can use the following general debugging features:

* **Restart**: Click the "Restart" button in the upper right corner of the chat preview window to clear the current chat history and state, allowing you to start a fresh round of testing.
* **Bubble Action Bar and Runtime Details**: Below each generated AI chat bubble, a row of auxiliary debugging tools is provided:
* **Copy**: Click to copy the text content of this AI response.
* **Read Aloud**: Click to convert the response text into speech and play it back.
* **Mark**: Allows developers to mark the question and expected answer, saving it to a designated dataset to correct and guide the model's future responses.
* **Retry**: Triggers the AI to regenerate the response for the last user input.
* **Runtime Details**: Click to expand the tree-structured decision-making chain. This logs the complete LLM reasoning process, internal plan updates, and tool execution logs. It also shows the unique Request ID, model model, response duration, and precise points consumption for performance auditing and cost control.
* **Response Duration** (e.g., `34.07 s`): Displays the total time in seconds spent from sending the request to receiving the full response.
* **Plan Card**: If the Agent initiates planning for a complex task, the chat interface will stream a visual "Plan Card." Color-coded steps and animations indicate the progress of each step (In Progress, Completed, Pending, Blocked). If a step is blocked, the card displays the cause of the blockage, helping you optimize your System Prompt or troubleshoot tool configurations.

### 2. Virtual Machine Debugging
For details on using the virtual machine sandbox for file management, dependency installations, and self-correcting debugging workflows, please refer to [Virtual Machine](./vm#virtual-machine-debugging).
file: ./content/guide/build/agentv2/debug.mdx
meta: {
"title": "辅助生成与调试",
"description": "详细介绍如何通过 AI 助手辅助生成 Prompt,以及如何利用调试预览窗口、运行详情等功能测试与排查 Agent。"
}
## 辅助生成
在“辅助生成”选项卡中,内置了针对 Agent V2 优化的 AI 协作助手,目前感知的范围有限,仍处于优化阶段,仅商业版开放。

### 1. 智能分析与方案优化
AI 协作助手能够综合读取您当前在左侧面板配置的模型、知识库及候选工具。您只需在对话框中以自然语言输入您的期望(例如:“我想要一个能够分析用户上传的 Excel 销售数据并直接生成可视化图表的助理”),协作助手将自动:
* **润色并生成结构化的 System Prompt**,设定合理的约束与引导规则。
* **智能筛选并推荐**适合当前任务的外部工具(Tools)或虚拟机命令。
* **推荐适配的知识库**以补充背景知识。
### 2. 配置自动应用
当协作助手根据您的诉求生成最佳方案后,系统会自动将推荐的提示词、工具列表、关联的知识库以及虚拟机开关状态,直接同步并填充到左侧的配置面板中。整个过程完全自动化,无需任何手动点击与二次配置,极大地简化了应用搭建流程。
***
## 调试预览
“调试预览”选项卡提供了一个实时的应用对话测试环境,让您能够像真实用户一样与 Agent 进行交互,在发布前验证应用效果。
### 1. 通用调试
无论是否启用虚拟机,您都可以使用以下通用调试功能:

* **重开对话**:点击调试窗口右上角的“重开对话”按钮,可以一键清空当前的聊天历史,方便您从头开始重新测试。
* **气泡操作栏与运行详情**:在每次对话生成的 AI 气泡下方,提供了一排辅助调试按钮:
* **复制**:点击可一键复制该条 AI 回复的文本内容。
* **朗读**:点击可将该条回复文本转化为语音进行播放。
* **标注**:允许开发者对该条对话的问答数据进行标注并保存至指定的知识库中,通过设定预期回答来纠错并引导模型在后续对话中给出更符合预期的回复。
* **重试**:让 AI 针对上一条用户输入重新生成一遍回复。
* **运行详情**:点击可展开树状的决策链路视图。这里完整记录了本次交互中 AI 的思考历程、内部参数动作和工具调用日志,并展示了单次请求的请求 Id、模型型号及精准的积分消耗,便于性能审计与成本核算。
* **响应耗时**(如 `34.07 s`):直观展示该次交互从发送请求到完全响应所耗费的秒数。
* **步骤规划卡片**:如果 AI 面对复杂任务启动了任务规划,对话中会实时流式渲染“计划步骤卡片”(Plan Card),并通过颜色(进行中、已完成、待处理、已阻塞)直观体现执行节点。若遇到阻塞,您可以根据卡片上的阻塞原因来优化您的 System Prompt 或排查工具配置。

### 2. 虚拟机调试
关于如何使用虚拟机进行调试、查看文件、文件注入以及运行状态查看等专属能力,请参考 [虚拟机](./vm#虚拟机调试)。
file: ./content/guide/build/agentv2/settings.en.mdx
meta: {
"title": "Configuration Panel",
"description": "Detailed configuration guide for LLM parameters, virtual machine environment, and skill integration in Chat Agent V2."
}
The Configuration Panel is used to configure and bind all the core capabilities and execution environments required by your Agent.

***
## AI Configuration & Virtual Machine
* **AI Model**: Choose the dialogue model and configure parameters, for more general LLM parameters, refer to [AI Settings](../general/ai_settings).
* **Prompt**: Define the core persona, objectives, and specific rules for the Agent. The editor supports rich text, and you can type `@` to quickly reference and bind tools, etc.
| | |
| :-------------------------------------------------------: | :---------------------------------------------------------------: |
|  |  |
* **Virtual Machine**: Once enabled, the system assigns a dedicated Linux sandbox environment for each session, supporting code execution, file operations, and startup command configuration. For architecture and debugging details, refer to [Virtual Machine](./vm). For startup script details and execution limits, see [Virtual Machine Lifecycle](./vm#virtual-machine-lifecycle).
***
## Associate SKILL & Tools
* **Associated SKILLs**: Select and bind published SKILL packages from the skill library. The entrypoint script in the SKILL package will execute automatically when the VM spins up. If you associate SKILLs without enabling the VM sandbox, a warning "Virtual Machine Not Ready" will be displayed. To learn how to write and package custom SKILLs, please refer to [Development & Debugging](../skill/development).

* **Tools**: You can choose to bind system built-in tools (e.g., search engines, charts), custom tools created by yourself or the team (including HTTP/MCP tools), or created applications.

***
## Knowledge Base & File Uploads
* **Knowledge Base**: Associate specific corporate documents and adjust search settings (Hybrid Search, Re-ranking, etc.). It also supports configuring team member authorization permissions.
* **File Uploads**: Toggle file uploads for end-users, permitting images, audio, video, or custom file extensions. File upload capabilities automatically adapt based on the multimodal features of the selected LLM. For detailed configurations, see [File Input](../general/fileInput).
file: ./content/guide/build/agentv2/settings.mdx
meta: {
"title": "专项配置",
"description": "详细介绍对话 Agent V2 中提示词、虚拟机与技能的专项配置。"
}
专项配置用于调整 Agent 的核心指令、运行环境与技能扩展能力。

***
## 提示词
提示词用于定义 Agent 的核心人设、工作目标和具体规则。编辑器支持富文本,并支持通过 `@` 快速唤起并绑定部分工具等上下文能力。
| | |
| :---------------------------------------------: | :------------------------------------------------: |
|  |  |
如果需要配置对话大模型、回复长度、推理内容展示等通用模型参数,请参考 [AI 配置说明](../general/ai_settings)。
***
## 虚拟机
虚拟机开启后,系统会为每个独立会话在后台分配一个专属的 Linux 沙盒运行环境,支持执行代码、读写文件及配置启动脚本等。
具体设计与联调操作请参考 [虚拟机](./vm),启动脚本去重与执行细节请参考 [虚拟机生命周期](./vm#虚拟机生命周期)。
***
## 技能
技能配置用于选择已发布的技能插件包。技能自带的入口脚本会随虚拟机拉起自动执行。
如果未开启虚拟机但关联了技能,系统会展示“虚拟机未就绪”的警告。若需了解如何自定义编写与打包技能,请参考 [开发与调试](../skill/development)。

file: ./content/guide/build/agentv2/vm.en.mdx
meta: {
"title": "Virtual Machine",
"description": "Understand the core concepts, runtime design, and debugging workflow for the Agent V2 Virtual Machine (Computer Sandbox)."
}
import { Alert } from '@/components/docs/Alert';

In FastGPT Agent V2, the **Virtual Machine** is a dedicated, physically isolated, and secure lightweight Linux running sandbox environment provisioned for each chat session. It equips the Agent with real-world computation, code execution, and file read/write capabilities, allowing the AI to not only "think" but also execute code to solve complex tasks like a human programmer.
***
## What is Virtual Machine
When you toggle the "Enable Computer" option, the system dynamically provisions and binds a dedicated sandbox container for each individual chat session in the background.

With the virtual machine, the Agent can:
* **Execute Dynamic Code**: Run Python, Node.js, or Shell scripts via the code executor to perform complex calculations and data manipulations.
* **Read and Write Local Files**: Create, modify, and read files in the isolated `/workspace` directory, including generating charts or processing uploaded CSV/Excel sheets.
* **Customize the Environment & Startup Script**: Dynamically customize the runtime environment by binding SKILL packages, or configuring custom [Startup Scripts](#virtual-machine-lifecycle) (which automatically execute specified Shell initialization commands, such as installing dependencies or setting environment variables, after the VM spins up but before the AI workflow starts).
***
## Virtual Machine Design
To balance security, latency, and resource footprint across high-concurrency and multi-tenant environments, the system features the following core designs:
### 1. Session-Level Isolation and Lifecycle Management
* The virtual machine is tightly coupled with the user's chat session. Different users and sessions run in entirely isolated environments.
* The system utilizes a keepalive mechanism to sustain active containers. When a session remains idle for too long (exceeding the configurable timeout, typically a few minutes), the container is automatically collected and destroyed to release host resources.
### 2. Session Persistence
The virtual machine remains active throughout the duration of a chat session. Dependencies downloaded, temporary files written, or environment variables set in the previous turns remain accessible in subsequent turns.
### 3. Security Constraints & Escape Prevention
* Strict resource quotas (CPU, Memory, Disk IO) and network firewall policies are applied to the sandbox to prevent malicious resource exhaustion or internal network access.
* All executed commands are encoded in Base64 and decoded securely inside the container to prevent Command Injection risks from string concatenations.
### 4. Instant Sandbox Reconstruction
Clicking "Restart" during testing will thoroughly destroy the current virtual machine. The next request will provision and initialize a clean, brand-new container to prevent historical files from contaminating the session.
***
## Virtual Machine Debugging
When you enable the **Computer** option in the left configuration panel, the Chat Preview window will automatically unlock the following sandbox-exclusive debugging features:
### Virtual Machine File Manager
Shortcut entries to the VM file manager are provided at the top of the chat window and below chat bubbles that involve VM operations. Clicking them pops up a modal to browse, edit, upload, or download files (such as charts, code, and HTML previews) inside the container, achieving a closed debugging loop.
| | |
| :------------------------------------------------------------: | :--------------------------------------------------------------: |
|  |  |
### Automatic Dialog File Injection
Any files you upload via the chat input box (such as CSV or Excel sheets) are automatically downloaded and written into the VM's `user_files/` directory before execution, allowing the AI to read and process them as local files via code.
### Code Execution and Self-Correction
The virtual machine provides a real execution environment. If code fails due to missing dependencies or syntax errors, the error logs are real-time fed back to the LLM, enabling the AI to self-correct and re-run within multi-turn planning.
### Pre-emptive Initialization and State Persistence
With the Computer option enabled, the system automatically spins up and prepares the sandbox environment before each chat session starts. All files, dependencies, and execution states are persisted across chat steps to support continuous multi-turn debugging.
### Real-time Startup Status Tracking
During testing, the chat bubble header streams real-time VM provisioning updates, allowing you to audit the startup progress and duration.
### Complete Sandbox Reconstruction upon Reset
Clicking the "Restart" button resets the test session, and the next interaction will provision and initialize a clean, brand-new VM container to prevent historical file contamination.
***
## Virtual Machine Lifecycle
When the **Computer** option is enabled, you can configure a **Startup Script** to automatically execute shell commands right after the sandbox environment spins up and before the AI workflow officially starts. This is commonly used for configuring environment variables, modifying software package sources, or installing Python packages (`pip`) and system-level utilities.
### Script Configuration & Lifecycle
Under the "Computer Configuration" section of the Agent Configuration Panel, you can write standard Shell commands directly inside the **Startup script (sh)** code editor.

#### Script Execution Sequence and Scope
When a new session starts or the virtual machine is reconstructed, the system executes the scripts sequentially in the background:
1. **Application Startup Script**: The custom Shell script configured in your Agent panel. It executes inside the virtual machine's working directory (usually `/workspace`) to prepare specific dependencies and runtime environments required by this application.
2. **Skill Entrypoint**: If your Agent is associated with skills, the [initialization entrypoint script](../skill/initialization) (e.g., `entrypoint.sh`) bundled inside the published skill package will be extracted and executed in the skill's deployment directory right after the application startup script completes.
**Transactional Skill Deployment**: During skill deployment, packages are first extracted to a temporary folder (e.g., `.tmp--`). Upon successful decompression, the folder is atomically renamed to the formal version directory to prevent corrupted partial extractions.
#### Lifecycle Flowchart

### Status Deduplication
To prevent latency from running commands repeatedly during subsequent turns (such as reinstalling packages via `pip`), the system employs an efficient **status deduplication mechanism**:
* **Execution State Record**: The system maintains an execution state file inside the sandbox at `~/.fastgpt/agent-skill-entrypoints/state.json`.
* **Hash-based Deduplication (Application Startup Script)**: For your custom "Startup script (sh)", the system computes a **SHA-256 hash value** based on the script text and compares it with the executed hashes in `state.json`. If the script remains unmodified, the system **automatically skips execution** on subsequent requests, ensuring fast starts. The script will only run again if you edit its content or click "Clear Chat" to completely rebuild the sandbox.
* **Version ID-based Deduplication (Skill Entrypoint)**: Associated skills are deduplicated using their immutable skill **Version ID**. Since published skill versions are read-only, the entrypoint script executes only once during the sandbox's cold start as long as the bound version remains unchanged.
* **State Lifecycle**: The deduplication state is managed along with the virtual machine instance. When the virtual machine is rebuilt (due to clicking "Clear Chat" or system reclamation), a fresh environment is allocated, and all scripts will run again during the next cold start.
### Execution Constraints & Fault Tolerance
To ensure sandbox stability and responsiveness, the startup script is subject to the following system rules:
* **Character Length Limit**: Due to front-end validation and input constraints, the startup script supports a maximum of **16,384 characters (approx. 16KB)**. Any script exceeding this limit is truncated on save. For complex initialization logic, write it inside a separate skill entrypoint or fetch and execute remote scripts.
* **Timeout Protection**: Script execution is protected by a timeout limit controlled by the environment variable `AGENT_SANDBOX_ENTRYPOINT_TIMEOUT_SECONDS`, with a **default timeout of 30 seconds** (clamped between 1 and 600 seconds). The process is forcefully terminated if execution exceeds this duration.
* **Non-blocking Workflow**: If your startup script errors out (exits with a non-zero code), times out, or fails to read/write the state file, the system **will not block** the main chat workflow. The AI continues executing subsequent workflow nodes or tools, though it may hit runtime exceptions later if critical dependencies are missing. You can troubleshoot these errors in the preview logs or through the Computer File Manager.
* **Log Truncation**: Combined standard output (stdout) and standard error (stderr) logs for the startup script are capped at approximately **8KB**. When logged to the system, output is truncated to 4,000 characters to prevent excessive resource utilization.
file: ./content/guide/build/agentv2/vm.mdx
meta: {
"title": "虚拟机",
"description": "深入了解 Agent V2 虚拟机(沙箱)的核心概念、运行设计与联调指南。"
}
import { Alert } from '@/components/docs/Alert';

在 FastGPT Agent V2 中,**虚拟机** 是专为每个会话分配的、物理隔离且安全的轻量级 Linux 运行沙盒环境。它为 Agent 提供了真实的计算、代码执行和文件读写操作能力,使得 AI 不仅仅能“思考”,还能像人类程序员一样通过实际运行代码来解决复杂任务。
***
## 什么是虚拟机
每当用户开启“启用虚拟机”选项,系统都会在后台为每一个对话会话(Session)动态预置并绑定一个专用的沙箱容器。

有了虚拟机,Agent 可以:
* **执行动态代码**:通过代码执行器运行 Python、Node.js 甚至 Shell 脚本,自主进行复杂计算或数据处理。
* **读写本地文件**:在独立的 `/workspace` 目录下创建、修改和读取文件,包括生成图表、处理上传的 CSV/Excel 电子表格等。
* **环境自定义与启动脚本**:通过关联 SKILL 包,或配置自定义的 [启动脚本](#虚拟机生命周期)(在虚拟机拉起后且 AI 正式开始前自动在后台执行的 Shell 命令,用于安装特定软件源、Python 依赖包或系统级工具等),动态准备专属于您应用的运行环境。
***
## 虚拟机设计
为了在多并发和多租户场景下兼顾安全性、响应速度与资源消耗,系统在底层采用了以下核心设计:
### 1. 会话级隔离与生命周期管理
* 虚拟机与用户的对话会话(Session)强绑定。不同用户、不同会话之间的运行环境完全物理隔离。(注意,未来版本将会改至用户级别隔离,从而减少资源消耗)
* 系统通过心跳(Keepalive)机制维持活动容器的存活。当会话长时间闲置(如超过数分钟无新请求)时,系统会自动回收并销毁该虚拟机实例以释放服务器资源。
### 2. 状态存续(Session Persistence)
在同一个会话的生命周期内,虚拟机是持续存活的。这意味着上一轮对话中下载的依赖、写入的临时文件以及配置的环境变量,在下一轮对话中依然有效,支持多轮交互的连续性。
### 3. 安全防逃逸与限制
* 沙盒环境采用了严格的资源配额限制(CPU、内存、磁盘 IO)和网络防火墙策略,防止恶意脚本消耗宿主机资源或访问内部网络。
* 所有执行的命令均通过 base64 编码在容器内安全解码运行,规避了 Shell 拼接造成的命令注入风险。
### 4. 一键快速重建
当用户在调试中点击“重开对话”或重新建立会话时,系统将彻底销毁当前的旧虚拟机,并为下一次请求分配一个纯净、全新的隔离容器,确保环境干净不被污染。
***
## 虚拟机使用
当您在左侧配置面板中**启用了虚拟机**时,调试预览窗口将自动激活以下沙盒专属调试能力:
### 虚拟机文件管理器
调试窗口顶部以及涉及虚拟机操作的对话气泡下方会提供“虚拟机文件管理器”入口。点击可弹出管理器弹窗,实时浏览、编辑、上传或下载容器内的文件(如图表、临时代码、HTML 预览等),实现调试闭环。
| | |
| :----------------------------------------: | :------------------------------------------: |
|  |  |
### 对话上传文件自动注入
对话输入框中上传的所有测试文件,会在对话开始前自动同步注入到虚拟机的 `user_files/` 目录下,使得 AI 可以通过代码以本地路径直接读取和处理这些文件。
### 代码执行与自我纠错
虚拟机提供了真实的代码运行环境。若代码由于依赖缺失或语法错误执行失败,报错信息会实时回传给大模型,AI 能够在多步规划中尝试自我修正并重新运行,实现闭环纠错调试。
### 前置拉起与状态存续
开启虚拟机后,系统会在每次对话启动前自动前置拉起并准备好沙盒环境。环境内的文件、依赖和运行状态在当前会话内跨步骤持久保留,支持连续的上下文联调。
### 启动状态实时展示
在对话调试过程中,气泡上方会流式展示虚拟机的创建与启动状态,方便感知沙箱所处阶段与启动耗时。
### 重置对话重建沙箱
点击“重开对话”按钮重置测试时,下一次交互会重新拉起并初始化一个干净、全新的虚拟机容器,避免历史测试生成的文件污染新一轮的调试。
***
## 虚拟机生命周期
在启用虚拟机后,您可以通过配置 **启动脚本**,在沙盒环境拉起后、AI 工作流正式开始执行前,自动执行指定的 Shell 命令。这通常用于配置环境变量、更换软件源、安装 Python 依赖(pip)或系统级工具等。
### 脚本配置与生命周期
在 Agent 配置面板的“虚拟机配置”中,您可以直接在 **启动脚本(sh)** 的代码编辑器中编写您的 Shell 脚本。

#### 脚本执行顺序与范围
当一个新会话启动或虚拟机重新拉起时,系统会按顺序在后台执行相应的脚本:
1. **应用启动脚本**:即您在配置面板中自定义的 Shell 脚本。该脚本会在虚拟机的工作目录(通常为 `/workspace`)下执行,用于准备当前应用所需的特定运行依赖和环境。
2. **技能入口脚本(Skill Entrypoint)**:若您的 Agent 关联了技能(Skills),技能发布包中自带的[初始化脚本](../skill/initialization)(如 `entrypoint.sh`)会在上述“应用启动脚本”执行完毕后,在每个技能包自身的部署目录下自动执行。
**技能包部署的原子性保障**:技能包在解压部署时,会先在临时目录(如 `.tmp--`)中解压,解压完全成功后,再以原子操作整体替换为正式版本目录,有效避免解压失败导致出现损坏的半截目录。
#### 生命周期流程图

### 状态去重
为了避免每次对话交互(热启动)时重复执行命令(例如重复通过 `pip install` 安装依赖包)带来等待延迟,系统设计了高效的 **状态去重机制**:
* **运行状态记录**:系统在虚拟机内部维护了一个状态文件:`~/.fastgpt/agent-skill-entrypoints/state.json`。
* **哈希去重(应用启动脚本)**:对于您手动编写的“应用启动脚本”,系统会计算其文本内容的 **SHA-256 哈希特征值(Hash)**,并与 `state.json` 中已执行过的哈希值进行对比。如果脚本内容没有任何修改,系统在后续交互中将 **自动跳过执行**,确保实现快速启动;仅当您修改了脚本内容,或在调试预览中点击了“重开对话”触发沙盒彻底重建时,脚本才会重新执行。
* **版本 ID 去重(技能入口脚本)**:关联技能对应的入口脚本则基于 **技能版本 ID** 进行比对去重。因为发布的技能版本是不可变的,只要绑定的技能版本未改变,其入口脚本也仅会在沙箱首次冷启动时执行一次。
* **状态的生命周期**:去重状态随虚拟机实例生命周期进行管理。当虚拟机重建(点击“重开对话”或闲置被系统回收)时,由于分配的是全新环境,所有的脚本都将在首次冷启动时重新执行。
### 执行限制与容错机制
为保障沙箱的稳定运行与响应时效,虚拟机启动脚本在运行时受以下系统规则约束:
* **字符长度限制**:由于前端校验与输入限制,虚拟机启动脚本最大支持 **16,384 个字符(约 16KB)**。超出此长度的脚本在保存时会被自动截断。对于复杂的初始化逻辑,建议编写在单独的技能入口脚本中,或在启动脚本中拉取远程脚本执行。
* **超时终止限制**:脚本执行存在超时保护限制,由系统环境变量 `AGENT_SANDBOX_ENTRYPOINT_TIMEOUT_SECONDS` 控制,**默认超时时间为 30 秒**(限制在 1 秒到 600 秒之间)。如果超过该时间脚本仍未执行完毕,系统将强制终止该进程。
* **非阻塞主流程**:即使您的启动脚本在执行时报错(退出状态码非 0)、超时终止,或是状态文件读写发生异常,系统也**不会阻断主对话流程**。AI 依然会继续执行后续的工作流或工具调用,但可能会因缺少特定依赖而在代码运行时抛出异常。您可以在调试预览的日志或虚拟机文件管理器中排查此类问题。
* **日志长度截断**:启动脚本标准输出(stdout)和标准错误(stderr)的最大日志输出量限制在 **8KB** 左右。在系统记录日志时,日志内容将被截断至 4,000 个字符,以防止过大的日志输出占用过多的系统与网络资源。
file: ./content/guide/build/general/ai_settings.en.mdx
meta: {
"title": "AI Settings",
"description": "FastGPT AI settings explained"
}
import { Alert } from '@/components/docs/Alert';
AI settings control how AI Chat nodes behave in apps and Workflows, including model selection, response length, multimodal recognition, response format, and reasoning display. This guide explains what each option in the settings modal means and how to choose values for common scenarios.
## Where to Find It
In the app editor, find the **AI Settings** section, select an AI model, and click the settings button on the right side of the model selector to open the AI settings modal.
In a Workflow, click the AI model configuration for the **AI Chat** node. You can open the same settings modal from the settings button on the right.
If you do not have specific requirements, selecting a suitable AI model and keeping the other
settings at their defaults is usually enough.
| | | |
| ------------------------------- | ------------------------------- | ------------------------------- |
|  |  |  |
## Why Some Options May Be Hidden
Not every option is always shown. The modal only displays settings supported by the selected model. For example, if a model does not support multimodal recognition, multimodal options are hidden. If a model does not support reasoning settings, those options are hidden as well.
## Basic Settings
### AI Model
Select the AI model used by the current app or node. Different models vary in response quality, cost, context length, tool calling capability, and multimodal capability.
The model section displays several types of information:
* **Credit cost**: A reference cost for model calls. Input content and model output are usually priced separately.
* **Max context**: The amount of content the model can reference in one request. A larger context window is better for long documents and long conversations.
* **Tool calling**: If supported, the model can use selected app tools to query data, run calculations, or call external capabilities.
* **Multimodal capability**: If the model supports image, audio, or video input, you can enable the corresponding multimodal recognition capability in AI Settings. Different models may support different media types. Use the capabilities shown in the settings modal as the source of truth.
### Max Histories
Controls how many previous conversation rounds the AI can reference when answering.
Higher values make it easier for the AI to use earlier context, but they also add more content to the request, which may increase cost and slow down responses. Lower values may prevent the model from using useful prior context.
If you do not have a specific requirement, use the default value. Customer support and Knowledge Base Q\&A apps usually only need a small number of history rounds.
### Max Tokens
Controls the maximum length of a single AI response.
When enabled, use the slider to limit response length. If the value is too low, the response may be cut off early. A higher value allows the model to generate more complete answers, but may also increase cost.
Use a lower value for concise answers. Use a higher value when generating plans, articles, or longer explanations.
### Temperature
Controls how stable the response is.
Lower values produce more stable responses and are better for customer support, Knowledge Base Q\&A, and scenarios with clear rules. Higher values produce more varied responses and are better for writing, brainstorming, and creative content.
Common choices:
* Customer support and Knowledge Base Q\&A: use a lower value.
* Copywriting, stories, and creative suggestions: use a higher value.
* If unsure: keep the default.
### Top\_p
Top\_p also controls response randomness, with some overlap with temperature.
In most cases, avoid adjusting temperature and Top\_p at the same time. If temperature already gives the result you want, keep Top\_p disabled or at its default.
### Stop
Stops the AI response when specified content appears.
Most chat scenarios do not need this setting. Use it only when the model should stop after outputting a fixed marker. Separate multiple stop strings with `|`, for example: `end|stop`.
### Response Format
Controls the format of the AI response.
For regular chat, customer support, and Knowledge Base Q\&A, keep the default. Change this only when a later step needs to read the response in a fixed format.
If you select `json_schema`, you also need to provide the corresponding schema. This option is suitable when the model must return content in a fixed structure.
### Multimodal Recognition
If the selected model is configured with multimodal capability, this setting controls whether the AI can read images, audio, or video from user input.
The available types depend on the model itself. If a model only supports images, only image recognition can be enabled. If it supports images, audio, or video, you can select the needed types.
When enabled, the AI Chat node converts matching uploaded files, or matching media links in the user's question, into model-readable input before sending the request. For example:
* Image recognition: for screenshots, table images, product images, posters, and similar image content.
* Audio recognition: for models that can understand uploaded audio content.
* Video recognition: for models that can understand uploaded video content.
Keep these limits in mind:
1. Even after a type is enabled, the request is filtered again by the actual model capability before it is sent. Unsupported media types are not sent to the model.
2. Media links in the user's question are only parsed when "Extract multimodal files from links" is enabled. Currently, extraction is attempted only when the user's question is under 500 characters, with at most 4 media links processed at a time.
3. Regular document files are not sent directly to the LLM as multimodal input. Documents still need to be parsed into text first.
4. Multimodal recognition depends on the model's own capability. If the modal says the model does not support multimodal recognition, switch to a model that supports the needed media type.
### Hide AI Output
When enabled, AI-generated content is not shown directly to the user, but it can still be passed to downstream nodes through the AI response output. For example, the AI can first organize an internal result, and the next node can rewrite it into the final response.
## Reasoning Settings
Some models can generate reasoning content before the final answer. When you select one of these models, the modal shows reasoning settings.
### Reasoning Effort
Controls how much reasoning the model performs.
* **Default**: Use the model's default behavior.
* **None**: Try to answer directly. This is suitable for simple questions.
* **Minimal / Low / Medium / High / Extra high**: Use stronger reasoning for more complex questions.
Reasoning effort follows OpenAI's `reasoning_effort` convention, with ai-proxy adapting it to the parameter format required by each model provider. For the full rules, see [ai-proxy reasoning compatibility](https://github.com/labring/aiproxy/blob/main/docs/REASONING_COMPATIBILITY.md).
OpenAI-Compatible Enum and Default Budget Mapping
| FastGPT option | OpenAI-compatible value | Default budget |
| -------------- | ----------------------------------------- | --------------------- |
| Default | Do not explicitly send `reasoning_effort` | Use the model default |
| None | `none` | `0` |
| Minimal | `minimal` | `1024` |
| Low | `low` | `2048` |
| Medium | `medium` | `8192` |
| High | `high` | `16384` |
| Extra high | `xhigh` | `32768` |
If an upstream provider only supports a token budget instead of discrete effort levels, ai-proxy uses the table above to convert effort to budget. When normalizing budget back to effort, `<=0` maps to `none`, `1-1024` maps to `minimal`, `1025-4096` maps to `low`, `4097-12288` maps to `medium`, `12289-24576` maps to `high`, and anything higher maps to `xhigh`.
OpenAI / OpenAI Responses
| Target format | Output field | Mapping |
| ------------------------- | ------------------ | ------------------------------------------------- |
| OpenAI Chat / Completions | `reasoning_effort` | Writes `none/minimal/low/medium/high/xhigh` as-is |
| OpenAI Responses | `reasoning.effort` | Writes `none/minimal/low/medium/high/xhigh` as-is |
OpenAI Chat / Completions only parses `reasoning_effort`. When Gemini, Claude, or other request formats are converted to an OpenAI-compatible format, they are first normalized to this field.
Google Gemini
Gemini native requests are parsed from `generationConfig.thinkingConfig`, including `thinkingLevel`, `thinkingBudget`, and `includeThoughts`. When writing to Gemini upstreams, ai-proxy chooses either `thinkingLevel` or `thinkingBudget` based on the model family.
| OpenAI-compatible value | Gemini 3+ Pro | Gemini 3+ non-Pro | gemini-2.5-pro | gemini-2.5-flash | gemini-2.5-flash-lite |
| ----------------------- | -------------------- | ----------------------- | ---------------------- | ---------------------- | ---------------------- |
| `none` | `thinkingLevel=low` | `thinkingLevel=minimal` | `thinkingBudget=128` | `thinkingBudget=0` | `thinkingBudget=0` |
| `minimal` | `thinkingLevel=low` | `thinkingLevel=minimal` | `thinkingBudget=1024` | `thinkingBudget=1024` | `thinkingBudget=1024` |
| `low` | `thinkingLevel=low` | `thinkingLevel=low` | `thinkingBudget=2048` | `thinkingBudget=2048` | `thinkingBudget=2048` |
| `medium` | `thinkingLevel=low` | `thinkingLevel=medium` | `thinkingBudget=8192` | `thinkingBudget=8192` | `thinkingBudget=8192` |
| `high` | `thinkingLevel=high` | `thinkingLevel=high` | `thinkingBudget=16384` | `thinkingBudget=16384` | `thinkingBudget=16384` |
| `xhigh` | `thinkingLevel=high` | `thinkingLevel=high` | `thinkingBudget=32768` | `thinkingBudget=24576` | `thinkingBudget=24576` |
Gemini 2.5 models clamp the budget to the model's supported range. Some Gemini models cannot fully disable thinking, so `none` falls back to the minimum supported level or budget.
Claude / Anthropic / Bedrock / Vertex AI
Claude native requests are parsed from `thinking` and `output_config`. When writing to Anthropic, AWS Bedrock Claude, or Vertex AI Claude, the payload still follows Claude's thinking format.
| OpenAI-compatible value | Legacy / budget mode | Adaptive mode |
| ----------------------- | ---------------------------------------------- | ---------------------------------------------------------------------- |
| `none` | `thinking.type=disabled` | `thinking.type=disabled`; may be removed for some adaptive-only models |
| `minimal` | `thinking.type=enabled`, `budget_tokens=1024` | `thinking.type=adaptive`, `output_config.effort=low` |
| `low` | `thinking.type=enabled`, `budget_tokens=2048` | `thinking.type=adaptive`, `output_config.effort=low` |
| `medium` | `thinking.type=enabled`, `budget_tokens=8192` | `thinking.type=adaptive`, `output_config.effort=medium` |
| `high` | `thinking.type=enabled`, `budget_tokens=16384` | `thinking.type=adaptive`, `output_config.effort=high` |
| `xhigh` | `thinking.type=enabled`, `budget_tokens=32768` | `thinking.type=adaptive`, `output_config.effort=max` |
Budget mode ensures `budget_tokens < max_tokens` and raises too-small budgets to the minimum accepted by the upstream provider.
Ali DashScope / Qwen / QwQ / GLM / Kimi-Compatible Models
| OpenAI-compatible value | Models with `thinking_budget` support | Models without budget support |
| ----------------------- | ------------------------------------------------- | ----------------------------- |
| `none` | `enable_thinking=false`; remove `thinking_budget` | `enable_thinking=false` |
| `minimal` | `enable_thinking=true`, `thinking_budget=1024` | `enable_thinking=true` |
| `low` | `enable_thinking=true`, `thinking_budget=2048` | `enable_thinking=true` |
| `medium` | `enable_thinking=true`, `thinking_budget=8192` | `enable_thinking=true` |
| `high` | `enable_thinking=true`, `thinking_budget=16384` | `enable_thinking=true` |
| `xhigh` | `enable_thinking=true`, `thinking_budget=32768` | `enable_thinking=true` |
ai-proxy currently treats `qwen3-*`, `qwq-*`, and Ali-compatible models whose names contain `glm` or `kimi` as supporting `thinking_budget`. Non-streaming `qwen3-*` requests are forced to disable thinking, while `qwq-*` requests are forced to streaming mode.
Zhipu / DeepSeek / Doubao / Moonshot Kimi
These providers currently preserve only the on/off meaning. They do not preserve budget or fine-grained effort levels.
| Provider | OpenAI-compatible value | Upstream field |
| --------------------------------------------- | ------------------------------- | --------------------------------------------------- |
| Zhipu / DeepSeek / Doubao | `none` | `thinking.type=disabled` |
| Zhipu / DeepSeek / Doubao | `minimal/low/medium/high/xhigh` | `thinking.type=enabled` |
| Moonshot / Kimi models with switch support | `none` | `thinking.type=disabled`; remove `reasoning_effort` |
| Moonshot / Kimi models with switch support | `minimal/low/medium/high/xhigh` | `thinking.type=enabled`; remove `reasoning_effort` |
| Moonshot / Kimi models without switch support | Any value | Remove `reasoning_effort` and omit `thinking` |
For Moonshot / Kimi, whether `thinking.type` can be written depends on the actual upstream model name after channel mapping.
Some models may not fully support every reasoning option. If an error occurs after switching the option, change it back to Default.
### Hide AI Reasoning
When enabled, users only see the final answer and do not see the AI's reasoning process. During app debugging, you can temporarily disable this option to inspect the model's intermediate reasoning.
file: ./content/guide/build/general/ai_settings.mdx
meta: {
"title": "AI 配置说明",
"description": "FastGPT AI 配置说明"
}
import { Alert } from '@/components/docs/Alert';
AI 配置用于调整应用或工作流中 AI 对话节点的模型、回复长度、多模态识别、回复格式和思考展示等行为。本文主要介绍配置弹窗中的各项含义,以及常见场景下的选择方式。
## 配置入口
在应用编辑页中,找到 **AI 配置** 区域,选择 AI 模型后,点击模型选择框右侧的设置按钮,即可打开 AI 配置弹窗。
在工作流中,点击 **AI 对话** 节点的 AI 模型配置项,也可以通过右侧设置按钮打开同一个配置弹窗。
如果暂时没有特殊要求,通常只需要选择合适的 AI 模型,其他参数保持默认即可。
| | | |
| ------------------------------- | ------------------------------- | ------------------------------- |
|  |  |  |
## 配置会不会都显示?
不会。弹窗会根据当前模型的能力显示可用配置。例如,模型不支持多模态识别时,不会提供多模态识别选项;模型不支持思考配置时,也不会显示对应选项。
## 基础配置
### AI 模型
用于选择当前应用或节点使用的 AI 模型。不同模型在回答能力、价格、可处理内容长度、工具调用能力、多模态能力等方面会有差异。
模型下面会显示几类信息:
* **积分价格**:模型调用时的积分消耗参考,通常会区分输入内容和模型回复。
* **最大上下文**:模型单次请求可参考的内容长度。数值越大,越适合长文档、长对话等场景。
* **工具调用**:如果显示支持,说明该模型可以配合应用中选择的工具完成查询、计算或外部能力调用。
* **多模态能力**:如果模型支持图片、音频或视频输入,可以在 AI 配置中开启对应的多模态识别能力。不同模型支持的媒体类型可能不同,具体以配置弹窗中显示的能力为准。
### 记忆轮数
控制 AI 回答时最多参考前面多少轮聊天。
数值越大,AI 越容易参考更早的对话,但也会带入更多内容,可能增加消耗并影响响应速度。数值过小,则可能无法利用前文信息。
如果没有明确需求,建议先使用默认值。客服、知识库问答类应用通常保留少量历史轮数即可。
### 回复上限
控制 AI 一次最多回答多长。
打开后,可以通过滑块限制模型回复长度。设置过低时,回复可能被提前截断;设置较高时,模型可以生成更完整的内容,但也可能增加消耗。
如果希望回答简短,可以适当调低;如果需要生成方案、文章或较长说明,可以适当调高。
### 温度
控制回答的稳定程度。
数值较低时,回答更稳定,更适合客服、知识库问答、规则明确的场景。数值较高时,回答更发散,更适合写作、头脑风暴、创意内容等场景。
常见选择:
* 客服、知识库问答:建议偏低。
* 文案、故事、创意建议:可以适当调高。
* 不确定时:建议保持默认。
### Top\_p
Top\_p 也是控制回复随机性的参数,作用和温度有一定重叠。
通常不建议同时调整温度和 Top\_p。如果已经通过温度获得了期望效果,可以保持 Top\_p 关闭或默认。
### 停止序列
当 AI 回复中出现指定内容时,会停止继续输出。
普通聊天一般不需要设置。只有在需要模型输出到某个固定标记就结束时才使用。多个停止词可以用 `|` 分隔,例如:`结束|stop`。
### 回复格式
控制 AI 的回答格式。
普通聊天、客服问答、知识库问答通常保持默认即可。只有当后续流程需要读取固定格式的内容时,才需要修改该配置。
如果选择 `json_schema`,还需要填写对应的格式要求。该选项适合需要模型按固定结构返回内容的场景。
### 多模态识别
如果当前模型配置了多模态能力,这里可以控制 AI 是否读取用户输入中的图片、音频或视频内容。
可选择的类型取决于模型本身的能力。模型只支持图片时,只能开启图片识别;模型同时支持图片、音频或视频时,可以按需选择对应类型。
打开后,AI 对话节点会在请求模型前,将用户上传的对应类型文件,或用户问题中的对应媒体链接,转换为模型可识别的输入。例如:
* 图片识别:用于识别截图、表格图片、商品图、海报等图片内容。
* 音频识别:用于让支持音频输入的模型理解用户上传的音频内容。
* 视频识别:用于让支持视频输入的模型理解用户上传的视频内容。
需要注意:
1. 即使开启了某类识别,请求发送前也会再次根据模型能力过滤,不支持的类型不会发送给模型。
2. 用户问题中的媒体链接需要开启“提取链接中的多模态文件”后才会尝试解析。当前仅在用户问题少于 500 字时尝试提取,且一次最多处理 4 个媒体链接。
3. 普通文档文件不会作为多模态输入直接发送给 LLM,文档内容仍需要通过文件解析转成文本。
4. 多模态识别依赖模型本身能力。如果弹窗里显示“该模型不支持多模态识别”,需要换成支持对应多模态输入的模型。
### 隐藏 AI 输出
打开后,AI 生成的内容不会直接展示给用户,但仍然可以通过 AI 回复输出交给后续节点继续处理。例如先让 AI 整理内部结果,再由下一个节点改写成最终回复。
## 思考配置
部分模型支持先生成思考过程,再输出最终回答。选择这类模型时,弹窗会显示思考配置。
### 思考配置
用于控制模型的思考强度。
* **默认**:使用模型默认配置。
* **不思考**:尽量直接回答,适合简单问题。
* **极简思考 / 轻量思考 / 标准思考 / 深度思考 / 极致思考**:问题越复杂,可以选择更高的思考强度。
思考强度配置对齐 OpenAI 规范中的 `reasoning_effort`,并通过 ai-proxy 适配不同模型平台的参数格式。完整规则可参考 [ai-proxy reasoning compatibility](https://github.com/labring/aiproxy/blob/main/docs/REASONING_COMPATIBILITY.zh.md)。
OpenAI 兼容枚举与默认 budget 映射
| FastGPT 选项 | OpenAI 兼容值 | 默认 budget |
| ---------- | ------------------------ | --------- |
| 默认 | 不显式传递 `reasoning_effort` | 使用模型默认值 |
| 不思考 | `none` | `0` |
| 极简思考 | `minimal` | `1024` |
| 轻量思考 | `low` | `2048` |
| 标准思考 | `medium` | `8192` |
| 深度思考 | `high` | `16384` |
| 极致思考 | `xhigh` | `32768` |
如果某个平台只支持 token budget,不支持离散档位,ai-proxy 会按上表把 effort 转成 budget。反向归一化时,`<=0` 会被视为 `none`,`1~1024` 视为 `minimal`,`1025~4096` 视为 `low`,`4097~12288` 视为 `medium`,`12289~24576` 视为 `high`,更高则视为 `xhigh`。
OpenAI / OpenAI Responses
| 目标格式 | 写入字段 | 映射方式 |
| ------------------------- | ------------------ | ----------------------------------------- |
| OpenAI Chat / Completions | `reasoning_effort` | `none/minimal/low/medium/high/xhigh` 原样写入 |
| OpenAI Responses | `reasoning.effort` | `none/minimal/low/medium/high/xhigh` 原样写入 |
OpenAI Chat / Completions 模式只解析 `reasoning_effort`。当 Gemini、Claude 等请求被转换为 OpenAI 兼容格式时,也会先归一化为该字段。
Google Gemini
Gemini 原生请求会从 `generationConfig.thinkingConfig` 中解析 `thinkingLevel`、`thinkingBudget` 和 `includeThoughts`。写给 Gemini 上游时,ai-proxy 会根据模型系列选择 `thinkingLevel` 或 `thinkingBudget`。
| OpenAI 兼容值 | Gemini 3+ Pro | Gemini 3+ 非 Pro | gemini-2.5-pro | gemini-2.5-flash | gemini-2.5-flash-lite |
| ---------- | -------------------- | ----------------------- | ---------------------- | ---------------------- | ---------------------- |
| `none` | `thinkingLevel=low` | `thinkingLevel=minimal` | `thinkingBudget=128` | `thinkingBudget=0` | `thinkingBudget=0` |
| `minimal` | `thinkingLevel=low` | `thinkingLevel=minimal` | `thinkingBudget=1024` | `thinkingBudget=1024` | `thinkingBudget=1024` |
| `low` | `thinkingLevel=low` | `thinkingLevel=low` | `thinkingBudget=2048` | `thinkingBudget=2048` | `thinkingBudget=2048` |
| `medium` | `thinkingLevel=low` | `thinkingLevel=medium` | `thinkingBudget=8192` | `thinkingBudget=8192` | `thinkingBudget=8192` |
| `high` | `thinkingLevel=high` | `thinkingLevel=high` | `thinkingBudget=16384` | `thinkingBudget=16384` | `thinkingBudget=16384` |
| `xhigh` | `thinkingLevel=high` | `thinkingLevel=high` | `thinkingBudget=32768` | `thinkingBudget=24576` | `thinkingBudget=24576` |
Gemini 2.5 系列会按模型允许范围 clamp budget。部分 Gemini 模型不能真正关闭 thinking,`none` 会退化为模型允许的最小 level 或 budget。
Claude / Anthropic / Bedrock / Vertex AI
Claude 原生请求会解析 `thinking` 和 `output_config`。写给 Anthropic 官方、AWS Bedrock Claude 或 Vertex AI Claude 时,字段形态仍遵循 Claude 的 thinking 规则。
| OpenAI 兼容值 | 旧式 / budget 模式 | adaptive 模式 |
| ---------- | ----------------------------------------------- | -------------------------------------------------------- |
| `none` | `thinking.type=disabled` | `thinking.type=disabled`,部分 adaptive-only 模型可能移除该字段 |
| `minimal` | `thinking.type=enabled` , `budget_tokens=1024` | `thinking.type=adaptive` , `output_config.effort=low` |
| `low` | `thinking.type=enabled` , `budget_tokens=2048` | `thinking.type=adaptive` , `output_config.effort=low` |
| `medium` | `thinking.type=enabled` , `budget_tokens=8192` | `thinking.type=adaptive` , `output_config.effort=medium` |
| `high` | `thinking.type=enabled` , `budget_tokens=16384` | `thinking.type=adaptive` , `output_config.effort=high` |
| `xhigh` | `thinking.type=enabled` , `budget_tokens=32768` | `thinking.type=adaptive` , `output_config.effort=max` |
budget 模式会保证 `budget_tokens < max_tokens`,并把过小的 budget 提升到上游可接受的最小值。
Ali DashScope / Qwen / QwQ / GLM / Kimi 兼容模型
| OpenAI 兼容值 | 支持 `thinking_budget` 的模型 | 不支持 budget 的模型 |
| ---------- | ------------------------------------------------ | ----------------------- |
| `none` | `enable_thinking=false`,移除 `thinking_budget` | `enable_thinking=false` |
| `minimal` | `enable_thinking=true` , `thinking_budget=1024` | `enable_thinking=true` |
| `low` | `enable_thinking=true` , `thinking_budget=2048` | `enable_thinking=true` |
| `medium` | `enable_thinking=true` , `thinking_budget=8192` | `enable_thinking=true` |
| `high` | `enable_thinking=true` , `thinking_budget=16384` | `enable_thinking=true` |
| `xhigh` | `enable_thinking=true` , `thinking_budget=32768` | `enable_thinking=true` |
当前 ai-proxy 会把 `qwen3-*`、`qwq-*`、模型名包含 `glm` 或 `kimi` 的 Ali-compatible 模型视为支持 `thinking_budget`。`qwen3-*` 非流式请求会被强制关闭 thinking,`qwq-*` 请求会被强制改为流式。
Zhipu / DeepSeek / Doubao / Moonshot Kimi
这些平台当前主要保留开关语义,不保留 budget 或细粒度 effort。
| 平台 | OpenAI 兼容值 | 写给上游的字段 |
| ------------------------- | ------------------------------- | ----------------------------------------------- |
| Zhipu / DeepSeek / Doubao | `none` | `thinking.type=disabled` |
| Zhipu / DeepSeek / Doubao | `minimal/low/medium/high/xhigh` | `thinking.type=enabled` |
| Moonshot / Kimi 支持开关的模型 | `none` | `thinking.type=disabled`,并移除 `reasoning_effort` |
| Moonshot / Kimi 支持开关的模型 | `minimal/low/medium/high/xhigh` | `thinking.type=enabled`,并移除 `reasoning_effort` |
| Moonshot / Kimi 不支持开关的模型 | 任意值 | 移除 `reasoning_effort`,不发送 `thinking` |
Moonshot / Kimi 是否能写入 `thinking.type` 取决于渠道映射后的实际上游模型名。
部分模型不一定完全支持所有思考选项。如果切换后出现报错,可以改回默认选项。
### 隐藏 AI 思考
打开后,用户只会看到最终回答,看不到 AI 的思考过程。调试应用时,可以临时关闭该开关,观察模型的中间思考内容。
file: ./content/guide/build/general/chat_input_guide.en.mdx
meta: {
"title": "Chat Input Guide",
"description": "FastGPT chat input guide"
}

## What is Custom Question Guidance?
You can preset questions for your app. As users type, the system dynamically searches these questions based on their input and displays them as suggestions, helping users ask questions faster.
You can configure the question list directly in FastGPT or provide a custom API endpoint.
## Custom Question List API
The endpoint must be accessible from the user's browser.
**Request:**
```bash
curl --location --request GET 'http://localhost:3000/api/core/chat/inputGuide/query?appId=663c75302caf8315b1c00194&searchKey=you'
```
Where `appId` is the application ID and `searchKey` is the search keyword (max 50 characters).
**Response**
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": [
"it's you",
"who are you",
"you're great",
"hello there",
"who are you!",
"hello"
]
}
```
`data` is an array of matched questions. Return at most 5 results.
**Parameters:**
* appId - Application ID
* searchKey - Search keyword
file: ./content/guide/build/general/chat_input_guide.mdx
meta: {
"title": "对话问题引导",
"description": "FastGPT 对话问题引导"
}

## 什么是自定义问题引导
你可以为你的应用提前预设一些问题,用户在输入时,会根据输入的内容,动态搜索这些问题作为提示,从而引导用户更快的进行提问。
你可以直接在 FastGPT 中配置词库,或者提供自定义词库接口。
## 自定义词库接口
需要保证这个接口可以被用户浏览器访问,需要注意允许跨域。
**请求:**
```bash
curl --location --request GET 'http://localhost:3000/api/core/chat/inputGuide/query?appId=663c75302caf8315b1c00194&searchKey=你'
```
其中 `appId` 为应用 ID,`searchKey` 为搜索关键字,最多是 50 个字符。
**响应**
```json
{
"code": 200,
"statusText": "",
"message": "",
"data": ["是你", "你是谁呀", "你好好呀", "你好呀", "你是谁!", "你好"]
}
```
data 是一个数组,包含了搜索到的问题,最多只需要返回 5 个问题。
**参数说明:**
* appId - 应用 ID
* searchKey - 搜索关键字
file: ./content/guide/build/general/fileInput.en.mdx
meta: {
"title": "File Input",
"description": "FastGPT file input feature overview"
}
Starting from version 4.8.9, FastGPT supports configuring file uploads in both `Basic Mode` and `Workflows`. This guide covers how to use file input and explains the difference between document parsing and multimodal file handling.
## Using in Basic Mode
When file upload is enabled in Basic Mode, it uses tool-calling mode — the model decides whether to read the file content.
Find the file upload option on the left panel and click the `Enable`/`Disable` toggle to open the configuration dialog.

Once enabled, a file selection icon appears in the chat input area. Click it to select files for upload.

**Behavior**
Starting from version 4.8.13, Basic Mode forces file parsing and injects the content into the system prompt, preventing cases where the model skips reading the file during multi-turn conversations.
## Using in Workflows
In Workflows, find the `File Input` option in the system configuration panel and click the `Enable`/`Disable` toggle to open the configuration dialog.

There are many ways to use files in Workflows. The simplest approach, shown below, connects document parsing via tool calling — achieving the same result as Basic Mode.
| | |
| ---------------------- | ---------------------- |
|  |  |
You can also use Workflows to extract or analyze document content, then pass the results to HTTP requests or other modules to build a document processing pipeline.

## How Document Parsing Works
Unlike multimodal recognition, LLMs currently cannot parse regular documents directly. All document "understanding" is achieved by converting documents to text and injecting it into the prompt. The following FAQs explain how this works — understanding the mechanics helps you use document parsing more effectively in Workflows.
### How are uploaded files stored in the database?
In FastGPT's chat history, messages with role=user store their value in this structure:
```ts
type UserChatItemValueItemType = {
type: 'text' | 'file';
text?: {
content: string;
};
file?: {
type: 'image' | 'audio' | 'video' | 'file';
name?: string;
key?: string;
url: string;
};
};
```
Uploaded files are stored as URLs — parsed document content is not stored.
### How are images, audio, and video handled?
The document parsing node does not parse multimodal files such as images, audio, or video. These files should be handled by an LLM that supports the corresponding multimodal capability, with multimodal recognition enabled in [AI Settings](./ai_settings).
In practice, file input has two different handling paths:
1. Document parsing: handles document files such as PDF, Word, Excel, Markdown, and HTML, converts their content to text, and provides that text to the AI.
2. Multimodal recognition: handles media files such as images, audio, and video. FastGPT converts them into model-readable input, and a model with the corresponding capability reads them.
### How does the document parsing node work?
The document parsing node accepts an `array` input (file URLs) and outputs a `string` (the parsed content).
* The node only parses URLs with document-type file extensions. If you upload both documents and multimodal files, multimodal files are ignored.
* **The document parsing node only processes files from the current workflow run, not files from chat history.**
* How multiple documents are concatenated:
Multiple files are concatenated using the following template — filename + content, separated by `\n******\n`:
```
File: ${filename}
${content}
```
### How to use document parsing in AI nodes
AI nodes (AI Chat / Tool Calling) have a document URL input that lets you reference document addresses directly.
It accepts an `Array` input. The URLs are parsed and injected into a system message using this prompt template:
```
Use the content in as reference for this conversation:
{{quote}}
```
# Changes to File Upload in Version 4.8.13
There are some differences from version 4.8.9. We've maintained backward compatibility to avoid breaking existing workflows, but please update your workflows to follow the new rules as soon as possible — compatibility code will be removed in future versions.
1. Basic Mode now forces file parsing instead of letting the model decide, ensuring documents are always referenced.
2. Document parsing: no longer parses files from chat history.
3. Tool Calling: supports direct document reference selection — no need to attach a document parsing tool. Automatically parses files from chat history.
4. AI Chat: supports direct document reference selection — no need to go through the document parsing node. Automatically parses files from chat history.
5. Standalone plugin execution: no longer supports global files. Plugin inputs now support file-type configuration as a replacement for global file upload.
6. **Workflow calling plugins: uploaded files are no longer automatically passed to plugins. You must manually specify the variable for plugin input.**
7. **Workflow calling sub-workflows: uploaded files are no longer automatically passed to sub-workflows. You can manually select which file URLs to pass.**
file: ./content/guide/build/general/fileInput.mdx
meta: {
"title": "文件输入功能介绍",
"description": "FastGPT 文件输入功能介绍"
}
从 4.8.9 版本起,FastGPT 支持在 `简易模式` 和 `工作流` 中,配置用户上传文件功能。下面先简单介绍下如何使用文件输入功能,最后介绍文档解析和多模态文件处理的区别。
## 简易模式中使用
简易模式打开文件上传后,会使用工具调用模式,也就是由模型自行决策,是否需要读取文件内容。
可以找到左侧文件上传的配置项,点击其右侧的 `开启` / `关闭` 按键,即可打开配置弹窗。

随后,你的调试对话框中,就会出现一个文件选择的 icon,可以点击文件选择 icon,选择你需要上传的文件。

**工作模式**
从 4.8.13 版本起,简易模式的文件读取将会强制解析文件并放入 system 提示词中,避免连续对话时,模型有时候不会主动调用读取文件的工具。
## 工作流中使用
工作流中,可以在系统配置中,找到 `文件输入` 配置项,点击其右侧的 `开启` / `关闭` 按键,即可打开配置弹窗。

在工作流中,使用文件的方式很多,最简单的就是类似下图中,直接通过工具调用接入文档解析,实现和简易模式一样的效果。
| | |
| ---------------------- | ---------------------- |
|  |  |
当然,你也可以在工作流中,对文档进行内容提取、内容分析等,然后将分析的结果传递给 HTTP 或者其他模块,从而实现文件处理的 SOP。

## 文档解析工作原理
不同于多模态识别,LLM 模型目前没有支持直接解析普通文档的能力,所有的文档“理解”都是通过文档转文字后拼接 prompt 实现。这里通过几个 FAQ 来解释文档解析的工作原理,理解文档解析的原理,可以更好的在工作流中使用文档解析功能。
### 上传的文件如何存储在数据库中
FastGPT 的对话记录存储结构中,role=user 的消息,value 值会按以下结构存储:
```ts
type UserChatItemValueItemType = {
type: 'text' | 'file';
text?: {
content: string;
};
file?: {
type: 'image' | 'audio' | 'video' | 'file';
name?: string;
key?: string;
url: string;
};
};
```
也就是说,上传的文件都会以 URL 的形式存储在库中,并不会存储 `解析后的文档内容`。
### 图片、音频、视频如何处理
文档解析节点不会解析图片、音频、视频等多模态文件。这类文件需要交给支持对应多模态能力的 LLM 处理,并在 [AI 配置说明](./ai_settings) 中开启多模态识别。
因此,文件输入中要区分两类处理方式:
1. 文档解析:处理 PDF、Word、Excel、Markdown、HTML 等文档文件,将内容转成文本后提供给 AI。
2. 多模态识别:处理图片、音频、视频等媒体文件,FastGPT 会将其转换为模型可接收的输入,再由支持对应能力的模型读取。
### 文档解析节点如何工作
文档解析依赖文档解析节点,这个节点会接收一个 `array` 类型的输入,对应的是文件输入的 URL;输出的是一个 `string`,对应的是文档解析后的内容。
* 在文档解析节点中,只会解析 `文档` 类型的 URL,它是通过文件 URL 解析出来的 `文件后缀` 去判断的。如果你同时选择了文档和多模态文件,多模态文件会被忽略。
* **文档解析节点,只会解析本轮工作流接收的文件,不会解析历史记录的文件。**
* 多个文档内容如何拼接的
按下列的模板,对多个文件进行拼接,即文件名+文件内容的形式组成一个字符串,不同文档之间通过分隔符:`\n******\n` 进行分割。
```
File: ${filename}
${content}
```
### AI 节点中如何使用文档解析
在 AI 节点(AI 对话/工具调用)中,新增了一个文档链接的输入,可以直接引用文档的地址,从而实现文档内容的引用。
它接收一个 `Array` 类型的输入,最终这些 URL 会被解析,并进行提示词拼接,放置在 role=system 的消息中。提示词模板如下:
```
将 中的内容作为本次对话的参考:
{{quote}}
```
# 4.8.13 版本起,关于文件上传的更新
由于与 4.8.9 版本有些差异,尽管我们做了向下兼容,避免工作流立即不可用。但是请尽快的按新版本规则进行调整工作流,后续将会去除兼容性代码。
1. 简易模式中,将会强制进行文件解析,不再由模型决策是否解析,保证每次都能参考文档。
2. 文档解析:不再解析历史记录中的文件。
3. 工具调用:支持直接选择文档引用,不需要再挂载文档解析工具。会自动解析历史记录中的文件。
4. AI 对话:支持直接选择文档引用,不需要进过文档解析节点。会自动解析历史记录中的文件。
5. 插件单独运行:不再支持全局文件;插件输入支持配置文件类型,可以取代全局文件上传。
6. **工作流调用插件:不再自动传递工作流上传的文件到插件,需要手动给插件输入指定变量。**
7. **工作流调用工作流:不再自动传递工作流上传的文件到子工作流,可以手动选择需要传递的文件链接。**
file: ./content/guide/build/general/voiceInput.en.mdx
meta: {
"title": "Voice Input",
"description": "FastGPT voice input configuration"
}
Voice Input lets users record speech in the chat UI and automatically convert it to text. It is useful on mobile devices, in customer support workflows, and in scenarios where typing is inconvenient.
## Configuration Entry
In the app editor, find **Voice Input** and click the settings button on the right to open the voice input configuration dialog.
| | |
| ---------------------------------------------------------- | ------------------------------------------------------------- |
|  |  |
## Enable Voice Input
After Voice Input is enabled, the chat input box shows a voice recording entry. Users can click it to start recording, and FastGPT converts the recording to text after it finishes.
If the browser or current environment does not support voice recording, the frontend will show a voice input unsupported message.
## Auto Send
When **Auto Send** is enabled, FastGPT automatically sends the recognized text after recording finishes. Users do not need to click the send button manually.
If users should review the recognized text before sending, keep Auto Send disabled.
## Auto Voice Response
When **Auto Voice Response** is enabled, questions sent through voice input will also receive AI responses that play automatically as audio.
This requires voice playback to be enabled. If voice playback is not enabled, the AI still returns a text response, but audio will not play automatically.
## Recommendations
* Enable Voice Input for mobile or on-site scenarios to improve input efficiency.
* For scenarios that require high recognition accuracy, disable Auto Send so users can verify the recognized text first.
* For continuous voice interactions, enable both Auto Send and Auto Voice Response.
file: ./content/guide/build/general/voiceInput.mdx
meta: {
"title": "语音输入",
"description": "FastGPT 语音输入配置说明"
}
语音输入支持用户在前台对话中进行语音录入,并自动识别转换为文字。该能力适合移动端、客服、现场记录等不方便打字的场景。
## 配置入口
在应用编辑页中,找到 **语音输入** 配置项,点击右侧的设置按钮,即可打开语音输入配置弹窗。
| | |
| ------------------------------------------------- | ------------------------------------------------- |
|  |  |
## 开启语音输入
开启后,前台对话输入框中会显示语音录入入口。用户点击后可以开始录音,录音完成后系统会将语音识别为文字。
如果浏览器或当前环境不支持语音录入,前台会提示浏览器不支持语音输入。
## 自动发送
开启 **自动发送** 后,用户完成语音录入并识别为文字后,系统会自动发送该内容,不需要用户再手动点击发送按钮。
如果希望用户在发送前检查识别结果,可以关闭自动发送,让用户确认文字内容后再发送。
## 自动语音回复
开启 **自动语音回复** 后,通过语音输入发送的问题,AI 的回复也会自动以语音形式播放。
该能力需要同时开启语音播报配置。若未开启语音播报,AI 仍会正常生成文字回复,但不会自动播放语音。
## 使用建议
* 面向移动端或现场场景的应用,可以开启语音输入提升输入效率。
* 对识别准确性要求较高的场景,建议关闭自动发送,让用户先确认识别文本。
* 需要连续语音交互时,可以同时开启自动发送和自动语音回复。
file: ./content/guide/build/general/welcomeText.en.mdx
meta: {
"title": "Welcome Text",
"description": "FastGPT welcome text configuration"
}
Welcome Text is the initial message sent automatically before each new conversation starts. Use it to introduce what the app can do, clarify the question scope, or provide common entry points.
## Configuration Entry
In the app editor, find **Welcome Text** and click the settings button on the right to edit the content.

## Markdown Support
Welcome Text supports standard Markdown syntax, including headings, lists, links, and bold text. This helps users quickly understand what the app can help with.
## Quick Questions
Welcome Text supports the special `[Quick Question]` format. FastGPT displays these items as clickable buttons, and users can send the preset question with one click.
Example:
```md
Hello, I can help you look up product information and support policies.
[How do I request support?]
[What scenarios does this product support?]
[Recommend a starter plan]
```
When users start a new conversation, they will see the welcome message and quick question buttons. After a user clicks a quick question, FastGPT sends that question as the user's message.

## Recommendations
* Keep the welcome text short and clear, focusing on what the app can solve.
* Use quick questions for common scenarios, and avoid adding too many options.
* If the app requires a specific input format, include an example in the welcome text.
file: ./content/guide/build/general/welcomeText.mdx
meta: {
"title": "对话开场白",
"description": "FastGPT 对话开场白配置说明"
}
对话开场白是每次新对话开始前由系统自动发送的欢迎词,适合用于介绍应用能力、说明提问范围,或提供常用入口。
## 配置入口
在应用编辑页中,找到 **对话开场白** 配置项,点击右侧的设置按钮,即可编辑开场白内容。

## Markdown 支持
开场白支持标准 Markdown 语法,可以使用标题、列表、链接、加粗等格式,让用户进入对话后快速理解当前应用可以处理的问题。
## 快捷问题
开场白支持使用 `[快捷问题]` 特殊格式。配置后,界面会将对应内容展示为可点击按钮,用户点击后即可直接发送预设问题。
例如:
```md
你好,我可以帮你查询产品信息和售后政策。
[如何申请售后?]
[产品支持哪些使用场景?]
[帮我推荐一个入门方案]
```
用户进入新对话后,会看到欢迎词和快捷问题按钮。点击快捷问题按钮后,系统会把该问题作为用户输入发送到对话中。

## 使用建议
* 开场白应简短明确,优先说明应用能解决什么问题。
* 快捷问题建议覆盖高频场景,避免一次配置过多导致用户难以选择。
* 如果应用依赖固定格式输入,可以在开场白中给出示例。
file: ./content/guide/build/publish/dingtalk.en.mdx
meta: {
"title": "DingTalk Bot Integration",
"description": "FastGPT DingTalk Bot Integration Tutorial"
}
Starting from version 4.8.16, FastGPT commercial edition supports direct DingTalk bot integration without additional APIs.
## 1. Create a DingTalk Internal Enterprise App
1. Create an internal enterprise app in the [DingTalk Developer Console](https://open-dev.dingtalk.com/fe/app).

2. Obtain the **Client ID** and **Client Secret**.

## 2. Add a Publishing Channel in FastGPT
In FastGPT, select the app you want to integrate. On the **Publishing Channels** page, create a new DingTalk bot publishing channel.
Enter the **Client ID** and **Client Secret** obtained earlier into the configuration dialog.

After creation, click the **Request URL** button and copy the callback address.
## 3. Add **Bot** Capability to the App
In the DingTalk Developer Console, click **Add App Capability** on the left sidebar, and add the **Bot** capability to the internal enterprise app you just created.

## 4. Configure Bot Callback Address
Click the **Bot** capability on the left sidebar, then set the **Message Receiving Mode** at the bottom to **HTTP Mode**, and paste the FastGPT callback address you copied earlier as the message receiving address.

After debugging, click **Publish**.
## 5. Publish the App
After the bot is published, you still need to publish the app version on the **Version Management and Publishing** page.

Click **Create New Version**, set the version number and description, then click save to publish.

Once the app is published, you can use the bot within your DingTalk enterprise. You can chat with the bot privately, or add the bot to a group and `@mention the bot` to start a conversation.

## FAQ
### How to start a new chat history
To reset your chat history, send a `Reset` message to the bot (case-sensitive), and the bot will start a new chat history.
file: ./content/guide/build/publish/dingtalk.mdx
meta: {
"title": "接入钉钉机器人教程",
"description": "FastGPT 接入钉钉机器人教程"
}
从 4.8.16 版本起,FastGPT 商业版支持直接接入钉钉机器人,无需额外的 API。
## 1. 创建钉钉企业内部应用
1. 在[钉钉开发者后台](https://open-dev.dingtalk.com/fe/app)创建企业内部应用。

2. 获取**Client ID**和**Client Secret**。

## 2. 为 FastGPT 添加发布渠道
在 FastGPT 中选择要接入的应用,在**发布渠道**页面,新建一个接入钉钉机器人的发布渠道。
将前面拿到的 **Client ID** 和 **Client Secret** 填入配置弹窗中。

创建完成后,点击**请求地址**按钮,然后复制回调地址。
## 3. 为应用添加**机器人**应用能力。
在钉钉开发者后台,点击左侧**添加应用能力**,为刚刚创建的企业内部应用添加 **机器人** 应用能力。

## 4. 配置机器人回调地址
点击左侧**机器人** 应用能力,然后将底部**消息接受模式**设置为**HTTP模式**,消息接收地址填入前面复制的 FastGPT 的回调地址。

调试完成后,点击**发布**。
## 5. 发布应用
机器人发布后,还需要在**版本管理与发布**页面发布应用版本。

点击**创建新版本**后,设置版本号和版本描述后点击保存发布即可。

应用发布后,即可在钉钉企业中使用机器人功能,可对机器人私聊。或者在群组添加机器人后`@机器人`,触发对话。

## FAQ
### 如何新开一个聊天记录
如果你想重置你的聊天记录,可以给机器人发送 `Reset` 消息(注意大小写),机器人会新开一个聊天记录。
file: ./content/guide/build/publish/feishu.en.mdx
meta: {
"title": "Lark Bot Integration",
"description": "FastGPT Lark Bot Integration Tutorial"
}
Starting from version 4.8.10, FastGPT commercial edition supports direct Lark bot integration without additional APIs.
## 1. Create a Lark App
Creating a free test enterprise makes debugging easier.
1. Create a custom enterprise app in the [Lark Open Platform](https://open.feishu.cn/app) developer console.

Add a **Bot** capability to the app.
## 2. Create a Publishing Channel in FastGPT
In FastGPT, select the app you want to integrate. On the Publishing Channels page, create a new Lark bot publishing channel and fill in the basic information.

## 3. Get App ID and App Secret
In the Lark Open Platform developer console, find the App ID and App Secret for the custom enterprise app you just created, and enter them in the FastGPT publishing channel dialog.

Enter both parameters in the FastGPT configuration dialog.

(Optional) In the Lark Open Platform developer console, go to Events & Callbacks -> Encryption Strategy to get the Encrypt Key, and enter it in the Lark bot integration dialog.

The Encrypt Key encrypts communication between Lark servers and FastGPT.
If using HTTPS, the Encrypt Key is not needed. If using HTTP, the Encrypt Key is recommended.
The Verification Token is generated by default for source verification. However, we use Lark's officially recommended, more secure verification method, so this configuration can be ignored.
## 4. Configure Callback URL
After creating the publishing channel, click **Request URL** and copy the corresponding request URL.
In the Lark console, click `Events & Callbacks` on the left sidebar, click the edit icon next to `Configure Subscription Method`, and paste the copied request URL into the input field.
| | | |
| --------------------------------- | --------------------------------- | -------------------------------- |
|  |  |  |
## 5. Configure Bot Callback Events and Permissions
* Add the `Receive Message` event
On the `Events & Callbacks` page, click `Add Event`.
Search for `Receive Message`, or directly search for `im.message.receive_v1`, find the `Receive Message v2.0` event, check it, and click `Confirm Add`.
After adding the event, add two permissions: click the corresponding permission, and a popup will prompt you to add permissions. Add the two permissions shown above.
| | |
| -------------------------------- | -------------------------------- |
|  |  |
It is not recommended to enable the two "legacy versions" shown above -- use the new version permissions instead.
* If "Read messages users send to the bot in private chats" is enabled, private messages sent to the bot will be forwarded to FastGPT
* If "Receive @bot message events in group chats" is enabled, messages @mentioning the bot in group chats will be forwarded to FastGPT
* If (not recommended) "Get all messages in groups" is enabled, all group chat messages will be forwarded to FastGPT
## 6. Configure Reply Message Permission
In the Lark console, click `Permission Management` on the left sidebar, enter `send message` in the search box, find the `Send messages as the app` permission, and enable it.

## 7. Publish the Bot
Click `Version Management & Publishing` on the left side of the Lark console to publish the bot.

You can then find your bot in the workspace. Next, add the bot to a group or chat with it privately.

## FAQ
### Sent a message but no response
1. Check if the Lark bot callback URL, permissions, etc. are configured correctly.
2. Check FastGPT chat logs to see if there is a corresponding question record.
3. If there is a record but Lark does not respond, the bot is missing the required permissions.
4. If there is no record, the app may have encountered an error. Try the simplest bot first. (Lark bots cannot accept global variables, files, or image content as input)
### How to start a new chat history
Lark bot chat history chatId comes from several sources:
1. Private chat window
2. Individual topics in Lark topic groups
3. In group chats, composed of group ID + personal ID.
To reset your chat history, send a `Reset` message to the bot (case-sensitive), and the bot will start a new chat history.
file: ./content/guide/build/publish/feishu.mdx
meta: {
"title": "接入飞书机器人教程",
"description": "FastGPT 接入飞书机器人教程"
}
从 4.8.10 版本起,FastGPT 商业版支持直接接入飞书机器人,无需额外的 API。
## 1. 申请飞书应用
开一个免费的测试企业更方便进行调试。
1. 在[飞书开放平台](https://open.feishu.cn/app)的开发者后台申请企业自建应用。

添加一个**机器人**应用。
## 2. 在 FastGPT 新建发布渠道
在fastgpt中选择想要接入的应用,在 发布渠道 页面,新建一个接入飞书机器人的发布渠道,填写好基础信息。

## 3. 获取应用的 App ID, App Secret 两个凭证
在飞书开放平台开发者后台,刚刚创建的企业自建应用中,找到 App ID 和 App Secret,填入 FastGPT 新建发布渠道的对话框里面。

填入两个参数到 FastGPT 配置弹窗中。

(可选)在飞书开放平台开发者后台,点击事件与回调 -> 加密策略 获取 Encrypt Key,并填入飞书机器人接入的对话框里面

Encrypt Key 用于加密飞书服务器与 FastGPT 之间通信。
建议如果使用 Https 协议,则不需要 Encrypt Key。如果使用 Http 协议通信,则建议使用 Encrypt Key
Verification Token 默认生成的这个 Token 用于校验来源。但我们使用飞书官方推荐的另一种更为安全的校验方式,因此可以忽略这个配置项。
## 4. 配置回调地址
新建好发布渠道后,点击**请求地址**,复制对应的请求地址。
在飞书控制台,点击左侧的 `事件与回调` ,点击`配置订阅方式`旁边的编辑 icon,粘贴刚刚复制的请求地址到输入框中。
| | | |
| ------------------------------ | ------------------------------ | ----------------------------- |
|  |  |  |
## 5. 配置机器人回调事件和权限
* 添加 `接收消息` 事件
在`事件与回调`页面,点击`添加事件`。
搜索`接收消息`,或者直接搜索 `im.message.receive_v1` ,找到`接收消息 v2.0`的时间,勾选上并点击`确认添加`。
添加事件后,增加两个权限:点击对应权限,会有弹窗提示添加权限,添加上图两个权限。
| | |
| ----------------------------- | ----------------------------- |
|  |  |
不推荐启用上图中的两个“历史版本”,而是使用新版本的权限。
* 若开启 “读取用户发给机器人的单聊消息”, 则单聊发送给机器人的消息将被送到 FastGPT
* 若开启 “接收群聊中@机器人消息事件”, 则群聊中@机器人的消息将被送到 FastGPT
* 若开启(不推荐开启)“获取群组中所有消息”,则群聊中所有消息都将被送到 FastGPT
## 6. 配置回复消息权限
在飞书控制台,点击左侧的 `权限管理` ,搜索框中输入`发消息`,找到`以应用的身份发消息`的权限,点击开通权限。

## 7. 发布机器人
点击飞书控制台左侧的`版本管理与发布`,即可发布机器人。

然后就可以在工作台里找到你的机器人啦。接下来就是把机器人拉进群组,或者单独与它对话。

## FAQ
### 发送了消息,没响应
1. 检查飞书机器人回调地址、权限等是否正确。
2. 查看 FastGPT 对话日志,是否有对应的提问记录
3. 如果有记录,飞书没回应,则是没给机器人开权限。
4. 如果没记录,则可能是应用运行报错了,可以先试试最简单的机器人。(飞书机器人无法输入全局变量、文件、图片内容)
### 如何新开一个聊天记录
飞书机器人的聊天记录 chatId 包含几种来源:
1. 私聊聊天框
2. 飞书话题群中单个话题
3. 群组聊天中,由群 id+个人id 组成。
如果你想重置你的聊天记录,可以给机器人发送 `Reset` 消息(注意大小写),机器人会新开一个聊天记录。
file: ./content/guide/build/publish/link.en.mdx
meta: {
"title": "Share Link Publishing",
"description": "FastGPT share link publishing"
}
## Introduction
A share link creates a temporary public URL that lets anyone on the internet use your app. FastGPT creates a temporary identity for each visitor to isolate users from each other. Usage is billed to the team that owns the app, so avoid sharing the link publicly unless intended.
## Usage Flow
### 1. Create a Link
Go to `App Details` -> `Publish Channels` -> `Share Link`, then create a new link.
Enter a name to create the link. The name is only used for display in the record list.

### 2. Copy the Link
Click Start Using to open the share link, then copy and share it as needed.

## Parameter Configuration
Some parameters are only available in the commercial edition.
* Name: The display name of the link record.
* Expiration time: The link becomes unavailable after this time.
* QPM: The maximum number of requests per minute for each user.
* Credit limit: The maximum billable usage generated by this link.
* Identity verification: Used to integrate with third-party systems for identity authentication and chat callbacks.
* Real-time running status: Whether to show currently running nodes.
* View quoted chunks: See the system introduction.
* View full quoted content: See the system introduction.
* Download/open original source: See the system introduction.
## Share Link Authentication
### Introduction
In FastGPT V4.6.4, we changed how share links read data. A `localId` is generated for each user to identify them and pull chat history from the cloud. However, this only works on the same device and browser -- switching devices or clearing browser cache will lose those records. Due to this limitation, we only allow users to pull the last `20` records from the past `30 days`.
Share link authentication is designed to quickly and securely integrate FastGPT's chat interface into your existing system with just 2 endpoints. This feature is only available in the commercial edition.
### Usage Guide
In the share link configuration, you can optionally fill in the `Identity Verification` field. This is the root URL for a `POST` request. Once configured, share link initialization, chat start, and chat completion will all send requests to specific endpoints under this URL. Below, we use `host` to represent the `identity verification root URL`. Your server only needs to return whether verification succeeded -- no other data is required. The format is as follows:
#### Unified Response Format
```jsonc
{
"success": true,
"message": "Error message",
"msg": "Same as message, error message",
"data": {
"uid": "Unique user identifier" // Required
}
}
```
`FastGPT` checks whether `success` is `true` to decide if the user can proceed. `message` and `msg` are equivalent -- you can return either one. When `success` is not `true`, this error message will be displayed to the user.
`uid` is the unique user identifier and must be returned. The ID format must be a string that does not contain `|`, `/`, or `\\` characters, with a length of 255 **bytes** or less. Otherwise, an `Invalid UID` error will be returned. The `uid` is used to pull and save chat history -- see the practical example below.
#### Flow Diagram

### Configuration Guide
#### 1. Configure the Identity Verification URL

Once configured, every time the share link is used, verification and reporting requests will be sent to the corresponding endpoints.
You only need to configure the root URL here -- no need to specify the full request path.
#### 2. Add an Extra Query Parameter to the Share Link
Add an extra parameter `authToken` to the share link URL. For example:
Original link: `https://share.fastgpt.io/chat/share?shareId=648aaf5ae121349a16d62192`
Full link: `https://share.fastgpt.io/chat/share?shareId=648aaf5ae121349a16d62192&authToken=userid12345`
This `authToken` is typically a unique user credential (such as a token) generated by your system. FastGPT will include `token=[authToken]` in the `body` of the verification request.
#### 3. Implement the Chat Initialization Verification Endpoint
```bash
curl --location --request POST '{{host}}/shareAuth/init' \
--header 'Content-Type: application/json' \
--data-raw '{
"token": "[authToken]"
}'
```
```json
{
"success": true,
"data": {
"uid": "Unique user identifier"
}
}
```
The system will pull chat history for uid `username123` under this share link.
```json
{
"success": false,
"message": "Authentication failed"
}
```
#### 4. Implement the Pre-Chat Verification Endpoint
```bash
curl --location --request POST '{{host}}/shareAuth/start' \
--header 'Content-Type: application/json' \
--data-raw '{
"token": "[authToken]",
"question": "User question",
}'
```
```json
{
"success": true,
"data": {
"uid": "Unique user identifier"
}
}
```
```json
{
"success": false,
"message": "Authentication failed"
}
```
```json
{
"success": false,
"message": "Content policy violation"
}
```
#### 5. Implement the Chat Result Reporting Endpoint (Optional)
This endpoint has no required response format.
The response data follows the same format as the [chat endpoint](../../../openapi/intro.en.mdx#response), with an additional `token` field.
Key fields to note: `totalPoints` (total AI credits consumed), `token` (total token consumption)
```bash
curl --location --request POST '{{host}}/shareAuth/finish' \
--header 'Content-Type: application/json' \
--data-raw '{
"token": "[authToken]",
"responseData": [
{
"moduleName": "core.module.template.Dataset search",
"moduleType": "datasetSearchNode",
"totalPoints": 1.5278,
"query": "导演是谁\n《铃芽之旅》的导演是谁?\n这部电影的导演是谁?\n谁是《铃芽之旅》的导演?",
"model": "Embedding-2(旧版,不推荐使用)",
"tokens": 1524,
"similarity": 0.83,
"limit": 400,
"searchMode": "embedding",
"searchUsingReRank": false,
"extensionModel": "FastAI-4k",
"extensionResult": "《铃芽之旅》的导演是谁?\n这部电影的导演是谁?\n谁是《铃芽之旅》的导演?",
"runningTime": 2.15
},
{
"moduleName": "AI 对话",
"moduleType": "chatNode",
"totalPoints": 0.593,
"model": "FastAI-4k",
"tokens": 593,
"query": "导演是谁",
"maxToken": 2000,
"quoteList": [
{
"id": "65bb346a53698398479a8854",
"q": "导演是谁?",
"a": "电影《铃芽之旅》的导演是新海诚。",
"chunkIndex": 0,
"datasetId": "65af9b947916ae0e47c834d2",
"collectionId": "65bb345c53698398479a868f",
"sourceName": "dataset - 2024-01-23T151114.198.csv",
"sourceId": "65bb345b53698398479a868d",
"score": [
{
"type": "embedding",
"value": 0.9377183318138123,
"index": 0
},
{
"type": "rrf",
"value": 0.06557377049180328,
"index": 0
}
]
}
],
"historyPreview": [
{
"obj": "Human",
"value": "使用 标记中的内容作为本次对话的参考:\n\n\n导演是谁?\n电影《铃芽之旅》的导演是新海诚。\n------\n电影《铃芽之旅》的编剧是谁?22\n新海诚是本片的编剧。\n------\n电影《铃芽之旅》的女主角是谁?\n电影的女主角是铃芽。\n------\n电影《铃芽之旅》的制作团队中有哪位著名人士?2\n川村元气是本片的制作团队成员之一。\n------\n你是谁?\n我是电影《铃芽之旅》助手\n------\n电影《铃芽之旅》男主角是谁?\n电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。\n------\n电影《铃芽之旅》的作者新海诚写了一本小说,叫什么名字?\n小说名字叫《铃芽之旅》。\n------\n电影《铃芽之旅》的女主角是谁?\n电影《铃芽之旅》的女主角是岩户铃芽,由原菜乃华配音。\n------\n电影《铃芽之旅》的故事背景是什么?\n日本\n------\n谁担任电影《铃芽之旅》中岩户环的配音?\n深津绘里担任电影《铃芽之旅》中岩户环的配音。\n\n\n回答要求:\n- 如果你不清楚答案,你需要澄清。\n- 避免提及你是从 获取的知识。\n- 保持答案与 中描述的一致。\n- 使用 Markdown 语法优化回答格式。\n- 使用与问题相同的语言回答。\n\n问题:\"\"\"导演是谁\"\"\""
},
{
"obj": "AI",
"value": "电影《铃芽之旅》的导演是新海诚。"
}
],
"contextTotalLen": 2,
"runningTime": 1.32
}
]
}'
```
**Full responseData Field Reference:**
```ts
type ResponseType = {
moduleType: FlowNodeTypeEnum; // Module type
moduleName: string; // Module name
moduleLogo?: string; // Logo
runningTime?: number; // Running time
query?: string; // User question / search query
textOutput?: string; // Text output
tokens?: number; // Total context tokens
model?: string; // Model used
contextTotalLen?: number; // Total context length
totalPoints?: number; // Total AI credits consumed
temperature?: number; // Temperature
maxToken?: number; // Model max tokens
quoteList?: SearchDataResponseItemType[]; // Citation list
historyPreview?: ChatItemMiniType[]; // Context preview (history may be truncated)
similarity?: number; // Minimum similarity threshold
limit?: number; // Max citation tokens
searchMode?: `${DatasetSearchModeEnum}`; // Search mode
searchUsingReRank?: boolean; // Whether rerank is used
extensionModel?: string; // Query expansion model
extensionResult?: string; // Query expansion result
extensionTokens?: number; // Query expansion total token length
cqList?: ClassifyQuestionAgentItemType[]; // Question classification list
cqResult?: string; // Question classification result
extractDescription?: string; // Content extraction description
extractResult?: Record; // Content extraction result
params?: Record; // HTTP module params
body?: Record; // HTTP module body
headers?: Record; // HTTP module headers
httpResult?: Record; // HTTP module result
pluginOutput?: Record; // Plugin output
pluginDetail?: ChatHistoryItemResType[]; // Plugin details
isElseResult?: boolean; // Conditional result
};
```
### Practical Example
We'll use [Laf as the server](https://laf.dev/) to demonstrate how these 3 endpoints work.
#### 1. Create 3 Laf Endpoints

In this endpoint, we require `token` to equal `fastgpt` to pass verification. (Not recommended for production -- avoid hardcoding values.)
```ts
import cloud from '@lafjs/cloud';
export default async function (ctx: FunctionContext) {
const { token } = ctx.body;
// Token decoding logic omitted
if (token === 'fastgpt') {
return { success: true, data: { uid: 'user1' } };
}
return { success: false, message: 'Authentication failed' };
}
```
In this endpoint, we require `token` to equal `fastgpt` to pass verification. Additionally, if the question contains a specific character, it returns an error to simulate content moderation.
```ts
import cloud from '@lafjs/cloud';
export default async function (ctx: FunctionContext) {
const { token, question } = ctx.body;
// Token decoding logic omitted
if (token !== 'fastgpt') {
return { success: false, message: 'Authentication failed' };
}
if (question.includes('你')) {
return { success: false, message: 'Content policy violation' };
}
return { success: true, data: { uid: 'user1' } };
}
```
The result reporting endpoint can handle custom logic as needed.
```ts
import cloud from '@lafjs/cloud';
export default async function (ctx: FunctionContext) {
const { token, responseData } = ctx.body;
const total = responseData.reduce((sum, item) => sum + item.price, 0);
const amount = total / 100000;
// Database operations omitted
return {};
}
```
#### 2. Configure the Verification URL
Copy any of the 3 endpoint URLs, e.g. `https://d8dns0.laf.dev/shareAuth/finish`, remove the `/shareAuth/finish` part, and enter the root URL `https://d8dns0.laf.dev` in the `Identity Verification` field.

#### 3. Modify the Share Link Parameters
Original share link: `https://share.fastgpt.io/chat/share?shareId=64be36376a438af0311e599c`
Modified: `https://share.fastgpt.io/chat/share?shareId=64be36376a438af0311e599c&authToken=fastgpt`
#### 4. Test the Result
1. Opening the original link or a link where `authToken` does not equal `fastgpt` will show an authentication error.
2. Sending content that contains the filtered character will show a content policy violation error.
### Use Cases
This authentication method is typically used to embed the `share link` directly into your application. Before opening the share link in your app, you should append the `authToken` parameter.
Beyond integrating with your existing user system, you can also implement a `balance` feature -- deduct user balance via the `result reporting` endpoint and check user balance via the `pre-chat verification` endpoint.
file: ./content/guide/build/publish/link.mdx
meta: {
"title": "免登录窗口发布",
"description": "FastGPT 免登录窗口发布"
}
## 介绍
免登录窗口可以创建一个临时可访问的地址,任何互联网上的用户可以通过该地址来使用应用。产生的费用产生在应用归属的团队下,所以请注意不要随意分享。
系统会为每个用户生成一个 localId,用于标识用户,从云端拉取对话记录。但是这种方式仅能保障用户在同一设备同一浏览器中使用,如果切换设备或者清空浏览器缓存则会丢失这些记录。这种方式存在一定的风险,因此我们仅允许用户拉取近 `30天` 的 `20条` 记录。
## 使用流程
### 1. 创建链接
在 `应用详情` - `发布渠道` - `免登录窗口` 中,创建新链接。
填写名称后即可创建,该名称仅用于记录展示。

### 2. 复制链接
点击开始使用,即可打开使用链接,复制后即可使用。

## 参数配置
部分参数仅商业版支持配置。
* 名称:仅用于记录该链接的名称,用于展示
* 过期时间:超过该时间后,链接无法使用。
* QPM: 每个用户每分钟最大访问次数。
* 积分上限:该链接产生的最大计费数据。
* 身份验证:便于接入第三方系统进行身份认证和对话回调。
* 实时运行状态:是否展示当前运行的节点。
* 查看引用片段:见系统介绍
* 查看引用全文:见系统介绍
* 下载/打开来源原文:见系统介绍
## 身份验证说明
分享链接身份验证设计的目的在于,将 FastGPT 的对话框快速、安全的接入到你现有的系统中,仅需 2 个接口即可实现。该功能目前只在商业版中提供。
### 使用说明
免登录链接配置中,你可以选择填写 `身份验证` 栏。这是一个 `POST` 请求的根地址。在填写该地址后,分享链接的初始化、开始对话以及对话结束都会向该地址的特定接口发送一条请求。下面以 `host` 来表示 `凭身份验证根地址`。服务器接口仅需返回是否校验成功即可,不需要返回其他数据,格式如下:
#### 接口统一响应格式
```jsonc
{
"success": true,
"message": "错误提示",
"msg": "同message, 错误提示",
"data": {
"uid": "用户唯一凭证" // 必须返回
}
}
```
`FastGPT` 将会判断 `success` 是否为 `true` 决定是允许用户继续操作。`message` 与 `msg` 是等同的,你可以选择返回其中一个,当 `success` 不为 `true` 时,将会提示这个错误。
`uid` 是用户的唯一凭证,必须返回该 ID 且 ID 的格式为不包含 "|"、"/“、"\\" 字符的、小于等于 255 **字节长度**的字符串,否则会返回 `Invalid UID` 的错误。`uid` 将会用于拉取对话记录以及保存对话记录,可参考下方实践案例。
#### 触发流程

### 配置教程
#### 1. 配置身份校验地址

配置校验地址后,在每次分享链接使用时,都会向对应的地址发起校验和上报请求。
这里仅需配置根地址,无需具体到完整请求路径。
#### 2. 分享链接中增加额外 query
在分享链接的地址中,增加一个额外的参数: authToken。例如:
原始的链接:`https://share.fastgpt.io/chat/share?shareId=648aaf5ae121349a16d62192`
完整链接: `https://share.fastgpt.io/chat/share?shareId=648aaf5ae121349a16d62192&authToken=userid12345`
这个 `authToken` 通常是你系统生成的用户唯一凭证(Token 之类的)。FastGPT 会在鉴权接口的 `body` 中携带 token=\[authToken] 的参数。
#### 3. 编写聊天初始化校验接口
```bash
curl --location --request POST '{{host}}/shareAuth/init' \
--header 'Content-Type: application/json' \
--data-raw '{
"token": "[authToken]"
}'
```
```json
{
"success": true,
"data": {
"uid": "用户唯一凭证"
}
}
```
系统会拉取该分享链接下,uid 为 username123 的对话记录。
```json
{
"success": false,
"message": "身份错误"
}
```
#### 4. 编写对话前校验接口
```bash
curl --location --request POST '{{host}}/shareAuth/start' \
--header 'Content-Type: application/json' \
--data-raw '{
"token": "[authToken]",
"question": "用户问题",
}'
```
```json
{
"success": true,
"data": {
"uid": "用户唯一凭证"
}
}
```
```json
{
"success": false,
"message": "身份验证失败"
}
```
```json
{
"success": false,
"message": "存在违规词"
}
```
#### 5. 编写对话结果上报接口(可选)
该接口无规定返回值。
响应值与 [chat 接口格式相同](../../../openapi/intro.mdx#响应),仅多了一个 `token`。
重点关注:`totalPoints` (总消耗 AI 积分),`token` (Token 消耗总数)
```bash
curl --location --request POST '{{host}}/shareAuth/finish' \
--header 'Content-Type: application/json' \
--data-raw '{
"token": "[authToken]",
"responseData": [
{
"moduleName": "core.module.template.Dataset search",
"moduleType": "datasetSearchNode",
"totalPoints": 1.5278,
"query": "导演是谁\n《铃芽之旅》的导演是谁?\n这部电影的导演是谁?\n谁是《铃芽之旅》的导演?",
"model": "Embedding-2(旧版,不推荐使用)",
"tokens": 1524,
"similarity": 0.83,
"limit": 400,
"searchMode": "embedding",
"searchUsingReRank": false,
"extensionModel": "FastAI-4k",
"extensionResult": "《铃芽之旅》的导演是谁?\n这部电影的导演是谁?\n谁是《铃芽之旅》的导演?",
"runningTime": 2.15
},
{
"moduleName": "AI 对话",
"moduleType": "chatNode",
"totalPoints": 0.593,
"model": "FastAI-4k",
"tokens": 593,
"query": "导演是谁",
"maxToken": 2000,
"quoteList": [
{
"id": "65bb346a53698398479a8854",
"q": "导演是谁?",
"a": "电影《铃芽之旅》的导演是新海诚。",
"chunkIndex": 0,
"datasetId": "65af9b947916ae0e47c834d2",
"collectionId": "65bb345c53698398479a868f",
"sourceName": "dataset - 2024-01-23T151114.198.csv",
"sourceId": "65bb345b53698398479a868d",
"score": [
{
"type": "embedding",
"value": 0.9377183318138123,
"index": 0
},
{
"type": "rrf",
"value": 0.06557377049180328,
"index": 0
}
]
}
],
"historyPreview": [
{
"obj": "Human",
"value": "使用 标记中的内容作为本次对话的参考:\n\n\n导演是谁?\n电影《铃芽之旅》的导演是新海诚。\n------\n电影《铃芽之旅》的编剧是谁?22\n新海诚是本片的编剧。\n------\n电影《铃芽之旅》的女主角是谁?\n电影的女主角是铃芽。\n------\n电影《铃芽之旅》的制作团队中有哪位著名人士?2\n川村元气是本片的制作团队成员之一。\n------\n你是谁?\n我是电影《铃芽之旅》助手\n------\n电影《铃芽之旅》男主角是谁?\n电影《铃芽之旅》男主角是宗像草太,由松村北斗配音。\n------\n电影《铃芽之旅》的作者新海诚写了一本小说,叫什么名字?\n小说名字叫《铃芽之旅》。\n------\n电影《铃芽之旅》的女主角是谁?\n电影《铃芽之旅》的女主角是岩户铃芽,由原菜乃华配音。\n------\n电影《铃芽之旅》的故事背景是什么?\n日本\n------\n谁担任电影《铃芽之旅》中岩户环的配音?\n深津绘里担任电影《铃芽之旅》中岩户环的配音。\n\n\n回答要求:\n- 如果你不清楚答案,你需要澄清。\n- 避免提及你是从 获取的知识。\n- 保持答案与 中描述的一致。\n- 使用 Markdown 语法优化回答格式。\n- 使用与问题相同的语言回答。\n\n问题:\"\"\"导演是谁\"\"\""
},
{
"obj": "AI",
"value": "电影《铃芽之旅》的导演是新海诚。"
}
],
"contextTotalLen": 2,
"runningTime": 1.32
}
]
}'
```
**responseData 完整字段说明:**
```ts
type ResponseType = {
moduleType: FlowNodeTypeEnum; // 模块类型
moduleName: string; // 模块名
moduleLogo?: string; // logo
runningTime?: number; // 运行时间
query?: string; // 用户问题/检索词
textOutput?: string; // 文本输出
tokens?: number; // 上下文总Tokens
model?: string; // 使用到的模型
contextTotalLen?: number; // 上下文总长度
totalPoints?: number; // 总消耗AI积分
temperature?: number; // 温度
maxToken?: number; // 模型的最大token
quoteList?: SearchDataResponseItemType[]; // 引用列表
historyPreview?: ChatItemMiniType[]; // 上下文预览(历史记录会被裁剪)
similarity?: number; // 最低相关度
limit?: number; // 引用上限token
searchMode?: `${DatasetSearchModeEnum}`; // 搜索模式
searchUsingReRank?: boolean; // 是否使用rerank
extensionModel?: string; // 问题扩展模型
extensionResult?: string; // 问题扩展结果
extensionTokens?: number; // 问题扩展总字符长度
cqList?: ClassifyQuestionAgentItemType[]; // 分类问题列表
cqResult?: string; // 分类问题结果
extractDescription?: string; // 内容提取描述
extractResult?: Record; // 内容提取结果
params?: Record; // HTTP模块params
body?: Record; // HTTP模块body
headers?: Record; // HTTP模块headers
httpResult?: Record; // HTTP模块结果
pluginOutput?: Record; // 插件输出
pluginDetail?: ChatHistoryItemResType[]; // 插件详情
isElseResult?: boolean; // 判断器结果
};
```
### 实践案例
我们以 [Laf 作为服务器为例](https://laf.dev/),简单展示这 3 个接口的使用方式。
#### 1. 创建 3 个 Laf 接口

这个接口中,我们设置了 `token` 必须等于 `fastgpt` 才能通过校验。(实际生产中不建议固定写死)
```ts
import cloud from '@lafjs/cloud';
export default async function (ctx: FunctionContext) {
const { token } = ctx.body;
// 此处省略 token 解码过程
if (token === 'fastgpt') {
return { success: true, data: { uid: 'user1' } };
}
return { success: false, message: '身份错误' };
}
```
这个接口中,我们设置了 `token` 必须等于 `fastgpt` 才能通过校验。并且如果问题中包含了 `你` 字,则会报错,用于模拟敏感校验。
```ts
import cloud from '@lafjs/cloud';
export default async function (ctx: FunctionContext) {
const { token, question } = ctx.body;
// 此处省略 token 解码过程
if (token !== 'fastgpt') {
return { success: false, message: '身份错误' };
}
if (question.includes('你')) {
return { success: false, message: '内容不合规' };
}
return { success: true, data: { uid: 'user1' } };
}
```
结果上报接口可自行进行逻辑处理。
```ts
import cloud from '@lafjs/cloud';
export default async function (ctx: FunctionContext) {
const { token, responseData } = ctx.body;
const total = responseData.reduce((sum, item) => sum + item.price, 0);
const amount = total / 100000;
// 省略数据库操作
return {};
}
```
#### 2. 配置校验地址
我们随便复制 3 个地址中一个接口: `https://d8dns0.laf.dev/shareAuth/finish` , 去除 `/shareAuth/finish` 后填入 `身份校验` : `https://d8dns0.laf.dev`

#### 3. 修改分享链接参数
源分享链接:`https://share.fastgpt.io/chat/share?shareId=64be36376a438af0311e599c`
修改后:`https://share.fastgpt.io/chat/share?shareId=64be36376a438af0311e599c&authToken=fastgpt`
#### 4. 测试效果
1. 打开源链接或者 `authToken` 不等于 `fastgpt` 的链接会提示身份错误。
2. 发送内容中包含你字,会提示内容不合规。
### 使用场景
这个鉴权方式通常是帮助你直接嵌入 `分享链接` 到你的应用中,在你的应用打开分享链接前,应做 `authToken` 的拼接后再打开。
除了对接已有系统的用户外,你还可以对接 `余额` 功能,通过 `结果上报` 接口扣除用户余额,通过 `对话前校验` 接口检查用户的余额。
file: ./content/guide/build/publish/mcp_server.en.mdx
meta: {
"title": "MCP Server",
"description": "A quick overview of FastGPT MCP Server"
}
## What is MCP Server?
MCP (Model Context Protocol) was released by Anthropic in early November 2024. It standardizes communication between AI models and external systems, simplifying integration. With OpenAI officially supporting MCP, more and more AI vendors are adopting the protocol.
MCP has two main components: Client and Server. The Client is the AI model consumer — it uses MCP Client to give the model the ability to call external systems. The Server provides and runs those external system integrations.
FastGPT's MCP Server feature lets you select `multiple` applications built on FastGPT and expose them via MCP protocol for external consumption.
Currently, FastGPT's MCP Server uses the SSE transport protocol, with plans to migrate to `HTTP Streamable` in the future.
## Using MCP Server in FastGPT
### 1. Create an MCP Server
After logging into FastGPT, open `Workspace` and click `MCP Server` to access the management page. Here you can see all your MCP Servers and the number of applications each one manages.

You can customize the MCP Server name and select which applications to associate.
| | |
| -------------------------- | -------------------------- |
|  |  |
### 2. Get the MCP Server URL
After creating an MCP Server, click `Start Using` to get the access URL.
| | |
| -------------------------- | -------------------------- |
|  |  |
#### 3. Use the MCP Server
Use the URL in any MCP-compatible client to call your FastGPT applications — for example, `Cursor` or `Cherry Studio`. Here's how to set it up in Cursor.
Open Cursor's settings page and click MCP to enter the MCP configuration page. Click the new MCP Server button to open a JSON configuration file. Paste the `integration script` from step 2 into the `JSON file` and save.
Return to Cursor's MCP management page and you'll see your MCP Server listed. Make sure to set it to `enabled`.
| | | |
| -------------------------- | -------------------------- | -------------------------- |
|  |  |  |
Open Cursor's chat panel and switch to `Agent` mode — only this mode triggers MCP Server calls.
After sending a question about `fastgpt`, you'll see Cursor invoke an MCP tool (described as: query fastgpt knowledge base), which calls the FastGPT application to process the question and return results.
| | |
| -------------------------- | --------------------------- |
|  |  |
## Self-Hosted MCP Server Setup
Self-hosted FastGPT deployments require version `v4.9.6` or higher to use MCP Server.
### Update docker-compose.yml
Add the `fastgpt-mcp-server` service to your `docker-compose.yml`:
```yml
fastgpt-mcp-server:
container_name: fastgpt-mcp-server
image: ghcr.io/labring/fastgpt-mcp_server:latest
ports:
- 3005:3000
networks:
- fastgpt
restart: always
environment:
- FASTGPT_ENDPOINT=http://fastgpt:3000
```
### Update FastGPT Container Environment Variables
Configure `SSE_MCP_SERVER_PROXY_ENDPOINT` in the FastGPT container. Set it to the client-accessible `fastgpt-mcp-server` URL without a trailing `/`. For example:
```yaml
environment:
SSE_MCP_SERVER_PROXY_ENDPOINT: https://mcp.fastgpt.cn
```
### Restart FastGPT
Restart FastGPT after changing the environment variable:
```bash
docker-compose down
docker-compose up -d
```
After restarting, the MCP Server option will appear in the Workspace.
file: ./content/guide/build/publish/mcp_server.mdx
meta: {
"title": "MCP 发布",
"description": "快速了解 FastGPT MCP server"
}
## MCP server 介绍
MCP 协议(Model Context Protocol),是由 Anthropic 在 2024 年 11 月初发布的协议。它的目的在于统一 AI 模型与外部系统之间的通信方式,从而简化 AI 模型与外部系统之间的通信问题。随着 OpenAI 官宣支持 MCP 协议,越来越多的 AI 厂商开始支持 MCP 协议。
MCP 协议主要包含 Client 和 Server 两部分。简单来说,Client 是使用 AI 模型的一方,它通过 MCP Client 可以给模型提供一些调用外部系统的能能力;Server 是提供外部系统调用的一方,也就是实际运行外部系统的一方。
FastGPT MCP Server 功能允许你选择 `多个` 在 FastGPT 上构建好的应用,以 MCP 协议对外提供调用 FastGPT 应用的能力。
目前 FastGPT 提供的 MCP server 为 SSE 通信协议,未来将会替换成 `HTTP streamable`。
## FastGPT 使用 MCP server
### 1. 创建 MCP server
登录 FastGPT 后,打开 `工作台`,点击 `MCP server`,即可进入管理页面,这里可以看到你创建的所有 MCP server,以及他们管理的应用数量。

可以自定义 MCP server 名称和选择关联的应用
| | |
| -------------------------- | -------------------------- |
|  |  |
### 2. 获取 MCP server 地址
创建好 MCP server 后,可以直接点击 `开始使用`,即可获取 MCP server 访问地址。
| | |
| -------------------------- | -------------------------- |
|  |  |
#### 3. 使用 MCP server
可以在支持 MCP 协议的客户端使用这些地址,来调用 FastGPT 应用,例如:`Cursor`、`Cherry Studio`。下面以 Cursor 为例,介绍如何使用 MCP server。
打开 Cursor 配置页面,点击 MCP 即可进入 MCP 配置页面,可以点击新建 MCP server 按钮,会跳转到一个 JSON 配置文件,将第二步的 `接入脚本` 复制到 `json 文件` 中,保存文件。
此时返回 Cursor 的 MCP 管理页面,即可看到你创建的 MCP server,记得设成 `enabled` 状态。
| | | |
| -------------------------- | -------------------------- | -------------------------- |
|  |  |  |
打开 Cursor 的对话框,切换成 `Agent` 模型,只有这个模型,cursor 才会调用 MCP server。\
发送一个关于 `fastgpt` 的问题后,可以看到,cursor 调用了一个 MCP 工具(描述为:查询 fastgpt 知识库),也就是调用 FastGPT 应用去进行处理该问题,并返回了结果。
| | |
| -------------------------- | --------------------------- |
|  |  |
## 私有化部署 MCP server 问题
私有化部署版本的 FastGPT,需要升级到 `v4.9.6` 及以上版本才可使用 MCP server 功能。
### 修改 docker-compose.yml 文件
在 `docker-compose.yml` 文件中,加入 `fastgpt-mcp-server` 服务:
```yml
fastgpt-mcp-server:
container_name: fastgpt-mcp-server
image: ghcr.io/labring/fastgpt-mcp_server:latest
ports:
- 3005:3000
networks:
- fastgpt
restart: always
environment:
- FASTGPT_ENDPOINT=http://fastgpt:3000
```
### 修改 FastGPT 容器环境变量
在 FastGPT 容器中配置 `SSE_MCP_SERVER_PROXY_ENDPOINT`,值为客户端可访问的 `fastgpt-mcp-server` 地址,末尾不要携带 `/`,例如:
```yaml
environment:
SSE_MCP_SERVER_PROXY_ENDPOINT: https://mcp.fastgpt.cn
```
### 重启 FastGPT 容器
修改环境变量后,需要重启 FastGPT 服务。启动后,可以在工作台看到 MCP server 服务选项。
```bash
docker-compose down
docker-compose up -d
```
file: ./content/guide/build/publish/official_account.en.mdx
meta: {
"title": "WeChat Official Account Integration",
"description": "FastGPT WeChat Official Account Integration Tutorial"
}
Starting from version 4.8.10, FastGPT commercial edition supports direct WeChat Official Account integration without additional APIs.
**Note: Currently only verified official accounts are supported (both Service Accounts and Subscription Accounts).**
## 1. Create a Publishing Channel in FastGPT
In FastGPT, select the app you want to integrate. On the *Publishing Channels* page, create a new WeChat Official Account publishing channel and fill in the basic information.

## 2. Get AppID, Secret, and Token
### 1. Log in to the WeChat Official Account Platform and select your account.
Open the WeChat Official Account website: [https://mp.weixin.qq.com](https://mp.weixin.qq.com)
**Only verified official accounts are supported. Unverified accounts are not currently supported.**
Developers can apply for a WeChat Official Account test account from this link for testing. Test accounts work normally but cannot configure AES Key.

### 2. Enter the 3 parameters into the FastGPT configuration dialog.

## 3. Add FastGPT IP to IP Whitelist

Self-hosted users can check their own IP address.
International edition users (cloud.fastgpt.io) can add the following IP whitelist:
```
35.240.227.100
34.124.237.188
34.143.240.160
34.87.51.146
34.87.79.202
35.247.163.68
34.87.102.86
35.198.192.104
34.126.163.205
34.124.189.116
34.143.149.171
34.87.173.252
34.142.157.52
34.87.180.104
34.87.20.189
34.87.110.152
34.87.44.74
34.87.152.33
35.197.149.75
35.247.161.35
```
China Mainland users (fastgpt.cn) can add the following IP whitelist:
```
47.97.1.240
121.43.105.217
121.41.178.7
121.40.65.187
47.97.59.172
101.37.205.32
120.55.195.90
120.26.229.115
120.55.193.112
47.98.190.173
112.124.41.79
121.196.235.183
121.41.75.88
121.43.108.48
112.124.12.6
121.43.52.222
121.199.162.43
121.199.162.102
120.55.94.163
47.99.59.223
112.124.46.5
121.40.46.247
120.26.145.73
120.26.147.199
121.43.125.163
121.196.228.45
121.43.126.202
120.26.144.37
```
## 4. Get AES Key and Select Encryption Mode


1. Randomly generate an AES Key and enter it into the FastGPT configuration dialog.
2. Select the encryption mode as Secure Mode.
## 5. Get URL
1. Confirm creation in FastGPT and get the URL.

2. Enter it in the URL field on the WeChat Official Account Platform, then submit and save.

## 6. Enable Server Configuration (Skip if already auto-enabled)

## 7. Start Using
Now when users send messages to the official account, messages will be forwarded to FastGPT, and conversation results will be returned through the official account.
## FAQ
### How to start a new chat history
To reset your chat history, send a `Reset` message to the bot (case-sensitive), and the bot will start a new chat history.
file: ./content/guide/build/publish/official_account.mdx
meta: {
"title": "接入微信公众号教程",
"description": "FastGPT 接入微信公众号教程"
}
从 4.8.10 版本起,FastGPT 商业版支持直接接入微信公众号,无需额外的 API。
**注意⚠️: 目前只支持通过验证的公众号(服务号和订阅号都可以)**
## 1. 在 FastGPT 新建发布渠道
在 FastGPT 中选择想要接入的应用,在 *发布渠道* 页面,新建一个接入微信公众号的发布渠道,填写好基础信息。

## 2. 获取 AppID 、 Secret和Token
### 1. 登录微信公众平台,选择您的公众号。
打开微信公众号官网:[https://mp.weixin.qq.com](https://mp.weixin.qq.com)
**只支持通过验证的公众号,未通过验证的公众号暂不支持。**
开发者可以从这个链接申请微信公众号的测试号进行测试,测试号可以正常使用,但不能配置 AES Key

### 2. 把3个参数填入 FastGPT 配置弹窗中。

## 3. 在 IP 白名单中加入 FastGPT 的 IP

私有部署的用户可自行查阅自己的 IP 地址。
国际版用户(cloud.fastgpt.io)可以填写下面的 IP 白名单:
```
35.240.227.100
34.124.237.188
34.143.240.160
34.87.51.146
34.87.79.202
35.247.163.68
34.87.102.86
35.198.192.104
34.126.163.205
34.124.189.116
34.143.149.171
34.87.173.252
34.142.157.52
34.87.180.104
34.87.20.189
34.87.110.152
34.87.44.74
34.87.152.33
35.197.149.75
35.247.161.35
```
中国大陆用户(fastgpt.cn)可以填写下面的 IP 白名单:
```
47.97.1.240
121.43.105.217
121.41.178.7
121.40.65.187
47.97.59.172
101.37.205.32
120.55.195.90
120.26.229.115
120.55.193.112
47.98.190.173
112.124.41.79
121.196.235.183
121.41.75.88
121.43.108.48
112.124.12.6
121.43.52.222
121.199.162.43
121.199.162.102
120.55.94.163
47.99.59.223
112.124.46.5
121.40.46.247
120.26.145.73
120.26.147.199
121.43.125.163
121.196.228.45
121.43.126.202
120.26.144.37
```
## 4. 获取AES Key,选择加密方式


1. 随机生成AESKey,填入 FastGPT 配置弹窗中。
2. 选择加密方式为安全模式。
## 5. 获取 URL
1. 在FastGPT确认创建,获取URL。

2. 填入微信公众平台的 URL 处,然后提交保存

## 6. 启用服务器配置(如已自动启用,请忽略)

## 7. 开始使用
现在用户向公众号发消息,消息则会被转发到 FastGPT,通过公众号返回对话结果。
## FAQ
### 如何新开一个聊天记录
如果你想重置你的聊天记录,可以给机器人发送 `Reset` 消息(注意大小写),机器人会新开一个聊天记录。
file: ./content/guide/build/publish/openapi.en.mdx
meta: {
"title": "Access App via API",
"description": "Access FastGPT app via API"
}
import { Alert } from '@/components/docs/Alert';
In FastGPT, the API entry under Publish Channels shows **API Keys** available to the current signed-in member. API Keys are member credentials for OpenAPI calls and are no longer created as app-scoped keys.
When calling this app through `chat/completions`, passing `appId` in the request body is recommended. If a third-party app only supports OpenAI SDK-style key configuration, you can use the `apiKey-appId` compatibility format. For details, [see the OpenAPI Introduction](../../../openapi/intro.en.mdx).
## Get an API Key
Go to App -> "Publish Channels" -> "API", then click "New" to create a key.
An API Key represents the current signed-in member's OpenAPI credential. Keep your key safe. To
copy it again later, use the copy button in the API list.

Tip: For security, you can set a quota or expiration time to prevent key abuse.
## Replace Variables in Third-Party Apps
```bash
OPENAI_API_BASE_URL: http://localhost:3000/api (replace with your deployed domain)
OPENAI_API_KEY = the key obtained in the previous step (passing appId in the request body is recommended; if the third-party app only accepts a key, use the apiKey-appId compatibility format)
```
**[ChatGPT Next Web](https://github.com/Yidadaa/ChatGPT-Next-Web) Example:**

**[ChatGPT Web](https://github.com/Chanzhaoyu/chatgpt-web) Example:**

file: ./content/guide/build/publish/openapi.mdx
meta: {
"title": "通过 API 访问应用",
"description": "通过 API 访问 FastGPT 应用"
}
import { Alert } from '@/components/docs/Alert';
在 FastGPT 中,发布渠道里的 API 入口展示当前登录成员可用的 **APIKey**。APIKey 是团队成员的开放接口调用凭证,不再按应用创建专属密钥。
调用当前应用的 `chat/completions` 接口时,推荐在请求体传入 `appId`。如果第三方应用只能配置 OpenAI SDK 风格的密钥,也可以使用 `apiKey-appId` 兼容格式。完整说明可以[查看 OpenAPI 介绍](../../../openapi/intro.mdx)。
## 获取 APIKey
依次选择应用 ->「发布渠道」->「API」,然后点击「新建」创建密钥。
APIKey 代表当前登录成员的开放接口调用凭证。请妥善保管密钥;如需再次复制,可在 API
列表中点击复制按钮。

Tips: 安全起见,你可以设置一个额度或者过期时间,防止 key 被滥用。
## 替换三方应用的变量
```bash
OPENAI_API_BASE_URL: http://localhost:3000/api (改成自己部署的域名)
OPENAI_API_KEY = 上一步获取到的密钥(推荐在请求体传 appId;如第三方应用只能配置密钥,可填 apiKey-appId 兼容格式)
```
**[ChatGPT Next Web](https://github.com/Yidadaa/ChatGPT-Next-Web) 示例:**

**[ChatGPT Web](https://github.com/Chanzhaoyu/chatgpt-web) 示例:**

file: ./content/guide/build/publish/wechat.en.mdx
meta: {
"title": "WeChat Personal Account Integration",
"description": "How to integrate FastGPT with a WeChat personal account"
}
## 1. Create a Publishing Channel
Open your FastGPT Agent, click the tab at the top to switch to **Publishing Channels**, select **WeChat Personal Account**, and click **Create**.

You can fill in the form fields as needed.
## 2. Scan QR Code to Log In
After creating the channel, a new entry will appear. For the first time, you need to scan a QR code to log in. Click **Scan QR Code to Log In** to display the login QR code.


## 3. Start Chatting
Once connected via QR code, a **WeChat ClawBot** will appear in your contacts list.

Click to open the conversation and start chatting!

## FAQ
### How to Reset a Chat
Type `Reset` or `/reset` in the input field to clear the chat history.
file: ./content/guide/build/publish/wechat.mdx
meta: {
"title": "接入微信个人号教程",
"description": "FastGPT 接入微信个人号教程"
}
## 1. 新建发布渠道
进入 FastGPT 搭建好的 Agent,点击顶部的 tab 切换到发布渠道,并选择`微信个人号`,点击新建。

表单内容随便填写即可。
## 2. 扫码登录
确认创建渠道后,会多出一条记录,首次需要扫码登录,点击扫码登录,即可跳出登录二维码。


## 3. 愉快玩耍
扫码连接后,好友列表就会多出一个`微信 ClawBot`的机器人可使用。

点击进入后,即可开始聊天啦!

## FAQ
### 微信没找到入口
目前仅支持 IOS 系统,并且需要升级最新版本微信。
### 如何重置聊天
输入框输入:`Reset`或者`/reset` 即可重置聊天记录。
file: ./content/guide/build/publish/wecom.en.mdx
meta: {
"title": "WeCom Bot Integration",
"description": "FastGPT WeCom Bot Integration Tutorial"
}
* Starting from version 4.12.4, FastGPT commercial edition supports direct WeCom bot integration without additional APIs.
* Starting from version 4.14.4, FastGPT cloud service edition supports WeCom intelligent bot integration through custom domain configuration.
## 1. (Required for Cloud Service Edition) Configure Custom Domain
WeCom requires intelligent bot message push addresses to use the enterprise's primary domain, so cloud service edition users must configure a custom domain before using WeCom bots.
* [Configure Custom Domain](../../workspace/customDomain.en.mdx)
If you are a commercial edition user, continue using your enterprise domain.
## 2. Create an Intelligent Bot
### 2.1 Super Admin Login
[Click to open WeCom Admin Console](https://work.weixin.qq.com/)
### 2.2 Find the Intelligent Bot Entry
On the "Security & Management" - "Management Tools" page, click "Intelligent Bot" (Note: Only the enterprise creator or super admin has permission to see this entry)

### 2.3 Select "API Mode Creation" for the Intelligent Bot
On the create bot page, scroll down and click "API Mode Creation"

### 2.4 Get Key Credentials
Randomly generate or manually enter Token and Encoding-AESKey, and record them

### 2.5 Create WeCom Bot Publishing Channel
In FastGPT, select the Agent you want to use. On the Publishing Channels page, select "WeCom Bot" and click "Create"

### 2.6 Configure Publishing Channel Information
Configure the publishing channel information. You need to enter the Token and AESKey recorded in step 2.4 (Token and Encoding-AESKey)

### 2.7 Copy Callback URL
After clicking "Confirm", select your configured custom domain, copy the callback URL, and paste it back into the WeCom intelligent bot configuration page.

## 3. Use the Intelligent Bot
In the WeCom platform's "Contacts", you can find the created bot and start sending messages

## FAQ
### Sent a message but no response
1. Check if the trusted domain is configured correctly.
2. Check if Token and Encoding-AESKey are correct.
3. Check FastGPT chat logs to see if there is a corresponding question record.
4. If there is no record, the app may have encountered an error. Try the simplest bot first.
file: ./content/guide/build/publish/wecom.mdx
meta: {
"title": "接入企微机器人教程",
"description": "FastGPT 接入企微机器人教程"
}
* 从 4.12.4 版本起,FastGPT 商业版支持直接接入企微机器人,无需额外的 API。
* 从 4.14.4 版本起,FastGPT 云服务版支持通过配置自定义域名的方式接入企微智能机器人。
## 1. (云服务版必须)配置自定义域名
企微要求智能机器人消息推送地址必须使用企业主体域名,因此云服务版本用户必须先配置自定义域名才能使用企微机器人。
* [配置自定义域名](../../workspace/customDomain.mdx)
若您是商业版用户,请继续使用您企业的域名。
## 2. 创建智能机器人
### 2.1 超级管理员登录
[点击打开企业微信管理后台](https://work.weixin.qq.com/)
### 2.2 找到智能机器人入口
在"安全与管理" - "管理工具"页面点击"智能机器人" ( 注意: 只有企业创建者或超级管理员才有权限看到这个入口 )

### 2.3 选择 “API模式创建” 智能机器人
在创建机器人页面, 下拉, 点击 "API模式创建"

### 2.4 获取关键密钥
随机生成或者手动输入 Token 和 Encoding-AESKey,并且记录下来

### 2.5 创建企微机器人发布渠道
在 FastGPT 中,选择要使用 Agent,在发布渠道页面,选择“企业微信机器人”,点击“创建”

### 2.6 配置发布渠道信息
配置该发布渠道的信息,需要填入 Token 和 AESKey,也就是第四步中记录下来的 Token 和 Encoding-AESKey

### 2.7 复制回调地址
点击“确认”后,选择您配置的自定义域名,复制回调地址,填回企微智能机器人配置页中。

## 3. 使用智能机器人
在企业微信平台的"通讯录",即可找到创建的机器人,就可以发送消息了

## FAQ
### 发送了消息,没响应
1. 检查可信域名是否配置正确。
2. 检查 Token 和 Encoding-AESKey 是否正确。
3. 查看 FastGPT 对话日志,是否有对应的提问记录。
4. 如果没记录,则可能是应用运行报错了,可以先试试最简单的机器人。
file: ./content/guide/build/skill/development.en.mdx
meta: {
"title": "Development & Debugging",
"description": "This guide details how to create or import a skill, and manage files, use the interactive terminal, and perform debug chat in Web IDE."
}
import { Alert } from '@/components/docs/Alert';
## 1. Creating & Importing Skills
Before writing code, you need to create a development project in the skills list. The platform supports two ways to create or load a skill:

### 1.1 Click the Create Area to Create a Skill
Click the "Create" card (the dashed box area with a plus icon) on the page. In the popup, set the skill name, icon, description, and **requirements**. When the system initializes the skill in the background, it takes different approaches based on your input:
* **Using Default Template**: If you leave the default "Goal/Process/Requirements" template unchanged, the system will use the built-in basic structure and boilerplate code to generate the skill workspace (without invoking AI models, consuming no points).
* **AI-Assisted Generation**: If you input custom functional requirements here (e.g., "help me write a skill that extracts all email addresses from a text"), the system will invoke the **configured default system LLM model** in the background to automatically generate the `SKILL.md` scheme and initialize the code, which will consume points.
### 1.2 Import an Existing Skill ZIP Archive
If you have a skill backed up or shared by others, click the "Import Skill" button at the top right of the page and upload the corresponding ZIP archive. The system will automatically unzip it and restore all code files and configurations in the background, allowing you to resume development immediately.
***
## 2. Workspace File Management
When you open the skill details page, the system initializes and provisions an isolated run workspace in a secure sandbox container, loading your project via the file tree on the right.

### 2.1 Multi-file Management
You can right-click or use action buttons on the file tree on the right to easily create, delete, rename, and move files or folders to organize your project structure.
### 2.2 Online Code Editing
Clicking any file in the tree opens it in the center editor:
* **Auto-Save**: The editor automatically saves and syncs your edits to the backend sandbox container as you type.
* **Real-Time Workspace Sync**: The file tree automatically monitors and syncs file changes. Whether you edit files, install package dependencies in the terminal, or background processes generate new files, the tree stays updated.
* **Change Isolation**: Any code changes made here only take effect instantly in the debugging environment, and will not directly affect live applications. To apply the latest code to production, you must click the "Publish" button to generate an official version. For details, please see [Versions & Publishing](/en/guide/build/skill/version).
### 2.3 Two-way Interactive Terminal
The command line terminal at the bottom right connects directly to the backend sandbox:
* **Running Commands**: You can enter various command line operations, such as installing required code dependencies online or executing various custom running and debugging scripts.
* **Log Feedback**: The terminal streams command execution logs in real time. If code or script execution errors occur, you can view the error messages directly in the terminal output to assist with debugging.
***
## 3. Agent Debug Chat
The agent debug panel on the left provides a testing environment, allowing you to test and call your custom skill logic in real time by chatting with the agent:

### 3.1 Immediate Effect
Every time you modify and save your code in the right editor, you don't need to manually compile, build, or redeploy. Simply send a new message in the chat box on the left, and the system will run the latest code in the background, allowing you to see the changes instantly.
### 3.2 Conversational Workspace Modification
You can directly chat with the agent to have it help you edit the file contents on the right (including creating, deleting, and modifying files). The file tree and editor on the right will reflect these changes in real time.
### 3.3 Real-time File Export
You can click "Export Config" in the top-right menu to package all code and configuration files in the current workspace into a ZIP archive and download it locally for backup or sharing.
file: ./content/guide/build/skill/development.mdx
meta: {
"title": "开发与调试",
"description": "详细介绍如何新建或导入技能,并在 Web IDE 中管理文件、使用交互终端以及进行对话调试。"
}
import { Alert } from '@/components/docs/Alert';
## 1. 新建与导入技能
在开始编写代码前,你需要先在技能列表中创建一个开发项目。系统支持以下两种方式来创建或载入技能:

### 1.1 点击新建区域创建技能
在页面中点击带有加号的“新建”卡片,在弹出的窗口中设置技能名称、图标、介绍以及**需求描述**。系统在后台初始化该技能时,会根据你的输入采取不同的方式:
* **使用默认模板**:如果你保留默认的“目标/流程/要求”模板未作修改,系统将直接使用内置的基础结构和样例代码生成技能工作区(不调用 AI 模型,不产生积分消耗)。
* **AI 辅助生成**:如果你在此输入了具体的功能需求(例如“帮我编写一个从文本中提取所有邮箱地址的技能”),系统在后台会调用**系统配置的默认大语言模型**,根据你的描述自动生成技能方案 `SKILL.md` 并完成代码初始化,这会产生相应的积分消耗。
### 1.2 导入已有技能 ZIP 压缩包
如果你手头有自己备份或他人分享的技能,可以点击页面右上角的“导入技能”按钮,上传对应的 ZIP 格式技能压缩包,系统会自动解压并在后台还原所有代码文件与配置,让你能够立即在此基础上继续开发。
***
## 2. 工作区文件管理
当你在后台打开技能详情页时,系统会在安全的沙盒容器中为你初始化并拉起一个独立的运行空间,同时在右侧通过文件树加载你的项目。

### 2.1 多文件管理
你可以在右侧的文件树上右键或点击按钮,轻松进行文件的新建、删除、重命名和移动,灵活组织项目的目录结构。
### 2.2 代码在线编辑
在文件树中点击任意文件即可在中央编辑器中打开并编辑代码:
* **自动保存**:编辑器自带自动保存机制,代码修改完成后,系统会自动同步并写入后台沙盒中。
* **实时状态刷新**:无论是你编辑保存、在终端安装依赖包,还是后台进程生成了新文件,右侧的文件树都会实时感知并刷新,自动同步展示最新状态。
* **变更隔离**:此处进行的所有代码修改仅在调试区即时生效,不会直接影响到线上已发布运行的应用。若需应用最新的代码,需要先点击“发布”生成正式版本,具体发布逻辑请详见 [版本与发布](/guide/build/skill/version)。
### 2.3 终端双向交互
右侧底部的命令行窗口(Terminal)直接连接到后台沙盒:
* **运行命令**:你可以在这里输入各种命令行操作,例如在线安装代码所需的依赖包,或是运行各类自定义测试与调试脚本等。
* **日志反馈**:终端会实时输出命令执行的过程和日志。如果代码或脚本运行出错,可以直接通过终端输出的报错信息来辅助定位问题。
***
## 3. 智能体对话调试
左侧的智能体调试面板提供了一个测试环境,允许你通过与智能体对话来实时测试和调用你编写的技能逻辑:

### 3.1 即改即生效
每次你在右侧编辑器中修改并保存代码后,无需手动进行编译、打包或重新部署。只需在左侧调试框中发送下一条消息,系统便会在后台自动运行你最新的代码,让你能够立即看到修改后的效果。
### 3.2 对话式工作区修改
你可以直接通过与智能体对话,让它帮你编辑右侧的文件内容(包括文件的创建、删除、修改等),右侧的文件树和编辑器会实时同步展示这些变化。
### 3.3 实时文件导出
你可以点击页面右上角菜单中的“导出配置”,将当前工作区中实时的所有代码与配置文件打包成 ZIP 压缩包下载到本地,方便进行本地备份或分享。
file: ./content/guide/build/skill/initialization.en.mdx
meta: {
"title": "Initialization Script",
"description": "Learn how to configure and execute initialization scripts in skill packages to prepare the skill running environment."
}
The skill initialization script is an optional pre-execution script provided by the skill developer. After the system successfully deploys and extracts your skill in an application, it automatically runs this script in an isolated virtual machine before executing the actual AI tasks.
Through the initialization script, you can automatically install third-party dependencies or perform necessary configurations before the skill code runs.
***
## 1. Skill Script Configuration and Execution Timing
To add an initialization script to your skill, simply place a Shell script named `entrypoint.sh` in the root directory of your skill package.

### Execution Timing
1. **Deployment and Extraction**: When a user runs an application referencing the skill, the system first deploys and extracts the skill package to the virtual machine at `./projects//`.
2. **Script Execution**: The system runs the `entrypoint.sh` script located in the root of the skill directory.
***
## 2. Smart Deduplication Mechanism
To prevent running environment initialization scripts repeatedly in subsequent conversations (e.g., executing dependency installation packages on every turn would cause severe latency), FastGPT designs a deduplication mechanism for skill scripts.
The execution state is stored in the `~/.fastgpt/agent-skill-entrypoints/state.json` file inside the virtual machine.
* **Version ID Deduplication**: Since the code and script content of a specific skill version (`versionId`) are immutable once published, the system tracks the successfully executed `versionId` inside the virtual machine.
* **Skipped Execution**: When the same virtual machine instance is reused in subsequent turns, if the corresponding `versionId` has already executed successfully, the system will **skip** running the script, enabling hot starts.
* **New Version Trigger**: Whenever a new skill version is published, the system will automatically deploy and run its initialization script upon the next conversation, regardless of whether it is an existing (old) or a new chat window.
***
## 3. Execution Constraints and Fault Tolerance
To ensure the smooth execution of the AI workflow, the skill initialization script must adhere to the same execution constraints and fault tolerance rules as the application startup script:
* **Execution Constraints and Non-blocking Fault Tolerance**: The timeout protection (default 30 seconds), non-blocking workflow (failures do not block main execution), and 8KB log truncation limits are identical to those of the application startup script. For detailed parameters, please refer to [Application Startup Script Execution Constraints](../agentv2/vm#execution-constraints--fault-tolerance).
* **Debug Mode Limitation**: In the skill edit mode, the virtual machine will not automatically execute the `entrypoint.sh` script. To verify the script's behavior, the skill developer can manually execute the commands inside the Workspace Terminal.
file: ./content/guide/build/skill/initialization.mdx
meta: {
"title": "初始化脚本",
"description": "了解如何在技能包中配置和执行初始化脚本,准备技能运行环境。"
}
技能初始化脚本是技能开发者提供的一个前置脚本。当应用成功部署并解压了您开发的技能后,系统在实际执行 AI 任务前,会在独立的虚拟机环境中自动运行该脚本。
通过初始化脚本,您可以在技能代码执行前,自动安装技能特有的第三方依赖,或进行必要的配置预处理。
***
## 1. 技能脚本配置与执行时机
要为您的技能添加初始化脚本,只需在技能压缩包的根目录下放置一个名为 `entrypoint.sh` 的 Shell 脚本。

### 执行时机
1. **技能包部署与解压**:当用户运行引用了该技能的应用时,系统会首先将技能包部署并解压到虚拟机的 `./projects//` 目录下。
2. **执行初始化脚本**:系统会在虚拟机中执行该技能根目录下的 `entrypoint.sh` 脚本。
***
## 2. 智能去重机制
为了避免在多次对话中重复运行环境初始化脚本(例如重复执行依赖包安装会导致每次对话产生严重的延迟),系统为技能脚本设计了去重机制。
去重状态记录在虚拟机内的 `~/.fastgpt/agent-skill-entrypoints/state.json` 状态文件中。
* **版本 ID 去重**:由于同一个技能版本(`versionId`)的代码和脚本内容在发布后是不可变的,系统会记录当前虚拟机中已成功运行过的技能 `versionId`。
* **跳过执行**:当同一个虚拟机实例在后续对话中被复用时,只要对应的 `versionId` 已经成功执行过,系统就会**直接跳过**该脚本的运行,实现秒级热启动。
* **新版本触发**:只要技能发布了新版本,无论是在旧的对话窗口还是新的对话窗口,在下一次对话触发时,系统都将在重新部署该技能后自动运行该版本的初始化脚本。
***
## 3. 执行约束与容错机制
为保障 AI 流程的流畅运行,技能初始化脚本需要遵循与应用启动脚本一致的执行限制与容错规则:
* **执行约束与非阻断容错**:技能入口脚本的超时时间限制(默认 30 秒)、非阻塞设计(执行报错或超时不阻断主流程)以及 8KB 日志输出截断规则,均与应用启动脚本保持一致。具体细节指标请参考 [应用启动脚本的执行限制](../agentv2/vm#执行限制与容错机制)。
* **调试预览限制**:在技能的编辑模式下,虚拟机不会自动执行技能的 `entrypoint.sh` 脚本。如果需要验证脚本效果,技能开发者可以直接在侧边栏调试区的控制台终端(Workspace Terminal)中手动执行相关命令。
file: ./content/guide/build/skill/integration.en.mdx
meta: {
"title": "Agent Integration",
"description": "How to bind published skills to AI agents and execute them."
}
import { Alert } from '@/components/docs/Alert';
## How to Bind a Skill to an Agent?
1. Go to the editing page of the **"Agent"** application where you want to integrate this skill (currently, only Agent applications support direct skill binding; simple apps and workflows do not support it).
2. In the configuration panel on the left, locate the **"Associated Skill"** section.
3. Click the **"Select"** button on the right, and in the popup list, select your published skill.

**Note:** Skill code needs to execute within a secure and isolated environment. Therefore, when
you associate a skill, the system will automatically enable the "Virtual Machine" for you; you
cannot disable the virtual machine while a skill remains associated.
***
## How Agents Call Skills
Once bound, the agent possesses this skill capability:
* **Multi-Skill Injection**: An agent can be bound to **multiple different skills** at the same time. When the sandbox (virtual machine) starts, all bound skill codes and configurations are automatically injected and deployed into the sandbox workspace. The skills are isolated from each other and will not conflict.
* **Automated Invocation**: **Provided that the LLM used is sufficiently intelligent**, you don't need to manually command the agent to run code. The AI will automatically judge whether to trigger the skill based on your input, and execute it securely in the background sandbox.
* **Virtual Machine File View**: You can click the **"Virtual Machine"** button at the bottom of the chat bubble (or the computer icon in the top right corner) to view all the latest files and code status in the virtual machine directly in the popup sidebar.

file: ./content/guide/build/skill/integration.mdx
meta: {
"title": "智能体集成",
"description": "如何将发布好的技能绑定到 AI 智能体中并执行。"
}
import { Alert } from '@/components/docs/Alert';
## 如何在应用中绑定技能?
1. 进入你想集成该技能的 **“智能体 (Agent)”** 应用编辑页面(目前仅智能体应用支持直接绑定技能,简易应用及工作流暂不支持)。
2. 在左侧的配置面板中,找到 **“关联 Skill”** 配置项。
3. 点击右侧的 **“选择”** 按钮,在弹出的选择窗口中,勾选你已经发布好的正式版本技能。

**注意:**
技能代码需要在安全隔离的环境中运行。因此,当你关联技能时,系统会自动为你开启“虚拟机”;并且在已关联技能的状态下,无法关闭虚拟机。
***
## 智能体如何调用技能?
绑定完成后,智能体即可获得该技能的执行能力:
* **多技能安全注入**:一个智能体支持同时绑定**多个不同的技能**。在沙盒(虚拟机)启动时,所有已绑定技能的代码和配置都会被自动注入并部署到沙盒工作区中,各技能间彼此独立、互不冲突。
* **智能自动调用**:**在使用的模型足够智能的前提下**,你无需手动命令智能体运行代码。AI 会根据你的输入,自动判断是否需要调用该技能,并在后台虚拟机中自动安全地执行代码。
* **虚拟机文件查看**:你可以点击聊天气泡底部的 **“虚拟机”** 按钮(或右上角的电脑图标),在弹出的侧边栏中直接查看虚拟机里当前最新的所有文件内容和代码状态。

file: ./content/guide/build/skill/intro.en.mdx
meta: {
"title": "Introduction",
"description": "The concept of AI Agent Skills, and how it is designed and implemented in FastGPT."
}
import { Alert } from '@/components/docs/Alert';
## What is an AI Agent "Skill"?
Under the latest ecosystem designs of mainstream AI providers, a **"Skill"** is defined as a **persistent, reusable, and modular workflow and capability package**.
For example, if you frequently need the AI to audit complex spreadsheets and write analysis
reports, you can package the 'audit code' and 'report template' into a Skill. In future chats, you
can simply upload your spreadsheet, and the AI will run the skill in the background to compute
results and format the report.
***
## Core Design Philosophy: From Tools to Skills
In the general cognitive framework of AI Agents, we typically divide capabilities into three layers:
* **The Brain (Brain)**: Responsible for reasoning and planning (the LLM itself).
* **Tools (Tools)**: Simple execution interfaces (such as sending a web request or running a temporary line of code), resembling the AI's "hands and feet".
* **Skills (Skills)**: Providing the complete **"operational knowledge and professional logic"** (Know-how).
A skill is typically a modular package encapsulating **instruction markdown (how to do it)** and **executable scripts (actually doing it)**.
If a tool is a "screwdriver" in your toolbox, then a skill is a **"furniture assembly guide"**. The AI can automatically grab this guide from its skill library based on the current context, execute the code inside a background sandbox, and complete the complex assembly.
***
## Skills in FastGPT
Following the industry-standard design of Skills, FastGPT provides you with a "**dedicated code execution workspace**" featuring the following core designs:

### 1. Isolated Secure Runtime Sandbox
Each created skill during editing runs in a fully isolated, secure sandbox environment (powered by Sealos Devbox, OpenSandbox, etc., in the backend). All operations are restricted within this workspace to ensure safety.
### 2. Instant Hot-Reloading Debugging
Provides an online debugging environment integrating a file tree, code editor, and console terminal. Equipped with an agent debug panel on the left supporting hot reloading, allowing you to troubleshoot the skill before publishing.
### 3. Isolation of Production & Debugging
Edits in the workspace will only take effect instantly in the "Debug Chat" area. Changes will only be applied to production agents once you click publish and snap a new version, ensuring service stability.
### 4. Auto-Sleep & Seamless Invocation
For long-inactive skills, the system automatically shuts down the sandbox and performs cold-archiving to storage. When edit or agent invocation resumes, the sandbox is automatically re-instantiated and restored from the archive in the background. You are only billed when the skill is active, dramatically reducing your runtime costs.
file: ./content/guide/build/skill/intro.mdx
meta: {
"title": "基础介绍",
"description": "AI 智能体技能概念,以及它在 FastGPT 中的设计与实现原理。"
}
import { Alert } from '@/components/docs/Alert';
## 什么是 AI 智能体的“技能”?
在当前主流 AI 厂商的最新生态设计中,**“技能”(Skills)** 被定义为一种**可持久保存、可复用的模块化专业流程与能力包**。
例如,如果你经常需要 AI
帮你核对两份复杂的财务表格并生成分析,你只需一次性把“计算代码”和“报告模板”放入技能中。在以后的对话中,你直接把表格丢给
AI,它就能自动在后台调用这个技能把数据算准、格式排好。
***
## 核心设计原理:从工具到技能
在 AI 智能体(Agent)的大众认知中,我们通常将它划分为三个层面:
* **大脑 (Brain)**:负责规划和推理,是大模型本身。
* **工具 (Tools)**:提供单纯的“动作接口”(例如:发送一段网络请求、运行一行临时代码),类似于 AI 的“手和脚”。
* **技能 (Skills)**:提供完整的“**做事章法与专业逻辑**”(Know-how)。
一个技能通常是由**说明文档(指明怎么做)**和**逻辑代码(真正去执行)**封装在一起的模块化包。如果工具是工具箱里的“螺丝刀”,那么技能就是一张**“家具组装手册”**,AI 能够自动根据当前对话任务,伸手从它的技能库里拿取这本手册,在后台沙盒中运行代码并完成复杂的装配任务。
***
## FastGPT 中的技能设计
承袭业界主流的技能(Skills)设计标准,FastGPT 支持你为智能体创建“**专属的独立代码空间**”,具备以下核心设计:

### 1. 独立的安全运行沙箱
创建出来的每个技能在编辑时都拥有一个完全隔离的安全运行沙箱(后台基于 Sealos Devbox、OpenSandbox 等沙盒服务运行)。所有操作都在此隔离空间内进行,保障技能执行的安全性。
### 2. 即改即生效的调试环境
提供了一个集成了文件管理、代码编辑器和交互式终端的在线调试环境。左侧配有智能体调试面板,支持“即改即生效”的热重载,方便你在发布前对技能进行充分的调试与排错。
### 3. 生产与调试环境隔离
在编辑区域直接修改的代码只在“调试区”即时生效。只有点击“发布”生成并保存正式版本后,改动才会正式应用到生产环境的智能体与工作流中。
### 4. 自动休眠与无感唤醒
针对长期闲置的技能,系统会自动将其从沙盒中清理并冷归档至存储。当需要再次编辑或被智能体调用时,会自动在后台重新拉起沙箱并复原。休眠期间不产生任何运行计费,大幅降低使用成本。
file: ./content/guide/build/skill/version.en.mdx
meta: {
"title": "Versions & Publishing",
"description": "Why you need to publish versions, how to save snapshots, and easy rollback to historical versions."
}
import { Alert } from '@/components/docs/Alert';
## Why do you need to "Publish"?
Edits in the editor only take effect in the "Debug Chat" panel. To formally apply your changes to your workflows or agents, you must click "Publish" to deploy a formal version.
**Note:** The debugging environment is isolated from the production deployment environment. This
ensures that when you edit or debug skill code, it will not affect the online agents and workflows
currently running.
***
## Saving Version Snapshots
1. Once testing is successful, click the **"Publish"** button in the top right corner of the editor.
2. In the modal, enter the **Version Name** (by default, the current time is prefilled, but you can customize it, e.g., `v1.0.0`).
3. Confirm, and the system will solidify the current code state as an "official version" and publish it.
**Note:** During publishing, the system automatically applies the ignore rules specified in the
`.gitignore` file at the project root (if not present, a default file ignoring `node_modules`,
`.venv`, `dist`, etc., will be created). Only files that are not ignored will be packaged, and you
must ensure the total size of these files does not exceed the limit, otherwise publishing may
fail.
***
## Version Rollback
To restore a previous version:
1. Click the **"Version History"** (clock) icon in the top right corner of the editor to view all published snapshots.
2. Hover over the version you want to restore, and click the **"Switch"** (return arrow) icon to instantly revert both your workspace files and the live production version back to that snapshot.
**Note:** Restored versions will not carry files that were ignored by `.gitignore` (for example,
local files like `node_modules` or `.venv` that were ignored cannot be recovered via rollback).

file: ./content/guide/build/skill/version.mdx
meta: {
"title": "版本与发布",
"description": "为什么需要发布版本,如何保存快照,以及历史版本的轻松回滚。"
}
import { Alert } from '@/components/docs/Alert';
## 为什么要“发布”?
在编辑器里直接修改的代码只在“调试区”即时生效。如果你想在工作流或者智能体当中正式应用你的改动,必须点击“发布”生成一个正式部署版本。
**注意:**
调试环境与生产部署环境是隔离的。这样可以确保你在调试、改写技能代码时,不会影响线上正在运行的智能体与工作流服务。
***
## 保存版本快照
1. 调试确认无误后,点击编辑器右上角的 **“发布”** 按钮。
2. 在弹出的窗口中,输入当前版本的**版本名称**(默认会自动填充当前时间作为名称,你也可以自定义修改,例如输入 `v1.0.0`)。
3. 确认后,系统会将当前的代码状态固化为一个“正式版本”发布上线。
**注意:** 发布时,系统会自动应用项目根目录下 `.gitignore`
文件的忽略规则(如不存在,系统会自动创建包含 `node_modules`、`.venv`、`dist`
等默认忽略项的文件)。只有未被忽略的文件才会被打包发布,请确保打包文件总体积未超限,否则可能导致发布失败。
***
## 历史版本回滚
如果需要恢复到以前的版本:
1. 点击编辑器右上角的 **“版本历史”**(时钟)图标,查看已发布的所有历史快照。
2. 将鼠标悬停在要恢复的历史版本上,点击 **“切换”**(返回箭头)图标,即可一键将工作区文件以及当前线上运行的版本同时切换回该历史版本。
**注意:** 回滚的版本不会携带被 `.gitignore` 忽略的文件(如依赖包 `node_modules`、虚拟环境 `.venv`
等已忽略的本地文件不会被恢复)。

file: ./content/guide/build/tools/mcp_tools.en.mdx
meta: {
"title": "MCP Tools",
"description": "A quick guide to integrating MCP tools with FastGPT"
}
Starting from FastGPT v4.9.6, a new application type called MCP Tools has been added. It lets you provide an MCP SSE URL to batch-create tools that models can easily call. Here's how to create MCP tools and have AI use them.
## Create an MCP Tools Collection
First, select "New MCP Tools Collection." We'll use the Amap (Gaode Maps) MCP Server as an example: [Amap MCP Server](https://lbs.amap.com/api/mcp-server/create-project-and-key)
You'll need an MCP URL, e.g., [https://mcp.amap.com/sse?key=xxx](https://mcp.amap.com/sse?key=xxx)

Enter the URL in the dialog and click Parse. The system will discover and list the available tools.
Click Create to finish setting up the MCP tools and collection.
## Test MCP Tools
Inside the MCP Tools collection, you can debug each tool individually.

For example, select the maps\_weather tool and click Run to see the weather data for Hangzhou.
## AI Calling Tools
### Call Individual Tools

Using maps\_weather and maps\_text\_search as examples, ask the AI two different questions. The AI intelligently selects the appropriate tool, retrieves the needed information, and responds based on the results.
| | |
| ------------------------- | ------------------------- |
|  |  |
### Call an Entire Tools Collection
FastGPT also supports calling an entire MCP Tools collection. The AI automatically picks the right tool to execute.
Click the MCP Tools collection to add a collection-type node, then connect it using the Tool Calling node.
| | |
| ------------------------- | ------------------------- |
|  |  |
The AI similarly selects the appropriate tool, retrieves the needed information, and responds based on the results.
file: ./content/guide/build/tools/mcp_tools.mdx
meta: {
"title": "MCP 工具集",
"description": "快速了解 MCP 工具接入 FastGPT"
}
FastGPT v4.9.6 版本开始,新增了 MCP 工具集 这种新的应用类型,允许传入一个 MCP 的 SSE URL 来批量创建可被模型轻松调用的 MCP 工具,下面就来看下如何创建 MCP 工具并且让 AI 调用
## 创建一个 MCP 工具集
首先选择新建 MCP 工具集,以对接高德地图的 MCP Server 为例,[高德地图 MCP Server](https://lbs.amap.com/api/mcp-server/create-project-and-key)
需要获取到一个 MCP 地址,例 [https://mcp.amap.com/sse?key=xxx](https://mcp.amap.com/sse?key=xxx)

然后填入到弹窗中的对应位置,点击后面的解析,会解析出对应的一系列工具
这时再点击创建就能轻松创建 MCP 工具和 MCP 工具集
## 测试 MCP 工具
进入到 MCP 工具集内部,能够对每个单独的 MCP 工具进行调试

以 maps\_weather 这个查询天气的工具为例,点击运行,可以看到能够获得杭州的具体天气
## 模型调用工具
### 调用单个工具

选中 maps\_weather 和 maps\_text\_search 这两个工具为例,分别问 AI 两个问题,可以看到 AI 智能地调用了相应的工具获得了需要的信息,然后根据获得的信息回答
| | |
| ------------------------- | ------------------------- |
|  |  |
### 调用工具集
FastGPT 也支持调用整个 MCP 工具集,AI 会自动选取需要的工具执行,
点击 MCP 工具集,会添加一个工具集类型的节点,使用工具调用节点连接
| | |
| ------------------------- | ------------------------- |
|  |  |
可以看到 AI 同样智能调用了相应的工具,获得了需要的信息,然后根据获得的信息回答
file: ./content/guide/build/workflow/intro.en.mdx
meta: {
"title": "Workflows & Plugins",
"description": "A quick overview of FastGPT Workflows and Plugins"
}
Starting from V4.0, FastGPT adopted a new approach to building AI applications. It uses Flow node orchestration (Workflows) to implement complex processes, improving flexibility and extensibility. This does raise the learning curve — users with development experience will find it easier to pick up.
[Watch the video tutorial](https://www.bilibili.com/video/BV1is421u7bQ/)

## What is a Node?
In programming terms, a node is like a function or API endpoint — think of it as a **step**. By connecting multiple nodes together, you build a step-by-step process that produces the final AI output.
Below is the simplest AI conversation, consisting of a Workflow Start node and an AI Chat node.

Execution flow:
1. The user inputs a question. The \[Workflow Start] node executes and saves the user's question.
2. The \[AI Chat] node executes. It has two required parameters: "Chat History" and "User Question." Chat history defaults to 6 messages, representing the context length. The user question comes from the \[Workflow Start] node.
3. The \[AI Chat] node calls the conversation API with the chat history and user question to generate a response.
### Node Categories
Functionally, nodes fall into 2 categories:
1. **System Nodes**: User guidance (configures dialog information) and user question (workflow entry point).
2. **Function Nodes**: Knowledge Base search, AI Chat, and all other nodes. These have inputs and outputs and can be freely combined.
### Node Components
Each node has 3 core parts: inputs, outputs, and triggers.

* AI model, prompt, chat history, user question, and Knowledge Base citation are inputs. Inputs can be manual entries or variable references, which include "global variables" and outputs from any previous node.
* New context and AI reply content are outputs. Outputs can be referenced by any subsequent node.
* Each node has four "triggers" (top, bottom, left, right) for connections. Connected nodes execute sequentially based on conditions.
## Key Concept — How Workflows Execute
FastGPT Workflows start from the \[Workflow Start] node, triggered when the user inputs a question. There is no **fixed exit point** — the workflow ends when all nodes stop running. If no nodes execute in a given cycle, the workflow completes.
Let's look at how workflows execute and when each node is triggered.

As shown above, nodes can "be connected to" and "connect to other nodes." We call incoming connections "predecessor lines" and outgoing connections "successor lines." In the example, the \[Knowledge Base Search] node has one predecessor line on the left and one successor line on the right. The \[AI Chat] node only has a predecessor line on the left.
Lines in FastGPT Workflows have these states:
* `waiting`: The connected node is waiting to execute.
* `active`: The connected node is ready to execute.
* `skip`: The connected node should be skipped.
Node execution rules:
1. If any predecessor line has `waiting` status, the node waits.
2. If any predecessor line has `active` status, the node executes.
3. If no predecessor lines are `waiting` or `active`, the node is skipped.
4. After execution, successor lines are updated to `active` or `skip`, and predecessor lines reset to `waiting` for the next cycle.
Walking through the example:
1. \[Workflow Start] completes and sets its successor line to `active`.
2. \[Knowledge Base Search] sees its predecessor line is `active`, executes, then sets its successor line to `active` and predecessor line to `waiting`.
3. \[AI Chat] sees its predecessor line is `active` and executes. The workflow ends.
## How to Connect Nodes
1. Each node has connection points on all four sides for convenience. Left and top are predecessor connection points; right and bottom are successor connection points.
2. Click the x in the middle of a connection line to delete it.
3. Left-click to select a connection line.
## How to Read Workflows
1. Read from left to right.
2. Start from the **User Question** node, which represents the user sending text to trigger the workflow.
3. Focus on \[AI Chat] and \[Specified Reply] nodes — these are where answers are output.
## FAQ
### How do I merge multiple outputs?
1. Text Processing: can merge strings together.
2. Knowledge Base Search Merge: can combine multiple Knowledge Base search results.
3. Other results: cannot be merged directly. Consider passing them to an `HTTP` node and merging them in your own service.
file: ./content/guide/build/workflow/intro.mdx
meta: {
"title": "工作流&插件",
"description": "快速了解 FastGPT 工作流和插件的使用"
}
FastGPT 从 V4.0 版本开始采用新的交互方式来构建 AI 应用。使用了 Flow 节点编排(工作流)的方式来实现复杂工作流,提高可玩性和扩展性。但同时也提高了上手的门槛,有一定开发背景的用户使用起来会比较容易。
[查看视频教程](https://www.bilibili.com/video/BV1is421u7bQ/)

## 什么是节点?
在程序中,节点可以理解为一个个 Function 或者接口。可以理解为它就是一个**步骤**。将多个节点一个个拼接起来,即可一步步的去实现最终的 AI 输出。
如下图,这是一个最简单的 AI 对话。它由用流程开始和 AI 对话节点组成。

执行流程如下:
1. 用户输入问题后,【流程开始】节点执行,用户问题被保存。
2. 【AI 对话】节点执行,此节点有两个必填参数“聊天记录”“用户问题”,聊天记录的值是默认输入的 6 条,表示此模块上下文长度。用户问题选择的是【流程开始】模块中保存的用户问题。
3. 【AI 对话】节点根据传入的聊天记录和用户问题,调用对话接口,从而实现回答。
### 节点分类
从功能上,节点可以分为 2 类:
1. **系统节点**:用户引导(配置一些对话框信息)、用户问题(流程入口)。
2. **功能节点**:知识库搜索、AI 对话等剩余节点。(这些节点都有输入和输出,可以自由组合)。
### 节点的组成
每个节点会包含 3 个核心部分:输入、输出和触发器。

* AI 模型、提示词、聊天记录、用户问题,知识库引用为输入,节点的输入可以是手动输入也可以是变量引用,变量引用的范围包括“全局变量”和之前任意一个节点的输出。
* 新的上下文和 AI 回复内容为输出,输出可以被之后任意节点变量引用。
* 节点的上下左右有四个“触发器”可以被用来连接,被连接的节点按顺序决定是否执行。
## 重点 - 工作流是如何运行的
FastGPT 的工作流从【流程开始】节点开始执行,可以理解为从用户输入问题开始,没有**固定的出口**,是以节点运行结束作为出口,如果在一个轮调用中,所有节点都不再运行,则工作流结束。
下面我们来看下,工作流是如何运行的,以及每个节点何时被触发执行。

如上图所示节点会“被连接”也会“连接其他节点”,我们称“被连接”的那根线为前置线,“连接其他节点的线”为后置线。上图例子中【知识库搜索】模块左侧有一根前置线,右侧有一根后置线。而【AI 对话】节点只有左侧一根前置线。
FastGPT 工作流中的线有以下几种状态:
* `waiting`:被连接的节点等待执行。
* `active`:被连接的节点可以执行。
* `skip`:被连接的节点不需要执行跳过。
节点执行的原则:
1. 判断前置线中有没有状态为 `waiting` 的,如果有则等待。
2. 判断前置线中状态有没有状态为 `active` 如果有则执行。
3. 如果前置线中状态即没有 `waiting` 也没有 `active` 则认为此节点需要跳过。
4. 节点执行完毕后,需要根据实际情况更改后置线的状态为 `active` 或 `skip` 并且更改前置线状态为 `waiting` 等待下一轮执行。
让我们看一下上面例子的执行过程:
1. 【流程开始】节点执行完毕,更改后置线为 `active`。
2. 【知识库搜索】节点判断前置线状态为 `active` 开始执行,执行完毕后更改后置线状态为 `active` 前置线状态为 `waiting`。
3. 【AI 对话】节点判断前置线状态为 `active` 开始执行,流程执行结束。
## 如何连接节点
1. 为了方便连接,FastGPT 每个节点的上下左右都有连接点,左和上是前置线连接点,右和下是后置线连接点。
2. 可以点击连接线中间的 x 来删除连接线。
3. 可以左键点击选中连接线
## 如何阅读?
1. 建议从左往右阅读。
2. 从 **用户问题** 节点开始。用户问题节点,代表的是用户发送了一段文本,触发任务开始。
3. 关注【AI 对话】和【指定回复】节点,这两个节点是输出答案的地方。
## FAQ
### 想合并多个输出结果怎么实现?
1. 文本加工,可以对字符串进行合并。
2. 知识库搜索合并,可以合并多个知识库搜索结果
3. 其他结果,无法直接合并,可以考虑传入到 `HTTP` 节点中,通过你的业务服务进行合并。
file: ./content/guide/dataset/third-party/api_dataset.en.mdx
meta: {
"title": "API File Library",
"description": "Introduction and usage of the FastGPT API File Library"
}
import { Alert } from '@/components/docs/Alert';
| | |
| ----------------------- | ----------------------- |
|  |  |
## Background
FastGPT supports local file imports, but in many cases users already have an existing document library. Re-importing files would create duplicate storage and complicate management. To address this, FastGPT offers an API File Library that connects to your existing document library through simple API endpoints, with flexible import options.
The API File Library lets you integrate your existing document library seamlessly. Implement a few endpoints that conform to FastGPT's API File Library specification, provide the service's baseURL and token when creating a knowledge base, and you can browse and selectively import files directly from the UI.
## How to Use the API File Library
When creating a knowledge base, select the API File Library type and configure the key parameters: the baseURL of your file service and the request header for authentication. As long as your endpoints conform to FastGPT's specification, the system will automatically fetch and display the complete file list for selective import.
You need to provide three parameters:
* baseURL: The base URL of your file service
* authorization: The authentication request header, sent as `Authorization: Bearer `
* basePath: Optional, the root directory path to specify the starting position of the file tree
## API Specification
Response format:
```ts
type ResponseType = {
success: boolean;
message: string;
data: any;
}
```
Data types:
```ts
// Single file item in the file list
type FileListItem = {
id: string;
parentId: string | null;
name: string;
type: 'file' | 'folder';
updateTime: Date;
createTime: Date;
hasChild?: boolean; // Optional, whether it has child nodes, defaults to true for folder type
}
```
### 1. Get File Tree
* parentId - Parent ID, optional. If not provided or null, the configured basePath will be used as the root directory
* searchKey - Search keyword, optional
```bash
curl --location --request POST '{{baseURL}}/v1/file/list' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"parentId": null,
"searchKey": ""
}'
```
```json
{
"success": true,
"message": "",
"data": [
{
"id": "xxxx",
"parentId": "xxxx",
"type": "file",
"name":"test.json",
"updateTime":"2024-11-26T03:05:24.759Z",
"createTime":"2024-11-26T03:05:24.759Z",
"hasChild": false
}
]
}
```
### 2. Get Single File Content (Text Content or Access Link)
```bash
curl --location --request GET '{{baseURL}}/v1/file/content?id=xx' \
--header 'Authorization: Bearer {{authorization}}'
```
```json
{
"success": true,
"message": "",
"data": {
"title": "Document Title",
"content": "FastGPT is an LLM-based knowledge base Q&A system with out-of-the-box data processing and model invocation capabilities. It also supports visual workflow orchestration via Flow for complex Q&A scenarios!\n"
}
}
```
* **title** - File title, optional. Used to display the file name. If not provided, the system will attempt to parse the filename from `previewUrl`.
* **content** - The text content of the file, optional. Returns the complete text content of the file directly, which the system will use for indexing and retrieval.
* **previewUrl** - The access link to the file, optional. Provides an accessible file URL, and the system will automatically request this address to download the file and extract its content. Supports various file formats (such as PDF, Word, Markdown, etc.).
**Important Notes:**
* Either `content` or `previewUrl` must be returned, **at least one is required**, otherwise an error will occur.
* If both `content` and `previewUrl` are returned, `content` takes priority and the system will use the `content` directly.
* When `previewUrl` is returned, the system will access the link to read and parse the document content, and will cache the parsing results to improve performance.
### 3. Get File Read Link (for Viewing the Original)
id is the file's ID.
```bash
curl --location --request GET '{{baseURL}}/v1/file/read?id=xx' \
--header 'Authorization: Bearer {{authorization}}'
```
```json
{
"success": true,
"message": "",
"data": {
"url": "xxxx"
}
}
```
* url - File access link; opens automatically once retrieved.
### 4. Get File Details
id is the file's ID.
```bash
curl --location --request GET '{{baseURL}}/v1/file/detail?id=xx' \
--header 'Authorization: Bearer {{authorization}}'
```
```json
{
"success": true,
"message": "",
"data": {
"id": "xxxx",
"name": "test.json",
"parentId": "xxxx",
"type": "file",
"updateTime": "2024-11-26T03:05:24.759Z",
"createTime": "2024-11-26T03:05:24.759Z"
}
}
```
* id - File ID
* name - File name
* parentId - Parent ID, null indicates root directory
* type - File type, file or folder
* updateTime - Update time
* createTime - Creation time
file: ./content/guide/dataset/third-party/api_dataset.mdx
meta: {
"title": "API 文件库",
"description": "FastGPT API 文件库功能介绍和使用方式"
}
import { Alert } from '@/components/docs/Alert';
| | |
| ----------------------- | ----------------------- |
|  |  |
## 背景
目前 FastGPT 支持本地文件导入,但是很多时候,用户自身已经有了一套文档库,如果把文件重复导入一遍,会造成二次存储,并且不方便管理。因为 FastGPT 提供了一个 API 文件库的概念,可以通过简单的 API 接口,去拉取已有的文档库,并且可以灵活配置是否导入。
API 文件库能够让用户轻松对接已有的文档库,只需要按照 FastGPT 的 API 文件库规范,提供相应文件接口,然后将服务接口的 baseURL 和 token 填入知识库创建参数中,就能直接在页面上拿到文件库的内容,并选择性导入
## 如何使用 API 文件库
创建知识库时,选择 API 文件库类型,然后需要配置两个关键参数:文件服务接口的 baseURL 和用于身份验证的请求头信息。只要提供的接口规范符合 FastGPT 的要求,系统就能自动获取并展示完整的文件列表,可以根据需要选择性地将文件导入到知识库中。
你需要提供三个参数:
* baseURL: 文件服务接口的 baseURL
* authorization: 用于身份验证的请求头信息,实际请求格式为 `Authorization: Bearer `
* basePath: 可选,根目录路径,用于指定文件树的起始位置
## 接口规范
接口响应格式:
```ts
type ResponseType = {
success: boolean;
message: string;
data: any;
}
```
数据类型:
```ts
// 文件列表中,单项的文件类型
type FileListItem = {
id: string;
parentId: string | null;
name: string;
type: 'file' | 'folder';
updateTime: Date;
createTime: Date;
hasChild?: boolean; // 可选,是否有子节点,默认 folder 类型为 true
}
```
### 1. 获取文件树
* parentId - 父级 id,可选。如果不传或传 null,则使用配置的 basePath 作为根目录
* searchKey - 检索词,可选
```bash
curl --location --request POST '{{baseURL}}/v1/file/list' \
--header 'Authorization: Bearer {{authorization}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"parentId": null,
"searchKey": ""
}'
```
```json
{
"success": true,
"message": "",
"data": [
{
"id": "xxxx",
"parentId": "xxxx",
"type": "file",
"name":"test.json",
"updateTime":"2024-11-26T03:05:24.759Z",
"createTime":"2024-11-26T03:05:24.759Z",
"hasChild": false
}
]
}
```
### 2. 获取单个文件内容(文本内容或访问链接)
```bash
curl --location --request GET '{{baseURL}}/v1/file/content?id=xx' \
--header 'Authorization: Bearer {{authorization}}'
```
```json
{
"success": true,
"message": "",
"data": {
"title": "文档标题",
"content": "FastGPT 是一个基于 LLM 大语言模型的知识库问答系统,提供开箱即用的数据处理、模型调用等能力。同时可以通过 Flow 可视化进行工作流编排,从而实现复杂的问答场景!\n"
}
}
```
* **title** - 文件标题,可选。用于显示文件名称,如果不提供,系统会尝试从 `previewUrl` 中解析文件名。
* **content** - 文件的文本内容,可选。直接返回文件的完整文本内容,系统会直接使用该内容进行索引和检索。
* **previewUrl** - 文件的访问链接,可选。提供一个可访问的文件 URL,系统会自动请求该地址下载文件并提取内容。支持各种文件格式(如 PDF、Word、Markdown 等)。
**重要说明:**
* `content` 和 `previewUrl` 二选一返回,**必须至少返回其中一个**,否则会报错。
* 如果同时返回 `content` 和 `previewUrl`,则 `content` 优先级更高,系统会直接使用 `content` 的内容。
* 返回 `previewUrl` 时,系统会访问该链接进行文档内容读取和解析,并会缓存解析结果以提高性能。
### 3. 获取文件阅读链接(用于查看原文)
id 为文件的 id。
```bash
curl --location --request GET '{{baseURL}}/v1/file/read?id=xx' \
--header 'Authorization: Bearer {{authorization}}'
```
```json
{
"success": true,
"message": "",
"data": {
"url": "xxxx"
}
}
```
* url - 文件访问链接,拿到后会自动打开。
### 4. 获取文件详情
id 为文件的 id。
```bash
curl --location --request GET '{{baseURL}}/v1/file/detail?id=xx' \
--header 'Authorization: Bearer {{authorization}}'
```
```json
{
"success": true,
"message": "",
"data": {
"id": "xxxx",
"name": "test.json",
"parentId": "xxxx",
"type": "file",
"updateTime": "2024-11-26T03:05:24.759Z",
"createTime": "2024-11-26T03:05:24.759Z"
}
}
```
* id - 文件 id
* name - 文件名称
* parentId - 父级 id,null 表示根目录
* type - 文件类型,file 或 folder
* updateTime - 更新时间
* createTime - 创建时间
file: ./content/guide/dataset/third-party/dingtalk_dataset.en.mdx
meta: {
"title": "DingTalk Knowledge Base",
"description": "How to connect DingTalk Knowledge Base to FastGPT"
}
FastGPT supports connecting DingTalk Knowledge Base through a DingTalk internal enterprise app. When creating the dataset, enter `App Key`, `App Secret`, and `User ID`. After creation, open the dataset detail page, click `Add file`, and select the DingTalk workspace, online documents, or folders to import.
Only DingTalk online document text is supported. Binary files such as PDF, Word, Excel, and PPT are not supported.
## 1. Create a DingTalk app

Open the [DingTalk developer app page](https://open-dev.dingtalk.com/fe/app?hash=%23%2Fcorp%2Fapp#/corp/app), then select an internal enterprise app under the target organization.
If you do not have an app yet, create an internal enterprise app from `Application Development`.
## 2. Get the FastGPT fields

| FastGPT field | Where to get it in DingTalk |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `App Key` | Open `Credentials and Basic Information` in the app detail page, then copy `Client ID (formerly AppKey and SuiteKey)`. |
| `App Secret` | Copy `Client Secret (formerly AppSecret and SuiteSecret)` from the same page. |
| `User ID` | Ask the organization contact administrator to open DingTalk admin. Path: [oa.dingtalk.com](https://oa.dingtalk.com/) -> `Contacts` -> `Member Management` -> select the operator member -> copy the member `User ID` from the detail page. |
Notes:
* `App Secret` is sensitive. Do not share it publicly.
* `User ID` is not a phone number, display name, or `unionId`.
* If the member detail page does not show `User ID`, ask the contact administrator to export the member list from `Contacts`; the exported sheet usually contains member `User ID`.
* We recommend using a dedicated DingTalk member as the FastGPT sync account and granting it read-only access to the target workspace.
* Workspaces that this member cannot access will not appear in FastGPT.
## 3. Enable DingTalk app permissions

Open `Permissions` in the DingTalk app detail page, then search for and enable:
| Permission | Purpose |
| --------------------- | ---------------------------------------------------- |
| `qyapi_get_member` | Get the operator ID from `User ID`. |
| `Wiki.Workspace.Read` | List DingTalk workspaces accessible to the operator. |
| `Wiki.Node.Read` | List folders and documents under a workspace. |
| `Storage.File.Read` | Read DingTalk online document content. |
Save and publish the app configuration after enabling permissions. If an error contains `requiredScopes`, enable the permissions listed there.
## 4. Create a DingTalk dataset in FastGPT
1. Open the FastGPT dataset list and click `New`.
2. Select `DingTalk Knowledge Base` under external document sources.
3. Enter:
* `App Key`
* `App Secret`
* `User ID`
4. Confirm creation.
You do not need to select a DingTalk workspace or root directory during creation.
## 5. Add files and sync
After creation:
1. Open the dataset detail page.
2. Click `Add file`.
3. Select the target DingTalk workspace.
4. Select online documents or folders to import.
5. Confirm the import.
When a folder is selected, FastGPT recursively imports supported online documents under that folder.
When DingTalk document content changes, click `Sync` from the imported file menu. FastGPT will read the latest content and update indexes.
file: ./content/guide/dataset/third-party/dingtalk_dataset.mdx
meta: {
"title": "钉钉知识库",
"description": "FastGPT 钉钉知识库功能介绍和使用方式"
}
| | |
| -------------------------------- | -------------------------------- |
|  |  |
FastGPT 支持通过钉钉企业内部应用接入钉钉知识库。创建时只需要填写 `App Key`、`App Secret`、`User ID`,创建完成后进入知识库详情页点击`添加文件`,再选择要导入的钉钉知识库、在线文档或文件夹。
当前仅支持钉钉在线文档文本,不支持 PDF、Word、Excel、PPT 等二进制文件。
## 1. 创建钉钉应用

打开 [钉钉开发者后台应用详情](https://open-dev.dingtalk.com/fe/app?hash=%23%2Fcorp%2Fapp#/corp/app),选择目标企业下的企业内部应用。
如果还没有应用,先进入`应用开发`创建一个企业内部应用。
## 2. 获取 FastGPT 要填写的参数

| FastGPT 字段 | 钉钉里去哪里拿 |
| ------------ | ------------------------------------------------------------------------------------------------------------------------------- |
| `App Key` | 应用详情页左侧进入`凭证与基础信息`,复制`Client ID(原 AppKey 和 SuiteKey)`。 |
| `App Secret` | 同一页面复制`Client Secret(原 AppSecret 和 SuiteSecret)`。 |
| `User ID` | 由企业通讯录管理员进入钉钉管理后台查看。路径:[oa.dingtalk.com](https://oa.dingtalk.com/) -> `通讯录` -> `成员管理` -> 找到作为操作人的成员 -> 点击成员详情,复制该成员的 `User ID`。 |
注意:
* `App Secret` 是密钥,不要公开发送。
* `User ID` 不是手机号、姓名,也不是 `unionId`。
* 如果成员详情页没有展示 `User ID`,让通讯录管理员在`通讯录`里导出成员列表,导出的表格中通常包含成员 `User ID`。
* 建议使用一个专门的钉钉成员作为 FastGPT 同步账号,并给它目标知识库的只读权限。
* 该成员没有权限访问的钉钉知识库,不会出现在 FastGPT 的添加文件列表里。
## 3. 配置钉钉应用权限

在钉钉应用详情页左侧进入`权限管理`,搜索并开通以下权限:
| 权限标识 | 用途 |
| --------------------- | --------------------------- |
| `qyapi_get_member` | 通过 `User ID` 获取接口需要的操作人 ID。 |
| `Wiki.Workspace.Read` | 获取当前操作人可访问的钉钉知识库列表。 |
| `Wiki.Node.Read` | 获取知识库下的文件夹和文档列表。 |
| `Storage.File.Read` | 读取钉钉在线文档正文。 |
权限配置完成后,保存并发布应用配置。若接口报错中出现 `requiredScopes`,按提示补开对应权限。
## 4. 在 FastGPT 中创建钉钉知识库
1. 进入 FastGPT 知识库列表,点击`新建`。
2. 选择`第三方知识库`下的`钉钉知识库`。
3. 填写:
* `App Key`
* `App Secret`
* `User ID`
4. 点击确认创建。
## 5. 添加文件和同步
创建完成后:
1. 进入该知识库详情页。
2. 右上角点击`添加文件`。
3. 选择目标钉钉知识库。
4. 选择要导入的在线文档或文件夹。
5. 确认导入。
选择文件夹时,FastGPT 会递归导入该文件夹下支持的在线文档。
钉钉文档内容更新后,可在已导入文件的更多菜单中点击`同步`,FastGPT 会重新读取最新正文并更新索引。
file: ./content/guide/dataset/third-party/lark_dataset.en.mdx
meta: {
"title": "Lark Knowledge Base",
"description": "Introduction and usage of the FastGPT Lark Knowledge Base"
}
| | |
| ------------------------------- | ------------------------------- |
|  |  |
Starting from FastGPT v4.8.16, commercial edition users can import from Lark knowledge bases. Configure a Lark app's appId and appSecret, then select a **top-level folder in a document space** to import. This feature is currently in beta — some interactions may still need refinement.
Due to Lark API limitations, you cannot directly access all document content. Currently, only files in shared space directories are accessible — personal spaces and wiki content are not supported.
Only cloud document types are supported for import.
## 1. Create a Lark App
Go to the [Lark Open Platform](https://open.feishu.cn/?lang=zh-CN), click **Create App**, select **Custom App**, and fill in the app name.
## 2. Configure App Permissions
After creating the app, configure the following **3 permissions**:
1. View the list of cloud documents in a folder
2. View new-format documents
3. View, comment, edit, and manage all files in the cloud space

## 3. Get the appId and appSecret

## 4. Grant Folder Permissions
Refer to the Lark tutorial: [https://open.feishu.cn/document/server-docs/docs/drive-v1/faq#b02e5bfb](https://open.feishu.cn/document/server-docs/docs/drive-v1/faq#b02e5bfb)
In summary:
1. Add the app you just created to a group chat
2. Grant directory permissions to that group
If your directory already has permissions granted to the "All Members" group, you can skip the steps above and go directly to getting the Folder Token.

## 5. Get the Folder Token
You can find the Folder Token in the page URL. Make sure not to include the question mark.

## 6. Create the Knowledge Base
Using the 3 parameters obtained from steps 3 and 5, create a knowledge base. Select the Lark file library type, fill in the parameters, and click Create.

file: ./content/guide/dataset/third-party/lark_dataset.mdx
meta: {
"title": "飞书知识库",
"description": "FastGPT 飞书知识库功能介绍和使用方式"
}
| | |
| ------------------------------- | ------------------------------- |
|  |  |
FastGPT v4.8.16 版本开始,商业版用户支持飞书知识库导入,用户可以通过配置飞书应用的 appId 和 appSecret,并选中一个**文档空间的顶层文件夹**来导入飞书知识库。目前处于测试阶段,部分交互有待优化。
由于飞书限制,无法直接获取所有文档内容,目前仅可以获取共享空间下文件目录的内容,无法获取个人空间和知识库里的内容。
目前只支持导入云文档类型的内容。
## 1. 创建飞书应用
打开 [飞书开放平台](https://open.feishu.cn/?lang=zh-CN),点击**创建应用**,选择**自建应用**,然后填写应用名称。
## 2. 配置应用权限
创建应用后,进入应用可以配置相关权限,这里需要增加**3个权限**:
1. 获取云空间文件夹下的云文档清单
2. 查看新版文档
3. 查看、评论、编辑和管理云空间中所有文件

## 3. 获取 appId 和 appSecret

## 4. 给 Folder 增加权限
可参考飞书教程: [https://open.feishu.cn/document/server-docs/docs/drive-v1/faq#b02e5bfb](https://open.feishu.cn/document/server-docs/docs/drive-v1/faq#b02e5bfb)
大致总结为:
1. 把刚刚创建的应用拉入一个群里
2. 给这个群增加目录权限
如果你的目录已经给全员组增加权限了,则可以跳过上面步骤,直接获取 Folder Token。

## 5. 获取 Folder Token
可以页面路径上获取 Folder Token,注意不要把问号复制进来。

## 6. 创建知识库
根据 3 和 5 获取到的 3 个参数,创建知识库,选择飞书文件库类型,然后填入对应的参数,点击创建。

file: ./content/guide/dataset/third-party/third_dataset.en.mdx
meta: {
"title": "Third-Party Knowledge Base Development",
"description": "How to integrate a third-party knowledge base with FastGPT",
"sidebarTag": "DEV"
}
import { Alert } from '@/components/docs/Alert';
There are many document libraries available online, such as Lark, Yuque, and others. Different FastGPT users may use different document libraries. FastGPT has built-in support for Lark and Yuque, but if you need to integrate other document libraries, follow this guide.
## Unified API Specification
To provide a unified interface for different document libraries, FastGPT defines a standard API specification with 4 endpoints. See the [API File Library endpoints](./api_dataset.en.mdx).
All built-in document libraries are extensions of the standard API File Library. Refer to the code in `FastGPT/packages/service/core/dataset/apiDataset/yuqueDataset/api.ts` to build extensions for other document libraries. You need to implement 4 endpoints:
1. Get file list
2. Get file content / file link
3. Get original file preview URL
4. Get file detail information
## Building a Third-Party File Library
For this walkthrough, we'll use adding a Lark Knowledge Dataset (FeishuKnowledgeDataset) as an example.
### 1. Add Third-Party Document Library Parameters
First, go to `FastGPT\packages\global\core\dataset\apiDataset.d.ts` in the FastGPT project and add the third-party document library server type. Design the fields based on your needs. For example, the Yuque knowledge base requires `userId` and `token` for authentication.
```ts
export type YuqueServer = {
userId: string;
token?: string;
basePath?: string;
};
```
If the document library supports a `root directory` selection feature, add a `basePath` field. [See the root directory feature](./third_dataset.en.mdx#adding-the-configuration-form)

### 2. Create the Hook File
Each third-party document library uses a Hook pattern to maintain a set of API endpoints. The Hook contains 5 functions to implement.
* Create a folder for your document library under `FastGPT\packages\service\core\dataset\apiDataset\`, then create an `api.ts` file inside it
* In `api.ts`, define the following 5 functions:
* `listFiles`: Get the file list
* `getFileContent`: Get file content / file link
* `getFileDetail`: Get file detail information
* `getFilePreviewUrl`: Get the original file preview URL
* `getFileId`: Get the original file's real ID
### 3. Add the Knowledge Base Type
In `FastGPT\packages\global\core\dataset\type.d.ts`, import your new knowledge base type.

### 4. Add Knowledge Base Data Retrieval
In `FastGPT\packages\global\core\dataset\apiDataset\utils.ts`, add the following content.

### 5. Add Knowledge Base Invocation Method
In `FastGPT\packages\service\core\dataset\apiDataset\index.ts`, add the following content.

## Adding the Frontend
Add your i18n translations in `FastGPT\packages\web\i18n\zh-CN\dataset.json`, `FastGPT\packages\web\i18n\en\dataset.json`, and `FastGPT\packages\web\i18n\zh-Hant\dataset.json`. Using Chinese translations as an example, you'll generally need the following:

In `FastGPT\packages\service\support\user/audit\util.ts`, add the following to support i18n translation retrieval.

The i18n translation content is stored in `FastGPT\packages\web\i18n\zh-Hant\account_team.json`, `FastGPT\packages\web\i18n\zh-CN\account_team.json`, and `FastGPT\packages\web\i18n\en\account_team.json`. The field format is `dataset.XXX_dataset`. For example, for the Lark knowledge base, the field value is `dataset.feishu_knowledge_dataset`.
Add your knowledge base icons under `FastGPT\packages\web\components\common\Icon\icons\core\dataset\`. You need two icons: `Outline` (monochrome) and `Color` (colored), as shown below.

In `FastGPT\packages\web\components\common\Icon\constants.ts`, register your icons. The `import` path points to where the icons are stored.

In `FastGPT\packages\global\core\dataset\constants.ts`, add your knowledge base type to both `DatasetTypeEnum` and `ApiDatasetTypeMap`.
| | |
| ----------------------------- | ------------------------------ |
|  |  |
The `courseUrl` field links to the relevant documentation — add it if available.
Documentation goes in `FastGPT/document/content/guide/build/workflow/nodes/knowledge_base_search_merge.mdx`.
The `label` value is the knowledge base name you added via i18n translations.
`icon` and `avatar` are the two icons you added earlier.
In `FastGPT\projects\app\src\pages\dataset\list\index.tsx`, add the following. This file handles the menu that appears when clicking the "New" button on the knowledge base list page. Your knowledge base must be added here to be creatable.

In `FastGPT\projects\app\src\pageComponents\dataset\detail\Info\index.tsx`, add the following. This configuration corresponds to the UI shown below.
| | |
| ------------------------------ | ------------------------------ |
|  |  |
## Adding the Configuration Form
In `FastGPT\projects\app\src\pageComponents\dataset\ApiDatasetForm.tsx`, add the following. This file handles the field input form when creating a knowledge base.
| | | |
| ------------------------------ | ------------------------------ | ------------------------------ |
|  |  |  |
The two components added in the code render the root directory selector, corresponding to the `getFileDetail` API method. If your knowledge base doesn't support this, you can omit them.
```
{renderBaseUrlSelector()} // Renders the `Base URL` field
{renderDirectoryModal()} // The `Select Root Directory` modal that appears when clicking `Select` (see image)
```
| | |
| ------------------------------ | ------------------------------ |
|  |  |
If the knowledge base needs root directory support, also add the following in the `ApiDatasetForm` file.
### 1. Parse the Knowledge Base Type
Parse your knowledge base type from `apiDatasetServer`, as shown:

### 2. Add Root Directory Selection Logic and `parentId` Assignment
Add root directory selection logic to ensure the user has filled in all required fields for the API methods, such as the Token.

### 3. Add Field Validation and Assignment Logic
Verify that all required fields are present before calling the API, and assign the root directory value to the corresponding field after selection.

## Tips
After creating the knowledge base, we recommend running a full test of all knowledge base features to check for issues. If you encounter problems that aren't covered in this documentation, it's likely that some configuration was missed. Do a global search for `YuqueServer` and `yuqueServer` to verify that your type has been added everywhere it's needed.
file: ./content/guide/dataset/third-party/third_dataset.mdx
meta: {
"title": "第三方知识库开发",
"description": "本节详细介绍如何在FastGPT上自己接入第三方知识库",
"sidebarTag": "DEV"
}
import { Alert } from '@/components/docs/Alert';
目前,互联网上拥有各种各样的文档库,例如飞书,语雀等等。 FastGPT 的不同用户可能使用的文档库不同,目前 FastGPT 内置了飞书、语雀文档库,如果需要接入其他文档库,可以参考本节内容。
## 统一的接口规范
为了实现对不同文档库的统一接入,FastGPT 对第三方文档库进行了接口的规范,共包含 4 个接口内容,可以[查看 API 文件库接口](./api_dataset.mdx)。
所有内置的文档库,都是基于标准的 API 文件库进行扩展。可以参考`FastGPT/packages/service/core/dataset/apiDataset/yuqueDataset/api.ts`中的代码,进行其他文档库的扩展。一共需要完成 4 个接口开发:
1. 获取文件列表
2. 获取文件内容/文件链接
3. 获取原文预览地址
4. 获取文件详情信息
## 开始一个第三方文件库
为了方便讲解,这里以添加飞书知识库( FeishuKnowledgeDataset )为例。
### 1. 添加第三方文档库参数
首先,要进入 FastGPT 项目路径下的`FastGPT\packages\global\core\dataset\apiDataset.d.ts`文件,添加第三方文档库 Server 类型。知识库类型的字段由自己设计,主要是自己需要那些内容。例如,语雀知识库中,需要提供`userId`、`token`两个字段作为鉴权信息。
```ts
export type YuqueServer = {
userId: string;
token?: string;
basePath?: string;
};
```
如果文档库有`根目录`选择的功能,需要设置添加一个字段`basePath`[点击查看`根目录`功能](./third_dataset.mdx#添加配置表单)

### 2. 创建 Hook 文件
每个第三方文档库都会采用 Hook 的方式来实现一套 API 接口的维护,Hook 里包含 5 个函数需要完成。
* 在`FastGPT\packages\service\core\dataset\apiDataset\`下创建一个文档库的文件夹,然后在文件夹下创建一个`api.ts`文件
* 在`api.ts`文件中,需要完成 5 个函数的定义,分别是:
* `listFiles`:获取文件列表
* `getFileContent`:获取文件内容/文件链接
* `getFileDetail`:获取文件详情信息
* `getFilePreviewUrl`:获取原文预览地址
* `getFileId`: 获取原文件真实Id
### 3. 添加知识库类型
在`FastGPT\packages\global\core\dataset\type.d.ts`文件中,导入自己创建的知识库类型。

### 4. 添加知识库数据获取
在`FastGPT\packages\global\core\dataset\apiDataset\utils.ts`文件中,添加如下内容。

### 5. 添加知识库调用方法
在`FastGPT\packages\service\core\dataset\apiDataset\index.ts`文件下,添加如下内容。

## 添加前端
`FastGPT\packages\web\i18n\zh-CN\dataset.json`,`FastGPT\packages\web\i18n\en\dataset.json`和`FastGPT\packages\web\i18n\zh-Hant\dataset.json`中添加自己的 I18n 翻译,以中文翻译为例,大体需要如下几个内容:

`FastGPT\packages\service\support\user/audit\util.ts`文件下添加如下内容,以支持获取 I18n 翻译。

此次 I18n 翻译内容存放在`FastGPT\packages\web\i18n\zh-Hant\account_team.json`,`FastGPT\packages\web\i18n\zh-CN\account_team.json`和`FastGPT\packages\web\i18n\en\account_team.json`,字段格式为`dataset.XXX_dataset`,以飞书知识库为例,字段值为`dataset.feishu_knowledge_dataset`
`FastGPT\packages\web\components\common\Icon\icons\core\dataset\`添加自己的知识库图标,一共是两个,分为`Outline`和`Color`,分别是有颜色的和无色的,具体看如下图片。

在`FastGPT\packages\web\components\common\Icon\constants.ts`文件中,添加自己的图标。 `import` 是图标的存放路径。

在`FastGPT\packages\global\core\dataset\constants.ts`中,添加自己的知识库类型,分别要在`DatasetTypeEnum`和`ApiDatasetTypeMap`中添加内容。
| | |
| ----------------------------- | ------------------------------ |
|  |  |
`courseUrl`字段是相应的文档说明,如果有的话,可以添加。
文档添加在`FastGPT/document/content/guide/build/workflow/nodes/knowledge_base_search_merge.mdx`
`label`内容是自己之前通过 i18n 翻译添加的知识库名称的。
`icon`和`avatar`是自己之前添加的两个图标
在`FastGPT\projects\app\src\pages\dataset\list\index.tsx`文件下,添加如下内容。这个文件负责的是知识库列表页的`新建`按钮点击后的菜单,只有在该文件添加知识库后,才能创建知识库。

在`FastGPT\projects\app\src\pageComponents\dataset\detail\Info\index.tsx`文件下,添加如下内容。此处配置对应ui界面的如下。
| | |
| ------------------------------ | ------------------------------ |
|  |  |
## 添加配置表单
在`FastGPT\projects\app\src\pageComponents\dataset\ApiDatasetForm.tsx`文件下,添加自己如下内容。这个文件负责的是创建知识库页的字段填写。
| | | |
| ------------------------------ | ------------------------------ | ------------------------------ |
|  |  |  |
代码中添加的两个组件是对根目录选择的渲染,对应设计的 api 的 getfiledetail 方法,如果你的知识库不支持,你可以不引用。
```
{renderBaseUrlSelector()} //这是对`Base URL`字段的渲染
{renderDirectoryModal()} //点击`选择`后出现的`选择根目录`窗口,见图
```
| | |
| ------------------------------ | ------------------------------ |
|  |  |
如果知识库需要支持根目录,还需要在`ApiDatasetForm`文件中添加如下内容。
### 1. 解析知识库类型
需要从`apiDatasetServer`解析出自己的知识库类型,如图:

### 2. 添加选择根目录逻辑和`parentId`赋值逻辑
需要添加根目录选择逻辑,来确保用户已经填写了调动的 api 方法所必需的字段,比如 Token 之类的。

### 3. 添加字段检查和赋值逻辑
需要在调用方法前再次检测是否以及获取完所有必须字段,在选择根目录后,将根目录值赋值给对应的字段。

## 提示
建议知识库创建完成后,完整测试一遍知识库的功能,以确定有无漏洞,如果你的知识库添加有问题,且无法在文档找到对应的文件解决,一定是杂项没有添加完全,建议重复一次全局搜索`YuqueServer`和`yuqueServer`,检查是否有地方没有加上自己的类型。
file: ./content/guide/dataset/third-party/yuque_dataset.en.mdx
meta: {
"title": "Yuque File Library",
"description": "Introduction and usage of the FastGPT Yuque File Library"
}
| | |
| ------------------------------- | ------------------------------- |
|  |  |
Starting from FastGPT v4.8.16, commercial edition users can import from Yuque file libraries by configuring a Yuque token and uid. This feature is currently in beta — some interactions may still need refinement.
## 1. Get the Yuque Token and UID
Go to the Yuque homepage > click your avatar > Settings to find the relevant parameters.

Follow the images below to get the Token and User ID. Make sure to assign the appropriate permissions to the Token:
**Personal Edition**:
| Get Token | Add Permissions | Get User ID |
| ------------------------------- | ------------------------------- | ------------------------------- |
|  |  |  |
**Enterprise Edition**:
| Get Token | Get User ID |
| -------------------------------- | -------------------------------- |
|  |  |
## 2. Create the Knowledge Base
Using the token and uid from the previous step, create a knowledge base. Select the Yuque file library type, fill in the parameters, and click Create.


## 3. Import Documents
After creating the knowledge base, click `Add File` to import from your Yuque document library and follow the on-screen guidance.
The Yuque knowledge base supports scheduled sync — it scans once daily at varying times. If documents have been updated, they will be synced automatically. You can also trigger a manual sync.

file: ./content/guide/dataset/third-party/yuque_dataset.mdx
meta: {
"title": "语雀文件库",
"description": "FastGPT 语雀文件库功能介绍和使用方式"
}
| | |
| ------------------------------- | ------------------------------- |
|  |  |
FastGPT v4.8.16 版本开始,商业版用户支持语雀文件库导入,用户可以通过配置语雀的 token 和 uid 来导入语雀文档库。目前处于测试阶段,部分交互有待优化。
## 1. 获取语雀的 token 和 uid
在语雀首页 - 个人头像 - 设置,可找到对应参数。

参考下图获取 Token 和 User ID,注意给 Token 赋值权限:
**个人版**:
| 获取 Token | 增加权限 | 获取 User ID |
| ------------------------------- | ------------------------------- | ------------------------------- |
|  |  |  |
**企业版**:
| 获取 Token | 获取 User ID |
| -------------------------------- | -------------------------------- |
|  |  |
## 2. 创建知识库
使用上一步获取的 token 和 uid,创建知识库,选择语雀文件库类型,然后填入对应的参数,点击创建。


## 3. 导入文档
创建完知识库后,点击`添加文件`即可导入语雀的文档库,跟随引导即可。
语雀知识库支持定时同步功能,每天会不定时的扫描一次,如果文档有更新,则会进行同步,也可以进行手动同步。

file: ./content/guide/workspace/team/invitation_link.en.mdx
meta: {
"title": "Invitation Links",
"description": "How to use invitation links to invite team members"
}
Starting from v4.9.1, team member invitations use the **invitation link** method, replacing the previous username-based approach.
After upgrading, any pending invitations that haven't been accepted will be automatically cleared. Please use invitation links to re-invite members.
## How to Use
1. **On the team management page, admins can click the "Invite Members" button to open the invitation dialog**

2. **In the invitation dialog, click "Create Invitation Link" to generate a new link**

3. **Fill in the details**

Link description: We recommend describing the intended use case or purpose. The description cannot be changed after creation.
Expiration: 30 minutes, 7 days, or 1 year
Usage limit: 1 person or unlimited
4. **Click "Copy Link" and send it to the people you want to invite**

5. **When a user visits the link, they will be redirected to the login page if not logged in or registered. After logging in, they will be taken to the team page to handle the invitation.**
> Invitation links look like: fastgpt.cn/account/team?invitelinkid=xxxx

Click "Accept" to join the team.
Click "Ignore" to close the dialog. The user can still accept the invitation by visiting the link again later.
## Link Expiration and Auto-Cleanup
### Why Links Expire
Links are manually disabled by an admin.
The invitation link reaches its expiration date and is automatically disabled.
A single-use link (1 person limit) has already been used.
Expired links cannot be accessed or re-enabled.
### Link Limits
Each user can have up to 10 **active** invitation links at a time.
### Auto-Cleanup
Expired links are automatically deleted after 30 days.
file: ./content/guide/workspace/team/invitation_link.mdx
meta: {
"title": "邀请链接说明文档",
"description": "如何使用邀请链接来邀请团队成员"
}
v4.9.1 团队邀请成员将开始使用「邀请链接」的模式,弃用之前输入用户名进行添加的形式。
在版本升级后,原收到邀请还未加入团队的成员,将自动清除邀请。请使用邀请链接重新邀请成员。
## 如何使用
1. **在团队管理页面,管理员可点击「邀请成员」按钮打开邀请成员弹窗**

2. **在邀请成员弹窗中,点击「创建邀请链接」按钮,创建邀请链接。**

3. **输入对应内容**

链接描述:建议将链接描述为使用场景或用途。链接创建后不支持修改噢。
有效期:30分钟,7天,1年
有效人数:1人,无限制
4. **点击复制链接,并将其发送给想要邀请的人。**

5. **用户访问链接后,如果未登录/未注册,则先跳转到登录页面进行登录。在登录后将进入团队页面,处理邀请。**
> 邀请链接形如:fastgpt.cn/account/team?invitelinkid=xxxx

点击接受,则用户将加入团队
点击忽略,则关闭弹窗,用户下次访问该邀请链接则还可以选择加入。
## 链接失效和自动清理
### 链接失效原因
手动停用链接
邀请链接到达有效期,自动停用
有效人数为1人的链接,已有1人通过邀请链接加入团队。
停用的链接无法访问,也无法再次启用。
### 链接上限
一个用户最多可以同时存在 10 个**有效的**邀请链接。
### 链接自动清理
失效的链接将在 30 天后自动清理。
file: ./content/guide/workspace/team/team_roles_permissions.en.mdx
meta: {
"title": "Teams, Groups & Permissions",
"description": "How to manage FastGPT teams, member groups, and permission settings"
}
# Teams, Groups & Permissions
## Permission System Overview
FastGPT's permission system combines **attribute-based** and **role-based** access control, providing fine-grained permission management for team collaboration. Through **members, departments, and groups**, you can flexibly configure access to teams, apps, and knowledge bases.
## Teams
Each user can belong to multiple teams. The system automatically creates an initial team for every user. Manual creation of additional teams is not currently supported.
## Permission Management
FastGPT offers three permission management levels:
**Member Permissions**: Highest priority, directly assigned to individuals
**Department & Group Permissions**: Use union logic, lower priority than member permissions
Permission evaluation follows this logic:
First, check the user's individual member permissions
Then, check permissions from the user's departments and groups (union)
Final permissions are the combination of the above
Authorization logic:

### Resource Permissions
Different **resources** have different permissions.
Resources refer to concepts like apps, knowledge bases, teams, etc.
The table below shows the manageable permissions for different resources.
Resource
Manageable Permissions
Description
Team
Create Apps
Create, delete, and other basic operations
Create Knowledge Bases
Create, delete, and other basic operations
Create Team APIKey
Create, delete, and other basic operations
Manage Members
Invite/remove users, create groups, etc.
App
Can Use
Allows conversation interaction
Can Edit
Modify basic info, workflow orchestration, etc.
Can Manage
Add or remove collaborators
Knowledge Base
Can Use
Can call this knowledge base in apps
Can Edit
Modify knowledge base content
Can Manage
Add or remove collaborators
### Collaborators
You must add **collaborators** before managing their permissions:

When managing team permissions, first select members/organizations/groups, then configure permissions.

For resources like apps and knowledge bases, you can directly modify member permissions.

Team permissions are set on a dedicated permissions page.

## Special Permissions
### Admin Permissions
Admins primarily manage resource collaboration relationships, with these limitations:
* Cannot modify or remove their own permissions
* Cannot modify or remove other admins' permissions
* Cannot grant admin permissions to other collaborators
### Owner Permissions
Each resource has a unique Owner with the highest permissions for that resource. Owners can transfer ownership, but will lose all permissions to the resource after transfer.
### Root Permissions
Root is the system's only super admin account, with complete access and management rights to all resources across all teams.
## Tips
### 1. Set Default Team Permissions
Use the "All Members Group" to quickly set baseline permissions for the entire team. For example, grant everyone access to an app.
**Note**: Individual member permissions override all-member group permissions. For example, if App A has all-member edit permissions, but User M is individually set to use-only, User M can only use the app, not edit it.
### 2. Batch Permission Management
Create groups or organizations to efficiently manage permissions for multiple users. Add users to a group, then grant permissions to the entire group.
### Developer Reference
> The following content is for developers. Skip if you're not doing custom development.
#### Permission Design Principles
FastGPT's permission system is inspired by Linux permissions, using binary storage for permission bits. A permission bit of 1 means the permission is granted, 0 means no permission. Owner permissions are specially marked as all 1s.
#### Permission Table
Permission information is stored in MongoDB's resource\_permissions collection, with these main fields:
* teamId: Team identifier
* tmbId/groupId/orgId: Permission subject (one of three)
* resourceType: Resource type (team/app/dataset)
* permission: Permission value (number)
* resourceId: Resource ID (null for team resources)
The system implements flexible and precise permission control through this data structure.
The schema for this table is defined in packages/service/support/permission/schema.ts:
```typescript
export const ResourcePermissionSchema = new Schema({
teamId: {
type: Schema.Types.ObjectId,
ref: TeamCollectionName
},
tmbId: {
type: Schema.Types.ObjectId,
ref: TeamMemberCollectionName
},
groupId: {
type: Schema.Types.ObjectId,
ref: MemberGroupCollectionName
},
orgId: {
type: Schema.Types.ObjectId,
ref: OrgCollectionName
},
resourceType: {
type: String,
enum: Object.values(PerResourceTypeEnum),
required: true
},
permission: {
type: Number,
required: true
},
// Resrouce ID: App or DataSet or any other resource type.
// It is null if the resourceType is team.
resourceId: {
type: Schema.Types.ObjectId
}
});
```
file: ./content/guide/workspace/team/team_roles_permissions.mdx
meta: {
"title": "团队&成员组&权限",
"description": "如何管理 FastGPT 团队、成员组及权限设置"
}
# 团队 & 成员组 & 权限
## 权限系统简介
FastGPT
权限系统融合了基于**属性**和基于**角色**的权限管理范式,为团队协作提供精细化的权限控制方案。通过**成员、部门和群组**三种管理模式,您可以灵活配置对团队、应用和知识库等资源的访问权限。
## 团队
每位用户可以同时归属于多个团队,系统默认为每位用户创建一个初始团队。目前暂不支持用户手动创建额外团队。
## 权限管理
FastGPT 提供三种权限管理维度:
**成员权限**:最高优先级,直接赋予个人的权限
**部门与群组权限**:采用权限并集原则,优先级低于成员权限
权限判定遵循以下逻辑:
首先检查用户的个人成员权限
其次检查用户所属部门和群组的权限(取并集)
最终权限为上述结果的组合
鉴权逻辑如下:

### 资源权限
对于不同的**资源**,有不同的权限。
这里说的资源,是指应用、知识库、团队等等概念。
下表为不同资源,可以进行管理的权限。
资源
可管理权限
说明
团队
创建应用
创建,删除等基础操作
创建知识库
创建,删除等基础操作
创建团队 APIKey
创建,删除等基础操作
管理成员
邀请、移除用户,创建群组等
应用
可使用
允许进行对话交互
可编辑
修改基本信息,进行流程编排等
可管理
添加或删除协作者
知识库
可使用
可以在应用中调用该知识库
可编辑
修改知识库的内容
可管理
添加或删除协作者
### 协作者
必须先添加**协作者**,才能对其进行权限管理:

管理团队权限时,需先选择成员/组织/群组,再进行权限配置。

对于应用和知识库等资源,可直接修改成员权限。

团队权限在专门的权限页面进行设置

## 特殊权限说明
### 管理员权限
管理员主要负责管理资源的协作关系,但有以下限制:
* 不能修改或移除自身权限
* 不能修改或移除其他管理员权限
-不能将管理员权限赋予其他协作者
### Owner 权限
每个资源都有唯一的 Owner,拥有该资源的最高权限。Owner
可以转移所有权,但转移后原 Owner 将失去对资源的权限。
### Root 权限
Root
作为系统唯一的超级管理员账号,对所有团队的所有资源拥有完全访问和管理权限。
## 使用技巧
### 1. 设置团队默认权限
利用"全员群组"可快速为整个团队设置基础权限。例如,为应用设置全员可访问权限。
**注意**:个人成员权限会覆盖全员组权限。例如,应用 A
设置了全员编辑权限,而用户 M 被单独设置为使用权限,则用户 M
只能使用而无法编辑该应用。
### 2. 批量权限管理
通过创建群组或组织,可以高效管理多用户的权限配置。先将用户添加到群组,再对群组整体授权。
### 开发者参考
> 以下内容面向开发者,如不涉及二次开发可跳过。
#### 权限设计原理
FastGPT 权限系统参考 Linux 权限设计,采用二进制方式存储权限位。权限位为
1 表示拥有该权限,为 0 表示无权限。Owner 权限特殊标记为全 1。
#### 权限表
权限信息存储在 MongoDB 的 resource\_permissions 集合中,其主要字段包括:
* teamId: 团队标识
* tmbId/groupId/orgId: 权限主体(三选一)
* resourceType: 资源类型(team/app/dataset)
* permission: 权限值(数字)
* resourceId: 资源ID(团队资源为null)
系统通过这一数据结构实现了灵活而精确的权限控制。
对于这个表的 Schema 定义在 packages/service/support/permission/schema.ts
文件中。定义如下:
```typescript
export const ResourcePermissionSchema = new Schema({
teamId: {
type: Schema.Types.ObjectId,
ref: TeamCollectionName
},
tmbId: {
type: Schema.Types.ObjectId,
ref: TeamMemberCollectionName
},
groupId: {
type: Schema.Types.ObjectId,
ref: MemberGroupCollectionName
},
orgId: {
type: Schema.Types.ObjectId,
ref: OrgCollectionName
},
resourceType: {
type: String,
enum: Object.values(PerResourceTypeEnum),
required: true
},
permission: {
type: Number,
required: true
},
// Resrouce ID: App or DataSet or any other resource type.
// It is null if the resourceType is team.
resourceId: {
type: Schema.Types.ObjectId
}
});
```
file: ./content/self-host/config/model/intro.en.mdx
meta: {
"title": "Model Configuration",
"description": "FastGPT model configuration guide"
}
import { Alert } from '@/components/docs/Alert';
import { Accordion, Accordions } from 'fumadocs-ui/components/accordion';
## Introduction
FastGPT uses the `AI Proxy` service to connect to different model providers. AI Proxy also provides load balancing, model logging, and analytics dashboards to help you monitor model usage.
Notes:
Only one speech recognition model can be active at a time, so you only need to configure one.
The system requires at least one language model and one embedding model to function properly.
### Architecture Diagram

### Model Types
1. Language Models - Text-based conversations; multimodal models also support image recognition.
2. Embedding Models - Index text chunks for semantic text retrieval.
3. Rerank Models - Reorder retrieval results to optimize search ranking.
4. Text-to-Speech (TTS) - Convert text to audio.
5. Speech-to-Text (STT) - Convert audio to text.
### Key Terminology
* Model ID: The value of the `model` field in the API request body. Must be globally unique.
* Model Name: The display name of the model, which can be customized.
* Model Channel: The protocol of different model providers, such as OpenAI, Anthropic, Google, etc. Most self-hosted channels follow the OpenAI protocol. A single model can be configured across multiple channels to enable load balancing.
* Custom Request URL / Key: Allows you to bypass Model Channels and send requests directly to a custom endpoint. You need to provide the full request URL and token. Generally not needed (not recommended as it's harder to manage).
## Adding Channels and Models
You can configure models from the `Account - Model Providers` page in FastGPT.
### 1. Create a Channel
Switch to the `Model Channels` tab. Note that you can only add models that already exist in `Model Configuration`. The system only includes mainstream models by default — if you need additional models, add them in `Model Configuration` first.

Click "Add Channel" in the top-right corner to open the channel configuration page.

Using Alibaba Bailian models as an example:

1. Channel Name: A display label for the channel, used for identification only.
2. Protocol Type: The API protocol for the model. Generally, select the provider that offers the model. Most providers support the OpenAI protocol, so you can also choose OpenAI as the protocol type.
3. Models: The specific models available in this channel. The system includes popular models by default. If the model you need isn't in the dropdown, click "Add Model" to [add a custom model](./intro.en.mdx#add-a-custom-model).
4. Model Mapping: Maps the model name in FastGPT requests to the actual model name at the provider. For example:
```json
{
"gpt-4o-test": "gpt-4o"
}
```
In FastGPT, the model is `gpt-4o-test`, and requests to AI Proxy also use `gpt-4o-test`. When AI Proxy forwards the request upstream, the actual `model` value becomes `gpt-4o`.
5. Proxy URL: Do not enter the full model request URL. Enter the `BaseUrl` instead, and check whether `/v1` needs to be appended.
6. API Key: The API credentials obtained from the model provider. Some providers require multiple keys — follow the on-screen prompts to enter them.
Click "Add" to save. The new channel will appear under "Model Channels".

### 2. Channel Testing
You can test the channel to verify that the configured models are working properly.

Click "Model Test" to see the list of configured models, then click "Start Test".

Once testing completes, you'll see the results and response times for each model.

### 3. Enable Models
The system includes models from major providers by default. If you're not familiar with the configuration, simply click `Enable`. The `Model ID` corresponds to the `Model` in `Model Channels`.
Click "Enable" to activate the model.
| Enable Models | Model ID Mapping |
| ------------------------------------------------- | -------------------------------------------------- |
|  |  |
### 4. Test Models
FastGPT provides simple tests for each model type on the UI to verify that models are working correctly. Each test sends an actual request using a template.

## Model Configuration
### Edit Model Configuration
Click the gear icon next to a model to open its configuration. Different model types have different configuration options.
| | |
| ------------------------------------------------- | ------------------------------------------------- |
|  |  |
### Add a Custom Model
If the built-in models don't meet your needs, you can add custom models. If the `Model ID` matches an existing built-in model ID, it will be treated as a modification rather than a new model.
1. **Add via Form**
| | |
| ------------------------------------------------- | ------------------------------------------------- |
|  |  |
2. **Add via Configuration File**
If you find it tedious to configure models through the UI, you can use a configuration file instead. This is also useful for quickly replicating the configuration from one system to another.
| | |
| ------------------------------------------------- | ------------------------------------------------- |
|  |  |
```json
{
"model": "Model ID",
"metadata": {
"isCustom": true, // Whether this is a custom model
"isActive": true, // Whether the model is enabled
"provider": "OpenAI", // Model provider, used for categorization. Built-in providers: https://github.com/labring/FastGPT/blob/main/packages/global/core/ai/provider.ts. You can submit a PR for new providers, or use "Other"
"model": "gpt-5", // Model ID (corresponds to the model name in the channel)
"name": "gpt-5", // Display name
"maxContext": 125000, // Maximum context length
"maxResponse": 16000, // Maximum response length
"quoteMaxToken": 120000, // Maximum citation content tokens
"maxTemperature": 1.2, // Maximum temperature
"charsPointsPrice": 0, // Credits per 1k tokens (commercial edition)
"censor": false, // Enable content moderation (commercial edition)
"vision": true, // Supports image input
"toolChoice": true, // Supports tool selection (used in classification, extraction, and tool calls)
"functionCall": false, // Supports function calling (used in classification, extraction, and tool calls). toolChoice takes priority; if false, functionCall is used; if also false, prompt mode is used
"customCQPrompt": "", // Custom text classification prompt (for models without tool/function call support)
"customExtractPrompt": "", // Custom content extraction prompt
"defaultSystemChatPrompt": "", // Default system prompt included in conversations
"defaultConfig": {}, // Default config sent with API requests (e.g., GLM4's top_p)
"fieldMap": {} // Field mapping (e.g., o1 models need max_tokens mapped to max_completion_tokens)
}
}
```
```json
{
"model": "Model ID",
"metadata": {
"isCustom": true, // Whether this is a custom model
"isActive": true, // Whether the model is enabled
"provider": "OpenAI", // Model provider
"model": "text-embedding-3-small", // Model ID
"name": "text-embedding-3-small", // Display name
"charsPointsPrice": 0, // Credits per 1k tokens
"defaultToken": 512, // Default token count for text splitting
"maxToken": 3000 // Maximum token count
}
}
```
```json
{
"model": "Model ID",
"metadata": {
"isCustom": true, // Whether this is a custom model
"isActive": true, // Whether the model is enabled
"provider": "BAAI", // Model provider
"model": "bge-reranker-v2-m3", // Model ID
"name": "ReRanker-Base", // Display name
"requestUrl": "", // Custom request URL
"requestAuth": "", // Custom request authentication
"type": "rerank" // Model type
}
}
```
```json
{
"model": "Model ID",
"metadata": {
"isActive": true, // Whether the model is enabled
"isCustom": true, // Whether this is a custom model
"type": "tts", // Model type
"provider": "FishAudio", // Model provider
"model": "fishaudio/fish-speech-1.5", // Model ID
"name": "fish-speech-1.5", // Display name
"voices": [
// Available voices
{
"label": "fish-alex", // Voice name
"value": "fishaudio/fish-speech-1.5:alex" // Voice ID
},
{
"label": "fish-anna", // Voice name
"value": "fishaudio/fish-speech-1.5:anna" // Voice ID
}
],
"charsPointsPrice": 0 // Credits per 1k tokens
}
}
```
```json
{
"model": "whisper-1",
"metadata": {
"isActive": true, // Whether the model is enabled
"isCustom": true, // Whether this is a custom model
"provider": "OpenAI", // Model provider
"model": "whisper-1", // Model ID
"name": "whisper-1", // Display name
"charsPointsPrice": 0, // Credits per 1k tokens
"type": "stt" // Model type
}
}
```
## Other
### Channel Priority
Range: 1–100. Higher values are prioritized.

### Enable / Disable Channels
In the control menu on the right side of each channel, you can enable or disable it. Disabled channels will no longer serve model requests.

### Model Call Logs
Model calls made through channels are logged on the `Call Logs` page. Logs include input/output tokens, request time, latency, request URL, and more. Failed requests show detailed parameters and error messages for debugging, but logs are retained for only 1 hour by default (configurable via environment variables).

### Self-Hosted Models
[See the ReRank model deployment tutorial](../../custom-models/bge-rerank.en.mdx)
### Custom Request URL
If you set a custom request URL, requests will bypass `Model Channels` and be sent directly to the specified endpoint. You must provide the full request URL, for example:
* LLM: \[host]/v1/chat/completions
* Embedding: \[host]/v1/embeddings
* STT: \[host]/v1/audio/transcriptions
* TTS: \[host]/v1/audio/speech
* Rerank: \[host]/v1/rerank
The custom request key is included as the `Authorization: Bearer xxx` header when sending requests to the custom URL.
All endpoints follow the OpenAI model format. Refer to the [OpenAI API documentation](https://platform.openai.com/docs/api-reference/guide) for details.
Since OpenAI does not provide a Rerank model, the Rerank endpoint follows the Cohere format. [See request examples](../../troubleshooting/model-errors.en.mdx)
### Migrating from OneAPI to AI Proxy
If you were using OneAPI in an older version, you can migrate your channel configuration to AI Proxy using a script.
Send the following HTTP request from any terminal. Replace `{{host}}` with the AI Proxy address and `{{admin_key}}` with the `ADMIN_KEY` value in AI Proxy.
The `dsn` parameter in the request body is the MySQL connection string for OneAPI.
```bash
curl --location --request POST '{{host}}/api/channels/import/oneapi' \
--header 'Authorization: Bearer {{admin_key}}' \
--header 'Content-Type: application/json' \
--data-raw '{
"dsn": "mysql://root:s5mfkwst@tcp(dbconn.sealoshzh.site:33123)/mydb"
}'
```
A successful response will return `"success": true`.
Note that the migration script performs a simple data mapping — it primarily transfers `proxy URLs`, `models`, and `API keys`. Manual verification after migration is recommended.
file: ./content/self-host/config/model/intro.mdx
meta: {
"title": "模型配置说明",
"description": "FastGPT 模型配置说明"
}
import { Alert } from '@/components/docs/Alert';
import { Accordion, Accordions } from 'fumadocs-ui/components/accordion';
## 介绍
FastGPT 借助 `AI Proxy` 服务,可以连接到不同的模型提供商。同时 `AI Proxy` 还提供了负载均衡、模型日志、数据看板等能力,方便检测模型调用情况。