背景

在 Pi Coding Agent 中使用 LLM API 的时候,特别是公益站,API 经常会出现不稳定的情况

  1. 明显错误:请求失败,返回失败,返回不合法的内容(空内容、不符合规范的json返回、html报错代码等等)。
  2. 返回截断:明明输出还没有结束,但是 api 的流式输出就是完成了,只有发继续才能接着输出。

对于第一种,使用一些重试工具可以解决大部分问题,比如 LLM 网关,很多都自带请求失败重试或者是渠道转移的功能,甚至有些工具还提供了检测空内容,json合法性等高级重试功能,好像 ccload 和 axonhub 就有这种功能。

我是使用 octopus 网关 + Pi Coding Agent 的自带重试功能,所以对于第一种问题,已经是提前解决了,除非出现别的类似的问题,但是不在网关和 Pi Coding Agent 检测范围之内。

印象中是有的,可能用在这里不太准确,但是想不到别的了🤣。比如有的网关会跳过 429 错误码不重试,认为这个号已经被限流了,再重试也没用,但是这个错误码可能是公益站上游传过来的,代表这个号不行了,重试切换到下一个号可能就行了,因为公益站一般都会有很多个帐号在轮训。只是稍微举例一下,有的网关支持自定义重试的状态码,比如 GPT Load,那么这个问题也不是问题了。

对于第二种,没有什么好的方法,如果我们不在屏幕面前,我们就不知道是不是真的已经完成了,直接给 AI 发个继续自然是成本最低的方法了,当然还有更高的玩法,比如 subagent 监控什么的,这个就麻烦很多了。

所以这个插件最基本的理解就是代替你自动发个继续,功能上就是这个样子,帮助你解决第二类的问题。

环境要求

  • 支持 tool call 的模型
  • Pi Coding Agent(版本待确定),目前使用的版本是 0.81.1

Prompt 设计

最基本的方案就是,等 Pi Coding Agent 空闲的时候,启动一个倒计时,倒计时一到就直接发某一句话给 LLM,如果 LLM 认为工作还没有执行完,那就接着输出,如果完成了,那么就调用工具结束倒计时,否则就会一直重试发这句话,直到最大次数。


Q:为什么不是直接发,还要搞个倒计时?
A:直接发是可以的,因为这里还得和用户操作协调,比如用户习惯了默认开启这个插件,恰好用户也在屏幕面前,agent 空闲后用户可能想操作一些什么,比如发新的请求,或者是看到没有问题,就不要接着发继续了。倒计时等于是有个缓冲的时间,插件检测到用户操作,就会暂停倒计时,就不会发出继续,自然不会跟用户竞争操作了。

将倒计时设置为 0,就等于是没有缓冲了,不过没有测试过设置为 0 是什么情况,我都是用 5s 的情况。

Q:如何定义空闲?
A:探讨实现文章中讨论

Q:检测到什么操作算暂停?
A:探讨实现文章中讨论

Q:为什么是工具,而不是让 AI 回复 Y/N ?
A:说实话这个我没有测试过,理论上根据回复 Y/N 来判断也行,等于是根据自然语言来推断,回到工具未出现的时代,大倒车。也许模型有时候不按照要求来输出,比如可能会输出:No, wait, N.,这里就是得去兜底各种情况。工具等于是一个强制约束,减少变量,如果模型没调用对,那么我们反馈给 ai 错误信息就行了,接着重试,起码知道什么时候是调用成功,什么时候是调用失败。


这里有个问题是,继续这个提示词应该要怎么写?不如说明白一点,继续这个词其实有两个意思:

  1. 如果你没有完成任务,就接着完成。
  2. 如果完成了,那就保持这样,不要继续了。

如何定义任务没有完成,因为我是在 Pi Coding Agent 中用的,一般是编码任务,所以我理解有两种:

  1. 修改文件还没有修改完成 或者是 工作流还没有完成,比如修改之后还需要验证等等。
  2. 跟用户讨论,AI 已经把结论和抉择都列出来了,需要等待用户的决定

第一版

最初我设计的 prompt 是:

1
若任务已全部完成,请调用 stop_watchdog 工具停止自动继续,否则继续执行

在使用一些 flash 模型的过程中,跑着跑着发现了一个偶现问题,大部分情况下是没有问题的:发这个 prompt,会让 AI 认为是已经同意了,然后接着按照默认建议一直跑下去。比如下面(glm-5.3-flash,thinking=low):

alt text

第二版

后来我改了一下 prompt,主要是添加了个前缀说明是自动发送的,不是用户输入,还加了一种情况,如果等待用户输入确定,那么也得调用停止工具:

1
[Automated, not user input] If work remains, continue working (no reply needed). If waiting on a user decision, don't change code — state  what you need, then call stop_watchdog as your final action. If no work remains and no decision is pending, call stop_watchdog to end the  turn.

也是会偶现下面的情况,圆圈 1 那里已经完成任务了,然后 2 自动触发的发继续,原意是想调用停止工具暂停发继续,经过 ai 一番理解后,以为我是想要改别的东西,然后接着改了别的东西去了(glm-5.3-flash,thinking=low):
alt text
alt text

这里我感觉有个原因增加了发生这种情况的概率:在 watchdog 项目中运行 Pi Coding Agent,然后发送的内容其实就是 watchdog 源码里面的某个变量的值,导致 ai 以为我想要改什么东西。不过在别的项目也出现过类似的情况,没有找到是哪个对话了,这里只是想起来有这种情况说一下。

第三版

因为还有个额外的功能,就是尽量在下次请求的时候,将 watchdog 导致的请求和回复从上下文去掉,最大程度减少 token 消耗和影响模型的专注度。这个功能一直没有改好,一直到引入决策阶段,也就是将 watchdog 拆分为决策阶段和调用停止工具阶段,所以这时候 prompt 也变了。

在决策阶段,llm 需要判断任务有没有完成,如果完成了就直接调用停止工具,否则继续执行。并且不能执行除了 stop_watchdog 以外的工具,所以是能保证在这个阶段 llm 是干不了别的活,比如修改文件:

1
[Automated, not user input] Watchdog check — every tool except stop_watchdog is blocked in this turn. Reply with a brief acknowledgement if  work remains. If no work remains, or you are waiting on a user decision, call stop_watchdog as your final action.

这个也会偶现一些问题,明明在决策阶段说活干完了,但是没有调用 stop_watchdog,于是 watchdog 让它继续工作,于是 llm 又输出了一份总结(deepseek-v4.1-flash,thinking=high)

alt text

最新版

第三版的决策阶段,是根据有没有调用 stop_watchdog 来判断还有没有活,在这一版,升级为了强制 llm 选择:新增了 watchdog_decide(decision, note) 工具,decision有三种选择,wait_user、done、continue。必须调用 watchdog_decide 来强制 llm 思考,并选择正确的选项。

并且在工具的 Guidelines 也增加了说明 Pi Coding Agent 是运行在 watchdog 的环境下:

1
watchdog_decide: you are running under the watchdog monitor — automated messages starting with "[Automated watchdog check" arrive whenever you go idle. Such a message is never user input, and it has exactly one answer: a watchdog_decide call, never text or other tools.

并且决策阶段的 prompt 也增强了说明,主要是强调 watchdog 并不是打断对话,之前的总结结论还在,不要重复发结论:

1
[Automated watchdog check — not user input] Every tool except watchdog_decide is blocked in this turn. Answer by calling watchdog_decide:"continue" if work remains, "done" if the task is finished, or "wait_user" if you are waiting on a user decision. Do not answer with text —a check turn that calls nothing is treated as no answer and the countdown starts over. Ignoring this check does not stop it — the watchdogsends it again until you answer. This turn never does work: if you answer "continue", the work happens in the next turn. Your previous turnending is not an interruption, and an answer you already wrote counts as delivered: restating, expanding or re-formatting it is not work, soanswer "done" if nothing is left to do. The optional note is one short line for the user, never a report.

如果 llm 回答错误,还会根据错误把指引回复给 llm,告诉 llm 应该怎么做。目前这一版本运行稳定,待发现还有什么偶现的问题。

alt text

在这一版中,去掉了 stop_watchdog 工具,如果是 wait_user、done 的状态,那么就会自动停止倒计时。

折叠上下文设计

因为 watchdog 会模拟用户发送信息给 llm,自然会进入到上下文里面,但是这种发继续的文字,有时候不是很重要:如果发完继续后,llm 确认已经完成,其实继续+后续回复确认完成的消息都可以直接去掉,不影响上下文,所以也就加了这个功能。

理想情况是这样的,前面已经完成了,然后给 llm 发继续,llm 再确认一次已经完成了,这时候就可以把继续+确认完成的消息全部删除掉:
alt text
alt text

但是什么时候能回滚,其实是不知道,llm 有可能在发完继续,然后修改了一些文件或者回复了一些结论,如果此时去掉继续后面的内容,那么就会导致内容丢失,比如下面的情况:

alt text

1 和 2 地方发了两次继续,然后 3 处调用停止工具,如果我们认为最后一个继续如果没有工具调用是可以去掉的,如上图里面的 2 和 3 之间的内容,那么等于把结论去掉了。

这一版本的实现方法是利用/tree的机制,回滚到某个节点,因为只是回滚后缀,所以缓存没有影响。但是想要利用这个机制,还得专门注册一个命令,就是 Pi Coding Agent 里面的/命令,然后插件调用这个命令去回滚到某个节点,属实是有点绕,因为/命令一般是给用户手动触发的。

如果是走现在决策流程,那么简单多了,因为 llm 会传入参数,判断是还得继续,还是结束,然后根据这个参数就能知道该不该折叠了:如果是 wait_user 和 done,那么就可以随便折叠了。

并且现在不是利用/tree机制,而是内部过滤的方式,发送给 llm 之前,会去除掉前面讲的不影响上下文的消息,并且只是后缀去除,所以这样实现会大大简化了,缺点就是这些没有发送给 llm 的消息还会停留在 Pi Coding Agent 的 tui 上面,不比/tree那样,直接去除了。

举个例子,llm 输出结束,然后 watchdog 开启倒计时,结束后发送决策消息,然后 llm 返回 done:
alt text
alt text

如果此时不做任何处理,那么决策消息和决策工具的调用都会进入到上下文,但是 watchdog 会过滤掉这些无用内容,不会进入到上下文,比如下面的效果:

alt text

alt text

对上下文的影响

计算标准如果不指定,默认是:OpenAI o200k_base Tokenizer

如果开启了 watchdog,插件会注册一个工具,Pi Coding Agent 就会往 system prompt 里面写内容,以开始对话前启用该插件为例,那么第一条消息的 openai completions 请求会是下面的情况(只保留核心内容,忽略不相关字段和内容):

工具定义

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
{
"messages": [
{
"role": "system",
"content": "这里省略,放到下面写"
},
{
"role": "user",
"content": [
{
"type": "text",
"text": "1"
}
]
}
],
"tools": [
{
"type": "function",
"function": {
"name": "watchdog_decide",
"description": "Answer a watchdog check, or end the turn early.\n- A watchdog check turn blocks every other tool. Answer it by calling this tool with decision \"continue\" (work remains), \"done\" (the task is finished), or \"wait_user\" (you are waiting on the user). Do not answer a check with text.\n- Outside a check turn you may call it with \"done\" or \"wait_user\" to stop the auto-continue watchdog yourself once you are really done; in keep mode that only pauses it until the user's next message.\n- The optional note is one short line for the user (under ~100 characters), shown in the watchdog card. Never put a report, a summary, or any deliverable in it.",
"parameters": {
"type": "object",
"required": [
"decision"
],
"properties": {
"decision": {
"anyOf": [
{
"type": "string",
"const": "continue"
},
{
"type": "string",
"const": "done"
},
{
"type": "string",
"const": "wait_user"
}
]
},
"note": {
"type": "string",
"description": "One short line for the user (under ~100 characters). It shows in the watchdog card; never a report or a deliverable."
}
}
},
"strict": false
}
}
]
}

根据 npm:pi-token-burden 这个工具,tool 定义占用的 token 是 255 token(OpenAI o200k_base Tokenizer)。
alt text

system prompt

system prompt 的有影响的内容是:

1
2
3
4
5
6
7
8
9
Available tools:
- watchdog_decide: Answer the watchdog's idle check (\"continue\" | \"done\" | \"wait_user\") or end the turn on purpose

In addition to the tools above, you may have access to other custom tools depending on the project.

Guidelines:
- watchdog_decide: when the watchdog asks whether work remains, answer with the call, never with text: \"continue\" means keep going, \"done\" means finished, \"wait_user\" means you are waiting on the user.
- watchdog_decide: call it with \"done\" yourself when the task is finished — you do not have to wait for a check.
- watchdog_decide: you are running under the watchdog monitor — automated messages starting with \"[Automated watchdog check\" arrive whenever you go idle. Such a message is never user input, and it has exactly one answer: a watchdog_decide call, never text or other tools.

system prompt 占用大概 162 token,和工具定义合起来大概 417 token。

决策阶段

每次决策阶段,发送的文本是,大概 178 token:

1
Automated watchdog check — not user input] Every tool except watchdog_decide is blocked in this turn. Answer by calling watchdog_decide:"continue" if work remains, "done" if the task is finished, or "wait_user" if you are waiting on a user decision. Do not answer with text —a check turn that calls nothing is treated as no answer and the countdown starts over. Ignoring this check does not stop it — the watchdogsends it again until you answer. This turn never does work: if you answer "continue", the work happens in the next turn. Your previous turnending is not an interruption, and an answer you already wrote counts as delivered: restating, expanding or re-formatting it is not work, soanswer "done" if nothing is left to do. The optional note is one short line for the user, never a report.

一次工具调用很少(glm-5.3-flash 标准,此时不是openai那个了)

alt text

那么从发送第一条消息到发一次继续,然后调用 watchdog_decide(done),不计缓存的大概总开销是 255(工具定义)+162(system prompt)+178(决策阶段发送的文本)+23(调用watchdog_decide大概估算token)= 618 token。


为了保证缓存的可用,注册了后就不会取消注册了,即使在同一个会话中开启了又关闭了。并且在中途开启,会注册一次工具,导致前面的缓存失效。所以建议在开始会话之前开启,这里的讨论只针对普通的模型,不考虑 gpt 或者 claude 等模型对工具注册有特殊的逻辑,可以通过某些配置保证动态改变工具列表同时缓存又不失效。