Coding Agent 不是陪聊机器人,它要能读代码、定位 Bug、改源码、跑测试,并依据真实退出码判断任务是否完成。原文用约 300 行纯 Python,不依赖 LangChain、CrewAI、AutoGen 等重框架,实现了一个可接入 OpenAI 兼容接口的生产级 Coding Agent。其核心不是堆提示词,而是“执行状态机 + 确定性工具集 + Token 预算”。
企业自研 Agent 常见三个翻车点:一是伪修复,Agent 声称修好但没有跑测试;二是破坏性全文件覆写,只改一行却重写整文件;三是上下文爆炸,大文件和长命令输出把 Prompt 撑满,导致首字延迟飙升甚至服务崩溃。原文的实现就是针对这三点收口。
一、三位一体架构
原文把 Claude Code、Hermes Agent 这类 Coding Agent 的执行流收敛为“三位一体”:规划、意图与决策中枢;Runtime Tool Engine,包括 Bash、Read、Patch、Rg;Memory & Context Budget,负责 Token 预算、修剪与状态回写。
对应到工程约束有三条:
1. 工具必须确定性。文件读取要带行号和截断保护;文件修改要走 Patch 模式,精准锚定替换,禁止整文件覆写;命令执行要捕获真实退出码 Exit Code。
2. 执行反馈必须真实回流。Agent 改完代码后,要通过命令行运行真实测试,例如 pytest 或 npm test;只有测试输出 PASS 且退出码为 0,才允许判定任务完成。
3. 上下文必须严格控盘。大文件读取和长命令输出要设置 Hard Cap,防止上下文瞬间被撑爆。
二、核心工具集实现
run_bash 负责执行 Shell 命令。它有一个危险命令黑名单,包含 rm -rf /、:(){ :|:& };:、mkfs、dd if=/dev。命中后直接返回安全拦截错误,退出码为 -1。
- def tool_run_bash(command: str, timeout: int = 60) -> str:
- dangerous_patterns = ['rm -rf /', ':(){ :|:& };:', 'mkfs', 'dd if=/dev']
- if any(p in command for p in dangerous_patterns):
- return json.dumps({'error': 'Security Alert: Command blocked by safety guardrails', 'exit_code': -1})
- try:
- proc = subprocess.run(
- command,
- shell=True,
- capture_output=True,
- text=True,
- encoding='utf-8',
- errors='replace',
- timeout=timeout,
- cwd=os.getcwd()
- )
- output = proc.stdout if proc.stdout else proc.stderr
- if len(output) > 8000:
- output = output[:4000] + '... [OUTPUT TRUNCATED DUE TO LENGTH] ...' + output[-4000:]
- return json.dumps({'output': output, 'exit_code': proc.returncode})
- except subprocess.TimeoutExpired:
- return json.dumps({'error': f'Execution timed out after {timeout} seconds', 'exit_code': 124})
- except Exception as e:
- return json.dumps({'error': str(e), 'exit_code': -1})
复制代码
这里有两个兼容性细节很重要。Windows 默认编码可能是 GBK,直接解码中文输出容易报错,所以代码强制 encoding='utf-8' 并使用 errors='replace',非法字节替换为占位符,保证不崩溃。输出超过 8000 字符时,保留头部 4000 字符和尾部 4000 字符,中间用截断标记连接,既能保留初始命令信息,也能保留最终报错栈和退出信息。
read_file 负责按行号读取文件。默认 offset=1、limit=300,返回 total_lines、showing_lines 和 content。每行格式为“行号| 内容”,便于模型定位。
- def tool_read_file(path: str, offset: int = 1, limit: int = 300) -> str:
- if not os.path.exists(path):
- return json.dumps({'error': f"File '{path}' not found."})
- with open(path, 'r', encoding='utf-8', errors='replace') as f:
- lines = f.readlines()
- total_lines = len(lines)
- start_idx = max(0, offset - 1)
- end_idx = min(total_lines, start_idx + limit)
- numbered_lines = [f'{i+1:4d}| {lines[i]}' for i in range(start_idx, end_idx)]
- return json.dumps({
- 'total_lines': total_lines,
- 'showing_lines': f'{start_idx+1}-{end_idx}',
- 'content': ''.join(numbered_lines)
- })
复制代码
patch_file 是精准局部替换工具。它读取原文件,统计 old_str 出现次数。次数为 0 说明目标没找到;次数大于 1 说明匹配不唯一,直接拒绝,并要求模型补充更多上下文。只有唯一匹配时才执行 replace(old_str, new_str, 1)。
- def tool_patch_file(path: str, old_str: str, new_str: str) -> str:
- if not os.path.exists(path):
- return json.dumps({'error': f"File '{path}' does not exist."})
- with open(path, 'r', encoding='utf-8') as f:
- content = f.read()
- count = content.count(old_str)
- if count == 0:
- return json.dumps({'error': 'Target old_str not found in file. Check line numbers and indentation.'})
- if count > 1:
- return json.dumps({'error': f'Target old_str matched {count} times. Please include more surrounding context lines for uniqueness.'})
- new_content = content.replace(old_str, new_str, 1)
- with open(path, 'w', encoding='utf-8') as f:
- f.write(new_content)
- return json.dumps({'status': 'success', 'message': f"Successfully patched '{path}'."})
复制代码
search_files 基于 os.walk 遍历目录,忽略 .git、node_modules、__pycache__、.venv,按文件名子串匹配,返回相对路径,最多 50 条。原文还定义了 TOOLS_SCHEMA,把 run_bash、read_file、patch_file、search_files 四个工具以 OpenAI function calling 格式暴露给模型,并通过 TOOL_MAP 做名称到函数的映射。
三、LLM 调用与 Agent 循环
call_llm 使用标准库 urllib.request 调用兼容 OpenAI 的 /chat/completions 接口。请求头带 Authorization: Bearer api_key,payload 中传入 model、messages、tools、tool_choice='auto',temperature 为 0.1。HTTP 超时为 120 秒,HTTPError 会读取错误体并抛出 RuntimeError。
CodingAgent 的 system_prompt 给出四条准则:先检查再修改,用 search_files 和 read_file 理解代码库;精准编辑,始终用 patch_file 并带足够唯一上下文,绝不整文件覆写;真实验证,改动后通过 run_bash 执行相关单元测试;持续迭代直到测试通过 exit_code=0,验证后才给出最终答案。
核心 run 方法维护 messages 数组,初始加入 system 和 user 消息。每一轮调用 call_llm,取 choices[0].message 并追加。如果 message 没有 tool_calls,说明模型给出自然语言回答,流程结束。如果有 tool_calls,就逐个解析函数名和 JSON 参数,执行真实工具,再把结果以 role='tool' 和 tool_call_id 追加回 messages。max_turns 默认 15,超过后返回“Task stopped: Max iterations reached without completion.”
上下文预算控制发生在循环内部:当 messages 长度超过 25 时,执行 messages = [messages[0]] + messages[-20:],保留首条 system 指令和最近 20 条消息。这样做的目的,是避免早期中间调试记录无限堆积,把 Token 开销和首字延迟稳定在可控区间。
- class CodingAgent:
- def __init__(self, api_key: str, base_url: str = 'https://api.openai.com/v1', model: str = 'gpt-4o'):
- self.api_key = api_key
- self.base_url = base_url
- self.model = model
- self.system_prompt = (
- 'You are an expert autonomous Software Engineering Agent.\n'
- 'Working Directory: ' + os.getcwd() + '\n'
- 'Guidelines:\n'
- '1. Inspect before modifying.\n'
- '2. Precision edits with patch_file. NEVER overwrite files completely.\n'
- '3. Ground truth verification with run_bash after changes.\n'
- '4. Keep working iteratively until tests pass exit_code=0.'
- )
- def run(self, user_goal: str, max_turns: int = 15) -> str:
- messages = [
- {'role': 'system', 'content': self.system_prompt},
- {'role': 'user', 'content': user_goal}
- ]
- for turn in range(1, max_turns + 1):
- if len(messages) > 25:
- messages = [messages[0]] + messages[-20:]
- resp = call_llm(messages, self.api_key, self.base_url, self.model)
- message = resp['choices'][0]['message']
- messages.append(message)
- if not message.get('tool_calls'):
- content = message.get('content', '')
- print(f'\n[Agent Finished]:\n{content}')
- return content
- for tool_call in message['tool_calls']:
- tool_name = tool_call['function']['name']
- tool_args = json.loads(tool_call['function']['arguments'])
- call_id = tool_call['id']
- fn = TOOL_MAP.get(tool_name)
- if fn:
- tool_result = fn(**tool_args)
- else:
- tool_result = json.dumps({'error': f"Tool '{tool_name}' not implemented."})
- preview = tool_result[:150] + '...' if len(tool_result) > 150 else tool_result
- messages.append({
- 'role': 'tool',
- 'tool_call_id': call_id,
- 'content': tool_result
- })
- return 'Task stopped: Max iterations reached without completion.'
复制代码
CLI 入口读取 sys.argv,如果参数不足则提示 Usage: python coding_agent.py '<task_goal>'。API Key 从 OPENAI_API_KEY 读取,base_url 从 OPENAI_BASE_URL 读取,模型从 AGENT_MODEL 读取,默认 gpt-4o。因此它可以接入任何兼容 OpenAI、DeepSeek、Claude 协议的 API。
四、为什么坚持 patch_file,而不是 write_file
如果文件有 500 行,模型只需改第 20 行,整文件覆写会带来三个问题:消耗大量输出 Token,速度慢;模型可能在长文件中间省略代码,导致源码损坏;还可能冲掉 Git 中其他人的并发修改。patch_file 强制模型提供 old_str 和 new_str,并严格校验 old_str 匹配次数必须等于 1。匹配多次就直接拒绝,让模型补充更多上下文。这样把修改操作变成确定性的、可验证的原子替换。
命令输出截断也是同样思路。pytest 或 npm run build 可能输出数万行依赖报错或堆栈日志。全部灌入 Prompt,一次工具调用就能打满上下文窗口。保留头尾各 4000 字符,信息密度更高,也保证后续推理不崩。
五、实战:自动排查并修复 calculator.py
原文构造了一个有 Bug 的业务文件 calculator.py:
- def calculate_tax(income: float, rate: float) -> float:
- if income < 0:
- return income * rate
- return income * (rate / 10)
复制代码
配套单元测试 test_calculator.py 要求 calculate_tax(10000, 0.2) 等于 2000.0,并要求负数收入抛出 ValueError:
- import pytest
- from calculator import calculate_tax
- def test_positive_tax():
- assert calculate_tax(10000, 0.2) == 2000.0
- def test_negative_income():
- with pytest.raises(ValueError):
- calculate_tax(-100, 0.2)
复制代码
运行命令为:
- python coding_agent.py "运行 pytest 查看测试失败原因,并修复 calculator.py 中的错误,直到所有测试通过"
复制代码
原文给出的真实交互环境是中文 Windows 11、Python 3.12、GLM-5.2-CL,并采用 OpenAI 兼容协议。控制台轨迹开头如下:
- [Agent Started] Goal: 运行 pytest 查看测试失败原因,并修复 calculator.py 中的错误,直到所有测试通过
- ======================================================================
- [Turn 1/15] Thinking...
- Tool Call: search_files({'pattern': 'calculator'})
- Result: {'matched_files': ['calculator.py', 'test_calculator.py']}
复制代码
原文日志在 search_files 处截断。按照这套代码的循环逻辑,后续会继续读取文件、用 patch_file 修改,再执行 pytest,直到退出码为 0。这里的关键不是控制台输出多漂亮,而是每一步工具调用都有真实结果回流:文件读取带行号,修改必须唯一锚定,测试必须真实执行。
六、这套 300 行实现的工程意义
它把 LLM 从“聊天角色”降为“决策器”,把文件系统、Shell、测试命令作为真值来源。对自研 Coding Agent 或 DevOps Agent 的团队来说,优先级应该是先定义确定性工具契约,再写执行状态机,最后才是接框架。框架能帮写 Demo,但协议、工具和状态机才是能上生产的骨架。
适用场景也很明确:代码库内 Bug 修复、测试失败排查、局部重构、命令行任务自动化、接入兼容 OpenAI 协议的模型服务。需要提醒的是,这套原型仍然只有 300 行,安全黑名单和输出截断只是基础护栏;真正企业级落地还要在权限、审计、并发、沙箱和更细粒度的上下文管理上继续补强。原文在结尾列出“从 300 行原型到企业级 Agent 的演进路线”这一方向,但没有展开更多代码细节;就已给出的实现而言,它已经覆盖了 Coding Agent 最核心的工具闭环、Patch 修改与 pytest 验证路径。 |