Files
paper-doc/backend/app/mcp/selfcheck.py
T
govin 2aff867641 backend: serve the writing domain to agents as an MCP server
The paper is written by a model now. A browser is the wrong client for that:
the work is "generate the paper, then put it in", and doing it through a form
means a person retyping what a model already produced. So the same domain is
served over the Model Context Protocol, which Claude Code, Codex and the
DeepSeek Harness all speak.

It is a third front door, not a second implementation. Every tool is three
lines around an `app.crud` call and validates through `app.schemas`, exactly
as the REST routes do, so a rule fixed in the CRUD layer is fixed on both
surfaces and a paper written by an agent is indistinguishable from one written
by hand. What `app/mcp/` adds is only what a model needs and a browser does
not:

- 29 tools, prefixed `paper_` `paragraph_` `sentence_` `template_` `field_`,
  because a model picks a tool out of a list by name rather than by reading 29
  descriptions;
- results as compact `None`-free JSON, since a tool result is paid for in
  context tokens and `PaperRead.model_dump()` carries four counts and two
  timestamps into every list row;
- paragraphs addressed by **heading** as well as by position. Storage is
  correct as it stands — a sentence remembers the position it sits at, which is
  what makes a template switch non-destructive — but nobody writing
  "1. Introduction" knows the template places it at `sort = 20`. The server
  translates, and refuses with the real heading list when it cannot, so a model
  that guessed wrong corrects itself in one retry;
- `paper_write` and `paper_write_text`: one intention, one call. The latter
  finds its own sections from Markdown headings or from lines that name a
  template heading, and reports every heading it could not place instead of
  writing half a paper;
- `sentence_search` across papers, for consistency rather than retrieval — a
  paper that says 洪水损失 should not be joined by one that says GUL;
- `paper_delete` refuses once, naming what would go with it. A cascading delete
  has no undo in a tool call.

Two transports, one build. `stdio` is what a client spawns — so nothing in the
process may print to stdout, and diagnostics go to stderr. `streamable-http` is
what a client on another machine connects to, optionally behind a bearer token;
binding a non-loopback address disables the SDK's DNS-rebinding allow-list,
because a LAN client sends whatever Host it knows the server by.

Tools register with `structured_output=False` on purpose: inferred from a
`-> str` annotation the SDK publishes a `{"result": ...}` envelope and sends
the JSON twice, once as `structuredContent` and once as text, and clients that
read only one of the two then disagree about what came back.

`scripts/smoke_mcp.py` drives the whole loop through a real MCP client — the
child process and JSON-RPC over stdin/stdout a client actually uses — and runs
unchanged against a running HTTP server via `--url`. 47 checks pass on both
transports; the REST suite still passes its 45.
2026-09-19 00:01:43 +08:00

65 lines
2.4 KiB
Python

"""Self-check: prove the server can start before a client has to find out.
Run with ``python -m app.mcp --check`` (or ``scripts/mcp_server.py --check``).
The failure this exists for is the quiet one. A stdio MCP client spawns the
server, the process starts, the handshake succeeds, and every tool call fails
with a database error the model reads as "the tool is broken". Connecting once
here, on demand, turns that into a message with a host and a port in it.
It is deliberately independent of the transports: it builds the same server
object the transports serve, so the tool count it prints is the tool count a
client will see.
"""
from __future__ import annotations
import sys
import anyio
from mcp.server.mcpserver import MCPServer
from sqlalchemy import func, select
from app.core.config import get_settings
from app.db.session import SessionLocal
from app.mcp.server import SERVER_VERSION
from app.models import Paper, PaperSentence, Template, TemplateFieldLibrary
def run_selfcheck(server: MCPServer) -> int:
"""Report the database, the tool surface, and where the tools come from."""
settings = get_settings()
print(f"paper-doc MCP {SERVER_VERSION} (name={server.name})")
print(f"database : {settings.safe_database_url}")
try:
with SessionLocal() as db:
counts = {
"paper": db.scalar(select(func.count(Paper.id))) or 0,
"template": db.scalar(select(func.count(Template.id))) or 0,
"paper_sentence": db.scalar(select(func.count(PaperSentence.id))) or 0,
"template_field_library": db.scalar(
select(func.count(TemplateFieldLibrary.id))
)
or 0,
}
except Exception as error: # noqa: BLE001 - the message is the point
print(f"database : 连接失败 — {error}", file=sys.stderr)
print(
"检查 backend/.env 的 DB_* 配置,或用 proxy_endpoint skill 确认网络可达。",
file=sys.stderr,
)
return 1
print("rows : " + ", ".join(f"{name}={count}" for name, count in counts.items()))
tools = anyio.run(server.list_tools)
groups: dict[str, int] = {}
for tool in tools:
groups[tool.name.split("_", 1)[0]] = groups.get(tool.name.split("_", 1)[0], 0) + 1
print(f"tools : {len(tools)} " + ", ".join(f"{k}={v}" for k, v in sorted(groups.items())))
for tool in tools:
print(f" - {tool.name}")
return 0