backend: serve the writing domain to agents as an MCP server

The paper is written by a model now. A browser is the wrong client for that:
the work is "generate the paper, then put it in", and doing it through a form
means a person retyping what a model already produced. So the same domain is
served over the Model Context Protocol, which Claude Code, Codex and the
DeepSeek Harness all speak.

It is a third front door, not a second implementation. Every tool is three
lines around an `app.crud` call and validates through `app.schemas`, exactly
as the REST routes do, so a rule fixed in the CRUD layer is fixed on both
surfaces and a paper written by an agent is indistinguishable from one written
by hand. What `app/mcp/` adds is only what a model needs and a browser does
not:

- 29 tools, prefixed `paper_` `paragraph_` `sentence_` `template_` `field_`,
  because a model picks a tool out of a list by name rather than by reading 29
  descriptions;
- results as compact `None`-free JSON, since a tool result is paid for in
  context tokens and `PaperRead.model_dump()` carries four counts and two
  timestamps into every list row;
- paragraphs addressed by **heading** as well as by position. Storage is
  correct as it stands — a sentence remembers the position it sits at, which is
  what makes a template switch non-destructive — but nobody writing
  "1. Introduction" knows the template places it at `sort = 20`. The server
  translates, and refuses with the real heading list when it cannot, so a model
  that guessed wrong corrects itself in one retry;
- `paper_write` and `paper_write_text`: one intention, one call. The latter
  finds its own sections from Markdown headings or from lines that name a
  template heading, and reports every heading it could not place instead of
  writing half a paper;
- `sentence_search` across papers, for consistency rather than retrieval — a
  paper that says 洪水损失 should not be joined by one that says GUL;
- `paper_delete` refuses once, naming what would go with it. A cascading delete
  has no undo in a tool call.

Two transports, one build. `stdio` is what a client spawns — so nothing in the
process may print to stdout, and diagnostics go to stderr. `streamable-http` is
what a client on another machine connects to, optionally behind a bearer token;
binding a non-loopback address disables the SDK's DNS-rebinding allow-list,
because a LAN client sends whatever Host it knows the server by.

Tools register with `structured_output=False` on purpose: inferred from a
`-> str` annotation the SDK publishes a `{"result": ...}` envelope and sends
the JSON twice, once as `structuredContent` and once as text, and clients that
read only one of the two then disagree about what came back.

`scripts/smoke_mcp.py` drives the whole loop through a real MCP client — the
child process and JSON-RPC over stdin/stdout a client actually uses — and runs
unchanged against a running HTTP server via `--url`. 47 checks pass on both
transports; the REST suite still passes its 45.
This commit is contained in:
2026-09-19 00:01:43 +08:00
parent 20cd63f9f9
commit 2aff867641
17 changed files with 3276 additions and 0 deletions
+40
View File
@@ -666,3 +666,43 @@ def template_paper_counts(
with the count that explains why.
"""
return _paper_counts_by_template(db, template_ids)
# --- search ------------------------------------------------------------------
def search_sentences(
db: Session,
*,
keyword: str,
paper_id: int | None = None,
limit: int = 50,
) -> list[tuple[PaperSentence, str]]:
"""Find sentences whose text contains ``keyword``, newest paper first.
Returns ``(sentence, paper_title)`` pairs, the title read alongside the row
so the caller does not pay a lazy load per sentence.
The MCP server's ``sentence_search`` tool is the caller. What it serves is
consistency rather than retrieval: a paper that calls it 洪水损失 should
not be joined by one that calls it GUL, and reading whole documents to find
one phrase would not fit in a context window.
"""
pattern = like_pattern(keyword)
stmt = (
select(PaperSentence, Paper.title)
.join(Paper, Paper.id == PaperSentence.paper_id)
.where(PaperSentence.content.like(pattern, escape=LIKE_ESCAPE))
.order_by(
Paper.updated_at.desc(),
PaperSentence.paper_id,
PaperSentence.paper_template_filed_sort,
PaperSentence.sort,
PaperSentence.id,
)
.limit(limit)
)
if paper_id is not None:
stmt = stmt.where(PaperSentence.paper_id == paper_id)
return [(sentence, title) for sentence, title in db.execute(stmt).all()]