Architecture¶
The three packages¶
flowchart TB
subgraph your["your machine"]
CLI["<b>resumix</b> (client)<br/>single executable"]
FILES["candidate_profile.json<br/>candidate_data.json<br/>candidate_preferences.md<br/>resume.tex.jinja · prompts · images"]
OUT["cv/ · discarded/ · error/ · working/<br/>applications.xlsx"]
FILES --> CLI
CLI --> OUT
end
subgraph wire["contracts"]
CT["resumix_contracts<br/>Envelope · CVStatus · RenderedCV<br/>JDAnalysis · JDDetection · RequestLog<br/>static_jd_guess"]
end
subgraph host["wherever you run it"]
API["<b>resumix-server</b><br/>FastAPI + uvicorn"]
PIPE["pipeline<br/>jd_validator · cv_generator · cv_validator<br/>cv_renderer · letter_generator"]
RES["resources<br/>5 prompts · resume.tex.jinja<br/>models.toml"]
STORE["RESUMIX_WORK_DIR<br/>one directory per request"]
API --> PIPE
RES --> PIPE
PIPE --> STORE
end
PROV["OpenAI-compatible<br/>/chat/completions"]
TEX["pdflatex"]
CLI -->|"multipart/form-data"| API
API -->|"JSON envelope"| CLI
CLI -.imports.-> CT
API -.imports.-> CT
PIPE --> PROV
PIPE --> TEX
client → contracts ← server. Neither half imports the other, and
tests/test_architecture.py fails the build if one ever does. The contracts
package depends on pydantic and nothing else: it is loaded both by a web
server and by a frozen executable, and has to stay cheap in both.
| Package | Holds | Ships as |
|---|---|---|
client/ |
Your profile, your contact data, your template and images, the output tree, the spreadsheet | A PyInstaller executable, ~1 file |
server/ |
The provider API key, the five prompts, the LaTeX template, the LaTeX toolchain | A container image, ~780 MB |
contracts/ |
The response envelope, the payload models, the free JD pre-check | A pydantic-only package, vendored into both |
What crosses the wire¶
Everything personal arrives with the request and dies with it. The server has no database, no accounts and no seeded state; the only thing it keeps is one directory per request:
$RESUMIX_WORK_DIR/<request_id>/
status.json a CV job's state — only CV jobs have one
result.json the finished RenderedCV — only once the job is done
log.json every line the request produced, model calls included
work/ the scratch directory, deleted when the work ends
Writes are atomic (temp file + os.replace), because a poll may read a file
while the worker is rewriting it. The newest 200 directories are kept; the
rest are dropped, except one still marked running.
Inside the server¶
flowchart TB
REQ["request"] --> MW["RequestContextMiddleware<br/>assigns request_id, binds the Run"]
MW --> AUTH{"RESUMIX_API_TOKEN set?"}
AUTH -->|"no match"| E401["401 unauthorized"]
AUTH -->|"ok / not set"| PARTS["multipart parts<br/>size, UTF-8, JSON, ≤10 images"]
PARTS --> SLOT{"a free job slot?"}
SLOT -->|"no"| E429["429 too_many_jobs<br/>Retry-After: 30"]
SLOT -->|"yes"| KIND{"which endpoint?"}
KIND -->|"/v1/jd/*, /v1/letter, /v1/cv/render"| POOL["threadpool<br/>answer when done"]
KIND -->|"/v1/cv"| THREAD["worker thread<br/>answer 202 now"]
POOL --> ENV["Envelope"]
THREAD --> STATUS["status.json rewritten<br/>at every step"]
STATUS --> POLL["GET /v1/cv/{id}/status"]
ENV --> LOG["log.json saved"]
STATUS --> LOG
- One exception handler.
describe_failure()is the single place that decides what an exception produces on the wire, so a failure body has the same shape as a success: arequest_id,ok: falseand one line naming the kind and the stage. See error model. - One semaphore.
RESUMIX_MAX_CONCURRENT_JOBS(10) bounds requests in flight. A full server refuses immediately rather than queueing a caller for minutes. - The pipeline is a library. Everything under
pipeline/takes its inputs in memory and returns its outputs; it reads no configuration, resolves no paths and writes nothing outside the scratch directory it is handed. That is what makes it safe to run per request. - The LaTeX subprocess is sandboxed. Shell escape off, reads and writes
confined to the scratch directory, no stdin, a timeout, its own
HOMEandTEXMFVAR. A request may supply both the template and a whole.tex, so both are treated as hostile input.
The five models¶
One provider, one API key, one endpoint — declared once in the [provider]
table of resources/models.toml. Five roles call it; a role that sets
use_alternate_provider = N (N ≥ 2) calls provider N instead — the
[provider] variable names with N appended (MODEL_API_KEY2 /
MODEL_BASE_URL2 for 2):
| Role | Used for | Shipped as |
|---|---|---|
detect |
JD detection: a one-word YES/NO on text that passed the free checks | qwen3.8-flash, thinking off, 100 max tokens, no JSON mode |
summary |
JD analysis, cover letters. The only role allowed server-side web search. | qwen3.8-flash, thinking on, low effort |
cv |
Writing the CV | qwen3.8-max, thinking on, 6500-token budget, strict JSON schema |
review |
Reviewing each CV against the master profile; answers in plain text, one violation per line | qwen3.8-max, thinking on, temperature 0, no JSON mode |
highlight |
The **bold** keyword pass over validated CV JSON |
qwen3.8-flash, thinking off |
Provider differences are declared as capabilities (web_search, thinking,
structured_output), never branched on by name. Each role's model,
temperature, thinking, reasoning_effort, thinking_budget,
structured_output and use_alternate_provider can be overridden with a
RESUMIX_<ROLE>_<FIELD> environment variable without editing the file. An
empty RESUMIX_<ROLE>_THINKING_BUDGET= unsets a budget declared here, which
is how RESUMIX_<ROLE>_REASONING_EFFORT wins back the request — a declared
budget otherwise always overrides reasoning effort.
Every model reply is re-validated: the JSON schema goes into the prompt and
into response_format, and the reply is parsed by pydantic. A weak
response_format costs a retry, never correctness.
Asynchrony, and why¶
Writing a CV takes minutes. POST /v1/cv answers 202 with a job id and
hands the work to a plain thread that outlives the response; the thread
reports itself by rewriting status.json, and a poll reads a file. Nothing
is held open for minutes, and a client that walks away costs nothing but its
slot.
The consequences are worth knowing before you deploy:
- A job id only means something to the instance holding its directory. Run
one instance per
RESUMIX_WORK_DIR, or give several a shared root and route by request id. WEB_CONCURRENCY=1. The slots and the workers are per-process, and a starting process marks every job still markedrunningas failed — correct for its own orphans, wrong for a sibling's live jobs.- A job cannot be cancelled yet. An abandoned one holds a slot until its
own budget runs out (
RESUMIX_REQUEST_BUDGET_SECONDS, 20 minutes).PLANNED-FEATURES.mddescribes theDELETE /v1/cv/{id}that fixes it.
Observability¶
Every response carries a request_id, and so does the X-Request-Id header.
That id is the handle for GET /logs/{request_id}, which returns what the
server did — one line per model call with its duration and token counts.
Nothing else rides on a reply: the commentary is fetched separately, so the
common case pays nothing for it and the debugging case gets all of it. The
client folds it into log.log next to each CV — always on failure, and for
every call with --debug.