Post-processing and quantities of interest
A campaign is only useful if you can get numbers out of it. ModelManager lets you attach a post_processor callback to run: a function invoked once per successfully completed simulation, whose return value is stored in a per-project sink and read back later as a table.
Use it for anything you want computed while the output is still on disk — final cell counts, time to extinction, fit residuals, summary statistics — and for side effects such as writing your own reduced-output files before a backend prunes the raw output.
Where your callback runs
After each simulation completes, ModelManager runs three steps in order:
- the backend's
postSimulationProcessinghook — simulator-specific, non-destructive work (e.g. standardizing output); - your
post_processor— invoked once per successfully completed simulation; - the backend's
postSimulationCleanuphook — simulator-specific, destructive cleanup/pruning.
Because your callback runs before cleanup, it always sees the intact (but processed) output folder:
run(sampling; post_processor = QoI("counts", sim -> (; final_count = countCells(simulationID(sim)))))The callback signature
The callback receives a Simulation — the same thing a QoI's compute receives, so one measurement function works in the sink, in sensitivity analysis, and in calibration alike. Use simulationID and pathToOutputFolder rather than reaching into its fields. The hook only fires for simulations that succeeded, and the owning monad is available as only(monadIDs(sim)) if you need it. Reading the actual simulation output into usable data is the backend's job — expect your simulator package to provide loaders keyed by simulationID.
What to return
The return value decides storage:
nothing→ nothing is stored (pure side effects — compute, write files, or clean up however you like).- a
NamedTupleorAbstractDictofname => scalar(where each value is aReal,Bool, orString) → one row (keyed bysimulation_id) is upserted into the project's post-processing sink atdata/outputs/postprocessing.db. Columns are added on demand, so a quantity not computed for a given simulation reads back asmissing; re-running overwrites that simulation's row. Anything else (including a non-scalar value such as a vector) raises anArgumentError.
A returned key does not become a bare column: it is prefixed with the name of the QoI that produced it, so QoI("counts", …) returning (; tumor = 3) writes counts.tumor. That is what lets two measurements both report a tumor without landing in one column, and it matches how sensitivity analysis labels the same spread.
A bare anonymous function therefore cannot write to the sink. Its name is derived from a gensym (anon_9) that changes between sessions, so the same script would write a fresh, half-empty set of columns on every run. Wrap it — QoI("counts", sim -> …) — or pass a named function, whose name is stable. A callback returning nothing is unaffected, since it stores nothing to name.
Storing nothing (side effects only)
If you only want side effects — writing your own output file, deleting data, logging — return nothing explicitly. This matters: a callback's value is its last expression, so a block that ends with a computation would store that value by accident. End with return nothing (or a bare nothing) to store nothing:
run(sampling; post_processor = function (sp)
writeCustomSummary(pathToOutputFolder(sp)) # your own file, in the sim's output folder
return nothing # <-- required; without it the summary would be stored
end)Storing a NamedTuple
The most concise form. Each field becomes a sink column:
run(sampling; post_processor = QoI("cells", sp -> (; final_count = countCells(simulationID(sp)),
mean_speed = meanCellSpeed(simulationID(sp)))))giving columns cells.final_count and cells.mean_speed.
Storing a Dict
Useful when the names are computed or come from data. Keys become the second half of each column name:
run(sampling; post_processor = QoI("viability", function (sp)
cells = loadCells(simulationID(sp)) # a loader from your simulator package
return Dict("n_alive" => countAlive(cells),
"n_dead" => countDead(cells))
end))giving columns viability.n_alive and viability.n_dead.
(countCells, meanCellSpeed, loadCells, … are stand-ins for whatever loaders your simulator package provides — see its documentation.)
Reading the quantities back
Read the collected quantities with postProcessingTable (or printPostProcessingTable); the result is keyed by :SimID:
postProcessingTable(sampling) # one row per simulation with stored quantitiesTo see the quantities alongside each simulation's parameters, pass post_processing=true to simulationsTable — it appends one column per quantity (missing where a quantity was not computed):
simulationsTable(sampling; post_processing=true)Quantities and tags are both keyed by simulation_id, so a recovery query can join them:
ids = findSimulationIDs(tags = ("project" => "immune-escape",), status = "Completed")
innerjoin(simulationsTable(ids; tags = true), postProcessingTable(ids), on = :SimID)Lifecycle and error handling
The sink stays consistent with the central database: deleting simulations (see Managing data) removes their sink rows, and resetDatabase removes the sink entirely.
If your post_processor (or a simulator hook) throws, run fails fast — it rethrows a clear error naming the stage (post_processor vs. a simulator hook) and the simulation, with the original stacktrace. It never hangs or silently swallows the exception, so a typo or a bad assumption in a callback surfaces immediately rather than parking a long HPC campaign.
For the storage layer itself (schema, postProcessingDBPath) see The database. For the runner API around the callback (SimulationProcess, SimulationSpec), see the Runner reference.