Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Scaling, memory, and hardware

Correctness narrows the feasible solver set; computational feasibility can still decide which member to use. Compare compilation, execution, peak memory, and accuracy on the hardware that will run the production model.

Sources of growth

A useful first inventory is:

AxisWhy it matters
State-grid productNumber of value-function cells
Action-grid productGrid-search candidates per state cell
Discrete branchesSeparate candidate surfaces and envelope work
Stochastic nodesContinuation evaluations before expectation
Outer candidatesComplete inner solves in a nested method
Refinement budgetAdaptive outer search, where node count is data-dependent
Period/regime shapesNumber of distinct programs JAX may compile
SubjectsSimulation batch size and compiled shape

For grid search, continuous action grids multiply. For a nested solver, outer candidates multiply the inner solve — a declared count under a finite outer grid, and a budget-bounded one under an adaptive mesh. For EGM, a savings-grid construction replaces one current state-by-action search but interpolation and envelope work remain.

These expressions predict direction, not wall time. Kernel fusion, memory traffic, compile reuse, and device occupancy can reverse a simple operation-count ranking.

CPU and GPU tendencies

Work shapeCPU tendencyGPU tendency
Dense static map-reduceViableUsually strong
Sequential topology scanOften strongOften poor
Query-side segmented envelopeViableOften preferable
Many small compiled shapesLower launch costCan be compile/launch bound
Large independent batchesCore parallelismStrong if memory permits

GridSearch is therefore not merely a slow fallback: a modest dense action grid can be excellent on a GPU. Conversely, an EGM method with sequential scans or many small shapes can lose despite lower arithmetic complexity.

A large GPU should not be treated as a faster small GPU automatically. Independent tests, regimes, branches, subjects, or candidate chunks can improve occupancy, but only if scheduling keeps aggregate device memory bounded. pylcm’s current public controls mainly stream individual numerical axes; cross-test concurrency remains a workflow decision.

Compilation and execution are different costs

JAX compiles programs for concrete shapes. Long lifecycle models with changing active regimes may produce many programs even when each executes quickly. pylcm enables a persistent compilation cache by default; repeated processes can reuse entries when the model and environment produce the same compilation key.

Measure at least:

  1. cold environment and cold compilation cache;

  2. warm environment and cold model compilation;

  3. warm compilation cache;

  4. repeated execution in one process;

  5. peak host and device memory.

A speed claim that reports only the fourth quantity does not answer whether the model is practical in estimation or CI.

What is known and what remains empirical

The candidate-growth relationships above are structural. Exact break-even grid sizes, the fastest envelope backend, useful batch sizes, and benefits from a larger GPU depend on model topology and hardware. The documentation should not manufacture universal thresholds.

Use the external LCM solver benchmarks for evolving comparisons, and benchmark your own model with the workflow in Performance and memory tuning. Package maintainers should use the distinct development benchmark suite.