Benchmarking Harness

FHOPS ships sample scenarios in examples/ (tiny7, med42, large84, and the synthetic tiers under examples/synthetic/) that cover increasing planning horizons. The Phase 2 benchmarking harness runs the MIP and heuristic solvers across these datasets, captures objectives/KPIs, and stores results for inspection.

Quick Start

fhops bench suite --out-dir tmp/benchmarks
fhops bench suite --scenario examples/tiny7/scenario.yaml --scenario examples/med42/scenario.yaml --out-dir tmp/benchmarks_med
fhops bench suite --scenario examples/large84/scenario.yaml --out-dir tmp/benchmarks_large --time-limit 180 --include-sa False
fhops bench suite --scenario examples/synthetic/small/scenario.yaml --out-dir tmp/benchmarks_synth --sa-iters 200 --include-mip False
fhops bench suite --include-ils --include-tabu --out-dir tmp/benchmarks_compare

This command:

  • loads each bundled scenario (tiny7 → med42 → large84 → synthetic-small by default),

  • solves them with the MIP (HiGHS) and simulated annealing using default limits, and

  • writes a summary table to tmp/benchmarks/summary.{csv,json} alongside per-solver assignment exports (mip_assignments.csv, sa_assignments.csv).

CLI Options

fhops bench suite accepts a number of flags:

  • --scenario / -s — add one or more custom scenario YAML paths. When omitted the built-in scenarios are used.

  • --time-limit — HiGHS time limit in seconds (default: 1800 for the large84 horizon). The quick-start example above still shows --time-limit 180 for a smoke run; omit that flag to use the higher default when you want optimal certificates on the largest instance.

  • --sa-iters / --sa-seed — simulated annealing iteration budget and RNG seed.

  • --driver — MIP driver (auto/highs-appsi/highs-exec/gurobi/gurobi-appsi/gurobi-direct) mirroring the solve-mip CLI.

  • --include-mip / --include-sa — toggle individual solvers when running experiments.

  • --out-dir — destination for summary files (default: tmp/benchmarks).

Watching heuristic progress

All heuristic entry points (fhops solve-heur, fhops solve-ils, fhops solve-tabu, fhops tune-*, and fhops bench suite) accept --watch/--no-watch plus --watch-refresh <seconds>. When enabled, FHOPS renders a Rich dashboard showing the shared metrics (scenario, solver, iteration, best/current/rolling objective, runtime, restarts/workers) and a solver-specific detail row (SA temperature/acceptance, ILS perturbations, Tabu tenure).

Example:

fhops solve-heur examples/med42/scenario.yaml \\
  --iters 200000 \\
  --cooling-rate 0.99999 \\
  --restart-interval 500 \\
  --watch \\
  --watch-refresh 0.5

The dashboard refreshes only when stdout is an interactive terminal. CI runs or redirected output print a single warning (Watch mode disabled: not running in an interactive terminal.) and continue normally. Adjust --watch-refresh (default 0.5 s) to control update cadence.

The workers column reports the requested parallel workers, but note that --parallel-workers currently uses Python threads—batch scoring remains GIL-bound. For true multi-core utilisation, prefer --parallel-multistart or process-level orchestration until the scoring loop is parallelised.

FAQ – Watch Mode

  • “Watch mode disabled: not running in an interactive terminal.” The dashboard renders only when stdout is a TTY. If you want both the live view and a log, wrap the command with script (or run inside tmux/screen) so the subprocess gets a pseudo-terminal:

    script -q -c "fhops solve-heur ... --watch" /tmp/fhops_watch.log
    

    The terminal shows the dashboard; the log captures the CLI text emitted after completion.

  • How do I capture a screenshot/GIF? Run a short watch-enabled command (e.g., fhops solve-heur examples/tiny7/scenario.yaml --watch --iters 500) and use a terminal recorder such as asciinema or ttystudio. The sparkline now renders below the main table so column widths stay stable while recording.

Standard Manuscript Pipeline

All experiments cited in the manuscript follow the same FHOPS CLI pipeline so readers can replay each stage without bespoke tooling:

  1. Validate or snapshot scenarios. Run fhops dataset validate <scenario.yaml> for each bundle (e.g., examples/tiny7 and examples/med42) and, when needed, snapshot synthetic tiers via fhops synth generate out/synthetic_{small,medium,large} with fixed RNG seeds.

  2. Benchmark solvers. Invoke fhops bench suite with explicit --scenario and --out-dir arguments (for the manuscript: docs/softwarex/assets/data/benchmarks/<slug>/) plus the desired heuristics (--include-ils, --include-tabu) and iteration budgets.

  3. Run the tuning harness. Launch python docs/softwarex/manuscript/scripts/run_tuning_benchmarks.py (wrapped by scripts/generate_assets.sh) to produce leaderboard/comparison tables in docs/softwarex/assets/data/tuning/.

  4. Replay schedules. Call fhops playback on the best SA/ILS assignments (deterministic and stochastic modes) so utilisation and downtime metrics land under docs/softwarex/assets/data/playback/<scenario>/<solver>/<mode>/.

  5. Estimate costs. Execute fhops dataset estimate-cost for each machine bundle to capture owning/operating/repair totals in docs/softwarex/assets/data/costing/.

  6. Synthetic scaling sweep. Use python docs/softwarex/manuscript/scripts/run_synthetic_sweep.py to regenerate runtime-vs-size CSV/JSON + plots under docs/softwarex/assets/data/scaling/.

Each command appends its parameters, commit hash, runtime, and SHA-256 digests to docs/softwarex/assets/benchmark_runs.log (when run through make manuscript-benchmarks or scripts/generate_assets.sh), providing a single provenance log for all artefacts.

Optional Gurobi backend (Linux)

HiGHS remains the default open-source MIP solver. If you have access to a Gurobi licence (e.g. academic named-user), install the optional extras and register the licence before selecting --driver gurobi:

# install gurobipy alongside FHOPS
pip install fhops[gurobi]

# download the lightweight licence tools bundle (version shown is illustrative)
wget https://packages.gurobi.com/lictools/licensetools13.0.0_linux64.tar.gz
tar xvfz licensetools13.0.0_linux64.tar.gz

# request your licence key (replace XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX)
./grbgetkey XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX

# accept the default path (typically $HOME/gurobi.lic) or specify a custom location
# if you store the licence elsewhere, export GRB_LICENSE_FILE so gurobipy can find it:
export GRB_LICENSE_FILE=/path/to/gurobi.lic

# quick sanity check
python -c "import gurobipy as gp; m = gp.Model(); m.setParam('OutputFlag', 0); m.optimize()"

Once the licence is active, any FHOPS command can use Gurobi by passing --driver gurobi (or the other Gurobi driver variants). Without an available licence the CLI will fall back to HiGHS.

Interpreting Outputs

The summary CSV/JSON records, per scenario/solver pair:

  • objective value (incorporating any objective weights),

  • runtime (wall-clock seconds),

  • number of assignments in the exported schedule,

  • key KPIs: total production, mobilisation cost, sequencing violation counts, etc.

  • For SA runs, iteration budget and RNG seed are included to help compare tuning parameters across experiments.

  • Comparison helpers:

    • solver_category labels exact vs heuristic solvers.

    • best_heuristic_solver / best_heuristic_objective identify the strongest heuristic per scenario.

    • objective_gap_vs_best_heuristic shows how far each solver trails the top heuristic (negative values mean the solver beats the best heuristic, e.g., MIP).

    • runtime_ratio_vs_best_heuristic reports runtime multiples relative to the quickest heuristic winner.

Shared KPI roll-up

+=================+========+==============+===========+=============+=============+=============+ | Scenario | Solver | Preset | Objective | Assignments | Runtime (s) | Best solver | +=================+========+==============+===========+=============+=============+=============+ | med42 | ils | | -1068.919 | 122 | 1.30 | ils | +—————–+——–+————–+———–+————-+————-+————-+ | med42 | sa | default | -631.573 | 106 | 2.98 | ils | +—————–+——–+————–+———–+————-+————-+————-+ | med42 | sa | diversify | -631.573 | 106 | 3.02 | ils | +—————–+——–+————–+———–+————-+————-+————-+ | med42 | sa | mobilisation | -176.113 | 80 | 12.39 | ils | +—————–+——–+————–+———–+————-+————-+————-+ | med42 | tabu | | 18.716 | 7 | 2.69 | ils | +—————–+——–+————–+———–+————-+————-+————-+ | tiny7 | ils | | 23.000 | 8 | 0.17 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | tiny7 | sa | default | 15.500 | 17 | 0.48 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | tiny7 | sa | diversify | 15.500 | 17 | 0.46 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | tiny7 | sa | mobilisation | 32.000 | 10 | 1.24 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | tiny7 | tabu | | 17.000 | 5 | 0.09 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | synthetic_small | ils | | 39.791 | 5 | 0.13 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | synthetic_small | sa | default | 39.791 | 5 | 0.31 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | synthetic_small | sa | diversify | 39.791 | 5 | 0.31 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | synthetic_small | sa | mobilisation | 39.791 | 5 | 0.64 | sa | +—————–+——–+————–+———–+————-+————-+————-+ | synthetic_small | tabu | | 39.791 | 5 | 0.05 | sa | +—————–+——–+————–+———–+————-+————-+————-+

The KPI roll-up highlights three solver behaviours we plan to discuss in the manuscript:

  • MiniToy: ILS beats the SA baselines on objective (23 vs. 15.5) with ~3x faster runtime by focusing on the last-mile improvements. Tabu is competitive on runtime but falls short on assignments.

  • Med42: The large production penalty showcases why we need heuristics beyond SA—only ILS hits the negative objective targets expected from contraction penalties. The mobilisation preset surfaces the landing-limited behaviour we outline in the narrative.

  • Synthetic tier: All solvers converge to the same objective, but Tabu demonstrates the runtime floor (0.05 s) while SA presets show how mobilisation shakes affect runtime.

These shared tables let us cite numbers in both the manuscript and the howto/benchmarks doc without copy/paste drift.

Interpretation tips:

  • Positive gaps mean the solver under-performs the best heuristic; negative gaps indicate an improvement (common for MIP or exploratory heuristics).

  • Runtime ratios greater than 1 are slower than the best heuristic; numbers below 1 are faster.

  • Combine these columns with telemetry logs to pinpoint operators that drive under-performance.

Example JSON snippet:

{
  "scenario": "tiny7",
  "solver": "sa",
 "objective": 9.5,
  "runtime_s": 0.02,
  "kpi_total_production": 42.0,
  "kpi_mobilisation_cost": 65.0,
  "kpi_mobilisation_cost_by_machine": "{\"H2\": 65.0}",
  "kpi_sequencing_violation_count": 0
}

Assignments are stored under <out-dir>/<scenario>/<solver>_assignments.csv. Feed these into fhops evaluate or project-specific analytics notebooks to dig deeper.

Mobilisation KPIs now include kpi_mobilisation_cost_by_machine (JSON string) so you can identify which machines drive the bulk of movement spend. The larger examples/large84 scenario demonstrates the effect at scale; the CLI example above runs the MIP solver alone to keep runtimes bounded.

Operator Telemetry

When running the suite with simulated annealing enabled, the summary CSV/JSON includes an operators_config column (final weights) and an operators_stats column (per-operator telemetry). operators_stats is a JSON object recording proposals, accepted moves, skips, weights, and acceptance rates for each registered operator. Example snippet:

{
  "swap": {"proposals": 200, "accepted": 200, "skipped": 0, "weight": 1.0, "acceptance_rate": 1.0},
  "move": {"proposals": 200, "accepted": 200, "skipped": 0, "weight": 1.0, "acceptance_rate": 1.0}
}

Use --show-operator-stats with fhops solve-heur for a human-readable table, or parse the benchmark summaries programmatically to monitor operator performance over time. Persistent telemetry logs (--telemetry-log) append the same structure to a newline-delimited JSON file for long-term analysis.

Visual Comparisons

The helper script scripts/render_benchmark_plots.py consumes a benchmark summary and renders comparison charts for documentation. For example:

fhops bench suite --include-ils --include-tabu --no-include-mip --out-dir tmp/bench_visuals
python scripts/render_benchmark_plots.py tmp/bench_visuals/summary.csv --out-dir docs/_static/benchmarks
Objective gap per heuristic solver

Objective gap versus the best heuristic solver for each scenario (negative bars indicate the solver beats the current best heuristic).

Runtime ratios per heuristic solver

Runtime ratios relative to the best heuristic solver (values > 1.0 are slower).

Note

Regenerate the comparison summary and plots after significant heuristic changes by rerunning fhops bench suite --include-ils --include-tabu and the plotting helper script.

Each assignment CSV includes machine_id, block_id, day, and shift_id columns. The shift label is derived from the scenario’s shift calendar or timeline definition, enabling sub-daily benchmarking and evaluation workflows.

Regression Fixture

tests/fixtures/benchmarks/tiny7_sa.json records the expected seed-42 SA output for the tiny7 scenario (200 iterations) and is exercised by tests/test_benchmark_harness.py. Use it as a reference when extending the harness or adjusting solver defaults.

Future Work

Phase 2 follow-up tasks include:

  • integrating benchmarking plots into the documentation,

  • adding Tabu/ILS runs as metaheuristics mature, and

  • calibrating mobilisation penalties against GeoJSON distance inputs.