This is a local product prototype for a familiar infrastructure moment: an AI batch becomes less efficient and the person on call needs to decide whether to keep capacity, investigate or pause it.
Start with the person who has to decide
Infrastructure dashboards can expose every available reading while leaving the important question unanswered. In this experiment, the job is not treated as a wall of charts. It is a situation with a delay, a likely cause and a decision that has consequences.
The interface begins with the job's state and the recommended next step. Supporting readings explain why that recommendation exists, rather than asking someone to assemble the story across several panels.
Replay one incident, not an entire cloud
The prototype keeps the scenario deliberately small: utilisation falls, the expected completion time changes and the cost at risk becomes visible.
The controls replay stable capacity, underutilised capacity and a blocked pipeline. They are a local incident simulation designed to test the explanation and interaction, not a claim about a live customer workload.
Use real tooling behind the concept
The companion stack is Docker-ready. Prometheus is the collector: every few seconds it asks the local workload service for readings such as worker progress, pressure and estimated time remaining, then stores them as time series. Grafana is the viewer: it asks Prometheus for those readings, turns them into panels and evaluates alert rules.
The local service produces the same worker-level signals that the product prototype translates into a decision. Grafana is intentionally not the hero. It is the inspectable technical layer behind the experiment: useful for verifying a query, following a signal or debugging an alert.
A technical evolution: test the evidence path
The job scenario remains a model because this Mac is not a GPU cluster. I later added a small real telemetry path to test the part that should not be fictional: how a signal arrives, stays fresh and becomes visible in the product.
A local Node process samples selected aggregate host readings every 15 seconds, sends them to Grafana Cloud through Prometheus remote write, and a private Vercel route queries the latest series. The Signal health panel shows that path without claiming that the Mac is the AI workload.
The decision surface, not the telemetry wall.
A model for deciding when an expensive training batch should keep capacity, be investigated or be paused. Live telemetry below validates the signal path, not this scenario.
Utilisation is stable and the next cost threshold is still distant. Extra capacity would add cost without improving completion time.
Worker health
all workers reportingA real signal path, kept separate from the scenario.
This second prototype verifies source, freshness and privacy boundaries. It publishes selected aggregate metrics from this Mac, not GPU or customer data.
Only selected aggregates are exposed. Host name, IP address, processes, Grafana workspace and credentials remain private.