Operating Kestrel means more than keeping a process online. You should be able to identify the affected session and run, see what happened, choose a safe action, and verify that the same user path works afterward.
Prepare and deploy
Start with local and remote runner connections, then configure environment and authentication and review the production operating model. These guides keep runner and provider credentials on trusted servers and establish the health checks required before serving users.
Investigate and recover
Use Observability to locate the request, Reliability to structure the incident, Replay to inspect persisted evidence, and Troubleshooting to start from a visible symptom.
Control and quality
Operator control acts on work already in flight. Ruhroh evaluations and quality gates help compare behavior before and after a change. Kestrel One model operators should also understand credential leases and model access decisions.