Keeping agents alive on long-horizon tasks with heartbeat logs
A thread on long-horizon reliability where one line per run — started, did X, finished or failed — beats restart policies, because a missing line is the alarm. The takeaway: reliability is less about keeping the agent alive and more about noticing quickly when it is not.