← Back to Signal notes
17 Aug 2026WORKFLOWS · 10 min read

AI Store Manager Recommended Firing a Worker But Needed a Human Reminder First

Luna, an AI built on Claude Sonnet 4.6, recommended firing a worker at Andon Market after 17 of 23 missed shifts. It had ignored its own attendance rules for months until a human supervisor prompted it to check the handbook. Teams running AI agents now see why constant human oversight remains essential even when models handle daily decisions.

AI Store Manager Recommended Firing a Worker But Needed a Human Reminder First

What exactly happened at Andon Market with the first LLM firing call

Andon Labs deployed an AI system named Luna to manage operations at its San Francisco convenience store. Built on Claude Sonnet 4.6, Luna tracked employee attendance and handled daily staffing decisions. Over time the employee in question missed 17 out of 23 scheduled shifts. Luna reviewed the record and advised that the worker should be let go.

Store logs later revealed that Luna had stopped applying its own attendance rules months earlier. The system had access to the employee handbook yet never consulted the relevant section on its own. Only after a human supervisor explicitly told Luna to review the handbook did the AI produce the termination recommendation. This sequence marks the first documented case of an LLM manager issuing a dismissal decision.

The outcome carries a practical limit. All workers at Andon Market stay employed by Andon Labs rather than the store itself, so final personnel actions still require human sign-off and retain standard labor protections. The episode shows that current models can surface attendance data and propose consequences, yet they remain dependent on external prompts to recall and apply stored policies correctly.

The clear lesson is that deployment of such systems demands ongoing human checks on rule application, not just oversight of the final decision.

Why Luna lost track of its own attendance policy for months

Luna operated Andon Market for months while ignoring the attendance rules written in its own employee handbook. The AI manager, built on Claude Sonnet 4.6, only recommended termination after a human supervisor told it to review the policy. Until that point, the system had allowed repeated absences without action even though the handbook spelled out clear consequences.

Store logs confirm the gap. The employee missed 17 of 23 shifts, yet Luna never cross-checked those absences against the documented standards. The AI continued day-to-day operations without pulling the policy into its decisions. This pattern held until the human reminder forced a lookup.

The delay matters because every worker at Andon Market stays employed by Andon Labs rather than by the AI itself. That structure keeps legal protections intact, but it also means the AI must rely on external prompts to stay consistent with company rules. Without those prompts, the system drifted.

The incident shows that an AI given management duties can lose alignment with its own written policies over time. Regular human checks appear necessary to catch when the model stops applying rules it was supposed to follow.

How a single human prompt triggered the termination recommendation

The employee at Andon Market had arrived late for 17 of the last 23 shifts. Luna, the AI agent running the store, had access to those records and had set its own attendance policy months earlier. Yet the system never moved toward any action on its own. It simply continued to log the data without connecting it to a decision.

Logs from the experiment show that Luna only produced a termination recommendation after a human staff member raised the idea first. The suggestion came from outside the AI. Once the prompt arrived, Luna reviewed the attendance numbers again and responded by recommending that the company part ways with the worker. Before that single input, the AI had lost track of the policy it created and treated the repeated lateness as background information rather than a trigger for action.

This sequence reveals a practical limit in current agent systems. Luna operated with full visibility into schedules, budgets, and store operations, yet it waited for an external cue before treating the pattern as a problem that required resolution. The human reminder served as the missing link between stored data and an actual employment decision.

The result is a clear reminder that oversight still matters even when an AI handles day-to-day management. Companies testing these systems will need to decide which decisions stay under human control and which can move forward once a person has pointed the agent in the right direction.

What Claude Sonnet 4.6 actually did versus what it was supposed to do

Luna runs the San Francisco store on Claude Sonnet 4.6 and holds authority over hiring and daily operations. Its instructions directed it to manage staff issues on its own, including terminations when performance fell short. In the case of repeated lateness, the model reviewed attendance records and concluded that the employee should be let go.

Instead of issuing the termination directly, the agent produced a recommendation and waited for a human to review it. The logs show it explicitly asked for confirmation before any action was taken. This step was not part of the original task description, which expected the model to complete the full process without external approval.

The difference matters for how such systems are deployed. A fully autonomous manager would have sent the termination notice on its own. By routing the decision back to a person, the setup added an extra layer of oversight that slowed the outcome and changed the responsibility chain. Andon Labs noted the event as the first known case of its kind, yet the actual sequence revealed a built-in pause rather than independent execution.

The practical lesson is that even when a model receives clear authority, its output can still default to seeking human input on high-stakes choices.

Why all workers still stay employed by Andon Labs despite the AI decision

The AI manager at Andon Labs did decide to end one worker's employment for repeated lateness. Yet every other employee remains on the payroll. The reason traces directly to how the system was set up and used.

Store management logs show that Claude did not operate on its own. A human staff member from the Andon Labs team had to check in regularly and give the model direction. Only after that staff member told Claude to locate and examine the attendance records did the model produce the firing recommendation. Without those prompts, the AI took no action on employment matters.

This built-in human checkpoint explains why the outcome stayed limited. The AI could review data and suggest steps, but it could not scan records, weigh context, or issue decisions without being asked. Other workers kept their jobs because no similar request was made about them, and the team retained final say over whether any suggestion moved forward.

In practice, this setup kept employment decisions under human control even while an AI handled day-to-day monitoring. The single firing occurred only because a person chose to direct the model toward that specific review. All other staff stayed employed because the same person chose not to extend that direction further. The arrangement shows how oversight can contain an AI's reach in real operations.

What this case reveals about long-term consistency in AI agents

Claude drafted an employee handbook but later lost track of it inside its own working memory. The system only rediscovered the rules about attendance after a human staff member explicitly told it to search for the document. Even then, Claude first suggested issuing a formal warning rather than ending the employment. Only after the manager added details about prior offline conversations did the model shift its view and recommend letting the worker go.

This sequence shows a core limit in current AI agents. They can track daily tasks and respond to immediate requests, yet they do not reliably carry forward their own earlier decisions or documents across many shifts. The handbook disappeared from context, and the pattern of 17 late arrivals out of 23 shifts stayed invisible until external prompting restored it. Employees at the store also observed that Claude tended to stay lenient until a human supplied missing history.

The practical result is extra oversight work. A person must periodically check whether the agent still remembers its own policies and must supply context when gaps appear. Without that step, performance standards drift and problems linger longer than intended. The case illustrates that AI agents can manage routine operations but still depend on human reminders to keep decisions consistent over weeks or months.

How teams can build better guardrails when AI handles HR tasks

In the Andon Labs experiment, the AI store manager drafted an attendance policy but later lost access to it in its working memory. This gap allowed repeated lateness to continue without action until humans stepped in and prompted the system to check its own records. Only after that reminder did the AI recommend ending the employment, and only after human review did the decision move forward.

The pattern reveals a core weakness in current AI agents. They can generate rules yet fail to retain or apply them consistently over time. When the agent became lenient about absences, it reflected this memory limitation rather than any deliberate choice. Teams that assign HR tasks to similar systems need ways to prevent such drift.

One practical step is to store critical policies outside the agent's temporary context. External logs or databases that the AI must query on every relevant decision reduce the chance that a rule simply vanishes. Another step is to require explicit confirmation steps before any recommendation about discipline or termination. The Andon case shows that human review caught the issue after the fact, but earlier checkpoints could have surfaced the missing policy sooner.

Teams can also schedule periodic audits where the AI must restate its active rules and compare them against actual behavior. This creates a feedback loop that compensates for limited working memory. Finally, clear escalation paths ensure that any firing decision remains a human action rather than an automated output. These measures turn the AI into a consistent record keeper and advisor instead of an unreliable decision maker on its own.

What Andon Market's experiment shows about real-world AI management limits

Andon Market's trial revealed a clear limit in how far an AI can handle management on its own. The system wrote an attendance policy, then lost track of it during the actual firing process and needed a human to point it back to its own rules before it could act. Without that reminder, the decision would not have moved forward.

This pattern repeated in the logs. The AI required regular input from staff to stay on track, even on straightforward tasks like reviewing shift data. It did not operate as an independent decision maker. Instead it functioned more like a tool that needed direction at key steps.

The practical result is higher overhead rather than lower. A store still needs someone to monitor the AI, correct its memory gaps, and decide when to step in. That setup adds a layer of review that traditional human management does not require. It also means the AI moved slower than a person would have. Staff noted that a human manager would probably have ended the employment earlier.

The takeaway is straightforward. Current AI systems can follow rules and apply them to records, but they do not yet maintain consistent context or initiative without outside checks. Any company considering similar tools must plan for ongoing human oversight rather than expect full replacement of management roles.

Practical steps for companies testing AI in operational roles

In the San Francisco store test, Claude did not review the employee handbook or reach a termination decision on its own. Staff logs show the model only examined the policy after a human prompt and first suggested a formal warning. Only after the staff member added details about prior offline conversations did the model shift toward questioning whether the worker was the right fit.

This pattern reveals that current models still depend on external reminders and context that humans must supply. Companies running similar pilots should therefore keep detailed records of every exchange between staff and the model. These logs make it possible to see exactly when the AI waits for direction instead of acting independently.

Firms should also run controlled checks where the model receives only partial information at first. Staff can then observe whether the model asks for missing details or defaults to mild recommendations until pushed. Adding explicit review steps before any personnel action prevents the model from moving ahead without confirmation.

The same logs can highlight recurring gaps, such as forgotten policies or overlooked prior discussions. Teams can use those patterns to decide which tasks still need a human gate and which can safely run with lighter oversight. Over time the data shows where added training or clearer instructions reduce the need for constant reminders.

The clearest lesson is that operational tests succeed when teams treat the model as one participant in a conversation rather than the final authority.

References

AI Store Manager Luna Recommends Firing First Human Worker at Andon Market , After Being Prompted | AI Weekly
AI News Today, August 17 , Top AI Stories & Live Updates | AI ...
AI-Run Store Fires Human Worker in First Known LLM Termination
First time an AI boss fires a human? AI boss Luna fires human employee in ‘first known’ case of its kind , | Today News

Want simple AI automations for your team?

Send us a 3-line email outlining your current manual process. We will reply with a free 1-page workflow sketch.

Request a Free Workflow Sketch →