Case study

A thousand identical POSTs: how we found and fixed an infinite loop in a vendor package in 20 minutes

A production case study from ORDR, a coffee pre-ordering app. Laravel 12, push notifications via Expo, observability by Unravel. Real trace IDs, real numbers, and a real bug in someone else's code.

TL;DR: A scheduled command from a third-party package got stuck in a loop. A task meant to finish in seconds kept running for hours (eight at its worst), hitting the Expo Push API every 200 ms the whole time. Unravel surfaced the symptom within a minute. The root cause was a bug in the latest release of a vendor package, and Claude Code took a few minutes to pin it down: it fetched the trace from Unravel over MCP, read it, and pointed at the exact line. Twenty minutes later a fix with regression tests was live in production.

The symptom: a wall of identical requests

It started with a routine glance at HTTP Client in Unravel, the page that lists every outgoing request the application makes. Normally it's an even mix: a payment provider, an SMS gateway, Telegram, Expo. This time, instead of a mix, there was a wall:

200  POST /--/api/v2/push/getReceipts   exp.host   201.4ms
200  POST /--/api/v2/push/getReceipts   exp.host   205.7ms
200  POST /--/api/v2/push/getReceipts   exp.host   203.8ms
200  POST /--/api/v2/push/getReceipts   exp.host   202.5ms
...screen after screen of this

The HTTP Client page in Unravel: a wall of identical POST getReceipts requests to exp.host

One endpoint, one host, ~200 ms per request, several per second. getReceipts is part of the Expo Push API: you send a push, get a ticket back, then ask for a receipt to learn whether it was delivered. Nobody intended to ask five times a second.

Who is doing this?

Second click: Scheduled Tasks. Unravel traces every run of every scheduler task, and the summary table points straight at the culprit:

TASK                                SCHEDULE          RUNS   FAILED   P95
artisan orders:advance-status       Every minute      2,733  0        3.67s
artisan expo:notifications:send     Every minute      1,437  0        3.41s
artisan expo:tickets:check          Every 10 minutes  103    0        416m 5s   ← !!!
artisan shifts:check-unclosed       Hourly            22     0        3.61s

The Scheduled Tasks page in Unravel: p95 of 416 minutes on expo:tickets:check

A p95 of 416 minutes, on a task that runs every ten minutes and should finish in seconds. And yet FAILED: 0. The task never crashed and no alerts fired; as far as classic monitoring is concerned, everything is fine. It just works. For hours.

Every run is red

Clicking the task opens its card: run statistics and Recent traces, the latest invocations with their durations. The DURATION column is red from top to bottom:

12m 2s · 59m 14s · 29m 26s · 10m 54s · 59m 6s · 74m 6s · 133m 15s · 49m 36s · 485m 52s

The scheduled task card in Unravel: Recent traces, every run red, from 10 to 133 minutes

Further down the list sits the record: 485 minutes. Eight hours straight of the task grinding away without a pause, while the scheduler stacked a fresh run on top every ten minutes.

Opening the trace

From Recent traces we open one of the runs. Unravel keeps the full trace of each:

The expo:tickets:check trace in Unravel: 29m 26s, 5001 events, ERROR, and a wall of identical N+1 insights

The header speaks for itself: 29m 26s, 5,001 events, ERROR. Time breakdown: SQL 0%, outgoing HTTP 12%, application code 88%. Below it, an Insights block with 825 signals, all identical: N+1 · repeated query · delete from expo_tickets where 0 = 1.

Scroll down to the timeline and hit the HTTP filter: of the 5,001 events exactly one thousand remain, and it's a thousand identical rows.

The trace timeline with the HTTP filter on: 1000 identical POST getReceipts at ~200 ms each

A thousand POST https://exp.host/--/api/v2/push/getReceipts at ~200 ms apiece, back to back with no pauses. The loop is visible to the naked eye. The only open question is why it never terminates, and that's a question for the code.

The agent takes over

Unravel ships an MCP server, and it's connected to Claude Code: traces, exceptions and stats are available to the agent directly. The prompt was literally this:

"take a look, something is stuck in a loop on production, trace 019f2ca4714d738c854a8bca77b2d3db, and the sends to POST /--/api/v2/push/getReceipts exp.host"

Claude called get_trace and got the full picture of the run:

Trace · 'artisan' expo:tickets:check
Status: ERROR · Duration: 1766.37s · Events: 5001
Time breakdown: sql 0% (3.93s) · http_out 12% (212.76s) · app 88% (1549.68s)

Half an hour of work, five thousand events, and inside it the same pattern repeating hundreds of times:

select count(*) as aggregate from `expo_tickets` -- count > 0? keep going
select * from `expo_tickets` limit 1000          -- fetch the same 1000 tickets
-- POST exp.host /getReceipts                    -- ask Expo (~200ms)
delete from `expo_tokens` where `value` in (?)   -- dead token removed
delete from `expo_tickets` where 0 = 1           -- ticket "removed"... nothing happens

The same pattern on the timeline, unfiltered. This is the anatomy of a single loop iteration:

The trace timeline unfiltered: the repeating cycle count → select → POST getReceipts → delete where 0 = 1

The smoking gun: delete from expo_tickets where 0 = 1. That's how Laravel compiles a whereIn with an empty array. Something in the loop builds a list of tickets to delete, the list comes out empty every time, the delete removes nothing, and the exit condition never fires. Unravel records SQL with real bindings and samples nothing, which is the only reason this half-millisecond query is visible at all. All five thousand times.

Root cause: a bug in the vendor package

Next, Claude opened the code. Not ours: the yieldstudio/laravel-expo-notifier package, which is where the expo:tickets:check command comes from.

public function handle(...): void
{
    while ($ticketStorage->count() > 0) {          // while the table isn't empty
        $tickets = $ticketStorage->retrieve();      // the same first 1000
        $response = $expoNotificationsService->receipts($ticketIds);
        if ($response->isEmpty()) {
            break;
        }
        $this->check($ticketStorage, $tickets, $response);
    }
}

protected function check(...): void
{
    $tickets->each(function (ExpoTicket $ticket) use (...) {
        // ...
        if ($receipt->details['error'] === 'DeviceNotRegistered') {
            event(new InvalidExpoToken($ticket->token));
            return;                                 // ← THE BUG: the ticket is NOT deleted
        }
        $ticketsToDelete[] = $ticket->id;
    });

    $ticketStorage->delete($ticketsToDelete);       // ← empty array → where 0 = 1
}

For a receipt with status DeviceNotRegistered (the user uninstalled the app, so the token is dead) the package fires a token-cleanup event and returns early, without ever adding the ticket to the delete list. The token disappears from expo_tokens; the row in expo_tickets stays forever.

The rest follows mechanically. Expo keeps receipts for about a day and dutifully returns them on every request: the response is never empty, the break never triggers, and count() > 0 holds forever. A single dead token, one uninstalled app, is enough to turn every run of the command into an hours-long bombardment of Expo.

What about upstream? We were already on the latest version of the package, and the bug was still there. Nothing to upgrade to. We had to fix it ourselves.

The fix

The immediate stop-gap was a single line of SQL. The tickets table is just bookkeeping for cleaning up dead tokens, so it's safe to wipe:

DELETE FROM expo_tickets;

The permanent fix is our own command replacing the package's. Three differences:

  1. The ticket is deleted on DeviceNotRegistered too. The token-cleanup event still fires, but the row no longer gets stuck.
  2. Cursor pagination by id instead of while (count() > 0): one pass over the table per run. An infinite loop is now structurally impossible, because the command terminates even if not a single row gets deleted.
  3. A 24-hour TTL, so tickets Expo will never answer for again are purged immediately, without any outbound request.

And a regression test that cites the incident right in its header:

// Acceptance (regression: production 2026-07-04, trace 019f2ca4714d738c854a8bca77b2d3db -
// the packaged expo:tickets:check looped forever because tickets with a
// DeviceNotRegistered receipt were never deleted from expo_tickets):
// A ticket with DeviceNotRegistered is deleted from expo_tickets, the invalid token
// is deleted from expo_tokens, and exactly one request goes out to Expo.
it('deletes the ticket and the invalid token on a DeviceNotRegistered receipt without looping', ...);

Http::assertSentCount(1): exactly one request to Expo. The exact opposite of the wall on the HTTP Client page.

As it happens, the "after" is already visible in the Scheduled Tasks screenshot above. The expo:tickets:process row, our new command, sits right below the broken expo:tickets:check with a p95 of 476.5ms against its 416 minutes. The screenshot was taken half an hour after the fix was deployed, so before and after ended up in the same frame.

Why this was findable at all

Unravel traces scheduled tasks and outgoing HTTP the same way it traces requests. A classic APM would have told you "this command is slow". Here you can see exactly what it does for all those thirty minutes: every SQL query with its bindings, every outgoing request, all 5,001 events of a single run.

It also helped that Unravel doesn't sample. delete ... where 0 = 1 is a half-millisecond query with a success status; any sampling tool would have thrown it away first. And it turned out to be the key piece of evidence.

And MCP turns observability into something an agent can work with. We didn't have to describe the symptoms to Claude Code or paste logs into the chat. It called get_trace itself, spotted the pattern itself, opened the vendor code itself and matched one against the other. From "take a look at what happened" to the exact buggy line in somebody else's package: minutes.

From the wall of requests in the screenshot to a fix deployed to production with tests took about 20 minutes. Most of that went into writing the tests, not finding the bug.


ORDR - a coffee pre-ordering app. Backend: Laravel 12, PHP 8.4. Observability: Unravel. Agent: Claude Code + Unravel MCP.

Debug it in production.

Connect your real app in two minutes. Free forever on a 3-day window: the full debugger, not a demo.

Start free