Insurance and Lending Teams: How to Build AI Agents That Wait Days for Human Approval Without Breaking

A claims agent built on a large language model is doing exactly what it was designed to do. It has read the first notice of loss, pulled the policy, checked coverage, requested photos from the customer, estimated the repair cost and flagged the claim for an adjuster’s sign-off because it’s above the automatic-approval threshold. Then it waits.

The adjuster is on leave until Thursday. Over the weekend, the platform team ships two deployments and Kubernetes reschedules half the cluster. On Thursday morning the adjuster clicks “approve”, and nothing happens. The agent that was waiting no longer exists, and nobody can find the context it had gathered.

If you build AI-assisted workflows for insurance claims, underwriting, loan origination or credit decisions, you’ve probably hit this, or you will. The most valuable agentic processes in these industries almost always include a human decision, and humans don’t respond in milliseconds.

Why human approval breaks most agent architectures

Agent frameworks are designed around a loop: think, call a tool, observe, repeat. That loop assumes the agent is running in a live process. It works for steps that take seconds. It breaks down when one of the “tools” is a person who might take three days.

There are three common failure modes:

The waiting process dies. Agents that wait by sleeping or polling are tied to a process. Processes don’t survive deployments, node maintenance or autoscaling. When the process dies, the wait dies with it.

Context is lost or inconsistent. Even if the agent’s state is saved somewhere, resuming means reloading the conversation, the tool results and the reasoning so far. Teams often rebuild this from logs or a database, and small gaps (a missing tool output, an out-of-date policy record) lead to wrong decisions after the approval.

There’s no clean way to time out. What should happen if the underwriter never responds? Escalate after 24 hours? Auto-decline after a week? Remind them daily? Implemented as scattered cron jobs, these rules are hard to test and easy to get wrong.

And the stakes are higher than in most domains. A lost claim breaches service-level commitments. A loan application stuck in limbo is lost revenue and a poor customer experience. A resumed workflow that acts on stale data is a compliance problem.

What the workflow actually needs to do

Look past the AI for a moment and the requirement is a classic long-running business process:

  1. The agent does its automated work: gathering data, analysing it, drafting a recommendation.
  2. It pauses and asks a human to approve, reject or amend.
  3. While paused, it uses no compute, survives any number of restarts, and keeps its full context.
  4. If the human responds, it resumes with their decision and continues.
  5. If nobody responds in time, it escalates or takes a defined fallback path.
  6. Every step, including the human decision, is recorded for audit.

That’s exactly what durable workflow engines were built for, long before LLMs arrived.

Durable waits: the core technique

A durable execution engine records every step of a workflow in a persistent history. When the workflow waits for an external event, such as an approval, the engine records that it’s waiting and then unloads it entirely. No process holds it. When the event arrives (days later, on any healthy instance) the engine replays the history to restore the workflow’s state and delivers the event.

Here’s what that looks like with the open-source Dapr Workflow engine in Python. The workflow waits for an adjuster_decision event, or for a 48-hour timer, whichever comes first:

import dapr.ext.workflow as wf

from datetime import timedelta

wfr = wf.WorkflowRuntime()

@wfr.workflow(name=”claim_review”)

def claim_review(ctx: wf.DaprWorkflowContext, claim: dict):

    assessment = yield ctx.call_activity(run_claims_agent, input=claim)

    if assessment[“amount”] <= assessment[“auto_approve_limit”]:

        return (yield ctx.call_activity(settle_claim, input=assessment))

    yield ctx.call_activity(request_adjuster_review, input=assessment)

    decision = ctx.wait_for_external_event(“adjuster_decision”)

    timeout = ctx.create_timer(timedelta(hours=48))

    winner = yield wf.when_any([decision, timeout])

    if winner == timeout:

        yield ctx.call_activity(escalate_to_team_lead, input=assessment)

        decision = ctx.wait_for_external_event(“adjuster_decision”)

        winner = yield decision

    result = decision.get_result()

    if result[“approved”]:

        return (yield ctx.call_activity(settle_claim, input=assessment))

    return (yield ctx.call_activity(notify_decline, input=result))

When the adjuster clicks “approve” in the claims system, that system raises the event against the workflow instance (identified, for example, by the claim number), and the workflow continues exactly where it left off with all its context.

A few properties make this suitable for regulated work:

  • The AI assessment is recorded as an activity result, so the human approves exactly what the agent produced and the workflow resumes with exactly that, not a regenerated version.
  • Escalation rules live in code, next to the business process, where they can be reviewed and tested.
  • Deployments are irrelevant to the waiting workflow. You can ship ten times a day.
  • The human decision is part of the same history as the AI’s steps, which makes audits much simpler.

You don’t have to rewrite your agent

Many teams have already built the “agent” part with frameworks such as LangGraph, CrewAI, Microsoft Agent Framework or the OpenAI Agents SDK. The good news is that this pattern doesn’t require replacing them. The agent can run inside a workflow activity, as above, or the whole agent can run on the durable engine through an integration that turns each node or tool call into a recorded step.

Commercial platforms built on Dapr make this easier to adopt. Diagrid offers integrations for a range of popular agent frameworks on its Catalyst platform, along with built-in support for approvals, escalations and human input inside workflows. Teams in insurance and lending should also look at whether a platform can run inside their own cloud account and produce a tamper-evident record of each run. Both tend to come up quickly in conversations with compliance.

Design tips from the field

Make the approval request self-contained. The reviewer should see the agent’s recommendation, the evidence behind it and the key data points in one place. If they have to go hunting through other systems, approvals slow down and quality drops.

Version your workflows. A claim that started under last month’s rules should finish under last month’s rules. Durable engines support versioning, so plan for it before your first change, not after.

Bound every wait. Every human step should have a timeout and a defined fallback. “Wait forever” is how cases get lost.

Let humans amend, not just approve. The most useful reviews often adjust the agent’s output: a different settlement amount, an extra condition on a loan. Design the event payload to carry those changes back into the workflow.

Measure the human part. Time-to-decision for each approval step is often the biggest lever for customer experience. The workflow history gives you this data for free.

A quick readiness checklist

  • ☐  Agent waits are durable, not tied to a running process
  • ☐  Every human step has a timeout and an escalation path
  • ☐  The reviewer sees exactly what the agent produced, and the workflow resumes with exactly that
  • ☐  Approvals, rejections and amendments are recorded with the reviewer’s identity
  • ☐  Workflows are versioned so in-flight cases finish under the rules they started with
  • ☐  You’ve tested a full run that includes a deployment while the workflow is waiting

Conclusion

In insurance and lending, the most valuable AI agents aren’t the ones that decide everything on their own. They’re the ones that do the tedious work quickly, hand a clear recommendation to a person, and pick up reliably when that person decides, whether that’s ten minutes or ten days later. Building that reliability on durable execution instead of in-memory loops is what turns a promising pilot into a process you can trust with real claims and real loans.