The case

Why this problem, where the
data came from, and how it's solved.

Recovery Loop investigates whether failure reasons and issuer health can improve retry timing under per-mandate limits. It does not claim access to Razorpay's internal retry implementation.

01 — Why this problem

Payments that fail without anyone choosing to leave.

A subscription charge fails because a balance was low, a bank timed out, or a card expired. The customer never cancelled. The industry calls this involuntary churn, and subscription-billing vendors such as Recurly and Chargebee commonly put it between a fifth and two-fifths of all cancellations. That range is vendor-published, not independently verified here.

What Razorpay does today

Their documentation is explicit about the retry model for subscriptions:

“In a T+3 days cycle, we will retry the payment thrice. That is, once every day for 3 days, excluding the date of the charge.” — after which the subscription moves to halted.razorpay.com/docs/payments/subscriptions/payment-retries/ · still live and quoted verbatim, 4 September 2026 · archived screenshot

This documented calendar is the evaluation baseline. A published schedule alone does not establish which internal signals a production system consumes.

Why that matters more in India than elsewhere

NPCI allows one attempt plus three retries per UPI AutoPay mandate, and those attempts are non-transferable — an attempt saved on one customer cannot be spent on another. Executions are also barred during peak windows. These caps took effect on 1 August 2025 under NPCI's UPI API guidelines, and the RBI rewrote India's recurring-payment rules on 21 April 2026 — repealing eight earlier circulars. The RBI framework this agent is built against is under five months old. Card rails elsewhere allow far more attempts, so Western retry engines optimise under a constraint India does not share.

The gap that made this worth building

Razorpay documents a specific limitation on Optimizer rules for recurring payments:

“Optimizer rules apply only for the first-time registration payment. All subsequent debits happen on the same terminal used for registration payment. Optimizer Rules will not be applicable for subsequent payments.”razorpay.com/docs/payments/optimizer/recurring-payments/ · still live and quoted verbatim, 4 September 2026

That restriction concerns gateway routing. Recovery Loop explores a related but different decision: when to retry, using failure reasons and issuer-health inputs. It does not demonstrate that Razorpay's own retries ignore downtime.

01b — Is this still a problem?

What the documentation establishes.

Published product descriptions motivate the experiment, but do not establish the absence of an internal capability.

1. Routing is different from retry timing

The Priority-based Routing section describes temporary gateway downtimes lasting twenty minutes when success rates fall below a threshold. Traffic moves to another gateway. Twenty minutes is the downtime duration, not a guaranteed detection or reaction time.

2. Subscription recovery is also a vendor product area

The project's archived Agent Studio reference describes a Subscription Recovery agent. That listing is not evidence that the capability is unavailable:

“Subscription Recovery— Analyzes failed subscription payments, apply smarter retry logic, and trigger targeted customer nudges.”razorpay.com/agent-studio · announced 12 March 2026 · catalogue checked 4 September 2026

Recovery Loop is an independently evaluated prototype, not proof of a missing vendor product.

3. The question this experiment tests

Under the stated simulation assumptions, does choosing retry times using failure reasons, issuer health and per-mandate opportunity cost improve recovery against a fixed ladder?

02 — Where the data came from

What is real, and what is not.

I could not find transaction-level retry data published anywhere public. Card network rules treat authorisation responses as confidential scheme data, and no processor publishes retry chains at transaction level. Every source below was checked and is labelled by tier.

Observed

Razorpay test-mode captures

20 real payment entities captured from the Razorpay Test API, redacted at capture time and committed. Five carry a DCBL issuer that joins to the downtime feed. They establish the real error tuple shape.

Captured 22 August – 1 September 2026 · Razorpay Test API

See the raw events →
Real · calibration

NPCI UPI AutoPay statistics

Monthly, per payer PSP: volume, approved %, business-decline % and technical-decline %. This is auto-debit mandate data specifically — the same object this agent operates on. Per-bank baselines are fitted from it.

Covers January 2025 – July 2026 · fetched 22 August 2026

NPCI moved its statistics to /product/ paths; the older /what-we-do/ URLs now 404. Both the live page and our capture are linked.

NPCI AutoPay statistics →Our captured data →
Real · calibration

NPCI NACH returns & incident log

Destination-bank returns split into financial and non-financial declines — the real soft/hard partition — plus response timing, and a reportable-incident log used to shape outage windows.

Covers January 2025 – June 2026 · fetched 22 August 2026

Source: NPCI, NACH Ecosystem Statistics — Destination Bankwise (npci.org.in). Captured and committed for the same reason.

See the captured data →
External benchmark

ONS Direct Debit failure rate

UK government statistics from real Bacs clearing traffic, monthly since 2019, by sector. Used only to sanity-check the order of magnitude of recurring-payment failure — a different rail, never fed into the simulator.

Series January 2019 – July 2026 · 2026 edition · downloaded 23 August 2026

ONS dataset (OGL v3.0) →

What had to be assumed

Repeat-attempt outcomes, transaction amounts, salary dates, exact outage placement, attempt costs, and the entire cards simulation. Each is listed in the limitations document, and each was perturbed by ±25% to check the result does not depend on it.

03 — How it is solved

Six steps, and one parameter that decides everything.

  1. 1
    Ingest

    Razorpay payment.failed and payment.downtime.* webhooks, HMAC-verified. Raw payloads are stored untouched.

  2. 2
    Diagnose

    The failure is classified on the tuple (error_source, error_step, error_reason) — not the error code alone, which is a generic 400 bucket. An unmapped tuple is refused, never guessed.

  3. 3
    Check the issuer

    Is this bank in an active outage, or declining well above its own baseline? Recovery Loop uses these inputs to evaluate whether to defer.

  4. 4
    Estimate

    Recovery probability from the failure class, issuer health, elapsed time, and proximity to the customer's salary credit.

  5. 5
    Price the attempt

    With one attempt and three retries and a closing horizon, spending one now means not having it later. That opportunity cost is computed by backward induction and falls to zero at the horizon.

  6. 6
    Act, or refuse

    Retry, wait and re-evaluate, or stop permanently. Hard stops are terminal; economic refusals are not. Every decision is logged with its inputs.

The parameter that turned out to decide the result

How long the agent may defer before being forced to act. The first design capped it at three days — and lost on both rails. Insufficient-funds failures cluster around salary day, so a three-day bound means the agent can never wait for the credit that actually funds the payment. It conserves attempts it can never spend, and roughly three thousand per cohort expire unused.

Raise the bound past a pay cycle and the same agent, unchanged, wins on 10 of 10 held-out seeds. The cap was chosen on one set of seeds, frozen, and validated once on ten it had never seen.

What it does better

  • Recovers ₹28.6 lakh net against the fixed ladder's ₹17.9 lakh — +₹10,68,274
  • Spends fewer attempts — 4,491.4 against 5,406.3 (means over 10 seeds)
  • Recovers more payments — 1,230.4 against 838.5
  • Wins on 10 of 10 held-out seeds, positive across a 14–35 day range

What it does not prove

  • Cards is inconclusive: 6 of 10 seeds, range crosses zero
  • Outcomes are simulated — no real merchant retry data exists publicly
  • Ground truth is authored, though independent of the model
  • No live A/B test; the project is locked to test mode

04

What this does not do yet

Real merchant outcome logs would make the next stage possible: learning from actual retry outcomes, validating cards, and testing recovery in production.

A learned policy, with merchant evidence.

The predictor is deliberately rule-based and the scheduler uses expected-value arithmetic. Training only on an authored simulator could optimise its assumptions; that would not establish a real-world improvement. With permissioned merchant logs, a budgeted contextual bandit is a possible next step: retry windows are the choices, each execution keeps its own limited retry allowance, and recovered revenue minus actual attempt costs provides the reward. The mandate-local pricing model provides a starting point for opportunity-cost accounting; sequential effects and selection bias would still need validation.

Resolve the cards result.

This project has not found a suitable public Indian card-authorisation decline baseline. Cards therefore remains uncalibrated and inconclusive at six of ten seeds. Merchant or card-network decline and retry-outcome data would support calibration and a new independent evaluation, not guarantee a win.

Validate with a controlled live test.

The project is locked to Razorpay test mode. With merchant permission and production safeguards, the next proof would be an A/B test against an agreed baseline, such as the fixed T+1/T+2/T+3 ladder where applicable. Measure total net recovered revenue alongside revenue per attempt, while enforcing execution limits and hard stops. No live trial has been run.

See it decide.

The console runs the same policy on a live decision stream. Take a bank down and watch pending retries become deferrals.