01 — Why this problem
Payments that fail without anyone choosing to leave.
A subscription charge fails because a balance was low, a bank timed out, or a card expired. The customer never cancelled. The industry calls this involuntary churn, and subscription-billing vendors such as Recurly and Chargebee commonly put it between a fifth and two-fifths of all cancellations. That range is vendor-published, not independently verified here.
What Razorpay does today
Their documentation is explicit about the retry model for subscriptions:
“In a T+3 days cycle, we will retry the payment thrice. That is, once every day for 3 days, excluding the date of the charge.” — after which the subscription moves to halted.razorpay.com/docs/payments/subscriptions/payment-retries/ · still live and quoted verbatim, 4 September 2026 · archived screenshot
This documented calendar is the evaluation baseline. A published schedule alone does not establish which internal signals a production system consumes.
Why that matters more in India than elsewhere
NPCI allows one attempt plus three retries per UPI AutoPay mandate, and those attempts are non-transferable — an attempt saved on one customer cannot be spent on another. Executions are also barred during peak windows. These caps took effect on 1 August 2025 under NPCI's UPI API guidelines, and the RBI rewrote India's recurring-payment rules on 21 April 2026 — repealing eight earlier circulars. The RBI framework this agent is built against is under five months old. Card rails elsewhere allow far more attempts, so Western retry engines optimise under a constraint India does not share.
The gap that made this worth building
Razorpay documents a specific limitation on Optimizer rules for recurring payments:
“Optimizer rules apply only for the first-time registration payment. All subsequent debits happen on the same terminal used for registration payment. Optimizer Rules will not be applicable for subsequent payments.”razorpay.com/docs/payments/optimizer/recurring-payments/ · still live and quoted verbatim, 4 September 2026
That restriction concerns gateway routing. Recovery Loop explores a related but different decision: when to retry, using failure reasons and issuer-health inputs. It does not demonstrate that Razorpay's own retries ignore downtime.
01b — Is this still a problem?
What the documentation establishes.
Published product descriptions motivate the experiment, but do not establish the absence of an internal capability.
1. Routing is different from retry timing
The Priority-based Routing section describes temporary gateway downtimes lasting twenty minutes when success rates fall below a threshold. Traffic moves to another gateway. Twenty minutes is the downtime duration, not a guaranteed detection or reaction time.
2. Subscription recovery is also a vendor product area
The project's archived Agent Studio reference describes a Subscription Recovery agent. That listing is not evidence that the capability is unavailable:
“Subscription Recovery— Analyzes failed subscription payments, apply smarter retry logic, and trigger targeted customer nudges.”razorpay.com/agent-studio · announced 12 March 2026 · catalogue checked 4 September 2026
Recovery Loop is an independently evaluated prototype, not proof of a missing vendor product.
3. The question this experiment tests
Under the stated simulation assumptions, does choosing retry times using failure reasons, issuer health and per-mandate opportunity cost improve recovery against a fixed ladder?
The results answer that experimental question. They do not audit Razorpay's internal scheduler or establish that no comparable production system exists.
02 — Where the data came from
What is real, and what is not.
I could not find transaction-level retry data published anywhere public. Card network rules treat authorisation responses as confidential scheme data, and no processor publishes retry chains at transaction level. Every source below was checked and is labelled by tier.
ObservedRazorpay test-mode captures
20 real payment entities captured from the Razorpay Test API, redacted at capture time and committed. Five carry a DCBL issuer that joins to the downtime feed. They establish the real error tuple shape.
Captured 22 August – 1 September 2026 · Razorpay Test API
See the raw events →Real · calibrationNPCI UPI AutoPay statistics
Monthly, per payer PSP: volume, approved %, business-decline % and technical-decline %. This is auto-debit mandate data specifically — the same object this agent operates on. Per-bank baselines are fitted from it.
Covers January 2025 – July 2026 · fetched 22 August 2026
NPCI moved its statistics to /product/ paths; the older /what-we-do/ URLs now 404. Both the live page and our capture are linked.
NPCI AutoPay statistics →Our captured data →Real · calibrationNPCI NACH returns & incident log
Destination-bank returns split into financial and non-financial declines — the real soft/hard partition — plus response timing, and a reportable-incident log used to shape outage windows.
Covers January 2025 – June 2026 · fetched 22 August 2026
Source: NPCI, NACH Ecosystem Statistics — Destination Bankwise (npci.org.in). Captured and committed for the same reason.
See the captured data →External benchmarkONS Direct Debit failure rate
UK government statistics from real Bacs clearing traffic, monthly since 2019, by sector. Used only to sanity-check the order of magnitude of recurring-payment failure — a different rail, never fed into the simulator.
Series January 2019 – July 2026 · 2026 edition · downloaded 23 August 2026
ONS dataset (OGL v3.0) → What had to be assumed
Repeat-attempt outcomes, transaction amounts, salary dates, exact outage placement, attempt costs, and the entire cards simulation. Each is listed in the limitations document, and each was perturbed by ±25% to check the result does not depend on it.
03 — How it is solved
Six steps, and one parameter that decides everything.
- 1
IngestRazorpay payment.failed and payment.downtime.* webhooks, HMAC-verified. Raw payloads are stored untouched.
- 2
DiagnoseThe failure is classified on the tuple (error_source, error_step, error_reason) — not the error code alone, which is a generic 400 bucket. An unmapped tuple is refused, never guessed.
- 3
Check the issuerIs this bank in an active outage, or declining well above its own baseline? Recovery Loop uses these inputs to evaluate whether to defer.
- 4
EstimateRecovery probability from the failure class, issuer health, elapsed time, and proximity to the customer's salary credit.
- 5
Price the attemptWith one attempt and three retries and a closing horizon, spending one now means not having it later. That opportunity cost is computed by backward induction and falls to zero at the horizon.
- 6
Act, or refuseRetry, wait and re-evaluate, or stop permanently. Hard stops are terminal; economic refusals are not. Every decision is logged with its inputs.
The parameter that turned out to decide the result
How long the agent may defer before being forced to act. The first design capped it at three days — and lost on both rails. Insufficient-funds failures cluster around salary day, so a three-day bound means the agent can never wait for the credit that actually funds the payment. It conserves attempts it can never spend, and roughly three thousand per cohort expire unused.
Raise the bound past a pay cycle and the same agent, unchanged, wins on 10 of 10 held-out seeds. The cap was chosen on one set of seeds, frozen, and validated once on ten it had never seen.
What it does better
- Recovers ₹28.6 lakh net against the fixed ladder's ₹17.9 lakh — +₹10,68,274
- Spends fewer attempts — 4,491.4 against 5,406.3 (means over 10 seeds)
- Recovers more payments — 1,230.4 against 838.5
- Wins on 10 of 10 held-out seeds, positive across a 14–35 day range
What it does not prove
- Cards is inconclusive: 6 of 10 seeds, range crosses zero
- Outcomes are simulated — no real merchant retry data exists publicly
- Ground truth is authored, though independent of the model
- No live A/B test; the project is locked to test mode
04
What this does not do yet
Real merchant outcome logs would make the next stage possible: learning from actual retry outcomes, validating cards, and testing recovery in production.
A learned policy, with merchant evidence.
The predictor is deliberately rule-based and the scheduler uses expected-value arithmetic. Training only on an authored simulator could optimise its assumptions; that would not establish a real-world improvement. With permissioned merchant logs, a budgeted contextual bandit is a possible next step: retry windows are the choices, each execution keeps its own limited retry allowance, and recovered revenue minus actual attempt costs provides the reward. The mandate-local pricing model provides a starting point for opportunity-cost accounting; sequential effects and selection bias would still need validation.
Resolve the cards result.
This project has not found a suitable public Indian card-authorisation decline baseline. Cards therefore remains uncalibrated and inconclusive at six of ten seeds. Merchant or card-network decline and retry-outcome data would support calibration and a new independent evaluation, not guarantee a win.
Validate with a controlled live test.
The project is locked to Razorpay test mode. With merchant permission and production safeguards, the next proof would be an A/B test against an agreed baseline, such as the fixed T+1/T+2/T+3 ladder where applicable. Measure total net recovered revenue alongside revenue per attempt, while enforcing execution limits and hard stops. No live trial has been run.
See it decide.
The console runs the same policy on a live decision stream. Take a bank down and watch pending retries become deferrals.