Platform capability
Automatic failover
Every provider has bad hours. Peneu watches each one you've connected, moves new payments away from a provider that's failing, and brings it back gradually once it recovers.
- Health measured from your live payments, not status pages
- Paused per method, not the whole provider
- Recovery tested on a small share before full restore
Traffic is split across providers A, B and C. Provider A's error rate rises, so its traffic is paused and shifted to B and C. Peneu sends a small amount of probe traffic to A; once A is healthy again, its share is restored.
Traffic share per provider over one incident
The problem
Status pages update after your customers have noticed
Provider incidents rarely begin as clean outages. Timeouts creep up on one method, one bank's netbanking slows down, or a UPI PSP starts returning errors at peak. By the time a status page changes, you've already lost a stretch of payments.
Switching by hand means someone has to notice, decide and change config, often late at night. Then someone has to remember to switch back.
Early signs that get missed
- Timeouts rising on one method while others look fine
- UPI payments sitting in “pending” longer than usual
- Response times doubling before errors appear
- Success rate slipping a few points with no incident declared
How it works
Measure, pause, probe, restore
- 01
Measure
Every attempt feeds a rolling health view per provider and method: error rate, timeout rate, response time and the share of payments still pending. The data comes from your own traffic, so it reflects what your customers are experiencing.
- 02
Pause
When a signal crosses your threshold, and there's enough volume for it to mean something, the provider is paused for that method. New payments route to the next eligible provider. Cards on the same provider can keep flowing while its UPI is paused.
- 03
Probe
After a cool-down, a small share of traffic goes back to the paused provider. If those payments succeed, it stays in rotation. If they don't, it stays paused and the cool-down starts again.
- 04
Restore and record
Traffic returns in steps rather than all at once. Every pause and restore is recorded with the signal that triggered it, so you can match it against the provider's own incident report later.
In-flight payments aren't moved
Failover changes where new payments go. Payments already created on the paused provider complete or fail there, and their status keeps updating as normal.
Thresholds
Signals, and example thresholds
The right thresholds depend on your volume. At low volume a handful of failures can look like a spike, so failover waits for a minimum number of attempts before acting.
| Signal | What it catches | Example threshold |
|---|---|---|
| Error rate | Provider returning 5xx or system errors | Above 3% over 5 minutes |
| Timeout rate | Slow or unreachable provider | Above 5% over 5 minutes |
| p95 response time | Degradation before outright failure | More than twice the provider's usual |
| Pending ratio | UPI or netbanking payments that don't resolve | Above 10% still pending after 2 minutes |
| Success-rate drop | Quieter problems, such as more declines than usual | 8 points below the provider's baseline |
Before and after
Switching by hand vs automatic failover
Switching by hand
- Someone notices a drop, often from customer complaints
- An engineer edits config or deploys
- All traffic moves at once, and comes back all at once
- Nothing records why, or for how long
Automatic failover
- Detected from live payment outcomes
- Paused per method, so UPI can fail over while cards keep flowing
- Recovery tested on a small share of traffic first
- Every pause and restore is logged with its trigger
Worth knowing
Failover needs somewhere to go
Failover only works if another connected provider supports the same method. Most businesses start with two providers for their main methods, usually UPI and cards, and add more where volume justifies it.
Evaluating it
Questions to ask about this, of any provider
Use these with any platform, including Peneu. For Peneu, what's available for your business is confirmed during onboarding.
- What triggers a failover, and who can override it?
- Automatic switching needs clear thresholds, and people need a manual switch for incidents.
- What happens to payments already in progress?
- In-flight payments on the failing provider still need their final status.
- How does traffic return when the provider recovers?
- Coming back gradually avoids swinging all traffic onto a provider that isn't fully well.
- How am I told?
- Failovers should be visible to your team as they happen, not discovered at month-end.
FAQ
Questions about automatic failover
What's the difference between failover and smart retry?
Retry recovers a single failed payment by trying another provider. Failover stops new payments going to a provider that's unhealthy. When a provider degrades, retry recovers the payments caught in the first minutes, and failover makes sure the next ones don't go there at all.
Can I fail over one payment method only?
Yes. Health is measured per provider and method, so a UPI problem on a provider doesn't take its cards out of rotation.
Can I take a provider out myself, for planned maintenance?
Yes. You can pause a provider or method manually, and routing treats it exactly as it would an automatic pause.
What if all my providers for a method are unhealthy?
Then there's no healthy route for that method. In your failover settings you decide whether the least-affected provider stays open or the method is reported as unavailable, so the customer can choose another one.
Works together with
Plan your second route
Tell us which methods carry most of your revenue. We'll help you work out where a second provider matters and what thresholds suit your volume.
Last reviewed . Samples on this page are illustrative.
