Financial Systems
Payment Settlement and Recovery
I built recovery workflows that finish processing merchant payments after staff record a final ledger result. For deposits and bank withdrawals, I connected workflow updates and partner notifications in one database transaction, so a failed queue write leaves the work ready for retry. I also added delivery records and controlled resends for operators.
- Role
- Software Engineer
- Employment
- to
- My Work
- Recovery jobs, database transactions and concurrency handling, callback delivery history, controlled resends, and failure-path tests.
When the Ledger and the Application Disagree
A recorded payment result and a completed application workflow are different things. Operations staff needed a way to finish the remaining work after a financial decision had been recorded, while keeping the payment result intact.
I built the recovery path between those two states. It finds unfinished deposit and bank-withdrawal workflows, checks the ledger result, closes the associated work, and queues the partner callback: the HTTP notification that tells an external system what happened to the payment.
Preserving the Financial Decision
The ledger was the authority for the payment result. My settlement job used that recorded result to finish the application workflow, without changing the transaction’s amounts, repricing the payment, or applying another balance movement.
That distinction shaped the retry behaviour. I rechecked the current payment and workflow states while holding database locks, so competing workers could not make decisions from an outdated view of the same records. Recovery had to respect the financial decision already made and stay within the remaining application work.
Making a Failed Queue Write Recoverable
Closing the workflow and then queueing a callback as a separate operation would leave a gap: if the process stopped between those writes, the application could mark work complete without recording the notification it still owed. I put both writes inside one database transaction so they would succeed or fail together.
- Lock and RecheckRead the final ledger result and confirm that the workflow is still open.
- Close the WorkflowApply that result to the deposit session or withdrawal request.
- Queue the CallbackRecord the notification work within the transaction.
If queue insertion fails, the workflow remains open for another attempt. I also checked that the queue configuration supported the transaction guarantee. This relies on PostgreSQL’s all-or-nothing transaction behaviour.
I kept the decision to claim work close to the records being changed. That made concurrent processing and recovery easier to reason about: a worker needed a current, protected view of the work before acting on it.
Keeping Retries Within Their Scope
I processed recovery work in bounded batches and isolated each payment’s transaction. One failed notification enqueue could be reported while processing continued for other payments. Scheduled runs gave interrupted work another opportunity to complete without turning the schedule into a delivery-time guarantee.
A repeated run leaves completed workflows alone. I kept routine recovery and operator-requested notification resends as separate actions, so each had a clear scope and an investigation could follow what had already happened.
Making Delivery Useful to Operators
I added callback delivery records and resend handling so operators could investigate what happened after a notification was queued. The history distinguishes pending work from recorded failures and gives operators useful delivery information without storing signing secrets.
Resends preserve the earlier delivery history and record the new attempt separately. I also applied partner access checks to the resend path. Operators can follow an investigation across attempts without losing the context of the original notification.
The database transaction guarantees that the local workflow update and queue write commit together. Receipt by a remote partner remains a separate boundary: a queued job is not proof of receipt, and retrying an HTTP request cannot by itself guarantee exactly-once delivery.
Testing the Interruptions
I wrote regression tests around the points where the workflow could diverge:
- Run settlement twice for confirmed and cancelled payments, then check that the ledger transaction and wallet balance are unchanged and only one callback job is queued.
- Force callback dispatch to fail, check that the workflow remains open, then rerun settlement and check that it completes.
- Fail one payment’s enqueue and verify that another payment can still settle.
- Reject queue configurations that cannot support the transaction guarantee, and leave work outside recovery’s scope untouched.
Delivery and resend tests also cover connection failures, retry handling, repeated processing, and exclusion of signing material from the delivery history. These checks tie the recovery behaviour to the states an operator would actually need to investigate.
Outcome
The backend can finish deposit and bank-withdrawal workflows after staff finalise a payment, preserving the ledger’s result and leaving interrupted work available for retry. Operators have a delivery history and an explicit resend path for notifications that need attention. The recovery job completes the remaining application work without replaying the financial movement.