I keep production apps running, so customers never notice when something breaks.
Site Reliability Engineer for fintech and regulated SaaS on AWS and Azure, with banking experience at a digital bank and ING. I prevent outages, fix them fast, and make every change safe and auditable.
Available for remote SRE contracts · UTC+8
| 130,000+ | stuck customer emails recovered during an outage, none lost |
| 3.7 min | median response to production alerts, vs the team's 6.0 min |
| 0 min | downtime during major platform upgrades |
| 69 | production changes, each with a change request and outcome on record, about 6% rolled back safely |
| 100% | of random error pages eliminated, no downtime |
What I do for a team
In plain terms, this is what hiring an SRE gets you.
I find weak spots before customers do and add monitoring that warns the team early.
When something breaks, I find the cause, fix it, and write down what happened so it doesn't repeat.
Every change gets a ticket, a window, and a way back: no downtime surprises, and the paper trail SOC 2 and ISO auditors ask for.
Tools I built
AI-assisted operations tools, built so they can't do damage without a human saying yes.
Watches multi-hour deployments, sorts failures by type, and retries only the safe ones. Real data errors are never rerun automatically.
Azure DevOps, AI agent toolingWorks support cases and change requests from a link, but can't change anything without my explicit approval.
Salesforce, tool-level permissionsPulls incident data read-only and reports only totals, never client data.
Squadcast APIHow I think
I use mental models, the latticework Charlie Munger describes, as everyday working tools. Here's where each one paid off in production.
Ask how it could fail, then prevent that.
Before platform upgrades, I listed how customer traffic could drop and capped how many servers could be replaced at once.
Upgrades with zero downtimeFind the real cause, not the symptom.
Random error pages came down to two systems with mismatched timeout settings. I fixed the root setting everywhere at once.
100% of those errors eliminatedAsk "and then what?" before acting.
Releasing 130,000 queued emails at once could overwhelm the systems receiving them, so I tuned throughput and released them in controlled batches.
All delivered in 90 minutes, none lostLeave room for being wrong.
My AI ops tools cannot change anything without a human approving it, and upgrades only run after a health check passes.
100+ servers upgraded with safety checksMake the safe path the default path.
Every production change went through the same change request: plan, window, rollback, recorded outcome.
69 changes shipped, about 6% rolled back safelyDashboards can be wrong. Check the source.
Before publishing my on-call numbers, I checked who acted on each incident against the raw logs.
On-call numbers checked against raw logsProblems I've solved
Each one in plain English, with the technical details underneath for engineers.
| Case | The problem | Why it mattered | What I did | Result |
|---|---|---|---|---|
| 01 · Email outage | A bank's transactional email stopped, and 130,000+ messages piled up. | Customers weren't getting OTPs and receipts. | Found the bottleneck and cleared the backlog in controlled batches.Postfix concurrency and worker tuning | Every email delivered in 90 min, none lost. |
| 02 · Error pages | Customers randomly hit "gateway timeout" errors. | Customers saw failed requests in a banking app. | Traced it to two systems with mismatched timeout settings and fixed it everywhere at once.ALB and Istio idle timeout, Terraform | Errors eliminated with no downtime. |
| 03 · Upgrades | Platform upgrades risked taking the app offline. | Banks can't schedule customer-facing outages. | Built an automated, staged upgrade process.EKS, Karpenter disruption budgets, Terraform | Upgrades shipped with no downtime. |
| 04 · On-call | Production alerts fired around the clock for a SaaS platform law firms use. | Every slow response is a longer outage for customers. | Took on-call and acknowledged 277 production alerts.Squadcast, Azure, Salesforce | 3.7 min median response, 38% faster than the team. |
What people say
I started as one of ten ULAP.org cloud scholars in the Philippines. Today I mentor the next cohort.
“I’m incredibly grateful to my mentor, Theron Bueno, for an insightful and inspiring six-month mentorship. Thank you for generously sharing your knowledge, not just on technical topics but also on essential soft skills like tailoring resumes, interview preparation, and confidence during an interview.”
Start here: a 2-week reliability review
A small, fixed-scope first step, so you can judge my work before committing to a monthly contract.
Experience
Regulated, customer-facing platforms in banking and legal tech.
- Dec 2025 – now
- Site Reliability Engineer, Digital bank
Keep a digital bank's apps and services online on Amazon Web Services.
AWS EKS, Karpenter, Istio, Terraform, Dynatrace, AWS DevOps Agent - Dec 2024 – Dec 2025
- Site Reliability Engineer, ING Hubs Philippines
Ran monitoring, disaster recovery drills, and upgrades for a global bank.
Azure, RHEL, OpenShift, Elastic Stack, LGTM, Ansible - Dec 2023 – Dec 2024
- Backend Engineer, ING Hubs Philippines
Built internal tools and took over a critical login service.
Java, Spring Boot, Node.js, Go - Contract
- DevOps / Support Engineer, Legal tech SaaS company
On call for a platform law firms rely on: resolved incidents, ran 69 production changes through change management, and handled 128 service requests.
Azure, Azure DevOps, Squadcast, Salesforce - 2018 – 2023
- Freelance full-stack developer
Built and ran websites and web apps for 30+ clients.
How I work
Tools I use
For engineers and technical reviewers.
| Area | Tools |
|---|---|
| Cloud and servers | Linux (RHEL), AWS (EKS, EC2), Azure, Kubernetes, OpenShift, Karpenter, Istio, Docker |
| Automation | Terraform, Ansible |
| Monitoring | Dynatrace, Elastic Stack, Loki, Grafana, Tempo, Mimir |
| Release pipelines | GitLab CI, GitHub Actions, Azure DevOps |
| Incident management | Squadcast, Salesforce |
| Languages | Python, Bash, Go, Java, Node.js |
| AI operations | AWS DevOps Agent, Claude, GitHub Copilot, Cursor |
Certifications
- Microsoft Certified: Azure Developer Associate (AZ-204)
- Microsoft Certified: Azure Fundamentals (AZ-900)
- AWS Certified Cloud Practitioner
- Google IT Support Professional Certificate
Education and community
- BS Computer Engineering, Pamantasan ng Lungsod ng Maynila (2023)
- Mentor at ULAP.org, helping early-career developers land their first engineering roles