I keep production apps running, so customers never notice when something breaks.

Site Reliability Engineer for fintech and regulated SaaS on AWS and Azure, with banking experience at a digital bank and ING. I prevent outages, fix them fast, and make every change safe and auditable.

Available for remote SRE contracts · UTC+8

Recent results
130,000+stuck customer emails recovered during an outage, none lost
3.7 minmedian response to production alerts, vs the team's 6.0 min
0 mindowntime during major platform upgrades
69production changes, each with a change request and outcome on record, about 6% rolled back safely
100%of random error pages eliminated, no downtime

What I do for a team

In plain terms, this is what hiring an SRE gets you.

Fewer outages

I find weak spots before customers do and add monitoring that warns the team early.

Faster recovery

When something breaks, I find the cause, fix it, and write down what happened so it doesn't repeat.

Safer, auditable changes

Every change gets a ticket, a window, and a way back: no downtime surprises, and the paper trail SOC 2 and ISO auditors ask for.

Tools I built

AI-assisted operations tools, built so they can't do damage without a human saying yes.

Deployment pipeline watcher

Watches multi-hour deployments, sorts failures by type, and retries only the safe ones. Real data errors are never rerun automatically.

Azure DevOps, AI agent tooling
Support ticket agent

Works support cases and change requests from a link, but can't change anything without my explicit approval.

Salesforce, tool-level permissions
On-call metrics report

Pulls incident data read-only and reports only totals, never client data.

Squadcast API

How I think

I use mental models, the latticework Charlie Munger describes, as everyday working tools. Here's where each one paid off in production.

Inversion

Ask how it could fail, then prevent that.

Before platform upgrades, I listed how customer traffic could drop and capped how many servers could be replaced at once.

Upgrades with zero downtime
First principles

Find the real cause, not the symptom.

Random error pages came down to two systems with mismatched timeout settings. I fixed the root setting everywhere at once.

100% of those errors eliminated
Second-order thinking

Ask "and then what?" before acting.

Releasing 130,000 queued emails at once could overwhelm the systems receiving them, so I tuned throughput and released them in controlled batches.

All delivered in 90 minutes, none lost
Margin of safety

Leave room for being wrong.

My AI ops tools cannot change anything without a human approving it, and upgrades only run after a health check passes.

100+ servers upgraded with safety checks
Checklists

Make the safe path the default path.

Every production change went through the same change request: plan, window, rollback, recorded outcome.

69 changes shipped, about 6% rolled back safely
The map is not the territory

Dashboards can be wrong. Check the source.

Before publishing my on-call numbers, I checked who acted on each incident against the raw logs.

On-call numbers checked against raw logs

Problems I've solved

Each one in plain English, with the technical details underneath for engineers.

CaseThe problemWhy it matteredWhat I didResult
01 · Email outageA bank's transactional email stopped, and 130,000+ messages piled up.Customers weren't getting OTPs and receipts.Found the bottleneck and cleared the backlog in controlled batches.Postfix concurrency and worker tuningEvery email delivered in 90 min, none lost.
02 · Error pagesCustomers randomly hit "gateway timeout" errors.Customers saw failed requests in a banking app.Traced it to two systems with mismatched timeout settings and fixed it everywhere at once.ALB and Istio idle timeout, TerraformErrors eliminated with no downtime.
03 · UpgradesPlatform upgrades risked taking the app offline.Banks can't schedule customer-facing outages.Built an automated, staged upgrade process.EKS, Karpenter disruption budgets, TerraformUpgrades shipped with no downtime.
04 · On-callProduction alerts fired around the clock for a SaaS platform law firms use.Every slow response is a longer outage for customers.Took on-call and acknowledged 277 production alerts.Squadcast, Azure, Salesforce3.7 min median response, 38% faster than the team.

What people say

I started as one of ten ULAP.org cloud scholars in the Philippines. Today I mentor the next cohort.

“I’m incredibly grateful to my mentor, Theron Bueno, for an insightful and inspiring six-month mentorship. Thank you for generously sharing your knowledge, not just on technical topics but also on essential soft skills like tailoring resumes, interview preparation, and confidence during an interview.”
Christian Ortiz · Full-Stack Software Developer, mentored through ULAP.org

Start here: a 2-week reliability review

A small, fixed-scope first step, so you can judge my work before committing to a monthly contract.

Week 1Access, an architecture walkthrough, and a close look at your alerts, incidents, and change process.
Week 2A written report ranking your top reliability risks by business impact, plus fixes for the quick wins.
PriceFixed and agreed up front. No open-ended hours.
AfterContinue monthly if it was useful. Either way, you keep the report and the fixes.

Experience

Regulated, customer-facing platforms in banking and legal tech.

Dec 2025 – now
Site Reliability Engineer, Digital bank

Keep a digital bank's apps and services online on Amazon Web Services.

AWS EKS, Karpenter, Istio, Terraform, Dynatrace, AWS DevOps Agent
Dec 2024 – Dec 2025
Site Reliability Engineer, ING Hubs Philippines

Ran monitoring, disaster recovery drills, and upgrades for a global bank.

Azure, RHEL, OpenShift, Elastic Stack, LGTM, Ansible
Dec 2023 – Dec 2024
Backend Engineer, ING Hubs Philippines

Built internal tools and took over a critical login service.

Java, Spring Boot, Node.js, Go
Contract
DevOps / Support Engineer, Legal tech SaaS company

On call for a platform law firms rely on: resolved incidents, ran 69 production changes through change management, and handled 128 service requests.

Azure, Azure DevOps, Squadcast, Salesforce
2018 – 2023
Freelance full-stack developer

Built and ran websites and web apps for 30+ clients.

How I work

EngagementRemote contract, long-term preferred.
HoursUTC+8: full overlap with Australia and Asia, EU mornings, and overnight coverage for US teams.
CommunicationClear written updates, plus runbooks and post-incident reports for everything I touch.

Tools I use

For engineers and technical reviewers.

AreaTools
Cloud and serversLinux (RHEL), AWS (EKS, EC2), Azure, Kubernetes, OpenShift, Karpenter, Istio, Docker
AutomationTerraform, Ansible
MonitoringDynatrace, Elastic Stack, Loki, Grafana, Tempo, Mimir
Release pipelinesGitLab CI, GitHub Actions, Azure DevOps
Incident managementSquadcast, Salesforce
LanguagesPython, Bash, Go, Java, Node.js
AI operationsAWS DevOps Agent, Claude, GitHub Copilot, Cursor

Certifications

  • Microsoft Certified: Azure Developer Associate (AZ-204)
  • Microsoft Certified: Azure Fundamentals (AZ-900)
  • AWS Certified Cloud Practitioner
  • Google IT Support Professional Certificate

Education and community

  • BS Computer Engineering, Pamantasan ng Lungsod ng Maynila (2023)
  • Mentor at ULAP.org, helping early-career developers land their first engineering roles

Need someone to keep your systems running?

Email me