Case Studies

Execution Narratives

Three production systems, told the way I think about them: the situation, the constraint that made it hard, the decisions and trade-offs I owned, and the measured result. Numbers are real; deeper architecture walkthroughs available on request.

@Mobecom / LifeIQ

Oct 2016 - Jun 2018

Rescued a 200K-client platform under a 14-day shutdown deadline — then helped scale the group to 6M+ users

Technical PM → CTO (LifeIQ) → Group Head of Technology (Mobecom)

0s
downtime relaunching the 200K-client platform under a 14-day deadline
~90%
infra cost cut (~$5.5K → ~$500/mo) at 5M-user scale
6M+
users across the group
~$30M
group revenue; acquired
CTX::SituationI joined LifeIQ as Technical Product Manager and was promoted to CTO within six weeks. An outsourced overseas vendor — billing US$30K/month — had built and hosted the Ruby-on-Rails platform serving 200K clients and, amid a commercial dispute, gave notice it would shut the platform down within 14 days. The business was weighing whether to close entirely.
LIM::ConstraintA hard 14-day deadline to stand up a replacement, 200K active clients who could not be disrupted, and the company’s content and operational data locked inside the vendor’s system.
DEC::Decisions & Trade-offs
  • 01Made the call to rebuild rather than wind the business down — securing the company’s content and data ahead of the deadline so a rebuild was even viable.
  • 02Hired and led a small team and personally owned architecture, code review, and technical oversight of a fast Node.js rebuild.
  • 03As Group Head of Technology, re-architected the group’s AWS infrastructure (Endless Rewards, 5M users) to demand-based auto-scaling — capacity on demand, not always-on 24/7.
OUT::ResultRelaunched the platform with zero downtime as the old system went dark. Cut infrastructure spend ~90% (~$5.5K → ~$500/month) at 5M-user scale, then expanded into multiple product lines on a shared backend. The group (Mobecom) grew to 6M+ users and ~$30M revenue and was acquired.
Node.jsAngularPostgreSQLMongoDBAWS EC2Auto-scaling

@Secure Code Warrior

Aug 2019 - Jul 2020

Recovered a global program 6 weeks behind — and 3x’d QA throughput

Senior Software Engineer / Senior Cloud Architect

6 wks
behind schedule → delivered on time
~300%
QA throughput increase
~50
engineers aligned across 10 locations
ISO 27001
certified — final technical authority
CTX::SituationBrought in as technical 2IC to the VP Engineering. Two coupled programs — ISO 27001 certification and the "Future Ready Platform" (ephemeral QA environments) — were 6 weeks behind schedule, with ~50 engineers across 10 global locations building against divergent standards.
LIM::ConstraintA slipping certification deadline with external auditors, no shared architecture baseline across sites, and QA bottlenecked on shared, long-lived test environments that serialised every team’s work.
DEC::Decisions & Trade-offs
  • 01Chaired an Enterprise Architecture Board to standardise architecture and tooling first — alignment before acceleration — rather than throwing hours at a moving target.
  • 02Introduced ephemeral, per-branch QA environments so testing ran in parallel instead of queuing on shared infra.
  • 03Architected zero-downtime CI/CD on Azure DevOps + AWS and drove the transition toward a monorepo to kill environment drift.
OUT::ResultRecovered the program and delivered on time. Ephemeral environments lifted QA throughput ~300%. Served as final technical authority during the ISO 27001 audit, which passed.
TypeScriptAngularMongoDBAWSAzure DevOpsMonorepo

@Keyin College

May 2023 - Jan 2025

Replaced fragmented manual ops with a platform — ~90% faster deploys

Project & Engineering Lead, Digital Operations

~90%
reduction in deployment cycle time
< 5 min
recovery time objective (RTO)
0
environment drift after standardisation
CTX::SituationAcademic and operational teams ran on fragmented, manual processes, with enrollment and admin work spread across disconnected tools. Every release risked environment drift, and there was no reliable recovery path if something broke.
LIM::ConstraintLive academic and operational teams that could not tolerate downtime during term, a graduate-heavy delivery team new to modern DevOps, and no existing CI/CD or backup discipline to build on.
DEC::Decisions & Trade-offs
  • 01Consolidated onto a single centralised platform (Node.js, PostgreSQL, React) instead of patching the existing tool sprawl.
  • 02Standardised on Docker and CI/CD to make environments reproducible — eliminating drift as a class of problem rather than fixing it case by case.
  • 03Engineered automated backup and recovery to a < 5-minute RTO, treating recoverability as a first-class requirement, not an afterthought.
OUT::ResultCut deployment cycle time ~90% and eliminated environment drift. Automated backup/recovery brought RTO under 5 minutes. Reliability, performance, and delivery velocity improved across academic and operational teams.
Node.jsPostgreSQLReactDockerCI/CD
OPEN_TO_WORK // AVAILABLE_NOW

Want the architecture walkthrough?

Open to contract, fractional, and full-time work — Sydney-based, remote worldwide, with work rights in Australia and Canada. Happy to walk through any of these systems — decisions, diagrams, and all.

Execute_Connect()