Blog
Notes from production estates rather than product documentation. Kubernetes at scale, platform engineering, cloud cost and the parts of compliance that actually shape architecture.
Buyer Guides3 min read
How to choose a Kubernetes consultant in the UKMost guidance on hiring a Kubernetes consultancy is written by consultancies. Here is the version that includes the questions we would find awkward.Read itPlatform Engineering5 min read
Six weeks to one day: what self-service infrastructure actually tookProvisioning took six weeks, and not because anyone was slow. Here is what it took to get it under a day across a hundred-plus cluster estate, and which part was hardest.Read itSecurity & Compliance3 min read
Making the secure path the easy pathVulnerabilities reaching production fell by around half. Almost none of that came from adding checks. It came from moving where the checks sat and making the governed route the fast one.Read itObservability3 min read
Federated observability: one view, separate tenantsRegulatory separation says telemetry cannot be pooled. On-call says they need one view. Federation is how both of those get to be true, and it is more work than a shared stack.Read itFinOps3 min read
Taking 80 percent out of a multi-cloud bill, in orderA portfolio bill across two clouds came down by around 80 percent. The order of the work mattered more than any individual technique, and most guidance gets that order backwards.Read itPlatform Engineering3 min read
Why golden paths get bypassed, and what to do about itA golden path only works while it is genuinely the easiest route. The moment it is missing something a team needs, they go around it, and now you maintain two platforms.Read itBuyer Guides3 min read
Questions to ask before signing a cloud migration contractMigrations rarely go wrong for technical reasons. They go wrong on what was assumed rather than agreed, and the assumptions surface at the least convenient moment.Read itSecurity & Compliance3 min read
SOC 2, ISO 27001, PCI DSS and GDPR at once: what actually overlapsFour regimes over one estate. A large share of the underlying controls are the same, and the cost is concentrated in the small number of places where they genuinely differ.Read itObservability3 min read
Hours to minutes: what actually moved MTTRMean time to resolution went from hours to minutes. The fixing was never the slow part. Almost all of the elapsed time went on working out where to look.Read itPlatform Engineering3 min read
Backstage is a product, not an installStanding up Backstage takes an afternoon. Getting engineers to open it twice is the actual project, and the gap between those two things is where most portal efforts quietly die.Read itFinOps3 min read
Committed use planning when the workloads are not yours to promiseCommitment discounts assume you can promise a baseline for one to three years. Across a portfolio where each company makes its own decisions, that promise is not yours to make.Read itKubernetes at Scale3 min read
Multi-tenancy when every tenant has a different regulatorMost multi-tenancy guidance assumes one organisation, one compliance regime and a shared appetite for risk. Remove those assumptions and the interesting decisions move somewhere else entirely.Read itBuyer Guides3 min read
Day rate, fixed price or retainer: which fits your problemThe right commercial model follows from how well you can describe the outcome. Pick it on that basis and the incentives line up; pick it on price and they usually do not.Read itSecurity & Compliance3 min read
SBOMs nobody reads, and how to make them load-bearingMost SBOM programmes produce a file per build that nobody ever opens. The generation is the easy part; making the result answer a question under time pressure is the actual work.Read itPlatform Engineering3 min read
Crossplane or Terraform modules? A decision record, not a comparisonThe honest answer is that we ran both, and the split between them was more useful than either tool would have been alone. Here is where the line fell and why.Read itObservability2 min read
Alerts that reach someone who can act on themMost alerting problems get diagnosed as threshold problems. They are usually routing problems: the page arrived somewhere that could see it but could not do anything about it.Read itFinOps3 min read
Cost allocation across tenants who must stay separateCost allocation inside one organisation is a tagging problem. Across tenants that must stay separate for regulatory reasons, it becomes an architecture problem with an accounting deadline.Read itKubernetes at Scale3 min read
Cutting deployment failures by 80 percent: the changes that matteredDeployment failures across the estate fell by around 80 percent. Not one change did that, and the ones that mattered most were not the ones we expected going in.Read itBuyer Guides3 min read
What a cloud architecture review should actually deliverAn architecture review is easy to sell and easy to do badly. The difference shows up in whether you can act on the output without the person who wrote it.Read itSecurity & Compliance3 min read
Policy as code that engineers do not route aroundA policy engine that blocks a deploy at four in the afternoon without explaining what to change does not produce compliance. It produces a workaround, and you will not hear about it.Read itObservability3 min read
Prometheus, Mimir, Loki and Tempo: what each is actually forThese get bundled into a single recommendation, which hides the fact that they solve four different problems and you probably do not need all of them yet.Read itKubernetes at Scale3 min read
A hundred clusters: what breaks that does not break at tenNothing dramatic happens between ten clusters and a hundred. Things that were merely tedious become impossible, and they do it gradually enough that nobody notices the threshold.Read itFinOps2 min read
Rightsizing is the last step, not the firstRightsizing is where most cost programmes begin, because tooling makes it easy to begin there. It is close to the worst place to start, and the reason is arithmetic.Read itKubernetes at Scale3 min read
Namespace, cluster or account? The isolation decision, costedThree boundaries, three very different running costs. The technical comparison is well covered elsewhere; what is usually missing is what each one costs you every month afterwards.Read it


