Operational Crises I Have Faced Running GitOps for Cloud…
The operational crises I keep running into when I manage cloud infrastructure with GitOps — and the patterns that have helped me avoid the worst of them.
1045 posts · Page 31/44 · 721-744 showing
Search runs on the posts loaded on this page. Use category or pagination for the deep archive.
The operational crises I keep running into when I manage cloud infrastructure with GitOps — and the patterns that have helped me avoid the worst of them.
Treating configuration like a product: feature flags, parameter store, schema, approval flow, audit log, and rollback discipline.
Take a deep look at Terraform plan's surprise resource deletions and the strategies for protecting your automation pipelines from these kinds of failures.
Discover the data consistency problems you run into when migrating from a monolithic database to a microservice architecture, plus solutions, in this…
Learn how to secure network traffic between pods using Kubernetes Network Policies. A from-A-to-Z guide with detailed examples for Network…
A deep look at database provisioning mistakes I keep running into on cloud platforms, the symptoms they cause, and the fixes that actually hold up in…
Why concurrent deployments matter on cloud-native platforms, and the role stress testing plays in keeping them from becoming incidents.
An in-depth look at the nature of intermittent errors in distributed systems, the stress they place on teams, and strategies for dealing with these 'ghosts'...
An exploration of the fear that comes with making the first change to a critical system and how automation makes the process easier.
A real outage story driven by unscalable cloud architecture, and the lessons we can take away from it.
Discover the critical role of leadership in architectural decision-making during crises in distributed systems, plus the strategies that work.
A deep look at vendor lock-in risk in database choices, the visible and hidden costs of migration, and the strategies you can use to avoid these traps…
Explore the limits of automation and the indispensable role that the human touch, critical thinking, and empathy play in crisis management when systems…
Discover the challenges that technical debt and legacy systems bring, plus the human cost behind them. Save your career and your projects with practical…
In distributed systems, badly designed retries make outages worse. An approach to limiting damage with timeout budgets, retry budgets, and backpressure.
An approach to building secure B2B file exchange using an object storage dropzone, short-lived access, and audit trails — instead of an SFTP bottleneck.
A real war story about an outage day in cloud architecture and why DNS failover strategies matter.
An in-depth look, from Mustafa Erbay's perspective, at the production issues caused by hidden dependencies in distributed systems and the 'backfire battles'…
The benefits of automation are undeniable, yet confronting its overlooked shadows and battling its unexpected side effects matter just as much…
A look at the security benefits of micro-segmentation, the unexpected network outages it triggers when applied incorrectly, the root causes, and how to fix…
Learn about the cache stampede problems that Origin Shield can cause in Cloud Native CDNs, and how to solve them.
How hidden dependencies in systems lead to unexpected production issues, and the architectural lessons we need to take away to reduce those risks…
Discover the journey from the engineer's nightmare of Pager Burnout to amplified system resilience and sustainability through SRE principles.
Explore the hidden traps and possible failure modes inside the auto-renewal process of certificates that are vital to digital security. Don't let your security…