Writing March 11, 2017 · 4 min read

The Fallout From the S3 Outage


What a week. I’ve been hearing from all sorts of people (customers, students, co-workers) about the S3 outage - some even referred to it as the #S3apocalypse.

I’ll let you all in on something. It’s all FUD (Fear, Uncertainty, and Doubt). Every single person who is writing articles, or tweeting about the effects of the outage is using the S3 outage to try and sell you something - myself included. They are telling you that you need multi-cloud, that you need to go back on-premise (because we all know this would have never happened on premise right?) or maybe even switch cloud providers.

Here’s what I’m going to sell you on - a rational response to an outage.

Before doing anything, ask yourself this one simple question. Does my application need to be architected in such a way that it can weather a region-wide failure without downtime?

If the answer to that question is ‘yes’ and you suffered an outage, then you have some work to do. If the answer is no (and I suspect that if you’re truthful with yourself, the vast majority of you will fall into this camp), then take a deep breath and relax. Now is not the time to make rash decisions because your co-workers are making you the butt of their jokes, or the boss is breathing down your neck because they told you public cloud was not the right way to go - FYI, they are wrong.

Technology fails - always has, always will. What separates organizations from each other is how well prepared they are for when the failures occur.

Here are my tips for those of you who suffered an outage, but don’t require 99.999% uptime

Before we talk about anything, let’s talk about your Recovery Time Objective (RTO) and your Recovery Point Objective (RPO). You have those defined right? Having these two objectives established is the first step in building your DR strategy. If you don’t have them defined, start with something simple and develop a plan to improve them over time - Continuous Service Improvement in action - I knew that ITIL certification would come in handy.

OK, so now that you have your RTO and RPO defined, what’s next?

Let’s talk about your DR plan. You have one, right? Grab the DR binder, dust it off and read it - the entire thing (twice). Now that you’ve read the DR plan and have RTO and RPO objectives documented go ahead and make adjustments to your plan which will let you achieve both your documented RTO and RPO goals.

I know what you’re thinking, when are we going to talk technology? I love technology - DR not so much. How well did that serve you this week? One more thing and we’ll get to some technology - don’t worry.

Now that we have defined our RTO and RPO, and updated (or created) our DR strategy, let’s talk about testing. The plan is only half of the equation; we need to develop a method of testing our DR strategy - regularly, not just one weekend a year.

Spend the next while building a testing plan and discuss this plan with the other stakeholders in your organization. Once you’ve come up with something you all like start testing it.

Here’s where the technology part comes in. Build notification systems, and automated responses when you detect failures. One caveat, don’t just stand up another virtual instance (EC2), use managed services and serverless solutions to build your system. Typically, they have high-availability built in and are more cost-effective than standing up more EC2 instances - we hate OS’ remember?

Now the (kind of) fun part, test - test your system every opportunity you get and use this testing to fine tune your DR plan, RTO and RPO objectives. Here’s the super secret tip - never stop testing!

Remember you can treat each application individually. For those that do require higher availability consider using a multi-region deployment strategy. Keep in mind this will raise the cost and complexity of the solution.

There you have it, my rational response to the S3 outage. Notice we talked technology for maybe two or three sentences? When technology fails we all scramble for a solution based on technology - the technology is reliable; it is our processes that are fallible, build them stronger!

All writing Start a conversation