Came into work at the hospital today after a 3-day weekend to find a bunch of emails from our IT department about which programs that the doctors use are down as a result of this outage.
My previous career was in IT. I remember having a heated argument with management about this scenario with a hospital who was our client. I got them to self host their active directory and backup using Azure. Hospitals should be able to island themselves if there’s a major outage. They do it with electricity, Internet/Phone and everything else.
Most had their own baskets and then ditched them, fairly recently too. Unless a company was founded in the last 10-15 years it probably wasn’t cloud native. And 9 times out of ten they had worse up times and more security vulnerabilities than AWS. Plus for most of the companies being down for a day is an inconvenience but it doesn’t kill anyone. No one dies for lack for Slack. So why put in all the extra effort making yourself more robust if it’s probably fine?
You are mistaken... aws does not have a monopoly. They have like 30% of the market. The real lesson to take from this is to add distributed redundancy to your networks. Its more complicated and expensive but helps with outages like this.
So, the AWS outage got me curious, so I decided to check on pricing. For a server comparable to the one I have in my closet, with the same amount of storage, it would cost more than the entire hardware cost of my server, 10GbE networking gear and redundant storage drives included, and the electricity to run it, in ONE FUCKING MONTH. I could buy a whole-ass new server EVERY MONTH and it would be cheaper than using AWS.
Oops, we outsourced the backbone of what was supposed to be a decentralized system to one private company that gets to break things for everyone. I say "we" cause I've built a half dozen systems on it, and they all broke too. Oops.
Also fuck trying to move everything to a subscription model. My copy of photoshop may be coming up on 20 years old but it doesn’t stop working if my computer can’t validate a subscription to a remote server
Sadly incompetence/software glitches are harder to track than motivated threat actors. It still happens though, its unlikely to be the cause of this though.
as the world of software gets more complicated and inundated with AI built functions that are only half understood I think we can expect this to be a newer normal
Not using them exclusively. Amazon doesn’t exclusively use theirnown services either. Never ever bet on one horse as the saying goes. It’s better to function at 50% than be out 100%.
I know, just being humorous. Customers often require second sourcing. Often what happens these days is you have multiple locations that have synchronized datastore’s if they lose synchronization it can cause chaos. Its bad enough when it happens with a single service like what happened to Github a few years ago but a full provider would be insane.
Yeah, this isn't an issue of "monopolies" its an issue of all these companies leaving the default setting of "US-East-1" when spinning up their vms and services, and putting zero thought or budget into disaster recovery.
A lot of web services use multiple hosts. The bigger issue is data integrity. You have to keep your data synchronized across services, sometimes a small outage can cause a cascade. Like say you edited something on a amazon database and it doesn’t show up on azure and you edit something that shows up on azure that does’t show up on amazon. You got a big problem, both databases have valid but asymmetric data. This has caused a lot of the big outages. Monopolies still suck though
Databases have ways to handle these kinds of problems it just takes a while if your database flags a problem and needs to verify things. Its funny how people think a few hours of down time is the system failing. Sometimes down time is better than bad data.
Yeah they can sync as long as both instances are still talking to each other. If they aren't, it becomes a bigger issue the longer they are out of sync.
For service providers like AWS, a few minutes of downtime is considered catastrophic if it impacts customers. This list is mostly just the services consumers see, but there were also impacts to B2B services which just cascade the problems.
The common 99.999 goal only equates to about 5 minutes of downtime a year. So yeah, hours are really bad.
All CEOs cared about was that fat bonus check after they could show shareholders "cost savings" by outsourcing such things...it doesn't matter to them they're putting the reputation of the brand online by putting all their eggs in someone else's basket, they've got golden parachutes while the rest of us get fucked.
I finally got it working on desktop but it took some doing. Most days I do one language on my phone and a second one on my desktop. Yesterday, the desktop was slower than usual, and it behaved a little oddly, but I was able to complete both of my language lessons. :)
I think I did my lesson _before_ it hit. Since it was limited to the default region us-east-1, it exposed not only a single point of failure at AWS, but also lack of diligence from some big companies in not having region redundancy.
I actually did get my lesson done, but this morning I found out that my perfect streak had been broken. No big deal, but evidence of malfunction just the same. Duo often seems to get confused about streaks when I switch between phone and desktop.
Many years ago a company I worked for was an early adopter of AWS. We had built a small environment with a few web servers, and started with one domain controller. One day (Sunday, of course) the domain controller disappeared before the 2nd was live. Amazon had taken down part of their infrastructure without warning and all of our authentication was toast because redundancy and failover costs extra. Relocating the server was minor but it took time to figure out the cause.
"The Cloud" is just Somebody Else's computer, and your business is going to have a bad time when that computer goes down because fixing it is out of your control.
They might provide better security, but they are a MUCH bigger target for hackers, which does not equate to more up time and more overall security. But the people in charge like buzz words and firing employees.
For the vast majority of businesses "The Cloud" is good enough and when it's not just have a backup. Like how hospitals have emergency generators but still use city power.
I remember 9/11, and how some companies had the foresight to have what's called a Dark NOC. They spent the money to have a physically separate location with servers containing a mirror of their data, plus space for employees desks, and PCs. The NOC (network operations center) would be "dark" (servers doing backups, but nothing else) until it was needed. They could walk in, turn on lights and PCs, and get to work. 1/?
The network at the NOC was a copy of their normal location, so employees could log in as if they were at the normal office, and everything was there: emails, SharePoints, databases, etc. All mirrored automatically right up until the towers went down. They were able to shift and keep working. The companies that DIDN'T do this struggled to survive, as all their data went POOF. Many didn't make it. 2/2
Yeah, for sure. The only difference is that it is (theoretically) more cost effective to have all those "somebody else's computer" in one place and sharing resources.
But yeah -- I work for a tech company. They use a third party cloud provider, even though it's something that would easily fit within their own product definition.
While true, it very likely goes down less frequently than it would if it were running on YOUR computers. And big businesses an sometimes get SLAs, so they get paid if the hosting provider fucks up for more than (for example) 0.0005% of a year.
Agreed. Many of our hospital are in Azure or AWS and it’s been so much easier for me. Hospital CTOs often don’t know how to run a proper datacenter and it totally makes my job so much harder.
Having a hosting provider is also "somebodye else's computer". Whether you have remote access to an OS running on bare metal or a VM doesn't really make a difference, here.
I can do something about it when it goes down, unlike AWS where you simply have to twiddle your thumbs up your ass until they decide to do something about it.
The story says "about 1000 sites" and then lists about 50. Do you think the other 900+ sites get the same response when these things go down? Do you think they have a person to call to find out what's going on? The important point is *big businesses* can get money back, everyone else gets fucked. And it was down for about 4 hours, 0.05% of a year. For one outage.
For your questions, I have no idea what the answer is to either of them. Other people's contracts are not within my knowledge. I'm not sure what point you're making about the outage duration. Yeah, it is 0.05%, meaning that any company with a higher than 99.95% uptime SLA will be getting compensation, as will any with a lesser guarantee if there were other outages this year.
But other than that... I mean, yeah, that's a risk you take, but for anyone that _isn't_ a big business, it's a hell of a lot easier, less risky, and/or cheaper to use some cloud hosting than it is to build your own, multiple datacenters with redundant power in different regions and to manage the personnel to maintain them. I don't have a hundred million dollars to blow on a personal project, for example.
My databases are reliable as fuck tyvm. I don’t run hundreds of PTs of synchronized data across multiple locations like cloud providers do. It’s not uncommon that they get a data integrity issues causing chaos like this. Its usually a single service but I don’t see why it couldn’t happen across a full provider.
do you have fire and natural disaster protection? as for redundancy. how many Internet links, generators, cooling unitsa? then there's multi-site redundancy
Colo literally has Fire and natural disaster protection built in. Feeds for both legs of power from two seperate power companies. Link has failover to second data provider. All of our hardware has redundancy. Networking/network appliances. Servers and services are both redundant. Yep. All of that you mentioned plus some. AND we regularly test failover.
That's vastly oversimplified. The full ramifications of abdication of control are never contemplated and the SLA never covers losses. AWS is not going to compensate for lost revenue, the whole "cloud" business model would fail under such liability. Rarely considered is exposure of data and the target value centralized data presents.
Never put critical or private data on anyone else's hardware, ever. Hardware rental services should be liable for data and revenue loss.
What part? They have liability for your lost revenue should their service fail? They have liability for your liabilities to your customers should their infrastructure be breached or compromised; whether by malice, incompetence, or act of god; whether by internal or external actions; all harms; immediate, projected, and goodwill; shall be made whole?
Career server engineer here. While in principle I agree on self-hosting your hardware, the world doesn't work like that. It's WAY too expensive to do so when starting out, and as much as I hate to admit it, AWS is in practice far too reliable and secure (when you configure correctly) to not use it at its current price. Lastly, if an SLA doesn't cover losses, what are they for? AWS will in fact lose money to several of their impacted clients.
I'd be very interested to see what the agreed remedies are for loss of utility. I have only negotiated what are in effect consumer click agreements, and I have run services on various cloud providers, even in my own startups because the usual deferring capital expenses and focusing burn on code. I did not host any customer data on other people's hardware nor any company confidential data. I'm being forced to weaken security with cloud enclaves to comply with CMMC 2.0 now, sucks....>
But back to AWS etc: if they are willing to accept operational liability, how do they bill against risk? Does a customer have to share their revenue per minute and then pay to insure against that? Do you have to incorporate the downtime risk in business continuity insurance? What if it is a life and death service, such as image stream processing for vehicles or remote surgery?
In my experience, SLAs are for upper management and senior leadership to pretend that they are negotiating something useful and to have a number to report against.
It would be very interesting to see a typical HIPAA compliance term. That came up recently in a related way considering remote processing of patient self-collected image data. The concern being ensuring both compliant and genuinely secure 3rd party processing of such data, which meant it could not be e2e on owned, premised hardware.
Corps should use a combination of on-site and cloud capacity, to account for sudden surges. If the surge becomes the norm, build out more on-site capacity. If it fades, great.
Instead, corps have put EVERYTHING in the "cloud" and that makes them vulnerable when one moron at said Cloud Computing Company trips over a network cable or fucks up a DNS entry.
Load balancing in a hybrid environment is really tricky and has to be done from the cloud side if you're actually dealing with meaningful user counts. What you're saying is smart for segmentation of internal applications with very predictable usage/growth patterns, but anything public facing pretty much needs to be cloud-based just to provide the UX we all expect. People tend to bounce in the time it takes to route to your on-prem infrastructure's gateway/load balancer and on to the real server.
If users leave in the time it takes to resolve to your on-site servers, you are doing something VERY wrong with your on-site servers.
Fiber optic connections are the norm now, and having your office in a major city should make it indistinguishable, in terms of response time, from AWS or MS Azure. If it's slower, you fucked up.
Those days are why cloud-first deployments are best practice for anything user facing. If your site goes down during a surge, what % of the potential business do you think the overload cost your employer? The other side of the coin, it's really hard to attract and retain people good with physical hardware management. When something does break, I'd rather it be AWS's best and brightest on it rather than Steve who got hired because his uncle used to work on the marketing team.
OceanCitySoul
Bell Telephone wasn't close to this big when they were made to split into the Baby Bells. WTF?
PoliticalWanderer
Came into work at the hospital today after a 3-day weekend to find a bunch of emails from our IT department about which programs that the doctors use are down as a result of this outage.
McKittyNuts
My previous career was in IT. I remember having a heated argument with management about this scenario with a hospital who was our client. I got them to self host their active directory and backup using Azure. Hospitals should be able to island themselves if there’s a major outage. They do it with electricity, Internet/Phone and everything else.
B2SteakSauce
Seems like a “all the eggs in one basket” situation. Maybe company’s need their own baskets?
Prometheusblu
Most had their own baskets and then ditched them, fairly recently too. Unless a company was founded in the last 10-15 years it probably wasn’t cloud native. And 9 times out of ten they had worse up times and more security vulnerabilities than AWS. Plus for most of the companies being down for a day is an inconvenience but it doesn’t kill anyone. No one dies for lack for Slack. So why put in all the extra effort making yourself more robust if it’s probably fine?
Orlandonuts
That's why I can't sign on to HBO Max
IAmNotNSAsodonotbeparanoid
They haven't learned from the lessons when the Cylons hacked the Colonial Defense mainframe
Datageek11
You are mistaken... aws does not have a monopoly. They have like 30% of the market. The real lesson to take from this is to add distributed redundancy to your networks. Its more complicated and expensive but helps with outages like this.
eurorubio3000
So why the heck did I still get 200 Teams calls that day? Cant even trust a service outage anymore.
Eldibs
So, the AWS outage got me curious, so I decided to check on pricing. For a server comparable to the one I have in my closet, with the same amount of storage, it would cost more than the entire hardware cost of my server, 10GbE networking gear and redundant storage drives included, and the electricity to run it, in ONE FUCKING MONTH. I could buy a whole-ass new server EVERY MONTH and it would be cheaper than using AWS.
z3253304
Heh
Stanistani
The dinky social media site I help run stayed up because I turned down the AWS 'free' offer a few years ago.
KevinStrexcorp
Oops, we outsourced the backbone of what was supposed to be a decentralized system to one private company that gets to break things for everyone. I say "we" cause I've built a half dozen systems on it, and they all broke too. Oops.
D3pleted
Electronic Arts was also affected, Battlefield 6 was unavailable. Reddit was also affected.
Darklinkinfinite
Also fuck trying to move everything to a subscription model. My copy of photoshop may be coming up on 20 years old but it doesn’t stop working if my computer can’t validate a subscription to a remote server
ArgentXero
Ironically, this means that workers have more impact in protesting than they realize.
If a small inconvenience can cause a major shutdown, think about how simply organizing in certain sectors can impact things.
gesel
Amazing how much can be solved by tossing a wooden shoe into the gears.
AFelineMassofEyes
Commenting to boost signal; this is a sound strategy
McKittyNuts
Sadly incompetence/software glitches are harder to track than motivated threat actors. It still happens though, its unlikely to be the cause of this though.
Firestar002
Plus it’s actually GETTING the people to DO the thing on X day.
NeverDownvoteMelBrooks
Outlook went down about an hour last week and AWS this week. Any bets out for next week's adventure?
TigerThong
as the world of software gets more complicated and inundated with AI built functions that are only half understood I think we can expect this to be a newer normal
McKittyNuts
Its somehow both surprising and not very surprising that Microsoft isn’t using their own services 😵
ToenailClippingsJar
Not using them exclusively. Amazon doesn’t exclusively use theirnown services either.
Never ever bet on one horse as the saying goes. It’s better to function at 50% than be out 100%.
McKittyNuts
I know, just being humorous. Customers often require second sourcing. Often what happens these days is you have multiple locations that have synchronized datastore’s if they lose synchronization it can cause chaos. Its bad enough when it happens with a single service like what happened to Github a few years ago but a full provider would be insane.
madjo
Cloudflare *fingers crossed*
downsouthfarm
This is great. It might teach people to host their own data.
blinkonceforyes
The outage was only us-east-1 region of aws. Billion dollar companies and they’re not using multi region on AWS. Sigh
HelpfulCorn
We're a Fortune 200 and we don't nor do we have a DR plan that's not laughable
blinkonceforyes
Sad
gumshoe99
The funny thing is this could be avoided if these companies would pay for regional balancing but instead chose to bind themselves to US East-1
HelpfulCorn
https://media1.giphy.com/media/v1.Y2lkPWE1NzM3M2U1a2wxcmN5aHhsZzZ5cHRobmVmb2szdmxldnJqNDhmemN3ajZmY21xaSZlcD12MV9naWZzX3NlYXJjaCZjdD1n/5xtDarmwsuR9sDRObyU/200w.webp
kris40k
Yeah, this isn't an issue of "monopolies" its an issue of all these companies leaving the default setting of "US-East-1" when spinning up their vms and services, and putting zero thought or budget into disaster recovery.
linkdk59
And I barely noticed.
tvstpq
Missing a lot of big names. Snowflake and Atlassian to start, just these two have taken so many orgs down with them.
Quebeker
Autodesk is not on the list but was impacted. I know revit didn’t work all day yesterday. And autocad had licence problems
iamgnat
That list seems to say an awful lot about Azure...
ralphmelish
AWS is an Amazon service.
McKittyNuts
No its owned by Alaska
ralphmelish
Or do you mean this because MS stuff is going down through an AWS outage?
iamgnat
And Microsoft is using it rather than their own competing service, Azure.
ralphmelish
I guess that was the point, "whoosh" to me.
MoralRectifier
It's more of a licensing thing that Microsoft allows customers to do, not something Microsoft does themselves. https://www.techtarget.com/searchvirtualdesktop/opinion/Amazon-WorkSpaces-finally-supports-Office-365-but-why-now
ToenailClippingsJar
Risk needs to be spread.
I’ll let you figure out how that works out with cloud providers okay?
McKittyNuts
A lot of web services use multiple hosts. The bigger issue is data integrity. You have to keep your data synchronized across services, sometimes a small outage can cause a cascade. Like say you edited something on a amazon database and it doesn’t show up on azure and you edit something that shows up on azure that does’t show up on amazon. You got a big problem, both databases have valid but asymmetric data. This has caused a lot of the big outages. Monopolies still suck though
Datageek11
Databases have ways to handle these kinds of problems it just takes a while if your database flags a problem and needs to verify things. Its funny how people think a few hours of down time is the system failing. Sometimes down time is better than bad data.
iamgnat
Yeah they can sync as long as both instances are still talking to each other. If they aren't, it becomes a bigger issue the longer they are out of sync.
For service providers like AWS, a few minutes of downtime is considered catastrophic if it impacts customers. This list is mostly just the services consumers see, but there were also impacts to B2B services which just cascade the problems.
The common 99.999 goal only equates to about 5 minutes of downtime a year. So yeah, hours are really bad.
McKittyNuts
I didn’t say data lose. A failure is downtime. Imagine how much money it costs when AWS fails?
loztriforce
All CEOs cared about was that fat bonus check after they could show shareholders "cost savings" by outsourcing such things...it doesn't matter to them they're putting the reputation of the brand online by putting all their eggs in someone else's basket, they've got golden parachutes while the rest of us get fucked.
PileOfWalthers
XBox was working fine, but I'm sure glad I took this week off.
PythonIndentAwwwYiss
No idea why xbox is on that list, Microsoft has its own datacenters and is a direct competitor to AWS
PileOfWalthers
(Company I work for is on that list, AND uses a bunch of services on that list.)
tinyfootprints
Is that why I couldn't use Duolingo yesterday?
faithydiesalot
Yes. They put up the maintenance break sign but it was an outage.
tinyfootprints
I finally got it working on desktop but it took some doing. Most days I do one language on my phone and a second one on my desktop. Yesterday, the desktop was slower than usual, and it behaved a little oddly, but I was able to complete both of my language lessons. :)
faithydiesalot
I think I did my lesson _before_ it hit. Since it was limited to the default region us-east-1, it exposed not only a single point of failure at AWS, but also lack of diligence from some big companies in not having region redundancy.
tinyfootprints
I actually did get my lesson done, but this morning I found out that my perfect streak had been broken. No big deal, but evidence of malfunction just the same. Duo often seems to get confused about streaks when I switch between phone and desktop.
faithydiesalot
one way to recover if midnight has passed is to change your timezone westward (say, LA, Honolulu or even Midway)
CyberWizard252
Many years ago a company I worked for was an early adopter of AWS. We had built a small environment with a few web servers, and started with one domain controller. One day (Sunday, of course) the domain controller disappeared before the 2nd was live. Amazon had taken down part of their infrastructure without warning and all of our authentication was toast because redundancy and failover costs extra. Relocating the server was minor but it took time to figure out the cause.
malachilenomade
I don't see Pornhub listed so
v
Targe0
I... I didn't know it had happened till after it was already over.
MisterLuminous
If anyone wants to declare war on "the"West", all they have to do is attack AWS and goodbye. This shows why decentralization is so important.
TuffyTDog
"The Cloud" is just Somebody Else's computer, and your business is going to have a bad time when that computer goes down because fixing it is out of your control.
KillingTlme
They might provide better security, but they are a MUCH bigger target for hackers, which does not equate to more up time and more overall security. But the people in charge like buzz words and firing employees.
AVerySillyPuppy
As far as I'm aware none of the major cloud providers have ever been hacked at the infrastructure level. Breaches are always user misconfig.
Bobbobbobobbananafanafobob
Don't underestimate the power of "well are all down too" on small and medium businesses.
wandermanspacebot
Ha! Fixing my own computer is also out of my control. Checkmate
barnwolf
For the vast majority of businesses "The Cloud" is good enough and when it's not just have a backup. Like how hospitals have emergency generators but still use city power.
tohmik
Yeah, about these emergency generators…
barnwolf
?
Hypothesist
its always a bad time when your own computer goes down.
CallMeCourierSix
I remember 9/11, and how some companies had the foresight to have what's called a Dark NOC. They spent the money to have a physically separate location with servers containing a mirror of their data, plus space for employees desks, and PCs. The NOC (network operations center) would be "dark" (servers doing backups, but nothing else) until it was needed. They could walk in, turn on lights and PCs, and get to work. 1/?
CallMeCourierSix
The network at the NOC was a copy of their normal location, so employees could log in as if they were at the normal office, and everything was there: emails, SharePoints, databases, etc. All mirrored automatically right up until the towers went down. They were able to shift and keep working. The companies that DIDN'T do this struggled to survive, as all their data went POOF. Many didn't make it. 2/2
ByThePowerOfSCIENCE
even superficial familiarity with setting up -or- maintaining a data center (let alone redundant ones) dispels that truism
Rockafella83
Office 365 ? Though they were on Azure
RevengeIsIceCream
/gallery/xkcd-cloud-hxSj3yw
UprootedGrunt
Yeah, for sure. The only difference is that it is (theoretically) more cost effective to have all those "somebody else's computer" in one place and sharing resources.
UprootedGrunt
But yeah -- I work for a tech company. They use a third party cloud provider, even though it's something that would easily fit within their own product definition.
n0n53n53
not to mention the redundancy which "should" help keep these type of issues at a minimum.
Corrodias
While true, it very likely goes down less frequently than it would if it were running on YOUR computers. And big businesses an sometimes get SLAs, so they get paid if the hosting provider fucks up for more than (for example) 0.0005% of a year.
Boatsntoes
Agreed. Many of our hospital are in Azure or AWS and it’s been so much easier for me. Hospital CTOs often don’t know how to run a proper datacenter and it totally makes my job so much harder.
malicart
I get that from my hosting provider and I also have full control over the server, AWS is a marketing scam and an expensive one at that.
Corrodias
Having a hosting provider is also "somebodye else's computer". Whether you have remote access to an OS running on bare metal or a VM doesn't really make a difference, here.
malicart
I can do something about it when it goes down, unlike AWS where you simply have to twiddle your thumbs up your ass until they decide to do something about it.
Corrodias
Well I don't think you're going to do much about it if they misconfigure the network your server is connected to, as happened in this outage...?
Evi1Gav
The story says "about 1000 sites" and then lists about 50. Do you think the other 900+ sites get the same response when these things go down? Do you think they have a person to call to find out what's going on? The important point is *big businesses* can get money back, everyone else gets fucked. And it was down for about 4 hours, 0.05% of a year. For one outage.
Corrodias
For your questions, I have no idea what the answer is to either of them. Other people's contracts are not within my knowledge. I'm not sure what point you're making about the outage duration. Yeah, it is 0.05%, meaning that any company with a higher than 99.95% uptime SLA will be getting compensation, as will any with a lesser guarantee if there were other outages this year.
Corrodias
But other than that... I mean, yeah, that's a risk you take, but for anyone that _isn't_ a big business, it's a hell of a lot easier, less risky, and/or cheaper to use some cloud hosting than it is to build your own, multiple datacenters with redundant power in different regions and to manage the personnel to maintain them. I don't have a hundred million dollars to blow on a personal project, for example.
McKittyNuts
My databases are reliable as fuck tyvm. I don’t run hundreds of PTs of synchronized data across multiple locations like cloud providers do. It’s not uncommon that they get a data integrity issues causing chaos like this. Its usually a single service but I don’t see why it couldn’t happen across a full provider.
Corrodias
You aren't hosting in multiple regions? So one BGP error could take you out entirely until it's fixed (just like this story)?
ByThePowerOfSCIENCE
do you have fire and natural disaster protection? as for redundancy. how many Internet links, generators, cooling unitsa? then there's multi-site redundancy
CatWithHands
Colo literally has Fire and natural disaster protection built in. Feeds for both legs of power from two seperate power companies. Link has failover to second data provider. All of our hardware has redundancy. Networking/network appliances. Servers and services are both redundant. Yep. All of that you mentioned plus some. AND we regularly test failover.
Corrodias
Now that's a company willing to invest a lot in its own infrastructure, and I applaud that.
gesel
That's vastly oversimplified. The full ramifications of abdication of control are never contemplated and the SLA never covers losses. AWS is not going to compensate for lost revenue, the whole "cloud" business model would fail under such liability. Rarely considered is exposure of data and the target value centralized data presents.
Never put critical or private data on anyone else's hardware, ever. Hardware rental services should be liable for data and revenue loss.
agonarch
Oof, clearly the cloud guys need to spin that aspect too then, how about "It's not a just a data breach, it's a surprise bonus backup!"
MathiasTolerain
It’s impossible to find an IOS mobile password manager that ISN’T a cloud based subscription service and it drives me BONKERS.
dlshark
Bitwarden
Boatsntoes
You can put that in a contract. I put it in all of my contracts with AWS.
gesel
What part? They have liability for your lost revenue should their service fail? They have liability for your liabilities to your customers should their infrastructure be breached or compromised; whether by malice, incompetence, or act of god; whether by internal or external actions; all harms; immediate, projected, and goodwill; shall be made whole?
Nifty255
Career server engineer here. While in principle I agree on self-hosting your hardware, the world doesn't work like that. It's WAY too expensive to do so when starting out, and as much as I hate to admit it, AWS is in practice far too reliable and secure (when you configure correctly) to not use it at its current price. Lastly, if an SLA doesn't cover losses, what are they for? AWS will in fact lose money to several of their impacted clients.
gesel
I'd be very interested to see what the agreed remedies are for loss of utility. I have only negotiated what are in effect consumer click agreements, and I have run services on various cloud providers, even in my own startups because the usual deferring capital expenses and focusing burn on code. I did not host any customer data on other people's hardware nor any company confidential data. I'm being forced to weaken security with cloud enclaves to comply with CMMC 2.0 now, sucks....>
gesel
But back to AWS etc: if they are willing to accept operational liability, how do they bill against risk? Does a customer have to share their revenue per minute and then pay to insure against that? Do you have to incorporate the downtime risk in business continuity insurance? What if it is a life and death service, such as image stream processing for vehicles or remote surgery?
Boatsntoes
Also interestingly, I’m seeing much better performance on azure blob vs S3.
MathiasTolerain
In my experience, SLAs are for upper management and senior leadership to pretend that they are negotiating something useful and to have a number to report against.
AllTheKitties
Does your experience include contract law?
Boatsntoes
This guy gets it. Hello fellow engineer. I use AWS for healthcare at tons of hospitals
gesel
It would be very interesting to see a typical HIPAA compliance term. That came up recently in a related way considering remote processing of patient self-collected image data. The concern being ensuring both compliant and genuinely secure 3rd party processing of such data, which meant it could not be e2e on owned, premised hardware.
AVerySillyPuppy
Sure, but without the cloud good luck scaling your service when your user count triples overnight because of a celebrity social media post.
CallMeCourierSix
Corps should use a combination of on-site and cloud capacity, to account for sudden surges. If the surge becomes the norm, build out more on-site capacity. If it fades, great.
Instead, corps have put EVERYTHING in the "cloud" and that makes them vulnerable when one moron at said Cloud Computing Company trips over a network cable or fucks up a DNS entry.
jaqque
The only use case for cloud that has made sense to me is a hybrid approach m: on prem for base load, cloud to scale to meet surge demand
AVerySillyPuppy
Load balancing in a hybrid environment is really tricky and has to be done from the cloud side if you're actually dealing with meaningful user counts. What you're saying is smart for segmentation of internal applications with very predictable usage/growth patterns, but anything public facing pretty much needs to be cloud-based just to provide the UX we all expect. People tend to bounce in the time it takes to route to your on-prem infrastructure's gateway/load balancer and on to the real server.
CallMeCourierSix
If users leave in the time it takes to resolve to your on-site servers, you are doing something VERY wrong with your on-site servers.
Fiber optic connections are the norm now, and having your office in a major city should make it indistinguishable, in terms of response time, from AWS or MS Azure. If it's slower, you fucked up.
jaqque
I’ll trust you that. We kept everything on prem and rarely had a surge problem. Hitting the front page of Reddit - oh, that was a day.
AVerySillyPuppy
Those days are why cloud-first deployments are best practice for anything user facing. If your site goes down during a surge, what % of the potential business do you think the overload cost your employer? The other side of the coin, it's really hard to attract and retain people good with physical hardware management. When something does break, I'd rather it be AWS's best and brightest on it rather than Steve who got hired because his uncle used to work on the marketing team.