It’s almost like monopolies are bad or something…

Oct 21, 2025 12:57 PM

McKittyNuts

Views

23956

Likes

492

Dislikes

7

Bell Telephone wasn't close to this big when they were made to split into the Baby Bells. WTF?

9 months ago | Likes 1 Dislikes 0

Came into work at the hospital today after a 3-day weekend to find a bunch of emails from our IT department about which programs that the doctors use are down as a result of this outage.

9 months ago | Likes 65 Dislikes 0

My previous career was in IT. I remember having a heated argument with management about this scenario with a hospital who was our client. I got them to self host their active directory and backup using Azure. Hospitals should be able to island themselves if there’s a major outage. They do it with electricity, Internet/Phone and everything else.

9 months ago | Likes 44 Dislikes 0

Seems like a “all the eggs in one basket” situation. Maybe company’s need their own baskets?

9 months ago | Likes 11 Dislikes 0

Most had their own baskets and then ditched them, fairly recently too. Unless a company was founded in the last 10-15 years it probably wasn’t cloud native. And 9 times out of ten they had worse up times and more security vulnerabilities than AWS. Plus for most of the companies being down for a day is an inconvenience but it doesn’t kill anyone. No one dies for lack for Slack. So why put in all the extra effort making yourself more robust if it’s probably fine?

9 months ago | Likes 1 Dislikes 0

That's why I can't sign on to HBO Max

9 months ago | Likes 3 Dislikes 1

They haven't learned from the lessons when the Cylons hacked the Colonial Defense mainframe

9 months ago | Likes 2 Dislikes 1

You are mistaken... aws does not have a monopoly. They have like 30% of the market. The real lesson to take from this is to add distributed redundancy to your networks. Its more complicated and expensive but helps with outages like this.

9 months ago | Likes 6 Dislikes 0

So why the heck did I still get 200 Teams calls that day? Cant even trust a service outage anymore.

9 months ago | Likes 1 Dislikes 0

So, the AWS outage got me curious, so I decided to check on pricing. For a server comparable to the one I have in my closet, with the same amount of storage, it would cost more than the entire hardware cost of my server, 10GbE networking gear and redundant storage drives included, and the electricity to run it, in ONE FUCKING MONTH. I could buy a whole-ass new server EVERY MONTH and it would be cheaper than using AWS.

9 months ago | Likes 1 Dislikes 0

Heh

9 months ago | Likes 1 Dislikes 0

The dinky social media site I help run stayed up because I turned down the AWS 'free' offer a few years ago.

9 months ago | Likes 1 Dislikes 0

Oops, we outsourced the backbone of what was supposed to be a decentralized system to one private company that gets to break things for everyone. I say "we" cause I've built a half dozen systems on it, and they all broke too. Oops.

9 months ago | Likes 1 Dislikes 0

Electronic Arts was also affected, Battlefield 6 was unavailable. Reddit was also affected.

9 months ago | Likes 1 Dislikes 0

Also fuck trying to move everything to a subscription model. My copy of photoshop may be coming up on 20 years old but it doesn’t stop working if my computer can’t validate a subscription to a remote server

9 months ago | Likes 1 Dislikes 0

Ironically, this means that workers have more impact in protesting than they realize.

If a small inconvenience can cause a major shutdown, think about how simply organizing in certain sectors can impact things.

9 months ago | Likes 48 Dislikes 1

Amazing how much can be solved by tossing a wooden shoe into the gears.

9 months ago | Likes 4 Dislikes 1

Commenting to boost signal; this is a sound strategy

9 months ago | Likes 10 Dislikes 1

Sadly incompetence/software glitches are harder to track than motivated threat actors. It still happens though, its unlikely to be the cause of this though.

9 months ago | Likes 12 Dislikes 0

Plus it’s actually GETTING the people to DO the thing on X day.

9 months ago | Likes 1 Dislikes 0

Outlook went down about an hour last week and AWS this week. Any bets out for next week's adventure?

9 months ago | Likes 21 Dislikes 0

as the world of software gets more complicated and inundated with AI built functions that are only half understood I think we can expect this to be a newer normal

9 months ago | Likes 1 Dislikes 0

Its somehow both surprising and not very surprising that Microsoft isn’t using their own services 😵

9 months ago | Likes 11 Dislikes 0

Not using them exclusively. Amazon doesn’t exclusively use theirnown services either.
Never ever bet on one horse as the saying goes. It’s better to function at 50% than be out 100%.

9 months ago | Likes 4 Dislikes 0

I know, just being humorous. Customers often require second sourcing. Often what happens these days is you have multiple locations that have synchronized datastore’s if they lose synchronization it can cause chaos. Its bad enough when it happens with a single service like what happened to Github a few years ago but a full provider would be insane.

9 months ago | Likes 2 Dislikes 0

Cloudflare *fingers crossed*

9 months ago | Likes 4 Dislikes 0

This is great. It might teach people to host their own data.

9 months ago | Likes 1 Dislikes 0

The outage was only us-east-1 region of aws. Billion dollar companies and they’re not using multi region on AWS. Sigh

9 months ago | Likes 4 Dislikes 0

We're a Fortune 200 and we don't nor do we have a DR plan that's not laughable

9 months ago | Likes 1 Dislikes 0

Sad

9 months ago | Likes 1 Dislikes 0

The funny thing is this could be avoided if these companies would pay for regional balancing but instead chose to bind themselves to US East-1

9 months ago | Likes 8 Dislikes 0

Yeah, this isn't an issue of "monopolies" its an issue of all these companies leaving the default setting of "US-East-1" when spinning up their vms and services, and putting zero thought or budget into disaster recovery.

9 months ago | Likes 6 Dislikes 0

And I barely noticed.

9 months ago | Likes 1 Dislikes 0

Missing a lot of big names. Snowflake and Atlassian to start, just these two have taken so many orgs down with them.

9 months ago | Likes 2 Dislikes 0

Autodesk is not on the list but was impacted. I know revit didn’t work all day yesterday. And autocad had licence problems

9 months ago | Likes 2 Dislikes 0

That list seems to say an awful lot about Azure...

9 months ago | Likes 26 Dislikes 1

AWS is an Amazon service.

9 months ago | Likes 10 Dislikes 3

No its owned by Alaska

9 months ago | Likes 3 Dislikes 1

Or do you mean this because MS stuff is going down through an AWS outage?

9 months ago | Likes 3 Dislikes 0

And Microsoft is using it rather than their own competing service, Azure.

9 months ago | Likes 16 Dislikes 1

I guess that was the point, "whoosh" to me.

9 months ago | Likes 15 Dislikes 0

It's more of a licensing thing that Microsoft allows customers to do, not something Microsoft does themselves. https://www.techtarget.com/searchvirtualdesktop/opinion/Amazon-WorkSpaces-finally-supports-Office-365-but-why-now

9 months ago | Likes 7 Dislikes 0

Risk needs to be spread.
I’ll let you figure out how that works out with cloud providers okay?

9 months ago | Likes 3 Dislikes 1

A lot of web services use multiple hosts. The bigger issue is data integrity. You have to keep your data synchronized across services, sometimes a small outage can cause a cascade. Like say you edited something on a amazon database and it doesn’t show up on azure and you edit something that shows up on azure that does’t show up on amazon. You got a big problem, both databases have valid but asymmetric data. This has caused a lot of the big outages. Monopolies still suck though

9 months ago | Likes 3 Dislikes 0

Databases have ways to handle these kinds of problems it just takes a while if your database flags a problem and needs to verify things. Its funny how people think a few hours of down time is the system failing. Sometimes down time is better than bad data.

9 months ago | Likes 2 Dislikes 1

Yeah they can sync as long as both instances are still talking to each other. If they aren't, it becomes a bigger issue the longer they are out of sync.

For service providers like AWS, a few minutes of downtime is considered catastrophic if it impacts customers. This list is mostly just the services consumers see, but there were also impacts to B2B services which just cascade the problems.

The common 99.999 goal only equates to about 5 minutes of downtime a year. So yeah, hours are really bad.

9 months ago | Likes 1 Dislikes 0

I didn’t say data lose. A failure is downtime. Imagine how much money it costs when AWS fails?

9 months ago | Likes 2 Dislikes 0

All CEOs cared about was that fat bonus check after they could show shareholders "cost savings" by outsourcing such things...it doesn't matter to them they're putting the reputation of the brand online by putting all their eggs in someone else's basket, they've got golden parachutes while the rest of us get fucked.

9 months ago | Likes 3 Dislikes 1

XBox was working fine, but I'm sure glad I took this week off.

9 months ago | Likes 5 Dislikes 0

No idea why xbox is on that list, Microsoft has its own datacenters and is a direct competitor to AWS

9 months ago | Likes 1 Dislikes 0

(Company I work for is on that list, AND uses a bunch of services on that list.)

9 months ago | Likes 4 Dislikes 0

Is that why I couldn't use Duolingo yesterday?

9 months ago | Likes 2 Dislikes 0

Yes. They put up the maintenance break sign but it was an outage.

9 months ago | Likes 1 Dislikes 0

I finally got it working on desktop but it took some doing. Most days I do one language on my phone and a second one on my desktop. Yesterday, the desktop was slower than usual, and it behaved a little oddly, but I was able to complete both of my language lessons. :)

9 months ago | Likes 2 Dislikes 0

I think I did my lesson _before_ it hit. Since it was limited to the default region us-east-1, it exposed not only a single point of failure at AWS, but also lack of diligence from some big companies in not having region redundancy.

9 months ago | Likes 2 Dislikes 0

I actually did get my lesson done, but this morning I found out that my perfect streak had been broken. No big deal, but evidence of malfunction just the same. Duo often seems to get confused about streaks when I switch between phone and desktop.

9 months ago | Likes 2 Dislikes 0

one way to recover if midnight has passed is to change your timezone westward (say, LA, Honolulu or even Midway)

9 months ago | Likes 1 Dislikes 0

Many years ago a company I worked for was an early adopter of AWS. We had built a small environment with a few web servers, and started with one domain controller. One day (Sunday, of course) the domain controller disappeared before the 2nd was live. Amazon had taken down part of their infrastructure without warning and all of our authentication was toast because redundancy and failover costs extra. Relocating the server was minor but it took time to figure out the cause.

9 months ago | Likes 4 Dislikes 0

I don't see Pornhub listed so v

9 months ago | Likes 3 Dislikes 0

I... I didn't know it had happened till after it was already over.

9 months ago | Likes 2 Dislikes 0

If anyone wants to declare war on "the"West", all they have to do is attack AWS and goodbye. This shows why decentralization is so important.

9 months ago | Likes 8 Dislikes 2

"The Cloud" is just Somebody Else's computer, and your business is going to have a bad time when that computer goes down because fixing it is out of your control.

9 months ago | Likes 230 Dislikes 3

They might provide better security, but they are a MUCH bigger target for hackers, which does not equate to more up time and more overall security. But the people in charge like buzz words and firing employees.

9 months ago | Likes 1 Dislikes 1

As far as I'm aware none of the major cloud providers have ever been hacked at the infrastructure level. Breaches are always user misconfig.

9 months ago | Likes 2 Dislikes 0

Don't underestimate the power of "well are all down too" on small and medium businesses.

9 months ago | Likes 1 Dislikes 0

Ha! Fixing my own computer is also out of my control. Checkmate

9 months ago | Likes 11 Dislikes 0

For the vast majority of businesses "The Cloud" is good enough and when it's not just have a backup. Like how hospitals have emergency generators but still use city power.

9 months ago | Likes 3 Dislikes 0

Yeah, about these emergency generators…

9 months ago | Likes 1 Dislikes 0

?

9 months ago | Likes 1 Dislikes 0

its always a bad time when your own computer goes down.

9 months ago | Likes 3 Dislikes 0

I remember 9/11, and how some companies had the foresight to have what's called a Dark NOC. They spent the money to have a physically separate location with servers containing a mirror of their data, plus space for employees desks, and PCs. The NOC (network operations center) would be "dark" (servers doing backups, but nothing else) until it was needed. They could walk in, turn on lights and PCs, and get to work. 1/?

9 months ago | Likes 1 Dislikes 0

The network at the NOC was a copy of their normal location, so employees could log in as if they were at the normal office, and everything was there: emails, SharePoints, databases, etc. All mirrored automatically right up until the towers went down. They were able to shift and keep working. The companies that DIDN'T do this struggled to survive, as all their data went POOF. Many didn't make it. 2/2

9 months ago | Likes 1 Dislikes 0

even superficial familiarity with setting up -or- maintaining a data center (let alone redundant ones) dispels that truism

9 months ago | Likes 2 Dislikes 0

Office 365 ? Though they were on Azure

9 months ago | Likes 1 Dislikes 0

Yeah, for sure. The only difference is that it is (theoretically) more cost effective to have all those "somebody else's computer" in one place and sharing resources.

9 months ago | Likes 2 Dislikes 0

But yeah -- I work for a tech company. They use a third party cloud provider, even though it's something that would easily fit within their own product definition.

9 months ago | Likes 1 Dislikes 0

not to mention the redundancy which "should" help keep these type of issues at a minimum.

9 months ago | Likes 2 Dislikes 0

While true, it very likely goes down less frequently than it would if it were running on YOUR computers. And big businesses an sometimes get SLAs, so they get paid if the hosting provider fucks up for more than (for example) 0.0005% of a year.

9 months ago | Likes 55 Dislikes 5

Agreed. Many of our hospital are in Azure or AWS and it’s been so much easier for me. Hospital CTOs often don’t know how to run a proper datacenter and it totally makes my job so much harder.

9 months ago | Likes 2 Dislikes 0

I get that from my hosting provider and I also have full control over the server, AWS is a marketing scam and an expensive one at that.

9 months ago | Likes 1 Dislikes 0

Having a hosting provider is also "somebodye else's computer". Whether you have remote access to an OS running on bare metal or a VM doesn't really make a difference, here.

9 months ago | Likes 1 Dislikes 0

I can do something about it when it goes down, unlike AWS where you simply have to twiddle your thumbs up your ass until they decide to do something about it.

9 months ago | Likes 1 Dislikes 0

Well I don't think you're going to do much about it if they misconfigure the network your server is connected to, as happened in this outage...?

9 months ago | Likes 1 Dislikes 0

The story says "about 1000 sites" and then lists about 50. Do you think the other 900+ sites get the same response when these things go down? Do you think they have a person to call to find out what's going on? The important point is *big businesses* can get money back, everyone else gets fucked. And it was down for about 4 hours, 0.05% of a year. For one outage.

9 months ago | Likes 3 Dislikes 1

For your questions, I have no idea what the answer is to either of them. Other people's contracts are not within my knowledge. I'm not sure what point you're making about the outage duration. Yeah, it is 0.05%, meaning that any company with a higher than 99.95% uptime SLA will be getting compensation, as will any with a lesser guarantee if there were other outages this year.

9 months ago | Likes 1 Dislikes 0

But other than that... I mean, yeah, that's a risk you take, but for anyone that _isn't_ a big business, it's a hell of a lot easier, less risky, and/or cheaper to use some cloud hosting than it is to build your own, multiple datacenters with redundant power in different regions and to manage the personnel to maintain them. I don't have a hundred million dollars to blow on a personal project, for example.

9 months ago | Likes 1 Dislikes 0

My databases are reliable as fuck tyvm. I don’t run hundreds of PTs of synchronized data across multiple locations like cloud providers do. It’s not uncommon that they get a data integrity issues causing chaos like this. Its usually a single service but I don’t see why it couldn’t happen across a full provider.

9 months ago | Likes 5 Dislikes 3

You aren't hosting in multiple regions? So one BGP error could take you out entirely until it's fixed (just like this story)?

9 months ago | Likes 1 Dislikes 0

do you have fire and natural disaster protection? as for redundancy. how many Internet links, generators, cooling unitsa? then there's multi-site redundancy

9 months ago | Likes 7 Dislikes 0

Colo literally has Fire and natural disaster protection built in. Feeds for both legs of power from two seperate power companies. Link has failover to second data provider. All of our hardware has redundancy. Networking/network appliances. Servers and services are both redundant. Yep. All of that you mentioned plus some. AND we regularly test failover.

9 months ago | Likes 3 Dislikes 1

Now that's a company willing to invest a lot in its own infrastructure, and I applaud that.

9 months ago | Likes 1 Dislikes 0

That's vastly oversimplified. The full ramifications of abdication of control are never contemplated and the SLA never covers losses. AWS is not going to compensate for lost revenue, the whole "cloud" business model would fail under such liability. Rarely considered is exposure of data and the target value centralized data presents.

Never put critical or private data on anyone else's hardware, ever. Hardware rental services should be liable for data and revenue loss.

9 months ago | Likes 16 Dislikes 5

Oof, clearly the cloud guys need to spin that aspect too then, how about "It's not a just a data breach, it's a surprise bonus backup!"

9 months ago | Likes 4 Dislikes 0

It’s impossible to find an IOS mobile password manager that ISN’T a cloud based subscription service and it drives me BONKERS.

9 months ago | Likes 1 Dislikes 0

Bitwarden

9 months ago | Likes 3 Dislikes 0

You can put that in a contract. I put it in all of my contracts with AWS.

9 months ago | Likes 1 Dislikes 0

What part? They have liability for your lost revenue should their service fail? They have liability for your liabilities to your customers should their infrastructure be breached or compromised; whether by malice, incompetence, or act of god; whether by internal or external actions; all harms; immediate, projected, and goodwill; shall be made whole?

9 months ago | Likes 1 Dislikes 0

Career server engineer here. While in principle I agree on self-hosting your hardware, the world doesn't work like that. It's WAY too expensive to do so when starting out, and as much as I hate to admit it, AWS is in practice far too reliable and secure (when you configure correctly) to not use it at its current price. Lastly, if an SLA doesn't cover losses, what are they for? AWS will in fact lose money to several of their impacted clients.

9 months ago | Likes 16 Dislikes 1

I'd be very interested to see what the agreed remedies are for loss of utility. I have only negotiated what are in effect consumer click agreements, and I have run services on various cloud providers, even in my own startups because the usual deferring capital expenses and focusing burn on code. I did not host any customer data on other people's hardware nor any company confidential data. I'm being forced to weaken security with cloud enclaves to comply with CMMC 2.0 now, sucks....>

9 months ago | Likes 2 Dislikes 0

But back to AWS etc: if they are willing to accept operational liability, how do they bill against risk? Does a customer have to share their revenue per minute and then pay to insure against that? Do you have to incorporate the downtime risk in business continuity insurance? What if it is a life and death service, such as image stream processing for vehicles or remote surgery?

9 months ago | Likes 2 Dislikes 0

Also interestingly, I’m seeing much better performance on azure blob vs S3.

9 months ago | Likes 1 Dislikes 0

In my experience, SLAs are for upper management and senior leadership to pretend that they are negotiating something useful and to have a number to report against.

9 months ago | Likes 4 Dislikes 2

Does your experience include contract law?

9 months ago | Likes 3 Dislikes 0

This guy gets it. Hello fellow engineer. I use AWS for healthcare at tons of hospitals

9 months ago | Likes 1 Dislikes 0

It would be very interesting to see a typical HIPAA compliance term. That came up recently in a related way considering remote processing of patient self-collected image data. The concern being ensuring both compliant and genuinely secure 3rd party processing of such data, which meant it could not be e2e on owned, premised hardware.

9 months ago | Likes 1 Dislikes 0

Sure, but without the cloud good luck scaling your service when your user count triples overnight because of a celebrity social media post.

9 months ago | Likes 8 Dislikes 4

Corps should use a combination of on-site and cloud capacity, to account for sudden surges. If the surge becomes the norm, build out more on-site capacity. If it fades, great.

Instead, corps have put EVERYTHING in the "cloud" and that makes them vulnerable when one moron at said Cloud Computing Company trips over a network cable or fucks up a DNS entry.

9 months ago | Likes 2 Dislikes 1

The only use case for cloud that has made sense to me is a hybrid approach m: on prem for base load, cloud to scale to meet surge demand

9 months ago | Likes 3 Dislikes 0

Load balancing in a hybrid environment is really tricky and has to be done from the cloud side if you're actually dealing with meaningful user counts. What you're saying is smart for segmentation of internal applications with very predictable usage/growth patterns, but anything public facing pretty much needs to be cloud-based just to provide the UX we all expect. People tend to bounce in the time it takes to route to your on-prem infrastructure's gateway/load balancer and on to the real server.

9 months ago | Likes 4 Dislikes 1

If users leave in the time it takes to resolve to your on-site servers, you are doing something VERY wrong with your on-site servers.

Fiber optic connections are the norm now, and having your office in a major city should make it indistinguishable, in terms of response time, from AWS or MS Azure. If it's slower, you fucked up.

9 months ago | Likes 2 Dislikes 0

I’ll trust you that. We kept everything on prem and rarely had a surge problem. Hitting the front page of Reddit - oh, that was a day.

9 months ago | Likes 3 Dislikes 0

Those days are why cloud-first deployments are best practice for anything user facing. If your site goes down during a surge, what % of the potential business do you think the overload cost your employer? The other side of the coin, it's really hard to attract and retain people good with physical hardware management. When something does break, I'd rather it be AWS's best and brightest on it rather than Steve who got hired because his uncle used to work on the marketing team.

9 months ago | Likes 2 Dislikes 0