AWS Tells Engineers to Cut CPU Waste Amid Crunch
Resumo
AWS enfrenta crise de capacidade de servidores e instrui engenheiros a reduzirem consumo de CPU; tempo de espera para provisionar capacidade aumentou de horas para dias, afetando prazos de projetos, enquanto demanda por chips de IA agrava escassez de recursos computacionais.

Amazon Web Services leaders met with engineers in May and delivered a sobering message: To ensure AWS has enough capacity for all its customers at its popular EC2 cloud server business in the future, engineers should save capacity however they can, according to a person with knowledge of the meeting.
That includes capacity on servers that run on central processing units, as well as those running on AI chips, which have long been in short supply, the person said. CPUs are the chips that have powered the last several decades of the internet age. Since then, some engineers within AWS have found it takes much longer to get capacity on CPU servers for their work.
One said it now takes a few days to get capacity for servers they used to get in a few hours, which could make it harder to meet deadlines on projects. The engineer said they’ve never experienced such long wait times during their several years of working at AWS.
AWS has given teams deadlines set for later this year to reduce the amount of compute used, according to someone with knowledge of the plans. Engineers are decommissioning idle EC2 virtual servers—known as instances—that they started using for software development but now no longer need, allowing them to reallocate this capacity to customers, the person said.
The compute crunch that has bedeviled the tech sector this year is increasingly expanding to older technologies. It’s no secret that skyrocketing demand for AI chips, such as Nvidia’s graphics processing units, has contributed to a shortage of AI capacity. But many companies are finding that even CPU-powered servers are in short supply, thanks to AI’s rising demand for CPUs. A lack of the memory chips that work with CPUs and of physical data center space are also contributing to the capacity crunch.
A big factor is that companies are using more CPUs as more employees are building software with agents, said Jing Xie, co-founder and managing director at Elendil Labs, a firm that helps financial services firms use AI. Xie said he’s seen clients’ IT spend per employee double because they’re doing more work using AI agents, which can require paying for more cloud compute and typically creates a need for CPUs.
Even in AI development, CPUs play a big role: For example, when AI firms prepare data for use in training models, CPUs read the raw data from documents, images and videos.
“Basically, a lot of what’s getting built and produced and run is requiring more CPUs than in the past,” Xie said.
The ratio of CPUs to GPUs used in AI inference—the running of models—was one to four in April, Intel CEO Lip-Bu Tan said in the company’s earnings call that month. In July, Intel’s finance chief, David Zinsner, said that ratio was approaching parity. AMD and Arm executives have made similar comments.
AWS disputes that anything has changed in how it runs, noting in a statement that even with “heavy demand, we continue to satisfy the overwhelming majority of compute needs for both our internal and external customers. We work closely with internal teams to meet their compute needs while ensuring they use EC2 resources as efficiently as possible…just as we’ve always done.”
Amazon says its engineers can operate more efficiently using the same guidance AWS gives its customers: for instance, turning off idle instances or switching to a type of instance more appropriate for a particular customer’s workload. If a customer is not using all of the capacity in a particular instance, they can switch to a smaller one.
AWS recommends that when a customer is using less than 40% of an EC2 server’s CPU and memory chip capacity over a four-week period, they should switch to a smaller machine.
AWS offers software enabling customers to run automated scans checking for underutilized servers. Amazon says its guidance to employees has not changed in light of the ongoing memory crunch.
Balancing internal capacity needs with those of customers has become a challenge for many big tech companies in the past couple of years. Last year, Google created a council of senior executives to decide how to allocate computing capacity between Google Cloud, its DeepMind AI research unit and its consumer businesses. Even so, tensions have arisen. Google’s star AI researcher, Noam Shazeer, left the lab earlier this summer over frustrations with his access to compute, The Information reported.
When companies can manage compute effectively, they can more quickly turn capacity into revenue. Microsoft Chief Financial Officer Amy Hood on the company’s most recent earnings call last month attributed some of the growth of its Azure cloud service to “efficiency gains” it made in managing its CPU and GPU fleets.
“We saw…good work this quarter in particular from our engineering teams to make as much of that available as we could,” Hood said. “And because of the supply-demand imbalance we’ve been talking about, when we can make efficiency gains, they are quickly monetized.”
A consultant who helps companies use AWS said he hasn’t seen operational shortages for the CPU computing capacity customers commit to via contracts. However, access to AWS spot instances—the overflow server capacity it sells at a deep discount but can reclaim with two minutes’ notice—has become more difficult to obtain in large quantities in recent months, the person said.