LinkedIn’s Unusual AI Strategy: No New Data Center Spending This Year

LinkedIn’s Unusual AI Strategy: No New Data Center Spending This Year
Advertisement

LinkedIn’s Unusual AI Strategy: No New Data Center Spending This Year

Contrary to the AI expansion frenzy gripping most major tech platforms, LinkedIn is taking a strikingly different approach: it has no plans to ramp up capital spending on new AI data centers for its current fiscal year.

Leaders of the professional social network told WIRED that the company will hold its existing level of GPU investment steady, with no planned growth to its overall compute and storage capacity for the 12-month period that started last month and wraps up next June. LinkedIn says it can skip the costly AI hardware splurge because engineering teams have managed to double the efficiency of its existing GPU fleet over just the past six months.

While the company’s current strategy could still shift as AI’s hardware requirements evolve rapidly, executives note they have already factored steep memory chip price hikes into their planning.

“One of our core goals is to keep our overall compute footprint flat, or as close to flat as possible, even as we launch more compute-intensive AI tools to production,” said Erran Berger, LinkedIn’s chief technology officer for engineering. “That’s a pretty bold stance to take in today’s AI climate.”

Berger and Raghu Hiremagalur, LinkedIn’s CTO for infrastructure, say the deliberate spending slowdown is a choice to prioritize fiscal prudence, and the new capacity constraints will push engineering teams to innovate more creatively as they build out the long list of new generative AI features LinkedIn plans to roll out. Berger adds that efficiency improvements will likely compound over time, allowing LinkedIn to get far more value out of future data center expansions when the company does eventually increase its capital budgets.

“I can’t stress enough that for a company of our scale, committing to an entire year without adding new storage or compute capacity is no small achievement—it took an enormous amount of work to get to this point,” Hiremagalur said.

Right now, firms including OpenAI, Meta, and Google are pulling together every available dollar and striking unusual new partnerships to build, stock, and run massive data centers packed with the latest AI chips. Labor and component shortages have delayed many of these projects, and dozens of companies have been forced to cap user access to some AI tools to manage capacity. At the same time, questions are growing about whether the relentless pace of AI infrastructure investment is sustainable.

With more than 1.3 billion global users, LinkedIn is likely the largest company so far to publicly address these spending concerns by stepping back from the industry-wide AI building boom.

“This is an encouraging shift for the entire industry,” said Songyee Yoon, managing partner of Principal Venture Partners and a board member at server manufacturer HP. “It signals that AI is starting to move out of the experimentation phase and into a phase of production discipline. The companies that come out on top won’t just be the ones that spend the most on infrastructure.”

Owning and Optimizing In-House Infrastructure

A few years after Microsoft bought LinkedIn in 2016, the company tested a move to its parent firm’s Azure cloud service, but ultimately found it didn’t make economic sense to fit the massive social network into Azure’s general-purpose data centers.

“Microsoft Azure was growing exponentially, with customer demand hitting all-time highs, and at the same time we were seeing explosive growth on LinkedIn’s side too,” Hiremagalur explained. In 2022, LinkedIn fully committed to building and operating its own data centers across Oregon, Texas, and Virginia. Owning its infrastructure outright gave LinkedIn full control over every detail of its tech stack, putting it in a strong position to adapt to the demands of the new AI era.

Around that same time, LinkedIn began developing AI-powered tools, including personal assistants that help users draft messages, search for jobs, and source job candidates. The push into AI didn’t come cheap. “Every query processed on our site has gotten more costly over time,” Hiremagalur said, noting that LinkedIn’s total stored data was doubling every year. “That trajectory was not sustainable long-term.”

To fix that, LinkedIn set out to optimize data center usage across every stage of the AI workflow, from initial model training to serving model outputs in response to user queries. Hiremagalur’s team built custom measurement tools to track exactly how much compute and storage each engineering team was consuming, then implemented a new system to allocate projects to data center hardware far more efficiently, cutting down on the amount of time GPUs and servers sat idle.

“Our allocation efficiency and GPU utilization on the training side is the best I’ve ever seen in any operation,” Hiremagalur said, noting that utilization rates sit above 95 percent.

LinkedIn also leveraged techniques like model distillation, which trains smaller, leaner AI models using knowledge gained from larger, more complex models. For example, the company’s job recommendation tool now uses a single small model that learned from two larger parent models to both identify relevant job openings and predict which users are most likely to click on a posting. While the smaller model is far cheaper to operate, Berger says it has not sacrificed any output quality.

“People are now finding jobs they couldn’t successfully discover before, because the model does a really good job of understanding what they’re looking for,” he said.

The model that selects which posts show up in users’ news feeds was also extremely costly to run when it was first launched, but LinkedIn has since refined it to run at a reasonable cost, according to Berger. The company made dozens of small, impactful improvements, including streamlining the model training process, reusing data generated from earlier recommendations, and balancing workloads more effectively between CPUs and GPUs.

LinkedIn even reworked core software running on Nvidia processors to let them handle larger tasks than their original design specifications allowed. The company also shifted some less intensive workloads off of expensive, hard-to-source, power-hungry Nvidia GPUs and onto CPUs to free up GPU capacity for more demanding AI tasks.

Altogether, LinkedIn estimates that its efficiency work has saved roughly $24 million over the past 12 months—equal to the cost of keeping roughly 1,100 GPUs running 24 hours a day for a full year. LinkedIn’s leaders acknowledge that this savings is a relatively small sum for a company that generates $18 billion in annual revenue. But Hiremagalur argues that intentional engineering craft matters, and the freed-up capacity delivers major agility benefits: engineers can start new projects sooner and add more AI capabilities without requiring LinkedIn to expand its overall computing footprint.

Berger adds that LinkedIn has managed to deliver better outcomes for job seekers and recruiters alike while controlling costs and delivering strong financial returns for parent company Microsoft. “We should be able to deliver better quality by deploying larger models and running deeper inference, and do it more cheaply if we can get the efficiency right,” he said.

Even with flat capital spending, LinkedIn’s data centers are not being left outdated. The company has committed to purchasing new servers to replace machines that age out or break down over the coming months, but it managed to lock in lower costs by purchasing hardware ahead of recent price hikes. “The cost of all this AI hardware has just gone through the roof,” Hiremagalur said, noting that prices for some servers have tripled in just the past few months. “It’s just insane.”

A Broader Industry Shift to Efficiency

LinkedIn’s push for greater AI efficiency is part of a broader industry trend sometimes called “AI tokenomics,” which involves deep, granular analysis of the cost of running generative AI tools.

“Enterprises are evolving from a ‘buy more’ era to a ‘do more’ era,” said Chirag Dekate, a Gartner consultant who helps companies develop AI cloud strategies. “Until recently, the industry mantra was buy more to save more. But buying more hardware only leads to ever-increasing costs.”

Dekate says he’s recently seen smaller companies—who don’t have the same level of control over their infrastructure that LinkedIn enjoys—find their own ways to cut AI costs. These businesses have cut unused software, reduced unnecessary headcount, leased data center capacity from newer “neocloud” providers that are often cheaper than traditional big cloud providers, and opted for the lowest-cost AI model that meets the needs of each specific project.

Dekate does worry that LinkedIn and other companies focused on capping spending could hit a breaking point, since demand for more compute and storage is ultimately inevitable as AI grows more capable. “At some point, something has to give,” he said. “Either you have to scale back your AI ambitions, or you have to walk back your mandate to freeze IT spending.”

LinkedIn has not ruled out expanding its data center footprint in the future, but it is changing how it allocates infrastructure capacity. Gone are the days of making “a wild ass guess” about how much computing power engineering teams will need over the coming year, Hiremagalur says.

Instead, the company has chosen “to embrace the chaos” and plan on a quarter-by-quarter basis while keeping its focus on long-term return on investment, according to Berger. And it’s entirely possible that the current flat spending plan won’t last as long as currently projected.