Apple embraces Nvidia GPUs to accelerate LLM inference via its open source ReDrafter tech

ReDrafter delivers 2.7x more tokens per second compared to traditional auto-regression
ReDrafter could reduce latency for users while using fewer GPUs
Apple hasn’t said when ReDrafter will be deployed on rival AI GPUs from AMD and Intel

Apple has announced a collaboration with Nvidia to accelerate large language model inference using its open source technology, Recurrent Drafter (or ReDrafter for short).

The partnership aims to address the computational challenges of auto-regressive token generation, which is crucial for improving efficiency and reducing latency in real-time LLM applications.

ReDrafter, introduced by Apple in November 2024, takes a speculative decoding approach by combining a recurrent neural network (RNN) draft model with beam search and dynamic tree attention. Apple’s benchmarks show that this method generates 2.7x more tokens per second compared to traditional auto-regression.

Could it extend beyond Nvidia?

Through its integration into Nvidia’s TensorRT-LLM framework, ReDrafter extends its impact by enabling faster LLM inference on Nvidia GPUs widely used in production environments.

To accommodate ReDrafter’s algorithms, Nvidia introduced new operators and tweaked existing ones within TensorRT-LLM, making the tech available for any developers looking to optimize performance for large-scale models.

In addition to the speed improvements, Apple says ReDrafter has the potential to reduce user latency while requiring fewer GPUs. This efficiency not only lowers computational costs but also lessens power consumption, a vital factor for organizations managing large-scale AI deployments.

While the focus of this collaboration remains on Nvidia’s infrastructure for now, it’s possible that similar performance benefits could be extended to rival GPUs from AMD or Intel at some point in the future.

Breakthroughs like this can help improve machine learning efficiency. As Nvidia says, “This collaboration has made TensorRT-LLM more powerful and more flexible, enabling the LLM community to innovate more sophisticated models and easily deploy them with TensorRT-LLM to achieve unparalleled performance on Nvidia GPUs. These new features open exciting possibilities, and we eagerly anticipate the next generation of advanced models from the community that leverage TensorRT-LLM capabilities, driving further improvements in LLM workloads.”

You can read more about the collaboration with Apple on the Nvidia Developer Technical Blog.

https://cdn.mos.cms.futurecdn.net/pBQSiTGru55Z7ghrsPMhxP-1200-80.jpg

Source link
waynewilliams@onmail.com (Wayne Williams)

Why volumetric video works for the Olympics – but not yet for cinema

Why you shouldn’t ask ChatGPT for relationship advice — it’ll just tell you you’re right and ‘may worsen rather than resolve conflict’

The Oppo Find X9 Ultra could be the world’s best camera phone — and it’s launching globally this month

12 irresistibly boxy and beige Apple accessories that’ll transform your shiny new gadgets into retro tech in seconds

Ex-Israeli Intelligence Official: Shockwaves of Trump’s “Take Over Gaza” Heard, Felt Across Region

What UK political parties are promising in the 2019 general election

Otto Warmbier’s parents want North Korea to suffer for their son’s death

Could a ‘youthquake’ cause Boris Johnson to lose the general election?

Two and a Half Men Cast, Charlie Sheen, Angus T. Jones: Where Are They Now?

Ask Lisi: Celebrity podcasts get failing grade

GTA 6 Official Release Date Update Is Sending Gamers Wild

Forget ‘Send Help,’ Your Next Favorite Survival Thriller Hits Streaming This April

Trump has no good options in Iran—here are 5 of them ahead of his speech to the nation tonight

Form 13D/A Real Messenger Corporation For: 1 April

Trump has no good options in Iran—here are 5 of them ahead of his speech to the nation tonight

Hargreave Hale AIM VCT allots 4.7 million shares at 31.78p

The YouTuber who has become one of Gen Z’s most beloved celebrities

26 last-minute holiday gifts that are still thoughtful and unique

Practicing gratitude regularly can make you less stressed and sleep better

8 things millennials wish you would just stop getting them for the holidays

Apple embraces Nvidia GPUs to accelerate LLM inference via its open source ReDrafter tech

Trump has no good options in Iran—here are 5 of them ahead of his speech to the nation tonight

Why volumetric video works for the Olympics – but not yet for cinema

Two and a Half Men Cast, Charlie Sheen, Angus T. Jones: Where Are They Now?

Form 13D/A Real Messenger Corporation For: 1 April