select between over 22,900 AI Tool and 17,900 AI News Posts.
AMD is buying Canadian startup Taalas, which hard-codes model weights directly into inference chips. That makes them extremely fast but locks each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach for Gemini.
The article AMD acquires Taalas, a startup that bakes AI models directly into silicon appeared first on The Decoder.
<p>AMD has reached a definitive agreement to acquire Taalas, a Toronto-based startup whose chips are custom-built around individual AI models, the company announced on August 6, 2026. The deal f [...]