Use as a dedicated GPU for encoding and decoding video.
Post processing like frame interpolation or superresolution. Use for GPGPU workloads. Run additional monitors independently. Use for GPU passthrough to virtual machines. Use as a backup GPU for troubleshooting. Use for test code without breaking the main GPU
I just bought a used Ayaneo 2 handheld, it has an old(er) mobile RDNA 2 GPU and I was blown away by how well this thing performed under Linux. Almost everything (that's not a recent AAA game) runs beautiful and a lot faster/smoother than it does under Windows. The experience has been so good that I'm considering switching my main pc (with a 9070XT) to Linux as well.
No doubt Timur contributed heavily to this given Valves Steamdeck (which uses a very similar but slower GPU).
Given the current hardware prices it's pretty awesome to see someone squeezing maximum performance out of old hardware!
I'm sure Valve probably still deserves the credit, since the Steam Deck uses a RDNA 2 gpu, but the article mentions that Timur specifically worked on gcn1.0 GPUs - e.g. radeon HD 7800/7900 from 2012 (!)
My experience has been that for Linux you are always better buying older mid tier hardware because all of the issues and optimisations have already been worked out and those changes have flowed through to your distro so you aren’t waiting on kernel updates.
I've switched my gaming PC last year to Linux and it's been flawless.
Originally ran a 3070 which matches your "older mid tier" description and then upgraded to a 9070XT maybe 6 months after launch so not so old or mid-tier.
Honestly, the folks writing inference drivers for Llama.cpp / GGML would sort of benefit from better compiler work like this.
In general, Valve's work has been exceptional and supplementing AMD's own ROCM/OpenCL and Vulkan team, they've gotten a lot of defaults right and people should work together with Valve to improve support for their chips.
One thing they should've learned from Nvidia is that it's really worth it for them to invest making their devices function as broadly as possible. Crypto and AI waves both benefited Nvidia much more than AMD partly due to their devices being more universally usable.
That "partly" was worth hundreds of billions of dollars but AMD were cheap/shortsighted enough to hire a few dedicated engineers.
Perhaps René, the maintainer of T/2 Linux, will end up vibe coding an NVIDIA driver one of these days. He’s been reverse engineering and vibe coding a bunch of drivers for old graphics lately and live streaming everything.
Maybe there’s a prestige angle to somewhat shame them into keeping up with nvidia. Though my sense is that they’re barely peers anymore in this space. :(
I have a Radeon RX 6900 XT (16 GB VRAM, originally released in 2021), and it's possible to run some lightweight models to have okayish performance and quality of output, but nothing I've tried has come anywhere close to the quality even of the models I can use for free from OpenCode Zen or the free tier of Openrouter. If you want to keep everything local on the same card I have, it requires putting up with a model that's noticeably worse in virtually every metric than what you can get for free elsewhere, and the GPUs this article are talking about are three times as old as mine.
It would be awesome if someone manages to figure out how to get small enough models to fit on older cards to be viable, but I'm not optimistic that it will come without some sort of fundamental architectural innovation rather than incremental improvements, and it's not clear if and when that will happen.
With an extra 8gb of vram you could run qwen 3.8 27b pretty comfortably, which isn't quite as good as frontier models but definitely on par with free models on openrouter and whatnot.
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
I have no clue, and I agree that it does not seem sustainable. Either someone needs to find a magic solution to making it a lot cheaper, or a lot of companies are going to need a lot of money from somewhere that isn't clear.
Most of these older cards are lacking the physical hardware for fp8 or other lower precisions that most quantized models use. Or the memory to run models at higher precision.
Use as a dedicated GPU for encoding and decoding video. Post processing like frame interpolation or superresolution. Use for GPGPU workloads. Run additional monitors independently. Use for GPU passthrough to virtual machines. Use as a backup GPU for troubleshooting. Use for test code without breaking the main GPU
No doubt Timur contributed heavily to this given Valves Steamdeck (which uses a very similar but slower GPU).
Given the current hardware prices it's pretty awesome to see someone squeezing maximum performance out of old hardware!
Originally ran a 3070 which matches your "older mid tier" description and then upgraded to a 9070XT maybe 6 months after launch so not so old or mid-tier.
In general, Valve's work has been exceptional and supplementing AMD's own ROCM/OpenCL and Vulkan team, they've gotten a lot of defaults right and people should work together with Valve to improve support for their chips.
Nvidia doesn't, and their Vulkan stack underperforms on Linux quite significantly.
That "partly" was worth hundreds of billions of dollars but AMD were cheap/shortsighted enough to hire a few dedicated engineers.
https://windowsforum.com/news/nvidia-ends-feature-support-fo...
It would be awesome if someone manages to figure out how to get small enough models to fit on older cards to be viable, but I'm not optimistic that it will come without some sort of fundamental architectural innovation rather than incremental improvements, and it's not clear if and when that will happen.
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
How long is this runway?