Can I run ollama on RTX 3060 and Inter iGPU to increase speed?

Jeena@piefed.jeena.net · 9 months ago

Can I run ollama on RTX 3060 and Inter iGPU to increase speed?

just_another_person@lemmy.world · 9 months ago

Nope.

theunknownmuncher@lemmy.world · 9 months ago

Models are computed sequentially (the output of each layer is the input into the next layer in the sequence) so more GPUs do not offer any kind of performance benefit

Blue_Morpho@lemmy.world · edit-2 9 months ago

More gpus do improve performance:

https://medium.com/@geronimo7/llms-multi-gpu-inference-with-accelerate-5a8333e4c5db

All large AI systems are built of multiple “gpus” (AI processers like Blackwell ). Really large AI models are run on a cluster of individual servers connected by 800 GB/s network interfaces.

However igpus are so slow that it wouldn’t offer significant performance improvement.

theunknownmuncher@lemmy.world · 9 months ago

What I am talking about is when layers are split across GPUs. I guess this is loading the full model into each GPU to parallelize layers and do batching

Blue_Morpho@lemmy.world · 9 months ago

No, full models are not loaded into each GPU to improve the tokens per second.

The full Gpt 3 needs around 640GB of vram to store the weights. There is no single GPU (ai processor like a100) with 640 GB of vram. The model is split across multiple gpus (AI processers).