Trending / repos
lyogavin/
airllm
AIRLLM 70B INFERENCE WITH SINGLE 4GB GPU QUICKSTART | CO
lyogavin/airllm
lyogavin/airllm · Jupyter Notebook
AirLLM 70B inference with single 4GB GPU Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB , and DeepSeek-V3 (671B) on ~12GB . AI Agents Recommendation: Best AI Game Sprite Generator Best AI Facial Expres
- Stars
- 24K
- Forks
- 2.7K
- Issues
- 89
- PRs
- 26
AirLLM 70B inference with single 4GB GPU Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB , and DeepSeek-V3 (671B) on ~12GB . AI Agents Recommendation: Best AI Game Sprite Generator Best AI Facial Expres