Trending / repos

lyogavin/

airllm

AIRLLM 70B INFERENCE WITH SINGLE 4GB GPU QUICKSTART | CO

24K Stars2.7K Forks89 Issues

lyogavin/airllm

lyogavin/airllm · Jupyter Notebook

AirLLM 70B inference with single 4GB GPU Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB , and DeepSeek-V3 (671B) on ~12GB . AI Agents Recommendation: Best AI Game Sprite Generator Best AI Facial Expres

Stars
24K
Forks
2.7K
Issues
89
PRs
26

Open on GitHubArchive stats

AirLLM 70B inference with single 4GB GPU Quickstart | Configurations | MacOS | Example notebooks | FAQ AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB , and DeepSeek-V3 (671B) on ~12GB . AI Agents Recommendation: Best AI Game Sprite Generator Best AI Facial Expres