Tracked Repositories

158 open-source AI inference repositories across 68 organizations.

158 repositories across 68 organizations

HuggingFace

5 repos·230.0k·103 commits this week

HuggingFace Transformers — state-of-the-art NLP/ML model library (~140K stars)

163.4k
75

HuggingFace Diffusers — diffusion model inference & training (Stable Diffusion, Flux, etc.)

34.2k
20

Minimalist Rust ML framework for inference — targets browser WASM and GPU, zero Python dependency

20.9k
3

HuggingFace TGI — LLM serving (archived March 2026, read-only)

10.9k

One-command HuggingFace Hub → OpenVINO IR export; INT4/INT8 quantization via NNCF for Intel NPU deployment

611
5

TensorFlow

3 repos·206.3k·279 commits this week

Industry-standard deep learning framework with XLA compilation backend

196.9k
276

TensorFlow Serving — high-performance gRPC/REST serving for TF models (multi-version, canary, batching)

6.4k
2

TensorFlow Lite for microcontrollers and embedded devices

3.0k
1

Ollama

1 repo·177.9k·15 commits this week

User-friendly local LLM runner built on llama.cpp (~167K stars)

177.9k
15

ggml-org

2 repos·175.5k·142 commits this week

High-performance LLM inference in C/C++ (CPU + GPU)

122.9k
107

High-performance Whisper speech recognition in C/C++

52.6k
35

Open WebUI

1 repo·148.1k

Self-hosted ChatGPT alternative with built-in RAG, offline-capable (~104K stars)

148.1k

Meta / PyTorch

3 repos·111.5k·382 commits this week

Primary ML framework; torch.compile + AOTInductor for production inference optimization

102.2k
313

PyTorch's portable execution framework for on-device inference

4.9k
69

TorchServe — production PyTorch model serving (archived August 2025)

4.3k

DeepSeek AI

1 repo·104.1k

Reference inference code for DeepSeek-V3 (671B MoE); includes FP8 training framework

104.1k

vLLM Project

2 repos·88.4k·285 commits this week

Most widely adopted open-source LLM serving engine; PagedAttention, continuous batching

88.4k
273

vLLM community plugin for Intel Gaudi accelerators

51
12

Google AI Edge

18 repos·83.2k·263 commits this week

Cross-platform ML pipeline framework (vision, audio, NLP)

36.5k
10

AI Edge model gallery

24.4k
17

LiteRT for language model inference

6.1k
49

Official Gemma model cookbook — recipes, fine-tuning, deployment guides

4.0k

Google's Lite Runtime (successor to TensorFlow Lite)

3.3k
73

Sample apps using MediaPipe

2.8k
4

Highly optimized neural network operators library (ARM, x86, WASM)

2.4k
34

Model visualization and exploration tool

1.5k
3

LiteRT integration with PyTorch

1.1k
14

Python API for Coral Edge TPU inference (archived)

405

Sample code for LiteRT

388
43

Quantization tooling for AI Edge models

188
2

C++ API for Coral Edge TPU inference (archived)

95

Web samples for MediaPipe

64
2

Command-line tooling for LiteRT

34
12

Sample models for AI Edge

25

Evaluation tooling for AI Edge models

9

Google AI Edge documentation site

Nomic AI

1 repo·77.4k

Desktop AI app + SDK for running LLMs locally (~73K stars)

77.4k

Miscellaneous

7 repos·64.3k·134 commits this week

MLC's universal LLM deployment engine (multi-backend)

23.0k
1

DwarfStar native DeepSeek V4 inference engine for Metal, CUDA, and ROCm

20.8k
80

Smartphone/PC LLM inference exploiting activation sparsity (PowerInfer-2: Qualcomm NPU, Mixtral MoE up to 47B); org transferred from SJTU-IPADS

9.7k

Tile-based ML language and compiler

7.1k
41

Community on-device LLM project

1.6k

vLLM-style inference on Apple silicon via MLX

1.5k
10

OpenVINO-based OpenAI-compatible inference server for Intel CPU/GPU/NPU

496
2

Apple / ML-Explore

13 repos·63.7k·81 commits this week

Array framework for ML on Apple silicon (Python)

27.9k
44

Example models and applications using MLX

8.9k

Reverse-engineered Apple Neural Engine (ANE) — hardware ops, memory layout, firmware interactions

7.2k

LLM inference and fine-tuning with MLX

6.5k
2

Tools for converting & running models with Core ML

5.4k
4

Example apps using MLX Swift

2.6k

Swift bindings for MLX

2.0k
1

Model export recipes, Python primitives, and Swift runtime utilities for on-device AI (Core AI)

1.5k
14

LLM inference in Swift via MLX

767
14

Efficient data loading for MLX

480

C bindings for MLX

229

Bridges PyTorch and Core AI — convert models to Core AI IR, composite ops, custom lowerings, inline Metal kernels

126

PyTorch model compression and optimizations for deployment via Core AI on Apple silicon

105
2

BerriAI

1 repo·55.7k·418 commits this week

Unified OpenAI-compatible proxy for 100+ LLM providers (vLLM, Ollama, Bedrock, Azure, etc.)

55.7k
418

Mudler (LocalAI)

1 repo·48.3k·152 commits this week

Free, open-source OpenAI drop-in replacement — runs locally, no GPU required (~36K stars)

48.3k
152

Oobabooga

1 repo·47.5k

Gradio web UI for LLMs — multi-backend (llama.cpp, ExLlamaV2, transformers) (~43K stars)

47.5k

Exo Explore

1 repo·46.7k

Run LLMs distributed across heterogeneous devices (Mac, iPhone, etc.)

46.7k

Microsoft / ONNX

3 repos·44.9k·103 commits this week

Microsoft's cross-platform, high-performance ONNX inference engine

21.3k
57

Open Neural Network Exchange format specification

21.3k
37

Model optimization tool (HF → ONNX → quantize → NPU deployment); used under the hood by Microsoft Foundry Local

2.4k
9

Ray Project

1 repo·43.5k·77 commits this week

Distributed AI compute engine; Ray Serve handles online and async batch inference (~39K stars)

43.5k
77

DeepSpeed AI

1 repo·42.9k·27 commits this week

Microsoft DeepSpeed — distributed training and inference (ZeRO, MII, FastGen)

42.9k
27

LM-Sys

1 repo·39.5k

LLM serving framework and home of Chatbot Arena (~37K stars)

39.5k

NVIDIA

4 repos·38.2k·171 commits this week

NVIDIA's optimized LLM inference library (GPU)

14.3k
163

NVIDIA's high-performance deep learning inference SDK (GPU)

13.2k
1

CUDA C++ templates for high-performance matrix-multiply (GEMM) and convolution kernels

10.2k
7

C++ LLM/VLM inference runtime for Jetson and NVIDIA edge devices

494

JAX (Google DeepMind)

1 repo·36.1k·149 commits this week

Composable NumPy transformations (JIT, grad, vmap) compiled via XLA to GPUs and TPUs — primary DeepMind research/production runtime

36.1k
149

SGLang

1 repo·31.4k·303 commits this week

High-throughput LLM/VLM serving with RadixAttention and structured generation

31.4k
303

Tencent

2 repos·28.3k·5 commits this week

High-performance neural network inference for mobile (Android/iOS)

23.6k
5

Tencent Neural Network — mobile and edge inference

4.6k

Modular

1 repo·26.7k·388 commits this week

Modular Platform monorepo — MAX inference server/framework + Mojo programming language for portable, high-performance AI on CPUs and GPUs

26.7k
388

Mozilla AI

1 repo·25.5k·2 commits this week

Single-file LLM executables via Cosmopolitan Libc — zero install, all platforms (~21K stars)

25.5k
2

Dao AI Lab

1 repo·24.6k·6 commits this week

Official FlashAttention — fast, memory-efficient exact attention (FA-2/FA-3) kernels for GPUs

24.6k
6

Triton Language (OpenAI)

1 repo·19.9k·54 commits this week

Python-like GPU kernel language used by vLLM FlashAttention and PyTorch inductor

19.9k
54

KVCache AI

1 repo·19.2k·3 commits this week

CPU-GPU hybrid inference; runs DeepSeek 671B on 14GB VRAM + 382GB DRAM with massive speedup over llama.cpp

19.2k
3

MLC AI

1 repo·18.5k·1 commits this week

High-performance LLM inference in web browsers via WebGPU

18.5k
1

jundot

1 repo·18.5k·80 commits this week

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

18.5k
80

Alibaba

1 repo·15.8k·14 commits this week

Alibaba's neural network inference framework for mobile & edge

15.8k
14

Blaizzy (Community MLX)

5 repos·14.4k·75 commits this week

Audio models (TTS, ASR) with MLX

7.7k
14

Vision-language models on Apple silicon via MLX

5.3k
61

Swift audio inference using MLX

747

Text embedding models with MLX

423

Video model inference with MLX

284

Apache

2 repos·14.1k·29 commits this week

Apache TVM ML compiler — auto-tunes models for any hardware target

13.7k
20

Apache TVM Foreign Function Interface for deep learning compilation

444
9

K2 / Next-gen ASR

1 repo·14.0k·8 commits this week

ONNX-based runtime for ASR, TTS, VAD, and keyword spotting

14.0k
8

OpenVINO Toolkit / Intel

3 repos·12.4k·90 commits this week

Intel's toolkit for optimizing & deploying deep learning on Intel hardware

10.6k
76

Neural Network Compression Framework — quantization, pruning, sparsity for OpenVINO

1.2k
3

OpenVINO GenAI — generative AI layer with speculative decoding & KV-cache opt

566
11

RunAnywhere

2 repos·11.8k·13 commits this week

RunAnywhere SDKs for on-device inference deployment

10.3k
13

RunAnywhere CLI tool

1.5k

Intel

2 repos·11.6k

Intel IPEX-LLM — local LLM acceleration on Intel hardware (archived Jan 2026, read-only)

8.9k

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity

2.7k

AMD Ryzen AI (XDNA NPU)

7 repos·11.5k·90 commits this week

Lemonade SDK — high-level, multi-vendor LLM inference SDK (OGA + llama.cpp), OpenAI-compatible server mode

5.3k
23

Xilinx/AMD AI Engine dev stack — source of the Vitis AI Execution Provider for XDNA

1.8k

AMD-backed Linux NPU runtime, used alongside Lemonade Server on Ryzen AI XDNA 2

1.7k
9

Open-source RAG/chat/agent app for Ryzen AI NPU, built on Lemonade SDK

1.5k
58

Official Ryzen AI Software stack — Vitis AI EP, quantization, LLM deployment; contains an experimental XDNA-NPU llama.cpp fork

864

Ingests any PyTorch model, optimizes, benchmarks/deploys across hardware targets incl. XDNA NPU

242

AMD's quantization toolkit (INT4/INT8 PTQ + QAT) for Ryzen AI NPU deployment

153

Triton Inference Server

1 repo·10.9k·1 commits this week

NVIDIA Triton — production multi-model inference server (HTTP/gRPC, multi-backend)

10.9k
1

Mistral AI

1 repo·10.8k

Official minimal inference library for all Mistral models (7B, Mixtral, Pixtral)

10.8k

Qualcomm

3 repos·9.9k·122 commits this week

On-device LLM/VLM SDK for Snapdragon NPU, GPU, and CPU (formerly Nexa AI)

8.3k
78

State-of-the-art ML models optimized for Qualcomm Snapdragon NPU/DSP/QNN deployment

1.2k
39

Sample apps and tutorials for deploying models on Qualcomm hardware (TFLite, ONNX, QNN)

443
5

Dusty-NV (NVIDIA Jetson)

1 repo·9.0k

DNN inference library & tutorials for NVIDIA Jetson

9.0k

BentoML

1 repo·8.8k

Unified serving framework: real-time APIs, task queues, batching, multi-model chains

8.8k

InternLM / Shanghai AI Lab

1 repo·8.0k·12 commits this week

High-throughput LLM serving with TurboMind engine (C++/CUDA)

8.0k
12

AI Dynamo (NVIDIA)

1 repo·7.7k·175 commits this week

Datacenter-scale distributed inference serving framework (Rust + Python, disaggregated prefill/decode, engine-agnostic)

7.7k
175

Osaurus

1 repo·7.5k·61 commits this week

Native macOS AI agent harness in Swift — any model, persistent memory, autonomous execution, MCP server, MLX + Apple Neural Engine, fully offline

7.5k
61

PaddlePaddle (Baidu)

1 repo·7.3k

Lightweight inference engine for mobile & embedded from PaddlePaddle

7.3k

ArgMax

4 repos·6.7k·6 commits this week

On-device Whisper inference for Apple platforms (Swift)

6.3k
5

Python tooling for WhisperKit model optimization

245

On-device AI benchmarking framework

88
1

Swift playground for ArgMax SDK

21

FlashInfer

1 repo·6.1k·51 commits this week

High-performance GPU kernel library for LLM serving — attention, sampling, and KV-cache primitives

6.1k
51

Cactus Compute

5 repos·5.9k

Cactus core edge inference framework

5.6k

React Native bindings for Cactus

178

Kotlin/Android bindings for Cactus

74

Flutter bindings for Cactus

71

Demo chat app using Cactus

28

OpenNMT

1 repo·4.6k·1 commits this week

Fast C++ inference for Transformer models; INT8/INT16 CPU quantization, multi-platform

4.6k
1

TurboDeRP (ExLlamaV2)

1 repo·4.6k

High-performance EXL2-quantized inference for consumer NVIDIA GPUs

4.6k

OpenXLA

1 repo·4.5k·193 commits this week

Compiler for JAX, TF, PyTorch targeting GPU, TPU, and CPU from a unified IR

4.5k
193

ModelTC

1 repo·4.2k·16 commits this week

Lightweight, high-throughput Python-based LLM inference and serving framework

4.2k
16

Predibase

1 repo·3.8k

Multi-LoRA inference server — serve thousands of fine-tuned adapters on a single GPU

3.8k

Liquid AI

5 repos·3.0k·2 commits this week

Examples, tutorials and apps for Liquid AI LFM + LEAP SDK

2.2k

Speech-to-Speech audio models by Liquid AI

555

Minimal fine-tuning repo for LFM2, fully open-source

185
1

Example apps for LeapSDK

71

Liquid AI documentation

28
1

Luminal AI

1 repo·2.9k·3 commits this week

Rust-based deep learning compiler with a small static graph IR for fast, portable inference (CUDA, Metal, CPU)

2.9k
3

Fluid Inference

3 repos·2.8k·4 commits this week

On-device audio inference framework

2.6k
4

Fluid Inference core runtime

85

Rust text processing library for inference

46

Try Mirai

2 repos·1.8k·23 commits this week

Mirai's on-device inference runtime

1.7k
20

Mirai's LLaMA-based on-device model

86
3

UbiquitousLearning

1 repo·1.6k·2 commits this week

Multimodal LLM inference framework for mobile & edge

1.6k
2

ARM Software

1 repo·1.3k

ARM Neural Network SDK for ARM & Mali devices

1.3k

AMD ROCm

4 repos·1.2k·100 commits this week

AI Tensor Engine for ROCm — centralized repo for high-perf AI operators on AMD Instinct GPUs

522
53

AMD's graph inference engine for MI-series GPUs

319
12

ROCm fork of FlashAttention with Composable Kernel (CK) and Triton backends

236

AiTer Optimized Model — lightweight vLLM-like server built on AITER kernels for ROCm

150
35

Rudrank Riyam

1 repo·1.2k

iOS/macOS workbench for Apple's Foundation Models framework — recipes, guided labs, prompt playground, AFM CLI, FMFBench eval suite, and reusable Swift packages for on-device Apple Intelligence

1.2k

NimbleEdge

2 repos·529

NimbleEdge's deliteAI on-device inference framework

527

NimbleEdge fork of ExecuTorch with edge optimizations

2

ThunderAgent

1 repo·405

A simple, fast and robust program-aware agentic inference system

405

Picovoice

1 repo·316·4 commits this week

Picovoice's on-device LLM inference engine

316
4

Zetic AI

5 repos·78

MLange sample applications

70

MLange extension library

5

iOS framework for MLange

2

iOS extension framework for MLange

1

MLange SDK documentation

0