Tracked Repositories

164 open-source AI inference repositories across 68 organizations.

164 repositories across 68 organizations

HuggingFace

5 repos·230.8k·62 commits this week

HuggingFace Transformers — state-of-the-art NLP/ML model library (~140K stars)

164.1k
37

HuggingFace Diffusers — diffusion model inference & training (Stable Diffusion, Flux, etc.)

34.3k
15

Minimalist Rust ML framework for inference — targets browser WASM and GPU, zero Python dependency

20.9k
1

HuggingFace TGI — LLM serving (archived March 2026, read-only)

10.9k

One-command HuggingFace Hub → OpenVINO IR export; INT4/INT8 quantization via NNCF for Intel NPU deployment

612
9

TensorFlow

3 repos·206.4k·373 commits this week

Industry-standard deep learning framework with XLA compilation backend

197.0k
370

TensorFlow Serving — high-performance gRPC/REST serving for TF models (multi-version, canary, batching)

6.4k

TensorFlow Lite for microcontrollers and embedded devices

3.0k
3

Ollama

1 repo·178.5k·34 commits this week

User-friendly local LLM runner built on llama.cpp (~167K stars)

178.5k
34

ggml-org

2 repos·176.7k·125 commits this week

High-performance LLM inference in C/C++ (CPU + GPU)

123.8k
108

High-performance Whisper speech recognition in C/C++

52.9k
17

Open WebUI

1 repo·148.7k

Self-hosted ChatGPT alternative with built-in RAG, offline-capable (~104K stars)

148.7k

Meta / PyTorch

3 repos·111.6k·408 commits this week

Primary ML framework; torch.compile + AOTInductor for production inference optimization

102.4k
319

PyTorch's portable execution framework for on-device inference

4.9k
89

TorchServe — production PyTorch model serving (archived August 2025)

4.3k

DeepSeek AI

1 repo·104.2k

Reference inference code for DeepSeek-V3 (671B MoE); includes FP8 training framework

104.2k

vLLM Project

2 repos·89.0k·335 commits this week

Most widely adopted open-source LLM serving engine; PagedAttention, continuous batching

89.0k
309

vLLM community plugin for Intel Gaudi accelerators

51
26

Google AI Edge

18 repos·83.6k·211 commits this week

Cross-platform ML pipeline framework (vision, audio, NLP)

36.6k
19

AI Edge model gallery

24.4k
10

LiteRT for language model inference

6.2k
49

Official Gemma model cookbook — recipes, fine-tuning, deployment guides

4.0k

Google's Lite Runtime (successor to TensorFlow Lite)

3.3k
78

Sample apps using MediaPipe

2.8k
1

Highly optimized neural network operators library (ARM, x86, WASM)

2.4k
23

Model visualization and exploration tool

1.5k
1

LiteRT integration with PyTorch

1.1k
4

Python API for Coral Edge TPU inference (archived)

405

Sample code for LiteRT

396
14

Quantization tooling for AI Edge models

189
2

C++ API for Coral Edge TPU inference (archived)

95

Web samples for MediaPipe

67
10

Command-line tooling for LiteRT

38

Sample models for AI Edge

25

Evaluation tooling for AI Edge models

10

Google AI Edge documentation site

Nomic AI

1 repo·77.4k

Desktop AI app + SDK for running LLMs locally (~73K stars)

77.4k

Miscellaneous

7 repos·64.9k·73 commits this week

MLC's universal LLM deployment engine (multi-backend)

23.1k

DwarfStar native DeepSeek V4 inference engine for Metal, CUDA, and ROCm

21.3k
31

Smartphone/PC LLM inference exploiting activation sparsity (PowerInfer-2: Qualcomm NPU, Mixtral MoE up to 47B); org transferred from SJTU-IPADS

9.7k

Tile-based ML language and compiler

7.2k
29

Community on-device LLM project

1.6k

vLLM-style inference on Apple silicon via MLX

1.5k
12

OpenVINO-based OpenAI-compatible inference server for Intel CPU/GPU/NPU

500
1

Apple / ML-Explore

13 repos·63.9k·165 commits this week

Array framework for ML on Apple silicon (Python)

27.9k
126

Example models and applications using MLX

8.9k

Reverse-engineered Apple Neural Engine (ANE) — hardware ops, memory layout, firmware interactions

7.2k

LLM inference and fine-tuning with MLX

6.6k

Tools for converting & running models with Core ML

5.4k
9

Example apps using MLX Swift

2.6k

Swift bindings for MLX

2.0k
1

Model export recipes, Python primitives, and Swift runtime utilities for on-device AI (Core AI)

1.5k
10

LLM inference in Swift via MLX

777
11

Efficient data loading for MLX

480

C bindings for MLX

230

Bridges PyTorch and Core AI — convert models to Core AI IR, composite ops, custom lowerings, inline Metal kernels

130
3

PyTorch model compression and optimizations for deployment via Core AI on Apple silicon

105
5

BerriAI

1 repo·56.3k·539 commits this week

Unified OpenAI-compatible proxy for 100+ LLM providers (vLLM, Ollama, Bedrock, Azure, etc.)

56.3k
539

Mudler (LocalAI)

1 repo·48.4k·80 commits this week

Free, open-source OpenAI drop-in replacement — runs locally, no GPU required (~36K stars)

48.4k
80

Oobabooga

1 repo·47.5k

Gradio web UI for LLMs — multi-backend (llama.cpp, ExLlamaV2, transformers) (~43K stars)

47.5k

Exo Explore

1 repo·46.8k

Run LLMs distributed across heterogeneous devices (Mac, iPhone, etc.)

46.8k

Microsoft / ONNX

3 repos·45.1k·77 commits this week

Microsoft's cross-platform, high-performance ONNX inference engine

21.4k
65

Open Neural Network Exchange format specification

21.3k
9

Model optimization tool (HF → ONNX → quantize → NPU deployment); used under the hood by Microsoft Foundry Local

2.4k
3

Ray Project

1 repo·43.5k·87 commits this week

Distributed AI compute engine; Ray Serve handles online and async batch inference (~39K stars)

43.5k
87

DeepSpeed AI

1 repo·42.9k·18 commits this week

Microsoft DeepSpeed — distributed training and inference (ZeRO, MII, FastGen)

42.9k
18

LM-Sys

1 repo·39.5k

LLM serving framework and home of Chatbot Arena (~37K stars)

39.5k

NVIDIA

4 repos·38.4k·213 commits this week

NVIDIA's optimized LLM inference library (GPU)

14.4k
206

NVIDIA's high-performance deep learning inference SDK (GPU)

13.2k

CUDA C++ templates for high-performance matrix-multiply (GEMM) and convolution kernels

10.2k
5

C++ LLM/VLM inference runtime for Jetson and NVIDIA edge devices

502
2

JAX (Google DeepMind)

1 repo·36.2k·151 commits this week

Composable NumPy transformations (JIT, grad, vmap) compiled via XLA to GPUs and TPUs — primary DeepMind research/production runtime

36.2k
151

SGLang

1 repo·31.8k·385 commits this week

High-throughput LLM/VLM serving with RadixAttention and structured generation

31.8k
385

Tencent

2 repos·28.3k·4 commits this week

High-performance neural network inference for mobile (Android/iOS)

23.7k
4

Tencent Neural Network — mobile and edge inference

4.6k

Modular

1 repo·26.8k·226 commits this week

Modular Platform monorepo — MAX inference server/framework + Mojo programming language for portable, high-performance AI on CPUs and GPUs

26.8k
226

Mozilla AI

1 repo·25.5k

Single-file LLM executables via Cosmopolitan Libc — zero install, all platforms (~21K stars)

25.5k

Dao AI Lab

1 repo·24.7k·7 commits this week

Official FlashAttention — fast, memory-efficient exact attention (FA-2/FA-3) kernels for GPUs

24.7k
7

Triton Language (OpenAI)

1 repo·19.9k·49 commits this week

Python-like GPU kernel language used by vLLM FlashAttention and PyTorch inductor

19.9k
49

KVCache AI

1 repo·19.2k·1 commits this week

CPU-GPU hybrid inference; runs DeepSeek 671B on 14GB VRAM + 382GB DRAM with massive speedup over llama.cpp

19.2k
1

jundot

1 repo·18.7k·79 commits this week

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

18.7k
79

MLC AI

1 repo·18.6k

High-performance LLM inference in web browsers via WebGPU

18.6k

Alibaba

1 repo·15.9k·44 commits this week

Alibaba's neural network inference framework for mobile & edge

15.9k
44

Blaizzy (Community MLX)

5 repos·14.5k·61 commits this week

Audio models (TTS, ASR) with MLX

7.7k
38

Vision-language models on Apple silicon via MLX

5.3k
23

Swift audio inference using MLX

753

Text embedding models with MLX

424

Video model inference with MLX

289

K2 / Next-gen ASR

1 repo·14.2k·20 commits this week

ONNX-based runtime for ASR, TTS, VAD, and keyword spotting

14.2k
20

Apache

2 repos·14.1k·12 commits this week

Apache TVM ML compiler — auto-tunes models for any hardware target

13.7k
10

Apache TVM Foreign Function Interface for deep learning compilation

445
2

OpenVINO Toolkit / Intel

3 repos·12.4k·110 commits this week

Intel's toolkit for optimizing & deploying deep learning on Intel hardware

10.6k
86

Neural Network Compression Framework — quantization, pruning, sparsity for OpenVINO

1.2k
3

OpenVINO GenAI — generative AI layer with speculative decoding & KV-cache opt

571
21

RunAnywhere

8 repos·11.9k·95 commits this week

RunAnywhere SDKs for on-device inference deployment

10.3k
46

RunAnywhere CLI tool

1.5k

Swift iOS/macOS starter example for the RunAnywhere SDK

12
13

Browser WASM reference app for the RunAnywhere Web SDK

10
10

React Native starter app using the RunAnywhere SDK

9

Android starter example for the RunAnywhere SDK

8
12

Flutter starter app for the RunAnywhere on-device SDK

6

Electron desktop app for on-device LLM, VLM, STT, TTS, VAD, and RAG

1
14

AMD Ryzen AI (XDNA NPU)

7 repos·11.7k·142 commits this week

Lemonade SDK — high-level, multi-vendor LLM inference SDK (OGA + llama.cpp), OpenAI-compatible server mode

5.3k
39

Xilinx/AMD AI Engine dev stack — source of the Vitis AI Execution Provider for XDNA

1.8k

AMD-backed Linux NPU runtime, used alongside Lemonade Server on Ryzen AI XDNA 2

1.7k
44

Open-source RAG/chat/agent app for Ryzen AI NPU, built on Lemonade SDK

1.5k
57

Official Ryzen AI Software stack — Vitis AI EP, quantization, LLM deployment; contains an experimental XDNA-NPU llama.cpp fork

868
2

Ingests any PyTorch model, optimizes, benchmarks/deploys across hardware targets incl. XDNA NPU

243

AMD's quantization toolkit (INT4/INT8 PTQ + QAT) for Ryzen AI NPU deployment

154

Intel

2 repos·11.6k

Intel IPEX-LLM — local LLM acceleration on Intel hardware (archived Jan 2026, read-only)

8.9k

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity

2.7k

Triton Inference Server

1 repo·10.9k·1 commits this week

NVIDIA Triton — production multi-model inference server (HTTP/gRPC, multi-backend)

10.9k
1

Mistral AI

1 repo·10.8k

Official minimal inference library for all Mistral models (7B, Mixtral, Pixtral)

10.8k

Qualcomm

3 repos·9.9k·104 commits this week

On-device LLM/VLM SDK for Snapdragon NPU, GPU, and CPU (formerly Nexa AI)

8.3k
71

State-of-the-art ML models optimized for Qualcomm Snapdragon NPU/DSP/QNN deployment

1.2k
29

Sample apps and tutorials for deploying models on Qualcomm hardware (TFLite, ONNX, QNN)

444
4

Dusty-NV (NVIDIA Jetson)

1 repo·9.0k

DNN inference library & tutorials for NVIDIA Jetson

9.0k

BentoML

1 repo·8.8k

Unified serving framework: real-time APIs, task queues, batching, multi-model chains

8.8k

InternLM / Shanghai AI Lab

1 repo·8.0k·16 commits this week

High-throughput LLM serving with TurboMind engine (C++/CUDA)

8.0k
16

AI Dynamo (NVIDIA)

1 repo·7.8k·153 commits this week

Datacenter-scale distributed inference serving framework (Rust + Python, disaggregated prefill/decode, engine-agnostic)

7.8k
153

Osaurus

1 repo·7.6k·52 commits this week

Native macOS AI agent harness in Swift — any model, persistent memory, autonomous execution, MCP server, MLX + Apple Neural Engine, fully offline

7.6k
52

PaddlePaddle (Baidu)

1 repo·7.3k

Lightweight inference engine for mobile & embedded from PaddlePaddle

7.3k

ArgMax

4 repos·6.7k

On-device Whisper inference for Apple platforms (Swift)

6.3k

Python tooling for WhisperKit model optimization

246

On-device AI benchmarking framework

88

Swift playground for ArgMax SDK

21

FlashInfer

1 repo·6.2k·48 commits this week

High-performance GPU kernel library for LLM serving — attention, sampling, and KV-cache primitives

6.2k
48

Cactus Compute

5 repos·6.1k·10 commits this week

Cactus core edge inference framework

5.7k
10

React Native bindings for Cactus

178

Kotlin/Android bindings for Cactus

74

Flutter bindings for Cactus

71

Demo chat app using Cactus

28

OpenNMT

1 repo·4.6k

Fast C++ inference for Transformer models; INT8/INT16 CPU quantization, multi-platform

4.6k

TurboDeRP (ExLlamaV2)

1 repo·4.6k

High-performance EXL2-quantized inference for consumer NVIDIA GPUs

4.6k

OpenXLA

1 repo·4.5k·266 commits this week

Compiler for JAX, TF, PyTorch targeting GPU, TPU, and CPU from a unified IR

4.5k
266

ModelTC

1 repo·4.2k·16 commits this week

Lightweight, high-throughput Python-based LLM inference and serving framework

4.2k
16

Predibase

1 repo·3.8k

Multi-LoRA inference server — serve thousands of fine-tuned adapters on a single GPU

3.8k

Liquid AI

5 repos·3.2k·4 commits this week

Examples, tutorials and apps for Liquid AI LFM + LEAP SDK

2.4k
3

Speech-to-Speech audio models by Liquid AI

560

Minimal fine-tuning repo for LFM2, fully open-source

197

Example apps for LeapSDK

74

Liquid AI documentation

30
1

Luminal AI

1 repo·2.9k·4 commits this week

Rust-based deep learning compiler with a small static graph IR for fast, portable inference (CUDA, Metal, CPU)

2.9k
4

Fluid Inference

3 repos·2.8k·8 commits this week

On-device audio inference framework

2.6k
8

Fluid Inference core runtime

85

Rust text processing library for inference

46

Try Mirai

2 repos·1.8k·22 commits this week

Mirai's on-device inference runtime

1.7k
18

Mirai's LLaMA-based on-device model

87
4

UbiquitousLearning

1 repo·1.6k·6 commits this week

Multimodal LLM inference framework for mobile & edge

1.6k
6

ARM Software

1 repo·1.3k

ARM Neural Network SDK for ARM & Mali devices

1.3k

AMD ROCm

4 repos·1.2k·126 commits this week

AI Tensor Engine for ROCm — centralized repo for high-perf AI operators on AMD Instinct GPUs

528
72

AMD's graph inference engine for MI-series GPUs

321
9

ROCm fork of FlashAttention with Composable Kernel (CK) and Triton backends

237
6

AiTer Optimized Model — lightweight vLLM-like server built on AITER kernels for ROCm

156
39

Rudrank Riyam

1 repo·1.2k·2 commits this week

iOS/macOS workbench for Apple's Foundation Models framework — recipes, guided labs, prompt playground, AFM CLI, FMFBench eval suite, and reusable Swift packages for on-device Apple Intelligence

1.2k
2

NimbleEdge

2 repos·528

NimbleEdge's deliteAI on-device inference framework

526

NimbleEdge fork of ExecuTorch with edge optimizations

2

ThunderAgent

1 repo·410

A simple, fast and robust program-aware agentic inference system

410

Picovoice

1 repo·316·1 commits this week

Picovoice's on-device LLM inference engine

316
1

Zetic AI

5 repos·78

MLange sample applications

70

MLange extension library

5

iOS framework for MLange

2

iOS extension framework for MLange

1

MLange SDK documentation

0