Directory

Runtimes

The inference layer changes performance, compatibility and deployment characteristics.

Runtime

Ollama

simple local inference + model management

WindowsmacOSLinux
Runtime

llama.cpp

GGUF inference across heterogeneous hardware

WindowsmacOSLinux
Runtime

vLLM

high-throughput server inference

Linux
Runtime

LocalAI

self-hosted OpenAI-compatible local inference

LinuxmacOSWindows via containers
Runtime

SGLang

high-performance LLM/VLM serving

Linux
Runtime

MLX LM

Apple Silicon local inference and fine-tuning

macOS
Runtime

ExLlamaV2

quantized NVIDIA GPU inference

LinuxWindows
Runtime

KoboldCpp

portable GGUF local inference server

WindowsLinuxmacOS