Ollama
simple local inference + model management
WindowsmacOSLinux
The inference layer changes performance, compatibility and deployment characteristics.
simple local inference + model management
GGUF inference across heterogeneous hardware
high-throughput server inference
self-hosted OpenAI-compatible local inference
high-performance LLM/VLM serving
Apple Silicon local inference and fine-tuning
NVIDIA optimized inference
quantized NVIDIA GPU inference
portable GGUF local inference server
multi-backend local model serving/UI