deltafin
by gavamedia · github.com/gavamedia/deltafin
Run Kimi K3, a 2.8T-parameter Mixture-of-Experts LLM, on a single Apple Silicon Mac. Streams MXFP4 experts on demand over HTTP into a local disk cache — fused NEON kernels, Metal/MPS compute, exact reproducible decoding, and an OpenAI-compatible API server for local chat and coding agents.
Stars
★ 328
Forks
29
Language
Python
Category
AI Agents
License
NOASSERTION
kimikimi-k3local-ailocal-llmpythonopenaiagentllm
Featured in
- deltafin @ 14:26 · Jul 29, 2026