Manage models, download new ones, and chat with your local AI
Load a different model in 10-15 seconds. The service will restart with the new model.
Open a real-time chat interface to interact with the currently active model.
Add new models that fit within 5-6GB RAM. Find direct download URLs from HuggingFace or Ollama libraries.
Use these endpoints to integrate with external tools and applications.
POST /v1/chat/completionsGET /healthGET /api/modelsThis hub runs llama.cpp with local GGUF models. Only one model loads at a time to fit within the 5-6GB RAM budget.
Installed Models: 4
Total size: 13.4 GB