⚡ LLM Admin Hub

Manage models, download new ones, and chat with your local AI

đŸŸĸ Active: capybarahermes-2.5-mistral-7b.Q3_K_S.gguf

📚 Your Models

capybarahermes-2.5-mistral-7b.Q3_K_S.gguf
3.16 GB
✓ Currently Active
Dolphin3.0-Llama3.2-3B-Q4_K_M.gguf
2.02 GB
Load
Llama-3.1-8B-abliterated-Q4_K_M.gguf
4.92 GB
Load
VibeThinker-3B-Q8_0.gguf
3.29 GB
Load

🔄 Switch Model

Load a different model in 10-15 seconds. The service will restart with the new model.

đŸ’Ŧ Chat

Open a real-time chat interface to interact with the currently active model.

âŦ‡ī¸ Download Model

Add new models that fit within 5-6GB RAM. Find direct download URLs from HuggingFace or Ollama libraries.

💡 Recommended uncensored/capable models (5-6GB fit):
â€ĸ Qwen2.5-3B Q8_0 (3.4GB) - Very capable, not jailbroken
â€ĸ Phi-4-mini Q8_0 (4.1GB) - Strong reasoning, minimal restrictions
â€ĸ Gemma-3-4B Q8_0 (4.3GB) - Good balance of capability and speed
â€ĸ Llama-3.2-8B Q4_K_M (3.8GB) - Larger model, still fits
â€ĸ Mistral-7B Q4 (3.5GB) - Fast and capable
Find GGUF files at:
🔗 LM Studio GGUF Collection
🔗 bartowski GGUF Repo

🔗 API Reference

Use these endpoints to integrate with external tools and applications.

Chat API:
POST /v1/chat/completions
Health:
GET /health
Models:
GET /api/models

â„šī¸ About

This hub runs llama.cpp with local GGUF models. Only one model loads at a time to fit within the 5-6GB RAM budget.

Installed Models: 4

Total size: 13.4 GB