Powered by WebLLM + WebGPU. Models run entirely in your browser — nothing is sent to a server.
1. Choose a model & click Load model
2. First load downloads weights (cached afterwards)
3. Chat offline with GPU acceleration