Expose a local Ollama API
Expose a local Ollama API on a public URL with a Horizon tunnel, including streaming responses and browser access with OLLAMA_ORIGINS.
Call your local Ollama models from another machine or a web app at https://my-app.hrzn.run. Streaming responses reach the caller as Ollama writes them.
The Ollama API has no authentication for local requests. Anyone who has your tunnel URL can run your models, use your GPU and read your model list. Horizon has no password protection or IP allowlist. Stop the tunnel when you aren't using it, and don't post the URL anywhere public.
Before you begin
- Node.js 18 or later
- A Horizon account and the CLI (see Getting started)
- A reserved subdomain for
-s. Reserve one on the Subdomains page. - Ollama installed, with one model pulled, for example
ollama pull llama3.2
Start Ollama
Ollama listens on 127.0.0.1:11434 by default. Start it if it isn't already running.
ollama serveLeave OLLAMA_HOST alone. The Horizon CLI connects from your machine, so Ollama doesn't need to listen on your network. Setting OLLAMA_HOST=0.0.0.0:11434 would also open the API to everyone on your local network.
Start a tunnel
Use -s with a subdomain you reserved. Without it, the subdomain is random and changes every run. Reserved subdomains are a paid feature, see Pricing.
hrzn tunnel http://localhost:11434 -s my-appHORIZON: Tunnel connected
URL https://my-app.hrzn.run (reserved)
Forwarding http://localhost:11434
Request log https://hrzn.run/dashboard/tunnels/my-appHorizon sets the Host header to the target, localhost:11434. Ollama sees a local request, so you don't need a host-header flag as you would with some other tunnel tools.
Allow browser clients
A web app on another origin needs permission to call Ollama from the browser. Ollama allows cross-origin requests from 127.0.0.1 and 0.0.0.0 by default. Add your app's origin with OLLAMA_ORIGINS, then restart Ollama.
OLLAMA_ORIGINS=https://my-web-app.example.com ollama serveIf you run the Ollama desktop app on macOS, set the variable with launchctl setenv OLLAMA_ORIGINS "https://my-web-app.example.com" and restart the app. On Linux with systemd, run systemctl edit ollama.service, add Environment="OLLAMA_ORIGINS=https://my-web-app.example.com" under [Service], then run systemctl daemon-reload and systemctl restart ollama.
Calls from curl or a server don't need this step.
Check it works
Ask for a streamed answer. -N stops curl from buffering the output.
curl -N https://my-app.hrzn.run/api/generate -d '{"model": "llama3.2", "prompt": "Why is the sky blue?"}'Ollama streams newline-delimited JSON (application/x-ndjson). You see one JSON object per line as the model writes, and the last one has "done":true:
{"model":"llama3.2","created_at":"...","response":"The","done":false}Your Horizon terminal prints one line for the request:
POST 200 /api/generateTo get one JSON response instead, add "stream": false to the request body.
Troubleshooting
The browser blocks the request with a CORS error
Your web app's origin isn't in OLLAMA_ORIGINS. Set it to the exact origin, scheme included, and restart Ollama. Check it with the browser's network tab, not with curl, which doesn't send an Origin header.
The response arrives all at once
Check that the request doesn't set "stream": false. Ollama streams by default.
The URL changed after a restart
You started the tunnel without -s, so Horizon gave you a new random subdomain. Restart with -s my-app and the URL stays the same. -s needs a subdomain you reserved, see Pricing.
Nothing reaches Ollama
- Check that the Horizon terminal is still running.
- Check that
ollama serveis running on port 11434.