Horizon

Expose a local Ollama API

Expose a local Ollama API on a public URL with a Horizon tunnel, including streaming responses and browser access with OLLAMA_ORIGINS.

Call your local Ollama models from another machine or a web app at https://my-app.hrzn.run. Streaming responses reach the caller as Ollama writes them.

Before you begin

  • Node.js 18 or later
  • A Horizon account and the CLI (see Getting started)
  • A reserved subdomain for -s. Reserve one on the Subdomains page.
  • Ollama installed, with one model pulled, for example ollama pull llama3.2

Start Ollama

Ollama listens on 127.0.0.1:11434 by default. Start it if it isn't already running.

ollama serve

Leave OLLAMA_HOST alone. The Horizon CLI connects from your machine, so Ollama doesn't need to listen on your network. Setting OLLAMA_HOST=0.0.0.0:11434 would also open the API to everyone on your local network.

Start a tunnel

Use -s with a subdomain you reserved. Without it, the subdomain is random and changes every run. Reserved subdomains are a paid feature, see Pricing.

hrzn tunnel http://localhost:11434 -s my-app
Output
HORIZON: Tunnel connected
  URL          https://my-app.hrzn.run (reserved)
  Forwarding   http://localhost:11434
  Request log  https://hrzn.run/dashboard/tunnels/my-app

Horizon sets the Host header to the target, localhost:11434. Ollama sees a local request, so you don't need a host-header flag as you would with some other tunnel tools.

Allow browser clients

A web app on another origin needs permission to call Ollama from the browser. Ollama allows cross-origin requests from 127.0.0.1 and 0.0.0.0 by default. Add your app's origin with OLLAMA_ORIGINS, then restart Ollama.

OLLAMA_ORIGINS=https://my-web-app.example.com ollama serve

If you run the Ollama desktop app on macOS, set the variable with launchctl setenv OLLAMA_ORIGINS "https://my-web-app.example.com" and restart the app. On Linux with systemd, run systemctl edit ollama.service, add Environment="OLLAMA_ORIGINS=https://my-web-app.example.com" under [Service], then run systemctl daemon-reload and systemctl restart ollama.

Calls from curl or a server don't need this step.

Check it works

Ask for a streamed answer. -N stops curl from buffering the output.

curl -N https://my-app.hrzn.run/api/generate -d '{"model": "llama3.2", "prompt": "Why is the sky blue?"}'

Ollama streams newline-delimited JSON (application/x-ndjson). You see one JSON object per line as the model writes, and the last one has "done":true:

Output
{"model":"llama3.2","created_at":"...","response":"The","done":false}

Your Horizon terminal prints one line for the request:

Output
  POST    200  /api/generate

To get one JSON response instead, add "stream": false to the request body.

Troubleshooting

The browser blocks the request with a CORS error

Your web app's origin isn't in OLLAMA_ORIGINS. Set it to the exact origin, scheme included, and restart Ollama. Check it with the browser's network tab, not with curl, which doesn't send an Origin header.

The response arrives all at once

Check that the request doesn't set "stream": false. Ollama streams by default.

The URL changed after a restart

You started the tunnel without -s, so Horizon gave you a new random subdomain. Restart with -s my-app and the URL stays the same. -s needs a subdomain you reserved, see Pricing.

Nothing reaches Ollama

  • Check that the Horizon terminal is still running.
  • Check that ollama serve is running on port 11434.

Next steps

  • Read the Ollama FAQ on OLLAMA_HOST and OLLAMA_ORIGINS.
  • Read the streaming reference.

On this page