Inference provider · Inference platform
Modal Inference
Verified Sep 2026
Quick answer: Code-first serverless inference where you deploy your own OpenAI-compatible endpoints (`modal endpoint create`) and scale GPU compute to zero — including Shared Endpoints for OpenCode/Codex.
Best for
Code-first serverless inference where you deploy your own OpenAI-compatible endpoints (`modal endpoint create`) and scale GPU compute to zero — including Shared Endpoints for OpenCode/Codex.
Skip if
You just need a paste-key multi-lab gateway (OpenRouter) or want zero endpoint provisioning (Together/Fireworks-class API).
Who it fits
- Code-first serverless inference where you deploy your own OpenAI-compatible endpoints (`modal endpoint create`) and scale GPU compute to zero — including Shared Endpoints for OpenCode/Codex.
- Builders who start in the browser without a local IDE setup
- Terminal-first engineers who live in the shell and git
Pros
Code-first serverless inference — `modal endpoint create --model <repo>` spins an OpenAI-compatible endpoint you own; scale to 1000+ GPUs then to zero; custom weights from Hugging Face or Modal Volume; OpenCode and Codex can point at Shared Endpoints (`https://inference.us-west.modal.direct/v1`); $30/month free compute on Modal (vendor claim).
Cons
Heavier than paste-key-into-Cursor (Baseten-class ops — you provision endpoints); Shared Endpoint plan credits no longer pay Shared Endpoint usage from 2026-09-01 (other credits still apply); not a multi-provider gateway like OpenRouter; not hands-on verified here.
Facts
- Locality:
- Cloud-native
- Surfaces:
- Web, CLI
- Maturity:
- Experimental
- Verified:
- Sep 2026
Agent standards & memory
Compare one neighbor first
Vendor pages sell hard. Read one alternative on this site, then leave for the official link if the fit still holds.
Official links
Alternatives
Someone asks you about vibe-coding software? Share this site with them — we are adding more useful data regularly. agents.dancingteeth.net