Self-hosted · Free · Your keys, your data

Every AI model.
One beautiful home.
Your hardware.

Promptly is a self-hosted AI workspace for you and your people — chat with any model, research with live web search, build shared Workspaces, learn with an AI tutor, talk out loud, and automate the boring parts. All of it runs on your own box.

No telemetry, ever Bring your own API keys One command — ./install.sh
promptly.yourdomain.com
Message Promptly… Web Claude Sonnet ▾
Bring your own keys

One picker, every model you care about — cloud or fully local.

Anthropic Claude OpenAI GPT Google Gemini OpenRouter · 300+ models DeepSeek Ollama · fully local Any OpenAI-compatible endpoint
Everything in the box

A full AI workspace, not just a chat window

Promptly ships the features you'd normally stitch together from five different subscriptions — integrated, consistent, and private.

New

One question. A whole research team.

Ask something broad and the model fans it out with the run_agents tool: up to four sub-agents research in parallel — each with its own web search, page reading and private working context — and hand back one merged brief with de-duplicated citations.

  • Genuinely concurrent — a comparison that would take a dozen sequential searches lands in one tool call.
  • Your chat stays lean — each agent reads pages into its own context and returns only a bounded digest, so raw page text never floods the conversation.
  • Watch it work — every agent reports live progress in the chat's Tool Activity card, exactly like below.
Tool Activity
run_agents Fanning out 3 research agents…
Agent 1 Pixel 10 Pro — camera & battery reviews
Agent 2 iPhone 17 — camera & battery reviews
Agent 3 Galaxy S26 — camera & battery reviews
Merged 3 briefs · 11 sources · citations de-duplicated
Workspaces

Chats, notes, canvas, boards & sheets — one shared brain

Everything inside a Workspace is indexed into the same retrieval layer. Ask a chat about your notes, your board, or a cell in your sheet, and it just knows. Click through the sidebar — the board really drags, and any item opens to the side.

Workspaces · Product Launch
Product Launch
Launch plan — v2
Launch plan — v2

Rich collaborative docs with live cursors (Y.js CRDT under the hood), tables, callouts and task lists. The AI can propose edits directly into a note.

  • Draft announcement post
  • Confirm pricing page copy
  • Record 90-second demo video
  • Prep launch-day support rota
💡 AI suggestion: the retro flagged checkout drop-off as the top risk — consider moving that fix before the launch date.
Idea backlog
Q3 launch scope
Ship it 🚀
Talk to 5 beta users before freeze
Pricing test: annual toggle default ON?
Demo video — keep under 90s
To do 3
Ship onboarding email sequence
launch
Fix checkout flow drop-off
buglaunch
Schedule pricing-page A/B test
In progress 2
Demo video storyboard
content
Docs refresh for new nav
Done 2
Announcement post draft ✍️
Q3 retro summarised by AI
ai
D5 =SUM(D2:D4)
ABCD
1ChannelLeadsConv %Revenue
2Newsletter1,2404.1%$18,420
3Product Hunt3,8672.3%$31,050
4Referrals6429.8%$22,310
5Total5,749$71,780
Team
MMia
KKai
AAva
LLeo
Drag a teammate onto a day
Mon 8
MMia09:00–17:00
Tue 9
KKai12:00–20:00
AAva09:00–17:00
Wed 10
LLeo09:00–17:00
Thu 11
MMia09:00–17:00
KKai12:00–20:00
Fri 12
AAva09:00–17:00
Which launch channel should we double down on, based on the sheet?
Claude Sonnet · workspace context
Referrals convert at 9.8% — more than four times Product Hunt — but only bring 642 leads. The sheet plus your retro note suggest the play is a referral incentive: highest conversion, most headroom, and it directly addresses the “growth felt rented” concern from the retro.
Launch metrics.sheet Q3 Retro.md
Automations

Your AI, on a schedule

Build node-graph flows that run while you sleep: cron and webhook triggers, HTTP calls to any API with a built-in credentials vault, AI steps, branching logic — and outputs that write straight back into your Workspaces.

  • Triggers your way — cron schedules, manual runs, or webhooks from CI, forms and monitors.
  • Credentials vault — API keys stored encrypted, referenced as {{secret.NAME}}, never echoed back.
  • Durable runs — a dedicated worker queue keeps flows alive through redeploys, with full run history.
  • Writes where you work — post results as board cards, notes, sheet rows or chat messages.
Flow · Morning competitor digest
Schedule
Every weekday
0 7 * * 1-5
HTTP request
GET pricing API
{{secret.RAPID_KEY}}
AI summarise
Compare vs yesterday, flag changes > 5%
Condition
Only continue if changes > 0
Board card
Post digest to Pricing watch board
Voice

Talk to it. It talks back. Nothing leaves your network.

Dictate messages with self-hosted Whisper speech-to-text and have replies read aloud by Kokoro text-to-speech — both running as local containers beside the app. Full voice conversations without a single audio byte going to a third party.

  • Dictation everywhere — tap the mic in any composer and speak your message.
  • Read-aloud replies — natural local TTS on any assistant message, with a hands-free voice mode.
  • Private by architecture — Whisper and Kokoro run on internal-only Docker networks; cloud STT is strictly opt-in.
Voice mode

Tap the orb to start the voice demo

Files

A real drive, wired into every conversation

Folders, starring, trash with undo, full-text search — plus everything the AI generates lands here automatically, linked back to the chat that created it. Mention any file in a message with @.

  • Paranoid on purpose — magic-byte validation, EXIF/GPS stripping, filename sanitisation and per-user quotas on every upload.
  • PDF, DOCX, images, data files — attach anything; text-only models even get AI-written captions for images via the vision relay.
  • Findable forever — Recent, Starred, Shared and Trash views plus search across your whole library.
Files · My Drive
NameKindModified
GeneratedFolderToday
UploadsFolderYesterday
Q3 board report.pdfPDF · AI-generated2 min ago
hero-illustration.pngImage · AI-generated1 hr ago
survey-results-clean.csvData · Code InterpreterTue
Q3 Retro.mdMarkdownMon
In the box

Plus everything else you’d expect

No add-ons, no per-seat upsells — the whole toolkit ships in the same one-command install.

Deep Research

An agentic pipeline that decomposes your question, runs parallel web searches, reads the actual pages, checks for gaps, and streams back a structured, cited report.

Code Interpreter

Models write and run real Python in a network-isolated sandbox. Attach a CSV, ask a question, get back charts and clean data — with the output files saved to your drive.

Generated artifacts

Ask for a PDF report, an image, or a data export and it lands in your Generated folder — linked back to the conversation that made it.

Memory that earns its keep

Promptly quietly captures durable facts across chats — pinned facts always apply, the rest are retrieved semantically. Capped, deduplicated, and fully yours to edit.

Web search, three ways

Self-hosted SearXNG out of the box — no API key needed. Prefer Brave or Tavily? Flip a switch. Per-conversation off / auto / on control.

Truly multi-user

Invite-only accounts for your family, team, or study group. Share chats read-only or invite collaborators right into the thread.

Context you can see

A live context-window meter on every chat, one-click compaction when it fills up, and truncation detection so a cut-off answer never slips past you.

Saved prompts

Your best prompts become slash commands. Type / in the composer and fire a template instantly.

Installable PWA

Add it to your phone's home screen for a real app feel — with push notifications when a long answer or scheduled task finishes while you're away.

Subchats

Chase a tangent in a floating throwaway side-chat without derailing the main thread. Get your answer, close it, carry on — nothing polluted.

Custom AI personas

Wrap any base model in a curated system prompt plus a knowledge library. "Support Bot on GPT-4o" shows up in everyone's picker, grounded in your docs.

Find anything, fast

Hit Ctrl+K and search every conversation you own or share — full-text with semantic re-ranking, snippets highlighted, one click to jump to the exact message.

Chat folders

Group your chats into collapsible folders — each with its own system prompt and default model that every chat inside quietly inherits.

Run it like you mean it

Multi-user security and admin tools that take it seriously

Not a toy deployment. Promptly ships the auth, audit and observability you'd expect from a product you pay for — because your household or team deserves it.

Real MFA

Authenticator TOTP, email one-time codes, hashed backup codes and 30-day trusted devices. Admins can require MFA org-wide.

Invite-only accounts

No open signup. One-time invites, account lockout after repeated failures, and rate limiting on every sensitive endpoint.

Usage analytics

Tokens, messages and dollar cost by day, user and model — so you always know exactly what your API keys are spending.

Full audit trail

Every login, lockout, MFA attempt and rate-limit trip is recorded and filterable in the admin panel.

Live console

Structured JSON logs with request context, error grouping and stack traces — streamed right into the admin UI.

Model governance

Enable or disable individual models per provider, set defaults, and see OpenRouter privacy badges (training & retention policies) before you let a model near your data.

Self-host

One command. Whole platform.

Clone the repo and run ./install.sh — it checks prerequisites, generates your secrets, and brings up the entire stack: app, Postgres with pgvector, Redis, local search, the code sandbox, voice services and the automation worker. Migrations run themselves. A first-run wizard in the browser creates your admin account.

  • Air-gap friendly — add --with-ollama and it all runs local: models, embeddings, the bundled SearXNG search and Whisper/Kokoro voice. Zero cloud, no API keys.
  • Everything bind-mounted — your database, uploads and models live in one ./data folder. Back it up, move it, own it.
  • Health-checked — a single endpoint pings Postgres, Redis and search so your uptime monitor always knows the truth.
Pick your install — copy grabs the full command
Recommended
Cloud-first — connect your API keys in the wizard. Auto-uses a native Ollama if one's running on the host.
./install.sh
With bundled Ollama
Run local models in-stack. GPU auto-detected: NVIDIA (CUDA), AMD (ROCm), else CPU.
./install.sh --with-ollama
Without bundled search
Skip SearXNG — add a Brave or Tavily key later in the admin panel.
./install.sh --no-search

0

models via OpenRouter alone

0

services in one compose file

0

command to deploy

0

bytes of telemetry sent anywhere

Warm by design

A calm, warm interface — day and night

No harsh blues, no clinical greys. Promptly's palette is warm paper and terracotta, tuned carefully for both themes. This whole page uses the app's exact design tokens — hit the toggle in the nav to feel it.

Light — warm paper
Dark — warm ember
Pricing

Own it, or let us run it

Self-host free forever on your own hardware — every feature, no seat limits. Or skip the ops entirely and let us run a private, managed instance for you.

Self-hosted
$0/ forever

Run the whole platform on your own box. Your data never leaves your network.

  • Every feature — nothing gated
  • Unlimited users on your hardware
  • Bring your own model keys, or run fully local with Ollama
  • One-command install & updates
  • Community support
Install now
Questions

Frequently asked

What does Promptly cost?
Nothing. Promptly is free, self-hosted software. The only bills you'll ever see are from the AI providers you choose to connect — and if you run local models with Ollama, even that is zero.
What do I need to run it?
Any machine that runs Docker — a home server, a NAS, a small VPS. Clone the repo and run ./install.sh: it checks prerequisites, generates secrets, brings up the full stack, and a first-run wizard walks you through creating the admin account. A GPU is optional and only matters if you want fast local models.
Can it run completely offline?
Yes — the default is cloud-first, but it's built to go fully local. Install with --with-ollama (or point it at an Ollama you already run) for chat models and embeddings; the SearXNG web search and Whisper + Kokoro voice containers are already local. In that setup, not a single request leaves your network.
Which AI models can I use?
Anthropic Claude, OpenAI GPT, Google Gemini, DeepSeek (including its reasoning modes), anything on OpenRouter (300+ models), local models via Ollama, and any server that speaks the OpenAI-compatible API — vLLM, LM Studio, LocalAI and friends. Admins choose exactly which models are enabled for the household.
Is my data used to train anything?
Promptly itself sends zero telemetry and stores everything in your own Postgres database. What cloud providers do with API traffic is governed by their policies — which is why the admin panel surfaces OpenRouter's per-model privacy badges (training, retention, ZDR) so you can make that call per model.
How many people can use one install?
It's designed for small groups — a family, a team, a study circle. Registration is invite-only, every user gets their own conversations, files, memory and quotas, and admins get per-user usage and cost breakdowns.
Promptly

Your models. Your data. Your box.

Stop renting your AI workspace. Promptly gives you the whole platform — chat, research, Workspaces, tutoring, voice and automations — on hardware you control.