
Offline AI Studio
On-device AI for iPhone & iPad — chat, image generation, voice, and Core ML Stable Diffusion. Every feature is free. Hosted models are optional.
Android version: OfflineGPT on Google Play
What's new
Latest in v2.2.6
Every feature is free
The Pro purchase is retired: on-device models, API mode, document tools, custom tools, and all themes are open to everyone. Early supporters keep their Legacy badge.
Power Mode — a bigger sandbox
For API models with your own key: capable models can use more bundled Python packages (numpy, pandas, matplotlib, pypdf) and attach workspace files with instant previews. Switch it on in the plus menu.
Optional hosted models
Another optional provider next to your own API keys. Gemma 4 31B on Swiss infrastructure with 100% renewable power, and server heat reused for heating. GPT-6 Luna is new on Android and coming soon on iPhone. GPT-4o mini and DeepSeek Flash are included too. No key of your own.
Custom API providers improved
Add several custom providers, with a clearer setup flow, better automatic detection, and model lists that match what the host actually offers.
More in this release
Grouped web-search cards, a rounded sidebar, a 24-round API tool loop, live Python in chat, and many stability fixes.
What Offline AI Studio can do
iPhone & iPad · iOS 17+ · Core ML image generation · Metal GPU acceleration
Offline by default
Chats, image generation, and voice input stay on your device. No ads, no tracking. Anything leaves your iPhone only if you choose hosted models or your own API keys.
Gemma 4 E2B — multimodal
Google's on-device model understands text, images, and audio in one conversation. Optional web search when you want it.
Thinking models
Qwen3 (0.6B – 4B) and SmolLM3-3B show step-by-step reasoning before answering — good for problems that need a second look.
Image generation — Core ML
Stable Diffusion runs natively via Core ML on Apple Silicon. Three tiers: SD 1.5 Fast, SD 2.1 Balanced, SDXL Quality — up to 1024px.
Vision models (image understanding)
LFM2 1.6B reads and describes images in chat, in a standard and a higher-quality build. Send a photo and ask questions about it — fully on-device.
Whisper voice input
Dictate prompts with on-device speech recognition — a compact Whisper model ships with the app, an enhanced version is an optional download.
Read aloud / TTS
Answers are read back through the iOS system voice with no setup, or through the optional Supertonic engine with ten neural voices.
Web search
DuckDuckGo Lite needs no key and is the default. Brave and Tavily work with a key of your own. Results arrive as grouped cards in the chat.
Power Mode — Python in chat
For API models with your own key: run Python in a per-chat sandbox with numpy, pandas, matplotlib, pypdf and more bundled in, and attach workspace files with previews. Not available for on-device models.
AI writing tools
Summarize, grammar, tone, simplify, translate, social posts, email drafts, and brainstorming — plus document processing, included for everyone.
Code Lab
Prototype HTML, CSS, and JavaScript with a built-in live preview and an AI coding assistant alongside.
Encrypted chat backup
Export and import your full chat history as an encrypted .ogpt file — password protected, ready to import on another device.
Documents and OCR
Summarize PDFs, CSVs, Excel files and images. Text is read from images on-device through Apple Vision — nothing is sent to a cloud.
Custom AI tools
Build your own tools with a custom name, icon, system prompt, and optional knowledge base. Pin them to the home screen.
Your own API keys
Connect OpenAI, DeepSeek, Gemini, Groq, Mistral, Grok, OpenRouter, or any OpenAI-compatible host. Several custom providers at once, with model lists discovered automatically.
Projects for your chats
Group conversations into project folders in the sidebar instead of scrolling one long list.
16 themes
Light, Dark, Aurora, Cyber and a dozen more. Since the Pro purchase was retired, every theme is free.
On-device model catalog
Download sizes are approximate. Custom GGUF imports also supported.
Text models (GGUF · llama.cpp + Metal)
LFM2.5 1.2B
Liquid AI · 4 GB RAM
Llama-3.2-1B
Meta · 4 GB RAM
Gemma-2B
Google · 6 GB RAM
Thinking models (GGUF · llama.cpp + Metal)
Qwen3 0.6B
Alibaba · 4 GB RAM
Qwen3 1.7B
Alibaba · 6 GB RAM
Qwen3 4B
Alibaba · 6 GB RAM
SmolLM3-3B
Hugging Face · 6 GB RAM
Multimodal — text + image + audio (LiteRT)
Gemma-4-E2B-it
Google (LiteRT) · 8 GB RAM
Vision models — image understanding (GGUF + mmproj)
LFM2 1.6B
Liquid AI (GGUF + mmproj) · 6 GB RAM
LFM2 1.6B Q6
Liquid AI (GGUF + mmproj) · 8 GB RAM
Image generation (Stable Diffusion · Core ML)
SD 1.5 Fast
Core ML · up to 512px · 4 GB RAM
SD 2.1 Balanced
Core ML · up to 768px · 6 GB RAM
SDXL Quality
Core ML · up to 1024px · 8 GB RAM
Voice input (whisper.cpp)
Whisper base Q8 (bundled)
whisper.cpp
Whisper base full (optional)
whisper.cpp
System voice (TTS)
AVSpeech — no download
Supertonic (TTS, optional)
10 neural voices
Everything in the app is free
The former Pro purchase is retired. Hosted models are an optional subscription.
Included
- Unlimited on-device models — text, thinking, vision, and image generation
- Custom GGUF imports and the in-app Hugging Face browser
- API mode with your own keys, and several custom providers
- Full context window your device can carry, and editable system prompts
- Document tools, custom AI tools, Code Lab, and Power Mode
- All 16 themes, and no ads or tracking anywhere
Hosted models
An optional monthly subscription. Gemma 4 31B is hosted in Switzerland on 100% renewable power, with server heat reused for heating. GPT-6 Luna is new on Android and coming soon on iPhone. GPT-4o mini and DeepSeek Flash are included too. No key of your own, no account, and unlimited personal use under fair use. Local models and your own API keys keep working without it.
Billed through the App Store; the price shown there for your region applies. Cancel anytime.
Learn more about hosted modelsFrequently asked questions
Do I need internet?
To download models and app updates, yes. Once downloaded, local chat, image generation, and voice input work fully offline. Optional: web search, API mode, and Power Mode need a connection when used.
Is Offline AI Studio free?
Yes, and every feature is. The former one-time Pro purchase is retired — on-device models, API mode with your own keys, document tools, custom tools, and all themes are open to everyone. People who bought Pro early keep a Legacy badge. The only paid option is the hosted-models subscription, and nothing depends on it.
What are hosted models?
An optional monthly subscription. Gemma 4 31B is hosted in Switzerland on 100% renewable power. GPT-6 Luna is new on Android and coming soon on iPhone. GPT-4o mini and DeepSeek Flash are included too. It needs no API key and no account, and it sits next to your own keys rather than replacing them. Billed through the App Store at the price shown for your region. More about hosted models.
Can the app run Python?
Yes, through Power Mode — but only when you chat with API models using your own key, not with on-device models. Scripts run in a sandbox per chat, with numpy, pandas, matplotlib, pypdf and more bundled in. Installing extra packages is not possible on iOS, and access to local network addresses is blocked.
What is Gemma 4 E2B?
Google's multimodal on-device model — understands text, images, and audio attachments in a single conversation. Runs via LiteRT-LM with an optional web search tool you can toggle on or off.
How does image generation work on iOS?
Stable Diffusion runs via Core ML — natively optimised for Apple Silicon and the Neural Engine. Three tiers are available: SD 1.5 Fast (~512px), SD 2.1 Balanced (~768px), and SDXL Quality (up to 1024px, requires iOS 17 and 8 GB RAM). No MNN, no cloud processing.
Is there an Android version?
Yes. On Google Play the app is called OfflineGPT — same offline-first philosophy, built for Android with llama.cpp, MNN, and Snapdragon NPU support.
Which text models are available?
LFM2.5 1.2B, Llama-3.2-1B, Gemma-2B (chat models); Qwen3 0.6B/1.7B/4B and SmolLM3-3B (thinking models); LFM2 1.6B (vision); Gemma 4 E2B (multimodal). You can also import your own GGUF files or browse Hugging Face inside the app.
How does privacy work?
By default, all inference, chat history, and voice input stay on your device. Data only leaves when you use optional features: model downloads, Gemma 4 web search, or API mode with your own keys. Full privacy policy.
