OfflineGPT

OfflineGPT

On-device AI for Android — chat, images, voice, TTS, and a Python sandbox for API models. OfflineGPT Core is free. Private Cloud is optional.

Get it on Google PlayDownload on the App Store

iOS version: Offline AI Studio on the App Store

NEW · v2.9.5

Python sandbox. With network.

When you chat with API models (your keys, your provider), OfflineGPT can run Python in a sandboxed runtime — with optional network access. Call APIs, pull live data, automate the boring stuff. Not for local downloads like Gemma 4 or GGUF models — those stay pure on-device.

What's new

Latest in v2.9.5

v2.9.5

Python sandbox — with network

When you use API models (your keys, your provider), OfflineGPT can run Python in a sandboxed runtime — with optional network access to call APIs, pull data, and automate workflows. Not available for local on-device models like Gemma 4 or GGUF downloads.

New model

Gemma 4 E4B

Google's multimodal on-device model: text, image, and audio understanding in one conversation. Runs via LiteRT-LM with optional web search.

What OfflineGPT can do

Android 8.0+, 64-bit ARM. Models run on-device — CPU, GPU (Vulkan/OpenCL), or Snapdragon NPU.

Offline by default

Chats, generations, and voice input stay on your device by default. No ads, no tracking. Cloud processing happens only if you choose Private Cloud or your own API keys.

Python sandbox + network

For API models only: a sandboxed Python runtime with optional network access. Hit APIs, fetch data, automate workflows — when you chat with your own cloud keys. Not available for downloaded on-device models.

Gemma 4 E4B — multimodal

Google's on-device model understands text, images, and audio in one conversation. Optional web search when you want it.

Thinking models

Qwen3 (0.6B – 4B) and SmolLM3-3B show step-by-step reasoning before answering. Good for problems that need a second look.

Vision & image generation

Describe images with InternVL or SmolVLM2. Generate art with Stable Diffusion (Anything V5, AbsoluteReality, ChilloutMix, CuteYukiMix).

Snapdragon NPU acceleration

Image generation runs on the Hexagon NPU on supported Snapdragon devices — faster, cooler, less battery drain.

Whisper voice input

Dictate prompts in any language with on-device speech recognition. Compact and enhanced Whisper models are optional downloads — pick what you need.

Read aloud / TTS

Answers are read back via system voice or the optional Supertonic 3 neural engine — 31 languages, 10 voices, streaming during generation.

AI writing tools

Summarize, grammar, tone, simplify, social posts, document processing, email drafting, and Idea Blast brainstorming.

Code Lab

Prototype HTML, CSS, and JavaScript with a built-in live preview and an AI coding assistant alongside.

Offline OCR

Extract text from photos and PDFs on-device using ML Kit — no cloud OCR.

Custom AI tools

Build your own tools with a custom name, icon, system prompt, and optional file knowledge base. Pin them to the home screen.

On-device model catalog

Download sizes are approximate. Custom GGUF imports also supported.

Text models (GGUF · llama.cpp)

LFM2.5 1.2B

LiquidAI

~731 MB

Llama-3.2-1B

Meta

~1.4 GB

Gemma-2B

Google

~1.85 GB

Qwen2.5-3B

Alibaba

~2.1 GB

Rocket-3B

Community

~3.1 GB

Thinking models (GGUF · llama.cpp)

Qwen3 0.6B Thinking

Alibaba

~640 MB

Qwen3 1.7B Thinking

Alibaba

~1.8 GB

Qwen3 4B Thinking

Alibaba

~2.6 GB

SmolLM3-3B

Hugging Face

~1.9 GB

Multimodal — text + image + audio (LiteRT)

Gemma-4-E2B-it

Google (LiteRT)

~2.58 GB

Gemma-4-E4B-it

Google (LiteRT)

~3.65 GB

Vision models — image understanding (MNN)

InternVL2.5 1B

OpenGVLab (MNN)

4 GB RAM min.

SmolVLM2 2.2B

Hugging Face (MNN)

6 GB RAM min.

Image generation (Stable Diffusion · MNN + NPU)

Anything V5

Anime — CPU + NPU

~1.25 GB

AbsoluteReality

Photorealistic — CPU + NPU

~1.25 GB

ChilloutMix

Photorealistic — CPU + NPU

~1.25 GB

CuteYukiMix

Anime — CPU + NPU

~1.25 GB

Voice input & text-to-speech

Whisper base (optional)

whisper.cpp

~70 MB

Whisper enhanced (optional)

whisper.cpp

~140 MB

Supertonic 3 (optional)

ONNX · 31 languages · 10 voices

~380 MB

OfflineGPT Core is free

Private Cloud is an optional Google Play subscription.

Core

Local chat, image generation, voice, writing tools, OCR, and custom tools — included, without a feature paywall.

Private Cloud

Two hosted models in Switzerland. CHF 12.90 per month in Switzerland, unlimited personal use. No OfflineGPT account and no cloud chat archive.

Learn more about Private Cloud

Frequently asked questions

Do I need internet?

To download models and app updates, yes. Once downloaded, local chat, image generation, and voice input work fully offline. Optional: Gemma 4 web search, API mode, and the Python sandbox (API models only) need a connection when used.

Is OfflineGPT free?

Yes. OfflineGPT Core is free: local chat, image generation, voice input, and writing tools. The optional Private Cloud subscription adds two hosted models in Switzerland (CHF 12.90 per month in Switzerland; the Google Play price for your region applies).

What is the Python sandbox?

A sandboxed Python runtime for API-model chats only — when you use cloud providers with your own keys. Enable network access so scripts can call APIs or fetch remote data. It is not available for local on-device models (Gemma 4, GGUF downloads, etc.).

What is Gemma 4 E4B?

Google's multimodal on-device model — it understands text, images, and audio attachments in a single conversation. It runs via LiteRT-LM, not llama.cpp, and supports an optional web search tool you can toggle on or off.

Is there an iPhone or iPad version?

Yes. On the App Store the app is listed as Offline AI Studio — same offline-first philosophy, built for iOS with Core ML.

How does image generation work?

Stable Diffusion models (Anything V5, AbsoluteReality, ChilloutMix, CuteYukiMix) run locally via MNN. On supported Snapdragon devices, NPU variants run on the Hexagon DSP for faster generation.

How does voice input work?

Download a Whisper model in the app (base or enhanced), grant microphone access, tap the mic, and dictate. Audio is transcribed on-device — nothing leaves your phone.

What is text-to-speech?

Answers can be read aloud automatically or on tap. The system voice works with zero setup. Supertonic 3 is an optional ~380 MB download with 31 languages and 10 neural voices, and supports streaming TTS during generation.

How does privacy work?

By default, prompts and chat history stay on your device. Data only leaves your phone when you explicitly use optional features: model downloads, Gemma 4 web search, your own API keys, the Python sandbox with network, or OfflineGPT Private Cloud. Full privacy policy.

Get OfflineGPT

Free download on Google Play · Android 8.0+

Get it on Google PlayDownload on the App Store