• Artificial Intelligence

Gemma 4 by Google: Specs, Benchmarks, Model Sizes, and How to Run It Locally (2026 Guide)

Published On: 3 April 2026.By .
Open Models  ·  Edge AI  ·  12B Unified  ·  MTP Speedups  ·  Developer Guide

Gemma 4 by Google: Specs, Benchmarks, All 5 Model Sizes & How to Run Locally — Complete 2026 Guide

Quick answer: Gemma 4 is Google DeepMind's most capable open AI model family, released March 31, 2026 under the Apache 2.0 license. It now comes in 5 sizes: E2B (2B params, for phones), E4B (4B, for edge), the new 12B Unified (multimodal with audio, added June 2026), 26B MoE (3.8B active, for consumer GPUs), and 31B Dense (for workstations). The 31B scores 85.2% on MMLU Pro, 89.2% on AIME 2026, and ranks #3 on Arena AI. New MTP drafters deliver up to 3x faster inference with identical output quality. The family supports text and image input, audio on E2B/E4B/12B, up to 256K context, and 140+ languages, and can be installed locally with one command: ollama run gemma4.

What if a model small enough to fit on your smartphone could outperform AI systems 20 times its size? That is no longer hypothetical. With Gemma 4, Google is pushing frontier-level AI into devices, laptops, workstations, and servers in a form developers can actually run, fine-tune, and deploy commercially. And the family keeps getting better: June and July 2026 brought the 12B Unified model, MTP drafters with up to 3x faster inference, and MTP support landing in llama.cpp and Ollama.

~24 min read
Published:  ·  By
New: 12B Unified
New: MTP 3x Speed
Gemma 4
Open Models
Benchmarks
On-Device AI
Deployment
Agentic AI
Quick Takeaways

Gemma 4 in 60 Seconds

Best short answer: Gemma 4 is Google's open-weight AI model family for local, edge, and server deployment. It is available in E2B, E4B, 12B Unified, 26B A4B MoE, and 31B sizes, supports up to 256K context, runs up to 3x faster with the new MTP drafters, and is licensed under Apache 2.0.

01
Best model for most developers
Use the 26B A4B MoE model when you want the strongest balance of quality, speed, and consumer-GPU deployment.
02
Best model for mid-range hardware
Use the new 12B Unified for multimodal work (text, image, audio) on 12–16GB GPUs or Apple Silicon laptops. Use E2B for phones and IoT, E4B for higher-end mobile.
03
Best setup path
Start with Ollama 0.31+ (auto-tuned MTP on Mac) or Google AI Edge Gallery for testing, then move to llama.cpp with MTP drafters, LiteRT-LM, vLLM, Vertex AI, or Google Cloud for production.
Auriga IT banner introducing Google Gemma 4 open AI model family for edge and cloud deployment
Google Gemma 4 model family — open AI for edge to server deployment. Credit: Auriga IT
500M+
Total Gemma Downloads
3x
Faster Inference with MTP Drafters
5
Model Sizes (12B Unified added June 2026)
256K
Max Context Window (tokens)
140+
Languages Supported
Apache 2.0
Commercially Open License
01 - What's New

What's New for Gemma 4 in June–July 2026: 12B Unified, MTP, and 3x Faster Local AI

Key takeaway: Since our last update, Google shipped a fifth model size (Gemma 4 12B Unified), released MTP drafters that speed up inference up to 3x with zero quality loss, and the local tooling ecosystem — llama.cpp, Ollama 0.31, MLX — added native MTP support. Gemma 4 is now meaningfully faster and more accessible than it was at launch.
1
Gemma 4 12B Unified — New Fifth Model Size
Released June 2026, the 12B Unified uses a unified multimodal architecture: no separate vision or audio encoders. Image patches and audio spectrograms flow directly through lightweight projection layers. It handles text, image, and audio, supports 256K context, and runs in ~16GB VRAM full precision or ~8GB quantized — RTX 4070-class cards and Apple Silicon laptops.
2
MTP Drafters — Up to 3x Faster Inference
Google released Multi-Token Prediction (MTP) drafters for the whole family under Apache 2.0. A lightweight drafter proposes several tokens; the main model verifies them in one parallel pass. Same output quality, up to 3x the speed. The drafters share the target model's KV cache, so overhead stays minimal.
3
llama.cpp MTP Support (June 7)
Official MTP support landed in llama.cpp b9549. Users report ~140 tok/s for the 12B on a 12GB RTX 4070 Super and ~2x speedups on dual RTX 3090s. Dense models benefit most; the 26B MoE sees ~1.2–1.3x.
4
Ollama 0.31 — Auto-Tuned MTP on Mac (June 29)
Ollama 0.31 ships auto-tuned MTP for Gemma 4 on MLX, on by default. On a real coding-agent benchmark (Aider polyglot), generation went from 50.2 to 95.0 tok/s on an M5 Max — nearly 90% faster with identical output.
5
DiffusionGemma — 1,100+ Tokens/Second
An experimental diffusion-based Gemma 4 variant generates 15–20 tokens per forward pass through parallel denoising, exceeding 1,100 tok/s per user on H100 hardware. It trades some benchmark accuracy for extreme speed — a preview of where open-model inference is heading.
6
Wafer-Scale Speed and New Competition
Cerebras now serves Gemma 4 31B at 1,851 tok/s — the first multimodal model at wafer-scale speed. Meanwhile, Alibaba's Qwen 3.6 (late June) became the new text-only local coding favourite, sharpening the competitive picture we cover below.

Why this matters: the biggest complaint about local open models has always been speed. With MTP now built into the mainstream local stack — llama.cpp, Ollama, MLX, vLLM, Transformers — Gemma 4 closes much of the responsiveness gap with cloud APIs while keeping data fully on your hardware.

02 - Google I/O 2026

Google I/O 2026: What Changed for Gemma and Google AI

Key takeaway: Google I/O 2026 (May 19–20) was one of the most consequential AI developer events in years. The major announcements — Gemini 3.5 Flash, Gemini Spark, and WebMCP — all affect how developers build with and around Gemma 4. Google also said commonly used open-weight models such as Gemma 4 are being added to the Android Bench LLM leaderboard.

At its annual I/O developer conference (May 19–20, 2026) at the Shoreline Amphitheatre in Mountain View, Google centred almost entirely on AI, with Gemini 3.5 Flash as the headline model release and a series of agentic AI announcements that directly change the landscape Gemma 4 operates in.

Here is a breakdown of every major I/O 2026 announcement relevant to Gemma 4 users and open-model developers:

1
Gemini 3.5 Flash — Faster Frontier Model
Google's new flagship model outperforms Gemini 3.1 Pro on almost all benchmarks while running 4x faster. It now powers the Gemini app by default, AI Overviews, Google Search, and the new Gemini Spark agent.
2
Gemini Spark — 24/7 Personal AI Agent
Gemini Spark is Google's new proactive AI agent that runs in the background inside the Gemini app and Workspace tools. It can monitor your inbox, handle multi-step tasks, and take actions autonomously — a direct signal of where agentic AI is heading in 2026.
3
Gemma 4 on Android Bench
Google describes Android Bench as an LLM leaderboard for Android development tasks and said it is adding commonly used open-weight models such as Gemma 4. This makes Gemma 4 directly comparable to proprietary models in Android-specific coding and reasoning tasks.
4
WebMCP — Browser-Native Tool Standard
Google proposed WebMCP as an open web standard, letting browser-based AI agents call JavaScript functions and interact with HTML forms via a standardized interface. This extends Anthropic's MCP concept into the browser layer. The origin trial starts in Chrome 149.
5
Gemini Omni — Video Generation Model
Google unveiled Gemini Omni, a new AI model built for cinematic video generation and editing via text prompts, images, and video clips. It supports conversational editing of characters, backgrounds, and scenes via voice commands.
6
Managed Agents in the Gemini API
A single Gemini API call now provisions a fully managed agent with a remote sandbox, removing the infrastructure setup friction from agentic workflows. This lowers the barrier for teams building Gemma 4-based agents that need cloud-scale execution.

What does I/O 2026 mean for Gemma 4 specifically? Google reinforced Gemma 4's place in the broader ecosystem through Android Bench updates, while the WebMCP announcements create new surfaces where Gemma 4 can be integrated as a local model backend for browser-native agent workflows.

03 - Why It Matters

Why Gemma 4 Is a Big Deal for Open AI in 2026

Since the first Gemma models launched, developers around the world have downloaded them over 500 million times and created more than 100,000 custom variants. That level of adoption tells you something important: people wanted open models that were practical, fast, and deployable beyond the cloud.

Gemma 4 is Google DeepMind's answer to that demand. It brings frontier-level intelligence into model families that can run on everything from a Raspberry Pi to a data-centre GPU, while remaining open enough for developers and businesses to actually build with. Built from the same research and technology behind Gemini 3, it is the most capable model family you can run on your own hardware — and Google keeps shipping: the 12B Unified and MTP drafter releases in June 2026 show the family is being actively expanded, not just maintained.

The open model landscape has shifted rapidly since Gemma 4's launch. Alibaba's Qwen 3.5 family dropped weeks later with competitive scores and was followed by Qwen 3.6 in late June, Meta continued expanding Llama 4 Scout's ecosystem, and Mistral pushed its own mid-size models. Yet Gemma 4 remains the only family that spans phones to servers under a fully permissive Apache 2.0 license with no MAU restrictions — a combination none of its competitors match as of July 2026.

Meanwhile, the broader AI ecosystem has moved decisively toward agentic AI, where models don't just answer questions but autonomously call tools, make decisions, and execute multi-step workflows. Anthropic's Model Context Protocol (MCP) has emerged as a standard for connecting AI models to external tools and data sources. Google's own Agent-to-Agent (A2A) protocol is gaining traction for multi-agent coordination. And the "vibe coding" movement has gone from novelty to mainstream workflow. Gemma 4 sits at the intersection of all three trends.

Gemma model family growth: from launch to 500 million downloads and 100,000 custom variants
Gemma model family milestones: 500M+ downloads, 100K+ community variants. Source: Google DeepMind
04 - Overview

What Is Gemma 4? Overview and Licensing

Key takeaway: Gemma 4 is Google's newest family of open AI models, released March 31, 2026 under the Apache 2.0 license. It now includes 5 model sizes built from Google Gemini research and technology, spanning smartphones to servers, with no commercial use restrictions.

Gemma 4 is Google's newest family of open AI models, released on March 31, 2026, with the fifth size — the 12B Unified — added in June 2026. The models are built from Gemini research and technology, but unlike Google's proprietary offerings, Gemma 4 is released openly for the community to use, modify, and deploy.

Google has published Gemma 4 under the Apache 2.0 license, which means developers and companies can use it commercially without restrictive licensing headaches. No monthly active user limits, no acceptable use policies, no special permissions needed.

This licensing distinction matters more in 2026 than ever before. As companies build AI agents that run continuously, process customer data, and integrate with internal tools via protocols like MCP and the new WebMCP standard announced at I/O 2026, the licensing terms of the underlying model become a strategic decision. Apache 2.0 means no surprises at scale.

05 - Model Sizes

Gemma 4 Model Sizes, Parameters, and Hardware Requirements

Key takeaway: Gemma 4 now comes in 5 sizes. The 26B MoE model is the sweet spot for most developers — it activates only 3.8B of its 26B parameters per token, delivering 97% of the 31B's quality at roughly 8x less compute, fitting on a single RTX 3090/4090. The new 12B Unified is the pick for multimodal work on mid-range hardware.

Google designed Gemma 4 to span edge devices, laptops, consumer GPUs, and production servers. The family now includes five options, each with a distinct deployment sweet spot.

Gemma 4 model size comparison: E2B, E4B, 12B Unified, 26B MoE, and 31B Dense with parameters, use cases, and context windows
Gemma 4 model size comparison table — E2B, E4B, 12B Unified, 26B MoE, 31B Dense.
Gemma 4 model sizes and minimum hardware requirements at Q4 quantization (updated July 2026)
Model Active Params Best For Context Min RAM (Q4)
Gemma 4 E2B~2.3B effectiveSmartphones, IoT, Raspberry Pi128K tokens~1.5 GB
Gemma 4 E4B~4.5B effectiveMobile apps, edge devices, laptops128K tokens~5 GB
Gemma 4 12B Unified NEW~12B (unified multimodal)Multimodal (text/image/audio) on 12–16GB GPUs, Apple Silicon256K tokens~8 GB
Gemma 4 26B A4B MoE3.8B of 26B totalConsumer GPUs (RTX 3090/4090), Mac256K tokens~14–18 GB
Gemma 4 31B Dense30.7B (all active)Maximum quality, research, fine-tuning256K tokens~20 GB

The Gemma 4 26B model uses a Mixture of Experts (MoE) architecture with 128 small experts, activating only 8 per token plus one shared expert. Instead of activating the full model every time, it selectively turns on the most relevant expert pathways, delivering near-31B quality at dramatically lower compute cost.

The 12B Unified, added in June 2026, takes a different route: it removes the separate vision transformer and acoustic encoder entirely. Image patches pass through simple linear layers directly into the transformer's embedding space, and audio spectrograms go through lightweight projection layers. The result is lower memory use, faster inference, and better multimodal alignment — a genuine architectural experiment shipped as a production model.

The E2B and E4B use Per-Layer Embeddings (PLE), giving them the representational depth of a much larger model while keeping memory usage low enough for smartphones and Raspberry Pi boards.

Gemma 4 device compatibility: E2B runs on smartphones and Raspberry Pi; E4B on laptops; 12B Unified on mid-range GPUs; 26B MoE on RTX 3090/4090; 31B Dense on workstations and servers
Gemma 4 hardware compatibility — from Raspberry Pi and smartphones to data centre GPUs.
06 - Capabilities

Gemma 4 Key Capabilities: Reasoning, Vision, Code, and Agentic AI

Gemma 4 is not just another general-purpose text model. It combines advanced reasoning, structured outputs, multimodal inputs, and long-context support in ways that make it genuinely useful for modern product development and for the new agentic surfaces opened up at Google I/O 2026.

01
Advanced Reasoning ("Thinking" Mode)
Built for multi-step planning, logic, and instruction-following. Use structured prompts, tool schemas, and evaluation tests to keep reasoning workflows reliable in production.
02
Agentic Workflows & Function Calling
Native support for function calling, JSON structured output, and system instructions. Strong fit for AI agents, including compatibility with MCP and the new WebMCP standard announced at I/O 2026.
03
Code Generation & Vibe Coding
Codeforces ELO jumped from 110 (Gemma 3) to 2150 (Gemma 4), reaching expert competitive programmer level. Works as a local offline coding assistant — now 2-3x faster with MTP drafters.
04
Multimodal: Vision, Video, and Audio
All models handle text + image input. E2B, E4B, and the new 12B Unified add native audio input. Video supported via frame sequences. Variable aspect ratio and resolution for images.
05
140+ Languages Supported
Natively trained across a very broad language set, valuable for global applications and multilingual content generation.
06
256K Token Context Window
Process huge documents, long conversations, or entire codebases in one go. The 26B MoE handles long context especially efficiently thanks to its hybrid attention architecture.
Gemma 4 key stats: 256K context window, 140+ languages, native audio, Apache 2.0 license
Gemma 4 key capability callouts: 256K context, 140+ languages, native audio, Apache 2.0.
07 - Benchmarks

Gemma 4 Benchmarks 2026: MMLU Pro, AIME, Codeforces, Arena AI

Key takeaway: The Gemma 4 31B Dense model ranks #3 among all open models on Arena AI (ELO 1452). On AIME 2026, it scores 89.2% — up from 20.8% for Gemma 3. Codeforces ELO jumped from 110 to 2150, the largest single-generation leap for any open model on record. And with MTP drafters, all of this now runs up to 3x faster.
Gemma 4 31B Dense benchmark results. Source: Google DeepMind
Benchmark Gemma 4 31B Gemma 3 27B Category
MMLU Pro85.2%General Knowledge
AIME 202689.2%20.8%Math Competition
GPQA Diamond84.3%42.4%Graduate-Level Reasoning
LiveCodeBench v680.0%29.1%Coding
Codeforces ELO2150110Competitive Programming
MMMU Pro76.9%Vision Understanding
Arena AI ELO1452 (#3 open)Human Preference

The Gemma 4 26B MoE model ranks #6 on Arena AI with an ELO of 1441, while only activating roughly 3.8 billion parameters during inference — achieving 97% of the 31B's quality at approximately 8x less compute per inference step. The new 12B Unified scores 77.2% on MMLU Pro — remarkable for a model that fits in 8GB quantized.

Inference speed in the MTP era (July 2026): with MTP drafters enabled, the 12B reaches ~140 tok/s on a 12GB RTX 4070 Super via llama.cpp, Gemma 4 on Apple Silicon jumps from ~50 to ~95 tok/s with Ollama 0.31's auto-tuned MTP, and dense models see 2–3x speedups broadly. On managed infrastructure, Cerebras serves the 31B at 1,851 tok/s. The old "local models are too slow" objection is aging fast.

08 - Comparison

Gemma 4 vs. Qwen 3.5/3.6 vs. Llama 4 — Full Comparison (2026)

Key takeaway: Gemma 4 and Qwen trade blows at the ~30B scale, within 1–2% on most benchmarks. Gemma 4 dominates math (AIME 89.2%), human preference (Arena AI #3), and multimodal local use. Qwen 3.5 leads on coding (SWE-bench 72.4%), and Qwen 3.6 (June 2026) is the new text-only local coding favourite. Llama 4 Scout trails on reasoning despite being 109B total. Both Gemma 4 and Qwen use Apache 2.0; Llama 4 has a 700M MAU restriction.
Gemma 4 31B vs Qwen 3.5 27B vs Llama 4 Scout — key dimensions as of July 2026
Dimension Gemma 4 31B Qwen 3.5 27B Llama 4 Scout
MMLU Pro85.2%86.1%
GPQA Diamond84.3%85.5%74.3%
AIME 2026 (Math)89.2%~48.7%*
Codeforces ELO2150
Arena AI ELO1452 (#3)~1404
LicenseApache 2.0Apache 2.0Meta License (700M MAU cap)
Context Window256K tokens128K tokens10M tokens
Smallest ModelE2B (2.3B) for phones0.8B109B total (no edge model)
Audio SupportYes (E2B/E4B/12B native)Omni variant onlyNo
Inference AccelerationMTP drafters (up to 3x, Apache 2.0)Generic speculative decodingGeneric speculative decoding

When to pick Gemma 4: Best for math-heavy reasoning, edge/on-device deployment, multimodal local work (the 12B Unified has no direct competitor), competitive programming, agentic AI tool-use workflows, and when you need the widest hardware coverage (phones to servers) under a fully open license.

When to pick Qwen 3.5/3.6: Best for text-only production coding workflows (SWE-bench leader; Qwen 3.6 27B is the current community favourite for local coding on high-memory Apple Silicon), when you need the largest available model (397B), or for real-time speech output via the Omni variant.

When to pick Llama 4 Scout: When you need massive context windows (10M+ tokens) and can accept Meta's licensing restrictions.

* Qwen 3.5 AIME score is from AIME 2025; direct numerical comparison across benchmark versions is directional, not exact.

What About Qwen 3.6, Mistral, DeepSeek, Phi-4, and Claude?

The open model space in 2026 is crowded and moving monthly. Qwen 3.6 (late June 2026) has become the community sweet spot for text-only local coding — pair it with llama.cpp and OpenCode on 48GB Apple Silicon and it is excellent. But Gemma 4 still wins multimodal local use, and with Ollama 0.31's MTP support, the speed argument for Qwen on Mac has narrowed considerably. Alibaba's flagship Qwen 3.7-Max is proprietary, so it competes with Gemini, not Gemma.

Beyond Qwen, developers are also evaluating Mistral's mid-size offerings for European data sovereignty use cases, DeepSeek V3 for cost-efficient Chinese-language tasks, Microsoft's Phi-4 for ultra-lightweight edge scenarios, and comparing open models against proprietary options like Anthropic's Claude and OpenAI's GPT families.

However, none of these match Gemma 4's combination of benchmark scores, edge-to-server hardware coverage, multimodal support (text, image, video, audio), MTP-accelerated inference, and Apache 2.0 licensing in a single model family.

09 - Agentic AI

Gemma 4 for Agentic AI, MCP, and the New WebMCP Standard

Key takeaway: 2026 is the year AI moved from chatbots to agents. Gemma 4's native function calling, structured JSON output, and long context make it one of the strongest open foundations for building agentic AI systems. Google I/O 2026's WebMCP proposal extends this further into browser-based agent workflows — and MTP speedups make local agents feel dramatically more responsive.

The biggest shift in AI during 2026 is not a new model; it is a new paradigm. Agentic AI — where models don't just answer questions but autonomously plan tasks, call external tools, make decisions, and execute multi-step workflows — has moved from research concept to production reality.

Anthropic's Model Context Protocol (MCP) has quickly become the standard for connecting AI models to external data sources and tools. At Google I/O 2026, Google extended this concept with WebMCP, a proposed open web standard that lets browser-based AI agents call JavaScript functions and interact with HTML forms through a standardized interface. The origin trial starts in Chrome 149.

Google also shipped Managed Agents in the Gemini API, allowing a single API call to provision a fully managed agent with a remote sandbox. And Gemini Spark — the new 24/7 background agent announced at I/O 2026 — signals that agentic AI is moving from developer experiments to mainstream consumer products.

Where Gemma 4 fits in agentic AI: Its native function calling, JSON structured output, configurable thinking modes, and 256K context window make it well-suited as the "brain" of agentic systems. Because it runs locally under Apache 2.0, you can deploy Gemma 4 agents on your own infrastructure without per-call API costs or data leaving your servers. And since agents live or die on decode speed — every tool call and retry is another generation pass — the MTP speedups shipped in June directly translate into faster agent loops.

10 - Vibe Coding

Gemma 4 for Vibe Coding: Local AI-Assisted Development

Key takeaway: Vibe coding has gone mainstream in 2026. Gemma 4's Codeforces ELO of 2150 and local deployment make it a strong foundation for private, offline AI-assisted development — and MTP now makes local coding agents up to 90% faster on real benchmarks.

The term "vibe coding" describes a new style of software development: instead of writing code line by line, you describe the intent and an AI model generates the implementation. What started as a playful concept has become a genuine productivity shift in 2026.

Tools like Cursor, Windsurf, Claude Code, GitHub Copilot, Bolt, Lovable, and Replit have made vibe coding accessible to millions of developers. But most of these tools rely on cloud-based proprietary models, which means your code, prompts, and context are sent to external servers.

Gemma 4 offers an alternative. With a Codeforces ELO of 2150 (expert competitive programmer level), 80% on LiveCodeBench v6, and the ability to run entirely on a single consumer GPU, it is one of the most capable coding models you can run locally. That means vibe coding with full privacy.

Speed used to be the catch — and it is disappearing. Code is exactly where MTP drafters shine: closing brackets, repeated identifiers, and boilerplate are patterns a draft model predicts well. Ollama 0.31 measured Gemma 4 going from ~50 to ~95 tok/s on the Aider polyglot coding-agent benchmark on Apple Silicon, and llama.cpp users report ~140 tok/s for the 12B on a 12GB RTX 4070 Super. Local coding agents built on Claude Code, OpenCode, or Continue.dev pointed at local Gemma 4 now feel genuinely responsive.

For teams building internal tools, prototyping features, or working with sensitive code, a local Gemma 4 instance combined with Continue.dev or LM Studio gives you the vibe coding experience without the data exposure risk.

11 - Getting Started

How to Download and Run Gemma 4 Locally (Ollama, llama.cpp, LM Studio)

Key takeaway: The fastest way to run Gemma 4 locally is with Ollama — use version 0.31+ for automatic MTP acceleration on Apple Silicon. For more control (and MTP on NVIDIA GPUs), use llama.cpp with the assistant drafter models. For a visual interface, use LM Studio. For production serving, use vLLM.

Step 1: Install with Ollama (Easiest Method)

# Install Ollama (macOS / Linux) — 0.31+ enables auto MTP on Apple Silicon
curl -fsSL https://ollama.com/install.sh | sh

# Run the default 26B MoE model (recommended for most developers)
ollama run gemma4

# Or choose a specific Gemma 4 size:
ollama run gemma4:e2b   # Edge — phones, Raspberry Pi (~1.5GB)
ollama run gemma4:e4b   # Edge — laptops, mobile apps (~5GB)
ollama run gemma4:12b   # NEW — multimodal unified, mid-range GPUs (~8GB Q4)
ollama run gemma4:26b   # MoE — best speed/quality balance (~14-18GB)
ollama run gemma4:31b   # Dense — maximum quality (~20GB)

Step 2: Visual Interface with LM Studio (GUI Option)

If you prefer a visual interface, LM Studio offers one-click download and chat for all Gemma 4 variants. Download from lmstudio.ai, search for "Gemma 4" in the model browser, select your preferred size and quantization, and start chatting. LM Studio also runs a local OpenAI-compatible API server.

Step 3 (Advanced): Maximum Control with llama.cpp — Now with MTP

# Build llama.cpp with GPU support (b9549+ includes MTP)
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j

# Run the Gemma 4 26B MoE model
./llama.cpp/llama-cli \
  -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL \
  --temp 1.0 --top-p 0.95 --top-k 64

# For up to 2-3x speed on dense models, add the MTP assistant drafter:
# download both the base GGUF and its -assistant GGUF, then serve with
#   --model-draft  --spec-type draft-mtp --spec-draft-n-max 4
# Tip: use the QAT checkpoints — they preserve MTP speedup when quantized.
# Note: MTP shines on dense models (12B/31B); the 26B MoE gains only ~1.2-1.3x.

On Apple Silicon Macs, use -DGGML_CUDA=OFF. Metal support is enabled by default. Known issue: quantizing the KV cache to q8_0 can break MTP acceptance rates — keep the KV cache at f16 when using drafters.

Step 4: Python Developers — Hugging Face Transformers

pip install transformers torch
# Load the model with model ID: google/gemma-4-31B-it
# MTP drafters: google/gemma-4-31B-it-assistant (same Apache 2.0 license)

Try Without Installing: Google AI Studio

Explore larger Gemma 4 models in Google AI Studio. For on-device variants, check Google's AI Edge resources.

Supported tools and frameworks: Hugging Face Transformers, vLLM, llama.cpp (MTP since b9549), MLX (Apple Silicon), LM Studio, Ollama (auto MTP in 0.31+), NVIDIA NIM & NeMo, Unsloth, SGLang, Keras, Docker, Baseten, and more.

12 - Edge Deployment

Can Gemma 4 Run on a Smartphone? Edge Deployment Guide

Yes. The Gemma 4 E2B and E4B models were specifically built for on-device use. They can run offline on smartphones, Raspberry Pi boards, and embedded hardware like NVIDIA Jetson devices. In 4-bit mode, E2B fits in approximately 1.5 GB RAM and E4B in roughly 5 GB — feasible for modern mobile and edge scenarios. The new 12B Unified extends this to laptops and mid-range GPUs at ~8 GB quantized, with full text + image + audio input.

E2B, E4B, and the 12B Unified all support native audio input, a capability that neither Llama 4 nor Qwen offer at these sizes. In the context of 2026's agentic AI trend, on-device models become even more valuable: an AI agent running locally on a warehouse scanner that reads barcodes, checks inventory via MCP, and triggers reorders without ever sending data to the cloud.

Gemma 4 RAM requirements by model and quantization level (updated July 2026)
Model 4-bit (Q4) RAM 8-bit RAM Full Precision
Gemma 4 E2B~1.5 GB~3 GB~10 GB
Gemma 4 E4B~5 GB~8 GB~15 GB
Gemma 4 12B Unified NEW~8 GB~13 GB~16 GB (BF16 VRAM)
Gemma 4 26B MoE~14–18 GB~28 GB~52 GB
Gemma 4 31B Dense~20 GB~34 GB~62 GB
13 - Fine-Tuning

Fine-Tuning Gemma 4 with QLoRA: Hardware and Tools Guide

The Apache 2.0 license allows unrestricted fine-tuning on proprietary data. Using QLoRA (Quantized LoRA) via tools like Unsloth, you can fine-tune the Gemma 4 31B model with as little as 16 GB VRAM — a single RTX 4090 or equivalent.

Fine-tuning Gemma 4 is supported on Google Colab, Vertex AI, Hugging Face TRL, Unsloth, and consumer GPUs. Full fine-tuning (all parameters) requires approximately 80 GB VRAM for the 31B model. For most custom tasks, QLoRA is sufficient and far more accessible.

A growing number of teams are fine-tuning Gemma 4 specifically for agentic tool use: training the model to reliably call the right MCP tools, parse structured responses, and handle multi-step workflows with minimal hallucination.

14 - Practical Impact

Why Businesses and Developers Should Choose Gemma 4

01
Privacy First — Data Stays On-Device
Running Gemma 4 locally means sensitive data stays on-device instead of being shipped to external APIs. Critical for healthcare, finance, and legal applications under HIPAA, GDPR, and data residency requirements.
02
Lower Cost at Scale
No per-token API billing. Run on hardware you already control. At scale, self-hosting Gemma 4 can cut inference costs by 60–80% vs. proprietary APIs — and MTP speedups improve throughput on the same hardware for free.
03
Full Customization
Fine-tune, adapt, and package for domain-specific tasks. No vendor lock-in or dependency on third-party roadmaps.
04
Apache 2.0 — No Usage Caps
No MAU limits, no commercial restrictions, no Gemma-specific monthly active user cap. Build products freely and redistribute without special permissions — unlike Llama 4's 700M MAU cap.
05
The Agentic AI Shift
Gemma 4's native function calling, JSON structured output, and long context window make it one of the strongest open foundations for building agentic systems compatible with MCP and the new WebMCP standard.
06
Vibe Coding Ready — Locally
With a Codeforces ELO of 2150 and MTP-accelerated generation, Gemma 4 powers local AI-assisted development workflows. Describe what you want; the model writes the code — without sending your proprietary codebase to external servers.

For teams looking to deploy AI at scale, Gemma 4's compatibility with major inference frameworks, MCP tool ecosystems, and cloud platforms makes the path from prototype to production cleaner than ever. Cygnus Alpha by Auriga IT can help integrate open models into real business workflows.

15 - Landscape

The Open AI Model Landscape in Mid-2026: Where Gemma 4 Fits

The pace of AI releases in 2026 has been extraordinary, and the weeks since Google I/O accelerated it further. To understand where Gemma 4 fits, it helps to zoom out and see the full picture.

Google's Dual-Track Strategy: Google's strategy is now clearly dual-track: Gemini 3.5 Flash for cloud-scale proprietary deployments, and the Gemma 4 family for open, on-device, and self-hosted use cases. The two are complementary, not competitive. Gemma 4 was added to Android Bench at I/O 2026, and the June releases of the 12B Unified and MTP drafters show Google is investing in the open track, not just the cloud one.

The Inference Speed Race: The frontier of open-model competition has shifted from raw benchmark scores to tokens-per-second. MTP drafters, DiffusionGemma's parallel denoising, Cerebras wafer-scale serving at 1,851 tok/s, and Ollama's auto-tuned speculative decoding all landed within weeks of each other. Expect speed, not just intelligence, to headline the next wave of releases.

Anthropic's MCP Ecosystem: Anthropic's Model Context Protocol has become the de facto standard for connecting AI to tools. Any model with function calling support, including Gemma 4, can participate in MCP ecosystems. Google extended this with WebMCP for browser-native agents at I/O 2026.

The Rise of AI Coding Tools: Cursor, Windsurf, GitHub Copilot Workspace, Claude Code, and other tools have made AI-assisted coding the default workflow for a growing number of developers. Open models like Gemma 4 are increasingly being plugged into these tools as local backends — and MTP makes those local backends feel responsive.

Enterprise AI Adoption: Companies are no longer asking "should we use AI?" but "which model, where, and under what terms?" The Apache 2.0 license, local deployment options, and agentic capabilities of Gemma 4 directly address the procurement, privacy, and compliance concerns that slowed enterprise adoption in previous years.

16 - Sources & Verification

Sources Used to Verify This Gemma 4 Guide

17 - FAQ

Frequently Asked Questions About Gemma 4 (Updated July 2026)

What is Gemma 4?

Gemma 4 is Google DeepMind's most capable family of open AI models, released March 31, 2026. Built from Google Gemini research, it now includes 5 model sizes (E2B, E4B, 12B Unified, 26B MoE, 31B Dense) under the Apache 2.0 license for unrestricted commercial use.

What is Gemma 4 12B Unified?

The fifth Gemma 4 model size, released June 2026. It uses a unified multimodal architecture with no separate vision or audio encoders — image patches and audio spectrograms pass directly through lightweight projection layers. It supports text, image, and audio input, a 256K context window, and runs in ~16GB VRAM full precision or ~8GB quantized. It scores 77.2% on MMLU Pro.

What is Gemma 4 MTP (Multi-Token Prediction)?

MTP drafters are lightweight companion models that use speculative decoding to deliver up to 3x faster Gemma 4 inference with identical output quality. A small drafter proposes several tokens; the main model verifies them in one parallel pass. Supported in llama.cpp (since June 7, 2026), Ollama 0.31+ (auto-tuned on Apple Silicon), Hugging Face Transformers, vLLM, and MLX. Released under Apache 2.0.

What did Google announce about Gemma at I/O 2026?

At Google I/O 2026 (May 19–20), Google added Gemma 4 to Android Bench, its LLM leaderboard for Android development tasks. The broader I/O announcements — including Gemini 3.5 Flash, WebMCP, and Managed Agents in the Gemini API — all affect how developers build with Gemma 4 as a local model backend.

Is Gemma 4 free to use commercially?

Yes. The Apache 2.0 license allows unlimited commercial use, modification, fine-tuning, and redistribution with no royalty payments, no MAU limits, and no restrictive use policies. This is more permissive than Meta's Llama 4 license.

What are the Gemma 4 model sizes?

Five sizes: E2B (~2.3B effective, for phones), E4B (~4.5B effective, for edge/laptops), 12B Unified (multimodal with audio, added June 2026), 26B A4B MoE (3.8B active of 26B total, for consumer GPUs), and 31B Dense (all parameters active, for maximum quality).

How do I download and run Gemma 4 locally?

The fastest method: install Ollama 0.31+ from ollama.com, then run ollama run gemma4 (or gemma4:12b for the new Unified model). Model weights are also on Hugging Face, Kaggle, and NVIDIA NIM.

What hardware do I need to run Gemma 4?

E2B: ~1.5 GB RAM (smartphones, Raspberry Pi). E4B: ~5 GB (laptops). 12B Unified: ~8 GB quantized or ~16 GB full precision (RTX 4070-class GPUs, Apple Silicon). 26B MoE at Q4: ~14–18 GB (fits on RTX 3090/4090 or Mac with 24GB unified memory). 31B Dense at Q4: ~20 GB. All models run on CPU too, though slower.

How does Gemma 4 compare to Qwen 3.5 and Qwen 3.6?

Within 1–2% on most reasoning benchmarks. Qwen leads on MMLU Pro (86.1% vs 85.2%) and SWE-bench coding, and Qwen 3.6 (June 2026) is the current community favourite for text-only local coding. Gemma 4 dominates on math (AIME 89.2%), competitive programming (Codeforces 2150), human preference (Arena AI #3, ELO 1452), and multimodal local use — Qwen has no equivalent to the audio-capable edge and 12B Unified models. Both use Apache 2.0.

How does Gemma 4 compare to Llama 4?

Gemma 4 31B outperforms Llama 4 Scout (109B total) on reasoning benchmarks like GPQA Diamond (84.3% vs 74.3%). Gemma 4 uses Apache 2.0 while Llama 4 has a 700M MAU restriction. Gemma 4 also covers edge deployment; Llama 4 has no small models for mobile or IoT.

What are the Gemma 4 31B benchmark scores?

MMLU Pro: 85.2%. AIME 2026: 89.2%. GPQA Diamond: 84.3%. LiveCodeBench v6: 80.0%. Codeforces ELO: 2150. MMMU Pro (vision): 76.9%. Arena AI: #3 with ELO 1452.

Can Gemma 4 run on a smartphone?

Yes. E2B and E4B are designed for on-device mobile deployment. E2B fits in ~1.5 GB RAM, runs on modern Android phones via Google AICore, operates completely offline, and supports native audio input.

What is the Gemma 4 context window?

E2B and E4B: 128K tokens. 12B Unified, 26B MoE, and 31B Dense: 256K tokens — sufficient for processing entire codebases, long documents, and extended conversations in a single inference pass.

Does Gemma 4 support images, video, and audio?

All models support text + image input with variable resolution. E2B, E4B, and the 12B Unified add native audio input. Video use cases can be handled through frame-sequence workflows.

Can Gemma 4 be fine-tuned?

Yes. Apache 2.0 allows unrestricted fine-tuning. Using QLoRA via Unsloth, the 31B can be fine-tuned with 16 GB VRAM. Full fine-tuning needs ~80 GB. Supported on Google Colab, Vertex AI, and consumer GPUs.

Can Gemma 4 be used for agentic AI workflows?

Yes. Gemma 4 has native support for function calling, JSON structured output, system instructions, and configurable thinking modes. Compatible with Anthropic's MCP and the new WebMCP standard announced at Google I/O 2026.

What is MCP and does Gemma 4 support it?

MCP (Model Context Protocol) is an open standard by Anthropic that lets AI models interact with external tools and data sources. Any model with function calling support — including Gemma 4 — works with MCP. Google extended this with WebMCP for browser-native agents at I/O 2026.

Can Gemma 4 be used for vibe coding?

Yes. With a Codeforces ELO of 2150 and 80% on LiveCodeBench v6, Gemma 4 powers AI-assisted coding workflows entirely offline — and MTP makes local coding agents up to 90% faster on real benchmarks (Ollama 0.31, Aider polyglot, Apple Silicon). Use with Ollama, LM Studio, or Continue.dev for a private local coding assistant.

Which Gemma 4 model should I use?

For most developers: the 26B MoE delivers 97% of 31B quality at ~8x less compute. For phones: E2B. For laptops: E4B. For multimodal work on a 12–16GB GPU or Apple Silicon laptop: the 12B Unified. For maximum quality with 20GB+ VRAM: 31B Dense.

What is the difference between Gemma 4 and Gemini?

Gemini is Google's proprietary cloud model (API-accessible). Gemma 4 is the open-weight version from the same Google Gemini research, designed to run locally on your own hardware with full data privacy. Gemini 3.5 Flash (announced at I/O 2026) is the latest Gemini; Gemma 4 is the open alternative.

What languages does Gemma 4 support?

Over 140 languages natively — one of the most multilingual open-weight model families available, making it suitable for global product development.

What is the best open source AI model in 2026?

As of July 2026: Gemma 4 (best all-around family, now 5 sizes from edge to server, AIME 89.2%, up to 3x faster with MTP), Qwen 3.6 (best for text-only production coding), and Llama 4 Scout (best ultra-long context at 10M tokens, Meta license with 700M MAU cap). The best depends on your use case, hardware, and licensing requirements.

18 - Bottom Line

The Bottom Line on Gemma 4 in Mid-2026

Three months after launch, Gemma 4 is not slowing down — it is compounding. Google I/O 2026 validated its position with Android Bench, and the June–July releases (12B Unified, MTP drafters, DiffusionGemma) closed the two biggest gaps open models had: mid-range multimodal hardware coverage and inference speed. While Gemini 3.5 Flash takes the cloud spotlight, Gemma 4 remains Google's answer for developers who need local, open, and private AI.

The bigger takeaway is not just that Gemma 4 is good. It is that open AI is increasingly becoming practical, competitive, and deployable in real-world products. With Apache 2.0 licensing, frontier-level benchmarks, edge deployment, up to 3x faster inference via MTP, agentic AI capabilities, MCP and WebMCP tool compatibility, and broad ecosystem support, Gemma 4 is the strongest open model family for developers who want to build without restrictions.

At Auriga IT, we help businesses turn AI breakthroughs like Gemma 4 into working products and scalable systems. From building intelligent applications to deploying them on strong cloud infrastructure, we work with the latest tools so teams can move faster with less uncertainty.

Build Smarter with Gemma 4 and Open AI

Whether you are exploring open models like Gemma 4, building AI agents with MCP, deploying private on-device intelligence, or integrating agentic workflows into your business, our team can help you turn the latest model advances into real outcomes.

Talk to Our AI Experts →
Gemma 4 Guide
Auriga IT

Gemma 4 by Google: Specs, Benchmarks, All 5 Model Sizes, MTP Speedups & How to Run It Locally  ·  © 2026 Auriga IT  · 

Related content

Stay Close to What We’re Building

Get insights on product engineering, AI, and real-world technology decisions shaping modern businesses.

suman yubraj
suman yubraj
Suman Yubraj is a Technical Writer at Auriga IT with a background in computer science and content writing. He translates complex technical topics into clear, accessible content for developers and business audiences alike.
Go to Top