Intro
Most AI email tools send your messages to OpenAI or Google. If you handle support, invoices, or customer data, that's a problem. This guide shows a fully private alternative.
A local AI email assistant uses Ollama to run a small LLM on your hardware and n8n to read, classify, and act on emails — without any data leaving your network. For a 16GB GPU, a 7B–8B model (Qwen 2.5 or Llama 3.1) classifies email reliably at 5–15 emails/minute.
What you need
- Ollama installed locally
- n8n (self-hosted or desktop)
- A Gmail app password (or IMAP credentials)
Step 1 — Pick a local LLM
For email classification, small fast models beat large slow ones.
ollama pull qwen2.5:7b
Step 2 — Connect n8n to Ollama
Use the HTTP Request node against http://host.docker.internal:11434/api/chat if n8n runs in Docker.
Step 3 — Build the classification flow
The video below walks through the full classification flow end to end — watch it here, then grab the ready-made files to follow along.
Complete workflow files for this guide — the n8n flow JSON you can import in one click, plus the prompts used in the video.
Q: Does this work without a GPU? A: Yes, but CPU inference is 5–10× slower. A quantized 7B model is usable on CPU for low volume.
Related posts
Escape the Average: Creative Prompt Techniques for Local LLMs
Local LLMs give generic answers by design. Fix it with Ollama sampling dials, ban lists, and constraint stacks baked into reusable Modelfiles.
Hire the Playbook: Claude Skills as Guided, Step-by-Step Workflows
Abandoned projects aren't waiting on motivation — they're waiting on structure. A skill file turns Claude from answer-dispenser into guide: plan decomposed, one step visible at a time, ELI5 on demand.
Stuck Between Two Options? Break Decision Paralysis Before the Deadline
A tied pro/con list means the list is out of answers. Break the binary, audit which beliefs you can verify before Friday, pre-live both futures — then hunt the one missing fact.



