RSSgate.
A self-hosted RSS feed manager with an AI-transcribed reader. One local Python web app, one YAML file, any OpenAI-compatible or hosted LLM.
What it is
RSSgate runs as a local Python web app (default http://0.0.0.0:8088), fully
configured from a single YAML file. It is desktop and mobile responsive and follows your OS
dark-mode setting automatically. There are two sides:
- Viewer (
/) — an endless reverse-chronological scroller that resumes where you left off, with a feed-filter sidebar, per-feed unread counts and a New/Since date toggle. Every article is replaced by an LLM transcription of the linked page: important text kept, advertising and boilerplate stripped. At the end of the stream you get “You've seen it all!” plus a date picker to jump back through time. - Admin (
/admin, gear icon top-right) — add feeds (RSS/Atom XML and bare web pages), manage categories, pick providers and models, tune polling and summarization, and watch your token spend.
Feed types
| Type | Examples | Behavior |
|---|---|---|
feed | Gizmodo, Ars Technica, NPR, your blog | Parsed with feedparser. Conditional GET (ETag/Last-Modified) plus guid
dedupe. Feed-declared <category> tags are captured and shown as chips. |
page | https://lite.cnn.com/ |
No structured feed: article links are extracted from the page with an LLM and treated like feed items. Skipped entirely when the page bytes haven't changed. |
auto | any URL | Probed once; result stored. |
Paste a homepage instead of a feed and RSSgate probes it first, verifies every candidate feed by fetching and parsing it as XML, and only then offers it. A bare page is never added silently.
Token efficiency is the design
- An article is summarized exactly once; guid dedupe means existing articles are never re-processed.
- Identical article text hits a
body_hashcache and is reused — zero tokens. - Bare pages are fetched with conditional GET and content-hash comparison, so the discovery call is skipped when nothing changed.
- Hero and gallery images are cached locally and served from your own instance — declared URLs are free, and extraction costs zero extra LLM tokens. Files are hash-named and sniffed for real image bytes, so no SSRF and no path traversal.
- Token usage per call is logged; the admin panel shows today / month / all-time, digest queue depth, seconds per article, cache hits and failed digests. Stats auto-refresh every 15 s.
- Digesting runs in parallel (
summarizer.concurrency, default 2), and items stuck inprocessingafter a crash are requeued automatically after 15 minutes.
LLM providers
llm.provider | Base URL | Notes |
|---|---|---|
local | any OpenAI-compatible server | vLLM, SGLang, llama.cpp, Ollama, LM Studio. A single-model server is auto-selected. |
openai | api.openai.com/v1 | needs api_key |
openrouter | openrouter.ai/api/v1 | needs api_key |
anthropic | Messages API | needs api_key |
llm.api_key accepts env:VARNAME so secrets stay out of the file, and
the key is never echoed back to the browser. Digesting is extractive work, so on local reasoning
servers RSSgate disables thinking tokens by default — on the author's hardware that took a
Gizmodo digest from 51.8 s / 3,127 output tokens to 7.4 s / 405 at equal quality.
Per-purpose model overrides let a cheap instruct model do the digesting while a reasoning model
handles bare-page discovery.
Also in the box
- Categories — post tags from the feed XML, feed-level categories, and your own labels per feed. Filters combine: chips within a box OR, different boxes AND.
- Raw mode (per feed, zero tokens) — still fetches, extracts, strips ads and dedupes; you just skip the digest for feeds that don't need one.
- Sponsored content filter (per feed, off by default) — a free title/URL pre-filter runs before the page is fetched, so sponsored posts never enter your reader and never reach the LLM.
- Resume and read state — your deepest scroll position is stored server-side,
each feed keeps its own read cursor, and view prefs live in
localStorage. - Maintenance — article retention, an image-cache cap, and orphan pruning; runs at startup and every 6 hours.
Install and run
python -m venv .venv && .venv/bin/pip install -r requirements.txt
cp config.example.yaml config.yaml # optional; run.py generates defaults
.venv/bin/python run.py # serves on 0.0.0.0:8088
Tests are offline by design:
.venv/bin/python -m pytest # 84 tests, no network required
Before you put it on the internet
RSSgate has no authentication and binds 0.0.0.0 by
default: it is meant for trusted networks. Set server.host: 127.0.0.1 or put it behind
a reverse proxy with auth if others can reach the port — the admin panel, and the LLM spend behind
it, is open to anyone who can.