Paperless-ngx
Description
Paperless-ngx is an open-source document management system that transforms scanned paper, PDFs, and email attachments into an indexed, full-text-searchable archive. Documents are OCR'd on ingestion, then classified with tags, correspondents, and document types using matching rules that can train themselves from your corrections. It is the actively maintained, community-supported successor to Paperless and Paperless-ng.
Features
- Automatic OCR: Recognizes text in scans and image PDFs in over 100 languages via Tesseract.
- Smart classification: Auto-assigns tags, correspondents, and document types with trainable matching.
- Full-text search: Fast search across content and metadata, with fuzzy matching and saved views.
- Consumption pipeline: Watch folders, IMAP email fetching, a REST API, and mobile uploads.
- Document lifecycle: Versioned edits, notes, custom fields, and configurable retention.
- Access control: Per-document and per-object permissions with users and groups.
Technology Stack
- Python / Django backend with a Celery task queue
- Angular single-page frontend
- PostgreSQL (or SQLite / MariaDB) plus Redis
- Tesseract OCR, Apache Tika, and Gotenberg for document processing
- Docker Compose for deployment
Requirements
- Docker and Docker Compose
- Minimum 2 GB RAM (4 GB recommended for multi-language OCR)
- Storage sized to your document archive plus originals
Categories
Topics
GitHub Metrics
Stars
44,936Forks
3,095Contributors
3,095Last Updated
9/8/2026Ad
Sponsor this page
Get seen without shouting for attention. Clean placement, real audience.
Learn more