Paperless-ngx

Paperless-ngx

Description

Paperless-ngx is an open-source document management system that transforms scanned paper, PDFs, and email attachments into an indexed, full-text-searchable archive. Documents are OCR'd on ingestion, then classified with tags, correspondents, and document types using matching rules that can train themselves from your corrections. It is the actively maintained, community-supported successor to Paperless and Paperless-ng.

Features

  • Automatic OCR: Recognizes text in scans and image PDFs in over 100 languages via Tesseract.
  • Smart classification: Auto-assigns tags, correspondents, and document types with trainable matching.
  • Full-text search: Fast search across content and metadata, with fuzzy matching and saved views.
  • Consumption pipeline: Watch folders, IMAP email fetching, a REST API, and mobile uploads.
  • Document lifecycle: Versioned edits, notes, custom fields, and configurable retention.
  • Access control: Per-document and per-object permissions with users and groups.

Technology Stack

  • Python / Django backend with a Celery task queue
  • Angular single-page frontend
  • PostgreSQL (or SQLite / MariaDB) plus Redis
  • Tesseract OCR, Apache Tika, and Gotenberg for document processing
  • Docker Compose for deployment

Requirements

  • Docker and Docker Compose
  • Minimum 2 GB RAM (4 GB recommended for multi-language OCR)
  • Storage sized to your document archive plus originals

Categories

Topics

GitHub Metrics

Stars
44,936
Forks
3,095
Contributors
3,095
Last Updated
9/8/2026
Ad

Sponsor this page

Get seen without shouting for attention. Clean placement, real audience.

Learn more
MohsenI built this — follow on X