December 20, 2025 · Damien B., Founder of Redact PDF AI

Best PDF Redaction APIs (2026): 9 Options Compared

The best PDF redaction API depends on one question: do you want to send a PDF and get a redacted PDF back, or do you want building blocks and assemble the pipeline yourself? For the first case, Redact PDF AI is the best overall choice in 2026: PDF or image in, flattened and irreversibly redacted PDF out, with AI detection of names, addresses, IBANs and other PII, OCR for scans, async jobs, and EU/Swiss hosting. Nutrient is the best pick if you already embed its viewer SDK, Private AI if you need an on-premise container with the widest entity and language coverage, Microsoft Presidio if you want open source, and Azure AI Language or AWS Comprehend if you are happy to build OCR, drawing and flattening around a text-only PII detector.

This guide compares nine options on what actually decides an integration: input and output formats, how sensitive data is found, OCR, hosting, and how much pipeline you have to operate yourself. Vendor details are from their public documentation and pricing pages as of September 2026; check them before you commit.

An API turns redaction into a step your pipeline runs over every incoming document.

What to evaluate in a redaction API

Before comparing vendors, get clear on the criteria that determine whether an API survives contact with production:

  • What goes in, what comes out. A "redaction API" that takes text and returns text with [PERSON] tokens leaves you to do OCR, map offsets back to page coordinates, draw the boxes and flatten the file. A PDF-in, PDF-out API does all of that.
  • How sensitive data is found. Regex and search strings catch what you already know (a phone-number pattern, a client's name). AI entity detection catches the names and addresses you did not know were there. Most real leaks are the second kind.
  • OCR. Scanned contracts, faxes and phone photos have no text layer. If the API cannot read them, the pipeline needs a second vendor.
  • Irreversible output. The output must be flattened or rasterized so nothing is recoverable by copy-paste, search, or glyph-position analysis. A box drawn over text is not a redaction.
  • Async job handling. OCR and detection take seconds to minutes on long documents. Look for a job id, status polling or webhooks, and idempotent retries.
  • Granular PII control. Choose categories per job, and exclude terms that must never be redacted (your own company name).
  • Retention and residency. Where the file is processed, how long it is kept, whether it trains anyone's models, and whether you can delete it on demand.
  • Errors, limits and a real spec. Explicit quota (402) and rate-limit (429) responses, and an OpenAPI definition you can generate clients from.

The APIs compared

APIInput → outputDetectionOCRHostingBest for
Redact PDF AIPDF, JPG, PNG → flattened PDFAI entities (8 categories) + always/never termsBuilt in, 100+ languagesEU & Swiss Azure, no training, auto-deletePDF-in, redacted-PDF-out with nothing to build
Nutrient AI Redaction APIPDF, Office, images → PDFAI entities + search/regexBuilt inCloud API or self-hosted Document EngineTeams already on the Nutrient SDK / viewer
Private AIText, PDF, DOCX, images → redacted file or entitiesAI, 50+ entity types, 50+ languagesBuilt inContainer in your VPC, or their cloudOn-premise, multi-format, multi-language
Azure AI Language (PII)Text → entities + redacted textAI entitiesSeparate (Document Intelligence)AzureBuilding your own pipeline on Azure
AWS Comprehend (PII)Text → entities / redacted textAI entitiesSeparate (Textract)AWSBuilding your own pipeline on AWS
Microsoft PresidioText and images → anonymized text / redacted imageRule + model recognizers, extensibleSeparate (bring your own)Self-hosted, open sourceOpen source, full control
Apryse SDKPDF → PDF, in-processSearch / regex, your own detectorAdd-onYour servers or client appsEmbedding redaction in your own app
pdfRest Redact PDF APIPDF → PDF (mark, then apply)Text strings and regex you supplyNoCloud or self-hostedRules-based redaction of known strings
ConvertAPI (PDF redact)PDF → PDFAI-basedYesCloudAdding redaction to an existing ConvertAPI workflow

1. Redact PDF AI — best overall for PDF-in, redacted-PDF-out

The Redact PDF AI API runs the same pipeline as the web app, exposed as a REST interface: Azure OCR, AI PII detection, and irreversible rasterized redaction. You upload one or many files, get a job id back immediately, poll until each document reaches redacted, and download a flattened PDF.

  • Detection. Choose per job from Person, Email, PhoneNumber, Address, Organization, Date, IBAN and CreditCard, or use your saved defaults. Add terms that must always be redacted and terms that never should be.
  • Scans and images. JPG and PNG are accepted directly; scanned PDFs are read with OCR in 100+ languages.
  • Output. Every page is rasterized and the text layer and metadata dropped, so the output cannot be reversed by copy-paste, search, or the glyph-position attacks that break text-layer redactions.
  • Retention modes. ephemeral (default) deletes originals after processing; studio keeps originals and masks so a human can review in the editor and re-export.
  • Reliability. X-Idempotency-Key makes retries safe; 402/429 are explicit; an OpenAPI spec and docs cover the full schema. An official MCP server exposes the same API to AI agents.
  • Hosting and compliance. Processed on Microsoft Azure in the EU and Switzerland, AES-256 at rest, TLS in transit, on SOC 2 Type II / ISO 27001 infrastructure, HIPAA-eligible; documents are never used to train models.
  • Price. Free credits on signup; Starter is US$50/month for 1,000 pages, Business US$250/month for 6,000 pages; pay-as-you-go credit packs for irregular volume.

Limits to know: it is a document redactor, not a general PDF toolkit (no merge/convert), and it is a hosted API rather than an in-process SDK. If files may not leave your network, look at Private AI or Presidio.

Best for: product and platform teams that need user-uploaded PDFs and images redacted reliably, with EU/Swiss residency and a no-training guarantee, without building OCR, detection, drawing and flattening themselves.

2. Nutrient AI Redaction API — best if you already use the Nutrient SDK

Nutrient (formerly PSPDFKit) sells document SDKs for web, mobile and server, and its AI redaction API sits on that stack: send a document, have entities detected, and receive a permanently redacted PDF. It supports native and scanned documents and can run as a cloud API or on your own infrastructure through Document Engine. The natural fit is a team that already renders and annotates PDFs with Nutrient's viewer and wants redaction from the same vendor. Pricing is credit-based per API call for the cloud service and licence-based for self-hosting, so budget for both if you need the on-premise option.

3. Private AI — best for on-premise and the widest coverage

Private AI is a PII detection and redaction service that runs as a container inside your VPC (or as their hosted API). It handles text, PDFs, Office files and images with built-in OCR, recognises 50+ entity types across 50+ languages, and returns either the redacted file or the entities with offsets. It is the strongest choice when data cannot leave your environment or when you redact many formats and languages. It is enterprise-priced and you operate the container, which is more work than calling a hosted endpoint.

4. Azure AI Language — PII detection as a building block

Azure's PII detection returns entities and a redactedText string for plain text, priced per 1,000 text records. It does not take a PDF and hand back a redacted PDF: you OCR with Azure Document Intelligence, map the detected spans back to coordinates, draw and flatten yourself. (A native-document mode has been in preview; check its current status before designing around it.) A sensible choice if you are all-in on Azure and want to own the pipeline. Redact PDF AI uses Azure AI services underneath and adds exactly that missing pipeline.

5. AWS Comprehend — the same, on AWS

Amazon Comprehend's DetectPiiEntities returns PII spans for text (English and a limited set of other languages), and its async jobs can write redacted text to S3. For PDFs you pair it with Amazon Textract for OCR and build the coordinate mapping, drawing and flattening. Billing is per unit of characters. Same trade-off as Azure: cheap building blocks, significant assembly.

6. Microsoft Presidio — best open source

Presidio is Microsoft's open-source PII framework: an analyzer with rule- and model-based recognizers, an anonymizer for text, and an image redactor that can black out detected text in images (including DICOM). You host it, tune the recognizers, and bring your own OCR for PDFs. Free, fully controllable, and a good fit for teams with ML and ops capacity who need to audit every rule; not a turnkey PDF service.

7. Apryse SDK — redaction inside your own application

Apryse (formerly PDFTron) provides an in-process SDK with a redaction module: search text or supply regions, then apply true redaction that removes the content. Detection is search and regex based unless you plug in your own detector. It is the option for embedding redaction into a desktop, mobile or server application you ship, licensed per deployment. Compare it with Nutrient if you are choosing an SDK rather than an API.

8. pdfRest Redact PDF API — rules-based redaction of known strings

pdfRest's API works in two calls: mark text you specify (exact strings or regular expressions) with a preview, then apply the redactions to produce a clean PDF. It is precise and predictable for the cases where you know what to remove (a case number, a list of names) and does not attempt to find PII it was not told about. Available as a cloud API on per-call pricing or self-hosted.

9. ConvertAPI — redaction as one step in a conversion workflow

ConvertAPI is a general file-conversion API that includes an AI-based PDF redaction endpoint. If your pipeline already converts, merges or compresses files through ConvertAPI, adding redaction there keeps one vendor. Detection controls and OCR behaviour are less extensive than the dedicated services above, so test on your real documents first.

Which one should you choose?

  • You want to send a PDF and get a redacted PDF back, with detection handled: Redact PDF AI. Nutrient if you are already on its SDK.
  • Files cannot leave your network: Private AI (commercial) or Presidio (open source).
  • You already run a document pipeline on Azure or AWS and want to own every step: Azure AI Language or AWS Comprehend, plus their OCR services, plus your own rendering and flattening.
  • You know exactly which strings to remove: pdfRest.
  • You ship an app and need redaction inside it: Apryse or Nutrient SDK.

Python quickstart: redact a PDF with Redact PDF AI

The whole flow is three calls: create a job, poll it, download each document.

import time
import requests

API = "https://www.redact-pdf.ai"
HEADERS = {"X-API-Key": "YOUR_API_KEY"}

# 1. Create a job (multipart upload, choose PII categories, ephemeral retention)
with open("contract.pdf", "rb") as f:
    job = requests.post(
        f"{API}/v1/jobs",
        headers={**HEADERS, "X-Idempotency-Key": "contract-2026-09-11"},
        files={"files": ("contract.pdf", f, "application/pdf")},
        data={
            "pii_categories": '["Person","Email","PhoneNumber","IBAN"]',
            "retention": "ephemeral",
        },
        timeout=60,
    ).json()

# 2. Poll until every document is terminal (redacted or error)
while True:
    job = requests.get(f"{API}/v1/jobs/{job['job_id']}", headers=HEADERS, timeout=30).json()
    if all(d["status"] in ("redacted", "error") for d in job["documents"]):
        break
    time.sleep(2)

# 3. Download the flattened, redacted output
for doc in job["documents"]:
    if doc["status"] == "redacted":
        pdf = requests.get(f"{API}/v1/documents/{doc['id']}/output", headers=HEADERS, timeout=60)
        with open(f"redacted-{doc['id']}.pdf", "wb") as out:
            out.write(pdf.content)

Treat 402 as quota exhausted and 429 as rate-limited (back off and retry with the same idempotency key). The Python quickstart and API reference cover error codes, retention and the studio hand-off.

Integration checklist

When you wire any redaction API into your stack, verify:

  • Keys are stored server-side and rotatable
  • Jobs are processed async with status polling or webhooks
  • PII categories are set per job to match each document type
  • Retention mode matches your data policy (delete vs. review)
  • Retries use an idempotency key
  • Backoff handles 429; quota handling covers 402
  • Output is verified to have no recoverable text layer (select, search, extract)
  • Data residency and no-training guarantees meet your compliance needs

FAQ

What is the best PDF redaction API? For PDF-in, redacted-PDF-out with AI detection and no pipeline to build, Redact PDF AI. Nutrient if you already use its SDK, Private AI for on-premise and the broadest language coverage, Presidio for open source, and Azure AI Language or AWS Comprehend if you want to assemble the pipeline yourself.

Is there a free PDF redaction API? Redact PDF AI gives free credits on signup and a no-key demo endpoint. Microsoft Presidio is free and open source but you host and operate it. Cloud providers offer free tiers for text PII detection, not for the full PDF pipeline.

Can a redaction API handle scanned PDFs? Only if it includes OCR. Redact PDF AI, Nutrient, Private AI and ConvertAPI do; Azure, AWS and Presidio need a separate OCR step; pdfRest works on the existing text layer.

Why async instead of a synchronous endpoint? OCR and PII detection take time on scanned or multi-page documents. Async jobs keep your request layer fast and let you process large batches without timeouts.

Is the output really irreversible? It is if the pages are rasterized or the content is removed from the page stream and the file is rewritten. Ask each vendor which they do; a drawn annotation is not a redaction. Redact PDF AI rasterizes every page and strips the text layer and metadata.

Can an AI agent call these APIs? Any REST API can be called from an agent. Redact PDF AI additionally ships an official MCP server (redact-pdf-mcp), an llms.txt and Markdown docs so agents can discover and use it without custom glue.

The bottom line

A redaction API lives or dies on what it returns (a redacted PDF or just entities), how it finds data, whether it reads scans, and whether the output can be undone. Redact PDF AI covers all four on an EU/Swiss-hosted pipeline; the alternatives above each win a specific situation. Get an API key and run a job, or read the developer docs first.