Best PDF Redaction APIs (2026): 9 Options Compared
The best PDF redaction API depends on one question: do you want to send a PDF and get a redacted PDF back, or do you want building blocks and assemble the pipeline yourself? For the first case, Redact PDF AI is the best overall choice in 2026: PDF or image in, flattened and irreversibly redacted PDF out, with AI detection of names, addresses, IBANs and other PII, OCR for scans, async jobs, and EU/Swiss hosting. Nutrient is the best pick if you already embed its viewer SDK, Private AI if you need an on-premise container with the widest entity and language coverage, Microsoft Presidio if you want open source, and Azure AI Language or AWS Comprehend if you are happy to build OCR, drawing and flattening around a text-only PII detector.
This guide compares nine options on what actually decides an integration: input and output formats, how sensitive data is found, OCR, hosting, and how much pipeline you have to operate yourself. Vendor details are from their public documentation and pricing pages as of September 2026; check them before you commit.
What to evaluate in a redaction API
Before comparing vendors, get clear on the criteria that determine whether an API survives contact with production:
- What goes in, what comes out. A "redaction API" that takes text and returns text with
[PERSON]tokens leaves you to do OCR, map offsets back to page coordinates, draw the boxes and flatten the file. A PDF-in, PDF-out API does all of that. - How sensitive data is found. Regex and search strings catch what you already know (a phone-number pattern, a client's name). AI entity detection catches the names and addresses you did not know were there. Most real leaks are the second kind.
- OCR. Scanned contracts, faxes and phone photos have no text layer. If the API cannot read them, the pipeline needs a second vendor.
- Irreversible output. The output must be flattened or rasterized so nothing is recoverable by copy-paste, search, or glyph-position analysis. A box drawn over text is not a redaction.
- Async job handling. OCR and detection take seconds to minutes on long documents. Look for a job id, status polling or webhooks, and idempotent retries.
- Granular PII control. Choose categories per job, and exclude terms that must never be redacted (your own company name).
- Retention and residency. Where the file is processed, how long it is kept, whether it trains anyone's models, and whether you can delete it on demand.
- Errors, limits and a real spec. Explicit quota (
402) and rate-limit (429) responses, and an OpenAPI definition you can generate clients from.
The APIs compared
| API | Input → output | Detection | OCR | Hosting | Best for |
|---|---|---|---|---|---|
| Redact PDF AI | PDF, JPG, PNG → flattened PDF | AI entities (8 categories) + always/never terms | Built in, 100+ languages | EU & Swiss Azure, no training, auto-delete | PDF-in, redacted-PDF-out with nothing to build |
| Nutrient AI Redaction API | PDF, Office, images → PDF | AI entities + search/regex | Built in | Cloud API or self-hosted Document Engine | Teams already on the Nutrient SDK / viewer |
| Private AI | Text, PDF, DOCX, images → redacted file or entities | AI, 50+ entity types, 50+ languages | Built in | Container in your VPC, or their cloud | On-premise, multi-format, multi-language |
| Azure AI Language (PII) | Text → entities + redacted text | AI entities | Separate (Document Intelligence) | Azure | Building your own pipeline on Azure |
| AWS Comprehend (PII) | Text → entities / redacted text | AI entities | Separate (Textract) | AWS | Building your own pipeline on AWS |
| Microsoft Presidio | Text and images → anonymized text / redacted image | Rule + model recognizers, extensible | Separate (bring your own) | Self-hosted, open source | Open source, full control |
| Apryse SDK | PDF → PDF, in-process | Search / regex, your own detector | Add-on | Your servers or client apps | Embedding redaction in your own app |
| pdfRest Redact PDF API | PDF → PDF (mark, then apply) | Text strings and regex you supply | No | Cloud or self-hosted | Rules-based redaction of known strings |
| ConvertAPI (PDF redact) | PDF → PDF | AI-based | Yes | Cloud | Adding redaction to an existing ConvertAPI workflow |
1. Redact PDF AI — best overall for PDF-in, redacted-PDF-out
The Redact PDF AI API runs the same pipeline as the web app, exposed as a REST interface: Azure OCR, AI PII detection, and irreversible rasterized redaction. You upload one or many files, get a job id back immediately, poll until each document reaches redacted, and download a flattened PDF.
- Detection. Choose per job from
Person,Email,PhoneNumber,Address,Organization,Date,IBANandCreditCard, or use your saved defaults. Add terms that must always be redacted and terms that never should be. - Scans and images. JPG and PNG are accepted directly; scanned PDFs are read with OCR in 100+ languages.
- Output. Every page is rasterized and the text layer and metadata dropped, so the output cannot be reversed by copy-paste, search, or the glyph-position attacks that break text-layer redactions.
- Retention modes.
ephemeral(default) deletes originals after processing;studiokeeps originals and masks so a human can review in the editor and re-export. - Reliability.
X-Idempotency-Keymakes retries safe;402/429are explicit; an OpenAPI spec and docs cover the full schema. An official MCP server exposes the same API to AI agents. - Hosting and compliance. Processed on Microsoft Azure in the EU and Switzerland, AES-256 at rest, TLS in transit, on SOC 2 Type II / ISO 27001 infrastructure, HIPAA-eligible; documents are never used to train models.
- Price. Free credits on signup; Starter is US$50/month for 1,000 pages, Business US$250/month for 6,000 pages; pay-as-you-go credit packs for irregular volume.
Limits to know: it is a document redactor, not a general PDF toolkit (no merge/convert), and it is a hosted API rather than an in-process SDK. If files may not leave your network, look at Private AI or Presidio.
Best for: product and platform teams that need user-uploaded PDFs and images redacted reliably, with EU/Swiss residency and a no-training guarantee, without building OCR, detection, drawing and flattening themselves.
2. Nutrient AI Redaction API — best if you already use the Nutrient SDK
Nutrient (formerly PSPDFKit) sells document SDKs for web, mobile and server, and its AI redaction API sits on that stack: send a document, have entities detected, and receive a permanently redacted PDF. It supports native and scanned documents and can run as a cloud API or on your own infrastructure through Document Engine. The natural fit is a team that already renders and annotates PDFs with Nutrient's viewer and wants redaction from the same vendor. Pricing is credit-based per API call for the cloud service and licence-based for self-hosting, so budget for both if you need the on-premise option.
3. Private AI — best for on-premise and the widest coverage
Private AI is a PII detection and redaction service that runs as a container inside your VPC (or as their hosted API). It handles text, PDFs, Office files and images with built-in OCR, recognises 50+ entity types across 50+ languages, and returns either the redacted file or the entities with offsets. It is the strongest choice when data cannot leave your environment or when you redact many formats and languages. It is enterprise-priced and you operate the container, which is more work than calling a hosted endpoint.
4. Azure AI Language — PII detection as a building block
Azure's PII detection returns entities and a redactedText string for plain text, priced per 1,000 text records. It does not take a PDF and hand back a redacted PDF: you OCR with Azure Document Intelligence, map the detected spans back to coordinates, draw and flatten yourself. (A native-document mode has been in preview; check its current status before designing around it.) A sensible choice if you are all-in on Azure and want to own the pipeline. Redact PDF AI uses Azure AI services underneath and adds exactly that missing pipeline.
5. AWS Comprehend — the same, on AWS
Amazon Comprehend's DetectPiiEntities returns PII spans for text (English and a limited set of other languages), and its async jobs can write redacted text to S3. For PDFs you pair it with Amazon Textract for OCR and build the coordinate mapping, drawing and flattening. Billing is per unit of characters. Same trade-off as Azure: cheap building blocks, significant assembly.
6. Microsoft Presidio — best open source
Presidio is Microsoft's open-source PII framework: an analyzer with rule- and model-based recognizers, an anonymizer for text, and an image redactor that can black out detected text in images (including DICOM). You host it, tune the recognizers, and bring your own OCR for PDFs. Free, fully controllable, and a good fit for teams with ML and ops capacity who need to audit every rule; not a turnkey PDF service.
7. Apryse SDK — redaction inside your own application
Apryse (formerly PDFTron) provides an in-process SDK with a redaction module: search text or supply regions, then apply true redaction that removes the content. Detection is search and regex based unless you plug in your own detector. It is the option for embedding redaction into a desktop, mobile or server application you ship, licensed per deployment. Compare it with Nutrient if you are choosing an SDK rather than an API.
8. pdfRest Redact PDF API — rules-based redaction of known strings
pdfRest's API works in two calls: mark text you specify (exact strings or regular expressions) with a preview, then apply the redactions to produce a clean PDF. It is precise and predictable for the cases where you know what to remove (a case number, a list of names) and does not attempt to find PII it was not told about. Available as a cloud API on per-call pricing or self-hosted.
9. ConvertAPI — redaction as one step in a conversion workflow
ConvertAPI is a general file-conversion API that includes an AI-based PDF redaction endpoint. If your pipeline already converts, merges or compresses files through ConvertAPI, adding redaction there keeps one vendor. Detection controls and OCR behaviour are less extensive than the dedicated services above, so test on your real documents first.
Which one should you choose?
- You want to send a PDF and get a redacted PDF back, with detection handled: Redact PDF AI. Nutrient if you are already on its SDK.
- Files cannot leave your network: Private AI (commercial) or Presidio (open source).
- You already run a document pipeline on Azure or AWS and want to own every step: Azure AI Language or AWS Comprehend, plus their OCR services, plus your own rendering and flattening.
- You know exactly which strings to remove: pdfRest.
- You ship an app and need redaction inside it: Apryse or Nutrient SDK.
Python quickstart: redact a PDF with Redact PDF AI
The whole flow is three calls: create a job, poll it, download each document.
import time
import requests
API = "https://www.redact-pdf.ai"
HEADERS = {"X-API-Key": "YOUR_API_KEY"}
# 1. Create a job (multipart upload, choose PII categories, ephemeral retention)
with open("contract.pdf", "rb") as f:
job = requests.post(
f"{API}/v1/jobs",
headers={**HEADERS, "X-Idempotency-Key": "contract-2026-09-11"},
files={"files": ("contract.pdf", f, "application/pdf")},
data={
"pii_categories": '["Person","Email","PhoneNumber","IBAN"]',
"retention": "ephemeral",
},
timeout=60,
).json()
# 2. Poll until every document is terminal (redacted or error)
while True:
job = requests.get(f"{API}/v1/jobs/{job['job_id']}", headers=HEADERS, timeout=30).json()
if all(d["status"] in ("redacted", "error") for d in job["documents"]):
break
time.sleep(2)
# 3. Download the flattened, redacted output
for doc in job["documents"]:
if doc["status"] == "redacted":
pdf = requests.get(f"{API}/v1/documents/{doc['id']}/output", headers=HEADERS, timeout=60)
with open(f"redacted-{doc['id']}.pdf", "wb") as out:
out.write(pdf.content)
Treat 402 as quota exhausted and 429 as rate-limited (back off and retry with the same idempotency key). The Python quickstart and API reference cover error codes, retention and the studio hand-off.
Integration checklist
When you wire any redaction API into your stack, verify:
- Keys are stored server-side and rotatable
- Jobs are processed async with status polling or webhooks
- PII categories are set per job to match each document type
- Retention mode matches your data policy (delete vs. review)
- Retries use an idempotency key
- Backoff handles
429; quota handling covers402 - Output is verified to have no recoverable text layer (select, search, extract)
- Data residency and no-training guarantees meet your compliance needs
FAQ
What is the best PDF redaction API? For PDF-in, redacted-PDF-out with AI detection and no pipeline to build, Redact PDF AI. Nutrient if you already use its SDK, Private AI for on-premise and the broadest language coverage, Presidio for open source, and Azure AI Language or AWS Comprehend if you want to assemble the pipeline yourself.
Is there a free PDF redaction API? Redact PDF AI gives free credits on signup and a no-key demo endpoint. Microsoft Presidio is free and open source but you host and operate it. Cloud providers offer free tiers for text PII detection, not for the full PDF pipeline.
Can a redaction API handle scanned PDFs? Only if it includes OCR. Redact PDF AI, Nutrient, Private AI and ConvertAPI do; Azure, AWS and Presidio need a separate OCR step; pdfRest works on the existing text layer.
Why async instead of a synchronous endpoint? OCR and PII detection take time on scanned or multi-page documents. Async jobs keep your request layer fast and let you process large batches without timeouts.
Is the output really irreversible? It is if the pages are rasterized or the content is removed from the page stream and the file is rewritten. Ask each vendor which they do; a drawn annotation is not a redaction. Redact PDF AI rasterizes every page and strips the text layer and metadata.
Can an AI agent call these APIs?
Any REST API can be called from an agent. Redact PDF AI additionally ships an official MCP server (redact-pdf-mcp), an llms.txt and Markdown docs so agents can discover and use it without custom glue.
The bottom line
A redaction API lives or dies on what it returns (a redacted PDF or just entities), how it finds data, whether it reads scans, and whether the output can be undone. Redact PDF AI covers all four on an EU/Swiss-hosted pipeline; the alternatives above each win a specific situation. Get an API key and run a job, or read the developer docs first.