Java で PDF Facades を使用する方法

Java で PDF Facades を使用する方法

注釈のフラット化

すべての注釈をページのコンテンツストリームにフラット化し、個別の注釈オブジェクトとして編集したり削除したりできなくします。これは、印刷や文書のアーカイブを行う前の一般的な手順です:

try (PdfAnnotationEditor editor = new PdfAnnotationEditor()) {
    editor.bindPdf("annotated.pdf");
    editor.flattenAnnotations();
    editor.save("flat.pdf");
}

ブックマークの削除

文書からすべてのアウトラインブックマークを削除します。ブックマーク(アウトラインとも呼ばれます)は、PDF リーダーのサイドバーに表示されるナビゲーションツリーのエントリです:

try (PdfBookmarkEditor editor = new PdfBookmarkEditor()) {
    editor.bindPdf("input.pdf");
    editor.deleteBookmarks();
    editor.save("no-bookmarks.pdf");
}

テキストの置換

文書のすべてのページで文字列を置換します。これは、文書を配布する前にテンプレートのプレースホルダーを埋める際に便利です:

try (PdfContentEditor editor = new PdfContentEditor()) {
    editor.bindPdf("template.pdf");
    editor.replaceText("{{Date}}", "2026-05-24");
    editor.save("output.pdf");
}

文書を暗号化

PdfFileSecurity を使用してアクセス権限付きの AES-256 暗号化を適用します。DocumentPrivilege.getForbidAll() プリセットはすべての操作を無効にします。その後、個々の権限はセッターメソッドで再度有効化できます:

try (PdfFileSecurity security = new PdfFileSecurity("input.pdf", "secured.pdf")) {
    DocumentPrivilege priv = DocumentPrivilege.getForbidAll();
    priv.setAllowPrint(true);
    security.encryptFile("user", "owner", priv, KeySize.x256);
}

テキストを抽出

PdfExtractor を使用してページテキストを反復処理します。まず extractText() を呼び出してすべてのページを処理し、次に getNextPageText() で各ページのテキストを取得します:

try (PdfExtractor extractor = new PdfExtractor()) {
    extractor.bindPdf("document.pdf");
    extractor.extractText();
    while (extractor.hasNextPageText()) {
        String text = extractor.getNextPageText();
        System.out.println(text);
    }
}

参照

 日本語