[Release] GroupDocs.Parser for Java 26.9

We are pleased to announce the release of GroupDocs.Parser for Java 26.9. This version brings the annotation extraction API to Java, extends HTML support, and restores metered licensing.

What’s new

Annotation extraction. Comments and annotations can now be extracted from PDF documents, word processing documents and presentations — either on their own through Parser.getAnnotations() and Parser.getAnnotations(pageIndex), or together with the document text through TextOptions.setIncludeAnnotations(true). Every item carries its type, timestamp, author name and author initials:

try (Parser parser = new Parser("annotated.docx")) {
    for (AnnotationItem annotation : parser.getAnnotations()) {
        System.out.println(annotation.getAuthorName() + ": " + annotation.getValue());
    }
}

More HTML support. getPageCount, getTextAreas and getTables are now implemented for HTML documents, so tables and positioned text can be extracted from HTML the same way as from word processing documents.

Smarter file type detection. When Parser.getFileInfo is given a file path, a known file
extension now narrows content-based detection to that format, which resolves formats that content sniffing alone cannot tell apart.

What’s fixed

  • Metered licensing. Metered.setMeteredKey failed with Signature encoding error in 26.5, so a metered license could not be applied at all. It works again, and the bundled components are now licensed independently of each other.
  • Memory usage on image-heavy PDF documents. A 52 MB document that previously could not be
    parsed in a 2 GB heap now parses in a 768 MB heap.
  • Encrypted documents. getFileInfo now reports the real file size and the encrypted flag
    instead of an empty result.
  • HTML images. getImages(pageIndex) now honours the page index.

All bundled Aspose components have been updated (Cells 26.3, Email 26.1, Imaging 25.12, PDF 26.2, Slides 26.2, Words 26.3, BarCode 26.4, ZIP 25.10).

Please note

Two changes affect text extraction from word processing documents:

  • comments are no longer part of the extracted text by default — call
    TextOptions.setIncludeAnnotations(true) if you need them;
  • whole-document extraction no longer emits the form feed character (\f) between pages; pages are separated by a line break.

Presentation page previews are now rendered through the current Aspose.Slides image API and may differ slightly from 26.5.

Getting the release

<repositories>
    <repository>
        <id>GroupDocsJavaAPI</id>
        <name>GroupDocs Java API</name>
        <url>https://releases.groupdocs.com/java/repo/</url>
    </repository>
</repositories>

<dependencies>
    <dependency>
        <groupId>com.groupdocs</groupId>
        <artifactId>groupdocs-parser</artifactId>
        <version>26.9</version>
    </dependency>
</dependencies>

The published JAR now comes with a GPG detached signature (.asc) signed with the GroupDocs public key, so the download can be verified with gpg --verify groupdocs-parser-26.9.jar.asc.

Useful links

As always, we are happy to hear your feedback — post your questions in this category and we will be glad to help.