Skip to content

Navigation Menu

Explore
By company size
By use case
By industry
View all solutions
Topics
- AI
- DevOps
- Security
- Software Development
- View all
Explore
- GitHub Sponsors
  Fund open source developers
- The ReadME Project
  GitHub community articles
Repositories
- Enterprise platform
  AI-powered developer platform
Available add-ons
Pricing

Search code, repositories, users, issues, pull requests...

Search

Clear

Search syntax tips

Provide feedback

We read every piece of feedback, and take your input very seriously.

Include my email address so I can be contacted

Saved searches

Use saved searches to filter your results more quickly

Name

Query

To see all available qualifiers, see our documentation.

You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session.

Dismiss alert

Unstructured-IO / unstructured Public

Notifications You must be signed in to change notification settings
Fork 828
Star 9.9k

Code
Issues 134
Pull requests 51
Discussions
Actions
Projects 1
Security
Insights

Additional navigation options

Code
Issues
Pull requests
Discussions
Actions
Projects
Security
Insights

Releases: Unstructured-IO/unstructured

Releases · Unstructured-IO/unstructured

0.16.16

27 Jan 23:30

christinestraub

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.16 Latest

Latest

0.16.16

Enhancements

Features

Vectorize layout (inferred, extracted, and OCR) data structure Using np.ndarray to store a group of layout elements or text regions instead of using a list of objects. This improves the memory efficiency and compute speed around layout merging and deduplication.

Fixes

Add auto-download for NLTK for Python Enviroment When user import tokenize, It will automatic download nltk data from tokenize.py file. Added AUTO_DOWNLOAD_NLTK flag in tokenize.py to download NLTK_DATA.
Correctly patch pdfminer to avoid PDF repair. The patch applied to pdfminer's parser caused it to occasionally split tokens in content streams, throwing PDFSyntaxError. Repairing these PDFs sometimes failed (since they were not actually invalid) resulting in unnecessary OCR fallback.

Drop usage of ndjson dependency

Assets 2

Loading

ErcinDedeoglu reacted with thumbs up emoji

ErcinDedeoglu reacted with hooray emoji

ErcinDedeoglu reacted with rocket emoji

All reactions

👍 1 reaction
🎉 1 reaction
🚀 1 reaction

1 person reacted

0.16.15

23 Jan 04:51

tbs17

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.15

Update unstructured-inference to 0.8.6 in requirements which removed layoutparser dependency libs
Update pdfminer-six to 20240706

Assets 2

Loading

All reactions

0.16.14

20 Jan 13:00

plutasnyy

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.14

Enhancements

Features

Fixes

Fix an issue with multiple values for infer_table_structure when paritioning email with image attachements the kwarg calls into partition to partition the image already contains infer_table_structure. Now partition function checks if the kwarg has infer_table_structure already

Assets 2

Loading

All reactions

0.16.13

13 Jan 15:40

plutasnyy

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.13

Enhancements

Add character-level filtering for tesseract output. It is controllable via TESSERACT_CHARACTER_CONFIDENCE_THRESHOLD environment variable.

Features

Fixes

Fix NLTK Download to use nltk assets in docker image
removed the ability to automatically download nltk package if missing

Assets 2

Loading

All reactions

0.16.12

05 Jan 22:06

cragwolfe

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.12

0.16.12

Enhancements

Prepare auto-partitioning for pluggable partitioners. Move toward a uniform partitioner call signature so a custom or override partitioner can be registered without code changes.
Add NDJSON file type support.

Features

Fixes

Base image has been updated.
Upgrade ruff to latest. Previously the ruff version was pinned to <0.5. Remove that pin and fix the handful of lint items that resulted.
CSV with asserted XLS content-type is correctly identified as CSV. Resolves a bug where a CSV file with an asserted content-type of application/vnd.ms-excel was incorrectly identified as an XLS file.
Improve element-type mapping for Chinese text. Fixes bug where Chinese text would produce large numbers of false-positive Title elements.
Improve element-type mapping for HTML. Fixes bug where certain non-title elements were classified as Title.

Assets 2

Loading

All reactions

0.16.11

10 Dec 00:51

scanny

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.11

Enhancements

Enhance quote standardization tests with additional Unicode scenarios
Relax table segregation rule in chunking. Previously a Table element was always segregated into its own pre-chunk such that the Table appeared alone in a chunk or was split into multiple TableChunk elements, but never combined with Text-subtype elements. Allow table elements to be combined with other elements in the same chunk when space allows.
Compute chunk length based solely on element.text. Previously .metadata.text_as_html was also considered and since it is always longer that the text (due to HTML tag overhead) it was the effective length criterion. Remove text-as-html from the length calculation such that text-length is the sole criterion for sizing a chunk.

Features

Fixes

Fix ipv4 regex to correctly include up to three digit octets.

Assets 2

Loading

TAJ2003 reacted with thumbs up emoji

All reactions

👍 1 reaction

1 person reacted

0.16.10

07 Dec 18:13

tbs17

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.10

0.16.10

Enhancements

Features

Fixes

Fix original file doctype detection from cct converted file paths for metrics calculation.

Assets 2

Loading

TAJ2003 reacted with thumbs up emoji

All reactions

👍 1 reaction

1 person reacted

0.16.9

02 Dec 21:52

vangheem

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.9

What's Changed

chore: fix CHANGELOG formatting by @cragwolfe in #3800
Use native ntlk download by @vangheem in #3796

Full Changelog: 0.16.8...0.16.9

Contributors

vangheem and cragwolfe

Assets 2

Loading

All reactions

0.16.8

26 Nov 19:39

plutasnyy

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.8

0.16.8

Enhancements

Metrics: Weighted table average is optional

Features

Fixes

Assets 2

Loading

All reactions

0.16.7

26 Nov 17:47

plutasnyy

This commit was created on GitHub.com and signed with GitHub’s verified signature.

GPG key ID: B5690EEEBB952194

Learn about vigilant mode.

Compare

Choose a tag to compare

Loading

0.16.7

0.16.7

Enhancements

Add image_alt_mode to partition_html Adds an image_alt_mode parameter to partition_html() to control how alt text is extracted from images in HTML documents for html_parser_version=v2 . The parameter can be set to to_text to extract alt text as text from <img> html tags

Features

Fixes

Assets 2

Loading

All reactions

Previous 1 2 3 4 5 … 16 17 Next

Previous Next

Footer

© 2025 GitHub, Inc.

Footer navigation

Terms
Privacy
Security
Status
Docs
Contact

You can’t perform that action at this time.