The multi-format text data processing system is based on document structure analysis and semantically rich text segmentation.

VN126787APending Publication Date: 2026-07-01CÔNG TY CỔ PHẦN AIZ
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
VN · VN
Patent Type
Applications
Current Assignee / Owner
CÔNG TY CỔ PHẦN AIZ
Filing Date
2026-06-08
Publication Date
2026-07-01

Smart Images

  • Figure VN1202604672_0
    Figure VN1202604672_0
Patent Text Reader

Abstract

The invention relates to a system for processing text data from multi-format documents based on document structure identification and semantically rich text segmentation. Accordingly, the input document is pre-processed and normalized, then its structure is determined to partition it into text areas, image areas, table or graph areas, and related data areas. Content within each data area is extracted using corresponding modules, where image and table data are described or summarized for integration back into the text content. Based on the document structure map, text content, and document information, the system performs text segmentation according to structure and content relationships, and identifies links between text segments.The text is then post-processed to remove noise, repeated headers and footers, and normalize the content, thereby creating structured, coherent text suitable for search, information retrieval, automated querying, document indexing, and knowledge management systems.
Need to check novelty before this filing date? Find Prior Art