The multi-format text data processing system is based on document structure analysis and semantically rich text segmentation.
VN126787APending Publication Date: 2026-07-01CÔNG TY CỔ PHẦN AIZ
0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- VN · VN
- Patent Type
- Applications
- Current Assignee / Owner
- CÔNG TY CỔ PHẦN AIZ
- Filing Date
- 2026-06-08
- Publication Date
- 2026-07-01
Smart Images

Figure VN1202604672_0
Abstract
The invention relates to a system for processing text data from multi-format documents based on document structure identification and semantically rich text segmentation. Accordingly, the input document is pre-processed and normalized, then its structure is determined to partition it into text areas, image areas, table or graph areas, and related data areas. Content within each data area is extracted using corresponding modules, where image and table data are described or summarized for integration back into the text content. Based on the document structure map, text content, and document information, the system performs text segmentation according to structure and content relationships, and identifies links between text segments.The text is then post-processed to remove noise, repeated headers and footers, and normalize the content, thereby creating structured, coherent text suitable for search, information retrieval, automated querying, document indexing, and knowledge management systems.
Need to check novelty before this filing date? Find Prior Art