A layout structure recovery method for a complex PDF form and a related device

CN122761401APending Publication Date: 2026-09-15NINGBO DIGITAL TWIN (EASTERN UNIV OF TECH) RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611053371.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0005]本发明提供一种面向复杂PDF表格的版面结构恢复方法及相关装置,采用本方法能够有效解决现有技术在面对无线框、弱边界、多行跨列、底色干扰、文本密排及中文连续排布等复杂PDF表格场景时,难以稳定恢复表格结构、准确识别行列边界及处理文本归属的问题,满足了后续结构化利用及业务应用的需求

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122761401A_ABST
    Figure CN122761401A_ABST
Patent Text Reader

Abstract

The application belongs to the field of complex PDF table layout structure recovery, and discloses a complex PDF table layout structure recovery method and related device. The method extracts table region text and coordinates standardization, combines line and background color signal to recover table reference structure, recovers text structure based on text position distribution, alignment and proximity relationship, fuses the two in the row and column direction to obtain the final structure, and finally performs text attribution arrangement and layout reconstruction. The method can stably recover the fine-grained structure of a complex table, accurately identify the row and column boundaries and text attribution, improve the robustness and accuracy of table analysis, and provide a reliable foundation for subsequent structured utilization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of PDF table layout structure restoration technology, and particularly relates to a method and related apparatus for restoring the layout structure of complex PDF tables. Background Technology

[0002] In various business scenarios such as government affairs, finance and taxation, customs, auditing, risk control, records management, and enterprise knowledge processing, key information such as policy documents, tax rules, and financial statements are increasingly being generated, transferred, archived, and stored in PDF format. This has led to a continuous increase in the demand for automatic recognition, table extraction, and structured utilization of PDF document content. Since a large amount of core information is not presented in plain text but embedded in PDF pages in the form of tables, the accurate restoration of the table row and column structure, cell boundaries, and text hierarchy directly determines the effectiveness of subsequent data governance, rule matching, information retrieval, risk identification, and intelligent analysis, and is a core prerequisite for achieving efficient utilization of PDF document information.

[0003] In real-world business scenarios, PDF tables are generated from fixed layout results and are essentially combinations of visual objects on the page. They exhibit characteristics such as implicit structure, complex layout, and mixed visual elements. This problem is particularly pronounced in Chinese document scenarios—Chinese text is continuously arranged, lacks natural word separation, and often includes vertical text and mixed simplified and traditional characters, significantly increasing the difficulty of column boundary judgment and text attribution identification. Existing PDF table extraction and parsing technologies mainly fall into four categories, all with significant limitations: methods based on line or border detection cannot adapt to implicit tables with no borders or weak lines; methods based on OCR results or text coordinate clustering are easily affected by the regularity of text layout, leading to column division errors and text adhesion problems in densely formatted Chinese text; unified extraction methods for multiple document types focus on comprehensive page content coverage, making it difficult to achieve fine-grained restoration of complex table structures; solutions for industry applications focus on downstream business result generation, failing to fundamentally solve the problem of high-precision table structure restoration, and upstream parsing errors can be transmitted to subsequent business processes, affecting the accuracy of business judgments.

[0004] It is evident that existing technologies struggle to reliably restore table structures, accurately identify row and column boundaries, and process text ownership when faced with complex PDF table scenarios such as wireless frames, weak boundaries, multiple rows spanning columns, background color interference, dense text layout, and continuous Chinese text arrangement. Consequently, they fail to meet the needs of subsequent structured utilization and business applications. Summary of the Invention

[0005] This invention provides a method and related apparatus for restoring the layout structure of complex PDF tables. This method can effectively solve the problems of existing technologies in dealing with complex PDF table scenarios such as no frame, weak boundaries, multiple rows spanning columns, background color interference, dense text layout, and continuous Chinese text layout. These problems make it difficult to stably restore the table structure, accurately identify row and column boundaries, and handle text ownership, thus meeting the needs of subsequent structured utilization and business applications.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A method for restoring the layout structure of complex PDF tables, comprising: Obtain the target table region in the target PDF document, extract the text objects in the target table region, and perform coordinate standardization on the text objects to obtain a standardized set of text objects under a unified image coordinate system. Based on the line signals and background color region signals in the target table area, the table reference structure is recovered; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; Based on the positional distribution, alignment, and proximity of a standardized set of text objects, the text structure is restored; wherein the text structure includes at least a text column structure and a visual row result; The table reference structure and the text structure are merged in the column direction and the row direction respectively to obtain the final column structure and the final row structure; Based on the final column structure and the final row structure, the text objects are reorganized into rows and columns and the layout is reconstructed to obtain a structured table result.

[0007] Further, the step of obtaining the target table region in the target PDF document, extracting the text objects in the target table region, and performing coordinate standardization processing on the text objects to obtain a standardized set of text objects in a unified image coordinate system includes: Obtain the target page in the target PDF document, determine the bounding box of the target table area corresponding to the table to be processed based on the page table detection results, and perform local cropping on the target page based on the bounding box to obtain a table sub-area containing only the content of the target table. The table sub-regions are rendered to generate table region images that retain the table's visual information and are used for subsequent structural skeleton detection; Extract text objects from a sub-region of the table to obtain a collection of text objects within the target table region; The position of the text object in the PDF coordinate system is converted into local coordinates relative to the top left corner of the target table area. Based on the ratio between the width and height of the target table area in the PDF coordinate system and the pixel width and height of the table area image, the local coordinates are mapped to the table area image coordinate system, thus completing the text object coordinate standardization. Boundary constraints are applied to the coordinate-standardized text objects to limit their left, right, top, and bottom boundaries to the range of the table area image, and abnormal text objects with width or height less than or equal to zero are filtered out, resulting in text objects that have undergone boundary constraints and filtering. Based on the text objects after boundary constraints and filtering, a standardized text object table is constructed, which includes at least text object identifiers, text content, coordinate information and size information, and finally a standardized text object set under a unified image coordinate system is obtained.

[0008] Furthermore, the process of restoring the table reference structure based on the line signals and background color region signals in the target table area includes: The image of the table region corresponding to the target table region is grayscaled and thresholded to separate the lines from the background. Vertical and horizontal lines are extracted through targeted morphological operations to obtain vertical and horizontal boundary candidates and corresponding line boxes and position sets, forming the initial space in the column and row directions. Merge adjacent line boundaries to form a stable set of vertical and horizontal boundary positions. Determine the minimum column width constraint and minimum row height constraint based on the effective spacing between adjacent boundaries. Eliminate unstructured short lines that cannot cross the minimum structural unit and retain and form line structure space in the column and row directions. Background reconstruction is performed on the table area image to remove text foreground interference, making the background color and background partition more continuous. Then, the background image is clustered and segmented to divide it into several background color regions. The decomposition positions between different background color regions are identified to form a color structure space in the column and row directions. The line structure space and the color structure space are merged in the column direction and the row direction respectively. The spatial correspondence between the two is determined. Based on the correspondence, the two types of structures are jointly confirmed, range segmented or independently preserved to form a merged structure space. Based on the fused structural space, output a column-direction and row-direction table reference structure that includes at least spatial identifiers, start and end boundaries, size information, and source information, and complete the restoration of the table reference structure.

[0009] Furthermore, the restoration of text structure based on the positional distribution, alignment, and proximity relationships of a standardized set of text objects includes: For a standardized set of text objects, the alignment relationship of the left and right boundaries and center coordinates of the text objects and the horizontal proximity relationship are comprehensively utilized in the column direction to divide the text objects into initial column bands. The main column band and the column band to be processed are distinguished according to the number of text objects in each initial column band. The column band to be processed is assigned and the main column band that meets the conditions is merged to obtain the text column structure. In the line direction, based on the upper boundary, lower boundary, and center ordinate of the text object, the text objects are clustered according to the principle of vertical proximity, and the text objects located in the same horizontal reading band are merged into the same visual line to complete the restoration of the text visual line. The system summarizes the results of the text column structure, text visual rows, and the attribution of text objects to the text column structure and text visual rows, outputs the table structure results on the text side, and finally completes the restoration of the text structure.

[0010] Furthermore, the fusion of the table reference structure and the text structure in the column direction and the row direction, respectively, to obtain the final column structure and the final row structure, includes: The column direction reference structure in the table reference structure is used as the parent column range, and the text column structure is used as the side column result of the text to be matched. Based on the boundary difference, overlap ratio, inclusion relationship and center position relationship between the two, the correspondence between the text column structure and the parent column range is determined. For the text objects to which the missing text column belongs, the attribution compensation is performed according to their position and overlap with the parent column range. The boundary correction is performed on the parent column range corresponding to the single final column structure to obtain the final column structure of the column direction. The row direction reference structure in the table reference structure is used as the parent row range, and the text visual row is used as the text side row result to be assigned. Based on the boundary difference, overlap ratio, inclusion relationship and center position relationship between the two, the correspondence between the visual row and the parent row range is determined, and the final row structure of the row direction is generated with the parent row range as the unit. The final column structure obtained by merging the columns and the final row structure obtained by merging the rows are combined to obtain the final column structure and final row structure of the table.

[0011] Furthermore, based on the final column structure and the final row structure, the text objects are reorganized into rows and columns and the layout is reconstructed to obtain a structured table result, including: Based on the final column structure and the final row structure, and combined with the standardized set of text objects, the column direction attribution results of text objects are uniformly organized, and the row direction attribution results are merged with the column direction attribution results according to the granularity of text objects, so that each text object can obtain the attribution information of both row direction and column direction, forming a unified text object-level attribution result. Text objects that have both row and column orientations are grouped according to their row and column identifiers. The text objects within the same group are then arranged and spliced ​​according to their page position order to obtain the text content of the corresponding cells, forming a cell-level layout result table. Expand the cell-level layout results table by row and column to generate a table-style layout reconstruction result, thus obtaining a structured table result.

[0012] Furthermore, the structured table results include one or more of the following: text object-level attribution results, cell-level layout result tables, table-style layout reconstruction results, and results of the unit to be inspected and its corresponding context.

[0013] A layout structure restoration system for complex PDF tables, comprising: The acquisition module is used to acquire the target table area in the target PDF document, extract the text objects in the target table area, and perform coordinate standardization processing on the text objects to obtain a set of standardized text objects under a unified image coordinate system. The first recovery module is used to recover the table reference structure based on the line signals and background color area signals in the target table area; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; The second recovery module is used to recover the text structure based on the positional distribution, alignment, and proximity relationships of a standardized set of text objects; wherein the text structure includes at least a text column structure and a visual row result; The fusion module is used to merge the table reference structure and the text structure in the column direction and the row direction, respectively, to obtain the final column structure and the final row structure. The sorting and reconstruction module is used to sort and reconstruct the layout of text objects based on the final column structure and the final row structure, resulting in a structured table.

[0014] A layout structure restoration device for complex PDF tables, comprising: Memory, used to store computer programs; A processor is used to implement the above-described method for restoring the layout structure of complex PDF tables when executing the computer program.

[0015] A computer-readable storage medium storing a computer program, which, when executed by a processor, is used to implement the above-described method for restoring the layout structure of complex PDF tables.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a method for restoring the layout structure of complex PDF tables. It extracts text from table areas and standardizes its coordinates, then combines line and background color signals to restore the table's reference structure. Simultaneously, it restores the text structure based on text position distribution, alignment, and proximity relationships. Finally, it merges these two methods along the row and column directions to obtain the final structure, and then performs text attribution sorting and layout reconstruction. This method complements the explicit visual cues provided by lines and background colors with the implicit layout information implied by text distribution. It overcomes the limitations of missing borders in frameless tables and alleviates the problems of text adhesion and column division errors in densely formatted Chinese text. Through fusion processing, it comprehensively considers the complexity of the layout and the continuity of the text, thus adapting to implicit structures and varied layouts. This method can stably restore the fine-grained structure of complex tables, accurately identify row and column boundaries and text attribution, improve the robustness and accuracy of table parsing, and provide a reliable foundation for subsequent structured utilization. Attached Figure Description

[0017] Figure 1 A flowchart illustrating the overall process of restoring the layout structure of complex PDF tables, as provided in this embodiment of the invention. Figure 2 A flowchart for obtaining target table areas and standardizing text objects provided in this embodiment of the invention; Figure 3 A flowchart for restoring the table reference structure provided in an embodiment of the present invention; Figure 4 A flowchart for restoring a table structure based on text objects provided in an embodiment of the present invention; Figure 5 A flowchart illustrating the fusion of table reference structure and text structure provided in this embodiment of the invention; Figure 6 A flowchart of page layout reconstruction based on row and column attribution provided for embodiments of the present invention; Figure 7 This is a core flowchart of a method for restoring the layout structure of complex PDF tables, provided in an embodiment of the present invention. Figure 8 This is a schematic diagram of a layout structure restoration system for complex PDF tables provided in an embodiment of the present invention. Detailed Implementation

[0018] To further understand the content of this invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.

[0019] The technical terms involved in this invention are explained below: PDF stands for Portable Document Format.

[0020] OCR stands for Optical Character Recognition.

[0021] As mentioned in the background section, the traditional techniques for extracting and parsing PDF tables, along with their corresponding shortcomings, are as follows: First, table extraction methods based on line or border detection: These methods typically recover the row and column structure of tables by detecting horizontal and vertical lines, border outlines, or grid relationships on a page. For tables with complete lines, clear boundaries, and regular grids, these methods achieve good recognition results. However, in real-world documents, tables often lack complete borders, or only have partial lines, weak lines, broken lines, or even implicit table structures formed by background blocks, text alignment, and white space. In such cases, methods relying solely on line detection are prone to problems such as missed horizontal and vertical boundaries, incomplete cell segmentation, broken row and column relationships, and structural misalignment, making them unsuitable for parsing tables without borders or with weak structures.

[0022] Second, table parsing methods based on OCR results or text coordinate clustering: These methods typically infer column boundaries, row groupings, and cell assignments in tables based on the coordinates of character boxes, word boxes, and text blocks, as well as their alignment, distance, or density distribution. While these methods somewhat eliminate reliance on explicit table lines, their effectiveness often depends on the regularity of text layout and the stability of spacing features. In English text, there are numerous inter-word spaces; relying solely on whitespace or horizontal distance for structural division can easily misinterpret normal cell spacing as column separation. In densely formatted Chinese tables, text is usually continuous with unstable inter-column spacing, often accompanied by line breaks, cross-column extensions, and the insertion of local descriptions, which can lead to inaccurate column boundary identification, text adhering to adjacent columns, incorrect splitting of multiple lines, and incorrect cell assignments. Therefore, methods relying solely on OCR or text coordinate clustering still have significant limitations in complex Chinese PDF table scenarios.

[0023] Third, a unified extraction method for multiple document types or multimodal content: These methods typically use PDFs as input in the document understanding process, uniformly identifying and extracting various elements such as text, tables, images, stamps, and titles from the page. While this approach has good versatility and is suitable for the overall extraction of complex document content, its focus is usually on comprehensive coverage of page content rather than fine-grained recovery of complex table structures. For tables containing multi-level headers, row / column spanning relationships, implicit column boundaries, background color interference, and abnormal explanatory text, these methods often struggle to fully characterize the table's hierarchical structure and cell semantic mapping relationships, thus affecting the effectiveness of subsequent structured utilization.

[0024] Fourth, document verification or information extraction solutions for industry applications: These solutions typically treat table extraction as a prerequisite for downstream entity identification, field comparison, rule validation, risk identification, event extraction, or business verification. Their focus is more on generating business results and making rule judgments than on high-precision recovery of the table structure itself. If there are issues in the upstream table parsing stage, such as incorrect column partitioning, abnormal row assignments, merging of cross-row text, or misplacement of explanatory text in cells, these errors will be further propagated to subsequent business identification and verification processes, leading to incomplete field extraction, rule matching deviations, and even incorrect business judgments. Therefore, while these solutions can cover specific business processes, they do not fundamentally solve the problem of unstable recovery of complex PDF table structures.

[0025] Therefore, while existing technologies can achieve a certain degree of PDF table extraction and parsing in regular tables or some simple scenarios, they still generally have the following shortcomings when facing complex table scenarios such as no frame, weak boundaries, multiple rows spanning columns, background color interference, dense text, and continuous Chinese text: First, it is difficult to stably recover the implicit table skeleton, and the accuracy of table row and column boundary recognition is insufficient; Second, it is difficult to correctly distinguish between table structure signals and text internal layout signals, which easily leads to column division errors or text sticking; Third, it is difficult to accurately handle cross-row, cross-column, and multi-line continuous text, which easily causes mis-splitting or incorrect attribution of cell content; Fourth, it is difficult to balance structural recovery and semantic integrity under complex visual interference conditions, affecting the effectiveness of subsequent structured processing and business applications.

[0026] To address the aforementioned issues, this embodiment provides a method for restoring the layout structure of complex PDF tables. Under conditions such as no frame, background color interference, multiple rows spanning columns, and continuous text arrangement, this method comprehensively utilizes multiple types of visual and text signals to improve the accuracy and stability of table skeleton restoration, row and column boundary recognition, cell content attribution, and structured expression, thereby providing a reliable foundation for subsequent rule parsing, knowledge extraction, data governance, and business intelligent processing.

[0027] For example, such as Figure 7As shown, this embodiment provides a method for restoring the layout structure of complex PDF tables, including: Obtain the target table region in the target PDF document, extract the text objects in the target table region, and perform coordinate standardization on the text objects to obtain a standardized set of text objects under a unified image coordinate system. Based on the line signals and background color region signals in the target table area, the table reference structure is recovered; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; Based on the positional distribution, alignment, and proximity of a standardized set of text objects, the text structure is restored; wherein the text structure includes at least a text column structure and a visual row result; The table reference structure and the text structure are merged in the column direction and the row direction respectively to obtain the final column structure and the final row structure; Based on the final column structure and the final row structure, the text objects are reorganized into rows and columns and the layout is reconstructed to obtain a structured table result.

[0028] The method for restoring the layout structure of complex PDF tables provided in this embodiment will now be described in detail with reference to the accompanying drawings: For example, this embodiment provides a method for restoring the layout structure of complex PDF tables, and the specific inventive concept is as follows: PDF tables are essentially table objects presented in the form of page layouts. Their structural information is usually not directly given as explicit row, column, and cell data, but rather dispersed in boundary lines, background areas, white space distribution, text position, and their layout relationships. For humans, these page elements can be directly recognized as tables, and based on this, they can understand the row and column relationships, correspondences, and hierarchical relationships between different content. However, for programs, it is still necessary to further determine which visual separators correspond to table boundaries and which text layouts correspond to row and column structures and cell relationships. Therefore, PDF table structure restoration is essentially not a direct reading of the existing structure, but a process of analyzing, constraining, and reconstructing multiple types of table structure signals within the page.

[0029] The purpose of tables is not merely to form a geometric grid, but to organize and present content in a structured way. The border lines, background areas, white space, and text layout within tables are not isolated visual elements, but rather page-level techniques that collectively serve to express the organizational relationships of the content. In other words, the division of areas in a table is not an end in itself, but a spatial method used to express the hierarchical, corresponding, parallel, and continuous relationships between content objects. That is, by dividing page areas, tables provide different content objects with relatively stable spatial assignments, and further, through positional and alignment relationships, express the organizational structure between different content items.

[0030] Based on the above understanding, the structural restoration of complex PDF tables should not involve mixing various page signals all at once. Instead, it should be carried out in layers, from overall reference to local attribution, and then to layout reconstruction. Specifically, first, based on visual page separation signals such as boundary lines, background areas, and white space distribution, the reference structure of the table is restored to determine the approximate row and column division. Then, combining the positional relationships, alignment relationships, proximity relationships, and layout rules between text objects, the column and row structures on the text side are restored. Next, the table reference structure is merged with the text structure to obtain the final column and row structures. Finally, based on the final column and row structures, the text objects are assigned to rows and columns, and the layout reconstruction result of the table is further formed.

[0031] In the table reference structure restoration stage, the approximate row and column division of the table is mainly restored using boundary lines and background areas. The restoration result at this stage is not directly equivalent to the final cell structure, but mainly serves to provide range constraints and attribution references for subsequent text structure restoration. In the text object-based structure restoration stage, the text column structure and visual row results are restored mainly based on the positional distribution, alignment, and layout rules of the text objects. The text column structure mainly reflects the actual organization of text objects in the column direction, while the visual row mainly reflects the direct arrangement of text objects on the page. In the table reference structure and text structure fusion stage, the table reference structure and text structure are matched and merged in both the column and row directions to obtain the final column structure and final row structure.

[0032] In the column direction, the final column structure of the table is primarily reflected in the horizontal organization of text objects. Therefore, column direction restoration should mainly be based on the text column structure formed by the text objects; boundary lines and background areas are mainly used to define the parent column range and boundary constraints. In the row direction, text objects first form directly observable page layout rows. The "visual row" referred to in this article is a page row aggregated based on the vertical proximity of text objects on the page. Since multiple visual rows do not necessarily correspond to multiple independent structural units, row direction restoration is based on visual rows and further determines the final row structure by combining the parent row range.

[0033] After restoring the final column and row structures, this embodiment further reconstructs the table layout based on the row and column assignments of the text objects. By organizing and merging text objects according to their row and column positions, a cell-level layout result table and a table-based layout reconstruction result can be formed, thus providing structured input for subsequent data processing, retrieval, and business rule execution. It should be noted that the layout reconstruction result is mainly used to restore the structured representation of the table, and it is not required to complete all semantic merging or business logic corrections in this step. Processing such as cross-row continuation, inheritance filling, or field docking can be further completed based on the layout reconstruction result in subsequent business processing.

[0034] Therefore, the core of this method does not lie in mechanically overlaying lines, background colors, and text objects, but rather in first distinguishing the role of different types of signals in table restoration, then setting processing methods according to the structural characteristics of the column and row directions, and achieving stable layout structure restoration of complex PDF tables through layered restoration and step-by-step reconstruction. Based on this overall technical concept, this embodiment can adapt well to scenarios such as no frame, weak boundaries, background color interference, dense Chinese text, multiple rows spanning columns, and complex local attribution.

[0035] Based on the above inventive concept, such as Figure 1 As shown, this embodiment provides a method for restoring the layout structure of complex PDF tables. The specific steps are as follows: like Figure 2 As shown, step S1: Obtaining the target table area and standardizing text objects: To ensure that subsequent table structure analysis is conducted within a unified and stable local area, this step first locates the target page in the target PDF document and further determines the area where the target table is located. Based on this, table area cropping, image rendering, and text object standardization are completed to form the basic input data required for subsequent row and column space construction.

[0036] S11. Obtaining the target table area: First, obtain the target page from the target PDF document and determine the target table region corresponding to the table to be processed. The bounding box of this target table region can be obtained based on the table detection results on the page. Then, the target page can be partially cropped according to the bounding box to obtain a table sub-region containing only the content of the target table.

[0037] This process can eliminate irrelevant text, headers, footers, and other layout distractions from the original PDF pages, allowing subsequent processing to focus on the target table itself, thereby improving the stability and relevance of the structure restoration process.

[0038] S12. Generation of table area images: After cropping the table region, the sub-regions are further rendered to obtain a table region image corresponding to the target table region. The table region image retains visual information such as the table lines, background color, background blocks, and text layout, and provides a unified input for subsequent image-based structural skeleton detection.

[0039] In other words, after this step, the target table area in the original PDF is converted into a local image space that can be both visually analyzed and mapped to the coordinates of text objects.

[0040] S13. Extracting text objects within the cropping area: In order to preserve the text content near the table boundaries as completely as possible, this embodiment does not extract the text from the entire page first and then perform region cropping. Instead, it directly performs text extraction within the cropped sub-region of the table to obtain a set of text objects located within the target table region.

[0041] The advantage of this approach is that it reduces omissions, truncation, or attribution shifts caused by extracting text objects close to the table edge, column header edge, or title edge, thereby improving the integrity of text objects within the table area.

[0042] S14. Text object coordinate standardization: After obtaining the collection of text objects, coordinate unification processing is further performed on each text object. Specifically, the position of the text object in the PDF coordinate system is first converted into local coordinates relative to the upper left corner of the target table area. Then, based on the proportional relationship between the width and height of the target table area in the PDF coordinate system and the pixel width and height of the table area image, the local coordinates are mapped to the table area image coordinate system.

[0043] This process maps text objects, originally in the PDF page coordinate system, to a pixel coordinate system consistent with the table area image. This allows subsequent comparisons, matching, and fusion of both the structure space obtained from image detection and the text space constructed from text objects, all under the same coordinate reference.

[0044] S15. Boundary constraints and invalid object filtering for text objects: After completing the coordinate mapping, boundary constraints and validity checks are performed on the text objects first. Specifically, the mapped left and right boundaries and top and bottom boundaries are restricted to the range of the table area image to avoid slight boundary overflows caused by floating-point conversion errors; at the same time, abnormal text boxes with a width or height less than or equal to zero are filtered out.

[0045] This process ensures that all text objects entering the subsequent analysis process have a clear and effective spatial range, providing a reliable foundation for subsequent text column construction, visual row clustering, and anomaly attribution compensation.

[0046] S16. Construction of Standardized Text Object Table: Ultimately, a unified, standardized text object table is constructed. The text object table preferably includes at least one or more of the following fields: text object identifier, text content, left boundary, right boundary, top boundary, bottom boundary, width, height, center x-coordinate, center y-coordinate, and the corresponding original PDF coordinate information.

[0047] Furthermore, text objects can be sorted according to their vertical and horizontal positions on the page, and their identifiers can be renumbered to form a unified input for subsequent text sub-column space construction, text visual row construction, structural space mapping, and layout restoration processing.

[0048] like Figure 3 As shown, step S2: Restoring the table reference structure: After obtaining the target table area and standardizing the text objects, it is necessary to further recover the reference structure that can represent the overall partitioning and parent boundaries of the table from the table area image. This step does not directly determine the final column structure, row structure, and cell structure. Instead, it first constructs a reference space for the table in the column and row directions based on visual separation signals such as boundary lines and background areas on the page. This provides a basis for boundary constraints, hierarchical attachment, and anomaly compensation for the subsequent structural recovery of the text objects.

[0049] S21. Structure detection based on line signals: First, line detection is performed on the table region image to obtain vertical and horizontal boundary candidates. In this embodiment, the image is first converted to grayscale and thresholded to separate the lines and other dark foreground elements in the table from the background. Then, morphological operations suitable for vertical and horizontal structures are used to extract the vertical and horizontal lines, respectively, and a set of line boxes and their positions is obtained accordingly, which constitutes the candidate lines and forms the initial space in the column and row directions.

[0050] S22. Original line candidate structural screening: Not all raw lines detected in an image have a tabular structure; they may contain various forms of noise, such as short local lines and decorative lines. Therefore, candidate lines need to be screened. In practice, spatially adjacent line boundaries are first merged to form a relatively stable set of vertical and horizontal boundary positions. Then, based on the effective spacing between adjacent vertical and horizontal boundaries, minimum column width and minimum row height constraints are determined. Using these constraints as a basis, unstructured short lines that cannot cross the minimum structural unit are removed from the candidate set, and the remaining lines are restored to their column and row structural space.

[0051] S23. Extraction of structural information based on the background color region: Considering that the structural information in actual PDF tables is not only presented through boundary lines, but also through background areas, background partitions, and color bands, this step extracts structural information from the background areas. First, background reconstruction is performed on the table area to remove foreground color interference such as text, making the background color and background partitions more continuous. Then, clustering and segmentation are performed on the background image to divide it into several background areas, and the decomposition positions between different background areas are identified to form a color structure space in the column and row directions.

[0052] S24. Integration of line structure and color structure: After obtaining the line structure space and color structure space separately, fusion is performed in the column and row directions respectively. It should be noted that the line results and color results are not considered as a single master-slave relationship that replaces each other, but rather as two types of structural signals with different sources and representational focuses, participating together in the reference structure recovery. Among them, the line results focus more on expressing boundary positions and segmentation relationships, while the color results focus more on expressing region partitioning and hierarchical blocks.

[0053] During the fusion process, the spatial correspondence between the line structure space and the color structure space is determined first, including inclusion, overlap, and adjacency relationships. Based on this correspondence, it is then determined whether the corresponding structure should undergo joint confirmation, range segmentation, or independent retention. For line structure spaces and color structure spaces that are basically corresponding and can corroborate each other, they can be considered to point to the same table structure, and the corresponding reference boundaries or reference partitions can be confirmed accordingly. In cases where one line structure space corresponds to multiple color structure spaces, the parent reference range corresponding to the line structure space can be segmented based on the boundary results of each color structure space, thus forming multiple finer-grained reference spaces. For line structure spaces or color structure spaces that cannot form a stable correspondence, they are allowed to be retained as independent reference structures for output.

[0054] S25, Output of table reference structure results: After the above processing, this step finally outputs the table reference structure in both column and row directions. The table reference structure preferably includes at least one or more of the following attributes: spatial identifier, start and end boundaries, width or height, center position, source information, and other auxiliary attributes. It should be noted that the table reference structure is only used to represent the approximate row and column division of the table and does not directly correspond to the final column structure, final row structure, or final cell structure. Instead, it serves as the basis for subsequent text object attribution determination and table structure restoration.

[0055] like Figure 4 As shown, step S3: Restoring the table structure based on text objects: After obtaining the table reference structure in step S2, it is necessary to further restore the table structure that more closely approximates the actual reading result based on the positional distribution and layout rules of the text objects themselves. It should be noted that the table reference structure is mainly used to represent the approximate row and column division of the table. For example, in the column direction, it can be determined that a certain area roughly belongs to the same parent column range; in the row direction, it can be determined that several text objects roughly fall within the same parent row range. However, for finer-grained structural relationships, such as whether there are multiple child columns within a certain parent column range, or whether multiple adjacent rows of text belong to the same content unit or different content units, it is usually difficult to determine directly based solely on the table reference structure. Further restoration is still required by combining the alignment, proximity, and layout rules of the text objects. Therefore, this step mainly restores the text column structure in the column direction and the visual row result in the row direction based on the text objects, and uses these as inputs for subsequent column direction fusion and row direction fusion.

[0056] S31. Text column structure restoration: In many complex PDF tables, column structure is not entirely defined by boundary lines, but rather by the horizontal alignment and hierarchy of text objects. For example, within a large parent column, the header might only provide a main title, but the content below it naturally divides into three horizontal columns (left, center, and right), appearing as three stably aligned text columns. In this case, the true column structure is not directly drawn by boundary lines, but rather reflected in the arrangement of text objects. Therefore, the text column structure primarily reflects the actual organization of text content along the column direction and serves as a crucial basis for subsequent final column structure reconstruction.

[0057] Therefore, in this embodiment, the text column structure is restored based on the horizontal organization of text objects in the column direction. Specifically, the alignment relationships of the left, right, and center boundaries of the text objects, as well as their horizontal proximity, are comprehensively utilized to initially divide the text objects into column bands. Subsequently, based on the number of text objects within each column band, they are divided into main column bands and column bands to be processed. A redistribution operation is performed on the column bands to be processed, and a merging operation is performed on the main column bands that meet the conditions, thereby obtaining a stable text column structure.

[0058] S32. Visual line restoration of text: In the line direction, this embodiment first restores the visual lines. In this embodiment, a "visual line" refers to a page line formed by aggregating text objects based on their vertical proximity. Specifically, text objects can be clustered according to their upper and lower boundaries and center ordinates, following the principle of vertical proximity. Text objects located within the same horizontal reading band on the page are grouped into the same visual line. The visual line reflects the direct arrangement of text objects on the page and is the basic input for restoring the line direction structure.

[0059] It's important to note that visual rows are merely the direct arrangement of text objects on the page and do not directly equate to the final row structure. For example, in some tables, a descriptive sentence within the same cell might be split into two visual rows due to page width limitations; conversely, two visual rows arranged vertically within the same parent row might each correspond to two parallel sub-items. Therefore, visual adjacency alone is usually insufficient to directly determine the final row structure and requires further reconstruction in subsequent steps, incorporating a table reference structure.

[0060] S33. Summary of Text Structure: After the above processing, this step finally outputs the table structure result on the text side. The result includes: text column structure, visual rows, and one or more of the following: the attribution result from text objects to text column structure and from text objects to visual rows.

[0061] It should be noted that the text column structure is mainly used to express the actual organizational relationship of text objects in the column direction, while the visual row is the manifestation of the same organizational relationship in the row direction. The text-side structure result is not directly equivalent to the final table structure, but serves as an important input for subsequent column-direction fusion and row-direction fusion. In other words, step S2 solves the table skeleton structure, while step S3 further solves how text objects are arranged into columns and rows within these areas.

[0062] like Figure 5 As shown, step S4: Integrating the table reference structure with the text structure: After obtaining the table reference structure in step S2 and the text side layout structure in step S3, it is necessary to further combine the two to restore a table structure that more closely approximates the actual reading result. It should be noted that the table reference structure is mainly used to provide the approximate row and column division of the table, while the text structure mainly reflects the actual organization of text objects in the column and row directions. Therefore, this step does not simply superimpose the two types of results, but rather performs joint processing according to the structural characteristics of each column and row direction to obtain the final column structure and the final row structure. In the current implementation, this step mainly corresponds to the column direction fusion and row direction fusion processes.

[0063] S41, Column Direction Merging: In the column direction, the column direction reference structure in the table reference structure is first merged with the text column structure recovered in step S3. Specifically, the column direction reference structure is first used as the parent column range, and then the text column structure is used as the text side column result to be matched. Further, the correspondence between the text column structure and the parent column range is determined based on the boundary differences, overlap ratio, inclusion relationship and center position relationship between the two.

[0064] In this process, if a text column structure is essentially identical to a parent column range, it can be considered a direct correspondence within that parent column range. If a text column structure falls entirely within a parent column range, it can be considered a child column under that parent column range. If a text column structure spans multiple parent column ranges or is outside of parent column ranges, it indicates that the text column structure cannot be simply categorized into a single parent column range and needs to be treated as an independent structure. Through the above processing, the hierarchical relationship between parent column ranges and text column structures can be gradually established.

[0065] Furthermore, for text objects lacking a column affiliation, affiliation compensation can be performed based on their positional relationship and overlap with the parent column range. If only one final column structure exists within a given parent column range, the boundaries of that final column structure can be corrected using the parent column range to maintain consistency in the column direction structure. After these processes, the final column structure in the column direction is obtained.

[0066] S42, Line Direction Fusion: Line-wise merging is more prone to ambiguity than column-wise merging. For example, two visually arranged lines may belong to the same content unit as cross-line continuation text, or they may correspond to two parallel sub-items within the same parent line range. Therefore, visual adjacency is not directly equivalent to the final line structure and still needs to be further determined in conjunction with the parent line range.

[0067] Based on this, this step uses the row direction table reference structure obtained in step S2 as the parent row range and the visual row result obtained in step S3 as the text side row result to be assigned. The correspondence between the visual row and the parent row range is determined based on the boundary differences, overlap ratio, inclusion relationship, and center position relationship between the two. If a visual row is basically consistent with a parent row range, or if the two have a high overlap ratio, then the visual row can be considered to directly correspond to that parent row range. If a visual row completely falls within the parent row range, then the visual row can be considered to belong to that parent row range. If a visual row spans multiple parent row ranges, or is located outside of a parent row range, then it is not directly assigned to a single parent row range, but is retained as an independent row structure. After completing the above correspondence determination, the final row structure is generated in units of parent row ranges. After the above processing, the final row structure of the row direction is finally obtained.

[0068] S43. Summary of Fusion Results: After column-wise and row-wise merging, this step yields the final column and row structures of the table. The final column structure represents the final allocation of text objects along the column direction, and the final row structure represents the final allocation of text objects along the row direction. It's important to note that the final column and row structures are not directly equivalent to the cell-level results, but rather serve as the direct basis for subsequent cell reconstruction. Through the above processing, a unified column and row structure foundation can be provided for subsequent text object-level layout table construction and cell-level layout reconstruction.

[0069] like Figure 6 As shown, step S5: Reconstructing the table layout based on row and column affiliation: After obtaining the final column and row structures in step S4, the table layout is further reconstructed based on the assignment of text objects in the row and column directions. It should be noted that step S4 addresses the question of "which columns and rows can serve as the final structure," while this step further addresses the question of "which row and column each text object ultimately belongs to," thereby forming a structured layout result suitable for subsequent data processing.

[0070] Specifically, in this step, the column direction attribution results are first uniformly organized, and the row direction attribution results are merged with the column direction attribution results according to the granularity of the text object, so that each text object can obtain its attribution information in both the row and column directions, thereby forming a unified text object-level attribution result.

[0071] Building upon this, text objects that possess both row and column orientation are grouped according to their row and column identifiers. The text objects within the same group are then organized and concatenated according to their page position order to obtain the corresponding cell text, thus forming a cell-level layout result table. Further, the cell-level layout result table is expanded according to row and column positions to generate a table-style layout reconstruction result, outputting a structured table result.

[0072] It should be noted that the layout reconstruction result is mainly used to restore the structured representation of the table, facilitating subsequent data processing, retrieval, and business rule execution, and does not require the completion of all semantic merging or business logic correction in this step. After the above processing, this step finally yields the layout reconstruction result. The result preferably includes at least one or more of the following: text object-level attribution results, cell-level layout result table, tabular layout reconstruction result, and results of the unit to be inspected and its corresponding context.

[0073] Furthermore, it should be noted that although there are existing methods for restoring the layout structure of PDF tables, most of them are based on the technical approach of "table detection and cell logic restoration". The core of the restoration method provided in this embodiment is to first restore the table reference structure, then restore the text structure, and determine the final structure by fusing the column and row directions respectively. This belongs to the technical approach of "complex PDF table layout structure restoration", rather than the traditional technique of cutting cells first and then filling in the content. It can be seen that this restoration method is particularly suitable for scenarios with weak boundaries, no frame, and complex Chinese PDF tables.

[0074] like Figure 8 As shown, exemplarily, this embodiment also provides a layout structure restoration system for complex PDF tables, including: The acquisition module is used to acquire the target table area in the target PDF document, extract the text objects in the target table area, and perform coordinate standardization processing on the text objects to obtain a set of standardized text objects under a unified image coordinate system. The first recovery module is used to recover the table reference structure based on the line signals and background color area signals in the target table area; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; The second recovery module is used to recover the text structure based on the positional distribution, alignment, and proximity relationships of a standardized set of text objects; wherein the text structure includes at least a text column structure and a visual row result; The fusion module is used to merge the table reference structure and the text structure in the column direction and the row direction, respectively, to obtain the final column structure and the final row structure. The sorting and reconstruction module is used to sort and reconstruct the layout of text objects based on the final column structure and the final row structure, resulting in a structured table.

[0075] The present invention also provides a layout structure restoration device for complex PDF tables, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the layout structure restoration method for complex PDF tables.

[0076] The present invention also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the layout structure restoration method for complex PDF tables.

[0077] When the processor executes the computer program, it implements the above-described steps for restoring the layout structure of complex PDF tables, for example: Obtain the target table region in the target PDF document, extract the text objects in the target table region, and perform coordinate standardization on the text objects to obtain a standardized set of text objects under a unified image coordinate system. Based on the line signals and background color region signals in the target table area, the table reference structure is recovered; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; Based on the positional distribution, alignment, and proximity of a standardized set of text objects, the text structure is restored; wherein the text structure includes at least a text column structure and a visual row result; The table reference structure and the text structure are merged in the column direction and the row direction respectively to obtain the final column structure and the final row structure; Based on the final column structure and the final row structure, the text objects are reorganized into rows and columns and the layout is reconstructed to obtain a structured table result.

[0078] For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing preset functions, the instruction segments describing the execution process of the computer program in the layout structure restoration device for complex PDF tables. For example, the computer program can be divided into an acquisition module, a first recovery module, a second recovery module, a fusion module, and a sorting and reconstruction module; the specific functions are as follows: the acquisition module is used to acquire the target table area in the target PDF document, extract the text objects in the target table area, and perform coordinate standardization processing on the text objects to obtain a standardized set of text objects under a unified image coordinate system; the first recovery module is used to recover the table reference structure based on the line signals and background color area signals in the target table area; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; the second recovery module is used to recover the text structure based on the position distribution, alignment relationship, and proximity relationship of the standardized text object set; wherein, the text structure includes at least the text column structure and the visual row result; the fusion module is used to fuse the table reference structure and the text structure in the column direction and the row direction respectively to obtain the final column structure and the final row structure; the sorting and reconstruction module is used to sort the row and column assignments of the text objects and reconstruct the layout based on the final column structure and the final row structure to obtain the structured table result.

[0079] The layout structure recovery device for complex PDF tables can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. This device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above are examples of layout structure recovery devices for complex PDF tables and do not constitute a limitation on such devices. They may include more components than described above, or combine certain components, or use different components. For example, the layout structure recovery device for complex PDF tables may also include input / output devices, network access devices, buses, etc.

[0080] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or any conventional processor. The processor is the control center for the layout structure restoration of complex PDF tables, connecting all parts of the layout structure restoration device for complex PDF tables via various interfaces and lines.

[0081] The memory can be used to store the computer program and / or modules. The processor implements various functions of the layout structure restoration device for complex PDF tables by running or executing the computer program and / or modules stored in the memory and calling the data stored in the memory.

[0082] The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0083] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method for restoring the layout structure of complex PDF tables.

[0084] If the modules / units integrated in the layout structure restoration system for complex PDF tables are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0085] Based on this understanding, the present invention can implement all or part of the processes in the above-described method for restoring the layout structure of complex PDF tables, or it can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described method for restoring the layout structure of complex PDF tables. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or preset intermediate forms, etc.

[0086] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0087] It should be noted that the content contained in the computer-readable storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0088] In summary, compared with existing methods for extracting and parsing PDF tables, this invention provides a method for restoring the layout structure of complex PDF tables, which has the following advantages: First, this invention can stably restore the table structure in complex PDF table scenarios such as no frame, weak boundaries, background color interference, dense Chinese text, and multi-row spanning columns.

[0089] Secondly, this invention comprehensively utilizes multiple table structure signals such as boundary lines, background color areas, text positions, and their layout relationships, avoiding the limitations of single line methods or single text methods in complex scenarios and improving the accuracy of table structure recovery.

[0090] Third, the present invention adopts different restoration strategies for the structural characteristics of the column direction and the row direction, so that the column direction can more accurately reflect the horizontal organization of text objects, and the row direction can more robustly restore the correspondence between the visual row and the final row structure.

[0091] Fourth, after the structure is restored, the present invention can further generate word-level attribution results, cell-level layout result tables, and table-based layout reconstruction results, providing a unified structured input for subsequent data processing, retrieval, and business rule execution.

[0092] The above embodiments are merely one of the implementation methods for achieving the technical solution of the present invention. The scope of protection claimed by the present invention is not limited to this embodiment, but also includes any variations, substitutions and other implementation methods that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for restoring the layout structure of complex PDF tables, characterized in that, include: Obtain the target table region in the target PDF document, extract the text objects in the target table region, and perform coordinate standardization on the text objects to obtain a standardized set of text objects under a unified image coordinate system. Based on the line signals and background color region signals in the target table area, the table reference structure is recovered; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; Based on the positional distribution, alignment, and proximity of a standardized set of text objects, the text structure is restored; wherein the text structure includes at least a text column structure and a visual row result; The table reference structure and the text structure are merged in the column direction and the row direction respectively to obtain the final column structure and the final row structure; Based on the final column structure and the final row structure, the text objects are reorganized into rows and columns and the layout is reconstructed to obtain a structured table result.

2. A method for restoring the layout structure of complex PDF tables according to claim 1, characterized in that, The process of obtaining the target table region in the target PDF document, extracting the text objects from the target table region, and performing coordinate standardization on the text objects to obtain a standardized set of text objects in a unified image coordinate system includes: Obtain the target page in the target PDF document, determine the bounding box of the target table area corresponding to the table to be processed based on the page table detection results, and perform local cropping on the target page based on the bounding box to obtain a table sub-area containing only the content of the target table. The table sub-regions are rendered to generate table region images that retain the table's visual information and are used for subsequent structural skeleton detection; Extract text objects from a sub-region of the table to obtain a collection of text objects within the target table region; The position of the text object in the PDF coordinate system is converted into local coordinates relative to the top left corner of the target table area. Based on the ratio between the width and height of the target table area in the PDF coordinate system and the pixel width and height of the table area image, the local coordinates are mapped to the table area image coordinate system, thus completing the text object coordinate standardization. Boundary constraints are applied to the coordinate-standardized text objects to limit their left, right, top, and bottom boundaries to the range of the table area image, and abnormal text objects with width or height less than or equal to zero are filtered out, resulting in text objects that have undergone boundary constraints and filtering. Based on the text objects after boundary constraints and filtering, a standardized text object table is constructed, which includes at least text object identifiers, text content, coordinate information and size information, and finally a standardized text object set under a unified image coordinate system is obtained.

3. A method for restoring the layout structure of complex PDF tables according to claim 1, characterized in that, The process of restoring the table reference structure based on line signals and background color region signals in the target table area includes: The image of the table region corresponding to the target table region is grayscaled and thresholded to separate the lines from the background. Vertical and horizontal lines are extracted through targeted morphological operations to obtain vertical and horizontal boundary candidates and corresponding line boxes and position sets, forming the initial space in the column and row directions. Merge adjacent line boundaries to form a stable set of vertical and horizontal boundary positions. Determine the minimum column width constraint and minimum row height constraint based on the effective spacing between adjacent boundaries. Eliminate unstructured short lines that cannot cross the minimum structural unit and retain and form line structure space in the column and row directions. Background reconstruction is performed on the table area image to remove text foreground interference, making the background color and background partition more continuous. Then, the background image is clustered and segmented to divide it into several background color regions. The decomposition positions between different background color regions are identified to form a color structure space in the column and row directions. The line structure space and the color structure space are merged in the column direction and the row direction respectively. The spatial correspondence between the two is determined. Based on the correspondence, the two types of structures are jointly confirmed, range segmented or independently preserved to form a merged structure space. Based on the fused structural space, output a column-direction and row-direction table reference structure that includes at least spatial identifiers, start and end boundaries, size information, and source information, and complete the restoration of the table reference structure.

4. A method for restoring the layout structure of complex PDF tables according to claim 1, characterized in that, The method of restoring text structure based on the positional distribution, alignment, and proximity relationships of a standardized set of text objects includes: For a standardized set of text objects, the alignment relationship of the left and right boundaries and center coordinates of the text objects and the horizontal proximity relationship are comprehensively utilized in the column direction to divide the text objects into initial column bands. The main column band and the column band to be processed are distinguished according to the number of text objects in each initial column band. The column band to be processed is assigned and the main column band that meets the conditions is merged to obtain the text column structure. In the line direction, based on the upper boundary, lower boundary, and center ordinate of the text object, the text objects are clustered according to the principle of vertical proximity, and the text objects located in the same horizontal reading band are merged into the same visual line to complete the restoration of the text visual line. The system summarizes the results of the text column structure, text visual rows, and the attribution of text objects to the text column structure and text visual rows, outputs the table structure results on the text side, and finally completes the restoration of the text structure.

5. A method for restoring the layout structure of complex PDF tables according to claim 1, characterized in that, The process of merging the table reference structure and the text structure in the column and row directions respectively to obtain the final column structure and the final row structure includes: The column direction reference structure in the table reference structure is used as the parent column range, and the text column structure is used as the side column result of the text to be matched. Based on the boundary difference, overlap ratio, inclusion relationship and center position relationship between the two, the correspondence between the text column structure and the parent column range is determined. For the text objects to which the missing text column belongs, the attribution compensation is performed according to their position and overlap with the parent column range. The boundary correction is performed on the parent column range corresponding to the single final column structure to obtain the final column structure of the column direction. The row direction reference structure in the table reference structure is used as the parent row range, and the text visual row is used as the text side row result to be assigned. Based on the boundary difference, overlap ratio, inclusion relationship and center position relationship between the two, the correspondence between the visual row and the parent row range is determined, and the final row structure of the row direction is generated with the parent row range as the unit. The final column structure obtained by merging the columns and the final row structure obtained by merging the rows are combined to obtain the final column structure and final row structure of the table.

6. A method for restoring the layout structure of complex PDF tables according to claim 1, characterized in that, The process of organizing and reconstructing the layout of text objects based on the final column and row structures to obtain structured table results includes: Based on the final column structure and the final row structure, and combined with the standardized set of text objects, the column direction attribution results of text objects are uniformly organized, and the row direction attribution results are merged with the column direction attribution results according to the granularity of text objects, so that each text object can obtain the attribution information of both row direction and column direction, forming a unified text object-level attribution result. Text objects that have both row and column orientations are grouped according to their row and column identifiers. The text objects within the same group are then arranged and spliced ​​according to their page position order to obtain the text content of the corresponding cells, forming a cell-level layout result table. Expand the cell-level layout results table by row and column to generate a table-style layout reconstruction result, thus obtaining a structured table result.

7. A method for restoring the layout structure of complex PDF tables according to claim 1, characterized in that, The structured table results include one or more of the following: text object-level attribution results, cell-level layout results, table-style layout reconstruction results, and results of the unit to be inspected and its corresponding context.

8. A layout structure restoration system for complex PDF tables, characterized in that, include: The acquisition module is used to acquire the target table area in the target PDF document, extract the text objects in the target table area, and perform coordinate standardization processing on the text objects to obtain a set of standardized text objects under a unified image coordinate system. The first recovery module is used to recover the table reference structure based on the line signals and background color area signals in the target table area; wherein, the table reference structure is used to represent the division range of the table in the column direction and / or row direction; The second recovery module is used to recover the text structure based on the positional distribution, alignment, and proximity relationships of a standardized set of text objects; wherein the text structure includes at least a text column structure and a visual row result; The fusion module is used to merge the table reference structure and the text structure in the column direction and the row direction, respectively, to obtain the final column structure and the final row structure. The sorting and reconstruction module is used to sort and reconstruct the layout of text objects based on the final column structure and the final row structure, resulting in a structured table.

9. A layout structure restoration device for complex PDF tables, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the layout structure restoration method for complex PDF tables as described in any one of claims 1-7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it is used to implement the layout structure restoration method for complex PDF tables as described in any one of claims 1-7.