Document layout reconstruction method, device and system and storage medium
Through deep learning and OCR technology, combined with instance segmentation and logical analysis, the lack of logical structure in document layout reconstruction is solved, and efficient and accurate document layout reconstruction and directory restoration are achieved.
Patent Information
- Application Number
- CN202411872346.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art ignores the logical structure in document layout analysis, making it difficult to restore paper documents to electronic document structures that conform to human reading habits.
Deep learning instance segmentation algorithm and OCR technology are used, combined with logical layout analysis, and end-to-end document layout reconstruction, detect and identify multiple layout elements, predict title levels, and reconstruct directory structure.
It realizes efficient and accurate document layout reconstruction, supports rapid analysis and high-precision recognition of various layout documents, and outputs document structures that conform to human reading habits.
Smart Images

Figure CN119962479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and natural language processing, and in particular to a document layout reconstruction method, device, system and storage medium. Background Art
[0002] With the development of computer technology, electronic documents are gradually replacing paper documents. As a common information carrier, documents are stored in electronic format, which is easy to store, carry, backup and share, and can significantly improve the efficiency of daily work and study. Therefore, document intelligence has gradually become a hot research topic, and document intelligence technology is also constantly developing and improving.
[0003] In OCR (optical character recognition) related research, document images taken from paper documents need to be analyzed and understood in layout, and decoded and reconstructed from multiple dimensions such as text information, table information, and image information before they can be converted into electronic documents. At present, research related to layout analysis is often limited to the detection and recognition of layout elements, that is, classifying and locating elements such as text, graphics, and tables, and then extracting element content using text recognition and other technologies. However, the limitation of the above method is that it ignores the logical structure of the document. For different document layouts, it is also necessary to consider providing more fine-grained semantic classification, predicting each title level, and finally restoring it to a structure that conforms to human reading habits. Summary of the invention
[0004] To this end, the technical solution of the present invention proposes a document layout reconstruction method, device, system and storage medium, which solves the extraction and restoration of physical layout elements and logical layout structures end-to-end, thereby realizing the reconstruction of document layout with directory structure. The method combines algorithms such as image processing, machine learning and deep learning, and realizes efficient recognition and reconstruction of document structure by understanding and extracting the logical structure and physical layout of the document, providing strong support for automated document processing and information extraction, and has the characteristics of generality, efficiency and high precision.
[0005] According to a first aspect of the technical solution of the present invention, a document layout reconstruction method is provided, comprising:
[0006] S1, document input and paging step: inputting an original document, and converting the original document into document images by paging;
[0007] S2, physical layout analysis step: based on the deep learning instance segmentation algorithm, locate and classify each layout element area in the document image to obtain the category and position of each layout element;
[0008] S3, input type determination step: determining whether the corresponding page of the document image can be directly parsed by code and whether it does not contain a table, if so, performing code parsing to obtain text-related information; otherwise, using OCR recognition to obtain text-related information;
[0009] S4, logical layout analysis step: according to the position of each layout element, the category of each layout element is matched with the text-related information, each layout element is sorted, and then hierarchical information is added to the layout elements whose categories are document titles and hierarchical titles, thereby realizing the reconstruction of the document layout with a directory structure.
[0010] Here, the document layout reconstruction with a directory structure means the document obtained after the document layout reconstruction, which realizes the detection and recognition of layout elements, that is, classifying and locating elements such as titles, paragraphs, graphics, tables, etc., and then extracting the element content by combining text recognition and other technologies, while considering more fine-grained semantic classification, predicting each title level, and finally restoring it to a structure that conforms to human reading habits.
[0011] Furthermore, in S1, the original document is a PDF document.
[0012] Furthermore, in S2, the instance segmentation algorithm based on deep learning is a yolov8-seg instance segmentation algorithm based on deep learning.
[0013] Furthermore, in S2, the categories of the layout elements include document title, table of contents, hierarchical title, paragraph, information block, table, picture, header, footer, page number, signature, seal, chart annotation, chart title, formula, and column.
[0014] Furthermore, the S2 specifically includes:
[0015] The yolov8-seg instance segmentation algorithm based on deep learning takes the document image as input, returns the category and corresponding mask area of each layout element in the document image, thereby obtaining the category and position of each layout element.
[0016] Furthermore, the S2 also includes: fitting a minimum circumscribed rectangle to the contour points of the mask area.
[0017] Furthermore, in S3, the OCR recognition specifically includes: table detection, table line segmentation and / or text line positioning, text line direction classification, text line font classification and text line recognition.
[0018] Furthermore, in S4, adding hierarchical information to layout elements of the document title and hierarchical title category specifically includes:
[0019] The layout elements of ordered document titles and hierarchical titles are traversed from back to front. Each hierarchical title and all its predecessor titles form a title pair from back to front. The pair is input into the deep learning network model based on the transformer architecture for prediction, and the output is a predefined prediction relationship.
[0020] Take the first title that is output as a subordinate relationship to form a parent-child pair in the title tree, and so on to find the parent title for each title element;
[0021] After traversing all parent-child pairs, since the root node is the only element in all parent-child pairs that has only child nodes but no parent nodes, the root node of the title tree is determined, which is the document title.
[0022] Thus, the title tree is constructed, and the depth of the title tree directly corresponds to the hierarchical information of the title.
[0023] Furthermore, the predefined prediction relationships include subordinate relationships, inclusion relationships, parallel relationships, and other relationships.
[0024] Furthermore, the relationship between the title pairs is distinguished according to the marked title levels. A level difference of 0 corresponds to a parallel relationship, a difference of 1 corresponds to an inclusion relationship, a difference of -1 corresponds to a subordinate relationship, and other differences correspond to other relationships.
[0025] According to a second aspect of the technical solution of the present invention, a document layout reconstruction device is provided, wherein the document layout reconstruction device operates based on the document layout reconstruction method according to any one of the above aspects, and comprises:
[0026] A document input paging unit, used for inputting an original document and converting the original document into document images by paging;
[0027] A physical layout analysis unit, used for locating and classifying each layout element region in the document image based on a deep learning instance segmentation algorithm, and obtaining a category and position of each layout element;
[0028] An input type determination unit is used to determine whether the corresponding page of the document image can be directly parsed by code and whether it does not contain a table. If so, code parsing is performed to obtain text-related information; otherwise, OCR recognition is used to obtain text-related information;
[0029] The logical layout analysis unit is used to match the category of each layout element with the text-related information according to the position of each layout element, sort each layout element, and then add hierarchical information to the layout elements whose categories are document titles and hierarchical titles, thereby realizing the reconstruction of the document layout with a directory structure.
[0030] According to a third aspect of the technical solution of the present invention, a document layout reconstruction system is provided, the system comprising: a processor and a memory for storing executable instructions; wherein the processor is configured to execute the executable instructions to execute the document layout reconstruction method as described in any of the above aspects.
[0031] According to a fourth aspect of the technical solution of the present invention, there is provided a computer-readable storage medium, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the document layout reconstruction method as described in any of the above aspects is implemented.
[0032] Beneficial effects of the present invention:
[0033] 1. More comprehensive layout element analysis: supports multiple fine categories of layout element detection, covering single-column, multi-column and other complex layout documents.
[0034] 2. Take into account both speed and accuracy: Support fast parsing of electronic text and high-precision OCR recognition of scanned files, solving the problem of layout element content recognition end-to-end.
[0035] 3. More universal sorting method: Optimize the sorting method of layout elements in complex multi-column layouts such as newspapers, papers, and magazines, and output them in a way that is more in line with human reading habits.
[0036] 4. Support restoring document directory: predict each title level through the logical layout analysis module and restore the document directory structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0038] Figure 1 A flow chart of a document layout reconstruction method according to an embodiment of the technical solution of the present invention is shown.
[0039] Figure 2 A schematic diagram showing information blocks in a layout element according to an embodiment of the technical solution of the present invention.
[0040] Figure 3 A schematic diagram showing a column (double column) in a layout element according to an embodiment of the technical solution of the present invention.
[0041] Figure 4 A logical layout analysis model architecture diagram of an embodiment of the technical solution of the present invention is shown.
[0042] Figure 5A schematic diagram showing word segmentation results of an embodiment of the technical solution of the present invention is shown.
[0043] Figure 6 A schematic diagram showing the sorting results of an embodiment of the technical solution of the present invention is shown.
[0044] Figure 7 A schematic diagram showing the title level prediction results of an embodiment of the technical solution of the present invention is shown.
[0045] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0046] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0047] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein, for example.
[0048] In addition, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.
[0049] Multiple includes two or more.
[0050] And / or, it should be understood that the term "and / or" used in this disclosure is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. It reduces manpower input and helps business automation, with the characteristics of universality, efficiency, and high precision.
[0051] Document layout reconstruction technology is a technology used to analyze and restore the original layout and structure of a document. It is mainly used to convert scanned paper documents, PDF files, pictures and other digital documents into editable electronic documents while retaining the layout and format of the original document as much as possible. It usually requires a combination of different technologies and algorithms to achieve comprehensive layout reconstruction.
[0052] The technical solution of the present invention first provides a document layout reconstruction method, comprising:
[0053] S1. Document input and paging step: input an original document and convert the original document into document images by paging.
[0054] In a preferred embodiment, in S1, the original document is a PDF document.
[0055] S2. Physical layout analysis step: Based on a deep learning instance segmentation algorithm, locate and classify the layout element areas in the document image to obtain the category and position of each layout element.
[0056] In a preferred embodiment, in S2, the deep learning-based instance segmentation algorithm is a deep learning-based yolov8-seg instance segmentation algorithm.
[0057] In a preferred embodiment, in S2, the categories of layout elements include document title, table of contents, level title, paragraph, information block, table, picture, header, footer, page number, signature, stamp, chart annotation, chart title, formula, and column.
[0058] In a preferred embodiment, S2 specifically includes:
[0059] The yolov8-seg instance segmentation algorithm based on deep learning takes the document image as input and returns the category of each layout element in the document image and the corresponding mask area.
[0060] In a preferred embodiment, the S2 further includes: fitting a minimum circumscribed rectangle to the contour points of the mask area.
[0061] S3, input type determination step: determine whether the corresponding page of the document image can be directly parsed by code and whether it does not contain a table. If so, perform code parsing to obtain text-related information; otherwise, use OCR recognition to obtain text-related information.
[0062] In a preferred embodiment, in S3, the OCR recognition specifically includes: table detection, table line segmentation and / or text line positioning, text line direction classification, text line font classification and text line recognition.
[0063] S4, logical layout analysis step: according to the position of each layout element, the category of each layout element is matched with the text-related information, each layout element is sorted, and then hierarchical information is added to the layout elements whose categories are document titles and hierarchical titles, thereby realizing the reconstruction of the document layout with a directory structure.
[0064] In a preferred embodiment, in S4, adding hierarchical information to layout elements of the document title and hierarchical title category specifically includes:
[0065] The layout elements of ordered document titles and hierarchical titles are traversed from back to front. Each hierarchical title and all its predecessor titles form a title pair from back to front. The pair is input into the deep learning network model based on the transformer architecture for prediction, and the output is a predefined prediction relationship.
[0066] Take the first title that is output as a subordinate relationship to form a parent-child pair in the title tree, and so on to find the parent title for each title element;
[0067] After traversing all parent-child pairs, since the root node is the only element in all parent-child pairs that has only child nodes but no parent nodes, the root node of the title tree is determined, which is the document title.
[0068] Thus, the title tree is constructed, and the depth of the title tree directly corresponds to the hierarchical information of the title.
[0069] In a preferred embodiment, the predefined prediction relationships include subordinate relationships, inclusion relationships, parallel relationships, and other relationships.
[0070] In a preferred embodiment, the relationship between title pairs is distinguished according to the marked title levels. A level difference of 0 corresponds to a parallel relationship, a difference of 1 corresponds to an inclusion relationship, a difference of -1 corresponds to a subordinate relationship, and other differences correspond to other relationships.
[0071] The technical solution of the present invention further provides a document layout reconstruction device, which operates based on the document layout reconstruction method described above, and includes:
[0072] A document input paging unit, used for inputting an original document and converting the original document into document images by paging;
[0073] A physical layout analysis unit, used for locating and classifying each layout element region in the document image based on a deep learning instance segmentation algorithm, and obtaining a category and position of each layout element;
[0074] An input type determination unit is used to determine whether the document image can be directly parsed and whether it does not contain a table. If so, code parsing is performed to obtain text-related information; otherwise, OCR recognition is used to obtain text-related information;
[0075] The logical layout analysis unit is used to match the category of each layout element with the text-related information according to the position of each layout element, sort each layout element, and then add hierarchical information to the layout elements whose categories are document titles and hierarchical titles, thereby realizing the reconstruction of the document layout with a directory structure.
[0076] The technical solution of the present invention further provides a document layout reconstruction system, which includes: a processor and a memory for storing executable instructions; wherein the processor is configured to execute the executable instructions to execute the document layout reconstruction method as described above.
[0077] The technical solution of the present invention further provides a computer-readable storage medium, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the document layout reconstruction method as described above is implemented.
[0078] Example
[0079] The whole process of document layout reconstruction in this embodiment is as follows Figure 1 , each step is introduced:
[0080] enter:
[0081] The input is a PDF document. First, try to parse each page of the document and save whether it can be parsed. Then convert the document pages into images.
[0082] Physical layout analysis:
[0083] The input is the converted document image. The instance segmentation algorithm based on deep learning is used to locate the layout element areas in the image and classify each element. The details of the physical layout analysis algorithm are as follows:
[0084] Definition of layout elements: The function of the physical layout analysis module is to classify and locate meaningful layout elements in document images. The predefined common layout element categories include document title, directory, level title, paragraph, information block, table, picture, header, footer, page number, signature, seal, chart annotation, chart title, formula, column, a total of 16 categories. Among them, "information block" and "column" are not common layout elements. The definition of information block is key-value pair information arranged in multiple lines, which is often seen in documents with layout formats such as contracts and legal documents (such as Figure 2 ), a column is defined as multiple parallel document blocks in the same horizontal area (such as Figure 3), the reading order between columns follows from left to right, and the reading order within a column follows from top to bottom. With the column restriction, documents in double-column or multi-column layouts can be sorted in accordance with human reading habits.
[0085] Physical layout analysis algorithm: The physical layout analysis module uses the yolov8-seg instance segmentation algorithm based on deep learning. The model takes the document image as input and returns the category of each layout element in the image and the corresponding mask area. Furthermore, the minimum circumscribed rectangle can be fitted to the contour points of the mask. The instance segmentation algorithm is chosen here instead of the target detection algorithm because the instance segmentation algorithm can return a more accurate target area when processing partially tilted and text-intensive document images. In addition, compared with other instance segmentation algorithms such as Mask-RCNN, the yolov8-seg algorithm greatly improves the inference speed while maintaining high accuracy, and is the mainstream algorithm in the current industrial computer vision field.
[0086] Input type judgment:
[0087] This step determines whether the current PDF page can be directly parsed by code. In addition, if it contains a table, the table structure needs to be retained. If the input is a parsable PDF and does not contain a table, the PDF text can be parsed directly. Otherwise, the OCR technology needs to be combined to extract text and tables from the image (i.e. the OCR recognition module in the figure).
[0088] OCR recognition:
[0089] This step recognizes text lines and table information for unparseable documents or document images, which involves table detection, table line segmentation, text line positioning, text line direction classification, text line font classification, and text line recognition algorithm. Since the OCR technology based on deep learning is relatively mature, this article will not explain it in detail.
[0090] Logical layout analysis:
[0091] The input is the parsing or recognition result of the entire document. The main task of the logical layout analysis module is to distinguish the levels of all titles and reconstruct the document tree structure. Combining the previous steps, we can obtain the position and category information of each layout element, the position and text information of the text line, the table structure and the cell text information. Therefore, each layout element can be matched with the corresponding text line based on the position information. Furthermore, each layout element can be sorted according to human reading habits. In order to adapt to the possible appearance of complex layout documents such as double columns and multiple columns, the technical solution of this application uses the layout element of the column in the physical layout analysis. First, the columns and elements outside the columns are sorted, and then the elements in each column are sorted, and finally the column elements are eliminated. The xy_cut algorithm is used here to sort the layout elements, and the sorting effect is as follows. Figure 6As shown in the figure. xy_cut is a top-down layout segmentation method that recursively performs horizontal and vertical projections to segment the page into a series of relatively independent rectangular areas according to the gaps. After obtaining an ordered list of layout elements, the document directory structure can be restored by simply adding hierarchical information to the elements with the categories of document title and hierarchical title. Therefore, the logical layout analysis module predicts the hierarchy of each title based on the model of the transformer architecture, thereby reconstructing the document directory structure.
[0092] Specifically, the role of the logical layout analysis module is to further restore the logical structure of the document. After the previous steps, an ordered list of layout elements has been obtained. In this module, only the elements of the document title and the level title need to be extracted as input.
[0093] Since it is difficult to accurately distinguish the levels of each title element only by visual features, especially in general scenarios compatible with various layouts. The technical solution of the present invention explores a deep learning algorithm based on the transformer architecture, combining text semantic features, coordinate box information (position, width and height) and other different dimensional information as model input. The following is a detailed introduction:
[0094] Data preprocessing
[0095] Considering the significant differences between different document formats, it is too difficult to directly predict the title level. This algorithm further transforms the title level prediction task into a relationship classification task between titles. The predefined relationships include subordinate relationships, inclusion relationships, parallel relationships, and other relationships. The relationship between title pairs is distinguished according to the marked title levels. The level difference of 0 corresponds to a parallel relationship, the difference of 1 corresponds to an inclusion relationship, the difference of -1 corresponds to a subordinate relationship, and the difference of others corresponds to other relationships.
[0096] Model Architecture
[0097] The overall structure of the logical layout analysis model is as follows: Figure 4 As shown:
[0098] The model takes the title information pair as input and the title pair relationship as output. It is essentially a classification model. The title information pair contains the input of different dimensions mentioned above (text, two-dimensional coordinates). The text input is converted into corresponding tokens by the word segmentation model and then aligned with the features of each dimension (such as Figure 5 ) are sent to the embedding layer together to obtain the fused embedded features, and then the corresponding classification results are obtained through the feature encoder and the relationship classifier.
[0099] All inputs in the figure are finally padded to the same length. The text information is converted into the segmentation result by the tokenizer segmentation model. The two texts of a set of title pairs are separated by sep_token, and the cls_token at the beginning represents the classification information of the sentence pair. The box information needs to be normalized to the range of (0, 1000) in advance, and the two boxes are represented separately as the text information. Token_type_ids uses 0 and 1 to represent the information corresponding to the two titles, which is convenient for the model to distinguish. Attention_mask uses 1 to represent the text in the input, which is convenient for the model to distinguish zero padding and meaningful parts in the input.
[0100] Hierarchical restoration
[0101] In order to restore the document tree, it is necessary to obtain the level corresponding to each title according to the title pair relationship list output by the model. Specifically, in the model prediction stage, in order to build the title tree, it is necessary to find the parent title for each title element. First, the input ordered title list is traversed from back to front, and each title is predicted from back to front with all its predecessor titles. The first title with a subordinate relationship is taken to form a parent-child pair of the title tree. Similarly, the parent title is found for each title element. After traversing all the parent-child pairs, since the root node is the only element in all parent-child pairs that has only child nodes but no parent nodes, the root node of the title tree can be determined. At this point, the title tree is constructed, and the depth of the title tree can be directly corresponded to the title level.
[0102] Therefore, the technical solution of this application realizes the reconstruction of the document layout with a directory structure, which realizes the detection and recognition of layout elements, that is, classifying and locating elements such as text, graphics, and tables, and then extracting the element content by combining text recognition and other technologies, while considering more fine-grained semantic classification, predicting each title level, and restoring the structure that conforms to human reading habits, and outputting the markdown preview result as follows Figure 7 As shown in the upper right. To show the difference, also attached Figure 7 The existing open source technology MinerU shown in the lower right is limited to the detection and recognition of layout elements and ignores the document's logical structure. All titles are output as first-level titles. Therefore, the technical solution of this application has the following characteristics:
[0103] 1. Versatility: It covers multiple categories of layout elements and is compatible with multi-format document images, and can meet the needs of document layout reconstruction in most scenarios;
[0104] 2. Simple and efficient: Based on the current mainstream deep learning model, it completes the tasks of extracting layout elements and restoring document layout in an end-to-end manner, opening up the entire process of document layout reconstruction;
[0105] 3. High precision: Combining the CV model and the NLP model, and utilizing the visual features and semantic features of different dimensions in the document image, high precision can be achieved in the tasks of extracting layout elements and predicting logical relationships in different layouts.
[0106] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0107] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0108] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0109] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. A document layout reconstruction method, characterized in that: include: S1, document input and paging step: inputting an original document, and converting the original document into document images by paging; S2, physical layout analysis step: based on the deep learning instance segmentation algorithm, locate and classify each layout element area in the document image to obtain the category and position of each layout element; S3, input type determination step: determining whether the corresponding page of the document image can be directly parsed by code and whether it does not contain a table, if so, performing code parsing to obtain text-related information; otherwise, using OCR recognition to obtain text-related information; S4, logical layout analysis step: according to the position of each layout element, the category of each layout element is matched with the text-related information, each layout element is sorted, and then hierarchical information is added to the layout elements whose categories are document titles and hierarchical titles, thereby realizing the reconstruction of the document layout with a directory structure.
2. The document layout reconstruction method according to claim 1, characterized in that: In S1, the original document is a PDF document.
3. The document layout reconstruction method according to claim 1, characterized in that: In S2, the instance segmentation algorithm based on deep learning is a yolov8-seg instance segmentation algorithm based on deep learning.
4. The document layout reconstruction method according to claim 1, characterized in that: In S2, the categories of the layout elements include document title, directory, level title, paragraph, information block, table, picture, header, footer, page number, signature, seal, chart annotation, chart title, formula, and column.
5. The document layout reconstruction method according to claim 3, characterized in that: The S2 specifically includes: The yolov8-seg instance segmentation algorithm based on deep learning takes the document image as input, returns the category and corresponding mask area of each layout element in the document image, thereby obtaining the category and position of each layout element.
6. The document layout reconstruction method according to claim 1, characterized in that: The S2 further includes: fitting a minimum circumscribed rectangle to the contour points of the mask area.
7. The document layout reconstruction method according to claim 1, characterized in that: In S3, the OCR recognition specifically includes: table detection, table line segmentation and / or text line positioning, text line direction classification, text line font classification and text line recognition.
8. The document layout reconstruction method according to claim 1, characterized in that: In S4, adding hierarchical information to layout elements of the document title and hierarchical title category specifically includes: The layout elements of ordered document titles and hierarchical titles are traversed from back to front. Each hierarchical title and all its predecessor titles form a title pair from back to front. The pair is input into the deep learning network model based on the transformer architecture for prediction, and the output is a predefined prediction relationship. Take the first title that is output as a subordinate relationship to form a parent-child pair in the title tree, and so on to find the parent title for each title element; After traversing all parent-child pairs, since the root node is the only element in all parent-child pairs that has only child nodes but no parent nodes, the root node of the title tree is determined, which is the document title. Thus, the title tree is constructed, and the depth of the title tree directly corresponds to the hierarchical information of the title.
9. The document layout reconstruction method according to claim 8, characterized in that: The predefined prediction relationships include subordinate relationships, inclusion relationships, parallel relationships, and other relationships.
10. The document layout reconstruction method according to claim 9, characterized in that: The relationship between title pairs is distinguished according to the marked title levels. A level difference of 0 corresponds to a parallel relationship, a difference of 1 corresponds to an inclusion relationship, a difference of -1 corresponds to a subordinate relationship, and other differences correspond to other relationships.
11. A document layout reconstruction device, characterized in that: The document layout reconstruction device operates based on the document layout reconstruction method according to any one of claims 1 to 10, including: A document input paging unit, used for inputting an original document and converting the original document into document images by paging; A physical layout analysis unit, used for locating and classifying each layout element region in the document image based on a deep learning instance segmentation algorithm, and obtaining a category and position of each layout element; An input type determination unit is used to determine whether the corresponding page of the document image can be directly parsed by code and whether it does not contain a table. If so, code parsing is performed to obtain text-related information; otherwise, OCR recognition is used to obtain text-related information; The logical layout analysis unit is used to match the category of each layout element with the text-related information according to the position of each layout element, sort each layout element, and then add hierarchical information to the layout elements whose categories are document titles and hierarchical titles, thereby realizing the reconstruction of the document layout with a directory structure.
12. A document layout reconstruction system, the system comprising: A processor and a memory for storing executable instructions; characterized in that the processor is configured to execute the executable instructions to perform the document layout reconstruction method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the document layout reconstruction method according to any one of claims 1 to 10 is implemented.
Citation Information
Cited By
Image processing method and device
CN120783353A
Multi-dimensional data processing method and system
CN121033876A
Parsing and editing system and device for high-frame-rate rendering of PDF (Portable Document Format) document
CN121328486A
A PDF document high frame rate rendering parsing and editing system and device
CN121328486B