Semantic segmentation model training method and system and table analysis method and system

By training the semantic segmentation model to process the table with missing cell lines, the problem of inaccurate and high cost in the existing technology is solved, and more efficient and accurate table analysis effect is achieved.

CN120126151APending Publication Date: 2025-06-10HUA DATA TECH (SHANGHAI) CO LTD
0 Cites 0 Cited by

Patent Information

Application Number
CN202510214131.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, when parsing tables in PDF documents, there are problems of inaccurate identification or high resolution costs, especially when dealing with tables with missing cell lines.

Method used

By providing a training method and system for semantic segmentation models, the semantic segmentation model is used to train the table images of missing cell lines to generate predicted images with complete cell lines, thereby improving the accuracy of table analysis and reducing the analysis cost.

Benefits of technology

This method improves the accuracy of table parsing, reduces errors during the analysis process, reduces analysis costs, and performs well especially when dealing with complex tables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126151A_ABST
    Figure CN120126151A_ABST
Patent Text Reader

Abstract

The invention provides a semantic segmentation model training method and system and a table analysis method and system, and the training method comprises the following steps: obtaining training data which comprises a plurality of table images; wherein the table image has a corresponding semantic segmentation tag, the semantic segmentation tag is constructed based on a cell line generated by simulation, the cell line generated by simulation is generated according to annotation information of the table image, and the annotation information comprises a textbox position and a cell position where the textbox is located; inputting the training data into a semantic segmentation model for training; wherein the trained semantic segmentation model is used for outputting a prediction image, and cells of the prediction image have complete cell lines; the semantic segmentation model can output and obtain a prediction image with complete cell lines according to an input target table image with missing cell lines; therefore, the precision of table analysis is improved, and the cost of table analysis is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of table parsing, and in particular, to a method and system for training a semantic segmentation model, and a method and system for table parsing. Background Art

[0002] When processing PDF document parsing tasks, traditional methods rely on a multi-model integration workflow, and these models together constitute an efficient processing pipeline (referring to an automated workflow). This process first involves layout analysis, which can accurately identify and locate various UI (User Interface) elements in the document, such as basic elements like titles, headers, footers, body text, images, formulas, and tables. Each element is carefully classified for subsequent processing. For different elements in the document, specialized processing methods are used. For example, text content is converted through an optical character recognition (OCR) model to convert the characters in the image into an editable text format. For table content, a professional table detection model is used to identify its rows, columns, and cells, thereby extracting table data. After all elements are processed separately, the system recombines these elements according to the position information of the UI elements in the normal reading order of people to restore the structure of the entire document.

[0003] However, any small error in the parsing process, such as row and column misalignment, may lead to the loss or unavailability of the entire table's information, especially for wireframe tables because they have no clear cell boundaries, and rows and columns may be incorrectly separated or merged due to large or small spacing between table elements.

[0004] With the progress of technology, multi-modal large models have shown increasingly powerful capabilities in dealing with complex problems. It can directly parse documents through manually predefined prompts. This method is not only more direct but also highly flexible. They perform particularly well in dealing with complex elements in documents. For example, they can accurately parse structurally complex tables and precisely match table images with identification content. However, despite the advantages of multi-modal large models in dealing with complexity, they face challenges in processing speed. In practical applications, it may take more than ten minutes to process a medium-sized document. In addition, when faced with long documents, these models may need to upload the document separately multiple times, and this batch processing method may disrupt the content coherence between page numbers in the document, thereby affecting the final parsing effect.

[0005] These existing table parsing models have obvious deficiencies in processing speed, which may be a problem for application scenarios that require quick responses. At the same time, the calling cost of multi-modal large models is relatively high, which may limit their wide adoption in cost-sensitive applications. Summary of the Invention

[0006] The technical problem to be solved by the present disclosure is to overcome the problems of inaccurate recognition or high parsing cost in parsing tables with missing cell lines in PDF documents in the prior art, and to provide a training method and system for a semantic segmentation model, a table parsing method and system.

[0007] The present disclosure solves the above technical problems through the following technical solutions:

[0008] In a first aspect, a training method for a semantic segmentation model is provided. The training method includes the following steps:

[0009] Obtain training data, where the training data includes a plurality of table images; wherein, the table images have corresponding semantic segmentation labels, the semantic segmentation labels are constructed based on simulated cell lines, and the simulated cell lines are generated according to the annotation information of the table images, and the annotation information includes the position of the text box and the position of the cell where the text box is located;

[0010] Input the training data into the semantic segmentation model for training; wherein, the trained semantic segmentation model is used to output a prediction image, and all cells in the prediction image have complete cell lines.

[0011] Optionally, the semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model, the training data includes first training data and second training data, and the table images include original table images with missing cell lines;

[0012] The training method further includes:

[0013] Remove the existing cell lines in the original table image according to line detection to obtain a first table image;

[0014] The step of inputting the training data into the semantic segmentation model for training specifically includes:

[0015] Input the first training data into the first semantic segmentation model for training; wherein, the first training data includes a plurality of first table images;

[0016] Input the second training data into the second semantic segmentation model for training; wherein, the second training data includes a plurality of original table images.

[0017] Second aspect, a table parsing method is provided. The table parsing method includes the following steps:

[0018] Obtain a target table image with missing cell lines;

[0019] Input the target table image into a trained semantic segmentation model to obtain a predicted image; the semantic segmentation model is obtained according to the training method of the semantic segmentation model described in the first aspect.

[0020] Optionally, the step of obtaining a target table image with missing cell lines specifically includes:

[0021] Obtain a table image in a document according to a target detection model;

[0022] In response to the table in the table image having missing cell lines, obtain the table image as the target table image.

[0023] Optionally, the table parsing method further includes:

[0024] Input the predicted image into an OCR model to obtain a corresponding text box;

[0025] Match the cells in the predicted image with the corresponding text boxes and obtain a matching result;

[0026] Output the parsing result of the target table image according to the matching result.

[0027] Optionally, the step of matching the cells in the predicted image with the corresponding text boxes and obtaining a matching result specifically includes:

[0028] Obtain the ratio of the part of the text box in the corresponding cell to the text box; and / or,

[0029] Obtain the distance between the target point of the text box and the target point of the corresponding cell;

[0030] In response to the ratio being greater than a preset ratio, and / or, the distance being less than a preset distance, determine that the cell matches the text box.

[0031] Optionally, the table parsing method further includes:

[0032] Obtain the target quantity of target text boxes, where the target text boxes have matching cells;

[0033] In response to the proportion of the target quantity in the total quantity of text boxes being greater than or equal to a preset threshold, fill the text in the text box into the corresponding cell to obtain the parsing result;

[0034] In response to the proportion of the target quantity in the total quantity of text boxes being less than the preset threshold, input the target table image into the multi-modal large model for parsing.

[0035] Optionally, if the semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model;

[0036] Then the step of obtaining the target quantity of the target text box specifically includes:

[0037] Select the larger value between the first quantity and the second quantity as the final matching result of the target quantity; wherein, the first quantity is used to represent the quantity of the target text box obtained according to the first semantic segmentation model, and the second quantity is used to represent the quantity of the target text box obtained according to the second semantic segmentation model.

[0038] Optionally, the table parsing method further includes:

[0039] Obtain the predicted image corresponding to the empty cell in the parsing result, input the predicted image corresponding to the empty cell into the target detection model for detection, and obtain the detection result;

[0040] In response to the detection result not being empty, fill the content in the detected predicted image into the corresponding cell.

[0041] In a third aspect, a training system for a semantic segmentation model is provided, and the training system includes: a training data acquisition module and a training module;

[0042] The training data acquisition module is used to acquire training data, and the training data includes a plurality of table images; wherein, the table image has a corresponding semantic segmentation label, the semantic segmentation label is constructed based on the simulated cell lines, and the simulated cell lines are generated according to the annotation information of the table image, and the annotation information includes the text box position and the cell position where the text box is located;

[0043] The training module is used to input the training data into the semantic segmentation model for training; wherein, the trained semantic segmentation model is used to output a predicted image, and all cells of the predicted image have complete cell lines.

[0044] Optionally, the semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model, the training data includes first training data and second training data, and the table image includes an original table image with missing cell lines;

[0045] The training system further includes:

[0046] The cell line removal module is used to remove the cell lines existing in the original table image according to line detection, and obtain a first table image;

[0047] The training module is specifically configured to input the first training data into the first semantic segmentation model for training; wherein, the first training data includes a plurality of first table images; input the second training data into the second semantic segmentation model for training; wherein, the second training data includes a plurality of original table images.

[0048] In a fourth aspect, a table parsing system is provided, and the table parsing system includes an acquisition module and a predicted image generation module;

[0049] The acquisition module is used to acquire a target table image with missing cell lines;

[0050] The predicted image generation module is used to input the target table image into the trained semantic segmentation model to obtain a predicted image; the semantic segmentation model is trained according to the training system of the semantic segmentation model as described in the third aspect.

[0051] Optionally, the acquisition module is specifically configured to acquire a table image in a document according to a target detection model; in response to the table in the table image having missing cell lines, acquire the table image as the target table image.

[0052] Optionally, the table parsing system further includes:

[0053] The first input module is used to input the predicted image into an OCR model to obtain a corresponding text box;

[0054] The matching module is used to match the cells in the predicted image with the corresponding text boxes and obtain a matching result;

[0055] The output module is used to output the parsing result of the target table image according to the matching result.

[0056] Optionally, the matching module includes:

[0057] The ratio acquisition unit is used to acquire the ratio of the part of the text box in the corresponding cell to the text box; and / or,

[0058] The distance acquisition unit is used to acquire the distance between the target point of the text box and the target point of the corresponding cell;

[0059] The determination unit is used to determine that the cell matches the text box in response to the ratio being greater than a preset ratio, and / or the distance being less than a preset distance.

[0060] Optionally, the table parsing system further includes:

[0061] A target quantity acquisition module, configured to acquire the target quantity of the target text box, where the target text box has a matching cell;

[0062] A filling module, configured to, in response to the proportion of the target quantity in the total quantity of text boxes being greater than or equal to a preset threshold, fill the text in the text box into the corresponding cell to obtain the parsing result;

[0063] A second input module, configured to, in response to the proportion of the target quantity in the total quantity of text boxes being less than the preset threshold, input the target table image into a multi-modal large model for parsing.

[0064] Optionally, if the semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model;

[0065] Then the acquisition module specifically includes:

[0066] A selection unit, configured to select the larger value of the first quantity and the second quantity as the final matching result of the target quantity; where the first quantity is used to represent the quantity of the target text box obtained according to the first semantic segmentation model, and the second quantity is used to represent the quantity of the target text box obtained according to the second semantic segmentation model.

[0067] Optionally, the table parsing system further includes:

[0068] A third input module, configured to acquire a predicted image corresponding to an empty cell in the parsing result, input the predicted image corresponding to the empty cell into the target detection model for detection, and obtain a detection result;

[0069] The filling module is further configured to, in response to the detection result not being empty, fill the content in the detected predicted image into the corresponding cell.

[0070] In a fifth aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and configured to run on the processor. When the processor executes the computer program, the training method of the semantic segmentation model described in the first aspect or the table parsing method described in the second aspect is implemented.

[0071] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the semantic segmentation model described in the first aspect or the table parsing method described in the second aspect is implemented.

[0072] In a seventh aspect, there is provided a computer program product including a computer program which, when executed by a processor, implements the training method of the semantic segmentation model as described in the first aspect or the table parsing method as described in the second aspect.

[0073] Based on common general knowledge in the art, the above preferred conditions may be combined arbitrarily to obtain various preferred examples of the present disclosure.

[0074] The positive and progressive effects of the present disclosure are as follows: To solve the problem that the recognition of tables with missing cell lines in PDF documents is not accurate enough, the present application obtains a semantic segmentation model through model training. The semantic segmentation model can output a predicted image with complete cell lines according to the input target table image with missing cell lines, thereby improving the accuracy of table parsing and reducing the cost of table parsing at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 It is a flowchart of a training method of a semantic segmentation model provided in Embodiment 1 of the present disclosure;

[0076] Figure 2 It is a schematic diagram of an original table image in Embodiment 1 of the present disclosure;

[0077] Figure 3 It is a schematic diagram of a table image for generating predicted cells in Embodiment 1 of the present disclosure;

[0078] Figure 4 It is a schematic diagram of a predicted cell in Embodiment 1 of the present disclosure;

[0079] Figure 5 It is a schematic diagram of a first table image in Embodiment 1 of the present disclosure;

[0080] Figure 6 It is a flowchart of a table parsing method provided in Embodiment 2 of the present disclosure;

[0081] Figure 7 It is a specific flowchart of step S21 provided in Embodiment 2 of the present disclosure;

[0082] Figure 8 It is a partial flowchart of a table parsing method provided in Embodiment 2 of the present disclosure;

[0083] Figure 9 It is a specific flowchart of step S24 provided in Embodiment 2 of the present disclosure;

[0084] Figure 10 It is a partial flowchart of another table parsing method provided in Embodiment 2 of the present disclosure;

[0085] Figure 11It is a partial flowchart of another table parsing method provided in Embodiment 2 of the present disclosure;

[0086] Figure 12 It is a flowchart of another table parsing method provided in Embodiment 2 of the present disclosure;

[0087] Figure 13 It is a schematic diagram of modules of a training system for a semantic segmentation model provided in Embodiment 3 of the present disclosure;

[0088] Figure 14 It is a schematic diagram of modules of a table parsing system provided in Embodiment 4 of the present disclosure;

[0089] Figure 15 It is a schematic diagram of modules of an electronic device provided in Embodiment 5 of the present disclosure. Detailed implementation manners

[0090] The present disclosure will be further described below by way of embodiments, but the present disclosure is not limited to the scope of the described embodiments.

[0091] In the embodiments of the present disclosure, prefix words such as "first" and "second" are only used to distinguish different described objects, and do not limit the position, order, priority, quantity or content of the described objects. The use of ordinal numbers and other prefix words for distinguishing described objects in the embodiments of the present disclosure does not limit the described objects. The description of the described objects refers to the description in the context of the embodiments, and no redundant limitation should be formed due to the use of such prefix words. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.

[0092] In the embodiments of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0093] Embodiment 1

[0094] Figure 1 It is a flowchart of a training method for a semantic segmentation model provided in this embodiment. The training method includes the following steps:

[0095] S11. Obtain training data, where the training data includes a plurality of table images; among them, the table images have corresponding semantic segmentation labels, the semantic segmentation labels are constructed based on simulated cell lines, and the simulated cell lines are generated according to the annotation information of the table images, and the annotation information includes the position of the text box and the position of the cell where the text box is located.

[0096] S12. Input the training data into the semantic segmentation model for training; wherein, the trained semantic segmentation model is used to output a predicted image, and all cells of the predicted image have complete cell lines.

[0097] In this embodiment, first, training data is obtained. Specifically, the SciTSR (Scientific Table Structure Recognition, a dataset dedicated to table structure recognition) dataset can be selected for preprocessing to obtain the required training data. The SciTSR dataset is an open-source dataset composed of table data in scientific papers, where the training set includes 13,000 tables and the test set includes 3,000 tables. The labels for table recognition contain multiple types of annotation information, and the simulated generated cell lines are generated according to the annotation information of the table image. The annotation information includes the position of the text box (i.e., coordinates) and the position of the cell where the text box is located (i.e., row and column positions), as follows:

[0098] {"pos": [139.74600219726562, 194.87005615234375, 447.20001220703125, 452.1813049316406], "text": "shallow-deep"}

[0099] {"id": 4, "text": "shallow-deep", "content": [ "shallow-deep"], "start_row": 1, "end_row": 4, "start_col": 0, "end_col": 0}

[0100] Among them, pos are respectively the abscissa of the upper left corner, the abscissa of the lower right corner, the ordinate of the lower right corner, and the ordinate of the upper left corner of the cell where the text box is located. In a specific example, if the lower left corner of the image is used as the coordinate origin, it can be transformed to use the upper left corner as the coordinate origin through coordinate transformation. text is the name of the text box. As Figure 2 shown in the table, when the name of the text box text is "shallow-deep", the abscissa of the upper left corner of the cell where "shallow-deep" is located is 139.74600219726562, the abscissa of the lower right corner is 194.87005615234375, the ordinate of the lower right corner is 447.20001220703125, and the ordinate of the upper left corner is 452.1813049316406.

[0101] id represents the number of the table; "id": 4 means Figure 2The table shown has the number 4. The "content" represents the content of the cell. "text": "shallow-deep", "content": ["shallow-deep"] means that the content corresponding to the text box named shallow-deep is also shallow-deep.

[0102] start_row and end_row represent the starting row and ending row of the cell respectively, and start_col and end_col represent the starting column and ending column of the cell respectively. If the start value and the end value are the same, it means the cell occupies 1 row / column. If the end value is greater than the start value, it means the cell occupies multiple rows and multiple columns. For example Figure 2 In the table shown, "start_row": 1, "end_row": 4, "start_col": 0, "end_col": 0 means that the starting row of the cell corresponding to shallow-deep is 1 (the first basic cell), the ending row is 4 (the fourth basic cell), the starting column is 0 (the zeroeth basic cell), and the ending column is 0 (the zeroeth basic cell), that is, this cell occupies 3 rows and 1 column of the basic cells.

[0103] Starting from the above two types of annotation information (the position of the text box and the position of the cell where the text box is located), generate cell lines as the semantic segmentation labels for the semantic segmentation model to be trained. According to the position of the text box and the position of the cell where the text box is located, calculate the upper and lower bounds of each row of cells and the left and right bounds of each column of cells, that is, the minimum value of the upper left coordinates and the maximum value of the lower right coordinates of all text boxes in each row / column; then, use the average value of the lower bound (right bound) of the i-th row / column and the upper bound (left bound) of the (i + 1)-th row / column as the row and column lines in the table. At this time, the table has been divided into m×n cells, where m represents the number of rows of cells and n represents the number of columns of cells; next, consider the processing of merged cells, and remove the row and column line segments inside the merged cells (excluding the boundaries) from the rows and columns to generate Figure 3 the image shown.

[0104] Construct semantic segmentation labels through the cell lines generated by simulation. The pixel points where the horizontal and vertical cell lines are located are marked as 1, and other non-cell line pixel points are marked as 0. Based on pixel point classification, construct a binary image, that is, the pixel point value is 0 or 255, the pixel points marked as 1 are 0, and the pixel points marked as 0 are 255, to obtain Figure 4 the image shown.

[0105] Subsequently, the obtained training data is input into the semantic segmentation model for training. The training data includes multiple table images. Among them, the table images have corresponding semantic segmentation labels. The trained semantic segmentation model is used to output a prediction image, and all cells in the prediction image have complete cell lines. Specifically, according to Figure 4 the simulated generated cell lines shown in Figure 3 the image shown in

[0106] are segmented to obtain multiple prediction images.

[0107] In a specific example, the training data can be input into the UNet model for training, that is, the semantic segmentation model is the UNet model. The UNet model is a deep learning model mainly used for semantic segmentation tasks of images, especially performing well in the field of medical image segmentation. It consists of an encoder-decoder architecture and has a symmetric "U" shape structure. The encoder part gradually reduces the spatial size of the image and increases the depth of the feature map through convolutional layers and pooling layers to extract image features. The decoder part then gradually restores the original size of the image through upsampling and convolutional operations, and at the same time combines the feature map information of the encoder to improve the segmentation accuracy. One of the key innovations of the UNet model is the skip connections, which directly connect the output of the encoder layer to the output of the decoder layer. This helps to retain more spatial information and reduce the problem of gradient disappearance during the training process.

[0108] In addition, the loss function of the UNet model is a weighted cross-entropy loss function, and the formula of this loss function is as follows:

[0109]

[0110] where L cross : cross-entropy loss value;

[0111] N: The total number of samples, i.e., how many samples are involved in calculating the loss;

[0112] K: The total number of classes, i.e., how many different classes;

[0113] : The weight of the i-th sample. This weight can be used to handle the problem of sample imbalance. That is, when the number of samples in some classes is much larger than that in other classes, the weight can be adjusted to give higher attention to the minority classes.

[0114] : The true label that the i-th sample belongs to the k-th class, which is an indicator function. If sample i belongs to class k, then = 1, otherwise = 0;

[0115] : The probability that the model predicts the i-th sample belongs to the k-th class.

[0116] In an alternative implementation, the semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model, the training data includes first training data and second training data, and the table image includes the original table image with missing cell lines;

[0117] The training method further includes:

[0118] Remove the existing cell lines in the original table image according to line detection to obtain a first table image.

[0119] Since the table in the original table image still contains some cell lines (such as three-line diagrams, etc., Figure 2 as shown in the example), the simulated generated cell lines may not be aligned with the existing cell lines, and there may be multiple cell lines between rows and columns, which affects model training. Therefore, it is necessary to remove the existing cell lines from the original table image to generate a first table image as shown in Figure 5 This operation is mainly achieved through line detection of the Hough transform. In order not to affect the table text content, the minimum detected line length is set to 2 / 3 of the shortest cell line in the table, which can basically remove the cell lines. There will be some omissions, but it has little impact on model training. The original table image is marked as A1; the first table image obtained by removing the existing cell lines in the original table image according to line detection is marked as A2.

[0120] In a specific example, considering that some cell lines in the original table image are helpful for semantic segmentation but may result in misalignment with the annotation results, the UNet model is trained using the original table image A1 and the first table image A2 obtained by Hough transform to remove the remaining cells, resulting in 2 UNet models; thereby improving the accuracy of semantic segmentation. Therefore, the step S12 specifically includes:

[0121] Input the first training data into the first semantic segmentation model for training; wherein, the first training data includes a plurality of first table images. In this embodiment, the first training data including a plurality of first table images A2 can be input into the first semantic segmentation model UNet1 for training, wherein the trained semantic segmentation model UNet1 is used to output a prediction image, and all cells in the prediction image have complete cell lines.

[0122] Input the second training data into the second semantic segmentation model for training; wherein, the second training data includes a plurality of original table images. In this embodiment, the second training data including a plurality of original table images A1 can be input into the second semantic segmentation model UNet2 for training, wherein the trained semantic segmentation model UNet2 is used to output a prediction image, and all cells in the prediction image have complete cell lines.

[0123] Embodiment 2

[0124] Figure 6 The flowchart of a table parsing method provided for this embodiment, the table parsing method includes the following steps:

[0125] S21. Obtain a target table image with missing cell lines.

[0126] In this embodiment, the table image can be input into the classification layer to determine whether the table in the table image is missing cell lines, so as to obtain the target table image. Specifically, the classification of the table image can be implemented by an existing convolutional network.

[0127] S22. Input the target table image into the trained semantic segmentation model to obtain a prediction image; the semantic segmentation model is obtained according to the training method of the semantic segmentation model described in Embodiment 1.

[0128] In an alternative embodiment, as Figure 7 shown, the step S21 specifically includes:

[0129] S211. Obtain the table image in the document according to the object detection model.

[0130] In this embodiment, the document can be a PDF document. A target detection model can be used to perform layout analysis on the PDF document. Specifically, the target detection model can be a YOLO target detection model, which locates and classifies various elements in the PDF document page. These elements include, but are not limited to: ['Text', 'Title', 'Figure', 'Equation', 'Table', 'Caption', 'Header', 'Footer', 'BibInfo', 'Reference', 'Content', 'Code', 'Other', 'Item', 'Author'].

[0131] Among them, Text: refers to the pure text content in the document, excluding any formatting or structural marks.

[0132] Title: the title of the document, usually a summary of the document content or a short description of the theme.

[0133] Figure: a chart or graph used to visually display data or information to help readers better understand the document content.

[0134] Equation: a mathematical formula or equation used to express mathematical relationships or scientific principles.

[0135] Table: a table, a way of organizing and presenting data, usually consisting of rows and columns.

[0136] Caption: a figure caption or table note, which is a short description or explanation of the content of a chart or table.

[0137] Header: the header, usually appearing at the top of each page, containing the title of the document, chapter titles, or other information.

[0138] Footer: the footer, usually appearing at the bottom of each page, which may contain page numbers, copyright information, or other auxiliary information.

[0139] BibInfo: reference information, including author, publication year, publication name, etc., used for citations in academic writing.

[0140] Reference: references, referring to other documents or materials cited in the document.

[0141] Content: the content, referring to the main information or materials in the document, excluding formatting or structural elements.

[0142] Code: code, referring to instructions or statements in a programming language used to implement specific functions or algorithms.

[0143] Other: Others, used to classify elements or content that do not belong to any of the above categories.

[0144] Item: An item, usually used for a single entry in a list or a bulleted list.

[0145] Author: The person who creates or writes the content of a document.

[0146] S212. In response to the absence of cell lines in the table in the table image, obtain the table image as the target table image.

[0147] In this embodiment, the obtained table image is classified, and the table image with missing cell lines is used as the target table image. The table image can be input into the classification layer to determine whether the table in the table image is missing cell lines, so as to obtain the target table image.

[0148] In an alternative embodiment, as Figure 8 shown, the table parsing method further includes:

[0149] S23. Input the prediction image into the OCR model to obtain the corresponding text box.

[0150] S24. Match the cells in the prediction image with the corresponding text boxes and obtain the matching result.

[0151] In this embodiment, the text box and the cell are matched: if the text box is completely surrounded by a certain cell area, that is, the text box is within a cell, it is called a perfect match and is assigned a value of 1; if no cell can be perfectly matched, the positional relationship between the text box and the cell needs to be calculated to obtain the matching result.

[0152] S25. Output the parsing result of the target table image according to the matching result.

[0153] In an alternative embodiment, the table parsing method further includes: before step S23, scaling the prediction image proportionally. Specifically, the prediction image can be scaled up by a factor of two and then input into the OCR model to output the corresponding text box and the recognized text sequence, alleviating the situation of incorrect merging of text in different cells and improving the text recognition effect.

[0154] In an alternative embodiment, as Figure 9 shown, step S24 specifically includes:

[0155] S241. Obtain the ratio of the part of the text box within the corresponding cell to the text box.

[0156] S242. Obtain the distance between the target point of the text box and the target point of the corresponding cell.

[0157] S243. In response to the ratio being greater than a preset ratio and the distance being less than a preset distance, determine that the cell matches the text box.

[0158] In this embodiment, (1) the ratio of the part of the text box in the corresponding cell to the text box; (2) the distance between the target point of the text box and the target point of the corresponding cell are used as the evaluation basis to judge the matching result of each text box, so as to improve the accuracy of the matching result. In addition, the IoU (Intersection over Union) value = area of regional intersection / area of regional union can also be calculated to obtain the positional relationship between the text box and the cell.

[0159] In other embodiments, only (1) the ratio of the part of the text box in the corresponding cell to the text box can also be used as the evaluation basis, that is, only execute step S241. Obtain the ratio of the part of the text box in the corresponding cell to the text box; then execute the step: in response to the ratio being greater than a preset ratio, determine that the cell matches the text box.

[0160] In other embodiments, only (2) the distance between the target point of the text box and the target point of the corresponding cell can also be used as the evaluation basis, that is, only execute step S242. Obtain the distance between the target point of the text box and the target point of the corresponding cell; then execute the step: in response to the distance being less than a preset distance, determine that the cell matches the text box.

[0161] In an alternative embodiment, as Figure 10 shown, the table parsing method further includes:

[0162] S26. Obtain the target quantity of the target text box, and the target text box has a matching cell.

[0163] S27. In response to the ratio of the target quantity in the total quantity of text boxes being greater than or equal to a preset threshold, fill the text in the text box into the corresponding cell to obtain the parsing result.

[0164] In this embodiment, if the ratio of the target quantity in the total quantity of text boxes is greater than or equal to a preset threshold, it indicates that the prediction of the cell lines of the target table image is successful. Fill the text in the text box into the corresponding cell to obtain the parsing result, improving the accuracy of the parsing result.

[0165] S28. In response to the proportion of the target quantity in the total quantity of text boxes being less than the preset threshold, input the target table image into a multi-modal large model for parsing.

[0166] In this embodiment, if the proportion of the target quantity in the total quantity of text boxes is less than the preset threshold, it indicates that the prediction of the cells in the target table image fails, and a multi-modal large model needs to be used for parsing. This situation rarely occurs. Therefore, it is acceptable in terms of speed and cost to call the multi-modal large model to parse the table with failed metal table parsing prediction.

[0167] In an alternative embodiment, if the semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model; then step S26 specifically includes:

[0168] Select the larger value between the first quantity and the second quantity as the final matching result of the target quantity; wherein, the first quantity is used to represent the quantity of the target text boxes obtained according to the first semantic segmentation model, and the second quantity is used to represent the quantity of the target text boxes obtained according to the second semantic segmentation model.

[0169] In this embodiment, input the first table image into the first semantic segmentation model, and input the original table image into the second semantic segmentation model; obtain the first quantity and the second quantity respectively, and take the result with more quantity, that is, more cell-to-text box matches, as the target quantity, so as to obtain the final matching result, improving the accuracy and precision of table parsing.

[0170] The processing of nested tables is a relatively complex problem in document parsing. A nested table means that some cells may contain non-text type elements, such as pictures, formulas, or other visual content. In the layout analysis stage, the traditional layout model (a model for document layout analysis) may encounter challenges and is difficult to accurately identify the special elements in these sub-cells. This usually results in the entire area being simply marked as a table type, while ignoring the rich sub-cell elements. This deficiency in recognition may lead to the loss of important information during the parsing process. For example, if a cell in a table contains a key picture or formula, and the model fails to correctly identify it, then this visual or mathematical information may be missing in the final parsing result. This not only reduces the usability of the document but may also affect the user's ability to understand and utilize the document content. Therefore, the following table parsing method is provided, as Figure 11 shown, the table parsing method further includes:

[0171] S29. Obtain the predicted image corresponding to the empty cell in the parsing result, input the predicted image corresponding to the empty cell into the target detection model for detection, and obtain the detection result;

[0172] S30. In response to the detection result not being empty, fill the content in the detected predicted image into the corresponding cell.

[0173] In this embodiment, after the text box and the cell are matched, there may be some empty cells. For these cells, there may be two cases: (1) The cell itself has no content. (2) The information in the cell is not of text type and cannot be detected. Therefore, it is necessary to extract the image of the cell area and input it into the target detection model again. If the detection result is empty, it means that the cell has no content; otherwise, the detected content is directly filled into the corresponding cell. The detected content includes at least one of images, formulas, and other elements. Thus, the problem of recognizing nested tables is solved. The target detection model can be a YOLO target detection model, and the YOLO target detection model can classify different elements and identify non-text elements in the table.

[0174] In an alternative embodiment, if a table image with complete cell lines is obtained, the table image with complete cell lines is directly input into the regression model. The regression model is used to predict the position coordinates of each cell, and then the table image is input into the OCR model to obtain text boxes and perform the matching between the text boxes and the cells.

[0175] In a specific example, the process of table parsing can be divided into three stages: table structure recognition, OCR text recognition, and post-processing of filling the recognized text back into the cells. Figure 12 It is a flowchart of a table parsing method.

[0176] As Figure 12 shown, first input the table image into the classification layer to obtain a wireless table (i.e., the original table image without cell lines) and a wired table (i.e., the table image with complete cell lines).

[0177] Among them, the processing of the wired table is as follows:

[0178] Directly input the wired table into the OCR model for parsing; then perform post-processing, that is, match each cell of the text box to obtain the matching degree (i.e., the matching result) of the wired table. If the matching degree is lower than the threshold, input the wired table into the multi-modal large model for parsing and then input the parsing result; if the matching degree is not lower than the threshold, directly output the parsing result.

[0179] The processing of the wireless table is as follows: directly input the wireless table into the UNet2 model (i.e., the second semantic segmentation model) to generate a second wired table with predicted cell lines; perform Hough transform on the wireless table to remove the remaining cell lines to obtain a wireless table without cell lines (i.e., the first table image), and then input the wireless table without cell lines into the UNet1 model (i.e., the first semantic segmentation model) to generate a first wired table with predicted cell lines.

[0180] Input both the first wired table and the second wired table into the OCR model for recognition, and post-process the recognition results to obtain a first matching degree corresponding to the first wired table and a second matching degree corresponding to the second wired table; take the larger matching degree between the first matching degree and the second matching degree as the final matching degree. If the final matching degree is lower than the threshold, input the wireless table into the multi-modal large model for parsing and then input the parsing result; if the matching degree is not lower than the threshold, directly output the parsing result; overall, the accuracy of table recognition is improved.

[0181] Embodiment 3

[0182] Corresponding to Embodiment 1 of the training method of the foregoing semantic segmentation model, the present disclosure also provides an embodiment of a training system for the semantic segmentation model.

[0183] Figure 13 It is a schematic diagram of the modules of a training system for a semantic segmentation model provided in this embodiment. The training system 300 includes: a training data acquisition module 31 and a training module 32;

[0184] The training data acquisition module is used to acquire training data, and the training data includes a plurality of table images; wherein, the table images have corresponding semantic segmentation labels, the semantic segmentation labels are constructed based on simulated cell lines, and the simulated cell lines are generated according to the annotation information of the table images, and the annotation information includes the text box position and the cell position where the text box is located.

[0185] The training module is used to input the training data into the semantic segmentation model for training; wherein, the trained semantic segmentation model is used to output a prediction image, and all cells of the prediction image have complete cell lines.

[0186] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution.

[0187] Embodiment 4

[0188] Corresponding to the foregoing Table Parsing Method Embodiment 2, the present disclosure also provides an embodiment of a table parsing system.

[0189] Figure 14 It is a schematic diagram of modules of a table parsing system provided in this embodiment. The table parsing system 40 includes an acquisition module 41 and a predicted image generation module 42;

[0190] The acquisition module is configured to acquire a target table image with missing cell lines.

[0191] The predicted image generation module is configured to input the target table image into a trained semantic segmentation model to obtain a predicted image; the semantic segmentation model is trained according to the training system of the semantic segmentation model as described in Embodiment 3.

[0192] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution.

[0193] Embodiment 5

[0194] Figure 15 It is a schematic diagram of the structure of an electronic device shown in an exemplary embodiment of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and configured to run on the processor. When the processor executes the computer program, it implements the training method of the semantic segmentation model described in Embodiment 1 above, or the table parsing method described in Embodiment 2. Figure 15 The shown electronic device 50 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0195] Such as Figure 15As shown, the electronic device 50 may be presented in the form of a general-purpose computing device. For example, it may be a server device. The components of the electronic device 50 may include, but are not limited to: the at least one processor 51 described above, the at least one memory 52 described above, and a bus 53 that connects different system components (including the memory 52 and the processor 51).

[0196] The bus 53 includes a data bus, an address bus, and a control bus.

[0197] The memory 52 may include volatile memory, such as random access memory (RAM) 521 and / or cache memory 522, and may further include read-only memory (ROM) 523.

[0198] The memory 52 may also include a program tool 525 (or utility) having a set (at least one) of program modules 524. Such program modules 524 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0199] The processor 51 executes various functional applications and data processing by running computer programs stored in the memory 52, such as the training method of the semantic segmentation model described in Embodiment 1 above, or the table parsing method described in Embodiment 2.

[0200] The electronic device 50 may also communicate with one or more external devices 54 (such as a keyboard, a pointing device, etc.). Such communication may be carried out through an input / output (I / O) interface 55. And, the electronic device 50 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 56. As shown in the figure, the network adapter 56 communicates with other modules of the electronic device 50 through the bus 53. It should be understood that although Figure 15 not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 50, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0201] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided and embodied by multiple units / modules.

[0202] Example 6

[0203] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the training method of the semantic segmentation model described in Embodiment 1 above, or the table parsing method described in Embodiment 2.

[0204] Among them, the readable storage medium can more specifically include but is not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0205] Example 7

[0206] An embodiment of the present disclosure also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the training method of the semantic segmentation model described in Embodiment 1 above, or the table parsing method described in Embodiment 2.

[0207] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0208] Although the specific implementation manners of the present disclosure have been described above, those skilled in the art should understand that this is only an example. The protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these implementation manners, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A method for training a semantic segmentation model, characterized in that: The training method comprises the following steps: Acquire training data, the training data comprising a plurality of table images; wherein the table images have corresponding semantic segmentation labels, the semantic segmentation labels are constructed based on simulated cell lines, the simulated cell lines are generated according to annotation information of the table images, the annotation information comprises a text box position and a cell position where the text box is located; The training data is input into a semantic segmentation model for training; wherein the trained semantic segmentation model is used to output a predicted image, and cells of the predicted image all have complete cell lines.

2. The training method according to claim 1, characterized in that: The semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model, the training data includes first training data and second training data, and the table image includes an original table image with missing cell lines; The training method further comprises: Removing cell lines existing in the original table image according to straight line detection to obtain a first table image; The step of inputting the training data into the semantic segmentation model for training specifically includes: Inputting the first training data into the first semantic segmentation model for training; wherein the first training data includes a plurality of first table images; The second training data is input into the second semantic segmentation model for training; wherein the second training data includes a plurality of original table images.

3. A table parsing method, characterized in that: The table parsing method comprises the following steps: Obtain a target table image with missing cell lines; The target table image is input into a trained semantic segmentation model to obtain a predicted image; the semantic segmentation model is obtained according to the training method of the semantic segmentation model as described in claim 1 or 2.

4. The table parsing method according to claim 3, characterized in that: The step of obtaining a target table image with missing cell lines specifically includes: Obtain table images in the document based on the object detection model; In response to a cell line being missing in a table in the table image, the table image is acquired as the target table image.

5. The table parsing method according to claim 4, characterized in that: The table parsing method further includes: Input the predicted image into the OCR model to obtain a corresponding text box; Matching the cells in the predicted image with the corresponding text boxes and obtaining a matching result; The parsing result of the target table image is outputted according to the matching result.

6. The table parsing method according to claim 5, characterized in that: The step of matching the cell in the predicted image with the corresponding text box and obtaining a matching result specifically includes: Obtaining the ratio of the portion of the text box in the corresponding cell to the text box; and / or, Obtaining the distance between the target point of the text box and the target point of the corresponding cell; In response to the ratio being greater than a preset ratio and / or the distance being less than a preset distance, it is determined that the cell matches the text box.

7. The table parsing method according to claim 6, characterized in that: The table parsing method further includes: Obtaining a target number of target text boxes, wherein the target text boxes have matching cells; In response to the proportion of the target quantity in the total quantity of the text boxes being greater than or equal to a preset threshold, filling the text in the text box into a corresponding cell to obtain the parsing result; In response to the proportion of the target quantity in the total quantity of the text boxes being less than the preset threshold, the target table image is input into the multimodal large model for parsing.

8. The table parsing method according to claim 7, characterized in that: If the semantic segmentation model includes a first semantic segmentation model and a second semantic segmentation model; Then the step of obtaining the target number of target text boxes specifically includes: Select the larger value of the first number and the second number as the final matching result of the target number; wherein the first number is used to represent the number of target text boxes obtained according to the first semantic segmentation model, and the second number is used to represent the number of target text boxes obtained according to the second semantic segmentation model.

9. The table parsing method according to any one of claims 5 to 8, characterized in that: The table parsing method further includes: Obtaining a predicted image corresponding to an empty cell in the analysis result, and inputting the predicted image corresponding to the empty cell into the target detection model for detection to obtain a detection result; In response to the detection result being not empty, the detected content in the predicted image is filled into the corresponding cell.

10. A training system for a semantic segmentation model, characterized in that: The training system comprises: a training data acquisition module and a training module; The training data acquisition module is used to acquire training data, wherein the training data includes a plurality of table images; wherein the table images have corresponding semantic segmentation labels, and the semantic segmentation labels are constructed based on simulated cell lines, and the simulated cell lines are generated according to the annotation information of the table images, and the annotation information includes the position of a text box and the position of the cell where the text box is located; The training module is used to input the training data into a semantic segmentation model for training; wherein the trained semantic segmentation model is used to output a predicted image, and cells of the predicted image all have complete cell lines.

11. A table parsing system, characterized in that: The table parsing system includes an acquisition module and a prediction image generation module; The acquisition module is used to acquire a target table image with missing cell lines; The predicted image generation module is used to input the target table image into a trained semantic segmentation model to obtain a predicted image; the semantic segmentation model is trained according to the semantic segmentation model training system as described in claim 10.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, it implements the training method of the semantic segmentation model described in claim 1 or 2, or the table parsing method described in any one of claims 3-9.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the training method of the semantic segmentation model described in claim 1 or 2, or the table parsing method described in any one of claims 3-9.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the training method of the semantic segmentation model as described in claim 1 or 2, or the table parsing method as described in any one of claims 3-9.