Table structure information identification method and device, computer equipment and storage medium
By inputting the table images into the cell detection model and the semantic segmentation model, and integrating the recognition results, the problem of low accuracy of table structure information recognition in the prior art is solved, and higher recognition accuracy and document integrity are achieved.
Patent Information
- Application Number
- CN202510280231.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the table structure information recognition method has the problem of low recognition accuracy, resulting in incomplete text documents.
By inputting the table image into the cell detection model and the semantic segmentation model, the cell recognition results and row-line recognition results are obtained respectively, and the two are fused to obtain the table recognition results, thereby extracting the target table structure information.
It improves the accuracy of table structure information identification, ensures the integrity of the generated text documents, and is suitable for financial technology and medical and health business fields.
Smart Images

Figure CN120107988A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, which is applied to the fields of financial technology and medical health business, and in particular to a method, device, computer equipment and storage medium for identifying table structure information. Background Art
[0002] In the fields of financial technology and medical health system business, the relevant management systems usually involve a large number of table images, which contain a large amount of information. In order to ensure the rapid processing of financial business or medical online business procedures, it is usually necessary to extract the table structure information and text information from the table image, and then generate editable Excel documents or markdown documents based on the table structure information and text information. For example, in the process of health insurance claims, table images can be table images of hospitalization settlement forms, medical invoices, etc. uploaded by customers in the business system.
[0003] At present, the existing technology usually detects cells from a table image through a cell detection model, and then splices the detected cells to obtain the table structure information. However, when the cell detection model detects the table image, it may not recognize all the cells, resulting in incomplete table structure information, and further resulting in incompleteness of the final generated text document.
[0004] Therefore, the existing table structure information recognition method has the problem of low recognition accuracy. Summary of the invention
[0005] The embodiments of the present invention provide a table structure information recognition method, device, computer equipment and storage medium to solve the problem of low recognition accuracy in the existing table structure information recognition method.
[0006] A method for identifying table structure information, comprising: Get a table image containing table data; Input the table image into a cell detection model to obtain a cell recognition result, and input the table image into a semantic segmentation model to obtain a row and column line recognition result; Merging the cell recognition result and the row and column line recognition result to obtain a table recognition result; Based on the table recognition result, target table structure information is extracted.
[0007] The above table structure information recognition method, optionally, inputting the table image into a cell detection model to obtain a cell recognition result includes: Inputting the table image into a cell detection model to obtain a plurality of candidate recognition results and a first confidence level corresponding to each of the candidate recognition results; The candidate recognition result whose first confidence level is greater than a preset first threshold is used as the cell recognition result.
[0008] The above table structure information recognition method may optionally include inputting the table image into a semantic segmentation model to obtain row and column line recognition results, including: Inputting the table image into a semantic segmentation model to obtain a plurality of candidate horizontal line recognition results and a second confidence level corresponding to each of the candidate horizontal line recognition results, and a plurality of candidate vertical line recognition results and a third confidence level corresponding to each of the candidate vertical line recognition results; The candidate horizontal line recognition results whose second confidence is greater than a preset second threshold value and the candidate vertical line recognition results whose third confidence is greater than a preset third threshold value are merged to obtain the row and column line recognition results.
[0009] In the above table structure information recognition method, optionally, extracting the target table structure information based on the table recognition result includes: Obtaining the cell coordinates of the table recognition result, wherein the cell coordinates are the coordinates of all intersections of horizontal and vertical lines in the table recognition result; The cell coordinates are converted based on a preset rule engine to obtain the target table structure information.
[0010] The above table structure information recognition method may, optionally, before inputting the table image into a cell detection model to obtain a cell recognition result, and inputting the table image into a semantic segmentation model to obtain a row and column line recognition result, the method may further include: Obtaining the first table structure information; Randomly splitting cells in the first table structure information to obtain a plurality of second table structure information; Filling each of the second table structure information with content respectively to generate a training image containing the table data; The semantic segmentation model and the cell detection model are trained respectively based on the training images to obtain the trained semantic segmentation model and the cell detection model.
[0011] Optionally, the above table structure information recognition method includes filling each of the second table structure information with content to generate a training image containing the table data, including: Fill each of the second table structure information with content to generate an initial image containing the table data; The initial image is distorted to generate the training image.
[0012] The above table structure information recognition method may, optionally, further include, after extracting the target table structure information based on the table recognition result: Input the table image into the target detection model to obtain a text image corresponding to each cell; Performing text recognition on each of the text images to obtain a filler text corresponding to each of the text images; A text file is generated based on the fill text and the target table structure information.
[0013] A table structure information recognition device, comprising: A table image acquisition module, used for acquiring a table image containing table data; A table image recognition module, used to input the table image into a cell detection model to obtain a cell recognition result, and input the table image into a semantic segmentation model to obtain a row and column line recognition result; A recognition result fusion module, used to fuse the cell recognition result and the row and column line recognition result to obtain a table recognition result; The structure information extraction module is used to extract the target table structure information based on the table recognition result.
[0014] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the above table structure information identification methods when executing the computer program.
[0015] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements any of the above table structure information identification methods.
[0016] The above table structure information recognition method, device, computer equipment and storage medium obtain cell recognition results by inputting the table image into the cell detection model, and obtain row and column line recognition results by inputting the table image into the semantic segmentation model; the cell recognition results and the row and column line recognition results are merged to obtain the table recognition results; based on the table recognition results, the target table structure information is extracted. It can be seen that the present invention performs table recognition on the table image in two ways to obtain cell recognition results and row and column line recognition results, and merges the cell recognition results and the row and column line recognition results to obtain the table recognition results, and then extracts the target table structure information from the table recognition results, realizing the complementary advantages of the two methods. Compared with only performing table structure information recognition through the cell detection model, it can achieve the purpose of improving the accuracy of table structure information recognition in the financial and medical fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0018] Figure 1 is a schematic diagram of an application environment of a table structure information identification method in one embodiment of the present invention; Figure 2 is a flow chart of a table structure information identification method in one embodiment of the present invention; Figure 3 is another flow chart of a method for identifying table structure information in one embodiment of the present invention; Figure 4 is another flow chart of a method for identifying table structure information in one embodiment of the present invention; Figure 5 is another flow chart of a method for identifying table structure information in one embodiment of the present invention; Figure 6 is another flow chart of a method for identifying table structure information in one embodiment of the present invention; Figure 7 is another flow chart of a method for identifying table structure information in one embodiment of the present invention; Figure 8 is another flow chart of a method for identifying table structure information in one embodiment of the present invention; Fig. 9 is a schematic diagram of a table structure information identification device in one embodiment of the present invention; Fig.10 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0021] It should also be understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0022] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]", depending on the context.
[0023] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0024] References to "one embodiment" or "some embodiments" etc. described in the present specification mean that one or more embodiments of the present invention include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0025] The present invention discloses a table structure information recognition method, device, computer equipment and storage medium, which obtains cell recognition results by inputting a table image into a cell detection model, and obtains row and column line recognition results by inputting the table image into a semantic segmentation model; the cell recognition results and the row and column line recognition results are fused to obtain a table recognition result; based on the table recognition result, the target table structure information is extracted. It can be seen that the present invention performs table recognition on the table image in two ways to obtain cell recognition results and row and column line recognition results, and fuses the cell recognition results and the row and column line recognition results to obtain the table recognition result, and then extracts the target table structure information from the table recognition result, realizing the complementary advantages of the two methods. Compared with only performing table structure information recognition through a cell detection model, the purpose of improving the accuracy of table structure information recognition in the financial and medical fields can be achieved.
[0026] The table structure information recognition method provided by the embodiment of the present invention can be applied as follows: Figure 1Specifically, the table structure information recognition method is applied in a table structure information recognition system, and the table structure information recognition system includes: Figure 1 The client and server shown in the figure communicate with each other through the network to improve the accuracy of identifying table structure information. The client, also known as the user end, refers to a program corresponding to the server that provides local services to customers. The client can be installed on, but not limited to, various personal computers, laptops, smart phones, tablets, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers.
[0027] In one embodiment, if Figure 2 As shown, a table structure information recognition method is provided, and the method is applied in Figure 1 The server in the example is used as an example to illustrate the process, including the following steps: S201: Acquire a table image containing table data.
[0028] In a specific implementation, the table image in this embodiment may be an image uploaded to the server by a user.
[0029] For example, when users are making claims for personal accident insurance and health insurance, they can use mobile phones and other electronic devices to take photos of paper information such as details of hospitalization medication, drug medical insurance types, total medical insurance reimbursement amounts, etc. to obtain a table image, and then upload the table image to the insurance company's server via the Internet through the insurance claims platform on the mobile phone. In this way, a table image containing the table data can be obtained.
[0030] S202: Input the table image into the cell detection model to obtain a cell recognition result, and input the table image into the semantic segmentation model to obtain a row and column line recognition result.
[0031] It can be understood that in this embodiment, the table image is input into the cell detection model and the semantic segmentation model respectively. The inputs of the cell detection model and the semantic segmentation model are the same, but the cell recognition results and row and column line recognition results output by the cell detection model may be different.
[0032] In a specific implementation, the cell detection model in this embodiment includes but is not limited to training based on any one of the CNN model, Faster model, R-CNN model, YOLO model, and table-transformer-detection model, etc. The semantic segmentation model in this embodiment includes but is not limited to training based on the FCN model, U-Net model, and DeepLab model.
[0033] S203: Merge the cell recognition result and the row, column and line recognition result to obtain a table recognition result.
[0034] Among them, in this embodiment, the cell recognition result includes a set of multiple cells, and the row and column line recognition result is a table line heat map.
[0035] In a specific implementation, in this embodiment, the column line recognition results can be visualized to obtain a table line heat map corresponding to the column line recognition results, and then the cells are drawn on the table heat map according to the positions of the cells in the table image to obtain the final table, that is, the table recognition result.
[0036] S204: Extracting target table structure information based on the table recognition result.
[0037] The target table structure information in this embodiment may be the HTML code of the table recognition result.
[0038] Optionally, in this embodiment, the target table structure information can be extracted based on the table recognition result through the following steps, such as Figure 3 As shown: S301: Acquire the cell coordinates of the table recognition result, where the cell coordinates are the coordinates of all intersections of horizontal and vertical lines in the table recognition result.
[0039] It can be understood that the table recognition result in this embodiment is a complete table. A table coordinate system is established based on the table recognition result, so that the coordinates of the intersection of each horizontal and vertical line in the table are determined in the table coordinate system. In this way, the cell coordinates of the table recognition result can be obtained.
[0040] S302: Convert the cell coordinates based on a preset rule engine to obtain target table structure information.
[0041] Among them, the rule engine in this embodiment includes a preset rule library, which stores business rules defined in some form (such as scripts, XML files, etc.) for converting cell coordinates to obtain target table structure information, that is, the html code of each cell coordinate.
[0042] In one embodiment, the business rules defined in the rule base of the rule engine of this embodiment may be as follows: 1. If the Y coordinates of two horizontal lines are the same and the number of vertical lines between them is greater than 0, a table row is defined between them; 2. If the X coordinates of two vertical lines are the same and the number of horizontal lines between them is greater than 0, a table column is defined between them; 3. Based on the intersection of the horizontal and vertical lines, the start and end coordinates of each cell can be determined.
[0043] It should be understood that the above is only an exemplary description of the business rules defined in the rule base of the rule engine, and does not limit the business rules defined in the rule base. Any technical solution formed by combining the business rules for the purpose of converting cell coordinates and obtaining target table structure information with the technical features in this embodiment is within the scope of protection of the present invention.
[0044] In summary, the present embodiment discloses a method, device, computer equipment and storage medium for identifying table structure information, which obtains cell recognition results by inputting a table image into a cell detection model, and obtains row and column line recognition results by inputting a table image into a semantic segmentation model; the cell recognition results and the row and column line recognition results are fused to obtain a table recognition result; and based on the table recognition result, the target table structure information is extracted. It can be seen that the present invention performs table recognition on a table image in two ways to obtain cell recognition results and row and column line recognition results, and fuses the cell recognition results and the row and column line recognition results to obtain a table recognition result, and then extracts the target table structure information from the table recognition result, realizing the complementary advantages of the two methods. Compared with only performing table structure information recognition through a cell detection model, the purpose of improving the accuracy of table structure information recognition can be achieved.
[0045] In one embodiment, Figure 4 As shown, in this embodiment, the table image can be input into the cell detection model to obtain the cell recognition result through the following steps: S401: Inputting a table image into a cell detection model to obtain a plurality of candidate recognition results and a first confidence level corresponding to each candidate recognition result.
[0046] S402: Taking the candidate recognition result whose first confidence is greater than a preset first threshold as the cell recognition result.
[0047] It can be understood that after the table image is input into the cell detection model, the cell detection model will identify the cells in the table image and output multiple alternative recognition results. The first confidence level corresponding to each alternative recognition result usually indicates the reliability of the corresponding alternative recognition result. The larger the value of the first confidence level, the higher the reliability of the alternative recognition result. Therefore, in this embodiment, the alternative cells are filtered based on the first confidence level by presetting the first threshold, that is, the alternative recognition results whose first confidence level is less than or equal to the preset first threshold are eliminated, and the alternative recognition results whose first confidence level is greater than the preset first threshold are retained as cell recognition results.
[0048] In one embodiment, Figure 5 As shown, in this embodiment, the table image can be input into the semantic segmentation model to obtain the row and column line recognition results through the following steps: S501: Inputting the table image into the semantic segmentation model to obtain multiple candidate horizontal line recognition results and a second confidence level corresponding to each candidate horizontal line recognition result, and multiple candidate vertical line recognition results and a third confidence level corresponding to each candidate vertical line recognition result.
[0049] It can be understood that in this embodiment, after the table image is input into the semantic segmentation model, the semantic segmentation model will simultaneously identify the horizontal and vertical lines in the table image. Therefore, the semantic segmentation model will simultaneously output multiple candidate horizontal line recognition results and multiple candidate vertical line recognition results, as well as multiple candidate horizontal line recognition results and the second confidence corresponding to each candidate horizontal line recognition result, and multiple candidate vertical line recognition results and the third confidence corresponding to each candidate vertical line recognition result. The second confidence corresponding to the alternative horizontal line recognition result represents the reliability of the alternative horizontal line recognition result, and the third confidence corresponding to the alternative vertical line recognition result represents the reliability of the alternative vertical line recognition result.
[0050] S502: Merge the candidate horizontal line recognition results whose second confidence is greater than the preset second threshold and the candidate vertical line recognition results whose third confidence is greater than the preset third threshold to obtain row and column line recognition results.
[0051] Based on the second threshold, the alternative horizontal line recognition results whose second confidence is less than or equal to the second threshold are eliminated, and the alternative horizontal line recognition results whose second confidence is greater than the second threshold are retained. Based on the third threshold, the alternative vertical line recognition results whose third confidence is less than or equal to the third threshold are eliminated, and the alternative vertical line recognition results whose third confidence is greater than the third threshold are retained by the third threshold.
[0052] It can be understood that the candidate horizontal line recognition results in this embodiment only include horizontal lines, and the candidate vertical line recognition results only include vertical lines. Therefore, the candidate horizontal line recognition results and the candidate vertical line recognition results need to be merged to obtain a table, that is, row and column line recognition results.
[0053] In one embodiment, Figure 6 As shown, in this embodiment, the following steps may also be included before step S101: S205: Obtain first table structure information.
[0054] S206: randomly splitting cells in the first table structure information to obtain a plurality of second table structure information.
[0055] It should be understood that the first table structure information is converted from cell coordinates through a rule engine. Therefore, the cells can be randomly split based on the cell coordinates to change the table structure corresponding to the first table structure information. Each time the cells are randomly split based on the cell coordinates, a new second table structure information will be obtained. This cycle can be repeated to obtain multiple second table structure information.
[0056] For example, taking the cell coordinates corresponding to the cell as (0, 0)(5, 0)(5, 2)(0, 2) as an example, by randomly splitting the cell, we can obtain two cells: (0, 0)(0, 2)(3, 2)(3, 0), and (3, 2)(3, 0)(5, 0)(5, 2), or we can obtain two cells: (0, 0)(0, 2)(2, 2)(2, 0), and (2, 2)(2, 0)(5, 0)(5, 2). In this way, the cells in the first table structure information can be randomly split to obtain multiple second table structure information.
[0057] S207: Fill each second table structure information with content to generate a training image containing table data.
[0058] In a specific implementation, in this embodiment, a corpus database containing a large amount of corpus can be pre-established. Each time a new second table result is generated, corpus is obtained from the corpus database to fill the second table structure information with content to generate a training image containing table data.
[0059] S208: Based on the training images, the semantic segmentation model and the cell detection model are trained respectively to obtain a trained semantic segmentation model and a cell detection model.
[0060] In a specific implementation, in this embodiment, the table structure in each training image can be annotated, and then the training image is input into the semantic segmentation model and the cell detection model for training to obtain a trained semantic segmentation model and a cell detection model.
[0061] In this embodiment, the semantic segmentation model can be trained based on the training image in the following manner: The training images are divided into a training set and a test set. The training images in the training set are input into the semantic segmentation model for iterative training. After each iterative training, the training images in the test set are input into the semantic segmentation model to obtain the row and column line recognition results output by the semantic segmentation model. The row and column line recognition results are converted into a table line heat map through methods such as Gradient-weighted ClassActivation Mapping (Grad-CAM). Then, the accuracy of the semantic segmentation model in recognizing the training images in the test set is judged based on the table line heat map. When the accuracy exceeds the preset accuracy threshold, it is determined that the semantic segmentation model has been fully trained.
[0062] To sum up, in this embodiment, by randomly splitting the cells in the first table structure information, multiple second table structure information are obtained, and then training images are generated based on each second table structure information, which effectively increases the number of training images and solves the problem of insufficient number of training images.
[0063] In one embodiment, Figure 7 As shown, step S107 in this embodiment can be implemented by the following steps: S701: Fill each second table structure information with content to generate an initial image containing table data.
[0064] S702: Perform distortion processing on the initial image to generate a training image.
[0065] In a specific implementation, in this embodiment, the initial image can be distorted by a distortion algorithm to generate a training image. The distortion algorithm in this embodiment includes but is not limited to any one of the thin plate spline algorithm (Thin-Plate-Spline, TPS), affine transformation (Affine Transformation) and elastic deformation (Elastic Deformation), which is not limited in this embodiment.
[0066] In this embodiment, the initial image is distorted to generate a training image so that the table in the training image is distorted to varying degrees, thereby solving the problem of lack of distorted training images. The semantic segmentation model and the cell detection model are trained using this training image, which can improve the recognition accuracy of the semantic segmentation model and the cell detection model for table structure information.
[0067] In one embodiment, Figure 8 As shown, in this embodiment, after step S104, the following steps are also included: S209: Input the table image into the target detection model to obtain the text image corresponding to each cell.
[0068] Optionally, the target detection model in this embodiment includes but is not limited to being trained based on any one of the R-CNN model and the YOLO model.
[0069] In the specific implementation, in this embodiment, the table image is input into the target detection model to obtain the target detection frame corresponding to each cell, and the table image is detected based on each target detection frame to obtain the text image corresponding to each target detection frame, that is, the text image corresponding to the cell.
[0070] S210: Perform text recognition on each text image to obtain the filling text corresponding to each text image.
[0071] S211: Generate a text file based on the fill text and the target table structure information.
[0072] In a specific implementation, in this embodiment, optical character recognition (OCR) can be used to perform text recognition on each text image to obtain the fill-in text corresponding to each text image, and then the recognized fill-in text is filled into the target table structure information to generate a text file. The format of the text file includes but is not limited to Excel format, markdown format, etc., which is not limited in this embodiment.
[0073] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0074] In one embodiment, a table structure information recognition device is provided, and the table structure information recognition device corresponds to the table structure information recognition method in the above embodiment. Fig. 9 As shown, the table structure information recognition device includes a table image acquisition module, a table image recognition module, a recognition result fusion module and a structure information extraction module. The detailed description of each functional module is as follows: The table image acquisition module 901 is used to acquire a table image containing table data; The table image recognition module 902 is used to input the table image into the cell detection model to obtain the cell recognition result, and input the table image into the semantic segmentation model to obtain the row and column line recognition result; The recognition result fusion module 903 is used to fuse the cell recognition result and the row and column line recognition result to obtain the table recognition result; The structure information extraction module 904 is used to extract the target table structure information based on the table recognition result.
[0075] The above-mentioned table structure information recognition device obtains cell recognition results by inputting a table image into a cell detection model, and obtains row and column line recognition results by inputting a table image into a semantic segmentation model; the cell recognition results and the row and column line recognition results are merged to obtain a table recognition result; and based on the table recognition result, the target table structure information is extracted. It can be seen that the present invention performs table recognition on a table image in two ways to obtain cell recognition results and row and column line recognition results, and merges the cell recognition results and the row and column line recognition results to obtain a table recognition result, and then extracts the target table structure information from the table recognition result, realizing the complementary advantages of the two ways. Compared with only performing table structure information recognition through a cell detection model, the purpose of improving the accuracy of table structure information recognition can be achieved.
[0076] In one embodiment, the table image recognition module 902 can be used to: Inputting the table image into the cell detection model to obtain a plurality of candidate recognition results and a first confidence level corresponding to each candidate recognition result; The candidate recognition result whose first confidence is greater than a preset first threshold is used as the cell recognition result.
[0077] In one embodiment, the table image recognition module 902 may also be used to: Inputting the table image into the semantic segmentation model to obtain a plurality of candidate horizontal line recognition results and a second confidence level corresponding to each candidate horizontal line recognition result, and a plurality of candidate vertical line recognition results and a third confidence level corresponding to each candidate vertical line recognition result; The candidate horizontal line recognition results whose second confidence is greater than the preset second threshold value and the candidate vertical line recognition results whose third confidence is greater than the preset third threshold value are merged to obtain the row and column line recognition results.
[0078] In one embodiment, the structure information extraction module 904 is used to Get the cell coordinates of the table recognition result. The cell coordinates are the coordinates of all intersections of horizontal and vertical lines in the table recognition result. The cell coordinates are converted based on the preset rule engine to obtain the target table structure information.
[0079] In one embodiment, a model training module is also included, which can be used to Obtaining the first table structure information; Randomly split cells in the first table structure information to obtain multiple second table structure information; Filling each second table structure information with content respectively to generate a training image containing table data; The semantic segmentation model and the cell detection model are trained respectively based on the training images to obtain the trained semantic segmentation model and the cell detection model.
[0080] In one embodiment, a model training module is also included, which can be used to Fill each second table structure information with content to generate an initial image containing table data; The initial image is distorted to generate the training image.
[0081] In one embodiment, a text file generation module is also included, which can be used to Input the table image into the object detection model to obtain the text image corresponding to each cell; Perform text recognition on each text image to obtain the filling text corresponding to each text image; Generate a text file based on the fill text and target table structure information.
[0082] For the specific definition of the table structure information identification device, please refer to the definition of the table structure information identification method above, which will not be repeated here. Each module in the above-mentioned table structure information identification device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0083] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.10 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a table structure information recognition method is implemented.
[0084] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the table structure information recognition method in the above embodiment is implemented, for example Figure 2 The steps of the table structure information identification method shown, or Figures 2 to 8Alternatively, when the processor executes the computer program, the functions of each module / unit in the embodiment of the table structure information identification device are realized, for example Fig. 9 To avoid repetition, it will not be described here.
[0085] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the table structure information recognition method in the above embodiment is implemented, for example Figure 2 The steps of the table structure information identification method shown, or Figures 2 to 8 Alternatively, when the processor executes the computer program, the functions of each module / unit in the embodiment of the table structure information identification device are realized, for example Fig. 9 To avoid repetition, it will not be described here.
[0086] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0087] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0088] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for identifying table structure information, characterized in that: include: Get a table image containing table data; Input the table image into a cell detection model to obtain a cell recognition result, and input the table image into a semantic segmentation model to obtain a row and column line recognition result; Merging the cell recognition result and the row and column line recognition result to obtain a table recognition result; Based on the table recognition result, target table structure information is extracted.
2. The table structure information recognition method according to claim 1, characterized in that: The step of inputting the table image into a cell detection model to obtain a cell recognition result includes: Inputting the table image into a cell detection model to obtain a plurality of candidate recognition results and a first confidence level corresponding to each of the candidate recognition results; The candidate recognition result whose first confidence level is greater than a preset first threshold is used as the cell recognition result.
3. The table structure information recognition method according to claim 1, characterized in that: The step of inputting the table image into a semantic segmentation model to obtain row and column line recognition results includes: Inputting the table image into a semantic segmentation model to obtain a plurality of candidate horizontal line recognition results and a second confidence level corresponding to each of the candidate horizontal line recognition results, and a plurality of candidate vertical line recognition results and a third confidence level corresponding to each of the candidate vertical line recognition results; The candidate horizontal line recognition results whose second confidence is greater than a preset second threshold value and the candidate vertical line recognition results whose third confidence is greater than a preset third threshold value are merged to obtain the row and column line recognition results.
4. The table structure information recognition method according to claim 1, characterized in that: The step of extracting target table structure information based on the table recognition result includes: Obtaining the cell coordinates of the table recognition result, wherein the cell coordinates are the coordinates of all intersections of horizontal and vertical lines in the table recognition result; The cell coordinates are converted based on a preset rule engine to obtain the target table structure information.
5. The table structure information recognition method according to claim 1, characterized in that: Before inputting the table image into a cell detection model to obtain a cell recognition result, and inputting the table image into a semantic segmentation model to obtain a row and column line recognition result, the method further includes: Get the first table structure information; Randomly splitting cells in the first table structure information to obtain a plurality of second table structure information; Filling each of the second table structure information with content respectively to generate a training image containing the table data; The semantic segmentation model and the cell detection model are trained respectively based on the training images to obtain the trained semantic segmentation model and the cell detection model.
6. The table structure information recognition method according to claim 5, characterized in that: The filling each of the second table structure information with content to generate a training image containing the table data includes: Fill each of the second table structure information with content to generate an initial image containing the table data; The initial image is distorted to generate the training image.
7. The table structure information recognition method according to claim 1, characterized in that: After extracting the target table structure information based on the table recognition result, the method further includes: Input the table image into the target detection model to obtain a text image corresponding to each cell; Performing text recognition on each of the text images to obtain a filler text corresponding to each of the text images; A text file is generated based on the fill text and the target table structure information.
8. A table structure information recognition device, characterized in that: include: A table image acquisition module, used for acquiring a table image containing table data; A table image recognition module, used to input the table image into a cell detection model to obtain a cell recognition result, and input the table image into a semantic segmentation model to obtain a row and column line recognition result; A recognition result fusion module, used to fuse the cell recognition result and the row and column line recognition result to obtain a table recognition result; The structure information extraction module is used to extract the target table structure information based on the table recognition result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the table structure information identification method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the table structure information identification method according to any one of claims 1 to 7 is implemented.