Table recognition method, device, equipment and storage medium
The affine transformation model corrects the table tilt, combined with semantic recombination and object detection model, solves the problem of low accuracy in wireless table correction and information extraction, and achieves high-accuracy table information recognition.
Patent Information
- Application Number
- CN202411909603.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The prior art cannot effectively correct the tilt of wireless tables, and the accuracy of table information extraction is low.
The affine transformation model is used to correct the tilt of the table image, and the table text is reorganized with the semantic recombination model, and the object detection model is used to perform semantic enhancement and encoding and decoding processing to identify the table information.
It improves the correction accuracy and generalization of borderless tables, enhances the recognition accuracy of cross-line text segments, and improves the recognition effect of complex cell structures.
Smart Images

Figure CN119360401B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a table recognition method, device, equipment and storage medium. Background Art
[0002] In documents of various fields, tables are an important form of displaying data information. In order to parse and extract data information from various tables, it is necessary to identify the tables first.
[0003] There are two methods for identifying tables in the prior art. One is to first use the Hough line detection method to correct the tilt of the table image, and then determine the table structure based on the corrected image to identify the table. The other is to detect the table based on the four-point detection method and perform deformation processing, then detect the cell coordinates of the table, and use the border clustering method to determine the minimum cell information of the table, so as to obtain the number of rows and columns across the table, and use the border clustering method to supplement the cells missed by the model, so as to achieve table recognition.
[0004] However, in the above two methods, the first method cannot effectively correct the tilt of the wireless table using the Hough line detection method, and the accuracy of table information extraction is low. In the second method, the four-point detection method cannot process wireless tables and has certain limitations. Summary of the invention
[0005] The purpose of the present application is to provide a table recognition method, device, equipment and storage medium to address the deficiencies in the above-mentioned prior art, so as to solve the problems in the prior art that the tilt of wireless tables cannot be effectively corrected and the accuracy of table information extraction is low.
[0006] To achieve the above objectives, the technical solutions adopted in this application are as follows:
[0007] In a first aspect, the present application provides a table recognition method, the method comprising:
[0008] The original table image is tilt-corrected based on the pre-trained affine transformation model to obtain a corrected table image;
[0009] Performing text detection on the corrected table image to obtain original table text information, wherein the original table text information at least includes: original table text;
[0010] The original table text is reorganized based on the pre-trained semantic reorganization model to obtain a plurality of reorganized table text features, wherein the reorganized table text features are used to characterize the reorganized table text, and if there are cross-line text segments in the original table text, the cross-line text segments are in the same paragraph in the reorganized table text;
[0011] Inputting the corrected table image and the plurality of reorganized table text features into a pre-trained target detection model, the target detection model performs semantic enhancement, encoding and decoding processing on the corrected table image based on the reorganized table text features, and obtains the start and end position coordinates and the start and end rows and columns of each cell in the corrected table image;
[0012] According to the start and end position coordinates and the start and end rows and columns of each cell, the table information in the original table image is identified.
[0013] In a second aspect, the present application provides a table recognition device, the device comprising:
[0014] A correction module, used for performing tilt correction on the original table image based on a pre-trained affine transformation model to obtain a corrected table image;
[0015] A detection module, used for performing text detection on the corrected table image to obtain original table text information, wherein the original table text information at least includes: original table text;
[0016] A reorganization module, used for performing text reorganization on the original table text based on a pre-trained semantic reorganization model to obtain a plurality of reorganized table text features, wherein the reorganized table text features are used to characterize the reorganized table text, and if there are cross-line text segments in the original table text, the cross-line text segments are in the same paragraph in the reorganized table text;
[0017] An enhancement module, used for inputting the corrected table image and the plurality of reorganized table text features into a pre-trained target detection model, and the target detection model performs semantic enhancement, encoding and decoding processing on the corrected table image based on the reorganized table text features to obtain the start and end position coordinates and start and end rows and columns of each cell in the corrected table image;
[0018] The recognition module is used to recognize the table information in the original table image according to the start and end position coordinates and the start and end rows and columns of each cell.
[0019] In a third aspect, the present application provides a model training method for training an affine transformation model in the table recognition method described in the first aspect, the method comprising:
[0020] Determine a plurality of matrix groups, each matrix group includes a forward transformation matrix and a target backward transformation matrix, wherein the forward transformation matrix and the target backward transformation matrix are inverse matrices of each other;
[0021] Transforming a plurality of original training images based on the forward transformation matrix in each of the matrix groups to obtain a plurality of deformed intermediate training images;
[0022] Preprocessing each of the intermediate training images to obtain a plurality of target training images;
[0023] Input a plurality of the target training images into an initial affine transformation model to obtain an actual backward transformation matrix output by the initial affine transformation model, and iteratively correct the initial affine transformation model according to the target backward transformation matrix in the matrix group corresponding to each of the target training images and the actual backward transformation matrix to obtain the affine transformation model.
[0024] In a fourth aspect, the present application provides a model training device, the device comprising:
[0025] A determination module, used to determine a plurality of matrix groups, each matrix group includes a forward transformation matrix and a target backward transformation matrix, wherein the forward transformation matrix and the target backward transformation matrix are inverse matrices of each other;
[0026] A transformation module, used for transforming a plurality of original training images based on the forward transformation matrix in each of the matrix groups to obtain a plurality of deformed intermediate training images;
[0027] A preprocessing model, used for preprocessing each of the intermediate training images to obtain a plurality of target training images;
[0028] A correction model is used to input multiple target training images into an initial affine transformation model to obtain an actual backward transformation matrix output by the initial affine transformation model, and iteratively correct the initial affine transformation model according to the target backward transformation matrix in the matrix group corresponding to each target training image and the actual backward transformation matrix to obtain the affine transformation model.
[0029] In a fifth aspect, the present application provides an electronic device comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the table recognition method described in the first aspect or the steps of the model training method described in the third aspect.
[0030] In a sixth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the table recognition method described in the first aspect or the steps of the model training method described in the third aspect are executed.
[0031] The beneficial effects of the present application are: the original table image is tilted and corrected based on the affine transformation model. Since the affine transformation does not pay attention to the borders of the cells in the table image, the tilt correction has a high accuracy, is applicable to borderless tables, and has strong generalization. The original table text detected in the corrected table image is reorganized based on the semantic reorganization model, so that the cross-line text segments in the original table text are reorganized into texts in the same paragraph, thereby improving the recognition accuracy of multi-line cell text. The target detection model performs semantic enhancement and encoding and decoding processing on the corrected table image based on the reorganized table text features. Since the target detection model combines the reorganized table text features with the corrected table image to achieve semantic enhancement, the recognition of complex cell structures is further improved. After encoding and decoding processing, the table information in the original table image is obtained based on the start and end position coordinates and the start and end rows and columns of each cell in the output corrected table image, and the start and end positions of each cell are used as logical relationship coordinates such as start and end rows and columns, thereby obtaining more accurate table information. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0033] Figure 1 A flowchart of a table recognition method provided in an embodiment of the present application;
[0034] Figure 2 A schematic diagram of a corrected table image provided in an embodiment of the present application;
[0035] Figure 3 A schematic diagram of the structure of the semantic reorganization model provided in the embodiment of the present application;
[0036] Figure 4 The table structure information and table information schematic diagram provided for the embodiment of the present application;
[0037] Figure 5 A schematic diagram of the training process of the affine transformation model provided in an embodiment of the present application;
[0038] Figure 6 A schematic diagram of a process for obtaining an affine transformation model provided in an embodiment of the present application;
[0039] Figure 7 A schematic diagram of a process for obtaining the start and end position coordinates and the start and end rows and columns of each cell provided in an embodiment of the present application;
[0040] Figure 8 A schematic diagram of the working structure of the semantic reconstruction model and the target detection model provided in the embodiments of the present application;
[0041] Fig. 9 A schematic diagram of a process for obtaining weighted image features provided in an embodiment of the present application;
[0042] Fig.10 A schematic diagram of a process for determining the starting and ending position coordinates and the starting and ending rows and columns of each cell provided in an embodiment of the present application;
[0043] Fig.11 A schematic diagram of a process for determining table information provided in an embodiment of the present application;
[0044] Fig.12 A module structure diagram of a table recognition device provided in an embodiment of the present application;
[0045] Fig.13 A module structure diagram of a model training device provided in an embodiment of the present application;
[0046] Fig.14 A schematic diagram of the structure of an electronic device 14 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.
[0048] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0049] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0050] In the current table recognition methods, table correction is usually performed in a traditional way, and cell detection is performed based on image input to achieve table structure recognition. Table correction methods, for example, use the Hough line detection method and the four-point detection method. However, for complex tables such as financial financial statements and research documents, the table style may have few lines or even no lines. The Hough line detection method and the four-point detection method cannot handle the tilt of wireless tables. Therefore, the table correction accuracy is low and the generalization is not high. In addition, the semantic information of the table text is not taken into account in the process of cell detection based on image input, making it difficult to handle the structural recognition of complex cells with multiple lines of text. When calculating the logical position of the cell to determine the table structure, the border clustering method is used to determine the minimum cell information, thereby obtaining the number of rows and columns across the table. At the same time, the border clustering method is used to supplement the cells missed by the model. This method relies on rule calculation and does not take into account the logical coordinates of the cell. Therefore, the determined table structure is not accurate enough.
[0051] Based on the above problems, the present application proposes a table recognition method, which performs tilt correction on the original table based on an affine transformation model. The affine transformation model does not correct the table based on the border, so the table correction is generalized. Then, the target detection model is used to semantically enhance, encode and decode the corrected table image based on the semantically reorganized table text features. Since the target detection model uses the table text features to semantically enhance the corrected table image, and the table text features have been semantically reorganized, the accuracy of cell recognition with line breaks in the table is further improved. The table information in the original table image is obtained based on the start and end position coordinates and the start and end rows and columns of each cell output by the target detection model, so that the table information is obtained using more specific cell position logic, thereby improving the accuracy of complex table structures.
[0052] Before introducing the technical solution of the embodiment of the present application, it is first necessary to explain that the form recognition method provided in the embodiment of the present application must obtain the user's consent in an explicit form before the personal information and important data collected and generated during the execution process can be applied to the solution of the embodiment of the present application. At the same time, the personal information and important data collected and generated during the execution of the form recognition method provided in the embodiment of the present application must comply with the principles of legality, legitimacy, and necessity, disclose the collection and use rules, and explicitly state the purpose, method, and scope of collecting and using information. The relevant data in the above examples do not include personal information that is not related to the relevant services provided by the above examples.
[0053] The embodiments of the present application can be applied to any field or scenario where table recognition is required. For example, in the financial field, the method of the embodiments of the present application can be used to recognize the structure of financial reports and research reports to obtain table information.
[0054] The table recognition method provided in the embodiments of the present application is described in detail below in combination with multiple embodiments.
[0055] Figure 1 The flowchart of the table recognition method provided in the embodiment of the present application is shown in FIG. 1 . The execution subject of the method can be any electronic device with processing capability. Figure 1 As shown, the method includes:
[0056] S101. Perform tilt correction on the original table image based on the pre-trained affine transformation model to obtain a corrected table image.
[0057] Optionally, the original table image may be a table in various forms, including but not limited to a table with a border, a table without a border, and a table with a partial border. Furthermore, the original table image may be obtained by scanning or photographing, and may be distorted or deformed.
[0058] Optionally, the affine transformation model can be a pre-trained DocScanner model.
[0059] S102: Perform text detection on the corrected table image to obtain original table text information, which at least includes: original table text.
[0060] As an optional implementation, an optical character recognition (OCR) algorithm may be used to perform text detection on the corrected table image to obtain original table text information.
[0061] Optionally, the original table text included in the original table text information may be a combination of multiple text segments in the corrected table image. Specifically, the text segments in the corrected table image are sorted from left to right or from right to left according to the starting horizontal coordinates of each text segment, to obtain a sorted text segment sequence, and each text segment in the text segment sequence is separated by a preset separator and then spliced to obtain the original table text. The preset separator may be, for example, "|". It is worth noting that in the process of sorting from left to right or from right to left according to the horizontal coordinates of each text segment in the corrected table image, if the starting horizontal coordinates of multiple text segments are the same, the order of the multiple text segments in the sorted sequence is not limited.
[0062] For example, Figure 2This is a schematic diagram of a corrected table image provided in an embodiment of the present application. Figure 2 As shown in the figure, the corrected table image includes 8 cells, some of which have no borders or only have partial borders. The text segments in the 8 cells are: { / }, {the Group}, {assets}, {derivative funds}, {financial assets}, {January}, {1000}, {February}, and {2000}. Among them, {derivative funds} and {financial assets} are two text segments obtained after line breaks in the same cell, and { / } is Figure 2 The empty text segment to the left of "The Group" in the table. Sort the above text segments based on their starting horizontal coordinates. For example, sort the text segments from left to right to get the original table text: { / }|{Derivative Finance}|{Financial Assets}|{Assets}|{1000}|{January}|{The Group}|{2000}|{February}.
[0063] As an optional implementation, when generating the original table text, the empty text segment can be placed at the very front or at the very end of the original table text. The above example shows the case where the empty text segment is placed at the very front.
[0064] S103. Performing text reorganization on the original table text based on the pre-trained semantic reorganization model to obtain a plurality of reorganized table text features, wherein the reorganized table text features are used to characterize the reorganized table text. If there are cross-line text segments in the original table text, the cross-line text segments are in the same paragraph in the reorganized table text.
[0065] Optionally, the semantic restructuring model may be a text-to-text language model, such as a T5 (Text to Text Transfer Transformer) model.
[0066] As an optional implementation, the semantic reorganization model may include a first encoder. The original table text is input into the first encoder of the semantic reorganization model for encoding and then text reorganization is performed, and a plurality of reorganized table text features are output.
[0067] Optionally, if a cell in the original table image includes multiple lines of text segments, the multiple lines of text segments are cross-line text segments. The table text represented by the reorganized table text feature is the text in each cell in each corrected table image, and the cross-line text segments in each cell are in the same paragraph. Among them, the paragraph is text divided according to semantics, and a paragraph can include one or more text segments.
[0068] Exemplarily, the original table text is: {derivative finance}|{financial assets}|{assets}|{1000}|{January}|{the Group}|{2000}|{February}, then the table text represented by the reorganized table text feature is: {derivative financial assets}|{assets}|{1000}|{January}|{the Group}|{2000}|{February}, wherein the cross-row text segments {derivative finance} and {financial assets} in the original table text are reorganized into {derivative financial assets}, that is, in the table text represented by the reorganized table text feature, the cross-row text segments that were originally separated in the same cell are reorganized into text in the same paragraph, so that the reorganized table text feature is consistent with the actual semantics in the original table.
[0069] Optionally, the number of text segments in the text represented by the reorganized table text feature is less than or equal to the number of text segments in the original table text.
[0070] The training process of the semantic reorganization model is as follows: obtain multiple training table texts in advance. The training table texts can be detected from multiple undeformed table images, and the cells in the above table images have multiple lines. Using manually annotated Hypertext Markup Language (HTML) tags, extract the text segments of all cells in the order of the HTML structure, and use preset delimiters to separate the text segments and then splice them to obtain labeled table texts. The labeled table text reorganizes multiple cross-line text segments in the same cell into texts in the same paragraph. The initial semantic reorganization model is trained based on multiple training table texts and multiple labeled table texts to obtain a semantic reorganization model.
[0071] S104. Input the corrected table image and multiple reorganized table text features into a pre-trained target detection model, and the target detection model performs semantic enhancement, encoding and decoding processing on the corrected table image based on the reorganized table text features to obtain the start and end position coordinates and start and end rows and columns of each cell in the corrected table image.
[0072] As an optional implementation, the target detection model first extracts image features from the corrected table image to obtain multiple image features, and then semantically enhances the corrected table image based on the reorganized table text, and then encodes and decodes the semantically enhanced features to obtain the start and end position coordinates and start and end rows and columns of each cell in the corrected table image. Among them, semantically enhancing the corrected table image based on the reorganized table text can significantly improve the robustness and accuracy of image recognition.
[0073] As an optional implementation, the encoding and decoding process can be implemented by an encoder and two decoders. Specifically, the semantically enhanced features are input into the encoder, the encoder encodes the semantically enhanced features and then inputs them into two decoders respectively, the two decoders decode the encoded features respectively, and output the start and end position coordinates of each cell and the start and end rows and columns of each cell respectively.
[0074] Among them, the starting and ending position coordinates of the cell include the upper left corner horizontal coordinate, the upper left corner vertical coordinate, the lower right corner horizontal coordinate and the lower right corner vertical coordinate, and the starting and ending rows and columns include the starting row, the starting column, the ending row and the ending column. Specifically, the coordinates of the upper left corner of the table can be used as the origin coordinates, and the cell starting and ending position coordinates (x0, y0, x1, y1) of each cell can be determined based on the origin coordinates, that is, the upper left corner horizontal coordinate x0, the upper left corner vertical coordinate y0, the lower right corner horizontal coordinate x1 and the lower right corner vertical coordinate y1. And take the cell in the upper left corner of the table as the 0th row and 0th column, so as to determine the starting and ending rows and columns (sr, sc, er, ec) of each cell, that is, the starting row sr, the starting column sc, the ending row er and the ending column ec. Exemplarily, for Figure 2 In the table, the cell with the content "Assets" starts at row 1, starts at column 0, ends at row 1, and ends at column 0.
[0075] Optionally, the training process of the target detection model includes: inputting the training samples into the initial target detection model to obtain the training results, determining the loss value based on the training results and the training labels based on the loss function, and iteratively correcting the initial target detection model based on the loss value to obtain the target detection model. The loss function is the sum of the loss calculations corresponding to the training results output by the two decoders in the initial target detection model. The loss function adopts the best matching idea involved in the end-to-end target detection model (Detection Transformer, DETR) and performs loss calculation based on the Hungarian algorithm.
[0076] As an optional implementation, the first encoder of the semantic recombination model can be frozen in the target detection model. When training the target detection model, in order to ensure the iterativeness of the semantic recombination model, it is adjusted in two stages to achieve joint training. Specifically, the target detection model is adjusted and optimized in the first stage, and the semantic recombination model is adjusted and optimized in the second stage.
[0077] As another optional implementation, the semantic reconstruction model may include a first encoder and a first decoder. Figure 3 This is a schematic diagram of the structure of the semantic reorganization model provided in the embodiment of the present application. Figure 3As shown, the original table text is input into the first encoder of the semantic reorganization model for encoding and text reorganization to obtain the reorganized table text features, and the reorganized table text features are input into the first decoder for decoding to obtain texts corresponding to multiple reorganized table text features. In the semantic enhancement process, the texts corresponding to multiple reorganized table text features are first encoded, and then the encoded features and the corrected table image are input into the pre-trained target detection model to obtain the start and end position coordinates and start and end rows and columns of each cell in the corrected table image.
[0078] S105. According to the start and end position coordinates and the start and end rows and columns of each cell, identify and obtain the table information in the original table image.
[0079] As an optional implementation, the table structure information corresponding to the original table image is aligned and generated according to the starting and ending position coordinates and the starting and ending rows and columns of each cell, and the table information is determined according to the original table text and the table structure information. The table structure information includes the border information of each cell in the original table image, and the border information can represent the position of the border line of each cell.
[0080] Optionally, the text segment corresponding to each cell may be determined based on the original table text, and the text segment in each cell may be filled into the table structure information to obtain the table information in the original table image.
[0081] For example, Figure 4 The table structure information and table information diagram provided in the embodiment of the present application. Figure 4 Part A, based on Figure 2 The corrected table image in can determine the table structure information. Figure 2 Taking the cell where {asset} is located in as an example, the target detection model outputs the start and end position coordinates and the start and end rows and columns of the cell where {asset} is located, where the start and end position coordinates can be (1, 0, 2, 1), and the start and end rows and columns can be (1, 0, 1, 0). According to the start and end position coordinates and the start and end rows and columns of the cell, the cell border lines that satisfy the positional logical relationship between the coordinates and the rows and columns are determined, such as Figure 4 The bold part in part a is shown in the figure. After determining the text segment corresponding to each cell based on the original table text, fill each text segment into Figure 4 From the table structure information in part a, we get Figure 4 The tabular information shown in part b.
[0082] In this embodiment, the original table image is tilted and corrected based on the affine transformation model. Since the affine transformation does not pay attention to the borders of the cells in the table image, the tilt correction has a high accuracy, is applicable to borderless tables, and has strong generalization. The original table text detected in the corrected table image is reorganized based on the semantic reorganization model, so that the cross-line text segments in the original table image are reorganized into texts in the same paragraph, thereby improving the recognition accuracy of multi-line cell text. The target detection model performs semantic enhancement and encoding and decoding processing on the corrected table image based on the reorganized table text features. Since the target detection model combines the reorganized table text features with the corrected table image to achieve semantic enhancement, the recognition of complex cell structures is further improved. After encoding and decoding processing, the table information in the original table image is obtained based on the start and end position coordinates and the start and end rows and columns of each cell in the output corrected table image, and the start and end positions of each cell are used as logical relationship coordinates such as start and end rows and columns, thereby obtaining more accurate table information.
[0083] Next, the step of performing tilt correction on the original table image based on the affine transformation model to obtain the corrected table image in the above step S101 is introduced.
[0084] Optionally, the original table image is input into an affine transformation model to obtain a target backward transformation matrix corresponding to the original table image.
[0085] Optionally, for an undeformed original table image, the output target backward transformation matrix is an identity matrix.
[0086] Optionally, the original table image is corrected according to the target backward transformation matrix to obtain a corrected table image.
[0087] Among them, the target backward transformation matrix is used to perform deformation corrections such as translation, rotation, scaling and shearing on the original table image, so as to obtain a corrected table image.
[0088] In this embodiment, the original table image is corrected according to the target backward transformation matrix output by the affine transformation model to obtain a corrected table image. Since the affine transformation model does not pay attention to the border lines of the table in the original table image, it is suitable for the correction of table images without border lines, thereby improving the accuracy and generalization of table correction.
[0089] Further, refer to Figure 5 The training process of the affine transformation model is introduced. Figure 5 A schematic diagram of the training process of the affine transformation model provided in an embodiment of the present application.
[0090] S501. Determine a plurality of matrix groups, each matrix group including a forward transformation matrix and a target backward transformation matrix, and the forward transformation matrix and the target backward transformation matrix are inverse matrices of each other.
[0091] Optionally, the more the number of matrix groups is, and the more transformation forms of the forward transformation matrix and the target backward transformation matrix in the matrix group are, the more accurate the output result of the trained affine transformation model is.
[0092] S502 : transform multiple original training images based on the forward transformation matrix in each matrix group to obtain multiple deformed intermediate training images.
[0093] The original training image is an image without deformation. After the forward transformation matrix transforms the original training image, multiple intermediate training images with deformation are obtained.
[0094] S503: Preprocess each intermediate training image to obtain a plurality of target training images.
[0095] As an optional implementation, the process of preprocessing each intermediate training image includes: normalizing the size of each intermediate training image to a first preset size, normalizing the pixel value of each intermediate training image to a value between 0 and 1, and cropping the pixel-normalized image using a center cropping method or a cropping method according to a second preset size to obtain multiple target training images. The first preset size may be, for example, 288*288, and the second preset size may be, for example, 224*224. The second preset size may be set according to the input size of the initial affine transformation model.
[0096] The purpose of cropping the pixel-normalized image using the center cropping method or the cropping method according to the second preset size is to remove the blank area around the image so that the target training image can highlight the table content.
[0097] S504, input multiple target training images into the initial affine transformation model to obtain the actual backward transformation matrix output by the initial affine transformation model, iteratively correct the initial affine transformation model according to the target backward transformation matrix in the matrix group corresponding to each target training image and the actual backward transformation matrix to obtain the affine transformation model.
[0098] Optionally, the number of actual backward transformation matrices outputted is the same as the number of target training images.
[0099] As an optional implementation, the target loss value is determined according to the target backward transformation matrix in the matrix group corresponding to each target training image and the actual backward transformation matrix, and the initial affine transformation model is iteratively corrected according to the target loss value to obtain the affine transformation model.
[0100] Optionally, multiple target training images can be divided into a training set, a validation set and a test set according to a preset ratio. After the initial affine transformation model is trained using the target training images in the training set, the trained model is fine-tuned using the validation set and the test set, respectively, to finally obtain the affine transformation model.
[0101] In this embodiment, the initial affine transformation model is iteratively corrected based on the forward transformation matrix and the target backward transformation matrix in each matrix group to obtain an affine transformation model. The forward transformation matrix and the target backward transformation matrix do not pay attention to the table border line, thereby improving the accuracy and generalization of the tilt correction of the affine transformation model.
[0102] Further, Figure 6 The following is a flow chart of obtaining an affine transformation model provided in an embodiment of the present application. Figure 6 As shown, the specific steps of iteratively correcting the initial affine transformation model according to the target backward transformation matrix in the matrix group corresponding to each target training image and the actual backward transformation matrix in the above step S504 to obtain the affine transformation model are introduced next.
[0103] S601. Determine a first loss value according to the cosine similarity between the target backward transformation matrix and the actual backward transformation matrix.
[0104] Specifically, the cosine similarity is the cosine value of the target backward transformation matrix and the actual backward transformation matrix, which can characterize the similarity between the target backward transformation matrix and the actual backward transformation matrix. The larger the cosine similarity, the more similar the target backward transformation matrix is to the actual backward transformation matrix, and the better the correction effect of the initial affine transformation model.
[0105] As an optional implementation, the exponent of the cosine value of the target backward transformation matrix and the actual backward transformation matrix is calculated, and the difference between the preset constant and the exponent is used as the first loss value. The preset constant may be 1. Since the larger the cosine similarity, the more similar the target backward transformation matrix is to the actual backward transformation matrix, the smaller the first loss value, the more similar the target backward transformation matrix is to the actual backward transformation matrix.
[0106] S602: Correct the target training image based on the actual backward transformation matrix to generate a training result image, and determine a second loss value according to the start and end height variances of each row of pixels in the training result image.
[0107] Optionally, the start and end height variance of each row of pixels may be the variance between the leftmost pixel height and the rightmost pixel height in each row of pixels. The smaller the start and end height variance is, the better the correction effect of the initial affine transformation model is.
[0108] As an optional implementation, the sum of the start and end height variances of each row of pixels in the target training image may be used as the second loss value.
[0109] S603: Determine a third loss value according to the target backward transformation matrix and the actual backward transformation matrix.
[0110] As an optional implementation, the third loss value may be a mean square error loss between the target backward transformation matrix and the actual backward transformation matrix.
[0111] As another optional implementation, the third loss value may also be the mean absolute value error loss of the target backward transformation matrix and the actual backward transformation matrix.
[0112] S604: Determine a target loss value according to the first loss value, the second loss value, and the third loss value.
[0113] Optionally, the sum of the first loss value, the second loss value and the third loss value is used as the target loss value.
[0114] S605. Iteratively correct the initial affine transformation model according to the target loss value to obtain an affine transformation model.
[0115] Optionally, the smaller the target loss value is, the more similar the actual backward transformation matrix output by the initial affine transformation model is to the target backward transformation matrix, and the better the correction effect of the initial affine transformation model is.
[0116] In this embodiment, the target loss value is determined by combining the cosine similarity of the target backward transformation matrix and the actual backward transformation matrix and the starting height variance of each row of pixels in the training result image, and the loss terms are richer, thereby improving the correction effect of the affine transformation model.
[0117] As an optional implementation, the start and end position coordinates and the start and end rows and columns of each cell in the corrected table image in the above step S104 can be determined by the following steps.
[0118] Figure 7 A schematic diagram of a process for obtaining the start and end position coordinates and the start and end rows and columns of each cell provided in an embodiment of the present application. Figure 7 As shown, the above step S104 may include:
[0119] S701, extracting features from the corrected table image to obtain multiple image features.
[0120] Specifically, a convolutional neural network (CNN) encoder based on the ViT (Vision Transformer) architecture can be used to extract features from the corrected table image to obtain multiple image features. The encoder can be a deep residual learning network.
[0121] S702: According to the similarity between the plurality of image features and the plurality of reorganized table text features, semantic enhancement is performed on each image feature to obtain a plurality of weighted image features.
[0122] Specifically, the similarities between multiple image features and multiple reorganized table text features are calculated, and then feature fusion is performed based on the similarities between the image features and each reorganized table text feature, so as to achieve semantic enhancement of each image feature and ensure the consistency between the reorganized table text feature and the image area. The similarity is used to characterize the correlation between the image feature and the reorganized table text feature.
[0123] S703 , encoding and decoding the multiple weighted image features to obtain the start and end position coordinates and the start and end rows and columns of each cell in the corrected table image.
[0124] As an optional implementation, the encoding and decoding process can be implemented by a second encoder, a second decoder, and a third decoder. Specifically, the weighted image features are input into the second encoder, the second encoder encodes the weighted image features and then inputs them into the second decoder and the third decoder respectively, the second decoder and the third decoder respectively decode the weighted image features, and respectively output the start and end position coordinates of each cell and the start and end rows and columns of each cell.
[0125] Optionally, Figure 8 The working structure diagram of the semantic reconstruction model and the target detection model provided in the embodiment of the present application is shown in FIG. Figure 8 As shown, the original table text is input into the first encoder in the semantic reorganization model to obtain the features of the reorganized table text, the corrected table image is input into the target detection model for feature extraction to obtain multiple image features, the image is semantically enhanced according to the similarity between the multiple image features and the multiple reorganized table text features to obtain multiple weighted image features, the multiple weighted image features are input into the second encoder for encoding, and then the encoding results are respectively input into the second decoder and the third decoder for decoding, the second decoder is used to output the start and end position coordinates of each cell in the corrected table image, and the third decoder is used to output the start and end rows and columns of each cell in the corrected table image.
[0126] In this embodiment, the image features are semantically enhanced according to the similarity between multiple image features and multiple reorganized table text features, thereby improving the accuracy of the output results of the target detection model based on the feature fusion of the two modalities, and because the reorganized table text features are features after text reorganization, semantic enhancement based on the reorganized table text features further improves the recognition accuracy of complex table text.
[0127] Further, refer to Fig. 9 The specific steps of obtaining multiple weighted image features in the above step S702 are introduced. Fig. 9 A schematic diagram of a process for obtaining weighted image features provided in an embodiment of the present application.
[0128] S901. Calculate the attention weight between a first image feature and each reorganized table text feature, wherein the first image feature is any image feature among the multiple image features.
[0129] Specifically, the product of each reorganized table text feature, the preset learnable parameter and the transposition of the first image feature can be used as the attention weight between the first image feature and each reorganized table text feature. Exemplarily, the attention weight between the first image feature and each reorganized table text feature can be calculated by the following formula (1):
[0130] (1)
[0131] Among them, S is the attention weight between the first image feature and each reorganized table text feature, is the text feature of each reorganized table, is the first image feature, is the transpose of the first image feature, and W is a preset learnable parameter.
[0132] S902. Determine a target attention weight of the first image feature according to the attention weights between the first image feature and each reorganized table text feature.
[0133] Specifically, the attention weights between the first image feature and each reorganized table text feature are normalized to obtain the target attention weight of the first image feature.
[0134] Exemplarily, the calculation formula (2) of the target attention weight of the first image feature is as follows:
[0135] (2)
[0136] Where A is the target attention weight of the first image feature, To normalize S.
[0137] S903. Weight the first image feature based on the target attention weight of the first image feature to obtain a weighted image feature corresponding to the first image feature.
[0138] For example, the target attention weight of the first image feature can be calculated by the following formula (3):
[0139] (3)
[0140] in, is the target attention weight of the first image feature.
[0141] Optionally, weighted image features corresponding to all image features may be calculated by the method of steps S801 to S803 above.
[0142] In this embodiment, by calculating the attention weight corresponding to the image feature to determine the weighted image feature corresponding to the image feature, the image feature can be semantically enhanced based on the reorganized table text feature to output accurate table information.
[0143] The following are the specific steps of obtaining the start and end position coordinates and the start and end rows and columns of each cell in the corrected table image in the above step S703. Fig.10 A schematic diagram of a flow chart for determining the starting and ending position coordinates and the starting and ending rows and columns of each cell provided in an embodiment of the present application.
[0144] S1001. Perform two-dimensional position encoding on a plurality of weighted image features to obtain a plurality of image feature vectors.
[0145] Specifically, multiple weighted image features are positionally encoded in the x-axis direction and positionally encoded in the y-axis direction, and then the positional encoding in the x-axis direction and the positional encoding in the y-axis direction are concatenated to obtain a one-dimensional image feature vector, wherein the image feature vector includes the positional information of the current image feature in the corrected table image.
[0146] S1002, splicing the multiple reorganized table text features into multiple image feature vectors to obtain multiple target vectors.
[0147] As an optional implementation, the reorganized table text feature is a one-dimensional feature vector, so multiple reorganized table text features can be concatenated after multiple image feature vectors to obtain multiple target vectors.
[0148] S1003, performing encoding and decoding processing according to multiple target vectors to determine the start and end position coordinates and the start and end rows and columns of each cell.
[0149] Specifically, refer to Figure 8In the structure of the target detection model, multiple target vectors are input into the second encoder for encoding, and the encoding results are input into the second decoder and the third decoder respectively. The second decoder and the third decoder respectively output the start and end position coordinates and the start and end rows and columns of each cell.
[0150] In this embodiment, by performing two-dimensional position encoding on the weighted image features and then splicing and reorganizing the table text features, semantic enhancement processing of the image features is achieved, thereby improving the accuracy of the logical coordinates of each cell output by the model.
[0151] Optionally, there are two implementations for obtaining the table information in the original table image according to the start and end position coordinates and the start and end row and column identification of each cell in the above step S105. The two implementations are respectively introduced below.
[0152] Optionally, the position coordinates of each text segment in the original table text may be the position coordinates of each text segment relative to the corrected table image.
[0153] Optionally, the position coordinates of each text segment in the original table text may be determined by an OCR algorithm.
[0154] As an optional implementation, the original table text information also includes: position coordinates of each text segment in the original table text.
[0155] Optionally, the table information in the original table image is identified based on the start and end position coordinates and the start and end rows and columns of each cell, and the position coordinates of each text segment in the original table text.
[0156] Specifically, the table structure information corresponding to the original table image can be generated by aligning each cell according to the starting and ending position coordinates and the starting and ending rows and columns of each cell. The text segment in each cell is determined according to the table structure information and the position coordinates of each text segment in the original table text, and the text segment in each cell is filled into the cell accordingly to obtain the table information in the original table image.
[0157] It is worth mentioning that, for text segments of multiple lines in the same cell, the order of each text segment can be determined according to the size of the longitudinal coordinates of each text segment, and the text segments can be filled into the corresponding cells according to the order of each text segment.
[0158] In this embodiment, the table information is determined according to the position coordinates of each text segment in the original table text, thereby obtaining the table information having the same structure and content as the table in the original table image.
[0159] Fig.11 A schematic diagram of a process for determining table information provided in an embodiment of the present application. As another optional implementation, Fig.11As shown, the method for determining the table information may also include the following steps:
[0160] S1101. Obtain multiple text segments based on multiple reorganized table text features.
[0161] Optionally, a plurality of reorganized table text features may be decoded by a decoder to obtain a plurality of text segments. It is worth mentioning that the text segments obtained after decoding are in the same paragraph in each cell.
[0162] S1102: Match the multiple text segments with the original table text to obtain the position coordinates of each text segment.
[0163] Optionally, the position coordinates of the text segments in the original table text corresponding to each text segment may be determined based on matching the text content in the text segment with the text content in the original table text, and the position coordinates may be used as the position coordinates of each text segment.
[0164] S1103. According to the starting and ending position coordinates and starting and ending rows and columns of each cell, and the position coordinates of each text segment, the table information in the original table image is identified.
[0165] Specifically, the table structure information corresponding to the original table image is determined according to the start and end position coordinates and the start and end rows and columns of each cell, and then each text segment is filled into the corresponding cell according to the position coordinates of each text segment to obtain the table information.
[0166] In this embodiment, the table information is obtained by identifying the text segment determined based on the features of the reorganized table text. Since the text segment determined based on the features of the reorganized table text is located in the same paragraph in each cell, the text in each cell in the identified table information is more coherent and will not be affected by the line breaks of the text in the cells in the original table image.
[0167] The embodiment of the present application also provides a model training method for training the aforementioned affine transformation model. The specific method steps of the model training method have been described in detail in the aforementioned embodiment, and can be referred to in the aforementioned embodiment, and will not be repeated here.
[0168] Based on the same inventive concept, a table recognition device corresponding to the table recognition method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned table recognition method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0169] Reference Fig.12 FIG. 1 is a module structure diagram of a table recognition device provided in an embodiment of the present application, wherein the device comprises:
[0170] The correction module 1201 is used to perform tilt correction on the original table image based on the pre-trained affine transformation model to obtain a corrected table image;
[0171] The detection module 1202 is used to perform text detection on the corrected table image to obtain original table text information, wherein the original table text information at least includes: original table text;
[0172] The reorganization module 1203 is used to perform text reorganization on the original table text based on the pre-trained semantic reorganization model to obtain a plurality of reorganized table text features, wherein the reorganized table text features are used to characterize the reorganized table text, and if there are cross-line text segments in the original table text, the cross-line text segments are in the same paragraph in the reorganized table text;
[0173] The enhancement module 1204 is used to input the corrected table image and the multiple reorganized table text features into a pre-trained target detection model, and the target detection model performs semantic enhancement, encoding and decoding processing on the corrected table image based on the reorganized table text features to obtain the start and end position coordinates and start and end rows and columns of each cell in the corrected table image;
[0174] The identification module 1205 is used to identify the table information in the original table image according to the start and end position coordinates and the start and end rows and columns of each cell.
[0175] As an optional implementation, the correction module 1201 is specifically configured to:
[0176] Inputting the original table image into the affine transformation model to obtain a target backward transformation matrix corresponding to the original table image;
[0177] The original table image is corrected according to the target backward transformation matrix to obtain a corrected table image.
[0178] As an optional implementation, the enhancement module 1204 is specifically configured to:
[0179] Performing feature extraction on the corrected table image to obtain a plurality of image features;
[0180] According to the similarity between the plurality of image features and the plurality of reorganized table text features, semantic enhancement is performed on each of the image features to obtain a plurality of weighted image features;
[0181] The plurality of weighted image features are encoded and decoded to obtain the start and end position coordinates and the start and end rows and columns of each cell in the corrected table image.
[0182] As an optional implementation, the enhancement module 1204 is specifically configured to:
[0183] Calculating an attention weight between a first image feature and each reorganized table text feature, wherein the first image feature is any image feature among the multiple image features;
[0184] Determining a target attention weight of the first image feature according to the attention weights between the first image feature and each reorganized table text feature;
[0185] The first image feature is weighted based on the target attention weight of the first image feature to obtain a weighted image feature corresponding to the first image feature.
[0186] As an optional implementation, the enhancement module 1204 is specifically configured to:
[0187] Performing two-dimensional position encoding on multiple weighted image features to obtain multiple image feature vectors;
[0188] splicing multiple reorganized table text features into multiple image feature vectors to obtain multiple target vectors;
[0189] Encoding and decoding processing is performed according to the plurality of target vectors to determine the start and end position coordinates and the start and end rows and columns of each cell.
[0190] As an optional implementation, the original table text information further includes: position coordinates of each text segment in the original table text;
[0191] The identification module 1205 is specifically used for:
[0192] The table information in the original table image is identified based on the start and end position coordinates and the start and end rows and columns of each cell, as well as the position coordinates of each text segment in the original table text.
[0193] As an optional implementation, the original table text information further includes: position coordinates of each text segment in the original table text;
[0194] The identification module 1205 is specifically used for:
[0195] According to the text features of the multiple reorganized tables, multiple text segments are obtained;
[0196] According to the matching of the plurality of text segments with the original table text, the position coordinates of each text segment are obtained;
[0197] The table information in the original table image is identified based on the start and end position coordinates and the start and end rows and columns of each cell, as well as the position coordinates of each text segment.
[0198] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference may be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.
[0199] Based on the same inventive concept, a model training device corresponding to the model training method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned model training method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0200] Reference Fig.13 As shown, it is a module structure diagram of a model training device provided in an embodiment of the present application, and the device includes:
[0201] A determination module 1301 is used to determine a plurality of matrix groups, each matrix group includes a forward transformation matrix and a target backward transformation matrix, and the forward transformation matrix and the target backward transformation matrix are inverse matrices of each other;
[0202] A transformation module 1302, configured to transform a plurality of original training images based on the forward transformation matrix in each of the matrix groups to obtain a plurality of deformed intermediate training images;
[0203] A preprocessing model 1303, used for preprocessing each of the intermediate training images to obtain a plurality of target training images;
[0204] The correction model 1304 is used to input multiple target training images into the initial affine transformation model to obtain the actual backward transformation matrix output by the initial affine transformation model, and iteratively correct the initial affine transformation model according to the target backward transformation matrix in the matrix group corresponding to each target training image and the actual backward transformation matrix to obtain the affine transformation model.
[0205] As an optional implementation, the correction model 1304 is specifically used to:
[0206] Determining a first loss value according to a cosine similarity between the target backward transformation matrix and the actual backward transformation matrix;
[0207] Correcting the target training image based on the actual backward transformation matrix to generate a training result image, and determining a second loss value according to the start and end height variances of each row of pixels in the training result image;
[0208] Determining a third loss value according to the target backward transformation matrix and the actual backward transformation matrix;
[0209] Determining a target loss value according to the first loss value, the second loss value, and the third loss value;
[0210] The initial affine transformation model is iteratively modified according to the target loss value to obtain an affine transformation model.
[0211] The present application embodiment also provides an electronic device 14, such as Fig.14 As shown, it is a schematic diagram of the structure of the electronic device 14 provided in the embodiment of the present application, including: a processor 141, a memory 142, and optionally, a bus 143. The memory 142 stores machine-readable instructions executable by the processor 141 (for example, Fig.12 The device comprises a correction module, a detection module, a recombination module, an enhancement module and a recognition module, and Fig.13 The execution instructions corresponding to the determination module, transformation module, preprocessing model and correction model in the device, etc.), when the electronic device 14 is running, the processor 141 communicates with the memory 142 through the bus 143, and the method steps in the aforementioned method embodiment are executed by the processor 141.
[0212] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above table recognition method are executed.
[0213] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0214] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or part of the technical solution that contributes to the prior art or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program code.
[0215] The above are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application.
Claims
1. A table recognition method, characterized in that: The method comprises: The original table image is tilt-corrected based on the pre-trained affine transformation model to obtain a corrected table image; Performing text detection on the corrected table image to obtain original table text information, wherein the original table text information at least includes: original table text; The original table text is reorganized based on the pre-trained semantic reorganization model to obtain a plurality of reorganized table text features, wherein the reorganized table text features are used to characterize the reorganized table text, and if there are cross-line text segments in the original table text, the cross-line text segments are in the same paragraph in the reorganized table text, and in the table text represented by the reorganized table text features, the cross-line text segments that were originally separated in the same cell are reorganized into texts in the same paragraph, so that the reorganized table text features are consistent with the actual semantics in the original table, wherein the paragraph is a text divided according to semantics, and one paragraph includes one or more text segments; The corrected table image and the multiple reorganized table text features are input into a pre-trained target detection model, and the target detection model performs feature extraction on the corrected table image to obtain multiple image features; semantic enhancement is performed on each of the image features according to the similarity between the multiple image features and the multiple reorganized table text features to obtain multiple weighted image features; encoding and decoding the multiple weighted image features to obtain the start and end position coordinates and the start and end rows and columns of each cell in the corrected table image; According to the start and end position coordinates and the start and end rows and columns of each cell, the table information in the original table image is identified.
2. The table recognition method according to claim 1, characterized in that: The method of performing tilt correction on the original table image based on the pre-trained affine transformation model to obtain a corrected table image includes: Inputting the original table image into the affine transformation model to obtain a target backward transformation matrix corresponding to the original table image; The original table image is corrected according to the target backward transformation matrix to obtain a corrected table image.
3. The table recognition method according to claim 1, characterized in that: According to the similarity between the plurality of image features and the plurality of reorganized table text features, semantic enhancement is performed on each of the image features to obtain a plurality of weighted image features, including: Calculating an attention weight between a first image feature and each reorganized table text feature, wherein the first image feature is any image feature among the multiple image features; Determining a target attention weight of the first image feature according to the attention weights between the first image feature and each reorganized table text feature; The first image feature is weighted based on the target attention weight of the first image feature to obtain a weighted image feature corresponding to the first image feature.
4. The table recognition method according to claim 1, characterized in that: The encoding and decoding processing of the plurality of weighted image features to obtain the start and end position coordinates and the start and end rows and columns of each cell in the corrected table image includes: Performing two-dimensional position encoding on multiple weighted image features to obtain multiple image feature vectors; splicing multiple reorganized table text features into multiple image feature vectors to obtain multiple target vectors; Encoding and decoding processing is performed according to the plurality of target vectors to determine the start and end position coordinates and the start and end rows and columns of each cell.
5. The table recognition method according to claim 1, characterized in that: The original table text information also includes: position coordinates of each text segment in the original table text; The step of identifying the table information in the original table image according to the start and end position coordinates and the start and end rows and columns of each cell includes: The table information in the original table image is identified based on the start and end position coordinates and the start and end rows and columns of each cell, as well as the position coordinates of each text segment in the original table text.
6. The table recognition method according to claim 1, characterized in that: The original table text information also includes: position coordinates of each text segment in the original table text; The step of identifying the table information in the original table image according to the start and end position coordinates and the start and end rows and columns of each cell includes: According to the text features of the multiple reorganized tables, multiple text segments are obtained; According to the matching of the plurality of text segments with the original table text, the position coordinates of each text segment are obtained; The table information in the original table image is identified based on the start and end position coordinates and the start and end rows and columns of each cell, as well as the position coordinates of each text segment.
7. A model training method, characterized in that: Used for training the affine transformation model in the table recognition method according to any one of claims 1 to 6, the method comprising: Determine a plurality of matrix groups, each matrix group includes a forward transformation matrix and a target backward transformation matrix, wherein the forward transformation matrix and the target backward transformation matrix are inverse matrices of each other; Transforming a plurality of original training images based on the forward transformation matrix in each of the matrix groups to obtain a plurality of deformed intermediate training images; Preprocessing each of the intermediate training images to obtain a plurality of target training images; Inputting a plurality of the target training images into an initial affine transformation model to obtain an actual backward transformation matrix output by the initial affine transformation model, and iteratively correcting the initial affine transformation model according to the target backward transformation matrix in the matrix group corresponding to each of the target training images and the actual backward transformation matrix to obtain the affine transformation model; Iteratively correcting the initial affine transformation model according to the target backward transformation matrix in the matrix group corresponding to each of the target training images and the actual backward transformation matrix to obtain the affine transformation model, including: Determining a first loss value according to a cosine similarity between the target backward transformation matrix and the actual backward transformation matrix; Correcting the target training image based on the actual backward transformation matrix to generate a training result image, and determining a second loss value according to the start and end height variances of each row of pixels in the training result image; Determining a third loss value according to the target backward transformation matrix and the actual backward transformation matrix; Determining a target loss value according to the first loss value, the second loss value, and the third loss value; The initial affine transformation model is iteratively modified according to the target loss value to obtain an affine transformation model.
8. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the table recognition method as described in any one of claims 1 to 6 or the steps of the model training method as described in claim 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the table recognition method according to any one of claims 1 to 6 or the steps of the model training method according to claim 7.
Citation Information
Patent Citations
Correction model training, correction and recognition method and device for non-front-view iris image
CN112651389A
Table image recognition method and device, equipment and medium
CN116912865A
Locomotive work order information intelligent identification method and system based on deep learning
CN117576699A
Information extraction method
CN118865420A