Table structure recognition method, system, device and storage medium based on graph neural network
By using a graph neural network-based method to construct table diagrams and perform feature interaction, the problem of decreased recognition accuracy caused by diverse table formats and scene changes is solved, and more efficient table structure recognition is achieved.
Patent Information
- Application Number
- CN202310891035.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-20
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2043-07-20
AI Technical Summary
Existing technologies are unable to effectively handle the problem of decreased accuracy in table structure recognition caused by diverse table formats, scene changes, and image degradation.
A graph neural network-based method is used to detect and recognize table images, construct a table graph with text rows as vertices, and use adaptive graph neural networks to perform intra-modal and inter-modal feature interaction. Graph convolutional networks and graph attention networks are combined to enhance feature representation, and finally the table structure is corrected through a classifier.
It improves the accuracy of table structure recognition, adapts to table recognition tasks in different forms and scenarios, and enhances the feature representation capability of table structure.
Smart Images

Figure CN117037201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and computers, and in particular to a table structure recognition method, system, electronic device and computer-readable storage medium based on a graph neural network. Background Art
[0002] Tables are widely present in various documents, including scientific literature, financial reports, newspapers and magazines, as well as in various scenes in daily life. The current difficulties in table structure recognition are due to the diverse forms and scenes, as well as image degradation. Tables include few-line tables, no-line tables, and color tables. They can be found in electronic documents such as PDFs, Excel spreadsheets, and scanned documents, as well as bills and invoices, and in natural scenes such as food packaging, brochures, and personal notes. All of these forms and scenes present challenges for table recognition technology. The diverse forms, changing scenes, and image degradation undoubtedly reduce table recognition accuracy.
[0003] Using graphs to model tables aligns with human intuition, and graphs are insensitive to scene changes and image degradation. Currently, most mainstream methods for recognizing table structures in images focus on image-based table structure recognition, with less research on modeling table graphs and exploring the potential of graph neural networks for table recognition. Therefore, graph neural network-based table recognition methods hold significant research significance in this area. Summary of the Invention
[0004] In order to address the deficiencies of the above-mentioned prior art, the present invention provides a table structure recognition method, system, electronic device and computer-readable storage medium based on graph neural network, which can solve the problems of table structure recognition caused by diverse table forms, scene changes, and image degradation. By effectively modeling the table graph and using graph neural network for feature enhancement, the accuracy of table structure recognition is effectively improved.
[0005] The first object of the present invention is to provide a table structure recognition method based on graph neural network.
[0006] The second object of the present invention is to provide a table structure recognition system based on graph neural network.
[0007] A third object of the present invention is to provide an electronic device.
[0008] A fourth object of the present invention is to provide a storage medium.
[0009] The first object of the present invention can be achieved by adopting the following technical solutions:
[0010] A table structure recognition method based on graph neural network, the method comprising:
[0011] Detecting and recognizing the table image to be recognized to obtain text rows and a bounding box corresponding to each text row, wherein a text row represents a row of text in a cell; and there are at least two text rows;
[0012] Constructing a tabular graph based on the text lines and the corresponding bounding boxes, wherein the text lines are vertices and the vertices have edge relationships;
[0013] Obtaining three modal features of vertices corresponding to the text line according to the text line and the corresponding bounding box;
[0014] Based on the tabular graph and the three modal features of the vertex, an adaptive graph neural network is used to perform feature interaction within each modality and between modalities to obtain an updated fusion feature of the vertex;
[0015] Obtain the features of each edge based on the updated fusion features of the two vertices corresponding to each edge in the table graph; input the features of each edge into the classifier and output the category of each edge;
[0016] According to the category of each edge, the table image is modified to obtain table structure information of the table image to be identified.
[0017] Furthermore, the adaptive graph neural network is used to perform feature interaction within and between modalities based on the table graph and the three modal features of the vertex to obtain the updated fusion features of the vertex, including:
[0018] For the three modal features, the same modal features of all vertices are stacked into a feature matrix and the edge matrix of the table graph are input into the adaptive graph neural network to complete the feature interaction within the modality and obtain the updated modal feature matrix;
[0019] The three updated modal feature matrices corresponding to the three modal features are fused, and the fused feature matrix and the edge matrix of the tabular graph are input into the adaptive graph neural network to complete the feature interaction between the modalities and obtain an updated fused feature matrix. Each row in the fused feature matrix represents the updated fused feature of each vertex; wherein the edge matrix is the set of all edges in the tabular graph.
[0020] Furthermore, the adaptive graph neural network includes a graph convolutional network and a graph attention network;
[0021] The same modal features of all vertices are stacked into a feature matrix and the edge matrix of the table graph is input into the adaptive graph neural network to complete the feature interaction within the modality and obtain the updated modal feature matrix, including:
[0022] Stacking the same modal features of all vertices into a feature matrix, inputting the feature matrix and the edge matrix of the table graph into the graph convolution network and the graph attention network respectively, to obtain a first feature matrix and a second feature matrix respectively;
[0023] The first feature matrix and the second feature matrix are fused, and the fused feature matrix is used as the updated modal feature matrix.
[0024] Furthermore, the tabular graph is constructed using a Delaunay triangulation algorithm.
[0025] Furthermore, constructing a table graph with text lines as vertices and edge relationships between vertices based on the text lines and the corresponding bounding boxes includes:
[0026] Uniformly sample the bounding boxes of the text lines, sample multiple points for each bounding box, and mark the bounding box to which each point belongs;
[0027] Apply the Delaunay triangulation algorithm to all the sampled points to obtain the Delaunay triangulation graph of all points;
[0028] Split each triangle in the Delaunay triangulation graph. Each edge of the triangle represents the edge between the two bounding boxes to which the two sampling points belong, and obtain the edge set of all sampling points.
[0029] Filter the duplicate edges and invalid edges in the edge set, and only retain one valid edge between the bounding boxes of every two text lines to obtain the valid edge set E;
[0030] Construct the vertex matrix V with the text lines as vertices and the valid edge set E as the edge matrix, and construct the table graph G = (V, E).
[0031] Furthermore, six points are sampled each time.
[0032] Furthermore, the three modal features of the vertices corresponding to the text line are obtained based on the text line and the corresponding bounding box, including:
[0033] Performing feature embedding on the spatial coordinates of the bounding box of the text line using a multi-layer perceptron to obtain spatial features of the text line;
[0034] Extracting image features of the area corresponding to the text line in the table image to be recognized to obtain image features of the text line;
[0035] Encoding the content of the text line to obtain text features of the text line;
[0036] Feature embedding is performed on the spatial features, image features and text features of the text line to obtain three modal features of the vertices corresponding to the text line.
[0037] Furthermore, the step of extracting image features of the area corresponding to the text line in the table image to be recognized to obtain the image features of the text line includes:
[0038] The deep neural network ResNet-50 is used to extract features of the table image to be identified;
[0039] According to the extracted feature image, an ROI Align algorithm based on bilinear interpolation is used to extract image features of the region corresponding to the text line, thereby obtaining image features of the text line.
[0040] The second object of the present invention can be achieved by adopting the following technical solutions:
[0041] A table structure recognition system based on graph neural network, the system comprising:
[0042] A recognition module, configured to detect and recognize the table image to be recognized, and obtain text rows and a bounding box corresponding to each text row, wherein a text row represents a row of text in a cell; and the number of text rows is at least two;
[0043] A construction module, configured to construct, based on the text lines and the corresponding bounding boxes, a table graph in which the text lines are vertices and the vertices have edge relationships;
[0044] An embedding module, configured to obtain three modal features of vertices corresponding to the text line based on the text line and the corresponding bounding box;
[0045] An interaction module is used to use an adaptive graph neural network to perform feature interaction within each modality and between modalities based on the table graph and the three modal features of the vertex to obtain an updated fusion feature of the vertex;
[0046] A classification module is used to obtain the features of each edge according to the updated fusion features of the two vertices corresponding to each edge in the table graph; input the features of each edge into the classifier and output the category of each edge;
[0047] The correction module is used to correct the table image according to the category of each edge to obtain the table structure information of the table image to be identified.
[0048] The third object of the present invention can be achieved by adopting the following technical solutions:
[0049] An electronic device includes a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the above-mentioned table structure recognition method is implemented.
[0050] The fourth object of the present invention can be achieved by adopting the following technical solutions:
[0051] A computer-readable storage medium stores a program, which, when executed by a processor, implements the above-mentioned table structure recognition method.
[0052] The present invention has the following beneficial effects compared to the prior art:
[0053] The present invention provides a table recognition method, system, device and storage medium based on graph neural network. The method detects and recognizes a table image to be recognized, obtains text lines and bounding boxes corresponding to each text line, wherein a text line represents a line of text in a cell; there are at least two text lines; based on the text lines and the corresponding bounding boxes, a table graph is constructed with text lines as vertices and with edge relationships between vertices; based on the text lines and the corresponding bounding boxes, three modal features of the vertices corresponding to the text lines are obtained; based on the table graph and the three modal features of the vertices, an adaptive graph neural network is used to perform feature interaction within and between each modality to obtain updated fused features of the vertex; based on the updated fused features of the two vertices corresponding to each edge in the table graph, the features of each edge are obtained; the features of each edge are input into a classifier to output the category of each edge; based on the category of each edge, the table graph is corrected to obtain table structure information of the table image to be recognized. By effectively constructing a table graph and using graph neural network for feature enhancement, the accuracy of table structure recognition is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0055] Figure 1 This is an overall flow chart of the table structure recognition method based on graph neural network according to embodiment 1 of the present invention;
[0056] Figure 2 The Delaunay triangulation graph G of Example 1 of the present invention is D Schematic diagram of;
[0057] FIG3( a ) is a schematic diagram of a table diagram generated by modeling based on the Delaunay triangulation algorithm according to Example 1 of the present invention, and FIG3( b ) is a schematic diagram of a table diagram after correction of FIG3( a );
[0058] Figure 4 This is a flowchart of the adaptive graph neural network according to embodiment 1 of the present invention;
[0059] Figure 5 is a table image to be recognized in Example 1 of the present invention;
[0060] Figure 6 Schematic diagram of a table structure in a table image to be recognized according to Example 1 of the present invention;
[0061] Figure 7 This is a structural block diagram of a table structure recognition system based on a graph neural network according to embodiment 2 of the present invention;
[0062] Figure 8 This is a structural block diagram of an electronic device according to embodiment 3 of the present invention. DETAILED DESCRIPTION
[0063] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. It should be understood that the specific embodiments described are only used to explain this application and are not used to limit this application.
[0064] Example 1:
[0065] like Figure 1 As shown, this embodiment provides a table structure recognition method based on a graph neural network, comprising the following steps:
[0066] S101. Identify all text rows in each cell in a table image to be identified and a bounding box corresponding to each text row, where a text row represents a line of text in a cell.
[0067] The table images to be recognized are scene images containing tables, including electronic documents such as PDF, Excel, scanned documents, etc., documents such as bills and invoices, as well as food packaging, promotional brochures, personal notes, etc. in natural scenes.
[0068] Use the OCR engine to detect and recognize the text lines of the cells in the table, and obtain all the text lines of each cell in the table and the spatial coordinates of the bounding box corresponding to each text line.
[0069] Optionally, the OCR engine is the open source project Paddle OCR, which can detect and recognize the bounding boxes and contents of all text lines in the table cells, and obtain the spatial coordinates of the bounding box corresponding to each text line and the content of the text line.
[0070] S102 : Construct a table graph based on the text lines and the corresponding bounding boxes, with the text lines as vertices and edges between the vertices.
[0071] According to the text lines and the corresponding bounding boxes, a table graph is constructed using the Delaunay triangulation algorithm to obtain a table graph with text lines as vertices and edge relationships between vertices.
[0072] In this embodiment, step S102 specifically includes:
[0073] (1) Uniformly sample the bounding boxes of the text lines, sampling 6 points for each bounding box and marking the bounding box to which each point belongs;
[0074] Specifically, two sampling points are taken on the long side of the bounding box and one on the short side to identify the four directions: up, down, left, and right. These six points are actually calculated using the four corner points of the rectangular box, for a total of ten points. Experiments have shown that using six points is similar to using more points (such as 10) in recalling valid edges, while also reducing the number of false edges.
[0075] (2) Apply the Delaunay triangulation algorithm to all the sampled points to obtain the Delaunay triangulation graph G of all points D , see Figure 2 ;
[0076] (3) Split each triangle in the Delaunay triangulation graph GD. Each edge of the triangle represents the edge between the two bounding boxes to which the two sampling points belong. The edge set E of all sampling points is obtained. D ;
[0077] (4) Filter the duplicate edges and invalid edges in the edge set, and only retain one valid edge between the bounding boxes of every two text lines to obtain the valid edge set E;
[0078] (5) Construct a vertex matrix with text lines as vertices and a valid edge set as an edge matrix, and construct a table graph G = (V, E);
[0079] By performing steps S101 and S102 on the image to be recognized, a table diagram modeled based on the Delaunay triangulation algorithm can be obtained, as shown in Figure 3(a). Each vertex with a sequence number represents a text line, and the line between each vertex represents an edge that may have a relationship between the two vertices.
[0080] This embodiment uses the Delaunay triangulation algorithm to model the table graph, thereby achieving a good balance between connectivity and sparsity in the table graph.
[0081] S103 : Obtain three modal features of vertices corresponding to the text line according to the text line and the corresponding bounding box.
[0082] The text line and the corresponding bounding box are input into the feature embedding network, and the three modal features of the vertices corresponding to the text line are output.
[0083] In this embodiment, step S103 specifically includes:
[0084] (1) Use a multi-layer perceptron to embed the spatial coordinates of the bounding box of the text line and obtain the spatial feature R of the text line. G ;
[0085] (2) The deep neural network ResNet-50 is used to extract features of the table image and the ROI Align algorithm based on bilinear interpolation is used to extract the image features of the area corresponding to the text line to obtain the image features R of the text line. I ;
[0086] (3) Encode the content of the text line using the pre-trained text embedding model Sentence Transformer to obtain the text feature R of the text line C ;
[0087] Spatial features, image features, and text features are the three modal features of a text line, that is, the three modal features of the vertices corresponding to the text line.
[0088] S104. Design an adaptive graph neural network that combines graph convolutional networks and graph attention networks.
[0089] In this embodiment, if Figure 4 As shown, step S104 specifically includes:
[0090] S1041. Using the modal feature matrix R and the edge matrix E of the table graph as input, the updated feature matrix R is obtained through the graph convolutional network. GCN ;
[0091] Specifically, the graph convolutional network adopts a two-layer convolution structure. The input data is processed in sequence by the graph convolution module, regularization module, activation function, graph convolution module, regularization module, and activation function to obtain the updated feature matrix R GCN ;
[0092] Among them, the graph convolution module formula is:
[0093]
[0094]
[0095] Among them, σ is the activation function, N iis the adjacent vertex of vertex i, c ij is the normalized weight, which depends on the number of neighbors of vertex i; w (l) is a learnable matrix; is the feature vector of vertex i output by the l-th layer network; R (l+1) is the feature matrix output by the l+1th layer network, which is formed by stacking the feature vectors of all vertices output by the l+1th layer network; R (0) Represents the feature matrix R, as the input of the first layer, Represents R (0) The i-th row of is the eigenvector of vertex i.
[0096] S1042: Using the modal feature matrix R and the edge matrix E of the table graph as input, the updated feature matrix R is obtained through the graph attention network. GAT ;
[0097] Specifically, the graph attention network adopts a two-layer attention structure. The input data is processed in sequence by the graph attention module, regularization module, activation function, graph attention module, regularization module, and activation function to obtain the updated feature matrix R GAT ;
[0098] Among them, the graph attention module formula is:
[0099]
[0100]
[0101] Among them, σ is the activation function, N i is the adjacent vertex of vertex i, α ij is the weight coefficient of vertex i to vertex j, which is calculated by the features of vertex i and vertex j, w (l) is a learnable matrix; is the feature vector of vertex i output by the l-th layer network, R (l+1) is the feature matrix output by the l+1th layer network, which is formed by stacking the feature vectors of all vertices output by the l+1th layer network; R (0) Represents the feature matrix R, as the input of the first layer, Represents R (0) The i-th row of is the eigenvector of vertex i.
[0102] S1043. The feature matrix R′ obtained by fusing the graph convolutional network branch and the graph attention network branch through the learnable weight α is used as the output of the adaptive graph neural network. The formula is expressed as:
[0103] R′=α·R GCN +(1-α)·R GAT ;
[0104] In this embodiment, an adaptive graph neural network is adopted, which combines the advantages of smoothness and standardization of the graph convolutional network and the advantages of the graph attention network in capturing global correlation and independence. This enables the adaptive graph neural network to adapt to the table structure recognition task through learning and enhances the feature representation ability of vertex features.
[0105] S105. Based on the tabular graph and the vertex features of each modality, an adaptive graph neural network is used to perform feature interaction within each modality and between modalities to obtain updated fusion features of each vertex.
[0106] In this embodiment, step S105 specifically includes:
[0107] (1) The spatial features of all vertices are stacked into a matrix, and the stacked matrix is input into the adaptive graph neural network as the feature matrix to complete the feature interaction within the spatial modality and obtain the updated spatial feature matrix R G′ ;
[0108] (2) The image features of all vertices are stacked into a matrix, and the stacked matrix is input into the adaptive graph neural network as the feature matrix to complete the feature interaction within the image modality and obtain the updated image feature matrix R I′ ;
[0109] (3) The text features of all vertices are stacked into a matrix, and the stacked matrix is input into the adaptive graph neural network as the feature matrix to complete the feature interaction within the text modality and obtain the updated text feature matrix R C′ ;
[0110] (4) The updated feature matrix is fused by the learnable weights β, γ, and δ to obtain the vertex fusion feature matrix R Fusion :
[0111] R Fusion =β·R G′ +γ·R I′ +δ·R C′
[0112] (5) The fusion feature matrix of the vertex is used as the feature matrix to input into the adaptive graph neural network to complete the feature interaction between the three modalities and obtain the updated fusion feature matrix R Fusion′ .
[0113] Each row in the updated fused feature matrix represents the updated fused feature of each vertex.
[0114] This embodiment enhances the intra-modal and inter-modal collaboration by making full use of the modal information provided by the table and using graph neural networks to perform intra-modal and inter-modal feature interaction on the spatial features, image features, and text features of the table, thereby effectively enhancing the feature representation of each modality at the table vertices.
[0115] S106. Obtain the features of each edge based on the updated fusion features of the two vertices of each edge; use the edge features as input, classify each edge using a classifier, and modify the table graph based on the edge classification results to obtain table structure information.
[0116] First, the network consisting of the feature embedding network, the adaptive graph neural network, and the classifier is trained: gradient backpropagation is performed based on the classification results and the parameters in the feature embedding network, the adaptive graph neural network, and the classifier are updated;
[0117] During the testing phase, the trained network is used to output the category of each edge of the table graph in the table image to be identified, and the table graph is corrected according to the edge classification results to directly obtain the table structure information.
[0118] In this embodiment, step S106 specifically includes:
[0119] (1) For the edge matrix of the table graph, combined with the updated fusion feature matrix R Fusion′ , the edge e ij The characteristic representation of Get the characteristics of each edge in the edge matrix; Represent the updated fusion features of vertex i and j respectively;
[0120] (2) The edge feature matrix E is used as input to the classifier for classification, and the output category of each edge is: same row, same column, same cell, or no relationship;
[0121] (3) During the training phase, the cross entropy is used as the loss function to calculate the loss between the classification results and the annotations, and gradient backpropagation is performed to update the parameters of the feature embedding network, the adaptive graph neural network, and the edge classifier. During the testing phase, the table graph constructed in step S102 is modified based on the edge classification results to obtain table structure information, as shown in Figure 3(b). During the testing phase, edges classified as irrelevant are deleted from the table graph to obtain accurate table graph information.
[0122] In this embodiment, the public table recognition dataset SciTSR is used to train the network composed of the feature embedding network, the adaptive graph neural network and the edge classifier, and the parameters in the feature embedding network, the adaptive graph neural network and the edge classifier are updated.
[0123] In terms of task definition, this embodiment defines the table recognition task as an edge classification task in the field of graph neural networks, and uses a classifier to classify the relationship between text lines into the same row, the same column, the same cell, or no relationship.
[0124] like Figure 5 For a rendering of an image to be recognized, after executing steps S101 to S106, a corrected table diagram is obtained, which is the actual structure of the table. Figure 6 As shown, the table has a structure of 10 rows and 2 columns, where each vertex with a sequence number represents a text line, the horizontal lines between vertices are edges classified into the same row, and the vertical lines are edges classified into the same column. It is worth noting that Figure 6 The vertical line between the two vertices numbered 17 and 18 is classified as an edge of the same cell because Figure 5 The two text lines in column 2 and row 9 belong to the same cell.
[0125] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above embodiments may be completed by instructing related hardware through a program, and the corresponding program may be stored in a computer-readable storage medium.
[0126] It should be noted that although the method operations of the above embodiments are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all of the illustrated operations must be performed to achieve the desired results. Rather, the depicted steps may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into a single step, and / or a single step may be broken down into multiple steps.
[0127] Example 2:
[0128] like Figure 7 As shown, this embodiment provides a table structure recognition system based on a graph neural network. The system includes a recognition module 701, a construction module 702, an embedding module 703, an interaction module 704, a classification module 705, and a correction module 706. The specific functions of each module are as follows:
[0129] Recognition module 701, configured to detect and recognize the table image to be recognized, and obtain text rows and bounding boxes corresponding to each text row, wherein a text row represents a row of text in a cell; and there are at least two text rows;
[0130] A construction module 702 is configured to construct a table graph with text lines as vertices and edges between vertices based on the text lines and the corresponding bounding boxes;
[0131] An embedding module 703 is configured to obtain three modal features of vertices corresponding to the text line based on the text line and the corresponding bounding box;
[0132] An interaction module 704 is configured to utilize an adaptive graph neural network to perform feature interaction within and between modalities based on the table graph and the three modal features of the vertex to obtain an updated fusion feature of the vertex;
[0133] The classification module 705 is configured to obtain the features of each edge based on the updated fusion features of the two vertices corresponding to each edge in the table graph; input the features of each edge into a classifier, and output the category of each edge;
[0134] The correction module 706 is used to correct the table image according to the category of each edge to obtain table structure information of the table image to be identified.
[0135] The specific implementation of each module in this embodiment can be found in the above-mentioned embodiment 1, and will not be described one by one here; it should be noted that the device provided in this embodiment is only illustrated by the division of the above-mentioned functional modules. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.
[0136] Example 3:
[0137] This embodiment provides an electronic device, which may be a computer or a server, etc. Figure 8 As shown, the system includes a processor 802, a memory, an input device 803, a display 804, and a network interface 805 connected via a system bus 801. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 806 and an internal memory 807. The non-volatile storage medium 806 stores an operating system, a computer program, and a database. The internal memory 807 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 802 executes the computer program stored in the memory, the table structure recognition method of the above-mentioned embodiment 1 is implemented as follows:
[0138] Detecting and recognizing the table image to be recognized to obtain text rows and a bounding box corresponding to each text row, wherein a text row represents a row of text in a cell; and there are at least two text rows;
[0139] Constructing a tabular graph based on the text lines and the corresponding bounding boxes, wherein the text lines are vertices and the vertices have edge relationships;
[0140] Obtaining three modal features of vertices corresponding to the text line according to the text line and the corresponding bounding box;
[0141] Based on the tabular graph and the three modal features of the vertex, an adaptive graph neural network is used to perform feature interaction within each modality and between modalities to obtain an updated fusion feature of the vertex;
[0142] Obtain the features of each edge based on the updated fusion features of the two vertices corresponding to each edge in the table graph; input the features of each edge into the classifier and output the category of each edge;
[0143] According to the category of each edge, the table image is modified to obtain table structure information of the table image to be identified.
[0144] Example 4:
[0145] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the table structure recognition method of the above embodiment 1 is implemented as follows:
[0146] Detecting and recognizing the table image to be recognized to obtain text rows and a bounding box corresponding to each text row, wherein a text row represents a row of text in a cell; and there are at least two text rows;
[0147] Constructing a tabular graph based on the text lines and the corresponding bounding boxes, wherein the text lines are vertices and the vertices have edge relationships;
[0148] Obtaining three modal features of vertices corresponding to the text line according to the text line and the corresponding bounding box;
[0149] Based on the tabular graph and the three modal features of the vertex, an adaptive graph neural network is used to perform feature interaction within each modality and between modalities to obtain an updated fusion feature of the vertex;
[0150] Obtain the features of each edge based on the updated fusion features of the two vertices corresponding to each edge in the table graph; input the features of each edge into the classifier and output the category of each edge;
[0151] According to the category of each edge, the table image is modified to obtain table structure information of the table image to be identified.
[0152] It should be noted that the computer-readable storage medium of the present embodiment may be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0153] In summary, the present invention provides a table structure recognition method, system, electronic device and computer-readable storage medium based on graph neural network. By using the Paddle OCR engine to detect and recognize the text lines of the image to be recognized, it is convenient to effectively extract the text line information in the table unit; by modeling the text lines into a table graph based on the Delaunay triangulation algorithm, a good balance between connectivity and sparsity in the table graph is achieved, which is conducive to the next step of vertex feature interaction; through the adaptive graph neural network, the graph neural network can adapt to the table structure recognition task through learning, and enhance the feature representation ability of vertex features; through the graph neural network for multimodal collaboration and fusion, the feature representation of each mode of the table is effectively enhanced; through the classifier to classify the modeled edges, the application of the graph neural network in table structure recognition is realized, the accuracy of table structure recognition is improved, and it is suitable for promotion and application. The present invention mainly solves the problems caused by the diversity of table forms, scene changes, and image degradation in table structure recognition, fills the research gap in how to effectively model table graphs and how to use graph neural networks for feature enhancement in the prior art, and has important research significance in the field of table recognition.
[0154] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A table structure recognition method based on graph neural network, characterized in that: The method comprises: Detecting and recognizing the table image to be recognized to obtain text rows and a bounding box corresponding to each text row, wherein a text row represents a row of text in a cell; and there are at least two text rows; Constructing a tabular graph based on the text lines and the corresponding bounding boxes, wherein the text lines are vertices and the vertices have edge relationships; Obtaining three modal features of vertices corresponding to the text line according to the text line and the corresponding bounding box; Based on the tabular graph and the three modal features of the vertex, an adaptive graph neural network is used to perform feature interaction within each modality and between modalities to obtain an updated fusion feature of the vertex; Obtain the features of each edge based on the updated fusion features of the two vertices corresponding to each edge in the table graph; input the features of each edge into the classifier and output the category of each edge; According to the category of each edge, the table image is modified to obtain table structure information of the table image to be identified; The step of constructing a table graph having text lines as vertices and edge relationships between vertices based on the text lines and the corresponding bounding boxes includes: Uniformly sample the bounding boxes of the text lines, sample multiple points for each bounding box, and mark the bounding box to which each point belongs; Apply the Delaunay triangulation algorithm to all the sampled points to obtain the Delaunay triangulation graph of all points; Split each triangle in the Delaunay triangulation graph. Each edge of the triangle represents the edge between the two bounding boxes to which the two sampling points belong, and obtain the edge set of all sampling points. Filter the duplicate edges and invalid edges in the edge set, and only retain one valid edge between the bounding boxes of every two text lines to obtain the valid edge set E ; Construct a vertex matrix with text lines as vertices V , valid edge set E For the edge matrix, construct a table graph G =( V , E ); The method uses an adaptive graph neural network to perform feature interaction within and between modalities based on the table graph and the three modal features of the vertex to obtain the updated fusion features of the vertex, including: For the three modal features, the same modal features of all vertices are stacked into a feature matrix and the edge matrix of the table graph are input into the adaptive graph neural network to complete the feature interaction within the modality and obtain the updated modal feature matrix; The three updated modal feature matrices corresponding to the three modal features are fused, and the fused feature matrix and the edge matrix of the tabular graph are input into the adaptive graph neural network to complete the feature interaction between the modalities and obtain an updated fused feature matrix. Each row in the fused feature matrix represents the updated fused feature of each vertex; wherein the edge matrix is the set of all edges in the tabular graph.
2. The table structure recognition method according to claim 1, characterized in that: The adaptive graph neural network includes a graph convolutional network and a graph attention network; The same modal features of all vertices are stacked into a feature matrix and the edge matrix of the table graph is input into the adaptive graph neural network to complete the feature interaction within the modality and obtain the updated modal feature matrix, including: Stacking the same modal features of all vertices into a feature matrix, inputting the feature matrix and the edge matrix of the table graph into the graph convolution network and the graph attention network respectively, to obtain a first feature matrix and a second feature matrix respectively; The first feature matrix and the second feature matrix are fused, and the fused feature matrix is used as the updated modal feature matrix.
3. The table structure recognition method according to claim 1, characterized in that: Six points were sampled at each location.
4. The table structure recognition method according to any one of claims 1 to 2, characterized in that: The method of obtaining three modal features of vertices corresponding to the text line based on the text line and the corresponding bounding box includes: Performing feature embedding on the spatial coordinates of the bounding box of the text line using a multi-layer perceptron to obtain spatial features of the text line; Extracting image features of the area corresponding to the text line in the table image to be recognized to obtain image features of the text line; Encoding the content of the text line to obtain text features of the text line; Feature embedding is performed on the spatial features, image features and text features of the text line to obtain three modal features of the vertices corresponding to the text line.
5. The table structure recognition method according to claim 4, characterized in that: The step of extracting image features of an area corresponding to the text line in the table image to be recognized to obtain image features of the text line includes: The deep neural network ResNet-50 is used to extract features of the table image to be identified; According to the extracted feature image, an ROI Align algorithm based on bilinear interpolation is used to extract image features of the region corresponding to the text line, thereby obtaining image features of the text line.
6. A table structure recognition system based on graph neural network, characterized in that: The system comprises: A recognition module, configured to detect and recognize the table image to be recognized, and obtain text rows and a bounding box corresponding to each text row, wherein a text row represents a row of text in a cell; and the number of text rows is at least two; A construction module, configured to construct, based on the text lines and the corresponding bounding boxes, a table graph in which the text lines are vertices and the vertices have edge relationships; An embedding module, configured to obtain three modal features of vertices corresponding to the text line based on the text line and the corresponding bounding box; An interaction module is used to use an adaptive graph neural network to perform feature interaction within each modality and between modalities based on the table graph and the three modal features of the vertex to obtain an updated fusion feature of the vertex; A classification module is used to obtain the features of each edge according to the updated fusion features of the two vertices corresponding to each edge in the table graph; input the features of each edge into the classifier and output the category of each edge; A correction module, configured to correct the table image according to the category of each edge, and obtain table structure information of the table image to be identified; The construction module is specifically used to: uniformly sample the bounding boxes of the text lines, sample multiple points in each bounding box, and mark the bounding box to which each point belongs; apply the Delaunay triangulation algorithm to all the sampled points to obtain the Delaunay triangulation graph of all the points; split each triangle in the Delaunay triangulation graph, and each edge of the triangle represents the edge between the two bounding boxes to which the two sampling points belong, to obtain the edge set of all the sampling points; filter the duplicate edges and invalid edges in the edge set, and only retain one valid edge between the bounding boxes of every two text lines to obtain the valid edge set. E ; Build a vertex matrix with text lines as vertices V , valid edge set E For the edge matrix, construct a table graph G =( V , E ); The interaction module is specifically used to: for three modal features, stack the same modal features of all vertices into a feature matrix and the edge matrix of the tabular graph, input them into the adaptive graph neural network, complete the feature interaction within the modality, and obtain an updated modal feature matrix; fuse the three updated modal feature matrices corresponding to the three modal features, input the fused feature matrix and the edge matrix of the tabular graph into the adaptive graph neural network, complete the feature interaction between the modalities, and obtain an updated fused feature matrix, where each row in the fused feature matrix represents the updated fused feature of each vertex; wherein the edge matrix is the set of all edges in the tabular graph.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the table structure recognition method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Apparatus and method for identifying textual image of structured layout
CN111492370A
Automatic delineation and extraction of tabular data in portable document format using graph neural networks
US20220180044A1