Table Structure Recognition Method, Device, Computer Equipment, and Storage Medium

By identifying and fusing the text area features of PDF tables, generating adjacency features and making predictions, the problem of inefficient recognition in traditional methods is solved, and efficient and accurate recognition of PDF tables is achieved.

CN114332893BActive Publication Date: 2025-07-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111020622.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-01
Publication Date
2025-07-01
Estimated Expiration
2041-09-01

AI Technical Summary

Technical Problem

The prior art cannot effectively identify blank fields in PDF format tables, and traditional methods require additional text detection networks, resulting in inefficient table recognition.

Method used

By obtaining the target table image area, identifying the text area and determining the image features and coordinate features, the text area element fusion features are generated after fusion, the adjacency features are determined, and feature stitching and classification prediction are performed, row-and-sequence relationship prediction results are generated, and the table structure is determined.

Benefits of technology

The overall recognition of PDF tables is realized, which improves the recognition accuracy and efficiency, reduces additional operations, and improves the accuracy and efficiency of table recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332893B_ABST
    Figure CN114332893B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, and storage medium for table structure recognition. The method includes: obtaining a target table image region, identifying text regions in the target table image region, determining the image features and coordinate features of each text region, fusing the image features and coordinate features to obtain the corresponding fused features of text region elements. According to the fused features of text region elements, determining the adjacency features of each node in the target table image region, performing feature splicing on the adjacency features of any two nodes, classifying and predicting the spliced adjacency matrix to generate a prediction result of the row-column relationship of the text regions corresponding to the two nodes, and based on the prediction result of the row-column relationship, determining the table structure corresponding to the target table image region. Using this method can achieve the overall recognition of the target table image region, determine the corresponding table structure according to the prediction result of the row-column relationship, and improve the accuracy and efficiency of table recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method, device, computer device, and storage medium for recognizing a table structure. Background Art

[0002] With the development of artificial intelligence technology, and the increasing requirements for the efficiency and accuracy of data information extraction, collation, and update, as a storage form of structured data, a table has the characteristic of standardization, which is more convenient for users to query, extract, or update and enter the data stored in the table. However, currently, it is usually published after converting the table into the PDF format, resulting in the inability to directly extract the data in the table or update the table. Therefore, there has emerged a technology for recognizing the structure and content of PDF-formatted tables.

[0003] Traditionally, for table recognition methods, text detection is mostly first performed on a PDF file to obtain the text regions in the image, which may include different text regions involved in the image. Then, a graph neural network is used to predict the relationship between each two text regions. According to the relationship between each two text regions, it is determined whether the corresponding text regions need to be merged or not. Finally, post-processing is performed on the predicted adjacency matrix to reproduce the table structure in the image, and then the content in the table is recognized.

[0004] However, traditional table recognition methods cannot directly solve the scenario where there are blank fields in the table, and the predicted adjacency matrix can only represent whether the text regions are merged, only considering the characteristics of domain nodes, and cannot cover the entire table to be recognized. An additional text detection network is required to locate the text positions in the image and then further organize them into row and column information. Therefore, traditional table recognition methods cannot perform an overall and global recognition of the table to be recognized, and an additional corresponding text detection network needs to be set, which is prone to problems of misidentifying the recognized content, resulting in relatively low table recognition efficiency. Summary of the Invention

[0005] Based on this, it is necessary to provide a method, device, computer device, and storage medium for recognizing a table structure that can perform an overall and comprehensive recognition of a PDF table to improve the accuracy and efficiency of table recognition for the above technical problems.

[0006] A method for recognizing a table structure, characterized in that the method includes:

[0007] Obtain a target table image region, and recognize the text regions in the target table image region;

[0008] Determine the image features and coordinate features of each of the text regions, and respectively fuse the image features and coordinate features to obtain the text region element fusion features corresponding to each of the text regions;

[0009] Determine the adjacency features of each node in the target table image region according to the text region element fusion features;

[0010] Perform feature splicing on the adjacency features of any two nodes, and perform classification prediction on the spliced adjacency matrix to generate the row-column relationship prediction result of the text regions corresponding to the two nodes;

[0011] Based on the row-column relationship prediction results of each of the text regions, determine the table structure corresponding to the target table image region.

[0012] In one embodiment, the generating the adjacency features of each node in the target table image region by performing feature aggregation based on the local features and the global features includes:

[0013] Obtain each gate parameter corresponding to the gate mechanism;

[0014] Based on a preset activation function and each of the gate parameters, perform feature aggregation on the local features and the global features to obtain the adjacency features of each node in the target table image region.

[0015] A table structure recognition device, characterized in that the device includes:

[0016] A text region recognition module, configured to obtain a target table image region and recognize the text regions in the target table image region;

[0017] A text region element fusion feature generation module, configured to determine the image features and coordinate features of each of the text regions, and respectively fuse the image features and coordinate features to obtain the text region element fusion features corresponding to each of the text regions;

[0018] An adjacency feature generation module, configured to determine the adjacency features of each node in the target table image region according to the text region element fusion features;

[0019] A row-column relationship prediction result generation module, configured to perform feature splicing on the adjacency features of any two nodes, and perform classification prediction on the spliced adjacency matrix to generate the row-column relationship prediction result of the text regions corresponding to the two nodes;

[0020] A table structure determination module, configured to determine the table structure corresponding to the target table image region based on the row-column relationship prediction results of each of the text regions.

[0021] A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0022] Obtain a target table image area, and identify text areas in the target table image area;

[0023] Determine the image features and coordinate features of each text area, and respectively fuse the image features and coordinate features to obtain text area element fusion features corresponding to each text area;

[0024] According to the text area element fusion features, determine the adjacency features of each node in the target table image area;

[0025] Perform feature splicing on the adjacency features of any two nodes, and perform classification prediction on the spliced adjacency matrix to generate a prediction result of the row-column relationship of the text areas corresponding to the two nodes;

[0026] Based on the prediction results of the row-column relationships of each text area, determine the table structure corresponding to the target table image area.

[0027] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0028] Obtain a target table image area, and identify text areas in the target table image area;

[0029] Determine the image features and coordinate features of each text area, and respectively fuse the image features and coordinate features to obtain text area element fusion features corresponding to each text area;

[0030] According to the text area element fusion features, determine the adjacency features of each node in the target table image area;

[0031] Perform feature splicing on the adjacency features of any two nodes, and perform classification prediction on the spliced adjacency matrix to generate a prediction result of the row-column relationship of the text areas corresponding to the two nodes;

[0032] Based on the prediction results of the row-column relationships of each text area, determine the table structure corresponding to the target table image area.

[0033] In the above table structure recognition method, device, computer device, and storage medium, by obtaining the target table image area, recognizing the text areas in the target table image area, and determining the image features and coordinate features of each text area, the image features and coordinate features are respectively fused to obtain the text area element fusion features corresponding to each text area. The overall recognition of the target table image area can be achieved by fusing the image features and coordinate features of different text areas within the target table image area, rather than the local recognition of a single text area. Furthermore, according to the text area element fusion features, the adjacency features of each node in the target table image area are determined, and the adjacency features of any two nodes are feature - stitched. By classifying and predicting the stitched adjacency matrix, the row - column relationship prediction results of the text areas corresponding to the two nodes are generated. Then, based on the row - column relationship prediction results of each text area, the table structure corresponding to the target table image area is determined. It is realized that according to the prediction results of the row - column relationships of each text area, the corresponding table structure can be determined without further recognition using an additional text detection network, which can reduce unnecessary cumbersome operations and thus improve the table recognition accuracy and recognition efficiency. Brief Description of the Drawings

[0034] Figure 1 It is an application environment diagram of the table structure recognition method in an embodiment;

[0035] Figure 2 It is a flowchart of the table structure recognition method in an embodiment;

[0036] Figure 3 It is a schematic diagram of the target table image area of the table structure recognition method in an embodiment;

[0037] Figure 4 It is a schematic diagram of the text area detection result of the table structure recognition method in an embodiment;

[0038] Figure 5 It is a schematic diagram of the row relationship prediction result of the table structure recognition method in an embodiment;

[0039] Figure 6 It is a schematic diagram of the column relationship prediction result of the table structure recognition method in an embodiment;

[0040] Figure 7 It is a flowchart of obtaining the text area element fusion features corresponding to each text area in an embodiment;

[0041] Figure 8 It is a flowchart of obtaining the local features of the nodes corresponding to the text area element fusion features in an embodiment;

[0042] Figure 9Schematic diagram of the overall process of the table structure recognition method in an embodiment;

[0043] Figure 10 Schematic diagram of the FLAG network structure for generating adjacent features in an embodiment;

[0044] Figure 11 Block diagram of the structure of the table structure recognition device in an embodiment;

[0045] Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0046] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] The table structure recognition method provided by the present application involves artificial intelligence technology. Among them, artificial intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Among them, artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0048] Among them, machine learning (ML) in artificial intelligence software technology is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0049] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, and smart classrooms. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0050] The table structure recognition method provided in the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. Among them, the server 104 obtains the target table image area, identifies the text areas in the target table image area, determines the image features and coordinate features of each text area, fuses the image features and coordinate features respectively, and obtains the text area element fusion features corresponding to each text area. Furthermore, the server 104 determines the adjacency features of each node in the target table image area according to the text area element fusion features, splices the adjacency features of any two nodes, and classifies and predicts the spliced adjacency matrix to generate the row-column relationship prediction result of the text areas corresponding to the two nodes. Furthermore, based on the row-column relationship prediction results of each text area, the table structure corresponding to the target table image area is determined, and the corresponding table structure is fed back to the terminal 102. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0051] In one embodiment, as Figure 2 shown, a table structure recognition method is provided. Taking the case where this method is applied to the Figure 1 server as an example, it includes the following steps:

[0052] Step S202, obtain the target table image area and identify the text areas in the target table image area.

[0053] Specifically, by performing object detection on the image to be recognized, the target table image area corresponding to the image to be recognized is obtained, and then the target table image area is further recognized to identify the text area in the target table image area.

[0054] In one embodiment, as Figure 3 shown, a target table image area of a table structure recognition method is provided. By performing object detection on the image to be recognized, a target table image area as Figure 3 shown is obtained, and then the target table image area is further recognized, and the text area detection result of the table structure recognition method as Figure 4 shown can be recognized.

[0055] Specifically, by using the Mask-RCNN network to perform object detection on the target table area, the text area within the target table area is determined. Among them, the text area can be represented by the text area detection result as Figure 4 shown, that is, within the target table area, there are text areas corresponding to each text box as Figure 4 shown, and then by using the Mask-RCNN network to perform object detection on the target table area, the positions of different text boxes can be determined.

[0056] Furthermore, the Mask-RCNN network represents a network that is compatible with general object detection and segmentation tasks, including a detection branch and a segmentation branch. In this embodiment, since only the text area needs to be recognized, only the detection branch of the Mask-RCNN network is used to detect the target table area.

[0057] Among them, the backbone network of the Mask-RCNN network is the Res50 network with FPN. Among them, FPN represents the Feature Pyramid Network, which is a neural network that can improve the object detection effect in a multi-scale manner, and the Res50 network represents a deep residual network with 50 layers, which belongs to the basic network type of convolutional neural networks. Among them, feature maps at different stages of the picture can be obtained through the Res50 network, and then a feature pyramid is established according to the feature maps at different stages, that is, the Res50 network with FPN is obtained.

[0058] In one embodiment, since there are still many redundant text regions in the RPN (Region Proposal Network) prediction results obtained after object detection by the Res50 network (Deep Residual Network) of the FPN (Feature Pyramid Network), the NMS algorithm is then used to filter all the text regions to filter out the redundant text regions, thereby reducing the computational complexity. Among them, the NMS algorithm (Non-Maximum Suppression) is the non-maximum suppression algorithm, which is used for searching for local maxima in the field of object detection and can filter the data values that do not meet the maximum requirements.

[0059] Step S204: Determine the image features and coordinate features of each text region, and fuse the image features and coordinate features respectively to obtain the text region element fusion features corresponding to each text region.

[0060] Specifically, by obtaining the position coordinates of each text region determined from the target table image region and performing dimension elevation on the position coordinates of each text region, the dimension-elevated coordinate features can be obtained. Further, the image content of the corresponding text region can be obtained according to the position coordinates of each text region, and then image feature alignment is performed based on the image content of the text region to obtain the aligned image features. Among them, the dimension of the aligned image features is the same as the dimension of the dimension-elevated coordinate features.

[0061] Further, by fusing the dimension-elevated coordinate features and the aligned image features, the text region element fusion features corresponding to each text region are obtained.

[0062] Step S206: Determine the adjacency features of each node in the target table image region according to the text region element fusion features.

[0063] Specifically, by obtaining the local features and global features of the nodes corresponding to each text region element fusion feature, and then performing feature aggregation based on the local features and global features, the adjacency features of each node in the target table image region are obtained.

[0064] Among them, by using the k-Nearest Neighbor algorithm, K nearest neighbors of the node corresponding to the fusion feature of each text region element are respectively determined, and the K nearest neighbors and the fusion feature of the text region element of this node are fused to obtain the adjacent fusion feature of each node. By integrating the adjacent fusion features of each node, the corresponding aggregation feature is obtained. Further, a FCN network (Full Connected Network, that is, a fully connected neural network) is used to perform dimensionality reduction processing on the aggregation feature of each node to obtain the dimensionality-reduced multi-head graph feature. Among them, the dimensionality-reduced multi-head graph feature is the local feature of the node corresponding to the fusion feature of the text region element.

[0065] Similarly, according to the multi-head attention mechanism, context feature aggregation is performed on the fusion feature of the text region element corresponding to each node to obtain the global feature of the node corresponding to the fusion feature of the text region element.

[0066] Further, by obtaining the preset activation function and each gate parameter corresponding to the AND gate mechanism, and then based on the preset activation function and each gate parameter, feature aggregation is performed on the local feature and the global feature to obtain the adjacent feature of each node in the target table image region.

[0067] In one embodiment, the following formula (1) is used to aggregate the local feature and the global feature to obtain the adjacent feature of each node in the target table image region:

[0068] F agg = Sigmoid(gate i ) * F global + (1 - Sigmoid(gate i )) * F local , i ∈ [1, 2,.., N]; (1)

[0069] Among them, F agg represents the aggregated adjacent feature, F global represents the global feature, F local represents the local feature, Sigmoid is the preset activation function, and gate i represents the gate parameter on the i-th head. The head represents different attention points in the Multi-head attention, and the number of heads is the same as the number of gates of the gate mechanism.

[0070] Step S208, perform feature splicing on the adjacent features of any two nodes, and perform classification prediction on the spliced adjacency matrix to generate a prediction result of the row-column relationship of the text region corresponding to these two nodes.

[0071] Specifically, by splicing the adjacency features of any two nodes, an adjacency matrix after splicing is obtained. Then, according to a fully connected neural network, binary classification prediction is performed on the adjacency matrix after splicing to obtain the prediction result of the row-column relationship of the corresponding text region.

[0072] Among them, the binary classification prediction includes row relationship prediction and column relationship prediction. Specifically, according to the fully connected network, binary classification prediction is performed on the adjacency matrix after splicing to determine whether the two text regions corresponding to the spliced adjacency matrix belong to the same row in the table, or to determine whether the two text regions corresponding to the spliced adjacency matrix belong to the same column in the table, and then the prediction result of the row-column relationship of the corresponding text region is obtained.

[0073] In one embodiment, as Figure 5 and Figure 6 shown, a schematic diagram of the row relationship prediction result of the table structure recognition method and a schematic diagram of the column relationship prediction result of the table structure recognition method are respectively provided. Referring to Figure 5 it can be seen that the row relationship prediction result can determine the text regions corresponding to each row in the table. In Figure 5 , text regions of different rows are represented by grayscales of different depths. The text regions included in the first row are "SE" and "POS tagging information", the text regions included in the second row are "NT", "adj", "verb", "idiom", "noun", "other", the text regions included in the third row are "pos", "1230", "734", "1026", "266", "642", the text regions included in the fourth row are "neg", "785", "904", "746", "165", "797", the text regions included in the fifth row are "neu", "918", "7569", "2016", "12668", "10214", and the text regions included in the sixth row are "sum", "2933", "9207", "3788", "13099", "11653".

[0074] Furthermore, referring to Figure 6 it can be seen that the column relationship prediction result can determine the text regions corresponding to each column in the table. In Figure 6Different shades of gray are used to represent text regions in different columns. Among them, the text regions included in the first column are "SENT", "pos", "neg", "neu", "sum", the text regions included in the second column are "POS", "adj", "1230", "785", "918", "2933", the text regions included in the third column are "tagg", "verb", "734", "904", "7569", "9207", the text regions included in the fourth column are "ing in", "idiom", "1026", "746", "2016", "3788", the text regions included in the fifth column are "form", "noun", "266", "165", "12668", "13099", and the text regions included in the sixth column are "ation", "other", "642", "797", "10214", "11653".

[0075] Step S210: Determine the table structure corresponding to the target table image region based on the prediction results of the row-column relationships of each text region.

[0076] Specifically, according to the prediction results of the row-column relationships of every two text regions, it can be further determined which specific text regions belong to the same row and which text regions belong to the same column. By further analyzing and arranging the row-column relationships and position coordinates of different text regions, the table structure corresponding to the target table image region can be determined.

[0077] In one embodiment, in order to objectively evaluate the accuracy of the table recognition result, a data set for measuring the accuracy of table structure recognition is constructed, as shown in Table 1. Among them, the recall rate of table structure recognition refers to the proportion of correctly predicted adjacency relationships among the tables existing in the test set, and the accuracy rate of table structure recognition refers to the proportion of correctly predicted adjacency relationships among the predicted results.

[0078] Table 1 Table Structure Recognition Metrics

[0079] Recall Precision Table Structure Recognition 97.91% 98.14%

[0080] In the above table structure recognition method, by obtaining the target table image area, recognizing the text areas in the target table image area, and determining the image features and coordinate features of each text area, the image features and coordinate features are respectively fused to obtain the text area element fusion features corresponding to each text area. The overall recognition of the target table image area can be achieved by fusing the image features and coordinate features of different text areas within the target table image area, rather than the local recognition of a single text area. Furthermore, according to the text area element fusion features, the adjacency features of each node in the target table image area are determined, and the adjacency features of any two nodes are feature - spliced. By classifying and predicting the adjacency matrix obtained by splicing, the row - column relationship prediction results of the text areas corresponding to the two nodes are generated. Then, based on the row - column relationship prediction results of each text area, the table structure corresponding to the target table image area is determined. It realizes that the corresponding table structure can be determined according to the prediction results of the row - column relationships of each text area, without the need to further identify using an additional text detection network, which can reduce unnecessary cumbersome operations and thus improve the table recognition accuracy and recognition efficiency.

[0081] In one embodiment, as Figure 7 shown, the step of obtaining the text area element fusion features corresponding to each text area, that is, the step of determining the image features and coordinate features of each text area and respectively fusing the image features and coordinate features to obtain the text area element fusion features corresponding to each text area, specifically includes:

[0082] Step S702, obtain the position coordinates of each text area determined from the target table image area, and perform dimension - elevation on the position coordinates of the text area to obtain the dimension - elevated coordinate features.

[0083] Specifically, by determining each text area from the target table image area and obtaining the position coordinates of each text area, and then using a FCN network (fully - connected network) to perform dimension - elevation on the position coordinates of each text area respectively to obtain the dimension - elevated coordinate features.

[0084] Among them, the position coordinates of each text area are 4 - dimensional, which can be the four - dimensional coordinates of (x, y, w, h). For subsequent fusion with image features, the FCN network is used to elevate the four - dimensional coordinate features to the same dimension as the image features.

[0085] In one embodiment, before obtaining the position coordinates of each text area determined from the target table image area and performing dimension - elevation on the position coordinates of the text area to obtain the dimension - elevated coordinate features, it further includes:

[0086] Calculate the intersection over union (IoU) of each text region within the target table image region and the preset labeled text region; filter out the text regions with an IoU greater than the preset IoU threshold.

[0087] Specifically, by obtaining the preset labeled text region, calculating the IoU of each text region within the target table image region and the preset labeled text region, and obtaining the preset IoU threshold, filter out the text regions with an IoU greater than the preset IoU threshold.

[0088] Among them, the preset labeled text region is a pre-labeled text region, and also carries the row-column relationship of the corresponding labeled text region, that is, which specific text regions the already labeled text region belongs to the same row or the same column. In this embodiment, the preset IoU threshold can be different values between 0.7 and 0.9. Preferably, the preset IoU threshold can be taken as 0.8.

[0089] Step S704, according to the position coordinates of each text region, obtain the image content of the corresponding text region.

[0090] Specifically, according to the position coordinates of the text region, determine the specific position of the text region within the target table image region, and then obtain the image content at the corresponding specific position, which is determined as the image content corresponding to the text region.

[0091] Step S706, perform image feature alignment based on the image content of the text region to obtain the aligned image features, and the dimension of the aligned image features is the same as the dimension of the upsampled coordinate features.

[0092] Specifically, use the Roi Align algorithm (that is, the algorithm that uses bilinear interpolation to fix the feature output of regions of interest with different sizes) to perform image feature alignment on the image content of the text region. Among them, the image features of the text region can be further obtained by identifying the image content within the text region according to the FPN network (Feature Pyramid Network) when using the Mask-RCNN network to perform object detection on the target table region to determine the text regions within the target table region.

[0093] Among them, when using the Roi Align algorithm to perform image feature alignment on the image content of the text region, the obtained aligned image features are 128-dimensional. Since the position coordinates of each text region are 4-dimensional, which can be the four-dimensional coordinates of (x, y, w, h), in order to be used for subsequent fusion with the image features, the four-dimensional coordinate features are then upsampled to the same dimension as the image features by using the FCN network (fully connected network), that is, the four-dimensional coordinate features are upsampled to 128-dimensional by the FCN network (fully connected network) to be consistent with the dimension of the aligned image features.

[0094] Step S708: Fuse the upscaled coordinate features and the aligned image features to obtain text region element fusion features corresponding to each text region.

[0095] Specifically, by using the method of element-wise addition, fuse the upscaled coordinate features and the aligned image features corresponding to each text region to obtain text region element fusion features corresponding to each text region.

[0096] In this embodiment, the position coordinates of each text region are obtained from the target table image region, and the position coordinates of the text region are upscaled to obtain upscaled coordinate features. According to the position coordinates of each text region, the image content of the corresponding text region is obtained, and image feature alignment is performed based on the image content of the text region to obtain aligned image features. The dimension of the aligned image features is the same as that of the upscaled coordinate features. By fusing the upscaled coordinate features and the aligned image features, text region element fusion features corresponding to each text region are obtained, enabling the overall recognition of all text regions in the target table image region instead of the local recognition of a single text region, thereby improving the accuracy of table recognition for the target table image region.

[0097] In one embodiment, as Figure 8 shown, the steps of obtaining the local features of the nodes corresponding to each text region element fusion feature specifically include:

[0098] Step S802: Obtain the nodes corresponding to each text region element fusion feature and determine a preset number of neighboring nodes for each node.

[0099] Specifically, by obtaining the nodes corresponding to each text region element fusion feature and using the K-Nearest Neighbor algorithm, K neighboring nodes corresponding to each text region element fusion feature are respectively determined. The K-Nearest Neighbor algorithm is used to determine the K nearest neighboring nodes to the current node.

[0100] Step S804: Fuse the text region element fusion features of each node and the neighboring nodes corresponding to this node to obtain neighboring fusion features corresponding to each node.

[0101] Specifically, by performing feature fusion on the text region element features of the K neighboring nodes and the text region element fusion feature of this node, neighboring fusion features corresponding to each node are obtained.

[0102] Among them, since the text region element feature of each node is 128-dimensional, after fusing the text region element features of K neighboring nodes with the fused feature of the text region elements of this node, the neighboring fused feature of the node is enhanced to 128K dimensions.

[0103] Step S806: Integrate the neighboring fused features of each node to obtain an aggregated feature.

[0104] Specifically, by using a FCN network (fully connected network) to integrate the neighboring fused features of the nodes, the corresponding aggregated feature is obtained. Among them, since the neighboring fused feature of the node is 128K-dimensional, by using the FCN network, the 128K-dimensional neighboring fused features of each node are aggregated into a 128-dimensional aggregated feature.

[0105] Step S808: Perform dimensionality reduction processing on the aggregated features of each node to obtain the dimensionality-reduced multi-head graph features. The dimensionality-reduced multi-head graph features are the local features of the nodes corresponding to the fused features of the text region elements.

[0106] Specifically, by using a preset number of parallel FCN networks to perform dimensionality reduction processing on the aggregated features of the nodes, the dimensionality-reduced multi-head graph features are obtained. Among them, the dimensionality-reduced multi-head graph features are the local features of the nodes corresponding to the fused features of the text region elements.

[0107] Furthermore, in this embodiment, it may be to use 8 parallel FCN networks to perform dimensionality reduction processing on the aggregated features of the nodes, and convert the 128-dimensional aggregated features into 8 16-dimensional multi-head graph features. When using parallel FCN networks to perform dimensionality reduction processing, the dimensionality reduction processing processes of each parallel FCN network are independent and do not affect each other.

[0108] In this embodiment, by obtaining the nodes corresponding to the fused features of each text region element, determining the preset number of neighboring nodes of each node, and then fusing the fused features of the text region elements of each neighboring node corresponding to this node, the neighboring fused features corresponding to each node are obtained. By integrating the neighboring fused features of each node, an aggregated feature is obtained, and dimensionality reduction processing is performed on the aggregated features of each node to obtain the dimensionality-reduced multi-head graph features. The obtained dimensionality-reduced multi-head graph features are the local features of the nodes corresponding to the fused features of the text region elements. It realizes the further integration and dimensionality reduction processing of the fused features of the text region elements, obtains the dimensionality-reduced multi-head graph features, which is convenient for further fusing with the global features obtained by aggregating the context features of the fused features of the text region elements through the multi-head attention mechanism later, so as to achieve the overall recognition of the target table image region, rather than the local recognition of a single text region, and improve the accuracy of table recognition.

[0109] In one embodiment, asFigure 9 As shown in the figure, an overall process of a table structure recognition method is provided, which specifically includes a P1 text detection part, a P2 feature aggregation part, and a P3 adjacency relationship prediction part, where:

[0110] 1. The P1 text region recognition part specifically includes:

[0111] 1) Use the Mask-RCNN network to perform object detection on the target table region to obtain the recognition results of the corresponding RPN network (region proposal network), and determine the text regions within the target table region according to the recognition results of the RPN network. Among them, the backbone network of the Mask-RCNN network is the Res50 network with FPN.

[0112] 2) Use the NMS algorithm (non-maximum suppression algorithm) to filter the recognized text regions to obtain the filtered text regions. Among them, the detection results of the text regions are as Figure 4 shown. Within the target table region, there are text regions corresponding to each text box as shown in Figure 4 the figure.

[0113] 2. The P2 feature aggregation part specifically includes:

[0114] 1) Determine the image features and coordinate features of each text region, and fuse the image features and coordinate features respectively to obtain the text region element fusion features corresponding to each text region.

[0115] In one embodiment, determining the image features and coordinate features of each text region, and fusing the image features and coordinate features respectively to obtain the text region element fusion features corresponding to each text region includes:

[0116] Obtain the position coordinates of each text region determined from the target table image region, and perform dimension elevation on the position coordinates of the text region to obtain the dimension-elevated coordinate features; according to the position coordinates of each text region, obtain the image content of the corresponding text region; perform image feature alignment based on the image content of the text region to obtain the aligned image features, and the dimension of the aligned image features is the same as the dimension of the dimension-elevated coordinate features; fuse the dimension-elevated coordinate features and the aligned image features to obtain the text region element fusion features corresponding to each text region.

[0117] Specifically, by determining each text region from the target table image region and obtaining the position coordinates of each text region, and then using a FCN network (fully connected network) to respectively perform dimensionality elevation on the position coordinates of each text region to obtain the dimensionality-elevated coordinate features. According to the position coordinates of the text region, determine the specific position of the text region within the target table image region, and then obtain the image content at the corresponding specific position, determine it as the image content corresponding to the text region, and use the RoiAlign algorithm (i.e., an algorithm that uses bilinear interpolation to fix the feature output of regions of interest with different sizes) to perform image feature alignment on the image content of the text region. By using the method of adding bit by bit, fuse the dimensionality-elevated coordinate features and the aligned image features corresponding to each text region to obtain the text region element fusion features corresponding to each text region.

[0118] 2) Determine the adjacency features of each node in the target table image region according to the text region element fusion features.

[0119] Specifically, determining the adjacency features of each node in the target table image region according to the text region element fusion features includes:

[0120] Obtain the local features and global features of the nodes corresponding to each text region element fusion feature; perform feature aggregation based on the local features and global features to obtain the adjacency features of each node in the target table image region.

[0121] In one embodiment, obtaining the local features of the nodes corresponding to each text region element fusion feature includes:

[0122] Obtain the nodes corresponding to each text region element fusion feature, and determine the preset number of neighboring nodes of each node; fuse the text region element fusion features of each node and the neighboring nodes corresponding to this node to obtain the neighboring fusion features corresponding to each node; integrate the neighboring fusion features of each node to obtain an aggregated feature; perform dimensionality reduction processing on the aggregated feature of each node to obtain the dimensionality-reduced multi-head graph feature, and the dimensionality-reduced multi-head graph feature is the local feature of the node corresponding to the text region element fusion feature.

[0123] Specifically, use the K-nearest neighbor algorithm to respectively determine the K neighboring nodes of the nodes corresponding to each text region element fusion feature, fuse the text region element features of the K neighboring nodes and the text region element fusion feature of this node to obtain the neighboring fusion features of each node, and use a FCN network (fully connected network) to integrate the neighboring fusion features of the nodes to obtain the corresponding aggregated feature. Further use a preset number of parallel FCN networks to perform dimensionality reduction processing on the aggregated feature of the node to obtain the dimensionality-reduced multi-head graph feature.

[0124] In one embodiment, obtaining the global features of the nodes corresponding to the fused features of each text region element includes:

[0125] According to the multi-head attention mechanism, perform context feature aggregation on the fused features of the text region elements corresponding to each node to obtain the global features of the nodes corresponding to the fused features of the text region elements.

[0126] Among them, the encoder of the transformer model (a language model based on the self-attention mechanism) is used to represent the fused features of the text region elements. The hidden layer size of the transformer model is 128, and the dimension of the fused features of the text region elements is 128-dimensional. Among them, the number of heads corresponding to the multi-head attention mechanism (Multi-head attention) can be 8. When performing context feature aggregation on the fused features of the text region elements corresponding to each node according to the multi-head attention mechanism, the dimension of the global features of the nodes corresponding to the fused features of the text region elements is 16-dimensional. And 8 parallel FCN networks are used to reduce the dimension of the 128-dimensional aggregated features of the nodes, and the dimension of the reduced multi-head graph features is also 16-dimensional.

[0127] In one embodiment, perform feature aggregation based on local features and global features to obtain the adjacency features of each node in the target table image region, including:

[0128] Obtain each gate parameter corresponding to the AND mechanism; based on a preset activation function and each gate parameter, perform feature aggregation on the local features and global features to obtain the adjacency features of each node in the target table image region.

[0129] In one embodiment, as Figure 10 shown, a schematic diagram of the FLAG network structure for generating adjacency features is provided. Among them, the FLAG network structure represents the structure of the self-attention mechanism incorporating graph features. Referring to Figure 10 it can be seen that the FLAG network structure is provided with a GNN branch (Graph Neural Networks, that is, graph neural network) for determining local features, including GNN1,..., GNN N , a self-attention (self-attention mechanism) branch for determining global features, including self-attention head1,..., self-attention head n , a gate mechanism (gate mechanism) for controlling the fusion of different types of features. The gate parameters corresponding to the gate mechanism include gate1,..., gate n, and an FFN (Feed-Forward Network) for enhancing the representation ability of the model. Among them, the number of GNN branches, the number of self-attention branches, and the number of gate mechanisms are the same.

[0130] Specifically, in the self-attention branch, the encoder of the transformer model (a language model based on the self-attention mechanism) is used to represent the fused features of the text region elements. Refer to Figure 10 It can be known that the attention function used by the transformer model includes: Q (query), K (key), and V (value). Among them, the query vector (Q) is responsible for finding the relevance between the target feature and other features, the key vector (K) is used to match the query vector (Q) to obtain a relevance score, and the value vector (V) is used to perform a weighted sum of the scores.

[0131] Furthermore, since the self-attention branch is a multi-head attention mechanism, there are correspondingly multiple heads, including head1, head2,..., head n , and the attention function used by the transformer model is separately set for each head, including K1, Q1, V1, K2, Q2, V2,..., K N , Q N , V N .

[0132] Specifically, by obtaining a preset activation function and the gate parameters corresponding to the gate mechanism, including gate1,..., gate n , and based on the preset activation function and each gate parameter, feature aggregation is performed on the local features and global features of each node to obtain the adjacent features of each node in the target table image area. Among them, an FFN (Feed-Forward Network) is used to further identify and analyze the adjacent features obtained by feature aggregation to enhance the representation ability of the model.

[0133] In this embodiment, a four-layer FLAG network structure is set, and the adjacent features output by each layer of the FLAG network structure are further subjected to feature fusion, and finally the adjacent features of each node in the target table image area are determined.

[0134] 3. The P3 adjacency relationship prediction part specifically includes:

[0135] (1) Concatenate the adjacency features of any two nodes, and perform classification prediction on the concatenated adjacency matrix to generate a prediction result of the row-column relationship of the text region corresponding to the two nodes.

[0136] In one embodiment, concatenating the adjacency features of any two nodes and performing classification prediction on the concatenated adjacency matrix to generate a prediction result of the row-column relationship of the text region corresponding to the two nodes includes:

[0137] Concatenate the adjacency features of any two nodes to obtain a concatenated adjacency matrix; perform binary classification prediction on the concatenated adjacency matrix according to a fully connected neural network to obtain a prediction result of the row-column relationship of the corresponding text region; wherein, the binary classification prediction includes row relationship prediction and column relationship prediction.

[0138] Specifically, perform binary classification prediction on the concatenated adjacency matrix according to the fully connected network to determine whether the two text regions corresponding to the concatenated adjacency matrix belong to the same row in the table, or determine whether the two text regions corresponding to the concatenated adjacency matrix belong to the same column in the table, so as to obtain a prediction result of the row-column relationship of the corresponding text region.

[0139] Further, the prediction results of the row-column relationships of each text region can be referred to Figure 5 the row relationship prediction result of the table structure recognition method shown, and Figure 6 the column relationship prediction result of the table structure recognition method shown.

[0140] (2) Based on the prediction results of the row-column relationships of each text region, determine the table structure corresponding to the target table image region.

[0141] Specifically, according to the prediction results of the row-column relationships of every two text regions, it can be further determined which text regions specifically belong to the same row and which text regions belong to the same column. By further analyzing and arranging the row-column relationships and position coordinates of different text regions, the table structure corresponding to the target table image region can be determined.

[0142] In the above table structure recognition method, by obtaining the target table image area, recognizing the text areas in the target table image area, and determining the image features and coordinate features of each text area, the image features and coordinate features are respectively fused to obtain the text area element fusion features corresponding to each text area. The overall recognition of the target table image area can be achieved by fusing the image features and coordinate features of different text areas within the target table image area, rather than the local recognition of a single text area. Furthermore, according to the text area element fusion features, the adjacency features of each node in the target table image area are determined, and the adjacency features of any two nodes are feature-stitched. By classifying and predicting the adjacency matrix obtained by stitching, the row-column relationship prediction results of the text areas corresponding to the two nodes are generated. Then, based on the row-column relationship prediction results of each text area, the table structure corresponding to the target table image area is determined. It realizes that the corresponding table structure can be determined according to the prediction results of the row-column relationships of each text area, without the need to further identify using an additional text detection network, which can reduce unnecessary cumbersome operations and thus improve the table recognition accuracy and recognition efficiency.

[0143] It should be understood that although the steps in the respective flowcharts involved in the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the respective flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0144] In one embodiment, as Figure 11 shown, a table structure recognition device is provided. This device can be a software module, a hardware module, or a combination of both to become a part of a computer device. Specifically, the device includes: a text area recognition module 1102, a text area element fusion feature generation module 1104, an adjacency feature generation module 1106, a row-column relationship prediction result generation module 1108, and a table structure determination module 1110, where:

[0145] The text area recognition module 1102 is used to obtain the target table image area and recognize the text areas in the target table image area.

[0146] The text region element fusion feature generation module 1104 is used to determine the image features and coordinate features of each text region, and fuse the image features and coordinate features respectively to obtain the text region element fusion features corresponding to each text region.

[0147] The adjacency feature generation module 1106 is used to determine the adjacency features of each node in the target table image region according to the text region element fusion features.

[0148] The row-column relationship prediction result generation module 1108 is used to splice the adjacency features of any two nodes, and perform classification prediction on the spliced adjacency matrix to generate the row-column relationship prediction result of the text region corresponding to the two nodes.

[0149] The table structure determination module 1110 is used to determine the table structure corresponding to the target table image region based on the row-column relationship prediction results of each text region.

[0150] In the above table structure recognition device, by obtaining the target table image region, identifying the text regions in the target table image region, and determining the image features and coordinate features of each text region, and fusing the image features and coordinate features respectively to obtain the text region element fusion features corresponding to each text region, the overall recognition of the target table image region can be achieved by fusing the image features and coordinate features of different text regions in the target table image region, rather than the local recognition of a single text region. Furthermore, according to the text region element fusion features, the adjacency features of each node in the target table image region are determined, and the adjacency features of any two nodes are spliced, and by performing classification prediction on the spliced adjacency matrix, the row-column relationship prediction result of the text region corresponding to the two nodes is generated. Furthermore, based on the row-column relationship prediction results of each text region, the table structure corresponding to the target table image region is determined. It is realized that according to the prediction results of the row-column relationships of each text region, the corresponding table structure can be determined without using an additional text detection network for further recognition, which can reduce unnecessary cumbersome operations and thus improve the table recognition accuracy and recognition efficiency.

[0151] In one embodiment, the text region element fusion feature generation module is further used for:

[0152] Obtain the position coordinates of each text region determined within the target table image region, and perform dimensionality elevation on the position coordinates of the text regions to obtain the dimensionality-elevated coordinate features; according to the position coordinates of each text region, obtain the image content of the corresponding text region; perform image feature alignment based on the image content of the text region to obtain the aligned image features, and the dimension of the aligned image features is the same as the dimension of the dimensionality-elevated coordinate features; fuse the dimensionality-elevated coordinate features and the aligned image features to obtain the text region element fusion features corresponding to each text region.

[0153] The above-mentioned text region element fusion feature generation module realizes the fusion of the dimensionality-elevated coordinate features and the aligned image features to obtain the text region element fusion features corresponding to each text region, and can achieve the overall recognition of all text regions within the target table image region, rather than the local recognition of a single text region, thereby improving the table recognition accuracy of the target table image region.

[0154] In one embodiment, the text region element fusion feature generation module further includes:

[0155] An intersection over union calculation unit for calculating the intersection over union of each text region within the target table image region and a preset labeled text region;

[0156] A text region screening module for screening out the text regions whose intersection over union is greater than a preset intersection over union threshold.

[0157] In one embodiment, the adjacent feature generation module is further used for:

[0158] Obtain the local features and global features of the nodes corresponding to each text region element fusion feature; perform feature aggregation based on the local features and global features to obtain the adjacent features of each node in the target table image region.

[0159] In one embodiment, the adjacent feature generation module further includes:

[0160] A neighboring node acquisition module for acquiring the nodes corresponding to each text region element fusion feature and determining a preset number of neighboring nodes for each node;

[0161] A neighboring fusion feature generation module for fusing each node with the text region element fusion features of each corresponding neighboring node of the node to obtain the neighboring fusion features corresponding to each node;

[0162] An aggregation feature generation module for integrating the neighboring fusion features of each node to obtain an aggregation feature;

[0163] A local feature generation module is used to perform dimensionality reduction processing on the aggregated features of each node to obtain the multi-head graph features after dimensionality reduction. The multi-head graph features after dimensionality reduction are the local features of the nodes corresponding to the fusion features of the text region elements.

[0164] The above adjacency feature generation module realizes the further integration and dimensionality reduction processing of the fusion features of the text region elements, obtains the multi-head graph features after dimensionality reduction, and is convenient for further fusing with the global features obtained by aggregating the context features of the fusion features of the text region elements through the multi-head attention mechanism, so as to achieve the overall recognition of the target table image region, rather than the local recognition of a single text region, and improve the accuracy of table recognition.

[0165] In one embodiment, the adjacency feature generation module further includes a global feature generation module, which is used for:

[0166] According to the multi-head attention mechanism, perform context feature aggregation on the fusion features of the text region elements corresponding to each node to obtain the global features of the nodes corresponding to the fusion features of the text region elements.

[0167] In one embodiment, the row-column relationship prediction result generation module is further used for:

[0168] Perform feature splicing on the adjacency features of any two nodes to obtain the spliced adjacency matrix; perform binary classification prediction on the spliced adjacency matrix according to the fully connected neural network to obtain the row-column relationship prediction result of the corresponding text region; wherein, the binary classification prediction includes row relationship prediction and column relationship prediction.

[0169] For the specific limitations of the table structure recognition device, reference can be made to the limitations of the table structure recognition method in the above text, which will not be elaborated here. Each module in the above table structure recognition device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0170] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 12As shown in the figure. The computer device includes a processor, a memory, and a network interface connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as text areas, image features, coordinate features, text area element fusion features, adjacency features, and row-column relationship prediction results. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a table structure recognition method.

[0171] Those skilled in the art can understand that Figure 12 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0172] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0173] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0174] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0175] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0176] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0177] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for recognizing a table structure, characterized in that, The method includes: Obtain a target table image region, and identify text regions in the target table image region; Determine the image features and coordinate features of each of the text regions, and respectively fuse the image features and the coordinate features to obtain text region element fusion features corresponding to each of the text regions; Obtain nodes corresponding to the text region element fusion features of each of the text regions, determine a preset number of neighboring nodes of each of the nodes, and fuse the text region element fusion features of each of the nodes and the neighboring nodes corresponding to the node to obtain neighboring fusion features corresponding to each of the nodes; Integrate the neighboring fusion features of each of the nodes to obtain an aggregated feature, and perform dimensionality reduction processing on the aggregated feature of each of the nodes to obtain a multi-head graph feature after dimensionality reduction; the multi-head graph feature after dimensionality reduction is the local feature of the node corresponding to the text region element fusion feature; Obtain the global feature of the node corresponding to the text region element fusion feature of each of the text regions, and perform feature aggregation based on the local feature and the global feature to obtain the adjacency feature of each of the nodes in the target table image region; Perform feature splicing on the adjacency features of any two nodes, and perform classification prediction on the spliced adjacency matrix to generate a row-column relationship prediction result of the text regions corresponding to the two nodes; Based on the row-column relationship prediction results of each of the text regions, determine a table structure corresponding to the target table image region.

2. The method according to claim 1, wherein The determining the image features and coordinate features of each of the text regions, and respectively fusing the image features and the coordinate features to obtain text region element fusion features corresponding to each of the text regions includes: Obtain the position coordinates of each of the text regions determined from the target table image region, and perform dimensionality increase on the position coordinates of the text regions to obtain coordinate features after dimensionality increase; According to the position coordinates of each of the text regions, obtain the image content of the corresponding text region; Perform image feature alignment based on the image content of the text region to obtain aligned image features, and the dimension of the aligned image features is the same as the dimension of the coordinate features after dimensionality increase; Fuse the coordinate features after dimensionality increase and the aligned image features to obtain text region element fusion features corresponding to each of the text regions.

3. The method according to claim 2, characterized in that, Before the obtaining the position coordinates of each of the text regions determined from the target table image region, and performing dimensionality increase on the position coordinates of the text regions to obtain coordinate features after dimensionality increase, it further includes: Calculate the intersection-over-union ratio of each of the text regions in the target table image region and a preset labeled text region; Filter out text regions with an intersection-over-union ratio greater than a preset intersection-over-union ratio threshold.

4. The method according to claim 1, wherein Obtaining the global feature of the node corresponding to the text region element fusion feature includes: According to the multi-head attention mechanism, perform context feature aggregation on the text region element fusion features corresponding to each of the nodes to obtain the global feature of the node corresponding to the text region element fusion feature.

5. The method according to claim 1, characterized in that Performing feature concatenation on the adjacency features of any two nodes, and classifying and predicting the concatenated adjacency matrix to generate a prediction result of the row-column relationship of the text region corresponding to the two nodes, including: Performing feature concatenation on the adjacency features of any two nodes to obtain a concatenated adjacency matrix; Performing binary classification prediction on the concatenated adjacency matrix according to a fully connected neural network to obtain a prediction result of the row-column relationship of the corresponding text region; wherein, the binary classification prediction includes row relationship prediction and column relationship prediction.

6. A table structure recognition device, characterized in that, The apparatus includes: A text region recognition module, configured to obtain a target table image region and recognize text regions in the target table image region; A text region element fusion feature generation module, configured to determine the image features and coordinate features of each text region, and respectively fuse the image features and coordinate features to obtain text region element fusion features corresponding to each text region; An adjacency feature generation module, configured to obtain nodes corresponding to the text region element fusion features, determine a preset number of neighboring nodes of each node, and fuse the text region element fusion features of each node and the neighboring nodes corresponding to the node to obtain neighboring fusion features corresponding to each node; integrating the neighboring fusion features of each node to obtain an aggregated feature, and performing dimensionality reduction processing on the aggregated feature of each node to obtain a dimensionality-reduced multi-head graph feature; the dimensionality-reduced multi-head graph feature is the local feature of the node corresponding to the text region element fusion feature; obtaining the global feature of the node corresponding to the text region element fusion feature, and performing feature aggregation based on the local feature and the global feature to obtain the adjacency feature of each node in the target table image region; A row-column relationship prediction result generation module, configured to perform feature concatenation on the adjacency features of any two nodes, and classify and predict the concatenated adjacency matrix to generate a prediction result of the row-column relationship of the text region corresponding to the two nodes; A table structure determination module, configured to determine a table structure corresponding to the target table image region based on the prediction results of the row-column relationships of the text regions.

7. The device according to claim 6, characterized in that, The text region element fusion feature generation module is further configured to: Obtain the position coordinates of each text region determined from the target table image region, and perform dimensionality increase on the position coordinates of the text region to obtain a dimensionality-increased coordinate feature; obtain the image content of the corresponding text region according to the position coordinates of each text region; perform image feature alignment based on the image content of the text region to obtain an aligned image feature, and the dimension of the aligned image feature is the same as the dimension of the dimensionality-increased coordinate feature; fuse the dimensionality-increased coordinate feature and the aligned image feature to obtain text region element fusion features corresponding to each text region.

8. The device according to claim 7, characterized in that, The text region element fusion feature generation module further includes: An intersection over union calculation unit, configured to calculate the intersection over union of each text region in the target table image region and a preset labeled text region; A text region screening module for screening out text regions with an intersection over union greater than a preset intersection over union threshold.

9. The device according to claim 6, characterized in that, The adjacent feature generation module further includes a global feature generation module for: According to the multi-head attention mechanism, perform context feature aggregation on the text region element fusion features corresponding to each of the nodes to obtain the global features of the nodes corresponding to the text region element fusion features.

10. The device according to claim 6, characterized in that, The row-column relationship prediction result generation module is further used for: Perform feature splicing on the adjacent features of any two nodes to obtain a spliced adjacent matrix; perform binary classification prediction on the spliced adjacent matrix according to a fully connected neural network to obtain the row-column relationship prediction result of the corresponding text region; wherein, the binary classification prediction includes row relationship prediction and column relationship prediction.

11. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image table extraction method and device, electronic equipment and storage medium

    CN111695517A