Complex table structure identification method, system and device for power industry and medium

By using multimodal fusion of visual, textual, and layout features, and contextual aggregation of graph neural networks, the problem of single feature representation and insufficient contextual relationship modeling in the recognition of complex tables in the power industry is solved, achieving efficient and accurate table structure recognition and improving the accuracy and adaptability of recognition.

CN120913231APending Publication Date: 2025-11-07GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511008542.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies for recognizing complex tables in the power industry suffer from problems such as limited feature representation, coarse modal fusion, insufficient modeling of contextual relationships, and weak generalization ability due to reliance on rules or templates, making it difficult to accurately identify tables with complex structures.

Method used

A multimodal fusion method combining visual, textual, and layout features is adopted. Features are extracted through a deep learning model and combined with a graph neural network for context aggregation and structural reasoning. Feature weights are dynamically adjusted to establish contextual relationships between table elements and improve recognition accuracy by utilizing local and global information.

Benefits of technology

It achieves efficient and accurate identification of complex tables in the power industry, improves the degree of automation and identification efficiency, enhances the adaptability and generalization ability of the method, and is suitable for diverse data processing in the power industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913231A_ABST
    Figure CN120913231A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of the power industry, and discloses a complex table structure identification method, system and device for the power industry and a medium, and the method comprises the steps: capturing key information, such as visual, text and layout features, of a target to-be-identified table image through a first feature extraction operation; and a data basis is provided for subsequent feature fusion and reasoning. The second feature fusion operation integrates different features, and the first stage strengthens the relation between vision and text features, so that the recognition is more accurate; in the second stage, layout features are fused, and the overall understanding of the table structure is enhanced. In the context aggregation step, a logic relation of table image elements is established, and recognition continuity and accuracy are improved. The line-column and cell relation between elements is analyzed through table image structure reasoning, and the recognition accuracy is ensured. According to the method, efficient and accurate identification of the complex table structure in the power industry is realized, and the automation degree and efficiency of table data processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of the power industry, and in particular to a complex table structure recognition method, system, device and medium for the power industry. BACKGROUND

[0002] In the power industry, there are a large number of complex table information processing scenarios in production, dispatching, operation and maintenance, management and other aspects. These scenarios not only include dispatching and operation and maintenance records, asset device account books, fault statistics reports, etc., but also may cover more types of table data. These tables usually have the characteristics of complex structure, variable format and information-intensive, which brings great challenges to automatic recognition and system integration.

[0003] In the prior art, the structure recognition method for table images mainly has the following deficiencies:

[0004] 1. Single feature expression, insufficient modal utilization: Current mainstream recognition methods usually rely only on visual information or text information for structure reasoning, lacking the ability to jointly model multi-modal features such as images, text and layout. This single feature expression method makes it difficult to accurately recognize tables with complex structure characteristics such as merged cells, cross-row and cross-column in complex power grid table scenarios, affecting the accuracy and reliability of recognition.

[0005] 2. Rough modal fusion method, unreasonable feature weight allocation: Some existing recognition methods directly concatenate or weight the features of different modalities, which cannot adaptively adjust the importance of each modality according to the specific table content. In the scenario of high diversity of power grid data, this rough modal fusion method is prone to misjudgment, reducing the recognition effect.

[0006] 3. Insufficient modeling of context relationships, limited structure understanding: Traditional table recognition methods fail to effectively model the complex relationships between cells in the table, such as the adjacency relationship between rows and columns, nested structure, etc. Especially in handling long-distance dependencies and recognizing table edge areas, it performs poorly and cannot meet the high requirements of table automation parsing and power grid data system integration.

[0007] 4. Relies on rules or templates, weak generalization ability: Some recognition methods rely on fixed layout templates or rules for judgment, which are difficult to adapt to real power industry documents with diverse table styles and complex structures. This dependency not only leads to poor adaptability of the method, but also requires high maintenance cost, making it difficult to be widely promoted in practical applications.

[0008] In summary, the prior art has many deficiencies in processing complex table information in the power industry, and it is urgent to develop more advanced and efficient multi-modal fusion and context modeling methods to improve the accuracy and generalization ability of table recognition and meet the actual needs of data processing in the power industry. SUMMARY

[0009] In view of the above existing problems, the present application is proposed.

[0010] Therefore, the present application provides a complex table structure recognition method, system, device and medium for the power industry, which can solve the problems of single feature expression, rough modal fusion, insufficient context relationship modeling and weak generalization ability caused by relying on rules or templates in complex table recognition in the power industry.

[0011] To solve the above technical problems, the present application provides the following technical solutions:

[0012] In a first aspect, the present application provides a complex table structure recognition method for the power industry, comprising:

[0013] Obtaining a target table image to be recognized in the power industry, and performing a first feature extraction operation on the target table image to be recognized;

[0014] The first feature extraction includes a visual feature extraction operation, a text feature extraction operation and a layout feature extraction operation;

[0015] Performing a second feature fusion operation on the result of the first feature extraction, the second feature fusion operation including first-stage feature fusion and second-stage feature fusion;

[0016] The first-stage feature fusion is used to fuse visual features and text features;

[0017] The second-stage feature fusion is used to fuse the first-stage feature fusion result and the layout features;

[0018] Performing context aggregation and table image structure reasoning on the result of the second feature fusion operation;

[0019] The context aggregation is used to establish the context relationship between the elements of the target table image to be recognized, and the table image structure reasoning is used to obtain the row, column and cell relationship between the elements in the target table image to be recognized.

[0020] As a preferred scheme of the complex table structure recognition method for the power industry according to the present application, further comprising:

[0021] Establishing a table structure recognition model based on the first feature extraction operation, the second feature fusion operation, the context aggregation and the table image structure reasoning;

[0022] According to the table structure recognition model, the target table image structure to be recognized in the power industry is recognized.

[0023] The table structure recognition model further comprises a loss function improved according to the row, column and cell relationship.

[0024] As a preferred scheme of the complex table structure recognition method for the power industry, the second feature fusion operation comprises:

[0025] A preset gate control parameter is used to control the contribution of different features.

[0026] A first-stage feature fusion expression is established based on the gate control parameter, the visual feature and the text feature.

[0027] The first-stage feature fusion result is subjected to region feature extraction, and a model representing the region feature extraction operation and the exchange relationship between layout features is established.

[0028] This preferred scheme can more finely regulate the weight of different features in the fusion process, thereby improving the accuracy and efficiency of feature fusion. Through the preset gate control parameter, the contribution of the visual feature and the text feature in the first-stage feature fusion can be dynamically adjusted, so that the fusion result is more consistent with the actual features of the target table image to be recognized. Meanwhile, the first-stage feature fusion result is subjected to region feature extraction, and the corresponding model is established, which can further mine and utilize the correlation information between features, thereby enhancing the robustness and accuracy of table structure recognition.

[0029] As a preferred scheme of the complex table structure recognition method for the power industry, the context aggregation comprises local context aggregation and global context aggregation.

[0030] The table elements after the second feature fusion operation are taken as vertices, and the connections between the table elements are taken as edges.

[0031] The local context aggregation is used to obtain edge features and edge weights, and the local context aggregation comprises an improved element distance solving strategy.

[0032] The global context aggregation takes reading order and spatial position as aggregation consideration items.

[0033] The preferred scheme can comprehensively consider the local and global information of the table elements, so as to more comprehensively understand the table structure. The local context aggregation accurately calculates the correlation strength between the table elements by improving the element distance solving strategy, and captures the detail features. The global context aggregation integrates the reading order and spatial position information, and grasps the overall layout and logical structure of the table from a macro perspective. This combination of local and global context aggregation not only improves the accuracy of table structure recognition, but also enhances the ability to adapt to different complex table structures, so that the method has higher practical value and wider applicability in the application of the power industry.

[0034] As a preferred scheme of the complex table structure recognition method for the power industry provided by the application, the table image structure reasoning comprises: based on the node pairs constructed by the k-nearest neighbor method, the results of the context aggregation of each pair of nodes are spliced in the channel dimension to form a table image structure vector.

[0035] As a preferred scheme of the complex table structure recognition method for the power industry provided by the application, the text feature extraction operation comprises:

[0036] Obtaining the text information of the target table image to be recognized;

[0037] Mapping the text information into an embedding graph of the same size as the original target table image to be recognized;

[0038] The embedding graph comprises a character embedding graph and a sentence embedding graph, and the character embedding graph and the sentence embedding graph are obtained through a word embedding model and a pre-trained language model respectively.

[0039] Fusing the character embedding graph and the sentence embedding graph to obtain text features.

[0040] As a preferred scheme of the complex table structure recognition method for the power industry provided by the application, the layout feature extraction operation comprises:

[0041] Obtaining the text box of the target table image to be recognized, and the layout feature is calculated based on the bounding box coordinates of the text box.

[0042] The layout feature includes but is not limited to the width, height and center point coordinate information of the bounding box.

[0043] In a second aspect, the application provides a complex table structure recognition system for the power industry, comprising:

[0044] A double-flow feature extraction network establishment module is configured to obtain a target table image to be recognized in the power industry, and perform a first feature extraction operation on the target table image to be recognized.

[0045] The first feature extraction includes a visual feature extraction operation, a text feature extraction operation and a layout feature extraction operation;

[0046] The two-stage multi-modal feature fusion module is configured to perform a second feature fusion operation on the result of the first feature extraction, and the second feature fusion operation includes a first-stage feature fusion and a second-stage feature fusion.

[0047] The first-stage fusion feature is used for fusing the visual feature and the text feature.

[0048] The second-stage feature fusion is used for fusing the first-stage feature fusion result and the layout feature.

[0049] The mixed context aggregation module and the relationship reasoning module are respectively configured to perform context aggregation and table image structure reasoning on the result of the second feature fusion operation.

[0050] The context aggregation is used for establishing a context connection between target table image elements to be recognized, and the table image structure reasoning is used for obtaining a row, column and cell relationship between elements in the target table image to be recognized.

[0051] In a third aspect, the present application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method as described above when executing the computer program.

[0052] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the method as described above.

[0053] Compared with the prior art, the present application has the following beneficial effects: the present application proposes a complex table structure recognition method for the power industry, through the first feature extraction operation, the present application can comprehensively capture the key information in the target table image to be recognized, including visual features, text features and layout features, providing a rich data basis for subsequent feature fusion and reasoning. The second feature fusion operation effectively integrates different features, wherein the first-stage feature fusion strengthens the connection between visual and text features, making the recognition process more accurate; the second-stage feature fusion further integrates the layout feature, enhancing the overall understanding of the table structure. The context aggregation step establishes the logical connection between the table image elements, which helps to improve the coherence and accuracy of the recognition. The table image structure reasoning further analyzes the row, column and cell relationship between the elements, further ensuring the accuracy of the recognition. The present application realizes efficient and accurate recognition of the complex table structure of the power industry, greatly improving the automation degree and efficiency of table data processing. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0055] Figure 1 A method flowchart of a complex table structure recognition method for the power industry is provided for an embodiment of the present application.

[0056] Figure 2 An internal structure diagram of an electronic device of a complex table structure recognition method for the power industry is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the above-mentioned objects, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0058] Embodiment 1, refer to Figure 1 For the first embodiment of the present application, the embodiment provides a complex table structure recognition method for the power industry, comprising:

[0059] In the prior art, there are some problems, for example, the traditional table structure recognition method is often difficult to accurately identify each element and its relationship in the table when processing the complex table of the power industry. Because the table of the power industry usually contains a large number of professional terms and data, the structure of these tables is also relatively complex, so the traditional recognition method is prone to errors or omissions when processing these tables. In addition, the existing recognition method is low in efficiency when processing large-scale data, and is difficult to meet the needs of practical application.

[0060] The present application provides a method that can effectively solve the above-mentioned problems, and the following will be described in detail how to realize the complex table structure recognition method for the power industry in combination with multiple embodiments;

[0061] Figure 1 A method flowchart of a complex table structure recognition method for the power industry is shown, comprising:

[0062] S101, acquiring a target table image to be recognized of the power industry, and performing a first feature extraction operation on the target table image to be recognized, wherein:

[0063] It should be noted that in the power industry, the complex table information processing scenarios involved in production, scheduling, operation and maintenance, management and other links, these tables often have complex structure, variable format, information-intensive and other characteristics, which brings challenges to automatic recognition and system integration. Therefore, a complex table structure recognition method that can realize efficient and accurate recognition is needed.

[0064] It should also be noted that before designing a complex table structure recognition method, it needs to be clear that the final designed table structure recognition method is used in the specific power industry related table to be recognized, so the target table image to be recognized in the power industry needs to be processed and the final structure is obtained.

[0065] In some specific embodiments, the target table image to be recognized in the power industry can come from various aspects of the power industry, such as power plant operation records, power grid dispatching center dispatching logs, and power trading market transaction data. These table images usually contain a large amount of text, numbers and symbols, as well as complex row and column structures and nested relationships, so an efficient and accurate recognition method is needed for processing.

[0066] In the present application, the target table image to be recognized is not limited, as long as the related table image in the power industry can realize structure acquisition through the present application.

[0067] After obtaining the target table image to be recognized in the power industry, the table images need to be preprocessed so that useful information can be obtained step by step, and finally the table structure recognition is realized.

[0068] In some specific embodiments, the first feature extraction operation on the target table image to be recognized can use a deep learning model, such as a convolutional neural network (CNN) or a residual network (ResNet), to capture low-level visual features in the target table image to be recognized, such as edges, textures and shapes. At the same time, combined with natural language processing (NLP) techniques such as word embedding and pre-trained language models, the text information in the table is encoded to extract character-level and sentence-level text features. In addition, layout analysis algorithms such as rule-based text box detection or deep learning-based object detection models are used to determine the position and size of table elements, thereby extracting layout features. These features together constitute a comprehensive description of the target table image to be recognized, providing a solid foundation for subsequent feature fusion and reasoning.

[0069] In other specific embodiments, the first feature extraction operation can also use feature extraction networks in deep learning frameworks such as Faster R-CNN or Mask R-CNN, which not only have strong feature extraction capabilities but also can achieve accurate target detection, further improving the accuracy and efficiency of table element recognition. Specifically:

[0070] Step 1.1: In the first feature extraction stage, visual features mainly focus on the appearance information of the table image, such as lines, borders, and fill colors, etc.

[0071] Step 1.2: Text features focus on the text content in the table, including character sequences, words and phrases, etc.

[0072] Step 1.3: Layout features focus on describing the spatial distribution and arrangement of table elements, such as row and column alignment, cell position and size, etc. These features are automatically learned and optimized by deep learning models, which can more accurately reflect the internal structure and information of the table, providing a more reliable foundation for subsequent feature fusion and reasoning.

[0073] However, in the table structure acquisition process of the present application, each table element (i.e. text block) in the text modality has discrete characteristics and cannot be directly input as an image for processing, so a new first feature extraction method is designed.

[0074] In an embodiment of the present application, the first feature extraction includes visual feature extraction operation, text feature extraction operation and layout feature extraction operation.

[0075] In an embodiment of the present application, the text feature extraction operation includes:

[0076] Obtaining the text information of the target table image to be recognized;

[0077] Mapping the text information into an embedding map of the same size as the original target table image to be recognized;

[0078] The embedding map includes a character embedding map and a sentence embedding map, which are obtained through a word embedding model and a pre-trained language model, respectively.

[0079] Fusing the character embedding map and the sentence embedding map to obtain text features.

[0080] In some specific embodiments, the text information of the target table image to be recognized can be obtained through optical character recognition (OCR) technology, which can efficiently convert the text content in the table image into editable text information.

[0081] In some specific embodiments, the text information of the target table image to be recognized can also be obtained through deep learning models for character detection and recognition. These models have been trained on a large amount of data and can accurately recognize various characters in the table and convert them into computer- processable text format.

[0082] In the embodiments of the present application, how to obtain the text information of the target table image to be recognized is not limited, and related technical personnel can select according to actual needs.

[0083] In some specific embodiments, the word embedding model and the pre-trained language model can be pre-trained through a large-scale corpus to capture the statistical rules and semantic information of the language. The character embedding graph maps each character to a high-dimensional vector, and these vectors can reflect the similarity and relevance between characters. The sentence embedding graph further considers the context information in the text and maps sentences or phrases to vector representations, thereby capturing higher-level semantic features. By fusing the character embedding graph and the sentence embedding graph, more rich and accurate text feature representations can be obtained, providing strong support for subsequent feature fusion and reasoning.

[0084] In other specific embodiments, the word embedding model and the pre-trained language model can also be fine-tuned through self-defined tasks to adapt to the recognition needs of specific fields or specific table structures. For example, in the power industry, the table may contain a large number of professional terms and specific data formats. By training the word embedding model and the pre-trained language model specifically, they can better understand these professional terms and data formats, thereby improving the accuracy and efficiency of table recognition. This fine-tuning process can be based on a small amount of labeled data, thereby reducing the dependence on a large amount of labeled data and reducing implementation costs.

[0085] In some specific embodiments, the detailed steps of pre-training the word embedding model and the pre-trained language model through a large-scale corpus can be as follows:

[0086] Step 2.1: Collect a large-scale power industry-related corpus. These corpora can come from power industry documents, reports, standards, specifications, etc., to ensure the richness and diversity of the corpus.

[0087] Step 2.2: Pre-train the word embedding model and the pre-trained language model using these corpora. Through unsupervised learning, the model automatically learns the information of words, phrases, and sentence structures in the corpus, forming a preliminary vector representation.

[0088] Step 2.3: Evaluate and adjust the pre-trained model. Through comparative experiments and cross-validation, optimize the parameters and structure of the model to improve the accuracy and generalization ability of the model.

[0089] Step 2.4: Integrate the pre-trained word embedding model and the pre-trained language model into the table structure recognition system, and use them to vectorize the text information in the table to provide strong support for subsequent feature extraction and classification recognition.

[0090] Specifically, the specific steps of the text feature extraction operation of the present application can be as follows:

[0091] Step 3.1: Adopt a character granularity and sentence granularity joint modeling strategy to map the text information into an embedding graph with the same size as the original target table image to be recognized.

[0092] Step 3.2: In the embedding graph, the character embedding graph and the sentence embedding graph are obtained through the word embedding model and the pre-trained language model (such as BERT) respectively, and are fused through normalization operation;

[0093] Step 3.3: Form a continuous text representation feature for subsequent neural network processing.

[0094] T = LayerNorm(T char + T sentce ) (1)

[0095] In the above formula, represents the character embedding mapping obtained by the word embedding layer, represents the sentence embedding mapping obtained by the pre-trained language model BERT, is the final text embedding graph, which will play a similar role to the original image, except that it is manually obtained from the original text. H and W represent the height and width of the original target table image to be recognized. C t represents the number of channels set. LayerNorm is used to normalize the combination of the two graphs.

[0096] It should be noted that the visual feature extraction operation of the present application can use a deep learning model, such as a convolutional neural network (CNN) or its variants, to capture low-level visual features such as edges, textures, shapes, etc. in the target table image to be recognized.

[0097] In the embodiment of the present application, a double-flow neural network structure is designed to simultaneously acquire text features and visual features. Considering the inherent characteristics of table data, a multi-scale convolution attention block (MSCA) is used to redesign the backbone of the double-flow neural network. Each component in the block is essential for tables. Stripe convolution helps to maintain horizontal and vertical cell distribution, while the use of multi-scale feature extraction solves the problem of inconsistent cell sizes. In addition, the attention mechanism allows processing of even remote dependent cells. The complete backbone of the double-flow neural network can be described as follows:

[0098] Step 4.1: Given input Firstly, get By applying a down-sampling module to reduce computational complexity;

[0099] Step 4.2: Apply MSCA to obtain multi-scale attention features.

[0100] These features do not need to be as fine as the semantic segmentation task, which requires pixel-level features. Therefore, the present application removes the feedforward network from the original structure.

[0101] Step 4.3: The present application restores the original input size by transposed convolution to obtain

[0102] C target represents the number of target channels.

[0103] In an embodiment of the present application, the layout feature extraction operation includes:

[0104] Obtain the text box of the target table image to be recognized, and the layout feature is calculated based on the boundary box coordinates of the text box;

[0105] The layout feature includes but is not limited to the width, height, and center point coordinate information of the boundary box.

[0106] It should be noted that the original target table image to be recognized and the text embedding mapping constructed above are fed into a double-flow neural network that uses the proposed special backbone network for table to obtain visual and text feature patterns.

[0107] It should also be noted that obtaining the target table image to be recognized in the power industry and performing the first feature extraction operation on the target table image to be recognized can significantly improve the accuracy and efficiency of table structure recognition. Through the combination of the improved deep learning model and natural language processing technology, the present application can comprehensively capture the visual, text, and layout features in the table image, providing a rich information foundation for subsequent feature fusion and reasoning. In addition, by fine-tuning the word embedding model and the pre-trained language model, the present application can adapt to the recognition needs of specific table structures in the power industry, further improving the accuracy and efficiency of recognition.

[0108] S102, performing a second feature fusion operation on the result of the first feature extraction, the second feature fusion operation including first-stage feature fusion and second-stage feature fusion, wherein:

[0109] It should be noted that after the completion of the individual feature extraction, the different features need to be subjected to feature fusion operation, and the feature fusion in this part is to organically combine the visual features, the text features and the layout features, so as to fully utilize the respective advantages and improve the accuracy and robustness of the table structure recognition.

[0110] In some specific embodiments, the feature fusion operation can be realized by a feature fusion network in a deep learning framework, for example, using an attention mechanism or a gating mechanism to dynamically adjust the weights of different features, so that important features can be highlighted and irrelevant features can be suppressed in the fusion process.

[0111] In some specific embodiments, the feature fusion operation can be realized by a feature fusion network in a deep learning framework, for example, using an attention mechanism or a gating mechanism to dynamically adjust the weights of different features, so that important features can be highlighted and irrelevant features can be suppressed in the fusion process.

[0112] It should be noted that because the table structure ultimately required by the present application is closely related to the layout features, the layout features can be used as the main fusion features, and therefore the present application preferentially fuses the visual features and the text features, and finally fuses the layout features, so that the final features obtained are closely related to the table structure.

[0113] Therefore, two-stage fusion operation can be set to realize the second feature fusion operation.

[0114] In the embodiment of the present application, the first stage feature fusion is used to fuse the visual features and the text features;

[0115] In the embodiment of the present application, the second stage feature fusion is used to fuse the first stage feature fusion result and the layout features;

[0116] In some specific embodiments, the first stage feature fusion and the second stage feature fusion can adopt the same feature fusion network or different feature fusion networks to realize. For example, in the first stage feature fusion, an attention mechanism can be used to dynamically adjust the weights of the visual features and the text features, so that important features can be highlighted and irrelevant features can be suppressed in the fusion process. In the second stage feature fusion, a feature splicing or feature addition method can be used to combine the first stage feature fusion result and the layout features, so as to form a more comprehensive and rich feature representation.

[0117] In some specific embodiments, the first-stage feature fusion and the second-stage feature fusion can specifically employ a convolutional neural network (CNN) or a recurrent neural network (RNN) in a deep learning model, or other network structures. In the first-stage feature fusion, the CNN can be used to extract spatial information in the visual features, while the RNN can be used to capture sequential information in the text features. By combining the two network structures, the visual features and the text features can be more effectively fused. In the second-stage feature fusion, a more complex deep learning model, such as a Transformer, can be employed to perform deep interaction between the first-stage feature fusion results and the layout features, so as to further improve the accuracy and robustness of the feature representation.

[0118] In the embodiments of the present application, the second feature fusion operation comprises:

[0119] a preset gate control parameter, which is used to control the contribution of different features;

[0120] establishing a first-stage feature fusion expression based on the gate control parameter, the visual features, and the text features;

[0121] performing region feature extraction on the first-stage feature fusion results, and establishing a model representing the exchange relationship between the region feature extraction operation and the layout features.

[0122] Specifically, in the first stage, the visual and text features merged by the present application are in the level of feature maps; in the second stage, the present application performs layout functions at the element level, and finally obtains fused functions for each element in three modalities.

[0123] 1) The first-stage feature fusion can be based on a learnable attention map, and an adaptive fusion module is used to fuse the features of the two modalities.

[0124] First feat = Att V feat + (1-Att) T feat (2)

[0125] wherein Att is a learnable parameter (i.e., the gate control parameter in the present application), which plays a gating role and controls the contribution of each modality. represents element-level multiplication. V feat and T feat respectively represent the feature maps obtained through the double-flow network. First, a feature vector f is generated by f = Att V + (1-Att) T

[0126] 2) Second stage feature fusion, in order to obtain the first stage fusion features corresponding to each element position, the RoIPooling method is used for region feature extraction. Subsequently, the Kronecker product is used to model the interaction between all possible first stage fusion features and layout features.

[0127]

[0128] In the above formula, L is a learnable linear transformation that converts the feature dimension to the target dimension of the application. is the Kronecker product operation. represents the layout feature of the i-th element. As known, the Kronecker product operation can be storage-intensive when considering all combinations of products.

[0129] It should be noted that the second feature fusion operation on the results of the first feature extraction can improve the accuracy and efficiency of complex table structure recognition. Through the second stage feature fusion, the visual features are combined with the text features and the layout features, which can more comprehensively capture the information in the table, providing more rich feature representation for subsequent classification, parsing and other steps. This fusion method helps the model better understand the structure and content of the table, thereby improving the overall recognition performance.

[0130] S103, context aggregation and table image structure inference are performed on the results of the second feature fusion operation, wherein:

[0131] It should be noted that after obtaining the fusion features after the second feature fusion operation, further data acquisition related to the table structure is needed, including the cell position in the table, the row and column relationship, and the correlation between cells, etc. In order to more accurately parse the table structure, the application adopts the method of context aggregation and table image structure inference.

[0132] In the embodiments of the application, context aggregation is used to establish the context relationship between the elements of the target table image to be recognized, and table image structure inference is used to obtain the row, column and cell relationship between the elements in the target table image to be recognized.

[0133] In some specific embodiments, context aggregation can be implemented using a graph neural network (GNN) or its variants, by constructing a graph structure of table elements and using the edge relationship between nodes to transfer and aggregate context information. This method can effectively capture the spatial relationship and dependency between table elements, providing strong support for subsequent structure inference.

[0134] In other specific embodiments, the context aggregation can also be implemented using attention mechanisms or self-attention mechanisms, by calculating attention weights between table elements to highlight key elements and suppress secondary elements, thereby enhancing the aggregation of context information. This method can more flexibly capture the relevance and importance between elements when processing complex tables, improving the accuracy of structure recognition. Whether using graph neural networks or attention mechanisms, the purpose of context aggregation is to better understand and parse table structures, providing a solid foundation for subsequent analysis and application.

[0135] In some specific embodiments, when using a graph neural network (GNN) for context aggregation, the specific steps can be as follows:

[0136] Step 5.1: Construct the graph structure of the target table image to be recognized, where nodes represent table elements (such as cells, rows, columns, etc.), and edges represent the relationship between elements (such as adjacency, inclusion, etc.).

[0137] Step 5.2: Use the node update mechanism of GNN to update the feature representation of each node by iteratively passing and aggregating the information of neighboring nodes.

[0138] Step 5.3: In each iteration, the node updates its own features based on the information of its neighbor nodes, gradually capturing more extensive context information.

[0139] Step 5.4: Aggregate the feature representations of all nodes to obtain the context representation of the entire table, providing input for subsequent structure reasoning.

[0140] However, in the present invention, in order to better adapt to the characteristics of complex table structures, an improved context aggregation method is proposed, which combines the advantages of graph neural networks and attention mechanisms. The context aggregation in the present invention includes local context aggregation and global context aggregation;

[0141] The table elements after the second feature fusion operation are taken as vertices, and the connections between the table elements are taken as edges;

[0142] Local context aggregation is used to obtain edge features and edge weights, and the local context aggregation includes an improved element distance solving strategy.

[0143] Global context aggregation considers reading order and spatial position as aggregation considerations.

[0144] Specifically, the present invention designs a hybrid context aggregator, which includes a local context module based on graph neural networks and a global context module based on multi-head attention blocks, wherein:

[0145] 1) Local Context Module (LCM): The table can be represented by a graph. Elements are considered as vertices and the relationships between elements are considered as edges. These relationships include "same row", "same column", "same cell", and "no connection". The k-nearest neighbor method is used to select the k-elements closest to the source element to form the element pair. When measuring the distance between two elements, the invention finds that directly using the Euclidean distance will cause the k-elements to be mainly distributed above and below the source element. The distribution of smaller row spacing and larger column spacing leads to this phenomenon, which can cause extreme class imbalance in the training data, resulting in significant differences in the prediction performance of the model in different classes. Therefore, the distance in the horizontal direction is compressed when calculating the Euclidean distance.

[0146] Further, the fused element features are directly used as vertex features, while the edge features and weights are calculated according to the spatial relationship of the node pair. In order to obtain edge features, physical relationship, logical relationship and relative shape information are considered. The formula is as follows:

[0147]

[0148] Further, formulas (4) and (5) show the details of calculating the physical relationship.

[0149] In formula (4), represents the position difference of element i and element j in the "left" direction; represents the position difference of element i and element j in the "top" direction; and represent the coordinate values of element i and element j in the "left" direction; and represent the coordinate values of element i and element j in the "top" direction.

[0150] In formula (5), is the physical relationship feature between element i and element j; respectively represent the position difference of element i and element j in the "left", "right", "top" and "bottom" directions; d is a normalization constant used to standardize the distance; || represents concatenation after normalizing the distance in each direction.

[0151]

[0152] Further, formulas (6) and (7) show the details of calculating the logical relationship.

[0153] In formula (6), represents the starting position difference of element i and element j in the "column" direction; represents the difference in the "left" direction between element i and element j; min(w i ,w j ) represents the minimum of the width of element i and element j, which is used to normalize the column spacing; represents the difference in the "row" direction between element i and element j; represents the difference in the "top" direction between element i and element j; min(h i ,h j ) represents the minimum of the height of element i and element j, which is used to normalize the row spacing.

[0154] In formula (7), is the logical relationship feature between element i and element j, which is used to measure the relative position relationship between the two, mainly reflecting the arrangement and organizational structure of the elements in the table; represents the difference in the starting position in the column direction; represents the difference in the ending position in the column direction; represents the difference in the starting position in the row direction; represents the difference in the ending position in the row direction; || represents splicing after normalizing the distance in each direction.

[0155] w i and h i represent the width and height of the i-th element.

[0156]

[0157] The above formula shows the detailed information for calculating the relative shape information. The present application obtains w and h represent the width and height of the element.

[0158]

[0159] As shown in formula (9), e ij is the edge feature between element i and element j; L() represents a linear function for fusing the features of physical relationship, logical relationship and shape information into a final edge feature; is the physical relationship feature between element i and element j; is the logical relationship feature between element i and element j; is the shape relationship feature between element i and element j.

[0160] It should be noted that in order to obtain the edge weight, the following formula is an important means for obtaining the edge weight, which looks a little similar to the Gaussian function.

[0161]

[0162] where, and are used to initialize row GNN and column GNN, respectively. H and W represent the height and width of the original table image. center center represent the coordinates of the element center

[0163] It should be noted that two types of graph neural networks are used in parallel in order to capture the shallow and deep neighborhood information of nodes. DeeperGCN receives node features and edge features, while GCN receives node features and edge weights.

[0164] 2) Global Context Module (GCM): Multi-head attention mechanism is the core of Transformer work, which enables the model to calculate the similarity between different tokens. In the method, each table element is regarded as a token, and the entire table is modeled as a sequence. The patent introduces a sequential-aware self-attention mechanism and a spatial-aware self-attention mechanism. The new attention score can be represented as follows:

[0165]

[0166] where, ij is the attention score between element i and element j; x i and x j are the feature vectors of element i and element j in the table, respectively; W φ and W K are the weight matrices of the element feature vectors; d k is the dimension of the feature space, used for normalization of the dot product; b 1D is a one-dimensional reading order bias, used to adjust the order relationship between elements, which can be obtained by the enhanced XY split algorithm; j -i is the index difference between element i and element j; are two-dimensional spatial position biases, used to consider the relative positions of elements in space, calculated according to their bounding box coordinates; are the x-coordinates of the top-left corners of the bounding boxes of element i and element j, respectively; are the y-coordinates of the top-left corners of the bounding boxes of element i and element j, respectively. The global context module uses only node features as input token features.

[0167] The final hybrid feature (i.e., the result of context aggregation in the patent) is obtained by fusing the outputs of the local context module (LCM) and the global context module (GCM).

[0168] ​​​​In the embodiment of the present application, the table image structure inference includes: based on the node pairs constructed by the k-nearest neighbor method, the results of the context aggregation of each pair of nodes are spliced in the channel dimension to form a table image structure vector.

[0169] Specifically, based on the node pairs constructed by the k-nearest neighbor method, the final mixed features of each pair of nodes are spliced in the channel dimension to form a vector.

[0170]

[0171] wherein, represents the final edge feature, N represents the number of elements in the table, K is the nearest neighbor parameter set, and d represents the dimension of the final mixed feature. represents the edge feature calculated between node i and its kth nearest neighbor node.

[0172] In the embodiment of the present application, the three-layer fully connected network is used to complete the binary classification prediction.

[0173] In summary, the present application proposes a complex table structure recognition method for the power industry. Through the first feature extraction operation, the present application can comprehensively capture the key information in the target table image to be recognized, including visual features, text features and layout features, providing a rich data basis for subsequent feature fusion and inference. The second feature fusion operation effectively integrates different features. The first stage feature fusion strengthens the connection between visual and text features, making the recognition process more accurate. The second stage feature fusion further integrates the layout features, enhancing the overall understanding of the table structure. The context aggregation step establishes the logical connection between the elements of the table image, which helps to improve the coherence and accuracy of the recognition. The table image structure inference further analyzes the row, column and cell relationships between the elements, further ensuring the accuracy of the recognition. The present application realizes efficient and accurate recognition of complex table structures in the power industry, greatly improving the automation and efficiency of table data processing.

[0174] In a preferred embodiment, the table structure acquisition can be realized by establishing a table structure recognition model based on the first feature extraction operation, the second feature fusion operation, the context aggregation and the table image structure inference.

[0175] The target table image structure to be recognized in the power industry is recognized according to the table structure recognition model.

[0176] The table structure recognition model further includes a loss function improved according to the row, column and cell relationships.

[0177] The table structure recognition model can be built and trained using deep learning frameworks such as TensorFlow or PyTorch. These frameworks provide rich neural network layers and optimization algorithms, making model building and tuning more efficient. During training, a large amount of power industry table data is used for supervised learning to ensure that the model can learn complex table structure features.

[0178] In some specific implementations, the proposed table structure recognition model can be trained in an end-to-end manner. The improved loss function consists of cell relationship loss, row relationship loss, and column relationship loss, and its expression is as follows:

[0179]

[0180] in, and These represent the cell relation loss, row relation loss, and column relation loss, respectively. These losses are all Focal Losses propagated back from the relation inference module. In each Focal Loss, if an edge edge(i,j) belongs to the same relation, it is labeled as category 1; otherwise, it is labeled as category 0. The weight parameters λ and γ are used to control the balance among the three types of losses.

[0181] It should be noted that this invention innovatively employs parallel feature extraction using visual and textual streams. The visual stream models the original image information, while the textual stream combines character-level and sentence-level semantic information to obtain high-quality text representations through a language model, thereby achieving a full perception of multimodal information in power grid table images. To address the differences in contribution of different modalities across different contexts, a first-stage fusion module based on an attention mechanism and a second-stage fusion module based on Kronecker product are designed to achieve deep fusion of visual, textual, and layout features at the image and element levels, significantly improving fusion performance and structural recognition accuracy. Furthermore, it innovatively combines local graph neural networks (GNNs) and global Transformer structures to model short-distance adjacency relationships and long-distance cross-row and cross-column dependencies, respectively, and introduces reading order bias and spatial bias to enhance attention mechanisms, effectively solving the problem of traditional methods struggling to handle complex structural relationships.

[0182] Example 3, referring to Figure 2 This embodiment also provides a complex table structure recognition system for the power industry, including:

[0183] A dual-stream feature extraction network establishment module is used to acquire target table images to be identified in the power industry and to perform the first feature extraction operation on the target table images to be identified.

[0184] The first feature extraction includes a visual feature extraction operation, a text feature extraction operation, and a layout feature extraction operation.

[0185] The two-stage multi-modal feature fusion module is configured to perform a second feature fusion operation on the result of the first feature extraction, and the second feature fusion operation includes a first-stage feature fusion and a second-stage feature fusion.

[0186] The first-stage fused feature is used to fuse the visual feature and the text feature.

[0187] The second-stage feature fusion is used to fuse the first-stage feature fusion result and the layout feature.

[0188] The mixed context aggregation module and the relationship reasoning module are respectively configured to perform context aggregation and table image structure reasoning on the result of the second feature fusion operation.

[0189] The context aggregation is used to establish a context relationship between the target table image elements to be recognized, and the table image structure reasoning is used to obtain a row, column, and cell relationship between the elements in the target table image to be recognized.

[0190] The above modules can be embedded in or independent of the processor in the electronic device in hardware form, or can be stored in the memory in the electronic device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above modules.

[0191] The embodiment also provides an electronic device, which can be a terminal, and an internal structure diagram of the electronic device can be as shown in Figure 2 The electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the electronic device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved through WIFI, a carrier network, NFC (near field communication), or other technologies. The computer program is executed by the processor to implement a complex table structure recognition method for the power industry. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or can be a key, trackball, or touchpad arranged on the shell of the electronic device. In addition, the input device can be an external keyboard, touchpad, or mouse, etc.

[0192] The embodiment also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by the processor to implement the following steps:

[0193] obtaining a target table image to be recognized in the power industry, and performing a first feature extraction operation on the target table image to be recognized;

[0194] The first feature extraction includes a visual feature extraction operation, a text feature extraction operation, and a layout feature extraction operation;

[0195] performing a second feature fusion operation on the result of the first feature extraction, the second feature fusion operation including a first stage feature fusion and a second stage feature fusion;

[0196] The first stage feature fusion is used to fuse the visual features and the text features;

[0197] The second stage feature fusion is used to fuse the first stage feature fusion result and the layout features;

[0198] performing context aggregation and table image structure reasoning on the result of the second feature fusion operation;

[0199] The context aggregation is used to establish the context relationship between the elements of the target table image to be recognized, and the table image structure reasoning is used to obtain the row, column, and cell relationship between the elements in the target table image to be recognized.

[0200] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and all modifications and equivalents should be included in the scope of the claims of the present application.

[0201] Although the preferred embodiments of the present application have been described, those skilled in the art can make further changes and modifications to the embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to be interpreted as including all the preferred embodiments and all the changes and modifications falling within the scope of the present application.

[0202] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A method for recognizing a complex table structure for the power industry, characterized by, The method comprises: acquiring a target table image to be recognized in the power industry, and performing a first feature extraction operation on the target table image to be recognized; the first feature extraction comprises a visual feature extraction operation, a text feature extraction operation, and a layout feature extraction operation; performing a second feature fusion operation on the result of the first feature extraction, the second feature fusion operation comprising a first-stage feature fusion and a second-stage feature fusion; the first-stage feature fusion is used for fusing visual features and text features; the second-stage feature fusion is used for fusing the first-stage feature fusion result and layout features; performing context aggregation and table image structure reasoning on the result of the second feature fusion operation; the context aggregation is used to establish a context relationship between elements of the target table image to be recognized, and the table image structure reasoning is used to acquire a row, column, and cell relationship between elements in the target table image to be recognized.

2. The method of claim 1, wherein the method is for a complex table structure recognition for the power industry. Further comprising: establishing a table structure recognition model based on the first feature extraction operation, the second feature fusion operation, the context aggregation, and the table image structure reasoning; performing table structure recognition of the target table image to be recognized in the power industry according to the table structure recognition model; the table structure recognition model further comprises a loss function improved according to the row, column, and cell relationship.

3. The method of claim 2, wherein the method is for a complex table structure recognition for the power industry. The second feature fusion operation comprises: a preset gate control parameter is used to control the contribution of different features; a first-stage feature fusion expression is established based on the gate control parameter, the visual feature, and the text feature; region feature extraction is performed on the first-stage feature fusion result, and a model representing the exchange relationship between the region feature extraction operation and the layout feature is established.

4. The method of claim 3, wherein the method is for a complex table structure recognition for the power industry. The context aggregation comprises local context aggregation and global context aggregation; the table elements after the second feature fusion operation are taken as vertices, and the relationship between the table elements is taken as edges; the local context aggregation is used to acquire edge features and edge weights, and the local context aggregation comprises an improved element distance solving strategy; the global context aggregation takes reading order and spatial position as aggregation considerations.

5. The method of claim 4, wherein the method is for a complex table structure recognition for the power industry. The table image structure reasoning comprises:

6. The method of claim 5, wherein the method is for a complex table structure recognition for the power industry. based on a node pair constructed by a k-nearest neighbor method, the result of the context aggregation of each node pair is spliced in a channel dimension to form a table image structure vector. The text feature extraction operation comprises: acquiring text information of the target table image to be recognized; mapping the text information into an embedding map of the same size as the original target table image to be recognized; the embedding map comprises a character embedding map and a sentence embedding map, which are acquired by a word embedding model and a pre-trained language model, respectively; 7. The method of claim 6, wherein the method is for a complex table structure recognition for the power industry. the character embedding map and the sentence embedding map are fused to obtain text features. The layout feature extraction operation comprises: acquiring a text box of the target table image to be recognized, and the layout feature is calculated based on the bounding box coordinates of the text box; 8. A complex table structure recognition system for the electric power industry, applying the method according to any one of claims 1 to 7, characterized in that, the layout feature comprises, but is not limited to, width, height, and center point coordinate information of the bounding box. The double-flow feature extraction network establishment module is configured to acquire a target table image to be recognized in the power industry and perform a first feature extraction operation on the target table image to be recognized. The first feature extraction includes a visual feature extraction operation, a text feature extraction operation, and a layout feature extraction operation. The two-stage multi-modal feature fusion module is configured to perform a second feature fusion operation on a result of the first feature extraction, and the second feature fusion operation includes a first-stage feature fusion and a second-stage feature fusion. The first-stage fused features are used to fuse visual features and text features. The second-stage feature fusion is used to fuse the first-stage feature fusion result and layout features. A hybrid context aggregation module and a relationship reasoning module are respectively configured to perform context aggregation and table image structure reasoning on a result of the second feature fusion operation. The context aggregation is used to establish context connections between elements of the target table image to be recognized, and the table image structure reasoning is used to acquire row, column, and cell relationships between elements in the target table image to be recognized. 9.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the computer program to implement the steps of the complex table structure recognition method for the power industry according to any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the complex table structure recognition method for the power industry according to any one of claims 1-7.

Citation Information

Cited By

  • Table structure identification method, model training method, system and equipment

    CN122049931A