Power grid field table recognition and information extraction method based on large model
By applying a multimodal fusion intelligent table analysis system based on large models in the power grid field, the problems of fuzzy data source and irregular table structure in complex tables are solved, efficient and accurate table information extraction is achieved, and the efficiency of grid management and emergency treatment is improved.
Patent Information
- Application Number
- CN202411901246.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
AI Technical Summary
When the prior art deals with complex and diverse power grid tables, it is often due to problems such as fuzzy data sources, irregular table structures and information across units, which leads to the inability to accurately identify or extract key data, affecting the safety and stability of power grid operation.
The multimodal fusion intelligent table analysis system based on large models is adopted, combined with visual big models and natural language processing big models, and through steps such as image preprocessing, table structure analysis, multimodal information fusion recognition, dynamic error correction and consistency inspection, structured data output and verification, etc., the spatial structure and semantic information of complex tables are accurately parsed to ensure the accuracy and completeness of the data.
It significantly improves the efficiency and accuracy of table analysis, reduces manual intervention, ensures rapid acquisition and accurate transmission of data, provides reliable support for timely scheduling and decision-making of the power grid, and improves the efficiency and emergency response capabilities of power grid management.
Smart Images

Figure CN119942574A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power grid technology, and in particular to a method for identifying power grid field tables and extracting information based on a large model. Background Art
[0002] Table recognition and information extraction in the power grid field based on large models refers to the use of large visual models (such as OCR and deep learning technology) to intelligently analyze complex tables in the power industry. Through the efficient recognition ability of the model, key information (such as power equipment parameters, operating data, scheduling information, etc.) can be extracted from the table image, and structured to generate digital data for subsequent analysis. This technology solves the problem of low efficiency in traditional manual table processing and has important application value in scenarios such as smart grid management, equipment monitoring and fault diagnosis.
[0003] The prior art has the following deficiencies:
[0004] In the prior art, when dealing with complex and diverse tables, table recognition and information extraction in the power grid field often fail to accurately identify or extract key data due to problems such as fuzzy data sources, irregular table structures, and cross-unit distribution of information. This limitation is particularly prominent in power grid emergency management scenarios. Especially in disasters or emergencies, information extraction errors may directly lead to inaccurate dispatching instructions and chaotic load distribution, thereby causing power supply delays or large-scale interruptions, causing serious social and economic losses. In addition, these erroneous data may also accumulate in subsequent analysis links, causing long-term interference to power grid operation optimization and fault prevention, and further threatening the safe and stable operation of the power system.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not constitute the prior art that is already known to one of ordinary skill in the art. Summary of the invention
[0006] The purpose of the present invention is to provide a method for table recognition and information extraction in the field of power grids based on a large model. Through a multimodal fusion intelligent table parsing system, combined with a large visual model and a large natural language processing model, the spatial structure and semantic information of complex tables can be accurately parsed, greatly improving the parsing efficiency and accuracy, and reducing manual intervention. With the support of deep learning and logical reasoning technology, the system can automatically detect and repair semantic and logical errors, ensure the integrity and reliability of table data, and is suitable for processing abnormal data in power grid operation. In addition, through structured conversion and standardized output, the system achieves seamless docking with the dispatching platform, supports automated dispatching and monitoring, and verifies the integrity of the data through advanced statistical analysis, significantly improving the efficiency of power grid management and emergency handling capabilities to solve the problems in the above-mentioned background technology.
[0007] In order to achieve the above-mentioned purpose, the present invention provides the following technical solution: a method for identifying a table in the field of power grid and extracting information based on a large model, which is used to solve the problems of fuzzy data sources, irregular table structures, and cross-unit distribution of information in complex tables, and avoid the negative impact of information extraction errors on power grid operation, comprising the following steps:
[0008] Image preprocessing: Obtain table images in the power grid area, improve image clarity through image enhancement algorithms, eliminate noise and background interference, and perform tilt correction and resolution optimization on the table to generate standardized images suitable for subsequent processing;
[0009] Table structure analysis: Use a large model based on deep learning to analyze the structure of the table, build a spatial layout model of the table cells, and automatically identify the table boundaries, cell locations, and merged areas through a row-column mapping algorithm;
[0010] Multimodal information fusion recognition: Combining the visual big model with the natural language processing big model, the multimodal deep fusion algorithm simultaneously processes the image information and text information in the table, accurately recognizes the text content, table units and special symbols, so as to achieve accurate information extraction;
[0011] Dynamic error correction and consistency check: A dynamic correction algorithm based on a neural network is used to check the semantic and logical consistency of the recognized table data, automatically detect and repair errors in information extraction, and ensure the accuracy and integrity of the table data;
[0012] Structured data output and verification: Convert the identified and corrected tabular information into a standardized structured data format, and use advanced statistical analysis algorithms to verify the extracted results to ensure that the output data meets the application requirements of power grid management and emergency dispatch.
[0013] Preferably, a table image of the power grid field is obtained, and the image clarity is improved by an image enhancement algorithm. The specific steps are as follows:
[0014] Use a high-resolution scanner, industrial camera or mobile device camera to obtain table images in the power grid field and input them into the system. The acquisition process must ensure the integrity of the table image to avoid information loss or distortion due to external factors such as shooting angle, insufficient light or table wrinkles. At the same time, the resolution and focal length of the acquisition device should be adjusted according to the complexity of the table content. For example, for tables containing tiny characters or complex lines, high-resolution devices should be selected to ensure clear presentation of details. In addition, in order to meet the needs of diverse table types, tables from different sources can be uniformly formatted and stored (such as JPEG, PNG, etc.) for subsequent processing. The core role of this step is to obtain high-quality original data and avoid the impact of low-quality input on subsequent processing.
[0015] After the image is input, the image is preprocessed using advanced image enhancement algorithms to eliminate noise interference and detect the edges of the table. This step usually uses denoising techniques such as Gaussian filtering and median filtering to deal with background noise, and uses adaptive edge detection algorithms (such as Canny edge detection) to identify the borders and lines of the table. Noise interference may come from the photosensitive noise of the shooting device, shadows in the background of the table, or impurities in the external environment. These noises will cover or blur the key information of the table. Through denoising and edge detection, the frame outline of the table can be effectively extracted, making the table content clearer and laying the foundation for subsequent cell segmentation. This stage plays a key role in improving the accuracy of table recognition.
[0016] After noise elimination, if the table is tilted or the light distribution is uneven during the acquisition process, it is necessary to further optimize the image through geometric correction and brightness equalization processing. The tilt correction technology analyzes the directionality of the edge lines of the table and uses the Hough transform or perspective transform algorithm to adjust the angle of the table to restore it to a horizontal or vertical state. For images with uneven brightness distribution, contrast stretching and histogram equalization techniques can enhance the contrast of the table content and make the text and background more clearly distinguishable. The purpose of these processing steps is to ensure the geometric integrity and visual consistency of the table and improve the usability and accuracy of the image in subsequent recognition.
[0017] Preferably, the table is tilt-corrected and resolution-optimized to generate a standardized image suitable for subsequent processing, and the specific steps are as follows:
[0018] During the process of table image acquisition, the image may be tilted to varying degrees due to shooting angles, uneven table placement, or improper equipment operation. To ensure the geometric alignment of the table image, first use an edge detection algorithm (such as Canny edge detection) to extract the outline information of the table border or lines. Then, use Hough Transform or Fast Fourier Transform (FFT) to analyze the direction angle of the edge lines and determine the tilt angle of the table. Next, based on the detected tilt angle, a rotation transformation technique (such as affine transformation) is used to rotate and correct the table image to restore it to a horizontal or vertical state. The purpose of tilt correction is to eliminate image geometric distortion and ensure that each cell of the table remains neatly arranged, thereby providing a standardized geometric basis for subsequent cell segmentation and information extraction. This step directly affects the recognition accuracy of the table structure by the subsequent algorithm.
[0019] After image correction, the image resolution is optimized to ensure that the table details are clearly visible. Resolution optimization is usually combined with an interpolation algorithm (such as bilinear interpolation or bicubic interpolation) to enlarge the low-resolution image, and a sharpening filter (such as Laplace sharpening or nonlinear enhancement) is used to improve the edge clarity of text and lines. In addition, to avoid blurring during the magnification process, super-resolution reconstruction technology (Super-Resolution) can also be used to generate high-quality details from low-resolution images through a deep learning model. The purpose of this optimization process is to improve the recognition of characters, lines and symbols in the image, especially when dealing with small fonts or complex patterns. The optimized high-resolution image can greatly improve the accuracy and robustness of subsequent table content recognition.
[0020] Preferably, a large model based on deep learning is used to parse the structure of the table, a spatial layout model of the table cells is constructed, and the table boundaries, cell positions and merged areas are automatically identified through a row-column mapping algorithm. The specific steps are as follows:
[0021] The table area in the image is detected through a large model based on deep learning (such as YOLO, Faster R-CNN, or a specific table detection model), and the table is separated from the non-table area. The model learns the characteristics of the table by training a large amount of labeled data, and can accurately identify the outer boundary of the table and its position in the image. For images containing multiple tables, the model can also mark and crop each table area separately. The core function of this step is to ensure that only table-related content is processed, excluding other irrelevant parts (such as comments, titles, or background patterns), and providing a clear input range for subsequent table structure analysis. The segmented table area ensures that the calculations in the subsequent steps are concentrated within the target range, improving processing efficiency and accuracy.
[0022] After the table area is determined, a convolutional neural network (CNN) based on deep learning is used to extract the row and column structure of the table. Specifically, the model generates a row and column mapping matrix by analyzing the line features (horizontal lines, vertical lines, or implicit alignment lines) in the table to identify the number, position, and range of rows and columns in the table. For tables without obvious grid lines, the alignment characteristics of the text can be further identified through the Attention mechanism or a model based on the Transformer architecture to infer the implicit row and column structure. The purpose of row and column structure extraction is to establish a framework model of the table so that the system can clearly distinguish the position and relationship of each cell, laying the foundation for subsequent cell content extraction and merged area analysis.
[0023] After extracting the row and column structure, further analyze the merged cell areas across rows or columns in the table. Through the semantic segmentation capability of the deep learning model, combined with the cell boundary features and text alignment mode, the actual boundaries of each cell and whether it is a merged area are identified. At the same time, a dynamic programming algorithm is used to verify the logical consistency of the merged cells to ensure that there is no overlap or missing in the recognition results. The purpose of this step is to accurately restore the actual layout structure of the table and ensure that the content of each cell can be correctly mapped to its corresponding area, which is especially critical when processing complex reports or heterogeneous tables. This link directly determines the completeness and accuracy of table parsing.
[0024] Preferably, the visual big model and the natural language processing big model are combined, and the image information and text information in the table are processed simultaneously through a multimodal deep fusion algorithm to accurately extract the information. The specific steps are as follows:
[0025] Use a large visual model (such as ViT, Swin Transformer or ResNet) to extract features from the table image and generate a visual feature vector. The model learns the spatial structure, border lines and visual characteristics of the cell content of the table through a multi-layer neural network. At the same time, use OCR (such as Tesseract OCR or Google Vision API) to recognize the text in the image, and use a large natural language processing model (such as BERT, GPT) to further perform semantic analysis and feature encoding on the extracted text information to generate a text feature vector. The extraction of visual features and text features ensures that the image and text information can be represented in a machine-processable form, laying the foundation for multimodal fusion. Through this step, the visual and semantic information of the table can be preliminarily separated to form an independent but related multimodal data representation.
[0026] After obtaining visual and text features, the visual features and text features are integrated through a multimodal deep fusion algorithm (such as a fusion model based on Cross-Attention or Concat-Layer). The fusion process uses visual information to provide spatial references for table structure and layout, and combines the semantic understanding ability of text features to enable the model to identify associations in complex table content. For example, visual features are used to locate cell positions, while text features help identify cell content and its meaning. The fused feature vector can capture the global information and local details of the table, forming a unified cross-modal representation. The purpose of this step is to break through the limitations of single-modal processing and improve the accuracy and robustness of the model in extracting information in complex tables.
[0027] The fused multimodal feature vector is further analyzed by a decoder module (such as Transformer Decoder or GRU) to generate a structured tabular data output. The decoder combines the visual characteristics and text content of the table to parse each cell one by one, identifying its position, content and relationship with other cells. For example, the decoder can mark "merged cells" as a whole area, or associate "header cells" with the corresponding column or row content. In addition, the extracted information can be semantically checked to ensure the logical consistency of the output results. The purpose of this step is to convert high-dimensional feature vectors into a specific, usable structured data format, achieve accurate information extraction, and provide reliable input for subsequent power grid data analysis and application.
[0028] Preferably, a dynamic correction algorithm based on a neural network is used to perform semantic and logical consistency checks on the identified table data, and errors in information extraction are automatically detected and repaired. The specific steps are as follows:
[0029] After the table data is recognized, the extracted cell content is first checked for semantic consistency using a neural network-based semantic analysis model (such as BERT, RoBERTa). The model analyzes the logical association of cell content through contextual understanding, for example, detecting whether the value in a row or column conforms to the expected unit, or whether the title and content match semantically. The model learns the semantic rules of specific fields through pre-trained industry data sets (such as power grid dispatch data, equipment operating parameters, etc.), and judges whether the recognition results are reasonable based on these rules. The purpose of this step is to discover possible semantic errors in the table content, such as unit errors, data type inconsistencies, or contradictions between fields and content, thereby improving the semantic accuracy of the recognition results.
[0030] After completing the semantic check, the logical consistency of the table data is further verified and identified through a dynamic correction algorithm. The neural network model is combined with a graph-based logic analysis algorithm to construct a dependency graph (DependencyGraph) between table cells. For example, it checks whether the increasing relationship of the values in a column conforms to the expected logic, or whether the total value is consistent with its components. The model automatically identifies logical contradictions in the table and marks abnormal cells through dynamic reasoning and iterative verification. The purpose of logical consistency verification is to ensure that the key logical relationships in the table conform to the expected application scenarios (such as the load distribution rules of power grid equipment), thereby eliminating hidden dangers caused by logical errors and providing reliable guarantees for subsequent data processing.
[0031] For detected semantic or logical errors, an error repair module based on a neural network is used for automatic correction. The repair module uses a generative adversarial network (GAN) or an autoregressive model to infer reasonable correction values by referring to the contextual information or associated fields of the table. For example, if it is detected that the value of a cell does not match the logic, the model will repair the error based on the numerical trend or historical data of the adjacent cells. At the same time, for semantic errors (such as unit mismatch), the repair module will call the preset standardized rules to automatically correct the unit and generate a correction record for user verification. The purpose of this step is to reduce manual intervention through an efficient automatic repair mechanism, achieve accuracy and completeness of table data, and provide high-quality input for subsequent power grid data analysis.
[0032] Preferably, the identified and corrected tabular information is converted into a standardized structured data format, and the extraction results are verified using advanced statistical analysis algorithms to ensure that the output data meets the application requirements of power grid management and emergency dispatch. The specific steps are as follows:
[0033] After the table data is identified and corrected, the table content is first converted into a standardized structured data format (such as JSON, CSV, SQL table, etc.). This process includes parsing the rows, columns, cells and their contents of the table, and mapping them to corresponding data fields, attributes and values. For example, the "Device Number", "Run Time" and "Load" columns in the table are converted into fields in the database, and the cell contents are filled in the corresponding positions. In addition, for merged cells or cells that span rows and columns, mapping rules are used to ensure that their contents can be accurately attributed to the standard structure. The purpose of this step is to convert unstructured table information into an easy-to-process digital form, provide a unified data interface for subsequent analysis and application, and improve processing efficiency and compatibility.
[0034] Import structured data into the advanced statistical analysis module, and use statistical methods and machine learning algorithms to verify the accuracy and completeness of the data. For example, use outlier detection algorithms (such as DBSCAN or Isolation Forest) to identify outliers that may be erroneous, or use time series analysis to check whether the fluctuations in operating parameters conform to historical patterns. At the same time, correlation analysis is used to evaluate whether the logical relationships between different fields are as expected, such as the correlation between equipment load and operating time. For detected anomalies, the system will generate a detailed verification report and mark suspected problem data. The purpose of this step is to further ensure the credibility of the data through scientific analysis methods, and provide an accurate and reliable basis for power grid management and dispatching.
[0035] The verified structured data is output to a specified standardized format (such as a dedicated dispatch file format or grid management system interface requirements) based on application requirements. This process includes the unification of naming rules for data fields, standardization of units (such as "kilowatt" converted to "kW"), and format adaptation (such as adjustment to XML or the structure required by a specific API). In addition, metadata tags (such as data source, timestamp, correction information) can be attached to enhance the traceability and use value of the data. The purpose of this step is to ensure that the data can be seamlessly connected to the grid management and emergency dispatch system, support multi-scenario applications such as real-time dispatch and equipment monitoring, so as to achieve efficient use of data and maximize its value.
[0036] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0037] The present invention uses a multi-modal fusion intelligent table parsing system to efficiently parse the spatial structure and semantic information of complex tables, especially when dealing with cross-cells, fuzzy characters and complex layouts. The visual large model accurately parses the layout and boundaries of the table, and the natural language processing large model performs a deep semantic understanding of the text content. The fusion of the two makes up for the shortcomings of single-modal processing, enabling the system to comprehensively parse the table content. Compared with traditional methods, the efficiency and accuracy of table parsing are significantly improved, and the need for manual intervention is greatly reduced. In power grid emergency management and daily operations, it ensures the rapid acquisition and accurate transmission of data, providing reliable support for timely scheduling and decision-making.
[0038] Through deep learning and logical reasoning technology, the present invention can automatically detect semantic and logical errors in tables and perform intelligent repairs based on contextual relationships and domain rules. This error correction mechanism is particularly suitable for handling abnormal problems caused by manual entry or recognition deviations in power grid operation data, ensuring the logical consistency and semantic accuracy of the data, effectively improving the reliability of table data, and providing a solid data foundation for efficient management of the power grid.
[0039] Through structured conversion and standardized output, the present invention can directly convert the parsed and corrected tabular information into a data format that can be used by the dispatching system, and enhance its traceability and compatibility by adding metadata, ensuring the seamless connection between the data and the power grid management platform, so that the dispatching, monitoring and optimization processes can be automated. The system's advanced statistical analysis further verifies the integrity and logical rationality of the data, avoiding operational interruptions or waste of resources caused by data defects. The beneficial effects of this standardization and intelligent adaptation provide important technical support for improving the efficiency of power grid operation and enhancing emergency handling capabilities, helping the power grid system to operate more stably and efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0041] Figure 1 The present invention is a flow chart of the method for identifying power grid field tables and extracting information based on a large model. DETAILED DESCRIPTION
[0042] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of the present disclosure will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art.
[0043] The present invention provides Figure 1 The method for identifying and extracting information from a table in the power grid field based on a large model is used to solve the problems of fuzzy data sources, irregular table structures, and cross-unit distribution of information in complex tables, and to avoid the negative impact of information extraction errors on power grid operation, including the following steps:
[0044] Image preprocessing: Obtain table images in the power grid area, improve image clarity through image enhancement algorithms, eliminate noise and background interference, and perform tilt correction and resolution optimization on the table to generate standardized images suitable for subsequent processing;
[0045] Table structure analysis: Use a large model based on deep learning to analyze the structure of the table, build a spatial layout model of the table cells, and automatically identify the table boundaries, cell locations, and merged areas through a row-column mapping algorithm;
[0046] Multimodal information fusion recognition: Combining the visual big model with the natural language processing big model, the multimodal deep fusion algorithm simultaneously processes the image information and text information in the table, accurately recognizes the text content, table units and special symbols, so as to achieve accurate information extraction;
[0047] Dynamic error correction and consistency check: A dynamic correction algorithm based on a neural network is used to check the semantic and logical consistency of the recognized table data, automatically detect and repair errors in information extraction, and ensure the accuracy and integrity of the table data;
[0048] Structured data output and verification: Convert the identified and corrected tabular information into a standardized structured data format, and use advanced statistical analysis algorithms to verify the extracted results to ensure that the output data meets the application requirements of power grid management and emergency dispatch.
[0049] Implementation method one: This implementation method designs a smart grid table parsing system based on multimodal fusion, which aims to accurately and efficiently extract information from complex tables and provide data support for power grid management and emergency dispatch. First, the system obtains table images in the power grid field from scanners, industrial cameras or mobile devices. To ensure data quality, the system performs a series of preprocessing operations on the image, including image enhancement, tilt correction and noise elimination. Specifically, the system uses Gaussian filtering and median filtering techniques to remove noise, uses Hough transform or adaptive threshold algorithm to correct the tilt angle of the table, and optimizes image contrast through histogram equalization. These processing steps ensure that the table content is clear and geometrically standardized, providing an ideal basis for subsequent parsing.
[0050] After the image processing is completed, the system uses large visual models (such as ViT, Swin Transformer) to extract the spatial layout features of the table and identify the row and column structure, cell boundaries and their positional relationships of the table. For complex tables, the system combines large natural language processing models (such as BERT, RoBERTa) to perform semantic analysis on the text information extracted by OCR and generate text feature vectors. Subsequently, the system uses a multimodal fusion algorithm (such as a model based on the Cross-Attention mechanism) to unify the visual features and text features, so as to accurately parse the content of each cell in the table and its semantic association. For example, the system can understand the meaning of the "load distribution" field from the text features, and confirm the corresponding cell position in combination with the visual features. The core advantage of multimodal fusion is that it breaks through the limitations of single-modal processing, enabling the system to take into account both the structural information and content semantics of the table.
[0051] Finally, the system generates a structured representation of the table, converting the images and text content in the complex table into a standardized JSON or SQL format. This output can be directly used in the power grid management system to achieve fast data docking. Compared with the traditional table parsing method, this implementation method can not only significantly improve the accuracy of complex table parsing, but also handle fuzzy information, merged cells and complex semantic relationships in the table, providing a reliable guarantee for the efficient management of power grid operation data. Through this solution, the efficiency of power grid operators in emergency management, equipment monitoring and dispatch optimization can be significantly improved, while reducing the delay and error risks caused by manual processing.
[0052] Implementation method two: This implementation method proposes a table data repair module for dynamic correction and consistency verification, which is used to perform semantic and logical checks on the identified table data to ensure the accuracy and completeness of the data. First, the system performs a semantic consistency check on the table data. Specifically, a pre-trained semantic analysis model (such as BERT or GPT) is used to perform inference analysis on the semantic relationship of the cell content. For example, for a cell labeled "equipment load (kW)", the model will check whether its unit and value match, and verify whether its context contains logical contradictions (such as the load exceeds the rated capacity). In addition, the system performs comparative analysis through a domain knowledge base (such as power grid equipment parameters or dispatching rules) to quickly identify semantic anomalies in the data. This inspection process can effectively detect field mismatches, unit inconsistencies, or content omissions between cells, thereby improving the semantic integrity of the table data.
[0053] Next, the system further verifies the logical consistency. Through a dynamic correction algorithm based on a neural network, the logical dependency relationship between cells is constructed in combination with a graph model. The system verifies the logical rules involved in the table (such as the increasing rules of operating parameters, the correctness of the total value, etc.) one by one. For example, in the equipment operation report, the system will check whether the load in each time period shows a reasonable change trend over time, or whether the value in the total row is consistent with the sum of the corresponding column. For cells that violate logical rules, the system will mark their abnormal status and pass them to the next repair module. Logical verification ensures the logical consistency between the table contents and further reduces the risk of data errors.
[0054] Finally, the system automatically repairs the detected abnormal cells through a repair module based on a generative adversarial network (GAN). The repair module combines contextual information, historical data, and domain knowledge to infer the most logical correction value. For example, when an abnormal load value of a device is detected, the module generates a reasonable alternative value based on the load trend of the adjacent time period. At the same time, the system records the repair operation and generates a log for manual verification. This implementation method effectively reduces manual intervention through an automated correction process, improves the accuracy of table data, and provides high-quality data support for power grid dispatch and operation optimization.
[0055] Implementation method three: This implementation method proposes a standardized data output and intelligent adaptation system for power grid dispatching, which realizes seamless connection and efficient use of data by standardizing the parsed and corrected tabular data. First, the system performs a structured conversion on the corrected tabular data and maps it to a standardized data format (such as JSON, CSV, or SQL table). Specific operations include: parsing the row and column structure of the table, mapping the content of each cell to a field name and value; splitting or merging the content across cells to ensure that the data matches the structural requirements of the database or API. In addition, the system also standardizes field names, units, etc. For example, "Load (kW)" is uniformly converted to "Load (kW)" to be compatible with the field specifications of the dispatching system. Through this step, the structured data generated by the system can be directly read and processed by downstream systems, significantly improving data circulation efficiency.
[0056] Next, the system performs advanced statistical analysis and verification on the structured data. Using statistical analysis tools (such as Pandas or Scipy) and machine learning models (such as anomaly detection algorithms DBSCAN or Isolation Forest), the system conducts a comprehensive check on the integrity and accuracy of the data. For example, it detects outliers in load data or analyzes whether the equipment operation time conforms to historical patterns. The system also combines time series analysis methods to verify whether the fluctuations in operating parameters are reasonable and check whether the timestamps of the data are continuous. Any problems found will be marked as anomalies by the system, and a detailed verification report will be generated for manual review or automatic correction.
[0057] Finally, the system adapts and outputs structured data according to the needs of power grid management and dispatching. For example, the data format is adjusted to the XML format or REST API transmission format that meets the requirements of the dispatching system interface, and metadata (such as generation time, source equipment, correction records, etc.) is added to the data to enhance the traceability and reliability of the data. In addition, the system supports direct connection of output data to the real-time monitoring platform of the power grid to achieve automated dispatching updates. Through this implementation method, the power grid management department can use data more efficiently for dispatching optimization, emergency handling and equipment maintenance, and improve the overall efficiency and safety of power grid operation.
[0058] Through the intelligent table parsing system of multimodal fusion, the present invention can efficiently parse the spatial structure and semantic information of complex tables, especially when dealing with cross-cells, ambiguous characters and complex layouts. The large visual model accurately parses the layout and boundaries of the table, and the large natural language processing model performs deep semantic understanding of the text content. The fusion of the two makes up for the shortcomings of single modality processing, enabling the system to comprehensively parse the table content. Compared with traditional methods, this solution significantly improves the efficiency and accuracy of table parsing and greatly reduces the need for manual intervention. In power grid emergency management and daily operations, this beneficial effect ensures the rapid acquisition and accurate transmission of data, providing reliable support for timely scheduling and decision-making.
[0059] The dynamic correction and consistency verification module uses deep learning and logical reasoning technology to automatically detect semantic and logical errors in tables, and perform intelligent repairs based on contextual relationships and domain rules. This error correction mechanism is particularly suitable for handling abnormal problems caused by manual entry or recognition deviations in power grid operation data, ensuring the logical consistency and semantic accuracy of the data. For example, for the repair of abnormal values of load data, the system combines the data trends and operation logic of adjacent time periods to generate correction values, effectively eliminating hidden dangers that may affect scheduling decisions. This effect improves the reliability of table data and provides a solid data foundation for efficient management of the power grid.
[0060] Through structured conversion and standardized output, the present invention can directly convert the parsed and corrected tabular information into a data format that can be used by the dispatching system, and enhance its traceability and compatibility by adding metadata. This process ensures the seamless connection between the data and the power grid management platform, enabling the dispatching, monitoring and optimization processes to run automatically. The system's advanced statistical analysis further verifies the integrity and logical rationality of the data, avoiding operational interruptions or waste of resources caused by data defects. The beneficial effects of this standardization and intelligent adaptation provide important technical support for improving the efficiency of power grid operation and enhancing emergency handling capabilities, helping the power grid system to operate more stably and efficiently.
[0061] The above description is only by way of illustration of certain exemplary embodiments of the present invention. It is undoubted that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A method for identifying power grid table and extracting information based on a large model, comprising the following steps: Obtain table images in the power grid area, improve image clarity through image enhancement algorithms, eliminate noise and background interference, and perform tilt correction and resolution optimization on the table to generate standardized images; Use a large model based on deep learning to analyze the structure of the table, build a spatial layout model of the table cells, and automatically identify the table boundaries, cell locations, and merged areas through a row-column mapping algorithm; Combining the visual big model with the natural language processing big model, the multimodal deep fusion algorithm simultaneously processes the image information and text information in the table to achieve accurate information extraction; Adopting a dynamic correction algorithm based on neural network to check the semantic and logical consistency of recognized table data, automatically detect and repair errors in information extraction, and ensure the accuracy and completeness of table data; The identified and corrected tabular information is converted into a standardized structured data format, and the extracted results are verified using advanced statistical analysis algorithms to ensure that the output data meets the application requirements of power grid management and emergency dispatch.
2. The method for identifying power grid domain tables and extracting information based on a large model according to claim 1 is characterized in that: Obtain a table image of the power grid area and improve the image clarity through an image enhancement algorithm. The specific steps are as follows: Acquire a table image of the grid area using a high-resolution scanner, industrial camera or mobile device camera and input it into the system; After the image is input, the advanced image enhancement algorithm is used to pre-process the image, eliminate noise interference and detect the edge of the table; After noise removal, the image is further optimized through geometric correction and brightness equalization to ensure the geometric integrity and visual consistency of the table.
3. The method for identifying power grid domain tables and extracting information based on a large model according to claim 1 is characterized in that: The table is tilted and resolution optimized to generate a standardized image suitable for subsequent processing. The specific steps are as follows: In the process of table image acquisition, the edge detection algorithm is used to extract the outline information of the table border or lines. Then, the Hough transform is applied to analyze the direction angle of the edge line to determine the tilt angle of the table. Then, based on the detected tilt angle, the rotation transformation technology is used to rotate and correct the table image to restore it to a horizontal or vertical state. After image correction, the image resolution is optimized to ensure that the table details are clearly visible.
4. The method for identifying power grid domain tables and extracting information based on a large model according to claim 1 is characterized in that: Use a large model based on deep learning to parse the structure of the table, build a spatial layout model of the table cells, and automatically identify the table boundaries, cell locations, and merged areas through a row-column mapping algorithm. The specific steps are as follows: Detect table areas in images and separate tables from non-table areas using a large model based on deep learning; After the table area is determined, the row and column structure of the table is extracted using a convolutional neural network based on deep learning; After extracting the row and column structure, we further analyze the merged cell areas across rows or columns in the table. Through the semantic segmentation capability of the deep learning model, combined with the cell boundary features and text alignment mode, we can identify the actual boundaries of each cell and whether it is a merged area. At the same time, we use a dynamic programming algorithm to verify the logical consistency of the merged cells to ensure that there is no overlap or missing in the recognition results.
5. The method for identifying power grid domain tables and extracting information based on a large model according to claim 1 is characterized in that: Combining the visual big model with the natural language processing big model, the multimodal deep fusion algorithm is used to simultaneously process the image information and text information in the table to accurately extract the information. The specific steps are as follows: Use the visual big model to extract features from the table image and generate a visual feature vector; After obtaining visual and text features, the visual features and text features are integrated through a multimodal deep fusion algorithm; The fused multimodal feature vector is further analyzed by the decoder module to generate structured tabular data output. The decoder combines the visual characteristics and text content of the table to parse each cell one by one to identify its position, content and relationship with other cells.
6. The method for identifying power grid domain tables and extracting information based on a large model according to claim 1 is characterized in that: A dynamic correction algorithm based on a neural network is used to check the semantic and logical consistency of the recognized table data, and errors in information extraction are automatically detected and repaired. The specific steps are as follows: After the table data recognition is completed, the semantic consistency of the extracted cell content is first checked using a semantic analysis model based on a neural network; After completing the semantic check, the logical consistency of the recognition table data is further verified through the dynamic correction algorithm; For detected semantic or logical errors, they are automatically corrected using a neural network-based error repair module.
7. The method for identifying power grid domain tables and extracting information based on a large model according to claim 1 is characterized in that: Convert the identified and corrected tabular information into a standardized structured data format, and use advanced statistical analysis algorithms to verify the extracted results to ensure that the output data meets the application requirements of power grid management and emergency dispatch. The specific steps are as follows: After the table data recognition and correction are completed, the table content is first converted into a standardized structured data format; Import structured data into the advanced statistical analysis module and use statistical methods and machine learning algorithms to verify the accuracy and completeness of the data; The verified structured data is output into a specified standardized format according to application requirements, ensuring that the data is seamlessly connected to the power grid management and emergency dispatch system, thereby achieving efficient use of the data and maximizing its value.
Citation Information
Cited By
Visual webpage data crawling method and system based on large model
CN120632181A
Document table structure identification method oriented to water conservancy large model retrieval enhancement
CN120633613A
Method and framework for optimizing scanning copy content recognition quality by using large model
CN120689894A
Multi-modal document content cross-platform analysis system
CN120726658A
PDF document intelligent identification and content extraction method based on deep learning
CN120808373A