Intelligent detection method and system for content consistency of multi-modal report materials

By automatically extracting indicators from multimodal reporting materials and verifying cross-modal variation coefficients through an intelligent detection system, the problems of low efficiency and poor accuracy of manual detection in existing technologies are solved, and efficient and accurate consistency detection of power grid reporting materials is achieved.

CN121301615BActive Publication Date: 2026-04-14STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, consistency checks on text, tables, and images in power grid reports and project materials rely on manual verification, which is inefficient and difficult to verify across modalities. This fails to meet the timeliness requirements of rapid decision-making and efficient project management in the power grid and increases the risk of erroneous data inflow.

Method used

This paper provides an intelligent method and system for content consistency detection of multimodal reporting materials. By acquiring and extracting indicators of text, table and image data, configuring consistency detection coefficients, and verifying cross-modal variation coefficients, it achieves automated consistency detection.

Benefits of technology

It enables automated consistency verification of multimodal data, improves detection efficiency and accuracy, reduces the time and subjective error of manual inspection, and enhances the intelligent level of quality control of report materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301615B_ABST
    Figure CN121301615B_ABST
Patent Text Reader

Abstract

The application provides a content consistency intelligent detection method and system for multi-modal report materials, and belongs to the technical field of data processing. The method comprises the following steps: obtaining a report material to perform index extraction, obtaining a plurality of text indexes, a plurality of table indexes and a plurality of picture indexes; configuring a first consistency detection coefficient, obtaining a plurality of intersection index table data and a plurality of table intersection index change coefficients, and configuring a second consistency detection coefficient; obtaining a plurality of picture intersection index change coefficients and configuring a third consistency detection coefficient; fusing the consistency detection coefficients and performing verification to obtain a consistency detection result. The application solves the technical problems that the consistency detection of the text, table and picture data of the report material in the prior art relies on manual verification, the detection efficiency is low, and it is difficult to perform cross-modal correlation verification, and achieves the technical effects of automatically extracting multi-modal intersection indexes and performing cross-modal change coefficient verification, improving the consistency detection efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for intelligent content consistency detection of multimodal reporting materials. Background Technology

[0002] With the continuous improvement of the digitalization and intelligence level of the power grid, various multimodal reporting materials have become important bases for power grid operation management, science and technology project management, and decision analysis. These multimodal reporting materials include operation and technical documents such as power grid operation analysis reports, equipment status assessment reports, and power grid planning reports, as well as project management documents such as power grid science and technology project contracts, task books, feasibility study reports, and project acceptance data. These reporting materials typically contain various forms of content, including text descriptions, data tables, and visualization charts, to present the power grid operation status, project implementation progress, and technological achievements from different dimensions.

[0003] However, the consistency verification of power grid reporting materials and project data currently relies mainly on professional personnel comparing each item one by one. Testing personnel need to repeatedly switch between text paragraphs, data tables, and visualization charts, manually identifying whether the values, descriptions, and trends of the same indicators or content are consistent across different formats. This manual testing method is inefficient, susceptible to subjective factors, and lacks cross-modal correlation verification mechanisms. It fails to meet the timeliness requirements of rapid decision-making and efficient project management in the power grid, and also increases the risk of erroneous data flowing into power grid dispatching decisions and project management processes. Summary of the Invention

[0004] This invention addresses the technical problems in existing technologies where consistency detection of text, tables, and images in report materials relies on manual verification, resulting in low detection efficiency and difficulty in cross-modal correlation verification. It provides an intelligent method and system for content consistency detection of multimodal report materials to solve these problems.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0006] In a first aspect, the present invention provides an intelligent content consistency detection method for multimodal report materials, comprising: acquiring report materials, extracting text data, table data, and image data and extracting indicators to obtain multiple text indicators, multiple table indicators, and multiple image indicators, wherein the indicators are variable data indicators; filtering to obtain multiple intersection indicators, configuring a first consistency detection coefficient, extracting table data to obtain multiple intersection indicator table data, performing change analysis to obtain multiple table intersection indicator change coefficients, and configuring a second consistency detection coefficient; identifying changes in multiple intersection indicators in image data to obtain multiple image intersection indicator change coefficients, and configuring a third consistency detection coefficient; fusing the configured consistency detection coefficients, extracting associated text data and identifying changes in intersection indicators to obtain multiple text intersection indicator change coefficients, verifying them with the multiple table intersection indicator change coefficients and the multiple image transaction indicator change coefficients to obtain a consistency detection result.

[0007] Secondly, this invention provides an intelligent content consistency detection system for multimodal report materials, comprising: a multimodal indicator extraction module, used to acquire report materials, extract text data, table data, and image data, and extract indicators to obtain multiple text indicators, multiple table indicators, and multiple image indicators, wherein the indicators are variable data indicators; a table consistency analysis module, used to filter and obtain multiple intersection indicators, configure a first consistency detection coefficient, extract table data to obtain multiple intersection indicator table data, perform change analysis to obtain multiple table intersection indicator change coefficients, and configure a second consistency detection coefficient; an image consistency analysis module, used to identify changes in multiple intersection indicators of image data, obtain multiple image intersection indicator change coefficients, and configure a third consistency detection coefficient; and a fusion verification detection module, used to fuse the configured consistency detection coefficients, extract associated text data and identify changes in intersection indicators to obtain multiple text intersection indicator change coefficients, and verify them with the multiple table intersection indicator change coefficients and multiple image transaction indicator change coefficients to obtain consistency detection results.

[0008] The beneficial effects of this invention are:

[0009] First, report materials are acquired, and textual, tabular, and image data are extracted and metrics are extracted to obtain multiple textual, tabular, and image metrics, thus establishing a multimodal data metric system. Next, multiple intersection metrics are selected, and a first consistency detection coefficient is configured. Tabular data is extracted to obtain multiple intersection metric tables, and change analysis is performed to obtain multiple table intersection metric change coefficients. A second consistency detection coefficient is configured to determine the change characteristics of common metrics in the table modality and establish a detection benchmark. Then, multiple intersection metrics are identified in the image data to obtain multiple image intersection metric change coefficients. A third consistency detection coefficient is configured to obtain the change characteristics of common metrics in the image modality. Finally, the configured consistency detection coefficients are fused, and associated textual data extraction and intersection metric change identification are performed to obtain multiple textual intersection metric change coefficients. These are then verified with multiple table and image intersection metric change coefficients to obtain consistency detection results, thus achieving automated consistency verification based on cross-modal change coefficient similarity.

[0010] The above technical solution solves the technical problems in the existing technology of relying on manual verification for consistency detection of text, table and image data in report materials, which is inefficient and difficult to verify across modalities. It achieves the technical effect of automatically extracting multimodal intersection indicators and verifying cross-modal change coefficients, thereby improving the efficiency and accuracy of consistency detection. Attached Figure Description

[0011] Figure 1 A flowchart illustrating the intelligent content consistency detection method for multimodal reporting materials provided by this invention;

[0012] Figure 2 This is a schematic diagram of the structure of the intelligent detection system for content consistency of multimodal reporting materials provided by the present invention.

[0013] In the attached diagram, the components represented by each number are as follows:

[0014] Multimodal index extraction module 11, table consistency analysis module 12, image consistency analysis module 13, fusion verification and detection module 14. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0017] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0018] Example 1, as Figure 1 As shown, embodiments of the present invention provide an intelligent method for content consistency detection of multimodal reporting materials, including:

[0019] S1. Obtain report materials, extract text data, table data and image data and extract indicators to obtain multiple text indicators, multiple table indicators and multiple image indicators, among which the indicators are change data indicators.

[0020] Specifically, the first step is to obtain the report materials to be tested. These materials can be operational technical documents such as daily power system operation reports, monthly load analysis reports, and annual grid operation reports, or project management documents such as grid technology project contracts, task books, feasibility study reports, and project acceptance data. All of these are power industry documents containing multimodal data. The report materials are usually in electronic document format, which may include, but is not limited to, PDF, Word, Excel, or combinations thereof. After obtaining the report materials, their content is analyzed and categorized, dividing the content into three categories based on data format: text data, tabular data, and image data.

[0021] Textual data refers to content presented in the report in the form of textual descriptions, such as "The system peak load this month was 5200MW, an increase of 8% compared to last month" or "The power grid operated smoothly without any major faults." Tabular data refers to structured power data organized in tabular form in the report, such as tables recording peak load, average load, and minimum load over multiple days, or statistical tables recording the load distribution of each substation. Visual data refers to visual content presented in graphical and chart form in the report, such as line graphs showing load change trends or pie charts showing the percentage of electricity consumption. In power grid technology project materials, textual data is the textual description of project technical indicators, investment scale, and construction period in the project contract or feasibility study report, such as "The total investment of the project is 35 million yuan, and the construction period is 24 months." Tabular data can be statistical tables recording project budget allocation, milestone nodes, and performance evaluation indicators. Visual data can be flowcharts showing the project's technical route, framework diagrams of the project implementation plan, project organizational structure diagrams, and other visual content.

[0022] After data extraction, indicator extraction processing was performed on text data, tabular data, and image data respectively. Indicators refer to specific power parameters involved in power system reports, and these indicators are variable data indicators, specifically those that change periodically, such as power operation parameters like "peak load," "average load," "load factor," "voltage level," "power factor," and "line loss rate." In power grid technology project data, indicators include project management parameters such as "project investment amount," "construction period," "completion rate of technical indicators," and "number of achievements." Through indicator extraction processing, multiple text indicators were identified and extracted from text data, multiple tabular indicators from tabular data, and multiple image indicators from image data, providing foundational data for subsequent consistency testing.

[0023] S2. Filter to obtain multiple intersection indicators, configure the first consistency detection coefficient, extract the table data to obtain multiple intersection indicator table data, perform change analysis to obtain the change coefficient of multiple table intersection indicators, and configure the second consistency detection coefficient.

[0024] Specifically, firstly, the intersection of multiple textual, tabular, and image indicators is analyzed to identify common indicators that exist simultaneously in the textual, tabular, and image data, thus obtaining multiple intersection indicators. For example, in a power system report, the textual data mentions five indicators: "peak load," "average load," "load factor," "daily maximum load," and "nighttime minimum load." The tabular data records four indicators: "peak load," "average load," "minimum load," and "voltage qualification rate." The image data shows the change curves of three indicators: "peak load," "load factor," and "average load." The intersection indicators of these three are "peak load" and "average load."

[0025] The number of intersection indicators reflects the degree of correlation between different modalities of data in the report material. More intersection indicators indicate a higher degree of overlap between the text, table, and image descriptions, providing a more solid foundation for data consistency verification and allowing for more lenient detection standards. Conversely, fewer intersection indicators require stricter detection standards to ensure accuracy. Based on this principle, the ratio of the number of intersection indicators to the number of text, table, and image indicators is calculated, and the minimum value is taken as the indicator consistency coefficient. For example, if there are 2 intersection indicators, 5 text indicators, 4 table indicators, and 3 image indicators, the calculated ratios are 2 / 5 = 0.4, 2 / 4 = 0.5, and 2 / 3 = 0.67, respectively. The minimum value of 0.4 is taken as the indicator consistency coefficient. Then, a first consistency detection coefficient is configured based on the indicator consistency coefficient. Specifically, the first consistency detection coefficient is calculated by subtracting the indicator consistency coefficient from the value 1. For example, when the indicator consistency coefficient is 0.4, the first consistency detection coefficient is 1 - 0.4 = 0.6. When the consistency coefficient of an indicator is large, the first consistency detection coefficient is small, indicating that a relatively lenient consistency detection can be performed; when the consistency coefficient of an indicator is small, the first consistency detection coefficient is large, indicating that a more stringent consistency detection is required.

[0026] After configuring the first consistency detection coefficient, data is extracted from the tabular data according to multiple intersection indicators. For example, if the intersection indicators include "peak load" and "average load," all data records corresponding to these two indicators are extracted from the tabular data to obtain multiple intersection indicator tabular datasets. For each intersection indicator tabular dataset, change analysis calculations are performed. Taking "peak load" as an example, if the table records peak load data for multiple consecutive days, such as 5000MW on the first day, 5200MW on the second day, 5100MW on the third day, 5300MW on the fourth day, and 5250MW on the fifth day, the change magnitude of this indicator over the time series is calculated. Calculate the absolute value of the rate of change between adjacent data points. The rate of change from day 1 to day 2 is |(5200-5000) / 5000| = 4%, from day 2 to day 3 is |(5100-5200) / 5200| = 1.92%, from day 3 to day 4 is |(5300-5100) / 5100| = 3.92%, and from day 4 to day 5 is |(5250-5300) / 5300| = 0.94%. Then calculate the average of all rates of change, i.e., (4% + 1.92% + 3.92%). +0.94%) / 4=2.7%. Similarly, a change analysis calculation is performed on the "average load". Assuming the table records 3800MW on the first day, 3900MW on the second day, 3850MW on the third day, 4000MW on the fourth day, and 3950MW on the fifth day, the absolute values ​​of the adjacent change rates are 2.63%, 1.28%, 3.90%, and 1.25%, respectively, and the average value is (2.63%+1.28%+3.90%+1.25%) / 4=2.27%. The change coefficient of the intersection index of the "average load" table is 2.27%.

[0027] Subsequently, the dispersion parameter of the change coefficients of the intersection indicators of multiple tables is calculated to obtain the table change dispersion parameter. The table change dispersion parameter reflects the degree of dispersion of the change amplitude of different intersection indicators in the table data, and can be measured by the standard deviation. For example, the change coefficient of the table intersection indicator for "peak load" is 2.7%, and the change coefficient of the table intersection indicator for "average load" is 2.27%. Then, the standard deviation of these two change coefficients, 0.215%, is calculated as the table change dispersion parameter. The smaller the table change dispersion parameter, the closer the change amplitude of each intersection indicator in the table, and the better the internal consistency of the table data.

[0028] Next, the ratio of the table's variation dispersion parameter to the maximum variation dispersion parameter of the report material is calculated as the second consistency detection coefficient. The maximum variation dispersion parameter is a preset empirical value or a maximum dispersion threshold obtained from historical report materials, used as a normalization benchmark. For example, assuming the maximum variation dispersion parameter of the report material is preset to 0.5%, then the second consistency detection coefficient = 0.215% / 0.5% = 0.43. The second consistency detection coefficient reflects the degree of consistency of the changes in various indicators within the table data. The smaller the second consistency detection coefficient, the more consistent the changing trends of the intersecting indicators in the table, and a more lenient standard can be used in subsequent detection.

[0029] S3. Identify changes in multiple intersection indicators of the image data, obtain the change coefficients of multiple image intersection indicators, and configure the third consistency detection coefficient.

[0030] Specifically, firstly, an image indicator change recognizer is acquired to identify the magnitude of changes in overlapping indicators within image data. This image indicator change recognizer is built upon a convolutional neural network and trained under supervised supervision using a set of sample indicators, a set of sample image data, and a set of sample indicator change coefficients. During training, a large number of chart images containing electricity indicators are used as sample image data, and the actual change coefficients of the corresponding indicators in each image are labeled. Through deep learning, the model is able to automatically identify and quantify the magnitude of indicator changes from various statistical charts such as line charts, bar charts, and pie charts.

[0031] After obtaining the image indicator change recognizer, the image data is combined with multiple intersection indicators and input into the image indicator change recognizer for identification. For example, for the intersection indicators "peak load" and "average load," image data containing these two indicators in the report material are input into the recognizer. Taking "peak load" as an example, if the image is a line graph showing the peak load change trend over five consecutive days, the image indicator change recognizer extracts the fluctuation characteristics of the line using image recognition technology, calculates the change amplitude between adjacent data points, and outputs the image intersection indicator change coefficient for this indicator, for example, 2.8%. Similarly, the chart for "average load" is recognized, and its image intersection indicator change coefficient is output, for example, 2.4%. By recognizing the image data of all intersection indicators, multiple image intersection indicator change coefficients are obtained.

[0032] Subsequently, the discrete parameters of the change coefficients of multiple image intersection indicators are calculated to obtain the image change discrete parameters. The image change discrete parameters reflect the degree of dispersion of the variation amplitudes of different intersection indicators in the image data. The calculation method is the same as that for the table change discrete parameters, and standard deviation can be used for measurement. For example, if the change coefficient of the image intersection indicator for "peak load" is 2.8% and the change coefficient of the image intersection indicator for "average load" is 2.4%, then the standard deviation of these two image intersection indicator change coefficients, 0.2%, is calculated as the image change discrete parameter. The smaller the image change discrete parameter, the closer the variation amplitudes of each intersection indicator are in the image, and the better the internal consistency of the image data.

[0033] Next, the ratio of the discrete parameter of image variation to the discrete parameter of maximum variation in the report material is calculated as the third consistency detection coefficient. For example, assuming the maximum discrete parameter of variation in the report material is preset to 0.5%, then the third consistency detection coefficient = 0.2% / 0.5% = 0.4. The third consistency detection coefficient reflects the degree of consistency of the changes in various indicators within the image data. The smaller the third consistency detection coefficient, the more consistent the changing trends of the intersecting indicators in the image, and a more lenient standard can be used in subsequent detection.

[0034] S4. Integrate the consistency detection coefficients, extract related text data and identify changes in intersection indicators to obtain multiple text intersection indicator change coefficients. Verify these coefficients with the multiple table intersection indicator change coefficients and multiple image transaction indicator change coefficients to obtain consistency detection results.

[0035] Specifically, firstly, the consistency detection coefficient is calculated based on the first, second, and third consistency detection coefficients. The calculation method can employ a weighted average or arithmetic average, among other fusion methods. For example, using the arithmetic average method, when the first consistency detection coefficient is 0.6, the second consistency detection coefficient is 0.43, and the third consistency detection coefficient is 0.4, the consistency detection coefficient = (0.6 + 0.43 + 0.4) / 3 ≈ 0.477. The consistency detection coefficient comprehensively reflects the consistency basis and internal consistency degree of different modal data in the report material, and is used to subsequently adjust the scale of textual data extraction.

[0036] Next, the preset text extraction scale is obtained. The preset text extraction scale refers to the default amount of text content extracted near tables and images, usually expressed in characters or words, for example, 1000 words. Based on the consistency detection coefficient, the preset text extraction scale is adjusted and calculated to obtain the final text extraction scale. The specific calculation method is: Text extraction scale = Preset text extraction scale × Consistency detection coefficient. For example, when the preset text extraction scale is 1000 words and the consistency detection coefficient is 0.477, the text extraction scale = 1000 × 0.477 = 477 words. The smaller the consistency detection coefficient, the better the data consistency, and the extracted text content can be appropriately reduced to save computing power; the larger the consistency detection coefficient, the more stringent the detection is required, and the extracted text content should be increased accordingly to improve detection accuracy.

[0037] Next, based on the text extraction scale, adjacent text content is extracted from the report material to obtain related text data. Adjacent text content refers to text paragraphs in the report material that are spatially adjacent to or close to the tables or images, typically including the table or image titles, explanatory text, and preceding and following paragraphs. This related text data usually contains textual descriptions and data explanations of the table and image content and is key text for consistency verification.

[0038] Subsequently, a text indicator change recognizer is obtained to identify the change coefficients of intersection indicators from text data. The text indicator change recognizer is built based on natural language processing algorithms and is obtained through supervised training on a set of sample text data and its labeled set of sample text indicator change coefficients. The associated text data is input into the text indicator change recognizer, which, through semantic analysis and numerical extraction, identifies the descriptions and changes of each intersection indicator in the text, and outputs multiple text intersection indicator change coefficients for multiple intersection indicators. For example, if descriptions such as "peak load increased by 2.6% compared to the previous average" and "average load showed a fluctuation range of 2.3%" are identified from the associated text data, the text intersection indicator change coefficient for "peak load" is output as 2.6%, and the text intersection indicator change coefficient for "average load" is output as 2.3%.

[0039] Next, the similarity of the change coefficients of multiple text intersection indicators, multiple table intersection indicators, and multiple image intersection indicators is calculated to obtain the consistency detection results. Specifically, for each intersection indicator, its change coefficients for text intersection indicators, table intersection indicators, and image intersection indicators are obtained, and the similarity among the three is calculated. The similarity calculation method is as follows: First, the standard deviation of the three change coefficients is calculated; the smaller the standard deviation, the closer the three are. Then, the formula: Similarity = 1 / (1 + Standard Deviation) is used for normalization, so that the similarity value ranges between 0 and 1. For example, the coefficients of variation for "peak load" in text, tables, and images are 2.6%, 2.7%, and 2.8%, respectively, with a calculated average of 2.7% and a standard deviation of approximately 0.082%. Therefore, the similarity is 1 / (1+0.082)≈0.924. The coefficients of variation for "average load" in text, tables, and images are 2.3%, 2.27%, and 2.4%, respectively, with a calculated standard deviation of approximately 0.055%. Therefore, the similarity is 1 / (1+0.055)≈0.948.

[0040] Then, the average similarity of all intersecting indicators is calculated, i.e., (0.924+0.948) / 2=0.936, to obtain the consistency coefficient, which serves as the consistency detection result. The higher the consistency coefficient, i.e., the closer it is to 1, the more consistent the descriptions of the same indicator are among the text, table, and image modalities, and the better the consistency of the report material; conversely, a lower coefficient indicates that there may be data inconsistencies or errors.

[0041] The above technical solution enables automated and intelligent detection of the consistency of multimodal report materials, solving the problems of low efficiency and poor accuracy of traditional manual detection, and achieving the following technical effects:

[0042] First, by extracting indicators from textual, tabular, and image data respectively, and then filtering to obtain intersecting indicators, a bridge of correlation between multimodal data was established, providing a reliable verification basis for consistency detection. The number of intersecting indicators directly reflects the degree of correlation between different modalities, making the detection process more targeted.

[0043] Second, a multi-layered consistency detection coefficient configuration mechanism was established. The first consistency detection coefficient dynamically adjusts the detection strictness based on the proportion of intersection indicators. The second and third consistency detection coefficients reflect the internal consistency of table data and image data, respectively. The three are combined to intelligently adjust the text extraction scale. This adaptive adjustment mechanism ensures detection accuracy while effectively saving computing resources, achieving a balance between detection efficiency and accuracy.

[0044] Third, by calculating the similarity of the variation coefficients of each intersection indicator across the three modalities of text, tables, and images, a quantitative consistency assessment of cross-modal data was achieved. The similarity calculation employed a standard deviation normalization method, ensuring that the detection results have a clear numerical expression, facilitating automated judgment and manual review. A high variation consistency coefficient indicates that the data descriptions in different forms within the report material are consistent; a low variation consistency coefficient allows for the timely detection of data inconsistencies or potential errors, providing strong support for report quality control.

[0045] Fourth, compared with the traditional manual item-by-item checking method, it greatly improves the testing efficiency, shortening the manual testing work that originally required several hours or even days to minutes. At the same time, it avoids the omissions and subjective judgment errors in manual testing, and improves the automation and intelligence level of the quality control of report materials.

[0046] Furthermore, the report materials were obtained, and textual, tabular, and image data were extracted and indicators were extracted to obtain multiple textual indicators, multiple tabular indicators, and multiple image indicators. Among these, the indicators are change data indicators, including:

[0047] S11. Obtain the report materials for the content consistency test;

[0048] S12. Extract content from the report material to obtain text data, table data, and image data;

[0049] S13. Input the text data, table data, and image data into the indicator extractor, and output multiple text indicators, multiple table indicators, and multiple image indicators. The indicator extractor includes a text indicator extraction branch, a table indicator extraction branch, and an image indicator extraction branch. The multiple text indicators, multiple table indicators, and multiple image indicators are variable data indicators.

[0050] In a preferred embodiment, the first step is to acquire the report materials to be tested. These report materials can be operational technical documents such as daily power system operation reports, monthly load analysis reports, and annual power grid operation reports, or project management documents such as power grid technology project contracts, task books, feasibility study reports, and project acceptance data. All of these are power industry documents containing multimodal data. The report materials are usually in electronic document format, such as PDF, Word, Excel, or a combination thereof.

[0051] Next, content extraction is performed on the report materials to obtain text data, tabular data, and image data. Specifically, document parsing technology is used to identify the format and separate the content of the report materials, extracting text paragraphs as text data, structured table content as tabular data, and visualizations such as charts and graphs as image data. The integrity and original format information of all types of data are maintained during the extraction process to ensure the accuracy of subsequent indicator extraction.

[0052] Subsequently, textual, tabular, and image data are input into the indicator extractor, which outputs multiple textual indicators, multiple tabular indicators, and multiple image indicators. The indicator extractor includes textual indicator extraction branches, tabular indicator extraction branches, and image indicator extraction branches. For example, for power system operation reports, the textual indicator extraction branch is used to identify and extract power indicator names from textual data, such as extracting indicators like "peak load" and "average load" from textual descriptions using natural language processing technologies like keyword matching and named entity recognition. The tabular indicator extraction branch is used to identify and extract indicator names from column or row headers in tabular data, determining the power parameters represented by each column or row through table structure analysis. The image indicator extraction branch is used to identify and extract indicator information from charts and graphs in image data, extracting indicator names from chart titles, legends, axis labels, and other locations using image recognition technology. For power grid technology project materials, the textual indicator extraction branch is used to extract project management indicators such as "project investment amount," "construction period," and "technical indicator completion rate" from textual data such as project contracts and feasibility study reports. The tabular indicator extraction branch is used to identify and extract indicator names such as "R&D funding" and "number of achievements" from tabular data such as project budget tables and schedules. The image indicator extraction branch is used to extract key nodes, phase goals, and other indicator information from image data such as project technical roadmap diagrams and organizational charts. The three extraction branches work in parallel, outputting multiple text indicators, multiple tabular indicators, and multiple image indicators respectively, providing a data foundation for subsequent intersection indicator selection.

[0053] Through the above steps, structured analysis of report materials and automatic extraction of multimodal indicators were achieved. The three-branch parallel processing architecture of the indicator extractor can simultaneously process three different data formats: text, tables, and images, improving the efficiency of indicator extraction. Each extraction branch employs corresponding technical methods tailored to the characteristics of different data formats, ensuring the accuracy and completeness of indicator identification and providing a data foundation for subsequent consistency testing.

[0054] Furthermore, the indicator extractor is constructed using the following steps:

[0055] S131. Based on the data processing of historical report materials, collect sample text data sets, sample table data sets, and sample image data sets, extract indicators from each sample text data, sample table data, and sample image data set to obtain multiple sample text indicator sets, multiple sample table indicator sets, and multiple sample image indicator sets.

[0056] S132. Based on machine learning, construct text indicator extraction branches, table indicator extraction branches, and image indicator extraction branches;

[0057] S133. Using the sample text data set and multiple sample text indicator sets, sample table data set and multiple sample table indicator sets, sample image data set and multiple sample image indicator sets, supervised training and testing are performed on the text indicator extraction branch, table indicator extraction branch and image indicator extraction branch respectively until convergence and test pass;

[0058] S134. Based on the text indicator extraction branch, table indicator extraction branch, and image indicator extraction branch that have passed the test, obtain the indicator extractor.

[0059] In a preferred embodiment, to construct the indicator extractor, firstly, based on the processing data of historical report materials, sample text data sets, sample table data sets, and sample image data sets are collected. Indicators are extracted from each sample text data, sample table data, and sample image data set, resulting in multiple sample text indicator sets, multiple sample table indicator sets, and multiple sample image indicator sets. Specifically, representative report samples are selected from historically accumulated power system report materials to obtain historical report material processing data, and these samples are then data extracted and labeled. For example, text paragraphs are extracted from sample reports as sample text data, and the indicator names contained therein, such as "peak load," "average load," and "load factor," are manually labeled to form a sample text indicator set. Similarly, tables from sample reports are extracted as sample table data, and the indicator names in the tables are labeled to form a sample table indicator set. Charts and images from sample reports are extracted as sample image data, and the indicator names displayed in the charts and images are labeled to form a sample image indicator set.

[0060] Then, based on machine learning, three branches are constructed for text indicator extraction, table indicator extraction, and image indicator extraction. Specifically, the text indicator extraction branch can be built using a named entity recognition network based on pre-trained language models such as BERT and GPT to identify indicator entities in text. The table indicator extraction branch can be built using a deep learning model based on table structure understanding, capable of recognizing the row and column structure of tables and extracting indicator information from the titles. The image indicator extraction branch can be built using a model based on convolutional neural networks combined with optical character recognition technology to identify and extract indicator text information from chart images. The three branches employ different network architectures to adapt to the characteristics of their respective data formats.

[0061] Subsequently, using sample text datasets and multiple sample text indicator datasets, sample table datasets and multiple sample table indicator datasets, and sample image datasets and multiple sample image indicator datasets, supervised training and testing were performed on the text indicator extraction branch, table indicator extraction branch, and image indicator extraction branch until convergence and passing the test. Specifically, sample text data was used as input, and the corresponding sample text indicators were used as labels. Supervised training was performed on the text indicator extraction branch, and the network parameters were continuously optimized through the backpropagation algorithm to make the model output indicators as consistent as possible with the labeled indicators. During training, the change of the loss function was monitored, and the model was considered to have converged when the loss function tended to stabilize. After convergence, performance was evaluated using an independent test set, and indicators such as accuracy and recall of indicator extraction were calculated. The test was considered to be qualified when the test indicators reached a preset threshold, for example, the preset threshold was accuracy ≥ 95%. The table indicator extraction branch and the image indicator extraction branch adopted the same training and testing process, and were trained using their respective sample datasets until they passed the test.

[0062] Next, the independently trained and tested text, table, and image indicator extraction branches are integrated to form a complete indicator extractor. The indicator extractor receives text, table, and image data as input. The text, table, and image indicator extraction branches process their respective data types in parallel, outputting text, table, and image indicators respectively. The integrated indicator extractor possesses comprehensive processing capabilities for multimodal data and can automatically extract all types of indicator information from report materials.

[0063] Through the above construction steps, the indicator extractor is trained on a large number of historical report samples, learning various patterns and features of indicator expression in power system reports, demonstrating good generalization ability and extraction accuracy. Supervised learning ensures the consistency between the model output and actual needs, and rigorous testing guarantees the reliability of indicator extraction, providing high-quality foundational data for subsequent consistency checks.

[0064] Furthermore, multiple intersection metrics are obtained through screening, and a first consistency detection coefficient is configured, including:

[0065] S21. Filter the intersection of multiple text indicators, multiple table indicators and multiple image indicators to obtain multiple intersection indicators;

[0066] S22. Calculate the ratio of the number of multiple intersection indicators to the number of multiple text indicators, multiple table indicators and multiple image indicators, and take the minimum value as the indicator consistency coefficient to calculate the first consistency detection coefficient.

[0067] In a preferred embodiment, specifically, a set operation is performed on the obtained multiple textual indicators, multiple tabular indicators, and multiple image indicators to find the common indicators that exist simultaneously in all three indicator sets. For example, the multiple textual indicators include "peak load," "average load," "load factor," "maximum daily load," and "minimum nighttime load"; the multiple tabular indicators include "peak load," "average load," "minimum load," and "voltage compliance rate"; and the multiple image indicators include "peak load," "load factor," and "average load." Based on the intersection of the multiple textual indicators, multiple tabular indicators, and multiple image indicators, two intersection indicators are obtained, namely "peak load" and "average load." The intersection indicators refer to the power parameters that are reflected in the textual descriptions, tabular records, and charts of the report materials. These indicators meet the conditions for cross-modal verification and are the core objects for content consistency detection.

[0068] Subsequently, the ratios of the number of intersection indicators to the number of text indicators, the number of intersection indicators to the number of table indicators, and the number of intersection indicators to the number of image indicators are calculated respectively. The minimum value among these three ratios is selected as the indicator consistency coefficient. For example, if there are 2 intersection indicators, 5 text indicators, 4 table indicators, and 3 image indicators, the three ratios are 2 / 5 = 0.4, 2 / 4 = 0.5, and 2 / 3 ≈ 0.67 respectively. The minimum value of 0.4 is taken as the indicator consistency coefficient. The indicator consistency coefficient reflects the minimum coverage of intersection indicators in each modality of data. The smaller the indicator consistency coefficient, the smaller the proportion of intersection indicators in a certain modality, the weaker the data correlation, and the more stringent the consistency verification process needs to be. Then, the indicator consistency coefficient is subtracted from the value 1, i.e., the first consistency detection coefficient = 1 - indicator consistency coefficient, to obtain the first consistency detection coefficient. For example, when the indicator consistency coefficient is 0.4, the first consistency detection coefficient = 1 - 0.4 = 0.6. The first consistency detection coefficient is used to adjust the stringency of subsequent testing. The smaller the first consistency detection coefficient, the more lenient the consistency detection standard can be adopted; the larger the first consistency detection coefficient, the more stringent the testing standard needs to be adopted.

[0069] Through the above steps, an adaptive detection mechanism based on the number of intersection indicators was established. When the overlap of indicators in different modalities in the report material is high, the detection coefficient is automatically reduced to improve detection efficiency; when the overlap is low, the detection coefficient is automatically increased to ensure detection accuracy. This dynamic adjustment mechanism allows consistency detection to flexibly adapt to the actual characteristics of the report material, avoiding both resource waste caused by over-detection and missed detection caused by under-detection.

[0070] Furthermore, the table data is extracted to obtain multiple intersection indicator tables. Change analysis is performed to obtain the change coefficients of the intersection indicators from multiple tables. A second consistency detection coefficient is then configured, including:

[0071] S23. Extract data from the table data according to multiple intersection indicators to obtain a table dataset of multiple intersection indicators.

[0072] S24. Based on the dataset of multiple intersection indicator tables, calculate the change of each type of intersection indicator table data to obtain multiple change rates, which are used as the change coefficients of the intersection indicators of multiple tables.

[0073] S25. Configure a second consistency detection coefficient based on the change coefficients of the intersection indicators of multiple tables.

[0074] In a preferred embodiment, firstly, multiple intersection indicators are extracted from the tabular data to obtain multiple intersection indicator tabular datasets. Specifically, based on the selected multiple intersection indicators, the data records corresponding to each intersection indicator are located and extracted from the tabular data. For example, if the intersection indicators include "peak load" and "average load," then the column or row labeled "peak load" is found in the table, and all numerical data in that column or row are extracted to form the intersection indicator tabular dataset for "peak load"; similarly, all numerical data corresponding to "average load" are extracted to form the intersection indicator tabular dataset for "average load." Each intersection indicator tabular dataset contains multiple data points arranged according to time series or other dimensions, and these data points reflect the values ​​of the indicator under different conditions.

[0075] Subsequently, based on multiple intersection indicator table datasets, the changes in each type of intersection indicator table dataset are calculated to obtain multiple rates of change, which serve as the change coefficients for the intersection indicators of multiple tables. Specifically, time series change analysis is performed on each intersection indicator table dataset to calculate the rate of change between adjacent data points. Taking "peak load" as an example, if its intersection indicator table dataset contains data for five consecutive days: Day 1 5000MW, Day 2 5200MW, Day 3 5100MW, Day 4 5300MW, and Day 5 5250MW, then the absolute values ​​of the rate of change for adjacent data points are calculated: from Day 1 to Day 2, it is |(5200-5000) / 5000|=4%; from Day 2 to Day 3, it is |(5100-5200) / 5200|=1.92%; from Day 3 to Day 4, it is |(5300-5100) / 5100|=3.92%; and from Day 4 to Day 5, it is |(5250-5300) / 5300|=0.94%. Then, calculate the average of all change rates: (4% + 1.92% + 3.92% + 0.94%) / 4 = 2.7%. Use this average as the change coefficient of the tabular intersection index for "peak load". Repeat the above process for all intersection indicators to obtain the tabular intersection index change coefficient for each intersection indicator, forming multiple tabular intersection index change coefficients. These tabular intersection index change coefficients quantify the fluctuation range of each intersection indicator in the tabular data.

[0076] Next, a second consistency detection coefficient is configured based on the change coefficients of the intersection indicators of multiple tables. Specifically, the dispersion parameter of the change coefficients of the intersection indicators of multiple tables is first calculated to obtain the table change dispersion parameter. The dispersion parameter is calculated using the standard deviation to measure the degree of dispersion of the change amplitudes of different intersection indicators. The smaller the table change dispersion parameter, the closer the change amplitudes of each intersection indicator in the table, and the better the internal consistency of the table data. Subsequently, the ratio of the table change dispersion parameter to the maximum change dispersion parameter of the report material is calculated as the second consistency detection coefficient. The second consistency detection coefficient reflects the degree of consistency of the changes of each indicator within the table data. The smaller the second consistency detection coefficient, the more consistent the change trend of each intersection indicator in the table, and a more lenient standard can be used in subsequent detection.

[0077] Through the above steps, a quantitative assessment of the internal consistency of the tabular data is achieved. By analyzing the magnitude and dispersion of the changes in each intersection indicator within the table, the reliability and consistency level of the tabular data are determined, and a second consistency detection coefficient is dynamically configured accordingly. When the changing trends of each indicator in the tabular data are consistent, it indicates that the table records are relatively standardized and reliable, and the intensity of subsequent detection can be appropriately reduced; when the changes in each indicator are significantly different, it indicates that the table data may contain anomalies, and the detection rigor needs to be increased, thereby improving the targeting and effectiveness of consistency detection.

[0078] Furthermore, based on the change coefficients of the intersection indicators of multiple tables, a second consistency detection coefficient is configured, including:

[0079] S251. Calculate the discrete parameters of the change coefficients of the intersection indicators of multiple tables to obtain the discrete parameters of the table changes.

[0080] S252. Calculate the ratio of the variation discrete parameter of the table to the maximum variation discrete parameter of the report material, and use it as the second consistency detection coefficient.

[0081] In one feasible implementation, firstly, the discrete parameters of the change coefficients of multiple table intersection indicators are calculated to obtain the table change discrete parameters. Specifically, the obtained change coefficients of multiple table intersection indicators are subjected to dispersion analysis, using standard deviation as the calculation method for the discrete parameter. The standard deviation reflects the degree of dispersion among multiple change coefficients; the smaller the value, the closer the change coefficients are, and the larger the value, the greater the difference between the change coefficients. For example, the change coefficient of the table intersection indicator for "peak load" is 2.7%, and the change coefficient of the table intersection indicator for "average load" is 2.27%. The standard deviation of the two is calculated to be 0.215%, which is the table change discrete parameter. The table change discrete parameter quantitatively describes the consistency of the change amplitude of different intersection indicators in the table data. The smaller the parameter, the more similar the change patterns of the intersection indicators, and the higher the internal correlation and reliability of the table data.

[0082] Subsequently, the ratio of the table variation dispersion parameter to the maximum variation dispersion parameter of the report material is calculated as the second consistency detection coefficient. Specifically, the calculated table variation dispersion parameter is compared with the preset maximum variation dispersion parameter to achieve normalization. The maximum variation dispersion parameter is an empirical threshold obtained from statistical analysis of historical report materials, or a maximum dispersion tolerance preset based on the characteristics of power system indicators. This maximum variation dispersion parameter represents the acceptable maximum degree of dispersion. For example, if the table variation dispersion parameter is 0.215% and the maximum variation dispersion parameter is preset to 0.5%, then the second consistency detection coefficient = 0.215% / 0.5% = 0.43. The value of the second consistency detection coefficient ranges from 0 to 1. The smaller the second consistency detection coefficient, the more consistent the changing trends of the intersecting indicators in the table data, and the better the consistency within the table. A relatively lenient standard can be used when conducting cross-modal consistency testing subsequently. The larger the second consistency detection coefficient, the greater the differences in the changes of the indicators within the table data, which may indicate data anomalies or non-standard recording. A stricter testing standard is needed to ensure the accuracy of the subsequent testing.

[0083] Through the above steps, the configuration of the second consistency detection coefficient was achieved. This second consistency detection coefficient not only reflects the internal quality of the tabular data, but also provides an important adjustment parameter for subsequent cross-modal consistency detection, enabling the detection process to adaptively adjust according to the actual characteristics of the data, thereby improving the intelligence and accuracy of the detection.

[0084] Furthermore, the image data undergoes change identification based on multiple intersection metrics to obtain change coefficients for these metrics. A third consistency detection coefficient is then configured, including:

[0085] S31. Obtain an image indicator change recognizer, wherein the image indicator change recognizer is obtained by training a sample indicator set, a sample image data set, and a sample indicator change coefficient set.

[0086] S32. The image data is combined with multiple intersection indicators and input into the image indicator change recognizer, and the multiple image intersection indicator change coefficients are output.

[0087] S33. Calculate the discrete parameters of the change coefficients of the intersection indicators of multiple images to obtain the discrete parameters of image changes;

[0088] S34. Calculate the ratio of the discrete parameter of the image change to the maximum discrete parameter of the report material, and use it as the third consistency detection coefficient.

[0089] In a preferred embodiment, firstly, an image indicator change recognizer is obtained. This image indicator change recognizer is a deep learning model built on a convolutional neural network, used to automatically identify the numerical change magnitude of specific indicators from chart images. During training, a large number of chart images from historical power system reports are first collected as a sample image dataset. These images include various forms such as line charts, bar charts, and pie charts. Simultaneously, the indicator names corresponding to these sample images are collected, forming a sample indicator set. Furthermore, through manual analysis or automatic calculation, the actual change magnitude of the corresponding indicator in each sample image is extracted, such as the average rate of change of adjacent data points in a line chart, forming a sample indicator change coefficient set as training labels. During training, the sample image data and corresponding sample indicators are used as input, and the sample indicator change coefficients are used as supervision signals to supervise the learning of the convolutional neural network. By learning the mapping relationship between image features and indicator change coefficients, the system can automatically output the indicator change coefficient given a chart image and indicator name. After training, the image indicator change recognizer is obtained. This recognizer can receive image data and indicator names as input and output the corresponding indicator change coefficients without manual reading and calculation.

[0090] Then, the image data is combined with multiple intersection indicators and input into the image indicator change recognizer, which outputs multiple image intersection indicator change coefficients. Specifically, the extracted image data and the obtained multiple intersection indicators are input into the image indicator change recognizer. For example, for the intersection indicator "peak load," the chart image containing this indicator, along with the indicator name, is input into the image indicator change recognizer. The fluctuation characteristics of the curve or bar corresponding to "peak load" in the chart are extracted, the change amplitude between data points is calculated, and the image intersection indicator change coefficient for this indicator is output, for example, 2.8%. Similarly, the intersection indicator "average load" is identified, and its image intersection indicator change coefficient is output, for example, 2.4%. By identifying and processing all intersection indicators separately, multiple image intersection indicator change coefficients are obtained. These image intersection indicator change coefficients quantify the fluctuation of each intersection indicator in the image data, providing a quantitative basis for subsequent consistency comparison.

[0091] Next, the discrete parameters of the change coefficients of multiple image intersection indicators are calculated to obtain the image change discrete parameters. Specifically, the dispersion analysis of the obtained change coefficients of multiple image intersection indicators is performed, and the calculation method is the same as that of the table change discrete parameters, using standard deviation as the measure. For example, the change coefficient of the image intersection indicator for "peak load" is 2.8%, and the change coefficient of the image intersection indicator for "average load" is 2.4%, with a calculated standard deviation of 0.2%, which is used as the image change discrete parameter. The image change discrete parameter reflects the consistency of the change amplitude of different intersection indicators in the image data. The smaller the image change discrete parameter, the more similar the change trends of each intersection indicator are shown in the chart, and the better the internal consistency of the image data.

[0092] Subsequently, the ratio of the discrete parameter of image variation to the maximum discrete parameter of variation in the report material is calculated as the third consistency detection coefficient. Specifically, the calculated discrete parameter of image variation is compared with the preset maximum discrete parameter of variation. For example, if the discrete parameter of image variation is 0.2% and the maximum discrete parameter of variation is preset to 0.5%, then the third consistency detection coefficient = 0.2% / 0.5% = 0.4. The value of the third consistency detection coefficient ranges from 0 to 1. The smaller the third consistency detection coefficient, the more consistent the changing trends of the intersection indicators in the image data, and the higher the data quality within the image, allowing for a more lenient standard to be used in subsequent detection. Conversely, the larger the third consistency detection coefficient, the greater the differences in the changes of the indicators within the image data, requiring a more stringent detection standard to be adopted.

[0093] Through the above steps, an intelligent assessment of the internal consistency of image data is achieved. Deep learning technology is used to automatically extract indicator change information from chart images, avoiding the errors and efficiency problems of manual chart reading. By calculating the discrete parameters of image changes and configuring the third consistency detection coefficient, an adjustment basis is provided for subsequent cross-modal consistency verification, enabling the detection process to comprehensively consider the characteristics of image data and improving the comprehensiveness and accuracy of the detection.

[0094] Furthermore, by integrating the consistency detection coefficients, extracting associated text data and identifying changes in intersection indicators, multiple text intersection indicator change coefficients are obtained. These coefficients are then verified against the multiple table intersection indicator change coefficients and multiple image transaction indicator change coefficients to obtain consistency detection results, including:

[0095] S41. Calculate the consistency detection coefficient based on the first consistency detection coefficient, the second consistency detection coefficient, and the third consistency detection coefficient;

[0096] S42. Obtain the preset text extraction scale;

[0097] S43. Based on the consistency detection coefficient, the preset text extraction scale is adjusted and calculated to obtain the text extraction scale;

[0098] S44. According to the text extraction scale, extract the text content adjacent to the table data and image data to obtain related text data;

[0099] S45. Input the associated text data into the text indicator change recognizer and output the multiple text intersection indicator change coefficients of multiple intersection indicators.

[0100] S46. Calculate the average similarity of the change coefficients of multiple intersection index table data, multiple image intersection indexes, and multiple text intersection indexes to obtain the change consistency coefficient, which is used as the consistency detection result.

[0101] In a preferred embodiment, firstly, a consistency detection coefficient is calculated based on a first consistency detection coefficient, a second consistency detection coefficient, and a third consistency detection coefficient. The configured first, second, and third consistency detection coefficients are then fused to obtain a comprehensive consistency detection coefficient. This fusion calculation can employ an arithmetic mean or a weighted average statistical method. For example, using an arithmetic mean, when the first consistency detection coefficient is 0.6, the second consistency detection coefficient is 0.43, and the third consistency detection coefficient is 0.4, the consistency detection coefficient = (0.6 + 0.43 + 0.4) / 3 ≈ 0.477. The consistency detection coefficient comprehensively reflects the coverage of intersection indicators in the report material, the internal consistency of tabular data, and the internal consistency of image data. It is a comprehensive assessment of the overall data quality and consistency foundation of the report material, used to guide adjustments to the scale of subsequent text data extraction.

[0102] Next, the preset text extraction scale is obtained. This preset text extraction scale represents the amount of text content extracted by default near tables and images, usually in units of characters or words. For example, the preset text extraction scale can be set to 1000 words. This preset text extraction scale serves as the baseline for text extraction and is subsequently dynamically adjusted based on the consistency detection coefficient to adapt to the detection needs of different report materials. Then, the preset text extraction scale is adjusted and calculated based on the consistency detection coefficient to obtain the final text extraction scale. Specifically, the preset text extraction scale is proportionally adjusted using the consistency detection coefficient, and the calculation formula is: Text Extraction Scale = Preset Text Extraction Scale × Consistency Detection Coefficient. For example, when the preset text extraction scale is 1000 words and the consistency detection coefficient is approximately 0.477, the text extraction scale = 1000 × 0.477 = 477 words. The principle behind this adjustment mechanism is as follows: when the consistency detection coefficient is small, it indicates that the intersection index coverage of the report materials is high, the internal consistency of the table and image data is good, and the basic data quality is relatively good. In this case, the amount of text extraction can be appropriately reduced to save computing resources. When the consistency detection coefficient is large, it indicates that more rigorous detection is required, and the amount of text extraction should be increased to improve the comprehensiveness and accuracy of the detection. Through adaptive adjustment, a dynamic balance between detection efficiency and detection accuracy is achieved.

[0103] Next, based on the text extraction scale, the text content adjacent to the table and image data is extracted to obtain related text data. Specifically, the positions of tables and images are located in the report material, and surrounding text content is extracted according to the determined text extraction scale, using these positions as the center. Adjacent text content refers to text paragraphs in the report document that are spatially adjacent to or close to the tables or images, specifically including the table or image titles, captions, explanatory text, preceding paragraphs, following paragraphs, and other related textual descriptions. For example, if a table has 200 words of explanatory text before and after it, when the text extraction scale is 477 words, the text content before the table (239 words) and after the table (238 words) can be extracted. This related text data typically contains textual explanations, data analysis, and conclusions of the table and image content, serving as a bridge for consistency verification among the three modalities of text, tables, and images.

[0104] Subsequently, the associated text data is input into the text indicator change recognizer, which outputs multiple text intersection indicator change coefficients for multiple intersection indicators. Specifically, the text indicator change recognizer is a model built based on natural language processing algorithms, obtained through supervised training on a set of sample text data and its labeled set of sample text indicator change coefficients. The acquired associated text data is input into the text indicator change recognizer, which automatically identifies the descriptions and changes of each intersection indicator in the text through semantic analysis, numerical extraction, and change pattern recognition. For example, from the associated text data, descriptions such as "peak load fluctuated by 2.6% compared to the previous average" and "average load showed a change of 2.3%" are identified, and the text intersection indicator change coefficient for "peak load" is extracted and output as 2.6%, and the text intersection indicator change coefficient for "average load" is 2.3%. All intersection indicators are identified and processed to obtain multiple text intersection indicator change coefficients, quantifying the change magnitude of the intersection indicators in the text description.

[0105] Next, the mean similarity of the variation coefficients of multiple intersection indicators (table data, images, and text) is calculated to obtain the consistency coefficient, which serves as the consistency detection result. Specifically, for each intersection indicator, its variation coefficients in the text, table, and image modalities are obtained, and the similarity among the three is calculated. The similarity calculation method is as follows: first, the standard deviation of the three variation coefficients is calculated; the smaller the standard deviation, the closer the three are. Then, the formula: Similarity = 1 / (1 + Standard Deviation) is used to ensure that the similarity value ranges between 0 and 1. For example, for the intersection indicator "Peak Load," its variation coefficients in text, table, and image are 2.6%, 2.7%, and 2.8%, respectively, and the similarity of this indicator is calculated to be 0.924. Similarly, for the intersection indicator "Average Load," its variation coefficients in text, table, and image are 2.3%, 2.27%, and 2.4%, respectively, and the similarity of this indicator is calculated to be 0.948. Then, the arithmetic mean of the similarity of all intersecting indicators is calculated, i.e., (0.924 + 0.948) / 2 = 0.936, to obtain the variation consistency coefficient. The variation consistency coefficient serves as the final consistency detection result. A higher coefficient indicates greater consistency in the description of the same indicator across the text, table, and image modalities, resulting in better content consistency and higher data quality in the report. Conversely, a lower coefficient suggests potential inconsistencies or errors between different modalities, requiring further manual verification and correction.

[0106] Through the above steps, intelligent detection of consistency in multimodal report materials is achieved. By dynamically adjusting the detection strategy through the integration of multi-level consistency detection coefficients, the system automatically extracts information on changes in indicators from the text, achieves quantitative consistency assessment through cross-modal similarity calculation, and outputs the change consistency coefficient as the consistency detection result. This fully utilizes the complementarity and correlation of text, tables, and images in the report materials, not only improving the accuracy and comprehensiveness of the detection but also optimizing the detection efficiency through an adaptive mechanism, thus providing support for the quality control of power system reports.

[0107] Furthermore, the training steps of the text indicator change recognizer include:

[0108] S451. Obtain the sample text data set, extract the data change coefficients of all indicators in it, and label the sample text indicator change coefficient set.

[0109] S452. Construct a text indicator change recognizer based on natural language processing algorithms;

[0110] S453. Input the associated text data into the text indicator change recognizer, recognize and output multiple text indicator change coefficients, and extract multiple text intersection indicator change coefficients of multiple intersection indicators.

[0111] In a preferred embodiment, firstly, a sample text data set is acquired, and the data change coefficients of all indicators within it are extracted, thus obtaining a sample text indicator change coefficient set. Specifically, a large number of text paragraphs containing indicator descriptions are collected from historical power system reports to form a sample text data set. This sample text data contains natural language descriptions of various power indicators and their changes, such as statements like "peak load increased by 8% compared to the previous month" and "average load fluctuation was 3.5%". For each sample text data, the indicator name and corresponding data change coefficient are extracted through manual annotation or automatic parsing. For example, from "peak load increased by 8% compared to the previous month", the indicator is extracted as "peak load", and the data change coefficient is 8%. All sample text data are annotated to obtain a sample text indicator change coefficient set, recording the correspondence between text descriptions and indicator change coefficients, providing a supervisory signal for training the text indicator change recognizer.

[0112] Next, a text indicator change recognizer is constructed based on natural language processing algorithms. Specifically, a deep learning model from the field of natural language processing is used to build the text indicator change recognizer. The model architecture can employ sequence labeling networks or text classification networks based on pre-trained language models such as BERT and GPT, capable of understanding the semantic information of text and extracting numerical change features. The model includes a text encoding layer, a semantic understanding layer, and a change coefficient prediction layer. The text encoding layer converts the input text data into vector representations, the semantic understanding layer analyzes the descriptive patterns of indicator changes in the text through an attention mechanism, and the change coefficient prediction layer outputs the corresponding change coefficient values ​​based on semantic features. The model is trained under supervised supervision using the obtained sample text data set and sample text indicator change coefficient set, and the model parameters are optimized through backpropagation to enable the model to accurately identify and extract the change coefficients of indicators from text descriptions. After training, the text indicator change recognizer is obtained.

[0113] Next, the associated text data is input into the text indicator change recognizer, which identifies and outputs multiple text indicator change coefficients, and extracts the change coefficients of multiple text intersection indicators. Specifically, the acquired associated text data is input into the trained text indicator change recognizer, which performs semantic analysis and numerical extraction on the text content, automatically identifying all indicators involved in the text and their corresponding change coefficients, and outputting multiple text indicator change coefficients. Then, based on the obtained multiple intersection indicators, the change coefficients corresponding to the intersection indicators are selected and extracted from the identified multiple text indicator change coefficients. For example, if the intersection indicators are "peak load" and "average load", the text intersection indicator change coefficients corresponding to these two indicators, 2.6% and 2.3%, are extracted for subsequent cross-modal consistency verification.

[0114] Through the training steps described above, the text indicator change recognizer learns the mapping relationship between text descriptions and indicator change coefficients, enabling it to automatically understand change information in various expression formats, including percentage expressions, multiple expressions, and absolute value changes. This text indicator change recognizer avoids the tedious work of manual reading and extraction, improves the automation and accuracy of text data processing, and provides reliable text indicator change data for multimodal consistency detection.

[0115] Example 2, as Figure 2 As shown, based on the same inventive concept as the intelligent content consistency detection method for multimodal reporting materials provided in Embodiment 1, this embodiment of the invention also provides an intelligent content consistency detection system for multimodal reporting materials, including:

[0116] The multimodal indicator extraction module 11 is used to obtain report materials, extract text data, table data and image data and extract indicators to obtain multiple text indicators, multiple table indicators and multiple image indicators, among which the indicators are change data indicators;

[0117] The table consistency analysis module 12 is used to filter and obtain multiple intersection indicators, configure the first consistency detection coefficient, extract table data to obtain multiple intersection indicator table data, perform change analysis to obtain the change coefficient of multiple table intersection indicators, and configure the second consistency detection coefficient.

[0118] Image consistency analysis module 13 is used to identify changes in multiple intersection indicators of image data, obtain the change coefficients of multiple image intersection indicators, and configure the third consistency detection coefficient.

[0119] The fusion verification and detection module 14 is used to fuse the consistency detection coefficients, extract related text data and identify the changes in intersection indicators, obtain multiple text intersection indicator change coefficients, verify them with the multiple table intersection indicator change coefficients and multiple image transaction indicator change coefficients, and obtain consistency detection results.

[0120] Furthermore, the execution steps of the multimodal indicator extraction module 11 include:

[0121] Obtain the report materials for which content consistency needs to be verified;

[0122] Content extraction was performed on the report materials to obtain text data, table data, and image data;

[0123] The text data, table data, and image data are input into the indicator extractor, which outputs multiple text indicators, multiple table indicators, and multiple image indicators. The indicator extractor includes a text indicator extraction branch, a table indicator extraction branch, and an image indicator extraction branch. The multiple text indicators, multiple table indicators, and multiple image indicators are variable data indicators.

[0124] Furthermore, the indicator extractor is constructed using the following steps:

[0125] Based on the data processing of historical report materials, sample text data sets, sample table data sets, and sample image data sets are collected. Indicators are extracted from each sample text data, sample table data, and sample image data to obtain multiple sample text indicator sets, multiple sample table indicator sets, and multiple sample image indicator sets.

[0126] Based on machine learning, we construct branches for extracting text metrics, table metrics, and image metrics.

[0127] Using the sample text data set and multiple sample text indicator sets, sample table data set and multiple sample table indicator sets, sample image data set and multiple sample image indicator sets, supervised training and testing are performed on the text indicator extraction branch, table indicator extraction branch and image indicator extraction branch respectively until convergence and test pass;

[0128] Based on the tested and approved text indicator extraction branches, table indicator extraction branches, and image indicator extraction branches, an indicator extractor is obtained.

[0129] Furthermore, the execution steps of the table consistency analysis module 12 include:

[0130] Filter the intersection of multiple text metrics, multiple table metrics, and multiple image metrics to obtain multiple intersection metrics;

[0131] Calculate the ratio of the number of multiple intersection indicators to the number of multiple text indicators, multiple table indicators, and multiple image indicators, and take the minimum value as the indicator consistency coefficient to obtain the first consistency detection coefficient.

[0132] Furthermore, the execution steps of the table consistency analysis module 12 include:

[0133] Based on multiple intersecting indicators, extract data from the tabular data to obtain a tabular dataset of multiple intersecting indicators;

[0134] Based on multiple intersection indicator tables, the changes in the data of each type of intersection indicator table are calculated to obtain multiple change rates, which are used as the change coefficients of the intersection indicators of multiple tables.

[0135] Configure a second consistency detection coefficient based on the change coefficients of the intersection indicators of multiple tables.

[0136] Furthermore, the execution steps of the table consistency analysis module 12 also include:

[0137] Calculate the discrete parameters of the change coefficients of the intersection indicators of multiple tables to obtain the discrete parameters of table changes;

[0138] The ratio of the variation discrete parameter of the table to the maximum variation discrete parameter of the report material is calculated and used as the second consistency detection coefficient.

[0139] Furthermore, the execution steps of the image consistency analysis module 13 include:

[0140] An image indicator change recognizer is obtained, wherein the image indicator change recognizer is trained using a sample indicator set, a sample image data set, and a sample indicator change coefficient set;

[0141] The image data is combined with multiple intersection indicators and input into the image indicator change recognizer, which outputs multiple image intersection indicator change coefficients.

[0142] Calculate the discrete parameters of the change coefficients of the intersection indicators of multiple images to obtain the discrete parameters of image change;

[0143] The ratio of the discrete parameter of the image change to the maximum discrete parameter of the report material is calculated and used as the third consistency detection coefficient.

[0144] Furthermore, the execution steps of the fusion verification and detection module 14 include:

[0145] Calculate the consistency detection coefficient based on the first consistency detection coefficient, the second consistency detection coefficient, and the third consistency detection coefficient;

[0146] Get the preset text extraction scale;

[0147] Based on the consistency detection coefficient, the preset text extraction scale is adjusted and calculated to obtain the text extraction scale;

[0148] Based on the stated text extraction scale, extract the text content adjacent to the table data and image data to obtain related text data;

[0149] The associated text data is input into the text indicator change recognizer, which outputs multiple text intersection indicator change coefficients for multiple intersection indicators.

[0150] The mean similarity of the change coefficients of multiple intersection index table data, multiple image intersection index change coefficients, and multiple text intersection index change coefficients is calculated to obtain the change consistency coefficient, which is used as the consistency detection result.

[0151] Furthermore, the training steps of the text indicator change recognizer include:

[0152] Obtain a sample text data set, extract the data change coefficients of all indicators in it, and label them to obtain a set of sample text indicator change coefficients;

[0153] A text indicator change recognizer was constructed based on natural language processing algorithms.

[0154] The associated text data is input into the text indicator change recognizer, which identifies and outputs multiple text indicator change coefficients, and extracts multiple text intersection indicator change coefficients of multiple intersection indicators.

[0155] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0156] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0157] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0158] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0159] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0160] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0161] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for intelligent content consistency detection of multimodal report materials, characterized in that: The method includes: Obtain report materials, extract text data, tabular data, and image data, and extract indicators to obtain multiple text indicators, multiple tabular indicators, and multiple image indicators, among which the indicators are change data indicators; Filter the intersection of multiple text metrics, multiple table metrics, and multiple image metrics to obtain multiple intersection metrics, configure the first consistency detection coefficient, extract table data to obtain multiple intersection metric table data, perform change analysis to obtain the change coefficient of multiple table intersection metrics, and configure the second consistency detection coefficient. The image data is subjected to multiple intersection indicators for change identification, the change coefficients of multiple image intersection indicators are obtained, and a third consistency detection coefficient is configured. By integrating consistency detection coefficients, extracting associated text data and identifying changes in intersection indicators, multiple text intersection indicator change coefficients are obtained. These coefficients are then verified against the multiple table intersection indicator change coefficients and multiple image intersection indicator change coefficients to obtain consistency detection results, including: Calculate the consistency detection coefficient based on the first consistency detection coefficient, the second consistency detection coefficient, and the third consistency detection coefficient; Get the preset text extraction scale; Based on the consistency detection coefficient, the preset text extraction scale is adjusted and calculated to obtain the text extraction scale; Based on the stated text extraction scale, extract the text content adjacent to the table data and image data to obtain related text data; The associated text data is input into the text indicator change recognizer, which outputs multiple text intersection indicator change coefficients for multiple intersection indicators. The mean similarity of the change coefficients of the intersection indicators of multiple tables, multiple images, and multiple texts is calculated to obtain the change consistency coefficient, which is used as the consistency detection result.

2. The intelligent content consistency detection method for multimodal reporting materials according to claim 1, characterized in that, Obtain report materials, extract textual, tabular, and image data, and extract metrics to obtain multiple textual metrics, multiple tabular metrics, and multiple image metrics, including: Obtain the report materials for which content consistency needs to be verified; Content extraction was performed on the report materials to obtain text data, table data, and image data; The text data, table data, and image data are input into the indicator extractor, which outputs multiple text indicators, multiple table indicators, and multiple image indicators. The indicator extractor includes a text indicator extraction branch, a table indicator extraction branch, and an image indicator extraction branch. The multiple text indicators, multiple table indicators, and multiple image indicators are variable data indicators.

3. The intelligent content consistency detection method for multimodal reporting materials according to claim 2, characterized in that, The indicator extractor is constructed using the following steps: Based on the data processing of historical report materials, sample text data sets, sample table data sets, and sample image data sets are collected. Indicators are extracted from each sample text data, sample table data, and sample image data to obtain multiple sample text indicator sets, multiple sample table indicator sets, and multiple sample image indicator sets. Based on machine learning, we construct branches for extracting text metrics, table metrics, and image metrics. Using the sample text data set and multiple sample text indicator sets, sample table data set and multiple sample table indicator sets, sample image data set and multiple sample image indicator sets, supervised training and testing are performed on the text indicator extraction branch, table indicator extraction branch and image indicator extraction branch respectively until convergence and test pass; Based on the tested and approved text indicator extraction branches, table indicator extraction branches, and image indicator extraction branches, an indicator extractor is obtained.

4. The intelligent content consistency detection method for multimodal reporting materials according to claim 1, characterized in that, Multiple intersection metrics were obtained through filtering, and a first consistency detection coefficient was configured, including: Filter the intersection of multiple text metrics, multiple table metrics, and multiple image metrics to obtain multiple intersection metrics; Calculate the ratio of the number of multiple intersection indicators to the number of multiple text indicators, multiple table indicators, and multiple image indicators, and take the minimum value as the indicator consistency coefficient to obtain the first consistency detection coefficient.

5. The intelligent content consistency detection method for multimodal reporting materials according to claim 1, characterized in that, Extract data from the tables to obtain multiple intersecting indicator tables, perform change analysis to obtain the change coefficients of the intersecting indicators from the multiple tables, and configure a second consistency detection coefficient, including: Based on multiple intersecting indicators, extract data from the tabular data to obtain a tabular dataset of multiple intersecting indicators; Based on multiple intersection indicator tables, the changes in the data of each type of intersection indicator table are calculated to obtain multiple change rates, which are used as the change coefficients of the intersection indicators of multiple tables. Configure a second consistency detection coefficient based on the change coefficients of the intersection indicators of multiple tables.

6. The intelligent content consistency detection method for multimodal reporting materials according to claim 5, characterized in that, Based on the change coefficients of the intersection indicators of multiple tables, a second consistency detection coefficient is configured, including: Calculate the discrete parameters of the change coefficients of the intersection indicators of multiple tables to obtain the discrete parameters of table changes; The ratio of the variation discrete parameter of the table to the maximum variation discrete parameter of the report material is calculated and used as the second consistency detection coefficient.

7. The intelligent content consistency detection method for multimodal reporting materials according to claim 1, characterized in that, The image data undergoes change identification based on multiple intersection metrics, yielding change coefficients for these metrics. A third consistency detection coefficient is then configured, including: An image indicator change recognizer is obtained, wherein the image indicator change recognizer is trained using a sample indicator set, a sample image data set, and a sample indicator change coefficient set; The image data is combined with multiple intersection indicators and input into the image indicator change recognizer, which outputs multiple image intersection indicator change coefficients. Calculate the discrete parameters of the change coefficients of the intersection indicators of multiple images to obtain the discrete parameters of image change; The ratio of the discrete parameter of the image change to the maximum discrete parameter of the report material is calculated and used as the third consistency detection coefficient.

8. The intelligent content consistency detection method for multimodal reporting materials according to claim 1, characterized in that, The training steps for the text indicator change recognizer include: Obtain a sample text data set, extract the data change coefficients of all indicators in it, and label them to obtain a set of sample text indicator change coefficients; A text indicator change recognizer was constructed based on natural language processing algorithms.

9. An intelligent detection system for content consistency of multimodal reporting materials, characterized in that: For implementing the intelligent content consistency detection method for multimodal reporting materials as described in any one of claims 1 to 8, the system comprises: The multimodal indicator extraction module is used to obtain report materials, extract text data, tabular data, and image data, and extract indicators to obtain multiple text indicators, multiple tabular indicators, and multiple image indicators, among which the indicators are change data indicators; The table consistency analysis module is used to filter and obtain multiple intersection indicators, configure the first consistency detection coefficient, extract table data to obtain multiple intersection indicator table data, perform change analysis to obtain the change coefficient of multiple table intersection indicators, and configure the second consistency detection coefficient. The image consistency analysis module is used to identify changes in multiple intersection indicators of image data, obtain the change coefficients of multiple image intersection indicators, and configure a third consistency detection coefficient. The fusion verification and detection module is used to fuse configuration consistency detection coefficients, extract associated text data and identify intersection index changes, obtain multiple text intersection index change coefficients, and verify them with the multiple table intersection index change coefficients and multiple image intersection index change coefficients to obtain consistency detection results.

Citation Information

Patent Citations

  • Intelligent data augmentation and cleaning system and method based on multi-modal consistency detection

    CN120179996A

  • Government affair digital human dynamic interaction method and system based on multi-modal large model

    CN120821813A