A data evaluation method based on power data report

Through deep learning and neural network processing power report text, multi-scale feature vectors are generated and abnormal identification combined with classifiers, the efficiency and accuracy of abnormal detection in power data reports are solved, ensuring the safe and efficient operation of the power system.

CN119296122BActive Publication Date: 2025-05-23WANSI INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411816643.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-23
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

现有技术在电力数据报表中难以准确识别异常数据,导致运营决策错误和安全隐患,传统方法耗时费力且误判漏判现象频发。

Method used

Using deep learning and neural network methods, the text content of power report is processed through semantic encoding, multi-scale feature vectors are generated, and exception recognition is combined with a classifier to realize automated and intelligent abnormal data detection.

Benefits of technology

It improves the efficiency and accuracy of abnormal detection of power data reports, reduces misjudgment and misjudgment, and ensures the safe and efficient operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119296122B_ABST
    Figure CN119296122B_ABST
Patent Text Reader

Abstract

The present application relates to the field of intelligent data analysis, and provides a data evaluation method based on power data reports, which first obtains power data reports from a power system, then preprocesses the power data reports to obtain the text content of the power reports, and then uses a natural language-based semantic analysis technology to perform semantic analysis on the text content of the power reports to determine whether there is abnormal data in the original power data reports. In this way, automated and intelligent abnormal data detection is achieved, greatly improving detection efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent data analysis, and more specifically, to a data evaluation method based on power data reports. Background Art

[0002] In the power system, the accuracy and reliability of data reports are crucial to ensure the safe and efficient operation of the system. However, in practice, due to problems that may occur in multiple links such as data collection, transmission and storage, there are inaccurate or abnormal data in the reports. These problems not only affect the operational decisions of the power system, but may also lead to waste of resources and even safety accidents.

[0003] Traditional methods may rely on manual review or simple statistical rules to detect abnormal data in power data reports. This method is not only time-consuming and laborious, but also prone to misjudgment or omission. In addition, traditional methods only stay at simple analysis of surface information such as numerical values ​​in reports, and it is difficult to deeply explore the semantic relationships hidden behind the data.

[0004] Therefore, an optimized data evaluation scheme based on power data reports is needed. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present application provides a data evaluation method based on power data reports, which includes:

[0006] Obtain power data reports from the power system;

[0007] Preprocessing the power data report to obtain power report text content;

[0008] Performing semantic coding processing on the electric power report text content to obtain a semantic understanding feature vector of the electric power report content;

[0009] Based on the semantic understanding feature vector of the power report content, an abnormality recognition result of the power data report is obtained.

[0010] In the above-mentioned data evaluation method based on electric power data reports, semantic encoding processing is performed on the text content of the electric power report to obtain a semantic understanding feature vector of the electric power report content, including: segmenting the text content of the electric power report and passing it through a word embedding layer to obtain multiple electric power report keyword embedding vectors; performing long-distance semantic encoding and medium-short-distance semantic encoding on the multiple electric power report keyword embedding vectors respectively to obtain a first-scale electric power report content semantic understanding feature vector and a second-scale electric power report content semantic understanding feature vector; performing feature internal matching alignment based on intrinsic feature decomposition on the first-scale electric power report content semantic understanding feature vector and the second-scale electric power report content semantic understanding feature vector to obtain the electric power report content semantic understanding feature vector.

[0011] In the above-mentioned data evaluation method based on electric power data reports, the multiple electric power report keyword embedding vectors are respectively subjected to long-distance semantic encoding and medium-short-distance semantic encoding to obtain a first-scale electric power report content semantic understanding feature vector and a second-scale electric power report content semantic understanding feature vector, including: passing the multiple electric power report keyword embedding vectors through a long-distance semantic encoder to obtain the first-scale electric power report content semantic understanding feature vector; passing the multiple electric power report keyword embedding vectors through a medium-short-distance semantic encoder to obtain the second-scale electric power report content semantic understanding feature vector.

[0012] In the above-mentioned data evaluation method based on electric power data reports, the long-distance semantic encoder is a converter-based semantic encoder, and the medium- and short-distance semantic encoder is a bidirectional long short-term memory neural network model.

[0013] In the above-mentioned data evaluation method based on electric power data reports, the first-scale electric power report content semantic understanding feature vector and the second-scale electric power report content semantic understanding feature vector are subjected to feature internal matching alignment based on intrinsic feature decomposition to obtain the electric power report content semantic understanding feature vector, including: performing normalization modulation based on vector norming on the first-scale electric power report content semantic understanding feature vector and the second-scale electric power report content semantic understanding feature vector to obtain a pre-aligned first-scale electric power report content semantic understanding feature vector and a pre-aligned second-scale electric power report content semantic understanding feature vector; performing intrinsic feature extraction on the pre-aligned first-scale electric power report content semantic understanding feature vector and the pre-aligned second-scale electric power report content semantic understanding feature vector to obtain a first electric power report content semantic understanding principal component feature vector and a second electric power report content semantic understanding principal component feature vector; performing intrinsic fine-grained alignment and fusion on the first electric power report content semantic understanding principal component feature vector and the second electric power report content semantic understanding principal component feature vector to obtain the electric power report content semantic understanding feature vector.

[0014] In the above-mentioned data evaluation method based on electric power data reports, an anomaly recognition result of the electric power data report is obtained based on the semantic understanding feature vector of the electric power report content, including: passing the semantic understanding feature vector of the electric power report content through a classifier-based anomaly identifier to obtain the anomaly recognition result, and the anomaly recognition result is used to indicate whether there is a data anomaly in the electric power data report.

[0015] In the above-mentioned data evaluation method based on electric power data reports, based on the electric power report content semantic understanding feature vector, an abnormality recognition result of the electric power data report is obtained, including: using the fully connected layer of the classifier to fully connect the electric power report content semantic understanding feature vector to obtain the electric power report content semantic understanding fully connected encoded feature vector; inputting the electric power report content semantic understanding fully connected encoded feature vector into the Softmax classification function of the classifier to obtain the probability value of the electric power report content semantic understanding feature vector belonging to each classification label, the classification label including one for indicating that the electric power data report has data anomalies and one for indicating that the electric power data report does not have data anomalies; the classification label corresponding to the largest of the probability values ​​is determined as the abnormality recognition result.

[0016] This application has significant technical effects due to the adoption of the above technical solutions:

[0017] The data evaluation method based on power data report provided by the present application first obtains power data report from the power system, then pre-processes the power data report to obtain the text content of the power report, and then uses the semantic analysis technology based on natural language to perform semantic analysis on the text content of the power report to determine whether there is abnormal data in the original power data report. In this way, automatic and intelligent abnormal data detection is realized, greatly improving the detection efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0019] Figure 1 The present invention is a flowchart of a data evaluation method based on a power data report according to an embodiment of the present application.

[0020] Figure 2 Flow chart of step S130 in the data evaluation method based on power data report according to an embodiment of the present application.

[0021] Figure 3 Flow chart of step S132 in the data evaluation method based on power data report according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0023] In the power system, the accuracy and reliability of data reports are crucial to ensure the safe and efficient operation of the system. The power system covers multiple interrelated and collaborative links such as power generation, transmission, transformation, distribution and power consumption. Each link is constantly generating massive amounts of data, and these data are collected, transmitted and stored, and finally gathered into data reports to guide system operation decisions.

[0024] However, in the actual operation process, the data collection process may be biased due to factors such as sensor failure, unreasonable installation location, or interference from complex external environments. For example, in some substations in remote areas, due to the harsh natural environment, some power parameter collection sensors may be affected by long-term wind and sand erosion, extreme temperature changes, etc., resulting in inaccurate data collection. In the data transmission stage, unstable network signals, transmission protocol compatibility issues, or communication line failures may cause data loss, misordering, or partial distortion, just like in some mountainous transmission line areas, due to insufficient coverage of communication base stations, data transmission is intermittent, which in turn affects the integrity and accuracy of the data. As for the data storage process, hardware failures, software vulnerabilities, or imperfect data management strategies of storage devices may also cause incorrect data storage or loss, resulting in inaccurate or abnormal data in the final data report.

[0025] The negative impact of these problems in data reports cannot be underestimated. On the one hand, it will seriously interfere with the operational decisions of the power system. The decisions made by operators based on inaccurate data reports, such as power generation allocation, equipment maintenance plan arrangement, and grid load adjustment, are likely to run counter to actual needs, and thus fail to achieve optimal allocation of power resources, resulting in frequent waste of resources. On the other hand, more seriously, these problems are very likely to cause safety accidents. For example, due to incorrect grid load data, the overloaded line cannot be warned and adjusted in time, which may cause line overheating, short circuit and other faults, thereby affecting the stability and safety of power supply, and even endangering the lives of personnel and the safety of equipment and property.

[0026] Traditional methods used in the past to detect abnormal data in power data reports mostly rely on manual review or simple statistical rules. Manual review often requires a lot of manpower and time costs. Staff need to carefully review, compare and analyze massive data reports one by one. In addition, long-term and high-intensity work can easily cause fatigue and negligence of reviewers, resulting in misjudgment or missed judgments. Simple statistical rules only judge based on surface information such as the size of the value and the range of change. They only stay at the superficial level of analysis of the values ​​in the report, and it is difficult to deeply explore the semantic relationships hidden behind the data, such as the correlation between the operating parameters of different equipment, the internal logic of data changes of the same equipment under different working conditions, etc. This makes traditional methods incapable of dealing with complex and changeable power data report anomaly detection. Therefore, an optimized data evaluation scheme based on power data reports is expected.

[0027] In recent years, deep learning and neural networks have been widely used in computer vision, natural language processing, text signal processing and other fields. In addition, deep learning and neural networks have also demonstrated a level close to or even beyond that of humans in image classification, object detection, semantic segmentation, text translation and other fields. The development of deep learning and neural networks has provided new solutions and solutions for data evaluation based on power data reports.

[0028] Figure 1 FIG. 1 is a flow chart of a data evaluation method based on a power data report according to an embodiment of the present application. Figure 1 As shown, according to the data evaluation method based on power data reports according to the embodiment of the present application, the method includes: S110, obtaining a power data report from a power system; S120, preprocessing the power data report to obtain the text content of the power report; S130, performing semantic encoding processing on the text content of the power report to obtain a semantic understanding feature vector of the power report content; S140, obtaining an abnormality recognition result of the power data report based on the semantic understanding feature vector of the power report content.

[0029] In step S110, an electric power data report is obtained from the electric power system. It should be understood that the electric power data report obtained in the electric power system includes information on power generation equipment such as the output power, voltage, current, and frequency of the generator set, information on transmission equipment parameters such as the voltage level, current carrying capacity, resistance value, and real-time operating current of the transmission line, data on transformer substation equipment such as the oil temperature, winding temperature, ratio, and load rate of the transformer, as well as power generation data, power consumption data, and power load data. By comprehensively analyzing these rich and specific data in the electric power data report, it can be known whether each indicator is within the normal operating range and logical relationship. For example, there is a certain corresponding relationship between the parameters of the power generation equipment and the power generation, and there is also a reasonable correlation logic between the power consumption and the power load. When these internal relationships are broken and unreasonable data combinations appear, it can be determined that there are data anomalies in the report.

[0030] In step S120, the power data report is preprocessed to obtain the text content of the power report. Accordingly, considering that the power data report comes from a wide range of sources, it covers data from multiple links, different equipment and various business scenarios in the power system. During the collection process, due to factors such as the accuracy differences of the collection equipment and environmental interference, the data may be noisy. For example, the collected voltage value may fluctuate slightly due to electromagnetic interference. If these noise data are not processed, it will affect the subsequent accurate understanding of the report semantics and interfere with abnormal judgment. In actual situations, there may be a phenomenon of missing data. For example, some monitoring data of the transmission line in a certain period of time is not collected, or a certain parameter of the equipment is not fully recorded due to sensor failure. This will lead to insufficient key information in subsequent analysis, affecting the accurate evaluation of the report as a whole. In addition, there are data in different formats in the power data report. In order to effectively eliminate problems such as data quality, data format and data integrity, in the technical solution of this application, it is necessary to preprocess the power data report. After preprocessing to remove noise data, supplement missing parts and standardize the format, subsequent semantic analysis can be based on more accurate and complete data, reducing misjudgments caused by data quality issues. This will enable more accurate detection of real anomalies in data reports and improve the accuracy of abnormal judgments on power data reports.

[0031] In step S130, semantic coding is performed on the power report text content to obtain a semantic understanding feature vector of the power report content. Specifically, Figure 2 FIG. 1 is a flow chart of step S130 in the data evaluation method based on the power data report according to an embodiment of the present application. Figure 2As shown, the step S130 includes: S131, segmenting the text content of the electricity report and passing it through a word embedding layer to obtain multiple electricity report keyword embedding vectors; S132, respectively performing long-distance semantic encoding and medium-short-distance semantic encoding on the multiple electricity report keyword embedding vectors to obtain a first-scale electricity report content semantic understanding feature vector and a second-scale electricity report content semantic understanding feature vector; S133, performing feature internal matching alignment based on intrinsic feature decomposition on the first-scale electricity report content semantic understanding feature vector and the second-scale electricity report content semantic understanding feature vector to obtain the electricity report content semantic understanding feature vector.

[0032] In step S131, the text content of the power report is segmented and then passed through the word embedding layer to obtain multiple power report keyword embedding vectors. It should be understood that the text content of the power report obtained after preprocessing is essentially a manifestation of natural language, which contains many sentences, words, etc. that describe various aspects of the power system. Natural language is usually a continuous stream of characters, and it is difficult for computers to directly process such text forms to mine semantic information. In order to better perform semantic analysis operations, it is necessary to segment the text content of the power report, that is, to split the original long text into meaningful basic units, so that the originally complex text semantic understanding problem can be transformed into the problem of semantic analysis of each word and the relationship between words, reducing the difficulty and dimension of the analysis. Next, the text content of the power report after word segmentation is input into the word embedding layer. The word embedding layer can map words in natural language to a low-dimensional vector space. In this vector space, each word is represented as a vector of a fixed dimension, for example, it is common to represent words as vectors of 100 dimensions, 300 dimensions, etc. In addition, the representation of vectors contains the semantic information of words. Words with similar semantics are often close to each other in the vector space. For example, the two semantically related words "generator" and "generator set" have vectors that are close to each other in space. Through this vector representation, computers can use numerical calculations to process the semantics of words that are difficult to quantify.

[0033] In step S132, the plurality of power report keyword embedding vectors are respectively subjected to long-distance semantic coding and medium-short-distance semantic coding to obtain a first-scale power report content semantic understanding feature vector and a second-scale power report content semantic understanding feature vector. Specifically, Figure 3 FIG. 1 is a flow chart of step S132 in the data evaluation method based on the power data report according to an embodiment of the present application. Figure 3As shown, the step S132 includes: S1321, passing the multiple electricity report keyword embedding vectors through a long-distance semantic encoder to obtain the first-scale electricity report content semantic understanding feature vector; S1322, passing the multiple electricity report keyword embedding vectors through a medium-short distance semantic encoder to obtain the second-scale electricity report content semantic understanding feature vector.

[0034] In step S1321, the multiple power report keyword embedding vectors are passed through a long-distance semantic encoder to obtain the semantic understanding feature vector of the first-scale power report content. It should be understood that when describing the power system-related situation, the power report text content often has long-distance semantic associations across multiple words or even multiple sentences. For example, in the content describing the impact of power generation equipment failure on the overall operation of the power grid, it may first be mentioned that "a certain generator set has a fault and shuts down", and then in the subsequent distant text part, it is mentioned that "the power supply of the regional power grid fluctuates, and the load rate of multiple substations changes significantly". Here, there is a causal relationship between "generator set failure and shutdown" and "substation load rate change", but it is a long-distance logical relationship in the text. It is difficult to capture this kind of semantic connection that is far apart by relying solely on the keyword embedding vector itself, so long-distance semantic encoding processing is needed to mine and integrate such long-distance semantic information to fully understand the deep semantics expressed by the power report text, so as to accurately determine whether there is data anomaly. Specifically, the multiple power report keyword embedding vectors are input into the long-distance semantic encoder to capture long-distance semantic associations. The long-distance semantic encoder described here is a semantic encoder based on a transformer. The transformer is a deep learning architecture based on the attention mechanism. It is mainly composed of a multi-head attention layer, a feedforward neural network layer, and some normalization and residual connections. The multi-head attention layer is its core part, which enables the model to focus on the association between information at different positions from multiple "perspectives" when processing input information (here, multiple power report keyword embedding vectors), that is, to calculate the correlation weights between different vectors, dynamically focus on the parts that need to be focused on according to these weights, and integrate information at different positions. The feedforward neural network layer is used to further perform nonlinear transformation on the information processed by the attention mechanism to enhance the expression ability of the model. Normalization and residual connections help improve the stability of the model and the training effect. The transformer-based semantic encoder is used to perform long-distance semantic encoding on multiple power report keyword embedding vectors, which helps to deeply mine the long-distance semantic information of the power report text, and thus provides strong support for accurately evaluating whether there are abnormalities in the power data report.

[0035] In step S1322, the multiple power report keyword embedding vectors are passed through a medium-short distance semantic encoder to obtain the second scale power report content semantic understanding feature vector. It should be understood that although the long-distance semantic encoder can capture the semantic associations across a large length in the power report text, the close semantic connection between words in the local range of the text should not be ignored. There are often many words in adjacent or close positions in the power report that jointly describe a specific power equipment state, local operation status, etc., such as "the oil temperature of the transformer is 80 degrees Celsius, and the oil temperature is slowly rising." Here, "transformer oil temperature", "80 degrees Celsius", "slowly rising trend" are a medium-short distance semantic relationship between these words, which together convey the current key operation information of the transformer. Medium-short distance semantic encoding of keyword embedding vectors helps to focus on and extract such local tight semantic features, thereby improving the semantic understanding of the power report text from a micro level, avoiding the omission of important local semantic details, and making the subsequent semantic analysis of the entire report content more comprehensive. The medium-short distance semantic encoder described here is a bidirectional long short-term memory neural network model. The bidirectional long short-term memory neural network model (Bi-LSTM) has the ability to process information in both directions. It can process multiple input power report keyword embedding vectors from both the forward and reverse directions. For a certain word in the power report text, it can not only consider the semantic relationship formed by the previous word with it (forward order), but also take into account the association between the following words and it (reverse order). For example, in the sentence "The output power of the generator is stable, and the voltage is also kept near the rated value", for the word "output power", it has a semantic connection with words such as "generator" from a forward perspective, and it is associated with "stable" and subsequent words such as "voltage" from a reverse perspective. Bi-LSTM can fully capture this short-distance semantic relationship in the forward and backward directions through bidirectional processing, thereby more comprehensively integrating the semantic information in the local range and generating a more accurate second-scale power report content semantic understanding feature vector that can better reflect the actual semantic situation.

[0036] In step S133, the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector are aligned by feature internal matching based on intrinsic feature decomposition to obtain the power report content semantic understanding feature vector. It should be understood that the power report text content contains multi-level and multi-dimensional semantic information. The first-scale power report content semantic understanding feature vector extracted by the long-distance semantic encoder focuses on capturing the macro and systematic semantic associations across a large text, such as the causal relationship between the failure of power generation equipment and the change of the operating status of the entire power grid, and the mutual influence of power equipment in different links and regions in the long-distance text description. The second-scale power report content semantic understanding feature vector obtained by the medium- and short-distance semantic encoder focuses on the close semantic connection between words in a local range, such as the micro-level descriptions of specific parameters and operating status of a single device and their mutual relationship. In order to comprehensively integrate the semantic information of the electricity report text content from macro to micro, avoid losing important semantic details by focusing on a single scale, and form a comprehensive feature representation that fully covers the semantic logic of all levels of the electricity report, so as to provide a comprehensive basis for accurately judging data anomalies, it is necessary to fuse the first-scale electricity report content semantic understanding feature vector and the second-scale electricity report content semantic understanding feature vector.

[0037] In particular, considering that the semantic understanding feature vector of the first-scale power report content is obtained by processing the long-distance semantic encoder, it may overemphasize the global context information and ignore the local details; the semantic understanding feature vector of the second-scale power report content is obtained by processing the medium-short distance semantic encoder, which will overemphasize the local information and fail to make full use of the global context. That is, the scale ranges of these two feature vectors are different. When the semantic understanding feature vectors of the first-scale and second-scale power report content are to be fused, due to the difference in their original scale ranges, the "granularity" of the semantic information carried by each is different. For example, a certain dimension in the feature vector obtained by the long-distance semantic encoder may represent the abstract features of the key semantic associations at the beginning and end of the entire power report text, while the corresponding dimension features obtained by the medium-short distance semantic encoder are only the semantic features within a small local text. In simple fusion (such as direct concatenation, weighted summation and other common fusion methods), if there is no appropriate normalization or adjustment strategy, it is easy to make the features that are originally very important from a long-distance perspective "equal" with the medium-short distance features after fusion, or the opposite situation occurs, thereby misestimating the importance of features of different scales. This will mislead the anomaly identifier and ultimately affect the accuracy of the anomaly judgment of the power data report data. Therefore, in the technical solution of the present application, it is necessary to perform feature internal matching alignment based on intrinsic feature decomposition on the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector to obtain the power report content semantic understanding feature vector.

[0038] Specifically, in an embodiment of the present application, the first-scale electricity report content semantic understanding feature vector and the second-scale electricity report content semantic understanding feature vector are subjected to feature internal matching alignment based on intrinsic feature decomposition to obtain the electricity report content semantic understanding feature vector, including: performing normalization modulation based on vector norming on the first-scale electricity report content semantic understanding feature vector and the second-scale electricity report content semantic understanding feature vector to obtain a pre-aligned first-scale electricity report content semantic understanding feature vector and a pre-aligned second-scale electricity report content semantic understanding feature vector; performing intrinsic feature extraction on the pre-aligned first-scale electricity report content semantic understanding feature vector and the pre-aligned second-scale electricity report content semantic understanding feature vector to obtain a first electricity report content semantic understanding principal component feature vector and a second electricity report content semantic understanding principal component feature vector; performing intrinsic fine-grained alignment and fusion on the first electricity report content semantic understanding principal component feature vector and the second electricity report content semantic understanding principal component feature vector to obtain the electricity report content semantic understanding feature vector.

[0039] In the embodiment of the present application, specifically, another implementation expression of performing feature internal matching alignment based on intrinsic feature decomposition on the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector to obtain the power report content semantic understanding feature vector can be: the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector are processed by the following formula to obtain the power report content semantic understanding feature vector; wherein the formula is:

[0040]

[0041]

[0042]

[0043]

[0044]

[0045] in, Represents the semantic understanding feature vector of the first-scale power report content, Represents the semantic understanding feature vector of the second-scale power report content, The first scale power report content semantic understanding feature vector The eigenvalues ​​at the positions, The first one represents the semantic understanding feature vector of the second scale power report content The eigenvalues ​​at the positions, represents the natural exponential function, represents the square of the vector's two-norm, represents the pre-aligned first-scale power report content semantic understanding feature vector, represents the pre-aligned second-scale power report content semantic understanding feature vector, represents intrinsic feature extraction, represents the sequence of first eigendecomposition vectors, represents the first diagonal matrix, and Represent the first position and the first diagonal matrix respectively. The eigenvalues ​​at the positions, represents the transpose of a vector, represents vector concatenation, , , The first, second and third sequences of the main component feature vectors of the semantic understanding of the first power report content are represented respectively. feature vectors, Represents the main component feature vector of the semantic understanding of the first power report content, represents the sequence of the second eigendecomposition vectors, represents the second diagonal matrix, and Represent the first and second positions of the second diagonal matrix respectively. The eigenvalues ​​at the positions, , , The first, second and third sequences of the main component feature vectors of the semantic understanding of the second power report content are represented respectively. feature vectors, Represents the main component feature vector of the semantic understanding of the second power report content, It means subtracting by position. It means point multiplication by position. It means adding by position. represents point convolution, Represents the feature vector of semantic understanding of power report content.

[0046] That is, in the technical solution of the present application, the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector are aligned by feature internal matching based on intrinsic feature decomposition to obtain the power report content semantic understanding feature vector, which first performs normalization modulation based on vector norming on the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector to obtain the pre-aligned first-scale power report content semantic understanding feature vector and the pre-aligned second-scale power report content semantic understanding feature vector. This process is not just for simple scale adjustment, but an important technology in feature engineering, which aims to solve the problem of inconsistent scales between features in the data set. By calculating the modulus of the feature vector and dividing each vector by its own modulus, it can be ensured that all feature vectors have the same length, thereby eliminating the impact of scale differences between different feature vectors. Normalization processing not only helps numerical stability in subsequent processing, but also makes the comparison between feature vectors more fair and reasonable, laying the foundation for subsequent feature alignment. In machine learning models, especially algorithms such as neural networks, the scale of input features directly affects the speed and direction of weight updates. Therefore, normalization can be seen as a regularization method that helps prevent overfitting and also helps avoid the problem of gradient vanishing or exploding.

[0047] Next, the pre-aligned first-scale power report content semantic understanding feature vector and the pre-aligned second-scale power report content semantic understanding feature vector are subjected to intrinsic feature extraction to obtain the first power report content semantic understanding principal component feature vector and the second power report content semantic understanding principal component feature vector. That is, after completing the normalization modulation, the pre-aligned first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector are subjected to intrinsic feature extraction. As an unsupervised linear dimensionality reduction method, intrinsic feature extraction is widely used in data compression and feature extraction. This process determines the principal components by finding the direction of the maximum variance of the data. These principal components are new coordinate axes in the original feature space, which can capture the structure of the data with minimal information loss. It is worth mentioning that by selecting the eigenvectors corresponding to the first few largest eigenvalues, the dimension of the feature space can be effectively reduced while maintaining the key information of the data. This dimensionality reduction can not only accelerate the model training process, but also help the model to generalize better, reduce the impact of noise, and thus improve the performance of the model.

[0048] Finally, the first electricity report content semantic understanding principal component feature vector and the second electricity report content semantic understanding principal component feature vector are intrinsically aligned and fused to obtain the electricity report content semantic understanding feature vector. The goal of this step is to find the best matching point or area between the principal component feature vectors of the two target analysis raw data to achieve alignment between the two. Fine-grained alignment goes beyond the traditional rough fusion methods such as simple weighted average or splicing, and needs to consider the time series characteristics or spatial distribution characteristics of the feature vector. The fusion after alignment is not a simple numerical operation, but requires the design of a specific mechanism to determine how to combine the information of the two feature vectors. This content-based intelligent fusion method can more effectively utilize multi-source information and improve the model's representation ability and decision-making quality.

[0049] In step S140, based on the semantic understanding feature vector of the power report content, the abnormality recognition result of the power data report is obtained. Specifically, in the embodiment of the present application, the step S140 includes: passing the semantic understanding feature vector of the power report content through a classifier-based abnormality identifier to obtain the abnormality recognition result, and the abnormality recognition result is used to indicate whether there is a data abnormality in the power data report. More specifically, in the embodiment of the present application, the step S140 includes: using the fully connected layer of the classifier to fully connect the semantic understanding feature vector of the power report content to obtain a fully connected encoded feature vector of the semantic understanding of the power report content; inputting the fully connected encoded feature vector of the semantic understanding of the power report content into the Softmax classification function of the classifier to obtain the probability value of the semantic understanding feature vector of the power report content belonging to each classification label, and the classification label includes a label for indicating that there is a data abnormality in the power data report and a label for indicating that there is no data abnormality in the power data report; the classification label corresponding to the largest of the probability values ​​is determined as the abnormality recognition result. It should be understood that after the previous operations of preprocessing, word embedding, long-distance and medium-short-distance semantic encoding, and feature vector fusion, the semantic understanding feature vector of the power report content is finally obtained. Although this vector contains rich semantic information about the power report text, it is still only a feature representation form and cannot directly represent whether there are data anomalies in the power data report. The classifier-based anomaly identifier can receive this feature vector as input, analyze and judge the feature vector according to the rules and patterns learned internally, and output the corresponding anomaly identification result, which realizes the key transformation from semantic features to clear judgment of anomalies. The most direct role of the anomaly identification result is to inform the relevant personnel of the power system whether there are data anomalies in the power data report. If the result shows that there are anomalies in the report, it means that the data in the current report may be inaccurate and unreliable, and the data needs to be further verified and corrected to ensure that the subsequent power system operation decisions, scheduling arrangements, etc. based on the report data are based on accurate data, avoiding problems such as waste of resources and safety accidents caused by data errors.

[0050] In summary, the data evaluation method based on the power data report according to the embodiment of the present application is explained, which first obtains the power data report from the power system, then pre-processes the power data report to obtain the power report text content, and then uses the natural language-based semantic analysis technology to perform semantic analysis on the power report text content to determine whether there is abnormal data in the original power data report. In this way, automatic and intelligent abnormal data detection is achieved, greatly improving the detection efficiency and accuracy.

Claims

1. A data evaluation method based on power data reports, characterized in that: include: Obtain power data reports from the power system; Preprocessing the power data report to obtain power report text content; Performing semantic coding processing on the electric power report text content to obtain a semantic understanding feature vector of the electric power report content; Based on the semantic understanding feature vector of the power report content, an abnormality recognition result of the power data report is obtained; The method of performing semantic coding on the electric power report text content to obtain a semantic understanding feature vector of the electric power report content includes: After the electric power report text content is segmented, it is passed through a word embedding layer to obtain multiple electric power report keyword embedding vectors; Performing long-distance semantic coding and medium-short-distance semantic coding on the plurality of power report keyword embedding vectors respectively to obtain a first-scale power report content semantic understanding feature vector and a second-scale power report content semantic understanding feature vector; Performing feature internal matching alignment based on intrinsic feature decomposition on the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector to obtain the power report content semantic understanding feature vector; The first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector are subjected to feature internal matching alignment based on intrinsic feature decomposition to obtain the power report content semantic understanding feature vector, including: Performing normalization modulation based on vector norming on the first-scale power report content semantic understanding feature vector and the second-scale power report content semantic understanding feature vector to obtain a pre-aligned first-scale power report content semantic understanding feature vector and a pre-aligned second-scale power report content semantic understanding feature vector; Performing intrinsic feature extraction on the pre-aligned first-scale power report content semantic understanding feature vector and the pre-aligned second-scale power report content semantic understanding feature vector to obtain a first power report content semantic understanding principal component feature vector and a second power report content semantic understanding principal component feature vector; Performing intrinsic fine-grained alignment and fusion on the first power report content semantic understanding main component feature vector and the second power report content semantic understanding main component feature vector to obtain the power report content semantic understanding feature vector; Among them, based on the semantic understanding feature vector of the power report content, an abnormality recognition result of the power data report is obtained, including: passing the semantic understanding feature vector of the power report content through a classifier-based abnormality identifier to obtain the abnormality recognition result, and the abnormality recognition result is used to indicate whether there is a data anomaly in the power data report.

2. The data evaluation method based on power data report according to claim 1 is characterized in that: The plurality of power report keyword embedding vectors are respectively subjected to long-distance semantic coding and medium-short-distance semantic coding to obtain a first-scale power report content semantic understanding feature vector and a second-scale power report content semantic understanding feature vector, including: Passing the plurality of power report keyword embedding vectors through a long distance semantic encoder to obtain a first scale power report content semantic understanding feature vector; The plurality of power report keyword embedding vectors are passed through a medium- and short-distance semantic encoder to obtain the second-scale power report content semantic understanding feature vector.

3. The data evaluation method based on power data report according to claim 2 is characterized in that: The long-distance semantic encoder is a converter-based semantic encoder, and the medium- and short-distance semantic encoder is a bidirectional long short-term memory neural network model.

4. The data evaluation method based on power data report according to claim 3 is characterized in that: Based on the semantic understanding feature vector of the power report content, an abnormality recognition result of the power data report is obtained, including: Using the fully connected layer of the classifier to perform fully connected encoding on the power report content semantic understanding feature vector to obtain a power report content semantic understanding fully connected encoding feature vector; Inputting the fully connected encoded feature vector of the semantic understanding of the power report content into the Softmax classification function of the classifier to obtain the probability value of the semantic understanding feature vector of the power report content belonging to each classification label, wherein the classification label includes a label for indicating that the power data report has data anomalies and a label for indicating that the power data report does not have data anomalies; The classification label corresponding to the largest probability value among the probability values ​​is determined as the abnormality recognition result.

Citation Information

Patent Citations

  • Intellectual property management method and system based on data lake

    CN115994177A