Data exception accurate identification method, system and device and medium
By combining multi-level composite neural networks and generative adversarial networks, the problem of dynamic anomaly detection in multi-source heterogeneous data environments is solved, achieving high-precision, real-time, and interpretable data correction capabilities, and improving the anomaly detection effect of power systems.
Patent Information
- Application Number
- CN202510842907.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-11-25
AI Technical Summary
Existing anomaly detection technologies are ill-suited to dynamic scenarios in multi-source heterogeneous data environments, lack interpretability, struggle to handle multimodal data, and lack data correction capabilities, resulting in insufficient detection accuracy and real-time performance.
A multi-level composite neural network structure is adopted, which combines wavelet decomposition and generative adversarial network for feature extraction and state modeling. The judgment threshold is adjusted through reinforcement learning, and data completion or replacement is performed when an anomaly is detected, so as to realize feature fusion and dynamic judgment of multi-source data.
It enables accurate identification and automatic repair of dynamic anomalies, improves the real-time performance, accuracy and intelligence of the detection system, reduces the false alarm rate and improves the interpretability of the results and the adaptability of the system.
Smart Images

Figure CN121009484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method, system, device, and medium for accurate identification of data anomalies. Background Technology
[0002] In data-driven applications, anomaly detection is a crucial means of ensuring system stability and operational efficiency. This technology is widely used in key areas such as power dispatching, equipment monitoring, intelligent manufacturing, and financial risk control, and is commonly used to identify potential system failures, signal interference, data transmission errors, or external attacks.
[0003] Most existing anomaly detection methods are based on fixed statistical rules, using preset thresholds to determine whether data is abnormal. These methods are simple to implement, computationally inexpensive, and perform well under conditions of stable data patterns and low noise levels. However, in real-world systems, data often exhibits non-stationarity, influenced by factors such as seasonal variations, user behavior, equipment malfunctions, or unexpected events, demonstrating significant dynamism and uncertainty. In such scenarios, fixed threshold strategies are prone to false alarms or missed detections, lacking adaptability to complex and dynamic environments.
[0004] In recent years, deep learning-based anomaly detection models have emerged, especially unsupervised learning methods, which can uncover latent patterns in data and identify anomalous behavior without explicit labeling. These methods have shown certain advantages in detection accuracy. However, due to their complex model structures, they often lack interpretability, making their output difficult for users to understand or trust, especially in practical applications requiring human intervention for decision-making, where deployment becomes challenging.
[0005] Most existing methods focus on anomaly detection, but lack the ability to correct errors after detection. When data is missing or measurement errors occur, the system can often only mark the anomaly, without being able to complete or replace the data. This design necessitates manual analysis and processing of the detection results, limiting its efficiency in high-real-time and highly automated scenarios.
[0006] In practical applications, data sources are becoming increasingly diverse, including both structured sensor data and continuous variables such as equipment parameters, as well as unstructured textual information such as inspection reports and operation logs. Traditional detection methods often rely on a single data source, making it difficult to integrate data from different modalities and lacking a unified information fusion mechanism. This limitation reduces the ability to identify complex events and restricts the generalization performance of algorithms in multi-source heterogeneous environments.
[0007] The existing technology has the following shortcomings:
[0008] 1. Lack of effective integration of multi-source heterogeneous data: Power system data comes from diverse sources, including SCADA systems, smart meters, PMUs, weather stations, etc., with different data formats and sampling frequencies. Existing methods are difficult to effectively integrate these heterogeneous data.
[0009] 2. Difficulty in capturing multi-scale time dependence: Power system data simultaneously contains short-term fluctuations (such as load changes, equipment start-up and shutdown) and long-term trends (such as seasonal changes, equipment aging), and existing methods have difficulty capturing these features at different time scales at the same time.
[0010] 3. Lack of deep integration with domain knowledge: There are complex physical and topological relationships between devices and parameters in power systems, and existing methods cannot fully utilize this domain knowledge to improve the accuracy of anomaly detection.
[0011] 4. Lack of interpretability of anomaly detection results: Existing deep learning methods are usually "black box" models, which make it difficult to explain the cause and scope of impact of anomalies, which is not conducive to the decision-making of operation and maintenance personnel.
[0012] 5. Difficulty in adapting to dynamic changes in data distribution: The operating status of the power system is affected by external factors (such as weather and electricity consumption behavior), and the data distribution changes over time. Existing methods are difficult to adapt to this non-stationary characteristic.
[0013] Therefore, existing anomaly detection technologies still have significant shortcomings in the following aspects: static rules cannot adapt to dynamic scenarios; deep models lack interpretability; most systems lack data correction capabilities; and traditional methods struggle to handle multimodal data. To address these issues, there is an urgent need to develop anomaly detection methods with dynamic judgment capabilities, interpretable results, support for automatic repair, and the ability to integrate structured and unstructured information, in order to meet the practical needs of complex systems for high accuracy, strong robustness, and high automation. Summary of the Invention
[0014] In view of the above-mentioned problems, the present invention is proposed.
[0015] Therefore, the technical problem solved by this invention is: how to achieve accurate identification and automatic repair of dynamic anomalies in a multi-source heterogeneous data environment, thereby improving the real-time performance, accuracy and intelligence level of the detection system.
[0016] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for accurate identification of data anomalies, comprising,
[0017] The system acquires multi-source data from the target system, including structured and unstructured data. It preprocesses the multi-source data, including denoising and missing value labeling for time-series structured data, and entity extraction and formatting for text-based unstructured data. The preprocessed data is then input into an anomaly detection model with time-series modeling and feature fusion capabilities for feature extraction and state modeling to obtain corresponding anomaly scores. Based on historical data statistical characteristics and the current anomaly score, a dynamic judgment threshold is determined, and the score is compared with the threshold to output an anomaly judgment result. If the anomaly judgment result is anomaly, a correction process is initiated, using a generative model to complete or replace the anomalous data to restore data continuity and credibility.
[0018] As a preferred embodiment of the data anomaly accurate identification method described in this invention, the preprocessing includes: performing wavelet decomposition on the target data to extract wavelet coefficients of different frequency bands; estimating the noise standard deviation based on the median absolute deviation of the high-frequency part of the wavelet coefficients; calculating a threshold based on the noise standard deviation and the number of sampling points; and performing threshold filtering on the wavelet coefficients.
[0019] As a preferred embodiment of the data anomaly accurate identification method described in this invention, the anomaly identification model includes a multi-level composite neural network structure, specifically including a feature extraction layer, a temporal modeling layer, a feature fusion layer, and an output layer.
[0020] As a preferred embodiment of the data anomaly accurate identification method described in this invention, the following steps are included: The judgment threshold is adaptively adjusted using a reinforcement learning algorithm; the reinforcement learning algorithm employs a dual-delay deep deterministic policy gradient algorithm and optimizes the policy based on a reward function that includes false positive rate, recall rate, and threshold change magnitude to update the judgment threshold for each period; the correction process uses a generative adversarial network (GAN) to complete or replace anomalous data; the GAN includes a generator and a discriminator, the generator generates missing values or replaces anomalous values based on the context information of the time series where the anomalous data is located; the discriminator is used to determine the distribution difference between the generated data and the real data.
[0021] In a preferred embodiment of the data anomaly accurate identification method described in this invention, the anomaly determination result includes an anomaly type label; the anomaly type label is used to identify any one of temperature anomaly, current anomaly, load anomaly, and data missing; the structured data and unstructured data undergo feature fusion before being input into the anomaly identification model; the unstructured data includes operation and maintenance logs or inspection reports; the feature fusion includes extracting and encoding text entities from the unstructured data and then concatenating them with the structured data.
[0022] Furthermore, a cross-modal feature mapping mechanism is used to establish the correlation between different data sources.
[0023] As a preferred embodiment of the data anomaly accurate identification method described in this invention, the wavelet decomposition adopts the Daubechies 6th order wavelet, and the decomposition level is 6 levels.
[0024] As a preferred embodiment of the data anomaly accurate identification method described in this invention, the feature extraction layer comprises a one-dimensional convolutional neural network, including a first convolutional layer with 32 convolutional kernels, a kernel size of 3, a stride of 1, and a ReLU activation function, and a second convolutional layer with 64 convolutional kernels, a kernel size of 3, a stride of 1, and a ReLU activation function. Each convolutional layer is followed by a batch normalization layer and a Dropout layer, and a residual connection is added between the two convolutional layers. The temporal modeling layer comprises a bidirectional long short-term memory network, including 64 hidden units, and outputs... The first layer of the BiLSTM has a dimension of 128 and 64 hidden units, while the second layer of the BiLSTM has an output dimension of 128. Each BiLSTM layer is followed by a Dropout layer. The feature fusion layer includes a multi-head self-attention mechanism, specifically including 8 attention heads, each with a dimension of 16. A query, key, and value matrix is generated through linear transformation, attention weights are calculated, and a weighted result is output. The output layer adopts a fully connected layer structure, including two hidden layers with dimensions of 128 and 64, respectively. Finally, an anomaly score is output through a Sigmoid activation function.
[0025] Another objective of this invention is to provide a system for accurately identifying data anomalies.
[0026] To address the aforementioned technical problems, this invention provides the following technical solution: a data anomaly accurate identification system, comprising: a data acquisition module for acquiring multi-source data from a target system, the multi-source data including structured and unstructured data; a preprocessing module for preprocessing the multi-source data, the preprocessing including denoising and missing value labeling of time-series structured data; entity extraction and formatting of text-based unstructured data; a feature fusion module for inputting the preprocessed data into an anomaly identification model with time-series modeling and feature fusion capabilities, performing feature extraction and state modeling to obtain corresponding anomaly scores; an anomaly evaluation module for determining a dynamic judgment threshold based on historical data statistical characteristics and the current anomaly score, comparing the score with the threshold, and outputting an anomaly judgment result; and an anomaly correction module for initiating a correction process when the anomaly judgment result is anomaly, supplementing or replacing the anomaly data based on a generative model to restore data continuity and credibility.
[0027] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the aforementioned method for accurate identification of data anomalies.
[0028] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the aforementioned method for accurate identification of data anomalies.
[0029] The beneficial effects of this invention are: it enables accurate detection and automatic correction of abnormal states in dynamic data environments, improving the system's practicality and responsiveness in complex application scenarios. Compared to existing static threshold judgment and detection mechanisms without correction capabilities, this invention offers significant improvements in detection accuracy, response speed, system intelligence, and data adaptability.
[0030] In the data preprocessing stage, a wavelet decomposition-based denoising mechanism is introduced, and the noise level is estimated using median absolute deviation, effectively eliminating the influence of high-frequency disturbances and providing stable input for subsequent model recognition. Through deep modeling and feature fusion of the temporal structure, the model can adapt to data patterns with abrupt changes, periodicity, and mixed fluctuations, improving the ability to identify anomalous events in non-stationary data. Experimental results show that the proposed method achieves an F1-score of 93.2% on a real dataset, which is a significant improvement over existing methods.
[0031] In terms of model structure, a one-dimensional convolutional neural network and a bidirectional long short-term memory network are used to collaboratively extract local and global temporal features. An attention mechanism is introduced to dynamically weight the contributions of different types of features, improving the model's expressive power and robustness to complex multimodal inputs. By combining reinforcement learning strategies to optimize the dynamic decision threshold, the model can adaptively adjust the decision boundary according to the real-time distribution of data, significantly reducing the false positive rate. On a typical test set, the average false positive rate decreased from 18.70% to 4.30%.
[0032] This invention introduces a closed-loop correction module after anomaly detection, employing a generative adversarial network (GAN) structure to generate missing values and replace erroneous values. The generator produces reliable data fragments based on the context before and after the anomaly segment, while the discriminator optimizes and controls the reconstruction quality. In test data with actual anomaly backgrounds, the repair accuracy reaches 98.3%, significantly reducing the workload of manual review and processing time.
[0033] This invention supports collaborative input of both structured and unstructured text data. Through feature encoding and cross-modal fusion mechanisms, it establishes a correlation between maintenance logs and equipment operating status, improving the overall system's adaptability to heterogeneous data sources. In practical applications, it supports high-concurrency data input of over 20,000 channels, with system response latency controlled within 50 milliseconds, meeting the real-time and scalability requirements of the industrial sector.
[0034] By outputting anomaly scoring probabilities, judgment labels, and anomaly type classification results, this invention also possesses result interpretation capabilities, assisting in manual review and fault tracing. Based on the feature response analysis mechanism integrated within the model, a contribution map of input data in the model can be generated to locate key indicators that generate abnormal results, thereby improving the transparency and credibility of the detection results.
[0035] This invention combines a data-driven modeling approach with a deep learning structure to achieve a high-precision, real-time, correctable, interpretable, and multi-source fusion intelligent anomaly identification process, which has good engineering feasibility and potential for expanded applications. Attached Figure Description
[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 The above is a flowchart of a method for accurately identifying data anomalies according to an embodiment of the present invention.
[0038] Figure 2 This is a schematic diagram of a dynamic threshold adjustment module for a data anomaly accurate identification method provided in one embodiment of the present invention.
[0039] Figure 3 This is a data flow diagram of the multimodal feature fusion process of a data anomaly accurate identification method provided in one embodiment of the present invention.
[0040] Figure 4 This is a structural diagram of a spatiotemporal feature extractor for a method for accurate identification of data anomalies provided in one embodiment of the present invention.
[0041] Figure 5 The diagram shows the closed-loop correction system of a data anomaly accurate identification method provided in one embodiment of the present invention.
[0042] Figure 6This is a schematic diagram of generative AI feature extraction for a method of accurately identifying data anomalies provided in an embodiment of the present invention.
[0043] Figure 7 This is a schematic diagram of a spatiotemporal feature extractor for a method for accurate identification of data anomalies provided in one embodiment of the present invention.
[0044] Figure 8 This is a schematic diagram of the network structure of a method for accurately identifying data anomalies provided in one embodiment of the present invention.
[0045] Figure 9 This is a schematic diagram illustrating the detection of abnormal data in equipment ledgers, which is part of an embodiment of the present invention for a method for accurately identifying data anomalies. Detailed Implementation
[0046] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0047] Example 1, referring to Figure 1 This is one embodiment of the present invention, which provides a method for accurate identification of data anomalies, including:
[0048] Step 1: Acquire multi-source data from the target system, including structured and unstructured data; Step 2: Preprocess the multi-source data, including denoising and missing value labeling for time-series structured data, and entity extraction and formatting for text-based unstructured data; Step 3: Input the preprocessed data into an anomaly detection model with time-series modeling and feature fusion capabilities for feature extraction and state modeling to obtain corresponding anomaly scores; Step 4: Determine a dynamic judgment threshold based on historical data statistical characteristics and the current anomaly score, compare the score with the threshold, and output the anomaly judgment result; Step 5: If the anomaly judgment result is anomaly, initiate a correction process to complete or replace the abnormal data based on the generative model to restore data continuity and credibility.
[0049] In one optional embodiment of the present invention, the structured data includes, but is not limited to, equipment operating parameters, sensor readings, load information, and environmental monitoring data; unstructured data includes text information such as maintenance logs, inspection reports, and repair records. The data acquisition module obtains structured data from SCADA systems, online monitoring devices, and distribution network automation terminals, while simultaneously calling a text processing interface to extract maintenance text data from the log system.
[0050] The data acquisition can be achieved through real-time or periodic uploading based on edge gateways, industrial Ethernet, or 5G communication, and organized according to a unified data tag format, including key identification fields such as timestamp, device ID, and data type.
[0051] In a preferred embodiment of the present invention, the structured data includes real-time collected information such as temperature, current, voltage, active power, reactive power, and equipment status in the power system, with a sampling frequency of once per minute, supporting continuous recording for 24 hours; the unstructured data comes from natural language text recorded by the dispatching platform, including equipment inspection conclusions, alarm handling procedures, anomaly analysis reports, etc.
[0052] Before entering the anomaly detection system, all data is uniformly converted into a standardized format. Structured data is tagged according to channel number, and unstructured text is purified and its metadata is supplemented using a log summary extraction tool. The system supports simultaneous access to 20,000 or more data channels and adds a unique identifier for each data stream, including the collection time, geographical location, and the device to which it belongs, to ensure spatial and temporal correlation in the subsequent modeling stage.
[0053] The beneficial effects of this preferred technical solution are as follows: through a refined data acquisition and labeling mechanism, this invention can achieve stable access to high-concurrency multi-source data, ensuring that the anomaly identification model has a comprehensive and accurate input foundation. The unified labeling design for structured and unstructured data enhances the collaborative capabilities between different data channels, enabling the system to possess good correlation and scalability in subsequent multimodal fusion and spatiotemporal modeling processes. In the Southern Power Grid's actual test environment, this solution achieved concurrent input from over 20,000 device points, with an average latency controlled within 50ms, effectively supporting real-time anomaly identification tasks in large-scale distributed power scenarios.
[0054] In step 2, wavelet decomposition is performed on the target data to extract wavelet coefficients of different frequency bands; the noise standard deviation is estimated based on the median absolute deviation of the high-frequency part of the wavelet coefficients; a threshold is calculated based on the noise standard deviation and the number of sampling points, and the wavelet coefficients are subjected to threshold screening.
[0055] In one optional embodiment of the present invention, the structured time-series data is first analyzed for missing data using a sliding window. A threshold is set to determine whether there are consecutive missing data points. When the number of consecutive missing points is greater than or equal to 3, it is marked as an abnormal segment. Subsequently, a multi-level wavelet decomposition algorithm based on Daubechies wavelets is used to transform the original data, with a decomposition level of 4 to 6 levels, and the wavelet coefficients of each level are extracted.
[0056] The median absolute deviation (MAD) of the high-frequency coefficients is used to estimate the background noise intensity in the data, and then a wavelet soft threshold is calculated for denoising. All coefficients below this threshold are processed according to a set strategy (soft threshold shrinkage), and wavelet reconstruction is performed to obtain smoothed time-series data.
[0057] For unstructured data, a pre-trained Chinese Named Entity Recognition (NER) model is used to extract core entity information such as time, equipment, faulty parts, and processing results from log texts, and output them in structured JSON format.
[0058] In a preferred embodiment of the present invention, a 6-level wavelet decomposition is used for processing structured data, with the mother wavelet function being db6 (Daubechies 6th order), which can effectively adapt to the typical short-term abrupt changes in power system data such as current, voltage, and temperature. The median MAD is calculated for the absolute value sequence of high-frequency coefficients. After obtaining an adaptive threshold, soft thresholding is performed on the wavelet coefficients of each level to complete signal reconstruction. This method can effectively filter out background disturbances and periodic high-frequency noise while retaining the signal characteristics of sudden events. In terms of missing value detection, the algorithm automatically identifies and marks minute-level sampling interruption segments, providing input boundaries for subsequent repair modules. For unstructured data, a text-summarized enhanced named entity extraction model is used, combined with industry keyword rules to improve the accuracy of entity recognition in inspection reports and logs. The extracted key fields include equipment ID, maintenance time, operation instructions, alarm event type, etc., and standardized representation is achieved through bidirectional word vector embedding, preparing for multimodal feature splicing. In addition, the wavelet decomposition uses the Daubechies 6th order wavelet with 6 decomposition levels, and the threshold calculation formula is: Where σ is the noise standard deviation and N is the number of sampling points. Based on this, the present invention optimizes and improves upon the characteristics of power system data: the Daubechies 6th order wavelet (db6) is selected to effectively adapt to typical short-term abrupt changes in power data, such as voltage drops and current surges, ensuring that key abnormal signals are preserved while denoising; the 6-level decomposition better separates high-frequency noise and low-frequency trends in power signals, improving the precision of denoising; the noise standard deviation σ is estimated using median absolute deviation (MAD), making the threshold calculation more robust to occasional large noise points that may exist in power system data, avoiding the problem of traditional standard deviation estimation being easily affected by outliers; at the same time, N is set to the number of sampling points per day (24 hours) (1440), allowing the threshold to better adapt to the periodic characteristics and sampling frequency of power system data, thereby achieving precise suppression of background noise in power data while maximizing the preservation of characteristic information of abnormal events, improving the accuracy and sensitivity of anomaly identification.
[0059] After obtaining the adaptive threshold, soft thresholding is performed on the wavelet coefficients of each layer to complete signal reconstruction. This method can effectively filter out background disturbances and periodic high-frequency noise while preserving the signal characteristics of sudden events.
[0060] In terms of missing value detection, the algorithm automatically identifies and marks minute-level sampling interruption segments, providing input boundaries for subsequent repair modules.
[0061] Unstructured data employs a text-summarized enhanced named entity extraction model, combined with industry keyword rules to improve entity recognition accuracy in inspection reports and logs. Key extracted fields include device ID, maintenance time, operation instructions, and alarm event types, and standardized representation is achieved through bidirectional word vector embedding, preparing for multimodal feature concatenation.
[0062] The beneficial effects of this preferred technical solution are that, by introducing a wavelet threshold denoising algorithm, the system can maintain the integrity of data features even in high-noise backgrounds, improving the model's ability to identify sudden anomalies. The MAD estimation method does not rely on prior training and can adapt to station signals with different noise levels, achieving stable and universal denoising processing. This mechanism reduces the mean square error before anomaly detection by approximately 47% in measured data, enhancing data quality stability.
[0063] The missing value labeling strategy boasts high temporal accuracy, effectively assisting the subsequent repair module in identifying reconstruction boundaries. The text data entity extraction module significantly improves the structuring of log information, laying a crucial foundation for cross-modal analysis and multi-source data fusion.
[0064] In the comprehensive testing environment, the preprocessing module has a processing latency of less than 10ms and significantly enhances data integrity and noise robustness, providing high-quality input assurance for subsequent model training and real-time inference.
[0065] In step 3, the anomaly detection model includes a one-dimensional convolutional neural network layer, a bidirectional long short-term memory network layer, and an attention mechanism fusion layer; the one-dimensional convolutional neural network layer is used to extract local temporal features of the input data; the bidirectional long short-term memory network layer is used to model the short-term and long-term dependencies of the time series; and the attention mechanism fusion layer is used to enhance the correlation expression ability between different feature channels.
[0066] In an optional embodiment of the present invention, the anomaly detection model may include multiple feature extraction sub-modules, such as convolutional neural networks (CNNs) for extracting local features, recurrent neural networks (such as LSTMs) for modeling time dependencies, and attention mechanisms for enhancing the expression weights between features. The model input data is a standardized and encoded multi-channel sequence tensor in the format (N×T×D), where N is the number of samples, T is the time length, and D is the feature dimension.
[0067] The model processes data using a sliding window approach, processing a time series at a time and outputting an anomaly score. The score represents the probability or severity of an anomaly in the current data segment, used to determine whether it constitutes an anomalous state.
[0068] For unstructured text information, after being transformed into a vector representation through the embedding layer, it is concatenated with structured numerical data and integrated into the network to learn the overall contextual representation.
[0069] In a preferred embodiment of the present invention, the anomaly recognition model specifically includes the following structure: the first layer is a one-dimensional convolutional neural network (1D-CNN) with a kernel size of 3 and a layer number of 2, which extracts local mutations and edge features; the second layer is a bidirectional long short-term memory network (BiLSTM) with 2 stacked layers, each containing 64 hidden units, used to capture the bidirectional dependencies of time-series data; the third layer is an attention fusion layer, which uses an additive attention mechanism to adaptively weight different feature channels to enhance the model's sensitivity to key input features.
[0070] During the training phase, the model uses a weighted cross-entropy loss function to handle class imbalance. During the inference phase, it outputs a normalized anomaly score (ranging from 0 to 1), representing the probability that the data is anomaly within that time period. This value will be used for subsequent comparison with a dynamic decision threshold.
[0071] Structured data such as temperature, current, and load sequences, along with unstructured text data that has undergone entity extraction and encoding, undergo uniform normalization processing before being input into the network. These are then concatenated to form a complete input tensor. This fusion strategy, combining power equipment operation data with maintenance records, helps the model accurately identify complex anomaly patterns associated with specific scenarios.
[0072] The beneficial effects of this preferred technical solution are as follows: the multi-layer composite neural network structure constructed in this step can effectively capture complex features such as abrupt changes, periodicity, and historical dependence in power data, while also taking into account the robustness and generalization ability of the model. Edge features are extracted through a one-dimensional convolutional network, enhancing the response capability to short-term anomalies such as lightning strikes and electric arcs; the bidirectional LSTM structure improves the modeling effect on slowly changing signals such as equipment aging trends and long-term drift; and the attention fusion mechanism realizes the dynamic allocation of weights among features, making the model pay more attention to key abnormal signals.
[0073] The multimodal data fusion mechanism supports joint modeling of structured and textual information, effectively improving the model's ability to distinguish anomaly types and its adaptability to different scenarios. Test results show that the model of this invention achieves an F1-score of 93.2% on typical anomaly datasets in the power industry, significantly outperforming traditional rule-based models and single LSTM network structures.
[0074] In step 4, the decision threshold is adaptively adjusted using a reinforcement learning algorithm. The reinforcement learning algorithm employs a dual-delay deep deterministic policy gradient algorithm and optimizes the policy based on a reward function that includes the false positive rate, recall rate, and threshold change magnitude to update the decision threshold for each cycle.
[0075] In an optional embodiment of the present invention, in this step, the system maintains a historical sequence of abnormal rating values over a past period using a sliding window approach, and determines a judgment threshold for the current time based on statistical indicators such as the mean and variance of the sequence. The abnormal rating value at the current time point is compared with the dynamic threshold; if it exceeds the threshold, it is judged as abnormal.
[0076] The length of the sliding window can be set according to data characteristics, such as 1 hour or 24 hours. The judgment result not only outputs a Boolean value indicating whether it is abnormal, but also includes the current score, threshold, and difference information for subsequent modules to use.
[0077] In a preferred embodiment of the present invention, to improve the adaptability and robustness of threshold determination, a reinforcement learning algorithm is introduced to dynamically optimize the threshold. Specifically, a dual-delay deep deterministic policy gradient algorithm (TD3) is used to construct the policy network, with the optimization objectives of minimizing the false positive rate and maximizing the recall rate.
[0078] Within each decision cycle (e.g., every 5 minutes), the system inputs the judgment performance indicators of the previous cycle (including true labels, false positive rate, and false negative rate) into the policy network as environmental feedback, and outputs the optimal judgment threshold θt for the current stage.
[0079] The final judgment logic is as follows: if the current score value p_t ≥ θ_t, it is judged as abnormal; otherwise, it is normal. The judgment result is generated into a structured output in the system, including the judgment time, the degree of abnormality, the abnormality trend, and the score-threshold difference value, for subsequent modules to perform correction or alarm operations.
[0080] The beneficial effect of this preferred technical solution is that this step realizes the transformation from static threshold judgment to adaptive dynamic judgment mechanism, overcoming the drawbacks of traditional fixed threshold method which is prone to high false alarms or missed detections when facing data fluctuations and distribution drift.
[0081] By introducing a reinforcement learning strategy, the system can automatically adjust its decision-making strategy based on real-time feedback, enabling the model to maintain high detection stability under various load conditions, climate changes, and equipment operating states. The reward function design combines false alarm control, recall optimization, and threshold smoothing to ensure robust adjustment and fast convergence.
[0082] Test results show that after adopting this dynamic threshold mechanism, the average false alarm rate of the anomaly identification system decreased from 22.3% to 4.3%, which greatly improved the reliability and usability of the detection results and provided a more reliable basis for subsequent intelligent repair.
[0083] In step 5, the correction process uses a generative adversarial network (GAN) to complete or replace anomalous data. The GAN includes a generator and a discriminator. The generator generates missing values or replaces anomalous values based on the context information of the time series in which the anomalous data is located. The discriminator is used to determine the distribution difference between the generated data and the real data.
[0084] In step 5, the anomaly determination result includes an anomaly type label; the anomaly type label is used to identify any one of temperature anomaly, current anomaly, load anomaly, and data missing.
[0085] In step 5, structured and unstructured data undergo feature fusion before being input into the anomaly detection model; the unstructured data includes operation and maintenance logs or inspection reports; the feature fusion involves extracting and encoding text entities from the unstructured data and then concatenating them with the structured data.
[0086] Furthermore, a cross-modal feature mapping mechanism is used to establish the correlation between different data sources.
[0087] In an optional embodiment of the present invention, after an anomaly is determined, the system immediately locates the corresponding time period or data segment and calls the correction module to process that portion of the data. The processing methods include:
[0088] For cases of missing data, interpolation, moving average, or simple time series prediction algorithms are used to fill in the missing data based on the sequence trends of adjacent data points. For outliers detected as mutations, replacements are made based on the contextual trends, such as using an ARIMA model to generate replacement values. The repaired results will be used as temporary data caches and will be processed together with subsequent data in subsequent business modules.
[0089] The corrected data will be marked as alternative or repaired data, and will include a confidence interval or prediction error estimate to ensure safe use in the future.
[0090] In a preferred embodiment of the present invention, the present invention preferably employs a generative adversarial network based on Wasserstein distance constraints (WassersteinGAN with Gradient Penalty, WGAN-GP) to achieve high-precision completion or replacement of anomalous data.
[0091] The generative model consists of a generator and a discriminator:
[0092] The generator takes contextual data before and after the abnormal time period as input and learns to generate complete segments that match the contextual trend. The discriminator is used to judge the difference between the generated result and the real data in terms of feature distribution. It adopts the WGAN-GP structure and improves training stability by introducing a gradient penalty term. The penalty coefficient λ is set to 10. During network training, an alternating optimization strategy is adopted, and the generator and discriminator are updated with learning rates of 1e-4 and 3e-5, respectively.
[0093] The generation process is not only used to complete missing data, but also to replace data points that have been identified as abnormal but are not missing, ensuring that the generated results are highly consistent with the actual data in terms of amplitude, fluctuation frequency and periodic trend.
[0094] The beneficial effects of this preferred technical solution are that, by introducing the WGAN-GP network with context modeling and generation capabilities, the present invention not only achieves automatic repair of abnormal data, but also significantly improves the authenticity and stability of the repaired data.
[0095] Actual test results show that, in the scenario of a lightning strike incident in the Southern Power Grid in 2024, the repair module achieved a 98.3% accuracy rate in correcting high-frequency abrupt anomaly data, an improvement of over 15 percentage points compared to the traditional LSTM interpolation method. Simultaneously, a discriminator is used to verify the distribution of the generated data, ensuring that the repair results are statistically insignificantly different from normal data (χ²). 2 (Pass rate > 97%)
[0096] In addition, the repair module significantly reduces the burden of manual review. Test results show that the proportion of manual review time decreased from 42% to 9.8%, which greatly improves the efficiency of system operation and maintenance and provides key technical support for achieving true closed-loop anomaly management.
[0097] In an optional embodiment of the present invention, while outputting the anomaly determination result, the system generates corresponding anomaly type labels based on the classification output module of the anomaly identification model. The labels can be represented using a multi-category numbering method, such as: 0 for normal, 1 for temperature anomaly, 2 for current anomaly, 3 for load anomaly, 4 for missing data, etc. The determination labels will be recorded together with the anomaly score and used for subsequent order dispatching, early warning grading, or statistical attribution.
[0098] In a preferred embodiment of the present invention, the anomaly type label is generated based on the probability distribution result output by the last layer of the softmax classifier in the model. The label determination not only depends on the current score value, but also incorporates contextual information after multimodal feature aggregation for auxiliary judgment. For anomalies whose type cannot be completely determined, a parallel label candidate set can be output, and a confidence score can be provided for each item.
[0099] The anomaly type label is also linked to the repair strategy; the system automatically calls the corresponding repair method module based on different labels. For example, for temperature anomalies, a generation strategy based on physical curve constraints is prioritized; if data is missing, a context-based adversarial generative network is invoked.
[0100] The beneficial effect of this preferred technical solution is that by introducing an anomaly type label output mechanism, the system can not only determine whether it is abnormal, but also achieve a higher level of intelligent understanding of what type of anomaly it is, which greatly enhances the interpretability and practicality of the anomaly identification system.
[0101] During on-site operation and maintenance, the label-based anomaly identification results can be directly used for work order generation, inspection scheduling, and anomaly root cause localization, replacing manual analysis processes. Actual test data shows that in typical power equipment anomaly samples, the anomaly classification accuracy reaches 91.5%, significantly outperforming traditional coarse classification methods that rely on threshold logic.
[0102] In an optional embodiment of the present invention, to improve recognition accuracy, the system performs feature fusion processing on structured data (such as temperature, current, and operating status) and unstructured data (such as maintenance logs and inspection reports) before inputting the data into the model. Unstructured text data is first processed by Named Entity Recognition (NER) and syntactic analysis to extract key entity words, then converted into vector representations. These vector representations are then aligned with the structured data through concatenation to form a unified input tensor.
[0103] In a preferred embodiment of the present invention, a cross-modal feature mapping mechanism is employed to construct a high-order association between structured and text-based unstructured data. The specific method includes:
[0104] Textual data is first processed by a BERT-like language model to extract embedded features, and then dimensionality reduction and linear projection mapping are performed to map it to the same dimension as the structured data. Subsequently, it is fused through residual connections and attention mechanisms, and the model can learn the collaborative expression patterns between different data sources during training. After structural fusion, it is input into a unified model structure to achieve unified processing of heterogeneous data and anomaly decision-making.
[0105] The beneficial effects of this preferred technical solution are that the model's recognition capability is significantly enhanced through the effective fusion of multi-source heterogeneous data. Structured data provides direct measurement information, while unstructured text carries operational context and historical experience information; the collaboration of the two enables more complete state modeling.
[0106] In actual testing, the fusion model improved the accuracy of identifying complex anomalies (such as load anomalies accompanied by alarm log prompts) to 92.7%, significantly improved the ability to identify false alarms, and provided a solid data foundation and generalization ability for the large-scale deployment of the system in complex industrial scenarios.
[0107] Example 2, refer to Figures 2-9 As one embodiment of the present invention, based on the previous embodiment, a method for accurate identification of data anomalies is provided, including:
[0108] This invention addresses the characteristics of high noise, abrupt changes, and multi-source time-series data in the power industry by employing the following composite AI structure, balancing accurate recognition capabilities with practical deployment effectiveness:
[0109] Input module: Receives multi-source structured data (such as temperature, current, load), semi-structured data such as alarm logs (converted to structured format before input), and spatiotemporal tags (such as device ID, collection time, location).
[0110] Data preprocessing module: Built-in wavelet denoising (6-level decomposition, automatic threshold selection), time series completion, anomaly marking and other functions.
[0111] Feature extraction layer: Convolutional Neural Network (1D-CNN): The first layer extracts local features from the original time-series signal. The kernel size is 3 and the number of layers is 2.
[0112] Bidirectional LSTM (BiLSTM) layer: captures short-term and long-term trends of signals, improves the ability to detect sudden anomalies (such as lightning strikes and power outages), hides 64 units, and stacks 2 layers.
[0113] Attention fusion layer: Adaptively strengthens the feature weights of multiple data types (environment, equipment, operation and maintenance) to improve multimodal correlation.
[0114] Anomaly detection output layer: Fully connected layer to Sigmoid / Softmax output, binary classification (anomaly / normal) or multi-class classification (anomaly type).
[0115] Model Connections: Original Input → Wavelet Denoising → 1D-CNN → BiLSTM → Attention → Fully Connected Layer → Output
[0116] Key parameters: Convolutional layer kernel number 32 / 64, kernel size 3, ReLU activation; LSTM unit number 64, layer 2; Data in each channel is normalized and feature selected (removing continuously zero-variable data columns); Dropout = 0.2 to prevent overfitting; The output layer takes Sigmoid as an example, and the output is the anomaly probability p∈[0,1].
[0117] We selected multi-source data from the past three years of actual operation and maintenance, equipment operation, and monitoring of China Southern Power Grid, covering 30 types of anomalies, including general operating conditions, special weather conditions, and typical anomalies. We also used manually reviewed / expert-judged and labeled data to significantly expand the coverage of scenarios unique to the power operation and maintenance industry (such as high photovoltaic penetration sites and severe weather).
[0118] Sample balancing: Oversampling with SMOTE or GAN-generated samples is used to augment rare anomaly types to improve the model's ability to identify rare anomalies; Training / validation / testing partitioning: The 8 / 1 / 1 principle is followed to ensure diverse scenarios, completeness, and generalization performance; Optimizer and learning rate adjustment: The Adam optimizer is used with an initial learning rate of 0.001, dynamically adjusted (ReduceLROnPlateau); Loss function: Weighted cross-entropy is used for rare anomalies (class weights are linked to actual industry accident risks); Early stopping prevents overfitting.
[0119] Noise suppression strategy: The wavelet denoising parameter σ and the decomposition layer L are optimized for power sampling time series to solve the problem of missed anomaly detection under high noise conditions; Anomaly level and maintenance suggestion output: The model output not only includes the anomaly type, but also automatically generates customized processing suggestions (such as automatic reset if manual maintenance is required); Industry-specific label fusion: Such as equipment type, city / region weather, dispatch level, etc., are used as auxiliary input features to effectively improve the ability to identify regional / operating condition specificity.
[0120] (1) Input data: Multi-channel splicing: temperature sequence, current sequence, operating status, equipment type (one input with multiple branches); maximum sampling frequency: 1 point per minute, the historical window length can be flexibly set (e.g., 24 hours equals 1440 points); the operation and maintenance log text content is input after entity extraction and encoding.
[0121] (2) Output data: One-dimensional probability sequence p(t)∈[0,1], threshold above 0.5 is judged as abnormal, and the optimal F1 split point can be adjusted; classification label (abnormal type number, such as 0-normal, 1-high temperature, 2-current abnormal, etc.); real-time abnormal list and key section distribution information, used for on-site alarm and operation and maintenance dispatch.
[0122] A standard system for anomaly types specifically designed for the power industry, encompassing 33 major categories of common and extreme anomalies, both on-site and historical. A multimodal modeling approach combining wavelet analysis, deep temporal data, and attention-based methods addresses challenges posed by sudden equipment changes and strong background noise. The integration of regional and meteorological multi-label information significantly reduces false alarms and missed alarms in specific areas (e.g., seasonal impacts from typhoons / rain / snow). Automatic anomaly cascading handling suggestions are generated to support automation in subsequent work order dispatching, shift maintenance scheduling, and other scenarios.
[0123] Adapting to new scenarios with high noise and high probability of sudden anomalies (such as short-term distortion of equipment in batches under the influence of extreme weather), the industry's false detection rate has decreased by 60% and the recall rate has increased by about 30%.
[0124] The model training and inference speeds are superior to traditional LSTM+rule combination schemes, meeting the real-time requirements of practical deployment.
[0125] The model structure and parameters are adjustable and easily transferable, adapting to various power application scenarios such as smart substations and new energy power plants.
[0126] For equipment temperature sampling data, determine whether there are 3 or more consecutive missing points (assuming a sampling interval of 1 minute): if there are 3 or more consecutive missing points in a certain time period, then an anomaly marker is activated for that point.
[0127] An improved wavelet thresholding denoising algorithm is used, and its core formula is as follows:
[0128]
[0129] Where L=6 represents the number of decomposition levels, indicating that the signal is decomposed into 6 levels of wavelet coefficients; T represents the threshold used for denoising; σ represents the standard deviation of noise (which can be estimated from the median of the high-frequency wavelet coefficients); N=1440 represents the number of sampling points in one day (24 hours), i.e., N=24×60; 3. Meaning of the characters: T represents the threshold used in the wavelet thresholding denoising process. All wavelet coefficients with an absolute value less than T are considered noise and will be set to zero or soft-thresholded; σ represents the standard deviation of noise in the original signal. In actual algorithms, it is generally estimated by the median absolute deviation (MAD) of the high-frequency wavelet coefficients; N represents the total number of sampling points, which actually refers to the total number of sampling points in one day. In this case, it is 14401440; L represents the number of wavelet decomposition levels. Setting it higher can capture more fine-grained information. Here, 66 levels are selected.
[0130] Step 2: Generative AI Feature Extraction, refer to Figure 6 Based on the TD3 reinforcement learning algorithm, the reward function is designed as follows:
[0131] R = α·(1-FPR) + β·TPR - γ|θ t -θ t-1
[0132] Where R represents the reward value, used to evaluate the effectiveness of the current threshold adjustment strategy; α represents the weighting coefficient, with a value of 0.6, indicating the emphasis on reducing false alarms; FPR represents the false alarm rate, i.e., the proportion of actually normal samples that are judged as abnormal; 1-FPR represents the proportion of normal samples that are correctly identified; β represents the weighting coefficient, with a value of 0.4, indicating the emphasis on improving recall; TPR represents the recall rate, i.e., the proportion of actually abnormal samples that are judged as abnormal; γ represents the weighting coefficient, with a value of 0.1, used to penalize excessive fluctuations in the threshold; θt represents the judgment threshold at the current moment; and θt-1 represents the judgment threshold at the previous moment. By adjusting the weights of α, β, and γ, false alarms, recall, and threshold stability can be balanced, making it more in line with the actual needs of power systems for anomaly detection.
[0133] Input: Preprocessed multidimensional time series data X∈R^(N×T×D) (N = number of devices, T = time steps, D = feature dimension).
[0134] Reference Figure 7 Parameter settings: temporal convolution kernel size = 3, dilation factor = 2^layer; number of graph attention heads = 3, LeakyReLU slope = 0.2.
[0135] The scoring formula is expressed as follows:
[0136]
[0137] Where S(t) represents the score at time t. This score is typically used to measure the degree of anomalousness or deviation from normal behavior of the system at time t. The higher the value, the greater the likelihood of anomaly; λ1 represents the weighting coefficient of the first term. As shown in the image, λ1 = 0.6. It determines the importance of the reconstruction error in the total score. This coefficient is usually determined through model training, cross-validation, or domain knowledge, aiming to balance the influence of the two components on the final score; This represents the reconstructed or predicted data at time t. This is typically the reconstructed or predicted result of the original data xt after processing by some model (e.g., autoencoder, prediction model, etc.). Ideally, if the system behaves normally... Should be with x t Very close; x tThis represents the raw input data or actual observation data at time t. It signifies the true state or characteristics of the system at time t. The square norm (usually L2 norm) represents the reconstruction error. It measures the reconstructed data. Compared with the original data x t The greater the difference, the worse the model's ability to reconstruct the current data, which may indicate that the current data is abnormal. This represents the weighting coefficient for the second term. As shown in the image, λ2 = 0.4. It determines the importance of the distribution variance (KL divergence) in the overall score. Similar to λ1, this coefficient also needs to be determined through tuning; KL(p t ||p hist ) represents the Kullback-Leibler (KL) divergence, used to measure the current feature distribution p t The difference between the KL divergence and the historical characteristic distribution phisp. The KL divergence is asymmetric, expressed by p. hist To approximate p t The amount of information lost during the process. If p t The greater the difference, the larger the KL divergence value, indicating a greater deviation between the current system behavior distribution and the historical normal behavior distribution, and a higher probability of anomaly; p t This represents the current feature distribution. This is typically a probability distribution obtained through statistical analysis or modeling of feature data near time t (e.g., a recent period). It reflects the system's behavioral pattern at the current moment; p hist This represents the historical characteristic distribution. It is typically a probability distribution obtained through statistical analysis or modeling of historical normal behavior data (e.g., data with a 7-day sliding window). It represents the system's behavior pattern under normal operating conditions; λ1 = 0.6 and λ2 = 0.4 are given values. In practical applications, determining these weighting coefficients is usually a hyperparameter tuning process aimed at finding the optimal combination for model performance. Here are some common methods for solving or determining these coefficients:
[0138] One method of determining coefficients involves empirical setting: these coefficients are directly set based on domain knowledge or past experience. For example, if reconstruction error is considered more important than distributional differences, a larger value can be assigned to λ1.
[0139] Another method for determining the coefficients is through a grid search: define the possible ranges and step sizes for λ1 and λ2 (e.g., λ1∈[0,1], step size 0.1; λ2∈[0,1], step size 0.1, and λ1+λ2==1). Then, iterate through all possible combinations, training and evaluating the model for each combination. Select the combination that performs best on the validation set.
[0140] Another method for determining coefficients is through random search: similar to grid search, but instead of traversing all combinations, it randomly samples parameter combinations within a defined range. In some cases, random search is more efficient than grid search, especially when the parameter space has high dimensionality.
[0141] Another way to determine the coefficients is through Bayesian optimization: a smarter optimization method that uses a probabilistic model to guide the search process, thus finding the optimal parameter combination more efficiently. Bayesian optimization selects the next most promising parameter combination to evaluate based on previous evaluation results.
[0142] Another way to determine the coefficients is through cross-validation: In the search methods described above, cross-validation is usually used to evaluate the performance of each set of parameters to ensure that the model performs well on unseen data and avoids overfitting.
[0143] In another way of determining the coefficients, genetic algorithms are used: optimization algorithms that simulate the process of natural selection can be used to search a complex parameter space to find the optimal or near-optimal combination of parameters.
[0144] The final choice of λ1 and λ2 should be closely aligned with specific business objectives. For example, if the business is more focused on anomaly recall (finding as many anomalies as possible), the coefficients may be adjusted to make the model more sensitive to anomalies; if the focus is more on precision (reducing false positives), a stricter threshold and a more balanced set of coefficients may be needed.
[0145] The adaptive threshold update rule is expressed as follows:
[0146] θ t+1 =α·θ t +(1-α)·(μ t +3σ t )
[0147] Where θt+1 represents the update threshold at time t+1. This is the critical value used to determine whether the rating S(t) is abnormal; θt represents the current threshold at time t; α represents the smoothing factor, which, as shown in the image, is 0.85. It determines the weight of the old and new thresholds in the update; μt represents the average rating over the past hour. This represents the average rating level of the system in the recent period; σt represents the standard deviation of the rating over the past hour. This represents the degree of rating fluctuation of the system in the recent period; the smoothing factor α = 0.85.
[0148] Reference Figure 8 The training strategy is as follows: generator learning rate = 1e-4, discriminator learning rate = 3e-5; gradient penalty coefficient λ = 10 (WGAN-GP framework).
[0149] Input data: Temperature sensor data for a 220kV transformer (sampling frequency 1Hz)
[0150] Anomaly detection: When 10 consecutive points exceed the dynamic threshold (θ = μ + 3σ, μ is calculated by the sliding window, and the window length is 60 seconds).
[0151] The characters (θt+1, θt, α, μt, σt) of the adaptive threshold update rule are compared with the previously discussed scoring formula S(t) and total loss function. The characters in the symbol may overlap to some extent (for example, they may all use Greek letters), but they represent different meanings and functions.
[0152] μt and σt: Although they do not appear directly in the scoring formula, they are the results of statistical calculations based on the score values generated by the scoring formula S(t). Therefore, there is an indirect, logical connection between them, but they are not direct characters in the same formula.
[0153] θt+1, θt, α: These characters are unique to the adaptive threshold update rule and are used to dynamically adjust the threshold for anomaly detection. They are not directly related to the meaning of the previous formula.
[0154] All the characters in the formula serve different purposes:
[0155] The scoring formula S(t) is used to calculate the anomaly score at a certain moment.
[0156] Adaptive threshold update rule: Used to dynamically adjust the threshold for judging anomalies.
[0157] Total loss function Used to train models so that they can better perform generation or prediction tasks.
[0158] Therefore, although they may be similar in symbolism, their roles in their respective formulas and the physical meanings they represent are independent.
[0159] Correction process: For outlier segments, LSTM-GAN is used to generate alternative sequences. The loss function includes:
[0160]
[0161] in, This represents the total loss function. In machine learning and deep learning, a loss function measures the difference between a model's predictions and the true values. The goal of model training is to minimize this loss function, enabling the model to better fit the data and make accurate predictions. This total loss function is a weighted sum of multiple sub-loss terms designed to constrain the model's behavior from different perspectives. This represents the Mean Squared Error (MSE) loss. It is typically used to measure the difference between continuous predicted values and true values. It is calculated as the mean of the squared differences between the predicted and true values. In image generation or data reconstruction tasks, MSE loss can help generate results that are as close as possible to the original data at the pixel or numerical level. This represents the Kullback-Leibler (KL) divergence loss. KL divergence measures the difference between two probability distributions. In generative models such as variational autoencoders (VAEs), KL divergence loss is often used to constrain the distribution of latent variables learned by the model to approximate a certain prior distribution (e.g., a standard normal distribution), thereby ensuring the diversity and quality of generated samples. represents the smoothing loss. This loss term is typically used to encourage the model to generate smooth outputs, avoiding sharp, discontinuous, or noisy results. For example, in image generation, smoothing loss can lead to images with better visual coherence; in time series prediction, it can make the prediction curve smoother. 0.7, 0.2, and 0.1 represent the weighting coefficients of each loss term. They determine the importance of each sub-loss term in the total loss function. By adjusting these weights, the model's emphasis can be balanced between different optimization objectives, such as prioritizing reconstruction accuracy (MSE), latent variable distribution (KL), or output smoothness. The MAE of the generated result relative to the original data should be ≤0.5°C; MAE stands for Mean Absolute Error. It measures the average absolute difference between the predicted and true values. Compared to MSE, MAE is less sensitive to outliers because it does not square the error. Here, MAE is used as a performance metric or constraint indicating that the mean absolute error between the generated result (which could be a predicted temperature value or other continuous data) and the original data must be less than or equal to 0.5 degrees Celsius. This is a hard requirement or evaluation standard for model performance.
[0162] Reference Figure 9 Scenario Description: A user's daily electricity consumption suddenly drops by 90% (normal baseline: 25±5kWh, measured value: 2.3kWh); Correlation Analysis: Based on the user's repair record text (meter replacement), a corrected value of 22.8kWh is generated; Performance Indicators: The corrected data is expressed as a χ² value. 2 The test (p>0.05) shows that it conforms to a normal distribution; structured parameter analysis: the impedance voltage field value was found to deviate from the average value of similar equipment (μ=13.2±0.5); text association analysis: the parameter correction rule was triggered by the event of replacing the cooler filter in the operation and maintenance log.
[0163] Dynamic Correction: The impedance voltage is automatically corrected to 13.1% (92.3% confidence level); Scenario Description: A certain industrial user's monthly electricity consumption suddenly increases by 300% (baseline: 1200±150MWh, measured value: 4892MWh); Equipment Correlation Analysis: 3 new electric arc furnaces (total power 85MW) are added under this user's name; Industry Comparison: The average monthly electricity consumption of metal processing enterprises of the same scale is 3800±600MWh; Correction Decision: Determined as normal data (98.1% confidence level), the user's electricity consumption baseline is updated to 4100MWh.
[0164] like Figure 2As shown in the figure, the key components of the dynamic threshold adjustment module are: State input: including historical data features and environmental state; TD3 reinforcement learning: optimizing the anomaly detection threshold through reinforcement learning algorithms (such as TD3); Dynamic threshold output: dynamically adjusting the detection threshold based on real-time data; Reward function: a reward mechanism designed based on detection accuracy, false alarm rate, and business rules; Business rule constraints: ensuring that the threshold adjustment conforms to business logic and actual needs.
[0165] like Figure 2 As shown, the input data includes text data, structured data, and time series data; feature extraction involves extracting features from each modality; cross-modal fusion uses attention mechanisms and graph embedding techniques to fuse features from different modalities; and the output fused features generate a comprehensive feature vector suitable for anomaly detection.
[0166] like Figure 4 As shown in the figure, the structure of the spatiotemporal feature extractor is as follows: Time series data: input multidimensional time series data; Temporal convolutional layer: extracts temporal features; Graph attention network: combines the device topology graph to extract spatial features; Feature concatenation layer: concatenates temporal and spatial features; Spatiotemporal feature output: outputs a comprehensive spatiotemporal feature vector.
[0167] like Figure 5 As shown in the figure, the workflow of the closed-loop correction system is as follows: Anomaly detection: Identify outliers in the data; Anomaly type classification: Determine the anomaly type (such as missing values or erroneous values); Generate correction candidate set: Generate correction values through a generative model; Business rule verification: Verify whether the correction values conform to business logic; Credibility assessment: Evaluate the credibility of the correction values; Data update or manual review: Correction values with high credibility are directly updated to update the data, while those with low credibility are reviewed manually.
[0168] Example 3 is an embodiment of the present invention, which provides a method for accurate identification of data anomalies. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0169] Refer to Table 1
[0170] Table 1 Comparison of Experimental Analysis
[0171] index Traditional methods This patent Increase Detection response time (ms) 120-250 38-52 67%↑ Multimodal data compatibility Structured data Structured + Text + Image New capabilities False Alarm Rate (FPR) 18.70% 4.30% 77%↓ Percentage of manual review time 42% 9.80% 76.7%↓
[0172] On the Southern Power Grid test dataset, the F1-score reached 93.2% (an improvement of 8.7pp compared to the literature method); it supports concurrent processing of 20,000+ sensors with an average latency of ≤50ms; it displays the basis for anomaly judgment through a feature contribution heatmap (Grad-CAM technology); in the 2024 lightning strike accident data, the anomaly correction accuracy rate was 98.3%, reducing manual review time by 76%.
[0173] Example 4 is an embodiment of the present invention, which provides a data anomaly accurate identification system, including:
[0174] The data acquisition module is used to acquire multi-source data from the target system, including structured data and unstructured data.
[0175] The preprocessing module is used to preprocess the multi-source data. The preprocessing includes performing noise reduction and missing value marking on time-series structured data; and performing entity extraction and formatting on text-based unstructured data.
[0176] The feature fusion module is used to input the preprocessed data into an anomaly recognition model with time series modeling and feature fusion capabilities, perform feature extraction and state modeling, and obtain the corresponding anomaly score value.
[0177] The anomaly assessment module is used to determine a dynamic judgment threshold based on the statistical characteristics of historical data and the current anomaly score, and compare the score with the threshold to output the anomaly judgment result.
[0178] The anomaly correction module is used to initiate a correction process when the anomaly determination result is abnormal, and to complete or replace the abnormal data based on the generation model in order to restore data continuity and credibility.
[0179] This embodiment also provides an electronic device applicable to a method for accurate identification of data anomalies, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for accurate identification of data anomalies as proposed in the above embodiment.
[0180] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements a method for accurate identification of data anomalies as proposed in the above embodiments.
[0181] The storage medium proposed in this embodiment and the method for accurately identifying data anomalies proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0182] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0183] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for accurately identifying data anomalies, characterized in that: include, Acquire multi-source data from the target system, including structured and unstructured data; The multi-source data is preprocessed, including denoising and missing value labeling of the time-series structured data. Perform entity extraction and formatting on unstructured text data; The preprocessed data is input into an anomaly recognition model with time series modeling and feature fusion capabilities to perform feature extraction and state modeling, and obtain the corresponding anomaly score. Based on the statistical characteristics of historical data and combined with the current anomaly score, a dynamic judgment threshold is determined, and the score is compared with the threshold to output the anomaly judgment result. If the anomaly determination result is abnormal, a correction process is initiated to complete or replace the abnormal data based on the generation model in order to restore data continuity and credibility.
2. The method for accurate identification of data anomalies as described in claim 1, characterized in that: The preprocessing includes performing wavelet decomposition on the target data and extracting wavelet coefficients for different frequency bands. The noise standard deviation is estimated based on the median absolute deviation of the high-frequency components of the wavelet coefficients. The threshold is calculated based on the noise standard deviation and the number of sampling points, and the wavelet coefficients are then subjected to threshold filtering.
3. The method for accurate identification of data anomalies as described in claim 2, characterized in that: The anomaly detection model includes a multi-level composite neural network structure, specifically including a feature extraction layer, a temporal modeling layer, a feature fusion layer, and an output layer.
4. The method for accurate identification of data anomalies as described in claim 3, characterized in that: The decision threshold is adaptively adjusted using a reinforcement learning algorithm; The reinforcement learning algorithm employs a dual-delay deep deterministic policy gradient algorithm and optimizes the policy based on a reward function that includes false positive rate, recall rate, and threshold change magnitude to update the decision threshold for each cycle. The correction process is based on generative adversarial networks to complete or replace anomalous data. The generative adversarial network includes a generator and a discriminator. The generator generates missing values or substitute outliers based on the context information of the time series in which the outlier data is located. The discriminator is used to determine the distribution differences between generated data and real data.
5. The method for accurate identification of data anomalies as described in claim 4, characterized in that: The anomaly determination result includes an anomaly type label; the anomaly type label is used to identify any one of temperature anomaly, current anomaly, load anomaly, and data missing. The structured and unstructured data are fused together before being input into the anomaly detection model; The unstructured data includes operation and maintenance logs or inspection reports; The feature fusion involves extracting and encoding text entities from unstructured data, and then concatenating them with structured data. Furthermore, a cross-modal feature mapping mechanism is used to establish the correlation between different data sources.
6. The method for accurate identification of data anomalies as described in claim 5, characterized in that: The wavelet decomposition uses the Daubechies 6th order wavelet with a decomposition level of 6.
7. The method for accurate identification of data anomalies as described in claim 6, characterized in that: The feature extraction layer includes a one-dimensional convolutional neural network, comprising a first convolutional layer with 32 convolutional kernels, each with a kernel size of 3 and a stride of 1, and an activation function of ReLU, and a second convolutional layer with 64 convolutional kernels, each with a kernel size of 3 and a stride of 1, and an activation function of ReLU. Each convolutional layer is followed by a batch normalization layer and a Dropout layer, and a residual connection is added between the two convolutional layers. The temporal modeling layer includes a bidirectional long short-term memory network, comprising a first BiLSTM layer with 64 hidden units and an output dimension of 128, and a second BiLSTM layer with 64 hidden units and an output dimension of 128. Each BiLSTM layer is followed by a Dropout layer. The feature fusion layer includes a multi-head self-attention mechanism, specifically including 8 attention heads, each with a dimension of 16. It generates a query, key, and value matrix through linear transformation, calculates attention weights, and outputs a weighted result. The output layer adopts a fully connected layer structure, including two hidden layers with dimensions of 128 and 64 respectively, and finally outputs the anomaly score value through the Sigmoid activation function.
8. A data anomaly accurate identification system, employing the data anomaly accurate identification method as described in any one of claims 1 to 7, characterized in that, include: The data acquisition module is used to acquire multi-source data from the target system, including structured data and unstructured data. The preprocessing module is used to preprocess the multi-source data, and the preprocessing includes performing noise reduction and missing value marking on the time-series structured data; Perform entity extraction and formatting on unstructured text data; The feature fusion module is used to input the preprocessed data into an anomaly recognition model with time series modeling and feature fusion capabilities, perform feature extraction and state modeling, and obtain the corresponding anomaly score value. The anomaly assessment module is used to determine a dynamic judgment threshold based on the statistical characteristics of historical data and the current anomaly score, and compare the score with the threshold to output the anomaly judgment result. The anomaly correction module is used to initiate a correction process when the anomaly determination result is abnormal, and to complete or replace the abnormal data based on the generation model in order to restore data continuity and credibility.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the data anomaly accurate identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the data anomaly accurate identification method according to any one of claims 1 to 7.
Citation Information
Cited By
Play switching method and device under server exception, terminal and medium
CN121531188A
Big data system operation monitoring method and device based on AI algorithm
CN121750510A