Power consumption data anomaly detection method and device based on channel independence and global-local channel dependence
By employing a method based on channel independence and global-local channel dependency, and utilizing the Mamba model and trend-residual decomposition technique, the complexity of channel correlation in multidimensional time series data is addressed, enabling efficient detection of anomalies in electricity consumption data and improving the accuracy and robustness of anomaly detection.
Patent Information
- Application Number
- CN202510959615.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-11
AI Technical Summary
Existing methods for detecting anomalies in electricity consumption data are unable to effectively handle the strong, weak, or no correlation between channels in multidimensional time-series data, resulting in insufficient anomaly detection performance, especially in complex electricity consumption scenarios where accurate anomaly identification is difficult to achieve.
We employ a method based on channel independence and global-local channel dependence, using the Mamba model for dual-channel variable processing. We combine trend-residual decomposition to construct a twin-assisted time series and build a learnable control parameter vector to flexibly adjust the relationship between channel independence and channel dependence components, thereby improving the anomaly detection performance of the model.
It improves the accuracy and robustness of anomaly detection in multidimensional time-series data, effectively identifies anomalies in data items such as power and energy readings, adapts to different anomaly types, and improves the recall and accuracy of anomaly detection.
Smart Images

Figure CN120910501A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric energy metering, and more particularly, to a power consumption data anomaly detection method and device based on channel independence and global-local channel dependence. BACKGROUND
[0002] High-quality data is the basis for the development of data elements. With data elements and data assetization becoming the top strategy for national economic development, power grid companies are facing huge opportunities for massive data assetization, but also higher regulatory requirements for data quality and reliability from the state.
[0003] As the "barometer" and "weather vane" of national economic development, accurate and reliable data can truly reflect the development of the country and society, and provide accurate basis for power supply and user service of State Grid Corporation. With the deepening of power market reform, a large number of new entities such as power selling companies, load aggregators, and virtual power plant operators have emerged. The demand for electricity is influenced by the interaction of various factors such as actual demand, electricity price policy, market rule game, distributed generation, and energy storage optimization coordination. This change has made the characteristics of electricity demand data unprecedentedly complex. Facing the characteristics of large volume, diverse structure, and complex characteristics of electricity demand data, the complexity and difficulty of electricity demand data analysis have increased, and higher requirements have been put forward for the quality of power source data.
[0004] Currently, the electricity information collection system has accessed all smart electric energy meters within the company's business scope, realizing the collection of daily, hourly, and other different time dimension electricity and power data. As the only source of company user electricity data, it has accumulated massive power user data resources, laid a solid foundation for professional applications such as user energy demand analysis and power grid operation state monitoring, and national power data analysis. However, due to objective factors such as data volume and collection conditions, the identification of massive data anomalies mostly relies on expert experience threshold judgment rules, and lacks intelligent identification tools with strong robustness and high sensitivity. Therefore, it is urgent to design and develop a robust and high-precision source collection data anomaly identification model to effectively identify data item anomalies such as electric energy indication value, power, voltage, and current, ensure that electricity data can be efficiently and conveniently used for analysis, decision-making, and business applications, and better serve government economic operation regulation, new power system construction, and company professional business development.
[0005] The research of multi-dimensional time series anomaly detection in electricity consumption scenarios highly depends on the characteristics of device operation and the distribution of data. Due to the strong robustness and stability of the electricity collection system, the data collected by the system is usually composed of a large amount of normal data. Therefore, the current mainstream methods mainly adopt two types of learning paradigms: one is a semi-supervised learning framework that only uses normal samples for training, and the other is an unsupervised learning framework that assumes that the training set is mainly composed of normal data. Most multi-dimensional time series anomaly detection methods usually calculate an anomaly score for each time point, and then compare the score with a certain threshold. In recent work on semi-supervised or unsupervised multi-dimensional time series anomaly detection, deep learning-based methods have achieved the best results in the authoritative and public multi-dimensional time series dataset benchmark test.
[0006] Deep learning methods can be mainly divided into two categories, including prediction-based methods and reconstruction-based methods. The prediction method uses past information to predict future values in the time series, and the prediction error is used as an indicator for anomaly detection. The reconstruction method uses an autoencoder or a generative model to encode the entire time series into a latent space, and infers the anomaly label according to the reconstruction error between the original data and the reconstructed data. However, multi-dimensional time series data usually has complex pattern changes, and future time series values may exhibit high uncertainty, which makes it challenging to accurately predict multi-dimensional time series data. In comparison, existing reconstruction-based methods can generally achieve more advanced results in complex multi-dimensional time series datasets.
[0007] In the monitoring scenario of the electricity collection system, multi-dimensional time series data generally presents nonlinear time-dependent characteristics, and the correlation patterns between channels are significantly different: when there is strong functional coupling between system sensors, the channel correlation tends to be close; otherwise, it presents a sparse or even unrelated state. Current modeling methods can be mainly divided into two categories: channel-dependent and channel-independent. However, research shows that in some real benchmark datasets, single-variable independent modeling methods that ignore channel correlation often outperform channel-dependent modeling. This phenomenon can be attributed to two factors: first, actual data may naturally have weak inter-sequence correlation; second, existing models lack the ability to capture complex channel relationships. Although channel-independent modeling avoids the difficulty of correlation modeling by decoupling multivariate analysis, its approach of completely discarding potential channel relationships essentially limits the optimization space of anomaly detection performance. It is worth noting that channel-dependent modeling is widely studied but is easily affected by data overfitting or time series pattern confusion. When the historical dependence relationship within multi-dimensional time series data occupies a dominant position, channel-dependent modeling may destroy the time dependence within the sequence by attempting to learn inter-sequence information. SUMMARY
[0008] In view of the deficiencies of the prior art, the present application provides a power consumption data anomaly detection method and device based on channel independence and global-local channel dependence.
[0009] According to one aspect of the present application, a power consumption data anomaly detection method based on channel independence and global-local channel dependence is provided, comprising:
[0010] obtaining multivariate long time series data of historical detection of the to-be-tested electric energy meter;
[0011] dividing the multivariate long time series data into a plurality of time window data of a preset window length;
[0012] inputting the plurality of time window data and adjacent time window data thereof into a pre-trained anomaly detection model to output reconstruction data corresponding to the time window data, wherein the anomaly detection model is used to realize generation of the reconstruction data based on channel independence and global-local channel dependence;
[0013] determining an anomaly score of each time point of each time window data according to the reconstruction data and original data of the time window data, and determining an anomaly degree of each time point of the to-be-tested electric energy meter according to the anomaly score.
[0014] According to another aspect of the present application, a power consumption data anomaly detection device based on channel independence and global-local channel dependence is provided, comprising:
[0015] an obtaining module configured to obtain multivariate long time series data of historical detection of the to-be-tested electric energy meter;
[0016] a dividing module configured to divide the multivariate long time series data into a plurality of time window data of a preset window length;
[0017] a reconstruction module configured to input the plurality of time window data and adjacent time window data thereof into a pre-trained anomaly detection model to output reconstruction data corresponding to the time window data, wherein the anomaly detection model is used to realize generation of the reconstruction data based on channel independence and global-local channel dependence;
[0018] a determining module configured to determine an anomaly score of each time point of each time window data according to the reconstruction data and original data of the time window data, and determine an anomaly degree of each time point of the to-be-tested electric energy meter according to the anomaly score.
[0019] According to still another aspect of the present application, a computer readable storage medium is provided, the storage medium storing a computer program, the computer program being used to execute the method according to any one of the above aspects of the present application.
[0020] According to a further aspect of the present application, there is provided an electronic device comprising: a processor; a memory for storing processor-executable instructions; the processor being arranged to read the executable instructions from the memory and execute the instructions to implement the method of any one of the above aspects of the present application.
[0021] Therefore, the present application proposes a multi-dimensional time series anomaly detection method based on channel independence and global-local channel dependence to solve the problem that existing methods are difficult to effectively model the coupling characteristics of strong correlation, weak correlation or no correlation between channels due to the complex and diverse distribution patterns between dimensions of multi-dimensional time series data. First, in the channel independence process, the proposed method designs a dual-channel variable processing module based on the Mamba model to effectively handle the historical dependence relationship in single variable data according to the different characteristics of continuous and discrete quantities of time series data. Then, in the channel dependence process, the proposed method constructs a twin auxiliary time series based on trend-residual decomposition, and through trend continuous loss and detail reconstruction loss, the model's global trend and detail representation ability are constantly improved by constraining the attention modeling process of the twin auxiliary time series. BRIEF DESCRIPTION OF DRAWINGS
[0022] The exemplary embodiments of the present application can be more fully understood by reference to the following drawings:
[0023] Figure 1 is a flowchart of an electricity consumption data anomaly detection method based on channel independence and global-local channel dependence provided by an exemplary embodiment of the present application;
[0024] Figure 2 is a schematic diagram of a multi-dimensional time series anomaly detection algorithm framework based on channel independence strategy and global-local channel dependence strategy provided by an exemplary embodiment of the present application;
[0025] Figure 3 is a schematic diagram of a CIGCD-Mamba structure provided by an exemplary embodiment of the present application;
[0026] Figure 4 is a schematic diagram of a trend attention mechanism and residual attention mechanism structure provided by an exemplary embodiment of the present application;
[0027] Figure 5 is a schematic diagram of an electricity consumption data anomaly detection device based on channel independence and global-local channel dependence provided by an exemplary embodiment of the present application;
[0028] Figure 6The electronic device is provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0029] Hereinafter, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. It should be apparent that these described embodiments are merely exemplary of the present application and should not be considered limiting the scope of the present application. Therefore, the disclosure of these exemplary embodiments is intended to be illustrative, and not to be limiting of the scope of the present application.
[0030] It should be noted that the relative arrangement of the components and steps, the numerical expressions, and numerical values set forth in these embodiments are not limiting of the scope of the application unless otherwise specifically indicated.
[0031] It should be understood by those skilled in the art that the terms "first", "second", and the like in the embodiments of the present application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they represent the inevitable logical order between them.
[0032] It should also be understood that in the embodiments of the present application, "a plurality of" can mean two or more, and "at least one" can mean one, two or more.
[0033] It should also be understood that for any component, data or structure mentioned in the embodiments of the present application, unless specifically limited or given the opposite implication by the context, it can be understood as one or more in general.
[0034] In addition, the term "and / or" in the present application is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in the present application generally represents an "or" relationship between the front and rear associated objects.
[0035] It should also be understood that the description of the embodiments of the present application emphasizes the differences between the embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0036] At the same time, it should be understood that for the sake of description, the size of each part shown in the drawings is not drawn in accordance with the actual proportion relationship.
[0037] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way limiting of the application or its use.
[0038] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail herein, but should be considered as part of the specification in appropriate circumstances.
[0039] It should be borne in mind, that the use of similar reference numerals in different drawings is intended to represent similar items. Thus, for example, it is understood that, when a part is given a designation in one drawing, that same part is meant to be indicated by a like designation in the other drawings, unless otherwise noted.
[0040] Embodiments of the present application can be applied to terminal devices, computer systems, servers, and the like electronic devices, which can operate with many other general or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and the like electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments that include any of the above systems, and the like.
[0041] Terminal devices, computer systems, servers, and the like electronic devices can be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform particular tasks or implement particular abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules can be located in local or remote computer system storage media including storage devices.
[0042] Exemplary method
[0043] Figure 1 is a flowchart of a power consumption data anomaly detection method based on channel independence and global-local channel dependence provided by an exemplary embodiment of the present application. The present embodiment can be applied to electronic devices, such as Figure 1 As shown in FIG. 1, the power consumption data anomaly detection method 100 based on channel independence and global-local channel dependence includes the following steps:
[0044] Step 101, obtaining multi-variable long time series data of historical detection of the to-be-tested electric energy meter;
[0045] Step 102, dividing the multi-variable long time series data into a plurality of time window data of a preset window length;
[0046] Step 103, inputting the plurality of time window data and adjacent time window data thereof into a pre-trained anomaly detection model, and outputting reconstruction data corresponding to the time window data, wherein the anomaly detection model is used to realize generation of the reconstruction data based on channel independence and global-local channel dependence;
[0047] Step 104: Determine the anomaly score for each time point of the data in each time window based on the reconstructed data and the original data, and determine the degree of anomaly of the energy meter under test at each time point based on the anomaly score.
[0048] Specifically, with the continuous development of electricity consumption data acquisition systems, the multidimensional time-series data they monitor often contains complex and diverse distribution patterns. Existing methods typically treat channels as independent variables or model them using only simple fully connected layers, ignoring the deep dependencies and dynamic interactions between channels. Furthermore, modeling with a purely channel-dependent strategy may lead to the transmission of direct erroneous information between channels, making it difficult to model fine-grained correlations between channels. Therefore, this invention proposes a channel-independent and global-local channel-dependent anomaly detection method, CIGCD. This section first describes the problem of multidimensional time-series anomaly detection and outlines the main framework of the proposed method, then details the key modules and anomaly detection process. The specific implementation is as follows:
[0049] 1. Problem Description
[0050] Multivariate time series data consists of multiple univariate time series data, containing dependencies between multiple features. Time series are typically observed at consecutive, equally spaced time stamps, where the input multivariate time series data is... x = {x1, x2, ..., x} T}, where T is the maximum length of the input timestamp, and M is the number of variables per timestamp, therefore x t ={x 1,t ,x 2,t ,....,x M,t}, Multidimensional temporal anomaly detection and reconstruction methods typically model the input data to obtain the output x′={x′1,x′2,....,x′ T The difference between the input x and the output x′ is calculated to obtain the anomaly score at each time point. The anomaly detection task uses the anomaly score as the evaluation metric to determine the positive anomaly state at that moment, thus obtaining the corresponding output vector y = {y1, y2, ..., y′}. T}, where y t ∈{0,1} (1 represents abnormal data, 0 represents normal data).
[0051] 2. Framework Overview
[0052] like Figure 2As shown, the present application proposes a power consumption data anomaly detection method CIGCD based on channel independence and global-local channel dependence. In view of the unknown correlation between channels of multi-dimensional time series data, CIGCD constructs a dual-link of channel independent modeling and channel dependent modeling of multi-dimensional time series correlation. In the channel independent modeling process, fully considering the different characteristics of continuous and discrete variables, CIGCD introduces the Mamba model into the field of time series anomaly detection, designs a dual-channel variable processing module based on the Mamba model, and processes the time series time dependence based on the excellent sequence modeling ability. In the channel dependent modeling process, in order to prevent the fusion of information between sequences from causing model overfitting or false fitting, CIGCD constructs a twin auxiliary time series based on trend-residual decomposition. In view of the different channel characteristics of trend residual continuous variables, CIGCD constructs a dual-channel attention mechanism to mine the channel correlation of multi-dimensional time series data, and through continuous loss constraint, the trend item is fitted to the characteristics more consistent with the trend dependence relationship, and the residual item is used to mine the characteristics of local points, thereby improving the global trend and detail representation ability of the model. In view of the problem that the existing method cannot correctly capture the weighted relationship between different channels in channel dependent modeling, CIGCD constructs a learnable regulation parameter vector to flexibly adjust the different channel component relationships of channel independent modeling and channel dependent modeling, further improving the model anomaly detection performance.
[0053] Overall, CIGCD mainly includes a channel independent module, a channel dependent module, and a model training and anomaly detection module. First, the standardized input data is split into discrete and continuous variables, which are respectively input into the Mamba model processing unit of the channel independent module to obtain the reconstruction output of the channel independent modeling process. Second, the continuous variables are decomposed into different items based on trend-residual decomposition, combined with the discrete variables to obtain trend and residual item inputs, and input into the continuous Mamba processing unit, and respectively construct the continuous item global attention modeling based on the continuous loss and the residual item local attention modeling based on the reconstruction loss, thereby obtaining the channel dependent module output. Then, through the model training and anomaly detection module, by constructing a controllable parameter vector, the CIGCD model reconstruction output is obtained, and through model training and weighted calculation, the data anomaly score can be obtained.
[0054] 3. Channel independent modeling: discrete and continuous variable analysis based on Mamba model
[0055] As Figure 3As shown, the CIGCD structure models the original sequence while building auxiliary time series ATS based on trend-residual decomposition, which can represent the relationship between sequences, and incorporates the ATS modeling results into the time series anomaly detection results to obtain the model output. When there is no ATS assistance, the CIGCD result is the channel-independent modeling result. Here, considering that the Transformer has a low computational efficiency when processing long sequences, and it does not compress each historical record, it cannot model any information outside the limited window. While the RNN cannot perform content-based reasoning in terms of data perception ability, that is, it cannot perform global perception like Attention, and it lacks the ability to model discrete data. Traditional state space models (SSMs) have very strong modeling capabilities for continuous data, but for discrete data, their performance is poor. The improved state space subsequently improved, but it encountered new problems in maintaining discrete modeling and selective processing, and could not improve the training efficiency through parallel methods. Therefore, the CIGCD introduces the Mamba model based on the Selective State Space Models (S6 model), which analyzes the single-variable time series data of the channel-independent modeling process based on its strong selective information learning ability and efficient training and reasoning ability.
[0056] The Mamba model is a new sequence modeling method that improves the efficiency and effectiveness of processing long sequence data by introducing the S6 model. In the Mamba model, the reconstruction process of the input data X is implemented through an end-to-end neural network architecture that does not rely on attention mechanisms or Multilayer Perceptron (MLP). For input data x, the CIGCD splits it into x i , i = {0, 1, 2,..., M-1} according to the channel. At this time, the Mamba model is used to reconstruct the independent x i data. It should be noted that the original Mamba model is a language sequence model. In order to better apply it to multi-dimensional time series data tasks, the CIGCD designs the CIGCD-Mamba to improve the model's time series data processing capability, such as Figure 2The CIGCD-Mamba includes a Tokenization Layer (TL), a Mamba block, and a Feed-Forward Network (FFN) layer. The TL mainly standardizes and scales the data, while the Mamba block models the dual-channel data of discrete and continuous quantities (CIGCD-Mamba-C and CIGCD-Mamba-D), and the FFN layer performs specific data embedding for subsequent reconstruction output.
[0057] The input of the TL layer is x i . The CIGCD first performs instance normalization on the input data, i.e., InstanceNorm(x i ), to stabilize the training process and reduce the impact of different data scales. Based on this, the CIGCD adopts the tokenization method of sequential text in natural language processing to standardize the time series data. InstanceNorm(x i ) is divided into patches of different scales, and the number of patches is P. These patches are subsets of the original data, used to capture local features:
[0058] x patch = Patch(InstanceNorm(x i ), P) (1)
[0059] The Mamba block includes discrete and continuous dual channels. The S6 model is used in this module to selectively learn multi-dimensional time series data, which effectively learns the interaction information of the data and avoids the accumulation of useless data. For input data is the continuous quantity patch, is the discrete quantity patch. The input data is transmitted into the Mamba dual channel for modeling, and then the intermediate quantity u patch in the Mamba residual network can be obtained, i.e.:
[0060] u patch = CIGCD-Mamba(x patch ) (2)
[0061] The output of the joint Mamba residual network is:
[0062] x′ patch = u patch + x patch (3)
[0063] To address the different characteristics of discrete and continuous data, CIGCD-Mamba incorporates a dual-channel processing module within the Mamba block. For continuous patches with smooth patterns and trends, the Mamba model uses a dual-attention mechanism to acquire representations within and between patches, thereby capturing local and global features in the data, and employs the S6 architecture for further data processing. For discrete patches, the Mamba model can capture specific patterns in the data through the S6 structure. The modeling process for the dual-attention mechanism for continuous data is as follows:
[0064]
[0065] a patch =Multi-HeadAttention({a in}) (5)
[0066] Among them, a in It is the attention representation within the patch, a patch This represents the attention representation between patches. Multi-HeadAttention is a multi-head attention mechanism; see details below. Figure 3 As shown. Place a patch By setting the channels of the normalization layer to match the input, we obtain the input for S6. S6 allows the model to selectively update the hidden states based on the current input. The state update equation is:
[0067] h′(p)=A′h(p)+B′x patch (p) (6)
[0068] y(p)=Ch′(p) (7)
[0069] Where h′(p) is the state at the p-th patch, h(p) is the state before the p-th patch, and x patch,p The input is at the p-th patch, and y(p) is the output at time step p. Matrices A′, B′, and C are obtained by discretizing the learnable matrices A, B, and C, and the step size Δ, respectively. Concatenating y(p) from all time steps yields u, which is the output of the aforementioned process. patch ={y(p),p=1,2,...,P}.
[0070] In the FFN layer, CIGCD further processes the output of the Mamba block. CIGCD then processes the x′ obtained from equation (3). patch Entering the FFN layer for detoxing and flattening yields the output x′ of the channel-independent modeling process. ind =FFN(x′) patch Therefore, the loss in the independent channel modeling process is:
[0071]
[0072] The CIGCD-Mamba pseudo code is shown in Algorithm 1, and the channel-independent modeling process pseudo code is shown in Algorithm 2.
[0073]
[0074]
[0075]
[0076] 4. Channel-dependent modeling: twin auxiliary time series based on trend-residual decomposition
[0077] Based on channel-independent modeling, CIGCD constructs a twin auxiliary time series based on trend-residual decomposition. It is worth noting that the decomposition of discrete quantities will bring greater training difficulty to the model, so CIGCD first separates continuous quantities and discrete quantities before trend-residual decomposition. Based on this, the trend item and the residual item after decomposing the continuous quantity are analyzed as two independent auxiliary time series, and the discrete quantity will be added in the process of Mamba model modeling. Considering that the Mamba model is only suitable for implicit channel analysis, a two-item channel attention mechanism is introduced here to show the analysis of the correlation of each channel.
[0078] As shown in Figure 4 , for input data x = {x con ,x dis}, x con is a continuous quantity, and x dis is a discrete quantity. CIGCD constructs a two-item ATS based on trend-residual decomposition:
[0079] ATS trend , ATS remain = STL(x con ) (9)
[0080] where ATS trend is the sum of the trend item and the seasonal item after STL decomposition, and ATS remain is the residual item. At this time, CIGCD has constructed a twin auxiliary time series. ATS trend , ATS remain are combined with x dis respectively, and are transformed into patch sequences x trend and x remain through a standard normalization layer:
[0081] x trend = Patch(ATS trend ,x dis ), xremain = Patch (ATS remain , x dis ) (10)
[0082] Both of them are passed into CIGCD-Mamba module, which can make the channel-dependent modeling process consistent with the channel-independent modeling process. Thus, we can get
[0083] z trend = CIGCD-Mamba (x trend ), z remain = CIGCD-Mamba (x remain ) (11)
[0084] where z trend and z remain are the Mamba residual network outputs of x trend and x remain . At this time, CIGCD-Mamba processes discrete and continuous quantities similarly to the channel-independent modeling process.
[0085] z trend and z remain are aligned with the original data channels through the FFN normalization layer, and channel attention mechanisms are performed on the two inputs based on trend attention and remaining, as shown in Figure 3 , the output data can be obtained as:
[0086] ATS' trend = TrendAttention (FFN (z trend )) (12)
[0087] ATS' remain = RemainAttention (FFN (z remain )) (13)
[0088] where ATS'trendis the reconstruction output of x trend through the Mamba model and trend attention modeling, and similarly, ATS' remain is the reconstruction output of x remain through the Mamba model and trend attention modeling. It should be noted that CIGCD is not just to have additional information input and help for auxiliary time series. CIGCD makes different loss constraints for trend ATS and remaining ATS. For the trend term, CIGCD hopes to learn the trend semantic features of multi-dimensional time series data, rather than the detailed changes at time points. Therefore, CIGCD smoothes the noise details by applying continuity loss L con , focusing on the overall trend that can be described by inter-sequence dependencies:
[0089]
[0090] Where is the standard deviation of the auxiliary time series in the m-th channel. By applying continuity loss, ATS trend This smooths out temporal details and noise, focusing on the overall trend that can be described by inter-series dependencies. This loss can be viewed as applying a low-pass filter to the ATS data, which effectively reduces high-frequency components in the ATS. Furthermore, ATS... remain The continuity of the data helps reduce the volatility of the output data, which enables it to minimize the impact of noise fluctuations and thus enhances its ability to capture long-term trends in multidimensional time series data.
[0091] For ATS remain CIGCD uses reconstruction error to make the remaining terms focus more on the detailed feature changes of multidimensional time series data:
[0092]
[0093] CIGCD aims to use continuity loss to enable the model to consider the trend semantic features of multidimensional time series data more comprehensively during the trend ATS construction process, and to use reconstruction loss to enable the model to consider the detailed feature changes of time series data during the residual ATS construction process. The model construction and loss constraints are as follows: Figure 2 As shown in Algorithm 3, the pseudocode for the channel dependency modeling process is as follows. At this point, the twin-aided time series output for channel dependency modeling is x′. dep =ξATS′ trend +(1-ξ)ATS′ remain ξ={ξ1,ξ2,...,ξ M} represents a learnable vector of trend weight parameters.
[0094]
[0095]
[0096] 5. Model Training and Anomaly Detection
[0097] After obtaining the outputs of channel-independent modeling and channel-dependent modeling, CIGCD designed a learnable parameter vector λ = {λ1, λ2, ..., λ...} M The magnitude of the channel correlation components is adjusted using a variable, and constrained by the final reconstruction loss. The final reconstruction output of the model is:
[0098]
[0099] At this time, x' = {x'(m, t) | m = {1, 2,..., M}, t = {1, 2,..., T}}, the model reconstruction loss is:
[0100]
[0101] If the result of channel independent modeling is more consistent with the original data, the parameter vector part value will be smaller, and if the component result considering channel correlation is more consistent, the parameter vector part value will be larger. The parameter will be fixed during testing. Then the overall model loss is:
[0102] Loss = L con + L rec + L ind + L io (17)
[0103] Then the model training process is shown in Algorithm 4.
[0104]
[0105]
[0106] 6、Reconstruction and anomaly detection
[0107] CIGCD takes the error between the original data and the reconstructed data as the anomaly score, and then compares the anomaly score output by the model with a given threshold to determine the normal anomaly state of the time series data. For test data (x t m-dimensional data), the anomaly score is calculated as follows:
[0108]
[0109] Where, AS t represents the anomaly score at time t in the preset time window, M is the number of feature dimensions of the multi-dimensional time series data, T is the window size of the time series data, x m , is the mth-dimensional univariate time series data and its corresponding reconstructed data, x i , is the multi-variate time series data at i time and its corresponding reconstructed data.
[0110] In addition, the present application first compares CIGCD with 17 more advanced models on the actual power consumption data set to verify the effectiveness and advancement of the CIGCD algorithm.
[0111] The present application selects AUC, Fc1 and F1 PA%K as evaluation indexes to verify the performance of the proposed method and the baseline model.
[0112] AUC (Area Under Curve) is one of the most commonly used methods for evaluating unsupervised anomaly detection tasks. The AUC evaluation metric primarily calculates the area under the ROC (Receiver Operating Characteristic) curve, ranging from 0 to 1. A perfect dataset will result in an AUC of 1, while random data will produce an AUC value close to 0.5. Compared to traditional evaluation metrics, the advantage of AUC is that it is not affected by threshold settings. However, AUC only reflects the number of time points where the method correctly detects anomalies; a high AUC does not necessarily mean that the method accurately detects all anomaly segments.
[0113] Fc1 (Composite F-score) is a recently proposed metric for time series anomaly detection. Unlike AUC, Fc1's advantage lies in its ability to comprehensively reflect all correctly detected anomaly segments in a multi-dimensional time series, focusing on the algorithm's ability to detect anomalous events. Fc1 calculates both recall within the anomalous time segment and precision at specific time points, thus avoiding the overestimation of algorithm performance by point adjustment strategies. Models with higher recall within anomaly segments and lower false positive rates across time steps receive higher Fc1 scores.
[0114] F1 PA%K (PointAdjustment%K). Similarly, we propose F1. PA%K This addresses the overestimation of model performance caused by point adjustment. It calculates a point-level F1 score, but performs point adjustments when the proportion of outliers detected by the model in consecutive outlier segments exceeds K%. To reduce dependence on the parameter K, the F1 score is... PA%K F1 can be adaptively calculated by adjusting the size of K. PA%K The area under the curve.
[0115] 2. Comparison Methods
[0116] The baseline comparison methods used in this experiment are the current influential mainstream methods, as shown in Table 1. These methods belong to different categories, including some classic methods: LOF, OCSVM and iForest. Channel-independent-based algorithms: DCdetector. Reconstruction-based algorithms: InterFusion. Generative adversarial models: BeatGAN, USAD. Models focusing on dimension or time analysis: GDN, GTA and MSCRED. The former two use graph structures to learn the relationship and coupling between different sensors, thereby achieving robust dimension correlation analysis of multivariate time series data. MSCRED uses long short-term memory networks and attention mechanisms to analyze data. The latest time anomaly detection algorithms: TranAD and AT. The former builds a Transformer-based anomaly detection model based on adaptive and adversarial training procedures. AT designs a method based on priori and sequence correlation based on attention mechanisms, and detects anomalies according to the different correlation differences in the positive anomaly state. Multidimensional time series anomaly detection reconstruction methods: CAE-AD, RAE, TSMAE, MOUT. CAE-AD is a method based on contrastive learning for implicit analysis of noise or anomalies, while the next three methods explicitly design modules for noise analysis. Mask reconstruction algorithm, ImDiffusion, implements unconditional generation-based time interpolation through diffusion models, and detects anomalies by masking reconstruction interpolation error.
[0117] Table 1 Baseline methods
[0118]
[0119]
[0120] 3. Implementation details
[0121] CIGCD is implemented based on PyTorch and trained on a server equipped with Intel(R) Xeon(R) Gold 6148 CPU and Nvidia Tesla V100 GPU. In CIGCD, the Mamba kernel size of the channel-independent modeling process is 3, and the Mamba kernel size of the trend term and the residual term modeling of the channel-dependent modeling process is 3 and 4, respectively. The block expansion factor E for input-output linear projection in the Mamba model is 2. The sliding window length T of the time series data is set to 100, and the dimension D of the embedding vector is 128. During training, the step size of the sliding window is set to 1, while during testing it is set to 100.
[0122] In addition, CIGCD is trained using the Adam optimizer with an initial learning rate of 1e-4 and a batch size of 32. The original training data is divided into training and validation sets in the ratio of 8:2. During training, the learning rate will be halved every time the loss on the validation set does not decrease for 3 epochs, and early stopping will be triggered if the loss on the validation set does not decrease for 6 epochs.
[0123] 4. Introduction of actual power consumption dataset
[0124] The specific data characteristics of the power consumption dataset (ELE) collected by the smart meter are shown in Table 2. The dataset is collected from 9 three-phase meters in multiple areas, each device including current (A phase, B phase, C phase), voltage (A phase, B phase, C phase), electric energy indication (forward active), electric energy indication (reverse active), electric energy indication (forward reactive), electric energy indication (reverse reactive), active power (A phase, B phase, C phase, total value), reactive power (A phase, B phase, C phase, total value), power factor (A phase, B phase, C phase, total value) 22 sensor values.
[0125] Table 2 Characteristics of actual power consumption dataset
[0126]
[0127] These smart meter devices have experienced various types of abnormalities such as power flow reversal, current loss, electric energy meter reverse, electric energy meter fly, electric energy indication uneven, electric energy meter start button box, power differential anomaly, etc. during their respective data recording periods. The dataset contains 9-16 months of continuous data collected by each meter entity at a daily sampling rate of 96 points. The experiment will only include normal data for training and use data containing abnormalities for testing. In addition, the actual power consumption dataset includes 9 complete entity devices, and the actual power consumption dataset exhibits different data sizes and uneven data abnormality proportions on different devices.
[0128] 5. Actual dataset result evaluation
[0129] To verify the universality of the proposed model, the performance of the proposed model is evaluated and analyzed on the actual power consumption dataset, and the comparison results are shown in Table 3. The table shows the AUC, Fc1, F1 PA%K performance of the proposed CIGCD and baseline methods. All results shown in the table are the average of 5 separate runs, which allows the study to evaluate the robustness of each baseline method. In addition, the method with the best performance is highlighted in bold, and the suboptimal performance is indicated in underlined form.
[0130] Notably, CIGCD performs well in all evaluation metrics, achieving an AUC value of 0.6388, an Fc1 score of 0.3804, and an F1 score of 0.4311, respectively. PA%K After analysis, it can be found that compared with the average results of other baseline methods, the AUC index of CIGCD is improved by 17.61%, which indicates that CIGCD has high accuracy in detecting abnormal time points and is less affected by data uncertainty. Similarly, in the Fc1 score, the performance of CIGCD is improved by 40.29% compared with the baseline method, which indicates that CIGCD has high recall rate for abnormal time period detection and high accuracy for abnormal time points. And in the F1 PA%K score, CIGCD is improved by 48.09%, so the detection accuracy of CIGCD algorithm for abnormal time period is also at a high level. The high level of the above three indicators also confirms the superiority and applicability of CIGCD algorithm in the task of multi-dimensional time series data anomaly detection.
[0131] Table 3 Comparative experimental results on actual power consumption data set
[0132]
[0133] From Table 2, it can be seen that the ELE data set has a relatively high abnormality ratio, and the types of anomalies in the data are rich. Therefore, the excellent performance of CIGCD on the actual power consumption data set reflects that the channel independent and channel joint decomposition combination module of the algorithm can better detect anomalies in different regions, thereby adapting to the influence of different anomalies on the model. In addition, since ELE is a multi-dimensional time series power consumption data collected by an actual intelligent three-phase power meter, its features are composed of multiple monitoring quantities such as voltage, current, and electric power distributed in different parts of the system, so its features are continuous quantities, and there is strong correlation between the dimensions of the data. Therefore, the excellent performance of CIGCD on this data set shows that it can learn the dimension correlation well. Overall, the discrete-continuous dual-channel Mamba architecture based on the state space model of CIGCD can effectively realize the information complementation of multi-modal features, and the dual-path attention mechanism based on trend-residual decomposition terms can fully capture the multi-level semantic representation of long-term trends and local details of multi-dimensional time series data. Through the combination of channel independent and channel dependent dual-path learning, CIGCD achieves excellent performance on the ELE data set.
[0134] Therefore, aiming at the problem that the existing method is difficult to effectively model the coupling characteristics of strong correlation, weak correlation or no correlation between dimensions of multi-dimensional time series data which presents complex and diverse distribution patterns, the application proposes a multi-dimensional time series anomaly detection method based on channel independence and global-local channel dependence. Firstly, in the channel independence process, the method designs a double-channel variable processing module based on the Mamba model to effectively process the historical dependence relationship in the single variable data according to the different characteristics of continuous and discrete time series data. Then, in the channel dependence process, the method constructs a twin auxiliary time series based on trend-residual decomposition, and through trend continuous loss and detail reconstruction loss, the attention modeling process of the twin auxiliary time series is constrained to continuously improve the global trend and detail representation ability of the model. Finally, on the basis of the channel independence and channel dependence double link of the multi-dimensional time series correlation relationship, the method constructs a learnable regulation parameter vector to flexibly adjust the component relationship of channel independence and channel dependence, and further improves the anomaly detection performance of the model.
[0135] The key point of the application is:
[0136] 1. Multi-dimensional time series anomaly detection algorithm based on fusion of channel independence and channel dependence
[0137] Aiming at the complexity and uncertainty of the correlation relationship between dimensions of power multi-dimensional time series data, a dual-path learning framework based on fusion of channel independence and channel dependence is proposed. Channel independence modeling avoids cross-channel noise interference by mining the time dependence pattern of each dimension through single variable time series analysis. Channel dependence modeling captures the global potential correlation and local dynamic interaction features between dimensions through an attention cross-correlation matrix. By introducing a learnable gating parameter vector, the weight distribution of the two modeling paths is dynamically adjusted to realize adaptive feature fusion from strong correlation dimensions to weak correlation dimensions.
[0138] 2. Double-channel Mamba processing module based on discrete-continuous
[0139] In order to solve the mixed modeling problem of continuous monitoring values and discrete state labels in power time series data, a discrete-continuous double-channel Mamba architecture based on state space model is proposed. For continuous time series data, the Mamba state space equation is used to model the continuous evolution process of time series in detail, and the long-range historical dependence features are captured through hidden state transmission. For discrete state labels, a discretization state transition module is designed to encode discrete quantities into differentiable probability distributions for embedding. Through the cross-modal interaction gate of the discrete-continuous double-channel, the information complementation is realized to improve the representation ability of the composite industrial time series data.
[0140] 3. Double-path attention mechanism based on trend-residual decomposition term
[0141] Based on the time series decomposition, the continuous monitoring data is parsed into low-frequency trend items and high-frequency residual items, and a two-path attention collaborative analysis framework is constructed. The trend item attention module focuses on the long-term evolution law of the equipment running state, and extracts the long-term influence of working condition switching on system stability through the association modeling of trend components and discrete state labels. The residual item attention module captures abnormal detail signals in short-term fluctuations, and locuses transient disturbances combined with the timestamp information of discrete events. Further, trend continuity loss and detail reconstruction loss are designed to constrain the smoothness requirement of the trend item and the information fidelity of the residual item respectively, and to ensure the balanced optimization of global trend and local details in the anomaly detection task.
[0142] Exemplary apparatus
[0143] Figure 5 FIG. 1 is a structural schematic diagram of an electricity data anomaly detection device based on channel independence and global-local channel dependence provided by an exemplary embodiment of the present application. As shown in Figure 5 the device 500 includes:
[0144] The acquisition module 510 is configured to acquire multivariate long time series data of historical detection of the to-be-tested electric energy meter.
[0145] The division module 520 is configured to divide the multivariate long time series data into a plurality of time window data of a preset window length.
[0146] The reconstruction module 530 is configured to input the plurality of time window data and adjacent time window data thereof into a pre-trained anomaly detection model, and output reconstruction data corresponding to the time window data, wherein the anomaly detection model is used to realize generation of the reconstruction data based on channel independence and global-local channel dependence.
[0147] The determination module 540 is configured to determine an anomaly score of each time point of each time window data according to the reconstruction data and the original data of the time window data, and determine an anomaly degree of each time point of the to-be-tested electric energy meter according to the anomaly score.
[0148] Exemplary electronic device
[0149] Figure 6 FIG. 6 is a structure of an electronic device provided by an exemplary embodiment of the present application. As shown in Figure 6 the electronic device 60 includes one or more processors 61 and a memory 62.
[0150] The processor 61 can be a central processing unit (CPU) or other forms of processing unit having data processing capability and / or instruction execution capability, and can control other components in the electronic device to perform desired functions.
[0151] The memory 62 can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 61 can run the program instructions to implement the methods of the software programs of the various embodiments of the present application described above and / or other desired functions. In one example, the electronic device can further include an input device 63 and an output device 64, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0152] In addition, the input device 63 can include, for example, a keyboard, a mouse, and / or the like.
[0153] The output device 64 can output various information to the outside. The output device 64 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and / or the like.
[0154] Of course, in order to simplify, Figure 6 Only some of the components of the electronic device related to the present application are shown in FIG. 1, and components such as a bus, an input / output interface, and the like are omitted. In addition to this, the electronic device can further include any other appropriate components according to the specific application.
[0155] Exemplary computer program product and computer readable storage medium
[0156] In addition to the above-mentioned methods and devices, embodiments of the present application can also be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.
[0157] The computer program product can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, and / or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote cloud device or server.
[0158] In addition, an embodiment of the present application can also be a computer readable storage medium, having stored thereon computer program instructions which, when executed by a processor, cause the processor to carry out the steps described in the above "Exemplary Method" section of this specification for the methods according to various embodiments of the present application.
[0159] The computer readable storage medium can be any combination of one or more computer readable medium(s). The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or apparatus or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0160] The above describes the basic principles of the present application in conjunction with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present application are only examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as necessary for each embodiment of the present application. In addition, the above specific details are only for the purpose of example and for the purpose of understanding, and are not limiting, and the above details do not limit the present application to the above specific details.
[0161] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts are described in the method embodiment.
[0162] The block diagrams of the devices, systems, apparatuses, systems involved in the present application are only illustrative examples and are not intended to require or imply that the connections, arrangements, configurations are as shown in the block diagrams. As those skilled in the art will recognize, these devices, systems, apparatuses, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, meaning "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.
[0163] The methods and systems of the present application can be implemented in a number of ways. For example, the methods and systems of the present application can be implemented via software, hardware, firmware, or any combination of software, hardware, and firmware. The above described order of steps for the methods is merely illustrative, and the steps of the methods of the present application are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, the present application can also be implemented as a program recorded on a recording medium, which includes machine readable instructions for implementing the methods according to the present application. Thus, the present application also covers recording media storing programs for executing the methods according to the present application.
[0164] It is also important to note that the systems, devices and methods of the present application can be embodied in a variety of forms without departing from the spirit of the application. Specifically, the systems, devices and methods of the present application can be implemented using hardware, software, firmware, or any combination thereof. In some embodiments, the systems, devices and methods of the present application can be implemented as a program tangibly embodied on a program carrier. It is therefore desirable to cover any and all modifications and variations of the application apparent to those skilled in the art. It is intended that the scope of the application should only be limited by the claims.
[0165] The above description has been presented for the purposes of illustration and description. Further, this description is not intended to limit the embodiments of the application to the form disclosed herein. Although various example aspects and embodiments have been discussed above, those of ordinary skill in the art will appreciate a variety of modifications, alternatives, permutations, additions, and sub-combinations, which fall within the scope of the application.
Claims
1. A power consumption data anomaly detection method based on channel independence and global-local channel dependence, characterized in that, The method comprises the following steps: obtaining the multivariate long time series data of the historical detection of the to-be-tested electric energy meter; dividing the multivariate long time series data into a plurality of time window data of a preset window length; inputting the plurality of time window data and its adjacent time window data into a pre-trained anomaly detection model to output the reconstruction data corresponding to the time window data, wherein the anomaly detection model is based on channel independence and global-local channel dependence to realize the generation of reconstruction data; determining the anomaly score of each time point of each time window data according to the reconstruction data and the original data of each time window data, and determining the anomaly degree of each time point of the to-be-tested electric energy meter according to the anomaly score.
2. The method of claim 1, wherein, The multivariate long time series data comprises A-phase current, B-phase current, C-phase current, A-phase voltage, B-phase voltage, C-phase voltage, positive active energy indication value, reverse active energy indication value, positive reactive energy indication value, reverse reactive energy indication value, A-phase active power, B-phase active power, C-phase active power, total active power value, A-phase reactive power, B-phase reactive power, C-phase reactive power, total reactive power value, A-phase power factor, B-phase power factor, C-phase power factor, and total power factor value.
3. The method of claim 1, wherein, After obtaining the multivariate long time series data of the historical detection of the to-be-tested electric energy meter, the method further comprises the following steps: standardizing each variable in the multivariate long time series data by using a maximum-minimum standardization method to obtain standardized multivariate long time series data, including continuous quantities and discrete quantities; wherein, the expression of the maximum-minimum standardization method is: In the formula, X i is multivariate long time series data, represents normalized X i , X i,max represents X i The maximum value of all sample data of each variable in X i,min represents X i The minimum value of all sample data of each variable in X 4. The method of claim 1, wherein, The training process of the anomaly detection model is as follows: obtaining a plurality of multivariate time series historical data samples of the historical detection of electric energy meters and merging them into a multivariate long time series historical data; standardizing each variable in the multivariate long time series historical data by using a maximum-minimum standardization method to obtain standardized multivariate long time series historical data, wherein the multivariate long time series historical data comprises continuous quantities and discrete quantities; windowing the standardized multivariate long time series historical data samples to divide them into a plurality of time window historical data samples of the preset time window; analyzing any time window historical data sample in the plurality of time window historical data samples based on a channel-independent model of a pre-constructed double-channel Mamba module to output channel-independent reconstruction data; analyzing any time window historical data sample in the plurality of time window historical data samples based on a channel-dependent model of a pre-constructed trend-residual decomposition to output channel-dependent reconstruction data; based on the channel-independent reconstruction data and the channel-dependent reconstruction data, forming reconstruction data samples; determining the total loss of the anomaly detection model according to a pre-constructed total loss function and the reconstruction data samples; updating the optimization network and parameters according to the total loss until convergence to determine the anomaly detection model.
5. The method of claim 4, wherein, The channel independent model based on the pre-constructed Mamba model analyzes any time window historical data sample in a plurality of time window historical data samples, and outputs channel independent reconstruction data, including: The tokenization layer of the channel independent model normalizes the discrete data and continuous data in the historical data sample respectively, to obtain continuous patches and discrete patches; The continuous patches and the discrete patches are input into the Mamba block of the channel independent model for modeling, to obtain Mamba residual network continuous intermediate quantities and discrete intermediate quantities; The dual-channel processing module is used to jointly process the continuous patches and the continuous intermediate quantities, and the discrete patches and the discrete intermediate quantities, to obtain joint output, wherein the dual-channel processing module includes a discrete quantity processing channel and a continuous quantity processing channel, and the continuous quantity processing channel uses a double attention mechanism to obtain intra-patch and inter-patch representations; The feedforward neural network layer of the channel independent model performs de-tokenization flattening on the joint output, to output the channel independent reconstruction data.
6. The method of claim 4, wherein, The channel dependent model based on the pre-constructed trend-residual decomposition analyzes any time window historical data sample in a plurality of time window historical data samples, and outputs channel dependent reconstruction data, including: trending and decomposing the continuous data in the historical data samples to obtain a two-term ATS, wherein the two-term ATS comprises a sum of a decomposed trend term and a season term ATS trend and a residual term ATS remain ; the sum of the trend item and the season item ATS trend and the residual item ATS remain are combined with the discrete data in the historical data sample respectively and transformed into a trend item patch sequence and a residual item patch sequence of the multi-dimensional time series data through a standard normalization layer; The trend item patch sequence and the residual item patch sequence are input into the CIGCD-Mamba module of the channel dependent model, to obtain trend item residual network output and residual item residual network output; The trend item residual network output and the residual item residual network output are input into the FFN normalization layer of the channel dependent model, to output trend reconstruction output and residual reconstruction output; According to the trend reconstruction output and the residual reconstruction output, the channel dependent reconstruction data is output.
7. The method of claim 6, wherein, The calculation expression of the reconstruction data sample is: x'(m, t) = (1 - λ m )x' ind (m, t) + λ m (x' dep (m, t)) = (1 - λ m )x′ ind (m,t) + λ m (ξ m ATS′ trend (m,t) + (1 - ξ m )ATS′ remain (m,t)) = (1 - λ m )x′ ind (m,t) + λ m ξ m ATS′ trend (m,t) + λ m (1 - ξ m )ATS′ remain (m,t) where x' = {x'(m, t) | m = {1, 2,...., M}, t = {1, 2,.... T}}, M is the number of feature dimensions of multi-dimensional time series data, and T is the size of time series data window; λ = {λ1, λ2,..., λ M} is a learnable parameter vector; x' ind is channel-independent reconstructed data; x' dep is channel-dependent reconstructed data; ATS trend (m, t) is the trend reconstruction output of m dimensions at time t; ATS remain (m, t) is the residual reconstruction output of m dimensions at time t; ξ = {ξ1, ξ2,...., ξ M} is a learnable trend weight parameter vector.
8. The method of claim 4, wherein, The expression of the total loss function Loss is: Loss = L con + L rec + L ind + L io Wherein, where L con is the continuity loss function of the channel-independent model; L rec is the dependency loss function of the channel-dependent model; L io is the reconstruction loss function; σ i is the standard deviation of the auxiliary time series in the i-th channel; M is the number of feature dimensions of the multi-dimensional time series data, T is the window size of the time series data, M trend is the dimension of the trend term, M remain is the dimension of the residual term; x′ ind is the channel-independent reconstructed data; x' dep is the channel-dependent reconstructed data; ATS trend (m, t) is the trend term and discrete quantity input of m dimensions at time t; ATS remain (m, t) is the trend term and discrete quantity input of m dimensions at time t; ATS' trend (m, t) is the trend reconstruction output of m dimensions at time t; ATS' remain (m, t) is the residual reconstruction output of m dimensions at time t; x is the model time series input data, x′ is the model reconstruction output data; λ = {λ1, λ2,..., λ M} is the learnable parameter vector, ξ = {ξ1, ξ2,..., ξ M} is the learnable trend weight parameter vector.
9. The method of claim 1, wherein, The calculation formula of the anomaly score is: In the formula, AS t represents the anomaly score at time t in the preset time window, M is the number of feature dimensions of the multi-dimensional time series data, T is the window size of the time series data, x m , is the single-variable time series data of the mth dimension and its corresponding reconstructed data, x i 、 is the multi-variable time series data at time i and its corresponding reconstructed data.
10. A power consumption data anomaly detection device based on channel independence and global-local channel dependence, characterized in that, Including: An acquisition module is configured to acquire multivariate long time series data of a to-be-tested electric energy meter; A division module is configured to divide the multivariate long time series data into a plurality of time window data of a preset window length; A reconstruction module is configured to input a plurality of time window data and adjacent time window data thereof into a pre-trained anomaly detection model, and output reconstruction data corresponding to the time window data, wherein the anomaly detection model is used to generate reconstruction data based on channel independence and global-local channel dependence; A determination module is configured to determine an anomaly score of each time point of each time window data according to reconstruction data and original data of the time window data, and determine an anomaly degree of each time point of the to-be-tested electric energy meter according to the anomaly score.
Citation Information
Patent Citations
Electric energy meter anomaly detection method and device based on diffusion model
CN118501795A
Multi-source ammeter data intelligent fusion and abnormity identification method
CN120257220A
Multivariate time series anomaly detection method for intelligent internet of things system
WO2024207627A1