A method and system for evaluating the operating state of an electric energy meter under small sample conditions
By using the combination of data augmentation and meta-learning model and LSTM model in the power meter operation state evaluation method, the difficulty of power meter operation state recognition under small sample data is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510301526.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The prior art is difficult to obtain real training data under the actual operating conditions of the power meter under the power outage detection, and the sample data is small, making it difficult to train the model, making it difficult to accurately identify the operating status of the power meter.
The power meter operation status evaluation method under small sample conditions is adopted to construct the original data set by collecting power meter operation data, and data enhancement is performed to generate and expand data sets, and the meta-learning model and LSTM model are used for training to improve the accuracy of the power meter operation status recognition.
Through data augmentation and model training, the number of data samples is expanded, the model training and prediction accuracy is improved, and the operating status of the power meter can be more accurately identified.
Smart Images

Figure CN119810608B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of on-line monitoring of electric power metering, and particularly relates to a method and system for evaluating the operation state of an electric energy meter under small sample conditions. Background Art
[0002] As a key metering device in the high-voltage power transmission system of the power system, the accuracy of the electric energy meter directly affects the billing and management of electric power. With the continuous development of the power market and the diversification of the needs of power users, ensuring the accuracy and reliability of high-voltage electric energy meters has become particularly important. Traditional methods for detecting electric energy meters mainly rely on regular off-line detection, which not only requires power outage, affects the normal power consumption of users, but also increases the detection cost and time cost. In recent years, on-line detection technology has gradually become the mainstream method for detecting the error state of high-voltage electric energy meters. This method allows detection to be carried out when the electric energy meter is operating normally, avoiding the impact of power outage on users, and at the same time reducing the labor and time costs. The advantage of on-line detection technology lies in its high efficiency and real-time performance, which can timely detect and repair potential problems of the electric energy meter, ensuring the accuracy of electric energy metering.
[0003] At present, the on-line detection methods for the state of electric energy meters mainly include the following several types: 1. Signal analysis method: By analyzing the waveforms and frequencies of current and voltage signals, and using technologies such as Fourier transform, the working state of the electric energy meter is monitored in real time. This method can effectively identify whether there are errors or faults in the electric energy meter. 2. Intelligent monitoring device: Combining with Internet of Things technology, the intelligent monitoring device installed on the electric energy meter can collect the operation data of the electric energy meter in real time and analyze it through the cloud platform to timely detect abnormal states. 3. Infrared thermal imaging technology: Using an infrared thermal imager to perform non-contact detection on the electric energy meter, and monitoring the temperature change of the electric energy meter during operation. Abnormal temperature may indicate equipment failure or error, and this method can detect problems in the early stage. 4. Data fusion and pattern recognition: By fusing the data of the electric energy meter with the data of other related devices, and using machine learning and pattern recognition technologies, the possible error states of the electric energy meter are identified. This method can improve the accuracy and efficiency of error detection.
[0004] In summary, the application of on-line detection technology makes the detection of the error state of electric energy meters more efficient and economical, providing an important guarantee for the management and operation of power enterprises. With the continuous progress of technology, on-line detection methods are expected to further improve the accuracy and reliability of electric energy meters, promoting the intelligent development of the power industry.
[0005] The literature "Liu Yuexiao, Yang Guanghua, Li Na, et al. A Method for Evaluating the Operating State of Electric Energy Meters Based on Transfer Learning [J]. Automation Application, 2024, 65(16): 142-146, 149. DOI: 10.19769 / j.zdhy.2024.16.044." proposed a method for evaluating the operating state of electric energy meters based on transfer learning. First, through experimental simulation, the operating data, environmental data, and attribute data of electric energy meters under different operating states were obtained, and the Adaboost algorithm was used to construct an association model between operating characteristics, environmental characteristics, attribute characteristics, and operating states. Secondly, through transfer learning, the Adaboost evaluation model established in the laboratory was adaptively adjusted to be applicable to evaluating the actual operating state of electric energy meters on-site. The above method uses transfer learning technology and requires a neural network model for transfer. When training the neural network model, a large amount of data is needed. However, it is difficult to obtain real training data through power outage detection under the actual operating conditions of electric energy meters, and the sample data volume is small, resulting in difficulty in training the model and thus difficult to accurately identify the operating state of electric energy meters. Summary of the Invention
[0006] The technical problem to be solved by the present invention is that in the prior art, it is difficult to obtain real training data through power outage detection under the actual operating conditions of electric energy meters, and the sample data volume is small, resulting in difficulty in training the model and thus difficult to accurately identify the operating state of electric energy meters.
[0007] The present invention solves the above technical problems through the following technical means: A method for evaluating the operating state of electric energy meters under small sample conditions, including:
[0008] S1. Collect the operating data of the electric energy meter and its corresponding evaluation results of the operating state of the electric energy meter to construct an original data set, and convert the original data set into a three-dimensional matrix of N× ×D, where N is the batch size, is the time window size, and D is the feature dimension, which are two dimensions of amplitude and phase respectively; input the three-dimensional matrix of N× ×D into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map; pass the multi-channel weighted fusion feature map through a normalization layer and a first fully connected layer in sequence, and output a hidden space vector; decompose the hidden space vector into a trend component and three periodic components, and flatten and splice the outputs of the trend component and the three periodic components into two-dimensional data to obtain an extended data set;
[0009] S2. Label a part of the data in the extended data set as training data, use the training data to train the meta-learning model, and input the extended data set into the trained meta-learning model to obtain the labeling results of the extended data set.
[0010] S3. Input the labeled extended dataset and the original dataset into the LSTM model to output the power meter status categories. Train the LSTM model, and input the real-time collected power meter data into the trained LSTM model to obtain the power meter operation status evaluation result.
[0011] Furthermore, convert the original dataset into a three-dimensional matrix of N × × D, including: with the time window size being , and the step size being , convert the original dataset into a three-dimensional matrix of N × × D.
[0012] Furthermore, input the three-dimensional matrix of N × × D into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map, including:
[0013] The multi-head self-attention mechanism includes h attention heads. The branch vector of the j-th attention head includes a query vector , a key vector , and a value vector , where j = 1, 2... h; calculate the attention weight of the j-th attention head according to the query vector and the key vector ; perform matrix multiplication on and to obtain weight feature vectors of h channels, multiply it with the three-dimensional matrix of N × × D to obtain the original feature maps of h channels, then use convolution and average pooling operations to obtain the pooled weight feature maps of h channels, then add the original feature maps of h channels and the pooled weight feature maps using the tensor broadcast mechanism to obtain the weighted feature maps of h channels, and finally fuse these h-channel weighted feature maps through the second fully connected layer to obtain the multi-channel weighted fusion feature map.
[0014] Furthermore, decompose the latent space vector into a trend component and three periodic components, including:
[0015] Use Tucker decomposition to decompose the latent space vector into a core tensor and three factor matrices. By performing low-pass filtering on the core tensor, it serves as the trend component; the three factor matrices are factor matrices corresponding to the batch size N, the time window size and the feature dimension D respectively, serving as the periodic components.
[0016] Further, training the meta-learning model using the training data in S2 means inputting the training data into the meta-learning model to train the meta-learning model. When the value of the loss function is minimized, the training stops, and a trained meta-learning model is obtained. Among them, the similarity between the th sample in the extended dataset and the corresponding sample in the original dataset and the adjustment parameter are used to calculate the th weight coefficient. Then, the cross-entropy loss between the true label of the th sample in the extended dataset and the probability distribution predicted by the meta-learning model and the product of the th weight coefficient are used to construct the loss function.
[0017] Further, S3 includes:
[0018] The LSTM model includes a forget gate, an input gate, and an output gate. Among them, the formula for the forget gate at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step, and the extended dataset at the current time step and the first weight matrix, plus the first bias term, and then multiplied by the logistic function; the formula for the input gate at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step, and the extended dataset at the current time step and the second weight matrix, plus the second bias term, and then multiplied by the Sigmoid activation function of the input gate; the candidate cell state at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step, and the weight matrix of the candidate cell state, plus the bias term of the candidate cell state as the input value and input into the hyperbolic tangent activation function; the result of the forget gate at the current time step is multiplied by the cell state at the previous time step, plus the product of the input gate at the current time step and the candidate cell state at the current time step to obtain the cell state at the current time step; the formula for the output gate at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step, and the weight matrix of the output gate, plus the bias of the output gate, and then multiplied by the Sigmoid activation function of the output gate; the hidden state of the last time step generated by the LSTM model is saved as the feature representation and sent to a third fully connected layer with a softmax activation function to output the electricity meter status category.
[0019] The present invention also provides an electricity meter operating status evaluation system under small sample conditions, including:
[0020] A data augmentation module for collecting electricity meter operation data and its corresponding electricity meter operation status evaluation results to construct an original dataset, and converting the original dataset into a three-dimensional matrix of N × × D, where N is the batch size, is the time window size, and D is the feature dimension, which are two dimensions of amplitude and phase respectively; N × The three-dimensional matrix of N×D is input into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map; the multi-channel weighted fusion feature map is successively passed through a normalization layer and a first fully connected layer to output a latent space vector; the latent space vector is decomposed into a trend component and three periodic components, and the outputs of the trend component and the three periodic components are flattened and concatenated into two-dimensional data to obtain an extended data set;
[0021] A data labeling module, which is used to label a part of the data in the extended data set as training data, train a meta-learning model using the training data, and input the extended data set into the trained meta-learning model to obtain the labeling result of the extended data set;
[0022] A state evaluation module, which is used to input the labeled extended data set and the original data set into the LSTM model, output the state category of the electricity meter, train the LSTM model, and input the real-time collected electricity meter data into the trained LSTM model to obtain the operation state evaluation result of the electricity meter.
[0023] Furthermore, the original data set is converted into a three-dimensional matrix of N× ×D, including: with the time window size of , and the step size of , the original data set is converted into a three-dimensional matrix of N× ×D.
[0024] Furthermore, inputting the three-dimensional matrix of N× ×D into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map, including:
[0025] The multi-head self-attention mechanism includes h attention heads. The branch vector of the j-th attention head includes a query vector , a key vector , and a value vector , where j = 1, 2... h; calculate the attention weight of the j-th attention head according to the query vector and the key vector ; perform matrix multiplication on and to obtain weight feature vectors of h channels, multiply it with the three-dimensional matrix of N× ×D to obtain the original feature maps of h channels, then adopt convolution and average pooling operations to obtain the pooled weight feature maps of h channels, and then add the original feature maps of h channels and the pooled weight feature maps using the tensor broadcasting mechanism to obtain the weighted feature maps of h channels. Finally, fuse the weighted feature maps of these h channels through a second fully connected layer to obtain a multi-channel weighted fusion feature map.
[0026] Further, the latent space vector is decomposed into a trend component and three periodic components, including:
[0027] The latent space vector is decomposed into a core tensor and three factor matrices using Tucker decomposition. By performing low-pass filtering on the core tensor, it serves as the trend component; the three factor matrices are factor matrices corresponding to the batch size N, the time window size and the feature dimension D, serving as the periodic components.
[0028] Furthermore, training the meta-learning model using the training data in the data labeling module means inputting the training data into the meta-learning model for training. When the value of the loss function is minimized, the training stops to obtain the trained meta-learning model; among them, using the similarity between the th sample in the extended dataset and the corresponding sample in the original dataset and the adjustment parameter to calculate the th weight coefficient, and then using the cross-entropy loss between the true label of the th sample in the extended dataset and the probability distribution predicted by the meta-learning model and the product of the th weight coefficient to construct the loss function.
[0029] Furthermore, the state evaluation module is also used for:
[0030] The LSTM model includes a forget gate, an input gate, and an output gate. Among them, the formula for the forget gate at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step, and the extended dataset at the current time step and the first weight matrix, plus the first bias term, and then multiplied by the logistic function; the formula for the input gate at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step, and the extended dataset at the current time step and the second weight matrix, plus the second bias term, and then multiplied by the Sigmoid activation function of the input gate; the candidate cell state at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step and the weight matrix of the candidate cell state, plus the bias term of the candidate cell state as the input value and input into the hyperbolic tangent activation function; multiplying the result of the forget gate at the current time step by the cell state at the previous time step and adding the product of the input gate at the current time step and the candidate cell state at the current time step to obtain the cell state at the current time step; the formula for the output gate at the current time step is the product of the hidden state at the previous time step, the original dataset at the current time step and the weight matrix of the output gate, plus the bias of the output gate, and then multiplied by the Sigmoid activation function of the output gate; saving the hidden state at the last time step generated by the LSTM model as the feature representation and sending it into a third fully connected layer with a softmax activation function to output the power meter state category.
[0031] The advantages of the present invention are as follows:
[0032] (1) The present invention inputs the original data set into the data generation model to achieve data augmentation and obtain an extended data set. Then, the generated extended data set is labeled, so that the generated extended data set can be used for the subsequent training of the LSTM model. Generally speaking, the number of data samples is expanded, the problem of small sample data volume is solved, which helps to train the model, and the accuracy rate is improved in subsequent model training and prediction, and the operation state of the electric energy meter is accurately identified.
[0033] (2) The present invention uses the channel multi-head attention mechanism to map the feature sequence in the channel domain into multiple subspaces to obtain richer channel relationships. Compared with the traditional channel attention mechanism, it captures more diverse channel relationships and can better help the network understand the relationships between different channels in the image, thereby improving the segmentation quality.
[0034] (3) The weight coefficient in the loss function set when training the meta-learning model of the present invention is a coefficient used to adjust the contribution degree of each sample in the loss function. The weight coefficient reflects the similarity or difference between each sample in the extended data set and the corresponding sample in the original data set. The purpose of the weight coefficient is to make the model pay more attention to those samples that are more similar or more important to the known original data set, thereby improving the accuracy of label assignment.
[0035] (4) The present invention inputs both the obtained extended data set and the original data set into the LSTM model, which is equivalent to passing through the activation function output twice. The two data sets jointly influence the output of the LSTM model to affect the results of the forgetting gate and the input gate, and the output result is more accurate. Description of the Drawings
[0036] Figure 1 It is a flowchart of a method for evaluating the operation state of an electric energy meter under small sample conditions disclosed in Embodiment 1 of the present invention;
[0037] Figure 2 It is a framework diagram of a data generation model in a method for evaluating the operation state of an electric energy meter under small sample conditions disclosed in Embodiment 1 of the present invention;
[0038] Figure 3 It is a schematic diagram of the LSTM model architecture in a method for evaluating the operation state of an electric energy meter under small sample conditions disclosed in Embodiment 1 of the present invention. Detailed Embodiments
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1
[0041] As Figure 1 shown, Embodiment 1 of the present invention provides a method for evaluating the operating state of an electric energy meter under small-sample conditions, including:
[0042] S1. Collect the operating data of the electric energy meter and its corresponding evaluation results of the operating state of the electric energy meter to construct an original data set , and convert the original data set into a three-dimensional matrix of N× ×D, where N is the batch size, is the time window size, and D is the feature dimension, which are two dimensions of amplitude and phase respectively; input the three-dimensional matrix of N× ×D into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map; pass the multi-channel weighted fusion feature map through a normalization layer and a first fully connected layer in sequence to output a hidden space vector; decompose the hidden space vector into a trend component and three periodic components, and flatten and splice the outputs of the trend component and the three periodic components into two-dimensional data to obtain an extended data set ; the specific process is as follows:
[0043] S11. Collect the historical voltage amplitude sequence data of the same model and same station of the electric energy meter, denoted as X, and its corresponding off-line evaluation data, which are divided into normal and abnormal, to construct an original data set . The original data set has a small number of samples, and each sample is a sequence data containing voltage data at the second level of unit time. The present invention designs a data generation model to generate new samples to achieve data augmentation, which mainly consists of two parts: the encoding layer (Encoder) maps the input data to the latent space and outputs the mean and variance of the latent variables; the decoding layer (Decoder) reconstructs the input data from the samples in the latent space. By introducing the multi-head self-attention mechanism into the encoding network to mine the deep features of the original data, comprehensive feature information is provided for the decoding network when generating data, realizing the generation of virtual data that is different from the original data but has a similar distribution. Finally, the number of data samples can be expanded, the problem of small sample size can be solved, and the accuracy can be improved in the subsequent network model training and prediction. The framework of the data generation model is as Figure 2 shown.
[0044] In the Encoder part, the original dataset is a sequence data containing voltage data in seconds per unit time, divided into two dimensions of amplitude and phase, both with a length of L. The vector data of each dimension is denoted as (x1, x2,..., xL). Let the size of the time window be , the step size be , and the window data batch size be N. Then:
[0045]
[0046] After simplification, we get:
[0047]
[0048] ;
[0049] Among them, represents the data length of the original dataset , represents the i-th data of the amplitude dimension data in the original dataset , represents the mean value of the amplitude dimension data in the original dataset , represents the standard deviation of the amplitude dimension data in the original dataset . According to the calculation results, with the time window size of , the step size of , the original dataset is transformed into a three-dimensional matrix of N× ×D.
[0050] S12. Import the three-dimensional matrix of N× ×D into the multi-head self-attention mechanism. The multi-head self-attention mechanism component can mine the deep features of the data. The multi-head self-attention mechanism is composed of multiple self-attention mechanisms. Compared with the self-attention mechanism, it can map the data to different subspaces to mine more comprehensive feature information of the data. By randomly generating weight matrices Wq, Wk, and Wv with appropriate sizes according to the uniform distribution, Q, K, and V are calculated through linear transformation. Q, K, and V are the query vector, key vector, and value vector respectively, and , , . To avoid the self-attention mechanism from over-concentrating on the current channel when encoding channel information, Q, K, and V are divided into h groups, so that the different encoding information generated by mapping the channel information to different subspaces. Therefore, the multi-head self-attention mechanism in the present invention includes h attention heads. The branch vectors of the j-th attention head include the query vector , the key vector , the value vector , where \(j = 1, 2, \cdots, h\); calculate the attention weights of the \(j\)-th attention head , where represents the transpose of a matrix, represents the key vector is the dimension of, used to scale the dot product to stabilize the gradient, represents function;
[0051] Each element in represents the correlation between different channels of the feature map. Multiply with and through matrix multiplication to obtain the weighted feature vectors of \(h\) channels. Multiply it with the three-dimensional matrix of \(N\times \times D\) to obtain the original feature maps of \(h\) channels. At this stage, the multi-head attention generates multiple branch vectors. Different branch vectors reflect different dependency relationships due to different feature expressions, thereby extracting diversified channel relationships. Therefore, the semantic of the generated attention matrix will be richer. On the other hand, since the multi-head attention only extracts features in the channel dimension, it will not generate an excessive amount of computation. Subsequently, the original feature maps of \(h\) channels are passed into the channel enhancement part. The channel enhancement part can be understood as an operation for feature enhancement. At this stage, convolution and average pooling operations are adopted to obtain the pooled weighted feature maps of \(h\) channels. Then, add the original feature maps of \(h\) channels and the pooled weighted feature maps using the tensor broadcasting mechanism to obtain the weighted feature maps of \(h\) channels. Finally, fuse these weighted feature maps of \(h\) channels through the second fully connected layer to obtain the multi-channel weighted fusion feature map. Tensor broadcasting addition is an operation method, which means automatically adapting to the addition of tensors with different sizes, that is, the size of the pooled weighted feature map may be different from that of the original feature map. The size of the pooled weighted feature map is adaptively extended to the size of the original feature Figure 1 map, and then the two are added. The channel multi-head attention mechanism maps the feature sequence in the channel domain to multiple subspaces to obtain richer channel relationships. Compared with the traditional channel attention mechanism, it captures more diversified channel relationships and can better help the network understand the relationships between different channels in the image, thereby improving the segmentation quality.
[0052] S13. Pass the multi-channel weighted fusion feature map through a normalization layer and a first fully connected layer in sequence, and output the latent space vector S;
[0053] S14. The input of the Decoder is the latent space vector S generated by the Encoder. Since time series data generally has trend and seasonal characteristics, a time structure is injected into the Decoder part. The time structure is divided into a trend component and a periodic component. For the latent space vector S in the form of a three-dimensional matrix of N×T×D, regarded as a tensor, the Tucker decomposition is used to decompose the latent space vector S into a core tensor and three factor matrices, namely:
[0054]
[0055] where, is the core tensor. By performing low-pass filtering on the core tensor to remove high-frequency noise and retain low-frequency components, which usually represent the long-term trend in the data, as the trend component; , , are factor matrices corresponding to the batch size N, the time window size and the feature dimension D respectively, representing the change direction of the data, as the periodic components.
[0056] Flatten and concatenate the outputs of the trend component and the three periodic components into two-dimensional data to obtain an extended dataset . Since the core tensor includes data in multiple dimensions and is mixed together, it is integrated into the required two-dimensional data by flattening and concatenating. Flattening and concatenating are processes in the prior art. Flattening means sorting the core tensor according to the batch size N and the time window size to form a data sequence under a time series. Then the data in this time series contains different dimensions. The data of the same dimension is placed in the same row (i.e., the concatenation process) to form two-dimensional data as the extended dataset .
[0057] S2. The extended dataset is a two-dimensional voltage feature data without labels. A small amount of data in this data is manually labeled with normal and abnormal probabilities by expert diagnosis as the training data for the subsequent model. The meta-learning model is trained using the training data. The extended dataset is input into the trained meta-learning model to obtain the labeling result of the extended dataset ; in this embodiment, the meta-learning model refers to a BP neural network. The BP neural network uses the process of meta-learning, so it is named the meta-learning model. The process of meta-learning refers to the loss function training process involved in S2 of the present invention. In actual applications, the meta-learning model can also use other neural network models combined with the loss function involved in S2 for model training. The specific process of S2 is as follows:
[0058] Based on each element in the extended dataset and the difference between the corresponding element in the original dataset to determine the th weight coefficient , the th weight coefficient is a coefficient used to adjust the contribution degree of the th sample in the loss function. The weight reflects the similarity or difference between each sample in the extended dataset and the corresponding sample in the original dataset . Setting the th weight coefficient aims to make the model pay more attention to those samples that are more similar or more important to the known data (the original dataset ), thereby improving the accuracy of label assignment. The specific formula is:
[0059]
[0060] where, represents the similarity between the th sample in the extended dataset and the corresponding sample in the original dataset , is the number of samples in the training data, represents the th element of the extended dataset , represents the th element of the original dataset .
[0061]
[0062] where, represents the th weight coefficient, is a tuning parameter, set manually, used to control the change range of the weight coefficient.
[0063] Next, set the loss function according to the weight coefficients of the samples. The loss function formula is
[0064]
[0065] where, are respectively the true label of the th sample in the extended dataset and the probability distribution predicted by the meta-learning model; represents the th in the extended dataset The cross-entropy loss between the true labels of the samples and the probability distribution predicted by the meta-learning model. The calculation method of the cross-entropy loss adopted in this embodiment is .
[0066] Continuously update the weight coefficients in the meta-learning model , train the meta-learning model, use the manually labeled data as the training data, the training data includes voltage data and the corresponding positive / abnormal probability, the input of the meta-learning model is the voltage data, the output is the normal / abnormal probability, randomly initialize the parameter set of the meta-learning model, for each iteration, use the current weight coefficient and the parameter set to calculate the value of the loss function, continuously update the parameter set of the meta-learning model to minimize the loss function value, finally obtain the optimal parameter set, after obtaining the optimal parameter set, the normal / abnormal probability of each unlabeled data can be predicted by the trained meta-learning model, and classify it as its final label according to the state with the largest probability in the positive / abnormal probability.
[0067] S3. Input the labeled extended data set and the original data set into the LSTM model, output the category of the electricity meter status, train the LSTM model, and input the real-time collected electricity meter data into the trained LSTM model to obtain the evaluation result of the electricity meter operation status; the specific process is as follows:
[0068] As Figure 3 shown, the LSTM model includes several gating units such as the forget gate, input gate, and output gate. The gating units help manage the information flow and determine which information should be retained or forgotten over time. Among them, the purpose of the forget gate is to determine which information to discard from the cell state. The formula for the forget gate at the current time step is , where is the logical function, is the first weight matrix, is the hidden state at the previous time step, the original data set at the current time step, and the extended data set at the current time step, is the first bias term;
[0069] The relevant formula for the input gate at the current time step is , where is the Sigmoid activation function of the input gate, is the second weight matrix, is the second bias term;
[0070] The candidate cell state at the current time step is , where is the weight matrix of the candidate cell state, is the bias term of the candidate cell state, and tanh is the hyperbolic tangent activation function;
[0071] Combine the results of the forget gate and the input gate to update the cell state, and the cell state at the current time step is , where is the cell state at the previous time step; Figure 3 Adding a "d" in the circle in represents the multiplication symbol.
[0072] The output gate determines the output hidden state, and the formula for the output gate at the current time step is , and the hidden state output by the output gate at the current time step is , where is the Sigmoid activation function of the output gate, is the weight matrix of the output gate, is the bias of the output gate;
[0073] The labeled extended dataset and the original dataset are input into the LSTM model, and the hidden state at the last time step generated by the LSTM model is saved as the feature representation, and then sent to a third fully connected layer with a softmax activation function to output the electricity meter status category. Using the above labeled extended dataset and the original dataset to train the LSTM model, stop training when the preset number of iterations is reached or the loss function of the LSTM model is minimized, and obtain the trained LSTM model. Input the real-time collected electricity meter data into the trained LSTM model to obtain the evaluation result of the electricity meter operation status.
[0074] Through the above technical solution, the present invention inputs the original dataset into the data generation model to implement data augmentation and obtain the extended dataset , and then label the generated extended dataset, so that the generated extended dataset can be used for subsequent training of the LSTM model. Overall, the number of data samples is expanded, the problem of small sample data volume is solved, which is helpful for training the model, and the accuracy is improved in subsequent model training and prediction, and the operation status of the electricity meter is accurately identified.
[0075] Embodiment 2
[0076] Based on Embodiment 1, Embodiment 2 of the present invention further provides an electricity meter operation status evaluation system under small sample conditions, including:
[0077] A data augmentation module, which is used to collect the electricity meter operation data and its corresponding electricity meter operation status evaluation results to construct the original dataset , and the original dataset Convert to an N× ×D three-dimensional matrix, where N is the batch size, is the time window size, and D is the feature dimension, which are two dimensions of amplitude and phase respectively; input the N× ×D three-dimensional matrix into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map; pass the multi-channel weighted fusion feature map through a normalization layer and a first fully connected layer in sequence to output a latent space vector; decompose the latent space vector into a trend component and three periodic components, and flatten and concatenate the outputs of the trend component and the three periodic components into two-dimensional data to obtain an extended dataset ;
[0078] Data labeling module, used to label a part of the data in the extended dataset as training data, use the training data to train the meta-learning model, and input the extended dataset into the trained meta-learning model to obtain the labeling result of the extended dataset ;
[0079] State evaluation module, used to input the labeled extended dataset and the original dataset into the LSTM model, output the electricity meter state category, train the LSTM model, and input the real-time collected electricity meter data into the trained LSTM model to obtain the electricity meter operation state evaluation result.
[0080] More specifically, convert the original dataset to an N× ×D three-dimensional matrix, including:
[0081] With the time window size being , and the step size being , convert the original dataset into an N× ×D three-dimensional matrix, where , , represents the data length of the original dataset , represents the i-th data of the amplitude dimension data in the original dataset , represents the mean of the amplitude dimension data in the original dataset , represents the standard deviation of the amplitude dimension data in the original dataset .
[0082] More specifically, for N× The three-dimensional matrix of ×D is input into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map, including:
[0083] The multi-head self-attention mechanism includes h attention heads. The branch vector of the j-th attention head includes a query vector , a key vector , and a value vector , where j = 1, 2... h; calculate the attention weight of the j-th attention head, where represents the transpose of the matrix, represents the dimension of the key vector , represents function;
[0084] Multiply with to obtain weight feature vectors of h channels, multiply it with the three-dimensional matrix of N× ×D to obtain the original feature maps of h channels, then adopt convolution and average pooling operations to obtain the pooled weight feature maps of h channels, and then add the original feature maps of h channels and the pooled weight feature maps using the tensor broadcast mechanism to obtain the weighted feature maps of h channels. Finally, fuse these h-channel weighted feature maps through the second fully connected layer to obtain the multi-channel weighted fusion feature map.
[0085] More specifically, the latent space vector S is decomposed into a trend component and three periodic components, including:
[0086] Use Tucker decomposition to decompose the latent space vector S into a core tensor and three factor matrices, namely:
[0087]
[0088] where is the core tensor. By performing low-pass filtering on the core tensor , it is used as the trend component; , , are factor matrices corresponding to the batch size N, the time window size and the feature dimension D respectively, and are used as periodic components.
[0089] Specifically, training the meta-learning model using the training data in the data labeling module means inputting the training data into the meta-learning model and training the meta-learning model. When the value of the loss function is the smallest, stop training to obtain the trained meta-learning model; where the formula of the loss function is:
[0090]
[0091] Among them, is the number of samples in the training data, are respectively the true label of the th sample in the extended data set and the probability distribution predicted by the meta-learning model, represents the cross-entropy loss between the true label of the th sample in the extended data set and the probability distribution predicted by the meta-learning model; represents the th weight coefficient, and , is a regularization parameter, represents the similarity between the th sample in the extended data set and the corresponding sample in the original data set , and , represents the th element of the extended data set , represents the th element of the original data set .
[0092] Specifically, the state evaluation module is further configured to:
[0093] The LSTM model includes a forget gate, an input gate, and an output gate. The formula for the forget gate at the current time step is , where is the logistic function, is the first weight matrix, is the hidden state at the previous time step, the original data set at the current time step, and the extended data set at the current time step, is the first bias term;
[0094] The formula related to the input gate at the current time step is , where is the Sigmoid activation function of the input gate, is the second weight matrix, is the second bias term;
[0095] The candidate cell state at the current time step is , where is the weight matrix of the candidate cell state, is the bias term of the candidate cell state, and tanh is the hyperbolic tangent activation function;
[0096] Combine the results of the forget gate and the input gate to update the cell state, and the cell state at the current time step is , where is the cell state at the previous time step;
[0097] The formula for the output gate at the current time step is , and the hidden state output by the output gate at the current time step is , where is the Sigmoid activation function of the output gate, is the weight matrix of the output gate, is the bias of the output gate;
[0098] Save the hidden state of the last time step generated by the LSTM model as the feature representation and send it into a third fully connected layer with a softmax activation function to output the category of the electricity meter state.
[0099] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating the operating status of an electric energy meter under small sample conditions, characterized in that: include: S1. Collect the operation data of the electric energy meter and its corresponding operation status evaluation results to construct the original data set, and convert the original data set into N× ×D three-dimensional matrix, where N is the batch size, is the time window size, D is the feature dimension, which is the amplitude and phase dimensions respectively; N× ×D three-dimensional matrix is input into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map; the multi-channel weighted fusion feature map is passed through a normalization layer and a first fully connected layer in sequence to output a latent space vector; the latent space vector is decomposed into a trend component and three period components, and the latent space vector is decomposed into a core tensor and three factor matrices using Tucker decomposition. The core tensor is low-pass filtered as a trend component; the three factor matrices correspond to the batch size N, the time window size and the factor matrix of feature dimension D as the periodic component; the outputs of the trend component and the three periodic components are flattened and concatenated into two-dimensional data to obtain an extended data set; S2, labeling a portion of the data in the extended data set as training data, using the training data to train the meta-learning model, and inputting the extended data set into the trained meta-learning model to obtain the labeling result of the extended data set; S3. Input the labeled expanded data set and the original data set into the LSTM model, output the electricity meter status category, train the LSTM model, and input the real-time collected electricity meter data into the trained LSTM model to obtain the electricity meter operation status evaluation result.
2. The method for evaluating the operating status of an electric energy meter under small sample conditions according to claim 1 is characterized in that: Convert the original dataset into N× ×D three-dimensional matrix, including: , the step length is , transforming the original data set into an N× A three-dimensional matrix of type .
3. The method for evaluating the operating status of an electric energy meter under small sample conditions according to claim 2 is characterized in that: N× The three-dimensional matrix of ×D is input into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map, including: The multi-head self-attention mechanism includes h attention heads, and the branch vector of the j-th attention head includes the query vector , key vector , value vector , j=1, 2...h; according to the query vector and key vector Calculate the attention weight of the jth attention head ;Will and Perform matrix multiplication to obtain the weight feature vector of h channels and add it to N× ×D three-dimensional matrices are multiplied to obtain the original feature maps of h channels, and then convolution and average pooling operations are performed to obtain the pooled weighted feature maps of h channels. The original feature maps of h channels and the pooled weighted feature maps are then added together using the tensor broadcast mechanism to obtain h channel weighted feature maps. Finally, the h channel weighted feature maps are fused through the second fully connected layer to obtain a multi-channel weighted fused feature map.
4. The method for evaluating the operating status of an electric energy meter under small sample conditions according to claim 1 is characterized in that: In S2, using training data to train the meta-learning model means inputting the training data into the meta-learning model, training the meta-learning model, and stopping the training when the value of the loss function is the smallest, thereby obtaining a trained meta-learning model. The similarity between the sample and the corresponding sample in the original data set and the adjustment parameters are calculated. weight coefficients, and then use the first The cross entropy loss between the true label of the sample and the probability distribution predicted by the meta-learning model is The loss function is constructed by multiplying the weight coefficients.
5. The method for evaluating the operating status of an electric energy meter under small sample conditions according to claim 1, characterized in that S3 include: The LSTM model includes a forget gate, an input gate, and an output gate, wherein the formula of the forget gate at the current time step is the product of the hidden state of the previous time step, the original data set of the current time step, and the expanded data set of the current time step with the first weight matrix plus the first bias term and then multiplied by the logic function; the formula of the input gate at the current time step is the product of the hidden state of the previous time step, the original data set of the current time step, and the expanded data set of the current time step with the second weight matrix plus the second bias term and then multiplied by the Sigmoid activation function of the input gate; the candidate cell state at the current time step is the hidden state of the previous time step, the original data set of the current time step, and the weight matrix of the candidate cell state, and then multiplied by the second bias term. The bias term of the candidate cell state is input as the input value into the hyperbolic tangent activation function; the result of the forget gate of the current time step is multiplied by the cell state of the previous time step, and then added to the product of the input gate of the current time step and the candidate cell state of the current time step to obtain the cell state of the current time step; the formula of the output gate of the current time step is the product of the hidden state of the previous time step, the original data set of the current time step and the weight matrix of the output gate, plus the bias of the output gate, and then multiplied by the Sigmoid activation function of the output gate; the hidden state of the last time step generated by the LSTM model is saved as the feature representation, sent to a third fully connected layer with a softmax activation function, and the energy meter state category is output.
6. A system for evaluating the operating status of an electric energy meter under small sample conditions, characterized in that: include: The data enhancement module is used to collect the operation data of the electric energy meter and its corresponding electric energy meter operation status evaluation results to construct the original data set, and convert the original data set into N× ×D three-dimensional matrix, where N is the batch size, is the time window size, D is the feature dimension, which is the amplitude and phase dimensions respectively; N× ×D three-dimensional matrix is input into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map; the multi-channel weighted fusion feature map is passed through a normalization layer and a first fully connected layer in sequence to output a latent space vector; the latent space vector is decomposed into a trend component and three period components, and the latent space vector is decomposed into a core tensor and three factor matrices using Tucker decomposition. The core tensor is low-pass filtered as a trend component; the three factor matrices correspond to the batch size N, the time window size and the factor matrix of feature dimension D as the periodic component; the outputs of the trend component and the three periodic components are flattened and concatenated into two-dimensional data to obtain an extended data set; The data labeling module is used to label a part of the data in the extended data set as training data, use the training data to train the meta-learning model, and input the extended data set into the trained meta-learning model to obtain the labeling result of the extended data set; The status evaluation module is used to input the marked extended data set and the original data set into the LSTM model, output the status category of the electric energy meter, train the LSTM model, and input the real-time collected electric energy meter data into the trained LSTM model to obtain the electric energy meter operation status evaluation result.
7. The system for evaluating the operation status of an electric energy meter under small sample conditions according to claim 6 is characterized in that: Convert the original dataset into N× ×D three-dimensional matrix, including: , the step length is , transforming the original data set into an N× A three-dimensional matrix of type .
8. The system for evaluating the operation status of an electric energy meter under small sample conditions according to claim 7 is characterized in that: N× The three-dimensional matrix of ×D is input into the multi-head self-attention mechanism to obtain a multi-channel weighted fusion feature map, including: The multi-head self-attention mechanism includes h attention heads, and the branch vector of the j-th attention head includes the query vector , key vector , value vector , j=1, 2...h; according to the query vector and key vector Calculate the attention weight of the jth attention head ;Will and Perform matrix multiplication to obtain the weight feature vector of h channels and add it to N× ×D three-dimensional matrices are multiplied to obtain the original feature maps of h channels, and then convolution and average pooling operations are performed to obtain the pooled weighted feature maps of h channels. The original feature maps of h channels and the pooled weighted feature maps are then added together using the tensor broadcast mechanism to obtain h channel weighted feature maps. Finally, the h channel weighted feature maps are fused through the second fully connected layer to obtain a multi-channel weighted fused feature map.
Citation Information
Patent Citations
Meta-learning model training method, system and equipment based on electrocardiogram diagnosis
CN109700434A
Electric energy metering fault diagnosis method and diagnosis terminal based on small sample learning
CN116842459A