Intelligent electric meter fault classification method based on global feature association mining
Through the global feature correlation mining and cross-attention mechanism, the sample imbalance problem of smart meter fault classification is solved, the classification accuracy and recall rate are improved, the operation and maintenance costs are reduced, and the stability of the power grid system is enhanced.
Patent Information
- Application Number
- CN202510681260.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-19
AI Technical Summary
It is difficult for the prior art to effectively classify various faults of smart meters, especially in the case of unbalanced samples, which leads to unreasonable decision-making and high operation and maintenance costs, and the impact of environmental differences in different regions increases the complexity of impacts.
The global feature correlation mining method is adopted, and each sample is spliced into a target-reference combination through the cross attention mechanism to construct multiple target-reference sample groups, calculate the importance of the feature point pair, and give the model the whole domain feature correlation mining ability to improve the accuracy of fault classification.
It improves the accuracy and recall rate of smart meter fault classification, reduces operation and maintenance costs, and improves the stability of the power grid system and the rationality of decision-making.
Smart Images

Figure CN120508858A_ABST
Abstract
Description
Technical field
[0001] The present invention relates to a smart meter fault classification method, and in particular to a smart meter fault classification method based on global feature association mining. [Background Technology]
[0002] Modern power systems feature a strong interactive relationship between users and the power grid. With the continuous increase in user electricity consumption and demand, and the large-scale integration of distributed power sources such as photovoltaic and wind power, traditional power grids face severe challenges. In this environment, smart grids have emerged. As terminal equipment in smart grids, smart meters, in addition to traditional meters for electricity consumption measurement, integrate data display, information transmission, and electricity theft prevention and control, playing a crucial role in supporting the stable operation of the power grid. However, with the widespread adoption of smart meters, the resulting failures have become more sudden, diverse, and complex, severely impacting users' electricity use. Currently, power grid maintenance relies primarily on operations and maintenance personnel, and subjective or biased decision-making during troubleshooting is unavoidable, leading to inappropriate and untimely handling of meter failures. Therefore, accurately classifying various smart meter failure types not only helps power grid companies make timely and appropriate decisions and dispatch experienced personnel to address them, but also plays a significant role in reducing human resource losses and operations and maintenance costs, and improving the overall stability of the power grid system.
[0003] With the widespread deployment of smart meters, the number of manufacturers is increasing. Each manufacturer varies greatly in the electronic components, process techniques, and overall design approaches used during manufacturing. This further increases the complexity of smart meters and inevitably leads to a wider range of fault types during use. Furthermore, given the significant differences in geographical environment and climate across regions, the operational status of smart meters is inevitably affected, further increasing the complexity of their faults. Classifying the various faults of smart meters is a typical machine learning problem. Smart meter fault data is diverse, has high-dimensional features, and exhibits complex coupling relationships between samples. Traditional algorithmic approaches employ methods such as loss design and weight matrix modification to prioritize the characteristics of minority class samples during model training, thereby mitigating the negative impact of data imbalance. However, this approach ignores the important information contained in majority class samples, and the losses and weights required for different tasks vary, making it less generalizable. While data-based approaches achieve balance between samples of different classes by increasing the number of minority classes, they fail to consider the relationships between samples and their neighbors when generating new samples. Although the generative method fits the distribution pattern of the original data set as much as possible through the feature extraction ability of the neural network, it cannot focus on local key features according to the characteristics of the data and the pattern, thereby reducing the overall performance. In order to give the model the ability to autonomously find correlations or key features, the present invention considers introducing a cross-attention mechanism, which enables the model to autonomously query the importance of each sample in the target-reference combination during the training process by splicing any sample with its surrounding neighbors into a target-reference combination, and then deeply mines the correlation information contained between fault samples and between each feature dimension. Based on the correlation relationship between different samples and combined with the known fault sample information, indirect discrimination of the meter fault data to be classified is achieved. The present invention proposes a smart meter fault classification method based on global feature association mining to improve the performance of smart meter fault classification. [Summary of the invention]
[0004] In view of this, the present invention proposes a smart meter fault classification method based on global feature association mining to improve the performance of smart meter fault classification.
[0005] The present invention proposes a smart meter fault classification method based on global feature association mining, which includes the following steps:
[0006] (1) The historical data of smart meters under different fault categories are divided to obtain multiple second-class data sets as input data, specifically:
[0007] The actual fault dataset of smart meters is processed. The samples in this dataset contain 10 characteristic variables, namely: working time, arrival batch number, power supply unit number, electric energy meter category, fault identification month, installation month, province, equipment specification, communication method, and equipment identification; its fault category labels contain 11 categories, namely: appearance fault, metering fault, storage unit fault, processing unit fault, display unit fault, control unit fault, power supply unit fault, communication unit fault, clock unit fault, other faults, and software fault; traverse each category of samples in the dataset, and regard all samples under this category as the minority class sample set, and the remaining fault category samples as the majority class sample set, so as to convert the original dataset into 11 two-category datasets; for each of these two-category datasets, the dataset can be described as follows:
[0008] X i =[X maj ,X min ]
[0009] Among them, X i is any one of the 11 transformed two-class data sets, i∈[1,11]; X min For the second-class dataset X i The minority class sample set in X, where each sample class label is set to 1; maj For the second-class dataset X i The majority class sample set in X, where each sample class label is set to 0; i The size is G i =len(X i ), len(X i ) represents the second-class dataset X i Through the above operations, the original unbalanced multi-classification problem is converted into multiple unbalanced binary classification problems for processing;
[0010] (2) Each individual sample in the original dataset is considered as a target, and the remaining samples are randomly shuffled and spliced to generate a diverse graph sample as a reference, and multiple target-reference sample groups are constructed, specifically:
[0011] Based on any one of the 11 imbalanced two-class datasets obtained in step (1) i , construct differentiated target-reference sample groups:
[0012] Combine all samples in the current data set and splice them into a two-dimensional image sample; in the spliced two-dimensional data, each row is an original single sample, and each column is an original set of feature dimensions; at this time, for ease of understanding, the spliced sample can be regarded as a special "image"; taking any data set as an example, assuming that it has M k-dimensional feature samples, so after splicing it, what is obtained is a two-dimensional image data of size M*k; after obtaining the imaged table data, for each single sample in the original data set, extract it from the image data and use it as a query sample in the subsequent training process; at the same time, for all the remaining samples, randomly shuffle their order to generate N different references; in this way, for each single sample in the previous data set, a set of diversified image sample references of size N can be obtained, which means that the original data set as a whole has been expanded by N times;
[0013] Where M represents the total number of samples in the dataset, k represents the number of features of the samples in the dataset, and N is the expansion factor, which represents the number of two-dimensional references generated for each query sample.
[0014] (3) By calculating the importance of each feature point in the differential graph sample to the task of determining the target sample category through cross-attention, the model is given the ability to mine global feature associations. Specifically:
[0015] Based on any target-reference sample group obtained in step (2), the cross-attention mechanism is used to calculate the correlation between its features and construct the corresponding multi-label trust discrimination network model i And train, the cross attention calculation formula and model optimization goal are:
[0016]
[0017] BCE(P,y)=-[ylog(p)+(1-y)log(1-p)]
[0018] Among them, Q comes from the query, which is obtained by passing the original query sample through a layer of feature extraction network. It contains the feature information that the model needs to pay attention to and represents the key features most relevant to the category in the target sample; K and V both come from the reference sample. In the proposed method, K represents the key value of each feature point of the reference sample, and V represents the abstract feature value of each feature point. The two are obtained by passing the reference sample through two different layers of feature extraction networks; d represents the dimension of the V vector. In this way, the calculated QK similarity matrix can be scaled to the same dimension as V; p represents the prediction result of the model, and y represents the actual true label of the sample.
[0019] (4) During the testing phase, the model analyzes the relationship between the current sample to be tested and multiple sets of known references to reasonably infer its category and obtain the fault category, specifically:
[0020] Based on each multi-label trust discriminant network model trained in step (3) i , for a test sample x test ,The model first calculates the cross attention of the current target sample and the reference sample, determines the sample in the reference that contributes the most to the current sample category inference, and sends this information and the reference to the classification layer at the same time, outputting the category inference of the current target sample;
[0021] Repeat the above process for each of the two-category datasets X i , the corresponding multi-label trust discrimination network model can be obtained i , i is the subscript of the pattern discrimination classifier, i∈[1,11]; for the sample to be tested x test , its predicted label The calculation is as follows:
[0022]
[0023] When the value is i, it means x test The predicted fault category is the i-th fault.
[0024] In the above method, in step (2), the value of N is 300.
[0025] In the above method, in step (3), the network model i The structure is as follows:
[0026]
[0027] Among them, in_features represents the feature dimension of the input data, num_classes represents the number of categories contained in the dataset, Linear() is the linear layer function; Sigmoid() is the activation function.
[0028] The smart meter fault classification method based on global feature association mining improves the accuracy and recall rate of smart meter fault classification.
[0029] It can be seen from the above technical solutions that the present invention has the following beneficial effects:
[0030] In the technical solution implemented by the present invention, each individual sample in the original data set is regarded as a target, and the remaining samples are randomly shuffled and spliced to generate diversified graph samples as references, thereby constructing multiple target-reference sample groups. The random splicing method can achieve exponential expansion of the original data set without introducing noise; the target-reference sample construction method and the two-dimensional graph sample construction method lay the foundation for subsequent global feature association mining. On this basis, the importance of each feature point in the differentiated graph sample to the task of determining the target sample category is calculated through cross-attention, thereby giving the model the ability to mine global feature associations, thereby improving the accuracy and recall rate of the smart meter fault classification results.
Brief Description of the Drawings
[0031] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0032] Figure 1 This is a schematic diagram of the framework flow of the smart meter fault classification method based on global feature association mining proposed in the present invention;
[0033] Figure 2 It is an unbalanced classification method based on diversified graph sample construction and global feature association mining;
[0034] Figure 3 It is an imbalanced data enhancement method based on the construction of diverse graph samples. [Specific implementation method]
[0035] In order to better understand the technical solution of the present invention, the present invention is described in detail below with reference to the accompanying drawings.
[0036] It should be understood that the embodiments of the invention described herein are only a portion of the embodiments of the invention, not all of the embodiments. Based on the embodiments of the invention, all other embodiments obtained by persons of ordinary skill in the art without creative effort are within the scope of protection of the invention.
[0037] An embodiment of the present invention proposes a smart meter fault classification method based on global feature association mining, including: dividing historical data under different fault categories of smart meters to obtain multiple second-category data sets as input data; treating each individual sample in the original data set as a target, randomly shuffling and splicing the remaining samples to generate diversified graph samples as references, and constructing multiple groups of target-reference sample groups; calculating the importance of each feature point in the differentiated graph sample to the task of judging the target sample category through cross-attention, thereby giving the model the ability of global feature association mining; in the testing phase, the model reasonably infers the category of the current sample to be tested by analyzing the correlation relationship between it and multiple groups of known references, and obtains the fault category.
[0038] Figure 1 The figure is a schematic diagram of the framework flow of the smart meter fault classification method based on global feature association mining proposed in the present invention. The method includes the following steps:
[0039] Step 101: historical data of smart meters under different fault categories are divided to obtain multiple second-category data sets as input data, specifically:
[0040] The actual fault dataset of smart meters is processed. The samples in this dataset contain 10 characteristic variables, namely: working hours, arrival batch number, power supply unit number, electric energy meter category, fault identification month, installation month, province, equipment specification, communication method, and equipment identification; its fault category labels contain 11 categories, namely: appearance fault, metering fault, storage unit fault, processing unit fault, display unit fault, control unit fault, power supply unit fault, communication unit fault, clock unit fault, other faults, and software fault; traverse each category of samples in the dataset, and regard all samples under this category as the minority class sample set, and the samples of the remaining fault categories as the majority class sample set, so as to convert the original dataset into 11 two-category datasets; for each of these two-category datasets, the dataset can be described as follows:
[0041] X i =[X maj ,X min ]
[0042] Among them, X i is any one of the 11 transformed two-class data sets, i∈[1,11]; X min For the second-class dataset X i The minority class sample set in X, where each sample class label is set to 1; maj For the second-class dataset X i The majority class sample set in X, where each sample class label is set to 0; i The size is C i =len(Xi ), len(X i ) represents the second-class dataset X i Through the above operations, the original unbalanced multi-classification problem is converted into multiple unbalanced binary classification problems for processing;
[0043] Step 102: Consider each individual sample in the original dataset as a target, randomly shuffle and concatenate the remaining samples to generate a diverse graph sample as a reference, and construct multiple target-reference sample groups, specifically:
[0044] Based on any one X in the 11 imbalanced two-class datasets obtained in step 101 i , construct differentiated target-reference sample groups:
[0045] Combine all samples in the current data set and splice them into a two-dimensional image sample; in the spliced two-dimensional data, each row is an original single sample, and each column is an original set of feature dimensions; at this time, for ease of understanding, the spliced sample can be regarded as a special "image"; taking any data set as an example, assuming that it has M k-dimensional feature samples, so after splicing it, what is obtained is a two-dimensional image data of size M*k; after obtaining the imaged table data, for each single sample in the original data set, extract it from the image data and use it as a query sample in the subsequent training process; at the same time, for all the remaining samples, randomly shuffle their order to generate N different references; in this way, for each single sample in the previous data set, a set of diversified image sample references of size N can be obtained, which means that the original data set as a whole has been expanded by N times;
[0046] Among them, M represents the total number of samples contained in the dataset, k represents the number of features of the samples in the dataset, and N is the expansion multiple, which represents the number of two-dimensional references generated for each query sample.
[0047] Step 103, through cross attention, calculates the importance of each feature point in the difference map sample to the task of determining the target sample category, thereby giving the model the ability to mine global feature associations, specifically:
[0048] Based on any target-reference sample group obtained in step 102, the cross-attention mechanism is used to calculate the correlation between its features and construct the corresponding multi-label trust discrimination network model i And train, the cross attention calculation formula and model optimization goal are:
[0049]
[0050] BCE(p,y)=-[ylog(p)+(1-y)log(1-p)]
[0051] Among them, Q comes from the query, which is obtained by passing the original query sample through a layer of feature extraction network. It contains the feature information that the model needs to pay attention to and represents the key features most relevant to the category in the target sample; K and V both come from the reference sample. In the proposed method, K represents the key value of each feature point of the reference sample, and V represents the abstract feature value of each feature point. The two are obtained by passing the reference sample through two different layers of feature extraction networks; d represents the dimension of the V vector. In this way, the calculated QK similarity matrix can be scaled to the same dimension as V; p represents the prediction result of the model, and y represents the actual true label of the sample.
[0052] Step 104: During the testing phase, the model analyzes the relationship between the current sample to be tested and multiple sets of known references to reasonably infer its category and obtain the fault category, specifically:
[0053] Based on each multi-label trust discrimination network model trained in step 103 i , for a test sample x test ,The model first calculates the cross attention of the current target sample and the reference sample, determines the sample in the reference that contributes the most to the current sample category inference, and sends this information and the reference to the classification layer at the same time, outputting the category inference of the current target sample;
[0054] Repeat the above process for each of the two-category datasets X i , the corresponding multi-label trust discrimination network model can be obtained i , i is the subscript of the pattern discrimination classifier, i∈[1,11]; for the sample to be tested x test , its predicted label The calculation is as follows:
[0055]
[0056] When the value is i, it means x test The predicted fault category is the i-th fault.
[0057] Figure 2This is a schematic diagram of an unbalanced classification method based on the construction of diverse graph samples and global feature association mining. To fully leverage the powerful feature extraction capabilities of deep networks, the original tabular data is spliced into two-dimensional graph data. Each spliced two-dimensional graph data is treated as an image and serves as an input sample for the deep network. Based on this approach, the proposed method transforms the label prediction task, which was previously focused on by algorithmic methods, into a task of analyzing the feature associations between the target sample and the constructed two-dimensional graph sample. Furthermore, the graphical representation of tabular data not only enables the model to focus on the associations between samples during training, but also lays the foundation for the exponential expansion of the sample size. Based on the graph sample construction, a cross-attention mechanism is introduced to calculate the importance between diverse reference feature points and the target sample label under the label prediction objective. The cross-attention mechanism is a feature extraction method in the field of image classification. By treating the test sample as the target sample and the constructed two-dimensional graph sample as the reference, the proposed algorithm introduces the cross-attention mechanism and fully mines the association information between different samples by comparing the associations between the target sample and its reference. During the testing phase, the original training samples were randomly spliced to form a diverse reference group, which was then input into the trained association extraction module together with the same sample to be tested. The prediction results of each reference group were integrated to infer the category of the sample to be tested, and the unbalanced data was effectively and accurately classified without generating new samples to achieve quantitative balance.
[0058] Figure 3 This is a schematic diagram of an unbalanced data augmentation method based on the construction of diversified graph samples. Before training begins, the current dataset is processed first. All samples in the current dataset are combined and spliced into a two-dimensional graph sample. In the spliced two-dimensional data, each row is an original single sample, and each column is an original set of feature dimensions. At this point, for ease of understanding, the spliced sample can be regarded as a special "image". Taking any dataset as an example, assuming that it has M k-dimensional feature samples, after splicing it, a two-dimensional image data of size M*k is obtained; after obtaining the imaged table data, for each single sample in the original dataset, it is extracted from the image data as a query sample in the subsequent training process; at the same time, for all the remaining samples, their order is randomly shuffled to generate N different references; in this way, for each single sample in the previous dataset, a set of diversified graph sample references of size N can be obtained, which means that the original dataset as a whole has been expanded by N times;
[0059] Algorithm 1 is the pseudo code of the tabular data visualization algorithm:
[0060]
[0061] Algorithm 2 is the pseudo code of the attention calculation process:
[0062]
[0063]
[0064] In a specific implementation, a fault history dataset of different smart meter categories was used for testing. The dataset collected data from smart meters covering 11 fault types across 25 provinces. Due to factors such as manual statistics and external conditions, data labels may be mislabeled or missing. Direct use of these labels can cause the model to learn incorrect features. After data cleaning using various techniques, including cluster analysis, missing value imputation, and outlier processing, a total of 15,885 fault sample data items were obtained. To reduce the randomness of the results, the dataset was randomly divided into training and test sets using a fixed random number seed in an 8:2 ratio.
[0065] Table 1 Datasets used in specific examples
[0066]
[0067] To verify the effectiveness of the proposed algorithm, this embodiment of the present invention uses the sample sampling methods SMOTE, Borderline-SMOTE, RCSMOTE, LDAS, DEBOHID, and SMOTE-NaN-DE, and the sample generation methods CWGAN-GP and ADA-INCVAE for comparative experiments. Considering the large sample size and complex classification characteristics of the smart meter fault dataset, the comparative method uses RF as the classifier to verify the sample balance effect. This embodiment of the present invention is represented in the table as DGSC-FDFA (An imbalanced classification framework via diversified graph-sample construction and full-domain feature association mining).
[0068] The embodiments of the present invention use the macro-F1 and G-mean metrics to evaluate the classification performance of the algorithm. The macro-F1 metric is the arithmetic mean of the F1-measure for each fault category and is used to comprehensively evaluate the model's precision and recall for each category. The G-mean metric is the geometric mean of the recall rates for each category and is used to evaluate the model's recall for each category. Both Macro-F1 and G-mean values range from 0 to 1, with larger values indicating better classification performance.
[0069] Table 2 shows a comparison of the experimental results of the present invention and mainstream oversampling methods for various smart meter fault categories using the F1-measure and macro-F1 metrics. Table 3 also shows a comparison of the experimental results for recall and G-mean metrics for each category. It can be seen that the present invention's smart meter fault classification method based on nearest neighbor sample pairs achieves higher F1-measure and recall rates than other methods for most categories, and achieves the highest macro-F1 and G-mean. The combined results of Tables 2 and 3 demonstrate that the present invention's method achieves better smart meter fault sample balancing than data-level balancing methods, achieving higher classification accuracy and recall.
[0070] Table 2 Experimental results of DGSC-FDFA and data-level methods on F1-measure and macro-F1 indicators under various fault categories of smart meters
[0071]
[0072] Table 3 Experimental results of DGSC-FDFA and data-level methods on recall rate and G-mean index for various fault categories of smart meters
[0073]
[0074] In order to further verify the effectiveness of the proposed method, the algorithm-level methods RF, GBDT, BRAF
[51] , CSSVM, DPHS-MDS and HUE were selected for comparative experiments. The experimental results of the embodiment of the present invention and the mainstream algorithm-level methods on the F1-measure and macro-F1 indicators under various fault categories of smart meters are shown in Table 4, and the experimental results on the recall rate and G-mean indicators of each category are shown in Table 5. It can be seen that the smart meter fault classification method constructed based on the nearest neighbor sample pairs of the present invention has achieved F1-measure and recall rates that exceed those of other methods in most categories, and has achieved the highest macro-F1 and G-mean. Combining the results of Tables 4 and 5, it is shown that the smart meter fault sample balancing effect of the embodiment of the present invention is better than that of the algorithm-level method, and can achieve higher classification accuracy and recall rate.
[0075] Table 4 Experimental results of DGSC-FDFA and algorithm-level methods on F1-measure and macro-F1 indicators under various fault categories of smart meters
[0076]
[0077]
[0078] Table 5 Experimental results of DGSC-FDFA and algorithm-level methods on recall rate and G-mean index for various fault categories of smart meters
[0079]
[0080] A large number of comparisons with mainstream oversampling methods and deep learning sample generation methods show that the present invention regards each individual sample in the original data set as a target, randomly shuffles the remaining samples and splices them to generate diversified graph samples as references, and constructs multiple groups of target-reference sample groups. The random splicing method can achieve exponential expansion of the original data set without introducing noise; the target-reference sample construction method and the two-dimensional graph sample construction method lay the foundation for subsequent global feature association mining. On this basis, the importance of each feature point in the differentiated graph sample to the task of judging the target sample category is calculated through cross-attention, thereby giving the model the ability to mine global feature associations, which can effectively improve the accuracy and recall rate of faulty meter classification.
[0081] In summary, the embodiments of the present invention have the following beneficial effects:
[0082] In the technical solution implemented by the present invention, historical data of smart meters under different fault categories are divided to obtain multiple second-category data sets as input data; each individual sample in the original data set is regarded as a target, and the remaining samples are randomly shuffled and spliced to generate diversified graph samples as references, and multiple groups of target-reference sample groups are constructed; the importance of each feature point in the differentiated graph sample to the task of judging the target sample category is calculated through cross-attention, thereby giving the model the ability to mine global feature associations; in the testing phase, the model reasonably infers the category of the current sample to be tested by analyzing the correlation relationship between it and multiple groups of known references, and obtains the fault category; experimental results in actual smart meter fault data sets prove that the proposed invention has achieved the best results on the vast majority of second-category data sets and has better universality.
[0083] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A smart meter fault classification method based on global feature association mining, characterized by: The steps include: (1) The historical data of smart meters under different fault categories are divided to obtain multiple second-class data sets as input data, specifically: The actual fault data set of smart meters was processed. The samples in this data set contained 10 characteristic variables, namely: working hours, arrival batch number, power supply unit number, meter type, fault identification month, installation month, province, equipment specifications, communication method, and equipment identification. The fault category labels include 11 categories: appearance fault, metering fault, storage unit fault, processing unit fault, display unit fault, control unit fault, power supply unit fault, communication unit fault, clock unit fault, other faults, and software fault. We traverse each category of samples in the dataset and use all samples in that category as the minority class sample set, and samples of the remaining fault categories as the majority class sample set. The original dataset is converted into 11 two-category datasets. For each of these two-category datasets, the dataset can be described as follows: X i =[X maj ,X min ] Among them, X i is any one of the 11 transformed two-class data sets, i∈[1,11]; X min For the second-class dataset X i The minority class sample set in X, where each sample class label is set to 1; maj For the second-class dataset X i The majority class sample set in X, where each sample class label is set to 0; i The size is G i =len(X i ), len(X i ) represents the second-class dataset X i Through the above operations, the original unbalanced multi-classification problem is converted into multiple unbalanced binary classification problems for processing; (2) Each individual sample in the original dataset is considered as a target, and the remaining samples are randomly shuffled and spliced to generate a diverse graph sample as a reference, and multiple target-reference sample groups are constructed, specifically: Based on any one of the 11 imbalanced two-class datasets obtained in step (1) i , construct differentiated target-reference sample groups: Combine all samples in the current dataset and splice them into a two-dimensional image sample; in the spliced two-dimensional data, each row is an original single sample, and each column is an original set of feature dimensions; at this time, for ease of understanding, the spliced sample can be regarded as a special "image"; taking any dataset as an example, assuming it has M k-dimensional feature samples, so after splicing it, what is obtained is a two-dimensional image data of size M*k; after obtaining the imaged table data, for each single sample in the original dataset, extract it from the image data and use it as a query sample in the subsequent training process; at the same time, for all the remaining samples, randomly shuffle their order to generate N different references; in this way, for each single sample in the previous dataset, a set of diversified image sample references of size N can be obtained, which means that the original dataset as a whole has been expanded by N times; Where M represents the total number of samples in the dataset, k represents the number of features of the samples in the dataset, and N is the expansion factor, which represents the number of two-dimensional references generated for each query sample. (3) By calculating the importance of each feature point in the differential graph sample to the task of determining the target sample category through cross-attention, the model is given the ability to mine global feature associations. Specifically: Based on any target-reference sample group obtained in step (2), the cross-attention mechanism is used to calculate the correlation between its features and construct the corresponding multi-label trust discrimination network model i And train, the cross attention calculation formula and model optimization goal are: BCE(p,y)=-[ylog(p)+(1-y)log(1-p)] Among them, Q comes from the query, which is obtained by passing the original query sample through a layer of feature extraction network. It contains the feature information that the model needs to pay attention to and represents the key features most relevant to the category in the target sample; K and V both come from the reference sample. In the proposed method, K represents the key value of each feature point of the reference sample, and V represents the abstract feature value of each feature point. The two are obtained by passing the reference sample through two different layers of feature extraction networks; d represents the dimension of the V vector. In this way, the calculated OK similarity matrix can be scaled to the same dimension as V; p represents the prediction result of the model, and y represents the actual sample true label; (4) During the testing phase, the model analyzes the relationship between the current sample to be tested and multiple sets of known references to reasonably infer its category and obtain the fault category, specifically: Based on each multi-label trust discriminant network model trained in step (3) i , for a test sample x test ,The model first calculates the cross attention of the current target sample and the reference sample, determines the sample in the reference that contributes the most to the current sample category inference, and sends this information and the reference to the classification layer at the same time, outputting the category inference of the current target sample; Repeat the above process for each of the two-category datasets X i , the corresponding multi-label trust discrimination network model can be obtained i , i is the subscript of the pattern discrimination classifier, i∈[1,11]; for the sample to be tested x test , its predicted label The calculation is as follows: When the value is i, it means x test The predicted fault category is the i-th fault.
2. The smart meter fault classification method based on multi-label confidence comparison according to claim 1 is characterized in that: In step (2), the value of N is 300.
3. The smart meter fault classification method based on multi-label confidence comparison according to claim 1 is characterized in that: In step (3), the network model i The structure is as follows: Among them, in_features represents the feature dimension of the input data, num_classes represents the number of categories contained in the dataset, Linear() is the linear layer function; Sigmoid() is the activation function.