Method and device for repairing abnormal verification data of verification line of electric energy meter
By combining Euclidean distance and Manhattan distance to improve the K-nearest neighbor algorithm, and using the Golden Jackal optimization algorithm to optimize parameters, the problems of low efficiency and poor adaptability in the data anomaly repair of electricity meter calibration lines are solved, achieving high-precision and adaptive data repair effect, and improving the accuracy and reliability of electricity meter calibration lines.
Patent Information
- Application Number
- CN202510682813.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-17
Smart Images

Figure CN120804533A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of electric energy metering, and particularly relates to an abnormal calibration data repairing method and device applied to an electric energy meter calibration line. BACKGROUND
[0002] The electric energy meter calibration line is a core facility for electric energy metering quality management and control, and its operation state directly affects the reliability and traceability of calibration data. In an automatic calibration system, the equipment system not only bears the accurate discrimination of the operation state of the electric energy meter, but also is a key link for guaranteeing the accuracy of the metering. However, in the actual operation process, the performance of the equipment components of the calibration line may drift due to environmental disturbance, mechanical wear and tear, aging of electronic components and other factors. Some fault chain reactions may directly cause quality problems such as time sequence missing and abnormal deviation of characteristic values of the calibration data, so that the subsequent equipment state evaluation based on data driving faces severe challenges. More notably, the multi-dimensional data generated by the calibration line is not isolated, and there is strong space-time correlation between the current / voltage sampling sequence, error curve, temperature vibration monitoring value and other parameters. The complex coupling relationship makes it difficult for the traditional data repairing method to accurately reconstruct the complete characteristics of the original signal, further increasing the difficulty of intelligent diagnosis of the equipment state. SUMMARY
[0003] The application provides an abnormal calibration data repairing method and device applied to an electric energy meter calibration line, to solve the problems of low efficiency, poor adaptability and insufficient complex data pattern recognition capability in the traditional electric energy meter calibration line data abnormality repairing method.
[0004] In a first aspect, the application provides an abnormal calibration data repairing method applied to an electric energy meter calibration line, comprising:
[0005] acquiring electric energy meter calibration line data, and dividing missing values in the electric energy meter calibration line data according to a missing value classification standard;
[0006] improving the K nearest neighbor algorithm by using a combination of Euclidean distance and Manhattan distance, obtaining an improved distance K nearest neighbor algorithm, and calculating each missing value after division by using the improved distance K nearest neighbor algorithm to obtain a weighted distance corresponding to each missing value;
[0007] optimizing the improved distance K nearest neighbor algorithm by using an improved golden jackal optimization algorithm to obtain an optimal parameter combination of the improved distance K nearest neighbor algorithm;
[0008] calculating a repairing value corresponding to each missing value according to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination.
[0009] In a second aspect, the application provides a device for repairing abnormal calibration data of a power meter calibration line, comprising:
[0010] a data acquisition module configured to acquire power meter calibration line data and divide missing values in the power meter calibration line data according to a missing value classification standard;
[0011] a distance calculation module configured to improve a K-nearest neighbor algorithm by using a combination of Euclidean distance and Manhattan distance, obtain an improved distance K-nearest neighbor algorithm, and calculate each missing value after division by using the improved distance K-nearest neighbor algorithm to obtain a weighted distance corresponding to each missing value;
[0012] a parameter calculation module configured to optimize the improved distance K-nearest neighbor algorithm by using an improved golden jackal optimization algorithm to obtain an optimal parameter combination of the improved distance K-nearest neighbor algorithm;
[0013] a repair value calculation module configured to calculate a repair value corresponding to each missing value according to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination.
[0014] The application provides a method and device for repairing abnormal calibration data of a power meter calibration line. The power meter calibration line data is acquired, and missing values in the power meter calibration line data are divided according to a missing value classification standard. A K-nearest neighbor algorithm is improved by using a combination of Euclidean distance and Manhattan distance to obtain an improved distance K-nearest neighbor algorithm. Each missing value after division is calculated by using the improved distance K-nearest neighbor algorithm to obtain a weighted distance corresponding to each missing value. The improved distance K-nearest neighbor algorithm is optimized by using an improved golden jackal optimization algorithm to obtain an optimal parameter combination of the improved distance K-nearest neighbor algorithm. A repair value corresponding to each missing value is calculated according to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination. The weighted distance corresponding to each missing value is calculated by improving the K-nearest neighbor algorithm by using a combination of Euclidean distance and Manhattan distance, which can significantly improve the robustness and adaptability of data repair, and further improve the accuracy and reliability of the power meter calibration line calibration process. The improved distance K-nearest neighbor algorithm is optimized by using an improved golden jackal optimization algorithm, which can effectively overcome the limitations of traditional parameter selection depending on experience or exhaustive search, and further significantly improve the generalization ability of dynamic abnormal data of the power meter calibration line, achieving high-precision and self-adaptive parameter optimization without manual intervention. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0016] Figure 1 is a flowchart of the method for repairing abnormal calibration data of an electric energy meter calibration line provided by the embodiments of the present application;
[0017] Figure 2 is a flowchart of the K nearest neighbor algorithm optimized by the golden jackal optimization algorithm provided by the embodiments of the present application;
[0018] Figure 3 is a structural schematic diagram of the device for repairing abnormal calibration data of an electric energy meter calibration line provided by the embodiments of the present application. DETAILED DESCRIPTION
[0019] In the following description, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments without these specific details. In other instances, well-known systems, devices, circuits, and methods have not been described in detail in order to avoid obscuring the description of the present application.
[0020] In order to make the purpose, technical solutions and advantages of the present application clearer, specific embodiments will be described below with reference to the drawings.
[0021] To solve the problems of low efficiency, poor adaptability, and insufficient ability to recognize complex data patterns in the traditional electric energy meter calibration line data anomaly repair method, the present application proposes a method for repairing abnormal calibration data of an electric energy meter calibration line, aiming to improve the accuracy and reliability of the calibration process of the electric energy meter calibration line. By combining the local data correlation characteristics of the K nearest neighbor algorithm (KNN) and the global search ability of the golden jackal optimization algorithm (GJO), the present application realizes intelligent repair of abnormal data and improves the quality and reliability of the electric energy meter calibration line data. Specifically, for the stronger correlation and time sequence of the electric energy meter calibration line data, the missing values and abnormal values in the calibration data are repaired. For the calculation of the distance of the nearest neighbor value in the repair method, a combination of multiple distance measurement methods is used, and the repair algorithm is optimized and improved using a bionic algorithm to realize automatic optimization of the K value. The present application can effectively reduce manual intervention, adapt to high-dimensional and multi-noise data characteristics in a big data environment, provide data support for subsequent calibration line and electric energy meter state evaluation, and at the same time reduce the evaluation error and economic loss risk caused by data anomalies.
[0022] Figure 1 The implementation flowchart of the method for repairing abnormal calibration data of an electric energy meter calibration line provided by the embodiments of the present application is described as follows:
[0023] In step 101, electric energy meter calibration line data is acquired, and missing values in the electric energy meter calibration line data are classified according to a missing value classification standard.
[0024] In the embodiments of the present application, electric energy meter calibration line data is collected in real time, including but not limited to error, voltage, current and power factor of the electric energy meter calibration line. Then, the electric energy meter calibration line data is preprocessed, that is, after the electric energy meter calibration line data is cleaned, the missing values in the electric energy meter calibration line data are classified according to a missing value classification standard. The missing value classification standard includes error missing type, voltage missing type, current missing type and power factor missing type.
[0025] In the embodiments of the present application, the preprocessing can include missing value identification, extreme abnormal value processing, data standardization, irrelevant feature elimination and relevant feature reservation.
[0026] In addition, the purpose of classifying the missing values in the electric energy meter calibration line data according to the missing value classification standard in the embodiments of the present application is to classify the data. Specifically, the electric energy meter calibration line data containing missing values is classified into specific categories such as error missing type, voltage missing type, current missing type and power factor missing type according to the type of missing values.
[0027] In a possible implementation manner, after the missing values in the electric energy meter calibration line data are classified according to the missing value classification standard, the method can further include:
[0028] According to a preset abnormality detection threshold, extreme abnormal values in the electric energy meter calibration line data are eliminated;
[0029] The missing values after the elimination of the extreme abnormal values are classified according to the missing value classification standard.
[0030] Optionally, after the missing values in the electric energy meter calibration line data are classified, the preset abnormality detection threshold of each parameter is defined to simply identify and eliminate the extreme abnormal values in the electric energy meter calibration line data. The effective data retained after the abnormal value processing is also classified according to the missing value classification standard, so as to maintain the consistency of the classification system of the electric energy meter calibration line data.
[0031] In a possible implementation manner, after the missing values in the electric energy meter calibration line data are classified according to the missing value classification standard, the method can further include:
[0032] The divided electric energy meter calibration line data is subjected to data standardization processing, and after the data standardization processing, the irrelevant features of different categories of missing values are sequentially removed and the relevant features are retained.
[0033] Optionally, after the electric energy meter calibration line data is divided, the divided electric energy meter calibration line data is subjected to data standardization processing, that is, all features are standardized to have a mean value of 0 and a variance of 1, so as to eliminate the influence of dimensional difference on distance calculation. The process of data standardization is as follows:
[0034] The mean value μ is calculated as shown in formula (1), and the standard deviation σ is calculated as shown in formula (2):
[0035]
[0036] Wherein, μ is the mean value, σ is the standard deviation, x i is a sample data point, N is the total amount of samples, and i is the sample serial number.
[0037] The standardized data is z, that is:
[0038]
[0039] Wherein, X is the sample data.
[0040] The normalized data is subjected to normalization processing, as shown in formula (4):
[0041]
[0042] Wherein, X ′ is the normalized data, X min is the minimum value of the sample data, and X max is the maximum value of the sample data.
[0043] In step 102, the K-neighbor algorithm is improved by using the combination of Euclidean distance and Manhattan distance, to obtain the K-neighbor algorithm of improved distance, and the K-neighbor algorithm of improved distance is used to calculate each missing value after division, to obtain the weighted distance corresponding to each missing value.
[0044] The K-neighbor algorithm of improved distance is obtained by combining the measurement methods of Euclidean distance and Manhattan distance, and then the K-neighbor algorithm of improved distance is used to calculate the distance of each missing value divided in step 101, to obtain the weighted distance corresponding to each missing value.
[0045] In the embodiment of the application, the K-neighbor algorithm of improved distance is obtained by combining the measurement methods of Euclidean distance and Manhattan distance, and then the K-neighbor algorithm of improved distance is used to calculate the distance of each missing value divided in step 101, to obtain the weighted distance corresponding to each missing value.
[0046] The distance calculation method of the K-neighbor algorithm used in the embodiment of the application is changed to the distance calculation method of the combination of Euclidean distance and Manhattan distance, which significantly improves the robustness and adaptability of data repair compared with the single distance calculation method. The Euclidean distance captures the global geometric features and effectively preserves the overall trend of the data, and the Manhattan distance enhances the stability of local dimensional changes and suppresses high-dimensional noise interference. The dynamic weight mechanism can adaptively adjust the contribution proportion of the two, avoiding the limitations of single distance and balancing global sensitivity and local noise resistance. At the same time, the embodiment of the application optimizes the calculation efficiency and considers the real-time requirements of edge devices, which is suitable for dynamic data repair scenarios under complex working conditions of electric energy meters, and provides reliable support for high-precision repair of abnormal data.
[0047] In a possible implementation, the calculation of each missing value after division by the K-neighbor algorithm of improved distance can include:
[0048] The Euclidean distance in the K-neighbor algorithm of improved distance is used to calculate each missing value after division to obtain the first distance corresponding to each missing value;
[0049] The Manhattan distance in the K-neighbor algorithm of improved distance is used to calculate each missing value after division to obtain the second distance corresponding to each missing value;
[0050] The first distance and the second distance corresponding to each missing value are used to obtain the weighted distance corresponding to each missing value.
[0051] The electric energy meter calibration line data usually contains multi-dimensional time sequence characteristics such as error, voltage, current, power factor, etc. The Euclidean distance is good at capturing global trends and can effectively reflect the systematic changes of the electric energy meter in the calibration process. However, for local fluctuations and abnormal noise in the data, the use of a single distance formula will cause a large error in the calculation of the nearest neighbor value. The Manhattan distance has stronger robustness to fluctuations and noise, and can avoid excessive interference of single-dimensional noise on distance calculation. By dynamically weighting the fusion of the two distances, the trend consistency of the key parameters in electric energy measurement can be retained, and high-frequency noise can be suppressed, so as to realize the balance between "global trend adaptation" and "local noise correction" in the repair process.
[0052] In the use of the K nearest neighbor algorithm using the improved distance, the weighted Euclidean-Manhattan distance is used, that is, the Euclidean distance and the Manhattan distance are combined. The electric energy meter calibration line data contains a large amount of continuous and discrete data. The Euclidean distance can better calculate the continuous data, and the Manhattan distance can better calculate the discrete data, that is, the first distance is calculated by using the Euclidean distance, and the second distance is calculated by using the Manhattan distance. The calculation formulas of the Euclidean distance and the Manhattan distance are as follows:
[0053]
[0054] Wherein, d1 is the first distance, d2 is the second distance, x i , x j is the predicted value, y i , y j is the actual value, and m and n are the sample quantities.
[0055] Then, the first distance and the second distance corresponding to each missing value are weighted and processed, and the K nearest neighbor algorithm is optimized to better adapt to the characteristics of the electric energy meter calibration line data.
[0056] In one possible implementation, the first distance and the second distance corresponding to each missing value are used to obtain the weighted distance corresponding to each missing value, which can include:
[0057] The sum of the first product and the second product of each missing value is taken as the weighted distance of the missing value. The first product is the product of the first distance and the first feature weight, and the second product is the product of the second distance and the second feature weight. The sum of the first feature weight and the second feature weight corresponding to each missing value is 1.
[0058] Optionally, the product of the first distance and the first feature weight is calculated to obtain a first product, and the product of the second distance and the second feature weight is calculated to obtain a second product. Then, the first product and the second product of each missing value are weighted and summed to obtain the weighted distance corresponding to each missing value. That is, the weighted Euclidean-Manhattan distance design of the first distance and the second distance is shown in formula (7):
[0059]
[0060] wherein d hybrid is the weighted distance, a is the first feature weight, i.e., the continuous feature weight; and β is the second feature weight, i.e., the discrete feature weight.
[0061] For example, according to the expert rule, the range of a can be set to 0.6-0.9, and β is fixed as 1-a. The constraint condition a+β=1 can ensure distance normalization. The embodiments of the application balance the fluctuation sensitivity of continuous quantity and the state representation ability of discrete quantity through the weight.
[0062] In step 103, the improved K nearest neighbor algorithm with improved distance is optimized by using the improved golden jackal optimization algorithm to obtain the optimal parameter combination of the improved K nearest neighbor algorithm with improved distance.
[0063] Since the data detected by the electric energy meter calibration line has multimodality and physical constraint, the data not only contains linear characteristics under steady state working condition, but also involves some nonlinear fluctuations. The traditional parameter selection method is difficult to balance the noise resistance and trend sensitivity in a complex high-dimensional space. The golden jackal optimization algorithm (GJO) converts parameter optimization into a global search problem by simulating the cooperative hunting mechanism of golden jackals. The dynamic population updating strategy and Cauchy reverse mutation mechanism can effectively avoid falling into local optimum, and are especially suitable for scenes with non-uniformly distributed abnormalities and multi-scale feature coupling in electric energy meter calibration line data.
[0064] In the embodiments of the application, the improved K nearest neighbor algorithm with improved distance is optimized by using the improved golden jackal optimization algorithm to obtain the optimal parameter combination of the improved K nearest neighbor algorithm with improved distance. The optimal parameter combination includes the number of nearest neighbors and the weight decay coefficient.
[0065] In one possible implementation, the improved K nearest neighbor algorithm with improved distance is optimized by using the improved golden jackal optimization algorithm to obtain the optimal parameter combination of the improved K nearest neighbor algorithm with improved distance, which can include:
[0066] The population in the optimization process, the maximum number of iterations, and the search space dimension are set, and population initialization is performed to randomly generate N individual positions. Each individual is a set of parameter combinations, and the total number of individuals is the same as the total number of populations.
[0067] calculate the fitness of each individual in the nearest neighbor repair scenario, arrange N fitness values in descending order, take the individual corresponding to the first fitness value as the male leader, and take the individual corresponding to the second fitness value as the female leader;
[0068] calculate the first position using the current position of the male leader and calculate the second position using the current position of the female leader;
[0069] determine whether the current iteration number is greater than the maximum iteration number;
[0070] If the current iteration number is not greater than the maximum iteration number, update the current position of the male leader using the first position, update the current position of the female leader using the second position, add 1 to the current iteration number, and return to the step of calculating the fitness of each individual in the nearest neighbor repair scenario for execution;
[0071] If the current iteration number is greater than the maximum iteration number, calculate the sum of the product of the first position and the first weight and the product of the second position and the second weight, and use the Cauchy inverse mutation as the disturbance strategy to optimize the sum of the product of the first position and the first weight and the product of the second position and the second weight, the first weight being greater than the second weight;
[0072] Calculate the optimal parameter combination using the optimized sum of the product of the first position and the first weight and the product of the second position and the second weight.
[0073] Optionally, the main principle of the improved golden cat optimization algorithm in the embodiment of the application includes that the golden cat usually hunts in pairs or small groups, and cooperates in tracking, surrounding and supplying prey. The golden cat optimization algorithm simulates the cooperation mechanism of the golden cat leader and the golden cat follower, and the process of dynamically adjusting strategy to capture prey. Referring to Figure 2 The main steps include initialization of parameters, population initialization, fitness evaluation, selection of male and female leaders, position updating, Cauchy inverse mutation, convergence judgment and output. The specific process is as follows:
[0074] First, initialize the parameters, set the population size, the maximum iteration number and the search space dimension. For example, the population N ∈ [50, 100], the maximum iteration number T max = 100, and since the main optimization parameters of the K nearest neighbor algorithm are the number of nearest neighbors k and the weight decay coefficient a, the search space dimension D can be set to 2. Then, initialize the population, randomly generate N individual positions, and each individual represents a parameter combination. That is, X i = [x i1 ,x i2 ,…,x iD ], i = 1, 2, …, N.
[0075] Fitness of each individual in the near neighbor repair scenario F(X i ) is calculated as follows:
[0076] The mean square error, time cost and relative entropy of N individuals are calculated using the positions of N individuals;
[0077] The fitness of each individual in the near neighbor repair scenario is calculated using the mean square error, time cost and relative entropy.
[0078] The specific calculation formula is as follows:
[0079]
[0080] Wherein, F(X i ) is the fitness, MSE repair is the mean square error, t repair is the time cost, KL(P repair ||P normal ) is the relative entropy, P repair is the distribution of repair data, and P normal is the distribution of normal data.
[0081] Wherein, the calculation formula of the mean square error MSE repair is as follows:
[0082]
[0083] Wherein, y i is the true value of the i-th data point, and N is the number of data points to be repaired.
[0084] The calculation formula of the relative entropy is as follows:
[0085]
[0086] After the fitness calculation is completed, the golden cat optimization algorithm will select the male and female leaders. The core advantage of setting the male and female dual leaders is to achieve an efficient balance between exploration and development through a division of labor and cooperation mechanism: the male leader dominates global exploration, quickly positioning the potential optimal region; the female leader focuses on local development, searching the current optimal neighborhood in detail, and both update the position of the solution, avoiding the population from falling into local optimum, and improving the convergence speed and accuracy through dynamic adjustment of search intensity, which can significantly enhance the robustness and solution quality of the algorithm. The selection criteria for the male and female leaders are the top two individuals in terms of fitness. The behavior of the leaders includes global search, local development, and disturbance strategy. Global search refers to the behavior of the leader guiding the population to search in a large range in the solution space to find potential prey (optimal region). The N fitness values are arranged in descending order, and the top two individuals corresponding to the fitness values are selected as the male and female leaders, respectively. Global search refers to the behavior of the leader guiding the population to search in a large range in the solution space to find potential prey (optimal region). The position update is as follows:
[0087] X new1 =X leader -E·|R·P prey -X current | (11)
[0088]
[0089] where X new1 is the first position, X leader is the current position of the female or male leader, P prey is the prey position, E is the energy factor, and t is the current iteration number.
[0090] After searching for the optimal region, the behavior of local development, i.e., the precise attack phase, is started. The trigger condition is whether the absolute value of the current energy factor is less than 1. That is, the current energy factor is calculated using the current iteration number and the maximum iteration number, and when the absolute value of the current energy factor is less than 1, the second position is calculated using the current position of the female leader.
[0091] The position update at this time is as follows:
[0092] X new2 =P prey -E·|R·P prey -X current | (13)
[0093] where X new2 is the second position.
[0094] Then it is judged whether the current iteration number is greater than the maximum iteration number, if not, the current position of the male leader is updated by the first position calculated above, the current position of the female leader is updated by the second position, the current iteration number is added by 1, and the step of fitness calculation is returned to continue execution; if greater, the product of the first position and the first weight is calculated, the product of the second position and the first weight is calculated, and the sum of the two products is optimized by taking the Cauchy inverse mutation as a disturbance strategy, and the optimal parameter combination is calculated by using the optimized sum of the products.
[0095] wherein the core of the golden jackal optimization algorithm is the position updating equation, which is calculated according to the mean value of the position vector of the prey updated by the male jackal and the female jackal. Since the male jackal has a greater impact on the update of the prey position in the iteration process, giving it a greater weight helps to improve the convergence speed of the algorithm. Therefore, by introducing an adaptive weight to replace the mean weight in the original algorithm, the size of the function value corresponding to the male jackal and the female jackal is used to update the weight, which can make the weight of the position updated according to the male jackal higher than the weight of the position updated according to the female jackal. The weight calculation and the sum of the products are as follows:
[0096]
[0097] wherein w male is the first weight, w female is the second weight, f female is the mean weight of the female leader, f male is the mean weight of the male leader, X i (t) is the position at the t time, and X i (t+1) is the position at the t+1 time.
[0098] In order to avoid falling into local optimum, a disturbance strategy is usually used to escape from local optimum. In the later stage of iteration, such as t>0.75T max , random noise is added to the leader position, that is:
[0099] X i (t+1)=X i (t+1)+N(0,σ 2 ) (16)
[0100]
[0101] wherein N(0,σ 2) is random noise, and σ is the noise variance. However, the effect of this random noise on preventing falling into local optimum is not ideal. Embodiments of the present application improve the position update formula of the golden jackal optimization algorithm by introducing a Cauchy variation term to replace the random noise. Due to the characteristics of high-dimensional time series, strong correlation of multiple parameters, and dynamic drift of the electric energy meter calibration line data, the parameter optimization space often presents the characteristics of multimodality and ruggedness. The heavy-tailed characteristics of Cauchy distribution give the variation operation stronger long-range disturbance ability. The probability density characteristics of low peak and long tail can generate a larger range of mutant solutions in iteration, thereby penetrating the local optimal barrier in the complex solution space, especially suitable for the "pseudo-optimal trap" formed by the aggregation of abnormal points in the electric energy meter data. Compared with other methods, Cauchy reverse mutation can more efficiently match the dual requirements of global sensitivity and dynamic adaptability for electric energy meter data repair while maintaining population diversity. The expression is as follows:
[0102] X i (t+1)=X i (t+1)+η(t)·Cauchy(0,1) (18)
[0103] η(t)=η0·e -βt (19)
[0104] wherein η0=0.1, β=0.02, Cauchy(0,1) is a standard Cauchy distribution, that is, a randomly generated value in the standard Cauchy distribution (the position parameter is 0, and the scale parameter is 1). The variation strength gradually decreases with the iteration number t. Cauchy variation significantly improves the robustness of the golden jackal optimization algorithm in complex optimization scenarios through its long-tail disturbance characteristics and dynamic decay mechanism. The essence of its improvement is to balance the exploration and development capabilities of the algorithm through controllable randomness injection. Therefore, the position update expression of the number of nearest neighbors k and the weight decay coefficient a in the K-nearest neighbor algorithm is as follows:
[0105]
[0106] Embodiments of the present application use the golden jackal optimization algorithm to optimize the number of nearest neighbors and the weight decay coefficient in the K-nearest neighbor algorithm, which can effectively overcome the limitations of traditional parameter selection relying on experience or exhaustive search. By simulating the group intelligence mechanism of golden jackal cooperative hunting, the golden jackal optimization algorithm dynamically searches for the optimal parameter combination in the global range, avoiding local optimal traps; after introducing adaptive weights and Cauchy reverse mutation strategies, the algorithm further balances the global exploration and local development capabilities, and adaptively adjusts the sensitivity of the parameters to the data distribution. This optimization method enables the number of nearest neighbors to accurately match the local density characteristics of the data, and the weight decay coefficient can adapt to noise and dimension changes, thereby significantly improving the generalization ability for dynamic abnormal data of the electric energy meter calibration line, and realizing high-precision and adaptive parameter optimization without human intervention.
[0107] In step 104, the repair value corresponding to each missing value is calculated according to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination.
[0108] In the embodiment of the present application, the core of the dynamic repair process is to realize real-time correction of the data stream through periodic parameter optimization and local weighted repair, that is, the weighted distance calculated in step 102 is corrected by using the optimal parameter combination calculated in step 103, and the repair value corresponding to each missing value is calculated.
[0109] In a possible implementation, the optimal parameter combination includes the number of nearest neighbors and the weight decay coefficient, and the repair value corresponding to each missing value is calculated according to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination, which can include:
[0110] For the missing value, the corresponding weight decay coefficient is substituted into the weighted distance of the missing value, and the weight coefficient of each nearest neighbor corresponding to the missing value is calculated by using the weighted distance after substituting the weight decay coefficient and the number of nearest neighbors corresponding to the missing value.
[0111] The repair value corresponding to the missing value is calculated by using the weight coefficient of each nearest neighbor corresponding to the missing value and the corresponding number of nearest neighbors.
[0112] Optionally, the optimal parameter combination of the K-nearest neighbor algorithm, that is, the number of nearest neighbors and the weight decay coefficient, is searched periodically by using the golden cat optimization algorithm, the weight decay coefficient is brought into formula (7) to calculate the weighted weighted distance, and the data is repaired by using the local weighting method. Its expression is as follows:
[0113]
[0114] Wherein, k is the number of nearest neighbors, w i is the weight coefficient of each nearest neighbor.
[0115] The generated repair value is replaced with the original abnormal value in the embodiment of the present application, that is, the repaired calibration data set is obtained. Finally, this process continues to circulate, and the parameter optimization is triggered again whenever new data accumulates to a certain scale, so that the K-nearest neighbor algorithm of improved distance can always adapt to the change of data distribution.
[0116] The application provides an abnormal calibration data repairing method applied to an electric energy meter calibration line, obtains electric energy meter calibration line data, and divides missing values in the electric energy meter calibration line data according to a missing value classification standard; a K nearest neighbor algorithm is improved by using a combination of Euclidean distance and Manhattan distance measurement methods to obtain an improved distance K nearest neighbor algorithm, each missing value is calculated by using the improved distance K nearest neighbor algorithm, and a weighted distance corresponding to each missing value is obtained; the improved distance K nearest neighbor algorithm is optimized by using an improved golden jackal optimization algorithm to obtain an optimal parameter combination of the improved distance K nearest neighbor algorithm; and a repairing value corresponding to each missing value is calculated according to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination. The application can significantly improve the robustness and adaptability of data repairing by using the combination of Euclidean distance and Manhattan distance measurement methods to improve the K nearest neighbor algorithm to calculate the weighted distance corresponding to each missing value, thereby improving the accuracy and reliability of the electric energy meter calibration line calibration process; and the improved distance K nearest neighbor algorithm can effectively overcome the limitations of traditional parameter selection depending on experience or exhaustive search, thereby significantly improving the generalization ability of dynamic abnormal data of the electric energy meter calibration line and realizing high-precision and self-adaptive parameter optimization without manual intervention.
[0117] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0118] The following is a device embodiment of the application. For details not described in detail, reference can be made to the corresponding method embodiments described above.
[0119] Figure 3 A structure diagram of the device for repairing abnormal calibration data of an electric energy meter calibration line provided by the embodiments of the application is shown. For ease of illustration, only the parts related to the embodiments of the application are shown, and the details are as follows:
[0120] As shown in Figure 3 The device for repairing abnormal calibration data of an electric energy meter calibration line 3 comprises:
[0121] A data acquisition module 31 is configured to obtain electric energy meter calibration line data, and divide missing values in the electric energy meter calibration line data according to a missing value classification standard;
[0122] A distance calculation module 32 is configured to improve a K nearest neighbor algorithm by using a combination of Euclidean distance and Manhattan distance measurement methods to obtain an improved distance K nearest neighbor algorithm, and calculate each missing value after division by using the improved distance K nearest neighbor algorithm to obtain a weighted distance corresponding to each missing value;
[0123] A parameter calculation module 33 is used to optimize the improved distance K-nearest neighbor algorithm using an improved golden jackal optimization algorithm to obtain an optimal parameter combination of the improved distance K-nearest neighbor algorithm;
[0124] The repair value calculation module 34 is used to calculate the repair value corresponding to each missing value according to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination.
[0125] The present application provides a device for repairing abnormal calibration data of an electric energy meter calibration line. The device obtains electric energy meter calibration line data and divides missing values in the electric energy meter calibration line data according to a missing value classification standard; improves a K-nearest neighbor algorithm by using a metric method that combines Euclidean distance and Manhattan distance to obtain an improved distance K-nearest neighbor algorithm, and calculates each missing value after division by using the improved distance K-nearest neighbor algorithm to obtain a weighted distance corresponding to each missing value; optimizes the improved distance K-nearest neighbor algorithm by using an improved golden jackal optimization algorithm to obtain an optimal parameter combination of the improved distance K-nearest neighbor algorithm; and calculates a repair value corresponding to each missing value based on the weighted distance corresponding to each missing value and the corresponding optimal parameter combination. This application improves the K-nearest neighbor algorithm by combining the Euclidean distance and the Manhattan distance to calculate the weighted distance corresponding to each missing value, which can significantly improve the robustness and adaptability of data repair, thereby improving the accuracy and reliability of the calibration process of the electricity meter calibration line; and uses the improved golden jackal optimization algorithm to optimize the K-nearest neighbor algorithm with improved distance, which can effectively overcome the limitations of traditional parameter selection relying on experience or traversal search, thereby significantly improving the generalization ability of dynamic abnormal data of the electricity meter calibration line, and realizing high-precision and adaptive parameter tuning without human intervention.
[0126] In a possible implementation, the device may further include an outlier elimination module, which may be used to:
[0127] Eliminate extreme outliers in the meter calibration line data based on the preset anomaly detection threshold;
[0128] After removing extreme outliers, the missing values are classified according to the missing value classification criteria.
[0129] In a possible implementation, the device may further include a standardization module, and the standardization module may be configured to:
[0130] The divided electric energy meter calibration line data is subjected to data standardization processing. After data standardization processing, irrelevant features are eliminated and relevant features are retained for missing values of different categories in turn.
[0131] Accordingly, the distance calculation module can be used to:
[0132] The K-Nearest Neighbor algorithm is improved by combining the Euclidean distance and Manhattan distance to obtain the improved distance K-Nearest Neighbor algorithm, and each missing value after irrelevant feature elimination and relevant feature retention is calculated by using the improved distance K-Nearest Neighbor algorithm to obtain the corresponding weighted distance of each missing value.
[0133] In a possible implementation, the distance calculation module can be configured to:
[0134] The Euclidean distance in the improved distance K-Nearest Neighbor algorithm is used to calculate each missing value after division to obtain the corresponding first distance of each missing value.
[0135] The Manhattan distance in the improved distance K-Nearest Neighbor algorithm is used to calculate each missing value after division to obtain the corresponding second distance of each missing value.
[0136] The corresponding first distance and second distance of each missing value are used to obtain the corresponding weighted distance of each missing value.
[0137] In a possible implementation, the distance calculation module can be configured to:
[0138] The sum of the first product and the second product of each missing value is taken as the weighted distance of the missing value, the first product is the product of the first distance and the first feature weight, and the second product is the product of the second distance and the second feature weight, and the sum of the first feature weight and the second feature weight corresponding to each missing value is 1.
[0139] In a possible implementation, the parameter calculation module can be configured to:
[0140] The population in the optimization process, the maximum number of iterations, and the search space dimension are set, and the population initialization is performed to randomly generate N individual positions, each individual is a combination of parameters, and the total number of individuals is the same as the total number of populations;
[0141] The fitness of each individual in the nearest neighbor repair scenario is calculated, and the N fitness values are arranged in descending order according to the numerical value, the individual corresponding to the first fitness value in the arranged N fitness values is taken as the male leader, and the individual corresponding to the second fitness value in the arranged N fitness values is taken as the female leader.
[0142] The first position is calculated by using the current position of the male leader, and the second position is calculated by using the current position of the female leader.
[0143] It is judged whether the current number of iterations is greater than the maximum number of iterations.
[0144] If the current iteration number is not greater than the maximum iteration number, the first position is used to update the current position of the male leader, the second position is used to update the current position of the female leader, the current iteration number is added by 1, and the step of calculating the fitness of each individual in the nearest neighbor repair scenario is returned to continue execution;
[0145] If the current iteration number is greater than the maximum iteration number, the sum of the product of the first position and the first weight and the product of the second position and the second weight is calculated, the sum of the product of the first position and the first weight and the product of the second position and the second weight is optimized by using the Cauchy inverse mutation as a disturbance strategy, the first weight is greater than the second weight;
[0146] The optimal parameter set is calculated by using the optimized sum of the product of the first position and the first weight and the product of the second position and the second weight.
[0147] In a possible implementation, the parameter calculation module can be configured to:
[0148] The current energy factor is calculated by using the current iteration number and the maximum iteration number, and the second position is calculated by using the current position of the female leader when the absolute value of the current energy factor is less than 1.
[0149] In a possible implementation, the parameter calculation module can be configured to:
[0150] The mean square error, the time cost and the relative entropy of the N individuals are calculated by using the positions of the N individuals.
[0151] The fitness of each individual in the nearest neighbor repair scenario is calculated by using the mean square error, the time cost and the relative entropy.
[0152] In a possible implementation, the optimal parameter combination includes the number of nearest neighbors and the weight decay coefficient, and the repair value calculation module can be configured to:
[0153] For the missing value, the corresponding weight decay coefficient is substituted into the weighted distance of the missing value, and the weight coefficient of each nearest neighbor number corresponding to the missing value is calculated by using the weighted distance substituted with the weight decay coefficient and the number of nearest neighbors corresponding to the missing value.
[0154] The repair value corresponding to the missing value is calculated by using the weight coefficient of each nearest neighbor number corresponding to the missing value and the number of nearest neighbors corresponding to the missing value.
[0155] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0156] Those skilled in the art can appreciate that the templates, units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0157] The modules / units, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the above-mentioned steps applied to the electric energy meter calibration line abnormal calibration data repair method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0158] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for repairing abnormal calibration data of an electric energy meter calibration line, characterized in that: include: Acquire electric energy meter calibration line data, and classify missing values in the electric energy meter calibration line data according to a missing value classification standard; The K-nearest neighbor algorithm is improved by using a metric method combining Euclidean distance and Manhattan distance to obtain an improved distance K-nearest neighbor algorithm, and the improved distance K-nearest neighbor algorithm is used to calculate each missing value after the division to obtain a weighted distance corresponding to each missing value; The improved golden jackal optimization algorithm is used to optimize the K-nearest neighbor algorithm with improved distance to obtain the optimal parameter combination of the K-nearest neighbor algorithm with improved distance; According to the weighted distance corresponding to each missing value and the corresponding optimal parameter combination, the repair value corresponding to each missing value is calculated.
2. The method for repairing abnormal calibration data of an electric energy meter calibration line according to claim 1 is characterized in that: After classifying the missing values in the electric energy meter calibration line data according to the missing value classification standard, the method further includes: Eliminate extreme abnormal values in the electric energy meter calibration line data according to a preset abnormality detection threshold; The missing values after the extreme outliers are eliminated are divided according to the missing value classification standard.
3. The method for repairing abnormal calibration data applied to an electric energy meter calibration line according to claim 1, characterized in that: After classifying the missing values in the electric energy meter calibration line data according to the missing value classification standard, the method further includes: The divided electric energy meter calibration line data is subjected to data standardization processing. After data standardization processing, irrelevant features are eliminated and relevant features are retained for missing values of different categories in turn. Accordingly, the K-nearest neighbor algorithm is improved by combining the Euclidean distance and the Manhattan distance to obtain the improved distance K-nearest neighbor algorithm, and the improved distance K-nearest neighbor algorithm is used to calculate each missing value after the division to obtain the weighted distance corresponding to each missing value, including: The K-nearest neighbor algorithm is improved by using a metric method that combines Euclidean distance and Manhattan distance to obtain an improved distance K-nearest neighbor algorithm. The improved distance K-nearest neighbor algorithm is used to calculate each missing value after irrelevant features are eliminated and relevant features are retained to obtain a weighted distance corresponding to each missing value.
4. The method for repairing abnormal calibration data applied to an electric energy meter calibration line according to claim 1, characterized in that: The improved distance K-nearest neighbor algorithm is used to calculate each missing value after the division to obtain the weighted distance corresponding to each missing value, including: Calculate each missing value after division using the Euclidean distance in the improved K-nearest neighbor algorithm to obtain a first distance corresponding to each missing value; Calculate each missing value after the division using the Manhattan distance in the improved K-nearest neighbor algorithm to obtain a second distance corresponding to each missing value; The first distance and the second distance corresponding to each missing value are used to obtain the weighted distance corresponding to each missing value.
5. The method for repairing abnormal calibration data of an electric energy meter calibration line according to claim 4 is characterized in that: The method of obtaining a weighted distance corresponding to each missing value by using the first distance and the second distance corresponding to each missing value includes: The sum of the first product and the second product of each missing value is used as the weighted distance of the missing value, where the first product is the product of the first distance and the first feature weight, and the second product is the product of the second distance and the second feature weight. The sum of the first feature weight and the second feature weight corresponding to each missing value is unit 1.
6. The method for repairing abnormal calibration data applied to an electric energy meter calibration line according to claim 1, characterized in that: The improved golden jackal optimization algorithm is used to optimize the improved distance K-nearest neighbor algorithm to obtain the optimal parameter combination of the improved distance K-nearest neighbor algorithm, including: The population, maximum number of iterations, and search space dimension in the optimization process are set, and the population is initialized to randomly generate N individual positions. Each individual is a set of parameter combinations, and the total number of individuals is the same as the total number of the population. Calculate the fitness of each individual in the neighbor repair scenario, and sort the N fitness values from large to small. The individual corresponding to the first fitness value among the sorted N fitness values is selected as the male leader, and the individual corresponding to the second fitness value among the sorted N fitness values is selected as the female leader. Calculating a first position using the current position of the male leader, and calculating a second position using the current position of the female leader; Determine whether the current number of iterations is greater than the maximum number of iterations; If the current number of iterations is not greater than the maximum number of iterations, the current position of the male leader is updated using the first position, the current position of the female leader is updated using the second position, the current number of iterations is increased by 1, and the process returns to the step of calculating the fitness of each individual in the neighbor repair scenario to continue. If the current number of iterations is greater than the maximum number of iterations, calculating the sum of the product of the first position and the first weight and the product of the second position and the second weight, and using Cauchy reverse mutation as the perturbation strategy to optimize the sum of the product of the first position and the first weight and the product of the second position and the second weight, where the first weight is greater than the second weight; An optimal parameter combination is calculated using the sum of the optimized product of the first position and the first weight and the optimized product of the second position and the second weight.
7. The method for repairing abnormal calibration data of an electric energy meter calibration line according to claim 6, characterized in that: The step of calculating the second position using the current position of the female leader comprises: The current energy factor is calculated using the current number of iterations and the maximum number of iterations, and when the absolute value of the current energy factor is less than 1, the second position is calculated using the current position of the female leader.
8. The method for repairing abnormal calibration data applied to an electric energy meter calibration line according to claim 6, characterized in that: The calculation of the fitness of each individual in the neighbor repair scenario includes: Using N individual positions, calculate the mean square error, time cost and relative entropy of N individuals; The fitness of each individual in the neighbor repair scenario is calculated using the mean square error, the time cost, and the relative entropy.
9. The method for repairing abnormal calibration data applied to an electric energy meter calibration line according to claim 1, characterized in that: The optimal parameter combination includes the number of nearest neighbors and the weight attenuation coefficient. The repair value corresponding to each missing value is calculated based on the weighted distance corresponding to each missing value and the corresponding optimal parameter combination, including: For the missing value, the corresponding weight decay coefficient is substituted into the weighted distance of the missing value, and the weight coefficient of each nearest neighbor number corresponding to the missing value is calculated using the weighted distance after substituting the weight decay coefficient and the number of nearest neighbors corresponding to the missing value; The repair value corresponding to the missing value is calculated using the weight coefficients of the nearest neighbors corresponding to the missing value and the corresponding number of nearest neighbors.
10. A device for repairing abnormal calibration data of an electric energy meter calibration line, characterized in that: include: A data acquisition module is used to acquire the electric energy meter calibration line data and classify the missing values in the electric energy meter calibration line data according to the missing value classification standard; A distance calculation module is used to improve the K-nearest neighbor algorithm by using a metric method that combines Euclidean distance and Manhattan distance to obtain an improved distance K-nearest neighbor algorithm, and use the improved distance K-nearest neighbor algorithm to calculate each missing value after division to obtain a weighted distance corresponding to each missing value; A parameter calculation module is used to optimize the improved distance K-nearest neighbor algorithm using an improved golden jackal optimization algorithm to obtain an optimal parameter combination of the improved distance K-nearest neighbor algorithm; The repair value calculation module is used to calculate the repair value corresponding to each missing value based on the weighted distance corresponding to each missing value and the corresponding optimal parameter combination.