Concentrator error calibration method based on machine learning
By dynamically allocating the influence weight of the decision tree in the random forest regression model, combined with the correlation analysis within the fluctuation period, the problem of the inability to fully consider the impact of different characteristics on electrical energy fluctuations in the existing technology is solved, and high-precision electrical energy error calibration and improved stability of the power system are achieved.
Patent Information
- Application Number
- CN202510205788.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The prior art cannot fully consider the different effects of different characteristics on electrical energy fluctuations, resulting in low calibration accuracy of electrical energy measurement errors.
The machine learning-based concentrator error calibration method is used to impart reasonable weights to different characteristics through the random forest regression model and dynamic allocation of each decision tree, combined with the correlation analysis during the fluctuation period.
Accurate calibration of electrical energy errors is achieved, the accuracy of power metering and system stability are improved, and efficient operation is maintained in a changing power system environment.
Smart Images

Figure CN119691703B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a concentrator error calibration method based on machine learning. Background Art
[0002] Smart concentrators are key devices for real-time data collection and remote monitoring in power systems. Their main functions include monitoring the power output of power equipment and other power system related parameters (such as voltage, current, etc.). These concentrators transmit the collected data to the central control system by supporting multiple communication protocols (such as PLC, GPRS / 3G, RF, etc.) for decision analysis and power management. However, the measurement process of the concentrator is affected by many factors, such as sensor drift, changes in environmental conditions, and equipment aging, which lead to errors in the measurement data, thus affecting the stability of the power system and the accuracy of billing.
[0003] Traditional error correction methods rely on manual intervention or fixed rules and cannot dynamically adapt to changing measurement environments and conditions. Existing error calibration techniques mainly rely on traditional mathematical modeling methods, such as linear regression and Kalman filtering. Although these methods can solve the error problem to a certain extent, they often perform poorly when faced with complex power systems and changing environmental conditions. With the development of machine learning technology, error calibration methods based on machine learning have become a new research direction.
[0004] Among them, the random forest algorithm, as an integrated machine learning method, has shown strong advantages in processing high-dimensional data and nonlinear problems. By constructing multiple decision trees and taking weighted average of the prediction results of each tree, the random forest can better fit the complex input-output relationship and has strong robustness. However, all decision trees of the random forest algorithm have the same weight in the model, which fails to fully consider the differences in the impact of different features on power fluctuations. Power measurement is affected by multiple factors (such as current, voltage, and load power), but the degree of influence and correlation of these factors are not exactly the same, which may ultimately lead to low accuracy of the random forest regression model. Summary of the invention
[0005] In view of the problem that all the above decision trees have the same weight, the present invention proposes a concentrator error calibration method based on machine learning, including: obtaining multiple parameter time series of the concentrator, the multiple parameter time series including the electric energy time series and the load power time series; multiple parameters at the same acquisition time constitute a data set, and the multiple data sets are input into the random forest regression model for training to obtain a trained regression model; obtaining the real-time data set of the concentrator, and inputting the real-time data set into the trained regression model to obtain the electric energy prediction value; the difference between the electric energy prediction value and the real-time electric energy constitutes the electric energy error, and the product of the electric energy error and the set correction coefficient and the sum of the real-time electric energy are used as the electric energy correction value; the random forest regression model also includes for each decision tree The decision tree allocates influence weights, specifically as follows: the electric energy fluctuation interval of the electric energy time series is obtained to obtain multiple fluctuation periods, and the similarity between the load power time series and the electric energy time series in each fluctuation period is calculated; the fluctuation period corresponding to the similarity less than the set threshold is recorded as a weakly correlated fluctuation period, and the remaining fluctuation periods are recorded as strongly correlated fluctuation periods; the average of the Pearson correlation coefficients between the parameter time series except the load power time series and the electric energy time series in all weakly correlated fluctuation periods is taken as the trend correlation of the corresponding parameter; the average of the Pearson correlation coefficients between the load power time series and the electric energy time series in all strongly correlated fluctuation periods is taken as the trend correlation of the load power; the sum of the trend correlations corresponding to all parameters contained in each decision tree constitutes the influence weight of the corresponding decision tree.
[0006] The machine learning-based concentrator error calibration method of the present invention solves the problem that the existing technology cannot consider the different effects of different features on power fluctuations by introducing a random forest regression model and dynamically assigning the influence weight of each decision tree. Compared with traditional machine learning methods such as linear regression and Kalman filtering, random forests can capture the complex nonlinear relationship of power fluctuations and assign reasonable weights to different features through correlation analysis within the fluctuation period. This method can not only calibrate measurement errors, but also dynamically adjust according to real-time data to ensure that high-precision power prediction results are provided in various power system environments, thereby improving the accuracy of power metering and the stability of the system.
[0007] Furthermore, the calculation method of the influence weight is specifically as follows:
[0008] ;
[0009] in Indicates The influence weight of each tree; Indicates The tree nodes; Represents parameter variables; Indicates The total number of nodes in a tree; Represents parameter variables The corresponding trend correlation.
[0010] The present invention further improves the accuracy of the model by calculating the influence weight of each decision tree and comprehensively evaluating the contribution of feature variables to the power prediction results. Unlike the traditional random forest algorithm, the present invention assigns weights to each tree, which can better handle the role of multiple features in power error calibration and ensure more accurate prediction results.
[0011] Furthermore, the calculation method of the trend correlation is specifically as follows:
[0012] ;
[0013] in Representation parameters Trend relevance of Indicates Weakly correlated volatility periods Corresponding parameter timing Energy Sequence Pearson correlation coefficient between ; Represents the total number of weakly correlated volatility periods.
[0014] By calculating the trend correlation, the present invention can effectively evaluate the long-term impact of each feature on power fluctuations, so that the model can adapt to different working environments. This method avoids the decrease in prediction accuracy of traditional algorithms under different load and environmental conditions, thereby improving the stability and reliability of the model.
[0015] Furthermore, the multiple parameter timings include: input current timing, input voltage timing, output current timing, output voltage timing, load power timing, ambient temperature timing, ambient humidity timing, and active power timing and reactive power timing.
[0016] Furthermore, the random forest regression model also includes taking input current, input voltage, output current, output voltage, load power, ambient temperature, and ambient humidity as input features; and taking active power and reactive power as feature labels.
[0017] Furthermore, obtaining the electric energy fluctuation interval of the electric energy time series to obtain multiple fluctuation time periods also includes: using a moving standard deviation algorithm to obtain the standard deviation of each sliding window in the electric energy time series; using the 3σ principle to screen the standard deviation greater than μ±3σ to obtain the corresponding multiple electric energy fluctuation intervals, and all sampling moments of each electric energy fluctuation interval constitute the corresponding fluctuation period.
[0018] The present invention uses the moving standard deviation and the 3σ principle to obtain the power fluctuation range, which enhances the model's sensitivity to power fluctuations, thereby more accurately identifying and calibrating measurement errors. Compared with existing methods, this technology can more accurately capture power fluctuations and avoid minor errors that may be ignored by traditional methods.
[0019] Furthermore, the setting correction coefficient is determined using an incremental learning method.
[0020] Furthermore, the set threshold is obtained using a maximum inter-class variance method.
[0021] The technical effects of the present invention are:
[0022] The present invention realizes accurate calibration of concentrator errors by combining the random forest regression model with feature correlation analysis. In particular, the present invention solves the problem that existing methods cannot fully consider the different impacts of different features on power fluctuations by dynamically allocating the influence weight of each decision tree. The accuracy and adaptability of power error correction are guaranteed by dividing the fluctuation period, calculating the trend correlation and applying the dynamic correction coefficient. In addition, the present invention also introduces the moving standard deviation method and the maximum inter-class variance method, making the error calibration process more accurate and robust, and able to maintain efficient operation in a changeable power system environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0024] Figure 1 It is a flowchart schematically showing a concentrator error calibration method based on machine learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0026] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0027] Concentrator error calibration method embodiment based on machine learning:
[0028] like Figure 1As shown, the concentrator error calibration method based on machine learning of the present invention includes:
[0029] S1. Obtain relevant data of the concentrator and perform preprocessing.
[0030] The intelligent concentrator is an intelligent device used for real-time data collection and remote monitoring in the power system. The concentrators mentioned later in this embodiment refer to this intelligent concentrator. Its core functions include real-time monitoring of the active power and reactive power output of the transformer, as well as other power system-related parameters such as voltage and current. The concentrator transmits the collected data to the central control system through various communication protocols (such as PLC, GPRS / 3G, RF, etc.) for data analysis and decision support. However, due to environmental factors, equipment aging or external interference, the concentrator may have measurement errors, affecting the accuracy of power data. In order to ensure the stability of the power system and the accuracy of billing, it is necessary to automatically calibrate the measurement errors of the concentrator.
[0031] Multiple relevant data collected by the concentrator can be obtained, including the input current of the power equipment , Input voltage , output current , output voltage , Load power , Ambient temperature , Ambient humidity And active energy Reactive Energy . Since the concentrator supports the IEC62056 protocol, in this embodiment, the current and voltage data of other power monitoring devices can be received through optical fiber, and the units are volts and amperes respectively; then the temperature data of the operating environment can be obtained through the built-in NTC thermistor of the concentrator, and the environmental humidity data can be obtained through the capacitive humidity sensor, and the unit is percentage; further, the load power of the power equipment can be obtained through the power sensor integrated in the concentrator, and the unit is kilowatt; finally, the power measurement module of the concentrator can calculate the active power and reactive power based on the phase difference and amplitude of the current and voltage of the sensor, and the units are kilowatt-hour and kilovar-hour, respectively. The specific calculation method of active power and reactive power belongs to the known technology and will not be repeated here.
[0032] After obtaining the relevant data from the above concentrators, preprocessing operations are required to ensure that the data meets the standards of the subsequent machine learning model, including:
[0033] The acquired initial data is subjected to outlier removal based on the 3σ principle, and missing values are filled by linear interpolation; each data is then standardized using mean-variance standardization and scaled to the range of [0,1] by minimum-maximum standardization to solve the problem of inconsistent dimensions of each parameter. The above-mentioned preprocessing algorithms belong to the known technology, and the specific implementation methods are not described here in detail.
[0034] At this point, the The standardized relevant data of the concentrator at the time of collection is expressed as: ;in Indicates Collect multi-dimensional data of the concentrator at the time of collection; and Respectively represent The input current and input voltage of the power equipment at the time of collection; and Respectively represent The output current and output voltage of the power equipment at the time of collection; Indicates Load power at the time of collection; Indicates The temperature of the concentrator operating environment at the time of collection; Indicates The humidity of the concentrator operating environment at the time of collection; and Respectively represent The active and reactive energy measured at the time of collection. It should be noted that the data collected in this step are all original data without errors after being screened by technicians and are used for subsequent model training.
[0035] S2. Split the concentrator multidimensional data obtained in step S1 into input features and feature labels; obtain training data and test data of the random forest regression model based on the model data set to complete the training of the random forest regression model.
[0036] In step S1, the Multi-dimensional data of the concentrator at the time of collection Active power and reactive power are two key power indicators collected by the concentrator. Active power is the main energy transmitted and consumed in the power system, and its accuracy directly affects the metering of power and the settlement of electricity charges; while reactive power is a key parameter for maintaining the voltage stability of the power system and the normal operation of power equipment. If errors occur, the operating efficiency of power equipment will be reduced, and it may even affect the safe operation of the power grid.
[0037] Although the concentrator collects data through sensors and performs the final calculation of active power and reactive power, the actual measured value may produce certain errors due to sensor drift, environmental factors (such as temperature and humidity changes) or load fluctuations of power equipment. Therefore, in this embodiment, the error can be predicted and corrected by a machine learning model, and the measurement error is often not linear, but is affected by multiple factors. The relationship between these factors may be very complex, so in this embodiment, a random forest algorithm can be used to perform nonlinear data prediction to calibrate the error of active power and reactive power.
[0038] Random forest is an integrated machine learning method that constructs multiple decision trees and performs weighted average on the prediction results of each decision tree to give the final prediction result. First, the model data set of random forest is obtained. In step S1, the first Multi-dimensional data of the concentrator at the time of collection Since in this embodiment only the active power and reactive power of the concentrator are predicted, the multi-dimensional data of the concentrator can be Split into input features and feature labels, that is, active energy Reactive Energy As feature labels, and the remaining parameters as input features, an exemplary explanation, for the Multi-dimensional data of the concentrator at the time of collection , and its corresponding input feature is recorded as , and there are:
[0039] ;
[0040] And the corresponding feature label is recorded as , and there are:
[0041] ;
[0042] That is to say, any concentrator multidimensional data Can be split into input features With feature tags In this embodiment, the scale of the model data set can be set to an empirical value of 5000, and the final model data set is:
[0043] ;
[0044] in Represents the model dataset; Indicates total The concentrator multidimensional data, in this embodiment, has , and subsequently in the model dataset 80% of the data is randomly selected as the training set And randomly select 20% of the data as a test set , the training set Used for model training and test set after training is complete Evaluate the predictive performance of the model. In this embodiment, during the model training process, the number of decision trees n_estimators is set to an empirical value of 100, the maximum depth of the decision tree max_depth is set to an empirical value of None, the minimum number of samples required to split internal nodes min_samples_split is set to an empirical value of 2, the minimum number of samples required for leaf nodes min_samples_leaf is set to an empirical value of 5, and the number of features considered when finding the best splitting point max_features is set to an empirical value of 2, that is, the total number of features is squared and rounded; the predictive performance of the model can also be evaluated based on the test set after the model training is completed using the root mean square error RMSE or the mean absolute error MAE. The specific training process of the above-mentioned random forest regression model belongs to the well-known technology and will not be repeated here.
[0045] S3. Obtain the power fluctuation period based on the moving standard deviation; calculate the Pearson correlation coefficient between the load power time series and the power time series corresponding to each power fluctuation period to obtain the load fluctuation correlation, and record the power fluctuation period corresponding to the load fluctuation correlation less than the set correlation threshold as a weakly correlated fluctuation period; obtain the trend correlation of each feature based on the Pearson correlation coefficient of the remaining feature time series and the power time series in each weakly correlated fluctuation period, and calculate the influence weight of the decision tree based on the trend correlation of the features corresponding to the nodes in each decision tree to obtain the weighted prediction value.
[0046] In step S2, the model data set is constructed and the random forest regression model is trained, and the final prediction value is obtained by averaging the prediction results of all decision trees. However, this equal weight calculation method based on all trees has a defect: all decision trees have the same impact on the prediction results, and the difference in the degree of correlation between different features in the decision tree and the power measurement (active power and reactive power) is not taken into account.
[0047] S3.1. Obtain the power fluctuation period based on the moving standard deviation.
[0048] In actual scenarios, the power data collected by the concentrator is affected by different features, such as voltage, current, ambient temperature, humidity, and load power. The correlation between these features and power varies greatly. Some features (such as voltage and current) are more directly and closely related to power, while other features (such as temperature and humidity) have a certain impact on power, but this impact is more indirect. Therefore, simply treating the prediction results of each tree as having equal weight may lead to some key features being underestimated, thereby affecting the accuracy of the prediction.
[0049] Based on the above analysis, the influence of each feature on electric energy can be evaluated in the follow-up, so that when each tree makes a prediction, the weight of the corresponding decision tree is determined according to the number of features selected by all leaf nodes on the tree. An exemplary explanation is as follows: if there are two decision trees, and each tree has 10 leaf nodes, 6 leaf nodes of the first tree are output current thresholds, and the remaining 4 leaf nodes are thresholds of other temperature or humidity features; the second tree has only two leaf nodes for output current thresholds, and the remaining 8 leaf nodes are thresholds of temperature or humidity features. Under preliminary evaluation, since the impact of the output current feature on electric energy is significantly greater than the impact of temperature or humidity on electric energy, it can be determined that the prediction result of the first tree with more leaf nodes containing the output current threshold is more accurate than that of the second tree, that is, the influence weight of the first tree on the final prediction result should be greater.
[0050] In step S1, the Multi-dimensional data of the concentrator at the time of collection , and in step S2, the size of the data set is set to 5000, that is, a total of 5000 multi-dimensional data of acquisition moments are obtained, then each characteristic parameter has a total of 5000 continuous data constituting a time series, which are: input current time series , Input voltage timing , Output current timing , Output voltage timing , Load power timing , Temperature timing , humidity time series , Active energy timing And reactive energy timing Since the calculation methods of active electric energy and reactive electric energy are similar, in this embodiment, only active electric energy is taken as an example to calculate the correlation between each characteristic parameter and active electric energy, specifically:
[0051] First, use the moving standard deviation algorithm to obtain the active energy time series The fluctuation range of the algorithm is the active energy time series. The step size of the sliding window can be set to an empirical value of 20 in this embodiment. The output of the algorithm is the standard deviation of each sliding window. In this embodiment, the sliding window with a larger standard deviation can be selected as the power fluctuation period based on the 3σ principle. The power fluctuation period is recorded as , will The active energy time sequence corresponding to each power fluctuation period is recorded as In addition to the empirical value of 20, the step size of the sliding window can also be selected as 10, 15 or 20 to balance the response speed of load fluctuations and the timeliness of data errors caused by environmental factors.
[0052] After obtaining the power fluctuation period, the fluctuation changes of the remaining feature time series corresponding to each power fluctuation period are calculated, and the similarity with the power fluctuation changes in the power fluctuation period is used to evaluate the impact of each feature on power. However, it is obvious that power itself is not static, especially for active power. With the connection, disconnection or change of the operating status of other load devices in the concentrator application scenario, the power metering will also change accordingly. If some high-resistance load devices are connected to the concentrator network, the current collected by the concentrator may change greatly but the voltage change is small. Therefore, compared with voltage or current, the change of load power can more directly affect the change of power, so the change of load power corresponding to each power fluctuation period is further obtained.
[0053] The present invention uses the moving standard deviation and the 3σ principle to obtain the power fluctuation range, which enhances the model's sensitivity to power fluctuations, thereby more accurately identifying and calibrating measurement errors. Compared with existing methods, this technology can more accurately capture power fluctuations and avoid minor errors that may be ignored by traditional methods.
[0054] S3.2. Calculate the Pearson correlation coefficient between the load power time series and the electric energy time series corresponding to each electric energy fluctuation period to obtain the load fluctuation correlation, and record the electric energy fluctuation period corresponding to the load fluctuation correlation less than the set correlation threshold as a weakly correlated fluctuation period.
[0055] In this embodiment, the Pearson correlation coefficient between the load power time series and the electric energy time series corresponding to each electric energy fluctuation period can be calculated and recorded as the load fluctuation correlation. A larger load fluctuation correlation indicates that the change in electric energy during this period is mainly caused by the change in load power. In other words, this period cannot well reveal the degree of influence of other weakly correlated features on electric energy. Therefore, it is necessary to obtain the electric energy fluctuation period with a smaller load fluctuation correlation, and analyze the influence of other features on electric energy based on such electric energy fluctuation period. The above-mentioned method for obtaining the Pearson correlation coefficient belongs to the well-known technology and will not be described in detail here. After obtaining the load fluctuation correlation between the load power time series and the electric energy time series corresponding to all electric energy fluctuation periods, the maximum inter-class variance method is used to obtain the correlation threshold, and the electric energy fluctuation period corresponding to the load fluctuation correlation less than the correlation threshold is recorded as a weakly correlated fluctuation period. An exemplary explanation:
[0056] There are five periods of power fluctuation. The active power time series corresponding to these five periods of power fluctuation are recorded as , , , as well as ; Then obtain the load power time series corresponding to these five power fluctuation periods respectively, , , , as well as ; Further calculate the Pearson correlation coefficient of the active power time series and the load power time series corresponding to each power fluctuation period to obtain a total of 5 load fluctuation correlations, recorded as , , , as well as ; Then the maximum inter-class variance method is used to calculate the correlation threshold of all load fluctuation correlations. Assume that the calculated correlation threshold is After comparing all load fluctuation correlations with the threshold, the values less than the correlation threshold are and ; Finally, based on the smaller load fluctuation correlation and Get weakly correlated fluctuation period as well as .
[0057] S3.3. The trend correlation of each feature is obtained based on the Pearson correlation coefficient of the remaining feature time series and the electric energy time series in each weakly correlated fluctuation period. The influence weight of the decision tree is calculated based on the trend correlation of the feature corresponding to the node in each decision tree to obtain the weighted prediction value.
[0058] So far, the weakly correlated fluctuation period with a small correlation of load fluctuation is obtained. By analyzing the correlation between the remaining feature time series and the active power time series in each weakly correlated fluctuation period, the trend correlation of each feature is obtained, which is as follows:
[0059] The first The weakly correlated fluctuation period is recorded as , and the corresponding timings are: , , , , , , , as well as , also taking active energy as an example, for the first The input current time series corresponding to the weakly correlated fluctuation period , and the Pearson correlation coefficient is used to calculate its relationship with the active energy time series The correlation between , then the trend correlation calculation method of the input current is:
[0060] ;
[0061] in Indicates the trend correlation of input current; Indicates Weakly correlated volatility periods The corresponding input current timing Active energy timing Pearson correlation coefficient between ; Represents the total number of weakly correlated fluctuation periods. According to the above input current trend correlation calculation method, the trend correlations of input voltage, output current, output voltage, temperature and humidity are obtained respectively. , , , as well as ; As for the trend correlation of load power, its calculation method is different from the above characteristics, which is as follows: the power fluctuation period corresponding to the load fluctuation correlation greater than or equal to the correlation threshold is recorded as the strong correlation fluctuation period, and the mean value of the load fluctuation correlation in all strong correlation fluctuation periods is calculated to obtain the trend correlation of load power .
[0062] By calculating the trend correlation, the present invention can effectively evaluate the long-term impact of each feature on power fluctuations, so that the model can adapt to different working environments. This method avoids the decrease in prediction accuracy of traditional algorithms under different load and environmental conditions, thereby improving the stability and reliability of the model.
[0063] So far, the trend correlation of all features compared to active power has been obtained, and the trend correlation of features compared to reactive power can also be obtained by the above method, which will not be repeated here. For the prediction value of any decision tree, its influence weight on the final prediction value depends on the features selected by each node in the tree. The specific method for calculating the influence weight of a tree is:
[0064] ;
[0065] in Indicates The influence weight of each tree; Indicates The tree nodes; represents characteristic variables, including input current in this embodiment , Input voltage , output current , output voltage , Load power ,temperature and humidity ; Indicates The total number of nodes in a tree; Represents feature variables The corresponding trend correlation. Finally, the weighted prediction value of active power obtained by the random forest regression model is:
[0066] ;
[0067] in Represents the weighted predicted value of active electric energy; Indicates The standardized influence weight of each tree is obtained by using the maximum and minimum value method to normalize the influence weights of all trees. Indicates The predicted value of active power by each tree; It represents the total number of decision trees, and in this embodiment, the empirical value is 100. The method for obtaining the weighted prediction value of reactive power is the same as that for obtaining the weighted prediction value of active power, and will not be described in detail in this embodiment.
[0068] The present invention further improves the accuracy of the model by calculating the influence weight of each decision tree and comprehensively evaluating the contribution of feature variables to the power prediction results. Unlike the traditional random forest algorithm, the present invention assigns weights to each tree, which can better handle the role of multiple features in power error calibration and ensure more accurate prediction results.
[0069] S4. The error calibration of the concentrator is realized based on the random forest algorithm of weighted prediction value.
[0070] In the previous three steps, the acquisition and preprocessing of the concentrator data, the training of the random forest model and the calculation of the feature correlation are completed, and finally the calculation method of the weighted prediction value is obtained. In step S4, the error of the concentrator is calibrated in real time based on the trained regression model.
[0071] First, for the newly collected data, the processing method of step S1 can be used to obtain the preprocessed data, which is recorded as , split it into input features And feature labels , specifically:
[0072] ;
[0073] ;
[0074] in Represents the input features corresponding to the new data after preprocessing, including: the newly collected input current , Input voltage , output current , output voltage , Load power ,temperature ,humidity ; Indicates the feature labels corresponding to the new data after preprocessing, including: active energy And reactive energy .
[0075] Then input the features Input the random forest model trained in step S2, and obtain the active power prediction value and reactive power prediction value of the new data based on the weighted prediction value improvement method in step S3, which are recorded as and , then the electric energy error can be obtained, taking active electric energy as an example:
[0076] ;
[0077] in Indicates active energy error; Represents the weighted predicted value of active electric energy; Indicates the actual measured value of active electric energy. The final corrected value of active electric energy is:
[0078] ;
[0079] in Indicates the active electric energy correction value; Indicates the actual measured value of active electrical energy; represents the correction factor. In one embodiment, it can be set The empirical value is 0.3. In another embodiment, it can be determined by an incremental learning method, that is, when the past error is relatively stable and small, the correction coefficient can be set to a smaller value to reduce the adjustment of the active power error; if the error is large or an emergency occurs, the correction coefficient can be increased to correct the active power error to a greater extent; Represents the active electric energy error. The above incremental learning method belongs to the well-known technology, and the specific implementation process will not be repeated here. At the same time, the error calibration of reactive electric energy is the same as the above active electric energy error calibration method, and will not be described in detail.
[0080] Relevant technicians can also determine the operating status of the concentrator by analyzing the changes in the power error. When the power error continues to increase or becomes abnormally large, the concentrator can be manually inspected to determine the cause of the data anomaly.
[0081] The present invention uses a correction factor Dynamically adjusting the electric energy error solves the problem that the error correction process cannot adapt to actual environmental changes. Compared with the existing technology, this method is more flexible and can adjust the correction coefficient in real time according to the actual error situation, thereby improving the accuracy and adaptability of error correction.
[0082] The machine learning-based concentrator error calibration method of the present invention solves the problem that the existing technology cannot consider the different effects of different features on power fluctuations by introducing a random forest regression model and dynamically assigning the influence weight of each decision tree. Compared with traditional machine learning methods such as linear regression and Kalman filtering, random forests can capture the complex nonlinear relationship of power fluctuations and assign reasonable weights to different features through correlation analysis within the fluctuation period. This method can not only calibrate measurement errors, but also dynamically adjust according to real-time data to ensure that high-precision power prediction results are provided in various power system environments, thereby improving the accuracy of power metering and the stability of the system.
[0083] In the description of this specification, "plurality" or "several" means at least two, such as two, three or more, etc., unless otherwise clearly and specifically defined.
[0084] Although this specification has shown and described a number of embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will conceive of many modifications, changes and alternatives without departing from the ideas and spirit of the present invention. It should be understood that in the practice of the present invention, various alternatives to the embodiments of the present invention described herein may be employed.
Claims
1. A concentrator error calibration method based on machine learning, characterized in that: The method comprises: Acquire multiple parameter time series of the concentrator, wherein the multiple parameter time series include electric energy time series and load power time series; multiple parameters at the same acquisition time constitute a data set, and input the multiple data sets into a random forest regression model for training to obtain a trained regression model; Acquire a real-time data set of the concentrator, input the real-time data set into a trained regression model to obtain an electric energy prediction value; the difference between the electric energy prediction value and the real-time electric energy constitutes an electric energy error, and the sum of the product of the electric energy error and a set correction coefficient and the real-time electric energy is used as an electric energy correction value; The random forest regression model also includes assigning an influence weight to each decision tree, specifically: Obtain the electric energy fluctuation interval of the electric energy time series to obtain multiple fluctuation time periods, and calculate the load fluctuation correlation between the load power time series and the electric energy time series in each fluctuation time period; record the fluctuation time period corresponding to the load fluctuation correlation less than the set threshold as a weakly correlated fluctuation time period, and record the remaining fluctuation time periods as strongly correlated fluctuation time periods; The average of the Pearson correlation coefficients between the parameter time series except the load power time series and the electric energy time series in all weakly correlated fluctuation periods is taken as the trend correlation of the corresponding parameters; the average of the Pearson correlation coefficients between the load power time series and the electric energy time series in all strongly correlated fluctuation periods is taken as the trend correlation of the load power; the sum of the trend correlations corresponding to all parameters contained in each decision tree constitutes the influence weight of the corresponding decision tree.
2. The concentrator error calibration method based on machine learning according to claim 1, characterized in that: The calculation method of the influence weight is specifically as follows: ; in Indicates The influence weight of each tree; Indicates The tree nodes; Represents parameter variables; Indicates The total number of nodes in a tree; Represents parameter variables The corresponding trend correlation.
3. The concentrator error calibration method based on machine learning according to claim 2, characterized in that: The calculation method of the trend correlation of the corresponding parameters is specifically as follows: ; in Indicates other parameters except load power The corresponding trend correlation; Indicates Weakly correlated volatility periods Corresponding parameter timing Energy Sequence Pearson correlation coefficient between ; Represents the total number of weakly correlated volatility periods; The calculation method of the trend correlation of the load power is specifically as follows: the power fluctuation period corresponding to the load fluctuation correlation greater than or equal to the correlation threshold is recorded as a strong correlation fluctuation period, and the average of the load fluctuation correlation in all strong correlation fluctuation periods is calculated to obtain the trend correlation of the load power. .
4. The concentrator error calibration method based on machine learning according to claim 1, characterized in that: The plurality of parameter timings include: Input current timing, input voltage timing, output current timing, output voltage timing, load power timing, ambient temperature timing, ambient humidity timing, active power timing and reactive power timing.
5. The concentrator error calibration method based on machine learning according to claim 4, characterized in that: The random forest regression model also includes taking input current, input voltage, output current, output voltage, load power, ambient temperature, and ambient humidity as input features; Active power and reactive power are used as feature labels.
6. The concentrator error calibration method based on machine learning according to claim 1, characterized in that: The electric energy fluctuation interval of the electric energy time series is obtained to obtain multiple fluctuation time periods, including: Use the moving standard deviation algorithm to obtain the standard deviation of each sliding window in the power time series; The 3σ principle is used to screen standard deviations greater than μ±3σ to obtain multiple corresponding power fluctuation intervals, and all sampling moments of each power fluctuation interval constitute the corresponding fluctuation period.
7. The concentrator error calibration method based on machine learning according to claim 1, characterized in that: The setting correction coefficient is determined using an incremental learning method.
8. The concentrator error calibration method based on machine learning according to claim 1, characterized in that: The set threshold is obtained using the maximum inter-class variance method.
Citation Information
Patent Citations
Water body dissolved oxygen concentration prediction system
CN116338819A
Method for generating and using analgesia prediction model of analgesia pump
CN116543866A