A real-time reconstruction and online repair method for missing sensor data in energy monitoring systems
By establishing an energy regulatory system model and applying the maximum expectation, maximum likelihood estimation algorithm and Bayesian data reconstruction method, the problem of missing data is solved, real-time reconstruction and online repair of data is realized, and the accuracy and completeness of data are improved.
Patent Information
- Application Number
- CN202211016709.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-08-24
AI Technical Summary
Data incompleteness and real-time problems caused by missing data in sensors in energy regulatory systems, existing filling methods are inefficient and inaccurate.
Establish an energy supervision system model, use the maximum expectation method, the maximum likelihood estimation algorithm and Bayesian data reconstruction method, combine the MCMC algorithm for data reconstruction and online repair, and establish local mathematical models and similarity analysis, and use the meteorological database for data filling and correction.
Real-time reconstruction and online repair of missing data of sensors is realized, ensuring the accuracy and completeness of data, improving filling efficiency, and improving the intelligent energy monitoring system.
Smart Images

Figure CN115408375B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underlying sensor data processing in energy consumption monitoring systems, and relates to a real-time reconstruction and online repair method for missing sensor data in energy monitoring systems, and specifically to a method for restoring abnormal missing data of different types of sensors in energy monitoring systems. Background Art
[0002] With the rapid development of cutting-edge technologies such as big data and artificial intelligence, research in the energy sector is also gaining momentum in the field of intelligent control. Energy monitoring systems are inherently complex systems, characterized by complex structures, numerous variables, large scale, and susceptibility to interference. The entire system is equipped with numerous and diverse sensors. Data is collected by the sensors and sent to a gateway, which then transmits it to a server. In his paper "Research on Fault Detection Technology for Data-Sensitive Sensor Nodes in Harsh Environments," Cao Chuansong argues that data loss can occur during transmission due to the wide distribution and large number of sensors, poor network communication quality in some areas, and unstable network transmission at any stage. The causes of data loss in energy monitoring systems are generally as follows: 1) mechanical failures leading to failures in data collection, transmission, and storage, such as memory damage, network transmission failures, and noise interference; 2) information loss due to human error; and 3) inherent limitations of sensors, including limited energy, computing power, and storage capacity. Sensor failure is a key cause of data loss. Missing data can cause energy systems to operate in an unknown state, causing control systems to fall far short of their intended performance and potentially leading to further energy waste and economic losses. In their paper "Missing Data in Statistics and Solutions," Tong Xin et al. address these issues, and numerous experts and scholars at home and abroad have conducted extensive research using various methods. These methods are generally categorized into three main approaches: deletion, non-processing, and filling in missing data. However, for missing data in energy regulatory systems, the first two methods, while simple, discard information in exchange for a complete historical database, severely impacting the real-time objectivity and integrity of regulatory data. Filling in missing data is currently the most promising approach to addressing missing variable data in energy regulatory systems. While numerous filling-in algorithms exist, they often suffer from limitations. For example, manual data filling methods lack a rationale and are inefficient and inaccurate when addressing large amounts of missing data. After studying and researching numerous data reconstruction methods, we discovered that because energy monitoring systems often have numerous variables, multiple states, strong coupling, and long-term variable operating conditions, the most suitable approach to addressing data loss failures lies in the use of quantitative physical models, which establish lumped parameter models based on various conservation laws and the interdependencies between devices and systems. This paper addresses the challenges encountered by these methods and proposes a method for real-time reconstruction and online repair of missing sensor data in energy monitoring systems. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method capable of real-time reconstruction and online repair of missing sensor data in an energy monitoring system.
[0004] The technical solutions of the present invention are as follows:
[0005] A method for real-time reconstruction and online repair of missing sensor data in an energy monitoring system, comprising the following steps:
[0006] S1. Establish an energy supervision system model based on the design principles and operation mechanisms of the energy supervision system
[0007] S1.1. Define the energy regulatory system model
[0008] S1.1.1. All-air single return air reheating system model
[0009] The energy management system is mainly composed of filters, surface coolers, reheaters, supply fans and return fans on the air side. After the outdoor air and return air are mixed and filtered by the filter, dehumidified by the surface cooler, and heated by the reheater, the supply fan provides power and sends it into the room to complete a return air reheating cycle; the outdoor temperature is T1, the mixed air temperature is T2, the surface temperature of the surface cooler is T3, the supply air temperature is T4, the return air temperature is T6, and the outdoor relative humidity is The relative humidity of mixed air is The humidity of the surface cooler is The relative humidity at the air supply point is The relative humidity at the return air state point is The fresh air ratio is R, the system air volume is M2, and the reheater heat exchange is Q h The heat exchange of the surface cooler is Q e ;
[0010] S1.1.2. Refrigeration station system model
[0011] The energy management system on the freezing station side is mainly composed of chillers, water pumps, water distributors, water collectors, evaporators and cooling towers. The chilled water system is powered by a water pump to transport the return water in the water collector to the chiller, which is cooled by heat exchange in the evaporator and then transported to the water distributor to transfer the cold energy to the terminal for heat exchange with the air; the cooling water system is powered by a water pump to bring the return water into the condenser of the chiller to exchange heat with the refrigerant, and then transport the heated water to the cooling tower to dissipate the heat to the air to complete a cycle; the chilled water supply and return water temperatures are T s1 、T s4 The pressure difference on both sides of the manifold is P1, and the supply and return water temperatures of the cooling water are T s5 、T s6 , the pressure difference on both sides of the chilled water pump is P2, the pipe network impedance is S1, the pressure difference on both sides of the cooling water pump is P3, the heat exchange of chilled water is Q1, and the chilled water flow rate is M s1 , the cooling water heat exchange is Q2, the cooling water flow is M s2 ;
[0012] S1.2. Define the input and output variables in the model:
[0013] According to the full air return air reheating system model established in step S1.1.1, determine the input variable T1 on the air side. T4, T6, R, M2; output variable T2, T3, Q h , Q e ;
[0014] According to the refrigeration station system model established in step S1.1.2, determine the input variable T of the refrigeration station system side s1 , T s4 , T s5 , T s6 , S1, Q1, Q2; output variables P1, P2, P3, M s1 , M s2 ;
[0015] S1.3. Define the mathematical model:
[0016] The energy regulation system is a thermodynamic system. Local mathematical models are established based on the thermodynamic relationships of several key state points. Each local mathematical model is constructed according to the corresponding variable relationship.
[0017] In the full air primary return air reheating system model, the following local mathematical model is established:
[0018] (1) Local model 1
[0019] In the thermodynamic relationship, the mixed air state point is mathematically modeled by the following formula:
[0020] T2=R·T1+(1-R)T6 (1)
[0021] (2) Local model 2
[0022] Local model 2 is mainly composed of the following formulas:
[0023] lg(P vs )=AB / (T+C) (2)
[0024]
[0025]
[0026] Where, P vs Indicates the saturated water vapor pressure at a certain temperature; T represents the air temperature; A, B, and C represent the constants corresponding to different substances at different temperatures; Relative humidity; Pv represents the partial pressure of water vapor in the air; w represents the moisture content in the air; P represents the atmospheric pressure;
[0027] w2=R·w1+(1-R)·w6 (5)
[0028] Where, w2 represents the humidity at the mixed air state point; w1 represents the humidity of the outdoor air; w6 represents the humidity of the return air;
[0029] (3) Local model 3
[0030] Local model 3 mainly uses heat transfer knowledge to establish the following relationship:
[0031] Q h =cM(T4-T3) (6)
[0032] Where, c represents the constant pressure specific heat capacity of air; M represents the air supply volume;
[0033] (4) Local model 4
[0034] Local model 4 adds the thermodynamic relationship between enthalpy and heat to local model 2 to form a closed loop. The specific new formula is as follows:
[0035] h=cT+w(h g +c v T) (7)
[0036] Q c =M(h2-h3) (8)
[0037] Where h represents the air enthalpy; c v represents the average specific heat of water vapor at constant pressure; h g Indicates the latent heat of vaporization of water at 0℃; Q c Indicates the heat exchange amount of air passing through the surface cooler; h2 indicates the air enthalpy value at the mixed air state point; h 31 Indicates the enthalpy of air after passing through the surface cooler;
[0038] In the refrigeration station system, the following local mathematical model is established:
[0039] (1) Local model 5
[0040] The local model 5 is established based on the following relationship between the supply and return water temperatures, chilled water flow and heat exchange capacity in the chilled water system;
[0041] Q1=c w M S1 (T S4 -T S1 ) (9)
[0042] Where cw represents the specific heat capacity of water;
[0043] (2) Local model 6
[0044] Local model 6 uses the characteristics of the pipe network to establish a system model according to the following formula:
[0045]
[0046] (3) Local Model 7 and Local Model 9
[0047] Local model 7 and local model 9 are mainly based on the characteristic curve of the water pump itself. However, since the characteristic curve does not have an exact formula, the least squares method is used to fit multiple sets of real experimental data to obtain a mathematical model as shown in the following formula:
[0048]
[0049] Where H i represents the pump head; a and b are the unknown coefficients of the model; represents the cooling water or chilled water flow rate; m represents the exponent in the Reppinson formula;
[0050] (4) Local model 8
[0051] The local model 8 is established based on the following relationship between the supply and return water temperatures, cooling water flow and heat exchange capacity in the cooling water system:
[0052] Q2=c w M S2 (T S5 -T S6 ) (12)
[0053] S2. Similarity Analysis
[0054] Select the real meteorological data of the corresponding month of the previous year from the meteorological database, use the temperature and humidity at certain time points every day of that month to calculate the corresponding enthalpy value, and then compare it with the enthalpy value of the outdoor environment at the corresponding time on the day when the current data is missing. Select the most similar day's data as historical backup data, and multiple historical backup data will form a historical backup data set;
[0055] h=(1.01+1.84d)t+2500d (13)
[0056] Where t is the air temperature, d is the air moisture content, 1.01 kJ / (kg·K) is the specific heat of air at constant pressure, 1.84 kJ / (kg·K) is the average specific heat of water vapor at constant pressure, and 2500 kJ / kg is the latent heat of vaporization of water at 0°C.
[0057] S3. Method for constructing a filled data set based on the maximum expectation method and maximum likelihood estimation algorithm
[0058] S3.1. There are two types of data types in the energy monitoring system: stable data types with relatively small numerical variations, including temperature, humidity, flow rate, pressure, and pressure differentials across manifolds; and fluctuating data types with relatively large numerical variations, including pressure differentials across chilled water pumps and cooling water pumps.
[0059] S3.2. Fill in the missing stable data set based on the maximum expectation method and maximum likelihood estimation algorithm;
[0060] S3.2.1. Determine the sensor, date, and time period where data is missing. Suppose that on day j, a sensor has lost n consecutive data points since time i. The endpoint data of this time period are X i,j and X i+n,j , and their two adjacent points are X i-1,j and X i+n+1,j ; Still sort all the remaining data of the day from small to large, and determine the X after sorting i-1,j and X i+n+1,j The new location, record n i-1 and n i+j+1 ;
[0061] S3.2.2. Select the most similar working condition k-th day historical data and sort them from small to large throughout the day, and find the corresponding and Two points, each of which is extended inward by h points to form a preliminary filling data set of two missing endpoints, where the number of h points must be less than n;
[0062] S3.2.3. Use the EM and MLE algorithms to reverse-estimate the normal distribution parameters of the historical backup dataset. Then, randomly sample the required amount of data from the normal distribution and the historical backup dataset to form the final filling dataset, ensuring the data volume requirements of the system model for the filling dataset.
[0063] S3.2.4. Repeat step 3.2.1 until the missing data segment is completely filled;
[0064] S3.3. Fill in missing data sets of fluctuating data based on the maximum expectation method, maximum likelihood estimation algorithm and the introduction of auxiliary variables;
[0065] S3.3.1. Determine the missing data point X i,k Or data segment X i,k ~X i+n,k And the corresponding variable data point Y at the corresponding time i,k or Y i,k ~Yi+n,k ;
[0066] S3.3.2. Sort the auxiliary variable Y data on the day of missing data from small to large, and find Y i,k or Y i,k ~Y i+n,k The position of each missing point in the new sequence is recorded as n i ;
[0067] S3.3.3. Extend h data points on the left and right sides of the missing point as needed, and find these points respectively. The value of the missing variable X at the original corresponding time; forming a preliminary filling data set;
[0068] S3.3.4. Use the EM and MLE methods to estimate the parameters of the preliminary populated dataset, then randomly extract an appropriate amount of data from the preliminary populated dataset after parameter estimation to obtain the complete populated dataset that is finally substituted into the likelihood function;
[0069] S4. Bayesian data reconstruction method based on MCMC algorithm
[0070] S4.1. Define the benchmark value: The benchmark value is a function value composed of the accurate values of other regulatory data in the system when a variable is missing data. b It is expressed as shown in the following formula (14):
[0071] Y b =f(Y o1 ,Y o2 ,…,Y on ) (14)
[0072] Among them, Y oi represents the supervised data values of other working sensors in the system at the moment when a variable is missing data; the undetected variables in the system at the moment when a variable is missing data are represented as unknown variables in formula (14) and substituted. The solution can be obtained under the condition that the number of data groups retrieved from steps S3.2.4 and S3.3.4 is greater than the number of sensors with missing data;
[0073] S4.2. Define the correction function: The correction function is composed of the historical data in the processed and filtered filling data set and the deviation to be estimated. c It is expressed as shown in the following formula (15):
[0074] Y c =f(Y i ,x) (15)
[0075] Among them, Y irepresents the i-th historical data in the filtered and processed filled data set; x represents an unknown number, which is the deviation between the true value of the missing data and the historical data value at a certain moment;
[0076] S4.3. The distance function corrects the missing data by minimizing the difference between the reference value and the correction function, that is, by repeatedly substituting the historical data in the filled data set, as shown in formula (16):
[0077]
[0078] S4.4, Substitute data: Substitute the multiple groups of steady-state measurement values T1 in S2, T4, T6, R, M2, T2, T3, Q h , Q e ;
[0079] S4.5. After defining the distance function, the distance function D(x) is made to satisfy a Gaussian distribution with a mean of 0. For the energy supervision system, since the data of each variable are measured in situ by various types of sensors in the system and uploaded to the database, the variance of the normal distribution is set to the accuracy value of the sensor, which constitutes the likelihood function of the Bayesian model of the energy supervision system, as shown in formula (17). By minimizing the distance function, a high degree of consistency between the estimated value and the true value is achieved.
[0080] Likelihood function:
[0081] Bayes' Theorem:
[0082] S4.6. When the value of the distance function is small, the likelihood function probability is the highest. According to the posterior distribution formula (18) of the Bayesian model of the energy regulatory system, it can be seen that P(Y b ) is a constant, the posterior P(Y b |x) also obeys the Gaussian distribution. In this case, the mean of the posterior distribution is the difference between the accurate value and the estimated value, which represents the distance between the true value and the historical data in the filled data set. Finally, the result x is added to the median of the sorted filled data set to obtain the Bayesian method reconstruction value of the missing data; as shown in formula (19), Y t represents the reconstructed value of missing data, Y c represents the estimated value of the missing data, and x is the difference between the reconstructed value and the estimated value obtained by solving the posterior distribution of Bayesian inference;
[0083] Y t =Y c +x (19)
[0084] The beneficial effects of the present invention are as follows: Based on data and models, the present invention can reconstruct and repair the data missing of one or multiple types of sensors in the energy monitoring system in real time and online, which can ensure the accuracy of the estimated value, and has high filling efficiency and large filling dimension. At the same time, it realizes the accuracy of historical data and the integrity of the database, which is of great significance to the improvement of the intelligent energy monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 This is a framework diagram of a real-time reconstruction and online repair method for missing sensor data in an energy monitoring system.
[0086] Figure 2 This is a flow chart of a method for real-time reconstruction and online repair of missing sensor data in an energy monitoring system.
[0087] Figure 3 This is a diagram of the full air single return air reheating system.
[0088] Figure 4 This is a diagram of the refrigeration station system.
[0089] Figure 5 Reconstruct the result map for the missing outdoor temperature scattered point data.
[0090] Figure 6 This is the reconstruction result diagram for the missing data of the mixed air temperature dispersion points.
[0091] Figure 7 Result graph of reconstruction for missing data of scattered points of chilled water supply temperature.
[0092] Figure 8 Reconstruction result graph for missing data of scattered points of cooling water supply temperature.
[0093] Specific implementation
[0094] The specific embodiments of the present invention are described in detail below in conjunction with the invention content, the accompanying drawings and the formulas of the specification.
[0095] Reference Figure 2 The present invention is a method for real-time reconstruction and online repair of missing sensor data in an energy monitoring system. Taking the real-time reconstruction and online repair of missing sensor data in a public central air-conditioning system in Dalian as an example, the specific steps are as follows:
[0096] S1. Establishment of the mathematical model of the central air-conditioning system. Based on the design principle and operation mechanism of the central air-conditioning system, the system model is established. The specific steps are as follows:
[0097] S1.1. Define the central air conditioning system model:
[0098] S1.1.1. All-air single return air reheating system model
[0099] Reference Figure 3 The central air conditioning system consists of a filter, a surface cooler, a reheater, a supply fan and a return fan on the air side. The outdoor temperature is T1, the mixed air temperature is T2, the surface temperature of the surface cooler is T3, the supply air temperature is T4, the return air temperature is T6, and the outdoor relative humidity is Mixed air relative humidity Humidity through the surface cooler Relative humidity at air supply point Relative humidity at return air state point Fresh air ratio R, system air supply volume M2, reheater heat exchange Q h , heat exchange of the surface cooler Q e The outdoor air and return air are mixed, filtered, dehumidified by the surface cooler, and heated by the reheater. Then, the air is powered by the blower and sent into the room, completing a single return air reheating cycle.
[0100] S1.1.2. Refrigeration station system model
[0101] Reference Figure 4 The central air-conditioning system is composed of chillers, water pumps, water distributors, and water collectors on the refrigeration station side. The supply and return water temperatures are T s1 、T s4 , pressure difference P1 on both sides of the manifold, cooling water supply and return water temperature T s5 、T s6 , pressure difference on both sides of the chilled water pump P2, pipe network impedance S1, pressure difference on both sides of the cooling water pump P3, chilled water heat exchange Q1 chilled water flow M s1 , cooling water heat exchange Q2, cooling water flow M s2 The chilled water system is powered by a water pump that transfers return water from the water collector to the chiller. After cooling through the evaporator, it is then transferred to the water distributor, where the cooling energy is transferred to the end point for heat exchange with the air. The cooling water system also has a circulating water pump that drives the return water into the chiller's condenser for heat exchange with the refrigerant. The heated water is then transferred to the cooling tower to dissipate the heat into the air, completing the cycle. The central air conditioning system is a thermodynamic system, and local mathematical models are established based on the thermodynamic relationships of several key state points.
[0102] S1.2. Define the input and output variables in the model:
[0103] According to the mathematical model established in S1.1.1, determine the input variable T1 on the air side. T4, T6, R, M2; output variable T2, T3, Q h , Q e .
[0104] According to the mathematical model established in S1.1.2, determine the input variable T of the refrigeration station system side s1 , T s4 , T s5 , T s6 , S1, Q1, Q2; output variables P1, P2, P3, M s1 , M s2 .
[0105] S1.3. Define the mathematical model:
[0106] The thermodynamic relationships of several key state points are used to establish local mathematical models of the central air-conditioning system. Each model is constructed according to the corresponding variable relationship. In the full air return air reheating system, local model 1 is composed of variables outdoor temperature T1, mixed air temperature T2, return air temperature T6 and fresh air ratio R; local model 2 mainly studies the outdoor relative humidity. Mixed air relative humidity Return air relative humidity Local model 3 for reheater heat transfer Q h Establish the relationship with the inlet and outlet temperatures T3 and T4. Local model 4 includes the inlet and outlet humidity of the reheater In the refrigeration station system, local model 1 is composed of the variable chilled water supply temperature T s1 , chilled water return temperature T s4 , chilled water flow M s1 , the chilled water heat exchange Q1; the main research objects of local model 2 include the pressure difference P1 between the distribution and collector, the pipe network impedance S1, the chilled water flow M s1 The main research objects of local model 3 include the pressure difference P2 on both sides of the chilled water pump, the chilled water flow rate M s1 ; Local model 4: cooling water supply temperature T s5 , cooling water return temperature T s6 , cooling water flow M s2 , cooling water heat exchange Q2; local model 5: pressure difference P3 on both sides of the cooling water pump, cooling water flow M s2 .
[0107] S2. Similarity analysis. The specific steps are as follows:
[0108] We select real meteorological data from the corresponding month of the previous year from the Dalian meteorological database, calculate the corresponding enthalpy value using the temperature and humidity at certain times of the day in that month, and then compare it with the enthalpy value of the outdoor environment at the corresponding time on the day of the missing data. We select the most similar day's data as historical backup data, and combine multiple historical backup data into a historical backup dataset.
[0109] S3. A method for constructing a filled data set based on the maximum expectation method and maximum likelihood estimation algorithm. The specific steps are as follows:
[0110] S3.1. In the case of missing stable data, based on the characteristics of small numerical span, the data to be studied are respectively reordered and two-point positioned.
[0111] S3.2. Fill in the data set with missing stable data segments based on the maximum expectation method and maximum likelihood estimation algorithm.
[0112] S3.2.1. Determine the sensor, date, and time period where data is missing. Suppose that on day j, a sensor has lost n consecutive data points since time i. The endpoint data of this time period are X i,j and X i+n,j , and their two adjacent points are X i-1,j and X i+n+1,j . All remaining data of the day are still sorted from small to large, and the sorted X is determined. i-1,j and X i+n+1,j The new location, record n i-1 and n i+j+1 .
[0113] S3.2.2. Select the kth historical data of the most similar working condition and sort them from small to large throughout the day. and Two points, each of which is extended inward by h points to form a preliminary filling data set of two missing endpoints, where the number of h points is less than n.
[0114] S3.2.3. Use the EM and MLE algorithms to reversely estimate the normal distribution parameters of the preliminary filled data set using the data set, and then randomly extract a certain amount of data from the distribution as needed and combine it with the preliminary data set to form the final filled data set. This ensures that the system model has the data volume requirements for the filled data set.
[0115] S3.2.4. Repeat the first step until the missing data segment is complete. For the type of missing data segment, select a small amount of data from the historical database so that the number of extended data points, h, is less than the number of missing data points, n. Leverage the two parameter estimates to fill in the small amount of data and identify the distribution that the missing data points follow.
[0116] S3.3. Based on the maximum expectation method and maximum likelihood estimation algorithm and the introduction of auxiliary variables, the dataset with missing fluctuation data is filled.
[0117] S3.3.1. Determine the missing data point X i,k Or data segment X i,k ~X i+n,kAnd the auxiliary variable data point Y at the corresponding moment i,k or Y i,k ~Y i+n,k .
[0118] S3.3.2. Sort the auxiliary variable Y data on the day of missing data from small to large, and find Y i,k or Y i,k ~Y i+n,k The position of each missing point in the new sequence is recorded as n i .
[0119] S3.3.3. Extend h data points to the left and right of the point as needed, and find these points respectively. The values of the missing variable X at the original corresponding time constitute the preliminary filling data set.
[0120] S3.3.4. Use the EM and MLE methods to estimate the parameters of the preliminary filled data set, and then randomly extract an appropriate amount of data from the estimated distribution as needed to obtain the complete filled data set that is finally substituted into the likelihood function.
[0121] S4. Bayesian data reconstruction method based on MCMC algorithm. The specific steps are as follows:
[0122] S4.1. Definition of baseline value: It is a function value composed of the accurate values of other regulatory data in the system when a variable is missing data. b It is expressed as shown in the following formula (1):
[0123] Benchmark function: Y b =f(Y o1 ,Y o2 ,…,Y on ) (1)
[0124] where Y oi Represents the monitoring data value of other working sensors in the system at that moment. The undetected variables in the system can be expressed as unknown variables in formula (1) and substituted into them. As long as the number of data sets is greater than the number of unknowns, the solution can be obtained.
[0125] S4.2. Define the correction function: The correction function is composed of the historical data in the processed and filtered filling data set and the deviation to be estimated. c It is expressed as shown in the following formula (2):
[0126] Correction function: Y c =f(Y i ,x) (2)
[0127] where Y irepresents the i-th historical data in the filtered and processed filled data set; x represents an unknown number, which is the deviation between the true value of the missing data and the historical data value at a certain moment.
[0128] S4.3. The distance function corrects the missing data by minimizing the difference between the reference value and the correction function, that is, by repeatedly substituting the historical data in the filled data set, as shown in formula (3):
[0129] Distance function:
[0130] S4.4, Substitute data: Substitute the multiple groups of steady-state measurement values T1 in S2, T4, T6, R, M2, T2, T3, Q h , Q e .
[0131] S4.5. After defining the distance function, the distance function D(x) is made to satisfy a Gaussian distribution with a mean of 0. For the specific energy regulation system, since the data of each variable is measured in situ by various types of sensors in the system and uploaded to the database, the variance of the normal distribution is set to the accuracy value of the sensor, which constitutes the likelihood function of the Bayesian model of the energy regulation system, as shown in Formula (4). By minimizing the distance function, a high degree of consistency between the estimated value and the true value is achieved.
[0132] Likelihood function:
[0133] Bayes' Theorem:
[0134] S4.6. When the value of the distance function is small, the likelihood function probability is the highest. According to the posterior distribution formula (5) of the Bayesian model of the energy regulatory system, it can be seen that P(Y b ) is a constant, the posterior P(Y b |x) also obeys Gaussian distribution. In this case, the mean of the posterior distribution is the difference between the accurate value and the estimated value, which represents the distance between the true value and the historical data in the filled data set. Finally, the result x is added to the median of the sorted filled data set to obtain the Bayesian method reconstruction value of the missing data. As shown in formula (6), Y t represents the reconstructed value of missing data, Y c represents the estimated value of the missing data, and x is the difference between the reconstructed value and the estimated value obtained by solving the posterior distribution of Bayesian inference.
[0135] Reconstructed value: Y t =Y c +x (6)
[0136] In the central air conditioning system, the outdoor temperature is T1, the mixed air temperature is T2, and the chilled water supply temperature is T s1 and cooling water supply temperature T s5 The data reconstruction results are shown in Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 The reconstruction results show that the minimum relative error can reach 0.004% and the maximum is only 3.548%, which preliminarily verifies the feasibility of this reconstruction method.
Claims
1. A method for real-time reconstruction and online repair of missing data of sensors in an energy monitoring system, characterized in that: Here are the steps: S1. Establish an energy supervision system model based on the design principles and operation mechanisms of the energy supervision system S1.
1. Define the energy regulatory system model S1.1.
1. All-air single return air reheating system model The energy management system is mainly composed of filters, surface coolers, reheaters, supply fans and return air fans on the air side. After the outdoor air and return air are mixed and filtered by the filter, dehumidified by the surface cooler, and heated by the reheater, the supply fan provides power and sends it into the room to complete a return air reheating cycle; the outdoor temperature is T1, the mixed air temperature is T2, the surface temperature of the surface cooler is T3, the supply air temperature is T4, the return air temperature is T6, and the outdoor relative humidity is The relative humidity of mixed air is The humidity of the surface cooler is The relative humidity at the air supply point is The relative humidity at the return air state point is The fresh air ratio is R, the system air volume is M2, and the reheater heat exchange is Q h The heat exchange of the surface cooler is Q e ; S1.1.
2. Refrigeration station system model The energy management system on the freezing station side is mainly composed of chillers, water pumps, water distributors, water collectors, evaporators and cooling towers. The chilled water system is powered by a water pump to transport the return water in the water collector to the chiller, which is cooled by heat exchange in the evaporator and then transported to the water distributor to transfer the cold energy to the terminal for heat exchange with the air; the cooling water system is powered by a water pump to bring the return water into the condenser of the chiller to exchange heat with the refrigerant, and then transport the heated water to the cooling tower to dissipate the heat to the air to complete a cycle; the chilled water supply and return water temperatures are T s1 、T s4 The pressure difference on both sides of the manifold is P1, and the supply and return water temperatures of the cooling water are T s5 、T s6 , the pressure difference on both sides of the chilled water pump is P2, the pipe network impedance is S1, the pressure difference on both sides of the cooling water pump is P3, the heat exchange of chilled water is Q1, and the chilled water flow rate is M s1 , the cooling water heat exchange is Q2, the cooling water flow is M s2 ; S1.2, define the input and output variables in the model: According to the full air return air reheating system model established in step S1.1.1, determine the input variable T1 on the air side. T4, T6, R, M2; output variable T2, T3, Q h , Q e ; According to the refrigeration station system model established in step S1.1.2, determine the input variable T of the refrigeration station system side s1 , T s4 , T s5 , T s6 , S1, Q1, Q2; output variables P1, P2, P3, M s1 , M s2 ; S1.
3. Define the mathematical model: The energy regulation system is a thermodynamic system. Local mathematical models are established based on the thermodynamic relationships of several key state points. Each local mathematical model is constructed according to the corresponding variable relationship. In the full air primary return air reheating system model, the following local mathematical model is established: (1) Local model 1 In the thermodynamic relationship, the mixed air state point is mathematically modeled by the following formula: T2=R·T1+(1-R)T6 (1) (2) Local model 2 Local model 2 is mainly composed of the following formulas: lg(P vs )=A-B / (T+C) (2) Where, P vs Indicates the saturated water vapor pressure at a certain temperature; T represents the air temperature; A, B, and C represent the constants corresponding to different substances at different temperatures; Relative humidity; P v represents the partial pressure of water vapor in the air; w represents the moisture content in the air; P represents the atmospheric pressure; w2=R·w1+(1-R)·w6 (5) Where, w2 represents the humidity at the mixed air state point; w1 represents the humidity of the outdoor air; w6 represents the humidity of the return air; (3) Local model 3 Local model 3 mainly uses heat transfer knowledge to establish the following relationship: Q h =cM(T4-T3) (6) Where, c represents the constant pressure specific heat capacity of air; M represents the air supply volume; (4) Local model 4 Local model 4 adds the thermodynamic relationship between enthalpy and heat to local model 2 to form a closed loop. The specific new formula is as follows: h=cT+w(h g +c v T) (7) Q c =M(h2-h3) (8) Where h represents the air enthalpy; c v represents the average specific heat of water vapor at constant pressure; h g Indicates the latent heat of vaporization of water at 0℃; Q c It represents the heat exchange amount of air passing through the surface cooler; h2 represents the air enthalpy value at the mixed air state point; h3 represents the air enthalpy value after passing through the surface cooler; In the refrigeration station system, the following local mathematical model is established: (1) Local model 5 The local model 5 is established based on the following relationship between the supply and return water temperatures, chilled water flow and heat exchange capacity in the chilled water system; Q1=c w M S1 (T S4 -T S1 ) (9) Where c w represents the specific heat capacity of water; (2) Local model 6 Local model 6 uses the characteristics of the pipe network to establish a system model according to the following formula: (3) Local Model 7 and Local Model 9 Local model 7 and local model 9 are mainly based on the characteristic curve of the water pump itself. However, since the characteristic curve does not have an exact formula, the least squares method is used to fit multiple sets of real experimental data to obtain a mathematical model as shown in the following formula: Where H i represents the pump head; a and b are the unknown coefficients of the model; represents the cooling water or chilled water flow rate; m represents the exponent in the Reppinson formula; (4) Local model 8 The local model 8 is established based on the following relationship between the supply and return water temperatures, cooling water flow and heat exchange capacity in the cooling water system: Q2=c w M S2 (T S5 -T S6 ) (12) S2. Similarity Analysis Select the real meteorological data of the corresponding month of the previous year from the meteorological database, use the temperature and humidity at certain time points every day of that month to calculate the corresponding enthalpy value, and then compare it with the enthalpy value of the outdoor environment at the corresponding time on the day when the current data is missing. Select the most similar day's data as historical backup data, and multiple historical backup data will form a historical backup data set; h=(1.01+1.84d)t+2500d (13)where t represents the air temperature, d represents the air moisture content, 1.01kJ / (kg·K) is the constant-pressure specific heat of air, 1.84kJ / (kg·K) is the average constant-pressure specific heat of water vapor, and 2500kJ / kg is the latent heat of vaporization of water at 0°C. S3. Method for constructing a filled data set based on the maximum expectation method and maximum likelihood estimation algorithm S3.
1. There are two types of data types in the energy monitoring system: stable data types, including temperature, humidity, flow, pressure, and pressure difference between the manifold and the manifold; and fluctuating data types, including pressure difference between the chilled water pump and the cooling water pump. S3.
2. Fill in the missing stable data set based on the maximum expectation method and maximum likelihood estimation algorithm; S3.2.
1. Determine the sensor, date, and time period where data is missing. Suppose that on day j, a sensor has lost n consecutive data points since time i. The endpoint data of this time period are X i,j and X i+n,j , and their two adjacent points are X i-1,j and X i+n+1,j ; Still sort all the remaining data of the day from small to large, and determine the X after sorting i-1,j and X i+n+1,j The new location, record n i-1 and n i+j+1 ; S3.2.
2. Select the most similar working condition k-th day historical data and sort them from small to large throughout the day, and find the corresponding and Two points, each of which is extended inward by h points to form a preliminary filling data set of two missing endpoints, where the number of h points must be less than n; S3.2.
3. Use the EM and MLE algorithms to reverse-estimate the normal distribution parameters of the historical backup dataset. Then, randomly sample the required amount of data from the normal distribution and the historical backup dataset to form the final filling dataset, ensuring the data volume requirements of the system model for the filling dataset. S3.2.
4. Repeat step 3.2.1 until the missing data segment is completely filled; S3.
3. Fill in missing data sets of fluctuating data based on the maximum expectation method, maximum likelihood estimation algorithm and the introduction of auxiliary variables; S3.3.
1. Determine the missing data point X i,k Or data segment X i,k ~X i+n,k And the corresponding variable data point Y at the corresponding time i,k or Y i,k ~Y i+n,k ; S3.3.
2. Sort the auxiliary variable Y data on the day of missing data from small to large, and find Y i,k or Y i,k ~Y i+n,k The position of each missing point in the new sequence is recorded as n i ; S3.3.
3. Extend h data points on the left and right sides of the missing point as needed, and find these points respectively. The value of the missing variable X at the original corresponding time; forming a preliminary filling data set; S3.3.
4. Use the EM and MLE methods to estimate the parameters of the preliminary populated dataset, then randomly extract an appropriate amount of data from the preliminary populated dataset after parameter estimation to obtain the complete populated dataset that is finally substituted into the likelihood function; S4. Bayesian data reconstruction method based on MCMC algorithm S4.
1. Define the benchmark value: The benchmark value is a function value composed of the accurate values of other regulatory data in the system when a variable is missing data. b It is expressed as shown in the following formula (14): AND b =f(Y o1 ,AND o2 ,…,AND on ) (14) Among them, Y oi represents the supervised data values of other working sensors in the system at the moment when a variable is missing data; the undetected variables in the system at the moment when a variable is missing data are represented as unknown variables in formula (14) and substituted. The solution can be obtained under the condition that the number of data groups retrieved from steps S3.2.4 and S3.3.4 is greater than the number of sensors with missing data; S4.
2. Define the correction function: The correction function is composed of the historical data in the processed and filtered filling data set and the deviation to be estimated. c It is expressed as shown in the following formula (15): Y c =f(Y i ,x) (15) Among them, Y i represents the i-th historical data in the filtered and processed filled data set; x represents an unknown number, which is the deviation between the true value of the missing data and the historical data value at a certain moment; S4.
3. The distance function corrects the missing data by minimizing the difference between the reference value and the correction function, that is, by repeatedly substituting the historical data in the filled data set, as shown in formula (16): S4.4, Substitute data: Substitute the multiple groups of steady-state measurement values T1 in S2, T4, T6, R, M2, T2, T3, Q h , Q e Substitute into the distance function D(x) defined in S4.3; S4.
5. After defining the distance function, the distance function D(x) is made to satisfy a Gaussian distribution with a mean of 0. For the energy supervision system, since the data of each variable are measured in situ by various types of sensors in the system and uploaded to the database, the variance of the normal distribution is set to the accuracy value of the sensor, which constitutes the likelihood function of the Bayesian model of the energy supervision system, as shown in formula (17). By minimizing the distance function, a high degree of consistency between the estimated value and the true value is achieved. Likelihood function: Bayes' Theorem: S4.
6. When the value of the distance function is the smallest, the likelihood function probability is the highest. According to the posterior distribution formula (18) of the Bayesian model of the energy regulatory system, it can be seen that P(Y b ) is a constant, the posterior P(Y b |x) also obeys the Gaussian distribution. In this case, the mean of the posterior distribution is the difference between the accurate value and the estimated value, which represents the distance between the true value and the historical data in the filled data set. Finally, the result x is added to the median of the sorted filled data set to obtain the Bayesian method reconstruction value of the missing data; as shown in formula (19), Y t represents the reconstructed value of missing data, Y c represents the estimated value of the missing data, and x is the difference between the reconstructed value and the estimated value obtained by solving the posterior distribution of Bayesian inference; AND t =And c +x (19)。
Citation Information
Patent Citations
Intelligent-power-grid-oriented missing data filling method
CN104133866A
Voltage missing value filling method based on historical data auxiliary scene analysis
CN111507412A