Method and device for continuously monitoring state of large-scale building dynamic model
By introducing a user comfort model and a Gaussian filter, and combining the Cook distance method to clean the building load curve data, the error problem in the monitoring of building dynamic model status was solved, and continuous monitoring of the status of large-scale building dynamic models was realized, improving the accuracy and stability of the analysis.
Patent Information
- Application Number
- CN202510870692.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
Smart Images

Figure CN120810568A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of power system load forecasting and analysis, and particularly relates to a method and device for continuously monitoring the state of a large-scale building dynamic model. BACKGROUND
[0002] With the development of economy, the demand for energy is increasing, and the pressure on the power system is increasing; in particular, the power load of building load is increasing, and the proportion in the overall power load is significantly improved. It is very important to model, analyze, and optimize the operation of buildings to realize the continuous monitoring of the state of the large-scale building dynamic model. As a key step, the modeling and data processing method of large-scale building is not only the basis of building-related power system load forecasting and analysis technology, but also a necessary step to obtain accurate results. It has become a key technology in power system load forecasting and analysis. Among them, the combination of building-related thermal models and power load models can more comprehensively reflect the building-related information of the power system load, and provide an important model and data basis for building-related research.
[0003] The commonly used data in building modeling is the load curve, which is a typical nonlinear time series data with obvious volatility and seasonality. However, the traditional load curve analysis method such as K-Means clustering mainly based on statistical principles often cannot accurately describe the complex characteristics of the load curve, making it difficult to accurately capture the nonlinear characteristics of the load curve, resulting in inaccurate analysis results.
[0004] At the same time, in order to realize the continuity of the results of building modeling and data processing, there are some researches on dynamic modeling and data processing, such as the research on building flexible load modeling and day-ahead peak shifting scheduling strategy in the paper "Integrated Intelligent Building Electric / Gas / Heat Regional Integrated Energy System Modeling and Operation Optimization Research", which proposes a coordinated optimization scheduling model and algorithm for intelligent building and regional integrated energy system (Integrated Community Energy Systems, ICES). However, it lacks balanced consideration of the consumer side and other aspects, and does not consider the economic factors of different subjects, resulting in a lack of attention to user comfort and a lack of consideration of economic factors, which is too high in cost and not conducive to the application and implementation of the scheme. At the same time, since the building is a complex system, there is a large error between the model results of the proposed scheme and the actual situation, which cannot obtain an accurate model and cannot realize the continuous monitoring of the state of the large-scale building dynamic model.
[0005] Therefore, how to consider multiple constraint factors on the basis of large-scale building modeling conforming to physical laws such as thermodynamics, and meanwhile find an effective data processing method, so as to realize continuous monitoring of the state of the large-scale building dynamic model is a problem to be solved at present. SUMMARY
[0006] To solve the above technical problems, the application provides a continuous monitoring method and device for the state of a large-scale building dynamic model, which introduces user comfort in modeling, considers random noise in the load curve, adds a chain equation on the basis of the existing Cook distance, and uses a Gaussian filter to set a corresponding threshold for each period for filtering and reducing data errors, so that the daily electricity consumption curve of a single user tends to be smooth, the stability of the result is improved, the defect that the current technology has a large deviation between the large-scale building dynamic model modeling and the actual operation is made up, and continuous monitoring of the state of the large-scale building dynamic model is realized.
[0007] The continuous monitoring method for the state of the large-scale building dynamic model comprises the following steps:
[0008] Based on the characteristics of indoor gas and walls, a large-scale building dynamic model comprising an indoor thermodynamic model and a temperature control load dynamic model is established;
[0009] Considering user comfort, cost constraints and equipment constraints, a user comfort model is constructed;
[0010] Based on the large-scale building dynamic model and the user comfort model, a single user daily load curve is extracted, the data quantity and quality are considered, and the extracted single user daily load curve data are used to realize continuous monitoring of the state of the large-scale building dynamic model.
[0011] Further, based on the characteristics of indoor gas and walls, an indoor thermodynamic model is established, which comprises: the indoor thermodynamic model is modeled with indoor air specific heat capacity and wall heat capacity as heat capacity variables, and wall thermal resistance and indoor gas thermal resistance as thermal resistance variables; the differential equation of the indoor thermodynamic model is represented as:
[0012]
[0013] wherein m a is the indoor air mass, H g and H l are the increased and decreased heat in the room, T1(t i ), T2(t i ) and T3(t i ) are the indoor air temperature, indoor wall temperature and outdoor air temperature at t iwhere T (t) is the indoor air temperature at time t, T (t) is the wall temperature at time t, T (t) is the indoor air temperature at time t, Q(t) is the cooling / heating load of the temperature-controlled load, C1 is the specific heat capacity of the indoor air, C2 is the heat capacity of the wall, R1 is the thermal resistance of the wall, R2 is the thermal resistance of the indoor air, d is a differential symbol, and dt is a differential with respect to time.
[0014] Further, a dynamic model of the temperature-controlled load is established based on the characteristics of the indoor air and the wall, including: modeling the dynamic model of the temperature-controlled load with the indoor air temperature as the air temperature variable, the working state of the single temperature-controlled load as the state carrier, and the refrigerating capacity as the working variable:
[0015]
[0016] where S Q (t) is the working state of the single temperature-controlled load at time t, 1 indicates normal operation, and 0 indicates standby; 3,t T (t) is the indoor air temperature at time t, T 3,set is the set value of the temperature-controlled load, and Ω represents the comfort of the human body.
[0017] p out (t) = P × S Q (t)
[0018] where p out (t) is the actual power output by the temperature-controlled load, and P is the rated power.
[0019] If the air conditioner in the temperature-controlled load is in a cooling state, the refrigerating capacity output by the temperature-controlled load is:
[0020]
[0021] where Q(t i ) is the actual refrigerating capacity output by the temperature-controlled load at time t i , η Q (t i ) is the actual energy efficiency ratio at time t i , and θ and δ are fitting coefficients.
[0022] Further, considering the user comfort, cost constraints, and equipment constraints, a user comfort model is constructed, wherein the user comfort is considered, the index PMV for evaluating the thermal environment is simplified and modified based on the index PMV, and the user comfort index is analyzed by dimension reduction:
[0023]
[0024] where n pmv represents the simplified PMV expression, n′ pmv represents the simple calculation formula after dimension reduction, and n′ averepresents overall comfort, M represents metabolic rate, L represents individual thermal load of user, t cl represents temperature at which user feels comfortable, t a represents room temperature, p cl represents clothing thermal resistance, R h represents relative humidity, n" pmv represents optimal comfort value; k represents that the entire data is divided into k parts, j represents the jth part of the data after division; n' pmv,j represents comfort degree obtained by the jth part of data.
[0025] Further, considering user comfort, cost constraints and equipment constraints, a user comfort model is constructed, wherein, considering cost constraints and equipment constraints, a single building is adopted to cooperate in a game, based on collected working states of devices in the building and dispatching cost, power dispatching and economic adjustment are performed by using temperature control load, and cost optimization under consideration of user comfort is realized:
[0026]
[0027] wherein, a i is a linear term coefficient, b i is a quadratic term coefficient, c i is a constant term, P t+1 represents total power required at a subsequent moment t+1, P t-1 represents total power required at t-1 moment, p i,max represents upper limit of power between building users and power grid, p i,min represents lower limit of power between building users and power grid, p' i,max represents updated upper limit of power between building users and power grid, p' i,min represents updated lower limit of power between building users and power grid, a e , β e respectively represent user power consumption benefit coefficient, represents user comfort degree, represents corresponding cost of user, f user represents user benefit.
[0028] Further, based on the large-scale building dynamic model and the user comfort model, a single user daily load curve is extracted; considering limited data quantity and quality, and the extracted single user daily load curve, continuous monitoring of the state of the large-scale building dynamic model is realized, including:
[0029] The collected single-user single-day power consumption data is subjected to fourth-order linear or GAM regression, and then the Cook distance method is used to search the data set to obtain the corresponding Cook distance. Whether the data point is an outlier is determined by setting a threshold, and the chain equation is used to interpolate the outliers and missing values to obtain the preprocessed data set;
[0030] The preprocessed data set is subjected to load curve extraction based on the Gaussian kernel function probability density distribution method, and the load probability density of the user historical load data at time t is calculated;
[0031] Based on the load probability density, for the case of limited data quantity and quality, a Gaussian filter is used to remove random noise in the load curve, and a corresponding threshold is set for each period for filtering, so that the single-user daily power consumption curve tends to be smooth.
[0032] Further, for the collected single-user daily load data set of large-scale buildings, the obtained data is first subjected to deletion of repeated time data during data analysis, and the data value missing users are marked. At the same time, the data units are standardized by string, and the missing data of the units within the threshold of long-term response state of the temperature control load cluster are supplemented to process the related data to ensure the completeness and standardization of the data used for subsequent fourth-order linear or GAM regression;
[0033] The Cook distance is represented as:
[0034]
[0035] wherein β j represents the jth actual value, represents the fitting value of the jth β calculated using regression prediction, represents the jth fitting value in all predictions except the predicted observation value i, SSE represents the error sum of squares without including the observation value i, p represents the number of parameters in the regression model, and D i represents the Cook distance, n is the number of observation values i, and SSE(i) represents the error sum of squares including the observation value i.
[0036] The data set is subjected to min-max standardization processing:
[0037]
[0038] wherein β j,min represents the minimum value of the sample data, β j,max represents the maximum value of the sample data, and β′ j represents the characteristic value after standardization processing.
[0039] Further, the single-user daily electricity load curve is subjected to fourth-order linear or GAM regression, and the Cook distance is used to retrieve outliers from the data set, and a threshold, i.e., the decision line of the Cook distance, is set to determine whether the data point is an outlier, and then the chain equation is used to interpolate the outliers and missing values, specifically:
[0040] 1) For the kth user with missing values, d interpolation models are specified for the missing days d of the user;
[0041] 2) The predicted values within the threshold during regression prediction are extracted for interpolation at the missing place;
[0042] 3) Based on the remaining variables, the single variable is subjected to regression, first, the estimated regression coefficient and DX-covX matrix are extracted from the pre-processing regression of the fourth-order linear or GAM regression, then the conditional distribution is determined by perturbing the regression coefficient of the single missing value, and finally, according to the conditional distribution, an interpolation value is selected for each missing data;
[0043] 4) Repeat 3) several times, and the final interpolation value is a data set;
[0044] 5) Repeat 3) and 4) N times to generate N interpolation data sets, i.e., data sets without missing values.
[0045] Further, step 3-2 is specifically:
[0046] The Gaussian kernel function is expressed as:
[0047]
[0048] where: ρ t is the kernel function, n0 is the sample size, d T,J is the Euclidean distance between sample T and sample J, d0 is the truncation distance; set as the average number of neighbors of the sample, determined by MISE, as follows:
[0049]
[0050] where m is the input parameter dimension, V T,J is the covariance matrix of the data set of sample T and sample J;
[0051] The load probability density at time t of the user historical load data t is calculated and expressed as:
[0052]
[0053] where: η is the weight parameter, h is the bandwidth coefficient, near(t) is the set of K nearest neighbors of sample T obtained by the KNN algorithm, n t is the number of observation values at time t, and x j,tis the load of the jth day at time t, x t,e is the rated load at time t, d t,j is the truncated distance of the jth day at time t; meanwhile, x t,e is the threshold value;
[0054] A series of vectors is formed using the load probability density and the Gaussian kernel function, and taking f t A series of x is taken with the maximum value t,e The matrix X is formed and transposed for further processing to improve the accuracy of the description of the user's power consumption characteristics by the typical load curve;
[0055]
[0056] i d is the number of optimal mean vectors of the matrix, V m is the matrix object with an input dimension of m, V n0 is the matrix center with a sample number of n0, M is the maximum value of the input parameter dimension, d is the Euclidean distance, i d is the number of mean vectors, so the optimal i value of the selected mean vector is calculated with the minimum i d The optimal i value of the selected mean vector is calculated with the minimum number of optimal mean vectors of the matrix as the target to achieve accurate input of the number of clusters, and i d samples are selected from X as mean vectors, and the samples X id are calculated to each mean vector 1≤m≤i d After clustering, the new mean vector is calculated:
[0057]
[0058] The weighted superposition is used:
[0059]
[0060] where: v is a parameter on [0, 1]; C t is a given cluster, which is composed of samples X i η k is the weight index, d k is the stage distance of the given cluster; the probability density function at each time is calculated iteratively until convergence, and the convergence error is taken as 2%, and the output is C = {C1, C2,..., C i}, which is finally formed by the new cluster center formed by the equal-weighted average superposition to more accurately describe the load curve of the users in the cluster to ensure the accuracy of the subsequent series of vectors, is the cluster divided by the sample based on the K-Means algorithm.
[0061] Further, based on the load probability density, although the data is cleaned, due to the limited amount of data and limited data quality, the processing amount required for data processing to reach the accuracy degree cannot be met, and there are noises and biases, etc. leading to the existence of some problems in the load curve. A more robust constraint method is used to reduce the influence of data on the related load curve, so for the case of limited data quantity and quality, a Gaussian filter is used to remove random noise in the load curve, and corresponding threshold values are set for each period for filtering, so that the daily power consumption curve of a single user tends to be smooth. Specifically:
[0062] Construct a Gaussian filter:
[0063]
[0064] Where: σ is the variance, is the mean, β j is a variable, k is the total number of samples, h(β) is a one-dimensional Gaussian function, and β is a random variable.
[0065] Using the Gaussian filter, the daily power consumption curve is divided into k equidistant intervals, each interval is considered to be normally distributed, and the mean and the standard deviation σ are calculated; a number of samples are randomly selected for each interval to calculate the mean of the selected samples and the standard deviation σ' of the selected samples, and the iteration is repeated until convergence, at which time the final curve is taken as the reference value for each interval.
[0066] The application also provides a continuous monitoring device for the state of a large-scale building dynamic model, which is used to implement the above continuous monitoring method, comprising: a data acquisition module, a data transmission module, a data analysis and processing module;
[0067] The data acquisition module is used to acquire single-user daily power consumption data of a building.
[0068] The data transmission module is used to transmit the collected data to a central processing platform through a wired or wireless network (such as TCP / IP, 4G module), to ensure the real-time and integrity of the data.
[0069] The data analysis and processing module is used to clean, aggregate and process historical data and real-time data, construct related models (such as large-scale building dynamic models and user comfort models), and ensure that the single-user daily power consumption curve tends to be smooth.
[0070] The beneficial effects of the application are that the method fully considers the indoor gas and wall characteristics, and makes up for the defects of the current technology that the dynamic model of the large-scale building is greatly deviated from the actual operation; the proposed model considering user comfort degree is analyzed by reducing the dimension of the user comfort degree index, and provides an evaluation and analysis method for considering the user comfort degree in actual analysis; the proposed user comfort degree model considering cost constraints and equipment constraints cooperates with the game of a single building, introduces the working state of each equipment in the building and the scheduling cost, uses the temperature control load for power scheduling and economic adjustment, and realizes the cost optimization considering the user comfort degree. The method considers the limited data quantity and quality, processes the building load data based on the Cook distance and chain equation, improves the reliability of statistical inference, reduces the deviation caused by missing data, uses a Gaussian filter to remove random noise in the load curve, sets corresponding threshold values for each period for filtering, makes the daily electricity consumption curve of a single user smooth, improves the stability of the results, and realizes continuous monitoring of the dynamic model state of the large-scale building. BRIEF DESCRIPTION OF DRAWINGS
[0071] Figure 1 is a flowchart of the method of the application;
[0072] Figure 2 is a thermodynamic ETP model diagram of room heat exchange of the application;
[0073] Figure 3 is a flowchart of the application considering cost constraints and equipment constraints;
[0074] Figure 4 is an algorithm flowchart of the K-means clustering method of the application;
[0075] Figure 5 is a curve diagram of the temperature change of a certain day in July in summer for testing;
[0076] Figure 6 is a diagram of the dynamic change of the room temperature value of the room to which the temperature control load cluster is in working state;
[0077] Figure 7 is a diagram of the change of the running power caused by the change of the set temperature of the air conditioner in the building;
[0078] Figure 8 is a diagram of the clustering analysis result of the user data;
[0079] Figure 9 is a diagram of the comparison of the temperature control load cluster regulation ability demand response of the same building in a day.
[0080] Figure 10 is a diagram of the temperature control load regulation ability of the large-scale building in multiple scenarios of the application. DETAILED DESCRIPTION
[0081] In order to make the content of the present application more easily understood, the present application is further described in detail below according to specific embodiments and in conjunction with the accompanying drawings.
[0082] As shown in the drawings, the continuous monitoring method for the state of the large-scale building dynamic model according to the present application comprises: Figure 1 Based on the indoor gas and the wall characteristics, a large-scale building dynamic model containing an indoor thermodynamic model and a temperature control load dynamic model is established.
[0083] Considering the user comfort, cost constraints and equipment constraints, a user comfort model is constructed.
[0084] Based on the large-scale building dynamic model and the user comfort model, a single user daily load curve is extracted, and the building load data is cleaned to remove random noise in the load curve and improve the stability of the results, thereby realizing continuous monitoring of the state of the large-scale building dynamic model.
[0085] The modeling of the indoor thermodynamic model in the large-scale building dynamic model takes the indoor air specific heat capacity and the wall heat capacity as the heat capacity variable, and takes the wall thermal resistance and the indoor gas thermal resistance as the thermal resistance variable, which conforms to the laws of thermodynamics such as the law of conservation of energy and the arrow law of thermodynamics. The thermodynamic model of room heat exchange is as shown in the drawings.
[0086] Figure 2 The equation is represented as:
[0087]
[0088] In the formula, m a is the indoor air mass, H g , H l are the heat added or subtracted in the room, T1(t i ), T2(t i ), T3(t i ) are the outdoor temperature, wall temperature and indoor air temperature at t i , Q(t) represents the temperature control load refrigeration / heat, C1 represents the indoor air specific heat capacity, C2 represents the wall heat capacity, R1 represents the wall thermal resistance, R2 represents the indoor gas thermal resistance, d usually represents the differential symbol, i.e. the meaning of difference, and dt represents the difference in time. When modeling and applying mathematics, the formula (1) needs to be specially modified to convert the discrete time points into a difference equation form for application.
[0089] The modeling of the dynamic model of the temperature-controlled load takes the indoor air temperature value as the air temperature variable, takes the working state of the single temperature-controlled load as the state carrier, and takes the refrigerating capacity as the working variable; the modeling conforms to the energy conservation law and the second law of thermodynamics and other thermodynamic laws;
[0090]
[0091] In the formula, S Q (t) is the working state of the single temperature-controlled load at t, the value 1 indicates normal operation, and the value 0 indicates standby, T 3,t is the indoor air temperature value at t, T 3,set is the temperature-controlled load set value, and Ω represents the human comfort degree and can be set as a constant value 1 ℃.
[0092] p out (t) = P x S Q (t) (3)
[0093] p out (t) is the actual power output by the temperature-controlled load, and P is the rated power.
[0094] If the air conditioner in the temperature-controlled load is in the refrigeration state, the refrigerating capacity output by the air conditioner is:
[0095]
[0096] In the formula, Q(t i ) is the actual refrigerating capacity output by the temperature-controlled load at t i , η Q (t i ) is the actual energy efficiency ratio at t i , and θ and δ are fitting coefficients.
[0097] Considering the user comfort degree, cost constraints and equipment constraints, a user comfort degree model is constructed, specifically: based on the most commonly used index PMV (Predicted Mean Vote) for evaluating the thermal environment, the index is taken as the starting point of the basic equation of human body heat balance and the grade of subjective thermal sensation of psychophysiology, and the comprehensive evaluation index of human body thermal comfort is considered, the user comfort degree index is analyzed by dimension reduction:
[0098]
[0099] In the formula, n pmv represents the simplified PMV expression, n′ pmv represents the simple calculation formula after dimension reduction, n′ ave represents the overall comfort degree, M represents the metabolic rate, L represents the individual heat load of the user, and t clt represents the temperature that the user feels comfortable, which is a change interval considering the actual metabolism level of the human body to maintain the sweat amount of the human skin, and is set as a constant value, which is determined according to the geographical location and the actual season, t a t represents the room temperature, p cl t represents the clothing thermal resistance, R h t represents the relative humidity, n" pmv t represents the optimal comfort value, n' pmv,j t represents the comfort degree obtained by the jth part of data, k represents that the entire data is divided into k parts, and j represents the jth part of the data after division.
[0100] When actually analyzing the user comfort, the overall comfort n' ave is taken as the target or constraint.
[0101] In addition, because the cost constraints and equipment constraints are considered, a single building is used for cooperative game, the working state of each device in the building and the scheduling cost are introduced, the temperature control load is used for power scheduling and economic adjustment, and the cost optimization considering the user comfort is realized, and the steps are as shown in Figure 3 .
[0102]
[0103] The cost coefficient refers to a proportion of the cost before and after the use of electric energy, which is obtained by quadratic function fitting of the day-ahead data of a single user, and the specific values of the coefficients a i , b i , c i of the quadratic function are obtained by the least square method and the maximum likelihood estimation method, so as to construct the cost curve, wherein a i is the first-order coefficient, b i is the second-order coefficient, c i is the constant term, P t+1 represents the total power required at the subsequent time t+1, p i,max represents the upper limit of the power between the building user and the power grid, p i,min represents the lower limit of the power between the building user and the power grid, p' i,max represents the updated upper limit of the power between the building user and the power grid, p' i,min represents the updated lower limit of the power between the building user and the power grid, a e , β e respectively represent the user electricity benefit coefficient, represents the user comfort, represents the corresponding cost of the user, f userThe user benefit is represented. Class I represents a rigid constraint, which is a device that must meet the minimum power requirement, otherwise it will cause building function paralysis or safety hazards; class II represents an upper limit constraint, which is a device whose power cannot exceed the dynamic upper limit, and exceeding the limit will significantly increase the power cost; class III represents a flexible constraint, which is a device whose power can be freely adjusted within the comfort range, and has the highest scheduling flexibility. The general steps are as follows: first, input the initial parameters: the power at the last time, the cost coefficient and the power threshold to provide data basis for subsequent optimization and ensure the continuity of scheduling; then calculate the comfort interval according to the above parameters and adjust the upper and lower limits of power; then aggregate all single device data, solve the optimal solution under cooperative game, minimize the total cost while maximizing the benefit, and ensure that the power is within the updated threshold range; if the power is equal to the updated minimum demand power at this time, it is class I, which needs to be forced to lower the power, adjust the power p, which may trigger the standby device; if the power is not equal to the updated minimum demand power at this time but meets the updated dynamic upper limit, it is class II, which needs to limit the power peak, adjust the power p (the real-time power value of the building interacting with the power grid), and avoid overload; if the power is neither equal to the updated minimum demand power nor equal to the updated dynamic upper limit at this time, it is class III, which is in the middle value, keeps the current scheduling, and does not need to be intervened; after classification judgment and power update, verify whether the total power meets P t+1 (Future demand), if not (N), the threshold or optimization strategy needs to be adjusted again; if it meets (Y), the process is terminated.
[0104] Based on the large-scale building dynamic model and user comfort model, the single user daily load curve is extracted to realize continuous monitoring of the state of the large-scale building dynamic model. If random noise appears in the extracted single user daily load curve data based on the large-scale building dynamic model and user comfort model, and the data is not processed, it may lead to incorrect conclusions in subsequent load prediction and analysis, resulting in temperature control adjustment errors and processing method errors. Data cleaning is performed on the building load data to extract typical user features, which can effectively deal with random noise. This method is divided into three parts:
[0105] 1) Data cleaning, this embodiment contains 216 households of electricity consumption in a city in 2015, the collection frequency is 15 min, and the complete and consistent of 646981 time series of electricity consumption data set are checked. In data analysis, delete repeated time data, and mark the data value missing user; At the same time, through the string, the unit of each data is standardized, and the unit missing data within the threshold value of long-term response state of temperature control load cluster is supplemented to process the related data to ensure the completeness and specification of the data used in the subsequent four-order linear or GAM regression; After four-order linear or GAM regression of single user daily electricity data, Cook distance is used to retrieve the data set to obtain outliers, and threshold value, i.e. decision line of Cook distance, is used to judge whether the data point is outlier, and chain equation is used to interpolate outliers and missing values.
[0106] Cook distance D i Can be expressed as:
[0107]
[0108] Where: β j The jth actual value, The jth fitting value calculated by using regression prediction, The jth fitting value of all predictions except the predicted observation value i, SSE represents the sum of squares of errors in the case of not containing the observation value i, p represents the number of parameters in the regression model, D i Indicates the Cook distance, n is the number of observation value i, SSE(i) represents the sum of squares of errors in the case of containing the observation value.
[0109] When applying Cook distance, a threshold value is usually set to judge whether the observation point is a strong influence point or an outlier. This threshold value is usually set to 4 / n, where n is the number of observation points. If the Cook distance of a certain observation point exceeds this threshold value, it is considered that the point is a strong influence point or an outlier. In this embodiment, 4 times of Cook distance is set as the decision line for convenience and other reasons, and the points exceeding the decision line are determined as outliers and deleted.
[0110] The Cook processed data set is subjected to min-max standardization processing:
[0111]
[0112] Where: β j The jth actual value, β j,min Indicates the minimum value of sample data, β j,max Indicates the maximum value of sample data, β' j Indicates the characteristic value after standardization.
[0113] At this time, the missing values are supplemented by using the multiple imputation principle of chain equation, and the steps are as follows:
[0114] a) For the kth user with missing values, specify d imputation models for the missing days d of the user;
[0115] b) Impute the predicted values within the threshold at the time of regression prediction at the missing place;
[0116] c) Regress a single variable based on the remaining variables; first extract the estimated regression coefficient and DX-covX matrix from the pre-processing regression of the fourth-order linear or GAM regression, then determine the conditional distribution of a single missing value by perturbing the regression coefficient, and finally select an imputation value for each missing data according to the conditional distribution;
[0117] d) Repeat c), in order to improve the accuracy and take into account the amount of calculation, the number of multiple imputations required is greater than 10 times, and the final imputation value is taken as a data set;
[0118] e) Repeat c) and d), and set N times to generate N imputed data sets, i.e. data sets without missing values, usually 3-10.
[0119] 2) Kernel density clustering, specifically as follows:
[0120] The method based on Gaussian kernel function probability density distribution is used for load curve extraction, wherein the Gaussian kernel function can be expressed as:
[0121]
[0122] wherein, ρ t is the kernel function, n0 is the sample size, d T,J is the Euclidean distance between sample T and sample J, d0 is the truncation distance, which can be set as the average number of neighbors, and MISE is used here, as shown below:
[0123]
[0124] wherein: m is the input parameter dimension, V T,J is the data set covariance matrix of sample T and sample J.
[0125] Therefore, when calculating the load probability density of the user historical load data t, it can be expressed as:
[0126]
[0127] where: η is the weight parameter, h is the bandwidth coefficient, near(t) is the set of K nearest neighbors of sample T using the K-Nearest Neighbor (KNN) algorithm, n t is the number of observations at time t, x j,t is the load at time t of the jth day, x t,e is the rated load at time t, d t,j is the cut-off distance at time t of the jth day. At the same time, the threshold value is determined according to the data set x t,e . The series vector is formed using the load probability density and the Gaussian kernel function, and the maximum value is taken as f t . The series of x t,e is formed, and after transposition, further processing is performed to improve the accuracy of the description of the user electricity consumption characteristics by the typical load curve.
[0128]
[0129] i d is the number of optimal mean vectors of the matrix, m is the input parameter dimension, n0 is the number of samples, V m is the matrix object with the input dimension m, is the matrix center with the number of samples n0, M is the maximum value of the input parameter dimension, d is the Euclidean distance, i d is the number of mean vectors, so the optimal i value of the selected mean vector is calculated with the minimum i d , that is, the number of optimal mean vectors of the matrix is minimized, to realize the accurate input of the number of clusters. i d samples are selected from X as the mean vectors, and the distances between the samples and each mean vector are calculated, and clustering is performed, and the new mean vector is calculated:
[0130]
[0131] Here, the weighted superposition is used:
[0132]
[0133] where: v is a parameter on [0, 1], C t is the given cluster, which is composed of samples X i , η k is the weight index, d k is the stage distance of the given cluster. The probability density function at each time is iteratively calculated until convergence, and the convergence error is taken as 2%. The output is C = {C1, C2,..., C iwhich is ultimately formed by equal-weighted average value superposition to form a new cluster center to ensure the accuracy of subsequent series vector formation, is the cluster to which the sample is divided based on the K-Means algorithm.
[0134] For the problem of limited data quantity and quality, a Gaussian filter is used to remove random noise in the load curve, and a corresponding threshold is set for each period for filtering, so that the single-user daily power consumption curve tends to be smooth.
[0135] The limited quantity of data refers to the limited amount of available data, which cannot meet the processing amount required for data processing to achieve accuracy; the limited quality of data refers to the limited quality of available data, such as noise and bias, which cannot meet the quality requirements required for data processing. This may be due to limitations of the data source, limitations of the data storage space, defects in the accuracy and integrity of the data source, or loss or tampering of the data during transmission and storage.
[0136] Therefore, after data preprocessing, a more robust constraint method needs to be used to reduce the impact of data quality and quantity.
[0137]
[0138] where σ is the variance, is the mean, β j is the variable, k is the total number of samples, h(β) is a one-dimensional Gaussian function, and β is a random variable.
[0139] Using the above formula (14) Gaussian filter to divide the daily power consumption curve into k equidistant intervals, regarding each interval as a normal distribution and calculating the mean and the standard deviation σ, randomly selecting 20 samples from each interval to calculate the mean of the selected samples and the standard deviation σ', setting an error of 5%, and repeating the iteration until convergence, at which time the final curve is taken as the reference value for each interval.
[0140] The simulation verification model designed by the invention is as follows:
[0141] To verify the feasibility and effectiveness of the multi-constraint factor dynamic modeling and data processing method for large-scale building proposed by the invention, the demand side large-scale building, i.e. the temperature control load cluster, is aggregated and modeled by modeling the room and air conditioner respectively, and the Monte Carlo simulation data method and the K-means clustering method are used to preprocess the building load response data, and then the large-scale building is analyzed, wherein the algorithm flowchart of the K-means clustering method is as shown in Figure 4 .
[0142] The state estimation performance comparison and analysis are as follows:
[0143] The actual data of a city is selected for the above test, and the temperature change curve of a certain day in July in summer is shown in Figure 5 The room area A is assumed to follow a normal distribution A ~ N(40, 20 2 ), that is, in the sample set, the average area of N rooms is 40 m2, and the sample standard deviation is 20 m2. It is assumed that each room has one air conditioner, so the initial set temperature of the N air conditioners is The value is randomly distributed in the interval 23-26℃. In addition, considering the corresponding relationship between the room area and the rated power of the air conditioner, the ratio of the rated power of the temperature control air conditioner to the corresponding room area is taken as 50 times, that is, a 2000W air conditioner is installed in a 40㎡ room.
[0144] The dynamic change of the room temperature value of the temperature control load cluster under the working state is shown in Figure 6 It is assumed that the number of air conditioners changes as follows: N = 100, 400, 800; and the temperature adjustment amount ΔT set is 1℃; and the control air conditioner load cluster participates in the response at 12 o'clock. Figure 6 (a), (b), (c) of respectively represent the dynamic change of the room temperature value when N is 100, 400, and 800, respectively, wherein different colors represent the individual temperature value change of each room. As can be seen from the figure, Figure 6 the line of (a) is sparse, the individual difference is obvious, and the cluster temperature fluctuation is large; Figure 6 the line of (b) is dense, the individual difference is reduced, and the consistency is enhanced; Figure 6 the line of (c) is highly dense, the cluster response tends to be synchronous, the temperature fluctuation is the most stable, and it reflects that the larger the scale, the more "uniform" the temperature migration to the new interval, the cluster regulation capacity improves with the expansion of the scale, and the target temperature control can be more accurately realized, thereby reflecting that the scale effect improves the overall response consistency by offsetting the individual randomness.
[0145] According to the simulation, Figure 6 it is shown that after the set temperature of the air conditioner aggregation group at 12 o'clock is adjusted, the hysteresis deviation fluctuation range of the room temperature value will migrate to the new temperature interval. Since the initial temperature of the N air conditioners in the cluster is randomly distributed in 23-26℃, the adjusted value is randomly distributed in 24-27℃. The change of the running power caused by the change of the set temperature of the air conditioner in the building is shown in Figure 7
[0146] Table 1 can be made from Figure 7 It can be seen from Table 1 that under different response scales, the regulation capacity RC A and the ramp rate RR A increase rapidly in a certain multiple with the increase of the number of temperature control loads in terms of value, and the response time RTA Duration DT A The numerical change is not obvious. This is quite different from the previous example of temperature adjustment capacity under different temperature adjustment amounts, in which the response time RT A is shortened rapidly with the increase of temperature adjustment amount, while the duration DT A is extended to a certain extent. This is because under different temperature adjustment amounts, the temperature adjustment value change can make the load cluster enter the standby state to a certain extent, and maintain the standby mode for a longer period of time. In this example, the increase of the size of the temperature control load has little effect on the response time RT A and the duration DT A , but because the size of the temperature control load cluster is approximately proportional to the numerical value of the regulation capacity, the increase of the response size of the temperature control load can proportionally increase the regulation capacity RC A .In addition, since the ramp rate is the ratio of the regulation capacity RC A to the response time RT A , in this simulation, with the increase of the size of the temperature control load, the equivalent ramp rate RR A of the temperature control load cluster response also increases linearly.
[0147] Table 1 Regulation capacity of cluster under different response sizes
[0148]
[0149] At the same time, the data of the demonstration project in the relevant city is taken to verify the effectiveness of the temperature control load cluster participating in demand response. The data set contains 215 user electricity data, including 15-minute level electricity collection from July 1 to August 1, 2022. The K-means method is used for clustering analysis of user data, and the clustering is divided into four clusters with similar electricity behaviors, i.e. four large-scale ideal buildings, as shown in Figure 8 .
[0150] Two weeks of data are taken for example analysis. Here it is assumed that in the first period, the users of each building normally use electricity, and the average power of a single user for seven days is taken as the control curve; in the second period, the command is sent at 15:00 every day, and the end command is sent at 16:00. The comparison of the demand response of the temperature control load cluster in the same building within a day is shown in Figure 9 . The regulation capacity of the temperature control load of the large-scale building in multiple scenarios is shown in Figure 10 .
[0151] Take scenario one as an example for analysis, process the electricity data of the first type of user in multiple scenarios, as shown in Figure 9shown. It is shown that in the second phase, due to the implementation of demand response, the total load power consumption of the temperature-controlled load cluster starts to decrease at 15:00 and reaches the lowest value of the period at about 15:41. From the simulation graph, it can be calculated that the regulation capacity RC A 492.21 kW, the response time RT A 41 min, the equivalent ramp rate RR A 40.61 kW / min, the duration DT A 20 min. After the end of demand response, the load power starts to increase at 16:00 and reaches the normal level at 16:11. At the same time Figure 10 (a) of FIG. 13 shows the regulation capacity dynamic change of a small-scale temperature-controlled load cluster (e.g., N = 100 air conditioners), Figure 10 (b) of FIG. 13 shows the regulation capacity dynamic change of a medium-scale temperature-controlled load cluster (e.g., N = 400 air conditioners), Figure 10 (c) of FIG. 13 shows the regulation capacity dynamic change of a large-scale temperature-controlled load cluster (e.g., N = 800 air conditioners), and Figure 10 (d) of FIG. 13 shows the comprehensive comparison of the cluster regulation capacity under multiple scenarios (e.g., different temperature adjustment amounts or user types). Among them Figure 10 (a) of FIG. 13 reflects that the regulation capacity of a small-scale cluster is limited, and the overall response consistency is low due to individual randomness; Figure 10 (b) of FIG. 13 shows that in a medium-scale, the cluster regulation capacity is enhanced, the temperature migration is more synchronized, and the scale effect is preliminarily manifested; Figure 10 (c) of FIG. 13 verifies that a large-scale cluster can significantly improve the regulation capacity and the ramp rate, and the individual randomness is offset, and the overall response is more accurate; and Figure 10 (d) of FIG. 13 reflects that a large-scale building can balance user comfort and cost through cooperative game and dynamic modeling, and realize efficient power dispatching.
[0152] It can be seen that the method provides an evaluation and analysis method for considering user comfort during actual analysis by performing dimensionality reduction analysis on the user comfort index. The temperature-controlled load cluster can guarantee the regulation needs during participation in demand response, effectively improve the regulation capacity of the power system, and realize cost optimization considering user comfort. Meanwhile, by extracting the typical user features at the beginning and extracting and cleaning the features of the data, the corresponding threshold values are set for filtering in each period, the daily power consumption curve of a single user tends to be smooth, the stability of the results is improved, and continuous monitoring of the dynamic model state of the large-scale building is realized.
[0153] The embodiment also provides a continuous monitoring device for the dynamic model state of the large-scale building, which is used to realize the continuous monitoring method and includes a data acquisition module, a data transmission module, and a data analysis and processing module.
[0154] The data collection module is configured to collect single-user daily power consumption data of the building.
[0155] The data transmission module is configured to transmit the collected data to a central processing platform through a wired or wireless network, so as to ensure real-time and integrity of the data.
[0156] The data analysis and processing module is configured to clean, aggregate and process historical data and real-time data, construct a related model and ensure that the single-user daily power consumption curve tends to be smooth.
[0157] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, such as object-oriented programming language Java and interpreted scripting language JavaScript.
[0158] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one flow or multiple flows and / or blocks
[0159] These computer program instructions can also be stored in a computer-readable memory that can guide the computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one flow or multiple flows and / or blocks
[0160] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The flowchart blocks Figure 1 The flowchart blocks
[0161] Although preferred embodiments of the application have been described herein, it will be apparent to those skilled in the art that various modifications can be made within the scope of the application without departing from the spirit of the application. Accordingly, it is intended that all such possible modifications be included within the scope of the application as described in the following claims. In compliance with the statute, the application has been described in language more or less specific to structural
[0162] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for continuously monitoring the state of a large-scale building dynamic model, characterized in that: include: Based on the characteristics of indoor gas and walls, a large-scale building dynamic model is established, including an indoor thermodynamic model and a temperature control load dynamic model; Considering user comfort, cost constraints and equipment constraints, a user comfort model is constructed; Based on the large-scale building dynamic model and user comfort model, a single-user daily load curve is extracted. Considering the limited data quantity and quality, as well as the data of the extracted single-user daily load curve, continuous monitoring of the status of the large-scale building dynamic model is achieved.
2. A method for continuously monitoring the state of a large-scale building dynamic model according to claim 1, characterized in that: Based on the characteristics of indoor gas and wall, an indoor thermodynamic model is established, including: modeling the indoor thermodynamic model with indoor airflow specific heat capacity and wall heat capacity as heat capacity variables, and with wall thermal resistance and indoor gas thermal resistance as thermal resistance variables; the differential equation of the indoor thermodynamic model is expressed as: Among them, m a is the indoor air quality, H g 、H l are the heat gain and loss in the room, T1(t i )、T2(t i )、T3(t i ) is t i where t is the outside temperature, wall temperature and indoor air temperature, Q(t) represents the temperature control load (cooling / heating), C1 represents the indoor air specific heat capacity, C2 represents the wall heat capacity, R1 represents the wall thermal resistance, and R2 represents the indoor gas thermal resistance; d represents the differential sign, i.e., difference; dt represents the difference with respect to time.
3. A method for continuously monitoring the state of a large-scale building dynamic model according to claim 2, characterized in that: Based on the characteristics of indoor gas and wall, a dynamic model of temperature control load is established, including: the modeling of the dynamic model of temperature control load uses the indoor temperature value as the temperature variable, the working state of the single temperature control load as the state carrier, and the cooling capacity as the working variable: Where: S Q (t) is the working state of the single temperature control load at time t, with a value of 1 indicating normal operation and a value of 0 indicating standby; T 3,t is the indoor temperature at time t, T 3,set is the temperature control load setting value, Ω represents human comfort; p out (t)=P×S Q (t) Among them, p out (t) is the actual power output of the temperature control load, and P is the rated power; If the air conditioner is in cooling state in the temperature control load, its output cooling capacity is: Where: Q(t i ) is t i The actual cooling capacity output by the temperature control load at the moment, η Q (t i ) is t i The actual energy efficiency ratio at the moment, θ and δ are fitting coefficients.
4. A method for continuously monitoring the state of a large-scale building dynamic model according to claim 3, characterized in that: Considering user comfort, cost constraints, and equipment constraints, a user comfort model is constructed. Considering user comfort, the PMV indicator for evaluating the thermal environment is simplified and modified to perform a dimensionality reduction analysis of the user comfort index: Where: n pmv represents the simplified PMV expression, n′ pmv Represents a simple calculation formula after dimensionality reduction, n′ ave represents the overall comfort, M represents the metabolic rate, L represents the user's individual heat load, t cl Indicates the temperature at which the user feels comfortable, t a represents room temperature, ρ cl Represents clothing thermal resistance, R h % indicates relative humidity, n" pmv represents the optimal comfort value; k represents that the entire data is divided into k parts, j represents the jth part of the divided data; n′ pmv,j Represents the comfort level obtained from the jth part of the data.
5. A method for continuously monitoring the state of a large-scale building dynamic model according to claim 4, characterized in that: Considering user comfort, cost constraints, and equipment constraints, a user comfort model is constructed. Taking cost constraints and equipment constraints into account, a cooperative game is adopted for a single building. Based on the collected working status and dispatching costs of each device in the building, power dispatch and economic adjustments are performed using temperature control loads to achieve cost optimization while considering user comfort: Among them, a i is the coefficient of the first-order term, b i is the coefficient of the quadratic term, c i is a constant term, P t+1 represents the total power required at the subsequent time t+1, P t-1 represents the total power required at time t-1, p i,max Indicates the upper limit of power between building users and the grid, p i,min Represents the lower limit of power between building users and the grid, p′ i,max represents the upper limit of the power between the building user and the grid after the update, p′ i,min represents the updated lower limit of the power between the building user and the grid, α e , β e They represent the user's electricity efficiency coefficient, Indicates user comfort, represents the corresponding cost of the user, f user Indicates user benefits.
6. The method for continuously monitoring the state of a large-scale building dynamic model according to claim 1, characterized in that: Based on the large-scale building dynamic model and the user comfort model, a single-user daily load curve is extracted. Considering the limited data quantity and quality, as well as the data of the extracted single-user daily load curve, continuous monitoring of the state of the large-scale building dynamic model is achieved, including: Perform a fourth-order linear or GAM regression on the collected single-user daily electricity consumption data. Then, use the Cook distance method to retrieve the data set to obtain the corresponding Cook distance. Set a threshold to determine whether the data point is an outlier. Use the chain equation to interpolate outliers and missing values to obtain the preprocessed data set. The load curve is extracted from the preprocessed data set using a method based on the probability density distribution of the Gaussian kernel function, and the load probability density of the user's historical load data at time t is calculated; Based on the load probability density and considering the limited data quantity and quality, a Gaussian filter is used to remove random noise in the load curve, and corresponding thresholds are set for each time period for filtering, so that the daily electricity consumption curve of a single user tends to be smooth.
7. The method for continuously monitoring the state of a large-scale building dynamic model according to claim 6, characterized in that: During data analysis, the collected single-user daily load dataset for large-scale buildings was first deduplicated and users with missing data values were marked. At the same time, each data unit was normalized using character strings, and missing data for units within the threshold for maintaining a long-term response state in the temperature-controlled load cluster was supplemented to ensure the integrity and standardization of the data used in subsequent fourth-order linear or GAM regression. Cook distance is expressed as: Among them, β j represents the jth actual value, represents the fitted value of the jth β calculated using regression forecast, represents the jth fitted value in all predictions except the predicted observation i, SSE represents the sum of squared errors when the observation i is not included, p represents the number of parameters in the regression model, D i It is expressed as Cook's distance, where n is the number of observations i, and SSE(i) represents the sum of squared errors when observation i is included. Perform min-max normalization on the dataset: Among them, β j,min represents the minimum value of sample data, β j,max Indicates the maximum value of the sample data, β′ j Represents the eigenvalue after normalization.
8. The method for continuously monitoring the state of a large-scale building dynamic model according to claim 7, characterized in that: After performing a fourth-order linear or GAM regression on the daily electricity load curve of a single user, the Cook distance is used to retrieve the data set to obtain outliers. A threshold, i.e., the decision line of the Cook distance, is set to determine whether a data point is an outlier. The chain equation is then used to interpolate outliers and missing values, specifically: 1) For the kth user with missing values, specify d imputation models for the number of missing days d; 2) For missing values, first extract the predicted values within the threshold when regressing the predictions and interpolate them; 3) Regress a single variable based on the remaining variables, first extract the estimated regression coefficients and DX-covX matrix from the preprocessed regression of the fourth-order linear or GAM regression, then use the perturbed regression coefficients to determine the conditional distribution of the single missing value, and finally select the imputed value for each missing data according to the conditional distribution; 4) Repeat 3) and perform multiple imputation several times, and the final imputed values are used as a data set; 5) Repeat 3) and 4) N times to generate N imputed data sets, i.e., data sets without missing values.
9. The method for continuously monitoring the state of a large-scale building dynamic model according to claim 7, characterized in that: The load curve is extracted from the preprocessed data set using a method based on the probability density distribution of the Gaussian kernel function; Among them, the Gaussian kernel function is expressed as: Where: t is the kernel function, n0 is the number of samples, d T,J is the Euclidean distance between sample T and sample J, d0 is the cutoff distance; it is set to the average number of neighbors of the sample, determined by MISE, as shown below: Among them: m is the input parameter dimension, V T,J is the data set covariance matrix of sample T and sample J; Calculate the load probability density of the user's historical load data at time t, expressed as: Where: η is the weight parameter, h is the bandwidth coefficient, near(t) is the set of K nearest neighbor points obtained by the KNN algorithm for sample T, and n t is the number of observations at time t, x j,t is the time series load on day j, x t,e is the rated sequence load at time t, d t,j is the time series cutoff distance of day j; at the same time, x is determined according to the data set. t,e Threshold; The load probability density and Gaussian kernel function are used to form a series of vectors, and f is set t Take the maximum value of a series of x t,e The matrix X is formed and transposed for further processing to improve the accuracy of describing the user's electricity consumption characteristics based on the typical load curve; i d is the optimal mean vector number of the matrix, V m is a matrix object with input dimension m, is the center of the matrix with sample number n0, M is the maximum value of the input parameter dimension, d is the Euclidean distance, i d is the number of mean vectors, with min i d That is, the minimum number of matrix optimal mean vectors is the optimal i value of the target calculation to achieve accurate input of the number of clusters, and i is selected from X d samples as the mean vector, calculate the sample To each mean vector 1≤m≤i d After the distance is calculated, clustering is performed and the new mean vector is calculated: Using weighted overlay: Where: v is the upper parameter [0,1]; C t For a given cluster, it is composed of sample X i constituted by; η k is the weight index, d k is the stage distance of a given cluster; the probability density function is iteratively calculated at each moment until convergence, with a convergence error of 2%, and the output is The final method is to use the new cluster center formed by superposition of equal weight average values to form the data division to more accurately describe the load curve of users in the cluster to ensure the accuracy of the subsequent series vector formation. The clusters into which the samples are divided are based on the K-Means algorithm.
10. A method for continuously monitoring the state of a large-scale building dynamic model according to claim 9, characterized in that: Based on the load probability density, and considering the limited data quality, a Gaussian filter is used to remove random noise from the load curve. A corresponding threshold is set for each time period to perform filtering, so that the daily electricity consumption curve of a single user tends to be smooth. Specifically: Construct a Gaussian filter: Where: σ is the variance, is the mean, β j is a variable, k is the total number of samples, h(β) is a one-dimensional Gaussian function, and β is a random variable; The Gaussian filter is used to divide the daily electricity consumption curve into k equally spaced intervals, each interval is considered as a normal distribution and the mean is calculated. and standard deviation σ; randomly select several samples from each interval and calculate the mean of the selected samples The standard deviation σ′ of the selected samples is repeated until convergence, at which point the final curve is taken as the reference value for each interval.
11. A device for continuously monitoring the state of a large-scale building dynamic model, characterized in that: include: Data acquisition module, data transmission module, data analysis and processing module; The data collection module is used to collect daily electricity consumption data of a single user in a building; The data transmission module is used to transmit the collected data to the central processing platform via a wired or wireless network to ensure the real-time and integrity of the data; The data analysis and processing module is used to clean, aggregate and process historical data and real-time data, build relevant models and ensure that the daily electricity consumption curve of a single user tends to be smooth.