Indoor thermal environment regulation and control technology based on personnel multi-mode thermal perception state
By integrating dynamic thermal sensory prediction model and reinforcement learning algorithm, combined with body domain network sensing equipment to obtain physiological characteristic parameters, solving the problems of precise and personalized control in building thermal environment regulation, and achieving rapid and accurate indoor temperature regulation and energy conservation.
Patent Information
- Application Number
- CN202410002734.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to achieve accurate and personalized indoor temperature control in building thermal environment regulation, cannot meet the dynamic thermal needs of personnel, and has a high energy consumption.
By combining the integrated dynamic thermal sensory prediction model, the behavioral intention probability model of personnel air conditioning regulation and the air conditioning system regulation method, the physical domain network sensing equipment is used to obtain physiological characteristic parameters, and combined with the reinforcement learning algorithm Q-learning, the optimal control of indoor temperature is achieved.
Personalized indoor temperature regulation is achieved, the accuracy of thermal perception prediction and model training efficiency is improved, fast and accurate indoor environment regulation is achieved, and energy consumption is reduced.
Smart Images

Figure BDA0004644848650000011 
Figure BDA0004644848650000012 
Figure BDA0004644848650000021
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent buildings, and relates to a technical method for regulating indoor environmental temperature, specifically to an indoor thermal environment regulation technology based on the multi-modal thermal perception state of personnel. Background Art
[0002] In actual life scenarios, people in buildings are always exposed to various different thermal environment conditions. The indoor thermal environment not only affects the thermal comfort and work efficiency of indoor people, but also has a significant impact on building energy consumption. Therefore, in the situation where the building energy consumption is increasing year by year and the energy reserve is becoming increasingly tense, the research direction of achieving an indoor thermal environment that not only meets the thermal comfort requirements of people but also maximally saves energy has become the focus of the industry.
[0003] However, the current technical means and solutions for the operation management and precise environmental control of building thermal environments are not sufficient. This is mainly reflected in the research on indoor thermal environment intelligent regulation technologies. Most existing research mainly constructs a thermal sensation or thermal comfort model of people based on subjective feelings and relatively single and traditional physiological parameters, while there are few reports on the research of introducing real-time monitoring information of people's thermal sensation states into the regulation process of building air conditioning systems. At the same time, people's perception of the thermal environment has thermal adaptability and thermal tolerance, and people's behavioral actions to regulate environmental equipment are also affected and controlled by "behavioral motivation". Therefore, without considering the behavioral thermal demand of people's thermal regulation behavior intention or motivation, it will be difficult to meet people's dynamic thermal demand and impossible to achieve intelligent on-demand management and precise regulation control of the indoor thermal environment.
[0004] In view of this, the present invention provides an indoor thermal environment regulation technology based on the multi-modal thermal perception state of personnel, which combines an integrated dynamic thermal sensation prediction model, a probability model of personnel air conditioning regulation behavior intention, and an air conditioning system regulation method to achieve the optimal control of indoor temperature, thereby achieving fast and precise indoor environment regulation performance. Summary of the Invention
[0005] The present invention proposes an indoor thermal environment regulation technology based on the multi-modal thermal perception state of personnel, including the following steps:
[0006] Step 1: Use a personnel body area network sensing device to obtain the physiological characteristic parameters of personnel in a thermal environment (physiological heat exchange system, cardiovascular system, brain nervous system). Obtain the thermal sensation votes of personnel in the form of questionnaire surveys.
[0007] Step 2: Take each physiological characteristic parameter and the thermal sensation vote (Thermal Sensation Vote, TSV) of personnel as inputs to establish a thermal sensation state model of different physiological systems of personnel under specific thermal environment conditions. The specific model is as follows:
[0008] PTS1 = a × SET + b (1)
[0009]
[0010]
[0011]
[0012] PTS4 = Y[XGBoost(e n (x i ))] (5)
[0013] Equation (1) represents a thermal sensation prediction model based on a physiological heat exchange system. In the equation, a is the gradient of the fitting line, b is the intercept, and SET represents the standard effective temperature.
[0014] Equation (2) represents a prediction model based on the relationship between the heart rate change rate and thermal sensation. In the equation, A is the limit coefficient of thermal sensation, and its value is the upper limit value of thermal sensation, i.e., 3, and the lower limit value of thermal sensation, i.e., -3; B is the slope coefficient. HR is the heart rate parameter response (unit: bpm), and HR n is the average heart rate baseline value when neutral (unit: bpm).
[0015] Equations (3) and (4) represent prediction models based on the relationship between heart rate variability and thermal sensation. Among them, Equation (3) is when the indoor temperature is greater than 26 °C, and Equation (4) is when the indoor temperature is less than or equal to 26 °C. m, n, and q are fitting coefficients.
[0016] Equation (5) represents a thermal sensation prediction model based on electroencephalogram signal characteristics. In the equation, Y represents the output dataset of the predicted value, and e n (x i ) where x i represents 4 eigenvalue extracted from the electroencephalogram signal. The superscript n represents the number of sample groups, and XGBoost(e n (x i )) represents the calculation of the sample parameter x i using the XGBoost-based model.
[0017] Step 3: Based on Step 2, combine and predict each physiological system with the thermal sensation state model to construct an integrated dynamic thermal sensation prediction model As shown in Equation (6).
[0018]
[0019] In the equation, α i = [α1, α2, α3, α4] T , represents the weight coefficient; PTS i= [PTS1, PTS2, PTS3, PTS4] T , represented as the above-mentioned step-by-step thermal sensation prediction models. The least squares method is used to assign the weight coefficient α to the integrated model i , to reflect the influence degree of each modal physiological parameter on the thermal sensation state of personnel; then the particle filter method is used to update the static estimation coefficient to eliminate the estimation error of the static method.
[0020] Step 4: Based on Step 3, through the integrated model to predict the thermal sensation at the next moment Finally, the predicted value of the integrated thermal sensation state in the time series is obtained as
[0021] Step 5: Construct the probability distribution of the personnel's air-conditioning adjustment behavior intention, initialize the Q-table and embed the probability distribution of the personnel's air-conditioning adjustment behavior intention, set parameters such as the learning rate, reward discount factor, exploration probability, and maximum number of iterations of the reinforcement learning model, and establish a reinforcement learning model considering the probability of the personnel's regulation behavior intention. The probability model of the personnel's air-conditioning behavior adjustment intention generated under the stimulation of the thermal environment conditions is as shown in the formula:
[0022]
[0023] In the formula, μ is the position coefficient; β is the scale coefficient, and S is the thermal sensation vote
[0024] Step 6: Based on Step 5, by monitoring the indoor temperature and the change of the thermal sensation state of personnel at this moment, calculate the reward obtained by the model and update the Q-table to train the reinforcement learning model considering the probability of the personnel's regulation behavior intention. The calculation formula is as follows:
[0025] q(s t , a t ) ← q(s t , a t ) + α[r t+1 + γmax(s t+1 , a t+1 ) - q(s t , a t )] (8)
[0026] In the formula, q(s, a) represents the goodness or badness of selecting the action a in the current state s, the subscript t represents the t moment, r represents the reward obtained after executing the action a in the state s, γ is the discount factor, indicating the importance of future rewards to the current reward, and γ ∈ [0, 1].
[0027] Step 7: Based on Step 6, continuously update the Q-table until the number of training times reaches the set maximum number of iterations, indicating that the model learning process has been completed. Output the optimal strategy for the indoor temperature set value, forming an indoor temperature control method based on Q-learning reinforcement learning, so as to achieve the best control of the indoor temperature.
[0028] Beneficial effects
[0029] (1) By integrating multi-modal perception signals of personnel under thermal environment conditions, the present invention constructs a personnel integrated dynamic thermal sensation prediction model, further improving the accuracy of personnel thermal sensation prediction.
[0030] (2) The present invention establishes a general model of the behavior of the probability of personnel's air-conditioning adjustment intention, obtains the personnel thermal sensation interval band, and realizes personalized indoor temperature regulation based on individuals.
[0031] (3) By embedding the initialization of the probability of personnel's air-conditioning adjustment behavior intention in the Q-learning algorithm of reinforcement learning, the present invention improves the efficiency of model training. Description of the drawings
[0032] Figure 1 is the flowchart of the method of the embodiment of the present invention;
[0033] Figure 2 Subjective questionnaire rating scale;
[0034] Figure 3 Comparison of the prediction performance of the step-by-step basic model, the integrated static model and the integrated dynamic model;
[0035] Figure 4 General model of the probability of air-conditioning adjustment behavior intention of personnel in different thermal sensation states;
[0036] Figure 5 Iteration steps of the Q-Learning model embedded with the probability model of personnel adjustment behavior intention and the non-embedded model. Detailed implementation manners
[0037] To make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates the present invention in detail with reference to specific embodiments and the accompanying drawings, including the following steps:
[0038] Step 1: In this case, the temperature is adjusted and changed in a drifting manner in the order of 30°C - 28°C - 26°C - 24°C - 22°C of the temperature set point to explore the influence of indoor air temperature on personnel's thermal perception. The specific experimental conditions are shown in the following table:
[0039] Table 1 Experimental environmental conditions
[0040]
[0041] The experimental duration for different experimental conditions was 40 minutes, and the intermittent stage for changing the indoor environmental parameter conditions was 10 minutes. At the beginning of each experimental condition, the subjects were prompted to fill out a subjective thermal sensation state questionnaire, and then they remained sitting still and entered the reading state. A questionnaire was filled out every 10 minutes. The specific subjective questionnaire scoring scale is as Figure 2 shown. After the physiological monitoring equipment of the subjects was installed and they got adapted to the environment, physiological parameter data such as electroencephalogram signals, electrocardiogram signals, blood pressure, and blood oxygen of the subjects were synchronously and real-time recorded.
[0042] Step 2: Obtain the SET parameters, heart rate change rate, LF / HF, and electroencephalogram signals under different thermal environment conditions through the experiments in Step 1. The specific expressions after fitting the original data of the SET parameters, heart rate change rate, and LF / HF are as follows:
[0043] PST group = 0.440 × SET - 11.293 (1)
[0044]
[0045] LF / HF = 0.079 × PTS group 2 + 0.043 × PTS group + 0.785 (3)
[0046] To obtain the electroencephalogram signals under different thermal environments through experiments, it is first necessary to perform preprocessing operations and feature extraction on the electroencephalogram data. The preprocessing process of the electroencephalogram signals is carried out in the EEGLAB toolbox in MATLAB. The specific process includes electroencephalogram electrode channel localization, removing useless electrodes, rereferencing, filtering, signal segmentation, independent component analysis, removing artifact components such as electrooculogram and electromyogram, so as to obtain the preprocessed multi-channel EEG data set. Then, based on the extracted feature data set, the XGBoost algorithm is used to build a thermal sensation prediction model. Based on the XGBoost model, a thermal sensation prediction model based on electroencephalogram signal features can be constructed, as shown in the following formula:
[0047] PTS4 = Y[XGBoost(e n (x i ))] (4)
[0048] Step 3: Take the above-mentioned various physiological systems and the thermal sensation model as inputs to construct an integrated dynamic thermal sensation model, which is the comprehensive thermal sensation output based on the multi-model integration method, as shown in formula (5).
[0049]
[0050] To identify the weight coefficients of the integrated model, the least squares method is used in the first stage to optimize the static weight values and minimize the objective function. The objective function is set as:
[0051]
[0052] where the objective function θ k represents the sum of the squared errors between the observed and predicted thermal sensations; m represents the number of samples used to estimate the state parameters.
[0053] Subsequently, the particle filter method is used to update the static estimation coefficients, including the prediction process, importance sampling process, and resampling process to eliminate the estimation errors of the static method, and finally the weight coefficients are obtained. The prediction process is expressed as:
[0054]
[0055] where α′ k,i represents the weight of the i-th particle at time k; α k-1,i represents the weight of the i-th particle at time k - 1, and r k-1,i represents the state noise of the i-th particle.
[0056] The importance sampling process based on the measurement equation calculates the particle weights by comparing the current observation value with the particle values predicted by the state equation and normalizes them based on the measurement equation.
[0057]
[0058] In the formula, q k,i represents the particle weight approximating p(α k |c 1:k ), p(α k |c 1:k ) represents the state posterior probability density at time k, α k represents the set of weights of all particles at time k, and c 1:k represents the measurement set of the sensor within the time range from 1 to k.
[0059] Based on the particle weight distribution, particles are screened for resampling. After resampling, finally, the average value of all particle weights is used as the weight vector α′ k at this time step. The α′ k obtained by the particle filter is used to integrate and combine the multi-modal parameter stepwise thermal sensation basic model.
[0060] Step 4: Predict the thermal sensation at the next moment through the integrated model Finally, the predicted value of thermal sensation under the time series is obtained as
[0061] Detect and evaluate the performance of the integrated dynamic thermal sensation prediction model. Based on the multi-parameter step-by-step thermal sensation prediction basic model constructed in Step 4, specifically analyze the performance of the integrated dynamic thermal sensation prediction model according to the experimental data samples, and introduce the coefficient of correlation (R 2 ) and the coefficient of variation of root mean square error (CV_RMSE) to calculate the difference between the predicted thermal sensation of the person by the integrated dynamic model and the true thermal sensation state, so as to evaluate the prediction accuracy of the model. The specific effect differences are as Figure 3 shown. It can be seen from the figure that the prediction accuracy of the integrated static thermal sensation prediction model (StaticCTS) is stable at 96.98%, the accuracy of the integrated dynamic thermal sensation prediction model (Dynamic CTS) can reach 98.24%, and the CV_RMSE is within 14.66%. Compared with the step-by-step thermal sensation model, the accuracy of the integrated thermal sensation coupling model can be improved by 7.03% - 77.23%, and the CV_RMSE can be reduced by 53.74% - 82.06%.
[0062] Step 5: The methodological construction of the probability model for the person's air-conditioning adjustment behavior intention can respectively obtain the probability expressions for the person to increase / decrease the air-conditioning behavior intention, and the probability distribution of the person not having the air-conditioning adjustment behavior intention. Through fitting the experimental data, its probability density distribution function is as shown in formula (9) and follows a normal distribution:
[0063]
[0064] According to the summary of experimental data, it is found that the general model of the air-conditioning adjustment behavior intention probability of the measured subjects in different thermal sensation states is as Figure 4 shown. In the figure, SI1 represents the conversion critical point between the intention to increase the air-conditioning behavior and the intention of no adjustment behavior, and SI2 represents the conversion critical point between the intention to decrease the air-conditioning behavior and the intention of no adjustment behavior. Initialize the Q-table and embed the person's air-conditioning adjustment behavior intention probability model in the reinforcement learning Q-learning model as the initial probability for model training. In this specific implementation case, set the learning rate of the reinforcement learning model to 0.5, the reward discount factor to 0.8, the exploration probability to 0.9, and other parameters such as the maximum number of iterations to 1000 times, and establish a reinforcement learning model considering the person's regulation behavior intention probability.
[0065] Step 6: Based on Step 5, by monitoring the changes in the indoor temperature and the thermal sensation state of the occupants at this moment, calculate the reward obtained by the model and update the Q-table to train the reinforcement learning model considering the probability of the occupants' regulation behavior intention. When updating the Q value, assign the initial probability value to the initial Q0(s,a), and update the Q value as the iteration progresses. This can make high-probability actions be selected in the initial iteration process, corresponding to larger Q0 values, thus guiding the agent to be more likely to choose these actions.
[0066] Specifically, assume that the current state is s, the action executed is a, the obtained reward is r, and the next state is s'. The update formula for the Q value can be expressed as:
[0067] Q0(s,a) = P0(s,a) (10)
[0068]
[0069] where α is the learning rate, γ is the discount factor, max(Q(s′,a′)) represents the maximum Q value selected in the next state s′, and P0(s,a) represents the probability of selecting action a in the initial state.
[0070] Step 7: Based on Step 6, continuously update the Q-table until the number of training times reaches the set maximum number of iterations, indicating that the model learning process has been completed. Output the optimal strategy for the indoor temperature set value, form an indoor temperature control method based on Q-learning reinforcement learning, so as to achieve the best control of the indoor temperature.
[0071] Discuss the regulation effect of the Q-Learning model embedded with the probability model of the occupants' adjustment behavior intention, as Figure 5 shown. Compared with not initializing the adjustment behavior probability, the number of steps for searching and training the temperature adjustment strategy to achieve the target thermal sensation state is less under the initialization of considering the probability of the occupants' air-conditioning adjustment behavior intention. Compared with the strategy without the embedded intention model, the number of training iteration steps of the embedded intention model strategy can be reduced by at least 50.0%, greatly improving the efficiency of model training, and being able to adjust the temperature to the target thermal sensation position at the fastest speed, thus realizing a fast and accurate indoor environment regulation strategy.
Claims
1. An indoor thermal environment regulation technology based on the multi-modal thermal perception state of personnel, characterized in that, It includes the following steps: Step 1): Use a personal body area network sensing device to obtain the physiological characteristic parameters (physiological heat exchange system, cardiovascular system, brain nervous system) of a person in a thermal environment. Obtain the thermal sensation votes of the person by means of questionnaire surveys. Step 2): Take each physiological characteristic parameter and the thermal sensation vote (Thermal Sensation Vote, TSV) of the person as inputs to establish a thermal sensation state model of different physiological systems of the person under specific thermal environment conditions. The specific model is as follows: PTS1 = a × SET + b (1) PTS4 = Y[XGBoost(e n (x i ))] Equation (5) Equation (1) represents a thermal sensation prediction model based on a physiological heat exchange system. PTS1 represents the thermal sensation voting value obtained from the physiological heat exchange system. In the equation, a is the gradient of the fitting line, b is the intercept, and SET represents the standard effective temperature (unit: °C). Equation (2) represents a prediction model based on the relationship between the heart rate change rate and thermal sensation. PTS2 represents the thermal sensation vote value obtained from the heart rate change rate. In the equation, A is the limit coefficient of thermal sensation, and its value is the upper limit value of thermal sensation, which is 3, and the lower limit value of thermal sensation, which is -3; B is the slope coefficient. HR is the heart rate parameter response (unit: bpm), and HR n is the average heart rate base value when it is neutral (unit: bpm). Equations (3) and (4) represent a prediction model based on the relationship between heart rate variability and thermal sensation. It represents the thermal sensation vote value obtained from heart rate variability. Among them, Equation (3) is applicable when the indoor temperature is greater than 26°C, and Equation (4) is applicable when the indoor temperature is less than or equal to 26°C. LF is the low-frequency power, HF is the high-frequency power, and m, n, and q are fitting coefficients. Equation (5) represents a thermal sensation prediction model based on EEG signal features. PTS4 represents the thermal sensation voting value predicted from the EEG signal. In the equation, Y represents the output data set of the predicted value, e n (x i ) where x i represents the i-th eigenvalue extracted from the EEG signal. The superscript n represents the number of sample groups. XGBoost(e n (x i )) represents the calculation of the sample parameter x i using the XGBoost-based model. Step 3): Based on Step 2), combine and predict each physiological system with the thermal sensation state model to construct an integrated dynamic thermal sensation prediction model As shown in Equation (6). where α i = [α1, α2, α3, α4] T , representing the weight coefficients; PTS i = [PTS1, PTS2, PTS3, PTS4] T , representing the above-mentioned stepwise thermal sensation prediction models. Step 4): Based on Step 3), use the integrated model to predict the thermal sensation at the next moment Finally, the predicted value of the integrated thermal sensation state in the time series is Step 5): Construct the probability distribution of the person's air-conditioning adjustment behavior intention, initialize the Q-table and embed the probability distribution of the person's air-conditioning adjustment behavior intention, and set parameters such as the learning rate, reward discount factor, exploration probability, and maximum number of iterations of the reinforcement learning model to establish a reinforcement learning model considering the probability of the person's regulation behavior intention. Step 6): Based on Step 5), by monitoring the current indoor temperature and the change in the thermal sensation state of the person, calculate the reward obtained by the model and update the Q-table to train the reinforcement learning model considering the probability of the person's regulation behavior intention. The calculation formula is as follows: q(s t ,a t ) ← q(s t ,a t ) + α[r t+1 + γ max(s t+1 ,a t+1 ) - q(s t ,a t )] (7) In the formula, q(s,a) represents the goodness or badness of selecting action a in the current state s. The subscript t represents the t-th moment, r represents the reward obtained after executing action a in state s, γ is the discount factor, indicating the importance of future rewards to the current reward, and γ ∈ [0,1]. Step 7): Based on Step 6), continuously update the Q-table until the number of training times reaches the set maximum number of iterations, indicating that the model learning process has been completed. Output the optimal strategy of the indoor temperature setting value to form an indoor temperature control method based on Q-learning reinforcement learning, so as to achieve the best control of the indoor temperature.
2. An indoor thermal environment regulation technology based on the multi-modal thermal perception state of personnel according to claim 1, characterized in that In step 3), based on the construction of the integrated dynamic thermal sensation by coupling the multi-modal information of human thermal physiology in step 2), the multi-modal thermal sensation model integration method is to use the weight coefficient α i to combine the step-by-step models Use the least squares method to assign the weight coefficient α to the integrated model i , to reflect the influence degree of each modal physiological parameter on the human thermal sensation state; then use the particle filter method to update the static estimation coefficient to eliminate the estimation error of the static method.
3. An indoor thermal environment regulation technology based on the multi-modal thermal perception state of personnel according to claim 1, characterized in that, The general behavior model of the probability of the person's air-conditioning adjustment intention established in Step 5), and obtain the thermal sensation interval band of the person. The probability model of the air-conditioning behavior adjustment intention generated by the person under the stimulation of the thermal environment conditions is as shown in the formula: In the formula, μ is the position coefficient; β is the scale coefficient, and S is the thermal sensation vote value. The thermal sensation interval band of the person can be obtained through fitting, which is divided into the thermal sensation interval TS Interval with weak behavior orientation Low , that is, a specific thermal sensation range that meets the comfort requirements of the person's thermal sensation state and has a very low probability of generating adjustment behavior; and the thermal sensation interval TS Interval with strong behavior orientation High , that is, a specific thermal sensation range with a very high probability of generating adjustment behavior.
4. An indoor thermal environment regulation technology based on the multi-modal thermal perception state of personnel according to claim 1, characterized in that, In step 6), the reinforcement learning control strategy that embeds the probability guidance mechanism of the personnel's adjustment behavior intention is considered. During the training process of the reinforcement learning model, the probability of the personnel's temperature adjustment behavior (raising or lowering) is taken into account. When setting the rewards in the Q-table, the guidance and tendency of temperature adjustment (raising and lowering) are given by setting the probability of the adjustment behavior intention, so that the strategy can reach the target position at the fastest speed, that is, the indoor temperature is regulated to the thermal sensation interval TS Interval pointed to by the weak behavior of the personnel. Low This goal is achieved by adding the behavior reward tendency probability to the reward function. As the number of learning times increases, the control system can achieve the optimal control of the room temperature for different individual thermal sensation states.