Indoor environment comfort level control system and method based on reinforcement learning

By introducing reinforcement learning technology into the indoor environmental control system, comprehensively considering multiple environmental variables, and formulating dynamic control strategies, it solves the problem that traditional control systems are difficult to achieve optimal overall comfort balance and energy waste, and achieves the goals of precise control and energy conservation and emission reduction.

CN119983459APending Publication Date: 2025-05-13HANGZHOU DIANZI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411939921.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional indoor environment control systems rely on simple threshold setting and feedback adjustment, making it difficult to achieve the best balance of overall comfort, resulting in waste of energy and inconsistent with modern energy-saving and environmental protection concepts.

Method used

A control system based on reinforcement learning is adopted, through the interaction between the agent and the environment, and based on the reward signal of environmental feedback, the control strategy is learned and optimized, and multiple environmental variables such as temperature, humidity, and air flow rate are comprehensively considered to formulate dynamic control strategies.

Benefits of technology

It achieves precise control of indoor environment comfort, improves overall comfort, reduces energy consumption, and meets energy conservation and environmental protection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119983459A_ABST
    Figure CN119983459A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor environment comfort level control system and method based on reinforcement learning. A processor is connected with a single-chip microcomputer through a line, and the single-chip microcomputer is connected with a humidity sensor, a human body existence sensing radar and a thermosensitive wind speed sensor through lines; the processor is in wireless communication with the humidifier, and the single chip microcomputer is in wireless communication with the air conditioner; the single-chip microcomputer receives environmental parameters collected by the humidity sensor, the human body existence sensing radar and the thermosensitive wind speed sensor and sends the environmental parameters to the processor. And the processor adopts a TD3 model to calculate a comfort PMV value through the environmental parameters, inputs the environmental parameters and the PMV value into the TD3 model, outputs set values of the air conditioner and the humidifier, and finally sends a control command according to the set values. Accurate control over the indoor environment comfort degree is achieved by integrating multi-sensor data through reinforcement learning, meanwhile, energy consumption is effectively optimized, and an effective solution is provided for intelligent control over the indoor environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent buildings, and in particular to an indoor environment comfort control system and method based on reinforcement learning. Background Art

[0002] With the rapid development of Internet of Things technology, more and more intelligent devices are integrated into modern buildings, forming so-called smart buildings. Smart buildings achieve real-time detection and control of indoor environment by connecting various sensors, actuators and central controllers. These sensors can collect environmental parameters such as temperature, humidity, air flow rate, etc., while actuators can control the working status of home appliances such as air conditioners and humidifiers.

[0003] Although the Internet of Things technology has brought convenience to indoor environmental control, traditional control methods still have some limitations. Most traditional indoor environmental control systems are based on simple threshold settings and feedback adjustment mechanisms. For example, the air conditioning system cools or heats according to a preset temperature range. These rely on pre-set rules and lack comprehensive consideration and coordinated optimization of multiple environmental factors, making it difficult to achieve the best balance of overall comfort. For example, when adjusting the temperature, humidity imbalance may occur, which in turn affects the body's perceived comfort. Due to the use of relatively fixed and simple control strategies, energy waste often occurs, such as excessive cooling or heating, unnecessary humidification, etc., which does not conform to the modern concept of energy conservation and environmental protection. For example, the air conditioner still maintains high power operation when no one is there, or the humidifier continues to work when the humidity has reached the standard.

[0004] In recent years, reinforcement learning, as an adaptive learning method, has shown great potential in the field of automatic control. Reinforcement learning continuously learns and optimizes its own behavior strategy through the interaction between the intelligent agent and the environment, according to the reward signal fed back by the environment, in order to achieve specific goals. In terms of indoor environmental comfort control, reinforcement learning has unique advantages: it can take multiple environmental variables such as temperature, humidity, air flow rate as state input, and comprehensively consider the complex relationship between various factors; through continuous learning and optimization, it can formulate dynamic control strategies for different scenarios, thereby effectively improving the overall comfort of the indoor environment; and it can avoid excessive energy consumption and achieve the goal of energy conservation and emission reduction while ensuring comfort. Therefore, the application of reinforcement learning to indoor environmental control has important practical significance and broad application prospects. It can overcome many shortcomings of traditional control methods and provide people with a more intelligent, comfortable and energy-saving indoor environment experience. Summary of the invention

[0005] The purpose of the present invention is to propose an indoor environment comfort control system and method based on reinforcement learning.

[0006] The invention discloses an indoor environment comfort control system based on reinforcement learning, comprising a processor, a single chip microcomputer, a humidity sensor, a human presence sensing radar and a thermal wind speed sensor, wherein the processor is connected to the single chip microcomputer through a line, and the single chip microcomputer is connected to the humidity sensor, the human presence sensing radar and the thermal wind speed sensor through a line; the processor communicates with the humidifier wirelessly, and the single chip microcomputer communicates with the air conditioner wirelessly; the single chip microcomputer receives environmental parameters collected by the humidity sensor, the human presence sensing radar and the thermal wind speed sensor, and sends the environmental parameters to the processor; the processor analyzes the environmental parameters by using a TD3 model, and then sends a control instruction to the humidifier and the single chip microcomputer, and the single chip microcomputer sends the received air conditioning control instruction to the air conditioner.

[0007] Preferably, the processor adopts Rockchip's RK3399 processor; the microcontroller adopts Lianshengde's W801 microcontroller; the microcontroller is connected to the human presence sensing radar through the UART interface, the microcontroller is connected to the temperature and humidity sensor through the I2C interface, and the microcontroller is connected to the thermal wind speed sensor through the RS485 interface; the microcontroller communicates with the processor through the RS485 bus, and the microcontroller communicates wirelessly with the air conditioner through an infrared transceiver. The processor communicates wirelessly with the humidifier through the Internet of Things API control.

[0008] Based on the above device, there is the following indoor environment comfort control method, which specifically includes the following steps: Step 1: Human presence detection By installing a human presence sensing radar in the indoor environment, it detects whether there is someone in the room. If someone is detected, the air conditioner and humidifier are started, and the system begins to control the indoor environment comfort.

[0009] Step 2: Environmental data collection The single-chip microcomputer receives the indoor temperature and humidity data collected by the temperature and humidity sensor and the indoor wind speed data collected by the thermistor wind speed sensor at the set frequency; the single-chip microcomputer waits for the processor request, and after receiving the processor request, sends the received temperature, humidity and wind speed data to the processor.

[0010] Step 3: Calculate the comfort PMV value Based on the collected temperature, humidity and wind speed data, the processor substitutes the temperature, humidity and wind speed into the corresponding PMV calculation model to obtain a PMV value that can quantitatively represent the current indoor environmental comfort conditions.

[0011] PMV=[0.303exp(−0.036M)+0.0275]*{M−W−3.05[5.733−0.007(M−W−3.05[5 .733−0.007(M−W)−Pa]−0.42(M−W−58.2)−0.0173M(5.867−Pa)−0.0014M(34− t a )−3.96∗10 f cl [ ( t cl +273) 4 − ( t r +273) 4 ]− f cl h c ( t cl − t a )} Where M is the metabolic rate of the human body, in W / m 2 , take 60W / m in working state 2 ; W is the mechanical work done by the human body, the unit is W / m 2 , 0W / m when stationary 2 ;t a is the air temperature in °C; f cl It is the area coefficient of human clothing, and the coefficient range is 0~2m 2 K / W; t r The average radiation temperature of the room where the human body is located is equal to the average indoor temperature, in degrees Celsius; t cl is the outer surface temperature of human clothing, in °C; t cl =35.7−0.028(M−W)− I cl {(3.96∗10^−8) f cl [ ( t cl +273) 4 − ( t r +273) 4 ]+ f cl h c ( t cl − t a )} ; Among them, I cl is the thermal resistance of the clothing, in m 2 K / W; h c Is the convective heat transfer coefficient, unit is W / (m 2 K); when hour, ;when hour, ; Where v is the wind speed in m / s; Pa is the partial pressure of water vapor around the human body, in Pa; , where φ is the relative humidity; Step 4: Reinforcement Learning Model Decision The collected temperature ,humidity , wind speed Data and calculated PMV values ​​as status information S=[ T in Φ in V in PMV , input to the trained TD3 algorithm's policy network, and output the control strategy A=[ T set Φ set V set ] ; in The indoor ambient temperature collected by the sensor, The indoor environment humidity collected by the sensor, is the indoor environmental wind speed collected by the sensor, and PMV is the comfort value calculated in step 3; Set the air conditioner temperature. is the wind speed setting value, Set the humidity value for the humidifier; The TD3 algorithm adopts a strategy-evaluation structure, including two Q networks, one strategy network, two target Q networks and one target strategy network. The strategy network realizes the mapping from state S to action A, and the Q network is used for quantitative evaluation of the state-action value function. When the PMV value is close to zero, the indoor environment is at a high comfort level, and a positive reward is given at this time; when the PMV value deviates from the comfort range, a negative reward is added according to the degree of deviation; the PMV range is set between -0.5 and +0.5 to indicate comfort, and outside this range is considered uncomfortable; the reward function R of TD3 model training is set as: ; Through 2 target Q networks and 1 target policy network, according to the above reward function and Bellman equation , update 2 Q networks and 1 policy network, and train the policy network.

[0012] in, and They represent the state and action at time t respectively, is the state-action value at time t, is the reward value at time t, is the depreciation factor, the value range is (0,1); The trained policy network is deployed on the processor to output the control policy.

[0013] Step 5: Send control commands The processor sends control commands to the air conditioner and humidifier. After receiving the set temperature and wind speed commands, the air conditioner adjusts the temperature and wind speed to the output values ​​of step 4, and the humidifier adjusts the humidification amount according to the set humidity value.

[0014] Step 6: Repeat the operations of steps 2 to 5 according to the sampling frequency of the single-chip microcomputer, so that the indoor environment always tends to the optimal comfort level; until the human presence sensing radar detects that there is no one in the room, and the air conditioner and humidifier are in working state, then turn off the air conditioner and humidifier, and end the comfort adjustment at the same time to avoid unnecessary energy consumption.

[0015] Preferably, the single-chip microcomputer communicates with the processor via an RS485 bus. The single-chip microcomputer controls the temperature, wind speed, start and shut down of the air conditioner via an infrared transceiver. The processor controls the humidity, start and shut down of the humidifier via an Internet of Things API.

[0016] The beneficial effects of the present invention are: The use of advanced automatic control systems can monitor and adjust indoor environmental parameters in real time without human intervention, reducing the complexity and uncertainty of manual operations and improving the system's response speed and control accuracy.

[0017] By real-time detection of the presence of human bodies in the environment, air conditioners, humidifiers and other electrical appliances are immediately turned off when no one is around to prevent high-power idling, thereby achieving energy-saving and consumption-reducing effects to meet energy-saving and environmental protection needs.

[0018] By introducing the PMV comfort evaluation index and simultaneously monitoring and adjusting temperature, humidity and wind speed, the system can more comprehensively reflect the environmental conditions, ensuring that the indoor environment always remains in the most suitable state, effectively avoiding the discomfort that may be caused by adjusting a single parameter, and improving the overall user experience.

[0019] Reinforcement learning is used to integrate multi-sensor data to achieve precise control of indoor environmental comfort, while effectively optimizing energy consumption, providing an effective solution for intelligent control of indoor environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a system hardware architecture diagram of the present invention; Figure 2 It is a flow chart of the control method of the present invention. DETAILED DESCRIPTION

[0021] The present invention is further described below in conjunction with the accompanying drawings.

[0022] like Figure 1The indoor environment comfort control system based on reinforcement learning is shown, including a processor, a single-chip microcomputer, a humidity sensor, a human presence sensing radar and a thermal wind speed sensor. The processor is connected to the single-chip microcomputer through a line, and the single-chip microcomputer is connected to the humidity sensor, the human presence sensing radar and the thermal wind speed sensor through a line; the processor communicates with the humidifier wirelessly, and the single-chip microcomputer communicates with the air conditioner wirelessly; the single-chip microcomputer receives environmental parameters collected by the humidity sensor, the human presence sensing radar and the thermal wind speed sensor, and sends the environmental parameters to the processor; the processor uses the TD3 model to analyze the environmental parameters, and then sends the control instructions to the humidifier and the single-chip microcomputer, and the single-chip microcomputer sends the received air conditioning control instructions to the air conditioner.

[0023] In this embodiment, Rockchip's RK3399 is used as the processor, running the Android system, using the RS485 bus to establish a stable connection with the sensor module, and accurately obtaining the environmental parameters collected by the sensor, including the presence of the human body, temperature and humidity, and air flow rate. At the same time, the processor has powerful wireless communication capabilities such as WiFi, and can use the Internet of Things API to send accurate control instructions to the humidifier device to achieve effective regulation of the indoor humidity environment.

[0024] The W801 single-chip microcomputer of Lianshengde is adopted, and the human presence sensing radar, temperature and humidity sensor, and thermal wind speed sensor are mounted on the board; in this embodiment, the human presence sensing radar of model LD2410B is adopted, which is connected to the single-chip microcomputer through the UART interface to capture the presence and activity of indoor personnel; the temperature and humidity sensor of model AHT10 is adopted, which is connected to the single-chip microcomputer through the I2C interface to monitor the indoor temperature and humidity changes in real time; the thermal wind speed sensor of model GT-RFS-12V is adopted, which is connected to the single-chip microcomputer through the RS485 interface to measure the indoor vacancy flow rate; the air conditioner communicates with the single-chip microcomputer through the infrared transceiver module YS-IRTM and the UART interface, and the single-chip microcomputer sends instructions issued by the processor to the air conditioner.

[0025] like Figure 2 As shown, based on the above device, the following indoor environment comfort control method is proposed, which specifically includes the following steps: Step 1: System hardware connection and initialization Complete the installation and configuration of the Android system of the processor to ensure that it can run stably and has the ability to communicate with peripherals. Connect the human presence sensing radar LD2410B, temperature and humidity sensor AHT10, thermal wind speed sensor GT-RFS-12V and infrared transceiver module YS-IRTM to the corresponding pins of the sensor module according to the electrical connection specifications, and realize the functions of the microcontroller reading the corresponding sensor and infrared transceiver. Use the RS485 bus to connect the controller module to the sensor module. The communication method adopts the client-server mode, such as Modbus-RTU. The sensor module acts as a Modbus server and the controller module acts as a Modbus client. The controller module obtains sensor data by requesting data from the server through the client. At the same time, initialize the Modbus-RTU communication parameters, such as baud rate, data bits, registers, etc., to ensure stable and efficient data transmission between the two.

[0026] Step 2: Human presence detection and electrical appliance startup control After the system is started and initialized, the controller module sends instructions to the sensor module to detect whether there is someone in the room. If a signal of someone's presence is detected, the sensor module transmits the information to the controller module. After receiving the information that someone is present, the controller module starts the air conditioner and humidifier to start working and prepare to enter the environmental data collection phase. If no one is detected, it will continue to be in the monitoring state, and the air conditioner and humidifier will not be started, keeping the system in low power standby mode.

[0027] Step 3: Environmental data collection In this embodiment, the data collection cycle is set to 5 minutes, and the timer is started. When the timer reaches the set cycle, the controller module sends a data collection instruction to the sensor module. After receiving the instruction, the sensor module integrates and preliminarily processes the temperature, humidity and wind speed data collected from the corresponding sensor according to the Modbus-RTU data specification (data format conversion, error checking, etc.), and transmits these data to the controller module via RS485.

[0028] Step 4: Calculate the comfort PMV value After the controller module receives the temperature, humidity and wind speed data from the sensor module, it calls the PMV value calculation program built into the software to calculate the comfort value PMV. PMV calculation formula:

[0029] PMV=[0.303exp(−0.036M)+0.0275]*{M−W−3.05[5.733−0.007(M−W−3.05[5 .733−0.007(M−W)−Pa]−0.42(M−W−58.2)−0.0173M(5.867−Pa)−0.0014M(34− t a )−3.96∗10 f cl [ ( t cl +273) 4 − ( t r +273) 4 ]− f cl h c ( t cl − t a )} Where M is the metabolic rate of the human body, in W / m 2 , generally 60W / m 2 ; W is the mechanical work done by the human body, the unit is W / m 2 , generally 0W / m 2 ;t a is the air temperature in °C; f cl It is the area coefficient of human clothing, and the coefficient range is 0~2m 2 K / W, generally 1.1m 2 K / W; t r is the average radiant temperature of the room where the human body is located, in °C; t cl is the outer surface temperature of human clothing, in °C; t cl =35.7−0.028(M−W)− I cl {(3.96∗10^−8) f cl [ ( t cl +273) 4 − ( t r +273) 4 ]+ f cl h c ( t cl − t a )} , where I cl is the thermal resistance of the clothing, in m 2 ·K / W.

[0030] h c Is the convective heat transfer coefficient, unit is W / (m 2 K); when hour, ;when hour, ; Where v is the wind speed in m / s.

[0031] Pa is the partial pressure of water vapor around the human body, in Pa; ; where φ is the relative humidity.

[0032] In the PMV calculation formula, environment-related variables include air temperature, mean radiation temperature, wind speed, and relative humidity. In a specific environment, the air temperature is not much different from the mean radiation temperature. Generally, the mean radiation temperature is treated as the average indoor temperature. Other uncontrollable variables such as human physiological parameters are treated as constants, as shown in Table 1.

[0033] Table 1 variable Numeric unit Human metabolic rate M 60 <![CDATA[W / m 2 ]]> The external mechanical work of the human body W 0 <![CDATA[W / m 2 ]]> <![CDATA[Area coefficient f of the human body's clothing cl > 1.1 <![CDATA[m 2 ·K / W]]> <![CDATA[Clothing thermal resistance I cl > 1.1 <![CDATA[m 2 ·K / W]]> Based on the internationally recognized PMV calculation model, the temperature, humidity and wind speed data as well as some preset human physiological parameters are input into the processor to obtain the comfort PMV value of the current indoor environment.

[0034] Step 5: Reinforcement Learning Model Decision The reinforcement learning model is trained using the Pytorch computing framework based on historical environmental data, and the model uses the TD3 algorithm. The TD3 algorithm adopts a policy-evaluation structure, including 2 Q networks, 1 policy network, 2 target Q networks, and 1 target policy network. The policy network realizes the mapping from state S to action A, and the Q network is used for the quantitative evaluation of the state-action value function.

[0035] Among them, the policy network takes state information as input and is expressed as S=[ T in Φ in V in PMV , including the indoor environmental parameters temperature collected by the sensor ,humidity , wind speed , and the comfort value PMV calculated in step 4. The output action is expressed as A=[ T set Φ set V set ] , including the temperature setting value of the air conditioner , Wind speed setting value and humidifier humidity setpoint The Q network takes state information and action strategy as input and outputs the state-action value Q, that is, All Q networks and policy networks are composed of 3 layers of MLP (Multi-layer Perceptron).

[0036] Based on the definition of PMV comfort index, when the PMV value is close to zero, it means that the indoor environment is at a high comfort level, and a positive reward is given at this time; when the PMV value deviates from the comfort range, a negative reward is added according to the degree of deviation. According to the current design specifications, a PMV range between -0.5 and +0.5 indicates comfort, and a range outside this range is considered uncomfortable. The reward function R for TD3 model training is set as: ;

[0037] In order to prevent the algorithm from ignoring the energy consumption problem, in this embodiment, an energy consumption penalty term is added to the reward function, and the energy consumption E of the appliance is used as the penalty basis. The energy consumption is obtained by calculating the operating power within the operating time window of the smart socket. ,in, is the electrical operating power at time t, is the time interval between two adjacent energy consumption statistics. There are m moments in the running time window, and a penalty is added: ;in, is the energy consumption weight.

[0038] Combining comfort and energy consumption, the total reward function is defined as: .in, Reward for comfort, For energy consumption rewards, is the penalty factor.

[0039] Through 2 target Q networks and 1 target policy network, according to the above reward function and Bellman equation , update the Q network and policy network, where, and They represent the state and action at time t respectively, is the state-action value at time t, is the reward value at time t, is the depreciation factor, and its value range is (0,1). Migrate the trained policy network to the processor and collect the temperature ,humidity , wind speed Data and calculated PMV values ​​as status information S=[ T in Φ in V in PMV , input to the trained policy network, output control policy A=[ T set Φ set V set ] .

[0040] The model training process is as follows: Initialize the Q network, policy network, target Q network, target policy network and experience pool parameters; Get the current status , get the action through the policy network ; Execute action , get the reward value and the state at the next moment ; The experience sample ( , , , ) is added to the experience pool, and the previous operation is repeated until the experience pool is full; Randomly sample small batches of samples in the experience pool; Add noise n, Obtained through the target strategy network , and the noise added Perform smoothing regularization; through two target Q networks, take the minimum value and process it with the Bellman equation to obtain the target Q value: ;in, is the depreciation factor, and its value range is (0,1).

[0041] Will and Through 2 Q networks, and the target Q value Calculate the mean square error to get the loss of the Q network and update the Q network gradient; state Get actions through the policy network , using Q network to evaluate Q value , the loss function is , update the policy network parameters; update the parameters of the target Q network and the target policy network; Reconstruct the experience pool for training until the training is completed.

[0042] TD3 is an algorithm in the field of continuous control, which is very suitable for temperature and humidity control, which has high requirements for comfort and continuous action space. The double-delay update strategy is introduced to reduce the deviation of future reward estimation and improve the stability and convergence speed of the algorithm in complex environments.

[0043] After the sensors in the room collect environmental information such as temperature and humidity data and human presence, the TD3 algorithm will process this state information in real time, and after normalization or standardization preprocessing, input it into the state space. Then, the action value, that is, the temperature, humidity, and wind speed setting value, is generated through the strategy network to achieve the optimal adjustment of comfort.

[0044] At the same time, the TD3 algorithm provides better direction guidance for the strategy network by evaluating the effectiveness of the current strategy, so that the system can significantly reduce energy consumption while meeting comfort requirements.

[0045] Step 6: Send control commands The controller module obtains the air conditioning set temperature calculated by the TD3 algorithm , Humidifier humidity setting And air conditioning set wind speed After that, for the humidifier, according to the IoT API specification provided by its manufacturer, the set humidity value is converted and packaged into a data format that meets the requirements. After establishing a connection with the WiFi capability, a control command request is sent to the humidifier according to the corresponding network protocol, and the humidifier adjusts the humidification working state accordingly. For the air conditioner, the infrared transceiver module of the sensor module converts the set temperature and wind speed values ​​into corresponding infrared coding signals according to the specific infrared coding protocol of the air conditioner manufacturer, drives the transmitting tube to send to the air conditioner, and the air conditioner receives and adjusts the set temperature and wind speed.

[0046] Step 7: System cycle operation After the system is started, the human presence detection mechanism will continue to run to detect the presence of human beings in the room in real time. Once the presence of human beings is detected in the room, the air conditioner and humidifier will be turned on immediately; if no one is detected, the air conditioner and humidifier will be turned off. After completing the sending of a control command and the execution of the device, the system will continue to execute the steps of environmental data collection, PMV value calculation, TD3 model decision and control command sending according to the set 5-minute cycle.

[0047] To ensure the effectiveness of the above method, temperature and humidity sensors and thermal wind speed sensors can be placed in specific sensitive areas in the space for data collection and real-time feedback. Environmental data sensitive areas include but are not limited to: the highest and lowest points of thermal comfort values, the location under the air conditioner (the area near the air outlet of the air conditioner, where the temperature and humidity often change greatly), and areas with frequent human activities, usually work areas, sleeping areas and other places where there are a lot of human activities, and the comfort and temperature and humidity requirements in these areas change greatly.

[0048] Human presence sensing radars are placed at the entrance of the room and in major activity areas to ensure accurate detection of the presence of people. This layout ensures that the system can obtain environmental data and personnel activity information in real time under different circumstances, providing a reliable basis for control strategies.

Claims

1. A method for controlling indoor environmental comfort based on reinforcement learning, characterized in that: The specific steps include: Step 1: Human presence detection By installing a human presence sensing radar in the indoor environment, it detects whether there is someone in the room; if someone is detected, the air conditioner and humidifier are started, and the system begins to control the indoor environment comfort; Step 2: Environmental data collection The single chip microcomputer receives the indoor temperature and humidity data collected by the temperature and humidity sensor and the indoor wind speed data collected by the thermal wind speed sensor according to the set frequency; The single-chip microcomputer waits for the processor request, and after receiving the processor request, sends the received temperature, humidity and wind speed data to the processor; Step 3: Calculate the comfort PMV value Based on the collected temperature, humidity and wind speed data, the processor substitutes the temperature, humidity and wind speed into the corresponding PMV calculation model to obtain a PMV value that can quantitatively represent the current indoor environmental comfort status; Where M is the metabolic rate of the human body, in W / m 2 , 60W / m in working state 2 ; W is the mechanical work done by the human body, the unit is W / m 2 , 0W / m when stationary 2 ;t a is the air temperature in °C; f cl It is the area coefficient of human clothing, and the coefficient range is 0~2m 2 K / W; t r The average radiation temperature of the room where the human body is located is equal to the average indoor temperature, in degrees Celsius; t cl is the outer surface temperature of human clothing, in °C; ; Among them, I cl is the thermal resistance of the clothing, in m 2 K / W; h c Is the convective heat transfer coefficient, unit is W / (m 2 K); when hour, ;when hour, ; Where v is the wind speed in m / s; Pa is the partial pressure of water vapor around the human body, in Pa; , where φ is the relative humidity; Step 4: Reinforcement Learning Model Decision The collected temperature ,humidity , wind speed Data and calculated PMV values ​​as status information , input to the trained TD3 algorithm's policy network, and output the control strategy ; in The indoor ambient temperature collected by the sensor, The indoor humidity collected by the sensor, is the indoor environmental wind speed collected by the sensor, and PMV is the comfort value calculated in step 3; Set the air conditioner temperature. is the wind speed setting value, Set the humidity value for the humidifier; The TD3 algorithm adopts a strategy-evaluation structure, including two Q networks, one strategy network, two target Q networks and one target strategy network. The strategy network realizes the mapping from state S to action A, and the Q network is used for quantitative evaluation of the state-action value function. When the PMV value is close to zero, the indoor environment is at a high comfort level, and a positive reward is given at this time; when the PMV value deviates from the comfort range, a negative reward is added according to the degree of deviation; the PMV range is set between -0.5 and +0.5 to indicate comfort, and outside this range is considered uncomfortable; the reward function R of TD3 model training is set as: ; Through 2 target Q networks and 1 target policy network, according to the above reward function and Bellman equation , update 2 Q networks and 1 policy network, and train the policy network; in, and They represent the state and action at time t respectively, is the state-action value at time t, is the reward value at time t, is the depreciation factor, the value range is (0,1); Deploy the trained policy network to the processor to output the control policy; Step 5: Send control commands The processor sends control commands to the air conditioner and humidifier. After receiving the set temperature and wind speed commands, the air conditioner adjusts the temperature and wind speed to the output values ​​of step 4, and the humidifier adjusts the humidification amount according to the set humidity value; Step 6: Repeat the operations of steps 2 to 5 according to the sampling frequency of the single-chip microcomputer, so that the indoor environment always tends to the optimal comfort level; until the human presence sensing radar detects that there is no one in the room, and the air conditioner and humidifier are in working state, then turn off the air conditioner and humidifier, and end the comfort adjustment at the same time to avoid unnecessary energy consumption.

2. The indoor environment comfort control method based on reinforcement learning according to claim 1, characterized in that: The single chip microcomputer communicates with the processor via the RS485 bus.

3. The indoor environment comfort control method based on reinforcement learning according to claim 1, characterized in that: The single chip microcomputer controls the temperature, wind speed, start and shut down of the air conditioner through the infrared transceiver.

4. The indoor environment comfort control method based on reinforcement learning according to claim 1, characterized in that: The processor controls the humidity, start and shut down of the humidifier through the Internet of Things API.

5. A system for running the indoor environment comfort control method based on reinforcement learning according to claim 1, characterized in that: It includes a processor, a single-chip microcomputer, a humidity sensor, a human presence sensing radar and a thermal wind speed sensor. The processor is connected to the single-chip microcomputer through lines, and the single-chip microcomputer is connected to the humidity sensor, the human presence sensing radar and the thermal wind speed sensor through lines; the processor communicates with the humidifier wirelessly, and the single-chip microcomputer communicates with the air conditioner wirelessly; the single-chip microcomputer receives environmental parameters collected by the humidity sensor, the human presence sensing radar and the thermal wind speed sensor, and sends the environmental parameters to the processor; the processor uses the TD3 model to analyze the environmental parameters, and then sends the control instructions to the humidifier and the single-chip microcomputer, and the single-chip microcomputer sends the received air conditioning control instructions to the air conditioner.

6. The system according to claim 5, characterized in that: The processor is Rockchip's RK3399 processor.

7. The system according to claim 5, characterized in that: The described single chip microcomputer adopts Lianshengde's W801.

8. The system according to claim 5, characterized in that: The single chip microcomputer is connected to the human presence sensing radar via a UART interface, the single chip microcomputer is connected to the temperature and humidity sensor via an I2C interface, and the single chip microcomputer is connected to the thermal wind speed sensor via an RS485 interface.

9. The system according to claim 5, characterized in that: The single chip microcomputer communicates with the processor via the RS485 bus, and the single chip microcomputer wirelessly communicates with the air conditioner via the infrared transceiver.

10. The system according to claim 5, characterized in that: The processor controls wireless communication with the humidifier through the Internet of Things API.

Citation Information

Cited By

  • Thermal comfort monitoring robot system and method based on multi-modal body language feature recognition and fusion

    CN120839810A