Artificial intelligence climate box adaptive control system based on reinforcement learning algorithm
Through the climate box adaptive control system based on reinforcement learning algorithm, real-time monitoring and dynamic adjustment of climate box environmental parameters is solved, and the problem that traditional systems cannot be dynamically adjusted is achieved, achieving efficient and accurate environmental control and energy consumption optimization.
Patent Information
- Application Number
- CN202510789762.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Traditional climate control systems are difficult to dynamically adjust control strategies according to crop growth stages and external environmental changes, resulting in environmental control lag and inaccuracy.
The artificial intelligence climate box adaptive control system based on reinforcement learning algorithm is adopted to monitor environmental parameters in real time through the data acquisition unit, and the reinforcement learning control unit is used to independently learn and generate control strategies based on crop growth stage, environmental changes and resource consumption state. The environmental parameters in the climate box are adjusted through the control execution unit to form closed-loop feedback to achieve adaptive adjustment.
It improves the accuracy and energy efficiency of environmental regulation, enhances the adaptability of climate chambers to complex environmental changes, ensures that crops obtain the best growth environment at different growth stages, and reduces the need for manual intervention.
Smart Images

Figure CN120295148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control, and particularly to an artificial intelligence climate chamber adaptive control system based on a reinforcement learning algorithm. Background Art
[0002] With the continuous progress of modern agricultural technologies, the role of precision and intelligent control in agricultural production has become increasingly prominent. Especially in the fields of facility agriculture such as greenhouse cultivation and climate chamber cultivation, the growth environment of crops (including parameters such as temperature, humidity, light intensity, and carbon dioxide concentration) has a direct and significant impact on the yield and quality of crops. To maintain a suitable growth environment for crops, traditional climate control systems usually rely on preset environmental parameter thresholds. When the detected environmental data exceeds or falls below the preset range, the control system activates corresponding adjustment devices, such as heaters, humidifiers, lighting fixtures, or ventilation devices, to restore the set environmental state.
[0003] However, there are significant differences in the environmental requirements of crops at different growth stages (such as germination stage, seedling stage, flowering stage, and fruiting stage). It is difficult for traditional control systems to dynamically adjust control strategies according to the crop growth cycle. At the same time, external climate conditions (such as seasonal changes, day-night temperature differences, and sudden weather changes) also have an important impact on the environment inside the greenhouse or climate chamber. Traditional systems lack the ability to adapt to external changes in real time, resulting in the lag and inaccuracy of environmental control.
[0004] Therefore, how to autonomously learn and adjust the climate chamber control strategy based on the crop growth stage, real-time environmental changes, and resource consumption status has become a technical problem to be solved urgently. Summary of the Invention
[0005] The main object of the present invention is to provide an artificial intelligence climate chamber adaptive control system based on a reinforcement learning algorithm, aiming to be able to autonomously learn and adjust the climate chamber control strategy based on the crop growth stage, real-time environmental changes, and resource consumption status.
[0006] To achieve the above object, the present invention proposes an artificial intelligence climate chamber adaptive control system based on a reinforcement learning algorithm, including: A data acquisition unit for real-time collecting a plurality of environmental parameters inside the climate chamber; A reinforcement learning control unit that autonomously learns and generates a control strategy based on the environmental parameters, crop growth stage, real-time environmental changes, and resource consumption status through a reward function, and the reward function comprehensively optimizes the environmental parameter proximity, energy consumption, and response speed; A control execution unit for adjusting the environmental parameters inside the climate chamber according to the control strategy; A communication unit for realizing data transmission and interaction among the units; Among them, the data acquisition unit, the reinforcement learning control unit, and the control execution unit form a closed-loop feedback through the communication unit. The reinforcement learning control unit generates a control strategy that minimizes energy consumption and meets the response speed according to the dynamic changes of environmental parameters and the target values of crop growth stages, and realizes the adaptive adjustment of the climate chamber environment through the control execution unit.
[0007] In an embodiment of the present application, the data acquisition unit includes a temperature sensor, a humidity sensor, a light sensor, a carbon dioxide concentration sensor, and an oxygen concentration sensor. The temperature sensor, the humidity sensor, the light sensor, the carbon dioxide concentration sensor, and the oxygen concentration sensor constitute a sensor group, and a sensor group is provided in each hierarchical structure of the cultivation area in the climate chamber.
[0008] In an embodiment of the present application, the control execution unit includes a heating device, a refrigeration device, a light adjustment device, a fan, an oxygen valve, and a carbon dioxide valve. The heating device and the refrigeration device realize the uniform diffusion of hot and cold air through a circulation fan, and the oxygen valve and the carbon dioxide valve adjust the gas concentration through the side plate air holes.
[0009] In an embodiment of the present application, the reward function of the reinforcement learning control unit is:
[0010] Where represents the immediate reward value obtained after executing the action in the state , represents the current environmental state, including real-time parameters such as temperature, humidity, light, and gas concentration; represents the control action; represents the environmental parameter deviation weight coefficient; is the current environmental parameter; is the ideal value corresponding to the crop growth stage; represents the number of types of monitored environmental parameters; represents the energy consumption weight coefficient; E ( a ) is the energy consumption of the action a ; represents the response speed weight coefficient; δ ( s ) is the response speed reward.
[0011] In an embodiment of the present application, the communication unit includes an Internet of Things gateway and an industrial control computer. The Internet of Things gateway is connected to each sensor through the Modbus protocol and communicates with the industrial control computer through the TCP protocol to realize the transmission of sensor data and control instructions.
[0012] In an embodiment of the present application, the reinforcement learning control unit further includes: A crop growth stage recognition module, configured to determine the current growth stage of the crop according to the historical data of environmental parameters and a preset crop growth model, and dynamically adjust the ideal value in the reward function and the weight coefficient , , ; Wherein, the growth stage recognition module cooperates with the data acquisition unit and the user interface unit, matches the real-time leaf age, plant height and environmental accumulation amount through the crop growth model, determines the stage switching threshold, and triggers the stage-by-stage adaptive update of the control strategy.
[0013] In an embodiment of the present application, the heating device is a resistance heating wire, the refrigeration device is a compressor and a condenser fan, the heating wire and the compressor waste heat are reused, and the environmental parameters are uniformly adjusted through a porous backplane and a circulation fan.
[0014] In an embodiment of the present application, a user interface unit is further included. The user interface unit provides real-time environmental parameter visualization, a manual control interface and an abnormal alarm function, and supports historical data export and comparative experiment configuration.
[0015] In an embodiment of the present application, the IoT gateway controls the power on and off of each device in the execution unit through a solid-state relay.
[0016] In an embodiment of the present application, the training process of the reinforcement learning control unit includes: Initializing model parameters based on historical data or a simulation environment; Generating an adaptive control strategy through environmental state acquisition, action generation, reward calculation and policy iteration optimization; When the reward value converges or reaches a preset number of iterations, output the optimal control strategy.
[0017] Adopting the above technical solution, the artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm can realize the dynamic adaptive adjustment of the internal environment of the climate chamber. Through a reward function that comprehensively considers the environmental parameter proximity, energy consumption minimization and response speed optimization, the reinforcement learning control unit can continuously optimize the control strategy, effectively improve the accuracy of environmental regulation and the energy efficiency utilization rate. At the same time, through the closed-loop feedback mechanism and the efficient data interaction of the communication unit, the highly automated operation and stability of the system are realized, the need for manual intervention is greatly reduced, the adaptability of the climate chamber to complex environmental changes is enhanced, the best growth environment conditions for crops at different growth stages are ensured, and the crop growth quality and experimental controllability are improved. Description of the Drawings
[0018] The present invention will be described in detail below in conjunction with specific embodiments and the accompanying drawings, where: Figure 1 is a schematic structural diagram of the first embodiment of the present invention; Figure 2 is a three-dimensional structural diagram of the climate chamber of the present invention.
[0019] 10. Cabinet area; 11. Upper groove area; 12. Cultivation area; 20. Display screen. Specific Embodiments
[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the following specific embodiments are only used to explain the present invention and do not constitute a limitation to the present invention.
[0021] As Figure 1 shown, in order to achieve the above objectives, the present invention proposes an adaptive control system for an artificial intelligence climate chamber based on a reinforcement learning algorithm, including: A data acquisition unit for real-time acquisition of multiple environmental parameters inside the climate chamber; A reinforcement learning control unit that autonomously learns and generates a control strategy based on the environmental parameters, crop growth stage, real-time environmental changes, and resource consumption status through a reward function, and the reward function comprehensively optimizes the environmental parameter proximity, energy consumption, and response speed; A control execution unit for adjusting the environmental parameters inside the climate chamber according to the control strategy; A communication unit for realizing data transmission and interaction between units; Among them, the data acquisition unit, the reinforcement learning control unit, and the control execution unit form a closed-loop feedback through the communication unit. The reinforcement learning control unit generates a control strategy that minimizes energy consumption and meets the response speed according to the dynamic changes of environmental parameters and the target values of the crop growth stage, and realizes the adaptive adjustment of the climate chamber environment through the control execution unit.
[0022] Specifically, this embodiment provides an adaptive control system for an artificial intelligence climate chamber based on a reinforcement learning algorithm, which specifically includes the following steps: First, multiple environmental monitoring sensors arranged inside the climate chamber are used to real-time acquire environmental parameters. The environmental parameters at least include air temperature, humidity, light intensity, carbon dioxide concentration, and oxygen concentration. Each environmental parameter is connected to the communication unit through the RS485 bus protocol and transmitted to the reinforcement learning control unit.
[0023] Secondly, the reinforcement learning control unit receives the environmental parameters and processes them by invoking the internal reinforcement learning model based on the preset crop growth stage database, real-time environmental change data, and resource consumption status data. The reinforcement learning model initializes parameters such as network weights, learning rates, and discount factors, and simultaneously sets a reward function with the goals of the environmental parameters approaching the target values, minimizing energy consumption, and having the optimal response speed. The reward function is defined as:
[0024] where represents the immediate reward value obtained after executing action in state , represents the current environmental state, including real-time parameters such as temperature, humidity, light, and gas concentration; represents the control action; represents the environmental parameter deviation weight coefficient; is the current environmental parameter; is the ideal value for the corresponding crop growth stage; represents the number of types of monitored environmental parameters; represents the energy consumption weight coefficient; E ( a ) is the energy consumption of action a ; represents the response speed weight coefficient; δ ( s ) is the response speed reward. The reinforcement learning control unit constructs a standardized state vector based on the current environmental state, inputs it into the reinforcement learning network, and generates a control strategy through neural network inference. The control strategy includes specific operation instructions for the internal devices of the climate chamber, such as adjusting the heating power of the heating plate, the cooling intensity of the condenser, the brightness of the lighting plate, the rotation speed of the fan, the opening degree of the oxygen valve, and the opening degree of the carbon dioxide valve.
[0025] Subsequently, the control execution unit receives the control strategy generated by the reinforcement learning control unit, parses it to generate specific execution signals, and controls the actions of each environmental regulation device through the solid-state relay module to ensure that each environmental parameter is adjusted towards the target value corresponding to the crop growth stage.
[0026] After the actions of each device are completed, the data acquisition unit collects the updated environmental parameters again. The updated environmental parameters are transmitted to the reinforcement learning control unit through the communication unit. The reinforcement learning control unit updates the state vector based on the new round of environmental parameter inputs and conducts learning iterations through the policy update process to form a closed-loop feedback mechanism.
[0027] The communication unit includes a Modbus RTU bus communication interface, a TCP / IP communication interface, and a WebSocket protocol module, which are responsible for ensuring real-time data transmission and interaction between the data acquisition unit, the reinforcement learning control unit, and the control execution unit. A buffer queue management mechanism is set inside the communication unit to ensure data integrity and timeliness in case of high-concurrency data requests.
[0028] In this embodiment, the data acquisition unit and the reinforcement learning control unit communicate using the TCP / IP protocol, and the reinforcement learning control unit and the control execution unit communicate using the Modbus RTU protocol, thus ensuring low-latency transmission of control signals. At the same time, the reinforcement learning control unit pushes environmental parameters and control status to the host display system in real time through the WebSocket protocol to achieve visual monitoring and manual intervention. When generating a control strategy, the reinforcement learning control unit analyzes the target environmental range corresponding to the crop growth stage in real time, adjusts the action amplitude according to the current deviation degree, and preferentially selects a low-energy consumption action path during the action adjustment process. In case of sudden environmental interference, the reinforcement learning control unit immediately re-evaluates the state and dynamically generates an emergency control strategy to ensure that the environment quickly returns to the target range.
[0029] With the above technical solutions, the artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm can achieve dynamic adaptive adjustment of the internal environment of the climate chamber. By comprehensively considering the reward function of environmental parameter proximity, energy consumption minimization, and response speed optimization, the reinforcement learning control unit can continuously optimize the control strategy, effectively improving the accuracy of environmental regulation and the energy efficiency utilization rate. At the same time, through the closed-loop feedback mechanism and the efficient data interaction of the communication unit, the highly automated operation and stability of the system are realized, greatly reducing the need for manual intervention, enhancing the adaptability of the climate chamber to complex environmental changes, ensuring that the crops obtain the best growth environment conditions at different growth stages, and improving the crop growth quality and experimental controllability.
[0030] In an embodiment of the present application, the data acquisition unit includes a temperature sensor, a humidity sensor, a light sensor, a carbon dioxide concentration sensor, and an oxygen concentration sensor. The temperature sensor, humidity sensor, light sensor, carbon dioxide concentration sensor, and oxygen concentration sensor form a sensor group, and the sensor group is provided in each hierarchical structure of the cultivation area inside the climate chamber.
[0031] Specifically, the data acquisition unit includes a temperature sensor, a humidity sensor, a light sensor, a carbon dioxide concentration sensor, and an oxygen concentration sensor. Among them, the temperature sensor, humidity sensor, light sensor, carbon dioxide concentration sensor, and oxygen concentration sensor form a sensor group through a unified data interface. Each sensor group is connected to the communication unit via a 485 bus, and the communication unit transmits the collected data to the reinforcement learning control unit.
[0032] There are multiple hierarchical cultivation structures inside the climate chamber, and a sensor group is fixedly installed inside each cultivation layer area. The specific installation method of each sensor group is as follows: The temperature sensor and the humidity sensor are respectively installed at the central position of each cultivation area to collect the temperature and humidity of the microenvironment around the crops inside the layer in real time.
[0033] The light sensor is installed in the middle area below the lamp panel above the cultivation area to monitor the actual light intensity received by the crops.
[0034] The carbon dioxide concentration sensor and the oxygen concentration sensor are installed near the air holes on the side plates of the cultivation area to detect the concentration of air components in the cultivation area.
[0035] Each sensor group is connected to the communication interface module in the cabinet area of the climate chamber through a shielded wire. The communication interface module uniformly transmits the signals to the gateway and then forwards them to the reinforcement learning control unit.
[0036] The data collected by each sensor group has a timestamp mark. The communication unit is equipped with a priority management mechanism to ensure the priority transmission of abnormal environment data. The communication process uses CRC check to ensure data integrity. By setting sensor groups in the hierarchical structures of each cultivation area in the climate chamber, accurate monitoring of the microenvironments of crops in different layers and different areas can be achieved. After the multi-layer environment data collected by the data acquisition unit is analyzed and processed by the reinforcement learning control unit, customized environment adjustment strategies can be formed, thereby improving the locality and adaptability of the climate chamber environment control.
[0037] Adopting the above technical solution, the data acquisition unit can realize the real-time acquisition of the microenvironment parameters of each layer inside the climate chamber by setting sensor groups inside each cultivation layer of the climate chamber, ensuring that the data received by the reinforcement learning control unit covers the environmental changes in different growth areas inside the climate chamber. The reinforcement learning control unit generates more reasonable control strategies based on more detailed data, improving the accuracy of environmental control and the consistency of the crop growth environment.
[0038] In an embodiment of the present application, the control execution unit includes a heating device, a refrigeration device, a light adjustment device, a fan, an oxygen valve, and a carbon dioxide valve. The heating device and the refrigeration device achieve uniform diffusion of hot and cold air through a circulation fan, and the oxygen valve and the carbon dioxide valve adjust the gas concentration through the air holes on the side plate.
[0039] Specifically, the control execution unit provided in this embodiment specifically includes a heating device, a refrigeration device, a light adjustment device, a fan, an oxygen valve, and a carbon dioxide valve. Among them, the heating device uses a high-power heating resistance wire, which is installed at the heating layer position above the refrigeration device. The heating resistance wire is controlled by a solid-state relay module for on / off and heating power adjustment.
[0040] The refrigeration device uses a compression refrigeration system, which is combined with a condensation fan for heat dissipation. The compressor is installed at the bottom of the climate chamber and is connected to the internal refrigeration area of the climate chamber through a condensation pipeline. The refrigeration compressor is controlled by a solid-state relay for on / off to achieve precise temperature control.
[0041] The circulation fan is installed above the heating device and the refrigeration device. The circulation fan adopts a variable-frequency control method and can adjust the rotation speed according to needs. The working principle of the circulation fan is to inhale the hot air generated by the heating device or the cold air generated by the refrigeration device and evenly diffuse it along the inside of the climate chamber through a porous back plate, realizing rapid and uniform distribution of the temperature inside the box and avoiding the influence of local temperature difference on crop growth.
[0042] The light adjustment device includes two detachable lamp boards installed at the top of the cultivation area. Each lamp board is equipped with an LED light source with adjustable brightness and spectral type. The light adjustment device is connected to a solid-state relay through a PWM dimming controller and can dynamically adjust the light intensity and spectral ratio according to the light strategy output by the reinforcement learning control unit to ensure the light requirements of crops at different growth stages.
[0043] The oxygen valve and the carbon dioxide valve are respectively installed at the corresponding positions of the air holes on the side plate of the climate chamber. Both the oxygen valve and the carbon dioxide valve adopt a proportional electric control valve structure and precisely control the gas flow by adjusting the opening and closing angles. The oxygen valve and the carbon dioxide valve receive instructions from the control execution unit and adjust the intake air volume in real time according to the oxygen concentration and carbon dioxide concentration inside the climate chamber, so that the gas concentration inside the climate chamber remains within the optimal range for crop growth. The oxygen valve and the carbon dioxide valve are connected to gas cylinders through independent pipelines, and the gas cylinders are equipped with pressure reducing valves to ensure stable gas supply; each part of the control execution unit is connected to the reinforcement learning control unit in real time through a communication unit. After receiving the control strategy generated by the reinforcement learning control unit, the communication unit analyzes the instructions and sequentially drives the heating device, the refrigeration device, the light adjustment device, the fan, the oxygen valve, and the carbon dioxide valve to act as required. After the action of the control execution unit is completed, the execution status is transmitted back through the communication unit for the reinforcement learning control unit to update the feedback information, ensuring the formation of a complete closed-loop control process.
[0044] With the above technical solution, the control execution unit realizes the uniform diffusion of hot and cold air inside the climate chamber through the cooperation of the heating device, the refrigeration device and the circulation fan, effectively avoiding the problem of uneven crop growth caused by local environmental fluctuations. At the same time, through the structure of the oxygen valve and the carbon dioxide valve combined with the air holes on the side plate of the climate chamber, the precise control of the gas concentration inside the climate chamber is realized, further ensuring the dynamic stability of the gas components in the crop growth environment.
[0045] In an embodiment of the present application, the reward function of the reinforcement learning control unit is:
[0046] wherein, represents the immediate reward value obtained after executing the action in the state , represents the current environmental state, including real-time parameters such as temperature, humidity, light, and gas concentration; represents the control action; represents the environmental parameter deviation weight coefficient; is the current environmental parameter; is the ideal value corresponding to the crop growth stage; represents the number of types of monitored environmental parameters; represents the energy consumption weight coefficient; E ( a ) is the energy consumption of the action a ; represents the response speed weight coefficient; δ ( s ) is the response speed reward.
[0047] Specifically, the reward function of the reinforcement learning control unit in this embodiment is set as:
[0048] wherein, represents the immediate reward value obtained after executing the control action in the current environmental state . The current environmental state s includes a set of parameters such as the temperature, humidity, light intensity, carbon dioxide concentration, and oxygen concentration inside the climate chamber collected in real time by the data acquisition unit. The control action represents the combination of environmental adjustment instructions generated by the reinforcement learning control unit according to the current state , including the adjustment of the heating power of the heating device, the adjustment of the refrigeration intensity of the refrigeration device, the adjustment of the brightness and spectrum of the light adjustment device, the change of the fan speed, the change of the oxygen valve opening, and the change of the carbon dioxide valve opening; in the reward function is the environmental parameter deviation weight coefficient, which is used to measure the importance of the deviation between the environmental parameter and the ideal target value. It is the current real-time monitored environmental values. is the ideal environmental parameter value set for the corresponding crop growth stage. The number of environmental parameters monitored includes five categories: temperature, humidity, light intensity, carbon dioxide concentration, and oxygen concentration. It is a penalty item for environmental parameter deviation. By finding the absolute deviation between each environmental parameter and the corresponding ideal value and taking a weighted sum, a negative reward corresponding to the degree of deviation is given to encourage the environment to quickly approach the target value.
[0049] is the energy consumption weight coefficient, used to evaluate the execution action The energy consumption caused by To control the action The corresponding energy consumption data includes the weighted sum of the energy consumption of the heating device, cooling device, light control device, fan, oxygen valve and carbon dioxide valve. It is an energy consumption penalty term, which encourages the reinforcement learning control unit to generate a lower energy consumption adjustment strategy by penalizing high energy consumption actions; is the response speed weight coefficient, which is used to measure the rate at which the environmental state converges to the target value. Reward for speed of response, It is defined as the degree of improvement of the environmental parameters approaching the target value per unit time. The third term The response speed reward item is used to reward control actions that quickly improve the environmental state, thereby improving the system's ability to respond to sudden environmental changes. During the training process, the reinforcement learning control unit obtains an immediate reward value after each control action is executed. To adjust the weight parameters of the internal strategy network and continuously optimize the control strategy, the internal environment of the climate chamber can minimize energy consumption and optimize response speed while meeting the needs of crop growth.
[0050] By adopting the above technical solution, the reinforcement learning control unit can realize the precise optimization of the climate chamber control strategy under different environmental changes by constructing a reward function that comprehensively considers environmental proximity, energy consumption and response speed, so that the system can effectively reduce energy consumption and improve the rapid response capability of environmental control while maintaining an ideal growth environment.
[0051] In one embodiment of the present application, the communication unit includes an Internet of Things gateway and an industrial computer. The Internet of Things gateway is connected to each sensor through the Modbus protocol and communicates with the industrial computer through the TCP protocol to achieve transmission of sensor data and control instructions.
[0052] Specifically, in this embodiment, the communication unit specifically includes an IoT gateway and an industrial control computer. The IoT gateway is connected to each sensor in a wired manner using the Modbus RTU protocol. The IoT gateway integrates multiple RS485 interfaces, which respectively correspond to the temperature sensors, humidity sensors, light sensors, carbon dioxide concentration sensors, and oxygen concentration sensors distributed on each layer inside the climate chamber. The IoT gateway collects real-time environmental parameter data of each sensor at a set time interval through a polling mechanism; the IoT gateway performs preliminary formatting processing on the collected sensor data, including adding timestamps, sensor identification codes, and data integrity verification, and then encapsulates the processed environmental parameter data into TCP data packets through the built-in TCP protocol stack and sends them to the industrial control computer.
[0053] The industrial control computer serves as the upper-level data processing center, receives the environmental parameter data transmitted from the IoT gateway, performs data parsing, standardization processing, and storage, and at the same time transmits the standardized environmental status data to the reinforcement learning control unit for subsequent control strategy generation; during the control instruction issuance process, the reinforcement learning control unit sends the generated control strategy to the industrial control computer in the format of a control instruction through the TCP protocol. After the industrial control computer parses the control instruction, it sends it to the IoT gateway again through the TCP protocol. The IoT gateway re-encodes the received control instruction according to the Modbus RTU protocol and issues it to the corresponding control devices, such as the execution modules of heating devices, refrigeration devices, light adjustment devices, fans, oxygen valves, and carbon dioxide valves, to ensure that the environmental parameters are adjusted in real time according to the control strategy generated by the reinforcement learning control unit.
[0054] To improve the stability and fault tolerance of communication, a buffer area and a disconnection reconnection mechanism are provided inside the IoT gateway. When network communication is abnormal, data caching and reissuing can be realized to ensure the continuity and reliability of data collection and instruction transmission. The industrial control computer is configured with a dual network interface structure to support redundant communication path design, further improving the stability of system operation and data security.
[0055] With the above technical solution, through the cooperation of the IoT gateway and the industrial control computer, and combining the Modbus protocol and the TCP protocol, the communication unit realizes the efficient collection of sensor data, standardized transmission, and accurate issuance of control instructions, ensuring the real-time data interaction and instruction execution synchronization between the modules inside the climate chamber, and greatly improving the communication reliability of the system.
[0056] In an embodiment of the present application, the reinforcement learning control unit further includes: A crop growth stage recognition module, which is used to judge the current growth stage of the crop according to the historical data of environmental parameters and a preset crop growth model, and dynamically adjust the ideal value in the reward function and the weight coefficient , , ; Among them, the growth stage recognition module collaborates with the data acquisition unit and the user interface unit. By matching the real-time leaf age, plant height, and environmental accumulation through a crop growth model, it determines the stage switching threshold and triggers the phased adaptive update of the control strategy.
[0057] Specifically, in this embodiment, the reinforcement learning control unit further includes a crop growth stage recognition module. The crop growth stage recognition module works in cooperation with the data acquisition unit and the user interface unit. First, it obtains the environmental parameter data of each cultivation area in the climate chamber in real time through the data acquisition unit, and conducts a comparative analysis in combination with the preset crop growth model data in the user interface unit. The crop growth model includes parameter information such as the typical environmental requirements characteristics, accumulated temperature, accumulated light amount, humidity change curve, and carbon dioxide concentration requirements corresponding to different growth stages of the crop. The crop growth stage recognition module calculates real-time environmental accumulation indicators, such as cumulative leaf age, cumulative plant height, cumulative light radiation amount, cumulative temperature time integral, etc., and matches them with the standard accumulation values in the growth model.
[0058] During the matching process, the crop growth stage recognition module sets a stage switching threshold. When the real-time accumulation reaches or exceeds the set stage switching threshold, the crop growth stage recognition module determines that the crop growth stage has switched, triggering the reward function parameter update process, which specifically includes dynamically adjusting the ideal value in the reward function , setting the target values of each environmental parameter to the standard environmental values corresponding to the new growth stage, and at the same time appropriately adjusting the weight coefficients according to the different requirements of the new stage for environmental stability, energy consumption control, and response speed , , , for example, increasing the weight of environmental parameter proximity in the seedling stage , increasing the weight of energy consumption optimization in the growth stage , increasing the weight of response speed in the flowering and fruiting stage , ensuring that the control strategy generated by the reinforcement learning control unit better meets the actual needs of crops in different growth stages; the crop growth stage recognition module can also receive the manual correction signal input by the user through the user interface unit. When the user inputs the measured leaf age or plant height data through the interface, the crop growth stage recognition module can correct the growth stage determination according to the manual data to improve the accuracy and flexibility of recognition.
[0059] After identifying the new growth stage, the crop growth stage recognition module sends the updated , , , Parameters are passed to the reinforcement learning control unit, and the reinforcement learning control unit regenerates the control strategy according to the new reward function parameters, realizing the phased adaptive update of the control strategy, so as to ensure that the internal environment of the climate chamber is always synchronously matched with the actual growth needs of the crops.
[0060] By adopting the above technical solution, the reinforcement learning control unit realizes the dynamic judgment of the crop growth stage and the adaptive adjustment of the reward function parameters by introducing the crop growth stage recognition module and combining the historical environmental parameter data with the preset crop growth model, enabling the system to optimize the environmental control objectives and control strategies in real time according to the actual growth process of the crops.
[0061] In an embodiment of the present application, the heating device is a resistance heating wire, the refrigeration device is a compressor and a condenser fan, the heating wire and the waste heat of the compressor are reused, and the environmental parameters are uniformly adjusted through a porous back panel and a circulation fan.
[0062] Specifically, in this embodiment, the heating device in the control execution unit is a resistance heating wire, which is made of a high-efficiency alloy material and has the performance of rapid heating and high-temperature aging resistance. The resistance heating wire is installed in a partition area above the refrigeration compressor inside the climate chamber. The refrigeration device includes a refrigeration compressor and a condenser fan. The refrigeration compressor adopts a variable-frequency compressor structure to reduce the temperature inside the climate chamber through refrigerant circulation. The condenser fan is used to assist the compressor in heat dissipation, improving the refrigeration efficiency and stability of the compressor.
[0063] During the operation of the system, when the refrigeration compressor completes the cooling, the heat generated by the long-term operation of the compressor body is conducted to the upper resistance heating wire area through physical contact. The resistance heating wire uses the waste heat of the compressor for preheating, so as to reduce the start-up energy consumption and heating response time when the heating mode needs to be started, realizing the secondary utilization of energy and improving the overall energy use efficiency of the climate chamber.
[0064] A porous back panel is arranged inside the climate chamber. The porous back panel is distributed at the intersection positions of the refrigeration area, the heating area and the air duct. The circulation fan is installed at the center position above the refrigeration compressor and the resistance heating wire. The circulation fan adopts variable-frequency control and can adjust the air volume according to the real-time environmental parameters. When the circulation fan operates, it inhales the cold air output by the refrigeration compressor or the hot air generated by the resistance heating wire into the air duct, and disperses it to each cultivation area of the climate chamber through the uniform pores on the porous back panel, so that the cold and hot air can be quickly and evenly diffused in the space, realizing the uniform adjustment of the temperature, humidity and gas components inside the climate chamber, and avoiding the impact of local cold and heat unevenness or humidity fluctuation on the growth of crops.
[0065] The resistance heating wire, the refrigeration compressor, the condensing fan, and the circulation fan are all connected to the solid-state relay control module of the control execution unit, and receive specific control instructions sent by the reinforcement learning control unit through the communication unit, so as to realize the dynamic intelligent control of the heating, refrigeration, and air circulation processes, and ensure that the internal environmental parameters of the climate chamber are stably maintained within the ideal range according to the requirements of the crop growth stage.
[0066] With the above technical solution, the heating device uses a resistance heating wire and combines with the waste heat reuse design of the compressor of the refrigeration device, which can effectively reduce the heating start-up energy consumption and shorten the heating-up response time, improve the energy utilization rate. The refrigeration device realizes efficient and stable cooling through the cooperation of the compressor and the condensing fan. The synergistic effect of the porous back panel and the circulation fan enables the internal environmental parameters of the climate chamber to spread quickly and evenly in space, avoiding the adverse impact on crop growth caused by local environmental fluctuations.
[0067] In an embodiment of the present application, a user interface unit is further included. The user interface unit provides functions of visualizing real-time environmental parameters, a manual control interface, and abnormal alarm, and supports the export of historical data and the configuration of comparative experiments.
[0068] Specifically, in this embodiment, the system further includes a user interface unit. The user interface unit is installed outside the vertical cabinet area of the climate chamber and realizes functional interaction by connecting to an industrial control computer. The user interface unit is configured with a high-definition display screen and a touch control module, and provides the function of visualizing real-time environmental parameters. The function of visualizing real-time environmental parameters receives the environmental data processed by the reinforcement learning control unit through the communication unit, and dynamically displays the temperature, humidity, light intensity, carbon dioxide concentration, and oxygen concentration in each cultivation area inside the climate chamber on the display screen in the form of charts, curves, and numerical values. The user interface unit is provided with a manual control interface. The manual control interface provides separate control switches and parameter adjustment sliders for the heating device, refrigeration device, light adjustment device, fan, oxygen valve, and carbon dioxide valve through a graphical operation panel. Users can temporarily adjust the working states or parameters of each device according to actual needs. The manual operation instructions are directly sent to the control execution unit through the communication unit to achieve immediate control. At the same time, the user interface unit is provided with a permission management system to ensure that the manual operation permission is limited by authorized users.
[0069] The user interface unit further includes an abnormal alarm function. The abnormal alarm function sets warning upper and lower limit thresholds for real-time environmental parameters. When it is detected that any environmental parameter exceeds the safe range, or communication failure, equipment abnormality, etc., it immediately notifies the user by means of pop-up windows, sound and light alarms, and log records, prompting the user to intervene and handle in time to ensure the safety of crop cultivation.
[0070] The user interface unit supports the function of exporting historical data. The historical data export function can export environmental parameter records by day, week, month or custom time period, and the data format supports common formats such as CSV and Excel, which is convenient for users to perform local archiving and subsequent analysis.
[0071] The user interface unit also provides the function of configuring comparative experiments. The comparative experiment configuration function allows users to set different groups of environmental target parameters in different layers of the cultivation area, set the differences in environmental conditions between the experimental group and the control group through the interface, and monitor and record the data change process of the two groups in real time for scientific experiments and data analysis; all interactive operations of the user interface unit are synchronized with the industrial control computer and the reinforcement learning control unit through the internal system bus to ensure that the user interface operations are consistent with the actual operating status of the climate chamber, avoiding data lag or operation failure.
[0072] By adopting the above technical solution, the user interface unit can effectively improve the operability and safety of the system by providing real-time visualization of environmental parameters, a manual control interface and an abnormal alarm function. Users can intuitively understand the internal environmental changes of the climate chamber and make intervention adjustments according to needs. The abnormal alarm mechanism ensures timely risk warning during the crop cultivation process. At the same time, the historical data export and comparative experiment configuration functions provide convenient data support and experimental basis for scientific research and environmental parameter optimization.
[0073] In an embodiment of the present application, the IoT gateway controls the power on and off of each device in the control execution unit through a solid-state relay.
[0074] Specifically, in this embodiment, the power on and off control between the IoT gateway and each device of the control execution unit is realized through a solid-state relay. The IoT gateway is internally provided with multiple digital output interfaces, and each digital output interface is connected to the input end of the corresponding solid-state relay. The output ends of the solid-state relays are respectively connected to the power supply circuits of the heating device, the refrigeration device, the light adjustment device, the fan, the oxygen valve and the carbon dioxide valve. After receiving the control instruction sent by the reinforcement learning control unit or the industrial control computer, the IoT gateway analyzes the instruction content and switches the high and low levels of the corresponding digital output interface according to the instruction requirements. When the IoT gateway outputs a high-level signal, the corresponding solid-state relay is activated to close the circuit and start the corresponding control execution unit device. When the IoT gateway outputs a low-level signal, the corresponding solid-state relay circuit is disconnected to turn off the corresponding control execution unit device, thereby realizing precise management of the power on and off of each control device.
[0075] To ensure the reliability of control signal transmission and the timeliness of execution actions, the digital output control of the IoT gateway adopts optoelectronic isolation design, effectively preventing misoperations caused by external electrical interference. At the same time, a status feedback module is set inside the IoT gateway to detect the output status of the solid-state relay in real time and feedback it to the industrial control computer and the reinforcement learning control unit through the communication unit to ensure the correct closed-loop confirmation of the execution action.
[0076] Adopting the above technical solution, the IoT gateway realizes the control of the power on and off of each device of the control execution unit through the solid-state relay, which not only greatly simplifies the control circuit structure of the climate chamber, improves the system response speed and stability, but also enhances the safety and lifespan of the overall system through the high reliability and protection function of the solid-state relay.
[0077] In an embodiment of the present application, the training process of the reinforcement learning control unit includes: Initializing model parameters based on historical data or simulation environment; Generating an adaptive control strategy through environmental state acquisition, action generation, reward calculation, and policy iteration optimization; When the reward value converges or reaches the preset number of iterations, output the optimal control strategy.
[0078] Specifically, the training process of the reinforcement learning control unit in this embodiment includes the following steps. First, initialize the model parameters based on historical data or simulation environment. The historical data consists of environmental parameters, control actions, and corresponding environmental change results recorded during the previous operation of the climate chamber. The simulation environment is constructed by a preset physical model of the climate chamber and a crop growth model. The reinforcement learning control unit performs pre-training according to historical data or simulation environment, initializes the core training hyperparameters such as the weight parameters, learning rate, and discount factor of the neural network, and sets the reward function:
[0079] As the training optimization goal; during the training process, the reinforcement learning control unit first obtains the current environmental state through the environmental state acquisition module , the environmental state includes normalized real-time parameters such as temperature, humidity, light intensity, carbon dioxide concentration, and oxygen concentration, and then generates a control action based on the current policy network , the control action corresponds to the specific control instructions of the heating device, refrigeration device, light adjustment device, fan, oxygen valve, and carbon dioxide valve; after the control action is executed in the simulation environment or the historical environment response model, the system calculates the new environmental state according to the execution result and calculates the immediate reward value according to the reward function , the reward value comprehensively considers the environmental parameter proximity, action energy consumption, and response speed. The reinforcement learning control unit forms training samples by combining the current state, action, reward value, and new state, and stores them in the experience replay pool. By sampling the training samples in the experience pool, the gradient of the policy network and value network is updated, and the reinforcement learning optimization algorithm performs iterative training to continuously adjust the network parameters to improve the policy performance.
[0080] The training process continues. When the cumulative reward value reaches the preset convergence criterion or the number of training rounds reaches the preset number of iterations, the reinforcement learning control unit determines that the training is completed and outputs the finally converged optimal control strategy. The optimal control strategy has the ability to adaptively generate control instructions according to the actual environmental state of the climate chamber, and can dynamically adjust the internal temperature, humidity, light, and gas concentration of the climate chamber in real time to meet the various requirements of the crop growth stage.
[0081] Adopting the above technical solution, the reinforcement learning control unit initializes the model parameters based on historical data or simulation environment, and forms an adaptive control strategy through environmental state acquisition, action generation, reward calculation, and policy iterative optimization. After the reward converges or reaches the preset number of iterations, the optimal control strategy is output. It can not only effectively improve the training efficiency and policy reliability of the climate chamber intelligent control system, but also ensure that the finally generated control strategy has good environmental adaptability.
[0082] The structure of the climate chamber of the present invention is as Figure 2 shown, and includes a vertical cabinet area 10, a groove area, an internal assembly area, and a cultivation area 12 respectively. Each area is connected to each other through wire routing to ensure the stable operation of the climate chamber. Further explanation, inside the vertical cabinet area 10, each sensor power supply unit, industrial computer, display screen 20, and Internet of Things gateway are installed and fixed, and it undertakes the wire routing task related to each control device.
[0083] The positive and negative power lines of each sensor are integrally supplied power by a power plug after integration, and the 485 communication AB lines used for data transmission are also integrated and sorted through a trough-type guide rail wiring box. All A lines are uniformly connected to port A of the Internet of Things gateway, and all B lines are connected to port B of the Internet of Things gateway. The overall wiring is neat, which is convenient for subsequent maintenance.
[0084] The power cord of the industrial computer connects the power plug to the groove area through a wiring box for independent power supply, and the network cable is directly connected to the Internet of Things gateway. The separated wiring method provides convenience for troubleshooting in the later stage. The power supply method of the display screen 20 is the same as that of the industrial computer, adopting an independent power supply method, and a normally open design is adopted to prevent the screen from being damaged due to frequent start and stop.
[0085] The gateway is clearly distinguished from the industrial control computer, the display screen 20, and the sensors, and is independently powered to ensure the stable operation of each control device. The multi-channel control line outputs are connected to the groove area through a wiring box and are connected to the live wire and neutral wire of each electrical appliance.
[0086] At the lower part inside the vertical cabinet area 10, there is a water storage bucket fixed by an iron frame, which is used to provide a stable water source for the humidifier to ensure the normal operation of the humidification function.
[0087] Further explanation, the groove area is mainly used to connect the live wire, neutral wire, ground wire of each electrical appliance and the control output DO line, and at the same time connect the solid-state relay protection lines required by each electrical appliance to realize the control of the electrical appliances by the gateway. By labeling the live, neutral, and ground wires of each electrical appliance, first connect them to the output end of the solid-state relay, and then connect the control output DO line led out from the vertical cabinet area 10 to the input end of the relay. After the circuit modification is completed, the opening and closing states of the DO port can be controlled through the Internet of Things gateway to realize the remote and intelligent control of each electrical appliance in the box.
[0088] The internal assembly area is mainly responsible for the installation of each electrical appliance to ensure the realization of the function of the climate chamber and its smooth operation. The specific installation sequence is as follows: (1) Refrigeration compressor and condenser fan: The climate chamber is equipped with two refrigeration compressors for cooling the box body. Since the power of the compressor is relatively large, it needs to be cooled in cooperation with the condenser fan during operation to ensure its long-term stable operation.
[0089] (2) Heating resistance wire: There are two groups of heating resistance wires for heating the box body, which are installed above the refrigeration compressor. After refrigeration is completed, it can be quickly switched to the heating mode to realize the constant control of temperature. This layout can also use the electrical waste heat of the compressor to preheat the resistance wire, thus saving energy and reducing consumption.
[0090] (3) Circulation fan: The circulation fan is installed above the refrigeration compressor and the heating resistance wire and is close to the bottom of the cultivation area 12 of the box body. Its function is to push the air from bottom to top, and at the same time part of the air flow is blown out through the small holes on the back panel to realize the rapid and uniform diffusion of cold air and hot air and keep the air in the box well circulated. (4) Humidifier: The humidifier is installed behind the porous back panel inside the box and runs through the entire back panel. Its water source is connected to the water storage tank in the right vertical cabinet through a water pipe. The coordinated operation of the fan can enhance the water mist diffusion effect and improve the humidification efficiency.
[0091] The cultivation area 12 is the planting area for experimental crops. This area can divide the experimental space by adding iron plates to realize comparative experiments. The specific composition is as follows: (1)Lamp board: There are two detachable lamp boards at the top of the cultivation area 12, and different light sources can be switched according to experimental needs. A spacing is reserved between the two lamp boards, and a black cloth partition can be added to form left and right control groups. Each iron plate is used to place experimental crops, and a lamp board can be installed below the iron plate to provide light for the next layer.
[0092] (2)Side plate hook: The hooks on the side plates are not only used to fix the iron plates to expand the crop planting area, but also can hang the extension probes of the sensors for environmental monitoring.
[0093] (3)Side plate air holes: The side plates are designed with air holes to facilitate the ventilation adjustment of carbon dioxide and oxygen to maintain the required gas concentration.
[0094] (4)Porous back plate: The porous back plate is used for the blower to send air, so that the air generated by humidification, refrigeration and heating is evenly diffused throughout the interior of the box. At the same time, the back plate channels also promote gas circulation, and the original gas can be discharged to adjust the environment of the box.
[0095] (5)Glass door: The opening glass of the climate chamber is pasted with a light-adjusting film, and the reflective or light-transmitting performance of the glass can be adjusted according to needs, which is convenient for observing the interior and can avoid the interference of external light on the experimental environment when observation is not required.
[0096] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural transformation made under the inventive concept of the present invention by using the content of the specification and drawings of the present invention, or directly / indirectly applied in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. An artificial intelligence climate chamber adaptive control system based on a reinforcement learning algorithm, characterized in that, Comprising: A data acquisition unit for real-time acquisition of multiple environmental parameters and crop growth stages inside the climate chamber; A reinforcement learning control unit that autonomously learns and generates a control strategy based on the environmental parameters, crop growth stages, real-time environmental changes, and resource consumption status through a reward function, where the reward function comprehensively optimizes the environmental parameter proximity, energy consumption, and response speed; A control execution unit for adjusting the environmental parameters inside the climate chamber according to the control strategy; A communication unit for realizing data transmission and interaction among the units; Wherein, the data acquisition unit, the reinforcement learning control unit, and the control execution unit form a closed-loop feedback through the communication unit. The reinforcement learning control unit generates a control strategy that minimizes energy consumption and meets the response speed based on the dynamic changes of environmental parameters and the target values of crop growth stages, and realizes the adaptive adjustment of the climate chamber environment through the control execution unit.
2. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 1, wherein The data acquisition unit includes a temperature sensor, a humidity sensor, a light sensor, a carbon dioxide concentration sensor, and an oxygen concentration sensor. The temperature sensor, humidity sensor, light sensor, carbon dioxide concentration sensor, and oxygen concentration sensor constitute a sensor group, and a sensor group is provided in each hierarchical structure of the cultivation area inside the climate chamber.
3. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 1, wherein The control execution unit includes a heating device, a refrigeration device, a light adjustment device, a fan, an oxygen valve, and a carbon dioxide valve. The heating device and the refrigeration device achieve uniform diffusion of hot and cold air through a circulation fan, and the oxygen valve and the carbon dioxide valve adjust the gas concentration through side plate air holes.
4. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 1, characterized in that, The reward function of the reinforcement learning control unit is: Among them, represents the immediate reward value obtained after performing the action in state . represents the current environmental state, including real-time parameters such as temperature, humidity, light, gas concentration, etc.; represents the control action; represents the environmental parameter deviation weight coefficient; is the current environmental parameter; is the ideal value corresponding to the crop growth stage; represents the number of types of monitored environmental parameters; represents the energy consumption weight coefficient; E ( a ) is the energy consumption of the action a . represents the response speed weight coefficient; δ ( s ) is the response speed reward.
5. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 2, wherein, The communication unit includes an Internet of Things gateway and an industrial computer. The Internet of Things gateway is connected to each sensor through the Modbus protocol and communicates with the industrial computer through the TCP protocol to realize the transmission of sensor data and control instructions.
6. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 4, wherein The reinforcement learning control unit further includes: A crop growth stage recognition module, which is used to determine the current growth stage of the crop according to the historical data of environmental parameters and a preset crop growth model, and dynamically adjust the ideal value in the reward function and the weight coefficient , , ; Wherein, the growth stage recognition module collaborates with the data acquisition unit and the user interface unit to determine the stage switching threshold by matching the real-time leaf age, plant height, and environmental accumulation through a crop growth model, and triggers the phased adaptive update of the control strategy.
7. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 3, characterized in that, The heating device is a resistance heating wire, the refrigeration device is a compressor and a condenser fan, the heating wire and the compressor waste heat are reused, and the environmental parameters are uniformly adjusted through a porous backboard and a circulation fan.
8. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 1, characterized in that It further includes a user interface unit that provides real-time environmental parameter visualization, a manual control interface, and an abnormal alarm function, and supports historical data export and comparative experiment configuration.
9. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 5, characterized in that, The Internet of Things gateway controls the power on and off of each device in the control execution unit through a solid-state relay.
10. The artificial intelligence climate chamber adaptive control system based on the reinforcement learning algorithm according to claim 1, characterized in that, The training process of the reinforcement learning control unit includes: Initializing model parameters based on historical data or a simulation environment; Generating an adaptive control strategy through environmental state acquisition, action generation, reward calculation, and policy iterative optimization; When the reward value converges or reaches a preset number of iterations, outputting the optimal control strategy.
Citation Information
Patent Citations
Greenhouse environment intelligent control method based on extreme learning machine network
CN108319134A
Agricultural greenhouse environment adjusting method, device, equipment and medium
CN116910490A
Irrigation decision-making method based on agricultural system model
CN119313095A
Intelligent temperature control system for greenhouse
CN119597075A
Intelligent LED spectrum regulation and control system based on multi-dimensional data fusion
CN119653555A
Cited By
Northern greenhouse tropical fruit microclimate adaptive control system based on reinforcement learning
CN121325625A
Reinforcement Learning-Based Adaptive Control System for Microclimate of Tropical Fruits in Northern Greenhouses
CN121325625B