Mushroom house environment control system and method based on reinforcement learning adaptive PID
By using a mushroom house environmental control system based on reinforcement learning adaptive PID, the problems of environmental parameter fluctuations and equipment wear in mushroom houses were solved, achieving precise control, extending equipment life, and improving system stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
- Filing Date
- 2026-04-16
- Publication Date
- 2026-07-24
Smart Images

Figure CN122018346B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural intelligent control technology, and in particular to a mushroom house environmental control system and method based on reinforcement learning adaptive PID. Background Technology
[0002] The mushroom cultivation industry is developing towards intensification and factory-style production, and the precision of environmental control directly determines the yield and quality of mushrooms. Because mushrooms have varying requirements for parameters such as temperature, humidity, and carbon dioxide concentration at different growth stages, traditional environmental control systems often use threshold switching control, leading to drastic fluctuations in environmental parameters and frequent equipment start-ups and shutdowns. Although some systems have introduced traditional PID control, the fixed parameters make it difficult to adapt to the large lag, strong coupling, and nonlinear characteristics of mushroom houses. End-to-end control schemes based on reinforcement learning, which have emerged in recent years, are prone to motion oscillations, causing mechanical damage to actuators and shortening equipment lifespan. Furthermore, the sensor placement and actuator layout of existing systems often lack scientific rigor, resulting in uneven distribution of the indoor environmental field.
[0003] Therefore, designing a mushroom house environmental control system that can adapt to complex environmental changes to achieve precise control, effectively suppress motion vibrations to extend equipment life, and incorporate a scientific hardware layout has become an urgent technical challenge. Summary of the Invention
[0004] The main objective of this invention is to provide a mushroom house environmental control system and method based on reinforcement learning adaptive PID. The aim is to design a mushroom house environmental control system that can adapt to complex environmental changes to achieve precise control, effectively suppress motion oscillations to extend equipment life, and is equipped with a scientific hardware layout.
[0005] To achieve the above objectives, this invention proposes a mushroom house environmental control system based on reinforcement learning adaptive PID, comprising: The environmental sensing module is used to collect environmental parameters inside the mushroom house in real time; An execution control module is used to respond to control quantities to adjust the environment inside the mushroom house; and The central control module is connected to both the environmental perception module and the execution control module. The central control module includes a cascaded reinforcement learning model and a PID controller; The reinforcement learning model is configured to output a reinforcement learning policy based on the current state vector, wherein the state vector represents the deviation state between the environmental parameters and the target value; The PID controller is configured to determine the PID control parameters for the current time step according to the reinforcement learning strategy, and to calculate the control quantity based on the PID control parameters to drive the execution control module. The reinforcement learning model is based on a time-smoothing regularization term. The time smoothing regularization term is trained to optimize the objective. Based on the local smoothing function in the continuous time domain of the control quantity The second-order Taylor expansion is constructed, and its expression is:
[0006] In the formula, , , These are the control variables for the current time step and the historical time step, respectively. The sampling period is and These are the weighting coefficients.
[0007] Preferably, the reinforcement learning model is based on a reward function. To optimize, its expression is:
[0008] In the formula, This represents the absolute deviation of environmental parameters from target values. The rate of change of deviation The total energy consumption for executing the control module, This represents the rate of change of the control quantity between two adjacent time steps. , , , These are the weighting coefficients.
[0009] Preferably, the specific value of the weighting coefficient is set as follows: , , , .
[0010] Preferably, the central control module further includes an uncertainty-aware adjustment module, and the reinforcement learning strategy consists of multiple parallel Actor networks. The uncertainty-aware adjustment module adjusts the strategy based on the statistical variance of the control output set of the multiple parallel Actor networks under the same state. Estimate the policy uncertainty in the current state, and scale the original control quantity based on the policy uncertainty. The scaled control quantity is... The expression is:
[0011] In the formula, To enhance the original control output of the learning strategy, This is the uncertainty adjustment coefficient.
[0012] Preferably, the final optimization objective of the reinforcement learning model is... for:
[0013] In the formula, To reinforce the learning of the basic objective function, For uncertainty regularization, For penalty weights.
[0014] Preferably, the state vector is formed by combining the normalized parameters, and its parameters include: the deviations of each environmental parameter from the target value. , the rate of change of deviation of each environmental parameter The PID control parameters at the current time step and the current operating status of each device in the execution and control module.
[0015] Preferably, the central control module is further configured to execute fault-tolerant control logic: monitor in real time the deviation between the actual operating state of the execution control module and the state of the control quantity command; when the deviation exceeds 15%, reissue the control quantity; when the deviation is detected to exceed 15% three times consecutively, switch to standby control mode.
[0016] Preferably, the mushroom house is equipped with five cultivation racks, which are arranged at equal intervals along the vertical direction; the environmental sensing module includes five sensor groups, which are respectively installed in the middle of each layer of the five cultivation racks; each sensor group includes a temperature sensor, a humidity sensor, a light sensor, a carbon dioxide concentration sensor, and an oxygen concentration sensor.
[0017] Preferably, the execution control module includes a heating system, which includes electric heating cables laid on the floor of the mushroom house, and the electric heating cables cover the floor of the mushroom house in an S-shaped wiring manner.
[0018] Preferably, the execution control module includes a humidification system, which includes ultrasonic spray heads installed at the bottom front end of each of the five cultivation racks, with the nozzle of each ultrasonic spray head tilted upward at 15° relative to the horizontal plane.
[0019] Preferably, the execution control module includes a ventilation system, which includes: an axial flow fan installed in the center of the mushroom house roof; and a first electric louver and a second electric louver, which are symmetrically arranged on the upper part of the side wall of the mushroom house; wherein the axial flow fan is linked and controlled with the first electric louver and the second electric louver.
[0020] Preferably, the execution control module includes a lighting system, which includes LED light strips installed on the upper edge of each of the five cultivation racks, wherein the ratio of red light to blue light spectrum of the LED light strips is 7:3.
[0021] Preferably, the execution control module includes a shading system, which includes a double-layered shading curtain installed on the inner side of the mushroom house roof, and the shading rate of the double-layered shading curtain is 95%.
[0022] This application also discloses a mushroom house environment control method based on reinforcement learning adaptive PID, applying the mushroom house environment control system based on reinforcement learning adaptive PID as described above, the method comprising the following steps: The environmental parameters inside the mushroom house are collected in real time through the environmental sensing module; The central control module outputs a reinforcement learning policy based on the current state vector through a reinforcement learning model; The central control module determines the PID control parameters for the current time step using the PID controller according to the reinforcement learning strategy, and calculates the control quantity based on the PID control parameters. The control module responds to the control quantity to adjust the environment inside the mushroom house.
[0023] The above technical solution has the following advantages: This invention employs a control architecture that cascades a reinforcement learning model with a PID controller in the central control module. By utilizing the real-time output parameter configuration strategy of the reinforcement learning model, it achieves adaptive and precise control of the complex environment of a mushroom-shaped control room. The reinforcement learning model is trained based on an optimization objective that includes a time-smoothing regularization term. This regularization term is constructed based on the second-order Taylor expansion of the control quantity. By jointly constraining the first and second derivatives of the control quantity, it effectively filters out fluctuations in high-frequency control signals without weakening the system's responsiveness. This mechanism suppresses the oscillations of the execution and control module from the algorithmic level, significantly reducing the mechanical wear of hardware equipment, extending the service life of core execution equipment, and improving the system's operational stability while ensuring adjustment accuracy. Attached Figure Description
[0024] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a schematic diagram of a mushroom house environmental control system based on reinforcement learning adaptive PID, provided as an embodiment of the present invention.
[0025] Figure 2 This is a schematic diagram of a mushroom house environment control method based on reinforcement learning adaptive PID provided in an embodiment of the present invention. Detailed Implementation
[0026] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited thereto. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention. In the following description, specific details such as particular system structures and technical details are set forth for illustration rather than limitation in order to provide a thorough understanding of the embodiments of the present invention. However, those skilled in the art will understand that the present invention can also be implemented in other embodiments without these specific details.
[0027] Example 1 This embodiment provides a mushroom house environmental control system based on reinforcement learning adaptive PID, the overall architecture of which is referenced. Figure 1 This system solves the problems of low precision, high energy consumption, and equipment damage caused by actuator vibration in mushroom house environmental control through deep integration of hardware-sensing layout and cascaded algorithm control. The system includes an environmental sensing module, an execution control module, and a central control module. The central control module is connected to both the environmental sensing module and the execution control module. Spatially, the mushroom house has five cultivation racks arranged at equal intervals along the vertical direction to maximize space utilization and achieve factory-style cultivation.
[0028] The environmental sensing module is used to collect environmental parameters inside the mushroom house in real time. To obtain accurate microenvironmental information, the environmental sensing module includes five sensor groups. These five sensor groups are installed in the middle of each of the five corresponding cultivation racks to reflect the environmental differences at different height levels. Each sensor group integrates a temperature sensor, a humidity sensor, a light sensor, a carbon dioxide concentration sensor, and an oxygen concentration sensor. In this embodiment, the temperature sensor is specifically model SHT35, and the carbon dioxide sensor is specifically model MH-Z19B, thereby achieving comprehensive monitoring of multidimensional environmental parameters affecting mushroom growth. The data lines of all sensors converge via an RS485 bus and are wired to the gateway using the Modbus protocol. The gateway encapsulates the data and synchronizes it to the central control module via the TCP protocol at a sampling frequency of 1Hz.
[0029] The control module is used to respond to control signals to regulate the environment inside the mushroom house. To ensure the uniformity and scientific nature of environmental regulation, the hardware layout of the control module has been specifically optimized. The heating system includes electric heating cables laid on the floor of the mushroom house, with the cables evenly covering the floor in an S-shaped wiring pattern. This layout utilizes the physical principle of natural rising hot air to achieve bottom-up heating, effectively eliminating vertical temperature gradients within the room. The humidification system includes five ultrasonic spray heads installed at the bottom front of each of the five cultivation racks. To prevent direct spraying of mist onto the mushroom surface and causing mold, the nozzles of each ultrasonic spray head are tilted upwards at a 15° angle relative to the horizontal plane. This tilted design, combined with the indoor airflow, allows the mist to slowly rise and diffuse between the cultivation rack layers. In addition, the system has a 50L water storage tank in the corner of the mushroom house, with an integrated water level sensor connected to the spray heads via water pipes, ensuring an adequate water supply and preventing dry burning. The ventilation system includes an axial flow fan installed in the center of the mushroom house roof, and a first and second motorized louver symmetrically arranged on the upper part of the side walls of the mushroom house. The axial flow fan, the first and second motorized louvers are linked and controlled to create negative pressure inside the room during exhaust, forcing fresh air to enter evenly from the side walls. The lighting system includes LED light strips installed on the upper edge of each of the five cultivation racks. The red to blue light spectrum ratio of the LED light strips is set at 7:3 to adapt to the photosynthetic and growth characteristics of mushrooms. In addition, the system also includes a shading system, specifically a double-layered shading curtain installed on the inside of the mushroom house roof. The curtain is made of polyester fiber with a light-blocking rate of 95% and is driven by a stepper motor with a torque of 0.5 N·m. This stepper motor is connected to the central control module by a driver and can complete a full closing or full opening action within 30 seconds, used to assist in adjusting the photoperiod or adjusting external strong light. All the above-mentioned actuators are connected to the power supply through solid-state relays, and the central control module controls the equipment by controlling the on and off of the relays.
[0030] The central control module is the decision-making core of the system, comprising a cascaded reinforcement learning model and a PID controller. The reinforcement learning model uses the Deep Deterministic Policy Gradient (DDPG) algorithm, which includes an Actor network and a Critic network. The reinforcement learning model is configured to output a reinforcement learning policy based on the current state vector. The state vector consists of the deviations of various environmental parameters from the target value. , the rate of change of deviation of each environmental parameter The PID control parameters for the current time step are... , , The state vector is composed of the normalized operating states of each device in the control module. This state vector comprehensively represents the deviation of the current environment from the ideal target and the operating load of the system itself.
[0031] During the control process, the Actor network generates a PID parameter configuration strategy based on the state vector, i.e., the adjustment amount of the PID parameters. , , The Critic network is configured to evaluate the value of the PID parameter configuration strategy based on the state vector and reward function, and update the parameters of the Actor network using a gradient descent algorithm. The PID controller determines the PID control parameters for the current time step based on the reinforcement learning strategy, i.e., the updated proportional, integral, and derivative coefficients, and calculates the final control input based on these parameters. This control quantity The signals are then converted into control signals to drive the control module. Control commands generated by the central control module and real-time device feedback status are communicated with the user interface module via the WebSocket protocol for low-latency interaction. This cascaded structure leverages the stability characteristics of traditional PID controllers while utilizing the online learning capabilities of reinforcement learning to achieve adaptive parameter optimization.
[0032] To mitigate the motion oscillation problem that may occur during the exploration process in reinforcement learning algorithms, reinforcement learning models are based on the inclusion of a time-smoothing regularization term. The optimization objective is used for training. To achieve structured constraints on the output of the reinforcement learning policy based on the smoothness constraints of the continuous-time control trajectory, it is assumed that the control quantity can be represented as a locally smooth function in the continuous-time domain. .right and exist Perform a second-order Taylor expansion at this point, and the expression is as follows:
[0033]
[0034] Therefore, the second-order consistency residual of the control quantity in the discrete-time domain is specifically as follows:
[0035] This term can be viewed as a discrete approximation of the second derivative of the continuous-time control trajectory, used to characterize the curvature change characteristics of the control signal in the time dimension. Based on the above analysis, the reinforcement learning model constructs a time smoothing regularization term based on Taylor expansion as follows:
[0036] In the formula, , , These are the control values for the current time step, the previous time step, and the time step before that, respectively. The sampling period is and The weighting coefficients are defined here. By jointly constraining the first and second derivative terms of the control quantity, this regularization term can effectively filter out high-frequency control jitter without weakening the system's low-frequency response capability. The aforementioned time consistency regularization term is then integrated with the reinforcement learning basic objective function. The unified strategy optimization objective is as follows:
[0037] Furthermore, the optimization of reinforcement learning models is also based on specific reward functions. The reward function comprehensively considers control accuracy, response speed, energy efficiency, and motion smoothness; its expression is:
[0038] In the formula, This represents the absolute deviation of environmental parameters from target values. The rate of change of deviation The total energy consumption for executing the control module, This represents the rate of change of the control quantity between two adjacent time steps. The weighting coefficient is specifically set as follows: , , , By adjusting the rate of change of the control quantity in the reward function. By imposing penalties, the model can be further guided to generate stable control sequences from the perspective of strategy selection.
[0039] During the model training phase, the system uses historical mushroom cultivation data to construct a simulation environment for continuous iteration. When the number of iterations reaches more than 1000 rounds, or the fluctuation range of the reward value within 50 consecutive rounds is less than 5%, the model training is considered complete and the optimal strategy is solidified.
[0040] Regarding system safety, the central control module also implements fault-tolerant control logic. The system monitors the actual operating status of the control module in real time, specifically the deviation between the current feedback from solid-state relays or the actual physical state of the equipment and the control command status. When the deviation exceeds 15%, the system determines that a single anomaly has occurred, reissues the control input, and triggers an alarm. If the deviation exceeds 15% three times consecutively, the system determines that the algorithm has failed or there is a hardware malfunction, and automatically switches to the backup control mode. The backup control mode uses pre-stored historically optimal PID parameters for degraded control to ensure that the environment inside the mushroom house does not fluctuate drastically.
[0041] Example 2 This embodiment, based on Embodiment 1, further details the adaptive adjustment mechanism of the central control module in response to model uncertainty. The central control module also includes an uncertainty-aware adjustment module. To quantify decision risk, the reinforcement learning strategy consists of multiple parallel Actor networks. These parallel Actor networks learn using different initial parameters or different subsets of empirical samples during the training phase.
[0042] The uncertainty-aware adjustment module is based on the set of control outputs of multiple parallel actor networks under the same state. Perform statistical analysis and calculate statistical variance. The specific statistical variance is as follows:
[0043] Based on the estimation results of policy uncertainty, the uncertainty-aware adjustment module performs uncertainty-aware scaling on the original control quantity, resulting in a scaled control quantity. The expression is as follows:
[0044] In the formula, To enhance the original control output of the learning strategy, This is the uncertainty adjustment coefficient. When the system is in a high-risk, high-uncertainty state, the output control quantity will be automatically suppressed, making the actuator action more conservative.
[0045] Meanwhile, to proactively guide the strategy to avoid high-risk areas during the model training phase, the central control module introduces an uncertainty regularization term, specifically:
[0046] In the formula The penalty weights. The ultimate optimization objective of the reinforcement learning model. Set as:
[0047] By adding a statistical variance term to the total loss function, the model tends to search for strategies that achieve a high degree of consensus among the sub-networks, thereby improving the reliability of the control logic.
[0048] Example 3 This embodiment illustrates the details of the collaborative operation between the control module and the environmental sensing module. The environmental sensing module includes five sensor groups distributed in the middle of each cultivation rack. Each sensor group contains sensors for temperature, humidity, light, carbon dioxide, and oxygen concentration, all using high-precision industrial-grade probes.
[0049] The humidification system in the control module features five ultrasonic spray nozzles tilted upwards at a 15° angle relative to the horizontal. This design, combined with the negative pressure airflow generated by the axial flow fan, causes the sprayed atomized particles to form an upward parabolic trajectory in the interlayer gaps. As the hot air rises and diffuses, it solves the problem of excessive moisture at the bottom and dryness at the top. When the carbon dioxide concentration exceeds a preset threshold, the ventilation system activates the axial flow fan installed in the center of the roof, simultaneously opening the first and second motorized louvers located on the upper side walls, creating negative pressure exhaust and drawing in fresh air. The lighting system provides 7:3 red-blue spectrum illumination through LED light strips installed on the upper edge of each cultivation rack, and uses double-layered shading curtains with a 95% shading rate to achieve precise control of the light cycle.
[0050] Example 4 like Figure 2 As shown in the figure, this embodiment illustrates a mushroom house environment control method based on reinforcement learning adaptive PID, which includes the following specific steps.
[0051] First, the sensing stage. Environmental parameters inside the mushroom house are collected in real time at a data acquisition frequency of 1Hz by environmental sensing modules installed on each cultivation rack.
[0052] Second, the feature preprocessing stage. The central control module receives environmental parameters and calculates the deviations. and the rate of change of deviation The current state vector is constructed by normalizing the equipment status information and the current PID control parameters.
[0053] Third, the intelligent decision-making stage. The central control module inputs the state vector into the reinforcement learning model, and outputs the PID parameter adjustment through the Actor network based on the DDPG algorithm. , , Simultaneously, the uncertainty perception and adjustment module calculates the policy uncertainty and adaptively scales and adjusts the control quantity.
[0054] Fourth, the control quantity calculation stage. The central control module updates the proportional, integral, and derivative parameters through the PID controller to calculate the final control quantity, which is specifically manifested as the duty cycle signal or voltage command of each actuator.
[0055] Fifth, the execution and fault tolerance phase. The execution control module responds to the control input, and the central control module monitors the actual status of the equipment in real time. If it detects that the actual operating deviation exceeds 15% for three consecutive times, it automatically switches to the standby control mode based on preset safety parameters and sends an alarm.
[0056] Sixth, the feedback and iteration phase. The system calculates a reward value that integrates control accuracy, energy consumption, and motion smoothness based on the environmental adjustment results, and stores the data in an experience pool for continuous online learning of the reinforcement learning model.
[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A mushroom house environmental control system based on reinforcement learning adaptive PID, characterized in that, include: The environmental sensing module is used to collect environmental parameters inside the mushroom house in real time; An execution control module is used to respond to control quantities to adjust the environment inside the mushroom house; as well as The central control module is connected to both the environmental perception module and the execution control module. The central control module includes a cascaded reinforcement learning model and a PID controller; The reinforcement learning model is configured to output a reinforcement learning policy based on the current state vector, wherein the state vector represents the deviation state between the environmental parameters and the target value; The PID controller is configured to determine the PID control parameters for the current time step according to the reinforcement learning strategy, and to calculate the control quantity based on the PID control parameters to drive the execution control module. The reinforcement learning model is based on a time-smoothing regularization term. The time smoothing regularization term is trained to optimize the objective. Based on the local smoothing function in the continuous time domain of the control quantity The second-order Taylor expansion is constructed, and its expression is: In the formula, , , These are the control variables for the current time step and the historical time step, respectively. The sampling period is and These are the weighting coefficients; The central control module further includes an uncertainty-aware adjustment module. The reinforcement learning strategy consists of multiple parallel Actor networks. The uncertainty-aware adjustment module adjusts the control output set of the multiple parallel Actor networks under the same state based on the statistical variance. Estimate the policy uncertainty in the current state, and scale the original control quantity based on the policy uncertainty. The scaled control quantity is... The expression is: In the formula, To enhance the original control output of the learning strategy, This is the uncertainty adjustment coefficient.
2. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The reinforcement learning model is based on a reward function. To optimize, its expression is: In the formula, This represents the absolute deviation of environmental parameters from target values. The rate of change of deviation The total energy consumption for executing the control module, This represents the rate of change of the control quantity between two adjacent time steps. , , , These are the weighting coefficients.
3. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 2, characterized in that, The specific values of the weighting coefficients are set as follows: , , , .
4. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The ultimate optimization objective of the reinforcement learning model for: In the formula, To reinforce the learning of the basic objective function, For uncertainty regularization, For penalty weights.
5. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The state vector is composed of normalized parameters and includes the deviations of each environmental parameter from the target value. , the rate of change of deviation of each environmental parameter The PID control parameters at the current time step and the current operating status of each device in the execution and control module.
6. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The central control module is also configured to execute fault-tolerant control logic: monitor in real time the deviation between the actual operating state of the execution control module and the state of the control quantity command; when the deviation exceeds 15%, reissue the control quantity; when the deviation is detected to exceed 15% three times in a row, switch to standby control mode.
7. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The mushroom house is equipped with five cultivation racks, which are arranged at equal intervals along the vertical direction. The environmental sensing module includes five sensor groups, which are installed in the middle of each layer of the corresponding five cultivation racks. Each sensor group includes a temperature sensor, a humidity sensor, a light sensor, a carbon dioxide concentration sensor, and an oxygen concentration sensor.
8. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The execution control module includes a heating system, which includes electric heating cables laid on the floor of the mushroom house. The electric heating cables cover the floor of the mushroom house in an S-shaped wiring pattern.
9. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The execution control module includes a humidification system, which includes five ultrasonic spray heads installed at the bottom front end of each of the five cultivation racks; the nozzle of each ultrasonic spray head is set at an upward tilt of 15° relative to the horizontal plane.
10. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The execution control module includes a ventilation system, which includes: an axial flow fan installed in the center of the mushroom house roof; and a first electric louver and a second electric louver, which are symmetrically arranged on the upper part of the side wall of the mushroom house; wherein the axial flow fan is linked and controlled with the first electric louver and the second electric louver.
11. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The execution control module includes a lighting system, which includes LED light strips installed on the upper edge of each of the five cultivation racks, with the ratio of red light to blue light spectrum of the LED light strips being 7:
3.
12. The mushroom house environmental control system based on reinforcement learning adaptive PID according to claim 1, characterized in that, The execution control module includes a shading system, which includes a double-layered shading curtain installed on the inside of the mushroom house roof, with a shading rate of 95%.
13. A mushroom house environment control method based on reinforcement learning adaptive PID, characterized in that, The mushroom house environmental control system based on reinforcement learning adaptive PID as described in any one of claims 1 to 12, the method comprising the following steps: The environmental parameters inside the mushroom house are collected in real time through the environmental sensing module; The central control module outputs a reinforcement learning policy based on the current state vector through a reinforcement learning model; The central control module determines the PID control parameters for the current time step using the PID controller according to the reinforcement learning strategy, and calculates the control quantity based on the PID control parameters. The control module responds to the control quantity to adjust the environment inside the mushroom house.