An intelligent environment control method for growing agaricus based on deep reinforcement learning
By employing an intelligent environmental control method based on deep reinforcement learning, the problems of reliance on manual labor and imprecise environmental control in traditional shiitake mushroom cultivation have been solved. This method achieves efficient and real-time environmental regulation, thereby improving the yield and quality of shiitake mushroom cultivation.
Patent Information
- Application Number
- CN202410033177.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-01-09
AI Technical Summary
Traditional methods of cultivating shiitake mushrooms rely on manual control of environmental conditions, resulting in low production efficiency, resource waste, and unstable yields. Furthermore, the application of deep reinforcement learning in shiitake mushroom cultivation faces challenges such as difficulty in data acquisition, high computational resource requirements, and insufficient hardware.
A deep reinforcement learning-based intelligent environmental control method is adopted. By establishing a growth environment model of mushroom shiitake, designing an action space and reward function, and combining expert experience strategies, the double Q learning algorithm is used to optimize the environmental control strategy to achieve intelligent environmental regulation.
It improves the efficiency and yield of shiitake mushroom cultivation, reduces costs, enhances the cultivation experience, and enables real-time perception and dynamic adjustment of the shiitake mushroom growth environment to meet the needs of different growth stages.
Smart Images

Figure CN117694183B_ABST
Abstract
Description
Technical Field
[0001] This invention patent belongs to the field of deep reinforcement learning and environmental control, and provides an intelligent environmental control method for mushroom cultivation based on deep reinforcement learning. Background Technology
[0002] Shiitake mushrooms are a widely cultivated edible fungus globally, and their high nutritional value and delicious taste make them popular in the food industry. However, shiitake mushroom growth is highly sensitive to environmental conditions, including temperature, humidity, light, and carbon dioxide concentration. To obtain high yields and high-quality shiitake mushrooms, these environmental parameters must be precisely controlled. However, traditional shiitake mushroom cultivation typically relies on manual control of environmental conditions such as temperature, humidity, and ventilation. This requires significant manpower and resources and is easily affected by human factors, such as differences in operator skills and experience. Therefore, traditional cultivation methods suffer from low production efficiency, resource waste, and unstable yields.
[0003] Reinforcement learning is a machine learning method, and deep reinforcement learning is a branch of reinforcement learning that combines deep neural networks and reinforcement learning algorithms. It has achieved significant results in the field of intelligent environmental control. It enables agents to autonomously learn and improve their behavioral strategies based on environmental feedback. In mushroom cultivation, deep reinforcement learning can be used to automatically control environmental parameters to optimize growth conditions and maximize yield and quality.
[0004] The intelligent environmental control method for shiitake mushroom cultivation based on deep reinforcement learning aims to overcome the limitations of traditional cultivation methods and improve the production efficiency and yield stability of shiitake mushrooms. This method utilizes deep reinforcement learning technology to make the control of the shiitake mushroom growth environment more intelligent and automated. Through continuous learning and optimization, this method can adapt to the growth needs of shiitake mushrooms at different stages, thereby achieving high-yield and high-quality shiitake mushroom production.
[0005] The shortcomings of existing technologies are summarized as follows:
[0006] 1. Limitations of existing technologies in solving specific problems
[0007] While traditional shiitake mushroom cultivation methods have achieved certain results in terms of production efficiency and yield, they still have many limitations. First, traditional methods rely on manual control of environmental conditions, and differences in operator skills and experience can affect the control effectiveness. Second, traditional methods still face significant challenges in dealing with complex environmental changes, such as temperature, humidity, light, and carbon dioxide concentration. Furthermore, traditional methods also have limitations in automating the control of environmental parameters, resulting in less precise environmental condition control, which also affects shiitake mushroom growth and yield.
[0008] 2. Shortcomings of existing technologies in achieving environmental control for shiitake mushrooms
[0009] Despite the significant achievements of deep reinforcement learning in the field of intelligent environmental control, some shortcomings remain in the cultivation of shiitake mushrooms. First, the deep reinforcement learning-based shiitake mushroom cultivation technology has not been fully validated at all stages of cultivation, requiring more experimental data to demonstrate its effectiveness and feasibility. Second, although deep reinforcement learning can make the control of the shiitake mushroom growth environment more intelligent and automated, manual adjustment and optimization of environmental parameters are still necessary, and the real-time performance and stability of the network need to be considered in practical applications.
[0010] 3. Difficulties of existing technologies in environmental control for treating shiitake mushrooms
[0011] Traditional methods of shiitake mushroom cultivation suffer from low production efficiency, resource waste, and unstable yields. This is primarily because these methods rely on manual control of environmental parameters, which is inherently imprecise and susceptible to subjective influences. Furthermore, traditional methods struggle to achieve high yields and high-quality shiitake mushrooms at different growth stages. These limitations restrict the effectiveness of traditional shiitake mushroom cultivation.
[0012] 4. Difficulties in applying deep reinforcement learning techniques to environmental control of shiitake mushrooms
[0013] The application of deep reinforcement learning in shiitake mushroom cultivation still faces several challenges. First, it requires a large amount of data for training, which is difficult to obtain for shiitake mushroom cultivation scenarios where environmental control parameters are unstable. Second, the learning process of deep reinforcement learning requires significant computational resources, making it challenging to implement in resource-constrained shiitake mushroom cultivation scenarios. Finally, the application of deep reinforcement learning requires suitable hardware support, which is often lacking in some shiitake mushroom cultivation scenarios. Summary of the Invention
[0014] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an intelligent environmental control method for shiitake mushroom cultivation based on deep reinforcement learning. This method applies modern computer science and deep reinforcement learning technology to the traditional agricultural field, achieving intelligent environmental control for shiitake mushroom cultivation to improve yield and quality. This objective is achieved through the following technical solutions:
[0015] This invention provides an intelligent environmental control method for mushroom cultivation based on deep reinforcement learning, comprising the following steps:
[0016] Step s1: Based on the growth process and control of the growth environment of shiitake mushrooms, establish a shiitake mushroom cultivation and growth environment model and design the state space of shiitake mushrooms. The state space includes the shiitake mushroom growth environment state space and the shiitake mushroom growth state space. The shiitake mushroom growth environment space includes the environmental state dimensions and value ranges in the shiitake mushroom growth environment. The shiitake mushroom growth state space includes the state dimensions of shiitake mushrooms and their value ranges.
[0017] Action space design: Define the environmental control equipment as an intelligent agent to obtain the action space of mushroom cultivation. The action space includes the adjustable environmental parameters and adjustable range in the environmental control equipment.
[0018] Step s2: Integrate existing expert experience strategies and planting data for shiitake mushroom cultivation, construct a sampling pool for the intelligent environmental control algorithm, and design the algorithm's sampling strategy;
[0019] Step s3: Based on the characteristics of the growth process of shiitake mushrooms, design a reward function to guide the intelligent control algorithm. The reward function is designed as a shiitake mushroom growth quality reward function. The reward function is used to evaluate whether the environmental control achieves the planting goal and whether it increases the planting yield.
[0020] Step s4: Design the training process of the intelligent environmental control algorithm, use the intelligent environmental control algorithm to learn the environmental control strategy, and change the growth environment of shiitake mushrooms according to the output environmental control strategy.
[0021] In the above technical solution, step s1 includes:
[0022] The state space S of shiitake mushroom includes the state space of the shiitake mushroom's growth environment S. env The growth state space of shiitake mushroom S hua ;
[0023] The aforementioned mushroom growth environment state space S env The state dimension of the mushroom cultivation environment is included, and is represented as follows:
[0024] S env =(T, H, C, L, G, D)
[0025] Where T represents the temperature of the growing environment, H represents the humidity of the growing environment, C represents the carbon dioxide concentration of the growing environment, L represents the light intensity of the growing environment, G represents the growth stage of the shiitake mushroom, with a value of 0 indicating that the shiitake mushroom is in the mycelium growth stage and a value of 1 indicating that the shiitake mushroom is in the color change stage, and D indicates whether the shiitake mushroom is in daytime or nighttime, with a value of 0 indicating daytime and a value of 1 indicating nighttime.
[0026] To optimize the implementation, temperature, humidity, and carbon dioxide concentration were discretized by dividing the area. Two environmental states, namely the mycelium growth period and the color change period, were designed for different stages of shiitake mushroom growth. Two environmental states, namely day and night, were added to account for the light intensity during shiitake mushroom growth.
[0027] The aforementioned mushroom growth state space S hua This includes the state dimension of shiitake mushroom growth, represented as follows:
[0028] S hua =(Hua_quality)
[0029] Here, Hua_quality is a score from 0 to 100, used to represent the current growth status of the shiitake mushroom. The higher the score, the better the growth status of the shiitake mushroom, with 100 being the ideal final growth state. The growth quality of the shiitake mushroom is evaluated using a deep learning method. Based on the diameter, crack size, and color characteristics of the shiitake mushroom, the growth status is mapped to a score of 0-100 for quantitative evaluation. This deep learning method first collects and labels image data of the shiitake mushroom growth process, and then uses a deep learning model for training and optimization. This deep learning model is used to perform the tasks of classifying the growth status and regressing the quality score. Finally, the quality score is converted into a score of 0-100 through a mapping function, thereby achieving an accurate quantitative evaluation of the growth quality of the shiitake mushroom.
[0030] In summary, the state space S of the shiitake mushroom is:
[0031] S=(T,H,C,L,G,D,Hua_quality)
[0032] The action space includes the equipment and control methods used to regulate the environment. The action space of the intelligent environmental control algorithm is composed of four items in the mushroom cultivation environment: light, temperature, humidity, and fan controller. Among them, light control includes two states: light enhancement and light reduction; humidity control includes two states: humidity increase and humidity decrease; temperature control includes two states: temperature increase and temperature decrease; and fan control includes two states: fan on and fan off. The intelligent environmental control algorithm selects the control actions of the four controllers based on the environmental status information. A total of 16 action combinations constitute the action space A of the intelligent environmental control algorithm, which is used to regulate the mushroom growth environment accordingly.
[0033] In the design of the action space, the intelligent environmental control algorithm presets a threshold for each action range based on experience, which represents the adjustment range of the action parameter. When the parameter exceeds the threshold, it will cause a large change in the environment, thereby affecting the growth of shiitake mushrooms. When the output action range exceeds the preset threshold, the action output stops, and the environment is corrected. The environmental control equipment is controlled to restore the various environmental parameters to a normal range.
[0034] In the above technical solution, step s2 includes:
[0035] Based on existing expert experience strategies and planting data for shiitake mushroom cultivation, a hierarchical sampling experience pool for intelligent environmental control algorithms is constructed. The hierarchical sampling experience pool is divided into an expert experience layer and a general information experience layer, and the sampling strategy of the algorithm is designed.
[0036] The stratified sampling experience pool consists of quadruples (s t a t r t s t+1 )constitute;
[0037] Where s t Indicates the current state of the mushroom, a. t r represents the action performed by the agent in the current state. t This represents the reward an agent receives from the environment after performing an action, and represents the agent's performance in the environment. t Take a specific action a in a state t The degree of good or bad, i.e., r t Only with s t With a t Related to, s t+1 Indicates that when performing action a t The new state after that indicates the next state that the mushroom will transition from its current state;
[0038] Design the maximum number of data entries for each experience layer and the update criteria, and initialize the expert experience layer and the general information experience layer.
[0039] Design sampling conditions for different layers of sampling pools, and determine the conditions under which the intelligent environmental control algorithm extracts samples from each experience layer for learning;
[0040] If the current mushroom state space S belongs to the expert experience layer, then sample the data in the expert experience layer and select the corresponding action a; if the current mushroom state space S does not belong to the expert experience layer, then compare and sample the data in the ordinary information experience layer. If the comparison exists, then select the corresponding action.
[0041] In the above technical solution, step s3 includes:
[0042] Based on the characteristics of the growth process of shiitake mushrooms, a reward function is designed to guide the intelligent control algorithm. The reward function is designed as a shiitake mushroom growth quality reward function. The reward function is used to evaluate whether the environmental control achieves the planting goal and whether it increases the planting yield.
[0043] The reward function for the growth quality of shiitake mushrooms is designed as follows: r t_q (st a t s t+1 )=(Hua_quality-target_quality)
[0044] Wherein, the subscript t_q is the defined subscript, Hua_quality is the quality evaluation index of shiitake mushrooms predicted by the deep learning model, representing the current growth status of the shiitake mushrooms; target_quality is the target shiitake mushroom quality, which is 100 here; r t_g (s t a t s t+1 ) is the reward function for the growth quality of shiitake mushrooms. The better the growth quality of shiitake mushrooms, the closer the growth quality of shiitake mushrooms is to the target quality of shiitake mushrooms, and the larger the corresponding function value.
[0045] The reward function r of intelligent control algorithms t (s t a t s t+1 )=λ·r t_q (s t a t s t+1 )
[0046] λ is a hyperparameter that controls the reward weight.
[0047] In the above technical solution, step s4 includes:
[0048] Design a training process for an intelligent environmental control algorithm, which is a double-Q learning algorithm based on an expert knowledge network. Let M be the total number of training rounds, N be the number of iterations within a round, and Q(s) be the value network. t a t The target value network is denoted as Q(s), where ω is the target value network. t a t ω - ), where ω represents the parameter of the current value network, ω - The meaning of target value network parameters, and the training process includes the following steps:
[0049] Step 5.1: Input the mushroom's state space S, action space A, discount rate y, learning rate θ, and time t.
[0050] Initialize the stratified sampling experience pool D with a capacity of W;
[0051] Step 5.2: Iterate through all rounds from episode = 1 to M, and execute:
[0052] Step 5.2.1: Initialize state S;
[0053] Step 5.2.2: Traverse the loop from t=1 to N, and execute:
[0054] Step 5.2.2.1, in state s t In this case, compare the stratified sampling experience pool and select the appropriate action;
[0055] Step 5.2.2.2, if state s t If it falls within the scope of the expert experience level, then select action a according to the expert experience level. t ;
[0056] Step 5.2.2.3, if state s t If it falls within the scope of the general information experience layer, then select action a according to the general information experience layer. t ;
[0057] Step 5.2.2.4, if state s t If neither of the above two options applies, then action a is selected using the ε-greedy strategy. t ;
[0058] Step 5.2.2.5: After the judgment is completed, execute the corresponding action a. t By observing the planting environment, a new reward r is calculated. t Update the new status s t+1 ;
[0059] Step 5.2.2.6: Update the value network parameters and add the quadruplets (s) t a t r t s t+1 The sample is placed into the stratified sampling experience pool D and continuously updated.
[0060] Step 5.2.3: The loop terminates when state S reaches the termination state; otherwise, the current loop ends and step 5.2.2 is executed.
[0061] Step 5.3, until the DQN value network Q(s) t a t If ω converges, the entire training will terminate; otherwise, the current iteration will end and step 5.2 will be executed.
[0062] For step 5.2.2.6, the typical steps for updating the value network parameters ω of the double-Q learning algorithm are as follows:
[0063] By performing forward propagation on DQN, the Q-value of the value network is obtained:
[0064]
[0065] formula Indicates in st+1 In the given state, select the action a with the highest Q value. * .
[0066] formula Indicates in s t+1 In this state, select action a. * , That is the corresponding value network Q value.
[0067] formula Indicates in s t In this state, select action a. t , This corresponds to the Q-value of the value network;
[0068] Define the corresponding target and error:
[0069] and
[0070] By backpropagating DQN, we obtain the gradient:
[0071]
[0072] Update the parameters of DQN using gradient descent:
[0073] ω new ←ω now -θ·δ t ·g t θ is the learning rate, which controls the size of the update step.
[0074] Let τ∈(0,1) be the hyperparameters that need to be manually tuned. Then, update the parameters of the target value network using a weighted average:
[0075]
[0076] For step 5.2.2.4, the ε-greedy strategy is:
[0077]
[0078] argmax a Q(s t , a; ω) indicates that in s t In the state, the action 'a' with the highest Q value is selected, which represents the best action under the current policy, where ε is a hyperparameter set between 0 and 1.
[0079] Because the present invention adopts the above-described technical solution, it has the following beneficial effects:
[0080] (1) The intelligent environmental control method for mushroom cultivation based on deep reinforcement learning adopted in this invention helps to improve the efficiency of mushroom cultivation and increase the quality and yield of mushroom cultivation.
[0081] (2) This invention introduces an expert experience strategy into the intelligent environmental control algorithm, avoiding the problem of poor stability and practicality of the early environmental control strategy learned by the intelligent environmental control algorithm, thereby reducing unnecessary exploration during the training process. It makes full use of expert experience, which usually comes from experienced professional growers and includes best practices in actual operation.
[0082] (3) By combining deep reinforcement learning technology, the intelligent environmental control method of this invention can perceive and adjust the growth environment of shiitake mushrooms in real time, adapting it to the needs of different growth stages. This dynamic regulation is impossible to achieve with traditional methods. In the process of implementation, technical challenges such as the nonlinearity, high interactivity, and uncertainty of environmental parameters were overcome. Through the training of the deep reinforcement learning model, an intelligent environmental control system was successfully constructed, which has the characteristics of adaptability and high intelligence, and can adjust parameters in real time according to the growth status of shiitake mushrooms and environmental feedback. This technological innovation provides a new intelligent solution for traditional shiitake mushroom cultivation, enabling producers to better cope with environmental changes, increase yield, and optimize the quality of shiitake mushrooms.
[0083] The intelligent environmental control method for mushroom cultivation based on deep reinforcement learning provided by this invention solves the following problems:
[0084] (1) Uncertainty: During the growth of shiitake mushrooms, the variables of the growth environment (such as temperature, humidity, carbon dioxide concentration, etc.) are difficult to control precisely, making it difficult to predict the growth status of shiitake mushrooms, thus affecting the quality of shiitake mushrooms.
[0085] (2) Complexity: In the existing methods of growing shiitake mushrooms, the control of the growing environment requires the manual setting of various parameters, including light, temperature, humidity, and fans. However, the control methods of these parameters are relatively complex and require shiitake mushroom growers to have a certain level of technical expertise. It is difficult to rely entirely on intelligent environmental control systems.
[0086] (3) Real-time: The growth process of shiitake mushrooms is real-time, and the growth environment needs to be effectively controlled in a short period of time in order to improve the growth quality and yield of shiitake mushrooms.
[0087] The technical solution provided by this invention improves efficiency, reduces costs, and enhances the planting experience for farmers.
[0088] (1) Improve efficiency: The intelligent environmental control system can adjust environmental parameters in real time according to the characteristics and needs of the shiitake mushroom growth process, so as to improve the growth efficiency and yield of shiitake mushrooms and reduce the grower's manpower and time costs.
[0089] (2) Reduced costs: Intelligent environmental control systems can significantly reduce the cost of growing shiitake mushrooms by utilizing existing technologies and data, as there is no need to purchase expensive monitoring equipment or shiitake mushroom growth expert experience strategies.
[0090] (3) Improved User Experience: The intelligent environmental control system can adjust environmental parameters in real time according to the characteristics and needs of the shiitake mushroom growth process, thereby improving the quality and yield of shiitake mushrooms and enhancing the grower's experience. In addition, the intelligent environmental control system can regularly collect and analyze data during the shiitake mushroom growth process to help growers identify problems in a timely manner and take measures, thus improving grower satisfaction. Attached Figure Description
[0091] Figure 1 The method flowchart provided in this application;
[0092] Figure 2 The flowchart for modeling the interaction between mushroom cultivation and the environment provided in this application;
[0093] Figure 3 Flowchart of the sampling experience pool for the intelligent environmental control algorithm provided in this application;
[0094] Figure 4 The diagram shows the structure of the intelligent environmental control double-Q learning algorithm model provided in this application. Detailed Implementation
[0095] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.
[0096] like Figure 1 As shown, this invention provides an intelligent environmental control method for mushroom cultivation based on deep reinforcement learning, comprising the following steps:
[0097] Step s1: Based on the growth process and environmental control of shiitake mushrooms, establish a shiitake mushroom cultivation and growth environment model. Design the state space of shiitake mushrooms, which includes the shiitake mushroom growth environment state space and the shiitake mushroom growth state space. The shiitake mushroom growth environment space includes the environmental state dimensions and their value ranges in the shiitake mushroom growth environment; the shiitake mushroom growth state space includes the state dimensions of the shiitake mushroom and their value ranges.
[0098] Action space design: Define the environmental control equipment as an intelligent agent to obtain the action space of mushroom cultivation. The action space includes the adjustable environmental parameters and adjustable range in the environmental control equipment.
[0099] Step s2: Integrate existing expert experience strategies and planting data for shiitake mushroom cultivation, construct a sampling pool for the intelligent environmental control algorithm, and design the algorithm's sampling strategy.
[0100] Step s3: Based on the characteristics of the shiitake mushroom growth process, design a reward function to guide the intelligent control algorithm. The reward function includes a shiitake mushroom growth quality reward function and a shiitake mushroom growth environment control reward function. The reward function is used to evaluate whether the environmental control achieves the planting goal and whether it increases the planting yield.
[0101] Step s4: Design the training process of the intelligent environmental control algorithm, use the intelligent environmental control algorithm to learn the environmental control strategy, and change the growth environment of shiitake mushrooms according to the output environmental control strategy.
[0102] like Figure 2 As shown, step s1 includes:
[0103] The state space S of shiitake mushroom includes the state space of the shiitake mushroom's growth environment S. env The growth state space of shiitake mushroom S hua ;
[0104] The aforementioned mushroom growth environment state space S env The state dimension of the mushroom cultivation environment is included, and is represented as follows:
[0105] S env =(T, H, C, L, G, D)
[0106] Where T represents the temperature of the growing environment, H represents the humidity of the growing environment, C represents the carbon dioxide concentration of the growing environment, L represents the light intensity of the growing environment, G represents the growth stage of the shiitake mushroom (0 represents the mycelium growth stage, 1 represents the color change stage), and D represents whether it is day or night (0 represents daytime, 1 represents nighttime). These are all environmental factors that affect the growth of shiitake mushrooms.
[0107] To optimize the implementation, temperature, humidity, and carbon dioxide concentration were discretized by dividing the area. Two environmental states, namely the mycelium growth period and the color change period, were designed for different growth stages of shiitake mushrooms. Two environmental states, namely day and night, were also added to the light intensity for the growth of shiitake mushrooms.
[0108] The aforementioned mushroom growth state space S hua This includes the state dimension of shiitake mushroom growth, represented as follows:
[0109] S hua =(Hua_quality)
[0110] Here, `Hua_quality` is a score from 0 to 100, representing the current growth status of the shiitake mushroom. A higher score indicates better growth, with 100 representing the ideal final growth state. This method uses deep learning to evaluate the growth quality of the shiitake mushroom. Based on features such as the mushroom's diameter, crack size, and color, the growth status is mapped to a 0-100 score for quantitative assessment. The method first collects and labels image data of the shiitake mushroom growth process, then trains and optimizes a deep learning model. This model performs tasks such as growth status classification and quality score regression. Finally, a mapping function converts the quality score into a 0-100 value, achieving a precise quantitative assessment of the shiitake mushroom's growth quality.
[0111] In summary, the state space S of the mushroom is:
[0112] S=(T,H,C,L,G,D,Hua_quality)
[0113] The action space design includes the equipment and methods used to regulate the environment. Four factors from the mushroom cultivation environment—light, temperature, humidity, and fan controllers—are used as the action space for the reinforcement learning model. Light control includes two states: light increase and light decrease; humidity control includes two states: humidity increase and humidity decrease; temperature control includes two states: temperature increase and temperature decrease; and fan control includes two states: fan on and fan off. Based on the environmental state information, the intelligent environmental control algorithm selects control actions from the four controllers, resulting in 16 possible action combinations that constitute the action space A of the intelligent environmental control algorithm, used to regulate the mushroom growth environment accordingly.
[0114] In the design of the action space, the intelligent environmental control algorithm may encounter problems such as significant environmental changes caused by the agent's actions during continuous trial and error, thus affecting the growth of shiitake mushrooms. Therefore, based on experience, a threshold is preset for each action range for various environmental information, representing the control range of that action parameter. When the parameter exceeds the threshold, it will cause significant environmental changes, thereby affecting the growth of shiitake mushrooms. When the output action range exceeds the preset threshold, the action output is stopped, and the environment is corrected, controlling the environmental control equipment to restore the various environmental parameters to a normal range.
[0115] like Figure 3 As shown, step s2 includes:
[0116] Based on existing expert experience strategies and planting data for shiitake mushroom cultivation, a hierarchical sampling experience pool for intelligent environmental control algorithms is constructed. The hierarchical sampling experience pool is divided into an expert experience layer and a general information experience layer, and a sampling strategy for the algorithm is designed.
[0117] The stratified sampling experience pool consists of quadruples (s t a t r t s t+1 )constitute.
[0118] Where s t Indicates the current state of the mushroom, a. t r represents the action performed by the agent in the current state. t This represents the reward an agent receives from the environment after performing an action; it indicates the agent's performance in the context of the environment. t Take a specific action a in a state t The degree of good or bad, i.e., r t Only with s t With a t Related to, s t+1 Indicates that when performing action a t The new state after that indicates the next state that the mushroom will transition from its current state.
[0119] Design the maximum number of data entries for each experience layer and the update criteria, and initialize the expert experience layer and the general information experience layer.
[0120] Design sampling conditions for different layers of sampling pools to determine the conditions under which the intelligent environmental control algorithm extracts samples from each experience layer for learning.
[0121] If the current mushroom state space S belongs to the expert experience layer, the data in the expert experience layer is sampled first, and the corresponding action a is selected; if the current mushroom state space S does not belong to the expert experience layer, the data in the ordinary information experience layer is compared and sampled. If the comparison exists, the corresponding action is selected.
[0122] In the above technical solution, step s3 includes:
[0123] Based on the characteristics of the shiitake mushroom growth process, a reward function is designed to guide the intelligent control algorithm. The reward function is designed as a shiitake mushroom growth quality reward function. The reward function is used to evaluate whether the environmental control achieves the planting goal and whether it increases the planting yield.
[0124] The reward function for the growth quality of shiitake mushrooms is designed as follows: r t_q (s t a t s t+1 )=(Hua_quality-target_quality)
[0125] Where Hua_quality is the quality assessment index of shiitake mushrooms predicted by the deep learning model, representing the current growth status of the shiitake mushrooms; target_quality is the target quality of the shiitake mushrooms, with a value of 100 here; rt_q (s t a t s t+1 ) is the reward function for the growth quality of shiitake mushrooms. The better the growth quality of shiitake mushrooms, the closer the growth quality of shiitake mushrooms is to the target quality, and the larger the corresponding function value.
[0126] The reward function r of intelligent control algorithms t (s t a t s t+1 )=α·r t_g (s t a t s t+1 )
[0127] α is a hyperparameter that controls the reward weight.
[0128] In the above-described scheme, step s4 includes:
[0129] Design a training process for an intelligent environmental control algorithm, which is a double-Q learning algorithm based on an expert knowledge network. Let M be the total number of training rounds, N be the number of iterations within a round, and Q(s) be the value network. t a t , where ω represents the parameters of the current value network, and the training process includes the following steps:
[0130] Step 5.1: Input the mushroom's state space S, action space A, discount rate y, learning rate θ, and time t.
[0131] Initialize the stratified sampling experience pool D with a capacity of W;
[0132] Step 5.2: Iterate through all rounds from episode = 1 to M, and execute:
[0133] Step 5.2.1: Initialize state S;
[0134] Step 5.2.2: Traverse the loop from t=1 to N, and execute:
[0135] Step 5.2.2.1, in state s t In this case, compare the stratified sampling experience pool and select the appropriate action;
[0136] Step 5.2.2.2, if state s t If it falls within the scope of the expert experience level, then select action a according to the expert experience level. t ;
[0137] Step 5.2.2.3, if state s t If it falls within the scope of the general information experience layer, then select action a according to the general information experience layer.t ;
[0138] Step 5.2.2.4, if state s t If neither of the above two options applies, then action a is selected using the ε-greedy strategy. t ;
[0139] Step 5.2.2.5: After the judgment is completed, execute the corresponding action a. t By observing the planting environment, a new reward r is calculated. t Update the new status s t+1 ;
[0140] Step 5.2.2.6: Update the value network parameters and add the quadruplets (s) t a t r t s t+1 The sample is placed into the stratified sampling experience pool D and continuously updated.
[0141] Step 5.2.3: The loop terminates when state S reaches the termination state; otherwise, the current loop ends and step 5.2.2 is executed.
[0142] Step 5.3, until the DQN value network Q(s) t a t If ω converges, the entire training will terminate; otherwise, the current iteration will end and step 5.2 will be executed.
[0143] For step 5.2.2.6, the typical steps for updating the value network parameters ω of the double-Q learning algorithm are as follows:
[0144] By performing forward propagation on DQN, the Q-value of the value network is obtained:
[0145]
[0146] formula Indicates in s t+1 In the given state, select the action a with the highest Q value. * .
[0147] formula Indicates in s t+1 In this state, select action a. * , That is the corresponding value network Q value.
[0148] Define the corresponding target and error:
[0149] and
[0150] By backpropagating DQN, we obtain the gradient:
[0151]
[0152] Update the parameters of DQN using gradient descent:
[0153] ω new ←ω now -θ·δ t ·g t
[0154] Let τ∈(0,1) be the hyperparameters that need to be manually tuned. Then, update the parameters of the target value network using a weighted average:
[0155]
[0156] For step 5.2.2.4, the ε-greedy strategy is:
[0157]
[0158] argmax a Q(s t , a; ω) indicates that in s t In the state, the action 'a' with the highest Q value is selected, which represents the best action under the current policy, where ε is a hyperparameter set between 0 and 1.
[0159] This invention utilizes deep reinforcement learning technology to achieve intelligent environmental control in shiitake mushroom cultivation, overcoming the shortcomings of existing technologies. It applies modern computer science and deep reinforcement learning technology to the traditional agricultural field to improve the yield and quality of shiitake mushroom cultivation.
[0160] The above description represents preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technical or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A smart environmental control method for mushroom cultivation based on deep reinforcement learning, characterized in that, Includes the following steps: Step s1: Based on the growth process and control of the growth environment of shiitake mushrooms, establish a shiitake mushroom cultivation and growth environment model and design the state space of shiitake mushrooms. The state space includes the shiitake mushroom growth environment state space and the shiitake mushroom growth state space. The shiitake mushroom growth environment space includes the environmental state dimensions and value ranges in the shiitake mushroom growth environment. The shiitake mushroom growth state space includes the state dimensions of shiitake mushrooms and their value ranges. Motion space design: The environmental control equipment is defined as an intelligent agent to obtain the motion space of mushroom cultivation. The motion space includes the adjustable environmental parameters and their adjustable range in the environmental control equipment. Step s2: Integrate existing expert experience strategies and cultivation data for shiitake mushroom cultivation, construct a sampling pool for the intelligent environmental control algorithm, and design the algorithm's sampling strategy; Step s3: Based on the characteristics of the shiitake mushroom growth process, design a reward function to guide the intelligent control algorithm. The reward function is designed as a shiitake mushroom growth quality reward function. The reward function is used to evaluate whether the environmental control achieves the planting goal and whether it increases the planting yield. Step s4: Design the training process of the intelligent environmental control algorithm, use the intelligent environmental control algorithm to learn the environmental control strategy, and change the growth environment of shiitake mushrooms according to the output environmental control strategy. Step s1 includes: The state space of shiitake mushrooms Including the growing environment and space of shiitake mushrooms Space for the growth of shiitake mushrooms ; The aforementioned mushroom growth environment state space The state dimension of the mushroom cultivation environment is included, and is represented as follows: =(T,H,C,L,G,D) Where T represents the temperature of the growing environment, H represents the humidity of the growing environment, C represents the carbon dioxide concentration of the growing environment, L represents the light intensity of the growing environment, G represents the growth stage of the shiitake mushroom, with a value of 0 indicating that the shiitake mushroom is in the mycelium growth stage and a value of 1 indicating that the shiitake mushroom is in the color change stage, and D indicates whether the shiitake mushroom is in daytime or nighttime, with a value of 0 indicating daytime and a value of 1 indicating nighttime. To optimize the implementation, temperature, humidity, and carbon dioxide concentration were discretized by dividing the area. Two environmental states, namely the mycelium growth period and the color change period, were designed for different stages of shiitake mushroom growth. Two environmental states, namely day and night, were added to account for the light intensity during shiitake mushroom growth. The aforementioned space for the growth state of shiitake mushrooms This includes the state dimension of shiitake mushroom growth, represented as follows: =(Hua_quality) Here, Hua_quality is a score from 0 to 100, used to represent the current growth status of the shiitake mushroom. The higher the score, the better the growth status of the shiitake mushroom, with 100 being the ideal final growth state. The growth quality of the shiitake mushroom is evaluated using a deep learning method. Based on the diameter, crack size, and color characteristics of the shiitake mushroom, the growth status is mapped to a score of 0-100 for quantitative evaluation. This deep learning method first collects and labels image data of the shiitake mushroom growth process, and then uses a deep learning model for training and optimization. This deep learning model is used to perform the tasks of classifying the growth status and regressing the quality score. Finally, the quality score is converted into a score of 0-100 through a mapping function, thereby achieving an accurate quantitative evaluation of the growth quality of the shiitake mushroom. In summary, the state space of the shiitake mushroom for: =(T,H,C,L,G,D,Hua_quality) The action space includes the equipment and control methods used to regulate the environment. The action space of the intelligent environmental control algorithm is composed of four items in the mushroom cultivation environment: light, temperature, humidity, and fan controller. Among them, light control includes two states: light enhancement and light reduction; humidity control includes two states: humidity increase and humidity decrease; temperature control includes two states: temperature increase and temperature decrease; and fan control includes two states: fan on and fan off. The intelligent environmental control algorithm selects the control actions of the four controllers based on the environmental status information. A total of 16 action combinations constitute the action space A of the intelligent environmental control algorithm, which is used to regulate the mushroom growth environment accordingly. In the design of the action space, the intelligent environmental control algorithm presets a threshold for each action range based on experience, which represents the adjustment range of the action parameter. When the parameter exceeds the threshold, it will cause a large change in the environment, thereby affecting the growth of shiitake mushrooms. When the output action range exceeds the preset threshold, the action output stops, and the environment is corrected. The environmental control equipment is controlled to restore the various environmental parameters to a normal range. Step s3 includes: Based on the characteristics of the shiitake mushroom's growth process, a reward function is designed to guide the intelligent control algorithm. The reward function is designed as a shiitake mushroom growth quality reward function, which is used to evaluate whether environmental control has achieved the planting objectives and whether it has increased the planting yield. The reward function for the growth quality of shiitake mushrooms is designed as follows: ( ) =(Hua_quality-target_quality) Among them, subscript The subscripts defined here are: Hua_quality, which is the quality assessment index of shiitake mushrooms predicted by the deep learning model, representing the current growth status of the shiitake mushrooms; and target_quality, which is the target quality of shiitake mushrooms, with a value of 100 here. ( The function ) is the reward function for the growth quality of shiitake mushrooms. The better the growth quality of the shiitake mushrooms, and the closer the growth quality is to the target quality, the larger the corresponding function value. Reward function of intelligent control algorithm ( ) ( ) It is a hyperparameter that controls the weight of rewards; Step s4 includes: Design a training process for an intelligent environmental control algorithm. The training process is a double-Q learning algorithm based on an expert knowledge network. Let M be the total number of training rounds, N be the number of iterations within a round, and the value network be denoted as... The target value network is denoted as To represent the parameters of the current value network, The meaning of target value network parameters, and the training process includes the following steps: Step 5.1: Input the mushroom state space S, action space A, and discount rate. Learning rate , time t Initialize the stratified sampling experience pool D with a capacity of W; Step 5.2: Iterate through all rounds. ,implement: Step 5.2.1: Initialize state S; Step 5.2.2, Traverse the loop ,implement: Step 5.2.2.1, in the state In this case, compare the stratified sampling experience pool and select the appropriate action; Step 5.2.2.2, if the state If it falls within the scope of expert experience level, then select the action according to the expert experience level. ; Step 5.2.2.3, if the state If it falls within the scope of the general information experience layer, then it is based on the general information experience layer; Select Action ; Step 5.2.2.4, if the state If it does not fall within either of the above two categories, then through Strategy Selection action ; Step 5.2.2.5: After the judgment is completed, execute the corresponding action. By observing the planting environment, new prizes can be calculated. Encourage Update the new status ; Step 5.2.2.6: Update the value network parameters and set the quadruples. Incorporate stratified sampling experience Pool D continuously updates the experience pool; Step 5.2.3: Terminate the loop until state S reaches the termination state; otherwise, end the current loop and execute the next step. Step 5.2.2; Step 5.3, until the DQN value network If convergence is achieved, the entire training session will terminate; otherwise, the current round will end. And proceed to step 5.2; For step 5.2.2.6, the typical update of the double-Q learning algorithm's value network parameters is... The steps of the plan are as follows: By performing forward propagation on DQN, the Q-value of the value network is obtained: formula Indicates in In this state, select the action with the highest Q value. formula Indicates in In this state, select an action. , That is the corresponding value network Q value. formula Indicates in In this state, select an action. , This corresponds to the Q-value of the value network; Define the corresponding target and error: and By backpropagating DQN, we obtain the gradient: Update the parameters of DQN using gradient descent: , The learning rate controls the size of the update step. set up (0, 1) are hyperparameters that need to be manually tuned to perform weighted average updates of the target value network parameters. For step 5.2.2.4, The strategy is: Indicates in In the given state, the action 'a' with the highest Q value is selected, representing the optimal action under the current policy. These are hyperparameters that can be set between 0 and 1.
2. The intelligent environmental control method for mushroom cultivation based on deep reinforcement learning according to claim 1, characterized in that, Step s2 includes: Based on existing expert experience strategies and planting data for shiitake mushroom cultivation, a hierarchical sampling experience pool for intelligent environmental control algorithms is constructed. The hierarchical sampling experience pool is divided into an expert experience layer and a general information experience layer, and the sampling strategy of the algorithm is designed. The stratified sampling experience pool consists of quadruples constitute; in This indicates the current state of the mushroom. This represents the action performed by the agent in the current state. This represents the reward an agent receives from the environment after performing an action; it also represents the agent's... Take a specific action in a certain state The degree of good or bad, that is Only follow and related, Indicates the execution of an action The new state after that indicates the next state that the mushroom will transition from its current state; Design the maximum number of data entries for each experience layer and the update criteria, and initialize the expert experience layer and the general information experience layer. Design sampling conditions for different layers of sampling pools, and determine the conditions under which the intelligent environmental control algorithm extracts samples from each experience layer for learning; Current mushroom state space If it falls within the scope of the expert experience layer, then sample the data from the expert experience layer and select the corresponding action a; if the current state space of the mushroom... If the data does not fall within the scope of the expert experience layer, then compare it with the data in the general information experience layer. If a match is found, then select the appropriate action.
Citation Information
Patent Citations
Embryo quality comprehensive evaluation device based on deep learning
CN111539308A
Physical layer security and rate maximization method
CN114124171A
Intelligent environment control method for crop planting in plant factory
CN115016413A