Intelligent control unit with built-in deep reinforcement learning model for temperature and humidity control
By using an intelligent control unit with built-in deep reinforcement learning model in cell culture environment, the time delay and strong coupling problems in existing temperature and humidity control are solved, and high-precision, fast-responsive temperature and humidity control and reduction of power consumption are achieved.
Patent Information
- Application Number
- CN202310190804.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-02
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-02-02
AI Technical Summary
There are time lag and strong coupling problems in the existing cell culture environment, making it difficult to achieve ideal control effects.
The intelligent control unit with built-in deep reinforcement learning model controls the opening and working hours of humidifiers, dryers, refrigerators and heaters based on real-time temperature and humidity, target temperature and humidity, and predetermined threshold range.
High-precision and fast response temperature and humidity control are achieved, reducing power consumption, and reducing the temperature and humidity fluctuation range after reaching the target.
Smart Images

Figure CN116126064B_ABST
Abstract
Description
[0001] This patent application is a divisional application of Chinese Patent No. 202110145889.2, whose invention name is "Temperature and humidity control method and system for cell culture chamber", and its entire text is incorporated herein by reference. Technical Field
[0002] The present invention relates to a cell culture device, and in particular to an intelligent control unit with a built-in deep reinforcement learning model for temperature and humidity control. Background Art
[0003] Cell culture refers to a method of simulating the in vivo environment (sterility, suitable temperature, pH and certain nutritional conditions, etc.) in vitro to allow it to survive, grow, reproduce and maintain its main structure and function. Cell culture technology can transform a cell into a simple single cell or a few differentiated multi-cells through mass culture. This is an indispensable part of cloning technology, and cell culture itself is a clone of cells. Cell culture technology is an important and commonly used technology in cell biology research methods. Through cell culture, a large number of cells can be obtained, and it can also be used to study cell signal transduction, cell anabolism, cell growth and proliferation, etc.
[0004] Take embryonic cell culture as an example. The culture of embryos has very strict requirements on environmental temperature and humidity. When the temperature is too low, the metabolic activity of the embryo decreases, the growth and classification are slow, or even death occurs, causing the cells to coagulate. When the temperature is too high, it causes the inactivation of enzymes, destroys lipids and nuclear division, produces coagulase, and denatures proteins. When the humidity is too high, it is easy to condense into small droplets and fall into the culture dish, contaminating the culture fluid. When the humidity is too low, the culture fluid is easy to volatilize, destroying the internal environment of cell culture. Therefore, a suitable temperature and humidity environment is crucial to the quality of cell culture.
[0005] The existing cell culture environment temperature and humidity joint control adopts conventional controllers, and conventional controllers have problems such as time lag and strong coupling, which are specifically manifested in: the heating of the heating tube will cause the temperature of a specified area of the incubator to change, and the water vapor content in the air will also change accordingly after heating. Similarly, although the humidification tube only plays a humidification role, it will also affect the temperature inside the box. The existing technology has the following defects: 1) The existing PID control technology actually regards temperature and humidity as two independent and unrelated invariant systems, and does not consider the coupling between temperature and humidity, so it is difficult to achieve a more ideal control purpose; 2) In addition, the PID control has a large overshoot, and it is difficult to meet higher requirements for accuracy and fluctuation; 3) Environmental modeling is very difficult, and it is difficult to fit complex environments based on a priori assumptions. The system transfer function and state function are difficult to fit complex environments.
[0006] Therefore, it is necessary to study an intelligent control unit with a built-in deep reinforcement learning model for temperature and humidity control to solve one or more of the above-mentioned technical problems. Summary of the invention
[0007] In order to solve at least one of the above technical problems, according to one aspect of the present invention, an intelligent control unit with a built-in deep reinforcement learning model for temperature and humidity control is provided, characterized in that the intelligent control unit controls the actuator to turn on or off at least one of the humidifier, dryer, refrigerator and heater in the cell culture chamber according to the received real-time temperature and humidity, target temperature, target humidity and predetermined threshold range of the cell culture chamber, and controls the working time of at least one of the humidifier, dryer, refrigerator and heater;
[0008] The deep reinforcement learning model is obtained by the following method:
[0009] a. Set the objective function and constraints to be optimized for the deep reinforcement learning model. The objective function to be optimized is shown in formula (1), which means minimizing the power consumed and the time t used to reach the target stable state. 0 , where p i represents the power consumed by the components actually involved in the work, λ is the harmonic coefficient; the constraint condition is as shown in formula (2), which means that the fluctuation range of temperature and humidity is within the predetermined threshold range after reaching the target stable state, T best RH best They represent the target temperature and target humidity respectively; Δt and ΔRH represent the temperature and humidity fluctuation range respectively. temp(t>t 0 ) represents the temperature after reaching the target stable state, RHumity(t>t 0 ) represents the humidity after reaching the target stable state, and t is the current time;
[0010]
[0011]
[0012] b. Training the deep reinforcement learning model
[0013] b1 sets the total number of iterations N of the deep reinforcement learning model e , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c ;
[0014] b2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c , where θ ais the parameter of the Actor network, θ c is the parameter of the Critic network, s is the current ambient temperature and humidity input state, a is the execution action and is a row vector;
[0015] b3 starts the first iteration, and count K=1;
[0016] b3.1 Start the first exploration, and count n = 1;
[0017] b3.2 According to the current ambient temperature and humidity status s t , the Actor network will s t As input, it passes through the network function A(s; θ a )|s=s t Next, a set of execution actions a is generated t ;
[0018] b3.3 After executing a t After that, the environmental state of the cell culture chamber changed, and the temperature and humidity detection point found that the new state was s t+1 , according to formula (3) we get a timely reward r t , r t is Reward(t);
[0019]
[0020] Where M 1 ,M 2 ,M 3 ,M 4 are the penalty factors for each item respectively;
[0021] b3.4 a t and the current ambient temperature and humidity status t The joint is used as input to the Critic network, and then passes through C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t .
[0022] b3.5 Calculate the Actor network A(s; θ) according to formula (4) a ) parameter θ a Gradient And update the parameter θ a ,
[0023]
[0024] b3.6 Calculate the critic network C(s,a;θ) according to formula (5) c ) parameter θc The gradient of θ is updated c ,
[0025]
[0026] in, is Reward(t), calculated by formula (3);
[0027] b3.7 Environment status update completed t ←s t+1 ;
[0028] b3.8 Update the exploration count n←n+1;
[0029] b3.9 Re-execute process b3.2-b3.8 until n>T, completing this exploration process;
[0030] b4 Update the iteration count, K←K+1;
[0031] b5 Re-execute b3.1-b3.9 and b4 until K>N e , complete the deep reinforcement learning model
[0032] DRL training;
[0033] c. Place the trained deep reinforcement learning model into the intelligent control unit.
[0034] According to another aspect of the present invention, there are multiple cell culture chambers, each of which is independent of each other and controlled by a separate intelligent control unit.
[0035] According to another aspect of the present invention, there are multiple cell culture chambers, each of which is independent of each other, and the intelligent control unit performs control according to the priority of each cell culture chamber.
[0036] According to another aspect of the present invention, gases from a humidifier, a dryer, a refrigerator and / or a heater are mixed via a mixing chamber and then input into the one or more cell culture chambers.
[0037] According to another aspect of the present invention, the humidifier, dryer, refrigerator and heater are connected to each cell culture chamber via independent pipelines.
[0038] According to another aspect of the present invention, the power consumption of the humidifier, dryer, refrigerator, and heater from the start to the stable state is p i Calculation formula:
[0039]
[0040] Where Ii (t),u i (t) represent the instantaneous current and voltage of each component respectively.
[0041] According to another aspect of the present invention, the Actor network has two input neurons, an intermediate layer and an output layer, and the two input neurons are represented by a row vector s=[s t ,s h ] indicates that each component in the row vector represents the temperature s of the current environmental state. t and relative humidity h ;
[0042] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains m i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0043] The output layer has 8 neurons, which are divided into two groups. The first group of 4 neurons represents the flag of the solenoid valve opening. The activation function is softmax, which is recorded as the row vector [flag 1 ,flag 2 ] and [flag 3 ,flag 4 ], indicating whether the solenoid valves of the humidifier and dryer are open, and whether the solenoid valves of the refrigerator and heater are open; the activation function of the second group of 4 neurons is linear y = x, and the 4 neural states are expressed by the row vector time = [time 1 ,time 2 ,time 3 ,time 4 ] indicates the opening time of the electromagnetic valve controlling the humidifier. 1 , the solenoid valve opening time of the dryer 2 , The solenoid valve of the refrigerator is open and running time 3 , heater solenoid valve opening time 4 .
[0044] According to another aspect of the present invention, the critic network has 10 input neurons, an intermediate layer and an output layer, wherein the 10 input neurons are respectively the temperature and relative humidity and the output of the Actor network, which are represented by a row vector as input = [s t ,s h ,flag 1 ,flag 2 ,flag 3,flag 4 ,time 1 ,time 2 ,time 3 ,time 4 ];
[0045] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains L i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0046] The output layer contains a linear neuron with an activation function of y=x, which evaluates the value of the Actor network action.
[0047] According to another aspect of the present invention, a temperature and humidity control system for a cell culture chamber is provided, characterized in that the cell culture chamber is connected to a humidifier, a dryer, a refrigerator and a heater through a gas channel, and the control system includes an intelligent control unit with a built-in deep reinforcement learning model, and the intelligent control unit controls an actuator to turn on or off at least one of the humidifier, the dryer, the refrigerator and the heater, and controls the working time of at least one of the humidifier, the dryer, the refrigerator and the heater according to the received real-time temperature and humidity of the cell culture chamber, the target temperature, the target humidity and the predetermined threshold range;
[0048] The deep reinforcement learning model is obtained by the following method:
[0049] a. Set the objective function and constraints to be optimized for the deep reinforcement learning model. The objective function to be optimized is shown in formula (1), which means minimizing the power consumed and the time t used to reach the target stable state. 0 , where p i represents the power consumed by the components actually involved in the work, λ is the harmonic coefficient; the constraint condition is as shown in formula (2), which means that the fluctuation range of temperature and humidity is within the predetermined threshold range after reaching the target stable state, T best RH best They represent the target temperature and target humidity respectively; Δt and ΔRH represent the temperature and humidity fluctuation range respectively. temp(t>t 0 ) represents the temperature after reaching the target stable state, RHumity(t>t 0 ) represents the humidity after reaching the target stable state, and t is the current time;
[0050]
[0051]
[0052] b. Training the deep reinforcement learning model
[0053] b1 sets the total number of iterations N of the deep reinforcement learning model e , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c ;
[0054] b2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c , where θ a is the parameter of the Actor network, θ c is the parameter of the Critic network, s is the current ambient temperature and humidity input state, a is the execution action and is a row vector;
[0055] b3 starts the first iteration, and count K=1;
[0056] b3.1 Start the first exploration, and count n = 1;
[0057] b3.2 According to the current ambient temperature and humidity status s t , the Actor network will s t As input, it passes through the network function A(s; θ a )|s=s t Next, a set of execution actions a is generated t ;
[0058] b3.3 After executing a t After that, the environmental state of the cell culture chamber changed, and the temperature and humidity detection point found that the new state was s t+1 , according to formula (5) we get a timely reward r t , r t is Reward(t);
[0059]
[0060] Where M 1 ,M 2 ,M 3 ,M 4 are the penalty factors for each item respectively;
[0061] b3.4 a t and the current ambient temperature and humidity status tCombined as input to the Critic network, after
[0062] C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t .
[0063] b3.5 Calculate the Actor network A(s; θ) according to formula (8) a ) parameter θ a Gradient And update the parameter θ a ,
[0064]
[0065] b3.6 Calculate the critic network C(s,a;θ) according to formula (9) c ) parameter θ c The gradient of θ is updated c ,
[0066]
[0067] in, is Reward(t), calculated by formula (5);
[0068] b3.7 Environment status update completed t ←s t+1 ;
[0069] b3.8 Update the exploration count n←n+1;
[0070] b3.9 Re-execute process b3.2-b3.8 until n>T, completing this exploration process;
[0071] b4 Update the iteration count, K←K+1;
[0072] b5 Re-execute b3.1-b3.9 and b4 until K>N e , complete the deep reinforcement learning model
[0073] DRL training;
[0074] c. Place the trained deep reinforcement learning model into the intelligent control unit.
[0075] According to another aspect of the present invention, there are multiple cell culture chambers, each of which is independent of each other and controlled by a separate intelligent control unit.
[0076] According to another aspect of the present invention, there are multiple cell culture chambers, each of which is independent of each other, and the intelligent control unit performs control according to the priority of each cell culture chamber.
[0077] According to another aspect of the present invention, gases from a humidifier, a dryer, a refrigerator and / or a heater are mixed via a mixing chamber and then input into the one or more cell culture chambers.
[0078] According to another aspect of the present invention, the humidifier, dryer, refrigerator and heater are connected to each cell culture chamber via independent pipelines.
[0079] According to another aspect of the present invention, the Actor network has two input neurons, an intermediate layer and an output layer, and the two input neurons are represented by a row vector s=[s t ,s h ] indicates that each component in the row vector represents the temperature s of the current environmental state. t and relative humidity h ;
[0080] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains m i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0081] The output layer has 8 neurons, which are divided into two groups. The first group of 4 neurons represents the flag of the solenoid valve opening. The activation function is softmax, which is recorded as the row vector [flag 1 ,flag 2 ] and [flag 3 ,flag 4 ], indicating whether the solenoid valves of the humidifier and dryer are open, and whether the solenoid valves of the refrigerator and heater are open; the activation function of the second group of 4 neurons is linear y = x, and the 4 neural states are expressed by the row vector time = [time 1 ,time 2 ,time 3 ,time 4 ] indicates the opening time of the electromagnetic valve controlling the humidifier. 1 , the solenoid valve opening time of the dryer 2 , The solenoid valve of the refrigerator is open and running time 3 , heater solenoid valve opening time 4 .
[0082] According to another aspect of the present invention, the critic network has 10 input neurons, an intermediate layer and an output layer, wherein the 10 input neurons are respectively the temperature and relative humidity and the output of the Actor network, which are represented by a row vector as input = [s t ,s h ,flag 1 ,flag 2 ,flag 3 ,flag 4 ,time 1 ,time 2 ,time 3 ,time 4 ];
[0083] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains L i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0084] The output layer contains a linear neuron with an activation function of y=x, which evaluates the value of the Actor network action.
[0085] The present invention can achieve one or more of the following technical effects:
[0086] 1. The deep learning model designed in the present invention can effectively solve the problem of strong coupling of temperature and humidity control according to the received real-time temperature and humidity of the cell culture chamber, the target temperature, the target humidity and the predetermined threshold range, and has high control accuracy and fast response;
[0087] 2. The target temperature and humidity can be quickly reached while minimizing power consumption;
[0088] 3. After reaching the target temperature and humidity, the fluctuation range of temperature and humidity can be reduced or minimized;
[0089] 4. The independent multi-chambers ensure that each culture activity does not interfere with each other, the environment is stable, and it supports a fixed temperature and humidity culture environment, which is more flexible. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0091] Figure 1 Schematic diagram of a temperature and humidity control system for a cell culture chamber according to a preferred embodiment of the present invention.
[0092] Figure 2FIG. 4 is a gas flow diagram of a cell culture chamber according to a preferred embodiment of the present invention.
[0093] Figure 3 The figure is a gas path diagram of the functional components and each culture chamber according to a preferred embodiment of the present invention.
[0094] Figure 4 An Actor network structure diagram of a deep learning model according to a preferred embodiment of the present invention.
[0095] Figure 5 A Critic network structure diagram of a deep learning model according to a preferred embodiment of the present invention.
[0096] Figure 6 This is a relationship diagram between a Critor network and an Actor network of a deep learning model according to a preferred embodiment of the present invention.
[0097] Figure 7 The present invention is a flowchart of a method for training a deep learning model according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0098] The best implementation mode of the present invention is described below through preferred embodiments in conjunction with the accompanying drawings. The specific implementation mode here is to illustrate the present invention in detail and should not be understood as limiting the present invention. Various deformations and modifications can be made without departing from the spirit and essential scope of the present invention, which should be included in the protection scope of the present invention.
[0099] Example 1
[0100] According to a preferred embodiment of the present invention, see Figure 1-7 , provides a method for controlling the temperature and humidity of a cell culture chamber, wherein the cell culture chamber is connected to a humidifier, a dryer, a refrigerator, and a heater through a gas channel, and the humidifier, the dryer, the refrigerator, and the heater are controlled by an intelligent control unit with a built-in deep reinforcement learning model. The control method comprises the following steps:
[0101] a. Set the objective function and / or constraints to be optimized for the deep reinforcement learning model. The objective function to be optimized is shown in formula (1), which means minimizing the power consumed and the time t used to reach the target steady state. 0 , where p i represents the power consumed by the components actually involved in the work, λ is the harmonic coefficient; the constraint condition is as shown in formula (2), which means that the fluctuation range of temperature and humidity is within the predetermined threshold range after reaching the target stable state, T best RH bestThey represent the target temperature and target humidity respectively; Δt and ΔRH represent the temperature and humidity fluctuation range respectively. temp(t>t 0 ) represents the temperature after reaching the target stable state, RHumity(t>t 0 ) represents the humidity after reaching the target stable state, and t is the current time;
[0102]
[0103]
[0104] b. Training the deep reinforcement learning model
[0105] b1 Figure 7 As shown, set the total number of iterations N of the deep reinforcement learning model e , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c ;
[0106] b2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c , where θ a is the parameter of the Actor network, θ c is the parameter of the Critic network, s is the current ambient temperature and humidity input state, a is the execution action and is a row vector;
[0107] b3 starts the first iteration, and count K=1;
[0108] b3.1 Start the first exploration, and count n = 1;
[0109] b3.2 According to the current ambient temperature and humidity status s t , the Actor network will s t As input, it passes through the network function A(s; θ a )|s=s t Next, a set of execution actions a is generated t ;
[0110] b3.3 After executing a t After that, the environmental state of the cell culture chamber changed, and the temperature and humidity detection point found that the new state was s t+1 , according to formula (5) we get a timely reward r t , r t is Reward(t);
[0111]
[0112] Where M 1 ,M 2 ,M 3 ,M 4 are the penalty factors for each item respectively;
[0113] b3.4 a t and the current ambient temperature and humidity status t The joint is used as input to the Critic network, and then passes through C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t .
[0114] b3.5 Calculate the Actor network A(s; θ) according to formula (8) a ) parameter θ a Gradient And update the parameter θ a ,
[0115]
[0116] b3.6 Calculate the critic network C(s,a;θ) according to formula (9) c ) parameter θ c The gradient of θ is updated c ,
[0117]
[0118] in, is Reward(t), calculated by formula (5);
[0119] b3.7 Environment status update completed t ←s t+1 ;
[0120] b3.8 Update the exploration count n←n+1;
[0121] b3.9 Re-execute process b3.2-b3.8 until n>T, completing this exploration process;
[0122] b4 Update the iteration count, K←K+1;
[0123] b5 Re-execute b3.1-b3.9 and b4 until K>N e , complete the training of the deep reinforcement learning model DRL;
[0124] c. placing the trained deep reinforcement learning model into the intelligent control unit, and the intelligent control unit controls the actuator to turn on or off at least one of the humidifier, dryer, refrigerator and heater, and controls the working time of at least one of the humidifier, dryer, refrigerator and heater according to the received real-time temperature and humidity of the cell culture chamber, the target temperature, the target humidity and the predetermined threshold range.
[0125] It can be understood that the power consumed to reach the target steady state and the time used t 0 The target temperature and humidity can be quickly reached while minimizing power consumption. When the culture chamber (culture chamber) is initially activated or when the chamber door is opened to place the embryos to be cultured, the temperature and humidity of the culture chamber often deviate from the target temperature and humidity.
[0126] Preferably, the multiple culture chambers are connected by independent control pipelines and temperature and humidity control components, and have independent gas flow environments, so that the temperature and humidity microenvironment of each culture chamber can be independent.
[0127] Preferably, the intelligent control unit can receive the preset ambient temperature and humidity target values from the main control system, receive the environmental parameter information of the temperature and humidity detection points in real time during operation, and output precise control instructions to control the opening and closing of the actuator and the working time parameters. The actuator can control the working state (open or close) and working time of the temperature and humidity control component, which generally includes a heater, a refrigerator, a humidifier, and a dryer.
[0128] Preferably, a main control system can also be set, which is a type of controller that can realize system logic control and data processing, such as ARM, etc. The main control system can select the number of the culture chamber to be used and set the target temperature and humidity value of the chamber. The main control system can also set the priority of multi-user culture.
[0129] Preferably, the main control system can accept customized environmental temperature and humidity culture requirements, allowing users to set temperature and humidity parameters and dynamic fluctuation range by themselves, so it is more flexible. In addition, each chamber of the incubator can operate under different temperature and humidity conditions, and can culture different types of cells to meet the culture needs of multiple users.
[0130] According to another preferred embodiment of the present invention, when controlling to minimize the power consumption and usage time to achieve the target temperature and humidity, the objective function to be optimized of the deep reinforcement learning model can be set separately. Accordingly, Reward(t)=-(M 2 |temp(t)-T best |+M 4 |RHumity(t)-RH best|), thereby providing a method for controlling the cell culture chamber to quickly and energy-savingly reach the target temperature and humidity; or, when controlling to reduce the fluctuation range after reaching a stable state, the constraints of the deep reinforcement learning model can be selected to provide a method for controlling the temperature and humidity fluctuation range of the cell culture chamber.
[0131] Specifically, a method for controlling a cell culture chamber to quickly and energy-efficiently reach a target temperature and humidity is provided, wherein the cell culture chamber is connected to a humidifier, a dryer, a refrigerator, and a heater through a gas channel, and the humidifier, the dryer, the refrigerator, and the heater are controlled by an intelligent control unit with a built-in deep reinforcement learning model. The control method comprises the following steps:
[0132] a. Set the objective function to be optimized for the deep reinforcement learning model. The objective function to be optimized is shown in formula (1), which means minimizing the power consumed and the time t used to reach the target stable state. 0 , where p i It represents the power consumed by the components actually involved in the work, and λ is the harmonic coefficient;
[0133]
[0134] b. Training the deep reinforcement learning model
[0135] b1 Figure 7 As shown, set the total number of iterations N of the deep reinforcement learning model e , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c ;
[0136] b2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c , where θ a is the parameter of the Actor network, θ c is the parameter of the Critic network, s is the current ambient temperature and humidity input state, a is the execution action and is a row vector;
[0137] b3 starts the first iteration, and count K=1;
[0138] b3.1 Start the first exploration, and count n = 1;
[0139] b3.2 According to the current ambient temperature and humidity status s t , the Actor network will st As input, it passes through the network function A(s; θ a )|s=s t Next, a set of execution actions a is generated t ;
[0140] b3.3 After executing a t After that, the environmental state of the cell culture chamber changed, and the temperature and humidity detection point found that the new state was s t+1 , according to formula (5) we get a timely reward r t , r t is Reward(t);
[0141] Reward(t)=-(M 2 |temp(t)-T best |+M 4 |RHumity(t)-RH best |) (5)
[0142] Where M 2 、M 4 are the penalty factors for each item, T best RH best Respectively represent the set target temperature and target humidity; temp(t) represents the current temperature, RHumity(t) represents the current humidity, and t represents the current time;
[0143] b3.4 a t and the current ambient temperature and humidity status t Combined as input to the Critic network, after
[0144] C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t .
[0145] b3.5 Calculate the Actor network A(s; θ) according to formula (8) a ) parameter θ a Gradient And update the parameter θ a ,
[0146]
[0147] b3.6 Calculate the critic network C(s,a;θ) according to formula (9) c ) parameter θ c The gradient of θ is updated c ,
[0148]
[0149] in, is Reward(t), calculated by formula (5);
[0150] b3.7 Environment status update completed t ←s t+1 ;
[0151] b3.8 Update the exploration count n←n+1;
[0152] b3.9 Re-execute process b3.2-b3.8 until n>T, completing this exploration process;
[0153] b4 Update the iteration count, K←K+1;
[0154] b5 Re-execute b3.1-b3.9 and b4 until K>N e , complete the deep reinforcement learning model
[0155] DRL training;
[0156] c. placing the trained deep reinforcement learning model into the intelligent control unit, and the intelligent control unit controls the actuator to turn on or off at least one of the humidifier, dryer, refrigerator and heater according to the received real-time temperature and humidity of the cell culture chamber, the target temperature and the target humidity, and controls the working time of at least one of the humidifier, dryer, refrigerator and heater.
[0157] According to another preferred embodiment of the present invention, a method for controlling the temperature and humidity fluctuation range of a cell culture chamber is also provided, characterized in that the cell culture chamber is connected to a humidifier, a dryer, a refrigerator and a heater through a gas channel, and the humidifier, the dryer, the refrigerator and the heater are controlled by an intelligent control unit with a built-in deep reinforcement learning model. The control method comprises the following steps:
[0158] a. Set the constraints of the deep reinforcement learning model. The constraints are shown in formula (2), which means that the fluctuation range of temperature and humidity is within the predetermined threshold range after reaching the target stable state, T best RH best They represent the target temperature and target humidity respectively; Δt and ΔRH represent the temperature and humidity fluctuation range respectively. temp(t>t 0 ) represents the temperature after reaching the target stable state, RHumity(t>t 0 ) represents the humidity after reaching the target stable state, and t is the current time;
[0159]
[0160] b. Training the deep reinforcement learning model. The specific method can be found in the training method in the aforementioned temperature and humidity control method of the cell culture chamber, which is omitted here.
[0161] c. placing the trained deep reinforcement learning model into the intelligent control unit, and the intelligent control unit controls the actuator to turn on or off at least one of the humidifier, dryer, refrigerator and heater, and controls the working time of at least one of the humidifier, dryer, refrigerator and heater according to the received real-time temperature and humidity of the cell culture chamber, the target temperature, the target humidity and the predetermined threshold range.
[0162] According to another preferred embodiment of the present invention, see Figure 1 There are multiple cell culture chambers, each of which is independent of each other and controlled by a separate intelligent control unit.
[0163] According to another preferred embodiment of the present invention, there are multiple cell culture chambers, each of which is independent of each other, and the intelligent control unit performs control according to the priority of each cell culture chamber.
[0164] According to another preferred embodiment of the present invention, see Figure 2 The gases from the humidifier, dryer, refrigerator and / or heater are mixed in a mixing chamber and then input into the one or more cell culture chambers.
[0165] Preferably, when the temperature and humidity at the temperature and humidity detection point do not meet the set expected values, the intelligent control unit makes a decision that several functional components need to work and last for different periods of time. The air intake pump control point, the exhaust pump control point and the incubator air intake pump control point are turned on, and the gas can enter the several functional components from the air intake pump, and then be output from the exhaust pump to the mixing chamber, and then enter the culture chamber from the mixing chamber, and the gas in the culture chamber enters the air intake pump again, and the above cycle is repeated until the concentration at the temperature and humidity detection point meets the requirements, and the air intake pump control point, the exhaust pump control point and the incubator air intake pump control point are immediately closed.
[0166] Preferably, when multiple users need to use the culture chamber at the same time, if the same temperature and humidity environment is used, the system will treat the gas path environment of each culture chamber as a whole and uniformly regulate it. The actions of each air intake pump control point, exhaust pump control point and incubator air intake pump control point will be consistent, and a balanced state can be reached quickly. If different temperature and humidity environments are used, the microenvironment temperature and humidity adjustment will have a priority order according to the priority, and the priority setting can be set by the main control system. Once the temperature and humidity of the current culture chamber microenvironment reaches equilibrium, the current air intake pump control point, exhaust pump control point and incubator air intake pump control point will be closed, and the next culture chamber microenvironment temperature and humidity adjustment will be turned on.
[0167] According to another preferred embodiment of the present invention, see Figure 3 The humidifier, dryer, refrigerator and heater are connected to each cell culture chamber through independent pipelines.
[0168] According to another preferred embodiment of the present invention, the power consumption generated by the humidifier, dryer, refrigerator, and heater from the start to the stable state is p i Calculation formula:
[0169]
[0170] Where I i (t),u i (t) represent the instantaneous current and voltage of each component respectively.
[0171] According to another preferred embodiment of the present invention, see Figure 4 , the Actor network has two input neurons, an intermediate layer and an output layer. The two input neurons are represented by a row vector s = [s t ,s h ] indicates that each component in the row vector represents the temperature s of the current environmental state. t and relative humidity h ;
[0172] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains m i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0173] The output layer has 8 neurons, which are divided into two groups. The first group of 4 neurons represents the flag of the solenoid valve opening. The activation function is softmax, which is recorded as the row vector [flag 1 ,flag 2 ] and [flag 3 ,flag4 ], indicating whether the solenoid valves of the humidifier and dryer are open, and whether the solenoid valves of the refrigerator and heater are open; the activation function of the second group of 4 neurons is linear y = x, and the 4 neural states are expressed by the row vector time = [time 1 ,time 2 ,time 3 ,time 4 ] indicates the opening time of the electromagnetic valve controlling the humidifier. 1 , the solenoid valve opening time of the dryer 2 , The solenoid valve of the refrigerator is open and running time 3 , heater solenoid valve opening time 4 .
[0174] According to another preferred embodiment of the present invention, see Figure 5-6 , the critic network has 10 input neurons, an intermediate layer and an output layer. The 10 input neurons are temperature and relative humidity, and the output of the Actor network is represented by a row vector as input = [s t ,s h ,flag 1 ,flag 2 ,flag 3 ,flag 4 ,time 1 ,time 2 ,time 3 ,time 4 ];
[0175] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains L i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0176] The output layer contains a linear neuron with an activation function of y=x, which evaluates the value of the Actor network action.
[0177] According to another preferred embodiment of the present invention, a temperature and humidity control system for a cell culture chamber is provided, characterized in that the cell culture chamber is connected to a humidifier, a dryer, a refrigerator and a heater through a gas channel, and the control system includes an intelligent control unit with a built-in deep reinforcement learning model, and the intelligent control unit controls an actuator to turn on or off at least one of the humidifier, the dryer, the refrigerator and the heater according to the received real-time temperature and humidity of the cell culture chamber, the target temperature, the target humidity and the predetermined threshold range, and controls the working time of at least one of the humidifier, the dryer, the refrigerator and the heater;
[0178] The deep reinforcement learning model is obtained by the following method:
[0179] a. Set the objective function and constraints to be optimized for the deep reinforcement learning model. The objective function to be optimized is shown in formula (1), which means minimizing the power consumed and the time t used to reach the target stable state. 0 , where p i represents the power consumed by the components actually involved in the work, λ is the harmonic coefficient; the constraint condition is as shown in formula (2), which means that the fluctuation range of temperature and humidity is within the predetermined threshold range after reaching the target stable state, T best RH best They represent the target temperature and target humidity respectively; Δt and ΔRH represent the temperature and humidity fluctuation range respectively. temp(t>t 0 ) represents the temperature after reaching the target stable state, RHumity(t>t 0 ) represents the humidity after reaching the target stable state, and t is the current time;
[0180]
[0181]
[0182] b. Training the deep reinforcement learning model
[0183] b1 sets the total number of iterations N of the deep reinforcement learning model e , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c ;
[0184] b2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c , where θ ais the parameter of the Actor network, θ c is the parameter of the Critic network, s is the current ambient temperature and humidity input state, a is the execution action and is a row vector;
[0185] b3 starts the first iteration, and count K=1;
[0186] b3.1 Start the first exploration, and count n = 1;
[0187] b3.2 According to the current ambient temperature and humidity status s t , the Actor network will s t As input, it passes through the network function A(s; θ a )|s=s t Next, a set of execution actions a is generated t ;
[0188] b3.3 After executing a t After that, the environmental state of the cell culture chamber changed, and the temperature and humidity detection point found that the new state was s t+1 , according to formula (5) we get a timely reward r t , r t is Reward(t);
[0189]
[0190] Where M 1 ,M 2 ,M 3 ,M 4 are the penalty factors for each item respectively;
[0191] b3.4 a t and the current ambient temperature and humidity status t The joint is used as input to the Critic network, and then passes through C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t .
[0192] b3.5 Calculate the Actor network A(s; θ) according to formula (8) a ) parameter θ a Gradient And update the parameter θ a ,
[0193]
[0194] b3.6 Calculate the critic network C(s,a;θ) according to formula (9) c ) parameter θc The gradient of θ is updated c ,
[0195]
[0196] in, is Reward(t), calculated by formula (5);
[0197] b3.7 Environment status update completed t ←s t+1 ;
[0198] b3.8 Update the exploration count n←n+1;
[0199] b3.9 Re-execute process b3.2-b3.8 until n>T, completing this exploration process;
[0200] b4 Update the iteration count, K←K+1;
[0201] b5 Re-execute b3.1-b3.9 and b4 until K>N e , complete the deep reinforcement learning model
[0202] DRL training;
[0203] c. Place the trained deep reinforcement learning model into the intelligent control unit.
[0204] According to another preferred embodiment of the present invention, there are multiple cell culture chambers, each of which is independent of each other and controlled by a separate intelligent control unit.
[0205] According to another preferred embodiment of the present invention, there are multiple cell culture chambers, each of which is independent of each other, and the intelligent control unit performs control according to the priority of each cell culture chamber.
[0206] According to another preferred embodiment of the present invention, the gases from the humidifier, the dryer, the refrigerator and / or the heater are mixed in a mixing chamber and then input into the one or more cell culture chambers.
[0207] According to another preferred embodiment of the present invention, the humidifier, dryer, refrigerator and heater are connected to each cell culture chamber via independent pipelines.
[0208] According to another preferred embodiment of the present invention, the Actor network has two input neurons, an intermediate layer and an output layer. The two input neurons are represented by a row vector s=[s t ,s h ] indicates that each component in the row vector represents the temperature s of the current environmental state.t and relative humidity h ;
[0209] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains m i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0210] The output layer has 8 neurons, which are divided into two groups. The first group of 4 neurons represents the flag of the solenoid valve opening. The activation function is softmax, which is recorded as the row vector [flag 1 ,flag 2 ] and [flag 3 ,flag 4 ], indicating whether the solenoid valves of the humidifier and dryer are open, and whether the solenoid valves of the refrigerator and heater are open; the activation function of the second group of 4 neurons is linear y = x, and the 4 neural states are expressed by the row vector time = [time 1 ,time 2 ,time 3 ,time 4 ] indicates the opening time of the electromagnetic valve controlling the humidifier. 1 , the solenoid valve opening time of the dryer 2 , The solenoid valve of the refrigerator is open and running time 3 , heater solenoid valve opening time 4 .
[0211] According to another preferred embodiment of the present invention, the critic network has 10 input neurons, an intermediate layer and an output layer, wherein the 10 input neurons are respectively the temperature and relative humidity and the output of the Actor network, which are represented by a row vector as input = [s t ,s h ,flag 1 ,flag 2 ,flag 3 ,flag 4 ,time 1 ,time 2 ,time 3 ,time 4 ];
[0212] The middle layer has several hidden layers, which are fully connected. Each hidden layer contains L ihidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer;
[0213] The output layer contains a linear neuron with an activation function of y=x, which evaluates the value of the Actor network action.
[0214] Example 2
[0215] This embodiment further describes the present invention in detail through examples based on Embodiment 1.
[0216] This embodiment provides a temperature and humidity control system for a cell culture chamber, which is divided into four parts: environment, intelligent control unit, actuator, and main control system.
[0217] 1) Environment is an abstract concept, which specifically includes all spaces through which temperature, humidity and gas circulate.
[0218] S11, such as Figure 1 As shown, the gas starts from each culture chamber to the temperature and humidity environmental monitoring point, then enters the temperature and humidity control component through each air intake pump, and then enters the culture chamber. The gas experiences the environment along the way.
[0219] S12. Each culture chamber has its own independent environment, and the temperature and humidity control components have independent pipes connecting each culture chamber. Even if multiple users use it at the same time, there will be no environmental fusion, and the environment can still be independent.
[0220] S13. The exhaust port of each chamber is provided with a temperature and humidity detection point for detecting the environmental parameter value and transmitting it to the intelligent control unit as important information.
[0221] S14. When multiple users use it at the same time and the preset values of the culture environment temperature and humidity are consistent, the various microenvironments will be integrated to accelerate the overall culture environment to reach a steady state of temperature and humidity.
[0222] 2) Intelligent control unit, which is a control unit with a built-in deep reinforcement learning model (DRL) and is solidified in the controller. The controller has a minimum operating system and has functions such as system information input, logic control, and data processing, such as STM32 microcontrollers.
[0223] S21, intelligent control unit, which can receive the culture environment concentration information preset by the main control system and use it as the final control target to meet the cultivation needs of diversified scenarios;
[0224] S22, when the intelligent control unit adjusts the environment, it needs to receive the temperature and humidity of the environment in real time after each execution structure generates an action. The built-in DRL model uses this as input, and the Actor network in the DRL makes precise adjustments.
[0225] S23. When multiple users are being trained, the intelligent control unit receives the control sequence instruction initiated by the main control system and decides whether it is its turn in the priority ranking. If so, it starts the temperature and humidity control. If not, it continues to wait for the next control sequence instruction.
[0226] 3) Actuator: It is the executor of the best decision made by the intelligent control unit after each environmental assessment. It mainly controls the solenoid valve through relays, can open the solenoid valve of each control node and adjust the on and off time of the temperature and humidity functional components.
[0227] S31. Functional components for regulating temperature and humidity generally include heaters, refrigerators, humidifiers, and dryers. They can be separate components or integrated components.
[0228] S32, the solenoid valve of the functional component for controlling temperature and the solenoid valve of the temperature intake pump control point operate synchronously, and similarly, the solenoid valve of the functional component for controlling humidity and the solenoid valve of the humidity intake pump control point operate synchronously.
[0229] 4) Main control system, which is a type of controller, including the minimum operating system, with functions such as system information input, logic control, data processing, etc., such as STM32 microcontroller.
[0230] S41. The controller of the main control system and the controller of the intelligent control unit with a built-in deep reinforcement learning model are connected through a bus to transmit the set temperature and humidity values to the intelligent control unit.
[0231] S42. The main control system can accept the temperature and humidity information of the culture chamber to be cultured set by the user and the priority of each culture chamber in the case of multi-user use.
[0232] S43, after completing the control, the intelligent control unit transmits information to the main control system and informs the main control system. In the case of multi-user use, the current priority is released based on this, and the environmental control task of the next priority culture chamber is started.
[0233] Preferably, see Figure 2 , which is the gas path condition of the culture chamber 1. When the temperature and humidity of the microenvironment in the culture chamber 1 do not reach the predetermined target, the intelligent control unit makes corresponding decisions based on the current temperature and humidity information to drive and control the actuators, that is, the various functional components.
[0234] S1. When the temperature and humidity are lower than the expected values, the intelligent control unit makes a decision that the heater and humidifier need to work and last for different periods of time. Control points 1, 2 and 3 are opened, and the gas can enter the heating and humidification functional components from the air intake pump. After a certain period of work, it is output from the exhaust pump to the mixing chamber, and then from the mixing chamber to the culture chamber 1. The gas in the culture chamber 1 enters this cycle again until the temperature and humidity detection reaches the expected value, and control points 1, 2, and 3 will be closed. At this time, it is considered that the temperature and humidity adjustment is completed.
[0235] S2. When the temperature and humidity are higher than the expected values, the intelligent control unit makes a decision that the refrigerator and dryer need to work and last for different periods of time. Control points 1, 2 and 3 are opened, and the gas can enter the refrigeration and drying functional components from the intake pump. After a certain period of work, it is output from the exhaust pump to the mixing chamber, and then from the mixing chamber to the culture chamber 1. The gas in the culture chamber 1 enters this cycle again until the temperature and humidity detection reaches the expected value, and control points 1, 2, and 3 will be closed. At this time, it is considered that the temperature and humidity adjustment is completed.
[0236] S3, when the temperature is higher and the humidity is lower than the expected value, the intelligent control unit makes a decision that the refrigerator and humidifier need to work and last for different periods of time. Control points 1, 2 and 3 are opened, and the gas can enter the refrigeration and humidification functional components from the intake pump. After a certain period of work, it is output from the exhaust pump to the mixing chamber, and then from the mixing chamber to the culture chamber 1. The gas in the culture chamber 1 enters this cycle again until the temperature and humidity detection reaches the expected value, and control points 1, 2, and 3 will be closed. At this time, it is considered that the temperature and humidity adjustment is completed.
[0237] S4. When the temperature is lower than and the humidity is higher than the expected value, the intelligent control unit makes a decision that the heater and dryer need to work and last for different periods of time. Control points 1, 2 and 3 are opened, and the gas can enter the heating and drying functional components from the air intake pump. After a certain period of work, it is output from the exhaust pump to the mixing chamber, and then from the mixing chamber to the culture chamber 1. The gas in the culture chamber 1 enters this cycle again until the temperature and humidity detection reaches the expected value, and control points 1, 2, and 3 will be closed. At this time, it is considered that the temperature and humidity adjustment is completed.
[0238] Preferably, each intelligent control unit has a built-in deep reinforcement learning model, which can, for example, perform offline training and online environmental temperature and humidity control. In combination with the specific requirements of the control task, here is an example of a deep reinforcement learning model DRL for the environmental temperature and humidity control of an incubator containing four culture chambers. The scenario assumes that there are four temperature and humidity control components in the incubator, namely a humidifier, a dryer, a refrigerator, and a heater. Then the optimization objective function of the DRL model becomes as shown in formula (3).
[0239] min p 1 +p2 +p 3 +p 4 +λt 0 (1)
[0240] p in the formula 1 、p 2 、p 3 、p 4 They are the power consumption of the above components from the start-up state to the stable state:
[0241]
[0242] Where I(t) and u(t) represent the instantaneous current and voltage of each component respectively.
[0243] The feedback function of DRL is defined as formula (5):
[0244]
[0245] Where M 1 ,M 2 ,M 3 ,M 4 are the penalty factors for each item, which measure the weight of each item. It should be noted that when the intelligent control unit does not reach a steady state, the corresponding time period is t<t 0 , then Reward(t)=-(M 2 |temp(t)-T best |+M 4 |RHumity(t)-RH best |), the goal is to force the environment to reach the specified temperature and humidity values as quickly as possible. After reaching the steady state, it is not only ensured to continue in the steady state, but also required that the fluctuation range is within a reasonable range.
[0246] The DRL model consists of an Actor network and a Critic network. The design of the Actor network is as follows:
[0247] S1.1.1. See Figure 5 , the Actor network in the DRL model has two input neurons, which can be represented by the row vector s = [s t ,s h ] indicates that each component in the row vector represents the current environmental state quantity temperature s t and relative humidity h .
[0248] S1.1.2, the middle layer has several hidden layers, which can be fully connected, and each layer contains m ihidden layer neurons, where i represents the hidden layer number, and its activation function is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer.
[0249] S1.1.3. The output layer has 8 neurons, which are divided into two groups in total.
[0250] S1.1.3.1 The first group of 4 neurons represents the flag of the solenoid valve opening. The activation function is softmax. They are grouped in groups of two, and each group is recorded as a row vector [flag 1 ,flag 2 ], each component in the row vector represents whether the solenoid valve of the humidifier and the dryer is turned on, and the other group is recorded as the row vector [flag 3 ,flag 4 ], each component in the row vector represents whether the solenoid valve of the refrigerator and heater is open.
[0251] S1.1.3.2 Another group of 4 neurons has a linear activation function y = x. The 4 neural states can be expressed by the row vector t = [time 1 ,time 2 ,time 3 ,time 4 ] indicates that each component in the row vector represents the opening time of the electromagnetic valve controlling the humidifier. 1 , the solenoid valve opening time of the dryer 2 , The solenoid valve of the refrigerator is open and running time 3 , heater solenoid valve opening time 4 .
[0252] See also Figure 6 , the Critic network is designed as follows:
[0253] S1.2.1. In combination with this task, the critic network needs to have 10 input neurons, which are the environmental state variables temperature and relative humidity and the output of the Actor network, represented by a row vector as input = [s t ,s h ,flag 1 ,flag 2 ,flag 3 ,flag 4 ,time 1 ,time 2 ,time 3 ,time 4 ].
[0254] S1.2.2, the middle layer has several hidden layers, which can be fully connected. Each layer contains L i hidden layer neurons, where i represents the hidden layer number, and its activation function is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer.
[0255] S1.2.3. The output layer contains a linear neuron with an activation function of y=x, which evaluates the value of the Actor network action.
[0256] Preferably, see Figure 6 , according to the current ambient temperature and humidity input state s t , Actor passes through the network mapping function A(s; θ a ) produces row a, a is a row vector, where θ a Represents the Actor network model parameters. Actor network output a and current ambient temperature and humidity input state s t Together as the Critic network C(s,a;θ c ) input, where θ c is the parameter state and outputs the evaluation value.
[0257] Furthermore, the Critic and the Actor work together to obtain the optimal deterministic strategy by solving the following joint optimization problem. The optimal parameters of the DRL model are obtained as shown in formulas (6)-(7). In formula (7): is the reward for the current state s and executing action a. It is calculated by formula (5).
[0258]
[0259]
[0260] The Actor network strives to maximize the evaluation of the Critic network, while the Critic network strives to make an accurate evaluation. The objective functions of the Actor network and the Critic network are both differentiable. By taking the derivatives of formulas (6)-(7) and using the chain rule, their gradients can be obtained, as shown in formulas (8)-(9).
[0261]
[0262]
[0263] Preferably, see Figure 7 , the training process of DRL is as follows:
[0264] S2.5.1 Set the total number of model iterations Ne , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c .
[0265] S2.5.2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c .
[0266] S2.5.3 starts the first iteration, and counts K=1.
[0267] S2.5.3.1 Start the first exploration, and count n=1.
[0268] S2.5.3.2 According to the current ambient temperature and humidity status t , the Actor network will s t As input, it passes through the network function A(s; θ a )|s=s t Next, a set of output actions a is generated t .
[0269] S2.5.3.3 The actuator completes the execution of a t After that, the environmental state changed, and the temperature and humidity detection point found that the new state was s t+1 , according to formula (5) we get a timely reward r t .
[0270] S2.5.3.4a t and the current ambient temperature and humidity status t The joint is used as input to the Critic network, and then passes through C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t .
[0271] S2.5.3.5 Calculate the Actor network A(s; θ) according to formula (8) a ) parameter θ a And update the parameter θ a ,
[0272] S2.5.3.6 Calculate the critic network C(s,a;θ) according to formula (9) c ) parameter θ c And update the parameter θ c ,
[0273] S2.5.3.7 At this time, the environment status is updated t ←s t+1 .
[0274] After S2.5.3.8 is completed, the exploration count is updated n←n+1.
[0275] S2.5.3.9 Re-execute process S2.5.3.2-S2.5.3.8 until n>T, completing this exploration process.
[0276] S2.5.4 Update the iteration count, K←K+1.
[0277] S2.5.5 Re-execute S2.5.3.1-S2.5.3.9 and S2.5.4 until K>N e , complete the DRL training.
[0278] Preferably, the real-time control of DRL is as follows:
[0279] S3.1 Once the DRL model is trained, the model structure and parameters are solidified on the control chip of the intelligent control unit.
[0280] When S3.2 is working, the intelligent control unit will generate the optimal control output according to the ambient temperature and humidity state s received in real time, and the Actor network will generate the optimal control output, as shown in formula (10).
[0281] a * =A(s;θ a ) (10)
[0282] S3.3 If the control uses the "ε-greedy strategy", on the basis of formula (10), add a random disturbance such as formula (12), and the final optimal control is as shown in formula (11). It should be noted that n in formula (11) represents random noise, and P max For adjustable boundaries, ensure the random intensity.
[0283] a=A(s;θ a )+n (11)
[0284] n~U(-P max ,P max ) (12)
[0285] The present invention can achieve one or more of the following technical effects:
[0286] 1. The deep learning model designed in the present invention can effectively solve the problem of strong coupling of temperature and humidity control according to the received real-time temperature and humidity of the cell culture chamber, the target temperature, the target humidity and the predetermined threshold range, and has high control accuracy and fast response;
[0287] 2. The target temperature and humidity can be quickly reached while minimizing power consumption;
[0288] 3. After reaching the target temperature and humidity, the fluctuation range of temperature and humidity can be reduced or minimized;
[0289] 4. The independent multi-chambers ensure that each culture activity does not interfere with each other, the environment is stable, and it supports a fixed temperature and humidity culture environment, which is more flexible.
[0290] It can be understood that the features in the above-mentioned embodiments can be combined with each other to produce new embodiments.
[0291] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of the present invention to be protected is defined by the attached claims and their equivalents.
Claims
1. An intelligent control unit with built-in deep reinforcement learning model for temperature and humidity control, Features The intelligent control unit controls the actuator to turn on or off at least one of a humidifier, a dryer, a refrigerator and a heater in the cell culture chamber according to the received real-time temperature and humidity, target temperature, target humidity and a predetermined threshold range of the cell culture chamber, and controls the working time of at least one of the humidifier, the dryer, the refrigerator and the heater; The deep reinforcement learning model is obtained by the following method: a. Set the objective function and constraints to be optimized for the deep reinforcement learning model. The objective function to be optimized is shown in formula (1), which means minimizing the power consumed and the time t used to reach the target stable state. 0 , where p i represents the power consumed by the components actually involved in the work, λ is the harmonic coefficient; the constraint condition is as shown in formula (2), which means that the fluctuation range of temperature and humidity is within the predetermined threshold range after reaching the target stable state, T best RH best They represent the target temperature and target humidity respectively; Δt and ΔRH represent the temperature and humidity fluctuation range respectively. temp(t>t 0 ) represents the temperature after reaching the target stable state, RHumity(t>t 0 ) represents the humidity after reaching the target stable state, and t is the current time; b. Training the deep reinforcement learning model b1 sets the total number of iterations N of the deep reinforcement learning model e , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c ; b2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c , where θ a is the parameter of the Actor network, θ c is the parameter of the Critic network, s is the current ambient temperature and humidity input state, a is the execution action and is a row vector; b3 starts the first iteration, and count K=1; b3.1 Start the first exploration, and count n = 1; b3.2 According to the current ambient temperature and humidity status s t , the Actor network will s t As input, it passes through the network function A(s; θ a )|s=s t Next, a set of execution actions a is generated t ; b3.3 After executing a t After that, the environmental state of the cell culture chamber changed, and the temperature and humidity detection point found that the new state was s t +1 , according to formula (3) we get a timely reward r t , r t is Reward(t); Where M 1 ,M 2 ,M 3 ,M 4 are the penalty factors for each item respectively; b3.4 a t and the current ambient temperature and humidity status t The joint is used as input to the Critic network, and then passes through C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t ; b3.5 Calculate the Actor network A(s; θ) according to formula (4) a ) parameter θ a Gradient And update the parameter θ a , b3.6 Calculate the critic network C(s,a;θ) according to formula (5) c ) parameter θ c The gradient of θ is updated c , in, is Reward(t), calculated by formula (3); b3.7 Environment status update completed t ←s t+1 ; b3.8 Update the exploration count n←n+1; b3.9 Re-execute process b3.2-b3.8 until n>T, completing this exploration process; b4 Update the iteration count, K←K+1; b5 Re-execute b3.1-b3.9 and b4 until K>N e , complete the training of the deep reinforcement learning model DRL; c. Place the trained deep reinforcement learning model into the intelligent control unit.
2. The intelligent control unit according to claim 1, Features The power consumption of humidifier, dryer, refrigerator and heater from start to stable state i Calculation formula: Where I i (t),u i (t) represent the instantaneous current and voltage of each component respectively.
3. The intelligent control unit according to claim 1 or 2, Features The Actor network has two input neurons, an intermediate layer and an output layer. The two input neurons are represented by a row vector s = [s t ,s h ] indicates that each component in the row vector represents the temperature s of the current environmental state. t and relative humidity h ; The middle layer has several hidden layers, which are fully connected. Each hidden layer contains m i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer; The output layer has 8 neurons, which are divided into two groups. The first group of 4 neurons represents the flag of the solenoid valve opening. The activation function is softmax, which is recorded as the row vector [flag 1 ,flag 2 ] and [flag 3 ,flag 4 ], indicating whether the solenoid valves of the humidifier and dryer are open, and whether the solenoid valves of the refrigerator and heater are open; the activation function of the second group of 4 neurons is linear y = x, and the 4 neural states are expressed by the row vector time = [time 1 ,time 2 ,time 3 ,time 4 ] indicates the opening time of the electromagnetic valve controlling the humidifier. 1 , the solenoid valve opening time of the dryer 2 , The solenoid valve of the refrigerator is open and running time 3 , heater solenoid valve opening time 4 .
4. The intelligent control unit according to claim 3, Features The critic network has 10 input neurons, an intermediate layer and an output layer. The 10 input neurons are temperature and relative humidity, and the output of the Actor network is represented by a row vector as input = [s t ,s h ,flag 1 ,flag 2 ,flag 3 ,flag 4 ,time 1 ,time 2 ,time 3 ,time 4 ]; The middle layer has several hidden layers, which are fully connected. Each hidden layer contains L i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer; The output layer contains a linear neuron with an activation function of y=x, which evaluates the value of the Actor network action.
5. A temperature and humidity control system for a cell culture chamber, Features The cell culture chamber is connected to a humidifier, a dryer, a refrigerator, and a heater through a gas channel, and the control system includes an intelligent control unit with a built-in deep reinforcement learning model, and the intelligent control unit controls the actuator to turn on or off at least one of the humidifier, the dryer, the refrigerator, and the heater according to the received real-time temperature and humidity of the cell culture chamber, the target temperature, the target humidity, and a predetermined threshold range, and controls the working time of at least one of the humidifier, the dryer, the refrigerator, and the heater; The deep reinforcement learning model is obtained by the following method: a. Set the objective function and constraints to be optimized for the deep reinforcement learning model. The objective function to be optimized is shown in formula (1), which means minimizing the power consumed and the time t used to reach the target stable state. 0 , where p i represents the power consumed by the components actually involved in the work, λ is the harmonic coefficient; the constraint condition is as shown in formula (2), which means that the fluctuation range of temperature and humidity is within the predetermined threshold range after reaching the target stable state, T best RH best They represent the target temperature and target humidity respectively; Δt and ΔRH represent the temperature and humidity fluctuation range respectively. temp(t>t 0 ) represents the temperature after reaching the target stable state, RHumity(t>t 0 ) represents the humidity after reaching the target stable state, and t is the current time; b. Training the deep reinforcement learning model b1 sets the total number of iterations N of the deep reinforcement learning model e , the number of explorations at each iteration point T, the learning rate of action network parameters η a , policy network parameter learning rate η c ; b2 uses a 0-1 Gaussian distribution to randomly initialize the Actor network A(s;θ a ) and Critic network C(s,a;θ c ) are denoted by θ a ,θ c , where θ a is the parameter of the Actor network, θ c is the parameter of the Critic network, s is the current ambient temperature and humidity input state, a is the execution action and is a row vector; b3 starts the first iteration, and count K=1; b3.1 Start the first exploration, and count n = 1; b3.2 According to the current ambient temperature and humidity status s t , the Actor network will s t As input, it passes through the network function A(s; θ a )|s=s t Next, a set of execution actions a is generated t ; b3.3 After executing a t After that, the environmental state of the cell culture chamber changed, and the temperature and humidity detection point found that the new state was s t +1 , according to formula (3) we get a timely reward r t , r t is Reward(t); Where M 1 ,M 2 ,M 3 ,M 4 are the penalty factors for each item respectively; b3.4 a t and the current ambient temperature and humidity status t The joint is used as input to the Critic network, and then passes through C(s,a;θ c )|s=s t ,a=a t After the action, an evaluation C is generated t ; b3.5 Calculate the Actor network A(s; θ) according to formula (4) a ) parameter θ a Gradient And update the parameter θ a , b3.6 Calculate the critic network C(s,a;θ) according to formula (5) c ) parameter θ c The gradient of Parameter θ c , in, is Reward(t), calculated by formula (3); b3.7 Environment status update completed t ←s t+1 ; b3.8 Update the exploration count n←n+1; b3.9 Re-execute process b3.2-b3.8 until n>T, completing this exploration process; b4 Update the iteration count, K←K+1; b5 Re-execute b3.1-b3.9 and b4 until K>N e , complete the deep reinforcement learning model DRL training; c. Place the trained deep reinforcement learning model into the intelligent control unit.
6. The temperature and humidity control system for a cell culture chamber according to claim 5, Features There are multiple cell culture chambers, each of which is independent of each other and controlled by a separate intelligent control unit.
7. The temperature and humidity control system for a cell culture chamber according to claim 5, Its characteristics are There are multiple cell culture chambers, each of which is independent of each other, and the intelligent control unit performs control according to the priority of each cell culture chamber.
8. The temperature and humidity control system for a cell culture chamber according to claim 5, Features The Actor network has two input neurons, an intermediate layer and an output layer. The two input neurons are represented by a row vector s = [s t ,s h ] indicates that each component in the row vector represents the temperature s of the current environmental state. t and relative humidity h ; The middle layer has several hidden layers, which are fully connected. Each hidden layer contains m i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer; The output layer has 8 neurons, which are divided into two groups. The first group of 4 neurons represents the flag of the solenoid valve opening. The activation function is softmax, which is recorded as the row vector [flag 1 ,flag 2 ] and [flag 3 ,flag 4 ], indicating whether the solenoid valves of the humidifier and dryer are open, and whether the solenoid valves of the refrigerator and heater are open; the activation function of the second group of 4 neurons is linear y = x, and the 4 neural states are expressed by the row vector time = [time 1 ,time 2 ,time 3 ,time 4 ] indicates the opening time of the electromagnetic valve controlling the humidifier. 1 , the solenoid valve opening time of the dryer 2 , The solenoid valve of the refrigerator is open and running time 3 , heater solenoid valve opening time 4 .
9. The temperature and humidity control system for a cell culture chamber according to claim 8, Features The critic network has 10 input neurons, an intermediate layer and an output layer. The 10 input neurons are temperature and relative humidity, and the output of the Actor network is represented by a row vector as input = [s t ,s h ,flag 1 ,flag 2 ,flag 3 ,flag 4 ,time 1 ,time 2 ,time 3 ,time 4 ]; The middle layer has several hidden layers, which are fully connected. Each hidden layer contains L i hidden layer neurons, where i represents the hidden layer number, the activation function of the hidden layer neurons is f(x)=max(wx+b,0), w represents the connection weight between neural network layers, x represents the output of the previous layer, and b represents the neuron bias of the current layer; The output layer contains a linear neuron with an activation function of y=x, which evaluates the value of the Actor network action.
Citation Information
Patent Citations
Cascade reservoir random optimization scheduling method based on deep Q learning
CN110930016A
System and method for realizing energy saving and temperature control of data center by applying artificial intelligence
CN111351180A