Steam sterilizer and intelligent control system based on deep reinforcement learning
By using deep reinforcement learning and a PID control system, the problem of not being able to determine whether water vapor is saturated steam in existing technologies has been solved, realizing intelligent control of the sterilizer, ensuring sterilization effect, and improving the reliability and efficiency of sterilization.
Patent Information
- Application Number
- CN202511263335.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing technologies cannot effectively determine whether water vapor is saturated steam, resulting in unsaturated or overheated steam affecting the sterilization effect.
An intelligent control system based on deep reinforcement learning is adopted. It acquires measured data through pressure and temperature sensors, uses a deep reinforcement learning module to determine whether the water vapor is saturated steam, and automatically adjusts the heating parameters through a PID controller to ensure that saturated steam is used in the sterilization chamber.
It achieves accurate judgment and automatic control of water vapor, ensuring sterilization effect and guaranteeing that saturated steam is always used in the sterilization chamber, thereby improving the reliability and efficiency of sterilization.
Smart Images

Figure CN120802602B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a steam sterilizer and intelligent control system based on deep reinforcement learning, belonging to the field of automatic control technology. Background Technology
[0002] Chinese utility model patent CN204521667U discloses a pulsed vacuum steam sterilizer, which includes a water inlet pipe, a water pump, a first solenoid valve, an evaporator, a sterilization chamber with an inner chamber and a jacket layer, a steam pipe, and a conductivity sensor. The water inlet pipe is sequentially connected to the conductivity sensor, the water pump, the first solenoid valve, the evaporator, and a microcontroller. The evaporator is connected to the jacket layer through the steam pipe. The conductivity sensor, the water pump, and the first solenoid valve are all connected to the microcontroller, which is a proportional-integral and derivative controller.
[0003] However, this utility model patent does not disclose how to determine whether the water vapor is saturated steam. Only saturated steam can cause the microbial proteins on the surface of sterilized items to denature and coagulate, rendering them irreversible and achieving sterilization. The required water vapor must be saturated steam; unsaturated steam and superheated steam will adversely affect the sterilization effect. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this invention provides a steam sterilizer and intelligent control system based on deep reinforcement learning. It can determine whether the steam is saturated steam and automatically control the equipment to adjust the heating parameters according to the determination result, so as to ensure that the steam participating in sterilization in the sterilization chamber is saturated steam and guarantee the sterilization effect.
[0005] To achieve the aforementioned objective, this invention provides a steam sterilizer based on deep reinforcement learning, comprising a sterilization chamber, a steam generator, and an intelligent control system. The sterilization chamber is equipped with a pressure sensor and a temperature sensor. The control system includes a deep reinforcement learning module, first and second PID controllers, first and second subtractors, and an electronic transfer switch. The first subtractor generates a first sequence error signal based on the measured temperature provided by the temperature sensor and the set temperature. The deep reinforcement learning module generates a first state based on the first sequence error signal. The second subtractor generates a second sequence error signal based on the measured air pressure provided by the pressure sensor and the water vapor saturation pressure determined by the set temperature. The deep reinforcement learning module then generates a second state based on this second sequence error signal. According to the first state Second state Generate the control strategy of the first PID controller at time t. , , The proportional, integral, and derivative coefficients of the first PID controller at time t are given respectively, and the control strategy of the second PID controller at time t is generated. , , These are the proportional, integral, and derivative coefficients of the second PID controller at time t; based on the second error signal at time t... Generate the control strategy of the electronic switching switch at time t B t The electronic transfer switch provides the control signal at time t; the first PID controller controls the operating state of the steam generator; the second PID controller selects either the electrically controlled intake valve or the first electrically controlled exhaust valve via the electronic transfer switch to control the operating state of the selected valve.
[0006] To achieve the aforementioned objective, the present invention also provides an intelligent control system, comprising a deep reinforcement learning module, first and second PID controllers, first and second subtractors, and an electronic transfer switch. The first subtractor generates a sequence of first error signals based on the measured temperature provided by a temperature sensor and a set temperature. The deep reinforcement learning module generates a first state based on the sequence of first error signals. The second subtractor generates a second sequence error signal based on the measured air pressure provided by the pressure sensor and the water vapor saturation pressure determined by the set temperature. The deep reinforcement learning module then generates a second state based on this second sequence error signal. According to the first state Second state Generate the control strategy of the first PID controller at time t. , , The proportional, integral, and derivative coefficients of the first PID controller at time t are given respectively, and the control strategy of the second PID controller at time t is generated. , , These are the proportional, integral, and derivative coefficients of the second PID controller at time t; based on the second error signal at time t... Control strategy of bioelectric switching at time t B t The electronic transfer switch provides the control signal at time t; the first PID controller controls the operating state of the steam generator; the second PID controller selects the electrically controlled intake valve and the first electrically controlled exhaust valve via the electronic transfer switch to further control the operating state of the selected valve.
[0007] Compared with existing technologies, the steam sterilizer and intelligent control system based on deep reinforcement learning provided by this invention can determine whether the steam is saturated steam and automatically control the equipment to adjust the heating parameters according to the judgment result, so as to ensure that the steam participating in sterilization in the sterilization chamber is saturated steam and guarantee the sterilization effect. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the composition of a steam sterilizer based on deep reinforcement learning provided in the first embodiment of the present invention.
[0009] Figure 2 This is a block diagram of the intelligent control system constructed according to the first embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] First Embodiment
[0012] Figure 1 This is a schematic diagram of the composition of a steam sterilizer based on deep reinforcement learning provided in the first embodiment of the present invention, as shown below. Figure 1 As shown, the steam sterilizer based on deep reinforcement learning includes: a sterilization chamber 1 and a steam generator. The exhaust port of the steam generator is connected to the sterilization chamber through an electrically controlled air inlet valve 2, and the sterilization chamber is connected to the outside through a first electrically controlled exhaust valve 3.
[0013] The steam generator includes a water tank 4 and a heating device 5. The heating device heats the water in the tank and converts it into gaseous steam. The heating device is fixed to a bracket 13. A level sensor 6 is installed inside the water tank to measure the water level. A pressure sensor 12 is installed inside the water tank to measure the pressure inside the tank.
[0014] The steam generator also includes a controller, an inlet valve 7, and an inlet pump 8. The inlet valve and the inlet pump operate synchronously. The controller controls the operation of the inlet valve and the inlet pump based on the water level in the tank detected by the level sensor. When the water level in the tank is lower than a first threshold, the inlet valve opens, and the inlet pump starts working, discharging water from the water source into the tank through the inlet valve. When the water level in the tank is higher than a second threshold, the inlet valve closes, and the inlet pump stops working. The second threshold is greater than the first threshold.
[0015] The steam generator also includes a second exhaust valve 9, which discharges some steam to the outside when the air pressure in the water tank is greater than the third threshold.
[0016] The sterilization chamber is also equipped with a pressure sensor 11 and a temperature sensor 10. The pressure sensor is used to measure the air pressure inside the sterilization chamber, and the temperature sensor is used to measure the temperature inside the sterilization chamber.
[0017] In the first embodiment, the deep reinforcement learning-based steam sterilizer also includes an intelligent control system, which will be discussed below. Figure 2 Please provide a detailed explanation.
[0018] Figure 2 This is a block diagram of the intelligent control system constructed according to the first embodiment of the present invention, as shown below. Figure 2 As shown, the intelligent control system includes a deep reinforcement learning module, a first PID controller, a second PID controller, a first subtractor 21, a second subtractor 22, and an electronic transfer switch K. The first subtractor generates a sequence of first error signals based on the measured temperature provided by the temperature sensor and the set temperature. The deep reinforcement learning module generates a first state based on the sequence of first error signals. The second subtractor generates a second sequence error signal based on the measured air pressure provided by the pressure sensor and the water vapor saturation pressure determined by the set temperature. The deep reinforcement learning module then generates a second state based on this second sequence error signal. According to the first state Second state Generate the control strategy of the first PID controller at time t. , , The proportional, integral, and derivative coefficients of the first PID controller at time t are given respectively, and the control strategy of the second PID controller at time t is generated. , , These are the proportional, integral, and derivative coefficients of the second PID controller at time t; based on the second error signal at time t... Control strategy of bioelectric switching B t The electronic transfer switch provides the control signal at time t; the first PID controller controls the operating state of the steam generator; the second PID controller selects either the electrically controlled intake valve or the first electrically controlled exhaust valve via the electronic transfer switch to control the operating state of the selected valve. The electrically controlled intake valve and the first electrically controlled exhaust valve are connected to the output terminal via adder 23, which can also be a connection node.
[0019] In the first embodiment, ideally assuming that the temperature and pressure in the sterilization chamber do not affect each other, the output of the first PID controller at time t is: The set temperature of the sterilization chamber at time t. and the measured temperature at time t The error is Then we have:
[0020]
[0021] In the formula, , ; ;
[0022] In the formula, e 1(t-1) Set the temperature y of the sterilization chamber at time t-1. 1d and the measured temperature y at time t-1 1(t-1) The error; e 1(t-2) The set temperature y of the sterilization chamber at time t-2. 1d and the measured temperature y at time t-2 1(t-2) Error;
[0023] Written in matrix form:
[0024] ,
[0025] In the formula, , .
[0026] The output of the second PID controller at time t is: The set temperature of the sterilization chamber at time t. saturated water vapor pressure The measured air pressure in the sterilization chamber at time t The error is Then we have:
[0027] ,
[0028] In the formula, , ; ;
[0029] In the formula, e 2(t-1) Let y be the set temperature of the sterilization chamber at time t-1. 1d saturated water vapor pressure at time and the measured air pressure y at time t-1 2(t-1) The error; e 1(t-2) The set temperature of the sterilization chamber at time t-2 is y. 1d saturated water vapor pressure at time And the measured air pressure y at time t-2 2(t-2) Error;
[0030] Written in matrix form:
[0031] ,
[0032] In the formula, , .
[0033] However, temperature and pressure are interdependent. Therefore, in the first embodiment, the control parameters of the first PID controller controlling the operating state of the steam generator are adjusted as follows:
[0034] In the formula, These are the proportional coefficient, integral coefficient, and derivative coefficient of the first PID controller at time t, respectively.
[0035] The operating parameters of the second PID controller controlling the electronically controlled intake valve or the first electronically controlled exhaust valve are adjusted as follows:
[0036] In the formula, These are the proportional coefficient, integral coefficient, and derivative coefficient of the second PID controller at time t, respectively.
[0037] In the first embodiment, the neural network of the deep reinforcement learning module includes an input layer, a hidden layer, and an output layer. The input layer includes 6 neurons, the hidden layer includes 6 neurons, and the output layer includes 8 neurons. The first to sixth neurons of the output layer output the proportional coefficient, integral coefficient, and derivative coefficient of the first and second PID controllers at time t, respectively.
[0038] ,
[0039] In the formula, , and These are the center and width of the j-th Gaussian function in the deep reinforcement learning module's neural network at time t; These are the weights between the j-th neuron in the hidden layer and the n-th neuron in the output layer of the deep reinforcement learning module at time t; j=1,…,6; n=1,…,6. Represents the L2 norm, Indicates splicing.
[0040] In the first embodiment, the deep reinforcement learning module is based on the first state. Second state Generating the temperature state value function at time t and the pressure state value function at time t Temperature state value function and pressure state value function The outputs are from the 7th and 8th neurons of the output layer of the deep reinforcement learning module's neural network, respectively:
[0041] ,
[0042] ,
[0043] In the formula, It represents the weights between the j-th neuron in the hidden layer and the 7th neuron in the output layer of a deep reinforcement learning neural network at time t. It represents the weights between the j-th neuron in the hidden layer and the 8th neuron in the output layer of the deep reinforcement learning neural network at time t.
[0044] In the first embodiment, the deep reinforcement learning model also constructs a cost function at time t according to the following formula:
[0045] ,
[0046] In the formula, ,
[0047] In the formula, , The measured temperature of the sterilization chamber at time t. The set temperature of the sterilization chamber at time t;
[0048] In the formula, The set temperature of the sterilization chamber at time t is The saturated water vapor pressure at that time The measured air pressure in the sterilization chamber at time t; The temperature state value function at time t; This is a function of the temperature state value at time t-1; The pressure state value function at time t; This is a function of the pressure state value at time t-1; , , , These are weighting coefficients used to adjust the dimensions of each value.
[0049] The deep reinforcement learning model updates its parameters according to the following formula:
[0050] , ,
[0051] ,
[0052] In the formula, , , , The learning coefficient; Let represent the weights between the j-th neuron in the hidden layer and the n-th neuron in the output layer of the neural network in the deep reinforcement learning module at time t+1, where j=1,…,6,n=1,…,6; and The weights between the j-th neuron in the hidden layer and the N-th neuron in the output layer of the deep reinforcement learning module at times t and t+1, j=1,…,6,N=7,8; Let j be the center of the j-th Gaussian function of the neural network in the deep reinforcement learning module at time t+1, where j=1,…,6; Let $\frac{j}{j}$ be the bandwidth of the $j$-th Gaussian function in the neural network of the deep reinforcement learning module at time $t+1$, where $j=1,...,6$.
[0053] In the first embodiment, ;
[0054] ;
[0055] ;
[0056] ;
[0057] ;
[0058] .
[0059] In the first embodiment, the deep reinforcement learning module also determines the cost function. Is it the smallest? If not, , , , Repeat the above steps. If yes, output... , , , As the optimal parameter for calculation , and make The proportional, integral, and derivative coefficients of the first PID controller are assigned values respectively; The proportional coefficient, integral coefficient, and derivative coefficient of the second PID controller are assigned values respectively.
[0060] Furthermore, although the parameters of the neural network in the deep reinforcement learning module are dynamically updated over time using gradient descent, the initial selection of these parameters is crucial for achieving the desired results.
[0061] In the first embodiment, based on the second error signal at time t Control strategy for generating electronic switching at time t Includes: determining the second error signal at time t. Is it a positive or negative number, that is, the set temperature of the sterilization chamber at time t? saturated water vapor pressure The measured air pressure in the sterilization chamber at time t is greater than the actual air pressure. At time t, the second error signal If the value is positive, the second PID controller is connected to the electrically controlled air intake valve via a switch, and the electrically controlled air intake valve operates to supply water vapor to the sterilization chamber; if the set temperature of the sterilization chamber at time t... saturated water vapor pressure The measured air pressure y in the sterilization chamber at time t is less than the actual air pressure y. 2t At time t, the second error signal When the value is negative, the second PID controller is connected to the first electrically controlled exhaust valve via a switch, and the first electrically controlled exhaust valve operates to expel water vapor from the sterilization chamber.
[0062] Compared with existing technologies, the steam sterilizer and intelligent control system based on deep reinforcement learning provided by this invention can determine whether the sterilization steam is saturated steam, and automatically control the equipment to adjust the heating parameters according to the judgment result, so as to ensure that the sterilization in the sterilization chamber is saturated steam and guarantee the sterilization effect.
[0063] Furthermore, although the parameters of the neural network in the deep reinforcement learning module are dynamically updated over time using gradient descent, the initial selection of these parameters is crucial for achieving the desired results.
[0064] Second Embodiment
[0065] The second embodiment of the present invention only describes the contents that are different from those of the first embodiment; the contents that are the same will not be described again.
[0066] The second embodiment of the present invention provides a computer program product, which uses a computer language to compile the method provided in the first embodiment into a computer program. The computer program can be stored in a storage medium and called by one or more processors to implement the method of multiple steps in the first embodiment.
[0067] The preferred embodiments of the present invention disclosed herein are merely for the purpose of illustrating the present invention. The preferred embodiments do not describe all the details exhaustively, nor do they limit the invention to specific implementation methods. Obviously, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A steam sterilizer based on deep reinforcement learning, characterized in that, The system includes a sterilization chamber, a steam generator, and an intelligent control system. The sterilization chamber is equipped with pressure and temperature sensors. The control system includes a deep reinforcement learning module, first and second PID controllers, first and second subtractors, and an electronic transfer switch. The first subtractor generates a first sequence error signal based on the measured temperature provided by the temperature sensor and the set temperature. The deep reinforcement learning module generates a first state based on the first sequence error signal. The second subtractor generates a second sequence error signal based on the measured air pressure provided by the pressure sensor and the water vapor saturation pressure determined by the set temperature. The deep reinforcement learning module then generates a second state based on this second sequence error signal. According to the first state Second state Generate the control strategy of the first PID controller at time t. , , The proportional, integral, and derivative coefficients of the first PID controller at time t are given respectively, and the control strategy of the second PID controller at time t is generated. , , These are the proportional, integral, and derivative coefficients of the second PID controller at time t; based on the second error signal at time t... Generate the control strategy of the electronic switching switch at time t B t The electronic transfer switch provides the control signal at time t; the first PID controller controls the operating state of the steam generator; the second PID controller selects either the electrically controlled intake valve or the first electrically controlled exhaust valve via the electronic transfer switch to control the operating state of the selected valve.
2. The steam sterilizer based on deep reinforcement learning according to claim 1, characterized in that, , , These are the proportional coefficient, integral coefficient, and derivative coefficient of the first PID controller at time t, respectively. These are the proportional coefficient, integral coefficient, and derivative coefficient of the second PID controller at time t, respectively.
3. The steam sterilizer based on deep reinforcement learning according to claim 2, characterized in that, , In the formula, , and These are the center and width of the j-th Gaussian function in the deep reinforcement learning module's neural network at time t; These are the weights between the j-th neuron in the hidden layer and the n-th neuron in the output layer of the deep reinforcement learning module at time t; j=1,…,6; n=1,…,6 This represents the L2 norm.
4. The steam sterilizer based on deep reinforcement learning according to claim 3, characterized in that, Deep reinforcement learning model based on the first state Second state Generating the temperature state value function at time t and the pressure state value function at time t : , , In the formula, It represents the weights between the j-th neuron in the hidden layer and the 7th neuron in the output layer of a deep reinforcement learning neural network at time t. It represents the weights between the j-th neuron in the hidden layer and the 8th neuron in the output layer of the deep reinforcement learning neural network at time t.
5. The steam sterilizer based on deep reinforcement learning according to claim 4, characterized in that, The deep reinforcement learning model also constructs a cost function at time t based on the following formula: , In the formula, , In the formula, , The measured temperature of the sterilization chamber at time t. The set temperature of the sterilization chamber at time t; In the formula, The set temperature of the sterilization chamber at time t is The saturated water vapor pressure at that time The measured air pressure in the sterilization chamber at time t; The temperature state value function at time t; This is a function of the temperature state value at time t-1; The pressure state value function at time t; This is a function of the pressure state value at time t-1; , , , These are the weighting coefficients; Update parameters according to the following formula: , , , In the formula, , , , The learning coefficient; Let represent the weights between the j-th neuron in the hidden layer and the n-th neuron in the output layer of the neural network in the deep reinforcement learning module at time t+1, where j=1,…,6,n=1,…,6; and These are the weights of the j-th neuron in the hidden layer of the deep reinforcement learning module and the N-th neuron in the output layer at times t and t+1, respectively, where j=1,…,6,N=7,8; Let j be the center of the j-th Gaussian function of the neural network in the deep reinforcement learning module at time t+1, where j=1,…,6; Let $\frac{j}{t+1}$ be the bandwidth of the $j$-th Gaussian function in the neural network of the deep reinforcement learning module at time $t+1$, where $j=1,...,6$.
6. An intelligent control system, characterized in that, The system includes a deep reinforcement learning module, first and second PID controllers, first and second subtractors, and an electronic transfer switch. The first subtractor generates a first sequence error signal based on the measured temperature provided by the temperature sensor and the set temperature. The deep reinforcement learning module generates a first state based on the first sequence error signal. The second subtractor generates a second sequence error signal based on the measured air pressure provided by the pressure sensor and the water vapor saturation pressure determined by the set temperature. The deep reinforcement learning module then generates a second state based on this second sequence error signal. According to the first state Second state Generate the control strategy of the first PID controller at time t. , , The proportional, integral, and derivative coefficients of the first PID controller at time t are given respectively, and the control strategy of the second PID controller at time t is generated. , , These are the proportional, integral, and derivative coefficients of the second PID controller at time t; based on the second error signal at time t... Control strategy of bioelectric switching at time t B t The electronic transfer switch provides the control signal at time t; the first PID controller controls the operating state of the steam generator; the second PID controller selects either the electrically controlled intake valve or the first electrically controlled exhaust valve via the electronic transfer switch to control the operating state of the selected valve.
7. The intelligent control system according to claim 6, characterized in that, , , These are the proportional coefficient, integral coefficient, and derivative coefficient of the first PID controller at time t, respectively. These are the proportional coefficient, integral coefficient, and derivative coefficient of the second PID controller at time t, respectively.
8. The intelligent control system according to claim 7, characterized in that, , In the formula, , and These are the center and width of the j-th Gaussian function in the deep reinforcement learning module's neural network at time t; These are the weights between the j-th neuron in the hidden layer and the n-th neuron in the output layer of the deep reinforcement learning module at time t; j=1,…,6; n=1,…,6 This represents the L2 norm.
9. The intelligent control system according to claim 8, characterized in that, Deep reinforcement learning model based on the first state Second state Generating the temperature state value function at time t and the pressure state value function at time t : , , In the formula, It represents the weights between the j-th neuron in the hidden layer and the 7th neuron in the output layer of a deep reinforcement learning neural network at time t. It represents the weights between the j-th neuron in the hidden layer and the 8th neuron in the output layer of the deep reinforcement learning neural network at time t.
10. The intelligent control system according to claim 9, characterized in that, The deep reinforcement learning model also constructs a cost function at time t based on the following formula: , In the formula, , In the formula, , The measured temperature of the sterilization chamber at time t. The set temperature of the sterilization chamber at time t; In the formula, The set temperature of the sterilization chamber at time t is The saturated water vapor pressure at that time The measured air pressure in the sterilization chamber at time t; The temperature state value function at time t; This is a function of the temperature state value at time t-1; The pressure state value function at time t; This is a function of the pressure state value at time t-1; , , , These are the weighting coefficients; Update parameters according to the following formula: , , , In the formula, , , , The learning coefficient; Let represent the weights between the j-th neuron in the hidden layer and the n-th neuron in the output layer of the neural network in the deep reinforcement learning module at time t+1, where j=1,…,6,n=1,…,6; and Let represent the weights between the j-th neuron in the hidden layer and the N-th neuron in the output layer of the neural network in the deep reinforcement learning module at times t and t+1, where j=1,…,6,N=7,8; The center of the j-th Gaussian function in the neural network of the deep reinforcement learning module at time t+1; Let be the bandwidth of the j-th Gaussian function of the neural network in the deep reinforcement learning module at time t+1.
Citation Information
Patent Citations
Pulsation vacuum steam sterilization ware
CN204521667U
Control system and control method of water bath type sterilizer
CN115639774A
Intelligent microwave-heat exchange composite milk sterilization system and control method thereof
CN119318356A