Steam sterilizer based on deep reinforcement learning and intelligent control system
Through deep reinforcement learning and PID control system, the problem of inaccurate water vapor state judgment in the existing technology is solved, and the automatic control of the sterilizer is realized, ensuring the use of saturated steam in the sterilization chamber and improving the sterilization effect.
Patent Information
- Application Number
- CN202511263335.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing technologies cannot effectively determine whether water vapor is saturated steam, resulting in unsaturated or superheated steam affecting the sterilization effect.
An intelligent control system based on deep reinforcement learning is used to determine whether the water vapor is saturated steam through pressure sensors and temperature sensors, and to automatically adjust the heating parameters using deep reinforcement learning modules and PID controllers to ensure the use of saturated steam in the sterilization chamber.
It realizes accurate judgment and automatic control of water vapor state, ensures the sterilization effect, ensures that saturated steam is always used in the sterilization room, and improves the reliability and efficiency of sterilization.
Smart Images

Figure CN120802602A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a kind of steam sterilizer and intelligent control system based on deep reinforcement learning, belong to automatic control technical field. BACKGROUND
[0002] The utility model discloses a pulsating vacuum steam sterilizer, it includes water inlet pipeline, water pump, first solenoid valve, evaporator, has the sterilization chamber of inner chamber and jacket layer, steam pipeline and electric conductivity sensor, the water inlet pipeline is connected electric conductivity sensor, water pump, first solenoid valve, evaporator and microcontroller in proper order, the evaporator is connected jacket layer by steam pipeline, electric conductivity sensor, water pump and first solenoid valve all connect microcontroller, and the microcontroller is proportional-integral and differential controller.
[0003] But the utility model does not disclose how to determine water vapor is saturated steam, only saturated steam can promote the denaturation coagulation of the microbial protein of surface pollutant of sterilization article, causes not to restore, reaches the purpose of sterilization.The required water vapor must be saturated steam, non-saturated steam and superheated steam will have adverse effects on sterilization effect. SUMMARY
[0004] To overcome the defects of prior art, the present application provides a kind of steam sterilizer and intelligent control system based on deep reinforcement learning, it can judge whether water vapor is saturated steam, and according to the judgment result, automatically control equipment adjusts heating parameter, ensure that saturated steam is involved in sterilization in sterilization chamber, guarantee sterilization effect.
[0005] To realize the present application purposes, the present application provides a kind of steam sterilizer based on deep reinforcement learning, it includes sterilization chamber, steam generator and intelligent control system, sterilization chamber is provided with pressure sensor and temperature sensor;Control system includes deep reinforcement learning module, first and second PID controllers, first and second subtractors and electronic switch, first subtractor generates sequence first error signal according to the measured temperature and set temperature provided by temperature sensor, deep reinforcement learning module generates first state According to sequence first error signal; Second subtractor generates sequence second error signal according to the measured gas pressure provided by pressure sensor and water vapor saturated gas pressure determined by set temperature, deep reinforcement learning module generates second state According to sequence second error signal; According to first state And second state , first PID controller control strategy at time t , , respectively are the proportional, integral and derivative coefficients of the first PID controller at time t, and generate a control strategy of the second PID controller at time t , , respectively are the proportional, integral and derivative coefficients of the second PID controller at time t; according to the second error signal at time t generate a control strategy of the electronic switch at time t , B t is the control signal of the electronic switch at time t; the first PID controller controls the working state of the steam generator; the second PID controller selects the electrically controlled air inlet valve or the first electrically controlled air outlet valve through the electronic switch to control the working state of the selected valve.
[0006] To achieve the object of the application, the application further provides an intelligent control system, which comprises a deep reinforcement learning module, first and second PID controllers, first and second subtractors and an electronic switch, the first subtractor generates a sequence of first error signals according to a measured temperature provided by a temperature sensor and a set temperature, the deep reinforcement learning module generates a first state according to the sequence of first error signals; the second subtractor generates a sequence of second error signals according to a measured air pressure provided by a pressure sensor and a saturated water vapor pressure determined by the set temperature, the deep reinforcement learning module generates a second state according to the sequence of second error signals, and generates a control strategy of the first PID controller at time t according to the first state and the second state , , , respectively are the proportional, integral and derivative coefficients of the first PID controller at time t, and generate a control strategy of the second PID controller at time t , , respectively are the proportional, integral and derivative coefficients of the second PID controller at time t; according to the second error signal at time t generate a control strategy of the electronic switch at time t , B t is the control signal of the electronic switch at time t; the first PID controller controls the working state of the steam generator; the second PID controller selects the electrically controlled air inlet valve and the first electrically controlled air outlet valve through the electronic switch to further control the working state of the selected valve.
[0007] Compared with the prior art, the steam sterilizer and the intelligent control system based on deep reinforcement learning provided by the application can determine whether the water vapor is saturated steam, and automatically control the equipment to adjust the heating parameters according to the determination result, so that the saturated steam is ensured to participate in sterilization in the sterilization chamber, and the sterilization effect is ensured. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 is a schematic diagram of the steam sterilizer based on deep reinforcement learning provided by the first embodiment of the application.
[0009] Figure 2 is a block diagram of the intelligent control system constructed by the first embodiment of the application. DETAILED DESCRIPTION
[0010] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0011] First embodiment
[0012] Figure 1 is a schematic diagram of the steam sterilizer based on deep reinforcement learning provided by the first embodiment of the application, as Figure 1 shown, the steam sterilizer based on deep reinforcement learning comprises a sterilization chamber 1 and a steam generator, the steam outlet of the steam generator is connected to the sterilization chamber through an electrically controlled air inlet valve 2, and the sterilization chamber is connected to the outside through a first electrically controlled steam outlet valve 3.
[0013] The steam generator comprises a water tank 4 and a heating device 5, the heating device is used for heating the water in the water tank and converting the water into gaseous water vapor, and the heating device is fixed on a support 13. A liquid level sensor 6 is arranged in the water tank, which is used for measuring the water level of the water in the water tank. A gas pressure sensor 12 is arranged in the water tank, which is used for measuring the gas pressure in the water tank.
[0014] The steam generator further comprises a controller, a water inlet valve 7 and a water inlet pump 8, the water inlet valve and the water inlet pump work synchronously, the controller controls the working state of the water inlet valve and the water inlet pump according to the water level in the water tank detected by the liquid level sensor, when the water level in the water tank is lower than a first threshold value, the water inlet valve is opened, and the water inlet pump starts to work, the water inlet pump discharges the water in the water source into the water tank through the water inlet valve. When the water level in the water tank is higher than a second threshold value, the water inlet valve is closed, and the water inlet pump stops working. The second threshold value is greater than the first threshold value.
[0015] The steam generator further comprises a second steam outlet valve 9, which discharges part of the steam to the outside when the gas pressure in the water tank is greater than a third threshold value.
[0016] The sterilization chamber is also provided with a pressure sensor 11 for measuring the air pressure in the sterilization chamber and a temperature sensor 10 for measuring the temperature in the sterilization chamber.
[0017] In the first embodiment, the steam sterilizer based on deep reinforcement learning further comprises an intelligent control system, which will be described in detail below. Figure 2
[0018] Figure 2 is the component block diagram of the intelligent control system constructed by the first embodiment of the present application, as shown in Figure 2 , the intelligent control system comprises a deep reinforcement learning module, a first PID controller, a second PID controller, a first subtractor 21, a second subtractor 22 and an electronic switching switch K, the first subtractor generates a sequence of first error signals according to the measured temperature provided by the temperature sensor and the set temperature, the deep reinforcement learning module generates a first state according to the sequence of first error signals; the second subtractor generates a sequence of second error signals according to the measured air pressure provided by the pressure sensor and the saturated vapor pressure determined by the set temperature, the deep reinforcement learning module generates a second state according to the sequence of second error signals, generates a control strategy of the first PID controller at time t , according to the first state and the second state , respectively, the proportional, integral and differential coefficients of the first PID controller at time t, and generates a control strategy of the second PID controller at time t , , , respectively, the proportional, integral and differential coefficients of the second PID controller at time t; generates a control strategy of the electronic switching switch according to the second error signal at time t , B t is the control signal of the electronic switching switch at time t; the first PID controller controls the working state of the steam generator; the second PID controller selects the electrically controlled air inlet valve or the first electrically controlled air outlet valve through the electronic switching switch to control the working state of the selected valve. The electrically controlled air inlet valve and the first electrically controlled air outlet valve are connected to the output end through an adder 23, and the adder can also be a connection node.
[0019] In the first embodiment, the output of the first PID controller at time t is under ideal conditions that the temperature and pressure in the sterilization chamber do not affect each other; the set temperature of the sterilization chamber at time t and the measured temperature at time t the error of then
[0020] where , ; ; where e 1(t-1) is the error of the set temperature y 1d of the sterilization chamber at time t-1 and the measured temperature y 1(t-1) at time t-1; e 1(t-2) is the error of the set temperature y 1d of the sterilization chamber at time t-2 and the measured temperature y 1(t-2) at time t-2; in matrix form is , where , .
[0021] the output of the second PID controller at time t is ; the set temperature y of the sterilization chamber at time t is the saturated water vapor pressure and the measured pressure y of the sterilization chamber at time t is the error then , where , ; ; where e 2(t-1) is the error of the set temperature y 1d of the sterilization chamber at time t-1 is the saturated water vapor pressure and the measured pressure y 2(t-1) at time t-1; e 1(t-2) is the error of the set temperature y 1d of the sterilization chamber at time t-2 is the saturated water vapor pressure and the measured pressure y 2(t-2) at time t-2; in matrix form is , where , .
[0022] However, temperature and pressure are mutually influenced, therefore, in the first embodiment, the control parameter adjustment of the first PID controller controlling the working state of the steam generator is: , wherein, are the proportional coefficient, the integral coefficient and the differential coefficient of the first PID controller at time t, respectively.
[0023] The working parameter adjustment of the second PID controller controlling the electrically controlled intake valve or the first electrically controlled exhaust valve is: , wherein, are the proportional coefficient, the integral coefficient and the differential coefficient of the second PID controller at time t, respectively.
[0024] In the first embodiment, the neural network of the deep reinforcement learning module includes an input layer, a hidden layer and an output layer, wherein the input layer includes 6 neurons, the hidden layer includes 6 neurons, and the output layer includes 8 neurons, the first to sixth neurons of the output layer respectively output the proportional coefficient, the integral coefficient and the differential coefficient of the first and second PID controllers at time t: , wherein, , and are the center and the width of the jth Gaussian function of the neural network of the deep reinforcement learning module at time t; is the weight between the jth neuron of the hidden layer of the neural network of the deep reinforcement learning module and the nth neuron of the output layer at time t; j = 1, …, 6; n = 1, …, 6, represents the two-norm, represents splicing.
[0025] In the first embodiment, the deep reinforcement learning module generates the temperature state value function at time t and the pressure state value function at time t according to the first state and the second state , wherein the temperature state value function and the pressure state value function are respectively output by the seventh and eighth neurons of the output layer of the neural network of the deep reinforcement learning module, and are respectively: , , wherein, is the weight between the jth neuron of the hidden layer of the neural network of the deep reinforcement learning module and the seventh neuron of the output layer at time t, is the weight between the jth neuron of the hidden layer of the neural network of the deep reinforcement learning module and the 8th neuron of the output layer at time t.
[0026] In the first embodiment, the deep reinforcement learning module further constructs the cost function at time t according to the following formula: , In the formula, , In the formula, , is the measured temperature of the sterilization chamber at time t, is the set temperature of the sterilization chamber at time t; In the formula, is the set temperature of the sterilization chamber at time t is the saturated water vapor pressure at time t, is the measured pressure of the sterilization chamber at time t; is the temperature state value function at time t; is the temperature state value function at time t-1; is the pressure state value function at time t; is the pressure state value function at time t-1; , , , , is a weight coefficient for adjusting the dimension of each value.
[0027] The deep reinforcement learning module updates the parameters according to the following formula: , , , In the formula, , , , is a learning coefficient; is the weight between the jth neuron of the hidden layer of the neural network of the deep reinforcement learning module and the nth neuron of the output layer at time t+1, j = 1, …, 6, n = 1, …, 6; and is the weight between the jth neuron of the hidden layer of the neural network of the deep reinforcement learning module and the Nth neuron of the output layer at times t and t+1, j = 1, …, 6, N = 7, 8; is the center of the jth Gaussian function of the neural network of the deep reinforcement learning module at time t+1, j = 1, …, 6; is the bandwidth of the j-th Gaussian function of the neural network of the deep reinforcement learning module at time t+1, j=1,…,6.
[0028] In the first embodiment, ; ; ; ; ; .
[0029] In the first embodiment, the deep reinforcement learning module also determines the cost function Is it the smallest? If not, , , , And repeat the above steps, if yes, output 、 、 、 As the optimal parameter to calculate , and make Assign the proportional coefficient, integral coefficient and differential coefficient to the first PID controller respectively; Assign values to the proportional coefficient, integral coefficient and differential coefficient of the second PID controller respectively.
[0030] Additionally, although the parameters of the neural network of the deep reinforcement learning module are dynamically updated over time using gradient descent methods, the initial choice of these parameters is crucial to achieving the desired results.
[0031] In the first embodiment, according to the second error signal at time t Generate the control strategy of the electronic switch at time t Including: determining the second error signal at time t Is it a positive or negative number, that is, the set temperature of the sterilization chamber at time t Saturated water vapor pressure Greater than the measured air pressure in the sterilization chamber at time t When the second error signal is If the temperature of the sterilization chamber is set at time t, the second PID controller is connected to the electronically controlled air inlet valve through the switch, and the electronically controlled air inlet valve works to provide water vapor to the sterilization chamber. Saturated water vapor pressure Less than the measured air pressure y in the sterilization chamber at time t 2t When the second error signal is For negative, the second PID controller is connected to the first electrically controlled exhaust valve through a switch, and the first electrically controlled exhaust valve is operated to make the water vapor in the sterilization chamber exhaust.
[0032] Compared with the prior art, the steam sterilizer and the intelligent control system based on deep reinforcement learning provided by the application can determine whether the sterilization steam is saturated steam, and automatically control the equipment to adjust the heating parameters according to the determination result, so as to ensure that the saturated steam is used for sterilization in the sterilization chamber and the sterilization effect is ensured.
[0033] In addition, although the parameters of the neural network of the deep reinforcement learning module are dynamically updated over time using the gradient descent method, the initial selection of these parameters is crucial to achieving the desired results.
[0034] Second embodiment
[0035] The second embodiment of the application only describes the different content from the first embodiment, and the same content is not described repeatedly.
[0036] The second embodiment of the application provides a computer program product which uses a computer language to compile the method provided by the first embodiment into a computer program, and the computer program can be stored in a storage medium and called by one or more processors to implement the method of the plurality of steps in the first embodiment.
[0037] The preferred embodiments of the disclosed application are only used to help explain the application, and the preferred embodiments do not describe all the details and limit the application to the specific embodiments. Obviously, according to the content of the specification, many modifications and changes can be made, and the specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the application, so that those skilled in the art can well understand and utilize the application. The application is limited by the claims and their entire scope and equivalents.
Claims
1. A steam sterilizer based on deep reinforcement learning, characterized in that: It includes a sterilization chamber, a steam generator and an intelligent control system. The sterilization chamber is provided with a pressure sensor and a temperature sensor. The control system includes a deep reinforcement learning module, a first and a second PID controller, a first and a second subtractor and an electronic conversion switch. The first subtractor generates a sequence first error signal according to the measured temperature and the set temperature provided by the temperature sensor. The deep reinforcement learning module generates a first state according to the sequence first error signal. The second subtractor generates a second error signal according to the measured air pressure provided by the pressure sensor and the water vapor saturation pressure determined by the set temperature, and the deep reinforcement learning module generates a second state according to the second error signal. , according to the first state and the second state Generate the control strategy of the first PID controller at time t , , are the proportional, integral and differential coefficients of the first PID controller at time t, and generate the control strategy of the second PID controller at time t , , are the proportional, integral and differential coefficients of the second PID controller at time t; according to the second error signal at time t Generate the control strategy of the electronic transfer switch at time t , B t is the control signal of the electronic switching switch at time t; the first PID controller controls the working state of the steam generator; the second PID controller selects the electronically controlled intake valve or the first electronically controlled exhaust valve through the electronic switching switch to control the working state of the selected valve.
2. The steam sterilizer based on deep reinforcement learning according to claim 1, characterized in that , , are the proportional coefficient, integral coefficient and differential coefficient of the first PID controller at time t respectively; are the proportional coefficient, integral coefficient and differential coefficient of the second PID controller at time t respectively.
3. The steam sterilizer based on deep reinforcement learning according to claim 2, characterized in that , Where, , and are the center and width of the j-th Gaussian function of the neural network of the deep reinforcement learning module at time t; is the weight between the jth neuron in the hidden layer of the neural network of the deep reinforcement learning module and the nth neuron in the output layer at time t; j=1,…,6; n=1,…,6 represents the two-norm.
4. The steam sterilizer based on deep reinforcement learning according to claim 3, characterized in that Deep reinforcement learning model based on the first state and the second state Generate the temperature state value function at time t and the pressure state value function at time t : , , Where, is the weight between the jth neuron in the hidden layer and the 7th neuron in the output layer of the neural network of the deep reinforcement learning model at time t, is the weight between the jth neuron in the hidden layer and the 8th neuron in the output layer of the neural network of the deep reinforcement learning model at time t.
5. The steam sterilizer based on deep reinforcement learning according to claim 4, characterized in that: The deep reinforcement learning model also constructs the cost function at time t according to the following formula: , Where, , Where, , is the measured temperature of the sterilization chamber at time t, is the set temperature of the sterilization chamber at time t; , where The set temperature of the sterilization chamber at time t is The saturated water vapor pressure at is the measured air pressure in the sterilization chamber at time t; is the temperature state value function at time t; is the temperature state value function at time t-1; is the pressure state value function at time t; is the pressure state value function at time t-1; 、 、 、 is the weight coefficient; Update the parameters according to the following formula: , , , Where, 、 、 、 is the learning coefficient; is the weight between the jth neuron in the hidden layer and the nth neuron in the output layer of the neural network of the deep reinforcement learning module at time t+1, j=1,…,6,n=1,…,6; and are the weights between the jth neuron in the hidden layer and the Nth neuron in the output layer of the neural network of the deep reinforcement learning module at time t and t+1, j=1,…,6, N=7,8; is the center of the j-th Gaussian function of the neural network of the deep reinforcement learning module at time t+1, j=1,…,6; is the bandwidth of the j-th Gaussian function of the neural network of the deep reinforcement learning module at time t+1, j=1,…,6.
6. An intelligent control system, characterized in that: It includes a deep reinforcement learning module, a first and a second PID controller, a first and a second subtractor, and an electronic conversion switch. The first subtractor generates a sequence first error signal according to the measured temperature and the set temperature provided by the temperature sensor. The deep reinforcement learning module generates a first state according to the sequence first error signal. The second subtractor generates a second error signal according to the measured air pressure provided by the pressure sensor and the water vapor saturation pressure determined by the set temperature, and the deep reinforcement learning module generates a second state according to the second error signal. , according to the first state and the second state Generate the control strategy of the first PID controller at time t , , are the proportional, integral and differential coefficients of the first PID controller at time t, and generate the control strategy of the second PID controller at time t , , are the proportional, integral and differential coefficients of the second PID controller at time t; according to the second error signal at time t Control strategy of the electronic switching switch at time t , B t is the control signal of the electronic switching switch at time t; the first PID controller controls the working state of the steam generator; the second PID controller selects the electronically controlled intake valve or the first electronically controlled exhaust valve through the electronic switching switch to control the working state of the selected valve.
7. The intelligent control system according to claim 6, characterized in that: , , are the proportional coefficient, integral coefficient and differential coefficient of the first PID controller at time t respectively; are the proportional coefficient, integral coefficient and differential coefficient of the second PID controller at time t respectively.
8. The intelligent control system according to claim 7, characterized in that: , Where, , and are the center and width of the j-th Gaussian function of the neural network of the deep reinforcement learning module at time t; is the weight between the jth neuron in the hidden layer of the neural network of the deep reinforcement learning module and the nth neuron in the output layer at time t; j=1,…,6; n=1,…,6 represents the two-norm.
9. The intelligent control system according to claim 8, characterized in that: Deep reinforcement learning model based on the first state and the second state Generate the temperature state value function at time t and the pressure state value function at time t : , , Where, is the weight between the jth neuron in the hidden layer and the 7th neuron in the output layer of the neural network of the deep reinforcement learning model at time t, is the weight between the jth neuron in the hidden layer and the 8th neuron in the output layer of the neural network of the deep reinforcement learning model at time t.
10. The intelligent control system according to claim 9, characterized in that: The deep reinforcement learning model also constructs the cost function at time t according to the following formula: , Where, , Where, , is the measured temperature of the sterilization chamber at time t, is the set temperature of the sterilization chamber at time t; , where The set temperature of the sterilization chamber at time t is The saturated water vapor pressure at is the measured air pressure in the sterilization chamber at time t; is the temperature state value function at time t; is the temperature state value function at time t-1; is the pressure state value function at time t; is the pressure state value function at time t-1; 、 、 、 is the weight coefficient; Update the parameters according to the following formula: , , , Where, 、 、 、 is the learning coefficient; is the weight between the jth neuron in the hidden layer and the nth neuron in the output layer of the neural network of the deep reinforcement learning module at time t+1, j=1,…,6,n=1,…,6; and is the weight between the jth neuron in the hidden layer of the neural network of the deep reinforcement learning module and the Nth neuron in the output layer at time t and t+1, j=1,…,6,N=7,8; is the center of the j-th Gaussian function of the neural network of the deep reinforcement learning module at time t+1; is the bandwidth of the j-th Gaussian function of the neural network of the deep reinforcement learning module at time t+1.
Citation Information
Patent Citations
Intelligent control method for pulsation vacuum sterilizer based on fuzzy control
CN104850010A
Self-adaptive adjustment method for PID (Proportion Integration Differentiation) controller of chlorine dioxide sterilizer
CN115356919A
Control system and control method of water bath type sterilizer
CN115639774A
Intelligent microwave-heat exchange composite milk sterilization system and control method thereof
CN119318356A
PID controller parameter self-tuning method based on reinforcement learning
CN120215250A
Cited By
Low-temperature steam heating control method and system
CN121386975A