Intelligent hospital clean room air supply adjusting method and system

By using a combination method of long and short-term memory neural network and deep Q network algorithm in hospital clean rooms, multi-source environment prediction and air supply adjustment strategy generation is solved, and the problem of insufficient multi-source environment prediction accuracy and air supply adjustment reliability in the prior art is solved, and more efficient and reliable air supply adjustment in clean room is achieved.

CN120176228AInactive Publication Date: 2025-06-20XIAN SITENG ENVIRONMENTAL TECH CO LTD

Patent Information

Application Number
CN202510652524.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has poor accuracy in predicting multi-source environments in hospitals and poor reliability of air supply regulation.

Method used

A prediction model based on long and short-term memory neural networks and a deep Q network algorithm are used to combine multi-source environmental data to predict and air supply adjustment strategies. The method includes collecting clean room multi-source environment data, using long-term and short-term memory neural networks to predict the multi-source environment, and then using the deep Q network algorithm to generate a air supply adjustment strategy, and performing air supply adjustment according to the strategy.

Benefits of technology

It improves the accuracy of multi-source environment prediction in clean rooms and the reliability of air supply adjustment, ensuring the optimization of air quality, energy consumption and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120176228A_ABST
    Figure CN120176228A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent hospital clean room air supply adjusting method and system, and relates to the technical field of air supply adjusting. The method comprises the following steps that multi-source environment data of a clean room are collected; obtaining multi-source environment prediction data according to the multi-source environment data by adopting a prediction model based on a long short-term memory neural network; a deep Q network algorithm is adopted to generate an air supply adjusting strategy according to the multi-source environment prediction data; and air supply adjustment is conducted on the clean room according to the air supply adjustment strategy. According to the method, the prediction model based on the long-short-term memory neural network and the deep Q network algorithm are combined to realize multi-source environment prediction and generate the air supply adjustment strategy, so that the accuracy of multi-source environment prediction in the clean room and the reliability of air supply adjustment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of air supply regulation, and particularly to an intelligent air supply regulation method and system for hospital clean rooms. Background Art

[0002] As the core medical environment for controlling particulate pollution, the clean room directly affects the success rate of surgeries and the infection risk of patients. Traditional air supply systems are difficult to dynamically adapt to real-time changes such as personnel flow and equipment startup and shutdown, which easily lead to local pollution or excessive energy consumption. At the same time, medical standards have strict regulations on air cleanliness, air flow direction, and pressure difference gradient, and deviations in air supply parameters may violate compliance requirements. In addition, unreasonable air supply will cause temperature and humidity fluctuations, affecting the operation of precision instruments and the concentration of medical staff. Therefore, it is necessary to study the air supply regulation method and system for hospital clean rooms.

[0003] In the prior art, Chinese Patent CN118729516A discloses a clean room air supply regulation method and system based on reinforcement learning, including: extracting personnel time series features and environmental time series features from the real-time obtained clean room sensing data respectively; performing linear data analysis and non-linear error correction on the personnel time series features and environmental time series features respectively to obtain analyzed environmental data and analyzed personnel data; generating analyzed sensing data according to the analyzed environmental data and analyzed personnel data, performing forward propagation on the analyzed sensing data to obtain an analyzed air supply strategy; performing strategy evaluation on the analyzed air supply strategy according to the analyzed sensing data to obtain a strategy reward; using the strategy reward to update the analyzed air supply strategy into a standard air supply strategy, and performing air supply regulation on the clean room according to the standard air supply strategy.

[0004] However, the above prior art has poor accuracy in predicting the multi-source environment in the clean room and poor reliability in air supply regulation. Summary of the Invention

[0005] This application provides an intelligent air supply regulation method and system for hospital clean rooms to solve the problems of poor accuracy in predicting the multi-source environment in the clean room and poor reliability in air supply regulation in the prior art.

[0006] On the one hand, this application provides an intelligent air supply regulation method for hospital clean rooms, including the following steps: Step 1, collect multi-source environment data of the clean room.

[0007] Step 2, use a prediction model based on a long short-term memory neural network to obtain multi-source environment prediction data according to the multi-source environment data.

[0008] Step 3, use a deep Q-network algorithm to generate an air supply regulation strategy according to the multi-source environment prediction data.

[0009] Step 4: Adjust the air supply to the clean room according to the air supply adjustment strategy.

[0010] In a possible implementation, in Step 1, the multi-source environmental data includes: temperature data, humidity data, air pressure data, carbon dioxide concentration data, and PM2.5 concentration data.

[0011] In a possible implementation, in Step 2, the prediction model is obtained by training a long short-term memory neural network with a multi-source environmental training set.

[0012] In a possible implementation, the long short-term memory neural network incorporates a self-attention mechanism.

[0013] In a possible implementation, Step 3 includes: Create a deep Q-network model, which includes a policy network and a target network.

[0014] Train the deep Q-network model with a multi-source environmental prediction training set.

[0015] Use the trained deep Q-network model to generate an air supply adjustment strategy based on the multi-source environmental prediction data.

[0016] In a possible implementation, in Step 3, the training of the deep Q-network model with a multi-source environmental prediction training set includes: Initialization phase: Initialize the policy network parameters and target network parameters, create an experience replay buffer, and set hyperparameters.

[0017] Action selection phase: Select an action using the ε-greedy strategy based on the current state and the policy network.

[0018] Execution and observation phase: Execute the selected action, obtain a new state, and observe the new state and the corresponding immediate reward.

[0019] Experience storage phase: Store the current state, current action, immediate reward, and new state in the experience replay buffer.

[0020] Policy network update phase: Randomly sample a batch of samples from the experience replay buffer, calculate the target Q-value for each sample, and update the policy network parameters to minimize the loss function.

[0021] Target network update phase: Update the target network parameters every preset number of update steps.

[0022] Single-round training check phase: Check whether the preset training steps or the end state have been reached. If so, proceed to the next step; otherwise, return to the action selection phase.

[0023] During the training completion check phase, check whether the preset number of training rounds has been reached. If so, the training is completed; otherwise, return to the initialization phase.

[0024] In a possible implementation, the immediate reward is obtained based on a multi-objective reward function.

[0025] The multi-objective reward function includes: an air quality objective reward function, an energy consumption objective reward function, and a comfort objective reward function.

[0026] In a possible implementation, the multi-objective reward function calculates the immediate reward by means of adaptive weight adjustment.

[0027] The adaptive weight adjustment maps the air quality objective deviation, the energy consumption objective deviation, and the comfort objective deviation to the weight coefficients of the multi-objective reward function.

[0028] In a possible implementation, the air supply regulation strategy is set with several constraint conditions, including: minimum air supply volume constraint, air quality safety constraint, minimum adjustment interval constraint, and sensor failure constraint.

[0029] On the other hand, the present application also provides an intelligent air supply regulation system for a hospital clean room, which adopts the above-mentioned intelligent air supply regulation method for a hospital clean room, and includes: an environmental data acquisition module, an environmental data prediction module, an air supply strategy generation module, and an air supply regulation execution module.

[0030] The environmental data acquisition module is used to acquire multi-source environmental data of the clean room.

[0031] The environmental data prediction module is used to obtain multi-source environmental prediction data according to the multi-source environmental data by using a prediction model based on a long short-term memory neural network.

[0032] The air supply strategy generation module is used to generate an air supply regulation strategy according to the multi-source environmental prediction data by using a deep Q-network algorithm.

[0033] The air supply regulation execution module is used to perform air supply regulation on the clean room according to the air supply regulation strategy.

[0034] An intelligent air supply regulation method and system for a hospital clean room in the present application have the following advantages: By combining a prediction model based on a long short-term memory neural network and a deep Q-network algorithm to achieve multi-source environmental prediction and generate an air supply regulation strategy, the accuracy of multi-source environmental prediction in the clean room and the reliability of air supply regulation are improved.

[0035] The proposed long short-term memory neural network incorporates a self-attention mechanism, which can adaptively focus on key time steps, thereby more accurately capturing complex non-linear relationships and improving the multi-source environment prediction accuracy of the prediction model.

[0036] The proposed multi-objective reward function calculates the immediate reward by means of adaptive weight adjustment, automatically focusing on the most optimized objectives without adding additional network structures, improving the flexibility and reliability of the immediate reward calculation, and thus enhancing the reliability of the air supply regulation strategy.

[0037] The proposed constraints include the minimum air supply volume constraint, the air quality safety constraint, the minimum adjustment interval constraint, and the sensor failure constraint, improving the safety of the air supply regulation in the clean room. Brief Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is a schematic flow chart of an intelligent air supply regulation method for a hospital clean room provided by an embodiment of the present application. Detailed Embodiments

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0041] As Figure 1 shown, an embodiment of the present application provides an intelligent air supply regulation method for a hospital clean room, including the following steps: Step 1, collect multi-source environment data of the clean room.

[0042] Step 2, use a prediction model based on a long short-term memory neural network to obtain multi-source environment prediction data according to the multi-source environment data.

[0043] Step 3, use the deep Q-network algorithm to generate an air supply regulation strategy according to the multi-source environment prediction data.

[0044] Step 4, perform air supply regulation on the clean room according to the air supply regulation strategy.

[0045] Exemplarily, in step one, the multi-source environmental data includes: temperature data, humidity data, air pressure data, carbon dioxide concentration data, and PM2.5 concentration data.

[0046] Specifically, in this embodiment, the temperature data, humidity data, air pressure data, carbon dioxide concentration data, and PM2.5 concentration data are respectively collected by a temperature sensor, a humidity sensor, a barometer, a carbon dioxide sensor, and a particulate matter sensor arranged in the clean room, and are all time-series numerical data (in this embodiment, the data of the past 30 minutes).

[0047] Exemplarily, in step two, the prediction model is obtained by training a long short-term memory neural network with a multi-source environmental training set.

[0048] Specifically, in this embodiment, the data in the multi-source environmental training set is a collection of multi-source environmental data collected historically. These data are aligned according to timestamps to form multi-dimensional time series data, and the dimensionality difference is eliminated through a normalization operation. A sliding time window is constructed to generate an input sequence and a target sequence (in this embodiment, the multi-dimensional time series data of the first 30 time steps is used to predict the values of the next 5 time steps, and each time step is set to 1 minute).

[0049] In this embodiment, the structure of the long short-term memory neural network is set as: an input layer, two LSTM layers, and an output layer. Among them, the input layer is used to input multi-dimensional time series data (30 time steps, a total of five dimensions including temperature data, humidity data, air pressure data, carbon dioxide concentration data, and PM2.5 concentration data), the LSTM layer is used to capture long-term dependencies (in this embodiment, each LSTM layer contains 128 LSTM units. The first LSTM layer outputs the hidden states of all time steps, and the second LSTM layer outputs the hidden state of the last time step), and the output layer is used to output the prediction results of multi-dimensional time series data (5 time steps, five dimensions).

[0050] In this embodiment, the optimizer of the long short-term memory neural network adopts the Adam optimizer, and the loss function adopts the weighted MES loss function.

[0051] Exemplarily, the long short-term memory neural network introduces a self-attention mechanism.

[0052] Specifically, in this embodiment, the self-attention mechanism is introduced into the long short-term memory neural network, including: a self-attention mechanism layer is set after the first LSTM layer of the long short-term memory neural network. The output of the first LSTM layer reaches the self-attention mechanism layer and the second LSTM layer respectively. The self-attention mechanism layer and the second LSTM layer are in a parallel processing relationship. The self-attention mechanism layer is used to generate Q, K, and V matrices respectively for the hidden states of all time steps output by the first LSTM layer through linear transformation, and calculate the weights of different time steps through the dot-product attention formula. Subsequently, the V matrix is weighted and summed according to the weights to output the attention-weighted features. A feature fusion layer is set after the second LSTM layer and the self-attention mechanism layer. The feature fusion layer is used to splice and fuse the hidden state output by the second LSTM layer and the attention-weighted features output by the self-attention mechanism layer, and output the fused features to the output layer. The output layer maps the fused features to the target dimension and outputs the prediction result.

[0053] The long short-term memory neural network is trained by using a multi-source environment training set. After several iterations, the trained long short-term memory neural network, that is, the prediction model, is obtained.

[0054] Exemplarily, step three includes: Create a deep Q-network model, and the deep Q-network model includes a policy network and a target network.

[0055] The deep Q-network model is trained by using a multi-source environment prediction training set.

[0056] The trained deep Q-network model is used to generate a air supply adjustment strategy according to the multi-source environment prediction data.

[0057] Specifically, in this embodiment, the state space and the action space of the deep Q-network model are defined. Among them, the state is defined as the current multi-source environment data, the corresponding multi-source environment prediction data, and the air supply adjustment parameters of the past 3 time steps; the action is defined as a combination of air supply adjustment parameters, including air supply volume, air supply speed, filtration level, etc. The input of the deep Q-network model is the current state, and the output is the Q value of each action. The action corresponding to the highest Q value is selected as the air supply adjustment strategy.

[0058] Exemplarily, in step three, the training of the deep Q-network model by using the multi-source environment prediction training set includes: In the initialization stage, the parameters of the policy network and the target network are initialized, an experience replay buffer is created, and hyperparameters are set.

[0059] In the action selection stage, an action is selected according to the current state and the policy network by using the ε-greedy strategy.

[0060] Execute the observation phase, perform the selected action to obtain a new state, and observe the new state and the corresponding immediate reward.

[0061] In the experience storage phase, store the current state, current action, immediate reward, and new state in the experience replay buffer.

[0062] In the policy network update phase, randomly sample a batch of samples from the experience replay buffer, calculate the target Q-value for each sample, and update the policy network parameters to minimize the loss function.

[0063] In the target network update phase, update the target network parameters every preset number of update steps.

[0064] In the single-round training check phase, check whether the preset training steps or the end state have been reached. If so, proceed to the next step; otherwise, return to the action selection phase.

[0065] In the training completion check phase, check whether the preset number of training rounds has been reached. If so, the training is completed; otherwise, return to the initialization phase.

[0066] Specifically, in this embodiment, the Q-value prediction of the policy network is denoted as Q(s, a; θ), and the Q-value prediction of the target network is denoted as Q(s, a; θ-). Here, s represents the state, a represents the action, θ represents the policy network parameters, and θ- represents the target network parameters.

[0067] In the initialization phase, initialize the policy network parameters θ, and use the policy network parameters θ to initialize the target network parameters θ-; create an experience replay buffer D, and set the capacity of the experience replay buffer to 1000 entries; set hyperparameters, including the learning rate α, discount factor γ, exploration rate ε, and preset number of update steps C.

[0068] In the action selection phase, randomly select an action with a probability of exploration rate ε, and select the optimal action output by the current policy network with a probability of 1 - ε.

[0069] In the experience storage phase, store the current state s t 、the current action a t 、the immediate reward r t 、the new state s t+1 as a set of experiences in the experience replay buffer, where t represents the current time step.

[0070] In the policy network update phase, each sample corresponds to a set of experiences, and the target Q-value is calculated as follows: y = r + γ * max a Q(s`, a; θ-).

[0071] Here, y represents the target Q-value, r represents the reward corresponding to the action a, and s` represents the next state of the state s.

[0072] The loss function of the deep Q-network model is as follows: L(θ)=E (s,a,r,s`)~D [(y - Q(s,a;θ))²].

[0073] Among them, L(θ) represents the loss function, and E (s,a,r,s`)~D represents the average loss of the samples (s,a,r,s`) sampled from the experience replay buffer D.

[0074] In the target network update stage, every preset update step C (set to 10 time steps in this embodiment) is used to update the target network parameter θ- by τθ+(1 - τ)θ-, where τ represents the update factor (set to 0.5 in this embodiment).

[0075] Exemplarily, the immediate reward is obtained based on a multi-objective reward function.

[0076] The multi-objective reward function includes: an air quality objective reward function, an energy consumption objective reward function, and a comfort objective reward function.

[0077] Specifically, in this embodiment, the air quality objective reward function is shown as follows: R1 = -(x1*z1 2 +x2*z2 2 ).

[0078] Among them, R1 represents the air quality objective reward, x1 represents the PM2.5 reward coefficient, z1 represents the PM2.5 concentration deviation, x2 represents the carbon dioxide reward coefficient, and z2 represents the carbon dioxide concentration deviation.

[0079] The energy consumption objective reward function is shown as follows: R2 = -x3*z3*x4*z4.

[0080] Among them, R2 represents the energy consumption objective reward, x3 represents the air supply volume reward coefficient, z3 represents the air supply volume, x4 represents the air supply wind speed reward coefficient, and z4 represents the air supply wind speed.

[0081] The comfort objective reward function is shown as follows: R3 = -(x5*z5 2 +x6*z6 2 +x7*z7 2 ).

[0082] Among them, R3 represents the comfort reward, x5 represents the temperature reward coefficient, z5 represents the temperature deviation, x6 represents the humidity reward coefficient, z6 represents the humidity deviation, x7 represents the air pressure reward coefficient, and z7 represents the air pressure deviation.

[0083] In this embodiment, the multi-objective reward function is as follows: r t = w1 * R1 + w2 * R2 + w3 * R3.

[0084] Among them, w1, w2, and w3 respectively represent the air quality weight coefficient, the energy consumption weight coefficient, and the comfort weight coefficient.

[0085] Exemplarily, the multi-objective reward function calculates the immediate reward by means of adaptive weight adjustment.

[0086] The adaptive weight adjustment maps the air quality target deviation, the energy consumption target deviation, and the comfort target deviation to the weight coefficients of the multi-objective reward function.

[0087] Specifically, in this embodiment, the air quality target deviation is as follows: H1 = g1 * z1 + g2 * z2.

[0088] Among them, H1 represents the air quality target deviation, and g1 and g2 respectively represent the PM2.5 deviation coefficient and the carbon dioxide deviation coefficient.

[0089] The energy consumption target deviation is as follows: H2 = g3 * z3 * g4 * z4 - P.

[0090] Among them, H2 represents the energy consumption target deviation, g3 and g4 respectively represent the air volume energy consumption coefficient and the air supply velocity energy consumption coefficient, and P represents the preset reference energy consumption.

[0091] The comfort target deviation is as follows: H3 = g5 * z5 + g6 * z6 + g7 * z7.

[0092] Among them, H3 represents the comfort target deviation, and g1, g2, and g3 respectively represent the temperature deviation coefficient, the humidity deviation coefficient, and the air pressure deviation coefficient.

[0093] In this embodiment, mapping the air quality target deviation, the energy consumption target deviation, and the comfort target deviation to the weight coefficients of the multi-objective reward function, first obtain the original weight coefficients as follows: w 1t = H 1t / (H 1t + H 2t + H 3t ), w 2t = H 2t / (H 1t + H 2t + H 3t ),w 3t = H 3t / (H1t +H 2t +H 3t )。

[0094] Among them, w 1t represents the original air quality weight coefficient at the current time step t, w 2t represents the original energy consumption weight coefficient at the current time step t, w 3t represents the original comfort weight coefficient at the current time step t, H 1t 、H 2t 、H 3t respectively represent the air quality target deviation, energy consumption target deviation, and comfort target deviation at the current time step t.

[0095] Subsequently, the original weight coefficients are adaptively adjusted as the time step changes, as shown in the following formula: w` 1t =λ*w 1t +(1-λ)*w 1(t-1) ,w` 2t =λ*w 2t +(1-λ)*w 2(t-1) ,w` 3t =λ*w 3t +(1-λ)*w 3(t-1) 。

[0096] Among them, w` 1t 、w` 2t 、w` 3t respectively represent the air quality weight coefficient, energy consumption weight coefficient, and comfort weight coefficient after adaptive adjustment at the current time step t, λ represents the reward function weight update factor (set to 0.3 in this embodiment), w 1(t-1) 、w 2(t-1) 、w 3(t-1) respectively represent the air quality weight coefficient, energy consumption weight coefficient, and comfort weight coefficient at the previous time step (t-1).

[0097] Then the multi-objective reward function is modified to: r t =w` 1t *R1+w` 2t *R2+w` 3t *R3。

[0098] Exemplarily, several constraint conditions are set for the air supply adjustment strategy, including: minimum air supply volume constraint, air quality safety constraint, minimum adjustment interval constraint, and sensor failure constraint.

[0099] In this embodiment, the minimum air supply volume constraint is to set a lower limit of the air supply volume (the air supply volume shall not be lower than this lower limit), the air quality safety constraint is to set an upper limit of the PM2.5 concentration (when exceeding this upper limit, the filtration level will be changed to high-efficiency filtration), the minimum adjustment interval constraint is to set a lower limit of the air supply adjustment time interval (in this embodiment, the air supply adjustment time interval shall not be lower than 5 minutes), and the sensor failure constraint is to adopt the maximum air supply volume, the maximum air supply wind speed, and high-efficiency filtration when the sensor is abnormal.

[0100] The embodiment of the present application also provides an intelligent air supply adjustment system for a hospital clean room, which adopts the above-mentioned intelligent air supply adjustment method for a hospital clean room, and includes: an environmental data acquisition module, an environmental data prediction module, an air supply strategy generation module, and an air supply adjustment execution module.

[0101] The environmental data acquisition module is used to acquire multi-source environmental data of the clean room.

[0102] The environmental data prediction module is used to obtain multi-source environmental prediction data according to the multi-source environmental data by using a prediction model based on a long short-term memory neural network.

[0103] The air supply strategy generation module is used to generate an air supply adjustment strategy according to the multi-source environmental prediction data by using a deep Q-network algorithm.

[0104] The air supply adjustment execution module is used to adjust the air supply of the clean room according to the air supply adjustment strategy.

[0105] The embodiment of the present application realizes multi-source environmental prediction and generates an air supply adjustment strategy by combining a prediction model based on a long short-term memory neural network and a deep Q-network algorithm, improving the accuracy of multi-source environmental prediction and the reliability of air supply adjustment in the clean room.

[0106] The proposed long short-term memory neural network introduces a self-attention mechanism, which can adaptively focus on key time steps, and thus more accurately capture complex non-linear relationships, improving the multi-source environmental prediction accuracy of the prediction model.

[0107] The proposed multi-objective reward function calculates the immediate reward by using an adaptive weight adjustment method, automatically focusing on the most optimized objective without adding additional network structures, improving the flexibility and reliability of immediate reward calculation, and thus improving the reliability of the air supply adjustment strategy.

[0108] The proposed constraint conditions include the minimum air supply volume constraint, the air quality safety constraint, the minimum adjustment interval constraint, and the sensor failure constraint, improving the safety of air supply adjustment in the clean room.

[0109] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0110] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. An intelligent hospital clean room air supply adjustment method, characterized in that: The following steps are involved: Step 1: Collect multi-source environmental data of the clean room; Step 2, using a prediction model based on a long short-term memory neural network to obtain multi-source environmental prediction data according to the multi-source environmental data; Step 3, using a deep Q network algorithm to generate an air supply adjustment strategy according to the multi-source environmental prediction data; Step 4: Adjust the air supply to the clean room according to the air supply adjustment strategy.

2. An intelligent hospital clean room air supply adjustment method according to claim 1, characterized in that: In step one, the multi-source environmental data includes: temperature data, humidity data, air pressure data, carbon dioxide concentration data, and PM2.5 concentration data.

3. The intelligent hospital clean room air supply adjustment method according to claim 1 is characterized in that: In step 2, the prediction model is obtained by training the long short-term memory neural network using a multi-source environment training set.

4. The intelligent hospital clean room air supply adjustment method according to claim 3 is characterized in that: The long short-term memory neural network introduces a self-attention mechanism.

5. The intelligent hospital clean room air supply adjustment method according to claim 1 is characterized in that: Step three includes: Creating a deep Q network model, wherein the deep Q network model includes a policy network and a target network; Using a multi-source environment prediction training set to train the deep Q network model; The trained deep Q network model is used to generate an air supply adjustment strategy according to the multi-source environmental prediction data.

6. An intelligent hospital clean room air supply adjustment method according to claim 5, characterized in that: In step 3, the use of a multi-source environment prediction training set to train the deep Q network model includes: In the initialization phase, the policy network parameters and target network parameters are initialized, the experience replay buffer is created, and the hyperparameters are set; In the action selection phase, the ε-greedy strategy is used to select actions based on the current state and the policy network; Execute the observation phase, execute the selected action, obtain the new state, observe the new state and the corresponding immediate reward; In the experience storage phase, the current state, current action, immediate reward, and new state are stored in the experience playback buffer; In the policy network update phase, a batch of samples are randomly selected from the experience replay buffer, the target Q value of each sample is calculated, and the policy network parameters are updated to minimize the loss function; In the target network update phase, the target network parameters are updated at intervals with a preset number of update steps; In the single-round training check phase, check whether the preset number of training steps or the end state has been reached. If so, proceed to the next step, otherwise return to the action selection phase; The training completion check phase checks whether the preset training round has been reached. If so, the training is completed, otherwise it returns to the initialization phase.

7. An intelligent hospital clean room air supply adjustment method according to claim 6, characterized in that: The instant reward is obtained based on a multi-objective reward function; The multi-objective reward function includes: an air quality objective reward function, an energy consumption objective reward function, and a comfort objective reward function.

8. An intelligent hospital clean room air supply adjustment method according to claim 7, characterized in that: The multi-objective reward function calculates the instant reward by adaptive weight adjustment; The adaptive weight adjustment is to map the air quality target deviation, the energy consumption target deviation, and the comfort target deviation into the weight coefficient of the multi-objective reward function.

9. The intelligent hospital clean room air supply adjustment method according to claim 1, characterized in that: The air supply adjustment strategy is set with several constraints, including: minimum air supply volume constraint, air quality safety constraint, minimum adjustment interval constraint, and sensor failure constraint.

10. An intelligent hospital clean room air supply adjustment system, using an intelligent hospital clean room air supply adjustment method as claimed in any one of claims 1 to 9, characterized in that: include: Environmental data collection module, environmental data prediction module, air supply strategy generation module, air supply regulation execution module; The environmental data acquisition module is used to collect multi-source environmental data of the clean room; The environmental data prediction module is used to obtain multi-source environmental prediction data according to the multi-source environmental data by adopting a prediction model based on a long short-term memory neural network; The air supply strategy generation module is used to generate an air supply adjustment strategy according to the multi-source environment prediction data using a deep Q network algorithm; The air supply adjustment execution module is used to adjust the air supply of the clean room according to the air supply adjustment strategy.

Citation Information

Patent Citations

  • Clean room air supply adjusting method and system based on reinforcement learning

    CN118729516A

  • Flexible regulation and control system for load of multi-split central air conditioner

    CN117989674A

  • Air conditioner control method, device and system, electronic equipment and air conditioner

    CN119737669A

  • Method, device and equipment for controlling clean room to save energy by deploying AI algorithm

    CN119802790A

  • Learning device, air conditioning control system, inference device, air conditioning control device, and trained model generation method

    US20250093065A1

Cited By

  • Clean room dynamic energy consumption management system and method based on AI

    CN120444710A

  • Method and system for controlling environmental parameters of electronic clean workshop

    CN121112463A

  • Method and system for controlling environmental parameters of an electronic clean room

    CN121112463B

  • Multi-area FFU differential pressure collaborative dynamic balance control method

    CN121163048A