Ventilation control method in anesthesia recovery period
By employing multimodal sequence modeling and safety constraints, the problems of information silos and response lags in ventilation management during anesthesia recovery were solved, enabling precise switching of ventilation modes, reducing patient-ventilator aggression, shortening recovery time, and improving patient safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE FIRST PEOPLES HOSPITAL OF XIAOSHAN DISTRICT HANGZHOU
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-01
AI Technical Summary
Current technologies for ventilation management during anesthesia recovery rely heavily on physician experience, leading to information silos and delayed responses, which can easily result in patient-ventilator asynchrony. Furthermore, they make it difficult to achieve accurate ventilation assessment and control, increasing the risk of hypoxemia, hypercapnia, or lung injury.
By employing multimodal sequence modeling and safety constraints, and combining variational mode decomposition, temporal skeleton alignment, cross-attention mechanism, and DecisionTransformer model with hard rule constraints, we can achieve precise and automatic switching of ventilation modes for processing multi-source physiological data of patients in the anesthesia recovery period.
It significantly reduced patient-ventilator asynchrony caused by hyperventilation, shortened postoperative recovery time, improved patient prognosis, and enhanced the accuracy and safety of ventilation control.
Smart Images

Figure CN121964082A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital medical technology, and more specifically to a method for ventilation control during the anesthesia recovery period. Background Technology
[0002] The recovery period after general anesthesia is considered one of the highest-risk phases in the perioperative period. This period involves not only central respiratory drive recalibration but is also closely related to neuromuscular function recovery. According to relevant initiatives from the World Health Organization and medical statistics, the complication rate in the post-anesthesia care unit (PACU) is as high as 54.8%, with about half of these related to the respiratory system. A survey of domestic hospitals showed that the incidence of residual muscle relaxant blockade at the time of extubation once reached as high as 57.8%, meaning that many patients were forced to attempt breathing before their muscle function had fully recovered, greatly increasing the possibility of hypoxemia, hypercapnia, or severe lung injury. During the anesthesia recovery phase, the process of the patient's respiratory drive recovering from drug suppression to spontaneous regulation has a highly non-linear characteristic. Therefore, accurate ventilation assessment and control are of irreplaceable clinical value in avoiding respiratory failure due to premature weaning, reducing patient-ventilator asynchrony caused by hyperventilation, and shortening postoperative recovery time.
[0003] However, current mainstream ventilation management models in hospitals still heavily rely on the anesthesiologist's personal experience and subjective judgment. They typically involve manual mode switching based on rough indicators such as cough reflexes or body movements, and suffer from significant information silos and response lags. Specifically, high-frequency respiratory waveforms, circulatory signs, and pharmacokinetics data are often displayed fragmentedly on different devices, forcing physicians to only provide passive, reactive interventions after a drop in blood oxygen saturation or abnormal end-tidal carbon dioxide levels. This makes patient-ventilator asynchrony more likely. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention aims to provide a ventilation control method for the anesthesia recovery period based on multimodal sequence modeling and safety constraints. This method involves collecting multi-source heterogeneous physiological data from patients during the anesthesia recovery period, and using variational mode decomposition and temporal skeleton alignment techniques for feature extraction and synchronous preprocessing. Then, a multimodal state fusion module based on cross-attention mechanism and a sequence decision model based on DecisionTransformer are constructed to capture long-term physiological dependence and drug metabolism lag. Furthermore, hard rule constraints and smoothing processes are used for clinical constraints, thereby achieving precise and automatic switching between mechanical ventilation and spontaneous breathing modes during the anesthesia recovery period.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for ventilation control during the anesthesia recovery period, comprising the following steps: Step 1: Perform waveform decomposition, spatiotemporal alignment, and first-order difference preprocessing on multi-source physiological data of patients in the anesthesia recovery period; Step 2: Construct a multimodal state fusion module based on the cross-attention mechanism. This module generates a fusion representation vector containing respiratory, circulatory, drug features and patient state through three parallel feature encoding branches: high-frequency respiratory feature embedding, circulatory and drug feature embedding, and discrete state encoding embedding, as well as an attention fusion unit. Step 3: Construct a sequence decision-making model based on DecisionTransformer, and use medical priors to construct a comprehensive reward function to jointly optimize the model in order to learn long-term contextual dependencies and ventilation strategies; Step four: In the inference phase, a target reward value representing smooth extubation is set, and hard rule constraints and smoothing are applied to the initial strategy generated by the model to finally obtain the real-time switching command between mechanical ventilation and spontaneous breathing.
[0006] As a further improvement of the present invention, the specific method for performing waveform decomposition, spatiotemporal alignment, and first-order difference preprocessing on the multi-source physiological data of patients in the anesthesia recovery period in step one is as follows: A variational mode decomposition algorithm is introduced to process the acquired high-frequency airway pressure waveform. By decomposing the complex waveform signal into several intrinsic mode functions, the low-frequency component representing the main wave of mechanical ventilation and the high-frequency component representing the patient's weak spontaneous breathing attempts are separated. Subsequently, to address the problem of inconsistent sampling of multi-source data such as airway pressure, drug concentration, and cardiovascular indicators, a unified 1Hz time skeleton is constructed to map data of different frequencies onto the same time axis for rigid spatiotemporal alignment. Then, the drug concentration data is calculated using first-order difference. Based on the current concentration, the concentration change rate is further obtained to clarify whether the drug is currently in the washout or maintenance phase.
[0007] As a further improvement of the present invention, the specific method for generating the fusion representation vector containing respiration, circulation, drug characteristics, and patient status in step two is as follows: The high-frequency respiratory feature embedding branch uses 1D-CNN to extract features from the preprocessed airway pressure waveform features and respiratory mechanics indices, and maps the captured local morphological changes into a high-dimensional respiratory embedding vector. As a core representation of transient respiratory state, the circulatory and drug feature embedding branch uses a multilayer perceptron to map low-frequency updated heart rate, blood oxygen saturation, and differentially processed drug effect-room concentration into circulatory drug embedding vectors. This provides the model with physiological context; the discrete state encoding branch targets intubation and spontaneous breathing states as well as cumulative apnea time, mapping discrete features and concatenating them into a dense, low-dimensional discrete state vector through a learnable embedding layer. Then, and Channel dimension aggregation is performed to form a global context vector. .
[0008] As a further improvement to the present invention, step three also uses a cross-attention mechanism to solve the modeling problem of complex coupling relationships in long time series, specifically as follows: using For query vector, Given key and value vectors, the model uses dot product attention scores to dynamically retrieve relevant contextual information from drug metabolism and patient state based on the current respiratory waveform morphology, ultimately obtaining the fusion result. .make If the scaling factor represents the dimension of the key vector, then this step can be expressed as follows: .
[0009] As a further improvement of the present invention, the specific method for constructing the sequence decision model based on DecisionTransformer in step three is as follows: Construct a sequence decision model based on DecisionTransformer, in which, for At that moment, For the patient's fusion state vector, Ventilation procedures were performed. For residual reward, the goal of the model is to predict the optimal action given the current state and expected reward. Then, the training sequence is constructed as follows. : ; Through the constructed training sequences Train the model.
[0010] As a further improvement of the present invention, the specific method for jointly optimizing the model by constructing a comprehensive reward function using medical priors in step three is as follows: The comprehensive reward function is defined as follows: Among them, security rewards The aim is to punish extreme physiological conditions that threaten the patient's life and to maintain the patient's blood oxygen saturation. and cycle stability; After that, This is an indicator function; it is set to 1 when the heart rate or systolic blood pressure deviates from the baseline by more than 20%. and If the weight is a hyperparameter, then the reward function for this item can be defined as follows: Among them, ventilation quality awards The aim is to dynamically assess the rationality of apnea based on the drug metabolism background and guide the model to stop mechanical ventilation in a timely manner; Next, order This refers to the partial pressure of carbon dioxide at the end of expiration. For the duration of apnea, This refers to the remifentanil effect-room concentration. , If the weights are hyperparameters, then the reward function can be defined as follows: Among them, respiratory mechanics and comfort reward The aim is to address the human-machine interaction problem and encourage smooth, spontaneous breathing; Furthermore, let Peak inhalation pressure, The variance of the high-frequency components obtained from variational mode decomposition. This is an indicator function for stable tidal volume, and it is set to 1 when the fluctuation range of tidal volume is less than 20% of the average value and the average tidal volume is greater than 300 ml. , , If the weights are hyperparameters, then the reward function can be defined piecewise according to the ventilation status as follows: Among them, motion smoothness penalty The aim is to prevent the model from frequently switching between mechanical ventilation and spontaneous breathing modes, and to encourage the model to only make changes when it is certain that a change in ventilation mode is necessary, thereby ensuring the consistency of clinical implementation. Finally, let If the weight hyperparameter is denoted as , then the penalty can be defined as follows: .
[0011] As a further improvement of the present invention, in step four, a target reward value representing smooth extubation is set during the inference stage, and hard rule constraints and smoothing are applied to the initial strategy generated by the model to finally obtain the real-time switching instruction between mechanical ventilation and spontaneous breathing. The specific method is as follows: First, 1.1 times the highest cumulative reward in the training set is used as the high target reward value that can represent "perfect and smooth extubation"; then, the model uses the self-attention mechanism of DecisionTransformer to perform autoregressive inference based on the fused state vector sequence, historical action sequence and the expected reward within the time period, and predicts whether mechanical ventilation should continue or spontaneous breathing should be switched at the current moment.
[0012] As a further improvement of the present invention, in step four, a rule constraint layer is introduced to verify the model output to ensure absolute clinical safety. Specifically, regardless of the model prediction results, if the patient's vital signs reach any of the following clinical red line conditions, the system will forcibly override the control command to mechanical ventilation to ensure immediate intervention in dangerous situations. These conditions include: blood oxygen saturation below 92% for more than 5 seconds, end-expiratory carbon dioxide partial pressure exceeding 55 mmHg, cumulative apnea time exceeding 30 seconds, or peak airway pressure exceeding 35 cmH2O.
[0013] As a further improvement of the present invention, in step four, a hysteresis comparator is also introduced, and asymmetric switching logic is used to smooth the action.
[0014] The beneficial effects of this invention are as follows: This invention utilizes variational mode decomposition technology to separate spontaneous breathing features from high-frequency airway pressure, and combines multimodal data preprocessing and fusion with temporal skeleton alignment and cross-attention mechanisms. Then, a sequence decision-making model based on DecisionTransformer is constructed, and a comprehensive reward function incorporating safety and comfort, along with a hard rule constraint layer, is applied. Under the premise of ensuring patient clinical safety and smoothness of movement, precise and automatic switching of ventilation modes during the anesthesia recovery period is achieved, thereby significantly reducing patient-ventilator asynchrony caused by hyperventilation, shortening postoperative recovery time, and improving patient prognosis. Attached Figure Description
[0015] Figure 1 This is a flowchart of the ventilation control method during the anesthesia recovery period of the present invention; Figure 2 This is a schematic diagram of the data processing procedure for the ventilation control method during the anesthesia recovery period of the present invention. Detailed Implementation
[0016] The present invention will now be described in further detail with reference to the embodiments shown in the accompanying drawings.
[0017] Example 1: As Figure 1-2 As shown, a method for ventilation control during the anesthesia recovery period based on multimodal sequence modeling and safety constraints is described, and the method steps are as follows: Step 1: Perform preprocessing operations such as waveform decomposition, spatiotemporal alignment, and first-order difference on the multi-source physiological data of patients in the anesthesia recovery period; Clinical experience shows that weak spontaneous breathing in patients during the anesthesia recovery period is often easily drowned out by the forced waveform of mechanical ventilation and is susceptible to interference from cardiac concussion and tubing fluid noise, leading to difficulties in identifying patient-ventilator asynchrony signals. Based on this, this solution introduces a variational mode decomposition algorithm to process the acquired high-frequency airway pressure waveform. By decomposing the complex waveform signal into several intrinsic mode functions, the low-frequency component representing the main mechanical ventilation wave and the high-frequency component representing the patient's weak spontaneous breathing attempts are separated. Subsequently, to address the issue of inconsistent sampling of multi-source data such as airway pressure, drug concentration, and cardiovascular indicators, a unified 1Hz time skeleton is constructed to map data of different frequencies onto the same time axis for rigid spatiotemporal alignment. Considering the smooth and strong inertia of effect-site concentration changes in anesthetic drugs such as propofol and remifentanil, first-order difference calculations are performed on the drug concentration data to obtain the current concentration... Based on this, the concentration change rate was further obtained. This is to determine whether the drug is currently in the washout or maintenance phase.
[0018] Step 2: Construct a multimodal state fusion module based on the cross-attention mechanism to generate a fusion representation vector containing respiration, circulation, drug features, and patient status; The data sources involved in the anesthesia recovery period are heterogeneous and vary greatly in time scale. The simple vector concatenation method widely used in current practice is insufficient to capture the differences in feature importance under different clinical scenarios. Based on the above situation, this solution constructs a state fusion module with a multimodal cross-attention network as its core, enabling the model to dynamically learn the association weights between different modalities of data and generate a latent vector representing the patient's current recovery potential.
[0019] This module is implemented through three parallel feature encoding branches: high-frequency respiratory feature embedding, circulatory and drug feature embedding, and discrete state encoding embedding, along with an attention fusion unit. Specifically, the high-frequency respiratory feature embedding branch uses 1D-CNN to extract features from the preprocessed airway pressure waveform and respiratory mechanics indices, mapping the captured local morphological changes into high-dimensional respiratory embedding vectors. As a core representation of transient respiratory state, the circulatory and drug feature embedding branch uses a multilayer perceptron to map low-frequency updated heart rate, blood oxygen saturation, and differentially processed drug effect-room concentration into circulatory drug embedding vectors. This provides the model with physiological context; the discrete state encoding branch targets intubation and spontaneous breathing states as well as cumulative apnea time, mapping discrete features and concatenating them into a dense, low-dimensional discrete state vector through a learnable embedding layer. Subsequently, and Channel dimension aggregation is performed to form a global context vector. .
[0020] After obtaining the feature vectors of each modality, this scheme utilizes a cross-attention mechanism to solve the modeling problem of complex coupling relationships over long time. For query vector, Given key and value vectors, the model uses dot product attention scores to dynamically retrieve relevant contextual information from drug metabolism and patient state based on the current respiratory waveform morphology, ultimately obtaining the fusion result. .make If the scaling factor represents the dimension of the key vector, then this step can be expressed as follows: Step 3: Construct a sequence decision-making model based on Decision Transformer, and use medical priors to construct a comprehensive reward function to jointly optimize the model in order to learn long-term contextual dependencies and ventilation strategies; The recovery process from anesthesia typically involves a drug metabolism lag of several minutes and a significant evolution of respiratory dynamics. This long-term dependence imposes considerable limitations on traditional reinforcement learning algorithms based on single-step state mapping. To address this, this proposal constructs a sequence decision-making model centered on the Decision Transformer, reconstructing ventilation control as a sequence prediction problem. Specifically, for At that moment, For the patient's fusion state vector, Ventilation procedures were performed. For residual reward, the goal of the model is to predict the optimal action given the current state and expected reward. First, construct the training sequence as follows: : To guide the model in learning clinically appropriate strategies, this approach designs a comprehensive reward function that incorporates safety, ventilation quality, comfort, and smoothness. Model parameters are optimized by maximizing the cumulative reward of this function. The comprehensive reward function is defined as follows: Among them, security rewards The aim is to punish extreme physiological conditions that threaten the patient's life and to maintain the patient's blood oxygen saturation. And cycle stability. Let This is an indicator function; it is set to 1 when the heart rate or systolic blood pressure deviates from the baseline by more than 20%. and If the weight is a hyperparameter, then the reward function for this item can be defined as follows: Ventilation quality reward The aim is to dynamically assess the rationale for apnea based on drug metabolism background and guide the model to discontinue mechanical ventilation in a timely manner. This refers to the partial pressure of carbon dioxide at the end of expiration. For the duration of apnea, This refers to the remifentanil effect-room concentration. , If the weights are hyperparameters, then the reward function can be defined as follows: Respiratory mechanics and comfort reward Aimed at solving the problem of human-machine interaction and encouraging smooth, spontaneous breathing, Peak inhalation pressure, The variance of the high-frequency components obtained from variational mode decomposition. This is an indicator function for stable tidal volume, and it is set to 1 when the fluctuation range of tidal volume is less than 20% of the average value and the average tidal volume is greater than 300 ml. , , If the weights are hyperparameters, then the reward function can be defined piecewise according to the ventilation status as follows: motion smoothness penalty The aim is to prevent the model from frequently switching between mechanical ventilation and spontaneous breathing modes, and to encourage the model to only implement changes when it is certain that a change in ventilation mode is necessary, thereby ensuring consistency in clinical implementation. If the weight hyperparameter is denoted as , then the penalty can be defined as follows: Step 4: In the inference stage, set the target reward value to represent smooth extubation, and apply hard rule constraints and smoothing to the initial strategy generated by the model to finally obtain the real-time switching command for mechanical ventilation or spontaneous breathing. This approach treats ventilation control as a conditional sequence generation problem. Unlike the random exploration in the training phase, the inference phase uses a high target reward value that is 1.1 times the highest cumulative reward in the training set to represent "perfect and smooth extubation". Subsequently, the model uses the self-attention mechanism of the Decision Transformer to perform autoregressive inference based on the fused state vector sequence, historical action sequence, and the expected reward over the time period, and predicts whether to continue mechanical ventilation or switch to spontaneous breathing at the current moment.
[0021] To mitigate the risk of directly applying uninterpretable end-to-end model outputs to clinical practice, this approach introduces a hard rule constraint layer to validate the model outputs and ensure absolute clinical safety. Specifically, regardless of the model's predictions, if a patient's vital signs trigger any of the following clinical red lines, the system will forcibly override the control command to mechanical ventilation to ensure immediate intervention in dangerous situations. These conditions include: oxygen saturation below 92% for more than 5 seconds, end-tidal carbon dioxide partial pressure exceeding 55 mmHg, cumulative apnea time exceeding 30 seconds, or peak airway pressure exceeding 35 cmH2O.
[0022] Furthermore, to prevent frequent ventilator start-ups and shutdowns due to second-level output fluctuations in the decision-making model, this solution introduces a hysteresis comparator and uses asymmetric switching logic to smooth the action. For the switch from mechanical ventilation to spontaneous breathing, the system only performs the weaning operation when the model outputs spontaneous breathing switching commands for 5 consecutive seconds with a confidence level higher than 80%. For the switch from spontaneous breathing to mechanical ventilation, once the model or hard rules issue a signal to switch to mechanical ventilation, a zero-delay switch is immediately executed to ensure smooth action while maximizing patient safety.
[0023] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for ventilation control during the recovery period from anesthesia, characterized in that: Includes the following steps: Step 1: Perform waveform decomposition, spatiotemporal alignment, and first-order difference preprocessing on multi-source physiological data of patients in the anesthesia recovery period; Step 2: Construct a multimodal state fusion module based on the cross-attention mechanism. This module generates a fusion representation vector containing respiratory, circulatory, drug features and patient state through three parallel feature encoding branches: high-frequency respiratory feature embedding, circulatory and drug feature embedding, and discrete state encoding embedding, as well as an attention fusion unit. Step 3: Construct a sequence decision-making model based on Decision Transformer, and use medical priors to construct a comprehensive reward function to jointly optimize the model in order to learn long-term contextual dependencies and ventilation strategies; Step four: In the inference phase, a target reward value representing smooth extubation is set, and hard rule constraints and smoothing are applied to the initial strategy generated by the model to finally obtain the real-time switching command between mechanical ventilation and spontaneous breathing.
2. The ventilation control method during the anesthesia recovery period according to claim 1, characterized in that: The specific method for waveform decomposition, spatiotemporal alignment, and first-order difference preprocessing of multi-source physiological data from patients in the anesthesia recovery period in step one is as follows: A variational mode decomposition algorithm is introduced to process the acquired high-frequency airway pressure waveform. By decomposing the complex waveform signal into several intrinsic mode functions, the low-frequency component representing the main wave of mechanical ventilation and the high-frequency component representing the patient's weak spontaneous breathing attempts are separated. Subsequently, to address the issue of inconsistent sampling of multi-source data such as airway pressure, drug concentration, and cardiovascular indicators, a unified 1Hz time skeleton is constructed to map data of different frequencies onto the same time axis for rigid spatiotemporal alignment. Then, first-order difference calculation is performed on the drug concentration data to obtain the current concentration... Based on this, the concentration change rate was further obtained. This is to determine whether the drug is currently in the washout or maintenance phase.
3. The ventilation control method during the anesthesia recovery period according to claim 1 or 2, characterized in that: The specific method for generating the fused representation vector containing respiratory, circulatory, drug features, and patient status in step two is as follows: The high-frequency respiratory feature embedding branch uses 1D-CNN to extract features from the preprocessed airway pressure waveform features and respiratory mechanics indices, and maps the captured local morphological changes into a high-dimensional respiratory embedding vector. As a core representation of transient respiratory state, the circulatory and drug feature embedding branch uses a multilayer perceptron to map low-frequency updated heart rate, blood oxygen saturation, and differentially processed drug effect-room concentration into circulatory drug embedding vectors. This provides the model with physiological context. The discrete-state encoding branch targets intubation and spontaneous breathing states, as well as cumulative apnea time. It maps discrete features and concatenates them into a dense, low-dimensional discrete-state vector through a learnable embedding layer. Then, and Channel dimension aggregation is performed to form a global context vector. .
4. The ventilation control method during the anesthesia recovery period according to claim 3, characterized in that: Step three also employs a cross-attention mechanism to address the modeling problem of complex, long-term coupling relationships, specifically as follows: Using... For query vector, Given key and value vectors, the model uses dot product attention scores to dynamically retrieve relevant contextual information from drug metabolism and patient state based on the current respiratory waveform morphology, ultimately obtaining the fusion result. .make If the scaling factor represents the dimension of the key vector, then this step can be expressed as follows: 。 5. The ventilation control method during the anesthesia recovery period according to claim 1 or 2, characterized in that: The specific method for constructing the sequence decision model based on Decision Transformer in step three is as follows: Construct a sequence decision model based on Decision Transformer. In the model, for At that moment, For the patient's fusion state vector, Ventilation procedures were performed. For residual reward, the goal of the model is to predict the optimal action given the current state and expected reward. Then, the training sequence is constructed as follows. : ; Through the constructed training sequences Train the model.
6. The ventilation control method during the anesthesia recovery period according to claim 5, characterized in that: The specific method for jointly optimizing the model using a comprehensive reward function constructed with medical priors in step three is as follows: The comprehensive reward function is defined as follows: Among them, security rewards The aim is to punish extreme physiological conditions that threaten the patient's life and to maintain the patient's blood oxygen saturation. and cycle stability; After that, This is an indicator function; it is set to 1 when the heart rate or systolic blood pressure deviates from the baseline by more than 20%. and If the weight is a hyperparameter, then the reward function for this item can be defined as follows: Among them, ventilation quality awards The aim is to dynamically assess the rationality of apnea based on the drug metabolism background and guide the model to stop mechanical ventilation in a timely manner; Next, order This refers to the partial pressure of carbon dioxide at the end of expiration. For the duration of apnea, This refers to the remifentanil effect-room concentration. , If the weights are hyperparameters, then the reward function can be defined as follows: Among them, respiratory mechanics and comfort reward The aim is to address the human-machine interaction problem and encourage smooth, spontaneous breathing; Furthermore, let Peak inhalation pressure, The variance of the high-frequency components obtained from variational mode decomposition. This is an indicator function for stable tidal volume, and it is set to 1 when the fluctuation range of tidal volume is less than 20% of the average value and the average tidal volume is greater than 300 ml. , , If the weights are hyperparameters, then the reward function can be defined piecewise according to the ventilation status as follows: Among them, motion smoothness penalty The aim is to prevent the model from frequently switching between mechanical ventilation and spontaneous breathing modes, and to encourage the model to only make changes when it is certain that a change in ventilation mode is necessary, thereby ensuring the consistency of clinical implementation. Finally, let If the weight hyperparameter is denoted as , then the penalty can be defined as follows: 。 7. The ventilation control method during the anesthesia recovery period according to claim 1 or 2, characterized in that: In step four, the target reward value representing smooth extubation is set during the inference phase, and hard rule constraints and smoothing are applied to the initial strategy generated by the model. The specific method for obtaining the real-time switching instruction between mechanical ventilation and spontaneous breathing is as follows: First, 1.1 times the highest cumulative reward in the training set is used as the high target reward value that can represent "perfect and smooth extubation". Then, the model uses the self-attention mechanism of the Decision Transformer to perform autoregressive inference based on the fused state vector sequence, historical action sequence and the expected reward within the time period, and predicts whether to continue mechanical ventilation or switch to spontaneous breathing at the current moment.
8. The ventilation control method during the anesthesia recovery period according to claim 7, characterized in that: In step four, a rule constraint layer is introduced to verify the model output to ensure absolute clinical safety. Specifically, regardless of the model's prediction results, if the patient's vital signs reach any of the following clinical red lines, the system will forcibly override the control command to mechanical ventilation to ensure immediate intervention in dangerous situations. These include: blood oxygen saturation below 92% for more than 5 seconds, end-tidal carbon dioxide partial pressure exceeding 55 mmHg, cumulative apnea time exceeding 30 seconds, or peak airway pressure exceeding 35 cmH2O.
9. The ventilation control method during the anesthesia recovery period according to claim 8, characterized in that: In step four, a hysteresis comparator is also introduced, and asymmetric switching logic is used to smooth the action.