Intelligent optimization method and system for personalized intraoperative blood pressure management

By collecting real-time multimodal physiological data to construct an individual hemodynamic state vector and using a deep learning model for closed-loop optimization, the problems of blood pressure regulation lag and individual differences in existing technologies have been solved, thereby improving the accuracy and safety of personalized intraoperative blood pressure management.

CN120977472AInactive Publication Date: 2025-11-18THE THIRD AFFILIATED HOSPITAL OF SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511353292.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have unresolved technical problems in intraoperative blood pressure management. They cannot effectively address the rapid adaptive mechanisms that address individual differences in anatomy, physiology, and medication response. This results in blood pressure regulation being highly dependent on human experience, with lag in adjustment and an inability to accurately match individual differences. Consequently, existing technologies cannot achieve true end-to-end intelligent optimization control.

Method used

By continuously collecting multimodal real-time physiological data, a state vector reflecting individual hemodynamics is constructed. Using an offline pre-trained and online adaptive blood pressure optimization decision model, regulatory instructions for vasoactive drugs and infusions are generated and adjusted through the actuator to form a closed-loop optimization control.

Benefits of technology

It improves the accuracy, adaptability and safety of personalized intraoperative blood pressure management, significantly reduces abnormal blood pressure fluctuations, and improves the real-time and stability of intraoperative blood pressure control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977472A_ABST
    Figure CN120977472A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent optimization method and system for personalized intraoperative blood pressure management, and relates to the technical field of intraoperative monitoring. The method comprises the following steps: continuously collecting multi-modal real-time physiological data for describing the circulation state of a patient in an operation process; based on the multi-modal real-time physiological data, constructing a state vector reflecting individual haemodynamics; inputting the state vector into a blood pressure optimization decision model subjected to offline pre-training and online self-adaption to obtain a regulation and control instruction aiming at the vasoactive medicine and / or infusion; adjusting the administration or infusion rate through an execution mechanism according to the regulation and control instruction; and updating the state vector according to the multi-modal real-time physiological data obtained after adjustment, and adaptively correcting the blood pressure optimization decision model to form closed-loop optimization control. According to the invention, the precision, the adaptability and the safety of intraoperative blood pressure control can be obviously improved, and the method has a good clinical application prospect and expansibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intraoperative monitoring technology, and in particular to an intelligent optimization method and system for personalized intraoperative blood pressure management. Background Technology

[0002] In recent years, the widespread clinical use of continuous invasive or non-invasive hemodynamic monitoring devices has enabled anesthesiologists to obtain intraoperative mean arterial pressure (MAP), cardiac output, and other indicators with second-level resolution. However, in current mainstream practice, anesthesiologists still need to rely on experience to adjust the infusion rate of vasoactive drugs or intravenous fluids in an "open-loop" manner. When the patient's circulatory status fluctuates rapidly, the patient-ventilator coupling lag often causes transient hypoperfusion or hyperperfusion, increasing the risk of ischemia-reperfusion injury to organs such as the heart, brain, and kidneys.

[0003] With advancements in artificial intelligence (AI) technology, academia and industry are exploring the introduction of deep learning and reinforcement learning models into intraoperative monitoring to achieve real-time trend prediction and decision support. Some studies have already shown promise in areas such as instrument recognition and intraoperative complication warning, demonstrating significant potential for improving surgical safety. Future development trends include integrating multimodal physiological data, digital twins, and closed-loop drug delivery to dynamically generate optimal blood pressure targets and dosing strategies for each patient.

[0004] Existing attempts mostly remain at the level of "group average" or empirical thresholds, lacking rapid adaptive mechanisms to address individual differences in anatomy, physiology, and drug response. Existing prototype systems are mostly "suggestive" semi-closed loops, and have not yet formed true end-to-end intelligent optimization control. Data drift, model interpretability, and failure protection are also key pain points hindering its large-scale clinical deployment. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide an intelligent optimization method and system for personalized intraoperative blood pressure management, which significantly improves the accuracy, adaptability and safety of intraoperative blood pressure control, and has good clinical application prospects and scalability.

[0006] To achieve the above objectives, the present invention provides the following solution: A smart optimization method for personalized intraoperative blood pressure management includes: During the surgery, multimodal real-time physiological data describing the patient's circulatory status were continuously collected; A state vector reflecting individual hemodynamics is constructed based on the aforementioned multimodal real-time physiological data; The state vector is input into a blood pressure optimization decision model that is pre-trained offline and adaptive online to obtain control instructions for vasoactive drugs and / or infusions. The administration or infusion rate is adjusted by the actuator according to the control instructions; The state vector is updated based on the adjusted multimodal real-time physiological data, and the blood pressure optimization decision model is adaptively corrected to form a closed-loop optimization control.

[0007] Preferably, the multimodal real-time physiological data includes: arterial blood pressure waveform, electrocardiogram, pulse wave conduction time, anesthesia depth index, and drug administration record.

[0008] Preferably, during the surgery, multimodal real-time physiological data describing the patient's circulatory status are continuously collected, including: After the patient enters the room, an arterial catheter is placed and connected to an invasive blood pressure monitoring module to collect the arterial blood pressure waveform in real time at a frequency of 1Hz to 20Hz. The electrocardiogram signal was acquired using standard electrocardiogram monitoring leads and simultaneously acquired with the arterial pressure signal. The pulse wave conduction time is calculated by comparing the time relationship between the electrocardiogram signal and the arterial pressure signal, and updated according to a set period. The anesthesia depth index was collected using a bifrequency electroencephalogram (EEG) index monitoring device, and a timestamp was recorded simultaneously. The drug administration record is obtained by extracting the information on vasoactive drugs, fluid types, and infusion rates recorded in the clinical workstation.

[0009] Preferably, constructing a state vector reflecting individual hemodynamics based on the multimodal real-time physiological data includes: The multimodal real-time physiological data obtained from each channel were resampled in the range of 1 Hz to 20 Hz according to a unified timestamp, and short-term missing values ​​were deleted or imputed. Bandpass filtering and artifact removal are applied to the arterial blood pressure waveform and the electrocardiogram signal to retain the effective physiological frequency band. Upper and lower limit checks are performed on the anesthesia depth index and the drug administration record to remove out-of-range or invalid inputs, thus obtaining preprocessed data. Within each sampling period, scalar features of the preprocessed data are extracted; the scalar features include: current mean arterial pressure, heart rate, pulse wave transit time, raw reading of the anesthesia depth index, and instantaneous infusion rate of each vasoactive drug and infusion fluid; The scalar features are concatenated according to a preset field order to form a pre-feature vector F with fixed dimensions. t ; Take the most recent N consecutive pre-eigenvectors {F t₋n₊1 … F t The input is a time-series coding network that has been pre-trained offline and allows for online fine-tuning, and the output is a state vector of a preset dimension to characterize the real-time hemodynamic state of an individual.

[0010] Preferably, the state vector is input into a blood pressure optimization decision model that has been pre-trained offline and adaptively adapted online to obtain regulatory instructions for vasoactive drugs and / or infusions, including: A blood pressure optimization strategy network is loaded and pre-trained offline using multi-center historical intraoperative data; the blood pressure optimization strategy network includes at least an actor-critic structure; wherein the actor subnet of the actor-critic structure outputs the control action, and the critic subnet of the actor-critic structure evaluates the reward value; Receive the state vector and the control action A from the previous cycle. t₋1 and rewards R t₋1 Push it into the observation buffer and maintain a fixed-length time window; The time-series prediction module built into the blood pressure optimization strategy network is invoked to perform rolling predictions of the mean arterial pressure at future ΔT time intervals, resulting in a prediction sequence. ; Based on individual target mean arterial pressure range [MAP] min MAP max Calculate the prediction bias vector D t = -Target interval center value; The current state vector S is processed according to the preset field order. t Prediction bias D t And the previous cycle action A t₋1 Concatenate into an enhanced state vector X t , used as joint input for actor-critic networks; Using the actor subnet for X t Perform forward inference and output an initial control action vector; wherein each component in the initial control action vector corresponds to the vasoactive drug dose increment and / or infusion rate adjustment. The initial control action vector is input into the safety gating module, and then trimmed component by component based on the maximum drug dose, single-step dose change limit, and specific contraindication parameters of the patient to obtain the final executable control command vector.

[0011] Preferably, after obtaining the regulatory instructions for the vasoactive drug and / or infusion, the method further includes: Execute A t Then, the actual mean arterial pressure response was collected to calculate the reward function. ;in, Indicates at time step Instant reward value; This represents the mean arterial pressure collected at the current time step; Target mean arterial pressure value set for individualization; Indicates the first The magnitude of dose or rate adjustment of the regulatory variable at the current time step; The total number of all controllable variables; It is a drug regulation penalty factor used to limit the drasticness of dose changes; For safety violation indication functions, when the current mean arterial pressure The value is 1 when the value exceeds the set safety range, otherwise it is 0. This is a safety penalty factor used to increase the penalty weight of the system for severe high and low blood pressure events; Based on time difference error The calculation strategy evaluates the residuals; among which, This indicates that the critic subnetwork has a view of the current state vector. Value estimation; This represents an estimate of the updated state; This is a discount factor used to control the impact of future returns on current strategy updates; by The parameters of the commentator subnetwork are updated using gradient descent as the loss term; Using the policy gradient algorithm, to To update the parameters of the actor subnetwork Incremental gradient adjustments are performed to enhance the current strategy's advantage in this state, enabling online adaptive optimization of the blood pressure regulation strategy network.

[0012] Preferably, adjusting the drug administration or infusion rate by an actuator according to the control command includes: The vector of the control command is converted into a control command frame conforming to the communication protocol format of the actuator by the data encoding module; the control command frame includes a channel number field for identifying the target channel and a corresponding target infusion rate or dose increment field; The control command frame is sent to the actuator connected to the system via a communication interface; the communication interface is a serial communication interface, a control bus interface, or an Ethernet communication interface; the actuator includes a micro-injection pump, an infusion pump, or an equivalent actuator. Based on the actuator, the target parameters are parsed according to the received control command frame, and the current drug administration rate or infusion rate is updated to complete the rate adjustment based on the control command. The actuator is equipped with a control limiting module, which is used to verify the adjustment value of each drug administration or infusion rate. When the adjustment value exceeds the preset dose change threshold, the adjustment value is automatically cut off to the threshold range. After the rate adjustment is completed, the actuator feeds back the current execution status to the central control unit; the current execution status includes the time of receiving the control command, the target parameters, the actual adjustment result, and the device response status code.

[0013] Preferably, the state vector is updated based on the adjusted multimodal real-time physiological data, and the blood pressure optimization decision model is adaptively corrected to form a closed-loop optimization control, including: Collect multimodal real-time physiological data acquired in the current cycle; The multimodal real-time physiological data is subjected to time alignment, missing data processing and numerical standardization operations, and instantaneous feature values ​​are extracted to construct the pre-feature vector for the current moment; The pre-feature vectors from multiple consecutive cycles are input into a temporal coding network, which outputs an updated state vector to reflect the current dynamic cyclic state of the patient. Based on the state vector and the actual regulatory actions and immediate rewards generated in the current cycle, a policy gradient signal and value function residual are constructed to update the blood pressure optimization decision model. Incremental parameter updates are performed on the strategy subnetwork and value evaluation subnetwork in the blood pressure optimization decision model, respectively, with the goal of enhancing the stability and effectiveness of the regulation strategy under the current state. The updated state vector and the blood pressure optimization decision model are used as inputs for the next cycle of strategy reasoning to achieve the periodic closed-loop execution of the blood pressure optimization control process.

[0014] A personalized intelligent optimization system for intraoperative blood pressure management includes: The data acquisition unit allows users to continuously collect multimodal real-time physiological data describing the patient's circulatory status during surgery. A vector construction unit is used to construct a state vector reflecting individual hemodynamics based on the multimodal real-time physiological data. The instruction generation unit is used to input the state vector into a blood pressure optimization decision model that has been pre-trained offline and adaptively adapted online to obtain control instructions for vasoactive drugs and / or infusions. A rate regulating unit is used to adjust the drug administration or infusion rate according to the control command via an actuator; The closed-loop optimization unit is used to update the state vector based on the adjusted multimodal real-time physiological data and adaptively correct the blood pressure optimization decision model to form closed-loop optimization control.

[0015] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: This invention provides an intelligent optimization method for personalized intraoperative blood pressure management. Addressing the problems of existing technologies, such as high reliance on human experience, regulatory lag, and inability to accurately match individual differences, this method introduces state modeling driven by multimodal real-time physiological data and deep reinforcement learning optimization strategies to achieve a closed-loop blood pressure regulation mechanism tailored to individual hemodynamic characteristics. This method not only integrates mean arterial pressure, ECG, pulse wave transit time, anesthesia depth index, and drug administration information during the data perception stage to improve the accuracy of state characterization, but also integrates short-term blood pressure trend prediction and online incremental learning during the decision-making process to continuously optimize the regulation strategy and effectively reduce the frequency of abnormal blood pressure fluctuations. Through automatic docking of control commands and actuators, the system can complete dose fine-tuning within safe limits at a rate of seconds, enhancing the real-time performance and stability of regulation. Compared with existing experience-based or semi-closed-loop models, this invention significantly improves the accuracy, adaptability, and safety of intraoperative blood pressure control, demonstrating promising clinical application prospects and scalability. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of the method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the system structure provided in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] The purpose of this invention is to provide an intelligent optimization method and system for personalized intraoperative blood pressure management, which significantly improves the accuracy, adaptability and safety of intraoperative blood pressure control, and has good clinical application prospects and scalability.

[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] Figure 1A flowchart of the method provided in the embodiments of the present invention, such as Figure 1 As shown, this invention provides an intelligent optimization method for personalized intraoperative blood pressure management, comprising: Step 100: Continuously collect multimodal real-time physiological data describing the patient's circulatory status during the operation; Step 200: Construct a state vector reflecting individual hemodynamics based on multimodal real-time physiological data; Step 300: Input the state vector into the blood pressure optimization decision model that has been pre-trained offline and adapted online to obtain the control instructions for vasoactive drugs and / or infusions; Step 400: Adjust the drug administration or infusion rate according to the control instructions via the actuator; Step 500: Update the state vector based on the adjusted multimodal real-time physiological data and adaptively correct the blood pressure optimization decision model to form a closed-loop optimization control.

[0022] Specifically, the multimodal real-time physiological data includes the following five key channels: Arterial blood pressure waveform (ABP): Acquired through invasive arterial catheterization, continuously sampled at a frequency of not less than 1 Hz. The sampling results include the original pressure waveform sequence and the mean arterial pressure (MAP) per cycle. As a core indicator for assessing circulatory perfusion status, it is also the main supervision target for subsequent reward function and prediction network training.

[0023] Electrocardiogram (ECG) signal: Acquire standard lead II ECG signals, and use the R wave identification time point for subsequent pulse wave conduction time calculation and heart rate variability analysis. The heart rate information obtained through this channel can dynamically reflect the level of sympathetic activity and indirectly infer the trend of vascular tone changes.

[0024] Pulse wave conduction time (PTT): Automatically calculated based on the time difference between the R-wave initiation and the rising edge of the MAP waveform on ECG. It serves as a key feature for assessing arterial compliance and vascular response delay, and is especially used to reflect the time delay characteristics of drug onset. It is an important physiological parameter that is missing from traditional monitoring methods.

[0025] Anesthesia depth index (such as BIS or entropy value): This quantitative indicator of anesthesia depth is collected from EEG monitoring equipment and time-aligned with blood pressure regulation behavior to avoid misjudgments caused by changes in the anesthesia level. This channel is used to distinguish whether blood pressure fluctuations are dominated by circulatory factors or central inhibitory factors, enhancing the causal modeling ability of the strategy network.

[0026] Dosing records: By interfacing with the anesthesia information system or smart infusion pump, the infusion information, such as the type, rate, and start time of vasoactive drugs and fluids, is obtained in real time, and timestamp calibration is performed locally to determine the actual window of action of the regulatory behavior.

[0027] Preferably, during the surgery, multimodal real-time physiological data describing the patient's circulatory status are continuously collected, including: After the patient enters the room, an arterial catheter is placed and connected to an invasive blood pressure monitoring module to collect the arterial blood pressure waveform in real time at a frequency of 1Hz to 20Hz. The electrocardiogram signal was acquired using standard electrocardiogram monitoring leads and simultaneously acquired with the arterial pressure signal. The pulse wave conduction time is calculated by comparing the time relationship between the electrocardiogram signal and the arterial pressure signal, and updated according to a set period. The anesthesia depth index was collected using a bifrequency electroencephalogram (EEG) index monitoring device, and a timestamp was recorded simultaneously. The drug administration record is obtained by extracting the information on vasoactive drugs and fluid infusion rates recorded in the clinical workstation.

[0028] In this embodiment, to support the state modeling and real-time evaluation of subsequent personalized blood pressure control strategies, a multi-source continuous data acquisition process is initiated immediately after the patient enters the operating room. Specifically, an arterial blood pressure waveform is acquired at a sampling frequency of at least 1 Hz via arterial catheterization and connection to an invasive blood pressure monitoring module, and the R-wave time point is extracted by combining it with a standard lead II electrocardiogram signal. This embodiment dynamically calculates the pulse wave conduction time using the time difference between the R-wave initiation and the rising edge of the pressure wave, and updates this value at a fixed period to reflect individual vascular compliance and circulatory response delay. The time references of all the above signals are unified to millisecond-level timestamps, achieving cross-channel alignment and providing a precise physiological basis for constructing an interpretable time-series state vector.

[0029] To further enhance the modeling capability for non-circulatory factors, this embodiment integrates an anesthesia depth monitoring device to collect bispectral indexes of electroencephalograms in real time and record corresponding timestamps, which are then sent to the data buffer along with circulatory data. Simultaneously, the types of vasoactive drugs and their corresponding infusion rates recorded in the clinical workstation are extracted in real time via a communication interface, forming a complete drug administration record channel. To ensure the input integrity of closed-loop control, this embodiment performs standardization, interpolation completion, and dynamic window pruning of all channel data at the local node, ensuring that each observation cycle has a synchronous and complete multimodal input structure for subsequent state vector calculations.

[0030] Preferably, constructing a state vector reflecting individual hemodynamics based on the multimodal real-time physiological data includes: The multimodal real-time physiological data obtained from each channel were resampled in the range of 1 Hz to 20 Hz according to a unified timestamp, and short-term missing values ​​were deleted or imputed. Bandpass filtering and artifact removal are applied to the arterial blood pressure waveform and the electrocardiogram signal to retain the effective physiological frequency band. Upper and lower limit checks are performed on the anesthesia depth index and the drug administration record to remove out-of-range or invalid inputs, thus obtaining preprocessed data. Within each sampling period, scalar features of the preprocessed data are extracted; the scalar features include: current mean arterial pressure, heart rate, pulse wave transit time, raw reading of the anesthesia depth index, and instantaneous infusion rate of each vasoactive drug and infusion fluid; The scalar features are concatenated according to a preset field order to form a pre-feature vector F with fixed dimensions. t ; Take the most recent N consecutive pre-eigenvectors {F t₋n₊1 … F t The input is a time-series coding network that has been pre-trained offline and allows for online fine-tuning, and the output is a state vector of a preset dimension to characterize the real-time hemodynamic state of an individual.

[0031] In this embodiment, to construct a state vector that accurately characterizes the patient's current hemodynamic state, the acquired multimodal real-time physiological data is first uniformly preprocessed. To eliminate alignment errors caused by inconsistent sampling frequencies across different channels, the embodiment resamples the data from each channel according to a unified timestamp, with the sampling frequency set between once per second and twenty times per second. Simultaneously, to ensure the stability of subsequent feature extraction, signal gaps within short time periods are filled using interpolation, and outliers caused by sudden artifacts or physical interference are deleted or corrected. For data cleaning, bandpass filtering and artifact removal algorithms are applied to pressure waveforms and ECG signals respectively to filter out interference from non-physiological frequency bands; for anesthesia depth index and drug infusion information, reasonable physiological range thresholds are set to automatically remove data points exceeding the range or with incorrect formats. The output of the above steps is a structured, cleaned data sequence.

[0032] Subsequently, within each sampling period, scalar features with clear clinical and physiological significance are extracted from the cleaned multi-channel data, including the current mean arterial pressure, heart rate estimated based on ECG intervals, pulse wave conduction time calculated based on waveform delay, real-time anesthesia depth index reading, and instantaneous infusion rates of various drugs and fluids. In this embodiment, the above scalar features are sequentially concatenated according to a uniform field order to generate a fixed-dimensional pre-feature vector sequence. To model the potential trend of the patient's circulatory state evolution over time, the pre-feature vectors from multiple consecutive time points are input into a pre-trained temporal coding network with online fine-tuning capabilities. This network outputs a state vector for strategy decision-making. This state vector, as the core expression structure characterizing the patient's hemodynamic state at the current moment, constitutes the central information input of the closed-loop control logic of this invention.

[0033] Preferably, the state vector is input into a blood pressure optimization decision model that has been pre-trained offline and adaptively adapted online to obtain regulatory instructions for vasoactive drugs and / or infusions, including: A blood pressure optimization strategy network is loaded and pre-trained offline using multi-center historical intraoperative data; the blood pressure optimization strategy network includes at least an actor-critic structure; wherein the actor subnet of the actor-critic structure outputs the control action, and the critic subnet of the actor-critic structure evaluates the reward value; Receive the state vector and the control action A from the previous cycle. t₋1 and rewards R t₋1 Push it into the observation buffer and maintain a fixed-length time window; The time-series prediction module built into the blood pressure optimization strategy network is invoked to perform rolling predictions of the mean arterial pressure at future ΔT time intervals, resulting in a prediction sequence. ; Based on individual target mean arterial pressure range [MAP] min MAP max Calculate the prediction bias vector D t = -Target interval center value; The current state vector S is processed according to the preset field order. t Prediction bias D t And the previous cycle action A t₋1 Concatenate into an enhanced state vector X t , used as joint input for actor-critic networks; Using the actor subnet for X t Perform forward inference and output an initial control action vector; wherein each component in the initial control action vector corresponds to the vasoactive drug dose increment and / or infusion rate adjustment. The initial control action vector is input into the safety gating module, and then trimmed component by component based on the maximum drug dose, single-step dose change limit, and specific contraindication parameters of the patient to obtain the final executable control command vector.

[0034] In this embodiment, to achieve the goals of individualized, predictive, and safe controllable regulation of intraoperative blood pressure, a data-driven pre-trained blood pressure optimization strategy network is used as the core of intelligent decision-making. This network receives the currently constructed state vector as input and generates regulatory outputs by combining historical behavior and predicted trends. The strategy network adopts a joint structure with both strategy output and value evaluation capabilities, including a sub-network for generating regulatory actions and a sub-network for evaluating regulatory effects. This structure is trained offline before surgery using large-scale, multi-center intraoperative monitoring data, enabling it to extract time-dependent features and establish a blood pressure response mechanism.

[0035] To enhance the network's sensitivity to short-term dynamic trends and improve response accuracy, this embodiment, in each control cycle, feeds the current state vector, the control action from the previous cycle, and the actual real-time feedback results into the network's observation buffer, constructing a fixed-time window within it. This window serves as a memory unit, providing a behavioral-response historical context for subsequent decisions. Simultaneously, the time-series prediction module integrated into the policy network is invoked to perform rolling predictions of mean arterial pressure over a future period. Combined with individually set blood pressure target intervals, a prediction bias reflecting expected risk is constructed. The aforementioned state vector, prediction bias, and behavioral history are concatenated using a unified field to form an enhanced state input, which is then fed into the policy network to generate the initial control action. Each dimension of this action vector corresponds to the rate adjustment increment of the target drug channel or fluid channel, guiding the control behavior in the next cycle.

[0036] To prevent the model output from containing excessive dosage or potential contraindications, this embodiment includes a safety gating module after the strategy output. This module limits and prunes each regulatory action to be executed. The pruning criteria include the maximum allowable dose threshold for the current patient, the limit on the dose variation range in a single cycle, and the individual drug contraindication list. This ensures that each regulatory output is both strategy-optimal and meets clinical safety boundaries. The pruned actions serve as the final regulatory instructions, which are then sent to the execution system to complete the actual drug administration and infusion, achieving continuous closed-loop optimization of intraoperative blood pressure control.

[0037] Preferably, after obtaining the regulatory instructions for the vasoactive drug and / or infusion, the method further includes: Execute A t Then, the actual mean arterial pressure response was collected to calculate the reward function. ;in, Indicates at time step Instant reward value; This represents the mean arterial pressure collected at the current time step; Target mean arterial pressure value set for individualization; Indicates the first The magnitude of dose or rate adjustment of the regulatory variable at the current time step; The total number of all controllable variables; It is a drug regulation penalty factor used to limit the drasticness of dose changes; For safety violation indication functions, when the current mean arterial pressure The value is 1 when the value exceeds the set safety range, otherwise it is 0. This is a safety penalty factor used to increase the penalty weight of the system for severe high and low blood pressure events; Based on time difference error The calculation strategy evaluates the residuals; among which, This indicates that the critic subnetwork has a view of the current state vector. Value estimation; This represents an estimate of the updated state; This is a discount factor used to control the impact of future returns on current strategy updates; by The parameters of the commentator subnetwork are updated using gradient descent as the loss term; Using the policy gradient algorithm, to To update the parameters of the actor subnetwork Incremental gradient adjustments are performed to enhance the current strategy's advantage in this state, enabling online adaptive optimization of the blood pressure regulation strategy network.

[0038] In this embodiment, after the regulatory action is executed, to achieve online adaptive optimization of the strategy network, an immediate reward function and strategy residual feedback need to be constructed based on the patient's actual physiological response. Specifically, the mean arterial pressure value at the current time step is collected in real-time at a rate of seconds by the invasive blood pressure monitoring module connected to the system; the individual target mean arterial pressure value is set by the anesthesiologist before surgery based on the patient's baseline blood flow status and is written as a static parameter into the strategy configuration table; the dose or rate adjustment range of each regulatory variable is recorded by the system control module, which is the difference between the strategy output and the parameters of the previous cycle; the total number of regulatory variables is determined by the number of channels defined during system deployment. The drug regulation penalty factor is set as a hyperparameter during the model training phase and can be individually adjusted intraoperatively based on the patient's weight and ASA classification. The safety violation indication function is automatically assigned a value by the system by comparing the real-time blood pressure value with the set safety range, and the safety penalty factor is tuned and fixed as a strategy constraint parameter during the strategy training phase. The state value estimate in policy update is calculated by the commentator subnetwork with the current state vector as input. The updated state is composed of the state vector of the next time step after the execution of the adjustment action, and its value is obtained by forward propagation of the same subnetwork. The discount factor is a fixed real number less than 1, which is set in the offline training phase to control the weight of the long-term reward.

[0039] In this embodiment, after executing the control instructions output by the blood pressure optimization strategy network, the network model is further updated online through a reinforcement learning mechanism to improve its adaptability to individual patient responses and the real-time optimality of the control strategy. Specifically, after each control cycle, the mean arterial pressure monitoring value is collected after the execution of the action, and an immediate reward value is calculated based on the control action taken in that cycle. The reward function consists of three parts: first, the deviation of the current mean arterial pressure from the center value of the individualized target range, used to measure the accuracy of the control target; second, the intensity of dose or rate change of all control variables in this cycle, used to quantify the control amplitude; and third, the identification index for whether an unsafe blood pressure range has been entered. This index is judged and assigned a value in real time by the safety gating module to strengthen the penalty feedback for overpressure or hypoperfusion states. The individual target mean arterial pressure value is derived from the preoperative setting by the anesthesiologist based on the patient's medical history, body type, and baseline circulatory status, and is written into the system parameter configuration table. The adjustment amplitude of the control variables is recorded in real time by the control module after each execution of the control instruction, and the safety violation judgment is automatically judged by the system based on the preset upper and lower blood pressure ranges to ensure that key parameters are quantifiable and traceable.

[0040] Based on the aforementioned reward values, a time difference error evaluation signal is constructed, and combined with the current state vector and the new state vector formed after adjustment, the update basis for the policy network is established. In this embodiment, the value assessment sub-network is invoked to score the value of the current state and the updated state respectively, and a residual term is constructed by combining it with a set future return discount factor. This residual, as the core component of the loss function, drives the gradient descent update of the value network, ensuring that its value prediction aligns with the actual return. Subsequently, this embodiment uses this residual as the advantage function multiplied by the gradient of the current policy to update the parameters in the policy output sub-network, achieving fine-tuning of the adjustment action selection tendency in the current state. All of the above parameter update processes are completed in the local edge computing unit, without relying on remote server calls, and the update magnitude in each round is limited by a preset learning rate and safety boundary constraints, ensuring the asymptotic stability of the behavior in this embodiment. Through this online incremental learning mechanism, the blood pressure optimization strategy can gradually match the dynamic changes in the individual's response to medication and infusion, thereby achieving continuous self-optimization of the personalized blood pressure control strategy.

[0041] Preferably, adjusting the drug administration or infusion rate by an actuator according to the control command includes: The vector of the control command is converted into a control command frame conforming to the communication protocol format of the actuator by the data encoding module; the control command frame includes a channel number field for identifying the target channel and a corresponding target infusion rate or dose increment field; The control command frame is sent to the actuator connected to the system via a communication interface; the communication interface is a serial communication interface, a control bus interface, or an Ethernet communication interface; the actuator includes a micro-injection pump, an infusion pump, or an equivalent actuator. Based on the actuator, the target parameters are parsed according to the received control command frame, and the current drug administration rate or infusion rate is updated to complete the rate adjustment based on the control command. The actuator is equipped with a control limiting module, which is used to verify the adjustment value of each drug administration or infusion rate. When the adjustment value exceeds the preset dose change threshold, the adjustment value is automatically cut off to the threshold range. After the rate adjustment is completed, the actuator feeds back the current execution status to the central control unit; the current execution status includes the time of receiving the control command, the target parameters, the actual adjustment result, and the device response status code.

[0042] In this embodiment, after generating the control instructions from the strategy model, these instructions need to be efficiently and securely transmitted to the actual drug infusion device to perform refined blood pressure regulation. To this end, the control instruction vector is first converted into a data frame structure conforming to the actuator's communication protocol via a data encoding module. The control command frame contains at least two core fields: a channel number field to identify the target drug or fluid infusion channel, and a target parameter field to characterize the required dose or infusion rate increment for the current cycle. The data frame is sent to the actuator via a system-integrated communication interface that supports serial communication, industrial control bus, or Ethernet protocols to ensure system compatibility with various device types. The actuator is a digitally controlled drug delivery device, including a micro-injection pump, an infusion pump, or a device with equivalent execution precision.

[0043] Upon receiving a control command frame, the actuator parses and extracts the channel identifier and target parameters from the command frame, updates the current output rate based on the parameter values, and achieves precise rate adjustment. To ensure that the system maintains a safety boundary during high-frequency control, this embodiment deploys a limiting module at the actuator end to perform real-time amplitude verification on each dose adjustment value. When the target adjustment value of a certain control channel exceeds the set single-cycle dose change threshold, the system automatically trims it to the maximum allowable range to prevent overdose caused by model output fluctuations or communication anomalies. After the adjustment is completed, the actuator sends the execution status back to the central control unit. The feedback information includes the control command reception timestamp, parsed target parameters, actual execution results, and response status code. This feedback mechanism supports control log recording and fault tracing, and can also provide input for state reasoning and strategy correction in the next cycle, realizing a complete closed loop of the system control chain.

[0044] Preferably, the state vector is updated based on the adjusted multimodal real-time physiological data, and the blood pressure optimization decision model is adaptively corrected to form a closed-loop optimization control, including: Collect multimodal real-time physiological data acquired in the current cycle; The multimodal real-time physiological data is subjected to time alignment, missing data processing and numerical standardization operations, and instantaneous feature values ​​are extracted to construct the pre-feature vector for the current moment; The pre-feature vectors from multiple consecutive cycles are input into a temporal coding network, which outputs an updated state vector to reflect the current dynamic cyclic state of the patient. Based on the state vector and the actual regulatory actions and immediate rewards generated in the current cycle, a policy gradient signal and value function residual are constructed to update the blood pressure optimization decision model. Incremental parameter updates are performed on the strategy subnetwork and value evaluation subnetwork in the blood pressure optimization decision model, respectively, with the goal of enhancing the stability and effectiveness of the regulation strategy under the current state. The updated state vector and the blood pressure optimization decision model are used as inputs for the next cycle of strategy reasoning to achieve the periodic closed-loop execution of the blood pressure optimization control process.

[0045] In this embodiment, to achieve dynamic closed-loop optimization of intraoperative blood pressure control, the system initiates a state vector update and adaptive correction process for the decision model after each control cycle. First, the latest multimodal real-time physiological data acquired in the current cycle is collected, and all channel data undergoes time synchronization, missing data completion, and standardization. Key physiological feature values ​​at the current moment are extracted to form a pre-feature vector. To fully capture the potential dynamics of blood flow evolution over time, the system selects a sequence of pre-feature vectors from several consecutive cycles and inputs them into a pre-established temporal coding network to generate an updated state vector. This state vector serves as a higher-order expression of the patient's circulatory system under a specific control context in the current cycle, providing information support for the generation and execution of the next round of strategies.

[0046] While updating the state vector, the system constructs a feedback signal for policy learning based on the actual control commands executed in that cycle, the patient's mean arterial pressure response value, and the calculated immediate reward. The policy gradient signal is used to assess the direction of contribution of the control behavior to reward enhancement, and the value function residual is used to measure the model's estimation error of the expected return in the current state. The system incrementally updates the parameters of both the policy subnetwork and the value evaluation subnetwork to ensure that the policy can gradually adapt to changes in the individual's cyclical response pattern, while improving the ability to predict and control potential dynamic trends. The updated model directly participates in policy inference in the next control cycle, forming a complete closed-loop control process with state perception – policy generation – control execution – feedback update as its core structure, continuously promoting personalized, precise, and safe intraoperative blood pressure control.

[0047] Corresponding to the above methods, such as Figure 2 As shown, this embodiment also provides an intelligent optimization system for personalized intraoperative blood pressure management, including: The data acquisition unit allows users to continuously collect multimodal real-time physiological data describing the patient's circulatory status during surgery. A vector construction unit is used to construct a state vector reflecting individual hemodynamics based on the multimodal real-time physiological data. The instruction generation unit is used to input the state vector into a blood pressure optimization decision model that has been pre-trained offline and adaptively adapted online to obtain control instructions for vasoactive drugs and / or infusions. A rate regulating unit is used to adjust the drug administration or infusion rate according to the control command via an actuator; The closed-loop optimization unit is used to update the state vector based on the adjusted multimodal real-time physiological data and adaptively correct the blood pressure optimization decision model to form closed-loop optimization control.

[0048] The beneficial effects of this invention are as follows: (1) This invention introduces a multimodal real-time physiological data acquisition mechanism, which integrates core parameters such as arterial blood pressure waveform, electrocardiogram, pulse wave conduction time, anesthesia depth index, and drug administration records into a synchronous and continuous data stream. Combined with high-frequency data cleaning and feature extraction strategies, it achieves high-fidelity modeling of the patient's circulatory state. Compared with traditional control methods that mainly rely on single blood pressure values ​​or intermittent indicators, this method can construct a more physiologically complete state expression, providing a solid foundation for the accuracy and individualization of subsequent control strategies.

[0049] (2) This invention utilizes a deep policy network pre-trained offline to perform policy reasoning by combining the patient's real-time status with predicted trends during surgery, and ensures the clinical executability of the output control commands with a safety limiting mechanism. By introducing a policy structure with a short-term blood pressure trend prediction module, the system not only has the ability to respond to the current status, but also can proactively intervene in potential blood pressure fluctuations, effectively reducing the risk of intraoperative hypoperfusion and hyperperfusion, and enhancing the preventiveness and stability of intraoperative control.

[0050] (3) After the strategy control is completed, the present invention constructs an instant reward function based on the actual blood pressure response and regulatory behavior, and performs online incremental learning in combination with the time difference error feedback of the strategy network, thereby realizing continuous adaptive optimization of the regulatory strategy. By introducing an actor-critic structure, the system dynamically corrects the behavior selection and value assessment respectively, which significantly improves the robustness and accuracy of the model in the face of different individual patient response differences and intraoperative dynamic changes, and has the ability to continuously evolve and self-adjust.

[0051] (4) This invention constructs a complete closed-loop blood pressure control chain of perception-reasoning-execution-feedback, covering the entire process of state acquisition, strategy generation, instruction execution and model update, and can achieve automated and high-frequency blood pressure regulation without human intervention. The system supports interface with standard communication interfaces of clinical execution equipment, has good deployment compatibility and practical implementation capability, is suitable for precise circulation management of various high-risk patients during surgery, and has broad clinical promotion value and application prospects.

[0052] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0053] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A smart optimization method for personalized intraoperative blood pressure management, characterized in that, include: During the surgery, multimodal real-time physiological data describing the patient's circulatory status were continuously collected; A state vector reflecting individual hemodynamics is constructed based on the aforementioned multimodal real-time physiological data; The state vector is input into a blood pressure optimization decision model that is pre-trained offline and adaptive online to obtain control instructions for vasoactive drugs and / or infusions. The administration or infusion rate is adjusted by the actuator according to the control instructions; The state vector is updated based on the adjusted multimodal real-time physiological data, and the blood pressure optimization decision model is adaptively corrected to form a closed-loop optimization control.

2. The intelligent optimization method for personalized intraoperative blood pressure management according to claim 1, characterized in that, The multimodal real-time physiological data includes: arterial blood pressure waveform, electrocardiogram, pulse wave conduction time, anesthesia depth index, and infusion and drug administration records.

3. The intelligent optimization method for personalized intraoperative blood pressure management according to claim 2, characterized in that, During the surgery, multimodal real-time physiological data describing the patient's circulatory status were continuously collected, including: After the patient enters the room, an arterial catheter is placed and connected to an invasive blood pressure monitoring module to collect the arterial blood pressure waveform in real time at a frequency of 1Hz to 20Hz. The electrocardiogram signal was acquired using standard electrocardiogram monitoring leads and simultaneously acquired with the arterial pressure signal. The pulse wave conduction time is calculated by comparing the time relationship between the electrocardiogram signal and the arterial pressure signal, and updated according to a set period. The anesthesia depth index was collected using a bifrequency electroencephalogram (EEG) index monitoring device, and a timestamp was recorded simultaneously. The drug administration record is obtained by extracting the information on vasoactive drugs, fluid types, and infusion rates recorded in the clinical workstation.

4. The intelligent optimization method for personalized intraoperative blood pressure management according to claim 2, characterized in that, Based on the aforementioned multimodal real-time physiological data, a state vector reflecting individual hemodynamics is constructed, including: The multimodal real-time physiological data obtained from each channel were resampled in the range of 1 Hz to 20 Hz according to a unified timestamp, and short-term missing values ​​were deleted or imputed. Bandpass filtering and artifact removal are applied to the arterial blood pressure waveform and the electrocardiogram signal to retain the effective physiological frequency band. Upper and lower limit checks are performed on the anesthesia depth index and the drug administration record to remove out-of-range or invalid inputs, thus obtaining preprocessed data. Within each sampling period, scalar features of the preprocessed data are extracted; the scalar features include: current mean arterial pressure, heart rate, pulse wave transit time, raw reading of the anesthesia depth index, and instantaneous infusion rate of each vasoactive drug and infusion fluid; The scalar features are concatenated according to a preset field order to form a pre-feature vector F with fixed dimensions. t ; Take the most recent N consecutive pre-eigenvectors {F t₋n₊1 … F t The input is a time-series coding network that has been pre-trained offline and allows for online fine-tuning, and the output is a state vector of a preset dimension to characterize the real-time hemodynamic state of an individual.

5. The intelligent optimization method for personalized intraoperative blood pressure management according to claim 1, characterized in that, The state vector is input into a blood pressure optimization decision model that has been pre-trained offline and adaptively adapted online to obtain regulatory instructions for vasoactive drugs and / or infusions, including: A blood pressure optimization strategy network is loaded and pre-trained offline using multi-center historical intraoperative data; the blood pressure optimization strategy network includes at least an actor-critic structure; wherein the actor subnet of the actor-critic structure outputs the regulatory action, and the critic subnet of the actor-critic structure evaluates the reward value; Receive the state vector and the control action A from the previous cycle. t₋1 and rewards R t₋1 Push it into the observation buffer and maintain a fixed-length time window; The time-series prediction module built into the blood pressure optimization strategy network is invoked to perform rolling predictions of the mean arterial pressure at future ΔT time intervals, resulting in a prediction sequence. ; Based on individual target mean arterial pressure range [MAP] min MAP max Calculate the prediction bias vector D t = -Target interval center value; The current state vector S is processed according to the preset field order. t Prediction bias D t And the previous cycle action A t₋1 Concatenate into an enhanced state vector X t , used as joint input for actor-critic networks; Using the actor subnet for X t Perform forward inference and output an initial control action vector; wherein each component in the initial control action vector corresponds to the vasoactive drug dose increment and / or infusion rate adjustment. The initial control action vector is input into the safety gating module, and then trimmed component by component based on the maximum drug dose, single-step dose change limit, and specific contraindication parameters of the patient to obtain the final executable control command vector.

6. The intelligent optimization method for personalized intraoperative blood pressure management according to claim 5, characterized in that, After receiving regulatory instructions for vasoactive drugs and / or infusions, the following is also included: Execute A t Then, the actual mean arterial pressure response was collected to calculate the reward function. ;in, Indicates at time step Instant reward value; This represents the mean arterial pressure collected at the current time step; Target mean arterial pressure value set for individualization; Indicates the first The magnitude of dose or rate adjustment of the regulatory variable at the current time step; The total number of all controllable variables; It is a drug regulation penalty factor used to limit the drasticness of dose changes; For safety violation indication functions, when the current mean arterial pressure The value is 1 when the value exceeds the set safety range, otherwise it is 0. This is a safety penalty factor used to increase the penalty weight of the system for severe high and low blood pressure events; Based on time difference error The calculation strategy evaluates the residuals; among which, This indicates that the critic subnetwork has a view of the current state vector. Value estimation; This represents an estimate of the updated state; This is a discount factor used to control the impact of future returns on current strategy updates; by The parameters of the commentator subnetwork are updated using gradient descent as the loss term; Using the policy gradient algorithm, to To update the parameters of the actor subnetwork Incremental gradient adjustments are performed to enhance the current strategy's advantage in this state, enabling online adaptive optimization of the blood pressure regulation strategy network.

7. The intelligent optimization method for personalized intraoperative blood pressure management according to claim 1, characterized in that, Adjusting the drug administration or infusion rate by means of the actuator according to the control instructions includes: The vector of the control command is converted into a control command frame conforming to the communication protocol format of the actuator by the data encoding module; the control command frame includes a channel number field for identifying the target channel and a corresponding target infusion rate or dose increment field; The control command frame is sent to the actuator connected to the system via a communication interface; the communication interface is a serial communication interface, a control bus interface, or an Ethernet communication interface; the actuator includes a micro-injection pump, an infusion pump, or an equivalent actuator. Based on the actuator, the target parameters are parsed according to the received control command frame, and the current drug administration rate or infusion rate is updated to complete the rate adjustment based on the control command. The actuator is equipped with a control limiting module, which is used to verify the adjustment value of each drug administration or infusion rate. When the adjustment value exceeds the preset dose change threshold, the adjustment value is automatically cut off to the threshold range. After the rate adjustment is completed, the actuator feeds back the current execution status to the central control unit; the current execution status includes the time of receiving the control command, the target parameters, the actual adjustment result, and the device response status code.

8. The intelligent optimization method for personalized intraoperative blood pressure management according to claim 1, characterized in that, The state vector is updated based on the adjusted multimodal real-time physiological data, and the blood pressure optimization decision model is adaptively corrected to form a closed-loop optimization control, including: Collect multimodal real-time physiological data acquired in the current cycle; The multimodal real-time physiological data is subjected to time alignment, missing data processing and numerical standardization operations, and instantaneous feature values ​​are extracted to construct the pre-feature vector for the current moment; The pre-feature vectors from multiple consecutive cycles are input into a temporal coding network, which outputs an updated state vector to reflect the current dynamic cyclic state of the patient. Based on the state vector and the actual regulatory actions and immediate rewards generated in the current cycle, a policy gradient signal and value function residual are constructed to update the blood pressure optimization decision model. Incremental parameter updates are performed on the strategy subnetwork and value evaluation subnetwork in the blood pressure optimization decision model, respectively, with the goal of enhancing the stability and effectiveness of the regulation strategy under the current state. The updated state vector and the blood pressure optimization decision model are used as inputs for the next cycle of strategy reasoning to achieve the periodic closed-loop execution of the blood pressure optimization control process.

9. An intelligent optimization system for personalized intraoperative blood pressure management, characterized in that, include: The data acquisition unit allows users to continuously collect multimodal real-time physiological data describing the patient's circulatory status during surgery. A vector construction unit is used to construct a state vector reflecting individual hemodynamics based on the multimodal real-time physiological data. The instruction generation unit is used to input the state vector into a blood pressure optimization decision model that has been pre-trained offline and adaptively adapted online to obtain control instructions for vasoactive drugs and / or infusions. A rate regulating unit is used to adjust the drug administration or infusion rate according to the control command via an actuator; The closed-loop optimization unit is used to update the state vector based on the adjusted multimodal real-time physiological data and adaptively correct the blood pressure optimization decision model to form closed-loop optimization control.