A method for controlling an ultrasonic actuated pump lung based on reinforcement learning and fuzzy control

By employing a method of ultrasound-driven pump-lung control based on reinforcement learning and fuzzy control, the pump-lung pulsation cycle and blood flow are dynamically adjusted, solving the problems of difficulty in dynamically adjusting oxygenation efficiency and device complexity in existing technologies, and achieving efficient and intelligent blood circulation support.

CN119846956BActive Publication Date: 2025-12-12SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411908252.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-12-12
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing technologies for improving oxygenation efficiency are difficult to dynamically adjust, and the equipment systems are relatively complex.

Method used

An ultrasound-driven lung pump control method based on reinforcement learning and fuzzy control is adopted. By constructing a fuzzy controller and a deep reinforcement learning model, the pump pulsation cycle and blood flow are dynamically adjusted. Combined with the precise control of the ultrasound motor, the blood flow pattern is optimized.

Benefits of technology

It significantly improves oxygenation efficiency, adapts to dynamic optimization under different working conditions, reduces blood damage, and achieves intelligent control and efficient blood circulation support for the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119846956B_ABST
    Figure CN119846956B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning and fuzzy control's ultrasonic actuator pump lung control method, it is related to extracorporeal membrane oxygenation field, including the following steps: according to expert system knowledge, construct the fuzzy rule between basic vital signs and pulsation frequency, obtain fuzzy controller, adjust average flow;According to the flow field condition of oxygenation device, construct the interactive environment of agent;The state space and action space of agent control strategy are constructed, and the movement of motor is kept stable;Establish and train the blood pump control strategy model based on deep reinforcement learning, obtain blood flow waveform curve and blood pump control strategy;Intelligent control is carried out to motor by extracting trained model, and the control performance of strategy model is evaluated.The application finds optimal flow waveform using deep reinforcement learning, regulates and controls blood flow field, significantly improves oxygenation efficiency without complex oxygenator design conditions, and different agents can be obtained according to different oxygenator structures, with strong portability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of extracorporeal membrane oxygenation (ECMO), and more particularly to an ultrasound-driven lung pump control method based on reinforcement learning and fuzzy control. Background Technology

[0002] Extracorporeal membrane oxygenation (ECMO) is an extracorporeal life support system primarily used to treat various acute circulatory and / or respiratory failures that are unresponsive to conventional life support. ECMO mainly consists of a blood pump and an oxygenator. The blood pump draws venous blood into the oxygenator, where gas exchange occurs before the blood is pumped into the arterial segment of the circulatory system. Integrated blood pump and oxygenator systems, often referred to as integrated pump-oxygenator systems or simply "pump-lungs," reduce the number of system components, simplifying operation and management, enhancing device portability, and minimizing the risk of blood-air contact, infection, and blood damage.

[0003] Natural lungs have a significantly higher gas exchange capacity than extracorporeal membrane oxygenation (ECMO) systems. This is primarily due to the large alveolar-capillary contact area of ​​natural lungs, approximately 100-150 square meters, and the extremely short diffusion distance between gas and blood, not exceeding 1-2 micrometers. At rest, the lungs can process approximately 200-250 ml of O2 and CO2 per minute on average for adults, and this amount can reach 10-20 times during exercise, all using indoor air. In contrast, the hollow fiber oxygenator membrane used in current cardiopulmonary bypass has a smaller area (1-4 square meters), with a much lower surface area to blood volume ratio than natural lungs. Furthermore, its gas diffusion distance is approximately 10-30 micrometers (Wnek G, Bowlin G. Lung, Artificial: Basic Principles and Current Applications / William J. Federspiel, Kristie A. Henchir. In: CRC Press; 2008: 1693-1704.), an order of magnitude greater than that of natural lungs. Therefore, while ensuring that the size of the instrument is appropriate, enhancing the oxygenation efficiency of the artificial lung becomes a crucial task.

[0004] To improve blood oxygenation efficiency, A. Martins Costa et al. investigated the impact of adjusting the arrangement of fiber tubes in the oxygenator on its performance. Researchers changed the number and arrangement of the ventilation fiber tubes, measured the oxygenation effect, and verified that the arrangement of the fiber tubes does indeed have a significant impact on oxygenation efficiency. However, this method of improving oxygenation efficiency is cumbersome in designing and manufacturing the fiber tubes, and adjusting the fiber arrangement to adapt to actual working conditions is difficult (Costa AM, Halfwerk FR, Thiel JN, et al. Effect of hollow fiber configuration and replacement on the gas exchange performance of artificial membrane lungs[J]. Journal of Membrane Science, 2023, 680: 121742.). In addition, Ryan A. Orizondo et al. increased gas transport in hollow fiber membranes (HFMs) by fiber oscillation, achieving a 40% enhancement in oxygenation without significantly increasing hemolysis, but the system structure was relatively complex (Orizondo RA, Gino G, Sultzbach G, et al. Effects of hollow fiber membrane oscillation on an artificial lung[J]. Annals of biomedical engineering, 2018, 46:762-771.).

[0005] Many existing technical solutions involve indirect regulation of the blood flow field, making it difficult to dynamically adjust pre-defined settings. Therefore, exploring technical solutions that directly regulate blood flow by altering the pump's motion pattern holds promise as a promising research area. Among blood pump drive methods, ultrasonic motors demonstrate significant potential. With their rapid and high-precision control capabilities, ultrasonic motors can flexibly adjust blood flow patterns. Furthermore, the compact design and low-noise characteristics of ultrasonic motors make them ideal for clinical applications, reducing patient discomfort. Simultaneously, their rapid response allows the system to flexibly adjust based on real-time monitoring data to dynamically adapt to the patient's physiological needs. This integrated pump-oxygenator system driven by an ultrasonic motor, namely the ultrasonic-actuated pump-lung, can provide more efficient and personalized circulatory support for clinical treatment.

[0006] Therefore, those skilled in the art are dedicated to developing a method for controlling ultrasound-driven lung pumps based on reinforcement learning and fuzzy control. Summary of the Invention

[0007] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is that the current methods for improving oxygenation efficiency are difficult to dynamically adjust, and the device system is relatively complex.

[0008] To achieve the above objectives, this invention provides a method for controlling an ultrasound-driven lung pump based on reinforcement learning and fuzzy control, the method comprising the following steps:

[0009] S101: Based on the knowledge of the expert system, a fuzzy rule between basic vital signs and pulsation frequency is constructed, and a membership function is designed to obtain a fuzzy controller. The fuzzy controller dynamically adjusts the pump-lung pulsation cycle according to the changes in heart rate and mean arterial pressure, thereby adjusting the mean flow rate.

[0010] S103: Based on the flow field conditions of the oxygenation device, construct an intelligent agent interaction environment, the environment including the flow field structure inside the oxygenation device and the gas oxygenation model;

[0011] S105: Construct the state space and action space of the intelligent agent control strategy. The objectives of the intelligent agent control strategy include maintaining stable motor motion, improving oxygenation efficiency, and reducing blood cell damage.

[0012] S107: Establish and train a blood pump control strategy model based on deep reinforcement learning. The strategy model is used to obtain blood flow waveform curves and blood pump control strategies that can significantly improve oxygenation efficiency.

[0013] S109: Extract the trained strategy model, perform intelligent control on the motor, and evaluate the control performance of the strategy model.

[0014] Furthermore, in step S101, the input and output of the fuzzy controller are modeled using Gaussian membership functions, including input variables and output variables. The input variables include heart rate and mean arterial pressure, and the output variables include pulse frequency. Each input and output variable is divided into three levels: low, normal, and high.

[0015] Further, in step S103, one working cycle of the oxygenation device includes a simulated diastolic cycle and a simulated contraction cycle, wherein,

[0016] During the simulated diastolic cycle, the motor drives the slider and the blood pump push plate to move downwards, the pressure inside the blood chamber decreases, and blood enters the blood chamber through the inlet valve and fills it; oxygen enters the hollow fiber membrane through the gas inlet, and the blood absorbs oxygen and releases carbon dioxide through the hollow fiber membrane, and the carbon dioxide is finally discharged from the device through the gas outlet.

[0017] During the simulated contraction cycle, the motor drives the slider and the blood pump push plate to move upward, the pressure inside the blood chamber increases, and the blood is oxygenated through the hollow fiber membrane and discharged from the blood outlet, providing the human body with oxygenated blood to maintain normal circulation.

[0018] Further, step S105 includes the following sub-steps:

[0019] S1051: Construct the state space of the ultrasound-actuated pump lung, wherein the state space is configured to store the calculation results of the oxygenation system simulation model on gas transport efficiency, hemolysis index and platelet activation level;

[0020] S1052: Construct the motion space of the ultrasound-driven pump lung, wherein the motion space automatically adjusts the excitation voltage frequency and duty cycle of the motor to optimize the smooth movement of the motor;

[0021] S1053: Construct the reward function for the ultrasound-driven lung pump, wherein the reward function adopts a hierarchical design, including immediate reward and periodic reward.

[0022] Furthermore, the state space is configured as follows:

[0023] s t ={OTE,SO2,HI,PA,x,T,U,I}

[0024]

[0025] Among them, s t Let be the state space, OTE be the oxygen transfer efficiency, SO2 be the oxygen saturation, HI be the hemolysis index, PA be the platelet activation level, x be the displacement of the blood pump push plate, T be the time required for the blood pump push plate to complete one round-trip periodic motion, i.e., the pulsation period, U be the voltage across the motor, and I be the current through the motor.

[0026] C oxypre C represents the amount of oxygen per unit volume of blood before oxygenation. oxypost C is the amount of oxygen per unit volume of blood after oxygenation. oxy Hb is the amount of oxygen per unit volume of blood, pO2 is the oxygen partial pressure, Hb is the hemoglobin concentration, and pO2 is the oxygen partial pressure.

[0027] Furthermore, the action space is configured as follows:

[0028] a t ={f,d}

[0029] Among them, a t Let f be the motor's excitation voltage frequency, and d be the motor's excitation voltage duty cycle.

[0030] Furthermore, the reward function is configured as follows:

[0031]

[0032] Among them, R total For the reward function, R instant (t) is the instantaneous reward function, reflecting the real-time motion characteristics of the motor, R cycle The periodic reward function is used to evaluate oxygenation efficiency and blood damage at the end of an oxygenation cycle, while controlling the pulsation frequency to meet the output requirements of fuzzy control. γ1 and γ2 are weighting coefficients, t is time, and T is a time period. During training, apart from the motor motion state and pulsation cycle, other state information is calculated and obtained by the constructed computational fluid dynamics simulation model. In practical applications, when using the trained strategy, the non-instantaneous feedback needs to be kept at fixed values, i.e., the blood damage index and oxygen transfer efficiency (OTE) need to be kept at fixed values, and oxygen saturation, heart rate, and mean arterial pressure should be acquired in real time through a monitor as input state information for the strategy.

[0033] Furthermore, the instant reward function is:

[0034] R instant (t)=λSO2-μ(C energy (t)+C instability (t))

[0035] C energy (t)=U(t)·I(t)

[0036]

[0037] Among them, C energy U(t) represents the energy consumption at the current time step, U(t) represents the voltage through the motor, U(t) represents the current through the motor, and C represents the current. instability (t) represents the penalty term for the motion instability of the blood pump push plate, which uses the acceleration of the displacement of the blood pump push plate inside the lungs of the ultrasound-actuated pump to reflect the oscillation, and λ,μ are the weighting coefficients.

[0038] Furthermore, the periodic reward function is:

[0039] R cycle =α·OTE-δ·HI-ε·PA-ζ|T-T0|

[0040] Where α, δ, ε, ζ are weighting coefficients, with values ​​ranging from [0, 1], and T0 is the current required time.

[0041] Further, step S107 includes the following sub-steps:

[0042] S1071: Initialize the policy network and the value network, wherein the policy network generates control signals based on the current state, the control signals are configured as actions, and the value network evaluates the value of the current state-action pair and outputs the Q value of the action;

[0043] S1072: Real-time acquisition of motor status data and corresponding flow waveform data, and import of one cycle of flow waveform data into the oxygenation device simulation system to obtain oxygenation efficiency index;

[0044] S1073: Combine the oxygenation efficiency index with the state data to form a complete empirical tuple, add the empirical tuple to the empirical replay pool, and keep the empirical replay pool updated.

[0045] S1074: Periodically draw small batches of experience samples randomly from the experience replay pool to train the policy network and the value network;

[0046] S1075: Repeat the data collection and training process, and achieve the iterative training process by continuously collecting new experience and optimizing strategies until the strategy model converges.

[0047] Furthermore, when training the policy network and the value network, the policy network is optimized using the policy gradient method, and the value network is optimized using the Bellman error minimization method.

[0048] In a preferred embodiment of the present invention, compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. This invention utilizes deep reinforcement learning to find the optimal flow waveform and improves oxygenation efficiency by directly regulating the blood flow field. In practical applications, it can be optimized for different working conditions, significantly improving oxygenation efficiency without the need for complex oxygenator design. It can also be dynamically optimized according to different working conditions. Furthermore, it incorporates fuzzy control, enabling the device to automatically adjust the average flow rate output of the pump lung according to the patient's different blood supply needs during actual use, so as to better adapt to the patient's actual situation.

[0050] 2. This invention utilizes deep reinforcement learning to optimize the smooth movement of the motor, blood oxygen levels, and blood destructiveness. It combines the SAC (Soft Actor-Critic) algorithm to train the reinforcement learning model, thereby achieving adaptive and intelligent control of the ultrasonic motor.

[0051] 3. This invention can be modeled and trained based on different oxygenator structures to obtain different intelligent agents, which can regulate the actual lung pump and have strong portability.

[0052] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description

[0053] Figure 1 This is a schematic diagram illustrating the implementation steps of the ultrasound-driven lung pump control method based on reinforcement learning and fuzzy control according to an embodiment of the present invention.

[0054] Figure 2 This is a schematic diagram of the deep reinforcement learning system architecture according to an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of the oxygenation device system structure according to an embodiment of the present invention.

[0056] Explanation of each number in the diagram:

[0057] 1-Blood chamber, 2-Hollow fiber membrane, 3-Blood inlet, 4-Valve, 5-Blood outlet, 6-Flow data collection and CFD simulation unit, 7-Ultrasonic flow meter, 8-Hose, 9-Ultrasonic linear motor slider and blood pump push plate, 10-Multi-parameter monitor, 11-Monitor signal conditioning circuit, 12-Control host, 13-MCU controller, 14-Drive circuit, 15-Housing, 16-Ultrasonic linear motor stator, 17-Displacement sensor, 18-Sensor signal conditioning circuit. Detailed Implementation

[0058] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0059] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.

[0060] like Figure 1 , Figure 2As shown, given the limitations of existing technologies in oxygenation efficiency, this invention, guided by the concept of layer-by-layer oxygenation, optimizes the blood flow waveform to achieve a significant improvement in oxygenation efficiency. The core of the layer-by-layer oxygenation concept lies in utilizing the precise micrometer-level displacement control capability of an ultrasonic linear motor to precisely control the amount of blood participating in oxygenation each time and its residence time in the oxygenator, ensuring that blood is promptly discharged after sufficient oxygenation, thereby maintaining a stable blood flow. Simultaneously, this method also ensures the most efficient utilization of gas in the fiber optic tube, improving overall oxygenation performance. The precise management of the amount of blood participating in oxygenation and its residence time by the ultrasonic motor is macroscopically manifested as dynamic regulation of the flow waveform during the oxygenation stage. This regulation makes the blood flow curve in the oxygenator more consistent with the optimal oxygenation requirements, thus achieving fine control of the blood oxygenation process at the microscopic level, highly consistent with the layer-by-layer oxygenation concept.

[0061] This invention designs fuzzy rules between a patient's basic vital signs and pulsation frequency, and designs a membership function to obtain a fuzzy controller. This controller can dynamically adjust the pump-lung pulsation cycle according to changes in the patient's heart rate and mean arterial pressure, thereby adjusting the mean flow rate to better adapt to the dynamic changes in human vital signs.

[0062] This invention proposes an ultrasound-driven lung pump control method based on reinforcement learning and fuzzy control. Based on the idea of ​​layered oxygenation, it designs the state space, action space, and reward function in the SAC algorithm network model, as well as the Actor and Critic network structures. It introduces an entropy regularization coefficient to enhance the exploratory nature of the strategy, helping the model to better balance exploration and utilization during the learning process, thereby maximizing the expected reward function value.

[0063] like Figure 1 As shown in the figure, this embodiment proposes an ultrasound-driven lung pump control method based on reinforcement learning and fuzzy control. The ultrasound-driven lung pump control method based on reinforcement learning and fuzzy control specifically includes the following steps.

[0064] Step 1: Based on the knowledge of the expert system, design fuzzy rules between the patient's basic vital signs and pulse rate, and design membership functions. Both input and output use Gaussian membership functions.

[0065] The fuzzy system consists of the following input variables: heart rate (HR), defined as 50-120 bpm; mean arterial pressure (MAP), defined as 50-120 mmHg; and pulse frequency (F), adjusted to 70-90 bpm. Because Gaussian membership functions possess smoothness, adjustability, and adaptability, making them suitable for changes in vital signs, the input and output of the fuzzy controller in this invention are modeled using Gaussian membership functions. Each input / output variable is divided into three levels: Low, Normal, and High. This greatly simplifies the calculation. Defuzzification uses the centroid method to obtain the final pulse frequency, and the pulse period T0 can then be calculated from the pulse frequency.

[0066] The Gaussian membership function is in the form of:

[0067]

[0068] It can also be expressed as:

[0069] gaussmf(x,[σ,c]).

[0070] The fuzzy universe of discourse for heart rate (HR) is [50, 120] bpm. The fuzzy subsets can be defined as Low, Normal, and High. The heart rate membership function is defined as follows:

[0071]

[0072] The fuzzy universe of discourse for mean arterial pressure (MAP) is 50-120 mmHg. The fuzzy subsets can be defined as low, normal, and high. The membership function for mean arterial pressure is defined as follows:

[0073]

[0074] The fuzzy domain of the output variable, pulse frequency, is 70-90 bpm. The fuzzy subsets can be defined as Low, Normal, and High. The pulse frequency membership function is defined as follows:

[0075]

[0076] Artificial pump-lungs combine oxygenation and pumping functions, requiring precise control of blood flow to coordinate adequate oxygenation and efficient delivery. This fuzzy rule design is based on heart rate and mean arterial pressure (MAP). Changes in heart rate (HR) and MAP reflect the patient's hemodynamic state and blood supply requirements (flow). An example of fuzzy control design is described below: Low heart rate and low blood pressure indicate a low-dynamic state of the circulatory system; reducing the pulse rate can avoid excessive shear stress and blood damage, while also reducing the uncoordinated burden between the device and the patient's cardiovascular system. High heart rate and high blood pressure indicate a high-metabolic or stress state; increasing the frequency can enhance the pump's output capacity, thereby improving hemodynamics in hypertension and supporting the perfusion needs of systemic tissues. Specific fuzzy rules are shown in Table 1.

[0077] Table 1 Fuzzy Rules

[0078] Rule Number HR MAP F R1 Low Low Low R2 Low Normal Low R3 Low High Normal R4 Normal Low Low R5 Normal Normal Normal R6 Normal High High R7 High Low Normal R8 High Normal High R9 High High High

[0079] Step 2: Based on the flow field conditions of the oxygenation device, construct an intelligent agent interaction environment, which includes the flow field structure inside the oxygenation device and the gas oxygenation model.

[0080] like Figure 2 As shown, in this embodiment, one working cycle of the oxygenation device includes a simulated diastolic cycle and a simulated systolic cycle, wherein,

[0081] During the simulated diastolic cycle, the motor drives the slider and the blood pump push plate to move downwards, reducing the pressure inside the blood chamber. Blood enters the blood chamber through the inlet valve and fills it. Oxygen enters the hollow fiber membrane through the gas inlet. Blood absorbs oxygen and releases carbon dioxide through the hollow fiber membrane. The carbon dioxide is finally discharged from the device through the gas outlet.

[0082] During the simulated contraction cycle, the motor drives the slider and blood pump pusher to move upward, increasing the pressure inside the blood chamber. After being oxygenated by the hollow fiber membrane, the blood is discharged from the blood outlet, providing the body with oxygenated blood to maintain normal circulation.

[0083] In order to achieve precise control of micron-level displacement, an ultrasonic linear motor is used in this embodiment.

[0084] Step 3: Construct the state space and action space of the intelligent agent control strategy. The goals of the intelligent agent control strategy include maintaining stable motor motion, improving oxygenation efficiency, and maintaining low damage to blood cells.

[0085] In this step, the state space and action space of the agent's corresponding strategy are determined to maintain stable motor motion, improve oxygenation efficiency, and reduce blood cell damage. The objectives of the state space and reward function are then defined. The reward function employs a hierarchical design, where immediate rewards reflect the real-time motion characteristics of the motor, while periodic rewards are used to evaluate the oxygenation efficiency at the end of an oxygenation cycle. Figure 2 As shown.

[0086] In this embodiment, step 3 further includes the following sub-steps:

[0087] S3.1: Construct the state space of the ultrasound-driven pump lung, and configure the state space to store the calculation results of gas transport efficiency, hemolysis index and platelet activation level of the oxygenation system simulation model.

[0088] In this step, the main program calls the oxygenation system simulation model of the computational fluid dynamics software to calculate the oxygen and carbon dioxide transport efficiency, as well as the hemolysis index and platelet activation level, and imports the output calculation results into the state space.

[0089] In this implementation, the state space takes the following form:

[0090] s t ={OTE,SO2,HI,PA,x,T,U,i}

[0091] Among them, s t Let be the state space, OTE be the oxygen transfer efficiency, SO2 be the oxygen saturation, HI be the hemolysis index, PA be the platelet activation level, x be the displacement of the blood pump push plate, T be the time required for the blood pump push plate to complete one round-trip periodic motion, i.e., the pulsation period, U be the voltage across the motor, and I be the current through the motor.

[0092] Oxygen transfer efficiency and blood damage indices were both solved using computational fluid dynamics simulations of the oxygenation system. Specifically, OTE, HI, and PA were calculated using an oxygenation system simulation model built with ANSYS finite element analysis software. These were updated and recorded as periodic indices after each lung pump completes one pulsation. The remaining indices were imported into the MATLAB state space as real-time measurements.

[0093] Oxygen transport efficiency is characterized by the amount of gas transported per unit blood flow, i.e.:

[0094]

[0095] Where OTE is the oxygen transfer efficiency, and C oxypre C represents the amount of oxygen per unit volume of blood before oxygenation. oxypost C is the amount of oxygen per unit volume of blood after oxygenation. oxyThe oxygen content per unit volume of blood is Hb, hemoglobin concentration is SO2, oxygen saturation is pO2, and oxygen partial pressure is pO2.

[0096] S3.2: Construct the motion space of the ultrasound-driven pump lung. The motion space automatically adjusts the excitation voltage frequency and duty cycle of the motor to optimize the smooth movement of the motor.

[0097] In this embodiment, the action space is:

[0098] a t ={f,d}

[0099] Among them, a t Let f be the motor's excitation voltage frequency, and d be the motor's excitation voltage duty cycle.

[0100] S3.3: Construct the reward function for ultrasound-driven lung pumping. The reward function adopts a hierarchical design, including immediate reward and periodic reward.

[0101] In this embodiment, the reward function is:

[0102]

[0103] Among them, R total For the reward function, R instant (t) is the instantaneous reward function, reflecting the real-time motion characteristics of the motor, R cycle γ1 and γ2 are the periodic reward function used to evaluate the oxygenation efficiency at the end of an oxygenation cycle, t is the time, and T is a time period.

[0104] The instant reward function is:

[0105] R instant (t)=λSO2-μ(C energy (t)+C instability (t))

[0106] C energy (t)=U(t)·I(t)

[0107]

[0108] Among them, R instant (t) is the immediate reward function, C energy (t) represents the energy consumption at the current time step, C. instability (t) is the penalty term for the motion instability of the blood pump push plate, which uses the acceleration of the displacement of the blood pump push plate inside the lung of the ultrasound-actuated pump to reflect the oscillation. U(t) is the voltage through the motor, I(t) is the current through the motor, and λ,μ are weighting coefficients.

[0109] The periodic reward function is:

[0110] R cycle =α·OTE-δ·HI-ε·PA-ζ|T-T0|

[0111] Here, α, δ, ε, and ζ are weighting coefficients, which can be continuously adjusted and optimized through training, and their values ​​range from [0,1]. Simultaneously, after the periodic motion ends, the period T of this motion is calculated based on the displacement sensor results and the controller's internal clock. This is compared with the fuzzy control output, i.e., the currently required time T0. If the deviation exceeds 5%, this strategy is directly abandoned.

[0112] Step 4: Establish and train a blood pump control strategy model based on deep reinforcement learning. The strategy model is used to obtain blood flow waveform curves and blood pump control strategies that can significantly improve oxygenation efficiency.

[0113] In this embodiment, by establishing and training a blood pump control strategy based on a deep reinforcement learning algorithm, a blood flow waveform curve and control strategy that can significantly improve oxygenation efficiency can be obtained.

[0114] This embodiment uses the SAC algorithm for policy training. The SAC algorithm maximizes cumulative reward while also increasing exploration capability by maximizing policy entropy. The algorithm is implemented using the MATLAB platform.

[0115] In this embodiment, step 4 includes a data acquisition phase and a model training phase, specifically including the following sub-steps:

[0116] S4.1: Initialize the policy network and value network. The policy network generates control signals based on the current state. The control signals are configured as actions. The value network evaluates the value of the current state-action pair and outputs the Q value of the action.

[0117] In this embodiment, the blood pump control strategy model includes a strategy network (Actor) and a value network (Critic). The Actor network is used to generate control signals (i.e., actions) based on the current state, and the Critic network is used to evaluate the value of the current state-action pair and output the Q value of the action. During one complete motion cycle of the motor, the motor's state and corresponding oxygenation efficiency indicators are recorded in real time.

[0118] S4.2: Real-time acquisition of motor status data and corresponding flow waveform data, and import of one cycle of flow waveform data into the oxygenation device simulation system to obtain oxygenation efficiency index and blood damage index.

[0119] When collecting real-time operating data of the motor, it is necessary to collect real-time data of a complete motion cycle of the ultrasonic linear motor, and record the state of the ultrasonic linear motor (including speed, voltage, and current) and the corresponding flow waveform data in real time.

[0120] The collected flow waveform data for one cycle is imported into the oxygenation device simulation system to obtain oxygenation efficiency and blood damage indicators.

[0121] During the training phase, MATLAB was used to simulate and generate heart rate and mean arterial pressure under different conditions. These input values ​​were then passed to a fuzzy controller, which output the corresponding pulse frequency to determine the pulse cycle T0. The agent began training under each pulse cycle condition. As the agent interacted with the environment, the replay pool was dynamically updated using acquired empirical data. During training, small batches of empirical samples were randomly drawn from the replay pool to optimize the Critic and Actor networks. The Bellman error minimization method was used to optimize the Critic network, while the policy gradient method was used to optimize the decision performance of the Actor network. Once the agent achieved optimal oxygenation and motor motion stability at the current pulse frequency, the input states of heart rate and mean arterial pressure were adjusted, and the above training process was repeated to further improve the system's robustness and adaptability.

[0122] S4.3: Combine the oxygenation efficiency index with the status data to form a complete empirical tuple, add the empirical tuple to the empirical replay pool, and keep the empirical replay pool updated.

[0123] In this step, the oxygenation efficiency index, blood damage index, and real-time data are combined to form a complete empirical tuple, and newly generated empirical tuples are continuously added to the empirical replay pool to keep the replay pool updated.

[0124] Meanwhile, during the training phase, whenever the agent interacts with the environment and gains experience, the experience is added to the replay pool.

[0125] S4.4: Periodically draw small batches of experience samples randomly from the experience replay pool to train the policy network and value network.

[0126] During the training phase, small batches of empirical samples are periodically and randomly drawn from the replay pool to train the Critic and Actor networks. During training, the Actor network is optimized using the policy gradient method, and the Critic network is optimized using the Bellman error minimization method.

[0127] S4.5: Repeated data collection and training process, through continuous collection of new experience and optimization strategies, to achieve an iterative training process until the policy model converges.

[0128] The process of repeated data collection and training is an iterative process of training, which involves continuously collecting new experience and optimizing strategies until the strategy converges.

[0129] To save resources, an appropriate number of iterations can be set based on the training results to end the training early.

[0130] In this embodiment, when training the policy network and the value network, the policy gradient method is used to optimize the policy network, and the Bellman error minimization method is used to optimize the value network.

[0131] Step 5: Extract the trained strategy model, perform intelligent control on the motor, and evaluate the control performance of the strategy model.

[0132] The trained strategy model is extracted and compared with traditional control methods. Under the premise of ensuring the same average flow rate, the comparison includes system stability and gas transfer efficiency to evaluate the performance of the strategy.

[0133] like Figure 3 The diagram shows the specific structure of the oxygenation device used in this embodiment. The oxygenation device includes a blood chamber 1, a hollow fiber membrane 2, a blood inlet 3, a valve 4, a blood outlet 5, a flow data collection and CFD simulation unit 6, an ultrasonic flow meter 7, a hose 8, an ultrasonic linear motor slider and blood pump push plate 9, a multi-parameter monitor 10, a monitor signal conditioning circuit 11, a control host 12, an MCU controller 13, a drive circuit 14, a housing 15, an ultrasonic linear motor stator 16, a displacement sensor 17, and a sensor signal conditioning circuit 18.

[0134] The specific contents of one working cycle of this oxygenation device include:

[0135] During simulated diastole, the motor drives the slider and blood pump pusher 9 to move downwards, reducing the pressure inside the blood chamber 1. Blood then enters the blood chamber 1 through the valve 4 of the blood inlet 3 and fills the blood chamber 1. Simultaneously, oxygen enters the hollow fiber membrane 2 through the gas inlet. The blood absorbs oxygen through the membrane and releases carbon dioxide, which is ultimately discharged from the device through the gas outlet.

[0136] During simulated contraction, the motor drives the slider and blood pump push plate 9 to move upward, increasing the pressure inside the blood chamber 1. After being oxygenated by the hollow fiber membrane 2, the blood is discharged from the blood outlet 5, providing the human body with oxygenated blood to maintain normal circulation.

[0137] In this embodiment, the computational fluid dynamics gas oxygenation model was established by referencing the method proposed by Kaesler A et al. (Kaesler A, Rosen M, Schmitz-Rode T, Steinseifer U, Arens J. Computational Modeling of Oxygen Transfer in Artificial Lungs. Artif Organs. 2018 Aug; 42(8):786-799. doi:10.1111 / aor.13146. Epub 2018 Jul 24. PMID:30043394.). This method, by simulating blood as a two-phase system and considering the interaction between hemoglobin and oxygen, can more accurately predict the oxygen transfer process in artificial lungs. Blood damage indices are solved using a power-law model of blood damage. The values ​​of some model parameters need to be verified and adjusted based on experimental data of the actual flow field to obtain their accurate values. Finally, using ANSYS finite element analysis software, the gas transfer efficiency is numerically calculated and analyzed to evaluate the gas exchange performance and blood destructiveness in the circulatory system.

[0138] This invention analyzes gas transport efficiency based on computational fluid dynamics and combines it with the proposed algorithm to achieve coupled control of motor motion characteristics and hemodynamic characteristics, thereby obtaining a specific flow waveform to improve oxygenation efficiency. Based on a comprehensive consideration of oxygenation efficiency and motor kinematics, this invention designs the state space, action space, and reward function in the SAC algorithm network model, proposing a control method for an ultrasound-actuated lung pump based on a deep reinforcement learning algorithm. This method can achieve precise control of the blood volume flowing into the oxygenation device and the holding time, adjust the blood side boundary layer thickness, and improve oxygenation efficiency.

[0139] Compared with existing technologies, the ultrasound-driven lung pump control method based on reinforcement learning and fuzzy control provided in this invention has the following characteristics:

[0140] 1. Addressing the challenges of dynamically adjusting current methods for improving oxygenation efficiency and the complexity of the related systems, this invention utilizes deep reinforcement learning to find the optimal flow waveform. By directly regulating the blood flow field, it enhances oxygenation efficiency and can be optimized for different operating conditions in practical applications. This invention precisely controls the amount of blood flowing into the oxygenator and its duration of stay by adjusting the blood-side boundary layer thickness to achieve a layer-by-layer, batch-based oxygenation process, thereby improving oxygenation efficiency. It significantly improves oxygenation efficiency without requiring complex oxygenator designs and can be dynamically optimized for different operating conditions.

[0141] 2. The working principle of ultrasonic motors involves complex vibration modes and nonlinear dynamics, making motion control complex and difficult to achieve precise and stable control using traditional methods. This invention utilizes deep reinforcement learning, with the smooth motion of the motor, blood oxygen saturation, and blood destructiveness as optimization objectives. Combined with the SAC algorithm, a reinforcement learning model is trained to achieve intelligent control of the ultrasonic motor. By training an agent, it can automatically adjust the excitation voltage frequency and duty cycle of the motor based on the real-time state of the motor and relevant blood indicators to optimize smooth motion, improve blood oxygen saturation, and reduce blood destructiveness. The agent uses an Actor network to generate control actions, and a Critic network evaluates the value of these actions. Learning is then performed based on a reward function (rewarding smooth motion and improved blood oxygenation, penalizing increased blood destructiveness) to ultimately achieve the control objectives, realizing adaptive and intelligent control of the ultrasonic motor.

[0142] 3. In practical applications, patients' physiological states are constantly changing, and their blood supply needs vary under different conditions. This invention monitors two key physiological parameters—heart rate (HR) and mean arterial pressure (MAP)—and utilizes data from an expert system's knowledge base to design fuzzy rules to dynamically adjust the average flow rate of the pump-lung machine, adapting to the patient's ever-changing blood supply needs. By modeling the input and output variables using Gaussian membership functions and employing the centroid method for fuzzification, this system can flexibly adjust the pump-lung flow rate. Furthermore, by combining deep reinforcement learning technology, the system can not only adjust the flow rate in real time according to the patient's state but also ensure high oxygenation efficiency and low blood damage during the adjustment process, thereby improving the adaptability and effectiveness of treatment.

[0143] Therefore, the ultrasonic-actuated pump-lung control method based on reinforcement learning and fuzzy control provided in this invention can significantly improve the oxygenation efficiency of artificial pump-lungs while reducing blood damage, which is of great significance for research on improving the oxygenation efficiency of oxygenators. Through deep reinforcement learning algorithms, precise control can be achieved without modeling the complex physical model of the ultrasonic motor, driving the pump to optimize the flow pattern of blood in the oxygenator, thereby improving the exchange efficiency between blood and gas at the interface and enhancing the overall efficiency of the oxygenation process. When facing different flow field conditions or different arrangements of hollow fibers, the method of this invention can be used to appropriately optimize and adjust parameters to improve the performance of the oxygenator. Furthermore, considering the difficulty of real-time feedback of blood gas indicators during training, this invention can also adopt an offline learning approach. In addition, since the experimental materials and environmental conditions required for collecting blood gas indicators through actual experiments are quite demanding, this invention chooses to use computational fluid dynamics simulation to obtain blood gas indicators. By importing the flow waveform signal of a complete cycle generated by the ultrasonic-actuated pump-lung into the simulation model, the corresponding blood gas indicators can be obtained, which greatly facilitates practical operation.

[0144] Meanwhile, the industrial application prospects of this invention are broad, with significant potential for improving the performance of artificial heart-lung machines. With the aging global population and advancements in medical technology, the demand for efficient and safe extracorporeal life support systems is increasing. This invention can optimize blood flow within the oxygenator using reinforcement learning algorithms within a limited device space, without physically modifying the internal structure or hollow fiber arrangement of the oxygenator. Furthermore, this invention can be modeled and trained based on different oxygenator structures to obtain different intelligent agents that can regulate actual lung pumping, exhibiting strong portability.

[0145] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A method for controlling an ultrasonically actuated pump lung based on reinforcement learning and fuzzy control, characterized in that, The method comprises the following steps: S101: According to the expert system knowledge, the fuzzy rules between the basic vital signs and the pulsation frequency are constructed, and the membership function is designed to obtain a fuzzy controller, which dynamically adjusts the pump lung pulsation period according to the heart rate and mean arterial pressure, and then adjusts the average flow; S103: According to the flow field conditions of the oxygenation device, an interactive environment of the agent is constructed, which includes the flow field structure inside the oxygenation device and the gas oxygenation model; S105: The state space and action space of the agent control strategy are constructed, and the target of the agent control strategy includes keeping the motor motion stable, improving the oxygenation efficiency, and reducing the blood cell damage; S107: A blood pump control strategy model based on deep reinforcement learning is established and trained, and the strategy model is used to obtain a blood flow waveform curve and a blood pump control strategy that can significantly improve the oxygenation efficiency; S109: The trained strategy model is extracted, the motor is intelligently controlled, and the control performance of the strategy model is evaluated Wherein, The output result of the fuzzy controller is the beat period After the periodic motion is finished, the period of this motion is calculated according to the result of the displacement sensor combined with the internal clock of the controller The output result of the fuzzy controller is the time needed at present In contrast, if the deviation exceeds 5%, the strategy is directly abandoned; The step S105 comprises the following sub-steps: S1051: The state space of the ultrasonic actuated pump lung is constructed, and the state space is configured to store the calculation results of the oxygenation system simulation model on the gas transmission efficiency, hemolysis index and platelet activation level; S1052: The action space of the ultrasonic actuated pump lung is constructed, which automatically adjusts the excitation voltage frequency and duty cycle of the motor to optimize the smooth motion of the motor; S1053: The reward function of the ultrasonic actuated pump lung is constructed, which adopts a hierarchical design including immediate reward and periodic reward; The reward function is configured as: wherein, is a reward function, is an instant reward function, reflecting real-time motion characteristics of the motor, is a periodic reward function, used to evaluate the oxygenation efficiency after the end of an oxygenation period, , is a weight coefficient, is time, is a time period; The immediate reward function is: wherein, is the energy consumption of the current time step, is the voltage across the motor, is the current through the motor, is the blood pump pusher plate motion instability penalty term, using the acceleration of the ultrasound-actuated pump pusher plate displacement inside the blood pump to reflect the oscillation, is the weight coefficient; The periodic reward function is: wherein, is a weight coefficient, with a value range of [0, 1], is the current required time.

2. The method of claim 1, wherein, In the step S103, one working period of the oxygenation device includes a simulated diastolic period and a simulated systolic period, wherein, In the simulated diastolic period, the motor drives the slider and the blood pump push plate to move downward, the pressure in the blood cavity decreases, the blood enters the blood cavity through the inlet valve and fills; oxygen enters the hollow fiber membrane through the gas inlet, the blood absorbs oxygen through the hollow fiber membrane and discharges carbon dioxide, and carbon dioxide is finally discharged from the device through the gas outlet; In the simulated systolic period, the motor drives the slider and the blood pump push plate to move upward, the pressure in the blood cavity increases, the blood is discharged from the blood outlet after oxygenation through the hollow fiber membrane to provide oxygen-containing blood for the human body to maintain normal circulation.

3. The method of claim 2, wherein, The state space is configured as: wherein, is the state space, is the oxygen transfer efficiency, is the oxygen saturation, is the hemolysis index, is the platelet activation level, is the displacement of the blood pump pusher plate, is the time required for the blood pump pusher plate to complete one round trip periodic motion, i.e., the beat period, is the voltage across the motor, is the current through the motor; C02 is the amount of carbon dioxide in the blood, C02 is the amount of carbon dioxide in the blood, C02 is the amount of carbon dioxide in the blood, C02 is the amount of carbon dioxide in the blood, C02 is the amount of carbon dioxide in the blood, 4. The method of claim 3, wherein, The action space is configured as: wherein, is the action space, is the excitation voltage frequency of the motor, is the excitation voltage duty cycle of the motor.

5. The method of claim 4, wherein, The step S107 comprises the following sub-steps: S1071: Initialize the policy network and the value network, the policy network generates a control signal according to the current state, the control signal is configured as an action, and the value network evaluates the value of the current state-action pair and outputs the Q value of the action; S1072: Real-time acquisition of state data of the motor and flow waveform data corresponding to the state data, and import of one period of the flow waveform data into the oxygenation device simulation system to obtain the oxygenation efficiency index; S1073: The oxygenation efficiency indicator is combined with the state data to form a complete experience tuple, and the experience tuple is added to the experience replay pool, and the experience replay pool is kept updated; S1074: Small batches of experience samples are randomly extracted from the experience replay pool at regular intervals for training the policy network and the value network; S1075: Repeat the data collection and training process, and through the continuous collection of new experiences and the optimization of the policy, realize the iterative process of training until the policy model converges.

6. The method of claim 5, wherein, When training the policy network and the value network, the policy gradient method is used to optimize the policy network, and the Bellman error minimization method is used to optimize the value network.

Citation Information

Patent Citations

  • New cardiac pump model-based variable rotation speed control method

    CN107045281A

  • ECMO centrifugal blood pump flow pulsatile control system based on RBF neural network

    CN116650827A