Pulsating type blood pump heat management control system based on rotating speed modulation
By combining the motor thermal power physical model and the thermal power correction prediction model of deep learning, the policy network parameters are updated in real time, and the complex nonlinear problem that traditional models are difficult to capture the heat loss of pulsating blood pumps is solved, and the stable operation of the blood pump and thermal management optimization is achieved, which improves the long-term safety and reliability of the blood pump.
Patent Information
- Application Number
- CN202510947590.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Traditional motor thermal loss models are difficult to capture the complex nonlinear and timing characteristics of pulsating blood pumps, resulting in deviations in the prediction of instantaneous thermal power and full-cycle cumulative thermal energy. The existing methods can only be planned in the short term, which may lead to excessive accumulation of thermal energy and affect the long-term and stable operation of the blood pump.
The thermal management control system based on speed modulation is adopted, and the motor thermal power physical model and thermal power correction prediction model are combined with deep learning to update the policy network parameters in real time to realize the accurate calculation of the motor thermal power and the optimization and adjustment of the speed, ensuring that the blood pump meets the temperature rise minimization and physiological requirements during long-term operation.
It minimizes the periodic accumulation of thermal energy under multiple constraints, improves the long-term operation safety and reliability of the pulsating blood pump, and prevents adverse biological effects caused by local overheating.
Smart Images

Figure CN120459519A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of implantable medical device control, and in particular relates to a pulsating blood pump thermal management control system based on speed modulation. Background Art
[0002] For patients with advanced heart failure, pulsatile blood pumps are a vital life-saving tool. Their operation requires not only sufficient blood flow but also the simulation of the pulsating characteristics of a heartbeat. Pulsatile blood pumps typically rely on brushless motors to deliver blood. Their speed regulation strategy often employs a rapid acceleration followed by rapid deceleration to simulate the heartbeat process. This causes a large amount of mechanical energy to be rapidly converted into heat energy in a short period of time, leading to local temperature rise and heat loss in the equipment. Motor overheating can cause protein denaturation in the blood, damage to blood cell membranes, and even hemolysis. The rapid temperature rise can also interfere with the local vascular temperature control mechanism, increasing the risk of tissue thermal damage, inflammation, and local tissue necrosis.
[0003] Traditional motor heat loss models are mainly calculated based on factors such as copper loss, iron loss, and mechanical loss. However, under the operating conditions of pulsating blood pumps, these models are difficult to accurately capture the actual heat loss due to complex nonlinear effects and fluctuations in environmental parameters. In addition, the pulsating speed regulation process exhibits significant time-varying characteristics, and the generation and dissipation of heat loss show a dynamic evolution trend. Traditional static models are unable to predict instantaneous heat loss (or thermal power) and the accumulated heat energy over the entire cycle in real time. In addition, existing control strategies mostly use fixed parameters or simple feedback control, which fails to fully consider the time-varying problems caused by individual differences and dynamic physiological states of patients. Existing methods (such as particle swarm optimization algorithms) can only perform short-term planning and are prone to falling into local optimality, which may lead to excessive heat energy accumulation under specific operating conditions, seriously affecting the long-term stable operation of the blood pump. Summary of the Invention
[0004] The purpose of the present invention is to provide a pulsating blood pump thermal management control system based on speed modulation to solve at least one of the following problems: the traditional motor heat loss model is difficult to capture the complex nonlinear and timing characteristics generated by pulsating speed regulation, resulting in deviations in the prediction of instantaneous thermal power and full-cycle accumulated thermal energy; and the existing method can only perform short-term planning, which may lead to excessive thermal energy accumulation under specific working conditions, seriously affecting the long-term stable operation of the blood pump.
[0005] The present invention solves the above technical problems through the following technical solutions: a pulsating blood pump thermal management control system based on speed modulation, wherein the control system is configured to perform the following steps:
[0006] Step 1: Generate the initial state vector based on the key parameters of the motor, blood pump output flow, blood temperature and aortic pressure difference at the historical moment;
[0007] Step 2: Using the strategy network to output a speed adjustment vector based on the initial state vector;
[0008] Step 3: Adjust the vector according to the speed to control the motor speed;
[0009] Step 4: Obtain the current key motor parameters, blood pump output flow, blood temperature, and aortic pressure difference;
[0010] Step 5: Generate the current state vector based on the current motor key parameters, blood pump output flow, blood temperature, and aortic pressure difference, and calculate the immediate reward;
[0011] Step 6: Based on the current state vector, loop through steps 2 to 5 for n-1 times to obtain the speed adjustment vector, state vector, and instant reward at n moments.
[0012] Step 7: Update the value network based on the state vector and immediate reward at n moments;
[0013] Step 8: Update the policy network based on the predicted reward error of the value network;
[0014] Step 9: Repeat steps 2 to 8 based on the new starting state vector until the entire cycle is traversed;
[0015] Step 10: Repeat steps 1 to 9 multiple times until the termination condition is met to obtain the trained policy network.
[0016] Furthermore, the key parameters of the motor include rotational speed and accumulated thermal energy, the accumulated thermal energy is calculated based on the motor thermal power, and the motor thermal power is equal to the sum of the basic thermal power and the thermal power correction value.
[0017] Furthermore, the basic thermal power is calculated using a motor thermal power physical model, and the expression of the motor thermal power physical model is:
[0018] ;
[0019] in, Indicates the basic thermal power of the motor; represents the electromagnetic torque, represents the torque constant, Indicates the stator winding resistance; represents the core hysteresis loss coefficient, Indicates the operating frequency, represents the magnetic flux density, and m represents the exponent of hysteresis loss in the core loss; represents the material parameter describing the eddy current loss of the core; Indicates the motor speed, represents the empirical coefficient related to the first-order term of the speed, represents the empirical coefficient related to the quadratic term of speed, and t represents time.
[0020] Furthermore, the thermal power correction value is obtained by predicting the motor operation index using a thermal power correction prediction model; the acquisition process of the thermal power correction prediction model includes:
[0021] Constructing a sample data set; wherein each sample in the sample data set includes an input quantity and an output quantity, the input quantity is a motor operation index, the output quantity is a thermal power correction value, and the thermal power correction value is equal to the difference between the actual thermal power measurement value and the basic thermal power;
[0022] Build an LSTM-TCN neural network model;
[0023] The sample data set is used to train the LSTM-TCN neural network model to obtain the thermal power correction prediction model.
[0024] Furthermore, the output of the LSTM-TCN neural network model is:
[0025] ;
[0026] ;
[0027] ;
[0028] in, represents the predicted thermal power correction value; and Represent the weight and bias of the fully connected layer respectively; represents the output of TCN; Indicates the position of the LSTM output sequence, , H represents the hidden layer dimension of LSTM; Represents the convolution kernel weight of TCN; represents the activation function; Indicates the convolution kernel size of TCN; represents the bias of TCN; Indicates At this moment, the hidden state vector output by LSTM is represents the expansion coefficient; represents the LSTM input sequence, and k represents the number of elements in the LSTM input sequence.
[0029] Furthermore, the calculation formula for the instant reward is: ;
[0030] in, Indicates the immediate reward at the current moment; Indicates the difference in motor thermal power between the current moment and the previous moment. , Indicates the thermal power of the motor at the current moment, Indicates the thermal power of the motor at the last moment; 、 and Respectively represent the penalty factors of the corresponding items; Indicates the blood pump output flow at the current moment, and Respectively represent the minimum and maximum output flow of the blood pump; Indicates the current blood temperature, Indicates the upper limit of safe temperature; Indicates the aortic pressure difference at the current moment, and Represent the minimum and maximum values of aortic pressure gradient, respectively; Indicates a conditional judgment symbol. When the condition is met, the value is 1; when the condition is not met, the value is 0.
[0031] Furthermore, the loss function of the value network is:
[0032] ;
[0033] ;
[0034] in, represents the loss function of the value network, represents the parameters of the value network; N represents the number of moments in the entire cycle; Represents the state vector of the value network at time t Output the estimated reward from time t to time t+N; represents the sum of the discounted cumulative reward calculated based on the immediate rewards at n moments and the estimated rewards at the next Nn moments; represents the discount factor; represents the immediate reward at time t+k; represents the estimated reward from time t+n to time t+N, determined by the value network.
[0035] Furthermore, the loss function of the policy network is:
[0036] ;
[0037] ;
[0038] in, represents the loss function of the policy network, Represents the parameters of the policy network; represents the predicted reward error of the value network; Represents the state vector Lower output speed adjustment vector probability; represents the speed adjustment vector; represents the sum of the discounted cumulative reward calculated based on the immediate rewards at n moments and the estimated rewards at the next Nn moments; Represents the state vector of the value network at time t Output the estimated reward from time t to time t+N.
[0039] Furthermore, the control system is further configured to perform the following steps:
[0040] Obtain key motor parameters, blood pump output flow, blood temperature, and aortic pressure difference at the start time;
[0041] Generate a state vector at the starting moment according to the key parameters of the motor, the blood pump output flow, the blood temperature and the aortic pressure difference at the starting moment;
[0042] Use the trained strategy network to output the speed adjustment vector based on the state vector at the starting time;
[0043] The motor speed is controlled by adjusting the vector according to the speed.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] The present invention minimizes the cycle accumulated heat energy based on the state vector and the designed instant reward online update strategy network parameters, while ensuring that the blood pump output flow, blood temperature and aortic pressure difference meet the predetermined requirements. That is, the temperature rise is minimized under the multiple constraints of blood pump output flow, blood temperature and aortic pressure difference, ensuring that the pulsating blood pump can provide stable blood flow support in long-term operation and effectively prevent adverse biological effects caused by local overheating, greatly improving the safety and reliability of the pulsating blood pump in long-term operation.
[0046] The present invention calculates the motor thermal power through a motor thermal power physical model and a thermal power correction prediction model. It not only utilizes the prior information provided by the physical model, but also captures complex nonlinear and timing characteristics through deep learning, without the need for a large amount of training data. It not only improves the prediction accuracy of the motor thermal power, but also solves the problem that traditional motor heat loss models are difficult to capture complex nonlinear and timing characteristics, resulting in deviations in the prediction of instantaneous thermal power and full-cycle accumulated thermal energy. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only one embodiment of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 This is a flow chart of modeling for the speed control of a pulsating blood pump motor in an embodiment of the present invention;
[0049] Figure 2 This is a flow chart of the speed control of the pulsating blood pump motor in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following is a clear and complete description of the technical solutions of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0051] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0052] The pulsating blood pump thermal management control system based on speed modulation provided by the embodiment of the present invention is configured to execute the pulsating blood pump motor speed control modeling step and the pulsating blood pump motor speed control step. Figure 1 As shown in FIG, the modeling steps for the pulsating blood pump motor speed control include:
[0053] Step 1: Generate an initial state vector based on the key parameters of the motor, blood pump output flow, blood temperature, and aortic pressure difference at the historical moment.
[0054] In order for the speed control to fully reflect the current operating state of the system, a state vector needs to be constructed. To minimize the thermal power of the blood pump motor, ensure temperature control safety, and ensure that the blood pump flow and pulsation meet physiological requirements, the state vector of the present invention is generated based on the key motor parameters, blood pump output flow, blood temperature, and aortic pressure difference. The key motor parameters include speed and accumulated heat energy. Therefore, the state vector can be expressed as:
[0055] (1)
[0056] in, represents the state vector at time t; Indicates the motor speed, which is collected by the speed sensor; represents electromagnetic torque; Indicates the accumulated heat energy of the motor; represents the blood pump output flow at time t (unit: L / min), which is fed back by the cardiovascular dynamics model; represents the blood temperature at time t (unit: °C), which is fed back by the cardiovascular dynamics model; It represents the aortic pressure difference at time t (unit: mmHg), which is fed back by the cardiovascular dynamics model. The accumulated heat energy is calculated based on the motor thermal power. The specific calculation formula is:
[0057] (2)
[0058] in, represents the cumulative heat energy from time zero to time t, represents the thermal power of the motor at time i, Indicates the speed adjustment interval.
[0059] Traditional motor heat loss models primarily calculate factors such as copper loss, iron loss, and mechanical loss. However, due to the complex nonlinear effects and fluctuating environmental parameters of pulsating blood pumps, these models struggle to accurately capture the actual heat loss. Furthermore, the pulsating speed regulation process exhibits significant time-varying characteristics, with the generation and dissipation of heat loss exhibiting a dynamic evolutionary trend. Traditional motor heat loss models are static and cannot predict instantaneous heat loss (or thermal power) or the accumulated heat energy over the entire cycle in real time.
[0060] In order to solve the above technical problems, the present invention realizes the motor thermal power through the motor thermal power physical model and the thermal power correction prediction model. Calculation:
[0061] (3)
[0062] in, Indicates the basic thermal power of the motor, which is calculated by the motor thermal power physical model; Represents the thermal power correction value, which is predicted by the thermal power correction prediction model.
[0063] In a specific embodiment of the present invention, the expression of the motor thermal power physical model is:
[0064] (4)
[0065] in, represents the electromagnetic torque, , represents the blood load determined by the cardiovascular dynamics model, represents the acceleration determined by the cardiovascular dynamics model, represents friction and fluid damping; Represents the torque constant, which is used to convert current into torque. Its specific value is 0.01Nm / A; Indicates the stator winding resistance; It represents the core hysteresis loss coefficient, and its specific value is 0.002; Represents the operating frequency, which is proportional to the speed, that is, , Represents the proportionality factor, which is related to design factors such as the number of motor pole pairs; It represents the magnetic flux density, which represents the magnetic flux intensity in the motor and depends on the motor design, such as is approximately 0.35T; m represents the index of hysteresis loss in the core loss, and its value range is 1.5~2.5. In this embodiment, the value of m is 2; represents the material parameter describing the eddy current loss of the core. In this embodiment, The value of is 1×10 -4 ; Indicates the motor speed; represents the empirical coefficient related to the first-order term of the speed. In this embodiment, The value of is 1×10 -3 W / (rpm); represents the empirical coefficient related to the quadratic term of the speed. In this embodiment, The value of is 5×10 -7 W / (rpm²).
[0066] The thermal power correction prediction model is used to correct the basic thermal power calculated by the motor thermal power physical model, improving the accuracy of motor thermal power prediction. The thermal power correction prediction model can be applied in multiple key scenarios: In real-time monitoring and early warning, it is deployed in the blood pump controller to output the corrected thermal power prediction value in real time. If it exceeds the safety threshold, a warning signal is issued promptly. In device design optimization, through joint training with CFD simulation data, it provides feedback on the local thermal risks of the motor structure and heat dissipation design, providing an optimization basis for material selection and structural layout. In clinical safety assessments, the thermal power correction prediction model can quantify the thermal risks of the motor during long-term operation, improve the accuracy of thermal safety assessments, and provide reliable support for implant risk assessments.
[0067] The thermal power correction prediction model is a data-driven deep learning model. In a specific embodiment of the present invention, the acquisition process of the thermal power correction prediction model includes:
[0068] Step 1.1: Build a sample dataset.
[0069] Each sample in the sample data set includes an input quantity and an output quantity. The input quantity is the motor operation index, and the output quantity is the thermal power correction value. The thermal power correction value is equal to the difference between the actual measured thermal power value and the basic thermal power.
[0070] Actual thermal power measurements under different motor operating parameters can be obtained through experiments or simulations. Experiments include animal experiments and in vitro experiments. Animal experiments involve implanting a blood pump in an animal after obtaining ethical approval to obtain actual thermal power measurements under different motor operating parameters. In vitro experiments involve constructing a simulated circulatory system, operating the blood pump under different motor operating parameters, and obtaining actual thermal power measurements. Simulations use CFD and finite element thermal analysis to generate thermal power data.
[0071] In a specific embodiment of the present invention, the motor operation indicators include speed, electromagnetic torque, magnetic flux density and operating frequency, so the input quantity of each sample can be expressed as:
[0072] (5)
[0073] Arrange the motor operating indicators at time t and the k moments before it to generate an input sequence, which can be expressed as:
[0074] (6)
[0075] in, Represents the input sequence. In order to reflect the number of elements and time periods in the input sequence, the input sequence can also be represented as .by As the input of the LSTM-TCN neural network model (i.e., long short-term memory network-temporal convolutional neural network), this method enables the LSTM-TCN neural network model to learn the impact of historical information on the current state.
[0076] Step 1.2: Build the LSTM-TCN neural network model.
[0077] The thermal power correction prediction model adopts a two-level network structure, namely the long short-term memory network (LSTM) and the temporal convolutional neural network (TCN). The LSTM is used to learn the long-term dependency information in the input sequence, and the TCN uses one-dimensional convolution to extract local (short-term) temporal features, thereby further strengthening the features of the LSTM output. Finally, the features extracted by the TCN are mapped into a single scalar output, namely the thermal power correction value, through a fully connected layer.
[0078] Utilizing the LSTM's memory cells and gating mechanism, we can capture long-term trends in indicators like speed within the input sequence. LSTM excels at capturing long-term dependencies and avoiding the vanishing gradient problem, making it suitable for describing data dependencies over long time spans. After the input sequence is fed into the LSTM, the output is a hidden state vector:
[0079] (7)
[0080] in, represents the hidden state vector output by LSTM, that is, the output sequence of LSTM; LSTM stands for Long Short-Term Memory Network; H represents the hidden layer dimension of LSTM. In this embodiment, the value of H is 100, which determines the size of the hidden state; Represents the field of real numbers.
[0081] TCN uses dilated convolution to amplify the receptive field, effectively extracting local features and capturing short-term transient fluctuations in the input. TCN complements LSTM's lack of local information capture, making the model more sensitive to short-term dynamic changes. The TCN's one-dimensional convolution operation can be expressed as a general formula:
[0082] (8)
[0083] in, represents the output of TCN; Indicates the position of the LSTM output sequence, ; Represents the convolution kernel weight of TCN, which is used for local feature extraction; represents the activation function. In this embodiment, the ReLU function is used to increase nonlinearity. represents the convolution kernel size of TCN, that is, the number of continuous signal points considered in each convolution operation. In this embodiment, The value of is 3; Represents the bias of TCN, the translation parameter used for compensation; Indicates At time t, the hidden state vector output by LSTM; Represents the expansion coefficient, which is used to control the sampling step of the convolution kernel, thereby expanding the receptive field. In this embodiment, The value of is 2. In this embodiment, the channel data of the output feature map of the TCN convolutional layer is set to 64.
[0084] The local features extracted by TCN After global aggregation, it is mapped into a single scalar output through the fully connected layer, which is used to correct the basic thermal power calculated by the motor thermal power physical model. The specific formula is:
[0085] (9)
[0086] in, represents the predicted thermal power correction value; and denote the weights and biases of the fully connected layer respectively.
[0087] Step 1.3: Use the sample data set to train the LSTM-TCN neural network model to obtain the thermal power correction prediction model.
[0088] The thermal power correction prediction model can be used to predict the thermal power correction value, and the motor thermal power at any moment can be calculated according to formula (3). The calculation of motor thermal power using the motor thermal power physical model and the thermal power correction prediction model not only utilizes the prior information provided by the physical model, but also captures complex nonlinear and temporal characteristics through deep learning, thereby improving the prediction accuracy of motor thermal power.
[0089] Step 2: Use the policy network (Actor) to output the speed adjustment vector based on the starting state vector.
[0090] In a blood pump control system, in order to ensure that the temperature rise is minimized while meeting physiological requirements such as blood flow, temperature control, and pulsation, it is necessary to accurately plan the speed according to the working cycle of the blood pump. This paper uses a parameterization method to discretize the working cycle, divides the entire working cycle into N discrete moments, and uses discrete parameters to form a speed parameter vector :
[0091] (10)
[0092] in, Represents the motor speed at the i-th discrete moment, and the speed range is 3000-9000rpm. In order to achieve smooth speed control, a continuous speed function can be constructed by linear or spline interpolation method: This parametric design can fully reflect the optimization requirements throughout the entire working cycle, ensuring smooth speed changes while providing a set of adjustable parameters for subsequent online optimization.
[0093] To achieve intelligent optimization of motor speed regulation during blood pump operation, the present invention proposes motor speed control based on reinforcement learning. This control method, based on state perception and reward feedback, automatically learns the optimal long-term speed scheduling strategy to ensure that the accumulated heat energy within the cycle is minimized while meeting physiological requirements such as blood pump output flow (i.e., blood flow), temperature control, and pulse pressure difference.
[0094] The reinforcement learning method of this invention uses a policy network-value network algorithm. The policy network (Actor) is responsible for outputting actions based on the state vector, namely, adjusting the discrete speed parameter vector (i.e., the speed adjustment vector). The value network (Critic) is used to evaluate the expected cumulative reward of the current state of the system in the future.
[0095] Assume that the entire working cycle is divided into N discrete moments, and the speed parameter vector is , the speed adjustment vector output by the strategy network at one time is:
[0096] (11)
[0097] in, Represents the adjustment amount for the i-th speed, usually limited to a certain range (for example, ±500 rpm).
[0098] The policy network of this embodiment adopts a feedforward multi-layer neural network structure, and its design parameters include:
[0099] State vector dimension :Right now The dimension of this embodiment =6;
[0100] Hidden layer dimensions : The number of nodes in the hidden layer of the strategy network, in this embodiment =50;
[0101] Number of hidden layers : The number of hidden layers of the policy network, in this embodiment =5;
[0102] Output dimension : is equal to the number of discrete speed planning moments, that is, the dimension of the speed adjustment vector. Equal to 20.
[0103] The optimization goal of the policy network is to maximize the probability of the advantageous action (that is, the action that minimizes the thermal power), so its loss function is Defined as:
[0104] (12)
[0105] (13)
[0106] in, represents the loss function of the policy network, Represents the parameters of the policy network, i.e., the weights and biases of the hidden layers; Represents the predicted reward error of the value network, that is, the predicted reward error at n moments, indicating the speed adjustment vector Advantages; Represents the state vector Lower output speed adjustment vector probability; represents the speed adjustment vector; represents the sum of the discounted cumulative reward calculated based on the immediate rewards at n moments and the estimated rewards at the next Nn moments; Represents the state vector of the value network at time t The output is the estimated reward from time t to time t+N, that is, the estimated reward of the value network from time t to the end of the cycle.
[0107] If the current action Make If it is a negative number, it means that the estimated reward of the current action is higher than the actual reward, then the vector is adjusted by increasing the speed. The probability of being selected is reduced. The initial state vector is input to the strategy network on the controller, and the strategy network outputs the speed adjustment vector .
[0108] Step 3: Adjust the vector according to the speed Control the motor speed.
[0109] Step 4: Obtain the current key motor parameters, blood pump output flow, blood temperature, and aortic pressure difference.
[0110] The key parameters of the motor include speed and accumulated heat energy. The controller adjusts the speed vector according to the output of the strategy network. Control the speed of the blood pump motor and initiate an interactive request to the thermal power prediction model (including the motor thermal power physical model and the thermal power correction prediction model) and the cardiovascular dynamics model. The thermal power prediction model feeds back the current motor thermal power to the controller, and then calculates the current cumulative thermal energy according to formula (2) The speed sensor feeds back the current motor speed to the controller The cardiovascular dynamics model feeds back the current blood pump output flow to the controller , blood temperature and aortic pressure gradient , the above parameters can also be directly measured through animal experiments to replace cardiovascular dynamics model simulation and realize feedback to the controller.
[0111] Step 5: Generate the current state vector (as shown in Formula (1)) based on the current motor key parameters, blood pump output flow, blood temperature, and aortic pressure difference, and calculate the immediate reward.
[0112] To guide the policy network to learn a speed scheduling strategy that meets both clinical safety and energy efficiency, this paper designs a multi-objective fusion instantaneous reward function. The instantaneous reward function scores the current action at each moment (i.e., time step). This serves as the instantaneous reward in reinforcement learning and is used to construct the discounted cumulative return or discounted cumulative reward within the cycle. The expression of the instantaneous reward function is: (14)
[0113] in, Indicates the immediate reward at the current moment; Represents the difference in motor thermal power between the current moment and the previous moment, that is, , Indicates the thermal power of the motor at the current moment, Indicates the thermal power of the motor at the previous moment. The lower the thermal power consumption, the higher the reward. 、 and They represent the penalty factors of the corresponding items respectively. In this embodiment, the penalty factors 、 and are trainable parameters, and their lower limits are limited to ; Indicates the blood pump output flow at the current moment, and Respectively represent the minimum and maximum output flow of the blood pump; Indicates the current blood temperature, Indicates the upper limit of safe temperature; Indicates the aortic pressure difference at the current moment, and Represent the minimum and maximum values of aortic pressure gradient, respectively; Indicates a conditional judgment symbol. When the condition is met, the value is 1; when the condition is not met, the value is 0.
[0114] The second term in formula (14) is the constraint term for the blood pump output flow: if , that is, if it exceeds the critical allowable safety range, a penalty will be given, with a value of 1. The second item is ; The third item is the upper temperature limit constraint (thermal damage protection): If , that is, if the temperature exceeds the upper limit of safety (set to 39°C in this embodiment), a penalty will be given, with a value of 1. The third item is ; The fourth item is the pulse pressure difference constraint item (cardiac load control): If , that is, if it exceeds the clinical safety range, a penalty will be given, with a value of 1. The third item is .
[0115] Step 6: Based on the current state vector, loop through steps 2 to 5 n-1 times to obtain the speed adjustment vector, state vector, and immediate reward at n moments.
[0116] The strategy network is used to output the speed adjustment vector based on the state vector at the current moment, and then control the motor speed to obtain the key motor parameters, blood pump output flow, blood temperature and aortic pressure difference at the next moment. Then, the state vector at the next moment is generated based on the key motor parameters, blood pump output flow, blood temperature and aortic pressure difference at the next moment, and the immediate reward is calculated. The cycle is executed n-1 times, thereby obtaining the speed adjustment vector, state vector and immediate reward at n moments.
[0117] Step 7: Update the value network based on the state vector and immediate reward at n moments.
[0118] To optimize the action selection direction output by the policy network, a value network is designed to evaluate the expected cumulative return that the current state of the system can obtain in the future, that is, the expected cumulative reward. The value network provides a benchmark for the policy network during training. By minimizing the difference between its estimated reward and the actual reward, it gradually learns the "reward capacity" corresponding to each state. Unlike the policy network, the value network outputs a scalar reward value instead of an action. , indicating the level of the current state "reward", The higher it is, the more "promising" the corresponding state is. For the task of the present invention, that is, the motor thermal power consumption is minimized under the conditions of satisfying multiple constraints such as blood pump output flow, blood temperature and aortic pressure difference.
[0119] Since it is impossible to wait for the system to run a full cycle before calculating the real reward, this paper adopts a compromise solution - n-step reward, which combines the immediate reward (discounted cumulative reward) of the previous n moments and the estimated reward of the future (Nn) moments to define the real reward :
[0120] (15)
[0121] in, It represents the sum of the discounted cumulative reward calculated based on the immediate rewards at n moments and the estimated rewards at the next Nn moments, that is, the real reward; Represents a discount factor, which emphasizes the importance of adjacent moments and weakens the impact of distant moments on the reward. In this embodiment, the discount factor is set to 0.8; represents the immediate reward at time t+k, that is, the real reward at the previous time t+k; It represents the estimated reward from time t+n to time t+N, that is, the estimated value of the subsequent cumulative return is determined by the value network, is regarded as a constant, and does not participate in the back propagation of the neural network. From formula (15), we can see that It is equal to the sum of the discounted cumulative reward calculated based on the immediate rewards at n moments and the estimated reward from the value network from moment t+n to moment t+N.
[0122] The learning goal of the value network is to output an estimated reward , as close as possible to the mathematical expectation of the real reward. Since it is impossible to wait for the system to run a complete cycle before calculating the real reward, the present invention uses As the real reward for the entire cycle, the loss function of the value network can be defined as:
[0123] (16)
[0124] in, represents the loss function of the value network, represents the parameters of the value network. The loss function makes the estimated reward As close as possible The expectation is to reduce the variance in the policy network update. The value network and the policy network of this embodiment use the same network structure, the only difference is the output dimension. Before training, the parameters of the policy network and the value network are randomly initialized, the cycle length is set to 1s, the discrete time N is set to 20, and the discount factor is set to 1s. Set to 0.8.
[0125] After accumulating the speed adjustment vector, state vector, and immediate reward for n moments, the real reward is constructed according to formula (15), and the loss value of the value network is calculated according to formula (16). The parameters of the value network are updated based on the loss value. In this embodiment, the parameters of the value network are updated using the gradient update method:
[0126] (17)
[0127] in, Represents the learning rate of the value network, and its initial value is set to 1e -3 ; Represents the gradient update direction of the value network.
[0128] Step 8: Predicted reward error based on the value network Update the policy network.
[0129] Calculate the predicted reward error of the value network according to formula (13) , and then maximize the advantage direction of the action probability according to formula (12), and then perform gradient update on the parameters of the policy network:
[0130] (18)
[0131] in, represents the learning rate of the policy network, and its initial value is set to 3e -4 ; Represents the gradient of the action with respect to the parameters of the policy network.
[0132] Step 9: Repeat steps 2 to 8 based on the new starting state vector until the entire cycle is traversed.
[0133] After each update (or training) is completed, the sliding window slides back one position on the cycle sequence (i.e., N moments), and uses the new moment as the starting point to generate a new starting state vector. Repeat steps 2 to 8 until the entire cycle is traversed, that is, the traversal of N moments is completed. For example, if the starting moment of the last update or training is the i-th moment (the starting state vector is ), then the starting time of the current update or training is the i+1th moment (the starting state vector is ), repeat steps 2 to 8 until the entire cycle is traversed.
[0134] During the training phase, the controller is not limited to sampling from the start of the cycle. Instead, it adopts a sliding window strategy, constructing a sliding sample trajectory from the state vector at any time as the starting state vector, and sampling the state-action sequence for n consecutive moments. This strategy can effectively cover the state changes within the complete cycle N, helps capture local differences and long-term dependency features, and achieves full learning of the state-action mapping, so that it can dynamically match the patient's heart rate and physiological changes during the deployment phase.
[0135] Step 10: Repeat steps 1 to 9 multiple times until the termination condition is met to obtain the trained policy network.
[0136] Repeat steps 1 to 9 multiple times to complete multiple cycles of updating or training. The termination conditions of this embodiment are that the output of the value network is stable, the reward converges, and the control strategy output by the policy network tends to be optimal.
[0137] The pulsating blood pump thermal management control system based on speed modulation of the present invention is also configured to perform a pulsating blood pump motor speed control step. Figure 2 As shown, the pulsating blood pump motor speed control steps include:
[0138] Step 11: Get the key parameters of the motor at the start time ( 、 ), blood pump output flow , blood temperature and aortic pressure gradient .
[0139] Step 12: According to the key parameters of the motor at the starting moment ( 、 ), blood pump output flow , blood temperature and aortic pressure gradient Generate the state vector at the start time , as shown in formula (1).
[0140] Step 13: Use the trained policy network to calculate the state vector at the starting time , output speed adjustment vector.
[0141] Deploy the trained policy network on the controller with the state vector at the starting time As the input of the trained policy network, the trained policy network directly outputs all speed adjustments in the current cycle.
[0142] Step 14: Adjust the vector according to the speed to control the motor speed.
[0143] The controller sequentially adjusts speed at each moment based on the speed adjustment vector, eliminating the need for sliding sampling or reward estimation. Therefore, during deployment, only the policy network is retained, and the value network is not used. The controller operates using single-cycle inference, offering fast response, low resource overhead, and strong adaptability, making it suitable for real-time blood pump speed regulation tasks. The policy network's parameters are fixed after deployment and are not updated.
[0144] The above disclosure is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or modifications within the technical scope disclosed in the present invention, and they should all be covered by the scope of protection of the present invention.
Claims
1. A pulsating blood pump thermal management control system based on speed modulation, characterized in that: The control system is configured to perform the following steps: Step 1: Generate the initial state vector based on the key parameters of the motor, blood pump output flow, blood temperature and aortic pressure difference at the historical moment; Step 2: Using the strategy network to output a speed adjustment vector based on the initial state vector; Step 3: Adjust the vector according to the speed to control the motor speed; Step 4: Obtain the current key motor parameters, blood pump output flow, blood temperature, and aortic pressure difference; Step 5: Generate the current state vector based on the current motor key parameters, blood pump output flow, blood temperature, and aortic pressure difference, and calculate the immediate reward; Step 6: Based on the current state vector, loop through steps 2 to 5 for n-1 times to obtain the speed adjustment vector, state vector, and instant reward at n moments. Step 7: Update the value network based on the state vector and immediate reward at n moments; Step 8: Update the policy network based on the predicted reward error of the value network; Step 9: Repeat steps 2 to 8 based on the new starting state vector until the entire cycle is traversed; Step 10: Repeat steps 1 to 9 multiple times until the termination condition is met to obtain the trained policy network.
2. The pulsating blood pump thermal management control system based on speed modulation according to claim 1, characterized in that: The key parameters of the motor include rotational speed and accumulated thermal energy. The accumulated thermal energy is calculated based on the thermal power of the motor. The thermal power of the motor is equal to the sum of the basic thermal power and the thermal power correction value.
3. The pulsating blood pump thermal management control system based on speed modulation according to claim 2, characterized in that: The basic thermal power is calculated using the motor thermal power physical model, and the expression of the motor thermal power physical model is: ; in, Indicates the basic thermal power of the motor; represents the electromagnetic torque, represents the torque constant, Indicates the stator winding resistance; represents the core hysteresis loss coefficient, Indicates the operating frequency, represents the magnetic flux density, and m represents the exponent of hysteresis loss in the core loss; represents the material parameter describing the eddy current loss of the core; Indicates the motor speed, represents the empirical coefficient related to the first-order term of the speed, represents the empirical coefficient related to the quadratic term of speed, and t represents time.
4. The thermal management control system of a pulsating blood pump based on speed modulation according to claim 2, characterized in that: The thermal power correction value is obtained by predicting the motor operation index using a thermal power correction prediction model; The process of obtaining the thermal power correction prediction model includes: Constructing a sample data set; wherein each sample in the sample data set includes an input quantity and an output quantity, the input quantity is a motor operation index, the output quantity is a thermal power correction value, and the thermal power correction value is equal to the difference between the actual thermal power measurement value and the basic thermal power; Build an LSTM-TCN neural network model; The sample data set is used to train the LSTM-TCN neural network model to obtain the thermal power correction prediction model.
5. The pulsating blood pump thermal management control system based on speed modulation according to claim 4, characterized in that: The output of the LSTM-TCN neural network model is: ; ; ; in, represents the predicted thermal power correction value; and Represent the weight and bias of the fully connected layer respectively; represents the output of TCN; Indicates the position of the LSTM output sequence, , H represents the hidden layer dimension of LSTM; Represents the convolution kernel weight of TCN; represents the activation function; Indicates the convolution kernel size of TCN; represents the bias of TCN; Indicates At this moment, the hidden state vector output by LSTM is represents the expansion coefficient; represents the LSTM input sequence, and k represents the number of elements in the LSTM input sequence.
6. The pulsating blood pump thermal management control system based on speed modulation according to claim 1, characterized in that: The calculation formula for the instant reward is: ; in, Indicates the immediate reward at the current moment; Indicates the difference in motor thermal power between the current moment and the previous moment. , Indicates the thermal power of the motor at the current moment, Indicates the thermal power of the motor at the last moment; 、 and Respectively represent the penalty factors of the corresponding items; Indicates the blood pump output flow at the current moment, and Respectively represent the minimum and maximum output flow of the blood pump; Indicates the current blood temperature, Indicates the upper limit of safe temperature; Indicates the aortic pressure difference at the current moment, and Represent the minimum and maximum values of aortic pressure gradient, respectively; Indicates a conditional judgment symbol. When the condition is met, the value is 1; when the condition is not met, the value is 0.
7. The pulsating blood pump thermal management control system based on speed modulation according to claim 1, characterized in that: The loss function of the value network is: ; ; in, represents the loss function of the value network, represents the parameters of the value network; N represents the number of moments in the entire cycle; Represents the state vector of the value network at time t Output the estimated reward from time t to time t+N; represents the sum of the discounted cumulative reward calculated based on the immediate rewards at n moments and the estimated rewards at the next Nn moments; represents the discount factor; represents the immediate reward at time t+k; represents the estimated reward from time t+n to time t+N, determined by the value network.
8. The thermal management control system of a pulsating blood pump based on speed modulation according to claim 1, characterized in that: The loss function of the policy network is: ; ; in, represents the loss function of the policy network, Represents the parameters of the policy network; represents the predicted reward error of the value network; Represents the state vector Lower output speed adjustment vector probability; represents the speed adjustment vector; represents the sum of the discounted cumulative reward calculated based on the immediate rewards at n moments and the estimated rewards at the next Nn moments; Represents the state vector of the value network at time t Output the estimated reward from time t to time t+N.
9. The thermal management control system of a pulsating blood pump based on speed modulation according to claim 1, characterized in that: The control system is further configured to perform the following steps: Obtain key motor parameters, blood pump output flow, blood temperature, and aortic pressure difference at the start time; Generate a state vector at the starting moment according to the key parameters of the motor, the blood pump output flow, the blood temperature and the aortic pressure difference at the starting moment; Use the trained strategy network to output the speed adjustment vector based on the state vector at the starting time; The motor speed is controlled by adjusting the vector according to the speed.
Citation Information
Patent Citations
Control method of intrusive miniature axial flow blood pump based on deep reinforcement learning algorithm
CN118211458A
Temperature control method and device
CN118466632A
Control method for autonomous flight of stratospheric airship
CN118672294A
Implementation method and device for bionic pulsating blood flow of interventional artificial heart
CN118846368A
Heart pump, and method for operating a heart pump
US20190351118A1