A left ventricular endocardium image segmentation method and system based on deep reinforcement learning
By integrating the deep reinforcement learning-based method in the existing technology, the data demand and precision problems existing in the existing technology are solved, and a more accurate decision-making strategy is provided in a complex environment, which solves the data demand and precision problems existing in the existing technology and improves the segmentation accuracy.
Patent Information
- Application Number
- CN202310843676.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-07-10
AI Technical Summary
The existing left ventricular image segmentation method requires a large amount of training data and has low segmentation accuracy, and the starting point positioning is not accurate enough.
A method based on deep reinforcement learning was used to divide the left ventricular endocardium image segmentation task into two Markov decision processes: starting point positioning and contour tracing. The preset D3QN network was used for iterative optimization and updating. Combined with the reward function design such as Rtrend+Rdist+Rclus, the decision-making of the intelligent agent was realized through DQN algorithm optimization.
It achieves more accurate decision-making strategies in complex environments, solves the data demand and accuracy problems existing in existing technologies, and improves segmentation accuracy.
Smart Images

Figure CN117036689B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer medical image processing, deep learning and reinforcement learning, and more specifically, to a left ventricular endocardial image segmentation method and system based on deep reinforcement learning. Background Art
[0002] Reinforcement learning is a branch of machine learning in which models learn autonomously, achieving optimal decisions through continuous trial and exploration within their environment. Deep Q-Network (DQN) is a reinforcement learning algorithm based on deep neural networks and the Q-learning algorithm. In DQN, a deep neural network is used to approximate the Q-function and use it as the value function, enabling the model to achieve optimal decisions by maximizing the Q-function. The Q-function defines the expected reward for performing an action in a given state. Therefore, the goal of DQN is to train the network to maximize the expected reward. At its core, the Q-learning algorithm is a reinforcement learning algorithm based on the Bellman equation. This algorithm trains the neural network by randomly sampling from an experience pool (Experience Replay) and using a fixed target network (Target Network) to solve problems. These techniques help improve learning speed and stability, making the model more accurate and reliable.
[0003] The DQN algorithm has achieved remarkable results in various fields, such as computer games, robotic control, and natural language processing. In computer games, DQN has achieved human-level performance in Atari games. In robotic control, by optimizing the parameters of deep neural networks, DQN enables robots to better learn control strategies and achieve efficient and precise movements. Furthermore, in natural language processing, DQN can be used to train intelligent dialogue systems, making them more natural, human-like, and fluent.
[0004] Double DQN (Double Deep Q-Network) and Dueling DQN are both improved versions of DQN, used to improve the stability and performance of the DQN algorithm in reinforcement learning.
[0005] Double DQN addresses the overestimation problem in the DQN algorithm by using two neural networks. Specifically, in DQN, the target Q value is calculated by the target network, and due to the maximization operation, this can lead to overestimation of the target Q value. Double DQN, on the other hand, uses a separate neural network to calculate the maximum Q value, reducing this estimation error and improving model stability and performance.
[0006] Dueling DQN addresses the instability issues inherent in the DQN algorithm by decomposing the Q function into two components: a state-value function and an advantage function. In traditional DQN, the Q function directly outputs the Q value of a given action. In Dueling DQN, however, the Q function first outputs a state-value function and an advantage function, which are then combined to produce the final Q value. The state-value function evaluates the value of the current state, while the advantage function assesses the relative merits of each action relative to other actions. This split improves the Q function's representation efficiency and makes the model more stable and reliable.
[0007] The prior art discloses a left ventricular image segmentation method, apparatus, device and storage medium, the method comprising: receiving a left ventricular image to be segmented, inputting the left ventricular image into a trained image segmentation network for segmentation, obtaining and outputting a left ventricular segmentation image processed by the image segmentation network, wherein the image segmentation network is a deep learning network comprising a downsampling part and an upsampling part, the upsampling part comprising a first convolutional layer and a downsampling network layer, and the upsampling part comprising a second convolutional layer and an upsampling network layer, thereby achieving automatic segmentation of the left ventricular image and improving the efficiency and effect of left ventricular image segmentation; the method in the prior art only uses a single neural network model for left ventricular image segmentation, which not only requires a large amount of training data, but also has low positioning accuracy of the starting point, and therefore low segmentation accuracy. Summary of the Invention
[0008] In order to overcome the defects of the above-mentioned existing technologies that require a large amount of training data and have low segmentation accuracy, the present invention provides a left ventricular endocardium image segmentation method and system based on deep reinforcement learning, which can improve the performance of the model and improve the segmentation accuracy.
[0009] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0010] A left ventricular endocardium image segmentation method based on deep reinforcement learning includes the following steps:
[0011] S1: Acquire human cardiac MRI dataset and perform preprocessing;
[0012] S2: Using the preset reinforcement learning agent and the preprocessed human cardiac MRI dataset, the starting point localization task is established and converted into the first Markov decision process;
[0013] S3: Use the preset first D3QN network to perform cyclic iterative optimization on the first Markov decision process, and use the last optimization result as the optimal starting point;
[0014] S4: The reinforcement learning agent uses the optimal starting point as the starting point of contour tracking, establishes a contour tracking task based on the preprocessed human heart MRI dataset and preset auxiliary information, and converts the contour tracking task into a second Markov decision process;
[0015] S5: Utilize the preset second D3QN network to cyclically update the second Markov decision process, save the contour points obtained in each update as the left ventricular endocardial image contour, and use the left ventricular endocardial image contour as the left ventricular endocardial image segmentation result.
[0016] Preferably, the specific method of pre-processing in step S1 is:
[0017] The left ventricle ROI and several surrounding ROIs are extracted from each cardiac MRI image in the human cardiac MRI dataset. All the extracted ROIs are resized to a preset size using the bilinear interpolation method. All the resized ROIs are normalized to complete the preprocessing.
[0018] Preferably, the starting point positioning task in step S2 is specifically:
[0019] In the starting point positioning task, the first state of the reinforcement learning agent is each preprocessed human heart MRI data, and any one point is selected as the starting point of the starting point positioning task;
[0020] The action of the reinforcement learning agent is: with the current point as the center, the reinforcement learning agent selects a direction from the direction set {0,1,2,3,4,5,6,7} and moves a certain number of pixels;
[0021] 0 to 7 in the direction set represent the eight directions defined by two-dimensional Cartesian coordinates: left, upper left, upper, upper right, right, lower right, lower and lower left;
[0022] The first reward r1 of the reinforcement learning agent is: set the target point and calculate the shortest Euclidean distance d between the current point position and the target point curr , and the shortest Euclidean distance d between the next point position and the target point next , then the first reward r1 is: r1 = d curr -d next .
[0023] Preferably, the first D3QN network in step S3 is specifically:
[0024] The first D3QN network includes: an input layer, a first 2D convolutional layer, a first 2D maximum pooling layer, a second 2D convolutional layer, a second 2D maximum pooling layer, a third 2D convolutional layer, a third 2D maximum pooling layer, a fourth 2D convolutional layer, an FC module, and an output layer connected in sequence;
[0025] The FC module includes two branches, one branch is a first advantage function FC layer and a second advantage function FC layer connected in sequence, and the other branch is a first value function FC layer and a second value function FC layer connected in sequence.
[0026] Preferably, the contour tracking task in step S4 is specifically:
[0027] In the contour tracking task, the second state of the reinforcement learning agent is pre-processed human heart MRI data with added edge detection auxiliary information, and the optimal starting point obtained in the starting point positioning task is used as the starting point of contour tracking;
[0028] The action of the reinforcement learning agent is: with the current point as the center, the reinforcement learning agent selects a direction from the direction set {0,1,2,3,4,5,6,7} and moves a certain number of pixels;
[0029] 0 to 7 in the direction set represent the eight directions defined by two-dimensional Cartesian coordinates: left, upper left, upper, upper right, right, lower right, lower and lower left;
[0030] The second reward r2 of the reinforcement learning agent is:
[0031] r2=R trend +R dist +R clus
[0032]
[0033]
[0034]
[0035] Among them, d1 is the distance between the current point and the target point expected to be reached in the next step, d2 is the distance between the point actually reached in the next step and the target point expected to be reached in the next step, step_size is the action step size of the agent, davg is the average of d1 and d2, D is the preset threshold; a is the preset discount; std is the standard deviation of the coordinates of the n trajectory points recently traversed by the agent.
[0036] Preferably, the edge detection auxiliary information is specifically Sobel operator edge detection information.
[0037] Preferably, the second D3QN network in step S5 is specifically:
[0038] The second D3QN network includes a Resnet18 network, a 2D adaptive average pooling layer and the FC module connected in sequence.
[0039] The present invention also provides a left ventricular endocardial image segmentation system based on deep reinforcement learning, which applies the above-mentioned left ventricular endocardial image segmentation method based on deep reinforcement learning, including:
[0040] Data preprocessing module: used to obtain and preprocess human cardiac MRI datasets;
[0041] Starting point localization task construction module: used to establish the starting point localization task using the preset reinforcement learning agent and the preprocessed human cardiac MRI dataset and convert it into the first Markov decision process;
[0042] Starting point positioning module: used to use the preset first D3QN network to perform cyclic iterative optimization on the first Markov decision process, and use the last optimization result as the optimal starting point;
[0043] Contour tracking task construction module: The reinforcement learning agent uses the optimal starting point as the starting point for contour tracking, establishes the contour tracking task based on the preprocessed human cardiac MRI dataset and preset auxiliary information, and transforms the contour tracking task into a second Markov decision process;
[0044] Contour tracking module: used to use the preset second D3QN network to cyclically update the second Markov decision process, save the contour points obtained in each update as the left ventricular endocardial image contour, and use the left ventricular endocardial image contour as the left ventricular endocardial image segmentation result.
[0045] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps in the above method when executed by a processor.
[0046] The present invention provides an electronic device, comprising a processor and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the above method are executed.
[0047] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0048] The present invention provides a left ventricular endocardial image segmentation method and system based on deep reinforcement learning. First, a human heart MRI data set is obtained and preprocessed; a starting point positioning task is established using a preset reinforcement learning agent and the preprocessed human heart MRI data set and converted into a first Markov decision process; the first Markov decision process is cyclically iteratively optimized using a preset first D3QN network, and the last optimization result is used as the optimal starting point; the reinforcement learning agent uses the optimal starting point as the starting point of contour tracking, establishes a contour tracking task based on the preprocessed human heart MRI data set and preset auxiliary information, and converts the contour tracking task into a second Markov decision process; finally, the second Markov decision process is cyclically updated using a preset second D3QN network, and the contour points obtained in each update are collectively saved as the left ventricular endocardial image contour, and the left ventricular endocardial image contour is used as the left ventricular endocardial image segmentation result;
[0049] The present invention provides a new method for locating the starting point of contour tracking. Unlike the existing supervised learning method that outputs contour features and then selects points on the feature map, the reinforcement learning decision-making process of this method is similar to the human thinking and decision-making process, starting from a certain position and moving until reaching the destination; the reward function is innovatively designed, among which the trend reward Rtrend is the most effective. This reward can enable the intelligent agent to move closer to the target contour, and also provides inspiration for the reward setting of similar tracking tasks; at the same time, the present invention introduces Double DQN and Dueling DQN to form D3QN, which has better convergence speed and stronger model expression ability than traditional DQN, can better handle the relationship between state value and action value, reduce the problem of over-estimation, and can provide more accurate and efficient decision-making strategies in complex environments; in addition, this method can solve the problem that deep learning-based methods require a large amount of data, because reinforcement learning can learn through trial and error, and can also use prior knowledge for learning, thereby improving the performance of the model and achieving higher segmentation accuracy compared with deep learning baseline methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flow chart of a left ventricular endocardium image segmentation method based on deep reinforcement learning provided in Example 1.
[0051] Figure 2 This is a flowchart of a left ventricular endocardium image segmentation method based on deep reinforcement learning provided in Example 2.
[0052] Figure 3 This is a schematic diagram of the direction set provided in Example 2.
[0053] Figure 4This is a diagram of the second D3QN network structure provided in Example 2.
[0054] Figure 5 This is a structural diagram of a left ventricular endocardium image segmentation system based on deep reinforcement learning provided in Example 3. DETAILED DESCRIPTION
[0055] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;
[0056] In order to better illustrate this embodiment, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product size;
[0057] It is understandable to those skilled in the art that some well-known structures and descriptions thereof may be omitted in the drawings.
[0058] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0059] Example 1
[0060] like Figure 1 As shown, this embodiment provides a left ventricular endocardium image segmentation method based on deep reinforcement learning, comprising the following steps:
[0061] S1: Acquire human cardiac MRI dataset and perform preprocessing;
[0062] S2: Using the preset reinforcement learning agent and the preprocessed human cardiac MRI dataset, the starting point localization task is established and converted into the first Markov decision process;
[0063] S3: Use the preset first D3QN network to perform cyclic iterative optimization on the first Markov decision process, and use the last optimization result as the optimal starting point;
[0064] S4: The reinforcement learning agent uses the optimal starting point as the starting point of contour tracking, establishes a contour tracking task based on the preprocessed human heart MRI dataset and preset auxiliary information, and converts the contour tracking task into a second Markov decision process;
[0065] S5: Utilize the preset second D3QN network to cyclically update the second Markov decision process, save the contour points obtained in each update as the left ventricular endocardial image contour, and use the left ventricular endocardial image contour as the left ventricular endocardial image segmentation result.
[0066] In the specific implementation process, a human heart MRI data set is first obtained and preprocessed; a starting point positioning task is established using a preset reinforcement learning agent and the preprocessed human heart MRI data set and converted into a first Markov decision process; the first Markov decision process is cyclically iteratively optimized using a preset first D3QN network, and the last optimization result is used as the optimal starting point; the reinforcement learning agent uses the optimal starting point as the starting point for contour tracking, establishes a contour tracking task based on the preprocessed human heart MRI data set and preset auxiliary information, and converts the contour tracking task into a second Markov decision process; finally, the second Markov decision process is cyclically updated using a preset second D3QN network, and the contour points obtained in each update are saved together as the left ventricular endocardial image contour, and the left ventricular endocardial image contour is used as the left ventricular endocardial image segmentation result;
[0067] The main difference between the starting point search process and the contour tracing process in this method is that the starting point search is to obtain the last position of the agent as the starting point of the contour tracing module, while the contour tracing module collects all the positions that the agent has passed and then forms a contour as the segmentation result;
[0068] This method divides the left ventricular endocardium image segmentation task into two Markov decision processes and trains them based on reinforcement learning and D3QN network, which can effectively improve the performance of the model and improve the segmentation accuracy.
[0069] Example 2
[0070] like Figure 2 As shown, this embodiment provides a left ventricular endocardium image segmentation method based on deep reinforcement learning, comprising the following steps:
[0071] S1: Acquire human cardiac MRI dataset and perform preprocessing;
[0072] S2: Using the preset reinforcement learning agent and the preprocessed human cardiac MRI dataset, the starting point localization task is established and converted into the first Markov decision process;
[0073] S3: Use the preset first D3QN network to perform cyclic iterative optimization on the first Markov decision process, and use the last optimization result as the optimal starting point;
[0074] S4: The reinforcement learning agent uses the optimal starting point as the starting point of contour tracking, establishes a contour tracking task based on the preprocessed human heart MRI dataset and preset auxiliary information, and converts the contour tracking task into a second Markov decision process;
[0075] S5: Using the preset second D3QN network to cyclically update the second Markov decision process, the contour points obtained in each update are saved together as the left ventricular endocardial image contour, and the left ventricular endocardial image contour is used as the left ventricular endocardial image segmentation result;
[0076] The specific method of pre-processing in step S1 is:
[0077] Extract the left ventricle ROI and several surrounding ROIs from each cardiac MRI image in the human cardiac MRI dataset, resize all extracted ROIs to a preset size using bilinear interpolation, and normalize all resized ROIs to complete preprocessing.
[0078] The starting point positioning task in step S2 is specifically:
[0079] In the starting point positioning task, the first state of the reinforcement learning agent is each preprocessed human heart MRI data, and any one point is selected as the starting point of the starting point positioning task;
[0080] The action of the reinforcement learning agent is: with the current point as the center, the reinforcement learning agent selects a direction from the direction set {0,1,2,3,4,5,6,7} and moves a certain number of pixels;
[0081] The direction 0 to 7 in the direction set represent the eight directions of left, upper left, upper, upper right, right, lower right, lower and lower left defined according to the two-dimensional Cartesian coordinates. Figure 3 As shown;
[0082] The first reward r1 of the reinforcement learning agent is: set the target point and calculate the shortest Euclidean distance d between the current point position and the target point curr , and the shortest Euclidean distance d between the next point position and the target point next , then the first reward r1 is: r1 = d curr -d next ;
[0083] The first D3QN network in step S3 is specifically:
[0084] The first D3QN network includes: an input layer, a first 2D convolutional layer, a first 2D maximum pooling layer, a second 2D convolutional layer, a second 2D maximum pooling layer, a third 2D convolutional layer, a third 2D maximum pooling layer, a fourth 2D convolutional layer, an FC module, and an output layer connected in sequence;
[0085] The FC module includes two branches, one branch is a first advantage function FC layer and a second advantage function FC layer connected in sequence, and the other branch is a first value function FC layer and a second value function FC layer connected in sequence;
[0086] The contour tracking task in step S4 is specifically:
[0087] In the contour tracking task, the second state of the reinforcement learning agent is pre-processed human heart MRI data with added edge detection auxiliary information, and the optimal starting point obtained in the starting point positioning task is used as the starting point of contour tracking;
[0088] The action of the reinforcement learning agent is: with the current point as the center, the reinforcement learning agent selects a direction from the direction set {0,1,2,3,4,5,6,7} and moves a certain number of pixels;
[0089] 0 to 7 in the direction set represent the eight directions defined by two-dimensional Cartesian coordinates: left, upper left, upper, upper right, right, lower right, lower and lower left;
[0090] The second reward r2 of the reinforcement learning agent is:
[0091] r2=R trend +R dist +R clus
[0092]
[0093]
[0094]
[0095] Where d1 is the distance between the current point and the target point expected to be reached in the next step, d2 is the distance between the point actually reached in the next step and the target point expected to be reached in the next step, step_size is the action step size of the agent, davg is the average of d1 and d2, D is the preset threshold; a is the preset discount; std is the standard deviation of the coordinates of the n trajectory points recently traveled by the agent;
[0096] The edge detection auxiliary information is specifically Sobel operator edge detection information;
[0097] like Figure 4 As shown, the second D3QN network in step S5 is specifically:
[0098] The second D3QN network includes a Resnet18 network, a 2D adaptive average pooling layer and the FC module connected in sequence.
[0099] In the specific implementation process, a human cardiac MRI dataset is first obtained and preprocessed. This embodiment uses two datasets, namely the Automated Cardiac Diagnosis Challenge 2017 (ACDC 2017) and the Sunnybrook Cardiac MR Left Ventricle Segmentation Challenge (Sunnybrook2009). The goal of this method is to segment the left ventricular endocardium. Since the left ventricle occupies only a small area in cardiac MRI, it is necessary to extract the ROI around the left ventricle. 1670 ROIs are extracted from the ACDC 2017 dataset, and the image size is resized to 368×368 using bilinear interpolation, and finally normalized as preprocessing. Similarly, 800 ROIs are extracted from the Sunnybrook2009 dataset, the image size is resized, and then normalized.
[0100] Then, the preset reinforcement learning agent and the preprocessed human cardiac MRI dataset are used to establish the starting point positioning task and transform it into the first Markov decision process;
[0101] In the starting point positioning task, the first state of the reinforcement learning agent is each preprocessed human heart MRI data, and any one point is selected as the starting point of the starting point positioning task;
[0102] The state is the input of the neural network, which should contain important image information so that the agent can make the best action. If the input image size is too large, the training time will be very long. For an original grayscale image (W×H), it needs to be cropped to (w×h) based on the current position of the agent, where w and h are both 64, representing the current field of view of the agent. Then, the historical 3-step position map needs to be spliced, and the dimension is (64×64×4). If the current number of steps is less than 4, a copy of the current position map is spliced. The agent starts from a random position P within 80% of the inner area of the image. The agent terminates navigation when it finds any point on the target contour. During the search process, the termination is triggered when the agent oscillates around the target contour point.
[0103] Given a state, the state is input into the neural network. The agent outputs a direction from the direction space {0, 1, 2, 3, 4, 5, 6, 7}. The output is a discrete value in the direction space, representing the direction of the next step. With the agent's current point as the center, eight actions corresponding to the eight "neighborhoods" of the current point are determined. These are defined based on the eight directions in two-dimensional Cartesian coordinates: left, upper left, upper, upper right, right, lower right, lower, and lower left. A "neighborhood" is defined as a distance of 5 pixels from the current point, representing a 5-pixel move per step.
[0104] The first reward r1 of the reinforcement learning agent is: set the target point and calculate the shortest Euclidean distance d between the current point position and the target point curr , and the shortest Euclidean distance d between the next point position and the target point next , then the first reward r1 is: r1 = d curr -d next The first reward indicates that if the agent's movement is towards the target contour point, this ensures a positive reward, otherwise a penalty is given;
[0105] Use the preset first D3QN network to perform cyclic iterative optimization on the first Markov decision process, and use the last optimization result as the optimal starting point;
[0106] The structure of the first D3QN network is shown in Table 1. The model takes a tensor of size (1×4×64×64) as input. It consists of four 2D convolutional layers and three 2D maximum pooling layers, and then connects the Dueling structure, namely the FC module: the FC module consists of two branches, one branch is the first advantage function FC layer and the second advantage function FC layer connected in sequence, and the other branch is the first value function FC layer and the second value function FC layer connected in sequence. Finally, the output is merged and the size is the same as the size of the action space;
[0107] Table 1: The first D3QN network structure
[0108]
[0109] The reinforcement learning agent then uses the optimal starting point as the starting point for contour tracking, establishes a contour tracking task based on the preprocessed human cardiac MRI dataset and preset auxiliary information, and transforms the contour tracking task into a second Markov decision process.
[0110] In the contour tracking task, the second state of the reinforcement learning agent is pre-processed human heart MRI data with added Sobel operator edge detection information, and the optimal starting point obtained in the starting point positioning task is used as the starting point for contour tracking;
[0111] Similarly, if the image size input to the neural network is too large, the training time will be very long. For an original grayscale image (W×H), it needs to be cropped to (w×h) based on the current position of the agent, where w and h are both 64, representing the current field of view of the agent. When the agent first starts to track the contour of the endothelium, the current position is the point P found last in the starting point positioning task, and then it is updated to the position of the next step. In addition, some auxiliary information needs to be added to provide more details. Here, the Sobel operator is used to detect edge information. There is also a map of historical trajectory points. Adding this can effectively prevent the agent from looking back or staying. The ROI grayscale image and these two images are spliced together, and then cut according to the current position of the agent, that is, the state dimension of the contour tracking task is (64×64×3). Finally, the agent terminates navigation when it meets the starting point again after taking a certain number of steps. Collecting all the trajectory points passed forms the contour of the object to be segmented as the segmentation result.
[0112] The action of the reinforcement learning agent is: with the current point as the center, the reinforcement learning agent selects a direction from the direction set {0,1,2,3,4,5,6,7} and moves a certain number of pixels;
[0113] 0 to 7 in the direction set represent the eight directions defined by two-dimensional Cartesian coordinates: left, upper left, upper, upper right, right, lower right, lower and lower left;
[0114] The second reward r2 of the reinforcement learning agent is:
[0115] r2=R trend +R dist +R clus
[0116]
[0117]
[0118]
[0119] Where d1 is the distance between the current point and the target point expected to be reached in the next step, d2 is the distance between the point actually reached in the next step and the target point expected to be reached in the next step, step_size is the action step size of the agent, which is 5 in this embodiment, davg is the average of d1 and d2, D is the preset threshold, which is 2 times the action step size in this embodiment; a is the preset discount, which is 0.05 in this embodiment; std is the standard deviation of the coordinates of the n trajectory points recently traversed by the agent, and n = 20 in this embodiment;
[0120] The second reward means that if the agent's movement is close to the target outline, it will receive a positive reward; Rtrend It is the trend reward, which means that if the next position is closer to the target contour than the current position, that is, (d1-d2)>0, then the current direction is a good trend and a positive reward is given, otherwise the reward is 0; R dist is the distance reward, which means that the closer the next position of the agent is to the target contour point, the greater the reward it can get; R clus For the aggregation reward, it means that if the most recent 20 trajectory points are clustered, a penalty will be given to encourage the agent to jump out of the cluster and keep walking around the target contour point;
[0121] Finally, the preset second D3QN network is used to cyclically update the second Markov decision process, and the contour points obtained in each update are saved as the left ventricular endocardial image contour, which is used as the left ventricular endocardial image segmentation result;
[0122] The second D3QN network in this embodiment includes a Resnet18 network, a 2D adaptive average pooling layer, and the FC module connected in sequence. The output of the 2D adaptive average pooling layer is flattened into a one-dimensional feature, which is then input into the first advantage function FC layer and the second advantage function FC layer to obtain an advantage value. The one-dimensional feature is then input into the first value function FC layer and the second value function FC layer to obtain a value. Finally, the value is added to the difference between the advantage value of each action and the mean of the advantage value to obtain an output.
[0123] This method divides the left ventricular endocardium image segmentation task into two Markov decision processes and trains them based on reinforcement learning and D3QN network, which can effectively improve the performance of the model and improve the segmentation accuracy.
[0124] Example 3
[0125] like Figure 5 As shown, this embodiment provides a left ventricular endocardial image segmentation system based on deep reinforcement learning, applying the left ventricular endocardial image segmentation method based on deep reinforcement learning described in Example 1 or 2, including:
[0126] Data preprocessing module 301: used to obtain human heart MRI data set and perform preprocessing;
[0127] A starting point positioning task construction module 302 is used to establish a starting point positioning task using a preset reinforcement learning agent and a preprocessed human heart MRI dataset and convert it into a first Markov decision process;
[0128] The starting point positioning module 303 is used to perform cyclic iterative optimization on the first Markov decision process using the preset first D3QN network, and use the last optimization result as the optimal starting point;
[0129] Contour tracking task construction module 304: The reinforcement learning agent uses the optimal starting point as the starting point for contour tracking, establishes a contour tracking task based on the preprocessed human heart MRI dataset and preset auxiliary information, and converts the contour tracking task into a second Markov decision process;
[0130] Contour tracking module 305: used to use the preset second D3QN network to cyclically update the second Markov decision process, save the contour points obtained in each update as the left ventricular endocardial image contour, and use the left ventricular endocardial image contour as the left ventricular endocardial image segmentation result.
[0131] In the specific implementation process, first, the data preprocessing module 301 obtains the human heart MRI data set and preprocesses it; the starting point positioning task construction module 302 uses the preset reinforcement learning agent and the preprocessed human heart MRI data set to establish the starting point positioning task and convert it into a first Markov decision process; the starting point positioning module 303 uses the preset first D3QN network to cyclically iteratively optimize the first Markov decision process, and uses the last optimization result as the optimal starting point; the reinforcement learning agent uses the optimal starting point as the starting point of contour tracking, and the contour tracking task construction module 304 establishes a contour tracking task based on the preprocessed human heart MRI data set and preset auxiliary information, and converts the contour tracking task into a second Markov decision process; finally, the contour tracking module 305 uses the preset second D3QN network to cyclically update the second Markov decision process, saves the contour points obtained in each update as the left ventricular endocardial image contour, and uses the left ventricular endocardial image contour as the left ventricular endocardial image segmentation result;
[0132] This system divides the left ventricular endocardium image segmentation task into two Markov decision processes and trains them based on reinforcement learning and D3QN network, which can effectively improve the performance of the model and improve the segmentation accuracy.
[0133] The same or similar reference numerals correspond to the same or similar components;
[0134] The terms used in the drawings to describe positional relationships are for illustrative purposes only and should not be construed as limiting this patent;
[0135] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A left ventricular endocardium image segmentation method based on deep reinforcement learning, characterized in that: The following steps are involved: S1: Acquire human cardiac MRI dataset and perform preprocessing; S2: Using the preset reinforcement learning agent and the preprocessed human cardiac MRI dataset, the starting point localization task is established and converted into the first Markov decision process; The starting point positioning task is specifically: In the starting point positioning task, the first state of the reinforcement learning agent is each preprocessed human heart MRI data, and any one point is selected as the starting point of the starting point positioning task; The action of the reinforcement learning agent is: with the current point as the center, the reinforcement learning agent selects a direction from the direction set {0,1,2,3,4,5,6,7} and moves a certain number of pixels; The directions 0 to 7 in the direction set represent the eight directions defined by two-dimensional Cartesian coordinates: left, upper left, upper, upper right, right, lower right, lower and lower left. The first reward of the reinforcement learning agent To: Set the target point and calculate the shortest Euclidean distance d between the current point and the target point curr , and the shortest Euclidean distance d between the next point position and the target point next , then the first reward for: ; S3: Use the preset first D3QN network to perform cyclic iterative optimization on the first Markov decision process, and use the last optimization result as the optimal starting point; S4: The reinforcement learning agent uses the optimal starting point as the starting point of contour tracking, establishes a contour tracking task based on the preprocessed human heart MRI dataset and preset auxiliary information, and converts the contour tracking task into a second Markov decision process; S5: Utilize the preset second D3QN network to cyclically update the second Markov decision process, save the contour points obtained in each update as the left ventricular endocardial image contour, and use the left ventricular endocardial image contour as the left ventricular endocardial image segmentation result.
2. The left ventricular endocardium image segmentation method based on deep reinforcement learning according to claim 1, characterized in that: The specific method of pre-processing in step S1 is: The left ventricle ROI and several surrounding ROIs are extracted from each cardiac MRI image in the human cardiac MRI dataset. All the extracted ROIs are resized to a preset size using the bilinear interpolation method. All the resized ROIs are normalized to complete the preprocessing.
3. The left ventricular endocardium image segmentation method based on deep reinforcement learning according to claim 2, characterized in that: The first D3QN network in step S3 is specifically: The first D3QN network includes: an input layer, a first 2D convolutional layer, a first 2D maximum pooling layer, a second 2D convolutional layer, a second 2D maximum pooling layer, a third 2D convolutional layer, a third 2D maximum pooling layer, a fourth 2D convolutional layer, an FC module, and an output layer connected in sequence; The FC module includes two branches, one branch is a first advantage function FC layer and a second advantage function FC layer connected in sequence, and the other branch is a first value function FC layer and a second value function FC layer connected in sequence.
4. The left ventricular endocardium image segmentation method based on deep reinforcement learning according to claim 3, characterized in that: The contour tracking task in step S4 is specifically: In the contour tracking task, the second state of the reinforcement learning agent is pre-processed human heart MRI data with added edge detection auxiliary information, and the optimal starting point obtained in the starting point positioning task is used as the starting point of contour tracking; The action of the reinforcement learning agent is: with the current point as the center, the reinforcement learning agent selects a direction from the direction set {0,1,2,3,4,5,6,7} and moves a certain number of pixels; The directions 0 to 7 in the direction set represent the eight directions defined by two-dimensional Cartesian coordinates: left, upper left, upper, upper right, right, lower right, lower and lower left. Second Reward for Reinforcement Learning Agents for: Where d1 is the distance between the current point and the target point expected to be reached in the next step, d2 is the distance between the point actually reached in the next step and the target point expected to be reached in the next step, step_size is the action step size of the agent, davg is the average of d1 and d2, and D is the preset threshold; is the preset discount; std is the standard deviation of the coordinates of the n trajectory points that the agent has recently walked through.
5. The left ventricular endocardium image segmentation method based on deep reinforcement learning according to claim 4, characterized in that: The edge detection auxiliary information is specifically Sobel operator edge detection information.
6. The left ventricular endocardium image segmentation method based on deep reinforcement learning according to claim 5, characterized in that: The second D3QN network in step S5 is specifically: The second D3QN network includes a Resnet18 network, a 2D adaptive average pooling layer and the FC module connected in sequence.
7. A left ventricular endocardial image segmentation system based on deep reinforcement learning, applying the left ventricular endocardial image segmentation method based on deep reinforcement learning according to any one of claims 1 to 6, characterized in that: include: Data preprocessing module: used to obtain and preprocess human cardiac MRI datasets; Starting point localization task construction module: used to establish the starting point localization task using the preset reinforcement learning agent and the preprocessed human cardiac MRI dataset and convert it into the first Markov decision process; Starting point positioning module: used to use the preset first D3QN network to perform cyclic iterative optimization on the first Markov decision process, and use the last optimization result as the optimal starting point; Contour tracking task construction module: The reinforcement learning agent uses the optimal starting point as the starting point for contour tracking, establishes the contour tracking task based on the preprocessed human cardiac MRI dataset and preset auxiliary information, and transforms the contour tracking task into a second Markov decision process; Contour tracking module: used to use the preset second D3QN network to cyclically update the second Markov decision process, save the contour points obtained in each update as the left ventricular endocardial image contour, and use the left ventricular endocardial image contour as the left ventricular endocardial image segmentation result.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
3D medical image detection method and system
CN114792311A
Method and system for detecting landmarks in medical images
US20210287363A1