Medical image segmentation method and system based on deep reinforcement learning assistance

By adopting deep reinforcement learning-assisted methods in medical image segmentation, and using SAM backbone network and PPO algorithm to optimize agent strategies, the problem of low accuracy of existing medical image segmentation methods is solved, and higher segmentation accuracy and better scalability are achieved.

CN120125816APending Publication Date: 2025-06-10INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510177726.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have the problem of insufficient accuracy, especially when the scale of medical image data sets is relatively limited, it is difficult to achieve good training results.

Method used

The medical image segmentation method based on deep reinforcement learning assisted is adopted, and the interaction between the SAM backbone network and the reinforcement learning agent is used to optimize the agent's strategy to improve the accuracy of image segmentation.

Benefits of technology

This method not only enhances the scalability and applicability of image segmentation, but also effectively improves the accuracy of image segmentation, which can reduce the workload of manual labeling while improving segmentation accuracy in clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125816A_ABST
    Figure CN120125816A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image segmentation method and system based on deep reinforcement learning assistance, and belongs to the technical field of image segmentation. The method comprises the steps of obtaining a state st and a reward rt of an image sample x on each time step based on interaction between an SAM backbone network and a reinforcement learning agent; performing back propagation according to the state st and the reward rt to obtain a trained agent; and based on the SAM backbone network and the trained intelligent agent, obtaining a segmentation result of the test image. According to the method, the expansibility and applicability of image segmentation are enhanced, and the accuracy of image segmentation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and particularly to a medical image segmentation method and system assisted by deep reinforcement learning. Background Art

[0002] In the existing field of using machine learning for medical image segmentation, there are generally two technical solutions: methods based on traditional machine learning and methods based on deep neural networks.

[0003] Methods based on traditional machine learning usually rely on feature engineering and classical machine learning algorithms. By artificially designing the features required for segmentation, image data is processed, which is suitable for early medical image segmentation tasks or specific problems. However, this type of method requires artificial feature design, is overly dependent on expert experience, and is difficult to handle complex problems in the field of medical images.

[0004] Image segmentation methods based on deep neural networks mainly rely on convolutional neural networks and their variants, such as fully convolutional neural networks and U-Net image segmentation networks. Due to the powerful expressiveness of the neural networks they rely on, these methods have become the mainstream image segmentation methods. However, the training process of neural networks requires a large amount of data for supervised learning, and the scale of the dataset that can be provided in the field of medical images is relatively limited, making it difficult to achieve good training results.

[0005] Therefore, the existing machine learning methods have the problem of insufficient accuracy. Summary of the Invention

[0006] The present invention proposes a medical image segmentation method and system assisted by deep reinforcement learning, which not only enhances the scalability and applicability of image segmentation, but also can effectively improve the accuracy of image segmentation.

[0007] To achieve the above object, the technical solution of the present invention includes the following content.

[0008] A medical image segmentation method assisted by deep reinforcement learning, the method comprising:

[0009] Based on the interaction between the SAM backbone network and the reinforcement learning agent, obtaining the state s of the image sample x at each time step t and the reward r t ; wherein, the SAM backbone network includes: an image encoder, a prompt encoder, and a mask decoder, and the agent includes: an actor network and a critic network;

[0010] According to the state s t and the reward r t perform backpropagation to obtain the trained agent;

[0011] Based on the SAM backbone network and the trained agent, obtain the segmentation result of the test image.

[0012] Furthermore, based on the interaction between the SAM backbone network and the reinforcement learning agent, obtain the state s of the image sample x at each time step t and the reward r t , including:

[0013] Divide the image sample x into small pieces, and use the center points of the small pieces as candidate actions to define the action space;

[0014] Initialize the prompt set P 0 , where the prompt set P 0 contains the positive sample a selected from the real sample y p and the negative sample a n ;

[0015] Use the image encoder to calculate the embedded representation E I (x);

[0016] The mask decoder calculates the initial mask M I based on the embedded representation E 0 (x) and the prompt set P 0 ;

[0017] Multiply the embedded representation E I (x) by the initial mask M 0 to obtain the initial state s 0

[0018] Obtain the state s of the image sample x at the t-th time step t , where t is a natural number;

[0019] The actor network selects an action a θ in the action space according to the current policy π t and the state s t ;

[0020] Add the action a t to the prompt set P t-1 to obtain the prompt set P t ;

[0021] The mask decoder calculates the mask M t based on the prompt set P I and the embedded representation E t ;

[0022] Use the mask M t and the embedded representation E IObtain the state s at the (t + 1)-th time step by (x). t+1 ;

[0023] According to the mask M t and the true sample y, obtain the reward r t+1 .

[0024] Furthermore, the actor network selects an action a in the action space according to the current policy π θ and the state s t , including: t :

[0025] Input the state s t into the actor network with the input policy π θ to obtain the distribution logits; where the distribution logits are a subset of actions in the action space;

[0026] Sample the action a from the distribution logits t .

[0027] Furthermore, perform backpropagation according to the state s t and the reward r t to obtain the trained agent, including:

[0028] The critic network estimates the value V t of the state s φ (s t ); where V φ represents the critic network with parameters φ;

[0029] According to the value V φ (s t ) and the reward r t , calculate the advantage at the t-th time step and obtain the probability ratio r t (θ) of the actor network between the old and new policies;

[0030] Based on the advantage and the probability ratio r t (θ), generate the loss L PPO (θ) of the actor network, and according to the value V φ (s t ) and the reward r t , obtain the loss L VF (φ) of the critic network;

[0031] Based on the loss L PPO (θ), the loss L VF(φ) and the policy entropy loss H(π θ ), to obtain the total loss L total ; wherein, the policy entropy loss H(π θ ) is used to encourage the agent to explore by increasing the entropy of the policy;

[0032] According to the total loss L total perform backpropagation on the agent to update the parameters of the agent.

[0033] Furthermore, according to the value V φ (s t ) and the reward r t , calculate the advantage at the t-th time step including:

[0034] Calculate the discounted cumulative return R from time step t to the last time step t ;

[0035] According to the discounted cumulative return R t and the value V φ (s t ), obtain the advantage at the t-th time step

[0036] Furthermore, the loss L VF (φ) = E t [(V φ (s t ) - R t ) 2 .

[0037] Furthermore, the probability ratio r t (θ) between the old and new policies = exp(log π θ (a t |s t ) - log π θold (a t |s t )); where π θold is the old policy.

[0038] Furthermore, the loss L PPO (θ) = E t [min(r t (θ)A k , clip(r t (θ), 1 - ∈, 1 + ∈)A k )]; where clip represents the clipping function and ∈ represents the hyperparameter threshold.

[0039] Further, based on the SAM backbone network and the trained agent, obtain the segmentation result of the test image, including:

[0040] Input the test image into the trained agent multiple times, and obtain a set of prompt points according to the best action selected by the trained agent each time;

[0041] Input the set of prompt points and the test image into the SAM backbone network, and use the mask output by the SAM backbone network as the segmentation result of the test image.

[0042] A medical image segmentation system assisted by deep reinforcement learning, the system includes:

[0043] A training module, configured to obtain the state s at each time step of the image sample x based on the interaction between the SAM backbone network and the reinforcement learning agent t and the reward r t ; wherein, the SAM backbone network includes: an image encoder, a prompt encoder, and a mask decoder, and the agent includes: an actor network and a critic network; perform backpropagation according to the state s t and the reward r t to obtain the trained agent;

[0044] A test module, configured to obtain the segmentation result of the test image based on the SAM backbone network and the trained agent.

[0045] Compared with the prior art, the present invention has at least the following beneficial effects.

[0046] 1) The present invention uses the SAM agent as the backbone network and integrates the Proximal Policy Optimization (PPO) algorithm to improve the effect of prompt-based medical image segmentation.

[0047] 2) The present invention uses reinforcement learning to effectively handle medical image segmentation through task-specific prompt optimization.

[0048] 3) The present invention uses the SAM model as a part of the network, making the overall framework a scalable framework that can be split and replaced, which not only enhances the scalability and applicability, but also reduces the manual annotation workload while improving the segmentation accuracy in clinical applications.

[0049] 4) The agent of the present invention is used in finding the prompts provided to SAM, making the work of the agent simpler and the efficiency of the overall framework higher. Description of the Drawings

[0050] Figure 1 Framework diagram of a medical image segmentation method assisted by deep reinforcement learning. Detailed Embodiments

[0051] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention and make the objectives, features, and advantages of the present invention more obvious and understandable, the following further elaborates on the core technology of the present invention in combination with examples.

[0052] The medical image segmentation method based on deep reinforcement learning assistance in the present invention uses Med-SA as the basic framework, and additionally introduces the reinforcement learning method of PPO (Proximal Policy Optimization) in this basic framework. It can not only address the problems in the field of medical image segmentation by further enhancing its structure, but also enable it to use optimized prompts to iteratively refine the segmentation boundary. Therefore, the present invention can dynamically adapt to complex scenarios, improve the segmentation accuracy without manual intervention, and can be adaptively optimized.

[0053] Figure 1 It is a framework diagram of the medical image segmentation method based on deep reinforcement learning assistance. This Figure 1 shows the interaction between the SAM backbone network and the reinforcement learning agent. The agent interacts with the SAM framework through the PPO algorithm to iteratively update the network. For simplicity of representation, the SAM backbone network is divided into three core components: an image encoder, a prompt encoder, and a mask decoder, and the agent is implemented based on the actor-critic framework. The following specifically discusses the implementation solution of the present invention.

[0054] Step 1: Based on the interaction between the SAM backbone network and the reinforcement learning agent, obtain the state s t and reward r t at each time step for the image sample x.

[0055] The reinforcement learning in the present invention uses PPO as the main structure, and the implementation of PPO is as follows:

[0056] Action space: The present invention divides the image into small blocks and defines the action space by taking the center points of the small blocks as candidate actions. Specifically, given an image x ∈ R H×W×C , the image is divided into small blocks x p ∈ R H′×W′×C , where H, W, and C represent the height, width, and number of channels of the image respectively, and H′, W′ represent the width and height of the divided small blocks. The total action space A contains the center points of each small block, and is formally defined as: A = {(w p , h p ) = center(x p )}, where x p represents the p-th small block, and w p , h p represent the coordinates of the center point of the p-th small block.

[0057] State Space: To represent the state at each time step, the present invention utilizes the intrinsic spatial features of the image and the mask predicted by the model in the previous step to calculate the state of the time step. Specifically, the state at the t-th step is defined as the product of the image embedding value and the mask from the previous step, with the expression: s t = E I (x) · M t-1 . Where, E I (x) represents the spatial features of the given image x extracted from the image encoder, and M t-1 represents the mask generated by the backbone network (the mask decoder in the SAM backbone network) in the previous step. This multiplication operation selectively retains the image embedding values corresponding to the foreground regions, ensuring that their features remain unchanged, while setting the embedding values of the background regions to zero, effectively discarding irrelevant information.

[0058] Reward Function: The design of the reward function aims to encourage actions that can improve the segmentation performance. The basic reward is defined as follows: if the selected action falls within the foreground region of the mask, the reward increases by 1; otherwise, the reward decreases by 1. To account for cases where, although the hint points are correct, they do not significantly improve the segmentation effect, an additional reward term based on the Dice coefficient is introduced. The final reward calculation formula is:

[0059] reward = Dice(M, G) + Correctness

[0060] Where, Dice(M, G) is the Dice coefficient calculated through the predicted mask M and the ground truth mask G, and Correctness is 1 when the action falls within the foreground region of the mask, and -1 otherwise. This reward design ensures that the model can improve the segmentation quality while preferentially selecting correct actions.

[0061] Training Process: In the framework of the present invention, Proximal Policy Optimization (PPO) is used to refine the segmentation by optimizing the interaction between the backbone network and the reinforcement learning agent. PPO uses a clipped surrogate objective function to limit the policy update within a trust region, thus ensuring stability and consistent improvement when selecting actions (such as generating accurate hints). By combining PPO with advantage estimation, the method of the present invention can dynamically optimize medical image segmentation in various scenarios.

[0062] In the actor-critic framework, an actor network and a critic network are included. The actor network represents the policy as the categorical distribution of actions given the state s t and is defined as:

[0063] π θ (a t |s t) = Categorical(logits = f θ (s t ))

[0064] Here, π θ represents the probability distribution of the policy π selecting the action a t at the given state s t . Here, θ are the parameters of the policy. f θ represents the function (here the actor network) executed under the parameter θ, and the policy f θ (s t ) represents the result calculated by the actor network with the input s t , which is called logits here, and the action a t is sampled from this distribution, denoted as:

[0065] a t ~ π θ (a t | s t )

[0066] And its log probability is calculated as:

[0067] log π θ (a t | s t )

[0068] Specifically, this step may include the following sub - steps:

[0069] Step 1.1: Divide the image sample x into small blocks, and define the action space with the center points of the small blocks as candidate actions;

[0070] Step 1.2: Initialize the prompt set P 0 , where the prompt set P 0 contains the positive sample a p selected from the real sample y and the negative sample a n ;

[0071] Step 1.3: Use the image encoder to calculate the embedded representation E I (x);

[0072] Step 1.4: The mask decoder calculates the initial mask M I based on the embedded representation E 0 (x) and the prompt set P 0 ;

[0073] Step 1.5: Multiply the embedded representation E I (x) by the initial mask M 0 to obtain the initial state s0

[0074] Step 1.6: Obtain the state s of the image sample x at the t-th time step t , where t is a natural number;

[0075] Step 1.7: The actor network selects an action a in the action space according to the current policy π θ and the state s t ; t ;

[0076] Step 1.8: Add the action a t to the prompt set P t-1 to obtain the prompt set P t ;

[0077] Step 1.9: The mask decoder calculates the mask M based on the prompt set P t and the embedding representation E I (x); t ;

[0078] Step 1.10: Use the mask M t and the embedding representation E I (x) to obtain the state s at the (t + 1)-th time step t+1 ;

[0079] Step 1.11: Obtain the reward r according to the mask M t and the real sample y t+1 .

[0080] Step 2: Perform backpropagation according to the state s t and the reward r t to obtain the trained agent

[0081] The critic network estimates the value function V t of the given state s φ (s t ), which is calculated through a convolutional feature extractor and a linear output layer. Its definition is:

[0082] V φ (s t ) = Linear(CNN(s t ))

[0083] Each image serves as an independent environment and contains a pair of positive and negative prompt points at initialization. The PPO algorithm optimizes the actor-critic network in E rounds of training, where the agent interacts with the environment for T time steps in each round. After the interaction ends, the collected observations, actions, log probabilities, and returns are used to update the network in K rounds of optimization.

[0084] This strategy is optimized by maximizing the surrogate objective of the clipping:

[0085]

[0086] where r t (θ) is the probability ratio between the old and new strategies, and E t represents the expectation at time step t, and the calculation formula is:

[0087] r t (θ) = exp(log π θ (a t |s t ) - log π θold (a t |s t ))

[0088] And is the advantage, and the calculation formula is:

[0089]

[0090] where R t represents the discounted cumulative return from time step t to the last time step:

[0091]

[0092] γ k-t is the discount factor at round k, indicating the contribution of the reward r k at the k-th round to the return at the t-th round. The farther the round is from the current round t, the lower the weight.

[0093] clip(r t (θ), 1 - ∈, 1 + ∈): This is the limit on the policy ratio, ensuring that its change does not exceed the range [1 - ∈, 1 + ∈]. ∈ is a very small positive number used to control the magnitude of the policy update.

[0094] The value function is optimized by the mean squared error loss:

[0095] L VF (φ) = E t [(V φ (s t ) - R t ) 2

[0096] where E t represents the expected value at time step t, which is over all possible states s t and the corresponding returns R t ​Calculate the average of the mean squared error above.

[0097] The total loss function combines the actor loss, value loss, and an entropy regularization term:

[0098] L total = L PPO (θ) - c 1 L VF (φ) + c 2 H(π θ )

[0099] where H(π θ ) is the policy entropy, which encourages the exploration of the model by increasing the entropy of the policy, and c 1 and c 2 are weight coefficients used to balance the influence of the value function loss and policy entropy in the total loss. c 1 controls the weight of the value function loss, and c 2 controls the weight of the policy entropy.

[0100] In each training epoch, the Actor network updates its parameters by backpropagating the gradient of L PPO (θ), while the Critic network uses L VF (φ) to update its parameters. This iterative optimization ensures stable policy updates while balancing exploration and exploitation.

[0101] Step 3: Based on the SAM backbone network and the trained agent, obtain the segmentation results of the test images.

[0102] After training, input the test image x into the agent (PPO network). The agent will select the best action (hint point) from the total action space (point set) multiple times to obtain a set of hint points, combine them, and finally give the content of the hint points. Input the hint point content and the original image into the SAM network, and the SAM network will obtain the final mask, where the hint encoder and mask decoder are implicitly used inside the SAM network.

[0103] Next, a specific experiment is used to illustrate the medical image segmentation method based on deep reinforcement learning assistance provided by the present invention.

[0104] Table 1 shows the comparison of the performance of the Med - SA framework and PPO proposed in the present invention with some current advanced methods on the ISIC2018 dataset.

[0105] Methods Iou Dice U-Net 77.85% 87.55% UNet++ 80.06% 87.93% TransFuse 80.92% 89.46% SAM 58.79% 74.04% MedSAM 70.45% 82.67% Med-SA 83.43% 90.16% The present invention 85.08% 91.31%

[0106] Table 1

[0107] Med-SA+PPO proposed by the present invention performs best among all evaluation methods on the ISIC2018 dataset, setting a new benchmark for medical image segmentation. Its IoU reaches 85.08%, and the Dice coefficient is 91.31%, outperforming task-specific methods and SAM-based methods. Compared with Med-SA, after adding PPO, the IoU is increased by 1.65%, and the Dice is increased by 1.15%, demonstrating the effectiveness of the reinforcement learning strategy in further optimizing the segmentation performance. Compared with the task-specific baseline method (such as TransFuse, whose IoU is 80.92% and Dice is 89.46%), Med-SA+PPO improves by 4.16% in IoU and 1.85 in Dice. Similarly, compared with the SAM-based method (such as MedSAM, whose IoU is 70.45% and Dice is 82.67%), the method of the present invention improves by 14.63% and 8.64% in IoU and Dice respectively. These results highlight the superiority of the method of the present invention in various segmentation frameworks.

[0108] Integrating proximal policy optimization (PPO) into the Med-SA framework plays a key role in these performance improvements. As a reinforcement learning strategy, PPO dynamically adjusts the model parameters to optimize long-term goals, which is particularly significant in complex segmentation tasks. This strategy enhances the model's ability to refine segmentation boundaries, especially when dealing with fuzzy lesion boundaries or small targets. The improvement in IoU and Dice scores emphasizes the value of PPO in enhancing the model's generalization ability and boundary sensitivity, enabling Med-SA+PPO to handle complex segmentation tasks more effectively than other methods.

[0109] Compared with task-specific methods (such as U-Net and UNet++), Med-SA+PPO takes advantage of SAM while addressing their limitations. Although U-Net and UNet++ are very suitable for medical image segmentation, their relatively simple architectures have difficulties in multi-scale feature extraction and boundary accuracy. Even advanced task-specific methods like TransFuse, although adopting attention mechanisms and feature fusion, are still surpassed by Med-SA+PPO due to the lack of a dynamic optimization strategy similar to reinforcement learning. This makes the method of the present invention more robust in dealing with diverse lesion morphologies and edge complexities.

[0110] The performance of Med-SA+PPO also outperforms other SAM-based methods. For example, MedSAM adapts the general SAM framework for medical applications but fails to match Med-SA+PPO. The method of the present invention is based on Med-SA, which has already demonstrated strong medical segmentation capabilities, and further improves the segmentation accuracy by adding PPO. In this way, Med-SA+PPO better handles data imbalance, fine boundary optimization, and better overall segmentation performance, making it stand out among other SAM-based adaptation methods.

[0111] In summary, Med-SA+PPO has established the most effective solution for the medical image segmentation task on the ISIC2018 dataset. By combining the efficient architecture of the SAM framework with the dynamic optimization ability of PPO, it achieves significant segmentation accuracy and robustness. This innovation bridges the gap between traditional task-specific methods and SAM-based frameworks, providing a scalable and adaptable approach for complex medical segmentation tasks. These results not only highlight the effectiveness of the method of the present invention but also demonstrate the potential of reinforcement learning strategies in promoting medical image analysis.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A medical image segmentation method based on deep reinforcement learning, characterized in that: The method comprises: Based on the interaction between the SAM backbone network and the reinforcement learning agent, the state s of the image sample x at each time step is obtained. t and reward r t ; Wherein, the SAM backbone network includes: an image encoder, a prompt encoder and a mask decoder, and the intelligent agent includes: an actor network and a critic network; According to the state s t and the reward r t Perform back propagation to obtain the trained agent; Based on the SAM backbone network and the trained agent, the segmentation results of the test image are obtained.

2. The method according to claim 1, characterized in that Based on the interaction between the SAM backbone network and the reinforcement learning agent, the state s of the image sample x at each time step is obtained. t and reward r t ,include: Divide the image sample x into small blocks and use the center points of the small blocks as candidate actions to define the action space; Initialize prompt set P0, which contains positive samples a selected from real samples y p and negative sample a n ; Use the image encoder to calculate the embedded representation E of the image sample x I (x); The mask decoder is based on the embedded representation E I (x) and prompt set P0 to calculate the initial mask M0; The embedding representation E I (x) is multiplied by the initial mask M0 to obtain the initial state s0 Get the state s of image sample x at the tth time step t , t is a natural number; The actor network follows the current strategy π θ and the state s t Select an action a in the action space t ; Action a t Add to prompt set P t-1 In the prompt set P, we get t ; The mask decoder is based on the prompt set P t and embedding representation E I (x) to calculate the mask M t ; Use mask M t and embedding representation E I (x) to get the state s of the t+1th time step t+1 ; According to the mask M t And the real sample y, get reward r t+1 .

3. The method according to claim 2, characterized in that The actor network follows the current strategy π θ and the state s t Select an action a in the action space t ,include: The state s t The input strategy is π θ The actor network is used to obtain distribution logits, wherein the distribution logits is a subset of actions in the action space; Sample action a from the distribution logits t .

4. The method according to claim 1, characterized in that: According to the state s t and the reward r t Perform back propagation to obtain the trained agent, including: The critic network estimates the state s t The value of V φ (s t ), where V φ represents the critic network with parameter φ; According to the value V φ (s t ) and the reward r t , calculate the advantage at the tth time step , and obtain the probability ratio r between the new and old strategies of the actor network t (θ); Based on the advantages and the probability ratio r t (θ), the loss L of the generated actor network PPO (θ), and according to the value V φ (s t ) and the reward r t , and get the loss L of the critic network VF (φ); Based on the loss L PPO (θ), the loss L VF (φ) and policy entropy loss H(π θ ), and the total loss L total ; Wherein, the strategy entropy loss H(π θ ) is used to encourage the agent to explore by increasing the entropy of the policy; According to the total loss L total Backpropagate the agent to update the agent's parameters.

5. The method according to claim 4, characterized in that According to the value V φ (s t ) and the reward r t , calculate the advantage at the tth time step ,include: Calculate the discounted cumulative return R from time step t to the last time step t ; According to the discounted cumulative return R t and the value V φ (s t ), and get the advantage of the tth time step .

6. The method according to claim 5, characterized in that The loss L VF (φ)=E t [(V φ (s t )-R t ) 2 ].

7. The method according to claim 4, characterized in that The probability ratio r between the new and old strategies t (θ) = exp(logπ θ (a t |s t )-logπ θold (a t |s t )); where π θold For the old strategy.

8. The method according to claim 4, characterized in that The loss L PPO (θ) = E t [min(r t (θ)A k ,clip(r t (θ),1-∈,1+∈)A k )]; where clip represents the clipping function and ∈ represents the hyperparameter threshold.

9. The method according to claim 1, characterized in that: Based on the SAM backbone network and the trained agent, the segmentation results of the test image are obtained, including: Input the test image to the trained agent multiple times, and obtain a set of cue points based on the best action selected by the trained agent each time; The cue point set and the test image are input into a SAM backbone network, and a mask output by the SAM backbone network is used as a segmentation result of the test image.

10. A medical image segmentation system based on deep reinforcement learning, characterized in that: The system comprises: The training module is used to obtain the state s of the image sample x at each time step based on the interaction between the SAM backbone network and the reinforcement learning agent. t and reward r t ; Wherein, the SAM backbone network includes: an image encoder, a prompt encoder and a mask decoder, and the intelligent agent includes: an actor network and a critic network; according to the state s t and the reward r t Perform back propagation to obtain the trained agent; The testing module is used to obtain the segmentation results of the test image based on the SAM backbone network and the trained agent.