A medical image anomaly detection method and system based on deep reinforcement learning
Patent Information
- Application Number
- CN202611045516.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-10-09
AI Technical Summary
[0008]本发明为了解决现有医学图像异常检测标注数据依赖强、泛化能力差、解释性不足、实时自适应能力弱的技术问题,提出了一种基于深度强化学习的医学图像异常检测方法及系统,以实现高精度、高效率、可解释、自适应的医学图像异常检测
[0044]1、在核心检测性能上,本发明通过将多模态医学影像数据预处理后输入深度强化学习智能体完成异常检测任务,能显著提升检测精度与召回率,有效降低假阳性、假阴性结果出现的概率,确保病灶识别既精准又全面。
Smart Images

Figure CN122887828A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical image processing and artificial intelligence technology, specifically to a method and system for detecting anomalies in medical images based on deep reinforcement learning. Background Technology
[0002] With the rapid development of medical imaging technology, medical images have become an indispensable basis for clinical disease diagnosis. Traditionally, the detection of abnormalities in medical images has relied heavily on manual interpretation by doctors, a method inherently flawed by its subjectivity, low efficiency, and susceptibility to missed diagnoses. In recent years, automated detection methods based on deep learning have gradually emerged, with convolutional neural networks (CNNs) achieving some success in tasks such as tumor detection and lesion segmentation. However, existing deep learning methods still face the following prominent problems:
[0003] I. Heavy reliance on labeled data: Supervised learning paradigms require a large number of accurately labeled training samples, but labeling is costly and high-quality labeled data is difficult to obtain in actual clinical practice;
[0004] 2. Limited generalization ability: The model has poor adaptability to new anomaly types outside the training set or data collected by different imaging devices;
[0005] Third, insufficient interpretability: Deep learning models are often regarded as "black boxes" and it is difficult to provide clear and credible decision-making basis, which restricts their reliable application in clinical practice;
[0006] Fourth, weak real-time performance and adaptive capabilities: Traditional models cannot dynamically adjust detection strategies based on new feedback during inference, and lack online learning and interactive optimization capabilities.
[0007] Deep reinforcement learning (DRL) organically integrates the feature representation capabilities of deep learning with the sequential decision-making mechanism of reinforcement learning, making it particularly suitable for intelligent decision-making problems in dynamic environments. Although DRL has made breakthrough progress in fields such as game control and robot navigation, its exploration in medical image anomaly detection is still in its early stages, and a complete solution that combines systematicity, interpretability, and high adaptability has not yet been formed. Summary of the Invention
[0008] To address the technical problems of existing medical image anomaly detection methods, such as strong reliance on labeled data, poor generalization ability, insufficient interpretability, and weak real-time adaptive ability, this invention proposes a medical image anomaly detection method and system based on deep reinforcement learning, so as to achieve high-precision, high-efficiency, interpretable, and adaptive medical image anomaly detection.
[0009] To address the aforementioned technical problems, the present invention adopts the following technical solution: a medical image anomaly detection method based on deep reinforcement learning, comprising the following steps:
[0010] Step S1: Acquire multimodal medical image data and preprocess it to output standardized image tensors;
[0011] Step S2: Input the standardized image tensor into the preset deep reinforcement learning agent to complete the anomaly detection task and output the detection result;
[0012] Step S3: After confirming or correcting the detection results using the preset result output and interpretation unit, generate labeled feedback medical image data and store it in the model update and storage unit.
[0013] Step S4: Update the deep reinforcement learning agent based on labeled feedback medical image data.
[0014] Furthermore, the preprocessing process includes performing size normalization, grayscale standardization, noise filtering, and data augmentation operations sequentially on the acquired raw medical image data, ultimately generating a unified standardized image tensor.
[0015] Furthermore, the pre-defined deep reinforcement learning agent has a built-in state encoding module, which is sequentially connected to a feature extraction network, a sequence modeling module, and a decision output module along the data flow direction.
[0016] Furthermore, step S2 specifically includes the following steps:
[0017] Step S21: Input multiple standardized image tensor codes into the state coding module. Through standardization processing and cross-modal feature alignment mechanism, image tensors of different modalities are mapped to the same feature space to form a unified state representation.
[0018] Step S22: Based on the unified state representation, a multi-scale feature extraction network based on a convolutional neural network is used to perform hierarchical feature extraction on the image tensor to obtain spatial features under multiple receptive fields.
[0019] Step S23: Input the extracted spatial features into the sequence modeling module, use a long short-term memory network to model the temporal image sequence or historical state sequence, extract temporal dependent features, realize the fusion expression of spatial and temporal information, and obtain the fused spatiotemporal state features.
[0020] Step S24: Input the fused spatiotemporal state features into the decision output module. The decision output module uses the Dueling DQN network to decompose and calculate the preset action value function and output the current optimal detection action.
[0021] Step S25: Based on the current optimal detection action output, update the medical image state detected by the deep reinforcement learning agent, and generate a new medical image state and detection result; at the same time, calculate the reward value based on the preset target reward function, and input the reward value as a feedback signal in the reinforcement learning training process to the decision output module to guide the deep reinforcement learning agent to optimize the detection strategy.
[0022] Furthermore, the action value function The expression is:
[0023] ;
[0024] In the formula, This indicates the detection value of the current medical image state s itself;
[0025] This represents the output value of the advantage function when performing a certain action a compared to other actions in a medical image state s.
[0026] Represents the set of candidate actions;
[0027] Indicates the number of elements in the action set;
[0028] Indicates the state Next, for candidate actions The output value of the advantage function.
[0029] Furthermore, step S3 specifically involves presenting the test results generated in step S2 to the clinical terminal through a preset result output and interpretation unit. The clinician then confirms or corrects the test results through the clinical terminal, generating labeled feedback medical imaging data.
[0030] Furthermore, the update of the deep reinforcement learning agent includes two stages: an offline pre-training stage and an online adaptive update stage. These two stages are based on the model parameters of the deep reinforcement learning agent and feedback medical image data with clinical annotations, enabling dynamic interaction and forming a continuously evolving closed-loop system.
[0031] Furthermore, step S4 includes the following steps:
[0032] Step S41, Offline pre-training stage: Obtain the historical labeled feedback medical image dataset stored in the model update and storage unit, and input the labeled medical image dataset into the deep reinforcement learning agent after standardization and preprocessing to generate labeled medical image state sequence, action sequence and reward value sequence.
[0033] Step S42, Online Adaptive Update Stage: The newly input medical image data is preprocessed and transmitted to the updated deep reinforcement learning agent to perform anomaly detection inference process, generate anomaly region location, category label and confidence information, and output the detection results to the clinical terminal for doctors to review; during the doctor's review process, the doctor's confirmation or correction operation on the detection results is recorded and the operation is converted into labeled feedback medical image data;
[0034] Step S43, Adaptive update trigger determination: Analyze the accumulated labeled feedback medical image dataset based on preset update conditions. When any update condition is met, the deep reinforcement learning agent update process is automatically started.
[0035] Step S44: After the update conditions are met and the deep reinforcement learning agent is triggered to update, the model update is fused with the historical interaction data stored in the experience pool of the storage unit and the newly added clinical labeled feedback medical image data to construct an incremental training dataset, and the deep reinforcement learning agent update operation is executed in an isolated training environment.
[0036] Step S45: Perform a comprehensive performance verification and evaluation on the updated deep reinforcement learning agent. When the performance of the deep reinforcement learning agent meets the preset deployment criteria, replace the currently deployed deep reinforcement learning agent with the updated deep reinforcement learning agent and update the model parameters simultaneously.
[0037] A deep reinforcement learning-based medical image anomaly detection system, which implements the aforementioned deep reinforcement learning-based medical image anomaly detection method, comprises:
[0038] The data acquisition module is used to acquire multimodal medical image data;
[0039] The preprocessing module is used to preprocess the acquired multimodal medical image data and output standardized image tensors.
[0040] A deep reinforcement learning agent is used to receive standardized image tensors to complete anomaly detection tasks and output detection results;
[0041] The results output and interpretation unit is used to confirm or correct the test results and generate labeled feedback medical image data.
[0042] The model update and storage unit is used to receive labeled feedback medical image data and update the deep reinforcement learning agent based on the labeled feedback medical image data.
[0043] The advantages of this invention over the prior art are as follows:
[0044] 1. In terms of core detection performance, this invention significantly improves detection accuracy and recall by inputting preprocessed multimodal medical image data into a deep reinforcement learning agent to complete the anomaly detection task, effectively reducing the probability of false positive and false negative results, and ensuring that lesion identification is both accurate and comprehensive.
[0045] 2. In terms of data adaptation capability, this invention integrates the historical interaction data stored in the experience pool of the model update and storage unit with the newly added clinical labeled feedback medical image data to construct an incremental training dataset, which greatly reduces the dependence on labeled data and can still maintain stable detection performance in data-scarce scenarios.
[0046] 3. In terms of clinical auxiliary value, during the online adaptive update stage, the newly input medical image data is preprocessed and transmitted to the updated deep reinforcement learning agent to perform anomaly detection reasoning process, generate abnormal area location, category label and confidence information, and output the detection results to the clinical terminal for doctors to review. By outputting interpretable detection results, the basis and logic of lesion identification are clearly presented, providing strong support for doctors' decision-making.
[0047] 4. In dynamic application scenarios, this invention analyzes the accumulated labeled feedback medical image dataset based on preset update conditions. When any update condition is met, the deep reinforcement learning agent update process is automatically started, which can realize real-time detection and adaptive update, quickly respond to the dynamic changes in the clinical environment, and ensure that the technology continues to play an efficient role in the clinical process. At the same time, it has a strong generalization ability and can be compatible with image data from different sources across devices and diseases, eliminating the detection limitations caused by differences in devices and diseases. Attached Figure Description
[0048] The present invention will be further described below with reference to the accompanying drawings:
[0049] Figure 1 This is a schematic diagram of the system structure of one embodiment of the present invention;
[0050] Figure 2 This is a diagram illustrating the internal working mechanism of a deep reinforcement learning agent according to an embodiment of the present invention.
[0051] Figure 3 This is a flowchart of a two-stage training and adaptive update process according to an embodiment of the present invention;
[0052] Figure 4 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0053] like Figures 1 to 4 As shown, this invention provides a medical image anomaly detection method based on deep reinforcement learning, comprising the following steps:
[0054] Step S1: Acquire multimodal medical image data and preprocess it to output standardized image tensors;
[0055] Specifically, the preprocessing process includes performing size normalization, grayscale standardization, noise filtering, and data augmentation operations sequentially on the acquired raw medical image data, ultimately generating a unified standardized image tensor.
[0056] The preprocessing process is executed in a pre-defined preprocessing module.
[0057] Step S2: Input the standardized image tensor into the preset deep reinforcement learning agent to complete the anomaly detection task and output the detection results.
[0058] The pre-defined deep reinforcement learning agent has a built-in state encoding module, which is connected sequentially to a feature extraction network, a sequence modeling module, and a decision output module along the data flow direction.
[0059] Step S2 specifically includes the following steps:
[0060] Step S21: Input multiple standardized image tensor codes into the state coding module. Through standardization processing and cross-modal feature alignment mechanism, image tensors of different modalities are mapped to the same feature space to form a unified state representation, so as to enhance the adaptability of deep reinforcement learning agents to cross-device and cross-modal data distribution differences.
[0061] Step S22: Based on the unified state representation, a multi-scale feature extraction network based on convolutional neural network (CNN) is used to perform hierarchical feature extraction on image tensors to obtain spatial features under multiple receptive fields, so as to enhance the discriminative representation ability of small lesions and complex backgrounds.
[0062] Step S23: Input the extracted spatial features into the sequence modeling module, use a long short-term memory network (LSTM) to model the time-series image sequence or historical state sequence, extract time-dependent features, realize the fusion expression of spatial and temporal information, and obtain the fused spatiotemporal state features to enhance the ability to identify dynamic lesion changes.
[0063] Step S24: Input the fused spatiotemporal state features into the decision output module. The decision output module uses the Dueling DQN network to decompose and calculate the preset action value function, and outputs the current optimal detection action using the action value function.
[0064] Specifically, the action value function Decomposed into state value function V(s) and action advantage function ;
[0065] The expression for the state value function V(s) is:
[0066] ;
[0067] In the formula, The detection value of the current medical image state s itself is calculated by the value stream network in Dueling DQN; This is the state feature vector extracted by a convolutional neural network / long short-term memory network (CNN / LSTM). These are the parameters of the value stream network, used to evaluate the value of the current medical image state s itself (i.e., whether the current medical image state as a whole is worth continuing to detect).
[0068] Action advantage function The expression is:
[0069] ;
[0070] In the formula, The output value of the dominance function represents the advantage function output value of performing a certain action a compared to other actions under the medical image state s. It is calculated by the dominance flow network and is used to determine "which detection action is better".
[0071] These are the parameters of the dominant flow network, used to measure the advantage of performing a certain action a compared to other actions in a medical image state s; action a may include abnormal region localization, detection window movement, detection termination, etc.
[0072] Action value function The expression is:
[0073] ;
[0074] In the formula, It is obtained by fusing the state value function and the action advantage function; where, Represents the set of candidate actions; Indicates the number of elements in the action set; Indicates the state Next, for candidate actions The advantage function output value; the action value function of a deep reinforcement learning agent. The action with the largest output is taken as the current optimal detection action.
[0075] Step S25: Based on the current optimal detection action output by the action advantage function, update the medical image state detected by the deep reinforcement learning agent and generate a new medical image state and detection result; at the same time, calculate the reward value Rt based on the preset target reward function. The target reward function includes at least three indicators: detection accuracy, model confidence and clinical consistency, so as to achieve synergistic optimization of detection performance and clinical rationality.
[0076] The expression for the target reward function is:
[0077] ;
[0078] In the formula, This represents the reward for detection accuracy (such as IoU or Dice coefficient), used to measure the degree of overlap between the predicted results and the actual annotations, and is calculated based on IoU or Dice coefficient; This represents the model confidence reward, which is used to weight the detection results based on the predicted probabilities output by the model. A higher reward is given when the detection result is correct and has a high confidence level. This indicates a clinical consistency reward, used to assess the degree of matching between test results and pre-set clinical diagnostic rules or physician annotation logic; This indicates a false positive penalty, which applies a penalty when a false positive is detected to suppress invalid detections; , , This is a weighting hyperparameter used to adjust the contribution ratio of each reward item to the total reward, and it satisfies... The reward value is input to the decision output module as a feedback signal during the reinforcement learning training process, which guides the deep reinforcement learning agent to optimize the detection strategy, thereby achieving a multi-objective balance between detection accuracy, confidence and clinical consistency.
[0079] Step S3: After confirming or correcting the detection results using the preset result output and interpretation unit, the system generates labeled feedback medical image data and stores it in the model update and storage unit.
[0080] Specifically, the detection results generated in step S2 are presented to the clinical terminal through a preset result output and interpretation unit. Clinicians confirm or correct the detection results through the clinical terminal, generating labeled feedback medical image data. Simultaneously, the state transition information during the detection process (current medical image state, action, target reward function, next medical image state) and the calculated reward value Rt are stored in the experience replay pool of the model update and storage unit as training samples for subsequent deep reinforcement learning agent updates.
[0081] Step S4: Update the deep reinforcement learning agent based on labeled feedback medical image data.
[0082] The update of a deep reinforcement learning agent consists of two stages: an offline pre-training stage and an online adaptive update stage. These two stages are based on the model parameters of the deep reinforcement learning agent and feedback medical image data with clinical annotations, enabling dynamic interaction and forming a continuously evolving closed-loop system.
[0083] More specifically, step S4 includes the following steps:
[0084] Step S41, Offline pre-training stage: Obtain the historical labeled feedback medical image dataset stored in the model update and storage unit, and input the labeled medical image dataset into the deep reinforcement learning agent after standardization and preprocessing to generate labeled medical image state sequence, action sequence and reward value sequence.
[0085] The actions include at least abnormal region location, candidate region movement, and detection termination.
[0086] The labeled reward value is calculated based on a real-labeled feedback medical image dataset and is used to guide the agent in learning the initial detection strategy.
[0087] During the offline pre-training phase, historical interaction data (state-action-reward) is stored in the experience pool of the model update and storage unit. Training batches are constructed through a random sampling mechanism, and the decision output module is iteratively updated based on the Dueling DQN algorithm until the deep reinforcement learning agent meets the preset convergence conditions (such as loss function stability or detection accuracy reaching the target). Finally, the model parameters of the deep reinforcement learning agent are output, and the updated deep reinforcement learning agent is obtained.
[0088] Step S42, Online Adaptive Update Stage: The newly input medical image data is preprocessed and then transmitted to the updated deep reinforcement learning agent to perform the anomaly detection inference process, generate the location, category label and confidence information of the abnormal area, and output the detection results to the clinical terminal for doctors to review;
[0089] During the doctor's review process, the doctor's confirmation or correction operations on the test results are recorded. The operations are converted into labeled feedback medical image data and stored together with the original test results and corresponding medical image status information for subsequent deep reinforcement learning agent optimization.
[0090] Step S43, Adaptive update trigger determination: Analyze the accumulated labeled feedback medical image dataset based on preset update conditions. When any update condition is met, the deep reinforcement learning agent update process is automatically started.
[0091] The update conditions include:
[0092] (1) The number of labeled feedback medical image data samples reaches the preset threshold;
[0093] (2) The detection performance of deep reinforcement learning agents on specific disease types or equipment data showed a significant decline;
[0094] (3) The application scenario of detection has changed significantly (such as the access of new equipment or the introduction of new diseases).
[0095] Step S44: After the update conditions are met and the deep reinforcement learning agent is updated, the model update is fused with the historical interaction data stored in the experience pool of the storage unit and the newly added clinical labeled feedback medical image data to construct an incremental training dataset. The deep reinforcement learning agent update operation is then executed in an isolated training environment. Update methods include:
[0096] (1) Online fine-tuning based on new data: The model parameters are updated rapidly and incrementally using newly added clinical feedback data, so that the model can adapt to the changes in data distribution in the current clinical scenario in a timely manner;
[0097] (2) Transfer learning for new diseases or new tasks: Transfer the general feature knowledge learned by the existing deep reinforcement learning agent on the original detection task to new diseases or new detection tasks, so as to reduce the cost of repeated training and improve the generalization ability of the deep reinforcement learning agent in new scenarios.
[0098] (3) By adopting federated learning technology, the original medical image data of each medical institution can be trained by multi-center joint deep reinforcement learning agent without leaving the local area, thereby effectively protecting data privacy and security.
[0099] Step S45: Perform a comprehensive performance verification and evaluation on the updated deep reinforcement learning agent. Evaluation indicators include detection accuracy, recall rate, and clinical consistency. When the performance of the deep reinforcement learning agent meets the preset deployment standards, replace the currently deployed deep reinforcement learning agent with the updated deep reinforcement learning agent and update the model parameters simultaneously.
[0100] By organically linking offline pre-training, online adaptive updates and feedback collection, triggered updates, and validation deployment, a closed-loop optimization mechanism is constructed, encompassing "offline training—online inference—feedback collection—deep reinforcement learning agent update." This mechanism enables the medical image anomaly detection system to continuously adapt to different data distributions in practical clinical applications, gradually improving the accuracy and robustness of anomaly detection.
[0101] The training strategy used in this embodiment is as follows:
[0102] A strategy combining curriculum learning and adversarial training is employed. Curriculum learning progressively constructs training tasks from simple to complex, specifically: first, using simple medical image data with obvious anomalies for initial training of the deep reinforcement learning agent; then, gradually introducing complex backgrounds, multiple lesions, and low-contrast anomaly data to improve the agent's ability to identify complex anomalies. During training, an adversarial example generation mechanism is introduced, generating adversarial examples by perturbing the input images. These adversarial examples, along with the original examples, are input into the deep reinforcement learning agent for training, enhancing its robustness to noise interference and distribution shifts. Simultaneously, the training process supports a federated learning framework, enabling collaborative training across multiple medical institutions through model parameter sharing. Each participating node independently updates its model parameters on local data, and a secure aggregation mechanism fuses the model parameters, thereby optimizing the deep reinforcement learning agent without sharing the original data, thus protecting multi-center data privacy.
[0103] During the anomaly detection inference process, the medical image data to be detected is input into the trained deep reinforcement learning agent. The deep reinforcement learning agent outputs the optimal action sequence based on the current state of the medical image and generates anomaly detection results. The results include the spatial location information of the abnormal region, which is represented by a bounding box or segmentation mask. At the same time, the class label and prediction confidence of the corresponding abnormal region are output. Furthermore, the result output and interpretation unit generates an interpretable heatmap based on the internal feature response and attention weight of the deep reinforcement learning agent. This heatmap is used to characterize the key regions in the image that contribute significantly to the detection results and extracts the corresponding key feature description information, thereby providing a visual interpretation basis for the detection results and improving the interpretability and credibility of the deep reinforcement learning agent output in clinical applications.
[0104] This invention collects feedback from doctors regarding confirmation or correction of detection results. When preset update conditions are met, it incrementally updates the deep reinforcement learning agent based on online fine-tuning, transfer learning, or federated learning mechanisms. This enables the deep reinforcement learning agent to continuously adapt to new diseases, new equipment, and new data distributions, achieving long-term self-evolution capability of the medical image anomaly detection system. Through an experience replay mechanism, historical medical image state sequences, action sequences, and reward value sequences are randomly sampled from the experience pool of the model update and storage unit. The Dueling DQN network parameters are then iteratively updated in batches, forming a stable reinforcement learning training closed loop.
[0105] Example
[0106] Taking CT scans of lung nodules as an example, the medical image anomaly detection method based on deep reinforcement learning includes the following steps:
[0107] Input CT sequence → Intelligent agent scans frame by frame and locates suspicious areas → Combines time series information for comprehensive judgment → Outputs nodule location, size, malignancy probability and interpretation map → Updates model based on doctor feedback.
[0108] Step S1: Acquire CT sequences and perform preprocessing to output standardized image tensors;
[0109] Step S2: Input the standardized image tensor into the preset deep reinforcement learning agent to complete the anomaly detection task and output the detection results, including nodule location, size, malignancy probability and interpretation map;
[0110] Step S3: After confirming or correcting the detection results using the preset result output and interpretation unit, generate labeled feedback medical image data and store it in the model update and storage unit.
[0111] Step S4: Update the deep reinforcement learning agent based on labeled feedback medical image data.
[0112] A deep reinforcement learning-based medical image anomaly detection system, which implements the aforementioned deep reinforcement learning-based medical image anomaly detection method, comprises:
[0113] The data acquisition module is used to acquire multimodal medical image data;
[0114] The preprocessing module is used to preprocess the acquired multimodal medical image data and output standardized image tensors.
[0115] A deep reinforcement learning agent is used to receive standardized image tensors to complete anomaly detection tasks and output detection results;
[0116] The results output and interpretation unit is used to confirm or correct the test results and generate labeled feedback medical image data.
[0117] The model update and storage unit is used to receive labeled feedback medical image data and update the deep reinforcement learning agent based on the labeled feedback medical image data.
[0118] This invention models the medical image anomaly detection process as a state, action, and reward-driven sequential decision-making process using a deep reinforcement learning agent. A multi-objective reward function is employed to jointly optimize detection accuracy, confidence, and clinical consistency. Simultaneously, a combination of curriculum learning and adversarial training strategies enhances the robustness of the deep reinforcement learning agent to complex anomalies and data perturbations. A federated learning mechanism ensures privacy protection and collaborative model optimization in multi-center data environments. Furthermore, an adaptive update mechanism based on physician feedback is introduced, enabling the deep reinforcement learning agent to continuously optimize parameters during practical applications, thereby improving its generalization ability to new diseases and data from different devices. Additionally, an attention mechanism generates interpretable detection results, enhancing the understandability and clinical credibility of the deep reinforcement learning agent's output. Therefore, this invention improves the system's adaptability and interpretability while maintaining detection accuracy, making it suitable for automated anomaly detection in multimodal medical imaging scenarios, thus improving clinical diagnostic efficiency.
[0119] Regarding the specific structure of this invention, it should be noted that the connection relationships between the various component modules used in this invention are definite and achievable. Except as specifically described in the embodiments, their specific connection relationships can bring about corresponding technical effects and solve the technical problems proposed by this invention without relying on the execution of corresponding software programs. The models of the components, modules, and specific components appearing in this invention, the connection methods between them, and the conventional usage methods and expected technical effects brought about by the above technical features, unless specifically described, are all publicly disclosed content in patents, journal articles, technical manuals, technical dictionaries, and textbooks that can be obtained by those skilled in the art before the application date, or belong to conventional technology, common knowledge, and other existing technologies in this field. There is no need to elaborate, which makes the technical solution provided in this case clear, complete, and achievable, and can reproduce or obtain corresponding physical products based on this technical means.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting anomalies in medical images based on deep reinforcement learning, characterized in that, Includes the following steps: Step S1: Acquire multimodal medical image data and preprocess it to output standardized image tensors; Step S2: Input the standardized image tensor into the preset deep reinforcement learning agent to complete the anomaly detection task and output the detection result; Step S3: After confirming or correcting the detection results using the preset result output and interpretation unit, generate labeled feedback medical image data and store it in the model update and storage unit. Step S4: Update the deep reinforcement learning agent based on labeled feedback medical image data.
2. The medical image anomaly detection method based on deep reinforcement learning according to claim 1, characterized in that, The preprocessing process includes performing size normalization, grayscale standardization, noise filtering, and data augmentation operations on the acquired raw medical image data in sequence, ultimately generating a unified standardized image tensor.
3. The medical image anomaly detection method based on deep reinforcement learning according to claim 1, characterized in that, The pre-defined deep reinforcement learning agent has a built-in state encoding module, which is connected sequentially to a feature extraction network, a sequence modeling module, and a decision output module along the data flow direction.
4. The medical image anomaly detection method based on deep reinforcement learning according to claim 3, characterized in that, Step S2 specifically includes the following steps: Step S21: Input multiple standardized image tensor codes into the state coding module. Through standardization processing and cross-modal feature alignment mechanism, image tensors of different modalities are mapped to the same feature space to form a unified state representation. Step S22: Based on the unified state representation, a multi-scale feature extraction network based on a convolutional neural network is used to perform hierarchical feature extraction on the image tensor to obtain spatial features under multiple receptive fields. Step S23: Input the extracted spatial features into the sequence modeling module, use a long short-term memory network to model the temporal image sequence or historical state sequence, extract temporal dependent features, realize the fusion expression of spatial and temporal information, and obtain the fused spatiotemporal state features. Step S24: Input the fused spatiotemporal state features into the decision output module. The decision output module uses the DuelingDQN network to decompose and calculate the preset action value function and output the current optimal detection action. Step S25: Based on the current optimal detection action output, update the medical image state detected by the deep reinforcement learning agent, and generate a new medical image state and detection result; at the same time, calculate the reward value based on the preset target reward function, and input the reward value as a feedback signal in the reinforcement learning training process to the decision output module to guide the deep reinforcement learning agent to optimize the detection strategy.
5. The medical image anomaly detection method based on deep reinforcement learning according to claim 4, characterized in that, Action value function The expression is: ; In the formula, This indicates the detection value of the current medical image state s itself; This represents the output value of the advantage function when performing a certain action a compared to other actions in a medical image state s. Represents the set of candidate actions; Indicates the number of elements in the action set; Indicates the state Next, for candidate actions The output value of the advantage function.
6. The medical image anomaly detection method based on deep reinforcement learning according to claim 1, characterized in that, Step S3 specifically involves presenting the test results generated in step S2 to the clinical terminal through a preset result output and interpretation unit. The clinician then confirms or corrects the test results through the clinical terminal, generating labeled feedback medical imaging data.
7. The medical image anomaly detection method based on deep reinforcement learning according to claim 1, characterized in that, The update of a deep reinforcement learning agent consists of two stages: an offline pre-training stage and an online adaptive update stage. These two stages are based on the model parameters of the deep reinforcement learning agent and feedback medical image data with clinical annotations, enabling dynamic interaction and forming a continuously evolving closed-loop system.
8. The medical image anomaly detection method based on deep reinforcement learning according to claim 7, characterized in that, Step S4 includes the following steps: Step S41, Offline pre-training stage: Obtain the historical labeled feedback medical image dataset stored in the model update and storage unit, and input the labeled medical image dataset into the deep reinforcement learning agent after standardization and preprocessing to generate labeled medical image state sequence, action sequence and reward value sequence. Step S42, Online Adaptive Update Stage: The newly input medical image data is preprocessed and transmitted to the updated deep reinforcement learning agent to perform anomaly detection inference process, generate anomaly region location, category label and confidence information, and output the detection results to the clinical terminal for doctors to review; during the doctor's review process, the doctor's confirmation or correction operation on the detection results is recorded, and the operation is converted into labeled feedback medical image data; Step S43, Adaptive update trigger determination: Analyze the accumulated labeled feedback medical image dataset based on preset update conditions. When any update condition is met, the deep reinforcement learning agent update process is automatically started. Step S44: After the update conditions are met and the deep reinforcement learning agent is triggered to update, the model update is fused with the historical interaction data stored in the experience pool of the storage unit and the newly added clinical labeled feedback medical image data to construct an incremental training dataset, and the deep reinforcement learning agent update operation is executed in an isolated training environment. Step S45: Perform a comprehensive performance verification and evaluation on the updated deep reinforcement learning agent. When the performance of the deep reinforcement learning agent meets the preset deployment criteria, replace the currently deployed deep reinforcement learning agent with the updated deep reinforcement learning agent and update the model parameters simultaneously.
9. A medical image anomaly detection system based on deep reinforcement learning that implements the medical image anomaly detection method based on deep reinforcement learning described in any one of 1-8 above, characterized in that, include: The data acquisition module is used to acquire multimodal medical image data; The preprocessing module is used to preprocess the acquired multimodal medical image data and output standardized image tensors. A deep reinforcement learning agent is used to receive standardized image tensors to complete anomaly detection tasks and output detection results; The results output and interpretation unit is used to confirm or correct the test results and generate labeled feedback medical image data. The model update and storage unit is used to receive labeled feedback medical image data and update the deep reinforcement learning agent based on the labeled feedback medical image data.