Intelligent agent-driven radiotherapy task decision-making method, computing device and storage medium
Through the agent-driven radiotherapy task decision-making method, the optimal model is dynamically selected and the execution strategy is output, which solves the problems of multimodal image data processing fragmentation and limited performance of a single model, and realizes efficient and accurate image data processing and task execution.
Patent Information
- Application Number
- CN202510302571.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
In the existing radiotherapy technology, multimodal image data processing is fragmented, lacking a unified framework, and a single model has limited performance, difficulty in generalization, and poor adaptability, which makes it difficult to take into account both task efficiency and accuracy.
Adopt the agent-driven radiotherapy task decision-making method, by identifying the current radiotherapy task, obtaining image data, processing data modal information and feature information, dynamically selecting the optimal model and outputting execution strategies, to achieve efficient collaborative processing of multimodal image data and task-model adaptation.
It realizes efficient collaborative processing of multimodal image data, improves image data processing efficiency, ensures the accuracy and stability of tasks at each stage, and enhances the generalization ability and real-time response ability of the system.
Smart Images

Figure CN120220968A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radiotherapy, and particularly relates to an intelligent agent-driven radiotherapy task decision-making method, a computing device, and a storage medium. Background Art
[0002] Radiotherapy (RT) is one of the main methods for current cancer treatment. By precisely controlling the radiation dose, it can kill cancer cells while minimizing damage to normal tissues. With the development of image-guided radiotherapy (IGRT), the application of multi-modal medical image data (such as CT, CBCT, plain films, etc.) has become particularly important in the radiotherapy process. These data run through the entire process from treatment planning to treatment implementation, providing a reliable basis for precise treatment.
[0003] Currently, data processing and analysis in the radiotherapy process usually face the following problems:
[0004] First, the processing of multi-modal image data is fragmented and lacks a unified framework. Image data such as CT, CBCT, and X-ray plain films have different physical characteristics and resolutions, and need to be processed specifically in the application process to achieve the goals of different stage tasks, such as organ segmentation, dose prediction, and positioning verification.
[0005] The processing of data in different modalities (such as CT, CBCT, X-ray plain films, etc.) is usually completed by independent algorithm modules. These modules lack a cooperation mechanism and are difficult to achieve efficient processing and integration in complex scenarios. Especially in tasks that require combining multi-modal information (such as positioning verification or dynamic plan adjustment), the adaptability of existing systems is poor, resulting in an unsmooth workflow.
[0006] Second, the performance of a single model is limited and it is difficult to generalize. Current image data processing tasks often use a single model to complete, such as a deep learning-based segmentation model or a dose prediction model. However, a single model is difficult to perform optimally in all tasks.
[0007] For example, a high-precision segmentation model has a large computational overhead and is not suitable for real-time scenarios during treatment; a low-complexity model is fast, but its performance is poor in high-noise CBCT images. The lack of a dynamic model selection mechanism makes it difficult to balance task efficiency and accuracy.
[0008] Furthermore, different stages of radiotherapy involve multiple tasks, such as treatment plan generation, dose distribution optimization, positioning error correction, etc. Each task may require a specific algorithm model. However, it is difficult for a single model to efficiently complete all tasks.
[0009] Third, it has poor adaptability. Existing technologies usually rely on fixed algorithms or preset parameter configurations and are difficult to dynamically adjust according to the task type and data characteristics. For example, during the treatment implementation process, changes in the patient's anatomical structure (such as tumor shrinkage, organ displacement) may require re-segmentation and dose adjustment, while existing systems often show lag and lack the ability to respond in real time when dealing with these changes.
[0010] In summary, traditional methods usually rely on static models or algorithms, lack flexibility and generalization ability when dealing with complex task scenarios, and are difficult to meet the requirements of modern radiotherapy for intelligence, precision, and personalization. Summary of the Invention
[0011] To solve the above technical problems, the present invention proposes an agent-driven radiotherapy task decision-making method, a computing device, and a storage medium.
[0012] To achieve the above object, the technical solution of the present invention is as follows:
[0013] In the first aspect, the present invention discloses an agent-driven radiotherapy task decision-making method, including:
[0014] Step S1: Identify the current radiotherapy task and obtain image data;
[0015] Step S2: Process the image data to obtain data modality information and data feature information;
[0016] Step S3: Based on the identified current radiotherapy task, obtain the models selected by historical tasks with high similarity to the current radiotherapy task and the performance scores of these models in historical tasks, and form reference information;
[0017] Step S4: Input the current radiotherapy task, data modality information, data feature information, and reference information into the trained agent. The agent can select one or more optimal models from the model library and output the execution strategies of one or more optimal models.
[0018] Based on the above technical solution, the following improvements can be made:
[0019] As a preferred solution, N candidate models are stored in the model library, specifically: M = m1, m2, …, m N};
[0020] If only one model is required to implement the current radiotherapy task, the agent uses the following formula to select an optimal model from the model library;
[0021]
[0022] Where:
[0023] α and β are weight factors;
[0024] P i is the prediction performance score of model m i for the current radiotherapy task;
[0025] C i is the computational cost of model m i ;
[0026] If multiple models are required to collaborate to achieve the current radiotherapy task, the agent selects multiple optimal models from the model library using the following formula;
[0027]
[0028] where:
[0029] A is the call order and combination strategy of multiple models;
[0030] T is the phased process of the radiotherapy task;
[0031] α and β are weight factors;
[0032] P t is the prediction performance score at the t-th stage;
[0033] C t is the computational cost at the t-th stage.
[0034] As a preferred solution, the agent is trained using the reinforcement learning method;
[0035] The state s includes: radiotherapy task, data modality information, data feature information, and reference information;
[0036] The action a includes: selecting one or more models;
[0037] The reward r includes: dynamically generated according to the radiotherapy task;
[0038] The goal of the agent is to maximize the cumulative reward by selecting the optimal policy π(a|s), and its objective function is:
[0039]
[0040] where:
[0041] γ t is the discount factor;
[0042] r t is the immediate reward at the t-th step.
[0043] As a preferred solution, DQN or PPO is used to optimize the agent, which specifically includes the following steps:
[0044] Step A: Initialize the agent policy π(a|s) and the model library;
[0045] Step B: Select an action a according to the state s;
[0046] Step C: After executing the action a, obtain the reward r and the new state s′;
[0047] Step D: Update the policy parameters;
[0048] For DQN, the optimization objective is as follows:
[0049]
[0050] Or,
[0051] For PPO, the optimization objective is as follows:
[0052] L PPO = E t [min(r t (θ)·A t , clip(r t (θ), 1 - ∈, 1 + ∈)·A t )];
[0053] Where:
[0054] A t is the advantage function;
[0055] clip() is the truncation function;
[0056] r t (θ) is the policy ratio.
[0057] As a preferred solution, the agent can adapt to new radiotherapy tasks or imaging data of new modalities through the meta - learning mechanism;
[0058] If there are n tasks in the radiotherapy task set, specifically: T = {T1, T2, …, T n};
[0059] For each radiotherapy task T i , the goal of the agent's optimization is to find the optimal initial parameter θ such that it can quickly adapt to the radiotherapy task after a small number of gradient updates;
[0060] The optimization objective is:
[0061]
[0062] Where:
[0063] is the loss function of the radiotherapy task T i ;
[0064] α is the learning rate for gradient update;
[0065] is the radiotherapy task T i the gradient under the current parameter θ.
[0066] As a preferred solution, the optimization of the agent includes inner-layer optimization and outer-layer optimization;
[0067] Inner-layer optimization, for each radiotherapy task T through the following formula i Use gradient update to adjust the parameters to obtain the updated parameter θ i ’;
[0068]
[0069] where: α is the inner-layer learning rate;
[0070] Outer-layer optimization, based on the performance of the updated parameter θ i ’on all radiotherapy tasks, optimize the initial parameter θ, and the formula is as follows:
[0071]
[0072] where: β is the outer-layer learning rate.
[0073] As a preferred solution, the radiotherapy task decision-making method further includes:
[0074] Step S5: Execute the radiotherapy task using the optimal model and execution strategy output by the agent;
[0075] Step S6: Obtain the feedback data after the radiotherapy task is completed and input it into the agent;
[0076] Step S7: The agent performs fine-tuning and optimization according to the feedback data.
[0077] As a preferred solution, the agent regularly evaluates the performance and call frequency of the models in the model library, and adjusts the models in the model library using a dynamic expansion and pruning mechanism.
[0078] In a second aspect, the present invention discloses a computing device, including:
[0079] One or more processors;
[0080] A memory;
[0081] And one or more programs, where one or more programs are stored in the memory and configured to be executed by one or more processors, and one or more programs include instructions for any of the above-mentioned radiotherapy task decision-making methods driven by an agent.
[0082] In a third aspect, the present invention discloses a storage medium storing one or more computer-readable programs, where the one or more programs include instructions adapted to be loaded and executed by a memory to perform any of the above-mentioned intelligent agent-driven radiotherapy task decision-making methods.
[0083] The present invention discloses an intelligent agent-driven radiotherapy task decision-making method, a computing device, and a storage medium, which can effectively solve the problems of efficient processing and dynamic adaptation of multi-modal image data in the radiotherapy process, and have the following beneficial effects:
[0084] First, the present invention introduces intelligent agent technology, enabling multi-modal image data to be efficiently processed collaboratively in the same system. The intelligent agent automatically selects appropriate execution strategies and models according to the requirements of radiotherapy tasks, breaking the current situation of fragmented multi-modal image data, improving the efficiency of image data processing, and ensuring the accuracy and stability of tasks at each stage.
[0085] Second, the intelligent agent disclosed in the present invention can automatically and dynamically select the optimal model from the model library according to the current radiotherapy task. Through the task-model adaptation mechanism, the balance between accuracy and efficiency for different tasks is ensured, and the generalization ability of the system is greatly improved.
[0086] Third, the present invention introduces an adaptive learning mechanism, enabling the intelligent agent to dynamically adjust strategies according to real-time feedback during operation, enhancing the real-time response ability.
[0087] Fourth, the present invention efficiently manages the model library through a dynamic expansion and pruning mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0089] Figure 1 It is a flowchart of the radiotherapy task decision-making method provided by an embodiment of the present invention.
[0090] Figure 2 It is a schematic diagram of Scenarios 1-5 provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0091] The following details the preferred embodiments of the present invention with reference to the drawings.
[0092] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0093] The expression of "including" an element is an "open-ended" expression, and this "open-ended" expression only means that there are corresponding components or steps, and should not be construed as excluding additional components or steps.
[0094] There are a wide variety of task types involved in the radiotherapy process, but the prior art lacks a model management mechanism for task characteristics. To solve the above problems, the present invention discloses an agent-driven radiotherapy task decision-making method. In some embodiments, such as Figure 1 shown, the radiotherapy task decision-making method includes:
[0095] Step S101: Identify the current radiotherapy task and obtain image data;
[0096] Step S102: Process the image data to obtain data modality information and data feature information;
[0097] Step S103: Based on the identified current radiotherapy task, obtain the models selected by historical tasks with high similarity to the current radiotherapy task and the performance scores of these models in historical tasks, and form reference information;
[0098] Step S104: Input the current radiotherapy task, data modality information, data feature information, and reference information into the trained agent. The agent can select one or more optimal models from the model library and output the execution strategies of one or more optimal models.
[0099] The present invention adopts a task-driven agent framework and dynamically selects the optimal model from the model library according to the requirements of specific radiotherapy tasks. This method can not only improve the image processing efficiency, but also ensure the accuracy and stability of tasks at each stage by introducing a model library management and adaptive optimization mechanism. For example, in the pre-treatment CT image, the agent can select a high-precision organ segmentation model; in the in-treatment CBCT image, the agent can quickly select a low-computation-overhead positioning correction model. Through the global scheduling and decision-making mechanism of the agent, the problems of algorithm selection and resource optimization in the multi-modal image data processing flow are solved.
[0100] The present invention can not only solve the adaptability problem of complex task scenarios in the radiotherapy process, but also improve the efficiency and accuracy of the entire treatment process, providing a new technical path for the intelligent development of modern radiotherapy.
[0101] The above steps are further elaborated below.
[0102] In step S101, the radiotherapy task is recognized to obtain image data.
[0103] The types of radiotherapy tasks can include but are not limited to: segmentation, dose prediction, positioning verification, etc.
[0104] In step S102, the image data is processed to obtain data modality information and data feature information.
[0105] The data modality information includes but is not limited to: CT, CBCT, plain film, etc.
[0106] The data feature information includes but is not limited to: image noise, resolution, patient data distribution, etc.
[0107] In step S103, based on the recognized current radiotherapy task, the model selected by the historical task with a high similarity to the current radiotherapy task and the performance score of the model in the historical task are obtained to form reference information.
[0108] The data obtained above can be represented by a state vector S, specifically as follows:
[0109] S = [T, M, F, H]
[0110] The meanings of each letter are as follows:
[0111] T is the type of the current radiotherapy task, such as: "segmentation", "dose prediction", etc., which can be represented by a one-hot encoding, e.g., T = [1, 0, 0] (indicating the segmentation task, and other tasks are 0).
[0112] M is the data modality information, which can be generated by a feature extraction method.
[0113] The data modality information can include but is not limited to: modality type, noise mean, noise variance, contrast, etc.
[0114] The modality type (such as: CT, CBCT, plain film) can be represented by a one-hot encoding, and the noise mean, noise variance, and contrast are around the image statistical features. Specifically, e.g., M = [modality type, noise mean, noise variance, contrast].
[0115] F is the data feature information, which can include but is not limited to: image size Size, image resolution Resolution, signal-to-noise ratio SNR, and histogram statistical feature HSF, etc. Specifically, e.g., F = [Size, Resolution, SNR, HSF].
[0116] H is the historical radiotherapy task and model performance data, including the model of the intelligent agent's historical selection for similar tasks and the performance score P of the model.
[0117] It can be stored in the form of time series, such as:
[0118] H = {(T1, m1, P1), (T2, m2, P2),...}.
[0119] In step S104, the current radiotherapy task, data modality information, data feature information, and reference information are input into the trained intelligent agent, which can select one or more optimal models from the model library and output the execution strategies of one or more optimal models.
[0120] The execution strategy can include: the selection of preprocessing steps (such as denoising, enhancement), and the order of collaborative processing of multiple models.
[0121] The execution strategy is represented by a sequence as: A = [a1, a2,..., a N ,
[0122] where: a t represents the operation at the t-th step (such as: calling a specific model or performing a specific process).
[0123] In the whole input-output process, the behavior of the intelligent agent can be modeled as a policy function π, which generates an output action A through the input state S:
[0124] where: θ is the internal parameter of the intelligent agent, which can be optimized through reinforcement learning or supervised learning.
[0125] The model selection problem can be regarded as a multi-armed bandit problem (Multi-Armed Bandit, MAB) or a reinforcement learning task scheduling problem. The intelligent agent needs to dynamically select models and optimize strategies in each task.
[0126] There are N candidate models stored in the model library, specifically: M = {m1, m2,..., m N}.
[0127] If only one model is required to implement the current radiotherapy task, the intelligent agent selects an optimal model from the model library using the following formula;
[0128]
[0129] where:
[0130] α and β are weight factors;
[0131] P i is the model m iThe predicted performance score of the current radiotherapy task can be evaluated through supervised learning (such as accuracy, Dice coefficient);
[0132] C i is the computational overhead of model m i (such as time, resource consumption, memory occupancy, etc.).
[0133] The performance P i of each of the above models m i and the computational overhead C i are represented by the following functions:
[0134] P i = f(T, M, F, H);
[0135] C i = g(T, M, F).
[0136] If multiple models are required to cooperate to achieve the current radiotherapy task, the agent selects multiple optimal models from the model library using the following formula;
[0137]
[0138] where:
[0139] A is the call order and combination strategy of multiple models (such as denoising first and then segmentation);
[0140] T is the phased process of the radiotherapy task;
[0141] α and β are weight factors;
[0142] P t is the predicted performance score at the t-th stage;
[0143] C t is the computational overhead at the t-th stage.
[0144] The agent can train a scheduling policy π(a|s) through reinforcement learning to dynamically determine the action a selected in state s (such as which model to select, whether to perform preprocessing, etc.).
[0145] The state, action, and reward are defined as follows:
[0146] The state s includes: radiotherapy task, data modality information, data feature information, and reference information;
[0147] The action a includes: selecting one or more models;
[0148] The reward r includes: dynamically generated according to the radiotherapy task;
[0149] If it is a segmentation task, the reward can be the Dice coefficient: r = Dice(i, GroundTruth);
[0150] If it is a positioning task, the reward can be the negative value of the positioning error: r = -|error|.
[0151] The goal of the agent is to maximize the cumulative reward by choosing the optimal policy π(a|s), and its objective function is:
[0152]
[0153] Where:
[0154] γ t is the discount factor;
[0155] r t is the immediate reward at the t-th step.
[0156] Furthermore, use DQN (Deep Q-Network) or PPO (Proximal Policy Optimization) to optimize the agent, which specifically includes the following steps:
[0157] Step A: Initialize the agent policy π(a|s) and the model library;
[0158] Step B: Select an action a according to the state s;
[0159] Step C: After executing the action a, obtain the reward r and the new state s';
[0160] Step D: Update the policy parameters;
[0161] For DQN, the optimization objective is as follows:
[0162]
[0163] Or,
[0164] For PPO, the optimization objective is as follows:
[0165] L PPO = E t [min(r t (θ)·A t , clip(r t (θ), 1 - ∈, 1 + ∈)·A t )];
[0166] Where:
[0167] A t is the advantage function;
[0168] clip() is the truncation function;
[0169] r t r(θ) is the policy ratio.
[0170] The above clip(r t (θ), 1 - ∈, 1 + ∈) is used to limit the policy ratio r t (θ) within the interval [1 - ∈, 1 + ∈]. Here, ∈ is a hyperparameter that can take 0.1 and is used to limit the policy update amplitude to prevent unstable training caused by drastic changes.
[0171] If the policy ratio r t (θ) does not change much (i.e., within [1 - ∈, 1 + ∈]), the optimization objective is equivalent to the original PPO objective (r t (θ)·A t ).
[0172] If the policy ratio r t (θ) changes too much (i.e., outside [1 - ∈, 1 + ∈]), the optimization objective will be clipped to avoid too fast policy changes, thereby improving training stability.
[0173] Furthermore, the agent can adapt to new radiotherapy tasks or imaging data of new modalities through the Meta - Learning mechanism;
[0174] Using the MAML (Model - Agnostic Meta - Learning) framework, if there are n tasks in the radiotherapy task set, specifically: T = {T1, T2, …, T n};
[0175] For each radiotherapy task T i , the optimization goal of the agent is to find the optimal initial parameter θ such that it can quickly adapt to the radiotherapy task after a small number of gradient updates;
[0176] The optimization objective is:
[0177]
[0178] Where:
[0179] is the loss function of radiotherapy task T i (such as the Dice loss for segmentation tasks and the MSE loss for dose prediction);
[0180] α is the learning rate for gradient update;
[0181] is the gradient of radiotherapy task T i at the current parameter θ.
[0182] Furthermore, the optimization of the agent includes inner-layer optimization and outer-layer optimization;
[0183] For inner-layer optimization, for each radiotherapy task T, the following formula is used i to adjust the parameters using gradient update to obtain the updated parameters θ i ';
[0184]
[0185] where: α is the inner-layer learning rate;
[0186] For outer-layer optimization, based on the performance of the updated parameters θ i ' on all radiotherapy tasks, the initial parameters θ are optimized, and the formula is as follows:
[0187]
[0188] where: β is the outer-layer learning rate.
[0189] Based on the above embodiments, further, the radiotherapy task decision method further includes:
[0190] Step S105: Execute the radiotherapy task using the optimal model and execution strategy output by the agent;
[0191] Step S106: Obtain the feedback data after the radiotherapy task is completed and input it into the agent;
[0192] Step S107: The agent performs fine-tuning and optimization based on the feedback data.
[0193] During the operation of the agent, its behavior is fine-tuned and optimized through real-time feedback. Its objective formula is the same as the above reinforcement learning objective function.
[0194] Adaptive optimization formula: By introducing the performance and feedback of the model, the model weights can be dynamically adjusted:
[0195]
[0196] where: R is the reward related to the performance of the model selected by the agent.
[0197] Furthermore, the above agent can also regularly evaluate the performance and call frequency of the models in the model library, and adjust the models in the model library using a dynamic expansion and pruning mechanism.
[0198] In the prior art, redundant models may lead to waste of system complexity and computing resources. The present invention manages and dynamically selects the model library.
[0199] When the agent is running, it may need to dynamically expand or prune the model library to improve efficiency and adaptability. The dynamic expansion formula is as follows:
[0200] New model m new When adding a new model m, its performance P needs to be evaluated new :
[0201] P new = f(T, M, F, H);
[0202] If P new meets the performance threshold P threshold , it will be added to the model library;
[0203] M ← M ∪ {m new}
[0204] The dynamic pruning formula is as follows:
[0205] Regularly evaluate the usage frequency U i and the recent performance P i of each model.
[0206] Where: U i is the number of calls divided by the total number of tasks;
[0207] P i is the recent average performance of the model.
[0208] If the model m i meets the following conditions, it will be removed from the model library:
[0209] U i < U threshold or P i < P threshold ;
[0210] M ← M\{m i}.
[0211] As Figure 2 shown below, several specific application scenarios of the present invention are introduced.
[0212] Scenario 1: Organ delineation after CT image acquisition
[0213] The newly acquired CT images of patients need to delineate the tumor region and organs at risk (OAR). The image quality of different patients varies greatly (for example, some images have high noise and some have low resolution), and the accuracy of delineation directly affects subsequent dose prediction.
[0214] First, task recognition: The agent receives the CT data of a new patient and recognizes the current task as "organ delineation".
[0215] Secondly, model selection:
[0216] 1) First, analyze the CT image quality:
[0217] If the image has high resolution and low noise, the agent selects a standard deep learning segmentation model (such as nnUNet, Swin-UNet, ViT, etc.).
[0218] If the image quality is poor (with high noise), the agent triggers the "preprocessing module" to first process the image with a denoising model (such as Denoising Autoencoder), and then send the processed image to the segmentation model.
[0219] 2) The agent selects a specially trained model according to the outlined target (such as tumor type, organ type). For example:
[0220] Lung tumor segmentation -> Select a segmentation model trained on lung CT datasets.
[0221] Abdominal tumor segmentation -> Select a model optimized for abdominal data.
[0222] Finally, execution and feedback:
[0223] Call the segmentation model to complete the outlining and output the results. If the outlining results show anomalies (such as missing key regions), the agent will automatically try alternative segmentation models or notify manual proofreading.
[0224] Scenario 2: Dose distribution prediction
[0225] Based on the organ segmentation and tumor regions of the patient, the agent needs to predict the optimal radiotherapy dose distribution. However, due to differences in patient body shape and organ position complexity, the requirements for the prediction model also vary.
[0226] First, task recognition: The agent receives the dose prediction task, and the inputs include CT images and the organ segmentation results of Scenario 1.
[0227] Secondly, model selection:
[0228] 1) According to the tumor location of the patient:
[0229] If it is a simple regular distribution (such as a small-volume tumor), the agent selects a lightweight dose regression model (such as 3D-CNN).
[0230] If the tumor is large in volume and complex in shape, the agent will select a more complex model (such as a Transformer-based model) to generate a high-precision dose distribution.
[0231] 2) According to the computing resources:
[0232] In the case of limited GPU computing resources, the agent selects a more efficient model (such as UNet).
[0233] If computing resources are sufficient, prioritize the invocation of complex models to pursue high precision.
[0234] Finally, execution and feedback:
[0235] The agent runs the model to generate the dose distribution. If the dose distribution output by the model deviates from the planned target (e.g., the dose to a critical organ exceeds the standard), the agent will trigger the "dose adjustment module" for optimization.
[0236] Scenario 3: Anomaly detection in plan verification
[0237] After the radiotherapy plan is generated (scenarios 1 and 2 are completed), it is necessary to verify whether the dose distribution complies with the specifications and whether there is excessive irradiation of critical organs.
[0238] First, task identification: The agent receives the radiotherapy plan and the dose distribution map, and identifies the task as "plan verification".
[0239] Secondly, model selection:
[0240] 1) First, analyze the complexity of the plan:
[0241] If the plan is simple (regular dose distribution), select a lightweight rule verification algorithm.
[0242] If the plan is complex (extensive organ protection), the agent selects a deep learning-based anomaly detection model (such as GAN anomaly detection).
[0243] 2) Regional analysis of the dose distribution:
[0244] Dose region of critical organs -> Use rule detection.
[0245] Irregular distribution region -> Use a deep learning model.
[0246] Finally, execution and feedback:
[0247] Invoke the model to verify the plan and detect abnormal doses. If problems are found (such as organ dose exceeding the standard), the agent notifies the plan optimization module and provides modification suggestions.
[0248] Scenario 4: Setup verification during treatment
[0249] After scenario 3 is completed and the treatment stage begins. During radiotherapy, the patient needs to perform setup verification through CBCT or plain film images. However, the tissue displacements of different patients are different, and appropriate methods need to be selected according to the image characteristics.
[0250] First, task identification: The agent receives the initial setup CBCT image and the CT reference image, and the task is "setup verification".
[0251] Secondly, model selection:
[0252] 1) If the input is CBCT:
[0253] For high-quality CBCT: Select an image registration model based on deep learning (such as VoxelMorph).
[0254] For low-quality CBCT: The agent will first call an enhancement model (such as GAN denoising) and then perform registration.
[0255] 2) If the input is a plain film image:
[0256] For plain film images, the agent selects a positioning verification algorithm based on key point detection (such as Siamese network).
[0257] 3) According to the positioning error threshold:
[0258] If the error is small, preferentially use a lightweight algorithm for quick verification.
[0259] If the error is large, select a more complex model for precise matching.
[0260] Finally, execution and feedback:
[0261] The agent runs the model and outputs positioning error data.
[0262] If the error exceeds the allowable range, the agent issues an alarm and recommends readjusting the patient's positioning.
[0263] Scenario 5: Follow-up analysis during dynamic treatment
[0264] During the treatment process, the patient's CBCT is collected weekly, and the agent needs to analyze the tissue change trend and dynamically adjust the treatment plan.
[0265] First, task identification: The agent receives the CBCT data and plain film data for one week, and the task is "tissue change analysis".
[0266] Secondly, model selection:
[0267] 1) The agent will call a time series analysis model (such as based on LSTM or VAE) to detect tissue changes.
[0268] 2) If the change amplitude is small, select a fast prediction model.
[0269] 3) If the change is obvious, select a multi-modal fusion model to perform joint analysis by combining the initial CT and the latest CBCT.
[0270] Finally, execution and feedback:
[0271] Output the change trend (such as tumor shrinkage, tissue displacement).
[0272] If it is found that the dose distribution needs to be adjusted, the intelligent agent will automatically trigger the plan optimization module.
[0273] In summary, the present invention designs a unified multi-modal image data processing and analysis framework by introducing the intelligent agent (Agent) technology, enabling efficient collaborative processing of various modal data such as CT, CBCT, and X-ray plain films in the same system. The intelligent agent automatically selects the appropriate processing path and model according to the task requirements, thus breaking the current situation of fragmented multi-modal data.
[0274] Based on the dynamic optimization mechanism of model library management, the intelligent agent can automatically select the optimal model from the model library according to the requirements of the current task (such as organ segmentation, dose prediction, or positioning correction). Through the task-model adaptation mechanism, the balance between accuracy and efficiency of different tasks is ensured, and the generalization ability of the system is greatly improved.
[0275] The present invention adopts reinforcement learning or meta-learning technology to enable the intelligent agent to dynamically adjust the strategy according to real-time feedback during operation. For example, during the patient treatment process, the intelligent agent can quickly adapt to anatomical structure changes and automatically update the segmentation and dose prediction schemes, thus realizing real-time and precise personalized treatment.
[0276] The present invention efficiently manages the model library through a dynamic expansion and pruning mechanism. The intelligent agent regularly evaluates the performance and call frequency of the models, eliminates inefficient models, and automatically introduces efficient models according to the requirements of new tasks, ensuring the flexibility and resource utilization rate of the system.
[0277] The present invention significantly improves the efficiency and accuracy of image data processing through an intelligent scheduling and optimization mechanism, reduces manual intervention, optimizes the entire radiotherapy process, and thus promotes the intelligent development of radiotherapy.
[0278] In some other embodiments, the present invention discloses a computing device, including:
[0279] One or more processors;
[0280] A memory;
[0281] And one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors, and one or more programs include instructions for any of the above intelligent agent-driven radiotherapy task decision-making methods.
[0282] In some other embodiments, the present invention discloses a storage medium storing one or more computer-readable programs, and one or more programs include instructions adapted to be loaded and executed by the memory for any of the above intelligent agent-driven radiotherapy task decision-making methods.
[0283] The present invention discloses an intelligent agent-driven radiotherapy task decision-making method, a computing device, and a storage medium, which can effectively solve the problems of efficient processing and dynamic adaptation of multi-modal image data in the radiotherapy process, and have the following beneficial effects:
[0284] First, the present invention introduces intelligent agent technology, enabling multi-modal image data to be efficiently processed collaboratively in the same system. The intelligent agent automatically selects appropriate execution strategies and models according to the requirements of radiotherapy tasks, breaking the current situation of fragmented multi-modal image data, improving the efficiency of image data processing, and ensuring the accuracy and stability of tasks at each stage.
[0285] Second, the intelligent agent disclosed in the present invention can dynamically select the optimal model from the model library according to the current radiotherapy task. Through the task-model adaptation mechanism, it ensures the balance between accuracy and efficiency for different tasks, and greatly improves the generalization ability of the system.
[0286] Third, the present invention introduces an adaptive learning mechanism, and the intelligent agent can dynamically adjust strategies according to real-time feedback during operation, enhancing the real-time response ability.
[0287] Fourth, the present invention efficiently manages the model library through a dynamic expansion and pruning mechanism.
[0288] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An agent-driven radiotherapy task decision-making method, characterized in that: include: Step S1: Identify the current radiotherapy task and obtain image data; Step S2: Process the image data to obtain data modality information and data feature information; Step S3: Based on the identified current radiotherapy task, obtain the model selected by the historical task with high similarity to the current radiotherapy task and the performance score of the model in the historical task to form reference information; Step S4: Input the current radiotherapy task, data modality information, data feature information and reference information into the trained intelligent agent, which can select one or more optimal models from the model library and output execution strategies of one or more optimal models.
2. The radiotherapy task decision method according to claim 1, characterized in that: The model library stores N candidate models, specifically: M = {m1, m2, ..., m N }; If the current radiotherapy task can be achieved with only one model, the agent selects an optimal model from the model library using the following formula: in: α and β are weight factors; P i For model m i predicted performance scores on the current radiotherapy task; C i For model m i The computational overhead of If the current radiotherapy task requires multiple models to be implemented collaboratively, the agent selects multiple optimal models from the model library using the following formula; in: A is the calling order and combination strategy of multiple models; T is the phased process of radiotherapy tasks; α and β are weight factors; P t is the predicted performance score at stage t; C t is the computational cost of the tth stage.
3. The radiotherapy task decision method according to claim 1, characterized in that: Use reinforcement learning methods to train intelligent agents; The state s includes: radiotherapy tasks, data modality information, data feature information and reference information; Action a includes: selecting one or more models; The reward r includes: dynamically generated according to the radiotherapy task; The goal of the agent is to maximize the cumulative reward by selecting the optimal strategy π(a|s), and its objective function is: in: γ t is the discount factor; r t is the immediate reward at step t.
4. The radiotherapy task decision method according to claim 3, characterized in that: Using DQN or PPO to optimize the agent involves the following steps: Step A: Initialize the agent strategy π(a|s) and model library; Step B: Select action a according to state s; Step C: After executing action a, obtain reward r and new state s′; Step D: Update strategy parameters; For DQN, the optimization objective is as follows: or, For PPO, the optimization objective is as follows: L PPO =E t [min(r t (θ)·A t ,clip(r t (θ),1-∈,1+∈)·A t )]; in: A t is the advantage function; clip() is the truncation function; r t (θ) is the strategy ratio.
5. The radiotherapy task decision method according to claim 1, characterized in that: The agent can adapt to new radiotherapy tasks or new modalities of imaging data through a meta-learning mechanism; If the radiotherapy task set has n tasks, specifically: T = {T1, T2, ..., T n (; For each radiotherapy task T i , the goal of the agent optimization is to find the optimal initial parameter θ, so that it can quickly adapt to the radiotherapy task after a small amount of gradient updates; The optimization goal is: in: For radiotherapy tasks T i The loss function of α is the learning rate, used for gradient update; For radiotherapy tasks T i The gradient at the current parameter θ.
6. The radiotherapy task decision method according to claim 5, characterized in that: The optimization of the intelligent agent includes inner layer optimization and outer layer optimization; The inner layer optimization is performed by the following formula for each radiotherapy task T i Use the gradient to update the adjustment parameters and get the updated parameters θ i '; Where: α is the inner layer learning rate; The outer layer optimizes the updated parameters θ on all radiotherapy tasks i ' is used as the basis for optimizing the initial parameter θ, and the formula is as follows: Where: β is the outer layer learning rate.
7. The radiotherapy task decision method according to any one of claims 1 to 6, characterized in that: The radiotherapy task decision method further includes: Step S5: using the optimal model and execution strategy output by the agent to perform the radiotherapy task; Step S6: Obtain feedback data after the radiotherapy task is completed and input it into the agent; Step S7: The agent performs fine-tuning and optimization based on the feedback data.
8. The radiotherapy task decision method according to any one of claims 1 to 6, characterized in that: The intelligent agent regularly evaluates the performance and calling frequency of the models in the model library, and uses a dynamic expansion and pruning mechanism to adjust the models in the model library.
9. A computing device, characterized in that include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and one or more of the programs include instructions for the agent-driven radiotherapy task decision-making method described in any one of claims 1-8 above.
10. A storage medium, characterized in that The storage medium stores one or more computer-readable programs, and the one or more programs include instructions, which are suitable for being loaded by the memory and executing the intelligent agent-driven radiotherapy task decision method described in any one of claims 1-8.
Citation Information
Cited By
Multi-agent optimization training method and system based on perceptual analysis
CN121052315A
Multi-agent optimization training method and system based on perception analysis
CN121052315B