Radiotherapy report generation method and device based on multi-modal agent and storage medium
By adopting multimodal agent collaboration method in the radiotherapy report generation system, problems such as information islands and low degree of automation in the existing system are solved, fusion and analysis of multimodal data are realized, and efficient and accurate radiotherapy reports are generated.
Patent Information
- Application Number
- CN202510318267.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-01
AI Technical Summary
The existing radiotherapy report generation system has problems such as information silos, low degree of automation, lack of intelligent analysis capabilities and insufficient real-time performance, making it difficult to effectively integrate multimodal data and reduce manual intervention.
The radiotherapy report generation method based on multimodal agents is adopted, and the multimodal fusion and analysis of image data and dose data are achieved through the collaboration of image agents, dose agents, fusion agents and report generation agents, and an automatically generated radiotherapy report is generated.
The comprehensive fusion and precise analysis of multimodal data is achieved, which reduces manual intervention, improves the accuracy and efficiency of reporting, and enhances the real-time and flexibility of the system.
Smart Images

Figure CN120236702A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radiotherapy, and particularly relates to a radiotherapy report generation method, device and storage medium based on a multi-modal intelligent agent. Background Art
[0002] Radiotherapy is a common method for cancer treatment, which uses high-energy radiation to destroy tumor cells while minimizing damage to surrounding normal tissues as much as possible. The core of radiotherapy includes accurate patient positioning, dose distribution control, and post-treatment effect evaluation. Among them, the positioning accuracy of the patient directly affects the treatment effect, and the dose distribution determines the irradiation intensity of the tumor area. Incorrect dose distribution may lead to treatment failure or side effects. Post-treatment imaging evaluation (such as soft tissue changes) is also an important part of treatment effect monitoring.
[0003] Currently, during the radiotherapy process, doctors will evaluate the treatment effect through imaging data (such as CT, CBCT, etc.), dose distribution data, and post-treatment imaging data (such as soft tissue changes). The processing of these data is usually carried out manually, involving a large amount of cumbersome calculations and image analysis, with a large workload, long time consumption, and prone to human errors, affecting the diagnosis and treatment efficiency and the accuracy of treatment.
[0004] With the continuous progress of technology, the large amount of data accumulation during the radiotherapy process urgently needs automated processing and report generation. Traditional methods for recording treatment processes and generating reports usually rely on manual intervention, which is not only inefficient but also prone to subjective differences and errors. Therefore, an automated radiotherapy process recording and report generation system has become an urgent need for radiologists and treatment planners.
[0005] With the rapid development of deep learning and artificial intelligence technologies, especially the maturity of multi-modal learning and natural language processing technologies, it has become more feasible to automatically generate accurate and standardized treatment reports.
[0006] In the field of medical image analysis, multi-modal learning refers to extracting information from multiple different modalities of data (such as images, doses, texts, etc.) and performing joint analysis. Multi-modal data contains information from different sources and can provide a comprehensive evaluation for the radiotherapy process. However, how to effectively integrate these multi-modal data, especially how to reduce human intervention while ensuring accuracy, remains a technical problem. Traditional deep learning models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), although capable of processing single-modal data (such as images or texts), still face many challenges when processing multiple modalities (such as containing both image and dose information simultaneously).
[0007] In summary, the following problems exist when generating existing radiotherapy reports.
[0008] First, the problem of information silos. Current radiotherapy data includes patient imaging data (CT, CBCT, plain films) and treatment dose data. These data are usually stored and processed separately, lacking a unified analysis framework. Existing systems are difficult to effectively integrate multi-modal data, resulting in the inability to fully utilize important information.
[0009] Second, low degree of automation. Traditional report generation mainly relies on manual operations, including image contrast analysis, dose assessment, and description of soft tissue changes, etc. It is time-consuming and laborious and is easily affected by subjective factors. This not only increases the workload of clinicians but also may lead to human omissions.
[0010] Third, lack of intelligent analysis ability. Currently, some automated tools are limited to single tasks, such as image registration or dose distribution analysis, and cannot achieve comprehensive processing of cross-modal data. Due to the lack of a global perspective, the system often has difficulty accurately describing the overall situation of the treatment process when generating reports.
[0011] Fourth, lack of real-time performance and flexibility. Existing single-task tools or models lack the ability of real-time feedback and dynamic adjustment and cannot adapt to changes in different treatment plans or data characteristics. At the same time, the generality of the model is poor, and it often needs to be retrained for specific tasks, increasing the difficulty of clinical application. Summary of the Invention
[0012] To solve the above technical problems, the present invention proposes a radiotherapy report generation method, device, and storage medium based on a multi-modal agent.
[0013] To achieve the above object, the technical solution of the present invention is as follows:
[0014] In the first aspect, the present invention discloses a radiotherapy report generation method based on a multi-modal agent, including:
[0015] Step S1: Input the imaging data and dose data of the patient;
[0016] Step S2: According to the type of input data and task description, decompose the task into multiple subtasks and assign them to the corresponding imaging agent, dose agent, fusion agent, and report generation agent;
[0017] Step S3: The imaging agent, dose agent, fusion agent, and report generation agent cooperate with each other;
[0018] Step S4: The imaging agent processes the imaging data and extracts imaging features;
[0019] The dose agent processes the dose data and extracts dose features;
[0020] The fusion agent integrates image features and dose features to establish integrated data with a multi-modal representation;
[0021] The report generation agent converts the integrated data into a natural language report to generate a radiotherapy report, which includes patient positioning information, dose distribution analysis, and a description of soft tissue changes.
[0022] Based on the above technical solution, the following improvements can be made:
[0023] As a preferred solution, step S2 includes:
[0024] Step S2.1: Decompose the task into multiple subtasks according to the type of input data and the task description;
[0025] Step S2.2: Calculate the priority P(Ti) of each subtask through the following formula;
[0026] P(Ti) = w1 * Criticality(Ti) + w2 * DataSize(Ti) - w3 * ProcessingTime(Ti);
[0027] Where:
[0028] w1, w2, and w3 are all weight coefficients;
[0029] Criticality(Ti) is the importance of the i-th subtask Ti;
[0030] DataSize(Ti) is the size of the input data of the i-th subtask Ti;
[0031] ProcessingTime(Ti) is the estimated processing time of the i-th subtask Ti;
[0032] Step S2.3: Assign each subtask with its priority to the basic models of the corresponding image agent, dose agent, fusion agent, or report generation agent.
[0033] As a preferred solution, the meta-learning mechanism is used to adaptively and dynamically adjust and optimize the selected basic model. The specific optimization objective function is as follows:
[0034]
[0035] Where:
[0036] is the model fine-tuned based on the i-th subtask Ti;
[0037] is the loss function of the i-th subtask Ti;
[0038] Let \(T\) be the set of tasks.
[0039] As a preferred solution, each agent has a model cache pool, which is used to cache multiple models fine-tuned based on subtasks.
[0040] As a preferred solution, the following model performance evaluation function is used to select the fine-tuned model with the highest score to execute the corresponding subtask;
[0041] \(S(M\) j ) = \(\alpha\cdot Accuracy(M\) j ) - \(\beta\cdot Latency(M\) j );
[0042] Where:
[0043] Both \(\alpha\) and \(\beta\) are weight parameters;
[0044] \(S(M\) j ) is the comprehensive score of the \(j\)-th fine-tuned model \(M\) j ;
[0045] \(Accuracy(M\) j ) is the accuracy score of the \(j\)-th fine-tuned model \(M\) j ;
[0046] \(Latency(M\) j ) is the execution latency score of the \(j\)-th fine-tuned model \(M\) j .
[0047] As a preferred solution, the image agent, dose agent, fusion agent, and report generation agent can communicate and cooperate in the following three ways;
[0048] Way 1: All agents can read and write data sharing based on the blackboard model;
[0049] Way 2: All agents can exchange data through a message passing protocol;
[0050] Way 3: Use a central controller to dynamically assign tasks to different agents.
[0051] As a preferred solution, when the data generated by the image agent and the dose agent do not meet the requirements, the fusion agent will send an adjustment signal to the image agent or the dose agent.
[0052] As a preferred solution, the base model of the image agent is an image model, which is used to extract the image semantic features of the image data;
[0053] The basic models of the dose agent are a physical simulation model and a deep learning regression model, which are used to analyze dose distribution, generate dose intensity maps, and compare planned doses with actual doses;
[0054] The basic model of the fusion agent is a multi-modal fusion model, which is used to integrate multi-modal data;
[0055] The basic model of the report generation agent is a GPT-based natural language generation model, which is used to convert integrated data into a standardized report.
[0056] In a second aspect, the present invention discloses a computing device, comprising:
[0057] One or more processors;
[0058] A memory;
[0059] And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for any of the above radiotherapy report generation methods based on multi-modal agents.
[0060] In a third aspect, the present invention discloses a storage medium storing one or more computer-readable programs, the one or more programs including instructions adapted to be loaded and executed by the memory for any of the above radiotherapy report generation methods based on multi-modal agents.
[0061] The present invention discloses a radiotherapy report generation method, device and storage medium based on multi-modal agents, having the following beneficial effects:
[0062] First, the present invention can achieve multi-modal data fusion and analysis. The present invention uses image agents, dose agents, and fusion agents to be responsible for image feature extraction, dose analysis, and feature fusion respectively, overcoming the problem of information isolation in traditional systems and generating more comprehensive and accurate treatment reports.
[0063] Second, the present invention can achieve automatic generation of radiotherapy reports. The present invention uses multi-agents (image agents, dose agents, fusion agents, report generation agents) to cooperate to achieve automatic recording and report generation of the entire radiotherapy process, including extraction of positioning information, dose distribution analysis, and description of soft tissue changes, greatly reducing manual intervention and solving the limitations of a single basic model in processing multi-modal data.
[0064] Third, the present invention can achieve intelligent dynamic model selection. The present invention introduces a dynamic model selection mechanism, automatically calls the most suitable basic model according to the characteristics of the input data, takes into account both efficiency and accuracy, and uses the meta-learning mechanism for adaptive adjustment and optimization. At the same time, the task allocation and communication strategies of the intelligent agent are optimized through reinforcement learning to improve real-time performance and flexibility.
[0065] Fourth, the present invention has strong clinical applicability. The present invention can flexibly expand different types of basic models to adapt to various radiotherapy schemes. At the same time, a standardized medical report is automatically generated, which is convenient for doctors to quickly understand the patient's treatment situation, assist in clinical decision-making, reduce manual participation, and improve work efficiency and report quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0067] Figure 1 It is a flowchart of the radiotherapy report generation method provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The preferred embodiments of the present invention will be described in detail below with reference to the drawings.
[0069] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0070] The expression "including" an element is an "open-ended" expression, which only means that there are corresponding components or steps, and should not be construed as excluding additional components or steps.
[0071] In order to achieve the purpose of the present invention, in some embodiments of a radiotherapy report generation method based on a multi-modal intelligent agent, as Figure 1 shown, the radiotherapy report generation method includes:
[0072] Step S101: Input the image data and dose data of the patient;
[0073] Step S102: According to the type of input data and the task description, decompose the task into multiple subtasks and assign them to the corresponding imaging agent, dose agent, fusion agent, and report generation agent;
[0074] Step S103: The imaging agent, dose agent, fusion agent, and report generation agent cooperate with each other;
[0075] Step S104: The imaging agent processes the imaging data and extracts imaging features, such as: features of bone position, tumor region, soft tissue morphology, etc.;
[0076] The dose agent processes the dose data and extracts dose features, such as: dose distribution map, treatment site, dose deviation, etc.;
[0077] The fusion agent integrates the imaging features and dose features to establish integrated data with a multi-modal representation;
[0078] The report generation agent converts the integrated data into a natural language report to generate a radiotherapy report, and the radiotherapy report includes patient positioning information, dose distribution analysis, and description of soft tissue changes.
[0079] The radiotherapy report includes:
[0080] 1) Patient positioning information: Record whether the positioning meets the expectation;
[0081] 2) Dose distribution analysis: Whether the dose distribution meets the planned requirements, and whether there is overdose or underdose;
[0082] 3) Soft tissue changes: The deformation of tissue morphology after treatment (such as tumor volume change, normal tissue displacement).
[0083] In the automated process of radiotherapy, the present invention adopts a multi-agent system (MAS, Multi-Agent System) for task allocation and coordination, which can improve the flexibility and processing efficiency of the system. Each agent is responsible for different subtasks (such as image processing, dose analysis, report generation, etc.), and the agents cooperate with each other through communication protocols to complete the overall task.
[0084] The above steps are elaborated in detail below.
[0085] In step S101, the collected imaging data is the patient positioning image of the patient, such as: CT, CBCT, kV / MV images, etc.
[0086] In some specific embodiments, since CBCT is usually used for positioning at the first treatment, and a flat film is used for positioning in subsequent treatments, the imaging data specifically includes: two combined situations:
[0087] Case 1: DRRs generated by CT and kV or MV plain films;
[0088] Case 2: CT and CBCT.
[0089] The collected dose data is the dose distribution record of the radiotherapy device.
[0090] Furthermore, step S102 includes:
[0091] Step S102.1: Decompose the task into multiple subtasks according to the type of input data and the task description;
[0092] Step S102.2: Calculate the priority P(Ti) of each subtask through the following formula;
[0093] P(Ti) = w1 * Criticality(Ti) + w2 * DataSize(Ti) - w3 * ProcessingTime(Ti);
[0094] Where:
[0095] w1, w2, and w3 are all weight coefficients;
[0096] Criticality(Ti) is the importance of the i-th subtask Ti (e.g., dose analysis takes precedence over report generation);
[0097] DataSize(Ti) is the size of the input data of the i-th subtask Ti;
[0098] ProcessingTime(Ti) is the estimated processing time of the i-th subtask Ti;
[0099] Step S102.3: Assign each subtask with its priority to the basic models of the corresponding image agent, dose agent, fusion agent, or report generation agent.
[0100] Specifically, in step S102.1, the agent can select an appropriate basic model for processing according to the type of input data. For example, a classifier (such as Logistic Regression or Transformer) is used to predict the type of input data. The loss function (classification task) is the cross-entropy loss.
[0101] The image agent can identify the differences between CT and CBCT and call the corresponding pre-trained basic model.
[0102] The agent also decomposes complex tasks into multiple subtasks (such as feature extraction, pattern matching, report generation). Through the task priority assignment method, the task processing order is optimized.
[0103] Furthermore, the selected base model is adaptively and dynamically adjusted and optimized using the meta - learning mechanism (MAML, Model - Agnostic Meta - Learning), with the goal of finding an initial model parameter θ that can quickly adapt in few - shot tasks. The specific optimization objective function is as follows:
[0104]
[0105] Where:
[0106] is the model fine - tuned based on the i - th sub - task Ti;
[0107] is the loss function of the i - th sub - task Ti;
[0108] T is the task set.
[0109] Ti belongs to the task set T, indicating the sum over all tasks Ti belonging to the task set T.
[0110] Furthermore, each agent has a model cache pool, which is used to cache multiple models fine - tuned based on sub - tasks.
[0111] Each agent is responsible for a specific type of task (image processing, dose calculation, etc.). Caching the fine - tuned models in the model cache pool of each agent can effectively reuse the dedicated models for their respective tasks and may improve performance by avoiding repeated model loading or fine - tuning.
[0112] Furthermore, the following model performance evaluation function is used to select the fine - tuned model with the highest score to execute the corresponding sub - task;
[0113] S(M j )=α·Accuracy(M j )-β·Latency(M j );
[0114] Where:
[0115] Both α and β are weight parameters;
[0116] S(M j ) is the comprehensive score of the j - th fine - tuned model M j ;
[0117] Accuracy(M j ) is the accuracy score of the j - th fine - tuned model M j ;
[0118] Latency(M j ) is the latency of the j - th fine - tuned model M jExecution delay score.
[0119] In step S103, the image agent, the dose agent, the fusion agent, and the report generation agent can communicate and cooperate in the following three ways;
[0120] Way 1: All agents can read and write data sharing based on the blackboard model;
[0121] Way 2: All agents can exchange data through a message passing protocol;
[0122] Way 3: Use a central controller to dynamically assign tasks to different agents.
[0123] For Way 1, the intermediate results generated by each agent are stored in the shared "blackboard" for other agents to read.
[0124] Agent Ai outputs result O i , the formula is as follows:
[0125] O i = f θi (X);
[0126] Where:
[0127] X is the input data;
[0128] f θi (X) is the processing function of agent Ai.
[0129] When the data generated by the image agent and the dose agent does not meet the requirements, the fusion agent will send an adjustment signal F i to the image agent or the dose agent.
[0130] F i = g(O i , T);
[0131] Where:
[0132] g() is the feedback function;
[0133] T is the target task requirement.
[0134] Each agent writes and reads information on the shared "blackboard", and the blackboard model maintains a global status table for storing the current task status.
[0135] Example:
[0136] Agent A writes the task result O i to the blackboard: Blackboard[Ti] = Oi;
[0137] Agent B reads and processes: Ri = h(Blackboard[Ti]).
[0138] For Method 2, a lightweight protocol (JSON) is used to exchange messages.
[0139] The message format is as follows:
[0140]
[0141]
[0142] For Method 3, the central controller dynamically allocates tasks and notifies relevant agents through message broadcasting.
[0143] The task scheduling formula is as follows:
[0144]
[0145] Where:
[0146] Load(A i ) is the current task load of Agent A i ;
[0147] Latency(A i ) is the execution latency.
[0148] The basic models of each agent are described below.
[0149] The basic model of the imaging agent is an image model, which is used to extract the image semantic features of imaging data. Such as: SwinTransformer (a variant of ViT), ResNet, UNet (for specific image segmentation tasks), etc.
[0150] The data training source of the basic model of the imaging agent can be: public medical imaging datasets (such as LUNA16, CT-ORG, NSCLC-Radiomics) or customized radiotherapy image data (annotating setup features, such as tumors, bone structures).
[0151] Training method: If there is enough labeled data, the model can be trained from scratch.
[0152] The loss function can be:
[0153] 1) For segmentation tasks: Dice Loss or Jaccard Loss;
[0154] 2) For keypoint detection: MSE or Smooth L1 Loss.
[0155] Use pre-trained medical image models (such as Med3D, SwinUNETR) and fine-tune them based on specific task data. Specifically, the fine-tuning can be carried out by freezing the first few layers and updating the weights of the subsequent layers.
[0156] The basic model of the dose agent is a physical simulation model and a deep learning regression model, which are used to analyze dose distribution, generate dose intensity maps, and compare planned doses with actual doses. Such as: DenseNet or Transformer-based Regression Models, etc.
[0157] The data training source of the basic model of the dose agent can be: dose distribution map (DICOM-RT format) or dose data generated by Monte Carlo simulation (used to expand the scale of training data).
[0158] The training method is training from scratch:
[0159] Model input: dose distribution map + treatment plan data;
[0160] Loss function: mean square error (MSE) or a customized loss function for a specific dose distribution.
[0161] The basic model of the fusion agent is a multimodal fusion model, which is used to integrate multimodal data. Such as: CLIP (Contrastive Language–Image Pretraining) or ALIGN or other cross-modal alignment models.
[0162] The data training source of the basic model of the fusion agent can be: paired image and dose data (annotating the registration relationship) or manually annotated high-risk regions (used for alignment supervision).
[0163] The training method is training from scratch:
[0164] Model input: image embedding + dose embedding;
[0165] Loss function: contrastive learning loss (Contrastive Loss), specifically as follows:
[0166]
[0167] Where: u i and v j are image and dose features;
[0168] τ is the temperature parameter.
[0169] The base model of the report generation agent is a GPT-based natural language generation model, which is used to convert integrated data into a standardized report. For example: GPT-based models (such as T5, GPT-4).
[0170] The data training source of the report generation agent can be: a medical report dataset (such as MIMIC-III) or a customized annotated radiotherapy report (including patient positioning, dose deviation, etc.).
[0171] The training method is training from scratch:
[0172] Use the standard generative model training method to generate text based on cross-modal features.
[0173] Loss function: Cross-Entropy Loss.
[0174] Specifically, most GPT layers can be frozen, and only the terminal generation layer is updated to adapt to specific domain reports for fine-tuning.
[0175] The present invention can break down complex tasks into multiple small tasks, thereby improving the modularity and scalability of the system. Each agent can operate, optimize, and update independently, adapt to different task requirements, and can be dynamically adjusted when necessary. In the automated report generation of radiotherapy, using a multi-agent system can flexibly call different base models according to the task characteristics to ensure the efficiency and accuracy of the entire system.
[0176] In addition, in some other embodiments, the present invention also discloses a computing device, including:
[0177] One or more processors;
[0178] A memory;
[0179] And one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for the method for determining the dominant vertices of the alternating group network disclosed in the above embodiments.
[0180] The processor may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computational operations related to machine learning.
[0181] The memory may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory is used to store at least one instruction, and the at least one instruction is used to be executed by the processor to implement the method for determining the dominating vertices of the alternating group network provided in the method embodiments of the present invention.
[0182] In addition, the computing device may optionally further include: a peripheral device interface and at least one peripheral device. The processor, the memory, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0183] Of course, the computing device may also include fewer or more components, and this embodiment does not limit this.
[0184] In addition, in some other embodiments, the present invention also discloses a storage medium, and the storage medium stores one or more computer-readable programs. The one or more programs include instructions, and the instructions are adapted to be loaded and executed by the memory to implement the method for determining the dominating vertices of the alternating group network disclosed in the above embodiments.
[0185] The present invention discloses a radiotherapy report generation method, device and storage medium based on a multi-modal intelligent agent, having the following beneficial effects:
[0186] First, the present invention can achieve multi-modal data fusion and analysis. The present invention uses an image intelligent agent, a dose intelligent agent, and a fusion intelligent agent to be responsible for image feature extraction, dose analysis, and feature fusion respectively, overcoming the problem of information isolation in traditional systems and generating a more comprehensive and accurate treatment report.
[0187] Second, the present invention can achieve the automatic generation of radiotherapy reports. The present invention uses multi-intelligent agents (image intelligent agent, dose intelligent agent, fusion intelligent agent, report generation intelligent agent) to cooperate to achieve automatic recording and report generation of the entire radiotherapy process, including the extraction of positioning information, dose distribution analysis, and description of soft tissue changes, greatly reducing manual intervention and solving the limitations of a single basic model in processing multi-modal data.
[0188] Third, the present invention can achieve intelligent dynamic model selection. The present invention introduces a dynamic model selection mechanism, automatically calls the most suitable basic model according to the characteristics of the input data, takes into account efficiency and accuracy, and uses a meta-learning mechanism for adaptive adjustment and optimization. At the same time, the task assignment and communication strategies of the intelligent agents are optimized through reinforcement learning to improve real-time performance and flexibility.
[0189] Fourth, the present invention has strong clinical applicability. The present invention can be flexibly extended with different types of basic models to adapt to various radiotherapy schemes. At the same time, it automatically generates standardized medical reports, facilitating doctors to quickly understand the patient's treatment situation, assisting clinical decision-making, reducing manual participation, and improving work efficiency and report quality.
[0190] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A method for generating a radiotherapy report based on a multimodal agent, characterized in that: include: Step S1: inputting the patient's imaging data and dose data; Step S2: According to the type of input data and task description, the task is decomposed into multiple subtasks, and the subtasks are assigned to the corresponding image agent, dose agent, fusion agent and report generation agent; Step S3: the image agent, the dose agent, the fusion agent and the report generation agent cooperate with each other; Step S4: the image agent processes the image data and extracts image features; The dose intelligent agent processes the dose data and extracts dose features; The fusion agent integrates the image features and the dose features to establish integrated data represented by multiple modalities; The report generation agent converts the integrated data into a natural language report to generate a radiotherapy report, which includes patient positioning information, dose distribution analysis, and a description of soft tissue changes.
2. The method for generating a radiotherapy report according to claim 1, characterized in that: The step S2 comprises: Step S2.1: Decompose the task into multiple subtasks according to the type of input data and task description; Step S2.2: Calculate the priority P(Ti) of each subtask by the following formula; P(Ti)=w1*Criticality(Ti)+w2*DataSize(Ti)-w3*ProcessingTime(Ti); in: w1, w2, and w3 are all weight coefficients; Criticality(Ti) is the importance of the i-th subtask Ti; DataSize(Ti) is the size of the input data of the i-th subtask Ti; ProcessingTime(Ti) is the estimated processing time of the ith subtask Ti; Step S2.3: Assign each subtask with its priority to the base model of the corresponding imaging agent, dose agent, fusion agent or report generation agent.
3. The method for generating a radiotherapy report according to claim 2, characterized in that: The meta-learning mechanism is used to adaptively and dynamically adjust and optimize the selected basic model. The specific optimization objective function is as follows: in: is the model fine-tuned based on the i-th subtask Ti; is the loss function of the i-th subtask Ti; T is the task set.
4. The method for generating a radiotherapy report according to claim 3, characterized in that: Each agent has a model cache pool, which is used to cache multiple models fine-tuned based on subtasks.
5. The method for generating a radiotherapy report according to claim 4, characterized in that: Use the following model performance evaluation function to select the fine-tuned model with the highest score to perform the corresponding subtask; S(M j )=α·Accuracy(M j )-β·Latency(M j ); in: α and β are weight parameters; S(M j ) is the jth fine-tuning model M j The overall score of Accuracy(M j ) is the jth fine-tuning model M j The accuracy score of Latency(M j ) The jth fine-tuned model M j Execution latency score.
6. The method for generating a radiotherapy report according to any one of claims 1 to 5, characterized in that: The image agent, dose agent, fusion agent and report generation agent can communicate and collaborate with each other in the following three ways: Method 1: All agents can read and write shared data based on the blackboard model; Method 2: All agents can exchange data through a message passing protocol; Method 3: Use a central controller to dynamically assign tasks to different agents.
7. The method for generating a radiotherapy report according to any one of claims 1 to 5, characterized in that: When the data generated by the image agent and the dose agent do not meet the requirements, the fusion agent will send an adjustment signal to the image agent or the dose agent.
8. The method for generating a radiotherapy report according to any one of claims 1 to 5, characterized in that: The basic model of the image agent is an image model, which is used to extract image semantic features of image data; The basic models of the dose agent are physical simulation models and deep learning regression models, which are used to analyze dose distribution, generate dose intensity maps, and compare planned doses with actual doses; The basic model of the fusion agent is a multimodal fusion model, which is used to realize the integration of multimodal data; The basic model of the report generation agent is a GPT-based natural language generation model, which is used to convert the integrated data into standardized reports.
9. A computing device, characterized in that include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and one or more of the programs include instructions for the radiotherapy report generation method based on a multimodal intelligent agent as described in any one of claims 1-8 above.
10. A storage medium, characterized in that The storage medium stores one or more computer-readable programs, and the one or more programs include instructions, which are suitable for being loaded by the memory and executing the radiotherapy report generation method based on a multimodal intelligent agent as described in any one of claims 1-8.