An intelligent adaptive radiotherapy system
The intelligent adaptive radiotherapy system utilizes a multimodal large language model to process radiotherapy data, enabling simple and efficient radiotherapy plan adjustments. This solves the problems of cumbersome processes and error accumulation in existing technologies, provides personalized treatment suggestions, and optimizes the use and adaptability of medical resources.
Patent Information
- Application Number
- CN202511012870.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing adaptive radiotherapy systems have cumbersome procedures and are prone to error accumulation, making it difficult to efficiently adjust radiotherapy plans when patients' anatomical structures change.
An intelligent adaptive radiotherapy system is adopted, which uses a multimodal large language model to directly extract key information from the data source during the radiotherapy process. The radiotherapy plan is adjusted in real time through visual representation, language representation, semantic alignment and decision modules. End-to-end processing is performed, including the encoder-decoder neural network of the visual representation module, the GPT or DeepSeek decoder architecture of the language representation module, the cross-attention structure of the semantic alignment module and the Qwen 32B model of the decision module.
It enables simple and efficient radiotherapy planning adjustments, reduces errors, provides personalized treatment recommendations, optimizes the use of medical resources, and has good scalability and adaptability to meet changing clinical needs.
Smart Images

Figure CN120526972B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical technology, and in particular to an intelligent adaptive radiotherapy system. Background Technology
[0002] In current cancer radiotherapy practice, an initial radiotherapy plan is typically designed based on the patient's localization CT scan, followed by several radiotherapy sessions. However, throughout the treatment cycle, a patient's anatomy may undergo significant changes, including but not limited to changes in tumor location and volume, alterations in body contour, and differences in organ filling status. These changes may render the static radiotherapy plan unsuitable for the patient's actual condition, thereby affecting treatment outcomes and increasing the risk of damage to normal tissues.
[0003] To address the aforementioned issues, traditional adaptive radiotherapy systems require physicians to redraw the target area and organs at risk based on the guiding images. The treatment plan is then redesigned and optimized by physicists and dosimeters using a commercial Treatment Planning System (TPS), a time-consuming and labor-intensive process. In response, a series of automated adaptive radiotherapy systems have emerged. However, these systems often involve multiple relatively independent algorithm modules working in sequence, resulting in a long and complex overall workflow and the potential for error accumulation. For example, in CN104117151A, rigid and deformation registration are first performed on the guiding and positioning images. Then, the target area and organs at risk are automatically drawn, and displacement and deformation are assessed to determine whether to adjust the radiotherapy plan. This multi-step process increases complexity and the likelihood of potential errors.
[0004] Therefore, proposing an intelligent adaptive radiotherapy system to overcome the difficulties of existing technologies is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides an intelligent adaptive radiotherapy system that extracts key information directly from various data sources in the radiotherapy process based on a multimodal large language model and adjusts the radiotherapy plan in real time to achieve true intelligent adaptive radiotherapy.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] An intelligent adaptive radiotherapy system includes: a visual representation module, a language representation module, a semantic alignment module, and a decision module;
[0008] The visual representation module is connected to the first input of the semantic alignment module to process the localization image and the guidance image, and extract the semantic representation of the image from them.
[0009] The language representation module is connected with a semantic alignment module second input end, and language representation of the original radiotherapy plan and patient examination information is extracted;
[0010] The semantic alignment module output end is connected with a decision module, and image language information is obtained by aligning image semantic representation and language representation;
[0011] The decision module judges whether the radiotherapy plan needs to be adjusted according to the image language information and gives the corresponding adjustment scheme.
[0012] Optionally, the visual representation module adopts an encoder-decoder neural network architecture based on image registration and segmentation technology to analyze and locate the differences between the positioning image and the guide image, specifically:
[0013] The positioning image and the guide image are divided into multiple patch blocks after channel splicing, and feature extraction is performed through a SwinTransformer module and a dimension conversion module;
[0014] In the decoding stage, the spatial resolution is gradually restored by using a 3D convolution layer and an upsampling module;
[0015] The encoder and the decoder are connected by a shortcut at each feature resolution scale to realize direct transmission of information;
[0016] The differences between the target region and the surrounding normal tissue in the guide image and the positioning image are analyzed.
[0017] Optionally, the visual representation module adopts a registration task of the positioning image and the guide image, and a segmentation task of the target region and the organ at risk in the positioning image and the guide image is used as a pre-training task to initialize the backbone network structure of the visual representation module. The visual representation module is optimized through the segmentation and registration pre-training tasks, and the weight is frozen.
[0018] Optionally, the language representation module converts the original radiotherapy plan and patient examination information into vector data format through the Tokenizer and Embedding two key sub-modules in the large-scale language model of the GPT or DeepSeek decoder architecture, specifically:
[0019] Through the Tokenizer sub-module, the text information of the original radiotherapy plan and patient examination information is segmented into a series of tokens;
[0020] The Byte Pair Encoding algorithm is used for word segmentation, and the generated token sequence enters the Embedding sub-module. Each token is mapped to a high-dimensional space through a linear transformation layer and converted into a vector data format.
[0021] The system, optionally, the semantic alignment module maps the image semantic representation and the language representation to a unified representation space, specifically:
[0022] The semantic alignment module performs position encoding and modality encoding on the token sequence obtained from the visual representation module and the language representation module; the position encoding adopts a rotation position encoding method to capture the relative position relationship between the internal elements of the sequence; the modality encoding adopts a one-hot encoding after normalization processing, and the data of different modalities is correctly identified and distinguished under a unified framework;
[0023] The tokens of the encoded image semantic representation and language representation are input into the alignment network to obtain the aligned tokens;
[0024] The alignment network includes a cross-attention structure, a residual connection and layer normalization, a feedforward neural network, and another residual connection and layer normalization.
[0025] The system, optionally, the semantic alignment module is fully fine-tuned, so that the semantic alignment module can capture data pattern changes.
[0026] The system, optionally, the decision module adopts Qwen 32B as the basic model, fine-tunes the Qwen 32B basic model, so that it can integrate and process the target area details in the positioning image and the guiding image in the radiotherapy process, the deformation of the surrounding normal tissue, the initial radiotherapy plan and the patient examination information;
[0027] The decision process of the decision module is divided into two stages:
[0028] In the first stage, the token sequence processed by the semantic alignment module is spliced with the token sequence processed by the semantic representation module after the prompt word "whether the current treatment plan of the patient needs to be adjusted?", and the spliced sequence is input into the large language model for comprehensive analysis. When the analysis result shows that the treatment plan needs to be adjusted, the decision process enters the second stage;
[0029] In the second stage, the token sequence processed by the semantic alignment module is spliced again with the token sequence processed by the semantic representation module after the prompt word "how to adjust the radiotherapy plan of the patient to maximize the therapeutic effect and avoid toxic side effects?", and the spliced token sequence is sent into the large language model to generate specific adaptive radiotherapy plan adjustment suggestions.
[0030] The system, optionally, fine-tunes the Qwen 32B basic model by selecting a low-rank adaptive method.
[0031] Compared with the prior art, the intelligent adaptive radiotherapy system has the following beneficial effects: the intelligent adaptive radiotherapy system is an end-to-end adaptive radiotherapy system, has a simple process and compact link, can output a radiotherapy plan adjustment scheme optimized for the current patient state based on input positioning and guided images, and effectively reduces error accumulation problems that may be caused by intermediate links; the visual representation module has strong deformation field extraction capability and can accurately segment a target region, and can efficiently and accurately perform complex tasks in actual application; the visual representation module can provide more accurate and personalized treatment adjustment suggestions for each patient, optimizes the use efficiency of medical resources, ensures that the model can maintain high efficiency while having good scalability and adaptability, and meets changing clinical needs; and by using differentiated fine-tuning strategies for different modules, such as freezing part of the weight and applying low-rank adaptation (LoRA), the model accuracy can be ensured while reducing the demand for computing resources, thereby realizing more efficient deployment and operation. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, brief introductions will be given to the drawings needed to be used in the embodiments or prior art descriptions. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of the provided drawings.
[0033] Figure 1 A structural diagram of the intelligent adaptive radiotherapy system provided by the present application;
[0034] Figure 2 A principle diagram of the language representation module in the intelligent adaptive radiotherapy system provided by the present application;
[0035] Figure 3 A semantic alignment network structure diagram in the intelligent adaptive radiotherapy system provided by the present application;
[0036] Figure 4 A decision module flowchart in the intelligent adaptive radiotherapy system provided by the present application. DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of protection of the present application.
[0038] In this application, the relationship terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of additional identical elements in the process, method, article or equipment including the element.
[0039] Referring to Figure 1 The application discloses an intelligent adaptive radiotherapy system, comprising a visual representation module, a language representation module, a semantic alignment module and a decision module.
[0040] The visual representation module is connected with the first input end of the semantic alignment module, processes the positioning image and the guide image, and extracts the image semantic representation therein.
[0041] The language representation module is connected with the second input end of the semantic alignment module, extracts the language representation of the original radiotherapy plan and the patient examination information.
[0042] The output end of the semantic alignment module is connected with the decision module, aligns the image semantic representation and the language representation, and obtains the image language information.
[0043] The decision module judges whether the radiotherapy plan needs to be adjusted according to the image language information and gives the corresponding adjustment scheme.
[0044] Further, the core function of the visual representation module is to accurately analyze the difference between the positioning image and the guide image to reflect the displacement of the patient in the positioning process and the change of the tumor target area and the surrounding tissue; the visual representation module adopts an encoder-decoder neural network architecture based on image registration and segmentation technology to analyze the difference between the positioning image and the guide image, specifically:
[0045] The positioning image and the guide image are divided into multiple patch blocks after being spliced through a channel, and feature extraction is performed through a series of SwinTransformer modules and dimension conversion modules.
[0046] In the decoding stage, the spatial resolution is gradually restored by using 3D convolution layers and up-sampling modules to ensure that the output result can accurately reflect the spatial structure information of the original input image.
[0047] In order to ensure the effectiveness of information transmission, the encoder and the decoder are connected by a shortcut connection at each feature resolution scale to realize direct transmission of information, thereby enhancing the learning ability and expression ability of the model.
[0048] In the adaptive radiotherapy system, it is necessary to analyze the differences between the target region and the surrounding normal tissue in the guide image and the positioning image, so the visual representation module should have the ability to extract the deformation field and the target region segmentation.
[0049] Further, the visual representation module not only needs to have strong deformation field extraction capability, but also needs to be able to accurately segment the target region, and the registration task of the guide image and the positioning image, as well as the segmentation task of the target region and the organ at risk in the guide image and the positioning image are used as pre-training tasks to initialize the backbone network structure of the visual representation unit, so as to ensure that it can efficiently and accurately perform the above complex tasks in actual application; the visual representation module has been optimized through pre-training tasks such as segmentation and registration, and its weights have been frozen, which can prevent the destruction of effective feature representation learned in the subsequent training process, while reducing the demand for computing resources.
[0050] Further, referring to Figure 2 As shown in the figure, the language representation module aims to convert the text information of the patient's radiotherapy plan and other related examinations into a vector data format suitable for processing by a large language model;
[0051] This process uses embedding technology in large-scale language models with only decoder (decoder-only) architecture such as GPT and DeepSeek, mainly including two key sub-modules, tokenizer and embedding, to convert the original radiotherapy plan and patient examination information into vector data format, specifically:
[0052] Firstly, through the tokenizer sub-module, the text information is divided into a series of minimum semantic unit sequences (tokens), which are the basic units for processing by large language models; in this invention, we use the current mainstream BytePair Encoding (BBPE) algorithm for tokenization to ensure efficient and accurate text segmentation;
[0053] Subsequently, these generated minimum semantic unit sequences (tokens) enter the embedding sub-module, where each minimum semantic unit sequence is mapped to a high-dimensional space through a linear transformation layer, thereby obtaining a minimum semantic unit sequence (tokens) vector data format suitable for input into a large language model.
[0054] Further, the semantic alignment module maps the image semantic representation and the language representation to a unified representation space, specifically:
[0055] The core goal of the semantic alignment module is to map visual information and natural language descriptions into a unified representation space and achieve deep fusion of the two, thereby achieving mutual conversion and understanding of the semantics between vision and language. This process is a key step in handling multi-modal problems and is a prerequisite for the implementation of current popular multi-modal technologies.
[0056] The semantic alignment module performs position encoding and modality encoding on the token sequences obtained from the visual representation module and the language representation module. Specifically, the position encoding uses the current industry mainstream rotation position encoding (RoPE) method, which can effectively capture the relative position relationship between elements in the sequence and thereby enhance the model's understanding of the sequence structure. For modality encoding, a normalized one-hot encoding is selected to ensure that data from different modalities can be correctly identified and distinguished under a unified framework.
[0057] The encoded image semantic representation and language representation tokens are input into the alignment network to obtain aligned tokens.
[0058] Referring to FIG. 4, Figure 3 The alignment network consists of four main parts: cross-attention structure (Cross Attention), residual connection and layer normalization (Add & Layer Normalization), feedforward neural network (FeedForward Network), and another residual connection and layer normalization (Add & Layer Normalization). Cross-attention mechanism allows the model to establish associations between different modalities, optimizing its own representation by focusing on relevant information from another modality. Then, the addition operation combined with layer normalization helps stabilize the training process and accelerate convergence. Finally, the feedforward neural network further refines the features, and the addition and layer normalization are applied again to ensure the quality and consistency of the output.
[0059] Further, the semantic alignment module, as one of the core components of the system, is responsible for accurately mapping information from different modalities to the same semantic space. It is fully fine-tuned to ensure that it can capture the most subtle data pattern changes and improve the overall performance of the model.
[0060] Further, the Token sequence processed by the semantic alignment module will be fed into a large language model (LLM) for in-depth decision analysis. Considering the balance between model performance and economy, Qwen 32B is selected as the basic model of the decision module. For the specific application requirements of intelligent adaptive radiotherapy, we have specially fine-tuned this basic model to enable it to effectively integrate and process various types of information involved in the radiotherapy process, including not only the details of the target area in the positioning images and the guiding images, but also the deformation of the surrounding normal tissues, the initial radiotherapy plan, and various related examination data of the patient. In this way, the system can provide more accurate and personalized treatment adjustment suggestions for each patient, while optimizing the use efficiency of medical resources. This process ensures that the model can maintain high efficiency while having good scalability and adaptability to meet changing clinical needs.
[0061] Further, referring to Figure 4 As shown in the decision-making process of the decision module is divided into two stages:
[0062] In the first stage, the multi-modal information (including the patient's image data, initial radiotherapy plan and other related examination data) processed by the semantic alignment unit is converted into a sequence of minimum semantic units (tokens), and is spliced with the Token sequence processed by the semantic representation unit after the prompt word 1 (“Does the current treatment plan of this patient need to be adjusted?”). This spliced sequence is then input into the large language model for comprehensive analysis. If the analysis result shows that the treatment plan needs to be adjusted, the process enters the second stage.
[0063] In the second stage, the above multi-modal minimum semantic unit sequence (token) is spliced again with the Token sequence processed by the semantic representation unit after the prompt word 2 (“How to adjust the radiotherapy plan of this patient to maximize the therapeutic effect and avoid side effects?”). Then, this updated Token sequence is sent to the large language model to generate specific treatment plan adjustment suggestions.
[0064] Further, Qwen 32B is used as the basic model, and Qwen 32B is fine-tuned to enable it to integrate and process the details of the target area in the positioning images and the guiding images, the deformation of the surrounding normal tissues, the initial radiotherapy plan, and the patient's examination information in the radiotherapy process.
[0065] In one embodiment, in the course of lung tumor treatment, the tumor size or central position changes as the treatment progresses, and the target region such as GTV also changes, and if the original plan is followed, the dose of radiation received by the target region may not be sufficient to kill the tumor cells; the dose received by the surrounding organs at risk may be too high, resulting in the death of cells in the organs at risk. According to the technical solution of the present application, the original positioning CT image, the guide CT image, and the original radiotherapy plan are input into the model; the visual representation module processes the positioning image and the guide image, and extracts the image semantic representation therein; the language representation module extracts the language representation of the original radiotherapy plan and the patient examination information; the semantic alignment module aligns the image semantic representation and the language representation, and obtains the image language information; the decision module analyzes and fine-tunes the original radiotherapy plan, such as adjusting the machine head angle, adjusting the shape of the multileaf collimator shielding, etc., so that the new radiotherapy plan ensures the therapeutic effect and maximally protects the organs at risk.
[0066] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment mainly describes the difference from other embodiments. In particular, for the system or system embodiments, since it is basically similar to the method embodiments, it is described more simply, and the related parts can be referred to the part of the method embodiments. The above-described system and system embodiments are only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to the actual needs. Those skilled in the art can understand and implement without creative labor.
[0067] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent adaptive radiotherapy system, characterized in that, include: Visual representation module, language representation module, semantic alignment module, and decision module; The visual representation module is connected to the first input of the semantic alignment module, processing the localization image and the guidance image to extract semantic representations. The visual representation module employs an encoder-decoder neural network architecture based on image registration and segmentation techniques to analyze the differences between the localization image and the guidance image, specifically: The localization image and the guidance image are stitched together by channels and divided into multiple patch blocks. Feature extraction is performed through the Swin Transformer module and the dimension transformation module. During the decoding stage, spatial resolution is gradually restored using 3D convolutional layers and upsampling modules; The encoder and decoder communicate directly with each feature resolution scale via shortcut connections; Analyze the differences between the target area and surrounding normal tissues between the guiding images and the localization images; The language representation module is connected to the second input of the semantic alignment module to extract the language representation of the original radiotherapy plan and patient examination information. The language representation module transforms the original radiotherapy plan and patient examination information into vector data format through the two key sub-modules, Tokenizer and Embedding, in a large-scale language model using a GPT or DeepSeek decoder architecture. Specifically: Through the Tokenizer submodule, the text information of the original radiotherapy plan and patient examination information is segmented into a series of tokens; The Byte Pair Encoding algorithm is used for word segmentation. The generated token sequence enters the Embedding submodule. Each token is mapped to a high-dimensional space through a linear transformation layer and converted into vector data format. The output of the semantic alignment module is connected to the decision module to align the semantic representation and linguistic representation of the image to obtain the image linguistic information. The semantic alignment module maps image semantic representations and language representations to a unified representation space, specifically as follows: The semantic alignment module performs positional and modal encoding on the token sequences obtained from the visual representation module and the language representation module; Positional encoding uses a rotational positional encoding method to capture the relative positional relationships between elements within a sequence; modal encoding uses a normalized one-hot encoding method, which correctly identifies and distinguishes data of different modalities within a unified framework. The encoded image semantic and linguistic representation tokens are input into the alignment network to obtain aligned tokens; The alignment network includes: a cross-attention structure, residual connections and layer normalization, a feedforward neural network and another residual connection and layer normalization; The decision-making module determines whether the radiotherapy plan needs to be adjusted based on image and language information and provides corresponding adjustment solutions.
2. The intelligent adaptive radiotherapy system according to claim 1, characterized in that, The visual representation module uses the registration task of localization image and guidance image, and the segmentation task of target area and organs at risk in localization image and guidance image as pre-training task to initialize the backbone network structure of visual representation module. After optimization by segmentation and registration pre-training task, the weights of visual representation module are frozen.
3. The intelligent adaptive radiotherapy system according to claim 1, characterized in that, The semantic alignment module was fully fine-tuned to enable it to capture changes in data patterns.
4. The intelligent adaptive radiotherapy system according to claim 1, characterized in that, The decision module uses Qwen 32B as the base model and fine-tunes the Qwen 32B base model to enable it to integrate and process target details in localization and guidance images during radiotherapy, deformation of surrounding normal tissues, initial radiotherapy plan and patient examination information; The decision-making process of the decision-making module is divided into two stages: In the first stage, the token sequence obtained after processing by the semantic alignment module is concatenated with the token sequence obtained after processing by the semantic representation module for the prompt "Does the current treatment plan for this patient need to be adjusted?" The concatenated sequence is then input into the large language model for comprehensive analysis. The analysis results indicate that when the treatment plan needs to be adjusted, the decision-making process enters the second stage. In the second stage, the token sequence obtained from the semantic alignment module and the token sequence after processing by the semantic representation module for the prompt "How to adjust the patient's radiotherapy plan to maximize efficacy and avoid toxic side effects?" are concatenated again. The concatenated token sequence is then fed into a large language model to generate specific adaptive radiotherapy plan adjustment suggestions.
5. The intelligent adaptive radiotherapy system according to claim 4, characterized in that, The low-rank adaptation method was selected to fine-tune the Qwen 32B basic model.
Citation Information
Patent Citations
Optimization method of online self-adaption radiotherapy plan
CN104117151A
Radiotherapy plan generation method and system based on large model
CN120199422A
Intelligent interaction method for patient in radiotherapy process, computing device and storage medium
CN120215711A