Dose prediction model, method and device based on deep learning, equipment and medium
By fusing CT image data and physical information of the radio source in the deep learning dose prediction model, local and global features are extracted, and a three-dimensional dose distribution map is generated, the problem of existing methods neglecting the physical characteristics of the radio source is improved, and the accuracy and real-time performance of dose prediction are improved.
Patent Information
- Application Number
- CN202510187877.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-27
AI Technical Summary
Existing deep learning methods ignore the physical properties of the radioactive source in dose prediction, resulting in inaccurate dose distribution, especially in dynamic or heterogeneous radiotherapy environments.
A deep learning-based dose prediction model is adopted, which includes a data input module, a deep convolution evolution module, a first fusion module and a prediction output module. By fusing CT image data with the physical information of the radioactive source, local and global context features are extracted to generate a three-dimensional dose distribution map.
By combining the physical characteristics of the radio source, the impact of the radio source on the dose distribution is accurately simulated, which improves the accuracy of dose prediction and meets the needs of clinical real-time prediction.
Smart Images

Figure CN120220966A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of dose prediction, and in particular, to a dose prediction model, method, device, equipment and medium based on deep learning. Background Art
[0002] The prediction of dose is a key step to ensure the dose effect. Traditionally, the prediction of radiotherapy dose mainly relies on physical models, Monte Carlo particle transport simulations or finite element methods. These methods are based on complex physical principles and mathematical calculations, but these traditional methods face significant challenges in practical applications. On the one hand, due to the extremely complex and time-consuming calculation process, these methods cannot meet the needs of real-time dose prediction, thus limiting their immediate application in clinical decision-making. On the other hand, these methods have high requirements for computing resources, increasing the treatment cost and time cost.
[0003] In recent years, deep learning has made certain progress in medical image processing and radiotherapy dose prediction. However, existing deep learning methods still have many limitations when dealing with dose prediction. A major problem is that traditional deep learning methods often ignore the physical characteristics of radiation sources (such as type, energy, irradiation angle, irradiation depth, etc.), which are crucial for accurate dose calculation. When existing models do not combine the physical information of radiation sources, they often cannot accurately simulate the dose distribution, especially in dynamic or heterogeneous radiotherapy environments. There is an urgent need for a deep learning-based dose prediction scheme. Summary of the Invention
[0004] In order to more accurately predict the dose distribution and improve the accuracy of dose prediction, the present application provides a dose prediction model, method, device, equipment and medium based on deep learning.
[0005] In a first aspect, the present application provides a dose prediction model based on deep learning, including a data input module, a deep convolutional evolution module, a first fusion module and a prediction output module;
[0006] The data input module inputs CT image data for the deep convolutional evolution module;
[0007] The deep convolutional evolution module is used to extract local features from the CT image data to obtain local feature map data;
[0008] The first fusion module is used to increase the number of channels describing the CT image data, and fuse the local feature map data with multiple radiation source physical information to obtain first comprehensive feature map data, where the radiation source physical information includes radiation source type, energy of the radiation source, irradiation angle or irradiation position;
[0009] The prediction output module is configured to predict and output a three-dimensional dose distribution map according to the first comprehensive feature map data, where the three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position.
[0010] The beneficial effects of this application are as follows: By taking the physical characteristics of the radiation source as part of the input data corresponding to the dose prediction model, the influence of the radiation source on the dose distribution can be accurately simulated, avoiding the problem that the traditional method ignores the characteristics of the radiation source and improving the accuracy of dose prediction.
[0011] Further, the prediction output module includes a sliding window transformation network module, a second fusion module, and a decoder module;
[0012] The sliding window transformation network module is configured to obtain global context feature data corresponding to the first comprehensive feature map data, where the global context feature data is a parameter characterizing the global information in the CT image data;
[0013] The second fusion module is configured to fuse the global context feature data and the local feature map data to obtain second comprehensive feature map data;
[0014] The decoder module is configured to gradually perform multiple upsampling operations on the second comprehensive feature map data to obtain the three-dimensional dose distribution map corresponding to the second comprehensive feature map data.
[0015] The beneficial effect of adopting the above further solution is as follows: By combining the deep convolutional evolution module and the sliding window transformation network module, the deep convolutional evolution module extracts anatomical features in the CT image data, and the sliding window transformation network module uses the anatomical features extracted in the previous step + the physical information of the radiation source to learn the spatial distribution of the dose. After local feature extraction and global context modeling, the second fusion module fuses the global context feature data and the local feature map data, and the second comprehensive feature map data obtained after fusion will be passed to the decoder module for final dose prediction, thereby improving the expression ability of the dose prediction model and helping to restore more fine details and structures in the subsequent upsampling stage.
[0016] Further, the deep convolutional evolution module includes a deep convolutional layer, a normalization layer, and a pointwise convolutional layer;
[0017] The deep convolutional layer is configured to independently perform three-dimensional convolution operations on each channel of the input CT image data to extract the current local feature data on each channel;
[0018] The normalization layer is configured to perform normalization processing on the current local feature data output by the deep convolutional layer;
[0019] The pointwise convolution layer is used to mix each of the current local feature data after normalization through 1×1×1 convolution.
[0020] The beneficial effect of adopting the above further solution is that the deep convolutional evolution module can optimize the calculation efficiency by simplifying the convolution design.
[0021] Further, the sliding window transformation network module includes a first transformation sub-module and a second transformation sub-module;
[0022] The first transformation sub-module includes a W-MSA module and a first MLP module, and the second transformation sub-module includes a SW-MSA module and a second MLP module.
[0023] The beneficial effect of adopting the above further solution is that the sliding window transformation network module significantly reduces the global computational complexity and supports efficient parallelism by using windowed attention, hierarchical structure, and shifted window mechanism. The combination method of the sliding window transformation network module and the deep convolutional evolution module can effectively reduce the calculation time, enabling the dose prediction scheme based on the dose prediction model to meet the requirements of clinical real-time prediction and providing timely and accurate dose prediction.
[0024] In a second aspect, the present application provides a dose prediction method based on a dose prediction model, including:
[0025] Obtain CT image data and multiple radiation source physical information, where the radiation source physical information includes radiation source type, energy of the radiation source, irradiation angle, or irradiation position;
[0026] Input the CT image data and multiple pieces of the radiation source physical information into the trained dose prediction model, so that the dose prediction model extracts local features from the CT image data based on the deep convolutional evolution module to obtain local feature map data, fuse the local feature map data with multiple pieces of the radiation source physical information to obtain first comprehensive feature map data, and based on the first comprehensive feature map data, predict and output a three-dimensional dose distribution map, where the three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position;
[0027] Real-time feedback the three-dimensional dose distribution map to the user for evaluation and adjustment of the radiotherapy plan.
[0028] Further, predicting the three-dimensional dose distribution map based on the first comprehensive feature map data includes:
[0029] Based on the sliding window transformation network module of the dose prediction model, obtain global context feature data corresponding to the first comprehensive feature map data, where the global context feature data is a parameter characterizing the global information in the CT image data;
[0030] Fuse the global context feature data and the local feature map data to obtain second comprehensive feature map data;
[0031] Perform progressive upsampling operations on the second comprehensive feature map data to obtain the three-dimensional dose distribution map corresponding to the second comprehensive feature map data.
[0032] Further, the fusing the local feature map data with multiple radiation source physical information to obtain first comprehensive feature map data includes:
[0033] For each piece of the radiation source physical information, perform numerical processing on the radiation source physical information, and organize the numerically processed radiation source physical information into a vector;
[0034] Concatenate the vectors corresponding to the respective radiation source physical information into a multi-dimensional vector;
[0035] Concatenate the multi-dimensional vector and the local feature map data in the channel dimension to form the first comprehensive feature map data.
[0036] In a third aspect, the present application provides a dose prediction device based on a dose prediction model, including:
[0037] An input data acquisition module, configured to acquire CT image data and multiple pieces of radiation source physical information, where the radiation source physical information includes a radiation source type, an energy of the radiation source, an irradiation angle, or an irradiation position;
[0038] A model input / output module, configured to input the CT image data and multiple pieces of the radiation source physical information into a trained dose prediction model, so that the dose prediction model performs local feature extraction on the CT image data based on a deep convolutional evolution module to obtain local feature map data, fuse the local feature map data with multiple pieces of the radiation source physical information to obtain first comprehensive feature map data, and predict and output a three-dimensional dose distribution map based on the first comprehensive feature map data, where the three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position;
[0039] A real-time feedback module, configured to real-time feedback the three-dimensional dose distribution map to a user for evaluation and adjustment of a radiotherapy plan.
[0040] In a fourth aspect, the present application provides an electronic device, including a processor and a memory, where the processor is coupled to the memory;
[0041] The processor is configured to execute a computer program stored in the memory, so that the electronic device executes the method according to any one of the second aspect.
[0042] Fifth aspect, the present application provides a computer-readable storage medium, including a computer program or instruction, which, when running on a computer, causes the computer to execute the method according to any one of the second aspect. Description of the Drawings
[0043] Figure 1 It is an architecture diagram of a dose prediction model based on deep learning according to an embodiment of the present application;
[0044] Figure 2 It is a schematic structural diagram of a ConvNeXt Block according to an embodiment of the present application;
[0045] Figure 3 It is a schematic structural diagram of a Downsample according to an embodiment of the present application;
[0046] Figure 4 It is a schematic structural diagram of a first transformation sub-module and a second transformation sub-module according to an embodiment of the present application;
[0047] Figure 5 It is a schematic flowchart of a dose prediction method based on a dose prediction model according to an embodiment of the present application;
[0048] Figure 6 It is a structural block diagram of a dose prediction device based on a dose prediction model according to an embodiment of the present application;
[0049] Figure 7 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed Embodiments
[0050] The following further describes the present application in detail with reference to the drawings.
[0051] As Figures 1 to 4 shown, an embodiment of the present application provides a dose prediction model based on deep learning, including a data input module, a deep convolutional evolution module, a first fusion module, and a prediction output module.
[0052] The data input module inputs CT image data for the deep convolutional evolution module.
[0053] The deep convolutional evolution module is used to perform local feature extraction on the CT image data to obtain local feature map data.
[0054] The first fusion module is used to increase the number of channels describing the CT image data, and fuse the local feature map data with multiple radiation source physical information to obtain first comprehensive feature map data, where the radiation source physical information includes radiation source type, energy of the radiation source, irradiation angle, or irradiation position.
[0055] The prediction output module is used to predict and output a three-dimensional dose distribution map according to the first comprehensive feature map data, where the three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position.
[0056] In this embodiment, the CT image data is the three-dimensional CT image of the patient, usually obtained by scanning before radiotherapy, and the size is a three-dimensional matrix (Depth, Height, Width). The CT image data provides information on the anatomical structure in the patient's body, especially the positions and morphologies of tumors and critical organs.
[0057] The depth convolution evolution module can be a ConvNeXt module, and the ConvNeXt module is responsible for extracting local features from the CT image data. The ConvNeXt module is based on the convolutional neural network (CNN) architecture, but by introducing deep convolutional layers and adaptive normalization techniques, it improves the ability to extract details and local structures. For the case where the boundaries of tumors and organs are not clear, the ConvNeXt module can effectively capture local features and make up for the deficiencies of traditional CNNs in detail processing. Based on the improvement and design of its architecture, the ConvNeXt module is optimized at the deep convolutional layer and adaptive normalization compared with the standard convolutional neural network.
[0058] The physical information of the radiation source is crucial for dose prediction because different radiation sources and irradiation parameters will affect the dose distribution in the body. The physical information of the radiation source can include the type of radiation source in the treatment plan (e.g., X-ray, proton, heavy particle, etc.), the energy of the radiation source, the irradiation angle, the irradiation position, and other treatment parameters. The treatment plan refers to the treatment parameters corresponding to the current CT image data, and the treatment plan can be formulated by a radiotherapy physicist using a TPS and can be exported by the TPS.
[0059] The data input module is responsible for receiving and preprocessing the CT image data and the physical information of the radiation source to ensure that the CT image data and the physical information of the radiation source can be input into the deep learning network simultaneously. To improve the training and prediction accuracy of the dose prediction model, the preprocessing of the CT image data includes normalization processing, and the normalization processing can be to normalize the pixel values of the CT image data to the interval [0, 1].
[0060] The preprocessing of the physical information of the radiation source includes converting it into a vector form through numerical processing. The physical information of the radiation source is a different type of input from the anatomical structure and is usually structured non-image data. For example, the energy, type, and irradiation angle of the radiation source are usually discrete numerical or categorical data. Therefore, numerical processing can convert these non-image data into an input form that can be processed by the network and make it suitable for combination with image data.
[0061] The physical information of different radiation sources is separately organized into a vector of a fixed length, and finally these information are concatenated into a multi-dimensional vector through concatenation. This multi-dimensional vector contains all relevant physical information and can be input into subsequent network modules together with the extracted local feature map data. The numerical physical information of the radiation source will be organized into a vector to facilitate the fusion and further processing with the local feature map data in the neural network.
[0062] The physical information of the radiation source is used as an additional input and combined with the CT image data at the early stage of the network. Specifically, after inputting the CT image data, local features are first extracted through the depth convolution evolution module of the network, and then the physical information of the radiation source is fused with the extracted local feature map data.
[0063] In this embodiment, the physical information of the radiation source can be fused with the local feature map data of the CT image in the way of "feature addition". Specifically, the multi-dimensional vector corresponding to the physical information of the radiation source is concatenated with the feature map extracted by the depth convolution evolution module in the channel dimension to form a comprehensive feature representation, which can ensure that the dose prediction model takes into account both the patient's anatomical structure and the physical parameters of the treatment during the prediction process.
[0064] The concatenation process includes: the shape of the local feature map data extracted by the depth convolution evolution module can be C×H×W, the multi-dimensional vector P is converted into a tensor with the shape of L×1×1, and is extended to L×H×W through the broadcast mechanism; after concatenation along the channel dimension, it is (C + L)×H×W.
[0065] By taking the physical characteristics of the radiation source (such as radiation source type, energy distribution, irradiation angle, etc.) as part of the input data corresponding to the dose prediction model, it can accurately simulate the influence of the radiation source on the dose distribution, avoid the problem of ignoring the radiation source characteristics in the traditional method, and improve the accuracy of dose prediction.
[0066] In this embodiment, the prediction output module includes a sliding window transformation network module, a second fusion module, and a decoder module. The sliding window transformation network module is used to obtain the global context feature data corresponding to the first comprehensive feature map data, and the global context feature data is a parameter characterizing the global information in the CT image data. The second fusion module is used to fuse the global context feature data and the local feature map data to obtain the second comprehensive feature map data. The decoder module is used to gradually perform multiple upsampling operations on the second comprehensive feature map data to obtain the three-dimensional dose distribution map corresponding to the second comprehensive feature map data.
[0067] The sliding window transformation network module can be the Swin Transformer module. In this embodiment, the network uses the Swin Transformer module to perform global context modeling on the fused first comprehensive feature map data to obtain global context feature data.
[0068] The Swin Transformer module can utilize the self-attention mechanism to capture long-range spatial relationships and can handle features of different scales, thereby enhancing the dose prediction model's comprehensive understanding of complex anatomical structures (such as the relationship between tumors and vital organs) and radiation source characteristics (such as energy distribution, irradiation angle, etc.).
[0069] After the ConvNeXt module extracts local feature map data, the Swin Transformer module is responsible for performing global context modeling. The Swin Transformer module uses the windowed self-attention mechanism and can effectively capture long-range spatial relationships and multi-scale information. The Swin Transformer module is particularly suitable for dealing with the complex spatial relationships between tumors and key organs, as well as the interaction between the radiation source and tissues during the treatment process. The Swin Transformer module ensures the spatial consistency and accuracy of dose prediction, especially in the case of complex cases and multi-site radiotherapy.
[0070] By combining the ConvNeXt module and the Swin Transformer module, the ConvNeXt module extracts anatomical features from the CT image data, and the Swin Transformer module uses the anatomical features + radiation source physical information extracted in the previous step to learn the spatial distribution of the dose. Through refined tumor boundary extraction and global dose distribution prediction, it helps to optimize the dose allocation between the tumor region and normal tissues and avoid over-irradiation of key organs.
[0071] After local feature extraction and global context modeling, the second fusion module fuses the global context feature data and the local feature map data. The second comprehensive feature map data obtained after fusion will be passed to the decoder module for final dose prediction, thereby improving the expression ability of the dose prediction model and helping to restore more fine details and structures in the subsequent upsampling stage.
[0072] The decoder module uses a transposed convolution layer to gradually restore the spatial resolution of the prediction result and generate the final three-dimensional dose distribution map. Specifically, the decoder module restores the spatial resolution through a step-by-step upsampling operation and converts the second comprehensive feature map data into the final three-dimensional dose prediction map. This stage combines local features and global information, thus ensuring the accuracy and spatial consistency of the dose distribution. The dose prediction model can output a three-dimensional dose distribution map, which represents the radiotherapy dose at each spatial position and can be directly used for the verification and optimization of clinical treatment plans.
[0073] In the radiotherapy plan, by predicting the dose distribution in real time, it is provided to the radiotherapy physician for plan evaluation and adjustment. The dose prediction model can, after the treatment plan is designed, combine CT images and radiation source physical information to predict the dose distribution in real time, ensuring the optimization and precise implementation of the treatment plan.
[0074] In this embodiment, a combined architecture of the ConvNeXt module and the Swin Transformer module is adopted, which can effectively capture the details of local structures while modeling the global spatial relationship between the tumor and key organs, significantly improving the accuracy of dose prediction, especially in the case of complex treatment plans and multiple radiation sources.
[0075] In this embodiment, the deep convolutional evolution module may include multiple ConvNeXt stages, that is, multiple processing stages. Exemplarily, the deep convolutional evolution module may include ConvNeXt stage1, ConvNeXt stage 2, ConvNeXt stage 3, and ConvNeXt stage 4. Each ConvNeXt stage includes multiple ConvNeXtBlocks, and each ConvNeXt Block includes a depth convolutional layer, a normalization layer, and a pointwise convolutional layer. The depth convolutional layer is used to independently perform a three-dimensional convolution operation on each channel of the input CT image data to extract the current local feature data on each channel. The normalization layer is used to perform normalization processing on the current local feature data output by the depth convolutional layer. The pointwise convolutional layer is used to mix the normalized current local feature data through 1×1×1 convolution.
[0076] The depth convolution layer can be Depthwise Conv3d, the normalization layer can be Layer Norm, and the pointwise convolution layer can be Pointwise Conv3d. The ConvNeXt Block can also include GELU, Layer Scale, and Drop Path. GELU is the Gaussian Error Linear Unit; Layer Scale is layer scaling, which is used to improve the training of deep networks; Drop Path is the path dropout layer, which is used to randomly drop certain paths in the network to reduce overfitting and improve the generalization ability of the dose prediction model. It is easy to understand that Figure 2 the cross mark in
[0077] In this embodiment, ConvNeXt stage 1 also includes a stem, which is a 3D convolution layer. The stem can use a larger convolution kernel (such as 7×7) instead of the traditional smaller convolution kernel (such as 3×3), so as to accelerate the calculation and increase the receptive field.
[0078] The CT image data input to the dose prediction model can be a CT image of 300x300x200, and the stem can downsample the CT image data to 150x150x100. ConvNeXt stage 1 can include 3 ConvNeXt Blocks, and each ConvNeXt Block can have 96 channels. The downsampled image data (150x150x100) is used as the input of the subsequent 3 ConvNeXt Blocks, and the output feature map size of the ConvNeXt Block can still be maintained as 150x150x100 (while the number of channels increases, the spatial dimension remains unchanged).
[0079] ConvNeXt stage 2 can include Downsample and 3 ConvNeXt Blocks. The Downsample in ConvNeXt stage 2 is used to perform 3D convolution on the feature map output by ConvNeXt stage 1, and each ConvNeXt Block in ConvNeXt stage 2 can have 192 channels.
[0080] ConvNeXt stage 3 may include Downsample and 9 ConvNeXt Blocks. The Downsample in ConvNeXt stage 3 is used to perform 3D convolution on the feature map output by ConvNeXt stage 2, and each ConvNeXt Block in ConvNeXt stage 3 may have 384 channels.
[0081] ConvNeXt stage 4 may include Downsample and 3 ConvNeXt Blocks. The Downsample in ConvNeXt stage 4 is used to perform 3D convolution on the feature map output by ConvNeXt stage 3, and each ConvNeXt Block in ConvNeXt stage 4 may have 768 channels. The feature map output by ConvNeXt stage 4 is the local feature map data.
[0082] As Figure 3 shown, in this embodiment, Downsample may include a normalization layer and a three-dimensional convolutional layer. The normalization layer is Layer Norm, and the three-dimensional convolutional layer is Conv3d. The three-dimensional convolutional layer (Conv3d) is located after the normalization layer (LayerNorm).
[0083] It is easy to understand that Figure 1 the "C" in is Concatenation, that is, the number of features (channels) describing the image itself increases, but the information under each feature does not increase. The local feature map data output by ConvNeXt stage 4 can be input to the Concatenation before the sliding window transformation network module, so as to realize the process of feature fusion between the local feature map data and the physical information of the radiation source.
[0084] In this embodiment, the dose prediction model further includes Patch Partition (image block segmentation). After PatchPartition partitioning, the local feature map data is divided into multiple small local windows, and self-attention is calculated separately within each window, that is, windowed attention. The data obtained by Patch Partition dividing the image blocks is input to the sliding window transformation network module.
[0085] The sliding window transformation network module may include a Linear Embedding, a first transformation sub-module, and a second transformation sub-module. Both the first transformation sub-module and the second transformation sub-module can be represented as Swin Transformer Blocks, and the first transformation sub-module and the second transformation sub-module are arranged consecutively. The Linear Embedding is used to map the local feature map data from a high-dimensional space to a low-dimensional space while maintaining the local structure of the local feature map data.
[0086] As Figure 4 shown, the first transformation sub-module includes a W-MSA module and a first MLP module, and the second transformation sub-module includes a SW-MSA module and a second MLP module. It is easy to understand that Figure 4 the cross marks in can represent residual connections, that is, there is a residual connection between the W-MSA module and the first MLP module, and there is also a residual connection between the SW-MSA module and the second MLP module. W-MSA and SW-MSA are respectively used for local and cross-window feature extraction. Combining the first MLP module and the second MLP module can improve the expressive ability of the dose prediction model.
[0087] The first transformation sub-module can be used to extract local information, and the second transformation sub-module can enhance the flow of global information through shifted windows.
[0088] The sliding window transformation network module can construct its hierarchical structure through multiple processing stages. In each processing stage, the dose prediction model processes smaller image patches and gradually reduces the resolution of each image patch.
[0089] In the calculation of each Block of the SW-MSA module, the position of the window will be shifted between adjacent layers, that is, the position of the window is offset by a certain pixel distance, so that the features between each window can interact, rather than being completely independent, to implement the shifted window mechanism.
[0090] In this embodiment, the depth convolution evolution module optimizes the calculation efficiency by simplifying the convolution design (such as depthwise separable convolution, large convolution kernel, LayerNorm), while the sliding window transformation network module significantly reduces the global computational complexity and supports efficient parallelism by using windowed attention, hierarchical structure, and shifted window mechanism. This combination method can effectively reduce the calculation time, so that the dose prediction scheme based on the dose prediction model can meet the needs of clinical real-time prediction. Especially when it is necessary to quickly evaluate the treatment plan effect during radiotherapy, it can provide timely and accurate dose prediction.
[0091] As Figure 1As shown in the figure, in this embodiment, the second fusion module includes another Concatenation different from the Concatenation before the sliding window transformation network module. The outputs of the depth convolution evolution module and the sliding window transformation network module are connected to the second fusion module, and the local feature map data output by the depth convolution evolution module and the global context feature data output by the sliding window transformation network module are fused at the second fusion module to obtain the second comprehensive feature map data.
[0092] As Figure 1 shown, the decoder module may include multiple upsampling modules. In this embodiment, the decoder module includes 5 upsampling modules, so that the second comprehensive feature map data can be gradually upsampled 5 times to restore to the size corresponding to the CT image data at the beginning of the input, and a three-dimensional dose distribution map corresponding to the second comprehensive feature map data is obtained.
[0093] In the radiotherapy plan, by predicting the dose distribution in real time, it is provided to the radiotherapy physician for plan evaluation and adjustment. Specifically, the full-space dose prediction result and the DVH result can be displayed in the TPS for the radiotherapy physician to evaluate and adjust the plan. Therefore, through this dose prediction model, it is possible to combine the CT image data and the physical information of multiple radiation sources to predict the dose distribution in real time after the treatment plan is designed, ensuring the optimization and precise implementation of the treatment plan.
[0094] Based on the same technical concept, the embodiment of the present application also provides a dose prediction method based on the above dose prediction model. This method can be executed by a device, which can be a server or a terminal device. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a desktop computer, etc., but is not limited thereto.
[0095] As Figure 5 shown, a dose prediction method based on the above dose prediction model, with an electronic device as the execution subject, the main process of the method is described as follows (Steps S101 to S103):
[0096] Step S101: Obtain CT image data and the physical information of multiple radiation sources, where the physical information of the radiation source includes the type of radiation source, the energy of the radiation source, the irradiation angle or the irradiation position.
[0097] Step S102: Input the CT image data and the physical information of multiple radiation sources into the trained dose prediction model, so that the dose prediction model performs local feature extraction on the CT image data based on the deep convolutional evolution module to obtain local feature map data, fuse the local feature map data with the physical information of multiple radiation sources to obtain first comprehensive feature map data, and predict and output a three-dimensional dose distribution map based on the first comprehensive feature map data. The three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position.
[0098] Step S103: Real-time feedback the three-dimensional dose distribution map to the user for the evaluation and adjustment of the radiotherapy plan.
[0099] In this embodiment, predicting the three-dimensional dose distribution map based on the first comprehensive feature map data includes:
[0100] Based on the sliding window transformation network module of the dose prediction model, obtain the global context feature data corresponding to the first comprehensive feature map data. The global context feature data is a parameter characterizing the global information in the CT image data;
[0101] Fuse the global context feature data and the local feature map data to obtain second comprehensive feature map data;
[0102] Perform progressive upsampling operations on the second comprehensive feature map data to obtain the three-dimensional dose distribution map corresponding to the second comprehensive feature map data.
[0103] In this embodiment, fusing the local feature map data with the physical information of multiple radiation sources to obtain first comprehensive feature map data includes:
[0104] For each piece of physical information of the radiation source, perform numerical processing on the physical information of the radiation source and organize the numerically processed physical information of the radiation source into a vector;
[0105] Concatenate the vectors corresponding to the respective physical information of the radiation sources into a multi-dimensional vector;
[0106] Concatenate the multi-dimensional vector and the local feature map data in the channel dimension to form the first comprehensive feature map data.
[0107] Based on the same technical concept, the present application also provides a dose prediction device based on a dose prediction model, as Figure 6 shown. The dose prediction device 200 based on the dose prediction model mainly includes:
[0108] An input data acquisition module 201 for acquiring CT image data and physical information of multiple radiation sources, where the physical information of the radiation sources includes the type of the radiation source, the energy of the radiation source, the irradiation angle or the irradiation position;
[0109] A model input / output module 202 for inputting the CT image data and the physical information of multiple radiation sources into a trained dose prediction model, so that the dose prediction model performs local feature extraction on the CT image data based on a deep convolutional evolution module to obtain local feature map data, fuses the local feature map data with the physical information of multiple radiation sources to obtain first comprehensive feature map data, and predicts and outputs a three-dimensional dose distribution map based on the first comprehensive feature map data, where the three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position;
[0110] A real-time feedback module 203 for real-time feedback of the three-dimensional dose distribution map to the user for evaluation and adjustment of the radiotherapy plan.
[0111] Optionally, the model input / output module 202 includes:
[0112] A global feature acquisition sub-module for acquiring global context feature data corresponding to the first comprehensive feature map data based on a sliding window transformation network module of the dose prediction model, where the global context feature data is a parameter characterizing the global information in the CT image data;
[0113] A local and global fusion sub-module for fusing the global context feature data and the local feature map data to obtain second comprehensive feature map data;
[0114] An upsampling sub-module for performing step-by-step upsampling operations on the second comprehensive feature map data to obtain the three-dimensional dose distribution map corresponding to the second comprehensive feature map data.
[0115] Optionally, the model input / output module 202 further includes:
[0116] A numerical processing sub-module for numerically processing each piece of physical information of the radiation source and organizing the numerically processed physical information of the radiation source into a vector;
[0117] A concatenation integration sub-module for concatenating the vectors corresponding to the physical information of each radiation source into a multi-dimensional vector;
[0118] A splicing sub-module for splicing the multi-dimensional vector and the local feature map data in the channel dimension to form the first comprehensive feature map data.
[0119] In one example, the modules in any of the above devices may be one or more integrated circuits configured to implement the above methods. For example: one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0120] For another example, when the modules in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call programs. For another example, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0121] In this application, names may be assigned to various objects such as various messages / information / devices / network elements / systems / devices / actions / operations / processes / concepts, etc. It can be understood that these specific names do not constitute a limitation on the relevant objects, and the assigned names may change with factors such as the scenario, context, or usage habits. The understanding of the technical meaning of the technical terms in this application should be mainly determined from the functions and technical effects they embody / perform in the technical solution.
[0122] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0123] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0124] Based on the same technical concept, this application also provides an electronic device, as Figure 7 shown. The electronic device 300 includes a processor 301 and a memory 302, and may further include one or more of an information input / output (I / O) interface 303, a communication component 304, and a communication bus 305.
[0125] Among them, the processor 301 is used to control the overall operation of the electronic device 300 to complete all or part of the steps in the above-mentioned dose prediction method based on the dose prediction model; the memory 302 is used to store various types of data to support the operation of the electronic device 300. These data may include, for example, instructions for any application or method operating on the electronic device 300, as well as application-related data. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, one or more of a magnetic disk or an optical disk.
[0126] The I / O interface 303 provides an interface between the processor 301 and other interface modules. The above-mentioned other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 304 is used to test the wired or wireless communication between the electronic device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G or 4G, or a combination of one or more of them. Therefore, the corresponding communication component 304 can include: a Wi-Fi component, a Bluetooth component, an NFC component.
[0127] The communication bus 305 may include a path to transmit information between the above components. The communication bus 305 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 305 can be divided into an address bus, a data bus, a control bus, etc.
[0128] The electronic device 300 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute the dose prediction method based on the dose prediction model given in the above embodiments.
[0129] The electronic device 300 can include, but is not limited to, mobile terminals such as digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc., and can also be a server, etc.
[0130] Based on the same technical concept, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned dose prediction method based on the dose prediction model are realized.
[0131] The computer-readable storage medium can include various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0132] The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0133] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0134] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0135] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A dose prediction model based on deep learning, characterized in that: It includes a data input module, a deep convolution evolution module, a first fusion module and a prediction output module; The data input module inputs CT image data to the deep convolution evolution module; The deep convolution evolution module is used to extract local features from the CT image data to obtain local feature map data; The first fusion module is used to increase the number of channels describing the CT image data, fuse the local feature map data with multiple radiation source physical information, and obtain first comprehensive feature map data, wherein the radiation source physical information includes the type of radiation source, the energy of the radiation source, the irradiation angle or the irradiation position; The prediction output module is used to predict and output a three-dimensional dose distribution map based on the first comprehensive feature map data, where the three-dimensional dose distribution map is a parameter that characterizes the dose at each spatial position.
2. The dose prediction model according to claim 1, characterized in that: The prediction output module includes a sliding window transformation network module, a second fusion module and a decoder module; The sliding window transformation network module is used to obtain global context feature data corresponding to the first comprehensive feature map data, wherein the global context feature data is a parameter representing global information in the CT image data; The second fusion module is used to fuse the global context feature data and the local feature map data to obtain second comprehensive feature map data; The decoder module is used to gradually perform multiple upsampling operations on the second comprehensive feature map data to obtain the three-dimensional dose distribution map corresponding to the second comprehensive feature map data.
3. The dose prediction model according to claim 2, characterized in that: The deep convolution evolution module includes a deep convolution layer, a normalization layer and a point-by-point convolution layer; The deep convolution layer is used to independently perform a three-dimensional convolution operation on each channel of the input CT image data to extract current local feature data on each channel; The normalization layer is used to normalize the current local feature data output by the deep convolution layer; The point-by-point convolution layer is used to mix the normalized current local feature data through 1×1×1 convolution.
4. The dose prediction model according to claim 3, characterized in that: The sliding window transformation network module includes a first transformation submodule and a second transformation submodule; The first transformation submodule includes a W-MSA module and a first MLP module, and the second transformation submodule includes a SW-MSA module and a second MLP module.
5. A dose prediction method based on the dose prediction model according to any one of claims 1 to 4, characterized in that: include: Acquiring CT image data and multiple radiation source physical information, wherein the radiation source physical information includes radiation source type, radiation source energy, irradiation angle or irradiation position; Input the CT image data and the multiple radiation source physical information into a trained dose prediction model, so that the dose prediction model extracts local features of the CT image data based on a deep convolution evolution module to obtain local feature map data, fuses the local feature map data with the multiple radiation source physical information to obtain first comprehensive feature map data, and predicts and outputs a three-dimensional dose distribution map based on the first comprehensive feature map data, wherein the three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position; The three-dimensional dose distribution map is fed back to the user in real time to evaluate and adjust the radiotherapy plan.
6. The dose prediction method according to claim 5, characterized in that: The predicting of a three-dimensional dose distribution map based on the first comprehensive characteristic map data includes: Based on the sliding window transformation network module of the dose prediction model, obtaining global context feature data corresponding to the first comprehensive feature map data, wherein the global context feature data is a parameter characterizing global information in the CT image data; Fusing the global context feature data and the local feature map data to obtain second comprehensive feature map data; A step-by-step upsampling operation is performed on the second comprehensive characteristic map data to obtain the three-dimensional dose distribution map corresponding to the second comprehensive characteristic map data.
7. The dose prediction method according to claim 6, characterized in that: The fusing of the local feature map data with multiple radiation source physical information to obtain first comprehensive feature map data includes: For each piece of the radioactive source physical information, perform digital processing on the radioactive source physical information, and organize the digitally processed radioactive source physical information into a vector; The vectors corresponding to the physical information of each radiation source are connected in series to form a multi-dimensional vector; The multidimensional vector and the local feature map data are concatenated in the channel dimension to form the first comprehensive feature map data.
8. A dose prediction device based on a dose prediction model, characterized in that: include: An input data acquisition module is used to acquire CT image data and multiple radiation source physical information, wherein the radiation source physical information includes the radiation source type, the energy of the radiation source, the irradiation angle or the irradiation position; A model input and output module, used for inputting the CT image data and the multiple radiation source physical information into a trained dose prediction model, so that the dose prediction model extracts local features of the CT image data based on the deep convolution evolution module to obtain local feature map data, fuses the local feature map data with the multiple radiation source physical information to obtain first comprehensive feature map data, and predicts and outputs a three-dimensional dose distribution map based on the first comprehensive feature map data, wherein the three-dimensional dose distribution map is a parameter characterizing the dose at each spatial position; The real-time feedback module is used to provide the user with real-time feedback of the three-dimensional dose distribution map so as to evaluate and adjust the radiotherapy plan.
9. An electronic device, characterized in that: comprising a processor and a memory, wherein the processor is coupled to the memory; The processor is configured to execute the computer program stored in the memory, so that the electronic device executes the method according to any one of claims 5 to 7.
10. A computer-readable storage medium, characterized in that: The method comprises a computer program or an instruction, which, when executed on a computer, causes the computer to execute the method according to any one of claims 5 to 7.