Liver tumor radiotherapy dose prediction method based on diffusion model

By adopting a diffusion model-based method in the dose prediction of radiotherapy of liver tumors, combining the beam field and multi-condition polymerization module, the problem of lack of high-frequency details and diversity of prediction results in the prior art is solved, and high-precision and robust dose prediction are achieved.

CN120072199AActive Publication Date: 2025-05-30CENT SOUTH UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510138456.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The prior art is difficult to capture complex dose distribution characteristics in the prediction of radiotherapy dose of liver tumors, resulting in a lack of high-frequency details and diversity in the prediction results. The existing methods are unstable during the training process and are difficult to adapt to the size and position differences of liver tumors.

Method used

Using a diffusion model-based approach, a noise prediction network including an encoder, an intermediate layer and a decoder is constructed, and the beam field is used as a model input, combining a multi-conditional aggregation module and an asymmetric fusion module, the model structure and parameter configuration are optimized to improve the accuracy of dose prediction.

Benefits of technology

High-precision prediction of the radiotherapy dose of liver tumors is achieved, which improves the robustness of the model and the diversity of predicted results, and enhances the adaptability to complex morphology of liver tumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072199A_ABST
    Figure CN120072199A_ABST
Patent Text Reader

Abstract

The invention discloses a liver tumor radiotherapy dose prediction method based on a diffusion model. The implementation scheme is as follows: 1) data collection; 2) pretreatment; 3) constructing a dose prediction model; 4) constructing a loss function; 5) training a prediction model; and 6) dosage prediction. According to the invention, by introducing a beam field, radiotherapy beam direction information and dose deposition information in a beam propagation process are provided for a dose prediction model; in order to efficiently utilize input condition information, in a noise prediction network, a multi-branch encoder and a multi-condition aggregation module are designed to realize feature extraction and aggregation of different input condition information, and an asymmetric fusion module is designed to reduce information loss in an input condition information processing process. Compared with a traditional deep learning method, the method can generate a more authentic dose prediction result, facilitates the improvement of the efficiency of radiotherapy plan design and the formulation of a personalized treatment plan, and has a great practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of radiotherapy dose prediction, and particularly relates to a method for predicting the radiotherapy dose of liver tumors based on a diffusion model. Background Art

[0002] Radiotherapy is one of the main means of combating cancer. The design of radiotherapy plans is the core link in the entire radiotherapy process. Its purpose is to ensure that the irradiation dose received by the tumor target area reaches the required coverage rate and intensity, and to minimize the irradiation of surrounding normal organs. Therefore, dose prediction is a key link in radiotherapy plan design. With the rapid development of deep learning technology, many researchers have begun to explore its application in cancer radiotherapy dose prediction. Existing methods generally use pixel-level loss functions (such as L1 or L2 loss), and the generated dose distribution maps are often too smooth, lacking high-frequency details and realism. Although the method based on generative adversarial networks (GANs) alleviates the problem of overly smooth results in dose distribution prediction to a certain extent, its training process is often unstable and prone to mode collapse, resulting in a lack of diversity in the generated results. In addition, existing research mainly focuses on regions with relatively consistent tumor morphologies such as the head and neck and pelvis, where the difficulty of dose prediction is relatively low. For liver tumors with significant differences in tumor size and location, the complexity of dose distribution prediction is relatively large, and existing research is relatively limited, and further in-depth exploration is needed to improve prediction accuracy and clinical applicability. Summary of the Invention

[0003] The present invention fully considers the problems existing in the prior art. Its purpose is to propose a method for predicting the radiotherapy dose of liver tumors based on a diffusion model. By optimizing the model structure and parameter configuration, high-precision prediction of the radiotherapy dose of liver tumors is achieved.

[0004] I. Technical Principle

[0005] At present, most dose prediction algorithms use CT images, tumor target volume (PTV) contour images, and organs at risk (OAR) contour images as model inputs. Due to the significant differences in the size and location of liver tumors, it is difficult to effectively train a model to fully capture the complex characteristics of the radiotherapy dose distribution of liver tumors based solely on these inputs, making it difficult to achieve accurate dose prediction. To obtain more information to improve the accuracy of dose prediction, the present invention constructs a beam field and uses it as a model input to provide the model with beam angle information and dose deposition information during the beam propagation process in liver cancer radiotherapy. To achieve high-precision prediction of the radiotherapy dose for liver tumors, the present invention designs a noise prediction network including an encoder, an intermediate layer, and a decoder and embeds it into a diffusion model. In the encoder part, feature extraction branches are designed for different input images to achieve targeted feature extraction. To capture the associations between input images, the present invention designs a multi-condition aggregation module (MAM) and uses it in the encoder. Using a Transformer encoder and a multi-head attention mechanism, this module can aggregate the features extracted by each branch to extract the dependencies between the features. In addition, to reduce the information loss caused during the image processing process, the present invention designs an asymmetric fusion module (AFM) and uses it in the decoder to reduce the feature loss during the input image processing by combining the features of the encoder and the decoder, thereby improving the accuracy of the model's dose prediction.

[0006] II. According to the above principle, the present invention is implemented through the following solutions:

[0007] A method for predicting the radiotherapy dose of liver tumors based on a diffusion model, comprising the following steps:

[0008] (1) Data collection

[0009] Collect three-dimensional radiotherapy data of historical liver tumor cases to form a dataset, including patient CT images, contour maps of the tumor target volume PTV, contour maps of the organs at risk OAR, dose maps of intensity-modulated radiation therapy (IMRT), and dose maps of conformal radiation therapy (CRT).

[0010] (2) Preprocessing, specifically including the following steps:

[0011] (2-a) Calculate the Signed Distance Map (SDM) based on the PTV contour map and OAR contour map obtained in step (1). The calculation formula is as follows:

[0012]

[0013] where SDM(u) represents the value at voxel u in the signed distance map SDM, and (i u , j u , k u ) represents the three-dimensional coordinate value of voxel u, and (i v , j v , k v ) represents the three-dimensional coordinate value of voxel v. V represents the voxel point on the surface of the region of interest ROI, that is, the voxel point on the surface of PTV or OAR; s , s i , s j represent the voxel spacings of the voxel in the three different dimensions of i, j, and k respectively; k k is an integer and k m ∈ {-1, 0, 1}. k m = 1, m

[0014] k m = -1, k m = 0 represent that voxel u is inside, outside, and on the surface of the region of interest ROI respectively.

[0015]

[0015] (2-b) Calculate the three-dimensional beam field BF based on the CRT dose map obtained in step (1). The calculation formula is as follows:

[0016]

[0017] where BF(i, j, k) represents the value of BF at the coordinate (i, j, k); CRT(i, j, k) represents the value of the CRT dose map at the coordinate (i, j, k), CRT min represents the minimum value in the CRT dose map, and CRT max represents the maximum value in the CRT dose map.

[0018] (2-c) Divide the CT images, PTV contour maps, OAR contour maps, IMRT dose maps, CRT dose maps of each patient in the dataset, as well as the SDM map obtained in step (2-a) and the BF map obtained in step (2-b), into two-dimensional images corresponding one-to-one to the original CT slices of the patient.

[0019] (3) Build a dose prediction model, which specifically includes the following steps:

[0020] (3-a) Noise addition process, gradually adding noise to the IMRT dose map according to the following rules:

[0021]

[0022] where t is the time step, representing different stages in the noise addition process, and T represents the maximum time step of noise addition; x t represents the image after noise addition at time step t, and x 0 represents the IMRT dose image without added noise; α t represents the noise addition intensity at time step t, and ∈ t is random noise following a standard normal distribution.

[0023] (3-b) Noise removal process, gradually removing noise starting from x T and finally generating the required dose prediction map according to the following rules:

[0024]

[0025] where x' T = x T , and ∈ θ (x′ t ,t,y) represents the noise predicted by the noise prediction network from the image x t ' at time step t during the noise removal process; z t is random noise following a standard normal distribution; y represents conditional information used to guide the model to generate the required dose prediction map, and y includes CT images, PTV contour images, OAR contour images, SDM images, and BF images.

[0026] (4) Construct the loss function:

[0027] The loss function is constructed as follows:

[0028]

[0029] where, represents the average of the mean square error between the predicted noise ∈ 0 , added noise ∈ t , time step t, and conditional information y, considering all IMRT dose maps x θ (x′ t ,t,y) and ∈ t ; MSE(·) represents calculating the mean square error.

[0030] (5) Training the prediction model: Input the data obtained in step (2) into the dose prediction model constructed in step (3), train it with the loss function constructed in step (4) as the optimization objective, and update the model parameters using the AdamW optimizer until the loss no longer decreases to obtain the trained model.

[0031] (6) Dose prediction: Use the trained dose prediction model in step (5) to perform dose prediction on the data to be tested to obtain the dose prediction result.

[0032] The noise prediction network in step (3 - b) consists of an encoder, an intermediate layer, and a decoder.

[0033] The encoder consists of three branches and a multi - conditional aggregation module MAM.

[0034] Branch 1 includes a 3×3 convolutional layer and five feature extraction modules. There is a downsampling layer between every two feature extraction modules, where the downsampling layer includes a 3×3 convolutional layer; The multi - input image composed of the CT image, SDM image, PTV contour image, and OAR contour image is first processed by the 3×3 convolutional layer, and then sequentially passes through five feature extraction modules to obtain the feature maps and X K and X K which is the output of Branch 1. Among them, the first four feature extraction modules consist of two residual layers, and each residual layer includes two group normalization layers, two Sigmoid linear units, two convolutional layers, an adaptive scaling bias layer, a linear layer, and a residual connection. The fifth feature extraction module is sequentially stacked by two groups of residual layer - multi - head self - attention layer, and each group includes a residual layer and a multi - head self - attention layer.

[0035] Branch 3 has the same structure as Branch 1. The BF image is first processed by the 3×3 convolutional layer, and then sequentially passes through five feature extraction modules to obtain the feature maps and X Q and X Q which is the output of Branch 3.

[0036] Branch 2 includes a 3×3 convolutional layer and five feature extraction modules. There is an element - wise addition operation and a downsampling layer between every two feature extraction modules. The structure of the five feature extraction modules and the downsampling layer is the same as that of Branch 1; The blurred image x t ' After being processed by the 3×3 convolutional layer, it sequentially passes through five feature extraction modules. The feature maps generated by the first four feature extraction modules are respectively added element - wise to the output feature maps of the corresponding feature extraction modules in Branch 1 and Branch 3 to obtain the feature maps and Finally, the output feature map X of branch 2 is obtained through the fifth module V .

[0037] The outputs X K , X V , and X Q obtained from branches 1, 2, and 3 respectively are input into the multi-condition aggregation module MAM together to obtain the output feature map X of the encoder E .

[0038] The intermediate layer is sequentially stacked by residual layer 1, multi-head self-attention layer, and residual layer 2, where the structures of residual layer 1 and 2 are the same as those of the residual layers in the encoder; the output feature map X of the encoder E is sequentially input into residual layer 1, multi-head self-attention layer, and residual layer 2 to obtain the output feature map X of the intermediate layer M .

[0039] The decoder includes five feature extraction modules and a 3×3 convolutional layer. There is an upsampling layer and an asymmetric fusion module AFM between every two feature extraction modules. The upsampling layer contains an interpolation layer and a 3×3 convolutional layer. The structures of the five feature extraction modules are the same as those of the five feature extraction modules of branch 1 in the encoder; the output feature map X of the intermediate layer M is input into the first feature extraction module of the decoder, and its output passes through the upsampling layer to obtain the feature map D 1 . D 1 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM i together, and the obtained output is sequentially input into the second feature extraction module and the second upsampling layer to obtain the feature map D 2 . D 2 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM 2 together, and the obtained output is sequentially input into the third feature extraction module and the third upsampling layer to obtain the feature map D 3 . D 3 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM 3 together, and the obtained output is sequentially input into the fourth feature extraction module and the fourth upsampling layer to obtain the feature map D 4 . D 4 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM 4 together, and the obtained output passes through the fifth feature extraction module and the 3×3 convolutional layer in sequence to obtain the predicted noise ∈ of the noise prediction network θ(x′ t , t, y).

[0040] The multi - condition aggregation module MAM includes data chunking and flattening operations, positional encoding embedding operations, a Transformer encoder, linear transformation and reshaping operations; among which, the flattening operation includes a linear layer and a flattening layer, and the positional encoding in the positional encoding embedding operation is randomly initialized and can be dynamically learned through network training; the output feature maps X Q , X K and X V of the three branches in the encoder are respectively obtained as feature maps X' Q , X' K and X' V after passing through the chunking and flattening operations and the positional encoding embedding operations. X' Q , X' K and X' V are input into the Transformer encoder together to obtain the feature map X, and the feature map X is obtained as the output feature map X E of MAM after linear transformation and reshaping operations.

[0041] The asymmetric fusion module AFM contains two identical - structured paths, and each path includes a 3×3 convolutional layer, an instance normalization layer and a Sigmoid unit; the feature map obtained from branch 2 in the encoder and the feature map D 5-i obtained in the decoder are respectively input into the two paths, where i = 1, 2, 3, 4. The feature maps processed by the two paths are added pixel - by - pixel and concatenated with D 5-i to obtain the output feature map of the asymmetric fusion module AFM i .

[0042] Compared with the prior art, the present invention has the following advantages:

[0043] (1) The present invention introduces the beam field into the model, enabling doctors to interact with the model, which helps to quickly obtain a dose prediction map that meets clinical standards and improves the efficiency and quality of plan design.

[0044] (2) The multi - condition aggregation module designed by the present invention obtains data sequences by chunking data, so as to capture the relationship between the input conditional information and the blurred images at each time step using the attention mechanism, which helps the noise prediction network generate more accurate prediction results and improves the robustness of the model.

[0045] (3) The asymmetric fusion module designed in the present invention integrates the features of multi-scale semantic information from the encoder into the decoder, preventing the loss of key information such as the shape, position, and boundary of the tumor target area and organs at risk during the downsampling process, facilitating the rapid convergence of the model, and improving the accuracy of dose prediction. Description of the Drawings

[0046] Figure 1 Flowchart of a method for predicting radiotherapy dose of liver tumors based on a diffusion model according to an embodiment of the present invention;

[0047] Figure 2 Structure diagram of the dose prediction model according to an embodiment of the present invention;

[0048] Figure 3 Structure diagram of the noise prediction network according to an embodiment of the present invention;

[0049] Figure 4 Structure diagram of the multi-condition aggregation module according to an embodiment of the present invention;

[0050] Figure 5 Structure diagram of the asymmetric fusion module according to an embodiment of the present invention;

[0051] Figure 6 Comparison chart of the dose prediction results of the present invention and the prediction results of other methods. Detailed Embodiments

[0052] The following describes the detailed embodiments of the present invention:

[0053] Example 1

[0054] Figure 1 Shown is a flowchart of a method for predicting radiotherapy dose of liver tumors based on a diffusion model according to an embodiment of the present invention. The specific steps are as follows:

[0055] Step 1, data collection:

[0056] Collect three-dimensional radiotherapy data of historical liver tumor cases to form a dataset, including patient CT images, contour maps of the tumor target area PTV, contour maps of organs at risk OAR, intensity-modulated radiotherapy IMRT dose maps, and conformal radiotherapy CRT dose maps.

[0057] Step 2, preprocessing, which specifically includes the following steps:

[0058] (2-a) Calculate the signed distance map SDM according to the PTV contour map and OAR contour map obtained in Step 1. The calculation formula is:

[0059]

[0060] Among them, SDM(u) represents the value at voxel u in the signed distance map SDM, and (i u , j u , k u ) represents the three-dimensional coordinate value of voxel u, and (i v , j v , k v ) represents the three-dimensional coordinate value of voxel v. v represents the voxel point on the surface of the region of interest ROI, that is, the voxel point on the surface of PTV or OAR; s , s i , s j , s k respectively represent the voxel spacings of the voxel in the three different dimensions of i, j, and k; k m is an integer and k m ∈{-1, 0, 1}, k m =1,

[0061] k m =-1, k m =0 respectively represent that voxel u is inside, outside, and on the surface of the region of interest ROI.

[0062] (2-b) Calculate the three-dimensional beam field BF according to the CRT dose map obtained in step 1. The calculation formula is:

[0063]

[0064] Among them, BF(i, j, k) represents the value of BF at the coordinate (i, j, k); CRT(i, j, k) represents the value of the CRT dose map at the coordinate (i, j, k), and CRT min represents the minimum value in the CRT dose map, and CRT max represents the maximum value in the CRT dose map.

[0065] (2-c) Divide the CT image, PTV contour map, OAR contour map, IMRT dose map, CRT dose map of each patient in the dataset, and the SDM map obtained in step (2-a) and the BF map obtained in step (2-b) into two-dimensional images corresponding one by one to the original CT slices of the patient. All images are sampled to a size of 128×128.

[0066] Step 3, construct a dose prediction model

[0067] Figure 2 The structure diagram of the dose prediction model according to the embodiment of the present invention is shown. This model mainly includes a noise addition process and a denoising process.

[0068] Noise addition process: Gradually add noise to the IMRT dose map according to the following rules:

[0069]

[0070] Where t is the time step, representing different stages in the noise addition process, and T represents the maximum time step of noise addition; x t represents the image after noise addition at time step t, and x 0 represents the IMRT dose image without added noise; α t represents the noise addition intensity at time step t, and ∈ t is a random noise obeying the standard normal distribution.

[0071] Denoising process: Starting from x T , the noise is gradually removed, and finally the required dose prediction map is generated. The rules are as follows:

[0072]

[0073] Where x' T = x T , and ∈ θ (x′ t , t, y) represents the noise predicted by the noise prediction network from the image x′ in the denoising process at time step t t ; z t is a random noise obeying the standard normal distribution; y represents the conditional information used to guide the model to generate the required dose prediction map. y includes CT images, PTV contour images, OAR contour images, SDM images, and BF images.

[0074] Figure 3 The following shows the structural diagram of the noise prediction network according to the embodiment of the present invention. This network is composed of an encoder, an intermediate layer, and a decoder.

[0075] The encoder is composed of three branches and a multi - conditional aggregation module MAM;

[0076] Branch 1 includes a 3×3 convolutional layer and five feature extraction modules. There is a downsampling layer between every two feature extraction modules. Among them, the downsampling layer includes a 3×3 convolutional layer; The multi - channel input image with a size of 17×128×128 composed of CT images, SDM images, PTV contour images, and OAR contour images first undergoes processing by the 3×3 convolutional layer, and then successively passes through five feature extraction modules to obtain feature maps and X K , X KThe output of Branch 1 with a size of 512×8×8. Among them, the first four feature extraction modules consist of two residual layers. Each residual layer includes two normalization layers, two Sigmoid linear units, two 3×3 convolutional layers, an adaptive scaling bias layer, a linear layer, and a residual connection. Here, the normalization operation uses group normalization to enable the model to be effectively trained on small batches of data. The residual connection is incorporated in the block to avoid the problem of gradient vanishing. The fifth feature extraction module is sequentially stacked by two groups of residual layer - multi - head self - attention layer. Each group includes a residual layer and a multi - head self - attention layer with 4 heads.

[0077] Branch 3 has the same structure as Branch 1. The size of the BF image is 1×128×128. First, it is processed by a 3×3 convolutional layer, and then sequentially passes through five feature extraction modules to obtain feature maps and X Q ,X Q The output of Branch 3 with a size of 512×8×8.

[0078] Branch 2 includes a 3×3 convolutional layer and five feature extraction modules. Between every two feature extraction modules, there is an element - wise addition operation and a downsampling layer. The structures of the five feature extraction modules and the downsampling layer are the same as those of Branch 1; The size of the blurred image x t ' is 1×128×128. After being processed by a 3×3 convolutional layer, it sequentially passes through five feature extraction modules. Among them, the feature maps generated by the first four feature extraction modules are respectively added element - wise to the output feature maps of the corresponding feature extraction modules in Branch 1 and Branch 3, and the feature maps and with sizes of 128×128×128, 128×64×64, 256×32×32, and 384×16×16 are obtained respectively. Finally, the output feature map X V of Branch 2 with a size of 512×8×8 is obtained through the fifth module.

[0079] The outputs X K 、X V 、X Q obtained from Branch 1, 2, and 3 respectively are input into the multi - condition aggregation module MAM together to obtain the output feature map X E of the encoder with a size of 512×8×8.

[0080] The intermediate layer is sequentially stacked by residual layer 1, multi - head self - attention layer, and residual layer 2. The structures of residual layer 1 and 2 are the same as those of the residual layers in the encoder; The output feature map X E of the encoder is sequentially input into residual layer 1, multi - head self - attention layer, and residual layer 2 to obtain the output feature map X MThe size is 512×8×8.

[0081] The decoder includes five feature extraction modules and a 3×3 convolutional layer. There is an upsampling layer and an asymmetric fusion module AFM between every two feature extraction modules. The upsampling layer contains an interpolation layer and a 3×3 convolutional layer. The structures of the five feature extraction modules are the same as those of the five feature extraction modules in branch 1 of the encoder. The output feature map X of the intermediate layer M is input into the first feature extraction module of the decoder, and the output after passing through the upsampling layer is the feature map D 1 , with a size of 512×16×16. D 1 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM i together, and the obtained output is successively input into the second feature extraction module and the second upsampling layer to obtain the feature map D 2 , with a size of 384×32×32. D 2 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM 2 together, and the obtained output is successively input into the third feature extraction module and the third upsampling layer to obtain the feature map D 3 , with a size of 256×64×64. D 3 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM 3 together, and the obtained output is successively input into the fourth feature extraction module and the fourth upsampling layer to obtain the feature map D 4 , with a size of 128×128×128. D 4 and the feature map obtained from encoder branch 2 are input into the asymmetric fusion module AFM 4 together, and the obtained output successively passes through the fifth feature extraction module and the 3×3 convolutional layer to obtain the predicted noise ∈ of the noise prediction network θ (x′ t , t, y), with a size of 1×128×128.

[0082] Figure 4 The figure shows the structure diagram of the multi-condition aggregation module according to the embodiment of the present invention, including data chunking and flattening operations, position encoding embedding operations, a Transformer encoder, linear transformation and reshaping operations. The flattening operation includes a linear layer and a flattening layer. The position encoding in the position encoding embedding operation is randomly initialized and can be dynamically learned through network training. The output feature maps X Q 、X K and X of the three branches in the encoderV After the slicing operation, each feature map consists of 4 non-overlapping blocks. The flattening operation first unfolds each block to obtain a vector with a dimension size of 8192, and then performs a linear transformation on it to change the dimension to 1024. To preserve the spatial information, the position encoding of each block is embedded into the vector to obtain the feature map X'. Q 、X' K and X' V , X' Q 、X' K and X' V are input into the Transformer encoder together to obtain the feature map X. The feature map X undergoes a linear transformation and a reshaping operation to obtain the output feature map X of the MAM E .

[0083] Figure 5 The structure diagram of the asymmetric fusion module according to the embodiment of the present invention is shown. The asymmetric fusion module includes two identical paths, and each path includes a 3×3 convolutional layer, an instance normalization layer, and a Sigmoid unit; the feature map obtained from branch 2 in the encoder 5-i and the feature map D obtained from the decoder 5-i are respectively input into the two paths, where i = 1, 2, 3, 4. The feature maps processed by the two paths are added pixel by pixel and concatenated with D i to obtain the output feature map of the asymmetric fusion module AFM

[0084] Step 4, construct the loss function

[0085] The loss function is constructed as follows:

[0086]

[0087] where represents the average of the mean square error between the predicted noise ∈ 0 、 adding noise ∈ t 、 time step t and conditional information y, considering all IMRT dose maps x without added noise θ (x′ t , t, y) and ∈ t ; MSE(·) represents calculating the mean square error.

[0088] Step 5, train the prediction model

[0089] The data obtained in step 2 is input into the dose prediction model constructed in step 3, trained with the loss function constructed in step 4 as the optimization objective, and the model parameters are updated using the AdamW optimizer until the loss no longer decreases, obtaining the trained model.

[0090] Step 6, Dose Prediction

[0091] Use the dose prediction model trained in Step 5 to perform dose prediction on the data to be tested, and obtain the dose prediction result.

[0092] Example 2

[0093] Use the method in Example 1 to perform a dose prediction experiment on the private liver cancer dataset. Our network implementation is based on PyTorch, and the model is trained on a server equipped with three RTX 2080Ti GPUs. The maximum time step T for adding noise is set to 1000.

[0094] In this embodiment, the Mean Absolute Error (MAE), HI (Homogeneity Index, HI), CI (Conformity Index, CI), D x , V x , D mean , D max Seven indicators are used to conduct experiments on six dose prediction networks, namely U-Net, DeepLabV3+, HD-Net2D, C3D, DoseDiff, and MD-Dose, and the present invention on the private liver cancer dataset.

[0095] MAE is defined as the average of the sum of the absolute values of the differences between all predicted values and true values, and the calculation formula is:

[0096]

[0097] where n represents the total number of pixels, represents the true dose value of pixel i. represents the predicted dose value of pixel i. represents the absolute error between the predicted value and the true value.

[0098] D x represents the minimum absorbed dose covering x% of the PTV or OAR volume. V x represents the percentage of the volume within the PTV or OAR where the absorbed dose is greater than or equal to a specific dose threshold x in the total volume of the region. D mean represents the average dose absorbed by the PTV or OAR. D max represents the highest dose absorbed by any voxel in the PTV or OAR.

[0099] The experimental results are shown in Tables 1 and 2. Table 1 shows the comparison of the indicator MAE, and Table 2 shows the indicators HI, CI, D x , V x , Dmean , D max The comparison with D. It can be found that, compared with other methods, in the 23 evaluation indicators for PTV and OARs in total of the present invention, 16 evaluation indicators reach the optimum.

[0100] Figure 6 The figure shows the comparison chart of the dose prediction results of the embodiment of the present invention and the dose prediction results of other methods. Dose distribution prediction Figure 1 and dose distribution prediction Figure 2 are the results obtained from the test data, and the differences Figure 1 and the differences Figure 2 respectively represent the differences between the dose distribution prediction Figure 1 and the dose distribution prediction Figure 2 and their IMRT manual planning results. The grayscale bar chart on the far right represents the correspondence between the grayscale in the figure and the numerical value. It can be seen that the prediction error of the method proposed by the present invention is the smallest, and the beam shape in the prediction result is clear, and the prediction of high-frequency details is relatively accurate.

[0101] The above embodiments are only the preferred embodiments of the present invention, and do not limit the implementation scope of the present invention. Therefore, all changes made according to the structure and principle of the present invention should be covered within the protection scope of the present invention.

[0102] Table 1

[0103]

[0104] Table 2

[0105]

Claims

1. A method for predicting liver tumor radiotherapy dose based on diffusion model, characterized in that The following steps are involved: (1) Data Collection Collect historical 3D radiotherapy data of liver tumor cases to form a data set, including patient CT images, contour images of tumor target volume PTV, contour images of organs at risk OAR, dose images of intensity modulated radiotherapy IMRT, and dose images of conformal radiotherapy CRT; (2) Pretreatment, specifically including the following steps: (2-a) Calculate the signed distance map SDM based on the PTV contour map and OAR contour map obtained in step (1), and the calculation formula is: Where SDM(u) represents the value of voxel u in the signed distance map SDM, (i u ,j u ,k u ) represents the three-dimensional coordinate value of voxel u, (i v ,j v ,k v ) represents the three-dimensional coordinate value of voxel v, and v represents the surface of the region of interest ROI voxel point, that is, the voxel point on the surface of PTV or OAR; i 、s j 、s k Respectively represent the voxel spacing in three different dimensions: i, j, and k; k m is an integer and k m ∈{-1,0,1}, k m =1, k m =-1, k m = 0 means that the voxel u is located inside, outside and on the surface of the region of interest ROI, respectively; (2-b) Calculate the three-dimensional beam field BF based on the CRT dose map obtained in step (1). The calculation formula is: Where BF(i,j,k) represents the value of BF at coordinate (i,j,k); CRT(i,j,k) represents the value of CRT dose map at coordinate (i,j,k). min Indicates the minimum value in the CRT dose map, CRT max Indicates the maximum value in the CRT dose map; (2-c) dividing the CT image, PTV contour image, OAR contour image, IMRT dose image, CRT dose image of each patient in the data set, as well as the SDM image obtained in step (2-a) and the BF image obtained in step (2-b) into two-dimensional images corresponding to the original CT slices of the patient; (3) Constructing a dose prediction model, which specifically includes the following steps: (3-a) Noise addition process: gradually add noise to the IMRT dose map. The rules are as follows: Where t is the time step, representing the different stages of the noise addition process, T represents the maximum time step of the noise addition; x t represents the image after adding noise at time step t, x0 represents the IMRT dose image without adding noise; α t represents the noise intensity at time step t, ∈ t is random noise that follows a standard normal distribution; (3-b) Denoising process, from x T Start to gradually remove the noise and finally generate the required dose prediction map. The rules are as follows: Among them, x' T =x T ,∈ θ (x' t ,t,y) represents the image x' from the denoising process at time step t. t The predicted noise; z t is random noise that obeys the standard normal distribution; y represents conditional information, which is used to guide the model to generate the required dose prediction map. y includes CT images, PTV contour images, OAR contour images, SDM images, and BF images; (4) Construct loss function: The loss function is constructed as follows: in, It means considering all IMRT dose maps x0 without adding noise, adding noise ∈ t , time step t and conditional information y, the prediction noise ∈ θ (x' t ,t,y) and∈ t The average value of the mean square error between ; MSE(·) means to find the mean square error; (5) Training the prediction model: inputting the data obtained in step (2) into the dose prediction model constructed in step (3), training with the loss function constructed in step (4) as the optimization target, and updating the model parameters using the AdamW optimizer until the loss no longer decreases, thereby obtaining a trained model; (6) Dose prediction: Use the dose prediction model trained in step (5) to perform dose prediction on the data to be tested to obtain a dose prediction result.

2. A method for predicting liver tumor radiotherapy dose based on a diffusion model as claimed in claim 1, characterized in that The noise prediction network in step (3-b) is composed of an encoder, an intermediate layer and a decoder; The encoder consists of three branches and a multi-condition aggregation module MAM; Branch 1 includes a 3×3 convolutional layer and five feature extraction modules. There is a downsampling layer between every two feature extraction modules, where the downsampling layer includes a 3×3 convolutional layer. The multi-channel input image consisting of CT images, SDM images, PTV contour images, and OAR contour images is first processed by the 3×3 convolutional layer, and then passes through five feature extraction modules in sequence to obtain feature maps. and X K And X K That is the output of branch 1, where the first four feature extraction modules are composed of two residual layers, each of which includes two group normalization layers, two Sigmoid linear units, two convolutional layers, an adaptive scaling bias layer, a linear layer, and a residual connection. The fifth feature extraction module is composed of two groups of residual layers-multi-head self-attention layers stacked in sequence, each group including a residual layer and a multi-head self-attention layer; Branch 3 has the same structure as branch 1. The BF image is first processed by a 3×3 convolutional layer, and then passes through five feature extraction modules in sequence to obtain feature maps. and X Q , and X Q This is the output of branch 3; Branch 2 includes a 3×3 convolutional layer and five feature extraction modules. Each feature extraction module includes a pixel-by-pixel addition operation and a downsampling layer. The structure of the five feature extraction modules and the downsampling layer is the same as that of branch 1. Blurred image x' t After being processed by the 3×3 convolutional layer, it passes through five feature extraction modules in sequence. The feature maps generated by the first four feature extraction modules are added pixel by pixel with the output feature maps of the corresponding feature extraction modules in branch 1 and branch 3, respectively, to obtain feature maps in sequence. and Finally, the output feature map X of branch 2 is obtained through the fifth module. V ; The output X obtained from branches 1, 2, and 3 respectively K , X V , X Q Input the multi-condition aggregation module MAM together to obtain the output feature map X of the encoder E ; The intermediate layer is composed of a residual layer 1, a multi-head self-attention layer and a residual layer 2 stacked in sequence, wherein the structures of the residual layers 1 and 2 are consistent with the residual layers in the encoder; the output feature map X of the encoder is E Input residual layer 1, multi-head self-attention layer and residual layer 2 in sequence to obtain the output feature map X of the intermediate layer M ; The decoder includes five feature extraction modules and a 3×3 convolution layer. Each two feature extraction modules include an upsampling layer and an asymmetric fusion module AFM, wherein the upsampling layer includes an interpolation layer and a 3×3 convolution layer. The structure of the five feature extraction modules is the same as that of the five feature extraction modules of branch 1 in the encoder. The output feature map X of the intermediate layer is converted into M Input the first feature extraction module of the decoder, and its output is passed through the upsampling layer to obtain the feature map D1. D1 and the feature map obtained in the encoder branch 2 are combined. Enter the Asymmetric Fusion Module AFM i The output is input into the second feature extraction module and the second upsampling layer in turn to obtain feature map D2. D2 and the feature map obtained in encoder branch 2 are combined. The output is input into the asymmetric fusion module AFM2, and the output is input into the third feature extraction module and the third upsampling layer in turn to obtain the feature map D3. D3 and the feature map obtained in encoder branch 2 are combined. The output is input into the asymmetric fusion module AFM3, and the output is input into the fourth feature extraction module and the fourth upsampling layer in turn to obtain the feature map D4. D4 and the feature map obtained in the encoder branch 2 are combined. The output is input into the asymmetric fusion module AFM4 together, and the output is passed through the fifth feature extraction module and the 3×3 convolution layer in turn to obtain the predicted noise ∈ θ (x' t ,t,y).

3. A method for predicting liver tumor radiotherapy dose based on diffusion model as claimed in claim 1, characterized in that The multi-condition aggregation module MAM includes data slicing and flattening operations, positional encoding embedding operations, Transformer encoders, linear transformation and reshaping operations; the flattening operation includes a linear layer and an unfolding layer, and the positional encoding in the positional encoding embedding operation is randomly initialized and can be dynamically learned through network training; the output feature maps X of the three branches in the encoder are Q , X K and X V After the slicing and flattening operations and the position encoding embedding operations, the feature map X' is obtained. Q , X' K and X' V , X' Q , X' K and X' V Input the Transformer encoder together to obtain the feature map X, and the feature map X is linearly transformed and reshaped to obtain the output feature map X of MAM E .

4. A method for predicting liver tumor radiotherapy dose based on diffusion model as claimed in claim 1, characterized in that The asymmetric fusion module AFM contains two pathways with the same structure, each of which includes a 3×3 convolutional layer, an instance normalization layer, and a Sigmoid unit; the feature map obtained by branch 2 in the encoder is And the feature map D obtained in the decoder 5-i Input two channels respectively, where i = 1, 2, 3, 4, add the feature maps processed by the two channels pixel by pixel, and add them to D 5-i Asymmetric fusion module AFM i The output feature map of .

Citation Information

Patent Citations

  • Liver tumor radiotherapy dose prediction method based on deep learning

    CN116870377A

  • Low-dose CT image denoising method and system based on fast diffusion model

    CN117952850A

  • Zero-sample low-dose CT image denoising method and device based on strip diffusion model

    CN118761929A

  • Automatic prediction method for dose distribution in radiotherapy plan based on improved FiT

    CN119215342A

  • Automatic prediction method for dose distribution in brachytherapy plan

    CN119295393A