Nasopharyngeal carcinoma radiotherapy dose automatic prediction method based on deep learning
The deep learning-based method for nose and throat cancer radiation dose prediction addresses inefficiencies by using a modified UNet architecture with VMamba modules and autoencoders, achieving fast and accurate dose prediction for improved treatment planning.
Patent Information
- Application Number
- CN202510744752.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing radiotherapy dose prediction system for nasopharyngeal carcinoma is slow to generate, low quality, low data processing efficiency, and insufficient calculation efficiency and accuracy, resulting in large prediction errors and difficult to widely use.
The variant network designed based on UNet and VMamba was used to automatically predict the dose of nasopharyngeal carcinoma radiation therapy, and the dose prediction was optimized by batch preprocessing and deep learning model training, and the autoencoder network was used to optimize the dose prediction, combined with the PyTorch framework and NVIDIA GPU for training optimization.
It realizes rapid generation of high-quality dose distributions, improves the degree of automation of radiotherapy plans, improves computing performance and prediction accuracy, reduces prediction errors, improves patient treatment effects and saves medical resources.
Smart Images

Figure CN120318220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nasopharyngeal carcinoma radiotherapy, and specifically to an automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning. Background Art
[0002] Nasopharyngeal carcinoma is a common head and neck malignant tumor. Due to the large number of critical organs in the head and neck, radiotherapy is the main treatment method for nasopharyngeal carcinoma. Before radiotherapy, it is necessary to automatically predict the radiotherapy dose. The general scheme for automatic prediction of radiotherapy dose is to input CT images and anatomical structures into a CNN to predict the dose distribution. The research significance lies in obtaining the expected radiotherapy dose distribution and helping radiotherapy physicists design radiotherapy plans. To improve the convenience of prediction, deep learning is a machine learning method based on artificial neural networks, which processes and analyzes data by simulating the structure and function of the human brain. The radiotherapy dose prediction of nasopharyngeal carcinoma can be realized based on deep learning.
[0003] The Chinese patent discloses a multi-objective dose prediction system for advanced nasopharyngeal carcinoma based on deep learning (publication number CN119252418A). This patented technology can improve the accuracy and personalization of nasopharyngeal carcinoma treatment dose prediction, enhance the automatic contouring accuracy and efficiency, optimize the clinical decision-making process, etc., and promote the development of related technologies. However, the above dose prediction system requires complex manual design rules, has a slow generation speed and low generation quality, is not conducive to prediction efficiency and accuracy, the extracted data is relatively messy, affecting the data processing efficiency, the input data and the reconstructed data are prone to errors, and gradient disappearance often occurs during the training process, hindering its wider application in general sequence modeling. Conventional hardware calculation methods have low calculation efficiency in dose prediction, do not delete the loss amount during prediction, and have large data errors after prediction, which is not conducive to the accuracy of prediction. Therefore, the technical personnel in this field provide an automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning to solve the problems raised in the above background art. Summary of the Invention
[0004] The purpose of the present invention is to provide an automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: An automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning, the steps are as follows: S1. Analyze the obtained original data; S2. Write a program to batch preprocess all the original data; S3. Input the preprocessed CT images into a variant network designed based on UNet and VMamba. The variant network uses the UNet encoder-decoder as the main architecture and embeds a VMamba module between the 3rd and 5th layers of the encoder. This module includes a position encoding layer, three parallel dilated convolution branches with dilation rates of 1 / 3 / 5 respectively, and a gating mechanism layer. The variant network inputs CT images with a size of 256×256×64 pixels and outputs a three-dimensional dose distribution matrix. S4. Under the same CT dataset of 200 nasopharyngeal carcinoma patients and the NVIDIA A100 GPU environment, compare the prediction results of the variant network with the existing dose prediction method based on 3D-CNN. The comparison metrics include dose prediction MAE, RMSE, and clinical compliance rate. When MAE ≤ 3 Gy and the clinical compliance rate ≥ 90%, verify the feasibility of this solution.
[0006] As a further solution of the present invention: The data in the S1 step includes the data parameters of the brainstem, left and right eyeballs, left and right lenses, larynx, mandible, optic chiasm, left and right optic nerves, left and right parotid glands, oral cavity, TMJ, submandibular gland, left and right middle ears, left and right inner ears, pituitary gland, spine, thyroid gland, and temporal lobe.
[0007] As a further solution of the present invention: The steps of batch preprocessing in the S2 step are as follows: S021. Format conversion: Use a 3D network structure to directly predict the three-dimensional dose distribution of patients, and convert the patient's DICOM image sequence into a three-dimensional NIFTI format for subsequent processing. S022. Extract anatomical contours: The information of organs at risk and target area delineation is stored in the RTSTRUCT file. Read each delineated contour as a separate binary image. Finally, integrate all the delineated contours into a file, and assign file names of 1, 2, 3, …, n to each organ at risk. The PTV contour is filled with its corresponding prescribed dose. S023. Extract dose distribution: The radiotherapy dose distribution of patients is stored in the RTDOSE file. First, convert the RD file into the NIFTI format, where the dose value is the product of the pixel value and the Dose Grid Scaling. Finally, resample the RTDOSE file to the same pixel spacing and shape as the CT image to ensure the alignment of the CT image and the dose distribution. S024. Resampling: The voxel spacing of different patients is different. For the dose prediction task, it is related to the pixel distance. Resample all the CT images and anatomical contours to 1mm×1mm×3mm. Use B-spline interpolation for CT images and nearest neighbor interpolation for anatomical binary contour images. S025. Image Cropping: All images are cropped to 256×256. In the depth direction, slices that do not contain contour information are cropped off, and finally the size of all images is adjusted to 128×256×256; S026. Data Normalization: The CT images are normalized using the maximum-minimum method, and the dose distribution is not normalized.
[0008] As a further aspect of the present invention: The content of calculating the corresponding dose volume histogram in the S3 step further includes: Training an autoencoder network based on the DVH matrix, which reduces the dimension of the DVH matrix and extracts the depth features of the DVH. After the autoencoder network is trained, a loss function is designed based on the extracted DVH deep learning features to optimize the model learning.
[0009] As a further aspect of the present invention: The deep learning model in the S4 step is built based on the PyTorch framework and trained on 2 NVIDIA GeForce 3090 GPUs, optimized using the Adam optimizer, with a learning rate of 1e-4 and a batch size of 1. The model is trained for 300 epochs to reach convergence.
[0010] As a further aspect of the present invention: The original feature extraction module in the UNet in the S4 step is uniformly replaced with a newly designed Mamba block, while retaining its original U-shaped structure and skip connections, which can take into account both local features and long-distance sequence features to achieve dose prediction.
[0011] Compared with the prior art, the beneficial effects of the present invention are: 1. The automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning in the present invention automatically performs dose prediction through a deep learning model without manual design of rules, and can quickly generate a high-quality dose distribution. Its high efficiency and accuracy improve the automation level of radiotherapy plans, which helps to improve the treatment effect of patients and save medical resources; 2. By batch preprocessing all the original data, it avoids the disappearance of gradients during the training process, enabling its wider application in general sequence modeling. By inputting it into the designed deep learning model for training and prediction of dose distribution, compared with conventional hardware calculation methods, it improves the computing performance, enhances the training and inference efficiency, deletes the loss amount according to the anatomical contour and patient dose distribution to calculate the corresponding dose volume histogram (DVH), reduces the data error after prediction, and improves the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a design diagram for the nasopharyngeal carcinoma radiotherapy dose prediction scheme; Figure 2The detailed network structure of UMDose on the left and the Mamba Block designed for this project on the right are shown in the figure; Figure 3 Feature extraction graph for stacked denoising autoencoders. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0014] Example The automatic prediction method of nasopharyngeal carcinoma radiotherapy dose based on deep learning is as follows: S1. Analyze the acquired raw data, including the parameters of brainstem, left and right eyeballs, left and right lenses, larynx, mandible, optic chiasm, left and right optic nerves, left and right parotid glands, oral cavity, TMJ, mandibular gland, left and right middle ears, left and right inner ears, pituitary gland, spine, thyroid gland and temporal lobe; S2. Write a program to batch preprocess all the raw data. The steps of batch preprocessing are as follows: S021. Format conversion: Use 3D network structure to directly predict the patient's 3D dose distribution and convert the patient's DICOM image sequence into 3D NIFTI format for subsequent processing; S022. Extract anatomical contours: The outline information of organs at risk and target areas is stored in the RTSTRUCT file. Each outline is read as a separate binary image, where 1 represents the foreground and 0 represents the background. 15 organs at risk that are of clinical concern are selected for model training. For details, see Figure 1 Design of radiotherapy dose prediction plan for nasopharyngeal carcinoma,Finally, all the outlines are integrated into one file, each organ at risk is given a file name of 1, 2, 3, …, n, and the PTV outline is filled with its corresponding prescription dose; S023. Extract dose distribution: The patient's radiotherapy dose distribution is stored in the RTDOSE file. First, convert the RD file to NIFTI format, where the dose value is the product of the pixel value and the Dose Grid Scaling. Finally, resample the RTDOSE file to the same pixel spacing and shape as the CT image to ensure that the CT image is aligned with the dose distribution. S024, Resampling: The voxel spacings of different patients are different. For the dose prediction task, which is related to the pixel distance, all CT images and anatomical contours are resampled to 1mm×1mm×3mm. B-spline interpolation is used for CT images, and nearest-neighbor interpolation is used for anatomical binary contour images; S025, Image Cropping: All images are cropped to 256×256. In the depth direction, the slices that do not contain contour information are cropped off. Finally, the sizes of all images are adjusted to 128×256×256; S026, Data Normalization: The CT images are normalized using the maximum-minimum method, and the dose distribution is not normalized; S3. Calculate the corresponding dose-volume histogram (DVH) based on the anatomical contour and the patient's dose distribution. The content of calculating the corresponding dose-volume histogram also includes: training an autoencoder network based on the DVH matrix, which reduces the dimension of the DVH matrix and extracts the depth features of the DVH. After the autoencoder network is trained, a loss function is designed based on the extracted DVH deep learning features to optimize model learning; S4. Input the preprocessed data into the designed deep learning model for training and predicting the dose distribution, and compare it with the existing dose prediction methods to verify the innovation and feasibility of this solution. The content of predicting the dose distribution includes: implementing dose prediction based on the state space sequence prediction model of Mamba. Using the Mamba network, the long sequence spatial relationship of the input image can be effectively extracted. For the radiotherapy dose automatic prediction task, it has a strong dependence on the relationship between each organ at risk and the anatomical contour and the dose distribution. The deep learning model is built based on the PyTorch framework and trained on 2 NVIDIA GeForce 3090 GPUs. The Adam optimizer is used for optimization, the learning rate is 1e-4, the batchsize is 1, and the model is trained for 300 epochs to reach convergence. Its overall structure is as Figure 2 shown.
[0015] The content of calculating the corresponding dose-volume histogram also includes: Stacked Denoising Autoencoder: Step 1: The length of the DVH matrix is 15×320. After passing through each autoencoder, the vector dimension is reduced by half. After passing through seven autoencoders, the dimension is 45. Subsequently, after passing through seven decoder modules, the dimension is restored to the original feature dimension and the reconstruction loss between the input and output is calculated; Step 2: After the autoencoder is trained, only the encoder part is retained to extract features, and the UMDose is optimized by calculating the DVH depth feature loss between the predicted dose distribution and the true dose distribution; See Figure 3Content of stacked denoising autoencoder feature extraction.
[0016] The dose volume loss function is as follows: The loss function of this project includes two parts:
[0017]
[0018]
[0019] Where represents the true dose distribution, represents the predicted dose distribution, represents the autoencoder, is a hyperparameter used to adjust the weight between the two loss functions.
[0020] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed claims.
[0021] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for automatically predicting the radiotherapy dose of nasopharyngeal carcinoma based on deep learning, characterized in that, The steps are as follows: S1. Analyze the acquired original data; S2. Write a program to batch preprocess all the original data; S3. Input the preprocessed CT images into a variant network designed based on UNet and VMamba. The variant network takes the UNet encoder-decoder as the main architecture, and embeds a VMamba module between the 3rd - 5th layers of the encoder. This module includes a position encoding layer, three parallel dilated convolution branches with dilation rates of 1 / 3 / 5 respectively, and a gating mechanism layer. The variant network inputs CT images with a size of 256×256×64 pixels and outputs a three-dimensional dose distribution matrix; S4. Under the same CT dataset of 200 nasopharyngeal carcinoma patients and NVIDIA A100 GPU environment, compare the prediction results of the variant network with the existing dose prediction method based on 3D-CNN. The comparison metrics include dose prediction MAE, RMSE, and clinical compliance rate. When MAE ≤ 3 Gy and the clinical compliance rate ≥ 90%, verify the feasibility of this solution.
2. The automatic prediction method for the radiotherapy dose of nasopharyngeal carcinoma based on deep learning according to claim 1, wherein In step S1, the data includes data parameters of the brainstem, left and right eyeballs, left and right lenses, larynx, mandible, optic chiasm, left and right optic nerves, left and right parotid glands, oral cavity, TMJ, submandibular gland, left and right middle ears, left and right inner ears, pituitary gland, spine, thyroid gland, and temporal lobe.
3. The automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning according to claim 1, characterized in that, The steps of batch preprocessing in step S2 are as follows: S021. Format conversion: Use a 3D network structure to directly predict the three-dimensional dose distribution of the patient, and convert the patient's DICOM image sequence into a three-dimensional NIFTI format for subsequent processing; S022. Extract anatomical contours: The information of organs at risk and target area delineation is stored in the RTSTRUCT file. Read each delineated contour as a separate binary image. Finally, integrate all the delineated contours into one file, and assign file names of 1, 2, 3, …, n to each organ at risk. The PTV contour is filled with its corresponding prescribed dose; S023. Extract dose distribution: The radiotherapy dose distribution of the patient is stored in the RTDOSE file. First, convert the RD file into the NIFTI format, where the dose value is the product of the pixel value and the Dose Grid Scaling. Finally, resample the RTDOSE file to the same pixel spacing and shape as the CT image to ensure the alignment of the CT image and the dose distribution; S024. Resampling: The voxel spacing of different patients is different. For the dose prediction task, it is related to the pixel distance. Resample all the CT images and anatomical contours to 1mm×1mm×3mm. Use B-spline interpolation for the CT images and nearest-neighbor interpolation for the anatomical binary contour images; S025. Image cropping: Crop all the images to 256×256. In the depth direction, crop the slices that do not contain delineation information; S026. Data normalization: Normalize the CT images using the maximum-minimum method, and do not normalize the dose distribution.
4. The automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning according to claim 1, wherein, The content of calculating the corresponding dose volume histogram in step S3 further includes: training an autoencoder network based on the DVH matrix, which reduces the dimension of the DVH matrix and extracts the deep features of the DVH. After the autoencoder network is trained, a loss function is designed based on the extracted deep learning features of the DVH to optimize the model learning.
5. The automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning according to claim 1, characterized in that The deep learning model in step S4 is built based on the PyTorch framework and trained on 2 NVIDIA GeForce 3090 GPUs. The Adam optimizer is used for optimization, the learning rate is 1e-4, the batch size is 1, and the model is trained for 300 epochs to reach convergence.
6. The automatic prediction method for nasopharyngeal carcinoma radiotherapy dose based on deep learning according to claim 1, wherein In step S4, the original feature extraction module in UNet is uniformly replaced with the newly designed Mamba block, while retaining its original U-shaped structure and skip connections, which can take into account both local features and long-distance sequence features to achieve dose prediction.
Citation Information
Patent Citations
Electroencephalogram feature optimization and epileptic seizure detection method based on deep self-coding
CN111387974A
Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model
CN119152205A
Nasopharynx cancer radiotherapy plan dose prediction method based on progressive feature fusion network
CN119379691A
Pancreas CT image segmentation method and device based on mamba
CN119478389A
Cited By
Cascade network nasopharynx cancer radiotherapy dose prediction system and method
CN121812071A
A nasopharyngeal carcinoma radiotherapy dose prediction system and method of a cascade network
CN121812071B