Tumor radiotherapy reaction prediction method, system and program product based on three-dimensional image and residual network
By adopting a three-dimensional image and residual network method in tumor radiotherapy response prediction, combined with a hard parameter sharing 3D CNN network and a multi-gated hybrid expert model, the coordinated prediction of the SUV average and its rate of change is achieved, solving the problem of insufficient prediction accuracy and learning efficiency in the prior art.
Patent Information
- Application Number
- CN202411950346.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-27
AI Technical Summary
The prior art is difficult to achieve dual-task joint prediction of the average SUV and its rate of change in tumor radiotherapy response prediction, resulting in insufficient learning efficiency and prediction accuracy.
Using a method based on three-dimensional image and residual network, three-dimensional features are extracted through a hard parameter shared 3D CNN network, and multi-task learning is performed in combination with a multi-gated hybrid expert model (MMoE) to achieve collaborative prediction of the SUV average value and its rate of change.
It improves the model's data utilization efficiency and prediction accuracy, enhances the model's generalization ability on different tasks, and can more effectively use limited data for prediction.
Smart Images

Figure CN120032857A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross-technical field of medical image processing and artificial intelligence, and more specifically, to a method, system and program product for predicting tumor radiotherapy response based on three-dimensional images and residual networks. Background Art
[0002] PET / CT is a fusion imaging technology that combines the functional information provided by positron emission tomography (PET) and the anatomical information provided by computed tomography (CT) to more accurately locate, evaluate and monitor various diseases, especially malignant tumors. With the development of precision radiotherapy for tumors, PET / CT is increasingly used in radiotherapy planning. Among them, the prediction of tumor radiotherapy response based on PET / CT is of great significance in cancer treatment. Accurate prediction of the tumor's response to radiotherapy can not only help doctors formulate the most optimized treatment plan, but also contribute to the development of personalized medicine. The tumor standard uptake value (SUV) is an important indicator for quantitatively evaluating tumor metabolic activity. By measuring the SUV value of the tumor area, the area of high metabolic activity can be effectively identified, which is directly meaningful for determining the aggressiveness of the tumor and the treatment response. Therefore, by analyzing the dynamic changes of the SUV value, the treatment effect of the patient can be accurately predicted.
[0003] With the continuous development of deep learning technology, its application in the field of medical image analysis has gradually expanded from early classification tasks such as disease identification and diagnosis to more complex regression tasks, including evaluating treatment effects and quantifying changes in lesion areas. In the field of tumor radiotherapy, accurately predicting the patient's response to treatment is the key to achieving personalized treatment plans.
[0004] In current research, when dealing with small sample medical data, researchers often use strategies including generative adversarial networks to synthesize data samples and transfer learning techniques to utilize knowledge learned from other large sample tasks. Although these methods have improved the effect of small sample data processing to a certain extent, they still have the problem of insufficient generalization ability for specific tasks. Due to the complexity of radiotherapy response, the prediction model of a single task is often difficult to accurately predict treatment effect indicators of different dimensions. For this special case, multi-task learning (MTL) provides a solution, which improves the learning efficiency and prediction accuracy of each task by learning multiple related tasks simultaneously in the same model. In the prediction of tumor radiotherapy response, multi-task learning can handle multiple related prediction tasks at the same time. This method can not only improve the efficiency of the model's use of data, but also help to better understand the intrinsic relationship between different radiotherapy responses when the sample size is limited. However, in the current research on the prediction of tumor radiotherapy response, multi-task learning methods are rarely used to deal with related regression problems, and there is no prediction model based on multi-task learning to simultaneously predict the average SUV and its rate of change.
[0005] The basic assumption behind MTL is that different tasks can share useful information that may not be obtained if learned in separate tasks. The sharing mechanism in MTL can generally be divided into two categories: hard parameter sharing and soft parameter sharing. In hard parameter sharing, multiple tasks share the weights of some neural network layers, such as sharing the convolutional layers of the first few layers, while the subsequent specific layers are designed independently for each task. In soft parameter sharing, each task has an independent model, but the similarity of model parameters is ensured by introducing additional regularization terms. In the soft parameter sharing framework, the Multi-gate Mixture-of-Experts (MMoE) model provides an advanced framework to optimize the performance of multi-task learning. The core design of the MMoE model includes multiple expert networks and multiple gated units, each of which is responsible for capturing a unique set of features from the input data, while the gated units dynamically adjust the contribution of each expert to each task. In addition, each task processes the fused information through an independent regression tower structure to produce the final output, ensuring the specialized processing of tasks and the independence of outputs.
[0006] Therefore, developing a tumor radiotherapy response prediction method based on three-dimensional images and residual networks to achieve dual-task joint prediction of SUV mean value and its rate of change is a technical problem that needs to be solved urgently. Summary of the invention
[0007] Due to the problems existing in the prior art, the present invention proposes a method, system and program product for predicting tumor radiotherapy response based on three-dimensional imaging and residual network. The method first fuses the two-dimensional PETpre image and the two-dimensional Dose image before radiotherapy, and then uses a 3D CNN network with hard parameter sharing to extract features from the fused image. Secondly, the extracted features are input into a multi-gated hybrid expert model, and the SUV average value and its rate of change are simultaneously predicted through a multi-task learning framework to solve the problems that the existing model cannot achieve dual-task joint prediction, and the learning efficiency and prediction accuracy are insufficient.
[0008] To achieve the above objectives, in a first aspect, the present invention provides a method for predicting tumor radiotherapy response based on three-dimensional images and residual networks, characterized in that it comprises the following steps:
[0009] Step S1, image acquisition and slicing: obtaining a three-dimensional PET pre image and a three-dimensional Dose image of the patient before radiotherapy, and performing continuous slicing processing on the three-dimensional PET pre image and the three-dimensional Dose image based on the Z axis; standardizing and normalizing the slice size to obtain a two-dimensional PET pre image and a two-dimensional Dose image;
[0010] Step S2, region of interest processing and image fusion: dividing the two-dimensional PETpre image and the two-dimensional Dose image into regions of interest and retaining the tumor region data, and setting the pixel values of the non-tumor volume area to zero through masking operation; performing image fusion on the PETpre image and the Dose image after the region of interest processing, and obtaining the fused image data set as the entire tumor input data set;
[0011] Step S3, three-dimensional feature extraction: extracting features from the fused image using a 3D CNN network with hard parameter sharing;
[0012] Step S4, construct a multi-gated hybrid expert model (3D ResMMoE model): ResNet is used as the expert network, and N parallel expert networks are constructed, each of which independently learns different feature representations of the input data; an adaptive gating mechanism is designed, and the weight coefficients of each expert network are calculated through the softmax function to achieve dynamic fusion of the outputs of the expert networks; an independent regression tower structure is constructed for each prediction task, the fused feature information is processed, and the prediction output of the corresponding task is generated;
[0013] Step S5, model training: training the multi-gated hybrid expert model in a gradient descent-based manner, using mean square error as a loss function to measure the difference between the predicted value and the true value; obtaining and saving the optimal model parameters through iterative optimization until the performance no longer improves or reaches a predetermined number of iterations;
[0014] Step S6, predicting the mean value and change rate of the tumor standard uptake value in the middle stage of radiotherapy: inputting the fused image into the 3D CNN network to extract the three-dimensional spatial features; taking the extracted three-dimensional features as the independent variable X, and the mean value and change rate of the tumor standard uptake value as the dependent variable Y, respectively, and inputting them into the trained multi-gated hybrid expert model to construct a dual-task regression problem; the last layer in the regression tower structure is set as a fully connected layer with an output dimension of 1, and using two independent regression tower structures, the mean value and change rate of the tumor standard uptake value can be predicted respectively;
[0015] Step S7, model performance evaluation: the leave-one-out cross-validation strategy is used to evaluate the performance of the multi-gated hybrid expert model.
[0016] The prediction method combines the 3D CNN network with hard parameter sharing to extract three-dimensional features, which not only improves the utilization efficiency of parameters, but also enhances the generalization ability of the model on different tasks; at the same time, combined with the multi-gated hybrid expert model in the soft parameter sharing framework, the multi-task learning paradigm is applied to the prediction of small sample data, which can not only realize the collaborative prediction of multiple regression prediction values and improve the computational efficiency, but also utilize the correlation of multiple tasks and the characteristics of shared information to improve the prediction performance and generalization ability of the model.
[0017] Furthermore, in step S1, the size of each slice is adjusted to the same size (H, W) as the largest tumor slice in the data set; in the normalization process, the x and y axis value ranges of each slice are adjusted to between 0 and 1, and the operation formula is: ,in, and Represent the minimum and maximum values in each slice, respectively.
[0018] Furthermore, in step S2, the image fusion adopts a multi-channel fusion method.
[0019] Furthermore, in step S3, the structure of the 3D CNN network is as follows: the convolution kernels of the first, second and third convolution layers all use three-dimensional convolution kernels of size 3*3*3, the padding parameter is set to 1, and the convolution step size is set to 2; each convolution layer is followed by a ReLU activation function, and the calculation formula of the ReLU activation function is: ; The pooling layers after each convolutional layer adopt the maximum pooling operation, the pooling kernel size is 2*2*2, and the step size is 2; a Dropout layer is set after the third pooling layer, and the drop rate is set to 0.3; an adaptive maximum pooling layer is set after the Dropout layer to adjust the feature map to a fixed size (1*10*10); finally, the three-dimensional feature map is converted into a one-dimensional feature vector with a feature dimension of 6400 through a flattening operation as the input of the next stage.
[0020] Furthermore, the multi-gated hybrid expert model in step S4 includes two parallel expert networks, each of which is composed of two residual blocks connected in series and a fully connected output layer; wherein the two residual modules have the same structure, and the specific structure of each residual block is: first, the input features are mapped to a 32-dimensional hidden space through the first fully connected layer, a nonlinear transformation is introduced through the ReLU activation function, and a Dropout layer with a dropout rate of 0.4 is used for regularization; then the features are mapped back to the original input dimension through the second fully connected layer; finally, the original input and the transformed features are added to realize the residual connection, and the output of the residual block is obtained through the ReLU activation function; finally, the expert network maps the features processed by the residual block to the target output dimension through the fully connected layer to complete the final transformation of the features.
[0021] Furthermore, the regression tower in step S4 is composed of two fully connected layers; the first fully connected layer maps the input features to a 10-dimensional latent space, introduces nonlinear transformation through the ReLU activation function, and uses a Dropout layer with a dropout rate of 0.4 for regularization; the second fully connected layer maps the features to a 1-dimensional output; each regression task is configured with an independent tower structure, which is used to predict the average value of tumor standard uptake value and its change rate.
[0022] Furthermore, in step S5, the loss function calculation formula is: ,in, It is a sample The true value of The model is a sample The predicted value of is the total number of samples.
[0023] Furthermore, in step S7, a leave-one-out cross-validation strategy is used to evaluate the model: each time, one patient is selected from the data set as a test set, and the remaining patients are used as training sets, and the cycle is iterated until all patients are evaluated once as test sets; in each round of training, in order to accurately select the best performing model, two patients are specially selected as validation sets, one of which responds to radiotherapy (high rate of change of SUV value) and the other does not respond significantly to radiotherapy (low rate of change of SUV value); the evaluation indicator of the model performance is the root mean square error (RMSE); by calculating the two-level evaluation indicators of the overall average RMSE and the RMSE of a single patient, the overall prediction performance and individual prediction accuracy of the model are comprehensively evaluated.
[0024] In a second aspect, the present invention provides a tumor radiotherapy response prediction system based on three-dimensional images and residual networks, comprising:
[0025] The image preprocessing module is used to perform continuous slicing processing based on the Z axis on the collected three-dimensional PETpre image and three-dimensional Dose image, and standardize and normalize the slice size to obtain a two-dimensional PETpre image and a two-dimensional Dose image; then divide the two-dimensional PETpre image and the two-dimensional Dose image into regions of interest and retain tumor region data; perform image fusion on the PETpre image and Dose image after the processing of the region of interest, and obtain a fused image data set as the entire tumor input data set;
[0026] The feature extraction module uses a 3D CNN network with hard parameter sharing to extract features from the fused image dataset;
[0027] The model building module uses ResNet as the expert network and builds N parallel expert networks. Each expert network independently learns different feature representations of the input data. An adaptive gating mechanism is designed to calculate the weight coefficients of each expert network through the softmax function to achieve dynamic fusion of the expert network outputs. An independent regression tower structure is built for each prediction task to process the fused feature information and generate the prediction output of the corresponding task.
[0028] A model training module, used to train the multi-gated hybrid expert model, obtain and save optimal model parameters;
[0029] A prediction module, including a trained multi-gated hybrid expert model and a 3D CNN network for feature extraction, is used to input the fused image into the 3D CNN network to extract three-dimensional features, and then input the extracted features into the trained multi-gated hybrid expert model to generate a prediction result;
[0030] The model evaluation module is used to evaluate the optimal model saved by the model training module and calculate its performance indicators on the test set.
[0031] In a final aspect, the present invention provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute the above-mentioned method for predicting tumor radiotherapy response based on three-dimensional images and residual networks.
[0032] Compared with the prior art, the present invention has the following technical effects:
[0033] (1) This paper uses a 3D CNN network with hard parameter sharing to extract three-dimensional features in images. Compared with traditional methods, it can more comprehensively extract and utilize spatial information and deep features in image data, thereby providing richer and more accurate feature representation. At the same time, hard parameter sharing not only improves the efficiency of parameter utilization, but also enhances the generalization ability of the model on different tasks;
[0034] (2) The present invention applies the multi-task learning paradigm to the prediction of small sample data, which not only realizes the collaborative prediction of multiple regression prediction values and improves the computational efficiency, but also utilizes the correlation and shared information characteristics of multiple tasks to improve the prediction performance and generalization ability of the model.
[0035] (3) The present invention constructs a tumor radiotherapy response prediction method and system based on three-dimensional imaging and residual network. By combining ResNet and 3D CNN technology, a multi-task learning strategy is used to achieve collaborative prediction of the SUV average value and its rate of change, which can more effectively utilize limited data and improve the prediction accuracy and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention and its features and advantages will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following accompanying drawings.
[0037] Figure 1 A flowchart of the steps of a method for predicting tumor radiotherapy response in one embodiment of the present invention;
[0038] Figure 2 It is a schematic diagram of the structure of the prediction model in the method for predicting tumor radiotherapy response in one embodiment of the present invention;
[0039] Figure 3 is a flow chart of tumor image data preprocessing in one embodiment of the present invention;
[0040] Figure 4 This is a schematic diagram of the structure of a tumor radiotherapy response prediction system in one embodiment of the present invention;
[0041] Figure 5 It is a line graph of the RMSE value of the average change rate of SUV of each model in one embodiment of the present invention;
[0042] Figure 6 It is a line graph of the average RMSE value of SUV of each model in one embodiment of the present invention. DETAILED DESCRIPTION
[0043] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0044] In the following detailed description, many specific details are described to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the well-known algorithms and models are not shown in detail to avoid blurring the main purpose of the present invention; and the multi-channel fusion method, 3D CNN network, ResNet network, gradient descent and other technologies involved in the following effect embodiments are prior arts that can be retrieved.
[0045] In addition, the order of execution of actions, steps, etc. in the devices and methods shown in the claims, specifications and drawings can be implemented in any order as long as there is no special explicit limitation on the order and the output of the previous processing is not used in the subsequent processing.
[0046] Example
[0047] See also Figure 1 and Figure 2 This embodiment provides a method for predicting tumor radiotherapy response based on three-dimensional images and residual networks, comprising the following steps:
[0048] Step S1, image acquisition and slicing: see Figure 3 , obtain the patient's three-dimensional PET pre image and three-dimensional Dose image before radiotherapy, perform continuous slicing based on the Z axis on the three-dimensional PET pre image and the three-dimensional Dose image; standardize and normalize the slice size to obtain a two-dimensional PET pre image and a two-dimensional Dose image. As an example, the size of each slice is uniformly 80 pixels * 80 pixels.
[0049] Normalization refers to normalizing the x and y axis values of each slice. This normalization process is achieved by adjusting the value range to between 0 and 1. The operation formula is: ,in, and Represent the minimum and maximum values in each slice, respectively.
[0050] Step S2: Region of interest processing and image fusion: see Figure 3 , divide the two-dimensional PETpre image and the two-dimensional Dose image into regions of interest and retain the tumor region data, set the pixel values of the non-tumor volume area to zero through mask operation; fuse the PETpre image and Dose image after the processing of the region of interest to obtain the fused image data set as the entire tumor input data set; wherein, the image fusion preferably adopts a multi-channel fusion method. As an example, see Figure 3 , the PET image before radiotherapy (PETpre) is used as the first channel input, and the dose (Dose) image is used as the second channel input for processing.
[0051] Step S3, three-dimensional feature extraction: extracting features from the fused image using a 3D CNN network with hard parameter sharing.
[0052] As a preferred technical solution, in step S3, the structure of the 3D CNN network is as follows: the convolution kernels of the first, second and third convolution layers all use a three-dimensional convolution kernel of size 3*3*3, the padding parameter is set to 1, and the convolution step size is set to 2; each convolution layer is followed by a ReLU activation function, and the calculation formula of the ReLU activation function is: ; The pooling layers after each convolutional layer adopt the maximum pooling operation, the pooling kernel size is 2*2*2, and the step size is 2; a Dropout layer is set after the third pooling layer, and the dropout rate is set to 0.3; an adaptive maximum pooling layer is set after the Dropout layer to adjust the feature map to a fixed size (1*10*10); finally, the three-dimensional feature map is converted into a one-dimensional feature vector with a feature dimension of 6400 through a flattening operation as the input of the next stage.
[0053] Step S4, constructing a multi-gated hybrid expert model: using the ResNet network as the expert network, constructing N parallel expert networks, each of which independently learns different feature representations of the input data; designing an adaptive gating mechanism, calculating the weight coefficients of each expert network through the softmax function, and realizing dynamic fusion of the outputs of the expert networks; constructing an independent regression tower structure for each prediction task, processing the fused feature information and generating the prediction output of the corresponding task;
[0054] As a preferred technical solution, the multi-gated hybrid expert model in step S4 includes two parallel expert networks, each of which is composed of two residual blocks connected in series and a fully connected output layer; the two residual modules have the same structure, and the specific structure of each residual block is: first, the input features are mapped to a 32-dimensional hidden space through the first fully connected layer, a nonlinear transformation is introduced through the ReLU activation function, and a Dropout layer with a dropout rate of 0.4 is used for regularization; then the features are mapped back to the original input dimension through the second fully connected layer; finally, the original input and the transformed features are added to realize the residual connection, and the output of the residual block is obtained through the ReLU activation function; finally, the expert network maps the features processed by the residual block to the target output dimension through the fully connected layer to complete the final transformation of the features.
[0055] Step S5, model training: train the multi-gated hybrid expert model in a gradient descent-based manner, using mean square error as a loss function to measure the difference between the predicted value and the true value; obtain and save the optimal model parameters through iterative optimization until the performance no longer improves or reaches a predetermined number of iterations. In step S5, the loss function calculation formula is: ,in, It is a sample The true value of The model is a sample The predicted value, is the total number of samples.
[0056] As an example, during the model training phase, the model training parameters were optimized, set to 300 epochs, starting with a learning rate of 0.001, and the Adam optimizer was used to adjust the weights, with the weight decay set to 10-4. To further avoid overfitting and optimize the training process, an early stopping strategy was implemented, with a patience value of 100, i.e., if the performance on the validation set did not improve for 100 consecutive epochs, the training was stopped.
[0057] Step S6, predicting the average value and change rate of the standardized uptake value of the tumor in the middle stage of radiotherapy: Input the fused image into the 3D CNN network to extract three-dimensional spatial features; use the extracted three-dimensional features as the independent variable X, and the average value and change rate of the standardized uptake value SUV of the tumor as the dependent variables Y, respectively, and input them into the trained multi-gated mixture of experts model to construct a dual-task regression problem; the last layer in the regression tower structure is set to a fully connected layer with an output dimension of 1, and two independent regression tower structures can be used to predict the average value and change rate of the standardized uptake value of the tumor, respectively.
[0058] Step S7, predicting model performance evaluation: The leave-one-out cross-validation strategy is used to evaluate the model performance.
[0059] The specific implementation process of the leave-one-out cross-validation strategy is as follows: Each time, select one patient from the dataset as the test set, and the remaining patients as the training set, and iterate until all patients have been evaluated as the test set once; during each round of training, in order to accurately select the best-performing model, two patients are specifically selected as the validation set, one patient with a response to radiotherapy (high change rate of SUV value), and the other patient with no obvious response to radiotherapy (low change rate of SUV value); the evaluation index of the model performance is the root mean square error (Root Mean Square Error, RMSE); by calculating the evaluation indexes at two levels of the overall average RMSE and the RMSE of individual patients, the overall prediction performance and individual prediction accuracy of the model are comprehensively evaluated.
[0060] See Figure 4 , in this embodiment, a tumor radiotherapy response prediction system based on multi-three-dimensional images and a residual network is used to implement the above prediction method. The prediction system includes:
[0061] The image preprocessing module is used to perform continuous slicing processing based on the Z axis on the collected three-dimensional PETpre image and three-dimensional Dose image, and standardize and normalize the slice size to obtain a two-dimensional PETpre image and a two-dimensional Dose image; then divide the two-dimensional PETpre image and the two-dimensional Dose image into regions of interest and retain tumor region data; perform image fusion on the PETpre image and Dose image after the processing of the region of interest, and obtain a fused image data set as the entire tumor input data set;
[0062] The feature extraction module uses a 3D CNN network with hard parameter sharing to extract features from the fused image dataset;
[0063] The model building module uses ResNet as the expert network and builds two parallel expert networks. Each expert network independently learns different feature representations of the input data. An adaptive gating mechanism is designed to calculate the weight coefficients of each expert network through the softmax function to achieve dynamic fusion of the expert network outputs. An independent regression tower structure is built for each prediction task to process the fused feature information and generate the prediction output of the corresponding task.
[0064] A model training module, used to train the multi-gated hybrid expert model, obtain and save optimal model parameters;
[0065] A prediction module, including a trained multi-gated hybrid expert model and a 3D CNN network for feature extraction, is used to input the fused image into the 3D CNN network to extract three-dimensional features, and then input the extracted features into the trained multi-gated hybrid expert model to generate a prediction result;
[0066] The model evaluation module is used to evaluate the optimal model saved by the model training module and calculate its performance indicators on the test set.
[0067] After the prediction system preprocessed the tumor imaging data and radiotherapy dose data through the prediction method, the average SUV value and its change rate were predicted. RMSE was selected as the model evaluation index, and the overall prediction performance and individual prediction accuracy of the model were comprehensively evaluated by calculating the overall average RMSE and the RMSE of a single patient.
[0068] The model performance was evaluated using a leave-one-out cross-validation strategy. The validation idea and process of the leave-one-out cross-validation are as follows: in each round of validation, one patient is selected from the data set as the test set, and the remaining patients are used as the training set, and the cycle is iterated until all patients are evaluated once as the test set; in each round of training, in order to accurately select the best performing model, two patients are specially selected as the validation set, one of whom is a patient who responds to radiotherapy (high rate of change of SUV value) and the other is a patient who has no obvious response to radiotherapy (low rate of change of SUV value). RMSE is used as the model evaluation indicator to comprehensively evaluate the performance of the model from two levels: the overall average RMSE and the RMSE of a single patient, so as to verify the prediction accuracy and effectiveness of the model in practical applications. At the same time, classic deep learning models and intermediate models (including 3D MMoE, 3D CNN, and 3D CNN with hard sharing mechanism) are used for comparison and analysis. The difference between 3D MMoE and the 3D ResMMoE model is that its expert network uses a basic fully connected neural network, which contains two fully connected layers and one activation layer. All control models adopt the same training parameter configuration as the model of the present invention, including hyperparameters such as training rounds (epochs) and learning rate. Figure 5 , Figure 6 Table 1 shows the RMSE values of the 3D ResMMoE model of the present invention and the control model in the prediction of the SUV mean value change rate and the SUV mean value prediction.
[0069] Table 1. Table of change rate of average SUV value and average RMSE value of SUV value for each model
[0070] RMSE 3D ResMMoE 3D MMoE 3D CNN Hard parameter sharing SUV average change rate 0.196 0.223 0.223 0.257 Average SUV 1.599 1.653 1.730 1.675
[0071] As can be seen from Table 1, the 3D ResMMoE model proposed in the present invention is significantly better than the other three models in predicting the average SUV change rate and average value. At the same time, the results of 3D MMoE are also better than 3D CNN, which also reflects the advantages of the MMoE multi-task learning framework.
[0072] The above-mentioned method for predicting tumor radiotherapy response based on three-dimensional images and residual networks can be embodied in the form of a computer program product or a software functional unit. If the above-mentioned method for predicting tumor radiotherapy response based on three-dimensional images and residual networks is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Therefore, the essence of the technical solution or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for enabling an electronic system (which can be a personal computer, a server, or a network system, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0073] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0074] In summary, the present invention provides a method, system and program product for predicting tumor radiotherapy response based on three-dimensional images and residual networks. The prediction method includes the steps of image acquisition and slicing, region of interest processing and image fusion, three-dimensional feature extraction, construction of a multi-gated hybrid expert model, model training, prediction of radiotherapy response regression value, model performance evaluation, etc.: first, the patient's three-dimensional PETpre image and three-dimensional Dose image are preprocessed; secondly, the preprocessed image is input into a 3D CNN network with hard parameter sharing to extract three-dimensional spatial features; finally, the extracted three-dimensional features are used as independent variables X, and the SUV average value and its change rate are used as dependent variables Y, respectively, and input into the trained multi-gated hybrid expert model for prediction. The present invention combines ResNet and 3D CNN technology and uses a multi-task learning strategy to achieve collaborative prediction of the SUV average value and its change rate, which can more effectively utilize limited data and improve the accuracy of prediction and the generalization ability of the model.
[0075] Those skilled in the art should understand that those skilled in the art can implement variations by combining the prior art and the above embodiments, which will not be described in detail here. Such variations do not affect the essential content of the present invention, and will not be described in detail here.
[0076] The above describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above-mentioned specific embodiments, and the systems and structures that are not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, or modify them into equivalent embodiments of equivalent changes, which does not affect the essential content of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention are still within the scope of protection of the technical solutions of the present invention.
Claims
1. A method for predicting tumor radiotherapy response based on three-dimensional images and residual networks, characterized in that: The following steps are involved: Step S1, image acquisition and slicing: obtaining a three-dimensional PET pre image and a three-dimensional Dose image of the patient before radiotherapy, and performing continuous slicing processing on the three-dimensional PET pre image and the three-dimensional Dose image based on the Z axis; standardizing and normalizing the slice size to obtain a two-dimensional PET pre image and a two-dimensional Dose image; Step S2, region of interest processing and image fusion: dividing the two-dimensional PETpre image and the two-dimensional Dose image into regions of interest and retaining the tumor region data, and setting the pixel values of the non-tumor volume area to zero through masking operation; performing image fusion on the PETpre image and the Dose image after the region of interest processing, and obtaining the fused image data set as the entire tumor input data set; Step S3, three-dimensional feature extraction: extracting features from the fused image using a 3D CNN network with hard parameter sharing; Step S4, constructing a multi-gated hybrid expert model: using ResNet as the expert network, constructing N parallel expert networks, each expert network independently learning different feature representations of input data; designing an adaptive gating mechanism, calculating the weight coefficients of each expert network through the softmax function, and realizing dynamic fusion of the expert network outputs; constructing an independent regression tower structure for each prediction task, processing the fused feature information and generating the prediction output of the corresponding task; Step S5, model training: training the multi-gated hybrid expert model in a gradient descent-based manner, using mean square error as a loss function to measure the difference between the predicted value and the true value; obtaining and saving the optimal model parameters through iterative optimization until the performance no longer improves or reaches a predetermined number of iterations; Step S6, predicting the average value and change rate of the tumor standard uptake value in the middle stage of radiotherapy: inputting the fused image into the 3D CNN network to extract the three-dimensional spatial features; The extracted three-dimensional features are used as independent variables X, and the mean value of tumor standard uptake value and its change rate are used as dependent variables Y, respectively, and are passed into the trained multi-gated hybrid expert model to construct a dual-task regression problem; the last layer in the regression tower structure is set as a fully connected layer with an output dimension of 1, and two independent regression tower structures are used to respectively predict the mean value of tumor standard uptake value and its change rate; Step S7, model performance evaluation: the leave-one-out cross-validation strategy is used to evaluate the performance of the multi-gated hybrid expert model.
2. The method for predicting tumor radiotherapy response based on three-dimensional images and residual networks according to claim 1, characterized in that: In step S1, the size of each slice is adjusted to the same size (H, W) as the largest tumor slice in the data set; in the normalization process, the x and y axis value ranges of each slice are adjusted to between 0 and 1, and the operation formula is: ,in, and Represent the minimum and maximum values in each slice, respectively.
3. The method for predicting tumor radiotherapy response based on three-dimensional images and residual networks according to claim 1, characterized in that: In the step S2, the image fusion adopts a multi-channel fusion method.
4. The method for predicting tumor radiotherapy response based on three-dimensional images and residual networks according to claim 1, characterized in that: In step S3, the structure of the 3D CNN network is as follows: the convolution kernels of the first, second and third convolution layers all use a three-dimensional convolution kernel of size 3*3*3, the padding parameter is set to 1, and the convolution step size is set to 2; each convolution layer is followed by a ReLU activation function, and the calculation formula of the ReLU activation function is ; The pooling layers after each convolutional layer adopt the maximum pooling operation, the pooling kernel size is 2*2*2, and the step size is 2; a Dropout layer is set after the third pooling layer, and the dropout rate is set to 0.3; an adaptive maximum pooling layer is set after the Dropout layer to adjust the feature map to a fixed size (1*10*10); finally, the three-dimensional feature map is converted into a one-dimensional feature vector with a feature dimension of 6400 through a flattening operation as the input of the next stage.
5. The method for predicting tumor radiotherapy response based on three-dimensional images and residual networks according to claim 1, characterized in that: The multi-gated hybrid expert model in step S4 includes two parallel expert networks, each of which is composed of two residual blocks connected in series and a fully connected output layer; the two residual modules have the same structure, and the specific structure of each residual block is: first, the input features are mapped to a 32-dimensional hidden space through the first fully connected layer, a nonlinear transformation is introduced through the ReLU activation function, and a Dropout layer with a dropout rate of 0.4 is used for regularization; then the features are mapped back to the original input dimension through the second fully connected layer; finally, the original input and the transformed features are added to realize the residual connection, and the output of the residual block is obtained through the ReLU activation function; finally, the expert network maps the features processed by the residual block to the target output dimension through the fully connected layer to complete the final transformation of the features.
6. The method for predicting tumor radiotherapy response based on three-dimensional images and residual networks according to claim 1, characterized in that: The regression tower in step S4 is composed of two fully connected layers; the first fully connected layer maps the input features to a 10-dimensional latent space, introduces nonlinear transformation through the ReLU activation function, and uses a Dropout layer with a dropout rate of 0.4 for regularization; the second fully connected layer maps the features to a 1-dimensional output; each regression task is configured with an independent tower structure, which is used to predict the average value of tumor standard uptake value and its change rate.
7. The method for predicting tumor radiotherapy response based on three-dimensional images and residual networks according to claim 1, characterized in that: In step S5, the loss function calculation formula is: ,in, It is a sample The true value of The model is a sample The predicted value of is the total number of samples.
8. The method for predicting tumor radiotherapy response based on three-dimensional images and residual networks according to claim 1, characterized in that: In step S7, the model is evaluated by using a leave-one-out cross-validation strategy: one patient is selected from the data set each time as the test set, and the remaining patients are used as the training set, and the cycle is iterated until all patients are evaluated once as the test set; in each round of training, in order to accurately select the model with the best performance, two patients are specially selected as the validation set, one of which responds to radiotherapy (high rate of change of SUV value) and the other does not respond significantly to radiotherapy (low rate of change of SUV value); the evaluation indicator of the model performance is the root mean square error; by calculating the two-level evaluation indicators of the overall average RMSE and the RMSE of a single patient, the overall prediction performance and individual prediction accuracy of the model are comprehensively evaluated.
9. A tumor radiotherapy response prediction system based on three-dimensional images and residual networks, characterized in that: include: The image preprocessing module is used to perform continuous slicing processing based on the Z axis on the collected three-dimensional PETpre image and three-dimensional Dose image, and standardize and normalize the slice size to obtain a two-dimensional PETpre image and a two-dimensional Dose image; then divide the two-dimensional PETpre image and the two-dimensional Dose image into regions of interest and retain tumor region data; perform image fusion on the PETpre image and Dose image after the processing of the region of interest, and obtain a fused image data set as the entire tumor input data set; The feature extraction module uses a 3D CNN network with hard parameter sharing to extract features from the fused image dataset; The model building module uses ResNet as the expert network and builds N parallel expert networks. Each expert network independently learns different feature representations of the input data. An adaptive gating mechanism is designed to calculate the weight coefficients of each expert network through the softmax function to achieve dynamic fusion of the expert network outputs. An independent regression tower structure is built for each prediction task to process the fused feature information and generate the prediction output of the corresponding task. A model training module, used to train the multi-gated hybrid expert model, obtain and save optimal model parameters; A prediction module, including a trained multi-gated hybrid expert model and a 3D CNN network for feature extraction, is used to input the fused image into the 3D CNN network to extract three-dimensional features, and then input the extracted features into the trained multi-gated hybrid expert model to generate a prediction result; The model evaluation module is used to evaluate the optimal model saved by the model training module and calculate its performance indicators on the test set.
10. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the method for predicting tumor radiotherapy response based on three-dimensional images and residual networks as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Abdomen multi-organ registration method based on adaptive multi-gating hybrid expert model
CN116993793A
Cross-anatomical-region organ increment segmentation method based on self-supervision and expert gating
CN117151162A
High-frequency electric appliance identification model construction method and system based on multi-gate control expert network
CN117574238A
Protein secondary structure prediction method based on multi-task deep learning
CN118280432A
Tumor radiotherapy reaction prediction method and system based on variational auto-encoder, and terminal
CN118781454A
Cited By
Three-dimensional medical image anomaly detection method and system based on self-supervised learning
CN120259787A