Longitudinal ct image assisted evaluation of response to neoadjuvant chemoradiation treatment for rectal cancer

By automatically segmenting and fusing deep learning features using a longitudinal CT image-assisted evaluation device, a multimodal prediction model is constructed, which solves the problem of accuracy in evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer, and realizes high-precision preoperative prediction of treatment response and individualized treatment plan planning.

CN121661045BActive Publication Date: 2026-04-14XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
Filing Date
2026-02-04
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, the assessment of response to neoadjuvant chemoradiotherapy for rectal cancer mainly relies on postoperative pathological examination, which cannot accurately predict the outcome before surgery. Traditional imaging assessment methods cannot capture dynamic changes in the tumor due to single-phase analysis, and manual segmentation is highly subjective and features are not fully extracted, resulting in low prediction accuracy.

Method used

A longitudinal CT image-assisted assessment device is used, including a lesion region segmentation module, a feature fusion module, and a prediction module. By automatically segmenting CT images, extracting deep learning features and performing weighted fusion, a multimodal prediction model is constructed, and the treatment response prediction results are output.

Benefits of technology

It significantly improved the accuracy of predicting pathological complete remission in rectal cancer patients, achieved high-precision preoperative assessment of treatment response, reduced reliance on postoperative pathological examination, and improved the scientific rigor and timeliness of treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661045B_ABST
    Figure CN121661045B_ABST
Patent Text Reader

Abstract

The application discloses a kind of longitudinal CT image auxiliary evaluation rectal cancer neoadjuvant radiotherapy treatment response device, the device includes focus area segmentation module obtains rectal cancer patient before neoadjuvant radiotherapy and before surgery CT enhanced image, carries out rectal focus area automatic segmentation to target image after pre-processing, obtains standardized three-dimensional tumor area;Feature fusion module extracts the deep learning feature of segmented focus area before neoadjuvant radiotherapy and before surgery from three-dimensional tumor area respectively, two time points deep learning features are weighted and fused by dynamic attention weight layer, obtain longitudinal comprehensive features representing tumor treatment response change;Predictive module fuses longitudinal comprehensive features and clinical baseline data to construct multimodal prediction model, outputs the treatment response prediction result of rectal cancer patient to neoadjuvant radiotherapy, can improve the prediction accuracy of rectal cancer patient pathological complete remission, improve patient prognosis and medical resource utilization efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image artificial intelligence analysis technology, and in particular to a device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images. Background Technology

[0002] Neoadjuvant chemoradiotherapy (NCRT) has been shown to effectively reduce the stage of locally advanced rectal cancer, with approximately 20% of patients achieving pathological complete response (pCR). This allows for a "wait and see" strategy, avoiding major surgical complications, preserving bowel function, and greatly improving quality of life.

[0003] Patient response to NCRT is particularly important for surgical planning and management decisions.

[0004] Currently, the main clinical assessment of rectal cancer patients' response to NCRT is through postoperative histopathological evaluation to determine pCR.

[0005] In addition, doctors can also visually observe the computed tomography (CT) images after NCRT and subjectively assess the patient's potential chemotherapy benefits by combining indicators such as tumor size and density.

[0006] Current research uses traditional machine learning algorithms (such as support vector machines and random forests) combined with manually extracted CT image features to predict pCR, but most of these are single-time-point predictions (using only CT images before NCRT).

[0007] The existing technology has the following problems: the evaluation method that takes postoperative pathological examination as the gold standard has limited accuracy in predicting pCR by preoperative imaging diagnosis (RECIST criteria, etc.). It is affected by factors such as the doctor's subjectivity and experience, and there is a misdiagnosis rate of about 20%-30% in distinguishing between "residual fibrosis" and "small viable tumors", thus resulting in "overtreatment" or "delayed evaluation".

[0008] In addition, traditional machine learning relies on manual extraction of CT features, which is time-consuming and cannot uncover deep features such as tumor heterogeneity and angiogenesis in images. Moreover, it is mostly a single time point prediction, which cannot assess the dynamic changes of tumors induced by NCRT, resulting in low prediction accuracy and difficulty in meeting clinical needs. Summary of the Invention

[0009] The main objective of this invention is to provide a longitudinal CT image-assisted device for assessing the response to neoadjuvant chemoradiotherapy for rectal cancer. This device aims to address the technical problems in the prior art where the assessment of the response to neoadjuvant chemoradiotherapy for rectal cancer mainly relies on postoperative pathological examination, which cannot accurately predict the response before surgery. It also addresses the low prediction accuracy caused by traditional imaging assessment methods, which suffer from the inability to capture dynamic changes in the tumor due to single-phase analysis, the strong subjectivity of manual segmentation, and incomplete feature extraction.

[0010] This invention provides a device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images. The device comprises:

[0011] The system comprises a lesion region segmentation module, a feature fusion module, and a prediction module; among which,

[0012] The lesion region segmentation module is used to acquire CT enhanced images of rectal cancer patients before neoadjuvant chemoradiotherapy and surgery, preprocess the CT enhanced images, and automatically segment the rectal lesion region of the preprocessed target image to obtain a standardized three-dimensional tumor region.

[0013] The feature fusion module is used to extract deep learning features of the segmented lesion regions at two time points before neoadjuvant chemoradiotherapy and before surgery from the three-dimensional tumor region, and to perform weighted fusion of the deep learning features at the two time points through a dynamic attention weight layer to obtain a longitudinal comprehensive feature characterizing the changes in tumor treatment response.

[0014] The prediction module is used to integrate the longitudinal comprehensive features with clinical baseline data to construct a multimodal prediction model, and output the predicted treatment response of the rectal cancer patient to neoadjuvant chemoradiotherapy through the multimodal prediction model.

[0015] Optionally, the lesion area segmentation module is also used to acquire T0 images reflecting the initial state of the tumor before neoadjuvant chemoradiotherapy and T1 images reflecting the changes in the tumor after treatment before surgery in rectal cancer patients.

[0016] The lesion region segmentation module is also used to perform standardized preprocessing on the T0 image and the T1 image respectively to obtain the preprocessed first target image and second target image.

[0017] The lesion region segmentation module is also used to automatically segment the rectal lesion region of the first target image and the second target image to obtain a standardized three-dimensional tumor region.

[0018] Optionally, the lesion region segmentation module is further configured to use an adaptive median filtering algorithm to eliminate electronic noise and artifacts in the T0 and T1 images using the following formula:

[0019]

[0020] in, For coordinate position The pixel value after adaptive median filtering. For median, The original image relative to the center point Offset The pixel value of the location, coordinates Belonging to Centered adaptive window Scope coordinates A centered adaptive change window;

[0021] The lesion region segmentation module is also used to perform image normalization processing on the filtered T0 image and the T1 image, and obtain the normalized pixel value by the following formula;

[0022]

[0023] in, These are the normalized pixel values. These are the original pixel values ​​after cropping. This is the lower limit of the pixel value. This represents the upper limit of pixel values.

[0024] The lesion region segmentation module is also used to generate a first target image corresponding to the T0 image and a second target image corresponding to the T1 image based on the normalized pixel values.

[0025] Optionally, the lesion region segmentation module is further configured to input the first target image and the second target image into a 3D U-Net deep learning segmentation model based on spatial attention mechanism to obtain a pixel-level tumor segmentation mask for the rectal lesion region.

[0026] The lesion region segmentation module is further used to multiply the pixel-level tumor segmentation mask with the first target image and the second target image pixel by pixel to extract the pure tumor volume, and uniformly resample to a preset standardized three-dimensional size to obtain a standardized three-dimensional tumor region.

[0027] Optionally, the lesion region segmentation module is further configured to input the first target image and the second target image into a 3D U-Net deep learning segmentation model based on spatial attention mechanism, and to perform registration, cropping and normalization processing on the first target image and the second target image with the gold standard mask.

[0028] The lesion region segmentation module is also used during the training phase to employ a joint loss function of Dice and cross-entropy using the following formula:

[0029]

[0030]

[0031] in, The value of the Dice loss function. For model prediction mask, For gold standard mask, To predict the sum of all voxel values ​​in the mask, This is the sum of all voxel values ​​in the gold standard mask. As a smoothing factor, For the total loss function, These are the weighting coefficients of the Dice loss. These are the weighting coefficients for the cross-entropy loss. Cross-entropy loss;

[0032] The lesion region segmentation module is also used to calculate the gradient of the parameters of the 3DU-Net deep learning segmentation model through backpropagation using the joint loss function, update the parameters using the Adam optimizer, and implement an early stopping strategy on the validation set to prevent overfitting.

[0033] The lesion region segmentation module is also used to dynamically adjust the weight ratio of Dice loss and cross-entropy loss during training, so that the 3D U-Net deep learning segmentation model can achieve a Dice similarity coefficient that meets the clinical diagnostic requirements on the validation set. The trained 3D U-Net deep learning segmentation model is then applied to the first target image and the second target image to obtain a pixel-level tumor segmentation mask for the rectal lesion region.

[0034] Optionally, the lesion region segmentation module is further configured to perform a pixel multiplication operation between the first pixel-level tumor segmentation mask corresponding to the first target image before the start of neoadjuvant chemoradiotherapy and the image itself, using the following formula:

[0035]

[0036] in, This corresponds to the first tumor region in the first target image. This is the first target image. This is the first pixel-level tumor segmentation mask;

[0037] The lesion region segmentation module is also used to perform a pixel multiplication operation between the corresponding second pixel-level tumor segmentation mask before surgery and the corresponding second target image using the following formula:

[0038]

[0039] in, This corresponds to the second tumor region in the second target image. This is the second target image. This is the second pixel-level tumor segmentation mask;

[0040] The lesion region segmentation module is also used to perform three-dimensional spatial resampling on the first tumor region and the second tumor region respectively, and to uniformly convert the tumor regions obtained by different patients and under different scanning conditions to a preset standardized three-dimensional size to obtain a standardized three-dimensional tumor region.

[0041] Optionally, the feature fusion module is further configured to input the first tumor region corresponding to the first target image and the second tumor region corresponding to the second target image in the three-dimensional tumor region into two parallel 3D ResNet-50 feature extraction branches that integrate channel attention mechanism in the 3rd to 5th convolutional blocks, to obtain feature vectors at two time points before neoadjuvant chemoradiotherapy and before surgery.

[0042] The feature fusion module is also used to adaptively weight and fuse the feature vectors of the two time points through a dynamic attention weight layer to obtain a longitudinal comprehensive feature characterizing the changes in tumor treatment response.

[0043] Optionally, the feature fusion module is further configured to input the first tumor region corresponding to the first target image and the second tumor region corresponding to the second target image in the three-dimensional tumor region into two parallel improved 3D ResNet-50 feature extraction branches. Each branch integrates a channel attention mechanism in the 3rd to 5th convolutional blocks to enhance the feature extraction capability of the 3D ResNet-50 feature extraction branch for tumor heterogeneity regions and suppress the weight of normal tissue features. The feature vectors at the two time points before neoadjuvant chemoradiotherapy and before surgery are obtained by the following formula:

[0044]

[0045] in, The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. The first tumor region, The 2048-dimensional feature vector before surgery This is the second tumor region.

[0046] Optionally, the feature fusion module is further configured to input the first feature vector at the time point before neoadjuvant chemoradiotherapy and the second feature vector at the time point before surgery into the dynamic attention weight layer, perform a linear transformation on the concatenated representation of the first and second feature vectors through a learnable weight matrix, and apply the softmax function to generate normalized weight coefficients using the following formula:

[0047]

[0048] in, Adaptive weights for the feature vectors at time points before neoadjuvant chemoradiotherapy. The adaptive weights are the feature vectors at the preoperative time points. For learnable weight matrix, Representing the eigenvector and splicing, The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. The 2048-dimensional feature vector before surgery;

[0049] The feature fusion module is further configured to perform a weighted summation of the first feature vector and the second feature vector based on the weight coefficients, and generate a vertically integrated feature vector using the following formula:

[0050]

[0051] in, This is a vertically integrated feature vector. Adaptive weights for the feature vectors at time points before neoadjuvant chemoradiotherapy. The adaptive weights are the feature vectors at the preoperative time points. The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. This is the 2048-dimensional feature vector before surgery.

[0052] Optionally, the prediction module is further configured to apply one-hot encoding to categorical clinical variables in the clinical baseline data:

[0053]

[0054] in, For a specific value of a clinical variable for a certain subtype, the position of 1 corresponds to the index position of the category among all possible categories;

[0055] Standardization was applied to the continuous clinical variables in the baseline clinical data:

[0056]

[0057] in, For the standardized variable values, The original variable value, Let be the mean of the variable in the training dataset. The standard deviation of the variable in the training dataset;

[0058] The prediction module is also used to combine the processed categorical clinical variables and continuous clinical variables to form a clinical feature vector;

[0059] The prediction module is further configured to concatenate and fuse the clinical feature vector with the longitudinal comprehensive feature using the following formula to obtain the concatenated and fused multimodal feature vector:

[0060]

[0061] in, For multimodal feature vectors, This is a vertically integrated feature vector. This is a clinical feature vector;

[0062] Based on the multimodal feature vectors, the predicted probability of treatment response is calculated using a single-layer fully connected network according to the following formula:

[0063]

[0064] in, To predict probabilities, It is the sigmoid activation function. For learnable weight matrix, For multimodal feature vectors, For bias terms;

[0065] The prediction module is further configured to generate a pathological complete remission prediction result for the rectal cancer patient under neoadjuvant chemoradiotherapy when the prediction probability is greater than a preset probability threshold, and otherwise generate a non-pathological complete remission prediction result.

[0066] This invention proposes a longitudinal CT image-assisted assessment device for evaluating the response to neoadjuvant chemoradiotherapy in rectal cancer. The device comprises: a lesion region segmentation module, used to acquire enhanced CT images of rectal cancer patients before neoadjuvant chemoradiotherapy and before surgery, preprocess the enhanced CT images, and automatically segment the rectal lesion region from the preprocessed target image to obtain a standardized three-dimensional tumor region; and a feature fusion module, used to extract deep learning features of the segmented lesion region from the three-dimensional tumor region at two time points: before neoadjuvant chemoradiotherapy and before surgery, and to perform weighted fusion of the deep learning features at the two time points through a dynamic attention weight layer to obtain a longitudinal comprehensive feature characterizing changes in tumor treatment response. The system includes a prediction module that integrates the longitudinal comprehensive features with clinical baseline data to construct a multimodal prediction model. This model outputs predictions of the treatment response to neoadjuvant chemoradiotherapy in rectal cancer patients. It effectively captures the dynamic changes in the tumor during neoadjuvant chemoradiotherapy, overcoming the limitations of single-phase image analysis in reflecting the dynamic process of treatment response. This significantly improves the accuracy of predicting pathological complete remission in rectal cancer patients, achieving high-precision preoperative treatment response prediction. This allows clinicians to accurately assess treatment effectiveness before surgery, plan individualized surgical strategies or adjust treatment plans in advance, reduce reliance on postoperative pathological examinations, improve the scientific rigor and timeliness of treatment decisions, and ultimately improve patient prognosis and the efficiency of medical resource utilization. Attached Figure Description

[0067] Figure 1 This is a functional block diagram of the first embodiment of the longitudinal CT image-assisted assessment device for neoadjuvant chemoradiotherapy response to rectal cancer according to the present invention;

[0068] Figure 2 A schematic diagram of the system architecture of a device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images;

[0069] Figure 3 A schematic diagram of the system workflow for a device that uses longitudinal CT images to assist in evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer.

[0070] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0071] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0072] The solution of this invention mainly includes: the longitudinal CT image-assisted assessment device for neoadjuvant chemoradiotherapy response in rectal cancer comprises: a lesion region segmentation module, used to acquire enhanced CT images of rectal cancer patients before neoadjuvant chemoradiotherapy and before surgery, preprocess the enhanced CT images, and automatically segment the rectal lesion region of the preprocessed target image to obtain a standardized three-dimensional tumor region; a feature fusion module, used to extract deep learning features of the segmented lesion region at two time points before neoadjuvant chemoradiotherapy and before surgery from the three-dimensional tumor region, and to perform weighted fusion of the deep learning features at the two time points through a dynamic attention weight layer to obtain longitudinal comprehensive features characterizing changes in tumor treatment response; and a prediction module, used to fuse the longitudinal comprehensive features with clinical baseline data to construct a multimodal prediction model, and to output the rectal cancer patient's response to neoadjuvant chemoradiotherapy through the multimodal prediction model. The predictive results of radiotherapy and chemotherapy treatment response can effectively capture the dynamic changes of tumors during neoadjuvant radiotherapy and chemotherapy, overcoming the limitations of single-phase image analysis in reflecting the dynamic process of treatment response. It significantly improves the prediction accuracy of pathological complete remission in rectal cancer patients, achieving high-precision preoperative prediction of treatment response. This allows clinicians to accurately assess treatment effects before surgery, plan individualized surgical strategies or adjust treatment plans in advance, reduce reliance on postoperative pathological examinations, improve the scientific nature and timeliness of treatment decisions, and ultimately improve patient prognosis and the efficiency of medical resource utilization. It solves the technical problems of existing technologies, such as the assessment of neoadjuvant radiotherapy and chemotherapy treatment response in rectal cancer mainly relying on postoperative pathological examinations and the inability to accurately predict preoperatively, as well as the low prediction accuracy caused by traditional image assessment methods due to the inability of single-phase analysis to capture dynamic changes of tumors, strong subjectivity of manual segmentation, and incomplete feature extraction.

[0073] Reference Figure 1 , Figure 1 This is a functional block diagram of the first embodiment of the longitudinal CT image-assisted assessment device for neoadjuvant chemoradiotherapy response to rectal cancer according to the present invention.

[0074] In a first embodiment of the longitudinal CT image-assisted assessment device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer of the present invention, the device comprises:

[0075] The module comprises a lesion region segmentation module 10, a feature fusion module 20, and a prediction module 30; among which,

[0076] The lesion region segmentation module 10 is used to acquire CT enhanced images of rectal cancer patients before neoadjuvant chemoradiotherapy and before surgery, preprocess the CT enhanced images, and automatically segment the rectal lesion region of the preprocessed target image to obtain a standardized three-dimensional tumor region.

[0077] The feature fusion module 20 is used to extract deep learning features of the segmented lesion regions at two time points, before neoadjuvant chemoradiotherapy and before surgery, from the three-dimensional tumor region, and to perform weighted fusion of the deep learning features at the two time points through a dynamic attention weight layer to obtain a longitudinal comprehensive feature characterizing the changes in tumor treatment response.

[0078] The prediction module 30 is used to integrate the longitudinal comprehensive features and clinical baseline data to construct a multimodal prediction model, and output the predicted treatment response of the rectal cancer patient to neoadjuvant chemoradiotherapy through the multimodal prediction model.

[0079] It should be noted that contrast-enhanced CT scans of rectal cancer patients were acquired before the start of neoadjuvant chemoradiotherapy (T0 time point) and before surgery (T1 time point). The T0 images reflect the initial state of the tumor, while the T1 images reflect the tumor regression after treatment. Subsequently, the contrast-enhanced CT images at these two time points were subjected to standardized preprocessing. Then, the rectal lesion area was automatically segmented from the preprocessed target images to obtain a standardized three-dimensional tumor region, providing a high-quality input basis for subsequent feature extraction and treatment response assessment.

[0080] Understandably, the system automatically extracts deep learning feature vectors at two time points, T0 and T1, from the standardized three-dimensional tumor region (T0 before neoadjuvant chemoradiotherapy and T1 before surgery). Subsequently, the feature vectors at the two time points are input into a dynamic attention weight layer, and the deep learning features at the two time points are weighted and fused through the dynamic attention weight layer to obtain a longitudinal comprehensive feature that characterizes the changes in tumor treatment response. This can accurately capture the dynamic changes of the tumor during neoadjuvant chemoradiotherapy and provide a high-order feature representation containing time dimension information for subsequent treatment response prediction, significantly improving the accuracy and clinical applicability of predicting pathological complete remission in rectal cancer patients.

[0081] It should be understood that the multimodal prediction model constructed by integrating the longitudinal comprehensive features with clinical baseline data, and the output of the predicted treatment response of rectal cancer patients to neoadjuvant chemoradiotherapy through the multimodal prediction model, significantly improves the accuracy and robustness of the prediction of neoadjuvant chemoradiotherapy response in rectal cancer patients. This enables clinicians to accurately assess the treatment effect before surgery, plan individualized treatment plans in advance, reduce reliance on postoperative pathological examinations, improve the scientific nature and timeliness of treatment decisions, and ultimately improve patient prognosis and the efficiency of medical resource utilization.

[0082] Furthermore, the lesion area segmentation module 10 is also used to acquire T0 images reflecting the initial state of the tumor before neoadjuvant chemoradiotherapy and T1 images reflecting the changes in the tumor after treatment before surgery in rectal cancer patients.

[0083] The lesion region segmentation module 10 is also used to perform standardized preprocessing on the T0 image and the T1 image respectively to obtain the preprocessed first target image and second target image.

[0084] The lesion region segmentation module 10 is also used to automatically segment the rectal lesion region of the first target image and the second target image to obtain a standardized three-dimensional tumor region.

[0085] It should be noted that the lesion region segmentation module first acquires T0 images reflecting the initial state of the tumor before neoadjuvant chemoradiotherapy (T0 time point) and T1 images reflecting the post-treatment changes of the tumor before surgery (T1 time point). These two time points form the basis of longitudinal analysis. Subsequently, the module performs standardized preprocessing on the T0 and T1 images, including using an adaptive median filtering algorithm to eliminate electronic noise and artifacts in the CT images, and normalizing the pixel values ​​of the CT images to the range of [-1000, 400] Hounsfield Units (HU) to highlight soft tissue contrast, thereby obtaining the preprocessed first and second target images, ensuring the comparability of images acquired by different devices and under different conditions. Finally, the module inputs the preprocessed first and second target images into a 3D U-Net deep learning segmentation model based on a spatial attention mechanism. This model is based on classic 3D... U-Net introduces spatial attention units into its encoder-decoder framework, enabling the network to focus more intently on the morphological information of tumors in local space within complex abdominal and pelvic environments. This allows for the automatic identification and precise segmentation of rectal lesion regions, generating pixel-level tumor segmentation masks. The segmentation masks are then multiplied pixel-by-pixel with the corresponding preprocessed images to extract the pure tumor volume. The images are then uniformly resampled to a standardized 3D size, eliminating spatial resolution differences between images from different patients and at different time points. Ultimately, this yields structured and comparable 3D tumor regions, providing high-quality input data for subsequent feature extraction and treatment response assessment.

[0086] Furthermore, the lesion region segmentation module 10 is also used to eliminate electronic noise and artifacts in the T0 image and the T1 image using an adaptive median filtering algorithm according to the following formula:

[0087]

[0088] in, For coordinate position The pixel value after adaptive median filtering. For median, The original image relative to the center point Offset The pixel value of the location, coordinates Belonging to Centered adaptive window Scope coordinates A centered adaptive change window;

[0089] The lesion region segmentation module 10 is also used to perform image normalization processing on the filtered T0 image and the T1 image, and obtain the normalized pixel value by the following formula;

[0090]

[0091] in, These are the normalized pixel values. These are the original pixel values ​​after cropping. This is the lower limit of the pixel value. This represents the upper limit of pixel values.

[0092] The lesion region segmentation module is also used to generate a first target image corresponding to the T0 image and a second target image corresponding to the T1 image based on the normalized pixel values.

[0093] It should be understood that the lesion region segmentation module first uses an adaptive median filtering algorithm to process noise in T0 and T1 images. This eliminates electronic noise and artifacts using the aforementioned formula, and the window size is dynamically adjusted: automatically expanding in noisy areas to effectively suppress noise, and automatically shrinking in areas rich in detail to preserve tumor boundary information. Subsequently, the module performs image normalization on the filtered images, highlighting the contrast between the rectal tumor and surrounding tissues. Finally, the module generates a first target image corresponding to the T0 image and a second target image corresponding to the T1 image based on the normalized pixel values. These preprocessed target images eliminate device differences and noise interference, possessing a uniform HU value range and spatial consistency. This provides high-quality input data for subsequent automatic segmentation of rectal lesion regions, ensuring the accuracy of tumor segmentation and the comparability of images at different time points. This ensures that the segmentation results meet the accuracy standards required for clinical diagnosis, laying a solid data foundation for longitudinal treatment response assessment.

[0094] Furthermore, the lesion region segmentation module 10 is also used to input the first target image and the second target image into a 3D U-Net deep learning segmentation model based on spatial attention mechanism to obtain a pixel-level tumor segmentation mask for the rectal lesion region.

[0095] The lesion region segmentation module 10 is further used to multiply the pixel-level tumor segmentation mask with the first target image and the second target image pixel by pixel to extract the pure tumor volume, and uniformly resample to a preset standardized three-dimensional size to obtain a standardized three-dimensional tumor region.

[0096] Understandably, the lesion region segmentation module inputs the preprocessed first target image (T0 time point before neoadjuvant chemoradiotherapy) and the second target image (T1 time point before surgery) into a 3D U-Net deep learning segmentation model based on spatial attention. This model introduces spatial attention units on the basis of the classic 3D U-Net encoder-decoder framework. By learning the weight distribution in the spatial dimension, the network can more accurately focus on the rectal tumor region in the complex abdominal and pelvic anatomical environment, effectively distinguishing tumor tissue from surrounding normal tissue and intestinal contents. The model first extracts multi-scale features from the input images, and then fuses low-level detail information with high-level semantic information through skip connections. Finally, it outputs a pixel-level tumor segmentation mask with the same spatial dimension as the input images. Each voxel value in the mask represents the probability that the corresponding location belongs to the tumor region. Subsequently, the module performs a pixel-by-pixel multiplication operation between the segmentation mask and the corresponding first and second target images to effectively filter out tumor regions. The module eliminates background tissue and organ interference, retaining only the pure tumor volume. Finally, it resamples the extracted pure tumor volume in three dimensions, uniformly converting tumor regions obtained from different patients and under different scanning conditions to a preset standardized three-dimensional size. This process uses a trilinear interpolation algorithm to ensure the continuity of HU values ​​and the integrity of spatial information, eliminating data inconsistencies caused by differences in the spatial resolution of the original image, scanning parameters, and individual patient differences. Through this series of processing steps, the module obtains a structured, spatially consistent standardized three-dimensional tumor region, accurately characterizing the morphological features of tumors before neoadjuvant chemoradiotherapy and surgery, providing high-quality, comparable input data for subsequent feature extraction and longitudinal treatment response assessment.

[0097] Furthermore, the lesion region segmentation module 10 is also used to input the first target image and the second target image into a 3D U-Net deep learning segmentation model based on spatial attention mechanism, and to perform registration, cropping and normalization processing on the first target image and the second target image with the gold standard mask.

[0098] The lesion region segmentation module 10 is also used during the training phase to employ a joint loss function of Dice and cross-entropy using the following formula:

[0099]

[0100]

[0101] in, The value of the Dice loss function. For model prediction mask, For gold standard mask, To predict the sum of all voxel values ​​in the mask, This is the sum of all voxel values ​​in the gold standard mask. As a smoothing factor, For the total loss function, These are the weighting coefficients of the Dice loss. These are the weighting coefficients for the cross-entropy loss. Cross-entropy loss;

[0102] The lesion region segmentation module 10 is also used to calculate the gradient of the parameters of the 3D U-Net deep learning segmentation model through backpropagation using the joint loss function, update the parameters using the Adam optimizer, and implement an early stopping strategy on the validation set to prevent overfitting.

[0103] The lesion region segmentation module 10 is also used to dynamically adjust the weight ratio of Dice loss and cross-entropy loss during the training process, so that the 3D U-Net deep learning segmentation model can achieve a Dice similarity coefficient that meets the clinical diagnostic requirements on the validation set. The trained 3D U-Net deep learning segmentation model is then applied to the first target image and the second target image to obtain a pixel-level tumor segmentation mask for the rectal lesion region.

[0104] It should be understood that the lesion region segmentation module first precisely registers the first target image (T0 time point before neoadjuvant chemoradiotherapy) and the second target image (T1 time point before surgery) with the gold standard mask annotated by experts, ensuring that the images and annotations are completely aligned in space. Then, it performs cropping to focus on the region of interest containing the rectal region and performs intensity normalization to eliminate HU value differences caused by different scanning conditions. During the training phase, the module optimizes model performance through a joint loss function of Dice and cross-entropy, effectively addressing the problems of small rectal tumor regions and class imbalance. The module uses this joint loss function to calculate the parameter gradient of the 3D U-Net deep learning segmentation model through backpropagation, employs the Adam optimizer for efficient parameter updates, and implements an early stopping strategy on the validation set (stopping training when the validation loss no longer decreases for several consecutive epochs) to prevent overfitting. During training, the module dynamically adjusts the weight ratio according to a predefined strategy. For example, it increases the Dice loss weight in the early stages of training to quickly establish sensitivity to small target regions, and appropriately increases the cross-entropy loss weight in the later stages to optimize boundary segmentation accuracy, ultimately achieving 3D... The U-Net deep learning segmentation model achieved a Dice similarity coefficient that meets clinical diagnostic requirements on the validation set. After training, the optimized model was applied to new first and second target images, automatically outputting a pixel-level tumor segmentation mask for the rectal lesion region. This mask accurately identifies the boundaries and morphology of the tumor in three-dimensional space, providing high-quality lesion region definition for subsequent treatment response assessment.

[0105] Furthermore, the lesion region segmentation module 10 is also used to perform a pixel multiplication operation between the first pixel-level tumor segmentation mask corresponding to the neoadjuvant chemoradiotherapy before the start of the first target image and the first pixel-level tumor segmentation mask corresponding to the start of the neoadjuvant chemoradiotherapy using the following formula:

[0106]

[0107] in, This corresponds to the first tumor region in the first target image. This is the first target image. This is the first pixel-level tumor segmentation mask;

[0108] The lesion region segmentation module 10 is further used to perform a pixel multiplication operation between the corresponding second pixel-level tumor segmentation mask before surgery and the corresponding second target image using the following formula:

[0109]

[0110] in, This corresponds to the second tumor region in the second target image. This is the second target image. This is the second pixel-level tumor segmentation mask;

[0111] The lesion region segmentation module is also used to perform three-dimensional spatial resampling on the first tumor region and the second tumor region respectively, and to uniformly convert the tumor regions obtained by different patients and under different scanning conditions to a preset standardized three-dimensional size to obtain a standardized three-dimensional tumor region.

[0112] It should be noted that the lesion region segmentation module achieves precise extraction and standardization of tumor regions through accurate mathematical operations. Specifically, the module performs a pixel-by-pixel multiplication operation between the first pixel-level tumor segmentation mask corresponding to the neoadjuvant chemoradiotherapy and the corresponding preprocessed first target image. This operation ensures that the tumor region (value close to 1) in the segmentation mask retains its original CT value, while the non-tumor region (value close to 0) is set to zero, thereby accurately extracting the first tumor region containing only tumor tissue and effectively filtering out interference from surrounding normal tissues such as intestinal wall, fat, and muscle. Similarly, the module performs a pixel-by-pixel multiplication operation between the second pixel-level tumor segmentation mask corresponding to the surgery and the corresponding preprocessed second target image to obtain the second tumor region containing only the post-treatment tumor tissue. Subsequently, the module performs three-dimensional spatial resampling processing on the extracted first and second tumor regions respectively, and uses a trilinear interpolation algorithm to uniformly transform tumor regions obtained under different patients, different scanning equipment, and different parameter conditions to a preset standardized three-dimensional size. This process strictly maintains the continuity of HU values ​​and the integrity of spatial information, eliminating the impact of differences in spatial resolution of the original images (such as slice thickness 1mm vs. 1mm). The module eliminates inconsistencies caused by factors such as 3mm, scanning conditions (e.g., different kilovolt values), and patient body size differences. Through this series of operations, the module ultimately obtains a structured, spatially consistent, standardized three-dimensional tumor region, accurately characterizing the morphological features of the tumor before neoadjuvant chemoradiotherapy and surgery, providing high-quality, comparable input data for subsequent feature extraction and longitudinal treatment response assessment.

[0113] Furthermore, the feature fusion module 20 is also used to input the first tumor region corresponding to the first target image and the second tumor region corresponding to the second target image in the three-dimensional tumor region into two parallel 3D ResNet-50 feature extraction branches that integrate channel attention mechanism in the 3rd to 5th convolutional blocks, to obtain feature vectors at two time points before neoadjuvant chemoradiotherapy and before surgery.

[0114] The feature fusion module 20 is also used to adaptively weight and fuse the feature vectors of the two time points through a dynamic attention weight layer to obtain a longitudinal comprehensive feature characterizing the changes in tumor treatment response.

[0115] Understandably, the feature fusion module first inputs the standardized first tumor region (corresponding to the T0 time point before neoadjuvant chemoradiotherapy) and the second tumor region (corresponding to the T1 time point before surgery) into two structurally identical 3D ResNet-50 feature extraction branches that integrate channel attention mechanisms in the 3rd to 5th convolutional blocks, to obtain feature vectors for the two time points before neoadjuvant chemoradiotherapy and before surgery. The feature vectors of the two time points are then adaptively weighted and fused through a dynamic attention weight layer to obtain a longitudinal comprehensive feature that characterizes the changes in tumor treatment response. This feature not only integrates tumor imaging information before and after treatment but also intelligently highlights more discriminative time point features (e.g., automatically assigning higher weights to T1 features when the tumor responds well to treatment), accurately capturing the dynamic changes of the tumor during neoadjuvant chemoradiotherapy.

[0116] Furthermore, the feature fusion module 20 is also used to input the first tumor region corresponding to the first target image and the second tumor region corresponding to the second target image in the three-dimensional tumor region into two parallel improved 3D ResNet-50 feature extraction branches. Each branch integrates a channel attention mechanism in the 3rd to 5th convolutional blocks to enhance the feature extraction capability of the 3D ResNet-50 feature extraction branch for tumor heterogeneous regions and suppress the weight of normal tissue features. The feature vectors at the two time points before neoadjuvant chemoradiotherapy and before surgery are obtained by the following formula:

[0117]

[0118] in, The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. The first tumor region, The 2048-dimensional feature vector before surgery This is the second tumor region.

[0119] It should be understood that the feature fusion module inputs the standardized first tumor region (corresponding to the T0 time point before neoadjuvant chemoradiotherapy) and the second tumor region (corresponding to the T1 time point before surgery) into two structurally identical improved 3DResNet-50 feature extraction branches. These two branches work in parallel without interfering with each other, each processing tumor data at one time point. Each feature extraction branch integrates a channel attention mechanism in the 3rd to 5th convolutional blocks of the standard 3D ResNet-50 network. This mechanism first performs global average pooling on the 3D feature map output by the convolutional block to obtain channel-level statistical information. Then, it generates channel weight vectors through a two-layer fully connected bottleneck structure, and obtains normalized weights through a sigmoid activation function. Finally, it multiplies these weights with the original feature map channel by channel to enhance the features of heterogeneous tumor regions (such as necrotic lesions, enhancement areas, and fibrotic areas) and suppress normal tissue features. This targeted improvement significantly enhances the 3DResNet-50 feature extraction branches. The network's ability to extract features from key regions in rectal cancer CT images allows it to adaptively focus on imaging features sensitive to treatment response, obtaining 2048-dimensional feature vectors before neoadjuvant chemoradiotherapy and before surgery, where represents the first tumor region and represents the second tumor region. These high-dimensional feature vectors automatically encode multi-scale information of the tumor region, including key imaging characteristics such as the distribution of HU values ​​in the tumor tissue, three-dimensional morphological features, edge sharpness, internal texture heterogeneity, and enhancement patterns, providing a rich discriminative feature base for subsequent longitudinal treatment response assessment.

[0120] Furthermore, the feature fusion module 20 is also used to input the first feature vector at the time point before neoadjuvant chemoradiotherapy and the second feature vector at the time point before surgery into the dynamic attention weight layer. A learnable weight matrix is ​​used to linearly transform the concatenated representation of the first and second feature vectors, and the softmax function is applied to generate normalized weight coefficients using the following formula:

[0121]

[0122] in, Adaptive weights for the feature vectors at time points before neoadjuvant chemoradiotherapy. The adaptive weights are the feature vectors at the preoperative time points. For learnable weight matrix, Representing the eigenvector and splicing, The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. The 2048-dimensional feature vector before surgery;

[0123] The feature fusion module is further configured to perform a weighted summation of the first feature vector and the second feature vector based on the weight coefficients, and generate a vertically integrated feature vector using the following formula:

[0124]

[0125] in, This is a vertically integrated feature vector. Adaptive weights for the feature vectors at time points before neoadjuvant chemoradiotherapy. The adaptive weights are the feature vectors at the preoperative time points. The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. This is the 2048-dimensional feature vector before surgery.

[0126] It should be noted that the feature fusion module inputs the first feature vector from the time point before neoadjuvant chemoradiotherapy and the second feature vector from the time point before surgery into the dynamic attention weight layer. This layer first performs a linear transformation on the concatenated representation of the two feature vectors using a learnable weight matrix to calculate the attention score; then, it applies the softmax function to generate normalized weight coefficients. This allows the model to automatically learn the relative importance of the features at the two time points based on the individualized treatment response patterns of different patients, without the need for manually setting fixed weights. For example, when the tumor responds well to treatment, shrinks significantly, or disappears, the model automatically assigns higher weights to the features at time point T1 (because the imaging features at T1 are more important for predicting the disease at this time). (A complete remission of the tumor is more discriminative), while when the tumor responds poorly to treatment, the features of T0 may be given higher weights. Finally, the module performs a weighted summation of the feature vectors at the two time points based on the calculated weight coefficients to generate a longitudinal comprehensive feature vector. This feature not only integrates tumor imaging information before and after neoadjuvant chemoradiotherapy, but also highlights the most discriminative time point features through dynamic weight allocation, effectively capturing the dynamic changes of the tumor during treatment. This adaptive mechanism enables the model to dynamically adjust the feature representation according to the tumor treatment response patterns of different patients, solving the problem that simple averaging or fixed weight fusion in traditional methods cannot adapt to individual differences.

[0127] Furthermore, the prediction module 30 is also used to apply one-hot encoding to the categorical clinical variables in the clinical baseline data:

[0128]

[0129] in, For a specific value of a clinical variable for a certain subtype, the position of 1 corresponds to the index position of the category among all possible categories;

[0130] Standardization was applied to the continuous clinical variables in the baseline clinical data:

[0131]

[0132] in, For the standardized variable values, The original variable value, Let be the mean of the variable in the training dataset. Let be the standard deviation of the variable in the training dataset.

[0133] The prediction module 30 is also used to combine the processed categorical clinical variables and continuous clinical variables to form a clinical feature vector.

[0134] The prediction module 30 is further configured to concatenate and fuse the clinical feature vector and the longitudinal comprehensive feature using the following formula to obtain the concatenated and fused multimodal feature vector:

[0135]

[0136] in, For multimodal feature vectors, This is a vertically integrated feature vector. This is a clinical feature vector;

[0137] Based on the multimodal feature vectors, the predicted probability of treatment response is calculated using a single-layer fully connected network according to the following formula:

[0138]

[0139] in, To predict probabilities, It is the sigmoid activation function. For learnable weight matrix, For multimodal feature vectors, For bias terms;

[0140] The prediction module 30 is further configured to generate a pathological complete remission prediction result for the rectal cancer patient to undergo neoadjuvant chemoradiotherapy when the prediction probability is greater than a preset probability threshold, and otherwise generate a non-pathological complete remission prediction result.

[0141] Understandably, the prediction module first applies one-hot encoding to the categorical clinical variables (such as T stage and N stage) in the clinical baseline data, where... For a specific value of a subcategorical clinical variable (such as T3 or N1), the position of 1 corresponds to the index position of the category among all possible categories (for example, when there are 4 categories for T stage, T3 will be encoded as [0, 0, 1, 0]). This process converts the classification information into a numerical vector that can be processed by the machine learning model; at the same time, for continuous clinical variables in the clinical baseline data (such as carcinoembryonic antigen...), ... The antigen (CEA) level is standardized, eliminating dimensional differences and making variables at different scales comparable. The module then combines the processed categorical and continuous clinical variables to form a unified clinical feature vector, which contains a structured representation of all key prognostic factors. The module sets a preset probability threshold (usually 0.5, but other values ​​are possible; this embodiment does not impose restrictions). When the predicted probability is greater than this threshold, a pathological complete response (pCR) prediction is generated, indicating a good response to neoadjuvant chemoradiotherapy; otherwise, a non-pathological complete response (non-pCR) prediction is generated, indicating a poor treatment response. This multimodal prediction method significantly improves the accuracy and robustness of predicting the response to neoadjuvant chemoradiotherapy in rectal cancer patients by fully exploring the correlation between longitudinal changes in imaging features and clinical baseline data. This allows clinicians to accurately assess treatment effects before surgery, plan individualized surgical strategies or adjust treatment plans in advance, reduce reliance on postoperative pathological examinations, improve the scientific rigor and timeliness of treatment decisions, and ultimately improve patient prognosis and the efficiency of medical resource utilization.

[0142] In the specific implementation, see Figure 2 , Figure 2 A schematic diagram of the system architecture of a device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images, as shown below. Figure 2 As shown, the system includes an image acquisition module, an image preprocessing module, a rectal lesion region segmentation module, a lesion feature extraction module, a multimodal prediction module, and a visualization module, enabling accurate prediction of the response to neoadjuvant chemoradiotherapy for rectal cancer. The specific technical solution is as follows:

[0143] Image acquisition module: It connects to the Picture Archiving and Communication System (PACS) via an interface to acquire enhanced CT images of rectal cancer patients at two time points: the most recent one before NCRT (recorded as T0, i.e., at the treatment baseline) and the most recent one before surgery (recorded as T1). These images are used as raw data for processing by the deep learning model.

[0144] Image preprocessing module: For CT enhanced images at two time points, T0 and T1, a synchronous preprocessing process is performed to ensure data consistency. An adaptive median filtering algorithm is used to remove electronic noise and artifacts in the CT images. Image normalization processing is performed to standardize the pixel values ​​of the CT images to [-1000, 400] HU, providing a standardized data format for subsequent model input.

[0145] The rectal lesion region segmentation module employs a 3DU-Net model based on a spatial attention mechanism to automatically segment 3D tumor regions from enhanced CT images at time points T0 and T1. The model structure introduces spatial attention units into the classic 3DU-Net encoder-decoder framework, enabling the network to focus more intently on the morphological information of the tumor in the local space within the complex abdominal and pelvic environment. To achieve stable segmentation performance, the model is pre-trained using a large amount of expert-annotated enhanced CT data: before training, images are registered, cropped, and normalized with the gold standard mask. During training, a joint loss function of Dice and cross-entropy is used, combined with 3D data enhancements such as random rotation, flipping, and intensity perturbation to improve generalization ability. Simultaneously, the Adam optimizer and early stopping strategy on the validation set control convergence, enabling the model to achieve a Dice similarity coefficient of over 0.85 in the rectal lesion segmentation task. During the inference phase, the model outputs tumor ROI masks for T0 and T1 respectively, and multiplies them pixel by pixel with the corresponding original images to obtain the pure tumor volume. Then, it is uniformly resampled to 64×64×64 pixels to provide standardized three-dimensional input for subsequent analysis modules.

[0146] Feature extraction module: An improved 3DResNet-50 convolutional neural network is used as the core of feature extraction. Channel attention mechanism is added to the 3rd to 5th layers of the convolutional blocks of the two parallel 3DResNet-50 branches to enhance the model's ability to extract features from heterogeneous tumor regions (such as necrotic lesions and enhancement areas) and suppress the weights of normal tissue features. The process of introducing channel attention includes: first, global average pooling is performed on the 3D feature map output by the convolutional block to obtain channel-level statistical information; then, channel weight vectors are generated through a two-layer fully connected bottleneck structure and normalized weights are obtained by sigmoid activation; finally, the weights are multiplied by the original feature map one by one according to the channel to achieve important feature enhancement and redundant feature suppression.

[0147] The preprocessed T0 and T1 tumor 3D ROIs are input into two parallel 3DResNet-50 feature extraction branches to obtain static feature vectors of dimension 2048. These static feature vectors are high-dimensional representations automatically extracted by the network from 3D CT images, including but not limited to density variation features, edge structure features, local texture distribution features, and 3D morphological structure features of the tumor region, reflecting the imaging state of the tumor tissue at the corresponding time point. In the feature fusion stage, a dynamic attention weight layer is set in the model to assign weights to the static features of T0 and T1. These weights are automatically learned by the model during training and reflect the relative contribution of the two time point features to the prediction task. The dynamic attention weight layer generates two weight coefficients based on the input features and normalizes them so that the sum of the weights is 1. Then, the static feature vectors of T0 and T1 are weighted and summed according to the obtained weights to obtain a longitudinal comprehensive feature vector for subsequent analysis. By setting the above dynamic weight allocation mechanism, the model can automatically highlight more discriminative time point features without increasing manual intervention, improving the effectiveness of longitudinal feature representation.

[0148] Multimodal prediction module: Patient's clinical baseline data includes categorical variables such as T-stage and N-stage, as well as continuous indicators such as CEA level. For categorical variables, one-hot encoding is used for vectorization, assigning independent binary bits to each possible category to transform the original classification information into a fixed-dimensional vector recognizable by the model. For continuous variables, standardization by subtracting the mean and dividing the standard deviation maps them to a numerically stable range to avoid dimensional differences affecting model training. After the above processing, the clinical feature vector and the longitudinal comprehensive feature vector are concatenated in dimensional order to form a unified input feature. The concatenated high-dimensional vector is input to a two-layer fully connected network, which generates the predicted probability of "pCR" or "non-pCR" treatment response through nonlinear activation and the final probability output layer. The probability is calculated using a normalization function to ensure that the output meets the 0–1 interval constraint and reflects the relative confidence of the binary classification result. This implementation improves the stability and reliability of prediction through joint modeling of deep features and clinical information.

[0149] Visualization module: Marks the tumor ROI region and displays heat maps of key features of the marked rectal tumor lesion region on T0 and T1 images.

[0150] In the specific implementation, see Figure 3 , Figure 3 A schematic diagram of the system workflow for a device that uses longitudinal CT images to assist in evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer, as shown below. Figure 3As shown, the process begins with data acquisition, collecting contrast-enhanced CT scans of rectal cancer patients before neoadjuvant therapy (T0 images) and after neoadjuvant therapy (T1 images) before surgery. Simultaneously, categorical clinical variables such as gender, carcinoembryonic antigen (CEA) levels, pathological type, and Tumor-Node-Metastasis (TNM) staging, as well as continuous clinical variables such as age, weight, and maximum tumor diameter, are acquired. The images then proceed to a preprocessing module, where denoising is performed to eliminate electronic noise and artifacts. Normalization is applied to map pixel values ​​to the [-1000, 400] HU range, and voxel alignment is performed to ensure spatial consistency between the two time points, ultimately generating high-quality first and second target images. In the lesion region segmentation module, the preprocessed images are input into a 3D image processing module based on spatial attention mechanisms. The U-Net deep learning segmentation model enhances the recognition of rectal tumor regions in complex abdominal and pelvic environments through spatial attention units, automatically outputting pixel-level tumor segmentation masks, and then extracting and standardizing them into structured 3D tumor regions. Next, the feature fusion module inputs the standardized tumor regions into two parallel improved 3D ResNet-50 feature extraction branches (High and Low level feature extraction). Each branch integrates channel attention mechanisms in the 3rd-5th convolutional blocks to enhance feature extraction capabilities for heterogeneous tumor regions, obtaining feature vectors (T0) and (T1) respectively. These features at these two time points are then adaptively weighted and fused through a dynamic attention weight layer to generate a longitudinal comprehensive feature that accurately represents tumor changes. Subsequently, the model enters the multimodal prediction module, which applies one-hot encoding to categorical variables in clinical data and standardizes continuous variables to form structured clinical feature vectors.

[0151] And it is spliced ​​and fused with vertical integrated features. );

[0152] The probability of predicting treatment response is calculated using a single-layer fully connected network. );

[0153] Finally, in the results output stage, the system generates a result as pCR (pathological complete remission) or non-pCR based on the comparison between the predicted probability and the preset threshold, providing clinicians with a preoperative treatment effect assessment and assisting in the formulation of individualized surgical plans or subsequent treatment strategies. The entire system significantly improves the accuracy and clinical applicability of predicting the response to neoadjuvant chemoradiotherapy for rectal cancer by integrating longitudinal imaging information and clinical data containing characteristics of time points before and after treatment.

[0154] Compared with the prior art, the technical advantages of this embodiment are mainly reflected in the following aspects, and it effectively solves the problems of the prior art through a combination of specific technical features:

[0155] Supporting preoperative decision-making: Current postoperative pathological assessments are lagging; this system can predict treatment response before surgery, providing a basis for clinical planning of surgical procedures in advance. For patients predicted to have pCR, the feasibility of sphincter-preserving surgery or delaying surgery can be assessed. For patients predicted not to have pCR, surgery can be arranged as early as possible, reducing treatment delays caused by lagging decision-making.

[0156] Reduced Subjective Interference: This system replaces traditional subjective assessment methods that rely on physician experience with an improved 3D convolutional neural network, reducing the impact of human factors on results and improving the stability of clinical assessments. Simultaneously, dynamic feature heatmaps visually display changes in areas before and after chemotherapy, providing physicians with interpretable decision-making evidence and lowering the barrier to entry for users.

[0157] Improved prediction accuracy: The system extracts tumor features before NCRT and surgery from longitudinal CT images, and constructs a multimodal model by combining clinical baseline data. This model comprehensively reflects the biological changes of tumors before and after chemotherapy, accurately identifies high-risk patients, and solves the problem of low prediction accuracy of existing models.

[0158] Enhanced clinical interpretability: The dynamic feature heatmap output by the system can intuitively mark areas with significant changes in tumor density and texture before and after chemotherapy, helping clinicians understand the basis of the model's decision-making and avoid black-box prediction.

[0159] In the specific implementation, the following example further illustrates this embodiment, taking a patient as an example (male, T3N1 stage, CEA 15ng / mL), the implementation process is as follows:

[0160] Image acquisition: Doctors import the patient's baseline (T0) and most recent (T1) enhanced CT images before NCRT via PACS, and the system automatically transmits them to the computing server.

[0161] Image preprocessing: The server processes the T0 and T1 images pixel by pixel to remove electronic noise and artifacts, and normalizes the pixel values ​​of the T0 and T1 images to [-1000, 400] HU.

[0162] Rectal lesion region segmentation: Preprocessed T0 and T1 images are input into the model, which automatically outputs a tumor ROI mask. By multiplying the mask pixel-by-pixel with the original image, pure tumor regions in the T0 and T1 images are extracted, with a uniform size of 64×64×64 pixels. The volume of the T0 tumor ROI is approximately 32.6 cm³. 3 The T1 tumor ROI volume is approximately 21.8 cm. 3 Preliminary results show that the tumor has shrunk after chemotherapy.

[0163] Lesion feature extraction: The segmented T0 and T1 tumor ROIs are input into two parallel 3DResNet-50 branches (T0 branch and T1 branch), respectively, outputting a 2048-dimensional T0 feature vector. Key features include: tumor morphology features (maximum diameter 3.8cm, corresponding to a "morphology dimension" value of 0.63 in the vector), density features (average normalized CT value 0.31, corresponding to a "density dimension" value of 0.58 in the vector), and texture features (gray-level co-occurrence matrix entropy value 1.8, corresponding to a "texture dimension" value of 0.42 in the vector). A 2048-dimensional T1 feature vector is also output, with key features including: tumor morphology features (maximum diameter 3.1cm, corresponding to a value of 0.51), density features (average normalized CT value 0.23, corresponding to a value of 0.39), and texture features (gray-level co-occurrence matrix entropy value 2.1, corresponding to a value of 0.55). Weights are assigned through a dynamic attention layer (T0 feature weight 0.29, T1 feature weight 0.34), and the weighted sum is used to obtain a 2048-dimensional vertical comprehensive feature vector.

[0164] Multimodal prediction: Manually input T3N1 and CEA 15ng / mL, load the patient's clinical feature vector (T3N1 phase is encoded as [0, 0, 1, 0, 0, 1], CEA is standardized to 0.8), fuse it with the longitudinal comprehensive feature vector of the image, input it into the prediction model, and output "non-pCR, prediction probability 91%".

[0165] Visualization: Generates a visual prediction report for this patient.

[0166] It should be noted that the deep learning-based longitudinal CT image-assisted assessment system for evaluating the treatment response of rectal cancer to neoadjuvant chemoradiotherapy in this embodiment includes an image acquisition module, an image preprocessing module, a rectal lesion region segmentation module, a lesion feature extraction module, a multimodal prediction module, and a visualization module. First, it acquires enhanced CT images and clinical baseline data of the patient before NCRT and before surgery through a standardized process. Image preprocessing is completed after denoising and CT value normalization. Then, an improved 3DU-Net model is used to automatically segment the rectal lesion region. Subsequently, an improved 3D ResNet-50 model is used to extract the longitudinal comprehensive features of the tumor at two time points, and clinical features are fused to construct a multimodal prediction model. Finally, the system outputs the predicted probability of pCR or non-pCR of the tumor lesion and generates visualization reports such as longitudinal ROI comparison maps and dynamic feature heatmaps. This system solves the technical problems of existing technologies, such as the inability to identify patients who will benefit from chemotherapy before surgery, reliance on doctors' subjective evaluation of CT images, lack of dynamic information, and low prediction accuracy. It provides objective and reliable decision support for clinical development of individualized treatment plans.

[0167] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0168] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0169] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images, characterized in that, The longitudinal CT image-assisted assessment device for the response to neoadjuvant chemoradiotherapy in rectal cancer includes: a lesion region segmentation module, a feature fusion module, and a prediction module; wherein... The lesion region segmentation module is used to acquire CT enhanced images of rectal cancer patients before neoadjuvant chemoradiotherapy and surgery, preprocess the CT enhanced images, and automatically segment the rectal lesion region of the preprocessed target image to obtain a standardized three-dimensional tumor region. The feature fusion module is used to extract deep learning features of the segmented lesion regions at two time points before neoadjuvant chemoradiotherapy and before surgery from the three-dimensional tumor region, and to perform weighted fusion of the deep learning features at the two time points through a dynamic attention weight layer to obtain a longitudinal comprehensive feature characterizing the changes in tumor treatment response. The prediction module is used to integrate the longitudinal comprehensive features with clinical baseline data to construct a multimodal prediction model, and output the predicted treatment response of the rectal cancer patient to neoadjuvant chemoradiotherapy through the multimodal prediction model; The lesion area segmentation module is also used to acquire T0 images reflecting the initial state of the tumor before neoadjuvant chemoradiotherapy and T1 images reflecting the changes in the tumor after treatment before surgery in rectal cancer patients. The lesion region segmentation module is also used to perform standardized preprocessing on the T0 image and the T1 image respectively to obtain the preprocessed first target image and second target image. The lesion region segmentation module is also used to automatically segment the rectal lesion region of the first target image and the second target image to obtain a standardized three-dimensional tumor region. The feature fusion module is also used to input the first tumor region corresponding to the first target image and the second tumor region corresponding to the second target image in the three-dimensional tumor region into two parallel 3D ResNet-50 feature extraction branches that integrate channel attention mechanism in the 3rd to 5th convolutional blocks, to obtain feature vectors at two time points before neoadjuvant chemoradiotherapy and before surgery. The feature fusion module is also used to adaptively weight and fuse the feature vectors of the two time points through a dynamic attention weight layer to obtain a longitudinal comprehensive feature characterizing the changes in tumor treatment response. The feature fusion module is further used to input the first feature vector at the time point before neoadjuvant chemoradiotherapy and the second feature vector at the time point before surgery into the dynamic attention weight layer. A learnable weight matrix is ​​used to linearly transform the concatenated representation of the first and second feature vectors, and the softmax function is applied to generate normalized weight coefficients using the following formula: in, Adaptive weights for the feature vectors at time points before neoadjuvant chemoradiotherapy. The adaptive weights are the feature vectors at the preoperative time points. For learnable weight matrix, Representing the eigenvector and splicing, The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. The 2048-dimensional feature vector before surgery; The feature fusion module is further configured to perform a weighted summation of the first feature vector and the second feature vector based on the weight coefficients, and generate a vertically integrated feature vector using the following formula: in, This is a vertically integrated feature vector. Adaptive weights for the feature vectors at time points before neoadjuvant chemoradiotherapy. The adaptive weights are the feature vectors at the preoperative time points. The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. This is the 2048-dimensional feature vector before surgery.

2. The device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images as described in claim 1, characterized in that, The lesion region segmentation module is also used to eliminate electronic noise and artifacts in the T0 and T1 images using an adaptive median filtering algorithm according to the following formula: in, For coordinate position The pixel value after adaptive median filtering. For median, The original image relative to the center point Offset The pixel value of the location, coordinates Belonging to Centered adaptive window Scope coordinates A centered adaptive change window; The lesion region segmentation module is also used to perform image normalization processing on the filtered T0 image and the T1 image, and obtain the normalized pixel value by the following formula; in, These are the normalized pixel values. These are the original pixel values ​​after cropping. This is the lower limit of the pixel value. This represents the upper limit of pixel values. The lesion region segmentation module is also used to generate a first target image corresponding to the T0 image and a second target image corresponding to the T1 image based on the normalized pixel values.

3. The device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images as described in claim 1, characterized in that, The lesion region segmentation module is also used to input the first target image and the second target image into a 3D U-Net deep learning segmentation model based on spatial attention mechanism to obtain a pixel-level tumor segmentation mask for the rectal lesion region. The lesion region segmentation module is further used to multiply the pixel-level tumor segmentation mask with the first target image and the second target image pixel by pixel to extract the pure tumor volume, and uniformly resample to a preset standardized three-dimensional size to obtain a standardized three-dimensional tumor region.

4. The device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images as described in claim 3, characterized in that, The lesion region segmentation module is also used to input the first target image and the second target image into a 3D U-Net deep learning segmentation model based on spatial attention mechanism, and to perform registration, cropping and normalization processing on the first target image and the second target image with the gold standard mask. The lesion region segmentation module is also used during the training phase to employ a joint loss function of Dice and cross-entropy using the following formula: in, The value of the Dice loss function. For model prediction mask, For gold standard mask, To predict the sum of all voxel values ​​in the mask, This is the sum of all voxel values ​​in the gold standard mask. As a smoothing factor, For the total loss function, These are the weighting coefficients of the Dice loss. These are the weighting coefficients for the cross-entropy loss. Cross-entropy loss; The lesion region segmentation module is also used to calculate the gradient of the parameters of the 3D U-Net deep learning segmentation model through backpropagation using the joint loss function, update the parameters using the Adam optimizer, and implement an early stopping strategy on the validation set to prevent overfitting. The lesion region segmentation module is also used to dynamically adjust the weight ratio of Dice loss and cross-entropy loss during training, so that the 3D U-Net deep learning segmentation model can achieve a Dice similarity coefficient that meets the clinical diagnostic requirements on the validation set. The trained 3D U-Net deep learning segmentation model is then applied to the first target image and the second target image to obtain a pixel-level tumor segmentation mask for the rectal lesion region.

5. The device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images as described in claim 3, characterized in that, The lesion region segmentation module is also used to perform a pixel multiplication operation between the first pixel-level tumor segmentation mask corresponding to the first target image before the start of neoadjuvant chemoradiotherapy and the image itself, using the following formula: in, This corresponds to the first tumor region in the first target image. This is the first target image. This is the first pixel-level tumor segmentation mask; The lesion region segmentation module is also used to perform a pixel multiplication operation between the corresponding second pixel-level tumor segmentation mask before surgery and the corresponding second target image using the following formula: in, This corresponds to the second tumor region in the second target image. This is the second target image. This is the second pixel-level tumor segmentation mask; The lesion region segmentation module is also used to perform three-dimensional spatial resampling on the first tumor region and the second tumor region respectively, and to uniformly convert the tumor regions obtained by different patients and under different scanning conditions to a preset standardized three-dimensional size to obtain a standardized three-dimensional tumor region.

6. The device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images as described in claim 1, characterized in that, The feature fusion module is further configured to input the first tumor region corresponding to the first target image and the second tumor region corresponding to the second target image in the three-dimensional tumor region into two parallel improved 3DResNet-50 feature extraction branches. Each branch integrates a channel attention mechanism in the 3rd to 5th convolutional blocks to enhance the feature extraction capability of the 3DResNet-50 feature extraction branch for tumor heterogeneous regions and suppress the weight of normal tissue features. The feature vectors at the two time points before neoadjuvant chemoradiotherapy and before surgery are obtained by the following formula: in, The 2048-dimensional feature vector before neoadjuvant chemoradiotherapy. The first tumor region, The 2048-dimensional feature vector before surgery This is the second tumor region.

7. The device for evaluating the response to neoadjuvant chemoradiotherapy for rectal cancer using longitudinal CT images as described in claim 1, characterized in that, The prediction module is also used to apply one-hot encoding to categorical clinical variables in the clinical baseline data: in, For a specific value of a clinical variable for a certain subtype, the position of 1 corresponds to the index position of the category among all possible categories; Standardization was applied to the continuous clinical variables in the baseline clinical data: in, For the standardized variable values, The original variable value, Let be the mean of the variable in the training dataset. The standard deviation of the variable in the training dataset; The prediction module is also used to combine the processed categorical clinical variables and continuous clinical variables to form a clinical feature vector; The prediction module is further configured to concatenate and fuse the clinical feature vector with the longitudinal comprehensive feature using the following formula to obtain the concatenated and fused multimodal feature vector: in, For multimodal feature vectors, This is a vertically integrated feature vector. This is a clinical feature vector; Based on the multimodal feature vectors, the predicted probability of treatment response is calculated using a single-layer fully connected network according to the following formula: in, To predict probabilities, It is the sigmoid activation function. For learnable weight matrix, For multimodal feature vectors, For bias terms; The prediction module is further configured to generate a pathological complete remission prediction result for the rectal cancer patient under neoadjuvant chemoradiotherapy when the prediction probability is greater than a preset probability threshold, and otherwise generate a non-pathological complete remission prediction result.

Citation Information

Patent Citations

  • Rectum CT tumor detection method based on multi-layer residual U-net

    CN113576505A

  • Rectum cancer neoadjuvant chemoradiotherapy curative effect prediction method and system based on magnetic resonance image analysis

    CN119153114A