A method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-14
AI Technical Summary
近年来,放射组学通过高通量提取影像特征构建诊断模型,为无创评估提供了新思路,但其传统方法往往将肿瘤视为单一整体,未能充分解析瘤内异质性对诊断的影响
1、本发明利用DECT尿路造影图像重建40keV、70keV、100keV虚拟单能图像,全程无需侵入性操作,实现了与膀胱镜活检等有创检查相比无创、快速、可重复的术前评估;同时克服了MRI对幽闭恐惧症患者及体内金属植入物患者的应用限制,CT扫描速度快、适用范围广,显著扩大了临床适用人群。
Smart Images

Figure CN122575643A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, specifically to a method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion. Background Technology
[0002] Bladder cancer, a highly prevalent malignant tumor of the urinary system, relies heavily on accurate assessment of muscle-invasive status in its clinical diagnosis and treatment. Significant differences exist in treatment strategies between muscle-invasive and non-muscle-invasive bladder cancer: the former requires radical cystectomy, while the latter typically involves transurethral resection of the bladder tumor combined with irrigation therapy. Currently, cystoscopic biopsy is used clinically for pathological confirmation, but this technique has inherent limitations, including invasiveness and potential misdiagnosis due to sampling bias. Imaging examinations such as CT have the advantages of being non-invasive and repeatable in preoperative assessment, especially dual-energy CT (DECT), which provides quantitative parameters to aid in tumor feature analysis through virtual single-energy image reconstruction and material decomposition techniques. In recent years, radiomics has provided new ideas for non-invasive assessment by extracting imaging features through high-throughput imaging to construct diagnostic models; however, its traditional methods often treat the tumor as a single entity, failing to fully analyze the impact of intratumoral heterogeneity on diagnosis.
[0003] While MRI can assist in assessing muscle layer invasion, its applicability is limited for patients with claustrophobia and those with metallic implants. Radiomics models based on conventional CT suffer from insufficient diagnostic accuracy because they ignore intratumoral functional heterogeneity. Habitat imaging techniques can partially address heterogeneity by dividing tumor subregions, but their application in bladder cancer is still immature, particularly lacking validation through integration with multi-parameter spectral CT images. Furthermore, while model fusion techniques can improve predictive performance, existing research lacks systematic comparisons of different fusion strategies (such as feature-level and decision-level fusion), hindering the optimization of collaborative analysis of multimodal imaging data. These shortcomings collectively limit the accuracy and clinical applicability of preoperative non-invasive assessment of bladder cancer muscle layer invasion. Summary of the Invention
[0004] To address the aforementioned technical challenges, such as the limited applicability of MRI to certain populations, the neglect of tumor heterogeneity by conventional CT radiomics, the immature application of habitat imaging in bladder cancer, and the lack of systematic comparison of multimodal fusion strategies, this invention provides a method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion. This invention primarily utilizes spectral CT habitat imaging technology to divide the tumor into multiple habitat subregions, extracts radiomics features from each subregion, and adaptively fuses multi-level, multi-subregion features using a Transformer fusion model. This achieves the technical effect of non-invasive, accurate, and universally applicable preoperative prediction of bladder cancer muscle invasion.
[0005] The technical means employed in this invention are as follows:
[0006] A method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion includes the following steps: Acquire DECT urography images of the patient and preprocess the DECT urography images into virtual monoenergetic images of high, medium and low energy levels; Delineate the region of interest of the tumor on the virtual monoenergetic image; Local features are calculated in the tumor region of interest, and the tumor region of interest is divided into different sub-regions based on the local features using the K-means clustering algorithm. Patients are randomly divided into training and testing sets, and the optimal features in the sub-regions are extracted. A preoperative bladder cancer prediction model is constructed, comprising a base model and a fusion model. The base model includes a traditional radiomics model based on a multilayer perceptron classifier and a habitat model. Optimal features are used to construct the base model. The traditional radiomics model is based on features in the tumor region of interest, while the habitat model is based on features in sub-regions. The model outputs the diagnostic efficacy of the base model and intermediate layer features. The base model is selected based on its superior diagnostic efficacy. The fusion model includes an early fusion model, an integrated fusion model, a stacked fusion model, and a Transformer fusion model. The fusion model fuses the intermediate layer features to obtain the diagnostic efficacy of the fusion model. The fusion model is selected based on its superior diagnostic efficacy. The preoperative bladder cancer prediction model was trained using the training set. The test set was input into the trained preoperative bladder cancer prediction model to obtain the probability and result of muscle-invasive bladder cancer.
[0007] Further, the step of calculating local features in the tumor region of interest, and dividing the tumor region of interest into different sub-regions based on the local features using a K-means clustering algorithm, includes: The intensity values of all voxels within the tumor region of interest are aggregated into a global feature matrix; In the global feature matrix, a 3×3×3 moving window is used to traverse each voxel, and a local feature vector is extracted for each voxel. The local feature vector includes: entropy, mean absolute deviation, median, information metric relevance, high gray level emphasis, and contrast. Based on the local feature vectors, the global feature matrix is divided into different sub-regions using the K-means clustering algorithm. The K-means clustering algorithm uses the silhouette coefficient, Davies-Bouldin index, and Calinski-Harabasz index to determine the optimal number of clusters, and uses the optimal number of clusters to obtain different sub-regions.
[0008] Further, extracting the optimal features from the sub-region includes: Extract radiomics features from the sub-regions, the radiomics features including first-order features, shape features and texture features; The radiomics features were then subjected to Z-score normalization, t-test, Pearson correlation coefficient screening, and LASSO model regression to select the optimal features.
[0009] Furthermore, the LASSO model selects the optimal regularization parameter through 10-fold cross-validation based on the minimum mean square error criterion, obtains the minimum mean square error based on the optimal regularization parameter, and finds the feature corresponding to the minimum mean square error as the optimal feature.
[0010] Further, the step of delineating the tumor region of interest on the virtual monoenergetic image includes: The tumor region of interest is delineated in a mid-energy virtual monoenergetic image using a segmentation model, and the tumor region of interest is simultaneously delineated in both a high-energy virtual monoenergetic image and a low-energy virtual monoenergetic image.
[0011] Furthermore, the workflow of the Transformer fusion model includes: The intermediate layer features are sequentially fed into the fully connected layer, the multi-head attention layer, and the fusion layer, and the probability and result of muscle-invasive bladder cancer are output.
[0012] Furthermore, the workflow of the preoperative bladder cancer prediction model includes: The traditional radiomics model is based on features in the tumor region of interest, while the habitat model is based on features in the sub-regions. The model outputs the diagnostic efficacy of the basal model and the intermediate layer features. The basal model is selected based on the superior diagnostic efficacy of the traditional radiomics model and the habitat model. The intermediate layer features are input into the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model, respectively, to obtain the diagnostic efficacy of the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model. The fusion model is selected based on the better diagnostic efficacy of the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model.
[0013] Furthermore, the energy levels of the high, medium, and low energy level virtual monoenergetic images are 100keV, 70keV, and 40keV, respectively.
[0014] Furthermore, the preoperative prediction model for bladder cancer was evaluated using AUC, sensitivity, specificity, accuracy, NRI, and IDI.
[0015] Compared with the prior art, the present invention has the following advantages: 1. This invention utilizes DECT urography images to reconstruct 40keV, 70keV, and 100keV virtual monoenergetic images, requiring no invasive procedures throughout the process. This achieves non-invasive, rapid, and repeatable preoperative assessment compared to invasive examinations such as cystoscopy and biopsy. At the same time, it overcomes the limitations of MRI for patients with claustrophobia and those with implanted metal devices. CT scanning is fast and has a wide range of applications, significantly expanding the clinically applicable population.
[0016] 2. This invention automatically segments the tumor region of interest into three habitat subregions with different signal intensity patterns using habitat segmentation technology. Within each subregion, 833 radiomics features (including first-order, shape, and texture features) are extracted. Combined with complementary information from three energy levels of 40keV, 70keV, and 100keV, this invention achieves a more comprehensive capture of the functional heterogeneity and spatial distribution information within the tumor compared to traditional whole-tumor radiomics feature extraction methods. This avoids the loss of key signals due to feature averaging and significantly improves diagnostic reliability and biological interpretability.
[0017] 3. This invention performs habitat segmentation on DECT images at three energy levels (40keV, 70keV, and 100keV). A 3×3×3 moving window is used to extract six local features from each voxel. K-means clustering is then used to automatically divide the tumor into three habitat sub-regions. A Transformer fusion model based on multi-head self-attention adaptively fuses features from multiple energy levels and sub-regions, achieving significantly superior diagnostic performance compared to traditional radiomics models and single DECT models: a training cohort AUC of 0.994, a testing cohort AUC of 0.877, a sensitivity of 0.923, and an accuracy of 0.800. This effectively improves the accuracy of preoperative prediction of bladder cancer muscle invasion.
[0018] Based on the above reasons, this invention can be widely applied in fields such as medical image processing. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the process for a method of preoperative prediction of bladder cancer muscle invasion based on energy spectrum CT habitat imaging and Transformer fusion according to the present invention.
[0021] Figure 2 This is a flowchart illustrating the workflow of a method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to the present invention.
[0022] Figure 3 A relevant map is generated for the habitat of this invention.
[0023] Figure 4 This is a structural diagram of the Transformer fusion model of the present invention.
[0024] Figure 5 Radar charts showing the diagnostic performance of different fusion strategies in this invention.
[0025] Figure 6 This is a performance comparison chart of the models of this invention.
[0026] Figure 7 This is a flowchart illustrating an embodiment of the present invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products or devices.
[0029] like Figure 1 and Figure 2 As shown, this invention provides a method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion, comprising the following steps: S1. Acquire DECT urography images of the patient and preprocess the DECT urography images into virtual monoenergetic images of high, medium and low energy levels.
[0030] DECT urography images of the patient were acquired using a 256-slice CT scanner. The energy levels of the virtual monoenergetic images for high, medium, and low energy levels were 100keV, 70keV, and 40keV, respectively. All images were normalized to voxel size of 1mm×1mm×1mm using ITK-snap software, with a window width of 400HU and a window level of 40HU.
[0031] S2. Delineate the region of interest for the tumor on the virtual monoenergetic image.
[0032] Specifically, all images were uploaded to the OnekeyAI analysis platform, where the 3DUnet automatic segmentation model automatically delineated the tumor region of interest (ROI) on the 70keV images. The ROI was then copied and pasted onto other images. The 3DUnet model training parameters were: initial learning rate of 0.1; epochs of 30.
[0033] S3. Calculate local features in the tumor region of interest, and divide the tumor region of interest into different sub-regions based on the local features using the K-means clustering algorithm.
[0034] Specifically, S3 includes: The first step is to aggregate the intensity values of all voxels within the region of interest of the tumor into a global feature matrix.
[0035] The second step is to use a 3×3×3 moving window to traverse each voxel in the global feature matrix, and extract local feature vectors for each voxel. The local feature vectors include...
[0036] The third step is to divide the global feature matrix into different sub-regions based on the local feature vectors using the K-means clustering algorithm. The K-means clustering algorithm uses the silhouette coefficient, Davies-Bouldin index, and Calinski-Harabasz index to determine the optimal number of clusters in the algorithm, and uses the optimal number of clusters to obtain different sub-regions.
[0037] The formula for calculating the profile coefficient is:
[0038] in, The contour coefficient has a value range of [-1, 1]. Let be the average distance between sample point i and all sample points in the nearest neighbor cluster, representing the separation degree of that point. The average distance between sample point i and all other sample points in the same cluster represents the density of that point.
[0039] The formula for calculating the Davies-Bouldin index is:
[0040] in, The Davies-Bouldin index. Let i be the density of cluster i. Let j be the density of cluster j. As the center of cluster i, As the center of cluster j, Let be the Euclidean distance between the centers of cluster i and cluster j. N This represents the number of clusters in the clustering process.
[0041] The formula for calculating the Calinski-Harabasz index is:
[0042] in, The Calinski-Harabasz index. The inter-cluster deviation matrix, The cluster-level deviation matrix, The total number of all sample points. This represents the number of clusters in the clustering process. Let be the trace of the inter-cluster separation matrix, representing the separation degree between clusters. Let be the trace of the cluster deviation matrix, representing the density of points within the cluster.
[0043] In habitat analysis based on CT images, the optimal number of clusters can be determined by comparing the CH, DB, and SC (profile coefficient) indices under different cluster numbers K. The specific procedure is as follows: Set a range for candidate cluster numbers, such as K=2~8, perform K-means clustering for each K, and calculate the three evaluation indices. The CH index reflects the ratio of inter-cluster separation to intra-cluster compactness; a higher value indicates better clustering. The DB index measures the similarity between different clusters; a lower value indicates more compact intra-cluster clusters and greater inter-cluster separation. The SC, or silhouette coefficient, reflects the relationship between a sample and its own cluster and nearest neighbors; a value closer to 1 indicates more reasonable clustering. Therefore, theoretically, the K corresponding to the maximum CH, minimum DB, and maximum SC should be selected. If the three indices point to the same cluster number, then this K can be considered the optimal number of clusters. If the results are inconsistent, the three indices can be normalized, with DB requiring inverse normalization, and then a comprehensive score can be calculated. The K with the highest comprehensive score is then selected. Ultimately, the final number of clusters should be determined by considering the spatial continuity of habitat subregions on CT images, the presence of excessive fragmentation, and medical interpretability. This method achieves fully automated habitat subregion division and is more flexible and simpler compared to other existing clustering techniques (Gaussian mixture model, hierarchical clustering).
[0044] Figure 2 In this context, DECT stands for Dual-Energy CT; ROC stands for Receiver Operating Characteristic (ROC); NRI stands for Net Weight Classification Index; and IDI stands for Integrated Discriminant Improvement Index.
[0045] like Figure 3 The image shows the habitat generation map, including local feature types, evaluation index curves for different cluster numbers, and habitat visualization, labeled with CH, DB, and SC indices. The generated habitat region and six features are shown in (a), including Calinski-Harabasz scores, Davies-Bouldin scores, and silhouette coefficients for different cluster numbers (b), and the habitat visualization map (c). CH: Calinski-Harabasz; DB: Davies-Bouldin; SC: Silhouette coefficient.
[0046] S4. Randomly divide patients into training and test sets, and extract the optimal features from the sub-regions.
[0047] Specifically, extracting the optimal features from a sub-region includes: The first step involved using Pyradiomics software to extract 833 radiomics features from the sub-regions. These features included first-order features, shape features, and texture features. First-order features included: energy, total energy, entropy, minimum, 10th percentile, 90th percentile, maximum, mean, median, interquartile range, range, mean absolute deviation, robust mean absolute deviation, root mean square, skewness, kurtosis, variance, and homogeneity. Shape features included: mesh volume, voxel volume, surface area, surface area-to-volume ratio, sphericity, maximum 3D diameter, maximum 2D diameter (slice), maximum 2D diameter (column), maximum 2D diameter (row), major axis length, minor axis length, shortest axis length, elongation, and flatness. Texture features included: gray-level co-occurrence matrix, gray-level run-length matrix, gray-level size region matrix, gray-level size region matrix, and neighborhood gray-level difference matrix. The gray-level co-occurrence matrix includes autocorrelation, joint average, cluster prominence, cluster shading, cluster trend, contrast, correlation, difference mean, difference entropy, difference variance, joint energy, joint entropy, IMC1, IMC2, inverse difference moment, MCC, inverse difference moment, inverse normalization, inverse variance, maximum probability, sum mean, sum entropy, and sum variance. The gray-level run-length matrix includes short run emphasis, long run emphasis, gray-level non-uniformity, normalized gray-level non-uniformity, run length non-uniformity, normalized run length non-uniformity, run percentage, gray-level variance, run variance, run entropy, low gray-level run emphasis, high gray-level run emphasis, short run low gray-level emphasis, short run high gray-level emphasis, long run low gray-level emphasis, and long run high gray-level emphasis. The grayscale size region matrix includes small region emphasis, large region emphasis, grayscale non-uniformity, normalized grayscale non-uniformity, region size non-uniformity, normalized region size non-uniformity, region percentage, grayscale variance, region variance, region entropy, low grayscale region emphasis, high grayscale region emphasis, small region low grayscale emphasis, small region high grayscale emphasis, large region low grayscale emphasis, and large region high grayscale emphasis. The grayscale dependency matrix includes small dependency emphasis, large dependency emphasis, grayscale non-uniformity, dependency non-uniformity, normalized dependency non-uniformity, grayscale variance, dependency variance, dependency entropy, low grayscale emphasis, high grayscale emphasis, small dependency low grayscale emphasis, small dependency high grayscale emphasis, large dependency low grayscale emphasis, and large dependency high grayscale emphasis. The neighborhood grayscale difference matrix includes roughness, contrast, busyness, complexity, and intensity.
[0048] The second step involves performing Z-score normalization, t-test, Pearson correlation coefficient (r<0.90) screening, and LASSO model regression to screen for the optimal features of the radiomics characteristics.
[0049] The formula for Z-score normalization is:
[0050] in, The result after normalization The original data values, The mean of the data. denoted as the standard deviation of the data.
[0051] The formula for the independent two-sample t-test is:
[0052] in, This is the t-test statistic. This represents the mean of the first group of samples (e.g., the muscle layer infiltration group). This represents the mean of the second group of samples (e.g., the non-muscle layer infiltration group). To combine standard errors, This represents the size (number of samples) of the first group. This represents the size (number of samples) of the second group.
[0053] The formula for calculating the Pearson correlation coefficient is:
[0054] in, For covariance, Let x be the standard deviation of the variable. Let y be the standard deviation of the variable. This is the Pearson correlation coefficient.
[0055] The loss function for LASSO regression is:
[0056] in, The loss function for LASSO regression is... Let i be the true label of the i-th sample. For the weight vector, For feature vectors, The regularization coefficient is . Let w be the L1 norm of the weight vector w.
[0057] The LASSO model selects the optimal regularization parameter through 10-fold cross-validation based on the minimum mean square error criterion, obtains the minimum mean square error based on the optimal regularization parameter, and finds the feature corresponding to the minimum mean square error as the optimal feature.
[0058] S5. Construct a preoperative bladder cancer prediction model. The preoperative bladder cancer prediction model includes a base model and a fusion model. The base model includes a traditional radiomics model based on a multilayer perceptron classifier and a habitat model. The optimal features are used to construct the base model. The traditional radiomics model is based on features in the tumor region of interest, while the habitat model is based on features in sub-regions. The diagnostic efficacy of the base model and intermediate layer features are output. The base model is selected based on the better diagnostic efficacy. The fusion model includes an early fusion model, an integrated fusion model, a stacked fusion model, and a Transformer fusion model. The fusion model fuses the intermediate layer features to obtain the diagnostic efficacy of the fusion model. The fusion model is selected based on the better diagnostic efficacy.
[0059] The workflow for preoperative bladder cancer prediction models includes: The first step involves using a traditional radiomics model based on features within the tumor region of interest, and a habitat model based on features within a sub-region. The outputs are the diagnostic efficacy of the basal model and the intermediate layer features. The basal model is selected based on the superior diagnostic efficacy of the traditional radiomics model versus the habitat model.
[0060] The second step involves inputting the intermediate layer features into the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model, respectively, to obtain the diagnostic efficacy of the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model. The fusion model is then selected based on the superior diagnostic efficacy of these three models.
[0061] Specifically, firstly, traditional omics models and habitat models were constructed in single-energy spectral images: traditional omics features and habitat features were extracted separately, used as inputs, filtered, and then modeled to output the diagnostic efficacy of the baseline models. The efficacy of the traditional omics model and the habitat model in each image was compared. Secondly, the best baseline models (traditional omics vs. habitat) in each image were fused, and four fusion techniques were used to generate four fused models (the four fusion methods have different principles: early fusion and transformer input the features of each baseline model; ensemble fusion and stacked fusion fuse the results of each baseline model). The best model for predicting bladder cancer diagnosis was compared and selected. Traditional radiomics models extract 107 radiomics features from the entire tumor region of interest (ROI) as input, and then select the best features for modeling through a series of feature selection methods. Habitat models, on the other hand, extract 107 radiomics features from each sub-region ROI separately, and then merge them (totaling 107). The system takes n features (where n is the number of sub-regions) as input, selects the best features for modeling through a series of feature selection methods, and the input is all features, not the optimal features.
[0062] like Figure 4 As shown, NMIBC represents non-muscle-invasive bladder cancer, while MIBC represents muscle-invasive bladder cancer. The workflow of the Transformer fusion model includes: The intermediate layer features are sequentially fed into the fully connected layer, the multi-head attention layer, and the fusion layer, and the probability and result of muscle-invasive bladder cancer are output.
[0063] Specifically, features with non-zero coefficients after LASSO regularization were used for model construction. Traditional radiomics models and habitat models for images at different energy levels were constructed based on MLP classifiers, and the habitat model with better performance was selected for fusion. Four fusion strategies were employed: early fusion (feature-level integration), ensemble (averaging predicted probabilities), stacking (combining meta-classifiers), and Transformer (feature splicing along the channel dimension, capturing location information through a multi-head self-attention mechanism). The Transformer model had an initial learning rate of 0.001, a batch size of 50, 100 training epochs, and SGD as the optimizer. The diagnostic efficacy of different model fusion strategies was compared. This technique effectively fills the gap in previous analyses of model fusion techniques, proposes a scheme for comparing model fusion, and analyzes and evaluates its efficacy.
[0064] The basis model calculates six probabilities based on virtual monoenergetic images of high, medium, and low energy levels.
[0065] Early fusion models increased the feature dimension. Features extracted from each keV image were stacked in terms of both quantity and dimension before being input together. After feature selection, modeling was based on the best features.
[0066] The fusion model directly averages the prediction probabilities of multiple base models to obtain the final prediction probability.
[0067] Stacking is a two-layer model fusion method. The first layer consists of multiple base models, with the prediction results of each base model used as features as input; the second layer is a meta-learner, which is used to synthesize the outputs of the first layer models to obtain the final prediction.
[0068] S6. Use the training set to train a preoperative bladder cancer prediction model.
[0069] S7. Input the test set into the trained preoperative bladder cancer prediction model to obtain the probability and result of muscle-invasive bladder cancer.
[0070] The preoperative prediction model for bladder cancer was evaluated using AUC, sensitivity, specificity, accuracy, NRI, and IDI.
[0071] The formula for calculating AUC is:
[0072] in, The area under the receiver operating characteristic curve. The predicted probability of a positive sample is ranked among all samples. This represents the total number of positive samples. This represents the total number of negative samples.
[0073] The formula for calculating sensitivity is:
[0074] in, It is a true positive. It is a false negative.
[0075] The formula for calculating specificity is:
[0076] in, It is a true negative. This is a false positive.
[0077] The formula for calculating accuracy is:
[0078] The formula for calculating NRI (Net Weight Classification Improvement Index) is as follows:
[0079] in, The number of events correctly classified by the new model. The number of events misclassified by the old model. The number of non-events correctly classified by the old model. The number of non-events misclassified by the new model. This represents the total number of people in the event group (positive group). This represents the total number of people in the non-event group (negative group).
[0080] The formula for calculating IDI (Comprehensive Discretionary Improvement Index) is as follows:
[0081] in, This represents the average predicted probability of the new model in the event group (positive group). This represents the average predicted probability of the old model in the event group (positive group). This represents the average predicted probability of the new model in the non-event group (negative group). This represents the average predicted probability of the old model in the non-event group (negative group).
[0082] like Figure 5 As shown, the radar charts for the diagnostic performance of different fusion strategies are presented, showing the comparison of AUC, sensitivity, specificity, and accuracy of each model in the training and testing queues.
[0083] like Figure 6 The chart shows a performance comparison of the models, including ROC curves, Delong test results, NRI, and IDI indices, visually demonstrating the advantages of the Transformer model compared to other models. It includes ROC curves, Delong tests, NRI, and IDI for all features in the training cohorts (a, b, c, d) and the test cohorts (e, f, g, h). The chart also includes the ROC receiver operating characteristic curves, the NRI net reclassification index, and the IDI comprehensive discriminant improvement index.
[0084] Example like Figure 7 As shown, this embodiment provides a specific detection method, the steps of which include: S1. A retrospective study included 200 bladder cancer patients who underwent DECTU examination between September 2019 and March 2025. They were divided into a training cohort (140 cases) and a test cohort (60 cases) in a 7:3 ratio. Among them, 141 cases were non-muscle-invasive bladder cancer and 59 cases were muscle-invasive bladder cancer.
[0085] S2. Acquire intravenous phase images 70 seconds after contrast agent injection. Tube voltage, tube current and other parameters are set according to standard clinical protocol, and multi-parameter (40keV, 70keV, 100keV) images are reconstructed.
[0086] S3. Two senior radiologists measured the iodine concentration in the tumor and external iliac artery, and calculated the iodine concentration, standardized iodine concentration, slope of the energy spectrum curve, and effective atomic number. All values were averaged after three measurements.
[0087] S4. Extract six local features of the ROI from the 40, 70 and 100 keB images respectively, and divide the tumor into three habitat sub-regions by K-means clustering. The optimal clustering results are CH index 1200000, DB index 0.1 and SC index 0.3.
[0088] S5. The model was trained on an NVIDIA GeForce RTX 4070 Ti SUPER GPU using the PyTorch framework. The Transformer model achieved a sensitivity of 0.923, a specificity of 0.766, and an accuracy of 0.800 for identifying muscle-invasive bladder cancer in the test cohort. The NRI and IDI indices demonstrated significant clinical net benefit.
[0089] S6. Compared with earlier fusion, integration, stacking models and the single DECT model, the Transformer model performed best in fusion at all energy levels. Its performance was not significantly improved after being combined with the DECT model, which verifies the superiority of habitat radiomics features.
[0090] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0091] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion, characterized in that, Includes the following steps: Acquire DECT urography images of the patient and preprocess the DECT urography images into virtual monoenergetic images of high, medium and low energy levels; Delineate the region of interest of the tumor on the virtual monoenergetic image; Local features are calculated in the tumor region of interest, and the tumor region of interest is divided into different sub-regions based on the local features using the K-means clustering algorithm. Patients are randomly divided into training and testing sets, and the optimal features in the sub-regions are extracted. A preoperative bladder cancer prediction model is constructed, comprising a base model and a fusion model. The base model includes a traditional radiomics model based on a multilayer perceptron classifier and a habitat model. Optimal features are used to construct the base model. The traditional radiomics model is based on features in the tumor region of interest, while the habitat model is based on features in sub-regions. The model outputs the diagnostic efficacy of the base model and intermediate layer features. The base model is selected based on its superior diagnostic efficacy. The fusion model includes an early fusion model, an integrated fusion model, a stacked fusion model, and a Transformer fusion model. The fusion model fuses the intermediate layer features to obtain the diagnostic efficacy of the fusion model. The fusion model is selected based on its superior diagnostic efficacy. The preoperative bladder cancer prediction model was trained using the training set. The test set was input into the trained preoperative bladder cancer prediction model to obtain the probability and result of muscle-invasive bladder cancer.
2. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 1, characterized in that, The process involves calculating local features within the tumor region of interest (ROI), and then dividing the ROI into different sub-regions using a K-means clustering algorithm based on these local features. The intensity values of all voxels within the tumor region of interest are aggregated into a global feature matrix; In the global feature matrix, a 3×3×3 moving window is used to traverse each voxel, and a local feature vector is extracted for each voxel. The local feature vector includes: entropy, mean absolute deviation, median, information metric relevance, high gray level emphasis, and contrast. Based on the local feature vectors, the global feature matrix is divided into different sub-regions using the K-means clustering algorithm. The K-means clustering algorithm uses the silhouette coefficient, Davies-Bouldin index, and Calinski-Harabasz index to determine the optimal number of clusters, and uses the optimal number of clusters to obtain different sub-regions.
3. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 1, characterized in that, Extracting the optimal features from the sub-region includes: Extract radiomics features from the sub-region, the radiomics features including first-order features, shape features and texture features; The radiomics features were then subjected to Z-score normalization, t-test, Pearson correlation coefficient screening, and LASSO model regression to select the optimal features.
4. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 3, characterized in that, The LASSO model selects the optimal regularization parameter through 10-fold cross-validation based on the minimum mean square error criterion, obtains the minimum mean square error based on the optimal regularization parameter, and finds the feature corresponding to the minimum mean square error as the optimal feature.
5. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 1, characterized in that, The step of delineating the tumor region of interest on the virtual monoenergetic image includes: The tumor region of interest is delineated in a mid-energy virtual monoenergetic image using a segmentation model, and the tumor region of interest is simultaneously delineated in both a high-energy virtual monoenergetic image and a low-energy virtual monoenergetic image.
6. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 1, characterized in that, The workflow of the Transformer fusion model includes: The intermediate layer features are sequentially fed into the fully connected layer, the multi-head attention layer, and the fusion layer, and the probability and result of muscle-invasive bladder cancer are output.
7. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 1, characterized in that, The workflow of the preoperative bladder cancer prediction model includes: The traditional radiomics model is based on features in the tumor region of interest, while the habitat model is based on features in the sub-regions. The model outputs the diagnostic efficacy of the basal model and the intermediate layer features. The basal model is selected based on the superior diagnostic efficacy of the traditional radiomics model and the habitat model. The intermediate layer features are input into the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model, respectively, to obtain the diagnostic efficacy of the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model. The fusion model is selected based on the better diagnostic efficacy of the early fusion model, the integrated fusion model, the stacked fusion model, and the Transformer fusion model.
8. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 1, characterized in that, The energy levels of the high, medium, and low energy level virtual monoenergetic images are 100keV, 70keV, and 40keV, respectively.
9. The method for preoperative prediction of bladder cancer muscle invasion based on spectral CT habitat imaging and Transformer fusion according to claim 1, characterized in that, The preoperative prediction model for bladder cancer was evaluated using AUC, sensitivity, specificity, accuracy, NRI, and IDI.