Multi-mode bladder cancer recognition system, medium and computer equipment
Through the multimodal bladder cancer identification system, combined with T2-weighted, diffusion weighted and apparent diffusion coefficient images, the ViT model is used for feature extraction and fusion, which solves the problem of distinguishing tumor expansion from benign diseases, and improves the accuracy and efficiency of bladder cancer recurrence monitoring.
Patent Information
- Application Number
- CN202510191097.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to distinguish between tumor expansion and benign diseases from T2-weighted images, resulting in unnecessary cystoscopy and diagnosis difficulties in bladder cancer recurrence monitoring.
The multimodal bladder cancer recognition system is adopted, combined with T2 weighted images, diffusion-weighted imaging and apparent diffusion coefficient images, and the image is mapped to the high-dimensional feature space through the projection module. The extraction module uses the ViT model for global feature extraction. The fusion module fuses the features of different modes, and the classification module for identification to output the recognition results of recurrence or non-recurrence of bladder cancer.
It improves the accuracy of identification of bladder cancer recurrence, reduces unnecessary cystoscopy, and enhances the robustness and reliability of monitoring.
Smart Images

Figure CN120088559A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bladder cancer identification, and more specifically, it relates to a multimodal bladder cancer identification system, medium, and computer device. Background Art
[0002] In the prior art, as of 2023 globally, bladder cancer is the fourth most common cancer in men, accounting for 6% of newly diagnosed cancers, and the cases of bladder cancer leading to death account for 4% of cancer-related deaths. At the first visit, 70 - 75% of patients have non-muscle-invasive bladder cancer (NMIBC), 20 - 25% of patients have muscle-invasive bladder cancer (MIBC), and the other 5% of patients have distant metastases. Approximately 90% of bladder cancer cases are urothelial cell carcinomas, and most of the rest are squamous cell carcinomas, adenocarcinomas, or neuroendocrine carcinomas. Transurethral cystoscopy combined with pathological biopsy is the gold standard for the initial diagnosis of bladder cancer. Multiparametric magnetic resonance imaging (mpMRI) provides excellent image contrast for the preoperative staging, recurrence monitoring, and treatment response assessment of bladder cancer. Multiple studies have shown that mpMRI has advantages in evaluating whether there is muscle invasion in patients with bladder cancer, so mpMRI is the preferred imaging examination method for bladder cancer. According to the treatment guidelines issued by the European Association of Urology, the American Urological Association, the Society of Urologic Oncology, and NCCN, NMIBC patients are recommended to undergo transurethral resection of a bladder tumor (TURBT) and receive intravesical adjuvant treatment according to the risk classification. Correspondingly, MIBC patients adopt a more aggressive treatment method, including neoadjuvant and adjuvant systemic chemotherapy combined with radical cystectomy or trimodality treatment (including TURBT, radiotherapy, and chemotherapy). Due to the high recurrence rate of bladder cancer, after NMIBC and some MIBC patients choose TURBT treatment, the subsequent treatment goal is to reduce stage progression and maintain the quality of life, while maintaining close monitoring to detect the occurrence of muscle invasion. The monitoring after the first treatment mainly relies on invasive cystoscopy, which causes a great economic burden on patients and reduces the compliance of follow-up.
[0003] It is reported that mpMRI can avoid unnecessary cystoscopy for 25% of patients during postoperative follow-up. However, mpMRI still faces challenges in the monitoring after treatment. Inflammation and fibrosis after TURBT can lead to irregular thickening of the bladder wall, accompanied by perivesical infiltration. These changes have similar T2-weighted signal intensities, and differentiating perivesical tumor extension from benign diseases is a complex task for radiologists. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a multi-modal bladder cancer recognition system, medium and computer device to overcome the disadvantages in the existing technology that it is difficult to distinguish tumor extension and benign diseases from T2-weighted images.
[0005] The above technical objective of the present invention is achieved through the following technical solutions: A multi-modal bladder cancer recognition system, including:
[0006] A projection module, configured to receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space, and generate a T2-weighted embedding vector; receive a diffusion-weighted imaging, project the diffusion-weighted imaging into a high-dimensional feature space to generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, and project the apparent diffusion coefficient image into a high-dimensional feature space to generate a diffusion coefficient embedding vector;
[0007] An extraction module, configured to input the T2-weighted embedding vector into a ViT model for global feature extraction to generate a T2-weighted feature vector; input the diffusion-weighted embedding vector into a ViT model for global feature extraction to generate a diffusion-weighted feature vector; input the diffusion coefficient embedding vector into a ViT model for global feature extraction to generate a diffusion coefficient feature vector;
[0008] A fusion module, configured to fuse the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector with each other to generate a fusion feature vector;
[0009] A classification module: configured to classify according to the fusion feature vector and output an identification result, and the identification result specifically includes: bladder cancer recurrence or no bladder cancer recurrence.
[0010] In one embodiment, the projection module specifically includes:
[0011] A first projection unit, configured to divide the T2-weighted image into multiple T2-weighted sub-images, flatten each T2-weighted sub-image to generate a corresponding T2-weighted one-dimensional vector; perform a linear transformation on each of the T2-weighted one-dimensional vectors and project them into a high-dimensional feature space to generate corresponding T2-weighted embedding vectors;
[0012] A second projection unit, configured to divide the diffusion-weighted image into multiple diffusion-weighted sub-images, flatten each diffusion-weighted sub-image to generate a corresponding diffusion-weighted one-dimensional vector, and perform a linear transformation on each of the diffusion-weighted one-dimensional vectors and project them into a high-dimensional feature space to generate corresponding diffusion-weighted embedding vectors;
[0013] A third projection unit is configured to divide the diffusion coefficient image into a plurality of diffusion coefficient sub-images, flatten each diffusion coefficient sub-image to generate a corresponding one-dimensional diffusion coefficient vector, and perform a linear transformation on each of the one-dimensional diffusion coefficient vectors respectively, and project them into a high-dimensional feature space to generate corresponding diffusion-weighted embedding vectors.
[0014] In one embodiment, after dividing the T2-weighted image into a plurality of T2-weighted sub-images, it further includes: adding spatial position encoding to each of the T2-weighted sub-images respectively;
[0015] After dividing the diffusion-weighted image into a plurality of diffusion-weighted sub-images, it further includes: adding spatial position encoding to each of the diffusion-weighted sub-images respectively;
[0016] After dividing the diffusion coefficient image into a plurality of diffusion coefficient sub-images, it further includes: adding spatial position encoding to each of the diffusion coefficient sub-images respectively.
[0017] In one embodiment, the extraction module specifically includes:
[0018] A first extraction unit is configured to input each of the T2-weighted embedding vectors into a ViT model one by one, utilize the multi-head attention mechanism of the ViT model, calculate the T2-weighted attention corresponding to the T2-weighted embedding vector by each attention head respectively, and splice all the T2-weighted attentions in the channel dimension to generate a T2-weighted feature vector;
[0019] A second extraction unit is configured to input each of the diffusion-weighted embedding vectors into a ViT model one by one, utilize the multi-head attention mechanism of the ViT model, calculate the diffusion-weighted attention corresponding to the diffusion-weighted embedding vector by each attention head respectively, and splice all the diffusion-weighted attentions in the channel dimension to generate a diffusion-weighted feature vector;
[0020] A third extraction unit is configured to input each of the diffusion coefficient embedding vectors into a ViT model one by one, utilize the multi-head attention mechanism of the ViT model, calculate the diffusion coefficient attention corresponding to the diffusion coefficient embedding vector by each attention head respectively, and splice all the diffusion coefficient attentions in the channel dimension to generate a diffusion coefficient feature vector.
[0021] In one embodiment, the step of calculating the T2-weighted attention corresponding to the T2-weighted embedding vector by each attention head respectively specifically includes:
[0022] Based on the T2-weighted embedding vector X T2 Calculate the corresponding T2-weighted query vector respectively T2-weighted key vector and T2-weighted value vector
[0023]
[0024] wherein, represents the projection matrix of the query vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; represents the projection matrix of the key vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; represents the projection matrix of the value vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix;
[0025] Based on the T2-weighted query vector T2-weighted key vector perform weighted processing on the T2-weighted value vector to generate T2-weighted attention:
[0026]
[0027] represents the T2-weighted attention output by the i-th attention head, represents the dimension of the i-th attention head.
[0028] In one embodiment, the splicing all the T2-weighted attentions in the channel dimension to generate a T2-weighted feature vector specifically includes:
[0029] Splice the T2-weighted attentions output by all attention heads in the channel dimension:
[0030]
[0031] Perform a linear transformation on the spliced vector to generate a T2-weighted feature vector Output T2 :
[0032]
[0033] In one embodiment, the fusion module specifically includes:
[0034] A convolutional unit for convolving the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector respectively using a preset convolution kernel to adjust the dimensions of the three vectors;
[0035] A splicing unit for splicing the adjusted three feature vectors in the channel dimension to generate a fused feature vector.
[0036] In one embodiment, the classification module specifically includes:
[0037] The first classification unit: for classifying based on the T2-weighted feature vector to obtain a first classification result;
[0038] The second classification unit: for classifying based on the diffusion-weighted feature vector to obtain a second classification result;
[0039] The third classification unit: for classifying based on the apparent diffusion coefficient feature vector to obtain a third classification result;
[0040] The fourth classification unit: for classifying based on the fusion feature vector to obtain a fourth classification result.
[0041] A multi-modal bladder cancer recognition method, the method includes:
[0042] S1. Receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space to generate a T2-weighted embedding vector; receive diffusion-weighted imaging, project the diffusion-weighted imaging to a high-dimensional feature space to generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, project the apparent diffusion coefficient to a high-dimensional feature space to generate an apparent diffusion coefficient embedding vector;
[0043] S2. Input the T2-weighted embedding vector into a ViT model for global feature extraction to generate a T2-weighted feature vector; input the diffusion-weighted embedding vector into a ViT model for global feature extraction to generate a diffusion-weighted feature vector; input the apparent diffusion coefficient embedding vector into a ViT model for global feature extraction to generate an apparent diffusion coefficient feature vector;
[0044] S3. Fuse the T2-weighted feature vector, the diffusion-weighted feature vector, and the apparent diffusion coefficient feature vector with each other to generate a fusion feature vector;
[0045] S4. Classify according to the fusion feature vector and output an identification result, where the identification result specifically includes: bladder cancer recurrence or no bladder cancer recurrence.
[0046] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0047] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0048] In summary, the present invention has the following beneficial effects: A multimodal bladder cancer recognition system includes: a projection module; an extraction module; a fusion module; a classification module; By using the multimodal bladder cancer recognition system of the present invention, combining the three modal features of T2-weighted images, DWI, and ADC, making full use of the complementary information of different modalities, and enhancing the model's recognition ability for bladder cancer recurrence. By using deep learning feature extraction and classification techniques, the dependence on manual feature design in traditional methods is reduced, and the recognition efficiency is improved. Classifying respectively on single-modal and multimodal fusion features enhances the robustness and reliability of recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flowchart of a multimodal bladder cancer recognition method of the present invention;
[0050] Figure 2 It is a structural diagram of a multimodal bladder cancer recognition system in an embodiment of the present invention;
[0051] Figure 3 It is an internal structural diagram of a computer device in an embodiment of the present invention;
[0052] Figure 4 It is a schematic structural diagram of a projection module in an embodiment of the present invention;
[0053] Figure 5 It is a schematic structural diagram of an extraction module in an embodiment of the present invention;
[0054] Figure 6 It is a schematic structural diagram of a fusion module in an embodiment of the present invention;
[0055] Figure 7 It is a schematic structural diagram of a classification module in an embodiment of the present invention.
[0056] In the figure: 1. Projection module; 11. First projection unit; 12. Second projection unit; 13. Third projection unit; 2. Extraction module; 21. First extraction unit; 22. Second extraction unit; 23. Third extraction unit; 3. Fusion module; 31. Convolution unit; 32. Stitching unit; 4. Classification module; 41. First classification unit; 42. Second classification unit; 43. Third classification unit; 44. Fourth classification unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings. Several embodiments of the present invention are shown in the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein.
[0058] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0059] Example 1
[0060] To solve the above problems, the present invention provides a multi-modal bladder cancer recognition method, as Figure 1 shown, the method includes:
[0061] S1. Receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space to generate a T2-weighted embedding vector; receive a diffusion-weighted imaging, project the diffusion-weighted imaging to the high-dimensional feature space to generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, project the apparent diffusion coefficient to the high-dimensional feature space to generate a diffusion coefficient embedding vector;
[0062] S2. Input the T2-weighted embedding vector into a ViT model for global feature extraction to generate a T2-weighted feature vector; input the diffusion-weighted embedding vector into a ViT model for global feature extraction to generate a diffusion-weighted feature vector; input the diffusion coefficient embedding vector into a ViT model for global feature extraction to generate a diffusion coefficient feature vector;
[0063] S3. Fuse the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector with each other to generate a fused feature vector;
[0064] S4. Classify according to the fused feature vector and output a recognition result, where the recognition result specifically includes: bladder cancer recurrence or no bladder cancer recurrence.
[0065] In practical applications, in order to improve the accuracy of bladder cancer recognition, this application proposes a multi-modal bladder cancer recognition method. After obtaining the T2-weighted image using multi-parameter nuclear magnetic resonance, the diffusion-weighted image (DWI) and the apparent diffusion coefficient image (ADC) are further obtained. Among them, the T2-weighted image is formed based on the T2 relaxation time (transverse magnetization relaxation time) of tissues. Different tissues have different T2 relaxation times. Tissues with high water content (such as cerebrospinal fluid, cystic lesions) have longer T2 relaxation times and appear bright in the T2-weighted image, while tissues with low water content (such as muscle, fibrotic tissue) appear dark. DWI forms a contrast based on the Brownian motion (diffusion motion) of water molecules in tissues. Different tissues have different degrees of restriction on the diffusion of water molecules: Tissues with restricted diffusion (such as tumors with high cell density or acute cerebral infarction areas) have stronger signals (bright) in the DWI image. Tissues with free diffusion (such as cystic lesions, cerebrospinal fluid) have weaker signals (dark). The apparent diffusion coefficient image (ADC) is a quantitative index calculated from DWI, which reflects the average ability of water molecule diffusion. The calculated ADC values are mapped into a grayscale image to generate the ADC image. Low ADC values (restricted diffusion) appear dark. High ADC values (free diffusion) appear bright.
[0066] After obtaining the three images, the three images are respectively mapped into a high-dimensional space. The original image is a matrix composed of pixel values. These values themselves are only low-level raw data and lack abstract semantic information. Directly processing the original pixels may lead to insufficient ability of the model to capture high-order features (such as texture, edge, shape). After being mapped into a high-dimensional space, each dimension can represent a combination of different features in the original image, such as color, edge direction, texture pattern, etc. The high-dimensional space allows the model to extract more discriminative features by learning weights. Similar pixel features in the original data may be mapped into more distinguishable features in the high-dimensional space, thereby enhancing the discriminative ability of the model.
[0067] After mapping the three images into a high-dimensional space to generate embedding vectors, feature extraction is also required for the three embedding vectors. ViT can model the relationships between features globally through the multi-head self-attention mechanism, without relying on local receptive fields like convolutional neural networks (CNNs); each part of the high-dimensional feature vector can interact with other parts, highlighting the features of important regions through attention weights; in the bladder cancer recognition task, this global property helps to capture the key pathological information in the imaging data. Different attention heads can focus on different patterns in the high-dimensional feature vector, improving the model's ability to express complex features and robustness. The results of each head are concatenated in the channel dimension, generating more diverse features that can capture information at different scales and levels.
[0068] Data of different modalities (such as T2-weighted, DWI, ADC) reflect different aspects of the sample: T2-weighted images: mainly provide structural and anatomical information of tissues; DWI images: capture the diffusion characteristics of water molecules and are related to cell density; ADC images: quantitatively reflect the degree of diffusion restriction and are complementary to DWI. A single modality cannot comprehensively describe the characteristics of the sample, and the fusion of multi-modal features can utilize their complementarity; the features of each modality may have advantages in different discriminant spaces. Through fusion, the discriminant ability range of the model can be expanded. The fused features can better capture complex patterns (such as the recurrence characteristics of bladder cancer), thereby improving the accuracy of classification or prediction. After feature extraction, it is necessary to fuse each feature vector to more accurately judge the recurrence possibility of bladder cancer.
[0069] In one embodiment, mapping the T2-weighted image to a high-dimensional feature space to generate a T2-weighted embedding vector specifically includes: dividing the T2-weighted image into multiple T2-weighted sub-images, flattening each T2-weighted sub-image to generate a corresponding one-dimensional T2-weighted vector; performing a linear transformation on each one-dimensional T2-weighted vector respectively and projecting it into the high-dimensional feature space to generate a corresponding T2-weighted embedding vector; mapping the diffusion-weighted imaging to the high-dimensional feature space to generate a diffusion-weighted embedding vector specifically includes: dividing the diffusion-weighted image into multiple diffusion-weighted sub-images, flattening each diffusion-weighted sub-image to generate a corresponding one-dimensional diffusion-weighted vector, performing a linear transformation on each one-dimensional diffusion-weighted vector respectively and projecting it into the high-dimensional feature space to generate a corresponding diffusion-weighted embedding vector; mapping the apparent diffusion coefficient image to the high-dimensional feature space to generate a diffusion coefficient embedding vector specifically includes: dividing the diffusion coefficient image into multiple diffusion coefficient sub-images, flattening each diffusion coefficient sub-image to generate a corresponding one-dimensional diffusion coefficient vector, performing a linear transformation on each one-dimensional diffusion coefficient vector respectively and projecting it into the high-dimensional feature space to generate a corresponding diffusion-weighted embedding vector.
[0070] In practical applications, after the image is segmented into sub-images, each sub-image corresponds to a local area of the image. In this way, finer-grained local information can be extracted, such as edges, textures, and brightness changes. Local features are particularly important in medical images. For example, the boundaries of tumors or lesion areas may be concentrated in certain local areas of the image. After the image is divided into sub-images, the data blocks processed each time are smaller, which is convenient for block-by-block calculation and reduces the complexity of processing the entire image at one time; the original pixel values are low-dimensional simple numerical values, which are difficult to directly capture complex image patterns. After being mapped to a high-dimensional feature space, more abstract and complex features can be represented. The original data of different modalities may have different distribution characteristics and dimensions. After being projected into the same high-dimensional feature space, multi-modal feature alignment and fusion can be conveniently carried out.
[0071] In one embodiment, after segmenting the T2-weighted image into a plurality of T2-weighted sub-images, it further includes: adding spatial position encoding to each of the T2-weighted sub-images respectively;
[0072] After segmenting the diffusion-weighted image into a plurality of diffusion-weighted sub-images, it further includes: adding spatial position encoding to each of the diffusion-weighted sub-images respectively;
[0073] After segmenting the diffusion coefficient image into a plurality of diffusion coefficient sub-images, it further includes: adding spatial position encoding to each of the diffusion coefficient sub-images respectively.
[0074] In practical applications, based on each sub-image, encoding information related to its spatial position (Positional Encoding) is added, so that the feature vector of each sub-image not only contains its own content information but also carries its position in the entire image. After the image is segmented into a plurality of sub-images, each sub-image only retains local content information and loses its global position information in the entire image. For example, the boundary of a tumor may span multiple sub-images, and it is difficult to understand its global significance based on a single sub-image. The spatial position encoding explicitly introduces the position information of each sub-image in the entire image. The model can combine the content features and position information to restore the spatial relationship between sub-images and enhance the understanding of the global structure.
[0075] In one embodiment, inputting the T2-weighted embedding vectors into a ViT model for global feature extraction to generate T2-weighted feature vectors specifically includes: sequentially inputting each of the T2-weighted embedding vectors into the ViT model, and using the multi-head attention mechanism of the ViT model, calculating the T2-weighted attention corresponding to each T2-weighted embedding vector by using each attention head respectively, and splicing all the T2-weighted attentions in the channel dimension to generate T2-weighted feature vectors; inputting the diffusion-weighted embedding vectors into the ViT model for global feature extraction to generate diffusion-weighted feature vectors specifically includes: sequentially inputting each of the diffusion-weighted embedding vectors into the ViT model, and using the multi-head attention mechanism of the ViT model, calculating the diffusion-weighted attention corresponding to each diffusion-weighted embedding vector by using each attention head respectively, and splicing all the diffusion-weighted attentions in the channel dimension to generate diffusion-weighted feature vectors; inputting the diffusion coefficient embedding vectors into the ViT model for global feature extraction to generate diffusion coefficient feature vectors specifically includes: sequentially inputting each of the diffusion coefficient embedding vectors into the ViT model, and using the multi-head attention mechanism of the ViT model, calculating the diffusion coefficient attention corresponding to each diffusion coefficient embedding vector by using each attention head respectively, and splicing all the diffusion coefficient attentions in the channel dimension to generate diffusion coefficient feature vectors.
[0076] In practical applications, the multi-head attention mechanism of the ViT model generates multiple attention heads respectively. Each attention head processes the input vector independently, calculates the attention weights and generates the output features, dynamically capturing the correlations between the features of the embedding vectors. Each attention head focuses on different feature relationships (such as local features, global features); the features output by all attention heads are spliced according to the channel dimension to integrate the feature results of multiple attention heads. Each attention head can independently calculate the correlations between different parts of the input vector and capture long-range dependencies. This mechanism enables the model to dynamically model the global relationships between features, rather than relying solely on fixed local windows. In bladder cancer image analysis, global information (such as the overall shape of the tumor, the characteristics of the surrounding tissues) is crucial for classification and recognition. Through ViT, both local features and global features can be considered simultaneously, improving the comprehensiveness of the analysis. The outputs of the multi-head attention are spliced together to integrate the feature results of each attention head. In this way, the feature learning results of the model at different scales and in different modes are synthesized to generate richer and more diverse feature representations. ViT has flexible requirements for the form of the input data, as long as it is an embedding vector, which is suitable for processing data of different modalities. This way of processing one by one and splicing makes full use of the structural advantages of ViT, enabling it to adapt to the analysis requirements of multi-modal medical images.
[0077] In one embodiment, calculating the T2-weighted attention corresponding to the T2-weighted embedding vector by using each attention head respectively specifically includes:
[0078] Based on the T2-weighted embedding vector X T2 Calculate the corresponding T2-weighted query vector respectively T2-weighted key vector and T2-weighted value vector
[0079]
[0080] wherein, represents the projection matrix of the query vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; represents the projection matrix of the key vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; represents the projection matrix of the value vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix;
[0081] Based on the T2-weighted query vector T2-weighted key vector weight the T2-weighted value vector to generate T2-weighted attention:
[0082]
[0083] represents the T2-weighted attention output by the i-th attention head, represents the dimension of the i-th attention head.
[0084] In one embodiment, concatenating all the T2-weighted attentions in the channel dimension to generate a T2-weighted feature vector specifically includes:
[0085] Concatenate the T2-weighted attentions output by all the attention heads in the channel dimension:
[0086]
[0087] Perform a linear transformation on the concatenated vector to generate a T2-weighted feature vector Output T2 :
[0088]
[0089] In practical applications, query vectors are used to dynamically find the most relevant key vectors for information matching. Through this mechanism, the model can flexibly capture the correlations among different parts of the embedded vectors. Each attention head focuses on different feature relationships. After concatenating the outputs of multiple attention heads, the generated feature representations are more diverse and have richer information. However, the dimension after concatenation is relatively large, and direct use may lead to redundancy. The feature dimension is adjusted through linear transformation to compress unnecessary redundant features while retaining important information. The attention mechanism can capture long-range dependencies, which is impossible for traditional convolutions (with fixed receptive fields). In medical image analysis, this global modeling ability helps to identify the overall shape of tumors and their associations with surrounding tissues. The flexibility of the input form of the ViT model enables it to easily process data of different modalities. Through this operation, the generated T2-weighted feature vectors can be seamlessly fused with feature vectors of other modalities.
[0090] In one embodiment, similar to the process of feature extraction for T2-weighted embedded vectors, feature extraction is performed on diffusion-weighted embedded vectors based on the diffusion-weighted embedded vector X T2 Calculate the corresponding diffusion-weighted query vectors respectively Diffusion-weighted key vectors And diffusion-weighted value vectors
[0091]
[0092] Wherein, Represents the projection matrix of the query vector of the i-th attention head in the ViT model for feature extraction of the diffusion weight matrix; Represents the projection matrix of the key vector of the i-th attention head in the ViT model for feature extraction of the diffusion weight matrix; Represents the projection matrix of the value vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix;
[0093] Based on the diffusion-weighted query vector Diffusion-weighted key vectors Perform weighted processing on the diffusion-weighted value vector To generate diffusion-weighted attention:
[0094]
[0095] Represents the diffusion-weighted attention output by the i-th attention head, Represents the dimension of the i-th attention head.
[0096] In one embodiment, the concatenating all the diffusion-weighted attentions in the channel dimension to generate a diffusion-weighted feature vector specifically includes:
[0097] Concatenate the diffusion-weighted attentions output by all the attention heads in the channel dimension:
[0098]
[0099] Perform a linear transformation on the concatenated vector to generate the diffusion-weighted feature vector Output DWI :
[0100]
[0101] Similar to the above two processes, when calculating the diffusion coefficient feature vector, finally generate the diffusion coefficient feature vector Output ADC :
[0102]
[0103] In one embodiment, the fusing the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector with each other to generate a fused feature vector specifically includes: using a preset convolution kernel to perform convolution on the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector respectively to adjust the dimensions of the three vectors; concatenating the adjusted three feature vectors in the channel dimension to generate a fused feature vector.
[0104] In practical applications, the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector are high-dimensional feature representations extracted by the ViT model respectively. The feature dimensions of each modality may be different, and these features cannot be directly concatenated or fused. Use a 1×1 convolution kernel to perform convolution on the feature vector of each modality. The 1×1 convolution kernel only acts on the channel dimension and does not change the spatial dimension. Adjust the feature dimensions of each modality to make them consistent. The T2-weighted feature provides organizational structure information (such as the boundary and shape of a tumor). The diffusion-weighted feature reflects the characteristics of water molecule diffusion and is sensitive to changes in cell density. The diffusion coefficient feature provides quantitative information for distinguishing the nature of a tumor. Each modality focuses on different features. After fusing them, these complementary information can be comprehensively utilized to generate a more comprehensive feature representation. The fused feature vector provides a more comprehensive input for tasks such as classification and segmentation.
[0105] In one embodiment, the classification module specifically includes: a first classification unit: for classifying based on the T2-weighted feature vector to obtain a first classification result; a second classification unit: for classifying based on the diffusion-weighted feature vector to obtain a second classification result; a third classification unit: for classifying based on the diffusion coefficient feature vector to obtain a third classification result; a fourth classification unit: for classifying based on the fusion feature vector to obtain a fourth classification result.
[0106] In practical applications, the first classification result, the second classification result, the third classification result, and the fourth classification result all include two categories, namely bladder cancer recurrence or no bladder cancer recurrence. In the actual use process, in order to ensure the life safety of bladder cancer patients, it is necessary to reduce the confidence level corresponding to the bladder cancer recurrence option and increase the confidence level corresponding to the no bladder cancer recurrence option. For example, set the confidence level corresponding to bladder cancer recurrence to 0.3 and the confidence level of no bladder cancer recurrence to 0.7. If the probability of bladder cancer recurrence is greater than 30%, it is recommended that the patient undergo a cystoscopy so that the patient can take timely treatment after obtaining a definite examination result.
[0107] On this basis, the model can identify the recurrence result of bladder cancer only through the fourth classification result, or it can also combine the four classification results for judgment. First, assign a corresponding confidence level to each classification result respectively, as shown in the following table:
[0108] Bladder cancer recurrence No recurrence of bladder cancer First classification result 32% 68% Second classification result 36% 64% Third classification result 33% 67% Fourth classification result 30% 79%
[0109] According to the above table, each classification result corresponds to an output judgment result. The above four judgment results can be used for voting to improve the comprehensiveness of detection, or voting can be carried out according to all classification results. For example, when two or more of the four detection results are considered to be bladder cancer recurrence, it means that the patient needs to undergo a cystoscopy to ensure that the patient can obtain timely treatment and ensure the patient's life safety.
[0110] Embodiment Two
[0111] Please refer to Figure 2 , a multi-modal bladder cancer recognition system, including:
[0112] A projection module 1, configured to receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space to generate a T2-weighted embedding vector; receive diffusion-weighted imaging, project the diffusion-weighted imaging into a high-dimensional feature space to generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, and project the apparent diffusion coefficient image into a high-dimensional feature space to generate a diffusion coefficient embedding vector;
[0113] An extraction module 2 is used to input the T2-weighted embedding vector into a ViT model for global feature extraction to generate a T2-weighted feature vector; input the diffusion-weighted embedding vector into the ViT model for global feature extraction to generate a diffusion-weighted feature vector; input the diffusion coefficient embedding vector into the ViT model for global feature extraction to generate a diffusion coefficient feature vector;
[0114] A fusion module 3 is used to fuse the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector with each other to generate a fused feature vector;
[0115] A classification module 4 is used to classify according to the fused feature vector and output an identification result, and the identification result specifically includes: bladder cancer recurrence or no bladder cancer recurrence.
[0116] In one embodiment, the projection module specifically includes:
[0117] A first projection unit 11 is used to divide the T2-weighted image into multiple T2-weighted sub-images, flatten each T2-weighted sub-image to generate a corresponding T2-weighted one-dimensional vector; perform a linear transformation on each of the T2-weighted one-dimensional vectors and project them into a high-dimensional feature space to generate corresponding T2-weighted embedding vectors;
[0118] A second projection unit 12 is used to divide the diffusion-weighted image into multiple diffusion-weighted sub-images, flatten each diffusion-weighted sub-image to generate a corresponding diffusion-weighted one-dimensional vector, and perform a linear transformation on each of the diffusion-weighted one-dimensional vectors and project them into a high-dimensional feature space to generate corresponding diffusion-weighted embedding vectors;
[0119] A third projection unit 13 is used to divide the diffusion coefficient image into multiple diffusion coefficient sub-images, flatten each diffusion coefficient sub-image to generate a corresponding diffusion coefficient one-dimensional vector, and perform a linear transformation on each of the diffusion coefficient one-dimensional vectors and project them into a high-dimensional feature space to generate corresponding diffusion-weighted embedding vectors.
[0120] In one embodiment, after dividing the T2-weighted image into multiple T2-weighted sub-images, it further includes: adding spatial position encoding to each of the T2-weighted sub-images;
[0121] After dividing the diffusion-weighted image into multiple diffusion-weighted sub-images, it further includes: adding spatial position encoding to each of the diffusion-weighted sub-images;
[0122] After dividing the diffusion coefficient image into multiple diffusion coefficient sub-images, it further includes: adding spatial position encoding to each of the diffusion coefficient sub-images.
[0123] In one embodiment, the extraction module specifically includes:
[0124] A first extraction unit 21, configured to input each of the T2-weighted embedding vectors into the ViT model one by one, and use the multi-head attention mechanism of the ViT model to calculate the T2-weighted attention corresponding to the T2-weighted embedding vector by using each attention head respectively, and splice all the T2-weighted attentions in the channel dimension to generate a T2-weighted feature vector;
[0125] A second extraction unit 22, configured to input each of the diffusion-weighted embedding vectors into the ViT model one by one, and use the multi-head attention mechanism of the ViT model to calculate the diffusion-weighted attention corresponding to the diffusion-weighted embedding vector by using each attention head respectively, and splice all the diffusion-weighted attentions in the channel dimension to generate a diffusion-weighted feature vector;
[0126] A third extraction unit 23, configured to input each of the diffusion coefficient embedding vectors into the ViT model one by one, and use the multi-head attention mechanism of the ViT model to calculate the diffusion coefficient attention corresponding to the diffusion coefficient embedding vector by using each attention head respectively, and splice all the diffusion coefficient attentions in the channel dimension to generate a diffusion coefficient feature vector.
[0127] In one embodiment, the step of calculating the T2-weighted attention corresponding to the T2-weighted embedding vector by using each attention head respectively specifically includes:
[0128] Based on the T2-weighted embedding vector X T2 calculate the corresponding T2-weighted query vector respectively T2-weighted key vector and T2-weighted value vector
[0129]
[0130] wherein, represents the query vector projection matrix of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; represents the key vector projection matrix of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; represents the value vector projection matrix of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix;
[0131] Based on the T2-weighted query vector T2-weighted key vector for the T2-weighted value vector Perform weighted processing to generate T2-weighted attention:
[0132]
[0133] Denote the T2-weighted attention output by the i-th attention head, Denote the dimension of the i-th attention head.
[0134] In one embodiment, the step of concatenating all the T2-weighted attentions in the channel dimension to generate a T2-weighted feature vector specifically includes:
[0135] Concatenate the T2-weighted attentions output by all the attention heads in the channel dimension:
[0136]
[0137] Perform a linear transformation on the concatenated vector to generate the T2-weighted feature vector Output T2 :
[0138]
[0139] In one embodiment, the fusion module specifically includes:
[0140] A convolution unit 31 for convolving the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector respectively using a preset convolution kernel to adjust the dimensions of the three vectors;
[0141] A concatenation unit 32 for concatenating the three adjusted feature vectors in the channel dimension to generate a fused feature vector.
[0142] In one embodiment, the classification module specifically includes:
[0143] A first classification unit 41: for classifying based on the T2-weighted feature vector to obtain a first classification result;
[0144] A second classification unit 42: for classifying based on the diffusion-weighted feature vector to obtain a second classification result;
[0145] A third classification unit 43: for classifying based on the diffusion coefficient feature vector to obtain a third classification result;
[0146] A fourth classification unit 44: for classifying based on the fused feature vector to obtain a fourth classification result.
[0147] For the specific limitations of a multimodal bladder cancer recognition system, reference may be made to the limitations of a multimodal bladder cancer recognition method in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned multimodal bladder cancer recognition system can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0148] Those skilled in the art can understand that Figure 2 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the solution of this application. Specifically, a multimodal bladder cancer recognition system may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements
[0149] Embodiment III
[0150] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a multimodal bladder cancer recognition method as described in Embodiment 1.
[0151] Embodiment IV
[0152] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it implements a multimodal bladder cancer recognition system method.
[0153] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. Specifically, the computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0154] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented: including:
[0155] S1. Receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space to generate a T2-weighted embedding vector; receive a diffusion-weighted imaging, project the diffusion-weighted imaging onto the high-dimensional feature space to generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, project the apparent diffusion coefficient onto the high-dimensional feature space to generate a diffusion coefficient embedding vector.
[0156] S2. Input the T2-weighted embedding vector into a ViT model for global feature extraction to generate a T2-weighted feature vector; input the diffusion-weighted embedding vector into a ViT model for global feature extraction to generate a diffusion-weighted feature vector; input the diffusion coefficient embedding vector into a ViT model for global feature extraction to generate a diffusion coefficient feature vector.
[0157] S3. Fuse the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector with each other to generate a fused feature vector.
[0158] S4. Classify according to the fused feature vector and output an identification result, where the identification result specifically includes: bladder cancer recurrence or no bladder cancer recurrence.
[0159] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0160] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope described in this specification is covered.
[0161] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A multimodal bladder cancer identification system, characterized in that: include: A projection module is used to receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space, and generate a T2-weighted embedding vector; receive a diffusion-weighted image, project the diffusion-weighted image to a high-dimensional feature space, and generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, and project the apparent diffusion coefficient image to a high-dimensional feature space, and generate a diffusion coefficient embedding vector; An extraction module, used for inputting the T2 weighted embedded vector into the ViT model for global feature extraction to generate a T2 weighted feature vector; inputting the diffusion weighted embedded vector into the ViT model for global feature extraction to generate a diffusion weighted feature vector; inputting the diffusion coefficient embedded vector into the ViT model for global feature extraction to generate a diffusion coefficient feature vector; A fusion module, used for fusing the T2-weighted feature vector, the diffusion-weighted feature vector and the diffusion coefficient feature vector to generate a fused feature vector; Classification module: used for classifying according to the fused feature vector and outputting a recognition result, wherein the recognition result specifically includes: bladder cancer recurrence or bladder cancer non-recurrence.
2. A multimodal bladder cancer identification system according to claim 1, characterized in that: The projection module specifically includes: A first projection unit is used to divide the T2-weighted image into a plurality of T2-weighted sub-images, flatten each T2-weighted sub-image to generate a corresponding T2-weighted one-dimensional vector; linearly transform each T2-weighted one-dimensional vector, and project it into a high-dimensional feature space to generate a corresponding T2-weighted embedding vector; A second projection unit is used to divide the diffusion weighted image into a plurality of diffusion weighted sub-images, flatten each diffusion weighted sub-image to generate a corresponding diffusion weighted one-dimensional vector, perform a linear transformation on each diffusion weighted one-dimensional vector, and project the result to a high-dimensional feature space to generate a corresponding diffusion weighted embedding vector; The third projection unit is used to divide the diffusion coefficient image into multiple diffusion coefficient sub-images, flatten each diffusion coefficient sub-image to generate a corresponding diffusion coefficient one-dimensional vector, linearly transform each of the diffusion coefficient one-dimensional vectors, and project them into a high-dimensional feature space to generate a corresponding diffusion weighted embedding vector.
3. A multimodal bladder cancer identification system according to claim 2, characterized in that: After said dividing said T2-weighted image into a plurality of T2-weighted sub-images, it further comprises: adding a spatial position code to each of said T2-weighted sub-images; After said dividing said diffusion weighted image into a plurality of diffusion weighted sub-images, it further comprises: adding a spatial position code to each of said diffusion weighted sub-images; After dividing the diffusion coefficient image into a plurality of diffusion coefficient sub-images, the method further includes: adding a spatial position code to each diffusion coefficient sub-image respectively.
4. A multimodal bladder cancer identification system according to claim 3, characterized in that: The extraction module specifically includes: A first extraction unit is used to input each of the T2 weighted embedding vectors into the ViT model one by one, use the multi-head attention mechanism of the ViT model, use each attention head to calculate the T2 weighted attention corresponding to the T2 weighted embedding vector, and concatenate all the T2 weighted attentions in the channel dimension to generate a T2 weighted feature vector; A second extraction unit is used to input each of the diffusion weighted embedding vectors into the ViT model one by one, use the multi-head attention mechanism of the ViT model, use each attention head to calculate the diffusion weighted attention corresponding to the diffusion weighted embedding vector, and splice all the diffusion weighted attentions in the channel dimension to generate a diffusion weighted feature vector; The third extraction unit is used to input each of the diffusion coefficient embedding vectors into the ViT model one by one, and use the multi-head attention mechanism of the ViT model to calculate the diffusion coefficient attention corresponding to the diffusion coefficient embedding vector using each attention head, and splice all the diffusion coefficient attentions in the channel dimension to generate a diffusion coefficient feature vector.
5. A multimodal bladder cancer identification system according to claim 4, characterized in that: The calculating the T2 weighted attention corresponding to the T2 weighted embedding vector using each attention head respectively specifically includes: Based on the T2 weighted embedding vector X T2 Calculate the corresponding T2 weighted query vectors respectively T2 weighted bond vector and T2 weighted value vector in, represents the projection matrix of the query vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; represents the projection matrix of the key vector of the i-th attention head in the ViT model for feature extraction of the T2 weight matrix; Represents the projection matrix of the i-th attention head value vector in the ViT model for feature extraction of the T2 weight matrix; Based on the T2 weighted query vector T2 weighted bond vector The T2 weighted value vector Perform weighted processing to generate T2 weighted attention: represents the T2-weighted attention output by the i-th attention head, represents the dimension of the i-th attention head.
6. A multimodal bladder cancer identification system according to claim 5, characterized in that: The method concatenates all T2-weighted attentions in the channel dimension to generate a T2-weighted feature vector, specifically including: Concatenate the T2 weighted attention outputs of all attention heads in the channel dimension: Perform linear transformation on the concatenated vector to generate the T2-weighted feature vector Output T2 :
7. A multimodal bladder cancer identification system according to claim 1, characterized in that: The fusion module specifically includes: A convolution unit, used to convolve the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector respectively using a preset convolution kernel to adjust the dimensions of the three vectors; The concatenation unit is used to concatenate the three adjusted feature vectors in the channel dimension to generate a fused feature vector.
8. A multimodal bladder cancer identification system according to claim 7, characterized in that: The classification module specifically includes: A first classification unit: used for performing classification based on the T2 weighted feature vector to obtain a first classification result; A second classification unit: used for performing classification based on the diffusion weighted feature vector to obtain a second classification result; A third classification unit: used for performing classification based on the diffusion coefficient feature vector to obtain a third classification result; The fourth classification unit is used to perform classification based on the fused feature vector to obtain a fourth classification result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a multimodal bladder cancer identification method is implemented, the method comprising: Receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space, and generate a T2-weighted embedding vector; receive a diffusion-weighted image, project the diffusion-weighted image to a high-dimensional feature space, and generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, and project the apparent diffusion coefficient to a high-dimensional feature space, and generate a diffusion coefficient embedding vector; Inputting the T2 weighted embedding vector into the ViT model for global feature extraction to generate a T2 weighted feature vector; inputting the diffusion weighted embedding vector into the ViT model for global feature extraction to generate a diffusion weighted feature vector; inputting the diffusion coefficient embedding vector into the ViT model for global feature extraction to generate a diffusion coefficient feature vector; Fusion the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector to generate a fused feature vector; Classification is performed according to the fused feature vector, and a recognition result is outputted. The recognition result specifically includes: bladder cancer recurrence or bladder cancer non-recurrence.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, a multimodal bladder cancer identification method is implemented, the method comprising: Receive a T2-weighted image, map the T2-weighted image to a high-dimensional feature space, and generate a T2-weighted embedding vector; receive a diffusion-weighted image, project the diffusion-weighted image to a high-dimensional feature space, and generate a diffusion-weighted embedding vector; receive an apparent diffusion coefficient image, and project the apparent diffusion coefficient to a high-dimensional feature space, and generate a diffusion coefficient embedding vector; Inputting the T2 weighted embedding vector into the ViT model for global feature extraction to generate a T2 weighted feature vector; inputting the diffusion weighted embedding vector into the ViT model for global feature extraction to generate a diffusion weighted feature vector; inputting the diffusion coefficient embedding vector into the ViT model for global feature extraction to generate a diffusion coefficient feature vector; Fusion the T2-weighted feature vector, the diffusion-weighted feature vector, and the diffusion coefficient feature vector to generate a fused feature vector; Classification is performed according to the fused feature vector, and a recognition result is outputted. The recognition result specifically includes: bladder cancer recurrence or bladder cancer non-recurrence.