Bladder cancer probability prediction system, medium and computer equipment
By dividing the MRI bladder cancer follow-up image into 3D image blocks and inputting the Transformer model for feature extraction, the problem that bladder wall inflammation and fibrosis affect the accuracy of tumor recurrence judgment after surgery is solved, and a more accurate prediction of bladder cancer recurrence and reducing the frequency of cystoscopy was achieved.
Patent Information
- Application Number
- CN202510191098.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, inflammation and fibrosis of the bladder wall after surgery lead to irregular thickening of the bladder wall, affecting the accuracy of judging whether the tumor recurs based on T2-weighted images.
A bladder cancer probability prediction system is adopted, including slicing units, processing units, extraction units and classification units. The system divides the MRI bladder cancer follow-up image into multiple 3D image blocks, flattened and projected. After adding position encoding, the image is embedded and input into the Transformer model for feature extraction, and finally classified according to the features to obtain the predicted probability of bladder cancer recurrence.
Through the capture and analysis of global information, the system can more accurately judge the probability of recurrence of bladder cancer, reduce the frequency of cystoscopy, and improve the quality of life of cancer patients.
Smart Images

Figure CN120088560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cancer identification, and more specifically, it relates to a bladder cancer probability prediction system, medium and computer device. Background Art
[0002] As of 2023 globally, bladder cancer is the fourth most common cancer in men, accounting for 6% of newly diagnosed cancers, and the cases of bladder cancer leading to death account for 4% of cancer-related deaths. At the first visit, 70 - 75% of patients have non-muscle-invasive bladder cancer (NMIBC), 20 - 25% of patients have muscle-invasive bladder cancer (MIBC), and another 5% of patients have distant metastases. Approximately 90% of bladder cancer cases are urothelial cell carcinomas, and most of the rest are squamous cell carcinomas, adenocarcinomas or neuroendocrine carcinomas.
[0003] Transurethral cystoscopy plus pathological biopsy is the gold standard for the initial diagnosis of bladder cancer. However, cystoscopy is an invasive procedure. The patient needs to insert a cystoscope through the urethra into the bladder, which may cause local pain, discomfort, and even in some cases, urinary tract infection or urethral injury. The cystoscope mainly provides a view of the bladder cavity, but it cannot well evaluate whether the tumor has invaded the muscular layer of the bladder wall.
[0004] Multiparametric magnetic resonance imaging (mpMRI) provides excellent image contrast for the preoperative staging, recurrence monitoring, and assessment of treatment response in bladder cancer. Multiple studies have shown that mpMRI has advantages in evaluating the presence of muscle invasion in patients with bladder cancer. Therefore, mpMRI is the preferred imaging examination method for bladder cancer. According to the treatment guidelines issued by the European Association of Urology, the American Urological Association, the Society of Urologic Oncology, and NCCN, patients with non-muscle-invasive bladder cancer (NMIBC) are recommended to undergo transurethral resection of a bladder tumor (TURBT) and receive intravesical adjuvant therapy according to risk stratification. Correspondingly, patients with muscle-invasive bladder cancer (MIBC) adopt more aggressive treatment methods, including neoadjuvant and adjuvant systemic chemotherapy combined with radical cystectomy or trimodality therapy (including TURBT, radiotherapy, and chemotherapy). Due to the high recurrence rate of bladder cancer, after NMIBC and some MIBC patients choose TURBT treatment, the subsequent treatment goal is to reduce stage progression and maintain quality of life while maintaining close monitoring to detect the occurrence of muscle invasion. The monitoring after the first treatment mainly relies on invasive cystoscopy, which causes a great economic burden on patients and reduces the compliance of follow-up. Some health economics scholars have also pointed out that bladder cancer is one of the most expensive malignant tumors. As described above, mpMRI performs excellently in the first preoperative evaluation of bladder cancer. It is reported that during postoperative follow-up, mpMRI can avoid unnecessary cystoscopy in 25% of patients.
[0005] T2-weighted signal is a type of image contrast in magnetic resonance imaging (MRI), which generates images based on the differences in the content of different types of water and the movement of water molecules in tissues. In MRI, T2-weighted imaging produces image contrast by emphasizing the T2 relaxation time in tissues. The T2 relaxation time reflects the time for the magnetization intensity of water molecules in tissues to decay after being stimulated. Different tissues have different T2 relaxation times, especially between tissues with different water contents. This difference can be used to distinguish different types of tissues and lesions. In T2-weighted imaging, tissues with abundant water content (such as fluid, cerebrospinal fluid, inflammatory areas, etc.) usually show high signal, that is, the bright areas in the image. Tissues with less water content, such as fibrosis or bone, show low signal and usually appear as dark areas in the image.
[0006] Inflammation and fibrosis of the bladder wall after surgery can lead to irregular thickening of the bladder wall. These changes usually produce T2-weighted signals similar to those of tumor tissue and may appear as high-signal regions. This is because the inflammatory area and certain fibrotic tissues also contain more water, resulting in a T2 relaxation time similar to that of tumor tissue. Due to the similar signal intensity of inflammation, fibrosis, and perivesical infiltration to the T2-weighted signal of the tumor, this poses a difficult task for radiologists to distinguish perivesical tumor extension from benign diseases. This similarity makes it challenging to determine tumor recurrence solely based on T2-weighted images. Summary of the Invention
[0007] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a bladder cancer probability prediction system, medium, and computer device to overcome the disadvantages in the existing technology that inflammation and fibrosis of the bladder wall after surgery can lead to irregular thickening of the bladder wall, affecting the accuracy of relying on T2-weighted images to judge tumor recurrence.
[0008] The above technical objectives of the present invention are achieved through the following technical solutions: In the first aspect, a bladder cancer probability prediction system includes the following steps:
[0009] A segmentation unit, configured to respond to the input of an MRI bladder cancer follow-up image and segment the MRI bladder cancer follow-up image into multiple 3D image blocks;
[0010] A processing unit, configured to flatten each 3D image block and project the flattened 3D image block onto a predetermined dimension to obtain a corresponding 3D image embedding;
[0011] An extraction unit, configured to add positional encoding to each 3D image embedding respectively, and then input all the 3D image embeddings into the encoder of a Transformer model for feature extraction to obtain 3D image features;
[0012] A classification unit, configured to classify according to the 3D image features to obtain the bladder cancer recurrence prediction probability.
[0013] In one embodiment, the segmentation of the MRI image into multiple 3D image blocks specifically includes: segmenting the MRI image into multiple non-overlapping 3D blocks.
[0014] In one embodiment, flattening each 3D image block and projecting the flattened 3D image block onto a predetermined dimension to obtain a corresponding 3D image embedding specifically includes:
[0015] Flattening each 3D image block into a one-dimensional vector;
[0016] Project the flattened 3D image patches onto a predetermined dimension using a fully connected layer to obtain corresponding 3D image embeddings.
[0017] In one embodiment, adding positional encoding to each of the 3D image embeddings specifically includes:
[0018] Obtain the position information of the 3D image patches in the original space;
[0019] Based on the position information, add positional encoding to the 3D image embeddings corresponding to each of the 3D image patches.
[0020] In one embodiment, inputting all the 3D image embeddings into the encoder of the Transformer model for feature extraction to obtain 3D image features specifically includes:
[0021] Combine a preset first CLS token and all the 3D image embeddings into a feature sequence;
[0022] Control the feature sequence to sequentially pass through multiple encoding layers for feature extraction to obtain 3D image features.
[0023] In one embodiment, controlling the feature sequence to sequentially pass through multiple encoding layers for feature extraction to obtain 3D image features specifically includes:
[0024] Use the multi-head attention in the encoding layer to calculate the global dependencies in the feature sequence, capture the correlation between each element and other elements in the feature sequence, and obtain the calculation result of the multi-head attention;
[0025] Perform the first residual connection and the first layer normalization on the calculation result output by the multi-head attention;
[0026] Use a feed-forward neural network to perform non-linear feature transformation on the calculation result after the first residual connection and the first layer normalization;
[0027] Perform the second residual connection and the second layer normalization on the calculation result after the non-linear feature transformation to obtain 3D image features.
[0028] In one embodiment, using the multi-head attention in the encoding layer to calculate the global dependencies in the feature sequence, capture the correlation between each element and other elements in the feature sequence, and obtain the calculation result of the multi-head attention specifically includes:
[0029] Use the query weight matrix, key weight matrix, and value weight matrix in the multi-head attention mechanism of the encoding layer to perform linear transformation on each element in the feature sequence respectively to obtain the query vector, key vector, and value vector corresponding to each element, including:
[0030] Q i = X i W Q , K i = X i W K , V i = X i W V ;
[0031] Among them, X i represents the i-th element in the said feature sequence; W Q represents the query weight matrix; Q i represents the i-th query vector; W K represents the key weight matrix; K i represents the i-th key vector; W V represents the value weight matrix; V i represents the i-th value vector;
[0032] Calculate the attention scores between each query vector Q i and each key vector K j specifically including:
[0033]
[0034] Among them, Q i represents the i-th query vector; K j represents the j-th key vector; represents the transpose matrix of the j-th key vector; d k represents the dimension of the key vector;
[0035] According to the attention scores between each query vector Q i and each key vector k k calculate the corresponding attention weights, including:
[0036]
[0037] Use the attention weight W(Q i , K j ) to calculate the weighted value vectors corresponding to each query vector Q i specifically including:
[0038]
[0039] Concatenate all the weighted value vectors to obtain a weighted value vector sequence; adjust the dimension of the weighted value vector sequence so that the dimension of the weighted value vector sequence is the same as that of the feature sequence to obtain the calculation result of the multi-head attention, specifically including:
[0040] MultiheadOutput = Concat(Head 1 , Head 2 , …, Head h )W O ;
[0041] Wherein, W O represents the linear transformation weight matrix, which is used to adjust the dimension of the weighted value vector sequence after concatenation; Concat represents the concatenation operation.
[0042] In one embodiment, the first residual connection and the first layer normalization processing are performed on the calculation result of the multi-head attention output, specifically including:
[0043] Output residual = LayerNorm[X + MultiheadOutput(X)];
[0044] Wherein, LayerNorm represents performing layer normalization processing; X represents the feature sequence; MultiheadOutput(X) represents the calculation result of the multi-head attention mechanism when the input is X.
[0045] In one embodiment, the non-linear feature transformation is performed on the calculation result after the first residual connection and the first layer normalization processing, specifically including:
[0046] FNN(Output residual ) = max(0, Output residual W 1 + b 1 )W 2 + b 2 ;
[0047] Wherein, W 1 represents the weight matrix of the first layer linear transformation in the feed-forward neural network; b 1 is the bias term of the first layer linear transformation; max(·) represents the ReLU activation function, which is used to change the negative part of the result after the first linear transformation to 0 and retain the positive part; W 2 represents the weight matrix of the second layer linear transformation in the feed-forward neural network; b 2 is the bias term of the second layer linear transformation.
[0048] In one embodiment, the second residual connection and the second layer normalization processing are performed on the calculation result after the non-linear feature transformation to obtain the 3D image feature, specifically including:
[0049] Output FNN= LayerNorm[Output residual + FNN(Output residual )];
[0050] Among them, LayerNorm represents performing layer normalization processing; Output residual represents the calculation result of the multi-head attention output; FNN(Output residual ) represents the conversion result of the feed-forward neural network.
[0051] In one embodiment, the classifying according to the 3D image features to obtain the bladder cancer recurrence prediction probability specifically includes:
[0052] After obtaining the feature sequence and generating 3D image features through feature extraction by passing through multiple encoding layers in sequence, obtaining the second CLS token generated by the feature extraction of the first CLS token through multiple encoding layers;
[0053] Inputting the second CLS token into the classification head for classification to obtain a score vector:
[0054] Logits = CLS · W cls + b cls ;
[0055] Among them, CLS represents the second CLS token; W cls represents the weight matrix in the classification head; b cls represents the bias term in the classification head, used to adjust the mapping result; Logits represents the vector composed of the scores of all results;
[0056] According to the score vector, calculate the probability corresponding to each prediction result respectively:
[0057]
[0058] Among them, P(class i ) represents the probability of the i-th prediction result; Logit i represents the score of the i-th prediction result; represents the exponential of e of the score of the i-th prediction result; represents the sum of the exponentials of e of the scores of all prediction results.
[0059] In the second aspect, a method for predicting the probability of bladder cancer includes:
[0060] In response to the input of the MRI bladder cancer follow-up visit image, segmenting the MRI bladder cancer follow-up visit image into multiple 3D image blocks;
[0061] Flatten each 3D image patch, project the flattened 3D image patch onto a predetermined dimension to obtain a corresponding 3D image embedding;
[0062] After adding positional encoding to each 3D image embedding respectively, input all the 3D image embeddings into the encoder of the Transformer model for feature extraction to obtain 3D image features;
[0063] Classify according to all the 3D image features to obtain the recurrence prediction probability of bladder cancer.
[0064] In a third aspect, a computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0065] In a fourth aspect, a computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0066] In summary, the present invention has the following beneficial effects: A bladder cancer probability prediction system, the system includes: a splitting unit for splitting the MRI bladder cancer follow-up image into multiple 3D image patches in response to the input of the MRI bladder cancer follow-up image; a processing unit for flattening each 3D image patch and projecting the flattened 3D image patch onto a predetermined dimension to obtain a corresponding 3D image embedding; an extraction unit for inputting all the 3D image embeddings into the encoder of the Transformer model for feature extraction to obtain 3D image features; a classification unit for classifying according to the 3D image features to obtain the recurrence prediction probability of bladder cancer; by using the prediction system of the present invention, during the recognition and judgment process of the model, it is possible to more comprehensively judge the recurrence probability of bladder cancer from global information, more effectively provide a diagnostic basis for doctors, ensure the life and health of patients, while reducing the usage frequency of invasive surgeries such as cystoscopy and improving the quality of life of cancer patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a flowchart of a method for predicting the probability of bladder cancer according to the present invention;
[0068] Figure 2 It is a structural diagram of a device for predicting the recurrence of bladder cancer in an embodiment of the present invention;
[0069] Figure 3 It is an internal structural diagram of a computer device in an embodiment of the present invention;
[0070] In the figure: 1. Splitting unit; 2. Processing unit; 3. Extraction unit; 4. Classification unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of specific embodiments of the present invention with reference to the accompanying drawings. Several embodiments of the present invention are shown in the accompanying drawings. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein.
[0072] In the present invention, unless otherwise clearly defined and limited, terms such as "installed", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention may be understood according to specific circumstances. The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0073] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may include the first and second features being in direct contact, or may include the first and second features not being in direct contact but being in contact through additional features therebetween. Moreover, the first feature being "above", "over", and "on top of" the second feature includes the first feature being directly above and diagonally above the second feature, or merely indicating that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below", and "beneath" the second feature includes the first feature being directly below and diagonally below the second feature, or merely indicating that the first feature has a lower horizontal height than the second feature. Terms such as "vertical", "horizontal", "left", "right", "up", "down", and similar expressions are only for the purpose of illustration and do not indicate or imply that the indicated device or element must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention.
[0074] The present invention will be described in detail below in conjunction with the accompanying drawings and embodiments. For ease of understanding of the embodiments, the related technologies will be described first. The traditional convolutional neural network (CNN) uses convolutional kernels to perform convolution on the input data. Each convolutional kernel can only process the pixels in a local area, and its ability to capture global features is weak. It is necessary to stack multiple layers of convolutions to gradually form an understanding of the global situation. This way of learning global features layer by layer results in some global dependencies not being well captured, and it is difficult to capture the global information of the MRI bladder cancer follow-up images during the bladder cancer prediction process. The core assumption of the convolution operation is local translational invariance, that is, no matter where the features in the image appear, the convolutional kernel can detect them. However, due to this characteristic, it is difficult to capture the long-range dependencies between local features. The design of the convolution assumes that the local features of the image (such as edges, textures, etc.) can share the same convolutional kernel at different positions. In complex visual tasks, this strongly structured local feature extraction may not be the optimal solution.
[0075] Embodiment 1
[0076] To solve the above problems, the present invention provides a method for predicting the probability of bladder cancer, as Figure 1 shown, including:
[0077] S1. In response to the input of the MRI bladder cancer follow-up image, the MRI bladder cancer follow-up image is sliced into multiple 3D image patches;
[0078] S2. Each 3D image patch is flattened, and the flattened 3D image patch is projected onto a predetermined dimension to obtain a corresponding 3D image embedding;
[0079] S3. After adding position encoding to each 3D image embedding respectively, all the 3D image embeddings are input into the encoder of the Transformer model for feature extraction to obtain 3D image features;
[0080] S4. Classification is performed according to the 3D image features to obtain the prediction probability of bladder cancer recurrence.
[0081] In this embodiment, in order to solve the disadvantages of traditional convolutional neural networks in the process of image classification, the applicant proposes a method for predicting bladder cancer recurrence based on ViTs (Visual Transformer). The ViTs model uses a self-attention mechanism, which can capture the dependencies between different positions in the image. For a task like bladder cancer recurrence prediction that requires global information, the ViTs model is better able to understand the complex pathological features in the bladder image as a whole, especially the global changes in the morphology, location, and edges of tumors. To use ViTs to determine whether there is a recurrence of bladder cancer in patients with MRI bladder cancer follow-up images, it is first necessary to obtain MRI bladder cancer follow-up images and segment the MRI bladder cancer follow-up images, which can decompose the overall MRI bladder cancer follow-up image into several separate small blocks. Since MRI generates three-dimensional images, but these three-dimensional images are usually represented by multiple two-dimensional slices, in the actual segmentation process, the small blocks formed are 3D image blocks. Since the Transformer model is used to process natural language tasks, it is necessary to convert the image data into a similar data format in order to process it accurately. First, the 3D image blocks are flattened into one-dimensional vectors to convert the image blocks into a representation similar to natural language, and then the one-dimensional vectors are mapped to a specific space to obtain 3D image embeddings, so that the image blocks can meet the input format requirements of the Transformer model. In order to enable the model to perceive the global information in the three-dimensional MRI image and capture the dependencies between different regions, it is also necessary to add corresponding position encodings to the segmented 3D image embeddings. To avoid the position encoding increasing the complexity that the model needs to process and wasting computing resources, first construct a position encoding matrix according to the position features and the dimension of the 3D image embeddings, and then add the position encoding matrix to the matrix corresponding to the 3D image blocks to obtain 3D image blocks containing position encoding information; input the 3D image blocks into the encoder of the ViTs model for encoding processing, and 3D image features can be obtained. By classifying the 3D image features, it can be determined whether there is a recurrence of bladder cancer according to the classification results, specifically including bladder cancer recurrence or no recurrence of bladder cancer. Using the technical method of this application, it is possible to more comprehensively determine whether there is a recurrence of bladder cancer from global information and reduce the frequency of use of invasive surgeries such as cystoscopy. When the model determines that there is a risk of bladder cancer recurrence, the patient is recommended to use cystoscopy to further examine the internal conditions of the bladder.
[0082] In one embodiment, the step of slicing the MRI image into multiple 3D image blocks specifically includes: slicing the MRI image into multiple non-overlapping 3D blocks.
[0083] In practical applications, the MRI image is sliced into multiple non-overlapping 3D blocks to reduce redundant and repetitive information. Since adjacent pixels in medical images may have strong correlations, the cutting method of overlapping 3D image blocks will lead to information redundancy and increase the computational difficulty of the model.
[0084] In one embodiment, each 3D image block is flattened, and the flattened 3D image block is projected onto a predetermined dimension to obtain a corresponding 3D image embedding, which specifically includes: flattening each 3D image block into a one-dimensional vector; using a fully connected layer to project the flattened 3D image block onto a predetermined dimension to obtain a corresponding 3D image embedding.
[0085] In practical applications, since the Transformer model is designed to process natural language processing tasks, the input is a series of discrete tokens, and each token is embedded as a fixed-length vector. To apply this sequence processing ability to images, the image data must be converted into a data format similar to a sequence. Therefore, the 3D image block needs to be flattened into a one-dimensional vector to convert the sliced 3D image block into a representation similar to natural language. To make the input data meet the requirements of the ViTs model, the flattened one-dimensional vector also needs to be mapped to a specific high-dimensional space to ensure that the representation of each 3D image block is a unified vector.
[0086] In one embodiment, adding position encoding to each 3D image embedding specifically includes: obtaining the position information of the 3D image block in the original space; based on the position information, adding position encoding to the 3D image embedding corresponding to each 3D image block.
[0087] In practical applications, after the data input to the Transformer model is segmented or embedded, it lacks explicit position information. To enable the model to understand the relative positions of each block in the input sequence, position encoding needs to be added to each 3D graphic embedding; position encoding is usually used to represent the order or spatial position information of the input data. Especially for image data, adding position encoding to the segmented small blocks can enable the model to understand the specific positions of each small block in the overall MRI image. In this embodiment, the MRI image is a 3D image represented in the form of planar slices. Therefore, the position encoding added to the 3D image embedding is also a position encoding for marking spatial coordinates. To avoid the position encoding added to the 3D image embedding increasing the complexity of the data and affecting the computational efficiency of the model, the position encoding is added in an additive manner, adding the original 3D image embedding and a position encoding vector of the same dimension, so as to add position encoding while avoiding increasing the complexity of the data.
[0088] In one embodiment, embedding all 3D images into the encoder of the Transformer model for feature extraction to obtain 3D image features specifically includes: combining a preset first CLS token and all 3D image embeddings into a feature sequence; controlling the feature sequence to sequentially pass through multiple encoding layers for feature extraction to obtain 3D image features.
[0089] In practical applications, the encoder of the Transformer model includes multiple encoding layers; each encoding layer performs self-attention calculation and feed-forward neural network processing on the input data, and multiple encoding layers can sequentially process the input data to capture more complex features layer by layer. In this application, the input data is specifically 3D image embeddings added with positional encoding, and the output calculation result of each encoding layer will be used as the input content of the next encoding layer. The encoding layers in the first few layers can usually capture low-level features in the input data, and the encoding layers located at the later positions can further capture more complex features. And each encoding layer includes a multi-head self-attention mechanism, which captures different correlations in the input sequence through multiple attention heads; the stacking of multiple encoding layers can comprehensively capture the features in the input data at different levels, different angles, and different spatial scales.
[0090] In one embodiment, controlling the feature sequence to sequentially pass through multiple encoding layers for feature extraction to obtain 3D image features specifically includes: using the multi-head attention in the encoding layer to calculate the global dependencies in the feature sequence, capturing the correlations between each element in the feature sequence and other elements, to obtain the calculation result of the multi-head attention; performing the first residual connection and the first layer normalization processing on the calculation result output by the multi-head attention; using a feed-forward neural network to perform non-linear feature transformation on the calculation result after the first residual connection and the first layer normalization processing; performing the second residual connection and the second layer normalization processing on the calculation result after the non-linear feature transformation to obtain 3D image features.
[0091] In practical applications, the main role of the multi-head attention mechanism is to capture the relationships and dependencies between various 3D image embeddings in the input sequence, especially the dependency between the 3D image embeddings and the overall MRI image. It enables the model to focus on different features in the input data from different subspaces and perspectives and integrate this information to generate better representations. The main role of the first residual connection is to directly add the input of the encoding layer to the output to form a skip connection, thereby preventing the common problems of vanishing gradients or exploding gradients in deep networks and enhancing the training stability of the model. The main role of layer normalization is to standardize the representation of each input token, prevent numerical anomalies, stabilize the training of the model, and accelerate the training process. The feed-forward neural network performs independent non-linear transformations on each input element to further abstract and extract features, helping the model capture deeper non-linear features.
[0092] In one embodiment, calculating the global dependencies in the feature sequence by using the multi-head attention in the encoding layer, capturing the correlation between each element and other elements in the feature sequence, and obtaining the calculation result of the multi-head attention specifically includes:
[0093] Using the query weight matrix, key weight matrix, and value weight matrix in the multi-head attention mechanism in the encoding layer to perform linear transformations on each element in the feature sequence respectively, obtaining a query vector, a key vector, and a value vector corresponding to each element, including:
[0094] Q i =X i W Q ,K i =X i W K ,V i =X i W V ;
[0095] Wherein, X i represents the i-th element in the feature sequence; W Q represents the query weight matrix; Q i represents the i-th query vector; W K represents the key weight matrix; K i represents the i-th key vector; W V represents the value weight matrix; V i represents the i-th value vector;
[0096] Calculating the attention scores between each query vector Q i and each key vector K j specifically includes:
[0097]
[0098] Among them, Q i represents the i-th query vector; K j represents the j-th key vector; represents the transpose matrix of the j-th key vector; d k represents the dimension of the key vector;
[0099] According to the attention scores between each query vector Q i and each key vector K j , calculate the corresponding attention weights, including:
[0100]
[0101] Use the attention weights W(Q i , K j ) to calculate the weighted value vectors corresponding to each query vector Q i respectively, specifically including:
[0102]
[0103] Specifically, the multi-head attention mechanism processes the input in different subspaces in parallel through multiple attention heads, enabling the model to capture different features in the input from multiple perspectives and subspaces. Each attention head independently processes the input from different subspaces, generating multiple feature representations. In each independent head, the respective query, key, and value vectors are calculated, and the self-attention mechanism is executed to obtain multiple different output representations. The output results of all attention heads are concatenated together and linearly transformed to generate the final attention output.
[0104] Concatenate all the weighted value vectors to obtain a sequence of weighted value vectors; adjust the dimension of the sequence of weighted value vectors so that the dimension of the sequence of weighted value vectors is the same as the dimension of the feature sequence to obtain the calculation result of the multi-head attention, specifically including:
[0105] MultiheadOutput = Concat(Head 1 , Head 2 , …, Head h )W O ;
[0106] where, W O represents the linear transformation weight matrix for adjusting the dimension of the concatenated sequence of weighted value vectors; Concat represents the concatenation operation.
[0107] In practical applications, after concatenating the outputs of multiple attention heads, the final representation is generated. This representation integrates the feature information from multiple subspaces and represents the global features of the input data.
[0108] In one embodiment, the first residual connection and the first layer normalization processing on the calculation result of the multi-head attention output specifically include:
[0109] Output residual = LayerNorm[X + MultiheadOutput(X)];
[0110] Wherein, LayerNorm represents performing layer normalization processing; X represents a feature sequence; MultiheadOutput(X) represents the calculation result of the multi-head attention mechanism when the input is X.
[0111] In practical applications, the main role of layer normalization is to standardize the representation of each input, prevent numerical anomalies, stabilize the training of the model, and accelerate the training process. By normalizing the activation of the layer, the output of each layer is maintained within a relatively stable numerical range, thereby improving the stability of training and the convergence speed of the model.
[0112] In one embodiment, the use of the feed-forward neural network to perform non-linear feature transformation on the calculation result after the first residual connection and the first layer normalization processing specifically includes:
[0113] FNN(Output residual ) = max(0, Output residual W 1 + b 1 )W 2 + b 2 ;
[0114] Wherein, W 1 represents the weight matrix of the first linear transformation in the feed-forward neural network; b 1 is the bias term of the first linear transformation; max(·) represents the ReLU activation function, which is used to change the negative part of the result after the first linear transformation to 0 and retain the positive part; W 2 represents the weight matrix of the second linear transformation in the feed-forward neural network; b 2 is the bias term of the second linear transformation.
[0115] In practical applications, the feed-forward neural network specifically includes two linear layers, and a non-linear activation function is included between the two linear layers. The first linear transformation, that is, Output residual W 1 + b 1 is used to expand the dimension of the input vector. The second linear transformation, that is, max(0, Output residual W 1 + b 1)W 2 +b 2 For shrinking the dimension back to its original size; the activation function introduces a non - linear transformation that can capture more complex patterns and features.
[0116] In one embodiment, performing a second residual connection and a second layer normalization on the calculation result after the non - linear feature transformation to obtain 3D image features, specifically including:
[0117] Output FNN =LayerNorm[Output residual +FNN(Output residual )];
[0118] Wherein, LayerNorm represents performing layer normalization processing; Output residual represents the calculation result of the output of the multi - head attention; FNN(Output residual ) represents the transformation result of the feed - forward neural network.
[0119] In practical applications, the second residual connection and the second layer normalization directly jump - connect the input data of the feed - forward neural network to the output of the feed - forward neural network, ensuring that the model still retains the original input information when processing complex transformations.
[0120] In one embodiment, classifying according to the 3D image features to obtain the recurrence prediction probability of bladder cancer, specifically including:
[0121] Obtaining the second CLS token generated by the first CLS token after feature extraction through multiple encoding layers after the feature sequence passes through multiple encoding layers to generate 3D image features;
[0122] Inputting the second CLS token into the classification head for classification to obtain a score vector:
[0123] Logits=CLS·W cls +b cls ;
[0124] Wherein, CLS represents the second CLS token; W cls represents the weight matrix in the classification head; b cls represents the bias term in the classification head, used to adjust the mapping result; Logits represents the vector composed of the scores of all results;
[0125] According to the score vector, calculating the probability corresponding to each prediction result respectively:
[0126]
[0127] Among them, P(class i ) represents the probability of the i-th prediction result; Logit i represents the score of the i-th prediction result; represents the exponential of e of the score of the i-th prediction result; represents the sum of the exponentials of e of the scores of all prediction results.
[0128] In practical applications, the prediction results mainly include: the probability of bladder cancer recurrence and the probability of no bladder cancer recurrence. Among them, the probability of bladder cancer recurrence can be used as a prediction result alone and fed back to the doctor for the diagnosis of bladder cancer recurrence, or it can be compared with the probability of no bladder cancer recurrence and then fed back to the doctor as a comprehensive result.
[0129] If the result determined by the model is bladder cancer recurrence, then the corresponding patient needs to be given a cystoscopy to obtain an accurate diagnosis result and avoid affecting the patient's life due to misdiagnosis. Therefore, in order to balance the relationship between the patient's life rights and interests and the health hazards brought by cystoscopy, it is necessary to correspondingly lower the threshold for determining bladder cancer recurrence. For example, when the probability of bladder cancer recurrence is greater than 0.3, it is determined that bladder cancer may recur, and the patient needs to be given a cystoscopy to avoid affecting the patient's life and health due to model errors.
[0130] Example Two
[0131] Please refer to Figure 2 , a bladder cancer probability prediction system, the bladder cancer recurrence prediction system includes:
[0132] The splitting unit 1 is used to respond to the input of the MRI bladder cancer follow-up image and split the MRI bladder cancer follow-up image into multiple 3D image blocks;
[0133] The processing unit 2 is used to flatten each 3D image block and project the flattened 3D image block onto a predetermined dimension to obtain a corresponding 3D image embedding;
[0134] The extraction unit 3 is used to add position encoding to each 3D image embedding respectively, and then input all the 3D image embeddings into the encoder of the Transformer model for feature extraction to obtain 3D image features;
[0135] The classification unit 4 is used to classify according to all the 3D image features to obtain the bladder cancer recurrence prediction probability.
[0136] For the specific limitations of the bladder cancer recurrence prediction system, reference can be made to the limitations of the bladder cancer recurrence prediction method in the above text, which will not be elaborated here. Each module in the above bladder cancer recurrence prediction system can be implemented in whole or in part by software, hardware, and their combinations. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0137] Those skilled in the art can understand that Figure 2 The structure shown in
[0138] Example 3
[0139] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the bladder cancer recurrence prediction method as described in Example 1 is implemented. The method includes:
[0140] S1. In response to the input of the MRI bladder cancer follow-up image, the MRI bladder cancer follow-up image is sliced into a plurality of 3D image blocks;
[0141] S2. Each 3D image block is flattened, and the flattened 3D image block is projected onto a predetermined dimension to obtain a corresponding 3D image embedding;
[0142] S3. After adding position encoding to each 3D image embedding respectively, all the 3D image embeddings are input into the encoder of the Transformer model for feature extraction to obtain 3D image features;
[0143] S4. Classification is performed according to the 3D image features to obtain the bladder cancer recurrence prediction probability.
[0144] Example 4
[0145] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 3As shown. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it implements a method for predicting the probability of bladder cancer.
[0146] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0147] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: including:
[0148] S1. In response to the input of the MRI bladder cancer follow-up image, segment the MRI bladder cancer follow-up image into a plurality of 3D image blocks;
[0149] S2. Flatten each 3D image block, and project the flattened 3D image block onto a predetermined dimension to obtain a corresponding 3D image embedding;
[0150] S3. After adding position encoding to each 3D image embedding respectively, input all the 3D image embeddings into the encoder of the Transformer model for feature extraction to obtain 3D image features;
[0151] S4. Classify according to the 3D image features to obtain the predicted probability of bladder cancer recurrence.
[0152] In one embodiment, the segmenting the MRI image into a plurality of 3D image blocks specifically includes: segmenting the MRI image into a plurality of non-overlapping 3D blocks.
[0153] In one embodiment, flattening each 3D image block and projecting the flattened 3D image block onto a predetermined dimension to obtain a corresponding 3D image embedding specifically includes:
[0154] Flatten each 3D image block into a one-dimensional vector;
[0155] Use a fully connected layer to project the flattened 3D image block onto a predetermined dimension to obtain a corresponding 3D image embedding.
[0156] In one embodiment, adding positional encoding to each 3D image embedding specifically includes:
[0157] Obtain the position information of the 3D image patch in the original space;
[0158] Based on the position information, add positional encoding to the 3D image embeddings corresponding to each 3D image patch respectively.
[0159] In one embodiment, inputting all the 3D image embeddings into the encoder of the Transformer model for feature extraction to obtain 3D image features specifically includes:
[0160] Combine a preset first CLS token and all the 3D image embeddings into a feature sequence;
[0161] Control the feature sequence to sequentially pass through multiple encoding layers for feature extraction to obtain 3D image features.
[0162] In one embodiment, controlling the feature sequence to sequentially pass through multiple encoding layers for feature extraction to obtain 3D image features specifically includes:
[0163] Use the multi-head attention in the encoding layer to calculate the global dependencies in the feature sequence, capture the correlation between each element and other elements in the feature sequence, and obtain the calculation result of the multi-head attention;
[0164] Perform the first residual connection and the first layer normalization processing on the calculation result output by the multi-head attention;
[0165] Use a feed-forward neural network to perform non-linear feature transformation on the calculation result after the first residual connection and the first layer normalization processing;
[0166] Perform the second residual connection and the second layer normalization processing on the calculation result after the non-linear feature transformation to obtain 3D image features.
[0167] In one embodiment, using the multi-head attention in the encoding layer to calculate the global dependencies in the feature sequence, capture the correlation between each element and other elements in the feature sequence, and obtain the calculation result of the multi-head attention specifically includes:
[0168] Use the query weight matrix, key weight matrix, and value weight matrix in the multi-head attention mechanism of the encoding layer to perform linear transformation on each element in the feature sequence respectively, and obtain the query vector, key vector, and value vector corresponding to each element, including:
[0169] Q i = Xi W Q , K i = X i W K , V i = X i W V ;
[0170] Among them, X i represents the i-th element in the said feature sequence; W Q represents the query weight matrix; Q i represents the i-th query vector; W K represents the key weight matrix; K i represents the i-th key vector; W V represents the value weight matrix; V i represents the i-th value vector;
[0171] Calculate the attention scores between each query vector Q i and each key vector K j specifically including:
[0172]
[0173] Among them, Q i represents the i-th query vector; K j represents the j-th key vector; represents the transpose matrix of the j-th key vector; d k represents the dimension of the key vector;
[0174] According to the attention scores between each query vector Q i and each key vector K j calculate the corresponding attention weights, including:
[0175]
[0176] Use the attention weights W(Q i , K j ) to calculate the weighted value vectors corresponding to each query vector Q i specifically including:
[0177]
[0178] Concatenate all the weighted value vectors to obtain a weighted value vector sequence; adjust the dimension of the weighted value vector sequence so that the dimension of the weighted value vector sequence is the same as the dimension of the feature sequence to obtain the calculation result of the multi-head attention, specifically including:
[0179] MultiheadOutput = Concat(Head1 , Head 2 , …, Head h )W O ;
[0180] Among them, W O represents the linear transformation weight matrix, which is used to adjust the dimension of the weighted value vector sequence after splicing; Concat represents the splicing operation.
[0181] In one embodiment, the first residual connection and the first layer normalization processing on the calculation result of the multi-head attention output specifically include:
[0182] Output residual = LayerNorm[X + MultiheadOutput(X)];
[0183] Among them, LayerNorm represents performing layer normalization processing; X represents the feature sequence; MultiheadOutput(X) represents the calculation result of the multi-head attention mechanism when the input is X.
[0184] In one embodiment, the non-linear feature transformation of the calculation result after the first residual connection and the first layer normalization processing specifically includes:
[0185] FNN(Output residual ) = max(0, Output residual W 1 + b 1 )W 2 + b 2 ;
[0186] Among them, W 1 represents the weight matrix of the first layer of linear transformation in the feed-forward neural network; b 1 is the bias term of the first layer of linear transformation; max(·) represents the ReLU activation function, which is used to change the negative part of the result after the first linear transformation to 0 and retain the positive part; W 2 represents the weight matrix of the second layer of linear transformation in the feed-forward neural network; b 2 is the bias term of the second layer of linear transformation.
[0187] In one embodiment, the second residual connection and the second layer normalization processing on the calculation result after the non-linear feature transformation are performed to obtain the 3D image feature, specifically including:
[0188] Output FNN = LayerNorm[Output residual + FNN(Outputresidual )];
[0189] Among them, LayerNorm represents performing layer normalization processing; Output residual represents the calculation result of the multi-head attention output; FNN(Output residual ) represents the conversion result of the feed-forward neural network.
[0190] In one embodiment, the classification based on 3D image features to obtain the bladder cancer recurrence prediction probability specifically includes:
[0191] After obtaining the feature sequence and generating 3D image features through feature extraction by passing through multiple encoding layers in sequence, the second CLS token generated by the feature extraction of the first CLS token through multiple encoding layers;
[0192] Input the second CLS token into the classification head for classification to obtain a score vector:
[0193] Logits = CLS · W cls + b cls ;
[0194] Among them, CLS represents the second CLS token; W cls represents the weight matrix in the classification head; b cls represents the bias term in the classification head, used to adjust the mapping result; Logits represents the vector composed of the scores of all results;
[0195] According to the score vector, calculate the probability corresponding to each prediction result respectively:
[0196]
[0197] Among them, P(class i ) represents the probability of the i-th prediction result; Logit i represents the score of the i-th prediction result; represents the exponential of the score of the i-th prediction result; represents the sum of the exponentials of the scores of all prediction results.
[0198] Embodiment Five
[0199] In one embodiment, the present application further provides a model training method. The trained model is used to execute the bladder cancer recurrence prediction method in Embodiments 1 to 4. The model training method specifically includes:
[0200] Data preprocessing: First, perform data standardization processing on the MRI medical image data of bladder cancer recurrence.
[0201] Loss function: We used the Cross Entropy Loss function to handle classification tasks.
[0202] Optimizer: Usually, the AdamW optimizer is selected.
[0203] Learning rate scheduling: A learning rate warm-up scheduler may be used to optimize the model training process.
[0204] Data augmentation: During the training process, to improve the generalization ability of the model, data augmentation techniques such as random rotation, scaling, and slicing were applied.
[0205] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0206] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0207] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A bladder cancer probability prediction system, characterized in that: include: A segmentation unit, configured to segment the MRI bladder cancer follow-up image into a plurality of 3D image blocks in response to an input of the MRI bladder cancer follow-up image; A processing unit, configured to flatten each 3D image block, and project the flattened 3D image block to a predetermined dimension to obtain a corresponding 3D image embedding; The extraction unit is used to add position encoding to each 3D image embedding, and then input all 3D image embeddings into the encoder of the Transformer model for feature extraction to obtain 3D image features; The classification unit is used to classify according to the 3D image features to obtain the predicted probability of bladder cancer recurrence.
2. A bladder cancer probability prediction system according to claim 1, characterized in that: The dividing the MRI image into a plurality of 3D image blocks specifically includes: dividing the MRI image into a plurality of non-overlapping 3D blocks; The step of flattening each 3D image block and projecting the flattened 3D image block to a predetermined dimension to obtain a corresponding 3D image embedding specifically includes: Flatten each 3D image block into a one-dimensional vector; The flattened 3D image block is projected to a predetermined dimension using a fully connected layer to obtain the corresponding 3D image embedding. The embedding and adding position codes to each 3D image respectively specifically includes: Obtaining position information of the 3D image block in the original space; Based on the position information, respectively embedding and adding position codes to the 3D images corresponding to the 3D image blocks; All 3D images are embedded and input into the encoder of the Transformer model for feature extraction to obtain 3D image features, specifically including: The preset first CLS token and all 3D image embeddings are combined into a feature sequence; The feature sequence is controlled to sequentially pass through multiple coding layers to perform feature extraction to obtain 3D image features.
3. A bladder cancer probability prediction system according to claim 2, characterized in that: The controlling the feature sequence to sequentially pass through multiple coding layers to extract features to obtain 3D image features specifically includes: The multi-head attention in the encoding layer is used to calculate the global dependency in the feature sequence, capture the correlation between each element in the feature sequence and other elements, and obtain the calculation result of the multi-head attention; Performing a first residual connection and a first layer normalization process on the calculation results of the multi-head attention output; A feedforward neural network is used to perform nonlinear feature transformation on the calculation results after the first residual connection and the first layer normalization processing; The calculation results after nonlinear feature conversion are subjected to a second residual connection and a second layer normalization process to obtain 3D image features.
4. A bladder cancer probability prediction system according to claim 3, characterized in that: The multi-head attention in the encoding layer is used to calculate the global dependency in the feature sequence, capture the correlation between each element in the feature sequence and other elements, and obtain the calculation result of the multi-head attention, specifically including: Using the query weight matrix, key weight matrix and value weight matrix in the multi-head attention mechanism in the encoding layer, each element in the feature sequence is linearly transformed to obtain the query vector, key vector and value vector corresponding to each element, including: Q i =X i W Q ,K i =X i W K ,V i =X i W V ; Among them, X i represents the i-th element in the feature sequence; W Q represents the query weight matrix; Q i represents the i-th query vector; W K represents the key weight matrix; K i represents the i-th key vector; W V Represents the value weight matrix; V i represents the i-th value vector; Calculate each query vector Q separately i And each key vector K j The attention scores between include: Among them, Q i represents the i-th query vector; K j represents the j-th key vector; represents the transposed matrix of the jth key vector; d k represents the dimension of the key vector; According to each query vector Q i And each key vector K j The attention scores between , calculate the corresponding attention weights, including: Using the attention weight W(Q i ,K j ) are calculated separately for each query vector Q i The corresponding weighted value vector includes: All weighted value vectors are concatenated to obtain a weighted value vector sequence; the dimension of the weighted value vector sequence is adjusted so that the dimension of the weighted value vector sequence is the same as the dimension of the feature sequence, and the calculation result of the multi-head attention is obtained, which specifically includes: Multih=(Head1,ead2,…,ead h )W O ; Among them, W O Represents a linear transformation weight matrix, which is used to adjust the dimension of the concatenated weighted value vector sequence; Concat represents a concatenation operation.
5. A bladder cancer probability prediction system according to claim 4, characterized in that: The first residual connection and the first layer normalization processing are performed on the calculation results of the multi-head attention output, specifically including: Output residual =LayerNorm[X+Multih(X)]; Among them, LayerNorm means layer normalization processing; X represents the feature sequence; Multih(M) represents the calculation result of the multi-head attention mechanism when the input is X.
6. A bladder cancer probability prediction system according to claim 5, characterized in that: The method of using a feedforward neural network to perform nonlinear feature conversion on the calculation results after the first residual connection and the first layer normalization processing specifically includes: FNN(Output residual )=ax(0,Output residual W1+1)W2+2; Wherein, W1 represents the weight matrix of the first layer linear transformation in the feedforward neural network; b1 is the bias term of the first layer linear transformation; max(·) represents the ReLU activation function, which is used to convert the negative part of the result after the first linear transformation to 0 and retain the positive part; W2 represents the weight matrix of the second layer linear transformation in the feedforward neural network; b2 is the bias term of the second layer linear transformation.
7. A bladder cancer probability prediction system according to claim 6, characterized in that: The calculation result after the nonlinear feature conversion is subjected to a second residual connection and a second layer normalization process to obtain the 3D image feature, specifically including: Output FNN =LayerNorm[Output residual +NN(Output residual )]; Among them, LayerNorm means layer normalization processing; Output residual Indicates the calculation result of multi-head attention output; FNN (Output residual ) represents the conversion result of the feedforward neural network.
8. A bladder cancer probability prediction system according to claim 2, characterized in that: The classification according to the 3D image features to obtain the predicted probability of bladder cancer recurrence specifically includes: After the feature sequence is sequentially subjected to multiple coding layers for feature extraction to generate 3D image features, a second CLS token is generated by extracting the first CLS token through the multiple coding layers; Input the second CLS token into the classification header for classification and obtain the score vector: Logits=CLS·W cls + cls ; Wherein, CLS represents the second CLS token; W cls represents the weight matrix in the classification head; b cls Represents the bias term in the classification head, which is used to adjust the mapping result; Logits represents the vector composed of the scores of all results; According to the score vector, the probability corresponding to each prediction result is calculated respectively: Among them, P(class i ) represents the probability of the i-th prediction result; Logit i Represents the score of the i-th prediction result; The e index represents the score of the i-th prediction result; Represents the sum of the e-index scores of all prediction results.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a method for predicting the probability of bladder cancer is implemented, the method comprising: In response to input of an MRI bladder cancer follow-up image, the MRI bladder cancer follow-up image is segmented into a plurality of 3D image blocks; Flatten each 3D image block, and project the flattened 3D image block to a predetermined dimension to obtain a corresponding 3D image embedding; After adding position encoding to each 3D image embedding, all 3D image embeddings are input into the encoder of the Transformer model for feature extraction to obtain 3D image features; Classification is performed based on 3D image features to obtain the predicted probability of bladder cancer recurrence.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, a method for predicting the probability of bladder cancer is implemented, the method comprising: In response to input of an MRI bladder cancer follow-up image, the MRI bladder cancer follow-up image is segmented into a plurality of 3D image blocks; Flatten each 3D image block, and project the flattened 3D image block to a predetermined dimension to obtain a corresponding 3D image embedding; After adding position encoding to each 3D image embedding, all 3D image embeddings are input into the encoder of the Transformer model for feature extraction to obtain 3D image features; Classification is performed based on 3D image features to obtain the predicted probability of bladder cancer recurrence.