A Multi-Expert-Based Tumor Staging and Classification Method for Prostate Cancer

By strengthening the learning of key region features through a multi-expert mechanism and a masking scoring network, the problems of low resolution and individual differences in ultrasound video diagnosis are solved, achieving higher diagnostic accuracy and robustness.

CN120823998BActive Publication Date: 2025-11-14TONGJI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511332412.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-11-14
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing ultrasound videos have low resolution and limited diagnostic accuracy in prostate cancer diagnosis, especially when dealing with small lesions or areas with unclear boundaries. Furthermore, existing models do not fully consider individual differences, diagnostic styles, and data quality fluctuations among multiple doctors' sampling, leading to unstable validation results.

Method used

A multi-expert mechanism is adopted, combining three medical expert models: MedSAM2, Deeplabv3, and nnUNet2++. The key region feature learning is enhanced by a masking scoring network, a simulated doctor consultation mechanism is constructed, and feature reconstruction and classification are performed using a masking scoring network and a visual converter encoder.

Benefits of technology

It significantly improves the diagnostic accuracy and robustness of the model in complex lesion areas, adapts to data from different doctors, and enhances the performance and practicality of ultrasound video diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823998B_ABST
    Figure CN120823998B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of representation learning, specifically involving a prostate cancer tumor staging classification method based on a multi-expert mechanism. It includes the following steps: Step 0: Constructing a prostate cancer tumor staging classification network based on a multi-expert mechanism; Step 1: Data preprocessing; Step 2: Masking score generation; Step 3: Pre-training the encoder; Step 4: Fine-tuning the encoder and classification head; Step 5: Using the trained network to predict the prostate ultrasound video staging. This method has the following advantages: (1) Targeted reconstruction during pre-training: The masking reconstruction pre-training strategy proposed in this invention, guided by prior knowledge of the medical segmentation model, has good adaptability to different data operated by different doctors. (2) Targeted additional prompts during downstream fine-tuning: The proposed method, based on adding important information to the masking score network trained during pre-training as additional prompts, also shows good adaptability when performing downstream classification on data operated by different doctors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of representation learning, specifically relating to a prostate cancer tumor staging and classification method based on a multi-expert mechanism. Background Technology

[0002] Prostate cancer is the second most common type of cancer among men worldwide and a leading cause of cancer-related deaths. As a crucial tool in modern medical diagnosis, medical imaging technology is widely used in disease detection and plays an irreplaceable role. Medical imaging technologies include magnetic resonance imaging (MRI), computed tomography (CT), X-ray imaging, and ultrasound imaging. These technologies provide valuable visual data to help doctors diagnose and assess diseases. In the diagnosis of prostate cancer, MRI is the most commonly used imaging tool. MRI has high specificity and can be used to predict the T stage of prostate cancer; however, its sensitivity is relatively low, and its widespread clinical application is limited by factors such as cost, contraindications, and equipment availability. Ultrasound imaging, another common diagnostic imaging tool, is also used in prostate cancer detection. Compared to MRI, ultrasound imaging has advantages such as lower cost, wider applicability, and less harm to the human body, making it more suitable for low-cost and efficient prostate cancer detection. However, the multifocal and scattered nature of prostate cancer reduces the diagnostic value of a single image. Ultrasound videos encompassing the entire prostate gland and surrounding tissues and organs compensate for the limitations of single-image diagnosis. However, ultrasound video imaging suffers from low resolution, making accurate diagnosis difficult. Furthermore, the sensitivity and specificity of ultrasound fluctuate significantly, especially in cases of small lesions and indistinct borders, where diagnostic accuracy relies heavily on the physician's experience. Therefore, effectively improving the resolution of ultrasound videos to provide clearer and higher-fidelity images, and enhancing the diagnostic accuracy for small lesions and complex areas with indistinct borders, has become a crucial research direction and key technological development area in prostate cancer diagnosis based on ultrasound imaging.

[0003] In the diagnosis of prostate cancer, ultrasound video, as an important imaging tool, offers advantages such as low cost and wide applicability. However, its low resolution and limited diagnostic accuracy, especially when dealing with small lesions or areas with indistinct borders, still present many challenges. Therefore, improving the diagnostic capabilities of ultrasound video, particularly in handling complex lesion areas, has become one of the core issues in the field of prostate cancer ultrasound diagnosis. Video classification has emerged as a crucial means to enhance the diagnostic effectiveness of ultrasound video in this context. Unlike traditional single static images, ultrasound video provides temporal features through continuous frame information, which is significant for capturing the dynamic changes of prostate cancer, lesion development, and its relationship with surrounding tissues. By analyzing multiple frames in ultrasound video, a more comprehensive assessment of the prostate tissue structure, blood flow, and tumor morphology can be achieved, thereby improving the detection rate and diagnostic accuracy.

[0004] MedSAM2, Deeplabv3, and nnUNet2++ are three open-source segmentation models. MedSAM2 is a medical image segmentation model based on the classic Segment Anything Model. By introducing spatial and semantic cue encoders, it significantly improves the segmentation accuracy and generalization ability of complex medical images. Deeplabv3 is a representative work of the DeepLab series. It uses dilated convolutions and ASPP modules to effectively capture multi-scale contextual information and is a classic architecture in the field of general semantic segmentation. nnUNet2++ is a powerful upgrade of nnUNet, continuing the "no major modifications" automated design philosophy. Through adaptive preprocessing and network optimization, it exhibits strong robustness in many medical segmentation tasks. All three models have some powerful pre-trained weights that can provide some lesion segmentation knowledge, thus they can be used as expert models for prostate cancer tumor staging and classification. However, using a single expert model may lead to misjudgments due to individual differences between doctors and patients, diagnostic styles, and data quality fluctuations.

[0005] In existing methods, the VideoMAE work employs a self-supervised pre-training method based on Masked Autoencoder (MAE). The video-based MAE pre-training process is as follows: Figure 2 As shown.

[0006] Using downsampled frames as input, a cross-line temporal sampling strategy is employed for more efficient video pre-training. Formally, a video clip consisting of t consecutive frames is first randomly sampled from the original video V. Then, temporal sampling is used to compress the clip into T frames, each with a resolution of [resolution missing]. And use cube embedding to obtain video tokens. Here we will use each with a size of... The cube is considered a token embedding. Therefore, the cube embedding layer obtains... Each token is assigned a 3D token and mapped to a channel dimension d. This design reduces the spatiotemporal dimension of the input, helping to alleviate spatiotemporal redundancy in the video.

[0007] For the obtained 3D tokens, VideoMAE randomly selects 90% of the information for masking. The unmasked tokens are then fed into the Vision Encoder, which uses a Vision Transformer (ViT) as its backbone. Finally, the decoder is used for reconstruction training to optimize the reconstruction loss. .

[0008] Subsequently, for the fine-tuning process, VideoMAE uses a pre-trained encoder to input the 3D token into the visual encoder of the previously pre-trained Vision Transformer (ViT) to obtain features. The deepest dimension features are taken as the input of the obtained features into the classification head, and the classification head and visual encoder are trained through classification loss.

[0009] However, prostate ultrasound data contains a large amount of irrelevant information, especially in lesion-free areas. Random masking can cause the model to ignore key lesion areas during learning, which is particularly important in medical imaging because local features often determine the early detection and accurate diagnosis of diseases. When MAE masks too many non-critical local features, the model's ability to reconstruct truly critical local features is relatively poor. This is the most significant drawback.

[0010] Existing models do not fully consider individual differences, diagnostic styles, and data quality fluctuations among multiple doctors' sampling, resulting in unstable validation results and failing to meet the high requirements for accuracy and robustness in medical-engineering integration applications. Summary of the Invention

[0011] To address the problems of existing technologies, this invention proposes a prostate cancer tumor staging and classification method based on a multi-expert mechanism. To address the issue that random occlusion during reconstruction pre-training in the MAE (Multi-Expert Analytical Process) workflow fails to enable the model to learn features of key regions, a occlusion scoring network based on lesion segmentation content is designed. This forces the model to learn features of key regions, i.e., crucial local features. In building the expert model, a multi-expert consultation mechanism simulating physician consultation is constructed, utilizing three expert models to reduce the probability of errors.

[0012] During downstream fine-tuning, a masking scoring network pre-trained during the reconstruction training phase is used to add blocks of important regions from the input video as additional cues to the feature-assisted classification head training. This addresses the model performance issues caused by individual differences among multiple doctor samples, diagnostic styles, and data quality fluctuations.

[0013] Technical solution

[0014] A multi-expert-based prostate cancer tumor staging and classification method includes the following steps:

[0015] Step 0: Construct a prostate cancer tumor staging and classification network based on a multi-expert mechanism, including: three pre-trained medical expert models, a masking scoring network, a visual converter encoder, and a classification head;

[0016] The three medical expert models are MedSAM2, Deeplabv3, and nnUNet2++ segmentation model. All three models have powerful pre-trained weights and can provide segmentation knowledge of lesions, so they can be used as medical expert models for prostate cancer tumor staging and classification methods.

[0017] There are three occlusion scoring networks, each corresponding to a medical expert model; each occlusion scoring network includes: a multi-head attention block (MHA) and a fully connected layer (FC);

[0018] The vision transducer encoder is the Vision Transformer-base (ViT-B) version, which includes a vision encoder and a decoder;

[0019] The classification head consists of a fully connected layer. The features extracted by the visual encoder are provided to the classification head to perform prostate cancer tumor staging and classification, thus completing the detection task.

[0020] Step 1: Data preprocessing;

[0021] The medical ultrasound video is divided into multiple clips of equal length with a fixed length. Each frame in the clip is then processed by three medical expert models to obtain the segmentation result (i.e., the segmented image).

[0022] The segmented image contains only black and white pixels, where the white part represents the lesion area and the black part represents the background.

[0023] Step 2: Obtain the 3D token of the clip by block embedding, and then use the masking scoring network to calculate its masking score;

[0024] The clip obtained in step 1 first obtains multiple 3D tokens through patch embedding;

[0025] The 3D tokens are input into three occlusion scoring networks to obtain the occlusion score for each token.

[0026] The occlusion scores of the three occlusion scoring networks are corrected using the three segmented images obtained in step 1: if the token is located within the segmented image (i.e., within the white lesion area), its occlusion score is increased by a small value for correction.

[0027] The three-way masking scores are input into the gating network. The weights output by the gating network and the three-way masking scores are weighted and averaged to obtain the final token score.

[0028] The subsequent steps involve masking the core lesion area of ​​the image, allowing the model to learn the representation of that area during the reconstruction process.

[0029] Step 3: Pre-train the visual encoder and the occlusion scoring network;

[0030] Based on the occlusion scores, the top 90% of the tokens are occluded and then input into the visual encoder to obtain encoded tokens.

[0031] During the pre-training phase, the encoded 10% of tokens are merged with the occluded 90% of tokens, and the result is input into the decoder to reconstruct the original image; the reconstruction loss is then calculated. Used to update the visual encoder and calculate the sampling loss. , used to update the masking score network.

[0032] Step 4: Further fine-tune the visual encoder and classification head using the labeled data;

[0033] After the visual encoder and occlusion scoring network are pre-trained, the occlusion scoring network is frozen, and the visual encoder and classification head in the model are fine-tuned using prostate ultrasound images.

[0034] In the fine-tuning stage, prostate ultrasound videos with tumor staging classification labels are selected. The video clips are processed through a masking scoring network to obtain the score of each token. The top 10% of tokens are added as additional prompts to the token sequence. After being input into a pre-trained visual encoder, the obtained features are input into the classification head for classification. The classification head and encoder are updated using cross-entropy loss.

[0035] Step 5: After completing the training in Steps 3 and 4, use a prostate cancer tumor staging classification network based on a multi-expert mechanism to predict the prostate ultrasound video staging.

[0036] Beneficial effects

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] (1) Targeted Reconstruction in the Pre-training Stage: This invention employs a pre-training strategy based on prior knowledge of the medical segmentation model to guide masking reconstruction, demonstrating good adaptability to different data from different doctors. Furthermore, the pre-training process can utilize large-scale unlabeled data, exhibiting scalability. This approach not only significantly improves model performance but, more importantly, enhances the model's robustness and practicality when facing complex, multi-source medical data in the real world.

[0039] (2) Targeted Additional Hints in Downstream Fine-tuning Stage: This invention, based on a masking scoring network, adds important information as additional hints, demonstrating good adaptability when classifying data from different doctors' operations downstream. This improves the performance of downstream classification tasks. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of the method framework of the present invention;

[0041] Figure 2 This is a diagram illustrating the application of an automatic masking encoder in video (VideoMAE).

[0042] Figure 3 These are visualizations of three medical segmentation models: (A) Segmentation result of MedSAM2 model; (B) Segmentation result of Deeplabv3 model; (C) Segmentation result of nnUNet2++ model; (D) Original ultrasound image.

[0043] Figure 4 This is a flowchart of the masking scoring network framework of the present invention;

[0044] Figure 5 This is a flowchart of step 3, pre-training parameter update, of the present invention;

[0045] Figure 6 This is a flowchart of step 4, fine-tuning parameter update, of the present invention;

[0046] Figure 7 This is a flowchart of the prostate ultrasound video staging prediction process using the network trained according to this invention. Detailed Implementation

[0047] The technical solution provided in this application will be further described below with reference to specific embodiments and accompanying drawings. The advantages and features of this application will become clearer from the following description.

[0048] To address the issue that random occlusion during reconstruction pre-training in the MAE process fails to enable the model to learn key region features, this invention designs an occlusion scoring network based on lesion segmentation content. This forces the model to learn the features of key regions, i.e., crucial local features.

[0049] In building expert models, due to the limitations of model generalization, their segmentation results on unseen datasets may contain errors for a subset of instances. Since different segmentation models encounter inconsistent data, different instances of segmentation errors may exist. However, the probability of three models simultaneously segmenting the same instance incorrectly is low. Visualization is as follows... Figure 3 As shown, the three medical expert models, MedSAM2, Deeplabv3, and nnUNet2++, are used to analyze the same image. Figure 3 D) gives different segmentation results ( Figure 3 (AC) segmentation results include information on both the lesion area and normal tissue. This invention constructs a multi-expert consultation mechanism that simulates a doctor's consultation, utilizing three expert models to make joint judgments to reduce the probability of errors.

[0050] like Figure 1 This is a schematic diagram of the method framework of the present invention (including steps 1 to 4).

[0051] A multi-expert-based prostate cancer tumor staging and classification method includes the following steps:

[0052] Step 0: Construct a prostate cancer tumor staging and classification network based on a multi-expert mechanism, including: three pre-trained medical expert models, a masking scoring network, a visual converter encoder, and a classification head;

[0053] The three medical expert models are MedSAM2, Deeplabv3, and nnUNet2++ segmentation model. All three models have powerful pre-trained weights and can provide segmentation knowledge of lesions, so they can be used as medical expert models for prostate cancer tumor staging and classification methods.

[0054] There are three occlusion scoring networks, each corresponding to a medical expert model; each occlusion scoring network includes: a multi-head attention block (MHA) and a fully connected layer (FC);

[0055] The vision transducer encoder is the Vision Transformer-base (ViT-B) version, which includes a vision encoder and a decoder;

[0056] The features extracted by the visual encoder are provided to the classification head to classify and stage prostate cancer tumors, thus completing the detection task.

[0057] Step 1: Data Preprocessing

[0058] The medical ultrasound video was divided into multiple equal-length clips, and multiple medical expert models were used for medical image segmentation.

[0059] The specific process is as follows:

[0060] The video with T frames is divided into multiple equal-length clips, each clip having a length of [missing information]. If the value is 16, then the number of clips is step= Each clip is sampled continuously, resulting in insufficient data at the end of the video. The remaining frames are discarded. This results in multiple equal-length clips:

[0061] (1)

[0062] The edited segments were processed by pre-trained segmentation models (medical expert models) MedSAM2, Deeplabv3, and nnUNet2++ to obtain the corresponding segmented content:

[0063] (2)

[0064] in This is an index for three different pre-trained segmentation models.

[0065] It should be noted that the segmentation models MedSAM2, Deeplabv3, and nnUNet2++ are pre-trained models and do not require additional doctor-annotated data for training and inference.

[0066] Step 2: Masking Score Generation ( Figure 4 )

[0067] The clip first employs a joint spatiotemporal cubic embedding to obtain multiple 3D tokens. Specifically, for a clip of size... Where H and W represent the height and width of the image, sampled 2 units in time and 2 units in space. The unit yielded N=( ) of size The tokens are mapped using a 3D convolution, resulting in... Dimensional tokens.

[0068]

[0069] The above methods can alleviate spatiotemporal redundancy.

[0070] Construct a masking scoring network, the network architecture of which is as follows: Figure 4 As shown. The masking scoring network consists of 3 independent multi-head attention blocks and a fully connected layer. The token ( The input occlusion scoring network is calculated as follows:

[0071] (3)

[0072] (4)

[0073] Here are the masking scores for all 3D tokens from the three masking scoring networks:

[0074] (5)

[0075] Segmentation content obtained using the three medical expert models in step 1 The occlusion score is then adjusted as follows:

[0076] (6)

[0077] in, The segmentation result of step 1, The corresponding numbers are for the three medical expert models;

[0078] The InRoI function checks if the token's location is within the segmented image; if so, it returns a small value. To increase the token's score, otherwise return 0. Details are as follows:

[0079] (7)

[0080] In this embodiment, The value is 2.

[0081] Ultimately, three independent camouflage scores were obtained.

[0082] (8)

[0083] The occlusion score is input into the gating network, as shown in the following formula:

[0084] (9)

[0085] Then and occlusion score vector ( Multiply the results and then add them together to get the final token score. .

[0086]

[0087] Where Sum represents a... Tensors of dimension added to Dimension.

[0088] Step 3: Pre-train the visual converter encoder and occlusion scoring network

[0089] The pre-training phase training and parameter update process is as follows: Figure 5 As shown.

[0090] Score for concealment The first 90% of the tokens are masked, and the remaining 10% of visible tokens are input into the Vision Transformers Encoder for encoding. The masked 90% of tokens are then added to the encoded tokens, and finally input into the Vision Decoder for reconstruction, yielding the reconstructed result. .

[0091] Calculate reconstruction loss This is used to update the weights of the encoder and decoder. The loss formula is as follows:

[0092] (10)

[0093] Where N represents the total number of tokens, and These represent the RGB (Red, Green, Blue) values ​​of the i-th token and the reconstruction result, respectively.

[0094] Calculate sampling loss The weights are used to update the masking score network. The loss formula is as follows:

[0095] (11)

[0096] In an image, different regions contain different amounts of information. For example, some regions of an image may contain more structured information (such as the outlines of objects, color variations, etc.), while other regions may be background areas containing less information.

[0097] In masked autoencoder (MAE) algorithms, high-information regions (such as objects and edges) are generally more critical for reconstruction than low-information background regions, and their reconstruction errors are typically higher. Therefore, maximizing the expected reconstruction error will result in a higher probability score for the network predicting high-information regions.

[0098] Step 4: Fine-tune the encoder and sorting head;

[0099] The fine-tuning phase training and parameter update process is as follows: Figure 6 As shown.

[0100] like Figure 1 The output of the visual encoder is provided to the classification head to perform classification and complete the prostate tumor staging detection.

[0101] The classification head consists of a fully connected layer.

[0102] After pre-training the visual encoder and occlusion scoring network in step 3, freeze the occlusion scoring network. This step only involves fine-tuning the visual encoder and classification head.

[0103] The specific process is as follows:

[0104] For a video with T0 / T1 class tags, split it into multiple Clips using the same splitting method as in step 1. The size of each Clip is [size missing]. ;

[0105] For each Clip, a masking score is generated through a masking score network, calculated in the same way as in step 2.

[0106] Then the tokens of the top 10% of the score Added as an extra hint after the embedded token:

[0107]

[0108] Inputting a pre-trained Vision Transformers Encoder into the input, the features are obtained after encoding. The obtained features are averaged across the feature dimensions to obtain the global features. :

[0109]

[0110] global features Input the classification header to obtain the final classification result:

[0111]

[0112] In classification tasks, the loss is cross-entropy loss. The formula is as follows:

[0113]

[0114] Where y represents the true label of the sample, which is 0 (T0) or 1 (T1). This represents the probability that the model's final output matches the true label of the sample. The parameters of the encoder and classifier head are updated using cross-entropy loss (existing technology).

[0115] Step 5: Use a prostate cancer tumor staging classification network based on a multi-expert mechanism to predict the prostate ultrasound video staging (e.g., Figure 7 );

[0116] After network training is completed, the prostate ultrasound video is divided into regions for prediction. The process is as follows:

[0117] The prostate ultrasound video was split into multiple clips, each clip being [size missing]. The splitting method is the same as step 1.

[0118] For each Clip, a masking score is generated using a masking scoring network; the calculation method is the same as in step 2.

[0119] Then, the model's classification result for the clip is obtained through a visual encoder and a classification head: The prediction method is the same as in step 4.

[0120] After obtaining the classification result for each clip, if one clip is classified as class T1, then the entire video is classified as class T1; otherwise, if the model classifies all clips as class T0, then the entire video is classified as class T0.

[0121] Example

[0122] Method Application

[0123] All experiments in this invention were performed on two NVIDIA RTX 3090 GPUs. The specific values ​​of the hyperparameters for this method are shown in Table 1.

[0124] Table 1. Value settings for various hyperparameters in this invention

[0125]

[0126] Dataset:

[0127] The experiment was conducted on the dataset from Shanghai Tenth People's Hospital for both training and validation. The data statistics are shown in Table 2 below. Table 2 presents the statistics of the training and validation sets in this embodiment:

[0128] Table 2. Statistical information of the dataset in this embodiment.

[0129]

[0130] Method application:

[0131] In this embodiment, the training batch size is set to 4, and training is performed on two RTX 3090 GPUs throughout the process. For our work, we use the AdamW optimizer, pre-training for 500 epochs at a learning rate of 1e-3 to minimize reconstruction and sampling losses. We then fine-tune for 300 epochs at a learning rate of 4e-4 to minimize the cross-entropy loss.

[0132] Effect Comparison

[0133] Comparison 1

[0134] The ablation experiment results for the mask are shown in Table 2. This embodiment compares the classification performance of the strategy of using random masks in the pre-trained visual encoder, i.e., directly and randomly masking the video, with the classification performance of the prior segmentation guided mask based on three medical segmentation models in this invention. Both also use the improved fine-tuning method of this invention: the evaluation metrics are ACC (Accuracy) and AUC (Area Under the Curve), which measure the accuracy and discriminative ability of the model in the prediction task, respectively.

[0135] As can be seen from Table 3, the multi-expert guided occlusion method of this invention has a stronger classification ability than the classification model fine-tuned using the same fine-tuning method by random occlusion.

[0136] Table 3 Ablation experiments for the mask

[0137]

[0138] Comparison 2

[0139] This embodiment conducted a comparative experiment on the fine-tuning method, comparing it with the fine-tuning method that directly inputs all tokens and extracts the deepest features of the visual encoder, and the fine-tuning strategy of this invention that uses the 10% most important additional tokens as input features of the visual encoder and extracts the deepest features. Both methods used the improved pre-training method of the prior segmentation guidance mask based on three medical segmentation models of this invention. As shown in Table 4, it can be seen that the fine-tuning method of the visual encoder of this invention is intuitively superior to the closest fine-tuning method on the market in terms of classification ACC and AUC metrics.

[0140] Table 4 Ablation experiments for the fine-tuning method

[0141]

[0142] Comparison 3

[0143] The method of this invention is compared with some existing typical methods, including ResNet50, US_UCL, RadFormer, PVTv2 (SOTA), VideoMAEv2, and m2clip (SOTA).

[0144] Among them, ResNet50, US_UCL, RadFormer, and PVTv2 (SOTA) are image-based methods. ResNet50 is a classic deep residual network that solves the gradient vanishing problem in deep models and is the backbone network for many visual tasks. US_UCL is a classification method based on ultrasound images, which uses uncertainty estimation and contrastive learning to improve model robustness. RadFormer is a Transformer model based on radiomics, which integrates local radiomics features with global contextual information. PVTv2 (SOTA), the Pyramid Vision Transformer v2, is an efficient visual backbone network that performs excellently on dense prediction tasks.

[0145] VideoMAEv2, m2clip (SOTA), and the method of this invention are video-based approaches. VideoMAEv2 is a video pre-training method based on a masked autoencoder, which learns strong representations by reconstructing a large number of masked video frames. m2clip (SOTA) is a model that aligns medical videos with reports, learning a shared multimodal embedding space that can be used for retrieval and classification.

[0146] Using a unified training and validation set partition, the results are shown in Table 5: It can be seen that the model of this invention has the best performance in related work based on images and videos.

[0147] Table 5 Comparison with other works

[0148]

[0149] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.

Claims

1. A prostate cancer tumor staging and classification method based on a multi-expert mechanism, characterized in that, Includes the following steps: Step 0: Construct a prostate cancer tumor staging and classification network based on a multi-expert mechanism, including: three pre-trained medical expert models, a masking scoring network, a visual converter encoder, and a classification head; The three medical expert models are MedSAM2, Deeplabv3, and nnUNet2++ segmentation model. There are three occlusion scoring networks, each corresponding to a medical expert model; each occlusion scoring network includes: a multi-head attention block (MHA) and a fully connected layer (FC); The encoder of the vision converter is the basic version of the vision converter, ViT-B, which includes a vision encoder and a decoder; The classification head consists of a fully connected layer. The features extracted by the visual encoder are provided to the classification head to perform prostate cancer tumor staging and classification, thus completing the detection task. Step 1: Data preprocessing; The medical ultrasound video is divided into multiple equal-length clips with a fixed length; each frame in the clip is processed by three medical expert models to obtain a segmented image of the frame. The segmented image contains only black and white pixels, where the white part represents the lesion area and the black part represents the background. Step 2: Obtain the 3D token of the clip by block embedding, and then use the masking scoring network to calculate its masking score; The clip obtained in step 1 is first used to obtain multiple 3D tokens through block embedding; The 3D tokens are input into three occlusion scoring networks to obtain the occlusion score for each token. The occlusion scores of the three occlusion scoring networks are corrected using the three segmented images obtained in step 1: if the token is located within the segmented image, i.e., within the white lesion area, its occlusion score is increased by a small value for correction. The three-way masking scores are input into the gating network. The weights output by the gating network and the three-way masking scores are weighted and averaged to obtain the final token score. Step 3: Pre-train the visual encoder and the occlusion scoring network; Based on the occlusion scores, the top 90% of the tokens are occluded and then input into the visual encoder to obtain encoded tokens. During the pre-training phase, the encoded 10% of tokens are merged with the occluded 90% of tokens, and the result is input into the decoder to reconstruct the original image; the reconstruction loss is then calculated. Used to update the visual encoder and calculate the sampling loss. , used to update the masking score network; Step 4: Further fine-tune the visual encoder and classification head using the labeled data; After the visual encoder and occlusion scoring network are pre-trained, the occlusion scoring network is frozen, and the visual encoder and classification head in the model are fine-tuned using prostate ultrasound images with tumor staging classification labels. In the fine-tuning stage, prostate ultrasound videos with tumor staging classification labels are selected. The video clips are processed through a masking scoring network to obtain the score of each token. The top 10% of tokens are added as additional cues to the token sequence and then input into a pre-trained visual encoder. The obtained features are then input into the classification head for classification. The classification head and encoder are updated using cross-entropy loss. Step 5: After completing the training in Steps 3 and 4, use a prostate cancer tumor staging classification network based on a multi-expert mechanism to predict the prostate ultrasound video staging.

2. The method for staging and classifying prostate cancer tumors based on a multi-expert mechanism according to claim 1, characterized in that, Step 1 specifically involves: dividing the video with T frames into multiple equal-length clips, with a clip length of... If the value is 16, then the number of clips is step = Each clip is sampled continuously, resulting in insufficient data at the end of the video. The remaining frames are discarded; ultimately, multiple clips of equal length are obtained: (1) The edited segments were processed by pre-trained segmentation models MedSAM2, Deeplabv3, and nnUNet2++ to obtain the corresponding segmented content: (2) in This is an index for three different pre-trained segmentation models.

3. The method for staging and classifying prostate cancer tumors based on a multi-expert mechanism according to claim 1, characterized in that, In step 2, the process of obtaining multiple 3D tokens through block embedding is as follows: For a clip, the Clip size is Where H and W represent the height and width of the image, sampled 2 units in time and 2 units in space. The unit obtained N=( ) of size The token; Each token is mapped using a 3D convolution, resulting in... Dimensional tokens: 。 4. The method for staging and classifying prostate cancer tumors based on a multi-expert mechanism according to claim 1, characterized in that, In step 2, the masking scoring network processes the process as follows: The masking scoring network consists of three independent multi-head attention blocks and a fully connected layer; it will store tokens ( The input occlusion scoring network is calculated as follows: (3) (4) For the masking scores of the three masking scoring networks for all 3D tokens (5) Segmentation content obtained using the three medical expert models in step 1 The occlusion score is adjusted as follows: (6) in, The segmentation result of step 1, The corresponding numbers are for the three medical expert models; The InRoI function checks if the token's location is within the segmented image; if so, it returns a small value. To increase the token's score, otherwise return 0, as follows: (7) Ultimately, three independent camouflage scores were obtained. (8) The occlusion score is input into the gating network, as shown in the following formula: (9) Then The token score is obtained by multiplying it by the camouflage score vector and then summing the results. ; Where Sum represents a... Tensors of dimension summed to Dimension.

5. The method for staging and classifying prostate cancer tumors based on a multi-expert mechanism according to claim 1, characterized in that, In step 3, the occlusion score is calculated. The first 90% of the tokens are masked, and the remaining 10% of visible tokens are input into the visual converter encoder for encoding. After encoding, the previously masked 90% of tokens are added to the encoded tokens, and then input into the visual decoder for reconstruction to obtain the reconstruction result. ; Calculate reconstruction loss The weights are used to update the encoder and decoder, and the loss formula is as follows: (10) Where N represents the total number of tokens, and These represent the RGB value and reconstruction result of the i-th token, respectively; Calculate sampling loss The weights used to update the masking score network are calculated using the following loss formula: (11)。 6. The method for staging and classifying prostate cancer tumors based on a multi-expert mechanism according to claim 1, characterized in that, Step 4 is as follows: Input a video with T0 / T1 class tags, and split it into multiple Clips, each Clip being [size missing]. ; For each Clip, a masking score is generated using a masking scoring network; Then the tokens of the top 10% of the score Added as an extra hint after the embedded token: Inputting a pre-trained visual converter encoder, the features are obtained after encoding. The obtained features are averaged across the feature dimensions to obtain the global features. : global features Input the classification header to obtain the final classification result: In classification tasks, the loss is cross-entropy loss. The formula is as follows: Where y is the true label of the sample, which is either T0 or T1. It represents the probability that the model's final output matches the true label of the sample; the parameters of the encoder and classifier head are updated using cross-entropy loss.

7. The method for staging and classifying prostate cancer tumors based on a multi-expert mechanism according to claim 1, characterized in that, Step 5 is as follows: The prostate ultrasound video was split into multiple clips, each clip being [size missing]. ; For each Clip, a masking score is generated using a masking scoring network; Then, the model's classification result for the clip is obtained through the encoder and classification head: ; After obtaining the classification result for each clip, if one clip is classified as class T1, then the classification result for the entire video is class T1; conversely, if the model classifies all clips as class T0, then the classification result for the entire video is class T0.

Citation Information

Patent Citations

  • Multi-type labeling process integration system for medical images

    CN117174261A

  • Medical image classification method based on prior knowledge enhanced mask and alignment modeling

    CN119399523A