Tumor prediction method and system for fusing multiple views through spatial perception mask prompt
A tumor prediction method that uses spatially aware mask cues to fuse multiple views, combined with a self-generated mask cues to fuse spatially aware coding modules, solves the problems of easy loss of small object features and high training complexity in three-dimensional ultrasound image processing, achieves efficient and accurate tumor detection and malignancy risk assessment, and improves the diagnostic performance and adaptability of the model.
Patent Information
- Application Number
- CN202510838929.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
Existing three-dimensional ultrasound image processing methods have problems in tumor lesion detection, such as easy loss of small object features, high training complexity, and strong dependence on large-scale data sets. These problems limit the model performance improvement and generalization ability, resulting in insufficient diagnostic efficiency and accuracy.
A tumor prediction method that uses spatially aware mask cues and fused multiple views is adopted. Through the combination of a segmentation network model, a fine-tuning module, a position information encoding module, a classification network model, and a prediction network model, ABUS/ABVS instruments are used to obtain three-dimensional ultrasound breast images for tumor classification and malignant risk prediction. The spatial perception and diagnostic accuracy of the model are improved by combining self-generated mask cues with the spatially aware encoding module.
It significantly improves the accuracy of benign and malignant tumor classification and the accuracy of benign tumor malignant risk assessment, simplifies the diagnostic process, reduces the workload of doctors, improves diagnostic efficiency and accuracy, and enhances the generalization ability of the model.
Smart Images

Figure CN120748716A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical ultrasound image processing, and in particular to a tumor prediction method and system for spatially perceptual mask-prompted fusion of multiple views. Background Art
[0002] Ultrasound examination is indispensable for breast assessment and is widely popular due to its lack of radiation, high comfort, and cost-effectiveness. However, traditional two-dimensional ultrasound relies heavily on the operator's experience and has a limited viewing angle, making it difficult to meet modern examination requirements. With technological advances, three-dimensional breast ultrasound (ABUS / ABVS) has become increasingly popular. It not only reduces reliance on operator proficiency but also provides detailed images in three dimensions, enhancing diagnostic accuracy and comprehensiveness. Therefore, three-dimensional ultrasound is becoming a vital tool for improving the detection of breast diseases.
[0003] Currently, mainstream algorithms in the field of 3D ultrasound image processing utilize convolutional neural networks (CNNs) and visual transformer (ViT) architectures. However, the limitations of these two approaches hinder the model's detection performance for small objects such as tumor lesions. Convolutional neural networks, in particular, tend to lose small features as network depth increases, and tumor lesions are typically small, which directly weakens the model's predictive capabilities. On the other hand, while the visual transformer can capture global information through a global attention mechanism, it has high computational complexity, performs poorly on small and medium-sized datasets, and is prone to overfitting, further complicating its application to benign and malignant classification tasks in other organs through transfer learning.
[0004] In summary, current 3D ultrasound image network models still have certain shortcomings in terms of flexibility, reliability, and efficiency. Current 3D ultrasound image processing methods are mainly based on convolutional neural networks and Vision Transformer architectures for training. Although they can complete specific tasks, they also expose numerous problems, such as the easy loss of small object features, high training complexity, and strong dependence on large-scale datasets. Therefore, these issues not only limit the improvement of model performance, but also significantly increase development costs and time. At the same time, they also reduce the model's ability to generalize to data from other organs or image types, thereby restricting its widespread application in the medical imaging field. Summary of the Invention
[0005] The present invention provides a tumor prediction method and system based on spatially aware mask-cued fusion of multiple views to solve the problems raised in the above background technology.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: A tumor prediction method based on spatially aware mask-cued fusion of multiple views, comprising the following steps: Step 1) obtaining a three-dimensional ultrasound breast image of the patient; Step 2) inputting the acquired 3D ultrasound breast image into a trained medical image analysis model; Step 3) Processing and analyzing the input 3D ultrasound breast image using a medical image analysis model to obtain a benign or malignant tumor classification result and a benign tumor malignant transformation risk prediction result in the 3D ultrasound breast image; Step 4) Present the benign tumor malignant transformation risk prediction results in structured text format to generate a risk assessment result for benign tumors.
[0007] Preferably, in step 1, the patient is examined using an ABUS / ABVS instrument to obtain a three-dimensional ultrasound breast image; The acquired three-dimensional ultrasound breast images include: anterior-posterior view, lateral view, and mediolateral oblique view of the patient's breast.
[0008] Preferably, in step 2, the medical image analysis model includes: A segmentation network model is used to segment breast-related tissue in each view of the 3D ultrasound breast image and generate a breast tissue mask image corresponding to each view of the 3D ultrasound breast image and save the feature map generated before the softmax layer. The breast tissue mask image includes: a tumor mask image, a dense tissue mask image, and a gland mask image; A fine-tuning module, used for medical personnel to manually perform fine-tuning operations on the breast tissue mask image generated by segmentation; A position information encoding module is used to perform three-dimensional spatial position encoding on the breast tissue mask image and the corresponding feature map after manual fine-tuning by medical personnel to generate a breast tissue mask image with spatial perception information; The classification network model is responsible for simultaneously receiving the breast tissue mask image with spatial perception information, the corresponding feature map, and the original 3D ultrasound breast image to extract the multi-scale spatial information of the breast tissue image. Based on the multi-scale spatial information of the breast tissue image, it classifies the tumor as benign or malignant in each view of the 3D ultrasound breast image. The prediction network model is used to perform weighted average fusion decision on the benign and malignant tumor classification results of each view in the three-dimensional ultrasound breast image, and is used to predict the risk of malignant transformation of benign tumors in the weighted average fusion decision in the next few years to generate a risk assessment result for benign tumors.
[0009] Preferably, the training strategy of the medical image analysis model is as follows: 1) Dataset construction; Collect a certain amount of public 3D breast ultrasound image data to construct a 3D breast ultrasound image dataset containing benign tumor samples and malignant tumor samples; The disclosed 3D breast ultrasound image data is obtained from patients undergoing 3D breast ultrasound examinations over the past few years. 2) Dataset delineation; Organize professional medical personnel to perform breast tissue contouring and review on the constructed three-dimensional breast ultrasound image dataset; 3) Dataset preprocessing; Performing targeted optimization and adjustment on the delineated and reviewed three-dimensional breast ultrasound image dataset, including image size adjustment, contrast adjustment, image intensity scaling adjustment, and horizontal flip adjustment; 4) Model training; Training the segmentation network model: The optimized and adjusted 3D breast ultrasound image dataset is used as the training input for the segmentation network model. Radiologists' annotated tumor regions, dense tissue regions, and glandular regions are introduced as gold standard supervisory signals. The segmentation network model is trained to learn the morphological features of breast-related tissues in the 3D breast ultrasound image dataset. Based on the learned morphological features, the model accurately segments the breast-related tissues in the 3D breast ultrasound image dataset. The model then outputs a mask image containing the segmentations of the tumor region, dense tissue region, and glandular region, as well as the feature map generated before the softmax layer. Training the classification network model: The spatially aware breast tissue mask image, the corresponding feature map, and the original 3D ultrasound breast image are used as training input for the classification network model. The spatially aware breast tissue mask image serves as prior knowledge to guide the classification network model's attention mechanism. This allows the classification network model to extract multi-scale spatially aware features, focusing on the characteristics of the tumor region in each view, and completing the benign and malignant tumor classification process for each view in the 3D ultrasound breast image. Prediction network model training: The benign and malignant tumor classification results for each view in the 3D ultrasound breast image are used as training input for the prediction network model. The 3D breast ultrasound image dataset and set thresholds are used as a benchmark to train the prediction network model to perform weighted average fusion decisions on benign tumors in multiple views. The prediction network model is then trained to predict the risk of malignant transformation of benign tumors in the next few years based on the benign tumor results in the weighted average fusion decision, thereby generating a risk assessment result for benign tumors presented in the form of structured text. 5) Model optimization; The segmentation network model that learns breast-related tissue morphological features during training is dynamically adjusted through the optimizer to obtain model weights with strong generalization capabilities. The classification network model focuses on the regional features of the tumor during training. By carefully selecting the loss function, optimizer, and learning rate, the classification network model can efficiently and accurately classify the regional features of the tumor into benign and malignant tumors. The prediction network model for multiple benign tumor fusion decisions during training is carefully selected through loss function, optimizer and learning rate, so that the prediction network model can deeply explore the details of the tumor lesion edge and the potential malignant signals of changes in the surrounding tissue environment.
[0010] Preferably, in the prediction network model training, the prediction method of the prediction network model is: A weighted average fusion decision is made based on the benign and malignant tumor classification results from multiple views. When the weighted average fusion decision result is the probability of a benign tumor, the prediction network model will conduct risk prediction training for benign tumors in the next few years based on the three-dimensional breast ultrasound image dataset, combined with the label information of whether the benign tumor probability will become cancerous and the corresponding feature map, to complete the cancer risk prediction of benign tumors at each time point in the next few years and generate corresponding results.
[0011] Preferably, after the medical image analysis model is trained, a comprehensive evaluation strategy is used to select the optimal model. The specific method of the comprehensive evaluation strategy is: The test was conducted on a hybrid dataset that combines a public 3D breast ultrasound image dataset and a private 3D breast ultrasound image dataset, where; The model with the highest DICE value on the mixed dataset will be selected as the weight of the final segmentation network model; The model with the highest AUC value on the mixed data set will be selected as the weight of the final classification network model; The model with the highest Time-dependent AUC value on the mixed dataset will be selected as the weight of the final prediction network model.
[0012] When the medical image analysis model is trained and used, when the doctor uses the ABUS / ABVS instrument to examine the patient, since there will be three views of each breast, the medical image analysis model will automatically analyze the key features in the three views and use the optimal weight parameters previously trained in the segmentation network model to accurately segment the tumor, dense tissue and glandular area in each view, and generate three breast tissue mask images of the corresponding views and save the feature map generated before the softmax layer.
[0013] To further improve the accuracy of diagnosis, doctors can review and make necessary fine-tuning on these breast tissue mask images before sending them to the final classification network model. This interactive design not only ensures the accuracy of tumor area delineation, but also provides a reliable data basis for subsequent benign and malignant classification and assessment of the risk of malignant transformation of benign tumors. After the doctor confirms or corrects the mask image, it enters the classification network model after the position information encoding module to generate a preliminary prediction result on the benign and malignant nature of the tumor. The prediction network model then makes a decision fusion on the results given by the three views and predicts the future risk of malignant transformation for the benign tumors in the decision results. The final malignant risk prediction result is presented to the doctor in the form of structured text for the doctor's reference.
[0014] Therefore, in this way, the present invention does not require doctors to spend too much time locating and measuring tumors and other tissue information. Instead, they only need to focus on the tumor area in the mask image automatically generated by the model and the corresponding feature map and make necessary subtle adjustments to it, which can greatly improve work efficiency and diagnostic accuracy.
[0015] Preferably, in step 3, the three-dimensional ultrasound breast image processing process is as follows: A segmentation network model is used to analyze breast tissue features in each view of a 3D ultrasound breast image. The model then accurately segments the tumor, dense tissue, and glandular regions in each view. The model then generates a tumor mask image, a dense tissue mask image, a glandular mask image, and a feature map generated before the softmax layer is saved for each view. Medical personnel are used to fine-tune the tumor mask image, dense tissue mask image, and gland mask image in each view; The fine-tuned tumor mask image, dense tissue mask image, gland mask image and corresponding feature map in each view are encoded in three-dimensional space through the position information encoding module to generate tumor mask images, dense tissue mask images, gland mask images and corresponding feature maps with spatial perception information.
[0016] Preferably, in step 3, the process of obtaining the benign or malignant tumor classification result is as follows: The classification network model is used to simultaneously receive the original input 3D ultrasound breast image, the tumor mask image with spatial perception information, the dense tissue mask image, the gland mask image and the corresponding feature map; After combining the simultaneously received image data, the convolutional neural network is used to extract multi-scale spatial information features of breast tissue images in three-dimensional ultrasound breast images. After extraction, the multi-scale spatial information is input into the encoder of the Transformer architecture. During this period, the decoder of the Transformer architecture is used to receive the tumor mask image with spatial perception information as the query object; Based on the multi-scale spatial information received by the encoder and the query object received by the decoder, the classification network model can perform the first-stage prediction processing on the benign and malignant nature of the tumor in each view in the three-dimensional ultrasound breast image to obtain the first-stage classification result of the benign and malignant tumor in each view.
[0017] Preferably, in step 3, the process of obtaining the prediction result of the risk of malignant transformation of benign tumors is as follows: The first-stage classification results of each view are further predicted for benign or malignant status based on the weighted average fusion decision to obtain the second-stage prediction results of the benign or malignant status of the tumor. If the second-stage prediction results determine that the tumor is benign, the prediction network model will predict the risk of malignant transformation of the benign tumor in the next few years to obtain the prediction results of the risk of malignant transformation of the benign tumor; The weighted average fusion decision-making method is as follows: each view in the 3D ultrasound breast image is given a certain weight, and a benign tumor probability threshold is set for the medical image analysis model; When the first-stage classification results of each view are fused through weighted average, and the probability of a benign tumor in the calculated output value is less than the set threshold, the tumor is determined to be malignant, and the prediction process ends; When the first-stage classification results of each view are fused through weighted average, and the probability of benign tumors in the calculated output value is greater than or equal to the set threshold, the tumor is judged to be benign, so as to predict the risk of malignant transformation of benign tumors in the next few years.
[0018] A computer device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, which enables the processor to perform operations corresponding to the above-mentioned spatially aware mask-cued multi-view fusion tumor prediction method.
[0019] A computer storage medium stores at least one executable instruction, wherein the executable instruction enables a processor to execute operations corresponding to the above-mentioned spatially-aware mask-cued fusion of multiple views for tumor prediction.
[0020] In this study, the proposed AutoSAMask-prompt module demonstrates excellent overall performance in the tasks of benign and malignant classification and malignant risk assessment of 3D ultrasound breast images. This module automatically segments and generates mask images of the tumor, dense tissue, and glandular regions, along with corresponding feature maps. This allows physicians to interactively fine-tune the generated mask images to ensure accurate delineation of the tumor region. Furthermore, the optimized mask images of the tumor, dense tissue, and glandular regions, along with their corresponding feature maps, are fed into a convolutional neural network (CNN) along with the original 3D ultrasound breast image to extract multi-scale spatial information. This rich feature information is then fed into an encoder and decoder to further analyze the key characteristics of the lesion. This approach significantly increases the amount of information processed by the model, enabling it to more accurately identify key features of the lesion region, significantly improving the accuracy of benign and malignant classification and providing a solid foundation for malignant risk assessment of benign lesions.
[0021] In addition, considering that three different views are generated during the ABUS (Automated Breast Ultrasound System) or ABVS (Automated Breast Volume Scanning) imaging process, the present invention needs to execute the processes of mask image generation, physician review and fine-tuning, and preliminary classification network prediction for each view separately, and finally fuse the results of the three views to provide a comprehensive prediction result. This not only enables the model to focus on the provided prompt information and efficiently extract key features in the ultrasound image, but also provides more accurate and rich sign references for subsequent assessment of the probability of benign tumors becoming malignant.
[0022] Overall, this invention not only provides doctors with a robust and reliable diagnostic reference, assisting them in making more accurate clinical decisions, but also, by enhancing the model's understanding and accuracy of lesions, strongly supports the medical decision-making process. This is particularly significant in assessing the subsequent risk of malignant transformation of benign tumors. This innovative approach significantly improves diagnostic efficiency and accuracy, providing patients with a better, more proactive healthcare experience.
[0023] Compared with existing technologies, this invention introduces an explicit prompt mechanism into the 3D ultrasound image network model. This prompt mechanism guides the model to more efficiently exploit subtle features in 3D breast ultrasound images, significantly improving the model's performance in classification tasks. Furthermore, combined with transfer learning methods, the proposed model can adapt to different organs and various medical imaging data types. This design not only improves the model's generalization capabilities but also has important clinical implications for AI-assisted diagnosis.
[0024] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the invention and to implement it according to the contents of the description, the following preferred embodiments of the present invention are described in detail with reference to the accompanying drawings. The specific implementation methods of the present invention are given in detail by the following embodiments and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Schematic diagram of the training process of the medical image analysis model in the present invention.
[0026] Figure 2 Schematic diagram of the network architecture and process of the medical image analysis model of the present invention. DETAILED DESCRIPTION
[0027] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0028] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways than those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0029] refer to Figure 1-Figure 2 As shown, the present invention first proposes a multi-task model for self-generated spatial-aware mask prompt fusion multi-view decision-making, that is, a medical image analysis model, which consists of a segmentation network model, a fine-tuning module, a position information encoding module, a classification network model and a prediction network model.
[0030] 1. Segmentation network model; The segmentation network model structure of this invention is a deep learning model based on the U-Net3D architecture. It can achieve accurate segmentation of tumor regions, dense tissue regions, and glandular regions through a supervised learning framework. Therefore, the segmentation network model can be used to accurately segment breast tissue in each view of a 3D ultrasound breast image, generate a breast tissue mask image corresponding to the 3D ultrasound breast image view, and save the feature map generated before the softmax layer.
[0031] Among them, the breast tissue mask images generated by segmentation in the segmentation network model include: tumor mask image, dense tissue mask image and gland mask image.
[0032] 2. Fine-tuning module; The fine-tuning module of the present invention allows medical personnel to manually fine-tune the segmented breast tissue mask image, enabling interactive fine-tuning of the breast tissue mask image. During interactive fine-tuning, medical personnel can perform secondary verification and fine-tuning of the breast tissue mask image as needed to ensure accurate delineation of the tumor area and a close match between the lesion area and the actual condition.
[0033] 3. Position information encoding module; The position information encoding module of the present invention is used to perform three-dimensional spatial position encoding on the breast tissue mask image and the corresponding feature map that have been manually fine-tuned by medical personnel, so that the breast tissue mask image can not only reflect the morphological characteristics of tumors, dense tissues and glands, but also carry their position information in three-dimensional space, thereby generating a breast tissue mask image and the corresponding feature map with spatial perception information.
[0034] The position information encoding module of the present invention provides additional spatial structural clues for the medical image analysis model, enabling the breast tissue mask image to have spatial perception capabilities and carry rich spatial semantic information, which helps to improve the medical image analysis model's understanding and recognition accuracy of tumor areas, dense tissue areas, and glandular areas.
[0035] 4. Classification network model; The classification network model structure of the present invention adopts a deep learning architecture, integrating a three-dimensional convolutional neural network (3DCNN) that can accept masks as prompt information and a Transformer architecture. It is responsible for simultaneously receiving a breast tissue mask image with spatial perception information, a corresponding feature map, and the original three-dimensional ultrasound breast image to extract the multi-scale spatial information of the breast tissue image. Based on the multi-scale spatial information of the breast tissue image, the model classifies the benign and malignant tumors in each view of the three-dimensional ultrasound breast image to obtain the classification result.
[0036] Among them, the present invention combines the breast tissue mask image that carries spatial perception information and has undergone data enhancement preprocessing with the original three-dimensional ultrasound breast image. After feature extraction through a convolutional neural network, it is integrated into a vector containing multi-scale features and multi-scale spatial information and passed to the encoder part of the Transformer architecture. In addition, the mask images of tumors, dense tissues, and glands and the corresponding feature maps are used as query objects (Object Queries) and combined with the decoder part of the model. This design helps guide the medical image analysis model to pay more attention to the spatial characteristics and contour details of tumors, dense tissues, and glands, thereby enhancing the effectiveness of classification and prediction tasks. The classification task is completed using a classification network model, and the prediction task is completed using a prediction network module.
[0037] 5. Prediction network model; The prediction network model of the present invention is used to perform weighted average fusion decision on the benign and malignant tumor classification results of each view in the three-dimensional ultrasound breast image, and is used to predict the risk of malignant transformation of benign tumors in the weighted average fusion decision in the next few years, so as to generate a risk assessment result for benign tumors.
[0038] Among them, after the weighted average fusion decision prediction results of benign tumors, the prediction network model will perform risk prediction on the benign tumors in the decision results to realize the canceration risk prediction of benign tumors at each time node in the next few years, and generate corresponding results after the prediction is completed.
[0039] It is worth noting that the medical image analysis model of the present invention innovatively introduces a decision fusion mechanism, which not only improves the reliability of diagnostic results and makes diagnostic conclusions more accurate and reliable, but also creates a more intuitive and easy-to-understand decision support tool for medical professionals, helping medical staff to make more efficient and accurate judgments in medical diagnosis work.
[0040] In summary, the medical image analysis model proposed in this invention can accept 3D ultrasound breast image data as input and generate mask images of tumors, dense tissue, and glands for secondary verification and fine-tuning by physicians to ensure spatial alignment with the actual lesion area. Subsequently, by positionally encoding the mask image and its corresponding feature map in 3D space, it acquires spatial awareness and carries rich spatial semantic information. This information is combined with the original 3D ultrasound breast image and, after multi-layer feature extraction using a convolutional neural network, is integrated into a collection of multi-scale image and spatial features that is fed into the encoder. Furthermore, the generated mask image and its corresponding feature map serve as query objects and are fed into the decoder, helping to guide the variable attention mechanism to understand the characteristic information of tumor, dense tissue, and gland regions. Furthermore, by using the spatially aware breast tissue mask image and its corresponding feature map as prior knowledge to guide the attention mechanism of the classification network model, the present invention can effectively improve classification performance and the prediction of the probability of malignant transformation of benign tumors. And by weighted fusion of the final classification results of the three views, the classification performance of the model, the performance of predicting the probability of malignant transformation of benign tumors, as well as the reference value and interpretability can be further improved.
[0041] It's worth emphasizing that this invention simplifies the diagnostic process to a certain extent. This streamlined process not only improves diagnostic efficiency but also allows doctors to fine-tune the generated mask image, eliminating the need to consider complex characteristics such as blood supply, density, or elasticity. This significantly reduces the workload and helps doctors diagnose conditions more accurately and quickly. It's also important to emphasize that the fine-tuned mask image files saved each time in this solution can be used to further expand the training model dataset, providing a data foundation for model iteration.
[0042] The present invention also proposes a training strategy for the above-mentioned medical image analysis model, the specific contents of which are as follows: 1) Dataset construction; The data collected can be the three-dimensional breast ultrasound image data publicly available from various hospitals or private three-dimensional breast ultrasound image data, wherein the public or private three-dimensional breast ultrasound image data are all image data obtained from patients who have continuously undergone three-dimensional breast ultrasound examinations in the past few years.
[0043] The training dataset used in this invention consists of a collection of publicly available 3D breast ultrasound image data, acquired from patients undergoing 3D breast ultrasound examinations over the past five years. Therefore, a 3D breast ultrasound image dataset containing both benign and malignant tumor samples can be constructed using this data.
[0044] The constructed three-dimensional breast ultrasound image dataset is based on the ultrasound data of patients over the past five years, which ensures that the medical image analysis model can fully learn the diverse pathological characteristics of benign tumors and the potential correlation between benign tumors and the risk of malignancy.
[0045] 2) Dataset delineation; Organize professional medical personnel to perform breast tissue contouring and review on the constructed three-dimensional breast ultrasound image dataset; After collecting a large dataset of 3D breast ultrasound images, experienced professionals are required to precisely delineate the contours of tumor areas, dense tissue areas, and glandular regions. These delineation results are then reviewed by senior experts. Therefore, through a series of rigorous quality control and optimization processes, we ensure that the delineated mask data is highly accurate and consistent, effectively improving the reliability and accuracy of the segmentation network model.
[0046] 3) Dataset preprocessing; To improve the segmentation performance of the medical image analysis model, targeted optimization and adjustment are required for the delineated and reviewed three-dimensional breast ultrasound image dataset at the initial stage of training. Optimization and adjustment include, but are not limited to, resizing, contrast adjustment, image intensity scaling, and horizontal flipping of all images.
[0047] The optimization and adjustment of the dataset preprocessing specifically includes adjusting all images to a fixed size to ensure consistent image dimensions; implementing random flipping to enrich the image perspective; performing image intensity scaling to adjust the overall brightness and darkness range of the image; and random contrast enhancement to flexibly improve the contrast differences between different areas in the image.
[0048] It's worth emphasizing that, given that the segmentation network model task requires mask annotation files, the aforementioned dataset preprocessing optimizations were also applied to the corresponding mask image files, including fixed image size and random flipping to maintain consistency between the disease annotation information and the image. This approach not only increases the diversity of the 3D breast ultrasound image dataset but also ensures that the segmentation network model fully learns the characteristics of breast-related tissues (tumors, dense tissue, glands) during training, thereby improving the accuracy and reliability of the generated mask images.
[0049] 4) Model training; After completing data collection and data preprocessing, save the preprocessed data in the specified format and file path, and then use code instructions to let the model learn and reason.
[0050] Training of the segmentation network model: The optimized and adjusted 3D breast ultrasound image dataset is used as the training input of the segmentation network model, and the tumor area, dense tissue area, and glandular area annotated by radiologists are introduced as the gold standard (Ground Truth Mask) supervision signal. The segmentation network model is trained to learn the morphological features of breast-related tissues in the 3D breast ultrasound image dataset, and accurately segment the breast-related tissues in the 3D breast ultrasound image dataset based on the learned morphological features. The model outputs a mask image containing the segmentation of the tumor area, dense tissue area, and glandular area, as well as the feature map generated before saving the softmax layer. The mask image format is NRRD format.
[0051] Training of the classification network model: The breast tissue mask image with spatial perception information, the corresponding feature map, and the original 3D ultrasound breast image are used as the training input of the classification network model. The breast tissue mask image with spatial perception information and the corresponding feature map are used as prior knowledge to guide the attention mechanism of the classification network model. The classification network model is trained to extract multi-scale spatial perception information features, focusing on the features of the tumor area in each view, and completing the benign and malignant tumor classification processing in each view of the 3D ultrasound breast image.
[0052] refer to Figure 2 As shown, during classification network model training, the mask image is encoded in 3D space by the position information encoding module. The fused image is then processed through a 3D convolutional neural network (3DCNN) for feature extraction, integrating multi-scale features to preserve small object details while also carrying multi-scale spatial perception information. This fusion is then input into the Transformer architecture's encoder for deep feature learning. Furthermore, the decoder receives the fused 3D spatial position encoded tumor mask image, dense tissue mask image, gland mask image, and corresponding feature maps as query inputs, guiding the model's variable attention mechanism to focus on features in the tumor region.
[0053] Training of the prediction network model: The benign and malignant tumor classification processing results of each view in the three-dimensional ultrasound breast image are used as the training input of the prediction network model. The three-dimensional breast ultrasound image dataset and the set benign tumor probability threshold are used as the benchmark to train the prediction network model to make weighted average fusion decisions on benign tumors in multiple views. The prediction network model is trained to predict the risk of malignancy of benign tumors in the fusion decision in the next few years based on the benign tumor results in the weighted average fusion decision, so as to generate a risk assessment result for benign tumors presented in the form of structured text.
[0054] Specifically, after using the predictive network model to classify benign and malignant stages, data predicted to be benign will undergo the next step of training. Based on the imaging data from the past five years and the labels indicating whether or not a sample has become cancerous, the risk prediction network model will be trained on the benign samples. A separate binary classification training for cancer will be conducted for each time point in the next five years. This allows the model to predict cancer risk at each time point in the next five years based on image information. Furthermore, through the AutoSAMask-prompt module, which integrates the self-generated mask prompts of the lesion area with spatially aware coding, the risk prediction network model is guided to deeply explore potential malignant signals, such as details of the tumor lesion edge and changes in the surrounding tissue environment, based on the analysis of multi-scale spatial features and lesion morphology information, thereby outputting accurate and comprehensive risk assessment results.
[0055] In the prediction network model training of the present invention, the prediction method of the prediction network model is: A weighted average fusion decision is made based on the benign and malignant tumor classification results from multiple views. When the weighted average fusion decision result is the probability of a benign tumor, the prediction network model will be trained to predict the risk of benign tumors in the next five years based on the three-dimensional breast ultrasound image dataset, combined with the label information of whether the benign tumor probability will become cancerous and the corresponding feature map, to complete the prediction of the canceration risk of benign tumors at each time point in the next five years and generate corresponding results.
[0056] 5) Model optimization; The segmentation network model that learns breast-related tissue morphological features during training is dynamically adjusted through the optimizer to obtain model weights with strong generalization capabilities. The classification network model focuses on the regional features of the tumor during training. By carefully selecting the loss function, optimizer, and learning rate, the classification network model can efficiently and accurately classify the regional features of the tumor into benign and malignant tumors. The prediction network model for multiple benign tumor fusion decisions during training is carefully selected through loss function, optimizer and learning rate, so that the prediction network model can deeply explore the details of the tumor lesion edge and the potential malignant signals of changes in the surrounding tissue environment.
[0057] After the training of the above medical image analysis model is completed, a comprehensive evaluation strategy needs to be used to select the optimal model. The specific method of the comprehensive evaluation strategy is as follows: The test was conducted on a hybrid dataset that combines a public 3D breast ultrasound image dataset and a private 3D breast ultrasound image dataset, where; The model with the highest DICE value on the mixed dataset will be selected as the weight of the final segmentation network model; The model with the highest AUC value on the mixed data set will be selected as the weight of the final classification network model; The model with the highest Time-dependent AUC value on the mixed dataset will be selected as the weight of the final prediction network model.
[0058] Based on the selection of the best model application display: In the medical image analysis model proposed in the present invention, the optimal segmentation, classification and risk prediction weights determined in the training phase are first utilized.
[0059] During the data acquisition phase, three views of each breast (i.e., anteroposterior, lateral, and medial-lateral oblique views) are acquired from ABUS (Automated Breast Ultrasound) or ABVS (Automated Breast Volume Imaging System). These views are then processed through an optimized segmentation network model to generate NRRD-format images containing three types of masks (tumor, dense tissue, and glandular tissue), along with corresponding feature maps. Physicians can fine-tune the generated masks as needed to ensure that the lesion areas closely match the actual conditions. After fine-tuning, the data enters the classification phase. Each view and its corresponding mask are input into the classification network model, and their respective classification results are obtained. Finally, the prediction network model performs a weighted average of the three classification results to arrive at a comprehensive prediction. The subsequent comprehensive prediction analysis report not only details the various tissue characteristics of the patient's breast but also clearly indicates whether a malignant lesion or a benign lesion is present. If the comprehensive prediction concludes that the test result is benign, the predictive network model will further assess the probability of malignant transformation of the benign tumor in the next few years, providing an important reference for subsequent follow-up and clinical decision-making. This structured document format makes diagnostic information clear at a glance, greatly facilitating the clinical decision-making process, not only reducing the workload of doctors, but also laying the foundation for more accurate diagnosis and treatment recommendations. Overall, this method significantly optimizes the medical imaging analysis process and improves the efficiency and accuracy of diagnosis and treatment.
[0060] It's worth noting that in real-world applications, due to the varying standards of different sampling devices, image data can vary. To ensure that the model performs well on new datasets, it's necessary to fine-tune the model. Fine-tuning involves further training a pre-trained model on a small dataset specific to a specific task to optimize its performance. This process allows the model to better adapt to sampling devices of varying standards and improves its performance in diverse environments.
[0061] Furthermore, fine-tuning empowers the model with cross-organ and cross-disease classification and prediction capabilities, expanding its application scenarios. For example, after appropriate fine-tuning, this model can be applied to disease diagnosis in other organs, greatly improving its generalization and applicability. During the fine-tuning process, appropriate fine-tuning strategies, data augmentation strategies, and optimization methods are selected to achieve the optimal model for the new task. It should be noted that this type of fine-tuning is well known and applicable to those skilled in the art based on their industry knowledge, and this proposal will not elaborate on this in detail.
[0062] refer to Figure 2 As shown, based on the above-trained medical image analysis model, the present invention also proposes a tumor prediction method based on spatially aware mask-cued fusion of multiple views, including the following steps: Step 1) obtaining a three-dimensional ultrasound breast image of the patient; Utilize ABUS / ABVS instruments to examine patients and obtain their three-dimensional ultrasound breast images; The acquired three-dimensional ultrasound breast images include: anterior-posterior view, lateral view, and mediolateral oblique view of the patient's breast.
[0063] Step 2) inputting the acquired 3D ultrasound breast image into a trained medical image analysis model; Step 3) Processing and analyzing the input 3D ultrasound breast image using a medical image analysis model to obtain a benign or malignant tumor classification result and a benign tumor malignant transformation risk prediction result in the 3D ultrasound breast image; In step 3, the three-dimensional ultrasound breast image processing process is as follows: A segmentation network model is used to analyze breast tissue features in each view of a 3D ultrasound breast image. The model then accurately segments the tumor, dense tissue, and glandular regions in each view. The model then generates a tumor mask image, a dense tissue mask image, a glandular mask image, and a feature map generated before the softmax layer is saved for each view. Medical personnel are used to fine-tune the tumor mask image, dense tissue mask image, and gland mask image in each view; The fine-tuned tumor mask image, dense tissue mask image, gland mask image and corresponding feature map in each view are encoded in three-dimensional space through the position information encoding module to generate tumor mask images, dense tissue mask images, gland mask images and corresponding feature maps with spatial perception information.
[0064] Since each breast has three views, the medical image analysis model automatically analyzes the key features in the three views and uses the optimal weight parameters trained in the above-mentioned segmentation network model to accurately segment the tumor, dense tissue, and glandular areas in each view. This generates three mask images for the corresponding views and saves the feature map generated before the softmax layer.
[0065] To further enhance diagnostic accuracy, medical professionals can review and fine-tune the segmented breast tissue mask before sending it to the classification network model. This interactive design not only ensures accurate delineation of the tumor region but also provides a reliable data foundation for subsequent benign and malignant classification and assessment of the risk of malignant transformation of benign tumors.
[0066] Moreover, through this interactive design approach, doctors do not need to spend too much time locating and measuring tumors and other tissue information. Instead, they only need to focus on the tumor area in the mask image automatically generated by the model and make necessary subtle adjustments to it, which can greatly improve work efficiency and diagnostic accuracy.
[0067] After the doctor confirms or corrects the breast tissue mask images of the three corresponding views, the present invention can perform position encoding on the mask images and the corresponding feature maps in three-dimensional space, so that the mask images of the three corresponding views have spatial perception capabilities and carry rich spatial semantic information. They are then entered into the classification network model to generate preliminary prediction results (i.e., the first-stage results) on the benign and malignant nature of the tumor and obtain the classification results in each view respectively.
[0068] The process of classifying benign and malignant tumors is as follows: The classification network model is used to simultaneously receive the original input 3D ultrasound breast image, the tumor mask image with spatial perception information, the dense tissue mask image, the gland mask image and the corresponding feature map; After combining the simultaneously received image data, the convolutional neural network is used to extract multi-scale spatial information features of breast tissue images in three-dimensional ultrasound breast images. After extraction, the multi-scale spatial information is input into the encoder of the Transformer architecture. During this period, the decoder of the Transformer architecture is used to receive the tumor mask image with spatial perception information and the corresponding feature map as the query object; Based on the multi-scale spatial information received by the encoder and the query object received by the decoder, the classification network model can perform the first-stage prediction processing on the benign and malignant nature of the tumor in each view in the three-dimensional ultrasound breast image to obtain the first-stage classification result of the benign and malignant tumor in each view.
[0069] It's worth noting that for tumors classified as benign in the first stage of the classification network model, a weighted average fusion decision can be used to perform a second-stage prediction of multiple benign and malignant tumor results to determine the overall benign and malignant outcome. If the overall decision results in a benign tumor, the prediction network model uses high-quality mask image information to further assess the probability of malignant transformation within the next few years, thereby providing comprehensive decision-making support for doctors to develop more precise follow-up and intervention strategies.
[0070] The process of predicting the risk of malignant transformation of benign tumors is as follows: The first-stage classification results of each view are further predicted for benign or malignant status based on the weighted average fusion decision to obtain the second-stage prediction results of the benign or malignant status of the tumor. If the second-stage prediction results determine that the tumor is benign, the prediction network model will predict the risk of malignant transformation of the benign tumor in the next few years to obtain the prediction results of the risk of malignant transformation of the benign tumor; Among them, the weighted average fusion decision-making method is: setting a certain weight for each view in the three-dimensional ultrasound breast image, and setting a benign tumor probability threshold for the medical image analysis model.
[0071] Specifically, the weighted average fusion decision is based on the weights assigned by medical professionals to the anteroposterior, lateral, and medial-lateral oblique views of the patient's breast. For example, the anteroposterior view is weighted 0.5, the lateral view 0.3, and the medial-lateral oblique view 0.2. The medical image analysis model sets a benign tumor probability threshold of 0.5, denoted as C.
[0072] When the first-stage classification results of each view are fused through weighted average, and the probability of a benign tumor in the calculated output value is less than the set threshold, the tumor is determined to be malignant, and the prediction process ends; When the first-stage classification results of each view are fused through weighted average, and the probability of a benign tumor in the calculated decision output value is greater than or equal to the set threshold, the tumor is judged to be benign, so as to predict the risk of malignant transformation of benign tumors in the next few years.
[0073] The first-stage classification results of each view are fused by weighted average, and the decision output value is calculated and recorded as [A, B]; where A+B equals 1, A represents the probability of benignity, and B represents the probability of malignancy.
[0074] When A<C, the tumor can be judged as malignant; when A≥C, the tumor is judged as benign.
[0075] like: When the calculated decision output value is [0.2, 0.8], 0.2 < 0.5, the tumor is determined to be malignant.
[0076] When the calculated decision output value is [0.8, 0.2], 0.8>0.5, the tumor is determined to be benign.
[0077] When the decision output value is calculated as [0.5, 0.5], 0.5=0.5, and the tumor is determined to be benign.
[0078] Step 4) Present the benign tumor malignant transformation risk prediction results in structured text format to generate a risk assessment result for benign tumors.
[0079] It is worth noting that the analysis report generated in structured text format can record in detail the various tissue characteristics of the patient's breast and clearly indicate whether there are malignant lesions or only benign lesions. This structured document format makes diagnostic information clear at a glance, greatly facilitating the clinical decision-making process, not only reducing the workload of doctors, but also laying the foundation for more accurate diagnosis and treatment recommendations. Overall, this invention significantly optimizes the medical image analysis process and improves the efficiency and accuracy of diagnosis and treatment.
[0080] In summary, the present invention constructs a CNN-Transformer multi-scale feature hybrid multi-task network driven by tumor annotation. In the model architecture, the CNN branch captures the microscopic morphological features of the tumor through local texture feature extraction, while the Transformer branch dynamically focuses on the global contextual associations of the tumor region through a deformable attention mechanism. In particular, by introducing the mask annotation of the tumor region as an explicit prompt for the classification network model, this prompt is embedded in the cross-attention layer of the encoder-decoder in the form of spatial encoding, guiding the model to accurately distinguish between benign and malignant representations of the tumor region, significantly suppressing background interference from non-tumor tissue (such as fat and blood vessels), and improving the interpretability of classification decisions and the probability of benign to malignant transformation.
[0081] In addition, the present invention proposes a module called Self-generated Mask Prompt Fusion Spatial Aware Coding (AutoSAMask-prompt), which can automatically segment and generate masks of tumors, dense tissues, glands and corresponding feature maps on datasets without prior mask labels or lacking mask annotations. Among them, the tumor mask and the corresponding feature map are used as prompts to input the encoder (Encoder) and decoder (Decoder) parts of the model, which enables the present invention to improve training efficiency and classification accuracy while maintaining stable and reliable performance in scenarios with insufficient data, thereby constructing an efficient training and accurate prediction model for benign and malignant breast ultrasound images, effectively improving clinical diagnosis efficiency.
[0082] In addition, due to the good flexibility of the model, it can be applied to the task of classifying benign and malignant tumor lesions in different organs through simple fine-tuning, greatly expanding its scope of application and practicality.
[0083] It should be noted that to help clinicians better assess subsequent risks, once the model predicts a benign lesion, the system will further provide an assessment of the potential probability of malignancy and a forecast of the risk of malignancy over the next few years. By constructing an additional regression branch for malignancy risk during the training phase and using supervised learning data from past cases, once the input is judged to be benign, the model will output the corresponding malignancy risk coefficient and a risk change curve over several years. This not only provides clinicians with more detailed diagnostic references, but also offers more forward-looking guidance for subsequent patient follow-up and intervention strategy formulation.
[0084] The present invention also proposes a computer device comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus, and the memory is used to store at least one executable instruction, wherein the executable instruction enables the processor to execute operations corresponding to the above-mentioned spatially-aware mask-cued fusion of multiple views tumor prediction method.
[0085] The present invention also proposes a computer storage medium, which stores at least one executable instruction, and the executable instruction enables a processor to perform operations corresponding to the above-mentioned spatially aware mask-cued fusion of multiple views of tumor prediction method.
[0086] The computer storage medium of the present invention can be a hard disk, a solid-state drive (SSD), network storage (such as cloud storage), etc. When training large-scale models or processing large-scale data sets, it is usually considered to use high-speed storage media to improve the efficiency of data reading and writing.
[0087] It is worth emphasizing that in the context of this article, "mask" and "mask" have the same meaning.
[0088] Throughout this specification, terms such as "one embodiment," "some embodiments," and "specific embodiments" mean that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0089] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and alterations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A tumor prediction method based on spatially aware mask-cued fusion of multiple views, characterized by: The following steps are involved: Step 1) obtaining a three-dimensional ultrasound breast image of the patient; Step 2) inputting the acquired 3D ultrasound breast image into a trained medical image analysis model; Step 3) Processing and analyzing the input 3D ultrasound breast image using a medical image analysis model to obtain a benign or malignant tumor classification result and a benign tumor malignant transformation risk prediction result in the 3D ultrasound breast image; Step 4) Present the benign tumor malignant transformation risk prediction results in structured text format to generate a risk assessment result for benign tumors.
2. The spatially aware mask-cued multi-view fusion tumor prediction method according to claim 1, characterized in that: In step 1, the patient is examined using an ABUS / ABVS instrument to obtain a three-dimensional ultrasound breast image; The acquired three-dimensional ultrasound breast images include: anterior-posterior view, lateral view, and mediolateral oblique view of the patient's breast.
3. The spatially aware mask-cued multi-view fusion tumor prediction method according to claim 1, characterized in that: In step 2, the medical image analysis model includes: A segmentation network model is used to segment breast-related tissue in each view of the 3D ultrasound breast image and generate a breast tissue mask image corresponding to each view of the 3D ultrasound breast image and save the feature map generated before the softmax layer. The breast tissue mask image includes: a tumor mask image, a dense tissue mask image, and a gland mask image; A fine-tuning module, used for medical personnel to manually perform fine-tuning operations on the breast tissue mask image generated by segmentation; A position information encoding module is used to perform three-dimensional spatial position encoding on the breast tissue mask image and the corresponding feature map after manual fine-tuning by medical personnel to generate a breast tissue mask image with spatial perception information; The classification network model is responsible for simultaneously receiving the breast tissue mask image with spatial perception information, the corresponding feature map, and the original 3D ultrasound breast image to extract the multi-scale spatial information of the breast tissue image. Based on the multi-scale spatial information of the breast tissue image, it classifies the tumor as benign or malignant in each view of the 3D ultrasound breast image. The prediction network model is used to perform weighted average fusion decision on the benign and malignant tumor classification results of each view in the three-dimensional ultrasound breast image, and is used to predict the risk of malignant transformation of benign tumors in the weighted average fusion decision in the next few years to generate a risk assessment result for benign tumors.
4. The spatially aware mask-cued multi-view fusion tumor prediction method according to claim 3, characterized in that: The training strategy of the medical image analysis model is as follows: 1) Dataset construction; Collect a certain amount of public 3D breast ultrasound image data to construct a 3D breast ultrasound image dataset containing benign tumor samples and malignant tumor samples; The disclosed 3D breast ultrasound image data is obtained from patients undergoing 3D breast ultrasound examinations over the past few years. 2) Dataset delineation; Organize professional medical personnel to perform breast tissue contouring and review on the constructed three-dimensional breast ultrasound image dataset; 3) Dataset preprocessing; Performing targeted optimization and adjustment on the delineated and reviewed three-dimensional breast ultrasound image dataset, including image size adjustment, contrast adjustment, image intensity scaling adjustment, and horizontal flip adjustment; 4) Model training; Training the segmentation network model: The optimized and adjusted 3D breast ultrasound image dataset is used as the training input for the segmentation network model. Radiologists' annotated tumor regions, dense tissue regions, and glandular regions are introduced as gold standard supervisory signals. The segmentation network model is trained to learn the morphological features of breast-related tissues in the 3D breast ultrasound image dataset. Based on the learned morphological features, the model accurately segments the breast-related tissues in the 3D breast ultrasound image dataset. The model then outputs a mask image containing the segmentations of the tumor region, dense tissue region, and glandular region, as well as the feature map generated before the softmax layer. Training the classification network model: The spatially aware breast tissue mask image, the corresponding feature map, and the original 3D ultrasound breast image are used as training input for the classification network model. The spatially aware breast tissue mask image serves as prior knowledge to guide the classification network model's attention mechanism. This allows the classification network model to extract multi-scale spatially aware features, focusing on the characteristics of the tumor region in each view, and completing the benign and malignant tumor classification process for each view in the 3D ultrasound breast image. Prediction network model training: The benign and malignant tumor classification results for each view in the 3D ultrasound breast image are used as training input for the prediction network model. The 3D breast ultrasound image dataset and set thresholds are used as a benchmark to train the prediction network model to perform weighted average fusion decisions on benign tumors in multiple views. The prediction network model is then trained to predict the risk of malignant transformation of benign tumors in the next few years based on the benign tumor results in the weighted average fusion decision, thereby generating a risk assessment result for benign tumors presented in the form of structured text. 5) Model optimization; The segmentation network model that learns breast-related tissue morphological features during training is dynamically adjusted through the optimizer to obtain model weights with strong generalization capabilities. The classification network model focuses on the regional features of the tumor during training. By carefully selecting the loss function, optimizer, and learning rate, the classification network model can efficiently and accurately classify the regional features of the tumor into benign and malignant tumors. The prediction network model for multiple benign tumor fusion decisions during training is carefully selected through loss function, optimizer and learning rate, so that the prediction network model can deeply explore the details of the tumor lesion edge and the potential malignant signals of changes in the surrounding tissue environment.
5. The spatially aware mask-cued multi-view fusion tumor prediction method according to claim 4, characterized in that: During the prediction network model training, the prediction method of the prediction network model is: A weighted average fusion decision is made based on the benign and malignant tumor classification results from multiple views. When the weighted average fusion decision result is the probability of a benign tumor, the prediction network model will conduct risk prediction training for benign tumors in the next few years based on the three-dimensional breast ultrasound image dataset, combined with the label information of whether the benign tumor probability will become cancerous and the corresponding feature map, to complete the cancer risk prediction of benign tumors at each time point in the next few years and generate corresponding results.
6. The spatially-aware mask-cued multi-view fusion tumor prediction method according to claim 4, characterized in that: After the medical image analysis model is trained, a comprehensive evaluation strategy is used to select the optimal model. The specific method of the comprehensive evaluation strategy is as follows: The test was conducted on a hybrid dataset that combines a public 3D breast ultrasound image dataset and a private 3D breast ultrasound image dataset, where; The model with the highest DICE value on the mixed dataset will be selected as the weight of the final segmentation network model; The model with the highest AUC value on the mixed data set will be selected as the weight of the final classification network model; The model with the highest Time-dependent AUC value on the mixed dataset will be selected as the weight of the final prediction network model.
7. The method for tumor prediction based on spatially aware mask-cued fusion of multiple views according to claim 1, characterized in that: In step 3, the three-dimensional ultrasound breast image processing process is as follows: A segmentation network model is used to analyze breast tissue features in each view of a 3D ultrasound breast image. The model then accurately segments the tumor, dense tissue, and glandular regions in each view. The model then generates a tumor mask image, a dense tissue mask image, a glandular mask image, and a feature map generated before the softmax layer is saved for each view. Medical personnel are used to fine-tune the tumor mask image, dense tissue mask image, and gland mask image in each view; The fine-tuned tumor mask image, dense tissue mask image, gland mask image and corresponding feature map in each view are encoded in three-dimensional space through the position information encoding module to generate tumor mask images, dense tissue mask images, gland mask images and corresponding feature maps with spatial perception information.
8. The spatially aware mask-cued multi-view fusion tumor prediction method according to claim 7, characterized in that: In step 3, the process of obtaining the benign and malignant tumor classification results is as follows: The classification network model is used to simultaneously receive the original input 3D ultrasound breast image, the tumor mask image with spatial perception information, the dense tissue mask image, the gland mask image and the corresponding feature map; After combining the simultaneously received image data, the convolutional neural network is used to extract multi-scale spatial information features of breast tissue images in three-dimensional ultrasound breast images. After extraction, the multi-scale spatial information is input into the encoder of the Transformer architecture. During this period, the decoder of the Transformer architecture is used to receive the tumor mask image with spatial perception information as the query object; Based on the multi-scale spatial information received by the encoder and the query object received by the decoder, the classification network model can perform the first-stage prediction processing on the benign and malignant nature of the tumor in each view in the three-dimensional ultrasound breast image to obtain the first-stage classification result of the benign and malignant tumor in each view.
9. The method for tumor prediction based on spatially aware mask-cued fusion of multiple views according to claim 8, characterized in that: In step 3, the process of obtaining the prediction result of the risk of malignant transformation of benign tumors is as follows: The first-stage classification results of each view are further predicted for benign or malignant status based on the weighted average fusion decision to obtain the second-stage prediction results of the benign or malignant status of the tumor. If the second-stage prediction results determine that the tumor is benign, the prediction network model will predict the risk of malignant transformation of the benign tumor in the next few years to obtain the prediction results of the risk of malignant transformation of the benign tumor; The weighted average fusion decision-making method is as follows: each view in the 3D ultrasound breast image is given a certain weight, and a benign tumor probability threshold is set for the medical image analysis model; When the first-stage classification results of each view are fused through weighted average, and the probability of a benign tumor in the calculated output value is less than the set threshold, the tumor is determined to be malignant, and the prediction process ends; When the first-stage classification results of each view are fused through weighted average, and the probability of benign tumors in the calculated output value is greater than or equal to the set threshold, the tumor is judged to be benign, so as to predict the risk of malignant transformation of benign tumors in the next few years.
10. A computer storage medium, characterized in that The computer storage medium stores at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the spatially-aware mask-cued multi-view fusion tumor prediction method according to any one of claims 1 to 9.