A breast cancer detection method based on multi-modal fusion and three-stage prediction
This breast cancer detection method, which combines multimodal fusion and three-stage prediction, integrates breast medical images and clinical data. It uses EfficientV2 and ResNet-34 networks for feature extraction and occlusion assessment, and incorporates a temporal multimodal deep prediction model. This approach solves the problems of high computational cost and occlusion impact in existing technologies, achieving efficient and accurate early detection of breast cancer.
Patent Information
- Application Number
- CN202510574044.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-05-06
AI Technical Summary
Existing computer-aided diagnostic methods for breast cancer screening are computationally expensive and easily affected by tissue obscuring, have limited diagnostic strategies, and cannot effectively combine imaging and clinical data for comprehensive evaluation.
A breast cancer detection method based on multimodal fusion and three-stage prediction is adopted. By acquiring breast medical images and structured clinical data, EfficientV2 neural network is used for symptom identification, ResNet-34 network is used for occlusion prediction, and a temporal multimodal deep prediction model is combined for risk prediction to generate breast cancer detection results.
It reduces computational costs, improves the robustness and accuracy of breast cancer detection, reduces the impact of tissue obscuring on diagnosis, and significantly improves the early detection rate of breast cancer, especially in high-risk populations.
Smart Images

Figure CN120511031B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical image analysis and computer-aided diagnosis, in particular to a breast cancer detection method based on multi-modal fusion and three-stage prediction. BACKGROUND
[0002] Breast cancer screening is a method of early detection of breast cancer in asymptomatic population through medical examination, aiming to reduce mortality and improve cure rate. Breast cancer screening mainly relies on mammogram analysis, combined with artificial intelligence to improve early detection accuracy.
[0003] In the prior art, computer-aided diagnosis (CAD) is widely used in breast cancer diagnosis. This method is based on image processing, pattern recognition, machine learning, etc. First, the input medical image is preprocessed, including image enhancement, noise reduction, etc. to improve image quality. Feature extraction algorithm is used to extract key features of the image, such as shape, size, density, etc. of the lesion. Finally, through the classifier or decision model, the extracted features are compared and analyzed with existing medical knowledge and case data to judge the nature and degree of the lesion, and provide diagnosis for doctors. CAD assisted diagnosis strategy is single, has limitations, high computing cost, and is easily affected by tissue shielding, which cannot combine actual image and clinical data to evaluate breast cancer screening, etc.
[0004] Therefore, there is an urgent need for a breast cancer detection method with lower computing cost, more diversified diagnosis strategy and less affected by tissue shielding. SUMMARY
[0005] Therefore, the present application discloses a breast cancer detection method based on multi-modal fusion and three-stage prediction to solve the above problems.
[0006] A breast cancer detection method based on multi-modal fusion and three-stage prediction, comprising: obtaining breast medical image data to be processed and structured clinical data of a patient; using a pre-trained breast cancer detection model based on multi-modal fusion and three-stage prediction to process the breast medical image data to be processed and the structured clinical data of the patient, to obtain a breast cancer detection result.
[0007] Further, the processing of the breast cancer detection model based on multi-modal fusion and three-stage prediction includes:
[0008] S1, pre-processing the breast medical image data to be processed and the structured clinical data of the patient to obtain an enhanced breast image and a structured clinical data vector;
[0009] S2, adopt the improved EfficientV2 neural network model to carry out symptom recognition to the enhanced breast image, obtain the probability density heat map of the lesion area, and calculate the tumor risk score according to the probability density heat map;
[0010] S3, adopt the ResNet-34 network based on adaptive optimization to carry out shielding prediction on the enhanced breast image, and obtain a shielding degree score;
[0011] S4, adopt a time series multi-modal deep prediction model to splice the tumor risk score, the shielding degree score and the structured clinical data vector, obtain a historical screening record, and model and predict the historical screening record based on the Transformer architecture to obtain a predicted risk probability;
[0012] S5, generate a breast cancer detection result according to the predicted risk probability.
[0013] The beneficial effects of the present application include:
[0014] By adopting a multi-element diagnosis strategy to process breast medical image visual features and structured clinical data, the limitations of a single data source are made up;
[0015] The ResNet-34 network based on adaptive optimization is adopted to evaluate the image shielding degree, the robustness is improved by carrying out shielding prediction, the influence of interference factors on diagnosis is reduced, so that the method proposed in the present application is not easily affected by tissue shielding;
[0016] By introducing the EfficientV2 network with high efficiency and light weight for feature extraction, and adaptively optimizing the ResNet-34 network to realize shielding degree evaluation, combined with the spatial attention mechanism and the local density region perception module, efficient feature modeling and resource optimization in the feature extraction and shielding prediction process are realized, and the calculation cost is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 It is a flowchart of the process of the breast cancer detection model based on multi-modal fusion and three-stage prediction in the present application for processing data;
[0018] Figure 2 It is a schematic diagram of iterative pre-training in the embodiment of the present application;
[0019] Figure 3 It is a flowchart of the process of the breast cancer detection method based on multi-modal fusion and three-stage prediction in the embodiment of the present application for processing CBIS data set;
[0020] Figure 4 It is a ROC curve diagram obtained by testing the breast cancer detection model based on multi-modal fusion and three-stage prediction in the embodiment of the present application;
[0021] Figure 5 The detection rate result chart obtained by testing the clinical data by the breast cancer detection model based on multi-modal fusion and three-stage prediction in the embodiments of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical scheme, characteristics and advantages of the present application more clear, the present application is further described below in combination with the drawings and embodiments.
[0023] The present embodiment includes a breast cancer detection method based on multi-modal fusion and three-stage prediction, comprising: obtaining breast medical image data to be processed and structured clinical data of a patient; using a pre-trained breast cancer detection model based on multi-modal fusion and three-stage prediction to process the breast medical image data to be processed and the structured clinical data of the patient, to obtain a breast cancer detection result.
[0024] Specifically, as shown in Figure 1 The processing of data by the breast cancer detection model based on multi-modal fusion and three-stage prediction includes:
[0025] S1, preprocessing the breast medical image data to be processed and the structured clinical data of the patient to obtain an enhanced breast image and a structured clinical data vector.
[0026] Specifically, the preprocessing includes uniform size and gray scale normalization processing. The structured clinical data includes patient age, family medical history, hormone level, breast density grade, etc.
[0027] Let the breast medical image be where H and W are the height and width of the image respectively. The image is uniformly adjusted to 1024x1024 size and gray scale normalized, and the formula is:
[0028]
[0029] where, represents the enhanced breast image, μ I and σ I respectively represent the image pixel mean and standard deviation, and gray scale normalization is used to enhance image contrast and reduce the influence of light changes.
[0030] The structured clinical data vector where each element c i represents a standardized clinical variable, such as age, breast density grade, estrogen level, etc. The categorical variables are one-hot encoded, and the numerical variables are normalized to the [0,1] interval.
[0031] S2, adopt improved EfficientV2 neural network model to enhanced breast image carries out symptom recognition, obtains the probability density heat map of lesion area, and calculates tumor risk score according to probability density heat map.
[0032] The method designed in the application introduces attention mechanism in the conventional EfficientV2 neural network to improve the sensitivity to the mass area, and adopts lightweight design to maintain the calculation efficiency.
[0033] Specifically, generating the tumor risk score includes:
[0034] Step 1, adopt standard EfficientV2 neural network, extract preliminary features from enhanced breast image, the formula is:
[0035]
[0036] Wherein, F s represents the preliminary feature, and EfficientV2() represents the standard EfficientV2 neural network.
[0037] Step 2, adopt SE attention mechanism, and adopt channel weighting to re-label the preliminary feature, obtain attention re-labeling feature, the formula is:
[0038] s c =σ(W2·ReLU(W1·GAP(F s )))
[0039]
[0040] Wherein, s c represents SE attention weight, σ represents Sigmoid activation function, ReLU() represents ReLU activation function, W2 and W1 represent two different weight matrices, GAP() represents average pooling, represents attention re-labeling feature.
[0041] Step 3, adopt two-dimensional convolution, calculate probability density heat map according to attention re-labeling feature, calculate tumor risk score according to probability density heat map, the formula is:
[0042]
[0043] Wherein, S tumor represents tumor risk score, the range is [0, 1], H represents the probability density heat map of lesion area, H i,j represents the i,j element in H, and w and h represent the column number and row number of H respectively.
[0044] S3, adopt ResNet-34 network based on adaptive optimization to carry out shielding prediction on the enhanced breast image, and obtain a shielding degree score.
[0045] Specifically, since the shielding risk is dominated by breast density, in order to simulate the shielding degree of the image region, the application adopts ResNet-34 network based on adaptive optimization. The input of the ResNet-34 network based on adaptive optimization is a standard breast image, and the network structure introduces a spatial attention mechanism and a local density region perception mechanism, which are used to judge the potential shielding region in the image caused by breast density, and the output is a shielding degree quantitative score, which is used as the input of subsequent multi-modal risk prediction. The shielding degree score includes:
[0046] Step 1, adopt a standard ResNet-34 network to extract basic features from the enhanced breast image, and the formula is:
[0047]
[0048] Wherein, F m represents the basic feature, and ResNet34() represents the standard ResNet-34 network.
[0049] Step 2, calculate the spatial attention mask according to the basic feature, associate the spatial attention mask to the basic feature, and obtain the spatial attention feature, and the formula is:
[0050]
[0051] Wherein, M s represents the spatial attention mask, Conv 7×7 () represents a 7x7 size convolution layer, AvgPool() represents an average pooling, MaxPool() represents a maximum pooling, represents the spatial attention feature.
[0052] Step 3, flatten the spatial attention feature, and adopt a full connection layer to process the flattened spatial attention feature, and obtain a shielding degree score:
[0053]
[0054] Wherein, S mask represents the shielding degree score, and the higher the score, the greater the risk of shielding the lesion region, Flatten() represents the flattening operation, W m represents the weight matrix of the full connection layer, and b m represents the bias vector of the full connection layer.
[0055] S4. Using a time-series multimodal deep prediction model, tumor risk scores, occlusion scores, and structured clinical data vectors are concatenated to obtain historical screening records. Based on the Transformer architecture, the historical screening records are modeled and predicted to obtain the predicted risk probability.
[0056] Specifically, the temporal multimodal deep prediction model employs a Transformer architecture to capture the temporal correlation between historical screening records and the tumor risk score and occlusion score of the current screening. The temporal multimodal deep prediction model can adaptively weight features and generate a final risk probability as the basis for screening decisions. The predicted risk probabilities include:
[0057] Step 1: Integrate tumor risk scores, masking scores, and structured clinical data to construct historical screening records. The formula for historical screening records is:
[0058] X = [S] tumor ,S mask ,c1,...,c k ]
[0059] Where X represents historical screening records, S tumor S represents the tumor risk score. mask The occlusion level is represented by scores c1,...,c k This represents structured clinical data.
[0060] Step 2: Encode the historical screening records using Transformer to obtain the time-series input sequence, as shown in the formula:
[0061] Z t =TransformerEncoder(X)
[0062]
[0063] Among them, Z t Z represents the t-th time element in the time-series input sequence, TransformerEncoder() represents the Transformer encoder, Z represents the time-series input sequence, and T represents the number of time elements in the time-series input sequence.
[0064] Step 3: Calculate the cancer risk probability based on the time-series input sequence. The formula is:
[0065] P cancer =σ(W r ·Z+b r )
[0066] Among them, P cancer W represents the probability of developing cancer. rdenotes a weight matrix, b r denotes a bias vector.
[0067] S5, generating a breast cancer detection result according to the predicted risk probability.
[0068] Specifically, a cancer probability threshold is set, when the cancer risk probability is greater than or equal to the cancer probability threshold, it is determined as a breast cancer positive detection result; when the cancer risk probability is less than the threshold, it is determined as a breast cancer negative detection result. The cancer probability threshold is preferably 0.5.
[0069] Further, the pre-training adopts a multi-task joint optimization strategy, which includes a lesion recognition loss (cross entropy), a masking prediction loss (mean square error), and a risk prediction loss (binary cross entropy). During the pre-training process, a multi-scale image enhancement strategy, a Dropout mechanism, and L2 regularization are introduced to avoid overfitting, and an AdamW optimizer and a cosine annealing strategy are used for optimization.
[0070] Specifically, the total loss function used in the pre-training is as follows:
[0071]
[0072] wherein, denote the lesion recognition loss, the masking prediction loss, and the risk prediction loss, respectively, and λ1, λ2, and λ3 denote the weight coefficients of the corresponding loss functions, and the preferred values are 1, 0.5, and 1, respectively.
[0073] Further, the formula of the lesion recognition loss function is as follows:
[0074]
[0075] wherein, y represents the true lesion label of the current sample, y = 1 represents the actual existence of breast lesions, and y = 0 represents the actual absence of breast lesions, i.e., normal state.
[0076] Further, the formula of the masking prediction loss function is as follows:
[0077]
[0078] wherein, m * represents the true value of the masking degree, which is derived from the breast density level labeled by doctors.
[0079] Further, the formula of the risk prediction loss function is as follows:
[0080]
[0081] As Figure 2 shown, the breast cancer detection model designed based on multi-modal fusion and three-stage prediction in the application adopts a loss function for iterative pre-training to obtain optimal parameters parameters, and the optimal parameters are set as initial parameters of the model after pre-training is completed.
[0082] As Figure 3 shown, the experimental data of the embodiment selects public breast X-ray image data sets INbreast and CBIS-DDS, and combines a real screening clinical data set (including review records and clinical labels), divides the samples into a training set and a test set in a ratio of 8:2, and adopts AUC, accuracy (ACC), sensitivity (SEN) and specificity (SPE) as evaluation indexes.
[0083] Specifically, as Figure 4 shown, the ROC curve of the breast cancer detection model based on multi-modal fusion and three-stage prediction is a Receiver Operating Characteristic Curve, which is used to evaluate the performance of a binary classification model. As can be seen from the figure, the breast cancer detection model based on multi-modal fusion and three-stage prediction proposed in the application corresponds to the purple curve, and the accuracy AUC is 0.72, which is the best among all the compared models, and the overall ROC curve is significantly better than the traditional breast density evaluation method.
[0084] Figure 4 The comparison methods in the table include: a prediction model based on percentage breast density (Percent density, AUC=0.51), a prediction model based on age-adjusted percentage density (Age-adjusted percent density, AUC=0.60), a prediction model based on dense area (Dense area, AUC=0.53), and a prediction model based on age-adjusted dense area (Age-adjusted dense area, AUC=0.61). As can be seen, the AUC of the traditional breast density related model is less than 0.62, because the traditional method only relies on breast density features, and it is difficult to effectively identify potential high-risk individuals. While maintaining a low false positive rate (FPR<0.2), the detection model proposed in the application can still maintain a high sensitivity, showing significantly better screening performance than the traditional method. The above experimental results further verify the effectiveness of the multi-modal fusion and three-stage prediction strategy of the application, which can improve the detection rate of early breast cancer screening, reduce the risk of missed detection, and help optimize the individualized screening process.
[0085] As Figure 5 shown, the breast cancer detection method based on multi-modal fusion and three-stage prediction can significantly improve the detection rate of breast cancer in clinical application. Specifically, among the 559 subjects who received supplementary MRI screening, a total of 36 cases of breast cancer were detected, with an overall cancer detection rate of 64.4 cases per 1000 MRI examinations, and a 95% confidence interval (CI) of 46.8-88.1, significantly better than traditional screening methods. Further combined with the detection results of different breast density levels of BI-RADS 3-5, in the BI-RADS 5 (highly suspected malignant) population, 12 cases of breast cancer were detected in 14 cases of examination, with a cancer detection rate of 85.7%, and the positive predictive values PPV1 and PPV3 were both 85.7% (95% CI: 53.5-96.9); in the BI-RADS 4 (suspiciously malignant) population, 17 cases of breast cancer were detected in 27 cases of examination, with a detection rate of 63.0% (95% CI: 42.8-79.4), PPV1 was 63.0%, and PPV3 was 65.4% (95% CI: 44.7-81.5); while in the BI-RADS 3 (possibly benign) population, 7 cases of cancer were detected in 54 cases of examination, with a detection rate of 13.0% (95% CI: 6.2-25.1), PPV1 was 13.0%, and PPV3 was 22.6% (95% CI: 10.8-41.2). At the same time, among the 464 cases without breast cancer detection, only 35 cases received benign biopsy, and 24 cases were temporarily unable to be finally confirmed due to incomplete MRI follow-up, indicating that the present application not only improves the detection rate but also effectively controls the benign biopsy rate, avoiding over-screening and over-intervention. Overall, the present application accurately selects the truly high-risk individuals through intelligent multi-modal fusion analysis, guides the supplementary MRI screening decision, and significantly improves the early detection rate of breast cancer, especially in the BI-RADS 4-5 high-risk population, which is particularly prominent. The present application has important clinical application value and promotion prospect in optimizing the screening path, reducing the missed detection rate, and improving the screening efficiency.
[0086] Finally, it should be noted that the above description only describes some embodiments of the present application, and those skilled in the art can make various changes, modifications, replacements and variations to these embodiments without departing from the principles and spirits of the present application. The protection scope of the present application is defined by the appended claims and their equivalents, and the above-mentioned behaviors should be covered within the protection scope of the present application.
Claims
1. A breast cancer detection method based on multimodal fusion and three-stage prediction, characterized in that, include: Acquire breast medical image data to be processed and patient structured clinical data; A pre-trained breast cancer detection model based on multimodal fusion and three-stage prediction was used to process breast medical image data and patient structured clinical data to obtain breast cancer detection results. The data processing for breast cancer detection models based on multimodal fusion and three-stage prediction includes: S1. Preprocess the breast medical image data to be processed and the patient's structured clinical data to obtain enhanced breast image and structured clinical data vectors; S2. An improved EfficientV2 neural network model is used to identify symptoms in enhanced breast images, obtaining a probability density heatmap of the lesion region, and calculating a tumor risk score based on the probability density heatmap; including: S21. Using the standard EfficientV2 neural network, preliminary features are extracted from the enhanced breast images; S22. Employ the SE attention mechanism and use channel weighting to recalibrate the initial features to obtain attention-recalibrated features; S23. Using two-dimensional convolution, a probability density heatmap is calculated based on the attention recalibration features, and a tumor risk score is calculated based on the probability density heatmap. S3. An adaptively optimized ResNet-34 network is used to predict occlusion in enhanced breast images, and an occlusion level score is obtained; including: S31. Use the standard ResNet-34 network to extract basic features from enhanced breast images; S32. Calculate the spatial attention mask based on the basic features, and associate the spatial attention mask with the basic features to obtain the spatial attention features. S33. Flatten the spatial attention features and process the flattened spatial attention features using a fully connected layer to obtain the occlusion score; S4. Using a time-series multimodal deep prediction model, tumor risk score, occlusion score and structured clinical data vector are concatenated to obtain historical screening records. Based on the Transformer architecture, the historical screening records are modeled and predicted to obtain the predicted risk probability. S5. Generate breast cancer detection results based on predicted risk probability.
2. The breast cancer detection method based on multimodal fusion and three-stage prediction according to claim 1, characterized in that, Preprocessing of the breast medical image data to be processed includes: adjusting the size of the breast medical image to 1024×1024 and performing grayscale normalization to obtain an enhanced breast image.
3. The breast cancer detection method based on multimodal fusion and three-stage prediction according to claim 1, characterized in that, The formula used for recalibrating the preliminary features is: in, Indicates the SE attention weight. This represents the Sigmoid activation function. Represents the ReLU activation function. and This represents two different weight matrices. Indicates average pooling. This indicates attention recalibration features.
4. The breast cancer detection method based on multimodal fusion and three-stage prediction according to claim 1, characterized in that, The formula used to calculate the tumor risk score is: in, Represents a probability density heatmap. express The first in One element, Indicates that the convolution kernel is Two-dimensional convolution, This indicates attention recalibration features. Indicates tumor risk score, and They represent The number of columns and rows.
5. The breast cancer detection method based on multimodal fusion and three-stage prediction according to claim 1, characterized in that, The formula used to obtain spatial attention features is: in, Represents a spatial attention mask. This represents the Sigmoid activation function. Indicates basic features, express Convolutional layers of varying sizes, Indicates average pooling. This indicates max pooling. This represents spatial attention characteristics.
6. The breast cancer detection method based on multimodal fusion and three-stage prediction according to claim 1, characterized in that, The predicted risk probabilities include: Step 1: Combine and integrate tumor risk scores, masking scores, and structured clinical data to construct historical screening records; Step 2: Encode the historical screening records using Transformer to obtain the time-series input sequence; Step 3: Calculate the probability of cancer risk based on the time-series input sequence.
7. The breast cancer detection method based on multimodal fusion and three-stage prediction according to claim 1, characterized in that, Pre-training employs a multi-task joint optimization strategy, which includes lesion identification loss, occlusion prediction loss, and risk prediction loss. The loss function formula used is as follows: in, Represents the loss function. This represents the lesion identification loss function. This represents the occlusion prediction loss function. This represents the risk prediction loss function. , representing the weighting coefficients of lesion identification loss, occlusion prediction loss, and risk prediction loss, respectively, with values of 1, 0.5, and 1. Indicates tumor risk score, The rating indicates the degree of occlusion. Indicates the predicted probability of risk. This represents the actual lesion label of the current sample, where y=1 indicates the actual presence of breast lesions, and y=0 indicates the actual absence of breast lesions. The actual value representing the degree of occlusion.
Citation Information
Patent Citations
Gynecological tumor image processing method and system based on AI multi-modal image analysis
CN120747029A