Breast cancer prognosis prediction method based on enhanced multi-modal feature fusion

By adopting an enhanced multimodal feature fusion method in breast cancer prognosis analysis, combining the Vision-LSTM model and the multimodal feature fusion module BASMF, the shortcomings of multimodal data integration and feature capture in the existing technology are solved, and a higher accuracy and stable breast cancer prognosis analysis is achieved.

CN120221105APending Publication Date: 2025-06-27HUZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201803.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate multimodal data to capture the heterogeneity of multimodal characteristics of breast cancer, resulting in insufficient accuracy and stability of breast cancer prognosis analysis.

Method used

Using an enhanced multimodal feature fusion method, histopathological image features were extracted through Vision-LSTM pre-trained model, and combined with deep feature extraction and dimensional matching, clinical data, gene expression data and DNA methylation data were integrated. Use the multimodal feature fusion module BASMF, a bidirectional attention block and self-attention block stack, to capture the cross information between multimodal features.

Benefits of technology

A more accurate and stable breast cancer prognosis analysis is achieved, which can effectively integrate multimodal data, capture the multimodal characteristic heterogeneity of breast cancer, and improve the accuracy and interpretability of prognostic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120221105A_ABST
    Figure CN120221105A_ABST
Patent Text Reader

Abstract

The invention provides a breast cancer prognosis prediction method based on enhanced multi-modal feature fusion. The breast cancer prognosis prediction method comprises the following steps: firstly, acquiring a histopathological image, clinical pathological parameters, gene expression data and DNA methylation data of a breast cancer patient; then preprocessing the tissue pathology image and the clinical pathology parameters, and performing feature screening on gene expression data and DNA methylation data; image features, clinical features, gene expression features and DNA methylation features are obtained through the steps, a multi-modal feature fusion module BASMF is created, the multi-modal features are sent to the BASMF for feature fusion, and comprehensive prognosis features are obtained for prognosis analysis. According to the scheme, the problems that the training effect is poor due to the fact that breast cancer samples are few and the resolution of histopathological images is too large, the model is over-fitted and unstable due to the high-dimensional characteristics of data, and the heterogeneity of breast cancer multi-modal characteristics is difficult to capture due to insufficient simple fusion are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical diagnosis, and particularly to a breast cancer prognosis prediction method based on enhanced multi-modal feature fusion. Background Art

[0002] Breast cancer is one of the malignant tumors with the highest incidence and mortality rates among women globally. Its high heterogeneity makes it difficult for traditional single-modal prognosis analysis methods (such as clinical data, gene expression data, etc.) to comprehensively capture tumor characteristics, and the prediction accuracy is limited. In the prior art, multi-modal data fusion methods mostly adopt simple connection or addition strategies, which cannot effectively extract cross-modal cross information and lack a feature screening mechanism for high-dimensional biological data. For example, existing methods such as ICSDA (fusing multi-modal features by connection), MGCN-CalRF (graph convolutional network combined with random forest), etc., have deficiencies in feature fusion effect and model generalization ability. In addition, the high-resolution characteristics of tissue pathological images lead to difficulties in traditional model training and low data utilization rate. Therefore, there is an urgent need for a breast cancer prognosis analysis method that can integrate multi-modal data, efficiently screen features, and capture cross-modal associations. Summary of the Invention

[0003] Aiming at the deficiencies in the prior art, the present invention provides a breast cancer prognosis prediction method based on enhanced multi-modal feature fusion, which can effectively solve the problems of poor training effect caused by few breast cancer samples and too large resolution of tissue pathological images, overfitting and instability of the model caused by the high-dimensional characteristics of data, and heterogeneity of multi-modal features of breast cancer that is difficult to capture due to insufficient simple fusion.

[0004] The above technical object of the present invention is achieved through the following technical solutions:

[0005] A breast cancer prognosis prediction method based on enhanced multi-modal feature fusion includes the following steps:

[0006] S1, obtaining tissue pathological images, clinicopathological parameters, gene expression data, and DNA methylation data of breast cancer;

[0007] S2, preprocessing the tissue pathological images, dividing them into several small images and saving the coordinate values and image feature data, then sending each small image into the Vision-LSTM pre-trained model for feature extraction, and obtaining a multi-modal feature after deep feature extraction and dimension matching;

[0008] S3, preprocessing the clinicopathological parameters, including performing maximum-minimum normalization on the age data, mapping the tumor stage, T stage, M stage, N stage, ER hormone receptor, and PR hormone receptor, and obtaining a multi-modal feature after deep feature extraction and dimension matching;

[0009] S4. Perform feature screening on the gene expression data. First, use a support vector machine to perform preliminary feature screening on the gene expression data, and then further screen the gene expression data through Mantel test and correlation analysis. Finally, determine the gene expression data features related to breast cancer prognosis, and then obtain a multi-modal feature through deep feature extraction and dimension matching;

[0010] S5. Perform feature screening on the DNA methylation data. First, use a support vector machine to perform preliminary feature screening on the DNA methylation data, and then further screen the gene expression data through Mantel test and correlation analysis. Finally, determine the DNA methylation data features related to breast cancer prognosis, and then obtain a multi-modal feature through deep feature extraction and dimension matching;

[0011] S6. Form four multi-modal features through steps S2 - S5: image features, clinical features, gene expression features, and DNA methylation features, and create a multi-modal feature fusion module BASMF. Send the multi-modal features into BASMF for feature fusion to obtain comprehensive prognostic features for prognostic analysis.

[0012] Furthermore, in step S2, when preprocessing the histopathological image, automatically identify and segment the tissue area through the CLAM library, divide it into several small images of 224×224, save the coordinate values and image feature data, send each small image into the Vision-LSTM pre-trained model for feature extraction, then divide the small image into n patch blocks of a fixed size, and divide them into non-overlapping patches through shared linear projection. The feature vector of each patch is X i ∈R m , where m is the feature dimension, and add a learnable position embedding vector E i ∈R n×d to each patch token, where d is the embedding dimension. The patch feature vector after adding the position embedding is Y i = X i + E i , and then send the patch feature vector Y i (i = 1, 2, …, n) into the encoder and linear projection to obtain a series of patch token sequences, and then send them into the alternating mLSTM blocks for processing. After processing, obtain the complete small image features, and connect all the small image features according to the coordinates to obtain the overall image features of the tissue area.

[0013] Furthermore, in steps S4 and S5, when using a support vector machine to perform preliminary feature screening on the gene expression data or DNA methylation data, assign weights to each feature, rank the importance based on the weights, and then identify and remove the features that are not related to breast cancer prognosis.

[0014] Furthermore, in step S6, the multimodal feature fusion module BASMF is stacked by a bidirectional attention block and two self-attention blocks; the bidirectional attention block is used to learn the multimodal joint feature representation by capturing the independent information within the features and the cross information between different features, and the self-attention block is used to learn and enhance the multimodal joint feature representation.

[0015] Furthermore, in step S6, the image features are taken as a part and input into N I , and the clinical features, gene expression features, and DNA methylation features are combined to form another part and input into M O ;

[0016] For the bidirectional attention block, first, the image features N I and the combined features N O are input to generate their respective Query matrices Q I and Q O , Key matrices K I and K O , and Value matrices V I and V O . The process of generating the matrices is as follows:

[0017] Q I , K I , V I = LP(Norm(N I ))(1)

[0018] Q0, K0, V0 = LP(Norm(N0))(2)

[0019] where Norm(·) and LP(·) represent normalization and linear projection respectively. The internal forward transmission process of the bidirectional attention mechanism is as follows:

[0020]

[0021] where Attention(Q I , K I , V O ) and Attention(Q O , K O , V O ) capture the independent information within the image features and the combined features respectively, and Attention(Q I , K O , V O ) takes the image features as the main information and captures the cross information between the image features and the combined features. Attention(Q O , KI , V I ) The main information is the combined features, capturing the cross - information between the combined features and the image features. λ is a hyperparameter used to balance the ratio between the independent information and the cross - information. The definition of the attention mechanism Attention(Q, K, V) is as follows:

[0022]

[0023] Among them, Softmax(·) is an activation function that can normalize a numerical vector into a probability distribution vector, and the sum of each probability is 1. T is the matrix transpose, and d k is the scaling factor;

[0024] The overall process of the bidirectional attention block is as follows:

[0025]

[0026] Among them, MLP(·) is a multi - layer perceptron, consisting of an input layer, one or more hidden layers, and an output layer. Mean(·) represents the average value, used to balance the two obtained multi - modal joint feature representations. The two multi - modal joint feature representations extracted after passing through the bidirectional attention block and are processed for balancing to obtain the overall multi - modal joint feature representation Then it is sent to the self - attention block to learn and enhance its feature representation. The overall process of the self - attention block is as follows:

[0027]

[0028] Among them, is the multi - modal feature representation after self - attention, and N is the final multi - modal joint feature representation.

[0029] The present invention has the following beneficial effects:

[0030] The present application provides a breast cancer prognosis prediction method based on enhanced multi-modal feature fusion, constructs a new comprehensive and all-round breast cancer prognosis prediction model, and is accurately and effectively used for breast cancer prognosis analysis. The model of this solution can integrate multi-modal data including histopathological images, clinical data, gene expression data, and DNA methylation data for breast cancer prognosis analysis; and introduces the Vision-LSTM pre-trained model to better extract features from histopathological images; a comprehensive feature screening strategy including support vector machine, Mantel test, and correlation analysis is designed to screen features from gene expression data and DNA methylation data, making the data more relevant to breast cancer prognosis, and optimizing the model in terms of analysis results and interpretability; and this solution fully considers the high heterogeneity of breast cancer and the independence and cross-interference of multi-modal features, and proposes a multi-modal feature fusion module BASMF. This module is stacked by bidirectional attention blocks and self-attention blocks, so that BASMF not only retains the independent information of multi-modal data, but also fully captures the cross-information between multi-modal features, enabling the fused comprehensive prognosis features to perform prognosis analysis more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is the PCGMMF structure diagram of an embodiment of the present invention; in the figure, A is a histopathological image, B is a divided tissue region, C is clinical data, D is gene expression data, E is DNA methylation data, F is feature extraction through the Vision-LSTM pre-trained model, G is a feature screening strategy, H is the multi-modal feature fusion module BASMF, and I is model evaluation and visualization.

[0032] Figure 2 It is the pre-processed histopathological image of an embodiment of the present invention.

[0033] Figure 3 It is the Vision-LSTM structure diagram of an embodiment of the present invention.

[0034] Figure 4 It is the BASMF structure diagram of an embodiment of the present invention; in the figure, A is the overall structure diagram of the multi-modal feature fusion module, B is the bidirectional attention block, C is the bidirectional attention mechanism, and D is the self-attention block.

[0035] Figure 5 It is the feature screening analysis diagram of an embodiment of the present invention; in the figure, A is the top 20 important features of gene expression, B is the Mantel test and correlation analysis results of gene expression data, C is the top 20 important features of 27 groups of DNA methylation, D is the Mantel test and correlation analysis results of 27 groups of DNA methylation, E is the top 20 important features of 450 groups of DNA methylation, and F is the Mantel test and correlation analysis results of 450 groups of DNA methylation.

[0036] Figure 6 This is the 5-fold cross-validation ROC graph in the embodiment of the present invention.

[0037] Figure 7 This is the ablation experiment result graph in the embodiment of the present invention.

[0038] Figure 8 This is the schematic diagram of the GO enrichment analysis in the experimental process in the embodiment of the present invention; in the figure, A is the biological process enrichment network graph, B is the biological process enrichment bar graph, C is the cellular component enrichment network graph, D is the cellular component enrichment bar graph, E is the molecular function enrichment network graph, and F is the molecular function enrichment bar graph.

[0039] Figure 9 This is the KM curve graph of the gene expression data in the embodiment of the present invention; in the figure, A is the Kaplan-Meier curves of all screened genes, B is the Kaplan-Meier curves of the USP4 gene, C is the Kaplan-Meier curves of the PNRC2 gene, D is the Kaplan-Meier curves of the SLC23A1 gene, E is the Kaplan-Meier curves of the EIF4E3 gene, F is the Kaplan-Meier curves of the SNX4 gene, G is the Kaplan-Meier curves of the PARP9 gene, H is the Kaplan-Meier curves of the WDR82 gene, and I is the Kaplan-Meier curves of the SLC1A4 gene.

[0040] Figure 10 This is the attention heat map in the interpretability analysis in the embodiment of the present invention; in the figure, A is the original histopathological image, B is the histopathological image heat map, and C is the high-importance area. Detailed implementation manners

[0041] The technical solutions in the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0042] The embodiment of the present invention provides a breast cancer prognosis prediction method based on enhanced multi-modal feature fusion, namely the PCGMMF method, and its overall framework is as Figure 1 shown, including the following steps:

[0043] S1, obtain the histopathological image, clinicopathological parameters, gene expression data, and DNA methylation data of breast cancer;

[0044] S2. Preprocess the histopathological image, divide it into several small images, save the coordinate values and image feature data, and then send each small image into the Vision-LSTM pre-trained model for feature extraction. After deep feature extraction and dimension matching, a multi-modal feature is obtained.

[0045] S3. Preprocess the clinicopathological parameters, including performing max-min normalization on the age data, mapping the tumor stage, T stage, M stage, N stage, ER hormone receptor, and PR hormone receptor. After deep feature extraction and dimension matching, a multi-modal feature is obtained.

[0046] S4. Screen the features of the gene expression data. First, use a support vector machine to perform preliminary feature screening on the gene expression data, and then further screen the gene expression data through Mantel test and correlation analysis. Finally, determine the gene expression data features related to breast cancer prognosis. After deep feature extraction and dimension matching, a multi-modal feature is obtained.

[0047] S5. Screen the features of the DNA methylation data. First, use a support vector machine to perform preliminary feature screening on the DNA methylation data, and then further screen the gene expression data through Mantel test and correlation analysis. Finally, determine the DNA methylation data features related to breast cancer prognosis. After deep feature extraction and dimension matching, a multi-modal feature is obtained.

[0048] S6. Through steps S2 - S5, four multi-modal features are formed: image feature, clinical feature, gene expression feature, and DNA methylation feature. Create a multi-modal feature fusion module BASMF, send the multi-modal features into BASMF for feature fusion, and obtain comprehensive and integrated prognosis features for prognosis analysis.

[0049] Regarding the histopathological image, in this embodiment, the hematoxylin and eosin-stained histopathological images are downloaded from the TCGA database (https: / / portal.gdc.cancer.gov). The histopathological images are preprocessed through the CLAM library, the tissue regions are automatically identified and segmented, and divided into several small images of 224×224, and the coordinate values and image feature data are saved. The preprocessed histopathological images in this embodiment are shown in Figure 2 . Regarding the CLAM library, refer to: Lu MY , Williamson DFK ,Chen TY, Chen RJ, Barbieri M, Mahmood F. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat Biomed Eng. 2021;5(6):555-570. doi:10.1038 / s41551-020-00682-w. Then each small image is fed into the Vision-LSTM pre-trained model for feature extraction. Vision-LSTM is a combination of Vision Transformer and the long short-term memory network LSTM, which is commonly used to process high-resolution images and analysis. For Vision Transformer, see Dosovitskiy A, Beyer L, Kolesnikov A, et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale[J]. 2020. DOI: 10.48550 / arXiv.2010.11929., and the pre-trained model architecture of Vision-LSTM is as Figure 3 shown. The small image is segmented into n non-overlapping patches of a fixed size, and divided into non-overlapping patches through a shared linear projection. The feature vector of each patch is X i ∈R m , where m is the feature dimension, and a learnable position embedding vector E i ∈R n×d is added to each patch token, where d is the embedding dimension. The patch feature vector after adding the position embedding is Y i = X i + E i . Then the patch feature vector Y i (i = 1, 2,..., n) is fed into the encoder and linear projection to obtain a series of patch token sequences, and then fed into the alternating mLSTM blocks for processing. The core of Vision-LSTM is the alternating mLSTM blocks, which are equipped with a memory storage and a covariance update rule. In the L mLSTM blocks, the odd-numbered blocks process the patch token sequence row by row from the upper left to the lower right, and the even-numbered blocks process the patch token sequence column by column from the lower right to the upper left. After processing, the complete small image features are obtained, and the image features of the overall tissue area are obtained by concatenating all the small image features according to the coordinates.

[0050] Regarding clinicopathological parameters, in the prognosis analysis of breast cancer, clinicopathological parameters are still the most commonly used and widely recognized prognostic factors for breast cancer. In this protocol, the age data is normalized by max-min normalization, and custom mappings are performed for tumor stage, T stage, M stage, N stage, ER hormone receptor, and PR hormone receptor. This protocol performs custom mappings on several data items including cancer digital stage, T stage, N stage, M stage, ER hormone receptor, and PR hormone receptor, digitizing these data. Table 1 below shows the specific clinical data distribution and the definition of custom mappings for all samples.

[0051] Table 1 Clinical data distribution and custom mapping

[0052]

[0053] Regarding gene expression data, gene expression data demonstrates powerful capabilities for breast cancer prognosis analysis and can serve as a strong prognostic factor in breast cancer prognosis analysis. However, the high-dimensional nature of gene expression data results in a large number of irrelevant and redundant features in the high-dimensional space, reducing the performance and stability of the method. An appropriate feature screening strategy can effectively reduce the impact of irrelevant and redundant features, while also reducing model parameters and the consumption of training resources. Finally, it can prevent model overfitting and improve interpretability. To this end, this protocol first uses support vector machines to perform preliminary feature screening on gene expression data, ranks the importance based on weights, and then determines the final gene features through Mantel test and correlation analysis.

[0054] Regarding DNA methylation data, as an emerging and powerful biomarker, DNA methylation data shows great promise in breast cancer prognosis analysis, can serve as an independent predictor, and can be used to construct an analysis model. The DNA methylation data used in this example is the β value of CpG sites, and its magnitude is related to the methylation level of the sites. Similar to gene expression data, this protocol also performs feature screening on DNA methylation data. In addition, due to the differences between HumanMethylation 27 and Human Methylation 450 in the acquired data, this protocol divides the samples into DNA methylation 27 group and 450 group respectively for feature screening.

[0055] Both gene expression data and DNA methylation data have high-dimensional characteristics. When conducting in-depth analysis, feature screening is required. First, this solution selects a support vector machine model with powerful high-dimensional data processing capabilities to obtain feature importance and rank them. The basic idea of the support vector machine is to find a hyperplane in the feature space that can separate different samples to the greatest extent. For non-linear problems, the support vector machine maps the data to a high-dimensional space by introducing a kernel function (in this paper, the radial basis function (RBF) is used as the kernel function), thereby constructing a linear hyperplane in the high-dimensional space. During the training process of the support vector machine, this solution assigns weights to each feature, and these weights reflect the contribution degree of the feature to the target variable. After model training, the importance ranking of the features is obtained, and then the features with lower importance are identified and removed to reduce the complexity of the data and improve the correlation between the features and breast cancer prognosis. Subsequently, in order to further refine the selected relevant features and ensure the biological significance and statistical significance of the selected features, this solution introduces the Mantel test and correlation analysis as auxiliary screening tools. The Mantel test is based on the idea of linear regression. By calculating the similarity measure between two matrices and performing a significance test to determine whether the two matrices are related. In addition, this application also conducts a correlation analysis between the screened features. After further screening, the features most relevant to breast cancer prognosis are finally determined, providing important feature information for the subsequent prognosis analysis model.

[0056] For the task of predicting the risk of recurrence and metastasis in breast cancer prognosis, sample selection is carried out. In this embodiment, a total of 207 samples are finally selected, and the histopathological images, clinical data, gene expression data, and DNA methylation data of all samples are downloaded. Table 2 below shows the prognostic status of the samples and the quantities of gene expression data and DNA methylation data before and after preprocessing.

[0057] Table 2 Information on sample data selected from breast cancer patients in the TCGA database

[0058]

[0059] Regarding the multi-modal feature fusion module BASMF, as Figure 4 shown, the multi-modal feature fusion module BASMF is stacked by a bidirectional attention block and two self-attention blocks; the bidirectional attention block is used to learn the multi-modal joint feature representation by capturing the independent information within the features and the cross information between different features, and the self-attention block is used to learn and enhance the multi-modal joint feature representation.

[0060] The multi-modal data input in this solution includes four parts: histopathological images, clinical data, gene expression data, and DNA methylation data, which are respectively represented as X image ,X clinical ,Xgene ,X methylation . After deep feature extraction and dimension matching, image features are obtained Clinical features Gene expression features DNA methylation features Finally, the image features are used as a part to input into N I , and the other three features are combined to form another part to input into N O .

[0061] The first layer of the multi-modal feature fusion module is a bidirectional attention block. For the bidirectional attention block, first input the image feature N I and the combined feature N O to generate their respective Query matrices Q I and Q O , Key matrices K I and K O , and Value matrices V I and V O . The process of generating the matrices is as follows:

[0062] Q I ,K I ,V I =LP(Norm(N I ))(1)

[0063] Q0,K0,V0=LP(Norm(N0))(2)

[0064] where Norm(·) and LP(·) represent normalization and linear projection respectively. The internal forward propagation process of the bidirectional attention mechanism is:

[0065]

[0066] where Attention(Q I ,K I ,V I ) and Attention(Q O ,K o ,V o ) capture the independent information inside the image features and the combined features respectively. Attention(Q I ,K O ,V O ) takes the image features as the main information and captures the cross information between the image features and the combined features. Attention(Q O ,K I ,V I) It takes the combined features as the main information and captures the cross-information between the combined features and the image features. λ is a hyperparameter used to balance the ratio between the independent information and the cross-information. The definition of the attention mechanism Attention(Q, K, V) is as follows:

[0067]

[0068] Among them, Softmax(·) is an activation function that can normalize a numerical vector into a probability distribution vector, and the sum of all probabilities is 1. T is the matrix transpose, and d k is the scaling factor;

[0069] The overall process of the bidirectional attention block is as follows:

[0070]

[0071] Among them, MLP(·) is a multi-layer perceptron composed of an input layer, one or more hidden layers, and an output layer. Mean(·) represents the average value, which is used to balance the two multi-modal joint feature representations obtained. The two multi-modal joint feature representations extracted after passing through the bidirectional attention block and are processed for balancing to obtain the overall multi-modal joint feature representation Then it is fed into the self-attention block to learn and enhance its feature representation. The overall process of the self-attention block is as follows:

[0072]

[0073] Among them, is the multi-modal feature representation after self-attention, and N is the final multi-modal joint feature representation.

[0074] Model training and optimization process:

[0075] PCGMMF is a classification network with multiple inputs and a single output. During the training process, the cross-entropy loss function is used as the objective function, and its definition is as follows:

[0076]

[0077] Among them, w1 represents the class weight of the positive samples, and w0 represents the class weight of the negative samples. γ is a regularization parameter to avoid model overfitting. θ represents all the parameters of the model. N represents the number of samples, and n class represents the number of sample classes, and N j represents the number of samples in class j.

[0078] In addition, the Stochastic Gradient Descent (SGD) algorithm is used for updating and optimizing the model parameters during the training process to minimize the loss function L(θ). To prevent overfitting, an early stopping strategy is introduced. That is, after training for more than 100 epochs, it is judged whether the stopping requirements are met. If the training loss does not decrease for 15 consecutive epochs, the training stops, and the maximum number of training epochs is 300. In addition, the learning rate is set to 6e-6, momentum is introduced in the stochastic gradient descent optimization and the momentum parameter is set to 0.9, weight decay (L2 regularization) is introduced, and the weight decay coefficient is set to 1e-5. Dropout is enabled and the ratio is set to 0.25.

[0079] Contents of the experimental process:

[0080] 1. Evaluation metrics:

[0081] Some existing common evaluation metrics such as the Receiver Operating Characteristic (ROC) curve, the Area Under the ROC Curve (AUC), etc. are adopted in this scheme to mainly evaluate the performance of the method. In addition, a series of evaluation metrics such as Accuracy, Precision, F1-Score, and Matthews Correlation Coefficient are used for comprehensive evaluation. The calculation formulas of these metrics are as follows.

[0082]

[0083] Among them, TPR represents the true positive rate, and FPR represents the false positive rate. TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative respectively. Recall is the recall rate, and its main function is to focus on how many positive samples the model can find. Its formula is

[0084] 2. Feature correlation evaluation

[0085] For gene expression data and DNA methylation data, this scheme conducts comprehensive feature screening. In the task of predicting the prognosis, recurrence, and metastasis risk of breast cancer, 60 most relevant features are finally selected for prognosis analysis. Figure 5 Figure A shows 20 relevant features in the gene expression data, their importance levels, and the distribution of their high and low expressions in the samples. It can be seen that these relevant genes are generally lowly expressed in samples with breast cancer prognosis recurrence or metastasis, which provides a reference basis for breast cancer prognosis analysis based on gene expression data. Figure 5 Figure C and Figure 5Figure E shows the distribution of 20 relevant features in the DNA methylation data, their importance levels, and the β-values of their high and low CpG sites in the samples. It can be seen that these relevant features generally have low CpG site β-values in samples of breast cancer prognosis recurrence or metastasis, which can also provide a reference for breast cancer prognosis analysis based on DNA methylation data.

[0086] To further screen features and verify the effectiveness of the screened features, this scheme uses the Mantel test and correlation analysis as auxiliary screening tools and verification references. Figure 5 Figure B shows the results of the Mantel test and correlation analysis of 20 relevant genes in the gene expression data. Different colors and line thicknesses are used in the figure to represent the results of the Mantel test. The analysis shows that the p-values of the Mantel test for most genes are less than the threshold of 0.05, indicating that they have a significant correlation with breast cancer prognosis recurrence and metastasis (Tumor in the figure), such as gene features like USP4, WDR82, and PNRC2. In addition, the correlations between relevant gene features and other factors of breast cancer prognosis are also analyzed. Genes such as CASP7, CLKLF, and USP4 have significant correlations with age, stage, and hormone receptor, providing evidence and reference for multimodal feature fusion. Finally, the correlations between gene features are also analyzed. It can be seen that there are also strong correlations between gene features. For example, there is a strong positive correlation between PARP9 and UBE2L6, TRANK1, etc., and a strong negative correlation between NDUFA4 and WDR82, PCGF5, etc. Screening the significant correlations between gene features not only confirms the effectiveness of the feature screening strategy in this study but also provides evidence for the discovery of breast cancer prognosis biomarkers.

[0087] Similarly, this scheme also performs the Mantel test and correlation analysis on the DNA methylation data. Figure 5 Figure D and Figure 5 Figure F show the analysis results of 20 relevant features in the DNA methylation data. The analysis shows that the p-values of the Mantel test for the relevant DNA methylation features after the feature screening strategy are mostly less than the threshold of 0.05. For example, at sites such as cg11564670, cg13086467, and cg14310890 in DNA methylation group 27, and in DNA methylation 450 group, except for the site cg08220793, the correlation with breast cancer prognosis may not be significant, and the rest of the sites have significant correlations, indicating that the screened features are related to breast cancer prognosis ( Figure 5There is a significant correlation between (Tumor in ), indicating that for the changes in data, it does not affect the good effect of the feature screening strategy in this study. In addition, the correlation between DNA methylation features and clinical features and gene expression features was also analyzed. The results showed that there is a significant correlation between DNA methylation features and other factors related to breast cancer prognosis, which indicates that there is a real correlation between multi-modal data, providing evidence and reference for multi-modal feature fusion. Finally, the correlation analysis was also carried out among various DNA methylation features. The analysis results showed that there is a strong correlation among the screened features, which also confirmed the effectiveness of the feature screening strategy in this study.

[0088] Based on the above analysis, it can be concluded that the feature screening strategy in this study can be well applied to the work of breast cancer prognosis analysis. At the same time, it can also provide reference for the discovery of breast cancer prognosis biomarkers.

[0089] 3. Method comparison

[0090] The PCGMMF model proposed in this scheme was compared with some main advanced methods to verify its effectiveness and advancement. The comparison methods include:

[0091] Liu et al. used image features and machine learning techniques in histopathological images to predict the risk of breast cancer prognosis recurrence and metastasis (based on the XGBoost method). See Liu X, Yuan P, Li R, et al. Predicting breast cancer recurrence and metastasis risk by integrating color and texture features of histopathological images and machine learning technologies. Comput Biol Med. 2022;146:105569. doi:10.1016 / j.compbiomed.2022.105569.

[0092] The recurrence and metastasis risk model (CNN-MCB) developed by Yang et al. for predicting breast cancer prognosis by integrating histopathological images and clinical data can be found in Yang J, Ju J, Guo L, et al. Prediction of HER2-positive breast cancer recurrence and metastasis risk from histopathological images and clinical information via multimodal deep learning. Comput Struct Biotechnol J. 2021;20:333-342. Published 2021 Dec 23. doi:10.1016 / j.csbj.2021.12.028.

[0093] The model DeepKEGG, a multi-omics integration framework provided by Lan et al. for predicting breast cancer prognosis recurrence, can be found in Lan W, Liao H, Chen Q, Zhu L, Pan Y, Chen YP. DeepKEGG: a multi-omics data integration framework with biological insights for cancer recurrence prediction and biomarker discovery. Brief Bioinform. 2024;25(3):bbae185. doi:10.1093 / bib / bbae185.

[0094] Yang et al. proposed the cancer prognosis prediction framework MMGCN by integrating gene expression data, copy number variation, and clinical data. It can be found in Yang P, Chen W, Qiu H. MMGCN: Multi-modal multi-view graph convolutional networks for cancer prognosis prediction. Comput Methods Programs Biomed. 2024;257:108400. doi:10.1016 / j.cmpb.2024.108400.

[0095] And ICSDA and MGCN-CalRF. For ICSDA, see Yao Y, Lv Y, Tong L, et al. ICSDA: a multi-modal deep learning model to predict breast cancer recurrence and metastasis risk by integrating pathological, clinical and gene expression data. Brief Bioinform. 2022;23(6):bbac448. doi:10.1093 / bib / bbac448. For MGCN-CalRF, see Palmal S, Arya N, Saha S, et al. Integrative prognostic modeling for breast cancer: Unveiling optimal multimodal combinations using graph convolutional networks and calibrated random forest[J]. Applied SoftComputing, 2024, 154111379.

[0096] The comparison results of the methods are shown in Table 3 below.

[0097] Table 3 Comparison results of different methods

[0098]

[0099] The comparison results show that the PCGMMF method proposed in this scheme demonstrates excellent performance in multiple key performance indicators. Specifically, PCGMMF achieved an AUC of 0.924, which ranked first among all methods and was significantly higher than the sub-optimal DeepKEGG method (4.8% higher), indicating that PCGMMF has a higher accuracy in distinguishing positive and negative samples. In terms of accuracy, PCGMMF also led far ahead with a result of 0.903, improving the performance by 10% compared to the best results of other methods, showing its superiority in the overall classification task. Although PCGMMF is not the best in terms of precision, the 0.862 achieved by PCGGMF is close to the 0.87 achieved by the highest ICSDA and far better than the results achieved by other methods. Notably, PCGMMF achieved the highest performance of 0.893 in the F1 score, higher than the 0.867 achieved by the sub-optimal MGCN-CalRF and far higher than other methods, indicating that it has achieved a good balance between precision and recall. In addition, PCGMMF also achieved the best result in the Matthews correlation coefficient.

[0100] In addition, to ensure the stability of the results, this scheme conducted 5-fold cross-validation and took the average value as the final result. The ROC curve of the 5-fold cross-validation is as Figure 6 shown. The detailed results of the 5-fold cross-validation are shown in Table 4 below. PCGMMF achieved an average AUC of 0.92, Acc of 0.861, Pre of 0.838, F1 value of 0.84, and Mcc of 0.719. The results show that PCGMMF still achieved advanced and stable performance.

[0101] Table 4 Results of 5-fold cross-validation

[0102]

[0103] 4. Ablation experiments

[0104] For the work of predicting the recurrence and metastasis risks of breast cancer prognosis, four ablation experiments were designed to verify the effectiveness of the modules, strategies, and multi-modal data in this method. The results of the ablation experiments are as Figure 7 shown.

[0105] First, while keeping other structures and parameters consistent, four pre-trained models, Resnet18, Resnet50, DenseNet101, and VGG16, were introduced for comparative experiments. The results are as Figure 7As shown in A. The Vision-LSTM pre-trained model used in this solution has better performance in image feature extraction than other mainstream feature extraction models, proving that Vision-LSTM has good effects on the processing of histopathological images and plays a crucial role in subsequent breast cancer prognosis analysis.

[0106] While keeping other structures and parameters the same, data combinations of four types of multimodal data were performed and ablation experiments were carried out. The experimental results of unimodal data (including histopathological images, clinical data, gene expression data, DNA methylation data) are as Figure 7 shown in B. The comparison results show that the results achieved by the multimodal data combination of this solution are far better than those of unimodal data. For bimodal data (including histopathological images and clinical data, histopathological images and gene expression data, histopathological images and DNA methylation data), the comparison results show that the results achieved by the multimodal data combination of this solution are also far better than those of bimodal data. For trimodal data (including histopathological images, clinical data and gene expression data, histopathological images, clinical data and DNA methylation data, histopathological images, gene expression data and DNA methylation data), the comparison results show that the results achieved by the multimodal data combination of this solution are also far better than those of trimodal data. The comparison with the above results of unimodal data, bimodal data and trimodal data proves the effectiveness and indispensability of the multimodal data combination of this solution.

[0107] While keeping other structures and parameters the same, several multimodal feature fusion methods such as matrix addition, concatenation, averaging and MCB were introduced for comparative experiments. The comparison results are as Figure 7 shown in E. Compared with other feature fusion methods, the multimodal feature fusion module BASMF proposed in this solution shows obvious advantages, strongly proving the effectiveness of BASMF.

[0108] While keeping other structures and parameters the same, experiments were also carried out on the number of stacked blocks inside the multimodal feature fusion module. The comparison results are as Figure 7 shown in F. The multimodal feature fusion module used in this solution, which is stacked by one bidirectional attention block and two self-attention blocks, achieves the best results. Further proving the effectiveness and applicability of the fusion module in this study.

[0109] Based on the above experiments, PCGMMF demonstrated powerful analytical and predictive capabilities, providing a reliable basis for assessing the risk of breast cancer prognosis recurrence and metastasis. It also proved the importance and wide applicability of the multi-modal data combination in this scheme for accurately predicting the risk of breast cancer prognosis recurrence and metastasis. In addition, the multi-modal feature fusion module BASMF proposed in this scheme can well capture the cross-information between each modality while maintaining the independent information of each modality, so as to obtain more comprehensive and complete comprehensive information on breast cancer patients' prognosis for prognosis analysis and clinical practice.

[0110] 5. Enrichment analysis

[0111] In this experiment, the functions of 60 related genes screened by this scheme were annotated based on Gene Ontology (GO) enrichment analysis to explore their potential roles in breast cancer prognosis work. The correlations with breast cancer prognosis were explored from three aspects: Biological Process, Cellular Component, and Molecular Function. Figure 8 The results are for GO enrichment analysis.

[0112] In the enrichment analysis of biological processes, as Figure 8 shown in A and B, multiple immune-related entries were highly significant. Among them, entries such as "positive regulation of defense response", "positive regulation of natural killer cell mediated immunity", "regulation of response to biotic stimulus", and "regulation of innate immune response" all showed extremely low P values, indicating that these genes play key roles in the immune regulation process. The above-enriched immune-related genes may affect breast cancer prognosis by influencing the functions of immune cells. In addition, some genes are also involved in the regulation processes of cell cycle and proliferation, such as "regulation of cell cycle process", etc. The regulation of these genes on the cell cycle may affect the proliferation rate of breast cancer cells and the progression and prognosis of the disease.

[0113] In the enrichment analysis of cellular components, as Figure 8As shown in C and D. Entries related to respiration and mitochondria such as "respiratory chain complex" and "mitochondrial respirasome" were significantly enriched. The metabolic reprogramming of breast cancer cells is closely related to mitochondrial function. Abnormalities in these genes may lead to mitochondrial dysfunction, affecting the viability and drug resistance of breast cancer cells, and thus affecting the prognosis. In addition, cellular components related to the endomembrane system such as "early endosome" were also enriched. Endosomes play a role in processes such as intracellular material transport and signal transduction. In breast cancer cells, endosome-related processes may be involved in the transport and signal transduction of cell surface receptors, affecting the cell's response to external signals, and further affecting cell behaviors such as growth and migration, which are closely related to the prognosis.

[0114] In the enrichment analysis of molecular functions, as Figure 8 shown in E and F. Functions related to transcriptional regulation such as "transcription corepressor activity" and "transcription coregulator activity" were significantly enriched. Transcription factors and transcriptional regulators play a key role in the regulation of gene expression in breast cancer. In addition, functions related to transporter protein activities such as "symporter activity" and "monovalentinorganic cation transmembrane transporter activity" were also enriched. The transport of intracellular substances is crucial for maintaining the normal physiological functions of cells. In breast cancer cells, abnormal ion transport may affect the intracellular ion balance, osmotic pressure, etc., and further affect the cell volume, activity, etc. For example, abnormal potassium ion transport may affect the cell membrane potential, thus affecting cell excitability and signal transduction, which are related to the processes of proliferation and migration of breast cancer cells, and ultimately affecting the prognosis.

[0115] Through the GO enrichment analysis of the screened features in this study, rich information was obtained in three aspects: biological processes, cellular components, and molecular functions. These enrichment results indicate that these genes are mainly involved in important biological processes such as immune regulation, mitochondrial function, transcriptional regulation, and material transport. These processes are closely related to the prognosis of breast cancer. They can determine the progression of the tumor and the prognosis of patients by affecting aspects such as the immunity, metabolism, gene expression regulation, and intracellular environmental stability of breast cancer cells. Further in-depth study of these genes and their related functional pathways is expected to provide important theoretical basis for breast cancer prognosis assessment, discovery of therapeutic targets, and development of new treatment strategies.

[0116] 6. Survival analysis

[0117] To verify whether the gene signature used in this study for predicting the prognosis, recurrence, and metastasis of breast cancer is an independent and comprehensive prognostic biomarker for breast cancer, this example also performed a survival analysis on it. First, the samples were divided into high-expression and low-expression groups according to the gene expression values, and then survival analysis was performed and Kaplan-Meier curves were plotted to intuitively understand the correlation of the biomarkers identified in this study with the prognosis analysis of breast cancer. Overall survival analysis was performed on all the selected important genes, and survival analysis was also performed on individual important genes. The results of the survival analysis are as Figure 9 shown.

[0118] In this protocol, the KM curve plotted through survival analysis demonstrated a close connection between the gene expression level and the risk of prognosis, recurrence, and metastasis of breast cancer. Figure 9 In A, the comparison of the survival of all the selected genes is presented. The results show that the P-value obtained through the log-rank test is less than 0.0001, indicating that there are significant differences in the survival status between the high-expression and low-expression groups, and it is unlikely that these differences are caused by random errors, reflecting the true differences in survival probabilities between the high-expression and low-expression groups, and preliminarily confirming that gene expression differences have a significant impact on the prognosis of breast cancer. Figure 9 In B to I, the results of the survival analysis of some important genes are presented. Their P-values are all less than the threshold of 0.05, and some important genes have even smaller P-values, such as the P-value of USP4 is 4e-4, the P-value of PNRC2 is 2e-4, etc. This indicates that there are also significant differences in the survival status of the expression levels of these important genes, and these differences are statistically significant. This means that the changes in the expression levels of these genes are closely related to the prognosis of breast cancer and can be used as potential prognostic biomarkers. Based on the above analysis, it can be known that when the expression levels of most important genes are low, it will lead to a worse survival status, which can provide a reliable reference for the prognosis analysis of breast cancer.

[0119] The gene expression level can affect the development process of breast cancer, which is of great significance for predicting the prognosis of breast cancer patients and helps clinicians formulate more precise treatment strategies and more accurate assessments of patient prognosis.

[0120] 7. Interpretability analysis

[0121] The histopathological images after feature extraction can not only be used for subsequent prediction tasks, but also verify whether they are consistent with the clinical diagnosis of pathologists. They can also show the regions highly correlated with breast cancer prognosis in the histopathological images and be used for clinical guidance. The attention mechanism is applied in the feature extraction of histopathological images, which can well aggregate the regions of high importance in the whole histopathological image and reduce the influence of regions with low correlation. In this scheme, the attention score is calculated for each image region, and the attention heatmap is drawn. At the same time, the 9 image regions with the greatest impact on the breast cancer prognosis outcome are also shown. The visualization results are as Figure 10 shown. As can be seen from Figure 10 , the cell nuclei in the regions of high importance of breast cancer prognosis recurrence and metastasis samples are relatively irregular in shape, different in size and darker in staining, which may be the characteristics of cancer cells. At the same time, the arrangement of cells is basically disordered and the cell - cell connections are loose, which conforms to the characteristics of cancer cells. These features imply that this sample may have a poor prognosis. While for the samples without recurrence and metastasis in breast cancer prognosis, the cell nuclei in the regions of high importance are relatively regular in shape, approximately the same in size and not deeply stained. At the same time, the cell sorting is relatively orderly and the cell - cell connections are compact. These features indicate that this sample may have a better prognosis. Through the interpretability study of the heatmap and regions of high importance of the histopathological images of breast cancer patients' prognosis, accurate results are obtained in the task of predicting the risk of breast cancer prognosis recurrence and metastasis, and it has reference significance for the guidance of breast cancer prognosis clinical practice and the formulation of personalized treatment plans.

[0122] In summary, this scheme proposes a new method PCGMMF based on multi - modal feature fusion to predict the risk of breast cancer prognosis recurrence and metastasis by integrating histopathological images, clinical data, gene expression data and DNA methylation data. The model first introduces the Vision - LSTM pre - trained model using the idea of transfer learning, effectively solving problems such as poor training effect caused by few breast cancer samples and too large resolution of histopathological images. At the same time, the comprehensive feature screening strategy proposed in this scheme can well solve problems such as model overfitting and instability caused by the high - dimensional characteristics of data. In addition, the multi - modal feature fusion module BASMF proposed in this study can effectively solve the heterogeneity problem of multi - modal data and obtain more comprehensive and overall prognosis features.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A breast cancer prognosis prediction method based on enhanced multimodal feature fusion, characterized by: The following steps are involved: S1, obtain histopathological images, clinical pathological parameters, gene expression data, and DNA methylation data of breast cancer patients; S2, preprocess the tissue pathology image, divide it into several small images and save the coordinate values ​​and image feature data, then send each small image to the Vision-LSTM pre-trained model for feature extraction, and then obtain a multimodal feature through deep feature extraction and dimension matching; S3, preprocessing of clinical pathological parameters, including maximum-minimum normalization of age data, custom mapping of tumor stage, T stage, M stage, N stage, ER hormone receptor, PR hormone receptor, and then obtaining a multimodal feature through deep feature extraction and dimension matching; S4, feature screening of gene expression data, first using support vector machine to perform preliminary feature screening of gene expression data, then further screening of gene expression data through Mantel test and correlation analysis, finally determining the gene expression data features related to breast cancer prognosis, and then obtaining a multimodal feature through deep feature extraction and dimension matching; S5, feature screening of DNA methylation data, first using support vector machine to perform preliminary feature screening of DNA methylation data, then further screening of gene expression data through Mantel test and correlation analysis, and finally determining the DNA methylation data features related to breast cancer prognosis, and then obtaining a multimodal feature through deep feature extraction and dimension matching; S6. Through steps S2-S5, four multimodal features are formed: image features, clinical features, gene expression features, and DNA methylation features. A multimodal feature fusion module BASMF is created, and the multimodal features are sent to BASMF for feature fusion to obtain comprehensive prognostic features for prognostic analysis.

2. The method for predicting breast cancer prognosis based on enhanced multimodal feature fusion according to claim 1, characterized in that: In step S2, when preprocessing the tissue pathology image, the CLAM library is used to automatically identify and segment the tissue area, and divide it into several 224×224 small images and save the coordinate values ​​and image feature data. Each small image is sent to the Vision-LSTM pre-trained model for feature extraction, and then the small image is divided into n patches of fixed size, which are divided into non-overlapping patches through shared linear projection. The feature vector of each patch is X i ∈R m , where m is the feature dimension, and a learnable position embedding vector E is added to each patch token i ∈R n×d , where d is the embedding dimension, and the patch feature vector after adding the position embedding is Y i =X i +E i , and then the patch feature vector Y i (i=1,2,…,n) is sent to the encoder and linear projection to obtain a series of patch token sequences, which are then sent to the alternating mLSTM blocks for processing. After processing, the complete small image features are obtained. After connecting all the small image features according to the coordinates, the overall image features of the tissue area are obtained.

3. The method for predicting breast cancer prognosis based on enhanced multimodal feature fusion according to claim 1, characterized in that: In step S4 and step S5, when using the support vector machine to perform preliminary feature screening on the gene expression data or DNA methylation data, a weight is assigned to each feature, and the importance is ranked based on the weight, and then features that are not related to the prognosis of breast cancer are identified and removed.

4. The method for predicting breast cancer prognosis based on enhanced multimodal feature fusion according to claim 1, characterized in that: In step S6, the multimodal feature fusion module BASMF is stacked by a bidirectional attention block and two self-attention blocks; the bidirectional attention block is used to learn the multimodal joint feature representation by capturing independent information from within the feature and cross-information between different features, and the self-attention block is used to learn and enhance the multimodal joint feature representation.

5. The method for predicting breast cancer prognosis based on enhanced multimodal feature fusion according to claim 4, characterized in that: In step S6, the image features are input as part of N I , clinical features, gene expression features and DNA methylation features form another part and input it into N O ; For the bidirectional attention block, first the image feature N I and the combined feature N O Input, generate their own Query matrix Q I and Q O 、Key matrix K I and K O 、Value matrix V I and V O , the process of generating the matrix is: Q I ,K I ,V I =LP(Norm(N I )) Q0, K0, V0 = LP (Norm (N0)) Where Norm(·) and LP(·) represent normalization and linear projection respectively. The internal forward propagation process of the bidirectional attention mechanism is: Where Attention(Q I ,K I ,V I ) and Attention(Q O ,K O ,V O ) capture the independent information within the image features and the independent information within the combined features, respectively. Attention(Q I ,K O ,V O ) takes image features as the main information and captures the cross information between image features and combined features. Attention (Q O ,K I ,V I ) takes the combined features as the main information and captures the cross information between the combined features and the image features. λ is a hyperparameter used to balance the ratio between independent information and cross information. The definition of the attention mechanism Attention(Q,K,V) is: Among them, Softmax(·) is an activation function that can normalize a numerical vector into a probability distribution vector, and the sum of each probability is 1, T is the matrix transpose, d k is the scaling factor; The overall flow of the bidirectional attention block is: Among them, MLP(·) is a multi-layer perceptron, which consists of an input layer, one or more hidden layers and an output layer. Mean(·) represents the average value, which is used to balance the two multimodal joint feature representations obtained after the bidirectional attention block. and Perform balancing to obtain the overall multimodal joint feature representation Then it is sent to the self-attention block to learn and enhance its feature representation. The overall process of the self-attention block is: in, is the multimodal feature representation after self-attention, and N is the final multimodal joint feature representation.