Multi-modal fusion cancer lifetime prediction system and storage medium
Through the effective integration and interaction of multimodal information, combined with the construction of survival prediction classifiers and loss functions, the problem of excessive dependence on partial modalities in the prior art is solved, and the performance and reliability of cancer survival prediction models are improved.
Patent Information
- Application Number
- CN202510110199.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
AI Technical Summary
The existing multimodal cancer survival prediction methods have excessive dependence on some modalities, resulting in insufficient training of other modalities, which affects the accuracy of the survival prediction model.
By obtaining clinical data, full-sliced pathological images and gene data of cancer patients, feature extraction and fusion are performed, loss functions are constructed using survival prediction classifiers, model parameters are adjusted, dependence on individual modalities is reduced, and uncertainty is quantified through a progressive cross-guidance mechanism and Dirichlet distribution.
It improves the performance of the survival prediction model, enhances the feature expression ability, improves the model's understanding of multimodal data and prediction accuracy, and provides decision uncertainty estimates, improving the reliability of diagnosis.
Smart Images

Figure CN119943430A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal learning, and in particular to a multimodal fusion cancer survival prediction system and a storage medium. Background Art
[0002] In cancer treatment and management, accurate prediction of patient survival prognosis is crucial for developing personalized treatment plans and evaluating treatment effects. Pathological image unimodal survival prediction methods and genetic unimodal survival prediction methods have shown good results, but they do not fully utilize diverse biomedical data. With the rapid development of medical imaging technology and molecular biology, cancer survival prediction methods based on multimodal data have emerged. These methods make comprehensive use of pathological images and genomic data to improve the accuracy of prediction.
[0003] The two types of data and sources are different, so how to effectively fuse heterogeneous data becomes a challenge. Both modalities contain rich information, but only a small part of the information can be correlated and used for survival prediction. Previous studies have adopted a one-time feature interaction after encoding the two modalities separately. The shallow interaction is not enough to cope with the high heterogeneity between modalities. In addition, the accuracy of pathological diagnosis is directly related to whether the patient can receive the correct treatment, so how to ensure the reliability of diagnosis becomes an important issue. Existing multimodal cancer survival prediction methods mainly focus on how to extract key features from pathological images and genetic data through image processing and bioinformatics techniques, and effectively integrate these multimodal biomedical data. These methods have made progress in utilizing the complementary information of different modalities, but they are still lacking in interpretability and reliability. In practical applications, the quality of data is not always reliable and stable, so it becomes particularly important to provide reliable decision uncertainty estimates, especially in high-risk medical fields. In addition, multimodal methods often show over-reliance on some modalities instead of considering all modalities fairly, resulting in insufficient training of other modalities. Summary of the invention
[0004] In view of the defects in the prior art, the present invention provides a multimodal fusion cancer survival prediction system and storage medium, which solves the problem that the multimodal cancer survival prediction in the prior art is overly dependent on some modalities, resulting in insufficient training of other modalities, so that the prediction effect of the survival prediction model is not accurate enough.
[0005] In order to achieve the above-mentioned purpose, one aspect of the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program includes program instructions, which, when executed by a processor, causes the processor to perform the following steps: obtaining clinical data, whole-slice pathological images and gene data of a cancer patient; obtaining image feature data and gene feature data based on the whole-slice pathological images and the gene data; performing feature fusion on the image feature data and the gene feature data to obtain fused feature data; introducing a survival prediction classifier, and constructing a loss function using the survival prediction classifier based on the fused feature data, the image feature data and the gene feature data; constructing a survival prediction model, training and evaluating the survival prediction model using the clinical data, the whole-slice pathological images and the gene data, and adjusting the parameters of the survival prediction model through the loss function.
[0006] The present invention reduces excessive reliance on a single modality through effective integration and interaction of multimodal information, and realizes parameter adjustment during survival prediction model training by constructing a loss function, thereby improving the performance of the survival prediction model.
[0007] Optionally, obtaining image feature data and gene feature data based on the full-slice pathology image and the gene data includes: performing semantic segmentation on the full-slice pathology image, and performing gene expression difference analysis on the gene data to obtain multiple pixel blocks and multiple mutated genes; extracting features from the pixel blocks and the mutated genes, respectively, and performing dimensionality reduction processing on the features to obtain the image feature embedding data and the gene feature embedding data; and performing feature encoding on the image feature embedding data and the gene feature embedding data, respectively, to obtain the image feature data and the gene feature data.
[0008] The present invention segments the data processing object through semantic segmentation and gene expression difference analysis, and performs dimensionality reduction processing on the segmented data, thereby reducing the complexity of calculation, and then performs feature encoding on the data, converting the feature data into a type suitable for machine learning, thereby improving data utilization efficiency.
[0009] Optionally, the performing feature encoding on the image feature embedding data and the gene feature embedding data respectively to obtain the image feature data and the gene feature data comprises: constructing a unimodal encoder; inputting the image feature embedding data and the gene feature embedding data into the unimodal encoder to obtain the image feature data and the gene feature data.
[0010] The present invention realizes simultaneous encoding of data of multiple modes by constructing a unimodal encoder, thereby improving encoding efficiency.
[0011] Optionally, constructing a unimodal encoder includes: introducing a self-attention mechanism and a feedforward neural network, and constructing a visual Transformer model based on the self-attention mechanism and the feedforward neural network; constructing a double-layer pathology image encoder and a gene encoder of the visual Transformer model based on the full-slice pathology image and the gene data; and constructing the unimodal encoder based on the pathology image encoder and the gene encoder.
[0012] The present invention constructs a visual Transformer model by introducing a self-attention mechanism and a feedforward neural network, which can effectively capture the complex patterns of images and genetic data. The pathological image encoder and gene encoder of the two-layer visual Transformer model are optimized for full-slice pathological images and genetic data, respectively, enhancing the feature extraction capability of the unimodal encoder.
[0013] Optionally, the feature fusion of the image feature data and the gene feature data to obtain the fused feature data includes: introducing a progressive cross-guidance mechanism; using the progressive cross-guidance mechanism to progressively perform cross-modal interaction between the image feature data and the gene feature data to obtain the fused feature data.
[0014] The present invention gradually fuses image and gene feature data through a progressive cross-guidance mechanism to achieve deep interaction of cross-modal information, thereby obtaining richer fusion features. This method not only enhances the feature expression capability, but also helps to capture the complex relationship between data, improve the model's understanding of multimodal data and prediction accuracy. In addition, progressive guidance helps alleviate the problem of information loss and makes feature fusion more sufficient and effective.
[0015] Optionally, the progressive cross-guidance mechanism is used to progressively perform cross-modal interaction between the image feature data and the gene feature data to obtain the fused feature data, including: performing cross-attention calculations on the image feature data and the gene feature data to obtain gene-guided embedding data and image-guided embedding data; performing attention pooling on the gene-guided embedding data and the image-guided embedding data, respectively, and inputting the results of the attention pooling into the unimodal encoder to obtain gene-guided feature data and image-guided feature data; and obtaining the fused feature data based on the gene-guided feature data, the image-guided feature data, the image feature data and the gene feature data.
[0016] The present invention realizes the deep fusion of image and gene feature data through cross-attention mechanism and attention pooling. Gene-guided embedding and image-guided embedding enhance the feature expression ability, and attention pooling further refines key information and improves the expressiveness of fused feature data for data fusion.
[0017] Optionally, constructing a loss function using the survival prediction classifier based on the fused feature data, the image feature data and the gene feature data includes: inputting the fused feature data, the image feature data and the gene feature data into the survival prediction classifier to obtain a prediction result; constructing a Dirichlet distribution based on the prediction result, and obtaining image uncertainty, gene uncertainty and fusion uncertainty based on the Dirichlet distribution; constructing an uncertain penalty loss function based on the image uncertainty, the gene uncertainty and the fusion uncertainty; and constructing the loss function based on the uncertain penalty loss function.
[0018] The present invention can accurately predict the survival result by inputting fusion feature data, image feature data and gene feature data into the survival prediction classifier, and uses Dirichlet distribution to quantify the uncertainty of image, gene and fusion features to improve the sensitivity of the loss function to feature uncertainty.
[0019] Optionally, the survival loss function satisfies the following formula: in, is the survival loss function, The patient review status, with a value of 0 or 1. The patient died of the disease. Represents patients who were alive at the last follow-up. For the survival prediction classifier, predict the patient in The raw fraction of deaths per time interval, For patients to survive The probability of a time interval, For patients in The probability of death in a time interval, For the survival prediction classifier, predict the patient in The raw fraction of deaths per time interval.
[0020] The survival loss function of the present invention provides a comprehensive evaluation standard for the survival prediction model by combining the patient's survival probability and death risk. It takes into account the patient's review status, can more accurately reflect the actual survival situation, and improve the accuracy of the survival loss function.
[0021] Optionally, the uncertain penalty loss function satisfies the following formula: in, is the uncertain penalty loss function, To calculate the maximum value function of the parameter, For the fusion uncertainty, For the gene uncertainty, is the image uncertainty.
[0022] The uncertainty penalty loss function of the present invention improves the accuracy of the uncertainty penalty loss function by quantifying the uncertainty of fusion features, gene features and image features.
[0023] Another aspect of the present invention provides a multimodal fusion cancer survival prediction system, comprising an input device, a processor, an output device and a memory, wherein the input device, the processor, the output device and the memory are interconnected, the memory comprises the computer-readable storage medium as described in the previous aspect of the present invention, the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions.
[0024] A multimodal fusion cancer survival prediction system of the present invention has a compact structure, stable performance, high integration and simple composition, and can stably execute the steps of program instructions in a computer-readable storage medium provided in the previous aspect of the present invention, further improving the overall applicability and practical application capabilities of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flowchart of program instructions in a computer-readable storage medium according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a multi-modal fusion cancer survival prediction system according to an embodiment of the present invention; Figure 3 This is the consistency index result on the lung adenocarcinoma dataset of an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are only for illustration and are not intended to limit the present invention. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that these specific details do not need to be adopted to implement the present invention. In other examples, in order to avoid confusing the present invention, known circuits, software or methods are not specifically described.
[0027] Throughout the specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment of the present invention. Therefore, the phrases "in one embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily all refer to the same embodiment or example. In addition, particular features, structures, or characteristics may be combined in one or more embodiments or examples in any suitable combination and / or subcombination. In addition, it should be understood by those of ordinary skill in the art that the figures provided herein are for illustrative purposes and that the figures are not necessarily drawn to scale.
[0028] See also Figure 1 In one embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor performs the following steps: Step S1, obtaining clinical data, whole-slice pathological images and genetic data of cancer patients.
[0029] In this embodiment, sources of clinical data, whole-slice pathological images, and genetic data of cancer patients include hospitals, research institutions, and public databases.
[0030] The clinical data of cancer patients include the patients’ basic information, medical history, treatment records, and survival status.
[0031] Full-slice pathology images are a technology that digitizes the entire pathology slice. It uses a scanner to perform high-resolution scanning and splicing of pathology slices to generate visual digital images. It can capture the tiny details of tissue slices and provide a reliable basis for pathology diagnosis.
[0032] Genetic data refers to an individual's genome sequence and related information, including the characteristics of DNA, RNA, and their expression products. These data are used to understand genetic characteristics, disease risks, biological characteristics and evolution. Genomic data can uniquely identify individuals and be associated with sensitive information such as diseases and blood relationships.
[0033] Step S2, obtaining image feature data and gene feature data according to the full-slice pathological image and the gene data.
[0034] Wherein, obtaining the image feature data and the gene feature data according to the full-slice pathological image and the gene data specifically includes the following sub-steps: Step S201, semantic segmentation is performed on the whole-slice pathological image, and gene expression difference analysis is performed on the gene data to obtain a plurality of pixel blocks and a plurality of mutated genes.
[0035] In this implementation, semantic segmentation is performed on the whole-slice pathology image to obtain multiple pixel blocks; gene expression difference analysis is performed on the gene data to obtain multiple mutant genes. This method essentially divides large data into multiple small data, which facilitates subsequent calculations and reduces the complexity of data calculations.
[0036] Each full-slice pathology image is semantically segmented, and its goal is to classify each pixel in each full-slice pathology image to distinguish different objects and backgrounds. Therefore, semantic segmentation is performed on each full-slice pathology image to remove the background and segment the full-slice pathology image after background removal into non-overlapping pixel blocks of 256*256 pixels.
[0037] Gene expression differential analysis is an important method to study the changes in gene expression levels under different conditions (such as different tissues, different developmental stages or different disease states). Gene expression differential analysis is performed on gene data by comparing each gene in the gene data with the control group using logFC to evaluate changes in gene expression relative to normal tissues and screen out the 2,000 genes with the largest changes. These 2,000 genes are divided into several gene function groups based on common features such as homology or biochemical activity to reflect the genomic embedding of unique biological functions, namely, cytokine and growth factor group, transcription factor group, cell differentiation marker group, protein kinase group, oncogene group and tumor suppressor group.
[0038] LogFC is the logarithmic fold change, which is used to measure the relative change of gene expression levels under two conditions. It is an important indicator in the analysis of gene expression differences, and is calculated by taking the logarithm (usually log2) of the ratio of gene expression in the experimental group and the control group. The value of LogFC can intuitively reflect the degree of upregulation or downregulation of gene expression.
[0039] The LogFC value is calculated as follows: in, is the gene data expression, is the gene expression data of the control group.
[0040] when , it means that the gene data expression level is higher than that of the control group, that is, the gene is upregulated.
[0041] when , it means that the gene data expression level is lower than that of the control group, that is, the gene is down-regulated.
[0042] when , it means that the expression level of the gene data is not significantly different from that of the control group.
[0043] Step S202: extract features from the pixel block and the mutated gene respectively, and perform dimensionality reduction processing on the features to obtain the image feature embedding data and the gene feature embedding data.
[0044] In this embodiment, the features of the pixel blocks are extracted by inputting each pixel block into the ResNet50 model pre-trained in ImageNet one by one, performing feature extraction, performing dimensionality reduction processing on the extracted features, and embedding them into a 1024-dimensional vector to form an image feature package vector. The image feature package vector is , where n represents the number of pixel blocks. In most cases, extracting features from pixel blocks will result in high-dimensional feature vectors, and dimensionality reduction can reduce the complexity of the calculation.
[0045] ImageNet is a large-scale annotated image database designed to support the research of visual object recognition software. The ResNet50 model pre-trained on ImageNet, usually referred to as a pre-trained model, refers to a ResNet50 model that has been trained using the ImageNet dataset. This model has learned the ability to extract features from images during the training process and can be directly used for feature extraction of pixel blocks.
[0046] The ResNet50 model is a deep residual network with 50 layers. When using ResNet50 for feature extraction, the last fully connected layer is usually removed, and the previous convolutional layer and pooling layer are retained to extract common image features. The feature map size output by the last convolutional layer of the ResNet50 model is 2048 dimensions, specifically 2048 channels. After global average pooling, the feature map is compressed into a 2048-dimensional vector. Subsequently, the ResNet50 model can embed the 2048-dimensional vector into 1024 dimensions through linear transformation or PCA dimensionality reduction. It should be noted that linear transformation is suitable for fast implementation, while PCA is more suitable for scenarios where the original information needs to be retained. When choosing, you can flexibly choose according to different actual conditions.
[0047] The feature extraction of the mutant gene is to input each mutant gene into the SNN one by one, perform feature extraction, perform dimensionality reduction processing on the extracted features, and embed them into a 1024-dimensional vector to form a gene feature package vector. The gene feature package vector is , where m is the number of genomes. Similarly, in most cases, extracting features from mutant genes will result in high-dimensional feature vectors, and dimensionality reduction can reduce the complexity of the calculation.
[0048] SNN is a neural network model that transmits and processes information by simulating the pulse emission process of biological neurons. It is different from traditional artificial neural networks (ANNs) in that SNNs use discrete pulse signals to encode information instead of continuous activation values.
[0049] Step S203, feature encoding is performed on the image feature embedded data and the gene feature embedded data respectively to obtain the image feature data and the gene feature data.
[0050] The steps of respectively encoding the image feature embedded data and the gene feature embedded data to obtain the image feature data and the gene feature data specifically include the following sub-steps: Step S20301, construct a unimodal encoder.
[0051] Among them, constructing a unimodal encoder specifically includes the following sub-steps: Step S2030101, introduce a self-attention mechanism and a feedforward neural network, and construct a visual Transformer model based on the self-attention mechanism and the feedforward neural network.
[0052] In this embodiment, the self-attention mechanism is an internal attention mechanism that allows the model to dynamically adjust the degree of attention to each element when processing sequence data. Unlike traditional attention mechanisms, the self-attention mechanism does not rely on external information, but directly focuses on the interactions between elements within the input sequence. The core of the self-attention mechanism is to map the input sequence into three vectors: query, key, and value.
[0053] Feedforward neural network is one of the most basic neural network architectures. Its core feature is that information flows unidirectionally between neurons without feedback connections. It is one of the most common network types in deep learning and is widely used in tasks such as classification, regression, image recognition, and speech recognition. Feedforward neural network consists of multiple layers, including input layer, hidden layer, and output layer. The input layer is used to receive input data. The hidden layer includes one or more intermediate layers to extract features of the input data, and each layer consists of multiple neurons. The output layer is used to produce the final output result.
[0054] Step S2030102, constructing a pathology image encoder and a gene encoder of the double-layer visual Transformer model respectively according to the full-slice pathology image and the gene data.
[0055] In this embodiment, the pathological image encoder is composed of two layers of visual Transformer models, the first layer of visual Transformer model is used to extract high-dimensional features of image features, and the second layer of visual Transformer model is used to capture global dependencies. The gene encoder is also composed of two layers of visual Transformer models, the first layer of visual Transformer model is used for high-dimensional features of gene features, and the second layer of visual Transformer model is used to capture global dependencies.
[0056] Step S2030103, constructing the unimodal encoder according to the pathological image encoder and the gene encoder.
[0057] In this embodiment, the unimodal encoder composed of the pathological image encoder and the gene encoder can simultaneously and efficiently process data of two different modalities, image and gene. The main function of the unimodal encoder is to further encode the input feature embedding data into a more expressive feature representation.
[0058] Step S20302: input the image feature embedding data and the gene feature embedding data into the unimodal encoder to obtain the image feature data and the gene feature data.
[0059] In this embodiment, the image feature embedding data is input into the unimodal encoder, which uses the pathological image encoder to extract the global feature representation of the image feature embedding data and then outputs the image feature data. At the same time, the gene feature embedding data is input into the unimodal encoder, which uses the gene encoder to extract the global feature representation of the gene feature embedding data and then outputs the gene feature data.
[0060] Step S3, performing feature fusion on the image feature data and the gene feature data to obtain fused feature data.
[0061] The step of fusing the image feature data and the gene feature data to obtain the fused feature data specifically includes the following sub-steps: Step S301, introducing a progressive cross-guidance mechanism.
[0062] In this embodiment, the progressive cross-guidance mechanism is a strategy for multimodal data fusion, which aims to enhance the complementarity and consistency between different modal features through gradual interaction and guidance. This mechanism usually combines the ideas of progressive fusion and cross-guidance, and achieves more efficient information integration and more accurate task performance through multi-stage feature interaction and optimization.
[0063] Progressive fusion refers to the gradual fusion of features of different modalities in stages to avoid information loss or redundancy problems caused by direct fusion.
[0064] Cross-guidance refers to using the features of one modality to guide the feature learning of another modality, enhancing the complementarity between modalities and reducing modality differences.
[0065] For the cross-attention mechanism, the input modality representation , , using the weight matrix , , , Perform linear mapping to obtain query matrix Q, key matrix K, value matrix , , calculate the attention matrix: Weight the input modal representation to obtain the modal cross representation , : Step S302: Using the progressive cross-guidance mechanism, the image feature data and the gene feature data are progressively cross-modally interacted to obtain the fused feature data.
[0066] Wherein, using the progressive cross-guidance mechanism to progressively perform cross-modal interaction on the image feature data and the gene feature data to obtain the fused feature data specifically includes: Step S30201, performing cross-attention calculation on the image feature data and the gene feature data to obtain gene-guided embedding data and image-guided embedding data.
[0067] In this embodiment, the cross-attention calculation requires the use of the cross-attention mechanism, which is a powerful attention mechanism for processing and fusing data from different modalities or sequences. It allows the features of one modality to be guided by the features of another modality, thereby enhancing the complementarity and relevance between modalities. This mechanism has been widely used in fields such as multimodal learning, machine translation, question-answering systems, and medical image analysis.
[0068] Two interaction models are established through the cross-attention mechanism. The first is the interaction model from pathological images to genes, and the second is the interaction model from genes to pathological images.
[0069] The interactive model from pathological images to genes takes the image feature data as the query and the gene feature data as the key-value pair. The attention score between the image feature data and the gene feature data is calculated by dot product or other similarity metrics, and the attention score is used as the first attention score. The first attention score is then used to perform weighted summation on the gene feature data to obtain the gene-guided embedding 1.
[0070] The interactive model from gene to pathological image is to use gene feature data as query and image feature data as key-value pair. By dot product or other similarity metrics, the attention score between gene feature data and image feature data is calculated, and the attention score is used as the second attention score. The image feature data is weighted summed using the second attention score to obtain the embedding guided by the pathological image 1.
[0071] Similarly, after cross-attention calculations are performed with the gene output and pathological image output of the first-layer transformer to obtain gene-guided embedding 2 and pathological image-guided embedding 2, cross-attention calculations are performed with the gene output and pathological image output of the second-layer transformer to obtain gene-guided embedding data and image-guided embedding data.
[0072] Therefore, gene-guided embedding data refers to the information related to gene feature data extracted from image feature data after being processed by the cross-attention mechanism, and image-guided embedding data refers to the information related to image feature data extracted from gene feature data after being processed by the cross-attention mechanism.
[0073] Step S30202, performing attention pooling on the gene-guided embedding data and the image-guided embedding data respectively, and inputting the result of the attention pooling into the unimodal encoder to obtain gene-guided feature data and image-guided feature data.
[0074] In this embodiment, gene-guided feature data and image-guided feature data are generated through a cross-attention mechanism, which contain key information of gene information and pathological image modality, respectively. However, these two data may still contain redundant information or underutilized details. In order to further extract and optimize the information in these two data, an attention pooling operation is used.
[0075] The attention pooling operation is a feature aggregation method based on the attention mechanism. It dynamically selects and aggregates key information by calculating the importance weights of gene-guided feature data and image-guided feature data. Specifically, attention pooling is applied to gene-guided feature data and image-guided feature data respectively to generate a more compact and representative global feature representation. This process not only retains the important information in gene-guided feature data and image-guided feature data, but also reduces noise and redundancy through weighted aggregation.
[0076] After the encoding of the two-layer Transformer model of the unimodal encoder is pooled with attention, the gene-guided global representation and the image-guided global representation are obtained. In order to further optimize these representations and extract deeper features, they are input into the unimodal encoder for encoding. The Transformer model in the unimodal encoder is a powerful sequence processing architecture that can capture long-distance dependencies between features through the self-attention mechanism. In this process, the first layer of Transformer is responsible for the preliminary encoding of the input global representation and extracting the interactive information of local and global features. The second layer of Transformer further optimizes the feature representation on this basis to enhance the discriminative power and expressive power of the features. Through the encoding of the two layers of Transformer, the gene-guided features and the pathological image-guided features are finally obtained. These features are processed by the multi-head self-attention mechanism and feedforward network of the Transformer, which can better capture the complex structure and semantic information within the modality.
[0077] Step S30203, obtaining the fusion feature data according to the gene-guided feature data, the image-guided feature data, the image feature data and the gene feature data.
[0078] In this embodiment, the gene-guided feature data and the image-guided feature data are added to the corresponding original modality features. Specifically, the gene-guided feature data is added to the gene feature data, and the image-guided feature data is added to the image feature data. The purpose of this process is to fuse the features optimized by the Transformer with the original modality features to retain the detailed information of the original modality. Finally, the fused gene features and image features are spliced to generate the final fused features. The splicing operation combines the features of the two modalities to form a comprehensive feature representation, which contains both the key information of the gene modality and the important features of the pathological image modality.
[0079] Step S4, introducing a survival prediction classifier, and constructing a loss function using the survival prediction classifier according to the fusion feature data, the image feature data and the gene feature data.
[0080] In an optional implementation, the survival prediction classifier is pre-trained, and the specific steps include: First, the training data is collected, which includes clinical data, image feature data, gene feature data and fusion feature data.
[0081] Next, a survival prediction classifier is constructed, and the survival prediction classifier is trained and evaluated using the clinical data, the image feature data, the gene feature data, and the fusion feature data.
[0082] Among them, introducing a survival prediction classifier, and constructing a loss function using the survival prediction classifier according to the fusion feature data, image feature data and gene feature data specifically includes the following sub-steps: Step S401, inputting the fusion feature data, the image feature data and the gene feature data into the survival prediction classifier to obtain a prediction result.
[0083] In this embodiment, the survival prediction classifier divides the survival time into four time periods. This division is based on the survival prediction needs of patients during treatment.
[0084] In terms of structure, the survival prediction classifier contains four types of outputs, corresponding to the four time periods mentioned above. This means that the survival prediction classifier needs to learn how to extract key information from the input features in order to assign each sample to the correct category.
[0085] In addition, the survival prediction classifier contains a fully connected layer. The fully connected layer is a common structure in neural networks. Its function is to perform weighted summation of input features and introduce nonlinearity through activation functions, so that it can learn complex feature combination relationships. In the survival prediction classifier, the fully connected layer is located in the last stage of the model, responsible for integrating the previously extracted features and outputting the final classification results. This structure enables the survival prediction classifier to better capture the relationship between features and improve the accuracy and reliability of predictions.
[0086] Step S402, constructing a survival loss function according to the prediction result; The survival loss function satisfies the following formula: in, is the survival loss function, The patient review status, with a value of 0 or 1. The patient died of the disease. Represents patients who were alive at the last follow-up. For the survival prediction classifier, predict the patient in The raw fraction of deaths per time interval, For patients to survive The probability of a time interval, For patients in The probability of death in a time interval, For the survival prediction classifier, predict the patient in The raw fraction of deaths per time interval.
[0087] Step, S402, constructing a Dirichlet distribution according to the prediction result, and obtaining image uncertainty, gene uncertainty and fusion uncertainty according to the Dirichlet distribution.
[0088] In this embodiment, in order to evaluate the uncertainty of model prediction, a Dirichlet distribution is constructed based on the prediction results. Dirichlet distribution is a probability distribution defined on a simplex and is often used to represent the uncertainty of multi-category probabilities.
[0089] For ease of understanding, the prediction results are decomposed into multiple parts, including the contribution of image features to the prediction, the contribution of gene features to the prediction, and the contribution of fusion features to the prediction. By constructing the Dirichlet distribution, the uncertainty of these contributions can be quantified.
[0090] Specifically, the parameters of the Dirichlet distribution can be set according to the confidence or probability distribution of the prediction results. For example, if the image feature data contributes more to the prediction results, the corresponding Dirichlet distribution parameters will reflect this information. By analyzing the shape of the Dirichlet distribution, three types of uncertainty can be obtained: image uncertainty, gene uncertainty, and fusion uncertainty. Image uncertainty reflects the degree of uncertainty of the image feature data on the prediction results. If the quality of the image feature data is high and highly correlated with the survival prediction, the image uncertainty will be low; conversely, if the image feature data is noisy or has a weak correlation with the prediction target, the image uncertainty will be high. Similarly, gene uncertainty reflects the uncertainty of the gene feature data on the prediction results. The complexity of gene data (such as heterogeneity of gene expression, diversity of mutations, etc.) may lead to high gene uncertainty. Fusion uncertainty comprehensively considers the integration effect of image features and gene features. It reflects the uncertainty of the fusion feature data on the prediction results. If the fusion algorithm can effectively integrate image and gene information, the fusion uncertainty will be low; otherwise, the fusion uncertainty will be high. By analyzing these three uncertainties, we can better understand the reliability of model predictions and provide a more comprehensive reference for clinical decision-making.
[0091] In practical applications, this uncertainty analysis based on Dirichlet distribution is of great significance. It can reveal which data sources contribute more to the prediction results and which data sources may contain noise or inaccurate information. In addition, by quantifying uncertainty, the model can be further optimized, such as by improving data preprocessing, adjusting fusion strategies, or retraining the model to reduce uncertainty, thereby improving the accuracy and reliability of the prediction.
[0092] Therefore, by inputting fusion feature data, image feature data, and gene feature data into the survival prediction classifier and using Dirichlet distribution to evaluate uncertainty, we can not only obtain survival prediction results, but also gain in-depth understanding of the reliability of model predictions and potential improvement directions. This combination of multimodal data fusion and uncertainty analysis provides a powerful tool for medical research and clinical applications, and helps promote the development of precision medicine.
[0093] The uncertainty satisfies the following formula: in, For uncertainty, are the parameters of the Dirichlet distribution constructed from the prediction results.
[0094] Step S403: construct an uncertainty penalty loss function according to the image uncertainty, the gene uncertainty and the fusion uncertainty.
[0095] The uncertain penalty loss function satisfies the following formula: in, is the uncertain penalty loss function, To calculate the maximum value function of the parameter, For the fusion uncertainty, For the gene uncertainty, is the image uncertainty.
[0096] Step S404: construct the loss function according to the uncertainty penalty loss function and the survival loss function.
[0097] The loss function satisfies the following formula: in, is the loss function, is the survival loss function, is the uncertain penalty loss function, is the weight parameter.
[0098] Step S5, constructing a survival prediction model, using the clinical data, the whole-slice pathological images and the gene data to train and evaluate the survival prediction model, and adjusting the parameters of the survival prediction model through the loss function.
[0099] In this embodiment, when machine learning trains the survival prediction model, the loss function is an important indicator to measure the difference between the prediction result of the survival prediction model and the true value, which provides direction and basis for the optimization of model parameters. For the survival prediction model, the role of the loss function is particularly critical because it needs to process complex survival time data and possible missing data, which means that the survival time of some samples is not fully observed.
[0100] There is a direct connection between the loss function and the parameters of the survival prediction model. The output value of the loss function reflects the prediction performance of the model under the current parameters, and the goal of the model is to minimize this loss value by adjusting the parameters. In other words, the loss function provides a feedback signal to the model, telling it which parameters need to be adjusted and how to adjust them, ultimately improving the accuracy of the survival prediction model.
[0101] The adjustment of the survival prediction model parameters is achieved through an optimization algorithm, and the most commonly used method is the gradient descent method. During the training process, the survival prediction model calculates the gradient (i.e., partial derivative) of the loss function with respect to each parameter, and then updates the parameters according to the direction and magnitude of the gradient. Specifically, if the gradient of a parameter is positive, it means that increasing the parameter will lead to an increase in loss, so the parameter needs to be reduced; conversely, if the gradient is negative, the parameter needs to be increased.
[0102] In short, in the training process of the survival prediction model, the relationship between the loss function and the parameters of the survival prediction model is interdependent. The loss function provides the optimization direction for the survival prediction model, and the model minimizes the loss value by adjusting the parameters. This can improve the performance of the survival prediction model and make it better adapt to complex survival data and practical application scenarios. In this embodiment, the present invention evaluates the survival prediction model through comparative experiments.
[0103] Experiments were conducted on cancer datasets to evaluate the effectiveness of the proposed method. The consistency index was used to evaluate the model performance, which evaluates the ability to rank the survival prediction time of multiple individuals.
[0104] The consistency index c-index is a metric used to evaluate the performance of survival analysis models. It measures the ability of a model to correctly sort pairs of individuals according to their predicted survival times. The consistency index satisfies the following formula: in, is the number of patients, is an indicator function, the value is 1 if the statement is true, and 0 otherwise. and For the and The survival time of the patients, For the The review status of each patient.
[0105] Cancer dataset: The cancer dataset TCGA is a dataset containing genomic data and clinical data of thousands of cancer patients. The prognosis data of lung adenocarcinoma (LUAD) (n=452) in TCGA is used to evaluate the performance of the survival prediction model of the present invention. In this experiment, five-fold cross validation is used to evaluate the performance of the survival prediction model of the present invention and other survival prediction models.
[0106] The experiment recorded the consistency index evaluation results of the unimodal survival prediction model and the multimodal survival prediction model as well as the consistency index evaluation results of the survival prediction model using the present invention.
[0107] like Figure 3 As shown, G is a gene unimodal survival prediction model, and W is a full-slice pathology image unimodal survival prediction model. The performance of the survival prediction model of the present invention is significantly improved compared with the unimodal survival prediction model. Specifically, compared with the best gene unimodal survival prediction model SNNTrans, the performance of the survival prediction model of the present invention is improved by 5.9%. Compared with the best exact slice pathology image unimodal survival prediction model CLAM-SB, the performance of the survival prediction model proposed by the present invention is improved by 9.6%. The survival prediction model of the present invention is superior to all compared multimodal survival prediction models. Compared with the best multimodal survival prediction model CMTA, the performance of the survival prediction model of the present invention is improved by 0.6%. This shows that the use of multi-level information progressive guidance for interaction and the use of uncertainty for calibration to promote full modality training are largely helpful for survival prediction.
[0108] In an optional implementation, the survival prediction model is evaluated by calculating the mean absolute error.
[0109] Obtain new clinical data, whole-slice pathology images, and gene data that are not involved in the training of the survival prediction model, input the whole-slice pathology images and gene data into the survival prediction model, and obtain the prediction results. Substitute the prediction results and clinical data into the mean absolute error expression to obtain the mean absolute error data set, and evaluate the survival prediction model through the mean absolute error data set. The smaller the data in the absolute error data set, the better the prediction effect of the model.
[0110] The mean absolute error expression satisfies the following formula: in, represents the mean absolute error, represents the number of samples, Indicates The true value of the survival time of samples, Indicates The survival time value of the samples.
[0111] In an alternative embodiment, the coefficient of determination is calculated to evaluate the survival prediction model.
[0112] Obtain new clinical data, whole-slice pathology images, and gene data that have not been involved in the training of the survival prediction model, input the whole-slice pathology images and gene data into the survival prediction model to obtain the prediction results. Substitute the prediction results and clinical data into the determination coefficient expression to obtain the determination coefficient data set, and evaluate the survival prediction model through the determination coefficient data set.
[0113] In the data set of the coefficient of determination, when When , it means that the model fits the data completely accurately. In the regression model, the predicted value is exactly the same as the true value, and all data points are on the regression line or regression plane. This is an ideal fit state, indicating that the model can perfectly explain the changes in the dependent variable.
[0114] In the data set of the coefficient of determination, when When , it means that there is an error between the model prediction result and the true value, which is normal. The closer the determination coefficient is to 1, the better the prediction effect of the model.
[0115] The determination coefficient expression satisfies the following formula: in, represents the coefficient of determination, represents the number of samples, Indicates The real value of the survival time of samples, Indicates The survival time value of the samples, Represents the mean of the real-valued survival times.
[0116] like Figure 2As shown, on the other hand, the present invention also provides a multimodal fusion cancer survival prediction system, comprising an input device, a processor, an output device and a memory, wherein the input device, the processor, the output device and the memory are interconnected, the memory comprises the computer-readable storage medium described above, the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions.
[0117] In this embodiment, the input device is used to provide input-related data or instructions to the system. In the cancer survival prediction system, the input device may include common human-computer interaction interface devices such as keyboards, mice, and touch screens. Through the input device, doctors or researchers can input necessary parameters.
[0118] The processor is the core component of the system, responsible for executing computer program instructions, processing and analyzing data. In the cancer survival prediction system, the processor analyzes and interprets the input test data by running pre-programmed algorithms and models. The processor can be a central processing unit (CPU), a graphics processing unit (GPU) or other dedicated processing units.
[0119] The memory is used to store computer programs, data and parameters required by the system. It can include random access memory (RAM) for temporary data storage and processing, and persistent memory (such as a hard disk or solid-state drive) for long-term storage and preservation of data.
[0120] The output device is used to present the results of system processing and analysis to users or external devices, and the output device can be a display, a printer, a chart drawing device, etc. Through the output device, the system can display the prediction results for reference by doctors, researchers or patients to assist decision-making and communication.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
Claims
1. A computer-readable storage medium, characterized in that: The computer readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor performs the following steps: Obtain clinical data, whole-slide pathology images, and genetic data of cancer patients; Obtaining image feature data and gene feature data according to the full-slice pathology image and the gene data; Performing feature fusion on the image feature data and the gene feature data to obtain fused feature data; Introducing a survival prediction classifier, and constructing a loss function using the survival prediction classifier according to the fusion feature data, the image feature data, and the gene feature data; A survival prediction model is constructed, the clinical data, the whole-slice pathological images and the gene data are used to train and evaluate the survival prediction model, and the parameters of the survival prediction model are adjusted by the loss function.
2. A computer-readable storage medium according to claim 1, characterized in that: The obtaining of image feature data and gene feature data according to the full-slice pathological image and the gene data comprises: Performing semantic segmentation on the full-slice pathological image and performing gene expression difference analysis on the gene data to obtain a plurality of pixel blocks and a plurality of mutated genes; Extracting features from the pixel block and the mutated gene respectively, and performing dimensionality reduction processing on the features to obtain the image feature embedding data and the gene feature embedding data; Feature encoding is performed on the image feature embedded data and the gene feature embedded data respectively to obtain the image feature data and the gene feature data.
3. A computer-readable storage medium according to claim 2, characterized in that: The performing feature encoding on the image feature embedding data and the gene feature embedding data respectively to obtain the image feature data and the gene feature data comprises: Construct a unimodal encoder; The image feature embedding data and the gene feature embedding data are input into the unimodal encoder to obtain the image feature data and the gene feature data.
4. A computer-readable storage medium according to claim 3, characterized in that: The constructing of a unimodal encoder comprises: A self-attention mechanism and a feedforward neural network are introduced, and a visual Transformer model is constructed according to the self-attention mechanism and the feedforward neural network; According to the full-slice pathological image and the gene data, constructing a pathological image encoder and a gene encoder of the double-layer visual Transformer model respectively; The unimodal encoder is constructed according to the pathological image encoder and the gene encoder.
5. The computer-readable storage medium according to claim 1, wherein: The performing feature fusion on the image feature data and the gene feature data to obtain fused feature data comprises: Introducing a progressive cross-guidance mechanism; The image feature data and the gene feature data are gradually cross-modally interacted by utilizing the progressive cross-guidance mechanism to obtain the fused feature data.
6. A computer-readable storage medium according to claim 5, characterized in that: The step of using the progressive cross-guidance mechanism to progressively perform cross-modal interaction on the image feature data and the gene feature data to obtain the fused feature data includes: Performing cross attention calculation on the image feature data and the gene feature data to obtain gene-guided embedding data and image-guided embedding data; Performing attention pooling on the gene-guided embedding data and the image-guided embedding data respectively, and inputting the results of the attention pooling into the unimodal encoder to obtain gene-guided feature data and image-guided feature data; The fusion feature data is obtained according to the gene-guided feature data, the image-guided feature data, the image feature data and the gene feature data.
7. A computer-readable storage medium according to claim 1, characterized in that: The constructing a loss function using the survival prediction classifier according to the fusion feature data, the image feature data and the gene feature data comprises: Inputting the fusion feature data, the image feature data and the gene feature data into the survival prediction classifier to obtain a prediction result; Constructing a survival loss function according to the prediction results; Constructing a Dirichlet distribution according to the prediction result, and obtaining image uncertainty, gene uncertainty and fusion uncertainty according to the Dirichlet distribution; Constructing an uncertainty penalty loss function according to the image uncertainty, the gene uncertainty and the fusion uncertainty; The loss function is constructed according to the uncertainty penalty loss function and the survival loss function.
8. A computer-readable storage medium according to claim 7, characterized in that: The survival loss function satisfies the following formula: in, is the survival loss function, The patient review status, with a value of 0 or 1. The patient died of the disease. Represents patients who were alive at the last follow-up. For the survival prediction classifier, predict the patient in The raw fraction of deaths per time interval, For patients to survive The probability of a time interval, For patients in The probability of death in a time interval, For the survival prediction classifier, predict the patient in The raw fraction of deaths per time interval.
9. A computer-readable storage medium according to claim 7, characterized in that: The uncertain penalty loss function satisfies the following formula: in, is the uncertain penalty loss function, To calculate the maximum value function of the parameter, For the fusion uncertainty, For the gene uncertainty, is the image uncertainty.
10. A multimodal fusion cancer survival prediction system, characterized in that: The invention comprises an input device, a processor, an output device and a memory, wherein the input device, the processor, the output device and the memory are interconnected, the memory comprises a computer-readable storage medium as claimed in any one of claims 1 to 9, the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions.
Citation Information
Cited By
Multi-modal data fusion-based interpretable cancer survival prediction method
CN120234764A
An interpretable cancer survival prediction method based on multimodal data fusion
CN120234764B