Multi-modal cancer survival prediction method based on potential differentiation variational auto-encoder

Through the multimodal cancer survival prediction method of potential differentiation variant autoencoder, WSI characterization is compressed and functionally specific genomic embedding is generated. Combined with PoE technology and alignment loss, high-precision survival risk prediction is achieved, solving the problems of computational redundancy and biological heterogeneity in the existing methods, and improving clinical applicability.

CN120354244AInactive Publication Date: 2025-07-22NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510838667.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing multimodal cancer survival prediction methods lack functional specificity when processing missing genomic data, resulting in inaccurate generation results and inefficient computational efficiency, making it difficult to scale to large-scale datasets.

Method used

Using a method based on a potential differentiation variational autoencoder, WSI characterization is compressed through the VIB-Trans module, LD-VAE is used to generate functionally specific genomic embedding, combined with PoE technology and alignment loss, multimodal joint distribution estimation is achieved, and high-precision survival risk prediction is output.

Benefits of technology

It improves the accuracy and clinical applicability of multimodal cancer survival prediction under missing data conditions, and solves the problems of computational redundancy and biological heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354244A_ABST
    Figure CN120354244A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal cancer survival prediction method based on a potential differentiation variational auto-encoder, and the method comprises the following steps: carrying out the tissue region segmentation of an input full-view digital slice, extracting the pathological features, and carrying out the grouping extraction of the grouping features of input genome data according to the function category; generating compressed pathological feature potential distribution through an information bottleneck theory and an attention mechanism; potential distribution of genome data is learned through global posteriori, specific potential variables are generated through a functional differentiation network, and missing genome features are reconstructed; integrating pathology and genome posteriori based on an expert product technology, and introducing alignment loss to constrain consistency of posteriori distribution; and screening survival related features through a co-attention mechanism, and outputting a survival probability and risk layering result. By adopting the multi-modal cancer survival prediction method based on the potential differentiation variational auto-encoder, the problem of calculation redundancy is solved, multi-modal joint distribution estimation under missing data is realized, and the clinical applicability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical artificial intelligence, and particularly to a multi-modal cancer survival prediction method based on a latent differentiation variational autoencoder. Background Art

[0002] The integration of histopathological images and genomic data for analysis has received increasing attention in human cancer survival prediction. Existing studies usually assume that all modal data are completely available, and significantly improve the prediction performance based on multi-modal fusion (such as co-attention mechanism, joint embedding learning).

[0003] However, in actual clinical scenarios, genomic data are often missing due to high detection costs, high technical thresholds, etc., resulting in a significant decline in the performance of existing multi-modal models in the test stage. In addition, existing methods face the problem of low computational efficiency when processing gigapixel-level whole-slide digital sections (WSIs). For example, traditional multi-instance learning models (such as TransMIL, ABMIL) need to directly process tens of thousands of image patches, and the self-attention computational complexity is as high as , making it difficult to scale to large-scale datasets. More critically, when existing generative models (such as VAE, GAN) reconstruct missing genomic data, they fail to consider the biological heterogeneity of different functional genomes, such as the opposite contribution directions of oncogenes and tumor suppressor genes to survival risk.

[0004] Therefore, existing methods generate all genomic features in a unified latent space, resulting in generation results lacking functional specificity and further reducing the reliability of prediction. These limitations seriously hinder the application of multi-modal survival prediction technology in clinical practice. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-modal cancer survival prediction method based on a latent differentiation variational autoencoder, which compresses WSI representations through the VIB-Trans module to solve computational redundancy, uses LD-VAE to generate functional-specific genomic embeddings, models biological heterogeneity, and combines the PoE technology and alignment loss to achieve multi-modal joint distribution estimation under missing data, and finally outputs high-precision and interpretable survival risk predictions through the co-attention mechanism to improve clinical applicability.

[0006] To achieve the above purpose, the present invention provides a multi-modal cancer survival prediction method based on a latent differentiation variational autoencoder, including the following steps: Step S1, preprocess and extract features from multi-modal data: input the whole-slide digital section WSI and genomic data of a patient, perform tissue region segmentation on the WSI to extract instance-level pathological features of non-overlapping patches, and group and extract grouped features from the genomic data according to functional categories; Step S2. Implement pathological feature compression based on the variational information bottleneck transformation theory: Compress the high-dimensional features of the WSI through the information bottleneck theory and the Nystrom attention mechanism to generate the latent distribution of the compressed pathological features. Step S3. Generate function-specific genomic embeddings through the latent differentiation variational autoencoder LD-VAE: Learn the latent distribution of genomic data through global posterior learning, use the function differentiation network to generate specific latent variables for different functional categories, and reconstruct the missing genomic features. Step S4. Multi-modal joint distribution learning and alignment: Integrate the pathological and genomic posteriors based on the product of experts PoE technique, and introduce the alignment loss to constrain the consistency of the posterior distribution. Step S5. Multi-modal fusion and survival prediction: Screen the survival-related features through the co-attention mechanism, and output the survival probability and risk stratification results based on the Cox model.

[0007] Preferably, in step S1, preprocess and extract features from the multi-modal data, and the specific process is as follows: Step S11. Input the whole-slide digital slice WSI of the patient, segment the tissue area of the WSI, and extract the instance-level pathological features of non-overlapping patches. Input the gigapixel-level WSI of the patient, use the OpenSlide library to detect the tissue area in the WSI, filter the non-tissue background, and at 20x magnification, segment the WSI into non-overlapping patches with a size of 224×224 pixels. Use the pre-trained Swin Transformer encoder CTransPath to extract the 1024-dimensional feature vector of each patch to obtain the instance-level feature set , as follows: ; where is the number of patches; is the pre-trained Swin Transformer encoder; is the j-th pathological slice patch of the k-th sample, i.e., the patient; is the m-th pathological slice patch of the k-th sample, i.e., the patient , and after being processed by the feature extraction function the obtained pathological feature; Step S12. Input the genomic data of the patient, group it according to functional categories, and extract the grouped features. Group the genomic data into 6 biological functions, including: tumor suppression, carcinogenesis, protein kinase, cell differentiation, transcription, and cytokines and growth; Perform Z-score normalization on each group of gene features, and extract the 256-dimensional embedding of each group of features through the self-normalizing neural network SNN to obtain the grouped feature set , as follows: ; Among them, is the number of biological function groups of genomic data; is the self-normalizing neural network SNN; is the original data of the i-th type of functional genome of the k-th patient; The 256-dimensional embedding features of the N-th type of functional genome extracted from the k-th patient by the self-normalizing neural network SNN.

[0008] Preferably, in step S2, based on the variational information bottleneck transformation theory, pathological feature compression is realized, and the specific process is as follows: Step S21: Compress the WSI feature through the variational information bottleneck transformation VIB-Trans module, and minimize the mutual information between the compressed feature and the survival objective. The objective function is as follows: ; Among them, is the loss function of the variational information bottleneck transformation module; is the latent variable of the pathological feature; is the set of input instance-level pathological features; is the posterior distribution of the latent variable when the pathological feature Y is given; is the survival loss; is the survival time of the patient; is the censoring status, 0 indicates censoring, and 1 indicates the occurrence of an event; is the information bottleneck trade-off coefficient, is adjusted from 0.1 to 1.0 by cosine annealing; is the divergence term; is the spherical Gaussian prior ; Step S22: Adopt the Nystrom attention mechanism to reduce the computational complexity; Adopt Landmark sampling to approximate the self-attention matrix, and input the set of instance-level pathological features , uniformly sample m = 64 positions from the input sequence as Landmark points, and construct the approximate key matrix and the approximate value matrix , and adopt the Nystrom attention mechanism to reduce the computational complexity to linear; among them, the attention score is as follows: ; Among them, , , are the query, key, and value matrices of the attention mechanism respectively; is the scaling factor; Step S23: Introduce learnable mean tokens and variance tokens , and perform global token aggregation; Generate global latent distribution parameters and through a Transformer encoder, which follow distribution.

[0009] Preferably, in step S3, generate function-specific genomic embeddings through a latent differentiation variational autoencoder LD-VAE, and the specific process is as follows: Step S31: Learn the latent distribution of genomic data through global posterior learning; Train LD-VAE, input grouped genomic features , where is the functional group category of gene set partitioning; generate global posterior distribution parameters and through a Transformer encoder, which follow distribution; Step S32: Use a functional differentiation network to generate specific latent variables for different functional categories; Let satisfy Markov chain independence, and assign independent latent variables to each functional category . Map the global posterior to function-specific distribution parameters through an MLP as follows: ; where, and are the mean parameter and covariance parameter of the i-th class of functional genomic latent variable respectively; and are two-layer MLPs with an input dimension of 128, a hidden layer of 64, and an output dimension of 128, which follow distribution; Step S33: Reconstruct missing genomic features; Sample from through the reparameterization technique as follows: ; where, is a random variable following a standard normal distribution, ; Use a fully connected decoder with an input of 128 and an output of 256, from Generate reconstructed features , and the optimization objective is the extended ELBO function, as follows: ; where, is the evidence lower bound loss function; is the variational posterior distribution of the latent variable z given the genomic feature X and the pathological feature Y; is all the latent variables learned from genomic data; is the likelihood function of reconstructing the genomic feature under the condition of the latent variable and the pathological feature Y; ; and are both standard Gaussian priors .

[0010] Preferably, in step S4, the multimodal joint distribution learning and alignment are as follows: Step S41: Integrate the pathological and genomic posteriors based on the product of experts (PoE) technique; Assume that the pathological and genomic data are conditionally independent, and the joint posterior distribution is as follows: ; The Gaussian distribution parameters are as follows: Mean parameter: ; Covariance matrix: ; where, and represent the mean and covariance of the prior distribution p(z) respectively, and the prior distribution = ; and are the mean and covariance parameters of the genomic latent distribution respectively; and are the mean and covariance parameters of the pathological latent distribution respectively; Step S42: Introduce an alignment loss to constrain the consistency of the posterior distribution; Introduce an alignment loss function to minimize the Wasserstein distance between the pathological and genomic posteriors; where, the alignment loss function is as follows: ; The total loss function is as follows: ; where, is the overall loss function; is the survival loss; is the variational information bottleneck loss; is the latent divergence variational autoencoder loss; is the alignment loss of the pathological modality - genomic modality latent feature distribution; is the alignment loss weight; The alignment loss function is optimized by dynamically adjusting the alignment loss weight during the training phase and the optimal weight is determined by grid search .

[0011] Preferably, in step S5, the multi-modal fusion and survival prediction are as follows: Step S51: Screen survival-related features through the co-attention mechanism; Through the co-attention mechanism, calculate the correlation weights between the WSI and genomic features as follows: ; where is the generated genomic feature; is the generated pathological feature; is the learnable parameter matrix; is the scaling factor; calculate the co-attention weights at the patch level, screen the Top-5 high-weight patches, and visualize the corresponding WSI regions; Step S52: Output the survival probability and risk stratification results; In survival prediction, generate the fusion feature through global attention pooling as follows: ; Then, based on the Cox proportional hazards model, define the survival function and risk function as follows: ; ; where is the survival time of the patient; t and are the specific time points of the survival time; is the multi-modal feature set of the patient; is the survival function; is the risk function; is the learnable parameter; is the baseline risk function, estimated non-parametrically by the Breslow estimator; is the fused multi-modal feature; Through the survival risk score and the median risk of the training set are compared, and the patients are divided into high-risk, i.e., and low-risk, i.e., Group; wherein, the threshold value Is determined by the median risk of the training set.

[0012] Therefore, the present invention adopts the above-mentioned multi-modal cancer survival prediction method based on the latent differentiation variational autoencoder, compresses the WSI representation through the VIB-Trans module, solves the computational redundancy, generates function-specific genomic embeddings using LD-VAE, models biological heterogeneity, and combines the PoE technology and alignment loss to realize the multi-modal joint distribution estimation under missing data. Finally, through the co-attention mechanism, a high-precision and interpretable survival risk prediction is output, improving the clinical applicability.

[0013] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Is a schematic flowchart of the multi-modal cancer survival prediction method based on the latent differentiation variational autoencoder of the present invention; Figure 2 Is a detailed model diagram of the multi-modal cancer survival prediction method based on the latent differentiation variational autoencoder of the present invention; Figure 3 Is a detailed architecture diagram of VIB-Trans and LD-VAE in the model of the present invention; wherein, (a) is the detailed architecture of VIB-Trans; (b) is the detailed architecture of LD-VAE; Figure 4 Is a schematic diagram of the consistency of the co-attention weights of the real and generated features of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0016] As Figure 1 and Figure 2 shown, the multi-modal cancer survival prediction method based on the latent differentiation variational autoencoder includes the following steps: Step S1, input the whole-slide digital slice (WSI) and genomic data of the patient, perform tissue region segmentation on the WSI to extract instance-level pathological features of non-overlapping patches, and group the genomic data according to functional categories to extract grouped features; Step S2, compress the high-dimensional features of the WSI through the information bottleneck theory and the Nystrom attention mechanism to generate a compressed latent distribution of pathological features; Step S3, learn the latent distribution of genomic data through global posterior learning, use the functional differentiation network to generate specific latent variables of different functional categories, and reconstruct the missing genomic features; Step S4: Integrate the pathology and genomic posterior using the Product of Experts (PoE) technique, and introduce an alignment loss to constrain the consistency of the posterior distribution. Step S5: Screen the survival-related features through a co-attention mechanism, and output the survival probability and risk stratification results based on the Cox model.

[0017] Embodiment

[0018] Step S1: Preprocess and extract features from the multi-modal data.

[0019] Step S11: Input the whole-slide digital sections (WSIs) of the patient, segment the tissue regions of the WSIs, and extract the instance-level pathological features of the non-overlapping patches.

[0020] Input the gigapixel-level WSIs of the patient, in the format of SVS or NDPI. Use the OpenSlide library to detect the tissue regions in the WSIs, filter the non-tissue background, and at 20x magnification, segment the WSIs into non-overlapping patches of size 224×224 pixels. On average, about 10,000 patches are generated for each patient. Use the pre-trained Swin Transformer encoder (CTransPath) to extract the 1024-dimensional feature vectors of each patch, and obtain the instance-level feature set , as follows: ; where is the number of patches; is the pre-trained Swin Transformer encoder; is the j-th pathological section patch of the k-th sample, i.e., the patient; is for the m-th pathological section patch of the k-th sample, i.e., the patient , and after being processed by the feature extraction function the obtained pathological feature.

[0021] Step S12: Input the genomic data of the patient, group it by functional category, and extract the grouped features.

[0022] According to the MSigDB Hallmark gene set, group the genomic data into 6 categories of biological functions, including tumor suppression, carcinogenesis, protein kinases, cell differentiation, transcription, and cytokines and growth. Perform Z-score normalization on each group of gene features to eliminate the dimension difference. Extract the 256-dimensional embeddings of each group of features through a self-normalizing neural network (SNN), and obtain the grouped feature set , as follows: ; where is the number of biological function groups of the genomic data; is a self - normalized neural network SNN; The original functional genomic data of the i - th type for the k - th patient; is the 256 - dimensional embedding feature of the N - th type of functional genome extracted by the SNN for the k - th patient.

[0023] Step S2: Based on the variational information bottleneck transformation theory, realize pathological feature compression.

[0024] Step S21: Compress the WSI feature through the variational information bottleneck transformation (VIB - Trans) module, as shown in Figure 3 (a) below.

[0025] Minimize the mutual information between the compressed feature and the survival objective based on the variational information bottleneck theory. The objective function is as follows: ; where, is the loss function of the variational information bottleneck transformation module; is the latent variable of the pathological feature; is the set of input instance - level pathological features; is the posterior distribution of the latent variable given the pathological feature Y; is the survival loss; t is the survival time of the patient; c is the censoring status, 0 indicates censoring, and 1 indicates the occurrence of an event; is the information bottleneck trade - off coefficient, which is adjusted from 0.1 to 1.0 by cosine annealing; is the divergence term; is the spherical Gaussian prior .

[0026] Step S22: Adopt the Nystrom attention mechanism to reduce the computational complexity.

[0027] Adopt Landmark sampling to approximate the self - attention matrix. Input the set of instance - level pathological features Y, uniformly sample m = 64 positions from the input sequence as Landmark points, and construct the approximate key matrix and the approximate value matrix to reduce the computational complexity to linear.

[0028] where the attention score is as follows: ; where, , , are the query, key, and value matrices of the attention mechanism respectively; is the scaling factor; and A subset of key-value pairs selected by uniform sampling.

[0029] Step S23: Introduce learnable mean tokens and variance tokens for global token aggregation.

[0030] Generate global latent distribution parameters and through a Transformer encoder, which follow a distribution.

[0031] Step S3: Generate function-specific genomic embeddings through a Latent Differentiable Variational Autoencoder (LD-VAE) as Figure 3 shown in (b) therein.

[0032] Step S31: Learn the latent distribution of genomic data through global posterior.

[0033] Train a Latent Differentiable Variational Autoencoder (LD-VAE) with input grouped genomic features , where is the functional group category of gene set partitioning; generate global posterior distribution parameters and through a Transformer encoder, which follow a distribution.

[0034] Step S32: Generate specific latent variables for different functional categories using a functional differentiation network.

[0035] Let satisfy Markov chain independence, and assign independent latent variables to each functional category . Map the global posterior to function-specific distribution parameters through an MLP as follows: ; where and are the mean parameter and covariance parameter of the i-th class functional genomic latent variable respectively; and are two-layer MLPs with an input dimension of 128, a hidden layer of 64, and an output dimension of 128, which follow a distribution.

[0036] Step S33: Reconstruct the missing genomic features.

[0037] Sample from using the reparameterization technique as follows: ; Among them, is a random variable subject to the standard normal distribution, .

[0038] Using a fully connected decoder , whose input is 128 and output is 256, to generate reconstructed features from , and the optimization objective is the extended ELBO function, as follows: ; Among them, is the evidence lower bound loss function; is the variational posterior distribution of the latent variable z given the genomic feature X and the pathological feature Y; is all the latent variables learned from genomic data; is the likelihood function of the reconstructed genomic feature under the conditions of the latent variable and the pathological feature Y; ; and are both standard Gaussian priors .

[0039] Step S4, Multi-modal joint distribution learning and alignment.

[0040] Step S41, Integrate the pathological and genomic posteriors based on the Product of Experts (PoE) technique. Assume that the pathological and genomic data are conditionally independent, and the joint posterior distribution is as follows: ; The Gaussian distribution parameters are as follows: Mean parameter: ; Covariance matrix: ; Among them, and respectively represent the mean and covariance of the prior distribution p(z), and the prior distribution = ; and are respectively the mean and covariance parameters of the genomic latent distribution; and are respectively the mean and covariance parameters of the pathological latent distribution.

[0041] Step S42, Introduce an alignment loss to constrain the consistency of the posterior distribution, as Figure 4 shown.

[0042] An alignment loss function is introduced to minimize the Wasserstein distance between the pathology and genomic posterior. The alignment loss function is as follows: ; The total loss function is as follows: ; where is the overall loss function; is the survival loss; is the variational information bottleneck loss; is the latent differentiation variational autoencoder loss; is the alignment loss of the latent feature distribution between the pathology modality and the genomic modality; is the alignment loss weight.

[0043] The alignment loss function is optimized by dynamically adjusting the alignment loss weight during the training phase and the optimal weight is determined through grid search .

[0044] Step S5: Perform multimodal fusion and survival prediction.

[0045] Step S51: Screen survival-related features through the co-attention mechanism.

[0046] Through the co-attention mechanism, calculate the correlation weights between the WSI and genomic features as follows: ; where is the generated genomic feature; is the generated pathology feature; is the learnable parameter matrix; is the scaling factor; calculate the co-attention weights at the patch level, screen the Top-5 high-weight patches, and visualize the corresponding WSI regions.

[0047] Step S52: Output the survival probability and risk stratification results.

[0048] In survival prediction, generate fusion features through global attention pooling as follows: ; Then, based on the Cox proportional hazards model, define the survival function and risk function as follows: ; ; where is the survival time of the patient; t and are the specific time points of the survival time; is the multi-modal feature set of the patient; is the survival function; is the risk function; is the learnable parameter; is the baseline risk function, non-parametrically estimated by the Breslow estimator; is the fused multi-modal feature.

[0049] Through the survival risk score and the median risk of the training set are compared, and the patients are divided into high-risk, that is and low-risk, that is groups; among them, the threshold is determined by the median risk of the training set.

[0050] Therefore, the present invention adopts the above-mentioned multi-modal cancer survival prediction method based on the latent differentiation variational autoencoder, compresses the WSI representation through the VIB-Trans module, solves the computational redundancy, generates function-specific genomic embeddings using LD-VAE, models biological heterogeneity, and combines the PoE technology and the alignment loss to achieve the multi-modal joint distribution estimation under missing data. Finally, through the co-attention mechanism, it outputs high-precision and interpretable survival risk prediction, improving the clinical applicability.

[0051] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multimodal cancer survival prediction method based on a latent differentiation variational autoencoder, characterized in that It includes the following steps: Step S1, preprocess and extract features from multimodal data: Input the whole-slide digital image (WSI) and genomic data of the patient, segment the tissue regions of the WSI to extract instance-level pathological features of non-overlapping patches, and group the genomic data by functional categories to extract grouped features; Step S2, achieve pathological representation compression based on the variational information bottleneck transformation theory: Compress the high-dimensional features of the WSI through the information bottleneck theory and the Nystrom attention mechanism to generate a compressed latent distribution of pathological features; Step S3, generate function-specific genomic embeddings through the latent differentiation variational autoencoder (LD-VAE): Learn the latent distribution of genomic data through global posterior, and use the function differentiation network to generate specific latent variables for different functional categories to reconstruct the missing genomic features; Step S4, multimodal joint distribution learning and alignment: Integrate the pathological and genomic posteriors based on the product of experts (PoE) technique, and introduce an alignment loss to constrain the consistency of the posterior distributions; Step S5, multimodal fusion and survival prediction: Screen survival-related features through the co-attention mechanism, and output the survival probability and risk stratification results based on the Cox model.

2. The multimodal cancer survival prediction method based on the latent differentiation variational autoencoder according to claim 1, wherein In step S1, preprocess and extract features from multimodal data, and the specific process is as follows: Step S11, input the whole-slide digital image (WSI) of the patient, segment the tissue regions of the WSI, and extract instance-level pathological features of non-overlapping patches; Input the gigapixel-level WSI of the patient, use the OpenSlide library to detect the tissue regions in the WSI, filter out the non-tissue background, at 20x magnification, segment the WSI into non-overlapping patches of size 224×224 pixels, use the pre-trained Swin Transformer encoder CTransPath to extract the 1024-dimensional feature vectors of each patch, and obtain the instance-level feature set , as follows: ; Among them, is the number of patches; is the pre-trained Swin Transformer encoder; is the j-th pathological slice patch of the k-th sample, i.e., the patient; is for the m-th pathological slice patch of the k-th sample, i.e., the patient , and the pathological feature obtained after being processed by the feature extraction function ; Step S12, input the genomic data of the patient, group it by functional categories, and extract grouped features; Grouping the genomic data into 6 categories of biological functions includes: tumor suppression, carcinogenesis, protein kinase, cell differentiation, transcription, and cytokines and growth; Z-score normalization is performed on each group of gene features, and 256-dimensional embeddings of each group of features are extracted through a self-normalizing neural network SNN to obtain a grouped feature set , as follows: ; Among them, is the number of biological function groups of genomic data; is the self-normalizing neural network SNN; is the original functional genomic data of the i-th type of the k-th patient; The 256-dimensional embedding feature of the N-th type of functional genome extracted by the self-normalizing neural network SNN for the k-th patient.

3. The multimodal cancer survival prediction method based on the potential differentiation variational autoencoder according to claim 1, wherein In step S2, based on the variational information bottleneck transformation theory, achieve pathological representation compression, and the specific process is as follows: Step S21, compress the WSI representation through the variational information bottleneck transformation (VIB-Trans) module, and minimize the mutual information between the compressed representation and the survival target. The objective function is as follows: ; Among them, is the loss function of the variational information bottleneck transformation module; is the latent variable of the pathological feature; is the set of input instance-level pathological features; is the posterior distribution of the latent variable when the pathological feature Y is given; is the survival loss; is the survival time of the patient; is the censoring status, 0 indicates censoring, and 1 indicates the occurrence of an event; is the information bottleneck trade-off coefficient, which is adjusted from 0.1 to 1.0 by cosine annealing; is the divergence term; is the spherical Gaussian prior ; Step S22, adopt the Nystrom attention mechanism to reduce the computational complexity; Adopt Landmark sampling to approximate the self-attention matrix, and input the set of instance-level pathological features , uniformly sample m = 64 positions from the input sequence as Landmark points, and construct an approximate key matrix and an approximate value matrix , and adopt the Nystrom attention mechanism to reduce the computational complexity to linear; among them, the attention scores are shown as follows: ; Among them, , , are the query, key, and value matrices of the attention mechanism respectively; is the scaling factor; Step S23: Introduce learnable mean tokens and variance tokens , and perform global token aggregation; Generate global latent distribution parameters through the Transformer encoder and , subject to distribution.

4. The multimodal cancer survival prediction method based on the latent differentiation variational autoencoder according to claim 1, wherein In step S3, generate function-specific genomic embeddings through the latent differentiation variational autoencoder (LD-VAE), and the specific process is as follows: Step S31, learn the latent distribution of genomic data through global posterior; Train LD-VAE with the input of grouped genomic features , where is the functional group category of gene set partitioning; generate the global posterior distribution parameters and , which follow distribution; Step S32, use the function differentiation network to generate specific latent variables for different functional categories; Let satisfy the Markov chain independence, and assign independent latent variables to each functional category , and map the global posterior to the functional-specific distribution parameters through the MLP as follows: ; Among them, and are the mean parameter and covariance parameter of the latent variable of the i-th class of functional genome respectively; ; and are two-layer MLP with an input dimension of 128, a hidden layer of 64, and an output dimension of 128, following distribution; Step S33, reconstruct the missing genomic features; Sampling is performed from using the reparameterization technique as follows: ; wherein, is a random variable subject to a standard normal distribution, ; Using a fully connected decoder , with an input of 128 and an output of 256, to generate reconstructed features , and the optimization objective is the extended ELBO function, as follows: ; Among them, is the evidence lower bound loss function; is the variational posterior distribution of the latent variable z given the genomic feature X and the pathological feature Y; is all latent variables learned from genomic data; is under the latent variable and the pathological feature Y, the likelihood function of reconstructing the genomic feature ; ; and are both standard Gaussian priors .

5. The multimodal cancer survival prediction method based on a latent differentiation variational autoencoder according to claim 1, wherein In step S4, multimodal joint distribution learning and alignment, and the specific process is as follows: Step S41, integrate the pathological and genomic posteriors based on the product of experts (PoE) technique; Assume that the pathological and genomic data are conditionally independent, and the joint posterior distribution is as follows: ; The Gaussian distribution parameters are as follows: Mean parameter: ; Covariance matrix: ; wherein, and respectively represent the mean and covariance of the prior distribution p(z), and the prior distribution = ; and are respectively the mean and covariance parameters of the genomic latent distribution; and are respectively the mean and covariance parameters of the pathological latent distribution; Step S42, introduce an alignment loss to constrain the consistency of the posterior distributions; Introduce an alignment loss function to minimize the Wasserstein distance between the pathological and genomic posteriors; where the alignment loss function is as follows: ; The total loss function is as follows: ; Among them, is the overall loss function; is the survival loss; is the variational information bottleneck loss; is the latent divergence variational autoencoder loss; is the alignment loss of the pathological modality-genomic modality latent feature distribution; is the alignment loss weight.

6. The multimodal cancer survival prediction method based on the latent differentiation variational autoencoder according to claim 1, characterized in that In step S5, multimodal fusion and survival prediction, and the specific process is as follows: Step S51: Screen survival-related features through the co-attention mechanism; Through the co-attention mechanism, calculate the correlation weights between the WSI and genomic features as follows: ; Among them, is the generated genomic feature; is the generated pathological feature; is the learnable parameter matrix; is the scaling factor; calculate the co-attention weights at the patch level, screen the Top-5 high-weight patches, and visualize their corresponding WSI regions; Step S52: Output the survival probability and risk stratification results; In survival prediction, generate fused features through global attention pooling as follows: ; Then, based on the Cox proportional hazards model, define the survival function and risk function as follows: ; ; wherein, is the survival time of the patient; t and are the specific time points of the survival time; is the multi-modal feature set of the patient; is the survival function; is the hazard function; are the learnable parameters; is the baseline hazard function, non-parametrically estimated by the Breslow estimator; is the fused multi-modal feature; By survival risk score Compared with the median risk of the training set Patients were divided into high-risk, i.e., And low-risk, i.e., Groups; among them, the threshold Is determined by the median risk of the training set.

Citation Information

Cited By

  • Aging prediction method and device based on age enhancement, equipment and storage medium

    CN120998505A

  • Clinical multi-mode cancer drug response prediction method based on feature reconstruction

    CN121528291A

  • Multi-modal fusion method based on optimal transmission and information difference guidance

    CN122153818A