Multi-omics tumor immune microenvironment panoramic analysis method based on HVAE and Cross-Attention
Through the multiomic data dimension reduction/fusion network and contrast learning method based on HVAE and Cross-Attention, the problems of data heterogeneity and cross-omic information loss in tumor immune microenvironment research were solved, and efficient fusion of multiomic data and accurate evaluation of panoramic characteristics of tumor immune microenvironment were achieved, which improved analysis efficiency and accuracy.
Patent Information
- Application Number
- CN202510257121.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
AI Technical Summary
Existing tumor immune microenvironment research faces the problems of data heterogeneity, lack of cross-omic information and systematic insufficient, and it is difficult to effectively integrate multi-omic data and learn the relationship between gene regulation, mutation information and clinical data.
A multi-omics data dimensionality reduction/fusion network based on HVAE and Cross-Attention is used to perform feature dimensionality reduction and fusion of multi-omics data through hierarchical variational autoencoder and cascade network modules to obtain optimized feature fusion latent variables, and optimize the discriminant ability of latent variables through comparative learning.
It enhances the ability of tumor cross-omic information interaction, avoids the lack of cross-omic information, improves the generalization ability of multi-omic data, makes the panoramic characteristics of tumor immune microenvironment more discriminative, and realizes comprehensive evaluation of immune cell infiltration, immune checkpoint gene expression, mutation load, etc., and improves the efficiency and accuracy of panoramic analysis of multi-omic tumor immune microenvironment.
Smart Images

Figure CN120089207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor immune microenvironment analysis, and particularly to a multi-omics tumor immune microenvironment panoramic analysis method based on HVAE and Cross-Attention. Background Art
[0002] The occurrence and development of cancer are affected by multi-level biological factors, among which the tumor immune microenvironment (TME) plays a crucial role in tumor progression, treatment response, and drug resistance.
[0003] However, the existing research on the tumor immune microenvironment faces the following challenges:
[0004] 1) Data heterogeneity: Different omics data have different data formats and statistical characteristics, making it difficult to directly fuse.
[0005] 2) Lack of cross-omics information: Existing methods generally have the problem of cross-omics information loss even if they can perform multi-omics fusion, and it is difficult to effectively learn the relationships between gene regulation (RNA-seq, DNA methylation), mutation information (SNV, CNV), and clinical data.
[0006] 3) Insufficient systematicness in TME analysis: Current tumor immune microenvironment analysis methods often only focus on a single data modality (such as RNA-Seq) and lack comprehensive and systematic analysis means. Summary of the Invention
[0007] The present invention provides a multi-omics tumor immune microenvironment panoramic analysis method based on HVAE and Cross-Attention to overcome the above technical problems.
[0008] To achieve the above object, the technical solution of the present invention is:
[0009] A multi-omics tumor immune microenvironment panoramic analysis method based on HVAE and Cross-Attention specifically includes the following steps:
[0010] S1: Obtain and preprocess multi-omics data on the tumor immune microenvironment;
[0011] And the multi-omics data at least includes DNA methylation data, RNA-Seq data, SNV data, CNV data, and Clinical clinical data;
[0012] S2: Construct a multi-omics data dimensionality reduction / fusion network based on HVAE and Cross-Attention to obtain feature fusion latent variables;
[0013] And the constructed multi-omics data dimensionality reduction / fusion network includes several encoder modules based on the hierarchical variational autoencoder HVAE and a cascaded network module based on Cross Attention;
[0014] Each encoder module is respectively used to perform feature dimensionality reduction encoding on different data types in the multi-omics data to obtain multi-omics latent variables, and store them in a pre-partitioned latent variable classification layer based on the multi-omics latent variables according to the data type;
[0015] The cascaded network module is used to perform feature fusion on the data in the latent variable classification layer to obtain feature fusion latent variables;
[0016] S3: Perform contrastive learning on the feature fusion latent variables to obtain optimized feature fusion latent variables;
[0017] S4: Extract the TME panoramic features regarding the tumor immune microenvironment from the optimized feature fusion latent variables;
[0018] And the TME panoramic features include immune cell infiltration features, immune checkpoint gene expression features, and tumor mutation burden features;
[0019] S5: Evaluate and obtain the immune cold / hot state of the tumor according to the TME panoramic features, and perform survival risk analysis and prediction on the patient's body based on the immune cold / hot state of the tumor;
[0020] S6: Based on the results of the survival risk analysis and prediction, customize a personalized immune microenvironment portrait for the individual patient for personalized identification of the tumor immune microenvironment state.
[0021] Furthermore, the method for obtaining the feature fusion latent variables in S2 specifically includes the following steps:
[0022] S21: Perform feature dimensionality reduction encoding on different data types in the multi-omics data through the encoder module to obtain multi-omics latent variables;
[0023] S22: Denote the DNA methylation latent variable obtained by feature dimensionality reduction encoding as z 1 、the RNA-Seq latent variable as z 2 、the SNV latent variable as z 3 、the CNV latent variable as z 4 and the Clinical latent variable as z 5 ; and store them in a pre-partitioned latent variable classification layer based on the multi-omics latent variables according to the data type, so as to store the multi-omics latent variables after feature dimensionality reduction encoding in the corresponding latent variable classification layer;
[0024] The latent variable classification layer includes a gene layer, a mutation layer, and a clinical layer;
[0025] Input the data of the latent variable classification layer into the cascaded network module;
[0026] S23: Use the RNA-Seq latent variable z 2 as the input of the Query vector function in the first Cross Attention module S1 in the cascaded network module; Use the DNA methylation latent variable z 1 as the input of the Key and Value vector functions in the first Cross Attention module S1 to obtain and output the gene regulation relationship vector z s1 ;
[0027] Use the CNV latent variable z 4 as the input of the Query vector function in the second Cross Attention module S2 in the cascaded network module; Use the SNV latent variable z 3 as the input of the Key and Value vector functions in the second Cross Attention module S2 to obtain and output the gene variation relationship vector z s2 ;
[0028] S24: Use the gene variation relationship vector z s2 as the input of the Query vector function in the third Cross Attention module S3 in the cascaded network module; Use the gene regulation relationship vector z s1 as the input of the Key and Value vector functions in the third Cross Attention module S3 to obtain and output the genomic variation fusion vector z s3 ;
[0029] S25: Use the Clinical latent variable z 5 as the input of the Query vector function in the fourth Cross Attention module S4 in the cascaded network module; Use the genomic variation fusion vector z s3 as the input of the Key and Value vector functions in the fourth Cross Attention module S4 to obtain and output the feature fusion latent variable Z fusion .
[0030] Furthermore, the expression for feature dimensionality reduction encoding in S1 is
[0031]
[0032] where: q(z i |X i ) represents multiple omics latent variables obtained after feature dimensionality reduction encoding; X iRepresents the multi-omics data regarding the tumor immune microenvironment after preprocessing; z i Represents the low-level latent variable of the multi-omics data; μ i Represents the mean of the latent variables of the encoder; σ i Represents the variance of the latent variables of the encoder; i represents the category of the multi-omics data, namely DNA methylation data, RNA-Seq data, SNV data, CNV data, and Clinical clinical data;
[0033] In the cascade network module based on Cross Attention in S22, each Cross Attention module obtains and outputs the gene regulation relationship vector z s1 Or obtains and outputs the gene variation relationship vector z s2 Or obtains and outputs the genomic variation fusion vector z s3 Or obtains and outputs the feature fusion latent variable Z fusion The expression of
[0034]
[0035] In the formula: Q represents the Query vector function; V represents the Value vector function; K represents the Key vector function; d represents the dimension; T represents the transpose.
[0036] Furthermore, in S3, the method for obtaining the optimized feature fusion latent variable by performing contrastive learning on the feature fusion latent variable is as follows:
[0037] Obtains several output feature fusion latent variables Z fusion As data samples;
[0038] Based on the TME scoring system, the data samples are divided to obtain positive sample pairs and negative sample pairs;
[0039] And the method for dividing the data samples based on the TME scoring system includes a division rule model, and its expression is
[0040] (Z i ,Z j ),where TME(i) = TME(j)
[0041] (Z i ,Z k ),where TME(i) ≠ TME(k)
[0042] In the formula: (Z i ,Z j ) represents the similar sample pair in the data sample, that is, the positive sample pair; (Z i ,Z k) represents the sample pairs with differences in the data samples, i.e., negative sample pairs; i, j, k represent the sample numbers in the data samples;
[0043] Based on the pre-trained contrastive learning network, obtain the clustering clusters of the feature fusion latent variables according to the positive sample pairs and negative sample pairs, that is, realize the optimized acquisition of the feature fusion latent variables to obtain the optimized feature fusion latent variables;
[0044] Among them, the contrastive loss function during the pre-training of the contrastive learning network is
[0045]
[0046] In the formula: represents the contrastive loss function; S(Z i , Z j ) represents the cosine similarity of the positive sample pair; S(Z i , Z k ) represents the cosine similarity of the negative sample pair; τ represents the temperature parameter that controls the weight of contrastive learning.
[0047] Furthermore, the specific steps of S5 are as follows:
[0048] S51: Evaluate and obtain the immune cold / hot state of the tumor according to the TME panoramic features;
[0049] And the evaluation method for obtaining the immune cold / hot state of the tumor according to the TME panoramic features is:
[0050] Obtain the evaluation score TME Score of the TME panoramic features, and based on the preset scoring threshold, divide the tumor state into immune cold state and immune hot state according to the evaluation score TME Score;
[0051] And the acquisition formula of the evaluation score TME Score is
[0052] TME Score = α × M immune + β × M checkpoint + γ × TMB
[0053] In the formula: M immune represents the extracted immune cell infiltration feature matrix; M checkpoint represents the extracted immune checkpoint gene expression feature matrix; TMB represents the number of mutations of the tumor mutation burden feature; α, β, γ represent adjustable weight coefficients;
[0054] S52: Perform survival risk analysis and prediction based on the immune cold / hot state of the tumor;
[0055] The pan - characteristics of the tumor microenvironment (TME) corresponding to the cold / hot state of tumor immunity will be evaluated as the input of a preset survival analysis model to analyze and predict the survival risk of patients and obtain the survival risk prediction results;
[0056] And the expression for analyzing and predicting the survival risk of patients is
[0057] h(t|Z t )=h 0 (t)exp(gZ t )
[0058] In the formula: h(t|Z t ) represents the individual survival risk corresponding to the feature vector Z of the pan - analysis characteristics of the tumor at time t t ; h 0 (t) represents the baseline hazard rate; g represents the regression coefficient related to the feature vector Z t ; exp(gZ t ) represents the hazard ratio.
[0059] Furthermore, the method for customizing the personalized immune microenvironment portrait of a patient individual in S6 includes formulating a rule model, and its expression is
[0060] A = [B, C, D, E, F]
[0061] In the formula: A represents the portrait descriptor of the personalized immune microenvironment portrait; B represents the cell infiltration ratio of the immune cell infiltration characteristics of the tumor; C represents the Z - score of the immune checkpoint gene expression characteristics; D represents the mutant tumor mutation burden (TMB) of the tumor mutation burden characteristics; E represents the cold / hot state of the tumor immunity; F represents the individual survival risk level of the survival risk analysis and prediction.
[0062] Beneficial effects: The present invention provides a multi - omics tumor immune microenvironment pan - analysis method based on HVAE and Cross - Attention. By using a multi - omics data dimensionality reduction / fusion network based on HVAE and Cross - Attention for tumor multi - omics data fusion, the cross - omics information interaction ability of the tumor is enhanced, and the loss of cross - omics information is avoided; by using a method based on contrast learning to optimize the feature fusion latent variables, the generalization ability of multi - omics data is improved, making the pan - characteristics of the tumor immune microenvironment TME more discriminative; a complete TME analysis framework is proposed, that is, the cold / hot state of the tumor immunity is evaluated and obtained according to the pan - analysis characteristics, and the survival risk analysis and prediction are carried out based on the cold / hot state of the tumor immunity, so as to realize the customization of the personalized immune microenvironment portrait of patients, achieve comprehensive evaluations such as immune cell infiltration, immune checkpoint gene expression, and mutation burden, and improve the efficiency and accuracy of the multi - omics tumor immune microenvironment pan - analysis. Description of the Drawings
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Figure 1 It is a flowchart of a multi-omics tumor immune microenvironment panoramic analysis method based on HVAE and Cross-Attention of the present invention;
[0065] Figure 2 It is an architecture diagram of a multi-omics data dimensionality reduction / fusion network based on HVAE and Cross-Attention in this embodiment. Specific embodiments
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0067] This embodiment provides a multi-omics tumor immune microenvironment panoramic analysis method based on HVAE and Cross-Attention, as Figure 1 shown, specifically including the following steps:
[0068] S1: Obtain and preprocess multi-omics data on the tumor immune microenvironment;
[0069] And the multi-omics data at least includes DNA methylation data, RNA-Seq data, SNV data, CNV data, and Clinical clinical data;
[0070] Among them: DNA methylation data: Characterize the methylation status of genes and affect gene expression;
[0071] RNA-Seq data: Reflect the gene expression level;
[0072] SNV (Simple Nucleotide Variation) data: Describe single nucleotide variations and can affect gene function;
[0073] CNV (Copy Number Variation) data: refers to the copy number changes of chromosomal segments, which affect gene expression;
[0074] Clinical data: the age, gender, cancer stage, tumor size, etc. of patients;
[0075] In this embodiment, considering that multi-omics data has problems such as high dimensionality, heterogeneous data sources, high noise, and missing values, if directly input into a deep learning model, the following problems may occur:
[0076] Inconsistent data distribution: The numerical ranges and distribution characteristics of different omics data are different, affecting model training; High data dimensionality: Omics data usually contains thousands of features, and direct use will lead to large computational overhead and high model complexity; Missing data: Some omics data may have missing values due to reasons such as sequencing quality and sample limitations, affecting the stability of the model.
[0077] Furthermore, in this embodiment, preprocessing operations are performed on multi-omics data regarding the tumor immune microenvironment, which includes: (1) Standardization of DNA methylation data:
[0078] The essence of DNA methylation data is the β value between 0 and 1, and its value is usually small, and the methylation level change ranges of different genes are different; In this embodiment, in order to ensure the stability of the data in the model, normalization is required:
[0079] Using the min-max normalization method, the methylation level is adjusted to between 0 and 1, and its expression is
[0080]
[0081] In the formula: X' DNA represents the normalized data of DNA methylation; X DNA represents the original DNA methylation data; X min represents the minimum value of DNA methylation data; X max represents the maximum value of DNA methylation data;
[0082] In this embodiment, by filtering low-variation CpG sites, only important methylation sites are retained to improve the representativeness of the data;
[0083] (2) Standardization of RNA-Seq data:
[0084] RNA-Seq data represents the gene expression level, usually expressed in the form of FPKM / RPKM / TPM, but its numerical distribution often shows a long-tailed distribution and needs to be numerically transformed:
[0085] Take the Log transformation of the RNA-Seq data to reduce numerical differences, and its expression is
[0086] X' RNA = log 2 (X RNA + 1)
[0087] And normalize the X' value by Z-score to ensure the comparability of the expression levels of different genes: RNA where μ and σ represent the mean and standard deviation of the RNA-Seq data respectively;
[0088]
[0089] Among them, μ and σ represent the mean and standard deviation of the RNA-Seq data respectively;
[0090] (3) Standardization processing of SNV data and CNV data:
[0091] SNV data and CNV data mainly describe the gene mutation status and require specific processing methods:
[0092] In this embodiment, the SNV data adopts One-Hot encoding to convert different mutation types into discrete features, and at the same time adopts variant frequency normalization to ensure the consistency of the numerical range of mutant genes;
[0093] And the feature description of converting different mutation types into discrete features is
[0094] No mutation → (1,0,0,0)
[0095] Missense mutation → (0,1,0,0)
[0096] Nonsense mutation → (0,0,1,0)
[0097] Splice site mutation → (0,0,0,1)
[0098] In this embodiment, the CNV data X CNV Adopts log transformation processing, and at the same time normalizes the output X' after log transformation processing CNV to the range of -1 to 1 to avoid the influence of extreme values on the model;
[0099] And the expression of the CNV data adopting log transformation processing is
[0100] X' CNV = log 2 (X CNV + 1)
[0101] (4) Standardization processing of Clinical clinical data:
[0102] In this embodiment, different processing methods are adopted for different variable types of Clinical clinical data:
[0103] Categorical variables (such as cancer stage, gender): One-Hot encoding is adopted;
[0104] Numerical variables (such as age, tumor size): Z-score standardization is adopted:
[0105]
[0106] where μ and σ respectively represent the mean and standard deviation of clinical numerical variables;
[0107] S2: Construct a multi-omics data dimensionality reduction / fusion network based on HVAE and Cross-Attention to obtain feature fusion latent variables;
[0108] As Figure 2 shown, and the constructed multi-omics data dimensionality reduction / fusion network includes several encoder modules (Encoder) based on the hierarchical variational autoencoder HVAE and a cascaded network module based on Cross Attention;
[0109] Each encoder module is respectively used to perform feature dimensionality reduction encoding on different data types in the multi-omics data to obtain multi-omics latent variables, and store them in a pre-partitioned latent variable classification layer based on the multi-omics latent variables according to the data type;
[0110] The cascaded network module is used to perform feature fusion on the data in the latent variable classification layer to obtain feature fusion latent variables;
[0111] In a specific embodiment, the method for obtaining feature fusion latent variables in S2 specifically includes the following steps:
[0112] S21: Through the encoder module, perform feature dimensionality reduction encoding on different data types in the multi-omics data to obtain multi-omics latent variables;
[0113] In this embodiment, HVAE is used to process high-dimensional multi-omics data and map it to a low-dimensional latent variable space; that is, each type of omics data is respectively input into an independent VAE encoder to learn its low-level latent variable Z i ; and the expression for feature dimensionality reduction encoding is
[0114]
[0115] In the formula: q(z i |X i ) represents the multi-omics latent variables obtained after feature dimensionality reduction encoding; X iRepresents the multi-omics data (RNA-Seq, DNA methylation, SNV, CNV, clinical data) regarding the tumor immune microenvironment after preprocessing; z i Represents the low-level latent variables of the multi-omics data; μ i Represents the mean of the latent variables of the encoder; σ i Represents the variance of the latent variables of the encoder; i represents the categories of the multi-omics data, namely DNA methylation data, RNA-Seq data, SNV data, CNV data, and Clinical clinical data;
[0116] In this embodiment, each type of omics data is respectively converted into the corresponding low-level latent variable Z i , but at this time, different omics data are still separate and have not been fused; at this time, the latent variables of multiple omics are still independent and need to be further fused through a cascaded network module based on CrossAttention;
[0117] S22: Denote the DNA methylation latent variable obtained by dimensionality reduction encoding of features according to the multi-omics latent variables as z 1 , the RNA-Seq latent variable as z 2 , the SNV latent variable as z 3 , the CNV latent variable as z 4 , and the Clinical latent variable as z 5 ; and store them in the pre-partitioned latent variable classification layer based on the multi-omics latent variables according to the data type, so as to store the multi-omics latent variables after dimensionality reduction encoding of features in the corresponding latent variable classification layer;
[0118] The latent variable classification layer includes a gene layer, a mutation layer, and a clinical layer;
[0119] Input the data of the latent variable classification layer into the cascaded network module;
[0120] Among them, in this embodiment, considering the complex biological relationships among the multi-omics data, feature fusion is performed on them to obtain the feature fusion latent variables that fuse the relationships among different modality data;
[0121] For example: RNA-Seq data is directly related to DNA methylation, and DNA methylation affects gene expression;
[0122] SNV data and CNV data jointly affect gene mutation and expression;
[0123] Clinical clinical data is correlated with all other omics data;
[0124] Therefore, in this embodiment, a cascaded network module based on Cross Attention is used to obtain the high-level fusion latent variables that fuse the interaction relationships among different modality data, namely the shared latent variable Zfusion ;
[0125] S23: Input the RNA-Seq latent variable z 2 as the input of the Query vector function in the first Cross Attention module S1 of the cascaded network module; input the DNA methylation latent variable z 1 as the input of the Key and Value vector functions in the first Cross Attention module S1 to obtain and output the gene regulation relationship vector z s1 ;
[0126] Input the CNV latent variable z 4 as the input of the Query vector function in the second Cross Attention module S2 of the cascaded network module; input the SNV latent variable z 3 as the input of the Key and Value vector functions in the second Cross Attention module S2 to obtain and output the gene mutation relationship vector z s2 ;
[0127] S24: Input the gene mutation relationship vector z s2 as the input of the Query vector function in the third Cross Attention module S3 of the cascaded network module; input the gene regulation relationship vector z s1 as the input of the Key and Value vector functions in the third Cross Attention module S3 to obtain and output the genome mutation fusion vector z s3 ;
[0128] S25: Input the Clinical latent variable z 5 as the input of the Query vector function in the fourth Cross Attention module S4 of the cascaded network module; input the genome mutation fusion vector z s3 as the input of the Key and Value vector functions in the fourth Cross Attention module S4 to obtain and output the feature fusion latent variable Z fusion ;
[0129] Specifically, in this embodiment, the expression for obtaining and outputting the gene regulation relationship vector z s1 or obtaining and outputting the gene mutation relationship vector z s2 or obtaining and outputting the genome mutation fusion vector z s3 or obtaining and outputting the feature fusion latent variable Z fusion in each Cross Attention module of the cascaded network module based on Cross Attention is
[0130]
[0131] In the formula: Q represents the Query vector function; V represents the Value vector function; K represents the Key vector function; d represents the dimension; T represents the transpose;
[0132] In this implementation, the cascaded network module based on Cross Attention also includes a decoder (Decoder). The decoders (Decoder1 to Decoder5) reconstruct the omics data of each category according to the fused latent variable Z fusion and the reconstruction outputs include: DNA methylation, RNA-Seq, SNV, CNV, and clinical data. These reconstructed data are used to train the cascaded network module based on Cross Attention, so as to optimize the generation process of the latent variable. The training process of the cascaded network module based on Cross Attention is to preprocess the multi-omics data on the tumor immune microenvironment as feature data, use the corresponding data categories as labels, input them into the cascaded network module for training, and use the data output by the corresponding decoder (Decoder) as the prediction output. And based on the mean square error between the prediction output and the real data as the loss function, the network weights of the cascaded network module are adaptively adjusted by the backpropagation method to realize the pre-training process of the cascaded network module based on Cross Attention, which will not be elaborated here;
[0133] S3: Perform contrastive learning on the feature fusion latent variable to obtain an optimized feature fusion latent variable;
[0134] In this embodiment, after executing the multi-omics data dimensionality reduction / fusion network based on HVAE and Cross-Attention, a shared latent variable Z fusion is obtained, which contains the fusion information of DNA methylation, RNA-Seq, SNV, CNV, and clinical data, and integrates the features related to the tumor immune microenvironment (TME); however, the latent variable directly obtained from HVAE training may have the following problems:
[0135] The latent variable lacks discriminability: The latent variables of different cancer subtypes may overlap with each other, resulting in unclear classification and affecting individualized treatment and survival prediction;
[0136] Poor generalization ability for small sample data: When the data is small, the model may not be able to correctly learn the distribution characteristics of the latent variable, affecting the analysis of the immune state of TME;
[0137] The traditional KL divergence constraint on the latent variable is too loose: When training VAE, the KL divergence only constrains the latent variable to be close to the normal distribution, but does not actively enhance the aggregation ability of samples of the same category;
[0138] In this embodiment, a contrastive learning method is adopted to optimize the discriminative ability of latent variables by comparing the similarities between samples, so that:
[0139] The latent variables of similar samples (same cancer subtype, similar TME status) are clustered;
[0140] The latent variables of different categories of samples (different cancer subtypes, immune cold / hot status) are far apart;
[0141] The contrastive learning method not only improves the accuracy of prognosis prediction, but also strengthens the differential expression of the tumor immune microenvironment among different cancer patients;
[0142] In step S3 of this embodiment, contrastive learning is performed on the feature fusion latent variables to obtain a method for optimizing the feature fusion latent variables, specifically:
[0143] Obtain a number of output feature fusion latent variables Z fusion As data samples;
[0144] Based on the TME scoring system, the data samples are divided to obtain positive sample pairs and negative sample pairs;
[0145] Among them, TME represents a scoring system, and the positive sample relationship is determined by the TME score (such as the immune infiltration level estimated by x Cell, TIMER, CIBERSORT);
[0146] And the method for dividing the data samples based on the TME scoring system includes a division rule model, and its expression is
[0147] (Z i , Z j ), where TME(i) = TME(j)
[0148] (Z i , Z k ), where TME(i) ≠ TME(k)
[0149] In the formula: (Z i , Z j ) represents the similar sample pairs in the data samples, that is, the positive sample pairs (the sample pairs with similar TME, that is, the patient samples with the same cancer subtype and similar immune microenvironment); (Z i , Z k ) represents the sample pairs with differences in the data samples, that is, the negative sample pairs (the sample pairs with significant TME differences, that is, patients with different cancer subtypes or different immune cold / hot status); i, j, k represent the sample numbers in the data samples;
[0150] Based on a pre-trained contrastive learning network, clustering clusters of feature fusion latent variables are obtained according to positive sample pairs and negative sample pairs, that is, the optimization of feature fusion latent variables is achieved to obtain optimized feature fusion latent variables;
[0151] Among them, the contrastive loss function during the pre-training of the contrastive learning network is
[0152]
[0153] In the formula: represents the contrastive loss function; S(Z i , Z j ) represents the cosine similarity of the positive sample pair; S(Z i , Z k ) represents the cosine similarity of the negative sample pair; τ represents the temperature parameter that controls the weight of contrastive learning;
[0154] In this embodiment, the cosine similarity is used to calculate the latent variable similarity S(Z i , Z j ) of the same-class samples, and the similarity range is [-1, 1]; where 1 represents complete similarity (positive sample pair); -1 represents complete dissimilarity (negative sample pair); the Info NCE loss is used as the contrastive loss function to optimize the contrastive relationship between samples, S(Z i , Z j ): the similarity of the positive sample pair (should be close to 1);
[0155] S(Z i , Z k ): the similarity of the negative sample pair (should be close to -1);
[0156] τ: the temperature parameter that controls the weight of contrastive learning (usually takes 0.07 - 0.2);
[0157] The goal of this embodiment is to maximize the similarity S(Z i , Z j ) of the positive sample, making it close to 1; at the same time, minimize the similarity of the negative sample (Z i , Z k ), making it close to -1; finally, the latent variable Z t has both generative ability and retains TME difference information;
[0158] S4: Extract the TME panoramic features of the tumor immune microenvironment from the optimized feature fusion latent variables;
[0159] And the TME panoramic features include immune cell infiltration features, immune checkpoint gene expression features, and tumor mutation burden features;
[0160] The panoramic analysis of the tumor immune microenvironment in this embodiment includes:
[0161] Immune cell infiltration levels: such as CD8+ T cells, NK cells, etc.;
[0162] Immune checkpoint gene expression: such as PD-L1, CTLA-4, LAG3, etc.;
[0163] Tumor mutation burden (TMB);
[0164] Immune cold / heat status;
[0165] And extract the TME panoramic features regarding the tumor immune microenvironment, including:
[0166] (1) Extraction of immune cell infiltration features:
[0167] RNA-Seq data is used to evaluate the infiltration levels of immune cells in tumor tissues, using the following algorithms:
[0168] TIMER: Estimate the infiltration of immune cells (such as B cells, CD8+ T cells, macrophages);
[0169] CIBERSORT: Infer the proportions of 22 immune cell subtypes;
[0170] X Cell: Provide abundance scores for 64 immune and stromal cells;
[0171] To obtain the immune infiltration matrix M immune , and each patient correspondingly obtains an immune infiltration feature vector for identifying the degree of immune activity;
[0172] Its immune infiltration matrix M immune The expression is
[0173] M immune =[CD8+ T cells, NK cells, Macrophages,...]
[0174] (2) Extraction of immune checkpoint gene expression features:
[0175] Extract the expression levels of immune checkpoint genes from RNA-Seq data, including:
[0176] Inhibitory checkpoints: PD-1, PD-L1, CTLA-4, LAG3, TIM-3, IDO1;
[0177] Stimulatory checkpoints: CD28, OX40, CD137;
[0178] Obtain the Z-score of each gene to get the immune checkpoint feature matrix M checkpoint, for evaluating the potential response of a patient to immunotherapy with immune checkpoint inhibitors;
[0179] and the immune checkpoint feature matrix M checkpoint has the expression
[0180] M checkpoint = [PD-L1, CTLA-4, LAG3, …]
[0181] (3) Tumor mutation burden feature extraction:
[0182] Based on the SNV data, obtain the number of mutations per million base pairs, i.e., the tumor mutation burden TMB; and patients with high TMB and MSI-H are usually more sensitive to immune checkpoint inhibitors;
[0183] and the formula for obtaining the tumor mutation burden TMB is
[0184]
[0185] S5: Evaluate and obtain the immune cold / hot state of the tumor based on the panoramic analysis features, and perform survival risk analysis and prediction on it based on the immune cold / hot state of the tumor;
[0186] Specifically, it includes the following steps:
[0187] S51: Evaluate and obtain the immune cold / hot state of the tumor based on the TME panoramic features, and
[0188] Immune hot state: High infiltration of immune cells, high expression of immune checkpoint genes, high TMB, MSI-H;
[0189] Immune cold state: Low immune infiltration, low expression of immune checkpoint genes, low TMB, MSI-L;
[0190] and the method for obtaining the immune cold / hot state of the tumor based on the TME panoramic features is:
[0191] Obtain the evaluation score TME Score of the TME panoramic features, and based on a preset scoring threshold, divide the tumor state into immune cold state and immune hot state according to the evaluation score TME Score;
[0192] and the formula for obtaining the evaluation score TME Score is
[0193] TME Score = α × M immune + β × M checkpoint + γ × TMB
[0194] where: M immune represents the extracted immune cell infiltration feature matrix; M checkpointIt represents the extracted immune checkpoint gene expression feature matrix; TMB represents the number of mutations of the tumor mutation burden feature; α, β, and γ represent adjustable weight coefficients;
[0195] S5: Perform survival risk analysis and prediction based on the immune cold / hot state of the tumor;
[0196] That is, evaluate the tumor panorama analysis features (TME panorama features) corresponding to the immune cold / hot state of the tumor, as the input of the preset survival analysis model, and use the Cox proportional hazards model to calculate the immune-related survival risk, so as to perform survival risk analysis and prediction on patients to obtain the survival risk prediction result;
[0197] And the expression for performing survival risk analysis and prediction on patients is
[0198] h(t|Z t ) = h 0 (t)exp(gZ t )
[0199] In the formula: h(t|Z t ) represents the individual survival risk corresponding to the feature vector Z of the tumor panorama analysis feature at time t t ; h 0 (t) represents the baseline hazard rate, that is, the basic risk of all patients at time t without considering the features; g represents the regression coefficient related to the feature vector Z t , that is, it represents the impact of each feature on the survival risk; exp(gZ t ) represents the hazard ratio, that is, it represents the relative impact of a certain feature on the survival risk;
[0200] S6: Based on the results of the survival risk analysis and prediction, customize the personalized immune microenvironment portrait for individual patients to be used for the personalized identification of the tumor immune microenvironment state;
[0201] In a specific embodiment, the method for customizing the personalized immune microenvironment portrait for individual patients in S6 includes formulating a rule model, and its expression is
[0202] A = [B, C, D, E, F]
[0203] In the formula: A represents the portrait descriptor of the personalized immune microenvironment portrait; B represents the cell infiltration ratio of the immune cell infiltration feature of the tumor; C represents the Z-score of the immune checkpoint gene expression feature; D represents the mutated TMB of the tumor mutation burden feature; E represents the immune cold / hot state of the tumor; F represents the individual survival risk level of the survival risk analysis and prediction;
[0204] In this embodiment, based on the TME panorama features of each patient, a personalized immune microenvironment portrait is generated, including:
[0205] B Immune infiltration level (such as the infiltration ratio of CD8+ T cells, NK cells);
[0206] C Immune checkpoint gene expression (such as the expression levels of PD-L1, CTLA-4);
[0207] D Mutation burden (TMB);
[0208] E Immune cold / hot tumor state;
[0209] F Risk level of survival prediction;
[0210] For example: Example output of the immune microenvironment portrait (A: Patient ID: P001);
[0211] B: CD8+ T cells: 18.2%;
[0212] C: PD-L1 expression: High (Z-score: 2.3);
[0213] D: TMB: 22 mutations / Mb;
[0214] E: TME Score: 7.8 (immune hot);
[0215] F: Survival risk level: Low;
[0216] Through the personalized immune microenvironment portrait customization in this example, it is possible to accurately reveal the tumor immune microenvironment state, support immunotherapy decision-making, and provide personalized survival prediction, realizing a precision medicine solution of "immunomics + survival prediction".
[0217] In this example, tumor multi-omics data fusion is performed through a multi-omics data dimensionality reduction / fusion network based on HVAE and Cross-Attention, enhancing the cross-omics information interaction ability of tumors and avoiding the loss of cross-omics information; through a method of optimizing feature fusion latent variables based on contrastive learning, the generalization ability of multi-omics data is improved, making the panoramic features TME of the tumor immune microenvironment more discriminative; a complete TME analysis framework is proposed, that is, the immune cold / hot state of the tumor is evaluated and obtained according to the panoramic analysis features, and survival risk analysis and prediction are performed based on the immune cold / hot state of the tumor, so as to realize the personalized immune microenvironment portrait customization of patients, realizing comprehensive evaluations such as immune cell infiltration, immune checkpoint gene expression, and mutation burden, and improving the efficiency and accuracy of the panoramic analysis of the multi-omics tumor immune microenvironment.
[0218] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-omics tumor immune microenvironment panoramic analysis method based on HVAE and Cross-Attention, characterized in that: The specific steps include: S1: Acquire and preprocess multi-omics data on the tumor immune microenvironment; The multi-omics data at least includes DNA methylation data, RNA-Seq data, SNV data, CNV data and Clinical data; S2: Construct a multi-omics data dimension reduction / fusion network based on HVAE and Cross-Attention to obtain feature fusion latent variables; The constructed multi-omics data dimension reduction / fusion network includes several encoder modules based on hierarchical variational autoencoder HVAE and cascade network modules based on Cross Attention; Each encoder module is used to perform feature dimension reduction encoding on different data types in the multi-omics data to obtain multi-omics latent variables, and store the multi-omics latent variables into pre-divided latent variable classification layers according to data types; The cascade network module is used to perform feature fusion on the data of the latent variable classification layer to obtain feature fusion latent variables; S3: Perform comparative learning on the feature fusion latent variables to obtain the optimized feature fusion latent variables; S4: Extract and optimize the TME panoramic features of the tumor immune microenvironment from the feature fusion latent variables; The TME panoramic features include immune cell infiltration features, immune checkpoint gene expression features, and tumor mutation load features; S5: Evaluate and obtain the immune cold / hot status of the tumor based on the panoramic characteristics of the TME, and perform survival risk analysis and prediction for individual patients based on the immune cold / hot status of the tumor; S6: Based on the results of survival risk analysis prediction, a personalized immune microenvironment profile is customized for each patient for personalized identification of the tumor immune microenvironment status.
2. The method of claim 1, wherein the method comprises: The method for obtaining feature fusion latent variables in S2 specifically includes the following steps: S21: The encoder module performs feature dimension reduction encoding on different data types in the multi-omics data to obtain multi-omics latent variables; S22: Based on the multi-omics latent variables, the DNA methylation latent variable encoded by the feature dimensionality reduction is recorded as z1, the RNA-Seq latent variable is recorded as z2, the SNV latent variable is recorded as z3, the CNV latent variable is recorded as z4, and the Clinical latent variable is recorded as z5; and based on the multi-omics latent variables, according to the data type, the multi-omics latent variables are stored in the pre-divided latent variable classification layer, so as to store the multi-omics latent variables encoded by the feature dimensionality reduction in the corresponding latent variable classification layer; The latent variable classification layer includes a gene layer, a mutation layer and a clinical layer; Input the data of the latent variable classification layer into the cascade network module; S23: Use the RNA-Seq latent variable z2 as the input of the Query vector function in the first Cross Attention module S1 in the cascade network module; use the DNA methylation latent variable z1 as the input of the Key and Value vector function in the first Cross Attention module S1 to obtain and output the gene regulation relationship vector z s1 ; The CNV latent variable z4 is used as the input of the Query vector function in the second Cross Attention module S2 in the cascade network module; the SNV latent variable z3 is used as the input of the Key and Value vector function in the second Cross Attention module S2 to obtain and output the gene variation relationship vector z s2 ; S24: Gene mutation relationship vector z s2 As the input of the Query vector function in the third Cross Attention module S3 in the cascade network module; the gene regulation relationship vector z s1 As the input of the Key and Value vector function in the third Cross Attention module S3, to obtain and output the genomic variation fusion vector z s3 ; S25: Use the Clinical latent variable z5 as the input of the Query vector function in the fourth Cross Attention module S4 in the cascade network module; use the genomic variation fusion vector z s3 As the input of the Key and Value vector function in the fourth Cross Attention module S4, to obtain and output the feature fusion latent variable Z fusion .
3. The method for panoramic analysis of the multi-omics tumor immune microenvironment based on HVAE and Cross-Attention according to claim 2, characterized in that: The expression for feature dimensionality reduction encoding in S1 is: Where: q(z i |X i ) represents the multi-omics latent variables obtained after feature dimension reduction encoding; X i Represents multi-omics data on tumor immune microenvironment after pretreatment; z i represents the low-level latent variables of multi-omics data; μ i represents the latent variable mean of the encoder; σ i represents the latent variable variance of the encoder; i represents the category of multi-omics data, namely DNA methylation data, RNA-Seq data, SNV data, CNV data and Clinical data; The acquisition and output of gene regulation relationship vector z of each Cross Attention module in the Cross Attention-based cascade network module in S22 s1 Or obtain and output gene mutation relationship vector z s2 Or obtain and output genomic variation fusion vector z s3 Or obtain and output feature fusion latent variables Z fusion The expression is Where: Q represents the Query vector function; V represents the Value vector function; K represents the Key vector function; d represents the dimension; T represents the transpose.
4. The method for panoramic analysis of the multi-omics tumor immune microenvironment based on HVAE and Cross-Attention according to claim 3, characterized in that: In S3, the feature fusion latent variables are compared and learned to obtain a method for optimizing the feature fusion latent variables, specifically: Get the feature fusion latent variable Z of several outputs fusion As a data sample; Divide data samples based on the TME scoring system to obtain positive sample pairs and negative sample pairs; And the method for dividing the data sample based on the TME scoring system includes a division rule model, which is expressed as (Z i ,Z j ),where TME(i)=TME(j) (Z i ,Z k ),where TME(i)≠TME(k) Where: (Z i ,Z j ) represents similar sample pairs in the data sample, i.e., positive sample pairs; (Z i ,Z k ) represents the sample pairs with differences in the data samples, i.e., negative sample pairs; i, j, k represent the sample numbers in the data samples; Based on the pre-trained contrastive learning network, the clustering clusters of the feature fusion latent variables are obtained according to the positive sample pairs and the negative sample pairs, that is, the optimization of the feature fusion latent variables is achieved to obtain the optimized feature fusion latent variables; Among them, the contrast loss function of the contrastive learning network during pre-training is Where: represents the contrast loss function; S(Z i ,Z j ) represents the cosine similarity of the positive sample pair; S(Z i ,Z k ) represents the cosine similarity of the negative sample pair; τ represents the temperature parameter that controls the contrastive learning weight.
5. The method for panoramic analysis of the multi-omics tumor immune microenvironment based on HVAE and Cross-Attention according to claim 4, characterized in that: The S5 specifically includes the following steps: S51: Assess and obtain the immune cold / hot status of the tumor based on the panoramic characteristics of the TME; And the evaluation method for obtaining the immune cold / hot state of the tumor based on the panoramic characteristics of the TME is: Obtain the TME Score of the TME panoramic feature, and based on the preset scoring threshold, divide the tumor state into immune cold state and immune hot state according to the evaluation score TMEScore; The evaluation score TME Score is obtained by: TME Score=α×M immune +β×M checkpoint +γ×TMB Where: M immune represents the extracted immune cell infiltration feature matrix; M checkpoint represents the extracted immune checkpoint gene expression feature matrix; TMB represents the number of mutations of the tumor mutation burden feature; α, β, γ represent adjustable weight coefficients; S52: Predicting survival risk based on the tumor's immune cold / hot status; The TME panoramic features corresponding to the cold / hot state of tumor immunity will be evaluated as the input of the preset survival analysis model to perform survival risk analysis and prediction on patients and obtain survival risk prediction results; And the expression for predicting the patient's survival risk analysis is h(t|Z t )=h0(t)ezp(gZ t ) Where: h(t|Z t ) represents the feature vector Z corresponding to the tumor panoramic analysis feature at time t t The individual survival risk; h0(t) represents the base risk rate; g represents the characteristic vector Z t The relevant regression coefficient; exp(gZ t ) represents the hazard ratio.
6. The method for panoramic analysis of the multi-omics tumor immune microenvironment based on HVAE and Cross-Attention according to claim 5, characterized in that: The method for customizing the personalized immune microenvironment profile of an individual patient in S6 includes a customized rule model, which is expressed as A=[B,C,D,E,F] In the formula: A represents the portrait descriptor of the personalized immune microenvironment portrait; B represents the cell infiltration ratio of the immune cell infiltration characteristic of the tumor; C represents the Z-score score of the immune checkpoint gene expression characteristic; D represents the mutation TMB of the tumor mutation load characteristic; E represents the immune cold / hot state of the tumor; F represents the individual survival risk level predicted by the survival risk analysis.