A method and system for predicting tumor evolution in spatial omics based on graph neural networks
Through the method based on graph neural network, the comprehensive spatial map is constructed and multimodal data is integrated, which solves the problem of data fragmentation and limited model adaptability in tumor evolution research, and improves the prediction accuracy and research efficiency of tumor evolution patterns.
Patent Information
- Application Number
- CN202410914627.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-07-09
AI Technical Summary
In the research on tumor evolution, the existing technology has problems such as fragmentation of data, insufficient multimodal data integration, limited model adaptability and high computational complexity, making it difficult to fully capture the spatial and temporal changes of tumors.
Using a spatial oscillary tumor evolution prediction method based on graph neural network, we use a comprehensive spatial map to integrate phenotype, genetic and imaging features, and use graph neural network models to learn features from spatial transcriptome data to perform spatiotemporal inference of tumor cloning evolution.
It improves the prediction accuracy of tumor evolution patterns, enhances the perception of tumor heterogeneity and subtle changes in the evolutionary process, has a high generalization and the effect of improving research efficiency, and supports the formulation of personalized treatment strategies.
Smart Images

Figure CN118969078B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene expression analysis, and in particular, to a method and system for predicting tumor evolution of spatial omics based on graph neural network. Background Technique
[0002] Tumor evolution is a complex multi-step process with genetic, phenotypic, and epigenetic heterogeneity. Traditional pathological methods mainly rely on morphological observations of tissue sections, diagnosing by observing the tissue structure and cytopathological features of patients under a microscope. However, pure pathological images cannot provide molecular information of lesions, and comprehensive exploration of diseases usually requires the combination of molecular diagnostic techniques.
[0003] In recent years, the rapid development of multi-omics technologies has exceeded the limitations of traditional pathology. These innovative methods help better understand tumors by analyzing DNA, RNA, and proteins in cancer cells. This not only helps accurately locate the type and stage of tumors, but also can identify specific cells, reveal important molecular markers, and is of great significance for the diagnosis and treatment of tumors.
[0004] Existing analysis methods include:
[0005] 1. Spatial omics technology: Spatial omics technology allows gene expression analysis while retaining tissue spatial information. By sequencing the gene expression at each position in a tissue sample, gene expression data with precise spatial position information can be obtained. Such data provides new opportunities for in-depth study of the spatial structure and function of biological tissues.
[0006] 2. Genetic information inference: Copy number variation (CNV) refers to the difference in the copy number of certain specific DNA fragments in the genomes of different individuals. These structural differences may be generated through duplication, deletion, or other changes. Gene amplifications and deletions common in cancer cells have important effects on the growth, drug sensitivity, and drug resistance of cancer cells. InferCNV is a tool for inferring large-scale chromosomal copy number alterations from single-cell RNA sequencing data. By comparing the gene expression intensities of tumor cells and normal cells, gene amplification or deletion regions can be identified.
[0007] 3. Pseudo-time-space inference (PSTS): PSTS is a spatial graph-based method for modeling and revealing the relationships between dynamically changing cellular transcriptional states across tissues. By combining spatial information and time information, more detailed trajectory mapping can be performed to study the dynamic changes of cells during processes such as development and disease progression.
[0008] Although these methods have achieved certain achievements in tumor research, they still have the following defects:
[0009] 1. Data fragmentation: Most existing methods can only provide static "snapshot" data and cannot capture the dynamic process of tumors at each time point. Although existing spatial transcriptome data can provide spatial information, it cannot directly infer information about temporal changes.
[0010] 2. Insufficient integration of multi-modal data: Although spatial omics technologies can provide gene expression and spatial location information, there are still deficiencies in integrating other phenotypic and imaging features. Traditional methods often analyze gene expression, morphological features, or imaging data separately and fail to fully utilize the comprehensive advantages of multi-modal data.
[0011] 3. Limited adaptability and generalization ability of models: The existing models have limited adaptability and generalization ability when dealing with different cancer types. The tissue clonal diversity of different cancer types varies greatly, and a general model that can adapt to multiple cancer types needs to be developed to provide a broader application prospect.
[0012] 4. Computational complexity and efficiency: As the data scale expands, existing methods have bottlenecks in computational complexity and efficiency. Especially when dealing with high-resolution spatial omics data, the computational speed and resource consumption of existing algorithms become the main obstacles restricting their application.
[0013] In summary, there are still problems in the existing technologies for tumor evolution research, such as data fragmentation, insufficient integration of multi-modal data, limited model adaptability, and high computational complexity. There is an urgent need for a new method to overcome these defects. Summary of the Invention
[0014] The purpose of the present invention is to provide a method and system for predicting tumor evolution based on graph neural networks using spatial omics data to clarify the spatial and temporal changes in tumor evolution, perform spatio-temporal inference of tumor clonal evolution, and improve prediction accuracy.
[0015] The purpose of the present invention can be achieved by the following technical solutions:
[0016] A method for predicting tumor evolution based on graph neural networks includes the following steps:
[0017] S1, Data acquisition and loading: Acquire different types of spatial transcriptome data, including gene expression data and spatial location information;
[0018] S2, Data preprocessing: Preprocess and normalize the acquired spatial transcriptome data;
[0019] S3, Gene annotation and grouping: Add annotations and sort the gene expression count data according to genomic positions, and then group the genes;
[0020] S4, Construct a spatial neighbor graph and calculate spatial statistics: Using the spatial location information and gene expression data in the spatial transcriptome data, construct an adjacency matrix reflecting the spatial relationship between cells, and calculate spatial statistics, where the adjacency matrix represents the connection relationship between each cell and its neighboring cells;
[0021] S5, Train the graph neural network model: Based on the adjacency matrix, construct a spatial neighbor graph as the input of the graph neural network model, train the model, and learn features from the spatial neighbor graph;
[0022] S6, Copy number estimation and diploid basic unit identification: Identify reliable diploid basic units and preliminarily estimate the copy number of each basic unit in spatial omics;
[0023] S7, Construct a comprehensive spatial graph: Add the estimated copy number to the spatial neighbor graph, construct a comprehensive spatial graph representing tumor heterogeneity, retrain the graph neural network model, capture the spatial distribution of genetic variations in tissue samples, and obtain potential embeddings;
[0024] S8, Breakpoint determination and copy number profile construction: According to the estimated copy number, use a segmentation algorithm to determine breakpoints and obtain the copy number profiles of each basic unit and spatial clone;
[0025] S9, Downstream task analysis: Use the potential embeddings for downstream task analysis, construct a tumor evolutionary tree, classify malignant and non-malignant cells, and analyze tumor heterogeneity;
[0026] S10, Cross-sample analysis: Extend to different spatial omics samples and conduct comparative analysis on primary tumors and metastases.
[0027] The S2 includes the following steps:
[0028] S21, Plot the transcript distribution map of each basic unit, filter out low-quality basic units according to the minimum transcript number threshold, and filter out low-expressed genes according to the minimum basic unit number threshold, where each basic unit represents a preset-sized area on the tissue section and provides the gene expression information of this area;
[0029] S22, Filter mitochondrial genes, HLA genes, and cell cycle genes;
[0030] S23, Standardize and logarithmically transform the gene expression in each basic unit.
[0031] The S3 includes the following steps:
[0032] S31, Map the ID to the genomic position and initialize the genomic position column;
[0033] S32, Filter out discarded gene IDs and null values;
[0034] S33, Traverse the filtered genes, fill in the chromosome position information, add annotations to the gene expression count data according to the genomic position, and sort them;
[0035] S34, Sort the genes according to their chromosome positions into several groups, each group contains a preset number of genes, and the gene expression values of each group are represented by the average expression level of the genes included in the group;
[0036] S35, Calculate the average count value of each group to obtain a count matrix, and the count matrix represents the gene expression data of each basic unit.
[0037] The S4 includes the following steps:
[0038] S41, Determine the adjacency relationship between cells through Delaunay triangulation or KNN algorithm, and generate a spatial relationship graph between cells;
[0039] S42, Calculate the centrality score of each cell in the spatial graph and calculate the cell co-occurrence probability;
[0040] S43, Conduct neighbor enrichment analysis to evaluate the enrichment degree of a certain type of cell among its neighbors;
[0041] S44, Calculate the Moran index to evaluate spatial autocorrelation and verify the degree of aggregation of gene expression in space.
[0042] The S5 includes the following steps:
[0043] S51, Based on the adjacency matrix, construct the spatial transcriptome data into a graph structure, where nodes represent each basic unit, edges represent the neighbor relationship between basic units, and input this graph structure into the graph neural network model for training;
[0044] S52, Use the graph neural network model to extract important features from the spatial neighbor graph, learn the latent representation of each basic unit through the multi-layer convolution operation of the model, and generate a more detailed feature map;
[0045] S53, Obtain the latent representation of each basic unit through the graph neural network model, and use the clustering method to cluster the latent representation.
[0046] The S6 includes the following steps:
[0047] S61, Calculate the autocorrelation score for each cluster obtained by clustering the latent representation in step S5, evaluate the stability of cell gene expression in each cluster, and the higher the autocorrelation score, the more consistent the gene expression of cells in the cluster. Among them, the calculation method of the autocorrelation score is:
[0048]
[0049] where: ρ m is the autocorrelation score of cluster m, x i and x j are the expression profiles of spot i and j in cluster m, μ m is the mean expression of cluster m, is the variance of cluster m;
[0050] S62. Based on the assumption that the gene expression diversity is smaller in diploid cells compared to tumor cells, select the cluster with the highest autocorrelation score as the diploid basic unit;
[0051] S63. Aggregate the gene expression data of the identified diploid basic units, calculate the average expression value of each group in the diploid basic unit, and form a baseline expression profile;
[0052] S64. Normalize the count matrix using the baseline expression profile to obtain a preliminary estimate of the copy number of each group in all basic units as the genetic feature.
[0053] The S8 includes the following steps:
[0054] S81. According to the obtained copy number estimates, use a likelihood-based segmentation algorithm to segment the copy number estimates of each basic unit, determine multiple breakpoints. The segment between each breakpoint represents a different gene region. By calculating the median copy number of all genes in the segment, obtain the copy number estimate of each segment and get the copy number profile of the basic unit. Among them, assuming that the observed data at each position i is independently and identically distributed, the segmentation algorithm determines the breakpoint by calculating the negative maximum log-likelihood function:
[0055]
[0056] where, represents the subsequence of data points from position τ a +1 to τ b θ represents the parameter for maximizing the probability of the observed data y i logP(y i ∣θ) is the log probability of the observed data y i under the given parameter θ, τ a and τ b represent two breakpoints;
[0057] S82. Aggregate the copy number estimates of each basic unit, calculate the average copy number of all basic units for each breakpoint segment, and combine the average values to form the overall copy number profile of the spatial clone.
[0058] The said S9 includes the following steps:
[0059] S91, perform clustering analysis using the cell latent embeddings generated in step S7 to identify tumor clones;
[0060] S92, automatically classify malignant and non - malignant cells: use the variational inference method to automatically classify cells into malignant and non - malignant cells according to the copy number variation;
[0061] S93, construct a tumor evolutionary tree for the tumor clones obtained by clustering: according to the clustering results obtained in step S8, group the tumor cells according to the clone structure, and use the grouped data to construct a tumor evolutionary tree to show the evolution process of different clones;
[0062] S94, analyze the tumor heterogeneity, identify the characteristics and evolution paths of different clones: perform a detailed analysis of the tumor heterogeneity in a specific sample to identify the unique characteristics and evolution paths of each clone; apply segmentation algorithms and variational inference techniques to refine the copy number variation of each clone and further identify the genomic structure changes of the tumor.
[0063] The said S10 includes the following steps:
[0064] S101, perform steps S1 - S9 on the spatial omics datasets of different tumor types to reveal the clone diversity of different cancer types;
[0065] S102, perform a comparative analysis of primary tumors and metastases to identify specific and shared genomic alterations and reveal the evolutionary path of metastasis.
[0066] A spatial omics tumor evolution prediction system based on a graph neural network for implementing the method as described above, the system includes:
[0067] Data acquisition module: acquire different types of spatial transcriptome data, including gene expression data and spatial location information;
[0068] Pre - processing module: pre - process and normalize the acquired spatial transcriptome data;
[0069] Annotation and grouping module: add annotations and sort the gene expression count data according to the genomic location, and then group the genes;
[0070] Spatial neighbor graph construction module: use the spatial location information and gene expression data in the spatial transcriptome data to construct an adjacency matrix reflecting the spatial relationship between cells, and calculate spatial statistics, where the adjacency matrix represents the connection relationship between each cell and its neighboring cells;
[0071] Graph Neural Network Training Module: Construct a spatial neighbor graph based on the adjacency matrix as the input of the graph neural network model, train the model, and learn features from the spatial neighbor graph;
[0072] Copy Number Estimation Module: Identify reliable diploid basic units and preliminarily estimate the copy number of each basic unit in spatial omics;
[0073] Comprehensive Feature Training Module: Incorporate the estimated copy number into the spatial neighbor graph to construct a comprehensive spatial graph representing tumor heterogeneity, retrain the graph neural network model, capture the spatial distribution of genetic variations in tissue samples, and obtain potential embeddings;
[0074] Breakpoint Segmentation Module: Determine breakpoints using a segmentation algorithm based on the estimated copy number to obtain the copy number profiles of each basic unit and spatial clone;
[0075] Analysis Module: Use the potential embeddings for downstream task analysis, construct a tumor evolution tree, classify malignant and non - malignant cells, and analyze tumor heterogeneity;
[0076] Extended Comparison Module: Extend to different spatial omics samples for comparative analysis of primary tumors and metastases.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] (1) Improve the prediction accuracy of tumor evolution patterns: The present invention constructs a comprehensive spatial graph by integrating multiple features in spatial omics technology, including phenotypic, genetic, and imaging features. This multi - feature integration method can comprehensively reflect the complex interactions in the tumor microenvironment, improve the accurate description of tumor evolution patterns, provide more detailed and realistic tumor evolution patterns, and is of great significance for the study of tumor heterogeneity and the understanding of tumor progression.
[0079] (2) More subtle feature perception: By introducing genetic information into the construction process of the graph network, the present invention enhances the ability to learn from multi - level spatial omics data. Compared with data analysis methods that only use a single feature, this method of constructing a multi - feature graph network can more accurately capture the subtle changes in tumor heterogeneity and the evolution process.
[0080] (3) Strong generalization ability: The method of the present invention has high generalization ability and can be applied to various cancer types. For different types of solid tumors, this method can provide valuable insights and reveal tissue clone diversity in different cancers. Therefore, this method has broad application prospects in the field of cancer research and can provide new tools and methods for the evolutionary research of different types of cancers.
[0081] (4)Improve research efficiency: Through automated and efficient data processing and analysis, the method of the present invention significantly improves research efficiency. Researchers can obtain high-quality analysis results in a shorter time, promoting the rapid development of tumor research.
[0082] (5)Facilitate the formulation of personalized treatment strategies: The method of the present invention can reveal the clonal diversity and evolutionary paths in different cancer types, which provides a scientific basis for formulating personalized cancer treatment strategies. By identifying the characteristics and evolutionary paths of different clones, it can assist doctors in designing treatment plans more targeted and improving the treatment effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is the flowchart of the method of the present invention;
[0084] Figure 2 is the system structure diagram of the present invention;
[0085] Figure 3 is the model training process diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0086] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0087] This embodiment provides a method for predicting tumor evolution of spatial omics based on graph neural network, as Figure 1 shown, including the following steps:
[0088] S1, Data acquisition and loading: Obtain different types of spatial transcriptome data, including gene expression data and spatial location information.
[0089] Specifically, S1 includes the following steps:
[0090] S11, According to the spatial omics types of different platforms, different methods are called in the process for data loading, and an AnnData object is established in Python.
[0091] S12, The applicable platforms covered by the method include: 10x Genomics Visium, 10x Genomics Xenium, Vigzen MERSCOPE (MERFISH), Nanostring CosMx, Stereo-seq, 4i (four-dimensional immunofluorescence imaging), Imaging Mass Cytometry, seqFISH, Slide-seqV2, etc.
[0092] The analysis processes for data from different platforms vary. In this embodiment, 10x Genomics Visium data is mainly used as an example. Among them, a "spot" is the basic unit for capturing and analyzing gene expression in the 10x Visium technology. Each spot (basic unit) represents a small area on the tissue section and provides gene expression information for that area.
[0093] In one embodiment, taking 10x Genomics Visium data as an example, the read_10x_h5 function in scanpy is used to read the HDF5 file, load the expression data, spatial coordinates, and scale factors, process the relevant files and metadata, and load the image information as needed.
[0094] S2, Data preprocessing: Preprocess and normalize the obtained spatial transcriptome data.
[0095] Specifically, S2 includes the following steps:
[0096] S21, Plot the transcript distribution map for each spot (basic unit), filter out low-quality basic units according to the minimum transcript number threshold, and filter out low-expressed genes according to the minimum basic unit number threshold.
[0097] Low-quality cells usually have fewer transcripts. By setting a minimum transcript number threshold, those cells with transcript numbers below this threshold can be filtered out. In one embodiment, it can be achieved by sc.pp.filter_cells(adata, min_counts = 10), and the threshold is set to 10. This function filters out those cells with a total transcript number less than 10, considering these cells to be of low quality.
[0098] Low-expressed genes are expressed in a small number of cells. By setting a minimum basic unit number threshold, those genes expressed in a small number of cells can be filtered out. In one embodiment, it can be achieved by sc.pp.filter_genes(adata, min_cells = 5), and the threshold is set to 5. This function filters out those genes expressed in fewer than 5 cells, considering these genes to have a low expression level.
[0099] S22, Filter mitochondrial genes (such as genes starting with 'MT-'), HLA genes, and cell cycle genes.
[0100] S23, Standardize the gene expression in each basic unit (normalize the total transcript number of each spot to 1e4), perform logarithmic transformation, and visualize the results.
[0101] S3, Gene annotation and grouping: Add annotations and sort the gene expression count data according to the genomic positions, and then group the genes (binning).
[0102] Specifically, S3 includes the following steps:
[0103] S31, Use the pyensembl library to map ENSG IDs to genomic positions and initialize the genomic position column.
[0104] S32, Filter out the discarded gene IDs and NaN values.
[0105] S33, Traverse the filtered genes, fill in the chromosome position information, add annotations and sort the gene expression count data according to the genomic positions; in this embodiment, the genomic data is obtained through the pyensembl library and added to the annotations of the gene expression count. By this method, each gene is attached with the corresponding genomic position information, so that sorting can be performed according to this information.
[0106] S34, For the convenience of subsequent analysis and to reduce the influence of dropout, sort the genes according to their chromosome positions into several groups, with each group containing 25 genes (i.e., 25 genes as a bin), and the gene expression values of each group (bin) are represented by the average expression level of the genes included in the group. Using bins to group genes helps to reduce the sparsity and noise of the data, improve the significance of statistical analysis, increase the computational efficiency, and standardize the data of different samples.
[0107] S35, Calculate the average count value of each bin, obtain the count matrix of spot×bin, and store it in adata.obs. The count matrix represents the gene expression data of each basic unit.
[0108] S4, Construct a spatial neighbor graph and calculate spatial statistics: Utilize the spatial position information and gene expression data in the spatial transcriptome data to construct an adjacency matrix reflecting the spatial relationship between cells, and calculate spatial statistics. The adjacency matrix represents the connection relationship between each cell and its neighboring cells.
[0109] Specifically, S4 includes the following steps:
[0110] S41, Determine the adjacency relationship between cells through Delaunay triangulation or the KNN (k-nearest neighbor) algorithm to generate a spatial relationship graph between cells; in this embodiment, the parameter k of the KNN algorithm is 6.
[0111] S42, Calculate the centrality score of each cell in the spatial graph and calculate the cell co-occurrence probability.
[0112] S43. Perform neighbor enrichment analysis to evaluate the enrichment degree of a certain type of cells among its neighbors.
[0113] S44. Calculate Moran's I to evaluate spatial autocorrelation and verify the degree of spatial aggregation of gene expression.
[0114] S5. Training of the graph neural network model: Construct a spatial neighbor graph based on the adjacency matrix as the input of the graph neural network model, train the model, and learn features from the spatial neighbor graph.
[0115] Specifically, S5 includes the following steps:
[0116] S51. Construct the spatial transcriptome data into a graph structure based on the adjacency matrix, where the nodes represent each basic unit spot, the edges represent the neighbor relationship between the basic units, and input this graph structure into the graph neural network model for training.
[0117] S52. Use a graph neural network (GNN) model (such as the Graph Transformer convolutional layer) to extract important features from the spatial neighbor graph, learn the latent representation of each basic unit through the multi-layer convolutional operation of the model, and generate a more detailed feature map. During the training process, adopt an appropriate loss function (such as cross-entropy loss) and optimization algorithm (such as the Adam optimizer, learning rate 0.01) to improve the performance of the model. Select a self-supervised learning loss function (such as reconstruction error) to measure the ability of the model to reconstruct the input data, and use an optimization algorithm (such as the Adam optimizer) for parameter update to improve the training effect and performance of the model.
[0118] S53. Obtain the latent representation of each basic unit through the graph neural network model, and use a suitable clustering method (such as k-means, mclust clustering) to cluster the latent representation.
[0119] S6. Copy number estimation and diploid basic unit identification: Identify reliable diploid spots (i.e., normal spots), and preliminarily estimate the copy number of each basic unit in the spatial omics.
[0120] Specifically, S6 includes the following steps:
[0121] S61. Calculate the autocorrelation score for each cluster obtained by clustering the latent representation in step S5 to evaluate the stability of cell gene expression in each cluster. The higher the autocorrelation score, the more consistent the gene expression of the cells in the cluster. Among them, the calculation method of the autocorrelation score is:
[0122]
[0123] Where: ρ mis the autocorrelation score of cluster m, x i and x j are the expression profiles of spot i and j in cluster m, μ m is the expression mean of cluster m, is the variance of cluster m.
[0124] S62. Based on the hypothesis that diploid cells (i.e., normal cells) have less gene expression diversity compared to tumor cells, select the cluster with the highest autocorrelation score as diploid spots. These cells usually exhibit less gene expression variability.
[0125] S63. Aggregate the gene expression data of the identified diploid spots, calculate the average expression value of each bin in the diploid spots, and form a baseline expression profile.
[0126] S64. Normalize the count matrix using the baseline expression profile to obtain a preliminary estimate of the copy number of each bin in all spots, which is used as a genetic feature.
[0127] S7. Comprehensive spatial map construction: Incorporate the estimated copy number into the spatial neighbor graph, integrate phenotypic, genetic, and imaging features to construct a comprehensive spatial map representing tumor heterogeneity, retrain the graph neural network model to capture the spatial distribution of genetic variations in tissue samples, and learn a finer estimate of the copy number of each bin in each spot to obtain a latent embedding for downstream tasks. The training process of the graph neural network model is as Figure 3 shown.
[0128] S8. Breakpoint determination and copy number profile construction: Based on the estimated copy number, use a segmentation algorithm to determine breakpoints and obtain the copy number profiles of each elementary unit and spatial clone.
[0129] Specifically, S8 includes the following steps:
[0130] S81. According to the obtained copy number estimates, use a likelihood-based segmentation algorithm to segment the copy number estimates of each elementary unit, determine multiple breakpoints, and the paragraph between each breakpoint represents a different gene region. By calculating the median copy number of all genes in the paragraph, obtain the copy number estimate of each segment and the copy number profile of the elementary unit.
[0131] Assume that the observed data at each position i is independently and identically distributed. The segmentation algorithm determines breakpoints by calculating the negative maximum log-likelihood function:
[0132]
[0133] where, represents from position τ afrom +1 to τ b a subsequence of data points, where θ represents the parameter for maximizing the probability of the observed data y i and logP(y i |θ) is the log probability of the observed data y given the parameter θ i , and τ a and τ b represent two breakpoints.
[0134] S82, Aggregate the copy number estimates of each basic unit, calculate the average copy number of all basic units for each breakpoint segment, and combine the averages to form the overall copy number profile of the spatial clone.
[0135] S9, Downstream task analysis: Use the latent embedding for downstream task analysis, construct a tumor evolutionary tree, classify malignant and non-malignant cells, and analyze tumor heterogeneity.
[0136] Specifically, S9 includes the following steps:
[0137] S91, Perform clustering analysis (such as Leiden, Mclust, Louvain methods) using the cell latent embedding generated in step S7 to identify tumor clones; downstream analysis and visualization can be customized according to user needs, such as performing pathway analysis, differential expression analysis, functional annotation, etc. on different clones.
[0138] S92, Automatically classify malignant and non-malignant cells: Use the variational inference method to automatically classify cells into malignant and non-malignant cells according to copy number variations.
[0139] S93, Construct a tumor evolutionary tree for the tumor clones obtained by clustering: According to the clustering results obtained in step S8, group the tumor cells according to the clone structure, and use the grouped data to construct a tumor evolutionary tree to show the evolution process of different clones.
[0140] S94, Analyze the heterogeneity of the tumor, identify the characteristics and evolution paths of different clones: Conduct a detailed analysis of the tumor heterogeneity in a specific sample, identify the unique characteristics and evolution paths of each clone; apply segmentation algorithms and variational inference techniques to refine the copy number variations of each clone and further identify the genomic structure changes of the tumor.
[0141] S10, Cross-sample analysis: Extend to different spatial omics samples and perform comparative analysis on primary tumors and metastases.
[0142] Specifically, S10 includes the following steps:
[0143] S101, Execute steps S1 - S9 on the spatial omics datasets of various different tumor types to reveal the clone diversity of different cancer types.
[0144] S102. Compare and analyze the primary tumor and metastatic tumors, identify specific and shared genomic alterations, and reveal the evolutionary path of metastasis.
[0145] This embodiment also provides a spatial omics tumor evolution prediction system based on a graph neural network for implementing the method described above. As Figure 2 shown, the system includes:
[0146] Data acquisition module: Acquire different types of spatial transcriptome data, including gene expression data and spatial location information;
[0147] Preprocessing module: Preprocess and normalize the acquired spatial transcriptome data;
[0148] Annotation and grouping module: Add annotations and sort the gene expression count data according to the genomic location, and then group the genes;
[0149] Spatial neighbor graph construction module: Utilize the spatial location information and gene expression data in the spatial transcriptome data to construct an adjacency matrix reflecting the spatial relationship between cells, and calculate spatial statistics. The adjacency matrix represents the connection relationship between each cell and its neighboring cells;
[0150] Graph neural network training module: Construct a spatial neighbor graph based on the adjacency matrix as the input of the graph neural network model, train the model, and learn features from the spatial neighbor graph;
[0151] Copy number estimation module: Identify reliable diploid basic units and preliminarily estimate the copy number of each basic unit in spatial omics;
[0152] Comprehensive feature training module: Add the estimated copy number to the spatial neighbor graph, construct a comprehensive spatial graph representing tumor heterogeneity, retrain the graph neural network model, capture the spatial distribution of genetic variations in tissue samples, and obtain potential embeddings;
[0153] Breakpoint segmentation module: Determine breakpoints using a segmentation algorithm based on the estimated copy number, and obtain the copy number profiles of each basic unit and spatial clone;
[0154] Analysis module: Use the potential embeddings for downstream task analysis, construct a tumor evolution tree, classify malignant and non - malignant cells, and analyze tumor heterogeneity;
[0155] Extended comparison module: Extend to different spatial omics samples and perform comparative analysis on primary tumors and metastatic tumors.
[0156] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0157] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A spatial omics tumor evolution prediction method based on graph neural network, characterized in that: The following steps are involved: S1, data acquisition and loading: Acquire different types of spatial transcriptome data, including gene expression data and spatial location information; S2, data preprocessing: preprocessing and normalizing the acquired spatial transcriptome data; S3, Gene Annotation and Grouping: Annotate and sort gene expression count data according to genomic location, and then group genes; S4, constructing a spatial neighbor graph and calculating spatial statistics: using the spatial location information and gene expression data in the spatial transcriptome data, constructing an adjacency matrix reflecting the spatial relationship between cells, and calculating spatial statistics, wherein the adjacency matrix represents the connection relationship between each cell and its neighboring cells; S5, graph neural network model training: construct a spatial neighbor graph based on the adjacency matrix as the input of the graph neural network model, train the model, and learn features from the spatial neighbor graph; The S5 comprises the following steps: S51, constructing the spatial transcriptome data into a graph structure based on the adjacency matrix, where the nodes represent each basic unit and the edges represent the neighbor relationship between the basic units, and inputting the graph structure into the graph neural network model for training; S52, uses a graph neural network model to extract important features from the spatial neighborhood graph, learns the potential representation of each basic unit through the model's multi-layer convolution operations, and generates a more detailed feature map; S53, obtaining the potential representation of each basic unit through the graph neural network model, and clustering the potential representation using a clustering method; S6, copy number estimation and diploid basic unit identification: identify reliable diploid basic units and preliminarily estimate the copy number of each basic unit in spatial omics; The S6 comprises the following steps: S61, calculating the autocorrelation score for each cluster obtained by the cluster potential representation in step S5, and evaluating the stability of the gene expression of the cells in each cluster. The higher the autocorrelation score, the more consistent the gene expression of the cells in the cluster. The autocorrelation score is calculated as follows: Where: m is the autocorrelation score of cluster M, x i and x j is the expression profile of spoti and j in cluster m, μ m is the mean expression value of cluster m, is the variance of cluster m; S62, based on the assumption that diploid cells have less gene expression diversity than tumor cells, the cluster with the highest autocorrelation score was selected as the diploid basic unit; S63, aggregating the gene expression data of the identified diploid basic units, calculating the average expression value of each group in the diploid basic unit, and forming a baseline expression profile; S64, normalize the count matrix by the baseline expression profile to obtain a preliminary estimate of the copy number of each group in all basic units as a genetic signature; S7, comprehensive spatial graph construction: add the estimated copy number to the spatial neighbor graph to construct a comprehensive spatial graph representing tumor heterogeneity, retrain the graph neural network model to capture the spatial distribution of genetic variations in tissue samples, and obtain latent embedding; S8, breakpoint determination and copy number profile construction: Based on the estimated copy number, the segmentation algorithm is used to determine the breakpoints and obtain the copy number profile of each basic unit and spatial clone; S9, downstream task analysis: Use latent embedding to perform downstream task analysis, construct a tumor evolution tree, classify malignant and non-malignant cells, and analyze tumor heterogeneity; S10, Cross-sample analysis: extended to different spatial omics samples, comparative analysis of primary tumors and metastases.
2. The spatial omics tumor evolution prediction method based on graph neural network according to claim 1, characterized in that: The S2 comprises the following steps: S21, drawing a transcript distribution map of each basic unit, filtering low-quality basic units according to a minimum transcript number threshold, and filtering low-expressed genes according to a minimum basic unit number threshold, wherein each basic unit represents an area of a preset size on a tissue section and provides gene expression information of the area; S22, filtering mitochondrial genes, HLA genes, and cell cycle genes; S23, Gene expression in each basic unit was normalized and log-transformed.
3. The spatial omics tumor evolution prediction method based on graph neural network according to claim 1, characterized in that: The S3 comprises the following steps: S31, mapping ID to genome position, initializing genome position column; S32, filtering discarded gene IDs and null values; S33, traverse the filtered genes, fill in the chromosome location information, add annotations to the gene expression count data and sort them according to the genome location; S34, sorting the genes according to their chromosomal positions into a plurality of groups, each group comprising a preset number of genes, and the gene expression value of each group is represented by the average expression level of the genes contained in the group; S35, calculating the average count value of each group to obtain a count matrix, wherein the count matrix represents the gene expression data of each basic unit.
4. The spatial omics tumor evolution prediction method based on graph neural network according to claim 1, characterized in that: The S4 comprises the following steps: S41, determine the adjacency relationship between cells through Delaunay triangulation or KNN algorithm to generate a spatial relationship map between cells; S42, calculate the centrality score of each cell in the spatial map and calculate the cell co-occurrence probability; S43, neighbor enrichment analysis was performed to evaluate the enrichment of a certain type of cell in its neighbors; S44, Moran's index was calculated to evaluate spatial autocorrelation and verify the degree of spatial clustering of gene expression.
5. The spatial omics tumor evolution prediction method based on graph neural network according to claim 1, characterized in that: The S8 comprises the following steps: S81, based on the obtained copy number estimate, the copy number estimate of each basic unit is segmented using a likelihood-based segmentation algorithm to determine multiple breakpoints, wherein the segments between each breakpoint represent different gene regions, and the copy number estimate of each segment is obtained by calculating the median copy number of all genes in the segment, thereby obtaining the copy number profile of the basic unit; wherein, assuming that the observed data at each position i are independent and identically distributed, the segmentation algorithm determines the breakpoints by calculating the negative maximum log-likelihood function: in, Represents the position τ a +1 to τ b θ represents the subsequence of data points used to maximize the observed data y i The probability parameter, logP(y i |θ) is the observed data y under given parameters θ i The logarithmic probability, τ a and τ b Indicates two breakpoints; S82, aggregate the copy number estimates for each basic unit, calculate the average copy number of all basic units for each breakpoint fragment, and combine the averages to form an overall copy number profile for the spatial clones.
6. The spatial omics tumor evolution prediction method based on graph neural network according to claim 1, characterized in that: The S9 comprises the following steps: S91, performing cluster analysis using the cell potential embeddings generated in step S7 to identify tumor clones; S92, automatic classification of malignant and non-malignant cells: using variational inference methods, cells are automatically classified into malignant and non-malignant cells based on copy number variation; S93, constructing a tumor evolution tree for the tumor clones obtained by clustering: according to the clustering results obtained in step S8, the tumor cells are grouped according to the clone structure, and a tumor evolution tree is constructed using the grouped data to show the evolution process of different clones; S94, analyze tumor heterogeneity and identify the characteristics and evolution paths of different clones: conduct detailed analysis of tumor heterogeneity in specific samples to identify the unique characteristics and evolution paths of each clone; apply segmentation algorithms and variational inference techniques to refine the copy number variation of each clone and further identify changes in the genomic structure of the tumor.
7. The spatial omics tumor evolution prediction method based on graph neural network according to claim 1, characterized in that: The S10 comprises the following steps: S101, performing steps S1-S9 on the spatial omics datasets of different tumor types to reveal the clonal diversity of different cancer types; S102, comparative analysis of primary tumors and metastases identifies specific and shared genomic alterations and reveals the evolutionary pathways of metastasis.
8. A spatial omics tumor evolution prediction system based on graph neural network, characterized in that: For implementing the method according to any one of claims 1 to 7, the system comprises: Data acquisition module: obtains different types of spatial transcriptome data, including gene expression data and spatial location information; Preprocessing module: preprocess and normalize the acquired spatial transcriptome data; Annotation and grouping module: annotate and sort gene expression count data according to genomic location, and then group genes; Spatial neighbor graph construction module: using the spatial location information and gene expression data in the spatial transcriptome data, construct an adjacency matrix reflecting the spatial relationship between cells and calculate spatial statistics. The adjacency matrix represents the connection relationship between each cell and its neighboring cells. Graph neural network training module: constructs a spatial neighbor graph based on the adjacency matrix as the input of the graph neural network model, trains the model, and learns features from the spatial neighbor graph; Copy number estimation module: Identify credible diploid basic units and preliminarily estimate the copy number of each basic unit in spatial omics; Comprehensive feature training module: Add estimated copy number to the spatial neighbor graph to construct a comprehensive spatial graph representing tumor heterogeneity, retrain the graph neural network model to capture the spatial distribution of genetic variations in tissue samples, and obtain latent embeddings; Breakpoint segmentation module: Based on the estimated copy number, the segmentation algorithm is used to determine the breakpoints and obtain the copy number profile of each basic unit and spatial clone; Analysis module: Use potential embeddings to analyze downstream tasks, build tumor evolution trees, classify malignant and non-malignant cells, and analyze tumor heterogeneity; Extended Comparative Module: Expanded to different spatial omics samples to conduct comparative analysis of primary tumors and metastases.
Citation Information
Patent Citations
Copy number variation prediction method and device, computer equipment and storage medium
CN111402951A
Method for analyzing active pathway in single cell multi-omics based on graph neural network
CN115240772A