Data processing method for geometric feature extraction of high-dimensional snapshot molecular data
Patent Information
- Application Number
- CN202411634760.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2044-11-15
AI Technical Summary
[0005]但这些研究大多集中在批量测序数据和脑成像数据上,针对快照型分子数据,尤其是细胞层次的几何状态量化工具仍然较为缺乏
Smart Images

Figure CN119694394B_ABST
Abstract
Description
Technical Field
[0001] This invention proposes a data processing method for extracting geometric features from high-dimensional snapshot molecular data. It identifies state differences, heterogeneity, and differentiation processes in the data by analyzing underlying geometric structures. This invention belongs to the interdisciplinary field of bioinformatics and geometric data analysis, and has wide applications in the processing and analysis of molecular data such as single-cell omics and genomics, especially in revealing cellular heterogeneity and cell differentiation dynamics. This invention relates to the field of processing high-dimensional snapshot molecular data. Background Technology
[0002] Snapshot molecular data refers to high-dimensional, static molecular data obtained from a single experiment, such as single-cell RNA sequencing (scRNA-seq) data, proteomics data, and multi-omics data. This type of data is widely used in fields such as single-cell omics, genomics, and proteomics. It provides researchers with detailed information on the expression of individual cells or molecules at a specific point in time, revealing the complex heterogeneity and potential state transition processes within biological systems. Single-cell RNA sequencing (scRNA-seq) technology is a representative example of snapshot molecular data, allowing researchers to study complex biological processes, such as cell fate decisions and cell differentiation, at the single-cell level. Unlike traditional bulk RNA sequencing (bulkRNA-seq) technology, scRNA-seq technology enables detailed analysis of heterogeneous cell populations. Snapshot molecular data typically exhibits high dimensionality and nonlinearity, posing significant challenges to data analysis and interpretation. To handle this high-dimensional data, various dimensionality reduction and visualization methods have been proposed in existing technologies, such as principal component analysis (PCA), t-distributed neighborhood embedding (t-SNE), and unified manifold approximation and projection (UMAP). These methods can reveal potential low-dimensional structures in data to some extent.
[0003] However, existing dimensionality reduction techniques have some limitations in analyzing the geometric structure of data. They often lose local and global geometric features during the dimensionality reduction process, leading to distortion of neighboring structures and failing to accurately preserve geodesic distances in the data. The geometric features in snapshot molecular data are closely related to molecular states, population distributions, and state transitions. Studying these geometric features helps to reveal underlying biological processes, such as cell differentiation and disease progression.
[0004] Recently, advances in network curvature theory and methods have provided new tools for high-dimensional data analysis, enabling the revelation of the intrinsic geometry and characteristics of data embedding manifolds. Network curvature has achieved some success in biological data analysis. For example, Farooq et al. found that curvature changes in brain networks can distinguish brain regions in healthy individuals and pathological states, and track changes in brain connectivity caused by aging and autism spectrum disorder (ASD) [1]. Simhal et al. revealed that curvature can identify inter-regional connectivity changes that are difficult to detect by traditional analysis methods by studying the brain networks of children with autism [2]. Sandhu et al. found that the curvature of cancer networks is significantly higher than that of normal networks by studying gene co-expression networks, and that it can effectively distinguish between normal and diseased samples [3]. Saucan et al. applied network curvature to metabolic and gene co-expression networks, revealing some novel general features in these networks [4].
[0005] However, most of these studies focus on batch sequencing data and brain imaging data, and tools for quantifying the geometric state of snapshot-type molecular data, especially at the cellular level, are still relatively lacking. To fill this gap, this invention proposes a geometric feature extraction-based method that can effectively analyze high-dimensional snapshot-type molecular data, and is particularly suitable for identifying stable states, transitional states, and population heterogeneity in the data, thereby revealing the intrinsic complex structure of molecules or cells.
[0006] The references are as follows:
[0007] [1]Hamza Farooq, Yongxin Chen, Tryphon T Georgiou, Allen Tannenbaum, and Christophe Lenglet. Network curvature as a hallmark of brain structural connectivity. Nature communications, 10(1):4937,2019.
[0008] [2]Anish K Simhal,Kimberly LH Carpenter,Saad Nadeem,Joanne Kurtzberg,Allen Song,Allen Tannenbaum,Guillermo Sapiro,andGeraldineDawson.Measuringrobustness of brain networks in autism spectrum disorder withricci curvature.Scientific reports,10(1):10819,2020.
[0009] [3]Romeil Sandhu,Tryphon Georgiou,Ed Reznik,Liangjia Zhu,IvanKolesov,Yasin Senbabaoglu,and Allen Tannenbaum.Graph curvature fordifferentiating cancer networks.Scientific reports,5(1):12323,2015.
[0010] [4]Emil Saucan,Areejit Samal,Melanie Weber,and Jürgen Jost.Discretecurvatures and network analysis.MATCH Commun.Math.Comput.Chem,80(3):605–622,2018.
[0011] [5]Sophie Petropoulos,Daniel Edsgrd, Reinius,Qiaolin Deng,SaritaPauliina Panula,Simone Codeluppi,Alvaro Plaza Reyes,Sten Linnarsson,RickardSandberg,and Fredrik Lanner.Single-cell rna-seq reveals lineage and xchromosome dynamics in human preimplantation embryos.Cell,165(4):1012–1026,2016.
[0012] [6] Varelas, Xaralabos, et al. The Crumbs complex couples cell densitysensing to Hippo-dependent control of the TGF-β-SMAD pathway. Developmental cell, 19.6: 831-844, 2010. Summary of the Invention
[0013] Existing technologies have significant limitations when processing snapshot-type molecular data, especially high-dimensional data such as single-cell RNA sequencing. This is primarily due to the inability to effectively quantify and analyze the geometric state at the cellular level. While existing network analysis has achieved some application in other biological data, tools for extracting geometric features from single-cell data remain scarce. This invention aims to address this issue by proposing a data processing method capable of efficiently extracting geometric features from high-dimensional snapshot-type molecular data. This method is used to analyze the underlying manifold structure within the data, thereby identifying key information such as cell state and cell fate differentiation.
[0014] To address these problems in existing technologies, this invention proposes a data processing method for geometric feature extraction from high-dimensional snapshot molecular data, applicable to geometric analysis of high-dimensional snapshot molecular data, especially high-dimensional single-cell data. The method includes the following steps:
[0015] Construct a K-Nearest Neighbor (KNN) graph based on single-cell gene expression data, where each node represents a cell and the edges between nodes are based on the differences in gene expression between cells.
[0016] Calculate the distance between cells in the graph using the shortest path.
[0017] Establish the probability distribution of cells and their neighbors.
[0018] By calculating Wasserstein distance, the geometric relationships between cells are analyzed, and the curvature characteristics of the edges between cells are quantified.
[0019] By calculating scalar curvature, we can identify the stable and transitional states of cells and quantify the heterogeneity of cell differentiation.
[0020] This technical solution can reveal embedded geometric structures that existing dimensionality reduction methods cannot capture. Unlike existing technologies, this invention can quantify the geometric state at the single-cell level without distorting geometric features, providing more detailed quantitative analysis of cells.
[0021] The technical solution of the present invention is further detailed below. The present invention provides a data processing method for geometric feature extraction of high-dimensional snapshot molecular data, used for geometric analysis of high-dimensional snapshot molecular data. Step A1: Construct a K-nearest neighbor graph from the snapshot molecular data, where each node represents a cell. Step A2: Calculate the shortest path between any two cells. Step A3: Generate the probability distribution for each cell. Step A4: Analyze the geometric relationship between cells by calculating Wasserstein distance, quantifying the curvature features of the edges between cells. Step A5: Calculate the scalar curvature of each cell by weighted averaging, thereby revealing cell state and heterogeneity. The input of the method is a high-dimensional molecular data matrix, especially a single-cell gene expression data matrix, and the output is the scalar geometric feature value for each cell. Through five steps (see... Figure 1 This method can calculate the intrinsic geometric features in snapshot molecular data, reflecting the underlying structure of complex cellular data.
[0022] Given single-cell sequencing data X = [x1, x2, ..., x n ]∈R r×n Where n is the number of cells and r is the number of genes, the method calculates the geometric feature value of each cell based on the following steps:
[0023] Step A1. Construct a graph of scRNA-seq data:
[0024] First, an undirected weighted graph G = (V, E, W) is constructed to approximate the manifold of the data embedding, where V is the vertex set, E is the edge set, and W is the edge weight matrix. In this graph, nodes represent cells, and each node is connected to the K nearest cells in terms of Euclidean distance from its gene expression level. The degree of each vertex is at least K. The edge weight matrix W = {W... ij} i,j∈V Calculated using Gaussian kernel distance:
[0025]
[0026] in, σ is the Euclidean distance between cells i and j. 2 It is a bandwidth parameter.
[0027] Step A2. Calculate the shortest path distance:
[0028] The distance d in the graph is defined as the weighted hop distance. Specifically, let p... ij p represents the path between nodes i and j through m+1 nodes. ij =i=v0~v1~…~v m =j, where the adjacent node v k ,v k+1 ∈p ij(k = 0, 1, ..., m-1) Corresponding edge e k =(v k ,v k+1 )∈E, each node appears only once. Let Let be the set of all possible paths between i and j (this set is finite because the graph is finite), and let be... It is a path The length of the associated edge, where, This is the Euclidean distance between nodes v_k and v{k+1}. Then, the corresponding path length is:
[0029]
[0030] The weighted hop distance d between nodes i and j ij The shortest path length among all paths connecting i and j:
[0031]
[0032] Assume the graph is a simple graph, which is connected and undirected, so there is at least one path between any two nodes i,j∈V.
[0033] Step A3. Assign a probability distribution to each cell:
[0034] For node (cell, data point) i, assign it a probability transition distribution μ. i The definition is as follows:
[0035]
[0036] Where N(i) is the neighborhood of node i in graph G.
[0037] Step A4. Calculate the local geometric features of the intercellular edges:
[0038] Based on the calculated shortest path distance matrix D={d ij} i,j∈V and the probability distribution of each vertex {μ i |i∈V}(or {μ j |j∈V}), the geometric eigenvalues of the edge (i,j)∈E are calculated as follows:
[0039]
[0040] Here, D defines the cost function for calculating the Wasserstein distance W1 (the shortest path distance matrix is used here).
[0041] Step A5. Calculate the scalar geometric eigenvalues of cell i using a weighted average, as shown in the following formula:
[0042]
[0043] This method applies the scalar curvature S of data point i. OR The definition of (i) uses the concept of convergence from the measure space to the real scalar curvature of the manifold.
[0044] The beneficial effects of this invention compared to the prior art are as follows:
[0045] This invention is applied to simulated data and actual single-cell sequencing data to explore the geometric features of high-dimensional data. In simulation studies, results show that this method has high accuracy and robustness in calculating curvature, and can recover the true curvature in simulated data.
[0046] In practical single-cell data analysis, this method was applied to human preimplantation embryo (HPE) and peripheral blood mononuclear cell (PBMC) datasets. In the HPE dataset, it was found that cells with high curvature were typically surrounded by neighboring cells in similar states, while cells with negative curvature acted as "bridges" between different populations. In other words, the calculated scalar curvature can serve as an important indicator for identifying cell states and transitional states, improving the interpretation and analysis capabilities of high-dimensional scRNA-seq data. In the PBMC dataset, this method successfully identified several novel cell subtypes by analyzing cell curvature distribution.
[0047] In summary, this invention provides a unique perspective for the geometric analysis of data manifolds by calculating the scalar curvature of each cell, revealing local geometric features that are difficult to detect using conventional methods. Its applications include identifying steady-state and transient states of cells, quantifying the heterogeneity of cell differentiation, and identifying immune cell subtypes. It provides a powerful tool for understanding complex structures in snapshot-type molecular data, and shows great potential, especially in revealing the dynamics of cell differentiation and development. Attached Figure Description
[0048] Figure 1 This is a flowchart of the method of the present invention.
[0049] Figure 2 Scalar curvature is used as an important indicator of cell type. (a) shows the density distribution of cell curvature in the HPE data. (b) is a schematic diagram of the visualization results of the HPE dataset using the EE dimensionality reduction method; red crosses indicate cells with small curvature values (<0). (c) is a schematic diagram of the visualization results of the HPE dataset using the EE dimensionality reduction method; red crosses indicate cells with large curvature values (>0.85). Detailed Implementation
[0050] Take, for example, the single-cell transcriptome sequencing data of the human preimplantation embryo (HPE) dataset. This dataset contains 1529 cells isolated from 88 human embryos, time-stamped from day 3 to day 7 of embryonic development (E3 to E7). HPE cells undergo a complex transition from morula to mature blastocyst. During this process, cells divide and differentiate into three main types: trophectoderm (TE), primitive endoderm (PE), and epiblast (EPI) [6]. We used the same preprocessing procedure as in reference [5].
[0051] Using the principal components of the preprocessed data as input to the method, and setting the nearest neighbor number K for each cell to 30, an undirected weighted graph G with 1529 vertices, 30713 edges, and an average degree of 40.17 is constructed. Each vertex represents a cell, and the edges represent gene expression differences between cells. For each pair of adjacent cells (i,j), the formula is used:
[0052]
[0053] Calculate edge weight w ij ,in This represents the Euclidean distance between cells i and j, with the bandwidth parameter σ set to 10.
[0054] Next, based on the graph constructed above, the shortest path between each pair of cells is calculated to estimate their distance on the embedded manifold. This is done using the weighted hop distance of all possible paths. Get the shortest path length d between cells i and j ij The calculation process uses the Floyd-Warshall algorithm.
[0055] Next, a probability distribution is defined for each cell in the graph to represent the propagation probability of each cell within its neighborhood. For each neighbor j of cell i, the probability is calculated using the formula... Calculate the transition probability. For cell j that is not in the neighborhood N(i), set μ. i (j) = 0.
[0056] Using the generated probability distribution, calculate the geometric relationships between cells. For each pair of adjacent cells (i,j), according to the formula:
[0057]
[0058] Calculate the Ollivier-Ricci curvature k(i,j). Where W1(μ) i ,μ j D) is the Wasserstein distance based on the shortest path distance matrix D, which can be calculated by solving a linear programming problem. In the actual implementation, the Python OT package is used to accelerate the calculation.
[0059] Then, by weighted averaging of the local curvatures, the scalar curvature of each cell is obtained. According to the formula:
[0060]
[0061] Calculate the scalar curvature S of each cell OR (i). The scalar curvature calculated using the above method exhibits a unimodal distribution, ranging from -0.508 to 0.913, with an average value of 0.488. Figure 2 (a).
[0062] We used the EE method to reduce the dimensionality of the HPE dataset to two-dimensional space, as shown in the scatter plot. Figure 2 As shown in (b), the color of the dot represents the number of developmental days of the cell at that point. As illustrated, positive cells with higher curvature are primarily located in the early and late stages of the HPE data. Specifically, cells with curvature greater than 0.85 mainly appear in stages E3, E4, and E7 (see Figure 1). Figure 2 (b)). In stages E3 and E4, the developmental state is relatively homogeneous, and the cells have not yet differentiated into different lineages. In contrast, cells in stage E7 are considered to have specialized into TE, PE, and EPI, forming mature blastocysts ready for implantation. Cells with higher curvature indicate that they are in an undifferentiated or specialized state, and visually suggest that they are surrounded by neighboring cells in a similar state.
[0063] Cells with negative curvature are mainly located between consecutive embryonic days, such as... Figure 2 (c) reflects the gradual development of cells during the transition. This is consistent with the major segregation factors reported in [5] during the developmental segregation process. Surprisingly, we also observed a large number of cells with negative curvature at the E5 stage. This finding is consistent with the formation of the blastocyst at the E5 stage. The transcriptional state at the E5 stage gradually diversifies, and cells begin to differentiate into different lineages, resembling a potential “tree-like” structure. Negative curvature indicates the transient state of cells and acts as a “bridge” between different populations.
[0064] Our results demonstrate that the scalar curvature of our method can serve as an important indicator of cell state, highlighting its potential to identify cell states and transitional states in single-cell sequencing data analysis, especially in complex biological processes.
Claims
1. A data processing method for extracting geometric features from high-dimensional snapshot molecular data, characterized in that: include: Step A1: Construct a K-nearest neighbor graph based on single-cell gene expression data, where each node represents a cell and the edges between nodes are constructed based on the differences in gene expression between cells; Step A2: Calculate the shortest path between any two cells; Step A3: Generate the probability distribution for each cell; Step A4: Analyze the geometric relationships between cells by calculating Wasserstein distance, and quantify the curvature characteristics of the edges between cells; Step A5 involves calculating the scalar curvature of each cell using a weighted average, thereby revealing cell state and heterogeneity; In step A1, an undirected weighted graph is constructed. To approximate the manifold of data embedding, For vertex set, For edge set, This is an edge weight matrix; nodes represent cells, and each node is connected to the K nearest cells in Euclidean distance from its gene expression level via an edge, with each vertex having a degree of at least K; edge weight matrix. Calculated using Gaussian kernel distance: ; in, It is a cell and The Euclidean distance between them It is a bandwidth parameter; In step A2, distance Defined as weighted hop distance; let Represents a node and Between The path of each node, i.e. Among them, adjacent nodes Corresponding edges Each node appears only once; among them, ; In step A3, for node Assign it a probability transition distribution The definition is as follows: ; in, It is a picture Middle node The neighborhood of.
2. The data processing method for geometric feature extraction of high-dimensional snapshot molecular data according to claim 1, characterized in that: make It is all possible and The set of paths between, and let { It is a path The length of the associated edge, where, It is a node and The Euclidean distance between them; then, the corresponding path length is: ; node and Weighted hop distance between For connection and The shortest path length among all paths: ; Any two nodes There is at least one path between them.
3. The data processing method for geometric feature extraction of high-dimensional snapshot molecular data according to claim 1, characterized in that: In step A4, based on the calculated shortest path distance matrix and the probability distribution of each vertex or Calculate the edges The geometric eigenvalues are as follows: ; in, Defines the method for calculating the Wasserstein distance. The cost function is the shortest path distance matrix used here.
4. The data processing method for geometric feature extraction of high-dimensional snapshot molecular data according to claim 1, characterized in that: In step A5, the cells are calculated by weighted average. The scalar geometric eigenvalues are given by the following formula: ; For data points scalar curvature The definition uses the concept of convergence from the measure space to the real scalar curvature of the manifold.
Citation Information
Patent Citations
Methods including geometric network analysis for selecting patients with high grade serous carcinoma of the ovary treated with immune checkpoint blockade therapy
WO2023239343A1