Single-cell trajectory inference methods and apparatuses, devices, media
By generating a correlation graph of single-cell transcriptome data and using a pre-defined knowledge graph to complete developmental information, the reliance on prior biological knowledge in existing technologies is resolved, and highly reliable inference of cell development trajectory without human intervention is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HUADA BIO & INFORMATION FUSION TECHNOLOGY RESEARCH CO LTD
- Filing Date
- 2025-12-01
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for inferring single-cell trajectories require researchers to provide prior biological knowledge, which increases the subjectivity and uncertainty of the analysis and makes it difficult to accurately interpret cell differentiation and development processes.
By acquiring single-cell transcriptome data to generate a correlation graph, and using a pre-defined knowledge graph to complete developmental information, a developmental relationship graph is generated, reducing reliance on prior knowledge and improving the objectivity and certainty of trajectory inference.
It can accurately infer cell development trajectory without requiring prior biological knowledge from humans, thus improving the objectivity and reliability of trajectory inference and generating reasonable analysis results.
Smart Images

Figure CN121237199B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of single-cell biological information technology, and in particular to a method, device, equipment, and medium for inferring single-cell trajectory. Background Technology
[0002] In single-cell data analysis, trajectory inference is a method used to reconstruct cell differentiation and developmental pathways from single-cell transcriptome data. By analyzing the similarities and differences between cells, trajectory inference arranges cells into a continuous trajectory, thereby revealing the processes of change in cell state.
[0003] In related technologies, Wolf et al. developed the PAGA algorithm and integrated it into the Scanpy toolkit. The main goal of the PAGA algorithm is to generate a clear diagram of cell state transitions by abstracting and simplifying complex cell atlases, thereby helping researchers better understand cell differentiation and development processes. However, it still has some limitations in interpreting the biological meaning of single-cell data: for example, it requires researchers to provide certain prior biological knowledge during interpretation, especially when defining the starting point of change or selecting a trajectory, which increases the subjectivity and uncertainty of the analysis. Summary of the Invention
[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method, apparatus, device, and medium for single-cell trajectory inference, which can improve the objectivity and determinism of single-cell trajectory inference and increase its reliability without requiring prior biological knowledge from humans.
[0005] To achieve the above objectives, a first aspect of this application proposes a single-cell trajectory inference method, the method comprising:
[0006] Obtain single-cell transcriptome data of the target sample, and generate an association map of multiple cells in the target sample based on the single-cell transcriptome data;
[0007] The developmental relationship graph is obtained by completing the developmental information based on the preset knowledge graph, which is constructed based on the developmental relationships between cells.
[0008] The single-cell trajectory inference result of the target sample is generated based on the developmental relationship diagram.
[0009] Optionally, the step of completing the developmental information of the relationship graph based on a preset knowledge graph to obtain a developmental relationship graph includes:
[0010] The relationships between cells in the relationship diagram are grouped to obtain multiple first paths; the multiple first paths are used to indicate different developmental lineages.
[0011] Query cell development information between cells in each of the first paths in the preset knowledge graph;
[0012] The first path is updated based on the cell development information to obtain the second path;
[0013] By integrating the various second paths, the developmental relationship diagram is obtained.
[0014] Optionally, querying cell development information between cells in each of the first paths in the preset knowledge graph includes:
[0015] The starting cell node is determined based on one of the two cells that are connected in the first path, and the ending cell node is determined based on the other cell.
[0016] Generate a developmental path query statement based on the starting cell node and the ending cell node;
[0017] Based on the developmental path query statement, the cell development information from the starting cell node to the ending cell node is obtained from the preset knowledge graph.
[0018] Optionally, generating a developmental path query statement based on the starting cell node and the ending cell node includes:
[0019] Obtain the potential developmental relationships between cells in the target sample; wherein, the potential developmental relationships include membership relationships and developmental relationships;
[0020] Obtain the path depth range followed by cells in the target sample; wherein the path depth range is used to define the path depth constructed by cell nodes;
[0021] The developmental path query statement is constructed based on the starting cell node, the ending cell node, the potential developmental relationship, and the path depth range.
[0022] Optionally, the developmental relationship diagram includes multiple developmental pathways;
[0023] The generation of single-cell trajectory inference results for the target sample based on the developmental relationship diagram includes:
[0024] Query the weights of the edges corresponding to each path depth in the developmental path; wherein the weights of the edges are constructed based on the number of times the edges are mentioned in literature evidence;
[0025] The path weight of the development path is calculated based on the weight of the edge.
[0026] The target path is obtained by filtering multiple developmental paths based on the path weights.
[0027] By integrating the target path, the single-cell trajectory inference result of the target sample is obtained.
[0028] Optionally, the method further includes at least one of the following steps:
[0029] A global interpretation problem is obtained for all paths in the single-cell trajectory inference result; a large language model is invoked to perform global analysis and interpretation on the global interpretation problem, all paths, and the preset knowledge graph to obtain the global interpretation result;
[0030] Obtain the local interpretation problem between cells in a single path for the single-cell trajectory inference result; call the large language model to perform local analysis and interpretation on the local interpretation problem, the corresponding cell type names between the cells, and the preset knowledge graph to obtain the local interpretation result.
[0031] Optionally, the preset knowledge graph is constructed based on literature recording intercellular relationships, including developmental and non-developmental relationships, and the method further includes:
[0032] Based on two first cells that are related in the relationship graph, query the developmental relationship graph;
[0033] If there are two second cells with two unknown developmental relationships in the developmental relationship diagram, then query the preset knowledge graph for literature information that records other non-developmental relationships between the second cells;
[0034] A large language model is invoked to generate sentiment analysis results based on the literature information; wherein, the sentiment analysis results are used to characterize the support rate or opposition rate for the existence of the association between two second cells.
[0035] To achieve the above objectives, a second aspect of this application provides a single-cell trajectory inference device, the device comprising:
[0036] The first graph construction module is used to acquire single-cell transcriptome data of the target sample and generate an association graph of multiple cells in the target sample based on the single-cell transcriptome data.
[0037] The second graph construction module is used to complete the developmental information of the relationship graph based on a preset knowledge graph to obtain a developmental relationship graph. The preset knowledge graph is constructed based on the developmental relationships between cells.
[0038] The trajectory generation module is used to generate single-cell trajectory inference results for the target sample based on the developmental relationship diagram.
[0039] To achieve the above objectives, a third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect.
[0040] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0041] The single-cell trajectory inference method, apparatus, device, and medium proposed in this application first acquire single-cell transcriptome data of the target sample, and then generate a correlation diagram of multiple cells in the target sample based on the single-cell transcriptome data. This correlation diagram can show whether there are connections between cells. Since the existence of connections between cells is relatively easy to identify, generating the correlation diagram does not require prior biological knowledge. Further, developmental information is supplemented into the correlation diagram based on a pre-defined knowledge graph, resulting in a developmental relationship diagram. This pre-defined knowledge graph is constructed based on the developmental relationships between cells. Thus, the developmental information between cells contained in the pre-defined knowledge graph can be used to complete the correlation diagram into a developmental relationship diagram. This developmental relationship diagram not only shows whether there are connections between cells, but also shows developmental information such as the developmental direction between connected cells. This step also does not require prior biological knowledge. Finally, the single-cell trajectory inference result of the target sample is generated based on the developmental relationship diagram. Since the developmental relationship diagram already contains developmental information such as the developmental relationships of multiple cells, the subjectivity and uncertainty in the trajectory inference process are reduced. In summary, this application improves the objectivity and certainty of single-cell trajectory inference without requiring prior biological knowledge from humans, resulting in higher reliability.
[0042] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0043] Figure 1 This is a flowchart of the single-cell trajectory inference method provided in the embodiments of this application;
[0044] Figure 2 This is a schematic diagram of the relationship diagram provided as an example in this application;
[0045] Figure 3 This is a schematic diagram of the correlation diagram of human peripheral blood provided in one example of this application;
[0046] Figure 4It is a schematic diagram of the developmental relationship diagram provided in the application example;
[0047] Figure 5 yes Figure 1 Flowchart for step 102;
[0048] Figure 6 This is a schematic diagram of the first path provided as an example in this application;
[0049] Figure 7 This is a schematic diagram of the first pathway of human peripheral blood provided as an example in this application;
[0050] Figure 8 This is a schematic diagram illustrating cell development information provided in one example of this application;
[0051] Figure 9 yes Figure 5 Flowchart for step 502;
[0052] Figure 10 yes Figure 1 Flowchart for step 103;
[0053] Figure 11 This is a schematic diagram of the single-cell trajectory inference results of human peripheral blood provided as an example in this application;
[0054] Figure 12 This is a flowchart of a single-cell trajectory inference method provided in this application;
[0055] Figure 13 This is a schematic diagram of the single-cell trajectory inference device provided in the embodiments of this application. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0057] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0059] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0060] High-throughput sequencing, also known as next-generation sequencing technology or massively parallel sequencing (MPS), differs from traditional Sanger sequencing (dideoxygenase method) in that it enables the parallel sequencing of a large number of nucleic acid molecules in a single operation. The core of high-throughput sequencing is obtaining sequence data of nucleic acid information through massive parallelization, covering applications ranging from genomes, transcriptomes, epigenetics, translational levels, and three-dimensional chromatin structure to microbial communities and immune diversity. In the field of transcriptomics, high-throughput sequencing can be specifically applied as single-cell transcriptome sequencing.
[0061] Single-cell RNA sequencing (scRNA-seq) is a sequencing method and analysis process that measures the transcriptome (i.e., the RNA expression profile within a cell) at the level of a single cell. Unlike traditional methods that obtain average expression at the population cell level, scRNA-seq can reveal differences, heterogeneity, and cell type composition between individual cells.
[0062] Single-cell trajectory inference is a method used in single-cell transcriptome data analysis to reconstruct cell differentiation and developmental pathways from single-cell transcriptome data. It analyzes the similarities and differences between cells, arranging them into a continuous trajectory to reveal changes in cell state.
[0063] The PAGA (Partition-based Graph Abstraction) algorithm addresses the challenges of analyzing cellular heterogeneity and continuous developmental trajectories in single-cell RNA sequencing data. It generates statistical models of cell population connectivity, representing continuous and discontinuous cellular structures at multiple resolutions, providing a means to explore cellular variations while preserving the data's topological structure.
[0064] Knowledge graphs are semantic networks that organize, represent, and store domain knowledge using a graph structure (nodes-edges-attributes). Nodes represent entities (concepts, objects, relationships, etc.), edges represent relationships between entities, and attributes describe the characteristics of nodes.
[0065] Large Language Models (LLMs): Models based on large neural networks (usually Transformer architecture) that have the ability to understand, generate, and reason about natural language through self-supervised learning on massive amounts of text data.
[0066] Currently, with the rapid development of high-throughput sequencing technology, especially the application of single-cell transcriptome sequencing, biomedical research has entered the era of big data. These technologies can provide high-resolution cellular state information, enabling researchers to delve into the dynamic changes in cell differentiation, development, and disease processes. Single-cell transcriptome sequencing technology provides the SMART-seq method. This method utilizes reverse transcription and directed amplification strategies to amplify RNA from a single cell and perform high-throughput sequencing. With subsequent technological developments and updates, single-cell transcriptome technology has become more sophisticated, and its cost has gradually decreased.
[0067] In single-cell transcriptome data analysis, trajectory inference methods analyze the similarities and differences between cells to arrange them into a continuous trajectory, thereby revealing the process of cell state changes. The development of trajectory inference methods has gone through several stages. Initially, researchers mainly relied on traditional clustering methods to classify cells, but this method could not capture continuous changes in cell states. To overcome this limitation, the Wishbone algorithm first proposed a pseudo-time-based concept, inferring cell differentiation trajectories by calculating the similarity between cells. Subsequently, various trajectory inference methods emerged, and Wolf et al. developed the PAGA algorithm and integrated it into the Scanpy toolkit. The main goal of this algorithm is to generate a clear cell state transition map by abstracting and simplifying complex cell atlases, thereby helping researchers better understand cell differentiation and development processes. Despite significant progress, these methods still have limitations in interpreting the biological meaning of single-cell data. For example, trajectory inference methods require researchers to provide certain prior biological knowledge (because existing trajectory inference techniques calculate the correlation between data using statistical models to infer developmental pathways. However, these existing techniques rely solely on the calculated correlation results, and choosing an incorrect starting point will lead to erroneous results, thus requiring users to have a certain biological background). This is especially true when defining the starting point of changes or selecting trajectories, which increases the subjectivity and uncertainty of the analysis. Furthermore, these methods mainly focus on the trajectory of changes in cell state, but have limited ability to interpret the underlying biological mechanisms and regulatory networks of these changes. Therefore, there is an urgent need for a method that optimizes trajectory inference and assists in interpretation.
[0068] Based on this, embodiments of this application provide a single-cell trajectory inference method, a single-cell trajectory inference device, an electronic device, and a computer-readable storage medium. Through these embodiments, it is possible to infer the direction of cell development trajectory from single-cell transcriptome data without prior biological knowledge, and to provide reasonable analytical results. These embodiments have three main features: 1. The trajectory inference results can be based on the standard PAGA algorithm, but with the addition of knowledge graph information, potential cell types in the single-cell data can be supplemented, increasing the rationality and directionality of the cell development trajectory. 2. The supplementary results can be traced back to the information source, making the inference results more reliable and highly interpretable. 3. Each developmental path can be further analyzed and judged by combining a large language model, such as identifying key gene expression changes and involved signaling pathways, and generating detailed analysis reports to help researchers more accurately understand the biological significance of cell differentiation trajectories.
[0069] The single-cell trajectory inference method provided in this application can be applied to either a terminal or a server, or it can be software running on the server. The server can be configured as an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The software can be an application that implements the single-cell trajectory inference method, but it is not limited to the above forms.
[0070] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0071] The single-cell trajectory inference method, single-cell trajectory inference device, electronic device and computer-readable storage medium provided in the embodiments of this application are specifically described through the following embodiments. First, the single-cell trajectory inference method in the embodiments of this application is described.
[0072] Please refer to Figure 1 , Figure 1 An optional flowchart of a single-cell trajectory inference method is shown. Figure 1 The method may include, but is not limited to, steps 101 to 103.
[0073] Step 101: Obtain single-cell transcriptome data of the target sample, and generate an association diagram of multiple cells in the target sample based on the single-cell transcriptome data;
[0074] Step 102: Complete the developmental information of the relationship graph based on the preset knowledge graph to obtain the developmental relationship graph. The preset knowledge graph is constructed based on the developmental relationships between cells.
[0075] Step 103: Generate the single-cell trajectory inference results of the target sample based on the developmental relationship diagram.
[0076] Steps 101 to 103, as illustrated in this embodiment, firstly, single-cell transcriptome data of the target sample is acquired, and a relationship diagram of multiple cells in the target sample is generated based on the single-cell transcriptome data. This relationship diagram can show whether there are connections between cells. Since the existence of connections between cells is relatively easy to identify, generating the relationship diagram does not require prior biological knowledge. Further, developmental information is supplemented into the relationship diagram based on a preset knowledge graph to obtain a developmental relationship diagram. This preset knowledge graph is constructed based on the developmental relationships between cells. In this way, the developmental information between cells contained in the preset knowledge graph can be used to complete the relationship diagram into a developmental relationship diagram. This developmental relationship diagram, in addition to showing whether there are connections between cells, further shows developmental information such as the developmental direction between connected cells. This step also does not require prior biological knowledge. Finally, the single-cell trajectory inference result of the target sample is generated based on the developmental relationship diagram. Since the developmental relationship diagram already contains developmental information such as the developmental relationships of multiple cells, the subjectivity and uncertainty in the trajectory inference process are reduced. In summary, this application improves the objectivity and certainty of single-cell trajectory inference without requiring prior biological knowledge from humans, resulting in higher reliability.
[0077] In step 101, the target sample refers to the research object / region to which trajectory inference is to be performed. For example, the target sample is a specified tissue slice or organ. In one example, the target sample is human peripheral blood. Human peripheral blood is the circulating blood in the human body, composed of red blood cells, white blood cells, and platelet cells. In another example, the target sample is a mouse colon, which can be obtained by slicing a cross-section of the mouse intestine.
[0078] Single-cell transcriptome data consists of sequencing data from the transcriptome of a single cell, representing the gene expression of that cell at a specific point in time. Unlike traditional "population-level" transcriptome data (which is the average expression obtained by sequencing a mixture of RNA from thousands of cells), single-cell transcriptome data captures the expression levels of individual cells, thus revealing intercellular heterogeneity, different cell types and states, and cell developmental or differentiation trajectories. For example, single-cell transcriptome data from human peripheral blood samples shows multiple cell types including: hematopoietic stem cells, Kupffer cells (macrophages in the liver), osteoclasts, granulocytes, plasmacytoid dendritic cells, central nervous system macrophages, myeloid suppressor cells, conventional dendritic cells, monocytes, adipose macrophages, and microglial cells. Trajectory inference is essentially the discovery of temporal relationships or developmental lineages among multiple cells in high-dimensional single-cell transcriptome data.
[0079] An association graph contains multiple cell nodes and connecting edges between them. The cell nodes indicate different cell types. The connecting edges indicate relationships between cells. If an edge connects two cell nodes, the cell types corresponding to those two nodes are associated. If no edge connects two cell nodes, the cell types corresponding to those two nodes are not associated. The association graph is undirected, meaning the developmental direction between two associated cells is not recorded.
[0080] In one embodiment, the PAGA algorithm can be used to construct association relationships in single-cell transcriptome data, resulting in an association map of multiple cells in the target sample. This association map is also called a PAGA map. Please refer to [link / reference needed]. Figure 2 The relationship graph includes 5 cell nodes: C1, C2, C3, C4, and C5. The graph also includes connecting edges between the cell nodes, specifically: connecting edge j.23 Connecting edge j 13 Connecting edge j 15 Connecting edge j 45 j mn This indicates that the connection between cell node Cm and cell node Cn is an undirected edge.
[0081] In one example, please refer to Figure 3 , Figure 3 This is a schematic diagram of a PAGA map of human peripheral blood. Figure 3 In the PAGA diagram, there are 11 cell nodes, numbered P1 to P11. Each cell node corresponds to a cell type: Kupffer cells (P1), central nervous system macrophages (P2), microglia (P3), osteoclasts (P4), adipose tissue macrophages (P5), conventional dendritic cells (P6), hematopoietic stem cells (P7), plasmacytoid dendritic cells (P8), monocytes (P9), myeloid suppressor cells (P10), and granulocytes (P11). The PAGA diagram includes connecting edges between cell nodes, specifically: connecting edge j P1P2 Connecting edge j P2P3 Connecting edge j P2P4 Connecting edge j P4P5 Connecting edge j P5P6 Connecting edge j P6P7 Connecting edge j P7P8 Connecting edge j P7P9 Connecting edge j P7P10 Connecting edge j P7P11 j PmPn The edge used to indicate the connection between cell type Pm and cell type Pn is an undirected edge.
[0082] In one embodiment, the PAGA graph can be generated through the following steps: First, the gene expression matrix of the single-cell transcriptome data is preprocessed by normalization and standardization. Then, the matrix is reduced in dimensionality (PCA or UMAP). Next, for each cell, its K nearest neighbors (Euclidean distance) are found to construct a KNN graph. Then, the KNN graph is clustered (Louvain algorithm or Leiden clustering algorithm) to obtain different cell subpopulations, which are then manually or automatically annotated. Subsequently, different cell subpopulations are scaled into different cell nodes according to the number of cells, and the connectivity between these cell nodes (connectivity can be understood as the gene expression similarity between cells) is calculated. Edges with high confidence are retained, and the gene expression similarity is used to construct a PAGA graph containing these edges.
[0083] The PCA mentioned above specifically refers to Principal Component Analysis (PCA). UMAP specifically refers to Uniform Manifold Approximation and Projection (UMAP), an algorithm for dimensionality reduction and data visualization, particularly suitable for handling high-dimensional datasets.
[0084] It should be noted that when single-cell transcriptome data contains a small number of cell types, they are difficult to annotate. This means the PAGA algorithm can only calculate and map the connections between annotated cell types, with the fineness of the calculation heavily dependent on the annotation results. Furthermore, it cannot provide path results with developmental direction information. Finally, the results calculated by the PAGA algorithm can only be analyzed and interpreted by experts with biological knowledge, which is time-consuming. Moreover, existing trajectory inference methods rely on certain biological knowledge to select appropriate starting points; incorrect selection can lead to the algorithm calculating drastically different developmental paths for different cell types, and accumulating biological knowledge requires a significant time investment. To overcome these shortcomings, this application introduces a knowledge graph and a large language model into the trajectory inference process. The specific implementation steps are described below.
[0085] In step 102, the preset knowledge graph is a knowledge graph (also known as a graph database) containing a large amount of biological data. The biological data includes developmental relationships between cells, and this preset knowledge graph can be constructed based on these relationships, thus containing developmental information for multiple cells. For example, the developmental relationships between cells can include the developmental path from cell type A to cell type B, thereby including information such as cell type A, cell type B, and connecting edges (also known as developmental edges) pointing from A to B in the preset knowledge graph.
[0086] In one embodiment, the preset knowledge graph may include, but is not limited to, the following information: (1) the developmental relationship between cells, such as cell type A can develop into cell type B, cell type B can develop into cell type C, etc.; (2) the non-developmental relationship between cells, such as cell type D and cell type E are in the same anatomical structure, cell type E and cell type F have an interaction relationship, etc.; (3) the developmental lineage of each cell type.
[0087] Based on the constructed pre-defined knowledge graph, the developmental relationship graph can be obtained by supplementing the developmental information in the relationship graph using the developmental information carried in the pre-defined knowledge graph. Please refer to [link to relevant documentation]. Figure 4The developmental relationship diagram can include seven cell nodes: C1, C2, C3, C4, C5, C6, and C7. The diagram also includes connecting edges between these cell nodes, specifically: connecting edge J. 17 Connecting edge J 73 Connecting edge J 32 Connecting edge J 56 Connecting edge J 64 J mn This indicates that the edge connecting cell node Cm and cell node Cn is a directed edge, and the direction is from cell node Cm to cell node Cn.
[0088] In one embodiment, please refer to Figure 5 Step 102 may include:
[0089] Step 501: Group the relationships between cells in the relationship graph to obtain multiple first paths;
[0090] Step 502: Query cell development information between cells in each first path in the preset knowledge graph;
[0091] Step 503: Update the first path based on cell development information to obtain the second path;
[0092] Step 504: Integrate the various second paths to obtain the developmental relationship diagram.
[0093] In steps 501 to 504 of the embodiments of this application, during the process of completing the association graph to obtain the developmental relationship graph using a preset knowledge graph, multiple first paths indicating different developmental lineages are first obtained by grouping based on the association graph. Then, each first path is updated individually using the preset knowledge graph. Finally, the updated second paths are integrated to obtain the developmental relationship graph. In this way, the complexity of information completion can be reduced, and the correlation between information completion and developmental lineage can be improved, thereby improving the accuracy of information completion.
[0094] In step 501, by grouping the relationships between cells in the relationship diagram, different first paths can be obtained. Multiple first paths are used to indicate different developmental lineages. A developmental lineage refers to the evolutionary path of cell states during the development of an organism, describing the connections from ancestral cells through a series of differentiation stages to gradually form different mature cell types.
[0095] Understandably, an association graph contains multiple cells and connecting edges between them, representing the relationships between the cells. An association graph is a connected graph, meaning that in an association graph, there exists a path (containing at least one connecting edge) from any cell to any other cell in the graph. By grouping the relationships between cells in the association graph, multiple first paths can be obtained. For example, please refer to... Figure 2 and Figure 6 ,right Figure 2 After grouping the relationships between cells in the relationship diagram shown, we can obtain the following: Figure 6 The first path shown 11 and the first path 12 First path 11 Includes cell nodes C1, C2, and C3, connected by edge j. 23 Connecting edge j 13 First path 12 Includes cell node C4, cell node C5, and connecting edge j 45 .
[0096] In one embodiment, the relationships between cells in the association graph can be grouped according to a threshold parameter of the PAGA algorithm to obtain a first path. For example, please refer to... Figure 3 and Figure 7 When the threshold parameter t = 0.05 in the PAGA algorithm (in the PAGA (Partition-based Graph Abstraction) algorithm, the threshold parameter controls the threshold for path grouping (or edge filtering). This threshold indicates that only when the connection significance (such as statistical confidence, edge weight) between two paths (or node groups) is not less than t is a "valid edge" or "valid connection" considered to exist between them), by... Figure 3 The PAGA diagram shown contains connecting edges j P1P2 Connecting edge j P2P3 Connecting edge j P2P4 Connecting edge j P4P5 Connecting edge j P5P6 Connecting edge j P6P7 Connecting edge j P7P8 Connecting edge j P7P9 Connecting edge j P7P10 Connecting edge j P7P11 By grouping, we can obtain Figure 7 The first path shown 13 (Including cell node P1), first path Path 14 (Including cell node P2, cell node P3, cell node P4, and connecting edge j)P2P3 Connecting edge j P2P4 ), First Path 15 (Cell node P5, cell node P6, cell node P7, cell node P8, cell node P9, cell node P10, cell node P11, connecting edge j) P5P6 Connecting edge j P6P7 Connecting edge j P7P8 Connecting edge j P7P9 Connecting edge j P7P10 Connecting edge j P7P11 Because of the first path 13 The first path contains only one cell node P1. 13 It can be deleted; subsequent steps only apply to the first path. 14 and the first path 15 Update.
[0097] In step 502, as mentioned above, the preset knowledge graph contains developmental information for multiple cells. At this point, cell developmental information between cells can be read from the developmental information contained in the preset knowledge graph.
[0098] In one example, cell development information is used to indicate attributes of developmental relationships between cells. Attributes of developmental relationships include belonging (is_a) and development (develop_into). For example, please refer to... Figure 6 and Figure 8 According to the first path 11 Cell nodes C2 and C3 in the graph query cell development information. The obtained cell development information can include the development of cell node C3 into cell node C2. Figure 8 (C3->C2 shown).
[0099] In one example, cell development information is used to indicate the cells traversed by intercellular developmental pathways and the properties of intercellular relationships. For example, see [link to relevant documentation]. Figure 6 and Figure 8 According to the first path 11 Cell nodes C1 and C3 in the preset knowledge graph query cell development information. The obtained cell development information can include the development of cell node C1 into cell node C7. Figure 8 The path shown is C1->C7->C1. 12 Cell nodes C4 and C5 in the preset knowledge graph query cell development information. The obtained cell development information can include the development of cell node C5 into cell node C4 through cell node C6. Figure 8 (C5->C6->C4 shown).
[0100] It should be noted that, Figure 6 The circles shown are used to indicate cell nodes in the first path. Figure 8 The hexagons shown are used to indicate cell nodes in the preset knowledge graph.
[0101] The cell development information described above can be organized into triplets, such as (Cell1, is_a / develop_into, Cell2). For example, for the case where cell node C3 develops into cell node C3, the triplets would be (C2, develop_into, C3). Similarly, for the case where cell node C5 develops into cell node C4 via cell node Ci, the triplets would be (C5, develop_into, Ci) and (Ci, develop_into, C4).
[0102] In one embodiment, please refer to Figure 9 Step 502 may include:
[0103] Step 901: Determine the starting cell node based on one of the two cells with a connection relationship in the first path, and determine the ending cell node based on the other cell;
[0104] Step 902: Generate a developmental path query statement based on the starting cell node and the ending cell node;
[0105] Step 903: Based on the developmental path query statement, retrieve the cell development information from the starting cell node to the ending cell node from the preset knowledge graph.
[0106] Steps 901 to 903 as shown in the embodiments of this application first identify the starting cell node and the ending cell node, then generate a developmental path query statement based on this, and then query the preset knowledge graph according to the statement. This simplifies the query process, enables high-concurrency queries, and improves query efficiency.
[0107] In step 901, one of the two cells with a connection relationship in the first path can be randomly selected as the starting cell node, and the other cell can be selected as the ending cell node.
[0108] For example, please refer to Figure 6 For the first path 11 There is a connecting edge j in the middle. 23 Given cell nodes C2 and C3, cell node C2 can be used as the starting cell node and cell node C3 as the ending cell node. For the first path... 12 There is a connecting edge j in the middle.45 Cell nodes C4 and C5 can be used as the starting cell node and the ending cell node.
[0109] In step 902, a developmental path query statement refers to a requesting expression used to describe, retrieve, or propose specific questions, conditions, or rules regarding the evolutionary path of cell states when analyzing developmental lineages or cell differentiation trajectories. A developmental path query statement includes a starting cell node, an ending cell node, and query conditions. Query conditions include at least one of the following: the potential developmental relationships followed by cells and the depth range of the paths followed by cells.
[0110] In one embodiment, step 902 may include: obtaining the potential developmental relationships between cells of the target sample; obtaining the path depth range between cells of the target sample; and constructing a developmental path query statement based on the starting cell node, the ending cell node, the potential developmental relationships, and the path depth range.
[0111] Specifically, potential developmental relationships include belonging relationships and developmental relationships. Path depth intervals are used to define the path depth constructed by cell nodes. Path depth intervals include minimum path depth and maximum path depth. Minimum path depth indicates the minimum path depth constructed by cell nodes. Maximum path depth indicates the maximum path depth constructed by cell nodes. For example, if the target sample is human peripheral blood, the path depth interval could be [1, 4].
[0112] The advantage of the above embodiments is that they can generate personalized developmental path query statements for target samples, thereby improving the flexibility of statement construction.
[0113] In one example, the developmental path query statement can be as follows:
[0114] {MATCH path = (n:Cell) - [r:is_a|develope_into*1..4] - (m:Cell) ,
[0115] WHERE n.name = $source AND m.name = $target ,
[0116] RETURN DISTINCT path, nodes(path),relationships(path)}.
[0117] In the developmental path query statement, "1..4" indicates a path depth range of [1,4], $source is the cell type name of the starting cell node, $target is the cell type name of the ending cell node, path is the queried developmental path, nodes(path) are the cells on the developmental path, and relationships(path) are the relationships between cells on the developmental path (either is_a or develop_into). This query statement is used to identify all developmental paths (differentiation / inheritance paths) with a length ≤ 4 from $source to $target in the knowledge graph and return cell development information, including the complete path, a list of nodes, and a list of edges.
[0118] In step 903, as can be seen from the previous example, cell development information may include the developmental path, the cells along the developmental path, and the relationships between cells along the developmental path.
[0119] It should be noted that for a single developmental path query statement, cell development information may include multiple developmental paths, with different cells on each path. For example, assuming the starting cell node is Cell n and the ending cell node is Cell m, the cell development information may include the following developmental paths: (1) Cell n->Cell k1->Cell m; (2) Cell n->Cell k2->Cell m; (3) Cell n->Cell k2->Cell k3->Cell m. i Used to characterize intermediate cell nodes on the developmental pathways of Cell n and Cell m.
[0120] In step 503, the first path is updated based on cell development information to obtain the second path. For example, please refer to... Figure 6 First path 11 There is a relationship between cell nodes C1 and C3, and a relationship between cell node C3 and cell node C2. Assuming the cell development information indicates that cell node C1 develops into cell node C7, then cell node C7 develops into cell node C3, and finally cell node C3 develops into cell node C2, then the second path is C1->C7->C3->C2. The first path is... 12 There is a relationship between cell node C4 and cell node C5. Assuming that the cell development information indicates that cell node C5 develops into cell node C6, and then cell node C6 develops into node cell C4, then the second path includes C5->C6->C4.
[0121] For example, there is a relationship between hematopoietic stem cells and myeloid suppressor cells in the relationship graph. After adding cell development information from the preset knowledge graph, a more comprehensive and directional developmental path can be obtained: hematopoietic stem cell -> leukocyte -> myeloid leukocyte -> myeloid suppressor cell.
[0122] In step 504, the various second paths are integrated to obtain a developmental relationship diagram. For example, each second path is used as a developmental path to obtain a developmental relationship diagram. Alternatively, the second paths can be filtered / deduplicated, and then the filtered second paths can be used as developmental paths to obtain a developmental relationship diagram.
[0123] In step 103, the single-cell trajectory inference results of the target sample can be generated based on the developmental relationship diagram obtained in step 102. The developmental relationship diagram includes multiple developmental paths. The single-cell trajectory inference results are generated by analyzing single-cell transcriptome data to infer the evolutionary trajectory of cell state over time (or pseudo-time), as well as information such as branches, differentiation nodes, and transformation probabilities.
[0124] In one embodiment, please refer to Figure 10 Step 103 may include:
[0125] Step 1001: Query the weight of the edge corresponding to the depth of each path in the development path;
[0126] Step 1002: Calculate the path weight of the development path based on the edge weights;
[0127] Step 1003: Filter multiple developmental paths according to path weights to obtain the target path;
[0128] Step 1004: Integrate the target path to obtain the single-cell trajectory inference result of the target sample.
[0129] In step 1001, the weight of the edge is used to indicate the probability / likelihood of a developmental relationship between cells.
[0130] In one example, the weights of the edges of a developmental path can be obtained from a pre-defined knowledge graph. Therefore, the process of querying the weights of the edges of a developmental path can include: querying the weights of the edges of the developmental path in the pre-defined knowledge graph. Specifically, the pre-defined knowledge graph is constructed based on literature recording the relationships between cells. These public database records and / or published research articles recording the relationships between cells can be referred to as literature evidence. The literature evidence contains relevant records of the edges of the developmental path; each record is counted as one mention. The weight of the edge is constructed based on the number of times the edge is mentioned in the literature evidence. The weight of the edge is directly proportional to the number of times the edge is mentioned in the literature evidence.
[0131] In another example, the weights of the edges of the aforementioned developmental path can be constructed in real time by the number of times the edges are mentioned in the literature. The process of querying the weights of the edges of the developmental path can include: obtaining the number of times the edges are mentioned in the literature; and performing weight mapping based on the number of mentions and a preset increasing function to obtain the weights of the edges.
[0132] In step 1002, the path weight is used to indicate the reliability of the development path. The average weight of the depths of each path in the development path can be used as the path weight. Other calculation methods can also be used, and this embodiment does not limit this. For example, for the development path C1->C7->C3, its path weight is... ,in, The weights for the path depths of C1 and C7, The weights for the path depths of C7 and C3.
[0133] In step 903, the development path with the highest path weight can be selected based on the path weight of each development path to obtain the target path.
[0134] In one embodiment, step 1003 may include: grouping developmental paths based on the presence of the same starting and ending cell nodes to obtain candidate path groups; and selecting the path with the highest path weight from each candidate path group to obtain the target path. This allows for the acquisition of target paths with more comprehensive developmental information and higher path reliability, thereby improving the accuracy of trajectory inference.
[0135] In step 1004, deduplication, merging, and other integration operations can be performed on the target path to obtain the single-cell trajectory inference result. For example, please refer to... Figure 4 The single-cell trajectory inference results include two single-cell trajectories: C1->C7->C3->C2 and C5->C6->C4.
[0136] For example, please refer to Figure 11The inferred single-cell trajectory results from human peripheral blood include: hematopoietic stem cells, blood cells, leukocytes, granulocytes, myeloid leukocytes, myeloid suppressor cells, monocytes, dendritic cells, plasmacytoid dendritic cells, conventional dendritic cells, adipose macrophages, and developmental edges e1, e2, e11, e21, e211, e212, e213, e22, e221, e222, and e2221. Each developmental edge indicates a developmental relationship between cells. It should be noted that... Figure 11 In the diagram, ellipses represent the origins of each cell from the relational graph, and hexagons represent the origins of cells from the preset knowledge graph.
[0137] In one embodiment, after step 103, the single-cell trajectory inference method may further include: calling a large language model to analyze and interpret the single-cell trajectory inference results and a preset knowledge graph to obtain an interpretation result. The large language model includes, but is not limited to, GPT, Gemini, Wenxin, Qianwen, and Claude3. The interpretation result can be a human-understandable text analysis and interpretation report. The large language model analyzes and interprets the single-cell trajectory inference results and the preset knowledge graph from multiple biological perspectives to obtain the interpretation result.
[0138] In one embodiment, the process of calling a large language model to analyze and interpret the single-cell trajectory inference results and the preset knowledge graph to obtain the interpretation results may include: obtaining a global interpretation problem for all paths of the single-cell trajectory inference results; calling a large language model to perform global analysis and interpretation of the global interpretation problem, all paths, and the preset knowledge graph to obtain a global interpretation result.
[0139] Specifically, for the single-cell trajectory inference results (hereinafter referred to as the developmental tree) obtained from the preset knowledge graph query, the large language model is invoked to provide judgment criteria and reasonable explanations for all paths in the developmental tree. The judgment criteria include the single-cell trajectory inference results and the cell development information queried from the preset knowledge graph. In providing reasonable explanations, the large language model interprets key regulatory genes involved in the cell development process, as well as the cell signal transduction pathways involved by these genes, such as cell differentiation and neuronal axon generation. These pathways may coincide with the developmental process, which can help users further judge the reasonableness of the developmental relationship.
[0140] The above embodiments analyze and interpret the entire pathway from a holistic perspective, focusing on key regulatory genes and related signaling pathways, resulting in more granular results and a more accurate and detailed global interpretation.
[0141] In one example, for the inference results of single-cell trajectories in human peripheral blood, the global interpretation results given by large language models (such as Qwen-max) can include the following: (1) When analyzing two cell development trajectories (#final_graph_p0# and #final_graph_p1#), we can understand these developmental pathways from multiple biological levels, including cell type, gene expression, signaling pathways and regulatory mechanisms. These biological pathways not only reveal the development of cells from one stage to another, but also reveal the relative probability of the branching structure and state transitions of the trajectory. (2) In #final_graph_p0#, we can see the developmental pathways from "hematopoietic stem cells" to "white blood cells" and "blood cells". The intercellular developmental relationships corresponding to these pathways can be found in existing biological mappings. Hematopoietic stem cells are pluripotent stem cells that can differentiate into multiple blood cell types. In this process, key genes such as GATA1, PU.1 and RUNX1 regulate the differentiation direction of stem cells. Signaling pathways such as JAK-STAT and Notch signal transduction also play a key role in regulating these differentiation processes. (3) The developmental pathways from “leukocytes” to “dendritic cells” and “myeloid leukocytes” are known in current biological knowledge. Dendritic cells, as antigen-presenting cells, play a key role in the immune response, and their differentiation involves the regulation of IRF8 and BATF3 genes. The development of myeloid leukocytes is regulated by transcription factors such as C / EBP-α and C / EBP-β. These differentiation processes depend not only on the expression of specific genes but are also regulated by external signals such as cytokines. (4) The developmental pathways from “dendritic cells” to “conventional dendritic cells” and “plasma cell-like dendritic cells” continue to show the process of further differentiation of dendritic cells. The development of conventional dendritic cells is regulated by Zbtb46 and ID2 genes, while plasma cell-like dendritic cells mainly depend on the expression of IRF7. These differentiation pathways reveal the specific roles of different dendritic cell subtypes in immune function. (5) In #finalgraph p1#, the path from “central nervous system macrophages” to “microglia” indicates that this developmental relationship is also known in biological knowledge. Microglia are resident macrophages in the central nervous system, and their development is regulated by genes such as CSFR and TREM2. Signaling pathways such as TGF-β and Wnt also play important roles in microglial differentiation. (6) The red border from “Kupffer cells” to “central nervous system macrophages” indicates that this developmental relationship was inferred from the dataset using the PAGA algorithm. Kupffer cells are specialized macrophages in the liver, and the transformation process may involve cross-organ signal transduction and gene regulation mechanisms. These inferences require further experimental verification to confirm their biological validity.(7) Overall, these cell developmental trajectories combine existing biological knowledge with new data inferences, demonstrating the complex developmental process from pluripotent stem cells to specific functional cell types. Gene expression, signaling pathways, and cell-cell interactions collectively determine the developmental fate of each cell. These mappings not only help us better understand the fundamental mechanisms of cell differentiation but also provide important references and directions for future research.
[0142] In one embodiment, the process of calling a large language model to analyze and interpret the single-cell trajectory inference result and the preset knowledge graph to obtain the interpretation result may include: obtaining the local interpretation problem between cells in a single path for the single-cell trajectory inference result; calling the large language model to perform local analysis and interpretation on the local interpretation problem, the corresponding cell type names between cells, and the preset knowledge graph to obtain the local interpretation result.
[0143] This embodiment analyzes from a local perspective, allowing users to ask multiple rounds of follow-up questions about the two cell types before and after any path, i.e., any edge in the developmental tree. For example, the local interpretation question is "What is the relationship between cell A and cell B in the developmental tree, and why?". The large language model generates the basis for judgment and reasonable explanations for the cell types before and after development, i.e., it generates the local interpretation result.
[0144] The advantage of the above embodiments is that they can provide users with local interpretation results about two cell types before and after development, improving the flexibility of interpreting single-cell trajectory inference results.
[0145] In one embodiment, after step 103, the single-cell trajectory inference method may further include: querying the developmental relationship graph based on two first cells that have a relationship in the relationship graph; if there are two second cells with an unknown developmental relationship in the developmental relationship graph, querying literature information in a preset knowledge graph that records other non-developmental relationships between the second cells; and calling a large language model to generate sentiment analysis results based on the literature information.
[0146] The pre-defined knowledge graph is constructed based on literature documenting inter-cell relationships. These relationships include developmental and non-developmental relationships. Developmental relationships include belonging relationships (is_a) and developmental relationships (develop_into). Non-developmental relationships include relationships within the same anatomical structure and relationships involving interaction. Sentiment analysis results are used to characterize the support or opposition rates for relationships between two second cells.
[0147] In one example, when two first cells with a relationship in the association graph do not have a corresponding developmental relationship in the developmental graph, it indicates that the two cells do not have a recognized direct developmental relationship (this direct developmental relationship includes is_a or develop_into), but other relationships may exist (such as being in the same anatomical structure, having an interaction relationship, etc.). In this case, these two first cells can be considered as second cells. For example, T lymphocytes and antigen-presenting cells originate from different lineages, but their interaction is key to initiating specific immunity. In this case, the large model can retrieve all literature information in the predefined knowledge graph that records other non-developmental relationships between two second cells. The large language model will then provide sentiment analysis results based on this literature information, such as the proportion of literature supporting the association between T lymphocytes and antigen-presenting cells, and the proportion of literature opposing the association between T lymphocytes and antigen-presenting cells.
[0148] Please see Figure 12 In one example, a single-cell trajectory inference method may include:
[0149] (1) Obtain single-cell transcriptome data of the target sample;
[0150] (2) Obtain the preset knowledge graph;
[0151] (3) Generate a correlation diagram of multiple cells in the target sample based on single-cell transcriptome data;
[0152] (4) Group the relationships between cells in the relationship diagram to obtain multiple first paths;
[0153] (5) Query cell development information between cells in each first path in the preset knowledge graph;
[0154] (6) Update the first path based on cell development information to obtain the second path, integrate the various second paths to obtain the developmental relationship diagram, and generate the single-cell trajectory inference result of the target sample based on the developmental relationship diagram;
[0155] (7) Obtain the global interpretation problem of all paths for the single-cell trajectory inference result, call the large language model to perform global analysis and interpretation of the global interpretation problem, all paths and the preset knowledge graph, and obtain the global interpretation result;
[0156] (8) Obtain the local interpretation problem of any two cells in a single path for the single-cell trajectory inference result, call the large language model to perform local analysis and interpretation of the local interpretation problem, the cell type name of any two cells, and the preset knowledge graph, and obtain the local interpretation result.
[0157] Specifically, this example uses the PAGA algorithm to process single-cell transcriptome data and generate a PAGA graph (i.e., a relational graph). Then, based on a pre-constructed knowledge graph containing a large amount of biological data, the cell types annotated in the PAGA graph are queried. Using the attributes `is_a` and `develop_into`, which represent inter-cell relationships, cell developmental information related to these cell types can be obtained from the pre-constructed knowledge graph. Next, the complete inter-cell developmental path, i.e., the single-cell trajectory inference result, is obtained by retaining the path with the highest weight among the cells. This single-cell trajectory inference result is then input into a large language model (such as GPT), allowing the large language model to analyze and interpret the path from multiple biological perspectives. This approach significantly reduces the time and workload for users to calculate and interpret inter-cell developmental paths. Compared to analysis and interpretation methods based solely on traditional PAGA algorithms or large language models, the results are traceable, interpretable, and more complete.
[0158] This application allows users to follow up on each developmental path in the single-cell trajectory inference results. By querying the literature information on intercellular developmental relationships recorded in the preset knowledge graph, it can provide statistical data on how many documents support or oppose the existence of intercellular developmental relationships, thereby obtaining more accurate and comprehensive interpretation results.
[0159] The benefits of the above example are: (1) A knowledge graph-based single-cell trajectory inference method has been developed. This method fully respects the authenticity of the data. Without tampering with the data of the association graph (which can be calculated based on the PAGA algorithm), it uses the biological information recorded in the knowledge graph to complete the developmental information of the association graph, ensuring that the inferred path is correct and authentic. (2) It has a visualization function for the single-cell trajectory inference results after knowledge graph completion, and allows users to use a large language model to conduct multi-round question and answer on the single-cell trajectory inference results, and generate human-understandable text analysis and interpretation reports for all paths or a single path in the single-cell trajectory inference results.
[0160] In conjunction with the above embodiments, the beneficial effects of this application are as follows: (1) Information completion is achieved by using a pre-set knowledge graph, which avoids the need for users to have certain biological background knowledge in advance, greatly shortens the time and workload for users to search and understand developmental relationships in the literature, and the results are traceable and have higher credibility. (2) The path can be analyzed and interpreted from multiple biological perspectives, both as a whole and in parts. (3) Reasonable questions can be raised about paths that exist in the relationship graph but whose developmental relationships are not recorded in the graph, and sentiment analysis can be performed on literature information and prompts can be provided to users.
[0161] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.
[0162] This application also discloses a single-cell trajectory inference device, referring to... Figure 13 The single-cell trajectory inference device includes: a first graph construction module 1301, a second graph construction module 1302, and a trajectory generation module 1303. The first graph construction module 1301 is used to acquire single-cell transcriptome data of the target sample and generate a relationship graph of multiple cells in the target sample based on the single-cell transcriptome data; the second graph construction module 1302 is used to complete the relationship graph with developmental information based on a preset knowledge graph to obtain a developmental relationship graph; and the trajectory generation module 1303 is used to generate the single-cell trajectory inference result of the target sample based on the developmental relationship graph.
[0163] In one embodiment, the single-cell trajectory inference device further includes a trajectory interpretation module, which is used to: call a large language model to analyze and interpret the single-cell trajectory inference results and a preset knowledge graph to obtain interpretation results.
[0164] In one embodiment, the single-cell trajectory inference device further includes a sentiment analysis module, used for: querying a developmental relationship graph based on two first cells that have a relationship in the relationship graph; when there is no developmental relationship corresponding to the two first cells in the developmental relationship graph, treating both first cells as unknown relationship cells; querying literature information in a preset knowledge graph that records other relationships between unknown relationship cells; and calling a large language model to generate sentiment analysis results based on the literature information; wherein the sentiment analysis results are used to characterize the support rate or opposition rate of the relationship between the two unknown relationship cells.
[0165] It should be noted that the specific implementation of this single-cell trajectory inference device is basically the same as the specific implementation of the single-cell trajectory inference method described above, and will not be repeated here.
[0166] This application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the single-cell trajectory inference method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0167] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the single-cell trajectory inference method described above.
[0168] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0169] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0170] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0171] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0172] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0173] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0174] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0175] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0176] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0177] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0178] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0179] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for inferring the trajectory of a single cell, characterized in that, The method includes: Obtain single-cell transcriptome data of the target sample, and generate an association map of multiple cells in the target sample based on the single-cell transcriptome data; The relationships between cells in the relationship diagram are grouped to obtain multiple first paths; the multiple first paths are used to indicate different developmental lineages. In a preset knowledge graph, query the cell development information between cells in each of the first paths. The preset knowledge graph is constructed based on the developmental relationships between cells. The cell development information is used to indicate the cells that the developmental path between cells passes through and the attributes of the relationships between cells. The attributes include belonging relationships and developmental relationships. The first path is updated based on the cell development information to obtain the second path; By integrating the various second paths, a developmental relationship diagram is obtained; The single-cell trajectory inference result of the target sample is generated based on the developmental relationship diagram.
2. The method according to claim 1, characterized in that, The step of querying cell development information between cells in each of the first paths in the preset knowledge graph includes: The starting cell node is determined based on one of the two cells that are connected in the first path, and the ending cell node is determined based on the other cell. Generate a developmental path query statement based on the starting cell node and the ending cell node; Based on the developmental path query statement, the cell development information from the starting cell node to the ending cell node is obtained from the preset knowledge graph.
3. The method according to claim 2, characterized in that, The step of generating a developmental path query statement based on the starting cell node and the ending cell node includes: Obtain the potential developmental relationships between cells in the target sample; wherein, the potential developmental relationships include membership relationships and developmental relationships; Obtain the path depth range followed by cells in the target sample; wherein the path depth range is used to define the path depth constructed by cell nodes; The developmental path query statement is constructed based on the starting cell node, the ending cell node, the potential developmental relationship, and the path depth range.
4. The method according to claim 1, characterized in that, The developmental relationship diagram includes multiple developmental pathways; The generation of single-cell trajectory inference results for the target sample based on the developmental relationship diagram includes: Query the weights of the edges corresponding to the depths of each path in the developmental path; wherein the weights are constructed based on the number of times the edges are mentioned in literature evidence; The path weight of the development path is calculated based on the weight of the edge. The target path is obtained by filtering multiple developmental paths based on the path weights. By integrating the target path, the single-cell trajectory inference result of the target sample is obtained.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes at least one of the following steps: The problem is to obtain a global interpretation of all paths in the single-cell trajectory inference results; The large language model is invoked to perform a global analysis and interpretation of the global interpretation problem, all paths, and the preset knowledge graph, to obtain the global interpretation result; The problem of local interpretation between cells in a single path for the single-cell trajectory inference result; The large language model is invoked to perform local analysis and interpretation on the local interpretation problem, the corresponding cell type names between the cells, and the preset knowledge graph to obtain the local interpretation results.
6. The method according to any one of claims 1 to 4, characterized in that, The preset knowledge graph is constructed based on literature recording intercellular relationships, including developmental and non-developmental relationships. The method further includes: Based on two first cells that are related in the relationship graph, query the developmental relationship graph; If there are two second cells with unknown developmental relationships in the developmental relationship diagram, then the literature information that records other non-developmental relationships between the second cells is queried in the preset knowledge graph. A large language model is invoked to generate sentiment analysis results based on the literature information; wherein, the sentiment analysis results are used to characterize the support rate or opposition rate for the existence of the association between two second cells.
7. A single-cell trajectory inference device, characterized in that, The device includes: The first graph construction module is used to acquire single-cell transcriptome data of the target sample and generate an association graph of multiple cells in the target sample based on the single-cell transcriptome data. The second graph construction module is used to group the relationships between cells in the relationship graph to obtain multiple first paths; the multiple first paths are used to indicate different developmental lineages; the module queries the cell development information between cells in each first path in a preset knowledge graph, which is constructed based on the developmental relationships between cells. The cell development information is used to indicate the cells traversed by the developmental path and the attributes of the cell relationships, including belonging relationships and developmental relationships; the module updates the first paths according to the cell development information to obtain second paths; and the module integrates the various second paths to obtain a developmental relationship graph. The trajectory generation module is used to generate single-cell trajectory inference results for the target sample based on the developmental relationship diagram.
8. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Single cell development trajectory reconstruction method and device, electronic equipment and storage medium
CN116453598A
Gene regulatory network construction method based on cell dynamic differentiation
CN116504314A
Intercellular signal transduction inference method and device based on optimal transmission of knowledge graph, equipment and medium
CN119339796A