Cell fate determinant identification method and related device
By constructing cell lineages and optimizing gene regulation models, cell fate determinants are identified, solving the problem that the sparsity and similarity of cell cluster gene regulation models are not considered in existing technologies, and achieving higher accuracy in identifying fate determinants.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUBEI UNIV OF TECH
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies cannot effectively take into account the sparsity and similarity of cell cluster gene regulation models, resulting in poor accuracy in identifying cell fate determinants.
By constructing cell lineages, the expression levels of transcription factors and target genes in single-cell transcriptome sequencing data were determined. Based on the cell lineages, a gene regulation model was constructed, and the regulatory coefficients were optimized through an objective function to identify differentially expressed transcription factors and determine cell fate determinants.
This improved the accuracy of gene regulatory network reconstruction, ensuring the continuity of gene regulatory models and the accuracy of fate determinant determination during development.
Smart Images

Figure CN121983134A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, and in particular to a method and related apparatus for identifying cell fate determinants. Background Technology
[0002] Biological development encompasses the process by which primitive stem cells divide, differentiate, and form fully functional cells. In the process of biological development, starting from a fertilized egg, cells grow and differentiate, with each cell expressing different genes in an orderly manner. Although they share the same genetic code, they ultimately meet vastly different cellular fates.
[0003] Identifying the genes and their regulatory functions that play a crucial role in cell fate at different stages of development, thereby determining cell fate determinants, has always been a fundamental and important problem in the life sciences. Research shows that cell fate determination during biological development is regulated by several key transcription factors and their related regulatory functions. Current techniques only construct the gene regulatory network of the entire cell "population" from a macroscopic perspective, neglecting the cell cluster-specific characteristics of gene expression regulation, ultimately leading to low accuracy in determinant identification. Considering that different cells share the same genetic code during development, the structure of the gene expression regulatory network within a cell cluster should change continuously rather than discontinuously between adjacent cell clusters. Therefore, how to integrate the topological structure of the cell lineage tree, infer the cell cluster-specific gene regulatory network, and then identify cell fate determinants at different developmental stages within the cell lineage is a challenging research area in bioinformatics.
[0004] Therefore, existing technologies cannot take into account the sparsity of gene regulation models in cell clusters and the similarity of gene regulation models between cell clusters in the identification of cell fate determinants, resulting in poor accuracy in the identification of cell fate determinants. Summary of the Invention
[0005] In view of this, it is necessary to provide a method and related device for identifying cell fate determinants, in order to solve the problem that the existing technology cannot take into account the sparsity of gene regulation models of cell clusters and the similarity of gene regulation models between cell clusters in the identification of cell fate determinants, resulting in poor accuracy in the identification of cell fate determinants.
[0006] To address the aforementioned problems, in a first aspect, the present invention provides a method for identifying cell fate determinants, comprising: Determine the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, and construct cell lineages based on single-cell transcriptome sequencing data; A gene regulation model was constructed based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster within a cell lineage. An objective function is constructed with the aim that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation model between each adjacent cell cluster is within a preset similarity range. Based on the objective function, the regulation coefficient of transcription factors on target genes in each cell cluster is determined. Differential transcription factors between adjacent cell clusters are identified based on regulatory coefficients, and cell fate determinants are determined based on these differential transcription factors.
[0007] In one possible implementation, determining the expression levels of transcription factors and target genes within cells from single-cell transcriptome sequencing data includes: Initial single-cell transcriptome sequencing data were obtained, and the genes were filtered according to the number of cells in which each gene was expressed to obtain transcription factors and target genes. The expression levels of transcription factors and target genes in different cells were normalized to obtain a gene expression matrix representing the expression levels of transcription factors and target genes in cells.
[0008] In one possible implementation, cell lineages are constructed based on single-cell transcriptome sequencing data, including: Obtain cell cluster information from single-cell transcriptome sequencing data; A directed tree relationship between cell clusters is constructed based on the cell differentiation process to obtain cell lineages.
[0009] In one possible implementation, a gene regulation model is constructed based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster within a cell lineage, including: For each cell cluster, a linear gene regulation model is constructed with the expression level of the target gene in the cell as the dependent variable, the expression level of the transcription factor in the cell as the independent variable, the regulatory coefficient of the transcription factor on the target gene as the slope, and Gaussian random noise as the error term.
[0010] In one possible implementation, the objective function is:
[0011] in, Let A represent the regulatory coefficient of transcription factors on target genes, B represent the precision range of gene regulation models for each cell cluster, C represent the sparsity range of gene regulation models for each cell cluster, and D represent the similarity range of gene regulation models between adjacent cell clusters. and These are the preset hyperparameters.
[0012] In one possible implementation, the regulatory coefficients of transcription factors on target genes in each cell cluster are determined based on an objective function, including: A pre-defined convex optimization algorithm was used to collaboratively optimize the objective function and determine the regulatory coefficients of transcription factors on the target gene in each cell cluster. Linear gene regulation models for each cell cluster were determined based on the regulatory coefficients of transcription factors on target genes.
[0013] In one possible implementation, differentially expressed transcription factors between adjacent cell clusters are determined based on regulatory coefficients, and cell fate determinants are determined based on these differentially expressed transcription factors, including: Based on a linear gene regulation model of each adjacent cell cluster, differentially expressed transcription factors that exhibit different regulation of target genes in each adjacent cell cluster were identified. The cell fate determinant type of differentially expressed transcription factors is determined based on their presence in adjacent cell clusters.
[0014] In a second aspect, the present invention also provides a cell fate determinant recognition device, comprising: The cell lineage determination module is used to determine the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, and to construct cell lineages based on single-cell transcriptome sequencing data. The model building module is used to construct gene regulation models based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster within a cell lineage. The regulation coefficient determination module is used to construct an objective function with the aim that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation model between each adjacent cell cluster is within a preset similarity range. Based on the objective function, the regulation coefficient of transcription factors on target genes in each cell cluster is determined. The fate determinant identification module is used to identify differentially expressed transcription factors between adjacent cell clusters based on regulatory coefficients, and to identify cell fate determinants based on differentially expressed transcription factors.
[0015] Thirdly, the present invention also provides an electronic device, including a memory and a processor, wherein, Memory, used to store programs; A processor, coupled to a memory, is used to execute a program stored in the memory to implement the steps in the cell fate determinant identification method of any of the above embodiments.
[0016] Fourthly, the present invention also provides a computer-readable storage medium for storing a computer-readable program or instructions, which, when executed by a processor, can implement the steps in the cell fate determinant identification method of any of the above embodiments.
[0017] The beneficial effects of this invention are as follows: The cell fate determinant identification method provided by this invention determines the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, constructs cell lineages based on single-cell transcriptome sequencing data, and constructs gene regulation models based on the relationship between the expression levels of each transcription factor and each target gene in each cell cluster within the cell lineage. This not only mines gene regulation information from single-cell transcriptome sequencing data but also integrates the topological structure of biological developmental cell lineages, thereby improving the accuracy of gene regulation network reconstruction and ensuring the continuity of gene regulation models during development. An objective function is constructed with the goals of ensuring that the accuracy of gene regulation models for each cell cluster is within a preset accuracy range, the sparsity of gene regulation models for each cell cluster is within a preset sparsity range, and the similarity of gene regulation models between adjacent cell clusters is within a preset similarity range. Based on the objective function, the regulatory coefficients of transcription factors on target genes in each cell cluster are determined. Based on the regulatory coefficients, differentially expressed transcription factors between adjacent cell clusters are identified, and cell fate determinants are determined based on the differentially expressed transcription factors. By analyzing the differences in gene regulation network structure between adjacent cell clusters in the cell lineage, transcriptional regulatory relationships and fate determinants that play a key role in cell fate determination at different developmental stages are identified, aiming to elucidate the cell fate determination mechanism in the developmental process. By controlling the sparsity of gene regulation models for each cell cluster and the similarity of gene regulation models between adjacent cell clusters, the accuracy of gene regulation models for each cell cluster and the accuracy of the determination of fate determinants can be ensured. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a method for identifying cell fate determinants provided in an embodiment of the present invention; Figure 2 A schematic flowchart illustrating a gene expression matrix construction method provided in an embodiment of the present invention; Figure 3 A schematic flowchart illustrating a method for constructing cell lineages according to an embodiment of the present invention; Figure 4A schematic diagram of embryonic development provided in an embodiment of the present invention; Figure 5 A schematic diagram of a cell lineage provided in an embodiment of the present invention; Figure 6 A flowchart illustrating a method for solving a gene regulation model provided in an embodiment of the present invention; Figure 7 A flowchart illustrating a method for determining fate determinants provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of a cell fate determinant recognition device provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0021] In the description of the embodiments of the present invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0022] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.
[0023] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0024] A specific embodiment of the present invention, such as Figure 1 As shown, a method for identifying cell fate determinants is disclosed, comprising: S101, determine the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, and construct cell lineages based on single-cell transcriptome sequencing data.
[0025] In this embodiment of the invention, single-cell transcriptome sequencing data can be downloaded from publicly uploaded transcriptome sequencing datasets by researchers or medical personnel in other laboratories via ArrayExpress or GEO public databases, or prepared using high-throughput sequencing technology. Single-cell transcriptome sequencing data refers to the collection of all messenger RNA abundance values of a cell at a specific time, usually displayed in the form of a numerical matrix, where rows represent genes and columns represent individual cells. Each value in the matrix represents the expression level (or expression count) of a certain gene in a certain cell; the higher the value, the more active the gene is in that cell. Furthermore, transcription factors and target genes are both genes in the cell. Transcription factors can regulate the expression of target genes in the cell. After obtaining single-cell transcriptome sequencing data, the expression levels of each transcription factor and target gene in the cell are statistically analyzed, and a cell lineage is constructed. The cell lineage refers to the family tree or developmental history of a cell from a fertilized egg, through multiple divisions and differentiations, to produce various specific types of cells (such as muscle cells and nerve cells). The process of constructing the cell lineage will be described in detail later in this invention.
[0026] S102, a gene regulation model was constructed based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster of the cell lineage.
[0027] In some possible embodiments of the present invention, the cell lineage contains multiple cell clusters, and for each cell cluster, the expression level of each gene in each cell is the same. In order to understand the effect of transcription factors on the expression of target genes, a gene regulation model can be constructed. This gene regulation model is used to represent the relationship between the expression level of transcription factors in cells and the expression level of target genes in cells. The specific construction process of the gene regulation model and the solution process of the model will be described in detail later in the present invention.
[0028] S103. An objective function is constructed with the aim of ensuring that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation model between adjacent cell clusters is within a preset similarity range. Based on the objective function, the regulation coefficient of transcription factors on target genes in each cell cluster is determined.
[0029] In this embodiment of the invention, in order to accurately solve the gene regulation model, the accuracy, sparsity, and similarity of the gene regulation models of each cell cluster can be controlled. Specifically, an objective function can be constructed with the aim of ensuring that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation models between adjacent cell clusters is within a preset similarity range. Then, the objective function is optimized to solve the gene regulation model of each cell cluster, thereby determining the regulation coefficient of the transcription factor on the target gene in each cell cluster.
[0030] S104 identifies differentially expressed transcription factors between adjacent cell clusters based on regulatory coefficients, and cell fate determinants are determined based on these differentially expressed transcription factors.
[0031] In this embodiment of the invention, after determining the regulatory coefficients of transcription factors on target genes in each cell cluster, the gene regulation model in the cell cluster can be calculated, thereby determining the regulatory coefficients of transcription factors controlling the expression level of target genes in each cell cluster. By comparing the differentially expressed transcription factors between adjacent cell clusters, cell fate determinants can be identified. The specific method for determining cell fate determinants will be described in detail later in this invention.
[0032] The cell fate determinant identification method provided by this invention determines the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, constructs cell lineages based on single-cell transcriptome sequencing data, and builds gene regulation models based on the relationship between the expression levels of each transcription factor and each target gene in each cell cluster within the cell lineage. This not only mines gene regulation information from single-cell transcriptome sequencing data but also integrates the topological structure of biological developmental cell lineages, thereby improving the accuracy of gene regulation network reconstruction and ensuring the continuity of gene regulation models during development. An objective function is constructed with the goals of ensuring that the accuracy of gene regulation models for each cell cluster is within a preset accuracy range, the sparsity of gene regulation models for each cell cluster is within a preset sparsity range, and the similarity of gene regulation models between adjacent cell clusters is within a preset similarity range. Based on the objective function, the regulatory coefficients of transcription factors on target genes in each cell cluster are determined. Based on the regulatory coefficients, differentially expressed transcription factors between adjacent cell clusters are identified, and cell fate determinants are determined based on the differentially expressed transcription factors. By analyzing the differences in gene regulation network structure between adjacent cell clusters in the cell lineage, transcriptional regulatory relationships and fate determinants that play a key role in cell fate determination at different developmental stages are identified, aiming to elucidate the cell fate determination mechanism in the developmental process. By controlling the sparsity of gene regulation models for each cell cluster and the similarity of gene regulation models between adjacent cell clusters, the accuracy of gene regulation models for each cell cluster and the accuracy of the determination of fate determinants can be ensured.
[0033] In some possible embodiments of the present invention, such as Figure 2 As shown, the expression levels of transcription factors and target genes in cells were determined from single-cell transcriptome sequencing data, including: S201: Obtain initial single-cell transcriptome sequencing data; filter genes based on the number of cells involved in gene expression in the initial single-cell transcriptome sequencing data to obtain transcription factors and target genes. S202 normalizes the expression levels of transcription factors and target genes in different cells to obtain a gene expression matrix that represents the expression levels of transcription factors and target genes in cells.
[0034] In this embodiment of the invention, the initial single-cell transcriptome sequencing data is typically obtained in the form of a raw counting matrix, where rows represent genes and columns represent individual cells. Each value in the matrix represents the unnormalized raw read counts or unique molecular identifier (UMI) count of a gene in a given cell. For each gene in the raw data, the number of cells in which its expression level is greater than zero is calculated. In other words, the expression level of the gene in each cell is counted to see if it exceeds a preset threshold (usually 0). Cells with expression levels greater than this threshold are considered to be cells "involved in gene expression." A minimum cell number threshold is set. This threshold can be determined based on a percentage of the total number of cells (e.g., requiring a gene to be expressed in at least 0.1% or 0.5% of the total number of cells) or an absolute value (e.g., requiring a gene to be expressed in at least 10 or 50 cells). Genes with a "number of cells involved in expression" greater than or equal to the threshold are retained, while genes below the threshold are removed. This step effectively removes low-quality and high-noise genes, significantly improving the signal-to-noise ratio. After filtering, the expression levels of the remaining genes in different cells still need to be normalized. The purpose of normalization is to eliminate the differences in total sequencing depth between cells due to technical reasons, so that the gene expression levels between different cells are comparable.
[0035] In a specific implementation example, single-cell transcriptome sequencing data of 15,000 cells were obtained. The original data contained approximately 25,000 genes. The filtering threshold was set to "genes expressed in at least 30 cells". After filtering, 12,000 genes remained, including 1,500 transcription factors and 10,500 target genes. Using a logarithmic normalization method, the gene count of each cell was first divided by the total count of that cell and multiplied by 10,000. Then, the result was logarithmically transformed to obtain a normalized gene expression matrix of size [15,000 cells × 12,000 genes]. The rows of the gene expression matrix correspond to genes, and the columns correspond to cells.
[0036] In this embodiment of the invention, a normalized gene expression matrix is obtained by preprocessing single-cell transcriptome sequencing data, which facilitates subsequent data processing.
[0037] In some possible embodiments of the present invention, such as Figure 3 As shown, cell lineages were constructed based on single-cell transcriptome sequencing data, including: S301, acquire cell cluster information from single-cell transcriptome sequencing data; S302, based on the cell differentiation process, constructs a directed tree relationship between cell clusters to obtain cell lineages.
[0038] In this embodiment of the invention, if the metadata provided by the transcriptome sequencing dataset contains information about the cell cluster to which each single cell belongs, then a cell lineage is established based on the cell differentiation process in the organism; if the information about the cell cluster to which a single cell belongs is not provided by the metadata, any one of the bioinformatics methods such as Monocle2, Slingshot, DPT, and Palantir can be used to construct the cell lineage. Figure 4 and 5 The diagram shows a schematic of embryonic development and a schematic of cell lineage. Furthermore, the cell lineage can be represented as a directed tree, i.e. During development, all K cell clusters form a vertex set. The directed edges between these K cell clusters form an edge set. The directed edges in this tree represent the differentiation relationships between cell clusters. During development, if cell clusters... Cell clusters differentiated So, in cell lineages, from arrive There exists a directed edge connecting them, which is an ordered pair. .
[0039] The embodiments of the present invention improve the accuracy of gene regulation model construction by constructing cell lineages and integrating the topological structure of biological developmental cell lineages.
[0040] In some possible embodiments of the present invention, a gene regulation model is constructed based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster within a cell lineage, including: For each cell cluster, a linear gene regulation model is constructed with the expression level of the target gene in the cell as the dependent variable, the expression level of the transcription factor in the cell as the independent variable, the regulatory coefficient of the transcription factor on the target gene as the slope, and Gaussian random noise as the error term.
[0041] In this embodiment of the invention, the k-th cell cluster in the cell lineage tree The relationship between the expression levels of transcription factors and target genes in cells is as follows:
[0042] in, This indicates that the target gene is located in the cell cluster. of A vector composed of expression levels within each cell Represents cell clusters The number of cells in Indicating transcription factors in cell clusters All A vector composed of expression levels within each cell This vector represents the regulatory coefficients of transcription factors on target genes. A positive coefficient indicates enhanced expression regulation of the target gene, a negative coefficient indicates repressive expression regulation, and a zero coefficient indicates no regulatory effect on the target gene. This represents Gaussian random noise.
[0043] The embodiments of the present invention construct a gene regulation model, which can solve the regulation coefficient of transcription factors on target genes, and thus determine the differential transcription factors between different cell clusters.
[0044] In some possible embodiments of the present invention, the objective function is:
[0045] in, Let A represent the regulatory coefficient of transcription factors on target genes, B represent the precision range of gene regulation models for each cell cluster, C represent the sparsity range of gene regulation models for each cell cluster, and D represent the similarity range of gene regulation models between adjacent cell clusters. and These are the preset hyperparameters.
[0046] In this embodiment of the invention, the expression levels of the target gene and transcription factor are observed in the established mathematical model, and the regulatory coefficients are... This is still to be speculated. The objective function for optimizing the model, which integrates the accuracy, sparsity, and similarity of gene regulation models across different cell clusters in the cell lineage, and incorporates the topological structure of the cell lineage, is as shown above.
[0047] Furthermore, in the objective function, the topological structure of cell lineages is incorporated into the design of the objective function through the similarity module C. If cell lineages... Differentiate into cell clusters Then there exists a directed edge connecting the cell lineages. and That is, ordered pairs ,at this time equivalent , Representing vectors Norm. Through the similarity module C, the similarity relationship of gene regulatory network structures between adjacent cell clusters in the cell lineage is constrained, ensuring that as the organism develops, the gene regulatory network structure among adjacent cell clusters in the cell lineage tree T exhibits gradual changes, rather than sudden and drastic changes.
[0048] This invention improves the accuracy of gene regulation model construction by setting an objective function.
[0049] In some possible embodiments of the present invention, such as Figure 6 As shown, the regulatory coefficients of transcription factors on target genes in each cell cluster are determined based on the objective function, including: S601 uses a pre-defined convex optimization algorithm to collaboratively optimize the objective function and determine the regulatory coefficients of transcription factors on the target gene in each cell cluster. S602, a linear gene regulation model for each cell cluster is determined based on the regulatory coefficients of transcription factors on target genes.
[0050] In this embodiment of the invention, because the collaborative optimization design of gene regulation models on different cell clusters in a cell lineage involves the selection of hyperparameter values, it is related to the selection of hyperparameter values. , Under the given conditions, the collaborative optimization of three different modules A, B, and C is performed. The objective function can be collaboratively optimized using any one of four convex optimization algorithms: SLEP, CVX, genlasso, and fGFL. Furthermore, for hyperparameters... and ,when At that time, it will lead to That is, the number of edges in the gene regulatory network on all inferred cell clusters is 0; when This will lead to the following consequences: for any They all This is valid if the structure of the gene regulatory network is identical across different cell clusters in the inferred cell lineage. Therefore, it is necessary to analyze the hyperparameters. and Optimize the values of hyperparameters in the objective function. and The method can be adjusted and determined by combining grid search and cross-validation.
[0051] In this embodiment of the invention, the hyperparameters in the objective function are evaluated using a combination of grid search and cross-validation. and Adjustments are made, and based on this, the objective function is co-optimized using a convex optimization algorithm to determine the regulatory coefficients in the gene regulation model of each cell cluster, thus completing the construction of the gene regulation model and determining the influence of each transcription factor in each cell cluster on the expression level of the target gene.
[0052] In some possible embodiments of the present invention, such as Figure 7 As shown, differentially expressed transcription factors between adjacent cell clusters are identified based on regulatory coefficients, and cell fate determinants are identified based on these differentially expressed transcription factors, including: S701, based on a linear gene regulation model of each adjacent cell cluster, identifies differentially expressed transcription factors that show differences in the regulation of target genes in each adjacent cell cluster. S702, based on the presence of differentially expressed transcription factors in adjacent cell clusters, determines the cell fate determinant type of differentially expressed transcription factors.
[0053] In this embodiment of the invention, the identification of key fate regulators at different developmental stages in cell lineages, if there are directed edges... So, in cell lineages, from cell clusters Differentiated into During this stage, the set of transcription factors that exhibit differences in regulating the target gene is as follows:
[0054] Among them, if
[0055] So, the first j The regulation of target genes by transcription factors in upstream cell clusters It is not present in the middle, but in the downstream cell clusters. This regulatory relationship was identified as affecting cell clusters. To cell clusters Differentiation plays a key role in determining fate.
[0056] if
[0057] So, the first j The regulation of target genes by transcription factors in upstream cell clusters It exists in the middle, but in the downstream cell clusters This does not exist. This regulatory relationship was identified as affecting cell clusters. Not towards cell clusters Differentiation plays a key role in determining fate.
[0058] This invention identifies the fate determinants that play a key role in the differentiation / undifferentiation of cells from upstream cell clusters to downstream cell clusters by observing the presence or absence of transcription factors between adjacent cell clusters.
[0059] To better implement the cell fate determinant identification method in the embodiments of the present invention, based on the cell fate determinant identification method, correspondingly, as follows: Figure 8 As shown, this embodiment of the invention also provides a cell fate determinant identification device, the cell fate determinant identification device 800 comprising: The cell lineage determination module 801 is used to determine the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, and to construct cell lineages based on single-cell transcriptome sequencing data. Model building module 802 is used to build a gene regulation model based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster in the cell lineage. The regulation coefficient determination module 803 is used to construct an objective function with the aim that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation model between each adjacent cell cluster is within a preset similarity range. Based on the objective function, the regulation coefficient of the transcription factor on the target gene in each cell cluster is determined. The fate determinant determination module 804 is used to determine differentially expressed transcription factors between adjacent cell clusters based on regulatory coefficients, and to determine cell fate determinants based on differentially expressed transcription factors.
[0060] The cell fate determinant identification device 800 provided in the above embodiments can realize the technical solutions described in the above cell fate determinant identification method embodiments. The specific implementation principles of each module or unit can be found in the corresponding content in the above cell fate determinant identification method embodiments, and will not be repeated here.
[0061] like Figure 9 As shown, the present invention also provides an electronic device 900. The electronic device 900 includes a processor 901, a memory 902, and a display 903. Figure 9 Only some components of the electronic device 900 are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0062] In some embodiments, processor 901 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in memory 902 or process data, such as the cell fate determinant identification method of the present invention.
[0063] In some embodiments, processor 901 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 901 may be local or remote. In some embodiments, processor 901 may be implemented on a cloud platform. In some embodiments, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, internal cloud, multi-cloud, etc., or any combination thereof.
[0064] In some embodiments, memory 902 may be an internal storage unit of electronic device 900, such as a hard disk or memory of electronic device 900. In other embodiments, memory 902 may also be an external storage device of electronic device 900, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 900.
[0065] Furthermore, the memory 902 may include both internal storage units of the electronic device 900 and external storage devices. The memory 902 is used to store application software and various types of data installed on the electronic device 900.
[0066] In some embodiments, display 903 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 903 is used to display information from electronic device 900 and to display a visual user interface. Components 901-903 of electronic device 900 communicate with each other via a system bus.
[0067] In some embodiments, when processor 901 executes the cell fate determinant identification program in memory 902, the following steps may be performed: Determine the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, and construct cell lineages based on single-cell transcriptome sequencing data; A gene regulation model was constructed based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster within a cell lineage. An objective function is constructed with the aim that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation model between each adjacent cell cluster is within a preset similarity range. Based on the objective function, the regulation coefficient of transcription factors on target genes in each cell cluster is determined. Differential transcription factors between adjacent cell clusters are identified based on regulatory coefficients, and cell fate determinants are determined based on these differential transcription factors.
[0068] It should be understood that when the processor 901 executes the cell fate determinant identification program in the memory 902, in addition to the functions mentioned above, it can also perform other functions, as can be found in the description of the corresponding method embodiments above.
[0069] Furthermore, this embodiment of the invention does not specifically limit the type of electronic device 900 mentioned. Electronic device 900 can be a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, laptop computer, or other portable electronic device. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The aforementioned portable electronic device can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the invention, electronic device 900 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0070] Accordingly, this application also provides a computer-readable storage medium for storing a computer-readable program or instruction. When the program or instruction is executed by a processor, it can implement the steps or functions of the cell fate determinant identification method provided in the above-described method embodiments.
[0071] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0072] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying cell fate determinants, characterized in that, include: The expression levels of transcription factors and target genes in cells were determined from single-cell transcriptome sequencing data, and cell lineages were constructed based on the single-cell transcriptome sequencing data. A gene regulation model was constructed based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster of the cell lineage. An objective function is constructed with the aim that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation model between each adjacent cell cluster is within a preset similarity range. Based on the objective function, the regulation coefficient of the transcription factor on the target gene in each cell cluster is determined. Based on the regulatory coefficients, differentially expressed transcription factors between adjacent cell clusters are determined, and cell fate determinants are determined based on the differentially expressed transcription factors.
2. The method for identifying cell fate determinants according to claim 1, characterized in that, The determination of the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data includes: Initial single-cell transcriptome sequencing data is obtained, and the genes are filtered according to the number of cells in which each gene is expressed in the initial single-cell transcriptome sequencing data to obtain transcription factors and target genes; The expression levels of the transcription factors and the target gene in different cells are normalized to obtain a gene expression matrix representing the expression levels of the transcription factors and the target gene in the cells.
3. The method for identifying cell fate determinants according to claim 2, characterized in that, The construction of cell lineages based on the single-cell transcriptome sequencing data includes: Obtain cell cluster information from the single-cell transcriptome sequencing data; A directed tree relationship between the cell clusters is constructed based on the cell differentiation process to obtain the cell lineage.
4. The method for identifying cell fate determinants according to claim 1, characterized in that, The gene regulation model constructed based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster of the cell lineage includes: For each cell cluster, a linear gene regulation model is constructed with the expression level of the target gene in the cell as the dependent variable, the expression level of the transcription factor in the cell as the independent variable, the regulatory coefficient of the transcription factor on the target gene as the slope, and Gaussian random noise as the error term.
5. The method for identifying cell fate determinants according to claim 4, characterized in that, The objective function is: in, Let A represent the regulatory coefficient of transcription factors on the target gene, B represent the precision range of the gene regulation model for each cell cluster, C represent the sparsity range of the gene regulation model for each cell cluster, and D represent the similarity range of gene regulation models between adjacent cell clusters. and These are the preset hyperparameters.
6. The method for identifying cell fate determinants according to claim 5, characterized in that, The determination of the regulatory coefficient of the transcription factor on the target gene in each of the cell clusters based on the objective function includes: The objective function is co-optimized using a pre-defined convex optimization algorithm to determine the regulatory coefficients of transcription factors on the target gene in each cell cluster. Based on the regulatory coefficients of the transcription factors on the target genes, a linear gene regulation model for each cell cluster was determined.
7. The method for identifying cell fate determinants according to claim 6, characterized in that, The determination of differentially expressed transcription factors between adjacent cell clusters based on the regulatory coefficient, and the determination of cell fate determinants based on the differentially expressed transcription factors, include: Based on a linear gene regulation model of each adjacent cell cluster, differentially expressed transcription factors that show differences in the regulation of the target gene in each adjacent cell cluster were identified. Based on the presence or absence of the differentially expressed transcription factor in adjacent cell clusters, the cell fate determinant type of the differentially expressed transcription factor is determined.
8. A cell fate determinant recognition device, characterized in that, include: The cell lineage determination module is used to determine the expression levels of transcription factors and target genes in cells from single-cell transcriptome sequencing data, and to construct cell lineages based on the single-cell transcriptome sequencing data. The model building module is used to build a gene regulation model based on the relationship between the expression levels of each transcription factor and the expression levels of each target gene in each cell cluster in the cell lineage. The regulation coefficient determination module is used to construct an objective function with the aim that the accuracy of the gene regulation model of each cell cluster is within a preset accuracy range, the sparsity of the gene regulation model of each cell cluster is within a preset sparsity range, and the similarity of the gene regulation model between each adjacent cell cluster is within a preset similarity range, and to determine the regulation coefficient of the transcription factor on the target gene in each cell cluster based on the objective function. A fate determinant determination module is used to determine differentially expressed transcription factors between adjacent cell clusters based on the regulatory coefficients, and to determine cell fate determinants based on the differentially expressed transcription factors.
9. An electronic device, characterized in that, Including memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the cell fate determinant identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store computer-readable programs or instructions, which, when executed by a processor, enable the implementation of the steps in the cell fate determinant identification method according to any one of claims 1 to 7.