A transcription factor regulatory network modeling method based on single-cell transcriptome data

By combining multiple nonnegative matrix factorization techniques with prior molecular interaction networks, the problem of integrating multidimensional molecular synergistic regulatory information in single-cell transcriptome data was solved, enabling effective identification and interpretation of phenotypic changes and improving the data's noise resistance and biological robustness.

CN116030878BActive Publication Date: 2026-02-13ZHONGSHAN OPHTHALMIC CENT SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111238534.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2026-02-13
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate multidimensional molecular synergistic regulatory information from single-cell transcriptome data, failing to intuitively reflect the synergistic regulatory mechanisms of transcription factors and genes related to phenotypic changes. Furthermore, single-cell omics data is characterized by data sparsity and high noise.

Method used

By employing multiple nonnegative matrix factorization techniques and combining them with prior intermolecular interaction networks, we can extract transcription factor-gene functional synergistic interaction information related to phenotypic changes from single-cell data through dimensionality reduction, and construct a multidimensional molecular synergistic regulatory mechanism.

Benefits of technology

It improves the noise resistance and biological robustness of single-cell omics data, enables better identification of multidimensional molecular synergistic regulatory mechanisms, and provides tools and means to explain the underlying biological molecular mechanisms behind phenotypic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030878B_ABST
    Figure CN116030878B_ABST
Patent Text Reader

Abstract

The application discloses a transcription factor regulation network modeling method based on single cell transcriptome data, comprising the following steps: S1, extracting multi-phenotype single cell omics data, performing data cleaning, and integrating the cleaned data; S2, based on a biological knowledge base, analyzing the data processed in S1, and constructing a prior inter-molecular interaction relationship network; S3, based on a multi-factor non-negative matrix factorization algorithm, establishing a multi-dimensional molecular synergistic interaction relationship module according to the data processed in S1 and the prior inter-molecular interaction relationship network in S2; S4, calculating the interaction relationship module related to the phenotype of the multi-dimensional molecular synergistic interaction relationship module; and S5, visualizing and exporting the interaction relationship between the multi-dimensional molecular synergistic interaction relationship module and the phenotype. With the aid of single cell sequencing technology and prior biological knowledge, the application extracts multi-dimensional molecular synergistic regulation relationship in high-dimensional single cell omics data, and obtains a transcription factor-gene function synergistic regulation interaction mechanism closely related to phenotype changes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of biotechnology industry, and relates to a method for analyzing multi-dimensional molecular synergistic regulation mechanism by using single cell omics data, in particular, a new method for identifying synergistic regulation interaction mechanism of transcription factor-gene function related to phenotype change based on multi-item non-negative matrix factorization algorithm framework by using multi-phenotype single cell omics data. BACKGROUND

[0002] In recent years, with the development of microfluidic chip technology, omics sequencing technology at the single cell level has been widely applied in more and more basic laboratories and clinical researches. Among them, single cell transcriptome sequencing technology (scRNA-seq) is most concerned by biomedical researchers because it can reflect the instantaneous physiological state of a single cell through the transcriptional expression profile. With the continuous optimization of sequencing depth and accuracy of scRNA-seq technology and the increasing number of cells that can be measured at one time, scRNA-seq is not only used for identification of different cell types, cell subpopulations, differentiation of different states of cells, but also used for tracking of cell lineage and capturing of key characteristics of cells in the transformation process of development and differentiation. Although the public data and tools of scRNA-seq are growing rapidly, only sporadic studies have integrated and extracted multi-dimensional molecular synergistic regulation information from scRNA-seq data, and there is no method and tool based on single cell omics data to extract phenotype-related transcription factor-gene function synergistic regulation modules.

[0003] Transcription factors are a class of DNA-binding proteins that regulate cell fate, developmental patterns and specific biological functions in living organisms. They can regulate the transcription or activation of downstream target genes by specifically recognizing eukaryotic cis-acting elements, and then promote the changes of biological phenotypes or physiological states. Many transcription factors, such as P53 and CREB, are closely related to the occurrence and development of important diseases such as tumors and inflammation. Understanding the changes of transcription factor-mediated signal pathway and target gene expression under specific phenotype changes is an important part of analyzing the molecular mechanism behind the changes. Existing multi-phenotype single cell transcriptome analysis often only performs differential gene analysis on the transcriptome of the same cell population under different phenotypes, and directly carries out downstream functional analysis based on the results. However, single cell omics data usually has the characteristics of data sparsity and high noise. The phenotype-related molecular information obtained only by comparative genomics cannot directly reflect the changes of regulatory activity related to phenotype change, nor can it fully reflect the changes of transcription factor-mediated signal pathway or gene function related to phenotype change.

[0004] To study the synergistic regulation between phenotype-related transcription factor-gene functions, the present application establishes a new method for identifying multi-dimensional molecular synergistic regulation mechanism based on multi-phenotype single cell data. The method introduces a multi-term non-negative matrix factorization technology. The technology is a derivative of the non-negative matrix factorization technology (NMF). It relies on the prior molecular interaction relationship network to utilize existing biological knowledge, extracts the transcription factor-gene function synergistic interaction information related to phenotype changes from single cell data by dimensionality reduction of high-dimensional features of single cell data. Compared with the common non-negative matrix factorization technology, the multi-term non-negative matrix factorization technology relies on the prior molecular interaction relationship, so the multi-dimensional molecular synergistic regulation mechanism recognition method based on the technology has stronger anti-interference to single cell data noise, and the recognized multi-dimensional molecular synergistic regulation mechanism has better biological robustness. SUMMARY

[0005] The purpose of the present application is to extract the multi-dimensional molecular synergistic regulation relationship in high-dimensional single cell omics data by means of single cell sequencing technology and existing prior biological knowledge, to obtain the transcription factor-gene function synergistic regulation interaction mechanism closely related to phenotype changes, and to present the multi-dimensional molecular interaction relationship related to phenotype changes in an easy-to-read and tightly integrated form.

[0006] The present application discloses a transcription factor regulation network modeling method based on single cell transcriptome data, comprising the following steps:

[0007] S1. Extracting multi-phenotype single cell omics data, performing data cleaning, and integrating the cleaned data;

[0008] S2. Based on the biological knowledge base, analyzing the data processed in S1 to construct a prior molecular interaction relationship network;

[0009] S3. Based on the multi-factor non-negative matrix factorization algorithm, establishing a multi-dimensional molecular synergistic interaction relationship module according to the data processed in S1 and the prior molecular interaction relationship network in S2;

[0010] S4. Calculating the multi-dimensional molecular synergistic interaction relationship module and the phenotype-related interaction relationship module;

[0011] S5. Visualizing and exporting the multi-dimensional molecular synergistic interaction relationship module and the phenotype-related interaction relationship module.

[0012] Further, the data cleaning in S1 comprises:

[0013] S101 sets the filtering conditions; the filtering conditions include at least one of the following: cells with low abundance in multiphenotype single-cell omics data, cells contaminated with cell debris, apoptotic or lysed cells, and multimers;

[0014] S102 filters multiphenotype single-cell omics data according to filtering conditions to obtain filtered data;

[0015] S103 performs feature recognition on the filtered data and integrates the recognized data; the feature recognition includes at least cell clustering and cell characteristic gene recognition.

[0016] Furthermore, the intermolecular interaction network in S2 includes the regulatory relationship network between transcription factors and target genes, and the functional association network between genes.

[0017] Furthermore, in the construction of the functional association network between genes, the associated genes participate in at least one of the following: regulating the same biological process, participating in the same gene pathway, or responding to the same phenotype.

[0018] The functional association network between genes includes, but is not limited to, one or more of the following: gene association networks with co-epigenetic modifications, receptor-ligand interaction networks of gene-encoded proteins, and protein-protein interaction networks of gene-encoded proteins.

[0019] Furthermore, the multi-factor nonnegative matrix factorization algorithm is as follows:

[0020] S301 sets the total number of observed cells as n, the total number of observed genes as m, and the total number of observed transcription factors as s, and establishes a [system / mechanism]. A nonnegative matrix of dimension, set as multi-phenotype single-cell gene expression profile data. ; establish A non-negative matrix of dimension, set as single-cell regulator activity matrix data. ;

[0021] S302 sets the number of all transcription factor-gene functional co-interaction modules observed in n cells to be k, and establishes a... A nonnegative matrix W of dimension 1; establish A nonnegative matrix of dimension 1, used to describe the weighted relationship between variables and genes in a low-dimensional space, is denoted as . ; establish A nonnegative matrix of dimension 1, used to describe the weight relationship between variables and transcription factors in low-dimensional space, is denoted as . ;

[0022] W satisfies , And the error in the square of the factorization is:

[0023]

[0024] wherein is the Frobenius norm of matrix , based on squared error, the objective function is constructed as:

[0025] .

[0026] Further, the S302 is The formula can be further written as:

[0027]

[0028] wherein is the trace coefficient, and is the trace matrix of matrix, A is the adjacency matrix of prior gene-gene functional relationship, B is the adjacency matrix of prior transcription factor-target gene functional relationship, is the i-th column vector in the matrix, is the j-th column vector in the matrix;

[0029] The maximum objective function of the prior relationship adjacency matrix is:

[0030]

[0031] .

[0032] Further, the S4 is calculating the interaction relationship between the multi-dimensional molecular synergistic interaction relationship module and the phenotype, specifically: using a difference significance statistical test method, detecting the correlation between the variables reflecting the multi-dimensional molecular synergistic interaction relationship in the low-dimensional space obtained after multi-factor non-negative matrix factorization and the phenotype changes, and obtaining the interaction relationship between the multi-dimensional molecular synergistic interaction relationship module and the phenotype.

[0033] Further, the difference significance statistical test method includes but is not limited to one of Student's t-test, Mann-Whitney U test, and analysis of variance.

[0034] Further, the visualization derivation in S5 is specifically: using the interaction relationship between the multi-dimensional molecular synergistic interaction relationship module and the phenotype obtained in S4, according to the corresponding transcription factor coefficient matrix and gene coefficient matrix, and cooperating with the prior inter-molecular interaction relationship network in S2, a visual transcription factor-gene functional synergistic relationship network is generated.

[0035] Further, the visual transcription factor-gene function synergistic relationship network comprises a plurality of synergistic relationship network nodes composed of a plurality of transcription factors, a plurality of genes and a plurality of biological functions, and edges of the visual transcription factor-gene function synergistic relationship network are composed of transcription factor-gene regulation relationships, gene-gene interaction relationships and gene-gene function relationships.

[0036] Compared with the prior art, the present application extracts multi-dimensional molecular interaction information of single-cell omics data based on prior biological knowledge, and the multi-dimensional molecular interaction information in the low-dimensional space obtained has strong anti-interference and good biological robustness to single-cell omics data noise; the application has strong scalability and good portability, and can flexibly adjust the prior interaction relationship network between molecules according to different research backgrounds, extract the multi-dimensional molecular regulation mechanism related to the research purpose in the single-cell data, and provide means and tools for explaining the potential biological molecular action mechanism behind the phenotype change and phenotype intervention. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 The present application is an embodiment flowchart.

[0038] Figure 2 The present application is an embodiment flowchart.

[0039] The rhombus node in the figure represents a transcription factor, the circular node represents a gene pathway, and the triangular node represents a gene. The connection between nodes represents the functional interaction between two molecules. DETAILED DESCRIPTION

[0040] In order to enable personnel in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present application, not all.

[0041] The present application discloses a transcription factor regulation network modeling method based on single-cell transcriptome data, as shown in Figure 1 The steps are as follows:

[0042] S1. Extract multi-phenotype single-cell omics data, perform data cleaning, and integrate the cleaned data;

[0043] Since the multi-phenotype single-cell omics data set used contains multiple phenotypes, single-cell data preprocessing is required, part of the useless cell data is removed after data cleaning, and data integration is performed to obtain the required data for downstream analysis;

[0044] S2. Based on the biological knowledge base, the data processed by S1 is analyzed to construct a prior intermolecular interaction relationship network;

[0045] Among them, as shown in the formula (I), according to the integrated data processed by S1, the gene-gene relationship network based on the annotated path and the single-cell level transcription factor-effective target gene relationship network are constructed through the single-cell level transcription factor activity spectrum, so as to generate the prior intermolecular interaction relationship network; Figure 1

[0046] S3. Based on the multi-factor non-negative matrix decomposition algorithm, the multi-dimensional molecular synergistic interaction relationship module is established according to the data processed by S1 and the prior intermolecular interaction relationship network of S2;

[0047] Among them, as shown in the formula (II), the standardized and high-quality annotated cell-gene expression matrix is constructed for the data processed by S1, and the multi-dimensional molecular synergistic interaction relationship module is identified in the important low-dimensional space based on the multi-factor non-negative matrix decomposition algorithm in cooperation with the prior intermolecular interaction relationship network of S2, and the multi-dimensional molecular synergistic interaction relationship module in this embodiment is specifically a transcription factor-gene function synergistic interaction relationship; Figure 1

[0048] S4. The interaction relationship module related to the phenotype of the multi-dimensional molecular synergistic interaction relationship module is calculated;

[0049] Among them, according to the multi-dimensional molecular synergistic interaction feature in the low-dimensional space identified by S3, the correlation between it and the target phenotype change is evaluated, the synergistic interaction feature closely related to the phenotype change is screened, and the transcription factor-gene-function synergistic interaction relationship module related to the phenotype is established according to the synergistic interaction feature;

[0050] S5. The multi-dimensional molecular synergistic interaction relationship module related to the phenotype is visualized and exported;

[0051] Among them, for the transcription factor-gene-function synergistic interaction relationship module related to the phenotype identified by S4, a multi-dimensional molecular feature network is generated to show the prior relationship between the transcription factor, the gene and the function in the multi-dimensional molecular synergistic regulation mechanism and the potential association driven by the data.

[0052] The embodiment of the application extracts the multi-dimensional molecular interaction information on the low-dimensional space based on the prior biological knowledge of the single-cell omics data, and the multi-dimensional molecular interaction information on the low-dimensional space has strong anti-interference and good biological robustness to the noise of the single-cell omics data; the application has strong scalability and good portability, and can flexibly adjust the prior interaction relationship network between molecules according to the different research backgrounds, extract the multi-dimensional molecular regulation mechanism related to the research purpose in the single-cell data, and provide means and tools for explaining the potential biological molecular action mechanism behind the phenotype change and the phenotype intervention. ​​

[0053] Optionally, the data cleaning in S1 includes:

[0054] S101 sets the filtering conditions; the filtering conditions include at least one of the following: cells with low abundance in multiphenotype single-cell omics data, cells contaminated with cell debris, apoptotic or lysed cells, and multimers;

[0055] S102 filters multiphenotype single-cell omics data according to filtering conditions to obtain filtered data;

[0056] S103 performs feature recognition on the filtered data and integrates the recognized data; the feature recognition includes cell clustering and cell characteristic gene recognition.

[0057] Optionally, the intermolecular interaction network in S2 includes the regulatory relationship network between transcription factors and target genes, and the functional association network between genes.

[0058] In particular, by using the functional association network between genes, genes that are a priori related to the target gene can be identified, thereby determining the association between transcription factors and multiple genes.

[0059] Optionally, in the construction of the functional association network between genes, the associated genes participate in at least one of the following: regulating the same biological process, participating in the same gene pathway, or responding to the same phenotype.

[0060] The functional association network between genes includes, but is not limited to, one or more of the following: gene association networks with co-epigenetic modifications, receptor-ligand interaction networks of gene-encoded proteins, and protein-protein interaction networks of gene-encoded proteins.

[0061] Optionally, the multi-factor nonnegative matrix factorization algorithm is as follows:

[0062] S301 sets the total number of observed cells as n, the total number of observed genes as m, and the total number of observed transcription factors as s, and establishes a [system / mechanism]. A nonnegative matrix of dimension, set as multi-phenotype single-cell gene expression profile data. ; establish A non-negative matrix of dimension, set as single-cell regulator activity matrix data. ;

[0063] S302 sets the number of all transcription factor-gene functional co-interaction modules observed in n cells to be k, and establishes a... A nonnegative matrix W of dimension 1; establish A nonnegative matrix of dimension 1, used to describe the weighted relationship between variables and genes in a low-dimensional space, is denoted as . ; establish The non-negative matrix factorization is used for describing the weight relationship between variables and transcription factors in a low-dimensional space, and is set as ;

[0064] W satisfies , ; and the square error of decomposition is:

[0065]

[0066] wherein is the Frobenius norm of the matrix , and the objective function is constructed based on the square error:

[0067] .

[0068] Optionally, the formula in the S302 can be further written as:

[0069]

[0070] wherein is a trace coefficient, and are trace matrices of the matrix, A is an adjacency matrix of prior gene-gene functional relationship, B is an adjacency matrix of prior transcription factor-target gene functional relationship, is the i-th column vector in the matrix , is the j-th column vector in the matrix ;

[0071] The maximum objective function of the prior relationship adjacency matrix is:

[0072]

[0073] .

[0074] In the embodiment S3, the multi-factor non-negative matrix factorization is carried out based on the multi-phenotype single-cell data and the prior molecular interaction relationship constructed in the S2, so as to identify a low-dimensional variable space reflecting the multi-dimensional molecular synergistic interaction relationship characteristics.

[0075] Based on the formula of the embodiment S3 of the application, the multi-factor non-negative matrix factorization formula can be constructed through the following sub-steps:

[0076] a) Randomly initializing , and matrices, so as to ensure that the matrices are non-zero matrices. The total iteration number of the non-negative matrix factorization is set as T, the minimum convergence residual is E, and the initial iteration number is t=0; ​

[0077] b) fixing and matrix, the objective function is:

[0078]

[0079] wherein the updating formula of the matrix is:

[0080]

[0081] find and , update matrix;

[0082] c) fixing matrix, the objective function is:

[0083]

[0084] wherein and the updating formula of the matrix is:

[0085]

[0086]

[0087] find and and matrix, update and matrix;

[0088] d) repeat steps b) to c) until the convergence condition is met (the number of running iterations reaches T times or the objective function value is less than E), to obtain and matrix, thereby determining the transcription factor-gene function synergistic interaction relationship.

[0089] Optionally, the module for calculating the multi-dimensional molecular synergistic interaction relationship in S4 and the module for calculating the interaction relationship related to the phenotype are specifically: using a difference significance statistical test method to detect the correlation between the variables reflecting the multi-dimensional molecular synergistic interaction relationship in the low-dimensional space obtained by the non-negative matrix factorization of the multi-factor and the phenotype changes, to obtain the module for calculating the multi-dimensional molecular synergistic interaction relationship and the module for calculating the interaction relationship related to the phenotype.

[0090] In particular, the difference significance statistical test method includes but is not limited to one of Student's t-test, Mann-Whitney U test, and analysis of variance.

[0091] Optionally, the visualization derived in S5 is specifically: using the multi-dimensional molecular synergistic interaction relationship module obtained in S4 and the interaction relationship module related to the phenotype, according to the corresponding transcription factor coefficient matrix and gene coefficient matrix, and cooperating with the prior inter-molecular interaction relationship network in S2, a visual transcription factor-gene function synergistic relationship network is generated.

[0092] In particular, the visual transcription factor-gene function synergistic relationship network includes a plurality of synergistic relationship network nodes, the synergistic relationship network nodes are composed of a plurality of transcription factors, a plurality of genes, and a plurality of biological functions, and the edges of the visual transcription factor-gene function synergistic relationship network are composed of transcription factor-gene regulation relationship, gene-gene interaction relationship, and gene-gene function relationship.

[0093] The embodiment of the present application is further described to illustrate the technical solution of the present application, taking a mouse experiment as an example, to identify the activity change of the regulatory subunit of the macrophage and monocyte from multiple tissues in the aging process of the mouse and the change of the signal pathway and target gene mediated by the transcription factor.

[0094] The aging of the body is rooted in the aging of the immune system. The dysfunction of macrophages and monocytes is the key reason for the aging of the immune system of the organism. Although many studies have reported that the proportion of macrophages and monocytes cell groups will be imbalanced in the aging process of the organism, and more pro-inflammatory factors will be gradually produced, thereby destroying the balance of pro-inflammatory and anti-inflammatory homeostasis in the body, but the specific regulatory mechanism such as the change of the transcription factor and the signal pathway mediated by the transcription factor in this process is not clear.

[0095] The embodiment first collects the single-cell transcriptome data of mouse liver, spleen, and kidney related to aging published and publicly available online by Jacob C. Kimmel et al. (GSE132901) and the single-cell transcriptome data of mouse liver, spleen, and kidney related to aging publicly published by Tabular Muris (Tabular Muris Senis, TMS).

[0096] According to the age of the sample mouse, the sample phenotype is divided into two categories of young mice (1-15 months old) and old mice (more than 15 months old), so as to study the transcription factor-gene function synergistic interaction mechanism related to aging of macrophages and monocytes in different tissues.

[0097] S1 first screens all macrophages (cells detected to express Fcgr1a / Mafb / Mertk / Csf1r / Adgre1 ) and monocytes (cells detected to express Cd14 / Lyz / Fcgr3a / Ms4a7 The raw data obtained by screening is re-cleaned and integrated, and the specific process is as follows:

[0098] Data cleaning: In the data cleaning process, a single sample needs to meet 1) total cell read number > 500; 2) total number of expressed genes detected by the cell > 250; 3) total number of macrophages and monocytes remaining after quality control after single sample screening > 10;

[0099] Data integration: After data cleaning, single sample uses open source R software package Seurat and harmony for data integration; the pre-dimension of dimension reduction used in integration is 50, the minimum number of adjacent cells of anchor point is set to 10, the reference number of adjacent cells of anchor point weight is set to 30, and the harmony algorithm runs using default parameters; the data after integration still uses the default process of Seurat for data standardization, dimension reduction, cell subpopulation classification, subpopulation characteristic gene identification and age-related differentially expressed gene identification.

[0100] In the step of measuring single cell regulatory activity, the open source pySCENIC software is called. Based on single cell transcriptome data, the software can predict which transcription factors are expressed in a single cell, the degree of regulatory activity of transcription factors, and the effective target gene set of the regulatory activity of transcription factors regulated by transcription factors. The mouse transcription factor annotation (v9) and transcription factor potential target gene information (mm10, TSS+ / -10kbp) used in the analysis are from the cisTarget database.

[0101] S2 constructs a prior gene-gene functional relationship network for genes involved in the same functional pathway.

[0102] Among them, the functional pathway database used in this embodiment is a multi-source integrated gene pathway database. The database is composed of data entries of multiple gene pathway databases, including entries of GO subset, KEGG subset, Reactome subset and Biocarta subset in MSigDB database (July 2021 version). In this embodiment, the data from different database sources is filtered, de-redundant and combined. First, in order to avoid that the definition of the pathway entry is too broad or too narrow, this embodiment only retains the entries with gene number in the range of 10-500. In this embodiment, the similarity of the annotated pathways from different data sources is evaluated by Tanimoto coefficient. The pathways with similarity score higher than 0.8 are combined into the same pathway, and the new pathway name inherits the shorter one of the combined pathway names. The similarity score calculation formula is as follows:

[0103]

[0104] wherein , are two sets of biologically meaningful genes, each of which represents a gene pathway in this embodiment, and each element in the gene pathway is composed of genes that have been reported to be involved in the work of the pathway. represents the number of genes common to the two sets, represents the square value of the number of genes in the set.

[0105] After S3 completes data preprocessing, based on the prior molecular relationship network and single-cell transcription factor expression profile, single-cell gene expression profile, the observation data is subjected to multi-factor non-negative matrix factorization, and transcription factor-gene function co-expression modules in a low-dimensional space are identified.

[0106] wherein this step is a semi-supervised learning method, and the co-expression relationship of the observation relative small community is set to 60. The iteration number is set to 5000, and the minimum convergence residual is set to 1e-8, 、 、 、 0.5.

[0107] S4 subsequently uses Mann-Whitney U test to evaluate the degree of association between all identified transcription factor-gene function co-expression modules and the aging phenotype, and calculates whether the module characteristic values of these modules in the same cell subpopulation have statistical differences between the young and aging phenotypes. After the p value is corrected using the Holm multiple test correction method, the co-expression modules with an adjusted p value less than 1e-5 are considered to be significantly related to the phenotype change. The interaction module.

[0108] This embodiment further analyzes the coefficient matrix of the above-mentioned transcription factor-gene function synergistic interaction module significantly related to aging, and respectively detects 3, 6 and 5 transcription factor-gene-gene pathway synergistic interaction modules significantly related to the aging phenotype in the kidney tissue, the spleen tissue and the lung tissue.

[0109] S5 can be visualized.

[0110] wherein the third transcription factor-gene function synergistic interaction module related to aging in the kidney macrophage in this embodiment is used for visual display and functional interaction relationship explanation. The initial relationship network of the molecular interaction relationship in the module is generated by the R open source software package ggraph, and then the network graph is imported into the open source software Cytoscape through the RCy3 package of R to optimize the network visualization style. The specific results of the embodiments are shown in the accompanying Figure 2 .

[0111] ​Figure 2 This image shows a network of transcription factor-gene functional interactions significantly associated with aging within renal red pulp macrophages (not all node information is displayed due to image resolution limitations). The network consists of various transcription factors (NFKB1, REL, RELA, ATF5, ATF3, IRF2, CREB1, FOSB, JUNB, FLI1, SREBF2, BMYC, AHR, etc.), multiple gene pathways (cell antigen receptor (TCR) pathway, glucocorticoid receptor (GCR) pathway, fMLP pathway, cleavage-activated protein kinase (MAPK) pathway, integrin signaling pathway), and multiple genes (…). Jun , Tgfb1 , Calm1 , Calm2 , Ncf2 , Itgb1 , Capns1 , Csnk1a1 , Actb It is composed of (etc.).

[0112] Among them, NFKB1, REL, RELA, ATF5, ATF3, IRF2, CREB1, FOSB, and JUNB are all involved in the regulation of inflammatory responses. In particular, NFKB1, REL, and RELA of the NF-kappa-B family are not only important regulators of macrophage activation, but also cross-linked with various biological processes that regulate lifespan, such as IGF-1, growth hormone pathway, SIRT, FOXO, and mTOR.

[0113] In addition to transcription factors of the NF-kappa-B family, transcription factors involved in interferon response regulation, such as ATF5, ATF3, and IRF2, are also core interacting molecules in this co-regulatory module.

[0114] The results support previous studies observing elevated NF-kappa-B expression levels in various aged tissues, and studies have shown that inhibiting NF-kappa-B expression in mouse models can delay the onset of aging phenotypes and age-related diseases. Similarly, studies have shown that aging causes uncontrolled interferon expression, leading to dysregulation of inflammatory and immune responses.

[0115] The above results demonstrate that the molecular synergistic regulatory mechanisms identified by the method developed in this invention, based on single-cell omics data, can indeed yield biologically significant results in real-world data applications. Furthermore, based on prior molecular relationships, it can be logically inferred that changes in the aging phenotype of macrophages in mouse kidneys are closely related to NF-κB and interferon-mediated inflammatory responses. These upstream transcription factors regulate intermediate target genes such as… Jun , Tgfb1The expression change of the gene can indirectly cause the change of TCR pathway, GCR pathway, fMLP pathway, MAPK pathway, etc.

[0116] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, but not to limit it. Although the present application has been described in detail with reference to the above examples, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced by equivalents after reading the present application. However, these modifications or changes are still within the scope of the present application.

Claims

1. A method for modeling transcription factor regulatory networks based on single-cell transcriptome data, characterized in that, The steps include the following: S1. Extract multiphenotype single-cell omics data, perform data cleaning, and integrate the cleaned data; S2. Based on a public biological knowledge base, the data processed by S1 is analyzed to construct a priori network of intermolecular interactions; S3. Based on the multi-factor nonnegative matrix factorization algorithm, a multi-dimensional molecular cooperative interaction module is established according to the data processed by S1 and the prior intermolecular interaction network of S2. S4. Phenotype-related interaction module in the multidimensional molecular cooperative interaction module; S5. Visualize and export the multidimensional molecular cooperative interaction module and the phenotype-related interaction module; The data cleaning in S1 includes: S101 sets the filtering conditions; The filtering conditions include at least one of the following: cells with low abundance in multiphenotype single-cell omics data, cells contaminated with cell debris, apoptotic or lysed cells, and multimers. S102 filters multiphenotype single-cell omics data according to filtering conditions to obtain filtered data; S103 performs feature recognition on the filtered data and integrates the recognized data; the feature recognition includes at least cell clustering and cell characteristic gene recognition. The intermolecular interaction network in S2 includes the regulatory relationship network between transcription factors and target genes expressed in single-cell data, and the functional association network between genes expressed in single-cell data. In the construction of the functional association network between genes, the associated genes are involved in at least one of the following: regulating the same biological process, participating in the same gene pathway, or responding to the same phenotype. The functional association network between genes includes, but is not limited to, one or more of the following: gene association network with co-epigenetic modification, receptor-ligand interaction network of gene-encoded protein, and protein-protein interaction network of gene-encoded protein. The S4 module for calculating the multidimensional molecular synergistic interaction relationship module and the phenotypic interaction relationship module specifically involves: using the statistical test of significant differences, detecting the correlation between variables reflecting multidimensional molecular synergistic interaction relationships in the low-dimensional space obtained after multi-factor nonnegative matrix factorization and phenotypic changes, and obtaining the multidimensional molecular synergistic interaction relationship module that is significantly related to the phenotypic relationship. The statistical tests for the significance of the differences include, but are not limited to, one of the following: Student's t-test, Mann-Whitney U test, and analysis of variance.

2. The method for modeling transcription factor regulatory networks based on single-cell transcriptome data according to claim 1, characterized in that, The multi-factor nonnegative matrix factorization algorithm is as follows: S301 sets the total number of observed cells as n, the total number of observed genes as m, and the total number of observed transcription factors as s, and establishes a [system / mechanism]. A nonnegative matrix of dimension, set as multi-phenotype single-cell gene expression profile data. ; establish A non-negative matrix of dimension, set as single-cell regulator activity matrix data. ; S302 sets the number of all transcription factor-gene functional co-interaction modules observed in n cells to be k, and establishes a... A nonnegative matrix W of dimension 1; establish A nonnegative matrix of dimension 1, used to describe the weighted relationship between variables and genes in a low-dimensional space, is denoted as . ; establish A nonnegative matrix of dimension 1, used to describe the weight relationship between variables and transcription factors in low-dimensional space, is denoted as . ; W satisfies , And the error in the square of the factorization is: in For matrix Based on the Frobenius norm and the square error, construct the objective function: 。 3. The method for modeling transcription factor regulatory networks based on single-cell transcriptome data according to claim 2, characterized in that, In S302 The formula can be further written as: in The trace coefficient, and Let A be the trace matrix of the matrix, A be the adjacency matrix of prior gene-gene function relationships, and B be the adjacency matrix of prior transcription factor-target gene function associations. for The vector in the i-th column of the matrix, for The vector in the j-th column of the matrix; The objective function for maximizing the prior relation adjacency matrix is: 。 4. The method for modeling transcription factor regulatory networks based on single-cell transcriptome data according to claim 1, characterized in that, The visualization export in S5 specifically involves using the multidimensional molecular synergistic interaction module and phenotype-related interaction module obtained in S4, and based on the corresponding transcription factor coefficient matrix and gene coefficient matrix, in conjunction with the prior intermolecular interaction network in S2, to generate a visualized transcription factor-gene functional synergistic relationship network.

5. The method for modeling transcription factor regulatory networks based on single-cell transcriptome data according to claim 4, characterized in that, The visualized transcription factor-gene function synergistic relationship network includes several synergistic relationship network nodes. Each synergistic relationship network node is composed of multiple transcription factors, multiple genes, and multiple biological functions. The edge of the visualized transcription factor-gene function synergistic relationship network is composed of transcription factor-gene regulatory relationships, gene-gene interaction relationships, and gene-gene function relationships.

Citation Information

Patent Citations

  • Single-cell network regulation relation construction method

    CN108090326A

  • Method for deducing gene regulation network by using single cell transcription and gene knockout data

    CN110517724A