Transcription factor activity inferring method, apparatus, storage medium, and computer device
Patent Information
- Application Number
- PCT/CN2024/080575
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-10-02
AI Technical Summary
Existing transcription factor activity inference methods have poor accuracy when based on spatial transcriptome sequencing data, mainly due to the lack of utilization of spatial location information.
By obtaining the spatial transcriptome data of the target slice, the nuclear density is estimated using the position information and gene expression information, the correlation score between the transcription factor and the target gene is calculated, the spatial co-localization module is constructed and motif enrichment analysis is performed, the spatial regulator of the transcription factor is determined, and finally a score is performed to infer the activity of the transcription factor.
The accuracy of transcription factor activity inference is improved by combining spatial location information and expression level information.
Smart Images

Figure CN2024080575_02102025_PF_FP_ABST
Abstract
Description
Transcription factor activity inference method, device, storage medium and computer equipment Technical Field
[0001] The present disclosure relates to the field of gene sequencing technology, and in particular to a method, apparatus, storage medium, and computer equipment for inferring transcription factor activity. Background Art
[0002] Transcription factors (TFs), also known as sequence-specific DNA-binding factors, control the rate of transcription of genetic information from DNA to messenger RNA by binding to specific DNA sequences. Transcription factors ensure that genes are expressed in the desired cells at the right time and in the right amount throughout the life of cells and organisms. Therefore, inferring transcription factor activity is a crucial task in bioinformatics analysis.
[0003] Currently, there is a mature transcription factor activity inference scheme for single-cell transcriptome data, but the accuracy of transcription factor activity inference for spatial transcriptome sequencing data based on this scheme is poor.
[0004] Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to propose a method, device, storage medium and computer equipment for inferring transcription factor activity, which can improve the accuracy of inferring transcription factor activity.
[0006] To achieve the above objectives, a first aspect of the embodiments of the present application provides a method for inferring transcription factor activity, the method comprising:
[0007] Acquiring spatial transcriptome data of a target slice, wherein the spatial transcriptome data includes location information and gene expression information of multiple categories of genes;
[0008] Screening genes in the gene expression information according to a reference regulatory network corresponding to the target slice, and constructing a target regulatory network based on the screened genes, wherein the target regulatory network includes a plurality of first target regulatory data pairs, each first target regulatory data pair including a transcription factor and a corresponding target gene;
[0009] performing kernel density estimation on the transcription factor and the target gene in each first target regulation data pair according to the position information and the gene expression information, and calculating a correlation score for each first target regulation data pair according to the kernel density estimation result;
[0010] constructing a spatial colocalization module for each transcription factor according to the correlation score, and determining a spatial regulator of each transcription factor according to a motif enrichment analysis result of the spatial colocalization module;
[0011] The spatial regulator of each transcription factor is scored, and the activity inference result of the corresponding transcription factor is determined based on the scoring result.
[0012] A second aspect of the present application provides a transcription factor activity inference device, comprising:
[0013] An acquisition unit, configured to acquire spatial transcriptome data of a target slice, wherein the spatial transcriptome data includes location information and gene expression information of multiple categories of genes;
[0014] a construction unit, configured to screen genes in the gene expression information according to a reference regulatory network corresponding to the target slice, and construct a target regulatory network based on the screened genes, wherein the target regulatory network includes a plurality of first target regulatory data pairs, each first target regulatory data pair including a transcription factor and a corresponding target gene;
[0015] a calculation unit, configured to perform kernel density estimation on the transcription factor and the target gene in each first target regulation data pair according to the position information and the gene expression information, and calculate a correlation score for each first target regulation data pair according to the kernel density estimation result;
[0016] a determination unit, configured to construct a spatial colocalization module for each transcription factor according to the correlation score, and determine a spatial regulator of each transcription factor according to a motif enrichment analysis result of the spatial colocalization module;
[0017] The inference unit is used to score the spatial regulator of each transcription factor and determine the activity inference result of the corresponding transcription factor based on the scoring result.
[0018] In one embodiment, the construction unit comprises:
[0019] A first acquisition subunit is configured to acquire a reference regulatory network corresponding to the target slice, wherein the reference regulatory network includes a plurality of reference regulatory edges and a reference transcription factor and a reference target gene connected to each reference regulatory edge;
[0020] an identification subunit, configured to identify an intersection of the genes in the gene expression information, the reference transcription factors, and the reference target genes, to obtain a plurality of first transcription factors and a plurality of first target genes;
[0021] A first determining subunit is configured to determine, based on the gene expression information, a plurality of second transcription factors and a plurality of second target genes whose corresponding capture sites are greater than a first preset threshold value among the plurality of first transcription factors and the plurality of first target genes;
[0022] The first construction subunit is used to construct a target regulatory network based on the multiple second transcription factors and the multiple second target genes.
[0023] Optionally, in some embodiments, the first construction subunit includes:
[0024] A search module, configured to search the reference regulatory network for at least one second transcription factor corresponding to each second target gene;
[0025] A construction module is used to construct a target regulatory network based on the multiple second target genes and at least one second transcription factor corresponding to each second target gene.
[0026] Optionally, in some embodiments, the computing unit includes:
[0027] A second determining subunit is configured to determine, for any of the first target regulation data pairs, a position matrix of a corresponding third target gene and a corresponding third transcription factor according to the position information and the gene expression information;
[0028] A first calculation subunit is used to calculate the kernel density value of the third target gene and the third transcription factor in each spatial site according to the position matrix to obtain a kernel density estimation result;
[0029] The second calculation subunit is used to calculate the correlation score of each first target regulation data pair according to the kernel density estimation result.
[0030] Optionally, in some embodiments, the second computing subunit includes:
[0031] a first determining module, configured to determine, in the target slice, a first region where the nuclear density value of the third target gene is greater than a second preset threshold, and to determine a second region where the nuclear density value of the third transcription factor is greater than a third preset threshold;
[0032] an identification module, configured to identify a target area where the first area intersects the second area;
[0033] A calculation module is used to calculate the Pearson correlation coefficient between the nuclear density of the third target gene and the third transcription factor in the target region to obtain a correlation score of the first target regulation data pair.
[0034] Optionally, in some embodiments, the determining unit includes:
[0035] a screening subunit, configured to screen the plurality of first target regulation data pairs according to the correlation scores to obtain a plurality of second target regulation data pairs;
[0036] a second construction subunit, configured to construct a spatial co-localization module for each transcription factor according to the plurality of second target regulation data pairs;
[0037] An analysis subunit, used to perform motif enrichment analysis on each spatial co-localization module to obtain motif enrichment analysis results;
[0038] The third determining subunit is used to determine the spatial regulator of each transcription factor according to the motif enrichment analysis results.
[0039] In some embodiments, the screening subunit comprises:
[0040] A first acquisition module, configured to acquire a fourth preset threshold for controlling the correlation score;
[0041] The second determining module is configured to determine, from among the plurality of first target regulation data pairs, a plurality of second target regulation data pairs having correlation scores greater than the fourth preset threshold.
[0042] Optionally, in some embodiments, the second construction subunit includes:
[0043] a first division module, configured to divide the plurality of second target regulation data pairs into a plurality of first categories according to different transcription factors, wherein the second target regulation data pairs of each first category contain the same transcription factor;
[0044] A second division module is used to divide the plurality of second target regulation data pairs into a plurality of second categories according to different target genes, wherein the second target regulation data pairs of each second category contain the same target gene;
[0045] a third determining module, configured to, for each transcription factor, determine, in each second target regulation data pair of the first category, a second target regulation data pair corresponding to a target gene having a correlation score with the transcription factor greater than a fifth preset threshold as a first subspace co-localization module;
[0046] The fourth determination module is used to determine, for each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the sixth preset threshold as the second subspace co-localization module
[0047] a fifth determination module, configured to, for each transcription factor, determine, in each second target regulation data pair of the first category, a second target regulation data pair corresponding to a target gene having a correlation score with the transcription factor greater than a seventh preset threshold as a third subspace co-localization module;
[0048] a sixth determination module, configured to, for each transcription factor, in each second category, determine a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in a preset first range with the target gene as a third regulatory data pair, and integrate all third regulatory data pairs found in the second category into a fourth subspace co-localization module;
[0049] a seventh determination module, configured to, for each transcription factor, in each second category, determine a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in a preset second range with the target gene as a fourth target regulatory data pair, and integrate all fourth target regulatory data pairs found in the second category into a fifth subspace co-localization module;
[0050] an eighth determination module, configured to, for each transcription factor, in each second category, determine a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in a preset third range with the target gene as a fifth target regulatory data pair, and integrate all fifth target regulatory data pairs found in the second category into a sixth subspace co-localization module;
[0051] The first construction module is used to construct a spatial colocalization module for each transcription factor based on the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module and the sixth subspace colocalization module corresponding to each transcription factor.
[0052] Optionally, in some embodiments, the analysis subunit includes:
[0053] The second acquisition module is used to obtain a preset motif enrichment analysis function;
[0054] An analysis module is used to perform motif enrichment analysis on each subspace colocalization module within each spatial colocalization module based on the preset motif enrichment analysis function to obtain a motif enrichment analysis result.
[0055] Optionally, in some embodiments, the third determining subunit includes:
[0056] a screening module, configured to screen, based on the motif enrichment analysis results, a seventh subspace colocalization module having an enrichment score greater than an eighth preset threshold and in which the enriched transcription factors are consistent with the transcription factors in the subspace colocalization module;
[0057] a fifth determination module, configured to determine a fourth target gene in the seventh subspace co-localization module whose enrichment contribution is greater than a ninth preset threshold;
[0058] The second construction module is used to construct a spatial regulator of each transcription factor based on the transcription factor in the seventh subspace co-localization module and the fourth target gene.
[0059] Optionally, in some embodiments, the inference unit includes:
[0060] A second acquisition subunit is used to obtain a preset scoring tool and score each spatial regulator based on the preset scoring tool to obtain a scoring result;
[0061] a fourth determining subunit, configured to determine a score distribution of each regulator in a plurality of cells based on the scoring result;
[0062] The fifth determining unit is configured to perform binary division on the plurality of cells based on a binary Gaussian mixture model, and determine the activity of the transcription factor in each cell according to the binary division result.
[0063] A third aspect of an embodiment of the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method for inferring transcription factor activity described in the first aspect is implemented.
[0064] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the transcription factor activity inference method described in the first aspect.
[0065] To achieve the above-mentioned purpose, the fifth aspect of the embodiment of the present application proposes a computer program product, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the transcription factor activity inference method described in the first aspect.
[0066] The transcription factor activity inference method, device, storage medium and computer equipment proposed in the embodiments of the present application are as follows: the method obtains spatial transcriptome data of a target slice, wherein the spatial transcriptome data includes position information and gene expression information of multiple categories of genes; the genes in the gene expression information are screened according to a reference regulatory network corresponding to the target slice, and a target regulatory network is constructed based on the screened genes, wherein the target regulatory network includes multiple first target regulatory data pairs, each first target regulatory data pair includes a transcription factor and a corresponding target gene; kernel density estimation is performed on the transcription factors and target genes in each first target regulatory data pair according to the position information and gene expression information, and the correlation score of each first target regulatory data pair is calculated according to the kernel density estimation result; a spatial co-localization module of each transcription factor is constructed according to the correlation score, and the spatial regulator of each transcription factor is determined according to the motif enrichment analysis result of the spatial co-localization module; the spatial regulator of each transcription factor is scored, and the activity inference result of the corresponding transcription factor is determined according to the scoring result.
[0067] Thus, this method uses spatial location information and gene expression levels in spatial transcriptome sequencing data to calculate correlations between transcription factors and corresponding target genes in the regulatory network, obtaining relatively accurate spatial correlations. Further transcription factor activity inference is then performed based on this spatial correlation. Compared to related techniques that perform correlation calculations and infer transcription factor activity based solely on gene expression levels, the transcription factor activity inference method provided by this disclosure can obtain more accurate transcription factor activity inference results. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0069] FIG1 is a flow chart of a method for inferring transcription factor activity provided in an embodiment of the present application;
[0070] FIG2 is a flow chart of step 102 in FIG1 ;
[0071] FIG3 is a schematic diagram of the constructed target regulatory network;
[0072] FIG4 is a flow chart of step 103 in FIG1 ;
[0073] FIG5 is a schematic diagram of the kernel density estimation of transcription factors and target genes;
[0074] FIG6 is a flow chart of step 104 in FIG1 ;
[0075] Figures 7a-7f are schematic diagrams of six subspace colocalization modules of transcription factors;
[0076] FIG8 is a flow chart of step 105 in FIG1 ;
[0077] FIG9 is a schematic diagram of a text file in gem format obtained in this embodiment;
[0078] FIG10 is a schematic diagram of the spatial correlation coefficient of the TF-target output in this embodiment;
[0079] FIG11 is a schematic diagram of the spatial transcriptome gene expression matrix provided in this embodiment;
[0080] FIG12 is a schematic diagram of the enrichment results obtained by the enrichment analysis in the present disclosure;
[0081] FIG13 is a schematic diagram of the spatial regulator obtained by integration in this embodiment;
[0082] FIG14 is a schematic diagram of the AUC scores calculated in this embodiment;
[0083] FIG15 is a schematic diagram of a binary matrix of AUC scores in this embodiment;
[0084] FIG16 is a schematic diagram of a clustering result based on AUC scores in this embodiment;
[0085] FIG17 is another schematic diagram of clustering results based on AUC scores in this embodiment;
[0086] FIG18 is a graph showing the spatial location and marker gene expression of hepatocyte subtypes identified by the AUC matrix of the present application;
[0087] FIG19 is a structural diagram of a transcription factor activity inference device provided in an embodiment of the present application;
[0088] FIG20 is a schematic diagram of the hardware structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0089] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0090] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:
[0091] Single-cell transcriptome: The total mRNA expression level within a single cell at a given moment reflects the overall characteristics of that cell. The emergence of single-cell transcriptomics technology has enabled the precision of research to be refined from the multicellular level of tissues down to the single-cell domain, enabling the study of the specific characteristics of individual cells or groups of cells. This technology plays a key role in understanding cell development, the tumor microenvironment, and single-cell mapping.
[0092] Spatial gene expression (ST-seq) is a technology that combines transcriptomics, single-cell sequencing technology and tissue section technology. Based on single-cell transcriptomics, spatial transcriptomics technology can further obtain spatial distribution information of different types of cells.
[0093] Regulatory network: Specifically, the transcription factor-target gene regulatory network (GRN) is a complex regulatory network composed of interactions between multiple target genes and multiple transcription factors. Within a cell, a single transcription factor typically regulates the expression of multiple genes, while the expression of a single target gene is often regulated by multiple transcription factors. Furthermore, a transcription factor itself can act as a target gene, with the expression of other transcription factors regulating its expression.
[0094] Regulons: Also known as regulators, a transcription factor and all target genes regulated by it are called a regulator. The regulators in the present disclosure eliminate false positive target genes from target genes.
[0095] Spatial transcriptomics technology has developed rapidly in recent years, and its application in clinical medicine and scientific discovery has continued to deepen. The most significant feature of spatial transcriptomics data is that it reveals the spatial positional relationship between molecules and cells in biological tissues. This spatial dimension of information provides a new dimension for understanding the biological behavior at the cellular and molecular levels within tissues and organs.
[0096] However, the current analysis method for spatial transcriptome data still needs to be improved, such as the analysis of most of the current transcription factor (Transcript factor, TF) activity inference still uses tools based on traditional transcriptome sequencing (Bulk-RNA seq) or single-cell transcriptome data development, such as, SCENIC (single-cell regulatory network inference and clustering) i.e. single-cell regulatory network reasoning and clustering, decoupleR etc., these tools do not utilize the spatial position information in spatial transcriptome data. In addition, the characteristics of spatial transcriptome data, in addition to including spatial position information, also have stronger data sparsity. This is because the capture of the RNA of the current spatial transcriptome sequencing scheme is affected by permeabilization conditions, and the heterogeneity of cells causes it to be impossible for all cells to reach ideal permeabilization conditions and release all mRNA. Therefore, the gene capture of spatial transcriptome is often inferior to single-cell transcriptome data, with stronger data sparsity. This means that the amount of gene expression information contained in spatial transcriptome data is lower than that in single-cell transcriptome data. If SCENIC, decoupleR and other schemes are still used for spatial transcriptome sequencing data, and TF activity is inferred based solely on expression information, the inaccuracy of the results will be greatly increased due to the lack of spatial information and insufficient expression information.
[0097] Therefore, there is an urgent need to develop a transcription factor activity inference tool that can utilize the spatial location information of spatial transcriptome data. To this end, the present disclosure provides a transcription factor activity inference method to improve the accuracy of transcription factor activity inference for spatial transcriptome data.
[0098] The following describes the method for inferring transcription factor activity provided in the examples of the present application.
[0099] 1 , in some embodiments, the transcription factor activity inference method provided in the embodiments of the present application includes but is not limited to steps 101 to 105 .
[0100] Step 101, obtaining spatial transcriptome data of a target slice;
[0101] Step 102, screening genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and constructing a target regulatory network based on the screened genes;
[0102] Step 103, performing kernel density estimation on the transcription factor and target gene in each first target regulation data pair based on the position information and gene expression information, and calculating the correlation score of each first target regulation data pair based on the kernel density estimation result;
[0103] Step 104, constructing a spatial co-localization module for each transcription factor based on the correlation score, and determining a spatial regulator of each transcription factor based on a motif enrichment analysis result of the spatial co-localization module;
[0104] Step 105 , scoring the spatial regulator of each transcription factor, and determining the activity inference result of the corresponding transcription factor based on the scoring result.
[0105] In the steps 101 to 105 shown in the embodiment of the present application, when the cells in the target slice are inferred to have transcription factor activity, the spatial transcriptome data of the target slice can be obtained first. The spatial transcriptome data is the data obtained by sequencing the target slice using spatial transcription sequencing technology, which includes the spatial position information of multiple sequencing sites (spot sites) and the quantity information of the messenger RNA captured by each spot site, and the quantity information of mRNA can be referred to as gene expression information. Wherein, the target slice can be a biological slice of any organism, such as a slice of a certain organ of an animal or a plant slice. The spatial transcriptome data can specifically be a gem file in text or compressed text format.
[0106] After obtaining the spatial transcriptome data of the target slice, the type of gene captured and the distribution of each type of gene can be determined according to the gene expression information therein, and the spatial position matrix of each type of gene is obtained. The gene name and spatial two-dimensional coordinates of each gene are included in the spatial position matrix. Then, the reference regulatory network corresponding to the target slice can be used to screen the genes captured in the spatial transcriptome data, and the genes not present in the reference regulatory network are eliminated to obtain the genes after preliminary screening, and a preliminary target regulatory network is constructed based on the relationship between the genes after these preliminary screening and the target genes in the reference regulatory network. The target regulatory network includes the connection edges between multiple transcription factors and multiple target genes, and between transcription factors and target genes. One of the transcription factors can correspond to multiple target genes, and the same target gene can also correspond to multiple transcription factors. A transcription factor and a corresponding target gene can form a regulatory data pair, which is called the first target regulatory data pair in the present embodiment. It is understandable that the process of determining the preliminary target regulatory network in step 102 is similar to the process of constructing a coexpression module in the SCENIC scheme, which is to first screen out some low-quality genes, thereby obtaining a coexpression module of transcription factors and relatively high-quality genes.
[0107] Further, in step 103, the position information and gene expression information in the spatial transcriptome data are used to calculate the correlation score between the transcription factor and the target gene in the first target regulation data pair. Wherein, this step is the core difference between the transcription factor activity inference method provided by the present disclosure and the SCENIC scheme. In the SCRNIC scheme, the correlation weight between the transcription factor and the target gene is calculated only according to the gene expression information of different genes, and then the false positive data pairs are eliminated based on the correlation weight to obtain final regulons, and finally the transcription factor activity is inferred based on the scoring of the regulons; whereas in the disclosed embodiment, not only gene expression information is used, but also the spatial position information of the gene is used to calculate the correlation score between the transcription factor and the target gene in the first target regulation data pair, and the score can therefore also be referred to as a spatial correlation score. Then, according to the spatial correlation score, the false positive data pairs are further proposed to obtain final regulons, which can also be referred to as spatial regulators (spatial regulons) herein, and finally the transcription factor activity is inferred based on the scoring of the spatial regulators.
[0108] Specifically, in this process, the position information and gene expression information in the spatial transcriptome data can be used to perform kernel density estimation on the transcription factors and target genes in the first target regulation data pair, and then the spatial correlation score between the transcription factors and target genes can be calculated based on the kernel density estimation results.
[0109] After calculating the spatial correlation score of each first target regulation data pair based on the spatial position information and gene expression information in the spatial transcriptome data, the spatial co-localization module of each transcription factor can be constructed according to the spatial correlation score of the first target regulation data pair, and the process is similar to the construction of the coexpression module in the SCENIC scheme.Then, motif (motif) enrichment analysis can be carried out to the spatial co-localization module obtained by building, and according to the motif enrichment analysis result, the target gene in the spatial co-localization module is screened again, to remove false positive target genes therein, and determine the spatial regulator of each transcription factor according to the transcription factor after screening and the corresponding target gene.Finally, the scoring tool (AUCell) provided by SCENIC can be used to score each spatial regulator, and determine the activity inference result of each transcription factor according to the scoring result.
[0110] That is, the transcription factor activity inference method provided in this case is a method for inferring transcription factor activity based on spatial transcriptome data. This method adaptively improves the SCENIC scheme of single-cell transcriptomes. When calculating the correlation between transcription factors and target genes, it introduces the spatial position information and expression information of each type of gene in the spatial transcriptome data, and calculates a more accurate spatial correlation between transcription factors and target genes, thereby improving the accuracy of transcription factor activity inference based on spatial transcriptome data.
[0111] Referring to FIG. 2 , in some embodiments, step 102 includes but is not limited to steps 201 to 204 .
[0112] Step 201: Obtain a reference regulatory network corresponding to a target slice, where the reference regulatory network includes a plurality of reference regulatory edges and reference transcription factors and reference target genes connected to each reference regulatory edge;
[0113] Step 202: identifying the intersection of genes in the gene expression information, reference transcription factors, and reference target genes to obtain a plurality of first transcription factors and a plurality of first target genes;
[0114] Step 203: Based on the gene expression information, determine, among the plurality of first transcription factors and the plurality of first target genes, a plurality of second transcription factors and a plurality of second target genes whose corresponding capture sites are greater than a first preset threshold;
[0115] Step 204 : constructing a target regulatory network based on the plurality of second transcription factors and the plurality of second target genes.
[0116] In the present embodiment, the screening process of genes in the acquired spatial transcriptome data can be implemented in two steps: a first step of screening based on the reference regulatory network and a second step of screening based on the number of captured sites of the gene, thereby obtaining genes that are widely distributed and relatively common in the slice.
[0117] Specifically, the reference regulatory network corresponding to the target slice can be first obtained, and the reference regulatory network can be specifically TF-target regulatory network public database (such as NicheNet database etc.) here.After obtaining the reference regulatory network, the information in the gene information and the spatial transcriptome data provided in the reference regulatory network can be taken intersection, the process of taking intersection can eliminate some uncommon genes in the spatial transcriptome data, thereby obtain multiple first transcription factors and multiple first target genes. Further, the capture site number that can be captured according to each first transcription factor and each first target gene in gene expression information, the gene that is less expressed or too concentrated in the first transcription factor and the first target gene is eliminated, thereby obtain multiple second transcription factors and multiple second target genes. For example, it is 100 that can be designed that the first preset threshold value is, when the capture site number corresponding to target gene is less than 100, then it is eliminated.Then build preliminary target regulatory network according to multiple second transcription factors and multiple second target genes.
[0118] In some embodiments, constructing a target regulatory network based on multiple second transcription factors and multiple second target genes includes:
[0119] Searching for at least one second transcription factor corresponding to each second target gene in the reference regulatory network;
[0120] A target regulatory network is constructed based on multiple second target genes and at least one second transcription factor corresponding to each second target gene.
[0121] After the genes in the spatial transcriptome data are pre-processed, that is, after screening to obtain multiple second target genes, the transcription factor corresponding to each second target gene can be further searched in the reference regulatory network, which can be referred to as the second transcription factor herein. Wherein, each second target gene can have one or more corresponding transcription factors. Then, a target regulatory network can be constructed based on at least one second transcription factor corresponding to multiple second target genes and each second target gene. As shown in Figure 3, a schematic diagram of the target regulatory network constructed is shown. The target regulatory network consisting of 2 transcription factors and 3 target genes is shown in the figure, wherein the number of transcription factors and target genes is only an example, and generally more transcription factors and target genes can be included in the regulatory network. As shown in the figure, multiple TF-target (target gene) data pairs are included in the target regulatory network, which can be referred to as the first target regulatory data pair in the present embodiment.
[0122] Referring to FIG. 4 , in some embodiments, step 103 includes but is not limited to steps 401 to 403 .
[0123] Step 401: for any first target regulation data pair, determine the position matrix of the corresponding third target gene and the corresponding third transcription factor according to the position information and gene expression information;
[0124] Step 402, calculating the kernel density value of the third target gene and the third transcription factor in each spatial site according to the position matrix to obtain a kernel density estimation result;
[0125] Step 403: Calculate the correlation score of each first target regulation data pair based on the kernel density estimation result.
[0126] In the disclosed embodiment, after constructing a preliminary target regulatory network, the correlation between the transcription factors and target genes in the TF-target data pairs in the target regulatory network can be further calculated to obtain the correlation score corresponding to each first target regulatory data pair.
[0127] Specifically, the third target gene and the third transcription factor can be estimated based on the spatial position matrix of each target gene (to distinguish it from the aforementioned target gene, referred to herein as the third target gene) and transcription factor (to distinguish it from the target gene shown elsewhere in this article, referred to herein as the third transcription factor) in the target regulatory network to obtain a kernel density value. Specifically, for any third transcription factor and its corresponding third target gene, the position matrix of each third target gene and the size of the entire transcriptome slice (the default is the range size of the x, y coordinates in the gem file divided by 10) can be input to calculate the density estimate of the occurrence of the third target gene at each spot site. As shown in Figure 5, it is a schematic diagram of the kernel density estimation of transcription factors and target genes. Among them, the horizontal and vertical coordinates in the figure represent the x, y coordinates in the position matrix.
[0128] After the kernel density estimation results of each third target gene and each third transcription factor are calculated, the correlation score of each first target regulation data pair can be further calculated based on the kernel density estimation results of each third target gene and each third transcription factor.
[0129] In some embodiments, calculating the correlation score of each first target regulation data pair according to the kernel density estimation result includes:
[0130] Determining, in the target slice, a first region where a nuclear density value of a third target gene is greater than a second preset threshold, and determining a second region where a nuclear density value of a third transcription factor is greater than a third preset threshold;
[0131] identifying a target area where the first area intersects the second area;
[0132] The Pearson correlation coefficient between the nuclear density of the third target gene and the first target transcription factor in the target region was calculated to obtain the correlation score of the first target regulation data pair.
[0133] In the embodiment of the present application, for each target regulatory data pair, its correlation score can be calculated by calculating the Pearson correlation coefficient (PCC) of the kernel density estimates of the transcription factor and the target gene contained therein in the intersection area of the kernel density key areas.
[0134] Specifically, for any target regulatory data pair, regions with estimated kernel density values greater than the N% quantile can be selected. For example, if N is 50, this region corresponds to the top 50% of spatial sites ranked by kernel density values. Thus, a second preset threshold for the kernel density values before and after the demarcation of the first and second 50% of sites can be calculated based on the kernel density of all sites of the third target gene in the target slice, and a third preset threshold for the kernel density values before and after the demarcation of the first and second 50% of sites can be calculated based on the kernel density of all sites of the first target transcription factor in the target slice. Then, a first region in the target slice where the kernel density of the third target gene exceeds the second preset threshold is identified, as well as a second region in the target slice where the kernel density of the first target transcription factor exceeds the third preset threshold. The intersection of the first and second regions is then determined, thereby determining the target region where the first and second regions intersect. Furthermore, the Pearson correlation coefficient between the kernel densities of the third target gene and the third transcription factor in the target region can be calculated to obtain a correlation score for the first target regulatory data pair.
[0135] Then, each first target regulation data pair is traversed, and the correlation score corresponding to each first target regulation data pair is calculated using the above method.
[0136] Referring to FIG. 6 , in some embodiments, step 104 includes but is not limited to the following steps 601 to 604 .
[0137] Step 601: screening multiple first target regulation data pairs according to the correlation scores to obtain multiple second target regulation data pairs;
[0138] Step 602 , constructing a spatial co-localization module for each transcription factor based on the plurality of second target regulatory data pairs;
[0139] Step 603, performing motif enrichment analysis on each spatial co-localization module to obtain motif enrichment analysis results;
[0140] Step 604: determine the spatial regulator of each transcription factor based on the motif enrichment analysis results.
[0141] In the disclosed embodiment, after calculating the correlation score of each first target regulation data pair, these first target regulation data pairs can be screened according to the correlation score to eliminate the first target regulation data pairs with lower correlation scores. For example, the first target regulation data pairs with a correlation score less than 0.03 can be eliminated, and the second target regulation data pairs with a correlation score greater than or equal to 0.03 can be retained. Then, the spatial co-localization module of each transcription factor can be constructed based on the multiple second target regulation data pairs obtained by screening. Among them, the spatial co-localization module corresponding to each transcription factor is composed of six subspace co-localization modules, and the six subspace co-localization modules can be constructed using different strategies. As shown in Figures 7a-7f, it is a schematic diagram of the spatial co-localization module of the transcription factor.
[0142] After building the spatial colocalization module corresponding to each transcription factor, motif enrichment analysis can be further carried out to each spatial colocalization module to obtain motif enrichment analysis result.Then the spatial regulator of each transcription factor can be determined according to the motif enrichment analysis result.Wherein, the structure of spatial regulator is similar to the subspace colocalization module in Fig. 7a-7f, just based on the motif enrichment analysis result, some false positive target genes in the spatial colocalization module are eliminated and removed, and final spatial regulator is obtained.
[0143] In some embodiments, a plurality of first target regulation data pairs are screened according to the correlation scores to obtain a plurality of second target regulation data pairs, including:
[0144] Obtaining a fourth preset threshold for controlling the relevance score;
[0145] A plurality of second target regulation data pairs having correlation scores greater than a fourth preset threshold are determined among the plurality of first target regulation data pairs.
[0146] When screening multiple first target control data pairs based on the correlation score, a fourth preset threshold for controlling the correlation score, such as the aforementioned 0.03 or other value, can be first obtained. Then, based on the fourth preset threshold, multiple second target control data pairs having correlation scores greater than the fourth preset threshold are determined from the multiple first target control data pairs.
[0147] In some embodiments, constructing a spatial co-localization module for each transcription factor based on a plurality of second target regulatory data pairs comprises:
[0148] Dividing the plurality of second target regulation data pairs into a plurality of first categories according to different transcription factors, wherein the second target regulation data pairs of each first category include the same transcription factor;
[0149] Dividing the plurality of second target regulation data pairs into a plurality of second categories according to different target genes, wherein the second target regulation data pairs of each second category include the same target gene;
[0150] For each transcription factor, in each second target regulation data pair of the first category, determining the second target regulation data pair corresponding to the target gene having a correlation score with the transcription factor greater than a fifth preset threshold as a first subspace co-localization module;
[0151] For each transcription factor, in each second target regulation data pair of the first category, determining the second target regulation data pair corresponding to the target gene having a correlation score with the transcription factor greater than a sixth preset threshold as a second subspace co-localization module;
[0152] For each transcription factor, in each second target regulation data pair of the first category, determining the second target regulation data pair corresponding to the target gene having a correlation score with the transcription factor greater than a seventh preset threshold as a third subspace co-localization module;
[0153] For each transcription factor, in each second category, a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked in the first preset range with the target gene is determined as a third regulatory data pair, and all third regulatory data pairs found in the second category are integrated into a fourth subspace co-localization module;
[0154] For each transcription factor, in each second category, a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in the preset second range with the target gene is determined as a fourth target regulatory data pair, and all fourth target regulatory data pairs found in the second category are integrated into a fifth subspace co-localization module;
[0155] For each transcription factor, in each second category, a second target regulation data pair corresponding to the target transcription factor and having a correlation score ranked first in the preset third range with the target gene is determined as a fifth target regulation data pair, and all fifth target regulation data pairs found in the second category are integrated into a sixth subspace co-localization module;
[0156] The spatial colocalization module of each transcription factor is constructed according to the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module and the sixth subspace colocalization module corresponding to each transcription factor.
[0157] Among them, when screening to obtain multiple second target regulatory data pairs and constructing the spatial colocalization module of each transcription factor based on the multiple second target regulatory data pairs, six different methods can be used to construct six subspace colocalization modules for each transcription factor, and then the six subspace colocalization modules together form a spatial colocalization module.
[0158] Specifically, multiple second target regulation data pairs can be divided into multiple first categories according to different transcription factors, wherein each first category of second target regulation data pairs contains the same transcription factor. In addition, multiple second target regulation data pairs can be divided into multiple second categories according to different target genes, wherein each second target regulation data pair of the second category contains the same target gene. Then, for the second target regulation data pairs of the first category, the first subspace co-localization module whose corresponding correlation score is greater than the fifth preset threshold, the second subspace co-localization module whose correlation score is greater than the sixth preset threshold, and the third subspace co-localization module whose correlation score is greater than the seventh preset threshold can be determined; for example, the fifth preset threshold can be set to 75%, the sixth preset threshold can be set to 90%, and the seventh preset threshold can be set to the value corresponding to the top 50 ranges of the correlation score ranking of the second target regulation data pairs of the first category.
[0159] In addition, in each second category of second target regulatory data pairs, the correlation scores of different transcription factors and corresponding target genes can be sorted, and then a second target regulatory data pair corresponding to the target transcription factor whose correlation score with the target gene is ranked in the top preset first range, preset second range and preset third range among the correlation scores of multiple second category second target regulatory data pairs corresponding to the target gene is determined, for example, the preset first range is ranked in the top 5, the preset second range is ranked in the top 10, and the preset third range is ranked in the top 50; the second target regulatory data pairs corresponding to these three preset ranges are respectively called the third, fourth, and fifth target regulatory data pairs, and then the third, fourth, and fifth target regulatory data pairs are found in all second categories. All third regulatory data pairs have the same target transcription factor and different target genes, and the third, fourth, and fifth target regulatory data pairs are respectively integrated into the fourth, fifth, and sixth subspace co-localization modules.
[0160] Furthermore, the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module and the sixth subspace colocalization module together constitute the final spatial colocalization module of each transcription factor.
[0161] In some embodiments, performing motif enrichment analysis on each subspace colocalization module in each spatial colocalization module to obtain a motif enrichment analysis result includes:
[0162] Get the preset motif enrichment analysis function;
[0163] Based on the preset motif enrichment analysis function, a motif enrichment analysis is performed on each subspace colocalization module in each spatial colocalization module to obtain a motif enrichment analysis result.
[0164] After constructing the spatial colocalization module for each transcription factor, a motif enrichment analysis function can be further obtained, specifically the prune2df function of the prune module in the SCENIC package. This enrichment analysis function is then used to perform motif enrichment analysis on each sub-spatial colocalization module within each spatial colocalization module to obtain motif enrichment analysis results. Specifically, TF motif enrichment analysis can be calculated using the RscisTarget algorithm to obtain enrichment analysis results.
[0165] In some embodiments, determining the spatial regulator of each transcription factor based on the motif enrichment analysis results includes:
[0166] According to the results of motif enrichment analysis, a seventh subspace colocalization module is screened out, whose enrichment score is greater than the eighth preset threshold and whose enriched transcription factors are consistent with the transcription factors in the subspace colocalization module;
[0167] determining a fourth target gene in the seventh subspace colocalization module whose enrichment contribution is greater than a ninth preset threshold;
[0168] A spatial regulator of each transcription factor was constructed based on the second target transcription factor and the fourth target gene.
[0169] Wherein, each subspace colocalization module in the spatial colocalization module of each transcription factor is carried out TF motif enrichment analysis using the above-mentioned motif enrichment analysis function, after obtaining enrichment analysis result, the seventh subspace colocalization module with an enrichment score greater than the eighth preset threshold and the transcription factor of enrichment and the transcription factor consistent in the subspace colocalization module can be retained, for example, the seventh subspace colocalization module with an enrichment score greater than 3 and the transcription factor of enrichment and the transcription factor consistent in the subspace colocalization module can be retained. Then, the contribution value of each gene to the enrichment score in each seventh spatial colocalization module can be determined, and then the target gene whose enrichment contribution to the seventh spatial colocalization module is greater than the ninth preset threshold is determined as the fourth target gene, and these target genes are retained. Subsequently, the df2regulons function is called to integrate the module with TF with identical TF in the motif enrichment result to obtain final spatial regulator.
[0170] Referring to FIG. 8 , in some embodiments, step 105 includes but is not limited to the following steps 801 to 803 .
[0171] Step 801: Obtain a preset scoring tool, and score each spatial regulator based on the preset scoring tool to obtain a scoring result;
[0172] Step 802, determining the score distribution of each regulator in multiple cells based on the scoring results;
[0173] Step 803 , performing binary division on the multiple cells based on a binary Gaussian mixture model, and determining the activity of the transcription factor in each cell according to the binary division result.
[0174] After constructing the spatial regulator of each transcription factor, a preset scoring tool can be obtained to score each spatial regulator and obtain a scoring result. Among them, the AUCell module in the SCENIC program package can be used to calculate the AUC (area under cure, area under the curve) score of all spatial regulators in each cell (or each spot) according to the expression value. Specifically, all genes in each cell are sorted according to the expression value, and the expression values are randomly sorted. Then, a recovery curve is drawn according to the sorting position of the genes in the regulator, and the area under the curve of the top 5% sorting position is randomly calculated as the AUC score. Then, for each regulator, according to its AUC score distribution in all cells, all cells are binary divided using a binary Gaussian mixture model. The cell value with a high score is set to 1, and it is believed that the corresponding regulator is in an activated state, that is, the activity of the transcription factor is activated; otherwise, it is considered to be in a closed state, that is, the activity of the transcription factor is not activated. Thus, the inference of the activity of all transcription factors is completed.
[0175] The following is a detailed introduction to the transcription factor activity inference method provided by the present application using a specific example.
[0176] First, spatial transcriptome data for human liver cancer sections can be obtained, specifically as a text file in GEM format (stereo-seq technology). This data was obtained from a published article. In this example, the invasion margin region of interest in the article was selected for analysis. Based on the preliminary cell annotation results and coordinates, the region ±750 μm from the invasion margin was selected for analysis.
[0177] Next, the aforementioned gem-formatted text file was loaded. The x- and y-coordinate ranges were calculated and divided by 10, rounded down, to obtain the spatial grid size for calculating the spatial kernel density. A human gene regulatory network file from the NicheNet database was read and intersected with the gene names in the gem file to obtain all possible TF-target regulatory relationships for that slice. Genes with low capture were then filtered out (the default threshold was 100, meaning fewer than 100 spatial sites captured the gene). To avoid repeated computations, the igraph package was used to construct a regulatory network for the filtered TF-target data frame and convert it into an adjacency list format. The ks package was then called, with the gene capture location and spatial grid size as input. Kernel density estimates were calculated for each TF and all its targets. For each TF-target pair, the regions where the kernel density estimates exceeded the 50th percentile were intersected. The Pearson correlation coefficient between the two kernel density estimates in the overlapping region was calculated as the spatial correlation coefficient for the TF-target pair. The spatial correlation coefficient for each TF-target pair was finally output as a data frame.
[0178] As shown in FIG9 , it is a schematic diagram of the gem format text file obtained in this embodiment, wherein x and y are the horizontal and vertical coordinates respectively, MIDCounts is the number of molecules detected at that position, and geneID is the gene name of the RNA molecule aligned to the reference genome.
[0179] Figure 10 is a schematic diagram of the spatial correlation coefficient of the TF-target output in this embodiment. As shown in the figure, the spatial correlation coefficient (spatial_pcc) of the transcription factor KLF2 and multiple target genes is shown.
[0180] Furthermore, the modules_from_adjacencies function of the utils module in the pySCENIC package can be called to input the spatial correlation coefficient of each pair of TF-target and the expression matrix of the spatial transcriptome data bin100 to construct a gene module. The default parameters are thresholds = (0.75, 0.9), top_n_targets = (50), and top_n_regulators = (5, 10, 50). That is, a spatial co-localization module is constructed for each TF in the following way: the TF and the target gene with a spatial correlation greater than 75% and 90% quantiles are included in the module; the TF and the 50 target genes with the largest spatial correlation are included in the module; and if the spatial correlation between the TF and the target gene is in the top 5, 10, or 50 of the spatial correlation ranking of the target gene and all other TFs, the target gene is included in the module.
[0181] Figure 11 shows a schematic diagram of the spatial transcriptome gene expression matrix provided in this embodiment, wherein the rows are the IDs of each bin 100 (which can also be spots or cells), the columns are the gene names, and the values are the expression values after scanpy standardization.
[0182] We then performed TF motif enrichment analysis on the colocalized modules and constructed spatial regulators. We first downloaded the motif binding energy ranking database and the corresponding species motif annotation database from the RcisTarget database. We then called the FeatherRankingDatabase method in the rnkdb module of the ctxcore package to construct the ranking database object: dbs. We then called the prune2df function in the prune module of pySCENIC, taking the dbs object, the path to all the spatially colocalized modules, and the motif annotation database file. We then performed motif enrichment analysis on all modules and output the enrichment results. We then used df2regulons to organize the enrichment results and integrate the target genes from multiple modules of each TF into spatial regulators.
[0183] Figure 12 shows a schematic diagram of the enrichment results obtained by the enrichment analysis in this disclosure. Each row represents an enrichment result, and the columns are the enriched motif ID, AUC score, NES score, enrichment Q value and motif annotation information, target gene and its weight, and the ranking of genes with high enrichment contribution (taking the maximum value).
[0184] Figure 13 is a schematic diagram of the spatial regulator obtained by integration in this embodiment, and also shows the weight corresponding to each target gene.
[0185] After obtaining the predicted regulators, call pySCENIC's AUCell function, inputting the expression matrix and spatial regulators. The AUC score is then calculated based on the expression of the target genes contained in the spatial regulators in each bin (or spot or cell). The AUC matrix can then be binarized using the binarize function in pySCENIC's binarization module. This completes the inference of transcription factor activity for all bins.
[0186] Figure 14 is a schematic diagram of the AUC scores calculated in this example, where the rows are the IDs of each bin 100 (or spot or cell), the columns are the names of the spatial regulators, and the values are the AUC scores.
[0187] Figure 15 shows a schematic diagram of the binary matrix of AUC scores in this embodiment. The rows are the IDs of each bin 100 (or spot or cell), the columns are the names of the spatial regulators, and the values are 0 or 1, where 0 represents the inactive state of the spatial regulator in the corresponding bin 100, and 1 represents the active state.
[0188] Then, the transcription factor activity inference results, i.e., the AUC matrix, can be used for downstream exploration. Call scanpy's sc.pp.neighbor and sc.tl.leiden functions to perform Leiden clustering on all bins.
[0189] Figure 16 shows a schematic diagram of the clustering results based on AUC scores in this example. The figure shows the tsne dimensionality reduction clustering results, with the left side showing the cell2location annotation results for each bin 100 (based on the most abundant cell type as the annotation result); the right side shows the four transcription factor clusters obtained by clustering.
[0190] Figure 17 shows another schematic diagram of clustering results based on AUC scores in this example. The figure shows the cell annotation results and spatial in situ maps of four transcription factor clusters. This diagram demonstrates that the proposed approach also has a certain effect in downstream clustering tasks.
[0191] In addition, this embodiment also uses data from published articles to verify the effect of the scheme of this embodiment. Among them, a group of hepatocyte subtypes (Hep1 and Hep2) were found in the published articles, among which Hep1 highly expressed SAA1 and SAA2, which were related to tumor progression. Using the published data, the area of ±750um of the tumor invasion edge was extracted to perform transcription factor activity inference analysis using this technical scheme. The number and spatial distribution of Hep1 hepatocyte subtypes identified by the scheme of the present invention are close to those reported in the article, and the difference in Hep1 and Hep2 transcription factor activity consistent with the trend of the article can also be inferred. Among them, as shown in Figure 18, the spatial position of the hepatocyte subtypes identified by the AUC matrix of this application and the expression of their marker genes (SAA1 and SAA2). The spatial distribution and marker gene expression of hepatocyte subtype 3 shown in the figure are close to those reported in the article. As can be seen from the figure, the effect achieved by the scheme provided in this embodiment in the downstream tasks is basically consistent with that reported in the article.
[0192] The following describes the transcription factor activity inference device provided in the embodiments of the present application.
[0193] 19 , in some embodiments, the present application further provides a transcription factor activity inference device, the transcription factor activity inference device comprising:
[0194] An acquisition unit 1901 is configured to acquire spatial transcriptome data of a target slice, where the spatial transcriptome data includes location information and gene expression information of multiple categories of genes.
[0195] Construction unit 1902 is used to screen genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and construct a target regulatory network based on the screened genes, wherein the target regulatory network includes multiple first target regulatory data pairs, each of which includes a transcription factor and a corresponding target gene;
[0196] A calculation unit 1903 is configured to perform kernel density estimation on the transcription factor and target gene in each first target regulation data pair based on the position information and the gene expression information, and calculate a correlation score for each first target regulation data pair based on the kernel density estimation result;
[0197] a determination unit 1904, configured to construct a spatial co-localization module for each transcription factor based on the correlation score, and determine a spatial regulator for each transcription factor based on a motif enrichment analysis result of the spatial co-localization module;
[0198] The inference unit 1905 is used to score the spatial regulator of each transcription factor and determine the activity inference result of the corresponding transcription factor according to the scoring result.
[0199] In one embodiment, the construction unit comprises:
[0200] A first acquisition subunit is used to obtain a reference regulatory network corresponding to the target slice, where the reference regulatory network includes multiple reference regulatory edges and reference transcription factors and reference target genes connected to each reference regulatory edge;
[0201] An identification subunit, configured to identify the intersection of genes in the gene expression information, reference transcription factors, and reference target genes, to obtain a plurality of first transcription factors and a plurality of first target genes;
[0202] A first determining subunit is configured to determine, based on gene expression information, a plurality of second transcription factors and a plurality of second target genes whose corresponding capture sites are greater than a first preset threshold value among the plurality of first transcription factors and the plurality of first target genes;
[0203] The first construction subunit is used to construct a target regulatory network based on multiple second transcription factors and multiple second target genes.
[0204] Optionally, in some embodiments, the first construction subunit includes:
[0205] A search module, configured to search for at least one second transcription factor corresponding to each second target gene in a reference regulatory network;
[0206] A construction module is used to construct a target regulatory network based on multiple second target genes and at least one second transcription factor corresponding to each second target gene.
[0207] Optionally, in some embodiments, the computing unit includes:
[0208] The second determining subunit is configured to determine, for any first target regulation data pair, a position matrix of a corresponding third target gene and a corresponding third transcription factor according to the position information and the gene expression information;
[0209] The first calculation subunit is used to calculate the kernel density value of the third target gene and the third transcription factor in each spatial site according to the position matrix to obtain a kernel density estimation result;
[0210] The second calculation subunit is used to calculate the correlation score of each first target regulation data pair according to the kernel density estimation result.
[0211] Optionally, in some embodiments, the second computing subunit includes:
[0212] A first determination module is configured to determine, in the target slice, a first region where the nuclear density value of the third target gene is greater than a second preset threshold, and to determine a second region where the nuclear density value of the third transcription factor is greater than a third preset threshold;
[0213] an identification module, configured to identify a target area where the first area intersects the second area;
[0214] The calculation module is used to calculate the Pearson correlation coefficient between the nuclear density of the third target gene and the third transcription factor in the target region to obtain the correlation score of the first target regulation data pair.
[0215] Optionally, in some embodiments, the determining unit includes:
[0216] a screening subunit, configured to screen the plurality of first target regulation data pairs according to the correlation scores to obtain a plurality of second target regulation data pairs;
[0217] a second construction subunit, for constructing a spatial co-localization module for each transcription factor according to a plurality of second target regulation data pairs;
[0218] An analysis subunit, used to perform motif enrichment analysis on each spatial co-localization module to obtain motif enrichment analysis results;
[0219] The third determining subunit is used to determine the spatial regulator of each transcription factor based on the motif enrichment analysis results.
[0220] In some embodiments, the screening subunit comprises:
[0221] A first acquisition module, configured to acquire a fourth preset threshold for controlling the correlation score;
[0222] The second determining module is configured to determine, from among the plurality of first target regulation data pairs, a plurality of second target regulation data pairs having correlation scores greater than a fourth preset threshold.
[0223] Optionally, in some embodiments, the second construction subunit includes:
[0224] A first division module is used to divide the plurality of second target regulation data pairs into a plurality of first categories according to different transcription factors, wherein the second target regulation data pairs of each first category contain the same transcription factor;
[0225] A second division module is used to divide the plurality of second target regulation data pairs into a plurality of second categories according to different target genes, wherein the second target regulation data pairs of each second category include the same target gene;
[0226] a third determining module, configured to, for each transcription factor, determine, in each second target regulation data pair of the first category, a second target regulation data pair corresponding to a target gene having a correlation score with the transcription factor greater than a fifth preset threshold as a first subspace co-localization module;
[0227] The fourth determination module is used to determine, for each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the sixth preset threshold as the second subspace co-localization module
[0228] a fifth determination module, configured to, for each transcription factor, determine, in each second target regulation data pair of the first category, a second target regulation data pair corresponding to a target gene having a correlation score with the transcription factor greater than a seventh preset threshold as a third subspace co-localization module;
[0229] a sixth determination module, configured to, for each transcription factor, in each second category, determine a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in a preset first range with the target gene as a third regulatory data pair, and integrate all third regulatory data pairs found in the second category into a fourth subspace co-localization module;
[0230] a seventh determination module, configured to, for each transcription factor, in each second category, determine a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in a preset second range with the target gene as a fourth target regulatory data pair, and integrate all fourth target regulatory data pairs found in the second category into a fifth subspace co-localization module;
[0231] an eighth determination module, configured to, for each transcription factor, in each second category, determine a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in a preset third range with the target gene as a fifth target regulatory data pair, and integrate all fifth target regulatory data pairs found in the second category into a sixth subspace co-localization module;
[0232] The first construction module is used to construct a spatial colocalization module for each transcription factor based on the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module and the sixth subspace colocalization module corresponding to each transcription factor.
[0233] Optionally, in some embodiments, the analysis subunit includes:
[0234] The second acquisition module is used to obtain a preset motif enrichment analysis function;
[0235] The analysis module is used to perform motif enrichment analysis on each subspace colocalization module within each spatial colocalization module based on a preset motif enrichment analysis function to obtain a motif enrichment analysis result.
[0236] Optionally, in some embodiments, the third determining subunit includes:
[0237] a screening module, configured to screen, based on the motif enrichment analysis results, a seventh subspace colocalization module having an enrichment score greater than an eighth preset threshold and having enriched transcription factors consistent with the transcription factors in the subspace colocalization module;
[0238] a fifth determination module, configured to determine a fourth target gene in the seventh subspace co-localization module whose enrichment contribution is greater than a ninth preset threshold;
[0239] The second building module is used to construct a spatial regulator of each transcription factor based on the transcription factor and the fourth target gene in the seventh subspace colocalization module.
[0240] Optionally, in some embodiments, the inference unit includes:
[0241] The second acquisition subunit is used to obtain a preset scoring tool and score each spatial regulator based on the preset scoring tool to obtain a scoring result;
[0242] a fourth determining subunit, configured to determine the score distribution of each regulator in a plurality of cells based on the scoring results;
[0243] The fifth determining unit is configured to perform binary division on the plurality of cells based on a binary Gaussian mixture model, and determine the activity of the transcription factor in each cell according to the binary division result.
[0244] It can be seen that the contents of the above-mentioned transcription factor activity inference method embodiment are all applicable to the embodiment of the present transcription factor activity inference device. The functions specifically implemented by the present transcription factor activity inference device embodiment are the same as those in the above-mentioned transcription factor activity inference method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned transcription factor activity inference method embodiment.
[0245] 20 , which illustrates a hardware structure of a computer device according to another embodiment, the computer device includes:
[0246] The processor 2001 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0247] The memory 2002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 2002 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 2002 and are called by the processor 2001 to execute the transcription factor activity inference method or the transcription factor activity inference method of the embodiments of this application;
[0248] Input / output interface 2003, used to implement information input and output;
[0249] Communication interface 2004, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0250] Bus 2005 , which transmits information between various components of the device (e.g., processor 2001 , memory 2002 , input / output interface 2003 , and communication interface 2004 );
[0251] The processor 2001 , the memory 2002 , the input / output interface 2003 and the communication interface 2004 are connected to each other in communication within the device via the bus 2005 .
[0252] The present application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the above-mentioned transcription factor activity inference method or transcription factor activity inference method.
[0253] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.
[0254] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0255] It should be understood that in the description of the embodiments of the present application, multiple (or multiple items) means more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.
[0256] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0257] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0258] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0259] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0260] It should also be understood that the various implementation methods provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0261] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A method for inferring transcription factor activity, characterized in that: The method comprises: Acquiring spatial transcriptome data of a target slice, wherein the spatial transcriptome data includes location information and gene expression information of multiple categories of genes; Screening genes in the gene expression information according to a reference regulatory network corresponding to the target slice, and constructing a target regulatory network based on the screened genes, wherein the target regulatory network includes a plurality of first target regulatory data pairs, each first target regulatory data pair including a transcription factor and a corresponding target gene; performing kernel density estimation on the transcription factor and the target gene in each first target regulation data pair according to the position information and the gene expression information, and calculating a correlation score for each first target regulation data pair according to the kernel density estimation result; constructing a spatial colocalization module for each transcription factor according to the correlation score, and determining a spatial regulator of each transcription factor according to a motif enrichment analysis result of the spatial colocalization module; The spatial regulator of each transcription factor is scored, and the activity inference result of the corresponding transcription factor is determined based on the scoring result.
2. The method according to claim 1, characterized in that The step of screening genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and constructing a target regulatory network based on the screened genes, includes: Obtaining a reference regulatory network corresponding to the target slice, the reference regulatory network comprising a plurality of reference regulatory edges and a reference transcription factor and a reference target gene connected to each reference regulatory edge; Identifying the intersection of the genes in the gene expression information, the reference transcription factors, and the reference target genes to obtain a plurality of first transcription factors and a plurality of first target genes; Based on the gene expression information, determining a plurality of second transcription factors and a plurality of second target genes whose corresponding capture sites are greater than a first preset threshold value among the plurality of first transcription factors and the plurality of first target genes; A target regulatory network is constructed based on the multiple second transcription factors and the multiple second target genes.
3. The method according to claim 2, characterized in that The constructing of a target regulatory network according to the plurality of second transcription factors and the plurality of second target genes comprises: Searching for at least one second transcription factor corresponding to each second target gene in the reference regulatory network; A target regulatory network is constructed based on the multiple second target genes and at least one second transcription factor corresponding to each second target gene.
4. The method according to claim 1, wherein The performing kernel density estimation on the transcription factor and the target gene in each first target regulation data pair according to the position information and the gene expression information, and calculating the correlation score of each first target regulation data pair according to the kernel density estimation result, includes: For any of the first target regulation data pairs, determining a position matrix of a corresponding third target gene and a corresponding third transcription factor according to the position information and the gene expression information; Calculating the kernel density value of the third target gene and the third transcription factor in each spatial site according to the position matrix to obtain a kernel density estimation result; The correlation score of each first target regulation data pair is calculated based on the kernel density estimation result.
5. The method according to claim 4, characterized in that Calculating the correlation score of each first target regulation data pair according to the kernel density estimation result includes: Determining, in the target slice, a first region where the nuclear density value of the third target gene is greater than a second preset threshold, and determining a second region where the nuclear density value of the third transcription factor is greater than a third preset threshold; identifying a target area where the first area intersects the second area; The Pearson correlation coefficient between the nuclear density of the third target gene and the third transcription factor in the target region is calculated to obtain a correlation score of the first target regulation data pair.
6. The method according to claim 1, characterized in that The step of constructing a spatial co-localization module for each transcription factor according to the correlation score, and determining a spatial regulator of each transcription factor according to a motif enrichment analysis result of the spatial co-localization module, comprises: screening the plurality of first target regulation data pairs according to the correlation scores to obtain a plurality of second target regulation data pairs; constructing a spatial co-localization module for each transcription factor according to the plurality of second target regulatory data pairs; Perform motif enrichment analysis on each spatial co-localization module to obtain motif enrichment analysis results; The spatial regulators of each transcription factor are determined based on the motif enrichment analysis results.
7. The method according to claim 6, characterized in that The step of screening the plurality of first target regulation data pairs according to the correlation scores to obtain a plurality of second target regulation data pairs comprises: Obtaining a fourth preset threshold for controlling the correlation score; A plurality of second target regulation data pairs having correlation scores greater than the fourth preset threshold are determined among the plurality of first target regulation data pairs.
8. The method according to claim 6, characterized in that The step of constructing a spatial co-localization module for each transcription factor according to the plurality of second target regulation data pairs comprises: Dividing the plurality of second target regulation data pairs into a plurality of first categories according to different transcription factors, wherein the second target regulation data pairs in each first category include the same transcription factor; Dividing the plurality of second target regulation data pairs into a plurality of second categories according to different target genes, wherein the second target regulation data pairs of each second category include the same target gene; For each transcription factor, in each second target regulation data pair of the first category, determining the second target regulation data pair corresponding to the target gene having a correlation score with the transcription factor greater than a fifth preset threshold as a first subspace co-localization module; For each transcription factor, in each second target regulation data pair of the first category, determining the second target regulation data pair corresponding to the target gene having a correlation score with the transcription factor greater than a sixth preset threshold as a second subspace co-localization module; For each transcription factor, in each second target regulation data pair of the first category, determining the second target regulation data pair corresponding to the target gene having a correlation score with the transcription factor greater than a seventh preset threshold as a third subspace co-localization module; For each transcription factor, in each second category, a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked in the first preset range with the target gene is determined as a third regulatory data pair, and all third regulatory data pairs found in the second category are integrated into a fourth subspace co-localization module; For each transcription factor, in each second category, a second target regulatory data pair corresponding to the target transcription factor and having a correlation score ranked first in the preset second range with the target gene is determined as a fourth target regulatory data pair, and all fourth target regulatory data pairs found in the second category are integrated into a fifth subspace co-localization module; For each transcription factor, in each second category, a second target regulation data pair corresponding to the target transcription factor and having a correlation score ranked first in the preset third range with the target gene is determined as a fifth target regulation data pair, and all fifth target regulation data pairs found in the second category are integrated into a sixth subspace co-localization module; A spatial colocalization module for each transcription factor is constructed according to the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module and the sixth subspace colocalization module corresponding to each transcription factor.
9. The method according to claim 6, characterized in that The motif enrichment analysis is performed on each spatial co-localization module to obtain a motif enrichment analysis result, including: Get the preset motif enrichment analysis function; Based on the preset motif enrichment analysis function, a motif enrichment analysis is performed on each subspace colocalization module within each spatial colocalization module to obtain a motif enrichment analysis result.
10. The method according to claim 6, characterized in that Determining the spatial regulator of each transcription factor according to the motif enrichment analysis results includes: Screening, according to the motif enrichment analysis results, a seventh subspace colocalization module whose enrichment score is greater than an eighth preset threshold and whose enriched transcription factors are consistent with the transcription factors in the subspace colocalization module; determining a fourth target gene in each seventh subspace co-localization module whose enrichment contribution is greater than a ninth preset threshold; A spatial regulator of each transcription factor is constructed based on the transcription factor in the seventh subspace colocalization module and the fourth target gene.
11. The method according to claim 1, wherein Scoring the spatial regulator of each transcription factor and determining the activity inference result of the corresponding transcription factor based on the scoring result include: Obtaining a preset scoring tool, and scoring each spatial regulator based on the preset scoring tool to obtain a scoring result; Determining the score distribution of each regulator in multiple cells based on the scoring results; The multiple cells are binary-divided based on a binary Gaussian mixture model, and the activity of the transcription factor in each cell is determined according to the binary division result.
12. A transcription factor activity inference device, characterized in that: The device comprises: An acquisition unit, configured to acquire spatial transcriptome data of a target slice, wherein the spatial transcriptome data includes location information and gene expression information of multiple categories of genes; a construction unit, configured to screen genes in the gene expression information according to a reference regulatory network corresponding to the target slice, and construct a target regulatory network based on the screened genes, wherein the target regulatory network includes a plurality of first target regulatory data pairs, each first target regulatory data pair including a transcription factor and a corresponding target gene; a calculation unit, configured to perform kernel density estimation on the transcription factor and the target gene in each first target regulation data pair according to the position information and the gene expression information, and calculate a correlation score for each first target regulation data pair according to the kernel density estimation result; a determination unit, configured to construct a spatial colocalization module for each transcription factor according to the correlation score, and determine a spatial regulator of each transcription factor according to a motif enrichment analysis result of the spatial colocalization module; The inference unit is used to score the spatial regulator of each transcription factor and determine the activity inference result of the corresponding transcription factor based on the scoring result.
13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method for inferring transcription factor activity according to any one of claims 1 to 11 is implemented.
14. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for inferring transcription factor activity according to any one of claims 1 to 11 is implemented.
15. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the method for inferring transcription factor activity according to any one of claims 1 to 11.