Transcription factor activity inferring method, apparatus, storage medium, and computer device
By incorporating location and expression information from spatial transcriptome data into the transcription factor activity inference method, a target regulatory network is constructed, and nuclear density estimation and motif enrichment analysis are performed. This solves the problem of insufficient accuracy in existing methods and achieves more accurate transcription factor activity inference.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- STOMICS TECH CO LTD
- Filing Date
- 2024-03-07
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods for inferring transcription factor activity are inaccurate when based on spatial transcriptome sequencing data, mainly due to the lack of utilization of spatial location information.
By acquiring spatial transcriptome data of the target slices, a target regulatory network is constructed using location information and gene expression level information. The nuclear density estimation and correlation score of transcription factors and target genes are calculated. A spatial colocalization module is constructed and motif enrichment analysis is performed to finally determine the inferred activity results of transcription factors.
It improves the accuracy of transcription factor activity inference by combining spatial location information and expression level information.
Smart Images

Figure CN2024080575_15052026_PF_FP_ABST
Abstract
Description
Methods, apparatus, storage media and computer equipment for inferring transcription factor activity Technical Field
[0001] This disclosure relates to the field of gene sequencing technology, specifically to a method, apparatus, storage medium, and computer equipment for inferring transcription factor activity. Background Technology
[0002] Transcription factors (TFs), also known as sequence-specific DNA-binding factors, control the rate of transcription of genetic information from DNA to messenger RNA by binding to specific DNA sequences. Transcription factors ensure that genes are expressed in the correct amounts and at the correct times throughout the lifespan of cells and organisms; therefore, inferring transcription factor activity is an important task in bioinformatics analysis.
[0003] Currently, there are mature schemes for inferring transcription factor activity from single-cell transcriptome data, but the accuracy of inferring transcription factor activity from spatial transcriptome sequencing data based on these schemes is poor.
[0004] Summary of the Invention
[0005] The main objective of this application is to provide a method, apparatus, storage medium, and computer device for inferring transcription factor activity, which can improve the accuracy of transcription factor activity inference.
[0006] To achieve the above objectives, a first aspect of this application provides a method for inferring transcription factor activity, the method comprising:
[0007] Acquire spatial transcriptome data of the target slice, the spatial transcriptome data including location information and gene expression level information of multiple gene categories;
[0008] Genes in the gene expression information are screened according to the reference regulatory network corresponding to the target slice, and a target regulatory network is constructed based on the screened genes. The target regulatory network includes multiple first target regulatory data pairs, and each first target regulatory data pair contains a transcription factor and a corresponding target gene.
[0009] Based on the location information and the gene expression level information, the nuclear density of transcription factors and target genes in each first target regulation data pair is estimated, and the correlation score of each first target regulation data pair is calculated based on the nuclear density estimation results.
[0010] Based on the correlation scores, a spatial colocalization module for each transcription factor is constructed, and the spatial regulators of each transcription factor are determined based on the motif enrichment analysis results of the spatial colocalization modules.
[0011] The spatial regulators of each transcription factor are scored, and the activity inference of the corresponding transcription factor is determined based on the scoring results.
[0012] A second aspect of this application provides a transcription factor activity inference device, the device comprising:
[0013] The acquisition unit is used to acquire spatial transcriptome data of the target slice, which includes location information and gene expression level information of multiple gene categories.
[0014] The construction unit is used to screen genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and construct a target regulatory network according to the screened genes. The target regulatory network includes multiple first target regulatory data pairs, and each first target regulatory data pair contains a transcription factor and a corresponding target gene.
[0015] The calculation unit is used to estimate the nuclear density of transcription factors and target genes in each first target regulation data pair based on the location information and the gene expression level information, and to calculate the correlation score of each first target regulation data pair based on the nuclear density estimation results.
[0016] A determination unit is used to construct a spatial colocalization module for each transcription factor based on the correlation score, and to determine the spatial regulator of each transcription factor based on the motif enrichment analysis results of the spatial colocalization module.
[0017] The inference unit is used to score the spatial regulators of each transcription factor and determine the inferred activity of the corresponding transcription factor based on the scoring results.
[0018] In one embodiment, the building unit includes:
[0019] The first acquisition subunit is used to acquire the reference regulatory network corresponding to the target slice. The reference regulatory network includes multiple reference regulatory edges and reference transcription factors and reference target genes connected to each reference regulatory edge.
[0020] The identification subunit is used to identify the intersection of the gene in the gene expression information with the reference transcription factor and the reference target gene, so as to obtain multiple first transcription factors and multiple first target genes;
[0021] The first determining subunit is used to determine, based on the gene expression level information, a plurality of second transcription factors and a plurality of second target genes whose corresponding capture sites are greater than a first preset threshold among the plurality of first transcription factors and a plurality of first target genes;
[0022] The first construction subunit is used to construct a target regulatory network based on the plurality of second transcription factors and the plurality of second target genes.
[0023] Optionally, in some embodiments, the first construction subunit includes:
[0024] The search module is used to search for at least one second transcription factor corresponding to each second target gene in the reference regulatory network.
[0025] A construction module is used to construct a target regulatory network based on the plurality of second target genes and at least one second transcription factor corresponding to each second target gene.
[0026] Optionally, in some embodiments, the computing unit includes:
[0027] The second determining subunit is used to determine the position matrix of the corresponding third target gene and the corresponding third transcription factor for any first target regulation data pair based on the position information and the gene expression level information.
[0028] The first calculation subunit is used to calculate the nuclear density values of the third target gene and the third transcription factor at each spatial site based on the location matrix, and obtain the nuclear density estimation result;
[0029] The second calculation subunit is used to calculate the correlation score of each of the first target regulation data pairs based on the kernel density estimation results.
[0030] Optionally, in some embodiments, the second computing subunit includes:
[0031] The first determining module is used to determine a first region in the target slice where the nuclear density value of the third target gene is greater than a second preset threshold, and to determine a second region in the target slice where the nuclear density value of the third transcription factor is greater than a third preset threshold.
[0032] The identification module is used to identify the target area where the first area and the second area intersect;
[0033] The calculation module is used to calculate the Pearson correlation coefficient between the nuclear density of the third target gene and the third transcription factor in the target region, and to obtain the correlation score of the first target regulatory data pair.
[0034] Optionally, in some embodiments, the determining unit includes:
[0035] A filtering subunit is used to filter the plurality of first target regulation data pairs according to the correlation score to obtain a plurality of second target regulation data pairs;
[0036] The second construction subunit is used to construct a spatial colocalization module for each transcription factor based on the multiple second target regulatory data;
[0037] The analysis subunit is used to perform sequence enrichment analysis on each spatial colocation module to obtain the sequence enrichment analysis results;
[0038] The third determining subunit is used to determine the spatial regulator of each transcription factor based on the motif enrichment analysis results.
[0039] In some embodiments, the filtering subunit includes:
[0040] The first acquisition module is used to acquire a fourth preset threshold for controlling the correlation score;
[0041] The second determining module is used to determine, among the plurality of first target control data pairs, a plurality of second target control data pairs whose correlation scores are greater than the fourth preset threshold.
[0042] Optionally, in some embodiments, the second building subunit includes:
[0043] The first partitioning module is used to partition the plurality of second target regulatory data pairs into a plurality of first categories according to different transcription factors, wherein each first category of the second target regulatory data pairs contains the same transcription factor;
[0044] The second partitioning module is used to partition the multiple second target regulation data pairs into multiple second categories according to different target genes, and each second category of the second target regulation data pairs contains the same target gene.
[0045] The third determination module is used to determine, for each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the fifth preset threshold as the first subspace colocalization module.
[0046] The fourth determination module is used, for each transcription factor, in each first category of second target regulation data pair, to determine the second subspace co-localization module for the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than a sixth preset threshold.
[0047] The fifth determination module is used to determine, for each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the seventh preset threshold as the third subspace colocalization module.
[0048] The sixth determination module is used to determine, for each transcription factor, in each second category, a second target regulatory data pair that is ranked first in the first preset range of correlation scores with the target gene and corresponds to the target transcription factor as the third regulatory data pair, and to integrate all the third regulatory data pairs found in the second category into the fourth subspace colocalization module.
[0049] The seventh determination module is used to determine, for each transcription factor, in each second category, a second target regulatory data pair that ranks first in the correlation score with the target gene within a preset second range and corresponds to the target transcription factor as the fourth target regulatory data pair, and to integrate all the fourth target regulatory data pairs found in the second category into the fifth subspace colocalization module;
[0050] The eighth determination module is used to determine, for each transcription factor, in each second category, a second target regulatory data pair that ranks first in the correlation score with the target gene within a preset third range and corresponds to the target transcription factor as the fifth target regulatory data pair, and to integrate all the fifth target regulatory data pairs found in the second category into the sixth subspace colocalization module;
[0051] The first construction module is used to construct a spatial colocalization module for each transcription factor based on the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module, and the sixth subspace colocalization module corresponding to each transcription factor.
[0052] Optionally, in some embodiments, the analysis subunit includes:
[0053] The second acquisition module is used to acquire the preset motif enrichment analysis function;
[0054] The analysis module is used to perform sequence enrichment analysis on each subspace colocalization module within each spatial colocalization module based on the preset sequence enrichment analysis function, and obtain the sequence enrichment analysis results.
[0055] Optionally, in some embodiments, the third determining subunit includes:
[0056] The screening module is used to screen out the seventh subspace colocalization module based on the motif enrichment analysis results, where the enriched transcription factors have an enrichment score greater than the eighth preset threshold and are consistent with the transcription factors in the subspace colocalization module.
[0057] The fifth determination module is used to determine the fourth target gene in the seventh subspace colocalization module whose enrichment contribution is greater than the ninth preset threshold;
[0058] The second construction module is used to construct a spatial regulator for each transcription factor based on the transcription factors in the seventh sub-spatial colocalization module and the fourth target gene.
[0059] Optionally, in some embodiments, the inference unit includes:
[0060] The second acquisition subunit is used to acquire a preset scoring tool and score each spatial adjustment element based on the preset scoring tool to obtain a scoring result.
[0061] The fourth determining subunit is used to determine the score distribution of each regulator in multiple cells based on the scoring results;
[0062] The fifth determining unit is used to perform binary division of the multiple cells based on a binary Gaussian mixture model, and determine the activity of transcription factors in each cell based on the binary division results.
[0063] A third aspect of this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the transcription factor activity inference method described in the first aspect.
[0064] To achieve the above objectives, a fourth aspect of the present application provides a storage medium storing a computer program that, when executed by a processor, implements the transcription factor activity inference method described in the first aspect.
[0065] To achieve the above objectives, a fifth aspect of this application provides a computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the transcription factor activity inference method described in the first aspect.
[0066] The present application proposes a method, apparatus, storage medium, and computer device for inferring transcription factor activity. The method involves acquiring spatial transcriptome data of a target slice, which includes location information and gene expression level information for multiple gene categories. Genes in the gene expression level information are screened according to a reference regulatory network corresponding to the target slice. A target regulatory network is constructed based on the screened genes, comprising multiple first target regulatory data pairs, each containing a transcription factor and a corresponding target gene. The nuclear density of the transcription factor and target gene in each first target regulatory data pair is estimated based on the location information and gene expression level information. A correlation score is calculated for each first target regulatory data pair based on the nuclear density estimation results. A spatial co-localization module for each transcription factor is constructed based on the correlation score. The spatial regulator of each transcription factor is determined based on the motif enrichment analysis results of the spatial co-localization module. The spatial regulator of each transcription factor is scored, and the activity inference result of the corresponding transcription factor is determined based on the scoring results.
[0067] Therefore, this method uses spatial location information and gene expression levels from spatial transcriptome sequencing data to calculate the correlation between transcription factors and their corresponding target genes in the regulatory network, obtaining a more accurate spatial correlation. Then, based on this spatial correlation, further transcription factor activity inference is performed. Compared to related techniques that only calculate correlations and infer transcription factor activity based on gene expression levels, the transcription factor activity inference method provided in this disclosure can obtain more accurate transcription factor activity inference results. Attached Figure Description
[0068] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0069] Figure 1 is a flowchart of the transcription factor activity inference method provided in an embodiment of this application;
[0070] Figure 2 is a flowchart of step 102 in Figure 1;
[0071] Figure 3 is a schematic diagram of the constructed target regulation network;
[0072] Figure 4 is a flowchart of step 103 in Figure 1;
[0073] Figure 5 is a schematic diagram of the nuclear density estimation of transcription factors and target genes;
[0074] Figure 6 is a flowchart of step 104 in Figure 1;
[0075] Figures 7a-7f are schematic diagrams of the six subspace colocalization modules of transcription factors;
[0076] Figure 8 is a flowchart of step 105 in Figure 1;
[0077] Figure 9 is a schematic diagram of the gem format text file obtained in this embodiment;
[0078] Figure 10 is a schematic diagram of the spatial correlation coefficient of the TF-target output in this embodiment;
[0079] Figure 11 is a schematic diagram of the spatial transcriptome gene expression matrix provided in this embodiment;
[0080] Figure 12 is a schematic diagram of the enrichment results obtained by enrichment analysis in this disclosure;
[0081] Figure 13 is a schematic diagram of the integrated spatial regulator in this embodiment;
[0082] Figure 14 is a schematic diagram of the AUC score calculated in this embodiment;
[0083] Figure 15 is a schematic diagram of the binary matrix of AUC scores in this embodiment;
[0084] Figure 16 is a schematic diagram of a clustering result based on AUC score in this embodiment;
[0085] Figure 17 is another schematic diagram of clustering based on AUC scores in this embodiment;
[0086] Figure 18 shows the spatial location and marker gene expression of hepatocyte subtypes identified by the AUC matrix in this application.
[0087] Figure 19 is a structural diagram of the transcription factor activity inference device provided in the embodiments of this application;
[0088] Figure 20 is a schematic diagram of the hardware structure of the computer device provided in an embodiment of this application. Detailed Implementation
[0089] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0090] Before providing a further detailed description of the embodiments of this application, the nouns and terms used in the embodiments of this application are explained, and the nouns and terms used in the embodiments of this application shall be interpreted as follows:
[0091] Single-cell transcriptomics refers to the total expression level of all mRNAs in a single cell at a specific point in time, reflecting the overall characteristics of that cell. The advent of single-cell transcriptomics technology has enabled the precision of research from the multicellular level of tissues down to the single-cell level, allowing for the study of specific characteristics of a particular cell or cell group. It plays a crucial role, especially in cell development, tumor microenvironment, and single-cell atlas mapping.
[0092] Spatial gene expression (ST-seq) is a technique that combines transcriptomics, single-cell sequencing, and tissue sectioning. Based on single-cell transcriptomics, spatial transcriptomics can further obtain spatial distribution information of different cell types.
[0093] Regulatory network: Specifically, it refers to the gene regulatory network (GRN), a complex regulatory network involving the interactions between multiple target genes and multiple transcription factors. In a cell, a single transcription factor typically regulates the expression of multiple genes, while the expression of a single target gene is usually regulated by multiple transcription factors. Furthermore, a transcription factor itself can also act as a target gene, being regulated by the expression of other transcription factors.
[0094] Regulators: Also known as regulators, a transcription factor and all its target genes are called a regulator. The regulators in this disclosure remove false positive target genes from the target genes.
[0095] Spatial transcriptomics technology has developed rapidly in recent years, and its application in clinical medicine and scientific discovery is constantly deepening. The most significant feature of spatial transcriptomics data is that it reveals the spatial relationship between molecules and cells in biological tissues. This spatial dimension of information provides a completely new dimension for understanding the biological behavior at the cellular and molecular levels within tissues and organs.
[0096] However, current analytical methods for spatial transcriptome data still need improvement. For example, most analyses for inferring transcription factor (TF) activity still rely on tools developed based on traditional bulk RNA sequencing (Bulk-RNA seq) or single-cell transcriptome data, such as SCENIC (single-cell regulatory network inference and clustering) and decoupleR. These tools do not utilize the spatial location information in spatial transcriptome data. Furthermore, spatial transcriptome data, in addition to containing spatial location information, also exhibits greater data sparsity. This is because RNA capture in current spatial transcriptome sequencing protocols is affected by permeation conditions; cellular heterogeneity prevents all cells from achieving ideal permeation conditions and releasing all mRNA. Therefore, gene capture in spatial transcriptome data is often inferior to that of single-cell transcriptome data, exhibiting greater data sparsity. This means that the amount of gene expression information contained in spatial transcriptome data is lower than that in single-cell transcriptome data. If SCENIC, decoupleR, and other methods are still used for spatial transcriptome sequencing data, and TF activity is inferred solely based on expression information, the inaccuracy of the results will be greatly increased due to the lack of spatial information and insufficient expression information.
[0097] Therefore, there is an urgent need to develop tools for inferring transcription factor activity from spatial transcriptome data using its spatial location information. To address this, this disclosure provides a method for inferring transcription factor activity from spatial transcriptome data, aiming to improve the accuracy of such inferences.
[0098] The following describes the method for inferring transcription factor activity provided in the embodiments of this application.
[0099] Referring to Figure 1, in some embodiments, the transcription factor activity inference method provided in this application includes, but is not limited to, steps 101 to 105.
[0100] Step 101: Obtain spatial transcriptome data of the target slice;
[0101] Step 102: Select genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and construct the target regulatory network based on the selected genes;
[0102] Step 103: Based on the location information and gene expression level information, perform nuclear density estimation on the transcription factors and target genes in each first target regulation data pair, and calculate the correlation score of each first target regulation data pair based on the nuclear density estimation results.
[0103] Step 104: Construct spatial colocalization modules for each transcription factor based on correlation scores, and determine the spatial regulators of each transcription factor based on the motif enrichment analysis results of the spatial colocalization modules.
[0104] Step 105: Score the spatial regulators of each transcription factor and determine the activity inference result of the corresponding transcription factor based on the scoring results.
[0105] In steps 101 to 105 of this application embodiment, when inferring transcription factor activity in cells of the target slice, spatial transcriptome data of the target slice can be obtained first. Spatial transcriptome data is data obtained by sequencing the target slice using spatial transcription sequencing technology. It includes spatial location information of multiple sequencing sites (spot sites) and the quantity of messenger RNA captured at each spot site. Here, the mRNA quantity information can be referred to as gene expression level information. The target slice can be a biological slice of any organism, such as a slice of an animal organ or a plant slice. Specifically, the spatial transcriptome data can be a text or compressed text format gem file.
[0106] After obtaining the spatial transcriptome data of the target slice, the types of captured genes and the distribution of each type can be determined based on the gene expression information, resulting in a spatial location matrix for each gene type. The spatial location matrix contains the gene name and two-dimensional spatial coordinates of each gene. Then, a reference regulatory network corresponding to the target slice can be used to screen the captured genes in the spatial transcriptome data, removing genes not present in the reference regulatory network to obtain preliminarily screened genes. A preliminary target regulatory network is then constructed based on the relationship between these preliminarily screened genes and target genes in the reference regulatory network. The target regulatory network contains connections between multiple transcription factors and multiple target genes, as well as between transcription factors and target genes. One transcription factor can correspond to multiple target genes, and the same target gene can correspond to multiple transcription factors. A transcription factor and its corresponding target gene can form a regulatory data pair, referred to in this embodiment as the first target regulatory data pair. It is understood that the process of determining the preliminary target regulatory network in step 102 is similar to the process of constructing co-expression modules in the SCENIC scheme; both involve first screening out some low-quality genes to obtain co-expression modules of transcription factors and relatively high-quality genes.
[0107] Further, in step 103, location information and gene expression level information from the spatial transcriptome data are used to calculate the correlation score between transcription factors and target genes in the first target regulatory data pair. This step is the core difference between the transcription factor activity inference method provided in this disclosure and the SCENIC scheme. In the SCENIC scheme, the correlation weight between transcription factors and target genes is calculated only based on the gene expression level information of different genes. False positive data pairs are then eliminated based on this correlation weight to obtain the final rules. Finally, transcription factor activity is inferred based on the scores of the rules. In this embodiment, not only gene expression level information but also the spatial location information of genes is used to calculate the correlation score between transcription factors and target genes in the first target regulatory data pair. This score can therefore also be called the spatial correlation score. False positive data pairs are then further eliminated based on this spatial correlation score to obtain the final rules, which can also be called spatial rules. Finally, transcription factor activity is inferred based on the scores of the spatial rules.
[0108] Specifically, in this process, the location information and gene expression information in the spatial transcriptome data can be used to estimate the nuclear density of the transcription factors and target genes in the first target regulatory data pair, and then the spatial correlation score between the transcription factors and target genes can be calculated based on the nuclear density estimation results.
[0109] After calculating the spatial relevance score for each first-target regulatory data pair based on spatial location and gene expression information from the spatial transcriptome data, a spatial co-localization module for each transcription factor can be constructed based on the spatial relevance score. This process is similar to the construction of co-expression modules in the SCENIC protocol. Then, motif enrichment analysis can be performed on the constructed spatial co-localization modules, and the target genes in the modules can be further screened based on the motif enrichment analysis results to remove false-positive target genes. The spatial regulators of each transcription factor are then determined based on the screened transcription factors and their corresponding target genes. Finally, the scoring tool (AUCell) provided by SCENIC can be used to score each spatial regulator, and the activity inference of each transcription factor can be determined based on the scoring results.
[0110] The transcription factor activity inference method provided in this case is a method for inferring transcription factor activity based on spatial transcriptome data. This method adapts and improves the SCENIC protocol for single-cell transcriptomes. When calculating the correlation between transcription factors and target genes, it introduces the spatial location information and expression level information of each type of gene in the spatial transcriptome data, and calculates a more accurate spatial correlation between transcription factors and target genes, thereby improving the accuracy of transcription factor activity inference based on spatial transcriptome data.
[0111] Please refer to Figure 2. In some embodiments, step 102 includes, but is not limited to, steps 201 to 204.
[0112] Step 201: Obtain the reference regulatory network corresponding to the target slice. The reference regulatory network includes multiple reference regulatory edges and reference transcription factors and reference target genes connected to each reference regulatory edge.
[0113] Step 202: Identify the intersection of genes in the gene expression information with reference transcription factors and reference target genes to obtain multiple first transcription factors and multiple first target genes;
[0114] Step 203: Based on gene expression level information, identify multiple second transcription factors and multiple second target genes whose corresponding capture sites are greater than a first preset threshold among multiple first transcription factors and multiple first target genes;
[0115] Step 204: Construct a target regulatory network based on multiple second transcription factors and multiple second target genes.
[0116] In this embodiment, the gene screening process in the acquired spatial transcriptome data can be implemented in two steps: a first step based on a reference regulatory network, and a second step based on the number of gene capture sites. This yields genes that are widely distributed and relatively common in the tissue sections.
[0117] Specifically, a reference regulatory network corresponding to the target slice can be obtained first. This reference regulatory network can be a public database of TF-target regulatory networks (such as the NicheNet database). After obtaining the reference regulatory network, the intersection of the gene information provided in the reference regulatory network and the information in the spatial transcriptome data can be taken. This intersection process can remove some uncommon genes in the spatial transcriptome data, thereby obtaining multiple first transcription factors and multiple first target genes. Further, based on the number of capture sites captured by each first transcription factor and each first target gene in the gene expression information, genes with low or overly concentrated expression in the first transcription factors and first target genes can be removed, thereby obtaining multiple second transcription factors and multiple second target genes. For example, a first preset threshold of 100 can be designed; if the number of capture sites corresponding to a target gene is less than 100, it is removed. Then, a preliminary target regulatory network can be constructed based on the multiple second transcription factors and multiple second target genes.
[0118] In some embodiments, a target regulatory network is constructed based on multiple second transcription factors and multiple second target genes, including:
[0119] Search the reference regulatory network for at least one second transcription factor corresponding to each second target gene;
[0120] A target regulatory network is constructed based on multiple second target genes and at least one second transcription factor corresponding to each second target gene.
[0121] After preprocessing the genes in the spatial transcriptome data (i.e., screening to obtain multiple second target genes), the transcription factor corresponding to each second target gene can be further searched in the reference regulatory network; these can be referred to here as second transcription factors. Each second target gene may have one or more corresponding transcription factors. Then, a target regulatory network can be constructed based on the multiple second target genes and at least one corresponding second transcription factor for each second target gene. Figure 3 shows a schematic diagram of the constructed target regulatory network. The figure shows a target regulatory network consisting of two transcription factors and three target genes. The number of transcription factors and target genes is just an example; generally, the regulatory network can contain more transcription factors and target genes. As shown in the figure, the target regulatory network contains multiple TF-target (target gene) data pairs, which can be referred to as the first target regulatory data pair in this embodiment.
[0122] Please refer to Figure 4. In some embodiments, step 103 includes, but is not limited to, steps 401 to 403.
[0123] Step 401: For any first target regulation data pair, determine the position matrix of the corresponding third target gene and the corresponding third transcription factor based on the position information and gene expression level information.
[0124] Step 402: Calculate the nuclear density values of the third target gene and the third transcription factor at each spatial site based on the location matrix to obtain the nuclear density estimation results;
[0125] Step 403: Calculate the correlation score for each first target regulation data pair based on the kernel density estimation results.
[0126] In this embodiment of the disclosure, after constructing a preliminary target regulation network, the correlation between transcription factors and target genes in the TF-target data pairs in the target regulation network can be further calculated to obtain the correlation score corresponding to each first target regulation data pair.
[0127] Specifically, the nuclear density of the third target gene and the third transcription factor can be estimated based on the spatial location matrix of each target gene (referred to as the third target gene here to distinguish it from the aforementioned target genes) and transcription factor (referred to as the third transcription factor here to distinguish it from the target genes shown elsewhere in this paper). Specifically, for any third transcription factor and its corresponding third target gene, the location matrix of each third target gene and the size of the entire transcriptome slice (defaulting to the range of x, y coordinates in the gem file divided by 10) can be input to calculate the density estimate of the third target gene at each spot site. Figure 5 shows a schematic diagram of the nuclear density estimation of transcription factors and target genes. In this figure, the horizontal and vertical axes represent the x, y coordinates in the location matrix.
[0128] After calculating the nuclear density estimates for each third target gene and each third transcription factor, the correlation score for each first target regulatory data pair can be calculated based on the nuclear density estimates for each third target gene and each third transcription factor.
[0129] In some embodiments, the correlation score for each first target regulation data pair is calculated based on the kernel density estimation results, including:
[0130] In the target slice, a first region where the nuclear density value of the third target gene is greater than a second preset threshold is identified, and a second region where the nuclear density value of the third transcription factor is greater than a third preset threshold is identified.
[0131] Identify the target area where the first and second regions intersect;
[0132] The Pearson correlation coefficient between the nuclear density of the third target gene and the first target transcription factor in the target region is calculated to obtain the correlation score of the first target regulatory data pair.
[0133] In the embodiments of this application, for each target regulatory data pair, its correlation score can be calculated by calculating the Pearson correlation coefficient (PCC) of the nuclear density estimates of the transcription factors and target genes within the intersection region of their nuclear density key regions.
[0134] Specifically, for any target regulatory data pair, regions with nuclear density estimates greater than the N% quantile can be selected. For example, if N is 50, this region corresponds to the spatial sites with nuclear density values ranking in the top 50%. Thus, a second preset threshold for dividing the nuclear density values into the top and bottom 50% can be calculated based on the nuclear density of the third target gene at all sites in the target slice, and a third preset threshold can be calculated based on the nuclear density of the first target transcription factor at all sites in the target slice. Then, a first region in the target slice where the nuclear density value of the third target gene is greater than the second preset threshold, and a second region where the nuclear density value of the first target transcription factor is greater than the third preset threshold, are identified. The intersection of the first and second regions is then determined, identifying the target region where the first and second regions intersect. Furthermore, the Pearson correlation coefficient between the nuclear densities of the third target gene and the third transcription factor in the target region can be calculated, thus yielding the correlation score for the first target regulatory data pair.
[0135] Then, for each first target control data pair, the correlation score corresponding to each first target control data pair is calculated using the above method.
[0136] Please refer to Figure 6. In some embodiments, step 104 includes, but is not limited to, steps 601 to 604 below.
[0137] Step 601: Based on the correlation score, multiple first-target control data pairs are filtered to obtain multiple second-target control data pairs;
[0138] Step 602: Construct a spatial colocalization module for each transcription factor based on multiple second-target regulatory data;
[0139] Step 603: Perform sequence enrichment analysis on each spatial colocalization module to obtain the sequence enrichment analysis results;
[0140] Step 604: Determine the spatial regulators of each transcription factor based on the motif enrichment analysis results.
[0141] In this embodiment, after calculating the correlation score of each first target regulatory data pair, these first target regulatory data pairs can be screened based on the correlation score to remove those with lower correlation scores. For example, first target regulatory data pairs with a correlation score less than 0.03 can be removed, while second target regulatory data pairs with a correlation score greater than or equal to 0.03 can be retained. Then, a spatial colocalization module for each transcription factor can be constructed based on the multiple second target regulatory data pairs obtained from the screening. Each transcription factor's spatial colocalization module consists of six sub-spatial colocalization modules, which can be constructed using different strategies. Figures 7a-7f illustrate the spatial colocalization modules of transcription factors.
[0142] After constructing the spatial colocalization module corresponding to each transcription factor, motif enrichment analysis can be performed on each spatial colocalization module to obtain the motif enrichment analysis results. Then, the spatial regulator of each transcription factor can be determined based on the motif enrichment analysis results. The structure of the spatial regulator is similar to the sub-spatial colocalization modules in Figures 7a-7f, except that some false positive target genes in the spatial colocalization module are removed and duplicates are eliminated based on the motif enrichment analysis results to obtain the final spatial regulator.
[0143] In some embodiments, multiple first-target regulation data pairs are filtered based on correlation scores to obtain multiple second-target regulation data pairs, including:
[0144] Obtain a fourth preset threshold for controlling the relevance score;
[0145] Among multiple first-target control data pairs, identify multiple second-target control data pairs whose correlation scores are greater than a fourth preset threshold.
[0146] When filtering multiple first-target control data pairs based on correlation scores, a fourth preset threshold for controlling the correlation score can be obtained first, such as the aforementioned 0.03 or other values. Then, based on this fourth preset threshold, multiple second-target control data pairs with correlation scores greater than the fourth preset threshold are identified from the multiple first-target control data pairs.
[0147] In some embodiments, constructing a spatial colocalization module for each transcription factor based on multiple second-target regulatory data includes:
[0148] Multiple second-target regulatory data pairs are divided into multiple first categories according to different transcription factors, and each first category of second-target regulatory data pairs contains the same transcription factor.
[0149] Multiple second-target regulatory data pairs are divided into multiple second categories according to different target genes, and each second category of second-target regulatory data pairs contains the same target gene;
[0150] For each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the fifth preset threshold is identified as the first subspace colocalization module;
[0151] For each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the sixth preset threshold is identified as the second subspace colocalization module.
[0152] For each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the seventh preset threshold is identified as the third subspace colocalization module.
[0153] For each transcription factor, in each second category, a second target regulatory data pair that ranks first in the correlation score with the target gene and corresponds to the target transcription factor is identified as the third regulatory data pair. All third regulatory data pairs found in the second category are integrated into a fourth subspace colocalization module.
[0154] For each transcription factor, in each second category, the second target regulatory data pair that ranks first in the correlation score with the target gene and corresponds to the target transcription factor is identified as the fourth target regulatory data pair. All fourth target regulatory data pairs found in the second category are integrated into the fifth subspace colocalization module.
[0155] For each transcription factor, in each second category, the second target regulatory data pair that ranks first in the correlation score with the target gene and corresponds to the target transcription factor is identified as the fifth target regulatory data pair. All fifth target regulatory data pairs found in the second category are integrated into the sixth subspace colocalization module.
[0156] The spatial colocalization module for each transcription factor is constructed based on the first, second, third, fourth, fifth, and sixth subspace colocalization modules corresponding to each transcription factor.
[0157] Among them, when screening to obtain multiple second-target regulatory data pairs and constructing a spatial colocalization module for each transcription factor based on multiple second-target regulatory data pairs, six different methods can be used to construct six sub-spatial colocalization modules for each transcription factor, and then the six sub-spatial colocalization modules together form a spatial colocalization module.
[0158] Specifically, multiple second-target regulatory data pairs can first be divided into multiple first categories based on different transcription factors, where each second-target regulatory data pair in each first category contains the same transcription factor. Similarly, multiple second-target regulatory data pairs can be divided into multiple second categories based on different target genes, where each second-target regulatory data pair in each second category contains the same target gene. Then, for the second-target regulatory data pairs in the first category, the corresponding first subspace co-localization modules with correlation scores greater than a fifth preset threshold, second subspace co-localization modules with correlation scores greater than a sixth preset threshold, and third subspace co-localization modules with correlation scores greater than a seventh preset threshold can be identified; for example, the fifth preset threshold can be set to 75%, the sixth preset threshold can be set to 90%, and the seventh preset threshold can be set to the values corresponding to the top 50 ranges of correlation scores in the second-target regulatory data pairs of the first category.
[0159] Furthermore, within each second category of second-target regulatory data pairs, the correlation scores between different transcription factors and their corresponding target genes can be ranked. Then, the second-target regulatory data pair corresponding to the target transcription factor is determined from the top of the correlation scores of the multiple second-category second-target regulatory data pairs corresponding to the target gene within a preset first range, a preset second range, and a preset third range. For example, the preset first range is ranked in the top 5, the preset second range is ranked in the top 10, and the preset third range is ranked in the top 50. The second-target regulatory data pairs corresponding to these three preset ranges are respectively called the third, fourth, and fifth-target regulatory data pairs. Then, the third, fourth, and fifth-target regulatory data pairs are found in all second categories. All third-target regulatory data pairs have the same target transcription factor but different target genes. The third, fourth, and fifth-target regulatory data pairs are integrated into the fourth, fifth, and sixth subspace co-localization modules, respectively.
[0160] Furthermore, the aforementioned first subspace colocalization module, second subspace colocalization module, third subspace colocalization module, fourth subspace colocalization module, fifth subspace colocalization module, and sixth subspace colocalization module together constitute the final spatial colocalization module of each transcription factor.
[0161] In some embodiments, a sequence enrichment analysis is performed on each subspace colocalization module within each spatial colocalization module to obtain the sequence enrichment analysis results, including:
[0162] Obtain the preset motif enrichment analysis function;
[0163] Based on the preset sequence enrichment analysis function, sequence enrichment analysis is performed on each subspace colocalization module in each spatial colocalization module to obtain the sequence enrichment analysis results.
[0164] After constructing the spatial colocalization module for each transcription factor, a motif enrichment analysis function can be obtained, specifically the `prune2df` function from the `prune` module in the SCENIC package. Then, the enrichment analysis function is used to perform motif enrichment analysis on each sub-spatial colocalization module within each spatial colocalization module, yielding the motif enrichment analysis results. Specifically, the RscisTarget algorithm can be used to calculate the TF motif enrichment analysis and obtain the enrichment analysis results.
[0165] In some embodiments, spatial regulators of each transcription factor are determined based on motif enrichment analysis results, including:
[0166] Based on the motif enrichment analysis results, the seventh subspace colocalization module was selected, whose enrichment score was greater than the eighth preset threshold and whose enriched transcription factors were consistent with the transcription factors in the subspace colocalization module.
[0167] Identify the fourth target gene in the seventh subspace colocalization module whose enrichment contribution is greater than the ninth preset threshold;
[0168] Spatial regulators for each transcription factor were constructed based on the second target transcription factor and the fourth target gene.
[0169] Specifically, after performing TF motif enrichment analysis on each sub-spatial co-localization module within the spatial co-localization module of each transcription factor using the aforementioned motif enrichment analysis function, and obtaining the enrichment analysis results, the seventh sub-spatial co-localization module with an enrichment score greater than the eighth preset threshold and whose enriched transcription factors are consistent with those in the sub-spatial co-localization module can be retained. For example, the seventh sub-spatial co-localization module with an enrichment score greater than 3 and whose enriched transcription factors are consistent with those in the sub-spatial co-localization module can be retained. Then, the contribution value of each gene in each seventh spatial co-localization module to the enrichment score can be determined, and the target genes whose enrichment contribution to the seventh spatial co-localization module is greater than the ninth preset threshold are identified as the fourth target genes and retained. Subsequently, the df2regulons function is called to integrate the modules with the same TF in the motif enrichment results with the TF to obtain the final spatial regulator.
[0170] Please refer to Figure 8. In some embodiments, step 105 includes, but is not limited to, steps 801 to 803 below.
[0171] Step 801: Obtain a preset scoring tool, and score each spatial adjuster based on the preset scoring tool to obtain the scoring result;
[0172] Step 802: Determine the score distribution of each regulator in multiple cells based on the scoring results;
[0173] Step 803: Divide multiple cells into binary groups based on a binary Gaussian mixture model, and determine the activity of transcription factors in each cell based on the binary division results.
[0174] After constructing the spatial regulators of each transcription factor, a pre-defined scoring tool can be used to score each spatial regulator and obtain the scoring results. Specifically, the AUCell module in the SCENIC package can be used to calculate the AUC (area under cure) score of all spatial regulators in each cell (or each spot) based on their expression values. Specifically, all genes in each cell are sorted according to their expression values, with those having the same expression value randomly sorted. Then, recovery curves are plotted based on the sorted positions of genes within the regulators, and the area under the curve of the top 5% of sorted positions is randomly calculated as the AUC score. Then, for each regulator, based on its AUC score distribution across all cells, a binary Gaussian mixture model is used to binary divide all cells. Cells with high scores are assigned a value of 1, indicating that the corresponding regulator is in an activated state, i.e., the transcription factor activity is activated; otherwise, it is considered to be in a deactivated state, i.e., the transcription factor activity is not activated. This completes the inference of the activity of all transcription factors.
[0175] The following will provide a detailed description of the transcription factor activity inference method provided in this application using a specific embodiment.
[0176] First, spatial transcriptome data of human liver cancer slices can be obtained, specifically in GEM format text files (obtained using stereo-seq technology). This data was obtained from published articles. In this embodiment, the invasive border region of interest from the article was selected for analysis. Based on preliminary cell annotation results and coordinates, a region within ±750 μm of the invasive border was selected for analysis.
[0177] Then, the aforementioned gem-formatted text file is read in, and the ranges of the x and y coordinates are calculated and divided by 10, rounded down to obtain the spatial grid size for calculating the spatial kernel density. The human gene regulatory network file from the NicheNet database is read, and its intersection with the gene names in the gem file is taken to obtain all possible TF-target regulatory relationships for the slice. Genes with low capture rates are then filtered out (the default threshold is 100, meaning fewer than 100 spatial sites of the gene are captured). To avoid redundant calculations, a regulatory network is constructed using the igraph package for the filtered TF-target data frame and converted to an adjacency list format. The ks package is called, inputting the gene capture location and spatial grid size, to calculate the kernel density estimate for each TF and all its targets. For each pair of TF-targets, regions with kernel density estimates greater than the 50th quantile are selected and their intersections are taken. Finally, the Pearson correlation coefficient of the kernel density estimates in the overlapping region is calculated as the spatial correlation coefficient for this pair of TF-targets. Finally, the spatial correlation coefficient of each pair of TF-targets is output in data frame format.
[0178] Figure 9 shows a schematic diagram of the gem format text file obtained in this embodiment, where x and y are the horizontal and vertical axes, respectively, MIDCounts is the number of molecules detected at this position, and geneID is the gene name of the RNA molecule that is aligned to the reference genome.
[0179] Figure 10 shows a schematic diagram of the spatial correlation coefficients of TF-targets output in this embodiment. The figure illustrates the spatial correlation coefficients (spatial_pcc) of transcription factor KLF2 and multiple target genes.
[0180] Furthermore, the `modules_from_adjacencies` function of the `utils` module in the `pySCENIC` package can be called. Inputting the spatial correlation coefficient of each TF-target pair and the expression matrix of the spatial transcriptome data bin100, gene modules can be constructed. The default parameters are `thresholds = (0.75, 0.9)`, `top_n_targets = (50)`, and `top_n_regulators = (5, 10, 50)`. That is, for each TF, a spatial co-localization module is constructed in the following ways: the TF and target genes with spatial correlation greater than 75% and 90% quantiles are included in the module; the TF and the 50 target genes with the highest spatial correlation are included in the module; and if the spatial correlation between the TF and the target gene is among the top 5, 10, or 50 of the spatial correlation ranking between the target gene and all other TFs, then the target gene is included in the module.
[0181] Figure 11 shows a schematic diagram of the spatial transcriptome gene expression matrix provided in this embodiment. The rows represent the ID of each bin100 (which can also be a spot point or cell), the columns represent gene names, and the values are the expression values after Scanpy normalization.
[0182] Subsequently, TF motif enrichment analysis was performed on the co-localization modules, and spatial regulators were constructed. First, the motif binding energy ranking database and the corresponding species' motif annotation database were downloaded from the RcisTarget database. The FeatherRankingDatabase method in the rnkdb module of the ctxcore package was used to construct a 'ranking database object': dbs. Then, the prune2df function in the prune module of pySCENIC was called, taking the dbs object, the paths to all spatial co-localization modules and the motif annotation database files as input, to perform motif enrichment analysis on all modules and output the enrichment results. Next, df2regulons was used to organize the enrichment results, integrating the target genes from multiple modules of each TF into spatial regulators.
[0183] Figure 12 shows a schematic diagram of the enrichment results obtained from the enrichment analysis in this disclosure. Each row represents an enrichment result, and the columns are the enriched motif ID, AUC score, NES score, enriched Q value, motif annotation information, target gene and its weight, and the ranking of genes with high enrichment contribution (taking the maximum value).
[0184] Figure 13 shows a schematic diagram of the spatial regulators integrated in this embodiment. The figure also shows the weight of each target gene.
[0185] After obtaining the predicted regulators, the AUCell function of pySCENIC is called, taking the expression matrix and spatial regulators as input. Then, the AUC score is calculated based on the expression of the target genes contained in each bin100 (or spot point or cell). The AUC matrix can then be binarized using the binarize function in the binarization module of pySCENIC. This completes the inference of transcription factor activity for all bins.
[0186] Figure 14 shows a schematic diagram of the AUC score calculated in this embodiment. The rows represent the ID of each bin100 (or spot point or cell), the columns represent the names of spatial regulators, and the values represent the AUC score.
[0187] Figure 15 shows a schematic diagram of the binary matrix of AUC scores in this embodiment. The rows represent the IDs of each bin100 (or spot point or cell), and the columns represent the names of spatial regulators. Values are 0 or 1, where 0 indicates the spatial regulator is off in the corresponding bin100, and 1 indicates it is active.
[0188] Then, the results of transcription factor activity inference, i.e., downstream exploration using the AUC matrix, can be performed. The `sc.pp.neighbor` and `sc.tl.leiden` functions of `scanpy` are called to perform Leiden clustering on all bins.
[0189] Figure 16 shows a schematic diagram of the clustering results based on AUC scores in this embodiment. As shown in the figure, the tsne dimensionality reduction clustering results are displayed, where the left side represents the annotation results of cell2location for each bin100 (with the cell type with the highest abundance as the annotation result); the right side shows the four types of transcription factor clusters obtained by clustering.
[0190] Figure 17 shows another schematic diagram of the clustering results based on AUC scores in this embodiment. The figure illustrates the cell annotation results and the spatial in-situ maps of the four transcription factor clusters. The diagram demonstrates that the proposed method also exhibits effectiveness in downstream clustering tasks.
[0191] Furthermore, this embodiment also utilizes data from published articles to verify the effectiveness of the proposed solution. Among these, published articles identified a group of hepatocyte subtypes (Hep1 and Hep2), with Hep1 highly expressing SAA1 and SAA2, which are associated with tumor progression. Using published data, a region of ±750 μm from the tumor invasion margin was extracted for transcription factor activity inference analysis using this technical solution. The number and spatial distribution of the Hep1 hepatocyte subtype identified by the proposed solution are close to those reported in the articles, and the differences in transcription factor activity between Hep1 and Hep2 can also be inferred in accordance with the trends observed in the articles. Figure 18 shows the spatial location of the hepatocyte subtypes identified by the AUC matrix of this application and the expression of their marker genes (SAA1 and SAA2). The spatial distribution and marker gene expression of hepatocyte subtype 3 shown in the figure are close to those reported in the articles. As can be seen from the figure, the solution provided in this embodiment achieves results in downstream tasks that are essentially consistent with those reported in the articles.
[0192] The transcription factor activity inference device provided in the embodiments of this application will be described below.
[0193] Referring to Figure 19, in some embodiments, this application also provides a transcription factor activity inference device, which includes:
[0194] Acquisition unit 1901 is used to acquire spatial transcriptome data of the target slice. The spatial transcriptome data includes location information and gene expression level information of multiple gene categories.
[0195] Construction unit 1902 is used to screen genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and construct a target regulatory network based on the screened genes. The target regulatory network includes multiple first target regulatory data pairs, and each first target regulatory data pair contains a transcription factor and a corresponding target gene.
[0196] The computing unit 1903 is used to estimate the nuclear density of transcription factors and target genes in each first target regulation data pair based on location information and gene expression level information, and to calculate the correlation score of each first target regulation data pair based on the nuclear density estimation results.
[0197] Unit 1904 is defined to construct spatial colocalization modules for each transcription factor based on correlation scores, and to determine the spatial regulators of each transcription factor based on the motif enrichment analysis results of the spatial colocalization modules.
[0198] Inference unit 1905 is used to score the spatial regulators of each transcription factor and determine the inferred activity of the corresponding transcription factor based on the scoring results.
[0199] In one embodiment, the building unit includes:
[0200] The first acquisition subunit is used to acquire the reference regulatory network corresponding to the target slice. The reference regulatory network includes multiple reference regulatory edges and reference transcription factors and reference target genes connected to each reference regulatory edge.
[0201] The identification subunit is used to identify the intersection of genes in gene expression information with reference transcription factors and reference target genes, thereby obtaining multiple first transcription factors and multiple first target genes.
[0202] The first determining subunit is used to determine, based on gene expression level information, multiple second transcription factors and multiple second target genes whose corresponding capture sites are greater than a first preset threshold among multiple first transcription factors and multiple first target genes;
[0203] The first building unit is used to construct a target regulatory network based on multiple second transcription factors and multiple second target genes.
[0204] Optionally, in some embodiments, the first building subunit includes:
[0205] The search module is used to search for at least one second transcription factor corresponding to each second target gene in the reference regulatory network;
[0206] The building module is used to construct a target regulatory network based on multiple second target genes and at least one second transcription factor corresponding to each second target gene.
[0207] Optionally, in some embodiments, the computing unit includes:
[0208] The second determining subunit is used to determine the position matrix of the corresponding third target gene and the corresponding third transcription factor for any first target regulation data pair based on position information and gene expression level information.
[0209] The first computational subunit is used to calculate the nuclear density values of the third target gene and the third transcription factor at each spatial site based on the position matrix, and obtain the nuclear density estimation results;
[0210] The second computational subunit is used to calculate the correlation score of each first target regulation data pair based on the kernel density estimation results.
[0211] Optionally, in some embodiments, the second computing subunit includes:
[0212] The first determining module is used to determine a first region in the target slice where the nuclear density value of the third target gene is greater than a second preset threshold, and to determine a second region in the target slice where the nuclear density value of the third transcription factor is greater than a third preset threshold.
[0213] The identification module is used to identify the target area where the first region and the second region intersect;
[0214] The calculation module is used to calculate the Pearson correlation coefficient between the nuclear density of the third target gene and the third transcription factor in the target region, and obtain the correlation score of the first target regulatory data pair.
[0215] Optionally, in some embodiments, the determining unit includes:
[0216] The filtering subunit is used to filter multiple first-target control data pairs based on correlation scores to obtain multiple second-target control data pairs;
[0217] The second construction subunit is used to construct the spatial colocalization module of each transcription factor based on multiple second target regulatory data;
[0218] The analysis subunit is used to perform sequence enrichment analysis on each spatial colocation module to obtain the sequence enrichment analysis results;
[0219] The third determining subunit is used to identify the spatial regulators of each transcription factor based on the results of motif enrichment analysis.
[0220] In some embodiments, the filtering subunit includes:
[0221] The first acquisition module is used to acquire the fourth preset threshold for controlling the correlation score;
[0222] The second determining module is used to determine multiple second target control data pairs among multiple first target control data pairs whose correlation scores are greater than a fourth preset threshold.
[0223] Optionally, in some embodiments, the second building subunit includes:
[0224] The first partitioning module is used to divide multiple second target regulation data pairs into multiple first categories according to different transcription factors, and each first category of second target regulation data pairs contains the same transcription factor.
[0225] The second partitioning module is used to divide multiple second target regulation data pairs into multiple second categories according to different target genes, and each second category of second target regulation data pairs contains the same target gene.
[0226] The third determination module is used to determine, for each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the fifth preset threshold as the first subspace colocalization module.
[0227] The fourth determination module is used, for each transcription factor, in each first category of second target regulation data pair, to determine the second subspace co-localization module for the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than a sixth preset threshold.
[0228] The fifth determination module is used to determine, for each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the seventh preset threshold as the third subspace colocalization module.
[0229] The sixth determination module is used to determine, for each transcription factor, in each second category, a second target regulatory data pair that is ranked first in the first preset range of correlation scores with the target gene and corresponds to the target transcription factor as the third regulatory data pair, and to integrate all the third regulatory data pairs found in the second category into the fourth subspace colocalization module.
[0230] The seventh determination module is used to determine, for each transcription factor, in each second category, a second target regulatory data pair that ranks first in the correlation score with the target gene within a preset second range and corresponds to the target transcription factor as the fourth target regulatory data pair, and to integrate all the fourth target regulatory data pairs found in the second category into the fifth subspace colocalization module;
[0231] The eighth determination module is used to determine, for each transcription factor, in each second category, a second target regulatory data pair that ranks first in the correlation score with the target gene within a preset third range and corresponds to the target transcription factor as the fifth target regulatory data pair, and to integrate all the fifth target regulatory data pairs found in the second category into the sixth subspace colocalization module;
[0232] The first building module is used to construct the spatial colocalization module of each transcription factor based on the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module, and the sixth subspace colocalization module corresponding to each transcription factor.
[0233] Optionally, in some embodiments, the analysis subunit includes:
[0234] The second acquisition module is used to acquire the preset motif enrichment analysis function;
[0235] The analysis module is used to perform sequence enrichment analysis on each subspace colocalization module within each spatial colocalization module based on a preset sequence enrichment analysis function, and obtain the sequence enrichment analysis results.
[0236] Optionally, in some embodiments, the third determining subunit includes:
[0237] The screening module is used to screen out the seventh subspace colocalization module based on the motif enrichment analysis results, which has an enrichment score greater than the eighth preset threshold and whose enriched transcription factors are consistent with the transcription factors in the subspace colocalization module.
[0238] The fifth determination module is used to determine the fourth target gene in the seventh subspace colocalization module whose enrichment contribution is greater than the ninth preset threshold;
[0239] The second building module is used to construct the spatial regulator of each transcription factor based on the transcription factors in the seventh sub-spatial colocalization module and the fourth target gene.
[0240] Optionally, in some embodiments, the inference unit includes:
[0241] The second acquisition subunit is used to acquire a preset scoring tool and score each spatial adjustment element based on the preset scoring tool to obtain the scoring result;
[0242] The fourth determining subunit is used to determine the score distribution of each regulator across multiple cells based on the scoring results;
[0243] The fifth determining unit is used to perform binary partitioning of multiple cells based on a binary Gaussian mixture model, and to determine the activity of transcription factors in each cell based on the binary partitioning results.
[0244] It is evident that the content of the above-described transcription factor activity inference method embodiments is applicable to the embodiments of this transcription factor activity inference device. The specific functions implemented by this transcription factor activity inference device embodiment are the same as those of the above-described transcription factor activity inference method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-described transcription factor activity inference method embodiments.
[0245] Referring to FIG20, FIG20 illustrates the hardware structure of a computer device according to another embodiment, the computer device including:
[0246] The processor 2001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0247] The memory 2002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 2002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 2002 and is called and executed by the processor 2001 using the transcription factor activity inference method or the transcription factor activity inference method of the embodiments of this application.
[0248] Input / output interface 2003 is used to implement information input and output;
[0249] The communication interface 2004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0250] Bus 2005 transmits information between various components of the device (e.g., processor 2001, memory 2002, input / output interface 2003, and communication interface 2004);
[0251] The processor 2001, memory 2002, input / output interface 2003 and communication interface 2004 are connected to each other within the device via bus 2005.
[0252] This application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the above-described transcription factor activity inference method.
[0253] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.
[0254] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0255] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.
[0256] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0257] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0258] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0259] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0260] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.
[0261] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.
Claims
1. A method for inferring transcription factor activity, characterized in that, The method includes: Acquire spatial transcriptome data of the target slice, the spatial transcriptome data including location information and gene expression level information of multiple gene categories; Genes in the gene expression information are screened according to the reference regulatory network corresponding to the target slice, and a target regulatory network is constructed based on the screened genes. The target regulatory network includes multiple first target regulatory data pairs, and each first target regulatory data pair contains a transcription factor and a corresponding target gene. Based on the location information and the gene expression level information, the nuclear density of transcription factors and target genes in each first target regulation data pair is estimated, and the correlation score of each first target regulation data pair is calculated based on the nuclear density estimation results. Based on the correlation scores, a spatial colocalization module for each transcription factor is constructed, and the spatial regulators of each transcription factor are determined based on the motif enrichment analysis results of the spatial colocalization modules. The spatial regulators of each transcription factor are scored, and the activity inference of the corresponding transcription factor is determined based on the scoring results.
2. The method according to claim 1, characterized in that, The step of screening genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and constructing a target regulatory network based on the screened genes, includes: Obtain the reference regulatory network corresponding to the target slice. The reference regulatory network includes multiple reference regulatory edges and reference transcription factors and reference target genes connected to each reference regulatory edge. Identify the intersection of the genes in the gene expression level information with the reference transcription factors and reference target genes to obtain multiple first transcription factors and multiple first target genes; Based on the gene expression level information, identify multiple second transcription factors and multiple second target genes whose corresponding capture sites are greater than a first preset threshold among the multiple first transcription factors and multiple first target genes; A target regulatory network is constructed based on the multiple second transcription factors and multiple second target genes.
3. The method according to claim 2, characterized in that, The construction of the target regulatory network based on the plurality of second transcription factors and the plurality of second target genes includes: In the reference regulatory network, at least one second transcription factor corresponding to each second target gene is located. A target regulatory network is constructed based on the plurality of second target genes and at least one second transcription factor corresponding to each second target gene.
4. The method according to claim 1, characterized in that, The step of estimating the nuclear density of transcription factors and target genes in each first target regulatory data pair based on the location information and the gene expression level information, and calculating the relevance score of each first target regulatory data pair based on the nuclear density estimation results, includes: For any first target regulation data pair, the location matrix of the corresponding third target gene and the corresponding third transcription factor is determined based on the location information and the gene expression level information. The nuclear density values of the third target gene and the third transcription factor at each spatial site are calculated based on the location matrix to obtain the nuclear density estimation results; The correlation score for each of the first target regulation data pairs is calculated based on the kernel density estimation results.
5. The method according to claim 4, characterized in that, The step of calculating the correlation score for each of the first target regulation data pairs based on the kernel density estimation results includes: In the target slice, a first region is identified where the nuclear density value of the third target gene is greater than a second preset threshold, and a second region is identified where the nuclear density value of the third transcription factor is greater than a third preset threshold. Identify the target area where the first region and the second region intersect; Calculate the Pearson correlation coefficient between the nuclear density of the third target gene and the third transcription factor in the target region to obtain the correlation score of the first target regulatory data pair.
6. The method according to claim 1, characterized in that, The process of constructing a spatial colocalization module for each transcription factor based on the correlation score, and determining the spatial regulator of each transcription factor based on the motif enrichment analysis results of the spatial colocalization module, includes: Based on the correlation score, the multiple first target regulation data pairs are filtered to obtain multiple second target regulation data pairs; Based on the multiple second-target regulatory data, a spatial colocalization module for each transcription factor was constructed. Sequence enrichment analysis was performed on each spatial colocalization module to obtain the sequence enrichment analysis results; The spatial regulators of each transcription factor were determined based on the motif enrichment analysis results.
7. The method according to claim 6, characterized in that, The process of filtering the multiple first-target regulation data pairs based on the correlation score to obtain multiple second-target regulation data pairs includes: Obtain a fourth preset threshold for controlling the correlation score; Among the plurality of first target control data pairs, a plurality of second target control data pairs with correlation scores greater than the fourth preset threshold are identified.
8. The method according to claim 6, characterized in that, The construction of a spatial co-localization module for each transcription factor based on the multiple second-target regulatory data includes: The multiple second target regulatory data pairs are divided into multiple first categories according to different transcription factors, and each first category of the second target regulatory data pairs contains the same transcription factor; The multiple second target regulatory data pairs are divided into multiple second categories according to different target genes, and the second target regulatory data pairs in each second category contain the same target gene; For each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the fifth preset threshold is identified as the first subspace colocalization module; For each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the sixth preset threshold is identified as the second subspace colocalization module. For each transcription factor, in each second target regulation data pair of the first category, the second target regulation data pair corresponding to the target gene whose correlation score with the transcription factor is greater than the seventh preset threshold is identified as the third subspace colocalization module. For each transcription factor, in each second category, a second target regulatory data pair that ranks first in the correlation score with the target gene and corresponds to the target transcription factor is identified as the third regulatory data pair. All third regulatory data pairs found in the second category are integrated into a fourth subspace colocalization module. For each transcription factor, in each second category, the second target regulatory data pair that ranks first in the correlation score with the target gene and corresponds to the target transcription factor is identified as the fourth target regulatory data pair. All fourth target regulatory data pairs found in the second category are integrated into the fifth subspace colocalization module. For each transcription factor, in each second category, the second target regulatory data pair that ranks first in the correlation score with the target gene and corresponds to the target transcription factor is identified as the fifth target regulatory data pair. All fifth target regulatory data pairs found in the second category are integrated into the sixth subspace colocalization module. A spatial colocalization module for each transcription factor is constructed based on the first subspace colocalization module, the second subspace colocalization module, the third subspace colocalization module, the fourth subspace colocalization module, the fifth subspace colocalization module, and the sixth subspace colocalization module corresponding to each transcription factor.
9. The method according to claim 6, characterized in that, The step of performing sequence enrichment analysis on each spatial co-location module to obtain sequence enrichment analysis results includes: Obtain the preset motif enrichment analysis function; Based on the preset sequence enrichment analysis function, sequence enrichment analysis is performed on each subspace colocalization module within each spatial colocalization module to obtain the sequence enrichment analysis results.
10. The method according to claim 6, characterized in that, The determination of the spatial regulators of each transcription factor based on the motif enrichment analysis results includes: Based on the motif enrichment analysis results, a seventh subspace colocalization module was selected, whose enrichment score was greater than the eighth preset threshold and whose enriched transcription factors were consistent with the transcription factors in the subspace colocalization module. Identify the fourth target gene in each seventh subspace colocalization module whose enrichment contribution is greater than the ninth preset threshold; Spatial regulators for each transcription factor are constructed based on the transcription factors in the seventh sub-spatial colocalization module and the fourth target gene.
11. The method according to claim 1, characterized in that, The process of scoring the spatial regulators of each transcription factor and determining the corresponding transcription factor activity inference result based on the scoring results includes: Obtain a preset scoring tool, and score each spatial adjuster based on the preset scoring tool to obtain the scoring result; The score distribution of each regulator in multiple cells is determined based on the scoring results; The multiple cells were binary-divided based on a binary Gaussian mixture model, and the activity of transcription factors in each cell was determined based on the binary division results.
12. A transcription factor activity inference device, characterized in that, The device includes: The acquisition unit is used to acquire spatial transcriptome data of the target slice, wherein the spatial transcriptome data includes location information and gene expression level information of multiple categories of genes. The construction unit is used to screen genes in the gene expression information according to the reference regulatory network corresponding to the target slice, and construct a target regulatory network according to the screened genes. The target regulatory network includes multiple first target regulatory data pairs, and each first target regulatory data pair contains a transcription factor and a corresponding target gene. The calculation unit is used to estimate the nuclear density of transcription factors and target genes in each first target regulation data pair based on the location information and the gene expression level information, and to calculate the correlation score of each first target regulation data pair based on the nuclear density estimation results. A determination unit is used to construct a spatial colocalization module for each transcription factor based on the correlation score, and to determine the spatial regulator of each transcription factor based on the motif enrichment analysis results of the spatial colocalization module. The inference unit is used to score the spatial regulators of each transcription factor and determine the inferred activity of the corresponding transcription factor based on the scoring results.
13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the transcription factor activity inference method according to any one of claims 1 to 11.
14. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the transcription factor activity inference method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the transcription factor activity inference method according to any one of claims 1 to 11.