Method and device for identifying spatial feature genes, electronic equipment and storage medium
By identifying spatially characteristic genes through screening and Moran index calculation, the problem of insufficient flexibility and applicability in existing methods is solved, enabling more effective gene identification and biological interpretation in spatial genome data with single-cell precision.
Patent Information
- Application Number
- CN202311647144.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-12-01
AI Technical Summary
Existing methods for identifying genes with spatial characteristics lack flexibility and applicability, especially in single 10x spatial datasets where it is difficult to effectively identify gene expression patterns across spatial locations.
By screening genes whose average expression level in the first spatial region is lower than a predetermined threshold, a second spatial region is defined. The Moran index, the assumed p-value, and the corrected q-value are used to identify spatial characteristic genes. The association between genes and spatial regions is established by combining a generalized logarithmic regression model.
It improves the flexibility and applicability of identifying spatially characteristic genes, enabling wider application in single-cell precision spatial genome data and enhancing biological interpretability.
Smart Images

Figure CN120089192B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of data processing, and particularly relates to a method and device for identifying spatial signature genes, an electronic device and a storage medium. BACKGROUND
[0002] In recent decades, transcriptomics technology has developed rapidly, especially the high-throughput spatial transcriptomics technology developed in recent years. Using this technology, gene expression characteristics and spatial distribution data can be obtained simultaneously in situ in tissues, further promoting the study of real gene expression of cells in situ in tissues. At the same time, combined with single-cell transcriptome and spatial transcriptome joint research, spatial location gene expression research can be combined with ultra-high resolution single-cell level.
[0003] A key task in spatial transcriptome research today is to identify spatially variant genes with different spatial expression patterns across spatial locations, so in order to elucidate the spatial variation of gene expression, spatial signature genes have been developed. Many computational methods, which can be divided into three categories according to the inherent principle: methods based on statistical modeling; methods based on machine learning; methods based on spatial grid.
[0004] However, the above-mentioned analysis methods for identifying spatial signature genes are mostly developed based on 10x spatial group data, and are all single spatial signature genes, which have limitations in flexibility and applicability. SUMMARY
[0005] The present disclosure provides a method and device for identifying spatial signature genes, an electronic device and a storage medium. The main purpose is to solve the problem of poor flexibility and applicability of the found spatial signature genes.
[0006] According to a first aspect of the present disclosure, a method for identifying spatial signature genes is provided, comprising:
[0007] Screening genes with average expression lower than a predetermined expression threshold in a first spatial region;
[0008] According to the need, a second spatial region for identifying spatial signature genes is demarcated from the first spatial region after screening genes;
[0009] Calculating the Moran index of each gene in the second spatial region;
[0010] Identifying and determining spatial signature genes in the second spatial region according to the Moran index.
[0011] Optionally, the second spatial region for identifying spatial signature genes is demarcated from the first spatial region after screening genes according to the need, comprising:
[0012] A second spatial region for spatial feature gene identification is demarcated from the first spatial region after screening genes by using a lasso tool based on spatial group clustering results in a jupyter notebook.
[0013] Optionally, the calculation of the Moran index for each gene in the second spatial region comprises:
[0014] The calculation of the Moran index for each gene in the second spatial region is based on a parallel strategy, and the Moran index corresponding to each gene is obtained.
[0015] Optionally, the identification of the spatial feature gene in the second spatial region according to the Moran index comprises:
[0016] The p-value value and the q-value value of each gene are calculated according to the Moran index of each gene.
[0017] According to the Moran index, the p-value value and the q-value value, genes meeting the spatial feature condition are screened from the genes as spatial feature genes; the spatial feature condition is that the Moran index is greater than a first threshold value, the p-value value is less than a second threshold value, and the q-value value is less than a third threshold value.
[0018] According to a second aspect of the present disclosure, a method for applying spatial feature genes is provided, which comprises:
[0019] Obtaining the identified spatial feature genes, which are genes identified according to the method of the first aspect described above;
[0020] Hierarchical clustering is performed on the spatial feature genes according to the expression values of the genes, and a predetermined number of spatially enriched gene modules are obtained.
[0021] Obtaining the representative genes in each spatially enriched gene module.
[0022] Optionally, the obtaining of the representative genes in each spatially enriched gene module comprises:
[0023] The linear correlation coefficient value between the expression amount of each gene in the spatially enriched gene module and the average expression amount of the genes in the enriched gene module is calculated.
[0024] Ranking is performed in descending order of the linear correlation coefficient values, and a predetermined number of genes with high rankings are taken as the representative genes of the spatially enriched gene module.
[0025] Optionally, the method for applying spatial feature genes further comprises:
[0026] acquire a spatial region in which spatial feature gene correlation needs to be performed;
[0027] establish a correlation relationship between the spatial region and the identified spatial feature gene based on a generalized log regression model.
[0028] According to a third aspect of the present disclosure, a device for identifying a spatial feature gene is provided, comprising:
[0029] a screening unit configured to screen genes with an average expression level lower than a predetermined expression level threshold in a first spatial region;
[0030] a delineation unit configured to delineate a second spatial region for spatial feature gene identification from the first spatial region after screening genes according to needs;
[0031] a calculation unit configured to calculate a Moran index for each gene in the second spatial region;
[0032] an identification unit configured to identify spatial feature genes in the second spatial region according to the Moran index.
[0033] According to a fourth aspect of the present disclosure, a device for application of a spatial feature gene is provided, comprising:
[0034] a first acquisition unit configured to acquire an identified spatial feature gene, the spatial feature gene being a gene identified according to the method of the first aspect;
[0035] a clustering unit configured to perform hierarchical clustering of spatial feature genes according to expression values of the genes to obtain a predetermined number of spatially enriched gene modules;
[0036] a second acquisition unit configured to acquire a representative gene in each spatially enriched gene module.
[0037] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:
[0038] at least one processor; and
[0039] a memory in communication connection with the at least one processor; wherein
[0040] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect or the method of the second aspect.
[0041] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method according to the first aspect or the method according to the second aspect.
[0042] According to a seventh aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method according to the first aspect or the method according to the second aspect.
[0043] The method and device for identifying spatial feature genes, the electronic device and the storage medium provided by the present disclosure filter genes with an average expression level lower than a predetermined expression threshold in a first spatial region; a second spatial region for identifying spatial feature genes is demarcated from the first spatial region after filtering genes according to needs; a Moran index is calculated for each gene in the second spatial region; and spatial feature genes in the second spatial region are identified and determined according to the Moran index. Compared with related technologies, the present application filters genes with an average expression level lower than a predetermined expression threshold in a first spatial region to obtain suitable genes for research, and then demarcates a spatial region of interest for identifying spatial feature genes in the first spatial region according to needs, and then calculates a Moran index for each gene in the spatial region of interest to obtain spatial feature genes. The spatial feature genes can be widely applied to spatial group data with single-cell precision, and have higher flexibility and applicability.
[0044] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0045] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0046] Figure 1 A flowchart of a method for identifying spatial feature genes according to an embodiment of the present disclosure is shown;
[0047] Figure 2 A flowchart of a method for applying spatial feature genes according to an embodiment of the present disclosure is shown;
[0048] Figure 3 A flowchart of a method for identifying and determining spatial feature genes in a second spatial region according to a Moran index according to an embodiment of the present disclosure is shown;
[0049] Figure 4 A flowchart of a method for obtaining representative genes in each spatially enriched gene module according to an embodiment of the present disclosure is shown;
[0050] Figure 5 A flowchart of another application method of a spatial feature gene provided by an embodiment of the present disclosure is shown in FIG. 12;
[0051] Figure 6 A clustering result diagram of a mouse brain slice provided by an embodiment of the present disclosure is shown in FIG. 13;
[0052] Figure 7 A diagram of expression of the top 3 genes ranked by moran i score provided by an embodiment of the present disclosure is shown in FIG. 14;
[0053] Figure 8 A diagram of a spatially enriched gene module and representative genes of each module provided by an embodiment of the present disclosure is shown in FIG. 15;
[0054] Figure 9 A gene expression distribution diagram for establishing a regression model of gene Mbp and cell type c, taking gene Mbp as an example, provided by an embodiment of the present disclosure is shown in FIG. 16;
[0055] Figure 10 A diagram of a Lasso result of a region of interest based on a spatial group clustering result provided by an embodiment of the present disclosure is shown in FIG. 17;
[0056] Figure 11 A structural diagram of an apparatus for identifying a spatial feature gene provided by an embodiment of the present disclosure is shown in FIG. 18;
[0057] Figure 12 A structural diagram of another apparatus for identifying a spatial feature gene provided by an embodiment of the present disclosure is shown in FIG. 19;
[0058] Figure 13 A structural diagram of an application apparatus of a spatial feature gene provided by an embodiment of the present disclosure is shown in FIG. 20;
[0059] Figure 14 A structural diagram of another application apparatus of a spatial feature gene provided by an embodiment of the present disclosure is shown in FIG. 21;
[0060] Figure 15 A schematic block diagram of an example electronic device 1500 is shown in FIG. 22. DETAILED DESCRIPTION
[0061] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.
[0062] The method and device for identifying spatial feature genes, the electronic device, and the storage medium are described below with reference to the accompanying drawings.
[0063] Figure 1 A flowchart of a method for identifying spatial feature genes is provided in the embodiments of the present disclosure. As shown in the figure, the method comprises the following steps: Figure 1
[0064] Step 101: screening genes with average expression lower than a predetermined expression threshold in a first spatial region.
[0065] The first spatial region can be a piece of tissue slice for research, or a cell region in a certain region. Specifically, the embodiments of the present disclosure do not limit the first spatial region; the average expression of the gene refers to the expression of the gene, which is the amount of protein finally translated from the expression of the gene, i.e., deoxyribo nucleic acid (DNA). Because many low-expression or non-expression genes may affect the analysis and judgment of data in subsequent analysis, the low-expression genes need to be screened out; the predetermined expression threshold refers to a value set by actual data to filter out genes with expression lower than the value.
[0066] Step 102: according to needs, demarcating a second spatial region for identifying spatial feature genes from the first spatial region after screening genes.
[0067] In the embodiments of the present disclosure, the second spatial region for identifying spatial feature genes from the first spatial region after screening genes refers to demarcating a research region on the entire tissue slice; the needs refer to using a lasso tool based on spatial group clustering results to arbitrarily demarcate a region as a sub-object for research on the entire tissue slice, and the number of demarcated regions can be determined by the user.
[0068] It should be noted that the spatial feature gene refers to a gene with special spatial differential expression.
[0069] Step 103: calculating a Moran index for each gene in the second spatial region.
[0070] In the embodiments of the present disclosure, the second spatial region refers to a region arbitrarily demarcated using a lasso tool based on spatial group clustering results; the calculation of the Moran index for each gene refers to the principle of the product of attribute and spatial relationship, to obtain the final correlation in space.
[0071] Step 104, identifying a spatial feature gene in the second spatial region according to the Moran index.
[0072] Figure 2 A flowchart of an application method of a spatial feature gene provided by an embodiment of the present disclosure is shown in FIG. 2. As shown in FIG. 2, the method comprises the following steps: Figure 2
[0073] Step 201, obtaining the identified spatial feature gene.
[0074] The spatial feature gene is the gene identified in steps 101-104. For related descriptions, refer to the related descriptions of steps 101-104, which will not be repeated here.
[0075] Step 202, performing hierarchical clustering on the spatial feature gene according to the expression value of the gene to obtain a predetermined number of spatial enrichment gene modules.
[0076] The spatial enrichment gene module refers to a set of genes with commonality in the expression value of the gene. Enrichment refers to the process of classifying genes according to prior knowledge, i.e., gene annotation information. After classification, the gene can help to find out whether the gene has certain commonality (such as function, composition, etc.).
[0077] The predetermined number is because a parameter must be set to define the number of modules for hierarchical clustering. The number of modules can be defined according to actual data, and a default value can be set, such as 20. If the classification is not fine enough, the parameter can be adjusted to increase the number of modules, i.e., the more the number of modules, the more fine the classification.
[0078] Step 203, obtaining a representative gene in each spatial enrichment gene module.
[0079] The representative gene in each spatial enrichment gene module refers to a linear correlation between each gene in each spatial enrichment gene module and the average value of the spatial enrichment gene module. The top-ranked gene is the representative gene. The representative gene can be one or more than one, which is taken according to actual needs, and the embodiments of the present disclosure do not limit this.
[0080] The method for identifying spatial feature genes provided by the present disclosure includes the following steps: screening genes with average expression lower than a predetermined expression threshold in a first spatial region; demarcating a second spatial region for spatial feature gene identification from the first spatial region after screening genes; calculating a Moran index for each gene in the second spatial region; and identifying spatial feature genes in the second spatial region according to the Moran index. Compared with related technologies, the present disclosure screens genes with average expression lower than a predetermined expression threshold to obtain suitable genes for research, then demarcates a spatial region of interest for spatial feature gene identification according to needs, and then calculates a Moran index for each gene in the spatial region of interest to obtain spatial feature genes. The spatial feature genes can be widely applied to spatial group data with single-cell precision, and have higher flexibility and applicability.
[0081] In an implementation manner of the present embodiment, when demarcating a second spatial region for spatial feature gene identification from the first spatial region after screening genes according to needs, the second spatial region for spatial feature gene identification is also called a region of interest (ROI), and can be implemented by using, but is not limited to, the following method: demarcating a second spatial region for spatial feature gene identification from the first spatial region after screening genes by using a lasso tool based on spatial group clustering results in a jupyter notebook. The lasso tool can be integrated in the jupyter notebook, can select a region of interest (ROI) for a spatial group analysis object, and then obtains a sub-object for downstream spatial feature gene analysis, and the default is cells in all regions on a chip. The jupyter notebook is a Web application that can combine explanatory text, mathematical equations, codes and visualizations into an easily shared document.
[0082] Further, when calculating a Moran index for each gene in the second spatial region, the following method can be used, but is not limited to the following method: calculating a Moran index for each gene in the second spatial region by using a parallel strategy.
[0083] The Moran index for each gene in the second spatial region is calculated based on a parallel strategy to obtain a Moran index corresponding to each gene.
[0084] In an implementation of the embodiment of the present disclosure, the calculation of the Moran index for each gene in the second spatial region is performed, i.e., the calculation of the Moran index for each gene in the region of interest (ROI) is performed; for the genes in the region of interest (ROI), a parallel strategy is adopted, and the Moran index morani for each gene x is calculated by the following formula 1, and the formula 1 is specifically as follows:
[0085]
[0086] wherein N is the number of cells (in the region of interest), x i is the expression amount of the gene x in the cell i, is the average expression value of the gene x in all cells, w ij is the distance weight of the cell i and the cell j, and the default value is w ij = 1 when the cells are adjacent to each other, and w ij = 0 otherwise, and W is the sum of all weights w ij .
[0087] In an implementation of the embodiment, when the spatial feature gene in the second spatial region is identified and determined according to the Moran index, the following method can be adopted, but is not limited to, as shown in the following formula 2, the method comprises the following steps: Figure 3
[0088] Step 301, calculating the assumed value p-value and the corrected value q-value of each gene according to the Moran index of each gene.
[0089] Step 302, screening the genes meeting the spatial feature condition from the genes as the spatial feature genes according to the Moran index, the assumed value p-value and the corrected value q-value; the spatial feature condition is that the Moran index is greater than a first threshold value, the assumed value p-value is less than a second threshold value, and the corrected value q-value is less than a third threshold value.
[0090] wherein the first threshold value in the Moran index greater than the first threshold value is any value in 0 to 1, which can be selected according to the actual situation. The closer the Moran index is to 1, the stronger the spatial specificity of the gene is; the second threshold value in the assumed value p-value less than the second threshold value is 0.05; and the third threshold value in the corrected value q-value less than the third threshold value is 0.01.
[0091] In the specific determination of the spatial feature gene, the z value of each gene can be calculated by the Moran index of each gene through formula 2, and the z value and the p-value value are used as a reference to determine whether the gene expression has a spatial feature. The z value is a statistical quantity calculated according to sample data under the assumption of normal distribution, which represents the distance between the sample mean and the population mean, and is measured in standard deviation units. In the hypothesis test of normal distribution, the z value can be calculated by the mean and standard deviation of the sample, and then the standard normal distribution table is looked up, and the p-value value is used as a reference to determine whether the gene expression has a spatial feature;
[0092]
[0093] Wherein, z is a z-score (z-score), also called a standard score (standard score), which is a process of dividing the difference between a number and an average by a standard deviation, wherein E[I] = -1 / (N-1), Var[I] = E[I 2 ]-E[I] 2 ; perform a permutation test (Monte-Carlo), with a default number of permutations of 599, to obtain a hypothetical value p-value; adopt the Benjamini-Hochberg method to control the false discovery rate (false discovery rate) to calculate a corrected value q-value; and screen out genes with a Moran index I>0, a hypothetical value p-value<0.05, and a corrected value q-value<0.01 as spatial feature genes. Specifically, an example is provided to illustrate the process of the permutation test: it is assumed in advance that all gene expression values are randomly distributed and have no spatial feature, then the positions of the cells in space are randomly disturbed, and the I value is calculated, for example: 599 times of repetition (the more the number of repetitions, the more stable the background distribution obtained), 599 permutation I values obtained by 599 permutation arrangements, at this time, the 599 permutation I values can represent the sampling population situation of this permutation experiment (approximately normal distribution), the number of times that the actual I value is greater than the 599 permutation I values is counted, and then the p-value value is calculated (the number of times that the actual I is greater than the number of permutations 599), when the p-value value is less than 0.05, it indicates that the original hypothesis is wrong, and it is proved that the gene expression has a spatial feature and is not randomly distributed.
[0094] In one implementation manner of the embodiment, in the process of obtaining the representative gene in each spatially enriched gene module, the following method can be used, but is not limited to the following method, for example, as shown in the following formula: Figure 4 The method comprises the following steps:
[0095] Step 401, calculating a linear correlation coefficient value between the expression of each gene in the space enrichment gene module and the average expression of the genes in the enrichment gene module.
[0096] In an embodiment of the present disclosure, the linear correlation can be Pearson correlation, or other linear correlation, and the specific embodiment of the present disclosure does not limit this. When it is Pearson correlation, the linear correlation coefficient value between the expression of each gene in the space enrichment gene module and the average expression of the genes in the enrichment gene module is calculated as the Pearson correlation coefficient value between the expression of each gene in the space enrichment gene module and the average expression of the genes in the enrichment gene module.
[0097] Step 402, ranking according to the linear correlation coefficient value from large to small, and taking the top pre-determined number of genes as the representative genes of the space enrichment gene module.
[0098] As shown above, ranking according to the Pearson correlation coefficient value from large to small, and taking the top several as the representative genes of the module.
[0099] Further, in one implementation manner of the present embodiment, the application method of the space feature gene includes: Figure 5 As shown, the application method of the space feature gene includes:
[0100] Step 501, obtaining a space region for which the space feature gene correlation needs to be performed.
[0101] The space region can be a cell type, a region in units of region, or a cell tissue in units of cell, and the specific embodiment of the present disclosure does not limit this.
[0102] Step 502, establishing the correlation between the space region and the identified space feature gene based on a generalized log regression model.
[0103] When the correlation between the space region and the identified space feature gene is established based on the generalized log regression model, the following method can be used, but is not limited to this. The method includes: establishing the correlation between the space region and the identified space feature gene by formula 3.
[0104]
[0105]
[0106] Wherein, P i is the expression value of gene P in cell i, α is a constant term, c kmay be any variable of interest, such as cell type or spatial domain, here refers to the kth spatial domain of the partition, w ij is the distance weight between cell i and cell j, the default value is w ij = 1 if the cells are neighbors, otherwise w ij = 0. Based on the above description of the embodiments, the following will take the mouse brain as an example, which analyzes the mouse brain spatial group data (stereo-seq) including 7765 cells, 25691 genes, and displays the results:
[0107] Based on the scheme of the present disclosure, as Figure 6 is the clustering result map of the mouse brain slice, a total of 451 spatial feature genes are found, and the top 3 genes ranked by morani score are as shown in Figure 7 :
[0108] Figure 6 indicates a clustering result map made by cell type, and different colors represent the distribution of different cell types. Figure 7 The top 3 genes ranked by morani score are used to visualize the genes obtained by morani, wherein the light gray area is the area where the gene is expressed, and it can be clearly seen from Figure 7 that the expression of the gene calculated by morani has spatial characteristics, which further verifies that the gene obtained by morani is a spatial feature gene.
[0109] For the spatially enriched gene module and the representative gene with the highest score of each module, please refer to Figure 8 , the cell types corresponding to the three gene modules are fiber tracts, thalamus, and striatum in turn.
[0110] Figure 8 module 1, module 2, and module 3 in FIG. are the visualization result maps of the results in step 202, wherein the parameter is set to 3 when defining the number of modules, and the related description can be referred to the related description of step 202. The embodiments of the present disclosure will not be repeated here.
[0111] Figure 8 Plp1, Ntng1, and Phactr1 in FIG. are the visualization results of the results in step 203, wherein one representative gene is selected for each module, and the related description can be referred to the related description of step 203. The embodiments of the present disclosure will not be repeated here.
[0112] In addition, when establishing the correlation between the region and the spatial feature gene, the region type c can be selected as the cell type, and the regression model of all 451 spatial feature genes obtained from the mouse brain and the cell type is taken as an example. The results of the gene Mbp are shown in Table 1, and the gene expression is shown in Figure 9 :
[0113] Table 1
[0114]
[0115] Figure 9 The gene expression distribution diagram of the regression model of the gene Mbp and the cell type c is taken as an example, wherein the light gray area is the expression distribution position of the gene Mbp. Table 1 shows the results of the regression model of the gene Mbp and the cell type, wherein it can be seen that the regression coefficient of the gene Mbp and the cell type fiber tracts is 0.264960907, which is the highest, indicating that the gene Mbp and the cell type fiber tracts are most relevant. Comparing the expression distribution position of the gene Mbp with the distribution results of the cell type fiber tracts in the above Figure 9 , the distributions are consistent, which fully proves the effect of the accuracy of the spatial feature gene identified by the present disclosure. Figure 6
[0116] In addition, the Lasso is performed on the region of interest based on the spatial group clustering results, the spatial feature gene is identified, and the downstream analysis is performed, as shown in Figure 10 :
[0117] In the embodiment of the present disclosure, the default is that all the slice tissues are studied, wherein Figure 10 shows that the region can be demarcated for study, and the dark region is the demarcated region.
[0118] The embodiment of the present disclosure screens the genes with an average expression lower than a predetermined expression threshold in the first spatial region to obtain suitable genes for study, and then demarcates a spatial research region of interest for spatial feature gene identification in the first spatial region according to needs. Then, the Moran index of each gene in the spatial research region of interest is calculated to obtain the spatial feature gene. This spatial feature gene can be widely applied to single-cell-precision spatial group data, and has higher flexibility and applicability. At the same time, the spatial enrichment gene module is determined by the spatial feature gene, the representative gene of the spatial enrichment gene module is obtained, and the cell type determined by the representative gene of the spatial enrichment gene module can be correlated with the spatial feature gene to establish a correlation model, so that the biological interpretability of the spatial feature gene obtained by the embodiment of the present disclosure is stronger.
[0119] Corresponding to the method for identifying a spatial feature gene and the application method of the spatial feature gene, the application further provides a device for identifying a spatial feature gene and an application device of the spatial feature gene. Since the device embodiment of the application corresponds to the method embodiment described above, for the details not disclosed in the device embodiment, the method embodiment described above can be referred to, and the application will not be described again.
[0120] Figure 11 A structural schematic diagram of a device for identifying a spatial feature gene provided by an embodiment of the application is shown in Figure 11 , and includes:
[0121] The screening unit 1101 is configured to screen genes with an average expression amount lower than a predetermined expression amount threshold in a first spatial region.
[0122] The demarcating unit 1102 is configured to demarcate a second spatial region for identifying a spatial feature gene from the first spatial region after the genes are screened according to a requirement.
[0123] The calculating unit 1103 is configured to calculate a Moran index for each gene in the second spatial region.
[0124] The identifying unit 1104 is configured to identify a spatial feature gene in the second spatial region according to the Moran index.
[0125] The device for identifying a spatial feature gene provided by the application screens genes with an average expression amount lower than a predetermined expression amount threshold in a first spatial region, demarcates a second spatial region for identifying a spatial feature gene from the first spatial region after the genes are screened according to a requirement, calculates a Moran index for each gene in the second spatial region, and identifies a spatial feature gene in the second spatial region according to the Moran index. Compared with the related art, the application embodiment screens genes with an average expression amount lower than a predetermined expression amount threshold to obtain suitable genes for research, then demarcates a spatial region of interest for identifying a spatial feature gene according to a requirement, and then calculates a Moran index for each gene in the spatial region of interest to obtain a spatial feature gene. The spatial feature gene can be widely applied to spatial group data with single-cell precision, and has higher flexibility and applicability.
[0126] Further, in a possible implementation manner of the embodiment, as shown in Figure 12 , the demarcating unit 1102 is specifically configured to demarcate a second spatial region for identifying a spatial feature gene from the first spatial region after the genes are screened by using a lasso tool based on a spatial group clustering result in a jupyter notebook.
[0127] Further, in a possible implementation of the embodiment, as shown in Figure 12 the calculation unit 1103 is specifically configured to calculate the Moran index of each gene in the second spatial region based on a parallel strategy, to obtain a Moran index corresponding to each gene.
[0128] Further, in a possible implementation of the embodiment, as shown in Figure 12 the identification unit 1104 further includes:
[0129] The calculation module 11041 is configured to calculate a p-value value and a q-value value of each gene according to the Moran index of each gene.
[0130] The screening module 11042 is configured to screen a gene satisfying a spatial feature condition from the genes as a spatial feature gene according to the Moran index, the p-value value and the q-value value; the spatial feature condition is that the Moran index is greater than a first threshold value, the p-value value is less than a second threshold value, and the q-value value is less than a third threshold value.
[0131] Figure 13 A structural schematic diagram of an application device of a spatial feature gene provided by an embodiment of the present disclosure is shown in Figure 13 and includes:
[0132] The first acquisition unit 1201 is configured to acquire the identified spatial feature gene, wherein the spatial feature gene is the gene identified in steps 101-104. For related descriptions, reference can be made to the related descriptions of steps 101-104, and the embodiments of the present disclosure will not be described here again.
[0133] The clustering unit 1202 is configured to perform hierarchical clustering on the spatial feature genes according to the expression values of the genes, to obtain a predetermined number of spatial enrichment gene modules.
[0134] The second acquisition unit 1203 is configured to acquire a representative gene in each spatial enrichment gene module.
[0135] Further, in a possible implementation of the embodiment, as shown in Figure 14 the second acquisition unit 1203 includes:
[0136] The third calculation module 12031 is configured to calculate a linear correlation value between the expression amount of the gene in each spatial enrichment gene module and the average expression amount of the genes in the enrichment gene module.
[0137] The acquisition module 12032 is configured to rank the linear correlation coefficient values in descending order, and take a predetermined number of genes with high ranking as the representative genes of the spatial enrichment gene module.
[0138] Further, in a possible implementation of the embodiment, as shown in Figure 14 The device further includes:
[0139] The third acquisition unit 1204 is configured to acquire a spatial region in which the spatial feature gene correlation needs to be performed.
[0140] The establishment unit 1205 is configured to establish a correlation between the spatial region and the identified spatial feature gene based on a generalized logit model.
[0141] It should be noted that the foregoing explanation of the method embodiment is also applicable to the device of the embodiment, and the principle is the same, which is not limited in the embodiment.
[0142] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0143] Figure 15 A schematic block diagram of an example electronic device 1500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0144] As shown in Figure 15 The device 1500 includes a computing unit 1501 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 1502 or a computer program loaded from the storage unit 1508 into a RAM (Random Access Memory) 1503. Various programs and data required for the operation of the device 1500 can also be stored in the RAM 1503. The computing unit 1501, the ROM 1502, and the RAM 1503 are connected to each other through a bus 1504. An I / O (Input / Output) interface 1505 is also connected to the bus 1504.
[0145] A number of components in the device 1500 are connected to the I / O interface 1505, including: an input unit 1506, such as a keyboard, a mouse, etc.; an output unit 1507, such as various types of displays, speakers, etc.; a storage unit 1508, such as a magnetic disk, an optical disk, etc.; and a communication unit 1509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1509 allows the device 1500 to exchange information / data with other devices through computer networks, such as the Internet, and / or various telecommunication networks.
[0146] The computing unit 1501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 1501 performs various methods and processes described above, such as the method of identifying spatial signature genes. For example, in some embodiments, the method of identifying spatial signature genes can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 1508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1500 via the ROM 1502 and / or the communication unit 1509. When the computer program is loaded onto the RAM 1503 and executed by the computing unit 1501, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 1501 can be configured to perform the aforementioned method of identifying spatial signature genes by any other appropriate means, such as by means of firmware.
[0147] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on a Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0148] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general or special purpose computer, such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0149] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electrical connection, a portable computer diskette, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, an optical fiber, a compact disc (CD) ROM, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0150] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0151] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (local area network), a WAN (wide area network), the Internet, and a blockchain network.
[0152] The computer system can include clients and servers. This relationship can be between a client and a server that are typically remote from each other and typically interact through a communication network. The relationship between client and server exists by virtue of computer programs running on the respective computer systems and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS (Virtual Private Server, or VPS for short) services. The server can also be a server of a distributed system, or a server combined with a blockchain.
[0153] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of people (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.
[0154] It should be understood that the various forms of the flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.
[0155] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for identifying spatial characteristic genes, characterized in that, The method comprises: screening genes with average expression lower than a predetermined expression threshold in a first spatial region; demarcating a second spatial region for spatial feature gene identification from the first spatial region after screening genes as needed; calculating the Moran index of each gene in the second spatial region; identifying spatial feature genes in the second spatial region according to the Moran index.
2. The method of claim 1, wherein, The demarcating a second spatial region for spatial feature gene identification from the first spatial region after screening genes as needed comprises: demarcating a second spatial region for spatial feature gene identification from the first spatial region after screening genes by using a lasso tool based on spatial group clustering results in a jupyter notebook.
3. The method of claim 2, wherein, The calculating the Moran index of each gene in the second spatial region comprises: calculating the Moran index of each gene in the second spatial region based on a parallel strategy to obtain the Moran index corresponding to each gene.
4. The method according to any one of claims 1-3, characterized in that, The identifying spatial feature genes in the second spatial region according to the Moran index comprises: calculating the p-value and q-value of each gene according to the Moran index of each gene; screening genes meeting spatial feature conditions as spatial feature genes from the genes according to the Moran index, the p-value and the q-value; the spatial feature conditions are that the Moran index is greater than a first threshold, the p-value is less than a second threshold, and the q-value is less than a third threshold.
5. A method of using a spatial feature gene, characterized by, The method comprises: obtaining identified spatial feature genes, which are genes identified according to any one of the methods in claims 1-4; performing hierarchical clustering on the spatial feature genes according to the expression values of the genes to obtain a predetermined number of spatially enriched gene modules; obtaining representative genes in each spatially enriched gene module.
6. The method of claim 5, wherein, The obtaining representative genes in each spatially enriched gene module comprises: calculating the linear correlation coefficient value between the expression of each gene in the spatially enriched gene module and the average expression of the genes in the enriched gene module; ranking the linear correlation coefficient values from large to small, and taking the top predetermined number of genes as the representative genes of the spatially enriched gene module.
7. The method of claim 5, wherein, The method further comprises: obtaining a spatial region for which the correlation of spatial feature genes needs to be determined; establishing the correlation between the spatial region and the identified spatial feature genes based on a generalized logit regression model.
8. A device for identifying spatial characteristic genes, characterized in that, The method comprises: a screening unit configured to screen genes with average expression lower than a predetermined expression threshold in a first spatial region; a demarcating unit configured to demarcate a second spatial region for spatial feature gene identification from the first spatial region after screening genes as needed; a calculating unit configured to calculate the Moran index of each gene in the second spatial region; an identifying unit configured to identify spatial feature genes in the second spatial region according to the Moran index.
9. An apparatus for applying spatial signature genes, characterized by The method comprises: a first obtaining unit, configured to obtain the identified spatial feature genes, the spatial feature genes being genes identified according to the method in any one of claims 1-4; a clustering unit, configured to perform hierarchical clustering on the spatial feature genes according to the expression values of the genes, to obtain a predetermined number of spatially enriched gene modules; a second obtaining unit, configured to obtain a representative gene in each spatially enriched gene module.
10. An electronic device, comprising: comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method in any one of claims 1-4 or the method in any one of claims 5-7.
11. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, the computer instructions are used to enable the computer to perform the method in any one of claims 1-4 or the method in any one of claims 5-7.
12. A computer program product, characterised in that, comprising a computer program, which, when executed by a processor, implements the method in any one of claims 1-4 or the method in any one of claims 5-7.
Citation Information
Patent Citations
Spatial feature analysis for digital pathology images
CN115668304A
Cell type unbiased positioning method and system based on space transcriptome
CN116364180A