A method and system for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptome
By acquiring multi-stage three-dimensional spatial transcriptome data and calculating a comprehensive tissue-specific score, the problem of low accuracy in three-dimensional tissue identification in traditional methods has been solved, and tissue-specific gene identification with high accuracy in dynamic developmental scenarios has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YAZHOUWAN NATIONAL LABORATORY
- Filing Date
- 2026-03-18
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional methods for identifying tissue-specific expressed genes are based on two-dimensional slice data, which cannot accurately reconstruct the spatial relationships of three-dimensional tissues, resulting in a decrease in recognition accuracy. This is especially true in scenarios with drastic dynamic changes, such as early embryonic development, where it is difficult to effectively identify tissue-specific expressed genes.
By acquiring three-dimensional spatial transcriptome data of target organisms at multiple consecutive developmental stages, and utilizing spatial adjacency maps, cell annotation data, and gene expression information, a comprehensive tissue-specific score is calculated, including spatial expression concentration, cell composition-corrected specificity parameters, cross-developmental phase stability parameters, and significance parameters, to screen for tissue-specific expressed genes.
It improves the accuracy of identifying tissue-specific expressed genes, eliminates interference from spurious differential signals, meets the needs of multi-scale structural analysis, and effectively identifies tissue-specific expressed genes related to tissues and substructures.
Smart Images

Figure CN122117019A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of gene recognition technology, and more specifically, relates to a method and system for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics. Background Technology
[0002] With the iteration of high-throughput sequencing and spatial positioning technologies, spatial transcriptome sequencing technology has overcome the limitations of traditional transcriptome sequencing in losing spatial location information, enabling the acquisition of three-dimensional spatial information on gene expression during embryonic and tissue development. However, the identification of traditional tissue-specific gene expression is limited to two-dimensional slice data or cannot reconstruct the spatial relationships of real three-dimensional tissues, which easily leads to a decrease in identification accuracy and affects the results of related scientific research. Summary of the Invention
[0003] The purpose of this application is to provide a method and system for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics, which can improve the accuracy of identifying tissue-specific expressed genes and is applicable to scenarios such as early embryonic development with dynamic changes in tissue boundaries, frequent cell migration, and significant fluctuations in cell proportions.
[0004] A first aspect of this application provides a method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics, including: The three-dimensional spatial transcriptome data of the target organism were obtained at multiple consecutive developmental stages. The three-dimensional spatial transcriptome data included cell three-dimensional coordinates, cell annotation data and gene expression information. The comprehensive tissue-specific score of each gene is calculated using at least two of the following parameters, and the tissue-specific expression genes of the target organism are identified based on the comprehensive tissue-specific scores of each gene: Based on the three-dimensional coordinates of cells, the three-dimensional spatial transcriptome data is transformed into a spatial adjacency graph. The spatial expression concentration of each gene is obtained by quantifying each gene using the Laplacian matrix corresponding to the spatial adjacency graph. Cell composition correction-specific parameters for each gene are obtained by grouping cells in the target organism according to cell type based on cell annotation data. Cross-developmental phase stability parameters were obtained by extracting the temporal characteristics of the tissue to which each gene in the target organism belongs based on three-dimensional spatial transcriptome data. The significance parameters of each gene are calculated based on the spatial adjacency graph.
[0005] A second aspect of this application provides a tissue-specific gene expression recognition system based on three-dimensional spatial transcriptomics, comprising: The data acquisition module acquires three-dimensional spatial transcriptome data of the target organism at multiple consecutive developmental stages. The three-dimensional spatial transcriptome data includes cell three-dimensional coordinates, cell annotation data, and gene expression information. The gene screening module is used to calculate the comprehensive tissue-specific score of each gene using at least two of the following parameters, and to identify tissue-specific expressed genes in the target organism based on the comprehensive tissue-specific scores of each gene: Based on the three-dimensional coordinates of cells, the three-dimensional spatial transcriptome data is transformed into a spatial adjacency graph. The spatial expression concentration of each gene is obtained by quantifying each gene using the Laplacian matrix corresponding to the spatial adjacency graph. Cell composition correction-specific parameters for each gene are obtained by grouping cells in the target organism according to cell type based on cell annotation data. Cross-developmental phase stability parameters were obtained by extracting the temporal characteristics of the tissue to which each gene in the target organism belongs based on three-dimensional spatial transcriptome data. The significance parameters of each gene are calculated based on the spatial adjacency graph.
[0006] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the above-described method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics.
[0007] In a fourth aspect of this application, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics.
[0008] The beneficial effects of the tissue-specific gene expression identification method and system based on three-dimensional spatial transcriptomics provided in this application are as follows: The embodiments of this application obtain three-dimensional spatial transcriptome data of the target organism at multiple consecutive developmental stages and perform subsequent analysis. This allows for full utilization of the adjacency relationships of cells in three-dimensional tissues, and accurate identification of specific genes enriched along the three-dimensional tissue structure.
[0009] This application also calculates a comprehensive tissue-specific score for each gene using at least two of the following parameters: spatial expression concentration, cell composition correction specificity parameter, cross-developmental phase stability parameter, and significance parameter. Based on the comprehensive tissue-specific score corresponding to each gene, the genes of the target organism are identified. This can eliminate the interference of spurious differential signals caused by differences in cell type ratios, fully consider the gene expression stability at different developmental stages, and verify the statistical significance of gene expression. It can meet the needs of multi-scale structural analysis, effectively identify tissue-specific expressed genes related to tissues, substructures, and developmental gradients, and improve the accuracy of tissue-specific expressed gene identification. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics, provided in an embodiment of this application; Figure 2 A structural block diagram of a tissue-specific gene recognition system based on three-dimensional spatial transcriptomics provided in an embodiment of this application; Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0013] To facilitate understanding, the main terms used in this application will be explained first.
[0014] Tissue-specific expressed genes refer to a class of functional genes that exhibit significantly and specifically different gene expression levels in a particular target tissue or cell type of a target organism compared to other tissues and cell types of the same organism. These genes can specifically mark the identity of the target tissue, regulate the differentiation and development of the target tissue, maintain the morphological structure and physiological homeostasis of the target tissue, and possess clear tissue-specific expression characteristics.
[0015] Multiple developmental phases refer to multiple different developmental stages of a target organism acquired sequentially over time during embryogenesis, organogenesis, tissue differentiation, and growth.
[0016] Cell annotation refers to labeling the cell types of single cells in a target organism at different developmental stages.
[0017] Tissue annotation refers to the labeling of tissue regions or anatomical structures at different developmental stages of a target organism.
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0019] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics, as provided in an embodiment of this application. The method may include S101 and S102.
[0020] S101: Obtain three-dimensional spatial transcriptome data of the target organism at multiple consecutive developmental stages. The three-dimensional spatial transcriptome data includes cell three-dimensional coordinates, cell annotation data, and gene expression information.
[0021] In this embodiment, the target organism is the organism to be studied. Cell three-dimensional coordinates represent the cell's coordinates in the x', y', and z' directions, and can be represented as (x', y', z'). Gene expression information refers to the quantitative data of transcript expression of each gene in each cell of the target organism, obtained through three-dimensional spatial transcriptome sequencing technology, including gene expression levels (such as the number of transcribed RNAs) and expression intensity, used to characterize the transcriptional activity of genes in the corresponding cells.
[0022] In one embodiment, the cell annotation data includes a cell type label corresponding to each single cell. For example, in the chicken embryo 3D spatial transcriptome, each single cell / spatial sequencing point can be labeled as: Neural crest cell, Mesoderm cell, Endoderm cell, Hepatocyte, and Endothelial cell.
[0023] In one embodiment, the three-dimensional spatial transcriptome data also includes tissue annotation data. Tissue annotation data refers to the anatomical or tissue region labels assigned to each spatial sequencing point or single cell, indicating the tissue affiliation of that sequencing point in a three-dimensional embryo or organ. For example, a batch of cells of the same cell type may be labeled as Liver, Heart, Neural tube, Somite, and Gut.
[0024] The three-dimensional spatial transcriptome data obtained in this embodiment includes data from each continuous developmental stage. Compared with the prior art which only uses data from a single time point for screening tissue-specific expressed genes, this method can effectively distinguish between "true tissue-specific expressed genes that are continuously / regularly expressed" and "non-robust genes whose transient gene expression levels are higher than a threshold." This threshold can be determined based on experimental data.
[0025] S102: Calculate the comprehensive tissue-specific score of each gene using at least two of the following parameters: spatial expression concentration, cell composition correction specificity parameter, cross-developmental phase stability parameter, and significance parameter. Based on the comprehensive tissue-specific score of each gene, identify the genes of the target organism to obtain tissue-specific expressed genes.
[0026] Among them, the three-dimensional spatial transcriptome data is transformed into a spatial adjacency graph based on the three-dimensional coordinates of the cell, and the spatial expression concentration of each gene is obtained by quantifying each gene using the Laplacian matrix corresponding to the spatial adjacency graph. Based on cell annotation data, cells in the target organism are grouped according to cell type to obtain cell composition correction-specific parameters for the cells to which each gene belongs; Based on three-dimensional spatial transcriptome data, the temporal characteristics of the tissue to which each gene belongs in the target organism are extracted to obtain cross-developmental phase stability parameters; The significance parameters of each gene are calculated based on the spatial adjacency graph.
[0027] In this embodiment, the comprehensive tissue-specific score of each gene is correlated with at least two of the following four parameters: spatial expression concentration, cell composition correction specificity parameter, cross-developmental phase stability parameter, and significance parameter.
[0028] First, when calculating the spatial expression concentration of each gene, it is necessary to first convert the three-dimensional spatial transcriptome data into a spatial adjacency graph, and then use the Laplacian matrix corresponding to the spatial adjacency graph to quantify each gene.
[0029] In this embodiment, gene expression levels are quantified using a Laplace matrix. This maps the originally high-dimensional, disordered transcriptome data into a numerical indicator reflecting the degree of spatial aggregation, thereby objectively distinguishing whether genes are regionally enriched or diffusely randomly expressed. Furthermore, quantifying gene expression levels using a Laplace matrix can suppress noise from random fluctuations between cells, enhance the true gene expression results with spatial structural characteristics, and improve the accuracy of subsequent tissue-specific gene screening.
[0030] Second, in order to eliminate spurious differential signals caused by differences in cell type ratios and retain only the differences in gene expression brought about by the tissue itself, this embodiment introduces cell composition correction-specific parameters.
[0031] In this embodiment, the cell composition correction specificity parameter is used to represent the distribution of gene expression along the three-dimensional tissue structure. The larger the value of the cell composition correction specificity parameter, the more the gene expression forms a continuous enrichment region along the three-dimensional tissue structure, rather than a random discrete distribution, thus intuitively reflecting its spatial specificity.
[0032] Third, in order to screen for tissue-specific genes that are stably or regularly expressed in multiple consecutive developmental stages and exclude transiently expressed genes, this embodiment introduces a cross-developmental phase stability parameter.
[0033] The cross-developmental phase stability parameter in this embodiment is used to ensure that the tissue-specific expressed genes selected are robust during dynamic development, rather than transiently expressed at a single time point.
[0034] Fourth, in order to screen for tissue-specific genes that exhibit statistical significance in their spatial distribution of gene expression and exclude randomly distributed non-specific genes, this embodiment introduces a significance parameter. The significance parameter in this embodiment is used to quantify the degree to which the gene expression distribution deviates from a random distribution.
[0035] In this embodiment, a comprehensive tissue-specific score for each gene is calculated by selecting at least two parameters from the aforementioned spatial expression concentration, cell composition correction specificity parameter, cross-developmental phase stability parameter, and significance parameter. Based on the relationship between the comprehensive tissue-specific score and a preset score threshold, genes exceeding the preset score threshold are selected, thus obtaining tissue-specific expression genes. The preset score threshold can be determined based on historical experimental data.
[0036] As can be seen from the above, this embodiment, by acquiring three-dimensional spatial transcriptome data of the target organism at multiple consecutive developmental stages and performing subsequent analysis, can make full use of the adjacency relationship of cells in three-dimensional tissues and accurately identify specific genes enriched along the three-dimensional tissue structure.
[0037] This embodiment also calculates a comprehensive tissue-specific score for each gene using at least two of the following parameters: spatial expression concentration, cell composition correction specificity parameter, cross-developmental phase stability parameter, and significance parameter. Based on the comprehensive tissue-specific score corresponding to each gene, the genes of the target organism are identified. This can eliminate the interference of spurious differential signals caused by differences in cell type ratios, fully consider the gene expression stability at different developmental stages, and verify the statistical significance of gene expression. It can meet the needs of multi-scale structural analysis, effectively identify tissue-specific expression genes related to tissues, substructures, and developmental gradients, and improve the accuracy of tissue-specific expression gene identification.
[0038] In one embodiment of this application, calculating the spatial expression concentration of each gene includes: The spatial distance between each pair of cells is calculated based on three-dimensional spatial transcriptome data. Based on the spatial distance, at least one neighboring cell is determined for each cell. A spatial adjacency graph is constructed based on the adjacency relationship between each cell. The spatial adjacency graph includes an adjacency matrix and a degree matrix. Each pair of cells consists of two cells. Calculate the Laplacian matrix of the spatial adjacency graph based on the adjacency matrix and degree matrix; The spatial expression concentration of each gene was obtained by quantifying each gene using the Laplace matrix.
[0039] In this embodiment, the spatial distance between each pair of cells can be represented by Euclidean distance. Determining at least one neighboring cell for each cell based on spatial distance can be understood as using the k-nearest neighbor algorithm, ε-neighborhood algorithm, Delaunay triangulation algorithm, or a combination of the above algorithms to determine at least one neighboring cell for each cell.
[0040] For example, for the k-nearest neighbor algorithm, for any cell i When the spatial distance between a cell and other cells is less than a preset spatial distance, it indicates that the cell is a cell. i The adjacent cells.
[0041] In this embodiment, the adjacency matrix A is N × N The adjacency matrix A is a set of matrices, where each data point represents whether two cells are adjacent. The degree matrix D can be calculated from the adjacency matrix A, where the diagonal elements of D represent the number of adjacent cells for each cell. Calculating the Laplacian matrix of the spatial adjacency graph based on the adjacency and degree matrices can be understood as directly subtracting the degree matrix from the adjacency matrix to obtain the non-normalized Laplacian matrix, or obtaining the normalized Laplacian matrix based on the following formula. : , Represents the identity matrix.
[0042] In one embodiment, the spatial expression concentration of each gene is obtained by quantifying each gene using the Laplace matrix, including: Obtain the original gene expression level of each gene in three-dimensional space; Calculate the variance of the original gene expression level; The original gene expression level was quantified using the Laplace matrix to obtain the quantified gene expression level, and the variance of the quantified gene expression level was calculated. The spatial expression concentration of each gene is obtained by comparing the variance of the quantified gene expression level with the variance of the original gene expression level.
[0043] In this embodiment, the spatial expression concentration of each gene is calculated based on the following formula. : =
[0044] Where x represents the original gene expression level. This represents the variance of gene expression levels quantified using the Laplace matrix. The variance representing the original gene expression level. This indicates the avoidance of minimum values where the denominator is 0.
[0045] In this embodiment, a spatial adjacency graph is constructed based on three-dimensional spatial transcriptome data, and the spatial expression concentration of each gene is obtained by quantifying each gene using the Laplacian matrix corresponding to the spatial adjacency graph. Finally, the genes in the target organism are screened based on the spatial expression concentration, and tissue-specific genes that form continuous enrichment regions along the three-dimensional tissue structure can be extracted.
[0046] In one embodiment of this application, calculating the cell composition correction-specific parameters of the cell to which each gene belongs includes: Based on cell annotation data, cells in the target organism are grouped according to cell type to obtain multiple sets of data; Based on multiple sets of data, the difference between the original gene expression levels of genes of the same cell type in the target tissue and in non-target tissues is calculated. The differences corresponding to all cell types are weighted and summed to obtain the cell composition correction specific parameter of the cell to which each gene belongs.
[0047] In this embodiment, the cell composition correction specificity parameter of the cell to which each gene belongs can be calculated based on the following formula: = ; in, This represents a cell composition correction-specific parameter, where c represents the c-th cell type. Indicates the number of cell types. This represents the weight corresponding to the c-th cell type. This represents the original gene expression level of the gene in cell type c of the target tissue. This represents the original gene expression level of cell type c in non-target tissue.
[0048] In this embodiment, grouped data is obtained by grouping three-dimensional spatial transcriptome data according to cell type based on cell annotation data. Cross-tissue comparison within the same cell type is achieved based on the grouped data, eliminating non-real specific signals introduced by differences in cell type ratios and retaining only the expression differences brought about by the tissue itself, thereby improving the accuracy of identifying tissue-specific expressed genes.
[0049] In one embodiment of this application, the temporal features include the average gene expression sequence at different developmental stages; the cross-developmental phase stability parameter obtained by calculating the temporal features of the tissue to which each gene belongs includes: Calculate the standard deviation and mean of the average gene expression sequence at different developmental stages for each gene, and calculate the coefficient of variation based on the standard deviation and mean; The cross-developmental phase stability parameter of each gene was obtained based on the coefficient of variation, and the coefficient of variation was negatively correlated with the cross-developmental phase stability parameter.
[0050] In this embodiment, the average gene expression sequence of each gene at different developmental stages can be μ={μ1,μ2,...,μt}, where μt is the average gene expression sequence of the gene in the target tissue at the t-th developmental stage, and t is the number of developmental stages.
[0051] In this embodiment, the cross-developmental stability parameter of each gene is calculated using the coefficient of variation (CV). The CV is calculated using the following formula: CV = ; Transphase stability parameters for: = ;in, The standard deviation of the average gene expression sequence. This represents the mean of the average gene expression sequence.
[0052] In this embodiment, the cross-developmental phase stability parameter is used to characterize the stability of gene expression within the same tissue across multiple developmental stages. The larger the value, the more stable the gene expression. In this embodiment, the cross-developmental phase stability parameter obtained by calculating the temporal characteristics of the target tissue to which each gene belongs can effectively screen for tissue-specific genes with stable expression.
[0053] In one embodiment of this application, the significance parameter of each gene is calculated, including: The spatial distance between each pair of cells is obtained based on the spatial adjacency graph, and at least one neighboring cell of each cell is determined based on the spatial distance. A spatial weight matrix is constructed based on the adjacency relationships between individual cells and the spatial distance between each pair of cells; the spatial weight matrix is used to quantify the strength of spatial associations between cells. The significance parameters of each gene were calculated based on three-dimensional spatial transcriptome data and spatial weight matrix.
[0054] In one embodiment, the significance parameter of each gene is calculated based on three-dimensional spatial transcriptome data and a spatial weight matrix, including: The average gene expression level of each gene was calculated based on three-dimensional spatial transcriptome data. The significance parameter of each gene is calculated based on the spatial weight matrix and the average gene expression level.
[0055] The significance parameters for each gene are calculated based on the spatial weight matrix and the average gene expression level, including: Based on the spatial weight matrix and the average gene expression level, the significance parameter of each gene is calculated using the following formula: = ×
[0056] in, The significance parameter is represented by N, which represents the total number of cells. This represents the spatial weight between cell i and cell j. This represents the average gene expression level across all cells. This indicates the gene expression level in cell i. This indicates the gene expression level in cell j.
[0057] In this embodiment, through Moran's I Indices can verify the statistical significance of the spatial distribution of gene expression, exclude randomly distributed non-specific genes, and make the tissue-specific expression genes obtained in the final screening more accurate.
[0058] In one embodiment of this application, the spatial expression concentration, cell composition-corrected specificity parameter, cross-developmental phase stability parameter, and significance parameter mentioned above can be weighted and summed to obtain a comprehensive tissue-specific score for each gene. :
[0059] Where p is the value obtained by performing a permutation test on I, used to quantify the degree to which the distribution deviates from a random distribution; , , , These are the weights corresponding to each parameter, which can be set according to experimental needs or automatically learned from training data, for example, all of which are 0.25.
[0060] In this embodiment, by combining data from four dimensions—three-dimensional spatial structure, cellular composition correction, cross-developmental phase stability, and spatial statistical significance—the tissue-specific expression genes screened based on a comprehensive tissue-specificity score are more accurate. The method in this embodiment is applicable to dynamic developmental scenarios and can significantly reduce the probability of false positives for tissue-specific expression genes.
[0061] In one embodiment, a comprehensive tissue-specific score based on individual genes is used. It can generate a "tissue × gene" TSscore matrix, and then output a list of tissue-specific expressed genes (sorted by TSscore) for all tissues, providing a basis for subsequent biological verification.
[0062] Corresponding to the above embodiment is a tissue-specific gene expression identification method based on three-dimensional spatial transcriptomics. Figure 2 This is a structural block diagram of a tissue-specific gene expression recognition system based on three-dimensional spatial transcriptomics, provided as an embodiment of this application. For ease of illustration, only the parts relevant to the embodiment of this application are shown. Reference Figure 2 The tissue-specific gene recognition system 20 based on three-dimensional spatial transcriptomics includes: a data acquisition module 21 and a gene screening module 22.
[0063] Among them, the data acquisition module 21 acquires three-dimensional spatial transcriptome data of the target organism at multiple consecutive developmental stages. The three-dimensional spatial transcriptome data includes cell three-dimensional coordinates, cell annotation data and gene expression information. Gene screening module 22 is used to calculate the comprehensive tissue-specific score of each gene using at least two of the following parameters, and to identify tissue-specific expressed genes in the target organism based on the comprehensive tissue-specific scores of each gene: Based on the three-dimensional coordinates of cells, the three-dimensional spatial transcriptome data is transformed into a spatial adjacency graph. The spatial expression concentration of each gene is obtained by quantifying each gene using the Laplacian matrix corresponding to the spatial adjacency graph. Cell composition correction-specific parameters for each gene are obtained by grouping cells in the target organism according to cell type based on cell annotation data. Cross-developmental phase stability parameters were obtained by extracting the temporal characteristics of the tissue to which each gene in the target organism belongs based on three-dimensional spatial transcriptome data. The significance parameters of each gene are calculated based on the spatial adjacency graph.
[0064] In one embodiment of this application, the gene screening module 22, when calculating the spatial expression concentration of each gene, is specifically used for: The spatial distance between each pair of cells is calculated based on three-dimensional spatial transcriptome data. Based on the spatial distance, at least one neighboring cell of each cell is determined. A spatial adjacency graph is constructed based on the adjacency relationship between each cell. The spatial adjacency graph includes an adjacency matrix and a degree matrix. Calculate the Laplacian matrix of the spatial adjacency graph based on the adjacency matrix and degree matrix; The spatial expression concentration of each gene was obtained by quantifying each gene using the Laplace matrix.
[0065] In one embodiment of this application, when the gene screening module 22 quantifies each gene using the Laplace matrix to obtain the spatial expression concentration of each gene, it is specifically used for: Obtain the original gene expression level of each gene in three-dimensional space; Calculate the variance of the original gene expression level; The original gene expression level was quantified using the Laplace matrix to obtain the quantified gene expression level, and the variance of the quantified gene expression level was calculated. The spatial expression concentration of each gene is obtained by comparing the variance of the quantified gene expression level with the variance of the original gene expression level.
[0066] In one embodiment of this application, the gene screening module 22, when calculating the cell composition correction-specific parameters of the cell to which each gene belongs, is specifically used for: Based on cell annotation data, cells in the target organism are grouped according to cell type to obtain multiple sets of data; Based on multiple sets of data, the difference between the original gene expression levels of genes of the same cell type in the target tissue and in non-target tissues is calculated. The differences corresponding to all cell types are weighted and summed to obtain the cell composition correction specific parameter of the cell to which each gene belongs.
[0067] In one embodiment of this application, the temporal features include the average gene expression sequences at different developmental stages; the gene screening module 22, when calculating the cross-developmental phase stability parameter obtained from the temporal features of the target tissue to which each gene belongs, is specifically used for: Calculate the standard deviation and mean of the average gene expression sequence at different developmental stages for each gene, and calculate the coefficient of variation based on the standard deviation and mean; The cross-developmental phase stability parameter of each gene was obtained based on the coefficient of variation, and the coefficient of variation was negatively correlated with the cross-developmental phase stability parameter.
[0068] In one embodiment of this application, the gene screening module 22, when calculating the significance parameter of each gene, is specifically used for: The spatial distance between each pair of cells is obtained based on the spatial adjacency graph, and at least one neighboring cell of each cell is determined based on the spatial distance. A spatial weight matrix is constructed based on the adjacency relationships between individual cells and the spatial distance between each pair of cells; the spatial weight matrix is used to quantify the strength of spatial associations between cells. The significance parameters of each gene were calculated based on three-dimensional spatial transcriptome data and spatial weight matrix.
[0069] In one embodiment of this application, when calculating the significance parameters of each gene based on three-dimensional spatial transcriptome data and a spatial weight matrix, the gene screening module 22 is specifically used for: The average gene expression level of each gene was calculated based on three-dimensional spatial transcriptome data. The significance parameter of each gene is calculated based on the spatial weight matrix and the average gene expression level.
[0070] In one embodiment of this application, when calculating the significance parameter of each gene based on the spatial weight matrix and the average gene expression level, the gene screening module 22 is specifically used for: Based on the spatial weight matrix and the average gene expression level, the significance parameter of each gene is calculated using the following formula: = ×
[0071] in, The significance parameter is represented by N, which represents the total number of cells. This represents the spatial weight between cell i and cell j. This represents the average gene expression level across all cells. This indicates the gene expression level in cell i. This indicates the gene expression level in cell j.
[0072] In one embodiment of this application, the cell annotation data includes a cell type label corresponding to each single cell.
[0073] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the data acquisition module 21 and the gene screening module 22 are shown.
[0074] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0075] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0076] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory.
[0077] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the embodiment of this application for a tissue-specific gene expression identification method based on three-dimensional spatial transcriptomics, or they can execute the implementation method of the electronic device described in the embodiment of this application, which will not be repeated here.
[0078] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0079] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0080] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0081] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0082] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or modules, or it may be an electrical, mechanical, or other form of connection.
[0083] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0084] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0085] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics, characterized in that, include: The three-dimensional spatial transcriptome data of the target organism at multiple consecutive developmental stages were obtained. The three-dimensional spatial transcriptome data included cell three-dimensional coordinates, cell annotation data, and gene expression information. The comprehensive tissue-specific score of each gene is calculated using at least two of the following parameters, and the tissue-specific expression genes of the target organism are identified based on the comprehensive tissue-specific scores of each gene: Based on the cell's three-dimensional coordinates, the three-dimensional spatial transcriptome data is converted into a spatial adjacency graph. The spatial expression concentration of each gene is obtained by quantifying each gene using the Laplacian matrix corresponding to the spatial adjacency graph. Cell composition correction specific parameters for each gene are obtained by grouping cells in the target organism according to cell type based on the cell annotation data. Cross-developmental phase stability parameters were obtained by extracting the temporal characteristics of the tissue to which each gene in the target organism belongs based on the three-dimensional spatial transcriptome data. The significance parameters of each gene are calculated based on the spatial adjacency graph.
2. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 1, characterized in that, Calculate the spatial expression concentration of each gene, including: The spatial distance between each pair of cells is calculated based on the three-dimensional spatial transcriptome data. At least one neighboring cell is determined for each cell based on the spatial distance. A spatial adjacency graph is constructed based on the adjacency relationship between each cell. The spatial adjacency graph includes an adjacency matrix and a degree matrix. Calculate the Laplace matrix of the spatial adjacency graph based on the adjacency matrix and degree matrix; The spatial expression concentration of each gene is obtained by quantifying each gene using the Laplace matrix.
3. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 2, characterized in that, The process of quantifying each gene using the Laplacian matrix to obtain the spatial expression concentration of each gene includes: Obtain the original gene expression level of each gene in three-dimensional space; Calculate the variance of the original gene expression level; The original gene expression level is quantified using the Laplacian matrix to obtain the quantified gene expression level, and the variance of the quantified gene expression level is calculated. The spatial expression concentration of each gene is obtained based on the ratio of the variance of the quantified gene expression level to the variance of the original gene expression level.
4. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 1, characterized in that, Calculate the cell composition-corrected specific parameters for the cell to which each gene belongs, including: Based on the cell annotation data, the cells in the target biological individual are grouped according to cell type to obtain multiple sets of data; Based on the multiple sets of data, the difference between the original gene expression levels of the same cell type in the target tissue and the original gene expression levels in non-target tissues is calculated. The differences corresponding to all cell types are weighted and summed to obtain the cell composition correction specific parameter of the cell to which each gene belongs.
5. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 1, characterized in that, The temporal features include the average gene expression sequence at different developmental stages; The cross-developmental phase stability parameters obtained by calculating the temporal characteristics of the tissue to which each gene belongs include: Calculate the standard deviation and mean of the average gene expression sequence at different developmental stages for each gene, and calculate the coefficient of variation based on the standard deviation and mean; The cross-developmental phase stability parameter of each gene is obtained based on the coefficient of variation, and the coefficient of variation is negatively correlated with the cross-developmental phase stability parameter.
6. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 1, characterized in that, Calculate the significance parameters for each gene, including: Based on the spatial adjacency graph, the spatial distance between each pair of cells is obtained, and based on the spatial distance, at least one neighboring cell of each cell is determined; A spatial weight matrix is constructed based on the adjacency relationships between individual cells and the spatial distance between each pair of cells; the spatial weight matrix is used to quantify the strength of spatial associations between cells. The significance parameter of each gene is calculated based on the three-dimensional spatial transcriptome data and the spatial weight matrix.
7. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 6, characterized in that, The calculation of the significance parameter for each gene based on the three-dimensional spatial transcriptome data and the spatial weight matrix includes: The average gene expression level of each gene was calculated based on the three-dimensional spatial transcriptome data. The significance parameter of each gene is calculated based on the spatial weight matrix and the average gene expression level.
8. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 7, characterized in that, The calculation of the significance parameter for each gene based on the spatial weight matrix and the average gene expression level includes: Based on the spatial weight matrix and the average gene expression level, the significance parameter of each gene is calculated using the following formula: = × in, The significance parameter is represented by N, which represents the total number of cells. This represents the spatial weight between cell i and cell j. This represents the average gene expression level across all cells. This indicates the gene expression level in cell i. This indicates the gene expression level in cell j.
9. The method for identifying tissue-specific expressed genes based on three-dimensional spatial transcriptomics as described in claim 1, characterized in that, The cell annotation data includes a cell type label for each individual cell.
10. A tissue-specific gene expression recognition system based on three-dimensional spatial transcriptomics, characterized in that, include: The data acquisition module acquires three-dimensional spatial transcriptome data of the target organism at multiple consecutive developmental stages. The three-dimensional spatial transcriptome data includes cell three-dimensional coordinates, cell annotation data, and gene expression information. The gene screening module is used to calculate the comprehensive tissue-specific score of each gene using at least two of the following parameters, and to identify tissue-specific expressed genes in the target organism based on the comprehensive tissue-specific scores of each gene: Based on the cell's three-dimensional coordinates, the three-dimensional spatial transcriptome data is converted into a spatial adjacency graph. The spatial expression concentration of each gene is obtained by quantifying each gene using the Laplacian matrix corresponding to the spatial adjacency graph. Cell composition correction specific parameters for each gene are obtained by grouping cells in the target organism according to cell type based on the cell annotation data. Cross-developmental phase stability parameters were obtained by extracting the temporal characteristics of the tissue to which each gene in the target organism belongs based on the three-dimensional spatial transcriptome data. The significance parameters of each gene are calculated based on the spatial adjacency graph.