A single-cell cross-species cell type identification method

By constructing a single-cell reference database and using homologous gene alignment methods, the problem of cross-species cell type identification in the existing technology is solved, and efficient and accurate cell type identification of large-scale single-cell data is achieved.

CN115064220BActive Publication Date: 2025-05-16ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210677411.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-05-16
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

Existing single-cell cell type identification tools are difficult to conduct cross-species analysis, and the calculation accuracy and efficiency of large-scale single-cell data are difficult to guarantee, and there is a lack of a unified reference database with comparable annotations.

Method used

By collecting single-cell map data of multiple species, a single-cell reference database was constructed, and the identification of cross-species cell types was achieved using homologous gene alignment, characteristic gene extraction and similarity alignment of gene expression patterns.

Benefits of technology

The identification of cell types across species has been achieved, the computational complexity is reduced, and it is suitable for cell type identification of millions of large-scale single-cell data, and an online platform for use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115064220B_ABST
    Figure CN115064220B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for cross-species cell type identification of a single cell, which belongs to the field of single-cell data analysis. The present invention integrates the full map of single-cell transcriptomes of multiple representative species, and constructs the largest single-cell reference database to date, which contains a variety of cell types without bias and with relatively comprehensive coverage. The present invention makes up for the defects of cell identification based on population gene expression data and the trouble of no standard reference database, and provides data support for cross-species analysis. Through homologous gene alignment, characteristic gene extraction and calculation of gene expression pattern similarity, the present invention eliminates the influence of single-cell data sequencing depth between different sequencing platforms, different species and even different cell types, and strikes a balance between high accuracy and fast calculation, and realizes the identification of cross-species cell types for the first time. The method provided by the present invention has important guiding significance for the identification of unknown cell types and cross-species analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of single-cell data analysis, relates to the processing and analysis of single-cell biological information data, and particularly relates to a cross-species cell type identification method for single cells. Background Art

[0002] Currently, large-scale single-cell sequencing has been widely used in various studies, and related single-cell analysis processes and tools have also rapidly developed and matured. The first step in single-cell data analysis is to perform raw processing on the sequencing data and process the sequencing sequence files into a digital gene expression (DGE) matrix. After obtaining the DGE matrix file, downstream analysis of the single-cell sequencing data can be performed.

[0003] Cell type recognition and identification is one of the most important issues in single-cell transcriptome sequencing (scRNA-seq) data analysis. At present, a variety of analysis tools such as Seurat, Monocle and scanpy perform unsupervised clustering of cells and calculate the genes specifically expressed by each cell group, and artificially identify cell types through a combination of marker genes. However, this artificial identification and identification of cell types is somewhat subjective. Although some studies have reported the identification of engineered cell types based on gene expression of bulk RNA-seq, these studies ignore the heterogeneity of cells within the population, and the reference cell types and gene expression profiles are not perfect, making them difficult to apply to accurate identification at the single-cell level. Therefore, tools for cell type identification at the single-cell level need to be developed.

[0004] Although some cell type identification and query tools have been developed, such as CellBlast, scmap, CellFishing.jl, CellAtlasSearch, etc. These tools can be divided into two categories according to the calculation method. They only directly find similar cell types from the single-cell gene expression spectrum, or rely on machine learning methods to generate models to predict the similarity between cells. However, both methods are highly dependent on reference databases. There is still no large-scale reference database with unified and comparable annotations. Users often need to search for suitable reference data by themselves, which poses a challenge to the accuracy of the comparison results. In addition, these methods are usually developed and tested based on smaller samples. For the current single-cell data of hundreds of thousands or even millions, the calculation accuracy and efficiency are difficult to guarantee. At the same time, some methods lack online tools and visual output results, which is not conducive to users using the tool.

[0005] Cell types are considered to be evolutionary units with independent evolutionary potential. The comparison and evolution of cell types are still limited to specific tissues, and systematic comparative analysis has not yet been performed in the whole species map. With the development of single-cell technology, the mapping of more and more single-cell maps of species has provided an unprecedented opportunity to overcome this biological problem. However, current cell type identification tools can only identify cell types of the same species based on reference databases, and cannot perform cross-species identification and analysis. This poses a challenge to the analysis and comparison of cross-species cell types, as well as the identification of cell types in some single-cell data sets that lack reference databases. Summary of the invention

[0006] In response to the deficiencies in the above-mentioned prior art, the present invention provides a single-cell cross-species cell type identification method. This method collects high-quality data resources without bias and covering comprehensive cell types from single-cell atlas data of multiple species, making up for the defects of cell identification based on population gene expression data and the trouble of no standard reference database, and providing data support for cross-species analysis. Through homologous gene alignment, characteristic gene extraction, and comparison of gene expression pattern similarity, cross-species cell type identification was achieved for the first time. The feature extraction method greatly reduces the computational complexity, making the present invention suitable for cell type identification of large-scale single-cell data of millions. An online use platform has been developed to promote and use the present invention.

[0007] The technical solution of the present invention is as follows:

[0008] The present invention provides a method for cross-species cell type identification of a single cell, comprising the following steps:

[0009] (1) Collect single-cell transcriptome atlas datasets of different species, organize cell types according to dataset annotation information, and unify naming methods and rules;

[0010] (2) Perform RPKM normalization on the gene expression matrix of single-cell atlas data of different species to obtain the standardized expression value of each gene in different cells; for each cell type, perform differential gene expression analysis based on the standardized expression value of the gene to obtain the characteristic genes that are specifically highly expressed in each cell type compared with other cell types;

[0011] (3) For each species, select the top several differentially expressed genes in each cell type and combine them into a characteristic gene set used as a reference data set, which is the cell type reference database;

[0012] (4) Identification of cell types in the test dataset:

[0013] (4.1) Collect annotated homologous genes between the species to be tested; for unknown homologous genes between the species to be tested, download the annotated gene coding regions of the species to be tested from the public database, or predict the coding regions of unknown homologous genes; input the gene coding region file of the species to be tested, and predict the homologous genes of the species to be tested;

[0014] (4.2) Calculation of cell type similarity score:

[0015] According to step (2), the test data set is standardized and logarithmically transformed;

[0016] According to the homologous genes between species in step (4.1), the expression matrix of the test data set is converted into homologous genes to obtain a new gene expression matrix;

[0017] Based on the cell type reference database of the query target species, the expression values ​​of the characteristic gene set are extracted; the similarity score of the cell type between the expression matrix of the test data set and the query target of the cell type reference database is calculated; based on the cell information contained in each cell type, the similarity score between the cell type in the test data set and the cell type in the cell type reference database is counted, and the cell type with the highest similarity score is used as the identification result of the cell type of the test data set.

[0018] Preferably, there are at least 15 species in step (1), covering species from lower to higher. The present invention has collected more than six million single-cell data, covering 15 species from lower to higher that are representative of multicellular animal evolution, and is the most comprehensive single-cell reference database to date. The whole-cell maps of these species cover most of the tissues and organs of the species and a relatively comprehensive range of cell types, and the data of the same species come from the same laboratory and use the same sequencing method, so that the collected cell types are non-biased and have a relatively comprehensive coverage.

[0019] Specifically, the single-cell transcriptome atlas dataset in step (1) is publicly available, and the dataset is divided into eight major lineages according to major cell categories: neural cells, muscle cells, immune cells, endothelial cells, epithelial cells, secretory cells, stromal cells, and proliferating cells; those not covered are marked as others.

[0020] Preferably, the calculation formula of the gene expression matrix in step (2) is:

[0021]

[0022] Where: The normalized expression value of the jth gene in the i-th cell in the gene expression matrix;

[0023] The number of unique identifiers for the jth gene in the ith cell in the gene expression matrix.

[0024] In step (2), the single-cell atlas data from different sequencing platforms are standardized by RPKM (Reads Per Kilobase per Million mapped reads), and 100 cells are randomly sampled three times to calculate the average value, so that the three sampling average data of each cell type have the expression of about 10,000 genes. Considering the sequencing depth of different platforms and the difference in the number of captured genes of different cell types on the same platform, the single-cell data sequencing depth and cumulative gene number curve are plotted, and it is found that the number of genes reaches a balance at about 10,000. The standardization and random sampling steps can eliminate the influence of the sequencing depth of single-cell data between different sequencing platforms, different species and even different cell types.

[0025] Specifically, 100 cells were randomly sampled from each cell type without replacement, and the results were repeated three times to obtain the average expression value of the gene, and the average value data of the three samples were obtained; if the total number of cells in a cell type was less than 100, all cells were selected;

[0026] Seurat software was used to take the average data of three samples of each cell type as the input gene expression matrix, and the seurat file was generated using the CreateSeuratObject function. The data was converted to log2(TPM / 100+1) using the NormalizeData function, and linearly transformed using the ScaleData function. The FindAllCluster function was run and the Wilcoxon rank sum test method was used to calculate the differentially expressed genes of each cell type. Finally, the generated gene list was sorted from high to low according to cell type, p_val_adj, and avg_log2FC.

[0027] Preferably, in step (3), when obtaining the reference expression value of gene expression of each cell type, the three sampling averages of each cell type are rounded, the rounded values ​​are averaged again, and logarithmic transformation is performed to further reduce the difference in sequencing depth, and finally the average standardized expression value of the characteristic gene set after logarithmic transformation is retained as the cell type reference database.

[0028] Comparative analysis showed that when the top 10 differentially expressed genes of each cell type were taken to establish a reference database, the comparison efficiency reached a high and stable level, achieving a balance between high accuracy and fast calculation, making the present invention suitable for cell type identification of large-scale single-cell data of millions.

[0029] Preferably, the predicted data set is a unique identifier, number of reads per kilobase of transcript per million mapped reads, or type of fragments per kilobase of transcript per million mapped reads.

[0030] Preferably, in step (4.1), one-to-one orthologous gene comparison is used for homologous gene comparison. In order to further improve the accuracy and efficiency of homologous gene comparison, the present invention only considers one-to-one homologous genes for further analysis, and discards one-to-many or many-to-many homologous genes.

[0031] Preferably, in step (4.2), when the expression of certain characteristic genes is not detected in the test data set, the expression values ​​of these characteristic genes in the sample are counted as zero, so that the test data set and the reference data set have the same number of rows.

[0032] Preferably, in step (4.2), the similarity score is the Pearson correlation coefficient.

[0033] The method of the present invention can be implemented in the online platform "http: / / bis.zju.edu.cn / cellatlas / annotation.html".

[0034] The present invention integrates the complete transcriptome maps of multiple representative species and constructs the largest single-cell reference database to date, which contains a variety of cell types without bias and with relatively comprehensive coverage. The present invention makes up for the defects of cell identification based on population gene expression data and the trouble of no standard reference database, and provides data support for cross-species analysis. Through homologous gene alignment, characteristic gene extraction and calculation of gene expression pattern similarity, the present invention eliminates the influence of single-cell data sequencing depth between different sequencing platforms, different species and even different cell types, and strikes a balance between high accuracy and fast calculation, and realizes the identification of cross-species cell types for the first time. The method provided by the present invention has important guiding significance for the identification of unknown cell types and cross-species analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flow chart of the technical route of the present invention.

[0036] Figure 2 This is a diagram showing the comparison results of the human and mouse cell maps predicted by the present invention.

[0037] Figure 3 This is a diagram showing the cross-species cell type identification results of the human adult kidney atlas predicted by the present invention.

[0038] Figure 4 This is a diagram showing the results of cross-species cell type identification of zebrafish intestinal epithelial cells predicted by the present invention. DETAILED DESCRIPTION

[0039] Based on the single-cell transcriptome profiles of multiple species, the present invention constructs a single-cell reference database, and predicts cross-species cell types by comparing homologous genes, extracting characteristic genes, and calculating the similarity of gene expression patterns. The following steps are included: Figure 1 As shown:

[0040] (1) Single-cell atlas data collection

[0041] Published single-cell transcriptome atlas datasets were collected from publicly available literature, including the Human Cell Atlas (~590,000 cells), the Mouse Development and Aging Cell Atlas (~1.12 million cells), the Crab-eating Macaque Cell Atlas (~1 million cells), the Rat Senescent Cell Atlas (~160,000 cells), the Zebrafish Developmental Cell Atlas (~1.03 million cells), the Drosophila Cell Atlas (~430,000 cells), the Salamander Cell Atlas (~1.07 million cells), the Xenopus Metamorphosis Cell Atlas (~500,000 cells), the Ciona Intestinalis Cell Atlas (~90,000 cells), the Earthworm Cell Atlas (~90,000 cells), the Nematode Cell Atlas (~50,000 cells), the Planarian Cell Atlas (~60,000 cells), the Star Anemone Cell Atlas (~10,000 cells), the Hydra Cell Atlas (~20,000 cells) and the Sponge Cell Atlas (~10,000 cells). A total of more than six million single-cell data were collected, covering 15 species from low to high levels that are representative of the evolution of multicellular animals. The cell types were sorted according to the annotation information of the dataset, and the naming methods and rules were unified. They were divided into 8 major lineages according to the major cell categories, including nerve cells, muscle cells, immune cells, endothelial cells, epithelial cells, secretory cells, matrix cells and proliferating cells. Those not covered were marked as others. At the same time, the period, gender, tissue and other information of each cell source sample in the dataset were statistically analyzed.

[0042] (2) Graph Data Feature Extraction

[0043] The single-cell atlas data of various species from different sequencing platforms were standardized by RPKM (Reads Per Kilobase per Million mapped reads), and the gene expression matrix was processed according to the following calculation formula:

[0044]

[0045] Where: The normalized expression value of the jth gene in the i-th cell in the gene expression matrix;

[0046] The number of unique identifiers for the jth gene in the ith cell in the gene expression matrix.

[0047] 100 cells were randomly selected and averaged, so that the three sampling average data of each cell type have the expression of about 10,000 genes. The operation process is: 100 cells were randomly sampled from each cell type without replacement (if the total number of cells in a cell type is less than 100, all cells were selected), repeated 3 times, and the average expression value of the gene was calculated to obtain the average data of 3 samples for subsequent analysis.

[0048] For each cell type, the data of three sampling averages were used to perform differential gene expression analysis to obtain the characteristic genes with high expression in each cell type. The three sampling averages of each cell type were used as the input gene expression matrix using Seurat software, and the CreateSeuratObject function was used to generate the seurat file. The data were converted to log2 (TPM / 100+1) using the NormalizeData function, and the data were linearly transformed using the ScaleData function. The FindAllCluster function was run and the Wilcoxon rank sum test method was used to calculate the differentially expressed genes of each cell type. Finally, the generated gene list was sorted from high to low according to cell type and avg_log2FC (average expression log difference fold).

[0049] (3) Establishment of cell type reference database

[0050] For each species, the top 10 differentially expressed genes (avg_log2FC top 10) of each cell type were selected and merged as the characteristic gene set used as reference data. Comparative analysis shows that when the top 10 differentially expressed genes of each cell type are taken to establish a reference database, the alignment efficiency reaches a high and stable level, achieving a balance between high accuracy and fast calculation, making the present invention suitable for cell type identification of millions of large-scale single-cell data. In order to better obtain the reference expression value of gene expression of each cell type, we rounded the 3 sampling averages of each cell type, and then averaged the rounded values ​​again, and performed logarithmic transformation to further reduce the difference in sequencing depth, and finally retained the average standardized expression value of the characteristic gene set after logarithmic transformation as the cell type reference database.

[0051] (4) Homologous gene comparison

[0052] Annotated homologous genes between species were collected from BioMart of the Ensembl database. For unknown homologous genes between species, the annotated gene coding regions were downloaded from the Ensembl database, or the coding regions of the genes were predicted using TransDecoder; then, the gene coding region files of the two species were used as inputs to OrthoFinder to predict the homologous genes of the two species. In order to further improve the accuracy and efficiency of homologous gene comparison, the present invention only considers one-to-one homologous genes for further analysis, and discards one-to-many or many-to-many homologous genes.

[0053] (5) Calculation of cell type similarity scores.

[0054] For the dataset to be tested, it can be a unique molecular identifier (UMI), reads per kilo base per million mapped reads (RPKM), or fragments per kilobase per million mapped fragments (FPKM).

[0055] According to step (1), the test data set is standardized and logarithmically transformed. According to the orthologous genes between the two species in step (4), the expression matrix of the test data set is transformed into homologous genes to obtain a new gene expression matrix, which is transformed into the genes of the query target species and listed as the cells of the test data set.

[0056] Based on the query of the cell type reference database of the target species, the expression values ​​of the characteristic gene set are extracted. If the expression of certain characteristic genes is not detected in the test data set, the expression values ​​of these characteristic genes in the sample are counted as zero, so that the test data set and the reference data set have the same number of rows.

[0057] The Pearson correlation coefficient between the expression matrix of the test data and the query target of the cell type reference database is calculated, and the correlation coefficient is the cell type similarity score. Based on the cell clustering information, that is, the cell information contained in each cell type, the similarity scores between the cell types in the test data set and the cell types in the cell type reference database are counted, and the cell type with the highest similarity score is used as the identification result of the cell type of the test data set.

[0058] The method of the present invention is described in detail below through examples.

[0059] Example 1: Comparison of human and mouse cell maps

[0060] By mapping the human cell atlas to the mouse development and aging cell atlas through one-to-one orthologous gene conversion, it was found that similar cell types in humans and mice showed the highest similarity ( Figure 2 Left). The eight cell lineages, secretory, epithelial, immune, neural, content, muscle, stromal, erythroid, and proliferative, are basically the same in both humans and mice. For example, C33, C42, and C79 are annotated as smooth muscle cells in the Human Cell Atlas, and they are highly related to the smooth muscle cell types in the Mouse Development and Aging Cell Atlas ( Figure 2 Right). Notably, C74 resembles ventricular cardiomyocytes from fetal and neonatal hearts in the Mouse Developing and Aging Cell Atlas ( Figure 2 Right). This cell type is annotated as a ventricular cardiomyocyte in the Human Cell Atlas, of which 99.6% are derived from fetal hearts. The accuracy of cell type and cell period further demonstrates the effectiveness of the present invention in identifying cell types across species.

[0061] Example 2: Human Adult Kidney Atlas Cross-species Cell Type Identification

[0062] Through one-to-one orthologous gene conversion, the human adult kidney map was mapped to the mouse development and aging cell map, and various parenchymal cells in the kidney were well annotated, including Henle loop, intercalated cells, principal cells, proximal tubules, etc. ( Figure 3 ).

[0063] Example 3: Cross-species cell type identification of zebrafish intestinal epithelial cells

[0064] Through one-to-one orthologous gene conversion, zebrafish intestinal epithelial cells were mapped to the mouse development and aging cell atlas. Common cell types in the intestinal epithelium were well matched with the corresponding cell lineages in the mouse atlas, including intestinal secretory cells, ionocytes, intestinal epithelial cells, epidermal cells, goblet cells, mesenchymal cells, etc. ( Figure 4 ).

Claims

1. A method for cross-species cell type identification of a single cell, characterized in that: The following steps are involved: (1) Collect single-cell transcriptome atlas datasets of different species, organize cell types according to dataset annotation information, and unify naming methods and rules; (2) Perform RPKM normalization on the gene expression matrix of single-cell atlas data of different species to obtain the standardized expression value of each gene in different cells; for each cell type, perform differential gene expression analysis based on the standardized expression value of the gene to obtain the characteristic genes that are specifically highly expressed in each cell type compared with other cell types; (3) For each species, select the top several differentially expressed genes in each cell type and combine them into a characteristic gene set used as a reference data set, which is the cell type reference database; (4) Identification of cell types in the test dataset: (4.1) Collect annotated homologous genes between the species to be tested; for unknown homologous genes between the species to be tested, download the annotated gene coding regions of the species to be tested from the public database, or predict the coding regions of unknown homologous genes; input the gene coding region file of the species to be tested, and predict the homologous genes of the species to be tested; (4.2) Calculation of cell type similarity score: According to step (2), the test data set is standardized; According to the homologous genes between species in step (4.1), the expression matrix of the test data set is converted into homologous genes to obtain a new gene expression matrix; Based on the cell type reference database of the query target species, the expression values ​​of the characteristic gene set are extracted; the similarity score of the cell type between the expression matrix of the test data set and the query target of the cell type reference database is calculated; based on the cell information contained in each cell type, the similarity score between the cell type in the test data set and the cell type in the cell type reference database is counted, and the cell type with the highest similarity score is used as the identification result of the cell type of the test data set.

2. The identification method according to claim 1, characterized in that The species in step (1) include at least 15 species, ranging from lower to higher species.

3. The identification method according to claim 1, characterized in that: The single-cell transcriptome atlas dataset in step (1) is publicly available. The dataset is divided into eight major lineages according to the major cell categories: neural cells, muscle cells, immune cells, endothelial cells, epithelial cells, secretory cells, stromal cells and proliferating cells; those not covered are marked as others.

4. The identification method according to claim 1, characterized in that: The calculation formula of the gene expression matrix in step (2) is: Where: The normalized expression value of the jth gene in the i-th cell in the gene expression matrix; The number of unique identifiers for the jth gene in the ith cell in the gene expression matrix.

5. The identification method according to claim 4, characterized in that: Randomly sample 100 cells from each cell type without replacement, repeat 3 times, calculate the average expression value of the gene, and obtain the average value data of 3 samples; if the total number of cells in a cell type is less than 100, select all cells; Seurat software was used to take the average data of three samples of each cell type as the input gene expression matrix, and the seurat file was generated using the CreateSeuratObject function. The data was converted to log2(TPM / 100+1) using the NormalizeData function, and linearly transformed using the ScaleData function. The FindAllCluster function was run and the Wilcoxon rank sum test method was used to calculate the differentially expressed genes of each cell type. Finally, the generated gene list was sorted from high to low according to cell type, p_val_adj, and avg_log2FC.

6. The identification method according to claim 1, characterized in that: In step (3), when obtaining the reference expression value of each cell type gene, the three sampling averages of each cell type are rounded, the rounded values ​​are averaged again, and logarithmic transformation is performed to further reduce the difference in sequencing depth. Finally, the average standardized expression value of the characteristic gene set after logarithmic transformation is retained as the cell type reference database.

7. The identification method according to claim 1, characterized in that: In step (4), the predicted data set is a unique identifier, the number of reads per kilobase of transcript per million mapped reads, or the fragment type per kilobase of transcript per million mapped reads.

8. The identification method according to claim 1, characterized in that: In step (4.1), a one-to-one orthologous gene alignment is used for homologous gene alignment.

9. The identification method according to claim 1, characterized in that: In step (4.2), when the expression of certain characteristic genes is not detected in the test data set, the expression values ​​of these characteristic genes in the test data set are counted as zero, so that the test data set and the reference data set have the same number of rows.

10. The identification method according to claim 1, characterized in that: In step (4.2), the similarity score is the Pearson correlation coefficient.

Citation Information

Patent Citations

  • Cross-species analysis method, equipment and medium for accessibility map of single-cell chromatin

    CN117542414A

  • High-throughput single-cell libraries and methods of making and of using

    US20220356461A1