Single cell marker gene identification method based on expression accumulation curve area
By calculating the score of single-cell marker genes based on the expression cumulative curve area, the scores of single-cell marker genes and screening out marker genes with high specificity were solved, and the problem of inadequate expression of marker genes in the prior art was solved, achieving stronger marker gene specificity and better downstream analysis capabilities.
Patent Information
- Application Number
- CN202311685828.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
In the existing single-cell marker gene identification methods, the highest-scoring marker genes are not expressed specifically in different cell types, making it difficult to perform downstream analysis as characteristic genes of certain cell populations.
Using a method based on the expression accumulation curve area, the cumulative curve area of a single gene in different cell types is calculated, and the gene scores are obtained, and marker genes that express specific expression in specific cell types are screened.
The expression specificity of marker genes in different cell classifications has been improved, making the screened marker genes more suitable for cell isolation screening and downstream analysis as characteristic genes of certain cell populations.
Smart Images

Figure BDA0004597637350000021 
Figure BDA0004597637350000071 
Figure FDA0004597637200000011
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and particularly to a method for identifying single-cell marker genes based on the area of the expression cumulative curve. Background Art
[0002] Marker genes can not only intuitively reflect the quality of cell clustering, but also help to deeply explore biological problems. For unknown cell types, identifying their marker genes helps to quickly and accurately annotate their categories; for known cell types, through similar analysis, some new marker genes can be discovered, providing more molecular markers for subsequent research.
[0003] Currently, commonly used methods for identifying marker genes include Seurat, Scanpy, sran, and COSG, etc. These methods mainly perform analysis based on t-tests, permutation tests, etc. The marker genes identified by these methods, especially the top few (usually the top 5) with the highest scores, are not specific enough in their expression in different cell types, and even show uniform expression in each type of cell, which is not conducive to using these genes as characteristic genes of certain cell populations for downstream analysis. Therefore, it is crucial to develop an identification method that can specifically identify marker genes. Summary of the Invention
[0004] In view of this, the technical problem to be solved by the present invention is to provide a method for identifying single-cell marker genes based on the area of the expression cumulative curve.
[0005] The present invention provides a method for identifying marker genes, including gene scoring and screening. The gene scoring includes the following steps:
[0006] Step 1: According to the clustering data, obtain the expression level information and cell classification information of genes in cells. Set the gene expression level threshold. After sorting the cells in descending order of gene expression level, calculate the cumulative expression curve A of a single gene A in cell type K k 、the cumulative expression curve A of a single gene A in other cell types outside cell type K k_except and the cumulative expression curve A of a single gene A in all cells all ;
[0007] Step 2: According to the following formula, obtain the score of a single gene A,
[0008] The formula is as follows:
[0009]
[0010] Wherein,
[0011] The single gene A is any one of the genes in the cell;
[0012] The cell type K is any one of the cell classifications.
[0013] Furthermore,
[0014] The gene expression level threshold is 0.
[0015] The screening criteria: Sort the scoring values score of all genes in each cell classification from high to low, and select the genes with high scoring values score in the corresponding cell classification.
[0016] Generally, the higher the scoring value score, the more likely it is to be a major marker in the corresponding cell type. In the present invention, the scoring value score corresponding to each gene in each cell classification of the clustering data is calculated, sorted according to the scoring value score (from high to low), and the top N single genes with the highest scores are taken as the marker genes of this cell type. The N includes 1, 2, 3, 4, 5, 6, 7, etc., and can be selected according to specific circumstances.
[0017] Even further,
[0018] In a specific embodiment of the present invention, the clustering data is the preprocessed and clustered pbmc3k data imported from the data of scanpy; in the present invention, the clustering data can also be the self-preprocessed and clustered pbmc3k data imported from the data of scanpy. The detailed steps of the preprocessing and clustering of the pbmc3k data can be referred to
[0019] https: / / scanpy-tutorials.readthedocs.io / en / latest / pbmc3k.html.
[0020] Clustering is to group cells with similar gene expression patterns for further study of the functions and differentiation states of different cell types. In the present invention, the software for performing clustering analysis on the data includes any one of Seurat, Scanpy, and / or SingleR. In a specific embodiment of the present invention, the preprocessed and clustered pbmc3k data in Scanpy is used.
[0021] The method for identifying the marker genes of the present invention is applicable to most publicly available single-cell, spatial transcriptomics, merfish, etc. data. By analyzing the above data using the identification method of the present invention, the specificity of the obtained marker genes is better than that of marker gene identification methods such as scanpy.
[0022] Furthermore, the single-cell, spatial transcriptomics, merfish, etc. data include V1_Adult_Mouse_Brain_Coronal_Section_2, PBMC10K, MERFISH (Conservation and divergence of cortical cell organization in human and mouse revealed by MERFISH), BGI spatial transcriptomics axolotl brain axolotlbrain 10DPI_1_left, etc.
[0023] The method for identifying marker genes provided by the present invention, compared with the existing scanpy and COSG methods, the obtained marker genes have stronger expression specificity in different cell classifications, and are more conducive to being used as characteristic genes of certain cell populations for cell separation screening and downstream analysis. And for data with large differences between cell types, the advantages of this method will be more obvious. The greater the difference between cell types, the more significant the difference in the cumulative curve area, and the more specific the expression of the screened Marker in specific cell types.
[0024] The present invention provides the marker genes obtained by identifying with the said identification method.
[0025] The present invention provides software and / or programs for identifying single-cell marker genes, and the execution code thereof includes part or all of the identification methods as described in the present invention.
[0026] The identification method described in the present invention is an algorithm for screening single-cell marker genes or the steps and rules for screening single-cell marker genes, which are transformed into an instruction set and / or programming language that can be read and executed by a computer, namely code; through the code, the said algorithm is transformed into instructions that can be executed by a computer, and the code uses the syntax and rules of the programming language to describe the algorithm or identification method described in the present invention.
[0027] The present invention provides a storage medium, which stores or burns the software and / or programs as described in the present invention. Specifically, the storage medium includes: various media that can store program codes such as USB flash drives, external hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, semiconductors, magnetic cores, magnetic drums, or optical discs, and the present invention does not make any limitations thereto.
[0028] The present invention provides an electronic device,
[0029] which can execute the software and / or programs as described in the present invention; and / or
[0030] is provided with the storage medium as described in the present invention.
[0031] Specifically, the electronic device may include one or more of a memory, a server, a network device, a computer, etc., and includes a system that assists in reading or executing the code and / or program in the storage medium.
[0032] When the electronic device runs the authentication method of the present invention, the operating system reads the program code from the storage medium and loads it into the memory of the electronic device. In the memory, the program code is converted into machine language instructions, and the CPU executes these instructions to complete the functions to be achieved by the program.
[0033] The present invention provides a platform for marker gene identification, which includes at least one of a sample processing platform and / or a data preprocessing platform and at least one of the following i) to iii):
[0034] i), the software and / or program of the present invention;
[0035] ii), the storage medium of the present invention;
[0036] iii), the electronic device of the present invention.
[0037] The platform of the present invention may be fully automated or semi-automated, and may be integral or integrated in parts. The present invention does not limit this.
[0038] The platform of the present invention may be fully automated or semi-automated. The present invention does not limit this.
[0039] In the present invention, the sample processing platform includes a sample pretreatment platform, a sample sequencing platform, etc.
[0040] The present invention provides the application of at least one of the following a) to f) in single-cell screening:
[0041] a), the authentication method of the present invention;
[0042] b), the marker gene of the present invention;
[0043] c), the software and / or program of the present invention;
[0044] d), the storage medium of the present invention;
[0045] e), the electronic device of the present invention;
[0046] f), the platform of the present invention.
[0047] The present invention provides a method for single-cell screening, including performing single-cell screening by using any one of the following A) to F):
[0048] A), the identification method described in the present invention;
[0049] B), the marker gene described in the present invention;
[0050] C), the software and / or program described in the present invention;
[0051] D), the storage medium described in the present invention;
[0052] E), the electronic device described in the present invention;
[0053] F), the platform described in the present invention.
[0054] The present invention provides an identification method for marker genes based on the area of the expression cumulative curve of single cells. Compared with other existing marker gene identification methods, the identified marker genes have stronger expression specificity in different cell classifications, and are more conducive to being used as characteristic genes of certain cell populations for cell separation screening and downstream analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 shows the situation of marker genes identified by the identification method of the present invention;
[0056] Figure 2 shows the situation of marker genes identified by the scanpy identification method;
[0057] Figure 3 shows the situation of marker genes identified by COSG. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The present invention provides an identification method for single-cell marker genes based on the area of the expression cumulative curve. Those skilled in the art can draw on the content of this article and appropriately improve the process parameters to achieve it. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art, and they are all considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and those related can obviously make changes or appropriate alterations and combinations to the methods and applications in this article without departing from the content, spirit and scope of the present invention to implement and apply the technology of the present invention.
[0059] The technical solution of the present invention:
[0060] 1. Calculate the area of the expression cumulative curve of the gene
[0061] (a), Extract the expression level and classification information of the gene in all cells. Assume that the total number of cell classifications is N, and count the total number of cells in each cell type, denoted as C0 , C 1 , ..., C N-1 ;
[0062] (b), Sort all cells according to the gene expression level from high to low, and then calculate the area of the cumulative curve:
[0063] i. If a gene is expressed in a cell (belonging to cell type i) (the expression level value is greater than the set threshold, generally 0), then accumulate 1 / C i ;
[0064] ii. For all cell types k (k ∈ [0, N - 1]), calculate the cumulative curves of cell type k and all cells outside cell type k respectively according to i), and the area of its cumulative curve is denoted as A k and A k_except ;
[0065] iii. Calculate the cumulative curve of all cells according to i), and the area of its cumulative curve is denoted as A all ;
[0066] iv. The final score is
[0067] (2), Calculate the scores of all genes according to 1), and obtain the scoring results of all genes;
[0068] (3), Select the top 5 (or top 10) results in each class for Dotplot result display.
[0069] The reagents and consumables used in the present invention are all ordinary commercially available products and can be purchased on the market.
[0070] The following further elaborates the present invention in conjunction with embodiments:
[0071] Example 1 Single-cell marker gene identification method based on the area of the expression cumulative curve
[0072] I. The analysis steps of the present invention are as follows
[0073] 1. Construct a gene expression matrix:
[0074] Download the PBMC3K data from the 10X Genomics official website, and then import the pre-clustered pbmc3k data from scanpy's datasets (the execution code is scanpy.datasets.pbmc3k_processed()). Finally, obtain the expression information of 1,838 genes in 2,638 cells, which are divided into 8 categories (B cells, CD14+ Monocytes, CD4 T cells, CD8 T cells, Dendritic cells, FCGR3A+ Monocytes, Megakaryocytes, NK cells);
[0075] This invention directly imports the pre-clustered pbmc3k data. In general procedures, the clustering of pbmc3k data can also be completed according to the following steps (the main steps of pbmc3k data clustering);
[0076] a) Import the pbmc3k data downloaded from the 10X official website
[0077] adata = sc.read_10x_mtx('data / filtered_gene_bc_matrices / hg19 / ', var_names='gene_symbols', cache=True)
[0078] b) Filter out cells with less than 200 expressed genes and genes expressed in less than 3 cells
[0079] sc.pp.filter_cells(adata, min_genes=200)
[0080] sc.pp.filter_genes(adata, min_cells=3)
[0081] c) Normalize the expression data
[0082] sc.pp.normalize_total(adata, target_sum=1e4)
[0083] sc.pp.log1p(adata)
[0084] d) PCA and clustering
[0085] sc.tl.pca(adata, svd_solver='arpack')
[0086] sc.pp.neighbors(adata, n_neighbors = 10, n_pcs = 40)
[0087] sc.tl.leiden(adata)
[0088] e), Identify marker genes and add cell type labels
[0089] sc.tl.rank_genes_groups(adata, 'leiden', method = 'wilcoxon')
[0090] new_cluster_names = ['CD4 T', 'CD14 Monocytes', 'B', 'CD8 T', 'NK', 'FCGR3A Monocytes', 'Dendritic', 'Megakaryocytes']
[0091] adata.rename_categories('leiden', new_cluster_names)
[0092] For the detailed steps of clustering the pbmc3k data, please refer to
[0093] https: / / scanpy-tutorials.readthedocs.io / en / latest / pbmc3k.html .
[0094] 2. For the above expression matrix, loop by column (gene) and record the scoring results of each gene:
[0095] a) Generate an expression matrix for a single gene, with rows as cells, columns as expression levels and cell classifications, and count the total number of cells in each class, denoted as C 0 , C 1 ,..., C N-1 ;
[0096] b) Sort the cells by expression level from high to low and calculate the area under the cumulative curve;
[0097] i. If the gene is expressed in a certain type of cell (belonging to cell type i) (expression level value is greater than the set threshold, generally 0), then accumulate 1 / C i ;
[0098] ii. For all cell types k (k ∈ [0, N - 1]), calculate the cumulative curves of cell type k and all cells outside cell type k respectively according to i), and the area under the cumulative curve is denoted as A k and A k_except ;
[0099] iii. Calculate the cumulative curve for all cells according to i), and denote the area of the cumulative curve as A all ;
[0100] iv. The final score of the cumulative curve area of this gene is
[0101]
[0102] 3. Screen the top 5 genes with the highest scores in each class and perform a Dotplot display.
[0103] 4. The results are as Figure 1 shown
[0104] II. Comparison with scanpy and COSG analysis and identification methods
[0105] Import the clustered pbmc3k data from scanpy's datasets (the execution code is scanpy.datasets.pbmc3k_processed()), which contains the expression information of 1838 genes in 2638 cells. Then, use the default parameters of rank_genes_groups in scanpy to identify marker genes for the above data. The marker genes found by scanpy are as Figure 2 shown. Take the top 3 genes in each class. The expression of these genes is basically relatively uniform and there is no significant specific expression in a certain class.
[0106] The marker genes found by the COSG method are as Figure 3 shown. The expression of these genes basically has a certain specificity, but the specificity of the marker genes obtained by the identification method of the present invention is stronger.
[0107] Moreover, for data with large differences between cell types, the advantages of this method will be more obvious. The greater the difference between cell types, the more significant the difference in the cumulative curve area, and the more specific the expression of the screened Markers in specific cell types.
[0108] III. Derivative product development
[0109] The identification method for single-cell marker genes described in the present invention can be further developed into software and / or programs. Embody the identification method in the form of code, and the software and / or programs can execute the code to complete the detection of marker genes.
[0110] Further, the software and / or program can be stored or burned on a storage medium, which includes: USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, semiconductor, magnetic core, magnetic drum, or optical disc, etc.
[0111] Furthermore, the storage medium can be set in an electronic device together with other software programs. This electronic device can be integral or separate. The other software programs are used to assist in reading and executing the code for completing the identification of marker genes, thereby realizing the identification of marker genes.
[0112] Specifically, the electronic device can include one or more of a memory, a server, a network device, a computer, etc.
[0113] The present invention can further develop a platform for marker gene identification, which includes the storage medium or electronic device described in the present invention, and some other platforms for single-cell marker gene identification; some other platforms for single-cell marker gene identification include: a sample processing platform and / or a data preprocessing platform, etc. Each component of the platform can be integral, separate, or part of it is separate and part of it is integral; the platform can be fully automatic or semi-automatic.
[0114] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. Method for identifying marker genes, characterized in that, it includes gene scoring and screening, and the gene scoring includes the following steps: Step 1. According to the clustering data, obtain the gene expression level information and cell classification information within the cells, set the gene expression level threshold, sort the cells in descending order of gene expression level, and then calculate the cumulative curve A of the expression level of a single gene A in cell type K k , the cumulative curve A of the expression level of a single gene A in other cell types outside cell type K k_except and the cumulative curve A of the expression level of a single gene A in all cells all ; Step 2: Obtain the score of a single gene A according to the following formula, the formula is as follows: wherein, the single gene A is any one of the intracellular genes; the cell type K is any one of the cell classifications.
2. The identification method according to claim 1, characterized in that, the gene expression threshold is 0.
3. The identification method according to claim 1, characterized in that, the screening criterion: sort the scoring values score of all genes in each cell classification from high to low, and select the genes with high scoring values score in the corresponding cell classification.
4. Marker genes identified by the identification method according to any one of claims 1 to 3.
5. Software and / or program for identifying single-cell marker genes, characterized in that, the execution code includes part or all of the identification methods according to any one of claims 1 to 3.
6. Storage medium, characterized in that, it stores the software and / or program according to claim 5.
7. Electronic device, characterized in that, it can execute the software and / or program according to claim 5; and / or it is provided with the storage medium according to claim 6.
8. Platform for identifying marker genes, characterized in that, it includes at least one of a sample processing platform and / or a data preprocessing platform and at least one of the following i) to iii): i) The software and / or program according to claim 5; ii) The storage medium according to claim 6; iii) The electronic device according to claim 7.
9. Application of at least one of the following a) to f) in single-cell screening: a) The identification method according to any one of claims 1 to 3; b) The marker gene according to claim 4; c) The software and / or program according to claim 5; d) The storage medium according to claim 6; e) The electronic device according to claim 7; f) The platform according to claim 8.
10. Method for single-cell screening, characterized in that, it includes performing single-cell screening by using any one of the following A) to F): A) The identification method according to any one of claims 1 to 3; B) The marker gene according to claim 4; C) The software and / or program according to claim 5; D) The storage medium according to claim 6; E) The electronic device according to claim 7; F) The platform according to claim 8.