A single cell transcriptome cell fragment and multi-cell filtration method, medium and device
By generating artificial multicellular structures and statistically analyzing their proportions within a neighborhood, and using bimodal coefficients to identify and remove low-quality cells, the problem of insufficient low-quality cell identification in single-cell transcriptome sequencing is solved, thus improving the accuracy and reliability of data filtering.
Patent Information
- Application Number
- CN202310181167.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-11-03
AI Technical Summary
Existing single-cell transcriptome sequencing technologies cannot effectively identify and filter low-quality cells, leading to biased analysis results and even contradictory biological implications.
By generating artificial multicellular structures and statistically analyzing the proportion of each real cell within a predefined neighborhood, the optimal neighborhood is determined using a bimodal coefficient. This process identifies and removes the real cells with the largest proportion of artificial multicellular structures, thereby improving data filtering standards and accuracy.
It enhances the reliability of single-cell transcriptome data, improves the accuracy of data filtering, and reduces the bias of analysis results.
Smart Images

Figure CN116805511B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is a divisional application based on application number 2022113673004, filed on November 3, 2022, entitled "A Method, Medium and Device for Filtering Low-Quality Cells from a Single-Cell Transcriptome". Technical Field
[0003] This invention relates to methods for processing biological data, and more particularly to a method, medium, and apparatus for filtering single-cell transcriptome cell debris and multi-cell data. Background Technology
[0004] Single-cell transcriptome sequencing based on microfluidic technology can quantify gene expression in tens of thousands of cells in a single experiment. It primarily identifies single cells based on sequence tags; its core technology involves adding a unique sequence tag to each cell, treating nucleic acid sequences carrying the same tag as originating from the same cell during sequencing. The 10X Genomics single-cell transcriptome sequencing platform is a widely used technology. This platform utilizes microfluidics, droplet encapsulation, and barcode tagging technologies to achieve high-throughput cell sorting and capture. It can isolate and label 500 to tens of thousands of single cells at once, and obtain transcriptome information for each cell after sequencing. It has advantages such as high cell throughput, low library preparation cost, and short capture cycle.
[0005] A typical single-cell transcriptome sequencing experimental procedure is as follows: First, a cell suspension is prepared. Using a microfluidic chip on a suitable platform instrument, the cell suspension is mixed with magnetic beads and then coated with oil droplets. Each microbead carries a unique nucleotide sequence, i.e., a barcode tag, which can label a single cell. Each barcode tag is also linked to a unique molecular identifier (UMI) composed of nucleotide sequences. Each UMI can label one mRNA transcript. After reverse transcription, PCR amplification, library generation, and sequencing, the sequence data can be analyzed based on the barcode and UMI tags to determine whether each sequence originates from the same cell and the same mRNA. This method reduces the bias of PCR towards different molecules. By matching and counting barcodes and UMIs, gene expression information is summarized in a counting matrix, thereby obtaining the transcriptome expression profile of a single cell.
[0006] Single-cell experiments often rely on the dissociation and fragmentation of biological tissues to obtain single cells in batches, which frequently results in numerous cell fragments or apoptosis. Droplet-based single-cell transcriptomics also involves two or more cells (or intact cells plus cell fragments) forming a single droplet. Single-cell transcriptomics data can contain information from hundreds of thousands or even millions of droplets, but the barcodes within these droplets do not automatically identify whether the droplet contains cells, or whether the contained cells are cell fragments, dead / dying cells, or multiple cells; in other words, they cannot automatically determine the quality of the cells. Cell quality significantly impacts subsequent analysis results, so it is necessary to determine the type of droplet represented by the barcode before data analysis. 10X Genomics' official software, CellRanger, can only determine if the barcode is an empty droplet, but cannot identify cell quality. This may lead to significant discrepancies between single-cell transcriptomics analysis results and actual conditions, even yielding biologically contradictory results. Currently, there is no systematic method for identifying low-quality cells and filtering cells. Summary of the Invention
[0007] To address at least one of the technical problems mentioned in the background art, the present invention aims to provide a method, medium, and device for filtering low-quality cells in single-cell transcriptome data, thereby identifying and filtering low-quality cells in single-cell transcriptome data, improving the filtering standards and accuracy of single-cell transcriptome data, and enhancing the reliability of the data.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A method for filtering low-quality cells from a single-cell transcriptome includes the following steps:
[0010] S101, based on real cell expression profiles, groups cells;
[0011] S1041, the cell expression profile of each cell population is calculated by taking the average expression level of each gene to generate the characteristic expression profile of each cell population.
[0012] S1042, Randomly combine the characteristic expression profiles of the cell population in pairs to generate a certain number of artificial multi-cells;
[0013] S1043, Combine the artificial multi-cell expression profile and the real cell expression profile, and calculate the distance between each cell;
[0014] S1044, Set up several equidistant neighborhoods within a specified range, and calculate the proportion of artificial multicellular cells in each neighborhood for each real cell;
[0015] S1045, statistically analyze the distribution of artificial multicellular proportions in each neighborhood, calculate its bimodal coefficient, and take the neighborhood with the largest bimodal coefficient as the optimal neighborhood.
[0016] S1046, in the optimal neighborhood, identify the specified number of real cells with the largest proportion of artificial multicellular cells as multicellular and remove them from the real cell expression profile.
[0017] Furthermore, the method for pairwise combination of the feature expression spectra is as follows:
[0018] Y = a1*X1 + a2*X2
[0019] Where Y represents the generated artificial multicellular structure, X1 and X2 represent the characteristic expression profiles of the cell population, and a1 and a2 are proportionality coefficients, with one of a1 and a2 set to 1 and the other set to a random value greater than 0 and less than 1.
[0020] Furthermore, the distance between the cells is Euclidean distance or Manhattan distance.
[0021] Furthermore, the artificial multi-cell ratio is the ratio of the number of artificial multi-cells to the total number of cells whose merged expression profile is located in that neighborhood.
[0022] Furthermore, the bimodal coefficient is:
[0023]
[0024] in, This is the bimodal coefficient; and These represent the skewness and kurtosis of the artificial multicellular proportion distribution, respectively. This represents the actual number of cells.
[0025] Furthermore, the method for determining the specified number is as follows: a multi-cell rate is set, and the product of the actual number of cells and the multi-cell rate is the specified number of actual cells identified as multi-celled.
[0026] Furthermore, in S1044, the neighborhood is set as follows: 100 equally spaced neighborhoods are set in the range of 0.0001 to 0.01.
[0027] A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the single-cell transcriptome low-quality cell filtering method as described above.
[0028] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the single-cell transcriptome low-quality cell filtering method as described above.
[0029] Compared with the prior art, the beneficial effects of the present invention are: the present invention can generate a certain amount of artificial multicellular cells, and for each set neighborhood, statistically analyze the distribution of the proportion of artificial multicellular cells of each real cell in the neighborhood to determine the optimal neighborhood. Then, under the optimal neighborhood, the real cells with the largest proportion of artificial multicellular cells are identified as multicellular cells and deleted from the real cell expression profile (low-quality cells), thereby improving the filtering standards and accuracy of single-cell transcriptome data and enhancing the reliability of the data. Attached Figure Description
[0030] Figure 1 This is an overall flowchart of an embodiment of the present invention.
[0031] Figure 2 This is a flowchart of a multi-cell filtration process according to an embodiment of the present invention. Detailed Implementation
[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Example 1:
[0034] Please see Figure 1 This embodiment provides a method for filtering low-quality cells from a single-cell transcriptome. The main steps of the process include steps S100 to S104, which are described in detail below:
[0035] Step S100: The original expression profile (also called the real cell expression profile) may contain non-cellular data. Therefore, the CellRanger software is first used to perform preliminary filtering on the original expression profile to remove non-cellular data and generate a filtered cell expression profile. If the original expression profile does not contain non-cellular data, filtering can be skipped, and the original expression profile can be used directly for the next step.
[0036] Table 1 below shows the original expression profile of this embodiment.
[0037] Table 1: Original Expression Profile
[0038]
[0039] In Table 1 above, C1, C301, C15001, etc. represent cell serial numbers, G1, G2, G53, etc. represent gene serial numbers, and the data in the main text of the table are gene expression levels.
[0040] Table 2 below shows the expression profile of cells filtered in this embodiment.
[0041] Table 2: Expression profile of cells filtered once
[0042]
[0043] Compared to Table 1 above, the single-filtered cell expression profile filtered out cells C10001 and C15001.
[0044] Step S101: Based on the first-filtered cell expression profile or the original expression profile, the cells are normalized, reduced in dimensionality, and clustered using Seurat software. The normalization method is "LogNormalize," with the scale.factor parameter set to 10000; the "vst" method is used for screening hypervariable genes; PCA dimensionality is reduced to 50 PC dimensions; and the clustering is performed using the FindNeighbors function, with the resolution parameter set to 0.8. At this point, all cells are divided into different cell populations. Table 3 below shows the clustering results of the first-filtered cell expression profile in this embodiment:
[0045] Table 3: Clustering results of single-filtered cell expression profiles
[0046]
[0047] In Table 3 above, cells C1 and C601 are classified into cell group A. Similarly, the other cells are classified into cell groups B, C, and D, respectively.
[0048] Filtering cell debris and dead / dying cells is completed in steps S102 to S103. It is worth mentioning that cell debris and dead / dying cells can be filtered together or removed separately.
[0049] Step S102: Set the genes into four gene sets A respectively. gene A mt A active A antioxi And calculate the expressed gene fraction S for each cell population. gene Mitochondrial fraction S mt Activity fraction S active Antioxidant fraction S antioxi .
[0050] Among them, A gene This is the total set of genes in the cell expression profile filtered once (including G1 to G53 in Table 3 above); S gene This is the average number of expressed genes (i.e., the number of genes with an expression level greater than 0) in all cells of the cell population.
[0051] A mtThis refers to the mitochondrial gene set of the corresponding species (including G1 and G2 in Table 3 above); S mt A for all cells in the cell population mt The average gene expression ratio, A mt Gene expression ratio = A in a single cell mt Total gene expression level / A in a single cell gene Total gene expression level * 100%.
[0052] A active This refers to the housekeeping gene set for the corresponding species (including G6 and G7 in Table 3 above); such as the ACTB and GAPDH genes in humans; S active A for all cells in the cell population active The average expression level of gene A active Average gene expression level = A in a single cell active Total gene expression level / A active Number of genes.
[0053] A antioxi This refers to the set of antioxidant genes for the corresponding species (including G20 and G21 in Table 3 above); such as human SOD1, SOD2, SOD3, CAT, GPX1, GPX2, GPX3, GPX4, GPX5, GPX6, GPX7, GPX8, NQO1, and NFE2L2 genes. antioxi A for all cells in the cell population antioxi The average expression level of gene A antioxi Average gene expression level = A in a single cell antioxi Total gene expression level / A antioxi Number of genes.
[0054] In summary, for each cell population, a corresponding gene expression score S is obtained. gene Mitochondrial fraction S mt Activity fraction S active Antioxidant fraction S antioxi The scores obtained are shown in Table 4 below.
[0055] Table 4: Display of Four Types of Fractions
[0056]
[0057] Step S103, set S gene S mt S active S antioxi The corresponding threshold G gene G mt G active G antioxiAnd determine the cell population type.
[0058] When S gene Score less than G gene At that time, the cell population was determined to be cell debris.
[0059] When S mt Greater than G mt And S antioxi Less than G antioxi At that time, the cell population was determined to be dead / dying cells.
[0060] When S active Less than G active At that time, the cell population was determined to be dead / dying cells.
[0061] G gene G mt G active G antioxi The values are usually set to 500, 25%, 2, 2.
[0062] As shown in Table 4 above, cell population D (corresponding to cell C8001) satisfies all three of the above conditions. Therefore, cell population D is identified as both cell debris and dead / dying cells.
[0063] In the first-stage filtered cell expression profile, cells identified as cell debris and dead / dying cells (meeting either of these criteria necessitates deletion) are removed, generating a second-stage filtered cell expression profile. Table 5 shows the second-stage filtered cell expression profile of this embodiment.
[0064] Table 5: Expression profile of secondary filter cells
[0065]
[0066] The cells in the first and second filtered cell expression profiles are all real cells, hence they are also called real cell expression profiles.
[0067] Step S104: Based on the simulated multicellular feature expression profile, multicellular cells in real cells are selected using the KNN algorithm, and then filtered from the secondary filtered cell expression profile to generate the final filtered expression profile. The filtering of multicellular cells can be performed independently or after filtering cell debris or dead / dying cells.
[0068] Please refer to Figure 2 Specifically, this is achieved through the following steps S1041 and S1046 (the grouping step is completed by the aforementioned S101):
[0069] S1041, The expression profile of each cell population is generated by taking the average expression level of each gene (rounded to the nearest whole number).
[0070] Taking cell population A as an example, if it only contains 2 cells in Table 5, its characteristic expression profile is as shown in Table 5 below:
[0071]
[0072] S1042. Randomly combine the characteristic expression profiles of all cell populations in pairs to generate artificial multicells with a certain proportion P N The proportion P N is usually set to 25%; the proportion P N of artificial multicells = number of artificial multicells / (number of artificial multicells + number of real cells) * 100%;
[0073] The method of pairwise combination of the characteristic expression profiles is as follows:
[0074] Y = a1 * X1 + a2 * X2
[0075] where Y is the generated artificial multicell, X1 and X2 are the characteristic expression profiles of cell populations; a1 and a2 are proportionality coefficients, one of a1 and a2 is set to 1, and the other is set to a random value greater than 0 and less than 1. That is, when a1 is 1, a2 is a random value between 0 and 1, and when a2 is 1, a1 is a random value between 0 and 1, 1 < a1 + a2 < 2. Each artificial multicell generated in this way contains a complete cell and a defective cell, that is, simulating the generation process of multicells in actual experiments.
[0076] S1043. Combine the artificial multicell expression profile and the secondary-filtered cell expression profile, and use the Seurat software to re-normalize and perform PCA dimensionality reduction. Based on the PCA dimensionality reduction results, calculate the distance between each cell, usually the Euclidean distance or the Manhattan distance, and other distance measurement methods can also play a similar role.
[0077] S1044. Set 100 equally spaced neighborhoods P kn within the range of 0.0001 to 0.01, kn calculate the proportion P kn of artificial multicells within each neighborhood P ANN for each real cell, that is, count the number N merge of artificial multicells among the nearest (total number of cells N kn * P ANN in the combined expression profile) cells to the real cell, and calculate P ANN = N ANN / (N merge * P kn ).
[0078] S1045. Statistically analyze P kn under each neighborhood P ANNDetermine the distribution and calculate its skewness s and kurtosis k. The formula for calculating skewness s is s = The formula for calculating kurtosis k is k= Where n is the same neighborhood P kn Next P ANN The number of p i For the same neighborhood P kn The i-th P ANN The value of M is the same as that of P in the neighborhood. kn All P ANN The average value of SD is the same as that of P in the neighborhood. kn All P ANN The standard deviation.
[0079] Then calculate each P ANN The bimodal coefficient BC of the distribution, BC= , where N real The actual number of cells is used. The neighborhood size where the bimodal coefficient BC is maximized is selected as the optimal neighborhood P to be used in the final calculation. K .
[0080] In the optimal neighborhood P K Within this neighborhood, the boundary between single-celled and multi-celled cells can be clarified to the greatest extent possible. That is, all cells will be divided into single-celled and multi-celled categories as much as possible, reducing the number of intermediate cells with ambiguous classification, so as to obtain the most accurate classification results possible.
[0081] S1046, set the multi-cell rate R doub Calculate the expected value E of the number of cells. doub =N real *R doub The neighborhood size is calculated as the optimal neighborhood P. K At that time, P of each real cell ANN , will P ANN The largest E doub Individual cells are identified as multicellular and removed from the secondary filtered cell expression profile to generate the final filtered cell expression profile.
[0082] Example 2:
[0083] A computer storage medium storing a computer program that, when executed by a processor, implements the single-cell transcriptome RNA contamination identification method as described in Example 1 and / or the single-cell transcriptome low-quality cell filtering method as described in Example 2.
[0084] Example 3:
[0085] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the single-cell transcriptome low-quality cell filtering method as described in Embodiment 1.
[0086] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.
Claims
1. A method for filtering cell debris and multi-cell transcriptome data, characterized in that, Includes the following steps: S101, based on real cell expression profiles, groups cells; The genes are set into four gene sets A. gene A mt A active A antioxi And calculate the expressed gene fraction S for each cell population. gene Mitochondrial fraction S mt Activity fraction S active Antioxidant fraction S antioxi A gene S is the total set of genes filtered from the cellular expression profile in one pass. gene A is the average number of expressed genes in all cells of the cell population; mt For the mitochondrial gene set of the corresponding species, S mt A for all cells in the cell population mt The average gene expression ratio; A active S represents the housekeeping gene set for the corresponding species. active A for all cells in the cell population active The average expression level of a gene; A antioxi For the corresponding species' antioxidant gene set, S antioxi A for all cells in the cell population antioxi The average gene expression level; set S gene S mt S active S antioxi The corresponding threshold G gene G mt G active G antioxi And determine the cell population type; in the real cell expression profile, when S gene Score less than G gene When a cell population is identified as cell debris, the cells identified as cell debris are deleted. S1041, the cell expression profile of each cell population is calculated by taking the average expression level of each gene to generate the characteristic expression profile of each cell population. S1042, Randomly combine the characteristic expression profiles of the cell population in pairs to generate a certain number of artificial multi-cells; S1043, Combine the artificial multi-cell expression profile and the real cell expression profile, and calculate the distance between each cell; S1044, Set up several equidistant neighborhoods within a specified range, and calculate the proportion of artificial multicellular cells in each neighborhood for each real cell; S1045, statistically analyze the distribution of artificial multicellular proportions in each neighborhood, calculate its bimodal coefficient, and take the neighborhood with the largest bimodal coefficient as the optimal neighborhood. S1046, In the optimal neighborhood, the real cells with the largest artificial multicellular proportion are identified as multicellular and removed from the real cell expression profile; The bimodal coefficient is: in, This is the bimodal coefficient; and These represent the skewness and kurtosis of the artificial multicellular proportion distribution, respectively. This represents the actual number of cells.
2. The method for filtering single-cell transcriptome cell debris and multiple cells according to claim 1, characterized in that, The method for pairwise combination of the feature expression spectra is as follows: Y = a1*X1 + a2*X2 Where Y represents the generated artificial multicellular structure, X1 and X2 represent the characteristic expression profiles of the cell population; one of a1 and a2 is set to 1, and the other is set to a random value greater than 0 and less than 1.
3. The method for filtering single-cell transcriptome cell debris and multiple cells according to claim 1, characterized in that, The distance between the cells is either Euclidean or Manhattan distance.
4. The method for filtering single-cell transcriptome cell debris and multiple cells according to claim 1, characterized in that, The artificial multi-cell ratio is the ratio of the number of artificial multi-cells to the total number of cells whose expression profiles are located in that neighborhood after merging.
5. The method for filtering single-cell transcriptome cell debris and multi-cell cells according to claim 1, characterized in that, The method for determining the specified number is as follows: Set a multi-cell rate, and use the product of the actual number of cells and the multi-cell rate as the specified number of actual cells to be identified as multi-celled.
6. The method for filtering single-cell transcriptome cell debris and multi-cell cells according to claim 1, characterized in that, In S1044, the neighborhood is set as follows: 100 equally spaced neighborhoods are set in the range of 0.0001 to 0.
01.
7. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the single-cell transcriptome cell debris and multi-cell filtering method as described in any one of claims 1 to 6.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the single-cell transcriptome cell debris and multi-cell filtering method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A
Cell development map and marker gene of mandible tissue of human three-month-old embryo
CN114438230A