A method and system for single-cell data quality control processing

By employing a systematic single-cell data quality control method that integrates multi-omics quality control processes, outliers, low-quality cells, and background noise are identified and removed. This solves the fragmentation problem of multi-omics quality control in existing technologies and achieves efficient, comprehensive, and flexible data quality control.

CN119920311BActive Publication Date: 2025-12-26GUANGZHOU NAT LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510120999.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-12-26
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

Existing single-cell quality control tools and methods lack comprehensive quality control for multi-omics, the quality control process is fragmented, and there is a lack of quality control methods for inter-sample comparison, which cannot effectively identify and remove outliers, low-quality cells and background noise in multi-omics data.

Method used

A systematic single-cell data quality control method is adopted, including sample level, cell level, feature level and batch effect assessment module. By calculating quality control indicators, identifying outliers, removing empty droplets and double cells, low-quality cells, correcting background noise and assessing batch effects, it integrates existing tools and processes to generate interactive reports.

Benefits of technology

It enables systematic and comprehensive quality control of multi-omics data, improves data reliability and accuracy, simplifies the quality control process, reduces computing resource requirements, and is suitable for large-scale single-cell data analysis on ordinary computers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920311B_ABST
    Figure CN119920311B_ABST
Patent Text Reader

Abstract

The present application relates to a single cell data quality control processing method, comprising the following steps: step one) calculating the quality control index according to the single cell data to be processed; then identifying outliers in the single cell data to be processed according to the quality control index and / or cell type composition, obtaining sample level quality control analysis results; step two) identifying empty droplets and double cell data in the single cell data to be processed; then identifying low-quality cell data in the single cell data to be processed, obtaining cell level quality control analysis results; step three) calculating and evaluating the average expression level, expression standardized variance and detection rate in all cells of each feature in the single cell data to be processed, obtaining feature level quality control analysis results; step four) evaluating the batch effect of the single cell data to be processed, obtaining batch effect evaluation analysis results. Compared with other existing process tools, the single cell multi-omics quality control process tool of the present application is more complete, flexible and friendly, and provides the most comprehensive and powerful quality control workflow for single cell multi-omics analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biological information, in particular to a single-cell data quality control processing method and system. BACKGROUND

[0002] The development of single-cell omics technology has played a key role in promoting people's understanding of health and disease, and has provided important evidence for discovering disease mechanisms and potential therapeutic targets. Quality control (QC) of single-cell omics is crucial to ensure the reliability and accuracy of data. To ensure data integrity, avoid mislabeling and misidentification, and prevent misleading results in downstream analysis, it is crucial to systematically evaluate low-quality items in samples, cells, and expression characteristics, as well as batch effects, environmental contamination, and background noise from various technical and human errors.

[0003] Although there have been significant advances in current single-cell QC strategies, such as scDblFinder and DoubletFinder, which can reliably identify and exclude doublets; tools such as CellBender and DecontX effectively alleviate system background noise; and SCTK-QC workflow can generate and visualize QC indicators for scRNA-seq data. However, existing quality control tools and methods have the following limitations: First, existing technologies mainly target a single omics (such as single-cell transcriptome), and lack comprehensive quality control for multiple omics. Moreover, the quality control process is fragmented, lacking a systematic and comprehensive quality control evaluation. Furthermore, there is a lack of quality control methods for inter-sample comparison, which is usually limited to single-sample quality control.

[0004] Therefore, in view of the defects of the current quality control method, the present application proposes a complete standardized single-cell data quality control processing method. SUMMARY

[0005] In order to perform systematic and comprehensive quality control of single-cell data omics and ensure the consistency and biological accuracy of multi-dimensional data, the present application proposes a single-cell data quality control processing method. The present application is implemented by using the following technical solutions:

[0006] A single-cell data quality control processing method, comprising the following steps:

[0007] Step 1) Calculate the quality control indicators in the single-cell data to be processed; identify outliers in the single-cell data to be processed based on the quality control indicators and / or cell type composition, and obtain sample-level quality control analysis results;

[0008] Step 2) Identify empty droplets and doublet data in the single-cell data to be processed; then identify low-quality cell data in the single-cell data to be processed, and obtain cell-level quality control analysis results;

[0009] Step three) calculating the average expression level, expression standardized variance and detection rate in all cells of each gene / protein in the single-cell data to be processed, and evaluating and removing the background noise of the single-cell data to obtain the feature level quality control analysis result;

[0010] Step four) evaluating the batch effect of the single-cell data to be processed to obtain the batch effect evaluation analysis result.

[0011] The method of the application can process single-cell single-omics data or single-cell multi-omics data.

[0012] Optionally, the single-cell multi-omics includes scRNA-seq, RNA+ADT-seq and scTCR / BCR-seq.

[0013] Preferably, the quality control indicators in step one) include cell number, sample cell UMI abundance median, sample cell detected gene / protein median, sample cell mitochondrial gene count percentage median, TCR% and BCR%, etc.

[0014] Optionally, the step one) identifies outliers in the single-cell data to be processed based on the quality control indicators, including the following steps:

[0015] The deviation of the value of each quality control indicator in the processed single-cell data from the median is calculated using the median absolute deviation, and outliers in each quality control indicator are identified;

[0016] Preferably, the quality control indicators include sample cell UMI abundance median, sample cell detected gene / protein median and sample cell mitochondrial gene count percentage median.

[0017] Optionally, the step one) identifies outliers in the single-cell data to be processed based on cell type composition, including the following steps:

[0018] 1) comparing and analyzing the cell type proportion of the sample in the single-cell data to be processed with the cell type proportion of the tissue atlas constructed internally, judging whether the cell proportion of the sample deviates from the overall range of its cell type, and identifying the deviated sample as an outlier;

[0019] 2) using PCA combined with confidence interval and DBSCAN method to identify samples with cell type proportion outliers in the overall sample group.

[0020] Optionally, the step one) further includes identity verification processing of the single-cell data to be processed, and identifying samples with identity verification errors;

[0021] Preferably, the identity verification process comprises setting a verification gene based on the sample type of the single-cell data to be processed, and performing identity verification by evaluating the cell detection proportion of the verification gene.

[0022] Optionally, the step two) adopts dropletUtils package to identify empty droplets.

[0023] Preferably, the step two) adopts at least one of the following methods to identify doublet cells: a single-cell transcriptome doublet cell identification method, a V(D)J sequencing doublet cell identification method, and an ADT data doublet cell identification method.

[0024] Preferably, the single-cell transcriptome doublet cell identification method comprises at least one of the following methods: DoubletFinder, scDblFinder, cxds, bcds, and hybrid.

[0025] Preferably, the ADT data doublet cell identification method is a method of identifying doublet cells based on ADT data in RNA+ADT joint sequencing technology by combining a cycle iteration COSG method and a common statistical classification method.

[0026] Optionally, the step of identifying low-quality cell data in the single-cell data to be processed adopts at least one of the following methods: a fixed threshold filtering method, a MAD filtering method, a miQC filtering method, and a ddqc filtering method.

[0027] Optionally, the step two) further comprises performing cluster analysis on the identified low-quality data, and analyzing the cell types of the identified low-quality cell data according to the cluster analysis results.

[0028] Then, it is determined whether to retain the part of the single-cell data according to the cell types.

[0029] The features in the step three) are genes or protein data in the single-cell data, and each gene and protein is a feature.

[0030] Optionally, the step three) further comprises performing background noise correction on the single-cell data to be processed, and the background noise correction comprises global background noise correction or local background noise correction.

[0031] Preferably, the global background noise correction method in the step three) comprises DecontX and DecontPro methods, so as to perform global background noise correction on RNA and ADT data.

[0032] Preferably, the local background noise correction method in the step three) comprises using scCDC to identify and remove highly contaminated RNA or protein.

[0033] Optionally, the evaluating the batch effect of the single-cell data to be processed comprises:

[0034] 1) performing UMAP analysis on the single-cell data to be processed, and performing PCA analysis on the pseudo-transcriptome data;

[0035] 2) performing singular value decomposition analysis and scater package analysis on the influence degree of batch-related variables.

[0036] Optionally, the method for quality control processing of single-cell data further comprises:

[0037] The sample-level quality control analysis result, the cell-level quality control analysis result, the feature-level quality control analysis result and the batch effect evaluation analysis result are combined to form a visual webpage report.

[0038] A method for quality control processing of single-cell data, comprising:

[0039] A sample-level quality control module is configured to calculate quality control indicators in the single-cell data to be processed, and identify outliers in the single-cell data to be processed based on the quality control indicators and / or cell type composition, to obtain a sample-level quality control analysis result.

[0040] A cell-level quality control module is configured to identify empty droplets and double-cell data in the single-cell data to be processed, then identify low-quality cell data in the single-cell data to be processed, and draw a quality control chart to obtain a cell-level quality control analysis result.

[0041] A feature-level quality control module is configured to calculate and evaluate the average expression level, expression standard deviation and detection rate in all cells of each gene / protein in the single-cell data to be processed, and evaluate and remove background noise of the single-cell data, to obtain a feature-level quality control analysis result.

[0042] A batch effect analysis module is configured to evaluate the batch effect of the single-cell data to be processed, to obtain a batch effect evaluation analysis result.

[0043] Compared with the prior art, the present application has the following advantages and effects:

[0044] The present application can simultaneously perform quality control on multiple omics sequencing data of multiple samples, including scRNA-seq, surface protein sequencing and immune repertoire sequencing, so that users can fully evaluate their data before downstream analysis. The present application integrates existing quality control steps, builds a bridge between various quality control processes and tools, solves the compatibility problem between different tools, connects multiple quality control processes into a workflow, and is built on the most popular Seurat framework to facilitate downstream analysis. At the sample level, the present application provides a unique outlier sample detection module, especially a cell proportion outlier detection module, which integrates thousands of single-cell data to help users determine the biological variation of the sample and assist in sample-level quality control. At the cell level, the present application provides low-quality cell filtering criteria and recommended reference ranges for sequencing characteristics for various tissues, so that users can select the most suitable QC reference threshold for their experiments. In addition, in order to avoid excessive filtering of non-low-quality cells, the present application also provides a re-clustering strategy for low-quality cells and marker gene enrichment analysis. In addition, the present application also develops a doublet identification strategy for ADT data in the omics RNA+ADT joint sequencing technology to make up for the deficiency of doublet identification on the RNA level.

[0045] The present application provides a large number of integrated and convenient calculations and QC charts, which greatly simplifies the complex quality control process. In particular, the present application can generate an interactive HTML report containing single-cell data system quality control information, so that users can quickly check the data quality before downstream analysis, and users who do not need to conduct in-depth programming experiments can easily perform comprehensive data quality control.

[0046] The present application is compatible with Seurat V5 and BPCells versions, can adapt to large single-cell data quality control, and reduces memory occupation. It has low requirements for computer memory, and a common notebook computer can perform quality control analysis on millions of cells.

[0047] Compared with other existing process tools, the single-cell multi-omics quality control process tool of the present application is more flexible and friendly, and provides the most comprehensive and powerful quality control workflow for single-cell multi-omics analysis. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0049] Figure 1 is a flowchart of the present application.

[0050] Figure 2 is the data basic situation, wherein A is the data part quality control index; B is the proportion distribution of TCR and BCR chains of V(D)J data; C is the receptor condition distribution of V(D)J data.

[0051] Figure 3 is the outlier sample identification condition, wherein A is the sample basic quality control index outlier condition; B is the violin plot of the sample basic quality control index; C is the PCA distribution plot of the sample cell type proportion.

[0052] Figure 4 is the doublet cell identification result, wherein A is the proportion condition of doublet cells identified by different doublet cell identification methods; B is the similarity and difference of the prediction results of different doublet cell methods.

[0053] Figure 5 is the proportion condition of low-quality cells identified by different low-quality cell identification methods.

[0054] Figure 6 is the result of low-quality cell re-clustering analysis, wherein A is the result of low-quality cell re-clustering ScType automatic annotation; B is the condition of some classic cell markers in different re-clustering cell populations.

[0055] Figure 7 is the feature level analysis result, wherein A is the expression condition of cell genes in cell populations; B is the expression condition of high-pollution genes identified by the scCDC method in cell populations; C is the expression condition of cell proteins in cell populations; D is the expression condition of high-pollution proteins identified by the scCDC method in cell populations.

[0056] Figure 8 is the partial analysis result of batch effect, wherein A is the PCA analysis at the sample level; B is the RNA omics UMAP plot at the cell level; C is the ADT omics UMAP plot at the cell level.

[0057] Figure 9 is the batch covariate analysis result, wherein A is the SVD analysis of the batch covariate and the PC1 of the pseudo-transcriptome data; B is the contribution degree of each batch covariate to the influence of the pseudo-transcriptome data at the PC1 level;

[0058] Figure 10 is an example of a report page generated by the system of the present application.

[0059] Figure 11Figure 1 shows the scatter plots of protein expression threshold division for uncorrected background condition, wherein A is the scatter plot of CD3-CD20 and CD19-CD3 mutually exclusive protein pair expression threshold division of TP1 sample, and the three figures correspond to different method strategies respectively, the red color represents the division threshold of each method, and the red dot ambiguous represents the T-B double cells identified at the VDJ level; B is the scatter plot of mutually exclusive protein pair expression threshold division of TP2 sample; C is the scatter plot of mutually exclusive protein pair expression threshold division of TP3 sample; D is the scatter plot of mutually exclusive protein pair expression threshold division of TP3-rep sample; E is the scatter plot of mutually exclusive protein pair expression threshold division of pbmc5K sample; F is the scatter plot of mutually exclusive protein pair expression threshold division of pbmc10K sample.

[0060] Figure 12 Figure 2 shows the scatter plots of protein expression threshold division for corrected background condition, wherein A is the scatter plot of CD3-CD20 and CD19-CD3 mutually exclusive protein pair expression threshold division of TP1 sample, and the two figures correspond to different method strategies respectively, the red color represents the division threshold of each method, and the red dot ambiguous represents the T-B double cells identified at the VDJ level; B is the scatter plot of mutually exclusive protein pair expression threshold division of TP2 sample; C is the scatter plot of mutually exclusive protein pair expression threshold division of TP3 sample; D is the scatter plot of mutually exclusive protein pair expression threshold division of TP3-rep sample; E is the scatter plot of mutually exclusive protein pair expression threshold division of pbmc5K sample; F is the scatter plot of mutually exclusive protein pair expression threshold division of pbmc10K sample.

[0061] Figure 13 Figure 3 shows the threshold variation of various methods at different clustering resolutions, wherein A is the threshold fluctuation of different methods for TP1 sample; B is the threshold fluctuation of different methods for TP2 sample; C is the threshold fluctuation of different methods for TP3 sample; D is the threshold fluctuation of different methods for TP3-rep sample; E is the threshold fluctuation of different methods for pbmc5K sample; F is the threshold fluctuation of different methods for pbmc10K sample. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical scheme of the present application more clear, the present application is further described in detail. In the following examples, the experimental methods are described, and if no special description, all are conventional methods: If no specific technology or condition is specified in the examples, the technology or condition described in the literature in the art or according to the product manual is used; if no special description, the reagents and materials can be obtained from commercial channels.

[0063] The application provides a single-cell data quality control processing method, which comprises the following steps:

[0064] Step one) sample level quality control, comprising: calculating the quality control indicators in the single-cell data to be processed; and identifying outliers in the single-cell data to be processed based on the quality control indicators and cell types. The specific process comprises:

[0065] 1-1) sample quality assessment:

[0066] This step is used for identifying samples that may be abnormal due to library construction, insufficient sequencing depth, amplification bias, pollution and the like. By integrating sequencing indicators generated by upstream software such as CellRanger, for example, Q30, saturation and the like, the sample is quickly preliminarily evaluated in quality according to the official standard. Then, other common quality control indicators are calculated, and the sample level quality control indicators are shown in Table 1, wherein some important indicators, such as cell number (nCell), median of nCount_RNA / nCount_ADT (median of sample cell UMI abundance), median of nFeature_RNA / nFeature_ADT (median of sample cell detected gene / protein), median of MT% (median of sample cell mitochondrial gene count percentage), TCR% and BCR%, etc., can be combined with quality control visualization charts to perform preliminary quality control on the sample.

[0067] Table 1 shows all optional quality control indicators in the method of the application

[0068] Table 1: Sample level quality control indicators

[0069]

[0070]

[0071]

[0072] 1-2) Outlier sample detection:

[0073] This step identifies samples that are significantly out of the quality control indicators or cell type composition in the sample group. The outlier sample detection includes two aspects of a) and b):

[0074] a) Identify samples deviating from the general distribution of various quality control indicators within the sample group. The assumption of this strategy is that the values of sample quality control indicators are generally similar. The default quality control indicators mainly include median of nCount_RNA / nCount_ADT, median of nFeature_RNA / nFeature_ADT and median of MT% common quality control indicators. Other quality control indicators can also be selected, see Table 1. The method of identifying outliers is to use the median absolute deviation (MAD) to measure the deviation of each value in the data set from the median, thereby identifying outlier samples for each key quality control indicator. Each quality control indicator will alert potential outlier samples, and the more sample quality control indicators that are alerted, the more likely it is that the sample is abnormal. This step facilitates users to quickly view potential abnormal samples, assisting users in sample control to avoid the impact of low-quality samples on biological analysis conclusions.

[0075] b) Identify samples with significantly deviated cell type composition. Deviation in cell type composition among samples from the same experimental group can indicate potential technical problems, such as prolonged processing time, changes in storage conditions, experimental handling errors, or batch processing effects, all of which can compromise data integrity. The strategy for identifying sample cell type composition includes two aspects:

[0076] First, the cell type proportion of the sample is compared with the cell type proportion of the existing tissue atlas. The specific steps include: first, download 2299 single-cell samples from DISCO data, involving 10 tissue types, including blood, bone, bone marrow, breast, intestine, kidney, liver, lung, pancreas and skin. Use ScType automatic annotation method and custom main type cell group to cluster and group annotate each sample, and further calculate the fluctuation of main type in different tissue samples, establish the reference fluctuation range of single-cell cell type of each tissue, including B cell, dendritic cell, endothelial cell, fibroblast, granulocyte, monocyte / macrophage, natural killer cell, other cell and T cell, and the reference fluctuation range is the maximum and minimum value of the cell type in all samples of the tissue (Table 2). The reference fluctuation range is the maximum and minimum value of the cell type in all samples of the tissue. For each sample to be detected, cluster and group annotation will be automatically run and further calculate the cell type proportion, compared with the established tissue fluctuation range, finally judge whether the cell proportion of the sample deviates from the overall range, if deviates, the sample is identified as an outlier, and the outlier sample is more likely to have abnormal situation, giving warning to the user.

[0077] Table 2: Reference fluctuation range of single-cell cell type of different tissues

[0078]

[0079]

[0080]

[0081] Second, the PCA combined with confidence interval and DBSCAN method is used to identify the sample group overall internal cell type proportion outliers. The hypothesis of this strategy is that the same group of samples, the cell proportion composition is more similar. Specifically, according to the existing annotation results according to ScType, the cell type proportion of each sample is calculated, and PCA analysis is performed on all samples, and the confidence ellipse and DBSCAN are used to identify potential outliers. These outliers represent the deviation of cell proportion, which to some extent represents the heterogeneity of such samples, and the meaning in different groups is different, and the user can analyze according to the situation of each data set. For example, different batches of samples of the same person, the theoretical situation is that the proportion should remain consistent, and if significant outliers are found, it means that the sample is abnormal.

[0082] 1-3) Sample identity verification module:

[0083] Purpose: Identify incorrect sample labels that may be caused by sample processing errors.

[0084] Solution: By evaluating the cell proportion of the expression verification gene, the expression pattern of the specific gene is evaluated, and the consistency of the data expression performance and the recorded label is compared through visual chart means. By default, the user can determine the sample gender according to the gender-specific gene expression, including XIST, DDX3Y, UTY and RPS4Y1, so as to better check the consistency with the recorded gender.

[0085] Step two) cell level quality control, including identifying empty droplets in single-cell data to be processed and double cells in double-cell data; then identifying low-quality cell data in single-cell data to be processed. The specific method includes the following steps:

[0086] 2-1) Identification and removal of empty droplets and double cells. It avoids the result errors caused by environmental RNA contamination, microfluidic cell capture errors and improper cell encapsulation. Among them, the empty droplets are identified using the dropletUtils package. The identification of double cells integrates five existing public single-cell transcriptome methods (DoubletFinder, scDblFinder, cxds, bcds and hybrid), a public method using V(D)J data, and a new double cell identification method for ADT data in RNA+ADT joint sequencing technology. Generally, the double cell detection strategy in V(D)J sequencing will mark the cells with multiple TCR or BCR chains (or both) as potential double cells. This is particularly effective in identifying homodouble cells by detecting multi-chain cells, which can make up for the shortcomings of the transcriptome method in identifying homodouble cells. For single-cell data, single-cell transcriptome strategies and other omics methods provided by the present process can be combined to identify double cells together.

[0087] Single-cell transcriptome analysis is prone to lose information due to the sparsity of data, resulting in a blind spot for quality control. The ADT data in the RNA+ADT joint sequencing technology is bimodal, and the missing rate is not so serious, so for the multi-omics RNA+ADT joint sequencing technology data, the identification of double cells can be further recognized and supplemented from the protein level. The double cell identification method developed internally based on the ADT data in the RNA+ADT joint sequencing technology mainly includes two parts, the first step of identification uses the cyclic iteration COSG strategy to identify negative and positive cell groups, and the second step uses statistical methods to identify the threshold level of high and low expression in the protein level. Specifically, the double cell identification method comprises the following steps:

[0088] a) Obtain a data set to be subjected to double cell identification, and then cluster and group the data set to obtain a plurality of cell groups. The data set is a single cell data set generated by the RNA+ADT joint sequencing technology. Clustering and grouping uses RNA clustering or ADT clustering. In the present application, RNA and ADT single-omics clustering is supported, and in the case of insufficient ADT omics proteins, RNA clustering can be used instead, and ADT clustering will be used by default.

[0089] b) In the cell groups obtained in step a), the expression negative cell group and the expression positive cell group corresponding to each protein in the mutually exclusive protein pair are determined by using the disclosed COSG method of cyclic iteration; the mutually exclusive protein pair is defined according to the biological cell type, for example, the mutually exclusive protein pair of B and T cells is CD19 and CD3; specifically comprising:

[0090] For each protein in the mutually exclusive protein pair, the percentage of proteins detected in each cell group and the average expression level are calculated; the product of the percentage of proteins and the average expression level is the penalty mean of the cell group;

[0091] The cell group with the highest penalty mean is used as the initial seed, and the cell groups are iteratively added from high to low according to the penalty mean of the cell groups; in each iteration, the COSG score of the combination of the added cell groups is calculated using the disclosed COSG method. This disclosed COSG method calculates the cosine similarity between the expression pattern of a feature in a certain cell group and its assumed marker pattern, thereby obtaining the specific COSG score of the feature in a certain cell group, and the specific formula is as follows:

[0092]

[0093] wherein g i represents each protein, G k represents the combination of cell groups, λ k represents G kExpression pattern of hypothetical marker genes of cell population combination, λ t Expression pattern of hypothetical marker genes of other cell population except G k u represents penalty factor; when iteration does not increase COSG score or the frequency of corresponding marker gene in current cell population is less than 0.2, stop iteration; the cell population in the added cell population combination is expression positive cell population, and the rest is expression negative cell population.

[0094] c) According to the data in the expression negative cell population and the expression positive cell population, calculate the positive expression and negative expression threshold of each protein in the mutually exclusive protein pair;

[0095] d) According to the threshold, identify double cell data in the data set to be double cell identified. Generally, the data of cells whose expression level of each protein in the mutually exclusive protein pair is higher than the respective threshold is identified as double cell data.

[0096] d) According to the data in the expression negative cell population and the expression positive cell population, calculate the positive expression and negative expression threshold of each protein in the mutually exclusive protein pair, including:

[0097] At least one of the ROC and the naive Bayes method is used to calculate the positive expression and negative expression threshold of each protein in the mutually exclusive protein pair; the ROC method uses the labels of the positive and negative populations of the protein generated in step c) and the expression value of the protein to generate the ROC curve and its object, and uses the Youden index to seek the optimal threshold of the ROC curve, that is, the positive expression and negative expression threshold of the protein. The naive Bayes method takes the labels of the positive and negative populations of the protein generated in step c) as the dependent variable of the model, and takes the expression value of the protein as the independent variable of the model, to construct a naive Bayes classification model. Then use the model to predict the expression value of the protein, get the new predicted label of the positive and negative cells, then calculate the minimum value of the protein expression of the new positive population and the maximum value of the protein expression of the negative population, and take the average of the minimum value and the maximum value as the threshold of the positive and negative expression of the protein.

[0098] COSG algorithm (Cosine Similarity-based Gene Identification) is a marker gene identification method based on cosine similarity. It is mainly used in single cell data analysis, which is used to identify cell marker genes more accurately and quickly.

[0099] The COSG algorithm utilizes cosine similarity to measure the relationship between gene expression vectors. Cosine similarity assesses the similarity between two vectors in a vector space by calculating the cosine of the angle between them. Unlike traditional statistical testing methods, cosine similarity compares the direction of two vectors, rather than their absolute values, which means that even if two genes have the same expression pattern but different expression abundances, cosine similarity will consider them equivalent.

[0100] The COSG algorithm is particularly suitable for cell type annotation in single-cell data analysis. In single-cell sequencing technology, cell type annotation is a key step, and traditional statistical methods may encounter challenges in identifying cell marker genes, as statistical tests tend to identify candidate genes with systematic differences between two groups rather than true cell markers. The COSG algorithm can more accurately identify cell marker genes by calculating the cosine similarity of gene expression vectors, thereby improving the accuracy of cell type annotation.

[0101] Compared with the existing MLtiplet CITE-seq method, the method is more stable in identifying positive and negative cell population modules. MLtiplet CITE-seq uses mean to distinguish positive and negative cell populations, which is easily affected by background noise, clustering, and expression levels. In addition, each mutually exclusive protein of MLtiplet CITE-seq needs to contain a matching gene to identify double cell populations. The method does not require matching proteins and genes, but only ADT data for double cell recognition analysis. In addition, it provides background correction strategies and various statistical discrimination methods, which can better adapt to different user groups.

[0102] 2-2) Identification and removal of low-quality cells: The identification and removal of low-quality cells can reduce interference with downstream data analysis. The method of the present invention integrates four low-quality cell removal strategies, including fixed threshold filtering of common quality control indicators and three publicly available data-driven low-quality strategies (MAD, miQC, and ddqc). For the fixed threshold filtering strategy, considering the variability of these indicators in different species and tissues, a fixed filtering threshold reference range is provided for each quality control parameter. These threshold ranges come from 240 single-cell studies covering 11 human tissues collected by this process, including cell UMI abundance, cell detected gene count, and cell mitochondrial gene count percentage. Further statistics on the most common, minimum, and maximum thresholds used in these studies are provided to users for reference.

[0103] It is noted that in the method of the present application, in order to minimize the false positive of low quality cells, the present procedure provides a low quality cell re-clustering procedure, combined with functional analysis and differential analysis, to assist in confirming the cell type, so as to retain functional or rare cell types for downstream analysis. We use a fixed threshold nFeature_RNA = 200 and percent.mt = 15 to determine low quality cells, which are used as input data for re-clustering. It can be seen that ScType and marker genes identify some characteristics of classic cells in the low quality cell population, such as platelets. If the user's research is aimed at platelet research, these cells should be re-included in the analysis.

[0104] Step three) feature level quality control: features are genes or protein data in single cell data, and each gene and protein is a feature. In feature level quality control, the average expression level, expression standardized variance and detection rate in all cells of each feature (gene / protein) in the single cell data to be processed are calculated; the system background noise is corrected. Specifically, it includes:

[0105] 3-1) Feature (gene / protein) quality assessment: provides a comprehensive overview of the features, filters potential low quality gene / protein data. First, the average expression level, expression standardized variance and feature detection rate of all cells of each feature are calculated. Visualization icons are used to visualize the distribution of these quality control indicators to provide a comprehensive overview of the features. In order to help filter potential low quality features, the present procedure provides reference thresholds for quality control indicators (feature detection number of all cells for each feature), and these threshold ranges come from 240 single cell studies covering 11 human tissues collected by the present procedure. Further statistics of the mode, minimum and maximum of these thresholds are provided to the user for reference. In addition, the present procedure also supports the use of the scater package to calculate the degree of influence of each feature on the batch factor, to assist users in excluding features that are greatly affected by the batch factor. In addition, for data of RNA+ADT joint sequencing technology, the present procedure additionally provides an evaluation of the correlation between RNA and corresponding surface protein features (e.g., CD3 RNA and CD3 protein).

[0106] 3-2) System background noise correction: this step can handle pollution from environmental RNA or non-specific antibody binding, etc. to avoid affecting downstream analysis. This step integrates DecontX and DecontPro methods to achieve global background noise correction for RNA and ADT data. In order to correct more specifically, this step also integrates the scCDC package to identify and remove highly contaminated RNA or protein, to minimize the risk of over-correction of unaffected features.

[0107] Step four) batch effect evaluation, batch effect evaluation includes the following steps:

[0108] 4-1) Assessing the impact of batch effect using high-dimensional dimensionality reduction techniques: UMAP analysis is performed on the original single-cell data at the cell level to visualize the extent of the impact of batch effects, which is the most commonly used batch assessment method at present. In addition, the present process further calculates the pseudo-transcriptome data for PCA analysis, and aggregates the count data to quickly identify batch effects between samples. The present process integrates the two functions, making it easier for users to use, reducing the complexity of existing processes, and supporting a wider user group.

[0109] 4-2) Batch covariate impact assessment: At present, batch assessment often overlooks the assessment of numerous batch covariates, and existing tools cannot analyze them well. Therefore, the present process integrates singular value decomposition (SVD) analysis and the scater package to comprehensively analyze the impact of batch-related variables. Specifically, it applies SVD analysis to the pseudo-transcriptome data of single-cell data of all cells or specific cell populations at the sample level, aggregates the count data to quickly identify batch-related variables significantly related to principal components, and further views the impact of the variables to assist in determining the importance of batch-related variables. In addition, the existing public scater package is used to calculate the percentage of variance explained by each batch variable on the RNA and surface protein features of the original single-cell data at the cell level, and the contributions of these batch variables are visualized by density.

[0110] Step five) Quality control web report

[0111] In order to facilitate users to execute the present process and better understand their own data set, the present process provides an integrated quality control interactive web report for the above four processes using Rmarkdown technology, which contains detailed parameter explanations and result explanations, provides a variety of interactive charts, and facilitates users to quickly explore data.

[0112] The present application also proposes a single-cell data quality control processing system based on the above method, comprising:

[0113] A sample-level quality control module is configured to calculate quality control indicators in the single-cell data to be processed, and identify outliers in the single-cell data to be processed based on the quality control indicators and / or cell type composition, to obtain sample-level quality control analysis results.

[0114] A cell-level quality control module is configured to identify and remove empty droplets and double-cell data in the single-cell data to be processed, and then identify and remove low-quality cell data in the single-cell data to be processed, to obtain cell-level quality control analysis results.

[0115] a feature level quality control module, configured to calculate and evaluate the average expression level, expression standardized variance and detection rate in all cells of each gene / protein in the single-cell data to be processed, and evaluate and remove background noise of the single-cell data, to obtain a feature level quality control analysis result;

[0116] a batch effect analysis module, configured to evaluate a batch effect of the single-cell data to be processed, to obtain a batch effect evaluation analysis result.

[0117] Test Example

[0118] In this test example, the quality control is performed on the single-cell data of PBMC of a healthy person. The samples of the single-cell data of PBMC of the healthy person are obtained from the same healthy donor, and the interval between each sampling is three days. Three peripheral blood mononuclear cell (PBMC) samples, TP1, TP2 and TP3, are obtained. The third node sample (TP3) fails to pass the QC initially, and then the data TP3-rep is obtained by re-measuring. In addition, two sets of public data of PBMC published on the website of 10x Genomics Chromium, i.e., pbmc5K and pbmc10K, are used in the analysis. All samples are generated by using the 10x Genomics Chromium 5' gene expression platform to generate single-cell data, including RNA-seq, surface protein profiling by antibody-derived tags (ADTs) and TCR / BCR sequencing. The multi-omics fastq files generated by sequencing are processed by using the 10x Genomics Cell Ranger software (v7.1.0) with the human reference genome sequence GRCh38. Among them, four samples (TP1, TP2, TP3 and TP3-rep), a total of 28,498 cells, are included in the subsequent quality control process analysis of the method of Example 1 of the present application.

[0119] Application Results

[0120] Part of the results of the sample level quality control:

[0121] Firstly, the basic situation of the multi-omics data set can be roughly explored by using the present process, such as the number of cells, the number of detected genes, the proportion of TCR chain cells detected, etc. Figure 2 In case 1, it can be found that the number of cells of the TP3 sample is less and the BCR% is greater than the TCR%, which can preliminarily determine that the TP3 sample is abnormal.

[0122] For the sample abnormality module, in this case, the outlying trend of the quality control indicators of TP3 can be found, and the outlying trend is also shown in the cell type proportion identification. Figure 3, Table 3), further can conclude that TP3 sample is abnormal. In the proportion comparison process, the proportion of B cells and T cells in TP3 sample deviates from the range of the reference database, and such sample is preferably excluded to avoid the impact on biological authenticity.

[0123] Table 3: Cell proportion of different samples

[0124]

[0125] The identity recognition module shows that the four samples are male, which is consistent with the registered gender, and there is no gender record error.

[0126] Part of the results of cell-level quality control:

[0127] All samples can successfully perform the operation of double cell and low quality cell quality control methods. In this test, the double cell prediction result shows that the proportion of double cell detection is different between different methods and different samples Figures 4-5 ).

[0128] In the determined low-quality cell re-clustering, it can be seen that some cells express platelet characteristics such as PPBP, PF4, etc. Such cells are filtered out, and can be further retained Figure 6 ).

[0129] Part of the results of feature quality control:

[0130] On the RNA level, it is observed that the expression of cell type genes is clear between groups, the high pollution genes predicted by scCDC algorithm are not typical genes, and there is no much improvement before and after correction, which shows that the background pollution on the RNA level is not serious, and does not need to be corrected Figure 7 A-B). In addition, on the ADT level, it can be seen that typical proteins are expressed in multiple cell groups, and the identified high pollution exists typical proteins, and there is a greater improvement after background correction, which shows that the ADT level can be further corrected for background pollution Figure 7 C-D).

[0131] Part of the results of batch effect evaluation:

[0132] Compared with other samples, TP3-rep shows obvious batch effect in ADT sequencing, which shows obvious deviation in PCA and UMAP graphs Figure 8 ), and batch correction is needed.

[0133] Meanwhile, batch covariate analysis can run smoothly, and due to the small sample size, the p value is not very accurate, but it can still be seen that the p value of the antibody reagent covariate is the smallest in most pseudo-transcriptome data, and the degree of influence on PC is high, indicating that the antibody batch covariant has a greater impact on the protein data of most cell types compared with other covariants. Figure 9

[0134] The quality control report result display is as follows:

[0135] The web quality control report can run and display on a notebook computer smoothly, and the quality control report includes four parts, and the samples with problems in each chapter can be seen, and the quality control report gives a yellow warning prompt and detailed explanation. Figure 10

[0136] The present application can simultaneously perform quality control on multiple omics sequencing data of multiple samples, including scRNA-seq, surface protein sequencing and immune repertoire sequencing, so that users can fully evaluate their data before downstream analysis. The present application integrates existing quality control steps, and builds a bridge between various quality control processes and tools, solves the compatibility problem between different tools, connects multiple quality control processes into a workflow. The present application is established on the most popular Seurat framework, which is convenient for users to perform downstream analysis. In addition, the present application is compatible with Seurat V5 and BPCells version, and can adapt to large single cell data quality control and reduce memory occupation. It has a very low requirement for computer memory, and a common notebook computer can perform quality control analysis on millions of cells.

[0137] At the sample level, the present application provides a unique outlier sample detection module, especially a cell proportion outlier detection module, which integrates thousands of single cell data samples to help users judge the biological variation of the sample and assist the quality control at the sample level. At the cell level, the present application integrates 240 single cell publications and provides recommended reference ranges of low-quality cell filtering criteria and sequencing characteristics for 11 kinds of tissues, so that users can select the most suitable QC reference threshold for their experiments. In order to avoid excessive filtering of non-low-quality cells, the present application also provides a re-clustering strategy for low-quality cells and marker gene enrichment analysis. In addition, the present application also develops a doublet identification strategy for ADT data in the multi-omics RNA+ADT joint sequencing technology, which makes up for the deficiency of RNA level identification of doublets.

[0138] In addition, the present application provides a large number of integrated and convenient calculations and QC graphs, which greatly simplify the complex quality control process. In particular, the present application can generate an interactive HTML report containing single cell data system quality control information, so that users can quickly check the data quality before downstream analysis, and users without deep programming experiments can easily perform comprehensive data quality control.​​

[0139] Compared with other existing process tools, the single cell multi-omics quality control process tool of the present application is more flexible and friendly, and provides the most comprehensive and powerful quality control workflow for single cell multi-omics analysis.

[0140] In order to better compare the pros and cons of the methods, the effect of the above ADT data in the RNA+ADT combined sequencing technology based on the COSG method for double cell recognition is also tested. A total of six sets of CITE-seq data sets are tested, and since the six sets of CITE-seq data sets all contain matching VDJ sequencing data, in order to better reflect the performance difference, we selected the common B and T cell mutual protein pairs CD19-CD3 and CD20-CD3, and it can be seen that the method of the present application can effectively divide the double peak characteristics of the proteins in different data sets and proteins, while MLtiplet CITE-seq can only perform well in part of the samples and proteins. Figure 11 Secondly, the method also maintains good division performance on the corrected data of the scCDC method. Figure 12

[0141] Furthermore, the present application further analyzes the COSG+statistical method strategy of the circulation iteration in the clustering dimension of different resolutions, and the threshold fluctuation of different proteins is more stable compared with MLtiplet CITE-seq. Figure 13 This avoids the influence brought by subjective selection of resolution.

[0142] The embodiment of the present application also provides a computer device having the above-mentioned double cell recognition device. The computer device includes one or more processors, memories, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are communicatively connected to each other by different buses, and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or memory to display GUI on an external input / output device (such as a display device coupled to the interface).

[0143] ​In some alternative embodiments, multiple processors and / or multiple buses can be employed with the multiple memories and the multiple memories if desired. Also, multiple computer devices can be connected, with each device providing portions of the necessary operations (e.g., as a server array, a group of blade servers, or a multi-processor system). The processor can be a central processing unit, a network processing unit, or a combination thereof. The processor can further include a hardware chip. The hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device can be a complex programmable logic device, a field programmable logic device, a general array logic, or any combination thereof. The memory stores instructions executable by the at least one processor to cause the at least one processor to perform the double cell recognition method as shown in the above embodiments.

[0144] The memory can include a program storage area and a data storage area. The program storage area can store an operating system, application programs required by at least one function, and the like. The data storage area can store data created according to the use of the computer device, and the like. In addition, the memory can include a high-speed random access memory, and can further include a non-transitory memory such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid state memory device. In some alternative embodiments, the memory can include a memory disposed remotely from the processor, and the remote memory can be connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0145] The memory can include a volatile memory such as a random access memory, and can include a non-volatile memory such as a flash memory, a hard disk, or a solid state disk. The memory can include a combination of the above-mentioned types of memory.

[0146] The computer device further includes an input device and an output device. The processor, the memory, the input device, and the output device can be connected through a bus or other means, Figure 5 For example, the connection through the bus is taken as an example.

[0147] The input device can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, and the like. The output device can include a display device, an auxiliary lighting device (e.g., an LED), a tactile feedback device (e.g., a vibration motor), and the like. The display device includes, but is not limited to, a liquid crystal display, a light emitting diode, a display, and a plasma display. In some alternative embodiments, the display device can be a touch screen.

[0148] The computer device also includes a communication interface for the computer device to communicate with other devices or communication networks.

[0149] The embodiments of the present application further provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium and downloaded to a local storage medium through network, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware.

[0150] The storage medium can be a disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state disk, etc. Further, the storage medium can also include a combination of the above-mentioned memories. It can be understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component that can store or receive software or computer code, when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0151] The embodiments of the present application can also provide a computer program product including computer program instructions, which, when executed by a processor, cause the processor to perform the steps of the above method. The computer program product can be written in any combination of one or more programming languages, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0152] The above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, but not limit them; although the above embodiments of the present application are described in detail, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some technical features can be replaced by equivalent ones; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for quality control processing of single-cell data, characterized in that, The method comprises the following steps: Step one) calculating quality control indicators according to the single-cell data to be processed; then identifying outliers in the single-cell data to be processed according to the quality control indicators and / or cell type composition, to obtain sample-level quality control analysis results; Step two) identifying empty droplets and double-cell data in the single-cell data to be processed; then identifying low-quality cell data in the single-cell data to be processed, and drawing a quality control chart to obtain cell-level quality control analysis results; Step three) calculating and evaluating the average expression level, expression standardized variance and detection rate in all cells of each gene / protein in the single-cell data to be processed, to obtain feature-level quality control analysis results; Step four) evaluating the batch effect of the single-cell data to be processed to obtain batch effect evaluation analysis results; The step one) comprises the following steps: Using the median absolute deviation to calculate the deviation of the value of each quality control indicator in the processed single-cell data from the median, and identifying outlier samples based on each quality control indicator; and The step one) comprises the following steps: 1) comparing and analyzing the cell type proportion of each sample in the single-cell data to be processed with the cell type proportion of the tissue to which the single-cell data to be processed belongs, to determine whether the cell type proportion of the sample deviates from the range of the cell types in the tissue, and identifying the samples deviating from the range as outliers; 2) using PCA combined with confidence interval and DBSCAN method to identify samples with deviated cell type proportion within the sample group as a whole; In the step two), ADT data based on RNA+ADT joint sequencing technology is used to identify double cells by combining the COSG method with cycle iteration.

2. The method of single-cell data quality control processing of claim 1, wherein, The single-cell data is single-cell multi-omics data.

3. The method of single-cell data quality control processing of claim 2, wherein, The single-cell multi-omics data comprises scRNA-seq data, RNA+ADT sequencing data and scTCR / BCR-seq data.

4. The method of single-cell data quality control processing of claim 1, wherein, The quality control indicators in the step one) comprise cell number, sample cell UMI median, sample cell detection gene / protein median, sample cell mitochondrial gene count percentage median, TCR% and BCR%.

5. The method of single-cell data quality control processing of claim 1, wherein, The tissues comprise blood, bone, bone marrow, thymus, intestinal tract, kidney, liver, lung, pancreas and skin. The cell types comprise B cells, dendritic cells, endothelial cells, fibroblasts, granulocytes, monocytes / macrophages, natural killer cells and T cells.

6. The method of single-cell data quality control processing of claim 1, wherein, The step one) further comprises identity verification processing of the single-cell data to be processed to identify samples with identity verification errors. The identity verification processing comprises setting verification genes based on the sample types of the single-cell data to be processed, and then performing identity verification according to the detection proportion of cells expressing the verification genes in the single-cell data to be processed.

7. The method of single-cell data quality control processing of claim 1, wherein, In the step two), the dropletUtils package is used for empty droplet identification.

8. The method of single-cell data quality control processing of claim 1, wherein, The then identifying low-quality cell data in the single-cell data to be processed adopts at least one of a fixed threshold filtering, a MAD filtering, a miQC filtering, and a ddqc filtering.

9. The method of single-cell data quality control processing of claim 1, wherein, In the step two), the identified low-quality data is further subjected to clustering analysis, and the cell type of the identified low-quality data is analyzed according to the clustering analysis result; and whether to retain the low-quality data in the single-cell data to be processed is determined according to the cell type.

10. The method of single-cell data quality control processing of claim 1, wherein, In the step three), the single-cell data to be processed is further subjected to background noise correction, and the background noise correction includes global background noise correction or local background noise correction.

11. The method of single-cell data quality control processing of claim 10, wherein, The global background noise correction method includes using the methods of DecontX and DecontPro to perform global background noise correction on RNA and ADT data.

12. The method of single-cell data quality control processing of claim 11, wherein, The local background noise correction method includes using the scCDC method to identify and remove highly contaminated RNA or protein.

13. The method of single-cell data quality control processing of claim 1, wherein, The batch effect evaluation of the single-cell data to be processed includes: 1) performing UMAP analysis on the single-cell data to be processed, and performing PCA analysis on pseudo-transcriptome data; 2) performing singular value decomposition analysis and scater package analysis on the influence degree of batch-related variables.

14. The method of single-cell data quality control processing of claim 1, wherein, The method of single-cell data quality control processing further includes: The sample-level quality control analysis result, the cell-level quality control analysis result, the feature-level quality control analysis result, and the batch effect evaluation analysis result are combined to form a visual web page report.

15. A system for single-cell data quality control processing, the system comprising: The system performs the method of single-cell data quality control processing according to any one of claims 1 to 14; including: A sample-level quality control module is configured to calculate a quality control index in the single-cell data to be processed, identify outliers in the single-cell data to be processed based on the quality control index and / or cell type composition, and obtain a sample-level quality control analysis result. A cell-level quality control module is configured to identify empty droplets and double-cell data in the single-cell data to be processed, and then identify low-quality cell data in the single-cell data to be processed, and obtain a cell-level quality control analysis result. A feature-level quality control module is configured to calculate and evaluate the average expression level, expression standard deviation, and detection rate in all cells of each gene / protein in the single-cell data to be processed, and evaluate and remove background noise of the single-cell data, and obtain a feature-level quality control analysis result. A batch effect analysis module is configured to evaluate the batch effect of the single-cell data to be processed, and obtain a batch effect evaluation analysis result.

16. An electronic device, comprising: The memory and the processor are connected; The memory stores computer instructions, and the processor executes the computer instructions to perform the method of single-cell data quality control processing according to any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a computer to perform the method of single-cell data quality control processing according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Cell data comprehensive analysis system

    CN116959583A

  • Application of reagent for detecting expression quantity of biomarker in preparation of product for evaluating prognosis of kidney cancer and kit for evaluating prognosis of kidney cancer

    CN119020490A