Systems and methods for single-cell omics analysis with data-driven batch inference

The method for batch inference in single-cell omics data addresses the challenges of unknown batch effects by clustering cells, computing local inverse Simpson index scores, and applying a Dixon’s Q-test, resulting in improved dataset integration and cell type annotation.

WO2025096862A1PCT designated stage expired Publication Date: 2025-05-08MT SINAI SCHOOL OF MEDICINE +2
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
PCT/US2024/054011
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current methods for single-cell omics analysis face challenges in batch inference and quality control, particularly due to unknown batch effects, which compromise dataset integration, cell type annotation, and the robustness of study conclusions.

Method used

The system and method for batch inference in single-cell or single nucleus omics data involve obtaining measurement values for each cell, clustering cells, computing sample-level local inverse Simpson index scores, and applying a Dixon’s Q-test to group samples and assign batches, thereby automating quality control and batch correction.

Benefits of technology

This approach enables robust batch inference and quality control, improving the accuracy of dataset integration and cell type annotation, and enhancing the reproducibility and robustness of single-cell omics analyses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024054011_08052025_PF_FP_ABST
    Figure US2024054011_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for batch correcting single-cell abundance data obtain a measurement value of each feature in a plurality of features originating from each cell in a plurality of cells. The cells collectively originate from a plurality of samples. Feature measurements for each cell are used to cluster the cells into clusters. Cluster assignments of each cell and an identity of a sample in the plurality of samples from which each cell originated are used to compute sample-level local inverse Simpson index (LISI) scores penalized by a measure of importance of cells in each sample. Scores are scaled and ordered before application of a Dixon's Q-test groups them into one or more groups. For each respective group, a common batch assignment is applied to each of the samples in the respective group.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR SINGLE-CELL OMICS ANALYSIS WITH DATA-DRIVEN BATCH INFERENCECROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to United States Provisional Patent Application No. 63 / 594,872, entitled “SYSTEMS AND METHODS FOR SINGLE-CELL OMICS ANALYSIS WITH DATA-DRIVEN BATCH INFERENCE,” filed October 31, 2023 which is hereby incorporated by reference.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under N6600119C4022, awarded by the Defense Advanced Research Projects Agency (DARPA), by R01DK46943 awarded by the National Institute of Health (NIH), and by R01GM071966 awarded by the NIH. The government has certain rights in the invention.TECHNICAL FIELD

[0003] This specification describes systems and methods for batch correcting single-cell abundance data.BACKGROUND

[0004] Single-cell sequencing technologies provide access to the molecular environment of individual cells in different tissues, species, and conditions. Recent advances in multimodal assays have expanded the ability to study the behavior of complex cellular systems at single cell resolution. For examplejoint profiling of scRNA-seq and scATAC-seq of CO VID-19 patients revealed disease severity associated dysregulation of immune responses in certain cell types, such as neutrophils and NK cells. See, Wilk et al., 2021, “Multi-omic profiling reveals widespread dysregulation of innate immunity and hematopoiesis in COVID- 19”, The Journal of experimental medicine 218(8) e20210582. Such insight cannot be achieved with traditional bulk assay technology. While interest in single-cell technologies by clinical and translational researchers has been rapidly growing, analysis of these datasets is not routine and requires a high level of bioinformatics expertise for proper interpretation.

[0005] Recovering a robust single-cell landscape and identifying differential patterns between cell types or conditions entails interrogating datasets from multiple samples. These analyses currently require the use of multiple software tools as well as Python or R scripting. The need for familiarity with rapidly evolving single-cell analysis methods creates a bottleneck for many biology and biomedical research groups in obtaining robust interpretations of publicly accessible or laboratory generated single-cell datasets. Importantly, the manual parameter optimization needed for the use of current tools may lead to different conclusions from the same data.

[0006] Furthermore, a key methodological challenge is that variations in the time of sample processing, sequencing depth, or unknown technical factors result in batch effects that compromise integration, analysis, and interpretation of single-cell studies. Thus, single-cell analyses need to account for batch labels at multiple steps in the processing pipeline. Identification of batches is a particularly difficult problem as the origin of most batch effects is unknown, yet existing single-cell integration methods (Haghverdi et al, 2018, “Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors,” Nature biotechnology 36(5), pp. 421-427; Stuart et al., 2019, “Comprehensive Integration of Single-Cell Data,” Cell 177(7), 1888-1902; Hie et al, 2019 “Efficient integration of heterogeneous single-cell transcriptomes using Scanorama,” Nature biotechnology 37(6), pp, 685-691; Korsunsky etal., 2019, “Fast, sensitive and accurate integration of single-cell data with Harmony,” Nature methods 16(12), pp.1289-1296) and benchmarking studies (Luecken et al., 2022, “Benchmarking atlas-level data integration in single-cell genomics,” Nature methods 19(1), pp. 41-50; and Tran et al., 2020, “A benchmark of batch-effect correction methods for single-cell RNA sequencing data,” Genome biology 21(1) 12-16) all assume that batch information is provided, accurate, and complete. As a result, dataset integration, cell type annotation, downstream analyses, and the robustness of study conclusions may be compromised due to the presence of unknown batch effects affecting the data.

[0007] Given the above-background, what is needed in the art are improved methods for batch inference and for automating quality control assessment of single-cell omics data.SUMMARY

[0008] Advantageously, the present disclosure provides systems and methods for batch inference of single-cell or single nucleus omics data and for automating quality control of these datasets. Such omics data includes, for example, single-cell and single nucleus RNA abundance data, chromatin accessibility data, and histone modifications. A measurement value of each such feature (e.g., gene abundance, amount of chromatin accessibility at a particular locus, count of a particular histone modification, count of a particular epigenetic feature such as methylation) in a plurality of features originating from each cell in a plurality of cells is obtained. The cells collectively originate from a plurality of samples. Feature measurement values for each cell are used to cluster the cells into clusters. Cluster assignments of each cell and an identity of a sample in the plurality of samples from which each cell originated are used to compute sample-level local inverse Simpson index (LISI) scores penalized by a measure of importance of cells in each sample. Scores are scaled and ordered before application of a Dixon’s Q-test groups them into one or more groups. For each respective group, a common batch assignment is applied to each of the samples in the respective group.

[0009] The following presents a summary of the invention in order to provide a basic understanding of some of the aspects of the invention. This summary is not an extensive overview of the invention. It is not intended to identify key / critical elements of the invention or to delineate the scope of the invention. Its sole purpose is to present some of the concepts of the invention in a simplified form as a prelude to the more detailed description that is presented later.

[0010] One aspect of the present disclosure provides a method for inferring batches in single-cell abundance data (e.g., at a computer system comprising one or more processors and a non-transitory computer-readable medium including computer-executable instructions). Once the batches are inferred, conventional batch correction methods can be used for correcting the batches.

[0011] There is obtained, for each respective cell in a plurality of cells, a corresponding measurement value of each respective feature in a plurality of features originating from the respective cell. In some embodiments, for example, the plurality of features comprises 30 or more features (e.g., 30 or more genes, 30 or more chromatin sites, 30 or more loci where histone modifications are measures, 30 or more sites where epigenetic modifications such asmethylation are measure, etc.). The plurality of cells collectively originates from a plurality of samples.

[0012] In some embodiments, the plurality of cells comprises 10 cells, 50 cells, 100 cells, 500 cells, or 1000 cells.

[0013] In some embodiments, the plurality of samples comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 samples.

[0014] In some embodiments, the corresponding measurement value of each respective feature in the plurality of features originating from the respective cell is determined using an scRNA, scRNA-seq or scATAC-seq based sequencing reaction that produces a plurality of sequence reads. Moreover, the plurality of sequence reads is mapped to a reference genome. In some such embodiments, the plurality of sequence reads is at least 1 x 106sequence reads. In some embodiments, each sequence read in the plurality of sequence reads is associated with a unique molecule identifier (UMI).

[0015] The method further comprises using the plurality of sequence reads to construct a first histogram of 100 equal percentile bins of a UMI count covariate. A left elbow point is found from a curve formed by the first histogram (e.g., using a Kneedie algorithm). Each cell that is in a portion of the first histogram to the left of the left elbow point is removed from the plurality of cells.

[0016] In some embodiments, each respective cell that has a number of unique UMI that exceeds a threshold percentile (e.g. a threshold percentile of 95 percent) for the plurality of cells is removed from the plurality of cells.

[0017] In some embodiments, the plurality of sequence reads is used to construct a second histogram of 100 equal percentile bins of a mitochondrial percentage. A right elbow point is found from a curve formed by the second histogram (e.g., using a Kneedie algorithm). Each cell that is in a portion of the histogram to the right of the right elbow point is removed from the plurality of cells.

[0018] The measurement value of each respective feature in the plurality of features of each respective cell in the plurality of cells is used to cluster the plurality of cells into a plurality of clusters.

[0019] In some embodiments, the corresponding measurement value of each respective feature in the plurality of features originating from each respective cell in the plurality of cells is dimension reduced prior to clustering.

[0020] In some embodiments, the plurality of clusters comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 clusters.

[0021] For each respective cell in the plurality of cells (i) a cluster assignment of the respective cell and (ii) an identity of a sample in the plurality of samples from which the respective cell originated is used to compute a plurality of sample-level local inverse Simpson index (LISI) scores. The plurality of sample-level LISI scores are penalized by a measure of importance of cells in each sample in the plurality of samples. In some such embodiments this is determined for each sample i by computing: sample-level local inverse Simpson indexwhere w(-7is a weight assigned to cluster j in sample i based on a measure of importance of cluster j to sample z, Pj is a proportion of cells in cluster j that originate from sample z, and N is the number of clusters in the plurality of clusters.

[0022] In some embodiments, the measure of importance of cluster j to sample z is a frequency of occurrence of cells in cluster j from sample z.

[0023] In some embodiments, the measure of importance of cluster j to sample z is a number of cells in cluster j from sample z.

[0024] The plurality of sample-level LISI scores is scaled to range from zero and one, thereby forming a plurality of scaled sample-level LISI scores.

[0025] The plurality of scaled sample-level LISI scores are numerically ordered (e.g., from small to large or vice versa), thereby forming an ordered plurality of scaled samplelevel LISI scores. In some embodiments, each respective score in the plurality of scaled sample-level LISI scores, a difference is obtained by subtracting the current score from its consecutive score (e.g., 0.1, 0.3, 0.35, 0.4, 0.480.2, 0.05, 0.05, 0.08), thereby forming an ordered plurality of differences of scaled sample-level LISI scores. A Dixon’s Q-test is applied to the ordered plurality of differences of scaled sample-level LISI scores thereby grouping the numerically ordered plurality of scaled sample-level LISI scores into one or more groups based on the location of the significant difference of scaled sample-level LISI scores.

[0026] A batch assignment for each sample in the plurality of samples is determined based on an assignment of corresponding scaled sample-level LISI scores in the plurality of scaled sample-level LISI scores common to the one or more groups.

[0027] In some embodiments, the one or more groups comprises three or more groups. Each group in the three or more groups is given a batch assignment. Each batch assignment is assigned at least one sample in the plurality of samples.

[0028] The batch assignment of each sample in the plurality of samples is used to batch correct a corresponding abundance of each respective feature in the plurality of features for each respective cell in the plurality of cells.

[0029] In some embodiments, the plurality of samples includes a first subset of samples that have a first condition and a second subset of samples that are free of the first condition and the method further comprises using the batch corrected corresponding measurement value of each respective feature in the plurality of features in each respective cell in the plurality of cells to associate one or more features in the plurality of features that have a statistically different measurement value between the first subset and the second subset with the condition.

[0030] In some such embodiments, the condition is a cancer, hematologic disorder, autoimmune disease, inflammatory disease, immunological disorder, metabolic disorder, neurological disorder, genetic disorder, psychiatric disorder, gastroenterological disorder, renal disorder, cardiovascular disorder, dermatological disorder, respiratory disorder, or a viral infection.

[0031] Another aspect of the present disclosure provides a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors, the one or more programs comprising instructions for performing any of the methods and / or embodiments disclosed herein.

[0032] Another aspect of the present disclosure provides a non-transitory computer readable storage medium storing one or more programs configured for execution by a computer, the one or more programs comprising instructions for carrying out any of the methods and / or embodiments disclosed herein.

[0033] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will berealized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0034] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entireties for all purposes to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The implementations disclosed herein are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. Like reference numerals refer to corresponding parts throughout the several views of the drawings.

[0036] FIGs. 1 A and IB illustrate an exemplary system topology, including a computer system, for batch inference of single-cell omics data, in accordance with an embodiment of the present disclosure.

[0037] FIGs. 2A, 2B, 2C, 2D, and 2E provide a flow chart of an example method for batch inference of single-cell omics data, in which dashed boxes represent optional elements in the flow chart, in accordance with an embodiment of the present disclosure.

[0038] FIG. 3 illustrates a one-step automated multiple sample single-cell analysis pipeline that does not require any parameter selection by the user, and includes a data-driven batch inference framework to improve the quality of integration. The framework takes the single-cell measurement data output and implements a workflow comprising quality control with automated parameter selection (Step 1), application of the data-driven batch inference pipeline (Step 2), data integration (Step 3), cell type annotation (Step 4), and optional downstream differential and pathway analyses (Step 5, and Example 6). The program returns an annotated, integrated data matrix (e.g., for single-cell RNA-seq and / or for single-cellATAC-seq), as well as selected analyses, in accordance with an embodiment of the present disclosure.

[0039] FIG. 4 illustrates a web interface for batch inference of single-cell abundance data in accordance with an embodiment of the present disclosure.

[0040] FIGs. 5A and 5B illustrate application of an information-theoretic approach to quantify sample distributions and identify local and global batch effects based on learned distributions in accordance with the present disclosure. The algorithm starts off with a low- resolution clustering that separates putative cell types. It then iterates through the clusters to determine if a group of samples is significantly different from the rest of the samples and assigns batch labels locally. In cases where a subset of samples within the significant batch has already been assigned a batch label in the previous iteration, the framework further divides the significant batch by giving the previously unassigned samples a new batch label. This process is repeated until all local batches are stable, which constitutes the final global batch assignment.

[0041] FIGs. 6A and 6B illustrate benchmarking the disclosed batch inference framework on public human lung atlas scRNA-seq dataset using four single-cell integration methods in accordance with an embodiment of the present disclosure. In a public atlas-level human lung scRNA-seq dataset comprising data from three different datasets, batches were labeled using either the sample ID, the dataset ID, or the data-driven method of the present disclosure. The three types of batch labels were used as input for four different batch correction and data integration packages (Harmony, Seurat CCA, Seurat RPC A, Scanorama). The results of the integration with different batch labeling approaches were compared using metrics for cell type coherence and for batch removal. (A) Evaluation of the disclosed batch inference in preserving cell type coherence among distinct cell types after integration. Each dot represents a sample score for a total of 16 samples. Each score quantifies the effectiveness in preserving the integrity of different cell types within that sample. A higher score indicates that the cells of a particular cell type are more distinctly separated from cells of other types after integration. Pairwise nonparametric Wilcoxon rank sum tests were performed, and p-values were Bonferroni-adjusted. (B) Evaluation of the disclosed batch inference in eliminating batch variants among samples after integration. Each dot represents a cell type for a total of 17 types. Each score represents how effectively batches are mitigated within the associated cell type. A higher score indicates that cells of the same type, regardless of their sample origin, are better inter-mixed after integration. Pairwisenonparametric Wilcoxon rank sum tests were performed, and p-values were Bonferroni- adjusted.

[0042] FIG. 7 illustrates application of batch inference in accordance with an embodiment of the present disclosure to a human peripheral blood mononuclear cell (PBMC) scRNA-seq study. Before integration, the poor registration of samples in the UMAP indicates the presence of severe batch effects. Batches were identified using either the automated data-driven batch inference framework implemented in accordance with the present disclosure or batch information supplied by the experimentalists.

[0043] FIGs. 8A and 8B illustrate strong batch effects present in the human PBMC scRNA-seq datasets across all major cell types. A) UMAP shows cells are segregated by sample IDs indicating strong batch effects present in all major cell types of the human PBMC scRNA-seq data. B) Louvain clustering at resolution of 0.2 produced 23 clusters. Adjusted Rand Index of cells between sample IDs and cell clusters is 0.323.

[0044] FIG. 9 illustrates UMAPs of a T / NK cell population after integration. In contrast to the results obtained with batch inference in accordance with an embodiment of the present disclosure (SPEEDI), the data integration based on supplied batches led to fragmentation of some cell type clusters (CD4 T memory, CD8 CTL) and failed to distinguish others (NK Proliferating).

[0045] FIGs. 10A and 10B illustrate integration and annotation data for all cell types in human PBMC scRNA-seq data in accordance with an embodiment of the present disclosure. A) UMAP of the batch-based integration and reference-based annotation of the present disclosure. B) UMAP of supplied batch-based integration and reference-based annotation.

[0046] FIG. 11 illustrates UMAPs of integrated T / NK lymphocytes colored by sample IDs. A) When integrating using supplied batch labels, UMAP colored by sample ID shows that some samples remain outliers even after integration. B) Outlier samples are minimal in the batch integration in accordance with the present disclosure.

[0047] FIG. 12 illustrates how UMAPs of gene expression of canonical markers confirm the inferred cell types made by the batch inference framework in accordance with some embodiments of the present disclosure.

[0048] FIG. 13 illustrates higher correlation of cell type proportions determined by scRNA-seq and by CyTOF mass cytometry on the same samples observed using the data- derived batches determined in accordance with an embodiment of the present disclosure ascompared to supplied batches. Correlation of cell type proportions determined by scRNA- seq and by CyTOF mass cytometry on the same samples using an embodiment of the disclosed batch labeling pipeline and supplied batches are shown. Each dot represents a sample that has both scRNA-seq and CyTOF data generated. Only cell types reported by both assays were used to calculate the proportions.

[0049] FIG. 14 illustrates PBMC T / NK lymphocytes cell type proportion Spearman correlations for each individual sample between cell types inferred from a batch inference framework in accordance with the present disclosure (SPEEDI) and supplied-batch versus CyTOF.

[0050] FIG. 15 illustrates how correlation of cell type proportions was determined by scRNA-seq and by CyTOF mass cytometry on the same samples was observed using data- derived batches inferred from a batch inference framework in accordance with the present disclosure (SPEEDI) compared to those determined by silhouette. Each dot represents a sample score, and each score quantifies the effectiveness in preserving the integrity of different cell types within that sample.

[0051] FIG. 16 illustrates how correlation of cell type proportions was determined by scRNA-seq and by CyTOF mass cytometry on the same samples was observed using data- derived batches inferred from a batch inference framework in accordance with the present disclosure (SPEEDI) compared to those determined by connectivity. Each dot represents a cell type score, and each score represents how effectively batches are mitigated within the associated cell type.

[0052] FIG. 17 illustrates that, for same-cell scRNA-seq and scATAC-seq multi ome datasets generated from 14 wild-type female murine pituitaries, a batch inference pipeline in accordance with the present disclosure identified consistent batches with both data types. Data from each assay were integrated and annotated by a batch inference pipeline in accordance with the present disclosure.

[0053] FIG. 18 illustrates how annotation of individual cells by cell type for the RNA- seq and for the ATAC-seq data of FIG. 17, after using a batch inference pipeline in accordance with the present disclosure. The figure shows a heatmap representation of the contingency table that compares the annotation of individual cells by cell type for the scRNA-seq and for the scATAC-seq data after using an embodiments of the disclosed batch inference framework. The rows represent the cell type annotation for scATAC-seq, and thecolumns represent those for scRNA-seq. Each cell ranges between 0 and 1, where the value indicates the percentage of overlapping barcodes (median cell subtype identification overlap = 0.96).

[0054] FIGs. 19A and 19B illustrate a summary of cell type representations across all samples in the human lung atlas data compendium.

[0055] FIG. 20 provides human PBMC study CyTOF cell type proportions used in validation experiments in accordance with the present disclosure.

[0056] FIG. 21 A illustrates a UMAP of integrated scRNA-seq data from a longitudinal SARS-CoV-2 study, in accordance with an embodiment of the present disclosure.

[0057] FIG. 2 IB illustrates functional module discovery on upregulated NK genes, in accordance with an embodiment of the present disclosure.

[0058] FIGs. 22A, 22B, 22C, 22D, 22E, 22F, and 22G. The present disclosure facilitates integrative cell type identification from scRNA-seq data. (A) The batch inference of the present disclosure was applied to a human peripheral blood mononuclear cell (PBMC) scRNAseq study with 20 subjects. Batches were identified using either the automated datadriven batch inference method implemented in SPEEDI (n = 12) or using sample IDs (n= 20). The UMAPs show the T / NK cell population before integration. (B) The UMAPs of the T / NK cell population after integration with both batch labeling strategies are shown. (C) Correspondence between the data-defined batch labels and sample IDs. The disclosed inference method identified a complex pattern of 12 data-defined batches that varied in size from 1 to 5 samples. (D) Score for cell type coherence. For cell type coherence measures, each dot represents a sample score, and each score quantifies the effectiveness in preserving the integrity of different cell types within that sample. Pairwise nonparametric Wilcoxon rank sum tests were performed. Scores for batch removal. For batch effect removal metrics, each dot represents a cell type score, and each score represents how effectively batches are mitigated within the associated cell type. Pairwise nonparametric Wilcoxon rank sum tests were performed. *p<0.05, *** p<0.001, n.s. Not-significant (p>0.05). Bonferroni corrected t-test. (E) The automated batch labeling of the present disclosure corrected the fragmentation of CD14 monocytes and Treg cells into different clusters as seen when using sample ID batch correction (integrated data for all cell types in human PBMC scRNA-seq data colored by cell types). Application of the four different batch correction quality scoring metrics of Example 5 showed that using the disclosed batch identification performed better than using sample IDsfor three of the metrics and comparably for the remaining metric. (F) When using sample ID labels for batch correction, CD8+ memory cells were more broadly distributed, intermingled with other T cells, and showed less overlap with CD8A versus using data-defined batch labels for batch correction (sample-based integration fails to recover distinct CD8+ T cell populations). (G) UMAPs of gene expressions of canonical marker genes (NKG7 for NK cells, SELL for CD4 naive cells, TNFRSF4 for CD4 memory cells, and FOXP3 for Treg cells) confirm the inferred cell types produced by the systems and methods of the present disclosure.DETAILED DESCRIPTION

[0059] The implementations described herein provide various technical solutions for inferring batches in single-cell omics data. As illustrated in Figure 3, embodiments of the present disclosure provide a fully automated end-to-end quality control, data-driven batch identification, data integration, and cell-type labeling pipeline that does not require any manual parameter selection or pipeline assembly. The present disclosure provides an automated data-driven batch inference framework, overcoming the problem of unknown or under-specified batch effects. It works with single-cell omics data, such as scRNA, scATAC, or sc multiome datasets and matches or exceeds benchmarking against existing manual, piecemeal methods. Unlike these other methods that also require a high level of bioinformatics expertise, the disclosed pipeline is also available in a web interface, as illustrated in Figure 4, making it accessible to a wide variety of researchers. The full automation of batch inference improves reproducibility and robustness of analyses, and democratizes the interpretation of single-cell omics datasets.

[0060] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0061] Definitions.

[0062] As used herein, the term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, in some embodiments “about” means within 1 or more than 1 standard deviation, per the practice in the art. In some embodiments, “about” means a range of ±20%, ±10%, ±5%, or ±1% of a given value. In some embodiments, the term “about” or “approximately” means within an order of magnitude, within 5-fold, or within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value can be assumed. All numerical values within the detailed description herein are modified by “about” the indicated value, and consider experimental error and variations that would be expected by a person having ordinary skill in the art. The term “about” can have the meaning as commonly understood by one of ordinary skill in the art. In some embodiments, the term “about” refers to ±10%. In some embodiments, the term “about” refers to ±5%.

[0063] As used herein, the term “subject,” “training subject,” or “test subject” refers to any living or non-living organism, including but not limited to a human (e.g., a male human, female human, fetus, pregnant female, child, or the like) and / or a non-human animal. Any human or non-human animal can serve as a subject, including but not limited to mammal, reptile, avian, amphibian, fish, ungulate, ruminant, bovine (e.g., cattle), equine (e.g., horse), caprine and ovine (e.g., sheep, goat), swine (e.g., pig), camelid (e.g., camel, llama, alpaca), monkey, ape (e.g., gorilla, chimpanzee), ursid (e.g., bear), poultry, dog, cat, mouse, rat, fish, dolphin, whale, and shark. The terms “subject” and “patient” are used interchangeably herein and can refer to a human or non-human animal who is known to have, or potentially has, a medical condition or disorder, such as, e.g., kidney disease. In some embodiments, a subject is a “normal” or “control” subject, e.g., a subject that is not known to have a medical condition or disorder. In some embodiments, a subject is a male or female of any stage (e.g., a man, a woman, or a child).

[0064] As used herein, the terms “control,” “healthy,” and “normal” describe a subject and / or an image from a subject that does not have a particular condition (e.g., kidney disease), has a baseline condition (e.g., prior to onset of the particular condition), or is otherwise healthy. In an example, a method as disclosed herein can be performed to diagnose a renal disease and / or a kidney graft failure in a subject having a renal disease using a trainedmodel, where the model is trained using one or more training images obtained from the subject prior to the onset of the condition (e.g., at an earlier time point), or from a different, healthy subject. A control image can be obtained from a control subject, or from a database.

[0065] The term “normalize” as used herein means transforming a value or a set of values to a common frame of reference for comparison purposes. For example, when one or more pixel values corresponding to one or more pixels in a respective image are “normalized” to a predetermined statistic (e.g., a mean and / or standard deviation of one or more pixel values across one or more images), the pixel values of the respective pixels are compared to the respective statistic so that the amount by which the pixel values differ from the statistic can be determined.

[0066] Clustering. In some embodiments, clustering refers to unsupervised clustering. In some embodiments, the model is a supervised clustering model. Clustering algorithms suitable for use in the present disclosure are described, for example, at pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York, (hereinafter “Duda 1973”) which is hereby incorporated by reference in its entirety. The clustering problem can be described as one of finding natural groupings in a dataset. To identify natural groupings, two issues can be addressed. First, a way to measure similarity (or dissimilarity) between two samples can be determined. This metric (e.g., similarity measure) can be used to ensure that the samples in one cluster are more like one another than they are to samples in other clusters. Second, a mechanism for partitioning the data into clusters using the similarity measure can be determined. One way to begin a clustering investigation can be to define a distance function and to compute the matrix of distances between all pairs of samples in a training dataset. If distance is a good measure of similarity, then the distance between reference entities in the same cluster can be significantly less than the distance between the reference entities in different clusters. However, clustering may not use a distance metric. For example, a nonmetric similarity function s(x, x') can be used to compare two vectors x and x'. s(x, x') can be a symmetric function whose value is large when x and x' are somehow “similar.” Once a method for measuring “similarity” or “dissimilarity” between points in a dataset has been selected, clustering can use a criterion function that measures the clustering quality of any partition of the data. Partitions of the data set that extremize the criterion function can be used to cluster the data. Particular exemplary clustering techniques that can be used in the present disclosure can include, but are not limited to, hierarchical clustering (agglomerative clustering using a nearest-neighboralgorithm, farthest-neighbor algorithm, the average linkage algorithm, the centroid algorithm, or the sum-of-squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, and Jarvis-Patrick clustering. In some embodiments, the clustering comprises unsupervised clustering (e.g., with no preconceived number of clusters and / or no predetermination of cluster assignments).

[0067] The terms “sequence reads” or “reads,” used interchangeably herein, refer to nucleotide sequences produced by any sequencing process described herein or known in the art. Reads can be generated from one end of nucleic acid fragments (“single-end reads”), and sometimes are generated from both ends of nucleic acids (e.g., paired-end reads, double-end reads). The length of the sequence read is often associated with the particular sequencing technology. High-throughput methods, for example, provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are of a mean, median or average length of about 15 bp to 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp. In some embodiments, the sequence reads are of a mean, median or average length of about 1000 bp or more. Nanopore sequencing, for example, can provide sequence reads that vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads vary to a lesser extent (e.g., where most sequence reads are of a length of about 200 bp or less). A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from part of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides of the entire nucleic acid fragment. A sequence read can be obtained in a variety of ways, e.g., using sequencing techniques or using probes (e.g., in hybridization arrays or capture probes) or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.

[0068] As disclosed herein, the terms “sequencing,” “sequence determination,” and the like refer generally to any and all biochemical processes that may be used to determine the order of biological macromolecules such as nucleic acids or proteins. For example,sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule such as a DNA fragment.

[0069] Several aspects are described below with reference to example applications for illustration. Numerous specific details, relationships, and methods are set forth to provide a full understanding of the features described herein. The features described herein can be practiced without one or more of the specific details or with other methods. The features described herein are not limited by the illustrated ordering of acts or events, as some acts can occur in different orders and / or concurrently with other acts or events. Furthermore, not all illustrated acts or events are used to implement a methodology in accordance with the features described herein.

[0070] Example Systems.

[0071] Figure 1 illustrates a computer system 100 for batch inference of single-cell omics data. In typical embodiments, computer system 100 comprises one or more computers. For purposes of illustration in Figure 1, the computer system 100 is represented as a single computer that includes all of the functionality of the disclosed computer system 100. However, the present disclosure is not so limited. The functionality of the computer system 100 may be spread across any number of networked computers and / or reside on each of several networked computers and / or virtual machines. One of skill in the art will appreciate that a wide array of different computer topologies is possible for the computer system 100 and all such topologies are within the scope of the present disclosure.

[0072] Turning to Figure 1 with the foregoing in mind, the computer system 100 comprises one or more processing units (CPUs) 52, a network or other communications interface 54, a user interface 56 (e.g., including an optional display 58 and optional input 60 (e.g. keyboard or other form of input device), a memory 92 (e.g., random access memory, persistent memory, or combination thereof), and one or more communication busses 94 for interconnecting the aforementioned components. To the extent that components of memory 92 are not persistent, data in memory 92 can be seamlessly shared with non-volatile memory (not shown) or portions of memory 92 that are non-volatile / persistent using known computing techniques such as caching. Memory 92 can include mass storage that is remotely located with respect to the central processing unit(s) 52. In other words, some data stored in memory 92 may in fact be hosted on computers that are external to computer system 100 but that can be electronically accessed by the computer system 100 over an Internet, intranet, orother form of network or electronic cable using network interface 54. In some embodiments, the computer system 100 makes use of models that are run from the memory associated with one or more graphical processing units in order to improve the speed and performance of the system. In some alternative embodiments, the computer system 100 makes use of models that are run from memory 92 rather than memory associated with a graphical processing unit.

[0073] The memory 92 of the computer system 100 stores:• an optional operating system 102 that includes procedures for handling various basic system services;• a batch inference framework module 104 for inferring batches in single-cell omics data;• an omnibus single-cell dataset 106 comprising measurement data for each of a plurality of cells, where the plurality of cells collectively originate from a plurality of different sources, and for each respective cell 108 (108-1, 108-2, . . ., 108-M) in the plurality of cells an indication of cell sample source 110 (110-1, 110-2, . . . 110-M) in a plurality of samples sources for the respective cell and a measurement value of each feature 112 in a plurality of features (e.g, 112-1-1, 112-1-2, 112-1-N) in the respective cell;• a single-cell dataset clustering 114 comprising a plurality of clusters 116 (116-1, 116- 2, ..., 116-L), each cluster including a subset (e.g., 118-1-1, 118-1-2, ..., 118-1-K) of the plurality of cells;• an identification 118 of sample sources that originated cells in the plurality of cells of the omnibus single-cell dataset 106, and for each sample source 120 (e.g., 120-1, 120- 2, ..., 120-W), a scaled sample-level LISA score 122 (e.g., 122-1, 122-2, ..., 122-W);• group assignments 124 for the plurality of cells, each group 126 (e.g., 126-1, . . ., 126- Q) including a subset 128 (e.g., 128-1, . . ., 128-Q) of the samples in the plurality of samples.

[0074] In some implementations, one or more of the above identified data elements or modules of the computer system 100 are stored in one or more of the previously mentioned memory devices, and correspond to a set of instructions for performing a function described above. The above identified data, modules or programs (e.g., sets of instructions) need not beimplemented as separate software programs, procedures or modules, and thus various subsets of these modules may be combined or otherwise re-arranged in various implementations. In some implementations, the memory 92 optionally stores a subset of the modules and data structures identified above. Furthermore, in some embodiments the memory 92 stores additional modules and data structures not described above.

[0075] Example methods.

[0076] Now that an exemplary system 100 for inferring batches in single-cell omics data has been described, exemplary methods for inferring batches in single-cell omics data are disclosed in conjunction with FIGs. 2A, 2B, 2C, 2D, and 2E. Although the methods discuss processing a plurality of single cells, it will be appreciated that the disclosed methods work for single nuclei data as well. Thus, it will be appreciated that all references to processing and batch inference of single-cell data in the present disclosure is also applicable to single nuclei data.

[0077] Referring to block 200, a method for inferring batches in single-cell omics data is provided (e.g., at a computer system comprising one or more processors and a non-transitory computer-readable medium including computer-executable instructions). In some embodiments some or all of the blocks disclosed in FIG. 2 and described below are performed by the batch inference framework module 104 of Figure 1 A.

[0078] Referring to block 202, in some embodiments, there is obtained, for each respective cell 108 in a plurality of cells, a corresponding measurement value of each respective feature 112 in a plurality of features originating from the respective cell. In some embodiments the plurality of features comprises 30 or more features (gene abundance, amount of chromatin accessibility at a particular locus, count of a particular histone modification, count of a particular epigenetic feature such as methylation) . In some embodiments the plurality of features comprises 10 or more features or 20 or more features. In some embodiments the plurality of features comprises 40 or more features, 60 or more features, 100 or more features, 500 or more features, 1000 or more features, 2000 or more features, 5000 or more features, or 10,000 or more features. The plurality of cells collectively originates from a plurality of samples.

[0079] For scRNA-seq data, in some embodiments raw data in unique molecular identified (UMI) count matrix format is obtained in block 202. For scRNA, in some embodiments sample data is filtered data, for example generated by Cell Ranger in the MEXformat. For scRNA, in some embodiments sample data is filtered data in H5 format files. In some embodiments, if working with MEX files, sample data follows the standard naming convention (three files with names barcodes.tsv.gz, features. tsv.gz, and matrix. mtx.gz). If working with H5 files, in some embodiments sample data file names end with filtered_feature_bc_matrix.h5. For scATAC-seq data, the fragment files are used instead in some embodiments.

[0080] In some embodiments, each sample consists of whole blood. In some embodiments, each sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid from the test subject. In some embodiments, each sample consists of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid from a different subject.

[0081] In some embodiments, each sample is a solid biopsy from a different subject.

[0082] In some embodiments, each sample is a tumor sample from a different subject.

[0083] In some embodiments, each sample is a different cell line.

[0084] Referring to block 204, in some embodiments, the plurality of cells comprises 10 cells, 50 cells, 100 cells, 500 cells, or 1000 cells. In some embodiments the plurality of cells comprises 2000 or more cells, 4000 or more cells, 6000 or more cells, 8000 or more cells, 10,000 or more cells, 20,000 or more cells, or 30,000 or more cells.

[0085] Referring to block 206, in some embodiments, the plurality of samples comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 samples. In some embodiments, each sample in the plurality of samples is from a different subject. In some embodiments, each sample in the plurality of samples is sequenced on a different day, week or month. In some embodiments, each sample in the plurality of samples is from a different hospital, medical facility, or geographic location. In some embodiments, each sample (also considered termed a source herein) comprises 10 cells, 50 cells, 100 cells, 500 cells, or 1000 cells. In some embodiments the plurality of cells comprises 2000 or more cells, 4000 or more cells, 6000 or more cells, 8000 or more cells, 10,000 or more cells, 20,000 or more cells, or 30,000 or more cells.

[0086] Referring to block 208, in some embodiments, the corresponding measurement value of each respective feature in the plurality of features originating from the respective cell is determined using an scRNA, scRNA-seq or scATAC-seq based sequencing reaction that produces a plurality of sequence reads. In some embodiments each measurement value is anabundance (e.g., abundance of a RNA mapping to a particular gene), a peak height (e.g., an area of a chromatin peak, etc.) or some other quantitative amount representative of a feature. In some embodiments the corresponding measurement value of each respective feature in the plurality of features originating from the respective cell is determined by any single cell or single nucleus assay that allows resolution of clusters of different cell types in an analysis of a mixed cell type tissue sample. This includes, but is not limited to, single-cell RNAseq, single-nucleus RNAseq, single nucleus ATACseq, single nucleus true multiomeRNAseq / ATACseq, and single nucleus assays for histone modifications (e.g., single nucleus Cut&Tag). Single-cell CUT&Tag is described in Bartosovic, 2021, “Single-cell CUT&Tag profiles histone modifications and transcription factors in complex tissues,” Nat. Biotechnol. 39(7), pp. 825-835, which is hereby incorporated by reference.

[0087] The plurality of sequence reads is mapped to a reference genome. In some embodiments, the reference genome is “hg38” and “mm 10” for human and mouse, respectively. More specifically, “hg38” represents the UCSC full genome sequence hg38 for Homo Sapiens available in Bioconductor package “BSgenome.Hsapiens.UCSC.hg38”, while “mmlO” comes from the “BSgenome.Mmusculus.UCSC.mmlO” package, which is the UCSC full genome sequences for Mus musculus version mmlO based on GRCm38.p6.

[0088] In some such embodiments, this mapping comprises aligning each respective sequence read in the plurality of sequence reads to a reference sequence such as a genome reference. In some embodiments, the reference sequence is a reference human genome. In In some embodiments, the mapping algorithm used to perform the mapping is TopHat2. See, Pertea et al., 2013, “TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions,” Genome Biol. 14(4):R36, which is hereby incorporated by reference. In some embodiments, the mapping algorithm used to perform the mapping is Bowtie or Bowtie2. See Friedel, 2012, “A comprehensive evaluation of alignment algorithms in the context of RNA-seq,” PLoS One 7(12):e52403, which is hereby incorporated by reference. In some embodiments, the mapping algorithm used to perform the mapping is a BWA (Burrows-Wheeler Aligner). See, for example, Durbin, 2009, “Fast and accurate short read alignment with Burrows-Wheeler transform,” Bioinformatics 25(14): 1754-60, which is hereby incorporated by reference. In some embodiments, the mapping algorithm used to perform the mapping is GMAP (Genomic Mapping and Alignment Program). GMAP is an alignment tool specifically designed for aligning RNAseq reads to complex genomes, handling both spliced and non-spliced alignments. See Wuand Watanabe, 2005, “GMAP: a genomic mapping and alignment program for mRNA and EST sequences,” Bioinformatics 21(9): 1859-75, which is hereby incorporated by reference.

[0089] Referring to block 210, in some embodiments, the plurality of sequence reads is at least 1 x 106sequence reads. In some embodiments, the plurality of sequence reads comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, at least 5 million, at least 6 million, at least 7 million, at least 8 million, at least 9 million, or more sequence reads. In some embodiments, the plurality of sequence reads comprises at least 1 x 107, at least 2 x 107, at least 3 x 107, at least 4 x 107, at least 5 x 107, at least 6 x 107, at least 7 x 107, at least 8 x 107, at least 9 x 107, at least 1 x 108, at least 2 x 108, at least 3 x 108, at least 4 x 108, at least 5 x 108, at least 6 x 108, at least 7 x 108, at least 8 x 108, at least 9 x 108, at least 1 x 109, or more sequence reads. In some embodiments, the plurality of sequence reads consists of no more than 5 x 107, no more than 1 x 107, no more than 5 x 106, no more than 4 x 106, no more than 3 x 106, no more than 2 x 106, no more than 1 x 106, no more than 500,000, no more than 100,000, no more than 50,000, no more than 30,000, no more than 20,000, no more than 10,000, no more than 9000, no more than 8000, no more than 7000, no more than 6000, no more than 5000, no more than 4000, no more than 3000, no more than 2000, no more than 1000, or less sequence reads.

[0090] In some embodiments, the plurality of sequence reads consists of between 1000 to 5000, from 1000 to 10,000, from 2000 to 20,000, from 5000 to 50,000, from 10,000 to 100,000, from 100,000 to 500,000 from 10,000 to 500,000, from 500,000 to 1 million, from 1 million to 30 million, from 30 million to 80 million, or from 10 million to 500 million sequence reads. In some embodiments, the plurality of sequence reads falls within another range starting no lower than 1000 sequence reads and ending no higher than 1 x 109sequence reads.

[0091] Referring to block 212, in some embodiments, each sequence read in the plurality of sequence reads is associated with a unique molecule identifier (UMI). In such embodiments, the method further comprises using the plurality of sequence reads to construct a first histogram of a UMI count covariate. A left elbow point is found from a curve formed by the first histogram (e.g., using a Kneedie algorithm). While the term “elbow” is typically applied to the inflection of a convex curve and “knee” to that of a concave curve, for simplicity the term “elbow” is used herein for both situations.

[0092] The Kneedie algorithm systematically finds the elbow or knee point in the distribution of transcript counts per cell. See Xue et al., 2022, “Liver tumour immune microenvironment subtypes and neutrophil heterogeneity,” Nature 612(7938), pp. 141-147, which is hereby incorporated by reference.

[0093] Conventionally, the elbow point is considered as the turning point where a distribution changes its behavior due to some latent factor. In single-cell data analysis, it can be interpreted as the point where the profiled barcodes transit into a different quality phase. In the case of a unimodal distribution, there are typically left and right elbow points that represent the lower and the upper threshold, respectively, for our algorithm. Data below this lower threshold are normally associated with broken cells where cytoplasmic mRNAs are escaping from the cell membrane.

[0094] To calculate the lower threshold of UMI counts LUMI, the batch inference framework begins by creating the first histogram as a histogram of UMI count covariate. For purposes of illustration the first histogram is a 100 group histogram, although smaller numbers (e.g., 50) or larger numbers (e.g., 200 can be used). The histogram effectively creates a set of discrete data points Di where:and Xi is the ithpercentile of UMI counts where i E N+and 0 < i < 100 (for the example of a 100 group histogram) and is the frequency of the zthpercentile of UMI counts and ytE N+.

[0095] To find the left elbow point from the curve formed by point set Dt, the subset D5. j is found between

[0096] The Kneedie algorithm first rotates the curve segment defined by Ds. clockwise for 0 degrees so that the curve is concave and the line formed betweenand ( Vx5.max ,' • VZ 5emax )' is horizontal. It then finds the local maximum ( V x5. I,m , 'Jy5, I,m J ) of the rotated curve. The lower elbow point is thus defined asEUMI — P(.sim) 1where P(sim) is the simthpercentile of UMI counts, 0 < sim< 100 (for the example of a 100 group histogram).

[0097] In some embodiments, each cell that is in a portion of the first histogram to the left of the left elbow point is removed from the plurality of cells.

[0098] In some embodiments, to restrict the algorithm from filtering too many cells, an upper limit (UL) of UMI counts is imposed, such that the final threshold of gene count is defined asLu MI = min(EUMI, UL').

[0099] In some embodiments, the upper limit is 1000 UMI counts. In such embodiments the final threshold of gene count is defined asIn other words, LUMIis equal to EUMIunless EUMIhas less than 1000 UMI counts. If EUMIhas less than 1000 UMI counts, than LUMIis set to 1000 UMI counts and those cells that have fewer then 1000 UMI counts are removed from the plurality of cells prior to the clustering described below. In some embodiments, a different upper limit is imposed, such as an upper limit that is between 200 UMI counts and 3000 UMI counts.

[0100] In some embodiments, the quality control (filtering) of block 212 is performed on individual sample datasets to avoid detecting false positive outliers that are distributed differently because of biological or technological factors e.g., rare cell types detected in a subset of samples or different sequencing depths). That is, although Figure 1 describes an omnibus single-cell dataset 106, this dataset may in fact be comprises of any number of component single-cell datasets and the quality control filtering of block 212 can be performed on each such dataset.

[0101] Referring to block 214, in some embodiments, each respective cell 108 that has a number of unique UMI that exceeds a threshold percentile for the plurality of cells (e.g. the threshold percentile is 95 percent) is removed from the plurality of cells. The purpose of this filter is to remove cells that have excessively high UMI counts. In the case where the threshold percentile is 95 percent:UUMI = ^(95).In some embodiments, a different threshold percentile is used, such as 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99 percent.

[0102] In some embodiments, the quality control (filtering) of block 214 is performed on individual sample datasets to avoid detecting false positive outliers that are distributed differently because of biological or technological factors (e.g., rare cell types detected in a subset of samples or different sequencing depths).

[0103] Referring to block 216, in some embodiments, the plurality of sequence reads is used to construct a second histogram of mitochondrial percentage. A right elbow point is found from a curve formed by the second histogram (e.g., using a Kneedie algorithm). Each cell that is in a portion of the histogram to the right of the right elbow point is removed from the plurality of cells. This is similar to the algorithm detailed in block 212. Thus, if EMT is defined to be the right elbow point calculated from the curve defined by the distribution of mitochondrial reads in the second histogram, the upper limit of mitochondrial content in percentage is defined as:UMT = max(EMT, P(75)'), in embodiments where EMT is used as the cutoff, provided that it does not exceed 75 percent. In alternative embodiments, EMT is used as the cutoff provided that it does not exceed a different percentile such as a percentile between 50 percent and 99 percent. In alternative embodiments, EMT is used as the cutoff, provided that it does not exceed a threshold percentile of 50 percent, 55 percent, 60 percent, 65 percent, 70 percent, 80 percent, 90 percent, or 95 percent. In the case where EMT exceeds the threshold percentile, the threshold percentile is used as the cutoff and all cells to the right of the threshold percentile are removed from the second histogram.

[0104] In some embodiments, the quality control (filtering) of block 216 is performed on individual sample datasets to avoid detecting false positive outliers that are distributed differently because of biological or technological factors (e.g., rare cell types detected in a subset of samples or different sequencing depths).

[0105] In some embodiments, quality control is performed for additional covariates. In some such embodiments, the 97th, 98thor 99thpercentile is used as the upper threshold. That is, a right elbow is found in the histogram of the covariate, and if the right elbow is less than a certain threshold percentage, that threshold percentage is used as the cutoff instead of the right elbow of the covariate. For example, consider the case where the right elbow for acovariate is at the 97thpercentile marker, meaning that 3 percent of the cells lie to the right of the right elbow in the histogram and 97 percent lie to the left of the right elbow. If the cutoff is the 99thpercentile, then the right elbow is not used and the 1 percent of the cells on the far right of the histogram of the covariate are removed from the plurality of cells. If, on the other hand, the cutoff is the 95thpercentile, then the right elbow is used and the 5 percent of the cells to the right of the right elbow of the histogram of the covariate are removed from the plurality of cells.

[0106] The quality control is performed on individual sample datasets to avoid detecting false positive outliers that are distributed differently because of biological or technological factors (e.g., rare cell types detected in a subset of samples or different sequencing depths). Nonlimiting examples of additional covariates include, but are not limited to the number of ribosomal RNA molecules (both large and small subunits) per barcode (see, Satopaa et al., 2011, ‘Finding a “Kneedie” in a Haystack: Detecting Knee Points in System Behavior,’ 31st International Conference on Distributed Computing Systems Workshops, Minneapolis, MN, USA, pp. 166-171, which is hereby incorporated by reference) as well as number of hemoglobins per barcode, which typically appears in PBMC (peripheral blood mononuclear cell) datasets as contamination.

[0107] In some embodiments, cell cycle scoring is also performed. In some embodiments, cell cycle scoring is performed by calling the “CellCycleScoring” function in Seurat with canonical cell cycle markers. See, Hao et al., 2021, “Integrated analysis of multimodal single-cell data,” Cell 184(13), pp. 3573-3587, which is hereby incorporated by reference.

[0108] After performing the above-identified quality control, the processed samples are merged into one dataset for normalization (e.g., the omnibus single-cell dataset 106 of Fig.1 A). This can be done for example, using Seurat’s “SCTransform” function with v2 regularization to normalize the merged count matrix. See Hafemeister et al, 2019, “Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression.” Genome biology 20(1) 296, which is hereby incorporated by reference.

[0109] In some embodiments, all quality control (QC) metrics described above as well as the cell cycle scores are used as regression covariates in normalization. In some embodiments, the normalized matrix is then scaled and centered.

[0110] In the case where the data is single-cell ATAC-seq data, in some embodiments the sample fragment files are processed with ArchR. Granja et al., 2021, “ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis,” Nature genetics 53(3), pp. 403-411, which is hereby incorporated by reference. In some such embodiments, each input read count by cell matrix is converted to a tile matrix by binning the genome-wide fragments into tiles of 500 bps. In some embodiments, scATAC-seq quality cells are selected with fixed thresholds. Cells with a number of fragments between 3,000 and 30,000, TSS enrichment > 12, and nucleosome ratio < 2 are retained in some embodiments. In some embodiments, doublets are removed using the “filterDoublets” function of ArchR under default settings.

[0111] Referring to block 218, the measurement value of each respective feature in the plurality of features of each respective cell in the plurality of cells is used to cluster the plurality of cells into a plurality of clusters. In some embodiments, Seurat is used for such clustering. See, Hao et al., 2021, “Integrated analysis of multimodal single-cell data,” Cell 184, pp. 3573-3587, which is hereby incorporated by reference. In some embodiments, ArchR is used for such clustering. See, Granja et al., 2021, “ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis.” Nature Genetics 53(3), pp. 403-411, which is hereby incorporated by reference. Such clustering is based on the measurement values of the plurality of features of each of the cells. In some embodiments, the measurement value of each feature 112 in the plurality of features of a respective cell 108 in the plurality of cells forms a vector that represents the cell. This vector is then compared to the similarly constructed vector from each of the other cells in the plurality of cells.

[0112] In some embodiments, the plurality of cells is clustered into a plurality of clusters by (i) computing a plurality of distances using the measurement data for the plurality of features for each unique pair of cells in the plurality of cells and (ii) evaluating the plurality of distances with a criterion function.

[0113] In some embodiments, the plurality of distances includes a separate distance for each unique pair of cells in the plurality of cells. Each respective distance in the plurality of distances represents a different pair of cells in the plurality of cells and quantifies a distance between (i) a respective first vector formed by the measurement data for the plurality of features for a respective first cell in the different pair of cells and (ii) a respective second vector formed by the single-cell measurement data for the plurality of features for a respective second cell in the different pair of cells, and each respective cluster 116 in theplurality of clusters represents a corresponding subset of cells of the plurality of cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of cells within the corresponding subset of cells with the criterion function.

[0114] Clustering is described at pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York, (hereinafter “Duda 1973”) which is hereby incorporated by reference in its entirety. As described in Section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings in a dataset. To identify natural groupings, two issues are addressed. First, a way to measure similarity (or dissimilarity) between two cells is determined. This metric (similarity measure) is used to ensure that the cells in one cluster are more like one another than they are to cells in other clusters based on the feature abundances. Second, a mechanism for partitioning the cells into clusters using the similarity measure is determined.

[0115] Similarity measures are discussed in Section 6.7 of Duda 1973, where it is stated that one way to begin a clustering investigation is to define a distance function and to compute the matrix of distances between all pairs of cells. If distance is a good measure of similarity, then the distance between cells in the same cluster will be significantly less than the distance between the cells in different clusters. However, as stated on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, a nonmetric similarity function s(x, x') can be used to compare two vectors x and x'. Conventionally, s(x, x') is a symmetric function whose value is large when x and x' are somehow “similar.” An example of a nonmetric similarity function s(x, x') is provided on page 218 of Duda 1973.

[0116] Once a method for measuring “similarity” or “dissimilarity” between vectors in a dataset has been selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. Partitions of the data set that extremize the criterion function are used to cluster the data. See page 217 of Duda 1973. Criterion functions are discussed in Section 6.8 of Duda 1973.

[0117] More recently, Duda el al., Pattern Classification, 2nd edition, John Wiley & Sons, Inc. New York, has been published. Pages 537-563 describe clustering in detail. More information on clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, N.Y.;Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, N.Y.; and Backer, 1995, Computer- Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, NewJersey, each of which is hereby incorporated by reference. Particular exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest-neighbor algorithm, farthest- neighbor algorithm, the average linkage algorithm, the centroid algorithm, or the sum-of- squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, Jarvis-Patrick clustering, Louvain clustering, or Leiden clustering. See, Blondel et al., July 25, 2008, “Fast unfolding of communities in large networks,” arXiv:0803.0476v2 [physical. coc-ph]; and Heumos et al., 2023, “Best practices for single-cell analysis across modalities,” Nature Review Genetics 24, 550-572, each of which is hereby incorporated by reference. Such clustering can be on the vectors formed directly from the single-cell measurement data for the plurality of features or vectors of principal components derived from such feature measurement data. In some embodiments, the clustering comprises unsupervised clustering where no preconceived notion of what clusters should form when dataset 106 is clustered are imposed.

[0118] In some embodiments before each respective first and second vector is generated, the single-cell measurement data for the plurality of features in each cell in the plurality of cells is subjected to principal component transformation to produce principal components. In such embodiments, the respective first vector formed by the single-cell measurement values for the plurality of features for a respective first cell in the different pair of cells is then the set of principal components derived for the respective first cell. Likewise, the respective second vector formed by the single-cell measurement value for the plurality of features for a respective second cell in the different pair of cells is the set of principal components derived for the respective second cell. Principal component analysis (PCA) algorithms that may be used to transform vectors of measurement data (e.g., abundance) of features (e.g., genes) to vectors of principal components are described in Jolliffe, 1986, Principal Component Analysis, Springer, New York, which is hereby incorporated by reference. PCA is also described in Draghi ci, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, which is hereby incorporated by reference. Principal components (PCs) are uncorrelated and are ordered such that the kth PC has the kthlargest variance among PCs. The kthPC can be interpreted as the direction that maximizes the variation of the projections of the data points such that it is orthogonal to the first k-1 PCs. The first few PCs capture most of the variation in dataset 106. In contrast, the last few PCs are often assumed to capture only the residual 'noise' in the dataset. In some embodiments, each respective vectorthat is clustered corresponds to a cell 106 in the plurality of cells and contains between four and five hundred principal components. In some embodiments, each respective vector that is clustered corresponds to a cell in the plurality of cells and contains between five and six hundred principal components. In some embodiments, each respective vector that is clustered corresponds to a cell in the plurality of cells and contains between six and second hundred principal components. In some embodiments, each respective vector that is clustered corresponds to a cell in the plurality of cells and contains at least 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 principal components. In some embodiments, the principal component analysis is performed on the plurality of features using the Seurat function RunPCA. In some embodiments, clusters are identified with Seurat function FindClusters, optionally using the shared nearest neighbor modular optimization based on the first 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or between 20 and 500 principal components identified by principal component analysis. In some such embodiments the clustering resolution is set to a value between 0.2 and 0.6, such as 0.4. See, Stuart et al., 2019, “Comprehensive integration of single-cell data,” Cell 177, 1888-1902, which is hereby incorporated by reference.

[0119] In some embodiments the clustering comprises k-means clustering of the plurality of cells into a predetermined number of clusters. The goal of k-means clustering is to cluster the cells based upon either the original feature measurement data or the principal components derived from the original feature measurement data for the plurality of cells into K partitions. In some embodiments, the k-means algorithm computes like clusters of cells from the higher dimensional data (where each dimension is either a different gene or a different principal component) and then after some resolution, the k-means clustering tries to minimize error. In this way, the k-means clustering provides cluster assignments. In some embodiments, K is a number between 2 and 50 inclusive. In some embodiments, the number K is set to a predetermined number such as 10. In some embodiments, the number K is optimized. In some embodiments, the number T is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more than 30. In some embodiments, the number K is at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100.

[0120] The clusters 116 represent a higher-level cell ontology similar to a cell type or state. An example of such clustering is provided in panel 502 of Figure 5 A. This clustering allows for the identification of batches on a per-cell-type basis, while maintaining sensitivityto local batch effects where only some cell types are affected by batch effects. In some embodiments a grid-search approach is used to find the optimal clustering resolution.

[0121] Referring to block 220, in some embodiments, as discussed above, the corresponding measurement value of each respective feature in the plurality of features originating from each respective cell in the plurality of cells is dimension reduced prior to clustering. In other words, the higher dimensional matrix (the measurement values of features in each of the cells) is projected into a lower dimensional orthogonal embedding of cells depending on the appropriate data type. In some embodiments, for scRNA-seq, PCA (principal component analysis) is performed on the normalized expression matrix. In some embodiments for scATAC-seq, LSI (latent semantic indexing) is performed on the tile matrix using the ArchR function “addlterativeLSI” with settings iterations = 2 and varFeatures = 20000. LSI is described in Deerwester et al., 1990, “Indexing by Latent Semantic Analysis.” Journal of the American Society for Information Science 41(6), pp. 391-407, which is hereby incorporated by reference.

[0122] Referring to block 222, in some embodiments, the plurality of clusters comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 clusters.

[0123] Batch inference.

[0124] In some embodiments, the disclosed batch inference framework makes the following assumptions. Define the input normalized matrix of measurement values for the plurality of features for the plurality of cells to be M 6 IRGxWswhere elementthe normalized of thehfeature in the nthcell of the sthsample. Within each sample, the separation among the cell types is preserved. Any sample can belong to any batch and a sample can be its own unique batch. The cell types contained within each sample do not have batch variations relative to each other. Batch effects do not intermix cells of distinct types or states from different samples. In other words, the measurement profiles of cells of cell type A in sample x are more correlated to cells of cell type A sample y than to cells of cell type B or C in sample y. Samples do not have to contain identical cell types.

[0125] Referring to block 223, in some embodiments, for each respective cell in the plurality of cells (i) a cluster assignment 116 of the respective cell and (ii) an identity of a sample 110 / 120 in the plurality of samples from which the respective cell originated is used to compute a plurality of sample-level local inverse Simpson index (LISI) scores. Theplurality of sample-level LISI scores are penalized by a measure of importance of cells in each sample in the plurality of samples.

[0126] In further detail, in some embodiments the batch inference framework iteratively finds differentially distributed groups of samples within clusters that contain at least a certain percentage (e.g., 1%, 2%, 3%, 4%, etc.) of all cells in the plurality of cells using Hill numbers Dq, an information-theoretic metric traditionally used to measure biological diversity. See Hill, 1973, “Diversity and evenness: a unifying notation and its consequences,” Ecology 54, pp. 427-432. Hill numbers have the form:When = 2, this formula becomes the Inverse Simpson Index (ISI) . Korsunsky et al. implemented a variant to ISI by adding a distance-based weight to the original index, namely Local Inverse Simpson’s Index (LISI). See Korsunsky et al., 2019, “Fast, sensitive and accurate integration of single-cell data with Harmony,” Nature methods 16(12), pp. 1289- 1296, which is hereby incorporated by reference.

[0127] Accordingly, in some embodiments, for each cluster 116, an LISI score that quantifies sample distributions is calculated. The unsealed LISI score measures the effective number of batches given observations and ranges from 0 to total number of samples 120 within the target cluster. However, given that the LISI metric is influenced by cell population size diversity, the original implementations of LISI are inappropriate for automated batch identification. Therefore, a penalized calculation of sample-level LISI (sLISI) scores is introduced using the frequency of occurrence of cells in each sample as weights. This penalized calculation is described in further detail in block 224.

[0128] Referring to block 224, in some embodiments, the using, for each respective cell in the plurality of cells (i) a cluster assignment 116 of the respective cell and (ii) an identity of a sample 120 in the plurality of samples from which the respective cell originated to compute the plurality of sample-level local inverse Simpson index (LISI) scores, where the plurality of sample-level LISI scores are penalized by a measure of importance of cells in each sample in the plurality of samples, is determined for each sample i by computing:1 sample-level local inverse Simpson index / = - -is a weight assigned to cluster j in sample i based on a measure of importance of cluster j to sample z, ptj is a proportion of cells in cluster j that originate from sample z, and N is the number of clusters in the plurality of clusters. Referring to block 226, in some embodiments, the measure of importance of cluster j to sample z is a frequency of occurrence of cells in cluster j from sample z. Referring to block 228, in some embodiments, the measure of importance of cluster j to sample z is a number of cells in cluster j from sample z.

[0129] Referring to block 229, in some embodiments, the plurality of sample-level LISI scores is scaled to range from zero and one, thereby forming a plurality of scaled samplelevel LISI scores 122. A score closer to 1 indicates sample well-mixing within the given cluster, while a score closer to 0 means the sample is locally enriched. As a result, samples with significantly small-scaled sample-level LISI scores 122 form putative batches.

[0130] Referring to block 230, the plurality of scaled sample-level LISI scores 122 is numerically ordered, thereby forming a numerically ordered plurality of scaled sample-level LISI scores.

[0131] Referring to block 232, a Dixon’s Q-test is applied to the numerically ordered plurality of scaled sample-level LISI scores thereby grouping the numerically ordered plurality of scaled sample-level LISI scores into one or more groups. Accordingly, given an ordered scaled sample-level LISI scores:where ls. is the sample-level LISI score 122 for sample source (120), the difference between each score is calculated and its subsequent neighbor, {tA}, where di — lSi+1~ IsThe goal is to find the outlier in DN with Dixon’s Q-test, which represents the most statistically significant gap between two consecutive scores in LN such thatsamples {si, . . ., sq} are considered to be one group, and samples {sq+i, ..., sn} form another group. In some embodiments statistical significance is considered to be / ? < 0.05, where p is a p-value. In some embodiments statistical significance is considered to be / ? < 0.01, or / ? < 0.001 or some other / ?-value threshold. In some embodiments, batch labels are iteratively assigned with an outlier detection method to ensure that all significant differences between each score are accounted for. Figure 5 illustrates how the disclosed algorithm iteratesthrough the clusters to determine if a group of samples is significantly different from the rest of the samples and assigns batch labels locally. For instance, as illustrated in panel 504, for cluster 1 of panel 502, sample 1 is determined to be significantly different from samples 2, 3, 4, and 5. As illustrated in panel 506, for cluster 2 of panel 502, samples S4 and S5 are determined to be significantly different than samples SI, S2, and S3. As illustrated in panel 508, for cluster 3 of panel 502, no additional local batch effects are found. As illustrated in panel 508, batch assignments consistent with the local batch effects found in the clusters dictates that sample SI belongs to a first batch, samples S4 and S6 belong to a second batch, and sample S2 and S3 belong to a third batch.

[0132] Referring to block 234, a batch assignment for each sample in the plurality of samples is determined based on an assignment of corresponding scaled sample-level LISI scores in the plurality of scaled sample-level LISI scores common to the one or more groups.

[0133] Referring to block 236, in some embodiments, the one or more groups comprises 3 or more groups. Each group 126 in the three or more groups is given a batch assignment. Each batch assignment is assigned at least one sample in the plurality of samples.

[0134] Referring to block 236, the batch assignment of each sample in the plurality of samples to batch is used correct a corresponding measurement value of each respective features in the plurality of features for each respective cell in the plurality of cells.

[0135] In some embodiments, the disclosed framework returns a finalized batch assignment for all samples. In some embodiments, the inferred batch covariate is used to subset the merged multi-sample dataset 106 into a list of batch datasets. In some embodiments, in the case of scRNA-seq data, the Seurat RPCA method is run by calling the “FindlntegrationAnchors” function followed by the “IntegrateData” function. See, Hao et al., 2021, “Integrated analysis of multimodal single-cell data,” Cell 184(13), pp. 3573-3587, which is hereby incorporated by reference. In some embodiments, for scATAC-seq data, if batch effects are determined to be present, the ArchR “addHarmony” function is called to integrate the lower embedding of each sample with the inferred batches. See, Granja et al., 2021, “ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis,” Nature genetics 53(3), pp. 403-411, which is hereby incorporated by reference.

[0136] In some embodiments, after samples are integrated into a batch-corrected matrix, cells are annotated with a reference-based approach. In some embodiments, for scRNA-seq,the default version of the disclosed pipeline implements the Seurat label transfer algorithm that projects an identity to each query cell using some prior annotated reference data by calling the Seurat “FindTransf er Anchors” function. Since ArchR also adopts Seurat label transfer, the same algorithm also applies to scATAC-seq data. Clustering is then performed with the ArchR “addClusters” function, and cell types are projected from a reference gene expression dataset onto the query chromatin accessibility dataset using the ArchR function “addGenelntegrationMatrix” .

[0137] The disclosed framework can be easily incorporated into other single-cell annotation tools such as SingleR. See Aran etal., 2019, “Reference-based analysis of lung single-cell sequencing reveals a transitional profibrotic macrophage,” Nature immunology vol. 20(2), pp. 163-172, which is hereby incorporated by reference.

[0138] On top of per-cell label transferring, in some embodiments the disclosed framework introduces a majority -voting regime that further refines cell labeling. The majority -voting regime begins with an over-clustering of the integrated dataset at a higher resolution, where each cluster represents a cell phenotype that serves as a surrogate to a parent cell type or state. It then loops through each cluster and finds the most represented predicted cell types and states among cells within the query cluster. Since the data is overclustered, this process eventually merges clusters representing the same cell identities together for a data-driven cell label annotation landscape.

[0139] Referring to block 238, in some embodiments, the plurality of samples includes a first subset of samples that have a first condition and a second subset of samples that are free of the first condition and the method further comprises using the batch corrected corresponding measurement value of each respective feature in the plurality of features in each respective cell in the plurality of cells to associate one or more features in the plurality of features that exhibits a statically significant differential measurement value (e.g., p-value less than 0.05) between the first subset and the second subset with the condition.

[0140] Referring to block 240, in some embodiments, the condition is a cancer, hematologic disorder, autoimmune disease, inflammatory disease, immunological disorder, metabolic disorder, neurological disorder, genetic disorder, psychiatric disorder, gastroenterological disorder, renal disorder, cardiovascular disorder, dermatological disorder, respiratory disorder, or a viral infection.

[0141] EXAMPLES

[0142] Example 1 - Validation of batch inference framework.

[0143] The disclosed batch inference framework was evaluated using four different real- world case studies that included different data types. First, in example 2, the performance of the batch inference framework applied to a public human lung scRNA-seq dataset was quantified using curated and published cell type annotation as the reference. Second, in example 3, batches were computationally inferred from a human PBMC scRNA-seq dataset using the disclosed framework and the cell type annotations were compared to those obtained from using metadata supplied batch information. Third, in example 4, the disclosed framework was shown to be robust across different data modalities by analyzing a mouse same-cell multiome dataset. Fourth, in example 7, the utility of the downstream analysis tools in the systems and method of the present disclosure were demonstrated by evaluating human PBMC scRNA-seq data from a longitudinal SARS-CoV-2 study.

[0144] Example 2 - Lung data. The human healthy lung single-cell expression data, originally collected by Braga et al, 2019, “A cellular census of human lungs identifies novel cell states in health and in asthma,” Nature medicine 25(7), pp. 1153-1163, includes three diverse datasets across different sequencing technologies and tissue locations. Two of the datasets were generated by 10X Chromium and the other one was generated by Drop-seq.

[0145] The Drop-seq dataset and one of the 10X datasets contain cells profiled from upper and lower airways, which were retrieved from deceased transplant donors. The other 10X dataset consists of parenchyma cells from resected lung tissue. These three datasets were further processed by Luecken et al. to include 16 samples composed of 32,472 cells and 15,148 genes, where a re-annotation of cells was performed since 10X and Drop-seq datasets used different annotation terms. Luecken etal., 2022, “Benchmarking atlas-level data integration in single-cell genomics,” Nature methods 19(1), pp. 41-50, which is hereby incorporated by reference.

[0146] The cell type annotation of the three datasets comprising 16 samples in total performed by Luecken et al 2022, “Benchmarking atlas-level data integration in single-cell genomics,” Nature methods 19(1) pp. 41-50, was used as an expert analysis standard for benchmarking. The raw count expression matrix was downloaded from Luecken et al. and was re-normalized using the batch inference framework in accordance with the present disclosure followed by automatic batch detection. Sample IDs were used as the supplied batch labels. The heterogeneity of samples resulted in the presence of strong batch effects inthe data that made integration particularly challenging. Other than intrinsic inter-individual variations, the different sequencing protocols and spatial locations resulted in highly variable cell type compositions between samples.

[0147] The heterogeneity of samples resulted in the presence of strong batch effects in the data that made integration particularly challenging. Other than intrinsic inter-individual variations, the different sequencing protocols and spatial locations resulted in highly variable cell type compositions between samples (Figures 19A and 19B). Such degree of complexity presented in this data compendium made itself an appropriate candidate for a benchmarking study.

[0148] The performance of the disclosed batch inference framework on this single-cell lung data compendium was evaluated. The disclosed pre-processing and normalization pipeline (blocks 212, 214, 216) was used to prepare the data for integration. The results obtained were evaluated with four common integration methods: Seurat CCA, Seurat RPC A, Harmony, and Scanorama. Using each of the four methods, the results using the standard approach of integration by sample IDs was compared with integration using the inferred batch labels determined in accordance with the present disclosure. To compare results, a panel of four metrics were developed to evaluate integration performance (preservation of cell type information and removal of dataset-specific variations) (Luecken et al., l^ , “Benchmarking atlas-level data integration in single-cell genomics,” Nature methods 19(1), pp. 41-50) leveraging the expert cell type annotation as further detailed in example 5. The systems and methods of the present disclosure exhibited a significant performance improvement in nearly all comparisons with the two Seurat and Harmony integrations. The improvement of the systems and methods of the present disclosure relative to Scanorama, which provides lower quality results for most metrics, was more modest. Overall, the benchmarking study indicates that the batch inference algorithm of the present disclosure significantly boosts integration performance.

[0149] Example 3 - Human PBMC scRNA-seq dataset

[0150] To further evaluate the reliability and robustness of the disclosed batch inference framework, it was applied to a 20 subject scRNA-seq dataset generated for a study of the responses in peripheral blood mononuclear cells (PBMC) to S. aureus infection. See, Chen et al., 2023, “Mapping disease regulatory circuits at cell-type resolution from single-cellmultiomics data,” Nature computational science doi: 10.1038 / s43588-023-00476-5, which is hereby incorporated by reference.

[0151] A total of 20 adult patients with culture-confirmed infection were selected with 10 MRS A patients and 10 MS SA patients from the scRNA-seq human S. aureus infection PBMC dataset. The sequencing batch labels were reported in an accompanying metadata file. The H5 files were extracted from 1 Ox Genomics Cell Ranger v3.1.0 and the raw count matrices were reprocessed using a batch inference framework in accordance with the present disclosure. For sample integration, two sets of batch labels were used: (a) the originally reported batch labels, and (b) the batch labels in inferred using an embodiment of the batch inference framework of the present disclosure (SPEEDI). Two different integrated datasets were subsequently obtained based on this input batch covariate. To assign cell types, a reference dataset was obtained by extracting data from 21 uninfected healthy individuals that were also reported and curated in Chen et al, Id. For illustration purposes, the analysis focused on a subset of cells that are T / NK leukocytes. The batch inference framework in accordance with an embodiment of the present disclosure identified eight major subtypes of the T and NK leukocytes in the integrated dataset following the batch inference framework in accordance with an embodiment of the present disclosure but only 7 subtypes were found in the reported-batch integrated dataset.

[0152] Before batch inference, a number of UMAP clusters were distinguished by sample, a pattern consistent with the presence of sever batch effects. See, Figures 7, 8A, and 8B. As can be seen in the T / NK lymphocyte subpopulations, the data-inferred batch labels determined in accordance with the systems and methods of the present disclosure led to better batch inference compared with batch labels obtained from metadata (Figs. 9, 10A, and 10B) for similar results obtained with all major cell types). The batch inference framework of the present disclosure corrected the fragmentation of the same cell type into different clusters as well as the presence of sample outliers that were seen with metadata-based batch labeling (Figs. 9 and 11). For example, UMAP of the expression of NKG7, a canonical marker gene for NK cells, shows that metadata-based batch inference failed to integrate the two NK batches together (FIG. 12). Metadata-based inference also failed to recover some rare cell populations that were identified using an embodiment of the batch inference framework of the present disclosure (SPEEDI), such as a small population of NK cells undergoing DNA replication (FIGs. 9, 10A, 10B, and 11).

[0153] To further compare the results following the batch inference framework of the present disclosure, an experimental gold standard for cell type proportion using CyTOF (Cytometry by Time of Flight) mass spectrometry was produced as summarized in Table 2.

[0154] Table 2 - Human PBMC study CyTOF sample metadata.

[0155] See, Bandura et al., 2009, “Mass cytometry: technique for real time single-cell multitarget immunoassay based on inductively coupled plasma time-of-flight mass spectrometry,” Analytical chemistry 81(16), pp. 6813-22, which is hereby incorporated by reference. The nonparametric Kendell’s correlation coefficient was computed for each of the matching cell types between the cell population composition of the dataset from the batch inference framework in accordance with an embodiment of the present disclosure and the CyTOF measurements, as well as reported-batch integrated and CyTOF datasets, respectively. The two lists of coefficients were used in the subsequent significance testing.

[0156] These new data were obtained using the same samples used to generate the previously reported S. aureus infection scRNA-seq data. See Chen et al., 2023, “Mapping disease regulatory circuits at cell-type resolution from single-cell multiomics data,” Nature computational science, doi: 10.1038 / s43588-023-00476-5, which is hereby incorporated by reference.

[0157] The correlation of cell type proportions in each sample obtained after integration were evaluated with the values from the CyTOF gold standard assay. Figure 20 provides the cell type proportions from CyTOF. The per-sample correlations of cell type proportions obtained from scRNA-seq using the batch inference framework in accordance with the present disclosure were higher than those obtained using the metadata batch label corrections (FIGs. 13 and 14). The integration after SPEEDI batch-labeling and metadata batch labeling in accordance with the present disclosure was also compared using silhouette score and cell type connectivity, two of the integration evaluation metrics that do not require reference cell type annotation data (Figs. 15 and 16). For both metrics, the batch-labeling in accordancewith the systems and methods of the present disclosure gave better performance, reflecting better preservation of biological variance among cell types and better overall batch effect reduction. These results demonstrate that, in comparison with metadata batch labels, the data-inferred batch labels assigned in accordance with the systems and methods of the present disclosure improve the accuracy of dataset integration and cell type annotation.

[0158] Further data showing the performance improvement of the disclosed batch inference systems and methods using this dataset is found in Wang et al., 2024, Automated single-cell omics end-to-end framework with data-driven batch inference,” bioRxiv 2023.11.01.564815, which is hereby incorporated by reference. See, in particular, the section entitled “Application of SPEEDI to heterogeneous single-cell RNA-seq data” in the Results section, in conjunction with Figures 3 A (Figure 22A of the present disclosure), 3B (Figure 22B of the present disclosure), S3 (Figure 22E), S4 (Figure 22F), S5 (Figure 22G), 3C (Figure 22C of the present disclosure), and 3D (Figure 22D of the present disclosure) of the reference.

[0159] Example 4 - Application of SPEEDI to single-cell multiome data

[0160] The data-inferred batch labels assignment pipeline in accordance with the systems and methods of the present disclosure can also be used for scATAC-seq and multiome datasets. For processing scATAC-seq data, the disclosed batch assignment in this example first used a standard ArchR(Granja et al., 2021, “ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis,” Nature genetics 53(3), pp. 403-411, which is hereby incorporated by reference) workflow to process input data. Sample batch labels were subsequently inferred, followed by integration and cell type annotation. If sample-paired or true multiome scRNA-seq data were also available, the batch inference framework of the present disclosure can process both data types.

[0161] To demonstrate the application of the batch inference framework in accordance with the present disclosure to single-cell multiome datasets, same cell scATAC and scRNA true multiome data generated from 14 wild-type female murine pituitary tissue samples was processed (GSE244132).

[0162] The pituitaries used in this study were collected from wild-type untreated female C57BL / 6 mice aged 11 weeks. All murine sample collection was conducted at McGill University (Montreal, Quebec, Canada). All animal experiments were performed in accordance with institutional and federal guidelines and were approved by the McGillUniversity and Goodman Cancer Centre Facility Animal Care Committee DOW- A (Protocol 5204). All ethical regulations and institutional protocols were complied with. Once dissected, the pituitaries were individually collected, immediately snap-frozen, and stored at - 80C until the assay was started.

[0163] Nuclei isolation was performed as described in previous literature. See Frederique et al., 2021 , “Single nucleus multi-omics regulatory landscape of the murine pituitary,” Nature communications 12(1), p. 2677; and Mendelev el al., 2022, “Multi-omics profiling of single nuclei from frozen archived postmortem human pituitary tissue,” STAR protocols 3.2, p. 101446, each of which is hereby incorporated by reference. Briefly, each snap-frozen pituitary was thawed individually on ice. RNAse inhibitor (NEB MO314L) was added to the homogenization buffer (0.32 M sucrose, 1 mM EDTA, 10 mM Tris-HCl, pH 7.4, 5mM CaC12, 3mM Mg(Ac)2, 0.1% IGEPAL CA-630), 50% OptiPrep (Stock is 60% Media from Sigma; cat# D1556), 35% OptiPrep and 30% OptiPrep right before isolation. Each pituitary was homogenized in a dounce glass homogenizer (1ml, VWR cat# 71000-514), and the homogenate filtered through a 40 mm cell strainer. An equal volume of 50% OptiPrep was added, and the gradient centrifuged (SW41 rotor at 9200rpm; 4C; 25min). Nuclei were collected from the interphase, washed, resuspended in IX nuclei dilution buffer (10X Genomics), and counted (Nexcelom K2 counter).

[0164] True (sn) multiome was performed following the Chromium Single Cell Multiome ATAC and Gene Expression Reagent Kits VI User Guide (lOx Genomics, Pleasanton, CA). Nuclei were counted as described above, transposition was performed in 10 ml at 37C for 60 min targeting 10,000 nuclei, before loading of the Chromium Chip J (PN- 2000264) for GEM generation and barcoding. Following post-GEM cleanup, the library was pre-amplified by PCR, after which the sample was split into three parts: one part for generating the snRNA-seq library, one part for the snATAC-seq library, and the rest was kept at -20C. snATAC and snRNA libraries were indexed for multiplexing (Chromium i7 Sample Index N, Set A kit PN-3000262, and Chromium i7 Sample Index TT, Set A kit PN-3000431 respectively).

[0165] The libraries were quantified by Qubit 3 fluorometer (Invitrogen), and quality was assessed by Bioanalyzer (Agilent). The libraries were sequenced first in a MiSeq (Illumina) to assess the reads and balance the sequencing pools and then sequenced in a NovaSeq 6000 (Illumina) at the New York Genome Center (NYGC) following 10X Genomics recommendations.

[0166] The batch inference framework in accordance with the present disclosure identified a total of 60,906 cells meeting quality control thresholds. Because the batch inference framework in accordance with the present disclosure independently processes the scATAC-seq and scRNA-seq components of the multiome data, comparison of the cell type annotation of each cell obtained using each data type provides an assessment of the reliability of the batch inference framework of the present disclosure. As set forth in Table 1 below, the batch labeling was consistent across data types, indicating that the batch inference framework in accordance with the present disclosure captured the technical variation between samples regardless of the type of data.

[0167] Table 1.

[0168] The cell type annotations obtained using the two data modalities were highly consistent (Fig. 17). Annotation of individual cells by cell type for the RNA-seq and for the ATAC-seq data after using the batch inference framework in accordance with the present disclosure showed a median cell subtype identification overlap of 0.96 as well as a median adjusted Rand index (ARI) of 0.85 (Fig. 18). These results support the applicability of the batch inference framework in accordance with the present disclosure to single-cell ATAC-seq and to single-cell multiome datasets.

[0169] Example 5 - Evaluation Metrics

[0170] A panel of metrics were developed to evaluate integration performance.

[0171] Leucken el al. defined two categories of metrics to evaluate integration results. See, Leucken etal., 2022, “Benchmarking atlas-level data integration in single-cell genomics,” Nature methods 19(1), pp. 41-50, which is hereby incorporated by reference. The first category measures the conservation of biological variance. The second category reflects the removal of batch effects. In this example, a total of four appropriate metrics were selected as evaluation criteria. Metrics from the first category considered how well cells of different identities were separated for each sample. Metrics from the second category measured the degree of overlapping of cells from each sample for each cell type.

[0172] Per-sample ASW.

[0173] The average silhouette width (ASW) is a classical metric that evaluates the validity of a clustering partition based on between-cluster proximities. In single-cell data science, a clustering of cells can be represented by labels such as sample IDs, sequencing protocols, cell type identities, etc. In the original definition, the ASW ranges between -1 and 1 such that -1 means strong overlapping and 1 means perfect separation between clusters.

[0174] The per-sample ASW score for cell type separation was computed based on sample IDs and scaled between 0 and 1 using the following equation for each sample. A score closer to 1 represents well-conserved cell types.

[0175] Per-sample graph connectivity.

[0176] The graph connectivity metric quantifies the average connectivity of the subgraph of cells in target cell type c E C in comparison to the largest connected component (LCC) for all cells of all types in a kNN graph. A score of 1 indicates that all cells of the same type are connected. A score is computed for each sample using the following equation:

[0177] Per cell-type GraphConn.

[0178] The per cell-type graph connectivity metric is defined similarly to the per-sample GraphConn metric, except for that it now averages the connectivity of all sample s E S. A score of 1 indicates that all cells of the same type are strongly connected post integration. A score is computed for each cell type using the following equation:

[0179] Per cell-type graph LISI.

[0180] The local inverse Simpson’s index (LISI), Korsunsky et al., 2019, “Fast, sensitive and accurate integration of single-cell data with Harmony,” Nature methods 16(12), pp. 1289-1296, which is hereby incorporated by reference, was further adapted by Leucken et al. to use the graph structure of the integrated data to calculate shortest path length as distance between two nodes. Luecken et al.,“Benchmarking atlas-level data integration in single-cell genomics,” Nature methods 19(1), pp. 41-50, which is hereby incorporated by reference. Like the original LISI, the graph LISI ranges from 1 to the total number of batches. The metric is then scaled with the following equation to a score between 0 and 1 for each cell type such that 1 indicates perfect cell type preservation.

[0181] Example 6 Optional downstream analyses 1.

[0182] In order to run downstream analyses, in this example, a metadata table is provided where row names are sample names and column names are metadata attributes of interest. Each metadata attribute contains exactly two unique labels (e.g., “disease” and “control”). The batch inference framework in accordance with an embodiment of the present disclosure has two different downstream analysis options. First, a cell-type specific differential expression analyses can be run. For each cell type, the pipeline in accordance with an embodiment of the present disclosure used the Wilcoxon Rank-Sum test (via Seurat’s FindMarkers method) to perform differential expression using the sample labels associated with a given metadata attribute. Results were filtered using an adjusted p-value threshold of 0.05, a log fold change threshold of 0.1, and a min. pct threshold of 0.1 (genes must be expressed in at least 10% of cells in one of the two groups associated with the metadata attribute). Next, pseudobulk analysis is performed using DESeq221. First, pseudobulk counts were calculated for each cell type and DESeq2 was used to find differentially expressed genes (via Wald test). Results were then filtered using a p-value of 0.05. Genes that pass both single-cell and pseudobulk differential analysis constituted the final list of differentially expressed genes. Differential expression results for each metadata attribute were written to a TSV file.

[0183] The disclosed pipeline in accordance with an embodiment of the present disclosure can also perform functional module discovery (FMD) to find functionally connected gene clusters using Human Base (Krishnan et al., 2016, “Genome-wide prediction and functional characterization of the genetic basis of autism spectrum disorder,” Nature neuroscience vol. 19(11), pp. 1454-1462, which is hereby incorporated by reference). By default, the pipeline in accordance with an embodiment of the present disclosure uses the lists of genes generated by the cell-type specific differential expression analysis described above. For each metadata attribute and cell type, the list of genes was first divided into positive fold change and negative fold change subsets. If at least 20 genes remain after fold change filtering, the pipeline in accordance with an embodiment of the present disclosure runs FMD. Results were written to a CSV file and included a URL that allows users to see their full results as well as a table that contains gene ontology (GO) enrichment results.

[0184] Example 7 - Optional downstream analyses 2.

[0185] To demonstrate the utility of the downstream analysis tools in the inference systems and methods of the present disclosure, human PBMC scRNA-seq data from a longitudinal SARS-CoV-2 study (Giroux et al., 2022, “Differential chromatin accessibility in peripheral blood mononuclear cells underlies COVID-19 disease severity prior to seroconversion,” Scientific reports 12.1 (2022), p. 11714, which is hereby incorporated by reference) was evaluated. The inference systems and methods of the present disclosure were used to integrate data from four severe seropositive COVID subjects and 5 healthy controls (9 samples in total, Figure 21 A) and then performed preliminary downstream analyses, including differential expression analysis and functional module discovery in the major immune cell types. The authors used bulk transcriptome data to note downregulation of chemokine CCL3 and cytokine IL1B and upregulation of KLF6 in COVID subjects. After processing the scRNA-seq data through the inference systems and methods of the present disclosure, the differential expression analysis option provided with the inference systems and methods of the present disclosure were applied to compare severe COVID subjects and healthy controls within each cell type. This analysis confirmed the regulatory changes reported in the published bulk RNA-seq analysis and also showed that CCL3 and IL1B were downregulated in CD 14 monocytes and that KLF6 was upregulated in CD4 TCM, CD8 Naive, CD8 TEM, B naive, MAIT, and pDC cells (Table S4 of Wang et al., 2024, Automated single-cell omics end-to-end framework with data-driven batch inference,” bioRxiv 2023.11.01.564815, which is hereby incorporated by reference). Furthermore, use of thefunctional module discovery tool (see Example 6) on upregulated and downregulated genes from different cell types revealed additional insights into the specific biological processes associated with severe SARS-CoV-2 infection. For example, upregulated genes in NK cells were found to be associated with positive regulation of the canonical Wnt signaling pathway (Vallee etal., 2021, “Interplay of opposing effects of the WNT / p-Catenin pathway and PPARy and implications for SARS-CoV2 treatment,” Frontiers in immunology 12, p. 666693, which is hereby incorporated by reference), positive regulation of lymphocyte activation (Malengier-Devlies etal., 2022, “Severe COVID-19 patients display hyperactivated NK cells and NK cell-platelet aggregates,” Frontiers in Immunology 13, p. 861251, which is hereby incorporated by reference), and the ERAD pathway (Movaqar et al., 2021, “Coronaviruses construct an interconnection way with ERAD and autophagy,” Future Microbiology 16.14, pp. 1135-1151, which is hereby incorporated by reference) (Figure 2 IB). Thus inference systems and methods of the present disclosure can provide useful automated analysis for single-cell datasets and, even when applied to data that has been previously analyzed, has the potential to facilitate addressing specific questions of interest to the user (in this case cell type-level transcriptome and pathway effects) not previously answered in a published analysis.

[0186] CONCLUSION

[0187] The terminology used herein is for the purpose of describing particular cases and is not intended to be limiting. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, to the extent that the terms “including,” “includes,” “having,” “has,” “with,” or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0188] Plural instances may be provided for components, operations or structures described herein as a single instance. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated inthe context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within the scope of the implementation(s). In general, structures and functionality presented as separate components in the example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the implementation(s).

[0189] It will also be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used to distinguish one element from another. For example, a first subject could be termed a second subject, and, similarly, a second subject could be termed a first subject, without departing from the scope of the present disclosure. The first subject and the second subject are both subjects, but they are not the same subject.

[0190] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting (the stated condition or event )” or “in response to detecting (the stated condition or event),” depending on the context.

[0191] The foregoing description included example systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative implementations. For purposes of explanation, numerous specific details were set forth in order to provide an understanding of various implementations of the inventive subject matter. It will be evident, however, to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures and techniques have not been shown in detail.

[0192] The foregoing description, for purpose of explanation, has been described with reference to specific implementations. However, the illustrative discussions above are not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many alterations, modifications, and variations will be apparent to those skilled in the art in light of the foregoing description without departing from the spirit or scope of the present disclosure and that when numerical lower limits and numerical upper limits are listed herein,ranges from any lower limit to any upper limit are contemplated. The implementations were chosen and described in order to best explain the principles and their practical applications, to thereby enable others skilled in the art to best utilize the implementations and various implementations with various modifications as are suited to the particular use contemplated.

Claims

What is claimed:

1. A method for inferring one or more batches in single-cell omics data comprising: at a computer system comprising one or more processors and a non-transitory computer-readable medium including computer-executable instructions:A) obtaining, for each respective cell in a plurality of cells, a corresponding measurement value of each respective feature in a plurality of features originating from the respective cell, wherein the plurality of features comprises 30 or more features, and wherein the plurality of cells collectively originates from a plurality of samples;B) using the measurement value of each respective feature in the plurality of features of each respective cell in the plurality of cells to cluster the plurality of cells into a plurality of clusters;C) using for each respective cell in the plurality of cells, (i) a cluster assignment of the respective cell and (ii) an identity of a sample in the plurality of samples from which the respective cell originated, to compute a plurality of sample-level local inverse Simpson index (LISI) scores, wherein the plurality of sample-level LISI scores are penalized by a measure of importance of cells in each sample in the plurality of samples;D) scaling the plurality of sample-level LISI scores to range from zero and one, thereby forming a plurality of scaled sample-level LISI scores;E) numerically ordering the plurality of scaled sample-level LISI scores, thereby forming a numerically ordered plurality of scaled sample-level LISI scores;F) subtracting each scaled sample-level LISI score in the numerically ordered plurality of scaled sample-level LISI scores from its consecutive score, thereby forming an ordered plurality of differences of scaled sample-level LISI scores;G) applying a Dixon’s Q-test to the ordered plurality of differences of scaled samplelevel LISI scores thereby grouping the numerically ordered plurality of scaled sample-level LISI scores into one or more groups; andH) determining a batch assignment for each sample in the plurality of samples based on an assignment of corresponding scaled sample-level LISI scores in the plurality of scaled sample-level LISI scores common to the one or more groups.

2. The method of claim 1, wherein the using, for each respective cell in the plurality of cells, (i) a cluster assignment of the respective cell and (ii) an identity of a sample in the plurality of samples from which the respective cell originated, to compute the plurality of sample-levellocal inverse Simpson index (LISI) scores, wherein the plurality of sample-level LISI scores are penalized by a measure of importance of cells in each sample in the plurality of samples, is determined for each sample z by computing: sample-level local inverse Simpson index / =wherein,Wij is a weight assigned to cluster j in sample i based on a measure of importance of cluster j to sample z,Ptj is a proportion of cells in cluster j that originate from sample z, andN is the number of clusters in the plurality of clusters.

3. The method of claim 2, wherein the measure of importance of cluster j to sample z is a frequency of occurrence of cells in cluster j from sample z.

4. The method of claim 2, wherein the measure of importance of cluster j to sample z is a number of cells in cluster j from sample z.

5. The method of any one of claims 1-4, the method further comprising determining the corresponding measurement value of each respective feature in the plurality of features originating from the respective cell using an scRNA, scRNA-seq or scATAC-seq based sequencing reaction that produces a plurality of sequence reads and that maps the plurality of sequence reads to a reference genome.

6. The method of claim 5, wherein the plurality of sequence reads is at least 1 x 106sequence reads.

7. The method of claim 5 or 6, wherein each sequence read in the plurality of sequence reads is associated with a unique molecule identifier (UMI) and the method further comprises: using the plurality of sequence reads to construct a first histogram of a UMI count covariate; finding a left elbow point from a curve formed by the first histogram;removing from the plurality of cells each cell that is in a portion of the first histogram to the left of the left elbow point prior to performing B).

8. The method of claim 7, the method further comprising removing from the plurality of cells each respective cell that has a number of unique UMI that exceeds a threshold percentile for the plurality of cells.

9. The method of claim 8, wherein the threshold percentile is 95 percent.

10. The method of any one of claims 7-9, wherein the finding the left elbow point from the curve formed by the first histogram is done using a Kneedie algorithm.

11. The method of any one of claims 7-10, the method further comprising: using the plurality of sequence reads to construct a second histogram of a mitochondrial percentage; finding a right elbow point from a curve formed by the second histogram; removing from the plurality of cells each cell that is in a portion of the histogram to the right of the right elbow point prior to the using B).

12. The method of any one of claims 1-11, wherein the plurality of samples comprises 2, 3,4, 5, 6, 7, 8, 9, or 10 samples.

13. The method of any one of claims 1-12, wherein the plurality of clusters comprises 2, 3, 4,5, 6, 7, 8, 9, or 10 clusters.

14. The method of any one of claims 1-13, wherein the plurality of cells comprises 10 cells, 50 cells, 100 cells, 500 cells, or 1000 cells.

15. The method of any one of claims 1-14, wherein the one or more groups comprises 3 or more groups, each group in the three or more groups is given a batch assignment, and each batch assignment is assigned at least one sample in the plurality of samples.

16. The method of any one of claims 1-15, wherein the corresponding measurement value of each respective feature in the plurality of features originating from each respective cell in the plurality of cells is dimension reduced prior to the using B).

17. The method of any one of claims 1-16, the method further comprising:H) using the batch assignment of each sample in the plurality of samples to batch correct a corresponding measurement value of each respective feature in the plurality of features for each respective cell in the plurality of cells.

18. The method of claim 17, wherein the plurality of samples includes a first subset of samples that have a first condition and a second subset of samples that are free of the first condition and the method further comprises using the batch corrected corresponding measurement value of each respective feature in the plurality of feature in each respective cell in the plurality of cells to associate one or more features in the plurality of features that exhibits a statistically significant differential abundance between the first subset and the second subset with the condition.

19. The method of claim 18, wherein the condition is a cancer, hematologic disorder, autoimmune disease, inflammatory disease, immunological disorder, metabolic disorder, neurological disorder, genetic disorder, psychiatric disorder, gastroenterological disorder, renal disorder, cardiovascular disorder, dermatological disorder, respiratory disorder, or a viral infection.

20. The method of any one of claims 1-19, wherein each feature in the plurality of features is a different gene and each corresponding measurement value is gene abundance.

21. The method of any one of claims 1-19, wherein each feature in the plurality of features is a different chromatin peak and each corresponding measurement value is a size of a chromatin peak.

22. The method of any one of claims 1-19, wherein each feature in the plurality of features is a different epigenetic modification and each corresponding measurement value is an extent of epigenetic modification.

23. A computer system for inferring one or more batches in single-cell omics data, the computer system comprising: one or more processors; and memory addressable by the one or more processors, the memory storing at least one program for execution by the one or more processors, the at least one program comprising instructions for:A) obtaining, for each respective cell in a plurality of cells, a corresponding measurement value of each respective feature in a plurality of features originating from the respective cell, wherein the plurality of features comprises 30 or more features, and wherein the plurality of cells collectively originates from a plurality of samples;B) using the measurement value of each respective feature in the plurality of features of each respective cell in the plurality of cells to cluster the plurality of cells into a plurality of clusters;C) using for each respective cell in the plurality of cells, (i) a cluster assignment of the respective cell and (ii) an identity of a sample in the plurality of samples from which the respective cell originated, to compute a plurality of sample-level local inverse Simpson index (LISI) scores, wherein the plurality of sample-level LISI scores are penalized by a measure of importance of cells in each sample in the plurality of samples;D) scaling the plurality of sample-level LISI scores to range from zero and one, thereby forming a plurality of scaled sample-level LISI scores;E) numerically ordering the plurality of scaled sample-level LISI scores, thereby forming a numerically ordered plurality of scaled sample-level LISI scores;F) subtracting each scaled sample-level LISI score in the numerically ordered plurality of scaled sample-level LISI scores from its consecutive score, thereby forming an ordered plurality of differences of scaled sample-level LISI scores;G) applying a Dixon’s Q-test to the ordered plurality of differences of scaled samplelevel LISI scores thereby grouping the numerically ordered plurality of scaled sample-level LISI scores into one or more groups; andH) determining a batch assignment for each sample in the plurality of samples based on an assignment of corresponding scaled sample-level LISI scores in the plurality of scaled sample-level LISI scores common to the one or more groups.

24. The computer system of claim 23, wherein the at least one program further comprises instructions for:H) using the batch assignment of each sample in the plurality of samples to batch correct a corresponding measurement value of each respective feature in the plurality of features for each respective cell in the plurality of cells.

25. A non-transitory computer readable storage medium, wherein the non-transitory computer readable storage medium stores instructions, which when executed by a computer system, cause the computer system to perform a method for inferring one or more batches in single-cell omics data, the method comprising:A) obtaining, for each respective cell in a plurality of cells, a corresponding measurement value of each respective feature in a plurality of features originating from the respective cell, wherein the plurality of features comprises 30 or more features, and wherein the plurality of cells collectively originates from a plurality of samples;B) using the measurement value of each respective feature in the plurality of features of each respective cell in the plurality of cells to cluster the plurality of cells into a plurality of clusters;C) using for each respective cell in the plurality of cells, (i) a cluster assignment of the respective cell and (ii) an identity of a sample in the plurality of samples from which the respective cell originated, to compute a plurality of sample-level local inverse Simpson index (LISI) scores, wherein the plurality of sample-level LISI scores are penalized by a measure of importance of cells in each sample in the plurality of samples;D) scaling the plurality of sample-level LISI scores to range from zero and one, thereby forming a plurality of scaled sample-level LISI scores;E) numerically ordering the plurality of scaled sample-level LISI scores, thereby forming a numerically ordered plurality of scaled sample-level LISI scores;F) subtracting each scaled sample-level LISI score in the numerically ordered plurality of scaled sample-level LISI scores from its consecutive score, thereby forming an ordered plurality of differences of scaled sample-level LISI scores;G) applying a Dixon’s Q-test to the ordered plurality of differences of scaled samplelevel LISI scores thereby grouping the numerically ordered plurality of scaled sample-level LISI scores into one or more groups; andH) determining a batch assignment for each sample in the plurality of samples based on an assignment of corresponding scaled sample-level LISI scores in the plurality of scaled sample-level LISI scores common to the one or more groups.

26. The non-transitory computer readable storage medium of claim 25, wherein the at least one program further comprises instructions for:H) using the batch assignment of each sample in the plurality of samples to batch correct a corresponding measurement value of each respective feature in the plurality of features for each respective cell in the plurality of cells.

Citation Information

Cited By

  • Precise enclosure method for stem cell detection and analysis

    CN121837206A

  • Inplanatable decision-making system for single-cell multi-omics data integration analysis

    CN122117020A

  • Sequencing data processing method and apparatus, electronic device, and storage medium

    CN122531473A