Machine learning based imputation of sequencing data

A machine learning-based imputation model iteratively corrects spurious zero expression values in single-cell sequencing data, enhancing data quality by differentiating between true and spurious zeros, thus improving analysis.

WO2026039759A1PCT designated stage Publication Date: 2026-02-19SYNLICO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/042218
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2025-08-15
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Single-cell sequencing data is prone to high sparsity due to a large proportion of spurious zero expression values, which hinders downstream analysis.

Method used

A machine learning-based imputation model, such as a transformer model, is trained to differentiate between true and spurious zero expression values by iteratively correcting zero expression values in sequencing data, using a dropout mask to generate artificial zeros and adjusting model parameters to preserve true zeros.

Benefits of technology

The model effectively reduces spurious zero expression values, improving the quality of sequencing data for downstream analysis by accurately distinguishing and correcting non-biological zeros while preserving biological zeros.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025042218_19022026_PF_FP_ABST
    Figure US2025042218_19022026_PF_FP_ABST
Patent Text Reader

Abstract

A degraded sequencing data sample may be generated to include one or more artificial zero expression values by replacing one or more expression values in a sequencing data sample with the one or more artificial zero expression values. An imputation model may be trained based on the degraded sequencing data sample before being applied to correct excess zero expression values in one or more gene expression profiles. The training of the imputation model may include leveraging the output of the imputation model to infer true zero expression values in the sequencing data sample for supervising further training of the imputation model. The imputation model may be trained to replace spurious zero expression values with an imputed expression value that is as close as possible to the original expression values while preserving any true zero expression values and non-zero expression values.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Ref.: 104631-888002MACHINE LEARNING BASED IMPUTATION OF SEQUENCING DATA CROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Application No. 63 / 684,091, entitled “MACHINE LEARNING ENABLED IMPUTATION OF SINGLE-CELL NUCLEIC ACID SEQUENCING DATA” and filed on August 16, 2024, the disclosure of which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The subject matter described herein relates generally to sequencing and more specifically to machine learning based techniques for imputing zero expression values in sequencing data.INTRODUCTION

[0003] Sequencing refers to various methodologies for determining the sequence of nucleotide bases in one or more polynucleotides including, for example, nucleic acid molecules such as deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and variants or derivatives thereof (e.g., single stranded DNA). When performed in bulk, sequencing may determine the nucleic acid sequence of polynucleotides extracted from a population of cells but without any differentiation between the different cell types present therein. Contrastingly, single-cell sequencing leverages next-generation sequencing (NGS) technologies to examine nucleic acid sequence information from individual cells. The resulting single-cell transcriptome may provide high resolution insights into the characteristics and functions of individual cell in the context of its microenvironment. For example, single-cell sequencing can reveal valuable insights into disease heterogeneity. This heterogeneity may include the cellular heterogeneity present within patients’ microenvironments such as tumor microenvironments (TME), which play a vitalPreprint. Under review.regulatory role in tumorigenesis (or the pathological process in which normal cells are transformed into malignant cells). In some cases, the clinical outcome of treatments, such as immunotherapy, chemotherapy, and radiotherapy in oncology, may be contingent upon correct and precise characterization of the tumor microenvironment. As such, in some cases, single-cell sequencing and bioinformatics may enable a type of cancer to be further categorized into different tumor microenvironment (TME) associated molecular subtypes. Moreover, the selection of treatments, including the detection of therapeutic resistances, may be made based on patients’ microenvironment (TME) associated molecular subtypes.SUMMARY

[0004] Systems, methods, and articles of manufacture, including computer program products, are provided for machine learning based imputation of single-cell nucleic acid sequencing data. In one aspect, there is provided a system for machine learning based imputation of single-cell nucleic acid sequencing data. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one data processor. The operations may include: generating a degraded sequencing data sample to include one or more artificial zero expression values, where the degraded sequencing data sample is generated by replacing one or more expression values in a sequencing data sample with the one or more artificial zero expression values; training, based at least on the degraded sequencing data sample, an imputation model; and applying the imputation model to correct one or more expression values in one or more gene expression profiles.

[0005] In another aspect, there is provided a computer-implemented method for machine learning based imputation of single-cell nucleic acid sequencing data. The method may include:2131251356v lgenerating a degraded sequencing data sample to include one or more artificial zero expression values, where the degraded sequencing data sample is generated by replacing one or more expression values in a sequencing data sample with the one or more artificial zero expression values; training, based at least on the degraded sequencing data sample, an imputation model; and applying the imputation model to correct one or more expression values in one or more gene expression profiles.

[0006] In another aspect, there is provided a computer program product for machine learning based imputation of single-cell nucleic acid sequencing data. The computer program product may include a non-transitory computer readable medium storing instructions that result in operations when executed by at least one data processor. The operations may include: generating a degraded sequencing data sample to include one or more artificial zero expression values, where the degraded sequencing data sample is generated by replacing one or more expression values in a sequencing data sample with the one or more artificial zero expression values; training, based at least on the degraded sequencing data sample, an imputation model; and applying the imputation model to correct one or more expression values in one or more gene expression profiles.

[0007] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0008] In some variations, the degraded sequencing data sample is generated by applying a dropout mask to replace one or more non-zero expression values in the sequencing data sample with the one or more artificial zero expression values.

[0009] In some variations, the training of the imputation model includes applying the imputation model to denoise the degraded sequencing data sample, identifying, based at least on3131251356v lan intermediate sequencing data sample generated by the imputation model denoising the degraded sequencing data sample, one or more true zero expression values present in the sequencing data sample, and training, based at least on the intermediate sequencing data sample, the imputation model to replace any artificial zero expression values remaining in the intermediate sequencing data sample while preserving the one or more true zero expression values.

[0010] In some variations, the sequencing data sample includes true zero expression values that are indistinguishable from spurious zero expression values. The intermediate sequencing data sample generated by the imputation model enables a differentiation between the true zero expression values and the spurious zero expression values.

[0011] In some variations, one or more lowest expression values in the intermediate sequencing data sample are identified as true zero expression values.

[0012] In some variations, the one or more lowest expression values comprise a threshold quantity that corresponds to an expected proportion of true zero expression values present in a type of the sequencing data sample.

[0013] In some variations, the imputation model is further trained to avoid replacing a non-zero expression value in the intermediate sequencing data sample with a different non-zero expression value.

[0014] In some variations, the training of the imputation model includes adjusting one or more parameters of the imputation model to reduce a difference between a non-zero expression value imputed by the imputation model for each remaining artificial zero expression value and a corresponding original non-zero expression value present in the sequencing data sample.4131251356v l

[0015] In some variations, a first plurality of subsets of gene expression profiles is identified in a set of gene expression profiles for a batch of cells that have undergone sequencing at a sequencing platform. The imputation model is applied to generate a first set of corrected gene expression profiles by at least correcting one or more expression values in each subset of gene expression profile from the first plurality of subsets of gene expression profiles

[0016] In some variations, a second plurality of subsets of gene expression profiles is identified in the set of gene expression profiles for the batch of cells. The imputation model is applied to generate a second set of corrected gene expression profiles by at least correcting the one or more expression values each subset of gene expression profile from the second plurality of subsets of gene expression profiles. Upon satisfying one or more criteria, a result of a current iteration of imputation is determined based at least on the first set and the second set of corrected gene expression profiles.

[0017] In some variations, one or more additional iterations of imputation are performed to correct one or more additional expression values present in the result of the current iteration of imputation.

[0018] In some variations, the imputation model corrects the one or more expression values in the one or more gene expression profiles by at least replacing, at a first timestep, at least a first expression value in the one or more gene expression profiles with a second expression value, and replacing, at a second timestep following the first timestep, at least the second expression value in the one or more gene expression profiles with a third expression value.5131251356v l

[0019] In some variations, the first expression value is a zero expression value, the second expression value is a non-zero expression value, and the third expression value is a different non-zero expression value.

[0020] In some variations, the third expression value is closer to a true expression value than the second expression value.

[0021] In some variations, the imputation model corrects the one or more expression values in the one or more gene expression profiles by at least encoding a gene expression profile to generate a token that represents the gene expression profile with a fewer quantity of variables than the gene expression profile in its original form, modifying the token to correct the one or more expression values in the gene expression profile, decoding an imputation token generated by the modifying of the token, the decoding generating an imputation vector including one or more non-zero expression values for replacing one or more corresponding zero expression values in the gene expression profile, and generating a corrected gene expression profile by at least applying the imputation vector to the gene expression profile.

[0022] In some variations, the imputation model comprises a transformer model, a first linear projection layer coupled to an input of the transformer model, and a second linear projection layer coupled to an output of the transformer model.

[0023] In some variations, the one or more gene expression profiles comprise single-cell sequencing data.

[0024] In some variations, the one or more gene expression profiles include true zero expression values and spurious zero expression values. The imputation model is applied to correct the spurious zero expression values while preserving the true zero expression values.6131251356v l

[0025] In some variations, the true zero expression values correspond to biological zero expression values and the spurious zero expression values correspond to non-biological zero expression values.

[0026] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non- transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.

[0027] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for7131251356v lillustrative purposes in relation to the imputation of single-cell nucleic acid sequencing data, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF DRAWINGS

[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,

[0029] FIG. 1A depicts a system diagram illustrating an example of a sequencing data analysis system, in accordance with some example embodiments;

[0030] FIG. IB depicts a schematic diagram illustrating an example of an imputation model, in accordance with some example embodiments;

[0031] FIG. 2A depicts a flowchart illustrating an example of a process for machine learning based imputation of sequencing data, in accordance with some example embodiments;

[0032] FIG. 2B depicts a flowchart illustrating another example of a process for machine learning based imputation of sequencing data, in accordance with some example embodiments;

[0033] FIG. 3A depicts a graph illustrating the discrepancy between training trajectory and inference trajectory of an imputation model, in accordance with some example embodiments;

[0034] FIG. 3B depicts a graph illustrating an example of a training trajectory of an imputation model, in accordance with some example embodiments;

[0035] FIG. 4A depicts a flowchart illustrating another example of a process for machine learning based imputation of sequencing data, in accordance with some example embodiments;8131251356v l

[0036] FIG. 4B depicts a flowchart illustrating another example of a process for machine learning based imputation of sequencing data, in accordance with some example embodiments;

[0037] FIG. 5 depicts graphs illustrating a comparison of the performance of different techniques for imputation of single cell nucleic acid sequencing data, in accordance with some example embodiments; and

[0038] FIG. 6 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.

[0039] When practical, similar reference numbers denote similar structures, features, or elements.DETAILED DESCRIPTION

[0040] Next-generation sequencing (NGS) is a form of massively parallel sequencing technique that provides ultra-high throughput, scalability, and speed by sequencing millions of nucleotide sequences at a time. Single-cell sequencing techniques, which leverages nextgeneration sequencing (NGS) technologies, are able to measure gene expression at the scale of individual cells, providing insights into the diversity of cellular characteristics and functions at an unprecedented resolution. Examples of single-cell sequencing techniques include single cell ribonucleic acid (RNA) sequencing, single cell transposase-accessible chromatin (ATAC) sequencing, single cell deoxyribonucleic acid (DNA) sequencing, single cell epigenetic sequencing (e.g., scCHIP-seq), single cell proteomics (e.g., mass cytometry (CyTOF)), single cell mass spectrometry, single cell spatial omics (e.g., spatial transcriptomics, spatial proteomics), single cell multi-omics, single cell imaging (e.g., MERFISH), single cell metabolomics, single cell functional assay, and / or the like.9131251356v l

[0041] The resulting gene expression profile of a cell, which includes values corresponding to the measurements of gene expression levels, may include zero expression values for unexpressed genes or for genes whose expression levels are not measured. While some of the zero expression values present in single cell sequencing data may be true biological zero expression values, others are spurious zero expression values. For example, true zero expression values in this context may include biological zero expression values attributable to genes that were not expressed at the time of sequencing (e.g., no mRNA for transcription at all or no mRNA due to faster mRNA degradation than transcription). Contrastingly, spurious zero expression values may include non-biological zero expression values, which can arise due to a number of reasons. Technical zeros are one type of biological zero expression values attributable to a gene whose mRNAs are not captured by complementary DNA (or cDNA) synthesis. Another type of biological zero expression values are sampling zeros, which may occur when a gene has too few copies of cDNAs to be captured by sequencing even though these cDNAs are in the sequencing library. A gene may have too few copies of cDNAs if the cDNAs are not sufficiently amplified (e.g., by polymerase chain reaction (PCR), in vitro transcription (IVT), and / or the like) or if the gene has too few copies of mRNA for amplification to yield a sufficient quantity of cDNAs.

[0042] The gene expression profiles generated by single-cell sequencing tend to be especially sparse. That is, the gene expression profiles generated by single-cell sequencing contain more zero expression values than those generated by bulk sequencing. For example, single-cell sequencing data may contain between 90-95% zero expression values while the proportion of zero expression values in bulk sequencing data performed on populations of cells is significantly lower at between 10-40%. The significantly higher proportion of zero values in10131251356v lsingle-cell sequencing data than in bulk sequencing data suggests that at least some of the zero expression values present in single-cell sequencing data are spurious, non-biological zero expression values. The sparsity of single-cell sequencing data, particularly the presence of a large proportion of spurious, non-biological zero expression values, may hinder downstream analysis. As described in more details below, the present disclosure provides a variety of machine-learning enabled imputation techniques to correct spurious, non-biological zeroexpression values.

[0043] In some example embodiments, an imputation model may be applied to correct one or more of the zero expression values present in sequencing data, including single-cell sequencing data such as the gene expression profiles of one or more individual cells. Although various examples of the imputation model are described as being applied to correct excess zero expression values in single-cell sequencing data, it should be appreciated that the same or similar imputation model may also be applied to correct zero expression values in other types of sequencing data, such as bulk sequencing data, without departing from the scope of the present disclosure. For example, in some cases, the imputation model may be trained to replace the one or more zero expression values present in the sequencing data (e.g., single-cell sequencing data and / or the like) with one or more non-zero expression values. In some cases, the sequencing data may be considered noisy or corrupted data containing noise (or corruption) in the form of non-biological zero expression values. As such, in some cases, the imputation model may be a machine learning model that has been trained to restore the sequencing data to its uncorrupted form by correcting one or more non-biological zero expression values. In some cases, the sequencing data may be restored or denoised iteratively over multiple, successive timesteps. For instance, at a first timestep, the imputation model may correct one or more expression values in11131251356v lthe sequencing data by replacing at least a first expression value in the sequencing data with a second expression value. At a second timestep, the imputation model may further correct one or more expression values in the intermediate sequencing data generated at the first timestep by replacing at least a third expression value in the intermediate sequencing data with a fourth expression value. It should be appreciated that the imputation model may correct the one or more expression values by replacing one or more zero expression values with one or more nonzero expression values. However, it is also possible for the imputation model to correct the one or more expression values by replacing one or more non-zero expression values with one or more different non-zero expression values.

[0044] In some example embodiments, the training data for training the imputation model may include sample sequencing data (e.g., sample single-cell sequencing data and / or the like) with one or more “masked” expression values, each of which being an expression value that has been replaced with an “artificial” zero expression value. For example, in some cases, a training sample may be generated by applying, to a sequencing data sample, a dropout mask configured to set one or more expression values in the sequencing data sample to an artificial zero expression value. It should be appreciated that both zero expression values and non-zero expression values originally found in the sequencing data sample may be set to artificial expression values by applying the dropout mask. The imputation model may be trained to replace an artificial zero expression value in the training sample with an imputed expression value that is as close as possible in value to the original expression value in the sequencing data sample. For instance, for an artificial zero expression value that masks a non-zero expression value in the sequencing data sample, the imputation model may be trained to replace the artificial zero expression value with the same (or nearly the same) non-zero expression value.12131251356v lContrastingly, the imputation model may be trained to avoid replacing true zero expression values in the sequencing data sample with non-zero expression values. As described in more details below, the output of the imputation model may be leveraged to identify the true zero expression values present in the sequencing data sample.

[0045] In some example embodiments, the training of the imputation model may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the imputation model to reduce the difference between the original expression values in each sample sequencing data and the expression values imputed by the imputation model. Although only non-biological zero expression values should be replaced while biological zero expression values should be preserved, there is in fact no way to differentiate between biological zero expression values and non-biological zero expression values in a sequencing data sample. Consequently, a sequencing data sample containing artificial zero expression values cannot be used to supervise the imputation model to preserve biological zero expression values and replace non-biological zero expression values. However, if this limitation is left unchecked during the training of the imputation model, the trained imputation model resulting therefrom may be incapable of recovering true (or clean) sequencing data at inference time. For example, in some cases, an improperly trained imputation model may exhibit a tendency to overcorrect and replace both biological zero expression values and non-biological zero expression values at inference time even though only the non-biological zero expression values constitute noise (or corruption). Accordingly, in some cases, the ability of the trained imputation model to recover true (or clean) sequencing data at inference time may be improved by leveraging the output of the imputation model to differentiate between biological zero expression values and non-biological zero expression values in for supervising the training of the imputation model. For instance, during13131251356v ltraining, the output of the imputation model may be used to identify one or more true zero expression values in a sequencing data sample and the imputation model may be further trained to avoid replacing these true zero expression values. In some cases, overcorrection may also be reduced by training the imputation model to avoid replacing a non-zero expression value in the sequencing data sample with a different non-zero expression value.

[0046] FIG. 1A depicts a system diagram illustrating an example of a sequencing data analysis system 100, in accordance with some example embodiments. Referring to FIG. 1 A the sequencing data analysis system 100 may include a sequencing controller 110, a sequencing platform 120, an analysis engine 130, and a client device 140. As shown in FIG. 1, the sequencing controller 110, the sequencing platform 120, the analysis engine 130, and the client device 140 may be communicatively coupled via a network 150. The client device 140 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable apparatus, and / or the like. The network 150 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like.

[0047] In some example embodiments, the sequencing platform 120 may be configured to determine the sequence of nucleotide bases in one or more polynucleotides (e.g., nucleic acid molecules such as deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and variants or derivatives thereof (e g., single stranded DNA)). In some cases, the sequencing data 120 may be a single-cell sequencing platform such that the sequencing data 125 that is generated by the sequencing platform 120 may be single-cell sequencing data, such as the gene expression profile of one or more individual cells. In some cases, in its raw form, the sequencing data 12514131251356v lgenerated by the sequencing platform 120 may contain corruption or noise. As the sequencing data 125 may include values corresponding to the measurements of gene expression levels, corruption or noise in this context may include incorrect measurements of gene expression levels. For example, in some cases, the corruption or noise that is present in the sequencing data 125 may include zero expression values that should be non-zero expression values instead. However, not all zero-expression values are considered noise or corruption. In the context of sequencing data, zero expression values may include biological zero expression values and non- biological zero expression values such as sampling zeros, technical zeros, and / or the like. Biological zero expression values are not considered corruption or noise because biological zero expression values are the true expression values of genes that were not expressed at the time of sequencing. Contrastingly, non-biological zero expression values are considered corruption or noise because non-biological zero expression values are the spurious expression values of genes that were expressed at the time of sequencing but are not measured due to technical and / or sampling errors.

[0048] It should be appreciated that single-cell sequencing data may be especially prone to containing an excess quantity of zero expression values. Sequencing at the resolution of individual cells requires amplification of minute quantities of nucleic acid molecules (e.g., messenger RNA (mRNA) in the case of single-cell RNA-sequencing (scRNA-seq)). This amplification gives rise to an excess of zero expression values in single-cell sequencing data, at least some of which is corruption or noise in the form of non-biological zero expression values. The presence of excess zero expression values may render the sequencing data 125 unsuitable, in its raw form, for downstream analysis (e.g., clustering, visualization, and / or the like) by the analysis engine 130. Accordingly, the sequencing data 125 may undergo imputation in order to15131251356v lreduce excess zero expression values in the sequencing data 125. For example, in some cases, the sequencing controller 110 may apply an imputation model 115. In some cases, the imputation model 115 may have been trained to reduce excess zero expression values in the sequencing data 125 by correcting one or more of the zero expression values present in the sequencing data 125.

[0049] In some example embodiments, the imputation model 115 may be a machine learning model, such as a transformer model, that has been trained to correct one or more of the zero expression values present in the sequencing data 125 iteratively, over multiple successive timesteps. For example, at a first timestep, the imputation model 115 may correct one or more expression values in the sequencing data 125 by replacing at least a first expression value with a second expression value. Furthermore, at a second timestep, the imputation model 115 may further correct one or more additional expression values in the sequencing data 125 by replacing at least a third expression value with a fourth expression value. In some cases, the imputation model 115 may correct an expression value by replacing a zero expression value with a non-zero expression value. However, it should be appreciated that the imputation model 115 may also, in some cases, correct an expression value by replacing a non-zero expression value with another non-zero expression value.

[0050] In some example embodiments, the sequencing controller 110 may train, based at least on a training dataset, the imputation model 115 to correct one or more of the zero expression values present in the sequencing data 125. In some cases, each training sample in the training dataset may include a sequencing data sample in which one or more of the expression values present therein have been masked or dropped out by being replaced with zero expression values. For example, in some cases, the sequencing platform 120 may generate, for inclusion in16131251356v la sequencing data sample 127, the gene expression profile of one or more individual cells. In some cases, a training sample may be generated by applying, to the sequencing data sample 127, a dropout mask containing one or more zero expression values to replace one or more expression values in the sequencing data sample 127. Doing so may create training samples with one or more artificial zero expression values. In some cases, each artificial zero expression value may be associated with a ground truth expression value in the form of its original expression value.

[0051] In some example embodiments, the imputation model 115 may be trained to impute, for each artificial zero expression value in the training sample generated from the sequencing data sample 127, the corresponding original expression value that is present in the sequencing data sample 127. For example, in some cases, the training of the imputation model 115 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the imputation model 115 to reduce (or minimize) a difference between the expression value recovered by the imputation model 115 and the corresponding original expression value. For an artificial zero expression value masking a non-zero expression value in the sequencing data sample 127, the imputation model 115 may be trained to impute a non-zero expression value that is as close as possible in value to that original non-zero expression value in the sequencing data sample 127. Contrastingly, for the zero expression values that are present in the training sample, the imputation model 115 may be trained to avoid replacing any true zero expression values with a non-zero expression value. However, as described in more details below, an accurate differentiation between true zero expression values, such as biological zero expression values, and spurious zero expression values, such as non-biological zero expression values, may not be possible. As such, training the imputation model 115 to avoid replacing true zero expression17131251356v lvalues may include leveraging the output of the imputation model 1 15 to infer which zero expression values in the sequencing data sample 127 are true zero expression values.

[0052] It should be appreciated that the zero expression values that are present in a sequencing data sample, such as the sequencing data sample 127, may include biological zero expression values as well as non-biological zero expression values (e.g., sampling zeros, technical zeros, and / or the like). Whereas non-biological zero expression values is considered corruption or noise due to being attributable to genes that were expressed at the time of sequencing but are not measured, biological zero expression values are not corruption or noise because biological zero expression values are associated with genes that were not expressed at the time of sequencing. In other words, correcting excess zero expression values should include correcting non-biological zero expression values while preserving biological zero expression values. However, in practice, there is no practicable way to differentiate between biological zero expression values and non-biological zero expression values. As such, during training, the imputation model 115 may be first applied to the training sample in order to infer one or more true zero expression values present there. For example, in some cases, a threshold proportion of the lowest expression values in the output of the imputation model operating on the training sample may be designated as true zero expression values. In some cases, further training of the imputation model 115 may be supervised based on the artificial zero expression values as well as the true zero expression values present in the training sample. For instance, the imputation model 115 may be trained to recover the original non-zero expression values of the artificial zero expression values while preserving any true zero expression values that have been inferred from the output of the imputation model 115. In doing so, the imputation model 115 may learn which zero expression values in the sequencing data 125 are corruptions (or noise) that should be18131251356v lreplaced with non-zero expression values and which zero expression values in the sequencing data 125 should be kept as such.

[0053] In some cases, excess zero expression values may be a type of corruption (or noise) associated with batch effects. Accordingly, the correction of excess zero expression values may be viewed as a form of batch correction, or the reduction (or elimination) of corruption (or noise) associated with batch effects. Generally speaking, batch effects arise from non-biological factors altering the outcome of an experiment. In the context of single-cell sequencing, batch effect may include the pattern of corruption (or noise) in the form of excess zero expression values that may be present in the gene expression profiles output by a sequencing platform. For example, in some cases, different batches of cells that undergo singlecell sequencing at the sequencing platform 120 may exhibit different batch effects in the form of distinct patterns of corruption (or noise). In some cases, the imputation model 115 may be applied to batch correct excess zero expression values by operating on individual subsets (or “bags”) of cells from each batch of cells. For instance, in some cases, each subset (or bag) of cells may contain multiple cells, such that the imputation model 115 is applied to batch correct excess zero expression values in the gene expression profiles of multiple cells from the same batch of cells at once rather than correcting excess zero expression values in the gene expression profile of one individual cell at a time. As described in more detail below, for a batch of cells, the imputation model 115 may be applied to batch correct the corresponding gene expression profiles over multiple successive iterations, with the imputation model 115 operating on different subsets (or bags) of cells selected from the batch of cells during each iteration. Applying the imputation model 115 to batch correct the gene expression profiles of multiple cells at once may be more effective than applying the imputation model 115 to correct the gene expression profiles19131251356v lof one individual cell at a time. In some cases, the correction may be performed incrementally, over multiple successive iterations, until the quantity of spurious zero expression values (e.g., non-biological zero expression values) present in the gene expression profiles reaches a threshold value.

[0054] To further illustrate, FIG. IB depicts a system diagram illustrating an example of the imputation model 115 for correcting excess zero expression values in sequencing data, in accordance with some example embodiments. As shown in FIG. IB, in some example embodiments, the imputation model 115 may be applied to batch correct the gene expression profiles X of a subset (or “bag”) of cells. In some cases, the subset (or bag) of cells may contain multiple cells, such that the imputation model 115 is applied to batch correct excess zero expression values in the gene expression profiles of multiple cells at once rather than correcting excess zero expression values in the gene expression profile of one individual cell at a time. For example, in the example shown in FIG. IB, the subset (or bag) of cells may include, for each batch of cells that have undergone sequencing (e.g., at the sequencing platform 120), a b quantity of cells. Furthermore, the gene expression profiles of each of the b quantity of cells may include values corresponding to the gene expression levels of each of a g quantity of possible genes. Accordingly, the imputation model 115 may be applied to batch correct excess zero expression values in the gene expression profiles X = [x1, x2, ••• , xb] of each of the b quantity of cells, with

[0055] Referring again to FIG. IB, in some cases, the imputation model 115 may batch correct the gene expression profiles X = [1(x2, ••• , xb] of a batch of b quantity of cells by at least encoding the X = [xltx2, --- , xb] into one or more corresponding tokens 162. The encoding ofthe = [x1, x2, ,xb] may be tantamount to projecting (e.g., by a first linear projection layer20131251356v l) the gene expression profiles X into a lower dimensional latent space. The tokens 162 resulting therefrom may represent the gene expression profiles X — [xltx2, • •• , xb] with fewer quantity of latent variables (or features) and thus occupy the lower dimensional latent space than the gene expression profiles X = [x1, x2, ••• , xb] in their original form. In some cases, a transformer model 163 may ingest the tokens 162 and predict one or more corresponding imputation tokens 164 therefrom. For instance, in some cases, the transformer model 163 may ingest a token corresponding to a lower dimensional representation of a gene expression profile xLand generate an imputation token by modifying the token to correct for one or more of the zero expression values present in the gene expression profile xt. As shown in FIG. IB, in some cases, a second linear projection layer 165 may then decode the imputation tokens 164, which projects the imputation tokens 164 from the lower dimensional space back to the original space of the gene expression profiles X = [xltx2, --- , xbJ. This decoding operation may generate, for the imputation tokens 164, a corresponding set of imputation vectors A= [<51(<S2, ••• , <5b] with AS R^.

[0056] In some cases, each imputation vector 6tmay include the expression values imputed by the transformer model 163 for correcting the excess zero expression values present in a corresponding gene expression profile xt. For example, in some cases, each imputation vector 8i may include one or more non-zero expression values imputed by the transformer model 163 for replacing one or more zero or non-zero expression values in the corresponding gene expression profile x;. Accordingly, as shown in FIG. IB, the batch corrected gene expression profiles X may be determined by applying the imputation vectors A to the original gene expression profiles X = [x±, x2, , xb] (e.g., X = X + A). Doing so may include replacing one or more of the zero or non-zero expression values in the original gene expression profiles X —21131251356v l[xltx2, ••• , xb] with one or more non-zero expression values in the corresponding imputation vectors A = [<5,, 82, , <?£,].

[0057] FIG. 2A depicts a flowchart illustrating an example of a process 200 for training an imputation model to correct excess zero expression values in sequencing data, in accordance with some example embodiments. Referring to FIGS. 1A-B and 2A, the process 200 may be performed, for example, by the sequencing controller 110 to train the imputation model 115. In some cases, the imputation model 115 may be trained to correct one or more expression values in sequencing data. For example, in some cases, the imputation model 115 may be trained to correct one or more expression values in single-cell sequencing data, which tends to contain an excess of zero expression values. It should be appreciated that the imputation model 115 may correct zero expression values as well as non-zero expression values in sequencing data. Moreover, it should be appreciated that the imputation model 115 trained in the manner described in FIG. 2A may be deployed to correct expression values in any type of sequencing data, including single-cell sequencing data as well as bulk sequencing data.

[0058] At 202, a degraded sequencing data sample may be generated to include one or more artificial zero expression values by replacing one or more expression values in a sequencing data sample with the one or more artificial zero expression values. In some example embodiments, the imputation model 115 may be trained based on a training dataset that includes one or more training samples. In some cases, each training sample may include a degraded sequencing data sample generated by replacing, with one or more artificial zero expression values, one or more expression values in a sequencing data sample. For example, in FIG. 1A, a degraded sequencing data sample may be generated by replacing one or more expression values in the sequencing data sample 127 with one or more artificial zero expression values. In some22131251356v lcases, the imputation model 115 may be trained to replace each masked expression value in the training sample with an imputed expression value that is as close as possible in value to the original expression value in the sequencing data sample. For instance, the imputation model 115 may be trained to replace an artificial zero expression value in the degraded sequencing data sample that is masking a non-zero expression value in the sequencing data sample 127 with the same (or nearly the same) non-zero expression value as the original non-zero expression value in the sequencing data sample 127. Furthermore, the imputation model 115 may be trained to avoid replacing any true zero expression values from the sequencing data sample 127. However, as described in more details below, true zero expression values are inferred from the output of the imputation model 115 at least because true zero expression values in the form of biological zero expression values are indistinguishable from spurious zero expression values in the form of non- biological zero expression values (e.g., technical zeros, sampling zeros, and / or the like).

[0059] In some example embodiments, each training sample may be generated by applying a dropout mask configured to replace one or more expression values in the corresponding sequencing data sample with one or more artificial zero expression values. To further illustrate, consider the example of the dropout mask m and the sequencing data sample x below. For the sake of clarity, the sequencing data sample x below includes integer values corresponding to the measurements of gene expression levels. However, it should be appreciated that, in practice, the measurements of gene expression levels may also be expressed as nonintegers such as decimals, fractions, and / or the like. In this particular example, the dropout mask m includes a sequence of binary values. For example, whereas a value of 1 in the dropout mask m preserves the corresponding expression value in the sequencing data sample , a value of 0 in the dropout mask m replaces the corresponding expression value in the sequencing data sample23131251356v lx with a zero expression value. As shown in this example, the dropout mask m may include a value of 0 where the sequencing data sample x includes non-zero expression values as well as zero-expression values. Applying this dropout mask m generates a training sample in which masked expression values may appear the sequencing data sample x includes non-zero expression values and zero expression values. x = [0,0, 3, 4, 7, 0,2,0] m = [0,1, 1,1,0, 1,0,0]

[0060] At 204, an imputation model may be trained based at least on the degraded sequencing data sample. In some example embodiments, the imputation model 115 may be trained to correct excess zero expression values iteratively, over multiple successive timesteps. For example, a training sample may be generated by at least replacing, with one or more artificial zero expression values, one or more of the expression values present in a sequencing data sample. In some cases, the imputation model 115 may be trained to restore, over multiple successive timesteps, the sequencing data sample by at least replacing the artificial zero expression values with an expression value that is as close as possible to the original expression value.

[0061] In some example embodiments, the training of the imputation model 115 may include adjusting one or more parameters (e g., weights, biases, and / or the like) of the imputation model 115 to reduce (or minimize) an error (or loss) in the output of the imputation model 115. As described in more details below, one example of error (or loss) that may be reduced (or minimized) during the training of the imputation model 115 is the difference (e.g., mean square error (MSE)) between the original sequencing data sample and the sequencing data sample with the imputed expression values determined by the imputation model 115. In some cases, the24131251356v ltraining of the imputation model 115 may fail to achieve convergence, or a point in the training process of the imputation model 115 where the error (or loss) in the output of the imputation model stabilizes and / or reaches a threshold level, if the imputation model 115 is only trained to increase artificial zero expression values to non-zero expression values. As such, in addition to being trained to replace artificial zero expression values, the imputation model 115 may be also be trained to avoid replacing a true zero expression value (e.g., biological zero expression value) with a non-zero expression value. In other words, in some cases, training the imputation model 115 to restore the original sequencing data sample may include training the imputation model to replace each artificial zero expression values with an imputed expression value that is as close as possible to the original non-zero expression value while preserving any true zero expression values. However, despite true zero expression values (e.g., biological zero expression values) are indistinguishable from spurious zero expression values (e.g., non-biological zero expression values), training the imputation model 115 to avoid replacing true zero expression values may require identifying the true zero expression values present in the original sequencing data sample. Accordingly, as described in more details below, the training of the imputation model 115 may be bootstrapped in order to leverage the output of the imputation model 115 to identify one or more true zero expression values for supervising further training of the imputation model 115.

[0062] As noted, in some example embodiments, the imputation model 115 may be trained to restore the original sequencing data sample iteratively, over multiple successive timesteps. For example, in some cases, the imputation model 115 may be trained to perform inversion by direct iteration (InDI), a variation of incremental data restoration in which the imputation model 115 learns to invert the full degradation applied to the data incrementally, over25131251356v lmultiple successive timesteps. In the context of correcting excess zero expression values in sequencing data, the imputation model 115 may be trained to learn a continuous forward degradation process in which noise in the form of artificial zero expression values are added iteratively to a sequencing data sample. For instance, in some cases, the forward degradation process may be defined by Equation (1) below, in which noise is added iteratively to a data sample x to generate a degraded data sample y. As shown in Equation (1), starting with the data sample x at t = 0, the degraded data sample y may be generated at t = 1, with one or more intermediate degraded data samples xtbeing generated at each intermediate timestep. In Equation (1), xt, which is indexed by time t, may denote an intermediate degraded data sample between the fully degraded data sample y and the original data sample x. xt= (1 — t)x + ty, with t e [0,1] (1)

[0063] In some example embodiments, the imputation model 115 may be trained to learn a restoration process corresponding to the inverse of the forward degradation process. In the case of inversion by direct iteration (InDI), the imputation model 115 may be trained to learn this restoration process directly. By contrast, denoising diffusion and score-based models rely on analytically defining a known degradation process first before leveraging the known analytical degradation at every timestep of the restoration process. As the forward degradation process is incremental, the recovery process may be incremental as well. For example, in some cases, the imputation model 115 may be trained to learn a restoration process that inverts, iteratively over multiple successive timesteps, the continuous forward degradation process defined in Equation (1) above. The restoration process may start with the fully degraded data sample xt= y (at t =1), and at each intermediate timestep t — 6, generate the best possible restoration xt-sof the26131251356v lcorresponding intermediate restored data sample xt, until reaching x0, shown as Equation (2) below.

[0064] In some example embodiments, the imputation model 115 may be trained to learn this restoration process by adjusting one or more parameters (e.g., weights, biases, and / or the like) of the imputation model 115 to reduce (or minimize) the error (e.g., mean-squared error (MSE)) present in each incremental restoration xt-s. In some cases, the one or more parameters (e.g., weights, biases, and / or the like) of the imputation model 115 may be adjusted to reduce (or minimize) the loss function shown as Equation (3) below. nhnty; t) - x||2(3) wherein fedenotes the imputation model 115 and 6 denotes the parameters (e.g., weights, biases, and / or the like) of the imputation model 115 which, as noted, are adjusted during the training of the imputation model 115.

[0065] In instances where the data sample x is sequencing data, such as single-cell sequencing data, the degradation that may be added thereto may include the one or more artificial zero expression values set by the application of the dropout mask m 6 {l,0}G. As such, in some cases, the fully degraded data sample y may correspond to y = m O x whereas the intermediate degraded data sample x at time t may be defined as x = (1 — t)x + t(m 0 x ). Where the data sample x is sequencing data, the training of the imputation model 115 to invert the degradation that is present in the fully degraded data sample y may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the imputation model 115 to reduce (or minimize) the loss function shown as Equation (4) below.27131251356v l

[0066] At 206, the imputation model may be applied to correct one or more expression values in one or more gene expression profiles. In some example embodiments, once trained, the imputation model 115 may be applied to correct excess zero expression values in the gene expression profiles of one or multiple cells. In some cases, the imputation model 115 may correct excess zero expression values by at least replacing one or more zero expression values in the gene expression profiles X with non-zero expression values. For example, in some cases, the imputation model 115 may replace a spurious zero expression value with a non-zero expression value imputed by the imputation model 115 but preserve any true zero expression values. In some cases, the imputation model 115 may be applied to correct excess zero expression values over multiple successive timesteps. For instance, the imputation model 115 may replace at least a first zero expression value with a first non-zero expression value at a first timestep before replacing at least a second zero expression value with a second non-zero expression value at a second timestep. In some cases, in addition to replacing one or more zero expression values with non-zero expression values, the imputation model 115 may correct excess zero expression values by replacing one or more non-zero expression values with one or more different non-zero expression values. In the earlier example, the imputation model 115 may, at a third timestep, replace the first non-zero expression value and / or the second non-zero expression value with a third non-zero expression value. In some cases, the imputation model 115 may replace a nonzero expression value with another non-zero expression value in a subsequent timestep in order to further refine the non-zero expression values present in the sequencing data such that the nonzero expression values that are included in the final corrected sequencing data generated by the imputation model 115 are as close to the true expression values as possible.28131251356v l

[0067] As noted, in some example embodiments, in order for the training of the imputation model 115 to achieve convergence, the imputation model 115 may be trained to not only replace zero expression values with non-zero expression values but also when to preserve a zero expression value. This behavior would be consistent with the observation that some but not all of the zero expression values in the data sample x constitute corruption (or noise). In particular, whereas non-biological zero expression values (e.g., technical zeros, sampling zeros, and / or the like) are considered corruption (or noise), biological zero expression values are not. Instead, biological zero expression values may be considered true zero expression values, which the imputation model 115 should avoid correcting to a non-zero expression value during the iterative restoration (or denoising) of the data sample x. However, in practice, there is no practicable way to differentiate between the biological zero expression values (e.g., true biological zeros) and non-biological zero expression values (e.g., noise or corruption) that are present in the data sample x. As such, while the imputation model 115 should be trained to remove noise (or corruption) from the data sample x, a true (or clean) version of the data sample xtruemay not be available for supervising the training of the imputation model 115 to do so. Instead, in some cases, the imputation model 115 may be trained to recover the data sample x from the degraded data sample y, which is generated by adding further corruption (or noise) in the form of the artificial zero value expression values set by the dropout mask m.

[0068] The aforementioned phenomenon is further illustrated in the graph 300 depicted in FIG. 3 A. As shown in FIG. 3 A, during training, the imputation model 115 may be trained to recover the data sample x from the degraded data sample y, which is generated by adding known corruption (or noise) to the data sample x. However, during inference when the trained imputation model 115 is applied to denoise the degraded data sample y, the objective should be29131251356v lto recover the true, noiseless version of the data sample x denoted as xtrueinthe graph 300. FIG. 3 A shows that the data sample x that the imputation model 115 is trained to recover may not be the same as the true (or clean) version of the data sample xtrue. Instead, as noted, the true (or clean) version of the data sample xtruecontains biological zero expression values but not noise in the form of non-biological zero expression values whereas the data sample x still contains noise in the form of non-biological zero expression values. While the true (or clean) version of the data sample xtruemaY be sparse, meaning that the true (or clean) version of the data sample xtruecontain a large proportion of zero expression values, the true extent of sparsity (or the proportion of true zero expression values) may be difficult to estimate to the necessary level of accuracy and precision. Thus, training the imputation model 115 to recover the data sample x is not tantamount to training the imputation model 115 to recover the true (or clean) version of the data sample xtrue. Put another way, when trained to recover the data sample x, the trained imputation model 115 resulting therefrom may be incapable of recovering the true (or clean) version of the data sample xtrue.

[0069] In some example embodiments, the limitations associated with training the imputation model 115 to recover data sample x may be reduced by bootstrapping the output of the imputation model 115 as its input during training. An example of this bootstrap training is illustrated schematically in the graph 350 show in FIG. 3B. As shown in FIG. 3B, instead of training the imputation model 115 to recover the data sample x from the degraded data sample y, the imputation model 115 may be trained to recover the data sample x from an intermediate data sample estimated from the degraded data sample y. This technique is called “bootstrapping” the imputation model 115 to learn from itself because the intermediate data sample x^ is generated by applying the imputation model 115 to denoise the degraded data sample y. In other30131251356v lwords, in some cases, the imputation model 1 15 may be trained based on its own output, in this case the intermediate data sampleFor example, in some cases, the degraded data sample y may be generated by applying the dropout mask m to the data sample x. When bootstrapping the imputation model 115 to train on its own output, the dropout mask m may initially only set non-zero expression values in the data sample x. According to Equation (5), the intermediate data sample x^ may be generated by applying the imputation model 115, denoted fein Equation (5), to the data sample y generated by applying the dropout mask m to the data sample y-

[0070] In some example embodiments, the intermediate data samplegenerated by the imputation model 115 may be used to infer which zero expression values in the data sample x are true zero expression values (e.g., biological zero expression values) and which zero expression values in the data sample x are noise (e.g., non-biological zero expression values such as sampling zeros and technical zeros). For example, while the dropout mask m sets only nonzero expression values in the data sample x to artificial zero expression values when generating the degraded data sample y, one or more of the lowest expression values in the intermediate data sample x^ generated by the imputation model 115 denoising the degraded data sample y may be considered true zero expression values (e.g., biological zero expression values). In some cases, a threshold quantity of the lowest expression values in the intermediate data sample x^ such as a percentage of expression values corresponding to the expected sparsity of sequencing data, may be designated as true zero expression values for supervising further training of the imputation model 115.31131251356v l

[0071] Referring again to FIG. 3B, the intermediate data sample x may also be considered less noisy and a closer match to the true (or clean) data sample xtruethan the original data sample x due to the denoising performed by the imputation model 115. In some cases, bootstrapped training may be performed over multiple training iterations, with the parameters (e.g., weights, biases, and / or the like) of the imputation model 115 undergoing incremental adjustments at each training iteration such that the intermediate data sample x^ at each iteration is less noisy than the intermediate data sample x^ from a previous iteration. As described in more details below, bootstrapping reduces the complete lack of differentiation between true zero expression values (e.g., biological zero expression values) and noise (e.g., non-biological zero expression values) in the data sample x. That these true zero expression values are identified based on the intermediate data sample x^ generated by the imputation model 115 denoising the degraded data sample y means that the imputation model 115 can be better trained using the intermediate data sample x^ to reduce noise while preserving the true zero expression values. Recall that in FIG. 3 A, merely training the imputation model 115 to restore the data sample x from the degraded data sample y does not enable the trained imputation model 115 resulting therefrom to then restore the true (or clean) data sample xtrueat inference time at least because the imputation model 115 has not learned to differentiate between true zero expression values (e.g., biological zero expression values) and noise (e.g., non-biological zero expression values).

[0072] In some example embodiments, the bootstrapped training of the imputation model 115 may include training the imputation model 115 to restore the data sample x by denoising the intermediate data sample x^\ In some cases, this denoising may be supervised based on the knowledge of true zero expression values present in the intermediate data sample x^. As shown in Equation (6), the training of the imputation model 115 may include minimizing the difference32131251356v l(e g., mean square error (MSE)) between the data sample x and the reconstruction of the data sample x generated by the denoising of the intermediate data sample x^\ In some cases, the imputation model 115 may be trained by adjusting one or more parameters (e.g., weights, biases, and / or the like) of the imputation model 115 such that the imputation model 115 replaces one or more artificial zero expression values in the intermediate data samplewith an imputed expression value that is as close to the true expression value present in the true (or clean) data sample xtrue. As noted, one or more of the lowest expression values present in the intermediate data sample x^ may be inferred as true zero expression values in the data sample x. Thus, the intermediate data sample x may include the artificial zero expression values set in the degraded data sample y as well as true zero expression values inferred from the output of the imputation model 115 denoising on the degraded data sample y. In some cases, the imputation model 115 may be trained to restore the data sample x from the intermediate data sample x^ by replacing the artificial zero expression values in the intermediate data sample x , which correspond to noise (or corruption) in the form of non-biological zero expression values (e.g., technical zeros, sampling zeros, and / or the like), while preserving the true zero expression values (e.g., biological zero expression values).

[0073] FIG. 2B depicts a flowchart illustrating an example of a process 250 for training an imputation model to correct excess zero expression values in sequencing data, in accordance with some example embodiments. Referring to FIGS. 1A-B and 2A-B, the process 250 may be performed, for example, by the sequencing controller 110 to train the imputation model 115. In some cases, the process 250 may implement operation 204 of the process 200 shown in FIG. 2A.For example, in some cases, the training of the imputation model 115 may be bootstrapped in33131251356v lorder to leverage the output of the imputation model 115 to identify one or more true zero expression values present in each training sample. Doing so may enable the imputation model 115 to be trained to replace non-zero expression values while preserving true zero expression values when applied to denoise sequencing data that contains true zero expression values (e.g., biological zero expression values) as well as spurious zero expression values (e.g., non- biological zero expression values). In some cases, the process 250 may be performed to train the imputation model 115 to correct one or more expression values in single-cell sequencing data, which tends to contain an excess of zero expression values. It should be appreciated that by performing the process 250, the imputation model 115 may be trained to replace zero expression values as well as non-zero expression values present in the sequencing data. Moreover, it should be appreciated that the imputation model 115 trained in the manner described in FIG. 2B may be deployed to correct zero expression values in any type of sequencing data, including single-cell sequencing data as well as bulk sequencing data.

[0074] At 252, an imputation model may be applied to denoise a degraded sequencing data sample generated by masking one or more non-zero expression values in a sequencing data sample with one or more artificial zero expression values. In some example embodiments, the training of the imputation model 115 may be bootstrapped in order to leverage the output of the imputation model 115 to identify one or more true zero expression values present in the sequencing data sample associated with each training sample. For example, the degraded data sample y may be generated by applying the dropout mask m to mask one or more non-zero expression values in the data sample x with artificial zero expression values. In some cases, the data sample x may contain both true zero expression values (e.g., biological zero expression values) and spurious zero expression values (e.g., non-biological zero expression values). Thus,34131251356v lwhile the masking of one or more non-zero expression values in the data sample x may emulate spurious zero expression values, there may be no practicable way to identify the true zero expression values that are present in the data sample x. Nevertheless, the training of the imputation model 115 may fail to reach convergence if the imputation model 115 is only trained to replace artificial zero expression values without also preserving true zero expression values. As described in more details below, the complete lack of differentiation between true zero expression values (e g., biological zero expression values) and spurious zero expression values (e.g., non-biological zero expression values) in the data sample x may be reduced by applying the imputation model 115 to denoise the degraded data sample y and recover the intermediate data sample x^ therefrom, thereby inferring one or more true zero expression values present in the data sample x.

[0075] At 254, one or more true zero expression values present in the sequencing data sample may be identified based at least on an intermediate sequencing data sample generated by the imputation model denoising the degraded sequencing data sample. In some example embodiments, one or more true zero expression values present in the data sample x may be inferred from the intermediate data sample x generated by the imputation model 115 denoising the degraded data sample y. For example, in some cases, a threshold quantity of the lowest expression values present in the intermediate data sample x^ may be designated as true zero expression values for supervising the imputation model 115 during subsequent training. In some cases, the threshold quantity of the lowest expression values may correspond to the expected sparsity of sequencing data. In the case of single-cell sequencing data, this expected sparsity may be greater than 0% to account for unexpressed genes but lower than the 90-95% sparsity typically associated with single-cell sequencing data. For instance, in some cases, the expected35131251356v lsparsity of single-cell sequencing data may correspond to the 10-40% sparsity observed with bulk sequencing data.

[0076] At 256, the imputation model may be trained, based at least on the intermediate sequencing data sample, to replace the one or more artificial zero expression values while preserving the one or more true zero expression values. In some example embodiments, the imputation model 115 may be further trained based on the intermediate data samplewhich includes one or more artificial zero expression values from the degraded data sample y and the true zero expression values inferred from the intermediate data sample x^\ In other words, the training of the imputation model 115 may include additional supervision to prevent overcorrection by the imputation model 115. For example, in some cases, the bootstrapped training of the imputation model 115 may be supervised based on three different types of ground truth annotations including the original non-zero expression value of each artificial zero expression value, the true zero expression values, and the non-zero expression values that were not masked with artificial zero expression values. Accordingly, in some cases, the imputation model 115 may be trained to replace each artificial zero expression value with a non-zero expression value that is as close as possible to the original non-zero expression value masked by the artificial zero expression value. In some cases, the imputation model 115 may also be trained to avoid replacing the one or more true zero expression values with a non-zero expression values. Furthermore, the imputation model 115 may be trained to avoid replacing a non-zero expression value with a different non-zero expression value. In some cases, the training of the imputation model 115 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the imputation model 115 such that the imputation model 115 replaces each artificial zero expression value while preserving the true zero expression values. Furthermore, in some cases,36131251356v lthe training of the imputation model 115 may be iterative, meaning that the parameters (e.g., weights, biases and / or the like) of the imputation model 115 are adjusted incrementally, over multiple successive training iterations.

[0077] FIG. 4A depicts a flowchart illustrating another example of a process 400 for machine learning based imputation of sequencing data, in accordance with some example embodiments. Referring to FIGS. 1A-B and 4A, the process 400 may be performed by the sequencing controller 110 to correct the excess zero expression values present, for example, in the sequencing data 125 from the sequencing platform 120. In some cases, the process 400 may implement operation 206 of the process 200 shown in FIG. 2A. For example, in some cases, the sequencing controller 110 may apply the imputation model 115, which has been trained to replace the spurious zero expression values (e.g., non-biological zero expression values) with an imputed non-zero expression value while preserving the true zero expression values (e.g., biological zero expression values) in the sequencing data 125. In some cases, the sequencing data 125 may be single-cell sequencing data, which may be especially prone to containing excess zero expression values. For example, in some cases, sequencing data 125 may include the gene expression profdes of one or more individual cells or, in some cases, a subset of cells (e.g., bag of cells) for facilitating the correction of batch effects. However, it should be appreciated that the process 400 may also be performed to reduce (or eliminate) spurious zero expression values present in other types of sequencing data, such as bulk sequencing data.

[0078] At 402, one or more subsets of cells from a batch of cells that have undergone sequencing at a sequencing platform may be identified. In some example embodiments, a subset (or “bag”) of cells containing multiple cells from a batch of cells that have undergone sequencing, such as single-cell sequencing at the sequencing platform 120, may be identified for37131251356v lbatch correction. In the context of single-cell sequencing, each individual cell from the batch of cells may be associated with a gene expression profile that includes values corresponding to measurements of gene expression levels. Batch correction in the context of single-cell sequencing may include the correction of spurious zero expression values, such as non-biological zero expression values (e.g., technical zeros, sampling zeros, and / or the like), that may be present in the gene expression profiles of cells that have undergone single-cell sequencing. In some cases, the batch correction may be performed on the gene expression profiles of multiple cells from the same batch of cells that have undergone single-cell sequencing rather than the gene expression profile of one individual cell at a time. For example, in some cases, the gene expression profiles X of a subset (or bag) of cells that includes a b quantity of cells may be identified for batch correction of excess zero expression values by the imputation model 115. In some cases, the b quantity of cells may be a quantity that optimizes the tradeoffs between batch correction performance and computational resources (e.g., size of the imputation model 1 15 such as the number of parameters, network layers, transformer blocks, and / or the like). For instance, in some cases, the b quantity of cells may be a minimum quantity of cells for conveying sufficient batch effect information to enable batch correction by the imputation model 115. In some cases, the b quantity of cells may be a minimum quantity of cells for achieving satisfactory batch correction performance using the available computational resources.

[0079] At 404, an imputation model may be applied to correct one or more expression values present in sequencing data that includes a gene expression profile for each cell of the one or more subsets of cells. In some example embodiments, the imputation model 115 may be applied to correct the excess zero expression values in the gene expression profiles of one or more cells. In some cases, the imputation model 115 may be applied to batch correct the excess38131251356v lzero expression values in the gene expression profiles of multiple cells at once. As noted, in some cases, the gene expression profile of a cell may include values corresponding to the gene expression levels of each of a g quantity of possible genes. For the gene expression profiles of a b quantity of cells, the imputation model 115 may be applied to batch correct excess zero expression values in the gene expression profiles X = [xt, x2, ••• , xb] of each of the b quantity of cells,

[0080] In some cases, the imputation model 115 may correct excess zero expression values by at least replacing one or more zero expression values in the gene expression profiles X with non-zero expression values. For example, the imputation model 115 may replace a zero expression value in the gene expression profiles X with a non-zero expression value imputed by the imputation model 115. Furthermore, in some cases, the imputation model 115 may also replace a non-zero expression value, which may be an original non-zero expression value in the gene expression profile X or a value imputed by the imputation model 115 at a previous timepoint, with a different non-zero expression value imputed by the imputation model 115. As described earlier, the imputation model 115 may have undergone bootstrapped training in order to learn to differentiate between true zero expression values (e.g., biological zero expression values) and spurious zero expression values (e.g., non-biological zero expression values). Accordingly, when applied to the gene expression profiles X of the b quantity of cells, the imputation model 115 may replace spurious zero expression values (e.g., non-biological zero expression values) in the gene expression profiles X while preserving the true zero expression values (e.g., biological zero expression values). In some cases, the imputation model 115 may correct the excess zero expression values in the gene expression profiles X over multiple successive timesteps. For instance, the imputation model 115 may replace at least a first zero39131251356v lexpression value with a first non-zero expression value at a first timestep followed by replacing at least a second zero expression value with a second non-zero expression value at a second timestep.

[0081] At 406, one or more downstream analytical tasks may be performed based on corrected sequencing data generated by the imputation model. In some example embodiments, the gene expression profiles of one or more cells may undergo further analysis upon correcting the excess zero expression values present therein. For example, once the imputation model 115 has been applied to batch correct the excess zero expression values in the gene expression profiles X of the b quantity of cells, the analysis engine 130 may perform one or more downstream analytical tasks based on the batch corrected gene expression profiles X generated by the imputation model 115. It should be appreciated that variety of downstream analytical tasks are contemplated within the scope of the present disclosure, some examples of which include molecular subtyping, antibody screening, and / or the like.

[0082] FIG. 4B depicts a flowchart illustrating another example of a process 450 for machine learning based imputation of sequencing data, in accordance with some example embodiments. Referring to FIGS. 1 A-B and 4B, the process 450 may be performed by the sequencing controller 110 to correct the excess zero expression values present, for example, in the sequencing data 125 from the sequencing platform 120. In some cases, the process 450 may implement operation 206 of the process 200 shown in FIG. 2A. For example, in some cases, the sequencing controller 110 may apply the imputation model 115, which has been trained to replace the spurious zero expression values (e.g., non-biological zero expression values) with an imputed non-zero expression value while preserving the true zero expression values (e.g., biological zero expression values) in the sequencing data 125. In some cases, the imputation40131251356v lmodel 115 may be applied to batch correct excess zero expression values incrementally over multiple iterations. For example, to batch correct the gene expression profiles of a batch of cells that have undergone sequencing (e.g., single cell sequencing, batch sequencing, and / or the like) at the sequencing platform 120, the imputation model 115 may be applied to operate on the gene expression profiles of different subsets of cells in each iteration, such that the quantity of excess zero expression values present in the gene expression profiles of the batch as a whole is reduced incrementally over multiple successive iterations.

[0083] At 452, a first plurality of subsets of gene expression profiles may be identified in a set of gene expression profiles for a batch of cells that have undergone sequencing at a sequencing platform. In some example embodiments, batch correction to remove excess zero expression values from the gene expression profiles of a batch of cells that have undergone sequencing (e.g., single cell sequencing, batch sequencing, and / or the like) at the sequencing platform 120 may include applying the imputation model 115 to operate on individual subsets of gene expression profiles over multiple successive iterations of imputation. During each iteration of imputation, the gene expression profiles of the batch of cells may undergo multiple rounds (or sub-iterations) of imputation, with the imputation model 115 being applied to operate on different subsets of gene expression profiles during each round. For example, during a first round of imputation, the gene expression profiles of the batch of cells may be split into two or more subsets of gene expression profiles. In some cases, the split may be random, meaning that each gene expression profile may be randomly assigned to at least one subset of gene expression profiles. Moreover, in some cases, the split may be even, meaning that a same quantity of gene expression profiles may be assigned to each subset of gene expression profiles. While it is possible for the gene expression profiles to be split evenly amongst the subsets of gene41131251356v lexpression profiles, in instances where the split is not even, a single gene expression profile may be assigned to multiple subsets of gene expression profile such that each subset of gene expression profiles contained a same quantity of gene expression profiles.

[0084] At 454, an imputation model may be applied to correct one or more expression values in each subset of gene expression profiles from the first plurality of subsets of gene expression profiles to generate a first set of corrected gene expression profiles. In some example embodiments, during each iteration and each constituent round of imputation, the imputation model 115 may be applied to correct zero and non-zero expression values in multiple gene expression profiles at once. For example, during the first round of imputation where the gene expression profiles of the batch of cells is split into two or more subsets of gene expression profiles, the imputation model 115 may be applied to correct one or more expression values, including zero expression values as well as non-zero expression values, in each subset of gene expression profiles. That is, in addition to replacing one or more zero-expression values with one or more non-zero expression values imputed by the imputation model 115, the imputation model 115 may also replace one or more non-zero expression values with one or more different non-zero expression values imputed by the imputation model 115. In instances where the imputation model 115 replaces a non-zero expression value with a different non-zero expression value, the imputation model 115 may be refining a non-zero expression value imputed by the imputation model 115 during a previous round or iteration of imputation. It should be appreciated that the imputation model 115 may simultaneously operate on multiple gene expression profiles at once. For instance, the imputation model 115 may ingest a data structure (e.g., a matrix and / or the like) including every gene expression profiles in a subset of gene42131251356v lexpression profiles before operating on the data structure to correct one or more expression values, including zero expression values as well as non-zero expression values, present therein.

[0085] At 456, a second plurality of subsets of gene expression profiles may be identified in the set of gene expression profiles for the batch of cells. In some example embodiments, in addition to multiple iterations of imputation, each iteration of imputation may include multiple rounds of imputation. In some cases, during each round of imputation, the imputation model 115 may be applied to operate on different subsets of gene expression profiles from the gene expression profiles of the batch of cells. For example, where the imputation model 115 is applied to correct zero expression values in two or more subsets of gene expression profiles during a first round of imputation, two or more different subsets of gene expression profiles may be identified during a second round of imputation. It should be appreciated that the two or more different subsets of gene expression profiles may include a different combination of gene expression profiles than the subsets of gene expression profiles from the first round of imputation. For instance, in some cases, the two or more different subsets of gene expression profiles may also be identified by splitting the gene expression profiles of the batch of cells by at least randomly assigning each gene expression profile to at least one subset of gene expression profiles such that the same quantity of gene expression profiles are assigned to every subset of gene expression profiles.

[0086] At 458, the imputation model may be applied to correct one or more expression values in each subset of gene expression profiles from the second plurality of subsets of gene expression profiles to generate a second set of corrected gene expression profiles. In some example embodiments, during the second round of imputation, the imputation model 115 may be applied to correct zero expression values in the two or more different subsets of gene expression43131251356v lprofiles. The imputation model 115 may be trained to detect a spurious zero expression values (e.g., a non-biological zero expression value) in a gene expression profile based a context that includes other expression values in the same gene expression profile as well other gene expression profiles from the same batch of cells. As such, applying the imputation model 115 to operate on different subsets of gene expression profiles during each round of imputation may improve imputation performance by enabling the imputation model 115 to correct zero expression values in different contexts. In some cases, the imputation model 115 may also replace non-zero expression values in the two or more different subsets of gene expression profiles. For example, in some cases, the imputation model 115 may replace a first non-zero expression value included in the two or more different subsets of gene expression profiles with a second non-zero expression value imputed by the imputation model 115. In some cases, the first non-zero expression value may be imputed by the imputation model 115 during the first round of imputation preceding the second round of imputation. Accordingly, in some cases, the second non-zero expression value may be a refinement of the first non-zero expression value, meaning that the second non-zero expression value may be closer to the true expression value in the corresponding gene expression profile than the first non-zero expression value.

[0087] At 460, upon satisfying one or more criteria, determining, based at least on the first set and the second set of corrected gene expression profiles, a result of a current iteration of imputation. In some example embodiments, the result of each iteration of imputation may correspond to the corrected gene expression profiles that are generated during each constituent round of imputation. In some cases, the sequencing controller 110 may continue applying the imputation model 115 to perform one or more additional rounds of imputation until one or more criteria are satisfied. For example, in some cases, one or more additional rounds of imputation44131251356v lmay be performed until a threshold quantity of rounds of imputation has been performed. Alternatively and / or additionally, one or more additional rounds of imputation may be performed until each gene expression profile has undergone a threshold quantity of rounds of imputation. In some cases, upon satisfying the one or more criteria, the sequencing controller 110 may determine the result for a corresponding round of imputation. Where an iteration of imputation includes a first round of imputation and a second round of imputation, the result for that iteration of imputation may be determined based on a first set of corrected gene expression profiles generated during the first round of imputation and a second set of corrected gene expression profiles generated during the second round of imputation. In some cases, the result for that iteration of imputation may include, for each gene expression value present in the gene expression profiles of the batch of cells, a value determined based on the corresponding value in each set of corrected gene expression profiles. For instance, each gene expression value in the gene expression profiles of the batch of cells may be a summary value (e.g., mean, median, mode, range, minimum, maximum, and / or the like) of a first gene expression value from the first set of corrected gene expression profiles and a second gene expression value from the second set of corrected gene expression profiles.

[0088] At 462, one or more additional iterations of imputation may be performed to further correct one or more additional expression values in the result of the current iteration of imputation. In some example embodiments, the result from one iteration of imputation may undergo one or more additional iterations of imputation, with each additional iteration of imputation including one or more rounds of imputation. In some cases, each additional iteration of imputation may include the imputation model 115 being applied to correct one or more additional expression values in the gene expression profiles of the batch of cells. For example, in45131251356v lsome cases, after one iteration of imputation in which the imputation model 115 is applied to correct expression values in the gene expression profdes of the batch of cells, the imputation model 115 may be applied to correct expression values in the corrected gene expression profdes of the batch of cells for at least one additional iteration of imputation. It should be appreciated that the expression values are correct during each iteration of imputation may include zero expression values as well as non-zero expression values. In some cases, incrementally fewer incorrect expression values, including spurious zero expression values (e.g., non-biological zero expression values), may remain after each iteration of imputation. As such, in some cases, the imputation model 115 may be trained to correct fewer expression values during each successive iteration of imputation. In some cases, the sequencing controller 110 may continue applying the imputation model 115 to perform one or more additional iterations of imputation until one or more criteria, such as a threshold quantity of iterations, are satisfied.

[0089] A comparison of the respective imputation performances of the conventional imputation method Markov affinity-based graph imputation of cells (MAGIC) and the imputation model 115 described herein is depicted in FIG. 5. Imputation performance is quantified in terms of the relative root mean square error (RMSE) in the corrected samples as a function of sample average library size (or the total sum of counts across all genes for each cell). Relative root mean square error in this context may be defined as a ratio between the root mean square error (RMSE) of each imputation method and a baseline root mean square error (RMSE). In the context of imputing single-cell sequencing data, the baseline root mean square error (RMSE) may be considered 100% (or universal zero expression values) whereas a perfect reconstruction would exhibit a root mean square error of 0%. Referring to FIG. 5, graph 500 illustrates the performance of Markov affinity-based graph imputation of cells (MAGIC) on46131251356v lsamples containing 10% dropouts (or 10% artificial zero expression values) and samples containing 30% dropouts (or 30% artificial zero expression values). Meanwhile, graph 550 illustrates the performance of the imputation model 115 on samples with 10% dropouts (or 10% artificial zero expression values) and samples with 30% dropouts (or 30% artificial zero expression values). As shown in FIG. 5, the imputation model 115 achieved significantly better relative root mean square error (RMSE) than Markov affinity -based graph imputation of cells (MAGIC) for both samples containing 10% dropouts (27.4% versus 56.8%) and samples containing 30% dropouts (30.5% versus 65.4%). Furthermore, FIG. 5 shows that the imputation model 115 is able to maintain a consistent level of imputation performance across different average library size, including low quality single-cell sequencing data with small library sizes, whereas the imputation performance of Markov affinity -based graph imputation of cells (MAGIC) is especially poor for lower quality data.

[0090] FIG. 6 depicts a block diagram illustrating an example of a computing system 600, in accordance with some example embodiments. Referring to FIGS. 1-6, the computing system 600 may be used to implement the sequencing controller 110, the sequencing platform 120, the analysis engine, the client device 140, and / or any components therein.

[0091] As shown in FIG. 6, the computing system 600 can include a processor 610, a memory 620, a storage device 630, and input / output devices 640. The processor 610, the memory 620, the storage device 630, and the input / output devices 640 can be interconnected via a system bus 650. The processor 610 is capable of processing instructions for execution within the computing system 600. Such executed instructions can implement one or more components of, for example, the sequencing controller 110, the sequencing platform 120, the analysis engine, the client device 140, and / or the like. In some example embodiments, the processor 610 can be a47131251356v lsingle-threaded processor. Alternately, the processor 610 can be a multi-threaded processor. The processor 610 is capable of processing instructions stored in the memory 620 and / or on the storage device 630 to display graphical information for a user interface provided via the input / output device 640.

[0092] The memory 620 is a computer readable medium such as volatile or non-volatile that stores information within the computing system 600. The memory 620 can store data structures representing configuration object databases, for example. The storage device 630 is capable of providing persistent storage for the computing system 600. The storage device 630 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 640 provides input / output operations for the computing system 600. In some example embodiments, the input / output device 640 includes a keyboard and / or pointing device. In various implementations, the input / output device 640 includes a display unit for displaying graphical user interfaces.

[0093] According to some example embodiments, the input / output device 640 can provide input / output operations for a network device. For example, the input / output device 640 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e g., a local area network (LAN), a wide area network (WAN), the Internet).

[0094] In some example embodiments, the computing system 600 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, the computing system 600 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e g., planning functionalities (e.g., generating, managing, editing of spreadsheet48131251356v ldocuments, word processing documents, and / or any other objects, etc ), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via the input / output device 640. The user interface can be generated and presented to a user by the computing system 600 (e.g., on a computer screen monitor, etc.).

[0095] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0096] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object- oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device,49131251356v lsuch as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid- state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.

[0097] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.50131251356v l

[0098] In the descriptions above and in the claims, phrases such as “at least one of’ or“one or more of’ may occur followed by a conjunctive list of elements or features. The term“and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at least one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone, B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.

[0099] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not51131251356v lnecessarily require the particular order shown, or sequential order, to achieve desirable results.Other implementations may be within the scope of the following claims.52131251356v l

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method, comprising: generating a degraded sequencing data sample to include one or more artificial zero expression values, where the degraded sequencing data sample is generated by replacing one or more expression values in a sequencing data sample with the one or more artificial zero expression values; training, based at least on the degraded sequencing data sample, an imputation model; and applying the imputation model to correct one or more expression values in one or more gene expression profiles.

2. The method of claim 1, further comprising: generating the degraded sequencing data sample by applying a dropout mask to replace one or more non-zero expression values in the sequencing data sample with the one or more artificial zero expression values.

3. The method of any of claims 1 to 2, wherein the training of the imputation model includes applying the imputation model to denoise the degraded sequencing data sample, identifying, based at least on an intermediate sequencing data sample generated by the imputation model denoising the degraded sequencing data sample, one or more true zero expression values present in the sequencing data sample, and training, based at least on the intermediate sequencing data sample, the imputation model to replace any artificial zero expression values remaining in the intermediate sequencing data sample while preserving the one or more true zero expression values.53131251356v l4. The method of claim 3, wherein the sequencing data sample includes true zero expression values that are indistinguishable from spurious zero expression values, and wherein the intermediate sequencing data sample generated by the imputation model enables differentiation between the true zero expression values and the spurious zero expression values.

5. The method of any of claims 3 to 4, wherein one or more lowest expression values in the intermediate sequencing data sample are identified as true zero expression values.

6. The method of claim 5, wherein the one or more lowest expression values comprise a threshold quantity that corresponds to an expected proportion of true zero expression values present in a type of the sequencing data sample.

7. The method of any of claims 3 to 6, wherein the imputation model is further trained to avoid replacing a non-zero expression value in the intermediate sequencing data sample with a different non-zero expression value.

8. The method of any of claims 3 to 7, wherein the training of the imputation model includes adjusting one or more parameters of the imputation model to reduce a difference between a non-zero expression value imputed by the imputation model for each remaining artificial zero expression value and a corresponding original non-zero expression value present in the sequencing data sample.

9. The method of any of claims 1 to 8, further comprising: identifying, in a set of gene expression profiles for a batch of cells that have undergone sequencing at a sequencing platform, a first plurality of subsets of gene expression profiles; and applying the imputation model to generate a first set of corrected gene expression profiles by at least correcting one or more expression values each subset of gene expression profiles from the first plurality of subsets of gene expression profiles.54131251356v l10. The method of claim 9, further comprising: identifying, in the set of gene expression profdes for the batch of cells, a second plurality of subsets of gene expression profiles; applying the imputation model to generate a second set of corrected gene expression profiles by at least correcting the one or more expression values each subset of gene expression profile from the second plurality of subsets of gene expression profiles; and upon satisfying one or more criteria, determining, based at least on the first set and the second set of corrected gene expression profiles, a result of a current iteration of imputation.

11. The method of claim 10, further comprising: performing one or more additional iterations of imputation to correct one or more additional expression values present in the result of the current iteration of imputation.

12. The method of any of claims 1 to 11, wherein the imputation model corrects the one or more expression values in the one or more gene expression profiles by at least replacing, at a first timestep, at least a first expression value in the one or more gene expression profiles with a second expression value, and replacing, at a second timestep following the first timestep, at least the second expression value in the one or more gene expression profiles with a third expression value.

13. The method of claim 12, wherein the first expression value is a zero expression value, wherein the second expression value is a non-zero expression value, and wherein the third expression value is a different non-zero expression value.

14. The method of any of claims 12 to 13, wherein the third expression value is closer to a true expression value than the second expression value.

15. The method of any of claims 1 to 14, wherein the imputation model corrects the55131251356v lone or more expression values in the one or more gene expression profiles by at least encoding a gene expression profile to generate a token that represents the gene expression profile with fewer quantity of variables than the gene expression profile in its original form, modifying the token to correct the one or more expression values in the gene expression profile, decoding an imputation token generated by the modifying of the token, the decoding generating an imputation vector including one or more non-zero expression values for replacing one or more corresponding zero expression values in the gene expression profile, and generating a corrected gene expression profile by at least applying the imputation vector to the gene expression profile.

16. The method of any of claims 1 to 15, wherein the imputation model comprises a transformer model, a first linear projection layer coupled to an input of the transformer model, and a second linear projection layer coupled to an output of the transformer model.

17. The method of any of claims 1 to 16, wherein the one or more gene expression profiles comprise single-cell sequencing data.

18. The method of any of claims 1 to 17, wherein the one or more gene expression profiles include true zero expression values and spurious zero expression values, and wherein the imputation model is applied to correct the spurious zero expression values while preserving the true zero expression values.

19. The method of claim 18, wherein the true zero expression values correspond to biological zero expression values, and wherein the spurious zero expression values correspond to non-biological zero expression values.

20. The method of any of claims 1 to 19, further comprising:56131251356v lperforming, based at least on one or more corrected gene expression profdes generated by the imputation model, one or more downstream analytical tasks.

21. A system, comprising: at least one data processor; and at least one memory storing instructions, which when execute by the at least one data processor, result in operations comprising the method of any of claims 1 to 20.

22. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1 to 20.57131251356v l