Systems and methods for high-throughput prediction

The high-throughput single-cell RNA sequencing method with supervised learning predicts LOH with high sensitivity and specificity, addressing existing challenges in throughput and sequencing errors, enabling efficient LOH prediction in individual and population cells.

JP2026500098APending Publication Date: 2026-01-06REGENERON PHARMACEUTICALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025528915
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-12-01
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing methods for predicting loss of heterozygosity (LOH) are labor-intensive, time-consuming, and lack high-throughput capabilities, and current systems face challenges such as low throughput, high dropout rates, sequencing errors, and allele expression imbalances in single-cell RNA sequencing.

Method used

A high-throughput method using single-cell RNA sequencing (inferLOH) that employs supervised machine learning to predict LOH with greater than 99% sensitivity and specificity, utilizing unique molecular identifiers (UMIs) and machine learning models trained without requiring pure LOH cells, and addressing imbalanced allele expression and sequencing errors.

Benefits of technology

Enables accurate, high-throughput prediction of LOH in individual cells and cell populations with minimal sampling error, overcoming limitations of existing methods by achieving high sensitivity and specificity without needing pure LOH cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500098000001_ABST
    Figure 2026500098000001_ABST
Patent Text Reader

Abstract

Disclosed are systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in a target cell population, the system including a memory and a processor configured to receive genetic data for a first reference cell population, the processor configured to sequence the genetic data of the first reference cell population to obtain first reference data, identify and remove heterozygous mutation locations with imbalanced allelic expression in the first reference data to generate second reference data, map an identifier for each cell of the target cell population to the second reference data, and apply the mapped identifier for each cell of the target cell population to a supervised machine learning model, the processor further configured to receive one or more outputs from the model, at least one of the one or more outputs including LOH for the target cell population.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 429,949, filed December 2, 2022, the entire contents of which are incorporated herein by reference.

[0002] Sequence Listing Reference This application contains a Sequence Listing, filed December 1, 2023, 3,506 bytes in size, and submitted electronically as an XML file entitled 381204006SEQ, which is incorporated herein by reference.

[0003] This application relates generally to predictive modeling, and more particularly to systems and methods for predicting loss of heterozygosity (LOH). [Background technology]

[0004] Existing systems and methods for predicting loss of heterozygosity (LOH) require labor and time and cannot be applied to all cell types. The overall LOH rate is reported to be low, ranging from approximately 5% to 6%. Currently, there are no high-throughput methods that can accurately assess LOH rates. Therefore, a high-throughput method for accurately assessing LOH rates is highly anticipated. Summary of the Invention

[0005] Existing systems and methods for zygosity assessment include single nucleotide polymorphism (SNP) arrays combined with array comparative genomic hybridization (aCGH), SNP genotyping combined with fluorescence in situ hybridization (FISH), single-cell DNA sequencing (scDNAseq), bulk DNA sequencing (DNAseq), and bulk RNA sequencing (RNAseq)-based assessments (Boutin et al., Nature Communications, 2021; Alanis-Lobato et al., PNAS, 2021; Groff et al., Genome Research, 2019; Groff et al., Genome Research, 2019). However, limitations associated with known zygosity assessment methods include the need for cell cloning, low throughput, cost, and low resolution. Challenges associated with single-cell RNA sequencing include determining the minimum unique molecular identifier (UMI) coverage requirements to include in zygosity assessment for a single mutation locus, how to perform zygosity measurements for DNA segments of cells with multiple mutation loci, the threshold for whether overall unique molecular identifier coverage allows assessment of zygosity of a cell, and the frequency of sequencing errors.

[0006] In one aspect, the present disclosure relates to a high-throughput method for inferring loss of heterozygosity in individual cells using single-cell RNA sequencing (inferLOH). In another embodiment, the present disclosure relates to a high-throughput method for detecting the prevalence of loss of heterozygosity in a cell population by inferring the zygosity of chromosomal regions susceptible to loss of heterozygosity (LOHR) at the single-cell level. An advantage provided by the present disclosure is a system and method for assessing loss of heterozygosity (LOH) using single-cell RNA sequencing that overcomes limitations commonly associated with single-cell RNA sequencing, such as high dropout rates, low coverage, and sequencing errors. The system and method also overcome potential drawbacks in inferring DNA zygosity using observed RNA sequences, such as allele expression imbalance and RNA editing.

[0007] For example, in one aspect, the systems and methods of the present disclosure can predict loss of heterozygosity with greater than 99% sensitivity and greater than 99% specificity. Another advantage provided by the present disclosure is that a pure population of loss of heterozygosity cells, which may not be available, is not required as part of the training data to build a model for determining loss of heterozygosity. As a result, the present disclosure fulfills a long-standing need for an accurate, high-throughput method for examining loss of heterozygosity at the single-cell level, as well as within a population of cells, while providing a method that avoids the need for loss of heterozygosity cells, which may not be available and are difficult to generate.

[0008] In various aspects, a system for predicting the prevalence of loss of heterozygosity (LOH) in a target cell population is provided. In some aspects, the system may include at least one memory that stores computer-executable instructions and at least one processor in communication with the at least one memory. In some aspects, the at least one processor may be configured to execute computer-executable instructions to receive genetic data of a first reference cell population; sequence the genetic data of the first reference cell population to obtain first reference data, the first reference data comprising chromosomal identifiers, nucleotide coordinates, nucleotide compositions, or a combination thereof; identify and remove heterozygous mutation locations with imbalanced allele expression in the first reference data to generate second reference data; map an identifier for each cell of the target cell population to the second reference data; apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the mapped identifier for each cell of the target cell population, the model previously trained using historical data, the historical data comprising the mapped identifier for each cell of the target cell population and their corresponding LOH; receive one or more outputs from the model, at least one of the one or more outputs comprising the LOH for the target cell population, thereby predicting a prevalence of LOH in the target cell population; update the historical data to include the genetic data of the target cell population and the corresponding one or more outputs; and re-train the model using the updated historical data.

[0009] In some aspects, the at least one processor comprises computer-executable instructions for receiving genetic data of a first reference cell population; sequencing the genetic data of the first reference cell population to obtain first reference data, the first reference data comprising a chromosome identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying and removing heterozygous variant positions with imbalanced allele expression in the first reference data based on single-cell RNA sequencing data generated from a second reference cell population to generate second reference data; establishing a set of homozygous positions a predetermined number of nucleotide distance away from each variant position in the second reference data to generate third reference data; mapping an identifier for each cell of the target cell population to the second reference data; and The method may be configured to execute computer-executable instructions to: map an identifier for each cell in the population to the second reference data and / or the third reference data; apply one or more inputs to a supervised machine learning model, where the one or more inputs include the mapped identifier for each cell in the target cell population and the mapped identifier for each cell in the second reference cell population, the model previously trained using historical data, the historical data including the mapped identifier for each cell in the target cell population and their corresponding LOH; receive one or more outputs from the model, where at least one of the one or more outputs includes LOH for the target cell population, thereby predicting the prevalence of LOH in the target cell population; update the historical data to include genetic data for the target cell population and the corresponding one or more outputs; and re-train the model using the updated historical data.

[0010] The system may include additional, fewer, or alternative features, including those described elsewhere herein. In various embodiments, computer-implemented methods for predicting the prevalence of loss of heterozygosity (LOH) in a target cell population are provided. These methods may be performed using a system including a computer device, including a processor communicatively connected to a memory device. Additionally or alternatively, the computer-implemented methods may be performed via one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearables, smart watches, smart contact lenses, smart glasses, augmented reality glasses, virtual reality headsets, mixed reality or augmented reality glasses or headsets, voice or chatbots, ChatGPT bots, and / or other electronic or electrical elements that can communicate with each other via wired or wireless communication.

[0011] In some aspects, the method may include receiving genetic data of a first reference cell population; sequencing the genetic data to obtain first reference data, the first reference data comprising chromosomal identifiers, nucleotide coordinates, nucleotide composition, or a combination thereof; identifying and removing heterozygous mutation locations with imbalanced allelic expression in the first reference data to generate second reference data; mapping an identifier for each cell of the target cell population to the second reference data; applying one or more inputs to a supervised machine learning model, the one or more inputs comprising the mapped identifier for each cell of the target cell population, the model previously trained using historical data, the historical data comprising the mapped identifier for each cell of the target cell population and their corresponding LOH; receiving one or more outputs from the model, at least one of the one or more outputs comprising the LOH for the target cell population, thereby predicting a prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more outputs; and re-training the model using the updated historical data.

[0012] In some aspects, the method includes receiving genetic data of a first reference cell population; sequencing the genetic data of the first reference cell population to obtain first reference data, the first reference data comprising chromosomal identifiers, nucleotide coordinates, nucleotide compositions, or a combination thereof; identifying and removing heterozygous variant positions with imbalanced allele expression in the first reference data based on single-cell RNA sequencing data generated from a second reference cell population to generate second reference data; establishing a set of homozygous positions a predetermined number of nucleotide distance away from each variant position in the second reference data to generate third reference data; mapping identifiers of each cell of the target cell population to the second reference data; and and applying one or more inputs to a supervised machine learning model, the one or more inputs including the mapped identifier of each cell of the target cell population and the mapped identifier of each cell of the second reference cell population, the model previously trained using historical data, the historical data including the mapped identifier of each cell of the target cell population and their corresponding LOH; receiving one or more outputs from the model, at least one of the one or more outputs including the LOH of the target cell population, thereby predicting a prevalence of LOH in the target cell population; and updating the historical data to include genetic data of the target cell population and the corresponding one or more outputs; and re-training the model using the updated historical data.

[0013] Such methods may include additional, fewer, or alternative features, including those described elsewhere herein. In various aspects, at least one persistent computer-readable storage medium having computer-executable instructions embodied therein is provided. In some aspects, the computer-executable instructions, when executed by at least one processor, may cause the at least one processor to receive genetic data of a first reference cell population; sequence the genetic data of the first reference cell population to obtain first reference data, the first reference data comprising chromosomal identifiers, nucleotide coordinates, nucleotide compositions, or a combination thereof; identify and remove heterozygous mutation locations with imbalanced allele expression in the first reference data to generate second reference data; map an identifier for each cell of the target cell population to the second reference data; apply one or more inputs to a supervised machine learning model, the one or more inputs comprising the mapped identifier for each cell of the target cell population, the model previously trained using historical data, the historical data comprising the mapped identifier for each cell of the target cell population and their corresponding LOH; receive one or more outputs from the model, at least one of the one or more outputs comprising the LOH for the target cell population, thereby predicting a prevalence of LOH in the target cell population; update the historical data to include the genetic data of the target cell population and the corresponding one or more outputs; and re-train the model using the updated historical data.

[0014] In some aspects, the computer-executable instructions, when executed by at least one processor, cause the at least one processor to receive genetic data of a first reference cell population; sequence the genetic data of the first reference cell population to obtain first reference data, the first reference data comprising a chromosomal identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identify and remove heterozygous variant positions having imbalanced allele expression in the first reference data to generate second reference data; establish a set of homozygous positions a predetermined number of nucleotide distance away from each variant position in the second reference data to generate third reference data; map an identifier for each cell of the target cell population to the second reference data; and generate third reference data. The method includes mapping identifiers of each cell of the cell population to the second reference data and / or the third reference data; applying one or more inputs to a supervised machine learning model, the one or more inputs including the mapped identifiers of each cell of the target cell population and the mapped identifiers for each cell of the second reference cell population, the model previously trained using historical data, the historical data including the mapped identifiers of each cell of the target cell population and their corresponding LOH; receiving one or more outputs from the model, at least one of the one or more outputs including the LOH of the target cell population, thereby predicting the prevalence of LOH in the target cell population; and updating the historical data to include genetic data of the target cell population and the corresponding one or more outputs; and re-training the model using the updated historical data.

[0015] Such storage media may include additional, fewer, or alternative features, including those described elsewhere herein. In some aspects, the systems and methods of the present disclosure allow for prediction of the prevalence of loss of heterozygosity (LOH) in a target cell population without the need for cell cloning.

[0016] In some aspects, the systems and methods provide unique molecular identifiers in place of traditional read count-based zygosity determination. In some embodiments, the systems and methods provide machine learning-based prediction of loss of heterozygosity. In some embodiments, the systems and methods do not require "purified" loss of heterozygosity cells to train the machine learning-based model. Alternatively, simulated loss of heterozygosity (mLOH) cell data from wild-type (WT) cells can be obtained.

[0017] In some aspects, the systems and methods provide predictive models of LOH for individual cells while minimizing sampling error. In some embodiments, the systems and methods are cell type, CRISPR editing site, and 10x genome single-cell RNA sequencing platform independent.

[0018] Challenges in developing models to detect loss of heterozygosity include the fact that models developed based on current data (cell type, CRISPR editing site, 10x Genomics) may not be applicable to single-cell RNA sequencing data generated from other CRISPR editing sites, different cell types, or different technologies; pure loss of heterozygosity cells cannot be used to train predictive models; and there may be imbalances in gene allele expression.

[0019] In some aspects, the present disclosure also provides systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in a CRISPR-edited cell population, the systems and methods including: bulk DNA sequencing a first cell population having the same DNA genotype as WT cells to obtain a heterozygosity reference (HetRef) comprising a chromosome identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying heterozygous mutation locations with unbalanced allele expression in the HetRef based on single-cell RNA sequencing data generated from the WT cells; and identifying heterozygous mutation locations with unbalanced allele expression in the HetRef. removing heterozygous mutation positions to generate HetRef2; mapping the UMIs from each cell in the test sample to the reference genome to generate UMI coverage correlated to the HetRef2 coordinates; and calculating a zygosity score based on the UMI coverage associated with the two nucleotides registered at each HetRef2 position; generating a range of UMI coverage thresholds; and predicting the prevalence of LOH in the CRISPR-edited cell population based on the percentage of cells predicted to be LOH using the model, the UMI coverage thresholds, or a combination thereof.

[0020] In some aspects, the present disclosure provides systems and methods in which a mutation position is included in a HetRef if the mutation position is covered by at least about 20 DNA sequencing reads.

[0021] In some aspects, the present disclosure provides systems and methods in which a heterozygous mutation position is defined as having unbalanced allelic expression if the same nucleotide covers 80% or more of the UMIs mapped to the HetRef position.

[0022] In some aspects, the present disclosure provides systems and methods in which scRNAseq data is excluded from the calculation of the zygosity score if it does not meet a minimum threshold of the number of UMIs detected at a HetRef position.

[0023] In some aspects, the present disclosure provides systems and methods in which a position is considered heterozygous if each of the two alleles in bulk DNA sequencing is represented by between about 20% and about 80% of the reads in a DNA sequencing dataset.

[0024] In one embodiment of the present disclosure, bulk DNA sequencing refers to random sequencing of a plurality of cells in a mixture of pooled cells.In one embodiment of the present disclosure, bulk RNA sequencing refers to random sequencing of pooled cells.

[0025] In some aspects, the present disclosure provides systems and methods in which HetRef2 is a subset of HetRef. In some aspects, the present disclosure provides systems and methods whereby HetRef2 removes mutation positions that have unbalanced allelic expression from HetRef.

[0026] In some aspects, the present disclosure provides systems and methods where the minimum threshold is at least 4 UMI. In some aspects, the present disclosure provides systems and methods for use in determining LOH in cancer cells.

[0027] In some aspects, the present disclosure further provides systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in a CRISPR-edited cell population, the systems and methods comprising: DNA sequencing bulk DNA from a first cell population heterozygous for a DNA segment of interest to obtain a first heterozygous variant reference (HetRef) comprising chromosomal coordinates and nucleotide composition; single-cell RNA sequencing cells from a second cell population to identify and remove HetRefs with imbalanced allele expression in the first heterozygous variant reference to generate a second heterozygous variant reference (HetRef2); and generating a homozygous variant reference (HomRef) by establishing a set of homozygous positions that are a specific number of nucleotides away from each variant position in HetRef2; and determining the HetRef2 or HomRef coordinates. Using unique molecular identifiers (UMIs) generated from single-cell RNA sequencing cells from a second population of wild-type (WT) cells mapped to , train a model as WT cells or mock LOH (mLOH) cells to detect LOH, perform coverage and zygosity assessments (coverage is equal to the sum of all UMIs covering the HetRef2 or HomRef coordinates, and zygosity score is equal to the sum of UMIs of the least frequently covered alleles at each position in all coordinates registered in HetRef or HomRef divided by their respective coverage), perform single-cell RNA sequencing of cells from a cell population of known genotype (WT or LOH), validate the model for detecting LOH, and predict the prevalence of LOH in the cell population of interest based on the model for detecting LOH.

[0028] In some aspects, the present disclosure provides systems and methods in which a mutation position is only included in HetRef if at least about 20 DNA sequencing reads cover the position. In some aspects, the present disclosure provides systems and methods in which a heterozygous mutation position is defined as having unbalanced allelic expression if the same nucleotide covers 80% or more of the UMIs mapped to the HetRef position.

[0029] In some aspects, the present disclosure provides systems and methods in which scRNAseq data is excluded from the calculation of the zygosity score if it does not meet a minimum threshold of the number of UMIs detected at a HetRef position.

[0030] In some aspects, the present disclosure provides systems and methods in which a position is considered heterozygous if each of the two alleles in bulk DNA sequencing is represented by between about 20% and about 80% of the reads in a DNA sequencing dataset.

[0031] In some aspects, the present disclosure provides systems and methods in which HetRef2 is a subset of HetRef. In some aspects, the present disclosure provides systems and methods whereby HetRef2 removes mutation positions that have unbalanced allelic expression from HetRef.

[0032] In some aspects, the present disclosure provides systems and methods where the minimum threshold is at least 4 UMI. In some aspects, the present disclosure provides systems and methods for use in determining LOH in cancer cells.

[0033] In yet another aspect, the disclosure provides systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in a CRISPR-edited cell population, the systems and methods comprising: bulk DNA sequencing a first population of cells heterozygous for a DNA segment of interest to obtain a heterozygosity reference (HetRef) comprising a chromosomal identifier, a nucleotide coordinate, and a nucleotide composition; single-cell RNA sequencing a second population of cells not treated with CRISPR (WT) to obtain a dataset comprising a plurality of unique molecular identifiers (UMIs) and a sequence of each UMI and its mapping to a chromosomal location; single-cell RNA sequencing a third population of cells (test sample) comprising CRISPR-treated cells, which may include cells with loss of heterozygosity induced by the CRISPR procedure, to obtain a sequence of each UMI and its mapping to a chromosomal location from the test sample; and identifying heterozygous mutation locations with imbalanced allele expression in the HetRef based on the single-cell RNA sequencing data generated from the second population of WT sample cells. to generate HetRef2, generate a homozygous mutant reference (HomRef) based on HetRef2, map UMIs from each cell of the WT sample to HetRef2 to generate UMI coverage, calculate a WT zygosity score based on the UMI coverage, map UMIs from each cell of the WT sample to HomRef to generate UMI coverage, calculate a simulated loss of heterozygosity (mLOH) zygosity score based on the UMI coverage, and map UMIs from each cell of the test sample to Het Mapping to Ref2 to generate UMI coverage, calculating zygosity scores based on the UMI coverage, generating a range of UMI coverage thresholds, generating zygosity prediction model cases using WT and mLOH zygosity scores generated from WT or mLOH cells that meet each UMI coverage threshold within the threshold range, calculating the number of cells in the test sample that meet each UMI coverage threshold, and selecting the best model based on model performance and the number of cells in the test sample that meet the corresponding UMI coverage threshold to ensure model accuracy.and predicting the prevalence of LOH in cells of the test sample based on the percentage of cells predicted to be LOH using the best-fitting model and a UMI coverage threshold, minimizing sampling error within the test sample.

[0034] In some aspects, the present disclosure provides systems and methods in which a mutation position is included in a HetRef if the mutation position is covered by at least about 20 DNA sequencing reads.

[0035] In some aspects, the present disclosure provides systems and methods in which a heterozygous mutation position is defined as having unbalanced allelic expression if the same nucleotide covers 80% or more of the UMIs mapped to the HetRef position.

[0036] In some aspects, the present disclosure provides systems and methods in which scRNAseq data is excluded from the calculation of the zygosity score if it does not meet a minimum threshold of the number of UMIs detected at a HetRef position.

[0037] In some aspects, the present disclosure provides systems and methods that consider a position to be heterozygous if each of the two alleles in bulk DNA sequencing is represented by between about 20% and about 80% of the reads in a DNA sequencing dataset.

[0038] In some aspects, the present disclosure provides systems and methods in which HetRef2 is a subset of HetRef. In some aspects, the present disclosure provides systems and methods whereby HetRef2 removes mutation positions that have unbalanced allelic expression from HetRef.

[0039] In some aspects, the present disclosure provides systems and methods where the minimum threshold is at least 4 UMI. In some aspects, the present disclosure provides systems and methods for use in determining LOH in cancer cells.

[0040] In various aspects, the present disclosure provides systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in a cell population, the systems and methods comprising DNA sequencing bulk DNA from a first cell population heterozygous for a DNA segment of interest to obtain a first heterozygous variant reference (HetRef) comprising chromosomal coordinates and nucleotide composition, single-cell RNA sequencing cells from a second cell population to identify and remove HetRefs with imbalanced allele expression in the first heterozygous variant reference, generating a second heterozygous variant reference (HetRef2), and single-cell RNA sequencing cells from the second population with loss of heterozygosity to learn a model for detecting loss of heterozygosity. This involves generating a mock loss of heterozygosity (mLOH) homozygous mutant reference (HomRef) to train the model, single-cell RNA sequencing wild-type and mLOH cells, performing coverage and zygosity assessments (coverage equals the sum of all UMIs covering the HetRef2 or HomRef coordinate, and zygosity score equals the sum of UMIs of the least frequently covered alleles at each position in all coordinates registered in HetRef or HomRef divided by their respective coverage), single-cell RNA sequencing of cells from cell populations of known genotype (WT or LOH), validating the model for detecting LOH, and predicting the prevalence of LOH in the cell population of interest based on the model for detecting LOH.

[0041] In some aspects, the present disclosure provides systems and methods in which a mutation position is included in a HetRef if the mutation position is covered by at least about 20 DNA sequencing reads.

[0042] In some aspects, the present disclosure provides systems and methods in which a heterozygous mutation position is defined as having unbalanced allelic expression if the same nucleotide covers 80% or more of the UMIs mapped to the HetRef position.

[0043] In some aspects, the present disclosure provides systems and methods in which scRNAseq data is excluded from the calculation of the zygosity score if it does not meet a minimum threshold of the number of UMIs detected at a HetRef position.

[0044] In some aspects, the present disclosure provides systems and methods that consider a position to be heterozygous if each of the two alleles in bulk DNA sequencing is represented by between about 20% and about 80% of the reads in a DNA sequencing dataset.

[0045] In some aspects, the present disclosure provides systems and methods in which HetRef2 is a subset of HetRef. In some aspects, the present disclosure provides systems and methods whereby HetRef2 removes mutation positions that have unbalanced allelic expression from HetRef.

[0046] In some aspects, the present disclosure provides systems and methods where the minimum threshold is at least 4 UMI. In some aspects, the present disclosure provides systems and methods for use in determining LOH in cancer cells.

[0047] In various aspects, the present disclosure provides systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in a cell population, the systems and methods including single-cell RNA sequencing a target cell population to generate a coverage estimate and a zygosity score (where coverage is equal to the sum of all unique molecular identifiers (UMIs) and the zygosity score is equal to the sum of UMIs of less frequently covered alleles in all registered coordinates divided by their respective coverages), single-cell RNA sequencing a simulated loss of heterozygosity population (mLOH) of cells to generate a range of UMI coverage and a UMI coverage threshold, generating a model for predicting the prevalence of LOH in the cell population based on the single-cell RNA sequencing of the target cell population and the mLOH cell population, and predicting the prevalence of LOH in the target cell population based on the model and the UMI coverage threshold.

[0048] In some aspects, the present disclosure provides systems and methods in which a mutation position is included in a HetRef if the mutation position is covered by at least about 20 DNA sequencing reads.

[0049] In some aspects, the present disclosure provides systems and methods in which a heterozygous mutation position is defined as having unbalanced allelic expression if the same nucleotide covers 80% or more of the UMIs mapped to the HetRef position.

[0050] In some aspects, the present disclosure provides systems and methods in which scRNAseq data is excluded from the calculation of the zygosity score if it does not meet a minimum threshold of the number of UMIs detected at a HetRef position.

[0051] In some aspects, the present disclosure provides systems and methods that consider a position to be heterozygous if each of the two alleles in bulk DNA sequencing is represented in between about 20% and about 80% of the reads in a DNA sequencing dataset.

[0052] In some aspects, the present disclosure provides systems and methods in which HetRef2 is a subset of HetRef. In some aspects, the present disclosure provides systems and methods whereby HetRef2 removes mutation positions that have unbalanced allelic expression from HetRef.

[0053] In some aspects, the present disclosure provides systems and methods where the minimum threshold is at least 4 UMI. In some aspects, the present disclosure provides systems and methods for use in determining LOH in cancer cells.

[0054] In various aspects, the present disclosure also provides systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in the genome of cancer cells, the systems and methods including: performing bulk DNA sequencing on a first cell population heterozygous for a DNA of interest to obtain a heterozygosity reference (HetRef) comprising a chromosome identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying heterozygous mutation positions with unbalanced allele expression in HetRef based on single-cell RNA sequencing data generated by the test sample cells; removing positions with unbalanced allele expression in HetRef to generate HetRef2; mapping UMIs from each cell in the test sample to HetRef2 coordinates to generate a UMI coverage; calculating a zygosity score based on the UMI coverage; generating a threshold range for the UMI coverage; and predicting the prevalence of LOH in the cell population based on the percentage of cells predicted to be LOH using a model, the UMI coverage threshold, or a combination thereof.

[0055] In some aspects, the present disclosure provides systems and methods in which a mutation position is included in a HetRef if the mutation position is covered by at least about 20 DNA sequencing reads.

[0056] In some aspects, the present disclosure provides systems and methods in which a heterozygous mutation position is defined as having unbalanced allelic expression if the same nucleotide covers 80% or more of the UMIs mapped to the HetRef position.

[0057] In some aspects, the present disclosure provides systems and methods in which scRNAseq data is excluded from the calculation of the zygosity score if it does not meet a minimum threshold of the number of UMIs detected for a HetRef position.

[0058] In some aspects, the present disclosure provides systems and methods that consider a position to be heterozygous if each of the two alleles in bulk DNA sequencing is represented by between about 20% and about 80% of the reads in a DNA sequencing dataset.

[0059] In some aspects, the present disclosure provides systems and methods in which HetRef2 is a subset of HetRef. In some aspects, the present disclosure provides systems and methods whereby HetRef2 removes mutation positions that have unbalanced allelic expression from HetRef.

[0060] In some aspects, the present disclosure provides systems and methods where the minimum threshold is at least 4 UMI. In some aspects, the present disclosure provides systems and methods for use in determining LOH in cancer cells.

[0061] In various aspects, the present disclosure provides systems and methods for predicting the prevalence of loss of heterozygosity (LOH) in the genome of natural cells, the systems and methods including bulk DNA sequencing a first cell population heterozygous for a DNA segment of interest to obtain a heterozygous reference (HetRef) comprising a chromosome identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying heterozygous mutation positions with unbalanced allele expression in HetRef based on single-cell RNA sequencing data generated by the test sample cells; removing positions with unbalanced allele expression in HetRef to generate HetRef2; generating a UMI coverage; calculating a zygosity score based on the UMI coverage; generating a range of UMI coverage thresholds; and predicting the prevalence of LOH in the cell population based on the percentage of cells predicted to be LOH using a model, the UMI coverage thresholds, or a combination thereof.

[0062] In some aspects, the present disclosure provides systems and methods in which a mutation position is included in a HetRef if the mutation position is covered by at least about 20 DNA sequencing reads.

[0063] In some aspects, the present disclosure provides systems and methods in which a heterozygous mutation position is defined as having unbalanced allelic expression if the same nucleotide covers 80% or more of the UMIs mapped to the HetRef position.

[0064] In some aspects, the present disclosure provides systems and methods in which scRNAseq data is excluded from the calculation of the zygosity score if it does not meet a minimum threshold of the number of UMIs detected at a HetRef position.

[0065] In some aspects, the present disclosure provides systems and methods that consider a position to be heterozygous if each of the two alleles in bulk DNA sequencing is represented by between about 20% and about 80% of the reads in a DNA sequencing dataset.

[0066] In some aspects, the present disclosure provides systems and methods in which HetRef2 is a subset of HetRef. In some aspects, the present disclosure provides systems and methods whereby HetRef2 removes mutation positions that have unbalanced allelic expression from HetRef.

[0067] In some aspects, the present disclosure provides systems and methods where the minimum threshold is at least 4 UMI. In one aspect, the present disclosure provides methods used to determine LOH in cancer cells.

[0068] The technical effects of the systems and methods described herein may be achieved by performing the following steps: receiving genetic data of a first reference cell population; sequencing the genetic data of the first reference cell population to obtain first reference data, the first reference data comprising chromosomal identifiers, nucleotide coordinates, nucleotide compositions, or a combination thereof; identifying and removing heterozygous mutation locations with imbalanced allele expression in the first reference data to generate second reference data; mapping identifiers of each cell of a target cell population to the second reference data; applying one or more inputs to a supervised machine learning model, the one or more inputs comprising the identifiers mapped to each cell of the target cell population, the model having been previously trained using historical data, the historical data comprising the identifiers mapped to each cell of the target cell population and their corresponding LOH; receiving one or more outputs from the model, at least one of the one or more outputs comprising the LOH of the target cell population, thereby predicting the prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more outputs; and re-training the model using the updated historical data.

[0069] At least one of the technical problems faced by the systems and methods disclosed herein includes (i) limitations associated with known methods of zygosity assessment, including the need for cell cloning, low throughput, cost, and low resolution; and (ii) challenges associated with single-cell RNA sequencing, including determining the minimum unique molecular identifier (UMI) coverage requirements to include in zygosity assessment of a single mutation position, how to perform zygosity measurements on DNA segments of cells with multiple mutation positions, the threshold for whether zygosity of a cell can be assessed by overall unique molecular identifier coverage, and the frequency of sequencing errors.

[0070] The resulting technical effects may include: (i) overcoming limitations commonly associated with single-cell RNA sequencing, such as high dropout rates, small coverage, and sequencing errors; (ii) overcoming potential drawbacks when using observed RNA sequences to infer DNA zygosity, such as allele expression imbalance or RNA editing; (iii) the ability to predict loss of heterozygosity cells with greater than 99% sensitivity and greater than 99% specificity; (iv) the ability to predict loss of heterozygosity without requiring a pure population of cells with loss of heterozygosity; and (v) fulfilling a long-standing need for an accurate, high-throughput method to interrogate loss of heterozygosity at the single-cell level and within cell populations, while avoiding the need for cells with loss of heterozygosity, which may not be available and may be difficult to generate.

[0071] These and other aspects of the present invention will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. The following description, while indicating various embodiments and many specific details thereof, is not intended to be limiting and is provided by way of example. Numerous substitutions, modifications, additions, or rearrangements may be made within the scope of the present disclosure. [Brief explanation of the drawings]

[0072] [Figure 1] 1 illustrates a block diagram of an example computer system according to an embodiment of the present disclosure. [Figure 2] 2 illustrates a block diagram of a loss of heterozygosity (LOH) prediction computer device that may be used in the exemplary computer system illustrated in FIG. 1. [Figure 3] 3 illustrates a block diagram of an artificial intelligence (AI) / deep learning (DL) module that may be used in the LOH prediction computer device illustrated in FIG. 2. [Figure 4] 1 illustrates an example of CRISPR editing-induced LOH according to an exemplary embodiment. Only heterozygous coordinates based on WT cells are illustrated. [Figure 5]According to an exemplary embodiment, it is illustrated that most unique molecular identifiers (UMIs) derived from single-cell RNA sequencing and mapping to chromosomal regions with potential loss of heterozygosity (LOHR) do not cover heterozygous mutation positions registered in HetRef (or HetRef2), and that the coverage of heterozygous mutation positions by UMIs derived from single-cell RNA sequencing is generally low. [Figure 6] 1 shows the distribution of heterozygous mutation positions on chromosome 16 in wild-type (top) or LOH (bottom) cells identified using single-cell RNA sequencing and using a first heterozygous mutant reference as a reference, according to an exemplary embodiment. [Figure 7] According to an exemplary embodiment, a histogram (top row) of the total number of UMIs covering heterozygous mutation positions in HetRef2 at the LOHR of each wild-type cell, a histogram (middle row) of the total number of heterozygous mutation positions in HetRef2 covered by UMIs at the LOHR of each wild-type cell, and the coverage of heterozygous mutation positions in HetRef2 by UMIs at the LOHR of all wild-type cells are shown. [Figure 8] According to an exemplary embodiment, we illustrate how zygosity information from all eligible heterozygous mutation positions in HetRef2 at LOHR is collected to obtain an overall zygosity score and UMI coverage for each cell. [Figure 9] According to an exemplary embodiment, the total unique molecular identifier coverage by at least four UMIs per heterozygous mutation position of HetRef2 at LOHR of chromosome 16 in wild-type cells (top row) and the number of heterozygous mutation positions of HetRef2 covered by at least four UMIs at LOHR of chromosome 16 in wild-type cells are shown. [Figure 10]According to an exemplary embodiment, we illustrate how a homozygous reference (HomRef) containing homozygous positions within ±100 chromosomal coordinates of each heterozygous mutation position in HetRef2 of wild-type cells is generated, and how mock LOH (mLOH) cells can be generated from WT cells by focusing on the HomRef coordinates. [Figure 11] According to an exemplary embodiment, this figure illustrates how impurity information from all eligible heterozygous mutation positions of HetRef2 at LOHR in wild-type cells, LOH cells, or mLOH cells was collected to form coverage and zygosity scores for each cell. [Figure 12] 1 shows a histogram of the sum of UMIs covering the HetRef2 coordinates of LOHR on chromosome 16 in wild-type cells with at least four UMIs for each mutation position, according to an exemplary embodiment. [Figure 13] In accordance with an exemplary embodiment, we illustrate that coverage requirements for a minimum number of UMIs can have adverse effects on model accuracy and sampling error. [Figure 14] According to an exemplary embodiment, we illustrate how dynamically determining the coverage requirements for determining a cell as zygosity evaluable can maximize model accuracy and minimize sampling error. [Figure 15] According to an exemplary embodiment, methods for generating, testing, and using logistic regression models to predict the prevalence of LOH in individual cells and in cell populations are outlined. [Figure 16] 1 illustrates how zygosity score thresholds and correct prediction rates correlate with UMI coverage requirements, according to an exemplary embodiment. [Figure 17]1 shows an example of a logistic regression model for predicting whether a cell is a wild-type cell or an LOH cell. According to an exemplary embodiment, by default, the probability cutoff is set to 0.5 (top panel), and the corresponding zygosity score (ΣMi / (ΣMi+ΣMa)) is defined as the zygosity score threshold. The zygosity score thresholds for various model instances are generated from different UMI coverage thresholds (ΣMi+ΣMa) (bottom panel). [Figure 18] 1 shows the predicted LOH prevalence for test sample 1 (top) and test samples 2 and 3 (bottom) using the model example shown therein, according to an exemplary embodiment. [Figure 19] 1 illustrates a workflow of the disclosed system and method according to an example embodiment. [Figure 20] According to an exemplary embodiment, the disclosed systems and methods are used to illustrate how LOH cells can undergo complete or partial loss of heterozygosity at LOH (top panel) and determine the effect of LOH on downstream gene expression (bottom panel). [Figure 21-1] The figures show examples of the number of cells that can be evaluated for zygosity that meet the minimum requirement of 1,000 cells for test sample 1 (TS1), test sample 2 (TS2), and test sample 3 (TS3). [Figure 21-2] The figures show examples of the number of cells that can be evaluated for zygosity that meet the minimum requirement of 1,000 cells for test sample 1 (TS1), test sample 2 (TS2), and test sample 3 (TS3). [Figure 21-3] The figures show examples of the number of cells that can be evaluated for zygosity that meet the minimum requirement of 1,000 cells for test sample 1 (TS1), test sample 2 (TS2), and test sample 3 (TS3). [Figure 21-4] The figures show examples of the number of cells that can be evaluated for zygosity that meet the minimum requirement of 1,000 cells for test sample 1 (TS1), test sample 2 (TS2), and test sample 3 (TS3). [Figure 21-5]The figures show examples of the number of cells that can be evaluated for zygosity that meet the minimum requirement of 1,000 cells for test sample 1 (TS1), test sample 2 (TS2), and test sample 3 (TS3). [Figure 21-6] The figures show examples of the number of cells that can be evaluated for zygosity that meet the minimum requirement of 1,000 cells for test sample 1 (TS1), test sample 2 (TS2), and test sample 3 (TS3). [Figure 22-1] According to an exemplary embodiment, we illustrate model cases that have been trained and validated using cells that meet UMI coverage requirements. [Figure 22-2] According to an exemplary embodiment, we illustrate model cases that have been trained and validated using cells that meet UMI coverage requirements. [Figure 22-3] According to an exemplary embodiment, we illustrate model cases that have been trained and validated using cells that meet UMI coverage requirements. [Figure 22-4] According to an exemplary embodiment, we illustrate model cases that have been trained and validated using cells that meet UMI coverage requirements. DETAILED DESCRIPTION OF THE INVENTION

[0073] Disclosed herein are machine learning methods. The computer-implemented methods described herein may include additional, fewer, or alternative actions, including those described elsewhere herein. These methods may be performed via computer-executable instructions stored on one or more local or remote processors, transceivers, servers, and / or persistently computer-readable medium(s).

[0074] In addition, the computer systems described herein may include additional, fewer, or alternative functionality, including functionality described elsewhere herein. The computer systems described herein may include or be executed by computer-executable instructions stored on a persistent computer-readable medium(s).

[0075] The processor or processing element can learn using supervised or unsupervised machine learning, and the machine learning program may employ a neural network, which may be a convolutional neural network, a deep learning neural network, a reinforcement learning or reinforcement learning module or program, or a hybrid learning module or program that trains in two or more fields or areas of interest. Machine learning may identify and recognize patterns in existing data to facilitate prediction of subsequent data. Models may be created based on example inputs to make valid and reliable predictions about new inputs.

[0076] Additionally or alternatively, machine learning programs may learn by inputting sample data sets or specific data, such as genetic and / or other data about cell populations. Machine learning programs may primarily utilize deep learning algorithms focused on pattern recognition and may learn after processing multiple examples. Machine learning programs may include Bayesian Program Learning (BPL), speech recognition and synthesis, image or object recognition, optical character recognition, and / or natural language processing, individually or in combination. Machine learning programs may also include natural language processing, semantic analysis, automated reasoning, and / or machine learning.

[0077] Both supervised and unsupervised machine learning techniques may be used. In supervised machine learning, a processing element is provided with example inputs and their associated outputs and discovers general rules that map the inputs to outputs, so that when a subsequent new input is provided, the processing element can accurately predict the correct output based on the discovered rules. In unsupervised machine learning, a processing element may be required to find unique structures within unlabeled example inputs. In some embodiments, machine learning techniques may be used to extract data regarding the specific prevalence of loss of heterozygosity (LOH) in a target cell population from genetic and / or other data of a reference cell population.

[0078] In some embodiments, the voicebots or chatbots described herein may be configured to utilize ML and / or AI techniques. For example, the voicebots or chatbots may be artificial intelligence (AI) bots (including generative AI bots). AI bots may employ supervised or unsupervised machine learning techniques, which may be followed by and / or used in conjunction with reinforcement learning or reinforcement learning techniques. AI bots may employ techniques used in ChatGPT.

[0079] As will be understood based on the foregoing description, the above-described embodiments of the present disclosure may be implemented using computer programming or engineering techniques, including computer software, firmware, hardware, or any combination or subset thereof. The resulting program having computer-readable code means may be embodied or provided in one or more computer-readable media, thereby creating a computer program product, i.e., an article of manufacture, in accordance with the embodiments discussed in this disclosure. The computer-readable medium may be, for example, but not limited to, a fixed (hard) drive, a diskette, an optical disk, a magnetic tape, a semiconductor memory such as a read-only memory (ROM), an SD card, a memory device, and / or any transmission / reception medium, such as the Internet or other communications network or link. An article of manufacture containing computer code may be created and / or used by executing the code directly from one medium, by copying the code from one medium to another, or by transmitting the code over a network.

[0080] These computer programs (also known as programs, software, software applications, "apps," or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or equipment (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, such as a machine-readable medium that receives machine instructions as machine-readable signals. However, "machine-readable medium" and "computer-readable medium" do not include transitory signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0081] As used herein, a processor may include any programmable system, including systems that use microcontrollers, reduced instruction set circuits (RISC), application specific integrated circuits (ASIC), logic circuits, and other circuits or processors that perform the functions described herein. The above examples are merely examples and are not intended to limit in any way the definition and / or meaning of the term "processor."

[0082] As used herein, "software" and "firmware" are used interchangeably and include any computer program stored in memory for execution by a processor, such as RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are exemplary only and are not intended to limit the types of memory that can be used to store computer programs.

[0083] In some embodiments, a computer program is provided, the program being embodied on a computer-readable medium. In an exemplary embodiment, the system runs on a single computer system without requiring connection to a server computer. In a further embodiment, the system runs in a Windows® (Windows is a registered trademark of Microsoft Corporation, Redmond, Washington) environment. In yet another embodiment, the system runs in a mainframe environment and a UNIX® (UNIX is a registered trademark of X / Open Company Limited, located in Reading, Berkshire, United Kingdom) server environment. The application is flexible and designed to run in a variety of environments without compromising its primary functionality.

[0084] In some embodiments, a system includes multiple elements distributed across multiple computing devices. One or more elements may be in the form of computer-executable instructions embodied on a computer-readable medium. The system and processes are not limited to the specific embodiments described herein. In addition, each system element and each process can be performed independently and independently of other elements and processes described herein. Each element and process can also be used in combination with other assembly packages and processes. The embodiments may enhance the functionality and capabilities of a computer and / or computer system.

[0085] The claims set forth at the end of this document are not intended to be construed under 35 U.S.C. § 112(f) unless traditional functional language is expressly recited in the claim(s), such as by expressly reciting "means for" or "step for" language.

[0086] This written description uses the disclosed examples to disclose the best mode, and also to enable any person skilled in the art to practice the disclosure, including making and using any devices or systems, and practicing the incorporated methods. The patentable scope of the disclosure is defined by the claims, and may include other examples that occur to those skilled in the art. If the embodiment has structural elements that do not differ from the literal language of the claims, or if the embodiment includes equivalent structural elements that are not substantially different from the literal language of the claims, then such other embodiments are intended to be within the scope of the claims.

[0087] An exemplary computer system for predicting loss of heterozygosity (LOH) is disclosed herein. For example, FIG. 1 shows a schematic diagram of an exemplary computer system 100. The computer system 100 is configured to predict LOH in a cell population. In one exemplary embodiment, the computer system 100 may include and / or facilitate communication between an LOH prediction computer device 110 and one or more user computer devices 130 (also known as "mobile devices"), and / or between the LOH prediction computer device 110 and one or more third-party devices 140 and / or LOH prediction servers 150.

[0088] The LOH prediction computing device 110 may be implemented as a server computing device with artificial intelligence and deep learning capabilities. Alternatively, the LOH prediction computing device 110 (and / or user computing device 130) may be implemented as any device capable of interconnecting to the Internet, including a mobile computing device or “mobile device,” such as a smartphone, “phablet,” or other web-enabled device or mobile device (e.g., one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearables, smart watches, smart contact lenses, smart lenses, augmented reality glasses, virtual reality headsets, mixed or augmented reality glasses or headsets, voice bots or chat bots, artificial intelligence (AI) bots (including generative AI bots), and / or other electronic or electrical components that may communicate with each other via wired or wireless communications).

[0089] LOH prediction computer device 110 may communicate with one or more user computer devices 130, third party devices 140, and / or LOH prediction server 150, such as through wireless communication or data transmission over one or more radio frequency links or wireless communication channels. In an exemplary embodiment, elements of computer system 100 may be communicatively connected to the Internet through a number of interfaces, including, but not limited to, at least one of the Internet, a local area network (LAN), a wide area network (WAN), or a network such as an integrated services digital network (ISDN), a dial-up connection, a digital subscriber line (DSL), a cellular communication connection (e.g., a 3G, 4G, 5G, etc. connection), a cable modem, a BLUETOOTH® connection, and the like.

[0090] Computer system 100 also includes one or more database(s) 120 containing information about various matters. For example, database 120 may include information such as genetic data and / or other information used, received, and / or generated by computer system 100 and / or its elements, such as information described herein. In one exemplary embodiment, database 120 may include a cloud storage device, where information stored therein may be securely stored but still accessible by one or more elements of computer system 100, such as LOH prediction computer device 110, user computer device 130, and / or LOH prediction server 150. In some embodiments, database 120 may be stored on LOH prediction computer device 110. In alternative embodiments, database 120 may be stored remotely from LOH prediction computer device 110, avoiding centralization.

[0091] In some embodiments, user computing device 130 may be a computer including a web browser or software application that allows user computing device 130 to access the functionality of LOH prediction computing device 110 using the internet or a direct connection, such as a cellular network connection. User computing device 130 may be any device that can access the internet, such as, but not limited to, a desktop computer, a mobile device (such as a laptop computer, personal digital assistant (PDA), mobile phone, smartphone, tablet, phablet, netbook, notebook, smartwatch or bracelet, smart glasses, wearable electronics, pager, virtual reality headset, augmented reality glasses, voice or chat bot, wearable, etc.), or other web-based, connectable device.

[0092] User computing device 130 may use data management app 112, for example, via user interface 132, to access data management app 112 maintained by LOH prediction computing device 110 when data management app 112 executes on user computing device 130. A user may use data management app 112 to provide input to LOH prediction computing device 110, view predictions generated by LOH prediction computing device 110, and perform other actions, including those described elsewhere herein, by LOH prediction computing device 110.

[0093] The third party device 140 may be a computing device associated with an external data source. The LOH prediction computing device 110 may request, receive, and / or otherwise access data from the third party device 140. The third party device 140 may be any device capable of interconnecting with the Internet, such as a server computing device, a mobile computing device or a "mobile device" such as a smartphone, or other web-enabled or mobile device.

[0094] An exemplary user analysis computing device is disclosed herein. For example, FIG. 2 illustrates an LOH prediction computing device 110 (shown in FIG. 1) according to an embodiment. In some embodiments, the LOH prediction computing device 110 may include a processor 202, a memory 204 (which may be similar to the database 120 also shown in FIG. 1), a communication interface 206, and a storage interface 208. The processor 202 is configured to execute instructions that may be stored in the memory 204. The processor 202 may include one or more processing units (e.g., a multi-core configuration) and be configured to execute multiple modules.

[0095] In some embodiments, processor 202 is operable to execute artificial intelligence / deep learning (AI / DL) module 210, LOH prediction module 212, and module 214 that maintains the functionality of data management app 112 (shown in FIG. 1 ). Modules 210, 212, and 214 may include specialized instruction sets and / or coprocessors. Database 120 and / or memory 204 may store data and / or instructions necessary for modules 210, 212, and 214 to function as described herein. In exemplary embodiments, database 120 may store genetic data 220, sequence data 222, and / or other information used, received, and / or generated by LOH prediction computing device 110.

[0096] The AI / DL module 210 may perform artificial intelligence and / or deep learning functions on behalf of the LOH prediction module 212. In particular, the AI / DL module 210 may include rules, algorithms, training datasets / programs, and / or other suitable data and / or executable instructions that enable the LOH prediction computing device 110 to predict LOH in a cell population using artificial intelligence and / or deep learning.

[0097] 3 illustrates an AI / DL module 210 (shown in FIG. 2) according to an embodiment. In some embodiments, the AI / DL module 210 includes a training set builder module 302 that is programmed to send one or more queries to database 120 (shown in FIGS. 1 and 2) to retrieve data and / or subsets of data and use those subsets to build a training dataset for generating a predictive model 308.

[0098] In an exemplary embodiment, the training set builder module 302 is programmed to obtain training datasets from a subset of the obtained data. Each training dataset corresponds to genetic data of a cell population and the corresponding LOH for each cell in the cell population. In some embodiments, the training data corresponds to identifiers mapped to each cell in the cell population and their corresponding LOH. The historical data may include pre-determined LOH. Each training dataset may include model input data along with outcome data representing LOH. The model input data may represent factors that may be expected to have some correlation with LOH or may be unexpectedly observed during model training.

[0099] The set builder module 302 generates training data sets and passes the training data sets to a model learning module 304, which is programmed to operate the model input data fields of each training data set as inputs to one or more machine learning models. Each of the one or more machine learning models is programmed to generate, for a respective training data set, at least one output that corresponds to or is intended to "predict" the value of at least one outcome data field of the training data set. Machine learning can include a variety of algorithms that can be used to identify and recognize patterns in existing data to train models that facilitate predictions for subsequent new input data.

[0100] For each training dataset, the model trainer module 304 is programmed to compare at least one output of the model with at least one outcome data field of the training dataset and apply a machine learning algorithm to adjust the model's parameters to reduce the difference, or "error," between the at least one output and the corresponding at least one outcome data field. In this way, the model trainer module 304 trains the machine learning models to accurately predict the LOH of the inputs. In other words, the model trainer module 304 iteratively applies one or more machine learning models through the training dataset, adjusting the model parameters until the error between the at least one output and the LOH falls below an appropriate threshold, and then uploads the at least one trained machine learning model to the predictive model module 308 for application to new structural data (e.g., primary amino acid sequences).

[0101] In some embodiments, the one or more machine learning models may include one or more neural networks, such as a convolutional neural network or a deep learning neural network. The neural network may have one or more layers of nodes, and the model parameters adjusted during training may be individual weight values ​​applied to one or more inputs for each node that generates a node output. In other words, the nodes in each layer receive one or more inputs and apply weights to the inputs to generate a node output. The node inputs for the first layer may correspond to model input data fields, and the node outputs for the final layer may correspond to at least one output of the model intended to predict at least one outcome data field. One or more intermediate layer nodes may be connected between the nodes in the first layer and the nodes in the final layer. As the model trainer module 304 processes the training data set in sequence, the model trainer module 304 applies a suitable backpropagation algorithm to adjust the weights of each node layer to minimize the error between at least one output and the corresponding outcome data field. In this manner, the machine learning model is trained to generate one or more outputs that reliably predict the LOH of the cell population. Alternatively, the machine learning model may have any suitable structure. In some embodiments, the model trainer module 304 provides the advantage of automatically detecting and appropriately weighting complex second- or third-order and / or other nonlinear interconnections between the model input data fields and at least one output, where without the machine learning model, such connections would be unexpected and / or undetectable by a human analyst.

[0102] Additionally or alternatively, the one or more machine learning models may include one or more multi-layer perceptron (MLP) classifiers, which may include an input layer, an output layer, and one or more hidden layers with a large number of neurons stacked on top of each other.

[0103] Additionally or alternatively, the one or more machine learning models may include one or more support vector machines (SVMs). SVMs are supervised learning models with associated learning algorithms that analyze data for classification and regression analysis. More specifically, SVMs construct a hyperplane or set of hyperplanes in a high- or infinite-dimensional space and can be used for classification, regression, or other tasks such as outlier detection.

[0104] In some embodiments, the predictive model module 308 compares the known LOH of the cell population with the output from the trained model and passes the comparison results to the model update module 306 of the AI / DL module 210. The model update module 306 is programmed to use the comparison results to update or "re-train" at least one machine learning model to improve performance. The re-trained machine learning model may be periodically re-uploaded to the predictive model module 308.

[0105] In some embodiments, the model trainer module 304 may update the training data set by creating one or more new historical records that include the new data and retraining the operator model using the updated training data set to further improve the accuracy of the operator model.

[0106] The LOH prediction module 212 may use the trained model to predict LOH in a cell population using the AI / DL module 210. More specifically, the LOH prediction module 212 may use output from the trained model to predict LOH in a cell population. The predicted LOH and other data may be viewed via the data management app 112.

[0107] The app module 214 is configured to facilitate maintaining the data management app 112 and providing its functionality to users. The app module 214 may store instructions that enable the download and / or execution of the data management app 112 on the user computing device 130. The app module 214 may store instructions regarding user interfaces, controls, commands, settings, etc., and may format data into a format suitable for transmission to and display on the user computing device 130.

[0108] In some embodiments, processor 202 is operatively coupled to communications interface 206, enabling LOH prediction computing device 110 to communicate with remote device(s), such as user computing device 130, third party device 140, and / or LOH prediction server 150 (all shown in FIG. 1 ), via wired or wireless connections. For example, communications interface 206 may receive genetic data, etc., from user computing device 130 and / or third party device 140. Communications interface 206 may include, for example, a wired or wireless network adapter and / or a wireless data transceiver for use in a mobile telecommunications network.

[0109] The processor 202 may be operatively coupled to the database 120 (and / or other storage devices) via the storage interface 208. The database 120 may be any computer operating hardware suitable for storing and / or retrieving data. In some embodiments, the database 120 may be implemented on the LOH prediction computer device 110. For example, the LOH prediction computer device 110 may include one or more hard disk drives as the database 120. In other embodiments, the database 120 is external to the LOH prediction computer device 110 and accessed by multiple computer devices. For example, the database 120 may include multiple storage units, such as a storage area network (SAN), a network-attached storage (NAS) system, a redundant array of independent disks (RAID) configuration of hard disks and / or solid-state disks, a cloud storage device, and / or other suitable storage devices.

[0110] Storage interface 208 is any element capable of providing processor 202 with access to database 120. Storage interface 208 may include, for example, an Advanced Technology Attachment (ATA) adapter, a Serial ATA (SATA) adapter, a Small Computer System Interface (SCSI) adapter, a RAID controller, a SAN adapter, a network adapter, and / or any other element that provides processor 202 with access to database 120.

[0111] Processor 202 may execute computer-executable instructions to carry out aspects of the present disclosure. In some embodiments, processor 202 may be converted into a special-purpose microprocessor by executing computer-executable instructions or being otherwise programmed. For example, processor 202 may be programmed with instructions.

[0112] Memory 204 may include, but is not limited to, random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), and non-volatile RAM (NVRAM). The above memory types are exemplary only and are not intended to limit the types of memory that may be used to store computer programs.

[0113] Figure 4 shows a chromosomal region of interest (LOHR) for zygosity according to an exemplary embodiment. In CRISPR editing-induced copy-neutral loss of heterozygosity (LOH), the LOHR is the DNA segment from the CRISPR editing site to the end of the telomere of the same chromosome arm. In different cells, the LOHR can be wild-type (WT) or LOH. The prevalence of CRISPR-induced copy-neutral LOH is about 5% to about 16%, but a high-throughput method that can accurately assess the prevalence of LOH in CRISPR-treated cell populations is needed.

[0114] Generating a model for predicting LOH can be complicated. A model that predicts LOH based on one dataset containing a first set of cell types, CRISPR editing sites, and 10x Genomics data cannot be applied to another dataset consisting of a different set of cell types, CRISPR editing sites, and 10x Genomics data. In some cases, pure LOH cells may not be available to train a model that predicts individual cell LOH.

[0115] Additionally, the coverage of heterozygous mutation positions within LOHRs covered by unique molecular identifiers (UMIs) is typically narrow. Figure 5 shows a chromosomal region of the wild-type DNA strand with a desired zygosity (LOHR) and further illustrates how few UMIs obtained from single-cell RNA sequencing map to heterozygous positions in LOHRs. Of the approximately 4,000 to 12,000 mRNA molecules sequenced using single-cell RNA sequencing, only a small fraction can map to LOHRs. While the sequence length of each sequenced mRNA molecule can be approximately 80 nucleotides, single-nucleotide polymorphisms were found to comprise approximately 0.1% of the human genome. Therefore, single-cell RNA sequencing can present various challenges, such as considering sequencing error. For example, it can be difficult to determine the minimum UMI coverage requirement for a single mutation position to be included in zygosity assessment. Furthermore, low overall UMI coverage can prevent assessment of zygosity in some cells, complicating the detection of a UMI coverage threshold. Generating a single LOHR zygosity measurement for cells from multiple mutation positions also poses challenges.

[0116] The disclosed systems and methods can detect the prevalence of LOH in a cell population by inferring LOHR zygosity at the single-cell level with greater than 99% sensitivity and specificity. In various embodiments, the systems and methods use a high-throughput single-cell RNA sequencing-based process that does not require cell cloning. These systems and methods can also be used independently of cell type, CRISPR edit site, and the 10x Genomics single-cell RNA sequencing platform. These systems and methods use UMI instead of traditional read count-based zygosity determination. These systems and methods can utilize machine learning-based LOH prediction and mLOH cell data for model training obtained from wild-type cells instead of purified LOH cells. Prediction model cases can be selected to balance model performance on individual cell predictions and minimize sampling error for estimating LOH prevalence in a cell population.

[0117] In some aspects, the present disclosure provides systems and methods whereby a mutation position is included in HetRef if the mutation position is covered by at least about 20 DNA sequencing reads and / or if each of two mutations covers about 20% or more of the DNA sequencing reads.

[0118] In other embodiments, HomRef can be generated by establishing a set of homozygous positions that are a specific nucleotide distance away from the variant positions registered in HetRef2. UMIs can be generated from single-cell RNA-sequenced cells of a second cell population (WT) mapped to HetRef2 or HomRef, which can be used as training and validation datasets for WT or mLOH cells, respectively, in training a logistic model to predict loss of heterozygosity. A predictive model using coverage and zygosity scores can be used to predict loss of heterozygosity. In various embodiments, suitable predictive models can include logistic regression models, tree-based methods, neural networks, support vector machines, k-nearest neighbor methods, or other conventional statistical methods. In one embodiment, a logistic regression model is used.

[0119] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing, particular methods and materials are described herein.

[0120] The term "a" refers to "at least one," and the terms "about" and "approximately" allow for standard variations that one of ordinary skill in the art would appreciate, and when ranges are expressed, the endpoints are inclusive. As used herein, the terms "include," "includes," and "including" are meant to be open-ended and refer to "comprise," "comprises," and "comprising," respectively.

[0121] The term "DNA segment" refers to a length of genomic DNA that can be defined by its ability to be sequenced using a DNA probe. The term "bulk DNA sequencing" or "bulk DNAseq" refers to sequencing DNA obtained from a population of cells.

[0122] The term "genome" refers to the DNA or RNA genetic material present in a cell, including the cell's chromosomal / nuclear genome and mitochondrial and / or plasmid genome. The nuclear genome includes protein-coding genes, non-coding genes, and other functional regions such as non-coding DNA or RNA, and, if present, junk DNA or RNA. In some embodiments, the genome is contained within a cell. In some embodiments, the genome is contained in a cell from an established cell line (e.g., 293T cells) or a primary cell cultured ex vivo (e.g., cells obtained from a subject and expanded in culture). In some embodiments, the genome is contained in a hematopoietic cell (e.g., a hematopoietic stem cell, white blood cell, or platelet), or a cell from solid tissue, e.g., a liver cell, kidney cell, lung cell, heart cell, bone cell, skin cell, brain cell, or other cell found in a subject. In some embodiments, the genome or a cell containing the genome is within a subject. Genome-containing subjects of the present disclosure include, but are not limited to, humans and / or other primates, mammals, such as, but not limited to, cattle, pigs, horses, sheep, cats, dogs, mice, and / or rats, and / or birds, including commercially relevant birds such as chickens, ducks, geese, and / or turkeys.

[0123] The term "protein-coding gene" refers to the sequence of a nucleic acid molecule (RNA or DNA molecule) that comprises a nucleotide sequence that encodes a protein. The coding sequence can further comprise initiation and termination signals operably linked to regulatory elements, including a promoter and polyadenylation signal, capable of directing expression in the cells of an individual or mammal to which the nucleic acid is administered. The coding sequence can be codon-optimized.

[0124] The term "non-coding" refers to nucleotide sequences that do not encode protein sequences within a cell. Non-coding DNA can be transcribed into functional non-coding RNA molecules, such as transfer RNA, microRNA, piRNA, ribosomal RNA, and regulatory RNA. Functional regions of non-coding DNA include regulatory nucleotide sequences that control gene expression, scaffold attachment regions, DNA replication origins, centromeres, and telomeres. Non-coding DNA regions, such as introns, pseudogenes, intergenic DNA, transposons, and viral fragments, may appear to have little or no function.

[0125] The term "single-cell RNA sequencing" or "scRNAseq" refers to a method for detecting and quantitatively analyzing messenger RNA molecules within individual cells. The term "chromosomal coordinate" refers to a location within a reference genome of a particular chromosome, or a range of locations, in the latter case including a start and end location.

[0126] The term "nucleotide composition" refers to the DNA or RNA nucleotide(s) found at a chromosomal coordinate. The term "genetic data" refers to information about the genetic material present in a cell, and may include information about the sequence of nucleotide bases that make up DNA or RNA, as well as the location and expression of genes.

[0127] The term "loss of heterozygosity" or "LOH" refers to the substitution of one DNA allele for another at a chromosomal region in a diploid organism. The term "mock loss of heterozygosity" or "mLOH" refers to single-cell RNA sequencing data of loss of heterozygosity digitally generated from wild-type cells by calculating the coverage and zygosity score of homozygous positions within a homozygous reference using Equation 1 and Equation 2, respectively.

[0128] The term "CRISPR" refers to an enzyme system comprising a guide RNA sequence, which comprises a nucleotide sequence complementary or substantially complementary to a target polynucleotide region, and a protein having nuclease activity. CRISPR-Cas systems include Type I, Type II, or Type III CRISPR-Cas systems and their derivative CRISPR-Cas systems, as well as engineered and / or programmed nuclease systems derived from naturally occurring CRISPR-Cas systems. CRISPR-Cas systems can include engineered and / or mutated Cas proteins. CRISPR-Cas systems can include engineered and / or programmed guide RNAs.

[0129] As used herein, the term "guide RNA" refers to an RNA that comprises a sequence that is complementary or substantially complementary to a region of a target DNA sequence. A guide RNA can comprise a nucleotide sequence that is not complementary or substantially complementary to a region of a target DNA sequence. A guide RNA can be a crRNA or a derivative thereof, such as a crRNA:tracrRNA chimera.

[0130] As used herein, the term "nuclease" refers to an enzyme capable of cleaving phosphodiester bonds between nucleotide subunits of nucleic acids. The term "endonuclease" refers to an enzyme capable of cleaving phosphodiester bonds within a polynucleotide chain. The term "nickase" refers to an endonuclease that cleaves only one strand of a DNA duplex. The term "Cas9 nickase" generally refers to a nickase derived from a Cas9 protein by inactivating one nuclease domain of the Cas9 protein.

[0131] The term "loss of heterozygosity region" or "LOHR" refers to a chromosomal region with a desired zygosity. With respect to CRISPR-induced copy number neutral LOH, LOHR can be the DNA segment from the CRISPR editing site to the end of the telomere on the same chromosome arm.

[0132] The terms "major allele" or "M a " refers to nucleotides detected quite frequently in scRNAseq at the variant position, or alleles with fairly high UMI coverage.

[0133] The term "minor allele" or "Mi" refers to an allele containing a nucleotide that is not frequently detected in scRNAseq at the variant position, or low UMI coverage.

[0134] The term "wild-type" or "wt" refers to a cell in which multiple copies of the same chromosome have a unique nucleotide sequence. For example, in a diploid cell, both copies of a chromosome have a unique nucleotide sequence. The term "wild-type" or "wt" also refers to a non-genome-edited, non-mutated cell, which serves as a standard for comparison to genome-edited cells.

[0135] A "unique molecular identifier" or "UMI" is a barcode used in single-cell RNA sequencing to refer to the specific mRNA molecule being sequenced. The number of UMIs mapped to a particular location is the same as the number of sequenced mRNA molecules covering the same location. UMI barcodes are added to nucleic acids in a nucleic acid library prior to sequencing. A specific UMI can, for example, help identify sequencing reads originating from the same specific cell.

[0136] The term "coverage" refers to the number of UMIs covering a particular chromosomal coordinate in HetRef2 or HomRef, or in the case of assessing LOHR zygosity of a cell, the sum of the number of UMIs covering all chromosomal coordinates registered in the LOHR of HetRef2 or HomRef in the cell. Equation 1 can be used by processor 202 to calculate the UMI coverage of the LOHR in a cell.

[0137]

number

[0138] In the formula, ΣM a and ΣM i is the sum of the number of UMIs covering major and minor alleles across all LOHR mutation positions registered in HetRef2 or HomRef, respectively.

[0139] The term "zygosity score" refers to the ratio of the sum of the number of UMIs covering the minor allele (Mi) to the total coverage calculated by Equation 1. Equation 2 can be used by the processor 202 to calculate the zygosity score of a cell.

[0140]

number

[0141] The term "zygosity score threshold" refers to a zygosity score derived from a model in which zygosity scores above the threshold predict heterozygosity and zygosity scores below the threshold predict homozygosity or LOH.

[0142] The term "read" refers to a sequence of nucleotides corresponding to all or a portion of a molecule comprising nucleotides generated using nucleic acid sequencing techniques, such as, for example, DNA sequencing, RNA sequencing, single-cell DNA sequencing, and single-cell RNA sequencing.

[0143] The term "heterozygous mutation location with unbalanced allele expression" or "heterozygous variant with unbalanced allele expression" refers to a heterozygous mutation location that is covered by at least about 100 UMIs and in which at least one of the two alleles appears in at least about 80% of all UMIs, which can be detected using single-cell RNA sequencing data from wild-type cells to generate pseudo-bulk gene expression data. These mutations can also be detected using actual bulk RNA sequencing.

[0144] The term "10x Genomics" refers to a single-cell RNA sequencing technology in which the Chromium Single Cell 3' solution uses microfluidic partitioning to capture single cells and prepare barcoded next-generation sequencing (NGS) cDNA libraries, enabling transcriptome analysis at the cell level. Single cells, reverse transcription (RT) reagents, gel beads containing barcoded oligonucleotides, and oil are combined on a microfluidic chip to form reaction vesicles called gel beads in emulsion, or GEMs. GEMs are formed in parallel within the chip's microfluidic channels, allowing users to process anywhere from hundreds to tens of thousands of single cells in a 7-minute Chromium instrument run. Cells are loaded at limiting dilution to maximize the number of GEMs containing single cells, maintaining low duplicate rates while maintaining high cell recovery rates of up to approximately 65%.

[0145] Each functional GEM contains a single cell, a single gel bead, and reverse transcription reagents. Within each gel bead in the emulsion reaction vesicle, the single cell is lysed, and the gel bead is dissolved, releasing reverse transcription oligonucleotides with identical barcodes into solution, allowing reverse transcription of polyadenylated mRNA. As a result, all cDNAs derived from a single cell will have the same barcode, allowing sequencing reads to be mapped to their original single cell of origin. Next-generation sequencing libraries from these barcoded cDNAs are then prepared in a highly efficient bulk reaction.

[0146] As used herein, "heterozygous variant reference" or "HetRef" refers to a first heterozygous variant reference, which includes the location of the heterozygous variant, the chromosomal coordinate of the heterozygous variant, and the nucleotide composition, and can be identified using whole genome bulk DNA sequencing of wild-type cells of LOHR. HetRef may not include haplotype information. HetRef can be generated using whole genome bulk DNA sequencing of wild-type cells. If the chromosomal coordinate covers a range of at least about 20 reads, and the relevant allele is represented by between about 20% and about 80% of the reads generated using whole genome bulk DNA sequencing, the chromosomal coordinate can be determined to be a heterozygous variant position.

[0147] As used herein, "second heterozygous variant reference" or "HetRef2" refers to a heterozygous variant reference created by removing heterozygous variant positions from HetRef where there may be an imbalance in allelic expression.

[0148] Processor 202 can collect impurity information from all eligible heterozygous mutation positions of HetRef2 within the LOHR and form artificial compound variants using Equation 1 (e.g., the sum of the number of UMIs covering frequently detected nucleotides at each mutation position and the number of UMIs covering less frequently detected nucleotides at each mutation position) to determine UMI coverage. HetRef2 heterozygous mutation positions may be included that are covered by at least about four UMIs at each position. Equation 2 can be used by processor 202 to determine a zygosity score, which is equal to the weighted average percentage of coverage across all eligible mutation positions, and the zygosity score can be used by processor 202 to construct a simple logistic regression model.

[0149] In one embodiment, a probability-based method can be used to assess zygosity by processor 202. Assuming that both alleles are transcribed at the same rate, Equation 3 can be used by processor 202 to calculate the probability of finding that r or more of n UMIs covering a chromosomal coordinate that is likely a heterozygous mutation location contain the same nucleotide.

[0150]

number

[0151] where n is the number of UMIs covering the chromosomal coordinate that is likely the mutation location, and r is the number of UMIs covering the commonly observed nucleotide at the mutation location. For example, assuming both alleles transcribe at the same rate, the probability of finding 9 or more UMIs with the same nucleotide (r=9) at a chromosomal coordinate that is a heterozygous mutation location with coverage of 10 UMIs (n=10) is 2.1484%. Equation 4 can be used by processor 202 to calculate the probability that all UMIs covering a chromosomal coordinate that is a heterozygous mutation location will contain the same nucleotide.

[0152]

number

[0153] And Table 1 shows the probability that all UMIs covering chromosomal coordinates that are heterozygous mutation positions consisting of adenine (A) and thymine (T) are found to contain the same nucleotide, A or T, when n = r and r = 2, 3, 4, 5, or 6.

[0154] [Table 1]

[0155] Probability-based methods are sensitive to coverage (n): when coverage is low, e.g., when there are fewer than four UMIs, there is a high chance (≥25%) that all UMIs will have the same nucleotide, even at truly heterozygous mutation positions.

[0156] Alternatively, nucleotide connectivity can be assessed by the processor 202 using impurity-based methods such as the Gini Index, Entropy, and Percentage. Equation 5 can be used by the processor 202 to calculate the Gini Index.

[0157] Equation 5: Gini Index:1-p 2 -(1-p) 2 where p is the UMI ratio of one of the two nucleotides covering the variant position. Equation 6 can be used by the processor 202 to calculate Entropy.

[0158] Equation 6: Entropy:-p * log2p-(1-p) * log2(1-p) where p is the UMI ratio of one of the two nucleotides covering the variant position.

[0159] Equation 7 can be used by processor 202 to calculate Percentage. Equation 7: Percentage:min(p,1-p) where p is the UMI ratio of the two nucleotides covering the variant position. When p≦0.5, all three impurity calculation methods are monotonic functions, with larger p values ​​corresponding to higher impurity assessment values. When processor 202 uses a rank-based method (such as logistic regression or decision tree) for the subsequent LOH prediction model, the prediction accuracy is the same whether using the Gini Index, Entropy, or Percentage Impurity-based method.

[0160] As used herein, a "homozygous reference" or "HomRef" refers to a reference that contains homozygous coordinates within ±100 nucleotides of each heterozygous variant coordinate registered in HetRef2. In some embodiments, HomRef is used to generate mLOH cells by calculating zygosity scores and coverage of HomRef positions. In some embodiments, HomRef is used to generate mLOH cells from wild-type cells for training an LOH model. mLOH cells can be digitally generated from wild-type cells, and their zygosity scores can be calculated using UMIs that cover homozygous variant reference positions within the LOHR of the wild-type cells. The zygosity scores of mLOH cells can be used together with the scores of wild-type cells to train an LOH prediction model.

[0161] The processor 202 collects impurity information from all eligible heterozygous mutation positions of HetRef2 in the LOHR of wild-type cells and other cells to form artificial compound mutants with a UMI coverage equal to the sum of the number of UMIs of nucleotides frequently detected at each mutation position, i.e., EA / a, and the number of UMIs of nucleotides infrequently detected at each mutation position, i.e., EM, as shown in Equation 1, and then, according to an exemplary embodiment, can determine the minimum coverage to be considered accessible to cell zygosity.

[0162] The processor 202 should apply a minimum UMI coverage requirement to select wild-type cells and other cells of interest for model training and model-independent validation, as well as to select cells from test samples for LOH prevalence detection. The evaluated chromosomal coordinates may be homozygous, but the zygosity score cannot be 0 in LOH or mLOH cells for various reasons, including, but not limited to, sequence errors, the homozygous position registered in HomRef and utilized is not actually homozygous in mLOH cells, or the LOH cells do not completely lose heterozygosity at LOHR. The minimum UMI coverage requirement is different from the UMI coverage threshold used to qualify a mutant position for inclusion in a cell.

[0163] Equation 8 can be used by the processor 202 to calculate the number of cells required to assess zygosity to establish the prevalence of LOH in a cell population.

[0164]

number

[0165] where p is the expected prevalence, d is the precision (d), and Z is the Z-value for the desired confidence level. In various embodiments of the systems and methods of the present disclosure, the target cell population can include cells that have been edited using genome editing tools such as CRISPR. In some embodiments, the target cell population can include cells that have lost heterozygosity through one or more genome editing procedures.

[0166] In various embodiments, the reference cell population can include cells that have not been treated with CRISPR, i.e., WT cells, cells with the same DNA genotype as WT cells, and / or cells heterozygous for a DNA segment of interest that can be used to generate HetRef.

[0167] In various aspects, the reference data may include HetRef, HetRef2, and / or HomRef. The present disclosure may be further understood with reference to the following examples, which should not be construed as limiting the scope of the disclosure. [Example]

[0168] method Three cell populations were sequenced. All three cell populations were derived from the same subject. The first cell population was used to perform bulk DNA sequencing and obtain a heterozygous reference (HetRef). The first cell population could be any haplotype cell from the subject. The second (training sample) and third (test sample) were used for scRNAseq. The second and third cell populations were of the same cell type(s). Because the second population was not involved in the CRISPR procedure, it was considered to be wild-type cells, whereas the third population was involved in the CRISPR procedure. In the case of LOH studies in cancer, the second population was a non-tumor sample, and the third population was a tumor sample from the same tissue / organ as the non-tumor sample. Using the scRNA-seq results from the second population, we i) identified positions within HetRef where the two alleles showed unbalanced expression, which were then deleted from HetRef to generate HetRef2, and then generated HomRef (a set of DNA coordinates a specific number of nucleotides away from the position registered in HetRef2) from HetRef2; ii) mapped the UMIs to HetRef2 to generate WT cell profiles (both coverage and zygosity scores); iii) mapped the UMIs to HomRef to generate mLOH cell profiles (both coverage and zygosity scores); and iv) used the WT and mLOH zygosity scores (which can generate multiple prediction model cases based on different coverage requirements) to train a zygosity prediction model (to predict the zygosity of individual cells).The UMIs generated from scRNA-seq of the third population were mapped to HetRef2 to generate profiles (both coverage and zygosity scores) for each cell. These were then used to: i) select predicted model cases generated from the second population by providing the number of cells in the third population that met the coverage requirements for each predicted model case (ideally, we select model cases that are sufficiently fit to be expected to have high prediction accuracy and minimize population sampling error, and also meet the coverage requirements); ii) use the selected model cases to validate the fit of each cell in the third population for predicted zygosity using their zygosity scores (fitness is based on coverage); and iii) calculate the percentage of LOH among eligible cells in the third population.

[0169] All cells were derived from human induced pluripotent stem cells (iPSCs). Clonal loss-of-heterozygosity cells were edited using CRISPR genome editing, targeting coordinate 31,191,431 (C → T) on chromosome 16 based on the GRCh38 / hg38 assembly. First, the zygosity scores of the loss-of-heterozygosity region were determined for each wild-type cell and each loss-of-heterozygosity cell. Next, a simple logistic regression model was trained to convert the cell zygosity scores into genotype probabilities, and cross-validated using 70% of the zygosity score data from wild-type and loss-of-heterozygosity cells. Next, the sensitivity, specificity, and accuracy of the simple logistic regression model were independently validated using 30% of the zygosity score data from wild-type and loss-of-heterozygosity cells. Finally, the prevalence of loss of heterozygosity was examined by assessing the zygosity scores of individual cells within the test sample using the simple logistic regression model.

[0170] Wild-type cells were assayed using whole-genome bulk DNA sequencing. If a chromosomal coordinate is covered by at least about 20 reads and the proportion of one allele is between about 20% and about 80% of the reads in the whole-genome bulk DNA sequencing dataset, the chromosomal coordinate was determined to be a heterozygous mutation position. The heterozygosity of each coordinate can also be established by existing algorithms developed for DNA sequencing data. According to an exemplary embodiment, bulk DNA sequencing identified 37,902 mutation positions on chromosome 16 of wild-type cells, and 15,465 mutation positions within the heterozygous loss region of chromosome 16.

[0171] Three test samples, including wild-type cells, clonal loss of heterozygosity cells, and wild-type cells with 3%, 10%, or 30% loss of heterozygosity, were assayed using single-cell RNA sequencing. Table 2 shows the number of wild-type cells, loss of heterozygosity cells, and cells in the three test samples assayed using single-cell RNA sequencing.

[0172] [Table 2]

[0173] The first heterozygous mutant reference was generated, containing the chromosomal coordinates and nucleotide composition of 37,902 heterozygous mutation locations identified on chromosome 16 in wild-type cells. 15,465 identified mutation locations were located in loss-of-heterozygosity regions on chromosome 16. To establish the feasibility of using RNA sequencing data instead of DNA sequencing data for zygosity determination, we treated single-cell RNA sequencing data generated from the second or third cell population as pseudo-bulk by ignoring cell identity information in the data and using UMIs instead of reads as coverage for zygosity determination. A chromosomal coordinate was determined to be at a heterozygous mutation location if it was covered by at least approximately 20 unique molecular identifiers, the nucleotide composition obtained using single-cell RNA sequencing was the same as the nucleotide composition identified using whole-genome bulk DNA sequencing, and the proportion of one allele was represented between approximately 20% and approximately 80% of all unique molecular identifiers.

[0174] result 6 shows the distribution of heterozygous mutation positions identified on chromosome 16 in wild-type (top) or LOH (bottom) cells using single-cell RNA sequencing and the first heterozygous mutant reference as a reference, according to an exemplary embodiment. Significantly more heterozygous positions exist in LOHR wild-type cells than in LOH cells, confirming the possibility of using RNA sequencing results generated from single-cell experiments to imply DNA zygosity. Possible reasons for the heterozygous mutation positions observed in LOH cells include, for example, (i) LOH cells did not lose heterozygosity at all heterozygous mutation positions in LOHR, or (ii) sequencing errors in single-cell RNA sequencing.

[0175] Using the first heterozygous mutant reference as a reference, potential allelic expression imbalances between the first heterozygous mutant reference positions were identified. A heterozygous mutant position in the first heterozygous mutant reference was determined to have potential allelic expression imbalances if the heterozygous mutant position was covered by at least approximately 100 unique molecular identifiers in the pseudobulk RNA sequencing data of wild-type cells (second population cells) and if one of the two alleles was represented by at least approximately 80% of all unique molecular identifiers. A total of 456 heterozygous mutant positions in the first heterozygous mutant reference were found to have allelic expression imbalances, of which 276 were within the region of loss of heterozygosity and 180 were outside the region of loss of heterozygosity.

[0176] A second heterozygous mutant reference was generated by deleting heterozygous mutant positions in the first heterozygous mutant reference where allele expression imbalances might exist. A total of 37,446 heterozygous mutation positions were identified on chromosome 16 in wild-type cells in the second heterozygous mutant reference, of which 15,189 were within the region of loss of heterozygosity and 22,257 were outside the region of loss of heterozygosity.

[0177] Figure 7 shows that, according to exemplary embodiments, coverage of the heterozygous mutation positions of the second heterozygous mutant reference by unique molecular identifiers was low in the region of heterozygosity loss on chromosome 16 of wild-type cells. The top panel of Figure 7 shows that, according to exemplary embodiments, the total number of unique molecular identifiers covering the heterozygous mutation positions of the second heterozygous mutant reference in the region of heterozygosity loss of each wild-type cell was low. The middle panel of Figure 7 shows that, according to exemplary embodiments, the number of heterozygous mutation positions of the second heterozygous mutant reference covered by unique molecular identifiers in the region of heterozygosity loss of each wild-type cell was low. The bottom panel of Figure 7 shows that, according to exemplary embodiments, coverage of the heterozygous mutation positions of the second heterozygous mutant reference by unique molecular identifiers was low in the region of heterozygosity loss of all wild-type cells.

[0178] Equation 3 is a probability test used to understand the minimum unique molecular identifier coverage required for a heterozygous mutation position in a second heterozygous mutant reference to be included in the zygosity assessment of a cell. The null hypothesis in Equation 3 was that the mutation position being assessed was heterozygous. Specifically, Equation 3 derived the probability of finding that at least r of the n unique molecular identifiers covering the mutation position contain the same nucleotide, given the assumption that the position being assessed is heterozygous and both alleles are transcribed at the same rate.

[0179] For example, if both alleles are transcribed at the same rate, at a heterozygous mutation position covered by four unique molecular identifiers (n=4), there is a 50% probability that three or more unique molecular identifiers with the same nucleotide (r=3) will be observed. Equation 4 derives the probability that all unique molecular identifiers covering a mutation position will contain the same nucleotide if both alleles are transcribed at the same rate.

[0180] 8 illustrates how zygosity information from all eligible heterozygous mutation positions of the second heterozygous variant reference at the loss of heterozygosity region was collected to obtain an overall zygosity score and unique molecular identifier coverage for each cell; however, according to an exemplary embodiment, a zygosity score=0 was not necessary or sufficient to determine that a cell is loss of heterozygosity, since sequencing errors, mapping errors, or both may prevent a complete determination of loss of heterozygosity or an overall low UMI coverage in LOHR. The sum of the number of unique molecular identifiers of nucleotides frequently found at each mutation position, Equation 1, i.e., ΣM a , and the sum of the number of unique molecular identifiers of infrequently detected nucleotides at each mutation position, i.e., ΣM iを The coverage of each cell was calculated using the following formula: Heterozygous mutation positions of the second heterozygous mutant reference that were covered by at least four unique molecular identifiers were included.

[0181] Figure 9 shows that, in the region of loss of heterozygosity of wild-type cells, heterozygous mutation positions of the second heterozygous mutant reference were covered by four or more unique molecular identifiers, according to an exemplary embodiment. The top panel of Figure 11 shows the total number of unique molecular identifiers covered by at least four unique molecular identifiers per position in the region of loss of heterozygosity of chromosome 16 of wild-type cells, according to an exemplary embodiment. The bottom panel of Figure 9 shows the number of heterozygous mutation positions of the second heterozygous mutant reference that were covered by at least four unique molecular identifiers in the region of loss of heterozygosity of chromosome 16 of wild-type cells, according to an exemplary embodiment. The coverage of each cell, which can be determined using Equation 1, was used to determine whether the zygosity of the cell was assessable.

[0182] 10 illustrates the generation of a homozygous reference containing homozygous chromosome coordinates by adding or subtracting approximately 100 to each heterozygous mutant chromosome coordinate in a second heterozygous mutant reference of a wild-type cell, according to an exemplary embodiment. The mock loss-of-heterozygosity cells were digitally generated from the wild-type cells by calculating the coverage and zygosity score of the homozygous positions in the homozygous reference using Equation 1 and Equation 2, respectively. The zygosity scores of the wild-type and mock loss-of-heterozygosity cells were equal to the coverage-weighted average of Mi / (Mi+Ma) covering all eligible heterozygous mutation positions, and were used to train a simple logistic regression model to predict heterozygous loss.

[0183] Figure 11 shows how impurity information is collected from all eligible heterozygous mutation positions in the loss of heterozygosity region of wild-type cells, loss of heterozygosity cells, and the second heterozygous mutant reference in mock loss of heterozygosity cells to calculate the total number of unique molecular identifiers of nucleotides frequently detected at each mutation position, i.e., ΣM a , and the number of unique molecular identifiers of infrequently detected nucleotides at each mutation position, i.e., ΣM i This illustrates creating an artificial compound mutant with a unique molecular identifier coverage equal to the sum of the number of unique molecular identifiers (as shown in Equation 1) and then determining the minimum coverage for a cellular zygosity to be considered accessible, according to an exemplary embodiment.

[0184] 12 shows the sum of unique molecular identifiers at loss of heterozygosity regions on chromosome 16 in wild-type cells where there are at least four unique molecular identifiers for each mutation position, according to an exemplary embodiment. The minimum unique molecular identifier coverage requirement is the sum of unique molecular identifier coverage across all positions registered in the second heterozygous variant reference with at least four unique molecular coverage per position, and should apply to the selection of wild-type and pseudo-loss of heterozygosity cells for model training, the selection of wild-type, pseudo-loss of heterozygosity, and actual loss of heterozygosity cells in model-independent validation, and the selection of cells from test samples for detecting the prevalence of loss of heterozygosity.

[0185] The accuracy of the model for detecting loss of heterozygosity in individual cells and sampling error in testing test samples can affect the accuracy of the model for detecting the prevalence of loss of heterozygosity. Figure 13 illustrates, according to exemplary embodiments, that the coverage requirement for a minimum number of unique molecular identifiers can have an inverse effect on the accuracy of the model and sampling error. Figure 13 also illustrates, according to exemplary embodiments, that as the minimum unique molecular identifier coverage requirement increases, the accuracy of the model for predicting loss of heterozygosity in individual cells can increase, thereby improving the accuracy of the model for predicting the prevalence of loss of heterozygosity.

[0186] However, Figure 13 also illustrates, in accordance with an exemplary embodiment, that as the number of test cells that meet the minimum unique molecular identifier coverage requirement decreases and the minimum unique molecular identifier coverage requirement increases, the model's sampling error increases, which may reduce the model's accuracy in predicting the prevalence of loss of heterozygosity. Figure 14 illustrates how, in accordance with an exemplary embodiment, model accuracy can be maximized and sampling error minimized by dynamically determining the coverage requirement for determining that a cell is assessable for zygosity.

[0187] Table 3 shows the sample size calculations for the resulting study using Equation 8 and the assumed prevalence (p), precision (d), and 90% confidence level (Z). The data in Table 3 indicate an accuracy of approximately 20% to approximately 25% for the assumed prevalence (heterozygosity rate). It was concluded that an initial estimate of loss of heterozygosity prevalence requires more than 1,000 zygosity-evaluable cells.

[0188] [Table 3]

[0189] Figure 15 illustrates an overview of an exemplary embodiment for generating, testing, and using logistic regression models to predict loss of heterozygosity in individual cells and the prevalence of loss of heterozygosity in cell populations. Figure 16 shows how the zygosity score threshold and percentage of qualified predictions correlate with unique molecular identifier coverage requirements for each row of Figure 22, where each row is a model example trained and validated using cells that met the indicated unique molecular identifier coverage requirement, according to an exemplary embodiment.

[0190] Figure 16 shows that as the coverage requirements for the unique molecular identifiers became more stringent, the zygosity score threshold and the percentage of correct predictions generally increased. The model was trained using 70% of wild-type and mock-loss-of-heterozygosity cells and tested on the remaining 30% of wild-type and mock-loss-of-heterozygosity cells, as well as on the remaining 30% of non-heterozygous cells. The model was generated using Python's CycItLearn with a LogisticRegressionCV classifier. The optimal L2 regularization hyperparameters were selected using stratified five-fold cross-validation. The final model was generated by fitting the entire training set using the optimal regularization hyperparameters from cross-validation. As the coverage requirements became more stringent, the accuracy, sensitivity, and specificity of the model increased. The model trained on wild-type and mock-loss-of-heterozygosity cells was able to accurately predict the zygosity of wild-type, mock-loss-of-heterozygosity, and non-heterozygous cells.

[0191] The specificity of major class predictions can have a significant impact on the overall prediction accuracy in highly imbalanced samples, where members of one class are significantly more abundant than members of other classes. For example, if the true prevalence of heterozygosity is 5% and the model's sensitivity and specificity for predicting heterozygosity are 95% and 100%, respectively, the predicted prevalence of heterozygosity is 4.75%, which is close to 5%. However, if the model's sensitivity and specificity for predicting heterozygosity are 100% and 95%, respectively, the predicted prevalence of heterozygosity is 9.75%, which is significantly different from 5%.

[0192] The top panel of Figure 17 shows that, according to an exemplary embodiment, the WT (wild-type) probability threshold for each model instance for predicting whether a cell is a wild-type cell or a loss-of-heterozygosity cell was set to 0.5. Cells with a WT probability of at least 0.5 were predicted to be wild-type cells. Cells with a WT probability of less than 0.5 were predicted to be loss-of-heterozygosity cells. The bottom panel of Figure 17 shows the zygosity score threshold for each logistic regression predictive model instance generated using various unique molecular identifier coverage requirements, according to an exemplary embodiment.

[0193] To minimize sampling error, each test sample contained at least 1,000 cells. The number of zygosity-evaluable cells that met the minimum requirement of 1,000 cells for Test Sample 1 (TS1), Test Sample 2 (TS2), and Test Sample 3 (TS3) is shown in Figure 21. For Test Sample 1, coverage requirements between 4 UMI and 12 UMI (inclusive) all met the minimum requirement of 1,000 zygosity-evaluable cells. Among the model cases using these coverage requirements, the most accurate model case used 12 UMI as the minimum coverage requirement. For Test Sample 2, coverage requirements between 4 UMI and 25 UMI (inclusive) all met the minimum requirement of 1,000 zygosity-evaluable cells. For Test Sample 3, any coverage requirement between 4 UMI and 30 UMI (inclusive) would satisfy the minimum 1,000 zygosity-evaluable cells requirement. Of the model cases using these requirements for Test Sample 2 or Test Sample 3, the most accurate model case used a minimum coverage requirement of 25 UMI for both Test Samples 2 and 3.

[0194] Figure 18 shows the predicted prevalence of loss of heterozygosity for three test samples, displayed at the top of the top and bottom panels of Figure 18, using the model case for test sample 1 shown in lines 4-24 of Figure 21 and the model case for test samples 2 and 3 shown in line 25 of Figure 21, according to an exemplary embodiment. The top panels of Figure 18 and Figure 21 show that the coverage and model zygosity score thresholds used to test sample 1 were 12 UMI and 14.0%, respectively, according to an exemplary embodiment.

[0195] The top panel of Figure 18 shows that, according to an exemplary embodiment, the predicted prevalence of loss of heterozygosity for test sample 1 using the model example shown therein was 3.70%, and the actual prevalence of loss of heterozygosity for test sample 1 was approximately 3%. The bottom panels of Figure 18 and Figure 21 show that, according to an exemplary embodiment, the coverage and model zygosity score thresholds used to test samples 2 and 3 were 25 UMI and 14.4%, respectively.

[0196] The bottom panel of Figure 18 shows that, in accordance with an exemplary embodiment, the predicted prevalence of loss of heterozygosity for test sample 2 using the model example shown therein was 12.36%, and the actual prevalence of loss of heterozygosity for test sample 2 was approximately 10%. The bottom panel of Figure 18 shows that, in accordance with an exemplary embodiment, the predicted prevalence of loss of heterozygosity for test sample 3 using the model example shown therein was 30.24%, and the actual prevalence of loss of heterozygosity for test sample 3 was approximately 30%.

[0197] FIG. 19 shows an exemplary workflow of the present disclosure, which worked because it 1) excluded variant positions with unbalanced allele expression from zygosity assessment, 2) used unique molecular identifiers instead of traditional read counts for zygosity assessment, 3) used coverage of at least four unique molecular identifiers as a threshold for variant positions within loss-of-heterozygosity regions to be included in zygosity assessment, 4) used cellular zygosity scores as the zygosity measure for building loss-of-heterozygosity models and predicting loss-of-heterozygosity, 6) used mock loss-of-heterozygosity cells derived from wild-type cells by generating zygosity scores and coverage from unique molecular identifiers mapped to homozygous positions within regions of loss-of-heterozygosity, and 7) established minimum coverage requirements for cells to be eligible for training the loss-of-heterozygosity model, where zygosity assessment was determined by generating an accurate loss-of-heterozygosity prediction model that could predict loss of heterozygosity in individual cells, and included enough cells in the test sample to avoid large sampling errors. In the figure, SAM represents a Sequence Alignment / Map format file, BAM represents a binary SAM file, and VCF represents a variant determination file.

[0198] Figure 20 illustrates a method for using an exemplary workflow of the present disclosure to determine whether a heterozygous loss cell has a complete or partial loss of heterozygosity at a heterozygous loss region and the impact of the loss of heterozygosity on downstream gene expression. The top panel of Figure 20 illustrates how an exemplary workflow of the present disclosure analyzes the overlay of pseudo-bulk unique molecular identifiers at the heterozygous mutation position of a second heterozygous mutant reference at the heterozygous loss region of a wild-type cell and a heterozygous loss cell to determine whether the heterozygous loss cell has a complete or partial loss of heterozygosity at the heterozygous loss region. The bottom panel of Figure 20 illustrates how an exemplary workflow of the present disclosure analyzes single-cell RNA sequencing differential gene expression data from a wild-type cell and a heterozygous loss cell to determine the impact of the loss of heterozygosity on downstream gene expression.

Claims

1. 1. A system for predicting the prevalence of loss of heterozygosity (LOH) in a target cell population, comprising: at least one memory storing computer-executable instructions; and at least one processor in communication with said at least one memory, receiving genetic data for a first reference cell population; sequencing the genetic data of the first reference cell population to obtain first reference data, wherein the first reference data comprises a chromosomal identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; Identifying and removing heterozygous mutation positions with unbalanced allele expression in the first reference data based on single-cell RNA sequencing data generated from a second reference cell population to generate second reference data; generating third reference data by establishing a set of homozygous positions that are a predetermined number of nucleotides away from each variant position in the second reference data; mapping an identifier for each cell in the target cell population to the second reference data to generate a first mapped identifier; mapping an identifier for each cell of the second reference cell population to the second reference data and / or the third reference data to generate a second mapped identifier; applying one or more inputs to a supervised machine learning model, the one or more inputs including the first mapped identifier for each cell of the target cell population and the second mapped identifier for each cell of the second reference cell population, the model having been previously trained using historical data, the historical data including mapped identifiers and their corresponding LOH for each cell of the target cell population; receiving one or more outputs from the model, wherein at least one of the one or more outputs comprises LOH for the target cell population; thereby predicting the prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more of the outputs; the at least one processor configured to execute computer-executable instructions to retrain the model using updated historical data.

2. The system of claim 1 , wherein the target cell population comprises genome-edited cells.

3. The system of claim 2 , wherein the genome-edited cells comprise CRISPR-edited cells.

4. The system according to any one of claims 1 to 3, wherein the target cell population comprises cancer cells.

5. The system of any one of claims 1 to 4, wherein the first reference cell population comprises the same genotype as wild-type cells of the first reference cell population.

6. 6. The system of claim 1, wherein at least one of the processors is configured to execute the computer-executable instructions to sequence the genetic data using bulk DNA sequencing.

7. The system of any one of claims 1 to 6, wherein the second reference cell population comprises cells that have not been treated with a genome editing tool.

8. The system of any one of claims 1 to 7, wherein the identifier comprises a UMI.

9. The system of any one of claims 1 to 8, wherein the model comprises a logistic regression model.

10. 1. A computer-implemented method for predicting the prevalence of loss of heterozygosity (LOH) in a target cell population, comprising: at least one memory storing computer-executable instructions; and at least one processor in communication with said at least one memory, receiving genetic data for a first reference cell population; sequencing the genetic data of the first reference cell population to obtain first reference data, wherein the first reference data comprises a chromosomal identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; Identifying and removing heterozygous mutation positions with unbalanced allele expression in the first reference data based on single-cell RNA sequencing data generated from a second reference cell population to generate second reference data; generating third reference data by establishing a set of homozygous positions that are a predetermined number of nucleotides away from each variant position in the second reference data; mapping an identifier for each cell in the target cell population to the second reference data to generate a first mapped identifier; mapping an identifier for each cell of the second reference cell population to the second reference data and / or the third reference data to generate a second mapped identifier; applying one or more inputs to a supervised machine learning model, the one or more inputs including the first mapped identifier for each cell of the target cell population and the second mapped identifier for each cell of the second reference cell population, the model previously trained using historical data, the historical data including mapped identifiers and their corresponding LOH for each cell of the target cell population; receiving one or more outputs from the model, wherein at least one of the one or more outputs comprises LOH for the target cell population; thereby predicting the prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more of the outputs; the at least one processor being configured to execute computer-executable instructions to retrain the model using updated historical data.

11. 11. The computer-implemented method of claim 10, wherein the target cell population comprises genome-edited cells.

12. 12. The computer-implemented method of claim 11, wherein the genome-edited cells comprise CRISPR-edited cells.

13. The computer-implemented method of any one of claims 10 to 12, wherein the target cell population comprises cancer cells.

14. 14. The computer-implemented system of claim 10, wherein the first reference cell population comprises the same genotype as wild-type cells of the first reference cell population.

15. 15. The computer-implemented method of any one of claims 10 to 14, comprising sequencing the genetic data using bulk DNA sequencing.

16. 16. The computer-implemented method of any one of claims 10 to 15, wherein the second reference cell population comprises cells that have not been treated with a genome editing tool.

17. The computer-implemented method of any one of claims 10 to 16, wherein the identifier comprises a UMI.

18. The computer-implemented method of any one of claims 10 to 17, wherein the model comprises a logistic regression model.

19. At least one persistent computer-readable storage medium embodied with computer-executable instructions, which when executed by at least one processor, cause the at least one processor to: receiving genetic data for a first reference cell population; sequencing the genetic data of the first reference cell population to obtain first reference data, wherein the first reference data comprises a chromosomal identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying and removing heterozygous mutation positions with unbalanced allelic expression in the first reference data to generate second reference data; generating third reference data by establishing a set of homozygous positions that are a predetermined number of nucleotides away from each variant position in the second reference data; mapping an identifier for each cell of the target cell population to the second reference data to generate a first mapped identifier; mapping an identifier for each cell of the second reference cell population to the second reference data and / or the third reference data to generate a second mapped identifier; applying one or more inputs to a supervised machine learning model, the one or more inputs including the first mapped identifier for each cell of the target cell population and the second mapped identifier for each cell of the second reference cell population, the model having been previously trained using historical data, the historical data including mapped identifiers and their corresponding LOH for each cell of the target cell population; receiving one or more outputs from the model, wherein at least one of the one or more outputs comprises LOH for the target cell population; thereby predicting the prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more of the outputs; The storage medium retrains the model using updated historical data.

20. 1. A system for predicting the prevalence of loss of heterozygosity (LOH) in a target cell population, comprising: at least one memory storing computer-executable instructions; and at least one processor in communication with said at least one memory, receiving genetic data for a first reference cell population; sequencing the genetic data of the first reference cell population to obtain first reference data, wherein the first reference data comprises a chromosomal identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying and removing heterozygous mutation positions with unbalanced allelic expression in the first reference data to generate second reference data; mapping an identifier for each cell in the target cell population to the second reference data to generate a mapped identifier; applying one or more inputs to a supervised machine learning model, the one or more inputs including the identifiers mapped for each cell in the target cell population, the model previously trained using historical data, the historical data including the identifiers mapped for each cell in the target cell population and their corresponding LOH; receiving one or more outputs from the model, at least one of the one or more outputs comprising LOH of the target cell population; thereby predicting the prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more of the outputs; the at least one processor configured to execute computer-executable instructions to retrain the model using updated historical data.

21. 21. The system of claim 20, wherein the target cell population comprises genome-edited cells.

22. 22. The system of claim 21, wherein the genome-edited cells comprise CRISPR-edited cells.

23. The system of any one of claims 20 to 22, wherein the target cell population comprises cancer cells.

24. The system of any one of claims 20 to 23, wherein the first reference cell population comprises cells of the same genotype as wild-type cells of the first reference cell population.

25. 25. The system of any one of claims 20 to 24, wherein at least one said processor is configured to execute computer-executable instructions to sequence genetic data using bulk DNA sequencing.

26. 26. The system of any one of claims 20-25, wherein at least one processor is configured to execute computer-executable instructions to identify allele expression imbalance heterozygous mutation locations in the first reference data based on single-cell RNA-sequencing data generated from the second reference cell population.

27. 27. The system of Claim 26, wherein the second reference cell population comprises cells that have not been treated with a genome editing tool.

28. The system of any one of claims 20 to 27, wherein the identifier comprises a unique molecular identifier (UMI).

29. The system of any one of claims 20 to 28, wherein the model comprises a logistic regression model.

30. 1. A computer-implemented method for predicting the prevalence of loss of heterozygosity (LOH) in a target cell population, comprising: receiving genetic data for a first reference cell population; sequencing the genetic data to obtain first reference data, wherein the first reference data comprises a chromosome identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying and removing heterozygous mutation positions with unbalanced allelic expression in the first reference data to generate second reference data; mapping an identifier for each cell in the target cell population to the second reference data to generate a mapped identifier; applying one or more inputs to a supervised machine learning model, the one or more inputs including the identifiers mapped for each cell in the target cell population, the model previously trained using historical data, the historical data including the identifiers mapped for each cell in the target cell population and their corresponding LOH; receiving one or more outputs from the model, at least one of the one or more outputs comprising LOH of the target cell population; thereby predicting the prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more of the outputs; retraining the model using updated historical data.

31. 31. The computer-implemented method of claim 30, wherein the target cell population is edited using a genome editing tool.

32. 32. The computer-implemented method of claim 31 , wherein the genome editing tool comprises a CRISPR-based genome editing tool.

33. 33. The computer-implemented method of any one of claims 30 to 32, wherein the target cell population comprises cancer cells.

34. 34. The computer-implemented method of any one of claims 30 to 33, wherein the first reference cell population comprises the same genotype as wild-type cells of the first reference cell population.

35. 35. The computer-implemented method of any one of claims 30 to 34, wherein the method comprises sequencing the genetic data using bulk DNA sequencing.

36. 36. The computer-implemented method of any one of claims 30-35, wherein identifying allele expression imbalance heterozygous mutation positions in the first reference data is based on single-cell RNA sequencing data generated from a second reference cell population.

37. 37. The computer-implemented method of Claim 36, wherein the second reference cell population comprises cells that have not been treated with a genome editing tool.

38. The computer-implemented method of any one of claims 30 to 37, wherein the identifier comprises a UMI.

39. 39. The computer-implemented method of any one of claims 30 to 38, wherein the model comprises a logistic regression model.

40. At least one persistent computer-readable storage medium embodied with computer-executable instructions, which when executed by at least one processor, cause the at least one processor to: receiving genetic data for a first reference cell population; sequencing the genetic data to obtain first reference data, wherein the first reference data comprises a chromosome identifier, a nucleotide coordinate, a nucleotide composition, or a combination thereof; identifying and removing heterozygous mutation positions with unbalanced allelic expression in the first reference data to generate second reference data; mapping an identifier for each cell in the target cell population to the second reference data to generate a mapped identifier; applying one or more inputs to a supervised machine learning model, the one or more inputs including the identifiers mapped for each cell in the target cell population, the model previously trained using historical data, the historical data including the identifiers mapped for each cell in the target cell population and their corresponding LOH; receiving one or more outputs from the model, at least one of the one or more outputs comprising LOH of the target cell population; thereby predicting the prevalence of LOH in the target cell population; updating the historical data to include the genetic data of the target cell population and the corresponding one or more of the outputs; The storage medium retrains the model using updated historical data.