Method for analyzing loss of heterozygosity (LoH) after whole genome amplification of defined restricted sites (DRS-WGA)
The method addresses the limitations of current LoH analysis techniques by using DRS-WGA for single-cell resolution, allowing for efficient LoH and copy number variant analysis from low-pass sequencing data without the need for high-coverage sequencing or normal controls.
Patent Information
- Application Number
- JP2022506443
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-30
- Filing Date
- 2020-07-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-07-29
AI Technical Summary
Current methods for analyzing loss of heterozygosity (LoH) in single cells from low-pass whole-genome sequencing data are limited by the need for high-coverage sequencing, the requirement of a normal control, and the inability to reliably re-analyze single cells for validation or additional genomic information.
A method that uses deterministic restriction site-based whole-genome amplification (DRS-WGA) to achieve single-cell resolution for LoH analysis, allowing for the inference of whole-genome LoH status without the need for high-coverage sequencing or normal controls, and enabling re-analysis of WGA products for additional information.
This method enables efficient analysis of LoH and copy number variants from low-pass sequencing data, providing single-cell resolution and reducing the need for extensive sequencing and controls, thereby enhancing the reliability and cost-effectiveness of LoH analysis.
Smart Images

Figure 0007684277000006 
Figure 0007684277000007 
Figure 0007684277000008
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This patent application claims priority to Italian Patent Application No. 102019000013335 filed on July 30, 2019, the entire disclosure of which is incorporated herein by reference.
[0002] The present invention relates to a method for analyzing loss of heterozygosity (LoH) in a sample from low - pass whole - genome sequencing data by deterministic restriction site - based whole - genome amplification (DRS - WGA) to achieve single - cell resolution, with or without the use of normal controls. This method can be applied in several single - cell applications, such as the analysis of circulating tumor cells, and in oncology, including single - cell heterogeneity in tissue samples, or in reproductive medicine, including pre - implantation genetic screening (PGS).
Background Art
[0003] Whole - genome amplification (WGA) of single - cell genomic DNA is often required to obtain more DNA for the purpose of facilitating and / or enabling various types of genetic analysis, including sequencing, SNP detection, etc. WGA by LM - PCR based on deterministic restriction sites (hereinafter DRS - WGA) is known from WO2000 / 017390.
[0004] Importantly, DRS - WGA has been shown to be the top - class WGA method from many perspectives, especially in terms of having less allelic drop - out from single cells (Borgstrom et al., 2017; Normand et al., 2016; Babayan et al., 2016; Binder et al., 2014).
[0005] The commercially available kit of DRS-WGA (Ampli1™ WGA kit, Silicon Biosystems) based on LM-PCR has been used in Hodgkinson C.L. et al., Nature Medicine 20, 897-903 (2014). In this study, copy number analysis by low-pass whole genome sequencing in single-cell WGA material was performed, where digestion and fragmentation of the WGA adapter were carried out before adapter ligation with Illumina barcodes for sequencing.
[0006] WO2017 / 178655 and WO2019 / 016401A1 teach a simplified method for preparing ultra-parallel sequencing libraries from DRS-WGA (e.g., Ampli1), or MALBAC for low-pass whole genome sequencing and copy number profiling. In Ferrarini et al., PLoSONE 13(3):e0193689, https: / / doi.org / 10.1371 / journal.pone.0193689, the method performance of WO2017 / 178655 using the Ion Torrent platform is detailed with reference to copy number profiling.
[0007] Ampli1™ WGA is compatible with array comparative genomic hybridization (aCGH). In fact, several groups (Moehlendick B. et al., 2013, PLoS ONE 8(6): e67031; Czyz ZT. et al., 2014, PLoS ONE 9(1): e85907) have shown that it is suitable for high-resolution copy number analysis. However, since the aCGH technology is expensive and labor-intensive, different methods such as low-pass whole genome sequencing (LPWGS) may be desirable for the detection of somatic copy number alterations (CNA).
[0008] DRS-WGA has been shown to be better than DOP-PCR for the analysis of copy number profiles from small amounts of microdissected FFPE material when using array CGH, metaphase CGH, and further for other genetic analysis assays such as loss of heterozygosity using targeted primers and PCR for the analysis of selected microsatellites (Stoecklein et al., Am J Pathol. July 2002; 161(1):43-51; Arneson et al., ISRN Oncol. 2012;2012:710692. doi: 10.5402 / 2012 / 710692. Epub Mar 14, 2012).
[0009] U.S. Patent No. 7,424,368 B2 teaches a method for estimating the copy number of genomic regions in an experimental sample, including the analysis of SNPs using a microarray. Microarray technology is less processable and flexible compared to next-generation sequencing and provides only relative signals, not absolute counts. Moreover, in contrast to next-generation sequencing (NGS), there are setup costs associated with probe synthesis and microarray manufacturing.
[0010] Zahn H. et al., Nature Methods, volume 14, pages 167-173 (2017) teach a method for preparing ultra-parallel single-cell libraries without pre-amplification and show the simultaneous inference of CNA and LoH on the bulk equivalent of the SA501X3F cell line. However, this approach requires a relatively large number of single cells (48). In addition, to perform the analysis using TITAN, the positions of heterozygous SNPs must be determined (Ha G. et al., 2014, Genome Research 24(11)).
[0011] This method has the following drawbacks.
[0012] 1. This is not compatible with the use of whole genome amplification libraries, but WGA is often actually desirable, for example when dealing with CTCs, for biomarker discovery, or for evaluating other known valid biomarkers that cannot be inferred by low-pass WGS alone, among other purposes. At the single cell level from each individual cell, for example, to obtain additional information regarding SNVs in oncogenes or tumor suppressor genes, reanalysis of different aliquots of the WGA product is necessary.
[0013] 2. In certain applications such as preimplantation genetic screening (PGS) or preimplantation genetic diagnosis (PGD), only single cells may be available, and thus the approach of Zahn et al. is clearly inapplicable.
[0014] 3. In certain applications, a large number of cells are available for analysis, but they may still be insufficient to provide enough information to use the approach of Zahn et al. For example, the number of CTCs collected from a 7.5 ml blood draw from metastatic patients using the CELLSEARCH system is less than 10 in most cases (see Allard WJ. et al., Clin Cancer Res., October 15, 2004; 10(20):6897 - 904, Table 2).
[0015] In oncology, the whole genome evaluation of LoH has been shown to be important in several contexts, including the assessment of the so-called BRCAness signature, in relation to the effectiveness of platinum-based therapies and poly(ADP - ribose) polymerase (PARP) inhibitors in some cancer types (e.g., Watkins et al., Breast Cancer Research 2014, 16:211). In addition, analysis of LoH at the BRCA1 and BRCA2 loci in tumors of individuals with germline mutations has been shown to be important for the effectiveness of treatment.
[0016] In preimplantation genetic screening (PGS) or preimplantation genetic diagnosis (PGD), it is desirable to determine uniparental disomy (UPD) that occurs when a person inherits two copies of a chromosome or part of a chromosome from one parent and no copies from the other parent. However, this type of information is not available from a standard LPWGS workflow when using conventional bioinformatics pipelines and analysis methods.
[0017] There is a need to provide a method that can infer whole-genome LoH status (and / or gene-specific LoH status) at single-cell resolution and overcome one or more of the following limitations inherent in the current state of the art: - The need for high-coverage whole-genome sequencing, or in other words, multiple single-cell low-pass sequencing to generate a high-coverage bulk equivalent; - A normal control is an essential requirement; - It is not possible to reliably re-analyze single cells for validation or additional targeted genomic information.
[0018] For CTC analysis and, more generally, other single-cell analysis applications, such as prenatal diagnosis in blastocysts and circulating fetal cells collected from maternal blood, it is desirable to have an efficient method that combines the reproducibility and quality of DRS-WGA and the ability to analyze whole-genome LoH together with copy number variants (CNVs) from the same low-pass sequencing data.
[0019] In addition, it is also desirable to determine whole-genome copy number profiles and LoH from similarly small amounts of cells, FFPE, or tissue biopsy samples.
[0020] Binder, V. et al., "A new workflow for whole-genome sequencing of single human cells", Human mutation, Vol. 35, No. 10, pp. 1260-1270, 2014, disclose a workflow that combines an efficient adapter-linker PCR-based WGA method with next-generation sequencing. This approach enables the comparison of single cells at base pair resolution. However, this method is based on genotyped SNPs, i.e., polymorphic genomic positions, for which sufficient coverage can be obtained to call genotypes with a certain confidence.
[0021] The above method and the method of the present invention have quite different purposes. The purpose of the present invention is to infer the whole-genome LoH state (and / or gene-specific LoH state) at the resolution of single cells, rather than at polymorphic gene positions as in Binder et al.
[0022] The method of Binder et al. implies a large number of reads, two orders of magnitude larger. Instead, according to the present invention, LoH can be called from a single sample starting from, for example, two million reads, corresponding to less than 1% of the reads used by Binder et al.
Prior art documents
Patent documents
[0023]
Patent document 1
Patent document 2
Patent document 3
Patent document 4
Non-patent documents
[0024]
Non-patent document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Non-Patent Document 6
Non-Patent Document 7
Non-Patent Document 8
Non-Patent Document 9
Non-Patent Document 10
Non-Patent Document 11
Non-Patent Document 12
Non-Patent Document 13
Non-Patent Document 14
Non-Patent Document 15
Non-Patent Document 16
Non-Patent Document 17
Summary of the Invention
Problems to be Solved by the Invention
[0025] Therefore, one object of the present invention is to provide a method for analyzing LoH that overcomes the drawbacks of prior art methods.
[0026] In particular, an object of the present invention is a method for analyzing LoH from a small number of cells with a resolution up to single cell after whole genome amplification, which involves using fewer cells, fewer normal controls, and fewer sequencing reads per cell for analysis than generally reported in the art.
Means for Solving the Problem
[0027] This object is achieved by the method according to claim 1.
Brief Description of the Drawings
[0028]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 14C
Figure 14D
Figure 14E
Figure 14F
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Best Mode for Carrying Out the Invention
[0029] Definitions Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Many methods and materials similar or equivalent to those described herein may be used in practicing or testing the present invention, but the preferred methods and materials are described below. Unless otherwise noted, the techniques described herein in connection with the present invention are standard techniques known to those of ordinary skill in the art.
[0030] The expression "ultra-parallel next-generation sequencing (NGS or MPS)" is intended to mean a method of sequencing DNA that involves the creation of a library of DNA molecules that are spatially and / or temporally separated and sequenced clonally (with or without prior clonal amplification). Examples include the Illumina platform (Illumina Inc), the Ion Torrent platform (Thermo Fisher Scientific Inc), the Pacific Biosciences platform, and MinIon (Oxford Nanopore Technologies Ltd).
[0031] The expression "low-pass whole-genome sequencing" is intended to mean whole-genome sequencing with an average sequencing depth of less than 3 when referenced against the entire reference genome.
[0032] As used herein, the expression "average sequencing depth" is intended to mean, for each sample, the total number of bases that are sequenced and mapped to the reference genome, divided by the size of the entire reference genome. The total number of bases that are sequenced and mapped can be approximated as the number of mapped reads × average read length.
[0033] The expression "reference genome" is intended to mean a reference DNA sequence for a particular species.
[0034] The term "locus" (plural "loci") is intended to mean a defined position on a chromosome (relative to the reference genome).
[0035] The expression "polymorphic locus" is intended to mean a locus that has two or more alleles with an observed frequency in the population that exceeds 1%.
[0036] The expression "heterozygous locus" is intended to mean a locus having two or more alleles observed in a particular sample.
[0037] The expression "genomic window" is intended to mean a region of the reference genome contained within a single chromosome, having a fixed or variable length.
[0038] The expression "genomic region" is intended to mean a region containing one or more adjacent genomic windows within the same chromosome.
[0039] The expression "covered genome" is intended to mean the portion of the reference genome covered by at least one read.
[0040] The term "read" is intended to mean a small piece of DNA that is sequenced ( "read") by a sequencer.
[0041] The expression "copy number region" is intended to mean a genomic region associated with the same copy number value.
[0042] The expression "segmented copy number region" is intended to mean a genomic region associated with the same copy number value as a result of bioinformatics analysis of CNA.
[0043] The expression "tumor suppressor gene" is intended to mean a gene, for example, in which loss of function due to a sequence variant - germline or somatic - is associated with an increased probability of tumor development.
[0044] The expression "reduction ratio" is intended to mean the total number of bases of the fragments obtained by in - silico digestion of the reference genome by the restriction enzyme employed in DRS - WGA, within a specified base pair range, divided by the total number of bases in the reference genome.
[0045] The expression "loss of heterozygosity" or "LoH" is intended to mean the loss of one allele in a genomic region.
[0046] The expression "LoH call" is intended to mean the assignment of the presence of LoH (in a genomic region).
[0047] The expression "allele content" is intended to mean the composition regarding the alleles detected at a locus.
[0048] For the sake of simplicity, in the description of the present invention, unless otherwise specified, a locus may be referred to as homozygosity or monoallelism when only one allele is detected, regardless of the actual genotype of the locus, and as heterozygosity or biallelism when at least two alleles are present.
[0049] Detailed Description of the Invention Referring to FIG. 1, a method according to the present invention for analyzing the loss of heterozygosity (LoH) in at least one sample containing genomic DNA includes the following steps.
[0050] In step a, at least one sample containing genomic DNA is prepared.
[0051] In step b, whole-genome amplification of definitive restriction sites (DRS-WGA) of the genomic DNA is performed.
[0052] In step c, a massively parallel sequencing library is prepared from the product of the DRS-WGA.
[0053] In step d, low-pass whole-genome sequencing is performed on the massively parallel sequencing library at an average coverage depth of less than 1, preferably less than 0.05, more preferably less than 0.01.
[0054] In step e, the reads obtained in step d are aligned onto a reference genome for the at least one sample.
[0055] In step f, the allelic content at multiple loci is extracted. The multiple loci include polymorphic loci and / or heterozygous loci.
[0056] In step g, for at least one genomic window of the reference genome for the at least one sample, a LoH score is assigned as a function of the number of loci having at least two different alleles at the multiple loci.
[0057] Preferably, the size selection step is performed before, during, or after step c of preparing the next-generation sequencing library, and the step of preparing the next-generation sequencing library does not include a random fragmentation step.
[0058] The size selection step preferably retains fragments in the range of 100 to 800 base pairs.
[0059] In a particular embodiment of the present invention, size selection preferably retains fragments in the range of 300 to 450 base pairs.
[0060] In a particular embodiment of the present invention, the peak of the fragment retained in the size selection step is preferably centered on a base pair range of 150 bp to 600 bp, and more preferably, the size selection step retains fragments in the range of 425 to 575 base pairs.
[0061] Preferably, at least one genomic window is: - having a constant width of base pairs, or - having a certain number of the multiple loci, or - selected from the group consisting of chromosomes, chromosome arms, and segmented copy number regions.
[0062] The plurality of loci preferably include polymorphic loci obtained from a database regarding the reference genome of said at least one sample, such as dbSNP, or obtained by genotyping a reference set of samples.
[0063] Alternatively, the plurality of loci preferably include known heterozygous loci regarding a control sample.
[0064] When the genomic window has a constant width in base pairs, or has a certain number of plural loci, or when the plural loci include polymorphic loci regarding the reference genome of said sample, the LoH score preferably corresponds to the number of heterozygous loci in said at least one genomic window.
[0065] Preferably, the LoH score corresponds to the ratio of heterozygous loci to the total number of polymorphic loci in at least one genomic window.
[0066] The LoH score preferably corresponds to the p-value of a statistical test.
[0067] The statistical test preferably determines the significance of over-representation of two-allele loci with respect to the error rates of sequencing and WGA, or the significance of under-representation of two-allele loci with respect to a control sample.
[0068] The control sample preferably includes at least one genomic region at the major ploidy from at least one sample.
[0069] The control sample is preferably at least one normal sample, which is more preferably obtained from the same individual under the test from which said at least one sample was obtained. In an oncology case, the control sample is preferably a normal (non-tumorous) sample.
[0070] In the case of circulating fetal cells, the control sample is preferably a maternal sample. Alternatively, if a paternal sample is available, it may be a paternal sample, or a combination of a maternal sample and a paternal sample. The availability of the maternal and / or paternal genotypes can be utilized to select a subset of loci known to be heterozygous in parental controls.
[0071] Preferably, when the LoH score meets the threshold for a genomic window, the genomic window is said to be in LoH. In this case, the method more preferably comprises the step of assigning an LoH state to at least one genomic region when the LoH score for each genomic window included in the region meets the threshold, or as a function of the LoH state of the genomic windows included in the region, the step of assigning an LoH state to at least one genomic region.
[0072] More preferably, at least one genomic region contains a tumor suppressor gene, and the tumor suppressor gene is even more preferably selected from the group consisting of BRCA1, BRCA2, PALB2, TP53, CDKN2A, RB1, APC, PTEN, CDKN1B, DMP1, NF1, AML1, EGR1, TGFBR1, TGFBR2 and SMAD4.
[0073] At least one sample preferably has a purity of at least 50%. More preferably, the at least one sample is a single cell.
[0074] The unambiguous relationship between loci and fragment lengths in DRS-WGA More specifically, the method according to the present invention utilizes, in DRS-WGA, for example Ampli1™ WGA, the fact that each locus in the genome is presented only within fragments having a base pair of a specific length in the WGA library. This property can be named as "unique relationship between locus and fragment length" (L2FLUR). Considering a general normal locus, for example, a locus related to polymorphic SNP, the locus is considered to be presented only in fragments of a given length equal to the size of the corresponding fragment plus twice the length of the universal WGA adapter (LIB1 primer length in the case of Ampli1 WGA) in addition to digestion with restriction enzymes (measured in either single strand). When WGA is sequenced after library preparation with the Ampli1 LowPass kit, a predictable additional length is introduced and ligated to the known sequencing adapter and barcode length.
[0075] Non-idealities, such as undigested restriction sites or sequence variants, and further other factors, can affect and distort the frequency of presentation of a given fragment in the WGA product relative to what is theoretically expected. These factors are typically of moderate degree, and in addition, as long as they are reproducible, their non-random nature is partially offset by canceling out their effects. Therefore, their effects will be ignored in this specification, unless otherwise indicated.
[0076] Reduction of genomic representation In the method according to the present invention, the L2FLURL property is utilized to generate a reduction of genomic representation, and the low-pass sequencing data for a given number of reads effectively reduces the size of the covered genome relative to the original size of the sample reference genome, thereby achieving a higher coverage of the covered genome. In other words, the size selection of the WGA fragments results in a deterministic subsampling of the reference genome. The term "deterministic" is essential in that as the number of reads increases, the same genomic loci will ultimately be resampled (see Figure 2).
[0077] Figure 2 shows the effect of genomic representation reduction on the observed coverage. Figure 2A shows the MseI fragment length distributions for three different approaches: Ampli1 LowPass for Ion Torrent with size selection to collect fragments of 300 - 450 bp (A1LP_ss), Ampli1 LowPass using selection guided by the sequencing process (A1LP), and a library obtained by performing random fragmentation and sequencing after Ampli1 WGA (A1_wFrg) (Binder V et al. 2014). These three different approaches correspond to different levels of reduction in genomic representation, ranging from the most stringent A1LP_ss to A1_wFrg, which is characterized by no selection. Figure 2B shows the Lorenz curves obtained using these different approaches, which shows a gradual decrease in coverage uniformity with increasing size selection level. The lower uniformity of A1LP_ss can be explained by saturation of the DNA template and repeated sequencing of the same fragments. Saturation of the template is confirmed by the plots in Figures 2C and 2D, which show the total number of bases covered and the average coverage per base in incremental intervals of mapped reads, respectively. These plots clearly show that the size selection step (A1LP_ss) has the effect of reducing the amount of available DNA and limiting the targets covered, but results in higher coverage.
[0078] It is worth noting that this approach is flexible in that various definitive enzymes may be suitable depending on the desired resolution and / or the sequencing platform and protocol used. For example, various high-frequency cutters can be used. In the case of Ampli1 WGA, the TTAA motif is the restriction site. Other 4-base cutters can also be used to cut at various restriction sites, such as GTAC, CTAG (Figure 3), to obtain different distributions of fragments. Figure 3 shows in silico digestion of the human genome using various restriction sites (4-base pairs or 6-base pairs). For a given range of fragment lengths (e.g., suitable for a particular sequencer and size selection method), different numbers of fragments are obtained depending on the various restriction sites.
[0079] If DRS-WGA is purified first after the first PCR, a first size selection is performed, where the shorter fragments of the WGA are removed along with the free primers. Advantageously, the method uses a further selection step. This additional selection step can be achieved either by size selecting specific fragments from the primary WGA and / or by a method of restricting the fragments that can be sequenced to generate an ultra-parallel sequencing library. For example, the Ampli1 LowPass kit includes an inherent size selection step, which is sufficient to have a positive impact on the process. In WO2017 / 178655, size selection on a gel is performed. In WO2019 / 016401, a first size selection is effectively performed by a continuous purification step using SPRI beads, where the base pair length is restricted to a range that is substantially dependent on the SPRI bead concentration. In addition, the sequencer itself may introduce size selection, because the longer the fragment, the lower the efficiency of generating sequence data (e.g., due to emulsion PCR efficiency in the case of Ion Torrent or due to bridge PCR for cluster formation in the case of the Illumina platform).
[0080] In DRS-WGA, there is also a definite relationship between the average size of the sequencing library and the subsampling ratio of the reference genome.
[0081] In the in silico analysis performed on the TTAA digest of the human reference genome hg19 (Figure 4), a total of approximately 19M fragments containing all chromosomal sequences were obtained, which are considered to be translated into 38M fragments on the normal diploid human genome. As an example, if in silico is selected, the fragments in the range of 175 - 225 bp are only 1,252,559, which cover a total of approximately 248M bases out of 3.09B bases, that is, 8.02% of the human reference genome. Referring to Table 1 below, the number of fragments, total base pairs, and reduction ratio (%) are listed for different selection ranges by size. This subsampling can be named the reduction ratio (RR).
[0082]
Table 1
[0083] Along with the reduction ratio, in DRS-WGA, there is also a definite relationship between the average interval and consecutive fragments according to the portion of the fragment length distribution selected for sequencing. In this regard, referring to Figure 5, Panel A shows a positive correlation between fragment length and interval, which is due to the decrease in the number of selected fragments when measured with a band of ±100 bp for three different fragment sizes of 200, 500, and 800; Panel B shows that three different bands (±50, ±100, ±150) were used for each fragment size, indicating an inverse correlation between band size and interval, which is also due to the decrease in the number of available fragments when the size range becomes narrower.
[0084] Generally, regarding the Ampli1 DRS-WGA fragment distribution, the following is found by in silico analysis of the human reference genome hg19: · The larger the average base pair length of the selected fragments, the fewer the number of fragments and the larger the intervals between them; · The narrower the range of the selected fragments, the fewer the number of fragments and the larger the intervals between them.
[0085] Fragment size selection Also, various size selection techniques can be used to achieve the desired reduction ratio, which depends on the selected number and / or resolution of sequencing reads per sample. Referring to Figure 4, for a given average fragment length, it is clear that by selecting smaller or larger bands centered around that average fragment length, fewer or more total fragments can be obtained.
[0086] Using an apparatus such as Pipping prep (Sage Science), the fragment length distribution can be more strictly controlled, using the analogy of a passband filter, Q = Fcenter / DeltaF = [(Fmin + FMAX) / 2] / (FMAX - Fmin) it is also possible to have a higher Q factor defined as, where, Fcenter = (Fmin + FMAX) / 2 is the average size of the fragments, DeltaF = FMAX - Fmin is the width of the range of fragment sizes.
[0087] Fmin is the size of the fragments below which the fragments are depicted as being at a conventional relative level (e.g., 1 / 10 = 10%) or less relative to the number of peaks within the normalized band of fragments per bin.
[0088] FMAX is the size of the fragments above which the fragments are represented as being at the same conventional relative level or less relative to the number of peaks within the normalized band of fragments per bin.
[0089] When using Illumina sequencing, the sequencing mode is preferably paired-end sequencing, because this increases the genome coverage, and thus the number of loci per million read pairs increases, enhancing the resolution. However, if the size selected for sequencing is below a certain size, the coverage is not considered to increase in paired-end sequencing because the two paired reads completely overlap.
[0090] When using Ion Torrent sequencing, a longer read length results in a proportional increase in the genome coverage, and thus the number of loci per million reads increases, enhancing the resolution. In the case of the Ampli1 LowPass Ion Torrent kit (Menarini Silicon Biosystems), barcoded pooled samples are size-selected on a gel or by other methods such as Pippin Prep. Different resolutions can be provided per million reads by selecting various Q factors and average fragment lengths.
[0091] One advantage of pooling samples and then performing size selection of the library for sequencing is that the fragment length distribution is considered to be the same for all samples, and thus the overlap of the genomes covered between different samples is maximized. This applies when using a control (e.g., a normal control or a maternal control)-based approach to identify potential heterozygous loci in the subject under test (SUT).
[0092] On the other hand, when using the Ampli1 LowPass for Illumina kit, if different LowPass libraries are first size-selected and then pooled, somewhat different size selections are made between different samples, resulting in a decrease in the genome covered between different samples per million reads. Size selection after pooling of the libraries is not mandated by the standard protocol, but may be employed to increase the overlap between samples, which may be beneficial for analysis based on controls.
[0093] According to the present invention, the combination of DRS-WGA and LPWGS unexpectedly results in a decrease in representation from the input sample. By sequencing using NGS, this decrease in representation of the library of the reference genome in turn reduces the genome covered within the selected (or sequenceable by any method) base pair range compared to alternative WGA methods using random priming or random shearing, and effectively high coverage of the genome covered per million reads is obtained.
[0094] According to the present invention, this effect can be utilized in various ways depending on the situation.
[0095] One example is when one or more control samples, such as "matched normal", are available and one or more subject samples (SUTs), such as tumor samples, are available. In this case, DRS-WGA increases the overlap of reads between the SUT and the control.
[0096] Another example is a situation without a control, such as in preimplantation genetic screening (PGS), where only a single sample corresponding to the SUT is available. In this case, DRS-WGA increases the number of loci covered by two or more reads.
[0097] Preferably, library preparation from DRS-WGA is one of the methods disclosed in WO2017 / 178655 and WO2019 / 016401, which results in a higher reduction ratio compared to digesting the WGA adapter, fragmenting the DNA, and then creating a sequencing-ready library as performed by Binder V. et al., 2014, or Hodgkinson C.L. et al., 2014. Indeed, DNA shearing increases the number of possible different fragments of the original DRS-WGA that can be found within a given base pair range selected for sequencing. However, once fragmented, longer fragments tend to fit within the above range, while smaller fragments tend to be less efficiently sheared compared to longer fragments, so that only a portion of the primary WGA fragments that were originally within the range are pushed out of the range due to fragmentation (see Figure 2).
[0098] LoH analysis Referring again to FIG. 1, the ultra-parallel sequencing library is preferably obtained using the Ampli1 LowPass kit (for Ion Torrent or for Illumina). Sequencing of the sample is performed using a compatible sequencer. The sequenced reads obtained from the library are mapped to a reference human genome, and alleles present at known loci and / or polymorphic genes are extracted. Preferably, such loci are covered by at least two sequencing reads. It should be noted that the detection of a single allele does not necessarily mean an actual homozygous genotype and may result from low sequencing coverage. The plurality of loci are preferably subdivided into genomic windows according to different criteria of genome partitioning. This partitioning is optional, and in certain embodiments, one may be interested only in the analysis of one or a few predetermined genomic windows, for example, a single chromosome or a single genomic locus containing one or more genes of interest. The allelic state of the loci detected within the genomic window is used to obtain a measurement. Such a measurement, hereinafter referred to as the LoH score, can be obtained by various methods according to the present invention, for example, by counting the number of heterozygous loci within the genomic window or by calculating the ratio of heterozygous loci. Further, it is preferable to apply a statistical test to determine the significance of the decrease in heterozygous loci corresponding to the LoH event, either by comparison with an internal control or by using an external control (from the same individual or from different individuals). Alternatively, it is preferable to apply a statistical test to determine the significance of the overrepresentation of heterozygous loci corresponding to genomic regions not subject to LoH relative to that expected based on the error rates of sequencing and WGA. Further, it is preferable to apply thresholding of the LoH score based on a fixed threshold calculated from a training dataset by known LoH events to define genomic regions corresponding to LoH events. The individual steps of the method are detailed below.
[0099] Genome partitioning Referring to FIG. 1, the optional step of splitting can be carried out in three alternative manners: i) a fixed base pair genomic window ii) a window of a fixed number of loci iii) a copy number segment.
[0100] In option i) shown in FIG. 6, the genomic window has a fixed width. Each genomic window contains a plurality of loci, the number of which depends on the position on the genome. This approach can be advantageous when comparing a sample to a set of control normal samples, as the reference genome is split in the same way for all samples, allowing for a direct comparison of LoH scores for each genomic window among a large number of samples. Since the number and proportion of heterozygous loci detected within a genomic window of a defined width increase with increasing read depth, it is preferable to normalize the number of mapped reads in each sample to a fixed number of reads in order to enable comparison of the sample to one (or a number of) control samples. Such normalization is done by randomly sampling reads and mapping them to the reference genome until the desired number is reached. The number of normalized reads may be, for example, 1 million or 2 million reads, preferably 3 million, 4 million, 5 million, 6 million, 7 million, 8 million or 9 million reads.
[0101] Figure 6 is a schematic diagram of an example of partitioning based on a fixed base pair genomic window. A pair of control samples (top) and test samples (bottom) are shown. The solid line represents the genome (a portion of it). The diamond markers define the boundaries of genomic windows of a fixed width, and known polymorphic loci are represented by dots (heterozygous loci: white-filled dots; homozygous loci: gray-filled dots). The number of loci detected per genomic window varies across the genome, but for a given window, it is expected that on average it will be similar between two different samples where the total read mapping has been normalized against a defined read count. A genomic window in LoH in the test sample is expected to show a decrease in heterozygous loci compared to the same window in a normal control sample. Due to the bias in SNP density along the genome, genomic windows located at different genomic positions on the same (or other) sample cannot be directly compared with the same window.
[0102] In option ii) shown in Figure 7, the genomic window has a fixed number of loci. This approach enables the normalization of the LoH score with respect to different SNP densities across the genome. This method can be advantageous when using a non - control approach, for example, because it allows the application of the same threshold to all genomic windows regardless of their positions in the genome and the underlying SNP density. This method can be disadvantageous when comparing a test sample to a control sample because depending on the distribution of loci sampled and detected by low - pass sequencing, different genomic windows may be generated for different samples.
[0103] Figure 7 shows a schematic diagram of an example of the division based on the fact that the number of loci per window is constant. It represents paired control samples (top) and test samples (bottom). The solid lines represent (a part of) the genome. The diamond markers define the boundaries of genomic windows containing a fixed number of loci. Known polymorphic loci are represented by dots (heterozygous loci: white-filled dots; homozygous loci: gray-filled dots). Since the sequencing coverage is low, not all loci within the genomic region are detected. Therefore, the ends of the genomic windows may differ between different samples based on locus sampling by sequencing reads, and thus the genomic windows detected in the test samples are not directly comparable to the corresponding genomic windows in other (control) samples. The genomic windows in LoH in the test samples are expected to show a decrease in heterozygous loci compared to the genomic windows of the same sample that are not in LoH.
[0104] The number and proportion of heterozygous loci detected in genomic windows using a fixed number of loci are thought to increase as the read depth increases (see Figure 8). Preferably, the number of reads mapped in each sample is normalized to a fixed number of reads so that the threshold setting of the LoH score can be set to a pre-calculated value. Such normalization is performed by randomly sampling reads and mapping them to the reference genome until the desired number is reached. The number of normalized reads may be, for example, 1 million or 2 million reads, preferably 3 million, 4 million, 5 million, 6 million, 7 million, 8 million or 9 million reads.
[0105] In option (iii) shown in FIG. 9, the genomic window is a genomic region segmented between two copy number breakpoints contained in the chromosomal arm, which is achieved by normalizing the raw copy number counts in the genomic window to be normalized by the GC content (Boeva, V. et al., 2011, Bioinformatics, 27(2), pp. 268-269), and applying a segmentation algorithm, such as an algorithm based on LASSO (Harchaoui, Z. et al., 2008, Adv. Neural Inform. Process. Syst., 20, pp. 617-624), circular binary segmentation (CBS) (Seshan VE. et al., 2019, DNAcopy: DNA copy number data analysis. R package version 1.58.0), or a similar algorithm for normalizing read counts. This method is based on the assumption that for the major "normal" ploidy of the sample, genomic regions showing changes in copy number levels are likely to be affected by a single genomic copy number aberration event and, therefore, are expected to have a uniform LoH state. Compared with options (i) and (ii), the genomic windows defined by this method are generally much larger (up to two to three orders of magnitude larger), contain more known heterozygous and / or polymorphic loci, and are therefore thought to provide higher statistical power of detection. Furthermore, by combining two different biological dimensions (copy number, LoH score), more accurate results can be achieved with a lower false positive rate using this method. However, this method may be disadvantageous when small LoH events are located within larger copy number events and are thought not to be detected by this method. Since it is not uncommon for chromosomal arms to undergo LoH events followed by replication, chromosomal arms are preferably used as segmentation units in chromosomes without copy number changes.This prevents the false calling (false positive) of LoH of a shorter chromosomal arm when only the longer arm is affected, or conversely, the false calling (false negative) of non-LoH for the shorter chromosomal arm when only the shorter one is affected.
[0106] More specifically, FIG. 9 provides an exemplary presentation of the copy number profiles of chromosomal arms (genome main ploidy = 2) affected by two copy number change events: a copy number loss segment with copy number = 1; a copy number gain with copy number = 3. A genomic window is defined as the region between two consecutive copy number breakpoints.
[0107] Also, to exclude false positives resulting from high-level amplifications, segmentation that utilizes copy number information can be employed. In fact, high-level amplifications are almost certainly derived from a single allele, and thus bias is introduced into the allele representation within that region, and even if a minor allele exists, it will be underrepresented and may cause false positive LoH calls.
[0108] The following Table 2 shows the main features, advantages, and disadvantages of each alternative step of the segmentation according to the present invention.
[0109] [Table 2]
[0110] LoH Scoring Step g of assigning an LoH score as a function of the number of loci having at least two different alleles at at least one genomic window of the reference genome for the at least one sample also includes an alternative preferred embodiment.
[0111] In one preferred embodiment, the LoH score corresponds to the number of heterozygous loci in the at least one genomic window. The genomic window in LoH is expected to show a low number of heterozygous loci as compared to regions or samples that are not LoH (see Figure 10).
[0112] In another preferred embodiment, for each genomic window, the LoH score is defined as the ratio of the number of heterozygous loci detected within that genomic window to the total number of polymorphic loci in the same genomic window (Figure 11). Similar to the above method, in the presence of a LoH event, a consistent decrease in the LoH score is expected. This method is considered advantageous when the window does not contain a uniform number of loci detected, for example, when using a fixed base pair genomic window or when using copy number segments to divide the genome.
[0113] LoH Scoring - Statistical Test Preferably, for each genomic window, the LoH score is determined by the result of a statistical test on the frequency of observed two - allele loci.
[0114] In one preferred embodiment, by performing a statistical test, the significance of the underrepresentation of heterozygous loci relative to internal / external controls can be determined. Specifically, for each genomic window, a contingency table is created considering the following two classifications: 1) sample type (test, control); 2) locus type (heterozygous, homozygous). A statistical test, such as Fisher's exact test, or an equivalent test for contingency table analysis (e.g., chi-square test, G test, Barnard's exact test, Fisher-Freeman-Halton test) is then applied. Preferably, the statistical test should be one-sided in order to limit detection to cases where there is underrepresentation of heterozygous loci due to LoH. In fact, in a given genomic segment, if there is a gain, i.e., an increase in copy number, the number of reads using low-pass WGS increases. This can result in a larger number of heterozygous loci than in the case where there is no LoH, and for the purposes of analysis, for the opposite reason, it may be marked as significant by a two-sided statistical test.
[0115] In an alternative preferred embodiment, the significance of overrepresentation of heterozygous loci can be tested against what is expected from the error rates of sequencing and WGA. This approach may be advantageous when testing for "gain of heterozygosity" (hereinafter GoH) in haploid single cells such as gametes. This can result, for example, from errors in meiotic nondisjunction that lead to chromosome gain.
[0116] Taking into account the multiple tests performed for each experiment (about 200, 400, 600 for a sample of 1 million reads with fixed windows of 500, 1000, and 1500 SNPs), multiple testing correction can be applied (see, for example, Benjamini Y. et al., 1995, Journal of the Royal Statistical Society. Series B (Methodological) Vol. 57, No. 1: pp. 289-300). Then, the LoH score is defined as the p-value obtained from the statistical test.
[0117] Control sample The control may be "internal" and can be defined, for example, by considering genomic regions having a ploidy equal to the major (average) genomic ploidy with the highest likelihood. This approach assumes that most genomic regions showing no copy number changes are not LoH.
[0118] Alternatively, the control may be "external" and can be generated, for example, by using one or more normal samples from the same individual during the test or from different individuals.
[0119] The use of an internal control can be advantageous for diploid or polyploid samples (e.g., tumor samples) and for damaged samples (e.g., FFPE samples) since it is independent of the number of reads (no normalization of the number of mapped reads is required). In fact, in damaged samples, the incidence of dropout where one of the two alleles at a locus is lost due to DNA damage may be shown to be higher compared to non-damaged ones, and thus the number of heterozygous sites may be lower than expected for genomic regions that are not LoH. This can interfere with the comparison of test samples with different levels of damage to external control samples. By using an internal control, such bias is removed as the genomic windows for control and test are considered to have the same dropout rate.
[0120] LoH threshold setting and LoH calling Optionally, a threshold can be set for the LoH score obtained from the previous step to define the genomic region at LoH. In most cases, the number and proportion of heterozygous loci detected within a genomic window with a constant number of loci increase with increasing read depth. To enable setting the threshold of the LoH score to a pre-calculated value, it is preferable to normalize the number of mapped reads in each sample against a fixed number of reads. Such normalization is performed by randomly sampling reads and mapping them to the reference genome until a desired number is reached (preferably within the range from 1,000,000 to 10,000,000 mapped reads). The above considerations do not apply when calculating the LoH score by performing a statistical test against an "internal" control.
[0121] Preferably, in the case of the LoH score calculated as the number of heterozygous loci, the data is first downsampled to 1,000,000 mapped reads. Loci covered by at least 1 read are divided using windows with a fixed number of detected loci (e.g., n = 500; n = 1000; n = 1500). Some preferred thresholds are 3, 6, and 9 heterozygous SNPs for 500, 1000, and 1500 loci, respectively (Figure 12). If the LoH score is lower than the selected threshold, LoH is called within a given genomic window.
[0122] More specifically, Figure 12 shows the ROC analysis used to define the LoH score threshold, defined as the number of two-allele SNPs within windows of (A) n = 500, (B) n = 1000, (C) n = 1500 SNPs covered by at least 1 read using 1,000,000 mapped reads. LoH detected in tumor cells by whole-genome high-pass sequencing and B-allele frequency analysis was used as a reference.
[0123] In the case of the LoH score calculated as the p-value obtained as a result of the application of a statistical test, some preferred thresholds are, for example, 5×10 -2 or 1×10 -2 . Subsequently, if the LoH score is lower than the selected threshold, LoH is called within the genomic window.
[0124] Once the LoH score threshold is set, the LoH state can be assigned to genomic regions according to the following different criteria. 1) Calling of LoH regions by merging windows. In this preferred embodiment, the LoH state is assigned to the genomic region if the LoH score for each genomic window contained in that region passes the threshold setting step. 2) Calling of LoH regions as a function of the LoH state within a genomic window. In this preferred embodiment, the LoH state is assigned to the genomic region if a given percentage / fraction of the genomic windows contained in that genomic region passes the threshold setting step. As an example, if more than 66%, 75%, 80%, 85%, 90%, 95% of the windows within the genomic region pass the threshold setting step, the LoH state is assigned to that genomic region. 3) Calling of LoH in genomic regions containing tumor suppressor genes. In this preferred embodiment, at least one genomic region contains a tumor suppressor gene.
[0125] Preferably, the gene is selected from the group consisting of BRCA1, BRCA2, PALB2, TP53, CDKN2A, RB1, APC, PTEN, CDKN1B, DMP1, NF1, AML1, EGR1, TGFBR1, TGFBR2, and SMAD4.
[0126] Sample purity LoH can be identified in DNA derived from a mixture of different types of cells (e.g., tumor cells and normal cells). Sample purity is defined as the percentage of the sample in the mixture belonging to the type of interest (e.g., tumor cells).
[0127] For example, when mixing clonal, i.e., genomically identical, and thus #TC tumor cells with the same LoH and CNA patterns, with #NC normal cells from the same individual, the purity of the resulting sample is #TC / (#TC+#NC), and the sample is considered to be uniform across the entire genome.
[0128] In general, purity, as used herein by the inventors, means a concept related to the LoH state in a given region of interest composed of one or more genomic regions. The region of interest may be as large as the entire reference genome (as in the example above) or as small as about 100 kbp.
[0129] For example, in the presence of a pool of tumor cells that are different clones derived from the same final common ancestor tumor cell, the purity can vary widely between different genomic regions, from a minimum value of 1 / number of cells in the pool when the LoH region is represented by only one cell, to a maximum value of 100% when the LoH state of the genomic region is common to all clones derived from the final common ancestor.
[0130] The sample analyzed for LoH preferably has a purity of at least 50%, more preferably at least 70%, as can be seen from FIG. 13 showing the area under the receiver operating characteristic (ROC) curve (AUC) values for LoH scores at various numbers of mapped reads (1,000,000 to 10,000,000 reads) and sample purity (10% to 90%). The LoH score is defined as the number of heterozygous SNPs within a window of n = 150 SNPs covered by at least 2 reads. Samples of different purities are obtained by in silico mixing of reads obtained from the analysis of tumor cells and normal cells at a ratio (tumor:normal) equal to the target purity. The LoH detected in tumor cells by high-pass whole-genome sequencing is used as a reference.
[0131] Effect of Size Selection on LoH Detection As already described above, size selection is preferably performed during or after step c of preparing the ultra-parallel sequencing library. The size of the fragments can be selected according to different criteria. The sequencing method can also be selected according to different criteria, which also depends on the fragment size. Generally, the more loci (polymorphisms or heterozygosities) that contribute to LoH analysis, the better the resolution (per million reads).
[0132] Figure 14 shows data obtained by selecting an in-silico subset of the sequenced fragments from data obtained from an actual single-cell sample in which Fcenter (sequencing library prepared using Ampli1 LowPass for Illumina) is increasing. Figure 14A shows the effect of size selection (bandwidth 100) on the coverage of DRS-WGA fragments with respect to the average fragment length when using 250,000 reads. Figure 14B shows the effect of size selection (bandwidth 100) on resolution with respect to base pairs (window of 150 SNPs covered by at least 2 reads) when using 250,000 reads. Figure 14C shows the effect of size selection bandwidth on the coverage of DRS-WGA fragments by 250,000 reads at a fixed average fragment length (500 bp). Figure 14D shows the effect of size selection bandwidth on resolution (bp) by 250,000 reads at a fixed average fragment length (500 bp). Figure 14E shows the effect of the number of reads on the coverage of DRS-WGA fragments at a fixed average fragment length (500 bp). The fraction of fragments covered by at least 2 reads and the total number of fragments covered increase proportionally to the number of mapped reads (dashed line). Figure 14F shows the effect of the number of reads on resolution (bp) at a fixed average fragment length (500 bp).
[0133] These data show that as the total number of DRS-WGA fragments decreases, on the one hand, the number of fragments covered by one or more reads useful for SNP calling increases and reaches a plateau at 500 bp (Figure 14A). As shown by the decrease in the length of the genomic window by a fixed number of SNPs (n = 150; Figure 14B), the resolution increases accordingly. When different bandwidths are applied to a given number of mapped reads and Fcenter, the fragment coverage and resolution increase as the bandwidth decreases (Figures 14C and 14D). The resolution also increases according to the number of mapped reads (Figures 14E and 14F).
Example
[0134] The following Table 3 summarizes the characteristics of the methods used in the three examples disclosed below.
[0135]
Table 3A
[0136]
Table 3B
[0137] (Example 1) In Example 1, one circulating tumor cell (CTC; test) and one white blood cell (WBC; control) obtained from a male patient with multiple myeloma were considered for the Ampli1 LowPass for Illumina DNA library. The sequenced reads were mapped to the hg19 reference human genome and downsampled at 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, and 9 million reads. Alleles present at dbSNP polymorphic loci (dbSNP150 common variants with minor allele frequency ≥ 5%) were extracted from both libraries. The loci were divided into genomic windows with a 10,000,000 bp fixed size. A one-sided Fisher's exact test was employed to determine the significance of the association between the two classifications (Table 4) under the null hypothesis that heterozygous and homozygous loci are equally likely in WBC (control) and CTC (test).
[0138]
Table 4
[0139] The results of the tests at each downsampling level are shown in Figure 15. This method starting from 2 million reads shows high sensitivity in detecting known LoH events on chromosome 11 and chromosome 113.
[0140] Specifically, Figure 15 shows at the top the copy number plot of CTCs from a patient with multiple myeloma. The x-axis is the chromosome and the y-axis is the copy number. Each dot represents a genomic window of a fixed size. The copy number segments are represented as solid lines. A reference (Ref) track is shown below the copy number plot, and the known LoH regions detected by whole-genome sequencing of the same CTCs are indicated by solid black lines. At the bottom, a track marked from 1M to 9M is shown: a heatmap of the logarithm p-values (base = 10) of the results of the Fisher's exact test at various numbers of reads (from 1 million to 9 million). More significant values are represented in darker gray.
[0141] (Example 2) In Example 2, the same single CTC data used in Example 1 was used as input, and the data was downsampled to 1 million reads. In this case, the loci were divided into windows having a fixed number (n = 1000) of loci covered by at least 1 read. For the identification of the LoH region, the LoH score was calculated as the number of heterozygous positions in each window.
[0142] Figure 16 shows the detection of LoH by using genomic windows having a constant number of loci. Specifically, the same copy number plot of CTC in Example 1 is shown at the top. The x-axis is the chromosome and the y-axis is the copy number. Each dot represents a genomic window of a fixed size. The copy number segments are represented as solid lines. Below the plot is a heatmap representing the heterozygosity count for each genomic window. Windows with a lower LoH score (fewer heterozygous loci), which are more likely to be in the LoH state, are represented by a darker gray. Chromosome 11, the large arm of chromosome 13, and the X chromosome (single copy in male individuals) show a low LoH score.
[0143] To determine the LoH score threshold for calling genomic windows in the LoH state, a training set of 9 single cells with known LoH regions was analyzed using the same method as the test sample (1,000,000 mapped reads and windows of n = 1000 SNPs). Subsequently, ROC analysis was performed, and a maximum LoH score threshold = 6 was determined as the best trade-off point between sensitivity and specificity (Figure 17, where the x-axis represents 1 - specificity (a smaller value means more specific detection), and the y-axis represents sensitivity. LoH detected in tumor cells by high-pass whole-genome sequencing was used as a reference).
[0144] By this method, LoH events on chromosomes 11 and 13 were successfully identified. The LoH status was also assigned to the X chromosome as expected in male individuals whose genome contains a single copy of the X chromosome (Figure 18, regions where the LoH score is below the fixed threshold (≤6) and larger than 10,000,000 bp are shown in black).
[0145] (Example 3) In Example 3, the Ampli1 LowPass for Illumina libraries of two Hodgkin Reed / Sternberg (HRS) single cells obtained from FFPE tissues of classical Hodgkin lymphoma samples from male patients were analyzed. The two HRS cells share the same copy number profile. Sequenced reads were mapped to the hg19 reference human genome, and alleles present at dbSNP polymorphic loci (dbSNP150 common variants with minor allele frequency ≥5%) were extracted from both libraries. Loci were segmented using copy number segments obtained using Control-FREEC software implementing GC-based normalization and copy number signal segmentation (Boeva, V. et al., Bioinformatics, 27(2), pp. 268 - 269, http: / / doi.org / 10.1093 / bioinformatics / btq635). An internal control defined by the union of all regions having a copy number equal to the ploidy of the cells (copy number = 2) was used. For each segment contained in a chromosomal arm as defined by the copy number analysis, a one-sided Fisher's exact test was performed to reject the null hypothesis that the observed two-allele and one-allele loci are equally likely to be present in the segment and the internal control (Figure 19, top: one representative HRS cell copy number profile. Bottom: -log of the p-value obtained as the output of the Fisher's test 10Heatmap. Only genomic regions with p-value ≤ 0.01 are shown. More significant values are represented by darker gray). As expected, all regions with copy number = 1 were correctly detected as LoH genomic regions. Despite having a copy number = 2, the long arm of the X chromosome was detected in the LoH state. This is presumably because the sample is from a male individual and thus the genome contains a single X chromosome. In addition, chromosome 9q was called as LoH, which is thought to be missed by using only copy number information (copy number = 2).
[0146] Advantages The method according to the present invention is suitable for analyzing data obtained from low-pass sequencing of genomic DNA from a test sample to detect LoH events. In contrast to other methods that need to infer LoH as a series of consecutive homozygous loci and extract the actual genotypes at a specific number of loci, the method of the present invention analyzes genomic windows containing a sufficient number of loci sequenced at low coverage and extracts the alleles observed at said loci that do not necessarily represent the genotype of the sample. Based on the principle that it is possible to detect LoH events as a reduction in the number of two-allele loci compared to what is observed by analyzing a normal diploid sample.
[0147] In contrast to other methods that infer LoH from alternative allele frequencies (B allele frequency or BAF) and require high coverage of the genome, e.g., 30-fold (Boeva et al., Bioinformatics, Vol. 28 no. 3 (2012), 423-42), the method according to the present invention operates on low-pass whole-genome sequencing data (less than 1-fold, or less, e.g., 0.05-fold or even 0.01-fold), with a corresponding cost reduction.
[0148] The method for analyzing LoH from a sample according to the present invention enables the inference of genome-wide LoH regions from low-pass whole-genome sequencing data with single-cell resolution using very small samples, such as when only a small number (up to single) of CTCs are available, and as an additional optional possibility, the analysis can also be performed without normal controls and with a relatively small number of reads.
[0149] Furthermore, a particular embodiment of the method can enhance the resolution of LoH calling by introducing specific processing steps into the library preparation process without incurring additional sequencing costs.
[0150] The method according to the present invention has capabilities that were previously considered unattainable by those skilled in the art and represents a startling advancement over the current state of the art. In particular, the method enables the following: - Identifying LoH on single cells by low-pass whole-genome sequencing with a low average coverage of 0.01 to 0.04 (250,000 to 1,000,000 single-end 150bp reads of the human genome); - Obtaining the above points without a control sample; - Obtaining the above points with the future possibility of obtaining additional genetic material for investigating other characteristics of the single cell, and furthermore, with the possibility of reliably re-analyzing the single cell for verification through the use of WGA inherent in the process.
[0151] In addition, the method according to the present invention enables the determination of the whole-genome copy number profile and LoH even from trace amounts of cells, FFPE, or tissue biopsy samples.
Claims
**Claim 1** A method for analyzing loss of heterozygosity (LoH) in at least one sample containing genomic DNA, comprising: b. performing whole-genome amplification of defined restriction sites (DRS-WGA) of said genomic DNA; c. preparing a massively parallel sequencing library from the product of said DRS-WGA; d. performing low-pass whole-genome sequencing on said massively parallel sequencing library at an average coverage depth of less than 1; e. aligning the reads obtained in step d onto a reference genome for said at least one sample; f. extracting allele content at a plurality of loci, said plurality of loci including polymorphic loci and / or heterozygous loci; g. assigning an LoH score as a function of the number of loci having at least two different alleles at said plurality of loci for at least one genomic window of said reference genome for said at least one sample. A method comprising the above steps. **Claim 2** The method according to claim 1, further comprising a size selection step of selecting fragments of the product of said DRS-WGA based on fragment length, said size selection step being performed before, during, or after said step c of preparing a massively parallel sequencing library, and said step of preparing a massively parallel sequencing library not including a random fragmentation step. **Claim 3** The method according to claim 2, wherein said size selection step retains fragments in the range of 100 to 800 base pairs. **Claim 4** The method according to claim 3, wherein said size selection step retains fragments in the range of 300 to 450 base pairs. **Claim 5** The method according to claim 3, wherein the peak of the fragments retained in said size selection step is centered on a base pair range of 150 bp to 600 bp. **Claim 6** The method according to claim 5, wherein said size selection step retains fragments in the range of 425 to 575 base pairs. **Claim 7** The method according to any one of claims 1 to 6, wherein said at least one genomic window has a constant width in base pairs. **Claim 8** The method according to any one of claims 1 to 6, wherein said at least one genomic window has a constant number of said plurality of loci. **Claim 9** The method according to any one of claims 1 to 6, wherein the at least one genomic window is selected from the group consisting of a chromosome, a chromosomal arm, and a segmented copy number region.
10. The method according to any one of claims 1 to 9, wherein the plurality of loci include polymorphic loci with respect to a reference genome for the at least one sample.
11. The method according to claim 7, 8 or 10, wherein the LoH score corresponds to the number of heterozygous loci in the at least one genomic window.
12. The method according to claim 10, wherein the LoH score corresponds to the ratio of heterozygous loci to the total number of polymorphic loci in at least one genomic window.
13. The method according to claim 10, wherein the LoH score corresponds to the p-value of a statistical test.
14. The method according to claim 13, wherein the statistical test determines the significance of overrepresentation of two-allele loci relative to the error rates of sequencing and WGA.
15. The method according to claim 13, wherein the statistical test determines the significance of underrepresentation of two-allele loci relative to a control sample.
16. The method according to claim 15, wherein the control sample includes at least one genomic region at the major ploidy level from the at least one sample.
17. The method according to claim 15, wherein the control sample is at least one normal sample, and the normal sample is a non-tumor sample.
18. The method according to claim 17, wherein the at least one normal sample is obtained from the same individual under the test from which the at least one sample is obtained.
19. The method according to claim 15, wherein the control sample is a maternal sample or a paternal sample for the at least one sample.
20. When the LoH score for a genomic window is less than a threshold value, the genomic window is called to be in LoH. The method according to any one of claims 11 to 13.
21. The method according to claim 20, further comprising the step of assigning an LoH state to the region when the LoH score for each genomic window included in at least one genomic region is less than the threshold value.
22. The method according to claim 20, further comprising assigning a LoH state to the region when more than a given percentage of the genomic windows included in at least one genomic region exhibit a LoH score smaller than the threshold value.
23. The method according to claim 21 or 22, wherein the at least one genomic region contains a tumor suppressor gene.
24. The tumor suppressor gene is a. BRCA1 b. BRCA2 c. PALB2 d. TP53 e. CDKN2A f. RB1 g. APC h. PTEN i. CDKN1B j. DMP1 k. NF1 l. AML1 m. EGR1 n. TGFBR1 o. TGFBR2 p. SMAD4 and is selected from the group consisting of, the method according to claim 23.
25. The method according to any one of claims 1 to 24, wherein the at least one sample is a single cell.
26. The method according to any one of claims 1 to 25, wherein in step d, low-pass whole-genome sequencing is performed at an average coverage depth of less than 0.05 on the ultra-parallel sequencing library.
27. The method according to any one of claims 1 to 26, wherein in step d, low-pass whole-genome sequencing is performed at an average coverage depth of less than 0.01 on the ultra-parallel sequencing library.
Citation Information
Patent Citations
Method and kit for generating DNA libraries for massively parallel sequencing
JP2019514360A
Methods for identifying DNA copy number changes
US7424368B2
DNA amplification of a single cell
WO2000017390A1
Method and kit for the generation of DNA libraries for massively parallel sequencing
WO2017178655A1
Improved method and kit for the generation of DNA libraries for massively parallel sequencing
WO2019016401A1