Method for analysing the copy-number profile of a sample using deterministic restriction-site whole genome amplification (DRS-WGA)

The method enhances DNA copy-number profiling by using a fragmentation-free library preparation and normalization techniques to improve resolution and accuracy in DRS-WGA data analysis, addressing limitations in existing technologies.

WO2026083244A1PCT designated stage Publication Date: 2026-04-23MENARINI SILICON BIOSYSTEMS SPA
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MENARINI SILICON BIOSYSTEMS SPA
Filing Date
2025-10-14
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for DNA copy-number profiling from low-pass whole genome sequencing data generated by deterministic restriction-site whole genome amplification (DRS-WGA) face limitations in resolution and accuracy, with marginal improvements beyond a certain data threshold, necessitating a more efficient analysis method.

Method used

A method involving fragmentation-free, sequencing-adaptor/WGA fusion-primer PCR reaction for library preparation, combined with binning and normalization techniques to compensate for amplification and sequencing biases, using in-silico expected fragments and normalized frequencies to enhance copy-number calling accuracy.

Benefits of technology

Improves the resolution and accuracy of DNA copy-number profiling by compensating for sequencing and amplification biases, providing a more precise estimation of copy-number variations and alterations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000038_0000
    Figure 00000038_0000
  • Figure 00000039_0000
    Figure 00000039_0000
  • Figure 00000040_0000
    Figure 00000040_0000
Patent Text Reader

Abstract

The present invention relates to a method for DNA copy- number profiling of a sample comprising genomic DNA First, the genomic DNA is amplified using a deterministic restriction-site whole genome amplification (DRS-WGA). Secondly, an aliquot of said DRS-WGA product is used to prepare a massively-parallel sequencing library with a fragmentation-free, sequencing-adaptor / WGA fusion-primer PCR reaction. Said library is then sequenced to obtain low- pass whole genome sequencing data. The sequencing reads, are finally normalized according to the invention in a way that is cognizant of the DRS-WGA and the sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] "METHOD FOR ANALYSING THE COPY-NUMBER PROFILE OF A SAMPLE

[0002] USING DETERMINISTIC RESTRICTION-SITE WHOLE GENOME AMPLIFICATION (DRS-WGA) "

[0003] Cross-Reference to Related Applications

[0004] This Patent Application claims priority from Italian Patent Application No . 102024000022833 filed on October 14 , 2024 , the entire disclosure of which is incorporated herein by reference .

[0005] Technical Field

[0006] The present disclosure relates to a method for DNA copynumber profiling of a sample , by analyzing data obtained by low-pass whole genome sequencing carried out on said sample .

[0007] The method according to the present disclosure can be used in several fields of application, including, but not limited to :

[0008] • Circulating Tumor Cell copy-number profiling

[0009] • Circulating Fetal Cells genome-wide copy-number profiling, and / or detection of microdeletions / microduplications

[0010] Prior Art

[0011] Whole Genome Ampli fication (WGA) of single cell genomic DNA is often required for obtaining more DNA in order to simpli fy and / or allow di f ferent types of genetic analyses , including sequencing, SNP detection etc . WGA with a Ligation- Mediated PCR ( LM-PCR) based on a Deterministic Restriction Site ( in the following DRS-WGA) is known from W02000 / 017390 .

[0012] DRS-WGA has been shown to be the best-in-class WGA method in many perspectives , in particular in terms of lower allelic drop-out from single cells (Borgstrom et al . , 2017 ; Normand et al . , 2016 ; Babayan et al . , 2016 ; Binder et al . , 2014) .

[0013] A LM-PCR based, DRS-WGA commercial kit (Amplil™ WGA kit, Silicon Biosystems) has been used in Hodgkinson C.L. et al., Nature Medicine 20, 897-903 (2014) . In this work, a Copy-Number Analysis by low-pass whole genome sequencing on single-cell WGA material was performed, carrying out digestion of the WGA adaptors and fragmentation prior to Illumina barcoded adaptor ligation for sequencing.

[0014] WO2017 / 178655 and W02019 / 016401A1 teach a simplified method to prepare massively parallel sequencing libraries from DRS-WGA (e.g. Amplil WGA) for low-pass whole genome sequencing and copy number profiling.

[0015] In Ferrarini et al., PLoSONE 13 (3) :e0193689 https: / / doi.org / 10.1371 / journal.pone.0193689, the method performance of W02017 / 178655 using the Ion Torrent Platform has been detailed with reference to copy number profiling, using Control-FREEC algorithm (Boeva et al., Bioinformatics.

[0016] 2011 Jan 15; 27 (2) : 268-269, doi: 10.1093 / bioinformatics / btq635) .

[0017] DRS-WGA has been shown to be better than DOP-PCR (Degenerate oligonucleotide-primed PGR) for the analysis of copy-number profiles from minute amounts of microdissected (formalin-fixed paraffin-embedded) FFPE material (Stoecklein et al., Am J Pathol. 2002 Jul; 161 ( 1 ) : 43-51 ; Arneson et al., ISRN Oncol. 2012;2012 : 710692. doi: 10.5402 / 2012 / 710692. Epub

[0018] 2012 Mar 14.) , when using array CGH (Comparative Genomic Hybridization) , metaphase CGH, as well as for other genetic analysis assays such as Loss of heterozygosity using targeted primers and PCR for analysis of selected microsatellites, however, it has been shown that depending on FFPE DNA quality, single-cell FFPE LP-WGS is possible but may become impractical for lower DNA quality scores (Mangano, C., Ferrarini, A., Forcato, C. et al. "Precise detection of genomic imbalances at single-cell resolution reveals intrapatient heterogeneity in Hodgkin's lymphoma". Blood Cancer J. 9, 92 (2019) . https: / / doi.org / 10.1038 / s41408-019-0256-y) .

[0019] In the present disclosure the term DRS-WGA also refers to the error-free sequencing of DNA method disclosed in W02015118077A1, also based on LM-PCR and similar to W02000 / 017390 . In both methods the DNA is digested with a sequence-specific restriction enzyme and a primer is ligated for the amplification, but - in W02015118077A1 - the ligated primer also contains a randomized sequence which enables error-free sequencing by introducing Unique Molecular Identifiers (UMI) in each fragment of the initial template digest, prior to the WGA primary PGR.

[0020] Both W02015118077A1 and W02000 / 017390 rely on "digesting the DNA to be amplified with a restriction endonuclease under conditions suitable to obtain DNA fragments of similar length".

[0021] In the context of the present disclosure, it is intended that the restriction endonuclease (RE) cleaves the DNA on, or at a substantially predictable distance from, the recognition sequence site. For example the RE is a Type II restriction endonuclease, as this property - combined with the knowledge of the sequence of the reference genome - allows to predict the sequence of DNA fragments obtained by the digestion of the DNA from the RE, and the resulting distribution of WGA fragment lengths (because the length of the universal WGA primers used in the whole genome amplification are known and can be added appropriately, i.e. any WGA fragment length in base pairs will be equal to the original un-ampli fied DNA restriction fragment plus twice the WGA universal primer base-pair length) .

[0022] The need is felt for a method for further improving the resolution and accuracy of DNA copy-number profiling from DRS-WGA products . The library preparation methods taught in W02017 / 178655 and W02019 / 016401 provide a suitable method for genome-wide copy-number DNA profiling with several advantages linked to ease of use and the possibility to carry out genome-wide LoH ( Loss of Heterozygosity) analysis (WO2021019459 ) and sample identi fication from the low-pass Whole Genome Sequencing data of paired samples (WO2023 / 042173 ) .

[0023] In general , the resolution ( intended as the minimum si ze measured in base-pairs for which copy-number alterations can be reliably detected) and accuracy of DNA Copy-Number profi ling increases with the amount of reads of the low-pass whole genome sequencing data, however the cost also increases , and anyway the marginal improvement vanishes ( saturation of the resolution and accuracy) above a certain number . It is thus desirable to improve on the analysis method by providing a bioinf ormatic pipeline that , with the same data (hence same cost devoted to generate the sequencing data ) , provides an improved resolution and accuracy obtained from the sequence data obtained using methods known in the prior-art such as Control-FREEC and the method taught in the above cited Ferrarini et al . , PLoSONE 13 ( 3 ) : e0193689 , which is based on the same Control-FREEC .

[0024] Summary

[0025] It is therefore an obj ect of the present invention to provide a method for DNA copy-number profiling which improves over the prior art methods in terms of copy-number calling from low-pass whole genome sequencing data generated by sequencing a massively-parallel sequencing library prepared with a fragmentation-free, sequencing-adaptor / WGA fusionprimer PCR reaction starting from a deterministic restriction-site whole genome amplification (DRS-WGA) of a s amp 1 e .

[0026] This object is achieved by the method as defined in claim 1.

[0027] Brief Description of the Drawings

[0028] Figs. 1A and IB show a whole genome copy number analysis workflow based either on partitioning of the reference genome in bins of variable length or on partitioning of the reference genome in bins of fixed length.

[0029] Figs. 2A-2C shows the single and combined relationship among the "GC" and "weight" factors and the read counts generated by the sequencing of a NGS library obtained from the DRS-WGA amplification (Msel) of DNA deriving from a single cell isolated from a Coriell GM17942 cell culture.

[0030] Fig. 3 shows quality statistics of copy number profiles generated from the sequencing of NGS libraries obtained from the amplification of (genomic DNA) gDNA from 4 tumor cell lines (NCI-H929, SK-MM2, JJN3, EJM) with 3 restriction enzymes (Bfal, Csp6I, Msel) and different normalization methods (gcw, gc, w) .

[0031] Fig. 4 shows quality statistics of copy number profiles generated from the sequencing of NGS libraries obtained from the amplification by DRS-WGA of gDNA from 4 tumor cell lines (NCI-H929, SK-MM2, JJN3, EJM) employing 3 distinct restriction enzymes (Bfal, Csp6I, Msel) , recognizing 3 different restriction site sequences on the gDNA sequence (CATAG, GATAC, TATAA) . Rows in the figure represent the different statistics calculated, while columns represent different cell lines, as indicated in the title of each bar plot .

[0032] Fig. 5 shows quality statistics of copy number profiles generated from the sequencing of NGS libraries obtained from the amplification of single cell DNA from NA12878 reference cell line with 3 restriction enzymes (Bfal, Csp6I, Msel) and different normalization methods (gcw, gc, w) .

[0033] Fig. 6 shows quality statistics on chromosome 1 (first row) and on chromosome 19 (second row) for profiles obtained from gDNA of 4 tumor cell lines (NCI-H929, SK-MM2, JJN3, EJM) .

[0034] Fig. 7 shows quality statistics on chromosome 1 (first row) and on chromosome 19 (second row) for profiles obtained from single cell DNA of NA12878 reference cell line.

[0035] Figs. 8A to 8C show the effect of the normalization method employed on 3 different chromosomes on copy number profiles generated by sequencing DNA libraries built from NCI-H929 cell line gDNA amplified by DRS-WGA employing the Bfal restriction enzyme for the gDNA fragmentation step.

[0036] Fig. 9 shows the effect of the normalization method employed on copy number profiles of chromosome 19 generated by sequencing DNA libraries built from NCI-H929 cell line gDNA amplified by DRS-WGA employing the Bfal restriction enzyme for the gDNA fragmentation step.

[0037] Figs. 10A-10C shows representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from NCI-H929 cell line gDNA amplified by DRS-WGA with Bfal restriction enzyme.

[0038] Figs. 11A-11C shows representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from NCI-H929 cell line gDNA amplified by DRS-WGA with Csp6I restriction enzyme.

[0039] Fig. 12A-12C shows representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from NCI-H929 cell line gDNA amplified by DRS-WGA with Msel restriction enzyme.

[0040] Fig. 13A-13C shows representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from EJM cell line gDNA amplified by DRS-WGA with Bfal restriction enzyme .

[0041] Fig. 14A-14C shows representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from JJN3 cell line gDNA amplified by DRS-WGA with Bfal restriction enzyme.

[0042] Fig. 15A-15C shows representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from SK- MM2 cell line gDNA amplified by DRS-WGA with Bfal restriction enzyme .

[0043] Definitions

[0044] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Although many methods and materials similar or equivalent to those described herein may be used in the practice or testing of the present disclosure, preferred methods and materials are described below. Unless mentioned otherwise, the techniques described herein for use with the present disclosure are standard methodologies well known to persons of ordinary skill in the art.

[0045] Copy-number variations (CNVs) is a general term used to describe a molecular phenomenon in which sequences of the genome are repeated, and the number of repeats varies between individuals of the same species.

[0046] Copy number alterations (CNAs) are a predominant source of genetic alterations in human cancer and play an important role in cancer progression. The term "copy-number variations" (CNVs) and "copy-number alterations" (CNAs) are used interchangeably in the description. In fact, the main difference in their routine use in scientific literature is that CNVs is mostly used to indicate germline copy-number changes, whereas CNAs is mostly used to indicate somatic copy-number changes, so that the most appropriate term depends substantially on the application of the method to a specific context (e.g. prenatal diagnosis versus oncology diagnostics) . To generalize, the term "copy-number profiling" is used to refer to the analysis of any of CNVs and CNAs and simplify the reading.

[0047] By the expression "extrachromosomal DNA" (ecDNA) there is intended any DNA that is found off the chromosomes, either inside or outside the nucleus of a cell. In fact, copy number alterations linked to oncogenes often occur in extrachromosomal DNA (e.g. Turner et al., Nature. 2017 Mar 2; 543 (7643) : 122-125, doi : 10.1038 / nature21356 "Extrachromosomal oncogene amplification drives tumor evolution and genetic heterogeneity") and are of particular interest in oncology. While in general scienti fic literature the expression "genomic DNA" normally refers to chromosomal DNA, for the sake of simplicity, in the present disclosure , with the expression "genomic DNA" there is intended any DNA that is found in the cell , including chromosomal DNA and extrachromosomal DNA which maps to the same reference genome as chromosomal DNA . In fact , regardless of its original location within a cell , for certain samples , such as cell- free DNA, it is not possible to ascertain where a DNA fragment may derive from .

[0048] By the expression "massive-parallel next generation sequencing (NGS or MPS ) " there is intended a method of sequencing DNA comprising the creation of a library of DNA molecules separated spatially and / or in time , clonally sequenced (with or without prior clonal ampli fication) . Examples include the I llumina platform ( I llumina Inc ) , the Ion Torrent platform ( Thermo Fisher Scienti fic Inc ) , the Paci fic Biosciences platform, the Minion ( Oxford Nanopore Technologies Ltd) .

[0049] By the expression " low-pass whole genome sequencing" there is intended a whole genome sequencing at mean sequencing depth lower than lx with reference to the entire Reference Genome , of a massively parallel sequencing library which has not been enriched for sequence-speci fic fragments . This definition explicitly excludes the case of PCR-based target enrichment or sequence-speci fic capture-baits target enrichment for a set of loci , such as for example Single- Nucleotide Polymorphisms ( SNPs ) and / or Short-Tandem Repeats ( STR) loci .

[0050] By the expression "mean sequencing depth" there is intended here , on a per-sample basis , the total number of bases sequenced, mapped to the reference genome , divided by the total reference genome si ze . The total number of bases sequenced and mapped can be approximated to the number of mapped reads time the average read length .

[0051] By the expression "reference genome" there is intended a reference DNA sequence for the speci fic species .

[0052] By the expression " in-silico expected fragments" there is intended a li st of non-overlapping subsequences of the reference genome obtained by the digestion of the reference genome with the restriction endonuclease employed for the DRS-WGA reaction .

[0053] By the term " genomic coordinate" there is intended a fixed position on a chromosome ( relative to the reference genome ) .

[0054] By the expression "covered genome" there is intended the portion of re ference genome covered by at least one read .

[0055] By the term "read" there is intended the piece of DNA that is sequenced ("read" ) by the sequencer .

[0056] By the expression " fragmentation- free, sequencing- adaptor / WGA fusion-primer and PGR reaction" massively- parallel sequencing library preparation, there is intended a massively-parallel sequencing library preparation on DRS- WGA products , without DNA fragmentation steps , whereby sequencing adaptors are added to the WGA product by fusion primers , e . g . according to patent applications ( W02017 / 178655 ) or ( W02019 / 016401A1 ) .

[0057] By the expression "cell-based non-invasive prenatal testing" there is intended carrying out genetic assays in order to evaluate fetal cells circulating in maternal blood .

[0058] By the expression "pre-implantation genetic testing / screening" there is intended carrying out genetic assays in order to evaluate embryos before trans fer to the uterus by genome-wide analysis of , for example , copy-number variations for determining the presence of aneuploidy ( either too many or too few chromosomes ) in a developing embryo .

[0059] By the expression "embryonic sample" , there is intended a sample containing DNA from an embryo , such as for example a blastocyst , a spent embryo-culture medium, a polar body .

[0060] Detailed Description

[0061] The method according to the present disclosure is aimed at the copy-number profiling of a sample comprising DNA, using reads , aligned to a reference sequence , obtained from low-pass whole genome sequencing data generated by sequencing a massively-parallel sequencing library prepared with a fragmentation- free , sequencing-adaptor / WGA fusionprimer PCR reaction starting from a deterministic restriction-site whole genome ampli fication ( DRS-WGA) of said s amp 1 e .

[0062] In a preferred embodiment the sample species is Homo sapiens , and unless otherwise noted this species will be referred to in the rest of the description, without limitation to the applicability to other species .

[0063] The method for DNA copy-number profiling of a sample comprising genomic DNA according to the present invention comprises the fol lowing steps which are schematically shown in Figures 1A and IB .

[0064] In step a ) , the sample comprising genomic DNA is ampli fied using a deterministic restriction-site whole genome ampli fication ( DRS-WGA) .

[0065] In step b ) a massively-parallel sequencing library is prepared from the ampli fied sample of step a ) with a fragmentation- free , sequencing-adaptor / WGA fusion-primer PCR reaction .

[0066] In step c ) , the massively-parallel sequencing library prepared in step b ) is sequenced by low-pass whole genome sequencing .

[0067] Low-pass whole genome sequencing i s preferably carried out at a coverage > O . O lx, more preferably at a coverage > 0 . 05x, even more preferably at a coverage > O . lx, even more preferably at a coverage > 0 . 5x .

[0068] In step d) the sequencing reads obtained in step c ) are aligned to a reference genome sequence .

[0069] In step e ) in-silico expected fragments of said DRS-WGA are computed as a list of non-overlapping subsequences of the reference genome sequence each starting and ending at the recognition sequence of the restriction enzyme used in said DRS-WGA.

[0070] The method according to the present invention advantageously further comprises the following steps .

[0071] In step f ) the number of read counts mapping to each in-silico expected fragment sequenced in step c ) is computed .

[0072] In step g) an expected normali zed frequency of each in- silico expected fragment of said reference genome sequence is computed as the sum of read counts for each in-silico expected fragment length divided by the total number of reads mapped to said reference genome sequence . More speci fically, the expected normali zed frequency of each in-silico expected fragment of said reference genome sequence is computed from the actual distribution of sequencing data, and is thus factoring in several factors which may skew the representation of DNA in the actual sequencing data, linked for example to ef ficiency of ampli fication during WGA, library preparation (e.g. effects of purifications and PCR amplification) and sequencing (e.g. bridge amplification and cluster formation) . The expected normalized frequency of a given fragment with a given length is computed by interrogating a look-up table comprising the frequency of reads mapped to said reference genome sequence as a function of fragments length.

[0073] If the read is a single-end read (not a paired-end read) , the corresponding DRS-WGA fragment length is computed determining the expected in silico fragment length for the DRS-WGA associated with that.

[0074] For example, this is the case when using the Amplil WGA kit in combination with a library preparation with Amplil LowPass kit for Ion Torrent (Menarini Silicon Biosystems SpA) , or with the Amplil LowPass kit for Illumina and sequencing with Illumina single-ended read chemistry (all Amplil kits are from Menarini Silicon Biosystems SpA) .

[0075] Conversely, when using the Amplil WGA kit in combination with a library preparation with the Amplil LowPass kit for Illumina and sequencing on Illumina with a paired-end reads chemistry, the actual read length may be computed. This may be different from the in-silico expected fragment length due to: undigested restriction sites or by the presence of SNPs in the DNA which alter the reference sequence by introducing (or removing) a restriction site.

[0076] However, since the number of undigested restriction sites or SNP introducing / removing restriction sites is normally below 15%, preferably below 10%, even more preferably below 5%, the first approach -neglecting the actual fragment length and relying only on the in-silico expected fragment length- is sufficiently precise. In step h) said reference genome sequence is partitioned in genomic bins comprising non-overlapping adj acent subsequences of said reference genome , either

[0077] I ) by determining fixed-length genomic bins , each bin comprising adj acent in-silico expected fragments so that the sum of said in-silico expected fragment lengths is substantially equivalent to a predefined number of basepairs sum-target value ( such as , by way of non-limiting example , 50kbp, l O Okbp, 500kbp, IMbp, 5Mbp ) , which can be chosen by the user depending on the application needs in terms of genomic resolution; or

[0078] I I ) by determining variable-length genomic bins , each bin comprising adj acent in-silico expected fragments so that the sum of expected normali zed frequencies of said in-silico expected fragments is substantially equivalent to a predefined sum-target value .

[0079] In a preferred embodiment , the predefined sum-target value of said in-silico expected fragments frequencies is determined as a predefined number of reads per bin ( such as , by way of non-limiting example, 50 reads , 100 reads , 200 reads , 300 reads ) divided by the total number of reads per sample . The higher the predefined number of reads per bin, the lower the stochastic variation of read counts from bin to bin, the lower the noise in copy-number estimate .

[0080] In fact , the number of reads per bin can be modeled as a Poisson distribution . Accordingly, the coef ficient of variation ( CV) of the number of reads per bin is equal to the inverse of the square root of the mean number of reads per bin . The resolution in base pairs across the genome will vary using a variable-length bin, however the mean value of genomic bin size can be predicted considering the genome size and a sum-target value. For example, a mean window size » IMbp may be expected when a genome of 3.2Gbp=3200Mbp (such as the human genome) is analyzed with a sum-target value of 155 reads per window / 500, 000 total reads = 3. l*10~4. In fact 3.1*10-4*3200 Mbp=0.99 Mbp .

[0081] Generalizing,

[0082] Target_Mean_Window_Size=Sum_Target_Value * Genome_Size or alternatively,

[0083] Sum_Target_Value= Target_Mean_Window_Size / Genome_Size where Target_Mean_Window_Size is the desired variablelength bin mean size and Genome_Size is the reference genome size, both expressed in basepairs.

[0084] As certain regions of the genome may be effectively hard to sequence due to poor mappability (e.g. low-complexity or high-repeat regions) or other non idealities (e.g. high GC content regions) , a more accurate estimate is computed substituting the Genome_size in the above formulas with Sequenceable_Genome_size, i.e. the portion of reference genome which can be unambiguously mapped with sequencers used. For human sequencing on short-read sequencers such as Illumina, a preferred value is

[0085] Sequenceable_Genome_size~2 .9Gbp .

[0086] In step i ) ,

[0087] I) In the case of a fixed length genomic bins, for each genomic bin:

[0088] 1) a bin-GC is computed as the mean GC content of the bin;

[0089] 2) a bin-weight is computed as the sum of normalized frequencies of in-silico expected fragments overlapping the bin;

[0090] 3) a bin-count is computed as the sum of read counts mapping to the bin;

[0091] 4) a normalized bin-count is computed as a function of bin-weight and bin-GC.

[0092] II) In the case of a variable length genomic bins, for each genomic bin:

[0093] 1) a bin-GC is computed as the mean GC content of the bin;

[0094] 2) a bin-count is computed as the sum of read counts mapping to the bin;

[0095] 3) a normalized bin-count is computed as a function of the bin-GC.

[0096] In step j) , a bin-ratio is computed as the normalized bin-count divided by the median normalized bin-count.

[0097] Genomic coordinate to fragment-length univocal relationship in DRS-WGA

[0098] The method according to the present invention exploits the fact that in DRS-WGA, such as the Amplil™ WGA, each genomic coordinate in the genome is represented in the WGA library only in fragments having a specific length in basepairs. Considering a general genomic coordinate, said genomic coordinate will be represented only in a fragment of a given length, equal to the size of the corresponding fragment following digestion by the restriction enzyme, plus double the length of the universal WGA adaptors (the length of the LIBI primer in case of Ampli l WGA) . When the WGA is sequenced following library preparation according to Ampli l LowPass kits , a predictable additional length is introduced linked to the sequencing adaptors and barcodes lengths , which are known .

[0099] Reproducibility

[0100] In the method according to the present invention, the property of DRS-WGA combined with the fragmentation- free library preparation is exploited to produce a predictable representation of the genome , whereby from the low-pass sequencing data, one can determine the probability to find certain fragments in a given genomic window . Thi s is a key property and advantage with respect to when a random process is inherent in the WGA ( e . g . as with WGA methods using Multiple Displacement Ampli fication or DOP-PCR) and / or in the sequencing library preparation ( e . g . by random fragmentation or tagmentation) .

[0101] In other words , a deterministic fragmentation of the reference genome occurs in the WGA, and predictability is maintained in the sequencing library preparation, obtaining a deterministic relationship between original DNA coordinates within the reference genome and the length of sequencing library DNA fragments into which it can be found in the sequencing data .

[0102] There is intended that the required predictability is based on the DNA sequence of the reference genome , so that restriction enzymes which are substantially insensitive to e . g . methylation ( or other DNA modi fications ) are preferably adopted in the DRS-WGA.

[0103] It is worth noting that the approach is flexible in that di f ferent deterministic restriction enzymes may be suitable depending on the desired resolution and / or sequencing platform and sequencing protocol used. For example, different frequent cutters may be used. In the examples of Amplil WGA, the TTAA motif is the Restriction Site (and Msel is the restriction enzyme) . Other four-base cutters may be used to cut at different Restriction Site, such as GTAC (e.g. using Csp6I as restriction enzyme) , CTAG (e.g. using Bfal as restriction enzyme) , obtaining a different distribution of fragments. By way of example, alternative restriction enzymes recognizing longer (e.g. 6 bases) motifs may be used.

[0104] Compensation of read length distribution

[0105] It is noteworthy that despite the deterministic fragmentation, the representation of each fragment in the sequencing reads may change due to amplification biases in WGA or library amplification, or in the sequencing method (e.g. bridge amplification and cluster formation in Illumina) .

[0106] In the method according to the present invention, a further optimization of the accuracy in copy-number calling is obtained by binning the genome and normalizing the bin count for each genomic region. Preferably, the normalization is carried out considering the actual distribution of fragments in the sequencing data. In fact, as commented above, depending on the characteristics and biases of the sequencing platform, the library preparation method and the WGA amplification efficiency, the distribution of fragments in the sequencing data can vary, in the sense that across different library prep or sequencing runs, and across different samples (e.g. containing more or less input material, with higher or lower genome integrity, or due to fixation which may hamper the availability and ampli fication of longer fragments ) certain bp lengths may be more or less represented . Nevertheless , according to the present invention, as each genomic bin considered for copy-number profiling is composed of fragments of deterministic, known lengths , those variations can be -at least in part- compensated using the information of the sample-speci fic observed distribution of DRS-WGA fragment lengths corresponding to the mapped sequenced reads , reducing the noisiness of the ensuing copy-number signal .

[0107] Indeed, even when using the same WGA type , Library Prep kit method and sequencer, the distribution of fragments may still change , due for example to lot-to-lot variations .

[0108] Or even when considering the same lots of reagents , the inevitable variability of the process may introduce di f ferences in the distribution of fragments sequenced .

[0109] However, all these variations can be advantageously compensated by first analyzing that distribution, after mapping the reads to the reference genome .

[0110] Fixed length bins

[0111] In a preferred embodiment each bin has a fixed length . The mean GC content of the bin is calculated (bin-GC ) , along with a bin count as the sum of read counts mapping to the bin . Based on the empirical distribution of fragment lengths in the analyzed sample , a bin-weight is computed as the sum of normali zed frequencies of in-silico expected fragments overlapping the bin . A normali zed bin count is then obtained as a function of bin-weight and bin-GC . In a preferred embodiment with fixed-length bins , the normali zation function is a tri-dimensional LOWESS function .

[0112] Variable length bins In another preferred embodiment each bin has a variable length, each bin comprising adj acent in-silico expected fragments so that the sum of expected normali zed frequencies , based on the distribution of fragment lengths computed from the actual sequencing data obtained from the sample, of said in-silico expected fragments is substantially equivalent to a predefined sum-target value , for each variable-length bin . The mean GC content of the bin is then calculated . Then, for each bin, a bin count is computed as the sum of read counts mapping to the bin . A normali zed bin count is then obtained as a function of bin-GC . In a preferred embodiment with variable-length bins , the normali zation function is a LOWESS function .

[0113] Bin ratio calculation

[0114] Both for fixed length and variable length bins , a bin ratio is then computed as the normali zed bin-count divided by the median normali zed bin-count .

[0115] In certain embodiments said bin-ratio obtained in step ) is directly used as being representative of the local DNA copy-number with respect to the mean ploidy of the sample . This is the simplest information about DNA copy-number which provides an estimate of fold-change (with respect to the median ploidy) .

[0116] In another preferred embodiment , said fold-change is further elaborated with a step k) to compute logged-bin ratios , calculated as the ratio logged on base 2 , making the values symmetrical with respect to the median ploidy, represented by the value 0 after the trans formation, with positive values representing gain of DNA copies and negative values representing loss of DNA copies .

[0117] In other preferred embodiments , said logged-bin ratios can be further elaborated with a step 1 ) to segment the genome , partitioning it into stretches of genome encompassed between two breakpoint events , comprising contiguous genomic bins which have the same copy number value , as further detailed in more detail in what follows . In a preferred embodiment , said copy-number value is determined to be equal to the median copy-number values o f said contiguous genomic bins . While an individual bin ratio is relatively noisy, calculating the median value across multiple contiguous bins comprised in the same segment provide a more reliable estimate of the local DNA copy-number ratio with respect to the main ploidy .

[0118] Another preferred embodiment comprises a further step m) to obtain an absolute copy number profile as further detailed in the examples below . In brief : for each bin, bincopy numbers are computed by multiplying the median binratio by a sampl e-specifi c factor, wherein said samplespeci fic factor is determined as the factor minimi zing the root mean squared deviation (RMSD) of the bin-copy number values from the corresponding segment median copy number value rounded to the nearest integer .

[0119] Reference sequence

[0120] The invention can be applied to several sample types and the reference sequence is chosen accordingly .

[0121] In a preferred embodiment , the reference sequence to which the reads are mapped is taken from a species belonging to the kingdom of Animalia .

[0122] In a preferred embodiment , the reference sequence to which the reads are mapped is a Human genome reference sequence .

[0123] In another preferred embodiment , the reference sequence to which the reads are mapped is a Mus musculus genome reference sequence.

[0124] Sample types

[0125] In a preferred embodiment, the sample comprising DNA is a single-cell or a pool of cells, selected from the group consisting of:

[0126] - Circulating Tumor Cells,

[0127] - Circulating Multiple Myeloma Cells,

[0128] - Circulating Melanoma Cells,

[0129] Circulating fetal cells in maternal blood, such as Circulating extravillous trophoblasts or Circulating fetal erythroblasts ,

[0130] Disseminated Cancer Cells obtained from one of the following biological samples: cerebro-spinal fluid, lymph nodes, bone marrow or urine;

[0131] - a blastocyst and

[0132] - a polar body.

[0133] In another preferred embodiment, the sample comprising DNA is a cell-free DNA sample, selected from the group consisting of circulating cell-free DNA sample obtained from plasma, such as from a cancer patient, an expecting mother, a healthy donor;

[0134] - cell-free embryonic DNA obtained from spent embryoculture medium, and

[0135] - DNA extracted and purified from a cellulated sample, a complex biological matrix, a tissue sample, such as a biopsy, an FFPE sample, or a cell culture.

[0136] Applications

[0137] The method according to the invention finds application in various fields, among which, by way of non-limiting example , in human diagnostics . In particular, it may be useful to improve the accuracy of copy-number profiling in

[0138] - oncology,

[0139] - cell-based non-invasive prenatal testing,

[0140] - pre-implantation genetic testing / screening .

[0141] Another application of the method is oncology research, such as the study of Circulating Tumor Cells , Circulating Melanoma Cells , Circulating Multiple Myeloma Cells , and Disseminated Cancer Cells .

[0142] Examples

[0143] Absolute copy number workflow

[0144] Figs . 1A and IB show the block diagram of a whole genome copy number analysis method according to the invention based on partitioning of the reference genome in either bins of fixed length or bins of variable length .

[0145] In detail , after sample processing, sequencing reads are aligned (mapped) to the reference genome and the position of WGA fragments on the reference genome sequence is computed in sili co based on the location of enzyme cutting sites on the reference genome sequence . Reads which cannot be mapped univocally with confidence are discarded . For example , reads with mapping quality equal to zero ( reads mapping to multiple locations ) are discarded . In a preferred embodiment only reads with a mapping quality equal or higher to 1 , even more preferably equal or higher than 5 , are retained . The read counts per in sili co expected fragment is computed and the expected normali zed frequency of each in sili co expected fragment is determined . The reference genome sequence is then partitioned either in bins of variable length or in bins of fixed length .

[0146] In the case of the fixed-length workflow, the genome is partitioned in bins of fixed length and the bin-gc is computed for each bin as the mean GC content of the bin . The bin-count is computed for each bin as the sum of reads mapping to each bin . The bin-weight is then computed for each bin as the sum of normali zed frequencies of fragments overlapping the bin . First , a normali zed bin-count is computed as a function of bin-weight and bin-gc by lowess fitting bin-count values against bin-gc and bin-weight . Then, normali zed counts are computed by dividing the bincount by the corresponding fitted values and by multiplying the resulting value by the average of all fitted values .

[0147] In the case of the variable-length workflow, the genome is partitioned into bins of variable length each comprising adj acent fragments so that the bin-weight , computed as the sum of normali zed frequencies of the fragments comprised in the bin, is substantially equivalent to a predefined sumtarget . The bin-gc is then computed for each bin as the mean GC content of the bin and the bin-count is computed as the sum of reads mapping to each bin . A normali zed bin-count is computed as a function of bin-gc by lowess fitting bin-count values against bin-gc . Normali zed counts are computed by dividing the bin-count by the corresponding fitted values and by multiplying the resulting value by the average of all fitted values .

[0148] A bin-ratio is computed as the normali zed bin-count divided by the median of bin-count across all bins . Reference genome is partitioned into segments comprising adj acent bins with similar bin-ratio by circular binary segmentation and in a further step m) absolute copy number values are determined as follows : for each bin, bin-copy numbers are computed by multiplying the median bin-ratio by a sampl e- specific factor comprised between 1.5 and 5.5, wherein said sample-specific factor is determined as the factor minimizing the root mean squared deviation (RMSD) of the bin-copy number values from the corresponding segment median copy number value rounded to the nearest integer.

[0149] For example, a multiplying factor is swept from 1.5 to 5.5 in step of 0.1, and for each value, for each genomic segment obtained from the binary circular segmentation, the median copy number value in the segment is first computed and then rounded to the nearest integer (segment absolute copy-number value) . Then the genome-wide RMSD is computed for the bin-copy numbers (including the multiplication for the multiplying factor being swept) with respect to the segment absolute copy-number value. The multiplying factor corresponding to the lowest RMSD is then selected as the sample-specific factor (which thus corresponds to the estimate of the median ploidy for the sample) . To each bin is then assigned an absolute copy-number value equal to the absolute copy-number value of the segment it pertains, computed with the sample-specific factor. The range for sweeping the multiplying factor (1.5-5.5 in the above example) , is chosen so as to comprise the original median ploidy of "normal" samples (e.g. non-tumor) . For example, for a human germline sample, the multiplying factor 2 (diploid genome) falls within the range. For a human gamete sample (haploid genome) , the range would be thus extended to be 0.5-5.5.

[0150] The upper bound for the sample-specific factor (5.5 in the examples above) is chosen on one hand at least as high as to comprise the possibility of hyperdiploid genomes, which may for example stem from whole-genome doubling (frequent in tumors) corresponding to a tetrapioid genome (multiplying factor=4) , on the other hand as low as to avoid overfitting artefacts which may occur, especially with noisy copy number profiles, for which a very high median ploidy may spuriously result in a lower RMSD, i.e. not corresponding to the real median ploidy of the sample.

[0151] Bin normalization with GC and fragment weight is superior to GC alone (Msel example)

[0152] Figs. 2A-2C show the single and combined relationship among the "GC" and "weight" factors and the read counts generated by the sequencing of a NGS library obtained from the DRS-WGA amplification (Msel) of DNA deriving from a single cell isolated from a Coriell GM17942 cell culture. The Mean Squared Error (MSE) was used to measure the average deviation of the true read count values from those predicted by the LOWESS regression for each individual model ("GC", "weight" and "GC+weight") and thus, identify the best fitting one. a) the plot shows the read count values (y-axis) corresponding to each 500Kb-long genomic bin stratified according to GC content (x-axis) . The black curve, representing the LOWESS regression, indicates that the counts can be largely normalized using the GC content as the main feature (MSE=4.275) . However, there are regions of the genome, mainly belonging to chromosomes 19 and represented by white triangles, which are suboptimally fitted by the regression line and that can be poorly normalized by GC content, b) the plot shows the relation between weight (x- axis) and read count values (y-axis) . This time the same regions of chromosome 19 depicted in a) have a clear nearly- linear relation with weight, in contrast to the other genomic regions which seem to have no relation at all. This demonstrates that , in addition to the GC factor, there are some other components , in some way related to DRS-WGA fragment lengths , which can contribute to the normali zation of the read counts and the improvement of the copy-number signal . However, the MSE calculated for this single feature is much higher, and therefore worse , than the MSE calculated for the GC, showing that the GC alone works better to fix the overall read count bias , c ) the plot shows the fitting surface obtained from the LOWESS regression using both the GC and weight components as independent variables . In this case the individual genomic bins are not shown for image clarity . Using this three-dimensional fitting, considering both GC and weight factors , the MSE decreases by approximately 50% compared to that obtained with the GC alone , showing that the combined model is superior to the single normali zation for both factors .

[0153] Bin normali zation with GC and fragment weight is superior to GC alone , regardless o f restriction enzyme used in DRS-WGA and input material .

[0154] Figures 3 and 4 show the quality statistics of copy number profiles generated from the sequencing of NGS libraries obtained from the ampli fication by DRS-WGA of gDNA from 4 tumor cell lines (NCI -H929 , SK-MM2 , JJN3 , EJM) across all lines and for each single cell line , respectively . Di f ferent restriction enzymes (Bfal , Csp6I , Msel ) were employed, recogni zing 3 di f ferent restriction site sequences on the gDNA sequence ( CATAG, GATAC, TATAA) . Each bar plot displays the mean and 95% confidence interval of the statistic across all cell lines and replicates for 3 methods of normali zation . The RMSD (Root Mean Square Deviation) a statistic known to the experts with ordinary skill in the art, is used as a measure of the similarity of the normalized copy number signal from the deduced copy number of the segment and is calculated as the square root of the quadratic sum of the deviations of the normalized copy number signal in each 500 Kbp genomic bin from the median of the segment rounded to the nearest integer divided by the number of bins in the segment. The DLRS (Derivative Log Ratio Spread) a statistic known to the experts with ordinary skill in the art, is a measure of bin-to-bin consistency, where lower values indicate a cleaner, less noisy, copy number profile and is defined as the standard deviation of n-th discrete difference of the logged ratio between the genomic bin normalized read counts and the median of the normalized counts across all genomic bins along consecutive genomic bins, divided by the root square of 2. The MAPD (Median Absolute Pairwise Difference) a statistic known to the experts with ordinary skill in the art, is, similarly to the DLRS, a measurement of the bin-to-bin variation in read coverage that is robust to the presence of copy-number variations (CNVs) , calculated as the median of absolute pairwise differences along consecutive bins between the logged value of the genomic bin normalized read counts divided by the median of the normalized counts across all genomic bins. Bars of different shades of gray in the plot represent different normalization methods employed where: gcw) uses a normalization function based on combination of bin-weight and bin-gc values which are fitted to bin-count values; gc) employs a function based on bin-gc and bin-count values; w) employs a function based on bin-weight and bincount values. In general the gcw method shows a synergistic effect of bin-gc and bin-weight features with improved statistics compared to both gc and w methods alone , independently of the restriction enzyme employed for the gDNA fragmentation step of DRS-WGA.

[0155] Fig . 5 shows the quality statistics of copy number profiles generated from the sequencing of NGS libraries obtained from the ampli fication by DRS-WGA of single cell DNA of the NA12878 reference cell line by employing 3 distinct restriction enzymes (Bfal , Csp6I , Msel ) , recogni zing 3 dif ferent restriction site sequences ( CATAG, GATAC, TATAA) . Analogously to what is observed for gDNA inputs the synergistic ef fect of bin-gc and bin-weight features is confirmed also for profiles generated from single cells with improved statistics compared to both gc and w methods alone .

[0156] More in detail , looking speci fically at certain individual chromosomes , Figures 6 and 7 show quality statistics of copy number profiles generated from the ampli fication of gDNA from 4 tumor cell lines (NCI-H929 , SK- MM2 , JJN3 , EJM) , and of single cell DNA from NA12878 reference cell line, respectively . In both figures , the first row shows statistics on chromosome 1 , while the second row shows statistics on chromosome 19 . While for chromosome 1 the gcw method shows a synergistic ef fect of bin-gc and binweight features with improved statistics compared to both gc and w methods alone for both gDNA and single cel l inputs , and gc method shows better statistics than w method, in agreement with observations with all chromosomes taken into consideration, for chromosome 19 method w shows better statistics than gc method and in some cases it equals the statistics of the combined method gcw, indicating an ef fect of the genomic region on the importance of di f ferent features in the normalization of the copy number signal.

[0157] Figs. 8A to 8C show, on 3 different chromosomes, the effect of the normalization method on copy number profiles generated from NCI-H929 cell line gDNA amplified by DRS-WGA employing the Bfal restriction enzyme for the gDNA fragmentation step. Dots in the scatter plots represent the copy number value in each genomic bin, while the black line represents the moving average of the copy number value in each bin along the genomic coordinates, using a window size = 11. Shaded areas represent the integral difference between the "normal" copy number level (=2) and the moving average. The normalization method gcw shows the cleaner profile with low difference from the copy number 2 for regions not showing copy number aberrations and a signal well centered to integer copy number levels (Y axis) for detected copy number aberrations. Normalization methods gc and w, instead show distortions of the copy number signal in different regions of the three chromosomes shown which cannot be explained as real biological signals as they either deviate from integer copy number values or show increasing / decreasing trends (black arrows) .

[0158] Fig. 9 shows the effect of the normalization method employed on copy number alteration calling in chromosome 19. DNA libraries were built from NCI-H929 cell line gDNA amplified by DRS-WGA employing the Bfal restriction enzyme for the gDNA fragmentation step. Dots in the scatter plots represent the copy number value in each genomic bin; solid black lines represent the median copy number value (shown above the line) of the segment rounded to the nearest integer; expected copy number levels for chromosome 19 on NCI-H929, based on public data (Menezes et al., The Journal of Molecular Diagnostics, Volume 22, Issue 9, September 2020, Pages 1179-1188, https: / / doi.Org / 10.1016 / j .jmoldx.2020.06.007) , are shown as dashed lines. For both methods gcw and w the observed rounded median copy number levels correspond to those expected (dashed and solid line overlapping and indistinguishable) ; for gc method a negative bias in the copy number level across all the chromosome causes a false negative call in correspondence of the gain at the start of chromosome 19 (black arrow) .

[0159] Figures 10A to IOC, 11A to 11C, 12A to 12C show representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from NCI-H929 cell line gDNA amplified by DRS-WGA with Bfal, Csp6I and Msel restriction enzymes. On the x axis are the genomic coordinates along the 22 autosomes and 2 sexual chromosomes, on the y axis is the copy number level. Dots in the scatter plots represent individual genomic bins. Solid lines represent genomic segments, obtained by segmentation of normalized bin-count with circular binary segmentation (CBS) . The copy number of each segment is calculated as the median of the copy number of all bins in the segment rounded to the nearest integer.

[0160] Figures ISA to 13C, 14A to 14C, 15A to 15C show representative copy number profiles obtained by processing with different normalization methods processing the sequencing data from a library built from gDNA from EJM, JJN3 and SK-MM2 cell lines amplified by DRS-WGA with Bfal restriction enzyme. On the x axis are the genomic coordinates along the 22 autosomes and 2 sexual chromosomes, on the y axis is the copy number level . Dots in the scatter plots represent individual genomic bins . Solid lines represent genomic segments , obtained by segmentation of normali zed bin-count with circular binary segmentation ( CBS ) . The copy number of each segment is calculated as the median of the copy number of all bins in the segment rounded to the nearest integer .

Claims

CLAIMS1 . A method for DNA copy-number profiling of a sample comprising genomic DNA, comprising the following steps : a ) ampli fying said sample comprising genomic DNA using a deterministic restriction-site whole genome ampli fication ( DRS-WGA) , b ) preparing a massively-parallel sequencing library with a fragmentation- free , sequencing-adaptor / WGA fusionprimer PGR reaction, c ) low-pass whole genome sequencing said massively- parallel sequencing library, d) aligning the sequencing reads obtained in step c ) to a reference genome sequence , e ) computing in-silico expected fragments of said DRS-WGA as a list of non-overlapping subsequences of the reference genome sequence each starting and ending at the recognition sequence of the restriction enzyme used in said DRS-WGA, the method further comprising : f ) computing the number of read counts mapping to each in- silico expected fragment sequenced at step c ) , g) computing an expected normali zed frequency of each in- silico expected fragment of said reference genome sequence as the sum of read counts for each in-silico expected fragment length divided by the total number of reads mapped to said reference genome sequence , h) partitioning said reference genome sequence in genomic bins comprising non-overlapping adj acent subsequences of said reference genome , eitherI ) determining fixed-length genomic bins , each bincomprising adjacent in-silico expected fragments so that the sum of said in-silico expected fragment lengths is substantially equivalent to a predefined number of base-pairs sum-target value, orII) determining variable-length genomic bins, each bin comprising adjacent in-silico expected fragments so that the sum of expected normalized frequencies of said in-silico expected fragments is substantially equivalent to a predefined sumtarget value, i) for each genomic bin, either:I) for fixed-length genomic bins1) a bin-GC is computed as the mean GC content of the bin;2) a bin-weight is computed as the sum of normalized frequencies of in-silico expected fragments overlapping the bin;3) a bin-count is computed as the sum of read counts mapping to the bin;4) a normalized bin-count is computed as a function of bin-weight and bin-GC.II) for variable length genomic bins:1) a bin-GC is computed as the mean GC content of the bin;2) a bin-count is computed as the sum of read counts mapping to the bin;3) a normalized bin-count is computed as a function of the bin-GC. j) computing a bin-ratio as the normalized bin-count divided by the median normalized bin-count.

2. The method according to claim 1, wherein the function for computing the normalized bin-count is a LOWESS function.

3. The method according to claim 1 or claim 2, wherein said genomic DNA can be chromosomal or extrachromosomal .

4. The method according to any of the preceding claims, wherein said sample comprising genomic DNA is a single cell.

5. The method according to any of claims 1 to 3, wherein said sample comprising genomic DNA is a cell-free DNA sample.

6. The method according to any of the preceding claims, further comprising a further step k) of computing a logged bin-ratio as the logarithm in base 2 of the bin-ratio.

7. The method according to any of the preceding claims, further comprising a further step 1) of computing a partition of the reference genome sequence comprising referencesequence segments wherein each segment comprises adjacent bins assigned to the same segmented DNA-copy number using circular binary segmentation.

8. The method according to claim 7, comprising a further step m) of computing a bin-copy number by multiplying the bin-ratio values to a sample-specific factor, preferably comprised between 1.5 and 5.5, wherein said sample-specific factor is determined as the factor minimizing the root mean squared deviation of the bin-copy number values from the corresponding segment median copy number value rounded to the nearest integer.

9. The method according to any of the preceding claims, wherein the deterministic restriction-site whole genome amplification (DRS-WGA) , is performed using a four-base cutter restriction enzyme.

10. The method according to claim 9, wherein said four- base cutter restriction enzyme recognizes one of thefollowing DNA sequences a) TTAA b) GTAC c) CTAG.

11. The method according to claim 10, wherein said four- base cutter restriction enzyme is Msel.

12. The method according to claim 10, wherein said four- base cutter restriction enzyme is Csp6I.

13. The method according to claim 10, wherein said four- base cutter restriction enzyme is Bfal .

14. The method according to any of claims 11-13, wherein said low-pass whole genome sequencing is carried out at a coverage > O.Olx, preferably at a coverage > 0.05x, more preferably at a coverage > O.lx, even more preferably at a coverage > 0.5x .

15. The method according to any of the preceding claims, wherein said sample is a human sample and said reference sequence is a human-genome sequence.

16. The method according to any of claims 1 to 4 and 6 to 15, wherein said sample comprising genomic DNA is a cell or a pool of cells, selected from the group consisting of: a) Circulating Tumor Cells; b) Circulating Multiple Myeloma Cells; c) Circulating plasma cells; d) Circulating Melanoma Cells; e) Circulating fetal erythroblast; f) Circulating extravillous trophoblasts; g) Disseminated Cancer Cells; h) Blastocyst cells; i) Polar bodies; and j) Formalin-Fixed Paraffin Embedded Cells.

17. The method according to any of claims 1 to 3 and 5 to 16, wherein said sample comprising genomic DNA is a cell- free DNA sample, selected from the group consisting of: a) circulating cell-free DNA sample obtained from plasma of a cancer patient; b) circulating cell-free DNA sample obtained from plasma of an expecting woman; c) circulating cell-free DNA sample obtained from plasma of a healthy donor; d) cell-free embryonic DNA obtained from spent embryoculture medium; and e) DNA extracted and purified from a cellulated sample.

Citation Information

Patent Citations

  • DNA amplification of a single cell

    WO2000017390A1

  • Error-free sequencing of DNA

    WO2015118077A1

  • Method and kit for the generation of DNA libraries for massively parallel sequencing

    WO2017178655A1

  • Improved method and kit for the generation of DNA libraries for massively parallel sequencing

    WO2019016401A1

  • Method for analysing the degree of similarity of at least two samples using deterministic restriction-site whole genome amplification (DRS-WGA)

    WO2023042173A1