Nucleic acid sample enrichment and screening methods
By using sample-specific nucleic acid libraries with unique markers for enrichment, the method addresses low fetal fraction challenges, enhancing sensitivity and specificity in genetic diagnostic assays, particularly in non-invasive prenatal screening.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MYRIAD WOMENS HEALTH INC
- Filing Date
- 2021-05-18
- Publication Date
- 2026-06-03
AI Technical Summary
Current methods for nucleic acid enrichment in samples, particularly for non-invasive prenatal screening, face challenges with low fetal fraction levels, leading to inefficiencies and reduced sensitivity and specificity in detecting genetic abnormalities, and lack a scalable method for pooling samples while maintaining sample identity.
A method involving sample-specific nucleic acid libraries, labeled with unique markers, are pooled and enriched for target nucleic acid fractions, ensuring equal representation and maintaining sample identity through techniques like size selection and sequencing, enhancing sensitivity and resolution of genetic diagnostic assays.
The method significantly increases the target nucleic acid fraction, improving sensitivity and specificity in genetic diagnostic assays, particularly in non-invasive prenatal screening, by maintaining sample identity and reducing variability.
Smart Images

Figure 0007869755000007 
Figure 0007869755000008 
Figure 0007869755000009
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 704,616, filed May 18, 2020, and U.S. Provisional Patent Application No. 63 / 113,730, filed November 13, 2020, both entitled "Pooled Nucleic Acid Sample Enrichment and Screening Methods".
[0002] Field of the Invention The present invention relates to methods for enriching test samples with respect to target nucleic acid molecules for further genetic screening.
Background Art
[0003] Genetic diversity is known to cause medical conditions, though not limited to, hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF). Often, the cause of the variation is a simple addition, substitution, or deletion of a single nucleotide base in a key gene. In addition, birth defects are caused by chromosomal abnormalities such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), monosomy X (Turner syndrome), and Klinefelter syndrome (XXY). Another form of genetic diversity is fetal sex, which can often be determined based on the sex chromosomes X and Y. Some genetic diversity may make an individual susceptible to any number of diseases, such as diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cancers (e.g., colorectal, breast, ovarian, lung). Identifying genetic variations can lead to the diagnosis of a medical condition or the determination of predisposition to it, and can inform medical decisions. Research and development efforts to find genetic abnormalities that cause harmful health effects have identified specific genes and / or important diagnostic markers for genetic disorders.
[0004] For example, cell-free DNA (cfDNA) has been described and developed as an effective biomarker for diseases and birth defects. For instance, the discovery of cfDNA in maternal plasma during pregnancy has led to the development of several non-invasive prenatal diagnostic (NIPS) techniques in the last decade. However, because only a small fraction of the cfDNA in maternal plasma originates from the fetus, detecting fetal aneuploidy and autosomal recessive disorders using cfDNA presents challenges. The primary driver of NIPS sensitivity for aneuploidy in a given maternal plasma sample is the fetal fraction (FF). Hui, L. et al., Prenat Diagn. 2020; 40:155-163. For most samples, the FF value is between 4% and 30%. Wang, E. et al., Prenat Diagn. 2013; 33:662-666. Many laboratories disqualify samples with less than 4% FF to mitigate the risk of false negative reports. Molecular and bioinformatics implementations of NIPS have evolved, diversified, and improved over time, so sensitivity at progressively lower FF levels depends on the platform and laboratory. Hui, L. et al., Prenat Diagn. 2020; 40:155-163; Artieri, CG et al., Prenat Diagn. 2017; 37:482-490. Indeed, recently published clinical empirical studies have demonstrated that NIPS based on customized whole-genome sequencing (WGS) that does not disqualify low-FF samples can have comparable accuracy at high and low FF levels for common aneuploidy of chromosomes 13, 18, and 21. Hancock, S. et al., Ultrasound Obstet Gynecol. 2020; 56:422-430. While common aneuploidy has long been a major focus of NIPS due to its frequency and highly penetrant phenotype, clinically problematic chromosomal abnormalities can occur across a wide range of sizes and anywhere in the genome.Wapner, RJ et al., Am J Obstet Gynecol. 2015; 212:332; Advani, HV et al., Prenat Diagn. 2017; 37:1067-1075; Pertile, MD et al., Sci Transl Med. 2017; 9:eaan1240. Therefore, a key frontier in NIPS development is increasing the resolution (i.e., detecting relatively small anomalies) and range (i.e., the number of regions) of screening.
[0005] cfDNA is a mixture of DNA that exhibits variation in properties (e.g., size, sequence, abundance) and origin tissue (e.g., maternal or fetal). For example, cfDNA obtained from pregnant women contains DNA of both maternal and fetal origin, cfDNA obtained from cancer patients contains DNA of both tumor and normal cell origin, and cfDNA obtained from transplant patients contains DNA of both host and graft origin. There are numerous known examples that statistically correlate cfDNA properties (e.g., size) with origin tissue. For example, in cfDNA derived from pregnant women, fetal DNA has a smaller fragment size distribution than maternal DNA. Fan, HC, et al., (2010). Analysis of the Size Distributions of Fetal and Maternal Cell-Free DNA by Paired-End Sequencing. Clinical Chemistry. 56(8): 1279-1286.
[0006] cfDNA has potential applications as a diagnostic biomarker for schizophrenia and other diseases such as cancer. In cancer patients, tumor DNA has a size distribution of smaller fragment sizes than DNA derived from normal tissue. Furthermore, cfDNA in schizophrenic patients has been shown to consist of relatively short DNA molecules and exhibit an apoptosis-like distribution pattern. Jang et al., Translational Psychiatry 2018; 8:104.
[0007] For assays involving nucleic acid molecules (e.g., cfDNA) with identifiable characteristics that can function as biomarkers for medical conditions or genetic abnormalities, such as fragment size, the primary factor influencing test performance is the relative proportion of the target DNA fraction (e.g., fetal, tumor, graft) in the sample. In the context of cfDNA, this is often referred to as the "fetal fraction or FF or cffDNA" or the "tumor fraction or cftDNA." For example, the fetal fraction is the ratio of fetal DNA to total DNA in a cfDNA sample (i.e., the proportion of cfDNA fragments derived from the placenta). A higher target DNA fraction makes it easier to detect target DNA characteristics such as chromosomal abnormalities. Nucleic acid methylation signatures are another identifiable feature that can be used. For example, differences in methylation signatures between mother and fetus have been observed. Hong-Dan, W., et al., Mol. Med. Rep. 2017;15(6): 3989-3998. Furthermore, a link has been shown between promoter hypermethylation and inactivation of genes involved in DNA repair that result in specific cancer types. Jin, B. et al., Adv. Exp. Med. Biol. 2013; 754: 3-29.
[0008] It has been shown that target DNA fractions with a relatively small fragment size distribution can be increased by enriching the mixture with respect to relatively small fragments using electrophoresis. (Liang, B., et al., Scientific Reports. 2018; 8:17675). In this process, relatively large fragments are discarded. The result is often a change in the proportion of target DNA in the mixture, but a decrease in the total amount. More than twofold enrichment of the target DNA fraction is possible. For example, a cfDNA mixture that was previously 10% fetal DNA becomes 20% fetal DNA after enrichment with respect to relatively small DNA fragments by electrophoresis. Such enrichment processes, in this case via electrophoresis, are often referred to as "size selection."
[0009] Reports and expert opinions expressing concerns about low FF samples are common, stemming from the usefulness of fetal fractions in diagnosis. (Artieri, CG et al., Prenat Diagn. 2017; 37:482-490; Gregg, AR et al., Genet Med. 2016; 18:1056-1065; Committee on Practice Bulletins - Obstetrics, Committee on Genetics, Society for Maternal-Fetal Medicine. Practice bulletin no. 163., Obstet Gynecol. 2016; 127:e123-e137.) Numerous publications have explored and discussed the merits of different approaches to handling low-FF samples: optimizing the NIPS algorithm to produce reliable results with low FF (Hancock, S. et al., Ultrasound Obstet Gynecol. 2020; 56:422-430), completely disqualifying such samples, or continuing mitigation strategies for disqualified low-FF samples (e.g., continuous re-drawing (Benn, P. et al., Obstet Gynecol. 2018;132:428-435; Hunkapiller, N. et al., Fetal Diagn Ther. 2016; 40:219-223) and FF-based risk scoring (McKanna, T. et al., Ultrasound Obstet Gynecol. 2019; 53:73-79; Benn, P. et al., J Genet Couns. 2019; 29:800-806)). However, consensus remains difficult to find. Therefore, the FX protocol represents progress in NIPS because samples that would have had low FF compared to standard NIPS are molecularly transformed into samples with high FF. With the increased test performance offered by the FX protocol, assay improvements will increase the confidence that suppliers and patients have in their results using NIPS.
[0010] While FF may appear to be an invariant and inherent feature of cfDNA samples, it can be altered, and strategies for increasing FF are revealed by factors correlated with FF. (Hui, L. et al., 2020; 40:155-163; Peng, XL et al., Int J Mol Sci. 2017; 18:453; Hestand, MS et al., Eur J Hum Genet. 2019; 27:198-202). For example, FF is known to increase with gestational age (Wang, E. et al., Prenat Diagn. 2013; 33:662-666), and therefore, blood sampling in late pregnancy results in relatively high FF, but the increase is less than 1% per week, so the impact is not significant. Livergood, MC et al., Am J Obstet Gynecol. 2017; 216:413 e411-413 e419; Yared, E. et al., Am J Obstet Gynecol. 2016; 215:370 e1-376 e6. FF also negatively correlates with first-term body mass index (BMI) and maternal age (Suzumori, N. et al., J Hum Genet. 2016; 61: 647-652), although these values are substantially constant for any given pregnancy. At the molecular level, fetal cfDNA fragments have been observed to be relatively short (Qiao, L., et al., Am J Obstet Gynecol. 2019; 221:345 e1-345 e11; Liang, B., et al., Sci Rep. 2018; 8:17675), hypermethylated (Sun, K. et al., Proc Natl Acad Sci USA. 2015; 112: E5503-E5512; Nygren, AO, et al., Clin Chem. 2010; 56:1627-1635; Lun, FM, et al., Clin Chem. 2013; 59:1583-1594), and to be enriched at different sites than maternal cfDNA fragments.Chan, KC, et al., Proc Natl Acad Sci USA. 2016; 113: E8159-E8168. Leveraging these biases at the molecular and bioinformatics levels has the potential to multiplicatively enhance the FF of each sample.
[0011] Furthermore, conventional selection techniques (e.g., size selection) are limited to parallel single samples, making them inefficient and time-consuming. For example, testing 100 parallel samples would inevitably involve a size selection procedure based on 100 parallel electrophoresis runs. To date, no method has been developed that would allow pooling samples from multiple patients before enrichment while maintaining the identity of sample origin. Significant technical barriers to such pooling and enrichment techniques may be the reason this has not been developed. For example, with respect to cfDNA, one challenge associated with any such solution is that the relative proportion of fragments by size differs among samples. Moreover, pooling equal amounts of samples before size selection results in uneven relative representation of the samples in the final size-selected pool. Furthermore, unlike the case of individual samples, running the protocol in large batches reduces the variability introduced in any sample size selection protocol. This is beneficial as many sequence anomaly detection tools rely on batch-level background correction.
[0012] A procedure is needed that can be scalably applied to samples undergoing screening processes such as NIPS, and that generates significantly higher target DNA fraction levels (e.g., FF levels), thereby increasing sensitivity and specificity for all abnormalities. [Prior art documents] [Non-patent literature]
[0013] [Non-Patent Document 1] Hui, L. et al., Prenat Diagn. 2020; 40:155-163 [Non-licensed document 2] Wang, E. et al., Prenat Diagn. 2013; 33:662-666 [Non-licensed document 3] Artieri, CG et al., Prenat Diagn. 2017; 37:482-490
Non-licensed Document 4
Non-licensed Document 5
Non-licensed Document 6
Non-licensed Document 7
Non-licensed literature 9
Non-licensed literature 10
Non-licensed Document 11
Non-licensed Document 12
Non-licensed Document 13
Non-licensed Document 14
Non-licensed Document 15
Non-licensed Document 16
Non-licensed Document 17
Non-licensed Document 18
Non-licensed Document 19
Non-licensed Document 20
Non-licensed Document 21
[0014] The object of the present invention is to provide a method for enriching and screening a target nucleic acid fraction (e.g., DNA or RNA) in a test sample. In another embodiment, the test sample is pooled from a plurality of test subjects.
[0015] Another object of the present invention is to provide a method for enhancing the sensitivity and resolution of genetic diagnostic assays of nucleic acid samples.
[0016] In some embodiments, the nucleic acid is genomic DNA. In other embodiments, the nucleic acid is cell-free DNA, and in other embodiments, the nucleic acid is FFPE DNA.
[0017] In some embodiments, the target nucleic acid fraction includes the fetal fraction (FF) (cffDNA) of cell-free DNA.
[0018] Another object of the present invention is to determine the appropriate volume / mass of a sample-specific nucleic acid library for use in at least one test sample such that equal amounts or concentrations of target nucleic acid fractions (e.g., within a specific size range) derived from each sample are represented in at least one test sample for a diagnostic assay or further selection / enrichment.
[0019] Another object of the present invention is to determine the appropriate volume / mass for use in at least one test sample by calculating a numerical offset value derived from each sample and then adjusting the volume / mass accordingly.
[0020] Another object of the present invention is to generate a test sample from a sample-specific nucleic acid library enriched with respect to a target nucleic acid fraction. In one embodiment, the test sample is enriched with respect to nucleic acid fragments within a specific size range and contains nucleic acids within a specific size range from each specific source sample at substantially equal concentrations. In another embodiment, the test sample enriched with respect to the target nucleic acid fraction also maintains the identity of the source sample. In one embodiment, the enriched test sample includes a fetal fraction (FF) of cell-free DNA. In yet another embodiment, the enriched test sample includes hypermethylated nucleic acid sequence fragments.
[0021] In one embodiment, a nucleic acid library is prepared / amplified with respect to each collected source sample. In this embodiment, each nucleic acid library corresponds to a source sample in that the nucleic acids contained within a particular library are obtained from a particular sample. In another embodiment, a unique marker (e.g., label, tag) can be used to label nucleic acid fragments in the library for the purpose of preserving the identity of the source sample. For example, a barcode based on a sequence that uniquely identifies the source sample can be attached to the nucleic acid fragment in each sample on one or both ends. In another embodiment, the labeled (e.g., barcoded) nucleic acid mixture contained in the sample-specific library is mixed in a 1:1 ratio by volume (μL) or mass (ng) to produce a first test sample. In another embodiment, the distribution of nucleic acid fragments that retain the desired features in the original sample (e.g., DNA fragment size distribution) is determined by a known methodology, followed by a sample-specific calculation of the relative amount of nucleic acid present (e.g., the relative abundance of nucleic acid present that retains the desired features). In another embodiment, using the sample-specific relative amount, a numerical offset is calculated using the ratio of the actual nucleic acid concentration in the library to the predicted nucleic acid concentration in the library. In another embodiment, a numerical offset value is used to calculate a weighted volume or weighted nucleic acid concentration from a sample-specific library, which is added to a second test sample for further selection or enrichment of fragments retaining the desired features. In another embodiment, a second test sample is prepared by mixing weighted volumes of the sample-specific library, each containing an equal amount of nucleic acid retaining the desired features. In yet another embodiment, selection (or enrichment) of the second test sample is performed using conventional techniques, thereby generating a third test sample containing nucleic acid fragments retaining the desired features. Fragments that do not possess the desired features can be discarded.In one embodiment, a third test sample is sequenced using known sequencing techniques, sample-specific nucleic acids are isolated using sample-specific labeling (e.g., barcodes), and screened for abnormalities.
[0022] In another embodiment, a nucleic acid library is prepared / amplified with respect to each collected source sample. In another embodiment, a unique marker (e.g., label, tag) can be used to label nucleic acid fragments in the library for the purpose of preserving the identity of the source sample. For example, a sequence-based barcode that uniquely identifies the source sample can be attached to the nucleic acid fragment in each sample on one or both ends. In one embodiment, the labeled (e.g., barcoded) nucleic acid mixture contained in the sample-specific library is mixed in a 1:1 ratio by volume (μL) or mass (ng) to produce a first test sample. In one embodiment, selection (or enrichment) is performed on the first test sample using conventional techniques, thereby producing a second test sample (e.g., a "target nucleic acid population") containing nucleic acid fragments that retain the desired features. Fragments that do not have the desired features can be discarded. In one embodiment, the relative amount of the target nucleic acid population to each source sample is determined using techniques including, but is not limited to, sequencing. In other embodiments, the relative amount can be determined using other techniques such as quantitative PCR (qPCR) or droplet digital PCR (ddPCR). In another embodiment, the numerical offset is calculated using a sample-specific relative amount, with the ratio of the actual nucleic acid concentration in the library to the predicted nucleic acid concentration in the library. In one embodiment, the numerical offset value is used to calculate a weighted volume or weighted nucleic acid concentration from the sample-specific library to be added to a third test sample for further selection or enrichment of fragments retaining the desired features. In some embodiments, weighted volumes of the sample-specific library are mixed to create a third test sample, which contains an equal amount of nucleic acid retaining the desired features from each sample-specific library. In some embodiments, the third test sample is sequenced and the target nucleic acid population is screened for genetic abnormalities.In some embodiments, a second selection (enrichment) is performed on a third test sample to isolate the target nucleic acid population from the suspension and form a fourth test sample that is enriched with respect to the target nucleic acid population and contains substantially equal proportions from each source sample. In some embodiments, the fourth test sample is sequenced and the target nucleic acid population is screened for genetic abnormalities.
[0023] Additional embodiments of the present invention include, but are not limited to, the following:
[0024] (1) a. A step of isolating and purifying nucleic acids from multiple test subjects to generate corresponding origin samples for producing at least one origin sample; b. A step of preparing a library for each test subject, wherein the nucleic acid fragments are barcoded and each library corresponds to a specific source sample; c. Adding a first number of nucleic acid units from each source sample to form a first pooled test sample; d. A step to determine the fragment size distribution within each source sample; e. A step of determining the abundance of the target nucleic acid population in each source sample; f. A step of calculating a unique numerical offset value for each origin sample; g. The step of adding a second number of nucleic acid units from each source sample based on a unique numerical offset value to form a second pooled test sample; and h. A step of performing fragment size selection on a second pooled test sample and isolating the target nucleic acid population from the suspension to form a third pooled test sample enriched with respect to the target nucleic acid population, wherein the third pooled test sample is ready for a diagnostic assay. A method for enhancing the sensitivity and resolution of a genetic diagnostic assay of pooled nucleic acid samples, including [the specified method].
[0025] (2) The method of (1), further comprising the step of sequencing the third pooled test sample and screening the target nucleic acid population for genetic abnormalities.
[0026] (3) The method of (1), wherein the fragment size distribution is determined by sequencing.
[0027] (4) The method of (3), wherein the sequencing is paired-end sequencing.
[0028] (5) The method of (1), wherein the fragment size distribution is determined by fluorescence correlation spectroscopy.
[0029] (6) The method of (1), further comprising the step of pairing nucleic acid fragments in a third pooled test sample with their respective source samples.
[0030] (7) The method of (1), wherein the nucleic acid is genomic DNA.
[0031] (8) The method of (1), wherein the nucleic acid is FFPE DNA.
[0032] (9) The method of (1), wherein the nucleic acid is RNA.
[0033] (10) The method of (1), wherein the nucleic acid is cell-free DNA.
[0034] (11) The method of (1) wherein the nucleic acid is isolated from whole blood.
[0035] (12) The method of (1), wherein the unique numerical offset value is calculated by dividing the abundance of the target nucleic acid population determined in step e by a first number of nucleic acid units.
[0036] (13) The method of (12), wherein the target nucleic acid population is the fetal fraction of the cell-free DNA.
[0037] (14) The method of (12), wherein the target nucleic acid population is the tumor fraction of cell-free DNA.
[0038] (15) The method of (12), wherein the target nucleic acid population is a nucleic acid fragment containing a specific methylation trace.
[0039] (16) The method of (15), wherein the methylation trace is hypermethylation or hypomethylation.
[0040] (17) The method of (1), wherein the target nucleic acid population is enriched with respect to fragments within a predetermined length range.
[0041] (18) The method of (1), wherein the target nucleic acid population is enriched with respect to fragments of a predetermined length.
[0042] (19) The method of (1), wherein the target nucleic acid population is enriched with respect to fragments containing specific methylation traces.
[0043] (20) The method of (19), wherein the methylation trace is hypermethylation or hypomethylation.
[0044] (21) The method of (1), wherein the selection of the fragment size is performed using gel electrophoresis.
[0045] (22) The method of (1), wherein the first and second numbers of nucleic acid units are selected from the group consisting of microliters, nanograms, and moles.
[0046] (23) The method of (1), further comprising the step of performing whole-genome sequencing.
[0047] (24) The method of (1), wherein the pooled test samples include 2 to 1000 different samples.
[0048] (25) a. A step of isolating and purifying nucleic acids from multiple test subjects to produce corresponding source samples; b. A step of preparing a library for each test subject, wherein the nucleic acid fragments are barcoded and each library corresponds to a specific source sample; c. Adding a first number of nucleic acid units from each source sample to form a first pooled test sample; d. Performing fragment size selection on the first pooled test sample and isolating the target nucleic acid population from the suspension to form a second pooled test sample enriched with respect to the target nucleic acid population; e. A step of determining the abundance of the target nucleic acid population in each source sample; f. A step of calculating a unique numerical offset value for each origin sample; g. A step of adding a second number of nucleic acid units from each source sample based on a unique numerical offset value to form a third pooled test sample enriched with respect to the target nucleic acid population; and h. A step of performing a second fragment size selection on a third pooled test sample and isolating the target nucleic acid population from the suspension to form a fourth pooled test sample that is enriched with respect to the target nucleic acid population and contains substantially equal proportions from each of the source samples, wherein the fourth pooled test sample is ready for a diagnostic assay. A method for enhancing the sensitivity and resolution of a genetic diagnostic assay of pooled nucleic acid samples, including [the specified method].
[0049] (26) The method of (25), wherein step f is performed by sequence determination.
[0050] (27) The method of (25), wherein the sequencing is paired-end sequencing.
[0051] (28) The method of (25), wherein step f is performed by quantitative PCR.
[0052] (29) The method of (25), wherein step f is performed by digital PCR.
[0053] (30) The method of (29), wherein the digital PCR is droplet digital PCR.
[0054] (31) The method of (25), further comprising the step of sequencing the fourth pooled test sample and screening the target nucleic acid population for genetic abnormalities.
[0055] (32) The method of (25), further comprising the step of pairing nucleic acid fragments in a fourth pooled test sample with their respective source samples.
[0056] (33) The method of (25), wherein the nucleic acid is genomic DNA.
[0057] (34) The method of (25), wherein the nucleic acid is FFPE DNA.
[0058] (35) The method of (25), wherein the nucleic acid is RNA.
[0059] (36) The method of (25), wherein the nucleic acid is cell-free DNA.
[0060] (37) The method of (25) wherein the nucleic acid is isolated from whole blood.
[0061] (38) The method of (25), wherein the unique numerical offset value is calculated by dividing the abundance of the target nucleic acid population determined in step e by a first number of nucleic acid units.
[0062] (39) The method of (25), wherein the target nucleic acid population is the fetal fraction of the cell-free DNA.
[0063] (40) The method of (25), wherein the target nucleic acid population is the tumor fraction of the cell-free DNA.
[0064] (41) The method of (25), wherein the target nucleic acid population is a fragment containing a specific methylation trace.
[0065] (42) The method of (25), wherein the methylation trace is hypermethylation or hypomethylation.
[0066] (43) The method of (25), wherein the target nucleic acid population is enriched with respect to nucleic acid fragments within a predetermined length range.
[0067] (44) The method of (25), wherein the target nucleic acid population is enriched with respect to fragments of a predetermined length.
[0068] (45) The method of (25), wherein the target nucleic acid population is enriched with respect to fragments containing specific methylation traces.
[0069] (46) The method of (25), wherein the methylation trace is hypermethylation or hypomethylation.
[0070] (47) The method of (25), wherein the selection of the fragment size is performed using gel electrophoresis.
[0071] (48) The method of (25), wherein the first and second numbers of nucleic acid units are selected from the group consisting of microliters, nanograms, and moles.
[0072] (49) The method of (25), further comprising the step of performing whole-genome sequencing.
[0073] (50) The method of (25), wherein the pooled test samples include 2 to 1000 different samples.
[0074] (51) a. A step of isolating and purifying nucleic acids from at least one test subject to produce at least one source sample; b. A step of preparing a nucleic acid library for at least one test subject, wherein the nucleic acid fragments are barcoded and the nucleic acid library corresponds to at least one source sample; c. Adding a first number of nucleic acid units from the nucleic acid library to form a first test sample; d. A step of determining the fragment size distribution within the nucleic acid library; e. A step of calculating the abundance of the target nucleic acid population in the nucleic acid library; f. A step of calculating a unique numerical offset value for the nucleic acid library; g. The step of adding a second number of nucleic acid units from the nucleic acid library based on a unique numerical offset value to form a second test sample; and h. A step of performing fragment size selection on a second test sample and isolating the target nucleic acid population from the suspension to form a third test sample enriched with respect to the target nucleic acid population, wherein the third test sample is ready for the diagnostic assay. A method for enhancing the sensitivity and resolution of a genetic diagnostic assay, including [specific method / feature].
[0075] (52) The method of (51), comprising multiple test subjects, multiple source samples, and multiple nucleic acid libraries corresponding to the multiple source samples.
[0076] (53) The method of (51), comprising a single test subject, a single source sample, and a single nucleic acid library corresponding to the source sample.
[0077] (54) The method of (52), wherein one or more of steps a to h are duplicated and performed on one or more microtiter plates (one well per source sample).
[0078] (55) The method of (54), wherein the plurality of nucleic acid libraries are pooled prior to the diagnostic assay.
[0079] Built-in by reference All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as each individual publication, patent, or patent application is specifically and individually indicated as being incorporated by reference. To the extent that any publications and patents or patent applications incorporated by reference conflict with any disclosures contained herein, this specification is intended to take precedence and / or supersede any such conflicting material.
[0080] Representative embodiments of the present invention are disclosed in more detail with reference to the following drawings. [Brief explanation of the drawing]
[0081] [Figure 1] This figure shows a flowchart illustrating an embodiment of the method described herein. [Figure 2] This figure shows a graph illustrating how the method described herein (FX protocol) increases the fetal fraction (FF) across all BMI levels. For 2,401 patients who indicated their BMI on the test request form, the fetal fraction levels measured before (circles) and after (triangles) the method are plotted as a function of the patient's BMI (vertical axis). The upper panel plots histograms of the samples with and without the application of the method. [Figure 3] This figure shows a graph illustrating the measurement of the abundance of cfDNA fragments from chrY plotted for male fetal pregnancies. The FX protocol is labeled "FFA" in this figure. [Figure 4] This figure shows a plot illustrating the change ratio in FF as a result of applying the FX protocol, as a function of the original FF without the FX protocol. The dotted line indicates no change in FF, while samples in the shaded area have increased FF when the FX protocol is used. [Figure 5]Figure 5A is a schematic diagram of the change in median depth per autosome as a result of the FX protocol. The degree of deviation from the background is itself a measure of FF and is shown as FF positive. Figure 5B illustrates that the increase in FF positivity with and without the FX protocol (circles) is shown for aneuploid samples with the indicated chromosomal abnormalities. Figure 5C illustrates the z-scores for the same samples as in Figure 5B, stratified by their screening results and summarized as a population distribution, with and without the FX protocol ("standard NIPS") and with the FX protocol ("NIPS with the FX protocol"). The distribution of screening-negative samples (negative; dotted line) has been scaled to a height comparable to the screening-positive distribution on the right (solid line). The vertical solid line indicates the z-score cutoff between screening-negative (left) and screening-positive (right) results. Regarding SCA, since z-scores are used to identify chrX aneuploidy, only female fetal pregnancies are shown (i.e., MX and TX), and a two-dimensional analysis without z-scores (not shown) is required to identify XXY and XYY (FF positivity increases in all XXY and XYY pregnancies tested using the FX protocol). NIPS: Non-invasive prenatal screening, RAA: Rare autosomal aneuploidy, SCA: Sex chromosome aneuploidy. Figure 5D illustrates the z-scores for the same samples as in Figure 5B, stratified by those screening results and summarized as separate samples, with and without the FX protocol (circles) and with the FX protocol (triangles). [Figure 6] This figure shows a plot illustrating how the FX protocol improves the coefficient of variation (CV) for mapped reads compared to a procedure that does not use the FX protocol. [Figure 7]This figure shows plots comparing the assay sensitivity for short microdeletions (approximately 3MB; shaded blue on the left side of each plot) under the FX protocol (labeled "FFA" in this plot) and standard NIPS conditions. The enhanced fetal fraction of FFA (left side) revealed microdeletions, while standard NIPS (right side) screened them as negative. The scattered dots are bin-level normalized depths; the black scattered dots indicate the rolling median of the blue dots (median across a 25-bin window). This particular deletion is shorter than a typical 5p deletion, and the median 3' breakpoint for it is shown (http: / / dbsearch.clinicalgenome.org / search / ). FF was 11% with FFA and 7% with standard NIPS. [Figure 8] This figure shows ROC curve graphs for various classes of chromosomal aberrations, demonstrating that the FX protocol (labeled "FFA" in this plot) enables near-perfect analytical sensitivity with near-perfect analytical specificity. The general sensitivity for aneuploidy is the aggregate sensitivity of RAA, which is higher when using the FX protocol. [Figure 9] This figure shows the distribution of FFchrY values for samples called female or male. For both standard NIPS (upper) and prequels using the FX protocol (lower), the solid lines show the raw data, and the dotted lines show the best-fit traces for the female (Gaussian) and male (Beta) populations. Only euploid samples are included. The arrow depicts one sample tested on both platforms that was called female in standard NIPS and called male when using the FX protocol (confirming the fetus was male). After minimizing the estimated number of miscalls on each platform, the miscalls in the analysis are predicted to be 318-fold lower using the FX protocol. [Figure 10]This diagram illustrates that fetal cfDNA is less abundant and more orderly and shorter than maternal cfDNA. The standard NIPS approach sequences all cfDNA samples regardless of length. Since a relatively large distribution of fetal cfDNA is retained compared to the maternal cfDNA distribution, the FX protocol (or "FFA" if labeled in this diagram) increases the FF by using agarose gel electrophoresis to select relatively short cfDNA fragments. The FX protocol increases the relative concentration of fetal cfDNA fragments through size selection. [Modes for carrying out the invention]
[0082] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Numerous variations, changes, and substitutions can be conceived by those skilled in the art without departing from the present invention. Substitutions for embodiments of the present invention described herein can be utilized. Where values are given as ranges, it will be understood that such disclosures include disclosures of all conceivable subranges within such ranges, as well as specific numerical values that fall within such ranges, whether or not specific numerical values or specific subranges are explicitly stated.
[0083] Unless a term is expressly defined in this Patent by the sentence “As used herein, the term ‘_____’ generally means ~” or a similar sentence, there is no intention to limit the meaning of that term, either expressly or implicitly, beyond its plain or ordinary meaning, and such term should not be construed as limiting in scope by any description made in any section of this Patent (except in the language of the claims). Any term used in the claims at the end of this Patent is referred to in this Patent in a manner consistent with a single meaning, solely for the purpose of preventing confusion for the reader, and such term in the claims is not intended to be limited in scope by any implied or otherwise. Finally, unless an element of a claim is defined by describing a function without the word “means” and any description of any structure, the scope of any element of a claim is not intended to be construed under the application of Section 112, paragraph 6 of the U.S. Patent Act.
[0084] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. Generally, the nomenclature used herein and the experimental procedures in cell culture, molecular genetics, organic chemistry, and nucleic acid chemistry and hybridization described below are well known and commonly used in the art. The techniques and procedures are generally carried out in accordance with methods commonly used in the art and the various general references provided throughout this document (see Sambrook et al. MOLECULAR CLONING: A LABORATORY MANUAL, 2d ed. (1989) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, generally incorporated herein by reference). The nomenclature used herein and the experimental procedures in analytical chemistry and organic synthesis described below are well known and commonly used in the art.
[0085] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly indicates otherwise.
[0086] As used herein, the terms “about” or “approximately” generally mean within an acceptable margin of error for a value as determined by those skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limits of the measuring system. For example, “about” may mean within 1 or more than 1 standard deviation, according to the practice in the relevant field. Alternatively, “about” may mean within 20%, 10%, 5%, or 1% of a given value.
[0087] As used herein, the term “subject” generally means an animal such as a mammal (e.g., human) or a bird (e.g., bird), or another living organism such as a plant. For example, a subject may be a vertebrate, a mammal, a rodent (e.g., mouse), a primate, an ape, or a human. Animals include, but are not limited to, livestock, sports animals, and pets. A subject may be a healthy or asymptomatic individual, an individual with or suspected to have a disease or a predisposition to a disease, and / or an individual that requires or is suspected to require treatment. A subject may be a patient. A subject may be a microorganism (or microbe) (e.g., bacteria, fungi, archaea, viruses).
[0088] As used herein, the term “genome” generally refers to the genetic information of a subject, which may be, for example, at least part or all of the information passed from the subject’s parents. The genome can be encoded in either DNA or RNA. The genome can include coding regions (e.g., protein-coding regions) as well as non-coding regions. The genome can include the sequences of all the chromosomes present in an organism. For example, the human genome typically has a total of 46 chromosomes. All of these sequences together can constitute the human genome.
[0089] As used herein, the terms “polynucleotide,” “nucleotide,” “nucleotide sequence,” “nucleic acid,” and “oligonucleotide” are interchangeable and generally refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, intergenetic DNA, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), small nucleolar RNA, ribozymes, cDNA, FFPE DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, aptamers, and primers. Polynucleotides may include modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be conjugated before or after polymer assembly. Nucleotide sequences can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with labeling components, tags, reactive moieties, or binding partners. Where provided, polynucleotide sequences are listed in the 5' to 3' direction unless otherwise noted.
[0090] As used herein, the term "gene" generally refers to a DNA segment that is involved in polypeptide production and includes regions preceding and following the coding region, as well as intervening sequences (introns) between individual coding segments (exons).
[0091] As used herein, the term “base pair” or “bp” generally refers to the interaction between adenine (A) and thymine (T), or between cytosine (C) and guanine (G) in a double-stranded DNA molecule (i.e., hydrogen bond pairing). In some embodiments, a base pair may include, for example, A paired with uracil (U) in a DNA / RNA double helix.
[0092] As used herein, the term “barcode” generally means a known nucleic acid sequence that enables the identification of certain features of the polynucleotide associated with the barcode. In some embodiments, the polynucleotide feature to be identified is the sample from which the polynucleotide originates. In some embodiments, the barcode is approximately or at least approximately 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more nucleotides long. In some embodiments, the barcode is less than 10, 9, 8, 7, 6, 5, or 4 nucleotides long. In some embodiments, the barcode associated with some polynucleotides is of a different length than the barcode associated with other polynucleotides. Generally, the barcodes are of sufficient length and contain sequences that are sufficiently distinct to enable the identification of a sample based on the barcodes they are associated with. In some embodiments, a barcode and the associated sample source can be precisely identified after mutations, insertions, or deletions of one or more nucleotides in the barcode sequence, such as mutations, insertions, or deletions of one, two, three, four, five, six, seven, eight, nine, ten, or more nucleotides. In some embodiments, each barcode in a plurality of barcodes differs from all the other barcodes in the plurality at at least three nucleotide positions, such as at least three, four, five, six, seven, eight, nine, ten, or more nucleotide positions. The plurality of barcodes can be represented in a sample pool, and each sample contains a polynucleotide containing one or more barcodes that are different from the barcodes included in the polynucleotide derived from other samples in the pool. Samples of polynucleotides containing one or more barcodes can be pooled based on the barcode sequences in which they are combined, so that the four nucleotide bases A, G, C, and T are represented approximately equally at one or more positions along each barcode in the pool (such as one, two, three, four, five, six, seven, eight, or more positions, or all positions, etc.).In some embodiments, the method of the present invention further includes the step of identifying the sample from which the target polynucleotide originates based on the barcode sequence to which the target polynucleotide is combined. Generally, the barcode includes a nucleic acid sequence that, when combined with the target polynucleotide, serves as an identifier for the sample from which the target polynucleotide originates. In some embodiments, a separate amplification reaction is performed for each separate sample using amplification primers containing at least one different barcode sequence for each sample, thereby ensuring that in a pool of two or more samples, there are no barcode sequences to combine with the target polynucleotide of more than one sample. In some embodiments, amplified polynucleotides from different samples and containing different barcodes are pooled before proceeding to further operations on the polynucleotide (e.g., before amplification and / or sequencing on a solid phase). The pool may include any fraction of the total component amplification reaction, including the entire reaction volume. The samples may be pooled equally or unevenly. In some embodiments, the target polynucleotides are pooled based on the barcode to which they are combined. The pools can accommodate approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 20, 25, 30, 40, 50, 75, 100 or more species, approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 20, 25 species, It may contain polynucleotides derived from 30, 40, 50, 75, 100 or more different samples, or approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 20, 25, 30, 40, 50, 75, 100 or more different samples.
[0093] As used herein, the term “sequencing” generally refers to methods and techniques for determining the sequence of nucleotide bases in one or more polynucleotides. For example, polynucleotides can be nucleic acid molecules such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) (including their variants or derivatives (e.g., single-stranded DNA)). Sequencing can be performed using a variety of systems currently available, including, but not limited to, sequencing systems from Illumina®, Pacific Biosciences (PacBio®), Oxford Nanopore®, or Life Technologies (Ion Torrent®). Such systems can provide multiple raw genetic data corresponding to the genetic information of a subject (e.g., a human) when generated by the system from a sample provided by the subject. In some examples, such systems provide sequencing reads (also referred to herein as “reads”). Reads may contain strings of nucleic acid bases corresponding to the sequence of the sequenced nucleic acid molecule.
[0094] As used herein, “next-generation sequencing” (NGS) generally refers to sequencing methods that enable large-scale parallel sequencing of cloned and single nucleic acid molecules, where multiple nucleic acid fragments, such as one million, are sequenced simultaneously from a single sample or multiple different samples. Non-exclusive examples of NGS include sequencing-by-synthesis, sequencing-by-ligation, real-time sequencing, and nanopore sequencing.
[0095] As used herein, the term “paired-end sequencing” generally refers to a high-throughput sequencing-based method that generates sequencing data from both ends of a nucleic acid molecule. The method typically involves a step of sequencing the nucleic acid sequence inward from the ends. Paired-end sequencing is useful for determining the length of a DNA segment located between two different sequences.
[0096] As used herein, the term “whole genome sequencing” means determining the complete DNA sequence of a genome in a single step. As used herein, “whole genome sequence” or WGS (also referred to in the art as “whole,” “complete,” or “entire” genome sequence) generally means encompassing a substantial portion of the subject’s genome, though not necessarily complete. In the art, the term “whole genome sequence” or WGS is used in some usages to mean a nearly complete genome of a subject, such as at least 95% complete. As used herein, the term “whole genome sequence” or WGS does not include “sequences” used in gene-specific techniques, such as single nucleotide polymorphism (SNP) genotyping, which typically cover less than 0.1% of the genome. As used herein, the term “whole genome sequence” or WGS does not require the genome to be aligned with any reference sequence, nor does it require the annotation of variants or other features.
[0097] The term “fragment size distribution” refers to one or a set of values that represent the length, mass, weight, or other measure of size of molecules (e.g., nucleic acid fragments originating from a particular chromosomal region) corresponding to a particular group. Various embodiments may utilize a variety of size distributions. In some embodiments, the size distribution relates to the ranking of the size (e.g., arithmetic mean, median, or mean) of a given chromosomal fragment compared to other chromosomal fragments. In other embodiments, the size distribution may relate to a statistical value of the actual size of a chromosomal fragment. In one implementation, the statistical value may include the arithmetic mean, mean, or median size of any of the chromosomal fragments. In another implementation, the statistical value may include the total fragment length below a cutoff value, which can be calculated by dividing the total length of all fragments, or at least the total length of fragments below a greater cutoff value.
[0098] As used herein, the terms “library” or “sequencing library” generally mean nucleic acids (e.g., DNA or RNA) that are processed for sequencing, for example, using large-scale parallel methods, such as NGS. The nucleic acids may optionally be amplified to obtain a collection of multiple copies of the processed nucleic acid, which can be sequenced by NGS or other preferred techniques.
[0099] As used herein, “fraction multiplier technology” or “FX technology” or “FX protocol” generally means a method described herein for increasing the yield of a target nucleic acid fraction (e.g., cffDNA) and thereby increasing sensitivity to detecting abnormalities, such as fetal abnormalities resulting from copy number variations of any size across the genome. Embodiments of methods utilizing the FX protocol are described in further detail below. In some embodiments, the FX protocol leverages a reduced size of the target nucleic acid molecule to increase the relative abundance of the target nucleic acid fraction. In such examples, the method may be referred to as “fetal fraction amplification” or “FFA.”
[0100] Referring to Figure 1, in one embodiment, a biological sample is collected from a test subject (e.g., plasma from a pregnant woman), and nucleic acids are isolated, purified, a library is prepared, and amplified using primers combined with sequence-based barcodes. In some embodiments, nucleic acids are extracted from formalin-fixed paraffin-embedded tissue (FFPE DNA). Formalin is commonly used as a fixative for long-term preservation and storage of tissue samples. The fixation process adequately preserves the ultrastructure of the tissue but results in various types of damage to the DNA within the tissue. DNA damage traces in FFPE DNA include hydrolysis of N-glycosyl bonds, deamination, oxidation, thymine dimerization, nicks, and double-strand breaks. Double-strand breaks produce genomic DNA fragments of varying lengths. A unique barcode sequence is used for each original sample ("sample of origin" or "original sample"). In this embodiment, a sequence-based barcode (also known as a nucleotide sequence identifier or NSI) is used to label nucleic acid fragments for the purpose of preserving the identity of the origin sample. Generally, high-throughput sequencing techniques often involve barcode ligation to nucleic acid fragments, which may include primer binding sites used for capturing, amplifying, and / or sequencing those nucleic acid fragments. This technique is particularly useful when samples from different origins are combined into a single high-throughput sequencing run. Sequence-based barcodes allow technicians to trace the origin of each sample from a pool of test samples. Such high-throughput sequencing systems that utilize nucleotide sequence identifiers to track the origin of each sample are known in the art. A non-limiting example of this technique is the Roche-designed FLX genome sequencer system, which utilizes multiplexed identifier sequences (MIDs). Similar nucleotide sequence identifiers are available for other sequencing systems using next-generation sequencing (NGS).
[0101] The following description includes the implementation of the method on pooled samples; however, it should be noted that the processes described herein are equally applicable to single samples and / or multiplexed samples. For example, a multiplex PCR reaction can be used to enrich a target nucleic acid. In one multiplexed embodiment, PCR primers can be designed for a barcode, and the PCR reaction can be performed to amplify the sequence containing the barcode. In another embodiment, the original samples are combined to a desired mixing ratio to produce a first pooled test sample. In one embodiment, a predetermined amount of the original sample is added to produce a first pooled test sample, thereby making the units of each sample substantially equivalent (e.g., 1:1:1:1, etc.) and equimixed. "Unit" can be defined as any appropriate unit of measurement, such as nanograms (ng), microliters (μL), or moles (mol).
[0102] In another embodiment, the distribution of nucleic acid fragments retaining specific characteristics (e.g., fragment size, molecular weight, methylation status) is assayed for a first pooled test sample. For example, in the case of cfDNA or FFPE DNA, paired-end sequencing can be used to estimate the fragment size distribution of each original sample within the first pooled test sample. Paired-end sequencing is well known in the art and was originally described in Smith, MW et al. (1994). Genomic sequence sampling: a strategy for high resolution sequence-based physical mapping of complex genomes. Nature Genetics. 7: 40-47. Paired-end sequencing obtains information for both ends of each DNA molecule. The length of a DNA fragment can be estimated by finding the coordinates of the two sequences relative to the genome via sequence alignment. A single sequencing experiment yields sequence and size information for 1 million to 1 billion DNA fragments. Sequencing can be performed using various currently available systems, including, but not limited to, those from Illumina®, Pacific Biosciences (PacBio®), Oxford Nanopore®, or Life Technologies (Ion Torrent®). It should be understood that other suitable techniques are available for determining the fragment size distribution. For example, fluorescence correlation spectroscopy as described in Jiang, J., et al. (2018). Analysis of the concentrations and size distributions of cell-free DNA in schizophrenia using fluorescence correlation spectroscopy. Translational Psychiatry. 8:104.
[0103] In one embodiment, once the fragment size distribution in a first pooled test sample is determined, the sample-specific relative amounts (i.e., in selected units) of DNA fragments within a target fragment size range (e.g., 100–165 bp) in the length vial, or alternatively, DNA fragments of a specific target size (e.g., 165 bp), can then be calculated for each original sample. The sample-specific relative amounts can be calculated using amplification procedures such as polymerase chain reaction (PCR), quantitative PCR (qPCR), droplet digital PCR (ddPCR), and isothermal amplification. In another embodiment, the sample-specific relative amounts can be used to calculate numerical offset values by determining the ratio of sample-specific relative amounts (units) to the total units in the first pooled test sample. Yet another embodiment uses numerical offset values to calculate the number of weighted units (e.g., μL, ng, etc.) from a sample-specific library to be added to a second pooled test sample for fragment size selection.
[0104] In one embodiment, a weighted number of units (e.g., based on mass, volume, etc.) of each source sample are mixed together to produce a second pooled test sample. In some embodiments, a weighted number of units are mixed so that the units are substantially equal in proportion. In some embodiments, unequal amounts (a predetermined number of sample-specific units) of one or more source samples are mixed together, depending on the assay being performed. Once the second pooled test sample is produced, fragment size selection can be performed to select with respect to a desired fragment length. In some embodiments, gel electrophoresis can be used to isolate, excise, and purify the desired nucleic acid size fraction and produce a third pooled test sample containing nucleic acid fragments within the target size range or a specific target size. In one embodiment, nucleic acid electrophoretic separation and subsequent recovery of the desired fragment length are used. Various known electrophoretic processes can be used for this purpose, but in one embodiment, Ranger Technology for high-throughput nucleic acid size selection is used. TMUse NIMBUS Select TM A workstation can be used.
[0105] It should be understood that alternative techniques for fragment selection may be used, such as bisulfite conversion techniques and subsequent column-based purification of methylated DNA; methylated DNA immunoprecipitation (based on nucleic acid methylation compared to other nucleic acids); solid-phase capture (e.g., affinity columns) (e.g., antibody-coated spin columns); synchronous (or non-synchronous) coefficient of drag alteration sizing (SCODA); solid-phase reversible solid-phase sizing (e.g., using carboxylated magnetic beads); affinity chromatography processes, or combinations of PCR amplification with amplicons of various lengths and microchip separation.
[0106] In another embodiment, following the preparation of a third pooled test sample comprising a nucleic acid mixture enriched with respect to target nucleic acid fragments of a specific size or size range, and containing predetermined proportions (equal or varying proportions) from each of the original samples, the third pooled test sample is sequenced by next-generation sequencing (NGS) or the like. Once the third pooled test sample is sequenced, the barcode can be deconvolved via a commercially available software program to pair the reads with the source samples based on the barcode sequence, as described in more detail above. Following the read pairing with the source samples, the sequenced reads can be analyzed, if desired, depending on the screening performed.
[0107] In an alternative embodiment, the numerical offset value is determined by performing fragment size selection on the first pooled sample, for example, by omitting the above fragment size distribution step via paired-end sequencing. In this embodiment, a predetermined amount of the first pooled test sample is used to isolate, extract, and generate a second pooled test sample containing nucleic acid fragments within the target size range or of a specific target size. In one embodiment, nucleic acid electrophoretic separation and subsequent recovery of nucleic acid fragments of the desired length are used as described above. In some embodiments, following recovery, the second pooled test sample can be sequenced to estimate the relative abundance of the target fragment in each source sample based on the number of reads assigned to the specific sample in the sample-specific bin. Subsequently, the sample-specific relative amount (i.e., in selected units) of DNA fragments within the target fragment size range (e.g., 100-165 bp) in the length bin, or alternatively, DNA fragments of a specific target size (e.g., 165 bp), can be calculated with respect to each original sample. Sample-specific relative amounts can be calculated using amplification procedures such as polymerase chain reaction (PCR), quantitative PCR (qPCR), droplet digital PCR (ddPCR), and isothermal amplification. In another embodiment, sample-specific relative amounts can be used to calculate numerical offset values by determining the ratio of sample-specific relative amounts (units) to the total units in the first pooled test sample. In yet another embodiment, numerical offset values are used to calculate the number of weighted units (e.g., μL, ng, etc.) from the sample-specific library to be added to the second pooled test sample for fragment size selection. In this embodiment, the numerical offset value is determined based on the relative abundance observed from the sample-specific sequencing reads. Aliquotes from each source sample, adjusted based on the numerical offset value, are pooled to produce a third pooled test sample, and the fragment size selection / sequencing step is repeated. Ideally, target fragments of equal relative abundance (based on molar concentration, molecular weight, etc.) are present in the third pooled test sample, which is also enriched with respect to the target fragments.In some embodiments, a second fragment size selection of a third pooled test sample can be performed using conventional techniques, and the target nucleic acid population is isolated in a suspension, thereby forming a fourth pooled test sample enriched with respect to the target nucleic acid population and containing substantially equal proportions from each of the source samples. In some embodiments, the third pooled test sample and / or the fourth pooled test sample are sequenced to screen the target nucleic acid population for genetic abnormalities.
[0108] In one embodiment, the FX protocol can be combined with NIPS based on whole-genome sequencing (WGS). For example, WGS-based NIPS (without the FX protocol) is configured to identify novel microdeletions somewhere in the genome, but its sensitivity and resolution are limited to microdeletions longer than 7 MB. Since many microdeletions are less than 7 MB, increased sensitivity for small regions across the genome could have significant clinical value. The resolution limits for detecting copy number variants across the genome are driven by the relative amount of signal in the sample (determined by the relative amount of FF and the size of CNVs in the sample) and the amount of noise present in the sample (determined by the depth to which the sample is sequenced) (for example, detecting small deletions in samples with low FF is more difficult). Attempts to increase resolution by deeper sequencing are screening tests that offer reduced results and rapid acquisition, but are economically unsustainable. Therefore, methods that increase FF (fetal fraction) are preferable under these circumstances; hence, the impact and applicability of the FX protocol are demonstrated. [Examples]
[0109] For each example discussed below, all samples were from patients who consented to an unspecified study and underwent testing using the NIPS prequel. The study was granted an Institutional Review Board (IRB) exemption by Advarra (Pro0042194).
[0110] Example 1 Plasma was separated from a 10 mL whole blood sample via centrifugation at 1600 g for 10 minutes using a two-step centrifugation process. The plasma was transferred to a microcentrifuge tube and centrifuged at 16,000 × g for 10 minutes. The plasma was stored at -80°C until DNA extraction. DNA fragments were extracted from 0.6 mL of cell-free plasma using a circulating nucleic acid kit (Qiagen, GE). The Ion Plus fragment library kit for the Ion Proton platform (Life Technologies, USA) was used to construct sequencing libraries for each plasma sample. The libraries were quantified using a Qubit fluorometer, where each sample-specific library contained substantially the same concentration of total DNA. The sample-specific libraries were barcoded to indicate the identity of the source sample. Each library contained different amounts of DNA within a specific size range (100–165 bp) or a specific target size (165 bp). As shown in Table 1, samples 1-5 (the remaining samples are not shown) were mixed in a 1:1 ratio (10 ng each) to produce the first pooled test sample (FPTS).
[0111] In some embodiments, alternative extraction and library preparation protocols involve extracting a target nucleic acid fraction (e.g., cffDNA) from plasma using silanol-coated magnetic beads (Dynabeads, ThermoFisher) to obtain a sample with relatively uniform concentration and fragment size or length (e.g., 165 bp). The target nucleic acid fraction was quantified (PicoGreen, ThermoFisher) and converted into a barcoded next-generation (NGS) qualified sequencing library suitable for the Illumina platform using the manufacturer's instructions. The library was amplified via 12 rounds of polymerase chain reaction (PCR) (KAPA HiFi HotStart PCR kit, Roche), followed by PCR purification based on magnetic beads and subsequent further quantification rounds.
[0112] The fragment size distribution of each sample in the FPTS was determined, and the relative amount (ng) of DNA within the target size range (100-165 bp) was obtained. 2 μL samples from each representative sample in the FPTS were analyzed for fragment size distribution using a fragment analyzer (Advanced Analytical Technologies, Ames Iowa). Referring to Table 1, the percentage of DNA within the target size range for each sample (1-5) was calculated compared to the total units (10 ng) added to the FPTS. Using these values, the total amount of DNA required from each of the five samples to have equal and predetermined DNA units (1 ng) within the target size range was calculated.
[0113] Based on calculations, the original library samples were remixed to a volume that would add 1 ng of DNA within the target size range for each sample, generating a second pooled test sample (SPTS). The SPTS were subjected to a fragment size selection procedure using a 2% E-Gel EX CloneWell agarose gel (Invitrogen, Carlsbad, CA, USA) as described by Qiao et al. and Liang et al. One E-Gel contains six effective wells, each capable of processing a mixed sample containing five different samples from the DNA sequencing library. DNA within the target size range was recovered from the bottom well on the gel, and the selected library (i.e., the third pooled test sample (TPTS)) was sequenced using the Ion Proton system (Life Technologies). Other sequencing strategies, such as processing via Illumina HiSeq 4000 and subsequent custom bioinformatics pipelines, can be used.
[0114] Another strategy for fragment size selection is electrophoresis on a 2% agarose cassette (BluePippin, Sage Science) following the manufacturer's instructions for the "range" mode. Short fragments are eluted from the gel until the desired target size of the eluted DNA is, for example, 140 nt. See Figure 10. With respect to fetal cfDNA, this size adequately preserves fetal cfDNA and depletes maternal cfDNA. Since fetal-derived fragments contain a relatively high fraction of the total size-selected cfDNA, the size-selected library had a relatively high FF (Full Fibre).
[0115] [Table 1]
[0116] The barcodes were deconvoluted via a commercially available software program to pair the reads with the source samples based on the barcode sequence, and the reads were subsequently screened for any associated medical conditions or chromosomal abnormalities, such as fetal aneuploidy.
[0117] Example 2 To analytically validate the FX protocol, plasma was extracted from 10 mL whole blood samples (1,264 NIPS patient samples and 66 controls tested for 11 batches) using a two-step centrifugation process. The plasma was transferred to a microcentrifuge tube and centrifuged at 16,000 × g for 10 minutes to remove residual cells and obtain cell-free plasma. The cell-free plasma was stored at -80°C until DNA extraction. DNA fragments were extracted from 0.6 mL of cell-free plasma using a circulating nucleic acid kit (Qiagen, GE). Sequencing libraries were constructed for each plasma sample using the Ion Plus fragment library kit for the Ion Proton platform (Life Technologies, USA), and the libraries were quantified using a Qubit fluorescence spectrometer, where each sample-specific library contained substantially the same concentration of total DNA. Sample-specific libraries were barcoded to indicate the identity of the source sample. Each library contained varying amounts of DNA within a specific size range (100–165 bp) or a specific target size (165 bp). The samples were mixed in a 1:1 ratio (10 ng each) to produce the first pooled test sample (FPTS).
[0118] Each patient sample was processed through two workflows: (1) NIPS based on standard WGS without the FX protocol, or (2) NIPS based on WGS with the FX protocol. The FX protocol leverages the reduced size of fetal-derived cfDNA molecules to increase the relative abundance of fetal cfDNA. The workflows were performed completely independently, each starting with cfDNA extraction from repeated plasma aliquots.
[0119] The fragment size distribution of each sample in the FPTS was determined, and the relative amount (ng) of DNA within the target size range (100-165 bp) was obtained. 2 μL samples from each representative sample in the FPTS were analyzed for fragmentation size distribution using a fragment analyzer (Advanced Analytical Technologies, Ames Iowa). The percentage of DNA within the target size range was calculated for each sample compared to the total units (10 ng) added to the FPTS. Using the obtained values, the total amount of DNA required from each of the five samples to have equal and predetermined DNA units (1 ng) within the target size range was calculated. According to the calculation, the original library samples were remixed to a volume that would add 1 ng of DNA within the target size range to each sample, generating a second pooled test sample (SPTS). The SPTS were subjected to a fragment size selection procedure, in this case using an E-Gel CloneWell agarose gel (Invitrogen, Carlsbad, CA, USA). Each E-Gel contains six effective wells, each capable of processing a mixed sample containing five different samples from a DNA sequencing library. DNA within the target size range was collected from the bottom well on the gel, and the selected library (i.e., the third pooled test sample (TPTS)) was sequenced using the Ion Proton system (Life Technologies).
[0120] A. The FX protocol increases the fetal fraction (FF). To directly measure the effect of the FX protocol, samples were tested using both the standard NIPS and FX protocols, with particular focus on the number of samples with FF > 4%. According to the American Society of Clinical Genetics and Genomics (ACMG), the threshold for low FF is less than 4%. As shown in Figure 2 (top), 3.7% of the samples tested using the standard NIPS protocol had FF less than 4% (e.g., fractions containing cffDNA). However, when using the FX protocol, zero samples had FF less than 4%. The lowest FF observed in the FX group was 4.9%. Since samples from patients with high BMI tend to have low FF and have been observed to cause an increase in test failures when using the standard NIPS protocol, samples were divided by their BMI class (Figure 2 (bottom)). Even among the highest BMI levels (Class III obesity), where 16% of the sample had low FF using standard NIPS, all samples tested using the FX protocol had FF > 4% (the lowest FF observed among Class III patients was 7.1%).
[0121] To confirm that the FX protocol did not artificially increase FF by damaging the inventors' FF inference regression model, the inventors verified that the density of reads from chrY in pregnancies with male fetuses increased proportionally. See Figure 3. Therefore, the FX protocol increases FF by directly increasing the relative abundance of fetal-derived cfDNA fragments in each sequenced sample.
[0122] The change in sample level at FF resulting from the FX protocol was examined to determine whether an upward shift in the overall FF distribution might mask downward-shifted FF in a subset of samples. See Figure 3. Figure 4 shows the relative increase in FF imposed by the FX protocol. In particular, 2395 out of 2401 samples tested (99.8%) showed an increase in FF when using the FX protocol, with an average increase of 2.3 times. The relative increase in sample level at FF changed as a function of FF (Figure 3): samples that had low FF (less than 4%) with standard NIPS showed the greatest increase in FF after receiving the FX protocol, with an average increase of 3.9 times. Consistent with the decrease in FF increase at higher original FF levels, the six samples in which FF decreased using the FX protocol had a median FF value of 27.8% (minimum 6.5%), while the FF remained high using the FX protocol (median: 25.4%, minimum: 6.4%).
[0123] B. The FX protocol increases NIPS sensitivity. Just as the fetal fraction (FF) can be directly measured in male fetal pregnancies from the relative NGS depth of chrX and chrY, it is also possible to measure the FF of aneuploid samples via the relative NGS depth to aneuploid chromosomes (FF 陽性 (Figure 5A). FF 陽性 The z-score is directly proportional to the aneuploidy region, with a higher z-score indicating that aneuploidy is more easily detected. Therefore, the FX protocol is effective in detecting aneuploidy in the aneuploidy region. 陽性 When increasing the FX protocol, the NIPS sensitivity also increases.
[0124] In all positive samples tested across common aneuploidy (e.g., sex chromosome aneuploidy (SCA)), rare autosomal aneuploidy (RAA), and microdeletions, the FX protocol resulted in increased FF (Figure 5B). FF without the FX protocol 陽性 The increase is indicated by the gray circle, and FF using the FX protocol 陽性The increase in z-scores is indicated by the purple triangle. This upward shift in the distribution of FF was not altered by FX. FX also increased the z-scores for all tested aneuploid samples, while the z-score distribution for euploid samples remained unchanged (Figure 5C-D). The relatively large z-score separation between positive and negative samples enhances the ability to distinguish such samples, thereby reducing the variability of false negatives and false positives. At the same time, these findings demonstrate that the FX protocol directly increases the concentration of fetal-derived reads in each sample and enhances the sensitivity and specificity of NIPS.
[0125] Samples that were screened negative for 5p microdeletions using standard NIPS but positive using the FX protocol (Figure 7, 5B, 5D) provided further support for the enhanced sensitivity for fetal chromosomal abnormalities conferred by the FX protocol. For this 3 MB microdeletion, copy number variation was prominently evident in the FX protocol data (Figure 7), converting z-scores below the call threshold to those above the threshold (Figure 3D, row of microdeletion).
[0126] To quantify the increase in sensitivity and specificity achievable using the FX protocol, the inventors analyzed the relationships between clinical and technical metrics such as z-score, depth, incidence, and FF (see Materials and Methods). ROC curves for different classes of chromosomal aberrations (Figure 8) show that the FX protocol enables near-perfect analytical sensitivity with near-perfect analytical specificity (Table 2). The sensitivity for common aneuploidies, which has been shown to be high without the FX protocol in the inventors' clinical experience, is the aggregate sensitivity of RAA, and is therefore higher when using the FX protocol (Figure 8). However, the increase in microdeletion sensitivity is substantial: with the FX protocol, the aggregate sensitivity for five common microdeletions is 97.2% with a joint specificity of 99.8%. In particular, for DiGeorge syndrome, the FX protocol has an analytical sensitivity of 95.6% with an analytical specificity of 99.95%. Table 2 (below).
[0127] [Table 2]
[0128] In addition to evaluating performance using the ROC analysis described above, the inventors also observed that all samples with confirmed aneuploidy or microdeletions were correctly identified using the FX protocol (Table 3).
[0129] [Table 3]
[0130] Furthermore, the results demonstrated repeatability and reproducibility both within and between batches (Tables 4 and 5). Simultaneously, these experiments establish the analytical reliability of the FX protocol.
[0131] [Table 4]
[0132] [Table 5]
[0133] Finally, the number of false negative (FN) results per sample screened using a prequel (labeled FFA herein) employing the standard NIPS or Fx protocol was estimated. The false negative rate (FNR) was calculated as (1-sensitivity), where sensitivity is the analytical sensitivity estimated from the ROC analysis. The number of FNs per screened sample is the product of FNR and prevalence. The prevalence number may vary based on age and other factors; that is, the prevalence value is an approximation expressed as 1 in x, where x is rounded to the nearest hundred (for common aneuploidy and RAA) or nearest 1000 (for common microdeletions and 22q11.2). The five common microdeletions are 1p, 4p, 5p, 15q11, and 22q11.2. The proportions of FNR and FN per screened sample are much lower when the FX protocol is used to enhance the fetal fraction. Please refer to Table 6 below.
[0134] [Table 6]
[0135] C. The FX protocol increases the accuracy of gender calling. Sex miscalls in NIPS arise from either biological (e.g., true fetal mosaicism, vanishing twin) or technical (e.g., low FF) limitations. The former has inherent difficulties (many sex miscalls occur at FF much greater than 4%), while the latter can be mitigated by the FX protocol due to its ability to increase FF in all samples, thus eliminating borderline calls. Figure 9 shows the distribution of FFchrY for male and female fetal pregnancies (i.e., FF as measured from NGS read density on the Y chromosome) as observed for standard NIPS and the FX protocol. In particular, the separation between male and female FFchrY distributions is relatively large when using the FX protocol, reducing the chance of sex miscalls due to borderline FFchrY values by up to approximately 318 times. To highlight the improvement, one sample tested in the validation study (Figure 9, orange arrow) was borderline and miscalled as XX in standard NIPS; however, the sample was clearly XY in screening using the FX protocol, and the pregnancy was orthogonally confirmed to be male via ultrasound.
[0136] Example 3 The goal of numerous next-generation sequencing (NGS) based tests is to integrate samples in equal amounts prior to sequencing. Ideally, all samples would receive the exact number of reads they require to maintain test performance. In many cases, for consistency, this number of reads will be equal across all samples, and the coefficient of variation (CV) for mapped reads will be zero. However, this is often not the case due to errors in the process (i.e., liquid handling, quantification, etc.). This is even more problematic when pooled samples are size-selected to isolate only specific size ranges of nucleic acids.
[0137] In this experiment, approximately 120 NGS libraries were combined (or pooled) at equimolar concentrations, and gel-based size selection was performed by gel electrophoresis to isolate fragments between 200 and 250 base pairs. Subsequently, the pooled NGS libraries were sequenced on an Illumina sequencer, and downstream analyses were performed to determine the number of reads mapped to the genome for all samples. Referring to Figure 6, the distribution of mapped reads centered on the mean is shown in the box plot on the left under the label "No in silico coefficient". Although all libraries were initially pooled at equimolar concentrations, the distribution of their reads was very broad because each sample had a different number of fragments within the 200–250 base pair size bins, with the lowest read count sample receiving one-third the number of reads as the highest read count sample. The CV for these samples was 0.16.
[0138] Next, the same equimolar pool was sequenced on a small Illumina sequencing platform, and the fragment length distribution for all samples in the pool was determined using paired-end data. The relative amount of DNA present in the 200–250 bp size range was determined and used to create a "coefficient" that can be applied to the original quantification values, which we here call the "in silico coefficient." Using these updated quantification values, the NGS libraries were reintegrated so that each library contained equimolar amounts of DNA in the 200–250 bp size range. Subsequently, this reintegrated pool was subjected to size selection based on the same 200–250 bp gel, Illumina sequencing, and analysis as outlined above. The distribution of reads mapped around the mean is shown in the box plot on the right. Since the samples were pooled based on the number of molecules present in the 200–250 bp size bins, the CV of the mapped reads for this integrated pool is quite low at 0.03.
[0139] Since many assays utilizing NGS have a minimum read threshold required for analysis, a lower CV (i.e., a tighter distribution) of mapped reads allows for lower costs and / or fewer disqualified samples. Often, libraries with broadly distributed mapped reads have samples that do not receive a sufficient number of reads. Mitigation strategies for this involve combining relatively small sample sizes with an integrated pool, but this comes at the cost of increased costs. This workflow allows for selection of DNA sample pool sizes while still maintaining a tight distribution (low CV) of mapped reads.
[0140] Here, the inventors verified and characterized the performance of NIPS with the FX protocol applied to all samples. For 99.8% of the tested samples, FF increased using the FX protocol, with an average increase of 2.3 times. Low FF samples underwent the largest FF scale change; of the 2401 samples tested, 3.7% had low FF before the FX protocol, but none had low FF after the FX protocol. The increase in FF was molecular, not algorithmic: the FX protocol distinguishes between maternal and fetal DNA and increases the relative proportion of fetal DNA in samples undergoing WGS. NIPS showed high sensitivity and specificity for general aneuploidy across the FF spectrum without using the FX protocol technique, but the application of the FX protocol increased performance for each type of aneuploidy. The increase was particularly significant for microdeletions.
[0141] The FX protocol has a dramatic impact on the performance of microdeletion screening in NIPS. For common microdeletions, the expected total sensitivity is increased (Table 1, Figure 8), reaching 97.2% when using the FX protocol. Previously, the American College of Obstetricians and Gynecologists (ACOG) did not recommend microdeletion screening due to low sensitivity and specificity for microdeletions; however, we anticipate that sensitivity of over 97% and specificity of over 99% for microdeletions will enable professional organizations to consider the clinical benefits of screening for microdeletions over the technical limitations.
[0142] Beyond common microdeletions, our data suggest that the FX protocol increases the resolution of gwCNV detection, enabling reliable identification of microdeletions below the current 7MB limit achievable with standard NIPS. Short microdeletions in samples with low FF can be difficult to detect with NIPS and may limit sensitivity, but the FX protocol increases the achievable sensitivity limit by reducing the frequency of low FF samples. In particular, the 22q11.2 microdeletion causing DiGeorge syndrome is most commonly around 2-3MB and has an expected sensitivity of 95.6% when using the FX protocol. While the resolution limit for novel gwCNV detection may need to be greater than 3MB to ensure that false positives are rare, dbVar includes more than 1000 unique pathogenic microdeletions ranging in size from 3MB to 7MB, whose number is associated with clinically severe phenotypes, and therefore any increase in resolution should increase the usefulness of NIPS for patients and suppliers.
[0143] Even when two NIPS laboratories must test the same plasma sample, the reported FF and aneuploidy sensitivity may differ due to variations in the molecular and computer protocols of each laboratory. For example, based on different methods of alignment, filtration, counting, and NGS read analysis, a laboratory reporting 8% FF may have higher aneuploidy sensitivity than a laboratory reporting 10% FF. In particular, since laboratories demonstrate performance on different sample sets and using different study designs (e.g., clinical empirical studies and analytical validation studies), these differences complicate comparisons of NIPS performance between laboratories. Therefore, it can be difficult to make definitive statements regarding relative NIPS performance. However, in this specification, we have demonstrated a clear increase in NIPS performance: two protocols (standard NIPS and FX protocol) were compared on a single set of samples within a single laboratory using a single aneuploidy calling algorithm. FF increased 2.3 times on average, and this increase in FF stemmed from a higher frequency of fetal-derived NGS reads. Without needing to provide evidence of a relative increase in performance, our ROC analysis yields an estimate of analytical sensitivity and specificity in an unbiased cohort that reflects a large population of clinical samples.
[0144] The FX protocol strategies described herein can increase the sample's FF (Free Fissure) at the molecular level through size selection upstream of sequencing, and further increase the FF downstream of sequencing through algorithmic size selection. Specifically, a bioinformatics pipeline could calculate the length of each fragment based on the mapping position of each paired-end read and weight relatively shorter fragments during analysis. However, a drawback of this bioinformatics approach is that considerable resources are consumed by sequencing relatively long fragments that are likely to be of maternal origin, even if they contribute little to fetal aneuploidy detection. In contrast, when molecular size selection is performed upstream of sequencing, all of the sequenced fragments have an increased likelihood of being of fetal origin.
[0145] While the present invention has been described in some detail with examples and embodiments for the purpose of clear understanding, it will be apparent to those skilled in the art that certain changes and modifications can be carried out without departing from the spirit and scope of the invention. Therefore, the description should not be understood as limiting the scope of the invention. The present invention includes, for example, the following embodiments: [1] A method for enhancing the sensitivity and resolution of a genetic diagnostic assay of pooled nucleic acid samples, comprising the following steps: a. A step of isolating and purifying nucleic acids from multiple test subjects to generate corresponding origin samples for producing at least one origin sample; b. A step of preparing a library for each test subject, wherein nucleic acid fragments are barcoded and each library corresponds to a specific source sample; c. Adding a first number of nucleic acid units from each source sample to form a first pooled test sample; d. A step to determine the fragment size distribution within each source sample; e. A step of determining the abundance of the target nucleic acid population in each source sample; f. A step of calculating a unique numerical offset value for each origin sample; g. The step of adding a second number of nucleic acid units from each source sample based on the unique numerical offset value to form a second pooled test sample; and h. A step of performing fragment size selection on the second pooled test sample and isolating the target nucleic acid population from the suspension to form a third pooled test sample enriched with respect to the target nucleic acid population, wherein the third pooled test sample is ready for a diagnostic assay. The method, including the method described above. [2] The method according to [1], further comprising the step of sequencing the third pooled test sample and screening the target nucleic acid population for genetic abnormalities. [3] The method according to [1], wherein the fragment size distribution is determined by sequencing. [4] The method according to [3], wherein the sequencing is paired-end sequencing. [5] The method according to [1], wherein the fragment size distribution is determined by fluorescence correlation spectroscopy. [6] The method according to [1], further comprising the step of pairing nucleic acid fragments in the third pooled test sample with their respective source samples. [7] The method according to [1], wherein the nucleic acid is genomic DNA. [8] The method according to [1], wherein the nucleic acid is FFPE DNA. [9] The method according to [1], wherein the nucleic acid is RNA.
[10] The method according to [1], wherein the nucleic acid is cell-free DNA.
[11] The method according to [1], wherein the nucleic acid is isolated from whole blood.
[12] The method according to [1], wherein the unique numerical offset value is calculated by dividing the abundance of the target nucleic acid population determined in step e by a first number of nucleic acid units.
[13] The method according to
[12] , wherein the target nucleic acid population is the fetal fraction of the cell-free DNA.
[14] The method according to
[12] , wherein the target nucleic acid population is the tumor fraction of cell-free DNA.
[15] The method according to
[12] , wherein the target nucleic acid population is a nucleic acid fragment containing a specific methylation trace.
[16] The method according to
[15] , wherein the methylation trace is hypermethylation or hypomethylation.
[17] The method according to [1], wherein the target nucleic acid population is enriched with respect to fragments within a predetermined length range.
[18] The method according to [1], wherein the target nucleic acid population is enriched with respect to fragments of a predetermined length.
[19] The method according to [1], wherein the target nucleic acid population is enriched with respect to fragments containing specific methylation traces.
[20] The method according to
[19] , wherein the methylation trace is hypermethylation or hypomethylation.
[21] The method according to [1], wherein the selection of the fragment size is performed using gel electrophoresis.
[22] The method according to [1], wherein the first and second numbers of nucleic acid units are selected from the group consisting of microliters, nanograms, and moles.
[23] The method according to [1], further comprising the step of performing whole genome sequencing.
[24] The method according to [1], wherein the pooled test samples include 2 to 1,000 different samples.
[25] A method for enhancing the sensitivity and resolution of a genetic diagnostic assay of pooled nucleic acid samples, comprising the following steps: a. A step of isolating and purifying nucleic acids from multiple test subjects to generate corresponding source samples; b. A step of preparing a library for each test subject, wherein nucleic acid fragments are barcoded and each library corresponds to a specific source sample; c. Adding a first number of nucleic acid units from each source sample to form a first pooled test sample; d. Performing fragment size selection on the first pooled test sample and isolating the target nucleic acid population from the suspension to form a second pooled test sample enriched with respect to the target nucleic acid population; e. A step of determining the abundance of the target nucleic acid population in each source sample; f. A step of calculating a unique numerical offset value for each origin sample; g. The step of adding a second number of nucleic acid units from each source sample based on the unique numerical offset value to form a third pooled test sample enriched with respect to the target nucleic acid population; and h. A step of performing a second fragment size selection on the third pooled test sample and isolating the target nucleic acid population from the suspension to form a fourth pooled test sample that is enriched with respect to the target nucleic acid population and contains substantially equal proportions from each of the source samples, wherein the fourth pooled test sample is ready for a diagnostic assay. The method, including the method described above.
[26] The method according to
[25] , wherein step f is performed by sequencing.
[27] The method according to
[25] , wherein the sequencing is paired-end sequencing.
[28] The method according to
[25] , wherein step f is performed by quantitative PCR.
[29] The method according to
[25] , wherein step f is performed by digital PCR.
[30] The method according to
[29] , wherein the digital PCR is a droplet digital PCR.
[31] The method according to
[25] , further comprising the step of sequencing the fourth pooled test sample and screening the target nucleic acid population for genetic abnormalities.
[32] The method according to
[25] , further comprising the step of pairing nucleic acid fragments in the fourth pooled test sample with their respective source samples.
[33] The method according to
[25] , wherein the nucleic acid is genomic DNA.
[34] The method according to
[25] , wherein the nucleic acid is FFPE DNA.
[35] The method according to
[25] , wherein the nucleic acid is RNA.
[36] The method according to
[25] , wherein the nucleic acid is cell-free DNA.
[37] The method according to
[25] , wherein the nucleic acid is isolated from whole blood.
[38] The method according to
[25] , wherein the unique numerical offset value is calculated by dividing the abundance of the target nucleic acid population determined in step e by a first number of nucleic acid units.
[39] The method according to
[25] , wherein the target nucleic acid population is the fetal fraction of the cell-free DNA.
[40] The method according to
[25] , wherein the target nucleic acid population is the tumor fraction of cell-free DNA.
[41] The method according to
[25] , wherein the target nucleic acid population is a fragment containing a specific methylation trace.
[42] The method according to
[25] , wherein the methylation trace is hypermethylation or hypomethylation.
[43] The method according to
[25] , wherein the target nucleic acid population is enriched with respect to nucleic acid fragments within a predetermined length range.
[44] The method according to
[25] , wherein the target nucleic acid population is enriched with respect to fragments of a predetermined length.
[45] The method according to
[25] , wherein the target nucleic acid population is enriched with respect to fragments containing specific methylation traces.
[46] The method according to
[25] , wherein the methylation trace is hypermethylation or hypomethylation.
[47] The method according to
[25] , wherein the selection of the fragment size is performed using gel electrophoresis.
[48] The method according to
[25] , wherein the first and second numbers of nucleic acid units are selected from the group consisting of microliters, nanograms, and moles.
[49] The method according to
[25] , further comprising the step of performing whole genome sequencing.
[50] The method according to
[25] , wherein the pooled test samples include 2 to 1,000 different samples.
[51] A method for enhancing the sensitivity and resolution of a genetic diagnostic assay, comprising the following steps: a. A step of isolating and purifying nucleic acids from at least one test subject to produce at least one source sample; b. A step of preparing a nucleic acid library for at least one test subject, wherein the nucleic acid fragments are barcoded and the nucleic acid library corresponds to at least one source sample; c. Adding a first number of nucleic acid units from the nucleic acid library to form a first test sample; d. A step of determining the fragment size distribution within the nucleic acid library; e. A step of calculating the abundance of the target nucleic acid population in the nucleic acid library; f. A step of calculating a unique numerical offset value for the nucleic acid library; g. The step of adding a second number of nucleic acid units from the nucleic acid library based on the specific numerical offset value to form a second test sample; and h. A step of performing fragment size selection on the second test sample and isolating the target nucleic acid population from the suspension to form a third test sample enriched with respect to the target nucleic acid population, wherein the third test sample is ready for a diagnostic assay. The method, including the method described above.
Claims
1. A method for enhancing the sensitivity and resolution of a genetic diagnostic assay of pooled nucleic acid samples, comprising the following steps: a. A step of isolating and purifying nucleic acids from multiple test subjects to generate corresponding origin samples for producing at least one origin sample; b. A step of preparing a library for each test subject, wherein nucleic acid fragments are barcoded and each library corresponds to a specific source sample; c. Adding a first number of nucleic acid units from each source sample to form a first pooled test sample; d. A step to determine the fragment size distribution within each source sample; e. A step of determining the abundance of the target nucleic acid population in each source sample; f. A step of calculating a unique numerical offset value for each source sample, wherein the unique numerical offset value is calculated by dividing the abundance of the target nucleic acid population determined in step e by the first number of nucleic acid units in the first pooled test sample; g. A step of adding a second number of nucleic acid units calculated from the unique numerical offset value from step f to each source sample to form a second pooled test sample; and h. A step of performing fragment size selection on the second pooled test sample and isolating the target nucleic acid population from the suspension to form a third pooled test sample enriched with respect to the target nucleic acid population, wherein the third pooled test sample is ready for a diagnostic assay. The method, including the method described above.
2. The steps further include sequencing the third pooled test sample and screening the target nucleic acid population for genetic abnormalities, or The third further step includes pairing nucleic acid fragments in the pooled test samples with their respective source samples, or The process further includes a step of performing whole-genome sequencing. The method according to claim 1.
3. The aforementioned fragment size distribution is determined by sequencing, and optionally, the sequencing is paired-end sequencing. The aforementioned fragment size distribution is determined by fluorescence correlation spectroscopy. The nucleic acid is genomic DNA. The nucleic acid is FFPE DNA. The nucleic acid is RNA. The nucleic acid is cell-free DNA, or The nucleic acid is isolated from whole blood. The method according to claim 1.
4. The aforementioned target nucleic acid population is the fetal fraction of cell-free DNA. The aforementioned target nucleic acid population is a tumor fraction of cell-free DNA, or The target nucleic acid population is a nucleic acid fragment containing a specific methylation trace, and optionally, the methylation trace is hypermethylated or hypomethylated. The method according to claim 1.
5. The target nucleic acid population is enriched with respect to fragments within a predetermined length range. The target nucleic acid population is enriched with respect to a fragment of a predetermined length. The target nucleic acid population is enriched with respect to fragments containing specific methylation traces, and optionally, the methylation traces are hypermethylated or hypomethylated. The aforementioned fragment size selection is performed using gel electrophoresis. The first and second numbers of the nucleic acid units are selected from the group consisting of microliters, nanograms, and moles, or The pooled test samples include 2 to 1000 different samples. The method according to claim 1.
6. A method for enhancing the sensitivity and resolution of a genetic diagnostic assay of pooled nucleic acid samples, comprising the following steps: a. A step of isolating and purifying nucleic acids from multiple test subjects to generate corresponding source samples; b. A step of preparing a library for each test subject, wherein nucleic acid fragments are barcoded and each library corresponds to a specific source sample; c. Adding a first number of nucleic acid units from each source sample to form a first pooled test sample; d. Performing fragment size selection on the first pooled test sample and isolating the target nucleic acid population in the suspension to form a second pooled test sample enriched with respect to the target nucleic acid population; e. A step of determining the abundance of the target nucleic acid population in each source sample; f. A step of calculating a unique numerical offset value for each source sample, wherein the unique numerical offset value is calculated by dividing the abundance of the target nucleic acid population determined in step e by the first number of nucleic acid units in the first pooled test sample; g. Adding a second number of nucleic acid units calculated from the unique numerical offset value from step f to each source sample to form a third pooled test sample enriched with respect to the target nucleic acid population; and h. A step of performing a second fragment size selection on the third pooled test sample and isolating the target nucleic acid population from the suspension to form a fourth pooled test sample that is enriched with respect to the target nucleic acid population and contains substantially equal proportions from each of the source samples, wherein the fourth pooled test sample is ready for a diagnostic assay. The method, including the method described above.
7. Step f is performed by array determination. The aforementioned sequencing is paired-end sequencing. Step f is performed by quantitative PCR. Step f is performed by digital PCR, and optionally, the digital PCR is droplet digital PCR. The method according to claim 6.
8. The steps further include sequencing the fourth pooled test samples and screening the target nucleic acid population for genetic abnormalities, or The fourth step further includes pairing nucleic acid fragments in the pooled test samples with their respective source samples. The method according to claim 6.
9. The nucleic acid is genomic DNA. The nucleic acid is FFPE DNA. The nucleic acid is RNA. The nucleic acid is cell-free DNA, or The nucleic acid is isolated from whole blood. The method according to claim 6.
10. The aforementioned target nucleic acid population is the fetal fraction of cell-free DNA. The aforementioned target nucleic acid population is a tumor fraction of cell-free DNA. The aforementioned target nucleic acid population is a fragment containing a specific methylation trace. The methylation trace is either hypermethylation or hypomethylation. The target nucleic acid population is enriched with respect to nucleic acid fragments within a predetermined length range. The target nucleic acid population is enriched with respect to a fragment of a predetermined length. The aforementioned target nucleic acid population is enriched with respect to fragments containing specific methylation traces. The methylation trace is either hypermethylation or hypomethylation. The aforementioned fragment size selection is performed using gel electrophoresis, or The first and second numbers of nucleic acid units are selected from the group consisting of microliters, nanograms, and moles. The method according to claim 6.
11. The method according to claim 6, further comprising the step of performing whole-genome sequencing.
12. The method according to claim 6, wherein the pooled test samples include 2 to 1,000 different samples.
13. A method for enhancing the sensitivity and resolution of a genetic diagnostic assay, comprising the following steps: a. A step of isolating and purifying nucleic acids from at least one test subject to produce at least one source sample; b. A step of preparing a nucleic acid library for at least one test subject, wherein the nucleic acid fragments are barcoded and the nucleic acid library corresponds to at least one source sample; c. Adding a first number of nucleic acid units from the nucleic acid library to form a first test sample; d. A step of determining the fragment size distribution within the nucleic acid library; e. A step of calculating the amount of the target nucleic acid population in the nucleic acid library; f. A step of calculating a unique numerical offset value for the nucleic acid library, wherein the unique numerical offset value is calculated by dividing the abundance of the target nucleic acid population determined in step e by the first number of nucleic acid units in the first pooled test sample; g. A step of adding a second number of nucleic acid units calculated from the specific numerical offset value from step f from the nucleic acid library to form a second test sample; and h. A step of performing fragment size selection on the second test sample and isolating the target nucleic acid population from the suspension to form a third test sample enriched with respect to the target nucleic acid population, wherein the third test sample is ready for a diagnostic assay. The method, including the method described above.