Methods for analyzing stored blood for subsequent analysis of cell-free DNA
Patent Information
- Application Number
- JP2024563757
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-17
- Filing Date
- 2023-01-17
- Publication Date
- 2026-02-20
AI Technical Summary
Existing methods for non-invasive prenatal testing (NIPT) face a high failure rate due to low fetal DNA concentrations, particularly in early pregnancy, high BMI, fetal or placental abnormalities, and advanced maternal age, and require specialized tubes with fixatives that are expensive and not always accessible, complicating sample processing and storage.
A method involving blood collection in a device without a fixative, allowing storage at room or refrigerated temperatures for up to 9 days, followed by size-selective isolation of DNA less than 300 bp, enabling effective analysis of cell-free DNA (CFDNA) for fetal or cancer screening without the need for expensive stabilizing agents.
This approach enhances the quality and quantity of CFDNA for sequencing, improving the detection of fetal abnormalities and cancer markers, reducing the reliance on expensive fixative tubes and ensuring sample integrity over extended storage periods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the storage of blood, e.g., maternal blood or blood from cancer patients or patients being screened for cancer, and testing for abnormalities in plasma, e.g., maternal blood or plasma-derived cell free DNA (cfDNA) in blood from cancer patients or patients being screened for cancer. [Background technology]
[0002] Since their introduction in clinical care in 2011 (see Ehrich, 2011 and Palomaki, 2011), noninvasive prenatal testing (NIPT) has provided many pregnant women in different countries with information about the risk of fetal aneuploidy. Whatever platform used and whatever protocol followed, most laboratories report failure rates between 0 and 25% (Badeau, 2017). Failure is usually due to fetal fractions being too low to be reliable (Blais, 2018). Some reasons for low fetal fractions are known, such as early gestational age, high body mass index, the presence of some fetal or placental anomalies, T18, and gestational age (Hou, 2019; Ashoor, 2013).
[0003] Although longer gestational age is associated with an increased fetal fraction, screening tests are useful in the early stages of pregnancy, and increased fetal fraction is mainly observed in the later stages of pregnancy (Wang, 2013). At the molecular level, fetal cfDNA can be distinguished from maternal cfDNA by the fact that the fragments are shorter (Fan, 2010; Shi, 2020) and are hypermethylated and enriched at specific genomic locations (Wang, 2013). Recently, several studies have reported that the fetal fraction can be meaningfully increased by using high-concentration agarose gel electrophoresis to select the shortest fragments of libraries prepared from cfDNA before sequencing (Welker, 2021; Qiao, 2019; Xue, 2020; Qiao, 2019; Hu, 2019). This procedure also appeared to improve the detection of rare mutant alleles in association with cancer (Underhill, 2021).
[0004] Tumor circulating cell-free DNA has also been shown to be shorter than cfDNA from healthy cells (Jiang, 2015). In both applications, various systems have been used to isolate or enrich such short fragments by electrophoresis using precast gels, such as Invitrogen's 2% E-gel EX (Thermofisher), BluePippin's 2-3% agarose cassettes (Sage Bioscience), and Yourgene Health Canada, Inc.'s 3% agarose cassettes (Underhill, 2021).
[0005] To obtain a sufficient proportion of cfDNA for meaningful sequencing, plasma must be prepared with some care to avoid contamination with host genomic DNA. Blood samples need to be processed within 4-8 hours of collection, and an additional high-speed centrifugation step for plasma preparation may be preferable (Barrett, 2011). However, this is not always possible, as collection facilities are not always equipped to perform high-speed centrifugation of the plasma sample and then freeze the sample for transport to a central laboratory. The solution adopted by most blood collection facilities is to use specialized tubes designed to maintain the integrity of the cfDNA, for example by stabilizing the cell membrane using a crosslinking agent. An example of such specialized tubes is the Streck® blood collection tube. Long-term storage of blood without specialized tubes leads to the decay of genomic material, resulting in an increase in cfDNA over time (Fernando 2018). Specialized tubes stabilize the blood sample, allowing time to transfer the sample to a laboratory with sample processing facilities. However, such specialized tubes are expensive.
[0006] It has recently been reported that it is possible to obtain NIPT-grade cfDNA by molecular diagnostics using Vacutainer K2EDTA with gel to collect plasma (Giroux, 2021). The use of such tubes followed by filtration has been shown to be cost-effective and efficient in reducing time-consuming steps in laboratories. However, despite the fact that most laboratories are capable of centrifuging blood tubes, many do not have the equipment or staff to open blood tubes and filter the plasma.
[0007] Therefore, there is a need for alternative methods for storing whole blood samples and processing the blood to perform meaningful cfDNA sequence analysis. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Ehrich, 2011 and Palomaki, 2011 [Non-Patent Document 2] Badeau, 2017 [Non-Patent Document 3] Blais, 2018 [Non-Patent Document 4] Hou, 2019 [Non-Patent Document 5] Ashooor, 2013 [Non-Patent Document 6] Wang, 2013 [Non-Patent Document 7] Fan, 2010 [Non-Patent Document 8] Shi, 2020 [Non-Patent Document 9] Welker, 2021 [Non-Patent Document 10] Qiao, 2019 [Non-Patent Document 11] Xue, 2020 [Non-Patent Document 12] Hu, 2019 [Non-Patent Document 13] Underhill, 2021 [Non-Patent Document 14] Jiang, 2015 [Non-Patent Document 15] Barrett, 2011 [Non-Patent Document 16] Fernando 2018 [Non-Patent Document 17] Giroux, 2021 Summary of the Invention [Means for solving the problem]
[0009] An embodiment of the present disclosure relates to a method for analyzing nucleic acid sequences obtained from a blood sample to provide sequence information, comprising the steps of: collecting a blood sample into a blood collection device, the blood collection device being fixative-free; storing the blood without fixative at a temperature above -20°C and below 35°C (e.g., room temperature or refrigerated temperature) prior to isolating cell-free DNA from the plasma; separating the plasma from blood cells present in the blood sample either before or after storage; isolating cell-free DNA from the plasma more than 24 hours after blood collection from the subject; separating the cell-free DNA by size of the cell-free DNA to isolate cell-free DNA less than 300 bp; and analyzing the nucleic acid sequence of the isolated cell-free DNA to detect sequence information. The blood can be stored without fixative for an extended period of time before isolating cfDNA from the plasma. This period can be up to 3, 4, 5, 6, 7, 8, or 9 days after blood collection.
[0010] A collection device suitable for use according to the invention can include a container having only one or more compositions therein, the one or more compositions consisting of one or more anticoagulants and metal ions selected from potassium, lithium, and sodium. An example of such a device is a K2EDTA tube. Alternatively, the blood collection device can include a container having only one or more compositions therein, the one or more compositions consisting of one or more anticoagulants and metal ions selected from potassium, lithium, and sodium, and a separation gel. In other embodiments, the collection device can include a container having only one or more compositions therein, the one or more compositions consisting of or consisting of a coagulant.
[0011] As described above, the separation step can be used to isolate cfDNA less than 300bp in the sample of isolated cfDNA. In some embodiments, the cfDNA isolated in the size separation step is 185bp or less or 165bp or less. In other embodiments, the cfDNA isolated in the size separation step is 155bp or less or 150bp or less. In further embodiments, the portion of cfDNA isolated from plasma sample is 50bp or more or 80bp or more. The separation step can be carried out by various suitable polynucleotide size selection techniques known in the art, including gel electrophoresis, chromatography, bead-based separation, or membrane filters with pore sizes that prevent the passage of polynucleotides with sizes larger than a desired threshold.
[0012] After size selection, the nucleic acid sequence of the isolated cell-free DNA can be sequenced and analyzed to determine the presence of genetic abnormalities in fetal DNA, or the presence of cancerous cells in a subject can be confirmed based on the presence of one or more genetic markers associated with tumorigenesis or malignancy, or by determining sequence information. To obtain sequence information for the isolated cfDNA, it may be necessary to prepare a sequence library, and it may be necessary to amplify the cfDNA. Once size selection and library preparation are complete, the sample can be analyzed to obtain sequence information. The sequence information can be analyzed to detect mutations and chromosomal abnormalities, including, but not limited to, translocations, transversions, monosomies, trisomies, and other aneuploidies, deletions, insertions, methylations, amplifications, fragments, translocations, and rearrangements or chromosoliposis. The sequence information can also be used to detect any changes in gene sequences by comparing them with a reference sequence.
[0013] Another aspect of the present disclosure is sequence validity testing, which can detect abnormality in the performance of size selection step or unacceptable sample degradation.In a method comprising size selection of DNA sample comprising cfDNA and sequencing of sample DNA after size selection, the method can comprise sequence validity testing of sequence information obtained from sequencing step.The validity testing comprises: generating a first normalized fragment size profile from the sequencing information associated with the sample; comparing one or more values in the fragment size profile with one or more corresponding ranges of values obtained from a set of reference parameters comprising a plurality of valid fragment size profiles; and if one of the one or more values in the fragment size profile associated with the sample is within the corresponding range of values, the sequence information and its analysis result are accepted; or if one of the one or more values in the fragment size profile associated with the sample is outside the corresponding range of values, the sequence information is rejected. The validation may also include the steps of calculating a first relative fragment size frequency value from sequencing information associated with the sample, comparing the first relative fragment size frequency value to a set of reference relative fragment size frequency values comprising a plurality of valid relative fragment size frequencies, and accepting the sequence information and its analysis results if the first fragment size frequency associated with the sample is within a range of values specified in the reference set, or rejecting the sequence information and its analysis results if the first fragment size frequency associated with the sample is outside the range of values specified in the reference set. [Brief description of the drawings]
[0014] [Figure 1] 1 depicts plots showing the fetal fraction of various samples at various time points for samples stored in plasma preparation tubes or Streck® tubes and stored for 2-9 days. [Figure 2A] 1 depicts a plot showing the fetal fraction calculated from loss of chromosome X fragment in male pregnancies. [Figure 2B] Fetal fraction assessed by chromosome Y fragment count in boys only is presented. [Diagram 3] 1 depicts a plot illustrating the sizes of sequencing fragments calculated from paired-end sequencing mapping, with grey being fragments before size selection and blue being fragments after size selection. [Figure 4] Figure 1 depicts a study schematic for the study described in Example 3. At enrollment, each participant donated 20 mL of blood. Blood was collected in 5 mL EDTA vacutainers, transported to a storage location, and processed to the intended testing facility. Blood was stored in the vacutainers at ambient temperature for 0-8 hours, 72 hours, 120 hours, or 168 hours, and then processed by two rounds of centrifugation to obtain EDTA plasma. Plasma was stored at -80°C and processed at intervals once sufficient samples were collected to run. [Figure 5-1] ~ [Figure 5-2] Figure 1 shows one-way ANOVA of raw estimated fetal fraction for 268 validation samples after sequencing, showing fetal fraction estimation data by timepoint, where the mean across all timepoints is shown by a horizontal line, green represents ANOVA, and blue represents ± standard deviation. Also shown are ANOVA, mean, and standard deviation analyses. Timepoint 1 = 0-8 hours, timepoint 2 = 72 hours, timepoint 3 = 120 hours, and timepoint 4 = 168 hours. [Figure 6-1] ~ [Figure 6-2] Means, standard deviations, and Welch ANOVA for sample M056 (time point 1; 2.49 ng / μL) from run 5, along with the 275 validation libraries in this study. Post-PCR library quantification (ng / μL) data are presented by time point. Note: Four equal variance tests demonstrated that the variances of the data were not equal across time points, and thus Welch ANOVA was performed in place of ANOVA. Time point 1 = 0-8 hours, time point 2 = 72 hours, time point 3 = 120 hours, and time point 4 = 168 hours. Welch ANOVA P<0.0001 indicates significant variance in the library concentration data across the four time points of the study (P<0.05). [Figure 7A]Size profile distribution and frequency at each time point for patient sample M047, where time point 1 = red / TP1 (YGL013-6), time point 2 = orange / TP2 (YGL015-6), time point 3 = blue / TP3 (YGL038-6), and time point 4 = green / TP4 (YGL012-6). [Figure 7B] Size profile distribution and frequency at each time point for patient sample M100, where time point 1=red / TP1 (YGL026R-12), time point 2=orange / TP2 (YGL003-12), time point 3=blue / TP3 (YGL021-12), and time point 4=green / TP4 (YGL012-12). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] A method for storing blood containing cfDNA to be sequenced is described herein.According to the present disclosure, after storing whole blood sample in a fixative-free blood collection device for an extended period of time of more than 24 hours, cfDNA is isolated from blood or plasma to obtain sequence information.The method includes carrying out a size selection step of the isolated DNA to select cfDNA that is less than a certain size, for example, 185bp or less, and then obtaining sequence information for the size-selected cfDNA.For example, by selecting cfDNA that has a nucleotide length of 185bp or less, the fraction of fetal cfDNA in cfDNA sample is increased, thereby improving the quality of the sample and enabling meaningful sequence analysis.
[0016] definition "Blood sample," as used herein, refers to a whole blood sample that has not been fractionated or separated into its component parts.
[0017] "Fixation" is a technique that helps maintain the structure of cells and / or subcellular components, such as organelles (e.g., nuclei), in a blood sample. Fixation modifies the chemical or biological structure of cellular components, for example, by cross-linking them. Fixation prevents lysis of whole cells and organelles, thereby inhibiting the release of cellular nucleic acids into the surrounding medium. For example, fixation can inhibit the release of nuclear DNA from white blood cells into the plasma fraction during centrifugation of whole blood.
[0018] A "fixative" is a substance used to fix a blood sample.
[0019] "Sequence of interest" refers to a nucleic acid sequence associated with a difference in sequence representation in healthy vs. patients. A sequence of interest may be a sequence on a chromosome that is misrepresented, i.e., over- or under-represented, in a pathological or genetic condition. A sequence of interest may be a portion of a chromosome, i.e., a chromosomal segment, or an entire chromosome. For example, a sequence of interest may be a chromosome that is over-represented in an aneuploidy condition, or a gene that encodes a tumor suppressor that is under-represented in cancer. A sequence of interest includes sequences that are over- or under-represented in the entire population or subpopulation of cells of a subject. A sequence of interest may also be referred to as a "marker" of a genotype associated with a particular phenotype, e.g., disease.
[0020] "Sequence tag" refers to the sequence of at least a 30 bp read from a cfDNA strand. In some embodiments, the tag is specifically assigned, i.e., mapped, by alignment to a larger sequence, e.g., a reference genome.
[0021] Methods for increasing and analyzing a fraction of cfDNA The present disclosure is directed to a method for elevating a fraction of plasma, e.g., a maternal blood sample or a blood sample for cancer diagnosis, and analyzing the nucleic acid sequence of the fraction to provide sequence information. The method includes collecting a blood sample in a blood collection device. The blood collection device does not include a fixative. In some embodiments, the blood collection device comprises a container having one or more anticoagulants and a metal ion selected from potassium, sodium, and lithium. In some embodiments, the blood collection device comprises a container having one or more anticoagulants and a metal ion, e.g., potassium, sodium, or lithium, and a separation gel.
[0022] The blood sample is centrifuged before or after storage without fixative to separate the plasma from the blood cells. Before or after separating the plasma, the blood sample can be stored at a temperature ranging from -20°C to 35°C, e.g., room or ambient temperature, refrigerated temperature, or room to refrigerated temperature, until the sample is ready for further processing. When the sample is ready for processing 24 to 216 hours after blood collection, e.g., 48 / 72 to 216 hours, or 48 / 72 to 192 hours, or 48 / 72 to 168 hours, or 48 / 72 to 144 hours, cfDNA is isolated from the plasma. To increase the fraction of cfDNA, the cfDNA is separated by size, e.g., by gel electrophoresis, to isolate cfDNA less than 300 bp (e.g., less than 185, 165, 160, 155, or 150 bp). Either before or after size selection, the cfDNA sample can be prepared and optionally amplified and sequenced to generate a sequence library containing cfDNA. The size-selected and isolated cfDNA can then be analyzed to detect sequence information. The steps described herein can be performed on a single entity or on multiple entities.
[0023] Blood collection device The blood sample is collected in a collection device. The collection device can be a vacuum collection container, typically a tube. According to the present disclosure, the blood collection device does not contain a fixative. Examples of suitable blood collection devices include K2 or K3 EDTA blood collection tubes with or without separation gel.
[0024] As defined above, a fixative is a substance used to fix a blood sample. Examples of fixatives include aldehydes (e.g., formaldehyde, glutaraldehyde, and their derivatives), formalin, alcohols, sulfhydryl-reactive crosslinkers, carbohydrate-reactive crosslinkers, carboxyl-reactive crosslinkers, photoreactive crosslinkers, cleavable crosslinkers, AEDP (3-[(2-aminoethyl)dithio]propionic acid HCl), APG (p-azidophenylglyoxal monohydrate), BASED, BM(PEO)3 (1,8-bis-maleimide), and the like. -PEO3), BM(PEO)4 (1,8-bis-maleimido-PEO4), BMB (1,4-bismaleimidobutane), BMDB (1,4-bis-maleimidyl-2,3-dihydroxybutane), BMH (bis-maleimidohexane), BMOE (bis-maleimidoethane), BS3 (bis(sulfosuccinimidyl)suberic acid), BSOCOES (bis(2-[succinimidoxycarbonyloxy]ethyl)sulfone), DFDNB (1-5- Difluoro-2,4-dinitrobenzene), DMA (dimethyladipimidic acid·2HCl), DMP (dimethylpimelimidic acid·2HCl), DMS (dimethylsuberimidic acid·2HCl), DPDPB (1,4-di-(3'-[2'pyridyldithio]propionamido)butane), DSG (disuccinimidyl glutaric acid), DSP (dithiobis(succinimidyl propionic acid) (Lomant's reagent)), DSS (disuccinimidyl suberic acid) , DST (disuccinimidyl tartrate), DTBP (dimethyl 3,3'-dithiobispropionimidate·2HC), DTME (dithiobis-maleimidoethane), DTSSP (3,3'-dithiobis(sulfosuccinimidyl propionate)), EGS (ethylene glycol bis(succinimidyl succinate)), HBVS (1,6-hexane-bis-vinyl sulfone), sulfo-BSOCOES, sulfo-DST, and sulfo-EGS.
[0025] Other fixatives include cell membrane stabilizers such as DMAE (dimethylaminoethanol), cholesterol, cholesterol derivatives, high concentrations of magnesium, vitamin E and vitamin E derivatives, calcium, calcium gluconate, taurine, niacin, hydroxylamine derivatives, bimoclomol, sucrose, astaxanthin, glucose, amitriptyline, isomer A hopantethalol phenylacetate, isomer B hopantethalol phenylacetate, citicoline, inositol, vitamin B, vitamin B complex, cholesterol hemisuccinate, sorbitol, calcium, coenzyme Q, ubiquinone, vitamin K, vitamin K complex, menaquinone, zonegran, zinc, ginkgo biloba extract, diphenylhydantoin, perftoran, polyvinylpyrrolidone, phosphatidylserine, tegretol, PABA, disodium cromoglycate, nedocromil sodium, phenyloin, zinc citrate, mexitil, dilantin, sodium hyaluronate, or poloxamer 188.
[0026] Other fixatives are formaldehyde-releasing preservatives such as diazolidinyl urea, imidazolidinyl urea, dimethylol-5,5-dimethylhydantoin, dimethylol urea, 2-bromo-2.-nitropropane-1,3-diol, oxazolidine, sodium hydroxymethylglycinate, 5-hydroxymethoxymethyl-1-1aza-3,7-dioxabicyclo[3.3.0]octane, 5-hydroxymethyl-1-1aza-3,7dioxabicyclo[3.3.0]octane, 5-hydroxypoly[methyleneoxy]methyl-1-1aza-3,7dioxabicyclo[3.3.0]octane, quaternary adamantines, and any combination thereof.
[0027] In further embodiments, other than EDTA and / or heparin, the blood collection device may be free of substances that reduce DNA degradation, such as nuclease inhibitors. Examples of nuclease inhibitors include diethylpyrocarbonate, ethanol, aurintricarboxylic acid (ATA), formamide, vanadyl-ribonucleoside complex, macaloid, proteinase K, hydroxylamine-oxygen-cupric ion, bentonite, ammonium sulfate, dithiothreitol (DTT), beta-mercaptoethanol, cysteine, dithioerythritol, tris(2-carboxyethyl)phosphine hydrochloride, and divalent cations, such as Mg+2, Mn+2, Zn+2, Fe+2, Ca+2, Cu+2. Additionally, other substances that reduce DNA degradation, such as zinc chloride, guanidine HCl, guanidine isothiocyanate, N-lauroyl sarcosine, and Na-dodecyl sulfate, may be omitted from the device.
[0028] The present disclosure contemplates excluding any particular combination of the above fixatives, or any particular combination of the above fixatives and / or agents that reduce DNA degradation, such as nuclease inhibitors.
[0029] In some embodiments, the collection device includes an anticoagulant. The anticoagulant may be selected from ethylenediaminetetraacetic acid (EDTA), heparin, citric acid, oxalic acid, and any combination thereof. In some embodiments, the anticoagulant may be spray dried onto the walls of the collection device. In other embodiments, the anticoagulant is in solution within the device. Suitable solvents include water, saline, dimethylsulfoxide, alcohol, and mixtures thereof. In an embodiment, the collection device includes a container having only one or more compositions, the one or more compositions consisting of an anticoagulant and a metal ion, such as an alkali or alkaline earth metal. In particular, the metal ion may be selected from calcium, sodium, lithium, and potassium.
[0030] In embodiments using EDTA, the EDTA may be chelated with a metal ion selected from potassium, lithium, and sodium, typically potassium. In some embodiments, the blood collection device includes a container having only one or more compositions, where the one or more compositions consist of chelated EDTA.
[0031] In some embodiments, the collection device includes a separation gel. Such gels are commonly used in serum separation tubes (SST) to perform plasma preparation (see, for example, BD Vacutainer® SST™ and PST™, Becton, Dickinson, and Company). The gel has an intermediate density relative to that of blood cells and the liquid phase of blood (densities of 1.09 g / cm3 and 1.03 g / cm3, respectively). Due to this density, the separation gel physically separates the liquid component of blood (plasma) from red blood cells and white blood cells after centrifugation and acts as a barrier between the liquid phase and the cells. The separation gel is an inert gel, for example, a polyester gel or a silicone gel. Other suitable gels include gelling agents based on sorbitol in diacrylic acid oligomers.
[0032] In some embodiments, the blood collection device includes a container having only one or more compositions, the one or more compositions consisting of an anticoagulant and a metal ion, e.g., chelated EDTA, and a separating gel.
[0033] In some embodiments, the collection device comprises a coagulant.
[0034] Preservation of blood collection devices and removal of large molecules from plasma In some embodiments, after taking a blood sample from a patient, the blood can be stored at a temperature between -20°C and 35°C, such as standard ambient temperature (25°C), room temperature (20-24°C), refrigerated temperature (1-8°C), or any temperature in between, for a period of up to 7-9 days, followed by centrifugation to separate the plasma from the cellular fraction and extract the DNA. The storage time before centrifugation can be 2, 3, 4, 5, 6, 7, 8, or 9 days. The storage temperature can be 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, or some value or range therebetween. In other embodiments, the storage temperature can be within the range of 9-23°C, 20-25°C, 25-30°C, or 30-35°C. In some embodiments, the storage time is 7 days or less, or 2-5 days, 2-6 days, 2-7 days, 3-5 days, 3-6 days, 4-6 days, 3-7 days, or 4-7 days. In some embodiments, the storage time is 2-7 days or 3-7 days, and is stored at refrigerated temperature, room temperature, or some temperature in between.
[0035] Upon removal from storage, e.g., from the refrigerator, the sample can be fractionated to remove large size particulates. Such techniques are described in Pos et al., Int. J. Mol. Sci. 2020, 21, 8634, entitled "Technical and Methodological Aspects of Cell-Free Nucleic Acids Analyzes", which is incorporated herein by reference. In one embodiment, this is done by centrifugation. Centrifugation can be a one- or two-step process. In some embodiments, centrifugation is done at refrigerated temperatures. In other embodiments, it is done at room temperature. In the first centrifugation step, the sample is subjected to low-speed centrifugation. Plasma is collected from the low-speed fraction sample. In the second step, the plasma is subjected to a low-speed or high-speed centrifugation step. Standard practice is to perform a high-speed step. The low-speed centrifugation parameters are those commonly used to separate plasma from cellular components in a blood sample. For example, in the low-speed step, the rotation speed may be 1,000-2,000×g, and the rotation time may be 8, 9, 10, 11, 12 minutes, or any part of the time therebetween or ranges thereof. In some embodiments, the rotation speed may be 1,500-1,900×g or 1,600-1,800×g, and the rotation time may be 9-11 minutes or 10 minutes. In the high-speed step, the rotation speed may be approximately 3-10 times that of the low-speed step for the same time. For example, the high-speed centrifugation may be 3,000-18,000×g.
[0036] Alternatively, instead of the second centrifugation step, a microfiltration step can be used to remove larger particles from the plasma. In some embodiments, the microfiltration step comprises a membrane separation process that removes particles with an average molecular weight of more than 200 kDa, more than 100 kDa, more than 75 kDa, or more than 65 kDa using a membrane with a pore size of less than 3, 2, or 1 μm. In some embodiments, the microfiltration membrane filter has a pore size of 0.55 μM, 0.50 μM, or 0.45 μM or less.
[0037] In other embodiments, the blood sample is collected in a collection device, for example a collection device having a separating gel, and centrifuged at low speed immediately after collection to separate the plasma from the cellular fraction of the blood sample. In embodiments, this step may be performed within 6-8 hours of sample collection, after which the collection device with the fractionated sample therein may be stored (i.e., no immediate isolation of the plasma is required). Alternatively, the plasma may be removed from the collection device and stored in a second collection device that does not contain a fixative. The low speed centrifugation parameters are the same as those described above. The storage temperature and storage time are the same as those described above. If not removed prior to storage, upon removal from storage, the plasma may be removed from the collection device and subjected to a second centrifugation or filtration step according to the parameters described above.
[0038] cfDNA isolation After collecting the plasma fraction as described above, cfDNA is isolated / extracted from the plasma. Extraction is actually a multi-step process, including separating DNA from plasma onto a column or other solid-phase binding matrix. The extracted cfDNA usually contains both maternal and fetal cfDNA. In some embodiments, isolation can be performed by the following operations: A. Denature and / or degrade proteins in the plasma (e.g., by contacting with a protease) and add guanidine hydrochloride or other chaotropic agent to the solution (to facilitate extraction of cfDNA from the solution). B. The treated plasma is contacted with a coated support matrix, such as beads or superparamagnetic particles within a membrane or column, to selectively bind nucleic acids. Such coatings may include silica, carboxyl, amine, imidazole, or combinations thereof. C. Washing the support matrix. D. Release the cfDNA from the matrix with a release solution and recover the cfDNA for downstream processing (e.g., preparation of indexed libraries) and statistical analysis.
[0039] The first part of this cfDNA isolation procedure involves denaturing or degrading nucleosomal proteins or otherwise releasing DNA from nucleosomes. A typical reagent mixture used to achieve this isolation includes detergents, proteases, and chaotropic substances, such as guanine hydrochloride. Proteases act to degrade nucleosomal proteins as well as background proteins in plasma, such as albumin and immunoglobulins. Chaotropic substances disrupt the structure of macromolecules by interfering with intramolecular interactions mediated by non-covalent forces, such as hydrogen bonds. Chaotropic substances also cause plasma components, such as proteins, to be negatively charged. The negative charge makes the medium somewhat energetically incompatible with negatively charged DNA. A method for using chaotropic substances to facilitate DNA purification is described in Boom et al., "Rapid and Simple Method for Purification of Nucleic Acids", J. Clin. Microbiology, v. 28, No. 3, 1990.
[0040] After this proteolytic treatment, which at least partially releases the DNA coils from the nucleosomal proteins, the resulting solution is passed through a column or otherwise exposed to a support matrix to which the cfDNA in the treated plasma selectively adheres. The negative charges imparted to the media components facilitate adsorption of the DNA into the pores of the support matrix.
[0041] After the treated plasma has passed through the support matrix, the support matrix with bound cfDNA is washed to remove proteins and other unwanted components of the sample. After washing, the cfDNA is released from the matrix and recovered. For example, for beads with amine or imidazole functional groups, in acidic solutions, cfDNA is adsorbed to the positively charged amine or imidazole moieties. In basic solutions, the amine and imidazole moieties are negatively charged and cfDNA is not adsorbed. Therefore, to remove unwanted components of the sample, the magnetic beads are immobilized with a magnet and the beads are washed with an acidic wash. After washing, the beads can be separated from the magnet and the cfDNA can be recovered by washing with a high pH solution. A similar process is used for carboxyl-coated beads. For carboxyl-coated beads, cfDNA is reversibly bound to the beads in the presence of polyethylene glycol (PEG), depending on its salt concentration.
[0042] Such DNA extraction processes typically result in a significant loss of a fraction of available DNA from the plasma. The support matrix has a high affinity for cfDNA, which limits the amount of cfDNA that can be easily separated from the matrix. As a result, the yield of the cfDNA extraction step can be quite low. Typically, the efficiency is much less than 50% (e.g., typical yields of cfDNA have been found to be 4-12 ng / ml of plasma from approximately 30 ng / ml of available plasma).
[0043] Commercially available kits for manual and automated cfDNA isolation can be used, such as QIAamp DNA Micro kit (Qiagen, Valencia, CA) or NucleoSpin Plasma kit (Macherey-Nagel, Duren, Germany). The process of isolating cfDNA is described in U.S. Patent No. 10,837,055, the entirety of which is incorporated herein by reference.
[0044] Preparation of sequence libraries The purified cfDNA obtained from the above isolation / extraction procedure can be used to prepare a library for sequencing. In order to sequence a population of double-stranded DNA fragments using a massively parallel sequencing system, the DNA fragments must be flanked by known adapter sequences. Such a collection of DNA fragments with adapters at either end is called a sequencing library. Two examples of methods suitable for generating a sequencing library from purified DNA are (1) ligation-based attachment of known adapters to either end of fragmented DNA, and (2) transposase-mediated insertion of adapter sequences.
[0045] Methods for generating sequence libraries are described in U.S. Patent Application Publication No. 2007 / 0128624 to Illumina Cambridge, Ltd., entitled "Methods for preparing libraries of template polynucleotides"; U.S. Patent Application Publication No. 2013 / 0123120 to Natera, Inc., entitled "Highly Multiplex PCR Methods and Compositions"; J.F. Hess, et al., Library preparation for next generation sequencing: A review of automation strategies, Biotechnology Advances, Volume 41, 2020, 107537, ISSN 0734-9750, https: / / doi.org / 10.1016 / j.biotechadv.2020.107537; and Head, Steven R et al. Library construction for next-generation sequencing: overviews and challenges, BioTechniques vol. 56,2 61-4, 66, 68, passim. 1 Feb. 2014, doi:10.2144 / 000114133, all of which are incorporated by reference in their entireties. Those of skill in the art will appreciate that there are many suitable massively parallel sequencing technologies.
[0046] Sequencing libraries can be normalized to facilitate volumetric pooling of samples using techniques known to those of skill in the art.
[0047] Amplification of cfDNA In some embodiments, the cfDNA present in the sample can be amplified either before or after preparing the sequence library. Amplification can be performed by any suitable technique, for example, a method based on polymerase chain reaction (PCR) or recombinase polymerase amplification (RPA). Such techniques are known to those skilled in the art. RPA is described in Lobato, et al., Recombinase polymerase amplification: Basics, applications and recent advances, Trends in analytical chemistry: TRAC vol. 98 (2018): 19-35. doi:10.1016 / j.trac.2017.10.015, the entirety of which is incorporated herein by reference.
[0048] To amplify a segment of DNA using PCR, a sample is first heated to denature the DNA or separate it into two pieces of single-stranded DNA. An enzyme called "Taq polymerase" then synthesizes, or builds, two new DNA strands using the original strand as a template. This process replicates the original DNA, with each new molecule containing one old and one new DNA strand. Each such strand can then be used to generate two new copies, and so on. The cycle of denaturing and synthesizing new DNA is repeated. In some embodiments, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification cycles are performed. PCR is described in more detail in Lorenz, Todd C., Polymerase chain reaction: basic protocol plus troubleshooting and optimization strategies, Journal of Visualized Experiments, 63 e3998. 22 May. 2012, doi:10.3791 / 3998, which is incorporated herein by reference in its entirety.
[0049] Size Selection Size selection is performed before or after amplification to increase the percentage of fetal fraction in the sample, thereby increasing the signal when sequencing. The desired maximum sequence length of cfDNA isolated by size selection can be 185bp, 180bp, 170bp, 175bp, 170bp, 165bp, 160bp, 155bp, 150bp, or 145bp, or some value therebetween. In some embodiments, the desired size of cfDNA is in the range of 80-160bp, 80-155bp, 80-150bp, 90-160bp, 90-155bp, 90-150bp, 100-150bp, 100-155bp, 100-160bp, 110-150bp, 110-155bp, 110-160bp, 80-170bp, or 80-185bp. In some embodiments, the desired size of the cfDNA is in the range of 100-200 bp, 150-250 bp, 200-300 bp, and 250-300 bp.
[0050] Size selection can be performed by any technique for separating macromolecules by size, such as gel electrophoresis, chromatography, bead-bound matrices, or membrane filtration. In the case of gel electrophoresis, the desired size range can be collected manually or automatically using an instrument, such as LightBench or NIMBUS Select from Yourgene Health Canada Inc.
[0051] Manual methods for recovering DNA from electrophoresis gels include (1) cutting a slice of the gel corresponding to a band of a DNA molecule of a particular length, followed by extraction of the DNA from the gel slice and precipitation of the DNA, and (2) taking an aliquot or set of aliquots directly from the electrophoresis gel at a pipette-accessible point of the gel, known in the art as an "extraction well," followed by precipitation of the DNA.
[0052] An automated electrophoresis system that can be used for size selection is described in U.S. Patent No. 10,775,344 owned by Yourgene Health Canada Inc., which is incorporated herein by reference in its entirety. The electrophoresis system may be configured to distribute power to each gel channel, and may adjust the power for each channel based on monitoring of the current feedback. In one embodiment, a system includes interfacing cassettes including a plurality of channels, the system including a robotic workstation for receiving at least one control signal from an external computing device for controlling an electrophoresis operation, the robotic workstation including an arm including at least one pipette and an on-board power module, a first modular platform for electrically coupling and receiving on the robotic workstation, an interfacing cassette stored on or located on the platform, the interfacing cassette including a plurality of channels, a pair of electrical contacts coupled to each of the plurality of channels, the first modular platform including a plurality of conductive cables or wires for distributing an independently controllable power signal from the power module to each channel, the platform including a processor and memory, the processor configured to receive power signals from the power module, to receive at least one of the control signals assigned to the interfacing cassettes of the first modular platform, and to adjust the power signals based on the control signals to generate independently controllable power signals defined for each channel of the interfacing cassettes of the first modular platform. In an embodiment, the processor is configured to further adjust the power signal based on electrophoretic parameters stored on the memory and associated with each of the cassette channels. The system may further include a power transfer module electrically coupled to at least a second modular platform and a power module, the power transfer module configured to broadcast and communicate power signals between the power module and the first modular platform and at least the second modular platform. The control signal is capable of uniquely identifying each of the cassette channels by multiplexing.The conditioning of the power signal may include conditioning using at least one of analog voltage control, pulse width conditioning, duty cycle control for pulse width conditioning, and frequency control for pulse width conditioning. In certain embodiments, the processor is further configured to (i.) monitor a current feedback of each of said channels, and (ii.) adjust a conditioned power signal of each of said channels based on said monitoring.
[0053] Other electrophoretic techniques are described in: capillary electrophoresis with self-coating low viscosity polymer matrices is described in Du M, Flanagan JH Jr, Lin B, Ma Y, Electrophoresis 2003 September; 24 (18): 3147-53), which is incorporated herein by reference in its entirety, selective extraction with microfabricated electrophoretic devices is described in Lin R, Burke DT, Burn MA, J. Chromatogr. A. 2003 Aug. 29; 1010(2): 255-68), which is incorporated herein by reference in its entirety, and microchip electrophoresis on reduced viscosity polymer matrices is described in Xu F, Jabasini M, Liu S, Baba Y, Analyst. 2003 June; 128(6): 589-92), which is incorporated herein by reference in its entirety.
[0054] Size selection of cfDNA can also be performed by chromatography, e.g., chromatography on agarose or polyacrylamide gels, ion-pair reversed-phase high performance liquid chromatography (IP RP HPLC, see Hecker KH, Green SM, Kobayashi K, J. Biochem. Biophys. Methods 2000 Nov. 20; 46(1-2): 83-93, which is incorporated herein by reference in its entirety), adsorption membrane chromatography (see Teeters MA, Conrardy SE, Thomas BL, Root TW, Lightfoot EN, J. Chromatogr. A. 2003 Mar. 7; 989(1): 165-73, which is incorporated herein by reference in its entirety), density gradient centrifugation (Raptis L, Menard HA, J. Clin. Invest. 1980 December; 66(6): 1391-9, the entirety of which is incorporated herein by reference), and nanotechnology-based methods, such as microfabricated entropy trap arrays (Han J, Craighead HG, Analytical Chemistry, Vol. 74, No. 2, Jan. 15, 2002, the entirety of which is incorporated herein by reference). Further usable bead-based techniques for size selection of cfDNA are described in Raymond et al., UltraPrep is a scalable, cost-effective, bead-based method for purifying cell-free DNA. PLoS ONE 15(6) (2020): e0231854. https: / / doi.org / 10.1371 / journal.pone.0231854, the entirety of which is incorporated herein by reference.
[0055] If a size selection step is performed after sequence library preparation, the desired cfDNA length maximum (e.g., less than 160 bp) or length range will be the same, but in order to isolate the appropriate bands resulting from the size selection device, the target maximum or range to be isolated must be adjusted to account for the adapters or other polynucleotides that are ligated onto or inserted into the cfDNA. For example, if the use of adapters adds 100 bp to the length of the cfDNA polynucleotides and the target maximum is less than 160 bp, the size selection parameter will be less than 260 bp.
[0056] By performing a size selection step on a sample from a pregnant woman, the percentage of fetal cfDNA in the sample analyzed for sequence information is increased by 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190% or 200% compared to the percentage before size selection. As shown in Example 1 and Figure 3, the percentage of fetal cfDNA estimated by SeqFF is nearly identical to or exceeds the percentage of fetal cfDNA in samples stored in Streck® tubes for the same period (SeqFF is the fetal fraction expressed as a percentage of the total cfDNA sequenced, calculated from sequence read counts. See Kim SK, et al.: Determination of fetal DNA fraction from the plasma of pregnant women using sequence read counts. Prenat Diagn 2015, 35:810-815, the entirety of which is incorporated herein by reference). Similarly, when examining samples from only male infants, size-selected samples have a high percentage of cfDNA for all samples.
[0057] Similarly, by performing a size selection step on a sample for cancer diagnosis, the percentage of cancer cell-derived cfDNA in the sample for which sequence information is analyzed increases by 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190% or 200% compared to the percentage before size selection.
[0058] Analyzing cfDNA to obtain sequence information After size selection and library preparation are completed, the sample can be analyzed to obtain sequence information.Sequence information can be analyzed to detect mutations and chromosomal abnormalities, including but not limited to translocations, transversions, monosomies, trisomies, and other aneuploidies, deletions, insertions, methylations, amplifications, fragments, translocations, and rearrangements or chromosomal lysis.Sequence information can also be used to detect any changes in gene sequences by comparing them with reference sequences.
[0059] In some embodiments, the sample is sequenced as part of a procedure to determine sequences of interest and to assess copy number variation (CNV). Any of a number of sequencing technologies can be used, including "first generation" sequencing, next generation sequencing, long read sequencing, or nanopore sequencing.
[0060] Some sequencing technologies are commercially available, such as the sequencing by hybridization platform from Affymetrix Inc. (Sunnyvale, Calif.), the sequencing by synthesis platform from Illumina, Inc. (Hayward, Calif.), and nanoball-based sequencing from Complete Genomics (San Jose, Calif.). Other single molecule sequencing technologies include, but are not limited to, Pacific Biosciences' SMRT™ technology (see Anthony Rhoads et al., PacBio Sequencing and Its Applications, Genomics, Proteomics & Bioinformatics, Volume 13, Issue 5, 2015, page 278-289, ISSN 1672-0229, https: / / doi.org / 10.1016 / j.gpb.2015.08.002, which is incorporated herein by reference in its entirety), ION TORRENT™ technology, and nanopore sequencing developed, for example, by Oxford Nanopore Technologies (see Soni GV and Meller A. Clin Chem 53: 1996-2001
[2007] , which is incorporated herein by reference in its entirety).
[0061] Although automated Sanger sequencing is considered to be a "first generation" technology, Sanger sequencing, including automated Sanger sequencing, can also be utilized in the methods described herein.Suitable additional sequencing methods include, but are not limited to, nucleic acid imaging techniques, such as atomic force microscopy (AFM) or transmission electron microscopy (TEM).
[0062] As an example, cfDNA can be sequenced using Illumina's sequencing technology. Illumina's technology relies on the binding of cfDNA to a planar, optionally transparent, surface to which an oligonucleotide anchor is attached. The template DNA is end-repaired to generate 5' phosphorylated blunt ends, and the polymerase activity of Klenow fragment is used to add a single A base to the 3' end of the blunt phosphorylated DNA fragment. This addition prepares the DNA fragment to be ligated to an oligonucleotide adaptor with a single T base overhang at the 3' end to enhance ligation efficiency. The adaptor oligonucleotide is complementary to the flow cell anchor. Under limiting dilution conditions, the adaptor-modified single-stranded template DNA is added to the flow cell and fixed by hybridization to the anchor. The bound DNA fragments are extended and bridge-amplified to create an ultra-high density sequencing flow cell with hundreds of millions of clusters, each of which contains 1,000 copies of the same template. In one embodiment, the cfDNA is amplified using PCR, and then subjected to cluster amplification. Alternatively, cfDNA is enriched using cluster amplification only using unamplified genomic library preparation (Kozarewa et al., Nature Methods 6:291-295 (2009)). The templates are sequenced using a robust four-color DNA sequencing by synthesis technology that utilizes reversible terminators with removable fluorescent dyes. Highly sensitive fluorescent detection is achieved using laser excitation and total internal reflection optics. Short sequence reads of about 20-300 bp, e.g., 150 bp, may be aligned to a reference genome with masked repetitive sequences, and unique mapping of the short sequence reads to the reference genome is identified using specially developed data analysis pipeline software. A reference genome that does not mask repetitive sequences may also be used. Only reads that uniquely map to the reference genome are counted, regardless of whether the repetitive sequences of the reference genome used are masked or not. After completion of the first read, the template can be regenerated in situ to allow a second read from the opposite end of the fragment. Thus, both single-end and paired-end sequencing of DNA fragments can be used.Partial sequencing of DNA fragments present in the sample is performed, and sequence tags containing reads of a given length, for example 36 bp, that are mapped to a known reference genome are counted. In one embodiment, the reference genome sequence is the NCBI36 / hg18 sequence, available on the World Wide Web at genome.ucsc.edu / cgi-bin / hgGateway?org=Human&db=hg18&hgsid=166260105. Alternatively, the reference genome sequence is GRCh37 / hg19, available on the World Wide Web at genome.ucsc.edu / cgi-bin / hgGateway. Other sources of public sequence information include GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (DNA Data Bank of Japan). Many computer algorithms are available for aligning sequences, including, but not limited to, BLAST (Altschul et al., 1990), BLITZ (MPsrch) (Sturrock & Collins, 1993), FASTA (Person & Lipman, 1988) or BOWTIE (Langmead et al., Genome Biology 10:R25.1-R25.10 (2009)), all of which are incorporated herein by reference in their entirety.
[0063] As another example, sequence information of nucleic acids in a test sample, e.g., cfDNA in a maternal test sample or cfDNA in a subject being screened for cancer, can be obtained using 454 sequencing (Roche) (e.g., Margulies, M. et al. Nature 437:376-380
[2005] , the entirety of which is incorporated herein by reference). 454 sequencing typically involves two steps. In the first step, DNA is sheared into fragments of approximately 300-800 base pairs and the ends of the fragments are blunted. Oligonucleotide adaptors are then ligated to the ends of the fragments. The adaptors act as primers for amplification and sequencing of the fragments. The fragments can be attached to DNA capture beads, e.g., streptavidin-coated beads, e.g., using adaptor B, which contains a 5' biotin tag. The bead-bound fragments are PCR amplified in droplets of an oil-water emulsion. This results in multiple copies of clonally amplified DNA fragments on each bead. In the second step, the beads are captured in wells (e.g. picoliter-sized wells). Pyrosequencing is performed in parallel for each DNA fragment. The addition of one or more nucleotides generates a light signal, which is recorded by a CCD camera in the sequencing instrument. The signal intensity is proportional to the number of nucleotides incorporated. Pyrosequencing utilizes pyrophosphate (PPi), which is released upon nucleotide addition. PPi is converted to ATP by ATP sulfurylase in the presence of adenosine 5' phosphosulfate. Luciferase uses ATP to convert luciferin to oxyluciferin, a reaction that generates light, which is measured and analyzed.
[0064] In another exemplary, but non-limiting embodiment, sequence information can be obtained by performing sequencing by ligation, in which genomic DNA is sheared into fragments and adapters are attached to the 5' and 3' ends of the fragments to generate a fragment library. Alternatively, internal adapters can be introduced by ligating adapters to the 5' and 3' ends of the fragments, circularizing the fragments, digesting the circularized fragments to generate internal adapters, and attaching adapters to the 5' and 3' ends of the resulting fragments to generate a mate pair library. A clonal bead population is then prepared in a microreactor containing beads, primers, templates, and PCR components. After PCR, the templates are denatured and the beads are enriched to separate beads with extended templates. The templates on selected beads are subjected to 3' modification, which allows them to be attached to a glass slide. The sequence can be determined by sequential hybridization and ligation of partially random oligonucleotides to a central determining base (or base pair) identified by a specific fluorophore. After recording the color, the ligated oligonucleotides are cleaved and removed, and the process is then repeated.
[0065] In another embodiment, sequence information of nucleic acids in test samples, for example cfDNA in maternal test samples, cfDNA in subjects to be screened for cancer, can be obtained using Pacific Biosciences' Single Molecule Real-Time (SMRT™) sequencing technology. In SMRT sequencing, the continuous incorporation of dye-labeled nucleotides is imaged during DNA synthesis. Single DNA polymerase molecules are attached to the bottom surface of individual zero mode wavelength detectors (ZMW detectors), which obtain sequence information while phosphate-linked nucleotides are incorporated into the growing primer strand. The ZMW detectors contain a confined structure, which allows the observation of the incorporation of a single nucleotide by the DNA polymerase against a background of fluorescent nucleotides that rapidly diffuse (e.g., within a few microseconds) in and out of the ZMW. Incorporation of a nucleotide into the growing strand typically takes several milliseconds. During this time, the fluorescent label is excited, generating a fluorescent signal and cleaving the fluorescent tag. The corresponding measurement of dye fluorescence indicates the incorporated base. This process is repeated to allow sequencing.
[0066] In another embodiment, sequence information of nucleic acids in a test sample, for example cfDNA in a maternal test sample or cfDNA in a subject to be screened for cancer, is obtained by using nanopore sequencing. Nanopore sequencing DNA analysis technology has been developed by many companies, including, for example, Oxford Nanopore Technologies (Oxford, UK), Sequenom, and NABsys. Nanopore sequencing is a single molecule sequencing technology, whereby a single molecule of DNA is sequenced directly as it passes through a nanopore. A nanopore is a small hole, typically on the order of one nanometer in diameter. When a nanopore is immersed in a conducting fluid and a potential (voltage) is applied to it, a small current is generated by ions conducting through the nanopore. The amount of current that flows is sensitive to the size and shape of the nanopore. As the DNA molecule passes through the nanopore, each nucleotide on the DNA molecule obstructs the nanopore to a different extent, changing the magnitude of the current passing through the nanopore to a different extent. This change in electrical current as a DNA molecule passes through the nanopore therefore allows the DNA sequence to be read.
[0067] In another embodiment, sequence information of nucleic acids in a test sample, for example cfDNA in a maternal test sample or cfDNA in a subject to be screened for cancer, can be obtained by using a chemical-sensitive field effect transistor (chemFET) array (e.g., as described in US Patent Application Publication No. 2009 / 0026082, the entirety of which is incorporated herein by reference). In one example of this technology, DNA molecules can be placed in a reaction chamber, and a template molecule can be hybridized to a sequencing primer that binds to a polymerase. The incorporation of one or more triphosphates into the 3' end of the sequencing primer of a new nucleic acid strand can be discerned by the chemFET as a change in current. The array can have multiple chemFET sensors. In another example, a single nucleic acid can be bound to a bead, the nucleic acid can be amplified on the bead, and individual beads can be transferred to individual reaction chambers on a chemFET array, each chamber having a chemFET sensor, and the nucleic acid can be sequenced.
[0068] In another embodiment, sequence information is obtained using Ion Torrent single molecule sequencing, where semiconductor technology is combined with simple sequencing chemistry to translate chemically coded information (A, C, G, T) into digital information (0, 1) directly on a semiconductor chip. In reality, when a nucleotide is incorporated into a DNA strand by a polymerase, a hydrogen ion is released as a by-product. Ion Torrent uses a high-density array of micromachined wells to perform this biochemical process in a massively parallel manner. Each well holds a different DNA molecule. Under the well is an ion-sensitive layer, and under that is an ion sensor. When a nucleotide, e.g., C, is added to a DNA template and then incorporated into the DNA strand, a hydrogen ion is released. The charge of this ion changes the pH of the solution, which can be detected by Ion Torrent's ion sensor. The sequencer, which is essentially the world's smallest solid-state pH meter, calls the bases and converts them directly from chemical information to digital information. The Ion personal Genome Machine (PGM™) sequencer then continuously pumps a huge number of nucleotides into the chip, one after the other. If the next nucleotide fed into the chip doesn't match, no voltage change is recorded and the base is not called. If there are two identical bases on the DNA strand, the voltage doubles and the chip records the two identical bases being called. Direct detection allows recording of nucleotide incorporation in seconds.
[0069] In another embodiment, the method of the present invention comprises obtaining sequence information of nucleic acid in a test sample, for example, cfDNA in a maternal test sample, by using hybridization sequencing. Hybridization sequencing comprises contacting a plurality of polynucleotide sequences with a plurality of polynucleotide probes, where each of the plurality of polynucleotide probes can be tethered to a substrate. The substrate can be a flat surface that includes an array of known nucleotide sequences. The pattern of hybridization to the array can be used to determine the polynucleotide sequences present in the sample. In another embodiment, each probe is tethered to a bead, for example, a magnetic bead, etc. Hybridization to the bead can be determined and used to identify a plurality of polynucleotide sequences in the sample.
[0070] In another embodiment, the method of the present invention includes obtaining sequence information of nucleic acids in a test sample, e.g., cfDNA in a maternal test sample, by massively parallel sequencing of millions of DNA fragments using Illumina's sequencing-by-synthesis and reversible terminator-based sequencing chemistry (e.g., as described in Bentley et al., Nature 6:53-59
[2009] , the entire contents of which are incorporated herein by reference).
[0071] Mapping of sequence tags is accomplished by comparing tag sequences to a reference sequence to determine the chromosomal origin of the sequenced nucleic acid molecule (in this case, cfDNA) and does not require specific gene sequence information. Small mismatches (0-2 mismatches per sequence tag) can be tolerated to account for minor polymorphisms that may exist between the reference genome and the genomes in the mixed sample.
[0072] A plurality of sequence tags are typically obtained per sample. The sequence tags include reads of 20-185 bp, e.g., 100 bp. In one embodiment, all sequence reads are mapped to all regions of the reference genome. In one embodiment, the tags mapped to all regions of the reference genome, e.g., all chromosomes, are counted to determine the CNV, i.e., over- or under-representation, of a sequence of interest, e.g., a chromosome or part thereof, in the mixed DNA sample. The method does not require differentiation between the two genomes.
[0073] The accuracy required for accurate determination of whether CNVs, such as aneuploidies, polyploidies, and deletions, are present in a sample is based on the variation between samples in the number of sequence tags that map to the reference genome during a sequencing run (inter-chromosomal variability), and the variation across different sequencing runs in the number of sequence tags that map to the reference genome (inter-sequencing variability). For example, variation may be particularly evident for tags that map to high GC or low GC reference sequences. Other variations may arise from using different protocols for nucleic acid extraction and purification, preparation of sequencing libraries, and use of different sequencing platforms. In the methods of the present invention, sequence dosage (chromosomal dosage or segment dosage) based on knowledge of normalizing sequences (normalizing chromosomal sequences or normalizing segment sequences) is used to essentially account for the resulting variability from inter-chromosomal (intra-run) and inter-sequencing (inter-run), as well as platform-dependent variability. The chromosomal dosage is based on knowledge of normalizing chromosomal sequences, which may consist of a single chromosome or two or more chromosomes selected from chromosomes 1-22, X, and Y. Alternatively, the normalizing chromosome sequence may consist of a single chromosome segment, or two or more segments of one chromosome or two or more chromosomes. The segment dosage is based on knowledge of the normalizing segment sequence, which may consist of a single segment of any one of chromosomes 1-22, X, and Y, or two or more segments of any two or more of chromosomes 1-22, X, and Y.
[0074] In one embodiment, the cfDNA is sequenced to determine the allelic sequence of the locus of interest in the sample. From this sequence, the heterozygous locus of interest is identified and the ratio of alleles is quantified. This ratio indicates the presence or absence of a chromosomal abnormality.
[0075] In some embodiments, the method also includes determining whether the fetus has a genetic disease from at least one sequence of interest of the fetus. In some embodiments, this includes determining whether the fetus is homozygous at a disease-causing allele in the sequence of interest when the mother is heterozygous at the same allele. In some embodiments, the disease-causing allele is a single nucleotide polymorphism (SNP) allele in the sequence of interest. In some embodiments, the disease-causing allele is a short tandem repeat (STR) allele in the sequence of interest.
[0076] In some embodiments, the at least one sequence of interest comprises the site of an allele associated with a disease. The at least one sequence of interest may comprise one or more of the following: single nucleotide polymorphisms, tandem repeats, microdeletions, insertions, and indels. In some embodiments, the method further comprises determining whether the fetus is homozygous or heterozygous for the disease-associated allele.
[0077] In some embodiments, the method further comprises counting sequence tags to determine copy number variations (CNVs) of chromosomal sequences and / or examining sequences of sequence tags to detect non-copy number variations of chromosomal sequences. Some embodiments further comprise determining the presence or absence of polymorphisms or chromosomal abnormalities in the fetus. In some embodiments, the method further comprises determining the presence or absence of SNPs, tandem repeats, insertions, deletions, indels, translocations, duplications, inversions, CNVs, partial aneuploidies, and / or complete aneuploidies in the fetus. Complete chromosomal aneuploidies can be duplications, multiple duplications, or complete chromosome losses. In some embodiments, the method further comprises localizing the partial aneuploidies. In some embodiments, the SNPs are selected from single SNPs and tandem SNPs. In some embodiments, the tandem repeats are selected from di-base repeats and tri-base repeats.
[0078] In some embodiments, the method further comprises determining a normalized chromosomal sequence value for the chromosomal sequence of interest and comparing the sequence value to upper and lower thresholds, where a sequence value that exceeds the upper and lower thresholds determines the presence or absence of a CNV, respectively.
[0079] In some embodiments, analyzing the plurality of sequence tags comprises determining the presence or absence of chromosomal abnormalities in chromosomes 1-22, X, and Y. Some embodiments further comprise determining whether the fetus is homozygous at an allele in the sequence of interest, or determining that the mother is heterozygous and the fetus is homozygous. Some embodiments further comprise determining the fetal fraction of the cfDNA.
[0080] Methods and systems for analyzing cfDNA are described in U.S. Patent No. 10,837,055, the entirety of which is incorporated herein by reference.
[0081] Another aspect of the present disclosure is a sequence validity test that can detect abnormalities in the size selection step or unacceptable sample degradation. An embodiment of this test is described in Example 4 and used to evaluate the suitability of the sample in Example 3. Although Example 3 describes the use of the related method with K2EDTA tubes, it is understood that this validity test can be used to evaluate the integrity of the sample when performing the size selection step, regardless of the type of collection tube or storage time or conditions. In a method that includes size selection of a DNA sample, including cfDNA, and sequencing the sample DNA after size selection, the method can include sequence validity testing of the sequence information obtained from the sequencing step. The validation includes generating a first normalized fragment size profile from sequencing information associated with the sample, comparing one or more values in the fragment size profile to one or more corresponding ranges of values obtained from a set of reference parameters comprising a plurality of validation fragment size profiles, and accepting the sequence information and its analysis results if one of the one or more values in the fragment size profile associated with the sample is within the corresponding range of values, or rejecting the sequence information if one of the one or more values in the fragment size profile associated with the sample is outside the corresponding range of values. The validation may also include the steps of calculating a first relative fragment size frequency value from sequencing information associated with the sample, comparing the first relative fragment size frequency value to a set of reference relative fragment size frequency values comprising a plurality of valid relative fragment size frequencies, and accepting the sequence information and its analysis results if the first fragment size frequency associated with the sample is within a range of values specified in the reference set, or rejecting the sequence information and its analysis results if the first fragment size frequency associated with the sample is outside the range of values specified in the reference set. [Example] EXAMPLES
[0082] This study is a comparative study evaluating the quality of cfDNA recovered from blood stored and transferred in K2EDTA tubes with gel (after low centrifugation) and non-centrifuged Streck tubes. Previous studies have shown that EDTA gel tubes are as good as Streck tubes for preparing cfDNA for NIPT without compromising the fetal fraction when blood is collected and processed on the same day (Giroux, 2021). The study will test the feasibility of keeping the centrifugation tubes at 4°C for 5 days and transporting them at low temperatures to a sequencing laboratory 4,800 km away.
[0083] Pregnant women were recruited exclusively in Vancouver after consenting to participate in this study. All recruited pregnant women were 10-14 weeks of gestational age. Details of the study population are shown in Table 1. [Table 1]
[0084] For each participant, 10 ml of peripheral blood samples were collected in two Streck tubes Cat#218997 (Streck, NB, USA) and 8 ml in molecular diagnostic K2EDTA gel blood tubes Cat#GR-455040 (Greiner, Austria). The EDTA gel tubes were centrifuged at 1800g for 10 min within 6 h of sample collection (recommended by the manufacturer) and then kept at 4°C until transported at low temperature to the testing facility in Quebec City once a week. Blood collected in Streck tubes was transported at room temperature to the same testing facility twice a week. A total of 61 samples were collected in Vancouver and transported over 4800 km to the testing facility in Quebec City.
[0085] Plasma was carefully collected from the Streck tubes using a 1 ml transfer pipette 2–5 days after blood collection. Plasma was filtered using a 5 ml syringe with Millipore 0.45 μM HPF-Millex-PVDF-Durapore (Merck, MA, USA). Plasma in the K2EDTA tubes was decanted into the cylinder of a 5 ml syringe and filtered using the same type of filter as previously reported (Giroux, 2021). For the EDTA gel tubes, it took 2–9 days from blood collection to filtration.
[0086] cfDNA was prepared from 2–5 ml of plasma using the QIAamp Circulating Nucleic Acid kit Cat# 55114 (Qiagen, Germany). 50 or 70 μl of DNA was collected and quantified fluorimetrically using the Qubit dsDNA HS Assay Kit Cat# Q32851 (Thermofisher, MA, USA). Libraries were prepared using a volume of 15–25 μl containing 2–15 ng of cfDNA using the KAPA Hyper Prep kit Cat# 07962363001 (Roche, Switzerland) according to the manufacturer's instructions with some modifications. At every step, the volume was reduced by half. Kapa Unique Dual-indexed Adapters Cat# 08861919702 (Roche, Switzerland) were used at a final concentration of 0.5 μM. Amplification was performed for 10 cycles and post-amplification cleanup was performed with 0.6× to 1× Kapa beads Cat #07983298001 (Roche, Switzerland).
[0087] In one set of samples, libraries were quantified fluorimetrically using the Qubit dsDNA HS Assay Kit as described above. Molar concentrations were calculated using an average size of 325 base pairs (bp). Pools of up to 17 equimolar libraries were prepared and subjected to paired-end sequencing for 150 cycles on an Illumina Next-Seq 550 instrument using a medium-output flow cell Cat#20024909 (Illumina, CA, USA).
[0088] For the second set of samples, libraries were size selected using a LightBench from Yourgene Health Canada Inc. (Manchester, UK). 25 μl of sample was spiked with loading buffer and 200bp / 300bp labeled markers, both run on a 3% precast gel. The software was programmed to extract DNA in the range of 229bp to 286bp. Recovered DNA was quantified on a Qubit and mass was converted to molar concentration using an average size of 275bp. Pools of up to 25 libraries were prepared and sequenced on a medium power flow cell as described above.
[0089] result Using standard procedures with Streck tubes, samples arrived at room temperature within 5 days. Seven pregnant women had a fetal fraction less than 4%, six of whom weighed more than 70 kg or had a BMI greater than 25.
[0090] For EDTA gel tubes, samples still arrived within 5 days, but had to be kept cold. Most of the samples arrived within 2 days by express delivery with the selected carrier. However, two shipments containing 14 samples were delayed by 4 and 5 days, and such samples were no longer kept cold. These 14 samples were damaged in shipping and had much lower fetal fractions than expected. These fetal fractions had an average decline of 4.6% compared to those obtained with Streck tubes. The fetal fractions for EDTA gel tubes (N=47) that arrived cold for 2 days or less were similar to those observed with Streck tubes, but were lower by an average of 1.1%. There was a trend between the time spent at 4°C (tubes were kept in a refrigerator until shipping) and the degree of decline in fetal fraction. We observed an increasing decline in fetal fraction over time prior to processing, from a decline of 0.75% at 2-3 days to a decline of 1.41% at 6 days or more.
[0091] All libraries prepared from DNA isolated from plasma in EDTA gel tubes were run on a 3% gel electrophoresis using a Yourgene Health Canada Inc. LightBench™ instrument. The software was set to harvest fragments from 229 base pairs to 286 base pairs, which constitute the adapters that add 136 bp to the cfDNA strand. The harvested material was sequenced and the fetal fraction was assessed by SeqFF as normal. An increase in fetal fraction was observed in all but five samples. The five that did not show an increase had an early high fetal fraction. The sample whose carriage was delayed for 5 days did not increase as much compared to the Streck tube results. Two samples had a lower fetal fraction of 0.4% and 3%, respectively, compared to the Streck tube, but the fetal fraction improved compared to the EDTA gel tube before size selection, and the other 12 samples had an increase of 0.1 to 1.5% (Figure 1, left sample). The remaining 47 samples all ranged from 0.8 to 10.2%, increasing by 5.4% on average, with the exception of three samples, which remained above 10%, but decreased compared to the Streck and even EDTA gels before size selection (Figure 1, samples marked with a star). Although the determination of the fetal fraction by SeqFF gives a rough estimate, this determination is nevertheless very useful in female pregnancies. The fetal fraction calculated from the decrease in ChrX fragments was also determined in male pregnancies, and an important increase was found for each sample analyzed (Figure 2A). Even samples with a low fetal fraction by SeqFF appeared to be much higher by this calculation method (blue stars in Figures 1 and 2A). An increase in reads from chromosome Y in male pregnancies (n=25) was also observed, so that the fetal fraction was not artificially increased by the tubes used and the selection on 3% agarose interfering with the counting by SeqFF (Appendix, Figure 2B).
[0092] Figure 3 shows the sizes of sequencing fragments calculated from mapping of paired-end sequencing. Grey indicates fragments before size selection, and blue indicates fragments after size selection. [Table 2] EXAMPLES
[0093] The purpose of this second study is to evaluate the effect of extended storage time before centrifugation of whole blood collected in K2EDTA tubes on the detection of fetal cfDNA. Separate tubes of single blood draws from each participant are incubated at a range of time points, centrifuged to obtain plasma, and then processed for DNA isolation and amplification, size selection, DNA libraries generated, and cfDNA sequencing. Fetal fraction percentages and chromosome ratios are used to determine the effect of storage time before centrifugation. This allows the measurement of the effect of variable blood storage times on fetal fraction percentages.
[0094] Samples are processed within the following time points: time point 1 - 0 to 8 hours, time point 2 - 72 hours, time point 3 - 120 hours, and time point 4 - 168 hours. During this time, a portion of the samples is stored at refrigerated temperature and a portion at room temperature. At the appropriate time points, plasma is separated from the whole blood by centrifugation. Within 6 hours of sample collection, all tubes are centrifuged at 1600g x 10 minutes. Plasma from the EDTA tubes is poured into a syringe cylinder and filtered through a 0.45 μm Millipore filter. The isolated plasma samples are stored at -80°C until at least 96 samples have been collected, at which time DNA extraction begins. Storage at this temperature preserves the samples until extraction. Extraction of cfDNA can be performed with the QIAamp Circulating Nucleic Acid kit Cat# 55114 (Qiagen, Germany). A specific size of cfDNA in the samples is selected using LightBench® from Yourgene Health Canada Inc. Size selection parameters were the same as in Example 1. NGS sequencing libraries were prepared and sequenced on an Illumina system. The fetal fraction was estimated using SeqFF. EXAMPLES
[0095] The objective of this third study is to evaluate the effect of extended storage time before centrifugation of whole blood collected in K2EDTA blood tubes (also known as K2EDTA vacutainers) on the detection of fetal cfDNA. Separate tubes of single blood draws from each participant will be incubated at time points 1-4, centrifuged twice to obtain plasma, and then processed for DNA isolation and amplification, DNA sequencing libraries will be generated, sequencing libraries will be pooled, size selection will be performed, and cfDNA will be sequenced. Fetal fraction percentages and chromosome ratios will be used to determine the effect of storage time before centrifugation. This will allow the measurement of the effect of variable blood storage times on fetal fraction percentages and the quality of NIPT testing with standard blood tubes with delayed plasma isolation.
[0096] Samples were stored at ambient temperature and processed within the following time points: time point 1 - 0 to 8 hours, time point 2 - 72 hours, time point 3 - 120 hours, and time point 4 - 168 hours. Ambient temperatures ranged from 20°C to 35°C. At the appropriate time points, plasma was separated from whole blood by two centrifugations. The plasma isolation centrifugation process included a first centrifugation step at 1600 rpm for 10 minutes and a second centrifugation step at 3000 rpm for 10 minutes. After the first centrifugation, the plasma from the K2EDTA tube was transferred to a second tube and subjected to a second centrifugation step. After the second step, the plasma was removed from the second tube, transferred to a third tube, and stored at -80°C until sufficient samples had been collected to start a run of less than 48 samples, at which time point DNA extraction was initiated. The isolated plasma samples were stored at -80°C until approximately 48 samples had been collected, at which time point DNA extraction process was initiated. Storage at this temperature preserves the samples until extraction.
[0097] Extraction of cfDNA was performed with the MagMAX™ Cell-Free DNA Isolation Kit. Despite the absence of Streck tubes, a lysis step was performed as part of this script and has no adverse effects on samples collected in EDTA tubes. Yourgene Health Canada Inc.'s QS250™ was used to select cfDNA of specific sizes in the samples. Size selection parameters were 106-146 bp (values do not include adapter length) and the recovery range was approximately 100-186 bp (mean fragment length value approximately 143-146 bp). NGS sequencing libraries were prepared and sequenced on an Illumina system. The fetal fraction was estimated using SeqFF.
[0098] Sample set details In this study, a matched set of 105 clinical samples was collected. Samples were collected with known gestational age, height, weight, and BMI. Separate vacutainers of single blood draws from each participant were incubated at time points 1-4, centrifuged to obtain plasma, and then run through the IONA® Nx workflow. Fetal fraction percentages and chromosomal ratios were used to determine the effect of storage time before centrifugation.
[0099] For nearly all participants, 4 5 mL blood samples (designated matched samples) were collected from each of the 105 patient volunteers (pregnant women with gestational age greater than 10+0 weeks) into BD Vacutainer™ plastic K2EDTA tubes. A 5 mL matched aliquot from each participant was subjected to centrifugation at each of time points 1-4 after incubation at ambient temperature. Table 3 below provides a summary of the samples collected. There are fewer than 105 sample deviating time points to obtain less than 20 mL from the participants. [Table 3]
[0100] The following samples, designated M029–M105, were mistakenly subjected to the standard single centrifugation in the IONA Nx protocol instead of two centrifugations: M029~M30 Time points 1~3 M031~M032 Time points 1~2 Such samples were excluded from this report because a single spin method was used for plasma preparation.
[0101] The remaining samples were subjected to two rounds of centrifugation and are summarized below. M029~M030 Time Point 4 M031~M032 Time point 3 and time point 4 M033~M105 Time point 1, Time point 2, Time point 3, and Time point 4 The samples listed above are those included in the analysis.
[0102] Exclusion and inclusion criteria for participants The inclusion and exclusion criteria were as follows: Inclusion criteria Pregnant women aged 18 or older Women who are able to understand and provide informed consent Singleton pregnancy: 10+0 weeks or more gestational age, no upper cutoff value for gestational age, i.e., until the due date Exclusion criteria Those with known maternal aneuploidy or maternal cancer Death of twins or one twin Those who are known to have a communicable disease, e.g., COVID-19 infection in circulation or in COVID-19 quarantine, HepB / Hepatitis B Those considered to be socially vulnerable Those known to have had an organ or stem cell transplant Those known to have received a blood transfusion in the past 12 months
[0103] result Sample Summary A summary of the QC adequacy status of all two rotations of patient samples at the four time points used in this study is shown in Table 4 below. [Table 4] JPEG2025502566000006.jpg247170 JPEG2025502566000007.jpg182170
[0104] Library Validity Library preparation (LP) validity rates were calculated for each run and each time point, where libraries with post-PCR quantification values of 2.5 ng / μL or greater were considered valid.
[0105] The initial library total validity and 1x retest validity rates for all samples at each time point in the study are shown in Table 5 below. [Table 5]
[0106] Summary statistics (means and standard deviations) and Welch ANOVA tests were performed on sample M056 (time point 1) from run 5, which was included in the size-selected sample pool because it nearly met the library validity criterion (2.49 ng / μL), in addition to the 275 validation libraries after 1× retesting (see 5). Data for all 276 samples are shown in FIG.
[0107] The Welch ANOVA test was performed instead of ANOVA because the unequal variance test demonstrated that the standard deviations of the data for each time point were not equal (see Figure 6 ); therefore, ANOVA was not an appropriate method of analysis as this test assumed equal variances.
[0108] Visually, the data demonstrates a change in high library concentration in the K2EDTA tubes prior to centrifugation with increasing sample storage time. Time point 4 demonstrated the highest library concentration mean value (11.31 ng / μL, FIG. 6), while time point 1 demonstrated the lowest concentration mean value (6.12 ng / μL, FIG. 6). Time point 2 presented the highest standard deviation (SD=2.90, FIG. 6), while time point 4 presented the lowest standard deviation (SD=1.48, FIG. 6). Time point 1 had a standard deviation of 2.48, time point 2 had a standard deviation of 2.90, and time point 3 had a standard deviation of 2.35, respectively, suggesting comparable variance across such time points. Such findings are consistent with the hypothesis that extended storage of samples in EDTA tubes results in increased DNA input to the library preparation reactions, and thus, later time points are expected to return higher library concentrations.
[0109] A Welch ANOVA p-value of <0.0001 in 6 indicates that there is statistically significant variance in the library quantification data between time points.
[0110] Note that any large DNA fragments (>1 kb) from the starting sample could theoretically be retained during the sample clean-up step of the library preparation workflow based on size alone and be independent of successful adapter ligation / amplification. Thus, such material could contribute to the fluorescent signal in the library fluorometric quantification assay, leading to increased library concentrations at later time points and longer sample storage times in this study (Figure 6).
[0111] The library concentration data of the 276 data points described above was processed using JMP's variability standard data quality assessment through main effects analysis to estimate the proportion of variability in the library concentration data that can be attributed to the test variables (EDTA storage time) and the runs performed. The variance components resulting from this test are shown in Table 6 below. [Table 6]
[0112] Variance component analysis showed that the study variable EDTA storage time (time point) contributed to 50% of the variance seen in the library concentration data (Table 6). Thus, 50% of the variance observed in the library concentration data comes from other workflow variables, 10% of this variance is due to run-to-run variability, and the remaining 40% is due to variables not investigated in detail in this analysis / study (Table 6). Possible variables may include, but are not limited to, the presence of multiple operators performing sample centrifugation and plasma isolation at different time points, laboratory / workflow variability across runs 1-9 (e.g., runs 1-3 by the service laboratory vs. runs 4-9 by R&D), and variability related to the patient themselves who collected the sample.
[0113] Sequencing Validity To assess the quality of the DNA samples and exclude any samples with problems in the sequencing process that would affect the quality of the sequencing results, each sequenced sample was analyzed for several parameters, mainly run control and sequencing quality control across runs, fragment counting, fragment size profile (compared to that expected for cfDNA size-selected sequencing libraries), GC profile fetal fraction, and consistency checks. All samples must have at least 2% fetal fraction. In addition, all samples with a risk of false negative or false positive results are assessed using our proprietary dynamic fetal fraction assessment, which adapts the level of fetal fraction required for a sample to its quality and the quality of the supporting sequencing data.
[0114] The overall sequencing validity rate for all samples tested (11 samples retested once) was 272 / 274 (99.27%). Eleven samples that initially failed sample validity were retested on subsequent runs, and nine of these samples returned valid results on retest.
[0115] The initial and 1x retest sequencing validity rates across each time point are presented in Table 7 below. [Table 7]
[0116] After 1x retesting, a total of 8 samples failed the sequencing validity check.
[0117] Six of these samples failed sequencing and were not retested. The 2x sequencing failure was due to failure of the fragment size profile at the 160 bp position (M080, time point 4 and M085, time point 4). The 4x sequencing failures were due to failure of the fragment size profile at the 180 bp position (M089, time point 4; M101, time point 3; M101, time point 4; and M102, time point 4). Two of these samples failed sequencing validity after 1x retest. The 1× sequencing failure was due to a failure of the fragment size profile at the 180 bp position (M068, time point 3). The 1× sequencing failure was due to failure of the fragment size profile at both the 120 bp and 180 bp positions (M076, time point 3).
[0118] The increased number of samples failing the sequencing QC checks (especially the "fragment size profile anomalies" check) at later time points could be due to increased DNA input as a result of extended sample storage in EDTA tubes, resulting in PCR plateaus and heteroduplex formation of library fragments in later PCR cycles. Sequencing of heteroduplex library fragments after sample enrichment by size selection (on the QS250) is therefore expected to increase the risk of samples failing the IONA software QC check, and indeed time point 4 is the most common with the highest frequency of samples failing the initial sequencing validity and the sequencing validity after 1x retest for "fragment size profile anomalies" (Table). Nevertheless, it was possible to produce a validity result by repeating the library preparation and sequencing of the samples.
[0119] An example of the change in fragment size frequency with increasing storage time before centrifugation can be observed in the two patients who returned validity results at each time point.
[0120] In Figures 7A and 7B, the representative shift in fragment profile to the right (i.e., larger DNA library insert size) demonstrated for patient samples M047 and M100 across all four time points is the expected result of heteroduplex formation, itself a result of increased DNA input into library preparation due to extended sample storage time in EDTA tubes. It should be noted that all aligned DNA fragments represented by the profiles in Figures 7A and 7B are within the expected size range of mononucleosomal cfDNA units and are not the expected DNA fragments arising from cell lysates as a result of extended storage time prior to centrifugation and plasma isolation.
[0121] Fetal Fraction Percentage The "IONA FF" method applies one of two algorithms to estimate the chromosome count depending on the value measured from the chromosome count analysis. IONA FF looks at the proportion of X and Y chromosomes. Then, - If there is a very high degree of confidence that the sample is from a male fetus, then the "E1" estimation algorithm is used, where a linear model is used to estimate FF from the X and Y chromosome data with the model calibrated during development. - Otherwise (i.e. if the sample is not confident that it is from a male fetus), use the "E2" algorithm for the sample. This method estimates from a model that associates the relative frequencies of long and short fragments with FF. This is also a linear model, but is calibrated using the E1 (X / Y based) estimate already calculated for each run. This improves the accuracy of the size-based estimate over the use of a static model by removing run-to-run variability in the measurement process.
[0122] The estimate is then output to the rest of the system along with an internal estimate of the measurement uncertainty that is also used by the plausibility checker.Another method of calculating the fetal fraction is by an algorithm known in the art as SeqFF.
[0123] Fetal fraction percentages were determined from the data of all 268 valid samples after analysis with IONA software. One-way ANOVA was also performed on the raw fetal fraction data (estimated enriched fetal fraction from raw sequencing data) to assess comparisons between each time point. ANOVA of all fetal fraction data is shown in Figure 5.
[0124] Visually, the data demonstrate a trend in change in low fetal fraction percentage with increasing storage time prior to centrifugation. All time points exhibited similar standard deviations when analyzing all 268 estimated fetal fractions, as supported by the unequal variance test described above, suggesting the presence of comparable variance across such time points.
[0125] The ANOVA p-value in FIG. 5 <0.0001 indicates that there is significant variance in fetal fraction percentages between time points.
[0126] Nevertheless, 267 of 268 samples had a fetal fraction above 4%, the common threshold applied for NIPT testing, and all exceeded the IONA Nx FF cutoff value.
[0127] Rare Autosomal Aneuploidy Analysis All 272 (100%) validity samples returned a validity RAA plugin result of "not detected," indicating that no aneuploidy was detected.
[0128] Match analysis A concordance analysis across time points for validity samples in this study was performed only if a validity baseline result (time point 1) was returned in the absence of a result for each of the other time points (i.e., results from library or sequencing failure), where the associated unsuitable result refers to a "not calculated" result for SCA / FSD determination. A concordance analysis was also performed for patient samples that returned validity results at all four time points.
[0129] RAA match analysis Concordance analysis for RAA status compared to baseline (Timepoint 1) is summarized below in Table 11. Samples were considered for which a valid RAA result at Timepoint 1 was available alongside at least one other timepoint. [Table 8]
[0130] Timepoint 2 - 60 / 60 (100%) patient samples returned a "not detectable" RAA result consistent with baseline. Timepoint 3 - 62 / 62 (100%) patient samples returned a "not detectable" RAA result consistent with baseline. Timepoint 4 - 59 / 59 (100%) patient samples returned an RAA result of "not detectable" consistent with baseline.
[0131] Concordance analysis for baseline comparable RAA status of 55 patient samples that returned validity results at all four time points is summarized in Table 12 below. [Table 9]
[0132] Timepoint 2 - 55 / 55 (100%) samples returned a "not detectable" RAA result consistent with baseline. Timepoint 3 - 55 / 55 (100%) samples returned a "not detectable" RAA result consistent with baseline. Timepoint 4 - 55 / 55 (100%) samples returned a "not detectable" RAA result consistent with baseline.
[0133] T13 / 18 / 21 match analysis Concordance analysis for T13 / 18 / 21 aneuploidy status compared to baseline (time point 1) is summarized in the table below. [Table 10]
[0134] Timepoint 2 - 60 / 60 (100%) patient samples returned a "low risk" result consistent with baseline. Timepoint 3 - 62 / 62 (100%) patient samples returned a "low risk" result consistent with baseline. Timepoint 4 - 59 / 59 (100%) patient samples returned a "low risk" result consistent with baseline.
[0135] Concordance analysis for baseline comparable T13 / 18 / 21 aneuploidy status for 55 patient samples that returned validity results at all four time points is summarized in the table below. [Table 11]
[0136] Timepoint 2 - 55 / 55 (100%) samples returned a "low risk" result consistent with baseline. Timepoint 3 - 55 / 55 (100%) samples returned a "low risk" result consistent with baseline. Timepoint 4 - 55 / 55 (100%) samples returned a "low risk" result consistent with baseline. EXAMPLES
[0137] This example provides an embodiment of a sequence validity check parameter that examines the fragment size profile and determines whether the profile is normal or abnormal. The purpose of this check is to reduce the risk of poor performance of the NIPT test due to reduced sample degradation and / or effectiveness of the fetal DNA size selection step (also called fetal DNA enrichment) by detecting cases of insufficient enrichment and not generating a test result for the corresponding sample analysis. This test can also be used to test for cancer cell sequences in cell-free DNA.
[0138] Fetal cell-free DNA is known to exist within a certain size range that is different from that of maternal cell-free DNA. Therefore, cases of ineffective enrichment and / or over-sample degradation can be detected by calculating a normalized fragment size profile from the sequencing data associated with the sample and comparing this parameter to a set of pre-calculated reference parameters that can distinguish effective enrichment from ineffective enrichment and / or over-sample degradation.
[0139] Size profile validity checks are designed, for example, to ensure that: 1. The enrichment of the fetal cell-free DNA component of maternal cell-free DNA is effective, and enrichment may be applied by size selection as described herein, ensuring preferential selection of fragments within a size range that are likely to be fetal; and / or 2. The sample sequencing data has a DNA fragment size profile consistent with the biological characteristics of cell-free DNA material extracted from plasma collected from pregnant women.
[0140] One or more size envelopes are defined at particular fragment size values of interest (in Example 3, the fragment size values of interest were 120 bp and 180 bp fragments). In particular, each envelope contains a fragment size profile of points that are considered to be valid. Thus, if an input fragment size profile satisfies the condition of inclusion by at least one of the envelopes so defined, the size profile validity check is considered to have passed for the corresponding sample.
[0141] Configuration parameters: To configure this embodiment of the validation parameters, the following parameters are required, which should be stored in an efficient hierarchical representation (e.g., using an XML-based format). - number of size envelopes to define - N env (integer, required). 108, - An array of size envelopes - SizeEnv[ ]. For each entry: - Number of size checkpoints - N chk (integer, required). - An array of size checkpoints - SizeChk[ ]. For each entry: - Define the fragment size - s chk (integer, required). - Validity relative frequency range - lower limit - z L (real number, any). Note: Value z UThis value may be omitted, provided that one substitutes in, in which case the lower limit is understood to be zero. - Validity relative frequency range - upper limit - z U (real number, any). Note: Value z L This value may be omitted, provided that one substitutes in, in which case the upper limit is understood to be 1.
[0142] Test input data: The input data is an array of measured fragment size values in the sequencing data.
[0143] Test Output: The method produces a single output: Size Profile Validity Status - V SP (Boolean, required).
[0144] Calculation of sample size profile: First, a relative frequency distribution is formed from the array of fragment size measurements generated for the sample by the analysis core pipeline, as generated by a method appropriate for the sequencing platform in use. This distribution is calculated as an array composed of the frequency of each size occurring, normalized according to the total number of fragments included in the distribution. Thus,
number
[0145] Any size values not shown (i.e.
number
[0146] Calculating the sample adequacy status: The relative frequency distribution Z calculated for the sample is compared in turn to each of the envelopes that define the set of constituent parameters. For each envelope, the comparison proceeds as follows: 1. At each size checkpoint, the corresponding size value s chk to find the value of the relative frequency of sizes in the distribution calculated for the sample: Z[s chk ]. 2. Compare the distribution value so obtained with the validity frequency range parameter: z L and z U . interval [z L ,z U ), i.e., Z[s chk ] <z L or Z[s chk ]≧z U If the distribution value does not fall within the interval [z L ,z U ), mark the comparison as passing.
[0147] The comparison proceeds to the next size checkpoint in this envelope until none remain.
[0148] The result of the envelope comparison is then calculated as the logical "and" of all constituent checkpoint results, i.e., the envelope comparison is considered to pass only if all constituent checkpoints pass.
[0149] If any single envelope comparison is found to pass, a validity result is generated for the sample validity by size profile validity check, and a value V SP = true. Only one envelope comparison is required to yield a validity result for the validity check as a whole.
[0150] After all envelope comparisons are completed, if none of the constituent check points are found to pass, a failure result for the sample validity by size profile validity check is generated, and a value V SP = False. This equates to an abnormal fragment size profile and the sample does not pass the sequence validity check and is rejected. A repeat blood draw is required.
[0151] Equivalent Although the present invention is described in detail with respect to this specific embodiment, it is understood that functionally equivalent variations exist within the scope of the present invention. Indeed, various modifications of the present invention in addition to those shown and described herein will become apparent to those skilled in the art from the foregoing specification and accompanying drawings. Such modifications are intended to fall within the scope of the appended claims. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the present invention described herein. Such equivalents are intended to be encompassed by the following claims.
[0152] All publications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication was specifically and individually indicated to be incorporated by reference in its entirety.
[0153] Many modifications and variations of this invention can be made without departing from the spirit and scope of the invention, as will be apparent to those skilled in the art. The specific embodiments described herein are offered by way of example only, and the present invention is limited only by the claims appended hereto, the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A method for analyzing a nucleic acid sequence obtained from a blood sample to provide sequence information, comprising the steps of: a. Analyzing the nucleic acid sequence of the size-selected cell-free DNA to detect sequence information, the size-selected cell-free DNA is obtained from a blood sample; the blood sample was collected in a fixative-free blood collection device; the blood sample was stored without a fixative at a temperature above -20°C and below 35°C prior to isolating the cell-free DNA from the plasma; the plasma being separated from blood cells present in the blood sample either before or after storage; wherein the cell-free DNA is isolated from the plasma at least 48 hours after collection of the blood sample from the subject; the size-selected cell-free DNA to be analyzed is obtained by separating the extracted cell-free DNA by size and isolating cell-free DNA of less than 300 bp from the extracted cell-free DNA;
2. A method for analyzing nucleic acid sequences obtained from a blood sample to provide sequence information, comprising the steps of: a. Analyzing the nucleic acid sequence of the size-selected cell-free DNA to detect sequence information, the size-selected cell-free DNA is obtained from a blood sample; the blood sample is collected in a blood collection device comprising a container having therein only one or more compositions, the one or more compositions consisting of one or more anticoagulants and a metal ion selected from potassium, lithium, and sodium; the blood sample was stored in the blood collection device at a temperature greater than -20°C and less than 35°C prior to isolating the cell-free DNA from the plasma; the plasma being separated from blood cells present in the blood sample either before or after storage; wherein the cell-free DNA is isolated from the plasma at least 48 hours after collection of the blood sample from the subject; the size-selected cell-free DNA to be analyzed is obtained by separating the extracted cell-free DNA by size and isolating cell-free DNA of less than 300 bp from the extracted cell-free DNA;
3. A method for analyzing nucleic acid sequences obtained from a blood sample to provide sequence information, comprising the steps of: a. Analyzing the nucleic acid sequence of the size-selected cell-free DNA to detect sequence information, the size-selected cell-free DNA is obtained from a blood sample; the blood sample is collected in a blood collection device comprising a container having only one or more compositions therein and a separation gel, the one or more compositions comprising one or more anticoagulants and a metal ion selected from potassium, lithium, and sodium; the blood sample was stored in the blood collection device at a temperature greater than -20°C and less than 35°C prior to isolating the cell-free DNA from the plasma; the plasma being separated from blood cells present in the blood sample either before or after storage; wherein the cell-free DNA is isolated from the plasma at least 48 hours after collection of the blood sample from the subject; the size-selected cell-free DNA to be analyzed is obtained by separating the extracted cell-free DNA by size and isolating cell-free DNA of less than 300 bp from the extracted cell-free DNA;
4. The method described in claim 1 or 3, wherein the cell-free DNA is isolated from plasma at least 3, 4, 5, 6, 7 or 8 days after blood collection.
5. A method according to any one of claims 1 to 3, wherein the plasma is separated from the blood after storing it for at least 2, 3, 4, 5, 6, 7 or 8 days after collection.
6. The method described in claim 5, wherein the plasma is separated from the blood after storage for at least 2, 3, 4, 5 or 6 days after collection, or after storage for 3 to 6 days, or after storage for 3, 4 or 5 days.
7. A method according to any one of claims 1 to 3, wherein the blood sample is stored in a blood collection device at a temperature above 0°C and below 35°C, or below 30°C, or below 25°C, or below 20°C before isolating cell-free DNA from the plasma.
8. A method according to any one of claims 1 to 3, wherein the blood sample is stored in a blood collection device at a temperature above 0°C and below 35°C or below 30°C or below 25°C or below 20°C before separating plasma from the blood.
9. A method according to any one of claims 1 to 3, wherein the blood sample is stored in a blood collection device at a temperature above 0°C or above 2°C and below 15°C or below 10°C or below 8°C or below 6°C before isolating cell-free DNA from the plasma.
10. A method according to any one of claims 1 to 3, wherein the blood sample is stored in a blood collection device at a temperature above 0°C or above 2°C and below 15°C or below 10°C or below 8°C or below 6°C before separating plasma from the blood.
11. A method according to any one of claims 1 to 3, wherein the blood sample is stored in a blood collection device at ambient temperature before isolating plasma from the blood.
12. A method according to any one of claims 1 to 3, wherein the nucleic acid sequence of the isolated cell-free DNA is analyzed to determine the presence of genetic abnormalities in the fetal DNA.
13. The method according to claim 1, wherein the sequence information comprises sequences of a plurality of sequence tags, The method further comprises analyzing the sequences of a plurality of the sequence tags to determine the presence of at least one sequence of interest in fetal DNA, wherein at least a portion of the plurality of sequence tags maps to at least one of the sequences of interest.
14. The sequence information includes genetic variations in the gene, The genes include BRCA1, BRCA2, MSH6, MSH2, MLH1, RET, PTEN, ATM, H-RAS, p53, ELAC2, CDH1, APC, AR, PMS2, MLH3, CYP1A1, GSTP1, GSTM1, AXIN2, CYP19, MET, NAT1, C DKN2A, NQ01, trc8, RAD51, PMS1, TGFBR2, VHL, MC4R, POMC, NROB2, UCP2, PCSK 1, PPARG, ADRB2, UCP3, glur1, cart, SORBS1, LEP, LEPR, SIM1, TNF, IL-6, IL- 1, IL-2, IL-3, IL1A, TAP2, THPO, THRB, NBS1, RBM15, LIE, MPL, RUNX1, Her-2, glucocorticoid receptor, estrogen receptor, thyroid receptor, p21, p27, K-RAS, N-RAS, retinoblastoma protein, Wiskott-Aldrich (WAS) gene, factor V Leiden, factor II (prothrombin), methylenetetrahydrofolate reductase, cystic fibrosis, LDL receptor, HDL receptor, superoxide dismutase gene, and SHOX gene; or The method of any one of claims 1 to 3, wherein the gene may be selected from genes involved in nitric oxide regulation, genes involved in cell cycle regulation, tumor suppressor genes, oncogenes, and genes associated with neurodegeneration.
15. The method described in claim 14, wherein the genetic mutation is a marker for a type of cancer.
16. The method of claim 1, wherein the separation step is performed using gel electrophoresis.
17. A method according to any one of claims 1 to 3, wherein the portion of cell-free DNA derived from the plasma sample is less than 185 bp, less than 180 bp, or less than 165 bp.
18. A method according to any one of claims 1 to 3, wherein the portion of cell-free DNA derived from the plasma sample is less than 155 bp or less than 150 bp.
19. The method described in claim 17, wherein the portion of the cell-free DNA derived from the plasma sample is greater than 50 bp, or greater than 80 bp, or greater than 90 bp, or greater than 100 bp.
20. The method of claim 1, wherein the fixative is a cross-linking agent, a metabolic inhibitor or a membrane stabilizer.
21. The method of claim 1, wherein the collection device contains metal ions selected from EDTA and potassium.
22. The method of claim 21, wherein the collection device further comprises a separation gel.
23. A method according to any one of claims 1 to 3, wherein the sample is obtained from a human source.
24. A method according to any one of claims 1 to 3, wherein the subject is a pregnant woman or a patient in need of cancer diagnosis.
25. A method according to any one of claims 1 to 3, wherein a sequence library containing the cell-free DNA is generated before or after isolating the cell-free DNA and separating it by size.
26. The method of claim 1, wherein the separation step is carried out using a bead-based binding matrix, column chromatography or a membrane filter.
27. A method according to any one of claims 1 to 3, wherein the step of analyzing the nucleic acid sequence of the isolated cell-free DNA to detect sequence information includes a step of sequencing the isolated cell-free DNA or sequencing at least a portion of the cell-free DNA to determine the nucleic acid sequence of one or more polynucleotides in the isolated cell-free DNA.
28. The sequence information comprises a nucleic acid sequence that has an insertion, deletion, chromosomal deletion, or methylation compared to a reference sequence; or The method according to any one of claims 1 to 3, wherein the nucleic acid sequence of the isolated cell-free DNA is analyzed for chromosomal numerical abnormalities or for the presence of XX chromosomes or XY chromosomes by calculating the chromosomal ratio.
29. A method according to any one of claims 1 to 3, further comprising the steps of generating a first normalized fragment size profile from sequencing information associated with a sample, comparing one or more values in the fragment size profile with one or more ranges of corresponding values obtained from a set of reference parameters comprising a plurality of plausible fragment size profiles, and accepting the sequence information and its analysis results if one of the one or more values in the fragment size profile associated with the sample is within the corresponding range of values, or rejecting the sequence information if one of the one or more values in the fragment size profile associated with the sample is outside the corresponding range of values.
30. The method described in claim 29, wherein the first normalized fragment size profile comprises or consists of fragment sizes less than 185 bp, or less than 180 bp, or less than 165 bp, or less than 155 bp, or less than 150 bp.
31. The method described in claim 30, wherein the first normalized fragment size profile comprises or consists of fragment sizes greater than 50 bp, or greater than 80 bp, or greater than 90 bp, or greater than 100 bp.
32. The method of claim 29, further comprising the step of sending, for a sample whose sequence information has been rejected, a request to re-collect the sample from the subject from whom the sample whose sequence information has been rejected was obtained.
33. A method according to any one of claims 1 to 3, further comprising the steps of calculating a first relative fragment size frequency value from sequencing information associated with a sample, comparing the first relative fragment size frequency value with a set of reference relative fragment size frequencies comprising a plurality of plausible relative fragment size frequencies, and accepting the sequence information and its analysis results if the first fragment size frequency associated with the sample is within a range of values specified in the reference set, or rejecting the sequence information and its analysis results if the first fragment size frequency associated with the sample is outside the range of values specified in the reference set.
34. The method described in claim 33, wherein the first relative fragment size frequency value is a fragment size of less than 185 bp, or less than 180 bp, or less than 165 bp, or less than 155 bp, or less than 150 bp.
35. The method described in claim 34, wherein the first relative fragment size frequency value is a fragment size of more than 50 bp, or more than 80 bp, or more than 90 bp, or more than 100 bp.
36. The method of claim 33, further comprising the step of sending, for a sample whose sequence information has been rejected, a request to re-collect the sample from the subject from whom the sample whose sequence information has been rejected was obtained.
37. The method of claim 33, further comprising the step of performing the calculation and comparison described in claim 33 for a second relative fragment size frequency value of a bp value different from the first fragment size frequency, wherein the sequence information and its analysis are accepted if the second fragment size frequency value associated with the sample is within a range of values specified in a reference set, or the sequence information is rejected if the second fragment size frequency value associated with the sample is outside the range of values specified in the reference set.