Identifying target molecules of interest in cell-free DNA
The method of using recognition elements to hybridize, circularize, and decode chromosomal sequences in cell-free DNA improves the accuracy of noninvasive prenatal screening, addressing false results and reducing the need for invasive procedures.
Patent Information
- Application Number
- PCT/US2025/027571
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-12
- Filing Date
- 2025-05-02
- Publication Date
- 2025-11-06
AI Technical Summary
Existing methods for noninvasive prenatal screening using cell-free DNA suffer from high false-positive, false-negative, and no-call results, necessitating invasive procedures like amniocentesis for fetal status determination.
A method involving recognition elements with complementary sequences to hybridize with non-polymorphic chromosomal sequences, followed by circularization, ligation, amplification, and decoding to accurately identify chromosomal anomalies and fetal fraction in cell-free DNA samples.
Enhances the accuracy of detecting chromosomal aneuploidies and fetal fraction, reducing the need for invasive procedures by providing robust and precise analysis of cell-free DNA.
Smart Images

Figure US2025027571_06112025_PF_FP_ABST
Abstract
Description
IDENTIFYING TARGET MOLECULES OF INTEREST IN CELL-FREE DNA
[0001] The present application claims the benefit of United States Provisional Application Serial No. 63 / 641,747, filed May 2, 2024, and United States Provisional Application Serial No. 63 / 670,514, filed July 12, 2024, each of which is incorporated herein by reference in its entirety.BACKGROUND
[0002] The ability to extract cell-free DNA from a blood sample has opened the door to multiple research opportunities, particularly in the field of cancer and noninvasive prenatal screening.
[0003] Noninvasive prenatal screening, or NIPT, utilizes cell-free DNA isolated from the blood of a pregnant female. NIPT offers tremendous potential as a screening method for fetal chromosomal aneuploidies, as approximately 3-13% of cell-free DNA isolated from the blood is of fetal origin. There are assays that utilize the fetal DNA from cell-free DNA in screening for the presence or the risk of fetal aneuploidies, such as by assaying for common trisomies such as Trisomy 13, 18 and 21. However, some disadvantages of existing methods include falsepositive, false-negative, or no-call results. Final fetal status determination using methods such as amniocentesis is still required for determining of the state of the fetus.
[0004] More accurate and robust tests and assays are needed to support further advances in this field of study.SUMMARY
[0005] Provided herein are methods, systems and compositions related to the analysis of cell- free DNA from a blood sample. In particular, the present disclosure provides accurate and robust methods for identifying chromosomal aneuploidies in a cell-free DNA blood sample taken from a pregnant female. Additionally, the present disclosure provides for the determination of a fetal fraction that makes up the cell-free DNA from the blood sample taken from a pregnant female. The present disclosure further provides for the detection of chromosomal aneuploidies, fetal fraction, or a combination thereof.
[0006] Aspects disclosed herein provide methods for identifying one or more chromosomal anomalies in cell-free DNA, comprising: a) providing a plurality of recognition elements, wherein a first recognition element of the plurality of recognition elements comprises a code, a 5 ’ region, and a 3 ’ region, wherein the 5 ’ region of the first recognition element comprises a sequence that is complementary to a first non-polymorphic chromosomal sequence of a first target chromosome in a cell-free DNA sample and wherein the 3 ’ region of the first recognitionelement comprises a sequence that is complementary to a second non -polymorphic chromosomal sequence of the first target chromosome in the cell -free DNA sample, and wherein the code of the first recognition element uniquely identifies a presence of the chromosomal sequences that are complementary to the sequence of the 5 ’ region and the sequence of the 3 ’ region of the first recognition element; b) hybridizing the first recognition element to the first non-polymorphic chromosomal sequence and the second non -polymorphic chromosomal sequence of the first target chromosome in the cell-free DNA sample to produce a hybridized first recognition element; c) ligating the 5’ region and the 3’ region of the hybridized first recognition element to a cell-free DNA in the cell-free DNA sample, thereby generating a circularized and ligated first recognition element; d) amplifying the circularized and ligated first recognition element to produce an amplified first recognition element; e) detecting the code of the amplified first recognition element; f) decoding the detected code of the amplified first recognition element; and g) concurrently performingb) through f) for a subsequent recognition element of the plurality of recognition elements, thereby identifying the one or more chromosomal anomalies in the cell-free DNA sample. In some embodiments, the first non- polymorphic chromosomal sequence is adjacently located to the second non-polymorphic chromosomal sequence in the first target chromosome. In some embodiments, the first non- polymorphic chromosomal sequence is not adjacently located to the second non-polymorphic chromosomal sequence in the first target chromosome, and wherein the method further comprises performing a gap -fill prior to ligating the 5’ region and the 3 ’ region of the hybridized first recognition element. In some embodiments, the code comprises one or more nucleic acid segments. In some embodiments, the one or more nucleic acid segments comprise two to ten nucleic acid segments. In some embodiments, the method further comprises hybridizing one or more detection polynucleotides to the amplified first recognition element and imaging the hybridizing the one or more detection polynucleotides to the amplified first recognition element, thereby detecting the two to ten nucleic acid segments of the code. In some embodiments, each of the one or more detection polynucleotides comprise a fluorescent moiety. In some embodiments, the decoding the detected code comprises soft decision decoding. In some embodiments, the cell-free DNA sample is extracted from a blood sample. In some embodiments, the blood sample is obtained or derived from a subject. In some embodiments, the subject is a pregnant female. In some embodiments, the method further comprises performing an exonuclease treatment after c). In some embodiments, the amplifying comprises polymerase chain reaction, primer extension, rolling circle amplification, or multiple strand displacement amplification. In some embodiments, the amplifying comprises generating concatemeric amplification products. In some embodiments, the one or more chromosomal anomaliescomprises a copy number variation of a chromosome, an insertion or a deletion in a chromosome, or a sub -chromosomal deletion. In some embodiments, the one or more chromosomal anomalies is in one or more of chromosome 1, chromosome 4, chromosome 5, chromosome 13, chromosome 16, chromosome 15, chromosome 18, chromosome 21, chromosome 22, chromosome X, or chromosome Y. In some embodiments, the method further comprises estimating a fetal fraction of the cell-free DNA sample. In some embodiments, the estimating the fetal fraction comprises identifying a plurality of single nucleotide polymorphisms in fetal DNA and in maternal DNA. In some embodiments, the plurality of single nucleotide polymorphisms comprises 50 single nucleotide polymorphisms to 300 single nucleotide polymorphisms. In some embodiments, the plurality of single nucleotide polymorphisms is located across a plurality of chromosomes in the cell -free DNA sample. In some embodiments, the plurality of chromosomes comprises chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome 23, or chromosome 24 in the cell-free DNA sample. In some embodiments, the method further comprises identifying the plurality of single nucleotide polymorphisms and the one or more chromosomal anomalies in the same reaction. In some embodiments, the one or more chromosomal anomalies are indicative of a fetal disease selected from the group consisting of DiGeorge syndrome, Prader-Willi syndrome, Angelman syndrome, Cri-du-chat syndrome, Down syndrome, Patau syndrome, Edward syndrome Trisomy 21, Turner’s syndrome, Triploidy, and Kleinfelter’s syndrome. In some embodiments, identifying the one or more chromosomal anomalies comprises utilizing decoded codes of a plurality of amplified recognition elements that align to the one or more chromosomal sequences in the cell-free DNA sample to determine a copy number of a first chromosome and a copy number of a second chromosome, and determining a ratio between the determined copy number of the first chromosome and the determined copy number of the second chromosome, and wherein the ratio identifies a presence or an absence of the one or more chromosomal anomalies present in the cell-free DNA sample.
[0007] Aspects disclosed herein provide methods for identifying one or more chromosomal anomalies in a cell-free DNA sample, comprising: a) hybridizing a plurality of recognition elements to chromosomal sequences present in the cell -free DNA sample to generate hybridized recognition elements, wherein each recognition element of the plurality of recognition elements comprises a code, a 5’ region, and a 3 ’ region, wherein each code of each recognition element ofthe plurality of recognition elements identifies the chromosomal sequences present in the cell- free DNA sample that are complementary to the corresponding recognition element of the plurality of recognition elements, wherein the presence of the chromosomal sequences indicates a presence of one or more chromosomal anomalies in the cell -free DNA sample; b) ligating the 5 ’ region and the 3 ’ region of the hybridized recognition elements, thereby generating circularized and ligated recognition elements; c) amplifying the circularized and ligated recognition elements to generate amplified recognition elements; d) detecting the codes of the amplified recognition elements; and e) decoding the codes that are detected in d), thereby identifying the one or more chromosomal anomalies in the cell -free DNA sample. In some embodiments, the cell-free DNA sample is extracted from blood. In some embodiments, the method further comprises performing an exonuclease treatment after b). In some embodiments, the 5’ region is adjacently located to the 3’ region of the hybridized recognition elements. In some embodiments, the 5’ region is not adjacently located to the 3’ region of the hybridized recognition elements, and wherein the method further comprises performing a gap -fill prior to ligating the 5’ region and the 3’ region of the hybridized recognition elements. In some embodiments, a code is the same for a subset of recognition elements of the plurality of recognition elements from a region of a chromosome. In some embodiments, the subset of recognition elements comprises 5 to 50 recognition elements. In some embodiments, the code comprises one or more nucleic acid segments. In some embodiments, the one or more nucleic acid segments comprises two to ten nucleic acid segments. In some embodiments, the method further comprises hybridizing one or more detection polynucleotides to the two to ten nucleic acid segments and imaging the hybridizing the one or more detection polynucleotides to the two to ten nucleic acid segments, thereby detecting the two to ten nucleic acid segments. In some embodiments, each detection polynucleotide of the one or more detection polynucleotides comprises a fluorescent moiety. In some embodiments, decoding the codes comprises soft decision decoding. In some embodiments, the amplifying comprises polymerase chain reaction, primer extension, rolling circle amplification, or multiple strand displacement amplification. In some embodiments, the amplifying comprises generating concatemeric amplification products. In some embodiments, the one or more chromosomal anomalies comprises a copy number variation of a chromosome, an insertion or a deletion in a chromosome, or a sub -chromosomal deletion. In some embodiments, the one or more chromosomal anomalies is in one or more of chromosome 1, chromosome 4, chromosome 5, chromosome 13, chromosome 16, chromosome 15, chromosome 18, chromosome 21, chromosome 22, chromosome X, or chromosome Y. In some embodiments, the method further comprises estimating a fetal fraction of the cell-free DNA sample. In some embodiments, the estimating the fetal fraction comprises identifying aplurality of single nucleotide polymorphisms in fetal DNA and in maternal DNA using the plurality of recognition elements. In some embodiments, the plurality of single nucleotide polymorphisms comprises 50 single nucleotide polymorphisms to 300 single nucleotide polymorphisms. In some embodiments, the plurality of single nucleotide polymorphisms are located across a plurality of chromosomes in the cell -free DNA sample. In some embodiments, the plurality of chromosomes comprises chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome ?, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome 23, or chromosome 24 in the cell-free DNA sample. In some embodiments, the method further comprises identifying the plurality of single nucleotide polymorphisms and the one or more chromosomal anomalies in the same reaction. In some embodiments, the one or more chromosomal anomalies are indicative of a fetal disease selected from the group consisting of DiGeorge syndrome, Prader-Willi syndrome, Angelman syndrome, Cri-du-chat syndrome, Down syndrome, Trisomy 13, Trisomy 18, Trisomy 21, Turner’s syndrome, Triploidy, and Kleinfelter’s syndrome.
[0008] Aspects disclosed herein provide compositions comprising: a) a first plurality of recognition elements, wherein each recognition element of the first plurality of recognition elements is hybridized to non -polymorphic target sequences on one of chromosome 1, chromosome4, chromosome 5, chromosome 13, chromosome 16, chromosome 15, chromosome 18, chromosome 21, chromosome 22, chromosome X, or chromosome Y, and b) a second plurality of recognition elements, wherein each recognition element of the second plurality of recognition elements is hybridized to a target sequence comprising a single nucleotide polymorphism sequence on one of chromosome 1 through chromosome 24. In some embodiments, the composition further comprises one or more of a ligase or an exonuclease. In some embodiments, the composition further comprises a first plurality of circularized and ligated recognition elements and a second plurality of circularized and ligated recognition elements. In some embodiments, the composition further comprises a polymerase enzyme and a plurality of concatenated amplification products.
[0009] Aspects disclosed herein provide kits, comprising a) a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements is capable of hybridizing to target chromosomal sequences in a cell -free DNA sample; b) one or more of a ligase and a DNA polymerase; c) a plurality of fluorescently labeled polynucleotides; and d) instructions for practicing any one of the methods herein. In some embodiments, the kit further comprises an extraction medium for extracting cell-free DNA from the blood sample. In someembodiments, the ligase comprises a thermostable ligase. In some embodiments, the kit further comprises a thermostable DNA polymerase comprising strand displacement activity. In some embodiments, the kit comprises one or more exonucleases.INCORPORATION BY REFERENCE
[0010] All publications, patents and patent applications mentioned in this specification are herein incorporated by reference to the extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS[OH] A better understanding of the features and advantages of the inventive concepts will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the inventive concepts are utilized, and the accompanying figures of which include:
[0012] Fig. 1 is an example of a recognition element used in methods of the present disclosure.
[0013] Fig. 2 is an example of a workflow for generating concatemeric amplification products used in methods of the present disclosure.
[0014] Fig. 3 A shows an example of a detection polynucleotide.
[0015] Fig. 3B shows an example of a detection polynucleotide hybridized to an example of a portion of a concatemeric amplification product.
[0016] Fig. 4 is a schematic diagram of an example of a soft decision decoding workflow for decoding a code of a recognition element.
[0017] Fig. 5 is an example instrument that comprises a computer system for use with the methods described herein.
[0018] Fig. 6 is an example of an application provision system for use with the methods described herein.
[0019] Fig. 7 is an example of an application provision system for use with the methods described herein.
[0020] Fig. 8 shows an example of a scenario where three chromosomes of different lengths are queried with recognition elements.
[0021] Fig. 9 shows an example of a scenario where an example of a chromosome is queried with recognition elements where the code may be shared among a number of recognitionelements but the 5’ and 3’ regions of each recognition element target different conserved sequences on the example of the chromosome.
[0022] Fig. 10 shows an example of a scenario for detecting single nucleotide polymorphisms (SNPs) present in both maternal and fetal DNA from a cell-free DNA sample for estimating a fetal fraction of the cell-free DNA sample.
[0023] Fig. 11 A shows a graph illustrating an example of data using spike-in DNA samples for proof of concept for an assay in determining correlation between fetal DNA concentration and Y chromosome ratio.
[0024] Fig. 1 IB shows a graph illustrating an example of a correlation between using data and fetal fraction from spike -in DNA SNP detection.
[0025] Fig. 12 shows examples of graphs illustrating the identification of different chromosomal anomalies in test samples using a method for copy-number variation (CNV) calling.
[0026] Fig. 13 A shows an example of a SNP detection prior to recognition element correction for performance in estimation of fetal fraction.
[0027] Fig. 13B shows an example of a SNP detection after recognition element correction for estimation of fetal fraction.
[0028] Fig. 14 provides a non-limiting example of genotypes used in a bioinformatics training module.
[0029] Fig. 15 shows an example of an analysis pipeline for determining error profiles for a recognition element, which can be used for data correction.
[0030] Fig. 16 shows an example of an analysis pipeline for estimating a fetal fraction of a cell- free DNA sample using an Expectation Maximization algorithm.
[0031] Fig. 17 shows an example of an analysis pipeline for determining chromosomal aneuploidy using intra-sample and inter-sample error mitigation strategies.
[0032] Fig. 18 shows non-limiting examples of results comparing detection of trisomies and euploid matched samples.
[0033] Fig. 19 shows a non-limiting example of the correlation between fetal fraction estimated by SNP detection and fetal fraction estimated from chromosome Y.
[0034] Fig. 20 shows a non-limiting example of trisomy 21 detection from cfDNA extracted from plasma.DETAILED DESCRIPTION
[0035] The present disclosure provides methods, systems and compositions for identifying targets of interest in cell-free DNA. Further disclosed herein are assays for identifying chromosomal anomalies in maternal and fetal DNA.
[0036] The methods and compositions as disclosed herein for determining a nucleic acid sequence may comprise an assay. In some embodiments, the assay is a solution -based assay. In some embodiments, the assay is a surface-bound assay. In some embodiments, the assay is a hybrid assay that includes a surface-bound component and a solution -based component. In some embodiments, the assay is performed in one or more tubes or in a plate-based format, such as a multi-well plate. A multi-well plate may include, for example, a 6, 8, 12, 24, 48, 96, 384, or 1536 well plate. In some embodiments, a multi -well plate may include, for example, an array of nanowells. In some embodiments, the assay may be performed on a microfluidics device. In some embodiments, the assay may be performed partially in one or more tubes and partially in a multi -well plate.
[0037] In some embodiments, a recognition element is used in the assay. A recognition element used in the assay may include sequences that are complementary to a target sequence of interest, wherein the complementary sequence can hybridize to a target sequence of interest, a code sequence that can be used to identify the target sequence of interest that has hybridized to its complement on the recognition element, and one or more functional sequences such as sequencing primer binding sites, one or more amplification primer binding sites, unique molecular identifier sequences (UMIs), sample indexes, or combinations thereof. In some embodiments, an amplification primer binding site may be adjacent to the code in a recognition element. The amplification primer binding site(s) may, in some cases, be universal primer sequence(s) that are common to all recognition elements in a set of recognition elements.Amplification primer binding site sequences may also be a code sequence or a portion thereof. A code sequence can be a combination of a number of subsequences (e.g., nucleic acid segments), wherein the combination can identify a target sequence of interest that has hybridized to a recognition element. In some embodiments, an amplification primer can be a nucleic acid segment or a portion thereof. Unique identifier sequences (UMIs) and sample indexes, which may find utility in next generation sequencing (NGS) reactions for counting, error correction and sample identification purposes, may also be part of the code such as one or more nucleic acid segments.
[0038] Once a recognition element has recognized and hybridized to its target molecule of interest, the recognition element may be circularized and ligated to generate a circular, ligated recognition element. The circular and ligated recognition element can be amplified in anticipation of a decoding event to identify the code associated with the original target molecule of interest that hybridized to the recognition element. The amplification may be by any method of amplification, including for example, nucleic acid extension, polymerase chain reaction (PCR), isothermal amplification, rolling circle amplification (RCA), and / or ultrarapidamplification, and the like. Surface based amplification may be performed using PCR with surface-anchored primers (e.g., Illumina bridge amplification technology), or recombinase polymerase amplification (RPA) (e.g., ExAmp technology).
[0039] In one embodiment, the amplification operation may comprise a rolling circle amplification (RCA) reaction to generate concatemeric amplification products.
[0040] In one embodiment, a recognition element may include a sequence which may prevent RCA of the recognition element while allowing for linear double -stranded PCR products. The non-extendable sequence may, for example, be located between a pair of amplification primer binding site sequences present on the recognition element. In one embodiment, a recognition element may include a restriction enzyme site that may be cleaved to yield a linear DNA molecule.
[0041] In some embodiments, a concatemeric amplification product may be sequenced to determine the nucleotide sequence of the code associated with the target molecule of interest. Any sequencing technology may be used to sequence the product. Non-limiting examples of sequencing technologies that may be used include sequencing by synthesis (SBS), avidity sequencing, sequencing by hybridization, sequencing by ligation, nanopore sequencing, and the like.
[0042] In some embodiments, a sequencing library may be generated from a set of recognition elements or complements or amplicons thereof. The sequencing library may be sequenced to determine the code of the recognition element associated with a target molecule of interest. The code sequence may be used as a digital count of the target molecule specific decoding event. In one embodiment, a sequencing library may be generated from a circularized recognition element. In another embodiment, a sequencing library may be generated from a concatemeric amplification product of a recognition element. In one embodiment, a concatemeric amplification product or a portion thereof that includes at least the code may be directly sequenced to determine the code associated with the target molecule of interest.
[0043] Fig. 2 shows an example of an encoded assay for use with the detection polynucleotides disclosed herein. A linear recognition element 210 may comprise a 5' end 220a which may be complementary to a portion of a target nucleic acid of interest 222, a 3 ' end 220b which may be complementary to another portion of a target nucleic acid of interest 222 from a sample, a code 216, and additional functional sequences 212, 214, 218 such as amplification primer binding sites, capture sequencing, cleavage sites, sequencing primer binding sites, unique molecule identifiers (UMIs), and the like. A target nucleic acid of interest 222 which may be complementary to the 5' 220a and 3' 220b ends of the linear recognition element 210 may hybridize to the linear recognition element, thereby bringing the ends in proximity for ligating togenerate a circular and ligated recognition element 225. The circular and ligated recognition element 225 may be subjected to extension amplification using one of the functional sequences 212, 214, 218, or the code 216 or a portion thereof, as a primer binding site. The result may be a concatemeric amplification product 230 which can be decoded using the detection polynucleotides disclosed herein for identifying and determining the presence of the target nucleic acid of interest from a sample.
[0044] Additional examples of encoded assays can be found in WO2022 / 109496 A2, which is incorporated herein by reference in its entirety.
[0045] The methods and compositions described herein may include providing recognition elements to an encoded assay for identifying the presence of a target molecule of interest from a sample. In some embodiments, a plurality of recognition elements is provided. In some embodiments, each recognition element in the plurality of recognition elements comprises one or more target recognition regions. The target recognition regions of the recognition elements may comprise one or more nucleic acid sequence(s) configured to hybridize to a target nucleic acid molecule. In some embodiments, the one or more nucleic acid sequences may hybridize to one or more target nucleic acid sequences of the target nucleic acid molecule. In some embodiments, the target recognition region may be configured to hybridize to one or more regions of the target nucleic acid molecule flanking a target of interest (e.g., SNP, insertiondeletion mutations (indel), and so on). In some embodiments, the target recognition region may be configured to hybridize to a variant nucleic acid of interest (e.g., the target recognition region base pairs with the SNP when the target of interest is a SNP).
[0046] As shown in Fig. 1, an example of a recognition element used in encoded assays described herein may comprise two target recognition regions, one at the 5’ prime end and another at the 3’ end. A recognition element may further comprise a code. As shown in Fig. 1, the example code may comprise four nucleic acid segments. However, the number of nucleic acid segments is not limiting. The number of segments may depend on the complexity of the assay (e.g., number of target molecules of interest identified from a sample). In some embodiments, a code may comprisemore than or equal to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 segments. In some embodiments, a code may comprise less than or equal to 30, 29, 28, 27, 26, 25, 24, 24, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 segments. In some embodiments, a target nucleic acid molecule of interest itself comprises the code, or an additional code such as a barcode that identifies another target molecule of interest such as a protein. In some embodiments, the code may be detected as a proxy for the target molecule, be it a nucleic acid, a protein, or both.
[0047] In some embodiments, the structure of the recognition element may vary. In some embodiments, the structure of the recognition element may configure into a specific structure when hybridized to a target nucleic acid. Non-limiting examples of a recognition element configuration may include a padlock probe, a molecular inversion probe, a hairpin oligonucleotide, a single-stranded oligonucleotide, a double-stranded oligonucleotide, a combination thereof, and the like. In some embodiments, the recognition element is linear prior to hybridization to its complementary target nucleic acid of interest. In some embodiments, the linear recognition element is circularized once hybridized to the respective target nucleic acid of interest and ligated thereafter. In some embodiments, the recognition element is circular prior to hybridization to the target nucleic acid of interest. In some embodiments, the target nucleic acid of interest may serve as a primer for an extension reaction, for example to initiate rolling circular amplification of the recognition element.
[0048] In some embodiments, the recognition element is configured to be a padlock probe once the recognition element is hybridized to the target nucleic acid of interest. Padlock probes may be referred to as linear oligonucleotides whose ends may be complementary to adjacent target sequences, or to non-adjacent target sequences thereby leaving a gap between the ends of the hybridized recognition element. Upon hybridization to a target nucleic acid, the two ends (e.g. , the 5’ end and the 3’ end) of the recognition element may be adjacently located, generating a padlock probe configuration for subsequent ligation. Alternatively, the two ends of the recognition element may be brought in proximity to, but not directly adjacent to, each other upon hybridization to a target nucleic acid. In this example, a gap may be left between the 5' and 3' hybridized ends of the recognition element which can be filled in several ways, for example by extension of the 5' end until it is adjacent to the 3' end, or by hybridizing a third oligonucleotide that fills the gap. In any scenario, the recognition element ends may be ligated together if hybridization, or hybridization and gap fill, occurs thereby generating circular and ligated recognition elements that are indicative of the hybridization event.
[0049] In some embodiments, a recognition element may further comprise one or more functional sequences. Functional sequences may include, but are not limited to, primer binding sites, cleavage sites, unique molecular identifiers (UMIs), capture sequences, or combinations thereof. In some embodiments, a primer binding site and / or a cleavage site may be universal in nature, such that a plurality of recognition elements may share the same sequence(s). A unique molecular identifier maybe included in a recognition element to identify a source of material, for example, for error correction.
[0050] The methods described herein may relate to the use of a code for identifying a target nucleic acid molecule of interest from a sample that hybridized to a recognition element toinitiate a ligation event. A code in a recognition element may be used to associate the recognition element 5' and 3' end regions with a target nucleic acid of interest, thereby determining the presence of a target nucleic acid molecule of interest from a sample without having to directly assay the target molecule itself. As such, a code in a recognition element may uniquely identify the presence of a target molecule from a sample. Using codes, any number of recognition elements can be multiplexed in one encoded assay as each code is unique and correlates to the presence of one target molecule. In some embodiments, the code is selected from a set of codes wherein the set of codes make up a “code space”. In some embodiments, the code comprises a plurality of nucleic acid segments, where each nucleic acid segment corresponds to one or more computational symbols, or colors, that are used in a decoding process. Fig- 1 shows an example of a recognition element where four nucleic acid segments make up the code of the recognition element, wherein each of the nucleic acid segments can be decoded using detection polynucleotides as disclosed herein and the combination of the decoded nucleic acid segments thereby builds the full code that is unique to the target nucleic acid of interest from a sample. The codes may be detected as proxies, thereby serving as an indirect analysis of the presence of a target molecule from a sample as the code correlates with the presence of the target molecule that hybridized to the recognition element allowing ligation, amplification and decoding. If there is no hybridization of a target sequence of interest to its complementary sequences of a recognition element, there may be expected to be no ligation (e.g., as the 5’ and the 3’ ends of the recognition element are not expected to be adjacent), no amplification and subsequently no amplification product to decode. As such, if there is no amplification product to decode, that may be an indication that the target molecule of interest may be potentially absent from the sample, or at such a low incidence that hybridization resulted in too few amplification products to cross a threshold for detection by decoding.
[0051] In some embodiments, each code from a set of codes may be from a predetermined set of codes. In some embodiments, each code from the set of codes may be selected to ensure that the selected code differs from other codes in the set of codes. As such, in some embodiments, several selection criteria may be implemented to generate a set of codes, wherein each code of a set of codes comprises from one to more than one nucleic acid segment. In some embodiments, selection of the codes, or nucleic acid segments that make up a code, may incorporate a Hamming distance criterion.
[0052] In some embodiments, to generate a code selected from a set of codes for use in a recognition element, a Hamming distance (HD) selection criterion may be implemented between any two codes of the set of codes, and also between any two nucleic acid segments that may be used in a code. A Hamming distance between two codes in a set of codes may refer to thenumber of symbols, or nucleotides, that differ between the two codes in the set of codes. The Hamming distance may measure the number of changes that may need to be made to a first code sequence to change the string of symbols, in this case nucleotides, to the second code. As such, a Hamming distance criterion used to select a code may not be greater than the length of the code. For example, if the length of a code is measured by the number of cycles or flows of decoding runs or queries and that number being eight cycles, and if each cycle corresponds to one symbol or color, therefore eight symbols or colors, then the maximum Hamming distance is eight. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance may be from about 2-10. In some embodiments, the Hamming distance may be from about 2-7. In some embodiments, the Hamming distance may be from about 3-5. The Hamming distance may increase as the number of codes that can be used decreases, as one purpose of the code is to impart a way to uniquely identify one target molecule from another target molecule.
[0053] The code may have a certain length in nucleotides. In some embodiments, the code has a length of greater than or equal to about three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 50, 60, 70, 80, 90, 100, 150, or 200 contiguous nucleotides. In some embodiments, the code has a length of fewer than or equal to about 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 contiguous nucleotides. In some embodiments, the length of the code is about 5 to 200, about 10 to 150, about 15 to 100, about 20 to 90, or about 30 to 80 contiguous nucleotides. The length of the code may be further divided into a number of discrete nucleic acid segments, as shown in Fig. 1.
[0054] In some embodiments, each code from a set of codes is generated using a 4 -ary nucleotide alphabet of A, C, G, and T. In some embodiments, each code of a set of codes is generated using a 3 -ary nucleotide alphabet of a set of three of A, C, G, and T. In some embodiments, the codes canbe generated from arbitrary symbols, colors, or numbers (e.g., 1, 2, 3, 4, etc.), corresponding to fluorophores associated with a unique string of nucleotides. The numbers (e.g., 1, 2, 3, 4, etc.) can serve as a means to numerate the colors or combinations of colors that are utilized to query the codes for decoding.
[0055] A code of a recognition element can comprise one or more nucleic acid segments. For example, as shown in Fig. 1, the example of the code comprises four nucleic acid segments.
[0056] In some embodiments, the recognition elements provided herein may comprise a code comprising one or more nucleic acid segments. The one or more nucleic acid segments, or thecomplements thereof, within the code may be used as a proxy for detection of the target molecule recognized by the recognition element.
[0057] The number of nucleic acid segments present in a code of a recognition element may be considered in the design of the recognition element. The number of segments in a code may help to determine the nucleotide length of the recognition element. For example, a recognition element that includes a code comprising five segments may comprise a greater nucleotide length than a recognition element that includes a code of two segments. A recognition element with a larger nucleotide length may run up against synthesis limits and may be at a greater risk of synthesis errors. Alternatively, a recognition element with a smaller nucleotide length may avoid synthesis limits and risks in synthesis errors. A recognition element with a larger nucleotide length may include less space for other portions of the recognition element, such as the target recognition regions, functional sequences, universal sequences, etc.
[0058] In some embodiments, the code comprises about 2 to 10 nucleic acid segments. In some embodiments, the code comprises about 2 to 8 nucleic acid segments. In some embodiments, the code comprises about 3 to 5 nucleic acid segments. In some embodiments, the code comprises at least 1 nucleic acid segments, at least 2 nucleic acid segments, at least 3 nucleic acid segments, at least 4 nucleic acid segments, at least 5 nucleic acid segments, at least 6 nucleic acid segments, at least 7 nucleic acid segments, at least 8 nucleic acid segments, at least 9 nucleic acid segments, at least 10 nucleic acid segments, at least 11 nucleic acid segments, at least 12 nucleic acid segments, at least 13 nucleic acid segments, at least 14 nucleic acid segments, or at least 15 nucleic acid segments.
[0059] In some embodiments, each nucleic acid segment may comprise a length in nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 10 to 30 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 10 to 25 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of about 15 to 20 nucleotides. In some embodiments, each nucleic acid segment may comprise a length of 2 or more nucleotides, 3 or more nucleotides, 4 or more nucleotides, 5 or more nucleotides, 6 or more nucleotides, 7 or more nucleotides, 8 or more nucleotides, 9 or more nucleotides, 10 or more nucleotides, 11 or more nucleotides, 12 or more nucleotides, 13 or more nucleotides, 14 or more nucleotides, 15 or more nucleotides, 16 or more nucleotides, 17 or more nucleotides, 18 or more nucleotides, 19 or more nucleotides, 20 or more nucleotides, 21 or more nucleotides, 22 or more nucleotides, 23 or more nucleotides, 24 or more nucleotides, or 25 or more nucleotides. In some embodiments, each of the nucleic acid segments in a code are of the same length. In some embodiments, each of the nucleic acid segments in a code are not the same length.
[0060] A Hamming distance selection criterion may be implemented between any two nucleic acid segments of a code. A Hamming distance between two nucleic acid segments in a code may refer to the number of symbols that differ between the nucleic acid segments. The Hamming distance may measure the number of changes that may need to be made to a first nucleic acid segment to change the string of symbols, or nucleotides, to a second nucleic acid segment. In some embodiments, the Hamming distance may be a minimum Hamming distance. In some embodiments, the Hamming distance may be a maximum Hamming distance. In some embodiments, a minimum Hamming distance maybe from about 2-20, about 3-19, about 4-18, about 5-17, about 6-16, about 7-15, about 8-14, about 9-13, or about 10-12. In some embodiments, a minimum Hamming distance may be greater than or equal to about 2, greater than or equal to about 3 , greater than or equal to about 4, greater than or equal to about 5, greater than or equal to about 6, greater than or equal to about 7, greater than or equal to about 8, greater than or equal to about 9, greater than or equal to about 10, greater than or equal to about 11, greater than or equal to about 12, greater than or equal to about 13 , greater than or equal to about 14, greater than or equal to about 15, greater than or equal to about 16, greater than or equal to about 17, greater than or equal to about 18, greater than or equal to about 19, or greater than or equal to about 20.
[0061] In some embodiments, a nucleic acid segment may comprise a universal primer binding site for amplification. In some embodiments, a nucleic acid segment may comprise multiple primer binding sites for amplification. For example, a segment may comprise an amplification primer binding site for performing rolling circle amplification (RCA) for generating a plurality of concatemeric amplification products.
[0062] In some embodiments, a nucleotide or nucleic acid sequence of each segment may correspond to one or more computational symbols, such as a detection color, for performing a decoding process. For example, one or more nucleic acid segments of a code may be detected with a first pool of detection polynucleotide complexes to produce one or more detectable binding complexes, for example by using a fluorescent label. In some embodiments, the one or more detectable binding complexes, once imaged, may produce one or more optical signals such as fluorescence in a particular wavelength. When all or substantially all segments of the code are detected by iteratively applying additional pools of detection polynucleotide complexes to the amplification products, a series of optical signals may be observed and collated.
[0063] The application of detection polynucleotides to amplification products for decoding can be called a “flow” or “cycle” or “query”, wherein a flow, cycle or query is the number of times a particular segment of an amplification product is queried, or the number of times a detection polynucleotide is flowed over an amplification product in order to detect a nucleic acid segm entsequence. If a nucleic acid segment is present, a detection polynucleotide that comprises a sequence complementary to that nucleic acid segment may hybridize to its complementary nucleic acid segment and the attached detectable label is detected, for example imaged. In some embodiments, one or more optical signals observed from querying an amplification product with detection polynucleotide complexes translates to one or more computational symbols such that each optical signal can be imaged and decoded. In some embodiments, a plurality of nucleic acid segments on a recognition element may correspond to at least three computational symbols. In some embodiments, the optical signal may be a color or a non -color. In some embodiments, the optical signal may be a combination of colors (e.g., when the detection polynucleotide complex comprises a plurality of detectable labels). In some embodiments, the computational symbols or colors can be referred to as numbers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.). In some embodiments, each detection polynucleotide complex comprises a detectable label, as such for example when four different fluorescent moieties are used as detectable labels there are four symbols, 1 to 4. However, the number of symbols can be larger depending on the combination of detectable labels with each unique detection polynucleotide complex. For example, for a set of 16 unique detection polynucleotide complexes wherein each detection polynucleotide complex has one offour fluorescent moieties, there may be 16 computational symbols used for decoding if all 16 unique detection polynucleotide complexes are used to decode an amplification product. However, additional ways to increase the number of computational symbols for decoding include, but are not limited to, adding levels of identifiability associated with a particular detectable signal such as whether a detectable signal is brighter or dimmer compared to a normal level of signal, or whether there is a combination of detectable colors that is used to identify a particular nucleotide or a nucleotide complex. As such, the number of computational symbols that may be used may be limited by practicality for any given assay.
[0064] In some embodiments, the methods described herein may use a number of computational symbols. The number of computational symbols used in the methods and systems described herein may be considered in the design of the recognition elements. For example, in some embodiments, a detection scheme using a larger number of computational symbols may lead to a larger code space and a greater number of codes that may be generated, which may allow for a greater amount of information that may be detected thereby allowing for a higher degree of assay target molecule multiplexing. In some embodiments, a detection scheme using a smaller number of computational symbols may be limited in the amount of information that can be detected. In some embodiments, using a larger number of computational symbols may result in a faster detection process (e.g., less time to determine a target molecule compared to using a smaller number of computational symbols). In some embodiments, a detection scheme using alarger number of computational symbols may require greater instrument complexity, which may lead to potential drawbacks such as color crosstalk, wherein the computational symbols used in the detection scheme may become difficult to distinguish from other computational symbols. In some embodiments, a greater number of computational symbols may require that a more complex detection tool be used.
[0065] In some embodiments, each nucleic acid segment may correspond to a combination of computational symbols. In some embodiments, each nucleic acid segment may correspond to one or more computational symbols, two or more computational symbols, three or more computational symbols, four or more computational symbols, five or more computational symbols, six or more computational symbols, seven or more computational symbols, eight or more computational symbols, nine or more computational symbols, or 10 or more computational symbols. In some embodiments, each nucleic acid segment may correspond to 10 or less computational symbols, nine or less computational symbols, eight or less computational symbols, seven or less computational symbols, six or less computational symbols, five or less computational symbols, four or less computational symbols, three or less computational symbols, or two or less computational symbols.
[0066] The methods described herein may include amplification of a circularized and ligated recognition element. In some embodiments, a target nucleic acid molecule is amplified. In some embodiments, the target nucleic acid molecule comprises a combination of a recognition element and a target nucleic acid molecule. In some embodiments, the amplification is selective amplification. For example, in some embodiments, the amplification is able to occur if a target recognition region of a recognition element recognizes and binds to a complementary target nucleic acid of interest. In some embodiments, amplification occurs if a primer is used that is complementary to one or more of a portion of a recognition element, a portion of a nucleic acid segment of a code, or another sequence in the recognition element that is complementary to a primer used for amplification. In some embodiments, the amplification is non -selective. For example, in some embodiments randomers can be used to prime amplification from a recognition element.
[0067] In some embodiments, the methods described herein may include selectively amplifying a subset of nucleic acids. For example, in some embodiments, a subset of a plurality of recognition elements hybridized to a plurality of target nucleic acid molecules may be amplified. The subset may comprise a percentage of the total amount of recognition elements hybridized to target nucleic acid molecules as described herein. In some embodiments, the subset may include 5% or more, 10% or more, 15% or more, 20% or more, 25% or more, 30% or more, 35% or more, 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70%or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more of the total amount of recognition elements hybridized to target nucleic acid molecules. In some embodiments, the subset may include 95% or less, 90% or less, 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the total amount of recognition elements hybridized to target nucleic acid molecules.
[0068] In some embodiments, the amplification may include rolling circle amplification (RCA). In some embodiments, the RCA may generate a concatemer as an amplification product, wherein the concatemer contains multiple copies of a circularized ligated recognition element, including associated codes, target recognition regions, and any other functional sequences that are included in the circularized and ligated recognition element. In some embodiments, RCA may be performed while the circularized and ligated recognition element is in solution. In some embodiments, RCA may be performed on a circularized recognition element while the circularized recognition element is immobilized, either reversibly or non -reversibly, on a substrate or surface. In some embodiments, the substrate or surface is a solid support and includes, but is not limited to, a bead, a flow cell, a microwell, a nanowell, a well, a slide. In some embodiments, the substrate is glass such as optical glass of imaging quality. In some embodiments, the substrate is plastic, polycarbonate, etc. In some embodiments, the substrate is positively charged or negatively charged. In some embodiments, the substrate is an anionic substrate. In some embodiments, the substrate is a cationic substrate. In some embodiments, the substrate comprises an immobilization composition, such as polyacrylamide, branched PEI, linear PEI, poly(P-aminoester) and poly(amidoamine), PEG, a gel, poly -L-ly sine, silane, agarose, muscle mimetic catecholamine polymer, and the like. In some embodiments, the substrate has no charge. In some embodiments, the substrate has no immobilization composition. In some embodiments, a substrate comprises a cationic polymer coated surface. An RCA reaction may be performed in the presence of a cationic polymer coated surface, resulting in simultaneous immobilization and amplification of a ligated recognition element. RCA primers may be supplied in solution or bound to the cationic polymer-coated surface prior to, or concurrent with, performing the RCA reaction.
[0069] In some embodiments, amplification may include on-surface polymerase chain reaction (PCR), isothermal amplification, RCA, ultrarapid amplification, or a combination thereof. In some embodiments, amplification may include polymerase chain reaction (PCR). In some embodiments, PCR is multiplexed PCR. The amplification methods disclosed herein may include isothermal amplification. Non-limiting examples of isothermal amplification include Nicking endonuclease amplification reaction (NEAR), Transcription mediated amplification(TMA), Loop-mediated isothermal amplification (LAMP), Helicase-dependent amplification (HD A), Nucleic Acid Sequence Based Amplification (NASBA), Strand displacement amplification (SDA), Multiple Displacement Amplification (MDA), Rolling Circle Amplification (RCA), bridge amplification, or Ramification (RAM) amplification method. In some embodiments, the amplification method is provided in Fakruddin M, Mannan KS, Chowdhury A, Mazumdar RM, Hossain MN, Islam S, Chowdhury MA. Nucleic acid amplification: Alternative methods of polymerase chain reaction. J Pharm Bioallied Sci. 2013 Oct;5(4):245-52, which is hereby incorporated by reference in its entirety.
[0070] The methods described herein may include introducing a detection probe to an amplified recognition element. A detection probe may be a single stranded oligonucleotide . Alternatively, a detection probe may be a partially single stranded and partially double stranded detection polynucleotide. A detection probe can comprise a detectable label.
[0071] In some embodiments, the methods described herein may include introducing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 50 or more, 100 or more, 200 ormore, 300 ormore, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1,000 or more detection probes to an amplification product. In some embodiments, the methods described herein may include introducing 1,000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, 50 or less, 25 or less, 20 or less, 15 or less, 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection probes to an amplification product.
[0072] In some embodiments, a detection probe may comprise a detectable label (e.g., fluorescent molecule). In some embodiments, a detection probe is a single stranded oligonucleotide with a portion that is complementary to a code, or a portion of a code, and a detectable label. In some embodiments, a detection probe comprises two oligonucleotides. Fig. 3A shows a non-limiting example of a structure of a detection probe 300 comprising two oligonucleotides. As shown in Fig. 3A, a first oligonucleotide 330 may comprise a detectable label 340. A second oligonucleotide 350 may comprise a portion that is complementary to the first oligonucleotide 310 and a second portion 320 that is complementary to a code or a portion of a code 360 (e.g., a nucleic acid segment of a code). The first oligonucleotide 330 may hybridize to the second oligonucleotide 350, thereby generating a detection polynucleotide. For decoding, a portion of the second oligonucleotide 320 may hybridize to its code complement 360 as seen in Fig- 3B, and a signal may be detected from the detectable label, therebyidentifying the code which in turn is correlated back to the presence of a target nucleic acid of interest from a sample.
[0073] The detection probe may comprise various nucleotide lengths. In some embodiments, the detection probe may comprise a length of about 5 to 25 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 20 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 15 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 10 nucleotides. In some embodiments, the detection polynucleotide may comprise a length of about 5 to 8 nucleotides.
[0074] In some embodiments, the detection probe may comprise a length of about 5-100 nucleotides, about 10-80 nucleotides, about 20-60 nucleotides, about 30-50 nucleotides, or about 15-30 nucleotides. In some embodiments, the detection probe may comprise one or more detectable labels. In some embodiments, the one or more detectable labels may comprise a fluorescent moiety. The fluorescent moiety may emit at red, far-red, near-red, yellow, green, or blue wavelengths. In some embodiments, the fluorescent moiety comprises one or more of 6- FAM (6 -carb oxy fluorescein), JOE (6-carboxy-4',5'-dichloro-2',7'-dimethoxyfluorescein), TAMRA (6-carboxytetramethylrhodamine), 5-Cy5 (5 -carboxyrhodamine), 5-Cy5.5 (5- carboxylic acid succinimidyl ester), 5 -Cyl (5-carboxyrhodamine), (hexachlorofluorescein), Alexa Fluor 488 (AF488), Alexa Fluor 514 (AF514), Texas Red, Cyanine 3, Cyanine 5, Pacific Blue, Tetramethyl rhodamine, Oxazole Yellow, Atto647N, and Rhodamine 6G (R6G). In some embodiments, the detection probe may comprise two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or 10 or more fluorescent moieties. In some embodiments, the detection polynucleotide may comprise 10 or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less fluorescent moieties.
[0075] In some embodiments, the fluorescent moiety may comprise an organic dye, a biological fluorophore, a quantum dot, or a combination thereof. In some embodiments, the organic dye may comprise an organic molecule. In some embodiments, the organic dye may comprise a coumarin, a cyanine, a benzofuran, a quinoline, a quinazolinone, an indole, a benzazole, a borapolyazaindacene, a xanthene, or a combination thereof. The organic dye may correspond to a color. For example, the organic dye may correspond to a green, a yellow, a blue, an indigo, a red, an orange a purple, a pink, a violet, or a combination thereof. In some embodiments, the organic dye may correspond to no color. In some embodiments, the organic dye may correspond to a black color. In some embodiments, the organic dye may correspond to a white color.
[0076] The detectable moiety can be identified by imaging. When the detectable label is a fluorophore, the fluorophore may emit a color in the visible light spectrum which can be captured by fluorescent imaging and associated filters. In some embodiments, the fluorophore may emit in a wavelength in the range from about 400 nanometers (nm) to 900 nm. In some embodiments, the fluorophore may emit in a wavelength between about 400 nm to 475 nm, about475 nm to 490 nm, about 490 nm to 530 nm, about 530 nm to 575 nm, about 575 nm to 600 nm, about 600 nm to 700 nm, or about 700 nm to 800 nm. In some embodiments, the fluorophore may emit a wavelength of 400 nm or more, 425 nm or more, 450 nm or more, 475 nm or more, 500 nm or more, 525 nm or more, 550 nm or more, 575 nm or more, 600 nm or more, 625 nm or more, 650 nm or more, 675 nm or more, 700 nm or more, 725 nm or more, 750 nm or more, 775 nm or more, 800 nm or more, 825 nm or more, 850 nm or more, 875 nm or more, or 900 nm or more. In some embodiments, the fluorophore may emit a wavelength of 900 nm or less, 875 nm or less, 850 nm or less, 825 nm or less, 800 nm or less, 775 nm or less, 750 nm or less, 725 nm or less, 700 nm or less, 675 nm or less, 650 nm or less, 625 nm or less, 600 nm or less, 575 nm or less, 550 nm or less, 525 nm or less, 500 nm or less, 475 nm or less, 450 nm or less, 425 nm or less, or 400 nm or less.
[0077] The detectable labels (e.g., fluorescent moieties) may be optically distinct. The number of optically distinct detectable labels used in the methods described herein can impact the amount of information that may be detected. For example, a detection scheme using a larger number of optically distinct detectable labels may allow for a higher amount of multiplexing of codes, which may in turn allow for a greater amount of target molecule related information to be detected and captured. A detection scheme using a smaller number of optically distinct detectable labels may allow for a lesser amount of target molecule related information to be detected and captured. In some embodiments, using a larger number of optically distinct detectable labels may lead to a detection process that identifies a target molecule in less time as compared to using a fewer number of optically distinct detectable labels when querying a concatemeric amplification product. In some embodiments, a detection scheme using a larger number of optically distinct detectable labels may lead to greater instrument complexity, which may lead to fluorescence detection crosstalk, whereby the fluorescence emission spectra of the optically distinct fluorescent moieties may not yield distinct fluorescence signals. In some embodiments, a detection scheme using a greater number of optically distinct fluorescent moieties may require use of a more complex detection tool.
[0078] In some embodiments, the detection polynucleotides may be provided in one or more detection pools. In some embodiments, the methods herein may use one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine ormore, ten or more, 15 or more, 20 or more, 25 or more, 30 or more, 40 or more, 45 or more, or 50 or more detection pools. In some embodiments, the methods herein may use 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection pools.
[0079] In some embodiments, each detection pool provided may comprise a number of detection polynucleotides. In some embodiments, each detection pool may comprise two or more, three or more, four or more, five or more, ten or more, 15 or more, 25 or more, 50 or more, 100 or more, 150 or more, 250 or more, 500 or more, 1,000 or more, 1,500 or more, 2,500 or more, or 5,000 or more detection polynucleotides. In some embodiments, each detection pool may comprise 5,000 or less, 2,500 or less, 1,500 or less, 1,000 or less, 500 or less, 250 or less, 150 or less, 100 or less, 50 or less, 25 or less, 15 or less, ten or less, nine or less, eight or less, seven or less, six or less, five or less, four or less, three or less, or two or less detection polynucleotides.
[0080] The number of detection pools and the number of detection polynucleotides in each detection pool may be considered in the design of the recognition elements. For example, one advantage to using a smaller number of detection pools and detection polynucleotides in the methods described herein may be to lower design costs. Conversely, one advantage to using a larger number of detection pools and detection polynucleotides in the methods described herein may be the need for a higher degree of multiplexing for target molecule detection and larger amounts of information that may be detected.
[0081] The methods described herein may include imaging a plurality of detection polynucleotides. The detection polynucleotides may have hybridized to their complementary code or a portion of a code in order to obtain identifiable signals which can be correlated back to the presence of a target molecule of interest. In some embodiments, the signals are associated with one or more segments of a code for each concatemeric amplification product. In some embodiments, the imaging is performed by an imaging system comprising a fluorescence detection system.
[0082] In some embodiments, the imaging may be conducted using an imaging system. The imaging system may comprise at the minimum a camera, a detector, an illuminator, a condenser, or a combination thereof. In some embodiments, the imaging may include images of fluorescence emission, luminescence, or a combination thereof. In some embodiments, the imaging system may comprise components or sub-systems of a larger system that may also include optics modules including when needed fluorescence filters, fluidics modules, temperature control modules, translation stages, robotic fluid dispensing and / or microplate handling, processors or computers, instrument control software, data analysis and displaysoftware, etc. In some embodiments, the imaging system may be a fluorescence imaging system. In some embodiments, the imaging may include fluorescent images from the fluorescent moieties present on the labeled probes.
[0083] In some embodiments, the image may comprise fluorescence information from one or more wavelengths. In some embodiments, the fluorescence information may comprise emission data from a wavelength from about 220-830 nanometers (nm), about 230-820 nm, about 240- 810 nm, about 250-800 nm, about 260-790 nm, about 270-780 nm, about 280-770 nm, about 290-760 nm, about 300-750 nm, about 310-740 nm, about 320-730 nm, about 330-720 nm, about 340-7 lO nm, about 350-700 nm, about 360-690 nm, about 370-680 nm, about 380-670 nm, about 390-660 nm, about 400-650 nm, about 410-640 nm, about 420-630 nm, about 430- 620 nm, about 440-610 nm, about 450-600 nm, about 460-590 nm, about 470-580 nm, about 480-570 nm, about 490-560 nm, about 500-550 nm, about 510-540 nm, about 520-530 nm, or a combination thereof.
[0084] Image detection and capture may relate to iteratively repeating the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof. In some embodiments, the iterative repetition of the operations is performed for each nucleic acid segment of a code one or more times.
[0085] In some embodiments, the iteratively repeating the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof may comprise two or more iterative repetitions. For example, the methods described herein may comprise about 2-50 iterative repetitions, about 2-10 iterative repetitions, or about 2-8 iterative repetitions. In some embodiments, the method described herein may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, or 50 or more iterative repetitions. Each iterative repetition may comprise : (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof.
[0086] In some embodiments, the number of iterative repetitions of the operations of : (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging thesignals of the hybridized detection polynucleotides to a code or a portion thereof may correspond to the number of nucleic acid segments present in a code of the recognition element.In some embodiments, each nucleic acid segment of the code of the recognition element may undergo a number of iterative repetitions of the operations of: (i) introducing detection polynucleotides to the concatemeric amplification products; (ii) hybridizing the detection polynucleotides to a code or a portion thereof; and (iii) imaging the signals of the hybridized detection polynucleotides to a code or a portion thereof. For example, the methods described herein may comprise iteratively repeating the operations described herein two or more times per segment, three or more times per segment, or four or more times per segment. In some embodiments, the methods described herein may comprise iteratively repeating the operations described herein two or more times per segment, three or more times per segment, four or more times per segment, five or more times per segment, six or more times per segment, seven or more times per segment, eight or more times per segment, nine or more times per segment, 10 or more times per segment, 11 or more times per segment, 12 or more times per segment, 13 or more times per segment, 14 or more times per segment, 15 or more times per segment, 16 or more times per segment, 17 or more times per segment, 18 or more times per segment, 19 or more times per segment, or 20 or more times per segment.
[0087] Additional methods for imaging a detection polynucleotide can be found in WO2023 / 158993A2, which is incorporated herein by reference in its entirety.
[0088] Several models may be used to identify a code that is associated with a target molecule based on the fluorescence signals generated and images captured from detection polynucleotide hybridization to code sequences. In one embodiment, decoding makes use of a hard decision decoding model. In another embodiment, decoding makes use of a soft decision decoding model.
[0089] For soft decision decoding, it may not be necessary to identify each base specifically. For example, signals generated during each detection event may be detected and recorded to produce a data set that may be used as input into a model to calculate a probability that a specific code is present without requiring that each base of a code be determined. Although it may not be necessary in a soft decision decoding model to make a hard decision about the identity of each nucleotide, a model may nevertheless include assigning a probability or identity to each nucleotide in the sequence of a code, wherein each nucleotide in the sequence of a code could be sequenced. Data gathered may include intensity readings for signals produced by the hybridized detection polynucleotide fluorescent moiety in various spectral bands. A set of intensity readings may be detected by imaging, stored and used as input into a soft decision decoding model fordetermining a probability that a particular code is present, and hence a target nucleic acid is present in the sample.
[0090] A model may be developed or trained using data from known codes, such as signal intensity data across a predetermined spectrum. The model may be used to calculate a set of probabilities across a set of one or more codes, indicating, for example, for each code, a probability that it is present in a concatemeric amplification product.
[0091] The probability that a particular code is present may be indicative of the probability that a particular target molecule associated with the code is present in the sample of interest. Data indicating the probability that a particular target is present may be, for example, to calculate probabilities relevant to diagnosis or screening of various medical conditions, or selection of drugs for treatment of various medical conditions.
[0092] A soft decoding decision model may include using an algorithm to predict the presence of target molecules from a sample. In some embodiments, the algorithm is a soft-decision decoding algorithm. In some embodiments, the algorithm is applied to the codes of the concatemeric amplification products for predicting the presence of a target molecule from a sample.
[0093] The methods disclosed herein may comprise soft decision decoding to predict the presence of the code in a recognition element or concatemeric amplification product thereof, wherein the presence of the code correlates and serves as a proxy for the presence of a target nucleic acid in a sample. In some embodiments, the methods described herein may use soft decision decoding. In some embodiments, the methods described herein may use hard decision decoding. For hard decision decoding, signals from queried concatemers may be extracted from images. This may be the same for soft decision decoding, in that signals that are generated and imaged are extracted from the images. For hard decision decoding, hard symbol calls may be generated from the intensities of the signals, whereas with soft decision decoding no hard symbol calls may be necessary as all of the signal range is retained. The code assignment for hard decision decoding may be determined by matching symbol-by-symbol readouts of codewords to codes, whereas with soft decision decoding, the signals may be cross correlated against the expected signals and a code assigned using a probabilistic methodology. When using soft decision decoding, it may not be necessary for the model to identify each symbol specifically. For example, signals (e.g., fluorescent signals) generated during each cycle of a detection process maybe detected and recorded to produce a data set that may be used as input into a model to calculate a probability that a specific code is present.
[0094] The permutation space on a recognition element is the totality of factors that determines the number of unique nucleotide possibilities at each nucleic acid segment. Factors maycomprise the number of segments present on a recognition element, the number of incubation periods or times a segment is queried with a detection pool comprising detection polynucleotides, and the number of computational symbols or colors.
[0095] Fig. 4 details an example of a soft decision decoding workflow for determining the presence of a target molecule from a sample based on detection and decoding of a code associated with the target molecule that originally hybridized to a recognition element. Images of the sample may be acquired, aligned, and processed to extract the intensity of the features, or signals of interest across the imaged field of view in multiple spectral channels. The corrected intensities of said features may be fed through a series of algorithms that make up the soft decoder. At first, the intensity profiles of the codes may be learned based on features of high confidence or high intensity. This trained model may provide a template for each code from which the rest of the features of interest may be compared to in the second operation. Third, a confidence score may be computed from the difference between the intensity profile of each feature and the trained profiles. Several filters may be applied to remove outliers, duplicates, and low confidence decoded concatemers. The output may comprise a table of decoded concatemers with an associated filter status, confidence score, and most likely assignment to one of the codes of one or more concatemeric amplification products.
[0096] In some embodiments, a recognition element comprises a larger code, for example a code with four segments instead of two or three. In some embodiments, a recognition element comprising a larger code may result in a detection scheme with better error correction as compared to a detection scheme having a recognition element comprising a smaller code . Additionally, in some embodiments, a larger code may result in a lower signal -to-noise ratio as compared to a detection scheme having a recognition element comprising a smaller code .
[0097] In some embodiments, a recognition element comprises a smaller code, for example a code with two segments, or one segment, as compared to a recognition element with four segments. In some embodiments, a recognition element comprising a smaller code may result in a detection scheme with lower error correction abilities as compared to a detection scheme with a larger code. Further, in some embodiments, a smaller code may result in a higher signal-to- noise ratio as compared to a larger code.
[0098] In some embodiments, the systems may comprise a solid substrate configured to immobilize one or more of a circularized and ligated recognition element, a concatemeric amplification product, a detection polynucleotide, and a hybridized complex of a concatemeric amplification product and a detection polynucleotide complex. In some embodiments, the systems comprise a welled plate or a flowcell. In some embodiments, the systems may comprisea fluid flow controller, a temperature controller, an imaging system, a computer system, or any combination thereof.
[0099] In some embodiments, the systems disclosed herein may include a solid substrate or a solid surface. The solid substrates and surfaces disclosed herein may be referred to as a substrate, a support, a solid support, or a surface. The substrate may be modified for immobilizing circularized and ligated recognition elements or concatemeric amplification products, or both. Example solid substrates include, but are not limited to, glass, modified or functionalized glass, plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, inorganic glasses, plastics, optical fiber bundles, optically clear glass, and other polymers. In some embodiments, the plastic solid substrates may include acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, or polyurethanes. In some embodiments, the silica -based solid substrates may include silicon or modified silicon.
[0100] In some embodiments, the substrate may be a welled plate comprising a number of wells. In some embodiments, the substrate may be a 96-well plate. In some embodiments, the substrate may be a 4-well plate, a 6-well plate, an 8-well plate, a 12-well plate, a 24-well plate, a 48-well plate, a 384-well plate, an 864-well plate, or a 1,536-well plate. In some embodiments, the substrate may have greater than or equal to 96 wells. In some embodiments, the substrate may have less than or equal to 96 wells.
[0101] In some embodiments, the substrate may be a flowcell. In some embodiments, the flowcell may have two or more lanes. In some embodiments, the flowcell may have two or less lanes.
[0102] In some embodiments, the substrate may be a microarray, a slide, a chip, a microwell, a tube, a column, a particle, a bead, or a paramagnetic bead.
[0103] In some embodiments, the substrate may comprise a coating. In some embodiments, the coating may comprise a layer that may be charged. In some embodiments, the coating layer may be positively charged. In some embodiments, the coating layer may be negatively charged. In some embodiments, the coating may be non-charged. In some embodiments, the substrate may comprise a surface comprising a cation-coating layer. In some embodiments, the substrate may comprise a surface comprising an anion-coating layer. In some embodiments, the substrate may comprise a surface comprising a neutral-charged layer. In some embodiments, the substrate may be coated with streptavidin. In some embodiments, the substrate may be coated with avidin. In some embodiments, the substrate may be coated with one or more antibodies.
[0104] The systems disclosed herein may comprise a fluidics system. The fluidics system may comprise a fluid flow controller. In some embodiments, the fluid flow controller may compriseone or more pumps, valves, mixing manifolds, reagent reservoirs, waste reservoirs, or any combination thereof. In some embodiments, the fluidic system and subcomponents of the fluidics system are fluidically connected to the reaction vessel of the present disclosure. In some embodiments, the fluidic system and subcomponents of the fluidics system iteratively flow in reagents (e.g., buffers, detector polynucleotides, anchor polynucleotides, detection oligonucleotide complexes, etc.) to the reaction vessel. In some embodiments, the reaction vessel comprises a solid substrate configured to immobilize the circularized and ligated recognition elements or concatemeric amplification products thereof.
[0105] The systems disclosed herein may comprise a temperature system. The temperature system may comprise a temperature controller. The temperature controller may be incorporated into the systems described herein to facilitate accuracy of the methods and systems described herein. In some embodiments, the temperature controller may comprise temperature control components. Non-limiting examples of temperature control components include resistive heating elements, infrared light sources, heating or cooling devices, heat sinks, thermocouples, thermistors, or a combination thereof. In some embodiments, the temperature controller may provide changes in temperature over specified time intervals. In some embodiments, the temperature controller may provide an increase in temperature. In some embodiments, the temperature controller may provide a decrease in temperature. In some embodiments, the temperature controller may provide for cycling of temperatures between two or more set temperatures so that thermocycling or amplification may be performed. In some embodiments, the temperature controller may provide a constant temperature.
[0106] The systems disclosed herein may comprise an imaging system. In some embodiments, signals produced by the labeled probes disclosed herein may be imaged by the imaging systems disclosed herein. The imaging system may comprise one or more light sources, one or more optical components, one or more filters, one or one or more imaging sensors for imaging and detection, or a combination thereof. In some embodiments, the one or more light sources may comprise light from a bulb. In some embodiments, the one or more optical components may comprise lenses, mirrors, digital mirror devices, prisms, optical filters, colored glass filters, narrowband interference filters, broadband interference filters, dichroic reflectors, diffraction gratings, apertures, optical fibers, optical waveguides, or a combination thereof. In some embodiments, the one or more imaging sensors may comprise a charge -coupled device (CCD) sensor or camera, a complementary metal-oxide-semiconductor (CMOS) imaging sensor or camera, a negative-channel metal-oxide semiconductor (NMOS) imaging sensor or camera, or a combination thereof.
[0107] Various operations of the methods and systems disclosed herein may be performed by a computer system of the present disclosure. Referring to Fig. 5, a non-liming example of a block diagram is shown depicting a non-limiting example of a machine that includes a computer system 500 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in Fig. 5 are non-limiting examples and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0108] Computer system 500 may include one or more processors 501, a memory 503, and a storage 508 that communicate with each other, and with other components, via a bus (solid lines). The bus may also link a display 532, one or more input devices 533 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 534, one or more storage devices 535, and various tangible storage media 536. All of these elements may interface directly or via one or more interfaces or adaptors to the bus. For instance, the various tangible storage media 536 can interface with the bus via storage medium interface 526. Computer system 500 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0109] Computer system 500 includes one or more processor(s) 501 (e.g., central processing units (CPUs), general purpose graphics processing units (GPGPUs), or quantum processing units (QPUs)) that carry out functions. Processor(s) 501 optionally contains a cache memory unit 502 for temporary local storage of instructions, data, or computer addresses. Processor(s) 501 are configured to assist in execution of computer readable instructions. Computer system 500 may provide functionality for the components depicted in Fig. 5 as a result of the processor(s) 501 executing non -transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 503, storage 508, storage devices 535, and / or storage medium 536. The computer-readable media may store software that implements particular embodiments, and processor(s) 501 may execute the software. Memory 503 may read the software from one or more other computer-readable media (such as mass storage device(s) 535, 536) or from one or more other sources through a suitable interface, such as network interface 520. The software may cause processor(s) 501 to carry out one or more processes or one or more operations of one or more processes described or illustrated herein. Carrying outsuch processes or operations may include defining data structures stored in memory 503 and modifying the data structures as directed by the software.[HO] The memory 503 may include various components (e.g., machine readable media) including, but not limited to, a random-access memory component (e.g., RAM 504) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phase - change random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 505), and any combinations thereof. ROM 505 may act to communicate data and instructions unidirectionally to processor(s) 501, and RAM 504 may act to communicate data and instructions bidirectionally with processor(s) 501. ROM 505 and RAM 504 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 506 (BIOS), including basic routines that help to transfer information between elements within computer system 500, such as during start-up, may be stored in the memory 503.[Hl] Fixed storage 508 is connected bidirectionally to processor(s) 501, optionally through storage control unit 507. Fixed storage 508 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 508 may be used to store operating system 509, executable(s) 510, data 511, applications 512 (application programs), and the like. Storage 508 can also include an optical disk drive, a solid-state memory device (e.g., flash -based systems), or a combination of any of the storage disclosed herein. Information in storage 508 may, in appropriate cases, be incorporated as virtual memory in memory 503.
[0112] In one example, storage device(s) 535 may be removably interfaced with computer system 500 (e.g., via an external port connector (not shown)) via a storage device interface 525. Particularly, storage device(s) 535 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 500. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 535. In another example, software may reside, completely or partially, within processor(s) 501.
[0113] Bus connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example, and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), aPeripheral Component Interconnect (PCI) bus, a PCI -Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0114] Computer system 500 may also include an input device 533. In one example, a user of computer system 500 may enter commands and / or other information into computer system 500 via input device(s) 533. Examples of an input device(s) 533 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 533 may be interfaced to bus via any of a variety of input interfaces 523 (e.g., input interface 523) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any of the input devices disclosed herein.
[0115] In particular embodiments, when computer system 500 is connected to network 530, computer system 500 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 530. Communications to and from computer system 500 may be sent through network interface 520. For example, network interface 520 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 530, and computer system 500 may store the incoming communications in memory 503 for processing. Computer system 500 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 503 and communicated to network 530 from network interface 520. Processor(s) 501 may access these communication packets stored in memory 503 for processing.
[0116] Examples of the network interface 520 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 530 or network segment 530 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 530, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0117] Information and data can be displayed through a display 532. Examples of a display 532 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 532 can interface to the processor(s) 501, memory 503, and fixed storage 508, as well as other devices, such as input device(s) 533, via the bus. The display 532 is linked to the bus via a video interface 522, and transport of data between the display 532 and the bus can be controlled via the graphics control 521. In some embodiments, the display is a video projector. In some embodiments, the display is a head- mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0118] In addition to a display 532, computer system 500 may include one or more other peripheral output devices 534 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus via an output interface 524. Examples of an output interface 524 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0119] In addition, or as an alternative, computer system 500 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more operations of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0120] Various illustrative logical blocks, modules, circuits, and algorithm operations described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations have been described herein generally in terms of their functionality.
[0121] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general -purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0122] The operations of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium. An example of a storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0123] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, notepad computers, set -top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Select televisions, video players, and digital music players with optional computer network connectivity may be suitable for use in the system described herein. Suitable tablet computers, in various embodiments, include those with booklet, slate, and convertible configurations.
[0124] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Suitable server operating systems may include, by way of non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Suitable personal computer operating systems may include, by way of non-limiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX- like operating systems such as GNU / Linux®. In some embodiments, the operating system isprovided by cloud computing. Suitable mobile smartphone operating systems may include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®. Suitable media streaming device operating systems may include, by way of non-limiting examples, Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung® HomeSync®. Suitable video game console operating systems may include, by way of non-limiting examples, Sony® PS3®, Sony® PS4®, Microsoft® Xbox 360®, Microsoft Xbox One, Nintendo® Wii®, Nintendo® Wii U®, and Ouya®.
[0125] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non -transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semi -permanently, or non-transitorily encoded on the media.
[0126] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, which perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, a computer program may be written in various versions of various languages.
[0127] The functionality of the computer readable instructions may be combined or distributed in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications,one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.
[0128] In some embodiments, the computer programs described herein may be used to perform at least one function. The computer programs described herein may perform functions related to storing data, receiving data, analyzing data, exporting data, or a combination thereof. In some embodiments, the computer programs described herein may perform functions related to applying selection criteria, including in silico selection criteria, functional selection criteria, or a combination thereof. In some embodiments, the computer programs may receive sequence information, including sequence information for nucleic acid segments. The sequence information may be configured as an array, a table, a list, or combination thereof. The sequence information may be formatted in a variety of ways, including, but not limited to a .txt file, a FASTA file, an .xls file, or a combination thereof. The computer programs described herein may apply selection criterion or selection criteria to a set of nucleic acid segments. The computer programs may sort the nucleic acid segments, determine or compute characteristics of the nucleic acid segments, perform calculations, reorder the nucleic acid segments, or a combination thereof. In some embodiments, the computer programs described herein may store information related to the nucleic acid segments. In some embodiments, the computer program may use information stored related to the nucleic acid segments to apply selection criteria to the nucleic acid segments. In certain embodiments, the computer program may receive information and / or data related to nucleic acid segments, selection criteria, or a combination thereof. In some embodiments, the computer programs may perform functions related to analyzing data from functional assays, including, but not limited to functional assays described herein. In some embodiments, analyzing data from functional assays may comprise image analysis, image quantification, intensity quantification, feature identification, or a combination thereof. The computer programs described herein may also export information. In some embodiments, the exported information may comprise images, files, data tables, documents, folders, or a combination thereof.
[0129] In some embodiments, a computer program includes a web application. In light of the disclosure provided herein, a web application, in various embodiments, may utilize one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non-limiting examples, relational, non-relational, object oriented, associative, XML, and document oriented database systems. In further embodiments, suitable relational database systems include, by way of non -limiting examples, Microsoft® SQL Server,mySQL™, and Oracle®. A web application, in various embodiments, may be written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof . In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In some embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM® Lotus Domino®. In some embodiments, a web application includes a media player element. In various further embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non-limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java™, and Unity®.
[0130] Referring to Fig. 6, in a particular embodiment, an application provision system comprises one or more databases 600 accessed by a relational database management system (RDBMS) 610. Suitable RDBMSs include Firebird, MySQL, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBMDB2, IBM Informix, SAP Sybase, Teradata, and the like. In this embodiment, the application provision system further comprises one or more application severs 620 (such as Java servers, .NET servers, PHP servers, and the like) and one or more web servers 630 (such as Apache, IIS, GWS and the like). The web server(s) optionally expose one or more web services via app application programming interfaces (APIs) 640. Via a network, such as the Internet, the system provides browser-based and / or mobile native user interfaces.
[0131] Referring to Fig. 7, in a particular embodiment, an example of an application provision system alternatively has a distributed, cloud-based architecture 700 and comprises elastically load balanced, auto-scaling web server resources 710 and application server resources 720, as well as synchronously replicated databases 730.
[0132] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobilecomputing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.
[0133] In view of the disclosure provided herein, a mobile application is created by techniques using, for example, hardware, languages, development environments, and the like. Mobile applications may be written in several languages. Suitable programming languages include, by way of non-limiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.
[0134] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non -limiting examples, Airplay SDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.
[0135] Several commercial forums may be available for distribution of mobile applications including, by way of non-limiting examples, Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop.
[0136] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Standalone applications may be compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.
[0137] In some embodiments, the computer program includes a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easilyaddingnew features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Several web browser plug-ins may be used in the methods and systems herein, including, for example, Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands.
[0138] In view of the disclosure provided herein, several plug-in frameworks may be available that enable development of plug-ins in various programming languages, including, by way of non-limiting examples, C++, Delphi, Java™, PHP, Python™, and VB .NET, or combinations thereof.
[0139] Web browsers (also called Internet browsers) are software applications, designed for use with network-connected computing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of nonlimiting examples, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, andKDEKonqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini-browsers, and wireless browsers) are designed for use on mobile computing devices including, by way of non- limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non -limiting examples, Google® Android® browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony® PSP™ browser.
[0140] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules may be created by techniques using, for example, machines, software, languages, and the like. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, aplurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non -limiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.
[0141] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, many databases maybe suitable for storage and retrieval of nucleic acid segment sequences or analysis thereof information. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity -relation ship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, andMongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.
[0142] Provided herein are kits related to the methods, compositions and systems described herein. In some embodiments, the kits may comprise a plurality of recognition elements, a plurality of detection polynucleotides, one or more buffers, one or more reagents, instructions for use, a manual, a protocol, or a combination thereof.
[0143] In some embodiments, a kit may comprise one or more buffers. In some embodiments, a kit may comprise two buffers. In some embodiments, a first buffer of a kit may be configured to promote hybridization. In some embodiments, a second buffer of a kit may be configured to promote de-hybridization, ligation, nucleic acid digestion, storage of a purified molecule. In some embodiments, a kit may comprise one or more reagents. In some embodiments, a kit comprises one or more enzymes. In some embodiments, a kit comprises one or more of a ligase, a DNA polymerase, and an exonuclease. In some embodiments, a kit may comprise instructions for use, a manual, a protocol, or a combination thereof. In some embodiments, a kit comprisesone or more 96 well plates. In some embodiments, one of the 96 well plates of a kit is configured to be assayed by an optical imaging device described herein.
[0144] The present disclosure and inventive concepts described herein relate to the detection of chromosomal aneuploidies from a cell-free DNA sample, for example a blood sample procured from a female, for example a pregnant female. The cell-free DNA samples can comprise DNA from a fetus, as such a cell-free DNA sample can include DNA from a mother and a fetus that can be utilized in methods for noninvasive prenatal testing for detecting fetal genetic anomalies. Genetic anomalies, which can be detected in a cell-free DNA sample includes, but are not limited to, chromosomal aneuploidies such as Trisomy 13 (also known as Patau syndrome), Trisomy 18 (also known as Edwards syndrome), and Trisomy 21 (also known as Down syndrome).
[0145] The methods, composition and system may relate to detection of targets of interest in cell-free DNA using a plurality of recognition elements. In some embodiments, each recognition element starts out as a linear oligonucleotide comprising a 5’ end region and a 3’ end region, a code that uniquely identifies each recognition element based on the 5’ and 3 ’ ends that are complementary to a target of interest, in this case sequences present in cell-free DNA, and optionally one or more functional sequences such as one or more amplification primer sequences, a cleavage sequence, one or more sequencing related sequences, and a sample index.
[0146] In some embodiments, the methods described herein relate to the detection chromosomal aneuploidies. In some embodiments, the methods described herein detect fetal fraction SNPs. In some embodiments, the methods described herein may relate to formation of circular DNA, wherein the circular DNA may comprise a code used as a proxy for a target of interest. In some embodiments, the circular DNA may comprise a recognition element. In some embodiments, the recognition element may hybridize to sequences of the target of interest, such as targets found in cell-free DNA.
[0147] The methods described herein may relate to the detection of targets of interest within cell-free DNA using recognition elements. The methods may comprise multiple operations, including but not limited to: i) receiving a sample comprising nucleic acids, wherein the nucleic acids are extracted from a sample such as a blood sample; ii) subjecting the extracted nucleic acids to one or more recognition elements, wherein a recognition element comprises (1) a code and (2) a 5’ and 3 ’ region that are complementary to target sequences in the extracted cell-free DNA or a known sequence associated with the target sequences, wherein the recognition element is configured to hybridize at the 5 ’ and 3 ’ ends to target sequences in the extracted cell- free DNA, such that the 5’ and 3 ’ regions are brought adjacent to each other and which the adjacent 5’ and 3 ’ ends can be ligated to form a circularized and ligated recognition element inthe presence of the target sequences; iv) amplifying the ligated recognition element to generate an amplified recognition element, wherein the amplified recognition element comprises more than one copy of the code; and v) detecting the more than one copy of the code, thereby detecting the target sequences of the cell-free DNA.
[0148] In some embodiments, nucleic acids for use in methods of the present disclosure may be extracted from a biological sample. In some embodiments, the biological sample may comprise whole blood, serum, or plasma. A whole blood sample may comprise veinous blood or capillary blood (e.g., obtained by a finger prick). In some embodiments, a whole blood sample may be obtained using a blood collection tube such as a Vacutainer tube, wherein the tube contains an anti-coagulant such as EDTA, citrate, heparin, and the like. When a blood collection tube comprises an anti -coagulant, the tube can be spun down to separate the red cells, white cells and plasma, wherein the plasma is collected, and cell-free DNA extracted therefrom. In some embodiments, a blood collection tube does not include an anti -coagulant. When a blood collection tube does not include an anti -coagulant, the blood may be allowed to clot, separating the cells, fibrin clot and the like from the serum, wherein the serum is collected, and cell-free DNA extracted therefrom.
[0149] The methods described herein may include amplification of ligated recognition elements. In some embodiments, amplification of ligated recognition elements may comprise rolling circle amplification which generates concatemeric amplification products. The concatemeric amplification products may comprise multiple copies of the ligated recognition element, which itself may comprise the complementary sequences of the target sequences of interest, a code made up one or more distinct nucleic acid segments that uniquely identifies the complementary sequence of the target sequences of interest, and optionally one or more functional sequences such as one or more amplification primer sequences, sequencing related sequences, cleavage sequences, sample index sequences, and the like.
[0150] In some embodiments, the 5’ end region of the recognition element and the 3 ’ end region of the recognition element may be able to hybridize to two adjacent sequences on a nucleic acid molecule from the cell-free DNA. Once the recognition element recognizes and hybridizes to the sequences on the nucleic acid molecule, the hybridized recognition element can assume a padlock probe configuration such that the 5’ end and 3 ’ end of the recognition element are adjacent to one another. In some embodiments, the 5’ end of the recognition element is phosphorylated. In some embodiments, the 3’ end of the recognition element may comprise a hydroxyl group. The sequence within the recognition element described herein that hybridizes to the complementary sequences from the cell-free DNA may be a variety of lengths. In some embodiments, the 5 ’ and / or 3 ’ regions of the recognition element that recognize and hybridize tocomplementary sequences from the cell-free DNA maybe about 5-50 nucleotides (nt), about 10- 60 nt, about 15-35 nt, about 20-30 nt, about 22-26 nt, about 5 nt, about 10 nt, about 15 nt, about 20 nt, about 22 nt, about 26 nt, about 30 nt, about 25 nt, about 40 nt or about 60 nt. In some embodiments, each of the 5’ region and the 3’ region comprise the same length. In some embodiments, the 5’ region is longer than the 3’ region. In some embodiments, the 3’ region is longer than the 5’ region.
[0151] In some embodiments, the adjacent 5’ end region and the 3’ end region s when hybridized to complementary sequences of cell-free DNA are ligated together. In some embodiments, the ligase may comprise T4 ligase, T3 ligase, T7 ligase, E. coli DNA ligase, Taq DNA ligase, a ligase from a Chlorella species, AmpLigase, RtcB ligase, or a combination thereof. In some embodiments, the ligase is a thermolabile ligase. In some embodiments, the ligase is a thermostable ligase.
[0152] In some embodiments, a ligated recognition element may be amplified using a polymerase. In some embodiments, the amplification is initiated by hybridization of a primer sequence, for example an amplification primer sequence found in the recognition element. In some embodiments, a portion of a code, for example a nucleic acid segment that makes up a code, can be used as an amplification primer sequence. In some embodiments, the amplification primer sequence is universal to every recognition element. In some embodiments, the polymerase may be DNA polymerase. In some embodiments, the polymerase may be an RNA polymerase. In some embodiments, the polymerase may be a stand -displacing polymerase. In some embodiments, the polymerase may be performed under isothermal conditions. In some embodiments, the polymerase may comprise Bst 2.0 DNA Polymerase, Bst 2.0 WarmStart® DNA Polymerase, Bst 3.0 DNA Polymerase, Bst DNA Polymerase, Full Length, Bst DNA Polymerase, Large Fragment, Klenow Fragment (3 ' — 5' exo-), phi29 DNA Polymerase, or a combination thereof. In some embodiments, the amplification may comprise additional components or reagents including but not limited to dNTPs, MgCl2, buffer, detergents, nucleic acid crowding agent(s), or a combination thereof. In some embodiments, the amplification of a ligated recognition element generated a concatemeric amplification product.
[0001] In some embodiments, each recognition element with complementary sequences to nucleic acid sequence of interest in the cell-free DNA comprises a code that uniquely identifies the recognition element and hence the nucleic acid sequence of interest in the cell-free DNA. The code that is there for associated with the target of interest in the cell-free DNA may comprise two or more nucleic acid segments, wherein the detection and decoding of the specific combination of nucleic acid segments serves as a proxy for the presence of the target of interest in the cell-free DNA. In some embodiments, every nucleic acid segment comprised within acode in a recognition element is detected as a part of the code. In some embodiments, a subset of nucleic acid segments comprised within a code of a recognition element is detected as part of the code. In some embodiments, one or more nucleic acid segments may be involved in other functions in addition to or alternative to its use as a part of a code described herein. In some embodiments, a nucleic acid segment may be used for example as a primer for amplification, a target recognition element, or a combination thereof.
[0153] The methods described herein may relate to the detection of chromosomal aneuploidies, including the detection of Trisomy 18, also known as Edwards Syndrome. In some embodiments, the detection of Trisomy 18 may comprise detecting sequences associated with chromosome 18. In some embodiments, one or more loci within chromosome 18 may be detected. In some embodiments, detection of the one or more loci within chromosome 18 may be preceded by selective amplification of the locus or loci. In some embodiments, the detection of the loci may be performed by any one of or combination of methods described herein. In some embodiments, additional genetic features may be analyzed in combination with analysis of chromosome 18 loci for detecting Trisomy 18. In some embodiments, the additional genetical features may be other chromosome sequences. In some embodiments, analyzing other genetic features may be performed for the purposes of normalizing data associating with chromosome 18.
[0154] The methods described herein may relate to the detection of chromosomal aneuploidies, including the detection of Trisomy 21, also known as Down Syndrome. In some embodiments, the detection of Trisomy 21 may comprise detecting sequences associated with chromosome 21. In some embodiments, one or more loci within chromosome 21 may be detected. In some embodiments, detection of the one or more loci within chromosome 21 may be preceded by selective amplification of the locus or loci. In some embodiments, the detection of the loci may be performed by any one of or combination of methods described herein. In some embodiments, additional genetic features may be analyzed in combination with analysis of chromosome 21 loci for detecting Trisomy 21. In some embodiments, the additional genetical features may be other chromosome sequences. In some embodiments, analyzing other genetic features may be performed for the purposes of normalizing data associating with chromosome 21.
[0155] The methods described herein may relate to the detection of chromosomal aneuploidies, including the detection of Trisomy 13, also known as Patau Syndrome. In some embodiments, the detection of Trisomy 13 may comprise detecting sequences associated with chromosome 13. In some embodiments, one or more loci within chromosome 13 may be detected. In some embodiments, detection of the one or more loci within chromosome 13 may be preceded by selective amplification of the locus or loci. In some embodiments, the detection of the loci maybe performed by any one of or combination of methods described herein. In some embodiments, additional genetic features may be analyzed in combination with analysis of chromosome 13 loci for detecting Trisomy 13. In some embodiments, the additional genetical features may be other chromosome sequences. In some embodiments, analyzing other genetic features may be performed for the purposes of normalizing data associated with chromosome 13.
[0156] The present disclosure may relate to identifying a change in chromosome number, either an increase or decrease, in cell-free DNA from a subject. In some embodiments, identifying a change in chromosome number is performed by querying non-polymorphic regions in one or more chromosomes. In some embodiments, the chromosomes of interest for querying non- polymorphic regions include, but are not limited to chromosome 13, chromosome 18, chromosome 21, chromosome 22, chromosome X, or chromosome Y. In some embodiments, all chromosomes are potential targets, for example target region(s) may be any non-polymorphic region in one or more of chromosomes 1 -24 which include 22 somatic chromosomes and 2 sex chromosomes. A human subject typically has 23 pairs of chromosomes, such that each pair is a pair of 2 chromosomes. In some embodiments, sub-chromosomal copy number variants may be detected. However, in some diseases a chromosome may have a pair plus 1 for 3 chromosomes, or a pair minus 1, or 1 chromosome. The present disclosure provides methods for identifying whether a subject has such aneuploidies wherein more than one or less than one of a chromosome pair are present in cell-free DNA, further in a mother and further in a fetus.
[0157] In some embodiments, a chromosome is queried with one or more recognition elements at conserved regions or non-polymorphic sequences located along the whole length, or a representative portion, of the chromosome. While not to scale, Fig. 8 represents an example of this query scenario, such that the longer (e.g., chromosome 13) is queried with recognition elements at more non-polymorphic sequences than chromosome 18 and chromosome 21, given the different lengths of the three chromosomes. In some embodiments, the number of recognition elements utilized to query non-polymorphic sequences is the same for each chromosome regardless of the length. In some embodiments, the number of recognition elements utilized in an assay to query one or more chromosomal conserved sequences is approximately 10-5,000 recognition elements, approximately 10-2,000 recognition elements, approximately, approximately 20-4,000 recognition elements, approximately, 30-3,000 recognition elements, approximately 200-1,000 recognition elements, approximately 30-700 recognition elements, approximately, 40-500 recognition elements, approximately 20-300 recognition elements, or approximately 50-500 recognition elements. In some embodiments, the recognition elements are spaced evenly across a chromosome. In some embodiments, the recognition elements are not spaced evenly across a chromosome. In some embodiments, therecognition elements are spaced at a higher density at one chromosomal location relative to another chromosomal location on the same chromosome. In some embodiments, there may be a distribution of recognition elements at one location on a chromosome but no recognition elements at another location on the sample chromosome. For example, DiGeorge syndrome comprises a microdeletion at 22ql 1 .2 as such recognition elements targeting this sub - chromosomal deletion may be more concentrated in that area of the chromosome with no or few recognition elements targeting other areas of the chromosome. The distribution of recognition elements on a chromosome can be different for each chromosome, or the same for each chromosome. As each chromosomal sequence is different, recognition element target spacing along each chromosome may be different as well. One goal for the determination of how many recognition elements might be used in an assay depends on the specificity and sensitivity of the assay for each chromosome. As such, there exists a sweet spot which can be empirically determined for how many recognition elements are needed, and where, to identify a chromosomal anomaly accurately and reproducibly in a cell-free DNA sample.
[0158] In some embodiments, a chromosome is queried with a plurality of recognition elements which each comprises different target 5’ and 3’ end regions but comprise the same identifying codes. For example, Fig. 9 shows an example of a chromosome with four groups of recognition elements wherein each group comprises a different code, but the code is the same within that group. However, each recognition element within the group may have different 5’ and 3 ’ end target sequences which collectively can be used to identify the chromosome and any existing chromosomal anomaly. In this example, recognition elements may have the same unique code (e.g., Code #1, Code #2, Code #3, Code #4) but the 5’ and 3 ’ regions of each recognition element may differ in their target sequences on the chromosome. As such, in this scenario the code may not be unique for each recognition element, but their detection and decoding in combination with other recognition elements with the same code collectively identify the presence or absence of a target sequence as the code is still associated with the 5 ’ and 3 ’ targeted regions of the recognition element. In some embodiments, each chromosome may be queried in any of the aforementioned scenarios, including the scenario where a subset of the recognition elements includes unique codes each in combination with a subset of recognition elements where two or more recognition elements comprise the same code, or any combination in between.
[0159] The target sequences in cell-free DNA may be in low abundance, such that additional upstream events can be applied to the cell-free DNA to boost abundance of low abundant target sequences. The fetal fraction of cell-free DNA in the plasma of a pregnant woman may be around 2-20% between 10 and 20 weeks of pregnancy. As such, the fetal fraction of cell-freeDNA in plasma of a pregnant female may be a minor proportion of the total cell-free DNA. Different factors can affect the amount of fetal DNA present in a plasma sample. For example, race, maternal age, mother body weight, physical activity, smoking, fetus gender, gestational age, levels of low-density lipoprotein (LDL), metformin, triglycerides, hemoglobinopathies, multiple fetuses to name a few have been found to have either positive or negative effects on the amount of fetal DNA present in a plasma sample. As such, in some instances, the methods may include enriching for fetal and / or maternal cell-free DNA from plasma prior to practicing the workflows as provided herein. As such, in some embodiments, maternal or fetal DNA found in a plasma sample are either pre amplified, enriched for, or both, from the plasma sample prior to downstream hybridization and detection for the presence of target molecules of interest from a cell-free DNA sample.
[0160] In some embodiments, target sequences of cell-free DNA are amplified, for example by PCR, whole genome amplification or linear amplification from the cell-free DNA, wherein the plurality of amplicons generated by PCR or linear amplification then serve as the target sequences for hybridizing to the 5’ and 3’ end regions of one or more recognition elements. In some embodiments, maternal and orfetai DNA from a cell-free DNA sample are enriched away from other DNA present in a cell-free DNA sample, for example contaminating DNA contributed from the buffy coat, and optionally amplified post -enrichment.
[0161] In some embodiments, the cell-free DNA may be extracted from plasma using SPRIselect beads as discussed herein. When the cell-free DNA is extracted from a plasma sample and / or amplified and / or enriched, the resulting DNA can be assayed for the presence of chromosomal anomalies that might be present in the maternal and / or fetal DNA.
[0162] In some embodiments, the cell-free DNA for use in an assay as described herein may be initially denatured such that the double stranded DNA is rendered single stranded. For example, heating most DNA to at least above 75 °C is sufficient to break the hydrogen bonds in the double helix of DNA to separate the two strands (e.g., denaturation of the double stranded target molecules). In some embodiments, linear single stranded recognition elements comprising a code and 5’ and 3 ’ ends complementary to targets in the maternal and / or fetal DNA are able to hybridize to their target sequences, thereby generating a target sequence / recognition element hybridized complex. In some embodiments, for example when a target sequence is present and has hybridized to the recognition element complementary sequences, the recognition element ends can be ligated to form a circularized and ligated recognition elements that comprises complementary sequences that are identified by the code in the recognition element. If there is no target hybridization to a recognition element, that recognition element may remain in its linear state (as compared, for example, to a padlock probe configuration) and ligation may notbe possible. In some embodiments, the hybridization and ligation events, including a denaturation event prior to each hybridization, may be performed for two or more cycles, three or more cycles, four or more cycles, five or more cycles, six or more cycles, seven or more cycles, eight or more cycles, nine or more cycles, or ten or more cycles.
[0163] In some embodiments, after hybridization and ligation (or a number of hybridization and ligation cycles) the reaction may be exonuclease treated to remove linear DNA, including cell- free DNA target sequences and linear recognition elements. In some embodiments, the exonuclease treated samples may be exposed to a polymerase and the recognition elements may be amplified. In some embodiments, the amplification is rolling circle amplification. In some embodiments, the amplification is multiple strand displacement amplification. In some embodiments, the amplification is linear amplification. In some embodiments, the amplification is polymerase chain reaction (PCR). Amplification of the recognition elements may provide for generation of multiple copies of a recognition element, including its code and complementary target sequences, which may serve as a proxy of the originally hybridized target sequences.
[0002] In some embodiments, the ligated recognition elements may be amplified on a substrate. For example, the recognition elements may be added to the wells of a plate and amplification may be performed on the surface of the well. In some embodiments, the wells of a plate can be coated with a composition that facilitates immobilization of DNA, thereby capturing recognition elements on the surface of a well, followed by amplification on the surface of the well. In some embodiments, where amplification is performed in a well, the ligated recognition elements may be allowed to pre-bind to the surface of the well before amplification reagents are added and amplification is performed. In some embodiments, amplification is performed in solution and the amplification products are added to wells which can be precoated with a composition that immobilizes DNA. In some embodiments, amplification is performed in solution and remains in solution for downstream detection of the recognition elements. In some embodiments, washing of the reactions to remove excess or remaining reagents occurs after hybridization of the target sequences to their complementary recognition elements, and / or after ligation of the recognition elements and / or after exonuclease treatment and / or after capture (e.g., capture on a substrate) of the recognition elements.
[0164] In some embodiments, there may not be washing after different stages in the workflow prior to detection of the recognition element amplification products. In some embodiments, compositions can be added to one or more operations of the workflow for enhancing hybridization of the nucleic acids, immobilization of nucleic acids, and the like such as betaine, BSA, propanediol, PEG, DMSO, polylysine, and molecular crowders, for example.
[0165] In some embodiments, amplification may be performed by a DNA polymerase or an RNA polymerase. In some embodiments, the DNA polymerase generates concatenated amplification products wherein each concatenated amplification product comprises multiple copies of a recognition element including the proxy sequences of the targeted nucleic acids and the code associated with the recognition element and thus the targeted nucleic acids. In some embodiments, the concatenated amplification products are detected by hybridizing one or more labeled detection polynucleotides followed by imaging of the hybridization events. In some embodiments, the detection polynucleotides are fluorescently labelled with one or more fluorophores. In some embodiments, the images are further decoded using soft decision decoding as described herein.
[0166] In some embodiments, detecting the presence of chromosomal copy number variants may be the output of the assay. In some embodiments, detecting the presence of SNPs is the output of the assay. In some embodiments, an assay measures both the presence of chromosomal copy number variants and the presence of SNPs and using both measurements outputs results related to the presence of chromosomal copy number variants for a mother or a fetus, and the estimated fetal fraction of the original cell-free DNA sample used as the input sample into the assay. In some embodiments, ligation cycles may differ depending on the detection focus, for example whether the data is predictive of chromosomal copy number variants or whether the data determines the presence of SNPs in the cell-free DNA. In some embodiments, a sample is first exposed to reagents that determine the presence of chromosomal copy number variants and then exposed, in a step wise manner, to reagents that determine the presence of SNPs.
[0167] In some embodiments, controls may be added to an assay in addition to cell-free DNA test samples. Controls can be added for a number of reasons, for example a positive control well with known chromosomal copy number variant and / or SNP composition can be run alongside the test samples and serve as an indicator of whether the reagents in the assay are performing as expected. Additionally, a negative control with no sample can be run concurrently with the cell- free DNA test samples to determine whether any nucleic acid contamination might be present as a no DNA sample may serve and provide negative results. In some embodiments, an internal control can also be included and run concurrently in the test cell-free DNA sample wells. An internal control could be used to monitor the assay conditions in each well, the environment in which the assays are performed, and the like, thereby controlling for those conditions. Positive controls, negative controls, and internal controls can help with identifying issues that might arise during an assay, as such their incorporation can provide insights into an assay that does not perform as expected or does not perform at all. In some embodiments, controls may also be used as calibrators. As controls may provide more qualitative or categorical information such as assaypass / fail or estimating the confidence of an assay, calibrators may be used for, for example, quantitatively correcting or altering raw data that is generated from an assay to improve accuracy of the assay as assessed by one or more external validation methods.
[0168] The methods described herein may relate to the detection of the sex of the fetus, or chromosomal aneuploidy associated with chromosome X and / or Y. In some embodiments, the determination of the sex of the fetus may comprise detecting sequences associated with chromosome X and / or chromosome Y. In some embodiments, one or more loci within chromosome X or chromosome Y may be detected. In some embodiments, detection of the one or more loci within chromosome X or chromosome Y may be preceded by selective amplification of the loci. In some embodiments, the detection of the loci may be performed by any one of or combination of methods described herein. In some embodiments, additional genetic features may be analyzed in combination with analysis of chromosome X loci or chromosome Y loci. In some embodiments, the additional genetical features may be other chromosome sequences. In some embodiments, analyzing other genetic features may be performed for the purposes of normalizing data associating with chromosome X or chromosome Y.
[0169] In some cases, the methods described herein may relate to the estimation of fetal fraction using single nucleotide polymorphism detection at SNP targets of interest in maternal and fetal cell-free DNA, or amplicons thereof. In some embodiments, the estimation of fetal fraction includes determining the presence of chromosome Y and the copy number status of chromosome Y. In some embodiments, the detection of fetal fraction single nucleotide polymorphisms may relate to detection of other genetic risk factors. In some embodiments, the detection of fetal fraction single nucleotide polymorphisms may be used to estimate the fetal DNA fraction of a cell-free DNA sample, comparative to the maternal DNA fraction of a cell-free DNA sample. In some embodiments, both SNP and CNV detection are utilized to estimate fetal fraction of a cell- free DNA sample. In some embodiments, cfDNA fetal fraction may be detected using methylation pattern targets that may be present in methylated fetal tissues while constitutively unmethylated in maternal tissues.
[0170] In some embodiments, estimating fetal fractionby identifying SNPs of interest in a cell- free DNA sample may comprise the use of approximately 100-1,000 recognition elements, approximately 200-800 recognition elements, approximately 300-700 recognition elements, or between approximately 400-600 recognition elements. In some embodiments, each recognition element comprises 5 ’ and 3 ’ end regions that identify the presence of a SNP target nucleotide of interest in maternal orfetai cell-free DNA. Additionally, each recognition element that targets a SNP target nucleotide of interest further comprises a code that aligns to the target SNP ofinterest. In some embodiments, the informative SNPs in the SNP data such that they can be used to estimating fetal fraction in a cell-free DNA sample are used as training data for training an algorithm to estimate fetal fraction. In some embodiments, at least 150 SNP loci identified from a plurality of chromosomal locations are utilized as informative loci for estimating fetal fraction of a cell-free DNA sample. Fig. 10 demonstrates an example of a single nucleotide polymorphism (SNP) detection scenario for two SNP loci that are queried to estimate fetal fraction. In some embodiments, between 150 and 300 SNPs of high minor allele frequency across a plurality of chromosomes, for example across chromosome 1 through chromosome 12, wherein the queried SNPs are widely spaced with minimal to no near neighbor SNPs, are queried with recognition elements. Querying a plurality of high minor allele frequency SNPs spread across multiple chromosomes of maternal and fetal cell-free DNA increases the ability to detect differences between maternal and fetal cell-free DNA contributions, and thus estimation of fetal fraction. In some embodiments, the data generated for estimating fetal fraction is first normalized and corrected for any potential bias found in the assay as a factor of environment, temperature, recognition element sequence, and the like.
[0171] In some embodiments, the presence of chromosome Y copy number can be used separately or in combination with SNP loci identification in estimating fetal fraction from cell- free DNA. In some embodiments, a plurality of recognition elements targeting a plurality of conserved regions with a plurality of chromosomes in cell-free DNA are queried. In some embodiments, a plurality of recognition elements targets at least conserved regions in chromosomes 13, 18, 21, 22, X and Y. In some embodiments, the recognition elements are unique for one target and include a unique code per recognition element as shown in Fig. 8. In some embodiments, the recognition elements are unique for one target, but a group of recognition elements include the same code, as shown in Fig. 9.
[0172] In some embodiments, cell-free DNA may not be quantified prior to practicing the chromosomal copy number determination methods described herein. In some embodiments, cell-free DNA may be quantified prior to practicing chromosomal copy number determination methods as described herein. In some embodiments, when a cell-free DNA sample is quantified, approximately 5 nanograms (ng), approximately 10 ng, approximately 15 ng, approximately 20 ng, approximately 25 ng, approximately 30 ng of cell-free DNA may be assayed. In some embodiments, after hybridization of the target sequences from the cell-free DNA sample to their complementary regions in one or more recognition elements, at least one ligation event is performed. In some embodiments, more than one hybridization and ligation event may be performed. In some embodiments, at least two hybridization / ligation cycles are performed, at least three hybridization / ligation cycles are performed, at least four hybridization / ligation cyclesare performed, at least five hybridization / ligation cycles are performed, at least six hybridization / ligation cycles are performed, at least seven hybridization / ligation cycles are performed, at least eight hybridization / ligation cycles are performed, at least nine hybridization / ligation cycles are performed, or at least ten hybridization / ligation cycles are performed. In some embodiments, increasing the number of hybridization and ligation cycles may increase the number of ligated recognition elements and therefore the number of concatemeric amplification products that can be detected and decoded for estimating fetal fraction of a cell-free DNA sample and / or chromosomal copy number variants in a cell-free DNA sample.
[0173] In some embodiments, the chromosomes targeted in the methods described herein include one or more of chromosome 13, chromosome 18, chromosome 21, chromosome 22, chromosome X, and chromosome Y. In some embodiments, for each chromosome a plurality of at least 500, at least 600, at least 700, at least 800, at least 900, at least 1 ,000, at least 1,500, at least 2,000, or at least 3,000 recognition elements target conserved sequences on each chromosome in a cell-free DNA sample. In some embodiments, a plurality of the total number of recognition elements comprises the same code, while maintaining the unique 5’ and 3 ’ ends of each recognition element. For example, recognition elements, each of which identifies a unique target sequence, may be binned with regards to the use of a single code, by region on a chromosome. For example, in some embodiments at least 1 ,000 recognition elements are designed to query 1 ,000 conserved regions on a chromosome. As a further example, for every 1,000 recognition elements that query a chromosome there are 100 unique codes, such that a region of 10 recognition elements share the same code. As such, the cumulative data can be used to enable analysis of copy number of a chromosome on a regional basis on that chromosome.
[0174] As such, in some embodiments, each conserved sequence of interest on a chromosome can be queried by a recognition element that may include a unique code that aligns to that conserved sequence. As such, each conserved target sequence may be determined, and high resolution may be achieved. Alternatively, in some embodiments each recognition element may be unique in targeting a chromosomal sequence of interest. However, all of the recognition elements may share the same code. In this scenario, regional information and resolution for each individual target sequence may be lost. In a third embodiment, in some embodiments, each recognition element may be unique in targeting a chromosomal sequence of interest. However, a certain number of the recognition elements in a defined region may share a code. In this third embodiment, finer resolution for deletions such as partial deletions may be realized and regional information aboutthe chromosome may be maintained. This third scenario might be importantfor identifying microdeletion anomalies such as DiGeorge syndrome that can be found querying chromosome 22.
[0175] Each recognition element may have distinct performance characteristics that can be collated in the recognition element’s profile. A recognition element profile may include: a background count of the concatenated amplification products that are not derived from real DNA, expected variant allele frequencies for homozygous reference or wildtype (e.g., homref), expected variant allele frequencies for homozygous alternate (e.g., homalt), expected variant allele frequencies for heterozygous genotypes (e.g., het), a concatenated amplification product count for each genotype (not considered background), or any combination thereof. Recognition element profiles may be useful in NIPT for estimating fetal fraction of a cell-free DNA sample. Recognition element profiles may be useful for determining the copy number variation (CNV) of a chromosome of interest. Fig. 15 shows an example of a pipeline for profiling recognition elements.
[0176] As shown in Fig. 15, a sample cohort with known genotypes are queried with designed recognition elements, and the results are collated for iterative learning to determine error profiles for each designed recognition element. Still referring to Fig. 15, the error profiles for a given recognition element are identified and applied as a performance metric to correct data for a given test sample, or as in the case of fetal fraction, error profile of a given recognition element used to query for a SNP may be utilized for adjusting the variant allele frequency of the SNP.
[0177] The following examples illustrate embodiments of the present disclosure in detail. It is to be understood that this present disclosure is not limited to the particular embodiments described herein and as such can vary. Those of skill in the art will recognize that there are numerous variations and modifications of this present disclosure, which are encompassed within its scope. Although various features of the present disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, although the present disclosure may be described herein in the context of separate embodiments for clarity, the present disclosure may also be implemented in a single embodiment.EXAMPLESExample 1 - Cell-free DNA target detection general workflow
[0178] This example describes several non-limiting general workflows that can be used for identifying the presence of a target sequence of interest in a cell-free DNA sample.
[0179] A blood sample is taken from a pregnant subject. For example, venousblood is collected in one or more blood collection tubes, such as cell-free blood collection tubes from Streck that are intended for the collection, stabilization and transport of venous whole blood samples for use with cell-free DNA assays. Blood tubes are stored according to the manufacturer’s recommendations.
[0180] The tube of blood is spun down, for example a 10 mL tube is centrifuged at 1 ,000-2,000 relative centrifugal force (ref) for 10 min. to separate the plasma from the blood cells. About two to four milliliters of the separated plasma are transferred to a new tube, taking care to avoid contacting and transferring any cells from the buffy coat. The plasma is centrifuged at 2,000 ref for 15 min. to pellet platelets and the supernatant is transferred to a new tube which is stored at less than -20°C, or the supernatant can proceed directly to cell-free DNA extraction.
[0181] The cell-free DNA in the plasma is largely contained in extracellular vesicles such as exosomes, as such the exosomes are lysed. Lipid and proteins are also removed from the cell- free DNA. There are two potential options for extracting the cell-free DNA from other plasma components.
[0182] The first option is a fully bead-based cell-free DNA extraction. Lysis reagents are added to the plasma to lyse the exosomes and the cell-free DNA is captured by magnetic particles, such as magnetic nanoparticles or beads. The magnetic particles to which the cell-free DNA is captured are themselves magnetically captured and pelleted. The supernatant is removed, the magnetic bead pellet washed and optionally resuspended, and the magnetic particles captured or recaptured (if resuspended). Washing can be repeated. After the final wash, the supernatant is removed and the pelleted beads with the captured cell-free DNA are allowed to dry for approximately 3 -5 min. The beads are removed from the magnet and the cell-free DNA is eluted from the beads by adding approximately 20-100 pL elution solution. The solutionis incubated at room temperature for approximately 1 -3 min. before the tube is applied to a magnet to pellet the magnetic beads. The supernatant with the cell-free DNA is collected and transferred to a new tube for storage, size selection and / or analysis.
[0183] The second option includes both bead and column-based cell-free DNA extraction. The exosomes are lysed, incubated, and magnetic particles added as in the first option. The magnetic beads with the captured cell-free DNA magnetically captured to form a pellet. The supernatant is removed from the pelleted beads and transferred to a binding column. The binding column is spun in a microcentrifuge for approximately 3 -5 min. at a maximum speed of around 20,000 ref to enable the cell-free DNA to bind to the column. The column flowthrough is discarded, the column washed, and the column centrifuged as before. More than one wash and centrifuge cycle can be performed. Following the final wash and spin, the column is allowed to dry briefly, forexample 3-5 min. and added to an elution tube, elution buffer added, and the tube is centrifuged to collect the eluted cell-free DNA from the column.
[0184] Regardless of whether the first or second option is followed, the cell-free DNA can be exposed to a size selection operation. By performing size selection for maternal and fetal cell- free DNA, negative consequences of hemolysis or operator error, for example when removing plasma from above the buffy coat, which could lead to contamination of the cell-free DNA sample with longer contaminating genomic maternal DNA can be mitigated. For example, a SPRIbead cleanup operation can remove longer than approximately 500 bp DNA fragments from the smaller matemal / fetal cell-free DNA. The cell-free DNA sample that has been size selected to remove contaminating genomic maternal DNA can then be used for analysis in the methods described herein.
[0185] To 5-100 nanograms (ng) of cell-free DNA sample is added to a pool of targeted and coded recognition elements. The 5 ’ and 3 ’ end region s of each recognition element is targeted to specific regions of the cell-free DNA, to which they are complementary to such that hybridization between the recognition element and the cell-free DNA target can occur. For example, a 5’ and 3 ’ end of one recognition element may be complementary to a non- polymorphic site on a chromosome of interest, such as chromosome 13, 18 or 21. A plurality of recognition elements at a concentration of approximately 0.075 nM of each targeted recognition element are added to the cell-free DNA sample, each one specific to a non-polymorphic sequence of interest in a chromosome. Targeted chromosomes can include chromosome 13, 18, 21, 22 and X, or any other chromosome of interest. Additionally, a plurality of known SNPs in the cell-free DNA are targeted by a plurality of recognition elements, the targeted SNP recognition elements at a concentration of approximately 0.5 nM. The SNP targeted recognition elements are designed to hybridize to any one of approximately 150 nucleic acid variants in fetal cell-free DNA. Figs. 8 and 9 show targeted recognition elements specific to non-polymorphic regions in chromosomes and SNP-related recognition element targets that can be used to estimate the amount of fetal fraction in the cell-free DNA sample relative to the maternal DNA contribution.
[0186] To the reaction mixture comprising a cell-free DNA sample aliquot and the targeted and coded recognition elements that is located in a reaction chamber such as a tube, well of a plate, or a flowcell, are added hybridization buffers, ligase and ligase buffers such that concurrent hybridization and ligation can occur. In this case, samples were added to wells of a 96 well plate and reactions were carried out in the wells of the 96 well plate. For example, to the reaction mixture in a well of a 96 well plate is added a buffer that facilitates both hybridization and ligation and a thermostable ligase that includes high specificity for ligating nicks in DNA, suchas AmpLigase, to a total reaction volume of 20 pL. For using AmpLigase, a buffer including KC1 may be used. Following addition of the hybridization and ligation buffers and enzymes, hybridization and ligation can occur. For example, the reaction mixture is heated to 95 °C for 1 min to denature the double stranded DNA, followed by 20 min at 60°C (e.g., in a thermocycler) to hybridize the targets and ligate the 5’ and 3’ ends of the hybridized recognition elements. The heating and annealing / ligating temperature reaction can be cycled, for example from 3 -24 cycles to maximize the hybridization and ligation events. Following hybridization and ligation, the reaction mixture can be stored at 4°C before proceeding with exonuclease digestion, or for longer storage placed at -20°C.
[0187] To remove any linear cell-free DNA or linear targeted and coded recognition elements from the reaction mixture after hybridization and ligation, an exonuclease and associated buffer can be added to the reaction mixtures in the wells. For example, a mixture of Exonuclease I and Exonuclease III, and an associated buffer, can be added to each reaction mixture for a reaction volume of 30 pL, followedby incubation at 37°C for 30 min. and enzyme inactivation at 95°C for 5 min. The exonuclease treated samples can be stored at 4°C for further processing or stored at -20°C overnight if needed. After exonuclease treatment for removing any unhybridized and unligated DNA fragments and recognition elements, the circularized and ligated recognition elements are amplified.
[0188] With an optically clear bottom, a 96 well plate that had been previously treated to help in immobilizing the recognition elements to the bottom of the wells, a 50 pL aliquot of the exonuclease treated reactions is transferred, a well for each reaction. The plate with the reaction mixtures is sealed and incubated at42°C for 1 hour, followed by the removal of all liquid from the wells. Amplification of the circularized and ligated recognition elements in the wells is performed by the addition of an amplification reaction mixture comprising 10 nM of an amplification primer, 0.15 U / pL of an EquiPhi polymerase, EquiPhi buffer, 1 mM of each dNTP (e.g., dATP, dTTP, dCTP, dGTP), and 1 mMDTT. The plate is re-sealed and incubated at 42°C for 2 hours, after which time the sample wells are washed and what remains are concatenated amplification products comprising multiple copies of the ligated recognition elements, as such multiple copies of representative target sequences of interest and multiple copies of the codes that uniquely identify the representative target sequences of interest.
[0189] For detecting the codes, and hence the presence or absence of the target of interest, a 50 pL master mix including detection polynucleotide components, hybridization buffer, blocking DNA and water is added to each well and the detection reaction is incubated for 15 min. at room temperature. Post incubation, the wells are washed and TE or TETS buffer is left in the wells for signal detection. The RAPTOR instrument is used to detect fluorescence in each well in theappropriate channel. The addition of the master mix and fluorescence detection is cycled a number of times, usingthe same or different detection polynucleotides targeting different code sequences, such that after a final cycle the fluorescence patterns from the different detection events are collated and decoded by the RAPTOR to identify the probability of a particular code being present in a well, which in turn is used as a proxy for the presence or a target of interest. The number of cycles of hybridization of detection polynucleotides and detection of the signal is determined by the complexity of the code and the number of different targets that are in need of decoding.Example 2 - Determining fetal fraction from a mother / son DNA mixture
[0190] In this proof -of-concept experiment, three mixtures of mother / son DNA samples were used to estimate the fetal fraction of the mixtures and correlate the SNP fetal fraction estimates to the presence of copies of the Y chromosome.
[0191] Genomic DNA samples from Covaris were sheared and used to generate sample mixtures of mother / son DNA with different ethnicities to mimic cell-free DNA samples extracted from a maternal blood sample, as outlined in Table 1.Table 1-Cell line ethnicities
[0192] Each of the three DNA samples further included a series of mother: son DNA concentration titrations, where son DNA was spiked into mother DNA as follows; 100% mother DNA:0% son DNA, 98% mother DNA:2% son DNA, 95% mother DNA:5% son DNA, 90% mother DNA: 10% son DNA, 80% mother DNA:20% son DNA, and 0% mother DNA: 100% son DNA.
[0193] The samples were queried for aneuploidy targeting non -polymorphic regions and for estimating fetal fraction (e.g., son contribution) as described in Example 1 herein and shown in Figs. 8 and 9. For copy number of chromosome Y, approximately 5,700 recognition elements spread across non -polymorphic regions across chromosome 13, 18, 21, 22, X and Ychromosomes were utilized with approximately 10 loci binned per unique code with approximately 100 unique codes per chromosome. Fig. 9 shows an example of the query strategy using four unique codes on a stretch of an example chromosome.
[0194] For SNP detection, approximately 300 recognition elements each with a unique code to identify the targeted SNP in the sample mixtures were queried to identify which SNPs could be used as informative loci in estimating the fetal fraction for estimating the fetal fraction.
[0195] The fetal fraction for the son was determined from the ratio of chromosome Y counts:total chromosome counts where the total chromosome counts include the sum counts of chromosomes 13, 18, 21, and 22. The estimate for son fetal fraction was generated by plotting the ratio of chromosome Y against the concentration of the son gDNA that was spiked into the maternal sample. In some instances, with low input gDNA concentrations such as when around 5-15ng of gDNA was used in the SNP workflow, subtracting signal to noise counts improved R2correlation (data not shown).
[0196] Figs. 11A-11B show examples of graphs illustrating fetal fraction estimation data from the dilution series for Sample 2. Fig. 11A shows an example of a graph illustrating example data when comparing the chromosome Y ratio to fetal fraction based on the son gDNA spike in concentrations. The results demonstrate that the chrY ratio can be used to identify the concentration of son gDNA that was spiked into maternal DNA in a correlated fashion for the dilution series generated. Similar results were seen with the dilution series for sample 1 and 3, with R2=0.995 for sample 1 and R2=0.998 for sample 3.
[0197] Fig. 1 IB shows an example of a graph illustrating an example data that SNP based fetal fraction estimation using the dilution series for sample 2 correlates with fetal fraction derived from the determined chromosome Y ratio for the sample 2 dilution series of 2-20% spiked in son gDNA from Fig. 11 A. Similar results were generated with the dilution series for sample 1 and 3, with R2=0.956 for sample 1 and R2=0.964 for sample 3.
[0198] The data demonstrate the proof of concept that the estimation of fetal fraction can be correlated with the determined ratio of chromosome Y.Example 3 - Chromosomal copy number variance determination from mother / son DNA mixture
[0199] In this experiment, the mock cell-free DNA samples from the Table 2 were used for identifying different chromosomal copy numbers.Table 2- Cell line and sample descriptions
[0200] The workflow was followed as described per Example 1. Chromosome ratio was determined by taking the total count of the chromosome of interest over the collective counts of the other chromosomes. Data was normalized.
[0201] Fig. 12 shows an example of graphs illustrating the results. As shown in Fig. 12, there is a high >95% correlation between the concentration of the spiked in gDNA son DNA and the identification and determination of Trisomy 21, Trisomy 18, Trisomy 13, Turner’s syndrome, Trisomy X and Klinefelter’s syndrome. The data suggest a clear separation of signal (e.g., chromosome ratio) between normal 0% DNA samples (normal) and 5% disease DNA spike in samples, suggesting that disease detection is specific and does not generate false-positives.Example 4 - Probe considerations for estimating fetal fraction by SNP detection in cell-free DNA
[0202] For estimating fetal fraction of a cell-free DNA sample from a pregnant woman, a set of SNPs that are informative for estimating the fetal fraction are needed. Two approaches can be implemented, either alone or together, that can help identify SNPs that can be queried that serve as informative SNPs.
[0203] A first approach can be to compute a likelihood ratio (LR) for each SNP and to use the likelihood ratios to identify the number and location of a plurality of SNPs that can be used as informative SNPs in determining fetal fraction. The likelihood ratio (LR) can be calculated as follows.
[0204] For a plurality of loci, the fetal fraction that maximizes the likelihood ratio is determined. Given a likelihood ratio threshold, loci exceeding this threshold can be defined as informativeloci. The median fetal fraction from the inferred informative loci can then be calculated to provide an estimate.
[0205] A second approach can be the expectation maximization approach. In this approach, assuming that there are several possible distribution states, the expectation maximization can be used to estimate which distribution each of a set of data points belongs to as well as the allele frequency of each distribution state. The relationship between the distribution states is bounded by the fetal fraction.
[0206] Figs. 13A-13B show examples of recognition element single nucleotide polymorphism (SNP) identification prior to and after recognition elements were corrected for performance issues. Fig. 13A shows SNP detection prior to correcting for the performance of recognition elements directed to SNP targeted sequences, whereas Fig. 13B shows performance after correcting the recognition elements directed to SNP targeted sequences. The corrected recognition elements can be used at this point for estimating fetal fraction.
[0207] The data from Figs. 13A-13B was generated using recognition elements that targeted known SNP genotypes samples, wherein the data was used in a bioinformatics training module. Representative data from all possible genotypes was reviewed to use in training the bioinformatics workflow for estimating fetal fraction. Genotypes of homozygous reference (0 / 0), heterozygous (0 / 1) and homozygous alternative (1 / 1) were utilized. An example of the genotypes used in the bioinformatics training module is provided in Fig. 14.
[0208] High minor allele frequency SNPs, initially 388 SNPs, from all populations were chosen for further consideration as informative SNPs. The SNPs were chosen such that there were no nearby interfering variants (as identified in dbSNP). As well, the SNPs were chosen that were widely spaced to rule out or minimize minimal haplotype linkage.
[0209] Roughly 2,552 individual DNA samples representing 327 SNPs from the 1 ,000 Genomes project were reviewed and aligned with the Coriell DNA catalog, resulting in a truth set of 2 ,533 individual DNA samples representing 327 SNPs. The truth set was evaluated, and it was determined that the truth set did not cluster by genotype, the SNPs were highly variable, the SNPs were not closely linked, and the samples came from a diverse population of individuals. The truth set was further analyzed to identify which samples from the set could be used that may cover all of the genotypes with the fewest samples, the smallest set found was 15 samples.
[0210] The 15 samples were ordered, sheared, quantified and normalized. Recognition elements were designed to hybridize to the test sample sequences and assays were optimized for, but not limited to, amount of shearedDNA for optimal ligation, optimal hybridization between the test samples and the designed recognition elements, off target hybridization, and the like. Screening the reaction conditions, and after further optimizations, it was determined that out of the original327 SNPs, 312 SNP targets were good in identifying HomRef truth, 315 SNP targets were good in identifying HomAlt truth, and 254 SNP targets were good in identifying Het truth.
[0211] The best 150 SNP targets were down selected for use from the main set of SNPs by additional experimentation with newly designed recognition elements. From the data, outlying SNPs were removed from the pool. SNPs with highly variable data were removed. SNPs that demonstrated non-conformity, outside the Het variable frequencies of 0.39-0.61 were removed. The additional screening results at 166 loci, of which 150 loci that were the closest to the variable frequency of 0.50 were chosen and became the bioinformatics teaching sample pool.
[0212] The final 150 loci sample pool was evaluated anew with recognition elements through detection and decoding. It was found that the set of 150 SNPs at 150 loci provided for better depth at the majority of the loci, thereby increasing the depth for most of the targets.
[0213] As such, approximately 150 SNPs were selected across chromosomes 1 to 12 for estimating the fetal fraction of a cell-free DNA sample. The SNPs were chosen for their high polymorphism across populations and lack of homology, adjacent SNPs and CNVs.
[0003] Results generated from the assays comprising cell-free DNA and recognition elements targeting the 150 SNPs were decoded, with each SNP being targeted by two recognition elements, each with a unique code; one recognition element targeting the wildtype nucleotide (e.g., reference) and the second recognition element targeting the variant nucleotide (e.g., alternate). Fig. 16 shows an example of an analysis pipeline for estimating fetal fraction in a cell-free DNA sample. The variant allele frequency of each SNP may be adjusted to account for performance differences in the recognition elements as previously described herein (Fig. 15), with the adjusted variant allele frequency being used to estimate the fetal fraction. The presence of fetal DNA is expected to cause a shift in the allele frequencies from expected values. An Expectation Maximization algorithm maybe used to assign each SNP to a state, either: 1) both maternal and fetal alleles are homozygous references, or 2) the maternal allele is a homozygous reference and the fetal allele is a heterozygous reference. The shifts between the states may be constrained by the fetal fraction. After several iterations, the state assignments may stabilize, allowing for estimation of the fetal fraction of the cell-free DNA sample.Example 5-Determining chromosomal aneuploidy
[0214] Conserved regions were selected on chromosomes with potential aneuploidy, for example chromosomes 13, 18, 21, 22, X and Y. The selected conserved regions were non- polymorphic across the population of samples and free of homology and SNPs. Each region had a specific recognition element designed against it, with several adjacent regions potentially sharing the same code among the adjacently designed recognition elements.
[0215] Results of an assay for determining CNV of one or more chromosomes are decoded, with each code corresponding to one or multiple adjacent regions. It is contemplated that if a chromosome has extra copies, then the concatemeric amplification product counts for that chromosome may increase. Given the potential biases that can be introduced during various operations of the assay (e.g., sample processing, enzymatic operations, imaging, etc.) variability in concatemeric amplification product counts can occur, as such aneuploidy detection aims to identify the true signal amidst the background noise.
[0216] Some regions are biased toward higher or lower coverage. One source of such bias is GC content. Within a sample, or intra-sample, correction involves applying GC correction to the codes to mitigate bias caused by differences in GC content among targeted genomic regions.
[0217] Other sources of bias can be mitigated so long as the bias is systemic. For example, if a bin has elevated counts relative to other bins reproducibly across all samples, even if there may be no mechanistic explanation for the elevated counts, it can be included in the analysis after a normalization operation. A normalization operation may include, for example, removing the signal from the first several principal components or calculating the median number of counts for the bin across a large sample cohort and subtracting that average from the bin for all samples.
[0218] To correct for unknown biases, a sample cohort can be used . Two strategies can be employed to correct for unknown biases between, or inter-samples: 1) Principal Component Analysis (PCA) or Singular Value Decomposition (SVD), and 2) Median Polishing. Fig. 17 shows an example of an analysis pipeline for aneuploidy determination including bias correction. PCA or SVD uses a cohort of normal samples, wherein PCA or SVD is conducted to derive the principal components that explain intrinsic variations in the sample cohort. For a test sample, a regression model is applied against the first several principal components to remove the unknown bias. Median polishing calculates the median number of concatenated amplification products for each code across the sample cohort and iteratively subtracts this value from all the codes from all the samples. This reduces sample-to-sample and code-to-code biases.
[0219] After bias normalization, variability across the samples can be quantified for each code. A weighted strategy may be used to consider variability, with bins showing higher variability receiving lower weights in the final chromosome summary, and bins with lower variability receiving higher weights. For each chromosome, a summary statistic such as the median or sum of corrected counts across the codes can be calculated. To test for aneuploidy, a z-score can be computed against the normal sample cohort and the thresholds and quality control measures, based on the z-scores, can be used to identify aneuploidy. The p-value can also be combinedwith other information (e.g., fetal fraction estimation from SNPs, mother’s age, etc.) to derive a statistical measurement of aneuploidy risk.Example 6- Detection of aneuploidies and microdeletions in a 4% fetal fraction scenario
[0220] DNA samples (Coriell Life Sciences) were sheared to approximate cfDNA lengths (e.g., less than 200nt lengths of DNA), including NA12878 (female, normal), NA02948 (male, trisomy 13), NA03769 (male, trisomy 18), NG09802 (male, trisomy 21), NA00857 (female, X0), NA03102 (male, XXY), NA04626 (female, XXX), NA05876 (male, 22ql l .2 1.5 Mb deletion) and NA17942 (male, 22ql 1.2 3Mb deletion). The sheared samples were quantified, and each sample was diluted to 2 ng / pl in TE buffer. From the sheared and diluted samples, spike-in samples were generated to approximate a fetal fraction cfDNA contribution of 4%.
[0221] Recognition element probe pools were generated, with the recognition elements targeting either copy number variation sequences in the one or more samples or single nucleotide polymorphism sequence in the one or more samples. Approximately 15 ng of each sample was aliquoted into wells of a 96 well plate, and approximately 5 pl of a ligation master mix (Ampligase, Ampligase buffer, 0.3 nM each CNV probes, 0.4 nM each SNP probes) was added to each sample. Hybridization and ligation, approximately 24 cycles, was carried out; 95°C for 1 min., 60°C for 20 min., followed by storage of the reactions at -20°C if not proceeding with exonuclease digestion.
[0222] In order to remove any non-hybridized and non -ligated linear recognition elements and cfDNA, the hybridization / ligation reactions were exonuclease treated by adding 30 pl of an exonuclease master mix (Exonuclease I, Exonuclease III, exonuclease buffer) to each reaction, followed by incubation at 37°C for 30 min, 95°C for 5 min., and transfer of the exonuclease treated samples to a new 96 well plate.
[0223] The circularized and ligated recognition elements were immobilized onto the well surfaces of the 96 well plate by incubating for 1 hr at 42°C, and as much liquid as possible was removed from each well. To each well was added 50 pl of an amplification master mix (EquiPhi DNA polymerase, EquiPhi buffer, dNTPs, DTT, 10 pM of an amplification primer), and the amplification reaction was incubated at 42°C for 2 hrs, following by washing the wells with TE buffer and adding TE to the wells post washing. The plate was placed at 4°C until the detection part of the assay was performed.
[0224] After the final wash, the test plate was placed in the imaging instrument where code detection and decoding were performed. Detection was carried out for a number of cycles, for example six to ten cycles depending on the complexity of the code. Briefly, liquid was aspirated from the reaction wells and dehybridization buffer was added to each reaction well and thedehybridization reactions were incubated for approximately 1 min. The wells were washed, the wash fluid aspirated from the wells and the first pool of detection polynucleotides was added, and each detection polynucleotide pair comprising a first oligonucleotide with a complementary sequence to a code or portion of a code and a second oligonucleotide with a complementary sequence to the first oligonucleotide and a fluorescent moiety. Four fluorescent moieties were utilized for the detection polynucleotides, with wavelengths ranging from 556 nm to 709 nm, which are detectable in one of four channels in the fluorescent imaging instrument. The detection oligonucleotides were allowed by hybridize to their target sequences for approximately five to nine minutes and the wells were washed several times to remove any unhybridized detection oligonucleotides or unhybridized detection polynucleotides.
[0225] After the final wash, the wells were imaged and the workflow started anew for the second readout flow, wherein residual liquid from each well was aspirated, dehybridization buffer was added to the wells, and a second set of detection oligonucleotides was added, etc. After all the detection flows were completed, secondary analysis was carried out on instrument wherein the instrument identified profiles of fluorescence for each well and following the soft decision decoding secondary analysis pipeline as previously described, the codes were decoded and the relative abundance of each target of interest was determined.
[0226] All of the images were stored and the detection profile for each well and amplified recognition element decoded by running the soft decision decoding secondary analysis pipeline, thereby determining the abundance of a code which is a proxy for the presence of the original target.
[0227] Results are reported in Fig. 18, which demonstrate the ability of the methods described herein to successfully detect aneuploidies and microdeletions as identified in the test samples compared to the euploid female normal sample, at a spike-in for approximating 4% fetal fraction. The y-axis represents the fraction of the target chromosome detected in each of the samples relative to other chromosomes in the panel, which is referred to as the “Chromosomal Ratio”. This Chromosomal Ratio was used to determine whether or not a spike-in of an aneuploidy or microdeletion can be detected versus a normal sample.Example 7-Fetal fraction estimation and correlation with chrY quantitation
[0228] This experiment was performed to determine whether fetal fraction that is estimated from detecting single nucleotide polymorphisms in cfDNA correlates with fetal fraction estimated from the percent of chromosome Y (chrY) in a reaction. In this experiment, recognition elements for SNP and CNV targets were combined during the hybridization and ligation reactions.
[0229] To demonstrate accurate estimation of fetal fraction and correlation with chrY quantitation in an integrated full-content panel, DNAused for preparing spiked-in samples were obtained from Coriell Institute and included 0%-16% spike-in of male into female DNA for chrY determination using DNA NA19701 and NA19702 (mother and son), to make spike-in samples.
[0230] For the assay, 15 ng of a mixed DNA sample in 10 pL TE buffer was added to a well of a multi-well plate. To each sample was added 10 pl of a hybridization and ligation buffer master mix (Ampligase, Ampligase buffer, approximately 0.075 nM targeted CNV recognition elements, 0.05 nM targeted SNP recognition elements, water). Hybridization of the targets to the recognition elements and ligation of the recognition elements was performed for approximately 24 cycles; 95°C for 1 min., 60°C for 20 min. followed by a hold at 4°C.
[0231] Following hybridization and ligation, the assay reactions were exonuclease treated by addition of 30 pl of an exonuclease master mix (Exonuclease I, Exonuclease III, exonuclease buffer, water) and incubation at 37°C for 30 min., followed by exonuclease inactivation at 95 °C for 5 min. and storage at 4 °C.
[0232] The exonuclease treated reactions were transferred to wells of a new multi -well plate and the plate was sealed and incubated for 1 hr at 42 °C, allowing the circularized recognition elements to become immobilized onto the surface of the wells. Following incubation, the liquid was removed from each well and to each reaction well was added 50 pl of an amplification master mix (EquiPhi DNA polymerase, EquiPhi buffer, dNTPs, DTT, 10 nM of an amplification primer, water) and the amplification reactions were incubated at 42°C for 2 hrs. The reactions were washed with TE buffer prior to detection of the codes of the circularized and amplified recognition elements.
[0233] Results shown in Fig. 19 demonstrate that fetal fraction estimation using SNP detection shows tight correlation with the presence of chromosome Y using the spike -in samples as a surrogate for cfDNA extracted from the blood of a pregnant female. The calculated correlation was high with a Pearson coefficient of R2> 0.95.Example 8-Assay for T21 using cfDNA purified from a plasma-like matrix
[0234] Samples were obtained from SeraCare and Coriell Institute for assay. Samples included 1) SeraSeq® trisomy 21 cfDNA male matched reference sample in a plasma-like matrix, 2) SeraSeq® euploid cfDNA male matched reference sample in a plasma-like matrix, 3) SeraSeq® circulating tumor DNA extracted reference sample, 4) Coriell NA12878 sheared normal cell -line DNA, and 5) Coriell NA02767 sheared Trisomy 21 cell -line DNA.
[0235] Cell-free DNA from the SeraSeq® samples were extracted from the plasma-like matrix using the Apostle MiniMax™ cfDNA Isolation Kit. Briefly, to reaction tubes including proteinase K and Sample Lysis Buffer was added 1 ml of the SeraSeq® samples. The tubes were incubated at 60°C for 20 min. and cooled to room temperature. The cfDNA was captured by adding magnetic nanoparticles in a cfDNA Lysis / Binding Solution to the reaction tubes and shaking the tubes with beads for approximately 10 min. on a shaker followed by magnetic capture of the beads until the solution cleared and discarding of the supernatants. The magnetic beads / cfDNA remaining after taking off and discarding the supernatant were washed several times with cfDNA Wash Solution per manufacturer’s protocol. After washing, the cfDNA was eluted from the magnetic nanoparticles by adding cfDNA Elution Solution and magnetic capture of the magnetic nanoparticles until the solution cleared. The supernatant with the eluted cfDNA was transferred to new tubes, the cfDNA quantified and kept at 4°C or stored at -20°C for long term.
[0236] For the assay, 20 ng of each DNA sample was added to a well of a multi -well plate. To each sample was added 10 pl of a hybridization and ligation buffer master mix (Ampligase, Ampligase buffer, 75 pM targeted recognition elements). Hybridization of the targets to the recognition elements and ligation of the recognition elements was carried for several cycles; 95°C for 1 min., 60°C for 20 min. followed by a hold at 4°C.
[0237] Following hybridization and ligation, the assay reactions were exonuclease treated by addition of 30 pl of an exonuclease master mix (Exonuclease I, Exonuclease III, exonuclease buffer) and incubation at 37°C for 30 min., followed by exonuclease inactivation at 95 °C for 5 min. and storage at 4°C.
[0238] The exonuclease treated reactions were transferred to wells of a new multi -well plate and the plate was sealed and incubated for 1 hr at 42°C, allowing the circularized recognition elements to become immobilized onto the surface of the wells. Following incubation, the liquid was removed from each well and to each reaction well was added 50 pl of an amplification master mix (EquiPhi DNA polymerase, EquiPhi buffer, dNTPs, 10 pM of an amplification primer) and the amplification reactions were incubated at 42°C for 2 hrs. The reactions were washed with TE buffer prior to detection of the codes of the circularized and amplified recognition elements.
[0239] After the final wash, the test plate was placed in the imaging instrument (Pleno, Inc. RAPTOR imaging instrument) where code detection and decoding were performed. Detection was carried out for a number of cycles, for example six to ten cycles depending on the complexity of the code. Briefly, liquid was aspirated from the reaction wells and dehybridization buffer was added to each reaction well and the dehybridization reactions were incubated forapproximately 1 min. The wells were washed, the wash fluid aspirated from the wells and the first pool of detection polynucleotides was added, each detection polynucleotide pair comprising a first oligonucleotide with a complementary sequence to a code or portion of a code and a second oligonucleotide with a complementary sequence to the first oligonucleotide and a fluorescent moiety. Four fluorescent moieties were utilized for the detection polynucleotides, with wavelengths ranging from 556 nm to 709 nm, which are detectable in one of four channels in the fluorescent imaging instrument. The detection oligonucleotides were allowed by hybridize to their target sequences for approximately five to nine minutes and the wells were washed several times to remove any unhybridized detection oligonucleotides or unhybridized detection polynucleotides.
[0240] After the final wash, the wells were imaged and the workflow started anew for the second query flow, wherein residual liquid from each well was aspirated, dehybridization buffer was added to the wells, and a second set of detection oligonucleotides was added, etc. After all the detection flows were completed, secondary analysis was carried out on instrument wherein the instrument identified profiles of fluorescence for each well and following the soft decision decoding secondary analysis pipeline as previously described, the codes were decoded and the presence, or absence, of a target of interest was determined.
[0241] All of the images were stored and the detection profile for each well and amplified recognition element decoded by running the soft decision decoding secondary analysis pipeline, thereby determining the presence of a code which is a proxy for the presence of the original target.
[0242] Results demonstrate that the Apostle cfDNA extraction kit was successful in extracting cfDNA from plasma, as seen in Fig. 20. The y-axis represents the fraction of chromosome 21 detected in each of the samples relative to other chromosomes in the panel, which is referred to as the “Chromosomal Ratio”. This Chromosomal Ratio was used to determine whether or not a spike-in of Trisomy 21 can be detected versus a normal sample. Trisomy 21 was accurately detected in both the SeraSeq sample and the shearedDNA sample, with fetal fractions of 7.5% and 5%, respectively.
[0243] While certain examples of methods and systems have been shown and disclosed herein, one of skill in the art will realize that these are provided by way of example only and not intended to be limiting within the specification. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the scope disclosed herein. Furthermore, it shall be understood that all aspects of the disclosed methods and systems are not limited to the specific depictions, configurations or relative proportions set forth herein whichdepend upon a variety of conditions and variables and the description is intended to include such alternatives, modifications, variations or equivalents.
Claims
CLAIMSWhat is claimed is:
1. A method for identifying one or more chromosomal anomalies in cell-free DNA, comprising: a) providing a plurality of recognition elements, wherein a first recognition element of the plurality of recognition elements comprises a code, a 5’ region, and a 3 ’ region, wherein the 5’ region of the first recognition element comprises a sequence that is complementary to a first non- polymorphic chromosomal sequence of a first target chromosome in a cell-free DNA sample and wherein the 3 ’ region of the first recognition element comprises a sequence that is complementary to a second non-polymorphic chromosomal sequence of the first target chromosome in the cell-free DNA sample, and wherein the code of the first recognition element uniquely identifies a presence of the chromosomal sequences that are complementary to the sequence of the 5 ’ region and the sequence of the 3 ’ region of the first recognition element; b) hybridizing the first recognition element to the first non-polymorphic chromosomal sequence and the second non-polymorphic chromosomal sequence of the first target chromosome in the cell-free DNA sample to produce a hybridized first recognition element; c) ligating the 5’ region and the 3’ region of the hybridized first recognition element to a cell-free DNA in the cell-free DNA sample, thereby generating a circularized and ligated first recognition element; d) amplifying the circularized and ligated first recognition element to produce an amplified first recognition element; e) detecting the code of the amplified first recognition element; f) decoding the detected code of the amplified first recognition element; and g) concurrently performing b) through f) for a subsequent recognition element of the plurality of recognition elements, thereby identifying the one or more chromosomal anomalies in the cell-free DNA sample.
2. The method of claim 1, wherein the first non-polymorphic chromosomal sequence is adjacently located to the second non-polymorphic chromosomal sequence in the first target chromosome.
3. The method of claim 1, wherein the first non-polymorphic chromosomal sequence is not adjacently located to the second non-polymorphic chromosomal sequence in the first target chromosome, and wherein the method further comprises performing a gap-fill prior to ligating the 5’ region and the 3 ’ region of the hybridized first recognition element.
4. The method of claim 1, wherein the code comprises one or more nucleic acid segments.
5. The method of claim 4, wherein the one or more nucleic acid segments comprise two to ten nucleic acid segments.
6. The method of claim 5, further comprising hybridizing one or more detection polynucleotides to the amplified first recognition element and imaging the hybridizing the one or more detection polynucleotides to the amplified first recognition element, thereby detecting the two to ten nucleic acid segments of the code.
7. The method of claim 6, wherein each of the one or more detection polynucleotides comprise a fluorescent moiety.
8. The method of any one of claims 1-7, wherein the decoding the detected code comprises soft decision decoding.
9. The method of any one of claims 1-8, wherein the cell-free DNA sample is extracted from a blood sample.
10. The method of claim 9, wherein the blood sample is obtained or derived from a subject.
11. The method of claim 10, wherein the subject is a pregnant female.
12. The method of any one of claims 1-11, further comprising performing an exonuclease treatment after c).
13. The method of any one of claims 1-12, wherein the amplifying comprises polymerase chain reaction, primer extension, rolling circle amplification, or multiple strand displacement amplification.
14. The method of any one of claims 1-13, wherein the amplifying comprises generating concatemeric amplification products.
15. The method of any one of claims 1-14, wherein the one or more chromosomal anomalies comprises a copy number variation of a chromosome, an insertion or a deletion in a chromosome, or a sub-chromosomal deletion.
16. The method of any one of claims 1-15, wherein the one or more chromosomal anomalies is in one or more of chromosome 1, chromosome4, chromosome 5, chromosome 13, chromosome 16, chromosome 15, chromosome 18, chromosome 21, chromosome 22, chromosome X, or chromosome Y.
17. The method of any one of claims 1-16, further comprising estimating a fetal fraction of the cell-free DNA sample.
18. The method of claim 17, wherein the estimating the fetal fraction comprises identifying a plurality of single nucleotide polymorphisms in fetal DNA and in maternal DNA.
19. The method of claim 18, wherein the plurality of single nucleotide polymorphisms comprises 50 single nucleotide polymorphisms to 300 single nucleotide polymorphisms.
20. The method of claim 19, wherein the plurality of single nucleotide polymorphisms is located across a plurality of chromosomes in the cell-free DNA sample.
21. The method of claim 20, wherein the plurality of chromosomes comprises chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome 23, or chromosome 24 in the cell-free DNA sample.
22. The method of claim 19, further comprising identifying the plurality of single nucleotide polymorphisms and the one or more chromosomal anomalies in the same reaction.
23. The method of any one of claims 1 -22, wherein the one or more chromosomal anomalies are indicative of a fetal disease selected from the group consisting of DiGeorge syndrome, Prader-Willi syndrome, Angelman syndrome, Cri-du-chat syndrome, Down syndrome, Patau syndrome, Edward syndrome Trisomy 21, Turner’s syndrome, Triploidy, and Kleinfelter’s syndrome.
24. The method of any one of claims 1-23, wherein identifying the one or more chromosomal anomalies comprises utilizing decoded codes of a plurality of amplified recognition elements that align to the one or more chromosomal sequences in the cell-free DNA sample to determine a copy number of a first chromosome and a copy number of a second chromosome, and determining a ratio between the determined copy number of the first chromosome and the determined copy number of the second chromosome, and wherein the ratio identifies a presence or an absence of the one or more chromosomal anomalies present in the cell-free DNA sample.
25. A method for identifying one or more chromosomal anomalies in a cell-free DNA sample, comprising: a) hybridizing a plurality of recognition elements to chromosomal sequences present in the cell-free DNA sample to generate hybridized recognition elements, wherein each recognition element of the plurality of recognition elements comprises a code, a 5’ region, and a 3’ region, wherein each code of each recognition element of the plurality of recognition elements identifies the chromosomal sequences present in the cell-free DNA sample that are complementary to the corresponding recognition element of the plurality of recognition elements, wherein the presence of the chromosomal sequences indicates a presence of one or more chromosomal anomalies in the cell-free DNA sample; b) ligating the 5’ region and the 3 ’ region of the hybridized recognition elements, thereby generating circularized and ligated recognition elements; c) amplifying the circularized and ligated recognition elements to generate amplified recognition elements; d) detecting the codes of the amplified recognition elements; and e) decoding the codes that are detected in d), thereby identifying the one or more chromosomal anomalies in the cell-free DNA sample.
26. The method of claim 25, wherein the cell-free DNA sample is extracted from blood.
27. The method of claim 25 or 26, further comprising performing an exonuclease treatment after b).
28. The method of any one of claims 25-27, wherein the 5’ region is adjacently located to the 3 ’ region of the hybridized recognition elements.
29. The method of any one of claims of 25-28, wherein the 5’ region is not adjacently located to the 3’ region of the hybridized recognition elements, and wherein the method further comprises performing a gap-fill prior to ligating the 5’ region and the 3 ’ region of the hybridized recognition elements.
30. The method of any one of claims 25-29, wherein a code is the same for a subset of recognition elements of the plurality of recognition elements from a region of a chromosome.
31. The method of claim 30, wherein the subset of recognition elements comprises 5 to 50 recognition elements.
32. The method of any one of claims 25-31, wherein the code comprises one or more nucleic acid segments.
33. The method of claim 32, wherein the one or more nucleic acid segments comprises two to ten nucleic acid segments.
34. The method of claim 33, further comprising hybridizing one or more detection polynucleotides to the two to ten nucleic acid segments and imaging the hybridizing the one or more detection polynucleotides to the two to ten nucleic acid segments, thereby detecting the two to ten nucleic acid segments.
35. The method of claim 34, wherein each detection polynucleotide of the one or more detection polynucleotides comprises a fluorescent moiety.
36. The method of any one of claims 25-35, wherein decoding the codes comprises soft decision decoding.
37. The method of any one of claims 25-36, wherein the amplifying comprises polymerase chain reaction, primer extension, rolling circle amplification, or multiple strand displacement amplification.
38. The method of any one of claims 25-37, wherein the amplifying comprises generating concatemeric amplification products.
39. The method of any one of claims 25-38, wherein the one or more chromosomal anomalies comprises a copy number variation of a chromosome, an insertion or a deletion in a chromosome, or a sub-chromosomal deletion.
40. The method of any one of claims 25-39, wherein the one or more chromosomal anomalies is in one or more of chromosome 1, chromosome 4, chromosome 5, chromosome 13, chromosome 16, chromosome 15, chromosome 18, chromosome 21, chromosome 22, chromosome X, or chromosome Y.41 . The method of any one of claims 25-40, further comprising estimating a fetal fraction of the cell-free DNA sample.
42. The method of claim 41, wherein the estimating the fetal fraction comprises identifying a plurality of single nucleotide polymorphisms in fetal DNA and in maternal DNA using the plurality of recognition elements.
43. The method of claim 42, wherein the plurality of single nucleotide polymorphisms comprises 50 single nucleotide polymorphisms to 300 single nucleotide polymorphisms.
44. The method of claim 43, wherein the plurality of single nucleotide polymorphisms are located across a plurality of chromosomes in the cell-free DNA sample.
45. The method of claim 44, wherein the plurality of chromosomes comprises chromosome 1, chromosome 2, chromosome 3, chromosome 4, chromosome 5, chromosome 6, chromosome 7, chromosome 8, chromosome 9, chromosome 10, chromosome 11, chromosome 12, chromosome 13, chromosome 14, chromosome 15, chromosome 16, chromosome 17, chromosome 18, chromosome 19, chromosome 20, chromosome 21, chromosome 22, chromosome 23, or chromosome 24 in the cell-free DNA sample.
46. The method of claim 45, further comprising identifying the plurality of single nucleotide polymorphisms and the one or more chromosomal anomalies in the same reaction.
47. The method of any one of claims 25-46, wherein the one ormore chromosomal anomalies are indicative of a fetal disease selected from the group consisting of DiGeorge syndrome, Prader- Willi syndrome, Angelman syndrome, Cri-du-chat syndrome, Down syndrome, Trisomy 13, Trisomy 18, Trisomy 21, Turner’s syndrome, Triploidy, and Kleinfelter’s syndrome.
48. A composition comprising: a) a first plurality of recognition elements, wherein each recognition element of the first plurality of recognition elements is hybridized to non-polymorphic target sequences on one of chromosome 1, chromosome 4, chromosome 5, chromosome 13, chromosome 16, chromosome 15, chromosome 18, chromosome 21, chromosome 22, chromosome X, or chromosome Y, and b) a second plurality of recognition elements, wherein each recognition element of the second plurality of recognition elements is hybridized to a target sequence comprising a single nucleotide polymorphism sequence on one of chromosome 1 through chromosome 24.
49. The composition of claim 48, further comprising one or more of a ligase or an exonuclease.
50. The composition of claim 48, further comprising a first plurality of circularized and ligated recognition elements and a second plurality of circularized and ligated recognition elements.
51. The composition of claim 50, further comprising a polymerase enzyme and a plurality of concatenated amplification products.
52. A kit, comprising: a) a plurality of recognition elements, wherein each recognition element of the plurality of recognition elements is capable of hybridizing to target chromosomal sequences in a cell-free DNA sample; b) one or more of a ligase and a DNA polymerase; c) a plurality of fluorescently labeled polynucleotides; and d) instructions for practicing the method of any one of claims 1-47.
53. The kit of claim 52, further comprising an extraction medium for extracting cell-free DNA from the blood sample.
54. The kit of claim 52, wherein the ligase comprises a thermostable ligase.
55. The kit of claim 52, wherein the kit further comprises a thermostable DNA polymerase comprising strand displacement activity.
56. The kit of claim 52, wherein the kit comprises one or more exonucleases.
Citation Information
Patent Citations
Encoded assays
WO2022109496A2
Multiscale lens systems and methods for imaging well plates and including event-based detection
WO2023158993A2
Methods for increasing fetal fraction in maternal blood
US20140065621A1
Systems and methods for prenatal genetic analysis
WO2014165267A2
US202463641747P