Synthetic nucleic acid spike-ins
Patent Information
- Application Number
- JP2024227628
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-01-27
- Filing Date
- 2024-12-24
- Publication Date
- 2025-08-04
AI Technical Summary
The prior art has insufficient efficiency and accuracy in detecting and quantifying low-abundance nucleic acids in complex samples, especially in detecting pathogenic nucleic acids in clinical samples.
The reduced diversity of synthetic nucleic acids is used to detect and quantify the target nucleic acid by adding at least 1,000 synthetic nucleic acids with unique variable regions to the sample and performing next-generation sequencing.
The detection and quantitative accuracy of target nucleic acids is improved, especially in low abundance, and the detection ability of pathogen nucleic acids in clinical samples is enhanced.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to synthetic nucleic acid spike-ins.
[0002] cross reference This application is a joint venture of U.S. Provisional Patent Application No. 62 / 313,668, filed March 25, 2016. No. 62 / 397,873, filed September 21, 2016, and U.S. Provisional Patent Application No. 201 No. 62 / 451,363, filed Jan. 27, 2007 (the entire disclosure of which is incorporated herein by reference). This application claims the benefit of the above-referenced US Pat. No. 6,399,443, which is incorporated herein by reference. [Background technology]
[0003] Next-generation sequencing collects large amounts of data about the genetic content of a sample. It can be used for the analysis of nucleic acids in complex samples such as clinical samples and It may be particularly useful for sequencing entire genomes. However, it is also useful for determining nucleic acids, particularly low abundance nucleic acids or patient More efficient and accurate methods for detecting and quantifying nucleic acids in samples are needed. There is a need in the technology field. Summary of the Invention
[0004] overview Next-generation sequencing assays and other assays using spike-in synthetic nucleic acids The present invention provides methods and compositions for the improved identification or quantification of nucleic acids in In some cases, the spike-in synthetic nucleic acid is provided with a specific sequence, length, GC content, The invention has special characteristics such as a degree of degeneracy, diversity and / or a known starting concentration. The methods provided herein are particularly useful for detecting pathogen nucleic acids in clinical samples such as plasma. However, it can also be used to detect other types of targets.
[0005] In one embodiment, the abundance of a nucleic acid in an initial sample containing a target nucleic acid is determined. The present invention provides a method for producing at least 1,000 synthetic nucleic acids. A volume of the at least 1,000 synthetic nucleic acids is added to the sample, each of which contains a unique variable region; and (b) Sequencing a portion of the target nucleic acid and a portion of the at least 1,000 synthetic nucleic acids. An assay is performed, thereby obtaining target and synthetic nucleic acid sequence reads. wherein the synthetic nucleic acid sequence reads comprise unique variable region sequences; and (c)(i) the synthetic nucleic acid sequence reads quantitating the number of different variable region sequences within the nucleic acid sequence reads to obtain a unique sequence determination value; ) comparing the starting amount of said at least 1,000 synthetic nucleic acids with said unique sequence determination values; By obtaining a diversity reduction of the at least 1,000 synthetic nucleic acids, (d) detecting a reduction in diversity of at least 1,000 synthetic nucleic acids; The diversity reduction of 0,000 synthetic nucleic acids was used to measure the abundance of the target nucleic acid in the initial sample. In some cases, the starting amount to compare is a starting concentration.
[0006] In some embodiments, the target nucleic acid comprises a pathogen nucleic acid. The nucleic acid includes pathogen nucleic acids from at least five different pathogens. The target nucleic acid includes pathogen nucleic acids from at least two different pathogens. The target nucleic acids include pathogen nucleic acids from at least 10 different pathogens.
[0007] In some cases, the at least 1,000 synthetic nucleic acids include DNA. In the case of the above, the at least 1,000 synthetic nucleic acids are RNA, ssRNA, dsDNA, In some cases, the nucleic acid may comprise at least one of the above. Each of the 1,000 synthetic nucleic acids is less than 500 base pairs or nucleotides in length. In some cases, each of the at least 1,000 synthetic nucleic acids is 200 bases long. In some cases, the length is at least 1,000 pairs or nucleotides. Each of the synthetic nucleic acids is less than 100 base pairs or nucleotides in length. In this study, samples were collected from blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, feces, and saliva. In some cases, the sample is from a human subject. In the case of, the sample is a sample of isolated nucleic acid.
[0008] In some cases, the method further comprises producing a sequencing library from the sample. ), wherein said at least one, 000 synthetic nucleic acids are added to the sample. In some cases, at least 1,000 of the above A reduction in diversity of 0 synthetic nucleic acids indicates a reduction in one or more nucleic acids during sample processing of the sample.
[0009] In some cases, each of the at least 1,000 synthetic nucleic acids comprises an identification tag sequence. In some cases, the quantification of the number of unique variable region sequences includes the number of sequences that contain the tag sequence. In some cases, detecting a sequence that is at least 1,0 in the first sequence read. Quantification of 00 unique sequences is based on the number of unique sequence reads in the first sequence read (read count). In some cases, determining at least 1,000 unique Synthetic nucleic acids are at least 10 4 The nucleic acid sequence comprises unique synthetic nucleic acids.
[0010] In some cases, the method further comprises adding a first additional synthetic nucleic acid population having a first length, a second additional synthetic nucleic acid population having a second length, a second additional synthetic nucleic acid group having a length of 1000; and a third additional synthetic nucleic acid group having a length of 1000; adding an acid group, wherein each of the first, second and third additional synthetic nucleic acid groups The method includes a synthetic nucleic acid having at least three different GC contents. The method further comprises using the additional synthetic nucleic acid to determine the absolute abundance value of the target nucleic acid in the sample. In some cases, the method further comprises using the additional synthetic nucleic acid to: The additional synthetic nucleic acid is classified into samples based on its length, GC content, or both length and GC content. The method includes calculating the absolute or relative abundance of the target nucleic acid in the sample.
[0011] In some cases, the first sample processing step comprises: Synthetic nucleic acids are added to the sample. In some cases, the method further comprises a second sample processing step. adding an additional pool of at least 1,000 unique synthetic nucleic acids to the sample; wherein the second sample processing step is different from the first sample processing step. In some cases, the method further comprises the step of: In some cases, the method further comprises calculating a diversity reduction resulting from at least 1, The diversity reduction for 1,000 synthetic nucleic acids was measured by additionally generating at least 1,000 synthetic nucleic acids. The samples showing relatively high diversity reduction were compared to the diversity reduction for the This includes specifying the processing steps.
[0012] In some cases, an additional pool of at least 1,000 unique synthetic nucleic acids. Each unique synthetic nucleic acid is a member of an additional pool of at least 1,000 synthetic nucleic acids. In some cases, the method further comprises: In some cases, (a) above comprises adding a sample-discriminating nucleic acid to the sample. Further, it involves adding a non-unique synthetic nucleic acid to the sample.
[0013] In some embodiments, the abundances calculated are relative abundances. In the morphology, the abundances calculated are absolute abundances.
[0014] In another embodiment, the relative abundance or initial abundance of pathogen nucleic acid in a sample is determined. The present invention provides a method for determining the amount of a pathogen present in a subject, the method comprising: (a) detecting a pathogen in a subject that is infected or infected with the pathogen; obtaining a sample from a subject suspected of having the disease, wherein the sample contains a plurality of pathogen nucleic acids; (b) adding a plurality of synthetic nucleic acids to the sample such that the sample contains a known initial abundance of the synthetic nucleic acid. wherein (i) the synthetic nucleic acid is less than 500 base pairs in length, and (ii) the synthetic The nucleic acids include synthetic nucleic acids having a first length, synthetic nucleic acids having a second length, and synthetic nucleic acids having a third length. wherein the first, second and third lengths are different; and (iii) a synthetic nucleic acid having Synthetic nucleic acids having a length of 1 include synthetic nucleic acids having at least three different GC contents. (c) performing a sequencing assay on a sample containing the plurality of synthetic nucleic acids; The final amount of the synthetic nucleic acid and the final amount of the multiple pathogen nucleic acids are determined by the above method. d) comparing the final abundance of the synthetic nucleic acid with the known initial abundance to determine a recovery profile for the synthetic nucleic acid; (e) obtaining a recovery profile for the synthetic nucleic acid, and (f) detecting the pathogen nucleic acid using the recovery profile for the synthetic nucleic acid. is compared to a synthetic nucleic acid having the closest GC content and length, thereby The pathogen nucleic acid relative abundance or initial abundance is determined to identify the pathogens. This involves normalizing the final abundance of the nucleic acid.
[0015] In some cases, the at least three different GC contents are between 10% and 40%. a first GC content, a second GC content between 40% and 60%, and a third GC content between 60% and 90%. In some cases, the at least three different GC contents are In some cases, the at least three different GC contents are 10% to 50%. In some cases, the synthetic nucleic acid is 200 base pairs or less. In some cases, the synthetic nucleic acid is less than 100 base pairs or nucleotides in length. In some cases, the at least three different GC contents are less than one nucleotide length. At least four, at least five, at least six, at least seven, or at least eight In some cases, the synthetic nucleic acid has at least a fourth length, at least At least the fifth length, at least the sixth length, at least the seventh length, at least the ninth length length, at least the 10th length, at least the 12th length, or at least the 15th In some embodiments, each length is at least 3, 4, 5, 6, 7 , 8, 9, 10 different GC contents, or synthetic nuclei with up to 50 different GC contents. Contains acid.
[0016] In some cases, the synthetic nucleic acid comprises double stranded DNA. In some cases, the method further comprises: In some cases, the synthetic nucleic acid is used to monitor the alteration of pathogen nucleic acid. In one embodiment, the method further comprises using a weighting factor to estimate the relative abundance or abundance of the pathogen nucleic acid. In some cases, the known concentration of the first synthetic nucleic acid and The raw measured value and the raw value of the first synthetic nucleic acid of the plurality of synthetic nucleic acids are compared with the known concentration of the second synthetic nucleic acid. and a weighting factor is determined by analyzing the raw measurements of the first and second synthetic nucleic acids of the plurality of synthetic nucleic acids. Get the number.
[0017] In another aspect, the present invention provides a method for detecting nucleic acid from a pathogen. The method includes (a) obtaining a first sample containing a first pathogen nucleic acid, wherein the first sample comprises: (b) obtaining a second sample from a second subject infected with a first pathogen; (c) a first set of nucleic acids, each of which comprises a different synthetic nucleic acid that cannot hybridize to the first pathogen nucleic acid; A sample identifier and a second sample identifier are obtained, and the first sample identifier is assigned to the first sample. (d) assigning the first sample identifier to the first sample; (e) adding a first sample identifier to the sample and a second sample identifier to the second sample; for the first sample that contains the child, and for the second sample that contains the second sample identifier A sequencing assay is performed, thereby determining sequence results for the first sample and the second sample. (f) obtaining a sequence result for the first sample, a sequence result for the second sample, (g) detecting the presence or absence of the sequence identifier and the first pathogen nucleic acid; (i) detecting a first sample identifier; and (ii) detecting a first pathogen nucleus in the first sample. and (iii) detecting no or less than a threshold level of the second sample identifier. If the second sample identifier is not detected, the first pathogen nucleic acid detected is in the first sample. This includes determining that it exists in the first place.
[0018] In another aspect, the present invention provides a method for detecting a nucleic acid, the method comprising: (a) obtaining a first nucleic acid sample comprising a first nucleic acid; and (b) obtaining a first control nucleic acid sample comprising a first positive control nucleic acid. (c) obtaining an acid sample, the first sample comprising a synthetic nucleic acid that is not capable of hybridizing to the first nucleic acid. (d) adding an identifier to the first control nucleic acid sample and and performing a sequencing assay on the first nucleic acid sample, thereby determining the first and control nucleic acid samples. (e) comparing the sequence reads for the first nucleic acid sample with a reference sequence; Aligning the first sample in the sequence reads for the first nucleic acid sample (f) detecting the presence or absence of a sequence identifier based on the alignment of the sequence reads. Determining whether a first positive control nucleic acid is present in the first nucleic acid sample.
[0019] In some cases, the synthetic nucleic acid of the first sample identifier is 150 base pairs or nucleotides. In some cases, the first positive control nucleic acid is a pathogen nucleic acid. In some embodiments, the first sample identifier comprises a modified nucleic acid. In some embodiments, the first sample identifier comprises In some cases, the sample contains acellular body fluid. The samples are from subjects infected with a pathogen.
[0020] In another aspect, the present invention provides a method for detecting a reagent in a sample. The method includes: (a) adding a first synthetic nucleic acid to the reagent, wherein the first synthetic nucleic acid has a unique sequence. (b) adding a reagent comprising a first synthetic nucleic acid to the nucleic acid sample; and (c) performing a sequencing assay. (d) performing a sequencing assay on the nucleic acid sample. (e) obtaining sequence results for the nucleic acid sample; Based on the result, the presence or absence of the first synthetic nucleic acid in the sample is determined. and detecting the reagent in the sample.
[0021] In some cases, the first synthetic nucleic acid is less than 150 base pairs or nucleotides in length. In some cases, the first synthetic nucleic acid is added to the first reagent lot, and the second synthetic nucleic acid is added to the second reagent lot. In some cases, detecting a reagent in a sample may involve adding the reagent to the lot. In some cases, the synthetic nucleic acid is a nucleic acid. In some cases, the reagent comprises an aqueous buffer. In some cases, the reagents include an extraction reagent, an enzyme, a ligase, a polymerase, or dNTPs.
[0022] In another aspect, the present invention provides a method for producing a sequencing library. The method includes: (a) providing (i) a target nucleic acid; (ii) a sequencing adaptor; and (iii) A sample containing at least one synthetic nucleic acid is obtained, wherein said at least one synthetic nucleic acid is (b) the sequencing adapter comprises DNA and is resistant to ligation to a nucleic acid; and and binding the target nucleic acid to the sample in preference to the synthetic nucleic acid. This includes carrying out a reaction.
[0023] In another aspect, the present invention provides a method for producing a sequencing library, the method comprising: The method includes (a) obtaining a sample containing a target nucleic acid and at least one synthetic nucleic acid; and (b) At least one synthetic nucleic acid is removed from the sample, thereby forming a sample that contains the target nucleic acid. (c) a sequencing sample that does not contain at least one synthetic nucleic acid of the sequencing adaptor; The method includes binding the nucleic acid sequence to a target nucleic acid in a sequencing sample.
[0024] In another aspect, the present invention provides a method for producing a sequencing library, the method comprising: The method includes (a) obtaining a sample containing a target nucleic acid and at least one synthetic nucleic acid, wherein said At least one of the synthetic nucleic acids is (i) a single-stranded DNA, and (ii) a nucleic acid sequence that inhibits the amplification of the synthetic nucleic acid. (iii) immobilization tags; (iv) DNA-RNA hybrids; (v) a nucleic acid having a length greater than the length of the target nucleic acid, or (vi) any combination thereof. (b) preparing a sequencing library from the sample for a sequencing reaction; (wherein at least a portion of the at least one synthetic nucleic acid is (not sequenced).
[0025] In another aspect, the present invention provides a method for producing a sequencing library, the method comprising: The method includes: (a) providing (i) a target nucleic acid; (ii) a sequencing adaptor; and (iii) at least obtaining a sample containing at least one synthetic nucleic acid, wherein said at least one synthetic nucleic acid is a DNA fragment; A, and is resistant to end repair; and (b) the target nucleic acid is more specific than the at least one synthetic nucleic acid. performing an end-repair reaction on the sample such that the ends are preferentially repaired.
[0026] In another embodiment, a kit for producing a sequencing library is provided according to the present invention. The kit includes: (a) a sequencing adaptor; and (b) at least one synthetic nuclease. wherein the at least one synthetic nucleic acid comprises DNA and a terminal acid to the nucleic acid. Resists repair.
[0027] In one embodiment, the absolute or relative concentration of the nucleic acid in the initial sample containing the target nucleic acid is Provided herein is a method for determining the abundance of a target molecule, the method comprising: (a) determining the abundance of a target molecule in a sample of at least 1.00 mg / mL; 1.00 unique synthetic nucleic acids are added to the sample, Each of the 0 unique synthetic nucleic acids comprises (i) an identification tag, and (ii) a variable region. (b) a portion of the target nucleic acid in the sample and the at least 1,000 uniplexes; A sequencing assay is performed on a portion of the target synthetic nucleic acid, thereby obtaining sequence reads, wherein the synthetic nucleic acid sequence reads include an identifier tag sequence and a variable region sequence; (c) (i) detecting sequence reads corresponding to at least a portion of the identification tag sequence to obtain a first obtaining a set of sequence reads; and (ii) quantifying the number of distinct variable region sequences within the first sequence read. (iii) obtaining a unique sequence determination value from said at least 1,000 unique combinations. The starting amount of synthetic nucleic acid is compared to the unique sequence determination value to identify the at least 1,000 unique By obtaining a diversity reduction of the nucleotide synthetic nucleic acid, the at least 1,000 unique (d) detecting a reduction in diversity of said at least 1,000 unique synthetic nucleic acids; Nucleic acid diversity reduction can be used to determine the absolute or relative abundance of a target nucleic acid in an initial sample. In some cases, the starting amount to compare is a starting concentration.
[0028] In some cases, the target nucleic acid includes a pathogen nucleic acid. The pathogen nucleic acid may include at least five different pathogens. In some cases, the pathogen nucleic acid may include at least five different pathogens. A total of 1,000 unique synthetic nucleic acids include DNA.
[0029] In some cases, each of the at least 1,000 unique synthetic nucleic acids comprises 5 In some cases, the length of the at least one Each of the 0,000 unique synthetic nucleic acids is less than 200 base pairs or nucleotides in length. In some cases, each of the at least 1,000 unique synthetic nucleic acids is 00 base pairs or nucleotides in length.
[0030] In some cases, samples were blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage The sample may be a sputum, urine, feces, saliva or nasal sample. In some cases, the sample is isolated. In some cases, the sample is from a human subject.
[0031] In some cases, the method further comprises producing a sequencing library from the sample. and wherein said at least 1,000 nucleic acids are isolated prior to preparing the sequencing library. In some cases, at least 1,000 unique synthetic nucleic acids are added to the sample. The diversity reduction of 0 unique synthetic nucleic acids is the reduction of one or more nucleic acids during sample processing of the sample. In some cases, the identification tag comprises a common sequence. In some cases, the first sequence Quantification of at least 1,000 unique sequences within a read is based on the number of unique sequences within the first sequence read. This includes determining the number of reads for the sequence.
[0032] In some cases, the at least 1,000 unique synthetic nucleic acids include at least 1 0 4 In some cases, the at least 1,000 unique synthetic nucleic acids. Unique synthetic nucleic acids are at least 10 5 In some cases, the nucleic acid sequence includes: The method further includes adding additional synthetic nucleic acids having at least three different lengths. .
[0033] In some cases, the method further comprises adding a first additional synthetic nucleic acid population having a first length, a second additional synthetic nucleic acid population having a second length, a second additional synthetic nucleic acid group having a length of 1000; and a third additional synthetic nucleic acid group having a length of 1000; adding an acid group, wherein each of the first, second and third additional synthetic nucleic acid groups The method includes a synthetic nucleic acid having at least three different GC contents. The method further comprises using the additional synthetic nucleic acid to detect absolute or relative abundance of a target nucleic acid in a sample. In some cases, the method further comprises calculating a specific abundance value of the additional synthetic nucleic acid. Based on the length, GC content, or both the length and GC content of the additional synthetic nucleic acid, The method includes calculating the absolute or relative abundance of the target nucleic acid in the sample based on the results of the analysis.
[0034] In some cases, the first sample processing step comprises: A unique synthetic nucleic acid is added to the sample. In some cases, the method further comprises: In the processing step, an additional pool of at least 1,000 unique synthetic nucleic acids is sampled. wherein the second sample processing step is different from the first sample processing step. In some cases, the method further comprises tracking at least 1,000 unique synthetic nucleic acids. In some cases, the method further comprises calculating the diversity reduction for the additive bins. and reducing the diversity of said at least 1,000 unique synthetic nucleic acids by at least 1. By comparing the diversity reduction with an additional pool of 1,000 unique synthetic nucleic acids, , including identifying sample processing steps that exhibit relatively high diversity reduction.
[0035] In some cases, an additional pool of at least 1,000 unique synthetic nucleic acids. Each unique synthetic nucleic acid is a member of an additional pool of at least 1,000 synthetic nucleic acids. In some cases, the method further comprises: In some cases, (a) above comprises adding a sample-discriminating nucleic acid to the sample. Further, it includes adding a non-unique synthetic nucleic acid to the sample. In some cases, variable sequence reads are detected by aligning the by aligning variable sequence reads to each other and removing duplicate sequence reads. The number of distinct variable sequence reads is quantified.
[0036] The present invention provides a method for determining the relative abundance or concentration of pathogen nucleic acid in a sample of nucleic acid. In some cases, the method includes administering to a subject infected or suspected of being infected with a pathogen. obtaining a sample from a subject to be screened, wherein the sample contains two or more pathogen nucleic acids, The two or more pathogen nucleic acids may include a first pathogen nucleic acid and a second pathogen nucleic acid having different lengths. and adding known concentrations of two or more synthetic nucleic acids to the sample (wherein the two or more synthetic nucleic acids The nucleic acid of the first pathogen is 65% to 135%, 75% to 125%, or 85% to 11% of the nucleic acid of the first pathogen. a first synthetic nucleic acid having a length of 5%, and 65% to 135%, 75% to 135% of the second pathogen nucleic acid; a second synthetic nucleic acid having a length of 125%, or 85% to 115%, the two or more synthetic nucleic acids do not hybridize to the first or second pathogen nucleic acid, and performing a sequencing assay to determine the two or more synthetic nucleic acids, the first pathogen nucleic acid, and and a raw measurement value for the second pathogen nucleic acid, and the raw measurement value of the first synthetic nucleic acid is compared with the previously measured value of the first synthetic nucleic acid. A recovery profile for the first synthetic nucleic acid is obtained by comparing the concentration of the first synthetic nucleic acid with the known concentration. The recovery profile is used to normalize the raw measurements for the first pathogen nucleic acid, thereby , determining the relative abundance or starting concentration of the first pathogen nucleic acid.
[0037] In some cases, the first pathogen nucleic acid and the second pathogen nucleic acid are derived from the same pathogen. In some cases, the first pathogen nucleic acid and the second pathogen nucleic acid are from different pathogens. In some cases, the methods described herein further include using weighting factors. and normalizing the relative abundance or starting concentration of the first pathogen nucleic acid by In some cases, compared to a known concentration of the first synthetic nucleic acid and a known concentration of the second synthetic nucleic acid. a raw measurement value of a first synthetic nucleic acid of said plurality of synthetic nucleic acids and a raw measurement value of a second synthetic nucleic acid of said plurality of synthetic nucleic acids Weighting factors are obtained by analyzing the raw measurements of the synthetic nucleic acid.
[0038] A method for determining the relative abundance or starting concentration of a nucleic acid in a sample of nucleic acid is provided herein. The method includes (a) obtaining a nucleic acid sample from a subject, the nucleic acid sample being a nucleic acid sequence of a different The first and second nucleic acids have a length of about 100 nm, and two or more synthetic nucleic acids of known concentrations are sampled. (wherein, (i) the two or more synthetic nucleic acids are 65% to 135%, 7% or more of the first nucleic acid A first synthetic nucleic acid having a length of 5% to 125%, or 85% to 115%, and a second nucleic acid have a length of 65% to 135%, 75% to 125%, or 85% to 115% of the length of (ii) a second synthetic nucleic acid, the first synthetic nucleic acid comprising a load domain of a particular length; and an identifier having a unique sequence coded to identify a particular length of the loading domain. and (iii) the two or more synthetic nucleic acids hybridize to the first or second nucleic acid. (b) performing a sequencing assay on the sample, thereby (c) obtaining raw measurements for the two or more synthetic nucleic acids, the first nucleic acid and the second nucleic acid; (d) comparing the raw nucleic acid measurements to the known concentrations of the first synthetic nucleic acid to obtain a recovery profile; The recovery profile is used to normalize the raw measurements for the first nucleic acid, thereby This involves determining the relative abundance or starting concentration of the nucleic acid.
[0039] In some cases, the first nucleic acid is a pathogen nucleic acid. The known concentrations of synthetic nucleic acids are 2 or more, 3 or more, 5 or more, 10 or more, 50 or more, 100 or more, or In some cases, the two or more synthetic nucleic acids may be combined to produce a single nucleic acid having a different concentration. The known concentrations are equimolar. In some cases, the two or more synthetic nucleic acids are DNA or In some cases, the two or more synthetic nucleic acids are RNA or modified RNA. In some cases, the two or more synthetic nucleic acids include 2 or more, 3 or more, 5 or more, 8 or more, 10 or more, 10 or more, 50 or more, 100 or more, or 1,000 or more different lengths of nucleic acid. In some cases, the two or more synthetic nucleic acids are two or more, three or more, five or more, eight or more, It contains 10 or more, 50 or more, 100 or more, or 1,000 or more different sequences of nucleic acid. In some cases, the two or more synthetic nucleic acids are less than 50 nucleotides in length, less than 100 nucleotides in length. 100 nucleotides or less, 200 nucleotides or less, 300 nucleotides or less, 350 nucleotides or less ≤ 400 nucleotides, ≤ 450 nucleotides, ≤ 500 nucleotides , 750 nucleotides or less in length, or 1,000 nucleotides or less in length. In the case of the above, the two or more synthetic nucleic acids are at least 10 nucleotides in length and at least 20 nucleotides in length, or at least 30 nucleotides in length, at least 50 nucleotides in length , at least 100 nucleotides in length, or at least 150 nucleotides in length. In some cases, the two or more synthetic nucleic acids are characterized as synthetic products. In some cases, two or more of the synthetic nucleic acids are identified as synthetic. The nucleic acid sequence is 10 nucleotides or less in length, 20 nucleotides or less in length, 30 nucleotides or less in length The following are the lengths: 40 nucleotides or less, 50 nucleotides or less, 100 nucleotides or less, It is 200 nucleotides or less in length, or 500 nucleotides or less in length. In some cases, the two or more synthetic nucleic acids include a nucleic acid sequence that specifies the length of the synthetic nucleic acid. In this case, the nucleic acid sequence specifying the length of the synthetic nucleic acid is 10 nucleotides or less, 20 nucleotides or less. 30 nucleotides or less, 40 nucleotides or less, 50 nucleotides or less 100 nucleotides or less, 200 nucleotides or less, or 500 nucleotides or less The length is less than or equal to 100 mm.
[0040] In some cases, the sample may be blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar fluid, or the like. selected from the group consisting of lavage fluid, urine, stool, saliva, nasal swab and any combination thereof In some cases, the sample comprises cell-free nucleic acid. In some cases, the sample comprises circulating In some cases, the subject is a human. In some cases, the pathogen is The infection is a bacterial, viral, fungal or parasitic infection. In some cases, the subject has sepsis. In some cases, the pathogen is associated with sepsis. In this case, the two or more pathogen nucleic acids are 3 or more, 5 or more, 10 or more, 50 or more, 100 or more. , 1,000 or more, 2,000 or more, 5,000 or more, 8,000 or more, 10,000 or more or more than 15,000, or more than 20,000 pathogen nucleic acid sequences.
[0041] In some cases, determining the relative abundance of the first pathogen nucleic acid comprises determining one or more genome copies. In some cases, producing one or more genome copies per volume In some cases, the methods described herein further comprise: In some cases, the extraction of nucleic acid from the sample is The extraction is performed using magnetic beads. In some cases, the methods described herein It also includes removing low quality sequencing reads. The described method further includes aligning or matching the sequence to a reference sequence for the species of interest. In some cases, the method includes removing the pinged sequencing reads. The disclosed method further includes determining the relative efficiency of recovering nucleic acids of one or more different lengths. In some cases, the methods described herein further include the step of synthesizing one or more synthetic nucleic acids. In some cases, the methods described herein include determining a measured concentration. and comparing the measured concentration of said one or more synthetic nucleic acids to a known concentration. The methods described herein may further comprise the step of: 3 or more, 5 or more, 10 or more, 50 or more, 100 or more, 1,000 or more, 2,000 or more, 5,000 or more, 8,000 or more, 10,000 or more, 15,000 or more, or 20, 000 or more pathogen nucleic acids. The method further comprises the step of detecting an antimicrobial, antibacterial, antiviral or antiviral compound in a sequencing assay. 1 or more, 2 or more, 3 or more, 5 or more, 10 or more, 50 or more, 100 or more, indicating fungal resistance ,000 or more, 2,000 or more, 5,000 or more, 8,000 or more, 10,000 or more, Some include detecting more than 15,000 or more than 20,000 pathogen nucleic acids. In the case of, the methods described herein further comprise the step of: This includes identifying the simultaneous presence of 5 or more, 10 or more, 50 or more, or 100 or more pathogens.
[0042] In some cases, two or more of the synthetic methods described above may be performed before or during the extraction of nucleic acids from a sample. Nucleic acid is added to the sample. In some cases, after extraction of nucleic acid from the sample and Prior to library preparation, two or more synthetic nucleic acids are added to the sample. In some cases, the length of the two or more synthetic nucleic acids differs by at least about 20 base pairs. The two or more synthetic nucleic acids are 3 or more, 5 or more, 8 or more, 10 or more, 20 or more, or 50 In some cases, the two or more synthetic nucleic acids include SEQ ID NO: 111. ~ SEQ ID NO: 118, and any combination thereof. In some cases, the two or more synthetic nucleic acids share a common forward sequence. The common forward sequence is about 20 base pairs in length or less. In some cases, the two The synthetic nucleic acids share a common reverse sequence. In some cases, the common reverse sequence The source sequence is about 20 base pairs or less in length.
[0043] In some cases, the methods described herein further include measuring the raw value of the second synthetic nucleic acid. A recovery profile for the second synthetic nucleic acid is obtained by comparing with a known concentration of the second synthetic nucleic acid. The recovery profile for the synthetic nucleic acid is used to normalize the raw measurements for the second pathogen nucleic acid. and thereby determining the relative abundance or starting concentration of the second pathogen nucleic acid. .
[0044] In some cases, the two or more pathogen nucleic acids are five or more pathogen nucleic acids having different lengths. The two or more synthetic nucleic acids each have a length equal to or greater than the length of each of the five or more pathogen nucleic acids. One or more with a length of 65% to 135%, 75% to 125%, or 85% to 115% and a synthetic nucleic acid, wherein the two or more nucleic acids do not hybridize to the five or more pathogen nucleic acids. performing a sequencing assay on the sample to identify the two or more synthetic nucleic acids and obtaining raw measurements for the five or more pathogen nucleic acids; and comparing the raw measurements to each of the five or more pathogen nucleic acids; obtaining a recovery profile for each synthetic nucleic acid by comparing with known concentrations of the synthetic nucleic acid. and / or using the recovery profile to generate a recovery profile for each synthetic nucleic acid. Normalize the raw measurements of each of the five or more pathogen nucleic acids using the determining the relative abundance or starting concentration of each of the five or more pathogen nucleic acids. In some cases, the five or more pathogen nucleic acids are ten or more, fifty or more, one hundred or more, 1,000 or more, 2,000 or more, 5,000 or more, 8,000 or more, 10,000 or more , 15,000 or more, or 20,000 or more pathogen nucleic acids. The methods described herein further include subjecting the two or more synthetic nucleic acids and the sample of nucleic acid to In some cases, the method includes extracting or purifying nucleic acids from two or more synthetic nucleic acids. The extraction or purification of nucleic acids in a sample of acid and nucleic acid includes the steps of: The relative concentration of nucleic acid in the sample of nucleic acid is changed. Values are number of reads.
[0045] Provided herein is a method for detecting nucleic acid from a pathogen, the method comprising: (a) detecting a first pathogen; obtaining a first nucleic acid sample containing pathogenic nucleic acid, the first nucleic acid sample being infected with or containing a first pathogen; (b) obtaining a first nucleic acid sample from a first subject suspected of being infected with a second pathogen nucleic acid; and obtaining a second nucleic acid sample comprising a second pathogen-infected or -infected ... (c) obtaining a second nucleic acid sample from a second subject suspected of having the pathogen nucleic acid; A first sample identifier and a second sample identifier each containing a different synthetic nucleic acid that cannot be replicated. and assigning a first sample identifier to a first nucleic acid sample and a second sample identifier to a second nucleic acid sample. (d) assigning a first sample identifier to the first nucleic acid sample and a second sample identifier to the second nucleic acid sample; (e) adding a sample identifier to the second nucleic acid sample, the first nucleic acid sample including the first sample identifier; and a sequencing uptake assay for a second nucleic acid sample comprising a second sample identifier. , thereby obtaining an array result for the first sample and the second sample, (f ) the presence of a first sample identifier, a second sample identifier, and a pathogen nucleic acid in the sequence result. or the absence thereof, and (g) said sequencing assay detects the presence or absence of the first sample identifier and the target nucleic acid. If the second sample identifier is not detected but the first sample identifier is not detected, the target nucleic acid is not present in the first sample. This includes determining that it exists in the first place.
[0046] In some cases, the synthetic nucleic acid is less than about 500 base pairs in length. The synthetic nucleic acid is less than about 100 base pairs in length. In some cases, the synthetic nucleic acid is at least about In some cases, the synthetic nucleic acid is at least about 100 base pairs in length. In some cases, the synthetic nucleic acid comprises DNA or modified DNA. In some cases, the synthetic nucleic acid comprises RNA or modified RNA. In some cases, the synthetic nucleic acid is a nucleic acid represented by SEQ ID NO: 1 to SEQ ID NO: 110 and the like. In some cases, the first sample comprises a sequence selected from the group consisting of any combination of: contains acellular body fluids.
[0047] A method for detecting a reagent in a sample is provided herein, the method comprising: Adding an acid to the reagent (wherein the first synthetic nucleic acid comprises a unique sequence) and eluting the first synthetic nucleic acid Adding a drug to the nucleic acid sample to prepare the nucleic acid sample for a sequencing assay; performing a sequencing assay on the nucleic acid sample, thereby obtaining a sequence result for the nucleic acid sample; Based on the sequence results for the nucleic acid sample, the presence or absence of the first synthetic nucleic acid in the sample is determined. involves detecting the reagent in the sample by determining its absence.
[0048] In some cases, adding the first synthetic nucleic acid to the reagent in step a may be a method for determining whether the first synthetic nucleic acid is a specific nucleic acid of the reagent. In some cases, the method includes adding a first synthetic nucleic acid to the lot. The method further includes detecting a particular lot of a reagent based on the sequence results for the nucleic acid sample. In some cases, the first synthetic nucleic acid hybridizes to a nucleic acid from a pathogen. In some cases, the methods described herein further include the step of comparing different lots of reagents. adding a second synthetic nucleic acid, wherein the second synthetic nucleic acid is a nucleic acid that is uniquely prepared by mixing a different lot of reagents. In some cases, the methods described herein further comprise: and detecting the target nucleic acid based on the results from a sequencing assay of the nucleic acid sample. In some cases, the methods described herein further include: (i) detecting a target nucleic acid accurately; If so, use that particular lot of reagent in future sequencing assays, or (ii) If the target nucleic acid is not accurately detected, the test This includes refraining from using a particular lot of a drug. In some cases, the agent is water soluble. In some cases, the synthetic nucleic acid is about 50 to about 500 base pairs in length. In some cases, the synthetic nucleic acid comprises DNA or modified DNA. In some cases, the synthetic nucleic acid comprises RNA or modified RNA. 110 and any combination thereof. In this case, the synthetic nucleic acid cannot be degraded by DNase.
[0049] Provided herein are methods for determining the diversity reduction or abundance of a nucleic acid in a sample. The method further comprises adding 1,000 unique synthetic nucleic acids of known concentration to a sample containing a target nucleic acid. Additionally, a sequencing assay is performed on the sample, thereby determining the number of sequence reads for the target nucleic acid. and obtaining sequence reads of at least a portion of the 1,000 unique synthetic nucleic acids; The number of sequence reads of at least a portion of the 1,000 unique synthetic nucleic acids is determined in step a. The sequences and alignments of 1,000 unique nucleic acids were added to the sample containing the target nucleic acid. The diversity of the number of aligned sequence reads was increased to 1,000 or more units. By comparing the diversity of the 1,000 unique synthetic nucleic acids, Detecting the diversity reduction and using the diversity reduction of the 1,000 unique synthetic nucleic acids, Calculating the diversity reduction in the target nucleic acid or the abundance of the target nucleic acid in the sample. Includes.
[0050] In some cases, the 1,000 unique synthetic nucleic acids are about 500 base pairs in length or less. or less than about 100 base pairs in length. In some cases, the 1,000 unique combinations In some cases, the 1,000 unique synthetic nucleic acids are added in equimolar concentrations. The acid is at least about 1 × 10 6 In some cases, the diversity of the 1,000 of unique synthetic nucleic acid is at least about 1×10 6 In some cases, The 1,000 unique synthetic nucleic acids are at least about 1×10 7 There is a diversity of In this case, the 1,000 unique synthetic nucleic acids are at least about 1×10 8 Diversity In some cases, the 1,000 unique synthetic nucleic acids have a randomized portion. In some cases, the 1,000 unique synthetic nucleic acids are DNA, modified In some cases, the 1,000 units may be selected from the group consisting of DNA, RNA, and modified RNA. The synthetic nucleic acid includes the sequences specified in SEQ ID NO:119 and SEQ ID NO:120. In some cases, the first sample processing step includes: Synthetic nucleic acids are added to the sample. In some cases, the method further comprises a second sample processing step. The method includes adding an additional pool of 1,000 unique synthetic nucleic acids to the sample. where the second sample processing step is different from the first sample processing step. calculates the diversity reduction for an additional pool of 1,000 unique synthetic nucleic acids. In some cases, the methods described herein may be used to generate 1,000 unique synthetic The diversity reduction for nucleic acids was compared with the diversity reduction for an additional pool of 1,000 unique synthetic nucleic acids. Identify sample processing steps that exhibit relatively high diversity reduction by comparing with the diversity reduction In some cases, the 1,000 unique synthetic nucleic acids include the 1, A domain that identifies the synthetic nucleic acid as a member of a pool of 0,000 unique synthetic nucleic acids. In some cases, the additional pool of 1,000 unique synthetic nucleic acids includes 1,00 A domain that identifies the synthetic nucleic acid as a member of an additional pool of 0 unique synthetic nucleic acids. In some cases, the 1,000 unique synthetic nucleic acids are used for the extraction of a target nucleic acid. In some cases, the 1,000 unique synthetic nucleic acids are added to the sample before the target nucleic acid is added. In some cases, the target nucleic acid is added to the sample prior to library preparation. The method further involves adding 5,000 unique combinations of known concentrations to a sample containing the target nucleic acid. The method includes adding synthetic nucleic acid.
[0051] Additionally, methods and compositions for analyzing molecules are disclosed herein. In one aspect, disclosed herein is a method for producing a sequencing library, the method comprising: a) (i) a target nucleic acid, (ii) a sequencing adaptor, and (iii) at least one Obtaining a sample containing synthetic nucleic acids, wherein said at least one synthetic nucleic acid comprises DNA. b) the sequencing adapter is resistant to ligation to the at least one synthetic nucleic acid; performing a ligation reaction on the sample such that the target nucleic acid is preferentially ligated to the target nucleic acid over the nothing.
[0052] In some cases, the at least one synthetic nucleic acid is linked via a phosphodiester bond. In some cases, the at least one synthetic nucleic acid is resistant to ligation to the nucleic acid. In another embodiment, the sequencing library is resistant to ligation to a sequencing adaptor. Disclosed herein is a method for producing a nucleic acid sequence comprising: a) reacting a nucleic acid sequence with a target nucleic acid; a) obtaining a sample containing at least one synthetic nucleic acid; and b) removing said at least one synthetic nucleic acid from said sample. thereby forming a sequence analysis sample that contains the target nucleic acid and does not contain the at least one synthetic nucleic acid. and c) attaching a sequencing adaptor to the target nucleic acid in the sequencing sample. In some cases, the removal of the at least one synthetic nucleic acid comprises removing an endonuclease. In some cases, the prior cleavage of the sample is not performed by cleavage. At least one of the synthetic nucleic acids is not linked to another synthetic nucleic acid. The at least one synthetic nucleic acid is resistant to end repair.
[0053] In another aspect, a method for producing a sequencing library is disclosed herein. The method includes: a) obtaining a sample containing a target nucleic acid and at least one synthetic nucleic acid; The sequencing adaptors are attached to the target nucleic acids in the sample, thereby forming a sequencing sample. c) obtaining a pool of the at least one synthetic nucleic acid by affinity-based depletion (removal). , RNA-induced DNase digestion, or a combination thereof, from the sequencing sample. wherein the removal of the at least one synthetic nucleic acid from the sequencing sample comprises: than the sequencing adaptor and than the multimer of the sequencing adaptor. (including preferentially removing at least one synthetic nucleic acid).
[0054] In some cases, the method further comprises endonuclease digestion, size-based removal or or a combination thereof, removing said at least one synthetic nucleic acid. In some cases, the sequencing adapter is a nucleic acid. Removal of another synthetic nucleic acid is performed by affinity-based depletion, In some cases, at least one of the synthetic nucleic acids comprises an immobilization tag. Removal of the RNA-guided DNase is performed by RNA-guided DNase digestion. The ase comprises a CRISPR-associated protein. In some cases, at least one of the The removal of synthetic nucleic acids is accomplished by endonuclease digestion. The removal of at least one synthetic nucleic acid is performed by size-based removal, The synthetic nucleic acid has a length greater than the length of the target nucleic acid. The removal of at least one synthetic nucleic acid is performed with RNase, and the removal of at least one synthetic nucleic acid is performed with RNase. is a DNA-RNA hybrid. In some cases, the sequencing adapter is inserted into the target nucleic acid. Binding to the nucleic acid includes ligating a sequencing adaptor to the target nucleic acid. In some cases, binding the sequencing adaptor to the target nucleic acid may be performed by binding the sequencing adaptor to the target nucleic acid. The method includes linking the nucleic acid to a target nucleic acid.
[0055] In another aspect, a method for producing a sequencing library is disclosed herein. The method includes: a) obtaining a sample comprising a target nucleic acid and at least one synthetic nucleic acid; The at least one synthetic nucleic acid comprises: (i) a single-stranded DNA; (ii) an amplification method for the synthetic nucleic acid; (iii) nucleotide modifications that suppress the binding of nucleotides to immobilized tags; and (iv) DNA-RNA hybridization. (v) a nucleic acid having a length greater than the length of the target nucleic acid, or (vi) any of these. b) preparing a sequencing library from the samples for a sequencing reaction. wherein at least a portion of the at least one synthetic nucleic acid is (not sequenced at the time of this writing).
[0056] In some cases, the at least one synthetic nucleic acid further comprises an endonuclease recognition In some cases, obtaining a sample includes extracting a target nucleic acid from a test sample. and extracting the target nucleic acid from the test sample, and then extracting the at least one In some cases, the method further comprises adding one or more synthetic nucleic acids to a test sample. includes extracting target nucleic acid from a test sample, and further includes extracting target nucleic acid from the test sample. adding said at least one synthetic nucleic acid to the test sample prior to extracting the acid. In some cases, the at least one synthetic nucleic acid may contain a blocking agent that inhibits ligation. In some cases, the blocking group comprises a modified nucleotide. In some cases, the modified nucleotide contains an inverted deoxy-sugar. In some cases, the inverted deoxy-base contains a 3' inverted deoxy-sugar. The thymine includes inverted thymine, inverted adenosine, inverted guanosine, or inverted cytidine. In some cases, the modified nucleotide contains an inverted dideoxy-sugar. The dideoxy-sugar comprises a 5' inverted dideoxy-sugar. In some cases, the modified nucleotide inverted dideoxythymidine, inverted dideoxy-adenosine, inverted dideoxy-guanosine In some cases, the modified nucleotide includes a dideoxy-cytidine or an inverted dideoxy-cytidine. In some cases, the at least one synthetic nucleic acid is a ligation reaction. In some cases, the blocking group includes a spacer. In some cases, the spacer comprises a C3 spacer or a spacer 18. At least one of the synthetic nucleic acids includes a blocking group that inhibits the ligation reaction. In some cases, the synthetic nucleic acid comprises at least one of the above. and a nucleotide modification that inhibits amplification of the synthetic nucleic acid of at least In some cases, the at least one free base site is At least one internal abasic site. In some cases, the nucleotide modification is 8-10 In some cases, the at least one free base site is a single free base site. In some cases, the at least one free base site is a modified ribose In some cases, the at least one free base site is present in a 1',2'- Dideoxyribose, locked nucleic acid, bridged nucleic acid, or twisted internucleotide In some cases, the at least one synthetic nucleic acid is a solid-state nucleic acid. The immobilization tag may be biotin, digoxigenin, polyhistidine or Ni- In some cases, the at least one synthetic nucleic acid is a DNA. In some cases, at least one of the synthetic Nucleic acids are removed from the sequencing sample using a uracil-specific excision reagent enzyme.
[0057] In some cases, the test sample is a biological sample. The biological sample may be whole blood, plasma, serum, or urine. In some cases, the target nucleic acid is acellular. In some cases, the cell-free nucleic acid is cell-free DNA. The cell-free nucleic acid is pathogen nucleic acid. In some cases, the cell-free nucleic acid is circulating cell-free nucleic acid. In some cases, the at least one synthetic nucleic acid comprises a double-stranded nucleic acid. In some cases, the at least one synthetic nucleic acid comprises a single-stranded nucleic acid. At least one of the synthetic nucleic acids may be DNA, RNA, a DNA-RNA hybrid or the like. This includes any analogues thereof.
[0058] In some cases, the method further comprises: (a) extracting a target nucleic acid from the sample; (c) purifying the target nucleic acid from the sample; (d) end-repairing the target nucleic acid; (e) fragmenting the target nucleic acid; (f) amplifying the target nucleic acid; and (g) sequencing adapters. and (g) binding to the target nucleic acid. In some cases, the method includes attaching a sequencing adaptor to the target nucleic acid; Furthermore, the sequencing sample is subjected to endonuclease prior to binding of the sequencing adaptor to the target nucleic acid. In some cases, the method further comprises treating the sequencing adapter with a target enzyme. and binding the sequencing adaptor to the target nucleic acid, and further comprising: This involves treating the sequencing sample with an endonuclease. The method includes end-repairing a target nucleic acid, wherein prior to end-repairing the target nucleic acid, In some cases, the method further comprises adding at least one synthetic nucleic acid to the sample. end-repairing the target nucleic acid, wherein after end-repairing the target nucleic acid, In some cases, the method further comprises adding a sequencing adaptor to the sample. and binding the sequencing adaptor to the target nucleic acid, prior to binding the sequencing adaptor to the target nucleic acid. At least one synthetic nucleic acid as described above is added to the sample. In some cases, The ratio of the concentration of the at least one synthetic nucleic acid to the concentration of the target nucleic acid in the sample is 1. :1~1000:1.
[0059] In some cases, the size of the at least one synthetic nucleic acid and the size of the target nucleic acid are The difference is that the at least one synthetic nucleic acid is separated from the target nucleic acid based on size. In some cases, synthetic nucleic acids contain blocking groups that inhibit ligation reactions, and and nucleotide modifications that inhibit the amplification reaction. In some cases, The blocking group contains a 3' inverted deoxy-T, a nucleotide modification that inhibits the amplification reaction. In some cases, the blocking group further comprises a 5' inverted dideoxy group. In some cases, the method further comprises treating the sample with endonuclease VII. In some cases, the sample is incubated with endonuclease I. In some cases, the method further comprises incubating the sample with ELISA kit (100 μg / mL) for up to 1 hour. and extracting the target nucleic acid from the sample, the extracting the target nucleic acid comprising: Higher yields compared to extracting target nucleic acids from samples that do not contain any synthetic nucleic acids. In some cases, the method includes end repairing the target nucleic acid, The end repair is performed by repeating the steps of: In some cases, the target nucleic acid is The nucleic acid may include a naturally occurring nucleic acid or a copy thereof. In some cases, the method further comprises: The method includes obtaining at least one sequence of the target nucleic acid using a computer.
[0060] In another aspect, a method for producing a sequencing library is disclosed herein. The method includes: (a) providing (i) a target nucleic acid; (ii) a sequencing adaptor; and (iii) at least one synthetic nucleic acid, wherein said at least one synthetic nucleic acid comprises DNA; b) obtaining a sample containing a nucleic acid sequence that is more specific than said at least one synthetic nucleic acid; performing an end repair reaction on the sample such that the target nucleic acid is preferentially end repaired. Includes.
[0061] In some embodiments, any of the above methods further comprise communicating the results of the method to a patient, caregiver, or or reporting the matter to another person.
[0062] In another embodiment, a kit for producing a sequencing library is provided herein. The kit comprises: a) a sequencing adaptor; and b) at least one synthetic Synthetic nucleic acids, wherein said at least one synthetic nucleic acid comprises DNA and no terminal modifications to the nucleic acid. In some cases, the amount and sequence of the at least one synthetic nucleic acid The ratio of the amount of determined adaptor is 1:1 or less.
[0063] The novel features of the disclosed subject matter are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed subject matter may be obtained by practicing the disclosed subject matter in accordance with exemplary embodiments thereof. By reference to the following detailed description and the accompanying drawings, which set forth exemplary embodiments, To be lied to.
[0064] References All publications, patents, and patent applications cited herein are hereby incorporated in their entirety by reference to each individual publication. Each patent or patent application is specifically and individually indicated to be incorporated by reference into this specification. Each of the above-referenced articles is hereby incorporated by reference as if fully set forth herein.
[0065] Detailed Description overview The present disclosure relates to improved methods for sequencing and quantification of nucleic acids in next generation sequencing assays and other assays. In general, the present invention provides a number of methods and approaches for identifying or quantifying The methods provided allow for the selection of specific sequences, lengths, GC content, degeneracy, degree of diversity and / or known This includes the use of spike-in synthetic nucleic acids with specific characteristics, such as known starting concentrations. The use of synthetic spike-in nucleic acids allows for the determination of absolute abundance, the determination of relative abundance, and the determination of abundance. Normalization, universal quantification, bias control, sample identification, cross-contamination detection, information Transduction efficiency, reagent tracking, normalization of diversity reduction, determination of absolute or relative loss, quality control and many other applications may be enabled and improved. The method also includes a specially designed carrier nucleic acid, which is This can increase the total concentration of nucleic acid in the sample but avoid detection by sequencing or other assays. Has the ability.
[0066] In a preferred embodiment, the present disclosure provides a set of spike-in synthetic nucleic acid species. where the length and / or GC content of each species is determined by the sequence of the target nucleic acid to be analyzed. Match or approximate the expected or observable length and / or GC content of the For example, the length of the spike-in synthetic nucleic acid is designed to be such that it is capable of infecting such pathogens. Disease-specific or pathogen-specific antibodies in samples (e.g., plasma) obtained from human patients The length of the specific cell-free nucleic acid can be approximated. In another preferred embodiment, the present disclosure provides , spike-ins containing sequences to uniquely identify samples, reagents, or reagent lots In yet another preferred embodiment, the present disclosure provides a method for the preparation of a synthetic nucleic acid comprising the steps of: A large number of spike-in synthetic nucleic acids (e.g., 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 or 10 10providing a pool containing 10 unique spike-in synthetic nucleic acids, This is especially true during the course of high-throughput sequencing assays, particularly sample processing steps, e.g. Reduced diversity of unique spike-in sequences during extraction and / or library production can be used to track absolute nucleic acid reduction in a sample via
[0067] The ability to track absolute nucleic acid depletion allows for the determination of the absolute abundance of the target nucleic acid in the initial sample. For example, the absolute amount of a pathogen in a clinical sample can be determined by the amount of the pathogen attributable to that pathogen. The therapeutic effect of the antibiotic or pharmaceutical composition may be determined based on the number of sequencing reads that Absolute pathogen abundance in clinical samples taken over time (pre-, post-, and post-transplant) By determining the presence of a particular pathogen, medical treatment can be monitored or adjusted. In addition to determining whether to treat or cure the disease, the degree or stage of infection or disease may also be determined. .
[0068] The method can be used to detect and elucidate clinical samples, processed samples (e.g., extracted nucleic acids, extracted cell-free DNA, extracted cell-free RNA, plasma, serum), unprocessed samples (e.g., whole blood) and any other type A wide variety of samples, including but not limited to samples containing nucleic acids, This may involve adding a spike-in synthetic nucleic acid to such a sample. The method includes the step of preparing a sample for analysis by reagents, in particular by sequencing (e.g., next generation sequencing). The spike-in synthetic nucleus is added to the test reagent (or a specific reagent lot) used at any stage. In a preferred embodiment, the method comprises adding a known concentration of synthetic nuclei. The method may include introducing an acid into the reagent and the sample. To detect, identify, monitor or quantify low abundance pathogens or nucleic acids derived from pathogens. The method may be particularly useful for increasing the accuracy and efficiency of assays designed to Also, errors in sample tracking or sample preparation may occur. Certain defects arise from the uneven depletion of nucleic acid sequences during nucleic acid purification or sequencing library preparation. Or the lack of an internal normalization standard when comparing analyses of different target nucleic acids or different samples? This can reduce undesirable effects arising from
[0069] FIG. 1 provides a general overview of the various steps of the methods provided herein, particularly with respect to abundance normalization. The method can include obtaining a sample from a subject 110 (e.g., a human patient). In some specific embodiments, the subject has an infection or is infected with a pathogen. The sample may be a blood sample 120 or a plasma sample, as shown. or any other type of biological sample, in particular body fluids, tissues and / or Alternatively, it may be a biological sample containing cells, or alternatively, a cell-free biological sample.
[0070] Nucleic acid (e.g., cell-free nucleic acid) from the sample 140 can be extracted and assayed, e.g., sequenced. It can be used in quantitative assays (e.g., next-generation sequencing assays). One or more types of synthetic nucleic acid 150 may be synthesized in one or more steps of the method, e.g. Add (or saturate) the blood sample 120, plasma sample 130 or sample nucleic acid 140 The synthetic nucleic acid is approximately the same length as the set of target nucleic acids to be analyzed. and / or approximate the GC content of the set of target nucleic acids to be analyzed. Generally, the synthetic nucleic acid also has a known starting concentration. The sample containing the synthetic nucleic acid is then subjected to a sequencing assay 160, e.g., a next generation sequencing assay. In some cases, the nucleic acid can be analyzed by sequencing assays. The amount of synthetic nucleic acid identified by (a) is compared with the known starting concentration of the synthetic nucleic acid to determine the number of reads. As a result, the abundance of the detected target nucleic acid is correlated with the length and concentration of the target nucleic acid. and / or GC content relative to the abundance of the synthetic nucleic acid that is closest to such target nucleic acid 170. By comparing the nucleic acid sequence, it is possible to identify or quantify the target nucleic acid in the sample nucleic acid. By use of such methods and other methods provided herein, the condition of a subject can be improved. It is possible to determine with a high level of accuracy and certainty. In this study, sequencing assays (e.g., next-generation sequencing assays) were performed using human patient-derived Detect pathogen nucleic acids in a sample of cell-free nucleic acid (e.g., DNA).
[0071] The steps may be performed in any order and in any combination. In some cases, new steps may be added to the steps shown, or It is interposed between the steps shown.
[0072] FIG. 2 shows a schematic diagram of a typical infection. The source of pathogen infection can be, for example, in the lungs. Cell-free nucleic acids derived from pathogens, such as cell-free DNA, are transported through the bloodstream. The nucleic acids in the sample can then be separated by It can be analyzed by sequencing assay as shown in FIG.
[0073] FIG. 3 shows a general scheme of some of the methods provided herein. The methods include: Obtaining a sample containing host (e.g., human) and non-host (e.g., pathogen) nucleic acids. The sample may be obtained from a subject, e.g., a patient. In some particular embodiments, In this case, the subject has an infectious disease or is suspected of being infected with a pathogen. a sample or plasma sample, or any other type of biological sample, in particular a body fluid, The sample may be a biological sample containing tissues and / or cells. For example, cell-free nucleic acid can be combined with a known amount of synthetic nucleic acid. The sample containing the nucleic acid is subjected to a sequencing assay, e.g., a next generation sequencing assay. Sequencing results can be analyzed against known host and non-host reference sequences. In some cases, sequences can be identified by sequencing assays. The amount of synthetic nucleic acid extracted is compared to a known starting concentration of the synthetic nucleic acid, and the number of reads is calculated based on the known starting concentration. As a result, the relative abundance of non-host sequences can be determined. The steps may be performed in any order and in any combination. In some cases, certain steps may be repeated several times. In some cases, certain steps may not be performed. In some cases, the steps shown may not be performed. New steps may be added to or interposed between the steps shown.
[0074] The methods provided herein are useful when the target nucleic acid is present in a sample at low abundance or when multiple Next-generation sequencing is particularly useful when comparing or tracking multiple samples or multiple target nucleic acids. For example, next-generation sequencing may allow for improved identification or quantification of target nucleic acids. - Target pathogens, tumor cells or tumorigenic markers in clinical samples by sequencing Accurate detection and quantification of the target nucleic acid may be compromised if the sample is improperly tracked or if the target nucleic acid is malformed. If the data is accurately normalized or quantified, it may be compromised or negatively affected. Thus, the methods provided herein are useful in sample tracking or identification or in the detection of nucleic acids. Avoiding pitfalls arising from errors in crowd-sourced analysis of quantitative or sequencing data This can help.
[0075] The methods and compositions provided herein are useful when the starting sample contains relatively small amounts of nucleic acid. In particular, in order to improve the yield, quality or efficiency of the sequencing library, It can be used to add and / or remove synthetic nucleic acids during library production. Generally, in some cases, the synthetic nucleic acid acts as a carrier nucleic acid in these applications. In addition, the concentration of total nucleic acid may be increased during the sample preparation process. The addition of One or more of the steps may be sensitive to nucleic acid concentration. For example, the yield of the step may be and / or the efficiency may depend on the concentration of nucleic acid in the sample. This may include extraction, purification, ligation and end repair. In some cases, the synthetic nucleic acid is The synthetic nucleic acid may be removed from the sequencing library. These compounds may contain certain features that prevent them from participating in one or more steps of the synthesis. The synthetic nucleic acid may not be sequenced in the sequencing step.
[0076] The methods and compositions are useful for analyzing samples from multiple subjects (e.g., These can be used to prepare a sequencing library from the target nucleic acid in the The concentration of target nucleic acid in each sample may vary from subject to subject. The addition of the synthetic nucleic acid to the pool may reduce concentration variation between samples, improving the accuracy of the analysis. .
[0077] The methods and compositions include the step of: extracting nucleic acids from a sample by adding at least one synthetic nucleic acid; The synthetic nucleic acid can be used to prepare a sequencing library. In some cases, the nucleic acid sequence may have one or more characteristics that prevent them from being sequenced. The synthetic nucleic acid may be used in one or more reactions in the production of a sequencing library, such as adapter ligation. For example, the nucleic acid may have an inversion at one or both ends, which inhibits ligation and nucleic acid amplification. It may contain sugars and / or one or more abasic sites.
[0078] In some cases, the synthetic nucleic acid may be removed from the sequencing library prior to sequencing. In some cases, the synthetic nucleic acid can be removed by enzymatic digestion. For example, the synthetic nucleic acid can be It may contain a restriction enzyme recognition site and can be digested by a restriction enzyme. Alternatively, the synthetic nucleic acid can be removed by affinity-based depletion. The acid can contain one or more immobilized tags and can be removed by affinity-based depletion. In some cases, the synthetic nucleic acid can be removed by size-based removal. The synthetic nucleic acid may also have a different size than other molecules in the sequencing library. It is possible, so that the synthetic nucleic acids can be removed by size-based removal. In the case of, the synthetic nucleic acid comprises a combination of features and / or modifications as described herein. so that they are not involved in one or more steps of the production of the sequence library, It may be removed prior to sequencing.
[0079] sample The methods provided herein may allow for improved analysis of a wide variety of samples. The synthetic nucleic acids provided in can be used to analyze such samples, which include Extracting cell-free nucleic acid from a sample or a processed form of the sample, e.g., a clinical plasma sample This may involve adding the synthetic nucleic acid directly to the
[0080] The samples analyzed in the methods provided herein are preferably from any type of clinical In some cases, the sample contains cells, tissue, or body fluid. In a preferred embodiment, the sample is a liquid or fluid sample. The sample may be a body fluid, e.g., whole blood, plasma, serum, urine, stool, saliva, lymphatic fluid, cerebrospinal fluid, synovial fluid, Contains bronchoalveolar lavage fluid, nasal swabs, respiratory secretions, vaginal fluid, amniotic fluid, semen or menstrual fluid. In some cases, the sample may consist, in whole or in part, of cells or tissue. In some cases, the cells, cell fragments, or exosomes are separated by, for example, centrifugation or The sample is removed by filtration. In this specification, the sample is a biological sample. It's possible.
[0081] The sample may contain any concentration of nucleic acid. The compositions and methods herein are directed to low concentrations. It may be useful for samples that contain a high degree of total nucleic acid. Both 100ng / μL, 50ng / μL, 10ng / μL, 5ng / μL, 2ng / μL , 1.5ng / μL, 1.2ng / μL, 1ng / μL, 0.8ng / μL, 0.4ng / μL, 0.2ng / μL, 0.1ng / μL, 0.05ng / μL, 0.01ng / μL L, 10ng / mL, 5ng / mL, 2ng / mL, 1ng / mL, 0.8ng / mL, Have a total concentration of nucleic acid of 0.6 ng / mL, 0.5 ng / mL or 0.1 ng / mL. In some cases, the sample has a concentration of at least 0.1 ng / mL, 0.5 ng / mL, 0. 6ng / mL, 0.8ng / mL, 1ng / mL, 2ng / mL, 5ng / mL, 10n g / mL, 0.01ng / μL, 0.05ng / μL, 0.1ng / μL, 0.2ng / μL, 0.4ng / μL, 0.8ng / μL, 1ng / μL, 1.2ng / μL, 1.5 ng / μL, 2ng / μL, 5ng / μL, 10ng / μL, 50ng / μL or 10 In some cases, the samples contained a total concentration of nucleic acid of about 0.1 ng / mL. ~10,000ng / mL (i.e., about 0.1ng / mL to about 10ng / μL) The total concentration of nucleic acid within the range.
[0082] The sample may include one or more controls. In some cases, the sample may include one or more negative controls. Typical negative controls include samples prepared to identify contaminants (e.g., plasma minus samples), plasma from healthy subjects, and low diversity samples (e.g., In some cases, the sample may be one or more A typical positive control is a healthy subject with genomic DNA from a known pathogen. Samples from subjects (e.g., plasma samples) include genomic DNA from known pathogens. It may be the complete genomic DNA. In some cases, the genomic DNA from a known pathogen may be For example, the material may be sheared to various average lengths. The shearing may be mechanical shearing (e.g., ultrasonic, hydrodynamic, etc.). mechanical shearing forces), enzymatic shearing (e.g., endonucleases), thermal degradation (e.g., high by incubation at room temperature), chemical fragmentation (e.g., alkaline solution, divalent ions) This can be done.
[0083] The sample may contain a target nucleic acid. Target nucleic acid refers to the nucleic acid to be analyzed in the sample. For example, the target nucleic acid can be naturally occurring in the sample, e.g., naturally occurring The sample may further comprise one or more synthetic nucleic acids as disclosed herein. In some cases, the target nucleic acid is a cell-free nucleic acid as described herein. For example, the target nucleic acid may be cell-free DNA, cell-free RNA (e.g., cell-free mRNA, cell-free miRNA, etc. A, cell-free siRNA), or any combination thereof. , cell-free nucleic acid is pathogen nucleic acid, e.g., nucleic acid from a pathogen. Cell-free nucleic acid is circulating nucleic acid, For example, it can be circulating tumor DNA or circulating fetal DNA. The sample can be a sample containing a pathogen, e.g. It may include nucleic acids from viruses, bacteria, fungi and / or eukaryotic parasites.
[0084] In some cases, the sample also contains an adaptor. The adaptor may be a known or unknown The adaptor may be a nucleic acid having a sequence. The adaptor may be at the 3' end, the 5' end, or both ends of the nucleic acid. The adaptor may comprise a known sequence and / or an unknown sequence. The adaptor may be double-stranded or single-stranded. In some cases, the adaptor is a sequencing adaptor. The sequencing adapter can bind to a target nucleic acid and aid in the sequencing of the target nucleic acid. For example, a sequencing adapter may contain a sequencing primer binding site, a unique identifier sequence, a non-unique and one or more of a unique identifier sequence, a sequence for immobilizing the target nucleic acid on a solid support. The target nucleic acid that is bound to the sequencing adaptor is fixed on a solid support on the sequencer. The sequencing primer hybridizes to the adapter and is then used in the sequencing reaction. In some cases, the target nucleic acid can be extended using the adaptor as a template. The identifiers in the target sequence are used to label sequence reads of different target sequences, thereby High-throughput sequencing of nucleic acids becomes possible.
[0085] The term "bond" and its grammatical equivalents refer to any bond used to connect two molecules. For example, binding can mean joining two molecules together by a chemical bond or other means. The binding of an adaptor to a nucleic acid can mean the synthesis of a new molecule. In some cases, this can mean forming a chemical bond between the adaptor and the nucleic acid. The attachment is performed, for example, by ligation using a ligase. For example, the nucleic acid adapter can be Ligation occurs by forming a phosphodiester bond catalyzed by ligase. It can bind to a target nucleic acid.
[0086] Sequencing libraries can be generated from samples using the methods and compositions provided herein. The sequencing library can be prepared from multiple sequences compatible with the sequencing system used. For example, the nucleic acids in the sequencing library may include one or more adaptor The process for producing a sequencing library may include a sample of the target nucleic acid bound to the target nucleic acid. The target nucleic acid is extracted from the sample, the target nucleic acid is fragmented, and an adapter is attached to the target nucleic acid. amplifying the target nucleic acid-adaptor complex; and -sequencing the adaptor complex.
[0087] A sample (especially a cell sample or tissue biopsy) from any part or area of the body. Exemplary samples can be, for example, blood, central nervous system, brain, spinal cord, bone marrow, pancreas, , thyroid gland, gallbladder, liver, heart, spleen, colon, rectum, lungs, respiratory system, throat, nasal passages, stomach, esophagus , ears, eyes, skin, limbs, uterus, prostate, genitals or any other organ or area of the body. can be obtained from.
[0088] Generally, the sample is from a human subject, particularly a human patient. However, the sample may also be or any other type of subject, for example any mammal, non-human mammal, non-human primate, Domestic animals (e.g., laboratory animals, household pets or livestock) or non-domestic animals (e.g., wild In some specific embodiments, the subject may be a dog, a cat, or a mammal. , rodents, mice, hamsters, cows, birds, chickens, pigs, horses, goats, sheep, a rabbit, ape, monkey or chimpanzee.
[0089] In a preferred embodiment, the subject is infected with or suffering from a pathogen. in a host organism (e.g., a human) that is at risk of infection or is suspected of harboring a pathogen infection. In some cases, a subject may be suspected of having a particular infection, e.g., a subject may have tuberculosis. In other cases, subjects are suspected of having an infection of unknown origin. In some cases, the host or subject may be infected with one or more microorganisms, pathogens, bacteria, viruses, or viruses. In some cases, the host or subject is infected with Have been diagnosed with one or more types of cancer or are at risk of developing one or more types of cancer In some cases, the host or subject may be infected with one or more microorganisms, pathogens, bacteria, not infected with a virus, fungus, or parasite. In some cases, the host or subject In some cases, the host or subject is susceptible to infection or is in a state where the infection is prevented or controlled. There is a risk.
[0090] In some cases, the subject is an antimicrobial, antibacterial, antiviral, or antiparasitic agent. The subject may be in the process of being treated or may be capable of being treated. In some cases, the patient may have an actual infection (bacterial, viral, fungal or parasitic). The subject is (e.g., infected with one or more microorganisms, pathogens, bacteria, viruses, fungi, or parasites). In some cases, the subject is healthy. In some cases, the subject is infected. susceptible to or at risk for infection (e.g., the patient is immunocompromised) The subject may have, or be at risk for, another disease or disorder. For example, the subject The patient may have or have a disease, such as cancer (e.g., breast cancer, lung cancer, pancreatic cancer, blood cancer, etc.). The patient may be at risk for or suspected of having such a disease.
[0091] The sample can be a nucleic acid sample. In some cases, the sample contains a quantity of nucleic acid. The nucleic acids in the sample can be double-stranded (ds) nucleic acids, single-stranded (ss) nucleic acids, DNA, and RNA. , cDNA, mRNA, cRNA, tRNA, ribosomal RNA, dsDNA, ssDNA A, miRNA, siRNA, circulating nucleic acid, circulating cell-free nucleic acid, circulating DNA, circulating RNA, no Cellular nucleic acid, cell-free DNA, cell-free RNA, circulating cell-free DNA, cell-free dsDNA, cell-free ssDNA, circulating cell-free RNA, genomic DNA, exosome, cell-free pathogen nucleic acid, circulating Pathogen nucleic acid, mitochondrial nucleic acid, non-mitochondrial nucleic acid, nuclear DNA, nuclear RNA, chromosome DNA, circulating tumor DNA, circulating tumor RNA, circular nucleic acid, circular DNA, circular RNA, circular The DNA may include single stranded DNA, circular double stranded DNA, plasmids, or any combination thereof. In some cases, the sample nucleic acid may include synthetic nucleic acid. , any type of nucleic acid disclosed herein, e.g., DNA, RNA, DNA-RNP A hybrid is included. For example, the synthetic nucleic acid can be DNA.
[0092] In some cases, different types of nucleic acids may be present in a sample. The samples can include cell-free RNA and cell-free DNA. This includes methods to analyze both RNA and DNA present in a sample, either alone or in combination. I can see it.
[0093] As used herein, the term "acellular" refers to a substance that is present in the body before a sample is obtained from the body. For example, circulating cell-free nucleic acid in a sample refers to the state of nucleic acid circulating in the bloodstream of a human body. They can originate from circulating cell-free nucleic acid. In contrast, they can be extracted from solid tissues such as biopsies. The released nucleic acid is generally not considered to be "cell-free."
[0094] In some cases, the sample is a processed sample containing cell-free or cell-associated nucleic acid. The sample may be a purified sample (e.g., serum, plasma) or an unprocessed sample (e.g., whole blood). In some cases, the sample may contain some type of nucleic acid, such as DNA, RNA, cell-free DNA, cell It is enriched for cellular RNA, cell-free circulating DNA, cell-free circulating RNA, etc. In some cases, the sample is subjected to a step of isolating nucleic acids or separating the nuclei from other components in the sample. In some cases, the samples are treated in some way to separate the acid. It is enriched for drug-specific nucleic acids.
[0095] Often the sample is a fresh sample. In some cases, the sample is a frozen sample. In some cases, the sample is, for example, formalin-fixed paraffin-embedded tissue. As such, the cells are fixed, for example, with a chemical fixative.
[0096] target nucleic acid The methods provided herein can be used to detect a large number of target nucleic acids. can be used to analyze whole or partial genomes, exomes, loci, genes, exons, and introns. , modified nucleic acids (e.g., methylated nucleic acids), and / or mitochondrial nucleic acids. Often, but not limited to, the methods provided herein involve the use of a pathogen target nucleic acid. In some cases, the pathogen target nucleic acid can be extracted from a subject's nucleic acid. Pathogen target nucleic acids are present in complex clinical samples containing acids. Influenza, tuberculosis, or any other known infectious disease or disorder (as described in more detail herein) In some cases, the The target nucleic acid described herein may be a target nucleic acid.
[0097] In some cases, the pathogen target nucleic acid is derived from a tissue sample, e.g., a tissue sample from a site of infection. In other cases, pathogen target nucleic acid has migrated from the site of infection and is present in the For example, it may be obtained from a sample containing circulating cell-free nucleic acid (e.g., DNA). do.
[0098] In some cases, the target nucleic acid is derived from cancer tissue. The target nucleic acid is extracted directly from the tissue or tumor. In some cases, the target cancer nucleic acid can be obtained from circulating cell-free nucleic acid or circulating tumor cells (CT C) is obtained.
[0099] In some cases, the target nucleic acid is present in only a very small portion of the total sample, e.g. Less than 1%, less than 0.5%, less than 0.1%, less than 0.01%, less than 0.0 Less than 0.01%, less than 0.0001%, less than 0.00001% or less than 0.0000001% In some cases, the target nucleic acid may comprise less than about 0.000 of the total nucleic acid in the sample. 0.1% to about 0.5%. Often, the total nucleic acid in the original sample can vary. For example, total cell-free nucleic acids (e.g., DNA, mRNA, RNA) are between 1 and 100 ng / ml. (e.g., in the range of about 1, 5, 10, 20, 30, 40, 50, 80, 100 ng / ml) In some cases, the total concentration of cell-free nucleic acids in the sample may be outside this range. (e.g., less than 1 ng / ml; in other cases, the total concentration exceeds 100 ng / ml This is because cell-free nucleic acids (e.g. This may be the case for DNA (DNA) samples. In such samples, The target nucleic acid or cancer target nucleic acid may be, for example, , may be insufficiently present compared to human or healthy nucleic acids. For example, pathogens The target nucleic acid may comprise less than 0.001% of the total nucleic acid in the sample, The target nucleic acid may constitute less than 1% of the total nucleic acid in the sample.
[0100] The length of the target nucleic acid can vary. In some cases, the target nucleic acid is at least about 20 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140 , 150, 160, 170, 180, 190, 200, 250, 300, 350, 400 , 450, 500, 750, 1000, 1500, 2000, 3000, 4000, 50 00, 10000, 15000, 20000, 25000 or 50000 nucleotides In some cases, the target nucleic acid may be at most about 20, 30, 40, 50, 60, 70, 80, 90, 100, 15 ...200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 1000 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 15 0, 160, 170, 180, 190, 200, 250, 300, 350, 400, 45 0, 500, 750, 1000, 1500, 2000, 3000, 4000, 5000, 10000, 15000, 20000, 25000 or 50000 nucleotides (also In some particular embodiments, the target nucleic acid is relatively short, For example, less than 500 base pairs (or nucleotides) in length or less than 1000 base pairs (or nucleotides). In some cases, the target nucleic acid is relatively long, e.g., less than 1000 nucleotides in length. More than 1500 base pairs (or nucleotides) long More than 2000 base pairs (or nucleotides) long, More than 2500 base pairs (or nucleotides) long more than 3,000 base pairs (or nucleotides) in length, or more than 5 000 base pairs (or nucleotides) in length. In some cases, the target nucleic acid is about 20 In some cases, the target nucleic acid may range from about 40 to about 100 base pairs. It may be a range of base pairs.
[0101] As with the sample nucleic acid, the target nucleic acid may be a double-stranded (ds) nucleic acid or a single-stranded (ss) nucleic acid. , DNA, RNA, cDNA, mRNA, cRNA, tRNA, ribosomal RNA, ds DNA, ssDNA, miRNA, siRNA, circulating nucleic acid, circulating cell-free nucleic acid, circulating DNA , circulating RNA, cell-free nucleic acid, cell-free DNA, cell-free RNA, circulating cell-free DNA, cell-free d sDNA, cell-free ssDNA, circulating cell-free RNA, genomic DNA, exosomes, cell-free Pathogen nucleic acid, circulating pathogen nucleic acid, mitochondrial nucleic acid, non-mitochondrial nucleic acid, nuclear DNA, nuclear RNA, chromosomal DNA, circulating tumor DNA, circulating tumor RNA, circular nucleic acid, circular DNA, circular RNA, circular single-stranded DNA, circular double-stranded DNA, plasmid, or any combination thereof The target nucleic acid may be any type of nucleic acid, including a nucleic acid sequence. The target nucleic acid is preferably a viral, bacterial, or bacterial vector. Bacteria, parasites and any other microorganisms, especially infectious microorganisms (including but not limited to The target nucleic acid is a nucleic acid derived from a pathogen, including those derived from a specific organ or tissue. In some cases, the target nucleic acid may be a nucleic acid derived directly from the subject rather than from a pathogen. Be guided.
[0102] Spike-in synthetic nucleic acid The present disclosure relates to various methods, particularly those related to high throughput or next generation sequencing assays. The present invention describes single synthetic nucleic acids and sets of synthetic nucleic acids for use in the present invention. When a spike-in synthetic nucleic acid is used in the described method, for example, it is derived The individual, pre-analytical sample treatment conditions, methods of nucleic acid extraction, molecular biology tools and methods Regardless of the nucleic acid manipulations, the method of nucleic acid purification, the measurement itself, storage conditions, and time course of the assay, Allows efficient normalization of nucleic acids across samples (e.g., disease-specific nucleic acids, pathogen nucleic acids) In some cases, the present disclosure provides a method for identifying specific features, such as a number of unique sequences. The set of synthetic nucleic acids is provided for sample analysis. can be used to monitor diversity reduction over the course of The synthetic nucleic acids provided herein can also be used to determine the abundance of nucleic acids in a sample. To track samples, to monitor cross-contamination between samples, to track reagents, They can be used to track reagent lots and for many other purposes. The design, length, quality, concentration, diversity level and sequence of the nucleic acid can be adapted for a particular application. In some cases, the spike-in synthetic nucleic acid may include a carrier synthetic nucleic acid ( For example, carrier synthetic nucleic acids are included.
[0103] The collection (or set) of synthetic nucleic acids provided herein includes several types of In some cases, the species may be the same in length, concentration and / or sequence. In some cases, the length, concentration and structure of the species may be similar. and / or the sequences may be different.
[0104] In a preferred embodiment, the synthetic nucleic acid species vary in length. The collection of nucleic acid species as a whole represents the observable range of lengths of a given target nucleic acid in a sample. or at least a portion of such observable range. In particular, samples obtained from subjects infected or suspected of being infected with a pathogen The length of disease-specific or pathogen-specific nucleic acid in the sample may be several In this case, the length of the disease-specific or pathogen-specific nucleic acid in the sample is about 40 to about 1 In some cases, the species may be in the range of 0.000 base pairs in the sample. A wide variety of disease-specific or pathogen-specific nucleic acid lengths may be involved. The species as a whole represent a particular pathogen-specific nucleic acid, e.g., a length of nucleic acid within a particular pathogen genome. In some cases, the nucleic acid may be a specific nucleic acid within a pathogen genome, e.g., Nucleic acids within virulence regions of a pathogen, antibiotic resistance regions of a pathogen, or other regions or It may be a specific nucleic acid or gene. In some cases, the length or nucleic acid is determined by the individual In other examples, the method may be specific to the type of tumor (e.g., acute, chronic, active, or latent). In this case, the species is generally a subset of a given target in a sample (e.g., from an infected subject). It may span the length of the nucleic acid and / or the pathogen nucleic acid.
[0105] The length of the synthetic nucleic acid seeds in the collection is determined by the length of the specific target nucleic acid (e.g., the length of the pathological site in the sample). In other cases, the range of observable drug-specific or disease-specific nucleic acids may be closely aligned. In the case of a synthetic nucleic acid, the length of the synthetic nucleic acid species in the synthetic nucleic acid collection exactly matches the length of the target nucleic acid, and For example, the length of the synthetic nucleic acid species may correspond to the length of the target nucleic acid. 50% to 150% of the length of the target nucleic acid, 55% to 145% of the length of the target nucleic acid, 60% to 140%, 65%-135% of the length of the target nucleic acid, 70%-130% of the length of the target nucleic acid, 75%-125% of the length of the target nucleic acid, 80%-120% of the length of the target nucleic acid, 85% to 115% of the length of the target nucleic acid, 90% to 110% of the length of the target nucleic acid, 95% to 1 05%, 96%-104% of the length of the target nucleic acid, 99%-101% of the length of the target nucleic acid, or is within the range of 99.5% to 100.5% of the length of the target nucleic acid. The length of the nucleic acid seed can be within the range of 50% to 150% of the length of the target nucleic acid. In some embodiments, the length of the synthetic nucleic acid species can be up to 2, 3, 4 or 5 times the length of the target nucleic acid. In some cases, the length of the synthetic nucleic acid seed is 1, 2, 3, 4, 5, ...2, 3, 4, 5, 1, 2, 3, 4, 5, 0, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150 or In some cases, the species of synthetic nucleic acid in the collection may be up to 200 nucleotides long. The length of the target nucleic acid was exactly matched at 65%, 75%, 80%, 85%, 90%, 92%, Greater than 95%, 97% or 99%.
[0106] Each or most of the nucleic acids in the collection (or pool) of synthetic nucleic acids disclosed herein Any nucleic acid "species" can contain one or more domains or regions of interest. The domain or region of interest is a length identifier sequence. It may contain a code that is predetermined to indicate or represent a length. Often, such a length identifier is The identifier can be a short sequence, e.g., 10 base pairs (bp), 9 bp, 8 bp , 7bp, 6bp, 5bp, 4bp or 3bp; <9bp, <8bp, <7bp or less than 6bp; or 6–15bp, 5–10bp, 4–8bp, or 6–9bp p. The species may contain one, two or more length identifier sequences. In the case of , the length identifier is present as a forward and / or reverse sequence.
[0107] In some cases, the domains in the nucleic acid species in the collection of synthetic nucleic acids are A specific length that generally corresponds to the length encoded by the length identification sequence in the synthetic nucleic acid. The length of the spike-in nucleic acid or load may vary. In some cases, the entire spike-in nucleic acid may be at least about 20, 30, 40, 50 , 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160 , 170, 180, 190, 200, 250, 300, 350, 400, 450 or 5 In some cases, the spike-in nucleic acid may be at most about 20 nucleotides long. , 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 14 0, 150, 160, 170, 180, 190, 200, 250, 300, 350, 40 In some cases, the spike-in nucleus may be 0, 450 or 500 nucleotides long. The acid may range from about 20 to about 200 base pairs, for example, from about 20 to about 120 base pairs. In the case of, the length of the loading sequence domain in the spike-in nucleic acid is at least about 20, 30 , 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 1 50, 160, 170, 180, 190, 200, 250, 300, 350, 400, 4 It may be 50 or 500 nucleotides in length. In some cases, The length of the loading sequence domain is at most about 20, 30, 40, 50, 60, 70, 80, 9 0, 100, 110, 120, 130, 140, 150, 160, 170, 180, 19 0, 200, 250, 300, 350, 400, 450 or 500 nucleotides in length. In some cases, the length of the loading sequence domain in the spike-in nucleic acid may range from 0 to about 2 00 bp.
[0108] Domains in a nucleic acid species within a collection of synthetic nucleic acids indicate that the nucleic acid is not part of the original sample. A synthetic nucleic acid identifier sequence containing a unique code that indicates that the nucleic acid is a spike-in [e.g., Spark Recognition Sequence, Spank Recognition Sequence. The unique code is a code that is not present in the original sample or in the pool of target nucleic acids. The synthetic nucleic acid identifier sequence may be a specific number of bp, e.g., 25 bp, 20 bp, 19 bp, 18 bp The species may include one, two, three, four, six, five, six, seven, seven, twelve, ten, or other lengths. The nucleic acid sequence may contain one or more synthetic nucleic acid identifier sequences or domains. The synthetic nucleic acid identifier sequence is present as a forward and / or reverse sequence.
[0109] In some cases, a domain in a nucleic acid species within a collection of synthetic nucleic acids may be a domain that is a whole of the synthetic nucleic acid. A diversity code domain may be a "diversity code" that is associated with a global pool or collection. can be a unique code that indicates the amount of diversity in the pool of synthetic nucleic acids. Each synthetic nucleic acid in the diverse pool is selected based on the degree of diversity of the pool (e.g., 10 8 Individual Uni In some cases, for example, two or more diversity pools may be used. When multiple pools are used for the same sample, the diversity code is used to identify two or more pools. This can be used to identify reduced diversity in a sample.
[0110] In some cases, the domains in the nucleic acid species in the collection of synthetic nucleic acids are, depending on the application, It may be a feature domain related to one or more of the characteristics of the sample or reagent. For example, A domain can be a specific reagent, a specific reagent lot, or a specific sample (e.g., sample no. number, patient number, patient name, patient age, patient sex, patient race, sample was obtained from patient The nucleic acid sequence may include a sequence coded to indicate the location of the nucleic acid sequence.
[0111] The domains or regions of interest may be present in any combination and number. The nucleic acid may include one or more length identifier sequences, one or more load sequences, one or more synthetic nucleic acid identifier sequences, one or more Any combination or ratio of the diversity codes and / or one or more feature domains listed above. For example, in some cases, the synthetic nucleic acid contains a length identifier sequence and a loading sequence. In some cases, the synthetic nucleic acid comprises a synthetic nucleic acid identifier sequence and a feature domain sequence. In some cases, the synthetic nucleic acid includes a synthetic nucleic acid identifier sequence, and in other cases, it It does not contain a sequence such as:
[0112] In some cases, synthetic nucleic acids may contain domains with overlapping purposes. For example, In some cases, the synthetic nucleic acid may include one or more length identifier sequences that also function as loading sequences. In some cases, the length identifier sequence and / or the load sequence are synthetic nucleic acid identifiers. It also functions as a separator array.
[0113] The synthetic or spike-in nucleic acids are selected or designed to fit into the nucleic acid library. In some cases, the synthetic nucleic acid or spike-in may include an adapter, a common sequence, a label, or a nucleic acid sequence. Random sequences, poly(A) tails, blunt or ragged ends or In some cases, the synthetic nucleic acid or spike-in may contain any combination of these. These or other properties are designed to mimic the nucleic acid in the sample. .
[0114] The synthetic nucleic acids provided herein (e.g., spike-in synthetic nucleic acids) can be any type of nucleic acid, or a combination of nucleic acid types. In a preferred embodiment, synthetic or spa The spike-in nucleic acid is DNA. In some cases, the synthetic or spike-in nucleic acid is single-stranded. In some cases, the synthetic or spike-in nucleic acid is double-stranded DNA. In some cases, the synthetic or spike-in nucleic acid is RNA. The synthetic or spike-in nucleic acid may contain modified or artificial bases. The spike-in nucleic acid may have blunt or recessed ends. In some cases, the synthetic nucleic acid may be double-stranded (ds) Nucleic acid, single-stranded (ss) nucleic acid, DNA, RNA, cDNA, mRNA, cRNA, tRNA , ribosomal RNA, dsDNA, ssDNA, snRNA, genomic DNA, oligonucleotides leotide, double-stranded oligonucleotide, longer assembled double-stranded DNA A (e.g., gBlock from Integrated DNA Technologies) s), plasmids, PCR products, in vitro synthesized transcripts, viral particles, fragments or Unfragmented genomic DNA, circular nucleic acid, circular DNA, circular RNA, circular single-stranded DNA, circular double-stranded DNA Synthetic nucleic acids may contain single stranded DNA, plasmids, or any combination thereof. For example, nucleic acid bases such as adenine (A), cytosine (C), guanine (G), and thymine (T) and / or uracil (U).
[0115] The synthetic nucleic acid can be any synthetic nucleic acid or nucleic acid analog, or any Synthetic nucleic acids may include nucleic acid analogs. Synthetic nucleic acids may have modified or altered phosphate backbones, modified peptides, intose sugars (e.g., modified ribose or deoxyribose), or modified or altered Nucleobases (e.g., modified adenine (A), cytosine (C), guanine (G), thymine (T In some cases, synthetic nucleic acids may contain one or more modified bases, e.g. For example, 5-methylcytosine (m5C), pseudouridine (Ψ), and dihydrouridine (D). , inosine (I) and / or 7-methylguanosine (m7G). In some cases, synthetic nucleic acids are peptide nucleic acids (PNAs), bridged nucleic acids (BNAs), analogous nucleic acids, glycerols, etc. Gol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), 2'-O -methyl substituted RNA, morpholino, or other synthetic polymers with nucleotide side chains In some cases, the synthetic nucleic acid may be DNA, RNA, PNA, LNA, BNA or In some cases, the synthetic nucleic acid may be a double helix or a triple helix. It may comprise a double helix or other structure.
[0116] Synthetic nucleic acids may contain any combination of any nucleotides. The nucleotides may be naturally occurring or can be synthetic. In some cases, the nucleotides are oxidized or methylated. Nucleotides may include, but are not limited to: Adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine GDP, guanosine triphosphate (GTP), thymidine monophosphate (UTP), Uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP ), cytidine diphosphate (CDP), cytidine triphosphate (CTP), 5-methylcytidine monophosphate phosphate, 5-methylcytidine diphosphate, 5-methylcytidine triphosphate, 5-hydroxymethylcytidine Cytidine monophosphate, 5-hydroxymethylcytidine diphosphate, 5-hydroxymethyl Cytidine triphosphate, cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (c GMP), deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (d ADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (d GMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (d GTP), deoxythymidine monophosphate (DTMP), deoxythymidine diphosphate (dTD P), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP) , deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), Deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP) and deoxycytidine diphosphate (dCDP) Deoxycytidine triphosphate (dCTP), 5-methyl-2'-deoxycytidine monophosphate, 5-Methyl-2'-deoxycytidine diphosphate, 5-Methyl-2'-deoxycytidine triphosphate Phosphate, 5-hydroxymethyl-2'-deoxycytidine monophosphate, 5-hydroxymethyl 5-Hydroxymethyl-2'-deoxycytidine diphosphate and 5-Hydroxymethyl-2'-deoxycytidine diphosphate Gin triphosphate.
[0117] Synthetic or spike-in nucleic acid can refer to any molecule that is added to a sample. For example, the present invention is not limited to molecules that are chemically synthesized on a column. Synthetic or spike-in nucleic acids can be prepared, for example, by PCR amplification, in vitro transcription, or template-based synthesis. In some cases, the synthetic or spike-in nucleic acid may be synthesized by other replication. It is or includes sheared or fragmented nucleic acid. The sheared or fragmented nucleic acid can include genomic nucleic acid, e.g., human or pathogen genomic nucleic acid. In some cases, the synthetic nucleic acid does not contain human nucleic acid. In some cases, synthetic nucleic acids do not contain nucleic acids that can be found in nature. Does not contain.
[0118] The guanine-cytosine content (GC content) of the spike-in or synthetic nucleic acid can vary In some cases, the GC content of the spike-in or synthetic nucleic acid is at least about 0%, 5% , 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55% , 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 100% In some cases, the GC content is at most about 5%, 10%, 15%, 20%, 2 5%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 7 It may be 5%, 80%, 85%, 90%, 95% or 100%. The GC content of the spike-in or synthetic nucleic acid is about 15% to about 85%, for example, about 20% to about 80%. The GC content of the species of synthetic nucleic acids in the collection can be in the range of 0.1% to 0.2% of the specific target nucleic acid (e.g., The GC content and severity of pathogen-specific or disease-specific nucleic acids in a sample (observable range) In other cases, the GC content of the species of synthetic nucleic acids within a collection of synthetic nucleic acids may be It can exactly match the GC content of the target nucleic acid, or it can substantially match such GC content. For example, the GC content of the synthetic nucleic acid seed is 75% to 125% of the GC content of the target nucleic acid, GC content 80% to 120% of the target nucleic acid, GC content 85% to 115% of the target nucleic acid 90%-110% of the amount, 95%-105% of the GC content of the target nucleic acid, 96%-104% of the GC content of the target nucleic acid, 99%-101% of the GC content of the target nucleic acid, or The content may be within the range of 99.5% to 100.5%.
[0119] The spike-in nucleic acid is conjugated to a different molecule, e.g., a bead, a fluorophore, a polymer, and Examples of fluorophores include fluorescent proteins, green fluorescent proteins, and the like. Photoprotein (GFP), Alexa dye, fluorescein, red fluorescent protein (RF Examples of fluorescent proteins include, but are not limited to, yellow fluorescent protein (YFP) and yellow fluorescent protein (YFP). Spike-in nucleic acids are proteins (e.g., histones, nucleic acid binding proteins, DNA binding proteins, etc.) In other cases, the spike may be bound to a ribosomal protein (ribosomal protein, RNA-binding protein). The spike-in nucleic acid is not bound to a protein. The spike-in nucleic acid can be protected with a particle (e.g. (similar to the nucleic acid in a virion). In some cases, the spike-in nucleic acid is packaged within the particle. In some cases, the particles are proteins, lipids, metals, These include oxides, plastics, polymers, biopolymers, ceramics or composites.
[0120] The spike-in nucleic acid is distinct from sequences potentially found in the sample or host. In some cases, the spike-in nucleic acid sequence is naturally occurring. In some cases, the spike-in nucleic acid sequence is not naturally occurring. The nucleic acid sequence is derived from a host. In some cases, the spike-in nucleic acid sequence is not derived from a host. In some cases, the spike-in or synthetic nucleic acid may be a nucleic acid that is a nucleic acid sequence that is a nucleic acid sequence that is a target nucleic acid (e.g., a pathogenic nucleic acid, disease-specific nucleic acid) and / or one or more sample nucleic acids (or not complementary).
[0121] The concentration of the spike-in nucleic acid in the sample can vary. The eluents can be added in a wide range of concentrations which can be useful in determining sample loss. In the case of 100,000, 500,000, 1 million, 2 million, 3 million, 4 million, 5 million, 60 0 million, 7 million, 8 million, 9 million, 10 million, 20 million, 30 million, 4000 10,000,000, 50,000,000, 60,000,000, 70,000,000, 80,000,000, 90,000,000, 100,0 ... or One billion molecules of each spike-in nucleic acid are added per mL of plasma or sample. In some cases, about 10 million to about 1 billion molecules of each spike-in nucleic acid are per mL of plasma or sample. In other cases, synthetic nucleic acids are The antibody is spiked into the sample at different concentrations.
[0122] The number of different spike-in nucleic acids added to the sample can vary. In some cases, at least about 1, 2, or 3 nucleic acids may be added to the sample or reagent. , 3, 4, 5, 6, 7, 8, 9 or 10 spike-in nucleic acids into the sample or reagent In some cases, at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, or Ten spike-in nucleic acids are added to the sample or reagent. The spike-in nucleic acids added to the pull or reagent are the same length. The spike-in nucleic acids added to the sample or reagent are of different lengths. The spike-in nucleic acid is a group consisting of SEQ ID NOs: 1 to 120 and any combination thereof. is selected from.
[0123] The level of uniqueness of the spike-in nucleic acid can vary. Essentially unlimited A number of spike-ins (eg, ID spikes) may be designed or used.
[0124] The step in the process at which the spike-in nucleic acid is added can vary. For the purpose of cloning, the earlier addition of the spike-in nucleic acid is better, and the later It may reduce the possibility of operator or system error. In some cases, The tube to which the sample (e.g., blood) is first added may already contain spike-in nucleic acid. The manufacture of these tubes is convenient for testing samples in a clinic or research laboratory. This allows for more systematic control and testing compared to the addition of equine nucleic acid. This can reduce the chance of sample mix-up. In some cases, ID spikes may replace all external labels ("white labels").
[0125] In some cases, the sequence reads are sequenced with a distinguishing nucleic acid marker, such that each sequence read contains the distinguishing marker. can be added to each nucleic acid fragment in the sample. This method is If the tagging of the fragments is sufficiently complete, it will allow the identification of cross-contamination. also allows for deliberate multiplexing of samples as soon as barcodes are added to the sample fragments. Methods for incorporating tags include transposons, terminal transfer These include cleavage at methylation sites, cleavage at demethylation sites, and However, the present invention is not limited to the above.
[0126] Applications including but not limited to those related to process quality control or development work For other applications, including those involving the manufacture of pharmaceutical compositions, spike-in nucleic acids may be added at various steps in the process. For example, in the case of RNA analysis, different concentrations, lengths, sequences and / or GC content can be used. Multiple RNA spike-ins, each with their own unique characteristics, can be added at the start of sample preparation. After the RNA is converted to DNA, DNA spike-ins can be added. In Lee's case, different forms of DNA were added at different steps in the library construction process. For example, to test the end repair process, a non-blunt-ended, 5'-phosphate-containing with or without (+ / - 5'-phosphate), and with or without 3'-adenine extension Adapters without (+ / - 3'-adenine stretch) DNA spike-ins can be used. To test the process of ligating the end-repaired fragment to the end-repaired fragment, (+ / - pre-adapted) spike-in may be used. Sequencing qPCR Sample loss at each step can be quantified. Spike-in qPCR can also be used to quantify This method may be used in conjunction with other library quantification methods for final library evaluation prior to determination.
[0127] "Spike-in", "spike-in synthetic nucleic acid", "spike" and "synthetic nucleic acid" The words are used interchangeably herein and shall have the same meaning unless the context indicates a different interpretation. The terms "ID spike" or "tracer" should be interpreted as , e.g. sample identification tracking, cross contamination detection, reagent tracking or reagent lot Generally, the term "spike" is used herein to refer to a discriminating spike that can be used for tracking. The term "Spark" refers to abundance normalization, development and / or provides nucleic acids that are size or length markers that can be used for analytical purposes as well as other purposes. "Spank" is a term commonly used in this specification to mean The term is generally understood to mean a degenerate pool, or a pool of nucleic acids having diverse sequences. and can often be used for diversity assessment and abundance calculations.
[0128] General normalization of nucleic acid results The present disclosure provides a method for detecting disease-specific Allows for efficient and improved normalization of the amount of nucleic acid, pathogen-specific nucleic acid or other target nucleic acid The set of spiked nucleic acids is described below. It is possible for the nucleic acid to contain several "species" of different nucleic acids, resulting in a collection of added nucleic acid species. The combination generally represents the length of the pathogen nucleic acid, disease-specific nucleic acid, or other target nucleic acid being measured. Within the observable range.
[0129] Spike-in synthetic nucleic acids can be used to normalize samples in a variety of ways. Often, normalization is performed based on the subject from which the sample was derived, preanalytical sample processing conditions, nucleic acid extraction, etc. Methods, manipulation of nucleic acids by molecular biology means and methods, methods for purifying nucleic acids, carrying out the measurements themselves; This may be across samples, regardless of storage conditions and / or age.
[0130] In some preferred embodiments, the spike-in nucleic acid is a disease-specific nucleic acid, a pathogen-specific Normalize across all methods and all samples that measure nucleic acids or other target nucleic acids In some cases, the spike-in is used to detect pathogen nucleic acid (or disease-specific Determine the relative abundance of a specific nucleic acid (allegorous nucleic acid or target nucleic acid) compared to other pathogen nucleic acids It can be used for:
[0131] In general, the methods provided herein involve the addition of one or more of a set of synthetic nucleic acids to a sample (a sample containing the synthetic nucleic acid). This addition step may be performed early, mid or late in the process. The method may be performed at any point throughout the process, including the end of the process. For example, synthetic nucleic acids may be prepared by isolating samples from a subject. When the sample is taken or immediately after, before or during sample storage, and during sample transport before or during nucleic acid extraction, before or during library preparation, before or during sequencing assays It may be introduced just before, or at any other step in the method. The method may involve the detection of pathogen-specific or disease-specific nucleic acids or other subunits that may be measured by the same method. Early in the process, a biological sample is prepared by adding a known amount of unique nucleic acid molecules that are easily distinguished from the sample nucleic acid. In some cases, a single step of the process (e.g., For example, when a sample is taken from a subject, when a sample is obtained for analysis, when a sample is During storage, before or during nucleic acid extraction, before or during library preparation, or during sequencing In other cases, the synthetic nucleic acid is added to the biological sample immediately prior to the assay. The same or different spike-in synthetic nucleic acids may be introduced at different steps in the process. For example, unique synthetic nucleic acids can be introduced early in the process (e.g., at the time of sample collection). and a different set of unique nucleic acids can be isolated at a later time point in the process (e.g. For example, it may be introduced at any time during the extraction, purification or library preparation. Alternatively, the same or in some way different collections of spike-in nucleic acids may be used to detect the It is also possible to repeat the addition step at different steps of the process.
[0132] Typically, a known concentration (or concentrations) of synthetic nucleic acid species can be added to each sample. In most cases, the synthetic nucleic acid species are added in equimolar concentrations. In this case, the concentration of the synthetic nucleic acid species is different.
[0133] Handling, preparation and measurement of samples, if they are to be processed and ultimately measured Due to inherent biases in the analysis, the relative abundance of the nucleic acid species may change. The recovery efficiency of each length of nucleic acid was determined by comparing the measured abundance of each to the initial amount added. This may provide a "length-based recovery profile."
[0134] Using a "length-based recovery profile," the length of the spiked molecule is compared to the closest length. is the abundance of disease-specific nucleic acids (or is the abundance of pathogen nucleic acid or other target nucleic acid) to normalize disease-specific nucleic acid , normalizing the abundance of all (or most or some) of the pathogen nucleic acids or other target nucleic acids This process is applicable to disease-specific nucleic acids and can be performed without sample addition. This can provide an estimate of the "original length distribution of all disease-specific nucleic acids" at a given time point. The process is also applicable to other target nucleic acids, e.g., pathogen-specific nucleic acids, and can be performed after sample addition. This can provide an estimate of the "original length distribution of all pathogen-specific nucleic acids" at a given time point. The "original length distribution of the target nucleic acid" refers to the distribution of the original length of the target nucleic acid (e.g., disease-specific nucleic acid) at the time of addition of the sample. , pathogen-specific nucleic acids). Try to recapitulate the added nucleic acid to achieve a reasonable abundance normalization. It is the distribution of these lengths that is taken as the distribution.
[0135] The relative presence of disease-specific nucleic acids, pathogen nucleic acids or other target nucleic acids in a particular sample It is not possible to spike a sample with a mixture of known nucleic acids that closely reproduces the abundance profile. (This may be due in part to the sample being exhausted or time being relative. (This may have altered the overall abundance profile of each spike-in species.) is weighted proportionally to its relative abundance in the original length distribution of all disease-specific nucleic acids All "weighting factors" can be used. The sum of can be equal to 1.0.
[0136] Normalization may involve a single step or a series of steps. In some cases, the nearest size The raw measurements of the abundance of the spiked nucleic acid are used to identify pathogen-specific nucleic acids (or pathogen nucleic acids or other The abundance of the target nucleic acid is normalized to obtain a “normalized disease-specific nucleic acid (or pathogen nucleic acid or Then, a "normalized disease-specific nucleic acid abundance" can be obtained. "(or pathogen nucleic acid or other target nucleic acid abundance) by a "weighting factor" and The weighted normalized disease-specific (or pathogen-specific) It is possible to obtain "specific or other target nucleic acid abundances." Advantages of this normalization method One of the reasons is that it is compatible with all (or most) methods that measure the abundance of disease-specific nucleic acids. Throughout, regardless of the method, the presence of a target nucleic acid (e.g., disease-specific nucleic acid, pathogen nucleic acid) This would allow for equivalent measurement of the amount of
[0137] Measurement of target nucleic acid abundance or relative abundance is useful in detection, prognostic, monitoring and diagnostic assays. Such assays may be particularly useful for detecting the presence of pathogens and is a method for detecting a target nucleic acid (e.g., The methods described herein may include measuring the amount of a nucleic acid sequence (such as a nucleic acid sequence of a disease-specific nucleic acid) in a nucleic acid sequence of a disease-specific nucleic acid. These measurements were performed using samples, measurement times, nucleic acid extraction methods, nucleic acid manipulation methods, nucleic acid measurement methods, and and / or across a variety of sample processing conditions.
[0138] The exact sequence of the added molecules, the exact number of "species", the length range of the "species", the concentration of the added molecules, The relative amounts of molecules, the actual amount of each molecule added, and the stage at which the molecules are added will be optimized based on the sample. The length can be optimized or adjusted based on GC content, nucleic acid structure, DNA damage or DNA modification status. can be substituted or analyzed.
[0139] In some cases, the methods provided herein may include (in some cases) short runs. A single length of nucleic acid, often with a largely fixed sequence composition (except for randomized portions) The method may include the use of an additive nucleic acid comprising a disease-specific nucleic acid, a pathogen-specific nucleic acid or This may work well if the other target nucleic acid is of approximately the same length as the added nucleic acid.
[0140] A single length of nucleic acid can be used alone, or the method can involve the use of multiple lengths of nucleic acid. For example, the method may be combined with another method involving the use of a nucleic acid extract when the sample is obtained or prior to extraction of the nucleic acid. In addition, a pool of nucleic acids of multiple lengths can be added to a sample, and nucleic acids of a single length can be Pools of acid are added at different points in the process (e.g., after nucleic acid extraction and after library extraction). It is possible to add single length and / or multiple length nucleotides to the sample (prior to preparation of the nucleotide sequence). When using a nucleic acid of length, the amount of disease-specific nucleic acid, pathogen nucleic acid or other target nucleic acid is It is possible to normalize to the amount of added nucleic acid measured at the end of the method.
[0141] In many cases, as described herein, the use of synthetic nucleic acids of multiple lengths will These may be preferred over methods involving the use of single length synthetic nucleic acids. Nucleic acids are particularly useful when the target nucleic acid has multiple lengths. For example, disease-specific (or Disease- or pathogen-specific nucleic acids can vary widely in length. The use of spike-in nucleic acids spanning the observable length of the The length of a disease-specific nucleic acid depends on the metabolism of the individual from which it is derived, the preanalytical sample processing conditions, the nucleic acid Methods of extraction, manipulation of nucleic acids by molecular biology means and methods, methods of nucleic acid purification, and the performance of the measurement itself They can also be dramatically affected by a number of factors, including application, storage conditions and the passage of time. The factors have differential effects on nucleic acids of different lengths and therefore the Acids do not adequately reflect the overall efficiency of processes carried out on nucleic acids of mixed lengths It's possible.
[0142] Calculating "genome copies per volume" The method and synthetic nucleic acid provided by the present invention can be used to identify samples from the results of next-generation sequencing. A computational method for determining the number of genome copies per volume of a microorganism or pathogen in a sample. Generally, genome copy numbers per volume can be used to aid in the determination of the number of copies of a given genome per volume of a fluid (e.g., blood). Target nucleic acid (e.g., target nucleic acid derived from a specific pathogen) per ml of plasma, urine, buffer, etc. It is possible to provide an absolute measure of the amount of a specific nucleic acid (specific nucleic acid) and thus the abundance or relative abundance of individual pathogens. It is often used as an expression to indicate abundance. The total number and / or magnitude of d readings may vary from sample to sample. It is advisable to report values that correspond to the levels of analytes and that may be useful for sample-to-sample comparisons. There are some interesting things.
[0143] In certain instances, the method comprises: per volume of pathogen nucleic acid in a sample obtained from a subject suspected of being infected The genome copy number per volume can be used to determine the genome copy number. The statistical framework can be determined or estimated using a framework. Providing a collection of non-human reads (e.g., pathogen reads) in the results It can be used to estimate the relative abundance of one or more genomes.
[0144] The spike-in synthetic nucleic acids provided herein can be used to detect one or more pathogens in a sample. An estimate of the "genome copy number per volume" of an organism can be calculated. Nucleic acids can be added to the sample at known concentrations. In some cases, The proportion of information from a sample that is observed when By comparing the number of acid-associated leads or the number of observed leads with the number of added leads The non-residential spike-in length can be observed for each spike-in length (by dividing by 1). It is also possible to back-calculate the original number of host or pathogen molecules (e.g., the number of spars at each length). (Inferred in part from the number of equinuclear reads). The genome length of each pathogen is known. In this case, this load can be converted into a measure of "genome copies per volume" .
[0145] In many cases, methods for detecting genome copies per volume (and those provided herein) are Other methods for elimination of low quality reads (e.g., elimination of low quality reads) may include elimination or isolation of low quality reads. The accuracy and reliability of the methods provided herein may be improved. In some cases, the methods include: Unmappable reads, derived from PCR duplicates reads, low quality reads, adapter dimer reads, sequencing adapter reads, non-uniform reads (Any combination of) the map-mapped reads and / or reads that map to uninformative sequences This may include removal or segregation (without
[0146] In some cases, sequence reads are mapped to a reference genome, and such reference Reads not mapped to genomes map to one or more target or pathogen genomes In some instances, the reads are mapped to a human reference genome (e.g., hg19). The remaining reads were identified as viral, bacterial, fungal and other eukaryotic pathogens (e.g. The genomes are then mapped against a curated reference database of genes (e.g., fungi, protozoa, parasites).
[0147] In some particular examples, the method includes DNA extraction (e.g., cell-free DNA extraction, either before cellular RNA extraction) or at different stages of the assay (e.g., post-extraction, library preparation) A known concentration of synthetic nucleic acid (e.g., DNA) is added to the sample (before sequencing, before analysis, during sample storage). Synthetic nucleic acids may include adding a nucleic acid to a sample (e.g., a plasma sample). In some cases, the control sample may also be added to the sample. The method may further comprise the steps of: , negative control), the library. is a sequencing device known in the art, in particular a device having the capability of next generation sequencing. The reads can be multiplexed and sequenced at 100 bp. The method can further include discarding low quality reads. , and aligning to a human reference sequence (e.g., hg19) to remove human reads. The remaining reads can then be aligned to a database of pathogen sequences. In some cases, a sequence corresponding to a target sequence of interest (e.g., a pathogen sequence) may be generated. Reads are quantified from the NGS read set. From this information, the target nucleic acid (e.g., pathogen The relative abundance of a nucleic acid (a single molecule) can be expressed as genome copies per volume. The genome copy number can be determined, for example, by measuring the number of copies of a known amount of oligonucleotides added to a sample (e.g., plasma). Determine the number of sequences present for each organism (e.g., pathogen) normalized to the otide Calculation of the number of genomes per volume can be done by calculating the relative lengths of the individual pathogen genomes. In some cases, the number of genome copies per volume may be calculated for each organism (e.g. quantifies the number of sequences present for a given pathogen (e.g., pathogens) and compares them against a known amount of synthetic nucleic acid added to the sample. The pathogen sequence can be normalized by length, Consider the synthetic nucleic acid that is closest to the pathogen sequence in terms of the sequence. Similarly, normalization can be performed for sequences of various lengths (e.g., 2, 3, 4, 5, 6, 10, 15, 20 or more different lengths of spike-ins It is possible to include the use of a collection of synthetic nucleic acids, where the pathogen nucleic acid is a spike-in nucleic acid. The spike-in nucleic acid is normalized to the closest spike-in nucleic acid in the collection. .
[0148] Spike-in for sample tracking and / or analysis Molecules are added to samples to generate unique identifiers and It is possible to obtain tracers. These molecules can become part of the sample. and can be read by a suitable measuring device (e.g. a sample read by a laser scanner). (similar concept to 1D or 2D barcodes on the outer surface of a tube). Optical, radioactive and and other tracers are possible, but for the analysis of nucleic acid samples, nucleic acid tracers are used. This may be the best choice because the nature of the spike-in determines how the nucleic acid in the sample is evaluated. This is because the results can be shown in the same process (e.g., DNA or RNA sequencing) that is used to be.
[0149] Nucleic acids of exogenous origin can include oligonucleotides, double-stranded oligonucleotides, and longer oligonucleotides. (assembled) double-stranded DNA (e.g., Integrated DNA Tech hnologies), plasmids, PCR products, in vitro synthetic transcription Products, viral particles, and fragmented or non-fragmented genomic DNA (including but not limited to (not necessarily) which can be added to a sample, e.g., a bodily fluid from a subject. Advantages of using equines include the fact that the nucleic acid sequence, length, diversity and concentration are This includes, but is not limited to, being adaptable to
[0150] Applications include, but are not limited to: Sample tracking (Tracking) [e.g. in addition to, or potentially instead of, regular label barcodes Alternatively, ID spikes may be used], sample cross-contamination (e.g., sample If the ID spike is not naturally found in any of the pulls and the sample If different ID spikes are added to the samples, a mixture of samples can be determined. ), reagent tracking [e.g., ID spikes can also be added to reagents. For example, each reagent lot Reagent tracking is traceable for each sample in which it is used, resulting in less error in reagent tracking Molecular Laboratory Information Management Systems (LIMS) may be implemented], quality control or development work [For example, library complexity (e.g., PCR duplicates) e) Different spike-ins are used to monitor sample loss or sensitivity. may be added at various points in the treatment process], normalization or yield [e.g., known input Comparing the force to the measured output of the spike-in allows the unknown input (e.g. in the sample) to be From these measurements and calculations, it may be possible to make inferences based on the measurement output. and increased nucleic acid concentration (e.g., the barcode is a nucleic acid) In some cases, it can be used in high concentrations for samples with limiting nucleic acid concentrations. , which may improve sample recovery).
[0151] In some preferred embodiments, the spike-in is a nucleic acid sequence that is , the likelihood that it originates from the observed sample, or its Whether their presence could be the result of cross-contamination or carry-over contamination from different samples The unique spike-in molecules can be used to estimate the number of unique spike-in molecules from a particular pathogen. (or other sequence class of interest) at a concentration higher than would reasonably be expected. introduction into the pool to prevent accidental introduction through cross-contamination or carry-over Any pathogen sequences (or other sequence classes of interest) are either contaminating or carry-over contaminants. It is likely that there will be a greater number of spike-in molecules from the source of the sequence. Pathogen sequence counts (or other classes) versus cross-contamination or carry-over spike-in molecule counts The ratio of the number of nucleotides in the sequence of the DNA sequence (i.e., the number of nucleotides in the sequence) that could be the result of cross-contamination or carryover between samples was used to determine the ratio of the number of nucleotides in the sequence of the DNA sequence (i.e., the number of nucleotides in the sequence of the DNA sequence) that could be the result of cross-contamination or carryover between samples. It is possible to identify any pathogen sequence. In some cases, cross-contamination or Utilizing the absence of contaminating spike-in molecules or their presence at levels below a threshold level to indicate that the sample is not contaminated.
[0152] For some applications, it may be desirable to determine the genotype of the subject from which the sample was derived, particularly for sample tracking purposes. In some cases, the genotype can be determined during the analysis procedure. Alternatively, an aliquot can be removed and subjected to a separate genotyping procedure. In some cases, the genotype of the sample is already known. The constant output can be compared to independently obtained genotypes. The advantage is that it is already part of the sample and is inherent to the sample. A novel orthogonal genotyping method is the detection of short tandem repeats. One such method is strain-specific transcriptome (STR) analysis. See, for example, ATCC testing services.
[0153] In some cases, phenotypic characteristics may help identify a sample. For example, the eye color of a subject. , blood type, sex, race and other traits could provide clues to the genotype.
[0154] ID Spike The unique sample identifier can be completely scrambled (e.g., DN Randomization of A, C, G, and T in A, or randomization of A, C, G, and U in RNA Alternatively, they may have several regions of common sequence. For example, Shared regions at each end can reduce sequence bias in ligation events. The region of interest is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 In some cases, the shared region is at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20 common base pairs. For typical sequences, see Table 1. I want to be.
[0155] Combinations of ID spikes increase diversity without using a huge number of ID spikes. The ID spikes can be added to well locations in a microtiter plate. can be used as an identifier for each well (e.g., for a 96-well plate, 96 distinct Another ID spike can be used as an identifier for the plate number. Yes (e.g., 24 different ID spikes for 24 different plates), slightly Using 96+24=120 arrays, we get 96×24=2,304 combinations. The use of three or more ID spikes per sample increases the achievable diversity even more dramatically. It can be great. [Table 1] TIFF2025060811000003.tif229167TIFF2025060811000004.tif229168TIFF20250608110 00005.tif228167TIFF2025060811000006.tif229167TIFF2025060811000007.tif229167 TIFF2025060811000008.tif227167TIFF2025060811000009.tif228168TIFF20250608110 00010.tif229167TIFF2025060811000011.tif229168TIFF2025060811000012.tif158168
[0156] Spark Bias Control Spike-in A set of nucleic acid sequences spanning multiple lengths ("Spark") is a size marker. These sequences can be added to the sample along with the sample nucleic acid and used for processing (e.g. For example, extraction, purification, and sequencing. Some processes may affect nucleic acids of different lengths differentially. For example, nucleic acid purification using silica membrane columns may affect the resolution of longer sequences. can be biased against or optimized to retain sequences of a particular length Nucleic acid sequencing is typically performed after nucleic acid extraction from a sample, The length frequency or distribution in the sequencing results may not be representative of the original sample. By adding spark sequences of known quantity and length, it is possible to generate samples of various lengths. It is possible to monitor and quantitate the effects of treatment and sequencing on sample nucleic acid. In addition, the sample nucleic acid and the spark size set The final number of sequencing reads for a nucleic acid is measured, based on the known amount of By normalizing to the Spark Size Set nucleic acid, the various It is possible to estimate the relative and / or absolute amount of sample nucleic acid of length.
[0157] In some cases, the spark size set is at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50 , 100, 200, 250, 300, 350, 400, 500, 600, 700, 800 In some cases, the Spark size set may contain 1000 or more nucleic acids. The number of t is at most about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, It may comprise 20, 25, 30, 35, 40, 45, 50, 100 or 200 nucleic acids. In some cases, the spark size set contains about 3 to about 50 nucleic acids, e.g., about 3 to about 30 In some cases, the nucleic acids in the spark size set include one or more different nucleic acids. The polypeptides may have different properties, such as different lengths, different GC contents and / or different sequences.
[0158] Spark Nucleic Acid consists of a length discrimination sequence, a load sequence, and a synthetic nucleic acid discrimination sequence (which is (which in this case would be a Spark identification sequence) and a signature domain, as described herein. The synthetic spike-in nucleic acid may include any of the features of the synthetic spike-in nucleic acid described herein. The nucleic acids in the clease set may be a fixed forward sequence and / or a fixed reverse sequence. It contains a fixed forward sequence and / or a fixed reverse sequence. The sequence can be common to all nucleic acids in the spark size set, and the sequence can be In some cases, a fixed forward sequence and / or The fixed reverse sequence is at least about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 6, 17, 18, 19, 20, 25, 30, 32, 40, 50, 60, 70, 80, 90 or 100 base pairs in length. In some cases, a fixed forward sequence and / or or the fixed reverse sequence is at most about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 32, 40, 50, 60, 70, 80, 9 0 or 100 base pairs in length. In some cases, a fixed forward sequence and / or Alternatively, the fixed reverse sequence may be from about 8 bp to about 50 bp, for example, from about 8 bp to about 20 bp. p, or in the range of about 16 bp to about 40 bp. The sequence is not found in the sample or does not occur naturally. The forward sequence fixed is different from the reverse sequence fixed.
[0159] In some cases, the nucleic acids in the Spark Size set are unique forward and reverse sequences. Contains a unique forward sequence and / or a unique reverse sequence. A nikreverse sequence can distinguish the sparks in the size set from each other. In the case of, the unique forward sequence and / or the unique reverse sequence are at least Also about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 32, 40, 50, 60, 70, 80, 9 0 or 100 base pairs in length. In some cases, a unique forward sequence and / or or unique reverse sequences are at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 32, 4 0, 50, 60, 70, 80, 90, 100, 200, 300, 306, 400 or 5 00 base pairs in length. In some cases, a unique forward sequence and / or a unique The cleavage sequence is in the range of about 4 to about 10 base pairs in length. Each nucleic acid in the size set has a different unique forward sequence and / or a unique linker sequence. In some cases, each nucleic acid in the spark size set has the same It has a unique forward sequence and / or a unique reverse sequence having a length. In some cases, each nucleic acid in the spark size set is a unique nucleic acid having a different length. It has a forward sequence and / or a unique reverse sequence.
[0160] In some cases, the nucleic acids in the Spark Size Set contain filler sequences. In some cases, the filler sequence distinguishes the sparks in the size set from each other. In some cases, the packing sequence may include at least about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 2 5, 30, 32, 40, 50, 60, 70, 80, 90 or 100 base pairs in length. In some cases, the packing sequence is at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 32, 4 0, 50, 60, 70, 80, 90, 100, 200, 300, 306, 400 or 5 00 base pairs in length. In some cases, the filler sequence is in the range of 0 to about 350 bp. In some cases, each nucleic acid in the spark size set is a filler sequence having a different length. In some cases, the length of the filler array is 0, 8, 31, 56, 81, 106 , 131 and 306 bp.
[0161] In some cases, the nucleic acids in the spark size set are at least about 10, 20, 30, 32, 40, 50, 60, 70, 80, 90 or 100 base pairs in length. In the case of, the nucleic acids in the spark size set are at most about 100, 200, 300 , 350, 400, 500, 600, 700, 800, 900 or 1,000 base pairs in length In some cases, the nucleic acids in the Spark size set are about 20 to about 500 bases long. Within the range of about 20 to about 400 base pairs in length, or within the range of about 20 to about 200 base pairs in length It is within range.
[0162] For example, eight double-stranded DNA sequences (Table 2, SEQ ID NO: 1 in FIG. 4) having the following characteristics: A set of 11 to 118) can be designed: 32 to 350 bp size range (e.g., Three filling sequence lengths of 0, 8, 31, 56, 81, 106, 131 and 306 bp were used. 2, 52, 75, 100, 125, 150, 175 and 350 bp fragments), immobilized A fixed 16-bp forward sequence and a fixed 16-bp reverse sequence different from the forward sequence sequence, as well as the unique 6 bp forward and reverse sequences. [Table 2] TIFF2025060811000014.tif234161TIFF2025060811000015.tif150162
[0163] GC Content Spike-in Panel A nucleic acid (e.g., DNA) that is added to a sample at a known concentration and then measured after processing can provide yield and other information about the process, which may include and can be used to infer additional properties about the sample itself. A nucleic acid spike inset containing the size range is added to a sample (e.g., plasma) and then Extraction and subsequent next generation sequencing (NGS) can be performed on each size spy. The yield of the PCR product is affected by a variety of factors, including deliberate size selection, temperature and other denaturing factors, and PCR biases. This information can be used to optimize recovery of the desired size range. To develop new methods to increase productivity or to monitor existing processes This may be useful for determining whether a particular
[0164] For double-stranded DNA library preparations, a relatively low melting temperature (T m ) DNA duplex The denaturation of Tm The yield of these double strands decreases inversely proportional to the given conditions (e.g. For example, salt concentration, temperature, pH, etc.) m Contributing factors that affect include: The sizes include length and GC content. Each size is represented by a single species with a single GC content. The size range of the duplex is shown in Fig. 1. m Provides only partial information about the response It is possible.
[0165] The length and / or GC content of the nucleic acid is m and how it affects processing For example, information on the recovery of short cell-free fragments from various pathogens in blood may be used to estimate the This may be important when using spike-ins as surrogates for pathogen nucleic acids. can vary dramatically in GC content and therefore are very sensitive to short fragment lengths. Various T m A number of cfDNA fragments of short length (e.g., 30, 40, 50 bp), they may be susceptible to denaturation during processing for example for NGS. . Extensive T m More detailed spike insets to track recovery across a range This may allow for a better estimation of the starting amount of an unknown sample.
[0166] A range of T m A panel of spike-in nucleic acids spanning GC and / or length is used to determine absolute for the determination of specific abundance and / or to allow detailed monitoring of denaturation For example, as shown in Table 3, four different lengths (e.g., 32, 42 , 52 and 75 bp) and seven different GC contents for each length (approximately 20, 30, 4 28 different nucleic acids (e.g., 0, 50, 60, 70, or 80% GC) For example, panels of one strand (double strand) may be used. Overall, the panels are made of a single strand for each size. In some cases, synthetic nucleic acids may provide a higher granularity than sets having a GC content of The panels (dsDNA, ssDNA, dsRNA, ssRNA) were and for each length, at least two different GC contents, at least three G C content, at least 4 GC content, at least 5 GC content, at least 7 GC content or or a nucleic acid with a GC content of at least 10. In some cases, synthetic nucleic acids (d The panel of DNA fragments (sDNA, ssDNA, dsRNA, ssRNA) was designed to contain at least five different length, and at least two different GC contents for each length, at least three GC contents amount, at least 4 GC content, at least 5 GC content, at least 7 GC content, or It may contain nucleic acids with a GC content of at least 10.
[0167] In some cases, the spike-in panel will include at least 3, 5, 10, 15, 20, 2 In some cases, spike-in panels contain as many as 5 or 30 unique nucleic acids. Each gene contains 15, 20, 25, 30, 35, 40, 45, 50 or 100 unique nucleic acids. include.
[0168] Spike-in nucleic acids with various GC contents can be used. The quine panels are approximately 40-60% GC, approximately 45-65% GC, and approximately 30-70% GC. , about 25-75% GC, or about 20-80% GC. In some cases, the spike-in panel comprises at least 2, 3, 4, 5, In some cases, the nucleic acid may have 6, 7, 8, 9, or 10 different GC contents. Spike-in panels may consist of at most 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20 In some cases, the spike-in panel includes nucleic acids having different GC content of GC. are at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20% different The percentage of GC is the percentage of G nucleotides and C nucleotides. The number of nucleotides in the sequence may be calculated by dividing the sum of the number of nucleotides in the sequence by the total number of nucleotides in the sequence. For example, For the sequence ACTG, the percentage of GC (% GC) is (1+1) / 4=50% GC. will be calculated.
[0169] Spike-in nucleic acids of various lengths can be used. The panel must have at least 3, 4, 5, 6, 7, 8, 9, 10 or 15 different lengths. In some cases, the spike-in panel includes at most 3, 4, 5, 6 , 7, 8, 9, 10, 15, 20, 25, 50 or 100 nucleic acids of different lengths In some cases, the spike-in panel includes about 40-50 bp, about 35-55 bp, p, about 30~60bp, about 35~60bp, about 35~65bp, about 35~70bp, about 3 5~75bp, approx. 30~70bp, approx. 30~80bp, approx. 30~90bp, approx. 30~10 0bp, about 25-150bp, about 20-300bp, or about 20-500bp In some cases, the spike-in panel includes nucleic acids having lengths ranging from at least Available in various lengths with 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20 bp difference In some cases, the spike-in panel includes nucleic acids that Nucleic acids having a length of 5 bp, or lengths of 27, 37, 47, 57, 62 and 67 bp Includes.
[0170] Spike-in nucleic acids having lengths and GC contents selected from a range of values may be used. For example, the set of synthetic nucleic acids can be selected from two or more lengths and two or more GC contents. The set of 28 synthetic nucleic acids in Table 3 (SEQ ID NO: 125 to SEQ ID NO: 152) consists of four different The sequences are of different lengths (e.g., 32, 42, 52 and 75 base pairs) and seven different GC contents ( For example, about 20, 30, 40, 50, 60, 70 and 80% GC). Different lengths (e.g., 27, 37, 47, 57, 62 and 67 bp) and different GC content (e.g., about 15, 25, 35, 45, 55, 65 and 75% GC), A similar set of synthetic nucleic acids can be obtained.
[0171] Various melting temperatures (T m ) may be used. , spike-in panels are approximately 40 to 50°C, approximately 35 to 55°C, approximately 30 to 60°C, approximately 35 to 60℃, about 35 to 65℃, about 35 to 70℃, about 35 to 75℃, or about 30 to 70℃ Melting temperature (T m In some cases, the nucleic acid may include a nucleic acid having the following structure: ,Spike-in panels must have at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 , various melting temperatures (T m ) is included.
[0172] In some cases, T m In addition to duplex length and GC content, duplex concentration, nucleic acid Nearest neighbor effects in nucleotide sequences, higher order DNA structure, and monovalent and / or divalent cation concentrations In some cases, T m is given Conditions, e.g., double-stranded DNA specific dye and gradual increase in temperature and detection of dye signal and can be calculated experimentally.
[0173] Spike-in nucleic acids with a variety of sequences can be used. A natural sequence or a sequence that cannot hybridize to the sample nucleic acid is used. In this case, the spike-in panel must consist of at least 3, 4, 5, 6, 7, 8, 9, 10 or 1 In some cases, the spike-in panel contains nucleic acids having as many as five different sequences. Also 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50 or 100 different The present invention includes nucleic acids having a sequence
[0174] Various numbers of spike-in nucleic acids can be used. In some cases, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40 or 50 nucleic acids are used. For example, a subset of the 28 sequences listed in Table 3 (e.g., 32 / 42 / 52) / 75bp × 20 / 50 / 80 % GC) can be used.
[0175] For RNA applications, an RNA panel may be used. RNA panels may consist of identical molecules or may differ in terms of length, GC content and / or other properties. The molecule may include a variety of molecules.
[0176] A set of 8 DNA sequences (each approximately 50% GC, SEQ ID NO: 11 in Table 2) 1–118) is the partial coverage of the 28-member GC panel listed in Table 3 ( coverage). [Table 3] TIFF2025060811000017.tif221169TIFF2025060811000018.tif143168
[0177] Degenerate Spike-in: Spank The spike-in synthetic nucleic acid is a degenerate pool of nucleic acids, or a pool of nucleic acids with high diversity. (sometimes referred to herein as "Spank"). Generally, a spank is a sample that leads to and / or contains a sequencing reaction. To determine the absolute or relative nucleic acid loss or diversity reduction that may occur during the nucleic acid processing steps, In the case of unique pools of spanked sequences, this can be used to reduce sequence diversity within the pool. should correspond directly to a decrease in nucleic acid abundance, in this case due to amplification or PCR bias. There is no need to consider the impact. For example, 8 Add unique spank sequences to the sample After sequencing, 4 If only one unique span sequence was recovered, the presence of nucleic acid Both abundance and nucleic acid diversity are 10 4 In some cases, Spank can be used to determine the degree of recovery of overlapping molecules. and library processing (which may involve PCR and potentially heterogeneous amplification of the various input molecules). ), sequencing and alignment of individual spans can indicate the degree of recovery of overlapping molecules. do.
[0178] The determined diversity reduction may then be used to determine the degree of diversity of one or more sample processing or sequencing steps. Prior to this, it is necessary to determine the absolute abundance of a nucleic acid (e.g., a target nucleic acid) in an initial sample. In some cases, the determined diversity reduction can be used to determine the Determine the relative abundance of nucleic acids. As shown in FIG. 5, the sample nucleic acids (S 1 , S 2 , ..., S m ) is one or more sample processing steps prior to Spank Spy Quin Synthetic Nucleic Acid (SP 1 , S.P. 2 , ..., S.P. n ) can be combined with, for example, about 10 8 A unique span may be added to the sample. Sample processing (e.g., nucleic acid extraction, During the purification, ligation and / or end repair, a portion of the sample nucleic acid and a portion of the synthetic nucleic acid After sample processing, the first 10 8 Approximately 10 of the unique sequences 6 Individual Uni Then, a portion of these sequences, e.g., 10 4 unique sequences The absolute diversity reduction can be calculated by sequencing or recovering the number of original unique sequences. This can be calculated as the number of unique sequences divided by the number of unique sequences found (e.g., 10 8 / 10 4 =10 4 Similarly, the recovery value is the number of unique sequences sequenced or recovered relative to the first unit. This can be calculated as the number of nucleotides divided by the number of nucleotides in the target gene (e.g., 10 4 / 10 8 =10 - 4 The calculated diversity reduction was used to determine the absolute abundance of nucleic acids in the initial sample. For example, sequencing links for spanked sequences and for sample sequences can be used The number of strands can be determined from the sequence analysis and is the initial concentration of the spanked sequence added to the sample. The degree or amount is known. The determined diversity reduction can be used to determine the degree or amount of the nucleic acid ( For example, the initial concentration or amount of nucleic acid from a particular organism, pathogen, tumor, or organ is determined. The absolute amount of sample nucleic acid in the original sample can be determined by the amount of sample nucleic acid and span nucleic acid. The final number of sequencing reads for an acid and / or the final diversity of spanked nucleic acids Measure and normalize to the known amount or diversity of spanked nucleic acid added to the original sample It can be estimated by
[0179] The number of unique sequence reads can be determined by a variety of methods. For example, The sequence reads can be identified. Then, a unique sequence within the sequence read with an identification tag is identified. The number of sequences is determined by de-duplicating duplicate sequences ("deduplicate"). For example, the sequences can be determined by Align sequences against a reference database or against each other to determine which overlap It is then possible to determine which ones are unique or different. Identification tags are typically conserved across sequences and therefore can be embedded randomly within each added molecule. In some cases, the spanked nucleic acid does not include an identification tag, and In such cases, the spank may be, for example, a reference to a database containing known sequences or can be determined by other methods such as alignment.
[0180] Spank sequences are used to monitor relative and / or absolute losses. In some cases, if the diversity of spank sequences is high enough, it may be possible to add It can be assumed that the spanked sequences obtained are substantially all non-unique. If a determined overlapping spank sequence is present, it is due to PCR amplification and is not a nucleotide sequence. This is unlikely to be due to multiple copies of the same Spank sequence added to the pull, and the analysis Also, if each spank sequence is unique, it can be added to the sample first. The total number of spanked sequences determined is based on the concentration and volume of nucleic acid added to the sample. The total number of unique spanned sequencing reads after sequencing is known. The values can be used together to calculate a diversity reduction or recovery value.
[0181] The method provided by the present invention is characterized by the population bottleneck effect. eck) or methods to identify steps during sample processing that are associated with reduced diversity. In some cases, bottleneck effects have been identified, but initially unknown effects in the starting population Correction factors may be applied to other molecules added. For example, substantially all of the input spank molecules are If the spunk is unique, but only 50% of the spunk collected is unique, this is a Bottleneck effects and other molecular diversity that may provide information for interpretation of the pool and reduced diversity.
[0182] To identify the steps where the bottleneck effect occurs, Thus, a collection of spanks can be added to the sample. For example, when a body fluid is taken from a subject, a first collection of spank can be introduced. Before or during further processing of the collected sample (e.g., removal of residual cells, storage), A second collection of punctures can be introduced into the sample and / or live Before the production of the rally, it is possible to introduce a third set of Spank. The aggregates of spanks added to the sample at different steps during sample processing are have the same or similar composition. In some cases, different collections of spanks are sampled. It is added to the sample at different steps during processing.
[0183] In some cases, each of the spanked nucleic acids comprises a randomized portion having a unique sequence. Spanks may contain one or more different domains. In some cases, spanks The Q includes one or more process codes, one or more diversity codes, one or more length identification sequences, one or more a load sequence, one or more synthetic nucleic acid identifier sequences (or spank identifier sequences), and In some cases, the spank may include an identification tag and / or one or more feature domains. and a unique nucleic acid sequence.
[0184] When various aggregates of sump are used, each aggregate may be used for a specific process (e.g., Identify the spank aggregates introduced into the sample during sample collection, extraction, and library processing. In such cases, the same process may be coded with a "process code" to distinguish Spanks with process codes were bioinformatically classified and analyzed for diversity reduction. The degree of diversity reduction associated with a particular step can then be determined and each sample treatment can then be run. The results can be compared across the entire processing steps.
[0185] Spanks are “diversity codes” that are associated with a whole pool or collection of synthetic nucleic acids or spanks. A diversity coding domain can include a unique sequence that indicates the amount of diversity in a pool of synthetic nucleic acids. In such a case, each synthetic nucleic acid in the diversity pool can be a primer code. The degree of diversity of the pool (e.g., 10 8 The nucleic acid sequence may be encoded by a sequence that represents one or more unique sequences. In some cases, for example when two or more diversity pools are used on the same sample, The diversity code is then used to identify the diversity reduction in those two or more pools. It can be used.
[0186] In some cases, Spanks are members of specific Spank pools or collectives. It may contain one or more codes that identify the spank (e.g., a process code). In this case, the spank is treated as a spank rather than as a nucleic acid originally present in the sample. The Spank identification domain may include one or more Spank identification domains that identify the Spank identification domain. As shown, Spank has a feature domain, a length discriminator domain, and a log domain. It may also include a load domain.
[0187] Spank can be used alone or in combination with other methods to calculate nucleic acid abundances or for other uses. In some cases, the spanks may be used in combination with other synthetic nucleic acids. For example, in some cases, the panels of the Spank and the A panel of nucleic acids (rk) may be added to the sample. In some cases, the sample discriminating nucleic acid may also be added to the sample. It can be added to the pool.
[0188] A spanked pool preferably contains a diverse mixture of nucleic acid sequences. Spank pools can be designed to maximize versatility. are derived from much larger spank pools. For example, in some cases, The oligonucleotide consists of two 8 bp strings of N (e.g., equal proportions of A Spank can be composed of (i) one or more identification tags and (i i) a unique nucleic acid sequence. In some cases, the unique nucleic acid sequence may be a synthetic nucleic acid comprising The columns can be multiple degenerate or random positions, for example, as shown in FIG. As shown, two of the degenerate positions of the 8 bp string are separated by one or more nucleotides. Two exemplary sequences are listed in Table 4. Oligonucleotide design with N of ring 4 16 =4.3×10 9 Different Cories A pool of oligonucleotides containing a total of 16 N. For example, 1 x 10 8 The molecules were added to 1 mL of plasma and the ID spike and spark were generated as described above. If processed as specified, almost all of the spanks will be unique. For example, ,In such cases, more than 90% (i.e., more than 90%) of the spanks, more than 95%; Over 99% are likely to be unique.
[0189] In some cases, the spanked nucleic acid has at least about 20, 30, 40, 50, 60, 7 0, 75, 80, 90, 100, 110, 120, 125, 130, 140, 150, 1 60, 170, 175, 180, 190, 200, 250, 300, 350, 400, 4 It can be 50, 500, 600, 700, 800, 900 or 1000 nucleotides in length. In some cases, the spanked nucleic acid is at most about 20, 30, 40, 50, 60, 7 0, 75, 80, 90, 100, 110, 120, 125, 130, 140, 150, 1 60, 170, 175, 180, 190, 200, 250, 300, 350, 400, 4 It can be 50, 500, 600, 700, 800, 900 or 1000 nucleotides in length. In some cases, the spanked nucleic acid may have a length within the range of about 20 to about 175 base pairs. In some cases, the nucleic acids in the spank set have the same length. In some embodiments, the nucleic acids in a spank set may be of two or more different lengths (e.g., 2, 3, 4, 5 or more bases). or longer).
[0190] In some cases, the spanked nucleic acid comprises at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 2 It may have 5, 26, 27, 28, 29 or 30 degenerate positions. Pancreatic nucleic acid is at most about 10, 11, 12, 13, 14, 15, 16, 17, 18, 1 9, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 degeneracy The spanked nucleic acid can have a number of degenerate positions in the range of about 5 to about 25. In some cases, the degenerate positions are contiguous, separate, or comprise groups of two or more, For example, the degenerate positions may be divided into 2, 3, 4 or 5 groups. In some cases where division is required, the degenerate positions may be divided evenly between the groups. (e.g., two groups of 8 bp strings of degenerate positions for a total of 16 degenerate positions) , or can be split unevenly between groups (e.g., one group One group has 10 degenerate positions, and the other group has 6 degenerate positions, for a total of 16 degenerate positions). In some cases where the overlapping positions are divided into groups, the groups may include one or more In some cases, the groups may be separated by at least Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40 or 50 pieces In some cases, the groups are separated by at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40 or 50 nuclei Separated by a punch hole.
[0191] In some cases, the spanked nucleic acid is at least 1×10 3 , 1×10 4 , 1×10 5 , 1×10 6 , 2×10 6 , 3×10 6 , 4×10 6 , 5×10 6 , 6×10 6 , 7×1 0 6 , 8×10 6 , 9×10 6 , 1×10 7 , 2×10 7 , 3×10 7 , 4×10 7 , 5 ×10 7 , 6×10 7 , 7×10 7 , 8×10 7 , 9×10 7 , 1×10 8 , 2×10 8 , 3×10 8 , 4×10 8 , 5×10 8 , 6×10 8 , 7×10 8 , 8×10 8 , 9×1 0 8 , 1×10 9 , 2×10 9 , 3×109 , 4×10 9 , 5×10 9 , 6×10 9 , 7 ×10 9 , 8×10 9 , 9×10 9 , 1×10 10 or 1×10 11 Unique sequences In some cases, the spanked nucleic acid may have a diversity of at most 1×10 6 , 2×1 0 6 , 3×10 6 , 4×10 6 , 5×10 6 , 6×10 6 , 7×10 6 , 8×10 6 , 9 ×10 6 , 1×10 7 , 2×10 7 , 3×10 7 , 4×10 7 , 5×10 7 , 6×10 7 , 7×10 7 , 8×10 7 , 9×10 7 , 1×10 8 , 2×10 8 , 3×10 8 , 4×1 0 8 , 5×10 8 , 6×10 8 , 7×10 8 , 8×10 8 , 9×10 8 , 1×10 9 , 2 ×10 9 , 3×10 9 , 4×10 9 , 5×10 9 , 6×10 9 , 7×10 9 , 8×10 9 , 9×10 9 , 1×10 10 or 1×10 11Each of the sequences may have a diversity of unique sequences. In some cases, the spanked nucleic acid is about 1×10 4 ~Approx. 1×10 11 Range of unique sequences There may be diversity within the range. [Table 4]
[0192] Tracer Array Laboratory-derived nucleic acids (e.g., pathogen genomic DNA) are essential for the development of infectious disease diagnostic tests. These standards are useful for development, verification, validation, assay control, etc. The same organisms may be present in clinical samples (e.g., pathogen-infected samples) and therefore laboratory-derived There is a risk that substances may cross-contaminate clinical samples during testing, resulting in false positive results. Not only can this provide inaccurate information to patients and physicians, but for some pathogen species: It may also give rise to statutory notification to health authorities. DNA from the original genome, cancer nucleic acid, tumor nucleic acid, or other disease-related nucleic acid) is useful as a positive control. It is possible, or even essential, that ordinary or even special care be taken when handling it. , especially in the case of highly sensitive assays such as next-generation sequencing (NGS), may not be sufficient to provide prevention.
[0193] Synthetic tracer nucleic acids that do not occur in nature or cannot hybridize to the sample nucleic acid are: It may be added to the positive control nucleic acid stock at an effective concentration at least as high as the positive control nucleic acid. The tracer and positive control nucleic acids are prepared in a form such that they are processed and detected in the same manner. Therefore, the endpoint (e.g., aligned sequences in the case of NGS) The sequence reads obtained from the tracer and positive control nucleic acids are the same for both the tracer and positive control nucleic acids. is at least as easily detected as the positive control nucleic acid due to its higher effective concentration. In some cases, the positive control nucleic acid is pathogen genomic DNA. Positive control nucleic acids include disease-associated nucleic acids such as oncogenes.
[0194] The tracer sequence may be characterized in one or more properties, such as sequence, length, concentration, GC content, etc. The sequence shown in Table 5 and used in Example 6 has a GC content of about 50%. However, the tracer sequence must be matched to the composition of the positive control or genome with which it is combined. For example, the GC content can be 30%, 35%, 40% GC content, etc. , 45% GC content, 50% GC content, 55% GC content, 60% GC content, 65% or 70% GC content.
[0195] In some cases, the tracer sequence may be, for example, a fragment, as described in Example 6. In some cases, a positive control nucleic acid or genomic DNA may be added to the positive control nucleic acid after amplification. To better represent the complete treatment performed on the acid or sample nucleic acid, a tracer The sequence may be added to a positive control nucleic acid or genomic DNA prior to fragmentation. Positive control nucleic acids that are rare and found in low concentrations in the human genome (e.g., pathogen DNA) are unlabeled nucleic acids. In order to minimize cross-contamination, the antibodies may be labeled with a tracer sequence as soon as possible.
[0196] In some cases, two or more tracer sequences are added to each positive control nucleic acid. In some cases, two or more, three or more, four or more, or five or more tracer sequences may be used at the same or different concentrations. The concentration of the amine is adjusted to 100%.
[0197] Depending on the application, different forms of tracer sequences may be used. For example, The length can correspond to the length of a reference sequence, e.g., the average or median length. In some cases, the length of the tracer sequence is 5 times the mean or median length of the control sequences. %, 10% or 20%.
[0198] For RNA applications, RNA tracer sequences may be used. [Table 5] TIFF2025060811000021.tif202166
[0199] molecular LIMS A Laboratory Information Management System (LIMS) is a method for tracking the consumption and use of consumables. In some cases, chemicals or reagents required for a given experiment and A method for ensuring that only chemicals or reagents necessary for an experiment are used in that experiment The LIMS tracks the lot numbers of the chemicals used in each iteration of an experiment. All of these functions (e.g., tracking lot numbers) can be useful for, for example, Problems of experimental failure when the quality of materials deteriorates or when the wrong reagents are used in the experiment It may help the solution.
[0200] The LIMS system records the catalog number and It is designed as an electronic or web application where laboratory personnel enter the drug and lot number. Typically, bar codes are used to facilitate and improve the accuracy of the process. However, human error can still result in incomplete recording of a given replicate of a reaction. do.
[0201] The present invention provides a method for molecularly labeling a reagent, particularly a reagent, reagent lot, aliquot, or shipment. In some cases, the method includes molecularly binding different containers of various reagents. This includes the use of spike-in synthetic nucleic acids to encode unique sequences (e.g. A spike-in nucleic acid or a short nucleic acid oligomer (e.g., a non-human, non-pathogen) 50-100bp) is added to each reagent, reagent lot, reagent aliquot, or reagent shipment. It is important to keep track of the stock of reagents used to produce the individual libraries. In some cases, one or more ID spikes, Spark or Spank sequences can be used in molecular LIMS The lot numbers and reagents used in processing each sample were then identified by sequencing. It is possible to automatically detect the lock used in successful implementations, for example. or to compare the missing or missing These can be used to troubleshoot problematic runs by identifying excess or redundant reagents. It is Noh.
[0202] Similarly, spikes associated with a particular reagent, reagent lot number, aliquot, or shipment Detection of nucleic acids is dependent on lot numbers, aliquots, and amounts of reagents used in successful sequencing runs. In some cases, nucleic acid or spikes may be used to identify the product or shipment. The intron can be detected by methods other than sequencing, for example by one or more fluorescent probes. A generic polymer labeled with a lobe can be detected using fluorescence.
[0203] DNA oligomers are effective in many aqueous solutions, but nucleic acids that are not susceptible to DNase action Oligomers (e.g., RNA, DNA oligomers with modified backbones) are incubated in a DNase-containing solution. Similarly, synthetic nucleic acids (e.g., DNA) that are resistant to RNases can be designed for use in , can be used to trace RNase-containing solutions.
[0204] Nucleic Acid Enrichment and Library Production In the methods provided herein, nucleic acids are synthesized by any means known in the art, For example, nucleic acids can be isolated from a sample by liquid extraction (e.g., Trizol). Nucleic acids can be extracted using commercially available kits. [QIAamp Circulating Nucleic Acid Kit] Kit), Qiagen DNeasy Kit, QIAamp Kit, Qiagen M idi kit, QIAprep].
[0205] Nucleic acids may be concentrated or precipitated by known methods including, by way of example only, centrifugation. The acid can be bound to a selective membrane (e.g., silica) for purification purposes. The nucleic acid can be purified to a desired length. Fragments, e.g., less than 1000, 500, 400, 300, 200 or 100 base pairs in length Fragment enrichment can also be achieved by size-based enrichment, for example by PEG-induced precipitation. Electrophoretic gel or chromatographic material (Huber et al. (1993) Nuclei c Acids Res.21:1061-6), gel filtration chromatography or T SK gel (Kato et al. (1984) J. Biochem, 95:83-86) (those (the entirety of which is incorporated herein by reference for all purposes) It can be done.
[0206] The nucleic acid sample is subjected to analysis for a target polynucleotide, particularly a target nucleic acid associated with inflammation or infection. In some preferred cases, the target nucleic acid may be a pathogen nucleic acid (e.g., a cellular In some preferred cases, the target nucleic acid is a nucleic acid derived from the uterus, heart, lung, kidney, or the like. , specific organs including, but not limited to, fetal brain, liver or cervical tissue. or tissue-associated cell-free RNA.
[0207] Target enrichment can be by any means known in the art. For example, nucleic acid samples Pulling is performed using target-specific primers (e.g., primers specific for pathogen nucleic acids). The target sequence can be enriched by amplifying it. Target amplification can be by any method known in the art. The method or system may be used in a form of digital PCR. Enrichment by capturing target sequences on arrays immobilized with selective oligonucleotides The nucleic acid sample can be prepared by mixing target-selective oligonucleotides free in solution or on a solid support. The oligonucleotides can be enriched by hybridizing to capture oligonucleotides. In some embodiments, the nucleic acid sequence may include a capture moiety that allows for capture by a reagent. The samples are not enriched with respect to the target polynucleotide, e.g., represent the entire genome.
[0208] In some cases, the target (e.g., pathogen organ) nucleic acid is subjected to, e.g., pull-down. -down) (e.g., a complementary oligonucleotide attached to a label such as a biotin tag) The target nucleic acid is hybridized to the solid support, for example, avidin or By using streptavidin, the target nucleic acid is preferentially isolated in a pull-down assay. First, pull down the markers in the sample, using targeted PCR or other methods. The nucleic acid sequence can be enriched against background (e.g., subject, healthy tissue) nucleic acids. Examples of enrichment techniques include (a) a majority population of nucleic acids in a sample self-cleaves more rapidly than a minority population of nucleic acids in the sample; (b) nucleofection from free DNA; (c) depletion of endothelial-associated DNA, to remove and / or isolate DNA of specific length intervals (d) exosome depletion or enrichment, and (e) strategic capture of regions of interest. However, it is not limited to these.
[0209] In some cases, the enrichment step includes (a) providing a sample of nucleic acid from the host (wherein Here, the sample of nucleic acid from the host is a sample of single-stranded nucleic acid from the host, and the host nucleic acid and (b) renaturing at least a portion of the single-stranded nucleic acid from the host; thereby generating a population of double-stranded nucleic acids in the sample; and (c) detecting the presence of nucleic acids in the sample. and removing at least a portion of the double-stranded nucleic acid in the sample using a nucleic acid sequence encoding the nucleic acid sequence of the host. In some cases, the enrichment step includes enriching for non-host sequences in a sample of nucleic acid from The enrichment step includes (a) providing a sample of nucleic acid from a host, (containing host and non-host nucleic acids associated with nucleosomes), and ( b) removing at least a portion of the host nucleic acid associated with nucleosomes, thereby removing the nucleic acid from the host; In some cases, enrichment includes enriching the non-host nucleic acid in a sample of nucleic acid from the host. The steps include (a) providing a sample of nucleic acid from a host, (b) a sample containing one or more length intervals of DNA; removing or isolating, thereby enriching for non-host nucleic acids in a sample of nucleic acid from a host. In some cases, the enrichment step includes (a) providing a sample of nucleic acid from the host. wherein the sample of nucleic acid from the host is a mixture of host nucleic acid, non-host nucleic acid and exosome nucleic acid. (b) removing or isolating at least a portion of the exosomes, thereby and enriching for non-host sequences in a sample of nucleic acid from a host. In one embodiment, the enrichment step preferentially removes nucleic acids from the sample that have a length of about 300 bases or greater. In some cases, the enrichment step includes preferentially isolating non-host nucleic acids from the sample. This includes amplifying or capturing.
[0210] The enrichment step may include enriching samples with nucleic acids greater than or equal to about 120, about 150, about 200, or about 250 bases in length. In some cases, the enrichment step may include preferentially removing from the pool about 10 bases to about 10 bases. Approximately 60 bases long, approximately 10 bases to approximately 120 bases long, approximately 10 bases to approximately 150 bases long, approximately 10 bases long ~300 bases long, ~30 bases long, ~60 bases long, ~120 bases long, ~30 bases long 10 to about 150 bases in length, about 30 to about 200 bases in length, or about 30 to about 300 bases in length In some cases, the enrichment step comprises preferentially enriching nucleic acids from a sample. , which involves preferentially digesting nucleic acids from a host (e.g., a subject). The enrichment step involves preferentially replicating non-host nucleic acid.
[0211] In some cases, the enrichment step increases the ratio of non-host nucleic acids to host (e.g., subject) nucleic acids. At least 2x, at least 3x, at least 4x, at least 5x, at least 6x times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least 14 times, at least 15 times, At least 16 times, at least 17 times, at least 18 times, at least 19 times, at least At least 20 times, at least 30 times, at least 40 times, at least 50 times, at least 60 times , at least 70 times, at least 80 times, at least 90 times, at least 100 times, at least The increase may be at least 1000-fold, at least 5000-fold, or at least 10,000-fold. In some cases, the enrichment step increases the ratio of non-host nucleic acid to host (e.g., subject) nucleic acid. In some cases, the enrichment step increases the abundance of the host (e.g., subject) nucleic acid by at least 10-fold. The ratio of non-host nucleic acid to the total nucleic acid is increased by about 10-fold to about 100-fold.
[0212] In some cases, a nucleic acid library is produced. A nucleic acid library is a single-stranded nucleic acid library. The nucleic acid library may be a library or a double-stranded nucleic acid library. In some cases, a single-stranded nucleic acid library may be used. The library is a single-stranded DNA library (ssDNA library) or an RNA library. In some cases, the double-stranded nucleic acid library can be a double-stranded DNA library. The method for producing a ssDNA library is to use double-stranded DNA. The NA fragment was denatured into a ssDNA fragment, and the primer docking sequence was ligated to a position of the ssDNA fragment. and hybridizing a primer to the primer docking sequence. The primers are then paired with the adapters to be sequenced using next-generation sequencing platforms. The method may further comprise extending the hybridized primer. and generating a duplex, wherein the duplex is a ssD The NA fragment and the extended primer strand are derived from the original ssDNA fragment. The extended primer strand can be recovered, where the extended primer strand is s The members of the sDNA library are used in the production of RNA libraries. The docking sequence is linked to one end of the RNA fragment, and the primer is linked to the primer docking sequence. The primer may be hybridized to a next-generation sequencing platform. The method may further include at least a portion of an adaptor sequence combined with the hybridization form. extending the redacted primer to generate a double strand; Here, the duplex comprises the original RNA fragment and the extended primer strand. The extended primer strand can be recovered, where the The extension primer strand is a member of the RNA library. The method of construction can include ligating an adapter sequence to one or both ends of the dsDNA fragment. .
[0213] In various embodiments, the dsDNA can be any of the nucleic acids known in the art or described herein. In some cases, dsDNA can be fragmented by any means known in the art. by mechanical shearing, spraying or sonication, enzymatic or chemical means. It can be fragmented.
[0214] In some embodiments, cDNA is generated from RNA. For example, random profiling is used. Using Lyme reverse transcription (RNase H+) to obtain cDNA of random size It is possible to generate cDNA by this method.
[0215] The length of the nucleic acid can vary. , or random size cDNA) are less than 1000bp, less than 800bp, less than 700 < bp, < 600bp, < 500bp, < 400bp, < 300bp, < 200 The DNA fragment is about 40 to about 100 bp, about 50 to about 125bp, about 100 to about 200bp, about 150 to about 400bp, about 300 to about 500b p, about 100 to about 500, about 400 to about 700bp, about 500 to about 800bp, about 700 about 900 bp, about 800 to about 1000 bp, or about 100 to about 1000 bp. In some cases, the nucleic acid or nucleic acid fragment (e.g., dsDNA fragment, RNA, or RNA) The standard size of the cDNA is within the range of about 20 to about 200 bp, for example, about 40 to about 100 bp. It can be in the range of p.
[0216] The ends of the dsDNA fragments can be polished (e.g., blunted). The ends of DNA fragments can be polished by treatment with a polymerase. Removal of 3' overhangs and filling in 5' overhangs (fill-in) The polymerase may comprise a proofreading polymerase, The enzyme may be an enzyme having, for example, a 3' to 5' exonuclease activity. Freeing polymerases include, for example, T4 DNA polymerase, Pol 1 cDNA polymerase, The polymerase may be any of the polymerases known in the art. This may include using a means to remove damaged nucleotides (eg, abasic sites).
[0217] The ligation of an adaptor to the 3' end of a nucleic acid fragment is carried out by connecting the 3' OH group of the fragment to the 5' OH group of the adaptor. Thus, the 5' phosphate from the nucleic acid fragment may be formed. Removal of the fragments can minimize aberrant ligation of two library members. In some embodiments, the 5' phosphate is removed from the nucleic acid fragment. In the form, the 5' phosphate is present in at least 50%, 5 5%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 95% In some embodiments, substantially all of the phosphate groups (phosphorus In some embodiments, substantially all of the phosphatase groups are removed from the nucleic acid fragments. The fragments are at least 50%, 55%, 60%, 65%, 70%, or 80% of the nucleic acid fragments in the sample. %, 75%, 80%, 85%, 90%, 95% or greater than 95% are removed from the nucleic acid sample. Removal of the phosphate group from the pool can be by any means known in the art. Removal of phosphate groups involves treating the sample with a heat-labile phosphatase. In some embodiments, the phosphate groups are not removed from the nucleic acid sample. In some embodiments, adaptors are ligated to the 5' ends of the nucleic acid fragments.
[0218] Sequencing The present disclosure provides methods for analyzing nucleic acids. Such methods include sequencing and sequencing of nucleic acids. This includes bioinformatics analysis of the results of the determination. The nucleic acids analyzed include genomic, epigenetic (e.g., methylation) and RNA expression. Methylation analysis can be used to obtain various types of information. RNA expression analysis can be performed using, for example, polyclonal antibodies. Ligand array hybridization, RNA sequencing technology, or RNA from This can be done by sequencing the cDNA generated.
[0219] In a preferred embodiment, sequencing is performed using a next generation sequencing assay. The term "next generation" as used herein is well understood in the art. and generally any high throughput sequencing method, including but not limited to one or more of the following: Defined approaches include: massively parallel signature sequencing, pyrosequencing (e.g., Roche 454 sequencer), Illumina ( Solexa sequencing, Illumina sequencing, Ionto Ion torrent sequencing, sequencing by ligation (e.g., SOLiD sequencing), single molecule real-time (SMRT) sequencing (e.g., Pacific Biology oscience, polony sequencing, DNA nanoballs ball sequencing, heliscope single molecule ecule sequencing (Helicos Biosciences) and nanopore (n anopore) sequencing (e.g., Oxford Nanopore). In some cases, the sequencing assay uses nanopore sequencing. Sequencing includes some form of Sanger sequencing. In some cases, sequencing is done by shot. In some cases, sequencing may involve bridge PCR, including cancer sequencing. In some cases, the sequencing is broad spectrum. The decision is targeted.
[0220] In some cases, the sequencing assay comprises Gilbert's sequencing method. In this approach, nucleic acids (e.g., DNA) are chemically modified and then cleaved at specific bases. In some cases, the sequencing assay detects dideoxynucleotide chain termination or Including Sanger sequencing.
[0221] In the methods provided herein, a sequencing by synthesis approach can be used. In some cases, fluorescently labeled reversible terminator nucleotides were used in glass flow cytometers. The DNA is introduced to clonally amplified DNA templates immobilized on the surface of the gel during each sequencing cycle. In this manner, a single labeled deoxynucleoside triphosphate (dNTP) can be added to the nucleic acid strand. Terminator nucleotides are imaged as they are added to identify the base. can then be enzymatically cleaved to allow incorporation of the next nucleotide All four reversible terminator-bound dNTPs (A, C, T, G) are generally Because they exist as one separate molecule, natural competition can minimize uptake bias.
[0222] In some cases, a method called single molecule real-time (SMRT) is used. In such an approach, nucleic acids (e.g., DNA) are introduced into a zero-mode waveguide (ZMW). The ZMW is a small well-like membrane that contains a capture means located at the bottom of the well. The vessel contains unmodified polymerase (bound to the bottom of the ZMW) and free-flowing in solution. Sequencing is performed using fluorescently labeled nucleotides that are incorporated into the DNA strand. Upon incorporation, the DNA strand is separated from the nucleotide, leaving an unmodified DNA strand. Such detectors can be used to detect luminescence. It is possible to obtain sequence information by
[0223] In some cases, a sequencing by ligation approach is used to identify Sequence the nucleic acid. One example is SOLiD [oligonucleotide ligation and detection Sequencing by Oligonucleotide Lig ation and Detection)] sequencing (Life Technology This next-generation technology can detect hundreds of millions to billions of DNA fragments. The sequencing method can generate small sequence reads simultaneously. This may involve producing a library of DNA fragments from the sample to be analyzed. uses the library to generate a unique marker on the surface of each bead (e.g., a magnetic bead). A population of clonal beads is prepared, each of which contains a single fragment. The fragments bound to the magnetic beads are One may have a universal P1 adaptor sequence attached such that the starting sequences of the pieces are known and identical. In some cases, the method may further include PCR or emulsion PCR. For example, Emulsion PCR may involve the use of microreactors that contain the reagents for PCR. The resulting bead-bound PCR products are then covalently attached to a glass slide. Sequencing assays, such as the SOLiD sequencing assay or ligation assay, are possible. Other sequencing by SEQ ID NO:1 may include a step involving the use of primers. The hybridization sequence may hybridize to an adapter sequence or to another sequence within the library template. In addition, four fluorescently labeled dinucleotide probes that compete for ligation to the sequencing primer were introduced. The specificity of the di-base probe can include inserting each of the first and second bases in each ligation reaction. This can be achieved by interrogating the second base. Multiple cycles of ligation, detection and cleavage are performed with the number of cycles determining the final read length. In some cases, after a series of ligation cycles, the extension products can be removed and a second round of For the ligation cycle, a primer complementary to position n-1 is used to reposition the template. For each sequence tag, multiple rounds (e.g., 5 rounds) of primer repositioning are completed. Through the primer relocation process, each base can be cleaved by two different primers. For example, at lead position 5, The bases in the ligation cycle are identified by primer number 2 in ligation cycle 2 and by primer number 3 in ligation cycle 3. 1 is assayed with primer number 3.
[0224] In any of the embodiments, the detection or quantitative analysis of the oligonucleotides includes sequencing. Subunit or fully synthetic oligonucleotides may be prepared according to the methods described herein. Any suitable method known in the art, including the sequencing methods described in Complete sequencing of all oligonucleotides using the umina HiSeq 2500 It can be detected.
[0225] Sequencing may be accomplished by classical Sanger sequencing methods well known in the art. Sequencing can also be accomplished using high-throughput systems; Some of these are characterized by the fact that the sequencing nucleotide occurs immediately or after its incorporation into the growing chain. is detected upon said capture (e.g., in real time or substantially in real time). In some cases, high-throughput sequencing allows for the detection of sequences in a single At least 1,000 per hour, at least 5,000, at least 10,000, At least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000 or at least 500,000 sequence In some cases, each read generates at least 50 reads. , at least 60, at least 70, at least 80, at least 90, at least 10 0, at least 120, or at least 150 bases. In some cases, each read at most 2000, at most 1000, at most 900, at most 800 per lead , at most 700, at most 600, at most 500, at most 400, at most 300 , at most 200 or at most 100 bases. Long read sequencing can be, for example, longer than 00 bases, longer than 800 bases, longer than 1000 bases, longer than 1500 bases , longer than 2000 bases, longer than 3000 bases, or longer than 4500 bases. This may include sequencing to provide a code.
[0226] In some cases, high-throughput sequencing is performed using Illumina's Genome Analyzer IIX, MiSeq personal sequencer, or HiSeq system systems such as HiSeq 2500, HiSeq 1500, and HiSeq 2000. or using the techniques available from the HiSeq 1000. The devices utilize reversible terminator-based sequencing by synthetic chemistry. The device can perform more than 200 billion DNA reads in eight days. Smaller systems include It can be used for implementation within 3, 2 or 1 day or less. Fast synthesis cycles can be used to minimize the time required to obtain a sequencing result. .
[0227] In some cases, high throughput sequencing is performed using the ABI Solid System This genetic analysis platform involves the use of technology available from The sequencing method may allow for massively parallel sequencing of cloned, clonally amplified DNA fragments. It is based on sequential ligation with labeled oligonucleotides.
[0228] Next generation sequencing is ion-based sequencing (e.g., Life Technologies, Inc., gies (Ion Torrent technology) Sequencing is achieved by releasing ions as nucleotides are incorporated into a DNA strand. The fact that high density arrays of microfabricated wells can be used to perform ion-semiconductor sequencing. Each well can hold a single DNA template. There may be an ion-sensitive layer, and below the ion-sensitive layer is an ion sensor. When a nucleotide is added to DNA, H+ is released. This can be measured as a change in pH. H+ ions are converted into a voltage, which The array chip is immersed in water one nucleotide at a time. Scans, lights and cameras may not be necessary. In some cases, IONPROT An ON™ sequencer is used to sequence the nucleic acid. The IONPGM™ Sequencer is used. Ion Torrent Personal Genome Machine The PGM can perform 10 million reads in two hours.
[0229] In some cases, high throughput sequencing is performed using Helicos BioScience. ces Corporation(Cambridge,Massachusetts) Technologies available today, such as Single Molecule Sequencing by Synthesis (SMS) Use of the Single Cell Sequencing by Synthesis (SMSS) method SMSS may make it possible to sequence the entire human genome within 24 hours. SMSS, like MIP technology, does not require a pre-amplification step prior to hybridization. SMSS would not require amplification. SMSS is described in U.S. Patent Application Publication No. 200 No. 60024711, No. 20060024678, No. 20060012793, No. 2 Nos. 0060012784 and 20050100932.
[0230] In some cases, high-throughput sequencing is performed using 454 Lifesciences Technologies available from Intel, Inc. (Branford, Connecticut), such as For example, the use of a Pico Titer Plate device, which includes a CCD camera within the device. A photofluorescence detector records the chemiluminescence signal generated by the sequencing reaction. This use of optical fiber results in at least 2,000 readings in 4.5 hours. This may allow detection of millions of base pairs.
[0231] A method for using bead amplification and subsequent fiber optic detection is described in Margui et al. les, M. et al., “Genome sequencing in microfabri cated high-density picolitre reactors”,N Nature, doi:10.1038 / nature03959, and U.S. patent application Publication No. 20020012930, No. 20030058629, No. 200301001 No. 02, No. 20030148344, No. 20040248161, No. 2005007 9510, 20050124022 and 20060078909 is.
[0232] In some cases, high-throughput sequencing is used to determine Clonal Single Molecular Using the lecule array (Solexa, Inc.) or reversible terminator These techniques are carried out using sequencing by synthesis (SBS) which utilizes star chemistry. , U.S. Patent Nos. 6,969,488, 6,897,023, and 6,833,246 , 6,787,308, and U.S. Patent Application Publication No. 20040106110, 2 No. 0030064398, No. 20030022207, and Constans, A. ,Partially described in The Scientist 2003,17(13):36 do.
[0233] In some cases, next generation sequencing is performed using nanopore sequencing. (See, e.g., Soni GV and Meller A. (2007) Clin Che (See, for example, U.S. Pat. No. 5,333,1996-2001.) Nanopores can be, for example, about 1 nanometer in diameter. The nanopore can be a small hole on the order of 1000 nm. The nanopore is immersed in a conductive fluid and a pore beyond it is When a potential is applied, a small current can be generated by the conduction of ions through the nanopore. The amount of current that flows can be sensitive to the size of the nanopore. When DNA is inserted into the nanopore, each nucleotide on the DNA molecule can block the nanopore to a different extent. When a DNA molecule passes through the nanopore, the change in the current passing through the nanopore indicates the reading of the DNA sequence. Nanopore sequencing technology is a technology developed by Oxford Nanopore Technology. ologies, for example the GridION system. A single nanopore can be inserted into the polymer membrane across the top of each microwell. The microwells can have electrodes for sensing up to 100,000 per chip. More than 100,000 (e.g., 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000 or 1,000, The chip can be fabricated into an array chip having microwells (0,000 or more). The devices (or nodes) can be used to analyze the data in real time. One or more devices can be operated simultaneously. The nanopore can be a protein nanopore, e.g., a heptapore. The nanopore may be a protein alpha-hemolysin, which is a polymer protein nanopore. It can be in a solid state, for example a synthetic membrane (e.g., SiNx, Or SiO 2 Nanopores can be nanometer-sized holes formed in a hybrid membrane. The nanopore can be a collection of nanopores (e.g., a protein pore incorporated into a solid membrane). Interaction sensors (e.g., tunneling electrode detectors, capacitive detectors, or graphene-based nano-gate Top or edge detectors (e.g., Garaj et al. (2010) Nature vol 67, doi:10.1038 / nature09379) The nanopore can be a nanopore that can detect the flow of a particular type of molecule (e.g., DNA, RNA or Nanopore sequencing can be functionalized to analyze proteins. In some cases, the method may involve "transcriptomic sequencing," in which an intact DNA polymer is As the DNA moves into the pore, it is sequenced in real time while the protein The enzyme separates the strands of double-stranded DNA and delivers the strands through the nanopore. The DNA can have a hairpin at one end and the system can read both strands. In some cases, nanopore sequencing is "exonuclease sequencing," which In the case of The nucleotide is then cleaved from the strand and allowed to pass through the protein nanopore. The characteristic disturbance in the current ( It is possible to identify bases using the method known as "disruption."
[0234] Nanopore sequencing technology from GENIA can be used. The nanopore can be embedded in a lipid bilayer membrane. Efficient nanopore-membrane assembly and DNA channeling To allow control of A movement, "active control" techniques can be used. In this case, the nanopore sequencing technology is from NABsys. Genomic DNA is approximately 10 The 100 kb fragments can be made single-stranded. which can then be hybridized to a 6-mer probe. The genome fragments can be threaded through the nanopore, which provides current versus time tracings. The current tracings can show the location of the probes on each genomic fragment. It is possible to align the genomic fragments of a genome and generate a probe map for the genome. The process can be performed in parallel on a library of probes. A search map can be created. The error is "Moving window sequencing by hive Moving Window Sequencing By Hyb These can be modified through a process called "multi-layered hysteresis homogenization (mwSBH)." In this case, the nanopore sequencing technology is from IBM / Roche. It is possible to generate nanopore-sized openings in microchips using The field can be used to pull or thread DNA through the nanopore. The DNA transistor device in the cell is made of alternating nanometer-sized layers of metal and dielectric. The isolated charges in the DNA backbone can be captured by the electric field in the DNA nanopore. By turning the gate voltage off and on, it may be possible to read the DNA sequence. .
[0235] Next generation sequencing is DNA nanoball sequencing (e.g., C For example, those performed by Drmanac et al. (2010) Science 327:78-81). A can be isolated, fragmented, and size selected. For example, DNA can be isolated, fragmented, and size selected (e.g., The DNA can be fragmented to an average length of about 500 bp (by sonication). The adapter is then attached to the end of the fragment, which is then hybridized to the anchor for sequencing reaction. The DNA with the adapters attached to each end is amplified by PCR. The adapter sequence is a sequence that allows the complementary single-stranded ends to be ligated together to form a circular DNA fragment. The DNA can be modified to form A. The adaptor (e.g., the right adaptor) may be methylated to protect it from cleavage by It is possible for the restriction recognition site to remain unmethylated. The unmethylated restriction sites in the adapter are recognized by restriction enzymes (e.g., AcuI). The DNA can be amplified by AcuI 13 bp to the right of the right adapter. The second round of right and left adaptors (A d2) can be ligated to either end of the linear DNA, and both adapters It is then possible to PCR amplify (eg, by PCR) all of the DNA to which it is attached. The Ad2 sequences can be modified so that they can join together to form a circular DNA. NA can be methylated, but the restriction enzyme recognition site is unmethylated on the left Ad1 adaptor. Restriction enzymes (e.g., AcuI) can be applied, and the DNA can be Ad1. The right and left amplicons of the third round can be cut 13 bp to the left to form a linear DNA fragment. Adaptors (Ad3) can be ligated to the right and left of the linear DNA, resulting in The resulting fragments can be amplified by PCR. Type III restriction enzymes (e.g., EcoP15) can be modified to form .gamma.-dimer DNA. EcoP15 is added 26 bp to the left of Ad3 and 26 bp to the right of Ad2. This cleavage removes large DNA segments and re-linearizes the DNA. The fourth round right and left adaptors (Ad4) are ligated to the DNA and the DNA is (e.g. By amplifying and modifying the DNA fragments (e.g., by PCR), they are then joined together to form the complete circular DNA template. It is possible to form the
[0236] Rolling circle replication (e.g., using Phi 29 DNA polymerase) ) can be used to amplify small fragments of DNA. The four adapter sequences are , a single strand may contain a hybridizable palindromic sequence, and may be complementary to itself. The folding results in a DNA nanobodies that can have an average diameter of about 200 to 300 nanometers. The DNA nanoballs can be used to form microarrays (sequencing frames). The flow cell can be attached (e.g., by adsorption) to a fluororesin (flow cell). The flow cell can be made of silicon dioxide, titanium, Silicon wafers coated with silane and hexamethyldisilazane (HMDS) The sequencing is performed by linking a fluorescent probe to DNA. This can be done by unchained sequencing. The fluorescent color of the terrogated area can be visualized by a high-resolution camera. The nucleotide sequence identity between the target sequences can be determined.
[0237] The methods provided herein can include the use of a system, e.g., is a nucleic acid sequencer (e.g., DNA sequencer, RNA sequencer) to obtain RNA sequence information. The system may include the use of a system including a sequencer (sequencer), which is capable of detecting DNA or RNA sequences. Software that performs bioinformatics analysis on sequence information Bioinformatics analysis may include, but is not limited to, a computer comprising: However, the construction of sequence data, genetic variants in samples [germline variants and body mass index variants], Cellular variants (e.g., genetic variants associated with cancer or precancerous conditions, genetic variants associated with infections) ) and quantification thereof.
[0238] Sequence data can be used to determine gene sequence information, ploidy state, and identity of one or more genetic variants. and quantitative measures of the variants (including relative and absolute relative measures) can be determined. It is Noh.
[0239] In some cases, genome sequencing may involve whole genome sequencing or partial genome sequencing. Sequencing can be unbiased and can be used to determine the Sequencing of all or substantially all (e.g., greater than 70%, 80%, 90%) of the nucleic acids in a The sequencing of the genome can be selective, e.g., to determine the It can be directed to a portion of the genome, e.g., a number of genes (and these Mutations in the gene (such as genomic DNA) are known to be associated with various cancers. Sequencing of selected genes or parts of genes may be sufficient. A polynucleotide that has been mapped to a particular locus within a genome can be, for example, For example, the nucleic acid can be isolated for sequencing by sequence capture or site-specific amplification.
[0240] Purpose The methods provided herein can be used for a variety of purposes, including treating a condition (e.g., an infection) Diagnosis or detection of a condition, predicting the onset or recurrence of a condition, monitoring treatment, or selecting a treatment regimen This approach can be used to improve treatment and and / or diagnostic regimens may be individualized and tailored according to data obtained at various time points throughout the course of treatment. and adjust the dosage, thereby providing an individually appropriate regimen.
[0241] Condition detection / diagnosis / prognosis The methods provided herein include detecting infection or malaria in a patient sample, e.g., a human blood sample. The method may be used to detect, diagnose or prognose a disease or condition. It can be used to detect rare microbial nucleic acid fragments in samples that consist primarily of For example, cell-free DNA (cfDNA) in blood is composed mainly of DNA fragments derived from the host. It also contains small fragments from the body's microorganisms. Extraction of cfDNA and subsequent detailed Sequencing (e.g., next generation sequencing or NGS) is a method for determining the identity of host and non-host genomes. Millions or billions of sequence reads that can be mapped to a database Similarly, the method may generate a readout of circulating or It can also be used to detect rare populations of cell-free RNA. In the case of a very small number of samples, the method provided by the present invention does not allow for the identification of different target nucleic acids. For comparison or with different samples (e.g., from different microorganisms or organisms) Sensitivity of the assay is compromised by the lack of an internal normalization standard to track assay parameters or reagents. The method provided herein can improve the specificity and the ability to detect target nucleic acids. may be used when the population comprises a larger proportion of the total population.
[0242] The methods provided herein are useful for detecting, monitoring, diagnosing, and prognosing a wide variety of diseases and disorders. In particular, the method may be used to detect, treat or prevent an infectious disease or disorder. It can be used to detect one or more target nucleic acids from a relevant pathogen. Diseases and disorders include any disease or disorder associated with an infection, such as sepsis, pneumonia, tuberculosis, , HIV infection, Hepatitis infection (e.g., Hep A, B, or C), Human Papillomavirus (HPV) infection, Chlamydia infection, Syphilis infection, Ebola infection, Staphylococcus aureus Staphylococcus aureus infection or influenza The methods provided herein are directed to the treatment of drug-resistant microorganisms, e.g., multidrug-resistant microorganisms, or It is particularly useful for detecting infections by microorganisms that cannot be cultured or typically tested for in vitro. Some non-limiting examples of diseases and disorders that can be detected by this method include: Included are: cancer, dilated cardiomyopathy, Guillain-Barre syndrome, multiple sclerosis, tuberculosis, anthrax, Sleeping sickness, dysentery, toxoplasmosis, ringworm, candidiasis, histoplasmosis, Ebola, Netobacter infection, actinomycosis, African sleeping sickness (African trypanosomiasis), AIDS ( Acquired immune deficiency syndrome, HIV infection, amebiasis, anaplasmosis, anthrax, alkane Bacterium haemolyticum (Arcanobacterium haemolyti cum) infection, Argentine hemorrhagic fever, ascariasis, aspergillosis, astrovirus infection, babesiosis, Bacillus cereus infection, Bacterial pneumonia, bacterial vaginosis (BV), Bacteroides infection, Balantidiosis, Baylissus Baylisascaris infection, BK virus infection, black sand hair, blastocyst infection , blastoderm infection, Bolivian hemorrhagic fever, Borrelia infection, botulism (and infant botulism) Toxicosis), Brazilian hemorrhagic fever, Brucellosis, Bubonic plague, Burkholderia infection, Buruli ulcer , Calicivirus infection (Norovirus and Sapovirus), Campylobacteriosis, Cannabidiol Diderasis (monilia; vaginitis), cat scratch disease, cellulitis, Chagas disease (American trypanosoma) varicella, chancroid, chickenpox, chikungunya fever, chlamydia, Chlamydophila pneumoniae Chlamydophila pneumoniae infection (Taiwan Acute Respiratory Syndrome) or TWAR), cholera, chromoblastomycosis, clonorchiasis, Clostridium Clostridium difficile infection, Coccidioides Colorado tick fever (CTF), common cold (acute viral nasopharyngitis; acute coryza), Itzfeldt-Jakob disease (CJD), Crimean-Congo hemorrhagic fever (CCHF), Cryptococcosis Cryptosporidiosis, CLM, Cyclosporosis, Cysticercosis , cytomegalovirus infection, dengue fever, diamoebiasis, diphtheria, diphyllobothriasis, Dinariasis, Ebola hemorrhagic fever, echinococcosis, ehrlichiosis, enterobiasis (pinworm infection), intestinal Cocci infection, enterovirus infection, epidemic typhus, erythema infectiosum (fifth disease), exanthema subitum ( 6th disease), fascioliasis, fascioliasis, filariasis, Clostridium perfringens Food poisoning caused by (Clostridium perfringens), a free-living amoeba Infection, Fusobacterium infection, Gas gangrene (Clostridial myonecrosis), Geotrichum, Gerstmann-Sträussler-Scheinker syndrome (GSS), giardiasis, glanders , Gnathostomiasis, Gonorrhea, Granuloma venereum (Donovanosis), Group A Streptococcus infection, Group B Streptococcus infection Dye, Haemophilus influenzae Infection, Hand, Foot and Mouth Disease (HFMD), Hantavirus Pulmonary Syndrome (HPS), Heartland artland) viral disease, Helicobacter pylori pylori) infection, hemolytic uremic syndrome (HUS), hemorrhagic fever with renal syndrome (HFRS), A Hepatitis B, Hepatitis C, Hepatitis D, Hepatitis E, Herpes simplex, Histoplasmosis, Hookworm infection, human bocavirus infection, human Ewingi ehrlichiosis human granulocytic anaplasmosis (HGA), human metanucleosis Virus infection, human monocytic ehrlichiosis, human papillomavirus (HPV) infection, Human parainfluenza virus infection, hymenococcosis, Epstein-Barr virus infection Mononucleosis (Mono), influenza (flu), isosporosis, Kawasaki disease, keratitis, Kingella kingae infection, kuru, Lassa fever, Legionella Legionnaires' disease, Legionnaires' disease (Pontiac fever), Leishmaniasis, Han Seneciovirus, leptospirosis, listeriosis, Lyme disease (Lyme borreliosis), lymphatic fibrosis Lariasis (elephantia), lymphocytic choriomeningitis, malaria, Marburg hemorrhagic fever (MHF) , measles, Middle East Respiratory Syndrome (MERS), melioidosis (Whitmore's disease), meningitis, meningococcal Sexual diseases, schizophrenia, microsporidiosis, molluscum contagiosum (MC), monkeypox, mumps, typhus (epidemic typhus), mycoplasma pneumonia, mycetoma, myiasis, neonatal conjunctivitis (neonatal Ophthalmia of newborns), (new) variant Creutzfeldt-Jakob disease (vCJD, nvCJD), Nocal Diarrhea, Onchocerciasis (river blindness), Paracoccidioides (South American blastomycosis) ), Paragonimiasis, Pasteurellosis, Head Lice Infestation (Head Lice), Body Lice Infestation pubic lice, crab lice, pelvic inflammatory disease (PID), Whooping cough (pertussis), plague, pneumococcal infection, Pneumocystis pneumonia (PCP), pneumonia, Polio, Prevotella infection, primary amebic meningoencephalitis (PAM), progressive multifocal leukoencephalopathy , Psittacosis, Q fever, rabies, respiratory syncytial virus infection, scysnopolydiosis, rhino Viral infection, Rickettsia infection, Rickettsialpox, Rift Valley fever (RVF), Rocky Mountain varicella Spotted fever (RMSF), rotavirus infection, rubella, salmonellosis, SARS (severe acute respiratory syndrome) syndrome), scabies, schistosomiasis, septicemia, shigellosis (bacterial dysentery), shingles (herpes zoster) smallpox, sporotrichosis, staphylococcal food poisoning, staphylococcal infection, strongyloides encephalitis, subacute sclerosing panencephalitis, syphilis, taeniasis, tetanus (biting spasm), tinea barbae (barber's itch) , Tinea of the scalp, Tinea corporis, Tinea cruris, White hands Tinea pedis (tinea of the hands), tinea nigricans, tinea pedis (athlete's foot), tinea unguium (onychomycosis), catfish rash, toxoplasmosis Ocular larva migrans (OLM), Toxocariasis (Visceral larva migrans (VLM)), Coma, Trinochccliasis, Trichinella hinlosis), Trichuriasis (whipworm infection), Tuberculosis, Tularemia, Typhoid, Ureaplasma Ureaplasma urealyticum infection, Valley fever, Bene Zuehl equine encephalitis, Venezuelan hemorrhagic fever, viral pneumonia, West Nile fever, white sand eagle (Chinea japonica) blanca (Tinea blanca), Yersinia pseudotuberculosis (Yer rsinia pseudotuberculosis infection, yersiniosis, yellow fever, Zika virus, and Zygomycosis.
[0243] In some cases, the methods described herein may be used to treat infections whether the infection is active or latent. In some cases, quantification of gene expression is used to detect active infection. In some cases, the present disclosure may provide methods for detecting, predicting, diagnosing, or monitoring a disease. The methods described in the literature include detecting active infection. Expression may be quantified by detection or sequencing of one or more target nucleic acids of interest. In some cases, quantification of gene expression may be useful for detecting, predicting, diagnosing, or monitoring latent infection. In some cases, the methods described herein may provide methods for detecting a latent infection. This includes detecting
[0244] The methods provided herein can be used to detect cancer, and in particular to detect such cancer. A subject having, at risk of having, or suspected of having such a cancer. The present invention can be used to detect cancer in subjects for whom the present invention is being studied. Examples of cancer include, but are not limited to, However, brain tumors, head and neck cancer, laryngeal cancer, oral cancer, breast cancer, bone cancer, blood cancer, leukemia, lymphoma, These include lung cancer, kidney cancer, pancreatic cancer, stomach cancer, colon cancer, rectal cancer, skin cancer, reproductive tract cancer, and prostate cancer. In some cases, the methods provided herein are directed to treating non-hematological cancers, such as cancers of solid organs (e.g., For example, it is particularly useful for detecting lung cancer, breast cancer, pancreatic cancer, etc.
[0245] The methods may also be useful for detecting any other type of disease or condition in a subject. Often it is useful to detect rare genetic variants or to identify specific Useful for detecting nucleic acid sequences that make up only a very small portion of the total nucleic acid population in a It is.
[0246] Detection of pathogen or organ nucleic acid may include determining the presence or absence and / or absence of pathogen or organ nucleic acid. or determining the amount of pathogen or organ nucleic acid by measuring the level of pathogen or organ nucleic acid. The level may be a qualitative or quantitative level. In some cases, the control or reference value may be a cell-free pathogen nucleic acid or a cell-free organelle-derived nucleic acid. A predefined absolute value that indicates the presence or absence of nucleic acid. For example, a cell-free value that exceeds a control value. Detection of levels of pathogen nucleic acid may indicate the presence of a pathogen or infection. The level may indicate the absence of a pathogen or infection. The control value may be the acellular level of a subject without an infection. The control value may be a value obtained by analyzing cellular nucleic acid levels. In some cases, the control value may be It can be a positive control value and is used to measure the specific infection or specific sensitivity of a specific organ. The nucleic acid sequence may be obtained by analyzing cell-free nucleic acid from a subject having the stain.
[0247] In some cases, to determine whether an infection is present, and often accurately To achieve the result, one or more of the following methods may be applied: (i) WO 20150700 As described in the 86 A1 patent, the reads obtained by sequencing The entire genome is compiled into a curated host genome reference database (which includes human, canine, and , feline, primate or any other host origin, e.g., Gen (ii) can be aligned against the human reference sequence (including the hg19 human reference sequence); Bioinformatics was used to further analyze only non-host sequences that contain host-associated sequences. (iii) the data processor for analysis may remove or isolate host sequences; The processor includes a curated set of example reference sequences from, for example, GenBank and Refseq. By aligning non-host sequences against a host microbial reference sequence database, (iv) the presence of one or more pathogens is statistically significant; A statistical analysis framework may be applied to determine whether (v) in some cases, the data processor may add a known concentration of ribosomal DNA to the sample prior to sequencing. The number of reads obtained for the pathogen compared to the number of reads obtained with the control molecule added Based on the number of reads, the amount of pathogen present can be quantified.
[0248] The control value is determined by measuring the level of the subject (e.g., a subject with an infection) at a different time point, e.g., a time point prior to the test time point. Acellular pathogen or organ-specific nuclei obtained from a subject or subject suspected of having an infection. In such cases, comparison of the levels at different times may indicate the presence of infection, It may indicate the presence of infection in a particular organ, improvement of the infection, or worsening of the infection. A steady increase in the amount of cellular pathogen nucleic acid over time can indicate the presence of an infection or a worsening of an infection. For example, at least 5%, 10%, 20%, 25%, 30%, 50%, 7% or more compared to the original value. 5%, 100%, 200%, 300% or 400% pathogen- or organ-specific acellular An increase in nucleic acid may indicate the presence of an infection or a worsening of an infection. Compared to at least 5%, 10%, 20%, 25%, 30%, 50%, 75%, 100% , 200%, 300% or 400% reduction in pathogen or organ-specific cell-free nucleic acid. It may indicate the absence of infection or amelioration of infection. Often, such measurements are performed over a specific period of time. The period may be, for example, daily, every other day, weekly, biweekly, monthly or bimonthly. An increase in pathogen or organ cell-free nucleic acid of at least 50% over a period of time indicates the presence of infection. It is possible.
[0249] The control or reference value may be measured as a concentration or as a number of sequencing reads. The reference value may be pathogen-dependent. For example, Escherichia coli i) The control value is Mycoplasma hominis The database of levels or control values may be different from the control values. Based on samples obtained from one or more subjects for one or more organs and / or one or more time points Such databases may be curated or proprietary. Recommended treatment options may be based on various threshold levels. For example, low levels may indicate sensitivity to Although it may indicate infection, treatment may not be necessary; moderate levels may warrant antibiotic treatment. high levels may require immediate or serious intervention.
[0250] The methods provided herein allow for the generation of sequencing data with high efficiency, high accuracy and / or high sensitivity. Often such methods involve plating or polymerase chain reaction (PCR). Pathogens that are not detectable or detectable by other methods such as PCR or The method generally has a very high sensitivity, e.g., 80%, 85%, 90%, The method may have a sensitivity of greater than 95%, 99% or 99.5%. False positive rates, e.g., 5%, 4%, 3%, 2%, 1%, 0.1%, 0.05%, 0.01 % false positive rate.
[0251] The methods provided herein have high specificity, high sensitivity, high positive predictive value and / or low The methods provided herein may provide a negative predictive value of at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, Specificity (or negative agreement) of 95%, 96%, 97%, 98%, 99% or higher In some cases, nominal The specificity is greater than 70%. The nominal negative predictive value (NPV) is greater than 95%. In this case, the NPV is at least 95%, 95.5%, 96%, 96.5%, 97%, and 9 7.5%, 98%, 98.5%, 99%, 99.5% or higher.
[0252] Sensitivity, positive agreement (PPA) or true positive rate (TPR) is TP / (TP + FN) or TP / (total number of infected subjects), where TP is (FN is the number of true positives, and FN is the number of false negatives). When calculating the denominator of the above formula, the value is Total number of infection results based on a specific independent infection detection method (e.g., blood culture or PCR) can reflect.
[0253] Specificity, negative agreement rate or true negative rate (specificity rate) is TN / (TN+FP) or T It can be related to a formula such as N / (total number of uninfected subjects), where TN is the true negative (FP is a false positive). When calculating the denominator of the formula, the value is calculated based on the independent detection of infection. Reflects the actual total number of "uninfected" individuals as determined by method (e.g., blood culture or PCR) It is possible.
[0254] In some cases, samples were 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, It is identified as infected with 99%, 99.5% or more accuracy. In some cases In some cases, samples are identified as infected with a sensitivity of over 95%. In some cases, samples are identified as infected with greater than 95% specificity. The sample was identified as infected with a sensitivity of over 95% and a specificity of over 95%. In some cases, the accuracy is calculated using a learned algorithm. As used herein, diagnostic accuracy refers to specificity, sensitivity, positive predictive value, negative predictive value and / or false positive. In some cases, the methods described herein include 70%, 75%, 80%, 90%, 100%, 120%, 140%, 160%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 3 0%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 9 Specificity of greater than 4%, 95%, 96%, 97%, 98%, 99%, or 99.5% is the sensitivity or at least 95%, 95.5%, 96%, 96.5%, 97%, 97. 5%, 98%, 98.5%, 99%, 99.5% or higher positive or negative predictive value It has a high degree of accuracy.
[0255] When classifying samples for the diagnosis of infection, typically four variables from a binary classifier are used: There are possible outcomes. If the outcome from the prediction is p and the actual value is also p, then it is true. It is called a positive (TP). However, if the actual value is n, it is a false positive (FP). Conversely, a true negative occurs when both the predicted result and the actual value are n, and the actual If the predicted result is n when the value of p is p, it is a false negative. In tests to detect a disease or disorder, subjects may show a positive test result but in reality False positives in this case may occur if the subject does not have an infection. However, false negatives can occur when a patient has a negative test result for such an infection.
[0256] Positive predictive value (PPV), or precision rate, or disease Post-test probability is the proportion of patients with a positive test result who are correctly diagnosed. It is PPV can be calculated by applying the formula: PPV=TP / (TP+FP). It may reflect the probability that a sexual test result reflects the underlying condition (disease) being tested for. The degree of significance may depend on the variable prevalence of the disease. The negative predictive value (NPV) is The negative predictive value is calculated by the following formula: TN / (TN+FN). The PPV and NPV measurements can be the proportion of patients with a negative test result. Prevalence estimates can be derived.
[0257] In some cases, the results of the sequencing analysis of the methods described herein are provided. It indicates the statistical confidence level that the diagnosis made is correct. In some cases, such statistical confidence The levels are 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 9 Greater than 8%, 99% or 99.5%.
[0258] Monitoring and Treatment The method may include monitoring whether the subject has an infection over time. For example, samples may be collected serially at various time points to determine the presence or absence of infection. In another example, the method may include monitoring the progress of the infection over time. In such cases, samples can be taken serially at various time points during the infection or disease. In some cases, sequential samples were compared to each other to determine whether the infection was improving. Determine whether the condition is improving or worsening.
[0259] The methods provided herein include administering to a subject, e.g., a subject having or suspected of having an infection, The present invention includes a method of treating a subject having an infection. Treatment may reduce, prevent, or eliminate an infection in a subject. In some cases, treatment may reduce, prevent or eliminate infection and / or inflammation.
[0260] Treatment involves administering drugs or other therapies to reduce or eliminate inflammation and / or infection. In some cases, for example, to prevent the onset of infection or inflammation. To achieve this, the subject is prophylactically treated with a drug.
[0261] For any therapy (including drugs) to improve or reduce the symptoms of infection or inflammation Exemplary drugs include, but are not limited to, antibiotics, anti-inflammatory drugs, and the like. Antivirals, ampicillin, sulbactam, penicillin, vancomycin, gentamicin Aminoglycosides, clindamycin, cephalosporins, metronidazole, thimerosine ciprofloxacin, ticarcillin, clavulanic acid, cefoxitin, antiretrovirals (e.g., Highly active antiretroviral therapy (HAART), reverse transcription inhibitors, nucleosides / Nucleotide reverse transcriptase inhibitors (NRTIs), non-nucleoside RT inhibitors and and / or protease inhibitors), antibody-drug conjugates, and immunoglobulins. Brynn included.
[0262] The method can include adjusting a treatment regimen. For example, the subject has a known infection. The present invention provides a method for treating the infection by administering a drug to the patient. The provided methods can be used to follow or monitor the effectiveness of drug treatments. In such cases, the treatment regimen may be adjusted depending on the results of such monitoring. For example, the methods provided herein demonstrate that the infection has not improved as a result of drug treatment. If a patient is experiencing symptoms of a chronic condition, change the type of medication or treatment given to the patient, or consider the use of the previous medication. discontinue use of the drug, continue use of the drug, or increase the dose of drug treatment. or by adding a new drug or other treatment to the subject's treatment regimen. In some cases, the therapeutic regimen may include a specific treatment. If the method indicates that the infection is improving or resolving, the adjustment is This may involve reducing or ceasing treatment.
[0263] The methods described herein may further include RNA sequencing (RNA-Seq). or can be combined with methods including RNA-Seq. Tissue damage or infection can This can result in the release of cell-free nucleic acids from specific organs or tissues. RNA can be released by apoptotic cells. RNA-Seq of cell-free RNA can be used to identify specific species in the body. These may indicate the health or condition of various tissues.
[0264] Methods involving RNA sequencing allow for detection of specific organs or tissues that are infected, and RNA-Seq can be used to detect or monitor the health of organs. can be used independently or by the methods described herein to assess the health of This can increase confidence that the infection detected is in a specific organ. may occur simultaneously with, after, or before the infection detection method. It is possible.
[0265] The present invention provides a method for detecting pathogens by RNA sequencing of cell-free RNA in body fluids. There are many potential scenarios that can be combined with the site detection methods provided herein. The method may be used to detect circulating cell-free nucleic acid from a pathogen. , an RNA-Seq study to detect increases in organ-specific cell-free RNA in subject blood The combination of test results may indicate that a pathogen is infecting the organ. It is possible to determine which organ tissue is infected. Ugh.
[0266] An RNA-Seq study (or a series of RNA-Seq studies) is sometimes referred to as The method may be performed after the procedure has demonstrated a positive test result (eg, detection of a pathogen infection). RNA-Seq testing is particularly useful for confirming or localizing infection. For example, the method can be used to determine whether a patient has a circulating cell-free nucleic acid in a subject. The presence of a pathogen may be detected, but the site of infection may be unknown. may further be used to detect increased levels of circulating cell-free RNA derived from organ tissue (e.g., by detecting increased levels of circulating cell-free RNA derived from organ tissue). Sequence cell-free RNA from the subject to confirm that the infection is present in the organ Then, the infection may worsen or improve in a particular organ or tissue. To determine whether the cancer has spread to different organs or tissues Similarly, the RNA sequencing test can be repeated over time. The test can be repeated over time.
[0267] In some cases, the pathogen detection methods described herein include RNA-Seq testing. For example, an increase in plasma levels of cell-free RNA associated with an organ may be In such cases, the method further comprises determining whether a circulatory system associated with an organ infection is indicative of a disorder such as an infection of the organ. This may include detecting the level of circular cell-free nucleic acid.
[0268] The methods described herein can be used, for example, to monitor infection or treatment over time. The methods described herein can be repeated for 1, 2, 3, 4, 5, 6, 7, 8 , every 9 or 10 days, every 1, 2, 3, 4, 5 or 6 weeks, or every 1, 2, 3, 4, It may be repeated every 5, 6, 7, 8 or 9 months.
[0269] In some cases, the methods described herein provide a negative test result (e.g., If no pathogen is detected, the method is performed over time to monitor pathogen nucleic acid in the subject. In some cases, negative pathogen tests can be performed sequentially. RNA-Seq assays were run serially over time after positive or negative RNA-Seq results. Repeat.
[0270] In some cases, the methods described herein may be used to detect a positive test result (e.g., pathogen If the subject exhibits a pulmonary function test (detection), a therapeutic regimen can be administered to the subject. , drug administration, antibiotic administration or antiviral administration, but are not limited to isn't it.
[0271] In some cases, if the methods described herein show a positive test result, the infection may be The method or test may be repeated serially over time to monitor progress. For example, treatment regimens can be adjusted depending on whether the infection is progressing upward or downward. In other cases, no treatment regimen is administered initially. For example, additional medical "Watchful waiting" or "wait and see" to see if the infection resolves without clinical intervention. In some cases, the infection may be monitored by a "watch and follow-up" method as described herein. If the method used gives a positive test result, the drug can be administered and the effectiveness of the drug can be evaluated. The course of the infection should be monitored to see how well the drug works or when to stop it. It is possible to monitor and, in some cases, modify therapy as necessary.
[0272] Computer Control System The present disclosure also relates to a computer control system programmed to carry out the methods of the present disclosure. FIG. 7 illustrates a method for implementing the disclosed method. 7 shows a computer system 701.
[0273] Computer system 701 includes a central processing unit (CPU; referred to herein as a “processor”). 705, which may also be referred to as a single core It may be a multi-core processor, or multiple processors for parallel processing. The computer system 701 also includes a memory or memory locations 710 (e.g., random access memory, read-only memory, flash memory), electronic storage devices 715 (e.g., a hard disk), a communications interface for communicating with one or more other systems 720 (e.g., a network adapter), and peripheral devices 725, e.g., a cache memory, other memory, data storage devices and / or electronic display adapters. The controller 710, the storage device 715, the interface 720 and the peripheral device 725 are connected to the motherboard. The memory device 71 is connected to the CPU 705 via a communication bus (real line) such as a 5 may be a data storage device (or data repository) for storing data. The computer system 701 communicates with a computer network using a communication interface 720. The network 730 may be operatively connected to an internetwork ("network") 730. Internet, Internet and / or Extranet, or Internet The network can be an intranet and / or an extranet connected to the Network 730 may, in some cases, be a telecommunications and / or data network. Network 730 includes one or more computer servers, which are distributed computing This may enable computing, for example, cloud computing. 0 may, in some cases, use a computer system 701 to establish a peer-to-peer network. The computer system 701 can execute the command may be able to act as a client or a server.
[0274] The CPU 705 executes a series of machine-readable instructions, which may be represented in a program or software. The instructions may be stored in a memory location, such as memory 710. The instructions may be 705, which then programs the CPU 705 to perform the methods of the present disclosure. Examples of operations that can be performed by the CPU 705 include fetch, decode, and This may include read, execute and writeback.
[0275] The CPU 705 may be part of a circuit such as an integrated circuit. In some cases, the circuit may be implemented as an application specific integrated circuit (AS IC).
[0276] The storage device 715 stores files, such as drivers, libraries, and saved programs. The storage device 715 may store user data, such as user preferences and user programs. The computer system 701 may, in some cases, store a It is possible for the system 701 to include one or more additional data storage devices external to the system 701. The public data storage device may be, for example, a computer connected to the internet via an intranet or the Internet. The data system 701 is located on a remote server connected to the data system 701 .
[0277] The computer system 701 communicates with one or more remote computers via a network 730. For example, the computer system 701 can communicate with a user (e.g., The remote computer system can communicate with the remote computer system of the Examples of data systems include personal computers (e.g., portable PCs), slates, or tablet PC (e.g., Apple (registered trademark) iPad, Samsung (registered trademark) Galaxy Tab), phones, smartphones (e.g., Apple )iPhone, Android-compatible devices, Blackberry (registered trademark) or a personal digital assistant. The user can A computer system 701 is accessible.
[0278] The methods described herein may be implemented using electronic storage locations, such as a computer system 701 For example, the memory 710 or electronic storage device 715 may store Machine executable code or machine readable code can be executed by a processor. The code may be provided in the form of software. During use, the code may be executed by the processor 705. In some cases, the code may be retrieved from storage 715 and executed by processor 716. 10. In some cases, Alternatively, electronic storage device 715 may be omitted, and machine executable instructions may be stored in memory 710. do.
[0279] The code may be used in conjunction with a machine having a processor adapted to execute the code. It can be precompiled and configured to run on the fly, or compiled on the fly. The code may be downloaded in precompiled or compiled form to make the code executable. The invention may be provided in a programming language that may be selected to implement the invention.
[0280] Aspects of the systems and methods provided herein, such as computer system 701, include: Various aspects of the technology may be embodied in programming. or processor) in the form of executable code and / or stored on a type of machine-readable medium A machine can be considered a "product" or "item of manufacture" in the form of associated data embedded in it. The executable code may be stored in electronic storage, such as memory (e.g., read-only memory, random access memory, etc.). It can be stored in a memory (e.g., a microprocessor, a USB flash memory, etc.) or on a hard disk. "Storage" type media includes tangible memory such as computers, processors, or or associated modules, such as various semiconductor memories, tape drives, disk drives, etc. It can contain anything and everything, and they are software programs. may provide non-transitory storage at any time for the processing of any part of the Software or Some, at times, may be connected to the Internet or various other telecommunications networks. Such communication may occur from one computer or processor to another. to a computer or processor, e.g., from a management server or host computer Allows software to be loaded onto application server computer platforms. Thus, other types of media that can carry software elements include wired and The physical interface between local devices is established through optical landline networks and various air links. This includes light waves, radio waves and electromagnetic waves, such as those used across a wide range of interfaces. The physical elements that carry the waves, such as wired or wireless links, optical links, etc., also carry the software. As used herein, a computer or machine "readable medium" may be considered a medium that retains information stored in the computer or machine readable medium. Such terms do not include those that require a processor for execution, unless they are limited to non-transitory tangible "storage" media. "Instruction" means any medium involved in giving instructions to a server.
[0281] Thus, the machine-readable medium, e.g., the computer executable code, may be any type of storage medium, such as a tangible storage medium, a portable It can take many forms, including but not limited to, a wave or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, any Any of the storage devices in a computer or the like, such as the database shown in the drawings Volatile storage media includes dynamic memory, e.g. For example, the main memory of such a computer platform. These include coaxial cable, copper wire and optical fiber (which make up the buses in computer systems). Carrier-wave transmission media include electric or electromagnetic signals, or acoustic or or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications Thus, typical forms of computer readable media include, for example, , floppy disk, flexible disk, hard disk, magnetic tape, any other Magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punch Card paper tape, any other physical storage medium with a pattern of holes, RAM, ROM , PROM and EPROM, FLASH-EPROM, any other memory chip or Cartridges, carrier waves carrying data or instructions, cables carrying such carrier waves or links, or computer-generated programming code and / or data Many of these forms of computer readable media may include any other medium that can be read by a computer. , which is involved in conveying one or more sequences of one or more instructions to a processor for execution. Ugh.
[0282] The computer system 701 may include an electronic display 735. or an electronic display 735, which displays the diagnosis of the subject. or a user interface that provides a report output that may include a therapeutic intervention for a subject. (UI) 740. Examples of UI include, but are not limited to, graphical user interfaces (GUIs). The system includes a GUI and a web-based user interface. The report may be provided to subjects, health care professionals, researchers, or other individuals. .
[0283] The methods and systems of the present disclosure may be implemented by one or more algorithms. The program may be implemented by software upon execution by the central processing unit 705. The algorithms may be used, for example, to enrich, sequence, and / or detect pathogens or other target nucleic acids. This can facilitate:
[0284] Patient or subject information, such as patient demographics, patient medical history or medical scans The computer system may be configured to include any of the methods described herein. To analyze results from the methods used or to report results to patients or physicians, or or may be used to formulate a treatment plan.
[0285] Reagents and Kits Also provided are reagents and kits thereof for carrying out one or more of the methods described herein. The reagents and kits thereof can be of many different types. Reagents of interest include those obtained from a subject. Identifying, detecting and / or quantifying one or more pathogens or other target nucleic acids in a sample. The kit includes reagents specifically designed for use in the method of the present invention. Nucleic acid extraction and / or detection using methods that are currently used, such as PCR and sequencing. The kit may further include a software package for data analysis. This may include a reference profile for comparison with the test profile. The kit may include reagents, e.g. For example, it may include buffer and water.
[0286] Such kits also contain information, such as scientific literature references, package inserts, clinical trials, and the like. The composition may include, for example, test results and / or summaries thereof, which may be used to Demonstrate or determine activity and / or benefits, and / or dosage, administration, side effects These information may include information about prescription drugs, drug interactions, or other information that may be helpful to your health care provider. The kit may also include instructions for accessing the database. The results of various studies, including studies using experimental animals including in vivo models and human clinical trials, The kits described herein may be used by physicians, May be offered, sold and / or promoted to healthcare providers, including nurses, pharmacists, prescribing personnel, etc. Kits may also be sold directly to consumers in some embodiments.
[0287] The present disclosure also provides a kit for producing a sequencing library. The kit includes: At least one synthetic nucleic acid as described herein, and a sequencing library reaction. In some cases, the kit may include one or more sequencing adaptors. and one or more carrier nucleic acids. The carrier nucleic acid in the kit comprises: i) one or more carrier nucleic acids that are resistant to end repair, ii) one or more carrier nucleic acids that are resistant to ligation. iii) one or more carrier nucleic acids that are resistant to amplification; and iv) one or more carriers that include an immobilization tag. a nucleic acid; v) one or more carrier nucleic acids having a size that allows for size-based depletion; and / or or vi) any combination thereof. For example, the kit may include one or more sequencing It may include an adaptor, and one or more carrier nucleic acids that are resistant to end repair.
[0288] The amount of sequencing library adaptor and the amount of one or more carrier nucleic acids in the kit are the same. In some cases, the ratio may be one or more of the amount of sequencing library adapters. The ratio of the amount of carrier nucleic acid to the amount of the nucleic acid was 1:10, 1:5, 1:1, 5:1, 10:1, 20:1, 5 0:1, 100:1, 500:1, or 1000:1 or less. The ratio of the amount of library adaptor to the amount of one or more carrier nucleic acids can be 1:1 or less.
[0289] Carrier Nucleic Acid (CNA) The present disclosure relates to carrier nucleic acids (CNAs), particularly those used in sequencing assays. A surreptitious The present disclosure also provides a method for the preparation of a nucleic acid sequence that avoids one or more of the steps of a sequencing assay. The present invention provides methods of using CNAs that can act covertly, but They are generally capable of increasing the total amount of nucleic acid in a sample, thereby Carrier nucleic acids are generally used to remove sequences from a sample. Increasing the amount of nucleic acid to improve yield and / or efficiency in producing a fixed library This may ultimately improve the accuracy and / or sensitivity of the sequencing assay. Addition of carrier nucleic acid containing modified CNAs is useful when samples contain small amounts (e.g., less than 1 ng) of target nuclei. This can be particularly useful when the nucleic acid is present in a small amount, because the library One or more steps of manufacturing (e.g., nucleic acid extraction, nucleic acid purification, nucleic acid end repair, adapter ligation, etc.) ) or the efficiency and / or yield of subsequent steps (e.g., amplification) in a sequencing assay. Nucleic acids based on DNA and / or RNA are known to be in any form and / or with or without one or more chemical modifications , can be added to a nucleic acid sample of interest as a CNA. Typically, the CNA is, for example, either by inhibition or by taking up a disproportionate portion of sequencing throughput. In some cases, DNA samples and / or RNA samples may be used for nucleic acid sequencing. In some cases, DNA samples and / or Add RNA CNA to the RNA sample. [Table 6]
[0290] The CNAs provided herein may be used in one or more steps of the sequencing library production, e.g., It may be designed or modified to avoid repair, fragmentation, amplification, ligation and sequencing. NAs can be added at one or more steps in the preparation of a sequencing library. For example, see FIG. As shown, the CNA may be used during or immediately after sample collection 802, sample preparation, or , for example during or after isolation of plasma 803, before nucleic acid isolation 804 or extraction 805; During or after, before, during or after nucleic acid purification, before or during end repair of nucleic acid 806 before or after ligation 807 or other procedures for attaching adaptors to nucleic acids. or after and / or before or during amplification 808. To identify CNAs, one or more of the following methods may be used: enzymatic digestion, affinity-based depletion and / or size-selection. Depletion based on the above method can be eliminated from the steps in the sequencing assay. The CNA provided should be included in the sequencing assay so that it is not included in the sequencing library. In some cases, CNAs can be physically removed from the sequencing process. They may be physically removed from the rally itself.
[0291] CNA Resistant to Binding The CNAs provided herein may comprise one or more sequencing adaptors and / or target nucleic acids. In some cases, the molecule may be resistant to binding or linking to other molecules such as In this case, the CNA is preferably ligated to the target nucleic acid in preference to the adaptor over the CNA. By avoiding ligation or binding to the adaptor or target nucleic acid, NAs may also be avoided from being sequenced.
[0292] In some cases, particularly those using ligation to attach adapters to nucleic acids in a sample, In some cases, the CNA may be designed to resist participating in the ligation reaction. Ligation involves joining two nucleic acids together via a phosphodiester bond. In some cases, CNAs may have secondary structures (e.g., single-stranded structures, hairpin structures) that resist ligation. The secondary structure can be RNA, DNA, ssDNA, dsDNA, It may include DNA-RNA hybrids and / or other features. The NA may contain blocking groups or other structures designed to prevent ligation.
[0293] The CNAs provided herein are designed to resist or reduce binding or ligation. A CNA may contain one or more single-stranded regions and / or double-stranded secondary structures. The single-stranded region may contain the nucleotide sequence of the CNA or may be entirely single-stranded. Although it may be located at any position, in some preferred cases the CNA is located near its terminus or It contains a single-stranded region at one or both ends. For example, a CNA may Within 50 nucleotides from the termini, e.g., 50 nt, 45 nt, or 50 nt from one or both termini t, 40nt, 35nt, 30nt, 25nt, 20nt, 15nt, 10nt or 5 In some preferred cases, the CNA may contain a single-stranded region within nt of its terminus. may contain a single-stranded region at one or both ends (e.g., the 5' end, the 3' end). In some cases, the CNA can be entirely double-stranded, or it can have regions that are double-stranded. The secondary structure (especially the hairpin loop) can be simply In some cases, the CNA may be a Y-shaped duplex. It is possible for the Y-shaped portion of the CNA to contain a nucleic acid, so that it can be linked to or attached to another nucleic acid. cannot be combined.
[0294] The hairpin structures that may be present in the CNAs provided by the present invention generally comprise loops and hybridization. For example, a hairpin may be a double-stranded hybridization region. Two complementary regions that form a hybridization region and a region that links the two complementary regions The complementary region may include at least 5, 10, 15, 20, 30, 40 The loop region may comprise at least 3, 4, 5, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 210, 220, 230, 240, In general, hairpin structures often contain 10, 30, 40, 50 nucleotides. Hairpins are single-stranded nucleic acids without any nucleic acid residues, and therefore can be relatively easy to manufacture. may contain DNA.
[0295] The CNAs provided herein have a cyclic structure that can resist or reduce binding or ligation. The circular structure may contain circular DNA, circular RNA or a circular DNA-RNA hybrid. In some cases, the circular structure is circular DNA. The circular structure may be double stranded or The circular structure can be of a particular length, e.g., at least 5 nt, 10 nt, 2 nt, 0nt, 30nt, 32nt, 40nt, 50nt, 60nt, 70nt, 80nt, 9 0nt, 100nt, 120nt, 140nt, 160nt, 180nt, 200nt, It can be 250nt, 300nt, 400nt, 500nt or 1000nt. In some cases, the circular structure comprises about 30 to about 100 nucleotides. In the present invention, the circular structure is in the range of about 10 nucleotides to about 10,000 nucleotides, for example about 1 The size of the circular structure may range from about 1,000 nucleotides to about 1,000 nucleotides. If double stranded, the circular structure must be at least 10 bp, 20 bp, 30 bp, 40 bp, 5 0bp, 60bp, 70bp, 80bp, 90bp, 100bp, 200bp, 250b The size of the fragment may be 300 bp, 400 bp, 500 bp or 1000 bp. In some cases, the double-stranded circular structure contains about 30 bp to 100 bp. The stranded circular structure is in the range of about 10 base pairs to about 10,000 base pairs, for example, about 100 base pairs. In some cases, the circular structure may have a size ranging from about 1,000 base pairs to about 1,000 base pairs. It is possible for CNAs to be resistant to digestion by certain enzymes, e.g. endonucleases. For example, CNAs can contain double-stranded circular structures, endonucleases, e.g., endonucleases that digest double-stranded linear DNA but not double-stranded circular DNA In some cases, CNAs may be largely or entirely resistant to digestion by nucleases. is circular, e.g., circular double-stranded DNA, circular single-stranded DNA. In some cases CNAs do not bind to endonucleases, e.g., secondary structures of CNAs, and / or contains secondary structures that resist digestion by endonucleases that do not recognize it. For example, CNAs are synthesized by endonucleases that recognize single-stranded DNA but not double-stranded DNA. In another example, a CNA may comprise double-stranded DNA that is resistant to digestion by Single stranded DNA that is resistant to digestion by endonucleases that recognize . It may contain DNA.
[0296] In some cases, the CNA is double stranded with one or more nicks. The gaps can be discontinuous in a double-stranded nucleic acid molecule, in which case one of the strands There is no phosphodiester bond between adjacent nucleotides. Nicks are formed by enzymes (e.g., In some cases, nicks can be generated by enzymes (e.g. In some cases, the nicks can be ligated by exonuclease digestion. and / or protected from interlocking.
[0297] A CNA may contain one or more modifications (e.g., modified nucleotides) that resist ligation. In some cases, the modification is a blocking group that prevents the CNA from ligating to the nucleic acid. For example, the CNA can have a blocking group at the 3' end, the 5' end, or both ends. The blocking group may comprise an inverted deoxy sugar. It can be an inverted deoxy sugar, an inverted dideoxy sugar, or another inverted deoxy sugar. The sugar can be a 3' inverted deoxy sugar or a 5' inverted dideoxy sugar. The bases are 3' inverted thymidine (dT), 3' inverted adenosine (dA), and 3' inverted guanosine (d G), 3' inverted cytidine (dC), 3' inverted deoxyuracil (dU), 5' inverted dideo 5'-inverted dideoxyadenosine (ddA), 5'-inverted dideoxy Deoxyguanosine (ddG), 5' inverted dideoxycytidine (ddC), 5' (inverted) dideoxy It may be xyluracil (ddU) or any analogue thereof. NA contains a 3' inverted thymidine. In some cases, CNA contains a 5' inverted dideoxythymidine. In some cases, the CNA contains a 3' inverted thymidine and / or a 5' inverted dideoxy acid. In some cases, the blocking group comprises dideoxycytidine. In some cases, modifications include uracil (U) bases, 2'OMe modified RNA, C3-18 spirochetes, (e.g., structures having 3 to 18 consecutive carbon atoms), biotin, di-deoxyribonucleic acid Crylene triphosphate, ethylene glycol, amines and / or phosphates (Phosphate).
[0298] Carrier nucleic acid that resists amplification CNAs inhibit nucleic acid amplification, preventing their amplification in sequencing reactions. In some cases, the modification may include one or more nucleic acid modifications that prevent, for example, the synthesis of a polymerase. By blocking or inhibiting (e.g., reducing) the function of the nucleic acid polymerase, In some cases, the modification may involve one or more abasic sites. An abasic site can refer to a position in a nucleic acid that does not have a base. The site may be at the 1' end with no abasic site. The amine structure may have a base analog or a phosphate backbone analog. The positions are linked by amide bonds to N-(2-aminoethyl)-glycine, tetrahydrofuran, It has furan or 1',2'-dideoxyribose (dSpacer). In some cases, the modification is an abasic site and a modified sugar residue, e.g., a sugar residue containing 3 carbon atoms. groups, e.g., partial ribose structures (e.g., only the 3', 4', and 5' terminal carbon atoms are retained) It is possible to include a set of nodes, which allows connectivity along the backbone It can be maintained.
[0299] Abasic sites can prevent polymerases from amplifying CNAs. The abasic site in NA induces a polymerase (e.g., Taq polymerase) to Each individual can inhibit by an order of magnitude.
[0300] The CNAs provided herein contain multiple abasic sites, e.g., multiple internal abasic sites and one The CNA may be prevented from participating in one or more library generation reactions. For example, CNAs may contain one or more internal abasic sites, a 3' inverted dT, and / or or 5' inverted ddT in any combination.
[0301] In some cases, CNAs may contain other modifications that inhibit nucleic acid amplification. Modifications that inhibit nucleic acid amplification include uracil (U) bases, 2'OMe modified RNA, C3 -18 spacer (e.g., having 3 to 18 consecutive carbon atoms, such as a C3 spacer Structure), ethylene glycol polymer spacer (e.g., spacer 18 (hexa-ethylene ethylene glycol spacer), biotin, di-deoxynucleotide triphosphate, ethylene glycol, amines and / or phosphates).
[0302] qualification CNAs must have at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more CNAs may contain multiple modifications (e.g., abasic sites) that inhibit nucleic acid amplification. In the case of a cluster of modifications, the modifications may be clustered (e.g., modifications that are adjacent to each other and are contiguous). In some cases, one or more modifications are present at the 5' end of the CNA. In some cases, one or more modifications are present at the 3' end of the CNA. The one or more modifications are present at both the 3' and 5' ends of the CNA. The one or more modifications are present at an internal position of the CNA. For example, the CNA may have one or more internal d May contain spacer (idsp).
[0303] Modifications described herein include 2-aminopurine, 2,6-diaminopurine, 5 -Bromo dU, deoxyuridine, inverted dT, inverted dideoxy-T, dideoxy-C, 5 -Methyl dC, deoxynosine, universal bases such as 5-nitroindole, 2'-O-methyl Chiral RNA bases, iso-dC, iso-dG, ribonucleotides, morpholinos, proteins Nucleotide Analogues, Sugar Nucleotide Analogues, Locked Nucleotides analogs, threose nucleotide analogs, chain terminating nucleotide analogs, thiouridine, pso Idouridine, dihydrouridine, queousine, wyosine nucleotides, abasic sites , functional groups such as alkyne functional groups, azide functional groups such as azide (NHS esters, non-natural natural linkages, e.g. phosphorothioate linkages, spacers, e.g. 2'-dideoxyribose (dSpacer), hexanediol, a photocleavable spacer, with various numbers of carbon atoms Spacers of various lengths can be used, such as C3 spacer phosphoramidites, C9 spacer phosphoramidites, etc. For example, triethylene glycol spacer, CI8 (18 atom hexaethylene glycol Such a spacer may be included at the 5' end of the CNA or adapter. The CNA may be incorporated at the end, at the 3' end, or internally. For example, either the 5' phosphate or the 3' phosphate (e.g., on the complementary strand) can be modified by phosphorylation to contain both.
[0304] Enzyme recognition site CNAs contain features that allow them to be removed from a sequencing library. Such features may include enzyme recognition sites. For example, CNAs are designed to allow the synthetic nucleic acid to be bound to an enzyme. In some cases, CNAs may contain one or more enzyme recognition sites so that they can be more easily degraded. may contain one or more enzyme recognition sites that are not present in the target nucleic acid and the adapter. Thus, the carrier nucleic acid can bind the recognition site without causing enzymatic degradation of the target nucleic acid or the adapter. It can be removed by targeted enzymes.
[0305] In some cases, the CNA may contain a nuclease recognition site. For example, The recognition site can be an endonuclease recognition site. Endonucleases can be type I, type II, (including Type IIS, IIG), Type III or Type IV endonucleases. In some cases, the endonuclease recognition site is a restriction endonuclease recognition site. For example, the endonuclease recognition sites are AatII, Acc65I, AccI, and AclI. , AatII, Acc65I, AccI, AclI, AfeI, AflII, AgeI, ApaI, ApaLI, ApoI, AscI, AseI, AsiSI, AvrII, Ba mHI, BclI, BglII, Bme1580I, BmtI, BsaHI, BsiEI , BsiWI, BspEI, BspHI, BsrGI, BssHII, BstBI, Bs tZ17I, BtgI, ClaI, DraI, EaeI, EagI, EcoRI, Eco RV, FseI, FspI, HaeII, HincII, HindIII, HpaI, K asI, KpnI, MfeI, MluI, MscI, MspA1I, MfeI, MluI , MscI, MspA1I, NaeI, NarI, NcoI, NdeI, NgoMIV, NheI, NotI, NruI, NsiI, NspI, PacI, PciI, PmeI, PmlI, PsiI, PspOMI, PstI, PvuI, PvuII, SacI, Sa cII, SalI, SbfI, ScaI, SfcI, SfoI, SgrAI, SmaI, SmlI, SnaBI, SpeI, SphI, SspI, StuI, SwaI, XbaI The enzyme recognition site may be a recognition site for XhoI, XmaI or XhoI. The enzyme recognition site is as listed above. The enzyme recognition site may be a site for a non-specific DNase, e.g., an exodeoxyribonuclease. The positions are uracil DNA glycosylase (UDG), DNA glycosylase-lyase (en nuclease VIII) or mixtures thereof (e.g., uracil-specific cleavage reagent (U For example, a CNA may contain one or more uracils (e.g., an internal uracil). The enzyme recognition site may include a site for an RNA-guided DNase, e.g., CRISPR. It may be the site of a nuclease associated protein, such as Cas9. The RNase recognition site is an RNase, e.g., an endoribonuclease, e.g., RNase A, RNase H, RNase III, RNase L, RNase P, RNase PhyM, RNase T1, RNase T2, RNase U2, RNase V, or exonuclease cleases, such as polynucleotide phosphorylase, RNase PH, RNase R, RNase D, RNase T, oligoribonuclease, exoribonuclease I, can be a recognition site for exonuclease II. In some specific examples, the CNA Restriction enzyme recognition sites may be included, and the methods provided herein include the step of extracting such sites. This can include digesting the CNA with a restriction enzyme that recognizes the CNA. In some cases, the CNA can be digested with an enzyme that recognizes the CNA. Enzymes (e.g., enzymes that bind to and / or degrade CNAs), ribozymes, Secondary or tertiary structures that can be recognized by tamers and DNA-based catalysts or binding polymers In some cases, the CNA comprises one or more specific bonds that can be recognized by an enzyme. The nucleic acid sequence includes
[0306] In some cases, CNAs are DNA fragments that can be degraded by DNases or RNases. In some cases, CNAs may include DNA-RNA-DNA hybrids. Such molecules may be double-stranded. The terminal regions of CNAs are deoxyribonucleic acid. The internal region may comprise ribonucleotides. In some cases, the internal region may comprise ribonucleotides. In this case, the DNA-RNA hybrid can be ligated to a target nucleic acid or an adaptor. The DNA-RNA hybrids are then depolymerized into RNA prior to sequencing (e.g., prior to the amplification step). In some specific cases, DNA-RNA hybrids can be digested by ( e.g., by RNase), while the target nucleic acid (e.g., DNA, e.g., cell-free DNA) is digested. NA) is not digested by RNases.
[0307] If the DNA portion of the CNA is long enough to resist amplification, it can be amplified prior to sequencing. An RNase digestion step to remove DNA-RNA hybrids may not be necessary in Alternatively, the DNA-RNA hybrid molecules are degraded by enzymatic digestion before amplification. In this case, the DNA-RNA hybrid must have a size or length that resists amplification. It may not be necessary.
[0308] CNA for size-based depletion CNAs are removed from sequencing libraries by size-based depletion. In some cases, the CNA may have a size that can be separated from the target nucleic acid. long or has a length longer than the average length of the target nucleic acid. For example, a CNA is at least 1.5, 2, 3, 4, 5, 10, 20 or 5 times the average length of the target nucleic acid The CNA may have a length of at least 150 bp, 200 bp, 300 bp, 400 bp, 00bp, 500bp, 600bp, 800bp, 1kb, 2kb, 5kb or 10k For example, the CNA may have a length of at least 500 bp. In this case, the CNA may have a size within the range of about 150 bp to about 1000 bp. In some cases, the CNA may have a size of up to 2 kb. The length is shorter than the length of the target nucleic acid or the average length of the target nucleic acid. For example, a CNA is or at most 99%, 95%, 90%, 80%, 60%, 50%, or less of the average length of the target nucleic acid. In some cases, the CNA may have a length of 40%, 20%, or 10% of the length of the target nucleic acid. The size may be at most 50% of the average size of the target nucleic acid. , the CNA has a length substantially the same as the target nucleic acid or the average length of the target nucleic acids.
[0309] CNAs having a size or length that allows for size-based depletion are described in this disclosure. To prevent any modification that has been performed, such as ligation, amplification, end repair, or a combination thereof. In some cases, one or both ends of the CNA may contain one or more of the modifications. In some cases, the modification may be an internal modification, such as an internal abasic site, or a terminal There may be a combination of terminal and internal modifications.
[0310] In some specific instances, the CNA is a longer CNA that allows for size-based depletion. They may have length and modifications such as inverted bases that prevent ligation (eg, terminal modifications, etc.). Other combinations of structures that prevent or impede ligation are also possible (e.g., hairpin loops, (combination of hairpin loop and terminal modification). In some cases, the CNA may contain one or more hairpin In some specific cases, the CNA may contain 500 bp, and may have a size or length of 3' inverted dT, 5' inverted dd T, C3 spacer, or spacer 18, or hairpin structure (present at one end In some specific cases, the CNA may have a size of more than 600 bp or length, 3' inverted dT, 5' inverted ddT (at one end) ) and one or more internal abasic sites.
[0311] Immobilization Tag CNAs can contain one or more immobilization tags. The immobilization tags can be used for affinity-based depletion. It is used to remove CNAs from solutions (e.g., sequencing library solutions) For example, the immobilized tag can be attached to a solid support, such as a bead or a plate. When the solution is contacted with the solid support, the CNA may be removed from the solution. The CNA can be shorter than the target nucleic acid. Alternatively, the CNA molecule can be, for example, To minimize carryover of CNAs into the sequencing reaction, It can be longer than the nucleic acid.
[0312] Immobilization tags include biotin, digoxigenin, Ni-nitrilotriacetic acid, and desthiobiotin. tin, histidine, polyhistidine, myc, hemagglutinin (HA), FLAG, fluorescence Tag, Tandem Affinity Purification (TAP) Tag, Glutathione S-Transferase (GST), polynucleotides, aptamers, polypeptides (e.g., antigens or antibodies) or derivatives thereof. For example, the CNA may include biotin, e.g., internal or terminal In some cases, the immobilization tag may comprise a magnetically susceptible material, e.g. In some specific examples, the biotinylated CNA Prior to the amplification step, the magnetic beads were used to separate CNAs from samples or sequencing libraries. Size-based depletion (e.g., with avidin-magnetic beads) may be possible. In some cases, the CNA may be a secondary or tertiary nucleotide that can be bound to a solid support or can be bound to an immobilized tag. The secondary structure is included.
[0313] In some cases, the target nucleic acid and / or the sequencing library nucleic acid may be one or more immobilized nucleic acids. In these cases, the CNA does not contain an immobilization tag or is different from the target nucleic acid. Thus, CNAs can be used in affinity chromatography using different immobilization tags. are separated from the target nucleic acid and / or the sequencing library nucleic acid by affinity-based depletion. For example, the target nucleic acid and / or the sequencing library nucleic acid can be immobilized on a solid support. In some cases, the CNA can be washed away. The CNA is directly or indirectly linked to an immobilization tag. In some cases, the CNA is linked to an immobilization tag. The user is disconnected from the group.
[0314] A CNA may include a combination of the features and structures disclosed herein. In this case, the CNA may have one or more modifications that inhibit nucleic acid amplification and one or more modifications that resist ligation. For example, CNAs contain one or more abasic sites (e.g., internal dspacers) and and inverted deoxy bases (e.g., 3' inverted thymidine). CNAs containing modifications may further include , enzyme recognition sites and / or immobilization tags. In some cases, a CNA may comprise one or more DNA-RNA hybrids having immobilization tags, e.g., biotinylated DNA-RNA- CNAs contain DNA hybrid molecules that bind to specific enzymes or proteins, non-amino acids Any catalytic or affinity entity based on a DNA molecule, such as a ribozyme, a DNA-based catalytic polymer, Secondary and / or tertiary structures of nucleic acids with high affinity for molecules and molecularly imprinted polymers It may also have the following structure:
[0315] Ratio of carrier nucleic acid to nucleic acid in the sample For example, to prepare a sequencing library from the nucleic acids in a sample, A specific amount of CNA can be added to a sample containing the CNA. In some cases, the total amount of CNA in the sample can be The ratio of the amount of nucleic acid to the amount of CNA added to the sample should be at least 1:100, 1:50, 1:10, 1:1, 10:1, 50:1, 100:1, 500:1, 1000:1, 20 In some cases, the ratio of the target nucleic acid in the sample to the total amount of the target nucleic acid is 0:1, or 5000:1. The ratio of the amount of CNA to the amount of CNA added to the sample should be at least 1:100, 1:50, or 1:1 0, 1:1, 10:1, 50:1, 100:1, 500:1, 1000:1, 2000: In some cases, the amount of total nucleic acid in the sample and the amount of sample The ratio of the amount of CNA added to the solution should be at most 10:1, 1:1, 1:10, 1:50, or 1:1 :100, 1:500, 1:1000, 1:2000 or 1:5000. In the case of , the ratio of the amount of target nucleic acid in the sample to the amount of CNA added to the sample is At most 10:1, 1:1, 1:10, 1:50, 1:100, 1:500, 1:100 0, 1:2000 or 1:5000. In some cases, the total nuclei in the sample The ratio of the amount of acid to the amount of CNA added to the sample ranges from about 1:1 to about 1:100. In some cases, the amount of target nucleic acid in the sample and the amount of CNA added to the sample are The ratio of the amount of is in the range of about 1:1 to about 1:100. In some cases, the ratio is a molar ratio. be.
[0316] How CNAs are used in producing sequencing libraries The disclosure herein includes a method for producing a sequencing library. In order to improve the efficiency and / or yield of the production of galleries, A sequencing library can include adding nucleic acid molecules to be subjected to sequencing. The method can refer to a population of target nucleic acids and / or adapters (e.g., sequencing The method may include obtaining a sample containing one or more CNAs, the ... The method may further include one or more steps for producing a sequencing library. The method may include sequencing one or more nucleic acids in the constant library. For example, CNAs cannot be physically removed from the library. or it is involved in one or more steps in the production of a sequencing library. It is possible to design it so that it does not affect the
[0317] The method comprises adding a CNA to a sample containing a target nucleic acid and / or an adaptor. The amount of CNA added to the sample may be at least 0.1 ng, 0.5 ng, 1 ng, 5ng, 10ng, 20ng, 30ng, 40ng, 50ng, 60ng, 70n g, 80ng, 90ng, 100ng, 150ng, 200ng, 300ng, 400n g or 500 ng. In some cases, the amount of CNA is 0.1 ng to 200 ng, 1ng-100ng, 5ng-80ng, 10ng-60ng, or 20ng-50ng The concentration of CNA in the sample must be at least 0.1 ng / mL, 0.5 ng / mL, or mL, 0.6ng / mL, 0.8ng / mL, 1ng / mL, 2ng / mL, 5ng / m L, 10ng / mL, 0.01ng / μL, 0.05ng / μL, 0.1ng / μL, 0 .2ng / μL, 0.4ng / μL, 0.8ng / μL, 1ng / μL, 1.2ng / μL L, 1.5 ng / μL, 2 ng / μL, 5 ng / μL or 10 ng / μL. In some cases, the amount of CNA added to the sample ranges from about 1 ng / 15 μL to about 5 ng / In some cases, the amount of CNA added to the sample can be in the range of about It can be in the range of 0.05 ng / μL to about 0.5 ng / μL.
[0318] The methods herein include any type of synthesis described throughout this disclosure. For example, the method can include adding a synthetic nucleic acid: Synthetic nucleic acids for the production of a constant library, synthetic nucleic acids for normalizing the relative abundance of target nucleic acids Synthetic nucleic acids (e.g., synthetic nucleic acids of known concentration), and / or nucleic acid diversity in the sample This may include adding one or more synthetic nucleic acids to determine the reduction.
[0319] Nucleic acid extraction The method can include extracting nucleic acid (e.g., target nucleic acid, cell-free nucleic acid) from the sample. Extraction removes other cellular components and contaminants that may be present in the sample, such as biological fluids or This may involve isolating nucleic acids from a phenol-based or tissue sample. Choleroform extraction or precipitation with organic solvents (e.g., ethanol or isopropanol) In some cases, extraction is performed using a nucleic acid binding column. In this case, commercially available kits, such as the Qiagen Qiamp Circulating Nucleic Acid Kit, are available. (Circulating Nucleic Acid Kit) Qiagen Q ubit dsDNA HS Assay Kit, Agilent™ DNA 1000 Kit, TruSeq™ Sequencing Library Preparation brary preparation), or nucleic acid binding spin columns (e.g., Qiagen Extraction is performed using a gen DNA miniprep kit. In some cases, cell-free The extraction of nucleic acids may involve filtration or ultrafiltration.
[0320] The nucleic acid may be added to the sample before or during extraction. For example, a carrier nucleic acid may be added to the sample before or during extraction. It may be added to the sample before being mixed with an extraction reagent, such as an extraction buffer. The carrier nucleic acid is added to an extraction reagent, e.g., an extraction buffer, which is then mixed with the sample. In some cases, the CNA may be mixed with the sample and an extraction reagent, e.g., an extraction buffer. In these cases, the target nucleic acid and the CNA can be extracted simultaneously. .
[0321] Addition of CNA to a sample can increase the yield of nucleic acid extraction. The extraction yield is, for example, at least 10%, 20%, 40%, 60%, 80%, 100%, 2x, 4x, 6x, 8x or 1 In some cases, the CNA may be 0 times higher than the target nucleic acid in the sample after nucleic acid extraction. The extract may contain at least 10 ng, 50 ng, 100 ng, 200 ng, 3 00ng, 400ng, 500ng, 600ng, 700ng, 800ng, 900ng Alternatively, 1000 ng of nucleic acid may be provided.
[0322] Nucleic acid purification The method can include purifying the target nucleic acid. Exemplary purification methods include ethanol precipitation. Purification by isopropanol precipitation, phenol-chloroform purification, and column purification (e.g., These include affinity-based column purification, dialysis, filtration or ultrafiltration.
[0323] The nucleic acid may be added to the sample before or during purification. For example, a carrier nucleic acid may be added to the sample before or during purification. It may be added to the sample before being mixed with a purification reagent, such as a purification buffer. The carrier nucleic acid is added to a purification reagent, such as a purification buffer, which is then mixed with the sample. In some cases, the CNA may be mixed with the sample and purification reagents, e.g., purification buffer. In these cases, the target nucleic acid and the CNA can be extracted simultaneously. .
[0324] Addition of CNA to a sample can increase the yield of nucleic acid purification. The purification yield is, for example, at least 10%, 20%, 40%, 60%, 80%, 100%, 2x, 4x, 6x, 8x or 1 In some cases, the CNA may be 0 times higher than the target nucleic acid after nucleic acid purification. In some cases, purification of nucleic acids in samples spiked with CNAs can be achieved by At least 1 pg, 10 pg, 50 pg, 100 pg, 500 pg, or 500 pg of total nucleic acid in the sample pg, 1ng, 5ng, 10ng, 50ng, 100ng, 200ng, 300ng, 4 00ng, 500ng, 600ng, 700ng, 800ng, 900ng or 100 In some cases, purification of nucleic acids in samples spiked with CNAs was At least 1 pg, 10 pg, 50 pg, 100 pg, 50 pg, 0pg, 1ng, 5ng, 10ng, 50ng, 100ng, 200ng, 300ng, 400ng, 500ng, 600ng, 700ng, 800ng, 900ng or 10 Give 00ng.
[0325] fragmentation The method can include fragmenting the target nucleic acid. Fragmentation of the target nucleic acid can be accomplished, for example, by mechanical by shearing, passing the sample through a syringe, sonication, heat treatment or a combination thereof. In some cases, fragmentation of the target nucleic acid can be achieved by the use of nucleases or transposases. The fragmentation is carried out by using enzymes including nucleases. Restriction endonucleases, homing endonucleases, nicking endonucleases The method may include a restriction enzyme, a high fidelity restriction enzyme, or any of the enzymes disclosed herein. The target nucleic acid is subjected to a certain length, e.g., at least 50, 60, 80, 100, 120, 14 0, 160, 180, 200, 300, 400, 500, 1000, 2000, 4000 This may include fragmenting the nucleic acid into fragments of 6000, 8000 or 10000 bp in length. The CNA may be added to the sample prior to fragmentation of the target nucleic acid. It can be added to the sample at a later time.
[0326] A-tailing The method may include A-tailing the target nucleic acid. The A-tailing reaction can be carried out using one or more A-tailing enzymes. For example, Nonproofreading DNA polymerases that add a single 3' adenine residue and Incubating DNA with ATP adds adenine (A) residues. It is possible to add CNA to the sample containing the target nucleic acid prior to A-tailing. Alternatively, after A-tailing, a sample containing the target nucleic acid can be added with CNA. It is possible to add
[0327] End Repair The method can include performing end repair on a target nucleic acid. For example, the target nucleic acid can have the sequence End repair can be performed on the target nucleic acid so that it can be compatible with other steps of the determination library. The end repair reaction can be carried out using one or more end repair enzymes. The enzymes for this purpose may include a polymerase and an exonuclease. For example, a polymerase may It is possible to fill in the missing bases of the DNA strand in the 5' to 3' direction. The strand DNA can have substantially the same length as the original longest DNA strand. The overhangs can be removed. The resulting double-stranded DNA is substantially identical to the original shortest DNA strand. It may have a length.
[0328] CNAs can be added to a sample containing the target nucleic acid prior to end repair. The addition of CNA increases the efficiency of the end repair reaction by, for example, at least 10%, 20%, 40%, 60%, 80% or 100% increase. In some cases, CNAs are In some cases, the addition of the CNA can be performed by the addition of an enzyme For example, the enzyme may maintain the activity and / or function of an end repair enzyme. In samples with low abundance, they may have reduced activity or abnormal function, The addition of CNA can increase the total amount of nucleic acid in the sample, thereby The element is capable of functioning normally in the sample.
[0329] Adapter Binding The method can include attaching one or more adaptors to the target nucleic acid. The adaptors can be , can be bound to the target nucleic acid by primer extension, reverse transcription or hybridization. In some cases, the adaptor is attached to the target nucleic acid by ligation. For example, the adaptor The adaptor can be attached to the target nucleic acid by a ligase. For example, the adaptor can be attached to the target nucleic acid by sticky end ligation or The adaptor can be attached to the target nucleic acid by blunt end ligation. In some cases, the adaptor is The target nucleic acid can be bound to the 3' end, 5' end or both by the spase. In some cases, the target nucleic acid may be linked to an adaptor at both ends. At the ends of the Alternatively, the target nucleic acid may be bound at one end to one or more adaptors.
[0330] The CNA can be added before the binding step. Alternatively, the CNA can be added after the binding step. CNAs can be resistant to ligation. For example, CNAs can be resistant to ligation reactions. In these cases, if CNA is added before the binding step, These are not ligated to either the target nucleic acid or the adapter and are not sequenced during the sequencing step. In other cases, CNAs may be removed from the sample prior to the binding step. Alternatively, CNAs may be removed after sample extraction and before the binding step.
[0331] Treating the sample with an enzyme prior to binding of the adaptor to the target nucleic acid in the sample For example, a sample can be treated with an endonuclease to remove ligation sites, e.g. For example, the adaptors can generate sticky or blunt ends. After binding to the nucleic acids, the sample can be treated with an enzyme.
[0332] amplification The method can include amplifying the target nucleic acid. Amplification increases the number of copies of a nucleic acid sequence. For example, amplification can refer to any method for increasing the number of copies of a gene by, for example, one or more polymerase chain reactions. The reaction may be carried out using a polymerase. Amplification may be carried out using methods known in the art. These methods often involve the preparation of multiple copies of a nucleic acid or its complement. One such method is the synthesis of polymers containing These include: AFLP (amplified fragment length polymorphism) PCR, allelic Child-specific PCR, Alu PCR, Assembly, Asymmetric PCR, Colony PCR, Helical enzyme-dependent PCR, hot start PCR, inverse PCR, in situ Inter-sequence specific PCR or IS-SR PCR, digital PCR, drop Droplet digital PCR, linear post-exponential PCR or rate (L ate PCR, long PCR, nested PCR, real-time PCR , duplex PCR, multiplex PCR, quantitative PCR or single cell PCR. Ligand chain reaction (LCR), nucleic acid sequence-based amplification (NASBA), linear amplification, isothermal linear amplification , Q-beta-replicase method, 3SR, transcription-mediated amplification (TMA), strand displacement (Stran d Displacement Amplification (SDA) or Rolling Circle Amplification (RCA) ) may also be used.
[0333] The CNA may be added before amplification. Alternatively, the CNA may be added after amplification. For example, the CNA may contain modifications that inhibit amplification. In these cases, if the CNA is added before amplification, it will not be amplified. can be absent from the sequencing library or not sequenced.
[0334] Removal of CNAs The method can further include removing CNAs from the sample, which can include: Often, CNAs are prevented from being sequenced. In some cases, the method This involves removing some or all of the CNAs from the sample to prepare the sequencing sample. The sequencing sample used may be free of CNAs and may be used as is for sequencing. In some cases, the method may also be used to detect other nucleic acids in the sample, such as the target At least one CNA, preferably over a nucleic acid, an adaptor or a multimer of adaptors. This includes removing.
[0335] Removal of CNAs can be accomplished using enzymes. For example, CNAs can be removed using enzymes, such as enzyme digestion. In some cases, the method uses a nuclease to remove CNAs. For example, the method includes removing endonucleases, such as type I, type II (II CNAs can be synthesized using type III or type IV endonucleases, including type S and type IIG. The method may include removing a restriction endonuclease, such as AatII, A cc65I, AccI, AclI, AatII, Acc65I, AccI, AclI, A feI, AflII, AgeI, ApaI, ApaLI, ApoI, AscI, AseI , AsiSI, AvrII, BamHI, BclI, BglII, Bme1580I, B mtI, BsaHI, BsiEI, BsiWI, BspEI, BspHI, BsrGI, BssHII, BstBI, BstZ17I, BtgI, ClaI, DraI, EaeI , EagI, EcoRI, EcoRV, FseI, FspI, HaeII, HincII , HindIII, HpaI, KasI, KpnI, MfeI, MluI, MscI, M spA1I, MfeI, MluI, MscI, MspA1I, NaeI, NarI, Nc oI, NdeI, NgoMIV, NheI, NotI, NruI, NsiI, NspI, PacI, PciI, PmeI, PmlI, PsiI, PspOMI, PstI, Pvu I, PvuII, SacI, SacII, SalI, SbfI, ScaI, SfcI, S foI, SgrAI, SmaI, SmlI, SnaBI, SpeI, SphI, SspI , StuI, SwaI, XbaI, XhoI, XmaI or any combination thereof The method may include removing CNAs using a DNA fragment not listed above. This may include removing CNAs using enzymes such as exodeoxyribonucleases. The method comprises the steps of: endonuclease (endonuclease VIII) or a mixture thereof (e.g., uracil-specific cleavage The method may include removing CNAs using a scission inhibitor (USER) enzyme. Site of NA-induced DNase, e.g., CRISPR-associated protein nuclease, e.g., C The method may include removing CNAs using RNase, such as RNase A. endoribonucleases, such as RNase A, RNase H, RNase I II, RNase L, RNase P, RNase PhyM, RNase T1, RNase T2, RNase U2, RNase V, or exoribonuclease, e.g., polynucleotide Leutidine phosphorylase, RNase PH, RNase R, RNase D, RNase T , oligoribonuclease, exoribonuclease I or exoribonuclease I I, or any combination thereof to remove carrier-synthesized nucleic acids. In some cases, the method includes the use of any nucleic acid degrading agent known in the art. In some cases, the method includes removing CNAs by physical treatment, e.g. This may include removing the CNAs by subjecting the material to heating, cooling or shearing. In some cases, the method for removing CNAs may involve the addition of CNAs to the target nucleic acid, the adapter, or the sequencing library. It does not remove any other molecules from the sample. In some cases, removal of CNAs is achieved b...
Claims
1. (a) adding a starting amount of at least 1,000 synthetic nucleic acids to a sample, wherein said each of the at least 1,000 synthetic nucleic acids comprises a unique variable region; (b) a portion of the target nucleic acid in the sample and at least 1,000 of the above-mentioned synthetic A sequencing assay is performed on a portion of the nucleic acid, thereby identifying the target and synthetic nucleic acid sequence reads. obtaining a synthetic nucleic acid sequence read comprising a unique variable region sequence; (c) (i) quantifying the number of distinct variable region sequences within the synthetic nucleic acid sequence reads to identify unique Obtaining a sequence determination value, (ii) deriving a starting amount of said at least 1,000 synthetic nucleic acids from said unique sequences; By obtaining a diversity reduction of said at least 1,000 synthetic nucleic acids compared to the determined value. detecting a decrease in diversity of the at least 1,000 synthetic nucleic acids; (d) using the diversity reduction of the at least 1,000 synthetic nucleic acids to generate a nucleic acid sequence similar to that of the initial sample; calculating the abundance of the target nucleic acid in an initial sample containing the target nucleic acid A method for determining the abundance of a nucleic acid.
2. The method of claim 1 , wherein the target nucleic acid comprises a pathogen nucleic acid.
3. Any of the preceding claims, wherein the target nucleic acid comprises pathogen nucleic acid from at least five different pathogens.
2. The method according to any one of claims 1 to 11.
4. 2. Any one of the preceding claims, wherein the at least 1,000 synthetic nucleic acids comprise DNA. The method described.
5. Each of the at least 1,000 synthetic nucleic acids is 500 base pairs or nucleotides long.
2. The method of claim 1, wherein the length of the first and second strands is less than 100 μm.
6. Each of the at least 1,000 synthetic nucleic acids is 200 base pairs or nucleotides long.
2. The method of claim 1, wherein the length of the first and second strands is less than 100 μm.
7. If the sample is blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage fluid, urine, stool, saliva or The method of any one of the preceding claims, wherein is a nasal sample.
8. 2. The method according to any one of the preceding claims, wherein the sample is a sample of isolated nucleic acid. 。
9. The method further comprises producing a sequencing library from the sample, wherein the sequencing library adding said at least 1,000 synthetic nucleic acids to the sample prior to preparing the library; A method according to any one of the preceding claims.
10. The diversity reduction of the at least 1,000 synthetic nucleic acids is achieved by the sample processing of the sample.
2. The method of any one of the preceding claims, which exhibits a reduction in one or more nucleic acids.
11. The method of claim 1, wherein each of the at least 1,000 synthetic nucleic acids comprises an identification tag sequence.
3. The method according to any one of the preceding claims.
12. Quantifying the number of unique variable region sequences includes detecting sequences that contain the tag sequence.
10. A method according to any one of the preceding claims.
13. 2. The method of any one of the preceding claims, wherein the sample is from a human subject.
14. Quantification of at least 1,000 unique sequences within the first sequence read The method of any one of the preceding claims, comprising determining the number of unique sequence reads of 。
15. The at least 1,000 unique synthetic nucleic acids are 4 Unique combination 2. The method of claim 1, comprising a synthetic nucleic acid.
16. The method of claim 1, further comprising adding additional synthetic nucleic acids having at least three different lengths.
2. The method of claim 1 .
17. A first additional synthetic nucleic acid group having a first length, a second additional synthetic nucleic acid group having a second length, and a third additional synthetic nucleic acid group having a third length, wherein each of the first, second and third additional synthetic nucleic acid groups comprises at least three different G 2. The method of claim 1, further comprising a synthetic nucleic acid having a C content.
18. The additional synthetic nucleic acid is used to calculate the absolute abundance of the target nucleic acid in the sample. The method of claim 15 or 16, further comprising:
19. The additional synthetic nucleic acid is used to determine the length, GC content, or length and and calculating the abundance of the target nucleic acid in the sample based on both the GC content and the abundance of the target nucleic acid in the sample. The method according to claim 15 or 16.
20. In a first sample processing step, the at least 1,000 synthetic nucleic acids are treated with a sample.
4. The method of claim 1, in addition to
21. In a second sample processing step, an additional 1,000 unique synthetic nucleic acids are The method further comprises adding the pool to a sample, wherein the second sample processing step comprises adding the pool to the first sample.
21. The method of claim 20, wherein the process is different from the ole processing step.
22. Calculating the diversity reduction for an additional pool of at least 1,000 synthetic nucleic acids.
21. The method of claim 20, comprising:
23. The diversity reduction for the at least 1,000 synthetic nucleic acids is at least 1,000 The relatively high diversity is compared to the reduced diversity associated with an additional pool of synthetic nucleic acids.
21. The method of claim 20, further comprising identifying sample processing steps that exhibit reduced activity.
24. Unique synthetic nucleic acids in an additional pool of at least 1,000 unique synthetic nucleic acids. Each of the synthetic nucleic acids is selected from the group consisting of at least 1,000 synthetic nucleic acids. The method of claim 20, further comprising a domain that identifies the synthetic nucleic acid.
25. 2. The method of claim 1, further comprising adding a sample-discriminating nucleic acid to the sample. How to.
26. Any of the preceding claims, wherein (a) further comprises adding a non-unique synthetic nucleic acid to the sample. The method according to any one of claims 1 to 5.
27. The method of claim 1 , wherein the calculated abundance is a relative abundance.
28. The method of claim 1 , wherein the calculated abundance is an absolute abundance.
29. (a) obtaining a sample from a subject infected or suspected of being infected with a pathogen, , the sample contains multiple pathogen nucleic acids, (b) adding a plurality of synthetic nucleic acids to a sample such that the sample contains a known initial abundance of the synthetic nucleic acid. In addition, here, (i) the synthetic nucleic acid is less than 500 base pairs in length; (ii) The synthetic nucleic acid is a synthetic nucleic acid having a first length, a synthetic nucleic acid having a second length, and a synthetic nucleic acid having a third length, wherein the first, second and third lengths are different. And, (iii) the synthetic nucleic acid having a first length has at least three different GC contents; a synthetic nucleic acid (c) performing a sequencing assay on a sample containing said plurality of synthetic nucleic acids; determining the final abundance of the synthetic nucleic acid and the final abundance of the plurality of pathogen nucleic acids; (d) comparing the final abundance of the synthetic nucleic acid with the known initial abundance to determine the recovery for the synthetic nucleic acid; Get a profile, (e) Using the recovery profile for the synthetic nucleic acid, the pathogen nucleic acid is identified as the closest GC containing and comparing the amount and length of the synthetic nucleic acid to the plurality of pathogen nucleic acids, thereby determining the relative By determining the abundance or initial abundance, the final abundance of said plurality of pathogen nucleic acids is determined. Normalizing the relative abundance or initial abundance of pathogen nucleic acid in a sample. How to decide.
30. a first GC content, wherein the at least three different GC contents are between 10% and 40%; 4 a second GC content that is between 0% and 60%, and a third GC content that is between 60% and 90%.
10. A method according to any one of the preceding claims.
31. The at least three different GC contents are each between 10% and 50%. The method according to any one of the preceding claims.
32. 2. The synthetic nucleic acid of claim 1, wherein the synthetic nucleic acid is less than 200 base pairs or nucleotides in length. The method described.
33. 2. The synthetic nucleic acid of claim 1, wherein the synthetic nucleic acid is less than 100 base pairs or nucleotides in length. The method described.
34. 2. The method of claim 1, wherein the synthetic nucleic acid comprises double-stranded DNA.
35. The method of any preceding claim further comprising using the synthetic nucleic acid to monitor the alteration of the pathogen nucleic acid. The method according to any one of claims 1 to 5.
36. Weighting factors are used to normalize the relative abundance or initial abundance of pathogen nucleic acids 2. The method of claim 1, further comprising:
37. The plurality of synthetic nucleic acids are compared with a known concentration of the first synthetic nucleic acid and a known concentration of the second synthetic nucleic acid. A raw measurement value of a first synthetic nucleic acid of the nucleic acid and a raw measurement value of a second synthetic nucleic acid of the plurality of synthetic nucleic acids are 37. The method of claim 36, wherein the analyzing derives the weighting factors.
38. (a) obtaining a first sample containing a first pathogen nucleic acid, wherein the first sample is from a first subject infected with (b) obtaining a second sample from a second subject; (c) first nucleic acids, each of which comprises a different synthetic nucleic acid that cannot hybridize to the first pathogen nucleic acid; Obtaining a sample identifier and a second sample identifier, and assigning the first sample identifier to the first sample. assigning a second sample identifier to the second sample; (d) assigning the first sample identifier to the first sample and the second sample identifier to the second sample. In addition to (e) for a first sample including a first sample identifier, and performing a sequencing assay on a second sample containing the first sample and the second sample containing the Obtain sequence results for two samples; (f) a first sample identifier and a second sample identifier in the sequence result for the first sample detecting the presence or absence of the first and second pathogen nucleic acids; (g) the sequencing assay comprises, in a first sample: (i) detecting a first sample identifier; (ii) detecting a first pathogen nucleic acid; and (iii) not detecting a second sample identifier or detecting a second sample identifier below a threshold level; If the nucleic acid identifier is not detected, the first pathogen nucleic acid detected is not originally present in the first sample. and determining that the pathogen is positive for the nucleic acid.
39. (a) obtaining a first nucleic acid sample comprising a first nucleic acid; (b) obtaining a first control nucleic acid sample comprising a first positive control nucleic acid; (c) combining a first sample identifier with a first pair, the first sample identifier comprising a synthetic nucleic acid that cannot hybridize to the first nucleic acid. In addition to nucleic acid identification, (d) a first control nucleic acid sample and a first nucleic acid sample comprising a first sample identifier. and performing a sequencing assay using the nucleic acid sequence of the first and control nucleic acid samples. Get the (e) aligning the sequence reads for the first nucleic acid sample with a reference sequence to identify the first nucleic acid sample; Detecting the presence or absence of a first sample identifier in sequence reads for the acid sample. 、 (f) determining whether the first positive control nucleic acid is a first nucleic acid sample based on the alignment of the sequence reads; 4. A method for detecting a nucleic acid comprising determining whether a nucleic acid is present in
40. the synthetic nucleic acid of the first sample identifier is less than 150 base pairs or nucleotides in length; 2. The method of claim 1 .
41. 2. The method of claim 1, wherein the first positive control nucleic acid is a pathogen nucleic acid.
42. 2. The method of any one of the preceding claims, wherein the first sample identifier comprises a modified nucleic acid.
43. 2. The method of claim 1, wherein the first sample identifier comprises DNA.
44. 2. The method of any one of the preceding claims, wherein the sample comprises an acellular body fluid.
45. 13. The method of claim 1, wherein the sample is from a subject infected with a pathogen. The method described.
46. (a) adding a first synthetic nucleic acid to the reagent, wherein the first synthetic nucleic acid comprises a unique sequence; (b) adding a reagent comprising a first synthetic nucleic acid to the nucleic acid sample; (c) preparing a nucleic acid sample for a sequencing assay; (d) performing a sequencing assay on the nucleic acid sample, thereby determining whether or not a sequence of interest is present on the nucleic acid sample; The resulting sequence is (e) determining the sequence of the first synthetic nucleic acid in the nucleic acid sample based on the sequence results for the sample; detecting the reagent in the sample by determining its presence or absence. A method for detecting a reagent in a sample.
47. 13. Any of the preceding claims, wherein the first synthetic nucleic acid is less than 150 base pairs or nucleotides in length.
2. The method according to claim 1.
48. A first synthetic nucleic acid is added to a first reagent lot, and a second synthetic nucleic acid is added to a second reagent lot.
2. The method of claim 1, comprising:
49. detecting the reagent in the sample includes detecting a particular lot of the reagent; A method according to any one of the preceding claims.
50. 2. The method according to any one of the preceding claims, wherein the synthetic nucleic acid is not degradable by nucleases. 。
51. 2. The method of claim 1, wherein the reagent comprises an aqueous buffer.
52. The method of any preceding claim, wherein the reagent comprises an extraction reagent, an enzyme, a ligase, a polymerase, or a dNTP. The method according to any one of claims 1 to 5.
53. (a) a nucleic acid sequence comprising: (i) a target nucleic acid; (ii) a sequencing adaptor; and (iii) at least one A sample is obtained that contains at least one synthetic nucleic acid, wherein the at least one synthetic nucleic acid is a DNA. and resists binding to nucleic acids; (b) the sequencing adaptor preferentially binds to the target nucleic acid over the at least one synthetic nucleic acid. performing a ligation reaction on the sample to ligate to a sequencing library. A manufacturing method of the present invention.
54. (a) obtaining a sample comprising a target nucleic acid and at least one synthetic nucleic acid; (b) removing said at least one synthetic nucleic acid from said sample, thereby obtaining a target nucleic acid. obtaining a sequencing sample that contains an acid and is free of said at least one synthetic nucleic acid; (c) attaching a sequencing adaptor to a target nucleic acid in the sequencing sample. A method for producing a sequencing library.
55. (a) obtaining a sample comprising a target nucleic acid and at least one synthetic nucleic acid; (b) binding a sequencing adaptor to a target nucleic acid in the sample, thereby determining a sequence Obtain a sequence decision sample, (c) subjecting said at least one synthetic nucleic acid to affinity-based depletion, RNA-guided depletion, or a combination thereof, The removal of the at least one synthetic nucleic acid from the sequencing sample comprises removing the at least one synthetic nucleic acid from the sequencing sample. and more preferably the at least one synthesis member than the adapter and more preferably the multimer of the sequencing adapter. and preferentially removing synthetic nucleic acids.
56. (a) obtaining a sample containing a target nucleic acid and at least one synthetic nucleic acid, wherein said at least one synthetic nucleic acid is At least one synthetic nucleic acid comprises: (i) single-stranded DNA, (ii) a nucleotide modification that inhibits amplification of the synthetic nucleic acid; (iii) an immobilization tag; (iv) DNA-RNA hybrids, (v) a nucleic acid having a length greater than the length of the target nucleic acid; or (vi) any combination thereof, (b) preparing a sequencing library from the sample for a sequencing reaction (herein); and wherein at least a portion of said at least one synthetic nucleic acid is sequenced in said sequencing reaction. (not specified)
57. (a) a nucleic acid sequence comprising: (i) a target nucleic acid; (ii) a sequencing adaptor; and (iii) at least one A sample is obtained that contains at least one synthetic nucleic acid, wherein the at least one synthetic nucleic acid is a DNA. Contains, resists end repair, (b) selecting a target nucleic acid such that the target nucleic acid is end-repaired preferentially over the at least one synthetic nucleic acid; and performing an end repair reaction on the sample. Law.
58. Any of the preceding claims including reporting the results of the method to a patient, caregiver or other person.
3. The method according to claim 1.
59. (a) a sequencing adaptor, and (b) at least one synthetic nucleic acid, wherein said at least one synthetic nucleic acid is DNA and resists end repair to nucleic acids) A kit for producing a sequencing library comprising: