Pathogen Diagnostic Tests

PCR and sequencing methods with direct lysis and synthetic nucleic acid addition facilitate rapid, sensitive, and cost-effective detection of viral infections, addressing the need for efficient SARS-CoV-2 diagnosis and monitoring.

JP7801240B2Active Publication Date: 2026-01-16OCTANT INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022560915
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-26
Filing Date
2021-04-05
Publication Date
2026-01-16
Estimated Expiration
2041-04-05

AI Technical Summary

Technical Problem

There is a need for rapid, sensitive, and cost-effective tests for viral infections, particularly for the novel coronavirus SARS-Cov-2, to efficiently identify and monitor outbreaks and guide healthcare responses.

Method used

Methods involving PCR and sequencing are used for highly specific and sensitive detection of viral genomes, including reverse transcription and amplification directly in a lysis agent without purification, with the addition of synthetic nucleic acid for accurate quantification and multiplexing capabilities to detect multiple pathogens.

Benefits of technology

The methods enable accurate detection of viral infections down to low copy numbers, allowing for rapid and cost-effective diagnosis of pathogens such as SARS-CoV-2 and other viruses, including influenza, with the potential to determine strain information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007801240000029
    Figure 0007801240000029
  • Figure 0007801240000030
    Figure 0007801240000030
  • Figure 0007801240000031
    Figure 0007801240000031
Patent Text Reader

Abstract

Described herein are methods useful for detecting and diagnosing pathogen infections using PCR and sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 005,996, filed April 6, 2020, U.S. Provisional Application No. 63 / 062,406, filed August 6, 2020, U.S. Provisional Application No. 63 / 136,449, filed January 12, 2021, and U.S. Provisional Application No. 63 / 154,571, filed February 26, 2021, each of which is incorporated herein by reference in its entirety. [Background technology]

[0002] The inventors and applicants intend to freely license certain of the subject matter described herein to parties advancing the common goal of ending the COVID-19 outbreak consistent with the Open Covid Pledge of March 31, 2020.

[0003] Viral diseases and outbreaks have plagued humanity for millennia. A key part of identifying, monitoring, and guiding viral diseases is the ability to efficiently and accurately test individuals for viral infections. The novel coronavirus SARS-Cov-2, which causes COVID-19, emerged in Wuhan Province, China, in late 2019. From there, through the first half of 2020, the virus spread rapidly around the world, overwhelming healthcare systems and paralyzing commerce. Rapid, sensitive, and cost-effective tests are needed. Summary of the Invention

[0004] Described herein are methods for diagnosing individuals with pathogen infections. The methods use PCR and sequencing to achieve highly specific and sensitive detection of viral genomes from biological samples. Features of the methods described herein that enable such detection include: 1) reverse transcription and / or amplification directly in a lysis agent or after lysis conditions without purification or isolation; 2) the presence of a synthetic nucleic acid spiked into the reverse transcription amplification mixture that is amplifiable by oligonucleotide primers targeting the viral sequence of interest but contains a distinct intervening sequence, allowing for more accurate quantification and a lower detection threshold; and 3) indexing to enable multiplexing with next-generation sequencing. In certain embodiments, the pathogen is a viral infection (e.g., SARS-CoV-2). The methods described herein can also be multiplexed to detect more than one viral pathogen (e.g., SARS-CoV-2 and influenza A or B, or both).

[0005] In one aspect, the present specification describes a method for diagnosing an individual with a pathogen infection, the method comprising: a) providing a biological sample from the individual; a) contacting the biological sample from the individual with a lysis agent to obtain a lysed biological sample; b) performing a polymerase chain reaction (PCR) on the lysed biological sample to obtain a PCR-amplified lysed biological sample, wherein the PCR reaction on the lysed biological sample is performed using a first set of PCR primers, and the first set of PCR primers amplifies a pathogen nucleic acid sequence; c) sequencing the PCR-amplified lysed biological sample using next-generation sequencing; and d) providing a positive diagnosis for the pathogen infection if a pathogen sequence is detected by the PCR or sequencing, or providing a negative diagnosis for the individual if a pathogen sequence is not detected by the PCR or sequencing. In certain embodiments, the individual is a human. In certain embodiments, the pathogen infection includes a bacterial infection, a viral infection, a fungal infection, and a combination thereof. In certain embodiments, the bacterial infection is an infection caused by Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof. In certain embodiments, the fungal infection is an infection caused by Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. In certain embodiments, the viral infection is an infection caused by a DNA virus. In certain embodiments, the DNA virus includes hepatitis A, hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or smallpox, and combinations thereof. In certain embodiments, the viral infection is an infection caused by an RNA virus.In certain embodiments, the RNA virus comprises an influenza virus, a coronavirus, a poliovirus, a measles virus, an Ebola virus, a retrovirus, or an orthomyxovirus. In certain embodiments, the viral infection is a coronavirus infection. In certain embodiments, the coronavirus infection is a SARS-COV-2 infection. In certain embodiments, the biological sample from the individual is from a blood sample, a plasma sample, a serum sample, a buccal swab, a urine sample, a semen sample, a vaginal swab, a stool sample, a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. In certain embodiments, the biological sample from the individual is from a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. In certain embodiments, the method comprises adding a synthetic nucleic acid to the lysis agent or the lysed biological sample. In certain embodiments, the synthetic nucleic acid is RNA. In certain embodiments, the synthetic nucleic acid is DNA. In certain embodiments, the synthetic nucleic acid comprises a set of sequences configured to be bound by the first set of PCR primers. In certain embodiments, the first set of primers amplifies both the pathogen nucleic acid sequence and the synthetic nucleic acid. In certain embodiments, the synthetic nucleic acid sequence comprises a nucleotide sequence that is not identical to the pathogen nucleic acid sequence. In certain embodiments, the method includes performing a reverse transcription reaction on the lysed biological sample. In certain embodiments, the reverse transcription reaction is performed before performing the polymerase chain reaction. In certain embodiments, the reverse transcription reaction is performed without further purification of the lysed biological sample. In certain embodiments, the reverse transcription reaction on the lysed biological sample produces viral cDNA. In certain embodiments, the viral cDNA is coronavirus cDNA. In certain embodiments, the coronavirus cDNA is SARS-COV-2 cDNA. In certain embodiments, the reverse transcription reaction and the PCR are a single-step reaction. In certain embodiments, the PCR is an end-point analysis. In certain embodiments, the PCR is not a real-time PCR reaction.In certain embodiments, the first set of PCR primers amplifies a coronavirus nucleic acid sequence. In certain embodiments, the coronavirus nucleic acid sequence is a SARS-COV-2 nucleic acid sequence. In certain embodiments, the SARS-COV-2 nucleic acid sequence comprises an N1 gene or an S2 gene. In certain embodiments, the method includes a second set of primers. In certain embodiments, the second set of PCR primers amplifies a human nucleic acid sequence. In certain embodiments, the second set of PCR primers amplifies a human nucleic acid sequence selected from GAPDH, ACTB, RPP30, and combinations thereof. In certain embodiments, the second set of PCR primers amplifies human RPP30. In certain embodiments, the second set of PCR primers includes a mixture of primers with sequencing adapter sequences and primers without sequencing adapter sequences. In certain embodiments, the ratio of primers with sequencing adapter sequences to primers without sequencing adapter sequences is about 1:1, about 1:2, about 1:3, or about 1:4. In certain embodiments, the PCR comprises 30 to 45 amplification cycles. In certain embodiments, the PCR comprises 35 to 45 amplification cycles. In certain embodiments, the PCR comprises 39 to 42 amplification cycles. In certain embodiments, the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise nucleic acid sequences comprising a variable nucleotide sequence. In certain embodiments, the variable nucleotide sequence is a sample ID unique to the individual. In certain embodiments, the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise adapter sequences for next-generation sequencing reactions. In certain embodiments, the method is capable of detecting fewer than 10 copies of a pathogen genome. In certain embodiments, the method is capable of detecting fewer than 5 copies of a pathogen genome. In certain embodiments, the pathogen genome is a coronavirus genome.In certain embodiments, the coronavirus genome is a SARS-COV-2 genome. In certain embodiments, if a coronavirus sequence is detected by the PCR, the positive diagnosis of coronavirus is a SARS-COV-2 diagnosis. In certain embodiments, the method determines the strain of coronavirus. In certain embodiments, the method determines the strain of COVID-19.

[0006] Also described herein is a method for diagnosing an individual with a pathogen infection, the method comprising: (a) providing a biological sample from the individual; (c) performing a polymerase chain reaction (PCR) on the biological sample to obtain a PCR-amplified biological sample, wherein the PCR reaction on the biological sample is performed using a first set of PCR primers, and the first set of PCR primers amplifies a pathogen nucleic acid sequence; (d) sequencing the PCR-amplified biological sample using next-generation sequencing; and (e) providing a positive diagnosis for the pathogen infection if a pathogen sequence is detected by the PCR or sequencing, or providing a negative diagnosis for the individual if a pathogen sequence is not detected by the PCR or sequencing. In certain embodiments, the individual is a human individual. In certain embodiments, the pathogen infection includes a bacterial infection, a viral infection, a fungal infection, and combinations thereof. In certain embodiments, the bacterial infection is an infection caused by Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof. In certain embodiments, the fungal infection is an infection caused by Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. In certain embodiments, the viral infection is an infection caused by a DNA virus. In certain embodiments, the DNA virus includes hepatitis A, hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or smallpox, and combinations thereof. In certain embodiments, the viral infection is an infection caused by an RNA virus. In certain embodiments, the RNA virus comprises an influenza virus, a coronavirus, a poliovirus, a measles virus, an Ebola virus, a retrovirus, or an orthomyxovirus.In certain embodiments, the viral infection is a coronavirus infection. In certain embodiments, the coronavirus infection is a SARS-COV-2 infection. In certain embodiments, the biological sample from the individual is a blood sample, a plasma sample, a serum sample, a buccal swab, a urine sample, a semen sample, a vaginal swab, a stool sample, a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. In certain embodiments, the biological sample from the individual is derived from a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. In certain embodiments, the method includes adding a synthetic nucleic acid to a lysis agent or the lysed biological sample. In certain embodiments, the synthetic nucleic acid is RNA. In certain embodiments, the synthetic nucleic acid is DNA. In certain embodiments, the synthetic nucleic acid comprises a set of sequences configured to be bound by the first set of PCR primers. In certain embodiments, the first set of primers amplifies both the pathogen nucleic acid sequence and the synthetic nucleic acid. In certain embodiments, the synthetic nucleic acid sequence comprises a nucleotide sequence that is not identical to the pathogen nucleic acid sequence. In certain embodiments, the method comprises performing a reverse transcription reaction on the biological sample. In certain embodiments, the reverse transcription reaction is performed before performing the polymerase chain reaction. In certain embodiments, the reverse transcription reaction is performed without further purification of the biological sample. In certain embodiments, the reverse transcription reaction on the biological sample produces viral cDNA. In certain embodiments, the viral cDNA is coronavirus cDNA. In certain embodiments, the coronavirus cDNA is SARS-COV-2 cDNA. In certain embodiments, the reverse transcription reaction and the PCR are a single-step reaction. In certain embodiments, the PCR is an end-point analysis. In certain embodiments, the PCR is not a real-time PCR reaction. In certain embodiments, the first set of PCR primers amplifies a coronavirus nucleic acid sequence. In certain embodiments, the coronavirus nucleic acid sequence is a SARS-COV-2 nucleic acid sequence. In certain embodiments, the SARS-COV-2 nucleic acid sequence comprises the N1 gene or the S2 gene.In certain embodiments, the method includes a second set of primers. In certain embodiments, the second set of PCR primers amplifies a human nucleic acid sequence. In certain embodiments, the second set of PCR primers amplifies a human nucleic acid sequence selected from GAPDH, ACTB, RPP30, and combinations thereof. In certain embodiments, the second set of PCR primers amplifies human RPP30. In certain embodiments, the second set of PCR primers includes a mixture of primers with sequencing adapter sequences and primers without sequencing adapter sequences. In certain embodiments, the ratio of primers with sequencing adapter sequences to primers without sequencing adapter sequences is about 1:1, about 1:2, about 1:3, or about 1:4. In certain embodiments, the PCR includes 30 to 45 amplification cycles. In certain embodiments, the PCR includes 35 to 45 amplification cycles. In certain embodiments, the PCR includes 39 to 42 amplification cycles. In certain embodiments, the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise nucleic acid sequences that include a variable nucleotide sequence. In certain embodiments, the variable nucleotide sequence is a sample ID unique to the individual. In certain embodiments, the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise adapter sequences for next-generation sequencing reactions. In certain embodiments, the method is capable of detecting fewer than 10 copies of a pathogen genome. In certain embodiments, the method is capable of detecting fewer than 5 copies of a pathogen genome. In certain embodiments, the pathogen genome is a coronavirus genome. In certain embodiments, the coronavirus genome is a SARS-COV-2 genome. In certain embodiments, if a coronavirus sequence is detected by the PCR, the positive diagnosis of coronavirus is a SARS-COV-2 diagnosis. In certain embodiments, the method determines the coronavirus strain.In certain embodiments, the method determines the strain of COVID-19.

[0007] In another aspect, the present specification describes a method for diagnosing an individual with a pathogen infection, the method comprising: amplifying nucleic acids from a biological sample from the individual using a first set of PCR primers to obtain amplified nucleic acids, wherein the first set of PCR primers amplify a pathogen nucleic acid sequence and a synthetic nucleic acid sequence from the biological sample, and the synthetic nucleic acid sequence differs from the pathogen nucleic acid sequence by at least one nucleotide. In certain embodiments, the individual is a human. In certain embodiments, the pathogen infection comprises a bacterial infection, a viral infection, or a fungal infection. In certain embodiments, the bacterial infection is an infection caused by Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, or a combination thereof. In certain embodiments, the fungal infection is an infection caused by Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. In certain embodiments, the viral infection is an infection caused by a DNA virus. In certain embodiments, the DNA virus includes hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, smallpox, or any combination thereof. In certain embodiments, the viral infection is an infection caused by an RNA virus. In certain embodiments, the RNA virus includes influenza virus, coronavirus, poliovirus, measles virus, Ebola virus, retrovirus, or orthomyxovirus. In certain embodiments, the viral infection is a coronavirus infection. In certain embodiments, the coronavirus infection is a SARS-COV-2 infection. In certain embodiments, the biological sample from the individual is a blood sample, a plasma sample, a serum sample, a buccal swab, a urine sample, a semen sample, a vaginal swab, a stool sample, a nasopharyngeal swab, a middle turbinate swab, or any combination thereof.In certain embodiments, the biological sample from the individual is derived from a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. In certain embodiments, the synthetic nucleic acid is RNA. In certain embodiments, the synthetic nucleic acid is DNA. In certain embodiments, the synthetic nucleic acid comprises a set of sequences configured to be bound by the first set of PCR primers. In certain embodiments, the synthetic nucleic acid sequence differs from the pathogen nucleic acid sequence by at least five nucleotides. In certain embodiments, the pathogen infection is diagnosed based on the ratio of the synthetic nucleic acid sequence to the pathogen nucleic acid sequence. In certain embodiments, the method comprises performing a reverse transcription reaction on the nucleic acid from the biological sample. In certain embodiments, the reverse transcription reaction on the nucleic acid from the biological sample produces coronavirus cDNA. In certain embodiments, the coronavirus cDNA is SARS-COV-2 cDNA. In certain embodiments, the amplification of the amplified nucleic acid comprises a PCR reaction. In certain embodiments, the PCR reaction is an end-point analysis. In certain embodiments, the PCR reaction is not a real-time PCR reaction. In certain embodiments, the first set of PCR primers amplifies a coronavirus nucleic acid sequence. In certain embodiments, the coronavirus nucleic acid sequence is a SARS-COV-2 nucleic acid sequence. In certain embodiments, the SARS-COV-2 nucleic acid sequence comprises an N1 gene or an S2 gene. In certain embodiments, the first set of PCR primers comprises a nucleic acid sequence comprising a variable nucleotide sequence. In certain embodiments, the variable nucleotide sequence is a sample ID unique to the individual. In certain embodiments, the first set of PCR primers comprises an adapter sequence for a next-generation sequencing reaction. In certain embodiments, the method includes amplifying nucleic acids from the biological sample using a second set of PCR primers, wherein the second set of PCR primers amplifies a human nucleic acid sequence. In certain embodiments, the second set of PCR primers amplifies a nucleic acid sequence selected from GAPDH, ACTB, RPP30, and combinations thereof.In certain embodiments, the second set of PCR primers amplifies human RPP30. In certain embodiments, the second set of PCR primers comprises a mixture of primers with sequencing adapter sequences and primers without sequencing adapter sequences. In certain embodiments, the ratio of primers with sequencing adapter sequences to primers without sequencing adapter sequences is about 1:1, about 1:2, about 1:3, or about 1:4. In certain embodiments, the PCR comprises 30 to 45 amplification cycles. In certain embodiments, the PCR comprises 35 to 45 amplification cycles. In certain embodiments, the PCR comprises 39 to 42 amplification cycles. In certain embodiments, the second set of PCR primers comprises a nucleic acid sequence comprising a variable nucleotide sequence. In certain embodiments, the variable nucleotide sequence is a sample ID unique to the individual. In certain embodiments, the second set of PCR primers comprises an adapter sequence for a next-generation sequencing reaction. In certain embodiments, sequencing of the amplified nucleic acids from the biological sample uses next-generation sequencing technology. In certain embodiments, the method is capable of detecting fewer than 10 copies of a pathogen genome. In certain embodiments, the method is capable of detecting fewer than 5 copies of a pathogen genome. In certain embodiments, the pathogen genome is a coronavirus genome. In certain embodiments, the pathogen genome is a SARS-COV-2 genome. In certain embodiments, the method determines the strain of a coronavirus. In certain embodiments, the method determines the strain of SARS-COV-2.

[0008] Also described herein in another aspect is a synthetic nucleic acid comprising a 5'-proximal region, a 3'-proximal region, and an intervening nucleic acid sequence. In certain embodiments, the synthetic nucleic acid comprises RNA. In certain embodiments, the synthetic nucleic acid comprises DNA. In certain embodiments, the 5'-proximal region comprises a viral nucleic acid sequence. In certain embodiments, the viral nucleic acid sequence comprises a coronavirus sequence. In certain embodiments, the viral nucleic acid sequence comprises a SARS-COV-2 sequence. In certain embodiments, the 3'-proximal region comprises a viral nucleic acid sequence. In certain embodiments, the viral nucleic acid sequence comprises a coronavirus sequence. In certain embodiments, the viral nucleic acid sequence comprises a SARS-COV-2 sequence. In certain embodiments, the 5'-proximal region, the 3'-proximal region, or both the 5'-proximal region and the 3'-proximal region are less than about 30 nucleotides in length. In certain embodiments, the 5'-proximal region, the 3'-proximal region, or both the 5'-proximal region and the 3'-proximal region are less than about 25 nucleotides in length. In certain embodiments, the 5'-proximal region, the 3'-proximal region, or both the 5'-proximal region and the 3'-proximal region are less than about 20 nucleotides in length. In certain embodiments, the 5'-proximal region is at the 5'-end of the synthetic nucleic acid. In certain embodiments, the 3'-proximal region is at the 3'-end of the synthetic nucleic acid. In certain embodiments, the intervening nucleic acid sequence is less than about 99%, 98%, 97%, 95%, 90%, 85%, 80%, or 75% identical to a viral nucleic acid sequence. In certain embodiments, the synthetic nucleic acid sequence is a coronavirus sequence. In certain embodiments, the synthetic nucleic acid sequence is a SARS-COV-2 sequence. In certain embodiments, the synthetic nucleic acid is for use in a method of detecting a pathogen infection in the individual. In certain embodiments, the pathogen infection is a coronavirus infection. In certain embodiments, the viral infection is a SARS-COV-2 infection.

[0009] In another aspect, described herein are compositions comprising synthetic nucleic acid molecules comprising a first nucleic acid sequence and a second nucleic acid sequence, wherein (1) the first nucleic acid sequence is identical to a sequence from a pathogen nucleic acid molecule, and (2) the second nucleic acid sequence is not identical to the sequence from the pathogen nucleic acid molecule. In certain embodiments, the first nucleic acid sequence is located 3' of the second nucleic acid sequence. In certain embodiments, the synthetic nucleic acid molecule further comprises a third nucleic acid sequence, wherein the third nucleic acid sequence is identical to the second sequence from the pathogen nucleic acid molecule. In certain embodiments, the third nucleic acid sequence is located 5' of the second nucleic acid sequence. In certain embodiments, the first nucleic acid sequence or the third nucleic acid sequence is less than 5, 10, 15, 20, 25, or 30 nucleotides. In certain embodiments, the second nucleic acid sequence comprises a total number of nucleotides of less than 25, 50, 100, 150, 200, or 500 nucleotides. In certain embodiments, the second nucleic acid sequence comprises a total number of nucleotides greater than 25, 50, 100, 150, 200, or 500 nucleotides. In certain embodiments, the synthetic nucleic acid molecule is a ribonucleic acid (RNA) molecule, a deoxyribonucleic acid (DNA) molecule, or an RNA-DNA hybrid molecule. In certain embodiments, the first nucleic acid sequence, the third nucleic acid sequence, or both the first and third nucleic acid sequences comprise a primer binding site. In certain embodiments, the composition further comprises a pathogen nucleic acid molecule. In certain embodiments, the pathogen nucleic acid molecule is derived from a pathogen, and the pathogen comprises a bacterium, a virus, a fungus, or a combination thereof. In certain embodiments, the bacterium is derived from the genera Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof. In certain embodiments, the fungus is Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. In certain embodiments, the virus is a DNA virus.In certain embodiments, the DNA virus includes hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or variola, and combinations thereof. In certain embodiments, the virus is an RNA virus. In certain embodiments, the RNA virus includes influenza virus, coronavirus, poliovirus, measles virus, Ebola virus, retrovirus, or orthomyxovirus. In certain embodiments, the virus is a coronavirus. In certain embodiments, the coronavirus is SARS-COV-2. In certain embodiments, the composition further comprises a plurality of primers, wherein one primer of the plurality of primers is configured to hybridize to a sequence of the synthetic nucleic acid molecule or a sequence of the pathogen nucleic acid molecule. In certain embodiments, the sequence of the synthetic nucleic acid molecule and the sequence of the pathogen nucleic acid molecule are identical. In certain embodiments, the synthetic nucleic acid molecule is amplified with the same efficiency as the pathogen nucleic acid molecule. In certain embodiments, the synthetic nucleic acid molecule is configured to generate an amplification product that is the same size as, or within 10 base pairs of, the amplification product of the pathogen nucleic acid molecule.

[0010] In another aspect, also described herein is a method of diagnosing an individual with a pathogen infection, the method comprising: (a) providing a biological sample from the individual; (b) contacting the biological sample from the individual with a lysing agent to obtain a lysed biological sample; (c) performing a polymerase chain reaction (PCR) on the lysed biological sample to obtain a PCR-amplified lysed biological sample, wherein the PCR reaction on the lysed biological sample is performed using a first set of PCR primers, wherein the first set of PCR primers amplify a pathogen nucleic acid sequence; (d) sequencing the PCR-amplified lysed biological sample using next-generation sequencing; and (e) providing a positive diagnosis for the pathogen infection if a pathogen sequence is detected by the PCR or the sequencing, or providing a negative diagnosis for the individual if a pathogen sequence is not detected by the PCR or the sequencing.

[0011] In another aspect, a composition is described herein comprising a plurality of synthetic nucleic acids having different nucleic acid sequences, wherein the plurality of synthetic nucleic acids having different nucleic acid sequences comprises a common 5' sequence identical to a pathogen nucleic acid sequence, a common 3' sequence identical to a pathogen nucleic acid sequence, and an intervening sequence that differs among the plurality of sequences. In certain embodiments, the plurality of synthetic nucleic acids having different nucleic acid sequences are single-stranded. In certain embodiments, the plurality of synthetic nucleic acids having different nucleic acid sequences are double-stranded. In certain embodiments, the plurality of synthetic nucleic acids having different nucleic acid sequences consists of or comprises RNA. In certain embodiments, the plurality of synthetic nucleic acids having different nucleic acid sequences consists of or comprises DNA. In certain embodiments, the common 5' sequence identical to a pathogen nucleic acid sequence is 30 nucleotides or less. In certain embodiments, the common 3' sequence identical to a pathogen nucleic acid sequence is 30 nucleotides or less. In certain embodiments, the intervening sequence is 50 nucleotides or less. In certain embodiments, the intervening sequence is 30 nucleotides or less. In certain embodiments, the pathogen nucleic acid sequence is derived from a bacterial pathogen, a fungal pathogen, or a viral pathogen. In certain embodiments, the bacterial pathogen is of the genera Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof. In certain embodiments, the fungal pathogen is Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. In certain embodiments, the viral pathogen is a DNA virus. In certain embodiments, the DNA virus includes hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or variola, and combinations thereof. In certain embodiments, the viral pathogen is an RNA virus.In certain embodiments, the RNA virus comprises an influenza virus, a coronavirus, a poliovirus, a measles virus, an Ebola virus, a retrovirus, an orthomyxovirus, or a combination thereof. In certain embodiments, the viral pathogen is a coronavirus. In certain embodiments, the coronavirus is SARS-COV-2. In certain embodiments, the pathogen nucleic acid sequence is a nucleic acid sequence encoding a coronavirus spike protein. In certain embodiments, the plurality of synthetic nucleic acids having different nucleic acid sequences comprises or consists of a sequence selected from any one or more of S2_001, S2_002, S2_003, and S2_004. In certain embodiments, the composition is for use in a method of diagnosing or detecting infection by a pathogen. In certain embodiments, the composition is for use in a method of normalizing pathogen next-generation sequencing reads.

[0012] Provided herein are methods for detecting a coronavirus infection in an individual, the method comprising: (a) providing a biological sample from the individual, the biological sample comprising a coronavirus synthetic RNA, the sequence of the coronavirus synthetic RNA being different from a naturally occurring coronavirus nucleic acid sequence; (b) lysing the biological sample, thereby producing a lysed biological sample; (c) performing a reverse transcription reaction on the lysed biological sample to obtain a lysed and reverse transcribed biological sample; (d) performing an amplification reaction on the lysed and reverse transcribed biological sample to obtain an amplified biological sample, the amplification reaction on the lysed and reverse transcribed biological sample being performed with a set of coronavirus primers specific to a coronavirus nucleic acid sequence, the coronavirus primer set amplifying the coronavirus nucleic acid sequence and the coronavirus synthetic RNA; and (e) sequencing the amplified biological sample using next-generation sequencing. In certain embodiments, the method further comprises providing a positive diagnosis of coronavirus infection if sequence reads for the coronavirus nucleic acid sequence are detected.

[0013] In certain embodiments, the coronavirus infection is a SARS-Cov-2 infection. In certain embodiments, providing the positive diagnosis for the coronavirus infection or SARS-Cov-2 infection is provided when the ratio of the sequence reads from the coronavirus nucleic acid sequence to the sequence reads from the coronavirus synthetic RNA, or its mathematical equivalent, is greater than about 0.1. In certain embodiments, providing the positive diagnosis for a coronavirus infection is provided when the ratio of the sequence reads from the coronavirus nucleic acid sequence to the coronavirus synthetic RNA is greater than about 100. In certain embodiments, the lysed biological sample is not isolated or purified prior to performing the reverse transcription reaction. In certain embodiments, lysing the biological sample and performing the reverse transcription reaction on the lysed biological sample are performed in the same well, tube, or reaction vessel. In certain embodiments, lysing the biological sample comprises thermal lysis. In certain embodiments, the thermal lysis comprises heating the biological sample to a temperature of at least about 50°C. In certain embodiments, the biological sample from the individual comprises a plurality of coronavirus synthetic RNA sequences, wherein the plurality of coronavirus synthetic RNA sequences comprises at least two different synthetic coronavirus RNA sequences. In certain embodiments, the plurality of synthetic coronavirus RNA sequences comprises at least four different synthetic coronavirus RNA nucleic acid sequences. In certain embodiments, the synthetic RNA nucleic acid or the plurality of synthetic nucleic acids comprises about 20% to about 30% guanine nucleotides, about 20% to about 30% adenine nucleotides, about 20% to about 30% cytosine nucleotides, and about 20% to about 30% uracil nucleotides. In certain embodiments, the synthetic coronavirus RNA nucleic acid or the plurality of synthetic coronavirus RNA nucleic acids comprises approximately equal ratios of guanine to cytosine to adenine to uracil. In certain embodiments, the synthetic coronavirus RNA nucleic acid or the plurality of synthetic coronavirus RNA nucleic acids comprises a synthetic SARS-Cov-2 RNA nucleic acid or a plurality of synthetic SARS-Cov-2 RNA nucleic acids.In certain embodiments, the method further comprises detecting influenza A infection, influenza B infection, or a combination thereof. In certain embodiments, the amplification reaction on the lysed biological sample is performed using a set of influenza A primers specific for an influenza A nucleic acid sequence or a set of influenza B primers specific for an influenza B nucleic acid sequence. In certain embodiments, the set of influenza A primers specific for the influenza A nucleic acid sequence comprises the sequences set forth in SEQ ID NO:24 or 25 and SEQ ID NO:26, or SEQ ID NO:27 and SEQ ID NO:28. In certain embodiments, the set of influenza B primers specific for the influenza B nucleic acid sequence comprises the sequences set forth in SEQ ID NO:29 or 30. In certain embodiments, the amplification reaction on the lysed biological sample is performed using a set of influenza A primers specific for an influenza A nucleic acid sequence and a set of influenza B primers specific for an influenza B nucleic acid sequence. In certain embodiments, the biological sample from the individual further comprises influenza A synthetic RNA, influenza B synthetic RNA, or a combination thereof, wherein the influenza A synthetic RNA, the influenza B synthetic RNA, or the combination thereof is distinct from naturally occurring influenza A or influenza B nucleic acid sequences. In certain embodiments, the method further comprises providing a positive diagnosis for influenza A infection when a ratio of sequence reads from influenza A to sequence reads of the influenza A synthetic RNA, or a mathematical equivalent thereof, is greater than about 0.1. In certain embodiments, the method further comprises providing a positive diagnosis for influenza B infection when a ratio of sequence reads from the influenza A to sequence reads of the influenza B synthetic RNA, or a mathematical equivalent thereof, is greater than about 0.1. In certain embodiments, the coronavirus nucleic acid sequence is an N1 sequence, an S2 sequence, or a combination thereof. In certain embodiments, the coronavirus nucleic acid sequence is a SARS-Cov-2 N1 sequence, a SARS-Cov-2 S2 sequence, or a combination thereof.In certain embodiments, the coronavirus primer set comprises the sequences set forth in SEQ ID NO: 13 and SEQ ID NO: 14, or SEQ ID NO: 18 and SEQ ID NO: 19. In certain embodiments, the coronavirus primer set comprises the sequences set forth in SEQ ID NO: 20 and SEQ ID NO: 21, or SEQ ID NO: 22 and SEQ ID NO: 23. In certain embodiments, the coronavirus primer set is present at a concentration of about 50 nanomolar to about 250 nanomolar. In certain embodiments, the coronavirus primer set is present at a concentration of about 100 nanomolar. In certain embodiments, the influenza A nucleic acid sequence is an influenza A matrix sequence, an influenza A nonstructural protein 1 sequence, an influenza A hemagglutinin sequence, an influenza A neuraminidase sequence, an influenza A nucleoprotein sequence, or a combination thereof. In certain embodiments, the influenza B nucleic acid sequence is an influenza B matrix sequence, an influenza B nonstructural protein 1 sequence, an influenza B hemagglutinin sequence, an influenza B neuraminidase sequence, an influenza B nucleoprotein sequence, or a combination thereof. In certain embodiments, any one or more of the coronavirus primer set, the influenza A primer set, or the influenza B primer set comprises one or more index sequences that enable sample multiplexing. In certain embodiments, the coronavirus synthetic RNA comprises a nucleic acid sequence at least about 90% identical to any one or more of the sequences set forth in any one of SEQ ID NOs: 1-12. In certain embodiments, the coronavirus synthetic RNA comprises a plurality of different nucleic acid sequences that are at least about 90% identical to any one or more of the sequences set forth in any one of SEQ ID NOs: 1-12. In certain embodiments, the influenza A synthetic RNA comprises an RNA sequence at least about 90% identical to any one or more of the sequences set forth in any one of SEQ ID NOs: 31 or 32. In certain embodiments, the influenza B synthetic RNA comprises an RNA sequence at least about 90% identical to the sequence set forth in SEQ ID NO: 33.In certain embodiments, the coronavirus synthetic RNA comprises a plurality of four different nucleic acid sequences at least about 90% homologous to the four sequences set forth in SEQ ID NOs: 1-4. In certain embodiments, the coronavirus synthetic RNA, the influenza A synthetic RNA, or the influenza B synthetic RNA is present at a concentration of about 10 copies / reaction to about 500 copies / reaction. In certain embodiments, the coronavirus synthetic RNA is present at a concentration of about 10 copies / reaction to about 500 copies / reaction. In certain embodiments, the coronavirus synthetic RNA, the influenza A synthetic RNA, or the influenza B synthetic RNA is present at a concentration of about 200 copies / reaction. In certain embodiments, the coronavirus synthetic RNA is present at a concentration of about 200 copies / reaction. In certain embodiments, the biological sample comprises a nasal swab or saliva sample. In certain embodiments, the biological sample comprises less than about 10 microliters of saliva from an individual or less than about 10 microliters of buffer inoculated with a nasal swab from an individual. In certain embodiments, the biological sample comprises less than about 10 microliters of buffer solution inoculated with a nasal swab. In certain embodiments, the amplification reaction of the lysed biological sample is carried out using a primer pair specific to a sample control. In certain embodiments, the sample control is a housekeeping gene. In certain embodiments, the primer pair specific to the sample control is specific to RPP30. In certain embodiments, the primer pair specific to the sample control comprises the sequences set forth in SEQ ID NO: 15 or 16 and SEQ ID NO: 17.

[0014] Described herein is a synthetic nucleic acid comprising a 5'-proximal region comprising a first nucleotide sequence from a virus, a 3'-proximal region comprising a second nucleotide sequence from the virus, and an intervening nucleotide sequence, wherein the intervening nucleotide sequence comprises about 20% to about 30% guanine nucleotides, about 20% to about 30% adenine nucleotides, about 20% to about 30% cytosine nucleotides, and about 20% to about 30% uracil or thymidine nucleotides, wherein the intervening sequence differs from a naturally occurring sequence of the virus. In certain embodiments, the synthetic nucleic acid comprises DNA. In certain embodiments, the synthetic nucleic acid consists of DNA. In certain embodiments, the synthetic nucleic acid comprises RNA. In certain embodiments, the synthetic nucleic acid consists of RNA. In certain embodiments, the virus is influenza A virus, influenza B virus, or coronavirus. In certain embodiments, the virus is a coronavirus. In certain embodiments, the coronavirus is SARS-COV-2. In certain embodiments, the intervening nucleotide sequence nucleic acid comprises approximately equal ratios of guanine, cytosine, adenine, and uracil or thymidine nucleotides. In certain embodiments, the 3'-proximal region and the 5'-proximal region comprise a nucleotide sequence at least 90% homologous to a coronavirus S2 gene sequence. In certain embodiments, the 5'-proximal region and the 3'-proximal region comprise a nucleotide sequence at least 95% homologous to a coronavirus S2 gene sequence. In certain embodiments, the 5'-proximal region and the 3'-proximal region comprise a nucleotide sequence identical to a coronavirus S2 gene sequence. In certain embodiments, the 5'-proximal region and the 3'-proximal region comprise a nucleotide sequence at least 90% homologous to a coronavirus N1 gene sequence. In certain embodiments, the 5'-proximal region and the 3'-proximal region comprise a nucleotide sequence at least 95% homologous to a coronavirus N1 gene sequence. In certain embodiments, the 5'-proximal region and the 3'-proximal region comprise a nucleotide sequence identical to a coronavirus N1 gene sequence. In certain embodiments, the coronavirus N1 gene sequence or the coronavirus S2 gene sequence is a SARS-CoV-2 gene sequence.In certain embodiments, the sequence nucleic acid comprises a sequence at least 90% homologous to any one or more of the sequences set forth in any one of SEQ ID NOs: 1-12. In certain embodiments, the sequence nucleic acid comprises a sequence at least 95% homologous to any one or more of the sequences set forth in any one of SEQ ID NOs: 1-12. In certain embodiments, the sequence nucleic acid comprises a sequence identical to any one or more of the sequences set forth in any one of SEQ ID NOs: 1-12. In certain embodiments, a plurality of synthetic nucleic acids is described herein, the plurality comprising synthetic nucleic acids comprising at least two different nucleotide sequences. In certain embodiments, the plurality of synthetic nucleic acids comprises synthetic nucleic acids comprising at least two different nucleotide sequences. In certain embodiments, the plurality of synthetic nucleic acids comprises at least four different nucleotide sequences. In certain embodiments, the four different nucleotide sequences are those set forth in SEQ ID NOs: 1-4. In certain embodiments, the four different nucleotide sequences are selected from those set forth in SEQ ID NOs: 5-12.

[0015] Described herein is a reaction mixture for determining the presence or absence of viral nucleic acid in a biological sample, the reaction mixture comprising a synthetic nucleic acid or a plurality of synthetic nucleic acids described herein, at least a portion of the biological sample, and one or more enzymes or reagents sufficient to amplify the viral nucleic acid in the biological sample, if present. In certain embodiments, the biological sample is a human biological sample. In certain embodiments, the biological sample comprises saliva, a buccal swab, a nasopharyngeal swab, or a middle turbinate swab. In certain embodiments, the biological sample comprises saliva or a nasopharyngeal swab. In certain embodiments, the viral nucleic acid is influenza A nucleic acid, influenza B nucleic acid, or coronavirus nucleic acid. In certain embodiments, the coronavirus nucleic acid is Sars-Cov-2 nucleic acid. In certain embodiments, the one or more reagents are selected from the list consisting of reverse transcriptase, dNTPs, a primer pair specific to the viral nucleotide sequence, a primer pair specific to a sample control nucleotide sequence, a magnesium salt, and combinations thereof. In certain embodiments, the primer pair specific to the sample control nucleotide sequence is specific to a human nucleotide sequence. In certain embodiments, the primer pair specific to the sample control nucleotide sequence is specific to a housekeeping gene. In certain embodiments, the primer pair specific to the sample control is specific to RPP30. In certain embodiments, the primer pair specific to the sample control comprises the sequences set forth in SEQ ID NO: 15 or 16 and SEQ ID NO: 17. In certain embodiments, the primer pair specific to the viral nucleotide sequence is specific to an influenza A nucleotide sequence, an influenza B nucleotide sequence, and a coronavirus nucleotide sequence. In certain embodiments, the primer pair specific to the viral nucleotide sequence is specific to a coronavirus S1 or N2 sequence. In certain embodiments, the coronavirus S1 or N2 sequence is a coronavirus S1 or N2 nucleic acid sequence.In certain embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence comprises the sequence set forth in any one of SEQ ID NOS: 13-30 or 100-605. In certain embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence comprises the sequence set forth in SEQ ID NOS: 13 and 14, 18 and 19, 20 and 21, 24 or 25 and 26, or 29 and 30. In certain embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence is present at a concentration of about 50 micromolar to about 250 micromolar. In certain embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence is present at a concentration of about 100 micromolar. In certain embodiments, the primer pair specific to the viral nucleotide sequence or the primer pair specific to the sample control nucleotide sequence is present at a concentration of about 200 micromolar. In certain embodiments, the coronavirus synthetic RNA is present at a concentration of about 10 copies / reaction to about 500 copies / reaction mixture. In certain embodiments, the coronavirus synthetic RNA is present at a concentration of about 200 copies / reaction mixture. In certain embodiments, the volume of the reaction mixture is about 10 microliters to about 100 microliters. In certain embodiments, the volume of the reaction mixture is about 20 microliters.

[0016] Also described herein are kits for determining the presence or absence of viral nucleic acid in a biological sample, the kits comprising a synthetic nucleic acid described herein or a plurality of synthetic nucleic acids described herein and one or more enzymes or reagents sufficient to amplify the viral nucleic acid from the biological sample. In certain embodiments, the viral nucleic acid is influenza A nucleic acid, influenza B nucleic acid, or coronavirus nucleic acid. In certain embodiments, the coronavirus nucleic acid is coronavirus nucleic acid. In certain embodiments, the one or more reagents are selected from the list consisting of reverse transcriptase, dNTPs, a primer pair specific for the viral nucleotide sequence, a primer pair specific for a sample control nucleotide sequence, magnesium salts, and combinations thereof. In certain embodiments, the primer pair specific for the sample control nucleotide sequence is specific for a human nucleotide sequence. In certain embodiments, the primer pair specific for the sample control nucleotide sequence is specific for a housekeeping gene. In certain embodiments, the primer pair specific for the sample control is specific for RPP30. In certain embodiments, the primer pair specific for the sample control comprises the sequences set forth in SEQ ID NO: 15 or 16 and SEQ ID NO: 17. In certain embodiments, the primer pair specific to the viral nucleotide sequence is specific to an influenza A nucleotide sequence, an influenza B nucleotide sequence, and a coronavirus nucleotide sequence. In certain embodiments, the primer pair specific to the viral nucleotide sequence is specific to a coronavirus S1 or N2 sequence. In certain embodiments, the coronavirus S1 or N2 sequence is a coronavirus S1 or N2 nucleic acid sequence. In certain embodiments, the primer pair specific to the viral nucleotide sequence or the primer pair specific to the sample control nucleotide sequence comprises a sequence set forth in any one of SEQ ID NOs: 13-30 or 100-605.In certain embodiments, the primer pair specific to the viral nucleotide sequence or the primer pair specific to the sample control nucleotide sequence comprises the sequences set forth in SEQ ID NO: 13 and SEQ ID NO: 14, SEQ ID NO: 18 and SEQ ID NO: 19, SEQ ID NO: 20 and SEQ ID NO: 21, SEQ ID NO: 24 or 25 and SEQ ID NO: 26, SEQ ID NO: 29 and SEQ ID NO: 30. [Brief explanation of the drawings]

[0017] The novel features described herein are set forth with particularity in the appended claims. A better understanding of the features and advantages thereof will be obtained by reference to the following detailed description that sets forth illustrative examples in which the principles of the features described herein are utilized, and the accompanying drawings in which: [Figure 1] 1 illustrates an exemplary schematic for diagnosing a viral infection. [Figure 2A] Illustrative data demonstrating detection of SARS-CoV-2 nucleic acid in a sample are provided. [Figure 2B] Illustrative data demonstrating detection of SARS-CoV-2 nucleic acid in a sample are provided. [Figure 3] 1 illustrates an exemplary priming scheme according to the present disclosure. [Figure 4] Illustrates the SwabSeq diagnostic testing platform for COVID-19. [Figure 5] Validation of clinical specimens is demonstrated, demonstrating limits of detection comparable to highly sensitive RT-qPCR reactions. [Figure 6] 1 shows the sequencing library design. [Figure 7] We demonstrate that the S2 primers exhibit comparable PCR efficiency in amplifying the COVID-19 amplicon and the synthetic S2 spike. [Figure 8] We demonstrate that SwabSeq maintains linearity at very high virus concentrations. [Figure 9] We show that sequencing performed on a MiSeq or NextSeq Machine exhibits similar sensitivity. [Figure 10] Preliminary and confirmatory detection limit data are shown for RNA purified samples using NextSeq550. [Figure 11] We demonstrate that traditional collection media and buffer-free extraction-free protocols require dilution to overcome the effects of RT and PCR inhibition. [Figure 12] We present an exemplary development of a lightweight sample receiving, collection, and processing system that allows scalable testing to thousands of samples per day. [Figure 13] Preheating saliva to 95°C for 30 minutes has been shown to improve RT-PCR. [Figure 14] This indicates that PCR inhibition has a significant effect on the amplification product. [Figure 15] We show that increasing the number of PCR cycles and running the tapestation on unpurified or inhibited sample types (e.g., saliva) are seen to increase the size of nonspecific peaks in library preparations. [Figure 16] TaqPath reduces the number of S2 reads in SARS-CoV2-negative samples compared to NEB Luna. [Figure 17] This indicates that carryover contamination from template lines in MiSeq contributes to cross-contamination. [Figure 18] Sequencing errors and potential amplicon mis-assignments in the amplicon reads are shown. [Figure 19] 1 shows a visualization of different indexing strategies. [Figure 20] 1 shows a computational correction for index misassignment using a mixed model. [Figure 21] We demonstrate quantifying the role of index misassignment as a source of noise in the S2 lead. [Figure 22]The effect of reducing primer concentration on primer dimers and non-specific amplification products is shown. [Figure 23] Diversified synthetic nucleic acid spike-in sequences are shown. [Figure 24] Data obtained with N1 spike-in (top) or S2 spike-in (bottom) and detection of the different amplicons N1 (top) or S2 (bottom) are shown. [Figure 25] We demonstrate that combining N1 and S2 can increase the sensitivity of SARS-CoV2 detection compared to either antibody alone. [Figure 26] 1 shows the results of detection using different volumes of saliva samples. [Figure 27] 4 shows the results of detection using different volumes of nasal swab samples. [Figure 28] Showing that SwabSeq can detect influenza A or influenza B. [Figure 29] 1 shows an exemplary algorithm for calling positive and negative samples. DETAILED DESCRIPTION OF THE INVENTION

[0018] The incidence of highly contagious or virulent diseases is increasing. One-third of global deaths are attributable to infectious diseases, making them the second leading cause of death and disability worldwide. Obtaining rapid and accurate readouts to identify and diagnose infectious diseases that cause human illness is a critical component of diagnostic medicine, particularly in the case of viral infections, for implementing effective public health responses and improving the delivery of human healthcare. Numerous methods for detecting viral infections have been developed for clinical diagnostic purposes, but most of these tests do not provide sufficiently rapid or high-throughput readouts and / or are not feasible due to the resource burden associated with rapidly testing a rapidly increasing number of subjects.

[0019] Additionally, methods and systems for diagnosing individuals with viral infections are provided herein. These methods and systems, as described herein, utilize polymerase chain reaction (PCR), library preparation strategies, and next-generation sequencing to provide readouts or information that accurately identify viral infections and can be effectively scaled to meet the challenges associated with the need for efficient testing of an increasing number of subjects. To facilitate the identification and diagnosis of viral infections in samples, the methods provide advances or solutions in (1) effective and efficient isolation of nucleic acid sequences from viruses, (2) effective and efficient processing of nucleic acid molecules corresponding to or derived from viruses, and (3) effective and efficient multiplexing of samples, allowing multiple samples to be tested in parallel, reducing the overall resource burden of testing.

[0020] Described herein are methods for detecting pathogen genomes in a sample, comprising amplifying nucleic acids from the sample using a first set of PCR primers to obtain amplified nucleic acids, wherein the first set of PCR primers amplifies a viral nucleic acid sequence and a synthetic nucleic acid sequence. The synthetic nucleic acid sequence provides a sample and amplification control in the assay, allowing for a lower limit of detection compared to amplification without synthetic nucleic acids. The synthetic nucleic acid can be added ("spiked") to a lysate contacted with a biological sample, or the synthetic nucleic acid can be present in the lysis buffer before contacting the lysis buffer with the biological sample. The sample can be a biological sample. The biological sample can be obtained from an individual. The synthetic nucleic acid can be added at any step along the way. More than one synthetic nucleic acid can be used in the methods described herein. For example, one synthetic nucleic acid can serve as a control for one or more viral nucleic acids, one synthetic nucleic acid can serve as a control (e.g., housekeeping control) for one or more human nucleic acids, or, for example, multiple synthetic nucleic acids can be added to serve as controls for multiple pathogen sequences.

[0021] Further described herein are synthetic nucleic acids comprising a 5'-proximal region, a 3'-proximal region, and an intervening nucleic acid sequence. The synthetic nucleic acid can include any sequence at the 5'- and 3'-ends that allows amplification of the pathogen sequence, provided that the sequence between the 5'- and 3'-ends is distinguishable by sequencing. In certain embodiments, the sequence between the 5'- and 3'-ends is identical to the pathogen sequence, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides that differ from the pathogen sequence.

[0022] Further described herein are synthetic nucleic acids comprising a 5'-proximal region, a 3'-proximal region, and an intervening nucleic acid sequence. The synthetic nucleic acid can include any sequence at the 5'- and 3'-ends that allows amplification of the pathogen sequence, provided that the sequence between the 5'- and 3'-ends is distinguishable by sequencing. In certain embodiments, the sequence between the 5'- and 3'-ends is identical to the pathogen sequence, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides that differ from the pathogen sequence.

[0023] Also described herein are methods for detecting coronavirus infection in an individual, the methods comprising: (a) providing a biological sample from the individual, the biological sample comprising coronavirus synthetic RNA, the sequence of the coronavirus synthetic RNA being different from a naturally occurring coronavirus nucleic acid sequence; (b) lysing the biological sample, thereby producing a lysed biological sample; (c) performing a reverse transcription reaction on the lysed biological sample to obtain a lysed and reverse-transcribed biological sample; (d) performing an amplification reaction on the lysed and reverse-transcribed biological sample to obtain an amplified biological sample, the amplification reaction on the lysed and reverse-transcribed biological sample being performed using a set of coronavirus primers specific to a coronavirus nucleic acid sequence, the coronavirus primer set amplifying the coronavirus nucleic acid sequence and the coronavirus synthetic RNA; and (e) sequencing the amplified biological sample using next-generation sequencing. In certain embodiments, the methods further comprise providing a positive diagnosis of coronavirus infection if sequence reads for the coronavirus nucleic acid sequence are detected. In certain embodiments, the coronavirus infection is SARS-COV-2 infection. In certain embodiments, providing the positive diagnosis for the coronavirus infection or SARS-Cov-2 infection is provided when the ratio of the sequence reads from the coronavirus nucleic acid sequence to the sequence reads for the coronavirus synthetic RNA, or its mathematical equivalent, is greater than a ratio of about 0.1. In certain embodiments, providing the positive diagnosis for a coronavirus infection is provided when the ratio of the sequence reads from the coronavirus nucleic acid sequence to the sequence reads for the coronavirus synthetic RNA is greater than about 100.

[0024] Also described herein are methods for detecting a coronavirus infection in an individual, the methods comprising: (a) providing a biological sample from the individual, the biological sample comprising a coronavirus synthetic RNA, the sequence of the coronavirus synthetic RNA being different from a naturally occurring coronavirus nucleic acid sequence; (b) lysing the biological sample, thereby producing a lysed biological sample; (c) performing a reverse transcription reaction on the lysed biological sample to obtain a lysed and reverse transcribed biological sample; and (d) performing an amplification reaction on the lysed and reverse transcribed biological sample to obtain an amplified biological sample, the amplification reaction on the lysed and reverse transcribed biological sample being performed using a set of coronavirus primers specific to a coronavirus nucleic acid sequence, the coronavirus primer set amplifying the coronavirus nucleic acid sequence and the coronavirus synthetic RNA. In certain embodiments, the coronavirus infection is a SARS-COV-2 infection.

[0025] Also described herein are methods for detecting a coronavirus infection in an individual, the methods comprising: (a) providing a biological sample from the individual, the biological sample comprising a coronavirus synthetic RNA, the sequence of the coronavirus synthetic RNA being different from a naturally occurring coronavirus nucleic acid sequence; (b) lysing the biological sample, thereby producing a lysed biological sample; and (c) performing a reverse transcription reaction on the lysed biological sample to obtain a lysed reverse transcribed biological sample. In certain embodiments, the coronavirus infection is a SARS-COV-2 infection.

[0026] Compositions containing synthetic nucleic acid molecules are also provided and are useful in the methods disclosed herein. Synthetic nucleic acids are generally in vitro transcribed or synthetic control RNAs that are identical to the viral sequence targeted for amplification, except for a short modified stretch that allows sequencing reads corresponding to the synthetic control to be distinguished from those corresponding to the pathogen sequence. For example, compositions are disclosed that include synthetic nucleic acid molecules containing a first nucleic acid sequence and a second nucleic acid sequence, where the first nucleic acid sequence is identical to a sequence from a pathogen nucleic acid molecule and the second nucleic acid sequence is not identical to the sequence from the pathogen nucleic acid molecule. In some embodiments, the first nucleic acid sequence is located in the 3' region of the synthetic nucleic acid molecule. In some embodiments, the synthetic nucleic acid molecule further includes a third nucleic acid sequence, where the first nucleic acid sequence is identical to the first sequence from a pathogen nucleic acid molecule and the third nucleic acid sequence is identical to the second sequence from the pathogen nucleic acid molecule. In some embodiments, the first nucleic acid sequence is located in the 3' region of the second nucleic acid molecule, and the third nucleic acid sequence is located in the 5' region of the second nucleic acid molecule. In some embodiments, the second nucleic acid sequence is less than 5, 10, 15, 20, 25, or 30 nucleotides. In some embodiments, the second nucleic acid synthetic nucleic acid molecule contains a total number of nucleotides less than 25, 50, 100, 150, 200, or 500 nucleotides. In some embodiments, the second nucleic acid synthetic nucleic acid molecule contains a total number of nucleotides greater than 25, 50, 100, 150, 200, or 500 nucleotides. In some embodiments, the synthetic nucleic acid molecule is a ribonucleic acid (RNA) molecule, a deoxyribonucleic acid (DNA) molecule, or an RNA-DNA hybrid molecule. In certain embodiments, the synthetic nucleic acid molecule is PCR amplified with an efficiency within about 10% of the corresponding pathogen sequence. In certain embodiments, the synthetic nucleic acid molecule is PCR amplified with an efficiency within about 5% of the corresponding pathogen sequence. In certain embodiments, the synthetic nucleic acid molecules are PCR amplified with the same efficiency as the corresponding pathogen sequences.

[0027] The use of multiple synthetic nucleic acids in accordance with the methods described herein can further increase sensitivity, reduce false positives, or improve the accuracy and / or precision of nucleic acid sequence quantification. The multiple synthetic nucleic acids may contain two, three, four, five, six, seven, eight, nine, or more distinct sequences that can be co-amplified with a set of primers specific to the pathogen sequence to be detected. The multiple synthetic nucleic acids may possess specific characteristics that are desirable for the multiple synthetic nucleic acids. In certain embodiments, the melting temperatures of the distinct sequences in the multiple synthetic nucleic acids may be substantially the same or may be within about 0.5°, 1°, 2°, 3°, 4°, or 5° Celsius of the average melting temperature of the multiple synthetic nucleic acids. In certain embodiments, the nucleotide composition of the targeted multiple synthetic nucleic acids is about 30% to about 20% A, about 30% to about 20% G, about 30% to about 20% C, and about 30% to about 20% T. In certain embodiments, the nucleotide composition of the targeted synthetic nucleic acids is about 25% A, about 25% G, about 25% C, and about 25% T. In certain embodiments, the nucleotide composition of the targeted synthetic nucleic acids is one or more of about 30% to about 20% A, about 30% to about 20% G, about 30% to about 20% C, and about 30% to about 20% T. In certain embodiments, the nucleotide composition of the targeted synthetic nucleic acids is one or more of about 25% A, about 25% G, about 25% C, and about 25% T. In certain embodiments, the synthetic nucleic acids are selected or designed to minimize secondary structure or dimerization between different sequences of the synthetic nucleic acids.

[0028] Also described herein is a method of diagnosing an individual with a viral infection, the method comprising the steps of: (a) providing a biological sample from the individual; (b) contacting the biological sample from the individual with a lysing agent to obtain a lysed biological sample; (c) performing a reverse transcription reaction on the lysed biological sample; and (d) performing a polymerase chain reaction (PCR) on the lysed biological sample to obtain a PCR-amplified lysed biological sample, wherein the PCR reaction on the lysed biological sample is performed using a first set of PCR primers and a second set of PCR primers. (e) sequencing the PCR-amplified lysed biological sample using next-generation sequencing, wherein the first set of PCR primers amplifies a viral nucleic acid sequence and the second set of PCR primers amplifies a genomic sequence of the species to which the individual belongs; and (f) optionally providing a positive diagnosis for the viral infection if a viral sequence is detected by the PCR or the sequencing, or a negative diagnosis for the human individual if a viral sequence is not detected by the PCR or the sequencing. In certain embodiments, the first set of PCR primers is capable of amplifying a synthetic nucleic acid sequence present in the lysed biological sample. The synthetic nucleic acid sequence contains the same primer binding sites as the viral nucleic acid sequence, except for a different intervening nucleic acid sequence, so that it can be distinguished by sequencing. In certain embodiments, the synthetic nucleic acid sequence is RNA.

[0029] Also described herein is a method of diagnosing an individual with a pathogen infection, the method comprising: (a) providing a biological sample from the individual; (b) contacting the biological sample from the individual with a lysing agent to obtain a lysed biological sample; (c) performing a polymerase chain reaction (PCR) on the lysed biological sample to obtain a PCR-amplified lysed biological sample, wherein the PCR reaction on the lysed biological sample is performed using a first set of PCR primers and a second set of PCR primers, wherein the first set of PCR primers amplify pathogen nucleic acid sequences and the second set of primers amplify nucleic acid sequences of the individual; (d) sequencing the PCR-amplified lysed biological sample using next-generation sequencing; and (e) optionally, providing a positive diagnosis for the pathogen infection if a pathogen sequence is detected by the PCR or the sequencing, or providing a negative diagnosis for the human individual if a pathogen sequence is not detected by the PCR or the sequencing. In certain embodiments, the first set of PCR primers can amplify a synthetic nucleic acid sequence present in a lysed biological sample. The synthetic nucleic acid sequence contains the same primer binding site as the viral nucleic acid sequence, except that the intervening nucleic acid sequence is different so that the synthetic nucleic acid sequence can be distinguished by sequencing. In certain embodiments, the synthetic nucleic acid sequence is RNA.

[0030] The methods described herein can be used to monitor for the presence of pathogens in any number of samples, including non-human samples. Monitoring can include monitoring livestock or wild animal herds, wildlife populations, or captive live animals (e.g., zoos, wildlife parks, or live animal markets where animals are sold for food or as pets).

[0031] Also described herein is a method for monitoring the presence of a pathogen in a sample, the method comprising: (a) providing a sample; (b) contacting the sample with the sample extractant to obtain an extracted sample; (c) performing a polymerase chain reaction (PCR) on the extracted sample to obtain a PCR-amplified extracted sample, wherein the PCR reaction on the extracted sample is performed using a first set of PCR primers; (d) sequencing the PCR-amplified lysed biological sample using next-generation sequencing; and (e) optionally providing a positive readout for the pathogen if a pathogen sequence is detected by the PCR or sequencing, or a negative readout if a pathogen sequence is not detected by the PCR or sequencing. In certain embodiments, the first set of PCR primers can amplify a synthetic nucleic acid sequence present in the extracted biological sample. The synthetic nucleic acid sequence contains the same primer binding sites as the pathogen nucleic acid sequence, except for a different intervening nucleic acid sequence, so that the synthetic nucleic acid sequence can be identified by sequencing. In certain embodiments, the synthetic nucleic acid sequence is RNA. In certain embodiments, the synthetic nucleic acid sequence is DNA.

[0032] To achieve identification, detection, and / or diagnosis of a viral infection, the methods and systems disclosed herein include the steps of: (a) providing a biological sample from the individual; (b) contacting the biological sample from the individual with a lysing agent to obtain a lysed biological sample; (c) performing an initial nucleic acid extension reaction on the lysed biological sample; and (d) performing a polymerase chain reaction (PCR) on the lysed biological sample to obtain a PCR-amplified lysed biological sample, wherein the PCR reaction on the lysed biological sample is performed using a first set of PCR primers and a second set of PCR primers. (e) sequencing the PCR-amplified lysed biological sample using next-generation sequencing; and (f) optionally, providing a positive diagnosis for the viral infection if a viral sequence or a derivative thereof is detected by the PCR or the sequencing, or providing a negative diagnosis for the human individual if a coronavirus sequence is not detected by the PCR or the sequencing.

[0033] Also disclosed herein is a method for nucleic acid processing for detecting a viral infection, which includes the steps of: (a) providing a sample containing a viral nucleic acid molecule and a host nucleic acid molecule; (b) performing a nucleic acid extension reaction on the viral nucleic acid molecule using a first primer containing a barcode sequence to generate a barcoded viral nucleic acid molecule; (c) performing a nucleic acid extension reaction on the host nucleic acid molecule using a second primer containing the barcode sequence to generate a barcoded host nucleic acid molecule; (d) sequencing the barcoded viral nucleic acid molecule and the barcoded host nucleic acid molecule to identify (i) the barcode sequence and (ii) a sequence corresponding to the viral nucleic acid molecule or its derivative, and the host nucleic acid molecule; and (e) providing a positive diagnosis of a viral infection if the sequence corresponding to the viral nucleic acid molecule is identified in (d).

[0034] As used herein, the term "barcode" generally refers to a label or identifier that conveys or can convey information about an analyte. A barcode can be part of the analyte. A barcode can be independent of the analyte. A barcode can be a tag attached to the analyte (e.g., a nucleic acid molecule) or a combination of a tag in addition to an intrinsic characteristic of the analyte (e.g., the size or terminal sequence of the analyte). A barcode can be specific. Barcodes can have a variety of different formats. For example, barcodes can include polynucleotide barcodes, random nucleic acid and / or amino acid sequences, and synthetic nucleic acid and / or amino acid sequences. Barcodes can be attached to analytes in a reversible or irreversible manner. For example, barcodes can be added to fragments of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) samples before, during, and after sequencing of the sample. Barcodes can enable identification and / or quantification of individual sequencing reads.

[0035] As used herein, the term "real-time" can refer to a response time of less than about one second, less than one tenth of a second, less than one hundredth of a second, less than one millisecond, or less. Response times may be greater than one second. In some examples, real-time can refer to simultaneous or substantially simultaneous processing, detection, or identification.

[0036] As used herein, the term "genome" refers to genomic information from a plant, animal, bacterium, fungus, or virus, which may be, for example, at least a portion or all of a subject's genetic information. A genome can be encoded in either DNA or RNA. A genome can include coding regions (e.g., regions that encode proteins) as well as non-coding regions. A genome can include the sequences of all chromosomes in an organism. For example, a human genome typically has a total of 46 chromosomes. All of these sequences can make up the human genome.

[0037] The terms "adaptor," "adapter," and "tag" may be used interchangeably. An adapter or tag can be attached to the polynucleotide sequence to be "tagged" by any approach, including ligation, hybridization, or other approaches.

[0038] As used herein, the term "sequencing" generally refers to methods and techniques for determining the sequence of nucleotide bases in one or more polynucleotides. Polynucleotides can be nucleic acid molecules, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single-stranded DNA). "Next-generation sequencing" refers to high-throughput sequencing methods that are not Sanger sequencing. Sequencing can be performed using a variety of currently available systems, including, but not limited to, sequencing systems from Illumina® (e.g., iSeq 100, MiniSeq, MiSeq, or NextSeq series machines), Pacific Biosciences (PacBio®), Oxford Nanopore®, or Life Technologies (Ion Torrent®). Alternatively or additionally, sequencing can be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real-time PCR), or isothermal amplification. Such a system may provide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., a human) generated by the system from a sample provided by the subject. In some examples, such a system provides sequencing reads (also referred to herein as "reads"). A read may include a string of nucleic acid bases corresponding to the sequence of a sequenced nucleic acid molecule. In some situations, the systems and methods provided herein may be used with proteomic information.

[0039] As used herein, the term "bead" generally refers to a particle. Beads may be solid or semi-solid particles. Beads may be gel beads. Gel beads may include a polymer matrix (e.g., a matrix formed by polymerization or cross-linking). The polymer matrix may include one or more polymers (e.g., polymers with various functional groups or repeating units). The polymers in the polymer matrix may be randomly arranged, e.g., in a random copolymer, and / or may have a regular structure, e.g., in a block copolymer. Cross-linking can be via covalent, ionic, or inductive bonds, interactions, or physical entanglements. Beads may be polymeric. Beads may be formed from nucleic acid molecules linked together. Beads may be formed through the covalent or non-covalent assembly of molecules (e.g., macromolecules), such as monomers or polymers. Such polymers or monomers may be natural or synthetic. Such polymers or monomers may be, for example, nucleic acid molecules (e.g., DNA or RNA) or may include such nucleic acid molecules. Beads may be formed of polymeric materials. Beads may be magnetic or non-magnetic. Beads may be rigid. The beads may be flexible and / or compressible. The beads may be frangible or dissolvable. The beads may be solid particles (e.g., metal-based particles, including but not limited to, iron oxide, gold, or silver) covered with a coating comprising one or more polymers. Such coatings may be frangible or dissolvable.

[0040] As used herein, the term "sample" is used broadly and can refer to an environmental sample (e.g., a water sample, a sewage sample), a raw or cooked food sample, a sample generated from a population of non-human individuals (e.g., wild or domestic animals), or a biological sample from a human or non-human animal. A sample may contain any number of macromolecules, for example, cellular macromolecules. A sample may be a cell sample. A sample may be a cell line or cell culture sample. A sample may include one or more cells. A sample may include one or more microorganisms. A biological sample may be a nucleic acid sample or a protein sample. A biological sample may also be a carbohydrate sample or a lipid sample. A biological sample may be derived from another sample. A sample may be a tissue sample, such as a biopsy, core biopsy, needle aspirate, or fine needle aspirate. A sample may be a fluid sample, such as a blood sample, urine sample, or saliva sample. A sample may be a skin sample. A sample may be a buccal swab. A sample may be a nasopharyngeal swab. A sample may be a plasma or serum sample. The sample may be acellular or acellular. The acellular sample may contain extracellular polynucleotides. The extracellular polynucleotides may be isolated from a bodily sample, which may be selected from the group consisting of blood, plasma, serum, urine, saliva, mucosal discharge, saliva, feces, and tears.

[0041] The term "biological particle," as used herein, generally refers to an individual biological system derived from a biological sample. A biological particle may be a macromolecule. A biological particle may be a small molecule. A biological particle may be a virus. A biological particle may be a cell or a derivative of a cell. A biological particle may be an organelle. A biological particle may be a rare cell from a population of cells. A biological particle may be any type of cell, including, but not limited to, a prokaryotic cell, a host cell, a bacterium, a fungus, a plant, a mammalian cell, or other animal cell type, mycoplasma, a normal tissue cell, a tumor cell, or other cell type, whether from a unicellular or multicellular organism. A biological particle may be a component of a cell. A biological particle may be or include DNA, RNA, an organelle, a protein, or any combination thereof. A biological particle may be or include a cell, or one or more components of a cell (e.g., a cell bead), such as a matrix (e.g., a gel or polymer matrix) containing DNA, RNA, an organelle, a protein, or any combination thereof from a cell. A biological particle may be obtained from a tissue of a subject. A biological particle may be a hardened cell. Such hardened cells may or may not include a cell wall or cell membrane. Biological particles may include one or more components of a cell, but may not include other components of a cell. Examples of such components are the nucleus or organelles. The cells may be living cells. Living cells may be culturable, for example, when encapsulated in a gel or polymer matrix or when containing a gel or polymer matrix.

[0042] As used herein, the term "pathogen" includes any organism or virus capable of causing disease in a population of individuals, which may include animals or plants. The term also encompasses pathogens present in a carrier individual or species, which do not cause disease in the carrier individual or species but can be transmitted to another individual or species and cause disease. As used herein, pathogens include, but are not limited to, bacteria, protozoa, fungi, nematodes, viroids, and viruses, or any combination thereof, each of which, by itself or in combination with another pathogen, can induce disease in vertebrates, including, but not limited to, mammals and humans. As used herein, the term "host" refers to an organism that can be infected by a pathogen, and includes plants, animals, vertebrates, mammals, rodents, cattle, horses, pigs, poultry, birds, geese, ducks, fish, crustaceans, etc.

[0043] As used herein, "bacteria" or "eubacteria" refers to the domain of prokaryotes. Bacteria include at least 11 different groups, including: (1) Gram-positive bacteria (gram+), which are divided into two major subdivisions: (i) the high G+C group (e.g., Actinomycetes, Mycobacteria, Micrococcus) and (ii) the low G+C group (e.g., Bacillus, Clostridia, Lactobacillus, Staphylococcus, Streptococcus, Mycoplasma); (2) Proteobacteria, e.g., purple photosynthetic + non-photosynthetic Gram-negative bacteria (including most "common" Gram-negative bacteria); (3) Cyanobacteria, e.g., oxygenic photosynthetic organisms; (4) Spirochetes and related species; (5) Planctomycetes; (6) Bacteroidetes; (7) Chlamydia; (8) Green sulfur bacteria; (9) Green non-sulfur bacteria (including anaerobic photosynthetic organisms); (10) Radioresistant Micrococcus and related species; and (11) Thermophiles of Thermotoga and Thermosipho. "Gram-negative bacteria" include cocci, non-enteric bacilli, and enteric bacilli. Examples of genera of Gram-negative bacteria include, for example, Neisseria, Spirillum, Pasteurella, Brucella, Yersinia, Francisella, Haemophilus, Bordetella, Escherichia, Salmonella, Shigella, Klebsiella, Proteus, Vibrio, Pseudomonas, Bacteroides, Acetobacteria, Aerobes, Agrobacterium, Azotobacter, Spirillum, Serratia, Vibrio, Rhizobium, Chlamydia, Rickettsia, Treponema, and Fusobacterium. "Gram-positive bacteria" include cocci, non-spore-forming bacilli, and spore-forming bacilli. Genera of Gram-positive bacteria include, for example, Actinomyces, Bacillus, Clostridium, Corynebacterium, Erysipelothrix, Lactobacillus, Listeria, Mycobacterium, Myxococcus, Nocardia, Staphylococcus, Streptococcus, and Streptomyces. "Pathogenic bacteria" or "pathogenic bacterium" are bacterial species that cause disease in another host organism (e.g., animals and plants) by directly infecting the other host organism or by producing a substance (e.g., bacteria that produce pathogenic toxins, etc.) that causes disease in the other organism.

[0044] As used herein, the term "macromolecular component" generally refers to a macromolecule contained within or derived from a bioparticle. The macromolecular component may include nucleic acids. In some cases, the bioparticle may be a macromolecule. The macromolecular component may include DNA. The macromolecular component may include RNA. The RNA may be coding or non-coding. The RNA may be, for example, messenger RNA (mRNA), ribosomal RNA (rRNA), or transfer RNA (tRNA). The RNA may be a transcript. The RNA may be small RNA less than 200 nucleobases in length or large RNA more than 200 nucleobases in length. Small RNAs may include 5.8S ribosomal RNA (rRNA), 5S rRNA, transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), small tRNA-derived RNA (tsRNA), and small rDNA-derived RNA (srRNA). The RNA may be double-stranded or single-stranded. The RNA may be circular. The polymeric component may comprise a protein. The polymeric component may comprise a peptide. The polymeric component may comprise a polypeptide.

[0045] As used herein, the term "molecular tag" generally refers to a molecule capable of binding to a macromolecular component. A molecular tag can bind to a macromolecular component with high affinity. A molecular tag can bind to a macromolecular component with high specificity. A molecular tag may comprise a nucleotide sequence. A molecular tag may comprise a nucleic acid sequence. The nucleic acid sequence may be at least a portion of or the entire molecular tag. A molecular tag may be a nucleic acid molecule or a portion of a nucleic acid molecule. A molecular tag may be an oligonucleotide or a polypeptide. A molecular tag may comprise a DNA aptamer. A molecular tag may be or comprise a primer. A molecular tag may be or comprise a protein. A molecular tag may comprise a polypeptide. A molecular tag may be a barcode.

[0046] As used herein, the term "housekeeping gene," "housekeeping control," or similar terms generally refers to a gene that is expressed in an organism under both normal and pathophysiological conditions, or a gene that is expressed by different tissues and cell types. In some cases, housekeeping genes are constitutive genes required for maintaining basic cellular functions. Housekeeping genes are generally expressed at a relatively constant rate under most normal and pathophysiological conditions. Specific examples of housekeeping genes include, but are not limited to, RPP30, β-actin, and / or GAPDH.

[0047] As used herein, the term "partition" generally refers to a space or volume that may contain one or more species or be suitable for carrying out one or more reactions. A partition may be a physical compartment, such as a droplet or a well. A compartment may isolate a space or volume from another space or volume. A droplet may be a first phase (e.g., an aqueous phase) in a second phase (e.g., an oil) that is immiscible with the first phase. A droplet may be a first phase in a second phase that is not phase-separated from the first phase, such as, for example, a capsule or liposome in an aqueous phase. A compartment may contain one or more other (internal) compartments. In some cases, a compartment may be a virtual compartment that can be defined and identified by an index (e.g., an indexed library) that spans multiple and / or distant physical compartments. For example, a physical compartment may contain multiple virtual compartments.

[0048] As used herein, the terms "a," "an," and "the" generally refer to singular and plural referents unless the context clearly dictates otherwise.

[0049] Whenever the terms "at least," "greater than," or "equivalent to" precede the first number in a series of two or more numbers, the terms "at least," "greater than," or "equivalent to" apply to each and every number in the series. For example, 1, 2, or 3 or more are equivalent to 1 or more, 2 or more, or 3 or more.

[0050] Whenever the terms "no more than," "less than," or "less than or equal to" precede the first number in a series of two or more numbers, the terms "no more than," "less than," or "less than or equal to" apply to each and every number in the series. For example, 3, 2, or 1 or less is equivalent to 3 or less, 2 or less, or 1 or less.

[0051] In the following description, specific details are set forth in order to provide a thorough understanding of various embodiments. However, one of ordinary skill in the art will understand that the provided embodiments may be practiced without these details. Unless the context otherwise requires, throughout the following specification and claims, the word "comprise" and variations thereof (e.g., "comprises" or "comprising") shall be construed in its open and inclusive sense, i.e., "including, but not limited to," unless the context otherwise requires. As used in the specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. It should also be noted that the term "or" is generally used in its sense including "and / or" unless the content clearly dictates otherwise. Additionally, the headings provided herein are for convenience only and do not interpret the scope or meaning of the subject embodiments.

[0052] As used herein, the term "about" refers to an amount close to the stated amount by 10% or less.

[0053] The terms "homologous," "homology," or "percent homology," as used herein to describe an amino acid sequence or a nucleic acid sequence compared to a reference sequence, can be determined using the formula described by Karlin and Altschul (Proc. Natl. Acad. Sci. USA 87:2264-2268, 1990, modified as in Proc. Natl. Acad. Sci. USA 90:5873-5877, 1993). Such formulas are incorporated into the basic local alignment search tool (BLAST) program of Altschul et al. (J. Mol. Biol. 215:403-410, 1990). Percent sequence homology can be determined using the most recent version of BLAST as of the filing date of this application.

[0054] As used herein, the terms "individual," "patient," or "subject" refer to an individual who has been diagnosed with, is suspected of suffering from, or is at risk for developing at least one disease for which the described compositions and methods are useful for detecting. In certain embodiments, the individual is a mammal. In certain embodiments, the mammal is a mouse, rat, rabbit, dog, cat, horse, cow, sheep, pig, goat, llama, alpaca, or yak. ​​In certain embodiments, the individual is a human.

[0055] pathogenic infection The methods described herein can allow for the detection of many different pathogens, including bacteria, viruses, fungi, protozoans, nematodes, and viroids.

[0056] For example, pathogens that may be detected or diagnosed include, but are not limited to, Bacillus anthracis (anthrax), Clostridium botulinum toxin (botulism), Yersinia pestis (plague), Variola major (smallpox) and other related poxviruses, Francisella tularensis (tularemia), viral hemorrhagic fevers, arenaviruses (e.g., Junin virus, Machupo virus, Ganarito virus, Chapare virus, Lassa virus, and / or Lujo virus), Bunyaviruses (e.g., hantaviruses that cause hantavirus pulmonary syndrome, Rift Valley fever, and / or Crimean-Congo fever), and / or other pathogens that may be detected or diagnosed. Hemorrhagic fever), flaviviruses, dengue fever, filoviruses (e.g., Ebola virus and Marburg virus), Burkholderia pseudomallei (melioidosis), Coxiella (Q fever), Brucella species (brucellosis), Burkholderia mallei (glanders), Chlamydia psittacosis (psittacosis), ricin toxin (castor bean), epsilon toxin (Clostridium perfringens), Staphylococcal enterotoxin B (SEB), typhus (Rickettsia typhi), food-borne and drinking-water-borne pathogens, diarrheagenic Escherichia coli, pathogenic Vibrio species, Shigella species, Salmonella, Listeria monocytogenes, Campylobacter jejuni, Enterocoli Mosquitoes, Calicivirus, Hepatitis A, Cryptosporidium parvum, Cyclospora cayetanensis, Giardia lamblia, Entamoeba histolytica, Toxoplasma gondii, Naegleria fowleri, Balamuthia mandrillus, fungi, Microsporidia, Mosquito-borne viruses (e.g., West Nile virus (WNV), La Crosse encephalitis (LACV), California encephalitis, Venezuelan equine encephalitis (VEE), Eastern equine encephalitis (EEE), Western equine encephalitis (WEE), Japanese encephalitis virus (JE), St. Louis encephalitis virus (SLEV), Yellow fever virus (YFV), Chikungunya virus Rux, Zika virus, Nipah virus and Hendra virus, additional hantaviruses, tick-borne hemorrhagic fever viruses, Bunyaviridae, severe fever with thrombocytopenia syndrome virus (SFTSV), Heartland virus, flaviviruses (e.g., Omsk hemorrhagic fever virus, Arkoura virus, Kyasanur Forest disease virus), tick-borne encephalitis complex flaviviruses, tick-borne encephalitis viruses, Powassan / deer tick virus, tuberculosis including drug-resistant tuberculosis, influenza virus, rabies virus, prions, streptococci, pseudomonas, shigella, campylobacter,Salmonella, Clostridium, Escherichia, Hepatitis B, Hepatitis C, Papillomavirus, Epstein-Barr virus, Chickenpox, Smallpox, Orthomyxovirus, Severe Acute Respiratory Syndrome-associated Coronavirus (SARS-CoV), SARS-CoV-2 (COVID-19), MERS-CoV, other highly pathogenic human coronaviruses, or any combination thereof.

[0057] The methods described herein can also be used to diagnose viruses. In certain embodiments, the viruses include DNA viruses. In certain embodiments, the viruses include RNA viruses.

[0058] Detection of viral nucleic acid is the basis of the methods described herein for detecting and diagnosing viral infections. For the detection and diagnosis of viral infections, the viral nucleic acid molecule is part of the genome of the virus being tested. In some embodiments, the virus is a coronavirus. In some embodiments, the coronavirus is selected from the group consisting of severe acute respiratory syndrome coronavirus 2 (COVID-19), severe acute respiratory syndrome coronavirus (SARS-CoV), and Middle East respiratory syndrome coronavirus (MERS-CoV). In some embodiments, the coronavirus is COVID-19. The disclosed methods are useful for classifying other RNA and DNA genome viruses. In some embodiments, the virus is an RNA virus. In some embodiments, the RNA virus comprises a double-stranded RNA genome. In some embodiments, the RNA virus comprises a single-stranded RNA genome. In some embodiments, the RNA virus is selected from the group consisting of coronavirus, influenza, human immunodeficiency virus, and Ebola virus. In some embodiments, the virus is a DNA virus. In some embodiments, the viral infection is COVID-19.

[0059] For example, the methods herein are useful for detecting coronaviruses. The methods for diagnosing coronavirus infections are also applicable to other viruses. Coronaviruses are members of the Coronaviridae family, the Coronavirinae subfamily, and the Nidovirales order (International Committee on Taxonomy of Viruses). The Coronavirinae subfamily consists of four genera: Alphacoronavirus, Betacoronavirus, Gammacoronavirus, and Deltacoronavirus. These genera are distinguished based on phylogenetic relationships and genome structure. Alphacoronaviruses and Betacoronaviruses generally infect mammals. Similarly, Gammacoronaviruses and Deltacoronaviruses generally infect birds but can also infect mammals. Alphacoronaviruses and Betacoronaviruses are associated with and can cause respiratory disease and gastroenteritis in humans. Highly pathogenic viruses, such as Severe Acute Respiratory Syndrome Coronavirus (SARS-CoV) and Middle East Respiratory Syndrome Coronavirus (MERS-CoV), can cause severe respiratory syndrome in humans. The other four coronaviruses (e.g., HCoV-NL63, HCoV-229E, HCoV-OC43, and HKU1) generally induce mild upper respiratory tract illness in immunocompromised hosts, infants, young children, and the elderly. Alphacoronaviruses and betacoronaviruses can impose a significant disease burden not only on humans but also on livestock. These viruses include porcine transmissible gastroenteritis virus, porcine enteric diarrhea virus (PEDV), and porcine acute diarrhea syndrome coronavirus (SADS-CoV). Based on current sequence databases, all human coronaviruses have zoonotic origins. For example, SARS-CoV, MERS-CoV, HCoV-NL63, and HCoV-229E are thought to have originated in bats. Livestock may play an important role as intermediate hosts, enabling virus transmission from natural hosts to humans. In addition, livestock themselves can be susceptible to coronavirus-induced diseases.

[0060] biological samples Provided herein are methods for using and processing biological samples from individuals to diagnose viral infections. Such biological samples can be from individuals who have previously been exposed to a virus, who have previously tested positive for a virus, or who are considered to be at risk of viral exposure. The methods described herein can also be useful for population surveillance, i.e., when using samples from multiple individuals to provide information on the amount of viral infection in a given population. The biological sample can be oral or nasal mucosa. The sample can be blood, serum, or plasma. In certain embodiments, the biological sample includes a nasal swab, a nasopharyngeal swab, a buccal swab, an oral fluid swab, a middle turbinate swab, or any combination thereof, where a swab is used to collect a sample from an individual.

[0061] In practice, biological samples are collected using a myriad of collection devices, all of which can be used with the device of the present invention. Collection devices are generally commercially available, but can also be specially designed and manufactured for a given application. For clinical samples, a variety of commercially available swab types are available, including nasal, nasopharyngeal, buccal, oral, stool, tonsillar, vaginal, cervical, and wound swabs. Sample collection devices vary in size and material, and may include special handles, caps, scores to facilitate and indicate breakage, and collection matrices.

[0062] Blood samples are collected in a wide variety of commercially available tubes with varying capacities, some of which contain additives (including anticoagulants such as heparin, citrate, and EDTA), a vacuum to facilitate sample flow, a stopper to facilitate needle insertion, and a cover to protect the operator from sample exposure. Tissues and bodily fluids (e.g., sputum, purulent material, aspirates) are also collected in tubes, but these are generally distinct from blood tubes. These clinical sample collection devices are generally sent to advanced hospitals or private clinical laboratories for testing (although certain tests, such as evaluation of throat / tonsil swabs for rapid streptococcal testing, can be performed at the point of care). Environmental samples may be present as filters or filter cartridges (e.g., from air respirators, aerosols, or water filtration devices), swabs, powders, or liquids.

[0063] Sample processing After collection, the biological sample is contacted with a lysis agent to release the viral nucleic acid present in cells obtained from the biological sample. Such lysis also releases the genomic DNA of the individual being tested. Such genomic DNA can serve as a sample / amplification control. Sample collection and lysis can be performed before the method described herein for amplifying viral sequences, or at the site of viral sequence amplification and sequencing.

[0064] According to the disclosed methods and systems, a sample may be collected or compartmentalized with a lysis reagent to release the contents of the sample within the compartment. Compartments include vials, tubes, and / or wells within a plate. In such embodiments, the lysis agent may be contacted with the sample suspension simultaneously with or prior to the addition of reagents used to extend and amplify nucleic acid molecules. In some embodiments, processing of nucleic acids within the sample (e.g., amplification, primer extension, reverse transcriptase, etc.) is performed under the same conditions as those used for cell lysis. Separate compartments may contain individual samples and / or one or more reagents. In some embodiments, the generated separate compartments may contain primers and enzymes for amplifying nucleic acids (e.g., reverse transcriptase and / or polymerase). In some embodiments, the generated separate compartments may contain barcoded oligonucleotides (e.g., primers containing barcode sequences). In some embodiments, the generated separate compartments may contain barcode-bearing beads. In some embodiments, the generated separate compartments may be unoccupied (e.g., no reagents, no sample).

[0065] Advantageously, when the lysis reagent and the sample are co-compartmentalized, the lysis reagent can facilitate the release of the contents of the sample in a compartment, which may remain separated from the contents of other compartments.

[0066] Examples of lysing agents include bioactive reagents, e.g., lytic enzymes used to lyse different cell types, e.g., Gram-positive or Gram-negative bacteria, plants, yeast, mammals, etc., such as lysozyme, achromopeptidase, lysostaphin, labiase, chitalase, lytic embodiments, and various other lytic enzymes available, e.g., from Sigma-Aldrich, Inc. (St. Louis, Mo.), as well as other commercially available lytic enzymes. Other lysing agents may additionally or alternatively be co-compartmentalized with the sample to cause the release of sample contents into the compartment. For example, in some embodiments, detergent-based lysing solutions may be used to lyse cells, although these may be less preferred for emulsion-based systems where detergents may interfere with stable emulsions. In some embodiments, the lysing solution may include a non-ionic detergent, e.g., Triton X-100 and / or Tween 20. In some embodiments, the lysis solution may include an ionic detergent, such as, for example, sarkosyl and sodium dodecyl sulfate (SDS). The lysis agents described herein may include one or more proteinases (e.g., proteinase K). Electroporation, thermal, acoustic, or mechanical cell disruption may also be used in certain embodiments, and non-emulsion-based compartmentalization may be used, such as encapsulation of the sample, which may be in addition to or instead of compartmentalization, where any pore size of the encapsulant is small enough to retain nucleic acid fragments of a given size after cell disruption.

[0067] The lysis agent may further comprise hot water or a reaction buffer that is heated before or after addition of the sample. For example, the lysis agent may be a heated PCR or reverse transcriptase reaction mixture that is heated to at least about 50, 55, 60, 65, 70, 75, 80, 85, 90, or 95°C. The amount of time may be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 minutes or more.

[0068] Alternatively, or in addition to the lysis agents compartmentalized with the sample described above, other reagents can also be compartmentalized with the sample, including, for example, DNase and RNase inactivators or inhibitors, e.g., proteinase K, chelating agents, e.g., EDTA, and other reagents employed in removing or otherwise reducing the negative activity or impact of different cell lysate components on subsequent processing of nucleic acids. Additionally, in encapsulated sample embodiments, the sample may be exposed to a suitable stimulus to release the sample or its contents from the co-compartmentalized microcapsules. For example, in some embodiments, a chemical stimulus may be co-compartmentalized with the encapsulated sample to enable disassembly of the microcapsules and release of the cells or their contents into larger compartments. In some embodiments, this stimulus may be the same as the stimulus described elsewhere herein for release of nucleic acid molecules (e.g., oligonucleotides) from respective microcapsules (e.g., beads). In alternative aspects, this may be a different, non-overlapping stimulus to allow release of the encapsulated sample into a compartment at a different time than release of the nucleic acid molecules into the same compartment.

[0069] The compartments may contain species (e.g., reagents) for carrying out one or more reactions. The species may include, for example, reagents for nucleic acid amplification reactions (e.g., primers, polymerase, dNTPs, cofactors (e.g., ionic cofactors), buffers), including those described herein, reagents for enzymatic reactions (e.g., enzymes, cofactors, substrates, buffers), reagents for reverse transcription (e.g., reverse transcriptase), reagents for nucleic acid modification reactions such as polymerization, ligation, or digestion, and / or reagents for template preparation. In some cases, a primer may bind to a precursor. A primer may be used for reverse transcription. The primer may contain a poly-T sequence or may be specific to a target nucleic acid molecule (i.e., complementary to the sequence of the target nucleic acid). In some embodiments, the primer hybridizes to a specific target nucleic acid sequence (e.g., a viral nucleic acid sequence).

[0070] Additional reagents may be compartmentalized with the sample, such as an endonuclease to fragment the sample's DNA, a DNA polymerase enzyme used to amplify the sample's nucleic acid fragments and attach barcode molecular tags to the amplified fragments, and dNTPs. Other enzymes may also be compartmentalized, including, but not limited to, polymerases, transposases, ligases, proteinase K, DNAse, etc. If the lysed sample contains an RNA-based virus, the additional reagents may also include a reverse transcriptase.

[0071] One advantage of the methods described herein is that a saliva sample can be used for the detection of one or more influenza A / B, coronavirus, or other pathogen sequences. A saliva sample for use in the methods described herein may contain less than about 100, 50, 25, 10, 9, 8, 7, 6, 5, 4, 3, or 2 microliters of saliva. In certain embodiments, the saliva sample is less than about 20 microliters. In certain embodiments, the saliva sample is less than about 10 microliters. In certain embodiments, the saliva sample is less than about 8 microliters. In certain embodiments, the saliva sample is less than about 7 microliters. In certain embodiments, the saliva sample is less than about 6 microliters. In certain embodiments, the saliva sample is less than about 5 microliters.

[0072] One advantage of the methods described herein is that a small volume of nasal swab sample can be used for the detection of one or more influenza A / B, coronavirus, or other pathogen sequences. The nasal swab may be inoculated into approximately 1 milliliter of a buffer solution, such as PBS or saline, followed by use of a small volume of sample to detect viral infection. A nasal swab sample for use in the methods described herein may comprise less than about 100, 50, 25, 10, 9, 8, 7, 6, 5, 4, 3, or 2 microliters of inoculated nasal swab. In certain embodiments, the inoculated nasal swab sample is less than about 20 microliters. In certain embodiments, the inoculated nasal swab sample is less than about 10 microliters. In certain embodiments, the inoculated nasal swab sample is less than about 8 microliters. In certain embodiments, the inoculated nasal swab sample is less than about 7 microliters. In certain embodiments, the inoculated nasal swab sample is less than about 6 microliters. In certain embodiments, the inoculated nasal swab sample is less than about 5 microliters.

[0073] Pathogen sequence detection Viruses generally contain genomes containing deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Such genomes can consist of single-stranded or double-stranded nucleic acids. Furthermore, viral genomes generally contain nucleic acid sequences that differ from those of the host they infect. Therefore, the distinct nucleic acid sequences within the viral genome provide targets for detection using molecular and gene amplification and sequencing techniques. Furthermore, differences in gene sequence and structure among viruses allow viral classes, subclasses, and individual strains to be distinguished from one another. Therefore, primers (e.g., probes) that recognize virus-specific nucleic acid sequences, elements, templates, or loci can be used to identify and detect viral infections.

[0074] The detection of viral sequences described herein involves amplifying viral nucleic acids using a PCR reaction and sequencing the results of the PCR amplification. In instances where the virus being detected is an RNA-based virus, the amplification further includes a reverse transcription reaction. The reverse transcription reaction may be a separate step prior to PCR amplification, or it may be a one-step reaction that occurs in the presence of a PCR enzyme and primers. PCR amplification is performed via an oligonucleotide primer pair. In addition to containing a target-specific portion, such primers may also contain an index sequence and / or a sequencing adapter sequence, as shown in Figure 3.

[0075] Amplification of viral and host nucleic acid molecules can be achieved through the use of enzymes that extend or amplify primers hybridized to viral and host nucleic acid molecules. In particular, reverse transcriptases are used to generate cDNAs corresponding to host and viral nucleic acids. For example, in the case of viruses with genomes containing RNA, the nucleic acid extension reaction is a reverse transcription reaction. The RNA nucleic acid product of DNA transcription within host cells can be processed using reverse transcriptases. Reverse transcriptases are readily available commercially. Avian myeloblastosis virus (AMV) reverse transcriptase and Moloney murine leukemia virus (M-MuLV, MMLV) reverse transcriptase are commonly used reverse transcriptases in molecular biology workflows. M-MuLV reverse transcriptase lacks 3'→5' exonuclease activity. ProtoScript II reverse transcriptase is a recombinant M-MuLV reverse transcriptase with reduced RNase H activity and improved thermostability. It can be used to synthesize single-stranded cDNA at higher temperatures than wild-type M-MuLV. This enzyme is active up to 50°C, providing greater specificity, higher cDNA yields, and fuller cDNA products up to 12 kb in length. The use of engineered RTs improves the efficiency of full-length product formation and ensures complete copying of the 5' end of mRNA transcripts, allowing for the propagation and characterization of faithful DNA copies of RNA sequences. The use of more thermostable RTs, which allow reactions to be performed at higher temperatures, is useful when working with RNAs containing significant amounts of secondary structure. In some embodiments, the nucleic acid extension reaction comprises a reverse transcriptase reaction, a polymerase chain reaction, or a combination thereof. In certain embodiments, the nucleic acid extension reaction comprises (i) hybridizing a primer to a viral nucleic acid molecule and (ii) using a reverse transcriptase enzyme to extend the primer.

[0076] After or during reverse transcription (if applicable), the sample is subjected to an amplification reaction. In certain embodiments, the amplification reaction is a PCR reaction. In certain embodiments, the PCR reaction is not a real-time PCR reaction. In certain embodiments, the PCR is performed for a set number of cycles sufficient to amplify the viral nucleic acid sequence. The lysed and reverse transcribed sample may be amplified for N cycles. In certain embodiments, N is greater than 30, 35, 40, or 45 cycles. In certain embodiments, N is between 30 and 50 cycles, between 40 and 50 cycles, between 35 and 45 cycles, between 36 and 44 cycles, between 37 and 43 cycles, between 38 and 42 cycles, between 39 and 41 cycles, or between 40 and 45 cycles. In certain embodiments, the lysed and reverse transcribed sample may be amplified for 40 cycles.

[0077] The primer pair can be added to the amplification PCR reaction at an optimized concentration. In certain embodiments, the primer set for amplifying viral nucleic acid / synthetic nucleic acid and / or host control can be added to the PCR reaction at a concentration of less than 1 micromolar. In certain embodiments, the concentration is about 800 nM, 400 nM, 200 nM, 100 nM, 150 nM, or 50 nM.

[0078] Polymerase chain reaction amplification can be used to incorporate additional functional sequences, such as variable nucleotide sequences (barcodes). In some embodiments, the polymerase chain reaction incorporates one or more additional sequences into one or both of the barcoded viral nucleic acid molecule and the barcoded host nucleic acid molecule, selected from the group consisting of a sample index sequence, an adapter sequence, a primer sequence, a primer binding sequence, a sequence configured to bind to a flow cell of a sequencer, and an additional barcode sequence.

[0079] The methods described herein involve synthetic nucleic acids that are co-reverse transcribed and / or amplified with viral nucleic acids. In certain embodiments, the set of oligonucleotide primers that amplify viral sequences also amplify the synthetic nucleic acids. A sample can be "spiked" with a synthetic nucleic acid molecule that provides information about the processing of nucleic acids in the sample. In certain embodiments, the synthetic nucleic acid molecule is added to a biological sample prior to processing or amplification. In certain embodiments, the synthetic nucleic acid is spiked into a biological sample at a known concentration or amount. In certain embodiments, the known amount is 1 x 10 4 , 1×10 3 , 1×10 2 , 10, 5, 4, 3, 2, or 1 copy of the synthetic nucleic acid. In certain embodiments, the synthetic control nucleic acid is RNA. Ideally, the length of the amplified portion of the synthetic nucleic acid matches the length of the viral nucleic acid target to be amplified, producing an amplicon that is the same length as the viral nucleic acid target or within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or 50 nucleotides. Similarly, the sequence of the synthetic nucleic acid must be highly homologous to the viral sequence, but desirably differs by at least one nucleotide so that the synthetic nucleic acid can be identified by sequencing. In certain embodiments, the synthetic nucleic acid differs by at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides compared to the sequence of the viral nucleic acid to be amplified.

[0080] Synthetic nucleic acid sequences may be used to normalize sequence reads and to account for losses during sample preparation or inefficiencies / biases introduced during amplification or sequencing.

[0081] The synthetic nucleic acid may contain multiple nucleic acids with different sequences to further improve the performance of the "spiked" synthetic nucleic acid for downstream processing and analysis. The plurality may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, or more different sequences. The plurality may contain 2, 3, 4, or 5 different sequences. The plurality may contain 4 different sequences. The plurality may be for use with primers that amplify the S2 spike sequence, the N1 nucleoprotein sequence, or a combination thereof. Exemplary sequences are shown in the table below (since the synthetic nucleic acid is RNA, all Ts are Us in the RNA version of the sequence). Any one or more of the synthetic nucleic acids or multiple synthetic nucleic acids may have at least about 90%, 95%, 97%, 98%, 99%, or identical homology to any of SEQ ID NOs: 1-12 shown below.

[0082] [Table 1]

[0083] The synthetic nucleic acid contains 5' and 3' proximal sequences that can be bound by the same primers that amplify the viral nucleic acid sequence being tested, but that contain different intervening sequences that are distinguishable from the viral nucleic acid sequence being tested by sequencing. In some embodiments, the methods described herein further include providing a synthetic nucleic acid molecule. In some embodiments, the methods described herein further include providing a synthetic nucleic acid molecule that lacks 5' and 3' proximal sequences that can be bound by the same primers that amplify the viral nucleic acid. In certain embodiments, the 5' proximal region, the 3' proximal region, or both the 5' proximal region and the 3' proximal region are less than about 30 nucleotides in length. In certain embodiments, the 5' proximal region, the 3' proximal region, or both the 5' proximal region and the 3' proximal region are less than about 25 nucleotides in length. In certain embodiments, the 5' proximal region, the 3' proximal region, or both the 5' proximal region and the 3' proximal region are less than about 20 nucleotides in length. In certain embodiments, the 5'-proximal region is at the 5'-end of the synthetic nucleic acid. In certain embodiments, the 3'-proximal region is at the 3'-end of the synthetic nucleic acid. In certain embodiments, the intervening nucleic acid sequence is less than about 99%, less than 98%, less than 97%, less than 95%, less than 90%, less than 85%, less than 80%, or less than 75% identical to the viral nucleic acid sequence. In some embodiments, the synthetic nucleic acid molecule comprises a synthetic sequence that differs from the viral nucleic acid molecule and the human nucleic acid molecule. In some embodiments, the synthetic nucleic acid molecule comprises no more than 10%, no more than 25%, no more than 50%, no more than 75%, no more than 90%, no more than 95%, or no more than 98% nucleotide sequence identity to the host and / or viral sequence.

[0084] In certain embodiments, synthetic nucleic acids are added to the lysate to achieve a concentration in the manner of about 1 copy / well to about 1,000,000 copies / well. In certain embodiments, synthetic nucleic acids are added to the lysate to achieve a concentration in the manner of about 1 copy / well to about 50 copies / well, about 1 copy / well to about 100 copies / well, about 1 copy / well to about 500 copies / well, about 1 copy / well to about 1,000 copies / well, about 1 copy / well to about 5,000 copies / well, about 1 copy / well to about 10,000 copies / well, about 1 copy / well to about 100,000 copies / well, about 1 copy / well to about 1,000,000 copies / well, about 50 copies / well to about 1 00 copies / well, about 50 copies / well to about 500 copies / well, about 50 copies / well to about 1,000 copies / well, about 50 copies / well to about 5,000 copies / well, about 50 copies / well to about 10,000 copies / well, about 50 copies / well to about 100,000 copies / well, about 50 copies / well to about 1,000,000 copies / well, about 100 copies / well to about 500 copies / well, about 100 copies / well to about 1,000 copies / well, about 100 copies / well to about 5,0 00 copies / well, about 100 copies / well to about 10,000 copies / well, about 100 copies / well to about 100,000 copies / well, about 100 copies / well to about 1,000,000 copies / well, about 500 copies / well to about 1,000 copies / well, about 500 copies / well to about 5,000 copies / well, about 500 copies / well to about 10,000 copies / well, about 500 copies / well to about 100,000 copies / well, about 500 copies / well to about 1,000,000 copies / well , about 1,000 copies / well to about 5,000 copies / well, about 1,000 copies / well to about 10,000 copies / well, about 1,000 copies / well to about 100,000 copies / well, about 1,000 copies / well to about 1,000,000 copies / well, about 5,000 copies / well to about 10,000 copies / well, about 5,000 copies / well to about 100,000 copies / well, about 5,000 copies / well to about 1,000,000 copies / well, about 10,000 copies / well to about 100,In certain embodiments, the synthetic nucleic acid is added at about 1 copy / well, about 50 copies / well, about 100 copies / well, about 500 copies / well, about 1,000 copies / well, about 5,000 copies / well, about 10,000 copies / well, about 100,000 copies / well, or about 1,000,000 copies / well. In certain embodiments, the synthetic nucleic acid is added at at least about 1 copy / well, about 50 copies / well, about 100 copies / well, about 500 copies / well, about 1,000 copies / well, about 5,000 copies / well, about 10,000 copies / well, about 100,000 copies / well, or about 1,000,000 copies / well. In certain embodiments, the synthetic nucleic acid is added at a maximum of about 50 copies / well, about 100 copies / well, about 500 copies / well, about 1,000 copies / well, about 5,000 copies / well, about 10,000 copies / well, about 100,000 copies / well, or about 1,000,000 copies / well.

[0085] After amplification and / or sequencing, the resulting ratio of detected pathogen nucleic acid molecules to synthetic oligonucleotides is useful for detecting and diagnosing pathogenic infections. The components of the ratio used for detecting and diagnosing infectious agents can be based on the number of reads for a particular sequence or set of sequences, the copy number of a particular sequence or set of sequences (e.g., the number of different unique molecular identifier sequences associated with a particular sequence), the number or amount derived from sequencing information, or any combination thereof. In some embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is greater than 1. In some embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is equal to or about 1. In some embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is less than 1.

[0086] In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is from about 0.00001:1 to about 0.5:1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is between about 0.00001:1 and about 0.00005:1, between about 0.00001:1 and about 0.0001:1, between about 0.00001:1 and about 0.0002:1, between about 0.00001:1 and about 0.0003:1, between about 0.00001:1 and about 0.0004:1, between about 0.00001:1 and about 0.0005:1, between about 0.00001:1 and about 0.001:1, between about 0.00001:1 and about 0.005:1, between about 0.00001:1 and about 0.01:1, between about 0.00001:1 and about 0.1:1, between about 0.00001:1 and about 0.1:1, or between about 0.00001:1 and about 0.00001:1. 1 to about 0.5:1, about 0.00005:1 to about 0.0001:1, about 0.00005:1 to about 0.0002:1, about 0.00005:1 to about 0.0003:1, about 0.00005:1 to about 0.0004:1, about 0.00005:1 to about 0.0005:1, about 0.00005:1 to about 0.001:1, about 0.00005:1 to about 0.005:1, about 0.00005:1 to about 0.01:1, about 0.00005:1 to about 0.1:1, about 0.00005:1 to about 0.5:1, about 0.0001:1 to about 0.0002:1, about 0.0001: 1 to about 0.0003:1, about 0.0001:1 to about 0.0004:1, about 0.0001:1 to about 0.0005:1, about 0.0001:1 to about 0.001:1, about 0.0001:1 to about 0.005:1, about 0.0001:1 to about 0.01:1, about 0.0001:1 to about 0.1:1, about 0.0001:1 to about 0.5:1, about 0.0002:1 to about 0.0003:1, about 0.0002:1 to about 0.0004:1, about 0.0002:1 to about 0.0005:1, about 0.0002:1 to about 0.001:1, about 0.0002:1 to about 0.005 :1, about 0.0002:1 to about 0.01:1, about 0.0002:1 to about 0.1:1, about 0.0002:1 to about 0.5:1, about 0.0003:1 to about 0.0004:1, about 0.0003:1 to about 0.0005:1, about 0.0003:1 to about 0.001:1, about 0.0003:1 to about 0.005:1, about 0.0003:1 to about 0.01:1, about 0.0003:1 to about 0.1:1, about 0.0003:1 to about 0.5:1, about 0.0004:1 to about 0.0005:1, about 0.0004:1 to about 0.001:1, about 0.0004:1 to about 0.0.0004:1 to about 0.01:1, about 0.0004:1 to about 0.1:1, about 0.0004:1 to about 0.5:1, about 0.0005:1 to about 0.001:1, about 0.0005:1 to about 0.005:1, about 0.0005:1 to about 0.01:1, about 0.0005:1 to about 0.1:1, about 0.0005:1 to about 0.5:1, about 0.001:1 to about 0.005:1, about 0.001:1 to about 0.01:1, about 0.001:1 to about 0.1:1, about 0.001:1 to about 0.5:1, about 0.005:1 to about 0.01:1, about 0.005:1 to about 0.1:1, about 0.005:1 to about 0.5:1, about 0.01:1 to about 0.1:1, about 0.01:1 to about 0.5:1, or about 0.1:1 to about 0.5:1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is about 0.00001:1, about 0.00005:1, about 0.0001:1, about 0.0002:1, about 0.0003:1, about 0.0004:1, about 0.0005:1, about 0.001:1, about 0.005:1, about 0.01:1, about 0.1:1, or about 0.5:1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is at least about 0.00001:1, about 0.00005:1, about 0.0001:1, about 0.0002:1, about 0.0003:1, about 0.0004:1, about 0.0005:1, about 0.001:1, about 0.005:1, about 0.01:1, or about 0.1:1.

[0087] In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is about 1:1 to about 1:100. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is about 1:100 to about 1:50, about 1:100 to about 1:25, about 1:100 to about 1:10, about 1:100 to about 1:5, about 1:100 to about 1:4, about 1:100 to about 1:3, about 1:100 to about 1:2, about 1:100 to about 1:1, about 1:100 to about 1:1, about 1:50 to about 1:25, about 1:50 to about 1:10, about 1:50 to about 1:5, about 1:50 to about 1:4, about 1:50 to about 1:3, about 1:50 to about 1:2, about 1:50 to about 1:1, about 1:50 to about 1:1, about 1:25 to about 1:10, about 1:25 to about 1:5, about 1:25 to about 1:4, about 1: 25 to about 1:3, about 1:25 to about 1:2, about 1:25 to about 1:1, about 1:25 to about 1:1, about 1:10 to about 1:5, about 1:10 to about 1:4, about 1:10 to about 1:3, about 1:10 to about 1:2, about 1:10 to about 1:1, about 1:10 to about 1:1, about 1:5 to about 1:4, about 1:5 to about 1:3, about 1:5 about 1:2, about 1:5 to about 1:1, about 1:5 to about 1:1, about 1:4 to about 1:3, about 1:4 to about 1:2, about 1:4 to about 1:1, about 1:4 to about 1:1, about 1:3 to about 1:2, about 1:3 to about 1:1, about 1:3 to about 1:1, about 1:2 to about 1:1, about 1:2 to about 1:1, or about 1:1 to about 1:1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is about 1:100, about 1:50, about 1:25, about 1:10, about 1:5, about 1:4, about 1:3, about 1:2, about 1:1, or about 1:1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is at least about 1:100, about 1:50, about 1:25, about 1:10, about 1:5, about 1:4, about 1:3, about 1:2, or about 1: 1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is at most about 1:50, about 1:25, about 1:10, about 1:5, about 1:4, about 1:3, about 1:2, about 1:1, or about 1:1.

[0088] In certain embodiments, the ratio of synthetic nucleic acid to pathogen nucleic acid is about 1:100 to about 1:50, about 1:100 to about 1:25, about 1:100 to about 1:10, about 1:100 to about 1:5, about 1:100 to about 1:4, about 1:100 to about 1:3, about 1:100 to about 1:2, about 1:100 to about 1:1, about 1:100 to about 1:1, about 1:50 to about 1:25, about 1:50 to about 1:10, about 1:50 to about 1:5, about 1:50 to about 1:4, about 1:50 to about 1:3, about 1:50 to about 1:2, about 1:50 to about 1:1, about 1:50 to about 1:1, about 1:25 to about 1:10, about 1:25 to about 1:5, about 1:25 to about 1:4, about 1: 25 to about 1:3, about 1:25 to about 1:2, about 1:25 to about 1:1, about 1:25 to about 1:1, about 1:10 to about 1:5, about 1:10 to about 1:4, about 1:10 to about 1:3, about 1:10 to about 1:2, about 1:10 to about 1:1, about 1:10 to about 1:1, about 1:5 to about 1:4, about 1:5 to about 1:3, about 1:5 about 1:2, about 1:5 to about 1:1, about 1:5 to about 1:1, about 1:4 to about 1:3, about 1:4 to about 1:2, about 1:4 to about 1:1, about 1:4 to about 1:1, about 1:3 to about 1:2, about 1:3 to about 1:1, about 1:3 to about 1:1, about 1:2 to about 1:1, about 1:2 to about 1:1, or about 1:1 to about 1:1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is about 1:100, about 1:50, about 1:25, about 1:10, about 1:5, about 1:4, about 1:3, about 1:2, about 1:1, or about 1:1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is at least about 1:100, about 1:50, about 1:25, about 1:10, about 1:5, about 1:4, about 1:3, about 1:2, or about 1: 1. In certain embodiments, the ratio of pathogen nucleic acid to synthetic nucleic acid is at most about 1:50, about 1:25, about 1:10, about 1:5, about 1:4, about 1:3, about 1:2, about 1:1, or about 1:1.

[0089] The ratio of pathogen reads to pathogen synthetic nucleic acids (having sequences distinguishable from the pathogen reads) can be used to indicate a positive diagnosis for a particular pathogen (e.g., coronavirus, influenza A, influenza B). Alternatively, a negative diagnosis is made if the ratio does not exceed a positive threshold. In some embodiments, a positive diagnosis for a pathogen is made if the sequence reads from the pathogen nucleic acid to the synthetic nucleic acid exceed a ratio of about 0.01 to about 0.5. In some embodiments, a positive diagnosis for the pathogen is made when the sequence read from the pathogen nucleic acid to the synthetic nucleic acid exceeds a ratio of about 0.01 to about 0.4, about 0.01 to about 0.3, about 0.01 to about 0.2, about 0.02 to about 0.5, about 0.02 to about 0.2, about 0.03 to about 0.5, about 0.03 to about 0.2, about 0.04 to about 0.5, about 0.03 to about 0.2, about 0.05 to about 0.5, about 0.05 to about 0.2, about 0.06 to about 0.5, about 0.06 to about 0.2, about 0.07 to about 0.5, about 0.07 to about 0.2, about 0.08 to about 0.5, or about 0.08 to about 0.2. In some embodiments, a positive diagnosis of a pathogen is made if the sequence read from the pathogen nucleic acid to the synthetic nucleic acid exceeds a ratio of about 0.01, about 0.02, about 0.03, about 0.04, about 0.05, about 0.06, about 0.07, about 0.08, about 0.09, or about 0.1.

[0090] The total amount of pathogen reads plus pathogen synthetic nucleic acids (having sequences distinguishable from the pathogen reads) can be used to indicate whether there is sufficient sequence data to attempt a positive or negative diagnosis for the presence of a pathogen (e.g., coronavirus, influenza A, influenza B). In some embodiments, a positive or negative diagnosis can only be made if the total number of sequence reads of the pathogen nucleic acids together with the synthetic nucleic acids exceeds a minimum threshold; otherwise, the result is inconclusive. In some embodiments, the minimum threshold is at least about 10 reads, at least about 20 reads, at least about 30 reads, at least about 40 reads, at least about 50 reads, at least about 60 reads, at least about 70 reads, at least about 80 reads, at least about 90 reads, at least about 100 reads, at least about 150 reads, at least about 200 reads, or at least about 250 reads.

[0091] In some cases, a sample control can be used to verify that sufficient genetic material is present in the sample to reliably call a positive or negative diagnosis. The amount of reads can be counted for the sample control. This sample control is usually a housekeeping gene for the species of the individual being tested. In certain embodiments, when the individual being tested is human, the sample control is β-actin, GAPDH, or RPP30. In certain embodiments, when the individual being tested is human, the sample control is RPP30. In certain embodiments, the amount of reads from the sample control present to deliver a positive or negative diagnosis is greater than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 128, 19, 20, 25, or 30 reads. In certain embodiments, the amount of reads from the sample control present to deliver a positive or negative diagnosis is equal to or greater than 10 reads.

[0092] The disclosed methods can also be combined with multiplexing strategies to enable effective and efficient testing of multiple samples in parallel. As used herein, "multiplexing" refers to simultaneously performing multiple assays on one or more samples. Multiplexing can also include simultaneously performing multiple assays on each of multiple separate samples. For example, the number of reaction mixtures analyzed can be based on the number of compartments, and the number of assays performed in each compartment can be based on the number of probes contacting the contents of each well. Multiplexing designs and strategies can be effectively employed in applying the disclosed methods. For example, one or more unique barcodes or adapter nucleic acid molecules can be attached to target nucleic acid molecules, where the unique barcodes or adapter nucleic acid molecules identify or barcode the target nucleic acid molecules. Once barcoded, the samples can be combined or pooled into a single sequencing library, and the barcoded target nucleic acids (e.g., pathogen nucleic acid molecules and / or synthetic nucleic acid molecules) can be distinguished from other barcoded samples. Thus, sample multiplexing and the use of nucleic acid barcodes can be used to identify sequencing reads from a first individual from a group or multiple individuals. In some embodiments, samples are pooled prior to sequencing. In certain embodiments, more than 10, 25, 50, 75, 100, or 500 samples are pooled and analyzed in a single sequencing library.

[0093] Furthermore, more than one pathogen can be tested in a single assay. For example, multiple respiratory viruses (or virus subtypes) can be analyzed in a single assay. Thus, methods disclosed for one pathogen can be combined for the detection and analysis of multiple pathogens in a single reaction. (For example, multiple primers specific to multiple pathogen nucleic acid molecules can be used, and multiple synthetic nucleic acid molecules can be added to a sample.) As a further example, in some embodiments, the first pathogen is a coronavirus and the second pathogen is an influenza virus. Examples of multiple pathogens also include different subtypes or clades of viruses (e.g., influenza A H1N1 and influenza A H3N2).

[0094] The methods and reaction mixtures described herein may utilize amplification and sequencing of multiple coronavirus sequences. In certain embodiments, the multiple coronavirus sequences amplified and sequenced include S2 sequences. In certain embodiments, the multiple coronavirus sequences amplified and sequenced include N1 sequences. In certain embodiments, the multiple coronavirus sequences amplified and sequenced include S2 sequences and N1 sequences. Such tests may utilize multiple synthetic nucleic acid spike ion controls for each viral target.

[0095] The methods, reaction mixtures, and kits described herein can be used to test for coronaviruses (e.g., SARS-CoV2) and influenza A and / or influenza B. Such tests would be useful for healthcare providers and health organizations to simultaneously monitor outbreaks of different common respiratory pathogens. The methods described herein can be used in triplex tests to simultaneously detect coronaviruses and influenza A and B. Such tests may utilize three different synthetic nucleic acid spiked ion controls, one for each viral target.

[0096] The methods described herein may also include a second set of oligonucleotide primers targeting nucleic acid sequences of the viral host (e.g., the individual being tested), providing a sample and a positive control for amplification. In certain embodiments, the second set of PCR primers amplifies a human nucleic acid sequence selected from GAPDH, ACTB, RPP30, and combinations thereof. In certain embodiments, the second set of PCR primers amplifies a human RPP30 sequence. In certain embodiments, the second set of PCR primers comprises a mixture of primers with and without a sequencing adapter sequence. In certain embodiments, the ratio of primers with and without a sequencing adapter sequence is about 1:1, about 1:2, about 1:3, or about 1:4. A mixture of primers with and without adapters allows for detection of viral host nucleic acids, but can allocate more sequencing reads to viral sequences.

[0097] In some embodiments, the diagnosis of pathogen infection is inconclusive if a predetermined minimum number of reads of the human nucleic acid sequence is not detected. In some embodiments, the predetermined minimum number of reads of the human nucleic acid sequence must exceed at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10 loads, at least about 15 reads, at least about 20 reads, or at least about 25 reads of the human nucleic acid sequence. In some embodiments, the predetermined minimum number of reads of the human nucleic acid sequence must exceed at least about 10 reads.

[0098] Coronaviruses form enveloped virus particles with diameters of 100–160 nm. Coronaviruses are single-stranded ribonucleic acid (ssRNA) viruses containing a positive-sense RNA genome of 27–32 kb in size. The 5′ region of the coronavirus genome encodes the polyprotein pp1ab, which is further cleaved into 16 nonstructural proteins involved in genome transcription and replication. The 3′ region encodes structural proteins, including the envelope glycoprotein spikes (e.g., viral spikes), envelope, membrane, and nucleocapsid. Genes encoding structural proteins can also function as species-specific accessory genes or may be dispensable for viral replication.

[0099] COVID-19 (also known as HCoV-19 or SARS-CoV-2) is a betacoronavirus and the seventh coronavirus known to infect humans. COVID-19 coronaviruses can cause severe disease (as observed in SARS-CoV and MERS-CoV). COVID-19 has demonstrated efficient targeting of the human receptor ACE2. The receptor-binding domain (RBD) of the spike protein is the most variable part of the coronavirus genome. Six RBD amino acids have been shown to be important for binding to the ACE2 receptor and determining the host range of SARS-CoV-like viruses. Based on the SARS-CoV protein sequence, the residues involved in ACE2 binding and contributing to receptor targeting are Y442, L472, N479, D480, T487, and Y4911 (these residues also correspond to L455, F486, Q493, S494, N501, and Y505 in SARS-CoV-2). In COVID-19, five of these six residues differ between COVID-19 and SARS-CoV, corresponding to L455, F486, Q493, S494, and N501 in the COVID-19 receptor-binding domain (RBD) protein. These sequence characteristics of COVID-19 can be used to identify, detect, and / or diagnose COVID-19. For example, the unique sequence of the COVID-19 receptor-binding domain (RBD) sequence can be used to identify, detect, or diagnose COVID-19.

[0100] COVID-19 also contains a functional polybasic (furin) cleavage site within the S1 and S2 subunits of the viral spike protein interface due to a 12-nucleotide insertion. This inserted sequence enhances the acquisition of three O-linked glycans. Both the polybasic cleavage site and the three adjacent predicted O-linked glycans are unique to COVID-19 and have not previously been identified in alphacoronaviruses or betacoronaviruses. Therefore, characterization of the COVID-19 viral spike sequence can be used to identify, detect, and / or diagnose COVID-19. For example, the unique sequence of the COVID-19 viral spike insert can be used to identify, detect, or diagnose COVID-19.

[0101] The methods described herein are useful for detecting and diagnosing viral infections in individuals. In certain embodiments, the viral infection is influenza. In certain embodiments, the viral infection is a coronavirus. In certain embodiments, the viral infection is COVID-19. In certain embodiments, the viral infection is MERS-CoV. In certain embodiments, the viral infection is SARS-CoV-2.

[0102] In principle, any suitable viral nucleic acid sequence can be targeted with an oligonucleotide primer pair. Such targets include nucleic acid sequences encoding viral nucleocapsid proteins, viral spike proteins, viral envelope proteins, or viral membrane proteins. Such targets include nonstructural proteins.

[0103] When the detected viral infection is SARS-CoV-2, the detected viral nucleic acid may encode one or more of the viral nucleocapsid protein (N1), viral spike protein (S2), viral envelope protein, or viral membrane protein. In certain embodiments, the detected nucleic acid encodes any one or more of the 16 SARS-CoV-2 nonstructural proteins: NSP1, NSP2, NSP3, NSP4, NSP5, NSP6, NSP7, NSP8, NSP9, NSP10, NSP111, NSP12, NSP13, NSP14, NSP15, or NSP16.

[0104] Exemplary primer pairs that can be used in the methods described herein can bind to the sequences set forth in the Sequence Appendix that follows this disclosure. In certain embodiments, any of the primers listed in the Appendix can be used in the methods described herein.

[0105] Primers containing barcoded oligos are useful in sample analysis (e.g., sample identification, tracking, quantification, etc.). For example, primers in a compartment containing barcoded oligonucleotides further comprising a common barcode sequence and a unique molecular identifier sequence can be used to (1) identify sequences belonging to a compartment through the use of the common barcode and (2) identify transcript copy numbers through the use of the unique molecular identifier. In some embodiments, the first primer set further comprises one or more additional functional sequences selected from the group consisting of a primer sequence, an adapter sequence, a primer annealing sequence, a unique molecular identifier sequence, and a capture sequence. In some embodiments, the second primer further comprises one or more additional functional sequences selected from the group consisting of a primer sequence, an adapter sequence, a primer annealing sequence, a unique molecular identifier sequence, and a capture sequence. Generally, primers are specific to (i.e., complementary to or comprise a sequence complementary to) a target sequence of a viral, host, or synthetic nucleic acid molecule.

[0106] Figures 4-21 further demonstrate and / or illustrate the methods and compositions disclosed herein and the advantages of their use and / or application. Figure 4 illustrates the SwabSeq diagnostic testing platform for COVID-19. In (A), the SwabSeq workflow is a five-step process that takes approximately 12 hours from start to finish. In (B), RT-PCR of a clinical sample is performed in each well. Each well contains two sets of indexed primers that generate cDNA and amplicons for the SARS-CoV-2 S2 gene and the human RPP30 gene. Each primer is synthesized with P5 and P7 adapters for Illumina sequencing, unique i7 and i5 molecular barcodes, and a unique primer pair. Importantly, every well contains a synthetic in vitro S2 standard, which is key to the scalability of this method. In (C), the in vitro S2 standard (abbreviated as S2-Spike) is complemented with the viral S2 gene, differing by six base pairs (underlined). In (D), reads are counted at various virus concentrations. In (E), ratiometric normalization allows for normalization within each amplicon well. In (F), each well has two internal well controls for amplification: an in vitro S2 standard and human RPP30. The RPP30 amplicon serves as a control for specimen collection. The in vitro S2 standard is essential for SwabSeq's ability to distinguish true negatives.

[0107] Figure 5 shows that clinical specimen validation demonstrates a detection limit comparable to that of a highly sensitive RT-qPCR reaction. (A) The limit of detection (LOD) of SARS-CoV2-free nasal swab samples was pooled and spiked with different concentrations of ATCC inactivated virus. The nasal swab samples were RNA purified and tested using SwabSeq, demonstrating a detection limit of 250 genome copy equivalents (GCE) per mL. (B) RNA-purified clinical nasal swab specimens obtained through the UCLA Health Clinical Microbiology Laboratory were tested using an FDA-approved platform based on clinical protocols and further tested using SwabSeq. 100% concordance is demonstrated with samples testing positive for SARS-CoV-2 (n=31) and negative for SARS-CoV-2 (n=35). (C) Tested RNA-purified samples from extraction-free nasopharyngeal swabs demonstrated a detection limit of 558 GCE / mL. (D) Clinical samples tested at the UCLA Health Clinical Microbiology Laboratory showed 100% agreement between negative (n=20) and positive (n=20) results. (E) Extraction-free processing of saliva specimens demonstrates a detection limit down to 1000 GCE / mL.

[0108] Figure 6 shows the sequencing library design. Amplicon designs are shown for the S2 (top) and RPP30 (bottom) amplicons. The amplicons were designed so that the i5 and i7 molecular indices uniquely identify each sample. SwabSeq was designed to be compatible with all Illumina platforms. Figure 7 shows that the S2 primers exhibit comparable PCR efficiency when amplifying the COVID-19 amplicon and the synthetic S2 spike. The slopes of PCR efficiency for primers with either S2_spike or SARS-CoV-2 virus (labeled in green as C19gRNA) input are as follows: S2_spike slope = -6.68e-6, and C19gRNA (Twist Control) slope = -6.74e-6. If the primers do not exhibit preferential amplification of S2 spike RNA versus C19gRNA, the slopes are expected to be comparable (parallel). This indicates that the S2 spike and C19 gRNA have comparable amplification efficiencies using the S2 primer pair. The bands represent 95% confidence intervals of predicted values ​​and do not overlap due to different intercepts, making them irrelevant for the analysis of this gradient.

[0109] Figure 8 shows that SwabSeq maintains linearity at very high virus concentrations. An internal well control, S2 Spike, is included to allow for calling negative samples even in the presence of heterogeneous sample types or PCR inhibition. In (A), reads attributed to S2 increase as virus concentration increases, while in (B), reads attributed to S2 Spike decrease. In (C), an additional level of ratiometric normalization is performed by the ratio of S2 to S2 Spike, demonstrating linearity up to at least 2 million copies / mL of lysate. Note that both axes are scaled on a log10 scale.

[0110] Figure 9 shows that sequencing performed on a MiSeq or NextSeq Machine exhibits similar sensitivity. Multiplexed libraries run on both MiSeq and NextSeq demonstrated linearity across a wide range of SARS-CoV2 viral copies in purified RNA backgrounds. Figure 10 shows preliminary and confirmatory limit of detection data for RNA purified samples using the NextSeq550. In (A), preliminary LOD data identified an LOD of 250 copies / mL, and in (B), confirmatory studies demonstrated an LOD of 250 copies / mL. In (C), exemplary result interpretation guidelines are provided for purified RNA.

[0111] Figure 11 shows that extraction-free protocols into traditional collection media and buffers require dilution to overcome the effects of RT and PCR inhibition. In (A), we tested the extraction-free protocol for nasopharyngeal (NP) swabs in viral transport medium (VTM). ATCC live-inactivated virus was spiked into pooled VTM at various concentrations, and the samples were diluted 1:4 with water before adding to the RT-PCR reaction. A detection limit of 5714 copies / mL was observed. In (B), nasopharyngeal (NP) swabs collected in normal saline (NS) were pooled and then spiked with various concentrations of ATCC live-inactivated virus. The artificial samples were diluted 1:4 with water. Here, earlier studies have shown similar detection limits between 2857 and 5714 copies / mL. In (C), we tested natural clinical samples collected in Amies buffer (ESwab). S gene Ct counts (x-axis) of positive samples were compared to the SwabSeq S2 to S2 spike ratio (y-axis). Samples were run in triplicate (color). High concordance was observed for Ct counts below 27, but greater variability was observed above 27, suggesting that RT and PCR inhibition impacted the detection limit.

[0112] Figure 12 illustrates the development of lightweight sample acceptance, collection, and processing, enabling scalable testing to thousands of samples per day. (A) To address sample collection challenges, a lightweight collection method was developed that allows samples to be collected directly into automatable tubes. Here, a funnel is used for individuals to deposit a small saliva sample (0.25 mL into the funnel and tube). This setup can accommodate multiple sample types. (B) To facilitate sample acceptance, a web-based app was developed for individuals to register their sample tubes using a barcode reader and transmit identifying information to a secure instance of Qualtrics. Individuals then collect their samples and place the tubes in a rack. This low-touch pre-analysis process enables thousands of samples per day without significant administrative burden. (C) The overall workflow streamlines laboratory processing. First, individuals collect samples into automatable tubes and place them in a 96-tube rack. Samples arrive at the lab in a 96-rack format, enabling efficient sample inactivation and processing, dramatically increasing sample flow through the platform.

[0113] Figure 13 shows that preheating saliva to 95°C for 30 minutes dramatically improves RT-PCR. Detection of viral genomes demonstrates improved robustness in control detection. In (A), without preheating, detection of the S2 spike is minimal, and counts of the control amplicon are low. In (B), robust detection of the S2 amplicon and synthetic S2 spike is observed with a 30-minute 95°C preheating step. Figure 14 shows that PCR inhibition significantly impacts amplification product. 2% agarose gels were run on a subset of wells from RT-PCR reactions. RT-PCR inhibition was observed from swabs of crude lysate (A1-A8) compared to purified RNA (A9-A12). Two bands were observed in this subset of wells, representing two amplicons for the S2 or S2 spike (177 bp) and RPP30 (133 bp) primer pairs. Figure 15 shows that increasing the number of PCR cycles and running the tapestation with unpurified or inhibited sample types (e.g., saliva) can increase the size of nonspecific peaks in library preparations. Representative results from the Agilent TapeStation for purified amplicon libraries are shown. A nonspecific peak (arrow) slightly larger than 100 bp was observed in both library traces, and this peak increased in size with unpurified samples and increased PCR cycle numbers. Importantly, as the size of this nonspecific peak increased, library quantification became inaccurate. Therefore, to optimize cluster density on Illumina sequencers, it is suggested to quantify the loading concentration of the final library based on the percentage of the desired peaks (RPP30 and S2).

[0114] Figure 16 shows that TaqPath reduces the number of S2 reads in SARS-CoV-2-negative samples compared to NEB Luna. (A) and (B) compare Luna One Step RT-PCR Mix (New England Biosciences) versus TaqPath™ 1-Step RT-qPCR Master Mix (Thermofisher Scientific). The presence of UNG in the TaqPath Mastermix significantly reduced the number of S2 reads in SARS-CoV-2-negative samples, potentially enabling more accurate diagnosis of SARS-CoV-2-positive versus SARS-CoV-2-negative samples.

[0115] Figure 17 shows the contribution of carryover contamination from template lines to cross-contamination in the MiSeq. In this experiment, RT-PCR was performed on four 384-well plates, but only three plates were pooled. On the left are the observed counts of each amplicon for each sample in a 384-well plate not included in the current run (but indexed in a previous run). Amplicon reads for the index used in the previous run were present at low levels (0-150 reads). Prior to the subsequent run, a bleach wash was performed in addition to the regular washes. In this subsequent run, three different plates were pooled, and the fourth 384-well plate was omitted. On the right are the observed counts of each amplicon for the sample index corresponding to the omitted plate (again, indexed in a previous run). A significant reduction in the amount of carryover contamination was observed, with carryover reads <10 per sample. Alternatively, in some embodiments, different combinations of reverse and forward primers are used in subsequent runs on the same instrument to reduce contamination and result confounds. Rotating the primers on the sequence allows us to determine whether the read is a true read from the sample or a crossover contaminant from a previous run. Figure 18 shows sequencing errors and potential amplicon misassignments in amplicon reads. Experiment v18 had lower than normal PhiX loading (11%), resulting in lower overall read quality. While the trend noted here continues in other runs, this run more clearly highlights potential issues caused by sequencing errors and overly permissive error correction. (A) Percentage of reads with a base quality score below 12 for each position in Read 1. Note that the first six bases in Read 1 distinguish S2 from S2 spike and have the highest percentage of low-quality base calls. (B) Hamming distances between each Read 1 sequence and either the expected S2 sequence (rows) or the S2 spike sequence (columns). Yellow indicates perfect match and edit distance 1 sequences that are clearly identifiable as S2 or S2 spike.In red are sequences that may have been misassigned (S2 spikes assigned as S2 are the most problematic in this assay).

[0116] Figure 19 shows a visualization of different indexing strategies. Here, the i5 index is depicted as a horizontal line, the i7 index as a vertical line, and the color represents the unique index. In combinatorial (or fully combinatorial) indexing, the i5 and i7 indexes are combined to create unique combinations, but each i5 and i7 index is used multiple times within a plate, allowing all possible i5 and i7 indices to be used. In unique dual indexing, the i5 and i7 indexes are each used only once per plate. This requires the synthesis of more oligos. In semi-combinatorial indexing, the combinations used are more limited, with indexes repeated for only a subset of wells, eliminating many possible combinations. In practice (not depicted here), a design was used in which the i7 index is unique, but the i5 index can be repeated up to four times across a 384-well plate. For the majority of Swabseq development, either semi-combinatorial indexing (384 × 96) or unique dual indexing (384UDI) was used, which allows for 1536 combinations or samples to be run.

[0117] Figure 20 shows computational correction for index misassignment using a mixture model. To expand the number of samples that can be tested, a combinatorial indexing strategy can be used. In this experiment, a single index on the i5 was used to uniquely identify the plate, and 96 i7 indexes were used to identify the wells. In (A), the ratio of S2 to S2 spike (y-axis) is plotted for clinical samples based on whether Covid was detected by RT-qPCR (x-axis). SARS-CoV-2-positive samples were filtered to have a Ct < 32. The impact of index misassignment across plates can be observed as a high i7 index (color) for the sum of S2 and S2 spike across all samples sharing the same i7 barcode across plates. In (B), the residuals of the best linear unbiased predictor are plotted for the data in A after computational correction for the log10(S2 + 1 / S2_spike + 1) ratio by treating the identity of the i7 barcode as a random effect (y-axis).

[0118] Figure 21 shows the quantification of the role of index misassignment as a source of noise in S2 reads. (A) Matching matrix of viral S2 + S2 spike counts for each pair of i5 and i7 index pairs from run v19 using a unique dual-index design. Index pairs along the diagonal correspond to the expected index pairs (expected matching indexes) for samples present in the experiment, while index pairs off the diagonal correspond to index misassignment events. (B) Distribution of viral S count to spike count ratios for samples with zero pre-existing viral RNA. The average ratio is 0.00028. (C) Number of i7 misassignment events vs. number of viral S2 + S2 spike counts for each sample. (D) Number of i5 misassignment events vs. number of viral S2 + S2 spike counts for each sample.

[0119] Barcode The variable nucleotide sequence (barcode) that functions as an index can be included in either the first set or the second set of PCR primers described herein.Furthermore, the barcode can be added in a separate library preparation reaction.The variable nucleotide sequence described herein can be used as a sample index to deconvolute the results obtained from the sequencing reaction used herein.

[0120] Once the cellular contents have been released into their respective compartments by the lysis agent, the macromolecular components contained therein (e.g., macromolecular components of the sample, such as RNA, DNA, or proteins) can be further processed within the compartment. According to the methods and systems described herein, the macromolecular component contents of individual samples can be provided with unique identifiers so that, upon characterization, they can be attributed as originating from the same sample or particle. The ability to attribute features to individual samples or groups of samples is provided by specifically assigning unique identifiers to individual samples or groups of samples. Unique identifiers, for example, in the form of nucleic acid barcodes, can be assigned or associated with individual samples or populations of samples to tag or label the macromolecular components of the sample (and, consequently, their features). These unique identifiers can then be used to attribute sample components or features to individual samples or groups of samples.

[0121] In some aspects, this is accomplished by dividing individual samples or groups of samples into compartments with unique identifiers or barcodes containing unique molecular identifier sequences (UMIs). In some aspects, the unique identifiers are provided in the form of nucleic acid molecules (e.g., oligonucleotides) containing nucleic acid barcode sequences that can bind to or otherwise associate with the nucleic acid content of the individual samples or other components of the samples, particularly fragments of those nucleic acids. The nucleic acid molecules are divided into compartments such that the nucleic acid barcode sequences contained therein are the same among the nucleic acid molecules in a given compartment, but the nucleic acid molecules in different compartments can and do have different barcode sequences, or alternatively, all of the compartments in a given analysis at least represent many different barcode sequences. In some aspects, only one nucleic acid barcode sequence may be associated with a given compartment, but in some embodiments, two or more different barcode sequences may be present.

[0122] Nucleic acid barcode sequences can comprise from about 6 to about 20 or more nucleotides within the sequence of a nucleic acid molecule (e.g., an oligonucleotide). Nucleic acid barcode sequences can comprise from about 6 to about 20, 30, 40, 50, 60, 70, 80, 90, 100, or more nucleotides. In some embodiments, the length of the barcode sequence can be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, the length of the barcode sequence can be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, the length of the barcode sequence can be up to about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or less nucleotides. These nucleotides may be completely contiguous, i.e., a single stretch of adjacent nucleotides, or may be separated into two or more separate subsequences separated by one or more nucleotides. In some embodiments, the separate barcode subsequences may be from about 4 to about 16 nucleotides in length. In some embodiments, the barcode subsequences may be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequences may be at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequences may be up to about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.

[0123] The co-compartmentalized nucleic acid molecules can also contain other functional sequences useful for processing nucleic acids from the co-compartmentalized samples. These sequences include, for example, targeting or random / universal amplification primer sequences for amplifying genomic DNA from individual samples within the compartments while binding associated barcode sequences, sequencing primers or primer recognition sites, hybridization or probe sequences for, for example, identifying the presence of a sequence or for pulling down barcode nucleic acids, or any of many other potential functional sequences. Other mechanisms for co-compartmentalizing oligonucleotides may also be employed, including, for example, combining two or more compartments, one of which contains an oligonucleotide, or microdispersing oligonucleotides into compartments, e.g., in a microfluidic system. In some embodiments, the primer comprises a barcode oligonucleotide. In some embodiments, the primer sequence is a targeting primer sequence complementary to a sequence within the template nucleic acid molecule. In some embodiments, the first nucleic acid molecule further comprises one or more functional sequences, and the second nucleic acid molecule comprises one or more functional sequences. In some embodiments, the one or more functional sequences are selected from the group consisting of an adapter sequence, an additional primer sequence, a primer annealing sequence, a sequencing primer sequence, a sequence configured to bind to a flow cell of a sequencer, and a unique molecular identifier sequence.

[0124] For example, the barcoded nucleic acid molecules (e.g., barcoded oligonucleotides) described above are added to a sample. In some embodiments, the compartments contain barcoded oligonucleotides having the same barcode sequence. In some embodiments, one compartment of the plurality of compartments contains barcoded oligonucleotides having the same barcode sequence, and each compartment within the plurality of compartments contains a unique barcode sequence. In some embodiments, the population of barcoded oligonucleotides provides a diverse barcode sequence library containing at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences, or more. Furthermore, each barcoded oligonucleotide can be provided with many nucleic acid (e.g., oligonucleotide) molecules attached thereto. In particular, the number of nucleic acid molecules comprising a barcode sequence on an individual barcoded oligonucleotide can be at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acids, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules, and in some embodiments, at least about 1 billion nucleic acid molecules, or more. The nucleic acid molecules of a given barcoded oligonucleotide can comprise identical (or common) barcode sequences, different barcode sequences, or a combination of both. The nucleic acid molecules of a given barcoded oligonucleotide can include multiple sets of nucleic acid molecules.The nucleic acid molecules of a given set can include the same barcode sequence, which can be different from the barcode sequence of the nucleic acid molecules of another set.

[0125] Furthermore, when the population of barcoded oligonucleotides is partitioned, the resulting population of partitions can further comprise a diverse barcode library comprising at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences. Further, each partition of the population can comprise at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acids, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules, and in some embodiments, at least about 1 billion nucleic acid molecules.

[0126] In some embodiments, it may be desirable to incorporate multiple different barcodes within a given compartment. For example, in some embodiments, the barcoded oligonucleotides within a compartment may include (1) a common barcode sequence shared by all barcoded oligonucleotides within the compartment, and (2) a unique molecular identifier or additional barcode sequence that differs among each barcoded oligonucleotide. The common barcode sequence may provide a stronger address or attribution of the barcode to a given compartment, for example, as a redundant or independent confirmation of the output from a given compartment, thereby making identification more reliable in subsequent processing.

[0127] In some embodiments, barcoded oligonucleotides are attached to beads, where all of the nucleic acid molecules attached to a particular bead contain the same nucleic acid barcode sequence, but many different barcode sequences are represented across the population of beads used. In some embodiments, for example, hydrogel beads comprising a polyacrylamide polymer matrix can carry many nucleic acid molecules and are therefore used as delivery vehicles for nucleic acid molecules to solid supports and compartments, and may be configured to release those nucleic acid molecules upon exposure to a particular stimulus, as described elsewhere herein.

[0128] Nucleic acid molecules (e.g., oligonucleotides) can be releasable from beads upon application of a specific stimulus to the beads. In some embodiments, the stimulus can be a light stimulus, e.g., by cleavage of a light-sensitive linkage to release the nucleic acid molecule. In other embodiments, a thermal stimulus can be used, where an increase in the temperature of the bead environment causes cleavage of the linkage or other release of the nucleic acid molecule to form the bead. In still other embodiments, a chemical stimulus can be used that cleaves the linkage of the nucleic acid molecule to the bead or otherwise results in release of the nucleic acid molecule from the bead. In one embodiment, such a composition comprises the polyacrylamide matrix described above for sample encapsulation, which can be degraded to release the bound nucleic acid molecule via exposure to a reducing agent such as DTT.

[0129] Supports that may be contemplated for use in the methods of the present disclosure may be, for example, wells, matrices, rods, containers, or beads. Supports may have any useful properties and characteristics, such as any useful size, surface chemistry, fluidity, solidity, density, porosity, and composition. In some embodiments, the support is the surface of a well on a plate. In some embodiments, the support may be beads, such as gel beads. Beads may be solid or semi-solid. Additional details of beads are provided elsewhere herein.

[0130] The support (e.g., a bead) may include an anchor sequence (e.g., as described herein) functionalized thereon. The anchor sequence may be attached to the support, for example, via a disulfide bond. The anchor sequence may include a partial read sequence and / or a flow cell functional sequence. Such a sequence may enable sequencing of the nucleic acid molecule attached to the sequence by a sequencer (e.g., an Illumina sequencer). Various anchor sequences may be useful for various sequencing applications. The anchor sequence may include, for example, a TruSeq or Nextera sequence. The anchor sequence may have any useful characteristics, such as any useful length and nucleotide composition. For example, the anchor sequence may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, the anchor sequence may include 15 nucleotides. The nucleotides of the anchor sequence may be naturally occurring or non-naturally occurring (e.g., as described herein). A bead may include multiple anchor sequences attached thereto. For example, a bead may include multiple first anchor sequences attached thereto. In some embodiments, a bead may include two or more different anchor sequences attached thereto. For example, a bead may include both multiple first anchor sequences (e.g., Nextera sequences) and multiple second anchor sequences (e.g., TruSeq sequences) attached thereto. For beads including two or more different anchor sequences attached thereto, the sequence of each different anchor sequence may be distinguishable from the sequence of each other anchor sequence at the end distal to the bead. For example, the different anchor sequences may include one or more nucleotide differences in the 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides furthest from the bead.

[0131] In some embodiments, multiple different barcode molecules (e.g., nucleic acid barcode molecules) may be generated on the same support (e.g., beads). For example, two different barcode molecules may be generated on the same support. Alternatively, three or more different barcode molecules may be generated on the same support. Different barcode molecules attached to the same support may contain one or more different sequences. For example, different barcode molecules may contain one or more different barcode sequences and / or other sequences (e.g., starter sequences). In some embodiments, different barcode molecules attached to the same support may contain the same barcode sequence. Different barcode molecules attached to the same support may contain the same or different barcode sequences. Similarly, different barcode molecules may contain the same or different unique molecular identifiers (UMIs).

[0132] Next-generation sequencing As described in the methods disclosed herein, sequencing of nucleic acid molecules is used and is useful for detecting and diagnosing pathogenic infectious diseases. Generally, sequencing refers to methods and techniques for determining the sequence of nucleotide bases in one or more polynucleotides. Polynucleotides can be, for example, deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single-stranded DNA). Sequencing can be performed using a variety of currently available systems, including, but not limited to, sequencing systems from Illumina, Pacific Biosciences, Oxford Nanopore, or Life Technologies (Ion Torrent). Such devices can provide multiple raw genetic data corresponding to a subject's (e.g., human) genetic information generated by the device from a sample provided by the subject. In some situations, the systems and methods provided herein can be used in conjunction with proteomic information. Alternatively or additionally, sequencing can be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real-time PCR), or isothermal amplification. Such a system may provide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., a human) generated by the system from a sample provided by the subject. In some examples, such a system provides sequencing reads (also referred to herein as "reads"). A read may include a string of nucleic acid bases corresponding to the sequence of a sequenced nucleic acid molecule. In some situations, the systems and methods provided herein may be used with proteomic information.

[0133] Next-generation sequencing encompasses many techniques capable of generating large amounts of sequence information, excluding Sanger sequencing or Maxam-Gilbert sequencing. Generally, next-generation sequencing encompasses single-molecule real-time sequencing, sequencing-by-synthesis, ion-semiconductor sequencing, and the like. Exemplary next-generation sequencing machines include Illumina, Inc.'s MiniSeq, iSeq100, NextSeq 1000, NextSeq 2000, NovaSeq 6000, and NextSeq 550 series, Ion Torrent machines from Thermo Fisher Scientific, and the Sequel system from Pacific Biosciences.

[0134] Next generation sequencing machines used with the methods described herein can generate at least 1, 5, 10, 15, 25, 50, 75, 100, 200, 300 gigabits of data or more from a single machine within 24 hours.

[0135] Next generation sequencing machines used with the methods described herein can generate data for at least 1, 1, 4, 10, 15, 25, 50, 75, 100, 200, 300, 500, or 1 billion sequence reads or more from a single machine within 24 hours.

[0136] Also included are computer programs, computing devices, or analysis platforms / systems for receiving and analyzing sequencing data and outputting one or more reports that can be transmitted or accessed electronically via a server, analysis portal, or email. The computing devices or analysis platforms can operate according to the algorithms and methods described herein.

[0137] Reaction mixture Also provided herein is a reaction mixture for determining the presence or absence of viral nucleic acid in a biological sample. In some embodiments, the reaction mixture comprises a synthetic nucleic acid provided herein, at least a portion of the biological sample, and one or more enzymes or reagents sufficient to amplify the viral nucleic acid in the biological sample, if present.

[0138] The biological sample may be any of the biological samples provided herein. In some embodiments, the biological sample is a human biological sample. In some embodiments, the biological sample comprises saliva, a buccal swab, a nasopharyngeal swab, or a middle turbinate swab. In some embodiments, the biological sample comprises saliva or a nasopharyngeal swab. In some embodiments, the biological sample comprises saliva. In some embodiments, the biological sample comprises a nasopharyngeal swab.

[0139] The synthetic nucleic acid may be any of the synthetic nucleic acids provided herein. In some embodiments, the synthetic nucleic acid is a SARS-CoV-2 synthetic RNA nucleic acid provided herein. In some embodiments, the SARS-CoV-2 synthetic RNA nucleic acid is present at a concentration of about 10 copies per reaction mixture to about 500 copies per reaction mixture. In some embodiments, the SARS-CoV-2 synthetic RNA nucleic acids are present at a concentration of about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, or 500 copies per reaction mixture.

[0140] The viral nucleic acid may be any of the viral nucleic acids provided herein. In some embodiments, the viral nucleic acid is influenza A nucleic acid, influenza B nucleic acid, or coronavirus nucleic acid. In some embodiments, the coronavirus nucleic acid is Sars-Cov-2 nucleic acid.

[0141] In some embodiments, the enzymes or reagents include reverse transcriptase, dNTPs, a primer pair specific to the viral nucleotide sequence, a primer pair specific to a sample control nucleotide sequence, magnesium salts, or a combination thereof.

[0142] In some embodiments, the reaction mixture includes one or more enzymes that can be used to amplify viral nucleic acids. In some embodiments, the enzyme is a reverse transcriptase. Non-limiting examples of reverse transcriptases include, but are not limited to, avian myeloblastosis virus (AMV) reverse transcriptase and Moloney murine leukemia virus (M-MuLV, MMLV), and variants thereof.

[0143] In some embodiments, the reaction includes deoxynucleotide triphosphates (dNTPs). In some embodiments, the kit includes a mixture of each of the dNTPs required for amplification of viral nucleic acids, as well as other desired nucleic acids (e.g., dATG, dCTP, dTTP, dGTP).

[0144] In some embodiments, the reaction mixture comprises a primer pair specific to a viral nucleotide sequence. The primer pair specific to a viral nucleotide sequence can be any of the primer pairs provided herein. The primer pair specific to the viral nucleotide sequence is specific to an influenza A nucleotide sequence, an influenza B nucleotide sequence, or a coronavirus nucleotide sequence. In some embodiments, the primer pair specific to the viral nucleotide sequence is specific to a coronavirus S1 or N2 sequence.

[0145] In some embodiments, the reaction mixture comprises a primer pair specific to a sample control nucleotide sequence. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to an endogenous nucleotide sequence expressed by the organism from which the biological sample is derived. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to a housekeeping gene. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to GAPDH, RPP30, or ACTB. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to RPP30.

[0146] In some embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence is present at a concentration of about 50 micromolar to about 250 micromolar. In some embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence is present at a concentration of about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 micromolar. In some embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence is present at a concentration of about 100 micromolar. In some embodiments, the primer pair specific for the viral nucleotide sequence or the primer pair specific for the sample control nucleotide sequence is present at a concentration of about 200 micromolar.

[0147] In some embodiments, the reaction mixture includes a magnesium salt. In some embodiments, the magnesium salt is included in an amount sufficient to allow the enzyme (e.g., reverse transcriptase) of the reaction to function and amplify the target nucleic acid. In some embodiments, the magnesium salt is magnesium chloride. In some embodiments, the reaction mixture includes a magnesium ion concentration of about 0.1 mM to about 50 mM. In some embodiments, the magnesium ion concentration is about 1 mM to about 10 mM.

[0148] In certain embodiments, the volume of the reaction mixture is about 10 microliters to about 100 microliters. In some embodiments, the volume of the reaction mixture is about 20 microliters to about 90 microliters, about 30 microliters to about 80 microliters, or about 40 microliters to about 60 microliters. In some embodiments, the volume of the reaction mixture is about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 microliters.

[0149] kit Kits containing oligonucleotide primer pairs are also described herein. In certain embodiments, kits include the oligonucleotide primer pairs described herein and synthetic nucleic acids according to the present disclosure. The synthetic nucleic acids may be for S1 and N2 amplification. The kits may also include sufficient reagents for amplification, such as buffers (provided at 10x, 5x, 2x, or 1x concentrations) for reverse transcription, PCR amplification, and / or sequencing reactions, dNTPs, reverse transcriptase, and PCR amplification enzymes (e.g., Taq polymerase and / or its variants). These components may be packaged individually or in combination as needed in vials or containers.

[0150] Also described herein is a kit for determining the presence or absence of viral nucleic acid in a biological sample. In some embodiments, the kit includes a synthetic nucleic acid provided herein and one or more enzymes or reagents sufficient to amplify the viral nucleic acid to form a biological sample.

[0151] In some embodiments, the enzymes or reagents include reverse transcriptase, dNTPs, a primer pair specific to the viral nucleotide sequence, a primer pair specific to a sample control nucleotide sequence, magnesium salts, or a combination thereof.

[0152] In some embodiments, the kit includes one or more enzymes that can be used to amplify viral nucleic acids. In some embodiments, the enzyme is a reverse transcriptase. Non-limiting examples of reverse transcriptases include, but are not limited to, avian myeloblastosis virus (AMV) reverse transcriptase and Moloney murine leukemia virus (M-MuLV, MMLV), and variants thereof.

[0153] In some embodiments, the kit includes deoxynucleotide triphosphates (dNTPs). In some embodiments, the kit includes a mixture of each of the dNTPs required for amplification of viral nucleic acids, as well as other desired nucleic acids (e.g., dATG, dCTP, dTTP, dGTP).

[0154] In some embodiments, the kit includes a primer pair specific to a viral nucleotide sequence. The primer pair specific to a viral nucleotide sequence can be any of the primer pairs provided herein. The primer pair specific to the viral nucleotide sequence is specific to an influenza A nucleotide sequence, an influenza B nucleotide sequence, or a coronavirus nucleotide sequence. In some embodiments, the primer pair specific to the viral nucleotide sequence is specific to a coronavirus S1 or N2 sequence.

[0155] In some embodiments, the kit includes a primer pair specific to a sample control nucleotide sequence. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to an endogenous nucleotide sequence expressed by the organism from which the biological sample is derived. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to a housekeeping gene. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to GAPDH, RPP30, or ACTB. In some embodiments, the primer pair specific to the sample control nucleotide sequence is specific to RPP30.

[0156] In some embodiments, the kit includes a magnesium salt. In some embodiments, the magnesium salt is included in an amount sufficient to allow an enzyme (e.g., reverse transcriptase) of the kit to function. In some embodiments, the magnesium salt is magnesium chloride.

[0157] The following embodiments detail non-limiting permutations of combinations of features disclosed herein. Other permutations of combinations of features are also contemplated. In particular, each of these numbered embodiments is contemplated as dependent or related to the preceding or following numbered embodiment, regardless of the order in which they are listed.

[0158] 1. A method of diagnosing an individual with a pathogen infection, comprising: (a) providing a biological sample from the individual; (b) contacting the biological sample from the individual with a lysing agent to obtain a lysed biological sample; (c) performing a polymerase chain reaction (PCR) on the lysed biological sample to obtain a PCR-amplified lysed biological sample, wherein the PCR reaction on the lysed biological sample is performed using a first set of PCR primers, and the first set of PCR primers amplify a pathogen nucleic acid sequence; (d) sequencing the PCR-amplified lysed biological sample using next-generation sequencing; and (e) providing a positive diagnosis for the pathogen infection if a pathogen sequence is detected by the PCR or the sequencing, or providing a negative diagnosis for the individual if a pathogen sequence is not detected by the PCR or the sequencing. 2. The method of embodiment 1, wherein the individual is a human individual. 3. The method of embodiment 1 or 2, wherein the pathogen infection comprises a bacterial infection, a viral infection, or a fungal infection, and combinations thereof. 4. The method of embodiment 1-3, wherein the lysis agent comprises water heated to at least 50°C or a PCR reaction buffer. 5. The method of embodiment 1-3, wherein the lysis agent comprises water heated to at least 90°C or a PCR reaction buffer. 6. The method of embodiment 1-5, wherein the bacterial infection is an infection caused by Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof. 7. The method of embodiment 1-5, wherein the fungal infection is an infection caused by Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. 8. The method according to any one of embodiments 1 to 5, wherein the viral infection is an infection caused by a DNA virus.9. The method of embodiment 8, wherein the DNA virus comprises hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or smallpox, and combinations thereof. 10. The method of any one of embodiments 1 to 5, wherein the viral infection is an infection caused by an RNA virus. 11. The method of embodiment 10, wherein the RNA virus comprises influenza virus, coronavirus, poliovirus, measles virus, Ebola virus, retrovirus, or orthomyxovirus. 12. The method of embodiment 11, wherein the viral infection is a coronavirus infection. 13. The method of embodiment 12, wherein the coronavirus infection is a SARS-COV-2 infection. 14. The method of any one of embodiments 1 to 13, wherein the biological sample from the individual is a blood sample, plasma sample, serum sample, buccal swab, urine sample, semen sample, vaginal swab, stool sample, nasopharyngeal swab, middle turbinate swab, or any combination thereof. 15. The method of embodiment 14, wherein the biological sample from the individual is derived from a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. 16. The method of any one of embodiments 1 to 15, comprising adding a synthetic nucleic acid to the lysis agent or the lysed biological sample. 17. The method of any one of embodiments 1 to 15, comprising adding a plurality of synthetic nucleic acids to the lysis agent or the lysed biological sample, wherein the plurality of synthetic nucleic acids comprises synthetic nucleic acids having different sequences. 18. The method of embodiment 17, wherein the plurality of synthetic nucleic acids comprises four. 19. The method of any one of embodiments 1 to 18, comprising adding a synthetic nucleic acid to the method. 20. The method of any one of embodiments 1 to 18, comprising adding a plurality of synthetic nucleic acids to the method, wherein the plurality of synthetic nucleic acids comprises synthetic nucleic acids having different sequences. 21. The method of embodiment 20, wherein the plurality of synthetic nucleic acids comprises four. 22. The method of embodiment 19, wherein the synthetic nucleic acid is RNA. 23. The method of embodiment 19, wherein the synthetic nucleic acid is DNA. 24. The method of any one of embodiments 19 or 23, wherein the synthetic nucleic acids comprise a set of sequences configured to be bound by the first set of PCR primers.25. The method of any one of embodiments 19 to 24, wherein the first set of primers amplifies both the pathogen nucleic acid sequence and the synthetic nucleic acid. 26. The method of any one of embodiments 19 to 24, wherein the synthetic nucleic acid sequence comprises a nucleotide sequence that is not identical to the pathogen nucleic acid sequence. 27. The method of any one of embodiments 1 to 26, comprising performing a reverse transcription reaction on the lysed biological sample. 28. The method of embodiment 27, wherein the reverse transcription reaction is performed before performing the polymerase chain reaction. 29. The method of any one of embodiments 1 to 26, wherein the reverse transcription reaction is performed without further purification of the lysed biological sample. 30. The method of any one of embodiments 1 to 26, wherein the reverse transcription reaction on the lysed biological sample produces viral cDNA. 31. The method of embodiment 30, wherein the viral cDNA is coronavirus cDNA. 32. The method of embodiment 31, wherein the coronavirus cDNA is SARS-COV-2 cDNA. 33. The method of any one of embodiments 1 to 31, wherein the reverse transcription reaction and the PCR are a single-step reaction. 34. The method of any one of embodiments 1 to 33, wherein the PCR is an end-point analysis. 35. The method of any one of embodiments 1 to 34, wherein the PCR is not a real-time PCR reaction. 36. The method of any one of embodiments 1 to 35, wherein the first set of PCR primers amplifies a coronavirus nucleic acid sequence. 37. The method of embodiment 36, wherein the coronavirus nucleic acid sequence is a SARS-COV-2 nucleic acid sequence. 38. The method of embodiment 37, wherein the SARS-COV-2 nucleic acid sequence comprises an N1 gene or an S2 gene. 39. The method of any one of embodiments 1 to 38, comprising a second set of PCR primers, wherein the second set of primers amplifies a nucleic acid sequence of the individual. 40. The method of embodiment 39, wherein the second set of PCR primers amplifies a human nucleic acid sequence. 41. The method of embodiment 40, wherein the second set of PCR primers amplifies a human nucleic acid sequence selected from GAPDH, ACTB, RPP30, and combinations thereof.42. The method of embodiment 41, wherein the second set of PCR primers amplifies human RPP30. 43. The method of any one of embodiments 40-42, wherein the second set of PCR primers comprises a mixture of primers with sequencing adapter sequences and primers without sequencing adapter sequences. 44. The method of embodiment 43, wherein the ratio of primers with sequencing adapter sequences to primers without sequencing adapter sequences is about 1:1, about 1:2, about 1:3, or about 1:4. 45. The method of any one of embodiments 1-44, wherein the PCR comprises 30 to 45 amplification cycles. 46. The method of embodiment 45, wherein the PCR comprises 35 to 45 amplification cycles. 47. The method of embodiment 45, wherein the PCR comprises 39 to 42 amplification cycles. 48. The method of any one of embodiments 1 to 47, wherein the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise nucleic acid sequences comprising a variable nucleotide sequence. 49. The method of embodiment 48, wherein the variable nucleotide sequence is a sample ID unique to the individual. 50. The method of any one of embodiments 1 to 49, wherein the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise adapter sequences for next-generation sequencing reactions. 51. The method of any one of embodiments 39-50, comprising adding a second synthetic nucleic acid to the method. 52. The method of embodiment 51, wherein the second synthetic nucleic acid is RNA. 53. The method of embodiment 51, wherein the second synthetic nucleic acid is DNA. 54. The method of any one of embodiments 51-53, wherein the second synthetic nucleic acid comprises a set of sequences configured to be bound by the second set of PCR primers. 55. The method of any one of embodiments 51-54, wherein the second set of primers amplifies both the human nucleic acid sequence and the synthetic nucleic acid. 56. The method of any one of embodiments 51-54, wherein the synthetic nucleic acid sequence comprises a nucleotide sequence that is not identical to the human nucleic acid sequence.57. The method of any one of embodiments 1 to 56, wherein the method is capable of detecting fewer than 10 copies of the pathogen genome. 58. The method of any one of embodiments 1 to 56, wherein the method is capable of detecting fewer than 5 copies of the pathogen genome. 59. The method of embodiment 57 or 58, wherein the pathogen genome is a coronavirus genome. 60. The method of embodiment 57 or 58, wherein the coronavirus genome is a SARS-COV-2 genome. 61. The method of any one of embodiments 1 to 60, wherein if a coronavirus sequence is detected by the PCR, the positive diagnosis of coronavirus is a SARS-COV-2 diagnosis. 62. The method of any one of embodiments 1 to 61, wherein the method determines the strain of coronavirus. 63. The method of any one of embodiments 1 to 61, wherein the method determines the strain of COVID-19.

[0159] 64. A method for diagnosing an individual with a pathogen infection, the method comprising: amplifying nucleic acid from a biological sample from the individual using a first set of PCR primers to obtain amplified nucleic acid, wherein the first set of PCR primers amplify a pathogen nucleic acid sequence and a synthetic nucleic acid sequence from the biological sample, and the synthetic nucleic acid sequence differs from the pathogen nucleic acid sequence by at least one nucleotide. 65. The method of embodiment 64, wherein the individual is a human individual. 66. The method of embodiment 64 or 65, wherein the pathogen infection comprises a bacterial infection, a viral infection, or a fungal infection. 67. The method of any one of embodiments 64 to 66, wherein the bacterial infection is an infection caused by Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, or a combination thereof. 68. The method of any one of embodiments 64 to 66, wherein the fungal infection is an infection caused by Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. 69. The method of any one of embodiments 64 to 66, wherein the viral infection is an infection caused by a DNA virus. 70. The method of embodiment 69, wherein the DNA virus includes hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, smallpox, and any combination thereof. 71. The method of any one of embodiments 64 to 66, wherein the viral infection is an infection caused by an RNA virus. 72. The method of embodiment 71, wherein the RNA virus includes influenza virus, coronavirus, poliovirus, measles virus, Ebola virus, retrovirus, or orthomyxovirus. 73. The method of embodiment 71, wherein the viral infection is a coronavirus infection. 74. The method of embodiment 73, wherein the coronavirus infection is a SARS-COV-2 infection.75. The method of any one of embodiments 64-74, wherein the biological sample from the individual is a blood sample, a plasma sample, a serum sample, a buccal swab, a urine sample, a semen sample, a vaginal swab, a stool sample, a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. 76. The method of embodiment 75, wherein the biological sample from the individual is derived from a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. 77. The method of any one of embodiments 64-76, wherein the synthetic nucleic acid is RNA. 78. The method of any one of embodiments 64-76, wherein the synthetic nucleic acid is DNA. 79. The method of any one of embodiments 64-78, wherein the synthetic nucleic acid comprises a set of sequences configured to be bound by the first set of PCR primers. 80. The method of any one of embodiments 64-79, wherein the synthetic nucleic acid sequence differs from the pathogen nucleic acid sequence by at least five nucleotides. 81. The method of any one of embodiments 64 to 79, wherein the pathogen infection is diagnosed based on the ratio of a synthetic nucleic acid sequence to a pathogen nucleic acid sequence. 82. The method of any one of embodiments 64 to 81, comprising performing a reverse transcription reaction on the nucleic acid from the biological sample. 83. The method of any one of embodiments 64 to 81, wherein the reverse transcription reaction on the nucleic acid from the biological sample produces coronavirus cDNA. 84. The method of embodiment 83, wherein the coronavirus cDNA is SARS-COV-2 cDNA. 85. The method of any one of embodiments 64 to 84, wherein the amplified nucleic acid comprises a PCR reaction. 86. The method of embodiment 85, wherein the PCR reaction is an end-point analysis. 87. The method of any one of embodiments 64 to 86, wherein the PCR reaction is not a real-time PCR reaction. 88. The method of any one of embodiments 64 to 87, wherein the first set of PCR primers amplifies a coronavirus nucleic acid sequence. 89. The method of embodiment 88, wherein the coronavirus nucleic acid sequence is a SARS-COV-2 nucleic acid sequence. 90. The method of embodiment 89, wherein the SARS-COV-2 nucleic acid sequence comprises the N1 gene or the S2 gene.91. The method of any one of embodiments 64 to 90, wherein the first set of PCR primers comprises a nucleic acid sequence comprising a variable nucleotide sequence. 92. The method of embodiment 91, wherein the variable nucleotide sequence is a sample ID unique to the individual. 93. The method of any one of embodiments 64 to 92, wherein the first set of PCR primers comprises an adapter sequence for a next-generation sequencing reaction. 94. The method of any one of embodiments 64 to 92, comprising amplifying nucleic acids from the biological sample using a second set of PCR primers, wherein the second set of PCR primers amplifies a human nucleic acid sequence. 95. The method of embodiment 94, wherein the second set of PCR primers amplifies a nucleic acid sequence selected from GAPDH, ACTB, RPP30, and combinations thereof. 96. The method of embodiment 95, wherein the second set of PCR primers amplifies human RPP30. 97. The method of any one of embodiments 94 to 96, wherein the second set of PCR primers comprises a mixture of primers with sequencing adapter sequences and primers without sequencing adapter sequences. 98. The method of embodiment 97, wherein the ratio of primers with sequencing adapter sequences to primers without sequencing adapter sequences is about 1:1, about 1:2, about 1:3, or about 1:4. 99. The method of any one of embodiments 64 to 98, wherein the PCR comprises 30 to 45 amplification cycles. 100. The method of embodiment 99, wherein the PCR comprises 35 to 45 amplification cycles. 101. The method of embodiment 99, wherein the PCR comprises 39 to 42 amplification cycles. 102. The method of any one of embodiments 94 to 101, wherein the second set of PCR primers comprises a nucleic acid sequence comprising a variable nucleotide sequence. 103. The method of embodiment 102, wherein the variable nucleotide sequence is a sample ID unique to the individual. 104. The method of any one of embodiments 94 to 103, wherein the second set of PCR primers comprises an adapter sequence for a next-generation sequencing reaction.105. The method of any one of embodiments 64 to 104, comprising sequencing the amplified nucleic acid from the biological sample using next-generation sequencing technology. 106. The method of any one of embodiments 64 to 105, wherein the method is capable of detecting fewer than 10 copies of a pathogen genome. 107. The method of any one of embodiments 64 to 104, wherein the method is capable of detecting fewer than 5 copies of a pathogen genome. 108. The method of embodiment 106 or 107, wherein the pathogen genome is a coronavirus genome. 109. The method of embodiment 106 or 107, wherein the pathogen genome is a SARS-COV-2 genome. 110. The method of any one of embodiments 64 to 109, wherein the method determines the strain of a coronavirus. 111. The method of embodiment 110, wherein the method determines the strain of SARS-COV-2.

[0160] 112. A synthetic nucleic acid comprising a 5'-proximal region, a 3'-proximal region, and an intervening nucleic acid sequence. 113. The synthetic nucleic acid of embodiment 112, wherein the synthetic nucleic acid comprises RNA. 114. The synthetic nucleic acid of embodiment 112, wherein the synthetic nucleic acid comprises DNA. 115. The synthetic nucleic acid of any one of embodiments 112-114, wherein the 5'-proximal region comprises a viral nucleic acid sequence. 116. The method of embodiment 115, wherein the viral nucleic acid sequence comprises a coronavirus sequence. 117. The method of embodiment 116, wherein the viral nucleic acid sequence comprises a SARS-COV-2 sequence. 118. The synthetic nucleic acid of any one of embodiments 112-114, wherein the 3'-proximal region comprises a viral nucleic acid sequence. 119. The method of embodiment 118, wherein the viral nucleic acid sequence comprises a coronavirus sequence. 120. The method of embodiment 119, wherein the viral nucleic acid sequence comprises a SARS-COV-2 sequence. 121. The synthetic nucleic acid of any one of embodiments 112-120, wherein the 5'-proximal region, the 3'-proximal region, or both the 5'-proximal region and the 3'-proximal region are less than about 30 nucleotides in length. 122. The synthetic nucleic acid of any one of embodiments 112-120, wherein the 5'-proximal region, the 3'-proximal region, or both the 5'-proximal region and the 3'-proximal region are less than about 25 nucleotides in length. 123. The synthetic nucleic acid of any one of embodiments 112-120, wherein the 5'-proximal region, the 3'-proximal region, or both the 5'-proximal region and the 3'-proximal region are less than about 20 nucleotides in length. 124. The synthetic nucleic acid of any one of embodiments 112-123, wherein the 5'-proximal region is at the 5'-end of the synthetic nucleic acid. 125. The synthetic nucleic acid of any one of embodiments 112 to 123, wherein the 3'-proximal region is at the 3'-end of the synthetic nucleic acid. 126. The synthetic nucleic acid of any one of embodiments 112 to 125, wherein the intervening nucleic acid sequence is less than about 99%, 98%, 97%, 95%, 90%, 85%, 80%, or 75% identical to a viral nucleic acid sequence. 127. The synthetic nucleic acid of embodiment 126, wherein the synthetic nucleic acid sequence is a coronavirus sequence. 128. The synthetic nucleic acid of embodiment 126, wherein the synthetic nucleic acid sequence is a SARS-COV-2 sequence.129. Use of the synthetic nucleic acid of any one of embodiments 112 to 128 in a method for detecting a pathogen infection in said individual. 130. The use of embodiment 129, wherein said pathogen infection is a coronavirus infection. 131. The use of embodiment 130, wherein said viral infection is a SARS-COV-2 infection.

[0161] 132. A method for nucleic acid processing for detecting a viral infection, comprising: (a) providing a sample containing a viral nucleic acid molecule and a host nucleic acid molecule; (b) performing a nucleic acid extension reaction on the viral nucleic acid molecule using a first primer comprising a first barcode sequence to generate a barcoded viral nucleic acid molecule; (c) performing a nucleic acid extension reaction on the host nucleic acid molecule using a second primer comprising a second barcode sequence to generate a barcoded host nucleic acid molecule; (d) sequencing the barcoded viral nucleic acid molecule and the barcoded host nucleic acid molecule to identify (i) the barcode sequence and (ii) a sequence corresponding to the viral nucleic acid molecule or a derivative thereof, and the host nucleic acid molecule; and (e) providing a positive diagnosis of viral infection if the sequence corresponding to the viral nucleic acid molecule is identified in (d). 133. The method of embodiment 132, wherein (b) and (c) are performed simultaneously. 134. The method of any one of embodiments 132-133, wherein the first primer further comprises one or more additional functional sequences selected from the group consisting of a primer sequence, an adapter sequence, a primer annealing sequence, a unique molecular identifier sequence, and a capture sequence. 135. The method of any one of embodiments 132-134, wherein the second primer further comprises one or more additional functional sequences selected from the group consisting of a primer sequence, an adapter sequence, a primer annealing sequence, a unique molecular identifier sequence, and a capture sequence. 136. The method of any one of embodiments 132-135, wherein (a) further comprises providing a synthetic nucleic acid molecule. 137. The method of embodiment 136, wherein the synthetic nucleic acid molecule comprises a synthetic sequence that is different from the viral nucleic acid molecule and the human nucleic acid molecule. 138. The method of embodiment 137, wherein (b) performs the nucleic acid extension using the first primer to generate a barcoded synthetic nucleic acid molecule. 139. The method of embodiment 137, wherein (c) uses the second primer to perform the nucleic acid extension reaction, thereby generating a barcoded synthetic nucleic acid molecule. 140. The method of any one of embodiments 132 to 139, wherein the nucleic acid extension reaction is a reverse transcription reaction.141. The method of any one of embodiments 132-139, wherein the nucleic acid extension reaction is a polymerase chain reaction. 142. The method of any one of embodiments 132-139, wherein the nucleic acid extension reaction comprises a reverse transcriptase reaction, a polymerase chain reaction, or a combination thereof. 143. The method of embodiment 142, wherein in (b), the nucleic acid extension reaction comprises (i) hybridizing the first primer to the viral nucleic acid molecule and (ii) using a reverse transcriptase enzyme to extend the primer. 144. The method of embodiment 142, wherein in (b), the nucleic acid extension reaction comprises (i) hybridizing the first primer to the viral nucleic acid molecule and (ii) using a reverse transcriptase enzyme to extend the primer. 145. The method of embodiment 142, wherein in (b), the nucleic acid extension reaction comprises (i) hybridizing the first primer to the viral nucleic acid molecule and (ii) using a polymerase enzyme to extend the primer. 146. The method of embodiment 142, wherein the polymerase chain reaction is an end-point polymerase chain reaction. 147. The method of any one of embodiments 132 to 146, wherein, prior to (d), the method further comprises amplifying the barcoded viral nucleic acid molecule and the barcoded host nucleic acid molecule. 148. The method of any one of embodiments 132 to 147, wherein, prior to (d), the method further comprises subjecting the barcoded viral nucleic acid molecule and the barcoded host nucleic acid molecule to N cycles of polymerase chain reaction, 149. The method of embodiment 148, wherein N is greater than 35 cycles. 150. The method of embodiment 148, wherein N is greater than 40 cycles. 151. The method of embodiment 148, wherein N is greater than 45 cycles. 152. The method of embodiment 148, wherein the polymerase chain reaction incorporates one or more additional sequences selected from the group consisting of a sample index sequence, an adapter sequence, a primer sequence, a primer binding sequence, a sequence configured to bind to a flow cell of a sequencer, and an additional barcode sequence into one or both of the barcoded viral nucleic acid molecule and the barcoded host nucleic acid molecule.153. The method of any one of embodiments 132-152, wherein after (b) and (c), the sample is combined with one or more samples after performing at least one nucleic acid extension reaction in (b) and (c). 154. The method of embodiment 153, wherein the sample is combined with one or more samples after performing one round of the nucleic acid extension reaction in (b) and (c). 155. The method of any one of embodiments 132-154, wherein the sample comprises one or more cells. 156. The method of embodiment 155, further comprising, before (b), releasing the viral nucleic acid molecules and the host nucleic acid molecules from the one or more cells. 157. The method of embodiment 155, further comprising, before (b), subjecting the sample to conditions sufficient to release the viral nucleic acid molecules and the host nucleic acid molecules from the one or more cells. 158. The method of embodiment 157, wherein (b) and (c) are performed under conditions sufficient to release the viral nucleic acid molecule and the host nucleic acid molecule from the one or more cells. 159. The method of embodiment 157, wherein the viral nucleic acid molecule and the host nucleic acid molecule are not purified from the sample prior to (b) and (c). 160. The method of any one of embodiments 132-159, further comprising using the barcode sequence in (d) to associate the viral nucleic acid molecule, or a derivative thereof, and the host nucleic acid molecule, or a derivative thereof, as associated with the sample. 161. The method of any one of embodiments 132-160, further comprising using the barcoded viral nucleic acid molecule and the barcoded host nucleic acid molecule in (e) to identify the sample, wherein the sample corresponds to a subject being tested for the viral infection. 162. The method of any one of embodiments 132-161, wherein the sample is obtained from a subject. 163. The method of embodiment 162, wherein the subject is a human. 164. The method of embodiment 162, wherein the subject is an animal. 165. The method of embodiment 162, wherein the host nucleic acid molecule is part of the genome of the subject. 166. The method of embodiment 162, wherein the host nucleic acid molecule is part of the transcriptome of the subject.167. The method of any one of embodiments 132 to 166, wherein the host nucleic acid molecule encodes a uniquely expressed protein. 168. The method of any one of embodiments 132 to 166, wherein the host nucleic acid molecule is a genomic DNA molecule. 169. The method of any one of embodiments 132 to 166, wherein the host nucleic acid molecule is a uniquely transcribed RNA molecule. 170. The method of any one of embodiments 132 to 169, wherein the host nucleic acid is GAPDH, ACTB, RPP30, or a combination thereof. 171. The method of any one of embodiments 132 to 170, wherein the viral nucleic acid molecule is part of a viral genome. 172. The method of embodiment 171, wherein the virus is a coronavirus. 173. The method of embodiment 172, wherein the coronavirus is selected from the group consisting of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), severe acute respiratory syndrome coronavirus (SARS-CoV), and Middle East respiratory syndrome coronavirus (MERS-CoV). 174. The method of embodiment 173, wherein the coronavirus is SARS-COV-2. 175. The method of embodiment 174, wherein the virus is an RNA virus. 176. The method of embodiment 175, wherein the RNA virus comprises a double-stranded RNA genome. 177. The method of embodiment 175, wherein the RNA virus comprises a single-stranded RNA genome. 178. The method of embodiment 175, wherein the RNA virus is selected from the group consisting of coronavirus, influenza, human immunodeficiency virus, and Ebola virus. 179. The method of embodiment 170, wherein the virus is a DNA virus. 180. The method of any one of embodiments 34 to 169, further comprising, prior to (a), dividing the sample into compartments. 181. The method of embodiment 180, wherein the compartment is a well. 182. The method of embodiment 180, wherein the compartment is one well of a plurality of wells. 183. The method of embodiment 182, wherein the barcode sequence is unique to the well of the plurality of wells. 184. The method of any one of embodiments 132-183, wherein (b) and (c) are performed simultaneously.185. The method according to any one of embodiments 132 to 184, wherein the viral infection is SARS-COV-2. 186. The method according to any one of embodiments 132 to 183, wherein the sample is obtained from the subject by nasopharyngeal swab or middle turbinate swab.

[0162] 187. A composition comprising a synthetic nucleic acid molecule comprising a first nucleic acid sequence and a second nucleic acid sequence, wherein (1) the first nucleic acid sequence is identical to a sequence from a pathogen nucleic acid molecule, and (2) the second nucleic acid sequence is not identical to a sequence from a pathogen nucleic acid molecule. 188. The composition of embodiment 187, wherein the first nucleic acid sequence is located 3' of the second nucleic acid sequence. 189. The composition of any one of embodiments 187-188, wherein the synthetic nucleic acid molecule further comprises a third nucleic acid sequence, wherein the third nucleic acid sequence is identical to the second sequence from a pathogen nucleic acid molecule. 190. The composition of embodiment 176, wherein the third nucleic acid sequence is 5' of the second nucleic acid sequence. 191. The composition of any one of embodiments 187-190, wherein the first nucleic acid sequence or the third nucleic acid sequence is less than 5, 10, 15, 20, 25, or 30 nucleotides. 192. The composition of any one of embodiments 187-191, wherein the second nucleic acid sequence comprises a total number of nucleotides less than 25, 50, 100, 150, 200, or 500 nucleotides. 193. The composition of any one of embodiments 187-191, wherein the second nucleic acid sequence comprises a total number of nucleotides greater than 25, 50, 100, 150, 200, or 500 nucleotides. 194. The composition of any one of embodiments 187-193, wherein the synthetic nucleic acid molecule is a ribonucleic acid (RNA) molecule, a deoxyribonucleic acid (DNA) molecule, or an RNA-DNA hybrid molecule. 195. The composition of any one of embodiments 187-194, wherein the first nucleic acid sequence, the third nucleic acid sequence, or both the first and third nucleic acid sequences comprise a primer binding site. 196. The composition of any one of embodiments 187-194, wherein the composition further comprises a pathogen nucleic acid molecule. 197. The composition of embodiment 196, wherein the pathogen nucleic acid molecule is derived from a pathogen, and the pathogen comprises a bacterium, a virus, a fungus, or a combination thereof. 198. The composition of embodiment 197, wherein the bacterium is derived from the genus Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof.199. The composition of embodiment 197, wherein the fungus is Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. 200. The composition of embodiment 197, wherein the virus is a DNA virus. 201. The composition of embodiment 200, wherein the DNA virus comprises hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or smallpox, and combinations thereof. 202. The composition of embodiment 197, wherein the virus is an RNA virus. 203. The composition of embodiment 202, wherein the RNA virus comprises influenza virus, coronavirus, poliovirus, measles virus, Ebola virus, retrovirus, or orthomyxovirus. 204. The composition of embodiment 189, wherein the virus is a coronavirus. 205. The composition of embodiment 190, wherein the coronavirus is a SARS-COV-2 virus. 206. The composition of any one of embodiments 196-205, wherein the composition further comprises a plurality of primers, wherein one primer of the plurality of primers is configured to hybridize to a sequence of the synthetic nucleic acid molecule or a sequence of the pathogen nucleic acid molecule. 207. The composition of embodiment 206, wherein the sequence of the synthetic nucleic acid molecule and the sequence of the pathogen nucleic acid molecule are identical. 208. The composition of any one of embodiments 187-207, wherein the synthetic nucleic acid molecule is amplified with the same efficiency as the pathogen nucleic acid molecule. 209. The composition of any one of embodiments 187-207, wherein the synthetic nucleic acid molecule is configured to generate an amplification product that is the same size as, or within 10 base pairs of, the amplification product of the pathogen nucleic acid molecule.

[0163] 210. A method for diagnosing an individual with a pathogen infection, the method comprising: (a) providing a biological sample from the individual; (c) performing a polymerase chain reaction (PCR) on the biological sample to obtain a PCR-amplified biological sample, wherein the PCR reaction on the biological sample is performed using a first set of PCR primers, and the first set of PCR primers amplify a pathogen nucleic acid sequence; (d) sequencing the PCR-amplified lysed biological sample using next-generation sequencing; and (e) providing a positive diagnosis for the pathogen infection if a pathogen sequence is detected by the PCR or the sequencing, or providing a negative diagnosis for the individual if a pathogen sequence is not detected by the PCR or the sequencing. 211. The method of embodiment 210, wherein the individual is a human individual. 212. The method of embodiment 210 or 211, wherein the pathogen infection includes a bacterial infection, a viral infection, or a fungal infection, and combinations thereof. 213. The method of any one of embodiments 210 to 212, wherein the bacterial infection is an infection caused by Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof. 214. The method of any one of embodiments 210 to 212, wherein the fungal infection is an infection caused by Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. 215. The method of any one of embodiments 210 to 212, wherein the viral infection is an infection caused by a DNA virus. 216. The method of embodiment 215, wherein the DNA virus includes hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or smallpox, and combinations thereof. 217. The method according to any one of embodiments 210 to 216, wherein the viral infection is an infection caused by an RNA virus.218. The method of embodiment 217, wherein the RNA virus comprises an influenza virus, a coronavirus, a poliovirus, a measles virus, an Ebola virus, a retrovirus, or an orthomyxovirus. 219. The method of embodiment 218, wherein the viral infection is a coronavirus infection. 220. The method of embodiment 219, wherein the coronavirus infection is a SARS-COV-2 infection. 221. The method of any one of embodiments 210-220, wherein the biological sample from the individual is from a blood sample, a plasma sample, a serum sample, a buccal swab, a urine sample, a semen sample, a vaginal swab, a stool sample, a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. 222. The method of embodiment 221, wherein the biological sample from the individual is from a nasopharyngeal swab, a middle turbinate swab, or any combination thereof. 223. The method of any one of embodiments 210-222, comprising adding a synthetic nucleic acid to the lysis agent or the lysed biological sample. 224. The method of any one of embodiments 210 to 222, comprising adding a synthetic nucleic acid to the method. 225. The method of embodiment 224, wherein the synthetic nucleic acid is RNA. 226. The method of embodiment 224, wherein the synthetic nucleic acid is DNA. 227. The method of any one of embodiments 224 to 226, wherein the synthetic nucleic acid comprises a set of sequences configured to be bound by the first set of PCR primers. 228. The method of any one of embodiments 224 to 227, wherein the first set of primers amplifies both the pathogen nucleic acid sequence and the synthetic nucleic acid. 229. The method of any one of embodiments 224 to 228, wherein the synthetic nucleic acid sequence comprises a nucleotide sequence that is not identical to the pathogen nucleic acid sequence. 230. The method of any one of embodiments 210 to 229, comprising performing a reverse transcription reaction on the lysed biological sample. 231. The method of embodiment 230, wherein the reverse transcription reaction is performed before performing the polymerase chain reaction. 232. The method according to any one of embodiments 210 to 231, wherein the reverse transcription reaction is carried out without further purification of the lysed biological sample.233. The method of any one of embodiments 210 to 231, wherein the reverse transcription reaction on the lysed biological sample produces viral cDNA. 234. The method of embodiment 233, wherein the viral cDNA is coronavirus cDNA. 235. The method of embodiment 233, wherein the coronavirus cDNA is SARS-COV-2 cDNA. 236. The method of any one of embodiments 210 to 235, wherein the reverse transcription reaction and the PCR are a single-step reaction. 237. The method of any one of embodiments 210 to 236, wherein the PCR is an end-point analysis. 238. The method of any one of embodiments 210 to 237, wherein the PCR is not a real-time PCR reaction. 239. The method of any one of embodiments 210 to 238, wherein the first set of PCR primers amplifies a coronavirus nucleic acid sequence. 240. The method of embodiment 239, wherein the coronavirus nucleic acid sequence is a SARS-COV-2 nucleic acid sequence. 241. The method of embodiment 240, wherein the SARS-COV-2 nucleic acid sequence comprises the N1 gene or the S2 gene. 242. The method of any one of embodiments 210-241, further comprising a second set of PCR primers, wherein the second set of primers amplifies a nucleic acid sequence of the individual. 243. The method of embodiment 242, wherein the second set of PCR primers amplifies a human nucleic acid sequence. 244. The method of embodiment 243, wherein the second set of PCR primers amplifies a human nucleic acid sequence selected from GAPDH, ACTB, RPP30, and combinations thereof. 245. The method of embodiment 244, wherein the second set of PCR primers amplifies human RPP30. 246. The method of any one of embodiments 242-245, wherein the second set of PCR primers comprises a mixture of primers with sequencing adapter sequences and primers without sequencing adapter sequences. 247. The method of embodiment 246, wherein the ratio of primers with sequencing adapter sequences to primers without sequencing adapter sequences is about 1:1, about 1:2, about 1:3, or about 1:4. 248. The method of any one of embodiments 210 to 247, wherein said PCR comprises 30 to 45 amplification cycles.249. The method of embodiment 248, wherein the PCR comprises 35 to 45 amplification cycles. 250. The method of embodiment 248, wherein the PCR comprises 39 to 42 amplification cycles. 251. The method of any one of embodiments 210 to 250, wherein the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise nucleic acid sequences comprising a variable nucleotide sequence. 252. The method of embodiment 251, wherein the variable nucleotide sequence is a sample ID unique to the individual. 253. The method of any one of embodiments 210 to 252, wherein the first set of PCR primers, the second set of PCR primers, or both the first set of PCR primers and the second set of PCR primers comprise adapter sequences for next-generation sequencing reactions. 254. The method of any one of embodiments 210 to 253, wherein the method is capable of detecting fewer than 10 copies of a pathogen genome. 255. The method of any one of embodiments 210 to 253, wherein fewer than 5 copies of a pathogen genome can be detected. 256. The method of embodiment 254 or 255, wherein the pathogen genome is a coronavirus genome. 257. The method of embodiment 254 or 255, wherein the coronavirus genome is a SARS-COV-2 genome. 258. The method of any one of embodiments 210 to 257, wherein if a coronavirus sequence is detected by the PCR, the positive diagnosis of coronavirus is a SARS-COV-2 diagnosis. 259. The method of any one of embodiments 210 to 258, wherein the method determines the strain of coronavirus. 260. The method of any one of embodiments 210 to 258, wherein the method determines the strain of COVID-19. 261. A composition comprising a plurality of synthetic nucleic acids having different nucleic acid sequences, wherein the plurality of synthetic nucleic acids having different nucleic acid sequences comprise a common 5' sequence identical to a pathogen nucleic acid sequence, a common 3' sequence identical to a pathogen nucleic acid sequence, and an intervening sequence that differs between the plurality of sequences. 262. The composition of embodiment 261, wherein the plurality of synthetic nucleic acids having different nucleic acid sequences are single-stranded.263. The composition of embodiment 261, wherein the plurality of synthetic nucleic acids having different nucleic acid sequences are double-stranded. 264. The composition of any one of embodiments 261 to 263, wherein the plurality of synthetic nucleic acids having different nucleic acid sequences consists of or comprises RNA. 265. The composition of any one of embodiments 261 to 263, wherein the plurality of synthetic nucleic acids having different nucleic acid sequences consists of or comprises DNA. 266. The composition of any one of embodiments 261 to 265, wherein the common 5' sequence identical to a pathogen nucleic acid sequence is 30 nucleotides or less. 267. The composition of any one of embodiments 261 to 266, wherein the common 3' sequence identical to a pathogen nucleic acid sequence is 30 nucleotides or less. 268. The composition of any one of embodiments 261 to 267, wherein the intervening sequence is 50 nucleotides or less. 269. The composition of embodiment 268, wherein the intervening sequence is 30 nucleotides or less. 270. The composition of any one of embodiments 261-269, wherein the pathogen nucleic acid sequence is derived from a bacterial pathogen, a fungal pathogen, or a viral pathogen. 271. The composition of embodiment 270, wherein the bacterial pathogen is of the genus Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Clostridium, or Escherichia, and combinations thereof. 272. The composition of embodiment 270, wherein the fungal pathogen is Candida, Blastomyces, Cryptococcus, Coccidoides, Histoplasma, Paracoccidioides, Sporothrix, or Pneumocystis, and combinations thereof. 273. The composition of embodiment 270, wherein the viral pathogen is a DNA virus. 274. The composition of embodiment 273, wherein the DNA virus comprises hepatitis B, hepatitis C, papillomavirus, Epstein-Barr virus, chickenpox, or smallpox, and combinations thereof. 275. The composition of any one of embodiments 1 to 3, wherein the viral pathogen is an RNA virus.276. The composition of embodiment 270, wherein the RNA virus comprises an influenza virus, a coronavirus, a poliovirus, a measles virus, an Ebola virus, a retrovirus, an orthomyxovirus, or a combination thereof. 277. The composition of embodiment 276, wherein the viral pathogen is a coronavirus. 278. The composition of embodiment 277, wherein the coronavirus is SARS-COV-2. 279. Embodiment 27, wherein the pathogen nucleic acid sequence is a nucleic acid sequence encoding a coronavirus spike protein. 7 or 278. 280. The composition of any one of embodiments 277-279, wherein the plurality of synthetic nucleic acids having different nucleic acid sequences comprises or consists of a sequence selected from any one or more of S2_001, S2_002, S2_003, and S2_004. 281. The composition of any one of embodiments 261-280 for use in a method for diagnosing or detecting infection by a pathogen. 282. The composition of any one of embodiments 261-280 for use in a method for normalizing pathogen next generation sequencing reads. [Example]

[0164] The following illustrative examples are representative of embodiments of the compositions and methods described herein and are not intended to be limiting in any way.

[0165] Example 1 - SwabSeq Diagnostics of Viral Infections Figure 1 shows an exemplary workflow for collecting and processing samples. Samples can be obtained from a human subject using a swab (e.g., saliva, nasopharynx, or middle turbinate), as illustrated at 102. The swab can be directly subjected to conditions sufficient to lyse cells within the sample, as illustrated at 106. RNA isolation is not required, which (1) reduces the time it takes to perform the assay, (2) increases sensitivity, and (3) reduces the cost of materials for the assay. Samples can optionally be stored in a buffer (e.g., TE buffer) and lysed at a later date, within a few days of obtaining the sample, as in steps 126 and 128. To facilitate accurate and sensitive identification of viral nucleic acids, the lysed sample can be spiked with synthetic nucleic acids that can benchmark subsequent nucleic acid processing and the resulting sequence information (illustrated at 108). The synthetic nucleic acid contains a sequence that can be amplified by virus-specific primers but has a known sequence that can be easily identified as distinct from the viral target in the sequence output. Samples are processed in compartments in microtiter plates that can be scaled between 1 and 32 384-well plates, allowing for testing of at least 384 to 12,288 samples.

[0166] Nucleic acid processing and library preparation are then performed within the compartments / wells by reverse transcription and PCR, as illustrated at 110. The primers used for detection contain sequences complementary to (a) the target of interest (e.g., a small region from coronavirus or influenza), (b) a human target (a control to confirm the presence of any human sequence; e.g., RPP30 (RNase P) is used by the CDC in their standard qPCR-based testing), and (c) an internal synthetic RNA molecule (as described above) to ensure the target assay amplifies and quantify the target, e.g., to allow for normalization of reads from pathogen sequences. PCR amplification is carried out to the endpoint, followed by sequencing. Endpoint PCR and synthetic templates add a level of normalization, avoiding expensive normalization steps. The design of the primers used for target amplification allows for sample deconvolution and the addition of an indicator that allows for identification of which samples contain positive sequencing reads for viral sequences.

[0167] The sample is then sequenced to identify the presence of barcoded sequences and viral target reads. The reads are then processed bioinformatically to determine positive or negative sequencing data for viral vector and viral load. This assay can be further developed to sequence polymorphic regions and determine pathogen strain. This can be achieved by comparing the ratio of target pathogen amplicons to human target amplicons and spikes in the amplicons.

[0168] Example 2 - Detection of SARS-CoV RNA in a sample Samples containing SARS-CoV-2 RNA were analyzed using the methods disclosed herein. The samples contained defined copy numbers of SARS-CoV-2 RNA molecules. Primers targeting the N1 region (nucleocapsid) and subunit S2 region were used to detect SARS-CoV-2 RNA molecules. The resulting sequencing products were compared to human RRP30 RNA. SARS-CoV:RRP30 as a function of copy number was analyzed for a range of target regions (Figures 2A-2B). Surprisingly, targeting subunit S2 resulted in distinguishable detection within samples containing three copies of SARS-CoV-2 RNA molecules. Targeting N1 demonstrated distinguishable detection with acceptable sensitivity, albeit at copy numbers 2-5 times higher than S2. Figure 2B demonstrates the analysis of different primer sequences 210 complementary to the N1 and S2 loci. Detection was dependent on the primer sequence for both N1 and S2 targets. Across all primers tested, targeting S2 demonstrated distinguishable detection at copy numbers 2-5 times lower than N1, suggesting that targeting S2 provides a higher resolution assay for the detection of SARS-CoV RNA in samples.

[0169] Example 3 - Viral diagnosis by amplification and sequencing An example of an amplification and sequencing protocol for use with the methods described herein is described herein: The sample in lysis buffer may be heated (e.g., at 75°C for 10 minutes) prior to addition to the QRT-PCR reaction.

[0170] QRT-PCR reactions were set up in individual reactions with a total volume of 20 μL, including 7 μL of sample lysate, 10 μL of Luna® Universal One-Step Reaction Mix, 1 μL of Enzyme Mix, 400 nM of viral nucleic acid target primer, housekeeping primer (e.g., RPP3 100 nM without adapter, 50 nM with adapter), and 100 copies of synthetic nucleic acid (same priming region as the target virus). By keeping the overall RPP3 primer concentration high while decreasing the RPP3 primer concentration with the sequencing adapter, low amounts of RPP3 were detected while limiting the number of sequencing reads dedicated to host RNA. Furthermore, this also suppressed primer dimer formation.

[0171] The cycling protocol was 55°C for 30 minutes (reverse transcription), 95°C for 1 minute, followed by 40 cycles of {95°C for 10 seconds, 60°C for 30 seconds}.

[0172] Pool 5 μL from each well for sequencing and purify using the [AxyPrep PCR Clean-up Kit] (www.fishersci.com / shop / products / axygen-axyprep-mag-pcr-clean-up-kits / 14223152). Quantify the library with the [deNovix dsDNA High Sensitivity Fluorescent Assay Kit] (www.denovix.com / denovix-dsdna-assays / ). Up to 50% PhiX control library DNA can be added. Sequence the library using Illumina Dual Index Single-Read Sequencing on NextSeq and MiSeq.

[0173] Below is a table (Table 1) listing primer pairs that can be used in the methods for detecting viral infections described herein. Primer pairs are disclosed as pairs; for example, o325_octN1_F and o326_octN1_R describe a primer pair designed to generate amplification of viral nucleic acids. The sequences listed below may further include sequencing adapters and / or index sequences.

[0174] [Table 2-1]

[0175] [Table 2-2]

[0176] [Table 2-3]

[0177] [Table 2-4]

[0178] [Table 2-5]

[0179] [Table 2-6]

[0180] [Table 2-7]

[0181] [Table 2-8]

[0182] [Table 2-9]

[0183] [Table 2-10]

[0184] [Table 2-11]

[0185] [Table 2-12]

[0186] [Table 2-13]

[0187] [Table 2-14]

[0188] [Table 2-15]

[0189] [Table 2-16]

[0190] [Table 2-17]

[0191] Example 4 - SwabSeq: A High-Throughput Platform for Large-Scale SARS-CoV-2 Testing In the absence of an effective vaccine, public health strategies remain the only means to control the spread of SARS-CoV-2, the cause of COVID-19. In contrast to SARS-CoV-1, whose infectivity is associated with symptoms, SARS-CoV-2 is highly infectious during the asymptomatic / presymptomatic phase. Therefore, suppression based on symptoms alone is not possible, making SARS-CoV-2 screening essential for pandemic control.

[0192] As local lockdowns are lifted and people return to work and resume other activities, infection rates are beginning to rise again. In many parts of the United States, the increase in cases is overwhelming the capacity of quantitative RT-PCR tests, which constitute the majority of FDA-approved tests for COVID-19. Even in areas that have expanded testing capacity, the cost of a single test, approximately $100, precludes the deployment of clinical testing for mass screening. Furthermore, long delays in returning results are due to insufficient testing capacity, not to test times of several hours, which renders tests useless for preventing viral transmission or suppressing local outbreaks. Frequent, inexpensive mass testing, combined with contact tracing and isolation of infected individuals, offers the best chance of stopping the spread of the virus. Described herein is SwabSeq, a SARS-CoV-2 testing platform that leverages next-generation sequencing to massively scale up testing capacity.

[0193] SwabSeq improves on the one-step reverse transcription and polymerase chain reaction (RT-PCR) approach in several key areas. SwabSeq uniquely labels each sample using molecular barcodes embedded in the RT and PCR primers. After the RT-PCR step, the barcoded samples are compiled into a single sequencing library, which allows for multiplexing of up to hundreds of thousands of samples on an Illumina sequencer. The test's readout is not fluorescence (thus eliminating the need for expensive qPCR machines) but rather a count of viral sequencing reads. The short read length (26 base pairs) is sufficient to identify the target sequence, limiting total sequencing time to approximately 5 hours.

[0194] Key to SwabSeq's robust ability to identify SARS-CoV-2 is its inclusion of an internal in vitro RNA standard, which allows for normalization of viral read counts within each well. This internal standard transforms the count-based assay into a ratiometric one. This is important because it allows accurate calling of positive and negative samples despite heterogeneity in RT and PCR inhibition in clinical samples. Because it relies on sequencing as a readout in addition to an in vitro RNA standard, SwabSeq can use amplification run for up to 50 cycles until the primers are consumed. This end-point PCR overcomes enzyme inhibition present in unextracted samples, further streamlining the workflow while minimizing loss of analytical sensitivity. These features enable testing at much higher throughput than current RT-quantitative PCR (RT-qPCR) diagnostic tests, which require RNA purification and RT-qPCR, a costly, labor-intensive, and reagent-intensive process that is limited by the availability of the necessary equipment. In contrast, SwabSeq uses only readily available and underutilized thermocyclers and DNA sequencers. SwabSeq can be performed at a scale of tens of thousands of samples per day without the extensive automation typically required at scales beyond a few hundred samples per day.

[0195] As demonstrated herein, SwabSeq has extremely high sensitivity and specificity for detecting viral RNA in purified samples. The disclosure and examples herein further demonstrate low limits of detection in extract-free lysates from mid-nasal swabs and oral fluids. These results demonstrate the potential for SwabSeq to be used for unprecedented scale SARS-CoV-2 testing.

[0196] result SwabSeq is a simple, scalable protocol consisting of five steps (Figure 4, Panel A): (1) sample collection, (2) reverse transcription and PCR using primers containing specific molecular indexes at the i7 and i5 positions (Figure 4, Panel B, Figure 6), (3) pooling of uniquely barcoded samples for library preparation, (4) sequencing of the pooled libraries, and (5) computational assignment of barcoded sequencing reads to each sample for counting and virus detection.

[0197] The assay consists of two primer sets that amplify two genes: the S2 gene of SARS-CoV-2 and the human ribonuclease P / MRP subunit P30 (RPP30). The assay contains a synthetic in vitro transcribed control RNA that is identical to the viral sequence targeted for amplification, except for a short modified stretch (Figure 4, panel C) that allows sequencing reads corresponding to the synthetic control to be distinguished from those corresponding to the negative sequence. The primers amplify both the negative and synthetic sequences with equal efficiency (Figure 7).

[0198] The S2 internal control serves two purposes. First, PCR and reverse transcription inhibitors, as well as other sources of well-to-well variation, can affect amplification in both controls and samples in the same way. The ratio of the number of negative reads to the number of control reads provides a more accurate and quantitative measure of viral load than the number of negative reads alone (Figure 4, panels E and D). Second, the internal control allows linearity to be maintained across a wide range of viral inputs, despite the use of end-point PCR (Figure 8). With this approach, the final amount of DNA in each well is largely determined by the total primer concentration, not the viral load; that is, negative samples are expected to have a similar total number of S2 reads as positive samples. The difference is that in negative samples, all S2 amplicon reads are mapped to the control synthetic sequence. In addition to viral S2, this assay / method also includes reverse transcribing and amplifying a human housekeeping gene to control for specimen quality and collection efficiency (Figure 4, panel F).

[0199] After RT-PCR, samples are combined in equal amounts, purified, and used to generate a single sequencing library. Both Illumina MiSeq and / or Illumina NextSeq can be used to sequence these libraries (Figure 9). Instrument sequencing time is minimized by sequencing only 26 base pairs of the amplicon (Methods). Each read is classified based on its sequence as originating from the negative, control S2, or RPP30 and assigned to a sample based on its associated index sequence (barcode). To maximize specificity and avoid false-positive signals resulting from misclassification or assignment, a conservative edit distance threshold is used for this matching operation (Methods and Supplementary Results). Sequencing reads that do not match one of the expected sequences are discarded. Read counts for the negative and control S2 and RPP30 are obtained for each sample and used for downstream analysis (Methods, below).

[0200] We observed that a few thousand reads were sufficient to detect and quantify the presence of viral RNA in a sample (conservatively 10,000 reads). This translates to 1,500 samples per run on a MiSeq v3 flow cell, 20,000 samples per run on a NextSeq, and up to 150,000 samples per run on a NovaSeq S2 flow cell. Computational analysis takes only a few minutes per run. SwabSeq is a reasonable and scalable protocol for COVID-19 testing. The SwabSeq protocol was further optimized by identifying and removing multiple sources of noise (Figure 21, panels A, B, C, and D).

[0201] Validation of SwabSeq as a diagnostic platform SwabSeq was first validated by the UCLA Clinical Laboratory on purified RNA samples previously tested with a standard RT-qPCR assay. To determine the analytical limit of detection, inactivated virus was diluted in pooled remaining clinical nasopharyngeal swab specimens. All remaining samples were confirmed negative for SARS-CoV-2. Serial two-fold dilutions of heat-inactivated SARS-CoV-2 (ATCCR VR-1986HK) were performed on these remaining samples, ranging from 8,000 to 125 genome copy equivalents (GCE) per mL. SARS-CoV-2 was detected in all samples up to 250 GCE per mL, and in most samples up to 125 GCE per mL (Figure 2A). These results demonstrate the high sensitivity of SwabSeq, with an analytical limit of detection (LOD) of 250 GCE per mL. This detection limit is lower than many currently approved highly sensitive RT-qPCR assays. This comparison demonstrates that SwabSeq is highly effective as a clinical diagnostic test.

[0202] SwabSeq detects viruses with high sensitivity and specificity. Confirmed positive (n=31) and negative (n=33) samples from the UCLA Clinical Microbiology Laboratory were retested. 100% concordance with RT-qPCR results was observed for all samples (Figure 5). Libraries were sequenced on both MiSeq and NextSeq550 (Figure 9), with 100% concordance between the different sequencing instruments.

[0203] One of the major obstacles to scaling up RT-qPCR diagnostic tests is the RNA purification step. RNA extraction is difficult to automate, and the supply chain has not been able to keep up with the demand for necessary reagents over the course of the pandemic. As a way to circumvent the obstacles of RNA purification, the ability of SwabSeq to detect SARS-CoV-2 directly from a variety of sample types without the need for extraction was investigated.

[0204] The CDC recommends several types of media for nasal swab collection: viral transport medium (VTM), Amies transport medium, and normal saline. The primary technical challenge arises from RT or PCR inhibition by components in the collection buffer. Diluting the specimen with water overcomes RT and PCR inhibition, enabling detection of viral RNA in confounded and positive clinical patient samples (Figure 11). Nasal swabs collected directly in Tris-EDTA (TE) buffer were diluted 1:1 with water and tested. This approach yielded a detection limit of 558 GCE / mL (Figure 5, panel C). Comparison of an extraction-free protocol for nasopharyngeal samples collected in normal saline with RT-qPCR performed by the UCLA Clinical Microbiology Lab showed 100% concordance for all samples (Figure 2D).

[0205] An extraction-free saliva protocol was also tested, in which saliva was collected directly into a matrix tube using a funnel-shaped collection device (Figure 12). The main technical challenges were preventing various components in saliva from degrading viral RNA and ensuring accurate pipetting of this heterogeneous and viscous sample type. Heating saliva samples at 95°C for 30 minutes reduced PCR inhibition and improved detection of the S2 amplicon compared to no preheating (Figure 13). After the heating step, the samples were diluted 1:1 by volume of 2x TBE with 0.1% Tween-20. Using this method, an LoD of 2000 GCE / mL was obtained (Figure 5, panel E).

[0206] Consideration SwabSeq has the potential to alleviate existing obstacles in diagnostic clinical testing. Furthermore, SwabSeq has even greater potential to enable testing at the scale necessary for pandemic containment through population surveillance. This technology describes a novel use of massively parallel next-generation sequencing for infectious disease surveillance and diagnosis. SwabSeq can detect SARS-CoV-2 RNA in clinical specimens from both purified RNA and extraction-free lysates, while maintaining high clinical and analytical sensitivity and specificity comparable to RT-qPCR performed in clinical diagnostic laboratories. SwabSeq will be further optimized to prioritize scale and low cost, as these are key elements missing from current COVID-19 diagnostic tests.

[0207] SwabSeq must be evaluated as a surveillance testing tool, not a clinical testing tool. While clinical testing requires high sensitivity and specificity to inform clinical decision-making, the most important factors for surveillance testing are the breadth, frequency, and turnaround time of testing. Sufficiently widespread and frequent testing, coupled with rapid result return, can effectively suppress viral outbreaks through selective isolation of infected individuals rather than blanket stay-at-home orders. Epidemiological modeling of surveillance testing on university campuses has shown that even diagnostic tests with a sensitivity of 70% can suppress infections when administered frequently with a short turnaround time. The major challenges to the practical application of frequent testing are the cost of testing and the logistics of collecting and processing hundreds to thousands of specimens per day.

[0208] The use of Illumina sequencing in diagnostic testing has raised concerns regarding turnaround time and cost. SwabSeq uses a short sequencing run that reads a molecular index and 26 base pairs of the target sequence in 5 hours, followed by computational analysis that can be performed in 5 minutes on a desktop computer. The cost of sequencing reagents for 1,000 samples analyzed in a single MiSeq run is less than $1 per sample. Running 10,000 samples on a NextSeq550, which generates 13x more reads per flow cell, can reduce this cost by approximately 10x. Further optimization to reduce reaction volumes and use less expensive RT-PCR reagents could further reduce the total cost per test.

[0209] Finally, scaling up SARS-CoV-2 testing requires high-throughput sample collection and processing workflows. The manual processes common in most academic clinical laboratories are not easily amenable to simple automation. Current protocols for placing nasopharyngeal swabs into VTMs, Amies, or NSs are collection methods that date back to the pre-molecular genetics era, when live virus cultures were used to identify cytotoxic effects on cell lines and testing was performed manually with low throughput. A new perspective on scaled-up collection methods would be extremely beneficial to the testing community.

[0210] Several groups are piloting "lightweight" sample collection approaches that push sample registration and patient information collection directly to individuals being tested via smartphone apps. Much of the sample request effort stems from a lack of interoperability between electronic health systems, leading to laboratory professionals manually entering all sample information. Developing a HIPAA-compliant registration process can streamline laborious sample requests. To facilitate scalability, sample collection protocols can be used that use small volumes of tubes amenable to simple automation, such as automated capper-decappers and 96-head liquid handlers. These approaches reduce the amount of hands-on work required in the laboratory to process and run tests.

[0211] The SwabSeq diagnostic platform complements the growing arsenal of traditional clinical diagnostic tests, as well as point-of-care rapid diagnostic platforms that have emerged for COVID-19, by increasing testing capacity to meet both diagnostic and widespread surveillance testing needs. In the future, SwabSeq can be easily expanded to accommodate additional pathogen and viral targets. This could further enhance the test's utility, particularly during the winter cold and flu season, when multiple respiratory pathogens circulate in populations and cannot easily be distinguished by symptoms alone. As society seeks to safely reopen educational, business, and recreational sectors, diagnostic testing will likely become part of the new normal.

[0212] method Sample Collection: All patient samples used in this study were de-identified. All samples were obtained with UCLA IRB Approval. Nasopharyngeal samples were collected by healthcare providers in patients with suspected COVID-19.

[0213] Contrived Sample Creation: For clinical detection limit experiments, remaining confirmed COVID-19-negative nasopharyngeal swab samples collected by the UCLA Clinical Microbiology Laboratory were pooled. Pooled clinical samples were spiked with the indicated concentrations of ATCC inactivated virus (ATCC 1986-HK) and extracted as described below. Clinically purified RNA samples were collected as nasopharyngeal swabs and purified with MagMax bead extraction using a KingFisherFlex (Thermofisher Scientific) instrument. All extractions were performed according to the manufacturer's protocol. For samples not requiring extraction, samples were first run at the indicated concentrations (contrived), and negative clinical samples were pooled. Sample dilutions were performed in TE buffer or water before adding to the RT-PCR master mix.

[0214] Extraction-free saliva specimen processing: Saliva is collected directly into a matrix tube using a small funnel (part number from Amazon). Saliva samples are collected into matrix tubes and heated to 95°C for 30 minutes. Samples are then either frozen or processed by dilution with 2x TBE containing 1% Tween-20 to a final concentration of 1x TBE and 0.5% Tween-2. 1x Tween containing Qiagen Protease and RNA Secure (ThermoFisher) was also tested. This was also successful, but had high sample-to-sample variability and required an additional incubation step.

[0215] Processing of extraction-free nasal swab lysates: All extraction-free lysates were inactivated using heat inactivation at 56°C for 30 minutes. Samples were then diluted 1:4 with water and added directly to the master mix. The dilution amount varied depending on the liquid medium used. Of the CDC-recommended media, normal saline was the most robust. Viral transport medium and Amies buffer exhibited significant PCR inhibition that was difficult to overcome, even with dilution with water. In one embodiment, swabs are placed directly into diluted TE buffer, which has no effect on PCR inhibition.

[0216] Barcode primer design: Barcode primers were selected from a set of 1,536 unique 10-bp i5 barcodes and a set of 1,536 unique 10-bp i7 barcodes. These 10-bp barcodes met the criteria of a minimum Levenshtein distance of 3 between any two indices (within the i5 and i7 sets) and that the barcodes contained no homopolymer repeats larger than two nucleotides. Additionally, barcodes were selected to minimize homodimerization and heterodimerization using helper functions in the Python API for Primer3.

[0217] Construction of S2 and RPP30 Synthetic Spikes: For construction of the S2 spike-in DNA template, RT-PCR was performed on SARS-CoV-2 gRNA (Twist BioSciences, #1) using the primers listed in Table 2. For construction of the RPP30 spike-in DNA template, RT-PCR (FP_1,R) and a second-round PCR (FP_2,R) were performed on HEK293T lysates. The products were run on a gel to identify a specific product of ~150 bp. DNA was purified using Ampure beads (Axygen) at a bead:sample ratio of 1.8. The mixture was vortexed and incubated at room temperature for 5 minutes. The beads were bound using a magnet for 1 minute, washed twice with 70% EtOH, and air-dried for 5 minutes. They were then removed from the magnet and eluted with 100 μL of IDTE buffer. The bead solution was returned to the magnet, and the eluate was removed after 1 minute. DNA was quantified using a Nanodrop (Denovix).

[0218] This prepared DNA template was used for standard HiScribe T7 in vitro transcription (NEB). IVT reactions were prepared according to the manufacturer's instructions, using 300 ng of template DNA per 20 μL reaction and incubating at 37°C for 16 hours. The IVT reactions were treated with DNAse I according to the manufacturer's instructions. RNA was purified using the RNA Clean & Concentrator-25 kit (Zymo Research) according to the manufacturer's instructions and eluted in water. The RNA spike-in was quantified using both the Nanodrop and the RNA Screen Tape kit for TapeStation (Agilent) according to the manufacturer's instructions to confirm that the RNA was the correct size (approximately 133 nt).

[0219] [Table 3]

[0220] One-step RT-PCR: RT-PCR was performed in 20 μL reactions using either the Luna® Universal One-Step RT-qPCR Kit (New England BioSciences E3005) or the TaqPath™ 1-Step RT-qPCR Master Mix (Thermofisher Scientific, A15300). Both kits were used according to the manufacturer's protocol. The final primer concentrations in the master mix were 50 nM for the RPP30 F and R primers and 400 nM for the S2 F and R primers. Synthetic S2 RNA was added directly to the master mix at a copy number of 500 copies per reaction. Samples were loaded into 20 μL reactions. All reactions were performed in 96- or 384-well formats, and thermocycler conditions were run according to the manufacturer's protocol. For purified RNA samples, 40 cycles of PCR were performed. For unpurified samples, 50 cycles of end-point PCR were performed.

[0221] Multiplex Library Preparation: After RT-PCR, samples were pooled using a multichannel pipette or an Integra Viaflow Benchtop liquid handler. 6 μL from each well was combined into a sterile reservoir, and the entire volume was transferred to a 15 mL conical tube and vortexed. A total of 100 μL was transferred to a 1.7 mL Eppendorf tube for double-sided SPRI cleanup. Briefly, 50 μL of Ampure XP beads (A63880) were added to the 100 μL pooled PCR volume and vortexed. After 5 minutes, the beads were collected using a magnet for 1 minute, and the supernatant was transferred to a new Eppendorf tube. An additional 130 μL of Ampure XP beads were added to the 150 μL supernatant and vortexed. After another 5 minutes, the beads were collected using a magnet for 1 minute, and the beads were washed twice with fresh 70% EtOH. DNA was eluted from the beads in 40 μL of Qiagen EB buffer. The beads were collected using a magnet for 1 minute and 33 uL of the supernatant was transferred to a new tube.

[0222] Sequencing protocol: Libraries were sequenced on either an Illumina MiSeq (2012) or NextSeq550. Before each MiSeq run, a bleach wash was performed using sodium hypochlorite solution (Sigma-Aldrich, 239305) according to the Illumina protocol. Pooled and quantified libraries were diluted to a concentration of 6 nM (based on the Qubit 4 Fluorometer and Illumina's ng / ul to nM conversion formula) and loaded onto the sequencer at 25 pM (MiSeq) or 1.5 pM (NextSeq). PhiX Control v3 (Illumina, FC-110-3001) was spiked into the libraries at an estimated 30–40% of the library volume. PhiX provides additional sequence diversity to Read 1, which aids in template registration and improves run and base quality.

[0223] For this application, the MiSeq requires two custom sequencing primer mixes: the Read1 primer mix and the i7 primer mix. Both mixes have a final primer concentration of 20 μM (10 μM of sequencing primer for each amplicon). The NextSeq requires an additional sequencing primer, the i5 primer mix, also at a final concentration of 20 μM. The MiSeq Reagent Kit v3 (150 cycles, MS-102-3001) is loaded into reservoir 12 with the Read1 sequencing primer mix and into reservoir 13 with 30 μL of the i7 sequencing primer mix. The NextSeq 500 / 550 Mid Output Kit is loaded into reservoir 20 with 52 μL of the Read1 sequencing primer mix, reservoir 22 with 85 μL of the i7 sequencing primer mix, and reservoir 23 with 85 μL of the i5 sequencing primer mix. Index 1 and Index 2 are 10 bp each, and Read 1 is 26 bp.

[0224] Analysis: Bioinformatics analysis involves standard conversion of BCL files to FASTQ sequencing files using Illumina's bcl2fastq software (v2.20.0.422). Demultiplexing and read counting per sample are performed using custom software. Here, read1 matches one of the three expected amplicons, allowing for the possibility of a single-nucleotide error in the amplicon sequence. Hamming distance is the number of positions at which corresponding sequences differ from each other and is a commonly used measure of distance between sequences. Samples are demultiplexed using two index reads to identify which sample the read originates from. The observed index read matches the expected index sequence, allowing for the possibility of a single-nucleotide error in one or both of the index sequences. If both index1 and index2 are at a Hamming distance greater than 1 from the expected index sequence, the set of three reads is discarded. Read counts are calculated for each amplicon and each sample. For this analysis, several custom scripts were written in R that relied on the ShortRead and stringdist packages to process Fastq files and calculate Hamming distances between observed and expected amplicons and indexes. This approach was very conservative and provided a very low level of control over the sequencing analysis. However, continued development of the kallisto and bustools SwabSeq analysis tools will provide a more user-friendly and computationally efficient solution for other groups performing SwabSeq.

[0225] Criteria for Classifying Patient Purified Samples: QC metrics for each type of specimen were developed for the analytical pipeline. For purified RNA, each sample was required to have at least 10 detectable reads for RPP30 and a combined total of S2 and S2 synthetic spike-in reads of more than 2,000 reads. If these criteria were not met, the sample was rerun once, and if the second run failed, a resample was requested. To determine whether SARS-CoV-2 was present, a ratio of S2 to S2 spikes greater than 0.003 was calculated. (Adding one count to both S2 and S2 spikes before calculating this ratio facilitates plotting the results on a logarithmic scale.) If the ratio was greater than 0.003, it was concluded that SARS-CoV-2 was detected for that sample; if the ratio was less than or equal to 0.003, it was concluded that SARS-CoV-2 was not detected (Figure 10, panels A, B, and C).

[0226] The same pair of primers amplifies both the S and S spike amplicons. Because the run is an endpoint assay, the primers are the limiting reagents for continued amplification. During the development of this assay, we observed that as the S count increased for a given sample, the S spike count decreased (Figure 8). At very high viral loads, the S spike read count decreased to less than 1,000 reads. Therefore, by considering the S and S spike together, QC can call SARS-CoV-2 even at very high viral titers. Therefore, it is important to also consider the scenario where the S amplicon count is very high and the S spike count is low because the sample contains a large amount of SARS-CoV-2 RNA (Figure 8).

[0227] Therefore, because S2 and S2 spike are derived from the same primer pair, QC required that the sum of the S2 count and the S2 spike count be greater than 2000. For example, if the S2 count is detected to be greater than 2000 and the S2 spike count is detected to be greater than 0, this is a definite SARS-CoV-2 positive sample, and consequently, SARS-CoV-2 will be detected.

[0228] Index misassignment analysis: Index misassignment was studied using a unique dual index and an amplicon-specific index. In this scheme, each sample was assigned two unique indexes for the S or spike amplicon and two unique indexes for the RPP30 amplicon, for a total of four unique indexes per sample. A count matrix with all possible pairwise combinations of each index pair (one i7 and one i5) was used to generate a match matrix. Counts on the diagonal of the match matrix correspond to actual samples, while off-diagonal counts correspond to index swap events. The extent of i7 and i5 index misassignment was determined by taking the row and column sums of the off-diagonal elements of the match matrix, respectively. The observed ratio of index swaps to wells with known zero amounts of existing viral RNA was calculated by taking the average of the viral S counts to spike ratios for those wells.

[0229] Supplementary results Improving the limit of detection requires minimizing noise sources: One of the major challenges in performing highly sensitive molecular diagnostic assays is that even a single contaminant or noise source can reduce the analytical sensitivity of the test. During the development of SwabSeq, S2 reads from control samples in which SARS-CoV-2 RNA was not present were observed (Figure 4, panel D). These reads are referred to as "no template control" (NTC) reads. A key part of SwabSeq optimization was understanding and minimizing the sources of NTC reads to improve the assay's limit of detection (LoD). Two sources of NTC reads were identified: molecular contamination and misassigned sequencing reads.

[0230] To minimize molecular contamination, we followed protocols and procedures commonly used in molecular genetic diagnostic laboratories. To limit molecular contamination, we used a dedicated hood for making dilutions of synthetic RNA controls and master mixes. At the start of each new run, we sterilized pipettes with dilution solution and PCR plates with 10% bleach, followed by a 15-minute UV light treatment.

[0231] To prevent high concentrations of post-PCR products from contaminating the pre-PCR process, the pre- and post-PCR steps were physically separated into two separate rooms, where no amplification plates were ever opened within the pre-PCR laboratory space. To further prevent post-PCR contamination, an RT-PCR master mix containing uracil-N-glycosylase (UNG) was compared with an RT-PCR master mix without uracil-N-glycosylase (UNG). The presence of UNG in the TaqPath™ 1-Step RT-qPCR Master Mix (Thermofisher Scientific) demonstrated a significant improvement in reducing post-PCR contamination of S2 reads present in negative patient samples compared with the Luna One Step RT-PCR Mix (New England Biosciences) (Figure 14). The RT-PCR master mix contains a mixture of dTTP and dUTP to ensure that the post-PCR amplicon is uracil-containing DNA. These post-PCR residues are remnants of previously performed SwabSeq experiments and can be selectively removed by UNG.

[0232] The third source of molecular contamination was carryover contamination on the Illumina MiSeq sequencer template line. Indexes from previous sequencing runs without bleach maintenance washes were identified in experiments that did not include those indexes. While some indexes had many S2 reads, the presence of carryover contamination impacts the sensitivity and specificity of the assay. After the extra maintenance and bleach washes, the amount of carryover reads present was substantially reduced to less than 10 reads (Figure 15).

[0233] Another source of NTC reads is amplicon misassignment. Amplicon misassignment occurs when sequencing (and, perhaps to a lesser extent, oligo synthesis) errors result in amplicon sequences that originate from S2 spikes but are incorrectly assigned to S2 sequences within a given sample. Only 6 bp at the beginning of read 1 distinguish S2 from S2 spikes. Sequencing errors can result in S2 spike reads that are incorrectly classified as S2 reads because the error rate appears higher at the beginning of the read (Figure 16, Panel A). If the error correction calculation for amplicon reads is too permissive, these reads may be inadvertently counted in the wrong category. To reduce this source of S2 read misassignment, a more conservative threshold for edit distance was used (Figure 16, Panel B). Future redesigns or extensions to additional viral amplicons should consider engineering longer regions of sequence diversity.

[0234] An additional source of NTC reads is when S2 amplicon reads are assigned to the wrong sample based on the indexing strategy. In the assay, individual samples are identified by pairs of index reads (Figure 4, panel B). Misassignment of a sample to the wrong index can occur due to contamination of the index primer sequence, synthesis errors in the index sequence, sequencing errors in the index sequence, or "index hopping."

[0235] Multiple indexing strategies were utilized in the development of SwabSeq, ranging from full combinatorial indexing (in which samples in the assay are tagged with each possible combination of i5 and i7 indexes) to unique dual indexing (UDI), in which each sample has a distinct and unrelated i7 and i5 index (Figure 17). However, the significant initial cost of developing many unique primers can limit the ability to scale up. A full combinatorial indexing approach significantly increases the number of unique primer combinations. A compromise strategy between full combinatorial indexing and UDI, in which a set of indexes is shared only among a small subset of samples, can also be used. Such a design mitigates the impact of sample misassignment and facilitates scale-up to tens of thousands of patient samples (Figure 17). With full combinatorial indexing (Figure 17), NTC read depth correlated with the total number of S2 reads summed across all samples sharing the same i7 sequence (Figure 18, panel A). This is consistent with an index-hopping effect from samples with high S2 viral reads to samples sharing the same index. If positive samples are randomized across indices, this effect can be computationally corrected for using, for example, a linear mixed model.

[0236] Finally, challenges associated with combinatorial and semi-combinatorial indexing strategies can be mitigated by using unique dual indexing (UDI), a strategy that reduces the number of index-hopped reads by two orders of magnitude. Consistently low S2 viral reads were observed in UDI of negative control samples. Index misassignment can also be quantified by counting reads with index combinations that should not occur in the assay (Figure 21, panels A and B). The number of index-hopping events correlated with the total number of S2 + S2 spike reads (Figure 21, panels C and D), indicating that hopped reads are more likely to come from wells with strong viral signals in the expected index. An overall rate of hopping of 1-2% was quantified with the MiSeq and is known to be higher with patterned flow cell instruments.

[0237] Many sources of noise exist in amplicon-based sequencing, ranging from environmental contamination during RT-PCR and sequencing steps to misassignment of reads based on computational corrections and "index hopping" on Illumina flow cells. Preventing and correcting these error sources dramatically improves the detection limit of the SwabSeq assay.

[0238] Example 5 - Effect of primer concentration on SARS-CoV-2 testing The effect of reducing primer concentration was examined for the S2 and N1 amplicons. Primer concentrations were reduced from 400 nM to 100 nM in 20 μL reactions using either saliva (diluted 1:1) or nasal swabs. The primer concentration for the RPP30 amplicon was the same at 50 nM. Lowering the primer concentration preserves quantitative capability. Furthermore, lowering the primer concentration reduces primer dimers and nonspecific amplification products that dominate sequencing reads. In the gel shown in Figure 22, this primer dimer band is ~120 bp in the no-template control (NTC) lane. Equal amounts of DNA were loaded on each gel. In the gel showing RT-PCR performed with 400 nM primers, the 120 bp primer dimer band in the NTC lane is much brighter than in the gel showing RT-PCR performed with 100 nM primers. A significant improvement and reduction of primer dimers and nonspecific products was also observed by using a primer concentration of 200 nM in the amplification reaction.

[0239] Example 6 - Diversified synthetic sequences The methods described herein are improved by the presence of a synthetic nucleic acid sequence that serves as an internal control sequence that is co-extracted, transcribed, and amplified with the sample.

[0240] S2 synthetic nucleic acid sequence Four synthetic RNA S2 oligos were designed to increase nucleotide base diversity during sequencing, allowing for sufficient base diversity to ensure robust base calling without the need for PhiX. The synthetic RNA oligos were pooled together in an equimolar formula and spiked into an RT-PCR master mix at a total concentration of 250 copies of RNA per reaction. Several combinations were tested, but Set #1 proved to work best. These sequences were derived from the wild-type S2 sequence with specific base pair modifications to evenly represent each of the four bases at each position (targeting 25% A, 25% T, 25% G, and 25% C). Achieving an even distribution of nucleotides required sequence randomization, resulting in a final diversified S2 spike sequence that resembled the original template over a 26-bp read region. The sequences were optimized to reduce secondary structure and all had similar melting temperatures. The flanking sequences shared homology with the S2 amplicon. The LoD using this spike was similar to that of the sequence shown in Figure 23. The primers used for this analysis were forward GCTGGTGCTGCAGCTTATTATGTGGGT (SEQ ID NO: 18) and reverse AGGGTCAAGTGCACAGTCTA (SEQ ID NO: 19).

[0241] N1 synthetic nucleic acid sequence This was also evident when using diversified synthetic sequences that could be amplified with N1-specific primers: 2019-nCoV_N1-F; GACCCCAAAATCAGCGAAAT (SEQ ID NO: 20) and 2019-nCoV_N1-R; TCTGGTTACTGCCAGTTGAATCTG (SEQ ID NO: 21).

[0242] A preliminary limit of detection experiment was performed using an artificial sample (using an ATCC heat-inactivated SARS-CoV-2 standard) spiked into negative saliva and serially diluted. This was performed using three primer conditions: both N1 and S2, N1 only, and S2 only. RPP30 was present at a final reaction concentration of 50 nM, S2 at 100 nM, and N1 at 100 nM. The sensitivity of our assay with S2 alone was determined to be 8000 GCE / mL. The preliminary LoD for N1 from this experiment was 6000 GCE / mL, and the preliminary LoD for N1 + S2 was 4000 GCE / mL, indicating that the addition of N1 improves the sensitivity of the SwabSEQ assay, as shown in Figure 25.

[0243] The inclusion of diversified synthetic sequences could improve the sensitivity and accuracy of the SwabSEQ test while reducing the reliance on PhiX during the sequencing process. See exemplary synthetic sequences in Table 3 below.

[0244] [Table 4]

[0245] Example 7 - SwabSeq can detect multiple different SARS-CoV-2 amplicons. SwabSeq can detect two SARS-CoV-2 genes: N1 and S2, as well as synthetic N1 and S2 controls, increasing the ability to call positive samples. Data for each amplicon is shown in Figure 24. N1 F: CAAGCAGAAGACGGCATACGAGAT XXXXXXXXXX ACCCCAAAATCAGCGAAAT (SEQ ID NO: 22), N1 R:AATGATACGGCGACCACCGAGATCTACAC XXXXXXXXXX TCTGGTTACTGCCAGTTGAATCTG (SEQ ID NO: 23), X represents the unique dual index (UDI) of each primer.

[0246] Example 8 - Performance of SwabSeq on different samples The ability to work from minimal or small samples is important for testing, especially when minimally invasive saliva samples are available. Figure 26 shows that up to 10 uL of saliva sample can be used during library preparation without inhibiting RT or PCR, and Figure 27 shows that up to 6 uL of nasal swab sample (nasal swab inoculated in 1 mL of 1X PBS) can be used during library preparation without inhibiting RT or PCR.

[0247] Example 9 - SwabSeq can be used to detect both influenza and SARS-CoV-2 viruses. Sequencing and detection systems and platforms capable of simultaneously detecting multiple viruses are advantageous. Experiments were conducted to detect SARS-CoV-2 and influenza A and B viruses simultaneously.

[0248] Specifically, SwabSeq's ability to identify target amplicons in the FluA MP, FluA NS, and FluB NS genome sequences was assessed by downloading them from the NCBI Influenza Database. Preliminary primers were then aligned to these sequences to extract target amplification regions for each of these genomes. This information was used to further optimize primers to target as many amplicons as possible. All unique amplification regions were then exported and added to a target list. The results, shown in Figure 28, demonstrate that SwabSeq can detect influenza A or B in artificial samples.

[0249] To perform Swabseq for other respiratory pathogens, universal influenza A and influenza B Swabseq primers were designed. For influenza A, two primer pairs targeting the influenza A gene were tested: M1 (FluA_5_f1 GACCAATCYTGTCACCTCTGAC (SEQ ID NO: 24), FluA_5_f2 GACCAATYCTGTCACCTYTGAC (SEQ ID NO: 25), FluA_5_r AGGGCATTTTGGAYAAAGCGTCTA (SEQ ID NO: 26)) and NS1 (FluA_12_f TTGGGGTCCTCATCGGAG (SEQ ID NO: 27), FluA_12_r TTCTCCAAGCGAATCTCTGTA (SEQ ID NO: 28)). For influenza B, a primer pair specific to the NS1 (FluB_11_f AAGATGGCCATCGGATCC (SEQ ID NO: 29) FluB_11_r GTCTCCCTCTTCTGGTGATAATC (SEQ ID NO: 30)) gene. The sequence of the synthetic nucleic acid spike is shown below. FluA_5_spike_004 GACCGATCCTGTCACCTCTGACTAAGGGTGGCACAACGCAGTGTGTTGAGCTCCCTTAGTAGGGTGACTAGCACCGGCAGCGTAGACGCTTTGTCCAAAATGCCCT (SEQ ID NO: 31) FluA_12_spike_002 TTGGGGTCCTCATCGGAGGACGCTTAAATAATAGTAGGAACGTTCGAGTCTCTAAAAATATACAGAGATTCGCTTGGAGAA (SEQ ID NO: 32) FluB_11_spike_004 AAGATGGCCATCGGATCCTCAACTCACTCCTGTAATGAGTCGTAGACAAGGATAAGAGGCCCGATCGGCAGTATTTCGCCCTTTCTAAAAATAATGTGACCTGGGACGCACTGCACCGATTATCACCAGAAGAGGGAGAC (SEQ ID NO: 33)

[0250] Example 10 - Comprehensiveness and cross-reactivity of S2 primer Given the high mutation rate of SARS-CoV-2, it is important that the primers utilized in sequencing tests be able to amplify many different variants. To assess the inclusiveness of the SwabSeq test, we performed an in silico analysis of the S2 primer set and S2 amplicon sequences. For primer analysis, we evaluated SARS-CoV-2 sequences available on GISAID. We first filtered out low-quality genomes, defined as having 1% or more unidentified nucleotides (N's). We then performed a BLASTn (NCBI) analysis on the remaining 324,355 high-quality genomes to quantify the level of primer identity between the two SARS-CoV-2 primer sequences by querying each of them against the downloaded SARS-CoV-2 sequence. The analysis showed that 91.24% of all analyzed strains had 100% identity to both primer sequences, while 28,398 strains (i.e., 8.76%) of the 324,355 total genomes had less than 100% identity to the primers. The forward S2 primer has 100% identity with 303,840 GISAID genomes (93.67%). The reverse S2 primer has 100% identity with 305,019 GISAID genomes (94.04%). The S2 amplicon is 26 base pairs long, and 303,934 GISAID genomes (93.70%) have 100% identity.

[0251] In silico analysis was performed to assess the cross-reactivity of SwabSeq primers with representative common respiratory pathogens. For primer analysis, 38 non-SARS-CoV-2 consensus genomes were downloaded from NCBI as a negative sample cohort. BLASTn (NCBI) analysis was performed to quantify the number of primer pairs with greater than 80% identity to each of the genomes in the cohort. None of the pathogens showed greater than 80% identity to any of the primers.

[0252] Example 11 - Analysis of sequencing data Presented herein is an example of an analytical platform for determining pathogen infection that can be combined with the extraction, amplification, and sequencing steps of SwabSeq.

[0253] First, all single-base substitutions, deletions, and insertions for all targeted amplicons and barcode indexes are generated. These are then used as keys in a nested data structure containing the original sequence, target name, and match type (exact match, mismatch, deletion, or insertion). Because the keys in this data structure are indexed, exact string matching can be performed against the sequencing data to quickly retrieve target information within a predetermined matching tolerance (single-base substitution, deletion, or insertion). All duplicate sequences are removed, and the match type of any remaining unique sequences is reclassified as "unconfirmed."

[0254] The sequencing data is then extracted from the FASTQ file and matched to the corresponding dictionary (index 1, index 2, or amplicon). Prior to matching, the instrument and kit information is read to determine which sequences need to be reverse complemented before matching.

[0255] The matching function returns either the original target sequence (used for indexing) or the name of the target (used for amplicons). This is an important step because many targets, including diversified spike and influenza targets, have several sequences that need to be aggregated together. Other data, such as match type, can be returned as well.

[0256] Matched sequences are then aggregated by their index and amplicon target, returning a table of total reads per index pair. This information is used downstream for quality control (it can be used to determine carryover contamination, index hopping, amplification failures, etc.), as well as the presence or absence of COVID-19 and influenza A / B.

[0257] This information is returned as an automatically generated PDF containing quality control plots and a table with information about each sample, such as QC pass / fail, COVID-19 positive / negative, Influenza A / B positive / negative, etc.

[0258] Dictionary-based perfect matching of amplicons and barcodes allows for fast analysis without relying on existing bioinformatics analysis libraries, while accounting for all single-base deletions, insertions, and mismatches.

[0259] An example of an algorithm for calling a positive sample when using S2 primers, S2 synthetic nucleic acids, and RPP30 primers is shown in Table 4 below and Figure 29, and requires the absolute number of RPP30 amplicon reads, the absolute number of S2+S2 spikes, and a ratio of S2 / S2 spikes greater than 0.1.

[0260] [Table 5]

[0261] Example 12 - Detection Limits Limit of detection (LoD) testing was performed by spiking heat-inactivated SARS-CoV-2 virus (ATCC, VR-1986) into negative saliva samples using a dilution series. Eleven extraction replicates were performed per concentration. The preliminary LoD was defined as the lowest concentration that tested positive in 11 of the 11 replicates. LoD determinations were performed independently for the Luna® Probe One-Step RT-qPCR 4X Mix with UDG and the TaqPath™ 1-Step RT-qPCR Master Mix. For both RT-qPCR master mixes, the preliminary LoD was determined to be 8,000 GCE / ml. The results are shown in Table 5.

[0262] [Table 6]

[0263] For both the Luna® Probe One-Step RT-qPCR 4X Mix with UDG and the TaqPath™ 1-Step RT-qPCR Master Mix, limit of detection validation studies were performed by spiking in heat-inactivated SARS-CoV-2 virus (ATCC, VR-1986) in negative saliva samples at final concentrations of 12,000 GCE / mL and 8,000 GCE / mL, corresponding to 1.5x and 1x the determined limits of detection. Ten extraction replicates were performed for each concentration. LoD validation was performed separately at these two concentrations for the Luna® Probe One-Step RT-qPCR 4X Mix with UDG and the TaqPath™ 1-Step RT-qPCR Master Mix. The minimum LoD for both RT-qPCR master mixes was confirmed to be 8000 GCE / mL for both the Luna® Probe One-Step RT-qPCR 4X Mix with UDG and the TaqPath™ 1-Step RT-qPCR Master Mix.

[0264] Example 13 - Clinical Evaluation A study was conducted to evaluate the performance of SwabSeq by comparing saliva clinical samples that had previously been analyzed using laboratory-developed tests (LDTs) and confirmed as positive or negative.

[0265] Saliva samples were collected and run using a MiniSeq Illumina sequencer. Results using the Luna Probe One-Step RT-qPCR 4x Mix with UDG are summarized in Table 6. Using the Luna Probe One-Step RT-qPCR 4x Mix with UDG, the positive agreement rate (25 / 25) was 100% and the negative agreement rate (93 / 93) was 100%. Results using the Taqpath 1-Step RT-qPCR Master Mix are summarized in Table 7. The positive agreement rate (25 / 25) was 100% and the negative agreement rate (93 / 93) was 100%.

[0266] [Table 7]

[0267] [Table 8]

[0268] To assess the reliability of SwabSeq, an intermediate precision test was performed on Day 2 using the same clinical samples. The testers and equipment used on Day 2 (Viaflo-96, manual pipette set, MiniSeq sequencer, and thermal cycler) were all different from Day 1. The saliva clinical samples were then compared to those previously analyzed using laboratory-developed tests (LDTs) that were confirmed to be positive or negative.

[0269] Results using SwabSeq are summarized in Tables 8 and 9. Clinical samples analyzed using Luna Probe One-Step RT-qPCR 4x Mix with UDG are summarized in Table 8. The positive concordance rate (25 / 25) was 100%, the negative concordance rate (93 / 93) was 100%, and the concordance rate between the two intermediate precision runs was 100%. Clinical samples analyzed using Taqpath 1-Step RT-qPCR Master Mix are summarized in Table 9. The positive concordance rate (25 / 25) was 100%, the negative concordance rate (93 / 93) was 100%, and the concordance rate between the two intermediate precision runs was 100%.

[0270] [Table 9]

[0271] [Table 10]

[0272] Two intermediate precision runs performed with SwabSeq comparing previously analyzed saliva clinical samples confirmed as positive or negative show 100% agreement.

[0273] Example 14 - Accuracy To assess intra- and inter-run precision, replicates were performed across three runs on a MiniSeq sequencer at two artificial virus concentrations per mL: 12,000 copies / mL and 24,000 copies / mL. This was replicated using both the Luna® Probe One-Step RT-qPCR 4X Mix with UDG (Table 10) and the TaqPath™ 1-Step RT-qPCR Master Mix (Table 11). These concentrations represent 1.5x and 3x the detection limit of our assay. We observed that the concordance rate for our third precision run using the Taqpath master mix was not 100% (see Table 11). We frequently observed that the artificial sample, consisting of a viral RNA SARS-CoV-2 standard spiked into negative control saliva, could deteriorate if not processed promptly. In the fourth precision run, both Luna and Taqpath confirmed 100% concordance for the fourth run, as the artificial sample was not left at room temperature for very long.

[0274] [Table 11]

[0275] [Table 12]

[0276] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous modifications, changes, and substitutions will occur to those skilled in the art without departing from the invention. It will be understood that various alternatives to the embodiments of the invention described herein may be utilized in practicing the invention.

[0277] All publications, patent applications, issued patents, and other documents mentioned herein are incorporated by reference herein as if each individual publication, patent application, issued patent, or other document was specifically and individually indicated to be incorporated by reference in its entirety. Definitions contained in text incorporated by reference are excluded to the extent they conflict with definitions in this disclosure.

Claims

1. 1. A method for detecting a viral infection in an individual, comprising: (a) lysing a biological sample to produce a lysed biological sample, the biological sample comprising synthetic viral RNA, wherein the sequence of the synthetic viral RNA differs from a naturally occurring nucleic acid sequence of the virus, the biological sample comprising a plurality of synthetic viral RNA sequences, wherein the plurality of synthetic viral RNA sequences comprises at least two different synthetic viral RNA sequences; (b) performing a reverse transcription reaction on the lysed biological sample to obtain a lysed reverse transcribed biological sample, wherein the lysed biological sample is not isolated or purified prior to performing the reverse transcription reaction; (c) performing an amplification reaction on the lysed and reverse transcribed biological sample to obtain an amplified biological sample, wherein the amplification reaction on the lysed and reverse transcribed biological sample is performed using a set of viral primers specific to a viral nucleic acid sequence, and the set of viral primers amplifies the nucleic acid sequence of the virus and the synthetic viral RNA; (d) sequencing the amplified biological sample using next generation sequencing; A method comprising:

2. 10. The method of claim 1, further comprising indicating a positive result for infection with the virus when a sequence read of the nucleic acid sequence of the virus is detected.

3. 3. The method of claim 1 or 2, wherein the synthetic viral RNA is synthetic coronavirus RNA.

4. The method according to any one of claims 1 to 3, wherein the viral infection is SARS-COV-2 infection.

5. 5. The method of any one of claims 2-4, wherein the positive diagnosis of infection with the virus is provided when the sequence reads from the nucleic acid sequence of the virus and the sequence reads of the synthetic viral RNA exceed about 100.

6. The method of any one of claims 1 to 5, wherein lysing the biological sample comprises thermal lysis.

7. 10. The method of claim 1, wherein the plurality of synthetic viral RNA sequences comprises at least four different synthetic viral RNA nucleic acid sequences.

8. 8. The method of any one of claims 1-7, wherein the synthetic viral RNA or synthetic viral RNAs comprise guanine nucleotides in an amount of about 20% to about 30%, adenine nucleotides in an amount of about 20% to about 30%, cytosine nucleotides in an amount of about 20% to about 30%, and uracil nucleotides in an amount of about 20% to about 30%.

9. 9. The method of any one of claims 1 to 8, wherein the synthetic viral RNA nucleic acid or plurality of synthetic viral RNA sequences comprises a synthetic SARS-Cov-2 RNA nucleic acid or plurality of synthetic SARS-Cov-2 RNA nucleic acids.

10. The method of any one of claims 1 to 9, wherein the method further comprises detecting influenza A infection, influenza B infection, or a combination thereof.

11. 11. The method of any one of claims 1 to 10, wherein the amplification reaction on the lysed biological sample is performed using a set of influenza A primers specific for influenza A nucleic acid sequences or a set of influenza B primers specific for influenza B nucleic acid sequences.

12. 12. The method of any one of claims 1 to 11, wherein the amplification reaction on the lysed biological sample is performed using a set of influenza A primers specific for influenza A nucleic acid sequences and a set of influenza B primers specific for influenza B nucleic acid sequences.

13. 13. The method of any one of claims 1 to 12, wherein the biological sample from the individual further comprises influenza A synthetic RNA, influenza B synthetic RNA, or a combination thereof, wherein the influenza A synthetic RNA, the influenza B synthetic RNA, or the combination thereof differs from a nucleic acid sequence of naturally occurring influenza A or influenza B.

14. 14. The method of any one of claims 1-13, further comprising indicating a positive result for influenza A infection if the ratio of sequence reads from influenza A to sequence reads of influenza A synthetic RNA, or its mathematical equivalent, is greater than a ratio of about 0.

1.

15. 15. The method of any one of claims 1-14, further comprising indicating a positive result for influenza B infection if the ratio of sequence reads from influenza B to sequence reads of influenza B synthetic RNA, or its mathematical equivalent, is greater than a ratio of about 0.

1.

16. The method of any one of claims 1 to 15, wherein the nucleic acid sequence of the virus is an N1 sequence, an S2 sequence, or a combination thereof.

17. 17. The method of any one of claims 1 to 16, wherein the nucleic acid sequence of the virus is a SARS-Cov-2 N1 sequence, a SARS-Cov-2 S2 sequence, or a combination thereof.

18. 18. The method of any one of claims 1 to 17, wherein the synthetic viral RNA is present at a concentration of from about 10 copies / reaction to about 500 copies / reaction.

19. The method of any one of claims 1 to 18, wherein the biological sample comprises a nasal swab or a saliva sample.

20. 20. The method of claim 19, wherein the biological sample comprises less than about 10 microliters of saliva from an individual or less than about 10 microliters of buffer solution inoculated with a nasal swab from an individual.

21. The method according to any one of claims 1 to 20, wherein the amplification reaction on the lysed biological sample is carried out using a primer pair specific for a sample control.

22. 22. The method of claim 21, wherein the sample control is a housekeeping gene.

23. 22. The method of claim 21, wherein the primer pair specific to the sample control is specific to RPP30.

Citation Information

Patent Citations

  • Compositions for Use in Identifying Adventitious Contaminating Viruses

    JP2009519010A

  • Direct amplification and detection of viruses and bacterial pathogens

    JP2014522646A

  • Improved virus detection method

    JP2017209036A

  • Methods and kits for detecting SARS-associated coronavirus

    US20040265796A1

  • Multiplex detection assay for influenza and RSV viruses

    US20090181360A1