Methods and systems for detection of COVID variants

JP2024523910A5Pending Publication Date: 2025-06-25LABORATORY CORPORATION OF AMERICA HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023578734
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-21
Filing Date
2022-06-21
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

There is a need to identify and track new variants of the SARS-CoV-2 virus causing COVID-19 and determine their geographic distribution to assist public health authorities in responding to the pandemic effectively.

Method used

A method involving nucleic acid sequencing of SARS-CoV-2 samples, using tiled primers to amplify a large portion of the viral genome, labeled with molecular barcodes, and aligning sequences with a reference genome to detect variants, which can be integrated with geographic identifiers for tracking.

Benefits of technology

Enables efficient identification and tracking of SARS-CoV-2 variants with high accuracy, allowing for correlation with infectivity and disease severity, and integration with public health databases for comprehensive pandemic response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods and systems are disclosed for detection of variants of the SARS-CoV-2 virus that causes COVID-19. For example, methods are disclosed for identifying and / or tracking variants of SARS-CoV-2, the methods including: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Field Methods and systems are disclosed for detection of variants of the SARS-CoV-2 virus that causes COVID-19 and the geographic location of individuals infected with any strain of the variants. [Background technology]

[0002] Introduction SARS-CoV-2 is an enveloped, single-stranded RNA virus of the family Coronaviridae, genus Beta coronavirus. All coronaviruses share similarities in the organization and expression of their genomes, which code for 16 nonstructural proteins and four structural proteins: spike (S), envelope (E), membrane (M) and nucleocapsid (N). Viruses of this family are of zoonotic origin. They cause diseases with symptoms ranging from a mild common cold to more severe ones such as Severe Acute Respiratory Syndrome (SARS), Middle East Respiratory Syndrome (MERS) and coronavirus disease 2019 (COVID-19). Other coronaviruses known to infect people include 229E, NL63, OC43 and HKU1. The latter are ubiquitous and infections generally cause cold or flu-like symptoms.

[0003] The 2019 novel coronavirus (SARS-CoV-2) is a beta-coronavirus that first emerged in December 2019 in Wuhan, China, as a pathogen with pandemic potential. Early reports suggested that limited person-to-person transmission occurred within China. However, in early 2020, additional cases of 2019-nCoV were recognized worldwide, thereby indicating persistent person-to-person transmission. Thus far, the clinical spectrum of SARS-CoV-2 ranges from mild self-limited upper respiratory tract infections to more severe lower respiratory tract disease resulting in significant morbidity and mortality. As the SARS-CoV-2 pandemic accelerates, greater attention has been paid to the diversity of viral genome sequences and how these variants affect the transmissibility of infection, the severity of infection, or viral evasion from innate or vaccine-induced immunity.

[0004] Viruses are constantly changing through mutation. Multiple variants of the virus that causes COVID-19 have been documented in the United States and globally. Some variants come and go, while others persist and exhibit increased transmissibility or increased severity of symptoms.

[0005] For example, as of June 2021, there were six notable variants in the United States: (1) B.1.1.7: This variant was first detected in the United States in December 2020. It was first detected in the United Kingdom. (2) B.1.351: This variant was first detected in the United States in late January 2021 and was first detected in South Africa in December 2020. (3) P.1: This variant was first detected in the United States in January 2021 - P.1 was first identified in a traveler from Brazil who was screened during routine screening at an airport in Japan in early January. (4) B.1.427 and (5) B.1.429: These two variants were first identified in California in February 2021. (6) B.1.617.2: This variant was first detected in the United States in March 2021. It was first identified in India in December 2020. CDC.gov / coronavirus / 2019-ncov / variants.

[0006] Thus, there is a need to identify and track new variants, and there is an additional need to track the geographic location of infected individuals to assist public health authorities in responding to the pandemic. Summary of the Invention [Means for solving the problem]

[0007] Abstract Methods and systems are disclosed for identifying and tracking variants of SARS-CoV-2 that can cause COVID-19. The methods and systems can be embodied in a variety of ways.

[0008] In certain embodiments, methods may include methods for identifying and / or tracking variants of SARS-CoV-2, the methods including: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating a sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence comprises a SARS-CoV-2 variant sequence.

[0009] In certain embodiments, the sequencing covers most of the viral genome. Thus, in certain embodiments, where the sample SARS-CoV-2 genome is amplified by RT-PCR, the resulting cDNA is then further amplified using tiled primers that bind at intervals along the viral genome. In certain embodiments, the tiled primers are spaced such that adjacent primers are 600 bp apart from each other. In this way, the SARS-CoV-2 genome is amplified very efficiently regardless of the presence or absence of new variants. For example, in certain embodiments, the nucleic acid sequencing includes sequencing at least 80%, or optionally at least 85%, or optionally at least 90%, or optionally at least 95% of the entire viral genome.

[0010] The amplified nucleic acid molecules can be labeled with a molecular barcode that identifies the sequence. For example, in certain embodiments, the tiling primers are primers that further include an adapter for the addition of a barcode sequence and / or a universal primer site for nucleic acid sequencing.

[0011] Also disclosed are systems for performing any of the disclosed method steps, and computer program products tangibly embodied in a non-transitory machine-readable storage medium that include instructions configured to execute any of the stations and / or components of the system and / or to perform a step or steps of the methods of any of the disclosed embodiments.

[0012] Also disclosed is a system including one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein; and a computer program product tangibly embodied in the non-transitory machine-readable storage medium, the computer program product including instructions configured to cause the one or more data processors to perform some or all of the methods disclosed herein.

[0013] Sequencing as described herein is advantageous for identifying variants.Various nucleic acid sequencing protocols can be used.In certain embodiments, nucleic acid sequencing comprises RT-PCR.For example, in certain embodiments, PacBio® sequencing protocol and / or device is used.

[0014] In further embodiments, methods and systems are disclosed for identifying the geographic location of an individual infected with a variant. For example, in certain embodiments, a barcode is associated with the individual's zip code or other geographic identifier. In addition, the present disclosure provides methods and / or systems for tracking the prevalence of variants in a population of infected individuals and / or in the general population. In either case, the geographic region may include a population. In further embodiments, the present disclosure provides methods and systems for correlating a particular variant with infectivity (virus transmission) and disease severity.

[0015] Data generated by the methods or systems of the present disclosure can be combined for analysis with other data of similar type from other sources and / or with other data of different types. In certain embodiments, the data can be deposited in a repository for analysis and / or combination with other data. In certain embodiments, the repository is the CDC database. Alternatively, other government or university or private databases can be used.

[0016] The present disclosure can be better understood with reference to the following non-limiting figures. [Brief description of the drawings]

[0017] [Figure 1] FIG. 1 shows a method for detection of SARS-CoV-2 variants according to certain step embodiments of the present disclosure.

[0018] [Diagram 2] FIG. 2 shows a method for preparing sample-specific SARS-CoV-2 nucleic acids for sequencing according to certain step embodiments of the present disclosure.

[0019] [Diagram 3] FIG. 3 shows a method for whole genome sequencing and variant identification according to an embodiment of the present disclosure.

[0020] [Figure 4] FIG. 4 illustrates method steps for analysis of sequence data according to one embodiment of the present disclosure.

[0021] [Diagram 5] FIG. 5 shows method steps for variant identification and phylogenetic assignment according to certain embodiments of the present disclosure.

[0022] [Figure 6]FIG. 6 shows method steps for revalidation of variant identification using in-house data and external databases according to an embodiment of the present disclosure.

[0023] [Figure 7] FIG. 7 shows a system for detection of SARS-CoV-2 variants according to an embodiment of the present disclosure.

[0024] [Figure 8] FIG. 8 illustrates a computing device for use with any of the methods or systems according to certain embodiments of the present disclosure.

[0025] [Figure 9] FIG. 9 shows a map of a 96-well plate used to distribute M13 forward (1001-1032) and M13 reverse primers (1049-1079, and 1082) in an amplification reaction according to one embodiment of the present disclosure.

[0026] [Figure 10] FIG. 10 shows a map of a 96-well plate with M13 forward (1001-1032) and reverse (1049-1051) barcoded primer combinations for use in amplification reactions according to certain embodiments of the present disclosure.

[0027] [Figure 11] FIG. 11 shows the distribution of average read coverage (NTC average CCS read depth) according to an embodiment of the present disclosure.

[0028] [Figure 12] FIG. 12 shows the mean CCS read counts of confirmed negative nucleic acid amplification (NAA) diagnostic samples according to certain embodiments of the present disclosure.

[0029] [Figure 13] FIG. 13 shows the distribution of strains according to certain embodiments of the present disclosure.

[0030] [Figure 14] FIG. 14 shows the rates for 90% genome coverage by NAA CT value according to certain embodiments of the present disclosure.

[0031] [Figure 15] FIG. 15 shows the average read counts for inter-assay samples used in a stability study of three separate sequencing runs (PBT5073, PBT5075, and PBT5080) according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0032] Detailed Description definition The terms sample or patient sample or biological sample or specimen are used interchangeably herein. Samples can include upper and lower respiratory specimens. Such specimens (samples) can include nasopharyngeal or oropharyngeal swabs, sputum, lower respiratory tract aspirates, bronchoalveolar lavage, and nasopharyngeal wash / aspirates or nasal aspirates. Other non-limiting examples of samples include tissue samples (e.g., biopsies), blood or blood products (e.g., serum, plasma, etc.), cell-free DNA, urine, liquid biopsy samples, or combinations thereof. The term "blood" encompasses whole blood, blood products, or any fraction of blood, such as serum, plasma, buffy coat, etc., as conventionally defined.

[0033] As used herein, the term subject or individual refers to a human or any non-human animal. A subject or individual may be a patient, which refers to a human who visits a health care provider for diagnosis or treatment of a disease, which in some cases may be any infectious disease caused by a pathogen. Also, as used herein, the term "individual", "subject" or "patient" includes all warm-blooded animals.

[0034] As used herein, SMRT refers to single molecule real-time sequencing using a zero mode waveguide (ZMW). A single DNA polymerase enzyme is attached to the bottom of the ZMW with a single molecule of DNA as a template. The ZMW creates an illuminated observation volume that is small enough to observe only a single nucleotide that is incorporated. Each of the four DNA bases is attached to one of four different fluorescent dyes. When a nucleotide is incorporated by the DNA polymerase, the fluorescent tag is cleaved and diffuses out of the observation area of ​​the ZMW, where its fluorescence can no longer be observed. A detector detects the fluorescent signal of the nucleotide incorporation, and base calling is performed according to the corresponding fluorescence of the dye.

[0035] As used herein, CT or ct refers to the cycle threshold, or the total number of cycles required to amplify and detect viral (e.g., SARS-CoV-3) nucleic acid by RT-PCR.

[0036] As used herein, locus loop capture is the process of using a molecular inversion probe to bind to and amplify a region of interest within a viral genome.

[0037] As used herein, a CCS or circular consensus sequencing read is a processed read in which errors in the raw sequencing data have been corrected by sequencing the length of the captured DNA fragment multiple times.

[0038] As used herein, repeatability (or intra-assay precision) describes the closeness of agreement between results of successive measurements of the same analyte, performed under the same measurement conditions. Intra-assay repeatability is a measure of the variability when the same specimen is analyzed within a single analytical run.

[0039] As used herein, reproducibility (or interassay precision) describes the closeness of agreement between results of successive measurements of the same analyte, performed under the same measurement conditions. Interassay repeatability is a measure of variability when the same analyte is analyzed in more than one run.

[0040] As used herein, agreement is a measure of the closeness of agreement between a measured value and a value conventionally accepted as true or an accepted reference value. This may require a "gold standard," or accepted methodology against which new methods can be compared.

[0041] As used herein, analytical validity involves establishing the probability that a test is positive when a particular sequence (analyte) is present (analyte sensitivity) and the probability that a test is negative when the sequence is absent (analyte specificity).In next generation sequencing (NGS), analytical sensitivity can be the likelihood that an assay will detect the target sequence variation when present, the nucleic acid sequence derived from the assay, and the reference sequence.For NGS, analytical specificity is defined as the probability that an assay will not detect a sequence variation when it is not present (false discovery rate is a useful measurement of sequencing assays).

[0042] As used herein, specificity defines the ability of a measurement procedure to measure only the analyte.

[0043] As used herein, assay tolerance to nucleic acid input is the tolerance to variations in the amount of analyte added to the reaction.

[0044] As used herein, GISAID is a global scientific initiative and primary source established in 2008 that provides open access to genomic data for influenza and coronavirus (e.g., COVID-19) data. The database is the world's largest repository of SARS-CoV-2 sequences. GISAID facilitates genomic epidemiology and real-time surveillance to monitor the emergence of new COVID-19 virus strains.

[0045] As used herein, when an action is "based on" something, this means that the action is based, at least to some extent, on at least a portion of that something.

[0046] Methods for NGS SARS-CoV-2 Strain Determination Methods and systems are disclosed for identifying and tracking variants of SARS-CoV-2 that can cause COVID-19. The methods and systems can be embodied in a variety of ways.

[0047] In certain embodiments, methods may include methods for identifying and / or tracking variants of SARS-CoV-2, the methods including: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating a sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence comprises a SARS-CoV-2 variant sequence.

[0048] The method may utilize a sample of unknown COVID status or may use a sample that has previously tested positive for COVID. In certain embodiments, a positive sample may be identified using an approved EUA approved COVID-19 RT-PCR Test (e.g., Labcorp EUA200011 and / or EUA203057). Thus, the result is of the identification of the SARS-CoV-2 strain infecting the individual following detection of viral RNA in the sample.

[0049] In certain embodiments, the sequencing covers most of the viral genome. Thus, in certain embodiments, where the sample SARS-CoV-2 genome is amplified by RT-PCR, the resulting cDNA is then further amplified using tiled primers that bind at intervals along the viral genome. In certain embodiments, the tiled primers are spaced such that adjacent primers are 600 bp apart from each other. In this way, the SARS-CoV-2 genome is amplified very efficiently regardless of the presence or absence of new variants. For example, in certain embodiments, the nucleic acid sequencing includes sequencing at least 80%, or optionally at least 85%, or optionally at least 90%, or optionally at least 95% of the entire viral genome.

[0050] The amplified nucleic acid molecules can be labeled with a molecular barcode that identifies the sequence. For example, in certain embodiments, the tiling primers are primers that further include an adapter for the addition of a barcode sequence and / or a universal primer site for nucleic acid sequencing.

[0051] In certain embodiments, the step of generating sample-specific SARS-CoV-2 nucleic acids comprises generating sample-specific SARS-CoV-2 cDNA using reverse transcriptase polymerase chain reaction (RT-PCR).

[0052] Also, in certain embodiments, generating sample-specific SARS-CoV-2 nucleic acids includes using targeted next-generation sequencing in combination with inverted molecular probes as a means to generate sample-specific SARS-CoV-2 nucleic acids (e.g., Molecular Loop SARS-CoV-2 Sequencing Panel). For example, in certain embodiments, generating sample-specific SARS-CoV-2 nucleic acids further includes hybridizing one strand of the sample SARS-CoV-2 cDNA to a single-stranded probe DNA template that includes a pair of SARS-CoV-2 probes, where a first probe is located at the 3' end of the probe DNA template and a second probe is located at the 5' end of the probe DNA template. In this manner, the 3' probe functions as a forward primer and the 5' probe functions as a reverse primer.

[0053] In certain embodiments, the probe sequences are selected as tiled probes that bind at intervals along the SARS-CoV-2 genome. In certain embodiments, the Wuhan-Hu-1 SARS-CoV-2 reference genome (NC_045512) (available at www.ncbi.nlm.nih.gov / nuccore / NC_045512) is used. Alternatively, other known reference genomes may be used. For example, in alternative embodiments, the probes may be spaced apart by about 100, or 200, or 300, or 400, or 500, or 600, or 700, or 800, or 900 base pairs, or more than 1,000 base pairs. Alternatively, intervals within this range (e.g., 450, 550, 650, or 750) may be used. The probes may be tiled across more than 99% (e.g., 99.6%) of the 30 kb SARS-CoV-2 viral genome. Probes can be tiled across the sequence and / or such that on average a given nucleotide provides 2X, 7X, 22X or more of the sequence.

[0054] Also, in certain embodiments, the single-stranded probe DNA template further comprises a universal sequencing primer (e.g., M13 primer) adjacent to the probe sequence. These may allow enrichment by matching the universal primer sequence with a unique sample-specific barcoding for downstream bioinformatics analysis. In addition, in certain embodiments, and as disclosed in more detail herein, the single-stranded probe DNA template further comprises an adapter sequence for the addition of a barcode sequence used to correlate the SARS-CoV-2 sample-specific nucleic acid with a sample number. In some cases, the barcode may be correlated with the zip code of the sample and / or patient's origin. The method may also include filling in the sequence between the two probes to generate a circular single-stranded probe DNA template containing a sequence specific to the sample SARS-CoV-2 cDNA between the two probe sequences, then liberating the circular single-stranded probe DNA template containing a sequence specific to the sample SARS-CoV-2 cDNA from the sample-specific SARS-CoV-2 DNA, and digesting the circular single-stranded probe DNA template containing a sequence specific to the sample SARS-CoV-2 cDNA to generate a linear DNA to be used as a template for nucleic acid sequencing. In certain embodiments, the linear probe DNA template is then modified to add an adapter and then PCR amplified (enriched) for DNA sequencing. In certain embodiments, the enrichment step includes a purification step (e.g., purification by beads). For example, in certain embodiments, substrates for sequencing are generated by RT-PCR, and then SARS-CoV-2 sequences are identified using approximately 1000 tiled Molecular Loop Inversion Probes (MIPS) designed to amplify RNA reverse transcribed into cDNA from 99.6% of the SARS-CoV-2 genome, with most bases covered by 22 MIPS. In certain embodiments, products synthesized during MIPS are enriched and have sample-specific molecular barcodes added by amplification and subsequently sequenced.

[0055] In certain embodiments, the method uses whole genome sequencing. In certain embodiments, next generation sequencing (NGS) is used. Alternatively, other types of sequencing may be used, such as, but not limited to, Sanger sequencing, shotgun sequencing, SMRT sequencing, pyrosequencing or nanopore sequencing. For example, in certain embodiments, PacBio whole genome sequencing may be used with corresponding SMRT link 9 software and analysis tools. For example, in one embodiment, the method may use PacBio whole genome sequencing test for SARS-CoV-2 strain identification using remaining total nucleic acid extract from positive samples. In certain embodiments, the nucleic acid sequencing includes sequencing at least 80%, or optionally 85%, or optionally 90% or more of the entire viral genome.

[0056] In certain embodiments, determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence includes aligning the sample SAR-CoV-2 sequence with a SARS-CoV-2 reference genome to generate a sample-specific assembly and consensus sequence. In addition, the method may include assessing the lineage of the sample. In certain embodiments, the method may include identifying the geographic location of the subject.

[0057] In addition, as disclosed herein, in certain embodiments, the method may include uploading the results of the step of determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence to a repository for further classification (e.g., lineage determination) if a variant is detected. The repository may be the CDC database. Alternatively, other public repositories may be used.

[0058] The method may further include a step of determining whether a depository update has occurred prior to the step of determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence.

[0059] The method may be automated at various steps in the procedure. In certain embodiments, the method may be used with a Hamilton Star robot for sample plate setup. Additionally and / or alternatively, a Formulatrix Mantis liquid handler or other automated device may be used for master mix distribution. Also, as disclosed herein, the method may be program-implemented and / or may include the use of a computer-program product tangibly embodied in a non-transitory machine-readable storage medium that includes instructions configured to perform any of the steps of the method.

[0060] For example, in certain embodiments, the remaining total nucleic acid extract from SARS-CoV-2 positive RT-PCR diagnostic test samples with Ct values ​​< 31 is cherry-picked, for example, using a Hamilton STAR from the RNA extraction plate to a 96-well plate containing only positive samples, as disclosed in more detail herein. The samples can then be aliquoted into a sequencing run plate of 95 samples with one water no template control (NTC). The method can be scaled as needed. For example, in certain embodiments, 8 plates, or 760 specimens, can be processed in one production batch.

[0061] FIG. 1 illustrates an example embodiment of the method 100 of the present disclosure. As shown in FIG. 1, a sample for testing, optionally positive for SARS-CoV-2, is obtained from a subject 102. In an embodiment, a SARS-CoV-2 cDNA sequence is generated by RT-PCT 104. The SARS-CoV-2 cDNA can then be incubated with a series of tiling probes 106. In a particular embodiment, the tiling probes are relatively evenly spaced across the SARS-CoV-2 genome. After binding, the region in between the two probes can be filled in with DNA polymerase and ligated to form a closed circular molecule (i.e., a circular probe template) with the sample-specific SARS-CoV-2 nucleic acid sequence located between the two probe sequences 108. Unbound or incomplete loops remain linear molecules and can be removed by exonuclease digestion 110. Also, at this point, the circular probe template molecule can be released from the sample cDNA (i.e., during denaturation in a PCR amplification reaction) 110. The probes can then be linearized and enriched by amplification using the 3' M13 universal sequence and the 5' sample-specific barcode 112. The samples are then pooled in equal volumes 114 and libraries are prepared for sequencing 116. At this point, the samples are sequenced as disclosed herein and the results are analyzed 300 and reported 400.

[0062] FIG. 2 provides an alternative illustration of a portion of a method for capturing SARS-CoV-2 specific sequences for sequence analysis. Thus, in certain embodiments, a custom Molecular Loop SARS-CoV-2 capture kit is used to prepare samples for sequencing. In certain embodiments, as disclosed herein, sequencing is accomplished using a PacBio Sequel II sequencer. Thus, reverse transcriptase is used to synthesize cDNA from RNA. SARS-CoV-2 cDNA is then used as a target for hybridization of a molecular loop probe (FIG. 2, step 1). The molecular loop probe can consist of two binding sites that are tiled across more than 99% (e.g., 99.6%) of the 30 kb SARS-CoV-2 viral genome and are approximately 600 bp apart. After binding, the region in between the two probes is synthesized with DNA polymerase and ligated to form a closed molecule (FIG. 2, step 2). Unbound or incomplete loops remain linear molecules and are removed by exonuclease digestion (Figure 2, step 3). Circular molecules are then released from the template cDNA (Figure 2, step 4), then linearized (e.g., by digestion with X1) and enriched by amplification with 3' molecular loop-specific M13 universal sequences and 5' sample-specific barcodes (Figure 2, step 5). P1 and P2 in Figure 2 are adapters that can be used for barcode addition. Samples are then pooled in equal volumes and libraries are prepared for PacBio sequencing. Library preparation can entail DNA damage repair, ligation of sequencing adapters, removal of unligated products by enzymatic digestion, and purification by beads. Libraries can then be sequenced in a PacBio Sequel II, for example, using a 15-hour movie.

[0063] At this point, as shown in FIG. 3, sequencing and data analysis steps can be performed 300. Thus, in certain embodiments, the method can include computer-implemented steps for sequence analysis. In certain embodiments, whole genome sequencing is performed 302. In certain embodiments, sequencing utilizes single molecule real-time (SMRT) long-read sequencing technology. Circular templates generated from library preparation are combined with polymerase and primers and loaded into an SMRTCell (sequencing cell). The single molecule product diffuses into one of 8 million zero mode waveguide (ZMW) wells with a polymerase immobilized at the bottom. A phosphate-linked nucleotide is then introduced into the ZMW, and the base can then be incorporated by the polymerase. When a given base pair is incorporated, its addition produces a nucleotide-specific emission that is detected by a camera for each well. This process is repeated for a period of time, or movie length, and the order of nucleotides on a given well is analyzed and translated into the corresponding nucleotide in the long sequence read output.

[0064] The data can then be assembled as a sequence file 303. For example, in certain embodiments, PacBio SMRT LINK software and custom molecular loop processing scripts can be used to generate a FASTQ file for each sample. The FASTQ can be analyzed using a genomics analysis pipeline implemented using CLC genomics server version 6.5.6. Alternatively, other sequencing analysis systems can be used. At this point, the sequencing primer sequences can be removed 304 and the sequences can be aligned with the SARS-CoV-2 reference genome (e.g., NC_045512v2) to generate an alignment bam file 306. In certain embodiments, Minimap2 can be used to generate the alignment. Alternatively, other alignment programs can be used. In certain embodiments, samples that meet the 50% minimum coverage are then used as input data for calling variants and for generating sample-specific genome assemblies to generate consensus sequences for each sample 308. Alternatively, other minimum coverage limits (e.g., 20, 30, 40, 60, 70, 80, 90 percent) may be used. In an embodiment, the consensus sequence may be generated using VCFcons (available at www.biorxiv.org / content / 10.1101 / 2021.02.26.433111v1). Or, another algorithm may be used. In an embodiment, there is a defined threshold for generating a consensus sequence. For example, in certain embodiments, when VCFcons calls a nucleotide sequence for genome assembly, it must have at least 4 circular consensus sequencing (CCS) reads covering that base pair, and an alternative allele frequency compared to the reference sequence of >50%. If a nucleotide has less than 4 reads, it is reported as an N (undefined nucleotide) in the consensus sequence.

[0065] The assignment of sample lineages may take into account certain experimental variables and / or controls 310. For example, in certain embodiments, evaluation of an external no template control (NTC) is used to assess the validity of the results 310. Additionally and / or alternatively, an external positive template control (PTC) may be added to demonstrate sufficient processing of the plate 310. Further, in certain embodiments, unique strains (if available) from successful runs may be pooled by strain type, and each unique pooled strain may be added to plates across batches (e.g., a set of eight sequencing plates) to ensure plate origin across plate processing. An external no template control (NTC) may be required to ensure that no master mix contamination events are present on a given amplification plate. The NTC may include water (e.g., molecular grade water) added to a defined position (e.g., A1 position) of every 96-well positive plate prior to sample addition. Alternatively, other NTCs (e.g., buffers) may be used. The NTCs can then be transferred to a sequencing run plate along with the positive samples and subjected to sequencing and (quality control) QC analysis.

[0066] In certain embodiments, after sequencing, the strain typing of the positive control of a given plate can be compared with the recorded strain added before processing. Discrepancies between the strain typing of the assigned plate can be further investigated to determine whether to proceed with the individual plate. For example, in certain embodiments, failure to match the positive control result can lead to the removal of all strains associated with the plate of a given control. In other embodiments, failure of the reaction of the positive control does not necessarily lead to the removal of the result, if the corresponding control in the other plate of the batch can eliminate the possibility of plate replacement.

[0067] In certain embodiments, after sequencing and NTC analysis, the mean value of median CCS reads may be computer analyzed for passing a criterion of 10 CCS reads 310. In certain embodiments, to release a positive sample result for a given 96-well sequencing plate, the NTC must return a defined level of median mean CCS reads (e.g., <10). If the mean value of median CCS reads for a given NTC on the plate is higher than a defined level of CCS reads, all corresponding samples on the plate may be scheduled to be repeated.

[0068] At this point, lineages for individual samples can then be assigned using the consensus sequence 312. In an embodiment, this is accomplished as an input to the Pangolin analysis package. Alternatively, another analysis can be used. In a particular embodiment, strain lineage results are released for samples with 90% genome coverage and / or an average median read coverage of >10 circular consensus sequence (CCS) reads across the entire genome 314. In an embodiment, different CCS read metrics are based on the nucleotide level (4 CCS reads) and on the genome level (10 CCS reads).

[0069] In certain embodiments, assessment of strain determination results is performed after NTC analysis and removal of any samples on plates that fail the NTC. Individual sample results are then computationally examined for a median CCS read average >10 CCS and percent genome coverage >90%. In certain embodiments, test results may be reported to health care providers and relevant public health agencies as required by local, state and federal governments. In certain embodiments, samples that do not meet these criteria will fail the analysis and no strain typing will be reported. Additionally and / or alternatively, when only positive samples are tested, the method is not used to detect SARS-CoV-2 infection status where infection status is not dependent on the viral whole genome sequencing results.

[0070] 4. Data Analysis Analysis of sequence data may, in certain embodiments, include pre-processing (i.e., upstream) and post-processing (i.e., downstream) steps. In certain embodiments, at least some of these steps include computer-implemented steps for data analysis. Upstream analysis may include monitoring sequencer runs for completion, demultiplexing to generate individual sample FASTQ files, and triggering each alignment with the SARS-CoV-2 reference genome to generate alignments and variant calls. Downstream analysis for samples in each SMRTCell may consist of generating all results, including phylogenetic classification, for each sample.

[0071] Upstream analysis An example method 400 for upstream analysis of sequencing data is shown in FIG. 4. Thus, in certain embodiments, for analysis initiation, PacBio / Molecular Loop raw sequencing data is deposited and the generated CCS BAM file can be copied for demultiplexing. In certain embodiments, samples that fail on the sequencer do not generate data files. Those samples that are designated to be repeated do not continue sequence analysis.

[0072] At this point, the generation of individual sample FASTQ files can be performed. In an embodiment, the generation of CCS BAM files, demultiplexing, and generation of FASTQ files are performed as disclosed in the examples herein. Or, other methods can be used. Thus, in a certain embodiment, preprocessing can include at least a part of the following steps: generating circular consensus sequence (CCS) BAM files (402); merging intermediate BAM files (404); demultiplexing using the individual generated BAM files corresponding to different barcode combinations (406); combining the demultiplexed output data by sample name and / or patient identifier (408); removing barcodes from sequences and generating individual sample FASTQ files (410); aligning sequences with barcodes and trimming barcodes (412); converting BAM files to FASTQ files and copying FASTQ and CCS BAM files to final locations (414); and starting the CLC workflow (416).

[0073] The CLC analysis workflow may be performed using the following steps. First, the NGS data analysis workflow may be performed for each sample using the current validated CLC Genomics Server version 418. Then, for each sample's FASTQ file, the following steps may be performed. First, the reads may be filtered to retain reads between 250 and 5000 bp in length 420. Next, the reads are aligned to the SARS-CoV-2 reference genome (e.g., "NC_045512v2") 422. This alignment may be performed using minimap2 to generate a BAM file. Or other alignment methods may be used. At this point, local realignment may be performed and variant calling may be made 424. This may be performed using the Low Frequency Variant Detection tool in the CLC Genomics Server. Or other methods may be used. At this point, both the assembly (BAM file) and the detected variants (cf) are input into downstream post-processing analyses 426. A script detects the completion of the CLC process, which results in the initiation of downstream analysis of the samples in each SMRTcell.

[0074] Downstream analysis An example of a downstream (post-processing) analysis flow chart 500 is shown in FIG. 5. The steps of post-processing part 1 (501) may be as follows in certain embodiments: Using the appropriate reference file, VCFCons may be used to generate a consensus sequence based on sequence alignment and variant calling for each sample 502. For this analysis, a minimum coverage of 4CCS reads and a minimum substitution frequency of 0.5 may be used to assign bases to each genomic position. Alternatively, a different threshold may be applied. In certain embodiments, positions that do not meet this criteria are assigned an ambiguous base "N". Next, a sequence base composition may be generated 504. In certain embodiments, this may be used later to determine the percentage of non-ambiguous bases. In certain embodiments, this analysis may be accomplished using Seqtk or an alternative algorithm. At this point, the consensus sequence may be used as input data to generate any one or all of (a) clade assignment, (b) variant calling, and (c) sample sequence quality checks, as desired 506. In certain embodiments, Nextclade is used for this analysis. Alternatively, in certain embodiments, other algorithms may be used. At this point, a lineage is assigned 508. In certain embodiments, Pangolin assigns lineage to the consensus sequence by generating a SARS-CoV-2 lineage (known as Pango nomenclature) and then assigning a SARS-CoV-2 genome sequence lineage (Pango lineage). In certain embodiments, Pangolin is set to only consider genomes with at least 50% unambiguous bases. Finally, coverage statistics may be generated 510. In certain embodiments, SummaryStat compiles the results from Nextclade, Pangolin, and Seqtk and generates coverage statistics required for subsequent QC, including the average median amplicon coverage and genome coverage percent. Alternatively, another algorithm may be used for compilation.In a particular embodiment, for this analysis, the median base coverage in 29 overlapping 1.2 kb regions spanning the entire SARS-CoV-2 genome is calculated for each of the samples. Alternatively, other thresholds can be used. The distribution statistics of these coverage values ​​(minimum, first quartile, mean, median, third quartile, and maximum) can then be calculated for each sample. Also, the genome coverage percentage can be calculated as the number of unambiguous bases (A, T, C, G) divided by the total sequence length, and the phylogenetic classification is aggregated, and only samples that produce Nextclade results and Pangolin phylogenetic calls are retained for further processing.

[0075] At this point, the second part of post-processing (503) can be started. Thus, QC is performed again using the appropriate reference files, strain surveillance specific metadata 509, 510 (demographic data, genome coverage percentage, and Ct values ​​from RT-PCR assays) and the data is added to the results 512. In an embodiment, samples lacking metadata are excluded from the result set 516. Also, no template QC is performed based on no template controls (NTCs) 516. Also, in a particular embodiment, if the average median coverage of the 29 genomic regions is >10 CCS reads, all sequenced samples on the same plate are removed 516. Finally, coverage QC is performed 516. In an embodiment, samples with genome coverage >=90% are retained in the results. Also, in an embodiment, samples with average median coverage >10 CCS reads were retained in the results. The results can then be transferred to a Report System location 514 to generate a patient report with the corresponding Pangolin lineage. In one embodiment, samples that fail to produce a result are reported as not being able to determine the lineage. SARS-CoV-2 virus may be detected but no lineage information may be reported.

[0076] In certain embodiments, the lineage calling criteria may be as follows: Inclusion criteria: (1) CT<31; (2) corresponding metadata (strain surveillance); (3)>90% genome coverage; (4) mean median coverage>10 CCS reads; (4) passed NTC controls; and (5) Nextclade results and Pangolin lineage calls. Exclusion criteria: (1) CT>31; (2) missing metadata (strain surveillance); (3)<90% genome coverage; (4) mean median coverage<10 CCS reads; and (4) failed NTC controls.

[0077] Revalidation In certain embodiments, the assay is revalidated in response to the emergence of new variants. In certain embodiments, at least some of these steps include computer-implemented steps for revalidation analysis. In certain embodiments, revalidating the classification accuracy of the Virseq assay 600 in response to the emergence of new variants (i.e., lineages) of the SARS-CoV-2 virus and the accompanying changes to the pangolin classification software can be performed as depicted in FIG. 6 (see also Example 5). The analysis depicted in FIG. 6 was developed for pangolin, but can be applied to other databases for phylogenetic assignment of viruses. The pangolin software is distributed by Dockerhub (at hub.docker.com / r / staphb / pangolin). Thus, in certain embodiments, the pangolin site can be monitored 602 and checked for updates by downloading and installing an updated docker container at regular intervals (e.g., weekly, biweekly, monthly) 604. In certain embodiments, if there are no updates, nothing needs to be done 606.

[0078] If there is an update, a regression analysis can be performed 601 using the in-house laboratory data. In an embodiment, a new pangolin version can be used to determine the lineage of the in-house reference samples 608 610. The reference sample set 608 can include data from various sets (e.g., based on date of outbreak and / or COVID type). For example, the data set can be defined to be primarily delta variant and / or omicron lineages. Or other types can be analyzed. In an embodiment, each sample in the reference set includes its consensus sequence, as well as its lineage classification history performed by previous pangolin versions. The reference sample set 608 can be updated periodically to include samples representative of newer, more prevalent lineages as pangolin versions are updated.

[0079] The format of the pangolin software output may then be compared to that of the previous version to determine if there have been any changes in the pangolin output format 612. If there have been any changes, these changes may be recorded and the laboratory pipeline may be modified to accommodate the changes. The modified version may then be deployed to the QC environment for testing 614. Any changes in lineage calls may then be assessed and compared to those expected from the software update change records 616. For example, in one particular embodiment, expected changes include reassignments between sub-lineages. If there have been any unexpected changes in lineages (e.g., changing delta sub-lineage to alpha), these are scrutinized and recorded 618.

[0080] At this point, a second regression test can be performed using the published (GISAID) sequences and their metadata603. or other public databases can be used. For this analysis, the latest GISAID sequences are downloaded, metadata for all GISAID sequences and the pangolin lineage are obtained, and the list of variants of concern (VOC) (i.e., variants being actively tracked by CDC and / or other health organizations) and variants of interest (VOI) (i.e., variants being monitored by CDC and / or other health organizations) can be updated based on the WHO updates and the latest complete list of lineages620. Next, a data simulator can be used to model the coverage and error characteristics of the in-house assay622. In an embodiment, the simulator uses the GISAID sequences as a starting point and imposes simulated coverage and error based on empirical coverage profiles and maximum-minor allele frequencies from a collection of in-house samples. The resulting simulated samples are run through pangolin and the lineage classification is compared to that of the original GISAID sequence. Classification stability is defined as the rate at which mutated sequences maintain their expected lineage classification. In an embodiment, two experiments on regression are performed to assess classification stability by simulation. Thus, the method can randomly sample up to 100 GISAID sequences for each VOC / VOI to assess the classification stability of these important lineages regardless of their frequency in available sequencing data624. Or, more or less GISAID sequences (e.g., 50, 200, 400, 500 or more) for each VOC / VOI can be sampled according to the needs of analysis. This can allow the classification stability of emerging variants and new sub-lineages of existing ones to be assessed. In addition, the method can randomly sample 10,000 GISAID sequences from the database for retrospective analysis based on frequency of lineage classification stability626.Alternatively, more or fewer retrospective GISAID sequences can be sampled, depending on the needs of the analysis. This may allow stability to be quantified relative to historical retention.

[0081] The output data of the data simulator experiment is then reviewed to check for unexpected changes in classification stability with respect to previous regression testing using GISAID data 628 and retrospective data 630 for the VOC / VOI data. In certain embodiments, any unexpected instability is investigated and recorded 632. In certain embodiments, an upgrade is accepted if certain parameters are met. In some cases, an upgrade is requested if the median VOC / VOI match between the simulation data and the reference sequence is at least 90% 640. If these criteria are not met, additional investigation may be required.

[0082] In certain embodiments, if the new mismatched lineage is new, the sample can be examined for confirmation. If the mismatched variant is not a new variant, the method can include further investigation to find the root cause of the mismatch. This can include examining the coverage of the reference sequence as well as the simulated sequence to ensure that it is not an undesired drop in base coverage in a particular region. Additionally and / or alternatively, this can include rerunning the simulation with a different seed to see if the mismatch reoccurs. If it does, the upgrade can be aborted.

[0083] At this point, the new variants can be assessed using the methods and systems disclosed herein650. For successful surveillance of emerging variants (lineages), it may be beneficial to review the potential impact on molecular loop inversion probe amplification by performing in silico analysis. Thus, the method may further include identifying the individual sequence variants of the emerging lineage and the location of the associated molecular loop probe to assess the possibility of interfering with probe binding. In an embodiment, a conservative assumption is used that new sequence variants that overlap with any probe will affect hybridization. Additionally and / or alternatively, adjacent probes within the region may be reconsidered to ensure coverage of the new sequence variant. For any sequence variants that may result in reduced coverage within a particular region, the affected probes in the pangolin lineage update validation summary are recorded.

[0084] NGS System for Determining SARS-CoV-2 Strains Systems for performing the methods herein are also disclosed. For example, the system may include a station or component(s) for performing various steps of the method. In certain embodiments, the station or component may include a robotic or computer controlled station or component for performing the step(s) of the method. In certain embodiments, systems are disclosed for performing at least some of the steps of: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence includes a SARS-CoV-2 variant sequence.

[0085] Thus, the system may include a station or component for obtaining a sample for testing. The sample may be of unknown COVID status or may be a sample that has previously tested positive for COVID. In certain embodiments, a positive sample may be identified using an approved EUA approved COVID-19 RT-PCR Test (e.g., Labcorp EUA200011 and / or EUA203057). Thus, the result is for the identification of the SARS-CoV-2 strain infecting the individual following detection of viral RNA in the sample.

[0086] In certain embodiments, the system may include a station or component for performing a step of generating sample-specific SARS-CoV-2 nucleic acids, including generating sample-specific SARS-CoV-2 cDNA using reverse transcriptase polymerase chain reaction (RT-PCR). The system may also include a station or component for hybridizing one strand of the sample SARS-CoV-2 cDNA with a single-stranded probe DNA template that includes a pair of SARS-CoV-2 probes, a first probe located at the 3' end of the probe DNA template and a second probe located at the 5' end of the probe DNA template. In certain embodiments, the probe sequences are selected as tiled probes that bind at intervals along the SARS-CoV-2 genome. For example, in alternative embodiments, the probes may be spaced apart by about 100, or 200, or 300, or 400, or 500, or 600, or 700, or 800, or 900 base pairs, or more than 1,000 base pairs. Alternatively, intervals within this range (e.g., 450, 550, 650, or 750) may be used. The probes may be tiled across more than 99% (e.g., 99.6%) of the 30 kb SARS-CoV-2 viral genome. Also, in certain embodiments, the single-stranded probe DNA template further comprises a universal sequencing primer (e.g., an M13 primer) located within the probe sequence. In addition, the single-stranded probe DNA template may further comprise an adapter sequence for the addition of a barcode sequence used to correlate SARS-CoV-2 sample-specific nucleic acids with sample numbers.The system may also include stations and / or components for filling in the sequence between the two probes to generate a circular single-stranded probe DNA template that contains a sequence specific to the sample SARS-CoV-2 cDNA between the two probe sequences, then liberating the circular single-stranded probe DNA template that contains a sequence specific to the sample SARS-CoV-2 cDNA from the sample-specific SARS-CoV-2 DNA, and digesting the circular single-stranded probe DNA template that contains a sequence specific to the sample SARS-CoV-2 cDNA to generate a linear DNA to be used as a template for nucleic acid sequencing. In certain embodiments, the system may include stations and / or components for modifying the linear probe DNA template to add adapters and then amplifying the linear DNA template for DNA sequencing. In certain embodiments, the enrichment step includes a purification step (e.g., purification by beads).

[0087] The system may further include a station and / or components for DNA sequencing. In certain embodiments, the method uses whole genome sequencing. In certain embodiments, next generation sequencing (NGS) is used. Or other types of sequencing, such as, but not limited to, Sanger sequencing, shotgun sequencing, SMRT sequencing, pyrosequencing or nanopore sequencing. For example, in certain embodiments, PacBio whole genome sequencing can be used with the corresponding SMRT link 9 software and analysis tools.

[0088] The system may further include a station and / or component for data analysis. Thus, the system may include a station and / or component for determining whether a nucleic acid sequence contains a SARS-CoV-2 variant sequence by aligning the sample SAR-CoV-2 sequence with the SARS-CoV-2 reference genome to generate a sample-specific assembly and consensus sequence, and / or for assessing the lineage of the sample. In certain embodiments, the system may include a station and / or component for identifying the geographic location of the subject.

[0089] Additionally, as disclosed herein, in certain embodiments, the system may include a station and / or component that may include uploading the results of the step of determining whether a nucleic acid sequence contains a SARS-CoV-2 variant sequence to a repository for further classification if a variant is detected. The repository may be the CDC database. Alternatively, other public repositories may be used.

[0090] As disclosed herein, the system may include a station and / or component for determining whether a depository update has occurred prior to the step of determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence.

[0091] The system may include stations and / or components for automating various steps of the procedure. In certain embodiments, a Hamilton Star robot may be used for sample plate setup. Additionally and / or alternatively, a Formulatrix Mantis liquid handler or other automated device may be used for master mix distribution.

[0092] FIG. 7 illustrates an embodiment of a system 700 for performing any of the method steps of the present disclosure. As shown in FIG. 7, the system may include a station or component 702 for obtaining a sample for testing. In certain embodiments, the sample is positive for SARS-CoV-2. In certain embodiments, the system includes a station 704 of a component for generating a SARS-CoV-2 cDNA sequence by RT-PCT. The system may further include a station or component 706 for incubating the SARS-CoV-2 cDNA with a series of tiled probes. In certain embodiments, the tiled probes are relatively evenly spaced across the SARS-CoV-2 genome. For example, in alternative embodiments, the probes may be spaced about 100, or 200, or 300, or 400, or 500, or 600, or 700, or 800, or 900 base pairs, or more than 1,000 base pairs apart. Alternatively, spacing within this range (e.g., 450, 550, 650, or 750) may be used. The probes may be tiled across more than 99% (e.g., 99.6%) of the 30 kb SARS-CoV-2 viral genome. After binding, the region in between the two probes may be filled in with DNA polymerase and ligated to form a closed molecule. This may be done in the same station as the step of incubating with the tiled probes or in a different station, using different components 708. Unbound or incomplete loops remain linear molecules and may be removed by exonuclease digestion. This may be done in the same station as the step of incubating with the tiled probes or in a different station, using different components. The system may further include a station 710 for release of circular molecules from the template cDNA and enrichment by amplification using the 3' M13 universal sequence and 5' sample-specific barcode. The system may further include stations and / or components 712 for sample pooling and library generation.

[0093] The system may further include stations and / or components for sequencing DNA 714, and stations and / or components for contig alignment and variant identification 716, using the methods disclosed herein. The system may also include stations and / or components for verifying and reporting results 718, as disclosed herein.

[0094] As illustrated herein, any of the method steps, stations or components may be automated, robotically controlled, and / or controlled, at least in part, by the computer 800 and / or programmable software. Thus, the system may include a computer-program product tangibly embodied in a non-transitory machine-readable storage medium that includes instructions configured to execute the system or any portion of the system (e.g., a station or component) and / or perform a step(s) of the method of any of the disclosed embodiments. In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods or processes disclosed herein and / or perform any of the portions of the system disclosed herein.

[0095] For example, a system is disclosed that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations that direct at least one of the following steps: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating a sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence.

[0096] Also disclosed is a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to execute the system and / or perform a step or steps of any of the methods of the disclosed embodiments. For example, in certain embodiments, the computer program product tangibly embodied in a non-transitory machine-readable storage medium includes instructions configured to cause one or more data processors to perform operations directing at least one of the following steps: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence includes a SARS-CoV-2 variant sequence. Additionally and / or alternatively, in certain embodiments, a computer program product tangibly embodied in a non-transitory machine-readable storage medium includes instructions configured to cause one or more data processors to perform operations directing at least one of: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating a sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence comprises a SARS-CoV-2 variant sequence.

[0097] The systems and computer products may perform any of the methods disclosed herein. One or more embodiments described herein may be implemented using a programmatic module, engine, or component. A programmatic module, engine, or component may include a program, a subroutine, a portion of a program, a software component, or a hardware component that can perform one or more defined tasks or functions. As used herein, a module or component may exist on a hardware component independent of other modules or components. Alternatively, a module or component may be a shared element or process of other modules, programs, or machines.

[0098] FIG. 8 shows a block diagram of an analysis system 800 used for detecting and / or quantifying pathogens. As shown in FIG. 8, one or more processor-executable modules, engines, or components (e.g., programs, codes, or instructions) may be used to implement various subsystems of the analyzer system according to various embodiments. The modules, engines, or components may be stored in a non-transitory computer medium. As needed, one or more of the modules, engines, or components may be loaded into a system memory (e.g., RAM) and executed by one or more processors of the analyzer system. The example depicted in FIG. 8 shows modules, engines, or components for implementing the methods of the present disclosure.

[0099] Thus, FIG. 8 illustrates an example of a computing device 800 suitable for use in systems and methods according to the present disclosure. The example computing device 800 includes a processor 805 in communication with a memory 810 and other components of the computing device 800 using one or more communication buses 815. The processor 805 is configured to execute processor-executable instructions stored in the memory 810 to perform one or more methods or operate one or more stations or components for detecting pathogen levels according to different examples such as those shown in FIGS. 1-7 or disclosed elsewhere herein. In this example, the memory 810 may store processor-executable instructions 825 that can analyze 820 the results of a sample or test unit confirmation as discussed herein.

[0100] The computing device 800 in this example may also include one or more user input devices 830, such as a keyboard, mouse, touch screen, microphone, etc., for accepting user input. The computing device 800 may also include a display 835, such as a user interface, for providing a video output to a user. The computing device 800 may also include a communication interface 840. In some examples, the communication interface 840 may enable communication using one or more networks, including a local area network ("LAN"); a wide area network ("WAN"), such as the Internet; a metropolitan area network ("MAN"); a point-to-point or peer-to-peer connection, etc. Communication with other devices may be accomplished using any suitable network protocol. For example, one suitable network protocol may include the Internet Protocol ("IP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), or a combination thereof, such as TCP / IP or UDP / IP. EXAMPLES

[0101] Certain embodiments of the disclosed methods and systems are provided in greater detail in the Examples herein below.

[0102] Example 1 Overall methodology and analysis of results Next generation sequencing (NGS) can be used to perform surveillance testing on a large number of samples and to generate a sufficient number of viral genomes to track mutations and variants of concern as they arise. The overall testing principle is as follows: First, cDNA is prepared from viral RNA using random priming for first strand synthesis. Then, an inverted probe is annealed to the target during a 16-hour hybridization. Then, the gap is filled by polymerization and ligation. Then, unreacted linear probe is removed and the probe is released from the target DNA. Then, the captured targets are enriched by PCR amplification using asymmetric barcodes. Then, the PCR products are pooled, quantified, and SMRTbell hairpin adapters are ligated to the amplicons and sequenced on a Pacific Biosciences Sequel II using a 15-hour movie.

[0103] At this point, lineage calls are made based on processing of the NGS sequence data. For this analysis, any condensation positives that are "cherry picked" (discussed in more detail herein) include a no template control (NTC) (i.e., molecular grade water) in well A1 of the 96-well condensation positive plate. The NTC results are reviewed prior to generation of the result file for a given SMRTcell. If an NTC is found to be invalid, results for all patient samples on the affected plate will not be reported. Upon completion of processing of the NGS results for a given sequencing cell, a result file is generated and saved. At this point, FASTQ files for each of the samples are generated using PacBio SMRTLNK software and a custom molecular loop processing script. The FASTQ results are analyzed using a genomic analysis pipeline implemented in CLC genomics server version 6.5.6. The workflow starts with the sample-level fastq files, primer trimming, and then aligning to the SARS-COVID19 reference genome ("NC_045512v2") using Minimap2 to generate an alignment bam file. After checking coverage, the bam files are used as input data for calling variants and for generating sample-specific genome assemblies. Consensus sequences for each sample are generated using "VCFcons", which requires coverage of 4CCS reads and 50% alternative allele frequency at each base. The Pangolin package is used to assign lineages for individual samples.

[0104] Example 2 VirSeq SARS-CoV-2 NGS Strain Determination Method and Determination Validation Genomic sequencing of SARS-CoV-2, the virus that causes COVID-19, can determine the specific strain of SARS-CoV-2. Strain information can potentially provide clinicians and epidemiologists with useful information to aid in the public health response or future clinical treatment of the virus. The determination of a given strain is based on a combination of multiple genomic changes detected from comparison of DNA sequencing results to the original Wuhan reference strain. This approach allows for the identification of any new and emerging strains of SARS-CoV-2 without revalidation as the virus changes over time. The intended use of this assay is to provide a SARS-CoV-2 lineage or strain call for samples that yield at least 90% genome coverage.

[0105] The overall test principle is as follows: Residual total nucleic acid extracts from residual SARS-CoV-2 NAA diagnostic test positive samples were cherry-picked from the running plate to a condensed positive plate using Hamilton STAR and aliquoted into 96 sequencing running plates, with 8 plates or 768 samples in one production batch. Samples were then processed to PacBio sequencing using Molecular Loop Viral RNA Target Capture in PacBio. First, cDNA was synthesized from RNA using the specific Thermo Fisher VILO reverse transcriptase from the Loop kit. SARS-CoV-2 cDNA was then used as a target to anneal to molecular loop probes as outlined in Table 1. Molecular loop probes were tiled across the entire 30KB SARS-CoV-2 genome. These probes contain two binding sites that are roughly 600 bp apart.

[0106] [Table 1]

[0107] After binding, an approximately 600 bp region in the middle of the two probes was synthesized with DNA polymerase and ligated to form a closed molecule for another 60 min using the hybridization conditions in Table 1. Unbound or incomplete loops remained linear molecules and were removed by exonuclease digestion (i.e., sample cleanup). Incubation times for cleanup are shown in Table 2. Samples were stored at -20°C if not used for the next step within approximately 2 hours.

[0108] [Table 2]

[0109] The resulting circular molecules (containing sample-specific SARS-CoV-2 nucleic acid inserted between the two probe sites) were then released from the template cDNA and PCR amplified using sample-specific barcodes. The conditions used for PCR amplification are shown in Table 3. 768 asymmetric barcode combinations are required to process one batch (i.e., 768 samples and controls). To do this, a plate of M13 barcoded primers was prepared (Figure 9), then ten 96-well plates were created by adding 20 μL of M13 forward barcoded primer and 30 μL of M13 reverse barcoded primer (Figure 10). As shown in Figure 9, columns 1-4 were M13 forward primers with barcodes 1001-1032 in the tail, and columns 7-10 were M13 reverse primers with barcodes 1049-1079 and 1082 in the tail. Each of the 32 forward primers was combined with a different reverse primer to create asymmetrically barcoded pairs as shown in FIG. 10 for one example plate.

[0110] [Table 3]

[0111] Samples were then pooled (or stored at -20°C until pooling was accomplished). For pooling, eight reaction plates were removed from storage, spun down, and an aliquot (e.g., 5 μL) of each reaction was transferred to an 8 mL tube. Typically, 768 samples plus controls were pooled prior to sequencing.

[0112] At this point, samples were purified using bead purification with AMPure PB beads. Using the 500 μL pool, AMPure PB Bead (0.70X) cleanup was accomplished by adding 350 μL PB AMPure beads, mixing, centrifugation to pellet the beads, incubating for 5 minutes at room temperature, centrifugation, and magnetic separation to recover the beads. The supernatant was removed, the beads washed with 80% ethanol, and the DNA was eluted from the beads with elution buffer and quantified.

[0113] At this point, 1000ng of pooled DNA was used to prepare an SMRTbell library. The pooled DNA was mixed with buffer (DNA Prep Buffer), NAD, DNA Damage Repair Mix v.2 and incubated at 37°C for 30 minutes. After returning to 4°C, end repair was accomplished by adding End Prep Mix, Reaction Mix 1, incubating at 20°C for 30 minutes, 65°C for 30 minutes, and then returning the reaction to 4°C. At this point, adapters were added using Reaction Mix 2, Overhand Adapter v3, Ligation Mix, Ligation Additive, and Ligation Enhancer, and the samples were incubated at 20°C for 60 minutes to ligate the probe constructs and at 65°C for 10 minutes to inactivate the ligase, then returned to 4°C. An enzymatic cleanup was then performed and the samples were purified with AMPure (0.6X) beads as above using 100uL of elution buffer. The AMPure bead cleanup was then repeated using a small volume (20 uL) of elution buffer and DNA was quantified.

[0114] Samples were sequenced on a PacBio Sequel II. Each 96-well plate in a batch of 10 requires a unique set of asymmetric barcodes. Figure 10 shows an example of a map for one plate. Other combinations were used for the other plates, such as using 10 plates with unique combinations. For example, in a second plate, wells A1-H4 would have a combination of M13 forward primers 1001-1032 and M13 reverse primer 1052, wells A5-H8 would have a combination of M13 forward primers 1001-1032 and M13 reverse primer 1053, wells A9-H12 would have a combination of M13 forward primers 1001-1032 and M13 reverse primer 1054, and so on for further plates.

[0115] After sequencing, FASTQ files were generated for each sample using PacBio SMRTLNK software and a custom molecular loop processing script. FASTQs were analyzed using the genomics analysis pipeline implemented in CLC genomics server version 6.5.6. The workflow started with the sample-level fastq file, primer trimmed, and aligned to the SARS-COVID19 reference genome ("NC_045512v2") using Minimap2 to generate an alignment bam file. After checking the coverage, this bam file was used as input data for calling variants and for generating sample-specific genome assemblies. A consensus sequence for each sample was generated using "VCFcons", which requires coverage of 4CCS reads and 50% alternative allele frequency at each base. The Pangolin package was then used to assign lineages for individual samples and obtain the results.

[0116] The following controls were included: A no template control (NTC) was included on each plate during every step run to demonstrate the absence of contamination across samples and reagents. This control was analyzed by sequencing. Failing NTCs were samples that resulted in strain calls with 90% genome coverage. A positive control was included on each plate of the run. For validation, a previously run sample was used as a positive control. Metrics for determining whether a sample passed or failed included percent genome coverage, minimum depth of coverage, and strain lineage call resolution rate.

[0117] Specimen requirements were as follows: extracted nucleic acid derived from a sample with a positive result from an EUA-authorized SARS-CoV-2, NAA test with a CT of less than 26 for an approximately 90% success rate. Metadata samples with a higher CT or no CT were deemed acceptable but increased the risk of failing to report a result.

[0118] Acceptable outcome metrics were: >90% genome coverage and mean median read coverage >10 CCS reads.

[0119] A. Results Precision (repeatability): Intra-assay Intra-assay repeatability was assessed in triplicate runs of 11 nucleic acid samples of various presumptive typing from current SARS-CoV-2 CDC surveillance testing. Samples ranged in C-values ​​and broad read counts in the original run. Additionally, samples were diluted 1:4 to allow sufficient total nucleic acid input for all intra- and inter-assay runs. The criteria was defined as ≥95% repeatability for all strains reaching the reporting threshold of ≥90% coverage of the SARS-CoV-2 genome.

[0120] Strain calls, percent genome coverage (expressed as percent missing), and read counts were compared (Table 4). Strain calls for all 11 samples were 100% concordant across three replicates, and all replicates met the 90% genome coverage and sufficient read depth criteria. The 95% accuracy criteria for strains with 90% genome coverage were met.

[0121] [Table 4]

[0122] Precision (reproducibility): Interassay Inter-assay repeatability was assessed in triplicate runs of 10 nucleic acid samples of various presumptive typing from ongoing SARS-COV-2 CDC surveillance. Samples were identical to those used in intra-assay runs, with one sample missing due to unintentional exclusion from the final run. Samples ranged in C-values ​​and a wide range of read counts in the original runs. Additionally, samples were diluted 1:4 to allow sufficient total nucleic acid input for all intra- and inter-assay runs. The criterion was defined as ≥95% repeatability for all strains reaching the reporting threshold of ≥90% coverage of the SARS-CoV-2 genome.

[0123] Strain calls, percent genome coverage (expressed as percent missing), and read counts were compared (Table 5). There were no discordant results for all 10 replicates. Nine samples produced the expected lineage call across triplicates. One sample, purposefully selected for borderline read coverage, failed to produce a result at any time due to lack of genome coverage. Overall, 93% of samples produced identical strain typing, and 100% of samples released accurate results that met the adjudication criteria.

[0124] [Table 5]

[0125] Matching degree Relative accuracy was established by direct comparison of results with those generated by alternative methods. There were two methods used for positive comparisons: Illumina sequencing and amplicon Pacbio sequencing. Pacbio is a validated sequencing technology, while the Molecular Loop method is mechanistically different in its presequencing step from traditional amplicon sequencing. Negatives were sequenced with Illumina only.

[0126] The individual comparative studies used are listed below. Illumina Artic sequencing 93 negative NAA samples 72 SARS-CoV-2 samples sequenced in the winter of 2020 29 samples previously sequenced with CMBP in the molecular loop 50 samples previously sequenced with DNA Identification in Molecular loop Pacbio Amplicon Sequencing 122 samples were amplicon sequenced at 90% coverage and run in duplicate on the Molecular Loop

[0127] To set a baseline for minimum read coverage, a minimum read coverage threshold set at 4CCS reads was established using 110 NTCs from the validation run and ongoing RUO strain surveillance run during the validation timeline. For Illumina concordance, 93 negatives were randomly selected from the CMBP NAA diagnostic production run and re-extracted after initial testing to ensure sufficient volume. 72 samples of strains circulated in the winter of 2020 and sequenced in January 2021 were re-sequenced in Molecular Loop. Additionally, 79 samples from CMBP and DNA were selected for Illumina parallel testing based on initial strain call, CT and read coverage to ensure diversity. 382 samples originally amplicon sequenced in Pacbio were re-processed in duplicate in Molecular Loop. The duplicate replicates differed only slightly in their composition of the Thermo Fisher VILO RT master mix, which we have previously shown to be equivalent. The initial amplicon sequencing run yielded 122 of 382 samples with >90% genome coverage. Only samples with an initial 90% coverage were used for further analysis of molecular loop results.

[0128] The criteria was ≥95% accuracy for all strains reaching a reporting threshold of ≥90% coverage of the SARS-CoV-2 genome.

[0129] Read Coverage Threshold: The mean, minimum and maximum average values ​​of median amplicon coverage, referred to here as mean read coverage, were analyzed for validation and production runs. The distribution of mean read coverage is shown in FIG. 11. The minimum, maximum and average were 0.24, 9 and 1.7 CCS reads, respectively. A typical threshold was set by 3X standard deviation plus the average, which is equal to 6.67 for this analysis. However, for rapid processing, the cutoff read coverage for base pairs used for strain typing was set at 10 CCS reads, or 10% plus maximum NTC.

[0130] Illumina Artic Sequencing: 93 samples previously determined to be negative were sequenced in parallel in duplicate in Molecular Loop and in Illumina. For the two samples resulting in a reported genome in Illumina and in Molecular loop, there was 98.3% concordance between the two technologies. Further investigation revealed that both samples obtained in Illumina were in fact positive for nucleic acid amplification (NAA) and were erroneously included in the validation. The average read count for the other 91 samples in duplicate further confirmed the conservative read depth threshold (Figure 12).

[0131] 72 samples previously sequenced in Illumina from strains circulating in January 2021 were resequenced in Molecular loop. 79 samples previously sequenced in Molecular loop at CMBP and DNA identification were resequenced in parallel in Molecular loop and Illumina to represent the current strains circulating. After removal of samples damaged during transport between testing sites or that failed the comparative sequencing reaction, 123 strain calls successfully generated on both platforms were analyzed (Figure 13).

[0132] Of the 72 samples, 51 met the QC thresholds of 10 CCS read depth and 90% genome coverage for a 71% success rate. This reportable genome rate was similar to the CDC strain surveillance reported genome rate of 72.5% DNA during the month of validation. All strain results of the 51 reportable results were 100% concordant (Table 6). Inclusion of results with a read depth of less than 10 CCS on average resulted in 10 / 13 (77%) matched strain results and 61 / 64, 95.3%, exact concordance. Analysis of samples with less than 90% genome coverage had only 2 / 7 identical strain results.

[0133] [Table 6]

[0134] Of the 72 samples originally sequenced with CMBP and DNA in Molecular Loop, 66 were successfully sequenced with sufficient read depth on Illumina and Molecular Loop after dilution. All 72 samples produced 90% coverage, of which 71 strain typings had identical strain calls to the Illumina reference method. Overall, all reported results were 100% concordant between parallel techniques (Table 7).

[0135] [Table 7]

[0136] Amplicon Pacbio Sequencing: Of 122 samples that had 90% coverage with amplicon sequencing, 116 were replicated at 90% coverage for both replicates. There was 100% concordance between the 116 molecular loop repeat strain typings. Overall, parallel testing between molecular loop and traditional amplicon sequencing was 98.2% concordance.

[0137] Analysis sensitivity / specificity The heat-inactivated SARS-CoV-2 strains B.1.1.7 (VR-3326HK™), Hong Kong / VM20001061 and Italy-INMI1 genomes have been characterized by ATCC. For analytical sensitivity, all variants were identified using an analytical pipeline and compared to the public ATCC strain variant dataset. Traditionally, in human genome sequencing, notable variants are analyzed, validated, and the false discovery rate (FDR), false positives (FP), are normalized to all positive calls (FP+TP, where TP=true positive) rather than all negatives. However, with Sars-Cov-2, there are 4-20+ variant combinations at defined positions in a given strain that result in a strain call. Also, due to full genome sequencing of viral RNA, there are multiple highly repetitive regions known to cause changes in sequencing data that are not relevant for current strain typing. Strain typing programs such as Nextclade can take this into account. Therefore, sensitivity was determined by dividing the number of called variants recorded for a strain by the total number of variants. In addition to FDR, specificity was calculated by the total number of false variants called compared to correctly sequenced base pairs. Assembled genomes were used as input data in Pangolin, which calls variants and outputs strain typing. No variants were called, and strains were output only for further analysis. Therefore, sensitivity and specificity were calculated using variant calls from another genome variant caller, CLC, and the Nextclade Sars-Cov-2 specific variant caller, which takes into account repetitive and difficult to sequence viral regions when making variant calls. The criteria were: (1) analytical sensitivity of ≥ 90% in control RNA for variants in segments above the minimum coverage; and (2) analytical specificity of ≥ 90% in control RNA for variants in segments above the minimum coverage.False discovery rates in control RNA for variants in segments above the minimum coverage were recorded, but no cutoff criteria were set.

[0138] Both variant calling platforms were sensitive in terms of their ability to detect variants, with a sensitivity of 96.23% for Nextclade and 98.11% for CLC (Table 8). When comparing the overall specificity of determining base pairs across the genome, both were >99.9% specific. However, Nextclade's ability to adjust for repetitive, difficult to sequence, regions was evident from the number of false variants detected at 3 compared to CLC without the adjustment process at 23. This led to a discrepancy in the variant false discovery rate, which was 5.36 for Nextclade and 30.26 for CLC, indicating that the false variant discovery was from difficult to sequence repetitive viral regions, which are not relevant for current strain typing. In summary, all criteria were met and the Molecular Loop process is sensitive and overall specific, with a high FDR from viral regions not analyzed by current strain typing algorithms.

[0139] [Table 8]

[0140] Assay Tolerance Assay tolerance for nucleic acid input can be thought of as the tolerance to changes in the amount of analyte added to the reaction. Although usually expressed in cp / μL, approximately 80% of the samples assayed will be from EUA NAA SARS-CoV-2 tests providing a corresponding cycle threshold (CT) value for each sample. Therefore, we used CT instead of cp / μL as the input metric for analysis and guidance. Sequencing viral genomes from residual NAA tests inherently has a high failure rate, which is directly related to the viral titer and RNA integrity of the specimen and can vary dramatically between samples. The failure rate is driven by the RNA titer (CT), but if the background is set conservatively, the increased failures observed in samples with higher CT will not result in discrepant results but simply increase the cost of the assay. The purpose of this validation assay tolerance experiment was to set a baseline of expected success rate at a given CT, but not restrict which samples will attempt to have their genomes sequenced.

[0141] 9,718 production results across three sites were analyzed for their success rate in yielding results based on their Nucleocapsid Target #1 (N1) CT value. First, samples were binned at 90% by their ability to yield genomes, and CT values ​​were rounded up to the nearest whole number in 1-integer increments. For example, a CT of 30.1 was calculated below the 31 bin. All samples with a CT of <16 were included in the 16 bin. All samples lacking CT metadata were removed from the analysis. The criteria were: (1) manufacturer recommendation for 10,000 copies of RNA for sequencing with acceptable variation in input concentration used meets the following criteria for analysis; and CT groups of at least 20 samples.

[0142] More than 8815 / 9718 producing samples had corresponding CT metadata. There was no drop in ability to generate approximately 90% genome coverage from the <16-24 CT analysis bins (Figure 14). Starting at 25 CT, there was a steep drop in the ability of samples at a given CT to generate genomes at 90% coverage, with the 30.1 → 31 CT bin reporting only 7.65% samples. If a CT cutoff for assay input data is required, it is recommended that samples should be <26 CT and they had a 78% genome production rate.

[0143] Analyte stability The samples in this validation were stored for a minimum of 4 weeks beyond the period the samples were tested in the clinical laboratory. Long-term stability should be determined by storing at least three aliquots under the same conditions as the study samples. The volume of the sample should be sufficient to analyze on three separate occasions. The stability of the analyte in the biological matrix at the intended storage temperature should be established.

[0144] Analyte stability under various storage conditions was established by measuring concordance at various storage lengths. After NAA diagnostic testing, extracted nucleic acid was shipped on dry ice to the testing laboratory and stored at -20°C until sequencing. All samples used for validation were residual production samples. The stability experiment described below is in addition to the process of collecting samples and shipping them to the sequencing laboratory. Analyte stability was measured in two separate experiments. In the first experiment, 10 samples used for interassay precision were thawed, assayed, and refrozen three times over a one-month time point. Samples represented a range of strains, CT values, and original read coverage. In the second experiment, 12 samples containing alpha, beta, and delta variants of concern (VOCs) with a range of original read sequence depths were resequenced after one month of storage at -20°C, which entailed three freeze-thaw cycles. Criteria were defined such that storage conditions were considered suitable if samples yielded the same strain detection after a defined length of storage and yielded results with ≧90% accuracy for the reportable strains.

[0145] Inter-assay results used in the stability study can be found in Table 5 above. Only two replicates of one sample, 1583805067, failed to yield 90% coverage and 10 CCS reads for 90% reproducibility. Additionally, no drop in sample-specific read counts was observed at the last stability time point (PBT5080) for three separate sequencing runs (PBT5073, PBT5075, and PBT5080) (Figure 15).

[0146] Reprocessing the VOC resulted in identical strain results for 9 / 12 samples. One sample yielded AY.3 when the original result was 1.617.2. Both AY.3 and 1.617.2 contained deltaVOC, and the original week 24 result is now classified as a separate substrain of deltaVOC. Further investigation revealed 38 CLC variants to be identical between the two results, and strain typing was attributed to differences in the Pangolin strain caller version. One sample was concordant but did not have sufficient read depth to report a result. The only true discordant was B.1.1.7, as originally reported, and B.1.621.1 in the repeat. It is believed that the discrepancy may have been caused by sample exchange, as both the original and stability sequencing resulted in sufficient read coverage with minimal sharing of called variants between runs. Overall accuracy was 91.6% and discrepant results were not believed to be due to stability issues.

[0147] Example 3 Cherry Picking A Hamilton MicroLab STAR liquid handler was used to transfer samples from a source plate containing both positive and negative patient samples to a condensed PCR plate containing only the positive samples for sequencing. Informally, this process is referred to as "cherry picking." Samples were total nucleic acid extracted from positive samples with a CT < 31.

[0148] Example 4 Analysis of sequencing data for SARS-CoV-2 strain determination Upstream analysis included monitoring sequencer runs for completion, demultiplexing to generate individual sample FASTQ files, and triggering each alignment with the SARS-CoV-2 reference genome to generate alignments and variant calls. Downstream analysis for samples in each SMRTCell included generating all results, including phylogenetic classification, for each sample.

[0149] Upstream analysis An example flow chart for upstream analysis is shown in Figure 4. To initiate the analysis, PacBio / Molecular Loop raw data was deposited from the sequencer into the AWS drop directory. The script detected when a run complete file was created and copied the data into a folder ready for demultiplexing. Samples that failed on the sequencer did not generate a data file. These samples were designated to be repeated and were not used for sequence analysis.

[0150] At this point, demultiplexing and generation of individual sample FASTQ files was accomplished using the following steps: (1) generating circular consensus sequence (CCS) BAM files using PacBio's SMRTLINK CCS program; (2) merging the intermediate BAM files using samtool; (3) demultiplexing using PacBio lima program to generate individual BAM files corresponding to the different barcode combinations in the run manifest; (4) combining the demultiplexed output data by sample name and / or patient identifier; (5) removing barcodes from sequences and generating individual sample FASTQ files; (6) aligning sequences with barcodes and trimming barcodes (e.g., using the PacBio trim script); (7) converting BAM files to FASTQ files (e.g., using bamtool); (8) copying the FASTQ and CCS BAM files to the final location; and (9) copying the FASTQ files and corresponding run manifest to the drop location to start the CLC workflow.

[0151] The CLC analysis workflow was performed using the following steps: (1) the NGS data analysis workflow was run for each sample using the current validated version of the CLC Genomics Server; (2) for each sample's FASTQ file, (a) reads were filtered to retain reads between 250 and 5000 bp in length; (b) reads were aligned to the SARS-CoV-2 reference genome ("NC_045512v2") using minimap2 to generate BAM files; (c) local realignment was performed and variant calling was performed using the Low Frequency Variant Detection tool in the CLC Genomics Server; and (d) both the assembly (BAM files) and the detected variants (cf) were input into downstream post-processing analysis. The script detected the CLC process completion, which triggered the initiation of downstream analysis of the sample in each SMRTcell.

[0152] Downstream analysis An example of a flow chart for downstream (post-processing) analysis is shown in Figure 5. Post-processing part 1 is represented by the first block in Figure 5. The steps of post-processing part 1 were as follows: (1) By using the appropriate reference file, VCFCons was used to generate a consensus sequence based on sequence alignment and variant calling for each sample. This analysis required a minimum coverage of 4CCS reads to assign a base to each genomic position, and a minimum alternative frequency of 0.5, and positions not meeting this criteria were assigned an ambiguous base "N". (2) Seqtk was used to generate sequence base composition, which was later used to determine the percentage of non-ambiguous bases. (3) Nextclade was used to generate (a) clade assignment, (b) variant calling, and (c) sample sequence quality check using the consensus sequence as input data. (4) Pangolin was then used to assign phylogeny to the consensus sequence by generating a SARS-CoV-2 phylogeny (known as Pango nomenclature) and then assigning a SARS-CoV-2 genome sequence phylogeny (Pango phylogeny). Pangolin only considers genomes with at least 50% unambiguous bases. (5) SummaryStat was used to compile the results from Nextclade, Pangolin and Seqtk and generate coverage statistics required for subsequent QC, including the mean median amplicon coverage and percent genome coverage. For this analysis, the median coverage of bases in 29 overlapping 1.2 kb regions spanning the entire SARS-CoV-2 genome was calculated for each of the samples. Distribution statistics of these coverage values ​​(minimum, first quartile, mean, median, third quartile and maximum) were calculated for each sample. Percent genome coverage was also calculated as the number of unambiguous bases (A, T, C, G) divided by total sequence length, and phylogenetic classification was aggregated; only samples that yielded Nextclade results and Pangolin phylogenetic calls were retained for further processing.

[0153] At this point, the second part of post-processing was started as shown in the "Combine Patient Metadata", "Quality Check" and "Generate Final Report" blocks in Figure 5. Therefore, QC was performed again using the appropriate reference files, strain surveillance specific metadata (demographic data, percent genome coverage, and Ct values ​​from RT-PCR assays) and the data was added to the results. Samples lacking metadata were excluded from the result set. Also, no-template QC was performed based on no-template controls. If the average median coverage of the 29 genomic regions was >10 CCS reads, all sequenced samples on the same plate were removed. Finally, coverage QC was performed. Samples with genome coverage >=90% were kept in the results, and samples with average median coverage >10 CCS reads were kept in the results. The results were then transferred to the Report System location to generate a patient report with the corresponding Pangolin lineage. Samples that could not produce a result were reported as lineage could not be determined. Even when the SARS-CoV-2 virus is detected, phylogenetic information may not be reported.

[0154] Lineage calling criteria were as follows: Inclusion criteria: (1) CT<31; (2) matching metadata (strain surveillance); (3) >90% genome coverage; (4) mean median coverage >10 CCS reads; (4) passed NTC controls; and (5) Nextclade results and Pangolin lineage calls. Exclusion criteria: (1) CT>31; (2) missing metadata (strain surveillance); (3) <90% genome coverage; (4) mean median coverage <10 CCS reads; and (4) failed NTC controls.

[0155] Example 5 Assessment of possible new variants and classification Revalidation of the classification accuracy of the Virseq assay in response to the emergence of new variants (i.e., lineages) of the SARS-CoV-2 virus and the concomitant changes to the pangolin classification software was performed as outlined in Figure 6. The pangolin software is distributed by Dockerhub (at hub.docker.com / r / staphb / pangolin). We monitored the Pangolin site and checked for updates by downloading and installing an updated docker container at regular intervals (e.g., once a week). If there were no updates, we determined that no action needed to be taken. The documentation docker container was updated with a changelog along with the release notes. The updated docker file contained the changelog as well as the latest versions of pangolin, pangoLEARN, pango-designation, scorpio, and constellation (see github.com / cov-lineages).

[0156] In the case of updates, regression analysis was performed using in-house laboratory data. Essentially, the steps were performed as follows: The new pangolin version was used to determine the lineage of samples contained within the reference set of historical Virseq sequences. The reference set included the initial SMRT cell from October 2021, which is composed primarily of the Delta lineage. It also included two updates of the Omicron lineage, made in December 2021 and March 2022. Each sample in the reference set included its consensus sequence, as well as its lineage classification history made by previous pangolin versions. The reference set was updated periodically to include samples representing newer, more prevalent lineages as pangolin versions were updated.

[0157] The format of the pangolin software output was then compared to that of the previous version to determine whether there had been any changes to the pangolin output format. If there were any changes to the CSV output data (i.e., additional columns, changes to column names), these were recorded and the laboratory Virseq pipeline was modified as necessary to accommodate the changes. The modified version was then deployed to the QA environment for testing.

[0158] Any changes in lineage calls were then assessed and compared to those expected from the software update change records. Expected changes typically include reassignments between sublineages. If there were any unexpected changes in lineage (e.g., from Delta sublineage to Alpha), these were investigated in detail and recorded.

[0159] The criteria and actions taken were as follows: Differences in lineage classification were primarily due to improvements in the pangoLEARN / pango-designation definition of variants in the newer version. Most of these were sublineage reassignments, but could also be due to changes in variant definitions by the model. Sublineage reassignments were reconsidered to ensure that they were expected changes under the parent lineage, such as AY-to-AY reassignments in the parent delta lineage. There may be other sources of discrepancies in samples with <90% genome coverage. Any discrepancies that could be explained by sublineage reassignments or genome coverage issues as described above were noted and further reviewed for approval. GISAID regression tests were then performed. If the discrepancies could not be explained as described above and no new pangolin lineages were added in the upgrade, the upgrade was stopped and production continued with the current version of pangolin. Discrepancies associated with the update as described above were noted and stored. Discrepancies were further investigated as new information became available and was noted, or the initiation of this protocol for the next release of pangolin could resolve the discrepancies.

[0160] At this point, a second regression test was performed using the published (GISAID) sequences and their metadata. The latest GISAID sequences were downloaded, metadata for all GISAID sequences and pangolin lineages were obtained, and the list of VOCs and VOIs was updated based on the WHO updates and the latest complete list of lineages. A data simulator was then used to model the coverage and error characteristics of the Virseq assay. The simulator used the GISAID sequences as a starting point and imposes simulated coverage and error based on empirical coverage profiles and maximum-minor allele frequencies from a collection of Virseq samples. The resulting simulated samples were run through pangolin and the lineage classification was compared to that of the original GISAID sequences. Classification stability was defined as the rate at which mutated sequences maintain their expected lineage classification. In this regression test, two experiments were performed to assess the stability of the simulated classification. First, up to 100 GISAID sequences were randomly sampled for each VOC / VOI to assess the taxonomic stability of these important lineages, regardless of their frequency in the available sequencing data. This allowed for the assessment of the taxonomic stability of emerging variants and new sublineages of existing ones. Second, 10,000 GISAID sequences were randomly sampled from the database for a frequency-based retrospective analysis of lineage taxonomic stability. This allowed for the quantification of stability relative to historical prevalence.

[0161] At this point, the output data of the data simulator experiment was reviewed and checked for unexpected changes in classification stability with respect to the previous regression tests using GISAID data for all known VOC / VOIs, as well as retrospective GISAID data. Any unexpected instabilities were investigated and recorded, and upgrades were accepted if certain parameters were met. In some cases, upgrades were requested if the median VOC / VOI agreement between the simulated data and the reference sequences was at least 90%. If these criteria were not met, further investigation was indicated.

[0162] If the new mismatched lineages were new, the new lineages were examined to determine whether they could be detected using the methods disclosed herein. If the mismatched variants were not new variants, they were investigated to find the root cause of the mismatch. This included examining the coverage of the reference sequence as well as the simulated sequence to ensure that there was no undesired drop in base coverage in certain regions. Also, the simulation was rerun with a different seed to determine whether the mismatches were reproduced. If so, the upgrade was aborted.

[0163] At this point, the new variants were assessed using the methods disclosed herein. For successful surveillance of emerging variants (lineages), the potential impact on molecular loop inversion probe amplification was reviewed by performing in silico analysis, such as by identifying the location of each sequence variant and associated molecular loop probe in the emerging lineage, to assess possible interference with probe binding. For example, a very conservative estimate will be used that any new sequence variants overlapping with any probe will affect hybridization. Also, all adjacent probes in the region were reviewed to ensure coverage of the new sequence variant. For any sequence variants that may result in reduced coverage in a particular region, the affected probes in the pangolin lineage update validation summary were recorded.

[0164] Example 6 Embodiment The present disclosure can be better understood with reference to the following non-limiting embodiments. A1. A method for identifying and / or tracking variants of SARS-CoV-2, comprising: (a) identifying a sample from the subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence. The method includes: A2. The method of any one of the preceding or following method embodiments, wherein the step of generating sample-specific SARS-CoV-2 nucleic acid comprises generating sample-specific SARS-CoV-2 cDNA using reverse transcriptase polymerase chain reaction (RT-PCR). A3. The method of any one of the preceding or following method embodiments, wherein the SARS-CoV-2 cDNA is then further amplified using tiled primers that bind at intervals along the viral genome. A4. The method of any one of the preceding or following method embodiments, wherein the tiled primers are spaced such that adjacent primers are approximately 600 bp apart from each other. A4.1 The method of any one of the preceding or following method embodiments, further comprising hybridizing one strand of the sample SARS-CoV-2 cDNA to a single-stranded probe DNA template comprising a pair of SARS-CoV-2 probes, wherein a first probe is located at the 3' end of the probe DNA template to function as a forward primer and a second probe is located at the 5' end of the probe DNA template to function as a reverse primer. A4.2 The method of any one of the preceding or following method embodiments, wherein the SARS-CoV-2 genome is amplified highly efficiently regardless of the presence or absence of new variants. A4.3 The method of any one of the preceding or following method embodiments, wherein the tiling primer is a primer further comprising an adapter for addition of a barcode sequence used to correlate the SARS-CoV-2 sample specific nucleic acid to a sample number and / or to a universal primer site for nucleic acid sequencing. A5. The method of any one of the preceding or following method embodiments, wherein the single stranded probe DNA template further comprises a universal sequencing primer located internal to the probe sequence. A6. The method of any one of the preceding or following method embodiments, wherein the single stranded probe DNA template further comprises an adapter sequence for addition of a barcode sequence used to correlate the SARS-CoV-2 sample-specific nucleic acid with a sample number. A7. The method of any one of the preceding or following method embodiments, further comprising filling in the sequence between the two probes to generate a circular single-stranded probe DNA template that contains a sequence specific to the sample SARS-CoV-2 cDNA between the two probe sequences. A8. The method of any one of the preceding or following method embodiments, further comprising the step of liberating a circular single-stranded probe DNA template comprising a sequence specific for the sample SARS-CoV-2 cDNA from the sample-specific SARS-CoV-2 DNA. A9. The method of any one of the preceding or following method embodiments, further comprising digestion of the circular single-stranded probe DNA template comprising a sequence specific for the sample SARS-CoV-2 cDNA to generate linear DNA that is used as a template for performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid. A10. The method of any one of the preceding or following method embodiments, further comprising uploading the results of the step of determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence to a depository for further classification if a variant is detected. A11. The method of any one of the preceding or following method embodiments, wherein the depository is a CDC database. A12. The method of any one of the preceding or following method embodiments, wherein the nucleic acid sequencing comprises sequencing at least 80%, or optionally 85%, or optionally 90% of the entire viral genome. A13. The method of any one of the preceding or following method embodiments, further comprising identifying the geographic location of the subject. A14. The method of any one of the preceding or following method embodiments, wherein nucleic acid sequencing comprises whole genome sequencing. A15. The method of any one of the preceding or following method embodiments, wherein the step of determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence comprises aligning the sample SAR-CoV-2 sequence to a SARS-CoV-2 reference genome to generate a sample-specific assembly and consensus sequence. A15.1 The method of any one of the preceding or following method embodiments, wherein sample SAR-CoV-2 nucleic acid sequences having a minimum coverage of at least 50% are used as input data for calling variants and / or for generating a sample-specific genome assembly, to generate a consensus sequence for each sample. A15.2 The method of any one of the preceding or following method embodiments, wherein there is a defined threshold for generating a consensus sequence. A15.3 The method of any one of the preceding or following method embodiments, wherein the defined threshold comprises at least 4 circular consensus sequencing (CCS) reads covering each individual base pair, and / or an alternative allele frequency compared to the reference of >50%. A15.4 The method of any one of the preceding or following method embodiments, further comprising evaluation of an external no template control (NTC) and / or an external positive template control (PTC) to assess the validity of the results. A15.5 The method of any one of the preceding or following method embodiments, wherein the sample SAR-CoV-2 nucleic acid sequencing reads are filtered to retain reads between 250 and 5000 bp in length. A15.6 The method of any one of the preceding or following method embodiments, wherein the sample SAR-CoV-2 nucleic acid sequencing reads are aligned to the SARS-CoV-2 reference genome (NC_045512v2). A15.7 The method of any one of the preceding or following method embodiments, wherein the sample SAR-CoV-2 nucleic acid sequencing reads are aligned to the SARS-CoV-2 reference genome, followed by local realignment and variant calling. A15.8 The method of any one of the preceding or following method embodiments, wherein the determination of the sample SAR-CoV-2 nucleic acid sequence base composition is produced to determine the percentage of non-ambiguous bases. A15.9 The method of any one of the preceding or following method embodiments, wherein the consensus sequence can be used as input data to optionally generate any one or all of: (a) clade assignment; (b) mutation determination; and (c) sample sequence quality checks. A16. The method of any one of the preceding or following method embodiments, wherein step (d) of determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence further comprises assessing the lineage for the sample. A16.1 The method of any one of the preceding or following method embodiments, wherein the phylogeny is assigned to a consensus sequence by generating SARS-CoV-2 phylogenies and then assigning the SARS-CoV-2 genomic sequence phylogeny. A16.2 The method of any one of the preceding or following method embodiments, wherein phylogenetic assignment is set to consider only genomes with at least 50% unambiguous bases. A16.3 The method of any one of the preceding or following method embodiments, wherein strain phylogenetic results are released for samples with 90% genome coverage. A16.4 The method of any one of the preceding or following method embodiments, wherein strain phylogenetic results are released for samples with an average median read coverage across the genome that is >10 circular consensus sequence (CCS) reads. A16.5 The method of any one of the preceding or following method embodiments, wherein the different CCS read metrics are based on the nucleotide level (4 CCS reads) and on the genome level (10 CCS reads). A16.6 The method of any one of the preceding or following method embodiments, wherein Pangolin is used to assign lineage. A16.7 The method of any one of the preceding or following method embodiments, further comprising the step of generating coverage statistics. A16.8 The method of any one of the preceding or following method embodiments, wherein the coverage statistics data is generated using SummaryStat. A16.9 The method of any one of the preceding or following method embodiments, wherein the median coverage of bases in 29 overlapping 1.2 kb regions spanning the entire SARS-CoV-2 genome is calculated for each of the samples. A16.10 The method of any one of the preceding or following method embodiments, wherein the average median coverage of 29 genomic regions is >10 CCS reads. A16.11 The method of any one of the preceding or following method embodiments, wherein samples with genome coverage >= 90% are retained in the outcome. A16.12 The method of any one of the preceding or following method embodiments, wherein samples with a median coverage average >10 CCS reads are retained in the result. A16.13 The method of any one of the preceding or following method embodiments, wherein QC is performed and the data added to the results using demographic data, percent genome coverage, and Ct values ​​from RT-PCR assays. A16.14 The method of any one of the preceding or following method embodiments, wherein the results are used to generate a patient report with corresponding lineage and / or geographic assignment. A16.15 The method of any one of the preceding or following method embodiments, wherein inclusion criteria include: (1) CT<31; (2) corresponding metadata (strain surveillance); (3) >90% genome coverage; (4) mean median coverage >10 CCS reads; (4) passing NTC controls; and (5) lineage calls. A16.16 The method of any one of the preceding or following method embodiments, wherein the exclusion criteria include: (1) CT>31; (2) missing metadata (strain surveillance); (3) <90% genome coverage; (4) mean median coverage <10 CCS reads; and (4) failed NTC controls. A17. The method of any one of the preceding or following method embodiments, further comprising the step of re-validating the lineage assignment by determining whether a depository update has occurred. A17.1 The method of any one of the preceding or following method embodiments, wherein the revalidating step is performed prior to the step of determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence. A17.2 The method of any one of the preceding or following method embodiments, wherein the revalidation includes a regression analysis using in-house data to determine whether to change the previously assigned lineage. A17.3 The method of any one of the preceding or following method embodiments, wherein the in-house data comprises a dataset defined by at least one of the following: date of emergence, SARS-CoV-2 lineage, geographic origin of the sample, history of phylogenetic typing, or updates to the algorithm used for phylogenetic typing. A17.4 The method of any one of the preceding or following method embodiments, wherein the update includes a lineage or sub-lineage change for in-house data. A17.5 The method of any one of the preceding or following method embodiments, wherein the revalidation comprises regression analysis using data from the depository. A17.6 The method of any one of the preceding or following method embodiments, wherein the depository is GISAID. A17.7 The method of any one of the preceding or following method embodiments, wherein GISAID sequences are downloaded, metadata and lineages for all GISAID sequences are obtained, and the list of VOCs and VOIs is updated based on WHO updates and the latest complete list of lineages. A17.8 The method of any one of the preceding or following method embodiments, further comprising the step of modeling the coverage and error characteristics of the in-house assay using a data simulator. A17.9 The method of any one of the preceding or following method embodiments, wherein the simulator uses the GISAID sequence as a starting point and imposes simulated coverage and error based on empirical coverage profiles and maximum-minor allele frequencies from a collection of samples, the resulting simulated samples are run through a phylogenetic algorithm and the phylogenetic classification is compared to that of the original GISAID sequence. A17.10 The method of any one of the preceding or following method embodiments, wherein taxonomic stability is defined as the rate at which variant sequences maintain their expected phylogenetic classification. A17.11 The method of any one of the preceding or following method embodiments, wherein for the simulation, 100 GISAID sequences are randomly sampled for each VOC and / or VOI. A17.12 The method of any one of the preceding or following method embodiments, wherein for the simulation, 10,000 GISAID sequences from the database are randomly sampled for frequency-based retrospective analysis of phylogenetic stability. A17.13 The method of any one of the preceding or following method embodiments, wherein upgrading is requested if the median VOC / VOI identity between the simulated data and the reference sequence is at least 90%. A18. The method of any one of the preceding or following method embodiments, wherein at least a portion of the steps are controlled by a computer-program product tangibly embodied in a computer and / or a non-transitory machine-readable storage medium. A18.1 At least some of the steps are one or more data processors; a non-transitory computer-readable storage medium containing instructions that, when executed on one or more data processors, cause the one or more data processors to perform a process that includes any of the method steps; The method of any one of the preceding or following method embodiments, wherein B1. A system including at least one station or component for performing any of the method embodiments described above or below. B2.(a) identifying a sample from the subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence. 23. A system comprising at least one station or component for performing any of the preceding or following method embodiments, comprising: B3. The system of any one of the preceding or following embodiments, wherein at least a portion of the steps are controlled by a computer-program product tangibly embodied in a computer and / or a non-transitory machine-readable storage medium. B.4 At least some of the steps: one or more data processors; a non-transitory computer-readable storage medium containing instructions that, when executed on one or more data processors, cause the one or more data processors to perform a process that includes any of the method steps; 4. The system of any one of the preceding or following method embodiments, wherein the system is controlled by C1. A computer program product tangibly embodied in a non-transitory machine-readable storage medium that, when executed on one or more data processors: (a) identifying a sample from the subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence; A computer program product for causing one or more data processors to perform processes including: D1. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising: (a) identifying a sample from the subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies to SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) determining whether the nucleic acid sequence contains a SARS-CoV-2 variant sequence. A computer program product comprising instructions configured to execute at least one station or component of the system to perform any of the above.

Claims

1. A method for identifying and / or tracking variants of SARS-CoV-2, comprising: (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or for antibodies against SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 cDNA from said sample; (c) performing nucleic acid sequencing on said sample-specific SARS-CoV-2 nucleic acid; and determining whether the sequence of said nucleic acid comprises a SARS-CoV-2 variant sequence A method comprising the steps of:

2. The method of claim 1, wherein said SARS-CoV-2 cDNA is further amplified using tile-type primers that hybridize at intervals along the viral genome.

3. The method of claim 2, wherein said tile-type primers are spaced such that adjacent primers are approximately 600 bp apart from each other.

4. The method according to any one of claims 1 to 3, wherein the step of generating sample-specific SARS-CoV-2 nucleic acid further comprises hybridizing one strand of said sample SARS-CoV-2 cDNA with a single-stranded probe DNA template comprising a pair of SARS-CoV-2 probes, a first probe being located at the 3' end of said probe DNA template so as to function as a forward primer, and a second probe being located at the 5' end of said probe DNA template so as to function as a reverse primer.

5. The method of claim 4, wherein said single-stranded probe DNA template further comprises a universal sequencing primer at a position adjacent to the sequence of said probe.

6. The method of claim 4, wherein said single-stranded probe DNA template further comprises an adapter sequence for the addition of a barcode sequence used to correlate said SARS-CoV-2 sample-specific nucleic acid with a sample number.

7. The method of claim 6, wherein said barcode is associated with a postal code or other geographical identifier for said sample.

8. The method of claim 4, further comprising the step of filling in the sequence between said two probes to generate a circular single-stranded probe DNA template comprising a sequence specific for said sample SARS-CoV-2 cDNA between the sequences of said two probes.

9. The method according to claim 1, wherein the nucleic acid sequencing comprises sequencing at least 90% of the entire viral genome.

10. The method according to claim 1, wherein the median base coverage in 29 overlapping 1.2 kb regions spanning the entire SARS-CoV-2 genome is calculated for each of the samples.

11. The method according to claim 1, further comprising uploading the results of step (d) to a repository for further classification if a variant is detected.

12. The method according to claim 11, wherein the repository is the CDC database.

13. The method according to claim 1, further comprising identifying the geographical location of the subject.

14. The method according to claim 1, wherein determining whether the nucleic acid sequence comprises a SARS-CoV-2 variant sequence comprises aligning the SARS-CoV-2 sequence of the sample to a SARS-CoV-2 reference genome to generate a sample-specific assembly and consensus sequence.

15. The method according to claim 1, wherein step (d) further comprises assessing the lineage for the sample.

16. The method according to claim 1, wherein the incorporation criteria for step (d) comprise > 90% genome coverage, average median coverage > 10 CCS reads; and lineage determination for the sample.

17. The method according to claim 11, further comprising determining whether the repository has been updated prior to determining whether the nucleic acid sequence comprises a SARS-CoV-2 variant sequence.

18. At least some of the steps are one or more data processors, and a non-transitory computer-readable storage medium containing instructions that cause the one or more data processors to perform a process including any of the steps of the method when executed by the one or more data processors and is controlled thereby. The method according to claim 1.

19. one or more data processors, and when executed by the one or more data processors, (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or antibodies against SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) a non-transitory computer-readable storage medium containing instructions for causing the one or more data processors to perform a process including determining whether the sequence of the nucleic acid contains a SARS-CoV-2 variant sequence A system comprising.

20. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, (a) identifying a sample from a subject as positive for SARS-CoV-2 nucleic acid and / or antibodies against SARS-CoV-2; (b) generating sample-specific SARS-CoV-2 nucleic acid from the sample; (c) performing nucleic acid sequencing on the sample-specific SARS-CoV-2 nucleic acid; and (d) A computer-program product comprising instructions configured to cause one or more data processors to perform a process including determining whether the sequence of the nucleic acid contains a SARS-CoV-2 variant sequence.