A dynamic clinical assay pipeline for detecting viruses
A dynamic clinical assay pipeline using molecular inversion probes and circular consensus sequencing addresses the challenges of virus detection and lineage assignment, achieving high accuracy and efficiency by ensuring sufficient sequencing depth and coverage.
Patent Information
- Application Number
- JP2025525221
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-19
- Filing Date
- 2023-11-01
- Publication Date
- 2025-12-09
AI Technical Summary
Existing methods for virus detection and lineage assignment are hindered by the lack of specific and sensitive diagnostic tests, challenges in distinguishing between similar viral strains, and the impact of rapid viral mutations, leading to inaccurate sequencing and false-negative results due to repetitive DNA sequences and insufficient sequencing depth.
A dynamic clinical assay pipeline using molecular inversion probes, PCR amplification, and circular consensus sequencing to capture and analyze viral nucleic acids, ensuring sufficient sequencing depth and coverage, and generating a report file for accurate virus detection and lineage assignment.
The method achieves 98.84% accuracy in virus detection and correct lineage assignment, improving sequencing efficiency, reducing costs, and allowing for dynamic adaptation to different viral lineages.
Smart Images

Figure 2025539721000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to clinical trials for viruses, and in particular to a dynamic clinical assay pipeline for detecting and / or assigning lineage to viruses (e.g., orthopoxviruses such as monkeypox virus) in samples. [Background technology]
[0002] Viruses are microscopic infectious agents that can replicate inside the living cells of organisms. While some viruses are harmless to humans, a wide variety of viruses can cause a wide range of illnesses and diseases, including life-threatening ones, in animals, plants, microorganisms, and humans. Understanding the properties of viruses, particularly their genetic material, and detecting their presence is essential for developing effective treatments and clinical protocols, as well as for preventing the spread of viral infections.
[0003] Traditional methods for virus detection involve diagnostic testing. However, virus detection can be hindered by a lack of specific and sensitive diagnostic tests. Many viral infections share similar symptoms, making it difficult to distinguish between different viruses based on symptoms alone. Therefore, highly specific and sensitive diagnostic tests are essential for accurate virus detection, but developing such tests can be a time-consuming and complex process.
[0004] Virus replication relies on its genetic material. With the development of testing, computing, and sequencing technologies, machines such as sequencers and associated software have been invented and manufactured to identify the presence of viruses in biological samples, such as blood, saliva, and tissue, based on the genetic material in those samples. Various methods, such as polymerase chain reaction (PCR) and next-generation sequencing (NGS), have also been developed to facilitate virus detection. However, detecting viral nucleic acids in samples can remain a challenge, in part due to rapid viral mutation. Viral mutations can create novel strains that trigger different immune responses in the host organism, potentially necessitating different treatments and protocols to prevent viral spread and infection. However, reliable methods and systems for detecting viruses in samples and assigning their lineages to those samples are lacking. Summary of the Invention
[0005] In various embodiments, the method includes obtaining nucleic acid from a sample obtained from the subject; capturing a target molecule in the nucleic acid using a molecular inversion probe under hybridization conditions; amplifying the target molecule using polymerase chain reaction (PCR) to obtain a plurality of amplified molecules; for each molecule of the plurality of amplified molecules, ligating an adapter to each end of the molecule to create a circular molecule; sequencing the circular molecules to obtain sequence reads; generating, using a computing system, a sequencing file including a sequence read for each molecule of the plurality of amplified molecules and a position of each sequence read in a reference genome of the virus by aligning the sequence reads to the reference genome of the virus; and generating a report file for the subject using the computing system and the sequencing file. The report file includes a predicted lineage of the virus in the sample, and the generating step includes generating a consensus sequence of the target molecule based on sequence reads of each of the plurality of amplified molecules, wherein if at least a predetermined number of sequence reads have nucleotide identity at a position in the consensus sequence, a nucleotide identity is assigned to the position, and if less than a predetermined number of sequence reads have nucleotide identity at a position in the consensus sequence, an "N" is assigned to the position; determining one or more scores of the consensus sequence based on a reference genome of the virus or a library of the virus, wherein the one or more scores are determined based on the distribution of mutations in the virus; and determining a predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence.
[0006] In some embodiments, the report file further includes the presence or absence of virus in the sample.
[0007] In some embodiments, the presence or absence of virus in a sample is determined using real-time PCR (RT-PCR) or based on one or more scores of consensus sequences.
[0008] In some embodiments, the subject tests positive for the virus and the sample contains viral nucleic acid.
[0009] In some embodiments, the virus is monkeypox (MPX) virus.
[0010] In some embodiments, the target molecule is a double-stranded DNA molecule.
[0011] In some embodiments, the molecular inversion probe consists of two binding sites separated by approximately 600-700 bp.
[0012] In some embodiments, the sequencing file is a binary alignment map (BAM) file.
[0013] In some embodiments, the report file is a variant call format (VCF) file.
[0014] In some embodiments, the method further comprises providing the subject with a treatment plan or clinical trial protocol.
[0015] In some embodiments, the treatment regimen includes administering an antiviral drug to the subject.
[0016] In some embodiments, the method further includes the steps of obtaining a prevalent lineage of a virus in a subject population determined based on a predicted lineage of the virus in the sample; updating a molecular inversion probe to capture a target molecule specific to the prevalent lineage; updating an adapter based on the updated molecular inversion probe; and obtaining a decision rule set specific to determining the prevalent lineage, wherein determining one or more scores of the consensus sequence is performed using the decision rule set.
[0017] In various embodiments, the method includes obtaining nucleic acid from a sample obtained from the subject; capturing a plurality of target molecules in the nucleic acid using a plurality of molecular inversion probes under hybridization conditions; amplifying each target molecule using polymerase chain reaction (PCR) to obtain a plurality of sets of amplified molecules, each corresponding to a target molecule; for each molecule in each set of amplified molecules, ligating an adapter to each end of the molecule to create a circular molecule; sequencing the circular molecules to obtain sequence reads; generating, using a computing system, a sequencing file including a sequence read for each molecule in each set of amplified molecules and a position of each sequence read in a reference genome of the virus by aligning the sequence reads to the reference genome of the virus; and generating, using the computing system and the sequencing file, a report file about the subject, wherein the report file includes a predicted lineage of the virus in the sample, the generating step generating a consensus sequence for each target molecule based on the sequence read for each molecule in the set of amplified molecules corresponding to the target molecule, wherein at least a nucleotide identity is assigned to a position if at least a first predetermined number of sequence reads have nucleotide identity to the position in the consensus sequence, and an "N" is assigned to the position if fewer than the first predetermined number of sequence reads have nucleotide identity to the position in the consensus sequence; generating a genomic construct of nucleic acids in the sample based on the consensus sequence of the target molecule, wherein a nucleotide identity is assigned to the position if at least a second predetermined number of consensus sequences have nucleotide identity to the position in the genomic construct, and a nucleotide identity is assigned to the position if the nucleotide identity at the position is present in at least 50% of sequence reads aligned to that position, and an "N" is assigned to the position if fewer than the second predetermined number of consensus sequences have nucleotide identity to the position in the genomic construct; determining one or more scores of the consensus sequences based on a viral reference genome or a viral library; and determining a predicted lineage of the virus in the sample based on the one or more scores of the consensus sequences.
[0018] In some embodiments, the report file further includes the presence or absence of virus in the sample.
[0019] In some embodiments, the presence or absence of virus in a sample is determined using real-time PCR (RT-PCR) or based on one or more scores of consensus sequences.
[0020] In some embodiments, the subject tests positive for the virus and the sample contains viral nucleic acid.
[0021] In some embodiments, the virus is monkeypox (MPX) virus.
[0022] In some embodiments, the target molecule is a double-stranded DNA molecule.
[0023] In some embodiments, the molecular inversion probe consists of two binding sites separated by approximately 600-700 bp.
[0024] In some embodiments, the sequencing file is a binary alignment map (BAM) file.
[0025] In some embodiments, the report file is a variant call format (VCF) file.
[0026] In some embodiments, the method further comprises providing the subject with a treatment plan or clinical trial protocol.
[0027] In some embodiments, the treatment regimen includes administering an antiviral drug to the subject.
[0028] In some embodiments, the method further includes the steps of obtaining a prevalent lineage of a virus in a subject population determined based on a predicted lineage of the virus in the sample; updating a plurality of molecular inversion probes to capture mutations specific to the prevalent lineage; updating adapters based on the updated plurality of molecular inversion probes; and obtaining a decision rule set specific to determining the prevalent lineage, wherein determining one or more scores of the consensus sequence is performed using the decision rule set.
[0029] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods or processes disclosed herein.
[0030] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0031] The terms and expressions used are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the illustrated and described features or portions thereof, but it is recognized that various modifications are possible within the scope of the claimed technology. Thus, although the present technology has been specifically disclosed by embodiments and optional features, it should be understood that those skilled in the art may make modifications and variations of the concepts disclosed herein, and that such modifications and variations are deemed to be within the scope of the present technology as defined by the appended claims. [Brief explanation of the drawings]
[0032] The invention will be better understood with reference to the following non-limiting drawings.
[0033] [Figure 1] FIG. 1 is a block diagram of a system capable of implementing virus detection and lineage assignment techniques according to various embodiments.
[0034] [Figure 2] 1 is an exemplary flowchart illustrating a process for generating a report file regarding the presence or absence of a virus in a sample according to various embodiments.
[0035] [Figure 3] 1 is an exemplary flowchart illustrating a process for implementing the disclosed techniques for determining the presence or absence of a virus in a sample according to various embodiments.
[0036] [Figure 4] 1 is a portion of a molecular loop probe set design and PCR pipeline for capturing target DNA molecules according to various embodiments.
[0037] [Figure 5] 1 is a portion of an adapter design, ligation, and sequencing pipeline for generating circular DNA molecules according to various embodiments.
[0038] [Figure 6] 1 is a summary of the complexity of the MPX genome according to various embodiments.
[0039] [Figure 7] 1 is an Integrative Genomics Viewer (IGV) plot showing genome coverage according to various embodiments.
[0040] [Figure 8] 8A-8C are high-resolution graphical representations of the IGV plots shown in FIG. 7 generated using software developed for various clinical samples according to various embodiments.
[0041] [Figure 9]1 is another method of showing genome coverage and read depth for various samples and controls according to various embodiments.
[0042] [Figure 10] 1 shows genome coverage and per-base coverage according to various embodiments.
[0043] [Figure 11] 1 shows exemplary virus calls and strain assignments in samples according to various embodiments.
[0044] [Figure 12A] FIG. 1 is a phylogenetic analysis of an initial pilot run showing predicted assignment of control and circulating strains to patient samples according to various embodiments. [Figure 12B] FIG. 1 is a phylogenetic analysis of an initial pilot run showing predicted assignment of control and circulating strains to patient samples according to various embodiments.
[0045] [Figure 13A] 11 is an analysis of a larger multiplex run including more control and patient samples according to various embodiments. [Figure 13B] 11 is an analysis of a larger multiplex run including more control and patient samples according to various embodiments. [Figure 13C] 11 is an analysis of a larger multiplex run including more control and patient samples according to various embodiments.
[0046] [Figure 14] 1 is a comparison of sequenced synthetic controls with published data according to various embodiments.
[0047] [Figure 15] 1 is a molecular loop sequencing process according to various embodiments.
[0048] [Figure 16] 1 is an example of a computing device suitable for use in systems and methods for determining the presence or absence of a virus in a sample obtained from a subject and determining its predicted lineage, according to various embodiments.
[0049] In the accompanying figures, similar components and / or features may be labeled with the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. When only a first reference label is used herein, the description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION
[0050] The following description is intended to provide preferred embodiments only and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred embodiments provides those skilled in the art with an enabling description for practicing various embodiments. It will be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.
[0051] In the following description, specific details are set forth to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other examples, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0052] It should also be noted that particular embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart or diagram may describe operations as a sequential process, many actions may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function. I. Introduction
[0053] Viruses pose a potentially life-threatening threat, making it necessary to accurately detect viruses in biological samples, distinguish between different viral strains, and assign lineages. Because viral replication relies on their genetic material, such as their nucleic acid, various methods have been developed to detect viruses in samples based on the nucleic acid from the sample. Detecting and confirming viral nucleic acid in biological samples can be challenging. One of the main reasons for this is that viruses mutate rapidly, with an average mutation rate of 10-4 to 10-8 mutations per replication site. These mutations alter the virus's genetic material, modifying surface proteins and other critical components, which can affect the host organism's immune response and the results of diagnostic tests.
[0054] Another challenge associated with virus detection is the presence of long stretches of repetitive DNA or RNA. In virus detection, sequencing is always involved in accurately identifying different virus strains, monitoring primer sites, and assigning lineages. However, without accurate sequencing depth and coverage, it is not only difficult to distinguish repetitive sequences in viral nucleic acids from host nucleic acid samples, but it may also mask the presence of mutations and variants, leading to false-negative results.
[0055] To address these and other challenges, the technology described herein is directed to methods and systems for building a dynamic clinical assay pipeline capable of capturing and analyzing viral nucleic acids to provide accurate virus detection and lineage assignment results. This dynamic clinical assay pipeline can be dynamically modified based on factors such as the virus's prevalence and / or lineage assignment history to provide more efficient and accurate virus detection. Specifically, molecular inversion probes can be designed, manufactured, and used to capture target nucleic acid molecules, and PCR methods can be used to amplify the captured molecules. To ensure sufficient sequencing depth and coverage for accurate virus detection and lineage assignment, the amplified molecules can be ligated to adapters of a selected size to obtain circular molecules. The circular molecules are sequenced using an appropriate sequencing method, such as circular consensus sequencing, and a consensus sequence for each target nucleic acid molecule is determined based on predetermined rules. A sequencing file is generated to store information including sequence reads and their locations in a reference genome. The sequencing file and stored data can then be used to adjust and generate a report file. The report file contains virus detection and / or lineage assignment information determined based on the consensus sequences. The virus detection and / or lineage assignment information can be used to dynamically modify molecular inversion probes and / or adapters to provide a dynamic clinical assay pipeline for efficient and accurate virus detection and lineage assignment. A system using the described technology can achieve 98.84% accuracy in virus detection and correct lineage assignment without removing uncovered regions or masking repeats.
[0056] There are several advantages associated with and achieved by the technology described herein. First, wet-lab techniques, including specifically designed molecular capture and ligation methods, improve the quality of molecules sequenced, while sequencing techniques such as circular consensus sequencing double to ensure uniform read coverage across the entire viral genome. Second, more qualified sequence reads can be obtained and secured per molecule, requiring fewer molecules per sample and allowing more samples to be sequenced simultaneously, thereby increasing sequencing efficiency and reducing sequencing costs. Third, the technology described herein can be implemented as a dynamic clinical assay pipeline for virus detection because it is universally applicable and / or easily modified to detect different lineages and types of viruses. Furthermore, because the entire system is specifically designed and the number of reads is predictable, the sequencing process can be optimized to ensure optimal costs. Finally, systems and methods using the described technology can achieve improved overall accuracy in virus detection and lineage assignment.
[0057] In various embodiments, the method includes obtaining nucleic acid from a sample obtained from the subject; capturing a target molecule in the nucleic acid using a molecular inversion probe under hybridization conditions; amplifying the target molecule using polymerase chain reaction (PCR) to obtain a plurality of amplified molecules; for each molecule of the plurality of amplified molecules, ligating an adapter to each end of the molecule to create a circular molecule; sequencing the circular molecules to obtain sequence reads; generating, using a computing system, a sequencing file including a sequence read for each molecule of the plurality of amplified molecules and a position of each sequence read in a reference genome of the virus by aligning the sequence reads to the reference genome of the virus; and generating a report file for the subject using the computing system and the sequencing file. The report file includes a predicted lineage of the virus in the sample, and the generating step includes generating a consensus sequence of the target molecule based on sequence reads of each of the plurality of amplified molecules, wherein if at least a predetermined number of sequence reads have nucleotide identity at a position in the consensus sequence, a nucleotide identity is assigned to the position, and if less than a predetermined number of sequence reads have nucleotide identity at a position in the consensus sequence, an "N" is assigned to the position; determining one or more scores of the consensus sequence based on a reference genome of the virus or a library of the virus, wherein the one or more scores are determined based on the distribution of mutations in the virus; and determining a predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence. II. Terms and Definitions
[0058] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. Unless otherwise defined, all terms (such as technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Furthermore, terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with the meaning in the context of this specification and the related art, and should not be interpreted in an idealized or overly formal sense unless expressly defined herein. Well-known functions or configurations may not be described in detail for the sake of brevity and / or clarity.
[0059] In this specification, terms such as "first," "second," and the like may be used to describe various elements, components, regions, layers, and / or sections; however, it will be understood that these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are used only to distinguish one element, component, region, layer, or section from another region, layer, or section. Thus, a first element, component, region, layer, or section discussed below could also be referred to as a second element, component, region, layer, or section without departing from the teachings of the present disclosure. The order of operations (or steps) is not limited to the order depicted in the claims or figures, unless otherwise indicated.
[0060] As used herein, the singular forms "a," "one," and "an" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, it will be understood that the terms "comprise" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "from about X to Y" mean "from about X to about Y."
[0061] As used herein, terms such as "home collection" or "self-collection" refer to the use of a kit that may be provided to a subject, which includes a swab, a container with a transport fluid (e.g., a buffer), and a container for returning the self-collected sample to a laboratory for testing.
[0062] As used herein, the terms "automated" and "automatic" mean that an operation can be performed with minimal or no manual input or input. The term "semi-automated" refers to operations that allow for some operator input or initiation, but where calculations, acquisitions, refinements, and other steps are performed electronically, usually by a program, without the need for manual input.
[0063] As used herein, when an action is "based on" something, this means that the action is based, at least in part, on at least a part of that something.
[0064] As used herein, the term "clade" or "lineage" refers to a subset of viral variants or species defined by genetic differences or their specific mutations or combinations of biomarkers.
[0065] As used herein, "CT" or "ct" refers to the cycle threshold, or the total number of cycles required to amplify and detect a nucleic acid target (e.g., a viral nucleic acid) by real-time PCR and / or PCR.
[0066] As used herein, the terms "patient" or "subject" are used broadly to refer to an individual who provides a sample for testing or analysis. An individual "patient" or "subject" from whom a sample is collected, obtained, and / or provided includes any and all warm-blooded mammalian subjects, such as humans and / or animals.
[0067] As used herein, the terms "probe," "probe oligonucleotide," "oligonucleotide," and "probe oligonucleotide sequence" may be used interchangeably. The terms "probe," "probe oligonucleotide," "oligonucleotide," and "probe oligonucleotide sequence" may be used to refer to any molecule or system used to detect a target molecule, and the length of the probe or probe oligonucleotide may vary, for example, from 4 nucleotides to about 200 nucleotides.
[0068] As used herein, the term "programmatic" means performed using a computer program and / or software, processor, or ASIC-directed operations. The term "electronic" and its derivatives refer to automatic or semi-automatic operations performed using devices with electrical circuits and / or modules, rather than mental steps, and typically refers to operations performed by a program.
[0069] As used herein, the term "protocol" refers to an automated electronic algorithm (usually a computer program) that comprises defined rules for mathematical calculation, data interrogation and analysis, and that operates a system to carry out a series of instructions.
[0070] As used herein, repeatability (or intra-assay precision) refers to the closeness of agreement between results from successive measurements of the same analyte under the same measurement conditions. Intra-assay reproducibility is a measure of the variability in the analysis of the same specimen within a single analytical run.
[0071] As used herein, reproducibility (or inter-assay precision) refers to the closeness of agreement between results from successive measurements of the same analyte under the same measurement conditions. Inter-assay repeatability is a measure of the variability when analyzing the same specimen during multiple runs.
[0072] As used herein, real-time PCR or quantitative PCR (qPCR) allows for real-time detection of PCR amplification products. Real-time polymerase chain reaction (PCR) assays use fluorescently labeled probes or intercalating dyes to visualize the PCR reaction and monitor the amount of double-stranded DNA product produced. A fluorescent 5' nuclease assay (i.e., TaqMan® assay) is a real-time PCR assay that uses a fluorescent probe consisting of an oligonucleotide with a reporter dye attached to its 5' end and a quencher dye attached to or near its 3' end. The probe anneals to a specific target sequence located between the forward and reverse primers. During the extension phase of the PCR cycle, the 5' nuclease activity of Taq polymerase degrades the probe, separating the reporter dye from the quencher dye and generating a fluorescent signal. With each cycle, additional reporter dye molecules are cleaved from each probe, and fluorescence intensity is monitored during the PCR reaction. The Taq polymerase used may be inactive at room temperature and may be activated by incubation at 95°C before starting the cycling portion of the assay. This minimizes the production of non-specific amplification products.
[0073] As used herein, the terms "sample," "patient sample," "biological sample," and "specimen" may be used interchangeably. Non-limiting examples of samples that may be used for analysis using the disclosed methods and systems include blood or blood products (e.g., serum, plasma, etc.), urine, nasal swabs, liquid biopsy samples, skin swabs, lesion swabs, or combinations thereof. In some cases, DNA may be extracted from lesion material, such as lesion fluid on a dry swab, lesion fluid swabs in viral transport medium, lesion fluid on a glass slide, eschar, or lesional maxilla. The term "blood" encompasses whole blood, blood products, or any fraction of blood as conventionally defined, such as serum, plasma, buffy coat, etc. Suitable samples include those that can be deposited on a substrate for collection and drying, such as, but not limited to, blood, plasma, serum, urine, saliva, tears, cerebrospinal fluid, organ, hair, muscle, pus, other tissue samples, or other liquid aspirates. The term "biological sample" further refers to a sample obtained from a living source, including, but not limited to, animals, cell cultures, organ cultures, tissues, and the like.
[0074] As used herein, the terms "substantially," "approximately," and "about" are defined as referring to most, but not necessarily all, of what is specified (including everything specified), as understood by one of ordinary skill in the art. In any disclosed embodiment, the terms "substantially," "approximately," or "about" may be substituted for "within a percentage" of what is specified, including 0.1, 1, 5, 10, 15, and 20%.
[0075] As used herein, the terms "swab" and "dry swab" refer to a physical vector for containing and / or obtaining a sample from an individual. The vector may be constructed of a variety of materials and may include cotton balls, tissue, plastic forceps, plastic inoculation loops, popsicle sticks, dry or wet swabs, and the like. III. Virus detection and linear allocation techniques
[0076] One or more embodiments described herein may be implemented using an operable, executable, or programmatic system, subsystem, module, block, or component. An operable or executable system, subsystem, module, block, or component may include in vitro or in silico operations or actions that can be performed by an operator (e.g., a researcher or practitioner), a machine (e.g., a computing device or sequencer), or a combination thereof. A programmatic system, subsystem, module, block, or component may include a program, subroutine, portion of a program, or software or hardware component that can perform one or more described tasks or functions. As used herein, a system, subsystem, module, block, or component may exist on a hardware component independent of other systems, subsystems, modules, blocks, or components. Alternatively, a system, subsystem, module, block, or component may be a shared element or process of other systems, subsystems, modules, blocks, or components.
[0077] FIG. 1 is a block diagram of a system 100 capable of implementing the virus detection and lineage assignment techniques described herein. System 100 is a dynamic clinical assay pipeline including a molecular capture subsystem 110, a sequencing subsystem 120, and an analysis subsystem 130. The systems and subsystems, modules, blocks, or components (e.g., programs, code, or instructions) illustrated in FIG. 1 may be executable by an operator (e.g., a human user), a dedicated machine (e.g., a sequencer), or one or more processors (e.g., a CPU, GPU, etc.) with the aid of or use of a computing device, according to various embodiments. In some cases, the virus to be detected is a coronavirus, influenza virus, or respiratory syncytial virus (RSV). In some cases, the virus to be detected is a non-smallpox orthopoxvirus, such as monkeypox or cowpox. In another example, the virus to be detected is a smallpox orthopoxvirus, such as variola. The virus typically has a length of several thousand to several million bases or base pairs. For example, the genome of monkeypox is approximately 190,000 base pairs, and the genome size of SARS-CoV-2 is approximately 30,000 bases (30,000 nucleotides).
[0078] The molecular capture subsystem 110 is a wet-lab subsystem that requires water, direct ventilation, and dedicated plumbing to test and analyze chemicals, drugs, other substances, or biological materials. The molecular capture subsystem 110 may include one or more blocks, modules, or services. For example, as shown in FIG. 1 , the molecular capture subsystem 110 includes a nucleic acid capture block 112, a probe oligonucleotide capture block 114, and a multiplexing module 116 that includes an annealing block 111, a gap-filling block 113, a probe removal / release block 115, a PCR amplification block 117, and a post-PCR wash block 118. The input to the molecular capture subsystem 110 is one or more samples 105, and the output of the molecular capture subsystem 110 is a plurality of amplified molecules 125.
[0079] Nucleic acid is acquired from one or more samples 105 in the nucleic acid acquisition block 112. The one or more samples 105 may be collected or acquired from a subject prior to acquisition in the nucleic acid acquisition block 112. An appropriate sample collection method may be used depending on the sample type and downstream application. For example, blood samples may be collected using venipuncture or finger prick, and saliva samples may be collected using commercially available collection kits. In some embodiments, the one or more samples 105 are collected using a sample self-collection kit, such as a monkeypox PCR test home collection kit. The monkeypox PCR test home collection kit is intended for individuals presenting with an acute generalized pustular or vesicular rash suspected of monkeypox to self-collect a lesion swab specimen in a culture medium at home. The swab specimen is placed in a culture medium and transported to a laboratory for testing of non-smallpox orthopoxvirus DNA extracted from the specimen.
[0080] Nucleic acids from one or more samples 105 may be obtained by isolation and extraction. For example, an appropriate test sample preparation method may be used to disrupt cells and tissues to release nucleic acids and remove contaminants such as proteins, lipids, and other cellular debris. Various sample preparation methods are available, such as phenol-chloroform extraction, column-based purification, and magnetic bead-based purification. Once the sample is prepared, nucleic acids must be extracted using an appropriate extraction method. The extraction method may be selected depending on the type of nucleic acid of interest (e.g., DNA or RNA) and the downstream application. Commonly used extraction methods include organic extraction, silica-based column purification, or magnetic bead-based extraction. Appropriate quality control methods may also be performed in the nucleic acid acquisition block 112 to confirm the quality of the obtained nucleic acids. In some cases, one or more samples 105 may be nucleic acid samples, eliminating the need for a preparation or extraction method. For example, one or more samples 105 may be nucleic acids obtained from a patient.
[0081] One or more probe oligonucleotides are prepared and acquired in the probe oligonucleotide acquisition block 114. The probe oligonucleotides may be pre-designed to be suitable for capturing target molecules and detecting the presence or absence of a virus in a sample. For example, if the virus to be detected is monkeypox virus, the probe oligonucleotides may be designed to be suitable for capturing monkeypox virus DNA molecules of a predetermined base pair length (e.g., approximately 675 base pairs). The pre-designed probe oligonucleotides may then be synthesized using an appropriate method, such as solid-phase synthesis, enzymatic synthesis, or chemical synthesis. The synthesized probe oligonucleotides are acquired in the probe oligonucleotide acquisition block 114. The probe oligonucleotide acquisition block 114 may perform only the function of acquiring the probe oligonucleotides. The design and synthesis / manufacture of the probe oligonucleotides may be performed by another system, subsystem, or block.
[0082] In the multiplexing module 116, functions such as annealing, gap filling, probe removal and release, and amplification may be performed in different blocks. For example, the annealing block 111 anneals the probe oligonucleotides obtained in block 114 to target molecules. Annealing is a key step in hybridization assays, allowing the probe oligonucleotides to specifically bind to their target sequences. In the annealing block 111, different probe oligonucleotides are designed and annealed to different target molecules. If the target molecule is a double-stranded DNA molecule, the corresponding probe oligonucleotide that captures the DNA molecule may be a double-stranded oligonucleotide. If the target molecule is a single-stranded RNA molecule, the corresponding probe oligonucleotide may be a single-stranded DNA molecule and annealed to the RNA molecule to obtain an RNA-DNA hybrid duplex in which the RNA and DNA strands are complementary to each other. The annealing process may form a double-stranded structure for subsequent steps or processes.
[0083] Gaps or missing sections in the DNA or RNA strands may be filled in gap filling block 113. Any suitable method may be used to perform gap filling, including polymerase mediated synthesis or homologous recombination.
[0084] In probe removal / release block 115, probes that do not react with any target molecules are removed, and the remaining probes are released from the hybrid molecules. Removal of unreacted probes may be accomplished using a variety of methods, including washing the mixture with a buffer that disrupts probe-target interactions, or using enzymatic digestion to degrade the probe or target sequence. Probe release may be achieved by heating, denaturation, or enzymatic digestion.
[0085] The target molecule may be amplified using PCR technology in PCR amplification block 117. To perform PCR amplification, primers complementary to adjacent regions of the target molecule, a polymerase enzyme, and deoxynucleotide triphosphates (dNTPs) may be used. The primers serve as starting points for the polymerase to synthesize new DNA or RNA strands during PCR amplification, and a different polymerase enzyme may be used to bind to the primers and synthesize the new DNA or RNA strands.
[0086] Unwanted components are removed in a post-PCR wash block 118. It is often necessary to perform a post-PCR wash step to remove unwanted reaction components such as primers, nucleotides, enzymes, and other impurities. Various post-PCR wash methods may be used in the post-PCR wash block, including gel electrophoresis, bead-based purification, and spin-column purification. Post-PCR washes can improve the accuracy and efficiency of downstream applications. After the post-PCR wash step, amplified molecules 125 are obtained and may be used in a subsequent sequencing / analysis subsystem. The target molecules are amplified thousands to millions of times. In some cases, the number of amplified molecules is approximately 16 million times the number of target molecules.
[0087] The sequencing subsystem 120 may be implemented within a sequencer, an automated system capable of sequencing DNA or RNA molecules and analyzing gene fragments for various applications. Examples of sequencing systems that may be used to implement the sequencing subsystem 120 include Sanger sequencers, Illumina sequencers, PacBio sequencers, Oxford nanopore sequencers, and Ion Torrent sequencers. For example, the sequencing subsystem 120 may be a capillary electrophoresis-based system in which probe-bound gene fragments migrate through a polymer and fluorescence emission is measured. A multiple capillary array allows sample loading in a multiwell microplate format. The sequencing subsystem 120 may also be a pyrosequencing technology-based system for rapid sequencing and analysis. Pyrosequencing is a method of DNA sequencing (and sometimes RNA sequencing) based on the "sequencing-by-synthesis" principle, in which sequencing is achieved by detecting nucleotides incorporated by DNA polymerase. Pyrosequencing relies on optical detection of pyrophosphate release, which is a chain reaction-based process. The modules, blocks, or components may be stored on non-transitory computer media. As needed, one or more of the modules, blocks, or components may be loaded into system memory (e.g., RAM) and executed by one or more processors in the sequencing subsystem 120. The sequencing subsystem 120 may include one or more blocks, modules, or services. For example, as shown in FIG. 1, the sequencing subsystem 120 includes an adapter connection block 122, a sequencing block 124, and a demultiplexing block 126.
[0088] In the adapter ligation block 122, adapters may be ligated to the molecules of the amplified molecule 125. During adapter ligation, adapters are ligated to the ends of the molecules using a ligase (e.g., DNA ligase). For example, the adapters used in the adapter ligation block 122 may be SMRTbell adapters. SMRTbell adapters are adapter sequences used in PacBio's Single Molecule Real-Time (SMRT) sequencing technology. SMRTbell adapters consist of two complementary oligonucleotide strands that can be annealed to the ends of a molecule and then ligated to each other to form a circular molecule. The circular molecule can then be sequenced using PacBio's SMRT sequencing technology. Other suitable adapters may also be used in the adapter ligation block 122.
[0089] In the sequencing block 124, the molecules (e.g., circular molecules) obtained from the adaptor ligation block 122 are sequenced using an appropriate sequencing technology to obtain sequence reads. In this process, thousands to millions, or even millions to billions, of sequence reads are typically obtained. In the sequencing block 124, circular consensus sequencing (CCS) technology may be used, which allows DNA polymerase to repeatedly pass through the same region of the circular molecule. The CCS sequencing used in the sequencing block 124 also ensures longer read lengths and higher consensus accuracy. In some cases, a consensus sequence read is generated based on the raw sequence reads. The consensus sequence read may be referred to herein as a "sequence read." It should be understood that other sequencing technologies, such as nanopore sequencing, may also be used in the sequencing block 124 to provide sequence reads suitable for applying the disclosed technology.
[0090] In the demultiplexing block 126, the sequence reads obtained in the sequencing block 124 may be separated by sample (or molecule). Demultiplexing is the process of separating pooled sequence reads based on sample-specific barcodes or indexes added to each sample during library preparation. When sequencing involves multiple samples, sequence reads from different samples are often pooled together and sequenced in a single run in the sequencing block 124. Demultiplexing is performed to separate the sequence reads from each sample. In some cases, demultiplexing also includes an adapter trimming step. In another example, adapters are trimmed in the sequencing block 124.
[0091] The analysis subsystem 130 may execute on a computing device such as a computer, a central processing unit (CPU), a graphics processing unit (GPU), or the like. The modules, blocks, or components of the analysis subsystem 130 may be stored on a non-transitory computer medium. As needed, one or more of the modules, blocks, or components may be loaded into system memory (e.g., RAM) and executed by one or more processors in the analysis subsystem 130. The analysis subsystem 130 may include one or more blocks, modules, or services. For example, as shown in FIG. 1 , the analysis subsystem 130 includes a sequencing file block 132, a determination module 134, and a reporting file block 136.
[0092] In the sequencing file block 132, a sequencing file is generated by a computing device on which the analysis subsystem 130 is executed. The sequencing file may be a FASTQ file, a BAM file, a SAM file, a FASTA file, a CRAM file, etc. The generation of the sequencing file is performed based on the output of the sequencing subsystem 120. The sequencing file generally includes sequencing information, such as sequence reads for each amplified molecule and positioning data associated with the sequence reads. The positioning data may include the start / end position of each sequence read in a reference viral genome. In some cases, the positioning data is determined by mapping or aligning each sequence read to the reference viral genome. Different alignment algorithms may be used for alignment. The sequencing file may be output to the determination module 134 for virus detection and / or lineage assignment.
[0093] In the determination module 134, functions including generating a consensus sequence, determining or obtaining a decision rule, and determining the presence or absence of a virus and / or determining the strain of the virus may be performed in different blocks. For example, as shown in FIG. 1 , the determination module 134 includes a consensus sequence block 131, a decision rule block 133, and a virus detection block 135.
[0094] In the consensus sequence block 131, a consensus sequence for each target molecule is generated. Different consensus sequence generation rules may be used to generate the consensus sequence in the consensus sequence block 131. In some cases, if at least a predetermined number of sequence reads have nucleotide identity at a position in the consensus sequence, the position is assigned a nucleotide identity; if fewer than a predetermined number of sequence reads have nucleotide identity at a position in the consensus sequence, the position is assigned an "N." In some cases, the predetermined number is four. Because the wet-lab subsystem and the sequencing subsystem guarantee read coverage, a generation method based on consistent reads works efficiently and effectively to determine the consensus sequence. In most cases, at least four-read coverage is guaranteed at 80% of positions in the reference virus genome.
[0095] In the decision rule block 133, a decision rule for virus detection and / or lineage assignment is obtained. The decision rule may be predetermined by a researcher, a physician, a medical practitioner, or an authorized person. The decision rule may also be determined using a computing device based on an algorithm or a machine learning model. In some cases, the consensus sequence is mapped to a reference genome or a portion of a reference genome, and one or more scores are determined based on the decision rule. In some cases, the one or more scores are determined based on the distribution of nucleotide comparisons (e.g., nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc.) between the consensus sequence and the reference genome or portion thereof. For example, the distribution indicates the number of nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc. The resolution of the nucleotide comparisons may vary. In some cases, the nucleotide comparison is a comparison of a single nucleotide. In another example, the nucleotide comparison is a comparison of multiple nucleotides (e.g., a comparison of three nucleotides). Different scores may be assigned to different types of nucleotide comparisons. In some cases, different scores may be assigned to subtypes of the same type of nucleotide comparison. For example, the score assigned to an "A" to "G" substitution may be different from the score assigned to an "A" to "C" substitution. In some cases, different scores are assigned based on the phylogenetic context of the nucleotide comparison (e.g., evolutionary relationships or phylogenetic classification). For example, nucleotide substitutions that are common in a particular lineage may receive a different score than rare ones. In some cases, nucleotide comparisons that result in changes to encoded amino acids within protein-coding regions may receive a higher score, especially if the change is known to affect protein function. In some cases, different types of nucleotide comparisons are assigned the same score.
[0096] The decision rule may include a similarity metric that indicates the relationship between the consensus sequence and the reference genome (or reference sequence). In some cases, the similarity metric assigns different weights to one or more scores associated with different consensus sequences. For example, the weight assigned to a perfect match between the consensus sequence and the reference genome is different from the weight assigned to a mismatch between the consensus sequence and the reference genome. In some cases, the similarity metric assigns a positive weight to the total number of mutations in the consensus sequence and the total number of mutations in the reference genome, and a negative weight to nucleotide comparisons between the consensus sequence and the reference genome. The decision rule may also include a threshold for determining the presence or absence of a virus and / or a threshold for determining a predicted lineage assignment. In some cases, if one or more scores are equal to or greater than the threshold, it is determined that a virus is detected or a predicted lineage is assigned to the virus in the sample. If one or more scores are less than the threshold, it is determined that a virus is not present in the sample or that no lineage is assigned to the virus in the sample.
[0097] The decision rules may comprise a decision tree model. In some cases, the decision tree model is an annotated phylogenetic tree with different lineages as nodes. In some cases, the decision rules are stored in a tree-like data structure, such as a binary tree, a binary search tree, a trie tree, an n-ary tree, a general tree, a B-tree, or an XML tree. Storing the decision rules in a tree-like data structure can improve the efficiency of data retrieval (e.g., decision-making) and also increase space efficiency. Tree structures can easily accommodate dynamic data, allowing for efficient updates, insertions, and deletions. This provides flexibility for data that changes over time and is compatible with dynamic clinical assay pipeline designs such as those disclosed herein. Tree structures also provide a feasibility for visualizing viral lineage assignments.
[0098] In virus detection block 135, a determination of the presence or absence of a virus in the sample and / or a determination of the predicted lineage of the virus is made. In some cases, both the determination of the presence or absence of a virus in the sample and the determination of the predicted lineage of the virus are made based on the consensus sequence of the target molecule generated in consensus sequence block 131 and a decision rule obtained in decision rule block 133. In some cases, the presence or absence of a virus in the sample is determined or pre-determined using a wet-lab method, such as real-time PCR, viral culture, immunofluorescence assay, electron microscopy, etc. If the presence of a virus in the sample has been pre-determined (e.g., using real-time PCR or based on a decision rule), a predicted lineage assignment of the virus in the sample is determined in virus detection block 135. The predicted lineage assignment of the virus in the sample may be determined by assigning lineages to the viruses in the sample based on a decision tree model (e.g., the decision tree model obtained in block 133).
[0099] In some cases, a predicted lineage is assigned based on a similarity metric. One or more scores for the consensus sequence are determined in virus detection block 135 based on the decision rule obtained in block 133, and the one or more scores are used to determine the presence or absence of the virus in the sample and / or to determine a predicted lineage assignment for the virus in the sample. In some cases, the one or more scores are compared to each other, and the known lineage associated with the highest or lowest score is assigned to the virus in the sample. In some cases, if each of the one or more scores exceeds a predetermined threshold, the virus in the sample is assigned an unknown lineage (or marked as a "novel" lineage), and further validation is performed to confirm the presence of this unknown lineage. Validation may be performed by peer review, or by replication using sequencing techniques, data analysis, systematic methods (e.g., the pipelines disclosed herein), or the like. In some cases, bioinformatics tools such as NextClade or Pangolin may be used or modified to determine the one or more scores, or to determine the presence or absence of the virus in the sample and / or to determine a predicted lineage assignment.
[0100] In a report file block 136, a report file is generated by the computing device on which the analysis subsystem 130 is executing. The report file may be a variant call format (VCF) file, a VCFCons file, a BCF file, a MAF file, a GFF file, etc. The generation of the report file is based on the output of the decision module 134. In some cases, the report file includes (i) the presence or absence of the virus in the sample and / or (ii) a predicted lineage of the virus in the sample. In some cases, the report file is output to a graphical user interface (GUI). In some cases, the report file also includes a display of the lineage assignment based on the decision rule obtained in block 133 (see, e.g., Figures 12A-12B and 13B-13C).
[0101] In some cases, the information generated by the determination module 134 is used to dynamically modify the molecular capture subsystem 110 and / or the sequencing subsystem 120. For example, the prevalence status of a strain can be determined based on strain assignments determined over a certain period of time. Based on the prevalence of the strain, the probes used in the molecular capture subsystem 110 may be reselected or redesigned to more accurately capture target molecules corresponding to the prevalence of the strain. The adapters used in the sequencing subsystem 120 may also be reselected or redesigned to adapt to changes in the probe / target molecules. Dynamic modification of the system 100 via feedback from the determination module 134 to the molecular capture subsystem 110 and / or the sequencing subsystem 120 helps to reconstruct a dynamic clinical assay pipeline in real time. Dynamic modification of the system 100 also improves the efficiency of capture and sequencing for cases using known or estimated prevalence status of the strain. Dynamic modification of the system 100 also contributes to providing accurate virus detection and strain assignment results.
[0102] For example, if lineage B.1 is found to be prevalent in a subject population (e.g., defined by a particular race, a particular age group, a particular geographic location, etc.), one or more probe oligonucleotides prepared in block 114 can be redesigned to capture target molecules specific to lineage B.1. The adapters used in adapter ligation block 122 can also be updated based on the target molecules specific to lineage B.1. Furthermore, the decision rules obtained in block 133 can be updated accordingly to reflect the prevalence of lineage B.1. These updates improve the accuracy of predicted outcomes. Furthermore, this optimization helps significantly increase the efficiency and cost-effectiveness of future predictions and assignments of viral lineages for samples collected from subject populations. These dynamic adjustments and enhancements represent substantial improvements in the fields of viral detection and lineage assignment, as well as in the area of viral assay procedures.
[0103] FIG. 2 is an exemplary flowchart illustrating a process 200 for generating a report file regarding the presence or absence of a virus or the predicted strain of a virus in a sample, according to various embodiments. In certain embodiments, the process or portions of the process illustrated in FIG. 2 may be performed by system 100 as described with respect to FIG. 1. Process 200 begins with block 210, which involves acquiring nucleic acid from a sample. The sample may be previously acquired from a subject or an organism, such as a plant, animal, or human. The sample may be a blood sample, a urine sample, a tissue sample, a saliva sample, a fecal sample, a hair sample, a bone sample, etc. The acquired nucleic acid may include deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA). Various nucleic acid acquisition methods may be used to acquire nucleic acid from the sample in block 210.
[0104] At block 220, target molecules in the nucleic acid are captured. In some embodiments, the target molecules are DNA molecules. In some embodiments, the target molecules are RNA molecules. In some embodiments, the target molecules are a mixture of DNA molecules and RNA molecules. In various embodiments, at block 220, multiple target molecules in the nucleic acid are captured. In other embodiments, a single target molecule in the nucleic acid may be captured. Various molecular capture methods may be used to capture the target molecules. In various embodiments, the target molecules are captured using probe hybridization. In other embodiments, the target molecules are captured using methods such as affinity capture, chromatography, microfluidics, etc.
[0105] In block 230, the captured molecules are amplified to obtain amplified molecules. In various embodiments, the captured molecules are amplified using polymerase chain reaction (PCR) technology. Other methods, such as isothermal amplification, multiple displacement amplification (MDA), rolling circle amplification (RCA), digital PCR (dPCR), reverse transcription polymerase chain reaction (RT-PCR), and digital droplet PCR (ddPCR), may also be used to amplify the captured molecules in block 230. Assigning substantially unique barcodes to the captured molecules before amplification enables tracking of the information. In some embodiments, the same barcode is assigned to target molecules in the same sample. In some embodiments, each target molecule has a unique barcode, which can be used to distinguish the target molecules. The barcode sequence may be predetermined to reduce laboratory costs and improve laboratory preparation efficiency. Alternatively, the barcode sequence may be randomly generated to ensure substantial uniqueness. The steps in blocks 220 and 230 may be implemented in the molecular inversion probe set design and PCR pipeline disclosed in detail in Section IV and FIG. 4.
[0106] In block 240, the amplified molecules are ligated to adapters for efficient sequencing. In some embodiments, the adapters are SMRTbell adapters of a selected size. SMRTbell adapters are adapters used in single molecule real-time (SMRT) sequencing technology. They have a hairpin structure in which the 5' and 3' ends of the adapters bind to each other to form a loop. This structure is important for SMRT sequencing because it allows the SMRTbell adapters to circularize, which allows polymerase to be passed through the same molecule multiple times. Each SMRTbell adapter may also contain a substantially unique barcode / sequence used for sample authentication and demultiplexing during data analysis. This allows multiple samples to be sequenced together in a single SMRT cell, which improves sequencing technology by increasing sequencing efficiency and reducing the cost per sample. Primer annealing and polymerase binding steps may also be performed in block 240. Ligated molecules are generated in block 240. In some embodiments, the ligated molecules are circular molecules.
[0107] In block 250, the linked molecules are sequenced using an appropriate sequencing method to obtain sequence reads. In some embodiments, the sequencing method is circular sequencing. In some embodiments, the sequencing is SMRT sequencing. Other sequencing methods, such as RCA sequencing or nanocircle sequencing, may also be used to sequence the linked molecules. The steps of blocks 240 and 250 may be implemented within the adapter design, ligation, and sequencing pipeline disclosed in detail in Section V and FIG. 5.
[0108] At block 260, a sequencing file is generated. The sequencing file may contain information related to the sequencing step at block 250. In some embodiments, the sequencing file contains a sequence read for each linked molecule obtained by aligning the sequence reads to the reference genome of the virus and the position of each sequence read in the reference genome of the virus. In some embodiments, the sequencing file is a BAM file. The sequencing file may be generated by a computing device or a sequencer.
[0109] In block 270, a report file is generated. The report file may be generated by the same computing device that generates the sequencing file. The report file includes information regarding the presence or absence of a virus in a sample and / or its predicted lineage. A consensus sequence for each target molecule may be determined in block 270. Determining the consensus sequence can increase the accuracy of determining the presence or absence of a virus in a sample and / or its predicted lineage. In some embodiments, a genome construct may be determined based on the consensus sequence of the target molecules, and the genome construct may be used to determine the presence or absence of a virus in a sample and / or its predicted lineage. In some embodiments, the sample is known to contain viral nucleic acid, and only the lineage of the virus in the sample is determined. One or more scores may also be determined in block 270 based on a decision rule that is predetermined or determined based on an algorithm or machine learning model. The determination of the presence or absence of a virus in a sample and / or its predicted lineage may be made in block 270 based on the one or more scores. In some embodiments, the report file is a VCF file or a modified VCF file. In some embodiments, the report file may be output to a GUI for a practitioner or a subject. The steps of blocks 260 and 270 may be implemented within the adjusted VirSeq pipeline disclosed in detail in Section VI.
[0110] In optional block 280, a treatment plan or clinical trial protocol is provided. The treatment plan or clinical trial protocol may be determined based on one or more of the scores, the presence or absence of the virus in the sample and / or its predicted lineage, or the report file generated in block 270. In some embodiments, the virus is monkeypox (MPX) virus and the treatment plan includes administering an antiviral drug to the subject.
[0111] FIG. 3 is an exemplary flowchart illustrating a process 300 for implementing the disclosed technology for predicting the presence or absence of a virus in a sample or predicting the lineage of a virus, according to various embodiments. In an embodiment, the process or portions of the process illustrated in FIG. 3 may be implemented by system 100, as described with respect to FIG. 1. Process 300 may be performed within three pipelines: in block 310, nucleic acids are obtained and sent to a molecular loop probe set design and PCR pipeline to obtain amplified molecules in block 320; the amplified molecules are then sent to an adapter design, ligation, and sequencing pipeline to obtain sequence reads in block 330; and the sequence reads are sent to a tuned VirSeq pipeline to obtain a personalized report; and the personalized report is output in block 350. Process 300 is described in more detail in Sections IV, V, and VI. IV. Molecular inversion probe set design and PCR pipeline
[0112] The molecular loop probe set design and PCR pipeline corresponds to subsystem 110 in FIG. 1 and is further illustrated in FIG. 4. The molecular loop probe set design and PCR pipeline is designed to capture target molecules and amplify the captured molecules so that a sufficient number of molecules can be sent to the adapter design, ligation, and sequencing pipeline to ensure a sufficient number of sequence reads and sufficient coverage of the sequence reads and reference genome. In some embodiments, a sufficient number of molecules / sequence reads is about 4-8 million, where sufficient coverage is (i) about 50% coverage or greater (e.g., 50%, 60%, 70%, 80%, 90%), or (ii) about 100,000 sequence reads or greater (e.g., 100,000 sequence reads, 1 million sequence reads, 10 million sequence reads).
[0113] To detect the presence or absence of a virus in a sample or the predicted lineage of the virus, nucleic acid must be obtained from the sample. Nucleic acid often contains both host genetic molecules and potential or suspected viral molecules. Therefore, the molecular loop probe set design and PCR pipeline must be able to capture viral molecules in nucleic acid. In some embodiments, the target molecules captured by the molecular loop probe set design and PCR pipeline are viral molecules. The viral molecules may be DNA molecules or RNA molecules. DNA molecules are double-stranded molecules, while RNA molecules are often single-stranded. In some embodiments, the virus to be detected is an MPX virus, which is a double-stranded DNA virus.
[0114] A probe hybridization method may be used to capture target molecules. The molecular loop probe set design and PCR pipeline may include a component for designing probes to capture target molecules. Appropriate design methods may be used to capture target molecules, and these methods may depend on various factors, including the type of molecule, sensitivity and specificity requirements, and downstream analysis needs. For example, if the virus to be detected is an MPX virus and the target molecule is an MPX molecule, the probe may be designed to have two binding sites separated by approximately 600 to 700 base pairs (bp). In some embodiments, a probe for detecting an MPX molecule consists of two binding sites separated by 675 bp. In some embodiments, multiple probes are designed to capture multiple target molecules. For example, to capture MPX virus molecules, more than 7,000 (e.g., 7,826) probes may be designed to capture 675 bp molecules uniformly tiled across the approximately 200 kilobase pairs (kb) MPX genome.
[0115] In some embodiments, the probe that captures the target molecule is a molecular inversion probe. Molecular inversion probes may provide higher specificity and sensitivity in capturing target molecules compared to other types of probes. For example, molecular inversion probes are designed to hybridize to specific molecules or genomic regions of interest, and their specificity reduces the likelihood of false positives or false negatives in downstream analysis. Due to their high signal-to-noise ratio, molecular inversion probes can also detect low levels of DNA or RNA molecules in nucleic acids. Furthermore, molecular inversion probes can be designed to simultaneously capture multiple target molecules or genomic regions, enabling high-throughput analysis of multiple targets. Furthermore, molecular inversion probes are cost-effective compared to some other target capture methods. Furthermore, molecular inversion probes can be easily designed and customized to meet specific research and industrial / clinical needs.
[0116] Figure 4 illustrates a portion of a molecular loop probe set design and PCR pipeline for capturing target DNA molecules, according to various embodiments. As shown in Figure 4, a molecular inversion probe may be designed to have two target-complementary regions, H1 and H2, two PCR primers, P1 and P2, and a probe-release cleavage site, X1. As shown in step 1 of Figure 4, in the probe hybridization process, multiple probes may be added to nucleic acids obtained from a sample, with each probe annealing to a target molecule on the H1 and H2 regions. The molecular inversion probe may be tiled across more than 99% (e.g., 99.6%) of the 200 kb MPX virus genome and may consist of two binding sites approximately 675 bp apart.
[0117] Molecular inversion probes can be beneficial in capturing target molecules. First, they enable tiling and prevent the dropout of novel genomic mutations. Second, unlike other probe-based assays that require genome reorientation to capture molecules, molecular inversion probes do not require upstream genome processing, making pipelines using molecular inversion probes much faster and relatively inexpensive with high throughput. Furthermore, they help optimize sequencer loading and analysis pipelines to ensure the expected product size. They also ensure not only complete (or sufficient) coverage of the genome, but also deep coverage of the genome at most base positions.
[0118] After binding, the region between the two probes is synthesized and ligated by DNA polymerase to form a closed molecule, as shown in step 2 of Figure 4. The gap is filled with free nucleotides by DNA polymerase, and the ends of the probes are joined by ligase, resulting in a fully circularized probe.
[0119] Unreacted probes remain linear because no gap-filling step is performed on them. Unbound or incomplete loops remain linear molecules and are removed by exonuclease digestion, as shown in step 3 of Figure 4.
[0120] The circular molecule is released and linearized, as shown in step 4 of Figure 4. In some embodiments, the probe-release cleavage site X1 is cleaved by a restriction enzyme so that the probe is linearized. In this linearized probe, the universal PCR primer sequences are located at the 5' and 3' ends, and the captured genomic target becomes part of the internal segment of the probe. In another embodiment, the probe remains a circularized molecule, and no cleavage is required.
[0121] The captured molecules can then be enriched by amplification using the 3' molecular loop-specific M13 universal sequence and the 5' sample-specific barcode, as shown in step 5 of Figure 4. If the probe is linearized, conventional PCR amplification can be performed using the probe's universal primer to enrich the captured molecules. Otherwise, circularized probes can be subjected to cyclic amplification methods such as RCA.
[0122] In some embodiments, adapters P1 and P2 in Figure 4 may be used for barcoding. Samples may then be pooled in equal amounts, and libraries prepared for sequencing (e.g., PacBio sequencing). Library preparation may involve DNA damage repair, ligation of sequencing adapters, enzymatic digestion to remove unligated products, and bead purification. The libraries may then be sequenced, for example, on a PacBio Sequel II with a 15-hour video recording.
[0123] In some embodiments, the target molecule is an RNA molecule. The molecular loop probe set design and PCR pipeline can also capture and amplify RNA molecules. In such embodiments, reverse transcriptase can be used to synthesize cDNA from the RNA. V. Adapter design, ligation, and sequencing pipeline
[0124] The adapter design, ligation, and sequencing pipeline may correspond to subsystem 120 in FIG. 1 and is further illustrated in FIG. 5. The adapter design, ligation, and sequencing pipeline may be designed to provide a circular sequencing pipeline that obtains stable sequence read coverage per site and across the genome. In various embodiments, SMRT sequencing technology may be used to generate sequence reads. To use SMRT sequencing technology, adapters may be designed and ligated to target molecules to generate circular molecules. SMRT sequencing technology may increase sequence read lengths to an average number substantially longer than conventional sequencing methods. In some cases, SMRT sequencing technology can produce sequence reads with an average length of 50 kb. With circularized DNA molecules, this sequence read length may provide many more consecutive passes (e.g., 70) of a single target molecule (e.g., in the size range of approximately 600-700 bp). The individual passes may be combined into a single consensus high fidelity consensus sequence (HiFi) read (e.g., 600-700 bp in length) with 99% accuracy.
[0125] Sequencing using single molecule real-time (SMRT) long-read sequencing technology requires a circular template (sequencing molecule). The circular template generated from library preparation is combined with a polymerase and primer and loaded onto an SMRT cell (sequencing cell). The single molecule product diffuses into one of eight million zero-mode waveguide (ZMW) wells, to which a polymerase is immobilized. A phosphate-linked nucleotide is then introduced into the ZMW, where the base can subsequently be incorporated by the polymerase. When a given base pair is incorporated, its addition produces a nucleotide-specific emission that is detected by a camera for each well. This process is repeated for a given time or video length, and the order of nucleotides in a given well is analyzed and translated into the corresponding nucleotide in the long-sequence read output.
[0126] Figure 5 illustrates a portion of an adapter design, ligation, and sequencing pipeline for generating circular DNA molecules, according to various embodiments. As shown in Figure 5 (see schematic workflow at https: / / ccs.how / , last visited March 22, 2023), the adapter design, ligation, and sequencing pipeline starts with a high-quality double-stranded DNA molecule. Hairpin adapters may be ligated to a linear DNA molecule to produce a circular DNA molecule. In some embodiments, the adapters are SMRTbell adapters. The adapters may vary in size. The circular DNA molecule enables the pipeline to produce long sequence reads that can be used to generate HiFi reads. After adapter ligation, selected primers and DNA polymerase may be added to sequence the circular DNA molecule and obtain sequence reads.
[0127] In some embodiments, as shown in the right part of Figure 5, adapters may be trimmed from long sequence reads to obtain subreads, and HiFi reads may be generated based on the subreads. There are several advantages associated with HiFi reads. First, HiFi reads provide both read length and accuracy suitable for various downstream analyses, such as variant calling. Furthermore, this technology also improves sequencing efficiency, i.e., it increases throughput and sequencing efficiency while reducing sequencing time and cost per sample. VI. Adjusted VirSeq pipeline
[0128] The tailored VirSeq pipeline may correspond to the analysis subsystem 130 of FIG. 1. The tailored VirSeq pipeline is designed to provide a personalized report on the viral information in a sample obtained from a subject. The tailored VirSeq pipeline takes as input sequence reads from the adapter design, ligation, and sequencing pipelines and outputs a personalized report. In some embodiments, the personalized report is displayed in a GUI.
[0129] In various embodiments, a sequencing file may be generated based on the sequence reads and a reference viral genome. The sequence reads may be raw sequence reads generated by the adapter design, ligation, and sequencing pipeline, or HiFi reads generated using the raw sequence reads. The reference viral genome may be obtained from a viral database or modified based on a viral genome obtained from a viral database. In some embodiments, the sequencing file is a BAM file. In some embodiments, the sequencing file is a FASTQ file. Software, scripts, and codes may be used to generate the sequencing file.
[0130] The tailored VirSeq pipeline and dynamic clinical assay pipeline are designed to be easily adaptable to the detection of different viruses with diverse genome lengths. In some embodiments, the tailored VirSeq pipeline and dynamic clinical assay pipeline are designed to detect MPX viruses. Figure 6 provides an overview of the complexity of the MPX genome (see Monkeypox Virus Genomics, https: / / nextstrain.org / monkeypox / hmpxv1, last visited March 22, 2023). The MPX genome is approximately 10 times larger than the SARS-CoV-2 genome and contains many long stretches of repetitive DNA regions and drop regions. In the molecular loop probe set design and PCR pipeline, more probes may be designed and used to capture target molecules in more complex viral genomes. For example, the number of probes required for MPX virus detection is approximately 7–10 times that required for SARS-CoV-2 detection.
[0131] In some embodiments, PacBio SMRT LINK software and custom molecular loop processing scripts may be used to generate sequencing files. The sequencing files may be analyzed using a genome analysis pipeline. In some embodiments, the genome analysis pipeline may be implemented using the CLC Genomics Server. It should be understood that other sequencing analysis systems may also be used. At this point, sequencing primer sequences can be removed, and the sequences are aligned to a reference viral genome (e.g., the MPX reference genome) to generate a BAM file of the alignment. In some embodiments, Minimap2 may be used for alignment. Other alignment programs or algorithms may also be used. In some embodiments, sequence reads that meet a minimum coverage of 50% are used as input for generating consensus sequences and / or genome constructs, as well as for detecting the virus. The upper limit of the minimum coverage may vary from 20% to 100% (e.g., 20%, 30%, 40%, 60%, 70%, 80%, 90%).
[0132] The design of a tailored VirSeq pipeline and a dynamic clinical assay pipeline can achieve better coverage. This superior coverage is achieved by the wet-lab methods and pipelines described in this disclosure. These methods and pipelines involve precise design of the molecular capture and ligation processes, ensuring high-quality molecular sequencing. Furthermore, sequencing techniques such as circular consensus sequencing further enhance uniform read coverage across the viral genome, thereby improving overall sequencing quality. Figure 7 shows an Integrative Genomics Viewer (IGV) plot with reads per nucleotide on the Y-axis and genomic coordinates on the X-axis. Figure 7 demonstrates highly uniform and complete coverage of the MPX genome. The dark vertical bars indicate the positions where variants were detected compared to the reference sequence. Variant detection is introduced later in this section.
[0133] Figure 8 is a high-resolution graphical representation of the IGV plot shown in Figure 7, generated using the developed software for various clinical samples. The genome coverage and CCS read coverage per base for eight samples are shown in Figure 8. When the minimum coverage threshold was predetermined to 4 (shown by the horizontal dashed line in Figure 8), seven of the eight samples reached the minimum coverage threshold, demonstrating that the disclosed technology provides a robust method for generating sequence reads.
[0134] Figure 9 shows another way of showing genome coverage and read depth for various samples and controls. Sequencing data for 16 samples is shown in Figure 9. The vertical dashed line indicates the minimum coverage threshold (per base coverage equal to 4). As shown in Figure 9, 15 of the 16 samples reached the minimum coverage threshold.
[0135] Figure 10 shows genome coverage and per-base coverage according to various embodiments. The top of Figure 10 shows that 32 of 48 samples have genome coverage greater than 80% at 50x base coverage. The bottom of Figure 10 shows the mean (solid line) and 25 / 75 quartile coverage (dashed line) for individual samples. Figure 10 demonstrates that the tuned VirSeq pipeline and dynamic clinical assay pipeline perform well with larger multiplex runs. This is achieved by obtaining and ensuring more high-quality sequence reads per molecule (e.g., using precise design of the molecule capture and ligation process). As a result, fewer molecules are required per sample, allowing more samples to be sequenced simultaneously. This also increases sequencing efficiency and reduces sequencing costs.
[0136] In some embodiments, generating sequencing files may include preprocessing such as demultiplexing. In an embodiment, preprocessing may include generating circular consensus sequence (CCS) BAM files, merging intermediate BAM files, using demultiplexing to generate individual BAM files corresponding to different barcode combinations, combining the demultiplexed output by sample name and / or subject identifier, removing barcodes from sequences to generate individual sample FASTQ files, aligning sequences to barcodes and trimming barcodes, converting BAM files to FASTQ files, and copying the FASTQ and CCSBAM files to a final location as sequencing files.
[0137] In some embodiments, generating the sequencing file may further include filtering the sequence reads based on length or quality and aligning the sequence reads to a reference viral genome (e.g., the MPX reference genome). In some embodiments, this alignment is a local alignment performed using tools within the CLC Genomics Server.
[0138] The sequencing files may then be used to generate a consensus sequence for each target molecule. In some embodiments, the consensus sequence may be generated using VCFcons (e.g., VCFcons v8.5.0). VCFcons is an improved version of VCF and is a general-purpose VCF-based consensus sequence generator for genomes (e.g., small genomes). In some embodiments, a predetermined threshold (e.g., 4) for generating a consensus sequence may be obtained. For example, in some embodiments, for VCFcons to call the nucleotide identity of a position in a target molecule, there must be at least four sequence reads covering that position. In some embodiments, the sequence reads are circular consensus sequencing (CCS) reads. If a nucleotide has fewer than four reads, it is reported as "N" (ambiguous, undefined nucleotide) in the consensus sequence. In some embodiments, if at least four sequence reads cover a position in the target molecule and the nucleotide identity at that position is present in at least 50% of the sequence reads aligned to that position, the nucleotide identity at that position is called. In some embodiments, a higher percentage (e.g., 60%) or a lower percentage (e.g., 40%) may be used. Methods for generating consensus sequences for target molecules are described, for example, in U.S. Patent Application No. 17 / 845,629, the entire contents of which are incorporated herein by reference for all purposes.
[0139] The sequencing file may also be used to generate a genome construct for each sample. For example, in some embodiments, when VCFcons calls the nucleotide identity of a position in a genome construct, there must be at least four consensus sequences covering that position. If a nucleotide has fewer than four consensus sequences, it is reported as "N" in the genome construct. In some embodiments, if at least four consensus sequences cover a position in a genome construct and the alternative allele frequency is greater than 50% compared to the reference, the nucleotide identity of that position is called.
[0140] After determining the sequence base composition, the percentage of unambiguous bases may also be determined. In some embodiments, this determination may be based on Seqtk or an alternative algorithm.
[0141] Virus detection and / or lineage assignment are determined using techniques disclosed herein (e.g., techniques disclosed with respect to block 135 of FIG. 1 ). Both virus detection and lineage assignment may be determined using a consensus sequence and / or genome construct and a decision rule. In some embodiments, virus detection is performed using wet-lab means, and lineage assignment is determined using a consensus sequence (or genome construct) and a decision rule. In some embodiments, the decision rule (e.g., the decision rule obtained in block 133 of FIG. 1 ) is determined in advance. In other embodiments, the decision rule is determined based on an algorithm or machine learning model. The machine learning model may be pre-trained using sequences or sequence reads of known lineages of the virus and / or sequences or sequence reads of common virus species. In some embodiments, one or more scores are assigned to the consensus sequence and / or assembled genome relative to a reference genome or viral library of the virus. In some embodiments, a decision rule is used to assign one or more scores to the consensus sequence and / or assembled genome.
[0142] One or more scores may be determined based on the distribution of mutations or nucleotide comparisons (e.g., nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc.) between the consensus sequence (or genome construct) and the reference genome or portion thereof. For example, the distribution may indicate the number of nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc. Different types of mutations or nucleotide comparisons may be assigned different scores. In some embodiments, subtypes of the same type of mutation or nucleotide comparison are assigned different scores. For example, the score assigned to an "A" to "G" substitution may be different from the score assigned to an "A" to "C" substitution. In some embodiments, different scores are assigned based on the phylogenetic context (e.g., evolutionary relationship or lineage) of the mutation or nucleotide comparison. For example, a nucleotide substitution that is common in a particular lineage may receive a different score than one that is rare. In some embodiments, a mutation or nucleotide comparison that leads to a change in an amino acid encoded in a protein-coding region may receive a higher score, especially if the change is known to affect the function of the protein. In some embodiments, different types of mutations or nucleotide comparisons are assigned the same score.
[0143] Once one or more scores are obtained, they may be used to determine the presence or absence of the virus in the sample and / or to determine its predicted lineage. In some embodiments, if one or more scores, or all scores, are below a predetermined number or above a predetermined range (or threshold), the virus in the sample is not assigned an existing lineage. Instead, a new lineage (different from existing or known lineages) is recommended and provided in the report file.
[0144] In some embodiments, NextClade is used for virus detection and lineage assignment. In some embodiments, Pangolin may be used to assign lineages to consensus sequences or genome constructs. In one embodiment, Pangolin is configured to consider only genomes with at least 50% unambiguous bases. In some embodiments, SummaryStat may be used to compile results from NextClade, Pangolin, Seqtk, and other algorithms, platforms, or programs to generate coverage statistics (e.g., median amplicon coverage and mean genome coverage percentage) required for subsequent quality control. The genome coverage percentage may be calculated by dividing the number of unambiguous bases (A, T, C, G) by the total sequence length, and the lineage classification is aggregated. Only samples that produce NextClade results and Pangolin lineage calls are retained for further processing.
[0145] For example, to detect the presence or absence of a virus, the aligned consensus sequence or genome construct is compared nucleotide by nucleotide with a reference viral genome or a portion of a reference viral genome. Differences are then scored and / or reported accordingly. In some embodiments, different types of mismatches are scored differently. For example, a nucleotide substitution (a change from one nucleotide to another) may have a higher or lower score than a nucleotide deletion (i.e., a gap). In some embodiments, different nucleotide substitutions may result in different scores. A threshold may be determined in advance or based on an algorithm or machine learning model to determine the presence or absence of a virus based on the total score or score distribution.
[0146] In some embodiments, the presence of the virus in the sample is known, and only the predicted strain of the virus is determined. In some embodiments, the presence of the virus in the sample is determined using PCR or real-time PCR. PCR technology can be used to detect the virus. Real-time PCR detection can be performed simultaneously with or prior to the dynamic clinical assay pipeline analysis. Based on the real-time PCR, a diagnostic positive / negative result can be provided to the subject. The diagnostic positive / negative result can also be used to determine whether to enter the sample into the dynamic clinical assay pipeline.
[0147] Real-time PCR can be used to detect viruses. The basic principle of real-time PCR is to amplify and detect specific regions of viral nucleic acids using fluorescent probes or fluorescent dyes. Virus detection using real-time PCR generally involves the following steps: (i) extracting viral nucleic acids by isolating them from a sample; (ii) designing primers and probes specific to the virus being detected and targeting conserved regions of the viral genome, enabling amplification and detection of the viral nucleic acids; (iii) preparing a PCR reaction mixture containing primers, probes, and a polymerase enzyme (different fluorescent dyes may be added to the reaction mixture); (iv) amplifying the viral nucleic acids by performing PCR in a thermal cycler that cycles through a series of temperature changes (measurement of fluorescent signals in real time at each cycle allows detection and quantification of the viruses); and (v) analyzing data generated during the PCR reaction using specific software to determine the presence and / or amount of viral nucleic acids in the sample. The amount of viral nucleic acids can be expressed as a cycle threshold (Ct value), which represents the cycle at which the fluorescent signal exceeds a predetermined threshold.
[0148] In some embodiments, to assign lineage, the variant of interest may be predetermined, and the genomic region related to the variant of interest may also be determined based on the variant of interest or based on an algorithm or machine learning model. The genomic region related to the variant of interest may be a single nucleotide polymorphism (SNP), a continuous portion in a reference genome, or multiple discontinuous portions in a reference genome. One or more scores are determined based on the alignment of the genomic region of the consensus sequence or genome construct with the genomic region of the reference genome or the variant of interest. To determine the predicted lineage based on one or more scores or score distributions, a threshold may be predetermined or determined based on an algorithm or machine learning model.
[0149] Figure 11 shows exemplary virus calls and lineage assignments in samples. The top five results (IDs 9, 10, 1, 3, and 4) are exemplary circulating strains (under the a4 synthetic control) compared to published genomes isolated in the United States. The bottom three results (IDs 8, 6, and 14) are genome controls based on previous outbreaks that are no longer circulating. Black vertical bars indicate variants relative to the reference, triangles indicate insertions, and gray boxes indicate gaps in the genome expected in the synthetic control.
[0150] Figures 12A and 12B show phylogenetic analyses of the initial pilot run showing the predicted assignment of patient samples to control and currently circulating strains. Figure 12A shows a phylogenetic tree that includes samples from the initial pilot run. Figure 12B is a blow-up view showing all samples under the B.1 subtree.
[0151] Figures 13A-13C show the analysis of a larger multiplex run, including more control and patient samples. The expected distribution is shown in Figures 13B-13C, which also confirms the multiple introduction of MPX into the human population, as previously reported. For comparison, Gisaid (database) entries for other detected lineages are shown. Figure 13A shows the NextClade lineage assignment results.
[0152] In some embodiments, more efficient and accurate virus detection and lineage assignment results can generally be achieved by dynamically modifying the dynamic clinical assay pipeline using virus detection and lineage assignment information. For example, the dynamic clinical assay pipeline may be used clinically for a certain period of time to collect virus detection and lineage assignment information. Based on the lineage assignment information collected during that period, the prevalence status of the lineage can be determined, and the reference genome or target molecule may be modified based on the prevalent lineage. In some embodiments, the adapters and decision rules may also be modified based on the prevalent lineage. The dynamic clinical assay pipeline may be designed to be automatically modified or updated in real time. The dynamic clinical assay pipeline improves the capture and sequencing of target molecules, contributing to the provision of accurate virus detection and lineage assignment results. VII. Downstream uses
[0153] The systems and methods disclosed herein are applicable to a variety of applications. One of the more important applications is providing a treatment plan or clinical trial protocol to a subject. If a virus is detected in a sample obtained from a subject, a computing system may be used to automatically generate a treatment plan or clinical trial protocol and provide it to the subject or a medical practitioner. The treatment plan or clinical trial protocol may vary based on the detected virus and / or predicted strain. For example, if the virus is an MPX virus or a cowpox virus, the treatment plan may include administering an antiviral drug to the subject.
[0154] Additionally, monitoring the emergence of novel variants and more virulent strains can be important because different variants or strains can elicit different responses or lead to different treatments. For example, some strains of MPX have been associated with deaths in up to 10% of patients, but in recent outbreaks, the number of cases was less than 1%. The disclosed technology's ability to assign lineages to samples allows for tracking the spread and transmission of different strains of virus. The output of the dynamic clinical assay pipeline can be used to estimate the size of an outbreak due to strain mutations, allowing for the development and implementation of plans to control the spread of the virus or strain. The dynamic clinical assay pipeline also enables primer site integrity monitoring. For example, MPX viruses / strains may have dropped regions / locuses, and the dropped regions / locuses may be primer sites. Primer site integrity monitoring provides feedback on the spread and transmission of the virus, allowing for dynamic changes to the dynamic clinical assay pipeline.
[0155] The disclosed techniques may also be used to design novel assays or machines. Novel assays or machines may utilize all or part of the disclosed techniques to perform efficient virus detection and / or predictive lineage assignment. In some embodiments, multiple viruses may be designed to be detected within the novel assay or device. VIII. Example
[0156] The systems and methods implemented in various embodiments will be better understood with reference to the following examples. 1. Example 1: Monkeypox analysis
[0157] In one example, 7826 evenly tiled molecular loop inversion probes were designed and used in the system described herein, which are expected to produce a 675 bp fragment from the 200 kb genome of the MPX virus. These probe pairs are trimmed from both ends of the read using custom code. In one example, 25 bp is trimmed from both the 5' and 3' ends of all sequences produced by the sequencing pipeline.
[0158] Figure 14 shows a comparison of the sequenced synthetic control with independently obtained sequencing data, highlighting the matching regions that are linearly related to mismapping. The synthetic control is a genome of known truth provided by the manufacturer. This is compared to the independently obtained sequencing data. A dot indicates that both sequences have identical bases at that genomic position. As shown in Figure 14, highly repetitive repeats at the 5' and 3' ends of the genome are expected to be double-mapped. 2. Example 2: SARS-CoV-2 analysis
[0159] The Labcorp VirSeq SARS-CoV-2 NGS test may be performed in laboratories designated by Labcorp that are certified under the Clinical Laboratory Improvement Act of 1988 (CLIA) (42 U.S.C. § 263a) and meet the requirements to perform high complexity testing as described in the Labcorp VirSeq SARS-CoV-2 NGS test standard operating procedure, which has been reviewed under FDA Emergency Use Authorization (EUA). Purpose of use
[0160] The Labcorp VirSeq SARS-CoV-2 NGS Test is a next-generation sequencing (NGS) test on the PacBio Sequel II sequencing system intended to identify and differentiate SARS-CoV-2, using the Phylogenetic Assignment of Named Global Outbreak (PANGO) phylogenetic nomenclature, from SARS-CoV-2-positive samples authenticated using Labcorp's COVID-19 RT-PCR Test or the Labcorp SARS-CoV-2 and Influenza A / B Assay, when clinically indicated. Testing is limited to Labcorp-designated laboratories certified under the Clinical Laboratory Improvement Act of 1988 (CLIA) (42 U.S.C. § 263a) and meeting the requirements for performing high-complexity testing. The Labcorp VirSeq SARS-CoV-2 NGS Test is intended for use in conjunction with patient history and other diagnostic information when clinically indicated, i.e., in situations where results may help determine appropriate clinical management. Results of this test are intended to be interpreted by the ordering healthcare professional. This test is not intended to be used as an aid in the primary diagnosis of SARS-CoV-2 infection or to confirm the presence or absence of SARS-CoV-2 infection, nor is it intended to identify specific SARS-CoV-2 genomic mutations. Results should not be used as the sole basis for treatment or other patient management decisions.
[0161] The Labcorp VirSeq SARS-CoV-2 NGS Test is intended for use by qualified clinical laboratory technicians who have received specific instruction and training in the operation of the PacBio Sequel II Sequencing System and next-generation sequencing workflow, as well as in vitro diagnostic procedures. The Labcorp VirSeq SARS-CoV-2 NGS Test is for use only under an Emergency Use Authorization by the U.S. Food and Drug Administration (FDA). Device description and inspection principle
[0162] The Labcorp VirSeq SARS-CoV-2 NGS test is a PacBio Sequel II-based whole-genome sequencing assay used to determine lineage by PANGO from RNA extracted from SARS-CoV-2-positive samples authenticated using Labcorp's COVID-19 RT-PCR test or Labcorp SARS-CoV-2 and Influenza A / B assays. The SARS-CoV-2 probe set used in this assay contains approximately 1,000 tiled molecular loop inversion probes (MIPS) designed to amplify RNA reverse-transcribed to cDNA from 99.6% of the SARS-CoV-2 genome, with most bases covered by 22 MIPS. Products synthesized between the MIPS are enriched and then sequenced to identify a sample-specific molecular barcode.
[0163] Residual total nucleic acid extracts from SARS-CoV-2-positive RT-PCR diagnostic test samples (N1 target cycle threshold (Ct) value less than 31) are tested using the Labcorp VirSeq SARS-CoV-2 NGS test. Residual nucleic acid extracts can be stored at -20°C for up to 30 days, after up to two freeze-thaw cycles. Using a Hamilton Microlab STAR, the residual total nucleic acid extracts are transferred to a 96-well plate containing only the positive RT-PCR diagnostic test samples. The samples are then aliquoted into a 94-sample sequencing run plate, including one water non-template control (NTC) and one positive control. Eight plates, or 752 specimens, are processed in one production batch.
[0164] A custom molecular loop SARS-CoV-2 capture kit is used to prepare samples for sequencing on the PacBio Sequel II instrument. First, cDNA is synthesized from RNA using Thermo Fisher's reverse transcriptase. The SARS-CoV-2 cDNA is then used as a target for hybridization of the molecular loop probe (Figure 15, information is available on the FDA website, "Emergency Use Authorization (EUA) Summary for the Labcorp VirSeq SARS-CoV-2 NGS Test" (downloaded March 22, 2023)). The molecular loop probe consists of two binding sites tiled across 99.6% of the 30 kb SARS-CoV-2 viral genome, separated by 600 bp. After binding, the 600 bp region between the two probes is synthesized with DNA polymerase and ligated to form a closed molecule (Figure 15, Step 2). Unbound or incomplete loops remain linear molecules and are removed by exonuclease digestion (Figure 15, Step 3). Next, circular molecules are released from the template cDNA (Figure 15, step 4) and enriched by amplification using a 3' molecular loop-specific M13 universal sequence and a 5' sample-specific barcode (Figure 15, step 5). Samples are then pooled in equal amounts, and libraries are prepared for PacBio sequencing. Library preparation involves DNA damage repair, ligation of sequencing adapters, removal of unligated products by enzymatic digestion, and bead purification. The libraries are then sequenced on a PacBio Sequel II with a 15-hour video recording (Figure 15, step 6).
[0165] After sequencing, PacBio SMRT LINK software and a custom molecular loop processing script were used to generate FASTQ files for each sample. The FASTQs were analyzed using a genome analysis pipeline implemented in CLC Genomics Server version 9.1.1. This workflow starts with sample-level FASTQ files, trims primers, and aligns them to the SARS-CoV-2 reference genome (NC_045512v2) using Minimap2 to generate alignment BAM files. VCFcons (v8.5.0) was used to generate consensus sequences for each sample. When VCFcons calls a nucleotide sequence for genome assembly, there must be at least four circular consensus sequencing (CCS) reads covering that base pair, and the alternative allele frequency must be greater than 50% compared to the reference. If a nucleotide is covered by fewer than four reads, it is reported as ambiguous (N) in the consensus sequence. The consensus sequence is used as input to the PANGOLIN (v3.1.20) analysis package to assign lineages to individual samples. Publish phylogenetic results for samples with genome coverage of at least 90% and overall genome coverage greater than 10 CCS reads. Overall genome coverage of a sample is defined as the average of the median read coverage across 29 contiguous regions (approximately 1 kb in length) spanning the entire viral genome. Equipment used for testing
[0166] The Labcorp VirSeq SARS-CoV-2 NGS test is used in conjunction with a Pacific Bioscience Sequel II sequencing instrument, a Mantis liquid handler, and the Labcorp VIRSEQ analysis pipeline for sequence analysis and lineage determination. The equipment and reagents required to perform the Labcorp VirSeq SARS-CoV-2 NGS test are listed in Table 1. Table 1. Reagent and Special Equipment Requirements [Table 1-1] [Table 1-2]
[0167] Designated laboratories will receive an FDA-approved equipment qualification protocol included as part of the Labcorp VirSeq SARS-CoV-2 NGS testing standard operating procedure (SOP) and will be instructed to implement the protocol prior to testing clinical samples. Designated laboratories must comply with the approved SOP, including the equipment qualification protocol, in accordance with the letter of mandate. Controls used in testing
[0168] External Positive Control: Add an external positive control to each plate containing 94 samples. The control consists of one of 10 Twist synthetic SARS-CoV-2 RNA controls with a given PANGO strain designation. Compare the strain designation of the external positive control to its known PANGO strain designation.
[0169] External Negative Control: An external non-template control (NTC) is required to ensure the absence of master mix contamination events on a given amplification plate. The control consists of molecular-grade water added to position A1 of each 96-well plate prior to sample addition. The NTC is then transferred to the sequencing run plate along with the positive samples and subjected to sequencing and quality control (QC) analysis.
[0170] Other controls: Publish lineage identification results only for samples with genome coverage ≥ 90% and overall genome coverage > 10 CCS reads. Interpretation of results
[0171] All laboratory controls should be reviewed before interpreting patient results. Patient results should not be interpreted if controls are not valid. 1) Labcorp VirSeq SARS-CoV-2 NGS Test Control - Positive:
[0172] After sequencing, the strain assignment of the positive control on a given plate is compared to the known strain assignment of the positive control. A positive control sample is considered to have failed if the strain determined by the assay differs from its known strain. Samples on plates where the external positive control failed are reanalyzed once more in the assay using material from the original total nucleic acid extract. If the residual total nucleic acid in the sample is depleted, it is reported as a failed sample. 2) Labcorp VirSeq SARS-CoV-2 NGS Test Control - NTC:
[0173] After sequencing, calculate the overall genome coverage of the NTC sample. If the overall genome coverage exceeds 10 CCS reads, the NTC sample is considered a failure. Samples on plates that fail the external negative control are reanalyzed once more in the assay using the original total nucleic acid extract. If the residual total nucleic acid in the sample is depleted, report it as a failed sample. Table 2. Expected results of external controls for the Labcorp VirSeq SARS-CoV-2 NGS test [Table 2] 3) Review and interpretation of patient sample results:
[0174] Perform external control analysis and grade Labcorp VirSeq SARS-CoV-2 NGS test results after removing any samples on plates that fail the external control. Calculate overall genome coverage and percent genome coverage for each patient sample. Publish phylogenetic results for samples with at least 90% genome coverage and overall genome coverage greater than 10 CCS reads. Interpretation and reporting of clinical specimens is summarized in Table 3. Table 3. Interpretation of results for patient samples [Table 3] Performance evaluation 1.Device Tolerance:
[0175] Residual nucleic acid extracts from 8,815 respiratory specimens that tested positive for SARS-CoV-2 by EUA200011-authorized Labcorp COVID-19 tests across three sites were sequenced with the Labcorp VirSeq SARS-CoV-2 NGS test. The number of samples that produced genome sequences that passed both of the following quality control criteria was determined for different N1 CT value categories: - Genome coverage is greater than 90% when compared to the SARS-CoV-2 reference genome ("NC_045512v2"). ●Overall genome coverage exceeds 10 CCS reads. Table 4. Results of tolerance studies [Table 4] 2. Precision (reproducibility): Within the assay
[0176] Intra-assay reproducibility was assessed by testing 11 nucleic acid samples, each in triplicate. The intra-assay reproducibility study assessed the assay's ability to accurately detect strains for replicates of the same sample within a single assay run. The N1 CT values of the samples tested ranged from 19.2 to 22.96.
[0177] Lineage designations for all 11 samples were concordant across three replicates each, and all replicates met QC criteria. 3. Precision (reproducibility): Between assays
[0178] Inter-assay reproducibility was assessed by testing the same nucleic acid samples assessed in the reproducibility study in three replicates across three assay runs. Of the 11 samples tested in the reproducibility study, one was unintentionally excluded from the final run. Six different technicians performed sequencing runs using three different lots of SMRT cells and two lots of sequencing reagents. The third run was performed three weeks after the first run and was multiplexed with approximately one-quarter the sample concentration of the previous two runs. Table 5. Reproducibility study description [Table 5]
[0179] The lineage assignments of 9 of the 10 samples were concordant across all three runs. The discordant samples did not meet the overall genome coverage QC criteria in runs 2 and 3. 4. Sample Stability (Freeze-Thaw)
[0180] The stability of patient samples under recommended storage conditions was assessed in a sample stability study. Following nucleic acid amplification (NAA) diagnostic testing using the Labcorp COVID-19 RT-PCR test, extracted nucleic acids were shipped to the laboratory on dry ice and stored at -20°C prior to sequencing.
[0181] Twelve samples containing α, β, and δ variants of concern (VOCs) were initially sequenced. These samples were then sequenced again after 5 weeks of storage at -20°C with three freeze-thaw cycles. Eleven of the 12 samples yielded concordant results between the initial and retest tests. The only discordant results were determined to be due to mechanical error and not related to sample stability. The results of this study confirm that the Labcorp VirSeq SARS-CoV-2 NGS test is stable for up to two freeze-thaw cycles when samples are stored at 20°C for up to 30 days. 5. Match rate
[0182] Lineage assignment results generated by the Labcorp VirSeq SARS-CoV-2 NGS test (PacBio molecular loop sequencing) were directly compared with lineage assignment results generated by Illumina COVIDSeq (RUO surveillance protocol using v3 primer pool) and PacBio amplicon sequencing. The following samples were included in the analysis: Illumina COVIDSeq (RUO) samples ○ 93 negative (NAA) samples 72 SARS-CoV-2 samples sequenced in winter 2020 29 samples sequenced at the Center for Molecular Biology and Pathology (CMBP) 50 samples sequenced through DNA testing PacBio amplicon sequenced samples 122 SARS-CoV-2 samples ampli-sequenced at 90% coverage Illumina COVIDSeq (RUO) comparison:
[0183] Ninety-three samples that tested negative using the Labcorp COVID-19 RT-PCR diagnostic test were sequenced in duplicate, one using the Labcorp VirSeq SARS-CoV-2 NGS test and one using the Illumina COVIDSeq (RUO). Of the 93 samples tested, one sample had a reportable SARS-CoV-2 genome using only the Illumina COVIDSeq (RUO), and one sample had a reportable SARS-CoV-2 genome using both the Illumina COVIDSeq (RUO) and the Labcorp VirSeq SARS-CoV-2 NGS test. Subsequent investigation revealed that both samples tested positive for SARS-CoV-2 using the Labcorp COVID-19 RT-PCR test and were erroneously included in the validation. For the 91 true-negative samples, the final concordance between the Illumina COVIDSeq (RUO) and the Labcorp VirSeq SARS-CoV-2 NGS test was 100%.
[0184] Of 72 samples sequenced using Illumina COVIDSeq (RUO) in winter 2020, 51 samples had reportable results when using the Labcorp VirSeq SARS-CoV-2 NGS test. Of the 51 reportable results, all lineage assignments were 100% concordant with the Illumina COVIDSeq (RUO) output. Table 6. Comparison of PANGO strain assignments between Illumina COVIDSeq (RUO) and Labcorp VirSeq SARS-CoV-2 NGS tests for winter 2020 circulating strains [Table 6]
[0185] Seven of the 79 samples were lost in transit or did not produce valid results on Illumina COVIDSeq (RUO). Seventy-two of the 79 samples collected for CMBP and DNA testing were successfully sequenced on both Illumina COVIDSeq (RUO) and the Labcorp VirSeq SARS-CoV-2 NGS test. Of these 72 samples, 66 samples that met the sequencing QC criteria for the Labcorp VirSeq SARS-CoV-2 NGS test all showed lineage assignment results consistent with Illumina COVIDSeq (RUO). Table 7. Comparison of PANGO lineage designations between Illumina COVIDSeq (RUO) and Labcorp VirSeq SARS-CoV-2 NGS tests in summer 2021 circulation lines [Table 7] Comparison of PacBio amplicon-based sequencing:
[0186] One hundred and twenty-two samples originally sequenced using the PacBio amplicon-based method were reprocessed in two replicates using the Labcorp VirSeq SARS-CoV-2 NGS test. The 122 samples with >90% coverage by amplicon sequencing were then reprocessed in two replicates using the Labcorp VirSeq SARS-CoV-2 NGS test, yielding a total of 244 results. Of the 244 lineage assignments using the Labcorp VirSeq SARS-CoV-2 NGS test, 234 results yielded genomes that met both QC criteria. Of these 234 results, 225 results were consistent with the lineage assignments made by the original PacBio amplicon assay. When compared with lineage assignments made using PacBio amplicon sequencing, 96% of samples yielded concordant lineage assignments. Table 8. Comparison of PANGO lineage assignments between PacBio amplicon sequencing and the Labcorp VirSeq SARS-CoV-2 NGS test [Table 8] 6. Reference sample testing
[0187] In this evaluation, heat-inactivated SARS-CoV-2 samples from the B.1.1.7 (VR-3326HK), Hong Kong / VM20001061, and Italian INMI1 lineages characterized by ATCC were used. Sequencing errors associated with the Labcorp VirSeq SARS-CoV-2 NGS test were assessed by comparing all mutations identified in the consensus sequence produced by the Labcorp VirSeq SARS-CoV-2 NGS test's analytical pipeline with the published ATCC reference sequence. The results are shown in the table below. Table 9. Number of nucleotide mismatches between the ATCC reference sequence and the sequence generated by the Labcorp VirSeq SARS-CoV-2 NGS test [Table 9] *b1117, B.1.1.7(VR-3326HK), HK: Hong Kong / VM20001061, ITLY: Italy INMI1
[0188] Overall, we observed an average of 0.012% sequence difference between the reference sequence and the consensus sequence produced by the Labcorp VirSeq SARS-CoV-2 NGS test. 7. Simulation Studies
[0189] A simulation study was conducted to assess the performance of the Labcorp VirSeq SARS-CoV-2 NGS test to identify samples of the PANGO lineage that were not tested in the concordance and analysis studies.
[0190] A sequencing error model simulating how the Labcorp VirSeq SARS-CoV-2 NGS test randomly introduces sequencing errors and ambiguous nucleotides into sequenced genomes was estimated based on the sequencing results of 760 clinical samples. The model estimated that, on average, the Labcorp VirSeq SARS-CoV-2 NGS test introduced one sequencing error per 33 sequenced SARS-CoV-2 genomes and 3,369 ambiguous nucleotides per sequenced SARS-CoV-2 genome. Simulation of variants of concern / variants of interest
[0191] A total of 23,400 reference sequences with known Pango lineage assignments representing 234 lineages (100 reference sequences per lineage) were downloaded from GISAID. Using a sequencing error model, sequencing errors were introduced into these 23,400 reference sequences to simulate hypothetical sequence output of the assay for these genomes. Each of these 23,400 sequences was used to generate multiple simulated sequences. The PANGO lineages of these simulated sequences were identified using the lineage identification software (PANGOLIN v3.1.20) used in the Labcorp VirSeq SARS-CoV-2 NGS test. The lineage identification results for each simulated sequence were compared to the known PANGO lineage assignment of the reference sequence used to generate the simulated sequence. The concordance results for simulated sequences with genome coverage of 90%, 95%, and 99% are shown in the table below. Table 10. In silico performance of the Labcorp VirSeq SARS-CoV-2 NGS test against 234 VOC / VOI lineages [Table 10]
[0192] We correctly identified 100% (95% CI: 99.55%–100.00%) of the 770 simulated omicron sequences at the sublineage level (BA.1, BA.2, and BA.3). Period Simulation
[0193] A total of 10,000 high-quality SARS-CoV-2 sequences were randomly sampled from GISAID. A total of 140,000 reference sequences with known PANGO lineage assignments were downloaded from GISAID. Using a sequencing error model, sequencing errors were introduced into these 140,000 reference sequences to simulate hypothetical sequence output of the assay for these genomes. Each of the 140,000 sequences was used to generate multiple simulated sequences. The PANGO lineages of these simulated sequences were identified using the lineage identification software (PANGOLIN v3.1.20) used in the Labcorp VirSeq SARS-CoV-2 NGS test. The simulated lineage identification results were compared with the known PANGO lineage assignments of the reference sequences used to generate the simulated sequences. The concordance results for simulated reads with genome coverage of 90%, 95%, and 99% are shown in the table below. Table 11. In silico performance of the Labcorp VirSeq SARS-CoV-2 NGS test on 140,000 sequences submitted to GISAID [Table 11]
[0194] All 43,328 simulated Omicron sequences were correctly identified as Omicron sequences. The subphylogenetic identity of the simulated Omicron sequences is shown in the table below. Table 12. In silico performance of the Labcorp VirSeq SARS-CoV-2 NGS test on 43,328 simulated Omicron sequences [Table 12] IX. Computing Systems and Additional Considerations
[0195] 16 illustrates an exemplary computing device 1600 suitable for use in systems and methods for determining the presence or absence of a virus and its predicted lineage in a sample obtained from a subject according to the present disclosure. The exemplary computing device 1600 includes a processor 1605 that communicates with memory 1610 and other components of the computing device 1600 using one or more communication buses 1615. The processor 1605 is configured to execute processor-executable instructions stored in memory 1610 to implement one or more methods for determining the presence or absence of a virus and its predicted lineage in a sample obtained from a subject, according to different examples such as some or all of the exemplary system 100, process 200, or process 300 described above with respect to FIGS. 1-3.
[0196] In this example, computing device 1600 also includes one or more user input devices 1630, such as a keyboard, mouse, touchscreen, microphone, etc., that receive user input. Computing device 1600 also includes a display 1635 that provides visual output, such as a user interface, to a user. Computing device 1600 also includes a communication interface 1640. In some examples, communication interface 1640 may enable communication with one or more networks, including a local area network (“LAN”), a wide area network (“WAN”) such as the Internet, a metropolitan area network (“MAN”), point-to-point or peer-to-peer connections, etc. Communication with other devices may be achieved using any suitable network protocol. For example, one suitable network protocol may include Internet Protocol (“IP”), Transmission Control Protocol (“TCP”), User Datagram Protocol (“UDP”), or a combination thereof, such as TCP / IP or UDP / IP.
[0197] In the above description, specific details are set forth to provide a thorough understanding of the embodiments. However, it is understood that the embodiments may be practiced without these specific details. For example, circuits may be shown in block diagrams in order not to obscure the embodiments in unnecessary detail. Furthermore, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order not to obscure the embodiments.
[0198] The implementation of the above-described techniques, blocks, steps, and means may be performed in various ways. For example, these techniques, blocks, steps, and means may be implemented in hardware, software, or a combination thereof. In the case of a hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processors (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the above-described functions, and / or combinations thereof.
[0199] Also, it should be noted that the embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart may describe operations as a sequential process, many operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to a return of the function to the calling function or the main function.
[0200] Furthermore, embodiments may be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting languages, and / or microcode, program code or code segments to perform the necessary tasks may be stored in a machine-readable medium such as a storage medium. Code segments or machine-executable instructions may represent procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, scripts, classes, or any combination of instructions, data structures, and / or program statements. Code segments may be coupled to other code segments or hardware circuits by receiving and / or passing information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, ticket passing, network transmission, etc.
[0201] For a firmware and / or software implementation, these methodologies may be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. Any machine-readable medium tangibly embodying instructions may be used for implementation of the methodologies described herein. For example, software code may be stored in memory. The memory may be implemented within the processor or external to the processor. The term "memory," as used herein, refers to any type of storage medium, whether long-term, short-term, volatile, non-volatile, or other, and is not limited to a particular type of memory, number of memories, or type of medium storing the memory.
[0202] Additionally, as disclosed herein, the terms "storage medium," "storage," or "memory" can refer to one or more memories for storing data, examples of which include read-only memory (ROM), random-access memory (RAM), magnetic RAM, core memory, magnetic disk storage media, optical storage media, flash memory devices, and / or other machine-readable media for storing information. The term "machine-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage media capable of storing instructions and / or data that contain or carry the instructions and / or data.
[0203] While the principles of the present disclosure have been described above in connection with specific apparatus and methods, it is to be clearly understood that this description is made only by way of example and not as a limitation on the scope of the disclosure.
Claims
1. obtaining nucleic acid from a sample obtained from a subject; capturing a target molecule in the nucleic acid using a molecular inversion probe under hybridization conditions; amplifying the target molecule using polymerase chain reaction (PCR) to obtain a plurality of amplified molecules; for each of the plurality of amplified molecules, Ligating an adaptor to each end of the molecule to create a circular molecule; sequencing the circular molecules to obtain sequence reads; generating, using a computing system, a sequencing file comprising the sequence read for each molecule of the plurality of amplified molecules and a position of each sequence read in a reference genome of the virus by aligning the sequence reads to the reference genome of the virus; generating a report file regarding the subject using the computing system and the sequencing file, the report file including a predicted lineage of the virus in the sample, the generating step comprising: generating a consensus sequence for the target molecule based on sequence reads for each of the plurality of amplified molecules, wherein if at least a predetermined number of the sequence reads have a nucleotide identity to a position in the consensus sequence, the nucleotide identity is assigned to the position, and if less than a predetermined number of the sequence reads have the nucleotide identity to the position in the consensus sequence, "N" is assigned to the position; determining one or more scores for the consensus sequence based on the reference genome of the virus or a library of the virus, wherein the one or more scores are determined based on a distribution of mutations in the virus; determining the predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence. Including, a method.
2. The method of claim 1 , wherein the report file further includes the presence or absence of the virus in the sample.
3. 3. The method of claim 2, wherein the presence or absence of the virus in the sample is determined using real-time PCR (RT-PCR) or based on the one or more scores of the consensus sequences.
4. 10. The method of claim 1, wherein the subject tests positive for the virus and the sample contains viral nucleic acid.
5. 2. The method of claim 1, wherein the virus is monkeypox (MPX) virus.
6. The method of claim 1 , wherein the target molecule is a double-stranded DNA molecule.
7. The method of claim 1, wherein the molecular inversion probe consists of two binding sites separated by about 600-700 bp.
8. The method of claim 1 , wherein the sequencing file is a binary alignment map (BAM) file.
9. 10. The method of claim 1, wherein the report file is a variant call format (VCF) file.
10. 10. The method of claim 1, further comprising providing the subject with a treatment plan or clinical trial protocol.
11. 11. The method of claim 10, wherein the treatment plan comprises administering an antiviral drug to the subject.
12. obtaining a prevalent lineage of the virus in a subject population, the prevalence lineage being determined based on the predicted lineage of the virus in the sample; updating the molecular inversion probe to capture the target molecule specific to the epidemic strain; updating the adapter based on the updated molecular inversion probe; and obtaining a decision rule set specific to determining the epidemic lineage, wherein the determining of the one or more scores of the consensus sequence is performed using the decision rule set.
13. obtaining nucleic acid from a sample obtained from a subject; capturing a plurality of target molecules in the nucleic acid using a plurality of molecular inversion probes under hybridization conditions; amplifying each target molecule using polymerase chain reaction (PCR) to obtain a plurality of sets of amplified molecules, each corresponding to a target molecule; For each molecule in each set of amplified molecules: Ligating an adaptor to each end of the molecule to create a circular molecule; sequencing the circular molecules to obtain sequence reads; generating, using a computing system, a sequencing file including the sequence reads for each molecule in each set of amplified molecules and the position of each sequence read in a reference genome of the virus by aligning the sequence reads to the reference genome of the virus; generating a report file regarding the subject using the computing system and the sequencing file, the report file including a predicted lineage of the virus in the sample, the generating step comprising: generating a consensus sequence for each target molecule based on sequence reads of each molecule in the set of amplified molecules corresponding to the target molecule, wherein if at least a first predetermined number of the sequence reads have a nucleotide identity to a position in the consensus sequence, the nucleotide identity is assigned to the position, and if less than the first predetermined number of the sequence reads have the nucleotide identity to the position in the consensus sequence, "N" is assigned to the position; generating a genomic construct of the nucleic acids in the sample based on the consensus sequence of the target molecule, wherein if at least a second predetermined number of the consensus sequences have a nucleotide identity to a position within the genomic construct, and the nucleotide identity at the position is present in at least 50% of the sequence reads aligned to the position, the nucleotide identity is assigned to the position, and if less than the second predetermined number of the consensus sequences have the nucleotide identity to a position within the genomic construct, the position is assigned "N"; determining one or more scores for the consensus sequence based on the reference genome of the virus or a library of the virus; determining the predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence. Including, a method.
14. The method of claim 13 , wherein the report file further includes the presence or absence of the virus in the sample.
15. 15. The method of claim 14, wherein the presence or absence of the virus in the sample is determined using real-time PCR (RT-PCR) or based on the one or more scores of the consensus sequences.
16. 14. The method of claim 13, wherein the subject tests positive for the virus and the sample contains viral nucleic acid.
17. 14. The method of claim 13, wherein the virus is monkeypox (MPX) virus.
18. 14. The method of claim 13, wherein the target molecule is a double-stranded DNA molecule.
19. The method of claim 13, wherein the molecular inversion probe consists of two binding sites separated by about 600-700 bp.
20. 14. The method of claim 13, wherein the sequencing file is a binary alignment map (BAM) file.
21. 14. The method of claim 13, wherein the report file is a variant call format (VCF) file.
22. 14. The method of claim 13, further comprising providing the subject with a treatment plan or clinical trial protocol.
23. 23. The method of claim 22, wherein the treatment plan comprises administering an antiviral drug to the subject.
24. obtaining a prevalent lineage of the virus in a subject population, the prevalence lineage being determined based on the predicted lineage of the virus in the sample; updating the plurality of molecular inversion probes to capture mutations specific to the epidemic strain; updating the adapters based on the updated plurality of molecular inversion probes; and obtaining a decision rule set specific to determining the epidemic lineage, wherein the determining of the one or more scores of the consensus sequence is performed using the decision rule set.
25. one or more data processors; a non-transitory computer-readable medium storing instructions that, when executed on said one or more data processors, cause said one or more data processors to perform any of the methods described in claims 1-24.
26. A computer program product tangibly embodied in a non-transitory machine-readable medium comprising instructions configured to cause one or more data processors to perform any of the methods according to claims 1 to 24.