Dynamic clinical assay pipeline for detecting viruses
By obtaining nucleic acids from the samples, using molecular inverted probes and PCR technology combined with sequencing technology, the virus lineage is accurately detected and allocated, solving the problem of inaccurate virus detection in the prior art, and achieving efficient and accurate virus detection and lineage allocation.
Patent Information
- Application Number
- CN202380076428.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-19
- Filing Date
- 2023-11-01
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to accurately detect viruses and their lineages in samples, especially when the virus mutations are rapid, resulting in inaccurate detection results or inability to effectively distinguish different viruses.
By obtaining nucleic acids from the sample, capturing the target molecule using a molecular inverted probe, PCR amplification, ligating adaptors to form a circular molecule, and sequencing, a sequencing file is generated to determine the predicted lineage of the virus.
Accurate detection and lineage allocation of viruses are achieved, the sensitivity and specificity of detection are improved, the pipeline can be dynamically adjusted to adapt to the epidemic situation of viruses, and the detection efficiency and accuracy are improved.
Smart Images

Figure CN120153087A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the clinical testing of viruses and, in particular, to a dynamic clinical assay pipeline for detecting viruses (e.g., orthopoxviruses such as monkeypox virus) in a sample and / or assigning a lineage to the virus. Background Art
[0002] A virus is a microscopic infectious agent that can replicate itself within the living cells of an organism. While some viruses are harmless to humans, various viruses can cause a variety of diseases and afflictions, including diseases with life-threatening effects on animals, plants, microorganisms, and humans. Understanding the nature of a virus, particularly its genetic material, and detecting the presence of the virus are crucial for developing effective treatments and clinical protocols and for preventing the spread of viral infections.
[0003] Traditional methods for virus detection include diagnostic tests. However, the lack of specific and sensitive diagnostic tests can also hinder virus detection. Many viral infections have similar symptoms, and it is difficult to distinguish between different viruses based solely on symptoms. Therefore, diagnostic tests with high specificity and sensitivity are crucial for accurately detecting viruses, but developing such tests can be a time-consuming and complex process.
[0004] The replication of a virus depends on its genetic material. With the development of laboratory techniques, computational techniques, and sequencing techniques, machines including sequencers and associated software have been invented and manufactured to identify the presence of a virus in a biological sample such as blood, saliva, or tissue based on the genetic material in the sample. Different methods, such as polymerase chain reaction (PCR) and next-generation sequencing (NGS), have also been developed to facilitate virus detection. However, detecting the nucleic acid of a virus in a sample can still be challenging, partly because viruses mutate rapidly. The mutations of a virus can give rise to new lineages that elicit different immune responses in their host organisms and may require different treatments and protocols to prevent their spread or infection. However, there is currently a lack of reliable methods or systems for detecting a virus in a sample and assigning a lineage to the virus in the sample. Summary of the Invention
[0005] In various embodiments, a method is provided that includes: obtaining nucleic acid from a sample obtained from a subject; using molecular inversion probes to capture target molecules in the nucleic acid under hybridization conditions; using polymerase chain reaction (PCR) to amplify the target molecules to obtain a plurality of amplified molecules; for each of the plurality of amplified molecules, ligating adapters to each end of the molecule to produce circular molecules; and sequencing the circular molecules to obtain sequence reads; using a computing system to generate a sequencing file that includes the sequence reads for each of the plurality of amplified molecules and the position of each sequence read in a reference genome of a virus by aligning the sequence reads with the reference genome of the virus; and using the computing system and the sequencing file to generate a report file for the subject, wherein the report file includes a predicted lineage of the virus in the sample, wherein the generating includes: generating a consensus sequence of the target molecules based on the sequence reads for each of the plurality of amplified molecules, wherein if at least a predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, the nucleotide identity is assigned to the position, and wherein if less than a predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, "N" is assigned to the position; determining one or more scores for the consensus sequence based on the reference genome of the virus or a library of the virus, wherein the one or more scores are determined based on the distribution of mutations of the virus; and determining the predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence.
[0006] In some embodiments, the report file further includes the presence or absence of the virus in the sample.
[0007] In some embodiments, the presence or absence of the virus in the sample is determined using real-time PCR (RT-PCR) or based on the one or more scores of the consensus sequence.
[0008] In some embodiments, the subject tests positive for the virus and the sample contains viral nucleic acid.
[0009] In some embodiments, the virus is a monkeypox (MPX) virus.
[0010] In some embodiments, the target molecule is a double-stranded DNA molecule.
[0011] In some embodiments, the molecular inversion probe consists of two binding sites that are approximately 600-700 bp apart.
[0012] In some embodiments, the sequencing file is a Binary Alignment / Map (BAM) file.
[0013] In some embodiments, the report file is a Variant Call Format (VCF) file.
[0014] In some embodiments, the method further comprises providing a treatment plan or a clinical test protocol to the subject.
[0015] In some embodiments, the treatment plan comprises administering an antiviral agent to the subject.
[0016] In some embodiments, the method further comprises: obtaining the epidemic lineage of the virus in a population of subjects, wherein the epidemic lineage is determined based on the predicted lineage of the virus in the sample; updating the molecular inversion probe to capture the target molecule, wherein the target molecule is specific to the epidemic lineage; updating the adaptor based on the updated molecular inversion probe; and obtaining a set of decision rules specific to determining the epidemic lineage, wherein the one or more scores for determining the consensus sequence are determined using the set of decision rules.
[0017] In various embodiments, a method is provided that includes: obtaining nucleic acids from a sample obtained from a subject; using a plurality of molecular inversion probes to capture a plurality of target molecules in the nucleic acids under hybridization conditions; using polymerase chain reaction (PCR) to amplify each target molecule to obtain multiple sets of amplified molecules, where each set of amplified molecules corresponds to a target molecule; for each molecule in each set of amplified molecules, ligating adapters to each end of the molecule to produce circular molecules; and sequencing the circular molecules to obtain sequence reads; using a computing system to generate a sequencing file that includes the sequence reads of each molecule in each set of amplified molecules and the position of each sequence read in a reference genome of a virus obtained by aligning the sequence reads with the reference genome of the virus; and using the computing system and the sequencing file to generate a report file for the subject, where the report file includes a predicted lineage of the virus in the sample, where the generation includes: generating a consensus sequence for each target molecule based on the sequence reads of each molecule in the set of amplified molecules corresponding to the target molecule, where if at least a first predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, the nucleotide identity is assigned to the position, and where if less than the first predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, "N" is assigned to the position; generating a genomic construct of the nucleic acids in the sample based on the consensus sequence of the target molecules, where if at least a second predetermined number of the consensus sequences have nucleotide identity at a position in the genomic construct and the nucleotide identity at the position is present in at least 50% of the sequence reads aligned to the position, the nucleotide identity is assigned to the position, and where if less than the second predetermined number of the consensus sequences have nucleotide identity at a position in the genomic construct, "N" is assigned to the position; determining one or more scores for the consensus sequence based on the reference genome of the virus or a library of the virus; and determining the predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence.
[0018] In some embodiments, the report file further includes the presence or absence of the virus in the sample.
[0019] In some embodiments, the presence or absence of the virus in the sample is determined using real-time PCR (RT-PCR) or based on the one or more scores of the consensus sequence.
[0020] In some embodiments, the subject tests positive for the virus, and the sample contains viral nucleic acid.
[0021] In some embodiments, the virus is the monkeypox (MPX) virus.
[0022] In some embodiments, the target molecule is a double-stranded DNA molecule.
[0023] In some embodiments, the molecular inversion probe consists of two binding sites that are approximately 600-700 bp apart.
[0024] In some embodiments, the sequencing file is a Binary Alignment / Map (BAM) file.
[0025] In some embodiments, the report file is a Variant Call Format (VCF) file.
[0026] In some embodiments, the method further comprises providing a treatment plan or a clinical test protocol to the subject.
[0027] In some embodiments, the treatment plan comprises administering an antiviral drug to the subject.
[0028] In some embodiments, the method further comprises: obtaining the epidemic lineage of the virus for a population of subjects, wherein the epidemic lineage is determined based on the predicted lineage of the virus in the sample; updating the plurality of molecular inversion probes to capture mutations specific to the epidemic lineage; updating the adaptor based on the updated plurality of molecular inversion probes; and obtaining a set of decision rules specific to determining the epidemic lineage, wherein the one or more scores for determining the consensus sequence are determined using the set of decision rules.
[0029] In some embodiments, a system is provided that includes: one or more data processors and a non-transitory computer-readable medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods or processes disclosed herein.
[0030] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0031] The terms and expressions employed are used in a descriptive rather than a restrictive sense, and are not intended to exclude any equivalents or portions of the features shown and described. However, it should be recognized that various modifications are possible within the scope of the claimed technology. Accordingly, it is to be understood that although the inventive techniques have been expressly disclosed by way of examples and optional features, those skilled in the art may make modifications and variations to the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The present invention will be better understood with reference to the following non - restrictive drawings, in which:
[0033] Figure 1 is a block diagram of a system capable of performing virus detection and lineage assignment techniques according to various embodiments.
[0034] Figure 2 is an exemplary flowchart showing a process for generating a report file regarding the presence or absence of a virus in a sample according to various embodiments.
[0035] Figure 3 is an exemplary flowchart showing a process for performing the disclosed techniques regarding the presence or absence of a virus in a sample according to various embodiments.
[0036] Figure 4 shows a portion of a molecular circular probe set design and PCR pipeline for capturing target DNA molecules according to various embodiments.
[0037] Figure 5 shows a portion of an adapter design, ligation, and sequencing pipeline for generating circular DNA molecules according to various embodiments.
[0038] Figure 6 shows an overview of the complexity of the MPX genome according to various embodiments.
[0039] Figure 7 is an Integrative Genomics Viewer (IGV) plot showing genome coverage according to various embodiments.
[0040] Figure 8 is according to various embodiments Figure 7 a higher - resolution graphical representation of an IGV plot generated using software developed for various clinical samples as shown.
[0041] Figure 9 shows alternative ways of demonstrating genome coverage and read depth for various samples and controls according to various embodiments.
[0042] Figure 10 Shows genomic coverage and per-base coverage according to various embodiments.
[0043] Figure 11 Shows exemplary virus calls and lineage assignments in samples according to various embodiments.
[0044] Figure 12A and 12B Shows lineage analysis of an initial pilot run according to various embodiments, showing the expected assignments of control and circulating strain patient samples.
[0045] Figures 13A - 13C Shows analysis of more multiplexed runs with more control and patient samples according to various embodiments.
[0046] Figure 14 Shows a comparison between sequencing-by-synthesis controls and publicly available data according to various embodiments.
[0047] Figure 15 Shows the molecular circular sequencing process according to various embodiments.
[0048] Figure 16 Shows an example computing device suitable for use with systems and methods for determining the presence or absence of a virus in a sample obtained from a subject and for determining its predicted lineage according to various embodiments.
[0049] In the drawings, similar components and / or features may have the same reference numerals. Further, various components of the same type may be distinguished by following the reference numeral with a dash and a second numeral that differentiates among the similar components. If only the first reference numeral is used in the specification, the description applies to any of the similar components having the same first reference numeral regardless of the second reference numeral. Detailed Description
[0050] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of the preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing the various embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.
[0051] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0052] In addition, note that a single embodiment may be described as a process, which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart or diagram may describe the operations as a sequential process, many operations may be performed in parallel or simultaneously. Additionally, the order of the operations may be rearranged. When the operations of a process are completed, it terminates, but it may have additional steps not included in the figure. A process may correspond to a method, a function, a program, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0053] I. Introduction
[0054] Viruses can be life-threatening and there is a need to be able to accurately detect viruses from biological samples, identify different viral strains, and assign lineages. Since the replication of a virus depends on its genetic material, such as its nucleic acid, various methods have been developed to detect viruses in a sample based on the nucleic acid in the sample. The detection and confirmation of the nucleic acid of a virus in a biological sample can be challenging. One of the biggest reasons is that viruses mutate rapidly, with an average mutation rate of 10 -4 to 10 -8 mutations per replication site. Such mutations alter the genetic material of the virus, change its surface proteins and other important components, and affect the immune response of the host organism as well as the results of diagnostic tests.
[0055] Another challenge associated with virus detection is the presence of long stretches of repetitive DNA or RNA. Sequencing is generally involved in virus detection for accurate identification of different viral strains, monitoring primer sites, and assigning lineages. However, in the absence of accurate sequencing depth and coverage, repetitive sequences in the nucleic acid of a virus may be difficult to distinguish from the host nucleic acid sample, mask the presence of mutations or variants, and result in false negative results.
[0056] To address these and other challenges, the techniques described herein relate to methods and systems for constructing a dynamic clinical assay pipeline that can capture and analyze the nucleic acids of a virus to provide accurate virus detection and lineage assignment results. The dynamic clinical assay pipeline can be dynamically modified based on factors such as virus prevalence and / or historical lineage assignment to provide more efficient and accurate virus detection. Specifically, molecular inversion probes can be designed, fabricated, and used to capture target nucleic acid molecules, and PCR methods can be used to amplify the captured molecules. To ensure sufficient sequencing depth and coverage for accurate virus detection and lineage assignment, the amplified molecules can be ligated to adapters of a selected size to obtain circular molecules. The circular molecules are sequenced using a desirable sequencing method, such as circular consensus sequencing, and the consensus sequence of each target nucleic acid molecule is determined based on predetermined rules. A sequencing file is generated to store information containing sequence reads and their positions in a reference genome so that the sequencing file and the stored data can be used to adjust and generate a report file. The report file contains virus detection and / or lineage assignment information determined based on the consensus sequence. The virus detection and / or lineage assignment information can be used to dynamically modify the molecular inversion probes and / or adapters to provide a dynamic clinical assay pipeline and provide efficient and accurate virus detection and lineage assignment. Systems using the described techniques can achieve 98.84% accuracy in detecting viruses and assigning the correct lineage without the need to remove uncovered regions or mask repetitive sequences.
[0057] There are various advantages associated with and achieved by the techniques described herein. First, wet laboratory techniques, including specially designed molecular capture and ligation methods, enable the quality of the molecules to be sequenced, and sequencing techniques such as circular consensus sequencing double ensure uniform read coverage across the viral genome. Second, since more qualified sequence reads of each molecule can be obtained and guaranteed, fewer molecules are required per sample, and more samples can be sequenced together, which improves sequencing efficiency and reduces sequencing costs. Third, the techniques described herein are generally applicable and / or can be easily modified to detect different lineages and types of viruses and can thus be implemented as a dynamic clinical assay pipeline for virus detection. In addition, since the entire system is specially designed and the read counts are predictable, the sequencing process can be optimized to ensure the best cost. Last but not least, the systems and methods using the described techniques can improve the overall accuracy of virus detection and lineage assignment.
[0058] In various embodiments, a method can include: obtaining nucleic acids from a sample obtained from a subject; using molecular inversion probes to capture target molecules in the nucleic acids under hybridization conditions; using polymerase chain reaction (PCR) to amplify the target molecules to obtain a plurality of amplified molecules; for each of the plurality of amplified molecules, ligating adapters to each end of the molecule to produce circular molecules; and sequencing the circular molecules to obtain sequence reads; using a computing system to generate a sequencing file that includes the sequence reads for each of the plurality of amplified molecules and the position of each sequence read in a reference genome of a virus by aligning the sequence reads with the reference genome of the virus; using the computing system and the sequencing file to generate a report file for the subject, where the report file includes a predicted lineage of the virus in the sample, where the generating includes: generating a consensus sequence of the target molecules based on the sequence reads for each of the plurality of amplified molecules, where if at least a predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, the nucleotide identity is assigned to the position, and where if less than a predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, "N" is assigned to the position; determining one or more scores for the consensus sequence based on the reference genome of the virus or a library of the virus; and determining the predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence.
[0059] II. Terms and Definitions
[0060] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of this specification and the relevant art and should not be interpreted in an idealized or overly formal sense. Well-known functions or constructions may not be described in detail for the sake of brevity and / or clarity.
[0061] It should be understood that although the terms first, second, third, etc. may be used herein to describe various elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer, or part from another region, layer, or part. Thus, without departing from the teachings of the present disclosure, the first element, component, region, layer, or section discussed below may be referred to as the second element, component, region, layer, or section. Unless otherwise specifically stated, the order of operations (or steps) is not limited to the order given in the claims or the drawings.
[0062] As used herein, unless the context clearly indicates otherwise, the singular forms "a / an" and "the" are also intended to include the plural forms. It will be further understood that when used in this specification, the term "comprises and / or comprising" specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "about X to Y" mean "about X to about Y".
[0063] As used herein, the terms "home collection" or "self - collection" etc. refer to the use of a kit that can be provided to a subject, the kit containing a swab, a container with a transport fluid (e.g., buffer), and a container for returning the self - collected sample to a laboratory for testing.
[0064] As used herein, the terms "automated" and "automatic" mean that an operation can be performed with little or no manual operation or input. The term "semi - automated" means that some input or activation is allowed by the operator, but calculations, collection, purification, and other steps are done electronically, usually programmed, without manual input.
[0065] As used herein, when an action is "based on" something, this means that the action is at least partially based on at least a part of something.
[0066] As used herein, the terms "clade" or "lineage" refer to a subset of viral variants or viral species, which are defined by a specific combination of genetic differences or their mutations or biomarkers.
[0067] As used herein, "CT" or "ct" refers to the cycle threshold, or the total number of cycles required to amplify and detect a nucleic acid target (e.g., viral nucleic acid) by real-time PCR and / or PCR.
[0068] As used herein, the terms "patient" or "subject" are used broadly and refer to an individual from whom a sample is provided for testing or analysis. The individual from whom a sample is collected, obtained, and / or provided, the "patient" or "subject", includes any and all warm-blooded mammalian subjects such as humans and / or animals.
[0069] As used herein, the terms "probe", "probe oligonucleotide", "oligonucleotide", and "probe oligonucleotide sequence" may be used interchangeably. The terms "probe", "probe oligonucleotide", "oligonucleotide", and "probe oligonucleotide sequence" can be used to refer to any molecule or system for detecting a target molecule, and the length of the probe or probe oligonucleotide can vary, e.g., from 4 nucleotides to about 200 nucleotides.
[0070] As used herein, the term "programming" means performing using computer programs and / or software, operations directed by a processor or ASIC. The term "electronic" and its derivatives refer to automated or semi-automated operations using a device having circuitry and / or modules rather than by mental steps, and generally refers to operations performed in a programmed manner.
[0071] As used herein, the term "protocol" refers to an automated electronic algorithm (usually a computer program) having defined rules of mathematical calculations, data interrogation, and analysis, which manipulates a system to execute a set of instructions.
[0072] As used herein, repeatability (or within-assay precision) describes the agreement between successive measurement results for the same analyte under the same measurement conditions. Intra-assay repeatability refers to the measurement of variability when analyzing the same sample during one analytical run.
[0073] As used herein, reproducibility (or between-assay precision) describes the agreement between successive measurement results for the same analyte under the same measurement conditions. Inter-assay repeatability refers to the measurement of variability when analyzing the same sample during more than one analytical run.
[0074] As used herein, real-time PCR or quantitative PCR (qPCR) allows for the real-time detection of PCR amplification products. Real-time polymerase chain reaction (PCR) assays use fluorescently labeled probes or intercalating dyes to visualize the PCR reaction and monitor the quantity of the resulting double-stranded DNA product. Fluorescent 5'-nuclease assays (i.e., The [assay name] is a real-time PCR assay that uses a fluorescent probe, which consists of an oligonucleotide with a reporter dye attached to the 5' end and a quencher group dye attached at or near the 3' end. The probe binds to a specific target sequence located between the forward and reverse primers. During the extension phase of the PCR cycle, the 5' nuclease activity of Taq polymerase degrades the probe, causing the separation of the reporter dye from the quencher group dye and generating a fluorescent signal. In each cycle, additional reporter dye molecules are cleaved from their respective probes, and the fluorescence intensity is monitored during PCR. The Taq polymerase used may be inactive at room temperature and can be activated by incubation at 95°C before starting the cycling part of the assay. This minimizes the production of non-specific amplification products.
[0075] As used herein, the terms "sample", "patient sample", "biological sample", and "specimen" may be used interchangeably. Non-limiting examples of samples that can be analyzed using the disclosed methods and systems include blood or blood products (e.g., serum, plasma, etc.), urine, nasal swabs, liquid biopsy samples, skin swabs, lesion swabs, or combinations thereof. In some cases, DNA can be extracted from lesion material, such as lesion fluid on a dry swab, lesion fluid swab in viral transport medium, lesion fluid on a slide, crust, or the top of a lesion. The term "blood" encompasses whole blood, blood products, or any fraction of blood, such as serum, plasma, buffy coat, etc., as traditionally defined. Suitable samples include samples that can be deposited on a substrate for collection and drying, including but not limited to: blood, plasma, serum, urine, saliva, tears, cerebrospinal fluid, organs, hair, muscle, face, or other tissue samples, or other liquid aspirates. The term "biological sample" further refers to samples obtained from biological sources, including but not limited to animals, cell cultures, organ cultures, tissues, etc.
[0076] As used herein, the terms "substantially", "about", and "approximately" are defined as mostly but not necessarily completely the specified content as understood by a person of ordinary skill in the art (and completely includes the specified content). In any of the disclosed embodiments, the terms "substantially", "about", or "approximately" can be replaced with "[percentage] within" the specified content, where the percentage includes 0.1%, 1%, 5%, 10%, 15%, and 20%.
[0077] As used herein, the terms "swab" and "dry swab" refer to a physical carrier used to contain and / or obtain a sample from an individual. The carrier can be composed of various materials and can encompass cotton swab balls, tissues, plastic forceps, plastic inoculation loops, popsicle sticks, dry or wet swab sticks, etc.
[0078] III. Virus Detection and Linear Assignment Techniques
[0079] One or more embodiments described herein may be implemented using operational, executable, or programmable systems, subsystems, modules, blocks, or components. The operational or executable systems, subsystems, modules, blocks, or components may include in vitro or computer-simulated operations or actions that may be performed by an operator (e.g., a researcher or practitioner), a machine (e.g., a computing device or a sequencer), or a combination thereof. The program systems, subsystems, modules, blocks, or components may include programs, subroutines, portions of programs, or software or hardware components capable of performing one or more of the tasks or functions. As used herein, a system, subsystem, module, block, or component may exist on a hardware component independent of other systems, subsystems, modules, blocks, or components. Alternatively, a system, subsystem, module, block, or component may be a shared element or process of other systems, subsystems, modules, blocks, or components.
[0080] Figure 1 is a block diagram of a system 100 capable of performing virus detection and lineage assignment techniques described herein. System 100 is a dynamic clinical assay pipeline that includes a molecular capture subsystem 110, a sequencing subsystem 120, and an analysis subsystem 130. According to various embodiments, Figure 1 The systems and subsystems, modules, blocks, or components (e.g., programs, code, or instructions) shown in may be performed by an operator (e.g., a human user), a dedicated machine (e.g., a sequencer), or one or more processors (e.g., a CPU, a GPU, etc.) with the assistance of or using a computing device. In some cases, the virus to be detected is a coronavirus, an influenza virus, or a respiratory syncytial virus (RSV). In some cases, the virus to be detected is a non-variola orthopoxvirus, such as monkeypox or vaccinia virus. In other cases, the virus to be detected is a variola orthopoxvirus, such as the smallpox virus. The length of the virus is typically thousands to millions of bases or base pairs. For example, the genome of monkeypox is approximately 190,000 base pairs, and the genome size of SARS-CoV-2 is approximately 30,000 bases (30,000 nucleotides).
[0081] The molecular capture subsystem 110 is a wet laboratory subsystem in which test and analyze chemicals, drugs, or other materials or biological substances, and the subsystem requires water, direct ventilation, and dedicated plumbing facilities. The molecular capture subsystem 110 may include one or more blocks, modules, or services. For example, as Figure 1As shown, the molecular capture subsystem 110 includes a nucleic acid acquisition block 112, a probe oligonucleotide acquisition block 114, and a multiplexing module 116, which includes an adhesion block 111, a gap filling block 113, a probe removal and release block 115, a PCR amplification block 117, and a post-PCR cleanup block 118. The input to the molecular capture subsystem 110 is one or more samples 105, and the output of the molecular capture subsystem 110 is a plurality of amplified molecules 125.
[0082] At the nucleic acid acquisition block 112, nucleic acids are obtained from the one or more samples 105. Before being obtained at the nucleic acid acquisition block 112, the one or more samples 105 may be collected or obtained from a subject. Depending on the sample type and downstream application, appropriate sample collection methods may be used. For example, venipuncture or finger stick methods may be used to collect blood samples, and saliva samples may be collected using available collection kits. In some embodiments, the one or more samples 105 are collected using a kit for self-collection of samples (e.g., the Monkeypox PCR Test Home Collection Kit). The Monkeypox PCR Test Home Collection Kit is intended for use by individuals presenting with an acute, generalized pustular or vesicular rash suspected of monkeypox disease to self-collect a lesion swab sample in a medium at home. The swab sample is placed in the medium and transported to a laboratory for testing of non-variola orthopoxvirus DNA extracted from the sample.
[0083] Nucleic acids from the one or more samples 105 can be obtained by isolation and extraction. For example, appropriate laboratory sample preparation methods can be used to break down cells and tissues to release nucleic acids and remove contaminants such as proteins, lipids, and other cellular debris. Various sample preparation methods are available, such as phenol-chloroform extraction, column-based purification, or bead-based purification. Once the sample is prepared, an appropriate extraction method is needed to extract the nucleic acids. The choice of extraction method may depend on the type of nucleic acid of interest (e.g., DNA or RNA) and the downstream application. Commonly used extraction methods include organic extraction, silica-based column purification, or bead-based extraction. Appropriate quality control methods may also be performed at the nucleic acid acquisition block 112 to check the quality of the obtained nucleic acids. In some cases, the one or more samples 105 may be nucleic acid samples and no preparation or extraction methods are required. For example, the one or more samples 105 may be nucleic acids obtained from a patient.
[0084] At the probe oligonucleotide acquisition block 114, one or more probe oligonucleotides are prepared and obtained. The probe oligonucleotides can be pre-designed to be suitable for capturing target molecules to detect the presence or absence of a virus in a sample. For example, if the virus to be detected is the monkeypox virus, the probe oligonucleotides can be designed to be suitable for capturing the DNA molecules of the monkeypox virus, and the DNA molecules have a predetermined number of base pairs in length (for example, about 675 base pairs in length). Then, the pre-designed probe oligonucleotides can be synthesized using appropriate methods such as solid-phase synthesis, enzymatic synthesis, or chemical synthesis. The synthesized probe oligonucleotides are obtained at the probe oligonucleotide acquisition block 114. The probe oligonucleotide acquisition block 114 can only perform the function of obtaining probe oligonucleotides. The design and synthesis / fabrication of the probe oligonucleotides can be carried out by separate systems, subsystems, or blocks.
[0085] At the multiplexing module 116, functions including ligation, gap filling, probe removal and release, and amplification can be performed at different blocks. For example, the ligation block 111 can ligate the probe oligonucleotides obtained at block 114 to the target molecules. Ligation is a key step in hybridization assays, which allows the probe oligonucleotides to specifically bind to their target sequences. Different probe oligonucleotides can be designed and ligated to different target molecules at the ligation block 111. When the target molecule is a double-stranded DNA molecule, the corresponding probe oligonucleotides that capture the DNA molecule can be double-stranded oligonucleotides. If the target molecule is a single-stranded RNA molecule, the corresponding probe oligonucleotides can be single-stranded DNA molecules and are ligated to the RNA molecule to obtain an RNA-DNA hybrid duplex in which the RNA strand and the DNA strand are complementary to each other. The ligation process can form a duplex structure for subsequent steps or processes.
[0086] At the gap filling block 113, the gaps or missing portions of the DNA or RNA strands can be filled. Appropriate methods can be used to perform gap filling, including polymerase synthesis and homologous recombination.
[0087] At the probe removal and release block 115, the probes that have not reacted with any target molecules are removed, and the remaining probes are released from the hybrid molecules. A variety of methods can be used to remove the unreacted probes, including washing the mixture with a buffer solution that disrupts the probe-target interaction, or using enzymatic digestion to degrade the probes or target sequences. Probe release can be carried out by heating, denaturation, or enzymatic digestion.
[0088] At the PCR amplification block 117, the target molecules can be amplified using PCR technology. Primers complementary to the flanking regions of the target molecules, polymerase, and deoxynucleotide triphosphates (dNTPs) can be used to perform PCR amplification. The primers can serve as the starting points for the polymerase to synthesize new DNA or RNA strands during PCR amplification, and different polymerases can be used to bind to the primers to synthesize new DNA or RNA strands.
[0089] At the post-PCR cleanup block 118, unwanted components are removed. A post-PCR cleanup step is typically required to remove unwanted reaction components such as primers, nucleotides, enzymes, and other impurities. Different post-PCR cleanup methods, including gel electrophoresis, bead-based purification, and spin column purification, can be used at the post-PCR cleanup block. Post-PCR cleanup can improve the accuracy and efficiency of downstream applications. After the post-PCR cleanup step, the amplified molecules 125 are obtained and can be used for subsequent sequencing and analysis subsystems. The target molecules are amplified thousands to millions of times. In some cases, the number of amplified molecules is about 16 million times the number of target molecules.
[0090] The sequencing subsystem 120 can be performed in a sequencer, which is an automated system capable of sequencing DNA or RNA molecules and analyzing genetic fragments for various applications. Examples of sequencers that can be used to perform the sequence of the sequencing subsystem 120 can include Sanger sequencers, Illumina sequencers, PacBio sequencers, Oxford Nanopore sequencers, and IonTorrent sequencers, etc. For example, the sequencing subsystem 120 can be a capillary electrophoresis-based system, where genetic fragments bound to probes migrate through a polymer and fluorescence emission is measured. Multiple capillary arrays allow for sample loading in a multi-well microplate format. The sequencing subsystem 120 can also be a system based on pyrosequencing technology for rapid sequencing and analysis. Pyrosequencing is a method of DNA sequencing (and in some cases RNA sequencing) based on the principle of "sequencing by synthesis", where sequencing is performed by detecting nucleotides incorporated by DNA polymerase. Pyrosequencing relies on light detection based on a chain reaction upon the release of pyrophosphate. Modules, blocks, or components can be stored on non-transitory computer media. As needed, one or more of the modules, blocks, or components can be loaded into system memory (e.g., RAM) and executed by one or more processors in the sequencing subsystem 120. The sequencing subsystem 120 can comprise one or more blocks, modules, or services. For example, as Figure 1 shown, the sequencing subsystem 120 includes an adapter ligation block 122, a sequencing block 124, and a demultiplexing block 126.
[0091] At the adapter ligation box 122, the adapter can be ligated to the molecules of the amplified molecule 125. During adapter ligation, a ligase (such as DNA ligase) is used to ligate the adapter to the ends of the molecules. For example, the adapter used at the adapter ligation box 122 can be a SMRTbell adapter. The SMRTbell adapter is an adapter sequence used in PacBio single molecule real-time (SMRT) sequencing technology. The SMRTbell adapter consists of two complementary oligonucleotide strands that can anneal to the molecule ends and then be ligated together to form a circular molecule. Then, the circular molecule can be sequenced using PacBio's SMRT sequencing technology. Other suitable adapters can also be used at the adapter ligation box 122.
[0092] At the sequencing box 124, the molecules (such as circular molecules) obtained from the adapter ligation box 122 are sequenced using an appropriate sequencing technology to obtain sequence reads. This process typically yields thousands to millions or millions to billions of sequence reads. At the sequencing box 124, circular consensus sequencing (CCS) technology can be used, which enables the DNA polymerase to repeatedly traverse the same region of the circular molecule. The CCS sequencing used at the sequencing box 124 also ensures longer read lengths and higher consensus accuracy. In some cases, consensus sequence reads are generated based on the raw sequence reads. The consensus sequence reads can be referred to herein as "sequence reads". It should be understood that other sequencing technologies, such as nanopore sequencing, can also be used at the sequencing box 124 to provide sequence reads suitable for applying the disclosed technology.
[0093] At the demultiplexing box 126, the sequence reads obtained at the sequencing box 124 can be separated by sample (or molecule). Demultiplexing is a process of separating pooled sequencing reads based on sample-specific barcodes or indices added to each sample during library preparation. When sequencing involves multiple samples, the sequencing reads from different samples are typically pooled together and sequenced in a single run at the sequencing box 124. To separate the sequence reads from each sample, demultiplexing needs to be performed. In some cases, demultiplexing also includes an adapter trimming step. In other cases, the adapters have been trimmed at the sequencing box 124.
[0094] The analysis subsystem 130 can be executed in a computing device such as a computer, a central processing unit (CPU), a graphics processing unit (GPU), etc. The modules, boxes, or components of the analysis subsystem 130 can be stored on a non-transitory computer medium. As needed, one or more of the modules, boxes, or components can be loaded into the system memory (such as RAM) and executed by one or more processors in the analysis subsystem 130. The analysis subsystem 130 can include one or more boxes, modules, or services. For example, asFigure 1 As shown, the analysis subsystem 130 includes a sequencing file frame 132, a decision module 134, and a reporting file frame 136.
[0095] At the sequencing file frame 132, a sequencing file is generated by the computing device on which the analysis subsystem 130 is executed. The sequencing file can be a FASTQ file, a BAM file, a SAM file, a FASTA file, a CRAM file, etc. The generation of the sequencing file is based on the output of the sequencing subsystem 120. The sequencing file typically contains sequencing information, e.g., sequence reads of each amplified molecule and position data associated with the sequence reads. The position data can include the start / end positions of each sequence read in the reference viral genome. In some cases, the position data is determined by mapping each sequence read to or aligning it with the reference viral genome. Different alignment algorithms can be used for alignment. The sequencing file can be output to the decision module 134 for virus detection and / or lineage assignment.
[0096] At the decision module 134, functions including generating a consensus sequence, determining or obtaining decision rules, and determining the presence or absence of a virus and / or determining the virus lineage can be performed in different frames. For example, as Figure 1 shown, the decision module 134 includes a consensus sequence frame 131, a decision rule frame 133, and a virus detection frame 135.
[0097] At the consensus sequence frame 131, a consensus sequence of each target molecule is generated. Different consensus sequence generation rules can be used to generate the consensus sequence at the consensus sequence frame 131. In some instances, if at least a predetermined number of sequence reads have nucleotide identity at a certain position in the consensus sequence, the nucleotide identity is assigned to that position, and if less than the predetermined number of sequence reads have nucleotide identity at a certain position in the consensus sequence, "N" is assigned to that position. In some cases, the predetermined number is four. The wet lab and the sequencing subsystem ensure read coverage such that the generation method based on consensus reads works efficiently and effectively to determine the consensus sequence. In most cases, at least 4-fold read coverage is ensured in 80% of the positions in the reference viral genome.
[0098] At decision rule block 133, decision rules regarding virus detection and / or lineage assignment are obtained. The decision rules can be predetermined by researchers, physicians, practitioners, or any authorized personnel. Decision rules can also be determined using a computing device based on an algorithm or a machine learning model. In some cases, the consensus sequence is mapped to a reference genome or a portion of the reference genome, and one or more scores are determined based on the decision rules. In some cases, the one or more scores are determined based on the distribution of nucleotide comparisons (e.g., nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc.) between the consensus sequence and the reference genome or a portion thereof. For example, the distribution shows the counts of nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc. The resolution of the nucleotide comparisons can vary. In some cases, the nucleotide comparison is a single nucleotide comparison. In other cases, the nucleotide comparison is a polynucleotide comparison (e.g., a trinucleotide comparison). Different scores can be assigned to different types of nucleotide comparisons. In some cases, within the same type of nucleotide comparison, different scores are assigned to subtypes. For example, the score assigned for substituting "A" with "G" can be different from the score assigned for substituting "A" with "C". In some cases, different scores are assigned based on the phylogenetic context (e.g., evolutionary relationships and lineage classifications) of the nucleotide comparison. For example, nucleotide substitutions that are common in a specific lineage may receive a different score than rare nucleotide substitutions. In some cases, nucleotide comparisons that result in a change in the encoded amino acid in the protein-coding region can receive a higher score, especially if it is known that such a change affects protein function. In some cases, the same score will be assigned to different types of nucleotide comparisons.
[0099] The decision rules can include a similarity metric that describes the relationship between the consensus sequence and the reference genome (or reference sequence). In some cases, the similarity metric assigns different weights to the one or more scores associated with different consensus sequences. For example, the weight assigned to an exact match between the consensus sequence and the reference genome is different from the weight assigned to a mismatch between the consensus sequence and the reference genome. In some cases, the similarity metric assigns positive weights to the total number of mutations in the consensus sequence and the total number of mutations in the reference genome, and assigns a negative weight to the nucleotide comparison between the consensus sequence and the reference genome. The decision rules can also include a threshold for determining the presence or absence of a virus and / or a threshold for determining the predicted lineage assignment. In some cases, if the one or more scores are greater than or equal to the threshold, it is determined that a virus is detected, or it is determined that the predicted lineage is assigned to the virus in the sample. If the one or more scores are less than the threshold, it can be determined that no virus is present in the sample, or it is determined that the lineage is not assigned to the virus in the sample.
[0100] Decision rules can include decision tree models. In some cases, the decision tree model is an annotated phylogenetic tree that includes different lineages as nodes. In some cases, the decision rules are stored in a tree data structure, such as a binary tree, binary search tree, trie, n-ary tree, general tree, B-tree, XML tree, etc. Storing the decision rules in a tree data structure can improve the efficiency of data search (e.g., decision-making) and can save space. The tree structure can easily adapt to dynamic data, enabling efficient updates, insertions, and deletions. This makes the tree structure suitable for data that changes over time and is suitable for dynamic clinical assay pipeline designs as disclosed herein. The tree structure also provides the feasibility of visualizing virus lineage assignments.
[0101] At the virus detection block 135, the presence or absence of a virus in the sample and / or the predicted lineage of the virus is determined. In some cases, both the presence or absence of a virus in the sample and the predicted lineage of the virus are determined based on the consensus sequence of the target molecule generated at the consensus sequence block 131 and the decision rules obtained at the decision rule block 133. In some cases, the presence or absence of a virus in the sample is determined or predetermined by wet laboratory methods, such as using real-time PCR, virus culture, immunofluorescence assay, electron microscopy, etc. If the presence of a virus in the sample has been predetermined (e.g., using real-time PCR or based on decision rules), then the predicted lineage assignment of the virus in the sample is determined at the virus detection block 135. The predicted lineage assignment of the virus in the sample can be determined by assigning a lineage to the virus in the sample based on a decision tree model (e.g., the decision tree model obtained at block 133).
[0102] In some cases, the predicted lineage is assigned based on a similarity metric. One or more scores for the consensus sequence can be determined at the virus detection block 135 based on the decision rules obtained at block 133, and the one or more scores are used to determine the presence or absence of a virus in the sample and / or the predicted lineage assignment of the virus in the sample. In some cases, the one or more scores are compared to each other, and the known lineage associated with the highest or lowest score is assigned to the virus in the sample. In some cases, when each of the one or more scores exceeds a predetermined threshold, an unknown lineage (or labeled as a "new" lineage) is assigned to the virus in the sample, and additional validation is performed to confirm the presence of this unknown lineage. The validation can be performed through peer review, reproducibility using sequencing techniques, data analysis, and phylogenetic methods (e.g., the pipeline disclosed herein), etc. In some cases, bioinformatics tools such as NextClade and Pangolin can be used or modified to determine the one or more scores, or the presence or absence of a virus in the sample and / or its predicted lineage assignment.
[0103] At the reporting file box 136, a reporting file is generated by the computing device on which the analysis subsystem 130 is executed. The reporting file can be a VCF (Variant Call Format) file, a VCFCons file, a BCF file, a MAF file, a GFF file, etc. The generation of the reporting file is based on the output of the decision module 134. In some cases, the reporting file includes (i) the presence or absence of a virus in a sample, and / or (ii) the predicted lineage of the virus in the sample. In some cases, the reporting file is output to a graphical user interface (GUI). In some cases, the reporting file also includes a display of the lineage assignment based on the decision rules obtained at block 133 (see, for example Figures 12A - 12B and 13B-13C).
[0104] In some cases, the information generated in the decision module 134 is used to dynamically modify the molecular capture subsystem 110 and / or the sequencing subsystem 120. For example, based on the lineage assignments determined over a period of time, the lineage prevalence can be determined. Based on the prevalent lineages, the probes used in the molecular capture subsystem 110 can be reselected or redesigned to better capture the target molecules corresponding to the prevalent lineages. The adapters used in the sequencing subsystem 120 can also be reselected or redesigned to accommodate the changes in the probes / target molecules. The dynamic modification of the system 100 through the feedback from the decision module 134 to the molecular capture subsystem 110 and / or the sequencing subsystem 120 helps to real-time reconstruct the dynamic clinical assay pipeline. The dynamic modification of the system 100 also improves the capture and sequencing efficiency using the known or estimated lineage prevalence. The dynamic modification of the system 100 also helps to provide accurate virus detection and lineage assignment results.
[0105] For example, if it is found that lineage B.1 is prevalent in a subject population (e.g., a subject population defined by a specific ethnicity, a specific age range, a specific geographical location, etc.), then the one or more probe oligonucleotides prepared and obtained at block 114 can be redesigned to be suitable for capturing target molecules specific to lineage B.1. The adapters used at the adapter ligation block 122 can also be updated based on the target molecules specific to lineage B.1. Additionally, the decision rules obtained at block 133 can be updated accordingly to reflect the prevalence of lineage B.1. These updates improve the accuracy of the prediction results. Furthermore, such optimization helps to significantly improve the efficiency and cost-effectiveness of future prediction and assignment of virus lineages for samples collected from the subject population. These dynamic adjustments and enhancements represent a substantial improvement in the field of virus detection and lineage assignment and in the field of virus assay procedures.
[0106] Figure 2FIG. 200 is an exemplary flowchart of a process for generating a report file regarding the presence or absence of a virus or a predicted lineage of a virus in a sample according to various embodiments. In some embodiments, Figure 2 the processing or a portion of the processing depicted in Figure 1 can be performed by the system 100 as described with reference to
[0107] Process 200 begins at block 210 where nucleic acids are obtained from a sample. The sample can be pre-obtained from a subject or a living organism such as a plant, an animal, or a human. The sample can be a blood sample, a urine sample, a tissue sample, a saliva sample, a fecal sample, a hair sample, a bone sample, etc. The nucleic acids obtained can include deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA). Various nucleic acid obtaining methods can be used at block 210 to obtain nucleic acids from the sample.
[0108] At block 220, target molecules in the nucleic acids are captured. In some embodiments, the target molecule is a DNA molecule. In some embodiments, the target molecule is an RNA molecule. In some embodiments, the target molecule is a mixture of a DNA molecule and an RNA molecule. In various embodiments, multiple target molecules in the nucleic acids are captured at block 220. In other embodiments, one target molecule in the nucleic acids can be captured. A variety of molecule capture methods can be used to capture the target molecules. In various embodiments, probe hybridization is used to capture the target molecules. In other embodiments, methods such as affinity capture, chromatography, microfluidics, etc. are used to capture the target molecules. Figure 4 are disclosed in detail in Section IV and
[0109] At block 240, the amplified molecules are ligated to adapters for efficient sequencing. In some embodiments, the adapter is a SMRTbell adapter of a selected size. A SMRTbell adapter is an adapter for single molecule real time (SMRT) sequencing technology. The adapter has a hairpin structure, where the 5' end and the 3' end of the adapter are combined together to form a loop. This structure allows the SMRTbell adapter to circularize, which is important for SMRT sequencing because it allows the polymerase to traverse the same molecule multiple times. Each SMRTbell adapter may also contain a substantially unique barcode / sequence for sample identification and demultiplexing during data analysis. This enables multiple samples to be sequenced together in a single SMRT cell, which improves the sequencing technology by increasing sequencing efficiency and reducing the cost per sample. The primer annealing and polymerase binding steps may also be performed at block 240. The ligated molecules are generated at block 240. In some embodiments, the ligated molecules are circular molecules.
[0110] At block 250, the ligated molecules are sequenced using an appropriate sequencing method to obtain sequence reads. In some embodiments, the sequencing method is circular sequencing. In some embodiments, the sequencing is SMRT sequencing. Other sequencing methods, such as RCA sequencing or nano-ring sequencing, may also be used to sequence the ligated molecules. The steps at block 240 and block 250 may be performed in an adapter design, ligation, and sequencing pipeline, which is disclosed in detail in Section V and Figure 5 herein.
[0111] At block 260, a sequencing file is generated. The sequencing file may contain information related to the sequencing step at block 250. In some embodiments, the sequencing file contains the sequence reads of each ligated molecule and the position of each sequence read in the reference genome of the virus obtained by aligning the sequence reads with the reference genome of the virus. In some embodiments, the sequencing file is a BAM file. The sequencing file may be generated by a computing device or a sequencer.
[0112] At block 270, a report file is generated. The report file can be generated by the same computing device that generated the sequencing file. The report file contains information related to the presence or absence of a virus in the sample and / or its predicted lineage. The consensus sequence of each target molecule can be determined at block 270. Determining the consensus sequence can improve the accuracy of determining the presence or absence of a virus in the sample and / or its predicted lineage. In some embodiments, a genomic construct can be determined based on the consensus sequence of the target molecule, and the genomic construct can be used to determine the presence or absence of a virus in the sample and / or to determine its predicted lineage. In some embodiments, it is known that the sample contains viral nucleic acid, and only the lineage of the virus in the sample is determined. At block 270, one or more scores can also be determined based on predefined or algorithm- or machine learning model-determined decision rules. The presence or absence of a virus in the sample and / or its predicted lineage can be determined at block 270 based on the one or more scores. In some embodiments, the report file is a VCF file or a modified VCF file. In some embodiments, the report file can be output to a GUI for presentation to a practitioner or a subject. The steps at blocks 260 and 270 can be performed in an adjusted VirSeq pipeline, which is disclosed in detail in Section VI.
[0113] At optional block 280, a treatment plan or a clinical test protocol is provided. The treatment plan or the clinical test protocol can be determined based on the one or more scores, the presence or absence of a virus in the sample and / or its predicted lineage, or the report file generated at block 270. In some embodiments, the virus is a monkeypox (MPX) virus, and the treatment plan includes administering an antiviral drug to the subject.
[0114] Figure 3 is an exemplary flowchart of process 300 that performs the disclosed techniques regarding the presence or absence of a virus in a sample or the predicted lineage of the virus according to various embodiments. In certain embodiments, Figure 3 the processing or a part of the processing depicted in Figure 1 can be performed by system 100 as described in reference
[0115] IV. Molecular Inversion Probe Set Design and PCR Pipelines
[0116] The molecular circular probe set design and the PCR pipeline can correspond to Figure 1 subsystem 110 in Figure 4 and is further shown in . The molecular circular probe set design and the PCR pipeline are designed to capture target molecules and amplify the captured molecules so that a sufficient number of molecules can be sent to the adapter design, ligation, and sequencing pipeline to ensure a sufficient number of sequence reads and sufficient coverage of the sequence reads and the reference genome. In some embodiments, the sufficient number of molecules / sequence reads is about 4 million to 8 million, and the sufficient coverage is (i) about or greater than 50% coverage (e.g., 50%, 60%, 70%, 80%, 90%), or (ii) about or greater than 100,000 sequence reads (e.g., 100,000 sequence reads, 1 million sequence reads, 10 million sequence reads).
[0117] To detect the presence or absence of a virus in a sample or to detect a predicted lineage of a virus, nucleic acids need to be obtained from the sample. Nucleic acids typically contain both host genetic molecules and potential or suspected viral molecules. Thus, the molecular circular probe set design and the PCR pipeline should be able to capture viral molecules in the nucleic acids. In some embodiments, the target molecules captured by the molecular circular probe set design and the PCR pipeline are viral molecules. The viral molecules can be DNA molecules or RNA molecules. DNA molecules are double-stranded molecules, while RNA molecules are typically single-stranded molecules. In some embodiments, the virus to be detected is the MPX virus, and the MPX virus is a double-stranded DNA virus.
[0118] To capture target molecules, probe hybridization methods can be used. The molecular circular probe set design and the PCR pipeline can include components for designing probes to capture target molecules. Appropriate design methods can be used to capture target molecules, and the methods may depend on various factors, including the molecule type, sensitivity and specificity requirements, and downstream analysis needs. For example, when the virus to be detected is the MPX virus and the target molecule is the MPX molecule, the probe can be designed to consist of two binding sites that are approximately 600 - 700 base pairs (bp) apart. In some embodiments, the probe for detecting the MPX molecule consists of two binding sites that are 675 bp apart. In some embodiments, multiple probes are designed to capture multiple target molecules. For example, to capture MPX virus molecules, more than 7,000 probes (e.g., 7,826) can be designed to capture 675 bp molecules evenly tiled in a ~200 kilobase pair (kb) MPX genome.
[0119] In some embodiments, the probe for capturing a target molecule is a molecular inversion probe. Compared to some other types of probes, molecular inversion probes can provide higher specificity and sensitivity in capturing target molecules. For example, molecular inversion probes are designed to hybridize to specific molecules or genomic regions of interest, and this specificity reduces the likelihood of false positives or false negatives in downstream analysis. Due to their high signal-to-noise ratio, molecular inversion probes are also capable of detecting low levels of DNA or RNA molecules in nucleic acids. Additionally, molecular inversion probes can be designed to capture multiple target molecules or genomic regions simultaneously, which allows for high-throughput analysis of multiple targets. Furthermore, compared to some other target capture methods, molecular inversion probes are cost-effective. Molecular inversion probes are also easy to design and can be customized to fit specific research and industrial / clinical needs.
[0120] Figure 4 Shows a portion of a molecular circle probe set design and a PCR pipeline for capturing target DNA molecules according to various embodiments. As Figure 4 shown, a molecular inversion probe can be designed to have two target complementary regions H1 and H2, two PCR primers P1 and P2, and a probe release cleavage site X1. During probe hybridization, as Figure 4 shown in step 1, multiple probes can be added to the nucleic acids obtained from a sample, where each probe binds to the target molecules on the H1 and H2 regions. The molecular inversion probe can tile more than 99% (e.g., 99.6%) of the 200 kb MPX viral genome and consists of two binding sites approximately 675 bp apart.
[0121] Molecular inversion probes can be beneficial in capturing target molecules. First, the molecular inversion probes allow tiling and prevent dropout due to novel genomic variations. Second, unlike other probe-based assays that require shearing the genome to capture molecules, the use of molecular inversion probes does not require upstream genomic processing, and the pipeline using molecular inversion probes can be processed faster and relatively inexpensively at high throughput. Additionally, this ensures the expected product size, which helps optimize sequencer loading and analysis pipelines. This also ensures complete (or sufficient) coverage of the genome and deep coverage of the genome at most base positions.
[0122] After binding, the region between the two probes is synthesized with a DNA polymerase and ligated to form a closed molecule, as Figure 4 shown in step 2. The gap is filled by the DNA polymerase using free nucleotides, and the ends of the probes are ligated by a ligase, thereby forming a fully circularized probe.
[0123] Since the non-reacted probes do not undergo the gap filling step, these probes remain linear. As Figure 4As shown in step 3, non-ligated or incomplete loops remain linear molecules and are removed by exonuclease digestion.
[0124] As Figure 4 shown in step 4, the circular molecules are released and linearized. In some embodiments, the probe release cleavage site X1 is cleaved by a restriction endonuclease such that the probe strand becomes linearized. In this linearized probe, the universal PCR primer sequences are located at the 5' and 3' ends, and the captured genomic target becomes part of the internal segment of the probe. In other embodiments, the probe can remain as a circularized molecule and no cleavage is required.
[0125] Then, the captured molecules can be enriched by amplification with the 3'-molecule-loop-specific M13 universal sequence and the 5'-sample-specific barcode, as Figure 4 shown in step 5. If the probe is linearized, traditional PCR amplification can be performed using the universal primers of the probe to enrich the captured molecules. Otherwise, a rolling circle amplification method such as RCA can be performed on the circular probe.
[0126] In some embodiments, Figure 4 the adaptors P1 and P2 in can be used to add barcodes. Then, the samples can be combined in equal volumes and a library can be prepared for sequencing (e.g., PacBio sequencing). Library preparation may require DNA damage repair, sequencing adaptor ligation, removal of unligated products by enzymatic digestion, and bead purification. Then, the library can be sequenced on a PacBio SequelII, for example, in the case of a 15-hour movie.
[0127] In certain embodiments, the target molecule is an RNA molecule. The molecular loop probe set design and PCR pipeline are also capable of capturing and amplifying RNA molecules. In such embodiments, reverse transcriptase can be used to synthesize cDNA from the RNA.
[0128] V. Adapter Design, Ligation, and Sequencing Pipelines
[0129] The adaptor design, ligation, and sequencing pipeline can correspond to Figure 1 subsystem 120 in and Figure 5Further shown in. The adapter design, ligation, and sequencing pipeline can be designed to provide a circular sequencing pipeline to obtain stable sequence read coverage at each locus and across the genome. In various embodiments, SMRT sequencing technology can be used to generate sequence reads. To use SMRT sequencing technology, adapters can be designed and ligated to target molecules to produce circular molecules. SMRT sequencing technology can increase the sequence read length to a much larger average than that of traditional sequencing methods. In some cases, the average length of sequence reads produced by SMRT sequencing technology can be 50 kb. In the case of circularized DNA molecules, this sequence read length can provide approximately more consecutive passes (e.g., 70 passes) of a single target molecule (e.g., in the size range of about 600 - 700 bp). Individual passes can be integrated into a single consensus high-accuracy consensus sequence (HiFi) read (e.g., 600 - 700 bp in length) with an accuracy of 99%.
[0130] Sequencing using single molecule real-time (SMRT) long-read sequencing technology requires circular templates (sequencing molecules). The circular templates generated by library preparation bind to polymerase and primers and are loaded onto an SMRT Cell (sequencing cell). The single molecule products diffuse onto one of 8 million zero-mode waveguides (ZMW) holes, where the polymerase is immobilized at the bottom. Then, phosphorylated ligated nucleotides are introduced into the ZMW, where bases can then be incorporated by the polymerase. When a given base pair is incorporated, its addition produces a nucleotide-specific light emission, which is detected by a camera by hole. This process is repeated for a given time or movie length, and the nucleotide sequence on a given hole is analyzed and translated into the corresponding nucleotides in the long sequence read output.
[0131] Figure 5 Shows a portion of the adapter design, ligation, and sequencing pipeline for generating circular DNA molecules according to various embodiments. As Figure 5 shown (see the schematic workflow at https: / / ccs.how / , last accessed on March 22, 2023), the adapter design, ligation, and sequencing pipeline starts with high-quality double-stranded DNA molecules. Hairpin adapters can be ligated to linear DNA molecules to produce circular DNA molecules. In some embodiments, the adapter is an SMRTbell adapter. The size of the adapter can vary. The circular DNA molecules enable the pipeline to produce long sequence reads that can be used to generate HiFi reads. After adapter ligation, selected primers and DNA polymerase can be added to sequence the circular DNA molecule and obtain sequence reads.
[0132] In some embodiments, as Figure 5As shown in the right part of, long sequence reads are trimmed of adapters to produce subreads, and HiFi reads can be generated based on the subreads. There are advantages associated with HiFi reads. First, HiFi reads provide both read length and accuracy suitable for various downstream analyses such as variant calling. In addition, the technology also improves sequencing efficiency: it reduces the sequencing time and cost per sample while increasing throughput and sequencing efficiency.
[0133] VI. Adjusted VIRSEQ Pipelines
[0134] The adjusted VirSeq pipeline can correspond to Figure 1 the analysis subsystem 130 in. The adjusted VirSeq pipeline is designed to provide a personalized report on viral information in a sample obtained from a subject. The adjusted VirSeq pipeline takes as input sequence reads from the adapter design, ligation, and sequencing pipelines and outputs a personalized report. In some embodiments, the personalized report can be displayed in a GUI.
[0135] In various embodiments, a sequencing file can be generated based on sequence reads and a reference viral genome. The sequence reads can be the raw sequence reads generated by the adapter design, ligation, and sequencing pipelines, or HiFi reads generated using the raw sequence reads. The reference viral genome can be obtained from a viral database or modified based on viral genomes obtained from the viral database. In some embodiments, the sequencing file is a BAM file. In some embodiments, the sequencing file is a FASTQ file. Software, scripts, and code can be used to generate the sequencing file.
[0136] The adjusted VirSeq pipeline and the dynamic clinical assay pipeline are designed to be easily adaptable to detect different viruses with various genomic length ranges. In some embodiments, the adjusted VirSeq pipeline and the dynamic clinical assay pipeline are designed to detect the MPX virus. Figure 6 An overview of the complexity of the MPX genome is shown (see Genomic epidemiology of monkeypox virus, https: / / nextstrain.org / monkeypox / hmpxv1, last accessed on March 22, 2023). The MPX genome is approximately 10 times that of the SARS-CoV-2 genome and contains a large number of long stretches of repetitive DNA and droplet regions. More probes in the molecular circular probe set design and PCR pipeline can be designed and used to capture target molecules of viral genomes with higher complexity. For example, the number of probes required to detect the MPX virus is approximately 7 to 10 times the number of probes required to detect SARS-CoV-2.
[0137] In some embodiments, PacBio SMRT LINK software and custom molecular ring processing scripts can be used to generate sequencing files. The sequencing files can be analyzed using a genomic analysis pipeline. In some embodiments, the genomic analysis pipeline can be implemented using a CLC Genomics Server. It should be understood that other sequencing analysis systems can also be used. At this point, the sequencing primer sequences can be removed and the sequences can be aligned with a reference viral genome (e.g., the MPX reference genome) to generate an aligned BAM file. In certain embodiments, Minimap2 can be used to generate the alignment. Other alignment programs or algorithms can be used. In some embodiments, sequence reads that meet a minimum coverage of 50% are used as input for generating consensus sequences and / or genomic constructs and for virus detection. The minimum coverage limit can vary from 20% to 100% (e.g., 20%, 30%, 40%, 60%, 70%, 80%, 90%).
[0138] The adjusted VirSeq pipeline and the dynamic clinical assay pipeline design can achieve better coverage. Better coverage can be achieved through the wet lab methods and pipelines described in the present disclosure. These methods and pipelines involve precise design of the molecular capture and ligation processes to ensure sequencing of high-quality molecules. Additionally, sequencing techniques such as circular consensus sequencing further improve the uniform read coverage across the viral genome, thereby improving the overall sequencing quality. Figure 7 is an Integrative Genomics Viewer (IGV) plot where the reads per nucleotide are on the Y-axis and the genomic coordinates are on the X-axis. It can be obtained from Figure 7 seen that there is a very uniform and complete coverage of the MPX genome. Compared to the reference sequence, the dark vertical bars are the positions where variants are detected. Variant detection will be introduced later in this section.
[0139] Figure 8 is Figure 7 a higher-resolution graphical representation of the IGV plot generated using software developed for various clinical samples. The genomic coverage and the CCS reads per base coverage for eight samples are shown in Figure 8 . When the minimum coverage threshold is predetermined to be four (as shown by the horizontal dashed line in Figure 8 ), seven out of eight samples reached the minimum coverage threshold, indicating that the disclosed technology provides a stable method for sequence read generation.
[0140] Figure 9 shows an alternative way to demonstrate the genomic coverage and read depth for various samples and controls. The sequencing data for sixteen samples are shown in Figure 9 . The vertical dashed line marks the minimum coverage threshold (where the coverage per base is equal to four). As shown in Figure 9As shown, fifteen out of the sixteen samples reached the minimum coverage threshold.
[0141] Figure 10 Genomic coverage and per-base coverage according to various embodiments are shown. Figure 10 At the top, it shows that at a base coverage of 50x, the genomic coverage of 32 out of 48 samples exceeded 80%. Figure 10 At the bottom, it shows the individual sample average (solid line) and the 25 / 75 quartiles of the coverage (dashed line). Figure 10 Indicates that the adjusted VirSeq pipeline and the dynamic clinical assay pipeline perform well in more multiplexed runs. This is achieved by obtaining and ensuring more high-quality sequence reads per molecule (e.g., using precise designs of the molecule capture and ligation processes). Thus, fewer molecules are required per sample, and more samples can be sequenced simultaneously. This also improves sequencing efficiency and reduces sequencing costs.
[0142] In some embodiments, the generation of the sequencing file may include preprocessing, such as demultiplexing. In certain embodiments, the preprocessing may include generating circular consensus sequence (CCS) BAM files, merging intermediate BAM files, using demultiplexing to generate individual BAM files corresponding to different barcode combinations, combining the outputs of the demultiplexed samples by sample name and / or subject identifier, removing barcodes from the sequences and generating individual sample FASTQ files, aligning the sequences with barcodes and trimming the barcodes, converting the BAM files to FASTQ files and copying the FASTQ and CCS BAM files to the final location as the sequencing files.
[0143] In some embodiments, generating the sequencing file may further include filtering sequence reads based on length or quality and aligning the sequence reads with a reference viral genome (e.g., the MPX reference genome). In some embodiments, this alignment is a local alignment using tools in the CLC Genomics Server.
[0144] Then, the sequencing file can be used to generate a consensus sequence for each target molecule. In some embodiments, VCFcons (e.g., VCFcons v8.5.0) can be used to generate the consensus sequence. VCFcons is a modified version of VCF and is a versatile VCF-based consensus sequence generator for genomes (e.g., small genomes). In some embodiments, a predetermined threshold (e.g., 4) can be obtained for generating the consensus sequence. For example, in certain embodiments, when VCFcons calls a nucleotide identity at a certain position of a target molecule, it must have at least 4 sequence reads covering that position. In some embodiments, the sequence reads are circular consensus sequencing (CCS) reads. If there are fewer than 4 reads of a nucleotide, it is reported as "N" (ambiguous, undefined nucleotide) in the consensus sequence. In some embodiments, when at least 4 sequence reads cover a certain position of a target molecule and the nucleotide identity at that position is present in at least 50% of the sequence reads aligned to that position, the nucleotide identity at that position is called. In some embodiments, a higher percentage (e.g., 60%) or a lower percentage (e.g., 40%) can be used. For example, the method for generating the consensus sequence of a target molecule is described in U.S. Patent Application No. 17 / 845,629, the entire content of which is incorporated herein by reference for all purposes.
[0145] The sequencing file can also be used to generate a genomic construct for each sample. For example, in certain embodiments, when VCFcons calls a nucleotide identity at a certain position of a genomic construct, it must have at least 4 consensus sequences covering that position. If there are fewer than 4 consensus sequences of a nucleotide, it is reported as "N" in the genomic construct. In some embodiments, when at least 4 consensus sequences cover a certain position of a genomic construct and have an alternative allele frequency greater than 50% compared to a reference, the nucleotide identity at that position is called.
[0146] After determining the sequence base composition, the percentage of non-ambiguous bases can also be determined. In some embodiments, the determination can be based on Seqtk or an alternative algorithm.
[0147] Use the techniques disclosed herein (e.g., the techniques disclosed regarding Figure 1 of box 135) to determine virus detection and / or lineage assignment. Both virus detection and lineage assignment can be determined using the consensus sequence and / or genomic construct and decision rules. In some embodiments, virus detection is performed using a wet laboratory approach, and lineage assignment is determined using the consensus sequence (or genomic construct) and decision rules. In some embodiments, the decision rules are predetermined (e.g., in Figure 1(the decision rule obtained at box 133). In other embodiments, the decision rule is determined based on an algorithm or a machine learning model. The machine learning model can be pre-trained using sequences or sequence reads of viruses with known lineages and / or sequences or sequence reads of general virus species. In some embodiments, one or more scores are assigned to the consensus sequence and / or the constructed genome based on a viral reference genome or a viral library. In some embodiments, the one or more scores are assigned to the consensus sequence and / or the constructed genome using the decision rule.
[0148] The one or more scores can be determined based on the distribution of mutations or nucleotide comparisons (e.g., nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc.) between the consensus sequence (or genome construct) and the reference genome or a portion thereof. For example, the distribution shows the counts of nucleotide substitutions, nucleotide deletions, nucleotide insertions, etc. Different scores can be assigned to different types of mutations or nucleotide comparisons. In some embodiments, within the same type of mutation or nucleotide comparison, different scores are assigned to subtypes. For example, the score assigned for substituting "A" with "G" can be different from the score assigned for substituting "A" with "C". In some embodiments, different scores are assigned based on the phylogenetic context (e.g., evolutionary relationships and lineage classifications) of the mutation or nucleotide comparison. For example, a nucleotide substitution that is common in a specific lineage may receive a different score than a rare nucleotide substitution. In some embodiments, a mutation or nucleotide comparison that results in a change in the encoded amino acid in a protein-coding region can receive a higher score, especially if it is known that such a change affects protein function. In some embodiments, the same score will be assigned to different types of mutations or nucleotide comparisons.
[0149] Once the one or more scores are obtained, they can be used to determine the presence or absence of the virus in the sample and / or to determine its predicted lineage. In some embodiments, when the one or more scores or all scores are below a predetermined number or exceed a predetermined range (or threshold), the existing lineage will not be assigned to the virus in the sample. In contrast, a new lineage (different from the existing or known lineage) is recommended and provided to the report file.
[0150] In some embodiments, NextClade is used for virus detection and lineage assignment. In certain embodiments, Pangolin can be used to assign lineages to consensus sequences or genomic constructs. In one embodiment, Pangolin is configured to consider only genomes with at least 50% non-ambiguous bases. In certain embodiments, SummaryStat can be used to compile the results of NextClade, Pangolin, Seqtk, and other algorithms, platforms, or programs to generate coverage statistics (e.g., mean of median amplicon coverage and percentage of genome coverage) required for subsequent quality control. The percentage of genome coverage can be calculated by dividing the number of non-ambiguous bases (A, T, C, G) by the total sequence length, and the lineage classifications are aggregated, and only the samples that produce NextClade results and Pangolin lineage calls are retained for further processing.
[0151] For example, to detect the presence or absence of a virus, the aligned consensus sequence or genomic construct is compared nucleotide by nucleotide to a reference virus genome or a portion of the reference virus genome. The differences are then scored and / or reported accordingly. In some embodiments, different types of mismatches are scored differently. For example, a nucleotide substitution (changing from one nucleotide to another) may be scored higher or lower than a nucleotide deletion (i.e., a gap). In some embodiments, different nucleotide substitutions may result in different scores. The threshold can be predefined or based on an algorithm or machine learning model to determine the presence or absence of the virus based on the total score or score distribution.
[0152] In some embodiments, the presence of a virus in a sample is known, and only the predicted lineage of the virus is determined. In some embodiments, PCR or real-time PCR is used to determine the presence of a virus in a sample. PCR techniques can be used to detect the virus. Real-time PCR detection can be performed simultaneously with dynamic clinical assay pipeline analysis or can be performed prior to dynamic clinical assay pipeline analysis. Diagnostic positive / negative results can be provided to a subject based on real-time PCR. The diagnostic positive / negative results can also be used to determine whether to input a sample into the dynamic clinical assay pipeline.
[0153] Real-time PCR can be used to detect viruses. The basic principle of real-time PCR is to amplify and detect a specific region of viral nucleic acid using a fluorescent probe or dye. Detecting a virus using real-time PCR technology generally involves the following: (i) extracting viral nucleic acid by isolating it from a sample; (ii) designing primers and probes that target conserved regions of the viral genome, which are specific to the detected virus and allow for the amplification and detection of viral nucleic acid; (iii) preparing a PCR reaction mixture including primers, probes, and polymerase, where different fluorescent dyes can be added to the reaction mixture; (iv) performing amplification by running PCR in a thermal cycler that cycles through a series of temperature changes to amplify the viral nucleic acid, where during each cycle, the fluorescent signal is measured in real time to enable the detection and quantification of the virus; and (v) analyzing the data generated during the PCR reaction using specific software to determine the presence and / or amount of viral nucleic acid in the sample. The amount of viral nucleic acid can be expressed as a cycle threshold (Ct) value, which represents the cycle at which the fluorescent signal exceeds a predetermined threshold.
[0154] In some embodiments, to assign lineages, a variant of interest can be predetermined, and genomic regions related to the variant of interest can also be determined based on the variant of interest or based on an algorithm or machine learning model. The genomic regions related to the variant of interest can be single nucleotide polymorphisms (SNPs), continuous portions in a reference genome, or multiple discontinuous portions in a reference genome. One or more scores are determined based on the alignment of genomic regions of a consensus sequence or genomic construct with those of a reference genome or the genomic region of the variant of interest. A threshold can be predetermined or based on an algorithm or machine learning model to determine a predicted lineage based on the one or more scores or score distribution.
[0155] Figure 11 Exemplary virus calls and lineage assignments in a sample are shown. The 5 results (IDs 9, 10, 1, 3, and 4) in the upper table are exemplary circulating strains (under an a4 synthetic control) of publicly available genomes isolated in the United States. The 3 results (IDs 8, 6, and 14) in the lower table are genomic controls based on previous outbreaks that are no longer prevalent. The black vertical bars represent variants against a reference, the triangles represent insertions, and the gray boxes represent gaps in the genome, which are expected in the synthetic control.
[0156] Figure 12A and 12B Shows a lineage analysis of the initial trial run, showing the expected assignments for controls and current circulating strain patient samples. Figure 12A Shows the phylogenetic tree where the initial trial run samples are located. Figure 12B Is an enlarged view showing all samples under the B.1 subtree.
[0157] Figures 13A - 13C shows the analysis of more multiplexed runs with more control and patient samples. The expected distribution is as Figures 13B - 13C shown, and also confirms the introduction of multiple MPXs into the population as previously reported. Gisaid (database) entries for other detected lineages are shown for comparison. Figure 13A shows the NextClade lineage assignment results.
[0158] In some embodiments, virus detection and lineage assignment information can be used to dynamically modify a dynamic clinical assay pipeline such that more efficient and accurate virus detection and lineage assignment results can generally be achieved. For example, a dynamic clinical assay pipeline can be put into clinical use for a certain period of time, and virus detection and lineage assignment information can be collected. Based on the lineage assignment information collected during a certain period of time, the prevalence of lineages can be determined, and reference genomes or target molecules can be modified based on the prevalent lineages. In some embodiments, adapters and decision rules can also be modified based on the prevalent lineages. The dynamic clinical assay pipeline can be designed to automatically modify or update in real time. The dynamic clinical assay pipeline provides improvements in capturing target molecules and sequencing and also helps to provide accurate virus detection and lineage assignment results.
[0159] VII. Downstream Applications
[0160] The systems and methods disclosed herein can have various applications. One of the more important applications is to provide a treatment plan or clinical test protocol for a subject. When a virus is detected in a sample obtained from a subject, a treatment plan or clinical test protocol can be automatically generated using a computing system and provided to the subject or a practitioner. The treatment plan or clinical test protocol may vary depending on the detected virus and / or the predicted lineage. For example, when the virus is an MPX virus or a vaccinia virus, the treatment plan can include administering an antiviral drug to the subject.
[0161] In addition, it may be crucial to monitor the emergence of new variants and more virulent strains. This is because different variants or strains may elicit different responses or result in different treatments. For example, in some cases of MPX strains, up to 10% of patients died, while in the recent outbreak, the number of deaths was less than 1%. The ability of the disclosed technology to assign lineages in samples can track the spread and transmission of the virus across different strains. The output of the dynamic clinical assay pipeline can be used to estimate the scale of the outbreak through changes in the strain, thereby generating and implementing a plan to control the spread of the virus or strain. The dynamic clinical assay pipeline can also perform primer site integrity monitoring. For example, the MPX virus / strain can lose regions / loci, and sometimes the lost regions / loci are primer sites, and the monitoring of primer site integrity can provide feedback on the spread and transmission of the virus, and the information can also be used to dynamically modify the dynamic clinical assay pipeline.
[0162] The disclosed technology can also be used to design new assays or machines. All or part of the disclosed technology can be deployed into new assays or machines to perform efficient virus detection and / or predicted lineage assignment. In some embodiments, more than one virus can be designed to be detected in the new assay or machine.
[0163] VIII. Examples
[0164] The systems and methods implemented in the various embodiments can be better understood by reference to the following examples.
[0165] 1. Example 1: Monkeypox Analysis
[0166] In one example, 7826 uniformly tiled molecular circle inversion probes were designed and used in the system described herein, with each probe expected to generate a 675bp fragment of the 200kb genome of the MPX virus. These probe pairs were trimmed from both ends of the reads using custom code. In one example, for each sequence generated at the sequencing pipeline, 25bp was trimmed from both the 5' end and the 3' end of the sequence.
[0167] Figure 14 A comparison between the sequenced synthetic control and independently obtained sequencing data was presented, highlighting the regions of concordance that are linearly related to the mismatches. The synthetic control is a genome with a known ground truth supplied by the manufacturer. It was compared with the independently obtained sequencing data. Dots indicate that the two sequences have the same base at the genomic position. As Figure 14 shown, the highly repetitive repeat sequences at the 5' and 3' genomic ends were expected to be double mapped.
[0168] 2. Example 2: SARS-CoV-2 Analysis
[0169] The Labcorp VirSeq SARS-CoV-2 NGS test can be performed in laboratories designated by Labcorp that are certified under the Clinical Laboratory Improvement Amendments (CLIA) of 1988, 42 U.S.C. § 263a, and meet the requirements for performing high-complexity tests as described in the standard operating procedures for the Labcorp VirSeq SARS-CoV-2 NGS test reviewed by the FDA under an Emergency Use Authorization (EUA).
[0170] Intended Use
[0171] The Labcorp VirSeq SARS-CoV-2 NGS test is a next-generation sequencing (NGS) test on the PacBio Sequel II sequencing system designed to identify and differentiate SARS-CoV-2 phylogenetically classified into the PANGO lineages from SARS-CoV-2 positive samples identified using Labcorp's COVID-19 RT-PCR test or Labcorp SARS-CoV-2 and Influenza A / B assays when clinically indicated. The test is limited to laboratories designated by Labcorp that are certified under the Clinical Laboratory Improvement Amendments (CLIA) of 1988, 42 U.S.C. § 263a, and meet the requirements for performing high-complexity tests. The Labcorp VirSeq SARS-CoV-2 NGS test is designed to be used in conjunction with patient history and other diagnostic information when clinically indicated, i.e., when the results may assist in determining appropriate clinical management. The results of this test are intended to be interpreted by the ordering healthcare professional. The test is not intended to aid in the initial diagnosis of SARS-CoV-2 infection or to confirm the presence of SARS-CoV-2 infection, and is also not intended to identify specific SARS-CoV-2 genomic mutations. The results should not be used as the sole basis for treatment or other patient management decisions.
[0172] The Labcorp VirSeq SARS-CoV-2 NGS test is intended for use by qualified clinical laboratory personnel who have received specific instruction and training in the operation of the PacBio Sequel II sequencing system and next-generation sequencing workflows, as well as in vitro diagnostic procedures. The Labcorp VirSeq SARS-CoV-2 NGS test is used only under an Emergency Use Authorization from the Food and Drug Administration.
[0173] Device Description and Test Principles
[0174] The Labcorp VirSeq SARS-CoV-2 NGS test is a PacBio Sequel II-based whole-genome sequencing assay used to determine the PANGO lineage from RNA extracted from SARS-CoV-2 positive samples identified using Labcorp's COVID-19 RT-PCR test or Labcorp SARS-CoV-2 and influenza A / B assays. The SARS-CoV-2 probe set used in this assay contains ~1000 tiled molecular inversion probes (MIPs) designed to amplify 99.6% of the SARS-CoV-2 genome's RNA reverse transcribed to cDNA, with most bases covered by 22 MIPs. The products synthesized between MIPs are enriched and amplified followed by the addition of sample-specific molecular barcodes by sequencing.
[0175] The residual total nucleic acid extract of SARS-CoV-2 positive RT-PCR diagnostic test samples with N1 target cycle threshold (Ct) < 31 is tested using the Labcorp VirSeq SARS-CoV-2 NGS test. The residual nucleic acid extract can be stored at -20°C for up to 30 days with up to 2 freeze-thaw cycles. The residual total nucleic acid extract is transferred to a 96-well plate containing only positive RT-PCR diagnostic test samples using a Hamilton Microlab STAR. The samples are then aliquoted into a sequencing run plate containing 94 samples, with one water non-template control (NTC) and one positive control. One production batch processes eight plates or 752 samples.
[0176] A custom molecular circular SARS-CoV-2 capture kit is used to prepare samples for sequencing on a PacBio Sequel II instrument. First, reverse transcriptase provided by Thermo Fisher Scientific synthesizes cDNA from RNA. Then, the SARS-CoV-2 cDNA is used as the target for molecular circular probe hybridization ( Figure 15 , step 1) (information on Figure 15 can be found on the FDA website, EMERGENCY USE AUTHORIZATION (EUA) SUMMARY Labcorp VirSeq SARS-CoV-2 NGS Test, downloaded on March 22, 2023). The molecular circular probes tile 99.6% of the 30 kb SARS-CoV-2 viral genome and consist of two binding sites 600 bp apart. After binding, the 600 bp region between the two probes is synthesized and ligated with DNA polymerase to form a closed molecule ( Figure 15, step 2). Non-circular or incomplete circles remain linear molecules and are removed by exonuclease digestion ( Figure 15 , step 3). Then the circular molecules are released from the template cDNA ( Figure 15 , step 4), and enriched by amplification with the 3'-molecule-circle-specific M13 universal sequence and the 5'-sample-specific barcode ( Figure 15 , step 5). Then, the samples can be combined in equal volumes and the library is prepared for PacBio sequencing. Library preparation requires DNA damage repair, sequencing adapter ligation, removal of unligated products by enzymatic digestion, and bead purification. Then, the library is sequenced on a PacBio Sequel II in the case of a 15-hour movie ( Figure 15 , step 6).
[0177] After sequencing, FASTQ files for each sample are generated using the PacBio SMRT LINK software and a custom molecule circle processing script. The FASTQ is analyzed using the genomic analysis pipeline implemented in CLC Genomics Server version 9.1.1. This workflow starts with the sample-level fastq files, trims the primers, and aligns them to the SARS-CoV-2 reference genome ("NC_045512v2") using Minimap2 to generate an aligned bam file. The consensus sequence for each sample is generated using VCFcons (v8.5.0). When VCFcons calls a nucleotide sequence for genome construction, it must have at least 4 circular consensus sequencing (CCS) reads covering the base pair, and the alternative frequency of the allele is greater than 50% compared to the reference. If a nucleotide is covered by fewer than 4 reads, it is reported as ambiguous (N) in the consensus sequence. Then, the consensus sequence is used as the input for the PANGOLIN (v3.1.20) analysis package to assign the lineage of individual samples. The lineage results of samples with a published genome coverage of at least 90% and an overall genome coverage of more than 10 CCS reads are reported. The overall genome coverage of a sample is defined as the average of the median read coverage of 29 (length ~1 kb) contigs across the entire viral genome.
[0178] Instruments to be Used with the Test
[0179] The Labcorp VirSeq SARS-CoV-2 NGS test is used with a Pacific Biosciences Sequel II sequencing instrument, a Mantis liquid handler, and the Labcorp VIRSEQ analysis pipeline for sequence analysis and lineage determination. Table 1 presents the instruments and reagents required to perform the Labcorp VirSeq SARS-CoV-2 NGS test.
[0180] Table 1: Requirements for Reagents and Specialized Instruments
[0181]
[0182]
[0183] The designated laboratory will receive an FDA-approved instrument qualification protocol, which is included as part of the LabcorpVirSeq SARS-CoV-2 NGS test standard operating procedure (SOP) and will be instructed to perform the protocol prior to testing clinical samples. The designated laboratory must comply with the authorized SOP, including the instrument qualification protocol, according to the letter of authorization.
[0184] Controls to be Used with the Test
[0185] External Positive Control: The external positive control is added to each plate of 94 patient samples. The control consists of one of ten Twist synthetic SARS-CoV-2 RNA controls with a predefined PANGO lineage designation. The lineage identification results of the external positive control will be compared to its known PANGO lineage designation.
[0186] External Negative Control: An external non-template control (NTC) is required to ensure the absence of master mix contamination events on a given amplification plate. The control consists of molecular grade water added to the A1 position of each 96-well plate prior to adding the samples. The NTC is then transferred to the sequencing run plate along with the positive samples and sequencing and quality control (QC) analyses are performed.
[0187] Other Controls: Lineage identification results are only published for samples with a genomic coverage of at least 90% and an overall genomic coverage greater than 10 CCS reads.
[0188] Result Interpretation
[0189] All test controls should be examined before interpreting patient results. If the controls are invalid, the patient's results cannot be interpreted.
[0190] 1) Labcorp VirSeq SARS-CoV-2 NGS Test Control - Positive:
[0191] After sequencing, the lineage designation of the positive control for a given plate is compared to the known lineage designation of the positive control. If the lineage determined by its assay is different from the known lineage, the positive control sample will be considered failed. Samples on plates with failed positive external controls are reanalyzed once using an assay with the original total nucleic acid extract. If the residual total nucleic acid of the sample is exhausted, the sample is reported as failed.
[0192] 2) Labcorp VirSeq SARS-CoV-2 NGS Test Control - NTC:
[0193] After sequencing, the overall genomic coverage of the NTC sample is calculated. If the overall genomic coverage of the NTC sample is greater than 10 CCS reads, the sample will be considered failed. Samples on plates with failed negative external controls are reanalyzed once using an assay with the original total nucleic acid extract. If the residual total nucleic acid of the sample is exhausted, the sample is reported as failed.
[0194] Table 2: Expected Results of Labcorp VirSeq SARS-CoV-2 NGS Test External Controls
[0195]
[0196] 3) Examination and Interpretation of Patient Sample Results:
[0197] After external control analysis and removal of any samples on plates with failed external controls, the Labcorp VirSeq SARS-CoV-2 NGS test results are evaluated. The overall genomic coverage and percentage of genomic coverage are calculated for each patient sample. The lineage results of samples with a genomic coverage of at least 90% and an overall genomic coverage greater than 10 CCS reads are reported.
[0198] Table 3 summarizes the interpretation and reporting of clinical samples.
[0199] Table 3: Result Interpretation of Patient Samples
[0200]
[0201] Performance Evaluation
[0202] 1. Device Tolerance:
[0203] The residual nucleic acid extracts of 8,815 respiratory samples that tested positive for SARS-CoV-2 using the Labcorp COVID-19 test authorized under EUA200011 and performed across 3 sites with the Labcorp VirSeq SARS-CoV-2 NGS test were sequenced. For different N1 CT value categories, the number of samples with generated genomic sequences passing the following two quality control criteria was determined.
[0204] ● When compared to the SARS-CoV-2 reference genome (“NC_045512v2”), the genomic coverage is greater than 90%.
[0205] ● The overall genomic coverage is greater than 10 CCS reads.
[0206] Table 4: Results of the tolerance study
[0207]
[0208] 2. Precision (Repeatability): Intra-assay
[0209] Intra-assay repeatability was evaluated by testing 11 nucleic acid samples in triplicate. Intra-assay repeatability studies assess the ability of an assay to accurately detect replicates of the same sample within a single assay run. The N1 CT values of the samples being tested were in the range of 19.2 to 22.96.
[0210] The lineage designations of all 11 samples were consistent among the three replicates, with all replicates meeting the QC criteria.
[0211] 3. Precision (Reproducibility): Inter-assay
[0212] Inter-assay reproducibility was evaluated by testing the same nucleic acid samples used in the repeatability study in triplicate across three assay runs. One of the 11 samples tested in the repeatability study was inadvertently excluded from the last run. The sequencing runs were performed by 6 different technicians using 3 batches of different SMRT cells and 2 batches of sequencing reagents. The third run was 3 weeks after the first run, and the concentration of the multiplexed samples was approximately 1 / 4 of the sample concentration in the previous two runs.
[0213] Table 5: Description of the reproducibility study
[0214]
[0215] The lineage designations of 9 out of 10 samples were consistent across all 3 runs. In runs 2 and 3, the inconsistent sample did not pass the overall genomic coverage QC criteria.
[0216] 4. Sample Stability (Freeze-Thaw)
[0217] In the sample stability study, the stability of patient samples under the recommended storage conditions was evaluated. After nucleic acid amplification (NAA) diagnostic testing using the Labcorp COVID-19 RT-PCR test, the extracted nucleic acids were transported to the testing laboratory on dry ice and stored at -20°C before sequencing.
[0218] Twelve samples, including alpha, beta, and delta-related variants (VOCs), were initially sequenced. These samples were re-sequenced after being stored at -20°C for 5 weeks (requiring 3 freeze-thaw cycles). The initial test results of 11 out of the 12 samples were consistent with the repeat test results. The only inconsistent result was determined to be due to mechanical error and was not related to sample stability. The study results indicate the stability of the Labcorp VirSeq SARS-CoV-2 NGS test for samples stored at -20°C for up to 30 days with up to 2 freeze-thaw cycles.
[0219] 5. Concordance
[0220] The lineage naming results generated by the Labcorp VirSeq SARS-CoV-2 NGS test (PacBio molecular circular sequencing) were directly compared with the lineage naming results generated by Illumina COVIDSeq (RUO monitoring protocol with v3 primer pool) and PacBio amplicon sequencing. The analysis included the following samples.
[0221] ● Illumina COVIDSeq (RUO) samples
[0222] ○ 93 negative (NAA) samples
[0223] ○ 72 SARS-CoV-2 samples sequenced in the winter of 2020
[0224] ○ 29 samples previously sequenced at the Center for Molecular Biology and Pathology (CMBP)
[0225] ○ 50 samples previously sequenced in DNA identification
[0226] ● Samples for PacBio amplicon sequencing
[0227] ○ 122 SARS-CoV-2 samples for amplicon sequencing at 90% coverage.
[0228] Illumina COVIDSeq (RUO) Comparison:
[0229] Ninety-three samples previously determined to be negative by the Labcorp COVID-19 RT-PCR diagnostic test were sequenced in duplicate using the Labcorp VirSeq SARS-CoV-2 NGS test and the Illumina COVIDSeq (RUO). Among the 93 samples tested, one sample produced a reportable SARS-CoV-2 genome using only the Illumina COVIDSeq (RUO), and another sample produced a reportable SARS-CoV-2 genome using both the Illumina COVIDSeq (RUO) and the Labcorp VirSeq SARS-Cov-2 NGS test. Additional investigation revealed that both of these samples were SARS-CoV-2 positive by the Labcorp COVID-19 RT-PCR test and were erroneously included in the validation. Among these 91 true-negative samples, the final concordance between the Illumina COVIDSeq (RUO) and the Labcorp VirSeq SARS-CoV-2 NGS test was 100%.
[0230] Based on 72 samples sequenced with the Illumina COVIDSeq (RUO) during the winter of 2020, 51 samples produced reportable results when tested with the Labcorp VirSeq SARS-CoV-2 NGS test. Among these 51 reportable results, all lineage naming results were 100% concordant with the Illumina COVIDSeq (RUO) output.
[0231] Table 6: Comparison of PANGO lineage naming of circulating lineages during the winter of 2020 between the Illumina COVIDSeq (RUO) and the Labcorp VirSeq SARS-CoV-2 NGS test.
[0232] Samples Tested Reportable Results Consistent Reportable Results Percent of Consistent Reportable Results 77 51 51 100%
[0233] Among 79 samples, a total of 7 samples were lost during shipping or did not produce valid results when using the Illumina COVIDSeq (RUO). Among the 79 samples collected and DNA-identified at the CMBP, 72 samples were successfully sequenced using both the Illumina COVIDSeq (RUO) and the Labcorp VirSeq SARS-CoV-2 NGS test. Among the 72 samples that passed the sequence quality QC criteria of the Labcorp VirSeq SARS-CoV-2 NGS test, the lineage naming of all 66 samples was concordant with the lineage naming of the Illumina COVIDSeq (RUO).
[0234] Table 7: Comparison of PANGO lineage naming for circulating lineages in Summer 2021 between Illumina COVIDSeq (RUO) and Labcorp VirSeq SARS-CoV-2 NGS testing.
[0235] Samples Tested Reportable Results Consistent Reportable Results Percent of Consistent Reportable Results 72 66 66 100%
[0236] PacBio Amplicons-Based Sequencing Comparison:
[0237] One hundred and twenty-two samples that were initially sequenced using a PacBio amplicon-based method were processed in duplicate using the Labcorp VirSeq SARS-CoV-2 NGS test. Each of the 122 samples with amplicon sequencing coverage > 90% was tested in duplicate using the Labcorp VirSeq SARS-CoV-2 NGS test; thus, a total of 244 test results were generated. Among the 244 Labcorp VirSeq SARS-CoV-2 NGS test lineage determination results, 234 results produced genomes that passed two QC metrics. Among these 234 results, 225 results had lineage naming that was consistent with the initial PacBio amplicon assay lineage naming. When compared to the PacBio amplicon sequencing lineage naming results, 96% of the samples yielded consistent lineage identification results.
[0238] Table 8: Comparison of PANGO lineage naming between PacBio amplicon sequencing and Labcorp VirSeq SARS-CoV-2 NGS testing
[0239] Number of Results Reportable Results Consistent Reportable Results Percent of Consistent Reportable Results 244 234 225 96%
[0240] 6. Reference sample testing
[0241] Heat-inactivated SARS-CoV-2 samples from B.1.1.7 (VR-3326HK), Hong Kong / VM20001061, and Italian-INMI1 lineages characterized by ATCC were used in this assessment. Sequencing errors associated with the Labcorp VirSeq SARS-CoV-2 NGS test were evaluated by comparing all mutations identified in the consensus sequences generated by the Labcorp VirSeq SARS-CoV-2 NGS test analysis pipeline with the published ATCC reference sequences. The results are shown in the table below:
[0242] Table 9: Number of nucleotide mismatches between the ATCC reference sequences and sequences generated by the Labcorp VirSeq SARS-CoV-2 NGS test
[0243]
[0244] * b1117, B.1.1.7 (VR-3326HK); HK, Hong Kong / VM20001061; ITLY, Italy - INMI1
[0245] Overall, an average sequence difference of 0.012% was observed between the reference sequences and the consensus sequences generated by the Labcorp VirSeq SARS-CoV-2 NGS test.
[0246] 7. Simulation study
[0247] A simulation study was conducted to evaluate the performance of the Labcorp VirSeq SARS-CoV-2 NGS test in identifying samples with PANGO lineages that were not tested in the concordance and analytical studies.
[0248] Based on the sequencing results of 760 clinical samples, a sequencing error model was estimated that simulated how sequencing errors and ambiguous nucleotides were randomly introduced into the sequenced genomes by the Labcorp VirSeq SARS-CoV-2 NGS test. The model estimated that the Labcorp VirSeq SARS-Cov-2 NGS test produced 1 sequencing error per 33 SARS-CoV-2 genomes sequenced on average, and 3,369 ambiguous nucleotides per SARS-CoV-2 genome sequenced.
[0249] Relevant Variant / Variant of Interest Simulations
[0250] A total of 23,400 reference sequences with known Pango-lineage designations representing 234 lineages (100 reference sequences per lineage) were downloaded from GISAID. The sequencing error model was used to introduce sequencing errors into these 23,400 reference sequences to simulate the putative sequence output of the determination of these genomes. Each of these 23,400 sequences was used to generate multiple simulated sequences. The PANGO lineages of these simulated sequences were identified using the lineage identification software (PANGOLIN v3.1.20) used in the Labcorp VirSeq SARS-CoV-2 NGS test. The lineage identification results of each simulated sequence were compared with the known PANGO lineage designations of the reference sequences used to generate the simulated sequences. The concordance results for simulated sequences with genome coverages of 90%, 95%, and 99% are shown in the following table:
[0251] Table 10: In silico performance of 234 VOC / VOI lineages of the Labcorp VirSeq SARS-CoV-2 NGS test.
[0252] Genome Coverage Number of Sequences Tested Number of Consistent Sequences Percent of Consistent Sequences (95% CI) 90% 17538 16762 95.58%(95.26%-95.86%) 95% 16274 15739 96.71%(96.42%-96.97%) 99% 11169 10929 97.85%(97.56%-98.10%)
[0253] At the sub-lineage level (BA.1, BA.2, and BA.3), 100% (95% CI 99.55% - 100.00%) of 770 simulated Omicron sequences were accurately identified.
[0254] Simulation Time Period
[0255] 10,000 high-quality SARS-CoV-2 sequences were randomly sampled from GISAID. A total of 140,000 reference sequences with known PANGO lineage designations were downloaded from GISAID. A sequencing error model was used to introduce sequencing errors into these 140,000 reference sequences to simulate the putative sequence output of the determination of these genomes. Each of these 140,000 sequences was used to generate multiple simulated sequences. The PANGO lineages of these simulated sequences were identified using the lineage identification software (PANGOLIN v3.1.20) used in the Labcorp VirSeq SARS-CoV-2 NGS test. The simulated lineage identification results were compared with the known PANGO lineage designations of the reference sequences used to generate the simulated sequences. The concordance results for simulated reads with genomic coverages of 90%, 95%, and 99% are shown in the table below:
[0256] Table 11: In silico performance of 140,000 sequences submitted to GISAID for the Labcorp VirSeq SARS-CoV-2 NGS test.
[0257] Genome Coverage Number of Sequences Tested Number of Consistent Sequences Percent of Consistent Sequences (95% CI) 90% 105930 102270 96.54%(96.43%-96.65%) 95% 96333 94222 97.81%(97.71%-97.89%) 99% 45929 45182 98.37%(98.25%-98.48%)
[0258] All 43,328 simulated Omicron sequences were accurately identified as Omicron sequences. The sub-lineage concordance rates of the simulated Omicron sequences are shown below:
[0259] Table 12: In silico performance of 43,328 simulated Omicron sequences of the Labcorp VirSeq SARS-CoV-2 NGS test.
[0260]
[0261] IX. Computing Systems and Additional Considerations
[0262] Figure 16Disclosed is an example computing device 1600 suitable for use with a system and method for determining the presence or absence of a virus in a sample obtained from a subject and determining its predicted lineage in accordance with the present disclosure. The example computing device 1600 includes a processor 1605 that communicates with a memory 1610 and other components of the computing device 1600 using one or more communication buses 1615. The processor 1605 is configured to execute processor-executable instructions stored in the memory 1610 to perform one or more methods for determining the presence or absence of a virus in a sample obtained from a subject and determining its predicted lineage, according to different instances, such as some or all of the exemplary system 100, process 200, or 300 described above with respect to Figures 1 - 3 to determine the presence or absence of a virus in a sample obtained from a subject and determine its predicted lineage, as described above with respect to some or all of the exemplary system 100, process 200, or 300.
[0263] In this instance, the computing device 1600 also includes one or more user input devices 1630, such as a keyboard, mouse, touch screen, microphone, etc., for receiving user input. The computing device 1600 also includes a display 1635 for providing a visual output to the user, such as a user interface. The computing device 1600 also includes a communication interface 1640. In some instances, the communication interface 540 may be capable of communicating using one or more networks, including a local area network (“LAN”); a wide area network (“WAN”), such as the Internet; a metropolitan area network (“MAN”); a point-to-point or peer-to-peer connection, etc. Communication with other devices may be accomplished using any suitable network protocol. For example, a suitable network protocol may include the Internet Protocol (“IP”), the Transmission Control Protocol (“TCP”), the User Datagram Protocol (“UDP”), or a combination thereof, such as TCP / IP or UDP / IP.
[0264] Specific details are given in the above description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments may be practiced without these specific details. For example, circuits may be illustrated in block diagrams so as not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0265] The implementation of the above-described techniques, blocks, steps, and devices may be carried out in various ways. For example, these techniques, blocks, steps, and devices may be implemented in hardware, software, or a combination thereof. For a hardware implementation, the processing unit may be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the functions described above, and / or combinations thereof.
[0266] In addition, note that embodiments may be described as a process, which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Additionally, the order of the operations may be rearranged. A process is terminated when its operations are completed, but may have additional steps not included in the figure. A process may correspond to a method, a function, a program, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to the function returning to the calling function or the main function.
[0267] In addition, embodiments may be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting languages, and / or microcode, the program code or code segments for performing the necessary tasks may be stored in a machine-readable medium such as a storage medium. A code segment or machine-executable instruction may represent any combination of a program, a function, a subroutine, a program, a routine, a subroutine, a module, a software package, a script, a class, or an instruction, a data structure, and / or a program statement. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. The information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted in any suitable manner, including memory sharing, message passing, ticket passing, network transmission, etc.
[0268] For firmware and / or software implementations, the method may be implemented using modules (e.g., programs, functions, etc.) that perform the functions described herein. Any machine-readable medium that tangibly embodies instructions may be used in implementing the methods described herein. For example, software code may be stored in a memory. The memory may be implemented inside or outside the processor. As used herein, the term "memory" refers to any type of long-term, short-term, volatile, non-volatile, or other storage medium, and is not limited to any particular type of memory or any particular number of memories or the type of medium in which the memories are stored.
[0269] In addition, as disclosed herein, the terms "storage medium", "storage device", or "memory" may refer to one or more memories for storing data, the one or more memories including read-only memory (ROM), random access memory (RAM), magnetic RAM, magnetic core memory, disk storage media, optical storage media, flash memory devices, and / or other machine-readable media for storing information. The term "machine-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage media capable of storing instructions and / or data.
[0270] Although the principles of the present disclosure have been described above in connection with specific devices and methods, it should be clearly understood that this description is made by way of example only and is not intended as a limitation on the scope of the present disclosure.
Claims
1. A method, comprising: Obtaining nucleic acid from a sample obtained from a subject; Using molecular inversion probes to capture target molecules in the nucleic acid under hybridization conditions; Using polymerase chain reaction (PCR) to amplify the target molecules to obtain a plurality of amplified molecules; For each of the plurality of amplified molecules, Ligating adapters to each end of the molecule to produce circular molecules; and Sequencing the circular molecules to obtain sequence reads; Using a computing system to generate a sequencing file, the sequencing file comprising the sequence reads of each of the plurality of amplified molecules and the position of each sequence read in the reference genome of the virus obtained by aligning the sequence reads with the reference genome of the virus; And Using the computing system and the sequencing file to generate a report file for the subject, wherein the report file comprises the predicted lineage of the virus in the sample, wherein the generating comprises: Generating a consensus sequence of the target molecules based on the sequence reads of each of the plurality of amplified molecules, wherein if at least a predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, the nucleotide identity is assigned to the position, and wherein if less than a predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, "N" is assigned to the position; Determining one or more scores of the consensus sequence based on the reference genome of the virus or a library of the virus, wherein the one or more scores are determined based on the distribution of mutations of the virus; And Determining the predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence.
2. The method according to claim 1, wherein the report file further comprises the presence or absence of the virus in the sample.
3. The method according to claim 2, wherein the presence or absence of the virus in the sample is determined using real-time PCR (RT-PCR) or based on the one or more scores of the consensus sequence.
4. The method according to claim 1, wherein the subject tests positive for the virus and the sample contains viral nucleic acid.
5. The method according to claim 1, wherein the virus is a monkeypox (MPX) virus.
6. The method according to claim 1, wherein the target molecule is a double-stranded DNA molecule.
7. The method according to claim 1, wherein the molecular inversion probe consists of two binding sites spaced about 600-700 bp apart.
8. The method according to claim 1, wherein the sequencing file is a binary alignment map (BAM) file.
9. The method according to claim 1, wherein the report file is a variant call format (VCF) file.
10. The method according to claim 1, further comprising providing a treatment plan or a clinical test protocol to the subject.
11. The method according to claim 10, wherein the treatment plan comprises administering an antiviral agent to the subject.
12. The method according to claim 1, further comprising: obtaining a prevalent lineage of the virus for a population of subjects, wherein the prevalent lineage is determined based on the predicted lineage of the virus in the sample; updating the molecular inversion probe to capture the target molecule, wherein the target molecule is specific for the prevalent lineage; updating the adaptor based on the updated molecular inversion probe; and obtaining a set of decision rules specific to determining the prevalent lineage, wherein the one or more scores for determining the consensus sequence are determined using the set of decision rules.
13. A method comprising: obtaining nucleic acid from a sample obtained from a subject; capturing a plurality of target molecules in the nucleic acid using a plurality of molecular inversion probes under hybridization conditions; amplifying each target molecule using polymerase chain reaction (PCR) to obtain a plurality of sets of amplified molecules, wherein each set of amplified molecules corresponds to a target molecule; for each molecule in each set of amplified molecules, ligating an adaptor to each end of the molecule to produce a circular molecule; and sequencing the circular molecule to obtain sequence reads; using a computing system to generate a sequencing file comprising the sequence reads for each molecule in each set of amplified molecules and the position of each sequence read in a reference genome of the virus by aligning the sequence reads with the reference genome of the virus ; and using the computing system and the sequencing file to generate a report file for the subject, wherein the report file comprises a predicted lineage of the virus in the sample, wherein the generating comprises: generating a consensus sequence for each target molecule based on the sequence reads for each molecule in the set of amplified molecules corresponding to the target molecule, wherein if at least a first predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, the nucleotide identity is assigned to the position, and wherein if less than the first predetermined number of the sequence reads have nucleotide identity at a position in the consensus sequence, "N" is assigned to the position; generating a genomic construct of the nucleic acid in the sample based on the consensus sequence of the target molecule, wherein if at least a second predetermined number of the consensus sequences have nucleotide identity at a position in the genomic construct and the nucleotide identity at the position is present in at least 50% of the sequence reads aligned to the position, the nucleotide identity is assigned to the position, and wherein if less than the second predetermined number of the consensus sequences have nucleotide identity at a position in the genomic construct, "N" is assigned to the position; determining one or more scores for the consensus sequence based on the reference genome of the virus or a library of the virus; and Determine the predicted lineage of the virus in the sample based on the one or more scores of the consensus sequence.
14. The method according to claim 13, wherein the report file further includes the presence or absence of the virus in the sample.
15. The method according to claim 14, wherein the presence or absence of the virus in the sample is determined using real-time PCR (RT-PCR) or the one or more scores of the consensus sequence.
16. The method according to claim 13, wherein the subject tests positive for the virus and the sample contains viral nucleic acid.
17. The method according to claim 13, wherein the virus is a monkeypox (MPX) virus.
18. The method according to claim 13, wherein the target molecule is a double-stranded DNA molecule.
19. The method according to claim 13, wherein the molecular inversion probe consists of two binding sites separated by about 600-700 bp.
20. The method according to claim 13, wherein the sequencing file is a binary alignment map (BAM) file.
21. The method according to claim 13, wherein the report file is a variant call format (VCF) file.
22. The method according to claim 13, which further includes providing a treatment plan or a clinical test protocol to the subject.
23. The method according to claim 22, wherein the treatment plan includes administering an antiviral drug to the subject.
24. The method according to claim 13, which further includes: Obtain the prevalent lineage of the virus for a population of subjects, wherein the prevalent lineage is determined based on the predicted lineage of the virus in the sample; Update the plurality of molecular inversion probes to capture mutations specific to the prevalent lineage; Update the adaptor based on the updated plurality of molecular inversion probes; And Obtain a set of decision rules specific to determining the prevalent lineage, wherein the one or more scores for determining the consensus sequence are determined using the set of decision rules.
25. A system, which includes: One or more data processors; and A non-transitory computer-readable medium storing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform the method according to any one of claims 1 to 24.
26. A computer program product tangibly embodied in a non-transitory machine-readable medium, the computer program product including instructions configured to cause one or more data processors to perform the method according to any one of claims 1 to 24.
Citation Information
Patent Citations
Methods and Systems for Detection of Covid Variants
US20220411886A1