Methods of analyzing viral nucleic acid
The method of generating DNA fragments, enriching with capture probes, and sequencing EBV DNA addresses inefficiencies in current sequencing methods, enabling comprehensive EBV genomic analysis and risk assessment.
Patent Information
- Application Number
- PCT/SG2025/050377
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-04
- Filing Date
- 2025-06-03
- Publication Date
- 2025-12-11
AI Technical Summary
Current methods for sequencing Epstein-Barr virus (EBV) genomes face challenges such as inefficiency due to low viral DNA content in samples, inability to capture repetitive sequences with short-read technologies, and reference bias from single reference genomes.
A method involving generating DNA fragments, enriching for viral DNA using specific capture probes, and performing long or short-read sequencing to analyze EBV genomic DNA, with thresholds for selecting sequencing type based on viral load and DNA integrity.
Enables comprehensive analysis of EBV genomic diversity, capturing repetitive sequences and reducing reference bias, allowing for detailed variant analysis and risk assessment.
Smart Images

Figure IMGF000041_0001 
Figure IMGF000063_0001 
Figure IMGF000064_0001
Abstract
Description
[0001] Methods of Analyzing Viral Nucleic Acid
[0002] Technical field
[0003] The present invention relates generally to methods for analyzing and sequencing nucleic acids and more specifically to methods of analyzing and sequencing viral genomic DNA in a sample.
[0004] Background
[0005] Epstein-Barr virus (EBV), also called human herpesvirus 4 (HHV-4), is one of the most disseminated viruses, with more than 90% of human populations having been infected. While EBV infections are largely asymptomatic, they are known to play a critical role in several malignancies such as nasopharyngeal carcinoma (NPC), lymphoma and gastric cancer, as well as in autoimmune diseases such as systemic lupus erythematosus (SLE) and multiple sclerosis (MS). A notable feature of EB V-associated malignancies is the remarkable ethnic and geographic variation in prevalence. Understanding the genetic diversity and evolution of EBV genomes in different patient cohorts from different parts of the world can help to elucidate the role of EBV genome variation in disease development.
[0006] EBV has a double-stranded DNA genome of around 170 kb, of which approximately 30%> consists of highly repetitive sequences, also known as repeat regions (RRs). Some of the longest repeats belong to the group known as the family of repeats (FR), and to the group of internal repeat regions (IR1, IR2, IR3 and IR4) and terminal repeat regions (TR). The ability to sequence these repeats is important as many of them encode proteins (or parts of proteins) that play major roles in virus biology, particularly in viral latency and persistence. However, current methods of EBV sequencing have several drawbacks. First, because patient samples typically contain only a small amount of EBV DNA (less than 1 %) relative to human DNA, direct sequencing of patient samples can be very inefficient for obtaining EBV genomic data. Second, attempts to sequence EBV isolates using next-generation sequencing (NGS) have mostly relied on short-read technologies (with each read up to 300 bp), but the repetitive sequences in the EBV genome are longer than typical short reads. As a result, short-read sequencing cannot fully capture the variability of the repetitive sequences. Third, conventional genome sequencing analysis of EBV is largely performed using referencebased mapping and variant calling analysis based on a single EBV reference genome. The use of a single reference sequence introduces “reference bias” into sequence analysis, because current EBV reference genomes (e.g., NC_007605) do not represent the genetic diversity of emerging and circulating EBV strains.
[0007] It would be desirable to overcome or alleviate at least one of the above-described problems, or at least to provide a useful alternative.
[0008] Summary
[0009] Disclosed herein is a method of analyzing viral genomic DNA in a sample; the method comprising: a) generating a population of DNA fragments from the sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample.
[0010] Disclosed herein is a method of sequencing a viral DNA genome; the method comprising: a) generating a population of DNA fragments from a sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to sequence the viral DNA in the sample.
[0011] Disclosed herein is a method of analyzing viral genomic DNA in a sample; the method comprising: a) detecting the viral copy number and / or DNA integrity in the sample; b) generating a population of DNA fragments from the sample; c) preparing a DNA library from the DNA fragments; and d) performing long or short read sequencing on the DNA library to analyze the viral genomic DNA in the sample, wherein long read sequencing is performed when the viral copy number is at least 3000 copies per 500 ng of sample DNA and / or the DNA integrity number (DIN) is at least 7 ; and wherein short read sequencing is performed when the viral copy number is greater than 250 copies but less than 3000 copies per 500 ng of sample DNA.
[0012] Disclosed herein is a method of detecting a subject who is at risk of a cancer; the method comprising: a) generating a population of DNA fragments from a sample obtained from the subject; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral-specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample, so as to determine whether the subject is at risk of a cancer.
[0013] Disclosed herein is a kit for analyzing EBV genomic DNA, the kit comprising one or more nucleic acid capture probes derived from a reference EBV genome comprising sequence information of EBV genomes from a plurality of EBV strains.
[0014] Disclosed herein is a method of analyzing EBV genomic DNA, the method comprising: a) obtaining a plurality of long or short read sequencing reads from a sample comprising viral genomic DNA, and b) comparing the plurality of sequence reads to a reference EBV genome.
[0015] Brief description of the drawings
[0016] Figure 1 shows the distribution of EBV read lengths across saliva samples with the long read capture and AS dataset.
[0017] Figure 2 shows the mapping coverage of 1 saliva sample against EBV reference NC_007605.1 with the three enrichment methods. The two innermost rings show the genomic position and GC content of the EBV reference. The following rings show the coverages generated from each method. The outermost ring highlights the Repeat Regions in EBV reference.
[0018] Figure 3 shows the comparative analysis for the performance of variant calling across the 3 different method data for one saliva sample.
[0019] Figure 4 shows the dots represent variants called in the 4 saliva samples across the 3 different methods. Figure 5 shows the SVs length and distribution of long read capture and AS data of one saliva sample.
[0020] Figure 6 shows the copy number diversity of the repetitive regions in saliva samples.
[0021] Figure 7 shows the distribution of EBV read lengths across eight EBV+ cell lines in a PacBio long read dataset.
[0022] Figure 8 shows the distribution of EBV reads’ coverage. Lines at the bottom of each graph indicate the coordinates of well-known large repeat regions (RRs).
[0023] Figure 9 shows the structure of contigs assembled using different algorithms.
[0024] Figure 10 shows the distribution of EBV read lengths across eight EBV+ cell lines in a ONT ultra-long read dataset.
[0025] Figure 11 shows a pairwise sequence comparison between the newly assembled and published sequences.
[0026] Figure 12 shows the distribution of A) IRl-spanning and B) TR-spanning reads in an ONT ultra-long read dataset.
[0027] Figure 13 shows a pan-genome graph reference drawn based on 8 newly assembled EBV genome sequences. Each bubble represents a structural variation, which can be multi-allelic as reflected by multiple paths through the bubble.
[0028] Figure 14 shows a comparison of variants called between the pan-genome and eight individual EBV genome sequence.
[0029] Detailed description
[0030] The inventors have developed an integrated analytical solution that enables sequencing of Epstein-Barr virus (EBV) and other viral genomes in various types of biological samples. The solution addresses various existing limitations of EBV genome analysis, namely: 1) the inability to read long repetitive genomic sequences; 2) low quantities of EBV DNA in isolates; and 3) the “reference bias” due to the employment of single reference sequence.
[0031] The analytical solution is flexible, allowing for both short read and long read sequencing approaches depending on the DNA quality and viral load in the sample. The inventors have identified specific thresholds for selecting long read and short read platforms for sequencing. For example, samples with a high viral load (e.g., > 3000 viral copies in 500 ng of total DNA) and / or good DNA integrity (e.g., DNA Integrity Number (DIN) > 7) may be analysed using long read sequencing, which provides more uniform genome coverage and more comprehensive variant analysis particularly in highly repetitive genomic regions. Samples with a lower viral load (e.g., between 250 and 3000 viral copies in 500 ng of total DNA) may still be sequenced using short read sequencing. For both long read and short read sequencing, virus-specific probes may be used to enrich for viral DNA fragments to improve sensitivity and to allow targeted sequencing. Sequencing reads are compared to a newdy generated graph reference that incorporates the genomic diversity of multiple viral genomes. This allow'S for more detailed analysis of various viral strains that may be present in a sample population.
[0032] Accordingly, this disclosure provides compositions and methods for analyzing viral genomic DNA. Disclosed herein is a method of analyzing viral genomic DNA in a sample; the method comprising: a) generating a population of DNA fragments from the sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral-specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample.
[0033] In one embodiment, there is provided a method of analyzing viral genomic DNA in a sample; the method comprising: a) processing a sample containing viral genomic DNA; and b) performing long read adaptive sampling (AS) sequencing of viral genomic DNA on the processed sample. Sample processing may comprise: detecting the viral copy number and / or DNA integrity in the sample; generating a population of DNA fragments from the sample; and preparing a DNA library from the DNA fragments for sequencing. The long read adaptive sampling (AS) sequencing may be long read ONT adaptive sampling (AS) sequencing.
[0034] In one embodiment, there is provided a method of analyzing viral genomic DNA in a sample; the method comprising: a) detecting the viral copy number and / or DNA integrity in the sample; b) generating a population of DNA fragments from the sample; c) preparing a DNA library from the DNA fragments; and d) performing long or short read sequencing on the DNA library to analyze the viral genomic DNA in the sample, wherein long read sequencing is performed when the viral copy number is at least 3000 copies per 500 ng of sample DNA and / or the DNA integrity number (DIN) is at least 7; and wherein short read sequencing is performed when the viral copy number is greater than 250 copies but less than 3000 copies per 500 ng of sample DNA.
[0035] Disclosed herein is a method of sequencing a viral DNA genome; the method comprising: a) generating a population of DNA fragments from a sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to sequence the viral DNA in the sample.
[0036] Also disclosed herein is a method of analyzing viral genomic DNA, the method comprising: a) obtaining a plurality of long or short read sequencing reads from a sample comprising viral genomic DNA, and b) comparing the plurality of sequence reads to a reference viral genome.
[0037] General definitions
[0038] As used herein, the term “sample” includes tissues, cells, body fluids and isolates thereof etc., isolated from a subject. Examples of samples include: whole blood, blood fluids (e.g. scrum and plasma), lymph and cystic fluids, sputum, stool, tears, mucus, hair, skin, ascitic fluid, cystic fluid, urine, nipple exudates, nipple aspirates, sections of tissues such as biopsy and autopsy samples, frozen sections taken for histologic purposes, archival samples, explants, and primary and / or transformed cells or cell lines derived from patient tissues. Samples also include nucleic acids (including crude and pure or substantially pure nucleic acid preparations) isolated from a cell, tissue or isolate from a subject.
[0039] The term “biological sample” or “sample” generally refers to a sample or part isolated from a biological entity. The biological sample may show the nature of the whole and examples include, without limitation, bodily fluids, dissociated tumor specimens, cultured cells, and any combination thereof. Biological samples can come from one or more individuals. One or more biological samples can come from the same individual. One non limiting example would be if one sample came from an individual’ s blood and a second sample came from an individual's tumor biopsy. Examples of biological samples can include but are not limited to, blood, serum, plasma, nasal swab or nasopharyngeal wash, saliva, urine, gastric fluid, spinal fluid, tears, stool, mucus, sweat, earwax, oil, glandular secretion, cerebral spinal fluid, tissue, semen, vaginal fluid, interstitial fluids, including interstitial fluids derived from tumor tissue, ocular fluids, spinal fluid, throat swab, breath, hair, finger nails, skin, biopsy, placental fluid, amniotic fluid, cord blood, emphatic fluids, cavity fluids, sputum, pus, microbiota, meconium, breast milk and / or other excretions. The samples may include nasopharyngeal wash. Examples of tissue samples of the subject may include but are not limited to, connective tissue, muscle tissue, nervous tissue, epithelial tissue, cartilage, cancerous or tumor sample, or bone. The sample may be provided from a human or animal. The sample may be provided from a mammal, including vertebrates, such as murines, simians, humans, farm animals, sport animals, or pets. The sample may be collected from a living or dead subject. The sample may be collected fresh from a subject or may have undergone some form of pre-processing, storage, or transport.
[0040] “Bodily fluid” generally can describe a fluid or secretion originating from the body of a subject. In some instances, bodily fluids are a mixture of more than one type of bodily fluid mixed together. Some non-limiting examples of bodily fluids are: blood, urine, bone marrow, spinal fluid, pleural fluid, lymphatic fluid, amniotic fluid, ascites, sputum, or a combination thereof.
[0041] As used herein, the term “nucleic acid”, and equivalent terms such as “polynucleotide”, refer to a polymeric form of nucleotides of any length, such as ribonucleotides, deoxyribonucleotides or peptide nucleic acids (PNAs), that comprise purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The nucleic acid may be double- stranded or single-stranded. References to single- stranded nucleic acids include references to the sense or antisense strands. The backbone of the polynucleotide can comprise sugars and phosphate groups, as may typically be found in RNA or DNA, or modified or substituted sugar or phosphate groups. A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. The sequence of nucleotides may be interrupted by non-nucleotide components. The terms nucleoside, nucleotide, deoxy nucleoside and dcoxynuclcotidc generally include complements, fragments and variants of the nucleoside, nucleotide, deoxynucleoside and deoxynucleotide, or analogs thereof.
[0042] “Nucleotide” can refer to a base-sugar-phosphate combination. Nucleotides are monomeric units of a nucleic acid sequence (e.g., DNA and RNA). The term nucleotide includes naturally and non-naturally occurring ribonucleoside triphosphates ATP, TTP, UTP, CTG, GTP, and ITP, for example and deoxyribonucleoside triphosphates such as dATP, dCTP, diTP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives can include, for example, [aSJdATP, 7-deaza-dGTP and 7-deaza-dATP, and, for example, nucleotide derivatives that confer nuclease resistance on the nucleic acid molecule containing them. The term nucleotide as used herein also refers to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxy ribonucleoside triphosphates include, ddATP, ddCTP, ddGTP, ddITP, ddUTP, ddTTP, for example. Other ddNTPs are contemplated and consistent with the disclosure herein, such as dd (2-6 diamino) purine.
[0043] The terms “identical” or percent “identity” in the context of two or more nucleic acids or polypeptides, refer to two or more sequences or sub-sequences that are the same or have a specified percentage of nucleotides or amino acid residues that are the same, when compared and aligned (introducing gaps, if necessary) for maximum correspondence. The percent identity can be measured using sequence comparison software or algorithms or by visual inspection. Various algorithms and software that can be used to obtain alignments of amino acid or nucleotide sequences are well-known in the ait. These include, but are not limited to, BLAST, Clustal Omega, T-COFFEE, and MAFFT. In some embodiments, two nucleic acids or polypeptides described herein are substantially identical, meaning they have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and in some embodiments at least 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% nucleotide or amino acid residue identity, when compared and aligned for maximum correspondence, as measured using a sequence comparison algorithm or by visual inspection.
[0044] “Polymerase chain reaction” or “PCR” can refer to a technique for replicating a specific piece of selected DNA in vitro, even in the presence of excess non-specific DNA. Primers are added to the selected DNA, where the primers initiate the copying of the selected DNA using nucleotides and, typically, Taq polymerase or the like. By cycling the temperature, the selected DNA is repetitively denatured and copied. A single copy of the selected DNA, even if mixed in with other, random DNA, is amplified to obtain thousands, millions, or billions of replicates. The polymerase chain reaction is used to detect and measure very small amounts of DNA and to create customized pieces of DNA.
[0045] “Amplified nucleic acid” or “amplified polynucleotide” is any nucleic acid or polynucleotide molecule whose amount has been increased at least two folds by any nucleic acid amplification or replication method performed in vitro as compared to its starting amount. For example, an amplified nucleic acid is obtained from a polymerase chain reaction (PCR) which can, in some instances, amplify DNA in an exponential manner (for example, amplification to 2n copies in n cycles). Amplified nucleic acid can also be obtained from a linear amplification.
[0046] “Amplification product” can refer to a product resulting from an amplification reaction such as a polymerase chain reaction.
[0047] A nucleic acid “library” herein refers to a collection of nucleic acids. The library can contain one or more target fragments. In some instances, the target fragments are amplified nucleic acids. In other instances, the target fragments are nucleic acids that have not been amplified. A library can contain nucleic acid that has one or more known oligonucleotide sequence(s) added to the 3’ end, the 5’ end or both the 3 ’and 5’ ends. The library may be prepared so that the fragments contain a known oligonucleotide sequence that identifies the source of the library (c.g., a molecular identification barcode identifying a patient or DNA source). In some instances, two or more libraries are pooled to create a library pool. Libraries may also be generated with other kits and techniques such as transposon mediated labelling, or “tagmentation” as known in the art. Kits may be commercially available for library preparation, such as the NEBNext® Ultra™ II DNA Library Prep Kit, SMRTbell Express Template Prep Kit 2.0 (PacBio), and the Ultra-Long DNA Sequencing Kit (Oxford Nanopore Technologies).
[0048] The terms “detecting”, “determining”, “measuring”, “evaluating”, “assessing” and “assaying” are used interchangeably herein to refer to any form of measurement, and include determining if an element is present or not. These terms include both quantitative and / or qualitative determinations. Assessing may be relative or absolute. The method as defined herein may comprise measuring or visualising the levels of one or more polynucleotide analytes in a sample.
[0049] “Sequencing,” “sequence determination,” and the like generally refers to any and all biochemical methods that may be used to determine the order of nucleotide bases in a nucleic acid.
[0050] The term “epigenetic modification” as used herein, may be any covalent modification of a nucleic acid base. In some cases, a covalent modification may comprise (a) adding a methyl group, a hydroxymethyl group, a formyl group, a carboxyl group, a carbon atom, an oxygen atom, or any combination thereof to one or more bases of a nucleic acid sequence, (b) changing an oxidation state of a molecule associated with a nucleic acid sequence, such as an oxygen atom, or (c) a combination thereof. A covalent modification may occur at any base, such as an adenine, a cytosine, a guanine, a thymine, a uracil, or any combination thereof. In some cases, an epigenetic modification may comprise an oxidation or a reduction. In some cases, an epigenetically modified base may comprise a methylated base, a hydroxymethylated base, a formylated base, or a carboxylic acid containing base or a salt thereof. An epigenetically modified base may comprise a 5-methylated base, such as a 5- methylated cytosine (5-mC). An epigenetically modified base may comprise a 5- hydroxymethylated base, such as a 5-hydroxy methylated cytosine (5-hmC).
[0051] The term “biotin,” as used herein, is intended to refer to biotin (5-[(3aS,4S,6aR)-2- oxohcxahydro-IH-thicno[3,4-d]imidazol-4-yl]pcntanoic acid) and any biotin derivatives and analogs. Such derivatives and analogs are substances which form a complex with the biotin binding pocket of native or modified streptavidin or avidin. Such compounds include, for example, iminobiotin, desthiobiotin and streptavidin affinity peptides, and also include biotin-.epsilon.-N-lysine, biocytin hydrazide, amino or sulfhydryl derivatives of 2- iminobiotin and biotinyl- s-aminocaproic acid-N-hydroxysuccinimide ester, sulfo- succinimide-iminobiotin, biotinbromoacetylhydrazide, p-diazobenzoyl biocytin, 3-(N- maleimidopropionyl) biocytin. “Streptavidin” can refer to a protein or peptide that can bind to biotin and can include: native egg-white avidin, recombinant avidin, deglycosylated forms of avidin, bacterial streptavidin, recombinant streptavidin, truncated streptavidin, and / or any derivative thereof.
[0052] The terms “subject”, “patient” or “host”, used interchangeably herein, refer to any subject, particularly a vertebrate subject, and even more particularly a mammalian subject, from which a sample may be obtained for sequencing analysis. Suitable vertebrate animals that fall within the scope of the invention include, but are not restricted to, any member of the subphylum Chordata including primates (e.g., humans, monkeys and apes, and includes species of monkeys such from the genus Macaca (e.g., cynomologus monkeys such as Macaca fascicularis, and / or rhesus monkeys {Macaca mulatto)) and baboon \ Paplo ursinus), as well as marmosets (species from the genus Callithrix), squirrel monkeys (species from the genus Saimiri) and tamarins (species from the genus Saguinus), as well as species of apes such as chimpanzees (Pan troglodytes)), rodents (e.g., mice rats, guinea pigs), lagomorphs (e.g., rabbits, hares), bovines (e.g., cattle), ovines (e.g., sheep), caprines (e.g., goats), porcines (e.g., pigs), equines (e.g., horses), canines (e.g., dogs), felines (e.g., cats), avians (e.g., chickens, turkeys, ducks, geese, companion birds such as canaries, budgerigars etc.), marine mammals (e.g., dolphins, whales), reptiles (snakes, frogs, lizards etc.), and fish. A preferred subject is a mammal with a viral infection, or one suspected to have a viral infection (including a latent infection), or one who is suffering from a disease or condition caused by or secondary to a viral infection. Tt will be understood that the aforementioned terms do not imply that symptoms are present in the subject.
[0053] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (or).
[0054] As used in this application, the singular form “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “an agent” includes a plurality of agents, including mixtures thereof. By “about” is meant a quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length that varies by as much 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 % to a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight or length.
[0055] Throughout this specification and the claims which follow, unless the context requires otherwise, the word “comprise”, and variations such as “comprises” and “comprising”, will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.
[0056] Throughout this specification and the claims which follow, unless the context requires otherwise, the phrase “consisting essentially of”, and variations such as “consists essentially of’ will be understood to indicate that the recited element(s) is / are essential i.e. necessary elements of the invention. The phrase allows for the presence of other non-recited elements which do not materially affect the characteristics of the invention but excludes additional unspecified elements which would affect the basic and novel characteristics of the method defined.
[0057] Viruses
[0058] The methods described herein can be used to sequence and / or analyze the genomes of various viruses, including both DNA and RNA viruses. In the case of RNA viruses, viral RNA may be reverse transcribed into complementary DNA (cDNA) prior to analysis. The methods that use long read sequencing arc particularly suitable for sequencing and analyzing viral genomes that contain repetitive elements or complex structural variations. Examples of such viruses include herpesviruses (e.g., Epstein-Barr virus, herpes simplex virus, varicella zoster virus, etc.), poxviruses (e.g., mpox virus), and African swine fever virus.
[0059] In some embodiments, the methods herein are used to sequence or analyze a DNA virus. DNA viruses include double- stranded DNA (dsDNA) viruses (e.g., herpesviruses, poxviruses, papillomaviruses, polyomaviruses, and adenoviruses), single-stranded DNA (ssDNA) viruses (e.g., parvoviruses), and viruses with a dsDNA genome that has an RNA intermediate (dsDNA-RT) (e.g., hepadnaviruses). In some embodiments, methods herein are used to sequence or analyze a herpesvirus, including but not limited to Epstein-Barr virus (EB V), herpes simplex virus (HS V), varicella zoster virus (VZV), cytomegalovirus (CMV), and human herpesvirus (HHV). In one embodiment, methods herein are used to sequence or analyze Epstein-Barr virus (EBV).
[0060] Samples and sample processing
[0061] A sample herein may be from a subject (such as a subject with a viral infection or suspected to be infected with a virus), or an environmental sample. Appropriate samples include any environmental or biological samples, including clinical samples obtained from a human or veterinary subject. Suitable samples include but are not limited to, bodily fluids (for example, blood, serum, urine, cerebrospinal fluid, middle ear fluids, bronchoalveolar lavage, tracheal aspirates, sputum, nasopharyngeal aspirates, oropharyngeal aspirates, or saliva), mid-turbinate swabs, nasopharyngeal swabs, oropharyngeal swabs, eye swabs, cervical swabs, vaginal swabs, rectal swabs, stool and stool suspensions, hair, cells (e.g., buccal cells), tissues, and autopsy samples. Suitable samples also include environmental samples including, but not limited to, food, water, surface swabs, or other materials that may contain or be contaminated with a virus.
[0062] The sample may undergo pre-treatment or pre-handling prior to nucleic acid extraction and / or nucleic acid fragmentation. Methods suitable for pre-processing of samples are known in the art and include, but are not limited to, sonication of cells with or without the presence of grinding particles, mixing (e.g., vortexing), high powered agitation with grinding particles (bead beating, ball mills) or the use of high pressure mechanical shearing (e.g., French pressure cell press, as is known in the art). Further, enzymatic methods that use particular enzymes such as zymolase to weaken the cell walls may be used. Further still, if a particular' cell type is being targeted, isolation of that particular' cell type from a larger population may be preferred.
[0063] The present methods arc also suitable for use on fixed cells and tissues. Non-limiting examples of fixatives and fixation procedures include, for example, crosslinking fixatives (e.g., aldehydes, such as glutaraldehyde, formaldehyde (formalin), etc.); precipitating (or denaturing) fixatives such as methanol, ethanol, acetic acid and acetone; and other noncrosslinking fixatives such as the PAXgene system from Qiagen. In one embodiment, the sample is a degraded nucleic acid sample. As used herein a “degraded” nucleic acid sample refers to a sample containing nucleic acid molecules which have undergone structural alteration (e.g., nicks, breaks, fragmentation) or chemical modification (e.g., crosslinking) prior to sample collection or as a result of sample processing, resulting in a loss of integrity relative to native cellular' or genomic nucleic acid. Degraded nucleic acid samples may be characterised by a DNA Integrity Number (DIN) of less than 7 as determined using the Tapestation system. Degraded nucleic acid samples include, without limitation, samples which have been chemically fixed (such as, e.g., formalin-fixed paraffin-embedded (FFPE) samples and PAXgene-fixed cyro-embedded (PFCE) samples); circulating cell-free DNA (cfDNA) samples (e.g., from plasma or lymph); urine samples; stool samples; sputum, oral and respiratory samples; mucosal swab samples; and samples with high microbial loads.
[0064] Methods herein may comprise detecting the DNA integrity in the sample. DNA integrity and the extent of DNA degradation may be determined by gel or capillary electrophoresis, as is well known in the art. A degraded DNA sample may exhibit fragment lengths predominantly below 10 kb, and may display electrophoretic profiles showing smearing or multiple lower molecular weight peaks. Commercial systems such as the TapeStation system from Agilent may be used to assess DNA integrity or degradation. The TapeStation system employs an algorithm to assess features of the electrophoretic trace, including fragment distribution and smear patterns, to generate a DNA Integrity Number (DIN) for the sample. The DIN may range from 1 (highly degraded) to 10 (highly intact).
[0065] In some embodiments, methods herein comprise detecting the viral copy number in the sample. Viral copy number may be determined by detecting the copy number of a viral gene using, for example, quantitative PCR (qPCR) or digital PCR (dPCR). The gene for copy number detection may be a conserved gene that is specific to the virus and is preferably present in single-copy in the viral genome. Examples of such genes include DNA or RNA polymerase genes, and genes encoding nuclcocapsid proteins, capsid proteins or envelope proteins. The inventors have found that qPCR assays based on amplification of the BALF5 gene (coding for a viral DNA polymerase subunit) is particularly useful for detecting EBV copy number. Samples herein may contain both viral DNA and DNA from a host (e.g., a vertebrate, mammalian or human host). In one embodiment, the methods of this disclosure comprise extracting or isolating DNA from the sample prior to DNA fragmentation. Generally, nucleic acids can be extracted from a sample by a variety of techniques such as those described by Maniatis, et al., Molecular Cloning: A Laboratory Manual, 1982, Cold Spring Harbor, NY, pp. 280-281, or Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3rdEd, Cold Spring Harbor Laboratory Press, 2001, Cold Spring Harbor, NY.
[0066] In some embodiments, viral DNA in the sample is amplified. The amplification reaction may be any amplification reaction known in the art that amplifies nucleic acid molecules, such as polymerase chain reaction (PCR), nested PCR, PCR-single strand conformation polymorphism, ligase chain reaction, multiple displacement amplification (MDA) (which encompasses and is used interchangeably with the term strand displacement amplification (SDA)), restriction fragment length polymorphism, transcription- based amplification system, rolling circle amplification, and hyper-branched rolling circle amplification. Further examples of amplification techniques that can be used include, but are not limited to, quantitative PCR, quantitative fluorescent PCR (QF-PCR), multiplex fluorescent PCR (MF- PCR), real time PCR (RTPCR), single cell PCR, restriction fragment length polymorphism PCR (PCR-RFLP), and RT-PCR-RFLP. Other suitable amplification methods include transcription amplification, self-sustained sequence replication, selective amplification of target polynucleotide sequences, consensus sequence primed polymerase chain reaction (CP-PCR), arbitrarily primed polymerase chain reaction (AP-PCR), degenerate oligonucleotide-primed PCR (DOP-PCR) and nucleic acid based sequence amplification (NABS A).
[0067] DNA is fragmented to produce fragments of suitable lengths for sequencing and analysis. In some embodiments, shear forces created during tissue homogenisation or cell lysis will mechanically generate nucleic acid fragments. Further fragmentation or shearing to give fragments of desired length may use a variety of mechanical, chemical and / or enzymatic methods known in the art. For example, DNA may be randomly sheared via sonication (e.g., using a Covaris ultra sonicator), hydrodynamic shearing (e.g., using a hydroshear instrument), brief exposure to a DNase, or using a mixture of one or more restriction enzymes, transposases or nicking enzymes. RNA may be fragmented by brief exposure to an RNase, heat plus magnesium, or by shearing. If fragmentation is employed, the RNA may be converted to cDNA before or after fragmentation. Generally, fragmented nucleic acid molecules can be from about 100 base pairs (bp) to about 500 kilobases (kb) or more.
[0068] Mechanical fragmentation methods have the advantage of producing fragments of a particular size range in a predictable manner. Enzymatic fragmentation methods can also be used to generate nucleic acid fragments, particularly shorter fragments of 300-500 bp in size. Enzymatic fragmentation methods typically include the use of endonucleases, and arc more amenable than mechanical fragmentation methods to multi-sample or high-throughput processing. However, enzymatic fragmentation methods are inherently prone to variability in the degree of fragmentation, because to achieve consistent fragment size distributions in such methods requires extremely careful control of enzyme activity, substrate amounts and concentrations, and digestion time.
[0069] In some cases, particularly when it is desired to isolate long fragments (such as fragments from about 10 to about 700 kb in length), cells may be lysed and the intact nuclei are pelleted with a gentle centrifugation step. The nucleic acid, usually genomic DNA, is released through enzymatic digestion, using for example proteinase K and RNase digestion over several hours. The resultant material is then dialyzed overnight or diluted directly to lower the concentration of remaining cellular waste. Since such methods of isolating the nucleic acid does not involve many disruptive processes (such as ethanol precipitation, centrifugation, and vortexing), the genomic nucleic acid remains largely intact, yielding a majority of fragments in excess of 100 kb.
[0070] In some embodiments, fragments of a particular size or in a particular size range are isolated. Such isolation methods are well known in the art. For example, gel fractionation can be used to produce a population of fragments of a particular size range, for example 1000 bp ± 50 bp.
[0071] In one embodiment, shorter fragments are used in the methods of this disclosure. Such shorter fragments may range in size from about 50 to about 1000 nucleotides in length, for example, about 50, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, or about 1000 nucleotides in length. In some embodiments, shorter fragments may be about 300 to about 500 nucleotides in length. Methods herein may comprise performing short read sequencing on shorter nucleic acid fragments. In one embodiment, longer fragments are used in the methods of this disclosure. Such longer fragments may range in size from about 1000 to about 100,000 nucleotides (100 kb) in length, for example, about 1000, about 2000, about 3000, about 4000, about 5000, about 6000, about 7000, about 8000, about 9000, about 10,000, about 15,000, about 20,000, about 25,000, about 30,000, about 35,000, about 40,000, about 45,000, about 50,000, about 55,000, about 60,000, about 65,000, about 70,000, about 75,000, about 80,000, about 85,000, about 90,000, about 95,000, or about 100,000 nucleotides in length. In further embodiments, longer fragments may be about 150 kb; about 200 kb; about 250 kb; about 300 kb; about 350 kb; about 400 kb; about 450 kb; about 500 kb; about 600 kb; about 700 kb; about 800 kb; about 900 kb; about 1000 kb; or about 1500 kb in length. Methods herein may comprise performing long read sequencing on longer nucleic acid fragments.
[0072] In one embodiment, the method comprises fragmenting DNA to a size distribution of about 300-500 bp for short read sequencing. In one embodiment, the method comprises fragmenting DNA to a size distribution of about 10 kb for long read sequencing.
[0073] Fragments from various samples may be pooled to increase sequencing throughput. Pooling may be based on, for example, the detected quantity or quality of viral DNA in the sample. In some embodiments, the detected viral copy number and / or DNA integrity of the sample is used for sample or fragment pooling. For example, fragments from samples with at least 3000 viral copies per 500 ng of sample DNA and / or a DNA integrity number (DIN) of at least 7 may be pooled for long read sequencing. Fragments from samples with viral copy number between 250 and 3000 per 500 ng of sample DNA may be pooled for short read sequencing.
[0074] Capture probes and fragment enrichment
[0075] The methods herein may include a step of enriching for virus-specific DNA fragments to reduce the contribution of host DNA to the sequence reads and thus increase the efficiency of sequence analysis. This is particularly advantageous when analysing samples containing only small quantities of viral genetic material. Capture probes may be utilized for fragment enrichment. As used herein, the term “capture probe” refers to a nucleic acid molecule comprising a sequence complementary to at least a portion of a target nucleic acid. The degree of complementarity between a probe and a corresponding target sequence can vary with the application. In some embodiments, the capture probe can be complementary or substantially complementary to a target sequence or portions thereof. For example, a capture probe can comprise a sequence having a complementarity to a corresponding target sequence of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%>, at least about 97%, or at least about 99%.
[0076] In certain embodiments, the viral- specific capture probe comprises at least about 50 nucleotides. In some embodiments, the viral- specific capture probe comprises at least about 100 nucleotides. In some embodiments, the viral-specific capture probe comprises at least about 150 nucleotides. In some embodiments, the viral- specific capture probe comprises at least about 200 nucleotides. In some embodiments, the viral-specific capture probe comprises at least about 250 nucleotides. In some embodiments, the viral- specific capture probe comprises at least about 300 nucleotides. In some embodiments, the viral-specific capture probe is about 50 to about 500 bp in length. In some embodiments, the viral-specific capture probe is about 50 to about 400 bp in length. In some embodiments, the viral-specific capture probe is about 50 to about 300 bp in length. In some embodiments, the viral-specific capture probe is about 50 to about 200 bp in length. In some embodiments, the viral-specific capture probe is about 100 to about 200 bp in length. In some embodiments, the viral-specific capture probe is about 100 to about 150 nucleotides in length. In one embodiment, the viral- specific capture probe is about 120 nucleotides in length.
[0077] A plurality of virus-specific capture probes may be designed to target various sections of the viral genome, such as particular coding or repeat regions, or specific genes. The probes may also be designed to target across the entire viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 50% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 60% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 70% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 75% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 80% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 85% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 90% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 91% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 92% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 93% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 94% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 95% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 96% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 97% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 98% of the viral genome. In some embodiments, the plurality of viral-specific capture probes target nucleic acid sequences across at least 99% of the viral genome. In one embodiment, the plurality of viral- specific nucleic acid capture probes target nucleic acid sequences across the entire viral genome (i.e., 100% of the viral genome).
[0078] The target sequences of a panel of probes may be substantially evenly distributed throughout the target region (e.g., substantially evenly distributed throughout the viral genome), or they may be concentrated in particular segments, such as in parts of coding or repeat regions of interest or in regions known to have low sequencing coverage.
[0079] The offset between probes refers to the distance between the start of one probe and the start of an adjacent probe. Panels of probes may be arranged with a fixed or with variable offsets. The probes may be end-to-end tiled (i.e., the offset equals the probe length); tiled with gaps (i.e., the offset is greater than the probe length); or overlapping (i.e., the offset is shorter than the probe length). The skilled person may design panels of probes to target portions of or the entirety of a given viral genome based on known principles of probe design. By way of nonlimiting example, probes used for fragment enrichment of short DNA fragments (e.g, fragments less than 500 bp) may be overlapping or end-to-end tiled with no gaps between adjacent probe target sequences. The enriched short fragments may be used for short read sequencing. Probes used for fragment enrichment of long DNA fragments (e.g., fragments longer than 1 kb) may be tiled with a gap between adjacent probes that is at least the length of the probe. The enriched long fragments may be used for long read sequencing.
[0080] The capture probe sequences may be designed based on a reference viral genome that contains sequence information from a plurality of viral genomes. For example, for EBV sequencing, such a reference viral genome may comprise sequence information from EBV genomes from two or more cell lines selected from B95-8, C17, NPC43, Raji, C666-1, M81, YCC-lO and SNU-719.
[0081] In some embodiments, nucleic acid capture probes for EBV sequencing comprise at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity (such as about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% sequence identity) to a nucleic acid sequence in SEQ ID NO: 9-76. The sequences set out in SEQ ID NO: 9-76 are derived from EBV genomes which are newly assembled by the inventors using long read sequencing (Table A).
[0082] In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 9-17. In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80%> sequence identity to a nucleic acid sequence in SEQ ID NO: 18-28. In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 29-33. In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 34-42. In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80%> sequence identity to a nucleic acid sequence in SEQ ID NO: 43-45. In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 46-54. In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 55-67. In some embodiments, a panel of nucleic acid capture probes is used for EBV sequencing that comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 68-76.
[0083] The panel of nucleic acid capture probes may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1 , 12, 13,
[0084] 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37,
[0085] 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61,
[0086] 62, 63, 64, 65, 66, 67 or 68 nucleic acid sequences selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 9-76.
[0087] Table A. Exemplary capture probe sequences
[0088] In certain embodiments, a capture probe includes a label. The label may be a reporter moiety or an affinity tag. In some embodiments, the capture probe comprises an affinity tag. Affinity tags can be useful for the bulk separation of target nucleic acids hybridized to capture probes. As used herein, the term “affinity tag” refers to a component of a multi-component complex, wherein the components of the multi-component complex specifically interact with or bind to each other. For example, an affinity tag can include biotin that can bind streptavidin. Other examples of multiple-component affinity tag complexes include ligands and their receptors; binding protcins / pcptidcs, including maltosc-maltosc binding protein (MBP), calciumcalcium binding protein / peptide (CBP); antigen-antibody, including epitope tags, such as c- MYC, HA, VSV-G, HSV, V5, and FLAG Tag™, and their corresponding anti-epitope antibodies; haptens, for example, dinitrophenyl and digoxigenin, and their corresponding antibodies; aptamers and their corresponding targets; fluorophores and anti-fluorophore antibodies; and the like.
[0089] In some embodiments, the affinity tag on the capture probe is biotin or a derivative thereof. Derivatives of biotin include but are not limited to 2-iminobiotin, desthiobiotin, Neutr Avidin (Molecular Probes, Eugene, Oreg.), CaptAvidin (Molecular Probes), and the like.
[0090] In certain embodiments, a capture probe comprises a reporter moiety. The skilled person will appreciate that many different species of reporter moieties can be used with the methods herein, either individually or in combination with one or more different reporter moieties. In certain embodiments, a reporter moiety can emit a signal. Examples of signals include a fluorescent, a chemiluminescent, a bioluminescent, a phosphorescent, a radioactive, a colorimetric, or an electrochemiluminescent signal. Example reporter moieties include fluorophores, radioisotopes, chromogens, enzymes, antigens including epitope tags, semiconductor nanocrystals such as quantum dots, heavy metals, dyes, phosphorescent groups, chemiluminescent groups, electrochemical detection moieties, binding proteins, phosphors, rare earth chelates, transition metal chelates, near-infrared dyes, electrochemiluminescent labels, and mass spectrometer-compatible reporter moieties, such as mass tags, charge tags, and isotopes. More reporter moieties that may be used with the methods and compositions described herein include spectral labels such as fluorescent dyes (c.g., fluorescein isothiocyanate, Texas red, rhodamine, and the like), radiolabels (c.g., 3H, 1251, 35S, 14C, 32P, 33P, etc.), enzymes (e.g., horse-radish peroxidase, alkaline phosphatase etc.) spectral calorimetric labels such as colloidal gold or colored glass or plastic (e.g. polystyrene, polypropylene, latex, etc.) beads; magnetic, electrical, thermal labels; and mass tags. Reporter moieties can also include enzymes (horseradish peroxidase, etc.) and magnetic particles. More reporter moieties include chromophores, phosphors and fluorescent moieties.
[0091] Capture probes may be prepared by a variety of methods. In some embodiments capture probes may be prepared in situ. For example, affinity tags, reporter moieties, and / or cleavable moieties can be incorporated into a capture probe hybridized to a nucleic acid. In some such embodiments, a capture probe hybridized to a target nucleic acid may be extended with nucleotides or nucleotide analogues that can comprise affinity tags, reporter moieties, or cleavable nucleotides.
[0092] In some embodiments, methods herein further comprise capturing viral DNA fragments that are hybridized to the viral- specific nucleic acid capture probe onto a substrate, for example, through interaction of an affinity tag on the capture probe with a complementary tag on the substrate. This can allow the hybridized fragments to be enriched from other unhybridized and / or unassociated nucleic acids. Examples of substrates include microspheres, planar surfaces, columns, and the like. By “microsphere” or “bead” or “particle” is meant a small discrete particle. The composition of the substrate will vary depending on the application. Suitable compositions include those used in peptide, nucleic acid and organic moiety synthesis, including, but not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, carbon graphite, titanium dioxide, latex or crosslinked dextrans such as Sepharose, cellulose, nylon, cross-linked micelles and Teflon may all be used. The beads need not be spherical; irregular particles may be used. In some embodiments, a substrate can comprises a metallic composition, e.g., ferrous, and may also comprise magnetic properties. An example embodiment utilizing magnetic beads includes capture probes comprising streptavidin-coated magnetic beads. In addition, the beads may be porous, thus increasing the surface area of the bead available for association with capture probes. The bead sizes can range from nanometers (e.g., 100 nm) to millimetres (e.g., 1 mm), with beads from about 0.2 pm to about 200 pm being preferred, and from about 0.5 pm to about 5 pm being particularly preferred, although in some embodiments smaller beads may be used. In some embodiments, the capture probe itself is associated with or immobilized on a substrate. In such embodiments, target nucleic acids may be directly captured onto a substrate through hybridization with the capture probe.
[0093] Some methods to enrich target nucleic acids that are associated with capture probes can include dissociating the target nucleic acid from at least a portion of a capture probe. In some embodiments, dissociating the target nucleic acid from at least a portion of a capture probe can be performed subsequent to removing nucleic acids not associated with a capture probe. As will be understood, methods to disassociate target nucleic acids from at least a portion of a capture probes will vary according to the type of association between the target nucleic acid and capture probe. In some embodiments target nucleic acids can be disassociated from at least a portion of a capture probe by denaturing nucleic acids, e.g., by increasing temperature. In some embodiments, a target nucleic acid can be disassociated from at least a portion of a capture probe by cleaving a cleavable linker. In some embodiments, a target nucleic acid can be disassociated from at least a portion of a capture probe by digesting at least a portion of the capture probe, e.g., RNA capture probes can be digested with RNAse. Some embodiments also include removing at least a portion of a capture probe disassociated from a target nucleic acid from the disassociated target nucleic acid by methods well known in the art, e.g., washing.
[0094] Some embodiments to enrich target nucleic acids associated with capture probes can include one or more rounds of enrichment.
[0095] Library preparation and sequencing
[0096] A skilled person may use known methods in the art to prepare a DNA sequencing library from the enriched viral DNA fragments. Library preparation may vary depending on the sequencing platform, the type of nucleic acid being sequenced (DNA or RNA), and the application (e.g., whole-genome sequencing, RNA-Seq, etc). Commercial kits may be available for specific sequencing platforms or technologies, and may be used with or adapted for the methods of this disclosure. Library preparation generally involves a combination of the following: a) end repair to create blunt 5’ and 3’ ends; b) creation of 3’ tails (typically containing one or a pair of A nucleotides) using, e.g., a non-proofreading polymerase (e.g., Taq polymerase) in preparation for ligation of a complementary-tailed adaptor; c) ligation of a sequencing adapter to the fragments using, e.g., a ligase, to allow for capture, identification, amplification and / or sequencing on the sequencing platform; d) library amplification to increase the available DNA for sequencing; e) quality control (QC) to assess fragment size distribution, contaminant levels and / or DNA concentration; and f) pooling of multiple indexed libraries (multiplexing) for sequencing. The above list is non-exhaustive.
[0097] In some embodiments, generating a population of DNA fragments comprises: a) fragmenting DNA in the sample; b) ligating fragmented DNA to nucleic acid adaptors; and c) amplifying ligated DNA to obtain a population of DNA fragments.
[0098] The sequencing adapters can include a unique molecular identifier (UMI) sequence (also known as an index or barcode sequence) to differentiate fragments derived from different samples and to allow simultaneous sequencing of multiple libraries (multiplexing) within the same sequencing inn. Sequencing adapters may also contain a universal priming site and / or one or more sequencing platform- specific sequences for use in subsequent cluster generation and / or sequencing (e.g., known P5 and P7 sequences). Library amplification may be carried out using known methods for nucleic acid amplification. Primers for amplification may be suitably designed to target priming sites in the ligated adapters.
[0099] The prepared DNA libraries are sequenced on long or short read sequencing platforms. As used herein “long read sequencing” refers to the production of sequencing reads with read lengths that arc 1 kilobasc (1 kb) or longer, such as several kilobascs to hundreds of kilobases. Platforms that arc capable of long read sequencing include but are not limited to the single-molecule real-time (SMRT) sequencing platform from Pacific Bioscicnccs, nanopore sequencing from Oxford Nanopore Technologies and Quantapore, the Complete Long Reads platform from Illumina, and sequencing by expansion (SBX) from Stratos Genomics. “Short read sequencing” refers to the production of sequencing reads with read lengths that are 500 nucleotides or shorter, such as 300 to 500 nucleotides. Platforms that are capable of short read sequencing include but are not limited to the Illumina and MGI Tech platforms.
[0100] Methods herein also provide for sequencing using platforms with adaptive sampling. “Adaptive sampling” in the context of nucleic acid sequencing refers to a technique that dynamically adjusts the sequencing depth or coverage based on the properties of the sample being sequenced, such as the presence of a specific region or sequence of interest in a nucleic acid sample. Examples of long read sequencing platforms capable of adaptive sampling include the Oxford Nanopore Adaptive Sampling (ONT-AS) system, which can be configured to reject nucleic acids with sequences that fall either within or outside of predetermined regions of interest.
[0101] The quality of the DNA in the sample may be used to determine the type of sequencing to be performed. In some embodiments, methods herein comprise performing short read sequencing on degraded nucleic acid samples. A degraded nucleic acid sample may be characterised by a DNA integrity number (DIN) that is less than 7, as determined using the Agilent TapeStation system. In some embodiments, methods herein may comprise performing long read sequencing on non-degraded nucleic acid samples. A non-degraded nucleic acid sample may be characterised by a DNA integrity number (DIN) that is 7 or higher, as determined using the Agilent TapeStation system.
[0102] The type of sequencing may also be selected based on viral copy number in the sample. The inventors have established copy number thresholds for determining when long read or short read sequencing would be suitable for EBV genome analysis. Thus, in some embodiments, methods herein may comprise performing long read sequencing when the viral copy number in the sample is at least 3000 copies per 500 ng of sample DNA. The inventors have found that a copy number of at least 3000 copies per 500 ng of DNA is sufficient to provide adequate genome coverage and to detect variants (especially structural and copy number variants) at sites throughout the EBV genome on most long read sequencing platforms. Specific long read sequencing platforms may be chosen based on the viral copy number and the quality of the nucleic acid in the sample. For example, samples with a viral load of at least 200,000 copies per 500 ng of DNA and a DNA integrity number (DIN) of at least 7 as determined using the Agilent TapeStation system may be sequenced using the Oxford Nanopore Adaptive Sampling (ONT-AS) platform, which does not require target enrichment for sequencing. Samples with a viral load of between about 3000 and 200,000 copies per 500 ng of DNA may be sequenced using long read sequencing without adaptive sampling, optionally in combination with target enrichment using viral-specific probes. Samples with a viral copy number that is greater than 250 copies but less than 3000 copies per 500 ng of sample DNA may be sequenced using short read sequencing, optionally in combination with target enrichment.
[0103] In some embodiments, samples are selected for short read sequencing when: a) the viral copy number in the sample is more than 250 copies but less than 3000 copies per 500 ng of sample DNA; or b) the viral copy number in the sample is more than 250 copies but less than 3000 copies per 500 ng of sample DNA, and the DNA integrity number (DIN) is less than 7.
[0104] In some embodiments, samples are selected for long read sequencing when: a) the viral copy number in the sample is at least 3000 copies per 500 ng of sample DNA; b) the DNA integrity number (DIN) is at least 7; or c) the viral copy number is at least 3000 copies per 500 ng of sample DNA, and the DNA integrity number (DIN) is at least 7.
[0105] In some embodiments, samples are selected for long read adaptive sampling when the viral copy number is at least 200,000 copies per 500 ng of sample DNA, and the DNA integrity number (DIN) is at least 7.
[0106] The DIN may be determined using the Agilent TapeStation system.
[0107] Reference genomes
[0108] The methods herein may comprise the use of a reference viral genome for genome analysis. In some embodiments, the reference viral genome comprises sequence information of viral genomes from a plurality of viral strains. The reference viral genome may be a pangenome reference or pangenome model constructed from the plurality of viral genome sequences. The terms “pangenome reference” and “pangenome model” are used interchangeably herein to refer to a data structure that represents genomic sequences and variations present in different samples or populations of a viral species. The pangenome reference includes sequences common to all input genome sequences as well as information about the position, alleles, and frequencies of each variant site within the input sequences. The genomic data may be represented graphically. The pangenome reference can thus serve as a central coordinating entity to describe the collection of viral genomes. Genome analysis based on a single reference genome (such as a linear consensus sequence) may be disadvantageous as it can lead to genome assemblies that appear to be more similar to the reference than they actually are, and may also underestimate the prevalence of structural variants. Pangenomic reference systems can reduce this bias by enabling a new genome to be directly related to all those represented in the pangenome.
[0109] The viral genomes used to construct a pangenome reference are preferably derived from a variety of sources, such as from different patient populations (e.g., patients with different severity of disease or different viral-induced secondary diseases, patients of different ethnicities, patients from different geographic locations, etc.), or from different viral strains, serotypes, subtypes or genetic variants.
[0110] In embodiments where EBV viral genomes are analyzed or sequenced, the reference viral genome may comprise sequence information from EBV genomes derived from EBV- infected cells or cell lines. By a viral genome that is “derived from” a cell or cell line is meant viral genomic DNA that is extracted from the cell or cell line. Sequence information from viral genomes may be obtained using methods for genome sequencing generally known the art. Alternatively, the methods of this disclosure may be used to sequence the viral genomes.
[0111] In some embodiments, the reference viral genome comprises sequence information from EBV genomes derived from two or more cell lines selected from B95-8, C17, NPC43, Raji, C666-1 , M81 , YCC-10 and SNU-719. These cell lines are derived from subjects with different EBV-induced diseases, including EBV-induced malignancies. Without being bound by theory, a pangenome reference containing EBV strains associated with EBV- induced diseases may be used to detect strains that predispose to particular' diseases.
[0112] The reference viral genome may contain sequence information from EBV genomes derived from 1 cell line, 2 cell lines, 3 cell lines, 4 cell lines, 5 cell lines, 6 cell lines, 7 cell lines or all 8 cell lines selected from B95-8, C17, NPC43, Raji, C666-1 , M81 , YCC-10 and SNU- 719. The reference viral genome may additionally contain sequence information from other EBV genomes, such as EBV genomes derived from other EBV-positive cells or cell lines. In some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the B95-8 cell line. Tn some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the C17 cell line. In some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the NPC43 cell line. In some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the Raji cell line. In some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the C666-1 cell line. In some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the M81 cell line. In some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the YCC-10 cell line. In some embodiments, the reference viral genome comprises sequence information from the EBV genome derived from the SNU-719 cell line.
[0113] In some embodiments, the reference viral genome comprises an EBV genome sequence selected from the group consisting of SEQ ID NO: 1-8. In some embodiments, the reference viral genome comprises two or more EBV genome sequences selected from SEQ ID NO: 1-8, such as 2, 3, 4, 5, 6, 7, or 8 EBV genome sequences selected from SEQ ID NO: 1-8. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 1. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 2. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 3. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 4. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 5. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 6. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 7. In some embodiments, the reference viral genome comprises the EBV genome sequence of SEQ ID NO: 8.
[0114] Genome analysis
[0115] Methods herein may be used for viral genome analysis. Genome analysis includes but is not limited to genome assembly (e.g., de novo genome assembly); genome annotation (e.g., annotation of coding sequences, non-coding regions, repetitive elements, regulatory elements, promoters, enhancers, and the like); genetic variation analysis (including, e.g., the identification and characterization of single nucleotide variations, structural variations, copy number variations, etc.); epigenomic analysis (including, e.g., methylation profiling, base modification detection, etc.); and evolutionary genomics (e.g., comparative genomic analyses, phylogenetic reconstructions, etc.).
[0116] In some embodiments, the method comprises aligning the sequencing reads to assemble a viral genome. For example, the long read sequencing reads may be aligned to assemble a viral genome de novo without using a reference genome. Aligned overlapping reads are also known as contigs. Longer sequencing reads are advantageous in helping to bridge gaps between contigs, resolve repetitive elements and segmental duplications, and generally improve the assembly of complex genomic regions.
[0117] In other embodiments, the method herein comprises comparing the sequencing reads to a reference viral genome to analyze the viral genomic DNA. In such embodiments, longer sequencing reads may be advantageous for detecting genetic features which may be missed or incorrectly characterized by short read sequencing methods. Such genetic features may include insertion / deletion (indel) mutations and otherstructural variations, regulatory elements, and epigenetic markers.
[0118] In some embodiments, analyzing the viral genomic DNA comprises detecting: i) the presence or absence, ii) the strain, iii) the subtype or iv) the genetic variant of the virus. The sequencing reads may be aligned to a reference viral genome to detect the presence or absence of a virus, or the strain, subtype or genetic variant of the virus. Certain viral strains, subtypes, serotypes or variants may predispose a subject to a secondary disease (such as a cancer), thus in some embodiments, the methods herein may be used to detect a subject who is at risk of a disease (such as a cancer) based on the presence of a particular viral strain, subtype, serotype or variant in a sample from the subject.
[0119] In some embodiments, analyzing the viral genomic DNA comprises detecting sequence variations in the genome. Sequence variations include but are not limited to single nucleotide polymorphisms (SNPs), structural variations (SVs) such as insertions, deletions, duplications, inversions and translocations, and other complex rearrangements. In some embodiments, analyzing the viral genomic DNA comprises detecting epigenetic variations in the genome. Epigenetic variations include but are not limited to methylated bases, hydroxymethyl ated bases, and oxidized bases.
[0120] Method of detecting virus-associated diseases
[0121] Disclosed herein is a method of detecting a subject who is at risk of a cancer; the method comprising: a) generating a population of DNA fragments from a sample obtained from the subject; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral-specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample, so as to determine whether the subject is at risk of a cancer. In one embodiment, the viral DNA is from EBV.
[0122] In some embodiments, the method comprises comparing the long read sequencing reads to a reference viral genome to detect: i) the presence or absence, ii) the strain, iii) the subtype or iv) the genetic variant of the virus to determine whether the subject is at risk of a cancer. In some embodiments, the virus is one which is associated with a cancer, such as a virus that is known to cause a cancer or predispose a subject to a cancer. In one embodiment, the virus is EBV. In some embodiments, the reference viral genome is an EBV reference genome. In some embodiments, the reference viral genome comprises sequence information of viral genomes from a plurality of viral strains, including but not limited to the EBV genome sequences set forth in SEQ ID NO: 1-8.
[0123] In some embodiments, the subject is one who is suffering from cancer or is suspected of having cancer. The subject may be one who has suffered from a previous viral infection, such as a viral infection that is associated with a cancer.
[0124] The cancer may be any type of cancer, and may be benign or malignant. In embodiments where the subject has suffered from a previous EBV infection, the cancer may be a lymphoma or a lymphoproliferative disorder, such as Hodgkin’s lymphoma, Burkitt’s lymphoma, T / NK cell lymphoma and B-cell lymphoma, or a carcinoma, such as nasopharyngeal cancer, gastric carcinoma, breast cancer, cervical cancer, thyroid cancer, salivary gland cancer, lung cancer, renal cancer, or bladder cancer.
[0125] Kits
[0126] Disclosed herein is a kit for performing a method of this disclosure.
[0127] In some embodiments, the kit comprises one or more viral- specific nucleic acid capture probes. In some embodiments, the kit comprises a plurality of viral-specific nucleic acid capture probes. In some embodiments, the plurality of viral-specific nucleic acid capture probes target nucleic acid sequences across at least 50% of the viral genome, such as across at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the viral genome. In some embodiments, the plurality of viral-specific nucleic acid capture probes target nucleic acid sequences across the entire viral genome.
[0128] In some embodiments, the viral- specific capture probe is about 100-140 nucleotides in length. In some embodiments, the viral- specific nucleic acid capture probe comprises an affinity tag. In one embodiment, the affinity tag is biotin or a biotin derivative. In some embodiments, the kit may also comprise a substrate comprising a moiety capable of binding to the affinity tag on the capture probe. For example, the kit may contain a streptavidin column or streptavidin-coated beads or particles for capturing biotinylated probes.
[0129] Disclosed herein is a kit for analyzing EBV genomic DNA, the kit comprising one or more nucleic acid capture probes derived from a reference EBV genome comprising sequence information of EBV genomes from a plurality of EBV strains. In one embodiment, the nucleic acid capture probes arc derived from a reference EBV genome comprising sequence information from EBV genomes derived from two or more cell lines selected from B95-8, C17, NPC43, Raji, C666-1, M81, YCC-10 and SNU-719.
[0130] In one embodiment, the kit comprises one or more nucleic acid capture probes comprising a nucleic acid sequence selected from a nucleic acid sequence having at least 80%> sequence identity to a nucleic acid sequence in SEQ ID NO: 9-76.
[0131] In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 9-17. In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 18-28. In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 29-33. In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 34-42. In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 43-45. In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 46-54. In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 55-67. In some embodiments, the kit comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 68-76.
[0132] The kit may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67 or 68 nucleic acid sequences selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 9-76.
[0133] The nucleic acid capture probes may comprise an affinity tag as described in this disclosure. In one embodiment, the affinity tag on the capture probe is biotin or a derivative thereof, such as 2-iminobiotin, dcsthiobiotin, NcutrAvidin (Molecular Probes, Eugene, Oreg.), CaptAvidin (Molecular Probes), and the like.
[0134] The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.
[0135] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. It is to be understood that the invention includes all such variations and modifications, which fall within the spirit and scope. The invention also includes all of the steps, features, compositions and compounds referred to or indicated in this specification, individually or collectively, and any and all combinations of any two or more of said steps or features.
[0136] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0137] Certain embodiments of the invention will now be described with reference to the following examples which are intended for the purpose of illustration only and are not intended to limit the scope of the generality hereinbefore described.
[0138] EXAMPLES
[0139] EXAMPLE 1: Short Read EBV sequencing library preparation and sequencing
[0140] Materials and methods
[0141] DNA extraction from tumor tissue samples, saliva samples, FFPE slides and Plasma samples
[0142] DNA was extracted from flash frozen tumor tissues using the QIAamp DNA Mini kit (Qiagen). The DNA in saliva samples was extracted with MGIEasy Magnetic Beads Genomic DNA Extraction Kit. The DNA in formalin-fixed paraffin-embedded (FFPE) samples was extracted with QIAamp DNA FFPE Tissue Kit (Qiagen). The DNA in plasma samples was extracted with Mag-Bind® cfDNA Kit (Omega BIO-TEK). The quantity of the extracted DNA was assessed using Qubit IX dsDNA HS Assay Kit (Thermofisher Scientific; Q33231). The quality of the extracted DNA was checked using Genomic DNA ScreenTape assay (Agilent, 5067-5365). All the DNA extracted from FFPE and plasma samples was used for short read targeting library preparation. For tumor and saliva samples, only the DNA samples with a DNA Integrity Number (DIN) >7 were selected for long read targeting library preparation. The DNA samples with DIN <7 were considered as degraded DNA and were only used for short read targeting library preparation.
[0143] Viral Load
[0144] Virus load was measured by a highly sensitive qPCR assay based on amplification of a 190 bp locus in the EB V BAL 5 gene. qPCR was conducted with 1 pl extracted DNA in triplicate in a 10 pl reaction. The EBV copy in 1 pl of DNA was calculated from the respective CT using the linear equation from the respective standard curve. The inventors determined that for tumor and saliva samples with more than 3,000 EBV copies in 500 ng DNA can be selected for long-read EBV target capture, and samples with more than 200,000 EBV copies in 500 ng DNA can be selected for ONT Adaptive sampling.
[0145] The schematic above shows the sequencing methods chosen based on the DNA quality and EBV copy in the sample.
[0146] Bait Design
[0147] 120-mer double-stranded baits spanning the length of the reference genomes (accession numbers NC_007605) were designed using IDT Algorithms (Integrated DNA Technologies). The bait panel included 1567 probes. Bait libraries were synthesized by IDT. Each DNA sample is hybridized with the panel to obtain the most uniformity of EBV genome coverage.
[0148] Target Enrichment and Illumina sequencing For short-read library preparation, about 200 ng of extracted DNA was first fragmented using a Covaris microTUBE to obtain an average size distribution of about 300-500 bp. For cfDNA extracted from plasma, the DNA can be used directly for library prep. For FFPE DNA, after fragmentation, the DNA fragments need to be purified with 1.4x AMPure XP beads to remove the chemicals which could inhibit the following library prep reactions. Library prepar ation of tumor and saliva samples was done using the NEBNext® Ultra™ II End Rcpair / dA-Tailing Module (NEB, #7546) and NEBNext™ Ultra II Ligation Module (NEB, #7595) following the manufacturer’s protocol with the UMI adaptors designed by inventor. The FFPE and cfDNA library prep was done with xGen™ cfDNA & FFPE DNA Library Prep v2 MC Kit (IDT) following the manufacture’s protocol. The quality of the library was checked using Agilent DNA Ik ScreenT ape. Library DNA in the expected region of 300-500 bp indicates that the library was good and suitable for capture.
[0149] In solution hybridization
[0150] Library preparations were pooled (6-8 libraries) based on the EBV copy determined by qPCR with biotin-labeled probes. Probe-hybridized complexes were then immobilized on streptavidin-coated magnetic beads. Enriched DNA fragments were then amplified over 12- 14 cycles to obtain the requested quantity for the Illumina sequencing.
[0151] Post-capture QC
[0152] The EBV copy in the post-capture (enriched) library was measured by qPCR and compared to the EBV copy in the pooled pre-capture library to calculate capture efficiency. Libraries with capture efficiency more than 20 are sent for Illumina sequencing.
[0153] Sequencing data analysis
[0154] Fastq data from demultiplexed paired-end sequencing results were merged using seqtk and trimmed for adaptors using trimadap. The trimmed sequences were aligned to the EBV B95- 8 reference genome (NC_007605.1) using the Burrows-Wheeler Aligner (BWA, version 0.7.17). Coverage and depth were calculated using mosdepth or Samtools (version 1.17). Variants were called following the GATK best practice workflows (version 4.2.4.1).
[0155] Results
[0156] Target enrichment and sequencing The large (172-kb) EBV genome combined with the low abundance of viral DNA relative to host DNA presents a sequencing challenge. A custom target enrichment library for EBV was designed, comprising 1567 of 120-nucleotide single-stranded oligonucleotide baits tiled across the EBV genome. Among the 6 FFPE samples, the EBV reads ratio were more than 85% and EBV genome coverage were more than 1000.
[0157] Variant calling
[0158] The number of single nucleotide polymorphism (SNP) identified in each sample ranged from 1000-2000 and the number of indels ranged from 300 to 700 (Table 1).
[0159] Table 1. Sequencing data and variants calling results of 6 FFPE samples
[0160] EXAMPLE 2: Long Read EBV target sequencing library preparation and sequencing
[0161] Bait Design
[0162] 120-mer double-stranded baits spanning the length of the reference genomes (accession numbers NC_007605) were designed using Twist Algorithms (Twist Bioscience). Two bait panels were designed: the first contains 997 probes covering the entire EBV genome, and the second contains 302 probes covering IR1 and other repetitive regions. Bait libraries were synthesized by Twist Bioscience. Each DNA sample is hybridized with the two panels separately to obtain the most uniformity of EBV genome coverage.
[0163] Target Enrichment and PacBio sequencing For long-read library preparation, about 300 ng of extracted DNA was first fragmented using a Covaris g-TUBE to obtain an average size distribution of about 10,000 bp. DNA fragments less than 3,000 bp were removed by a sizing selection step performed with 0.4x AMPure PB magnetic beads. The fragmented DNA was used as input for the library preparation. Library preparation was done using the Twist Library Preparation Kit 1, Mechanical Fragmentation, 96 Samples (100876) following the manufacturer’s protocol. The quality of the library was checked using Agilent Genomic DNA ScrccnTapc. A single peak in the expected region of -10,000 bp indicates that the library was good and suitable for capture.
[0164] Tn solution hybridization
[0165] Library preparations were pooled (7-8 libraries) based on the EBV copy determined by qPCR with biotin-labeled probes. Due to the abundant GC-rich regions in the EBV genome, a temperature gradient hybridization setup (5 min at 95°C, 90°C, 85°C, 80°C, 75°C and 70°C, followed by 16 hr at 70°C) was applied to increase probe hybridization to EBV genomic fragments. Probe-hybridized complexes were then immobilized on streptavidin- coated magnetic beads. Enriched DNA fragments were then dehybridized from the beads and amplified over 14-20 cycles to obtain the requested quantity for the PacBio sequencing library preparation.
[0166] Post-capture QC
[0167] The EBV copy in the post-capture (enriched) library was measured by qPCR and compared to the EBV copy in the pooled pre-capture library to calculate capture efficiency. Libraries with capture efficiency more than 20 are sent for PacBio sequencing.
[0168] PacBio Sequencing
[0169] Enriched libraries were pooled together as input DNA for sequencing library preparation. Sequencing libraries were created using the SMRTbell Express Template Prep Kit 2.0 (PacBio) per manufacturer’s instructions. Sequencing libraries were sequenced on a PacBio Sequel II System. After sequencing, CCS analyses were inn using SMRTLink software vlO to produce HiFi reads.
[0170] Sequencing data analysis
[0171] Once the PacBio Capture sequencing data was generated, the PCR duplicates were removed using Pbmarkdup. Then the data was demultiplexed using in-house scripts. Briefly, a read is assigned to a specific sample if at least one of its Twist index pair (or its complementary sequences) is found from either end of the read. The reads which were assigned to multiple samples were filtered. To remove all possible adapters, the 100 bases from both ends of the reads were trimmed. Subsequently, the reads were aligned to EBV reference using Pbmm2, and SNPs and small Indels were called using deepvariant v 1.5.0 with PacBio mode.
[0172] Results
[0173] Target enrichment and sequencing
[0174] The large (172-kb) EBV genome combined with the low abundance of viral DNA relative to host DNA presents a sequencing challenge. A custom target enrichment library for EBV was designed, comprising two series of 120-nucleotide double-stranded oligonucleotide baits tiled across the EBV genome (997 baits covering entire EBV and 302 covering repetitive regions). Starting with only 300 ng of DNA, up to 300-1900 fold enrichment of EBV DNA was achieved using this method, with 58 to 96% of reads mapping to the EBV genome after enrichment, compared to about 0.03-0.55 % of reads without enrichment. Multiplexing up to 22 enriched samples per run on a PacBio Sequel II System produced sufficient reads from each sample with an average coverage depth of more than 1,000 reads per nucleotide (Table 1 and Figure 1).
[0175] Variant calling
[0176] The number of single nucleotide polymorphism (SNP) identified in each sample ranged from 891-1558 and the number of indels ranged from 84 to 143 (Table 2).
[0177] Table 2. Long read target sequencing statistics for each sample
[0178]
[0179] HC: healthy control
[0180] Table 3. SNPs and indels identified in each sample.
[0181] EXAMPLE 3: Long read EBV target sequencing with Oxford Nanopore ONT Adaptive Sampling (AS)
[0182] Oxford Nanopore sequencing allows real-time decoding of the region of the genome being sequenced. This characteristic allows decisions to be made in real time on whether a particular strand is of interest or not. This called adaptive sampling, and it can perform realtime selection of reads when the sequencing software (MinKNOW) is supplied with a .bed file containing the regions of interest (ROI) and a FASTA reference file.
[0183] Adaptive sampling can run in two different modes: enrichment and depletion. In enrichment mode, ROIs have to be uploaded to MinKNOW, which then rejects strands that fall outside of these regions. Tn depletion mode, targets that are not of interest (e.g., host DNA) are uploaded to MinKNOW, which then rejects strands that fall within these regions.
[0184] The inventors used the enrichment mode and uploaded the EBV reference genome (NC_007605.1) as ROI to MinKNOW, which then rejected strands that fall outside of these regions.
[0185] DNA preparation
[0186] For tumor samples, when DNA extraction is performed using column-based methods, such as the Qiagen AllPrep DNA / RNA Kit, the resulting DNA typically has a fragment size of <60 kb and is of good quality (i.e., no degradation). The DNA can then proceed directly to library preparation.
[0187] For saliva samples, a shorter fragment tail was present in most of the extracted DNA samples. Therefore, 1-2 pg of DNA input were subjected to mechanical fragmentation using Megaruptor 3 (Diagenode, B06010003) with settings targeted at size range of 55 - 60 kb for samples required shearing. The DNA was purified with Ampurc XP beads at 0.4x to remove DNA fragments <3 kb.
[0188] Library preparation and sequencing
[0189] The tumor and purified saliva DNA were used as input for the library preparation. Library preparation was done using the Oxford ONT SQK-LSK114 kit and followed the manufacturer’s protocols. The barcoded library DNA was pooled and sequenced on PromethlON 24 with R10.4.1 flowcells. Sequencing run was conducted with live- basecalling in fast mode, with adaptive sampling selected.
[0190] Data analysis
[0191] The output POD5 files from MiniKNOW were re-basecalled using Dorado (v0.7.0) with the SUP model and reads shorter than 1 kbps were filtered out. The filtered reads (>1 kbps) were mapped to a combined reference of human (CHM13) and EBV (NC_007605) using Minimap2 (v2.28). Reads aligning to the EBV genome were extracted as EBV-like reads and subsequently mapped to an EBV-only reference for downstream analysis.
[0192] Results:
[0193] As recommended by Oxford ONT, Adaptive Sampling (AS) can only be used when the target DNA percentage in the host DNA exceeds 1 %. However, most clinical samples, such as saliva and tumor samples, have much lower EBV content than 1%. To better understand whether Adaptive Sampling can be applied for EBV-targctcd sequencing, the inventors compared ONT Whole Genome Sequencing (WGS) and ONT Adaptive Sampling (AS) using four EBV cell lines: C17, C666, NPC43, and Raji. Among these, C17, C666, and NPC43 have low EBV copy numbers, while Raji has a high EBV copy number. Table 4 displays the sequencing results from both Adaptive Sampling and WGS sequencing. Table 4. Sequencing results of the hybrid runs with four EBV cell lines
[0194] From the table, it is clear that while direct sequencing yields more reads over 1 kb, Adaptive Sampling (AS) outperforms in terms of the number of EBV reads and genome coverage. This result demonstrates that even when the EBV DNA content is much lower than 1%, AS can still achieve more than 50x genome coverage for EBV in three of the cell lines examined.
[0195] To comprehensively understand the performance of short read capture, long read capture and AS, the inventors compared the three methods in terms of EBV genome coverage, variants calling and structural variations when used to analyse saliva samples. The number samples which had been included in the comparison testing are summarised in Table 5.
[0196] Table 5. Number of samples sequenced using the three methods.
[0197] A) The median read length of AS and long read capture
[0198] The median read length of AS and long-read capture were similar, both around 8 kb (Figure 1). However, AS had a higher proportion of longer reads compared to long-read capture, which could be beneficial for analyzing repetitive regions, such as copy number variations (CNVs) and structural variations (SVs), as demonstrated below.
[0199] B) The uniformity of EBV coverage among the three different methods
[0200] The uniformity of EBV genome coverage was compared across the three different EBV enrichment sequencing methods. As shown in Figure 2, AS exhibited the best uniformity, followed by long-read capture and short-read capture. The low coverage observed in two of the target sequencing datasets was primarily in repetitive and GC-rich regions. For the long- read sequencing protocol, the low coverage in regions with high GC content is likely due to the challenges PCR faces in amplifying these areas. Similarly, for the short-read sequencing protocol, the low coverage can be attributed to both the difficulty PCR encounters when amplifying GC-rich regions and the challenge short reads face in being uniquely mapped within repetitive regions.
[0201] C) Comparative analysis for the performance of variant calling across the three different method data for one saliva sample.
[0202] Using a single saliva sample, the variants called by the three different sequencing methods were compared using the Deep Variant pipeline. In the non-repetitive regions of the EBV genome, 99% of SNPs were detected by all three methods, with only a few SNPs identified uniquely by each method (Figure 3A). However, the common indels called from non- repetitive regions were only 73%, and only one indel was uniquely identified by the shortread capture method (Figure 3B). This is expected, as short-read methods have lower power to detect indels. When comparing variants called from repetitive regions, only 16% of SNPs and 6%> of indels were detected by all three methods. The two long-read methods identified significantly more variants than the short-read method (Figure 3C and D). Among these, AS demonstrated the highest capacity for variant detection in repetitive regions. This is reasonable, as short-read capture methods often struggle to map reads to repetitive regions, further limiting their ability to detect variants in these regions.
[0203] More detailed analysis showed that most of the variants identified by AS but missed by the other two methods, are mainly the ones within high GC and / or repeat regions (Figure 4). As indicated by the coverage analysis, these regions were poorly analyzed by long and short read capture methods (above). D) Comparative analysis for the performance of SVs calling between AS and long read capture.
[0204] The inventors did not include the short-read capture data in this analysis, as short-read methods typically have low capability in detecting large structural variants (SVs). Surprisingly, both the long-read capture and AS methods demonstrated the same ability to call SVs, both in terms of number and size (Figure 6). The locations of the SVs were also identified by the two methods and it was found that the SVs were located in exactly the same regions. This suggests that the SVs identified by both methods are reliable, as the results are supported by the findings from the other method.
[0205] E) Analysis of the copy number diversity of the repetitive regions with the AS data.
[0206] As mentioned earlier, compared to the other two capture methods, AS data contains longer reads that can span the entire repetitive regions of the EBV genome. The inventors selected four repetitive regions as target regions and performed diversity analysis using AS data. In the EBV reference genome, each repetitive region has a unique copy number, such as 7.6 copies in IR1, 12.3 in IR2, 24.9 in IR4, and 3.9 in the TR regions. However, the results clearly show both intra-sample (within the same sample) and inter-sample (across different samples) diversity in all four regions (Figure 6). While the biological significance of this diversity remains unclear, it is believed this is the first time such diversity has been observed, a finding that is not achievable with short-read or capture methods.
[0207] In summary, a comprehensive approach is provided for analyzing the entire EBV genome isolated from various sample types, including FFPE, saliva, plasma, and tumor tissue. Depending on the DNA quality and EBV load, three different sequencing methods can be chosen for whole EBV genome sequencing. In general, AS offers the best performance in terms of EBV genome coverage uniformity, variant calling, and copy number diversity analysis of repetitive regions, and long read sequencing in general is able to identify more variants, in particular structural variants, compared to short sequencing.
[0208] EXAMPLE 4: Construction of pan-genome graph reference of EBV genome
[0209] Materials & Methods
[0210] Collection and processing of eight EBV cell lines Eight EBV+ cell lines that were established from patients with different diseases were collected from different collaborators. Specifically, B95-8 cell line is a house-keeping cell line, Raji and C666-1 cell lines were contributed by National Cancer Center of Singapore (NCCS), C17 and NPC43 cell lines were contributed by National University of Singapore (NUS). The YCC-10 and SNU-719 cell lines were contributed by Duke-NUS Medical School, Singapore. The remaining M81 cell lines was contributed by the German Cancer Research Center, Elcidclbcrg.
[0211] B95-8, originating from a monkey and linked to transfusion-induced mononucleosis, differs from the other seven human cell lines. Raji comes from an African Burkitt’s lymphoma (BL) patient, YCC-10 and SNU-719 from EBV+ Gastric Cancer (GC) patients, and C666-1, M81, C17 and NPC43 from NPC patients. C17 was isolated from a French patient, while the remaining three NPC cell lines were isolated from Chinese patients.
[0212] High Molecular Weight (EIMW) DNA extraction
[0213] HMW DNA used for both Illumina and PacBio sequencing was extracted using the GentraPuregene kit (Qiagen; #158043) according to manufacturer’s instruction. Briefly, 1 x 107frozen cell pellets from respective EBV cell lines were used as input to extraction. All vortexing steps were replaced with gentle inversion throughout the process and 300 pl of Qiagen EB buffer was used for elution. Eluted DNA was incubated at 12°C with gentle shaking over a period of 7 to 10 days. To avoid shearing the high molecular weight DNA, wide bore tips with gentle pipetting were used during handling. DNA was stored at 4°C to prevent freeze and thaw cycle. The quantity and purity of extracted HMW DNA were assessed using triplicate concentration measurements from top, middle and bottom sections of sample volume, using Qubit dsDNA BR (Broad-Range) assay (Thermofisher Scientific; Q32853) and NanoDrop 2000 spectrophotometer (ThermoFisher Scientific; ND-2000), according to manufacturer’s instructions.
[0214] Short-read library preparation and sequencing
[0215] About 500ng of DNA of the samplc(s) were first fragmented using the Covaris to obtain an average distribution size of ~350bp. However, only ~150ng of the fragmented DNA was used for the actual input for the library preparation. Library preparation was done using a commercially available kit, NEBNext® Ultra™ IT DNA Library Prep Kit for Illumina® following the manufacturer’s protocol. During the library construction, the fragmented dsDNA then underwent end repair, adapter ligation, size selection and PCR enrichment to generate the final sequencing ready library. The quality of the library was checked using Agilent DI 000 ScreenTape. A single peak in the expected region of -470 bp should be observed indicating that the library was good and suitable for sequencing. The different libraries were then pooled together and QC using Agilent D1000 ScreenTape. They were then submitted for sequencing on 2x lanes of Standard NOVASEQ S4 PE151.
[0216] Long-read library preparation and sequencing
[0217] After DNA extraction, DNA fragment lengths were then measured using TapeStation 4200 (Agilent). Sequencing libraries for each sample were created using the SMRTbell Express Template Prep Kit 2.0 (PacBio) per manufacturer’s instructions. Libraries were sequenced with one sample per run on a Sequel 11 System (PacBio). After sequencing, CCS analyses were run using SMRTLink software vlO to produce HiFi reads for each sample.
[0218] Ultra-High Molecular Weight (UHMW) DNA extraction
[0219] UHMW DNA was extracted using the Nanobind CBB Big DNA Kit (Circulomics; NB-900- 001-01), in combination with Nanobind UL Library Prep Kit (Circulomics; NB900-601-01) and UHMW DNA Auxiliary Kit (Circulomics; NB-900-101-01). In brief, 6 x 106frozen cell pellets from respective EBV cell lines were used as input for extraction, with final elution volume in 760pl Buffer EB+, according to the recommended Circulomics protocol. The quantity and purity of extracted UHMW DNA were assessed using the same triplicate concentration measurements as above. Quality-assessed UHMW DNA from each EBV cell line was carried forward to Nanopore ultra-long read library preparation and sequencing.
[0220] Nanopore Ultra-Long (UL) read library preparation and PromethlON sequencing Quality-assessed UHMW DNA from each EBV cell line in 750p 1 Buffer EB+ was used as input to 3x scaled Nanopore ultra-long read library preparation using the Ultra-Long DNA Sequencing Kit (Oxford Nanoporc; SQK-ULK001). In brief, for each library prepared, 6pl ONT Fragmentation Mix (FRA) was diluted with 244pl FRA Dilution Buffer (FDB) and added to 750pl UHMW DNA, with immediate vortcxing for 5 seconds at low setting, followed by gentle pipetting with P1000 wide bore tips. The DNA tagmentation reaction was incubated at room temperature for 5 minutes, followed by inactivation at 75°C, for 5 minutes, and equilibration to room temperature for 20 minutes. Rapid sequencing adaptor attachment was carried out with the addition of 5p 1 of Rapid Adapter F (RAP-F) to the equilibrated tagmentation reaction, followed by gentle pipetting with P1000 wide bore tips and incubation at room temperature for 35 minutes. Library purification was carried out using the Nanobind UL Library Prep Kit (Circulomics; NB900-601 -01 ), according to manufacturer’s protocol, with overnight elution of purified UL read library. Final UL read libraries were assessed for recovery using the Qubit dsDNA BR assay (Thermofisher Scientific; Q32853) and divided into three equal aliquots for PromethlON sequencing. Each UL read library aliquot was sequenced on R9.4.1 flowcells (Oxford Nanoporc Technologies; FLO-PROOD2), with MUX scan intervals set at 6 hours. Nuclease flushes were performed according to manufacturer’s instructions, using the flowcell wash kit (Oxford Nanopore Technologies; EXP-WSH004), at 24-hour intervals, followed by reloading of the UL read library, for a total of two flushes and three loads per library. All sequencing runs were conducted with live-basecalling in super accurate (SUP) mode, on PromethlON release versions 21.05.24, with base-calling performed under Guppy versions 5.0.11.
[0221] De novo assembly analysis of EBV genome sequences and construction of EBV pan-genome reference
[0222] PacBio HiFi reads with accuracy higher than 99% (>Q20) were extracted, and then aligned against EBV reference genome (NC_007605.1) using minimap2_v2.22. Those EBV-related reads with less than 70% of their length mapped to reference EBV genome were discarded using Ram_v2.1.0. De novo assembly analysis based on the remaining EBV-specific reads was conducted using multiple assemblers including Hifiasm_vO.13-r3O8, Flye_v2.8.2- bl689, Raven_vl.5.0 and Hifiasm-meta_vl. Detailed structures of the assembled EBV genomes / contigs were then viewed under Bandage, and the pairwise sequence comparison between newly assembled EBV genomes and the corresponding public sequences were investigated using Gepard_vl.4. Those newly assembled genomes / contigs were converted in their form of complementary reversion if needed, and then manually rearranged to share the same starting position as the reference NC_007605.1 for the sake of straightforward downstream sequence analysis.
[0223] The newly assembled sequences generated from Raven were set as the by-dcfault genome sequences for each EBV strain. Eventually, these eight newly assembled genomes (including B95-8, C17, NPC43, Raji, C666-1, M81, YCC-10 and SNU-719) were incorporated into constructing a pan-genome reference for EBV using minigraph-0.20. Again, the detailed structure of this EBV pangenome reference was viewed under Bandage and further interpretations were investigated using an in-house customized script.
[0224] Extraction of RRs-spanning reads and estimation of copy numbers
[0225] Ikb regions upstream and downstream of the 1R1 region were extracted based on EBV reference NC_007605.1 and k-mer hash tables were constructed. A spanning read was defined if it overlapped k-mcrs with both the upstream and downstream regions. This information was generated via commands such as “Yak count” and “Yak triobin” by Yak_V0.1-r69, respectively.
[0226] To estimate copy numbers, a new reference consisting of multiple copies of the 1st unit of RR was constructed and the spanning reads were mapped to it. The copies of RR were estimated by dividing the alignment block length by unit length (unit length is 3071 -bp for IR1, and 538-bp for TR). Each value of the copy numbers was confirmed by Gepard_V2.1 dotplot.
[0227] Quality evaluation of new EBV genome assemblies
[0228] For each individual EBV+ cell line, detailed EBV-related reads were summarized using “flagstat” under samtools_vl.9, the total reads’ numbers and mapping rates were subsequently calculated. Coverage distributions were examined using bedtools and presented using Rstudio_v2021.09. Distribution of reads’ lengths were summarized using an in-house script and plotted using Rstudio_v2021.09. Quast_v5.0.2 was also applied to evaluate the performance and completeness of new genome assemblies. To carefully investigate the distribution of copy numbers over different RRs, another in-house tool was designed to extract those PacBio HiFi long reads which enable to transcend through the entire TR1 and TR region. Once these reads were retrieved, two-end sequences covering flanking regions were trimmed. Detailed copy numbers were eventually calculated in terms of the well-defined sequence length of single copy unit over IR1 and TR region.
[0229] Comparison between public and newly assembled genome sequences
[0230] Newly assembled genome sequences of five EBV strains were compared to their corresponding public sequences (when available) using mafft_v7.487. In terms of the pairwise sequence comparison, an in-house script was used to curate the SNPs and INDELs, which were then manually examined and validated using the corresponding PacBio HiFi long- and Illumina short-read dataset in Integrative Genomics Viewer (IGV). Functional annotations of these variants were performed according to the genomic annotation of B95-8 strain (NC_007605.1 ). Comparison of newly assembled EBV genome sequences / contigs from different assemblers also followed the same way.
[0231] Results
[0232] De novo assembly analysis of EBV genomes in eight EB V+ cell lines
[0233] EBV genomes in eight EBV+ cell lines were sequenced using PacBio HiFi and ONT ultralong read sequencing technologies. Among these cell lines, B95-8, Raji, C666-1 , M81 and SNU-719 are commonly used as “model strains” for EBV research, while C17, NPC43 and YCC-lOare more recently established (Table 6). B95-8, originating from a monkey and linked to transfusion-induced mononucleosis, differs from the other seven human cell lines. Raji comes from an African BL patient, YCC-10 and SNU-719 from EBV+ GC patients, and C666-1, M81, C17 and NPC43 from NPC patients. C17 was isolated from a French patient, while the remaining three NPC cell lines were isolated from Chinese patients.
[0234] EBV genomic DNAs from eight EBV+ cell lines were sequenced using PacBio SMRT technology, generating varying numbers of HiFi reads (508,248 to 2,875,932) with read quality > Q20. These HiFi reads were aligned against the EBV reference genome (NC_007605.1), resulting in 120 to 5,079 EBV-related reads (average length ~15kb) (Table 6). Selected EBV-rclatcd reads (>70% identity to NC_007605.1) were defined as EBV- specific reads and used for de novo assembly, providing coverage ranging from 10X for C 17 (104 reads) to 520X for M81 (4,968 reads) (Table 6). Among these eight EBV+ cell lines, B95-8 had the highest percentage of EBV-specific reads (0.67%, N=3,398), while C17 and NPC43 had the lowest (4.7e-5%, N=104 and 0.3%, N=449), showcasing variations in EBV DNA abundance in each cell line, but not the sequencing coverage. Despite differing EBV- specific read counts, sequencing coverage was distributed evenly across the entire EBV genome for all eight EBV+ cell lines (Figure 3).
[0235] De novo assembly of the EBV genome employed three algorithms (Hifiasm, Flye and Raven) (Table 8), with Raven being the sole algorithm to yield a single circular contig for each of the eight strains. Hifiasm and Flye generated single contigs for five and six strains, respectively. Notably, Hifiasm produced multiple contigs for B95-8, Raji, and M81, with varying lengths (Figure 4). Flye generated one long circular contig and one short circular contig for both Raji and YCC-10, while no contig was produced for M81 (Figure 4). Given these eight strains, the circular contigs generated by the three algorithms were different in length. For Raji and SNU-719, Hifiasm's circular contigs surpassed those of Raven and Flye, exceeding the current reference EBV genome (NC_007605.1). In contrast, Flye's circular contigs for B95-8, C666-1, and SNU-719 were notably shorter than those of Raven and Hifiasm, falling below the size of the reference genome. Hifiasm-meta, optimized for mctagcnomc analysis, performed poorly, generating circular contigs for only three strains (Raji, C666-1, and NPC43).
[0236] Comparison of circular contigs generated by the three assembly algorithms revealed differences mainly in repeat regions (RRs), especially in the copy numbers of IR1 and TR repeats (Table 8). For B95-8, the contigs generated by Hifiasm and Raven had singular long IR1 sequences with 13 (174,968-bp) and 9 repeat copies (162,698-bp) respectively (Table 8) , while Flye could not fully resolve the IR1 region, yielding two related sequences (1,217- bp and 1,855-bp) composing one complete IR1 repeat copy (3072-bp) (Figure 4). In C666- 1, Hifiasm and Raven contigs carried eight copies of IR1 repeats, while Flye had unresolved paths with two incomplete alternative sequences (4,923-bp and 4,295-bp) (Figure 4). The contigs from three assemblers also showed various copy numbers of TR repeats (Table 8). For B95-8, Flye and Raven yielded a single TR sequence carrying three copies, but Hifiasm assembled a circular' contig (#1) with three copies and two linear' contigs (#2 & 3) with different repeat copy numbers. In Raji, Hifiasm and Flye contigs had sixteen TR repeats, while Raven had nine. Non-RR sequences showed no differences among contigs from the three assemblers.
[0237] Analysis of RR sequence diversity within individual EBV strains
[0238] To clarify the discordant copy numbers of RR repeats in circular contigs identified by the three algorithms, individual reads spanning the entire RRs (referred to as RR-spanning reads) were isolated using Yak_v0.1-r69. These reads were selected based on sequences at both ends successfully mapping to RR flanking regions. The lengths of EBV-specific PacBio HiFi and ONT ultra-long reads were compared. For each cell line, ONT ultra-long reads yielded between 140 to 21,459 EBV-specific ultra-long reads (>70% identity to EBV reference NC_007605.1) (Table 7 and Figure 5). The longest 25% of EBV-specific ultralong reads had lengths ranging from 18,232 to 171 ,497 bps, which surpassed the lengths of PacBio HiFi reads, where the longest 25% ranged from 17,083 to 21,033 bps (Table 9). Because PacBio HiFi reads are much shorter than ONT ultra-long reads, the copy number analysis of eight EBV+ cell lines were conducted using ONT ultra-long reads to avoid any potential bias resulting from the length limitations.
[0239] Across the eight cell lines, 10 to 1,074 IRl-spanning reads and 9 to 819 TR-spanning reads were identified (Table 10), as summarized in Figure 7. Notably, RR-spanning reads from the same sample exhibited varying copy numbers for both IR1 and TR. The 1R1 copy numbers varied among strains, with B95-8 ranging from -1 to -18 and the most abundant reads carried -13 copies. M81 displayed a wide range from -1 to -14 copies, with the most abundant reads carrying -8 copies. Other strains such as C666-1 , NPC43, SNU-719, and YCC-10 also exhibited high diversity in IR1 copy numbers, ranging from -3 to -8 copies. TR-spanning reads showed lower copy number diversity compared to IRl-spanning reads, with B95-8 carrying -2 to -5 TR copies, M81 carrying -1 to -9 TR copies and NPC43 carrying -4 to -7 TR copies. These RR-spanning ONT reads highlight high intra- sample diversity of IR1 and TR copy numbers within each EB V cell line.
[0240] IR1 copy numbers from the new assembly were compared with IRl-spanning ONT (Figure 7). Raven's assembled IR1 sequences predominantly matched the copy numbers from IRl- spanning reads obtained from ONT ultra-long sequencing, except for C17, which had extremely low coverage and rare IRl-spanning reads detected. Considering the alignment between IR-spanning read-level and genome-level assemblies achieved by Raven, and the unique capability of Raven to produce a single circular contig for all eight strains. Raven assemblies were chosen as the EBV genomes for subsequent analyses.
[0241] Quality evaluation of new genome assemblies and comparison with public sequences
[0242] The new genome assemblies’ quality and completeness were assessed with Quast. Genome fractions, representing the total numbers of aligned bases with the current EBV reference (NC_007605.1) divided by the genome size of the reference, showed a high level of consistency (92.772% - 99.964%) across the eight contigs assembled by Raven (Table 11). Assembly quality was further evaluated by comparing non-RR sequences directly with the reference NC_007605.1. Variants between the new assembled genome sequences and the reference were manually checked using HiFi reads on Integrative Genomics Viewer (IGV), confirming Raven's good performance and the high quality of the newly assembled genomes. Five newly assembled EBV genomes were systematically compared with their corresponding public sequences: B95-8 (partially NC_007605.1), Raji (KF717093), C666- 1 (KJ41 1974), M81 (KF373730) and SNU-719 (MH590579). C17, NPC43 and YCC-10 were excluded due to the absence of published sequences. After the necessary reverse- complementary conversion, linear representations of the genomes were created and aligned with the reference NC_007605.1. First, the completeness of the genomes were compared. It was observed that many RR sequences were missing in published genomes, particularly for C666-1 and SUN-719. For example, there are a total of 7,971 ambiguous bases (N) within the IR2 / 3 / 4 and TR regions of the published C666-1 genome sequences and 22,150 ambiguous bases (N) within the TR1 / 2 / 3 / 4 and TR regions of the published SNU-719 genome sequence (Figure 6A). Notably, the new assemblies filled gaps in RR sequences absent in published genomes, particularly in C666-1 and SNU-719. The completeness of RRs was enhanced in the assemblies (Figure 6B). For the IR1 region, it was observed that Raji, C666-1 and M81 all carried 8 copies of repeats in both newly assembled and published sequences. The newly assembled SNU-719 sequence revealed 7 copies of IR1 repeats, which were missing in the corresponding published sequence. The newly assembled sequence of B95-8 carried 9 copies of IR1 repeats instead of 8 in the published sequence.
[0243] Beyond the RRs, the new genome assemblies exhibited sequence differences (INDELs and SNPs) compared to published sequences. Examples include a 1-bp deletion in BCLF1 (new B95-8 assembly), four SNPs in EBNA-LP and one SNP in B VLF1 (new Raji assembly), one SNP in BFLF1, BHLF1, LF2, BALF5 and Qp-EBNAl (new M81 assembly), and a 4-bp deletion in Qp-EBNAl (new SNU-719 assembly). Analysis of both long and short reads confirmed the biallelic or multi-allelic states, with the new assemblies representing major alleles, while the traditional reference showed minor alleles. Taken together, the high accuracy of long-read sequencing technology effectively captured variants in both RRs and non-RRs across all cell lines.
[0244] Construction of a EBV pan-genome reference
[0245] By combining the newly assembled genome sequences of eight strains, an EBV pan-genome reference was built and a graphical representation of the pan-genome was created (Figure 8B). The pan-genome graph is composed of the chains of bubbles wi th the newly assembled B95-8 genome sequence as a backbone. Each bubble represents a structural variation, which can be multi-allelic as reflected by multiple paths through the bubble. For instance, a bubble within the IR1 region has three sequence segments of 3 -bp (SI), 3,072-bp (S2) and 3076-bp (S3) (Figure SA) that represented three alternative paths including SI (4-bp insertion), S2 (3,072-bp insertion, ~1 IR1 copy unit) and S2+S3 (6,148-bp insertion, ~2 IR1 copy units). Among these three paths, SI is contributed by the genome sequence of SNU-719 strain while S2 & S3 is represented by the sequence of B95-8 strain. The pan-genome reference consists of a core region (excluding structure variation) with the length of 147,376-bp and the variable regions (including structural variation) with the total length of 34,147-bp. Table 12 summarizes the total numbers of sequence segments (3 to 26) carried by each variable region, as well as the corresponding numbers of potential alternative paths (2 to 1,081). These variable regions mainly overlap with RRs and key latent gene regions (LMP-1 / 2A / 2B, EBNA-l / LP), with the longest path (13,231-bp) at position 142,930-bp against the reference (NC_007605.1), coinciding with the 1R4 deletion in B95-8.
[0246] The utility of the new EBV pan-genome reference for analyzing NGS-based genome sequencing data was further explored using 28 EBV isolates. First, the new pan-genome graph was converted into a linear reference by combining all the consensus sequences of the core regions and the multiple sequences of variable regions (representing different alternative paths) into a single sequence file. By following GATK Best Practices, joint variant calling of these 28 EBV isolates was performed by using either the traditional reference (NC_007605.1) or the newly assembled pan-genome reference. Overall, 3,405 variants were called by using the newly assembled pan-genome reference, whereas 3,185 variants were called by using the traditional reference (NC_006705.1). Of these variants called, 3,057 variants were shared between two analyses. Among those 128 variants that were exclusively called by the traditional reference, 83 (64.8%) were TNDELs (65 within the RRs and 18 within the non-RRs), 45 (35.2%) were SNPs (42 within RRs and 3 within non-RRs). Among the 348 variants that were exclusive to the new pan-genome reference, 208 (59.8%) out of 348 variants were INDELs (176 within the RRs and 32 within the non- RRs), while the remaining 140 (40.2%) variants were SNPs (128 within RRs and 12 within non-RRs). Using the same pipeline, variant calling analyses of 28 EBV isolates from NPC patients was also performed by using each of eight newly assembled EBV genome sequence as reference. Eight sets of variants called were then separately compared to the results of variant calling based on the new pan-genome reference, with detailed results illustrated via venn diagram (Figure 9). As shown, all the analyses using single genome reference called less numbers of variants than the one by using the pan-genome reference. Discussion
[0247] Here, long-read sequencing technology was used to sequence eight EBV+ cell lines and complete genome sequences were generated for these EBV strains through de novo assembly. As the most comprehensive long-read sequencing analysis of EBV genomes, this study has created a new EBV pan-genome reference. Compared to the single genome references, the new pan-genome reference captured sequence diversity of these eight EBV strains and outperformed in empowering NGS-based variant calling.
[0248] The long reads, capable of transcending through all repeat regions (RRs), enabled a thorough interrogation of the repetitive sequences constituting about 30% of the EBV genome. Typically masked and excluded in NGS-based EBV genome sequencing analysis, these RR sequences, particularly in the IR1 region, play a crucial role in viral transformation: 1) the ability of EBV-transformed LCLs dropped progressively with the viruses carrying four or fewer copes of IR1, and 2) deletion of a largely inverted repeat sequence (invR) within IR1 abolished its transforming capability. Consistent with the other observations, the result suggests that, when analysing DNA samples with diverse copy numbers of repeat sequences, PacBio HiFi sequencing showed a bias in favour of sequencing shorter DNA fragments with less copy number of repeats than ONT ultra-long sequencing. It needs to be point out that other factors, including distinct library preparation protocols, DNA damage, DNA quality (that could potentially decline during prolonged sequencing at room temperature) and non- canonical DNA structures may also contribute to variations between these two commonly used long read sequencing technologies. This technical bias will impose limitations on analysing DNA samples with substantial genetic heterogeneity, such as microbial and cancer DNA samples, by PacBio HiFi sequencing. The current study has clearly demonstrated the capability of ultra-long read sequencing to do the comprehensive EBV genome sequence analysis and quantify the intra-host genetic diversity, which provides the opportunity to investigate quasi-species evolution of EBV within a host.
[0249] The study revealed substantial differences among four algorithms in analyzing sequencing reads with significant genetic diversity. Raven was the only algorithm able to generate a complete genome sequence (circular contig) for all eight strains and yield copy numbers of repeats supported by sequencing data. Hifiasm-meta, designed for analyzing metagenomics data, failed to generate complete genome sequences for five of the eight strains. It remains to be demonstrated whether these algorithms will face similar limitations in analyzing cancer DNA samples, where genetic diversity of sequences is also significant. De novo assembly analysis of such DNA samples may require new algorithms capable of handling the high- level genetic diversity of sequencing reads.
[0250] In summary, complete genome sequences for eight EB V strains were successfully generated using long-read sequencing and de. novo assembly analysis. Utilizing these eight EBV genome sequences, a pan-genome reference for EBV was constructed, showcasing its advantages over single genome references in empowering reference-based sequencing and variant calling analysis of the EBV genome. This study also demonstrated the capability of long-read sequencing technologies to analyze repetitive sequences and inter- and intra-host genetic diversities of the EBV genome. This would provide deeper insights into understanding EBV evolution and pathogenesis in various human diseases.
[0251] Table 6: Swmnary of EBV-rdated&pecific long reads distributed amsg eight EBV4 ceB lines.
[0252] Table Summary of EBV-related / speciBc ultra-long reads distributed among eight EBV* cell lines.
[0253] Table 8. Performance evaluation of different assemblers.
[0254] Table Comparison of BBV-«pecific HiFi and ONT uhra-loog read®.
[0255] Table 10. Summary of ONT ultra-long RRs-related reads.
[0256] Table 11. Summary of assembly QV and completeness. Table 12. Details description of bubbles (structural variations) over 8 newly assembled EBV genome sequences.
[0257] It will be appreciated that many further modifications and permutations of various aspects of the described embodiments are possible. Accordingly, the described aspects are intended to embrace all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims.
Claims
CLAIMS1 . A method of analyzing viral genomic DNA in a sample; the method comprising: a) generating a population of DNA fragments from the sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample.
2. The method of claim 1, wherein analyzing the viral genomic DNA comprises detecting: i) the presence or absence, ii) the strain, iii) the subtype or iv) the genetic variant of the virus.
3. The method of claim 1 or 2, wherein the viral DNA is from Epstein-Barr virus (EB V).
4. The method of any one of claims 1 to 3, wherein the method comprises comparing the long or short read sequencing reads to a reference viral genome to analyze the viral genomic DNA.
5. The method of claim 4, wherein the reference viral genome comprises sequence information of viral genomes from a plurality of viral strains.
6. The method of claim 5, wherein the reference viral genome comprises sequence information from EBV genomes derived from two or more cell lines selected from B95-8, C17, NPC43, Raji, C666-1, M81, YCC-10 and SNU-719.
7. The method of any one of claims 1 to 6, wherein the plurality of viral- specific nucleic acid capture probes is derived from a reference viral genome.
8. The method of claim 7, wherein the reference viral genome comprises sequence information of viral genomes from a plurality of viral strains.
9. The method of claim 8, wherein the reference viral genome comprises sequence information from EBV genomes derived from two or more cell lines selected from B95-8, C17, NPC43, Raji, C666-1 , M81 , YCC-10 and SNU-719.
10. The method of any one of claims 1 to 9, wherein the plurality of viral- specific nucleic acid capture probes comprises at least one nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ ID NO: 9-76.1 1. The method of any one of claims 1 to 10, wherein the sample is a biological sample from a subject.
12. The method of claim 11, wherein the subject is one who is suffering from cancer or is suspected of having cancer.
13. The method of any one of claims 1 to 12, wherein generating a population of DNA fragments comprises: a) fragmenting DNA in the sample; b) ligating fragmented DNA to nucleic acid adaptors; and c) amplifying ligated DNA to obtain a population of DNA fragments.
14. The method of claim 13, wherein the method comprises fragmenting DNA to a size distribution of about lOkb.
15. The method of any one of claims 1 to 14, wherein the plurality of viral- specific nucleic acid capture probes are about 100-140 nucleotides in length.
16. The method of any one of claims 1 to 15, wherein the plurality of viral- specific nucleic acid capture probes target nucleic acid sequences across the entire viral genome.
17. The method of any one of claims 1 to 16, wherein the plurality of viral- specific nucleic acid capture probes comprises an affinity tag.
18. The method of claim 17, wherein the affinity tag is biotin or a derivative thereof.
19. The method of any one of claims 1 to 18, further comprising capturing viral DNA fragments that are hybridized to the viral-specific nucleic acid capture probes onto a substrate.
20. The method of any one of claims 1 to 19. wherein the method comprises: a) generating a population of DNA fragments from the sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample.
21. The method of claim 20, wherein the sample is characterized by: i) a viral copy number that is at least 3000 copies per 500 ng of sample DNA, and / or ii) a DNA integrity number (DIN) that is at least 7.
22. The method of any one of claims 1 to 19, wherein the method comprises: a) generating a population of DNA fragments from the sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing short read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample.
23. The method of claim 22, wherein the sample is characterized by: i) a viral copy number that is greater than 250 copies but less than 3000 copies per 500 ng of sample DNA.
24. The method of any one of claims 1 to 23, wherein the method further comprises detecting the viral copy number and / or DNA integrity in the sample.
25. A method of sequencing a viral DNA genome; the method comprising:a) generating a population of DNA fragments from a sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral-specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to sequence the viral DNA in the sample.
26. The method of claim 25, wherein the method comprises a) generating a population of DNA fragments from a sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long read sequencing on the viral DNA library to sequence the viral DNA in the sample.
27. A method of analyzing viral genomic DNA in a sample; the method comprising: a) detecting the viral copy number and / or DNA integrity in the sample; b) generating a population of DNA fragments from the sample; c) preparing a DNA library from the DNA fragments; and d) performing long or short read sequencing on the DNA library to analyze the viral genomic DNA in the sample, wherein long read sequencing is performed when the viral copy number is at least 3000 copies per 500 ng of sample DNA and / or the DNA integrity number (DIN) is at least 7; and wherein short read sequencing is performed when the viral copy number is greater than 250 copies but less than 3000 copies per 500 ng of sample DNA.
28. The method of claim 27, further comprising enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes prior to step (c).
29. The method of claim 27, wherein long read adaptive sampling sequencing is performed when the viral copy number is at least 200,000 copies per 500 ng of sample DNA and the DNA integrity number (DIN) is at least 7.
30. A method of detecting a subject who is at risk of a cancer; the method comprising: a) generating a population of DNA fragments from a sample obtained from the subject; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long or short read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample, so as to determine whether the subject is at risk of a cancer.
31. The method of claim 30, wherein the method comprises: a) generating a population of DNA fragments from the sample; b) enriching the population of DNA fragments for viral DNA fragments by hybridizing viral DNA fragments to a plurality of viral- specific nucleic acid capture probes; c) preparing a viral DNA library from the enriched viral DNA fragments; and d) performing long read sequencing on the viral DNA library to analyze the viral genomic DNA in the sample.
32. The method of claim 30 or 31, wherein the cancer is an EB V-associated cancer.
33. A kit for analyzing EBV genomic DNA, the kit comprising one or more nucleic acid capture probes derived from a reference EBV genome comprising sequence information of EBV genomes from a plurality of EBV strains.
34. The kit of claim 33, wherein the reference EBV genome comprises sequence information from EBV genomes derived from two or more cell lines selected from B95-8, C17, NPC43, Raji, C666-1, M81 , YCC-10 and SNU-719.
35. The kit of claim 34, wherein the one or more nucleic acid capture probes comprise a nucleic acid sequence selected from a nucleic acid sequence having at least 80% sequence identity to a nucleic acid sequence in SEQ TD NO: 9-76.
36. A method of analyzing EB V genomic DNA, the method comprising: a) obtaining a plurality of long or short read sequencing reads from a sample comprising viral genomic DNA; and b) comparing the plurality of sequence reads to a reference EBV genome.
37. The method of claim 36, wherein the reference EBV genome comprises sequence information from EBV genomes of B95-8, C17, NPC43, Raji, C666-1, M81, YCC-10 and SNU-719.
Citation Information
Patent Citations
Method for rapidly identifying eight herpes viruses based on nanopore sequencer
CN113637796A
Detection of circulating tumor DNA using double stranded hybrid capture
WO2021046655A1