Genome sequencing and detection methods

The real-time quality control method using k-mer alignment and primer trimming in sequencing devices addresses the challenge of detecting low-frequency nucleic acids, enhancing accuracy and efficiency in pathogen detection.

JP2026063015APending Publication Date: 2026-04-10ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2026-01-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Next-generation sequencing technologies face challenges in accurately and rapidly detecting low-frequency nucleic acid sequences due to errors and noise from sample defects and PCR biases, making it difficult to identify low-concentration viral or bacterial nucleic acids.

Method used

A real-time quality control method using a sequencing device with a hash table to identify k-mers and generate quality metrics, and a sequencing device with a computer to perform k-mer alignment and generate quality metrics during the sequencing run, along with methods to trim primer sequences and detect variants.

Benefits of technology

Enables rapid and accurate detection of nucleic acid sequences by minimizing noise and improving sequencing accuracy, allowing for efficient resource allocation and real-time detection of pathogens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063015000001_ABST
    Figure 2026063015000001_ABST
Patent Text Reader

Abstract

We disclose nucleic acid sequencing methods. [Solution] Sequence data generated by a sequencing device can be analyzed to scan for k-mers of fixed size n in individual reads within the sequence data. The exact match of k-mers in the sequence data with a reference k-mer is identified. Using k-mer matching, alternative alleles in sequence data with anomalous distributions associated with contamination or other quality issues can be identified, and quality metrics can be determined in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] (Cross - reference to related applications) This application claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 022,296, filed on May 8, 2020, the disclosure of which is incorporated herein by reference.

[0002] The disclosed technology generally relates to nucleic acid characterization, e.g., sequencing techniques. In some embodiments, the disclosed technology includes rapid and accurate methods for virus detection from sequence data based on genomic sequencing, e.g., whole - genome sequencing.

Background Art

[0003] The subject matter considered in this section should not be assumed to be prior art merely as a result of mention in this section. Similarly, problems mentioned in this section, or problems associated with the subject matter provided as background, should not be assumed to have been previously recognized in the prior art. The subject matter of this section merely represents different approaches and may itself also correspond to embodiments of the claimed technology.

[0004] Next - generation sequencing technologies have made sequencing faster and enabled deeper sequencing depths. However, sequencing accuracy and sensitivity are affected by errors and noise from various sources, e.g., sample defects or PCR biases during library preparation. Thus, detection of very low - frequency sequences, such as host samples containing low - concentration viral or bacterial nucleic acids, can be complex. Therefore, there is a desire to develop methods for detecting and / or sequencing nucleic acid molecules that are present in small amounts in a rapid and accurate manner.

Summary of the Invention

Means for Solving the Problems

[0005] In one embodiment, the present disclosure relates to a real-time quality control method. The method includes generating sequence data from a biological sample using a sequencing device that performs a sequencing run; identifying k-mers in the sequence data that have a perfect match in a hash table initialized with a set of k-mers including a reference allele k-mer and an alternative allele k-mer of the reference allele; determining the distribution of the reference allele and alternative allele in the sequence data based on the number of perfect matches; and generating a quality metric for the biological sample based on the distribution during the sequencing run of the biological sample.

[0006] In another embodiment, the present disclosure relates to a sequencing device comprising a substrate loaded with a sequencing library prepared from a sample. The sequencing device also comprises a computer programmed to perform a sequencing run to generate sequence data from the sequencing library, identify k-mers in the sequence data that have a perfect match in a hash table initialized with a set of k-mers including a reference allele k-mer and surrogate allele k-mers of the reference allele, determine the distribution of the reference allele and surrogate allele in the sequence data based on the number of perfect matches, and generate a quality metric for the biological sample in the sequencing device based on the distribution during the sequencing run.

[0007] In another embodiment, the present disclosure relates to a method for detecting variants in a biological sample. This method includes generating an amplicon from a biological sample using a primer pair; preparing a sequencing library from the generated amplicon; generating sequence data from the sequencing library; identifying sequence reads in the sequence data that start within the primer region of the primer of each primer pair and are in the same direction as the primer; trimming the identified sequence reads in the same direction as the primer to exclude sequences within the primer region; identifying variant sequences in untrimmed sequence reads that extend into the primer region or are in a different direction from the primer, and variant sequences at locations in untrimmed sequence reads that correspond to or are complementary to the primer region.

[0008] The foregoing description is provided to enable the fabrication and use of the disclosed technology. Various modifications to the disclosed embodiments are apparent, and the general principles defined herein may be applied to other embodiments and uses without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but rather to be given the broadest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims.

[0009] These features, aspects, and advantages of the present invention, as well as other features, aspects, and advantages, will be better understood by reading the following detailed description with reference to the accompanying drawings, where similar features are represented in similar parts across the drawings. [Brief explanation of the drawing]

[0010] [Figure 1] This is a schematic diagram of the k-mer alignment workflow according to the aspects of this disclosure. [Figure 2]This is a schematic diagram of an exemplary k-mer of a genome according to an aspect of this disclosure. [Figure 3] This is a schematic diagram of a method for detecting viruses from sequencing data according to the aspects of this disclosure. [Figure 4] This is a schematic diagram of an alignment-based virus detection method according to the embodiments of this disclosure. [Figure 5] This is a schematic diagram of a target region or k-mer coverage in alignment-based virus detection according to the embodiments of this disclosure. [Figure 6] This is a schematic diagram of a method for generating a set of pathogen-specific k-mers and control k-mers for pathogen detection according to an aspect of the present disclosure. [Figure 7] This is a block diagram of a system configured to acquire sequencing data and perform alignment-based detection, according to the aspects of this disclosure. [Figure 8] This figure shows an exemplary workflow for sample preparation for pathogen detection. [Figure 9] This figure shows the sequencing results of the amplicon in the workflow shown in Figure 8. [Figure 10] This figure shows the identification of variants after primer trimming. [Modes for carrying out the invention]

[0011] The following considerations are presented to enable those skilled in the art to fabricate and use the disclosed technology and are provided in relation to specific uses and their requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and uses without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but is given the broadest scope consistent with the principles and features disclosed herein.

[0012] Various methods and configurations enabling the characterization of nucleic acids are described herein. In embodiments, the disclosed methods are used as part of sequence analysis of sequence data generated from a biological sample to rapidly and accurately detect a genome sequence of interest. In embodiments, the disclosed methods utilize an ultrafast hash-based aligner to generate low-error or error-free subsequences from sequence data. One application of the disclosed methods is the rapid detection of viral genomes present in a sequenced library. The method operates by scanning each k-mer of a fixed size "n" in all sequence reads of the sequenced library and checking for presence / absence in a hash table. The hash table is initialized with all n k-mers of the viral genome or a curated subset thereof. For example, curation can be used to remove k-mers that are not specific to the pathogen(s) of interest. The success of the sequence's k-mer match against the hash table is counted per viral k-mer.

[0013] In embodiments, pathogen infection is detected by a human positive control amplicon using a specialized aligner that employs rapid and complete k-mer matching of a complete set of virus-specific k-mers or a reduced (e.g., curated) set of k-mers. However, the disclosed method may be used in other applications such as the detection of germline variants in biological samples, microbiome characterization, and the detection of pooled or complex input samples in environmental monitoring (e.g., wastewater monitoring). Furthermore, the disclosed method may be used to detect a single pathogen of interest (e.g., SARS-CoV-2) or one or more pathogens in pathogen panels, e.g., respiratory pathogen panels (SARS-CoV-2, RSV, pneumonia, influenza), or strain tracking panels containing k-mers representing different strains of a particular pathogen.

[0014] Figure 1 shows an exemplary workflow 12 including a sample preparation step by sequencing analysis that may be used in conjunction with the disclosed method. The sample 20 undergoes processing or sample preparation 24 to generate a sequencing library containing a plurality of nucleic acid fragments suitable for sequencing step 28 to generate sequence data 30. The sequence data 30 may undergo certain primary analytical steps, such as quality control or filtering, before being sent for k-mer scanning and k-mer alignment, as generally provided herein.

[0015] The generated sequence data 30 is scanned to identify k-mers of fixed size n, and these identified k-mers are provided to a k-mer aligner 36. The k-mer aligner 36 may contain a hash table initialized with a set 34 of known k-mers of size n derived from a reference genome. The reference genome may be all k-mers of size n of the subject of the pathogen genome (or a curated subset thereof) or sequences of other subjects as provided herein.

[0016] The sequence data 30 may be fed to the k-mer aligner 36 in real time or sequentially (rolling basis), so that when the k-mer aligner 36 receives the sequence data 30, it operates on the additional sequence data 30 available in block 40 to detect the target k-mer in the sequence data 30. The k-mer aligner 36 identifies the k-mer in the sequence data 30 that perfectly match the set of target k-mers 34. A perfect match may relate to the total number of matches in the sample 20. When the sample 20 exceeds a threshold number of perfect matches for the identified k-mers, the workflow 12 provides a detection output 42. In an embodiment, individual samples 20 may be characterized as positive or negative with respect to the detection of sequences in set 34. Because the k-mer aligner 36 operates on data flowing in real time, the detection function can quickly determine the state of the sample 20 using perfect matches for k-mers as soon as the threshold number is exceeded. Furthermore, k-mer-based detection is less computationally intensive than conventional alignment-based methods, and in embodiments, less computationally intensive than other k-mer methods. In one example, the disclosed method uses a fixed k-mer size n. Thus, k-mer matching is based only on the matching of k-mers of size n, and not on the matching of all k-mers of all possible sizes or all k-mers within a certain range of k-mer sizes. In another example, within the set of all possible k-mers of fixed size n, the method evaluates matching only for a known subset based on known sequences in the reference genome.

[0017] The number of k-mers obtained for each sample 20 is used, as provided herein, to characterize the sample and provide a detection output 42 to determine, for example, the pathogen infection status. For example, a number of k-mers above a threshold indicates a positive result for the presence of a pathogen in the sample. A negative result indicates that the number of k-mers in the sample is not at or below the threshold level. The number of k-mers may be evaluated against a global threshold that reflects the total number of matching k-mers for each sample 20. In other embodiments, as disclosed herein, the number of k-mers may be evaluated against a target area criterion and / or undergo a quality metric before relating the number of k-mers to the detection of a pathogen, e.g., a positive or negative result.

[0018] The detection output 42 may, in embodiments, include providing a notification, message, or report indicating the characteristics of the sample 20, e.g., a positive detection result, a negative detection result. The detection output 42 may, in embodiments, control subsequent processing steps of the sequence data 30. In contrast to conventional alignment-based detection that sends all or most of the input data to secondary analysis, workflow 12 can restrict additional processing to a subset of samples that are positive for the pathogen or other genome / target sequence. That is, once identified, only positive samples 20 may be sent for additional or secondary sequencing analysis. In this way, workflow 12 improves the allocation of processing resources by not directing resources to secondary analysis of samples that may not contain the target sequence based on k-mer matching. Additional sequencing analysis may include determining the sequence of the biological sample in block 46 to generate a variant calling output 48. Thus, analyses that may be time-consuming, namely alignment and variant calling against a reference genome, may be limited to identified positive (e.g., infection) samples. Furthermore, samples 20 that have not yet been identified as positive can continue to be evaluated by the k-mer aligner 36 until sufficient data is obtained to confirm a negative or positive result. A further advantage of the disclosed method is that k-mer-based detection occurs in real time based on relatively rapid analysis. Therefore, initiating secondary analysis on relevant subsets of positive samples is possible with improved processing efficiency without significant delay. Moreover, depending on the analysis performed, the workflow 12 may terminate after the detection output 42 without proceeding to subsequent analysis or variant calling in block 46.

[0019] FIG. 2 is a schematic diagram of k-mers 64 of nucleic acids 60 that form a set 34 of target k-mers of a k-mer aligner 36 (see FIG. 1). The nucleic acid 60 can represent all or part of a reference genome or a previously characterized target genome, such as a pathogen genome. Thus, the disclosed techniques may be reference-free in the sense that the reference genome need not be sequenced along with the sample 20, and the set 34 can be computationally constructed based on the conserved or accessed reference sequence data of the nucleic acid 60. In embodiments, the nucleic acid 60 can be the reverse complement and / or cDNA copy of a single-stranded reference genome.

[0020] As provided herein, a k-mer or k-mers refers to a continuous substring or substrings of length "k" contained within a biological sequence such as a nucleic acid sequence. A set of k-mers can refer to all or only a portion of the sequences contained within a nucleic acid of length L. A known or identified sequence of length L will have all k-mers, and an un-identified or unknown sequence can have xk possible or potential k-mers, where x is the number of possible monomers (e.g., 4 for DNA or RNA).

[0021] In embodiments, the k-mer is used with a fixed size n, such that for a given operation, all k-mers used to construct the set 34 of k-mers and to scan the sequence data are of the same fixed size relative to each other. However, different k-mers of the same size represent different sequence strings at different or shifted locations relative to each other. In certain embodiments, k-mers having a length = 32 (which can be efficiently analyzed on a 64-bit CPU) are used for k-mer matching, although any size k-mer having a fixed length greater than 24 may be used. Thus, the fixed length of the k-mer can be 25, 26, 27, 28, 29, 30, etc.

[0022] Nucleic acid 60 may include previously identified sequences, but may also include further sequences such as known or predicted variants 70. The disclosed reference-less method has the advantage of the fact that variants in the viral genome are very few relative to the overall size of the virus. During k-mer alignment, k-mers from sequence data of samples containing / overlapping with variants are considered "lost" because they cannot have a perfect match in a hash table initialized with a set 34 that does not contain the reference k-mer variant. However, since variants are very few relative to the overall size of the virus, this simply minimizes the loss of sensitivity. In some methods, known variants present in the population may also be included as one or more "variant k-mers" 34 added to the set 34 of k-mers in the k-mer aligner 36.

[0023] Figure 3 shows an exemplary method 100 for detecting a viral pathogen in a human sample. The sequence data 102 of the human sample in the exemplary embodiment is provided as FASTQ format data, enabling secondary analysis and alignment of the sequence reads, for example, using DRAGEN or another secondary analysis tool. Alignment 104 of the sequence reads is performed using a k-mer aligner 36 (see Figure 1) to identify a complete k-mer match of fixed size n, using a set of reference k-mers based on the genome of the viral pathogen. Alignment 104 may also include identifying a complete k-mer match in the sequence data 102 with respect to one or more human control amplicons (e.g., 2 to 15 amplicons) used as a measure of sample quality. In some embodiments, alignment 104 may be a standard DRAGEN alignment to a reference genome containing a virus, e.g., SARS-CoV-2, and one or more human control amplicons.

[0024] The human reads 110 and virus reads 112 are subjected to additional metrics, as provided herein, to evaluate sample quality based on human amplicon coverage 114 and generate a control detection output 120. The metrics also include a virus amplicon coverage metric 130 for providing a virus detection output 132. Positive samples based on both the virus detection output and the control detection output 120 can be sent to variant calling 124 to generate a virus sequencing output 128.

[0025] Once the sequence reads are aligned / matched using the k-mer aligner 36, metrics related to the specified virus are interpreted, and a determination is made for the detection of the virus and the internal (human) control, as shown in Figure 4. In some methods, the number of unique reads 160 that map to the target region (or detected k-mer) of each amplicon can be counted.

[0026] As shown in Figure 5, the “target region” may, in embodiments, be defined as the sequence 184 of an amplicon 184, excluding primers and any overlap with other amplicons 184. This can be done either by a) aligning reads to the viral genome 180 and counting the number of reads 188 that map to the location of each amplicon (with any possible overlaps removed), or by b) counting the number of k-mers 190 from the sequence 184 of each amplicon observed in the reads. The number of k-mers or reads is compared to a threshold per amplicon coverage, and each amplicon 184 is referred to as “covered” or “not covered.” If more viral amplicons 184 are covered than a second set threshold, the call or virus detection output is that the virus is detected. The total number of amplicons depends on the assay used. In the example in Figure 5, the amplicons 184 are non-overlapping. However, it should be understood that more overlapping amplicons 184 can be used to achieve coverage of the entire viral genome.

[0027] Returning to Figure 4, after alignment and / or k-mer identification for human amplicon 162 and viral amplicon 164, the coverage 170 for each individual human amplicon and the coverage 172 for each viral amplicon are counted. The number of reads per amplicon (or the number of k-mers detected) is compared to a target threshold to determine the covered amplicons. The number of covered amplicons is then used to detect virus 178 (by positive detection results based on covered amplicons above the virus threshold) and internal (human) control 174 (by positive control detection results based on covered amplicons above the human control threshold). The thresholds for detecting positive amplicons, and the thresholds for the number of amplicons required to detect controls and / or viruses, may vary. In some embodiments, the detection threshold may be as little as two amplicons or as much as two, for example, three, four, or more amplicons. In embodiments, the threshold number of covered amplicons may be at least 1%, at least 10%, or at least 50% of the total number of amplicons. In embodiments, the threshold number of covered amplicons may range from 1% to 5% of the total number of amplicons in the assay. Since detection is designed to provide fast results for real-time sequencing data as additional sequencing data is generated from the sample, setting a percentage threshold allows detection to be performed based on any combination of positive amplicons. Therefore, detection is independent of sample variability in the location of sequenced clusters or other detection-specific variables that differ by sample.

[0028] Figure 4 shows an exemplary virus detection performed on human controls. For human control amplicons 170, control 1 had 25 unique reads, and control 3 had 64 unique reads exceeding the target threshold, and these amplicons were determined to be covered amplicons. In the next step, the two positive amplicons for the human controls were compared to a human control threshold set to 2 or higher, resulting in a control detection threshold exceedance determination 174. Thus, human control detection 174 involved a two-step analysis that determined the coverage of individual human amplicons based on the amplicon coverage threshold, and then evaluated the number of amplicons that exceeded the coverage threshold. Similarly, virus detection 178 involved a first step in which the number of unique reads for each virus amplicon (e.g., virus 1, virus 2, virus 3, etc.) was counted. Virus 1 had 34 unique reads, Virus 2 had 21 unique reads, and Virus Amplicon 3 had 64 unique reads; all were considered covered amplicons. However, Virus Amplicon 98, which had only one unique read, was not considered a covered amplicon. In the next step, the three covered amplicons were compared to a virus threshold set to 3 or higher to obtain virus detection results.

[0029] The disclosed method includes quality and control parameters for establishing a set of reference k-mers and / or control k-mers to be used in a k-mer aligner (e.g., k-mer aligner 36) for k-mer-based alignment. Figure 6 is a schematic diagram of method 200 for generating a set of pathogen-specific k-mers and control k-mers for pathogen detection. A given pathogen genome contains a set of all potential k-mers of fixed size n, where n can be more than 24 bases. However, certain k-mers may have a perfect match within a control genome (e.g., the human genome). In block 204, potential pathogen k-mers may be run against a control genome, and in block 206, specific k-mers may be removed to generate a final set of pathogen k-mers in block 208. In one example, k-mers in a set of potentials that have a perfect match against a control genome are removed. In another example, k-mers exceeding a similarity threshold with the control genome are removed. For example, k-mers exceeding the similarity threshold may include k-mers that have 1 to 3 bases (continuous or discontinuous) that differ from the control genome. For example, in the case of a fixed-size 32 k-mer, potential k-mers with 31 / 32 or 30 / 32 sequence matches with the control genome, which would correspond to potential base call errors that could lead to false positive detection, are removed. Thus, the k-mers retained in the final set in block 208 may include k-mers that do not have a perfect match with the control genome and / or k-mers that have sufficient dissimilarity with the control genome (e.g., 1 to 3 bases differ within the k-mer).

[0030] A set of control k-mers can be selected from a pool of potential k-mers based on a metric in block 210. In an assay that sequences RNA in a human sample to detect the presence of an RNA virus, the human sample will also contain human RNA, e.g., mRNA. Therefore, a set of human control k-mers may be based on mRNA sequences that are likely to be consistently expressed in the tissue of the sample. The set of control k-mers may be selected to be smaller than the reference set, for example, containing fewer amplicons. In block 214, the potential set of control k-mers is run against each other, in embodiments against the reference genome, and control k-mers that are a perfect match or too similar to each other and to the control genome (e.g., differing by 1-3 bases but otherwise a perfect match) are removed in step 216 to produce the final set of control k-mers in block 218. In block 220, the final set of pathogen k-mers and the final set of control k-mers are provided to the k-mer aligner.

[0031] Figure 7 is a schematic diagram of a sequencing device 260 that may be used in conjunction with the disclosed embodiments for obtaining sequence data from a sample, as provided herein. The sequencing device 260 can perform a sequencing run on the sample to obtain sequence data. The sequencing device 260 may be implemented according to any sequencing method, such as those incorporating the sequencing-by-synthesis method described in U.S. Patent Publications 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2005 / 0100900, U.S. Patent No. 7,057,026, and International Publications 05 / 065814, 06 / 064199, and 07 / 010251, the entire disclosure of which is incorporated herein by reference. Alternatively, sequencing by a ligation method may be used in the sequencing device 260. Such a method involves incorporating oligonucleotides using a DNA ligase and identifying the incorporation of such oligonucleotides, and is described in U.S. Patents 6,969,488, 6,172,218, and 6,306,597, the entire disclosures of which are incorporated herein by reference. In some embodiments, nanopore sequencing can be utilized, in which nucleic acid strands or nucleotides of a sample are removed from the nucleic acids of the sample by an exonuclease and pass through nanopores. As nucleic acids or nucleotides of a sample pass through nanopores, each base species can be identified by measuring the variation in the electrical conductance of the pores (the entire disclosure is incorporated herein by reference: U.S. Patent No. 7,001,792, Soni & Meller, Clin. Chem. 53, 1996-2001 (2007), Healy, Nanomed. 2, 459-481 (2007), Cockroft et al. J. Am. Chem. Soc. 130, 818-820 (2008)). Further embodiments include the detection of protons released upon incorporation of nucleotides into the extension product.For example, sequencing based on the detection of released protons may utilize electrodetectors and related technologies commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies), or sequencing methods and systems described in U.S. Patent Publications 2009 / 0026082(A1), 2009 / 0127589(A1), 2010 / 0137143(A1), or 2010 / 0282617(A1), the entirety of which are incorporated herein by reference. Certain embodiments may utilize methods that include real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interaction between fluorophore-supported polymerase and γ-phosphate-labeled nucleotides, or using zero-mode waveguides as described in, for example, Levene et al. Science 299, 682-686 (2003), Lundquist et al. Opt. Lett. 33, 1026-1028 (2008), and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the entirety of which is incorporated herein by reference. Other suitable alternatives include, for example, fluorescence in situ sequencing (FISSEQ) and massively parallel signature sequencing (MPSS). In certain embodiments, the sequencing device 260 may be an iSeq manufactured by Illumina (La Jolla, CA). In other embodiments, the sequencing device 260 may be configured to operate using a CMOS sensor with nanowells fabricated on photodiodes so that DNA deposition is aligned one-to-one with each photodiode.

[0032] In the described embodiment, the sequencing device 260 comprises a separate sample substrate 262, e.g., a flow cell or sequencing cartridge, and an associated computer 264. However, as described above, these may be implemented as a single device. In the described embodiment, a biological sample may be loaded onto the substrate 262, which is imaged to generate sequence data. For example, a reagent interacting with the biological sample may fluoresce at a specific wavelength in response to an excitation beam generated by the imaging module 272, thereby returning radiation for imaging. For example, the fluorescent component may be generated by a fluorescently tagged nucleic acid, which hybridizes to a complementary molecule of the component or to a fluorescently tagged nucleotide incorporated into an oligonucleotide using polymerase. As will be understood by those skilled in the art, the wavelength at which the dyes in the sample are excited, and the wavelength at which they fluoresce, will depend on the absorption and emission spectra of the particular dye. Such returned radiation may propagate through a directional optical system. This retrobeam may be directed roughly to a detection optical system of the imaging module 272, which may be a camera or other optical detector.

[0033] The imaging module detection optics may be based on any preferred technique, for example, a charged coupled device (CCD) sensor that generates pixelated image data based on photons affecting location within the device. However, it will be understood that any of a variety of other detectors may be used, including but not limited to detector arrays configured for time delay integration (TDI) operation, complementary metal oxide semiconductor (CMOS) detectors, avalanche photodiode (APD) detectors, Geiger-mode photon counters, or any other suitable detectors. TDI mode detection can be coupled with line scanning, as described in U.S. Patent No. 7,329,860, incorporated herein by reference. Other useful detectors are described, for example, in the references previously provided herein in the context of various nucleic acid sequencing methodologies.

[0034] The imaging module 272 may be under processor control, for example, by a processor 274, and may include an I / O control unit 276, an internal bus 278, a non-volatile memory 280, RAM 282, and any other memory configuration in which the memory can store executable instructions, as well as other suitable hardware components, which may be similar to those described with respect to Figure 7. Furthermore, the associated computer 264 may also include a memory architecture including a processor 184, an I / O control unit 286, a communication module 294, and RAM 288 and non-volatile memory 290, so that the memory architecture can store executable instructions 292. The hardware components may be connected by an internal bus 294, which can also be connected to a display 296. In embodiments in which the sequencing device 260 is implemented as an all-in-one device, certain redundant hardware elements can be eliminated.

[0035] A processor (e.g., processors 274, 284) may be programmed to assign individual sequencing reads to a sample based on the relevant index sequence or sequence by the method provided herein. In certain embodiments, based on image data acquired by the imaging module 272, the sequencing device 260 may be configured to generate sequencing data including sequence reads for individual clusters, each sequence read associated with a specific location on the substrate 270. Each sequence read may originate from a fragment containing an insert. The sequencing data includes base calls for each base of the sequencing read. Furthermore, for sequentially performed sequencing reads based on image data, individual reads may also be linked to the same location, and therefore the same template strand, via the image data. In this way, index sequencing reads may be associated with sequencing reads of insert sequences before being assigned to the original sample. The processor 274 may also be programmed to perform downstream analysis on the sequences of a particular sample following the assignment of sequencing reads to the sample.

[0036] ). In certain embodiments, executable instruction 292 causes the processor to perform one of more actions of the method disclosed herein. The processor (e.g., processors 274, 284) may be a highly reconfigurable field-programmable gate array (FPGA) technology. The processor (e.g., processors 274, 284) may be programmed to receive user input for a particular analytical workflow in order to access a hash table containing an appropriate set of reference k-mers and / or reference k-mers stored in memory (e.g., memory 280, 290). In one example, device 260 receives user input to select a run or panel of interest, and the k-mer aligner uses the hash table associated with the user input to align the streaming sequence to identify a perfect match of k-mers in the sequence data. The memory may store several different sets of k-mers or different initialization hash tables that are specifically selected based on the user input. In embodiments, the selection may also include the selection of reference k-mers. For example, control k-mers may include control k-mers from humans, mammals, or other host organisms.

[0037] The disclosed methods may be used to characterize samples, for example, biological samples. Samples may originate from any in vivo or in vitro source, including one or more cells, tissues, organs, or organisms, whether living or dead, or any biological or environmental source (e.g., water, air, soil). For example, in some embodiments, the nucleic acids of the sample include or consist of dsDNA of eukaryotes and / or prokaryotes originating from or derived from humans, animals, plants, fungi (e.g., mold or yeast), bacteria, viruses, viroids, mycoplasmas, or other microorganisms. In some embodiments, the nucleic acid of the sample includes or comprises genomic DNA, subgenomic DNA, chromosomal DNA (e.g., derived from an isolated chromosome or a portion of a chromosome, e.g., from one or more genes or loci derived from a chromosome), mitochondrial DNA, chloroplast DNA, plasmid or other episome-derived DNA (or recombinant DNA contained therein), or double-stranded cDNA prepared by reverse transcribing RNA using RNA-dependent DNA polymerase or reverse transcriptase to produce a first-strand cDNA, and then extending primers annealed to the first-strand cDNA to produce dsDNA. In some embodiments, the nucleic acid of the sample includes a plurality of dsDNA molecules in or prepared from nucleic acid molecules (e.g., a plurality of dsDNA molecules in or prepared from cDNA prepared from or derived from biological (e.g., cells, tissues, organs, organisms) or environmental (e.g., water, air, soil, saliva, sputum, urine, feces) sources of genomic DNA or RNA). In some embodiments, the nucleic acid of the sample is derived from an in vitro source. For example, in some embodiments, the nucleic acid of the sample includes or consists of single-stranded DNA (ssDNA) or single-stranded or double-stranded RNA (dsDNA prepared in vitro using methods well known in the art, such as primer extension using a suitable DNA-dependent and / or RNA-dependent DNA polymerase (reverse transcriptase)).In some embodiments, the nucleic acid of the sample includes or consists of dsDNA prepared from all or part of one or more double-stranded or single-stranded DNA or RNA molecules using any method known in the art, including methods for: amplification of DNA or RNA (e.g., PCR or reverse transcriptase PCR (RT-PCR), transcription-mediated amplification, involving amplification of all or part of one or more nucleic acid molecules); molecular cloning of all or part of one or more nucleic acid molecules in a plasmid, fosmid, BAC, or other vector to be replicated in a suitable host cell; or capture of one or more nucleic acid molecules by hybridization, such as hybridization to a DNA probe on an array or microarray.

[0038] The advantages of the disclosed method include the suppression of noise (e.g., cross-contamination) that appears as uniformly scattered reads throughout the viral genome, in contrast to the actual signals clustered by the amplicons. The method is adaptable to different amplicons with varying PCR performance by setting a variable threshold per amplicon (higher for strongly amplified amplicons). The disclosed method closely corresponds to existing qPCR tests that also report the number of positive amplicons and therefore provides output results readily convertible for clinical use. Detection output may be reported per sample, or the detection output may undergo downstream quality control.

[0039] In some embodiments, variant calling data for any positive sample may also be reported. In some embodiments, a positive sample may be identified, and this method includes providing notification or recommendations for treatment based on the diagnosis of the positive sample. In embodiments, treatment for the detected pathogen is administered to the patient from whom the sample was taken, based on a diagnosis of pathogen detection or non-detection by the disclosed method, which is used as a point-of-care detection system. For example, if the detected pathogen is based on the detection of the SARS-CoV-2 genome, treatment for SARS-CoV-2 is administered or a surveillance protocol is initiated. If the SARS-CoV-2 genome is not detected, a SARS-CoV-2 vaccine may be administered based on a diagnosis that the patient does not have an active infection.

[0040] Further advantages of the disclosed method include real-time quality metrics and variant detection generated on the instrument. In the example in Figure 7, real-time quality metrics are generated on the sequencing device 260 and not as part of a cloud-based secondary analysis. In certain embodiments, sequence data may be analyzed based on the presence and distribution of variants, which may include alternative alleles or single nucleotide polymorphisms (SNPs). The assay may include amplicon generation and / or targeted sequencing based on the analysis of desired variants or SNPs. For any specific allele to be detected, the distribution of the allele in the sequence reads can be evaluated to obtain quality metrics on the instrument. Allele detection may be alignment-based or may use k-mer matching as provided herein. The k-mer approach may be extended to detect known variants by including alternative allele versions of k-mers, where the allele k-mer may represent a reference or an alternative allele. The reference allele k-mer and alternative allele k-mer may include their respective sets of k-mers extending to the variant sequence or sequence location. These modifications allow the algorithm to run without alignment (and therefore faster) while the sequencing device 260 is still generating data.

[0041] For a given variant and a given individual sample, the distribution of alleles can follow a predictable level. For example, a particular germline variant allele is likely to have a 50% distribution (50% of sequenced reads where one allele is present, and the remaining 50% where the other allele is present) or 100% in the reads, if present. Furthermore, if the germline variant is not detected, it is likely to be 0% in the sequencing reads. Thus, variant-to-reference ratios of 1:1 and 1:0 can be considered within the expected distribution of detected germline variants. However, distributions of 80%–20% or 95%–5% in the reads for an individual sample are biologically unlikely and are therefore potentially the result of error or contamination. Therefore, ratios deviating from a 1:1 or 1:0 ratio (e.g., within an acceptable range of 5%–10% considering sequencing errors) are likely sequencing artifacts and / or based on sample contamination. Therefore, the sequencing device 260 can evaluate the sample with respect to the germline variant allele distribution based on variant detection within the sequenced reads.

[0042] For a given variant panel, e.g., an SNP panel, only some of the variants may be matched to a particular sample. However, for variants detected that deviate from the expected allele distribution, the anomalous distribution may indicate sample contamination, patient identification errors or sample identification errors in assigning sample reads, or problems with sample preparation. Therefore, samples containing variants with anomalous or low-frequency distributions (e.g., 95%–5%) may be flagged. In response to flagging, the sequencing device 260 may display an error message on a graphical user interface (e.g., a displayed notification) that identifies potentially contaminated samples in real time. Thus, the disclosed method includes a real-time sample quality metric of the sequencing device 260. Samples may be indicated as pass or fail depending on one or more evaluated allele distributions. In embodiments, it is sufficient to flag a sample if only one allele distribution is failing. In the case of k-mer-based detection, the k-mer set generated by the calculation of variants or alternative alleles may be updated when new variants or strains are tracked.

[0043] By identifying flagged or unacceptable samples based on abnormal allele distributions, the sequencing device 260 can stop the transmission of relevant sequence data from a sample to cloud-based secondary analysis. Therefore, for multiple samples or multiplexed runs, the sequencing device 260 can transmit only the samples that pass for further analysis to the cloud. If multiple samples all contain the same abnormal allele distribution, the entire multiplexed run may be flagged as potentially contaminated.

[0044] In embodiments, the disclosed method includes an improvement in the detection of variants that may be masked in sequencing data based on primer design or location. For example, variants in genomic regions corresponding to primer regions may be identified based on overlapping amplicon designs, thereby covering the primer regions with genomic reads derived from the overlapping amplicons. Figure 8 shows an exemplary workflow for sample preparation for pathogen detection that may be used in conjunction with the disclosed variant detection method. In the illustrated example, the sample is processed in block 300 to extract RNA. RNA can be extracted from a sample such as a nasopharyngeal swab.

[0045] The extracted RNA is converted to cDNA, and the cDNA is used to generate amplicons using an assay-specific primer set. For example, for COVIDSeq applications, the cDNA is split into two parts, and two different primer pools are used to generate different overlapping amplicons 304 between the two parts. Each sample is indexed via tagmentation, for example, in step 308, and sequenced in step 310.

[0046] Figure 9 shows the sequencing results of amplicons from the workflow in Figure 8, illustrating overlapping coverage for amplicon 314 generated from a first primer pool and amplicon 316 generated from a second primer pool, both indexed together, for example, as originating from the same sample. Read 320 from pool 1 and read 324 from pool 2 contain overlapping portions at their edges within the primer region. Forward read 326 and reverse read 328 are present within read 320 from pool 1 and read 324 from pool 2. Post-PCR fragmentation has the effect of partially depleted primers, resulting in an edge effect where reads cluster toward the primer side in both the forward and reverse directions. Sequencing reads in the overlapping region contain a heterogeneous mixture of genomic reads with variants and primer reads with reference sequences.

[0047] However, because primer reads represent the expanded portion of the mixture due to edge effects, clustering of primer reads toward the amplicon edges can reduce the observed fraction of alternative alleles. For example, forward primer 330 overlaps with the internal region of another amplicon 334. All reads 320 in pool 1 are forward reads 326 derived from primer 330. Reads 324 in pool 2 contain both forward reads 326 and reverse reads 328. Reads 320 in pool 1 derived from primer 330 are considered to be a perfect match of the primer and therefore are not considered to contain variants present in the genomic region covered by primer 330.

[0048] To improve the sensitivity of variant detection, the disclosed method includes a primer trimming step of hard clipping, masking, or removing primer sequences from the reads. The filter trims 1) reads that start in the primer region and 2) reads that are consistent with the primer orientation. That is, any sequence read having a first nucleotide that starts in the region covered by the primer and is a forward read in the forward primer direction or a reverse read in the reverse primer direction is trimmed. However, coverage in the primer region remains from overlapping amplicons extending to the read and any opposite (complementary) reverse reads. Figure 10 shows an example of trimmed reads in the primer-covered region of the reactant. The read mixture includes untrimmed or retained reads 350 and trimmed reads 352. Reads 352 are trimmed only in sequences corresponding to the primer region 354, indicated by start X and end X, and only forward reads are trimmed. As shown, retained reads 350 mainly contain variants from G to T. The T variant is not observed in most trimmed reads. In one example, the variant is called based on a threshold percentage (e.g., at least 50%) of 350 untrimmed reads that contain the variant.

[0049] Table 1 shows an example of improved detection of a single G-to-T variant. After filtering and trimming, the proportion of the remaining alleles converged toward nearly 100% allele proportions in the remaining reads, which is considered to be the expected biological distribution.

[0050] [Table 1]

[0051] Although the described embodiments demonstrate trimming of a single primer, primer trimming can be used to cover all, i.e., both forward and reverse primers, in the reactant to improve variant identification in any region covered by the primers. For whole-genome sequencing of pathogens where multiple, e.g., 50 or more primer pairs are used, primer trimming can significantly improve variant detection.

[0052] This written description uses examples, including best-case examples, of the embodiments of the disclosure and enables any person skilled in the art to practice the disclosed embodiments, including by fabricating and using any device or system and by performing any incorporated methods. The patentable scope of the disclosure is defined by the claims and may include other examples that a person skilled in the art may conceive. Such other examples are intended to be within the claims if they include structural elements that are no different from the literal words of the claims, or if they include equivalent structural elements that differ only slightly from the literal words of the claims. [Explanation of symbols]

[0053] 12 Workflows 20 samples 24. Sample preparation 28 Sequencing 30 Sequence Data 36 k-mer Alaina 42 Detection Output 48 Variant Calling Output 60 Nucleic acids 70 Known or predicted variants 102 Sequence Data 104 Alignment 110 human leads 112 Virus Leads 114 Human amplicon coverage 120 Control Detection Output 124 Variant Calling 128 Virus Sequencing Output 130 Virus Amplicon Coverage Metrics 132 Virus detection output 160 unique reads 162 Human amplicons 164 Virus Amplicon 170 Human amplicon coverage 172 Virus Amplicon Coverage 174 Human control detection 178 viruses detected 180 viral genomes 260 Sequencing Devices 262 Sample Substrate 264 Computers 270 Base material 272 Imaging Module 274 processors 276 I / O Control Unit 278 Internal Bus 280 Non-volatile memory 284 processors 286 I / O Control Unit 290 Non-volatile memory 292 Executable Instructions 294 Communication Module 296 displays 314 Amplicons generated from the first primer pool 316 Amplicon generated from the second primer pool 320 Reeds 324 Reeds 326 Forward Lead 328 Reverse Lead 330 Forward Primer 350 Untrimmed or retained leads 352 Trimmed Lead 354 Primer region

Claims

1. A real-time quality control method, The steps include generating sequence data from a biological sample using a sequencing device to perform a sequencing run, and A step of identifying a k-mer in the sequence data that has a perfect match in a hash table initialized with a set of k-mers including a reference allele k-mer and alternative alleles k-mer of the reference allele, A step of determining the distribution of the reference allele and the substitute allele in the sequence data based on the number of perfect matches, During the sequencing run of the biological sample, the steps include generating a quality metric for the biological sample based on the distribution, Methods that include...

2. The method according to claim 1, comprising the step of flagging the biological sample as contaminated based on the quality metric.

3. The method according to claim 2, wherein the alternative allele is present in 5% or less of the sequence reads of the sequence data in the contaminated sample.

4. The method according to claim 1, further comprising the step of indicating that the biological sample has passed the quality metric based on the fact that the ratio of the reference allele to the surrogate allele in the sequencing data is within the expected range.

5. The method according to claim 1, wherein the quality metric is generated on the sequencing device.

6. The method according to claim 1, wherein the alternative allele includes a previously characterized single nucleotide polymorphism.

7. A sequencing device, A substrate loaded with a sequencing library prepared from a sample, It is a computer, The sequencing device is made to perform a sequencing run to generate sequence data from the sequencing library, Identifying k-mers in the sequence data that have a perfect match in a hash table initialized with a set of k-mers including a reference allele k-mer and alternative alleles k-mer of the reference allele, Based on the number of perfect matches, the distribution of the reference allele and the substitute allele in the sequence data is determined. During the sequencing run, a quality metric for the biological sample in the sequencing device is generated based on the distribution. A computer programmed to do this, A sequencing device equipped with [a specific feature].

8. The sequencing device according to claim 7, further comprising a display for displaying the quality metrics.

9. The sequencing device according to claim 7, further comprising a communication circuit for transmitting the generated sequence data to a cloud computing environment based on the fact that the quality metric of the biological sample is associated with passing.

10. The sequencing device according to claim 9, wherein the quality metric of the biological sample is associated with the ratio of the reference allele to the substitute allele being within the expected range.

11. The sequencing device according to claim 7, further comprising a communication circuit that stops the transmission of the generated sequence data to a cloud computing environment based on the fact that the quality metric of the biological sample is associated with failure.

12. The sequencing device according to claim 11, wherein the quality metric of the biological sample is associated with failure based on the presence of the alternative allele in 5% or less of the sequenced reads of the sequence data.

13. A method for detecting variants in biological samples, A step of generating an amplicon from a biological sample using a primer pair, The steps include preparing a sequencing library from the generated amplicons, The steps include generating sequence data from the aforementioned sequencing library, A step of identifying sequence reads in the sequence data that start within the primer region of the primer of each primer pair and are in the same direction as the primer, The steps include trimming the identified sequence reads that are in the same direction as the primer to exclude the sequence within the primer region, The steps include identifying variant sequences in untrimmed sequence reads that extend to the primer region or are in a direction different from the primer, and identifying variant sequences at locations in the untrimmed sequence read that correspond to or are complementary to the primer region, Methods that include...

14. The method according to claim 13, comprising the steps of extracting RNA from the biological sample and converting the RNA to cDNA before generating the amplicon.

15. The method according to claim 13, wherein the amplicon includes a duplicate portion of the reference genome.

16. The method according to claim 15, wherein the reference genome is a pathogen genome.

17. The method according to claim 15, wherein the reference genome is the SARS-CoV-2 genome.

18. The method according to claim 16, wherein the reference genome is the human genome.

19. The method according to claim 13, further comprising the step of calling the identified variant sequence based on the fact that the identified variant sequence is present in at least 50% of the untrimmed sequence reads at the location.

20. The method according to claim 13, comprising the steps of: identifying sequence reads in the sequence data that start within the reverse primer region of the reverse primer of the primer pair and are in the same direction as the reverse primer; and trimming the identified sequence reads to exclude sequences within the reverse primer region.