Genome sequencing and detection methods

JP7918098B2Active Publication Date: 2026-09-09ILLUMINA INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022567152
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-08
Filing Date
2021-05-07
Publication Date
2026-09-09
Estimated Expiration
2041-05-07

Smart Images

  • Figure 0007918098000001
    Figure 0007918098000001
  • Figure 0007918098000002
    Figure 0007918098000002
  • Figure 0007918098000003
    Figure 0007918098000003
Patent Text Reader

Abstract

A nucleic acid sequencing method is described. For example, sequence data generated by a sequencing device can be analyzed by scanning k-mers of a fixed size n in each read in the sequence data. Perfect matches of the k-mers in the sequence data with reference k-mers are identified. The number of perfect matches, the distribution of perfect matches in the reference genome, and / or the number of sequence reads in the sequence data that map to different target regions can be used to determine the characteristics of the sample. In one example, the characteristic is the presence of a pathogen in the sample.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 022,296, filed on May 8, 2020, the disclosure of which is incorporated herein by reference. BACKGROUND ART

[0002] The disclosed technology relates generally to nucleic acid characterization, for example, sequencing methods. In some embodiments, the disclosed technology includes rapid and accurate methods for virus detection from sequence data based on genomic sequencing, for example, whole genome sequencing.

[0003] The subject matter discussed in this section should not be construed as prior art merely by virtue of its mention in this section. Similarly, any problems mentioned in this section or problems associated with subject matter provided as background should not be construed as having been previously recognized in the prior art. The subject matter in this section merely represents different approaches, and may itself also correspond to embodiments of the claimed technology.

[0004] Next-generation sequencing technologies have made sequencing increasingly faster and enabled deeper sequencing depth. However, sequencing accuracy and sensitivity are affected by errors and noise from various sources, for example, sample defects during library preparation or PCR bias. Accordingly, detection of very low-frequency sequences such as host samples containing low concentrations of viral or bacterial nucleic acid can be complex. Therefore, it is desirable to develop methods for detecting and / or sequencing nucleic acid molecules present in small amounts in a rapid and accurate manner. SUMMARY OF THE INVENTION

[0005] In one embodiment, the present disclosure relates to a method for detecting pathogens in a biological sample. The method includes: receiving sequence data derived from a biological sample; identifying k-mers in the sequence data that have a perfect match in a hash table initialized with a first set of k-mers containing pathogen k-mers and a second set of k-mers containing control k-mers in the genome of a pathogen; and providing a detection output for the biological sample based at least in part on a first number of perfect matches of k-mers in the sequence data with the first set, and a second number of perfect matches of k-mers in the sequence data with the second set, wherein if the first number exceeds a first set threshold and the second number exceeds a second set threshold, the detection output includes a positive result for pathogen detection; and if the first number falls below the first set threshold, the second number falls below the second set threshold, or both, the detection output includes a negative result for pathogen detection.

[0006] In another embodiment, the present disclosure relates to a method for detecting pathogens in a biological sample. The method includes: generating sequence data from a sequencing library prepared from a biological sample; identifying k-mers in the sequence data that have a perfect match in a hash table initialized with a set of k-mers containing pathogen k-mers in the pathogen genome of the pathogen; determining the coverage in the sequence data for individual target regions of the pathogen genome based on either or the number of identified k-mers in the sequence data containing the identified k-mers corresponding to each individual target region of the pathogen genome, wherein each target region is determined to be covered when the number of identified k-mers or the number of sequence reads corresponding to the individual target region exceeds a threshold number; determining that the number of covered individual target regions exceeds a detection threshold; and providing a detection output indicating that the biological sample is positive for the presence of a pathogen.

[0007] In another embodiment, the present disclosure relates to a sequencing device comprising a substrate loaded with a sequencing library prepared from a sample. The sequencing device also comprises a computer programmed to cause the sequencing device to generate sequence data from the sequencing library, to scan k-mers of fixed size n in individual reads in the sequence data, to access a hash table stored in the computer's memory, the hash table being initialized with a set of reference k-mers of fixed size n, and to use the hash table to identify a perfect match of a k-mer with the set of reference k-mers, and to determine the characteristics of the sample based on whether the number of identified perfect matches exceeds a threshold.

[0008] The foregoing description is provided to enable the fabrication and use of the disclosed technology. Various modifications to the disclosed embodiments are apparent, and the general principles defined herein may be applied to other embodiments and uses without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but rather to be given the broadest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims. [Brief explanation of the drawing]

[0009] These features, aspects, and advantages of this disclosure, as well as other features, aspects, and advantages, will be better understood by reading the following detailed description with reference to the accompanying drawings, where similar features are represented in similar parts across the drawings. [Figure 1] This is a schematic diagram of the k-mer alignment workflow according to the aspects of this disclosure. [Figure 2] This is a schematic diagram of an exemplary k-mer of a genome according to an aspect of this disclosure. [Figure 3] This is a schematic diagram of a method for detecting viruses from sequencing data according to the aspects of this disclosure. [Figure 4] This is a schematic diagram of an alignment-based virus detection method according to the embodiments of this disclosure. [Figure 5] This is a schematic diagram of a target region or k-mer coverage in alignment-based virus detection according to the embodiments of this disclosure. [Figure 6] This is a schematic diagram of a method for generating a set of pathogen-specific k-mers and control k-mers for pathogen detection according to an aspect of the present disclosure. [Figure 7] This is a block diagram of a system configured to acquire sequencing data and perform alignment-based detection, according to the aspects of this disclosure. [Modes for carrying out the invention]

[0010] The following considerations are presented to enable those skilled in the art to fabricate and use the disclosed technology and are provided in relation to specific uses and their requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and uses without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but is given the broadest scope consistent with the principles and features disclosed herein.

[0011] Various methods and configurations enabling the characterization of nucleic acids are described herein. In embodiments, the disclosed methods are used as part of sequence analysis of sequence data generated from a biological sample to rapidly and accurately detect a genome sequence of interest. In embodiments, the disclosed methods utilize an ultrafast hash-based aligner to generate low-error or error-free subsequences from sequence data. One application of the disclosed methods is the rapid detection of viral genomes present in a sequenced library. The method operates by scanning each k-mer of a fixed size "n" in all sequence reads of the sequenced library and checking for presence / absence in a hash table. The hash table is initialized with all n k-mers of the viral genome or a curated subset thereof. For example, curation can be used to remove k-mers that are not specific to the pathogen(s) of interest. The success of the sequence's k-mer match against the hash table is counted per viral k-mer.

[0012] In embodiments, pathogen infection is detected by a human positive control amplicon using a specialized aligner that employs rapid and complete k-mer matching of a complete set of virus-specific k-mers or a reduced (e.g., curated) set of k-mers. However, the disclosed method may be used in other applications such as the detection of germline variants in biological samples, microbiome characterization, and the detection of pooled or complex input samples in environmental monitoring (e.g., wastewater monitoring). Furthermore, the disclosed method may be used to detect a single pathogen of interest (e.g., SARS-CoV-2) or one or more pathogens in pathogen panels, e.g., respiratory pathogen panels (SARS-CoV-2, RSV, pneumonia, influenza), or strain tracking panels containing k-mers representing different strains of a particular pathogen.

[0013] Figure 1 shows an exemplary workflow 12 including a sample preparation step by sequencing analysis that may be used in conjunction with the disclosed method. The sample 20 undergoes processing or sample preparation 24 to generate a sequencing library containing a plurality of nucleic acid fragments suitable for sequencing step 28 to generate sequence data 30. The sequence data 30 may undergo certain primary analytical steps, such as quality control or filtering, before being sent for k-mer scanning and k-mer alignment, as generally provided herein.

[0014] The generated sequence data 30 is scanned to identify k-mers of fixed size n, and these identified k-mers are provided to a k-mer aligner 36. The k-mer aligner 36 may contain a hash table initialized with a set 34 of known k-mers of size n derived from a reference genome. The reference genome may be all k-mers of size n of the subject of the pathogen genome (or a curated subset thereof) or sequences of other subjects as provided herein.

[0015] The sequence data 30 may be fed to the k-mer aligner 36 in real time or sequentially (rolling basis), so that when the k-mer aligner 36 receives the sequence data 30, it operates on the additional sequence data 30 available in block 40 to detect the target k-mer in the sequence data 30. The k-mer aligner 36 identifies the k-mer in the sequence data 30 that perfectly match the set of target k-mers 34. A perfect match may relate to the total number of matches in the sample 20. When the sample 20 exceeds a threshold number of perfect matches for the identified k-mers, the workflow 12 provides a detection output 42. In an embodiment, individual samples 20 may be characterized as positive or negative with respect to the detection of sequences in set 34. Because the k-mer aligner 36 operates on data flowing in real time, the detection function can quickly determine the state of the sample 20 using perfect matches for k-mers as soon as the threshold number is exceeded. Furthermore, k-mer-based detection is less computationally intensive than conventional alignment-based methods, and in embodiments, less computationally intensive than other k-mer methods. In one example, the disclosed method uses a fixed k-mer size n. Thus, k-mer matching is based only on the matching of k-mers of size n, and not on the matching of all k-mers of all possible sizes or all k-mers within a certain range of k-mer sizes. In another example, within the set of all possible k-mers of fixed size n, the method evaluates matching only for a known subset based on known sequences in the reference genome.

[0016] The number of k-mers obtained for each sample 20 is used, as provided herein, to characterize the sample and provide a detection output 42 to determine, for example, the pathogen infection status. For example, a number of k-mers above a threshold indicates a positive result for the presence of a pathogen in the sample. A negative result indicates that the number of k-mers in the sample is not at or below the threshold level. The number of k-mers may be evaluated against a global threshold that reflects the total number of matching k-mers for each sample 20. In other embodiments, as disclosed herein, the number of k-mers may be evaluated against a target area criterion and / or undergo a quality metric before relating the number of k-mers to the detection of a pathogen, e.g., a positive or negative result.

[0017] The detection output 42 may, in embodiments, include providing a notification, message, or report indicating the characteristics of the sample 20, e.g., a positive detection result, a negative detection result. The detection output 42 may, in embodiments, control subsequent processing steps of the sequence data 30. In contrast to conventional alignment-based detection that sends all or most of the input data to secondary analysis, workflow 12 can restrict additional processing to a subset of samples that are positive for the pathogen or other genome / target sequence. That is, once identified, only positive samples 20 may be sent for additional or secondary sequencing analysis. In this way, workflow 12 improves the allocation of processing resources by not directing resources to secondary analysis of samples that may not contain the target sequence based on k-mer matching. Additional sequencing analysis may include determining the sequence of the biological sample in block 46 to generate a variant calling output 48. Thus, analyses that may be time-consuming, namely alignment and variant calling against a reference genome, may be limited to identified positive (e.g., infection) samples. Furthermore, samples 20 that have not yet been identified as positive can continue to be evaluated by the k-mer aligner 36 until sufficient data is obtained to confirm a negative or positive result. A further advantage of the disclosed method is that k-mer-based detection occurs in real time based on relatively rapid analysis. Therefore, initiating secondary analysis on relevant subsets of positive samples is possible with improved processing efficiency without significant delay. Moreover, depending on the analysis performed, the workflow 12 may terminate after the detection output 42 without proceeding to subsequent analysis or variant calling in block 46.

[0018] Figure 2 is a schematic diagram of the k-mer 64 of nucleic acid 60 that forms the set 34 of target k-mers in k-mer aligner 36 (see Figure 1). Nucleic acid 60 can represent a reference genome or a pre-characterized target genome, for example, all or part of a pathogen genome. Therefore, the disclosed method may be reference-free in the sense that the reference genome does not need to be sequenced together with the sample 20, and set 34 can be constructed computationally based on stored or accessed reference sequencing data of nucleic acid 60. In embodiments, nucleic acid 60 may be the reverse complement and / or cDNA copy of a single-stranded reference genome.

[0019] As provided herein, a k-mer or k-mers refers to a sequence of "k" substrings of length "k" contained within a biological sequence, such as a nucleic acid sequence. A set of k-mers may refer to all or only some of the sequences contained within a nucleic acid of length L. A known or identified sequence of length L will have all k-mers, while an unidentified or unknown sequence will have x k It may have n possible k-mers or potential k-mers, where x is the number of possible monomers (e.g., 4 for DNA or RNA).

[0020] In some embodiments, the k-mer is used with a fixed size n, and as a result, for a given operation, all k-mers used to construct a set 34 of k-mers and to scan the sequence data are of the same fixed size relative to one another. However, different k-mers of the same size represent different sequence strings at different or shifted locations relative to one another. In certain embodiments, k-mers with length = 32 (which can be efficiently analyzed on a 64-bit CPU) are used for k-mer matching, but k-mers of any size with a fixed length greater than 24 may be used. Thus, the fixed length of the k-mer may be 25, 26, 27, 28, 29, 30, etc.

[0021] The nucleic acid 60 may comprise a previously identified sequence, and may also comprise additional sequences such as known or predicted variants 70. The disclosed reference-free approach exploits the fact that the number of variants in the viral genome is very small relative to the overall size of the virus. During k-mer alignment, k-mers from the sequence data of a sample that contain / overlap with a variant cannot have an exact match in the hash table initialized with the variant-free set 34 of reference k-mers, and are therefore considered "lost". However, since variants are very few relative to the overall size of the virus, this simply results in minimal loss of sensitivity. In some methods, known variants present in a population can also be included as one or more "variant k-mers" 34 added to the set 34 of k-mers in the k-mer aligner 36.

[0022] Figure 3 shows an exemplary method 100 for detecting viral pathogens in human samples. In the illustrated embodiment, sequence data 102 of the human sample is provided as data in FASTQ format that enables secondary analysis and alignment of sequencing reads, performed for example using DRAGEN or another secondary analysis tool. Alignment 104 of sequencing reads is performed using the k-mer aligner 36 (see FIG. 1) to identify exact matches of k-mers of fixed size n using the set of reference k-mers based on the genome of the viral pathogen. Alignment 104 may also include identifying exact matches of k-mers in the sequence data 102 for one or more human control amplicons (e.g., 2 to 15 amplicons) used as a measure of sample quality. In some embodiments, alignment 104 can be a conventional DRAGEN alignment to a reference genome comprising a virus, e.g., SARS-CoV2, and one or more human control amplicons.

[0023] Human reads 110 and viral reads 112 are subjected to additional metrics as provided herein to evaluate sample quality based on human amplicon coverage 114, thereby generating control detection output 120. The metrics also include viral amplicon coverage metrics 130 for providing viral detection output 132. Positive samples based on both the viral detection output and the control detection output 120 can be sent to variant calling 124 to generate viral sequence output 128.

[0024] After sequence read alignment / matching using a k-mer aligner 36 is performed, metrics associated with the designated virus are interpreted, and as shown in Figure 4, a determination is made regarding the detection of the virus and the internal (human) control. In some methods, the number of unique reads 160 mapping to the target region (or detected k-mer) of each amplicon can be counted.

[0025] As shown in Figure 5, the “target region” may, in embodiments, be defined as the sequence 184 of an amplicon 184, excluding primers and any overlap with other amplicons 184. This can be done either by a) aligning reads to the viral genome 180 and counting the number of reads 188 that map to the location of each amplicon (with any possible overlaps removed), or by b) counting the number of k-mers 190 from the sequence 184 of each amplicon observed in the reads. The number of k-mers or reads is compared to a threshold per amplicon coverage, and each amplicon 184 is referred to as “covered” or “not covered.” If more viral amplicons 184 are covered than a second set threshold, the call or virus detection output is that the virus is detected. The total number of amplicons depends on the assay used. In the example in Figure 5, the amplicons 184 are non-overlapping. However, it should be understood that more overlapping amplicons 184 can be used to achieve coverage of the entire viral genome.

[0026] Returning to Figure 4, after alignment and / or k-mer identification for human amplicon 162 and viral amplicon 164, the coverage 170 for each individual human amplicon and the coverage 172 for each viral amplicon are counted. The number of reads per amplicon (or the number of k-mers detected) is compared to a target threshold to determine the covered amplicons. The number of covered amplicons is then used to detect virus 178 (by positive detection results based on covered amplicons above the virus threshold) and internal (human) control 174 (by positive control detection results based on covered amplicons above the human control threshold). The thresholds for detecting positive amplicons, and the thresholds for the number of amplicons required to detect controls and / or viruses, may vary. In some embodiments, the detection threshold may be as little as two amplicons or as much as two, for example, three, four, or more amplicons. In embodiments, the threshold number of covered amplicons may be at least 1%, at least 10%, or at least 50% of the total number of amplicons. In embodiments, the threshold number of covered amplicons may range from 1% to 5% of the total number of amplicons in the assay. Since detection is designed to provide fast results for real-time sequencing data as additional sequencing data is generated from the sample, setting a percentage threshold allows detection to be performed based on any combination of positive amplicons. Therefore, detection is independent of sample variability in the location of sequenced clusters or other detection-specific variables that differ by sample.

[0027] Figure 4 shows an exemplary virus detection performed on human controls. For human control amplicons 170, control 1 had 25 unique reads, and control 3 had 64 unique reads exceeding the target threshold, and these amplicons were determined to be covered amplicons. In the next step, the two positive amplicons for the human controls were compared to a human control threshold set to 2 or higher, resulting in a control detection threshold exceedance determination 174. Thus, human control detection 174 involved a two-step analysis that determined the coverage of individual human amplicons based on the amplicon coverage threshold, and then evaluated the number of amplicons that exceeded the coverage threshold. Similarly, virus detection 178 involved a first step in which the number of unique reads for each virus amplicon (e.g., virus 1, virus 2, virus 3, etc.) was counted. Virus 1 had 34 unique reads, Virus 2 had 21 unique reads, and Virus Amplicon 3 had 64 unique reads; all were considered covered amplicons. However, Virus Amplicon 98, which had only one unique read, was not considered a covered amplicon. In the next step, the three covered amplicons were compared to a virus threshold set to 3 or higher to obtain virus detection results.

[0028] The disclosed method includes quality and control parameters for establishing a set of reference k-mers and / or control k-mers to be used in a k-mer aligner (e.g., k-mer aligner 36) for k-mer-based alignment. Figure 6 is a schematic diagram of method 200 for generating a set of pathogen-specific k-mers and control k-mers for pathogen detection. A given pathogen genome contains a set of all potential k-mers of fixed size n, where n can be more than 24 bases. However, certain k-mers may have a perfect match within a control genome (e.g., the human genome). In block 204, potential pathogen k-mers may be run against a control genome, and in block 206, specific k-mers may be removed to generate a final set of pathogen k-mers in block 208. In one example, k-mers in a set of potentials that have a perfect match against a control genome are removed. In another example, k-mers exceeding a similarity threshold with the control genome are removed. For example, k-mers exceeding the similarity threshold may include k-mers that have 1 to 3 bases (continuous or discontinuous) that differ from the control genome. For example, in the case of a fixed-size 32 k-mer, potential k-mers with 31 / 32 or 30 / 32 sequence matches with the control genome, which would correspond to potential base call errors that could lead to false positive detection, are removed. Thus, the k-mers retained in the final set in block 208 may include k-mers that do not have a perfect match with the control genome and / or k-mers that have sufficient dissimilarity with the control genome (e.g., 1 to 3 bases differ within the k-mer).

[0029] A set of control k-mers can be selected from a pool of potential k-mers based on a metric in block 210. In an assay that sequences RNA in a human sample to detect the presence of an RNA virus, the human sample will also contain human RNA, e.g., mRNA. Therefore, a set of human control k-mers may be based on mRNA sequences that are likely to be consistently expressed in the tissue of the sample. The set of control k-mers may be selected to be smaller than the reference set, for example, containing fewer amplicons. In block 214, the potential set of control k-mers is run against each other, in embodiments against the reference genome, and control k-mers that are a perfect match or too similar to each other and to the control genome (e.g., differing by 1-3 bases but otherwise a perfect match) are removed in step 216 to produce the final set of control k-mers in block 218. In block 220, the final set of pathogen k-mers and the final set of control k-mers are provided to the k-mer aligner.

[0030] Figure 7 is a schematic diagram of a sequencing device 260 that may be used in conjunction with the disclosed embodiments for obtaining sequence data from a sample, as provided herein. The sequencing device 260 can perform a sequencing run on the sample to obtain sequence data. The sequencing device 260 may be implemented according to any sequencing method, such as those incorporating the sequencing-by-synthesis method described in U.S. Patent Publications 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2005 / 0100900, U.S. Patent No. 7,057,026, and International Publications 05 / 065814, 06 / 064199, and 07 / 010251, the entire disclosure of which is incorporated herein by reference. Alternatively, sequencing by a ligation method may be used in the sequencing device 260. Such a method involves incorporating oligonucleotides using a DNA ligase and identifying the incorporation of such oligonucleotides, and is described in U.S. Patents 6,969,488, 6,172,218, and 6,306,597, the entire disclosures of which are incorporated herein by reference. In some embodiments, nanopore sequencing can be utilized, in which nucleic acid strands or nucleotides of a sample are removed from the nucleic acids of the sample by an exonuclease and pass through nanopores. As nucleic acids or nucleotides of a sample pass through nanopores, each base species can be identified by measuring the variation in the electrical conductance of the pores (the entire disclosure is incorporated herein by reference: U.S. Patent No. 7,001,792, Soni & Meller, Clin. Chem. 53, 1996-2001 (2007), Healy, Nanomed. 2, 459-481 (2007), Cockroft et al. J. Am. Chem. Soc. 130, 818-820 (2008)). Further embodiments include the detection of protons released upon incorporation of nucleotides into the extension product.For example, sequencing based on the detection of released protons may utilize electrodetectors and related technologies commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies), or sequencing methods and systems described in U.S. Patent Publications 2009 / 0026082(A1), 2009 / 0127589(A1), 2010 / 0137143(A1), or 2010 / 0282617(A1), the entirety of which are incorporated herein by reference. Certain embodiments may utilize methods that include real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interaction between fluorophore-supported polymerase and γ-phosphate-labeled nucleotides, or using zero-mode waveguides as described in, for example, Levene et al. Science 299, 682-686 (2003), Lundquist et al. Opt. Lett. 33, 1026-1028 (2008), and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the entirety of which is incorporated herein by reference. Other suitable alternatives include, for example, fluorescence in situ sequencing (FISSEQ) and massively parallel signature sequencing (MPSS). In certain embodiments, the sequencing device 260 may be an iSeq manufactured by Illumina (La Jolla, CA). In other embodiments, the sequencing device 260 may be configured to operate using a CMOS sensor with nanowells fabricated on photodiodes so that DNA deposition is aligned one-to-one with each photodiode.

[0031] In the described embodiment, the sequencing device 260 comprises a separate sample substrate 262, e.g., a flow cell or sequencing cartridge, and an associated computer 264. However, as described above, these may be implemented as a single device. In the described embodiment, a biological sample may be loaded onto the substrate 262, which is imaged to generate sequence data. For example, a reagent interacting with the biological sample may fluoresce at a specific wavelength in response to an excitation beam generated by the imaging module 272, thereby returning radiation for imaging. For example, the fluorescent component may be generated by a fluorescently tagged nucleic acid, which hybridizes to a complementary molecule of the component or to a fluorescently tagged nucleotide incorporated into an oligonucleotide using polymerase. As will be understood by those skilled in the art, the wavelength at which the dyes in the sample are excited, and the wavelength at which they fluoresce, will depend on the absorption and emission spectra of the particular dye. Such returned radiation may propagate through a directional optical system. This retrobeam may be directed roughly to a detection optical system of the imaging module 272, which may be a camera or other optical detector.

[0032] The imaging module detection optics may be based on any preferred technique, for example, a charged coupled device (CCD) sensor that generates pixelated image data based on photons affecting location within the device. However, it will be understood that any of a variety of other detectors may be used, including but not limited to detector arrays configured for time delay integration (TDI) operation, complementary metal oxide semiconductor (CMOS) detectors, avalanche photodiode (APD) detectors, Geiger-mode photon counters, or any other suitable detectors. TDI mode detection can be coupled with line scanning, as described in U.S. Patent No. 7,329,860, incorporated herein by reference. Other useful detectors are described, for example, in the references previously provided herein in the context of various nucleic acid sequencing methodologies.

[0033] The imaging module 272 may be under processor control, for example, by a processor 274, and may include an I / O control unit 276, an internal bus 278, a non-volatile memory 280, RAM 282, and any other memory configuration in which the memory can store executable instructions, as well as other suitable hardware components, which may be similar to those described with respect to Figure 7. Furthermore, the associated computer 264 may also include a memory architecture including a processor 184, an I / O control unit 286, a communication module 294, and RAM 288 and non-volatile memory 290, so that the memory architecture can store executable instructions 292. The hardware components may be connected by an internal bus 294, which can also be connected to a display 296. In embodiments in which the sequencing device 260 is implemented as an all-in-one device, certain redundant hardware elements can be eliminated.

[0034] A processor (e.g., processors 274, 284) may be programmed to assign individual sequencing reads to a sample based on the relevant index sequence or sequence by the method provided herein. In certain embodiments, based on image data acquired by the imaging module 272, the sequencing device 260 may be configured to generate sequencing data including sequence reads for individual clusters, each sequence read associated with a specific location on the substrate 270. Each sequence read may originate from a fragment containing an insert. The sequencing data includes base calls for each base of the sequencing read. Furthermore, for sequentially performed sequencing reads based on image data, individual reads may also be linked to the same location, and therefore the same template strand, via the image data. In this way, index sequencing reads may be associated with sequencing reads of insert sequences before being assigned to the original sample. The processor 274 may also be programmed to perform downstream analysis on the sequences of a particular sample following the assignment of sequencing reads to the sample.

[0035] ). In certain embodiments, executable instruction 292 causes the processor to perform one of more actions of the method disclosed herein. The processor (e.g., processors 274, 284) may be a highly reconfigurable field-programmable gate array (FPGA) technology. The processor (e.g., processors 274, 284) may be programmed to receive user input for a particular analytical workflow in order to access a hash table containing an appropriate set of reference k-mers and / or reference k-mers stored in memory (e.g., memory 280, 290). In one example, device 260 receives user input to select a run or panel of interest, and the k-mer aligner uses the hash table associated with the user input to align the streaming sequence to identify a perfect match of k-mers in the sequence data. The memory may store several different sets of k-mers or different initialization hash tables that are specifically selected based on the user input. In embodiments, the selection may also include the selection of reference k-mers. For example, control k-mers may include control k-mers from humans, mammals, or other host organisms.

[0036] The disclosed methods may be used to characterize samples, for example, biological samples. Samples may originate from any in vivo or in vitro source, including one or more cells, tissues, organs, or organisms, whether living or dead, or any biological or environmental source (e.g., water, air, soil). For example, in some embodiments, the nucleic acids of the sample include or consist of dsDNA of eukaryotes and / or prokaryotes originating from or derived from humans, animals, plants, fungi (e.g., mold or yeast), bacteria, viruses, viroids, mycoplasmas, or other microorganisms. In some embodiments, the nucleic acid of the sample includes or comprises genomic DNA, subgenomic DNA, chromosomal DNA (e.g., derived from an isolated chromosome or a portion of a chromosome, e.g., from one or more genes or loci derived from a chromosome), mitochondrial DNA, chloroplast DNA, plasmid or other episome-derived DNA (or recombinant DNA contained therein), or double-stranded cDNA prepared by reverse transcribing RNA using RNA-dependent DNA polymerase or reverse transcriptase to produce a first-strand cDNA, and then extending primers annealed to the first-strand cDNA to produce dsDNA. In some embodiments, the nucleic acid of the sample includes a plurality of dsDNA molecules in or prepared from nucleic acid molecules (e.g., a plurality of dsDNA molecules in or prepared from cDNA prepared from or derived from biological (e.g., cells, tissues, organs, organisms) or environmental (e.g., water, air, soil, saliva, sputum, urine, feces) sources of genomic DNA or RNA). In some embodiments, the nucleic acid of the sample is derived from an in vitro source. For example, in some embodiments, the nucleic acid of the sample includes or consists of single-stranded DNA (ssDNA) or single-stranded or double-stranded RNA (dsDNA prepared in vitro using methods well known in the art, such as primer extension using a suitable DNA-dependent and / or RNA-dependent DNA polymerase (reverse transcriptase)).In some embodiments, the nucleic acid of the sample includes or consists of dsDNA prepared from all or part of one or more double-stranded or single-stranded DNA or RNA molecules using any method known in the art, including methods for: amplification of DNA or RNA (e.g., PCR or reverse transcriptase PCR (RT-PCR), transcription-mediated amplification, involving amplification of all or part of one or more nucleic acid molecules); molecular cloning of all or part of one or more nucleic acid molecules in a plasmid, fosmid, BAC, or other vector to be replicated in a suitable host cell; or capture of one or more nucleic acid molecules by hybridization, such as hybridization to a DNA probe on an array or microarray.

[0037] The advantages of the disclosed method include the suppression of noise (e.g., cross-contamination) that appears as uniformly scattered reads throughout the viral genome, in contrast to the actual signals clustered by the amplicons. The method is adaptable to different amplicons with varying PCR performance by setting a variable threshold per amplicon (higher for strongly amplified amplicons). The disclosed method closely corresponds to existing qPCR tests that also report the number of positive amplicons and therefore provides output results readily convertible for clinical use. Detection output may be reported per sample, or the detection output may undergo downstream quality control.

[0038] In some embodiments, variant calling data for any positive sample may also be reported. In some embodiments, a positive sample may be identified, and this method includes providing notification or recommendations for treatment based on the diagnosis of the positive sample. In embodiments, treatment for the detected pathogen is administered to the patient from whom the sample was taken, based on a diagnosis of pathogen detection or non-detection by the disclosed method, which is used as a point-of-care detection system. For example, if the detected pathogen is based on the detection of the SARS-CoV-2 genome, treatment for SARS-CoV-2 is administered or a surveillance protocol is initiated. If the SARS-CoV-2 genome is not detected, a SARS-CoV-2 vaccine may be administered based on a diagnosis that the patient does not have an active infection.

[0039] This written description uses examples, including best-case examples, of the embodiments of the disclosure and enables any person skilled in the art to practice the disclosed embodiments, including by fabricating and using any device or system and by performing any incorporated methods. The patentable scope of the disclosure is defined by the claims and may include other examples that a person skilled in the art may conceive. Such other examples are intended to be within the claims if they include structural elements that are no different from the literal words of the claims, or if they include equivalent structural elements that differ only slightly from the literal words of the claims.

Claims

1. A method for detecting pathogens in a biological sample, The sequencing data originates from a sequencing device and is generated from a biological sample. Identifying k-mers in sequencing data that have a perfect match in a hash table initialized with a first set of k-mers containing pathogen k-mers in the genome of the pathogen, and a second set of k-mers containing control k-mers derived from a control genome, wherein the control genome is a human genome, the first set of k-mers is a subset of all k-mers in the genome of the pathogen, and the subset is based on sufficient dissimilarity to the control genome. A positive detection output for the biological sample is provided based at least partially on a first number of perfect matches of the k-mer in the sequence data with the first set exceeding a first set threshold, and a second number of perfect matches of the k-mer in the sequence data with the second set exceeding a second set threshold. A method comprising terminating the identification of the k-mer of the biological sample in response to the provision of the positive detection output, wherein the positive detection output is provided while sequence data derived from the biological sample is being generated.

2. The method according to claim 1, wherein the k-mer in the sequence data, the first set, and the second set is of a fixed size greater than 24 nucleotides.

3. The method according to claim 1, wherein the first set of k-mers includes a variant in the genome of the pathogen.

4. The method according to claim 1, wherein the first set is larger than the second set.

5. The method according to claim 1, wherein the first set comprises k-mer derived from a plurality of different pathogens, and the positive detection output for the biological sample comprises pathogen detection of one of the plurality of different pathogens.

6. The method according to claim 1, comprising aligning the sequence data with the genome of the pathogen based on the positive result for pathogen detection.

7. The method according to claim 6, comprising identifying sequence variants of the pathogen in aligned sequence data.

Citation Information

Patent Citations

  • Pathogen operation set detecting method, device, computer equipment and storage medium

    CN109949866A

  • Microorganism detection method and device based on targeted amplification and sequencing

    CN110875082A

  • Characterization of biological material in a sample or isolate using unassembled sequence information, probabilistic methods and trait-specific database catalogs

    US20140288844A1

  • Using k-mers for rapid quality control of sequencing data without alignment

    US20190172553A1