Systems and methods for secondary analysis of nucleotide sequencing data

By using a sequencing-by-synthesis system and iterative alignment technology, secondary analysis of DNA sequencing is performed in real time, solving the problem of time-consuming variant identification in existing technologies and achieving rapid and accurate detection of genetic differences and variants.

CN115810396BActive Publication Date: 2026-07-24ILLUMINA INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ILLUMINA INC
Filing Date
2017-10-06
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing DNA sequencing technologies are time-consuming in the mutation identification process and cannot achieve rapid, real-time mutation detection.

Method used

A sequencing-by-synthesis system is used, combined with iterative alignment and variant identification technologies, to perform secondary analysis in real time. Nucleotide subsequences are processed through the first and second alignment paths respectively, improving computational efficiency and enabling real-time detection of genetic differences and variations.

Benefits of technology

This technology enables real-time detection of genetic differences and variations during sequencing, reducing data bandwidth requirements, shortening specific response times, and improving the efficiency and accuracy of variation identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115810396B_ABST
    Figure CN115810396B_ABST
Patent Text Reader

Abstract

Disclosed herein are systems and methods for performing secondary analysis of nucleotide sequencing data in a time-efficient manner. Some embodiments include iteratively performing secondary analysis as sequence reads are generated by a sequencing system. Secondary analysis can include alignment of sequence reads to a reference sequence (e.g., a human reference genome sequence) and use of the alignment to detect differences between a sample and the reference. Secondary analysis can be capable of detecting genetic differences, variant calling and genotyping, identifying single nucleotide polymorphisms (SNPs), small insertions and deletions (indels), and structural changes in DNA, such as copy number variations (CNVs) and chromosomal rearrangements.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application based on Chinese Patent Application No. 201780040788.0.

[0002] Related applications

[0003] This application claims priority to U.S. Provisional Application No. 62 / 405824, filed October 7, 2016, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0004] This disclosure generally relates to the field of DNA sequencing, and more specifically to systems and methods for real-time secondary analysis for next-generation sequencing applications. Background Technology

[0005] Genetic mutations can be identified by identifying variations in a sequence read—relative to a reference sequence. To identify variations, a sample from a subject is fully sequenced using sequencing instruments to obtain sequence reads. After obtaining the sequence reads, they can be assembled or aligned before variant identification. Therefore, variant identification involves different steps performed sequentially and can be time-consuming, especially after the sequencing process is complete. Summary of the Invention

[0006] This document discloses systems and methods for sequencing polynucleotides. In one embodiment, the system includes: a memory including a reference nucleotide sequence; and a processor configured to execute instructions comprising: receiving a first nucleotide subsequence of a read from a sequencing system; processing the first nucleotide subsequence using a first alignment path to determine a first plurality of candidate positions of the read on the reference sequence; determining whether the first nucleotide subsequence aligns to the reference sequence based on the determined candidate positions; receiving a second nucleotide subsequence from the sequencing system; and processing the second nucleotide subsequence using a path that aligns to the reference sequence to determine a second plurality of candidate positions of the read: if the read aligns to the reference sequence, then using the second alignment path, and if not, then using the first alignment path, wherein the second alignment path is computationally more efficient than the first alignment path in determining the second plurality of candidate positions of the read.

[0007] In one embodiment, the method includes: receiving a first nucleotide subsequence from a sequencing system during a sequencing run; and performing secondary analysis of the first nucleotide subsequence of the read based on a reference sequence using a first analysis path or a second analysis path, wherein the second analysis path is computationally more efficient than the first processing path in performing secondary analysis.

[0008] As a non-limiting example, this application provides the following implementation scheme:

[0009] Implementation Plan 1. A system for sequencing polynucleotides:

[0010] A memory, which includes a reference nucleotide sequence;

[0011] A processor configured to execute instructions that include the following methods:

[0012] Receive the first nucleotide sequence of the read from the sequencing system;

[0013] The first nucleotide subsequence is processed using a first alignment path to determine a first plurality of candidate positions of the reads on the reference sequence;

[0014] Based on the identified candidate positions, determine whether the first nucleotide subsequence is aligned to the reference sequence;

[0015] Receive the second nucleotide sequence from the sequencing system;

[0016] The second nucleotide subsequence is processed using the following path to determine a second plurality of candidate positions for alignment of the read to the reference sequence:

[0017] If the read segment aligns to the reference sequence, then the second alignment path is used, and

[0018] If not, then use the first comparison path.

[0019] The second comparison path determines the second plurality of candidate positions of the read segment with higher computational efficiency than the first comparison path.

[0020] Implementation Scheme 2. The system as described in Implementation Scheme 1, wherein the second nucleotide subsequence is processed using the first alignment path or the second alignment path based on an alignment quality metric.

[0021] Implementation Scheme 3. The system as described in Implementation Scheme 1, wherein the length of the first nucleotide subsequence is one or more nucleotides.

[0022] Implementation Scheme 4. The system as described in Implementation Scheme 1, wherein the length of the second nucleotide subsequence is one or more nucleotides.

[0023] Implementation Scheme 5. The system as described in Implementation Scheme 1, wherein the second comparison path is more computationally efficient than the first comparison path in terms of memory usage or operand computation.

[0024] Implementation Scheme 6. The system of Implementation Scheme 1, wherein if the first nucleotide subsequence aligns to the reference sequence, the processor is further configured to store data corresponding to at least one of the first plurality of candidate positions.

[0025] Implementation Scheme 7. The system of Implementation Scheme 6, wherein if the read segment remains aligned to the reference sequence, the processor is further configured to store data corresponding to at least one of the second plurality of candidate positions.

[0026] Implementation Scheme 8. The system as described in Implementation Scheme 1, wherein processing the second nucleotide subsequence using the second alignment path includes performing a simple alignment to determine a simple alignment score.

[0027] Implementation Scheme 9. The system as described in Implementation Scheme 8, wherein performing the simple comparison includes: comparing the second nucleotide subsequence with the corresponding sequence of the second nucleotide subsequence on the reference sequence based on the first plurality of candidate positions.

[0028] Implementation Scheme 10. The system of Implementation Scheme 8, wherein processing the second nucleotide subsequence using the second processing path further includes: determining a mapping quality (MapQ) score for each of the second plurality of candidate positions of the read.

[0029] Implementation Scheme 11. The system as described in Implementation Scheme 10, wherein the simple comparison scoring includes the MapQ scoring.

[0030] Implementation Scheme 12. The system of Implementation Scheme 1, wherein the processor is further configured to perform variant identification on the output of the first or second alignment path, which includes at least one of the first plurality of candidate positions or at least one of the second plurality of candidate positions.

[0031] Implementation Scheme 13. The system as described in Implementation Scheme 12, wherein mutation identification of the output of the first or second comparison path includes:

[0032] The output of the first or second alignment path is used to identify mutations using either a first mutation identification path or a second mutation identification path, wherein the second mutation identification path is computationally more efficient than the first mutation identification path in identifying mutations in the second subsequence.

[0033] Implementation Scheme 14. The system as described in Implementation Scheme 12, wherein mutation identification is performed using the output of the first or second alignment path based on a mutation identification metric.

[0034] Implementation Scheme 15. The system as described in Implementation Scheme 14, wherein the variation identification metric is determined based on the number of different base types identified at the position of the reference sequence.

[0035] Implementation Scheme 16. The system as described in Implementation Scheme 1, wherein the processing of the first nucleotide subsequence is completed before the sequencing system determines the second nucleotide subsequence during sequencing operation.

[0036] Implementation Scheme 17. The system as described in Implementation Scheme 1, wherein the sequencing system implements a sequencing-by-synthesis method to determine a first subsequence.

[0037] Implementation Scheme 18. A method for sequencing polynucleotides, comprising:

[0038] During the sequencing run, the first nucleotide sequence of the read is received from the sequencing system; and

[0039] Using either a first or a second analysis path, a second analysis of the first nucleotide subsequence of the read is performed based on a reference sequence, wherein the second analysis path is computationally more efficient than the first processing path in performing the second analysis.

[0040] Implementation Scheme 19. The method of Implementation Scheme 18, wherein performing the secondary analysis includes processing the first nucleotide subsequence to determine a first plurality of candidate positions of the read aligned to the reference sequence:

[0041] If the read segment does not align to the reference sequence (from a previous iteration), then the first alignment path is used.

[0042] If not, then use the second comparison path.

[0043] The second comparison path determines the first plurality of candidate positions of the read segment with higher computational efficiency than the first comparison path.

[0044] Implementation Scheme 20. The method of Implementation Scheme 19, wherein processing the second nucleotide subsequence using the second alignment path includes performing a simple alignment to determine an alignment score.

[0045] Implementation Scheme 21. The method as described in Implementation Scheme 19, wherein the result of the secondary analysis includes the output of the first alignment path, the output of the second alignment path, or any combination thereof.

[0046] Implementation Scheme 22. The method of Implementation Scheme 18, wherein performing the secondary analysis includes identifying variations in the first nucleotide subsequence, comprising:

[0047] The output of the first or second alignment path is used to identify mutations using either a first mutation identification path or a second mutation identification path, wherein the second mutation identification path is computationally more efficient than the first mutation identification path in identifying mutations in the first subsequence.

[0048] Implementation Scheme 23. The method as described in Implementation Scheme 22, wherein the result of the secondary analysis includes the output of the first variant identification path, the output of the second variant identification path, or any combination thereof.

[0049] Implementation Scheme 24. The method as described in Implementation Scheme 18, further comprising providing the results of the secondary analysis to the user during the sequencing run.

[0050] Implementation Scheme 25. The method as described in Implementation Scheme 24, wherein the results of the secondary analysis are provided to the user at fixed intervals.

[0051] Implementation Scheme 26. The method as described in Implementation Scheme 24, wherein the results of the secondary analysis are provided to the user as requested by the user.

[0052] Implementation Scheme 27. The method of Implementation Scheme 18, wherein performing the secondary analysis includes performing a secondary analysis of the first nucleotide subsequence of the read based on the results of a previous sequencing interval from the sequencing run. Attached Figure Description

[0053] Figure 1 This is a schematic diagram illustrating an example sequencing system used for real-time analysis.

[0054] Figure 2 A functional block diagram of an example computer system for performing real-time analysis is shown.

[0055] Figure 3 This is a flowchart of an example method for sequencing while synthesizing.

[0056] Figure 4 This is a flowchart of an example method for base identification.

[0057] Figure 5A and Figure 5B Example iterative alignment and variant identification are shown.

[0058] Figure 6 This is a flowchart of an example method for performing real-time secondary sequence analysis.

[0059] Figure 7A and Figure 7B It is a traditional method of two-dimensional analysis ( Figure 7A Iterative methods for secondary analysis () Figure 7B A diagram for comparison.

[0060] Figure 8 This is a schematic diagram of a read segment generated by 16 base intervals.

[0061] Figure 9AThis is a flowchart of an example method for performing real-time secondary analysis. Figure 9B This is a line graph showing the predictions for data processed according to K-Mer. Figure 9C It is a bar chart showing the runtime.

[0062] Figure 10 This is another flowchart of an example method for performing real-time secondary analysis.

[0063] Figure 11A and Figure 11B The existing variant caller ( Figure 11A ) and variant identifiers using high-confidence, low-processing paths as described in this article ( Figure 11B (Compare) Detailed Implementation

[0064] In the following detailed description, reference is made to the accompanying drawings, which form a part of the description. In the drawings, similar symbols generally identify similar components unless the context otherwise indicates. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter set forth herein. It will be readily understood that aspects of this disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are explicitly contemplated herein.

[0065] This document discloses systems and methods for performing secondary analysis of nucleotide sequencing data in a time-sensitive manner. In some embodiments, the method includes iteratively performing secondary analysis while sequence reads are generated by the sequencing system. Secondary analysis may include aligning sequence reads with a reference sequence (e.g., a human reference genome sequence) and using that alignment to detect differences between the sample and the reference. Secondary analysis can enable the detection of genetic differences, variant detection, and genotyping, identifying single nucleotide polymorphisms (SNPs), small insertions and deletions (indels), and structural changes in DNA, such as copy number variations (CNVs) and chromosomal rearrangements.

[0066] By performing secondary analysis simultaneously with the generation of sequence reads, the system and method can iteratively determine preliminary variant identification in real time (or with zero or low latency). The final results of variant identification can be obtained shortly after (or immediately after) the sequencing run ends. Alternatively, the sequencing run can be terminated early if variant identification is obtained with sufficient confidence during the run. In some embodiments, only information relevant to variant identification (e.g., variant recognition) is transmitted from the sequencing system. This can reduce or minimize the required data bandwidth compared to variant identification performed on an external system. Additionally, variant information can be sent only to a computing system (e.g., a cloud computing system) for further processing. In this embodiment, the sequencing run can be terminated before the entire sequencing process is completed. For example, if the identity of the target pathogen is determined after multiple sequencing cycles of the sequencing run, the sequencing run can be terminated. Therefore, the time for specific responses (e.g., pathogen identification) can be reduced. In one embodiment, the system's output and intermediate results can include histograms of the following: duplicates, exact matches, single and double SNPs, and single and double insertions / deletions.

[0067] definition

[0068] Unless otherwise defined, the technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology, 2nd Edition, J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). For the purposes of this disclosure, the following terms are defined as follows.

[0069] Sequencing instruments used for real-time secondary analysis

[0070] This document discloses systems and methods for iteratively performing secondary analyses in a time- and / or computationally resource-efficient manner. Secondary analyses can include aligning sequence reads with a reference sequence (e.g., a human reference genome sequence) and using that alignment to detect differences between the sample and the reference. Secondary analyses can be used to detect genetic differences, variant detection, and genotyping, identifying single nucleotide polymorphisms (SNPs), small insertions and deletions (indels), and structural changes in DNA, such as copy number variations (CNVs) and chromosomal rearrangements. Secondary analyses can be performed over a sequencing cycle while simultaneously generating sequencing data for the next sequencing cycle.

[0071] Figure 1 This is a schematic diagram illustrating an example sequencing system 100 for performing real-time secondary analysis. Non-limiting examples of sequencing methods utilized by the sequencing system 100 may include sequencing-by-synthesis and Heliscope single-molecule sequencing. The sequencing system 100 may include an optical system 102 configured to generate raw sequencing data using sequencing reagents supplied by a jet system 104, which is part of the sequencing system 100. The raw sequencing data may include fluorescence images captured by the optical system 102. A computer system 106, which is part of the sequencing system 100, may be configured to control the optical system 102 and the jet system 104 via communication channels 108a and 108b. For example, a computer interface 110 of the optical system 102 may be configured to communicate with the computer system 106 via communication channel 108a.

[0072] During the sequencing reaction, the jet system 104 can guide a reagent stream through one or more reagent tubes 112 into and out of the flow cell 114 located on the mounting stage 116. The reagents can be, for example, fluorescently labeled nucleotides, buffers, enzymes, and lysis reagents. The flow cell 114 may include at least one fluid channel. The flow cell 114 may be a patterned array flow cell or a random array flow cell. The flow cell 114 may include multiple single-stranded polynucleotide clusters for sequencing in at least one fluid channel. The length of the polynucleotides may vary, for example, from 200 bases to 1000 bases. The polynucleotides may be attached to one or more fluid channels of the flow cell 114. In some embodiments, the flow cell 114 may include multiple beads, each bead comprising multiple copies of the polynucleotide to be sequenced. The mounting stage 116 may be configured to allow the flow cell 114 to be properly aligned and moved relative to other components of the optical system 102. In one embodiment, the mounting stage 116 may be used to align the flow cell 114 with a lens 118.

[0073] The optical system 102 may include multiple lasers 120 configured to generate light of a predetermined wavelength. The light generated by the lasers 120 can pass through fiber optic cable 122 to excite a fluorescent tag in flow cell 114. A lens 118 mounted on a focuser 124 can be moved along the z-axis. The focused fluorescence emission can be detected by a detector 126, such as a charge-coupled device (CCD) sensor or a complementary metal-oxide-semiconductor (CMOS) sensor.

[0074] The filter assembly 128 of the optical system 102 can be configured to filter the fluorescence emission of a fluorescent tag in the flow cell 114. The filter assembly 128 may include a first filter and a second filter. Each filter can be a long-pass filter, a short-pass filter, or a band-pass filter, depending on the type of fluorescent molecules used in the system. The first filter can be configured to detect the fluorescence emission of a first fluorescent tag by the detector 126. The second filter can be configured to detect the fluorescence emission of a second fluorescent tag by the detector 126. Using the two filters in the filter assembly 128, the detector 126 can detect two different wavelengths of fluorescence emission.

[0075] In some embodiments, the optical system 102 may include a dichroic separator configured to separate fluorescence emission. The optical system 102 may include two detectors: a first detector coupled to a first filter for detecting fluorescence emission at a first wavelength, and a second detector coupled to a second filter for detecting fluorescence emission at a second wavelength.

[0076] In use, a sample containing the polynucleotide to be sequenced is loaded into flow cell 114 and placed in stage 116. Computer system 106 then activates jet system 104 to begin the sequencing cycle. During the sequencing reaction, computer system 106 instructs jet system 104 to supply reagents, such as nucleotide analogs, to flow cell 114 via communication channel 108b. Computer system 106 is configured via communication channel 108a and computer interface 110 to control laser 120 of optical system 102 to generate light of a predetermined wavelength and irradiate the nucleotide analog coupled to a fluorescent tag incorporated into growth primers that hybridize with the polynucleotide to be sequenced. Computer system 106 controls detector 126 of optical system 102 to capture the emission spectrum of the nucleotide analog in a fluorescence image. Computer system 106 receives the fluorescence image from detector 126 and processes the received fluorescence image to determine the nucleotide sequence of the polynucleotide to be sequenced.

[0077] Computer System

[0078] As discussed above, the computer system 106 of the sequencing system 100 can be configured to control the optical system 102 and the jet system 104. The computer system 106 can have many configurations. Figure 2 An implementation scheme is shown. For example... Figure 2 As shown, computer system 106 may include processor 202, which is in electrical communication with memory 204, storage device 206, and communication interface 208. In one embodiment, computer system 106 includes a field-programmable gate array (FPGA), a graphics processing unit (GPU), and / or a vector central processing unit (CPU) to perform sequence alignment and generate variant identification.

[0079] Processor 202 can be configured to execute instructions that cause jet system 104 to supply reagents to flow cell 114 during sequencing reactions. Processor 202 can execute instructions to control laser 120 of optical system 102 to generate light of a predetermined wavelength. Processor 202 can execute instructions to control detector 126 of optical system 102 and to receive data from detector 126. Processor 202 can execute instructions to process data received from detector 126, such as fluorescence images, and to determine the nucleotide sequence of polynucleotides based on the data received from detector 126.

[0080] The memory 204 can be configured to store instructions for configuring the processor 202 to perform the functions of the computer system 106 when the sequencing system 100 is powered on. When the sequencing system 100 is powered off, the memory 206 can store instructions for configuring the processor 202 to perform the functions of the computer system 106. The communication interface 208 can be configured to facilitate communication between the computer system 106, the optical system 102, and the jet system 104.

[0081] Computer system 106 may include a user interface 210 configured to communicate with a display device (not shown) for displaying sequencing results (including results of secondary analyses, such as variant identification) from sequencing system 100. User interface 210 may be configured to receive input from a user of sequencing system 100. Optical system interface 212 and jet system interface 214 of computer system 106 may be configured to communicate via… Figure 1 The communication channels 108a and 108b shown control the optical system 102 and the jet system 104. For example, the optical system interface 212 can communicate with the computer interface 110 of the optical system 102 via the communication channel 108a.

[0082] Computer system 106 may include a nucleic acid base determiner 216 configured to determine the nucleotide sequence of a polynucleotide using data received from detector 126. Nucleic acid base determiner 216 can generate a template for the location of polynucleotide clusters in flow cell 114 using a fluorescence image captured by detector 126. Nucleic acid base determiner 216 can record the location of the polynucleotide clusters in flow cell 114 in the fluorescence image captured by detector 126 based on the generated location template. Nucleic acid base determiner 216 can extract the intensity of fluorescence emission from the fluorescence image to generate an extracted intensity. Nucleic acid base determiner 216 can determine the bases of the polynucleotide from the extracted intensity. Nucleic acid base determiner 216 can determine a quality score for the determined polynucleotide bases.

[0083] Computer system 106 may include an iterative aligner 218 and a variant identifier 220, such as a Strelka variant identifier (sites.google.com / site / strelkasomaticvariantcaller / home / faq). During sequencing cycles, iterative aligner 218 may align a sequence read determined by nucleic acid base determiner 216 with a reference sequence. The aligned sequence read may have an associated score. The score may be the probability that the sequence read has been correctly aligned to the reference sequence (e.g., mismatch percentage). In some embodiments, computer system 106 may include hardware, such as a field-programmable gate array (FPGA) or graphics processing unit (GPU), for aligning sequence reads with a reference sequence and for determining variant identification. In some embodiments, iterative aligner 218 and variant identifier 220 may be implemented by a different computer system than computer system 106. In some embodiments, computer system 106 may be an integrated component of sequencing system 100. In some implementations, the optical system 102, the jet system 104, and / or the computer system 106 can be integrated into a single machine.

[0084] Synthesis and sequencing

[0085] Figure 3 This is a flowchart of an example method 300 for sequencing-by-synthesis using sequencing system 100. After starting method 300 at box 305, a flow cell 114 comprising fragmented double-stranded polynucleotide fragments is received at box 310. The fragmented double-stranded polynucleotide fragments can be generated from a deoxyribonucleic acid (DNA) sample. The DNA sample can be from a variety of sources, such as biological samples, cell samples, environmental samples, or any combination thereof. The DNA sample may include one or more of the patient's biological fluids, tissues, and cells. For example, the DNA sample may be derived from or include blood, urine, cerebrospinal fluid, pleural fluid, amniotic fluid, semen, saliva, bone marrow, biopsy samples, or any combination thereof.

[0086] DNA samples may include DNA from target cells. Target cells may vary and, in some embodiments, express a malignant phenotype. In some embodiments, target cells may include tumor cells, bone marrow cells, cancer cells, stem cell endothelial cells, virus-infected pathogen cells, parasitic somatic cells, or any combination thereof.

[0087] The length of the fragmented double-stranded polynucleotide fragment can be from 200 to 1000 bases. When a flow cell 114 comprising the fragmented double-stranded polynucleotide fragment is received at box 310, method 300 proceeds to box 315, wherein the double-stranded polynucleotide fragment is bridged to the cluster of double-stranded polynucleotide fragments attached to the inner surface of one or more channels of the flow cell (e.g., flow cell 114). The inner surface of one or more channels of the flow cell may include two types of primers, such as a first primer type (P1) and a second primer type (P2), and the DNA fragment can be amplified by well-known methods.

[0088] After clusters are generated in flow cell 114, method 300 can begin the sequencing-by-synthesis process. The sequencing-by-synthesis process may include determining the nucleotide sequence of the clusters of single-stranded polynucleotide fragments. To determine the sequence of a cluster of single-stranded polynucleotide fragments having the sequence 5'-P1-F-A2R-3', a primer having the sequence A2F (which is the complementary sequence to sequence A2R) can be added, and at box 320, a nucleotide analog with 0, 1, or 2 tags can be extended by DNA polymerase to form a growth primer-polynucleotide.

[0089] During each sequencing cycle, four types of nucleotide analogs can be added and incorporated into the growing primer-polynucleotide. These four types of nucleotide analogs can have different modifications. For example, the first type of nucleotide can be a deoxyguanosine triphosphate (dGTP) analog without any fluorescent tag conjugation. The second type of nucleotide can be a deoxythymidine triphosphate (dTTP) analog conjugated to a first-type fluorescent tag via a linker. The third type of nucleotide can be a deoxycytidine triphosphate (dCTP) analog conjugated to a second-type fluorescent tag via a linker. The fourth type of nucleotide can be a deoxyadenosine triphosphate (dATP) analog conjugated to a first-type and a second-type fluorescent tag via one or more linkers. The linker can contain one or more cleavage groups. The fluorescent tag can be removed from the nucleotide analog before subsequent sequencing cycles. For example, the linker to which the fluorescent tag is attached to the nucleotide analog can contain azide and / or alkoxy groups, e.g., on the same carbon, such that the linker can be cleaved by a phosphine reagent after each incorporation cycle, thereby releasing the fluorescent tag from subsequent sequencing cycles.

[0090] Nucleotide triphosphates can be reversibly blocked at the 3' position, allowing controlled sequencing, and no more than one nucleotide analog can be added to each extended primer-polynucleotide in each cycle. For example, the 3' ribose position of the nucleotide analog can contain alkoxy and azide functional groups, which can be removed by cleavage with a phosphine reagent, yielding nucleotides that can be further extended. After incorporation of the nucleotide analog, the jet system 104 can wash one or more channels of the flow cell 114 to remove any unincorporated nucleoside analogs and enzymes. Reversible 3' blocks can be removed before subsequent sequencing cycles, allowing another nucleotide analog to be added to each extended primer-polynucleotide.

[0091] At box 325, a laser, such as laser 120, can excite two fluorescent tags at a predetermined wavelength. At box 330, signals from the fluorescent tags can be detected. Detecting the fluorescent tags can include, for example, capturing fluorescence emission at a first wavelength and a second wavelength in two fluorescence images using a detector 126 with two filters. The fluorescence emission of the first fluorescent tag may be at or around the first wavelength, and the fluorescence emission of the second fluorescent tag may be at or around the second wavelength. The fluorescence images can be stored for later offline processing. In some embodiments, the fluorescence images can be processed to determine the sequence of primer-polynucleotides grown in each cluster in real time.

[0092] In online real-time fluorescence imaging processing, a fluorescence image containing the detected fluorescence signal can be processed in box 335, and the bases of the incorporated nucleotides can be determined. For each determined nucleotide base, a quality score can be determined in box 340. A decision can be made in box 345 whether to detect more nucleotides based on, for example, the quality of the signal or after a predetermined number of bases. If more nucleotides are to be detected, nucleotide determination for the next sequencing cycle can be performed in box 320. In some embodiments, labeled nucleotides can be added to one end of the DNA strand corresponding to a cluster. Labeled nucleotides can also be added to the other end of the DNA strand corresponding to a cluster. Reads at one end of the DNA strand are typically referred to as read set 1, and reads at the other end of the DNA strand are typically referred to as read set 2. Sequencing techniques that allow the determination of two or more reads from two positions on a single polynucleotide duplex are called paired-end (PE) sequencing. Two or more reads from two positions on a single polynucleotide duplex are referred to as read set 1, read set 2, etc. Paired-end sequencing has been described in U.S. Patent Application No. 14 / 683,580, the contents of which are incorporated herein by reference in their entirety. The advantage of the paired-end method is that sequencing two segments from a single template yields more information than sequencing each of two independent templates in a random manner.

[0093] Before the next sequencing cycle, the fluorescent tag can be removed from the nucleotide analog, and the reversible 3' block can be removed, allowing another nucleotide analog to be added to each extended primer-polynucleotide. After processing all fluorescence images, method 300 can terminate at box 350.

[0094] Base recognition

[0095] Base recognition can refer to the process of determining whether the base of a nucleotide incorporated into a growth primer-polynucleotide cluster being sequenced is guanine (G), thymine (T), cytosine (C), or adenine (A). Figure 4 This is a flowchart of an example method 400 for base identification using a sequencing system 100. Figure 3 Processing the detected signal at box 335 may include performing base recognition as described in method 400. After starting at box 405, light of a predetermined wavelength can be generated using a laser. The generated light can then be directed onto the nucleotide analog at box 410. For example, computer system 106 can use its optical system interface 212 and communication channel 108a to enable laser 120 to generate light of the predetermined wavelength.

[0096] The laser-generated light can illuminate nucleotide analogs that are incorporated into nucleotide analogs in a growth primer-polynucleotide attached to the inner surface of one or more channels of a flow cell (e.g., flow cell 114). The primer-polynucleotide may comprise clusters of single-stranded polynucleotide fragments hybridized to sequencing primers. Each nucleotide analog may contain 0, 1, or 2 fluorescent tags. The two fluorescent tags may be a first fluorescent tag and a second fluorescent tag. Upon excitation by the laser-generated light, the fluorescent tags may emit fluorescence. For example, the first fluorescent tag may emit fluorescence at a first wavelength, which may be captured, for example, in a first fluorescence image. The second fluorescent tag may emit fluorescence at a second wavelength, which may be captured, for example, in a second fluorescence image.

[0097] Nucleotide analogs can include type I, type II, type III, and type IV nucleotides. Type I nucleotides, such as analogs of deoxyguanosine triphosphate (dGTP), do not conjugate to either the first or second fluorescent tag. Type II nucleotides, such as analogs of deoxythymidine triphosphate (dTTP), may conjugate to a type I fluorescent tag but not to a type II fluorescent tag. Type III nucleotides, such as analogs of deoxycytidine triphosphate (dCTP), may conjugate to a type II fluorescent tag but not to a type I fluorescent tag. Type IV nucleotides, such as analogs of deoxyadenosine triphosphate (dATP), may conjugate to both type I and type II fluorescent tags.

[0098] At box 415, at least one detector can be used to detect the fluorescence emission of the nucleotide analog at a first wavelength and a second wavelength. For example, detector 126 can capture two fluorescence images: a first fluorescence image at the first wavelength and a second fluorescence image at the second wavelength. After receiving the two fluorescence images from optical system 102, nucleic acid base determiner 216 can determine the presence or absence of fluorescence emission in the two fluorescence images.

[0099] Because the first type of nucleotide is not conjugated to either the first or second fluorescent tag, it produces little or no fluorescence emission at either the first or second wavelength. At decision box 420, if no fluorescence emission is detected, the nucleotide can be identified as a first type nucleotide, such as dGTP. If any or more than the minimum fluorescence emission is detected, method 400 can proceed to decision box 425.

[0100] Because the second type of nucleotide is conjugated to a first type of fluorescent tag but not to a second type of fluorescent tag, the second type of nucleotide can produce fluorescence emission at the first wavelength but not at or with minimal fluorescence emission at the second wavelength. At decision box 425, if no fluorescence emission at the second wavelength is detected in the second fluorescence image, and fluorescence emission at the first wavelength is detected in the first fluorescence image from decision box 420, the nucleotide can be identified as a second type of nucleotide, such as dTTP. If fluorescence emission is detected at the second wavelength, method 400 can proceed to decision box 430.

[0101] Because the third type of nucleotide is conjugated to the second type of fluorescent tag but not to the first type of fluorescent tag, the third type of nucleotide can produce fluorescence emission at the second wavelength while producing little or no fluorescence emission at the first wavelength. At decision box 430, if no fluorescence emission at the first wavelength is detected in the first fluorescence image, and fluorescence emission at the second wavelength is detected in the second fluorescence image from decision box 425, the nucleotide can be identified as the third type of nucleotide, such as dCTP.

[0102] Because the fourth type of nucleotide is conjugated to both the first and second type of fluorescent tags, it can produce fluorescence emission at either the first or second wavelength. At decision box 430, if fluorescence emission is detected at the first wavelength in the first fluorescence image, and fluorescence emission at the second wavelength can be detected in the second fluorescence image from decision box 425, the nucleotide can be identified as a fourth type of nucleotide, such as dATP.

[0103] Flow cell 114 may include clusters of growth primer-polynucleotides to be sequenced. At decision box 435, for a given sequencing cycle, if at least one more cluster with fluorescent emission remains to be processed, method 400 may continue at box 410. If no more single-stranded polynucleotide clusters remain to be processed, method 400 may terminate at box 440.

[0104] sequencing methods

[0105] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing technologies. Particularly suitable technologies are those in which nucleic acids are attached to fixed positions in the array such that their relative positions do not change and where the array is repeatedly imaged. Embodiments in which images are obtained in different color channels (e.g., consistent with different tags used to distinguish one nucleotide base type from another) are particularly suitable. In some embodiments, the process of determining the nucleotide sequence of the target nucleic acid can be automated. Preferred embodiments include sequencing-by-synthesis (“SBS”) technology.

[0106] Sequencing-by-synthesis (“SBS”) technology typically involves the enzymatic extension of nascent nucleic acid chains by iteratively adding nucleotides to a template strand. In conventional SBS methods, a single nucleotide monomer can be provided to the target nucleotide in the presence of a polymerase in each delivery. However, in the method described herein, more than one type of nucleotide monomer can be provided to the target nucleic acid in the presence of a polymerase during delivery.

[0107] Iterative alignment and mutation identification

[0108] Figure 5A and Figure 5B An example iterative alignment and variant identification process according to one implementation is illustrated. After imaging a minimum number of sequencing cycles, a preliminary analysis can be performed in real time to determine base identification and quality scoring for each unaligned read. Figure 5A The minimum sequencing cycle number shown is 3. In some implementations, the minimum sequencing cycle number can be 16, 32, or more. (See above for reference.) Figure 3 This describes base identification and quality scoring. Each read can select the most probable alignment to align to a reference sequence, and then the reads can be stacked in a pile-up for variant identification.

[0109] exist Figure 5AThe primary analysis involves identifying unaligned sequence reads from 16 clusters displayed on the flow cell, such as CCA 504a, TTA 504d, and TAG 504k. Under the primary analysis headings, each cluster is represented by a line of letters, with each letter representing a sequenced polynucleotide. When sequencing is performed with a minimum number of cycles (e.g., 3 cycles), the secondary analysis may include comparing the 16 sequence reads with… Figure 5A The sequence (GATTACATAAGATTCTTTCATCG 508) shown under the heading "Secondary Analysis" is aligned. In the Secondary Analysis diagram, the sequences aligned under the reference sequence form a polynucleotide stack. As an example, sequence reads CCA 504a (row 1 under the heading "Primary Analysis"), TTA 504d (row 4), and TAG 504k (row 11) can be aligned to sequences ACA, TTA, and TAC within the TTACAT 512 subsequence of reference sequence 508, respectively, with 1, 0, and 1 mismatch. Therefore, the third position of the TTACAT 512 subsequence can be correctly identified as C516a instead of A in reference sequence 508 with a certain probability, and the fourth position of the TTACAT 512 subsequence can be correctly identified as G 516b instead of C in the reference sequence with a certain probability. Other variations of the reference sequence can be identified similarly.

[0110] When a new sequencing cycle is performed and base recognition is determined, alignment probabilities can be improved, and read alignments can shift to new, most probable alignments. This shift will trigger new variant recognition in the affected regions. Figure 5B In the sequence, after the fourth sequencing cycle, the sequencing reads CCA 504a, TTA 504d, and TAG 504k from the third sequencing cycle become CCAT 504a' (line 1 under the heading "Level 1 Analysis"), TTAC 504d' (line 4), and TAGG 504k' (line 11), respectively. The sequence reads CCAT 504a' and TTAC 504d' can still be aligned with the TTACAT 512 subsequence of reference sequence 508, showing 1 and 0 mismatches, respectively. For the sequence reads CCAT 504a' and TTAC 504d', the alignment position is... Figure 5A The iterations shown Figure 5B The iterations shown do not change; the third position of the TTACAT 512 subsequence can be determined as C 516a instead of A in the reference sequence. Two mismatches are required for the read TAGG 504k' to align to the TTACAT 512 subsequence. However, the read TAGG 504k' can align to TAAG 520 of the reference sequence 508 with a higher probability because this alignment has only one mismatch. Figure 5A and Figure 5B Examples show that alignment positions can change as sequencing runs continue, and this can improve variant identification.

[0111] In one implementation, aligning a sequence read with a reference sequence involves maintaining a list of the most likely alignments for each read as leaves on a node. Each leaf can have an associated probability. Leaves with probabilities below a certain threshold can be pruned.

[0112] Real-time secondary analysis

[0113] Figure 6 This is a flowchart of an example method 600 for performing real-time secondary sequence analysis. After method 600 begins at box 605, imaging data from the sequencing cycle can be received at box 610. For example, computer system 106 can receive imaging data from detector 126. At box 615, bases can be identified and a quality score for the bases can be determined. (Reference) Figures 3-4 This describes the generation of imaging data and the determination of the identified bases and quality. After each sequencing cycle, the length of the sequencing read becomes one nucleotide longer. For example, after the 31st sequencing cycle, the sequencing read is 31 nucleotides long, and after the 32nd sequencing cycle, the sequencing read becomes one nucleotide longer, increasing to a length of 32 nucleotides.

[0114] At decision box 620, it can be determined whether a certain number of minimum sequencing cycles have been performed. The minimum sequencing cycles can be 16, 32, or more. If the number of sequencing cycles performed is less than the required minimum sequencing cycles, method 600 proceeds to box 610. If the number of sequencing cycles performed is at least the required minimum sequencing cycles, method 600 proceeds to box 625.

[0115] At box 625, the identified sequence read can be aligned to a reference sequence. Method 600 can use different alignment methods in different implementations. Non-limiting examples of alignment methods include global alignment (such as the Needleman–Wunsch algorithm), local alignment, dynamic programming (such as the Smith–Waterman algorithm), heuristic algorithms or probabilistic methods, incremental methods, iterative methods, motif discovery or spectral analysis, genetic algorithms, simulated annealing, pairwise alignment, and multiple sequence alignment.

[0116] At box 630, variants can be identified. Initial variants can only be identified after a predetermined variant threshold has been reached. The variant threshold can be important due to possible PCR or sequencing errors. The variant threshold can be based on a base alignment with a reference sequence position (different from the corresponding base in the reference sequence).

[0117] exist Figure 5AIn this context, the variant threshold is an observation. Therefore, the third position of TTACAT can be determined as C instead of A in the reference sequence. If the variant threshold is two or more, the C variant will not be identified at box 630 of a specific sequencing cycle. Figure 5B In this approach, if the mutation threshold is at most two observations, the third position of TTACAT can be determined as C instead of A in the reference sequence. In some implementations, the mutation threshold can be a percentage of all bases aligned to a specific position in the reference sequence, such as 1%, 5%, 10%, 25%, 50%, or greater. As described further below, for each sequence read, the most likely alignments can be stored as leaves on the node. Each leaf can have an associated probability. Leaves with probabilities below a certain threshold can be pruned. Therefore, the mutations identifying nucleotide positions on the reference sequence can be improved or reduced during subsequent cycles.

[0118] At decision box 635, it can be determined whether there are more nucleotides to be read or whether all sequencing cycles are complete. This determination can be based on, for example, the quality of the signal or after a predetermined number of bases. If there are more nucleotides to be read and not all sequencing cycles are complete, method 600 proceeds to box 610, where sequencing data can be generated for the next sequencing cycle. If there are no more nucleotides to be read and all sequencing cycles are complete, method 600 ends at box 650.

[0119] In some implementations, boxes 625 and 630, as well as boxes 610 and 615, can be performed in parallel after a minimum number of sequencing cycles have been completed. For example, after 32 sequencing cycles, the method can proceed to box 625 for alignment of a 32-nucleotide sequence read. Although method 600 performs alignment at box 625 and variant identification at box 630, the next sequencing cycle (i.e., the 33rd sequencing cycle) can be performed. Therefore, variants can be identified at box 630 before the 33rd sequencing cycle is completed. Furthermore, method 600 can perform alignment and variant identification in real-time (or with zero or low latency) during sequencing cycles. Additionally, variants identified during earlier sequencing cycles can be improved during subsequent cycles. Therefore, Figure 6 The variant identification shown can be an iterative process. For example, a variant identified after the 32nd sequencing cycle or during the 33rd sequencing cycle can be the initial variant identified. During subsequent sequencing cycles, the identified variants can be improved (including preventing previously identified variants at specific nucleotide positions from being identified and reducing their number). As another example, such as Figure 5A and Figure 5B As shown, the variant at the fourth position of TTACAT is identified as G after the third cycle, while no variant at that position is identified after the fourth position.

[0120] In another implementation, the sequencing process can be terminated before all sequencing cycles are completed. For example, if a specific target variant is identified before all sequencing cycles are completed, the sequencing process can be terminated. This makes the system more cost-effective in terms of reagents and provides the desired results earlier than a system that requires all cycles to be completed before target variant identification.

[0121] In some implementations, alignment may not be performed at box 625, and variants may be identified at box 630 in each sequencing cycle. Alignment may be performed, and variants may be identified every nth sequencing cycle, where n is 1, 2, 3, 4, 5, 10, 20, or more sequencing cycles. In some implementations, the frequency of alignment performed at box 625 and the frequency of variants identified at box 630 may be based on the number of variants identified in previous sequencing cycles. For example, if a large number of variants are identified in a sequencing cycle, alignment and variant identification may be performed more frequently (e.g., in the next cycle) or less frequently. As another example, if no variants are identified or no new variants are identified in a sequencing cycle, alignment and variant identification may be performed more frequently or less frequently (e.g., not in the next cycle).

[0122] In some implementations, variant identification at frame 630 can be selectively performed on regions of the reference sequence. The alignment portion of the reference sequence can vary in different implementations. For example, variant identification can be selectively performed on regions of the reference sequence where the alignment of the sequence read with the reference sequence has changed during a previous sequencing cycle (e.g., an immediately preceding sequencing cycle). As another example, the alignment region of the reference sequence can be determined based on known single nucleotide polymorphism (SNP) locations.

[0123] In some implementations, the method 600 for performing real-time secondary sequence analysis can be based on a tree structure for each read. The root of the tree can be marked with a "$" to indicate the start of the sequence. The child nodes of the root correspond to four possible base identifications: 'A', 'C', 'G', and 'T'. Each node in the tree can have three variables associated with it: the total number of differences in the sequence from the root to the current branch (called sequence S), the bases from the current read (called sequence W), and the start and stop indices in the Burrows-Wheeler transform (BWT) of the reference sequence for all positions in the reference that matches sequence S. An important feature of BWT is that it guarantees that all rows with a common start sequence are continuous in the transform, rather than keeping a list of individual indices in the reference that matches sequence S, which is sufficient to track the start and stop indices. This is valuable in the case of mapping reads to a human reference genome because there are many repetitive regions.

[0124] Then, each child node of the root will also have its own four child nodes, corresponding to the four possible bases 'A', 'C', 'G', and 'T'. Similarly, the number W of differences from the current read sequence can be tracked. For example, if the reads of the first two cycles are 'C' followed by 'T', the reads could have a path through the tree defined by root->C->T. Therefore, for the final T node, the total cumulative difference will be zero. Conversely, for the path defined by root->A->G, the total cumulative difference at the G node will be 2, because neither A nor G matches the corresponding cycle in the current read.

[0125] In some implementations, a limit can be defined on the number of differences from an acceptable reference. Once this limit is reached, the branch dies and is no longer analyzed in subsequent loops. A BWT transformation with an appropriate exponent can be used to perform the computations necessary for each node in a constant O(1) time. The amount of memory required for computation and the number of nodes in the tree are influenced by the total allowed error threshold. In some implementations, support for small insertions and missing values ​​can be implemented.

[0126] In some implementations, more complex rearrangements are processed through multiple seeds. That is, if a particular read is found to be mismatched anywhere, the process can be restarted in some later loop, with the expectation that another part of the read will be plotted somewhere. All these reads can be tracked, and more complex analyses (e.g., dynamic programming methods like the Smith-Waterman algorithm) can be performed when available computational power.

[0127] Alternative implementation schemes

[0128] Another implementation is a system and method for secondary analysis, which includes iterative processing of sequencing reads. Secondary analysis may include comparing the sequence reads with a reference sequence (e.g., a human reference genome sequence) and using this alignment to detect differences between the sample and the reference, such as variant detection and identification. In one implementation, alignment and variant identification results can be obtained before the sequencer completes its run. For example, these results can be provided at time intervals depending on available computing resources. This can be achieved by using alignment results from the current iteration to extend intermediate alignment results from previous iterations. Alignment results from the current iteration are generated by comparing the bases of the newly sequenced sequences in the current iteration with the bases of the reference sequence at the previously aligned positions. The comparison results are combined with the alignment results from previous iterations, and the combined output is stored for the next iteration.

[0129] Figure 7A and Figure 7B It is a traditional two-dimensional analysis method ( Figure 7ASecondary analysis of the implementation scheme disclosed herein () Figure 7B A diagram for comparison. Figure 7A This explains that in traditional two-stage analysis methods, alignment is not performed until the complete set of bases in the read is sequenced. The alignment process can include multiple alignment processing steps. The first alignment processing step waits for the complete set of sequenced bases in the read to be available. After the alignment process is complete, the variant identifier process can begin, which includes multiple variant identifier processing steps. The first variant identifier processing step waits for the complete set of alignment data to be available.

[0130] Figure 7B An iterative method for secondary analysis according to one embodiment of this disclosure is described. As shown, alignment and variant identification run in real time and generate provisional results. Processing can be arranged at fixed intervals. A fixed interval can include the arrival of a subsequence of N bases, where N is a positive integer, such as 16. For example, processing can occur at intervals of 16 bases. As another example, processing can occur at intervals of 1, 2, 4, 8, 16, 32, 64, 128, 151, or more bases. In one embodiment, processing can occur at any number of intervals between 1 and 152, most preferably at intervals of 16 + / - 8. In one embodiment, the interval can be changed from one iteration to another. Figure 8 As shown, sequencing systems, such as Figure 1 The sequencing system 100 can generate sequence reads with 16-base intervals. Alternatively, the number of bases in each processing interval can be different. For example, a first interval can be processed after sequencing 16 bases, and a second iteration can be processed after sequencing 18 bases. The number of bases in the iteration can be as low as 1 or as high as the number of bases in the read.

[0131] When using paired end sequencing technology, Figure 7B The process described can be applied to either read set 1 or read set 2. Furthermore, information captured while processing read set 1 can be applied to read set 2. For example, it is possible that an alignment step is performed using standard methods during or after sequencing read set 1, and this information can be used to process read set 2 when sequencing read set 2 polynucleotides.

[0132] Now for reference Figure 8Multiple reads (804a-804d) of single-stranded polynucleotides can be generated from the sequencing instrument. These single-stranded polynucleotides can be 151 bases in length, referred to as bases 0 to 150. The sequences of these single-stranded polynucleotides can be determined using sequencing-by-synthesis as described above. After iteration 0 (first iteration) of 16 sequencing cycles, 16-base reads are determined by the sequencing system. For example, for read 0 (804a), 15-base reads are generated, and for read 1 (804b), 15-base reads are determined, and so on. After iteration 1 (second iteration) of another 16 sequencing cycles, 16 additional bases are determined for each read. For example, 16-base 31 are generated for read 0 (804a). The sequencing system can continue generating reads at 16-interval intervals until, at iteration 8, 128-143-base reads are generated for each cluster. The sequencing system can generate reads of 144 to 151 bases per cluster at iteration 9 (the last iteration). In an alternative implementation, the number of bases generated at each iteration can vary, determined by available computational resources. For example, the first processing interval can consist of 16 bases, while the second processing interval can consist of 18 bases. The minimum number of bases in a processing interval is 1, and the maximum number of bases in a processing interval is equal to the length of the read.

[0133] refer to Figure 7B As shown, alignments can occur at 16-base intervals. Variation identification can also occur at 16-base intervals after alignment. For example, a sequencing system for real-time secondary analysis can output 16-base sequence reads every 1.3 hours. For read-time secondary analysis, the total time required for alignment and variation identification should be within 1.3 hours so that the user can access variation identification performed before the next 16 bases of the sequence read are available.

[0134] In one implementation, processing can occur continuously as quickly as possible using available computing resources, without fixed iterative steps. Analysis can be self-adjusting and occur as close as possible to sequencing progress. Alignment and variant identification results can be generated on demand.

[0135] Alternative Implementation Schemes - Comparison

[0136] Figure 9A This is a flowchart of an example method 900 for performing real-time secondary analysis. Method 900 includes two paths: a low-confidence, high-computation processing path of conventional secondary analysis methods and a high-confidence, low-computation processing path according to an embodiment of this disclosure. The low-confidence, high-processing path and the high-confidence, low-processing path are referred to herein as the blue path and the yellow path, respectively.

[0137] A low-confidence, high-computational-processing path may involve aligning each read to a reference sequence. For this path, all bases from available iterations of the read are used to align the read to the reference sequence. For example, if iteration 0 and iteration 1 each consist of 16 bases, the aligner will process 32 bases. One of many common alignment techniques can be used for low-confidence, high-computational-processing paths. Once sequence alignment is complete, plotting and alignment positions can be stored and scored. Variations can be identified after all reads have been aligned.

[0138] Method 900 improves upon conventional two-stage analysis by adding a high-confidence, low-computation processing path. At iteration 0, Method 900 waits for multiple sequencing cycles to complete to generate multiple bases for each read. For example, Method 900 may wait for 16 sequencing cycles to complete to generate 16 bases for each read. During iteration 0, the 16 bases for each read are analyzed and processed according to the low-confidence, high-computation processing path. This is referred to as the blue path in this paper. During iteration 1 and any subsequent iterations, the next 16 bases for each read are analyzed according to either the low-confidence, high-computation processing path or the high-confidence, low-computation processing path. If the read was aligned with sufficient confidence in an immediately preceding iteration, the 16 bases for the current iteration are analyzed according to the high-confidence, low-computation processing path. Otherwise, the 16 bases for the current iteration are analyzed according to the high-confidence, low-computation processing path.

[0139] If the read was aligned with sufficient confidence in an immediately preceding iteration, the 16 bases of the current iteration are aligned with the next 16 bases of the reference sequence. This alignment is referred to as simple alignment in this paper, and it requires less processing compared to common sequence alignments. Instead of aligning with the entire reference sequence, the number of mismatches between the 16 bases of the current iteration and the next 16 bases of the reference sequence can be determined. If the number of mismatches is higher than a threshold, processing of the 16 bases can be reverted to a low-confidence, high-computation processing path. When reverting to a low-confidence, high-processing path, the alignment (isAligned) variable can be set to 0 or error. The number of mismatches can be determined relative to the 16 bases of the current iteration or to all bases of the current iteration and one or more previous iterations.

[0140] If the number of mismatches is below a threshold, processing of 16 bases can remain within a high-confidence, low-computation processing path, and alignment results for specific reads can be stored. Alternative metrics can be developed to determine whether the isAligned variable is set to 0 or is incorrect. For example, if the number of mismatches is below a threshold, a MapQ score (for mapping quality) can be calculated. The MapQ score can be equal to -10log[...]. 10 Pr{plot position error}, rounded to the nearest integer. Therefore, if the plot is correct... Figure 1If the probability of a random read segment is 0.99, then the MapQ score should be 20 (i.e., logarithmic). 10 (0.01*-10). If the probability of a correct match increases to 0.999, the MapQ score will increase to 30. Conversely, as the probability of a correct match tends to zero, the MapQ score will also increase.

[0141] When the 16-base processing remains in the high-confidence, low-computation processing path, reads can contribute to stacking (when multiple reads align to similar positions in the reference sequence, these reads "stack" on top of each other on top of the reference sequence). When the 16-base processing returns to the low-confidence, high-computation processing path, reads can be removed from the stack. In one implementation, reads are processed in the low-confidence, high-computation processing path only when the number of candidates or the total number of sequence alignment positions is below a threshold such as 1000. While readings are being processed, the alignment results are stored.

[0142] Figure 9B Is using Figure 9A The method 900 shown is a conceptual diagram of the amount of data processed by two processing paths. After 16 sequencing cycles, 16 bases for each read are generated by the sequencing system. All reads are processed in the low-confidence, high-computation processing path during iteration 0. After 32 sequencing cycles, approximately 75% of the candidates are considered aligned after iteration 1. These candidates are processed in the high-confidence, low-computation processing path during iteration 2. After iteration 2, approximately 90% of the candidates are considered aligned and are processed in the high-confidence, low-computation processing path during iteration 3. When processing reads in the high-confidence, low-computation processing path, less computation and processing are required because only simple alignment is needed. Due to the large amount of data processed in the high-confidence, low-computation processing path and the less processing required in this path, the total time required is lower than if reads were processed only in the low-confidence, high-computation processing path. Therefore, alignment and variant identification results can be obtained before the sequencer finishes its run. These results can be provided to the user at time intervals based on available computational resources. Therefore, Method 900 can perform secondary analysis in a time-efficient manner to achieve real-time secondary analysis.

[0143] Figure 9C It shows Figure 10 Improvements to the prediction runtime of the comparator described in [the document]. Using only [the technology / method]. Figure 10The “Existing Processing” (regular or blue path) generates “base” data. The “Load Read 1” data illustrates the reduced processing loops when data from Read 1 is aligned, pre-stored, and then used to accelerate data processing in Read 2. Method 900 can implement one of two types of simple aligners for high-confidence, low-computation processing paths: a simple aligner that skips exact matches or a simple aligner that skips single mismatches. A simple aligner that skips single matches allows 0 or 1 mismatches. The “Skip Exact Match” data illustrates the reduced processing loops when the regular (blue) path is skipped if the 16 bases of the current iteration perfectly match the 16 bases of the reference sequence at a previously determined reference position. The “Skip Single Mismatch” data illustrates the reduced processing loops when the regular (blue) path is skipped if the 16 bases of the current iteration align to the 16 bases of the reference sequence at a previously determined reference position and have at most one mismatch. Figure 9C The results show that, compared to the baseline, method 900 reduces runtime by a factor of three when it utilizes a simple comparator that skips conventional processing upon detecting a single mismatch in a high-confidence, low-computation processing path. Note that these figures were generated by a prototype processor that does not include all processing steps and is therefore a predictive forecast.

[0144] Figure 10 This is another flowchart of Example Method 1000 for performing real-time secondary analysis. Figure 9A Methods 1000 and 900 shown can implement the same low-confidence, high-computation processing path and different high-confidence, low-computation processing paths. In method 1000, the high-confidence, low-computation processing path generates a MapQ score after a simple comparison, and the MapQ score is used to determine whether to continue processing in the high-confidence, low-computation processing path or return to the low-confidence, high-computation processing path.

[0145] A high percentage of runtime occurs on a small percentage of reads. In one implementation, if the confidence of success determined using a metric is low, the low-confidence, high-computation processing path of Method 900 or 1000 can skip the alignment and storage steps. In one implementation, a metric can be generated that indicates the number of candidate positions to which the subsequence can be aligned to the reference sequence. If the number of candidate sites is high, the confidence of alignment success will be low. In a second implementation, the confidence of alignment success is low if the base diversity in the sequence is low. Base diversity can be determined, for example, by calculating the number of unique n-mers in the subsequence, where an n-mer is a base sequence in the subsequence whose length is less than or equal to the length of the subsequence itself.

[0146] Alternative implementation scheme – variant identifier

[0147] Figure 11A and Figure 11B Existing mutation identification methods are illustrated, including the Strelka small mutation identifier (…). Figure 11A ) and the variant identification method disclosed herein ( Figure 11B A simplified flowchart of ). Figure 11A The diagram illustrates a small variant identifier that uses stacking information generated from the alignment unit as input. From the stacking, the small variant identifier identifies sequence variant regions, known as active regions. Next, de novo reassembly can be applied to the active regions. At each genomic location, probabilities are generated to determine the likelihood that the sequenced polynucleotide at that location is A, C, T, or G. Based on these probabilities, variants can be detected.

[0148] Figure 11B An embodiment of a variant identifier as disclosed in this invention is shown. In this embodiment, a metric is generated to determine whether a polynucleotide at a genomic location can be identified with high confidence. For example, a high-confidence decision can be generated if all polynucleotides at a given genomic location are identical. Alternatively, a high-confidence decision can be generated if the number of polynucleotides of the same type at a genomic location is above a threshold. Alternative metrics for determining high confidence can also be implemented. If a polynucleotide can be identified with high confidence, the probability formula can be skipped and a simple variant identification step can be performed. For example, a simple variant identifier can identify any variant detected with high confidence.

[0149] The probability generation and mutation identification steps of existing mutation identification methods can be combined, requiring up to 40% of the computation and processing of the mutation identifyer. Figure 11B A mutation identification method 1100 is illustrated, which implements both low-confidence, high-computation processing paths and high-confidence, low-computation processing paths of existing mutation identification methods. By adding a high-confidence, low-computation processing path, the Strelka mutation identifyer is optimized, and the processing is reduced by nearly 40%. High-confidence, low-computation processing paths can be added to alternative mutation identifyers.

[0150] like Figure 7B As shown, the mutation identifier can be executed within the iterative processing window. Figure 11A or Figure 11B The mutation identifier can be executed iteratively within the iterative processing window. Furthermore, more than one type of mutation identifier can be executed within the iterative processing window. For example, small mutation identifiers, such as Strelka, and alternative mutation identifiers, such as structural mutation identifiers or copy number mutation identifiers, can be executed within the iterative processing window.

[0151] In at least some of the previously described embodiments, one or more elements used in the embodiments may be used interchangeably in another embodiment, unless such substitution is technically impractical. Those skilled in the art will understand that various other omissions, additions, and modifications can be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and variations are intended to fall within the scope of the subject matter defined by the appended claims.

[0152] Regarding the use of virtually any plural and / or singular terms in this document, those skilled in the art can appropriately convert from plural to singular and / or from singular to plural depending on the context and / or application. For clarity, various singular / plural permutations may be explicitly described herein.

[0153] Those skilled in the art will understand that, generally, the terminology used herein and, in particular, in the appended claims (e.g., the subject matter of the appended claims) is intended to be “open-ended” (e.g., the term “comprising” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “at least having,” the term “comprising” should be interpreted as “including but not limited to,” etc.). Those skilled in the art will further understand that if there is an intent to enumerate a particular number of introduced claims, such intent will be explicitly stated in the claims, and without such a statement, such intent does not exist. For example, to aid understanding, the appended claims may contain the use of the introductory phrases “at least one” and “one or more” to introduce the claim statement. However, the use of these phrases should not be construed as implying that introducing the claim statement with the indefinite article “a” or “an” limits any particular claim containing such an introduced claim statement to embodiments containing only one such statement, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted as meaning “at least one” or “one or more”); the same applies to the use of definite articles used to introduce the claim statement. Furthermore, even when a specific number of introduced claims are explicitly stated, those skilled in the art will recognize that such a statement should be interpreted as referring to at least the number stated (e.g., simply stating "two statements" without any other modifiers implies at least two statements or two or more statements). Additionally, in cases where the convention of "at least one of A, B, and C" is used, the intended construction is generally the same as that which those skilled in the art would understand (e.g., "a system having at least one of A, B, and C" will include, but is not limited to, systems having a single A, a single B, a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In cases where the convention of "at least one of A, B, or C" is used, the intended construction is generally the same as that which those skilled in the art would understand (e.g., "a system having at least one of A, B, or C" will include, but is not limited to, systems having a single A, a single B, a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Those skilled in the art will further understand that any extractive word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to include the possibility of including one, any one, or both of these terms. For example, the phrase “A or B” will be understood to include the possibility of “A” or “B” or “A and B”.

[0154] Furthermore, when the features or aspects of this disclosure are described in accordance with the Markush Group, those skilled in the art will recognize that this disclosure is therefore also described in the form of any single member or subgroup of members of the Markush Group.

[0155] As those skilled in the art will understand, for any and all purposes, in relation to providing a written description, all scopes disclosed herein also encompass any and all possible subscopes and combinations thereof. Any listed scope can be readily identified as sufficiently descriptive and such that the same scope can be decomposed into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each scope discussed herein can be readily decomposed into a lower third, a middle third, and an upper third, etc. As those skilled in the art will also understand, all terms such as “up to,” “at least,” “greater than,” “less than,” and similar terms include the stated numbers and refer to scopes that can subsequently be decomposed into subscopes as discussed above. Finally, as those skilled in the art will understand, a scope includes each individual member. Thus, for example, a group having 1-3 items means a group having 1, 2, or 3 items. Similarly, a group having 1-5 items means a group having 1, 2, 3, 4, or 5 items, and so on.

[0156] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The aspects and embodiments disclosed herein are for illustrative purposes and not restrictive; the true scope and spirit are indicated by the appended claims.

Claims

1. A system for sequencing polynucleotides, comprising: A sequencing device configured to determine the nucleotide sequence of a polynucleotide; as well as A processor configured to control the sequencing device and execute instructions to perform methods including: Receive the first nucleotide subsequence of the polynucleotide; The first process is used to determine whether the first nucleotide subsequence has been aligned to the reference sequence at a first plurality of candidate positions exceeding a threshold confidence level. The sequencing device receives a second nucleotide subsequence of the polynucleotide, wherein the second nucleotide subsequence comprises the first nucleotide subsequence plus one or more additional nucleotides; If the alignment of the first nucleotide subsequence to the reference sequence exceeds a threshold confidence level, then one or more additional nucleotides in the second nucleotide subsequence are compared with the reference sequence, partly based on a first plurality of candidate positions, or If the alignment of the first nucleotide subsequence to the reference sequence does not exceed the threshold confidence level, the first process is repeated by aligning the entire second nucleotide subsequence to the reference sequence.

2. The system of claim 1, wherein the threshold confidence level depends on the number of mismatches or the probability of a correct match.

3. The system of claim 1, wherein the length of the first nucleotide subsequence is one or more nucleotides.

4. The system of claim 1, wherein the length of the second nucleotide subsequence is one or more nucleotides.

5. The system of claim 1, wherein comparing one or more additional nucleotides in the second nucleotide subsequence with a reference sequence includes a simple alignment process that is computationally more efficient than the first process in terms of memory usage or operand computation.

6. The system of claim 5, wherein the processor is further configured to determine a simple alignment score based on the simple alignment processing.

7. The system of claim 1, wherein if the first nucleotide subsequence aligns to the reference sequence, the processor is further configured to store data corresponding to at least one of the first plurality of candidate positions.

8. The system of claim 1, wherein the processor is further configured to store data corresponding to at least one of a second plurality of candidate positions generated by comparing the second nucleotide subsequence with the reference sequence.

9. The system of claim 1, wherein comparing one or more additional nucleotides in the second nucleotide subsequence with the reference sequence comprises comparing the second nucleotide subsequence with a corresponding sequence of the second nucleotide subsequence on the reference sequence based on the first plurality of candidate positions.

10. The system of claim 8, wherein the processor is further configured to determine a MapQ score for the mapping quality of each of the second plurality of candidate locations.

11. The system of claim 1, wherein determining whether the first nucleotide subsequence is aligned to the reference sequence begins before the sequencing reaction is complete.

12. The system of claim 1, wherein the processor is further configured to identify variations in the first nucleotide subsequence or the second nucleotide subsequence.

13. The system of claim 12, wherein mutation identification comprises: Variation identification is performed using either a first variation identification process or a second simple variation identification process, wherein the second variation identification process is computationally more efficient than the first variation identification process in identifying variations in the second nucleotide subsequence.

14. The system of claim 12, wherein the variant identification is performed based on a variant identification metric, using the output of the first processing or the output of a processing for comparing one or more additional nucleotides in the second nucleotide subsequence with the reference sequence.

15. The system of claim 14, wherein the variation identification metric is determined based on the number of different base types identified at the position of the reference sequence.

16. The system of claim 1, wherein, prior to the completion of the sequencing reaction, one or more additional nucleotides in the second nucleotide subsequence are compared with the reference sequence.

17. The system of claim 1, wherein the sequencing device enables sequencing-while-synthesizing.

18. A computer-implemented method for efficient sequencing of multinucleotides, comprising: During the sequencing run of the first nucleotide sequence, the first nucleotide sequence of the read is received from the sequencing device; Secondary analysis of the first nucleotide subsequence of a read is performed based on a reference sequence using either a first or second processing method, wherein the first nucleotide subsequence contains one or more additional nucleotides compared to a previous iteration, wherein the second processing method is computationally more efficient than the first processing method in performing secondary analysis, wherein the first processing method aligns the entire first nucleotide subsequence with the reference sequence, wherein the second processing method aligns one or more additional nucleotides with the reference sequence in part based on results from previous iterations, and wherein the secondary analysis includes: The first nucleotide subsequence is processed using the following method to determine a first plurality of candidate positions of the read that aligns to the reference sequence: If the read segment did not align to the reference sequence in the previous iteration, then the first processing is used. If not, the second processing is used, wherein the second processing determines the first plurality of candidate positions of the read segment in a more computationally efficient manner than the first processing; The first nucleotide subsequence is compared with the reference sequence to determine a first subsequence of the reference sequence that is highly similar to the first nucleotide subsequence; and Determine whether the sequencing device should generate additional nucleotide reads.

19. The method of claim 18, wherein performing secondary analysis of the first nucleotide subsequence using the second processing includes performing a simple alignment to determine a simple alignment score.

20. The method of claim 18, wherein the result of the secondary analysis includes the output of the first processing or the output of the second processing.

21. The method of claim 18, wherein performing secondary analysis includes variant identification of the first nucleotide subsequence, the variant identification comprising: The output of the first processing or the second processing is used to identify variants using a first variant identification process or a second variant identification process, wherein the second variant identification process is computationally more efficient than the first variant identification process in identifying variants in the first nucleotide subsequence.

22. The method of claim 21, wherein the result of the secondary analysis includes the output of the first variant identification process and the output of the second variant identification process.

23. The method of claim 18, further comprising providing the results of the secondary analysis to the user during the sequencing run.

24. The method of claim 23, wherein the results of the secondary analysis are provided to the user at fixed intervals.

25. The method of claim 23, wherein the results of the secondary analysis are provided to the user as requested by the user.

26. The method of claim 18, wherein the secondary analysis is performed based on whether the alignment of the first nucleotide subsequence to the reference sequence exceeds a threshold confidence level from previous iterations.

27. A computer-readable recording medium having a program recorded thereon for implementing the functions of the system of any one of claims 1-17 in a computer.

28. A computer-readable recording medium having a program that causes a computer to perform the method of any one of claims 18-26.

29. A system for sequencing polynucleotides, comprising: A sequencing device configured to determine the nucleotide sequence of a polynucleotide; as well as The processor is configured as follows: The instructions to perform a method comprising: Receive the first nucleotide sequence of the polynucleotide from the sequencing device; The first nucleotide subsequence is compared with a reference sequence to determine whether the alignment of the first nucleotide subsequence to the reference sequence has a first confidence level. The sequencing device receives a second nucleotide subsequence of the polynucleotide, wherein the second nucleotide subsequence comprises the first nucleotide subsequence plus one or more additional nucleotides; If the first nucleotide subsequence aligns to the reference sequence at the first confidence level, then one or more additional nucleotides in the second nucleotide subsequence are compared to the reference sequence; and The sequencing process is terminated based on the comparison of one or more additional nucleotides.

30. The system of claim 29, wherein the first confidence level depends on the number of mismatches, the probability of a correct match, the number of candidate positions for alignment of the first nucleotide subsequence with the reference sequence, or the diversity of bases in the first nucleotide subsequence.

31. The system of claim 29, wherein comparing one or more additional nucleotides in the second nucleotide subsequence with the reference sequence includes a simple alignment process that is computationally more efficient than the process used to align the first nucleotide subsequence with the reference sequence in terms of memory usage or computational operands.

32. The system of claim 31, wherein the processor is further configured to determine a simple alignment score based on the simple alignment processing.

33. The system of claim 31, wherein the simple alignment process is based in part on a first plurality of candidate positions obtained by comparing the first nucleotide subsequence with the reference sequence, and aligns only one or more additional nucleotides in the second nucleotide subsequence to the reference sequence.

34. The system of claim 29, wherein the processor is further configured to store data corresponding to at least one of a first plurality of candidate positions generated by comparing the first nucleotide subsequence with the reference sequence.

35. The system of claim 34, wherein comparing one or more additional nucleotides in the second nucleotide subsequence with the reference sequence comprises comparing the second nucleotide subsequence with a corresponding sequence of the second nucleotide subsequence on the reference sequence based on the first plurality of candidate positions.

36. The system of claim 29, wherein the processor is further configured to store data corresponding to at least one of a second plurality of candidate positions generated by comparing the second nucleotide subsequence with the reference sequence.

37. The system of claim 36, wherein the processor is further configured to determine a Map Quality (MapQ) score for each of the second plurality of candidate locations.

38. The system of claim 29, wherein determining whether the first nucleotide subsequence is aligned to the reference sequence begins before the sequencing reaction is complete.

39. The system of claim 29, wherein the processor is further configured to identify variations in the first nucleotide subsequence or the second nucleotide subsequence.

40. The system of claim 39, wherein mutation identification comprises: Variation identification is performed using either a first variation identification process or a second variation identification process, wherein the second variation identification process is more time-efficient or computationally resource-efficient than the first variation identification process in the identification of variations in the second nucleotide subsequence.

41. The system of claim 39, wherein the variant identification is based on a variant identification metric and is performed using the following outputs: a process for aligning the first nucleotide subsequence with the reference sequence, or a process for comparing one or more additional nucleotides in the second nucleotide subsequence with the reference sequence.

42. The system of claim 41, wherein the variation identification metric is determined based on the number of different base types identified at the position of the reference sequence.

43. The system of claim 29, wherein processing of the second nucleotide subsequence begins before the sequencing reaction is complete.

44. The system of claim 29, wherein the sequencing device enables sequencing-while-synthesizing.

45. The system of claim 29, further comprising aligning the entire second nucleotide subsequence to the reference sequence if the first nucleotide subsequence does not align to the reference sequence at the first confidence level.

46. ​​The system of claim 41, wherein the metric indicates whether a predetermined confidence level has been reached in the identification of the variant.

47. The system of claim 41, wherein the metric indicates whether a particular target variant has been identified.

48. The system of claim 41, wherein the processor is further configured to terminate receiving nucleotide subsequences from the sequencing device based on the metric.

49. The system of claim 41, wherein the processor is further configured to terminate the sequencing operation of the sequencing device based on the metric.

50. A computer-implemented method for determining the sequence of a polynucleotide having a nucleotide subsequence, the method comprising: During a sequencing run, a first nucleotide subsequence of a read is received from the sequencing device, wherein the first nucleotide subsequence contains one or more additional nucleotides compared to a previous iteration; Secondary analysis of the first nucleotide subsequence of a read is performed based on a reference sequence using either a first or a second processing method, wherein the second processing method is more time-efficient or computationally resource-efficient than the first processing method in performing the secondary analysis, wherein the first processing method aligns the entire first nucleotide subsequence to the reference sequence, wherein the second processing method aligns one or more additional nucleotides to the reference sequence in part based on results from previous iterations, and wherein the secondary analysis includes: The first nucleotide subsequence is processed using the following method to determine a first plurality of candidate positions of the read that aligns to the reference sequence: If the read segment did not align to the reference sequence in the previous iteration, then the first processing is used. If not, then the second process is used, wherein the second process determines the first plurality of candidate positions of the read segment in a more computationally efficient manner than the first process. The first nucleotide subsequence is compared with the reference sequence to determine a first subsequence in the reference sequence that is highly similar to the first nucleotide subsequence; and The sequencing process is terminated based on the comparison of one or more additional nucleotides.

51. The method of claim 50, wherein processing the first nucleotide subsequence using the second method includes performing a simple alignment to determine a simple alignment score.

52. The method of claim 50, wherein the result of the secondary analysis includes the output of the first processing, the output of the second processing, or any combination thereof.

53. The method of claim 50, wherein performing the secondary analysis includes variant identification of the first nucleotide subsequence, the variant identification comprising: The output of the first processing or the second processing is used to identify variants using a first variant identification process or a second variant identification process, wherein the second variant identification process is more time-efficient or computationally resource-efficient than the first variant identification process in identifying variants in the first nucleotide subsequence.

54. The method of claim 53, wherein the result of the secondary analysis includes the output of the first variant identification process, the output of the second variant identification process, or any combination thereof.

55. The method of claim 50, further comprising providing the results of the secondary analysis to the user during the sequencing run.

56. The method of claim 55, wherein the results of the secondary analysis are provided to the user at fixed intervals.

57. The method of claim 55, wherein the results of the secondary analysis are provided to the user as requested by the user.

58. The method of claim 50, wherein performing secondary analysis comprises performing secondary analysis of the first nucleotide subsequence of the read based on results from previous sequencing intervals from the sequencing run.