Basecaller with dilated convolutional neural networks

Dilated convolutional neural networks enhance base calling accuracy in capillary electrophoresis systems by accurately identifying base sequences and quality values, addressing low resolution issues and reducing sequencing costs.

JP7791844B2Active Publication Date: 2025-12-24LIFE TECHNOLOGIES CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022576416
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-11
Filing Date
2021-06-10
Publication Date
2025-12-24
Estimated Expiration
2041-06-10

AI Technical Summary

Technical Problem

Conventional base callers struggle with low accuracy in calling mixed bases due to sequencing artifacts, mobility shifts, and low resolution at the 5' and 3' ends, especially for amplicons shorter than 150 base pairs, leading to increased error rates and high sequencing costs.

Method used

Implementing dilated convolutional neural networks for base calling in capillary electrophoresis systems, utilizing deep learning to accurately identify base sequences and quality values by converting fluorescent signals into digital data, and using connectionist temporal classification to minimize loss and enhance base call precision.

Benefits of technology

Improves base calling accuracy for both pure and mixed bases, enhances read length, reduces sequencing costs, and increases fidelity of Sanger sequencing data, particularly at the 5' and 3' ends.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007791844000004
    Figure 0007791844000004
  • Figure 0007791844000005
    Figure 0007791844000005
  • Figure 0007791844000006
    Figure 0007791844000006
Patent Text Reader

Abstract

A method for automatically sequencing or base calling one or more DNA (deoxyribonucleic acid) molecules of a biological sample is described. The method includes measuring the biological sample using a capillary electrophoresis genetic analyzer to obtain at least one input trace containing digital data corresponding to fluorescence values ​​of multiple scans. Scan labeling probabilities for the multiple scans are generated using a trained artificial neural network including multiple layers, including a convolutional layer. A base call sequence is determined, including multiple base calls, for the one or more DNA molecules based on the scan labeling probabilities for the multiple scans.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Patent Application No. 16 / 899,545, filed June 11, 2020. These and all other external materials discussed herein are incorporated by reference in their entirety. [Background technology]

[0002] The present disclosure relates generally to systems, devices, and methods for base calling, and more particularly to systems, devices, and methods for base calling using deep learning for DNA sequence analysis using capillary electrophoresis.

[0003] In capillary electrophoresis (CE), a biological sample, such as a nucleic acid sample, is injected into a denaturing separation medium within a capillary at the inlet end of the capillary, and an electric field is applied to the end of the capillary. Different nucleic acid components in the sample, such as a polymerase chain reaction (PCR) mixture or other sample, migrate to the detector point at different speeds due to differences in their electrophoretic properties. As a result, they arrive at a photodetector (usually a fluorescence detector or ultraviolet (UV) absorbance detector operating in the visible light range) at different times. The result is displayed as a series of detected peaks, each ideally representing one nucleic acid component or species in the sample.

[0004] The magnitude of any given peak, including artifact peaks, is most often determined optically based on either UV absorption by nucleic acids, e.g., DNA, or fluorescence emission from one or more labeling dyes associated with the nucleic acids. UV and fluorescence detectors applicable to nucleic acid CE detection are well known in the art.

[0005] The CE capillary itself is often quartz, although other materials known to those skilled in the art can also be used. Several CE systems with both single and multiple capillary capabilities are commercially available. The methods described herein are applicable to any device or system for denaturing CE of nucleic acid samples.

[0006] Historically, Sanger sequencing using capillary electrophoresis (CE) genetic analyzers has been considered the gold standard for DNA sequencing. It offers high accuracy, long read times, and flexibility to support diverse applications in many research fields. The accuracy of base calls and quality values ​​(QVs) from Sanger sequencing on CE genetic analyzers is considered essential for the success of sequencing projects. Traditional base callers were previously developed to provide complete, integrated base calling solutions supporting sequencing platforms and applications. They were originally designed to base call long plasmid clones (pure bases) and were subsequently expanded to base call mixed-base data to support variant identification.

[0007] However, apparent mixed bases may be called pure bases even with high predicted QVs. False positives, where pure bases are mistakenly called mixed bases, are also relatively frequent due to sequencing artifacts such as dye blobs, polymerase slippage and primer impurities resulting in n-1 peaks, and mobility shifts. Clearly, improved base calling and QV accuracy for mixed bases is necessary to support sequencing applications for identifying variants such as single nucleotide polymorphisms (SNPs) and heterozygous insertion-deletion variants (het-indels). The base calling accuracy of conventional base callers at the 5' and 3' ends is also relatively low due to mobility shifts and low resolution at the 5' and 3' ends. Conventional base callers also struggle to call amplicons shorter than 150 base pairs (bps), especially those shorter than 100 bps, and may be unable to estimate the average peak spacing, average peak width, spacing curve, and / or width curve, resulting in an increased error rate.

[0008] Therefore, improving the accuracy of base calling for pure and mixed bases, especially at the 5' and 3' ends, is highly desirable, so that base calling algorithms can increase the fidelity of Sanger sequencing data, improve variant discrimination, increase read length, and also save sequencing costs for sequencing applications.

[0009] Modern base callers frequently use recurrent neural network-based models to identify base-calling sequences based on raw input data. Although recurrent neural networks can adequately model time-series data in base calling due to their recurrent structure, the speed of recurrent network-based base callers can be severely limited, especially when dealing with longer sequence reads, because computation at a given time point must wait for the results of previous time points. Summary of the Invention

[0010] The systems and methods are described for use in capillary electrophoresis deep learning based base calling systems, such as in convolutional neural network based base calling systems that use microfluidic separations (separation occurs through microchannels etched into glass, silicon or other substrates) or capillary electrophoresis genetic analyzers based on separation by capillary electrophoresis using single or multiple cylindrical capillary tubes.

[0011] Convolutional architectures, such as dilated convolutional neural networks implemented in embodiments of the invention described herein, perform well in gene sequence modeling tasks, outperform recurrent networks, and potentially reach state-of-the-art accuracy in a wide range of sequence modeling tasks. Training and inference of convolutional neural networks is much faster than recurrent networks, such as long short-term memory (LSTM) networks. In particular, dilated convolutional neural networks have the potential to achieve exponentially larger receptive fields with fewer parameters and fewer layers.

[0012] A method for automatically calling bases of one or more DNA (deoxyribonucleic acid) molecules of a biological sample is described. The method includes converting a plurality of fluorescent signals of the biological sample, each of which is measured by a capillary electrophoresis genetic analyzer and converted into at least one input trace containing digital data corresponding to the fluorescent values ​​in a plurality of scans. Scan labeling probabilities are generated for each of the plurality of scans. The scan labeling probabilities are generated using a trained deep neural network including multiple layers, including a convolutional layer. A base call sequence including multiple base calls for the one or more DNA molecules is determined based on the one or more scan labeling probabilities for each of the plurality of scans.

[0013] In some embodiments, the base call position of each base call in the base call sequence is also determined, and the base call position corresponds to the scan position of the peak scan label probability associated with the base call. In some embodiments, the search for the peak probability associated with a given base call is made more efficient by first identifying the first and last scans in the scan range of the scan that corresponds to the scan label probability associated with the given base call, and then searching only within that scan range.

[0014] In some embodiments, the quality value of each base call is determined using feature values ​​derived from the scan probability values ​​associated with the base call, rather than using the value of the image trace associated with the base call. Further, in some embodiments of the invention, the neural network is trained to call two mixed bases per base call position. In some embodiments, the neural network can be trained to call more than two mixed bases per base call position.

[0015] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. [Brief explanation of the drawings]

[0016] [Figure 1] 1 illustrates a capillary electrophoresis sequencing system according to one embodiment of the present invention. [Figure 2] 1 shows an exemplary electropherogram that may be displayed in accordance with one embodiment of the present invention. [Figure 3] 1 illustrates a capillary electrophoresis genetic analysis process according to some embodiments of the present invention. [Figure 4] 1 shows a diagram of exemplary input and output data that may be displayed in accordance with one embodiment of the present invention; [Figure 5]1 illustrates a deep learning based base calling workflow process according to one embodiment of the present invention. [Figure 6] 1 illustrates a deep neural network architecture according to one embodiment of the present invention. [Figure 7] 1 illustrates a residual block architecture according to one embodiment of the present invention. [Figure 8] 1 illustrates a method for generating base call sequences according to one embodiment of the present invention. [Figure 9] 1 illustrates a method for generating scan ranges and scan positions for one or more base calls in a base call sequence, according to one embodiment of the present invention. [Figure 10] 1 illustrates a method for training a scan landmark model according to an embodiment of the present invention. [Figure 11] 1 illustrates a method for constructing a trained quality value lookup table according to one embodiment of the present invention. [Figure 12] 1 illustrates a block diagram of an exemplary computing device that may incorporate embodiments of the present invention.

[0017] While the invention has been described with reference to the above-described drawings, it is to be understood that the drawings are intended to be illustrative and that other embodiments are consistent with the spirit and scope of the invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] Various embodiments will now be described in more detail below with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific examples in which the embodiments may be practiced. This specification may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this specification will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. Among other things, the specification may be embodied as methods or devices. Accordingly, any of the various embodiments herein may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Therefore, the following specification should not be construed in a limiting sense.

[0019] This patent application contains material related to PCT Application No. PCT / US2019 / 065540, filed December 10, 2019, with a priority date of December 10, 2018, which is incorporated herein by reference in its entirety. These and all other technical publications, patent applications, scientific publications, and all other external materials discussed herein are incorporated by reference in their entirety.

[0020] The embodiments of the invention discussed herein utilize the principles of DNA replication used in Sanger dideoxy sequencing, a process that harnesses the ability of DNA polymerases to incorporate 2',3'-dideoxynucleotides - nucleotide base analogs that lack the 3'-hydroxyl group essential for phosphodiester bond formation.

[0021] Sanger dideoxy sequencing requires a DNA template, sequencing primers, DNA polymerase, deoxynucleotides (dNTPs), dideoxynucleotides (ddNTPs), and a reaction buffer. Sanger dideoxy sequencing was originally designed to set up four separate reactions, each containing a radiolabeled nucleotide and either ddA, ddC, ddG, or ddT. The annealing, labeling, and termination steps are performed in separate heat blocks. DNA synthesis is performed at 37°C, the temperature at which DNA polymerase has optimal enzymatic activity. DNA polymerase adds either a deoxynucleotide or the corresponding 2',3'-dideoxynucleotide at each step of chain elongation. Whether a deoxynucleotide or a dideoxynucleotide is added depends on the relative concentrations of both molecules. Once a deoxynucleotide (A, C, G, or T) is added to the 3' end, chain elongation can continue. However, once a dideoxynucleotide (ddA, ddC, ddG, or ddT) is added to the 3' end, chain elongation terminates. Sanger dideoxy sequencing results in the formation of extension products of various lengths that terminate at the 3' end with a dideoxynucleotide.

[0022] The extension products are then separated by electrophoresis. During electrophoresis, an electric field is applied to cause negatively charged DNA fragments to migrate toward a positive electrode. The speed at which DNA fragments migrate through the medium is inversely proportional to their molecular weight. This process of electrophoresis can separate extension products by size with a resolution of one base.

[0023] The automated DNA fluorescence-based cycle sequencing system manufactured and used in embodiments of the present invention by Applied Biosystems, Inc. is an extension and improvement of Sanger dideoxy sequencing. Applied Biosystems' automated DNA sequencing generally follows the steps of DNA template preparation, cycle sequencing, post-cycle sequencing purification, capillary electrophoresis, and data analysis. An exemplary fluorescence-based cycle sequencing system that can be used in embodiments of the present invention is described in "DNA Sequencing by Capillary Electrophoresis Chemistry Guide" (3), published by Thermo Fisher Scientific, Inc., which is incorporated herein by reference in its entirety. rd Edition, 2016).

[0024] Similar to Sanger sequencing, fluorescence-based cycle sequencing requires a DNA template, sequencing primers, a thermostable DNA polymerase, deoxynucleoside triphosphates / deoxynucleotides (dNTPs), dideoxynucleoside triphosphates / dideoxynucleotides (ddNTPs), and a buffer. However, unlike the Sanger method, which uses radioactive materials, cycle sequencing uses fluorescent dyes to label extension products and combines the components in a reaction that undergoes cycles of annealing, extension, and denaturation in a thermal cycler. Thermal cycling of the sequencing reaction creates and amplifies extension products that terminate at one of four dideoxynucleotides. The ratio of deoxynucleotides to dideoxynucleotides is optimized to generate a balanced population of long and short extension products.

[0025] The automated cycle sequencing procedure used in some embodiments of the present invention incorporates fluorescent dye labeling using dye-labeled dideoxynucleotides (dye terminators) using four different dyes, each of which emits a unique wavelength when excited by light, allowing the fluorescent dyes in the extension products to identify the 3'-terminal dideoxynucleotide as A, C, G, or T.

[0026] In dye terminator chemistry, each of the four dideoxynucleotide terminators is tagged with a different fluorescent dye. A single reaction containing the enzyme, nucleotides, and all of the dye-labeled dideoxynucleotides is performed. The products from this reaction are injected into a single capillary.

[0027] In one embodiment of the present invention, the cycle sequencing reaction is directed by a highly modified, thermostable DNA polymerase selected to enable the incorporation of dideoxynucleotides, process GC-rich and other difficult sequence stretches, and generate peaks of variable height. The modified DNA polymerase is also formulated with pyrophosphatase to prevent reversal of the polymerization reaction (pyrophosphorolysis).

[0028] In one embodiment of the present invention, Applied Biosystems Cycle Sequencing Kits available for dye terminator chemistry include BigDye Terminator v1.1 and v3.1 Cycle Sequencing Kits, dGTP BigDye Terminator v1.0 and v3.0 Cycle Sequencing Kits, and BigDye Direct Cycle Sequencing Kits. The fluorescent dyes used in BigDye terminators, BigDye primers, and BigDye Direct have narrower emission spectra and less spectral overlap than the rhodamine dyes used in previous sequencing kits. As a result, the dyes may tend to produce less noise.

[0029] Historically, DNA sequencing products were separated using polyacrylamide gels manually poured between two glass plates. Capillary electrophoresis, using a denaturing flow polymer, has largely replaced the use of gel separation techniques due to significant improvements in workflow, throughput, and ease of use. Fluorescently labeled DNA fragments are separated according to molecular weight. Because capillary electrophoresis does not require gel pouring, DNA sequence analysis using CE is more easily automated and can process many more samples at once.

[0030] During capillary electrophoresis, the extension products of a cycle sequencing reaction enter the capillary as a result of electrokinetic injection. Application of a high-voltage charge to a buffered sequencing reaction forces negatively charged fragments into the capillary. The extension products are separated by size based on their total charge. The electrophoretic mobility of samples can be affected by run conditions (buffer type, concentration, and pH, run temperature, amount of applied voltage, and type of polymer used).

[0031] Shortly before reaching the positive electrode, fluorescently labeled DNA fragments, separated by size, move across the path of a laser beam. The laser beam causes dyes on the fragments to fluoresce. In one embodiment of the present invention, the optical detection system of an Applied Biosystems genetic and / or DNA analyzer detects the fluorescence. Data acquisition software used in one embodiment of the present invention converts the fluorescent signal into digital data and then records the data in an AB1 (.ab1) file. Because each dye emits light at a different wavelength when excited by the laser, all four colors, and therefore all four bases, can be detected and distinguished in a single capillary injection.

[0032] 1 illustrates a system 100 according to an exemplary embodiment of the present invention. System 100 includes a capillary electrophoresis ("CE") instrument 101, one or more computers 103, and a user device 107.

[0033] 1 , in one embodiment, a CE instrument 101 includes a source buffer 118 that contains a buffer and receives a fluorescently labeled sample 120, a capillary 122, a destination buffer 126, a power supply 128, and a controller 112. The source buffer 118 is in fluid communication with the destination buffer 126 via the capillary 122. The power supply 128 applies a voltage to the source buffer 118 and the destination buffer 126, generating a voltage bias across an anode 130 of the source buffer 118 and a cathode 132 of the destination buffer 126. The voltage applied by the power supply 128 is configured by the controller 112, which is operated by the computing device 103. The fluorescently labeled sample 120 near the source buffer 118 is drawn through the capillary 122 by the voltage gradient, and optically labeled nucleotides of DNA fragments within the sample are detected as they pass an optical sensor 124 on their way to the destination buffer 126. Different sized DNA fragments within the fluorescently labeled sample 120 are drawn through the capillary at different times due to their size.

[0034] The optical sensor 124 detects the fluorescent labels on the nucleotides as an image signal and communicates the image signal to the computing device 103. The computing device 103 aggregates the image signals as sample data and utilizes a base calling computer program product 104 to operate the deep neural network 102 to convert the sample data into processed data including base call sequences and quality values, and to generate an electropherogram that can be displayed on the display 108 of the user device 107.

[0035] Instructions for implementing deep neural network 102 reside on computing device 103 in computer program product 104 stored in storage 105, where the instructions are executable by processor 106. When processor 106 is executing the instructions of computer program product 104, the instructions, or portions thereof, are typically loaded into working memory 109, from which the instructions are readily accessed by processor 106. In one embodiment, computer program product 104 is stored on storage 105 or other non-transitory computer-readable medium (which may include being distributed across media in different devices and locations). In an alternative embodiment, the storage medium is transitory.

[0036] In one embodiment, processor 106 actually includes multiple processors, which may include additional working memory (additional processors and memory not separately shown), including a graphics processing unit (GPU) that includes at least thousands of arithmetic logic units to support massively parallel computations. GPUs are frequently utilized in deep learning applications because they can perform associated processing tasks more efficiently than general-purpose processors (CPUs). Other embodiments include one or more specialized processing units, including systolic arrays and / or other hardware configurations that support efficient parallel processing. In some embodiments, such specialized hardware operates in conjunction with a CPU and / or GPU to perform various processes described herein. In some embodiments, such specialized hardware includes an application-specific integrated circuit (which may refer to a portion of an application-specific integrated circuit), a field-programmable gate array, or the like, or a combination thereof. However, in some embodiments, a processor, such as processor 106, may be implemented as one or more general-purpose processors (preferably having multiple cores) without necessarily departing from the spirit and scope of the present invention.

[0037] User device 107 includes display 108 for displaying the results of the processing performed by neural network 102. In alternative embodiments, a neural network, such as neural network 102, or portions thereof, may be stored in a memory device and executed by one or more processors present in CE equipment 101 and / or user device 107. Such alternatives do not depart from the scope of the present invention.

[0038] FIG. 2 shows an exemplary electropherogram 200 that can be displayed according to one embodiment of the present invention. Electropherogram 200 includes a graph (Y-axis: relative fluorescence units (RFU) and X-axis: scans) that displays image signals of detected fluorescent labels on nucleotides as a series of peaks, e.g., 210, 211, 212, and 213. Signals corresponding to fluorescently labeled nucleotides can be displayed in four different colors, which can be represented in FIG. 2 and other figures herein as color, grayscale, or different variations of black and white diagonal lines representing various colors. Each color represents the base required for that peak (e.g., in IUPAC-IUB notation, T=red, C=blue, G=black, and A=green, respectively). Two or more (e.g., three or four) peaks may occur at one position, in which case they can be called mixed bases (e.g., a mixed base of two peaks can be represented in IUPAC-IUB notation as follows: A+C=M, A+G=R, A+T=W, C+G=S, C+T=Y, G+T=K).

[0039] Referring to FIG. 3 , a CE process 300 utilized in one embodiment of the present invention includes configuring the operating parameters of a capillary electrophoresis instrument to sequence at least one fluorescently labeled sample (block 302). Configuring the instrument may include creating or importing a plate setup for running a series of samples and assigning labels to plate samples to assist in processing collected image data. This process may also include communicating configuration controls to a controller to initiate the application of voltages at predetermined times. In block 304, the CE process 300 loads the fluorescently labeled sample into the instrument. After the sample is loaded into the instrument, the instrument transfers the sample from the plate well into a capillary tube and then positions the capillary tube in a starting buffer at the start of the capillary electrophoresis process. In block 306, the CE process 300 begins running the instrument after the sample is loaded into the capillary, applying a voltage to buffer solutions positioned at both ends of the capillary to form an electric gradient that transports DNA fragments of the fluorescently labeled sample from the starting buffer to a destination buffer and across an optical sensor. In block 308, the CE process 300 detects individual fluorescent signals on the nucleotides of the DNA fragments as they move through the optical sensor toward the destination buffer and communicates the image signal to a computing device. In block 310, the CE process 300 aggregates the image signals from the optical sensor in the computing device, analyzes the aggregated image signal, and generates sample data corresponding to the fluorescent intensities of the nucleotides of the DNA fragments. In block 312, the CE process 300 processes the sample data using both a deep learning neural network and a sequence analysis algorithm to help identify the bases called within the DNA fragments at a specific time point (corresponding to a specific scan number in multiple scans). In block 314, the CE process 300 displays the processed data on a display device as analyzed traces and base call sequences displayed in an electropherogram.

[0040] FIG. 4 shows a diagram of exemplary input and output data 400 that may be generated and / or displayed in accordance with one embodiment of the present invention. The input data includes an analyzed trace 410 generated using the CE process 300, which may be displayed in an electropherogram similar to that shown in FIG. 2. The output data includes a plurality of base call positions 420, a plurality of base call labels 430, and a plurality of quality values ​​440. FIG. 4 also shows intermediate data CTC scan label probabilities 450 corresponding to each base call, which may comprise the output of a dilated convolutional neural network, as described below and implemented in embodiments of the present invention described herein. In certain embodiments, the base call positions 420, base call labels 430, and quality values ​​440 are typically displayed in exemplary embodiments, while the CTC scan label probabilities are generally not displayed in the electropherogram 200 of FIG. 2.

[0041] In some embodiments of the present invention, the user can select whether the input data contains only pure bases or mixed bases. Base calling is the interpretation of dye data used to draw an electropherogram. This determines which nucleotide (represented by base call label 430) belongs to which position (represented by base call position 420). Each color shown in the input analysis trace 410 and base call label 430 represents a base (here, instead of the standard color notation, each base may be rendered in grayscale and / or with different / separate dotted / dashed lines). In FIG. 4, the input analysis trace 410 and base call label 430 are rendered with the colors called for each peak: T = red, C = blue, G = black, A = green. As mentioned above, base calls may also be mixtures of two or more (e.g., three or four) nucleotides, showing two or more peaks that overlap or are slightly offset from each other, possibly with different peak heights.

[0042] In embodiments of the invention, a quality value 440 of Figure 4 is also generated for each base call. The quality values ​​440 are shown in the output data 400 of Figure 4 as vertical bars that vary in height for each base 430 called depending on the calculated estimated probability of error or quality value. The calculation of quality values ​​performed in embodiments of the invention is described further herein.

[0043] FIG. 5 shows a diagram of a deep base calling workflow process 500 according to one embodiment of the present invention. The input data includes an analyzed trace 510 generated using the CE process 300, which may be displayed in an electropherogram similar to that shown in FIG. 2. The input trace 510 may be a sequence of dye relative fluorescence units (RFUs) collected from a capillary electrophoresis (CE) instrument, or raw spectral data collected directly on the CE instrument. The input trace 510 may be divided into several windows, each containing multiple scans. In one embodiment of the present invention, the scan window size determines the number of scans to scan into the label model 520.

[0044] The scan label model 520 receives an input scan window and generates scan label probabilities for all scans within the scan window. The scan label model 520 may include one or more trained models. A model may be selected to utilize to generate the scan label probabilities. In one embodiment of the present invention, a deep learning model including a neural network 520 is trained to learn an optimal mapping function from the analyzed traces 510 and scan label probabilities 530. In one embodiment of the present invention, the neural network 520 includes a dilated convolutional neural network trained to minimize the loss between the target sequence of bases and the corresponding predicted scan label probabilities 530 using a connectionist temporal classification (CTC) loss function, as described further herein below. The deep learning model may be trained according to the process shown in FIG. 10.

[0045] The decoder 540 receives the scan label probabilities for the assembled scan window. The decoder 540 then decodes the scan label probabilities into base calls for the input trace sequence. The decoder 540 can utilize a prefix beam search or other decoder on the assembled label probabilities to find base calls for the sequencing sample.

[0046] The CTC scan labeling probabilities 530 are combined using a CTC decoder and segmentation module 540, which walks through the scan labeling probabilities 530 for all scans and generates the sequence with the maximum labeling probability as the final result. The CTC decoder and segmentation module 540 further finds the scan range and then finds the scan position of the peak labeling probability within the scan range for each called base to generate a base call (label) and base call position 550 for the sequence. The output data generated by the CTC decoder and segmentation module 540 is then used by a base call quality value (QV) predictor 560 to calculate a quality value (QV) 570, i.e., a quality score for each called base, as described further herein below. The base call quality value (QV) predictor 560 uses the features calculated from the CTC scan labeling probabilities as keys to find a quality score for each called base from a trained QV lookup table.

[0047] Dilated Convolutional Neural Networks Recent studies have shown that convolutional neural network architectures outperform recurrent neural networks and can achieve state-of-the-art accuracy in speech synthesis, word-level language modeling, and machine translation. For example, the generic temporal convolutional network (TCN) architecture described in the following reference (Bai, Shaojie, Kolter, J. Zico and Koltun, Vladlen, An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling, arXiv:1803.01271v2[cs.LG], April 19, 2018) ("Bai et al.") has been evaluated on a wide range of sequence modeling tasks beyond Sanger sequencing using CE, including polyphonic music modeling, word-level sequence modeling, and character-level sequence modeling. Bai et al.'s results show that TCNs outperform canonical recurrent networks such as LSTMs while demonstrating a longer effective memory.

[0048] Embodiments of the present invention utilize a neural network architecture similar to TCN, but in some embodiments, the neural network architecture utilized has some important differences. In one embodiment of the present invention, the network architecture differs from TCN in that a one-dimensional (1D) fully expanded convolution is used instead of a one-dimensional (1D) fully expanded causal convolution. TCN uses causal convolution. In this case, the output at time t is convolved only with elements of the previous layer before time t. However, in CE-based calling, since the entire input scan trace is available during base calling, past, present, and future scan data can be utilized. Some embodiments of the present invention utilize a one-dimensional non-causal fully expanded convolutional network where the output at time t is convolved not only with elements before time t but also with elements after the previous layer. The length of the subsequent layer is the same as the previous layer with zero padding added. In one embodiment of the present invention, expanded convolution is used to achieve an exponentially larger receptive field with fewer parameters and fewer layers.

[0049] FIG. 6 shows a deep neural network architecture 600 according to one embodiment of the present invention. In one embodiment shown in FIG. 6, the deep neural network architecture 600 is trained to learn an optimal mapping function from an input analysis trace 610 to an output of scan label probabilities 670.

[0050] The input analysis trace 610 may include a plurality of scans that segment a plurality of fluorescence signals of the input analysis trace 610. Since the migration speed of DNA is unstable and may be slower than the measurement speed of fluorescence signals, the lengths of the base sequences may vary and may be much shorter than the segments of fluorescence signal measurement. Therefore, the main task of the model is to convert scans of fluorescence signal measurements of fixed length T into base sequences of non-uniform length M (0 < M < T).

[0051] Network architecture 600 includes multiple hidden layers in each of four residual blocks, shown as blocks 620, 630, 640, and 650, each of which includes one or more non-causal convolutional layers. In one embodiment of the present invention, a filter size k=9 is used for all residual blocks 620-650. The expansion factor for each residual block is, in one embodiment of the present invention, d=2. (i-1) where i is the depth of the residual block in the neural network, and i = 1, 2, 3, and 4 for the residual blocks 620, 630, 640, and 650, respectively. The feature map sizes are given as w = 32, 48, 64, and 128 for the residual blocks 620, 630, 640, and 650, respectively. The stacked residual blocks function as a feature extractor that maps the fluorescent signal measurements to a feature space. In some embodiments of the present invention, the extended acausal convolution is performed in the time dimension, so that the extracted features indicate the correlation of the fluorescent signal measurements at different time points. Subsequently, a 1 × 1 convolutional reduction layer 655 is added after the last residual block to reduce the number of extracted features to match the number of output labels, and a softmax function layer 660 is added after the 1 × 1 convolutional reduction layer. In one embodiment of the present invention, a Softmax function 660 converts the output of the 1x1 convolutional reduction layer 655 into a probability matrix, where each matrix row indicates the probability of a base occurring at that time to generate a number of scan label probabilities 670.

[0052] FIG. 7 illustrates a residual block architecture according to some embodiments of the present invention.

[0053] In some embodiments of the residual block architecture shown in Figure 7, two one-dimensional (1D) fully dilated convolutional layers 702 and 708 of Figure 7 are stacked within the residual block. Layer normalization 704 and 710 of Figure 7 and spatial dropout 706 and 712 of Figure 7 may also be added after each dilated convolution for effective training and regularization in some embodiments of the present invention.

[0054] Nonlinearities, such as one or more rectified linear units (ReLUs), shown here as ReLUs 714, 716, and 718 in FIG. 7, can also be included after the dilation convolution. Within the residual block, a skip connection can be used to direct the input 701 of block 700 in FIG. 7 to the output 730 of the block, which is useful for deep network training. If the input and output have different widths, an additional optional 1×1 convolution 720 in FIG. 7 can be applied to the input to match the width of the output. Multiple residual blocks can be stacked together, as shown in an exemplary manner in FIG. 6, to reach the desired receptive field by increasing the dilation factor d exponentially with the depth of the network.

[0055] Connectionist Temporal Classification For tasks such as automatic speech recognition (ASR), the process is often divided into a series of subtasks, such as speech segmentation, acoustic modeling, and language modeling. Each of these subtasks is solved independently by a separate trained model. In 2006, connectionist temporal classification (CTC) was introduced by Alex Graves (see Graves, Alex, Supervised Sequence Labeling with Recurrent Neural Networks, volume 385 of Studies in Computational Intelligence, Springer 2012), which makes it possible to train deep learning networks end-to-end for tasks such as ASR.

[0056] CTC is an objective function that allows deep learning models to be trained for sequence-to-sequence tasks without requiring prior alignment between input and target sequences. More specifically, we use CTC as the loss function to train a dilated convolutional neural network to minimize the loss between the target sequence of bases and the predicted scan label probabilities, which are the network's output but normalized by a Softmax function.

[0057] In addition to base labels (pure bases with a single nucleotide, A, C, G, or T, or mixed bases with two nucleotides), additional "blank" labels are introduced into CTCs. The blank labels have two important functions. First, they can separate bases, especially consecutive repeating bases such as AAAA. This allows us to label scans that do not belong to valid bases and predict sequences of various lengths of bases.

[0058] Each input scan can be labeled as a base or a blank. A CTC path is the sequence of all scan labels of bases or blanks. The probability of a CTC path is the product of the scan label probabilities of all scans within that CTC path. A CTC path is converted to a base call sequence by collapsing consecutively repeated labels and then removing blanks. Because many possible CTC paths can be converted to a single base call sequence, the total probability of a base call sequence is the sum of all the probabilities of all possible CTC paths for that base call sequence. For a given input scan sequence x and a target base call sequence y*, the probability of y* given x is written as Pr(y*|x). The CTC loss function is defined as the negative logarithm of the probability, -log(Pr(y*|x)). A dilated convolutional neural network is trained to minimize the CTC loss.

[0059] Segmentation by CTC decoder and prefix beam search Because many possible CTC paths can be converted to a single base call sequence, the probabilities of all possible paths that generate the same base call sequence are calculated and summed to obtain the base call sequence probability. The final base call result can be obtained by selecting the base call sequence with the highest probability. In the embodiment of the invention described herein, a CTC prefix beam search was used to efficiently decode the CTC output (see Graves, Alex. "Towards End-to-End Speech Recognition with Recurrent Neural Networks," Proceedings of the 31st International Conference on Machine Learning, Beijing, China, 2014. JMLR:W&CP volume 32). In one embodiment of the present invention, a CTC decoder algorithm is used to decode the scan label probabilities and generate the final base call sequence. This algorithm is then extended to discover the scan range and locate the scan position of each base call within the final sequence.

[0060] 8 shows a method 800 for generating base call sequences according to one embodiment of the present invention. The prefix beam search starts with an empty base call sequence as the first candidate in step 810. The method 800 then iterates through all scans within a window of input traces to determine CTC scan labeling probabilities in step 820.

[0061] In step 830, for each scan within the scan window, all candidate subsubsequences are extended with all possible labels (all possible options for bases (pure or mixed), or blank labels) and scored by incorporating the scan label probabilities of the extended labels in that scan. In step 840, a fixed-size subset B of the K highest-scoring extended candidates is saved and extended in the next scan. Typically, the candidates in each scan are called prefixes, and the number of saved candidates K is called the beam width. In step 850, a separate candidate subset C is created to save all the top candidates for each scan during the beam search. If the highest-scoring candidate for scan t differs from the top candidate for the previous scan t-1, it is saved in subset C, and scan t is assigned to this candidate. After the final scan, the best-scoring base call sequence is returned as the final base call result, and the candidate subset of the top candidates saved for each scan during the beam search is also returned in step 860.

[0062] 9 illustrates a method 900 for generating scan ranges and scan positions for one or more base calls in a base call sequence, according to one embodiment of the present invention. In method 900, a base call sequence y, which is the most likely sequence found by the decoder in method 800, is denoted as a sequence of length L, and the base calls at positions i=1,...,L in the sequence are denoted as y i The method 900 calculates a base call at position i in sequence y at scan position t i Find where i=1,...,L.

[0063] The method starts at step 910 with the first base call of the final base call sequence, where i=1. At iteration i, the subsequence y 1 with the first i base calls in sequence y is (1...i) is examined in step 920. The method then examines the subsequence y in the candidate subset C in step 930. (1...i)Iterates through all base calls in base call sequence y as shown until each base call in the entire base call sequence y has been examined by searching for

[0064] If the examined subsequence is in the candidate subset C determined in method 800, method 900 proceeds to step 940, where subsequence y (1...i) The scan assigned to i is used as the starting scan of the next base call and is extended by prefix scan number nt until the starting scan of the next base call. i End scan of discovered. i Once the scan range for y is determined, scan positions with peak scan labeling probabilities within the defined scan range are assigned the base call y, as shown in step 950. i In step 960, the scan position and start and end scan of each base call in the entire base call sequence y are returned.

[0065] The pseudocode in Algorithm 1 describes the CTC decoder and segmentation procedure for a CTC network implemented in one embodiment of the present invention. The blank probability Pb(y,t) is the probability of an output sequence y at a particular time t that comes from one or more CTC paths that end in a blank indicator. The non-blank probability Pnb(y,t) is the probability of an output sequence y at a particular time t that accounts for all CTC paths that end in a non-blank indicator. The total probability Pt(y,t) is the sum of Pb(y,t) and Pnb(y,t).

[0066] Given an input scan sequence x, the probability of emitting a tag (or blank) with index k at time t is denoted as Pr(k,t|x). The expansion probability of y with tag k at time t, Pr(k,y,t), is defined as:

number

[0067] where y e is the final indicator of y. Furthermore, y ← Define y as a prefix of y with the last label removed, and y 1,...,i is the subsequence of y containing only the first i labels, and

number

number

[0068] Referring to Figure 10, a scan label model training method 1000 in one embodiment of the present invention receives a sequencing dataset (block 1002). The dataset may include pure base datasets and mixed base datasets. In some embodiments of the present invention, the data in these datasets is annotated and manually reviewed, and the correct base call sequences ("ground truth") are written to each data file (e.g., an .ab1 data file). In one embodiment, the sequencing dataset may use representative data files compiled from data generated using a wide variety of CE genetic and DNA analysis instruments and instrument configurations (e.g., voltage, temperature, chemistry type, capillary array length, etc.).

[0069] For example, in one embodiment of the present invention, a pure base dataset may include approximately 49 million base calls, and a mixed base dataset may include approximately 13.4 million base calls. A mixed base dataset may consist primarily of pure bases and occasional mixed bases. For each sample in the dataset, the entire trace is divided into scan windows or segments (block 1004). Each scan window may have 500 scan segments. The trace may be a sequence of pre-processed or processed dye RFUs. Furthermore, the scan window for each sample may be shifted by 250 scans to minimize bias in scan position during training. Next, annotated base calls are listed for each scan window (block 1006). These are used as target sequences during training. Next, training samples are constructed (block 1008). Each of them may include a scan window containing 500 scans and their respective annotated base calls. The CNN is initialized (block 1010). In one embodiment of the present invention, the CNN may include one or more residual blocks and one 1x1 convolutional reduction layer, as shown in Figure 7. A Softmax layer may be used as the output layer of the CNN, outputting scan label probabilities for all scans of the input trace.

[0070] Next, a mini-batch of training samples is selected (block 1011). The mini-batch can be randomly selected from the training dataset at each training step. The mini-batch of training samples is then applied to the CNN (block 1012). All scan label probabilities within the input scan window are output (block 1014). The loss between the output scan label probabilities and the target annotated base calls is calculated.

[0071] A connectionist temporal classification (CTC) loss function can be used to calculate the loss between the output scan label probabilities and the target annotated base calls. The network weights are updated to minimize the CTC loss for a mini-batch of training samples (block 1020). The weights can be updated using an Adam optimizer or other gradient descent optimizer. The network is then saved as a model (block 1022). In some embodiments, the model is saved during a particular training step. The saved model is evaluated using a validation dataset, an independent subset of samples not included in the training process. The scan label model training method 1000 then determines whether the validation loss and error rate have stopped decreasing or a predetermined number of training steps have been reached, whichever occurs first (decision block 1024). If not, the scan label model training method 1000 is re-run from block 1012 (i.e., the next iteration of the network) using the network with updated weights. Once the validation loss and error rate stop decreasing or a predetermined number of training steps have been performed, the saved model is evaluated (block 1026). The best trained models are then selected based on the minimum validation loss or error rate from the trained models, which can then be utilized by the CTC decoder and segmentation-based calling system 540.

[0072] In some embodiments, the scan label model training method 1000 can be used to generate two scan label models / neural networks: one model for the pure base category of the data, and a second model for the mixed base category of the data.

[0073] Embodiments of the present invention can also be trained to call mixed bases, e.g., two, three, or four bases at a single position. However, training data from diploid organisms, such as human samples, with two mixed bases per position is generally more common than training data from samples with more than two bases per position, such as some bacterial samples. Mixed base calling is more difficult than pure base calling because peaks at mixed-base positions (e.g., two bases at a single position) often do not align exactly over one another. Typically, two peaks may be slightly offset from each other. Furthermore, in Sanger sequencing, peak heights are often not uniform, so two peaks may have different peak heights, or even significantly different peak heights.

[0074] In some embodiments, data augmentation techniques such as adding noise, spikes, dye blobs, or other data artifacts, or simulated sequencing traces, can be used to improve the robustness of the model. Other techniques such as dropout or weight decay can be used during training to improve the generality of the model. Generative adversarial nets (GANs) can be used to implement these techniques.

[0075] Transfer learning for customized or application-specific models Transfer learning has been successfully used to reuse existing neural models in image classification, object recognition, translation, speech synthesis, and many other fields. Using transfer learning, some embodiments of the present invention allow networks already trained for general pure or mixed basecalling to be retrained for customized, application-specific models. Generic models learned from existing training datasets can be reused. Furthermore, only the final 1x1 convolutional reduction layer can be retrained with additional customer or application data to generate specific models for different customers and applications. Because trained features stored in previous layers are reused and only the weights in the final layer are updated, much less customer or application training data is required for training. Transfer learning, as used in some embodiments of the present invention, allows customers to leverage annotated data to optimize general deep basecalling neural networks to improve the basecalling performance of specific applications.

[0076] Because only the weights of the final 1x1 convolution reduction layer need to be retrained on the customer dataset, and the number of weights in the final layer ranges from a few hundred to a few thousand, a customer- or application-specific dataset of a few thousand examples may be sufficient, which is much less than the number of examples used to train a baseline model, which can be in the hundreds of thousands.

[0077] The process a user goes through to retrain a model is similar: First, select annotated training, validation, and test datasets. Next, train a model using the training dataset and monitor the training. Next, select the best trained model using the validation dataset and test the selected model on the test dataset. However, because training does not start from scratch but from a common trained model, a much smaller number of samples are required and the training time (perhaps a few minutes) is much shorter compared to the training time required to train a baseline model.

[0078] Base call quality value (QV) The quality value model 560 in Figure 5 receives the assembled scan window, base calls, and scan labels for peak scan label probabilities. The quality value model 560 then generates estimated base calling error probabilities. The estimated base calling error probabilities can be converted to Phred-style quality scores by the following formula: QV=-10xlog(probability of error).

[0079] As another example, Phred (Ewing & Green, 1998) proposed that the Phred Basecaller estimates the probability of an error or quality value for each base call, a negative log-transformed error probability, using a function of certain parameters calculated from the trace data (see Brent Ewing and Phil Green, "Base-Calling of Automated Sequencer Traces Using Phred. II. Error Probabilities," Genome Res. 1998 8:186-194).

[0080] A similar strategy has been applied to gene sequence analysis software, such as KB Basecaller, manufactured by the Applied Biosystems division of Thermo Fisher Scientific, Inc., to calculate the QV of each base call (see Labrenz, James, Sorenson, Jon M., and Gehman, Curtis. Methods and systems for the analysis of biological sequence data, WO 2004 / 113557 A2, 2004-12-29). However, compared to the original Phred basecaller, different parameters calculated from trace data are used in the KB basecaller's QV calculation. Similarly, the deep learning basecaller described herein as an embodiment of the present invention also calculates a quality value to provide an estimate of the confidence of every called base. Unlike the Phred and KB basecallers described above, the parameters or functions used in the QV calculation in embodiments of the present invention are based on CTC scan labeling probabilities rather than trace data. Specifically, a feature vector with the four parameters or features listed below is calculated from a local window of CTC scan labeling probabilities around the base call scan location for each base call.

[0081] (1) CTC scan labeling probability: CTC scan labeling probability of the called base at the base call scan position (2) Noise-to-signal ratio: the ratio of the maximum scan label probability from an uncalled base, or the noise scan label probability within a local window, to the scan label probability of a called base at the base call scan position. (3) Base call interval ratio: the ratio of base intervals between a base call and its adjacent base call (4) Resolution: The ratio of the local base spacing to the width of the scan label probability peak of the called base.

[0082] 11 shows a method 1100 for constructing a trained quality value lookup table according to one embodiment of the present invention. As shown in step 1110, a large annotated dataset is required in QV training to generate one or more QV lookup tables.

[0083] In one embodiment of the present invention, in step 1120, a convolutional neural network trained using the method shown in FIG. 6, a decoder using the method shown in FIG. 8, and a base call position detector using the method shown in FIG. 9 are used to calculate the CTC scan labeling probability, base calls, and base call scan position for each base call in the training dataset. Whether a base call in this QV training dataset is considered correct depends on the alignment between the called sequence and the annotated sequence for each sample. In step 1130, all base calls in the training dataset can be assigned to one of two categories: correct base calls and incorrect base calls. Base calls can be characterized by a feature vector p with the four features described above. The feature vector for each base call is calculated in step 1140. All features must be positively and monotonically related to the probability of error. In step 1150, the base calls used for QV training are grouped into many cuts that equalize the histogram of each feature. In step 1150, an empirical error rate is also calculated for each cut. In step 1160, a lookup table is constructed. The cut with the lowest error rate is added to the lookup table first. The feature vector defining the cut and the QV(p) corresponding to the error rate of that cut are i ,q i ) is used to add a new row to the lookup table for that scene. When a scene is added to the lookup table, the calls contained in that scene are also removed from all remaining scenes. This process is repeated until all scenes have been added to the QV lookup table. The QV lookup table is now complete.

[0084] Multiple trained QV lookup tables can then be used to assign a QV to each base call. Embodiments of the invention utilize three separate trained QV tables: one for pure bases in the pure base data category, one for pure bases in the mixed base data category (i.e., samples that are almost entirely pure bases with occasional mixed bases), and one for mixed bases in the mixed base data category. In some embodiments of the invention, QV lookup table training may be performed twice: once using the pure base data set to create the pure base data category QV table, and a second time using the mixed base data set to create the pure bases in the mixed base data category QV table and the mixed bases in the mixed base category QV table.

[0085] For a called base, a feature vector p for that base call is computed. The feature vector is then used as a query key to search the lookup table row by row, in order, until a row is found that has all feature values ​​greater than or equal to the corresponding feature for that base call. The QV associated with that line is then assigned to that base call. Base calls for which no row is found are assigned QV=0.

[0086] Exemplary Computing Device Embodiments 12 is an exemplary block diagram of a computing device 1200 that can incorporate embodiments of the present invention. FIG. 12 is merely an example of a mechanical system for performing aspects of the technical processes described herein and is not intended to limit the scope of the claims. Those skilled in the art will recognize other variations, modifications, and alternatives. In one embodiment, the computing device 1200 typically includes a monitor or graphical user interface 1202, a data processing system 1220, a communication network interface 1212, input device(s) 1208, output device(s) 1206, etc.

[0087] 12, data processing system 1220 may include one or more processor(s) 1204 that communicate with several peripheral devices via a bus subsystem 1218. These peripheral devices may include input device(s) 1208, output device(s) 1206, a communication network interface 1212, and storage subsystems such as volatile memory 1210 and non-volatile memory 1214. Volatile memory 1210 and / or non-volatile memory 1214 may store computer-executable instructions that, when applied to and executed by processor(s) 1204, thus form logic 1222 that implements embodiments of the processes disclosed herein.

[0088] Input device(s) 1208 include devices and mechanisms for inputting information into data processing system 1220. These may include keyboards, keypads, touchscreens integrated into monitor or graphical user interface 1202, voice recognition systems, audio input devices such as microphones, and other types of input devices. In various embodiments, input device(s) 1208 may be embodied as a computer mouse, trackball, trackpad, joystick, wireless remote, drawing tablet, voice command system, eye tracking system, etc. Input device(s) 1208 typically allow a user to select objects, icons, control areas, text, etc. displayed on monitor or graphical user interface 1202 via commands such as clicking a button.

[0089] Output device(s) 1206 include devices and mechanisms for outputting information from data processing system 1220. These may include a monitor or graphical user interface 1202, speakers, printers, infrared LEDs, etc., as is well understood in the art.

[0090] The communications network interface 1212 provides an interface to communications networks (e.g., communications network 1216) and devices external to the data processing system 1220. The communications network interface 1212 may serve as an interface for receiving data from other systems and transmitting data to other systems. Embodiments of the communications network interface 1212 may include an Ethernet interface, a modem (telephone, satellite, cable, ISDN), an (asynchronous) digital subscriber line (DSL), a wireless communications interface such as FireWire, USB, Bluetooth, or WiFi, a short-range communications wireless interface, a cellular interface, etc. The communications network interface 1212 may be coupled to the communications network 1216 via an antenna, cable, etc. In some embodiments, the communications network interface 1212 may be physically integrated onto a circuit board of the data processing system 1220, or in some cases, may be implemented in software or firmware, such as a “soft modem.” The computing device 1200 may include logic to enable communication over a network using protocols such as HTTP, TCP / IP, RTP / RTSP, IPX, UDP, etc.

[0091] Volatile memory 1210 and nonvolatile memory 1214 are examples of tangible media configured to store computer-readable data and instructions that form logic for implementing aspects of the processes described herein. Other types of tangible media include removable memory (e.g., plug-in USB memory devices, mobile device SIM cards), optical storage media such as CD-ROMs and DVDs, semiconductor memory such as flash memory, non-transitory read-only memory (ROM), battery-backed volatile memory, networked storage devices, etc. Volatile memory 1210 and nonvolatile memory 1214 may be configured to store the basic programming and data constructs that provide the functionality of the disclosed processes and other embodiments within the scope of the present invention. Logic 1222 that implements embodiments of the present invention may be formed by volatile memory 1210 and / or nonvolatile memory 1214 storing computer-readable instructions. Such instructions may be read from volatile memory 1210 and / or nonvolatile memory 1214 and executed by processor(s) 1204. Volatile memory 1210 and nonvolatile memory 1214 may also provide repositories for storing data used by logic 1222. Volatile memory 1210 and nonvolatile memory 1214 may include several memories, including main random access memory (RAM) for storing instructions and data during program execution and read-only memory (ROM) in which read-only, non-transient instructions are stored. Volatile memory 1210 and nonvolatile memory 1214 may include a file storage subsystem that provides persistent (non-volatile) storage for program and data files. Volatile memory 1210 and non-volatile memory 1214 may include removable storage systems, such as removable flash memory.

[0092] Bus subsystem 1218 provides a mechanism for allowing the various components and subsystems of data processing system 1220 to communicate with each other as intended. Although communications network interface 1212 is shown schematically as a single bus, some embodiments of bus subsystem 1218 may utilize multiple separate buses.

[0093] It will be readily apparent to one skilled in the art that computing device 1200 may be a device such as a smartphone, a desktop computer, a laptop computer, a rack-mounted computer system, a computer server, or a tablet computer device. As is generally known in the art, computing device 1200 may be implemented as a collection of multiple networked computing devices. Furthermore, computing device 1200 will typically include operating system logic (not shown), the type and nature of which are well known in the art.

[0094] An embodiment of the present invention includes a system, a method, and a non-transitory computer-readable storage medium(s) tangibly storing computer program logic that can be executed by a computer processor.

[0095] Those skilled in the art will appreciate that computer system 1200 represents just one example of a system in which a computer program product according to an embodiment of the present invention may be implemented. As an example of an alternative embodiment, execution of instructions included in a computer program product according to an embodiment of the present invention may be distributed across multiple computers, such as, for example, computers in a distributed computing network.

[0096] While the present invention has been particularly described with reference to exemplary embodiments, it is intended that various changes, modifications, and adaptations may be made based on this disclosure and are within the scope of the present invention. While the present invention has been described in connection with what are presently considered to be the most practical and preferred embodiments, it is understood that the invention is not limited to the disclosed embodiments, but on the contrary, the various embodiments referenced above and below are intended to cover various modifications and equivalent arrangements that fall within the scope of the basic principles underlying the invention as described.

[0097] term Terms used herein with reference to the embodiments of the invention disclosed herein should be given their ordinary meaning by those of ordinary skill in the art, unless otherwise indicated explicitly or by context.

[0098] "Quality value" in this context refers to an estimate (or prediction) of the likelihood that a given base call will be in error. Typically, quality values ​​are scaled according to rules established by the Phred program: QV = -10 log10(Pe), where Pe represents the estimated probability that the call is in error. See Brent Ewing and Phil Green, "Base-Calling of Automated Sequencer Traces Using Phred. II. Error Probabilities," Genome Res. 1998 8:186-194. Quality values ​​are a measure of the reliability of base-calling and consensus-calling algorithms. The higher the value, the lower the likelihood of algorithmic error. Sample quality values ​​refer to the quality value per base of the sample, and consensus quality values ​​are the quality value per consensus.

[0099] "Sigmoid function" in this context refers to a function of the form f(x)=1 / (exp(-x)). Sigmoid functions are used as activation functions in artificial neural networks. They have the property of mapping a wide range of input values ​​to the range 0 to 1, and sometimes -1 to 1.

[0100] In this context, "capillary electrophoresis genetic analyzer" or "capillary electrophoresis DNA analyzer" refers to an instrument that applies an electric field to a capillary filled with a biological sample, causing negatively charged DNA fragments to migrate toward a positive electrode. The speed at which DNA fragments migrate through the medium is inversely proportional to their molecular weight. This process of electrophoresis can separate extension products by size with single-base resolution.

[0101] "Image signal" in this context refers to the fluorescence intensity reading from one of the dyes used to identify bases during a data run. In one embodiment of the present invention, the signal intensity numbers are shown in the annotation view of the sample file.

[0102] "Exemplary commercially available CE devices" in this context include, among others, Applied Biosystems, Inc. (ABI) genetic analyzer models 310 (single capillary), 3130 (4 capillaries), 3130xL (16 capillaries), 3500 (8 capillaries), 3500xL (24 capillaries), and SeqStudio genetic analyzer model, DNA analyzer models 3730 (48 capillaries), and 3730xL (96 capillaries), as well as the Agilent 7100 device, Prince Technologies, Inc.'s PrinceCE™ Capillary Electrophoresis System, Lumex, Inc.'s Capel-105™ CE system, and Beckman Coulter's P / ACE™ MDQ system.

[0103] "Base pair" in this context refers to complementary nucleotides in a DNA sequence: thymine (T) is complementary to adenine (A), and guanine (G) is complementary to cytosine (C).

[0104] "ReLU" in this context refers to a rectified linear activation function unit, which is a piecewise linear function that outputs a direct input if it is positive, and zero otherwise. This is also known as a ramp function and is similar to a half-wave rectifier in electrical signal theory. ReLU is a common activation function in deep neural networks.

[0105] A "heterozygous insertion-deletion variant" (or "het indel") in this context refers to a polymorphism in which one copy of a DNA sequence has an insertion or deletion relative to the other copy that is sequenced together at the same time. The result of sequencing a het indel is the presence of two peaks (also called mixed bases) at most positions downstream of the heterozygous insertion or deletion.

[0106] "Mobility shift" in this context refers to the change in electrophoretic mobility imposed by the presence of different fluorophores associated with different labeling reaction extension products.

[0107] "Variant" in this context refers to bases where the consensus sequence differs from the provided reference sequence.

[0108] "Polymerase slippage" in this context results in the presence of a small peak 3' to the homopolymer. The polymerase may "slip" when sequencing a long homopolymer stretch, skipping one or more bases within the homopolymer. This creates truncated products that differ in length by one to a few bases, appearing as small peaks within and downstream of the homopolymer.

[0109] "Amplicon" in this context refers to the product of a PCR reaction. An amplicon is typically a short fragment of DNA.

[0110] "Base calling" in this context refers to the assignment of a nucleotide base to each peak in the fluorescent signal (in IUPAC-IUB notation, A, C, G, T). Base calls can be mixed, with two peaks per position (in IUPAC-IUB notation, R=A and G, Y=C and T, S=G and C, W=A and T, K=G and T, and M=A and C) or three peaks per position (in IUPAC-IUB notation, B=C, G, and T, D=A, G, and T, H=A, C, and T, and V=A, C, and G).

[0111] "Raw data" or "input analysis trace" in this context refers to the multicolor graph displaying the collected fluorescence intensities (signals) for each of the four fluorescent dyes, and / or the data used to populate or create such a graph.

[0112] "Base spacing" in this context refers to the number of data points from one peak to the next. Negative spacing values ​​or spacing values ​​displayed in red indicate that the base caller used the default spacing value rather than one calculated based on the current data.

[0113] "Separation or sieving media" in this context refers to non-gel liquid polymers such as linear polyacrylamide, hydroxyalkyl cellulose (HEC), agarose, and cellulose acetate, which can be used. Other separation media that can be used in capillary electrophoresis include, but are not limited to, poly(N,N'-dimethylacrylamide) (PDMA), polyethylene glycol (PEG), poly(vinylpyrrolidone) (PVP), polyethylene oxide, polysaccharides, and water-soluble polymers such as pluronic polyols, various polyvinyl alcohol (PVAL)-based polymers, polyether-water mixtures, and lyotropic polymer liquid crystals, among others.

[0114] "Adam optimizer" in this context refers to an optimization algorithm that can be used instead of classical stochastic gradient descent, which iteratively updates network weights based on training data. Stochastic gradient descent maintains a single learning rate (called alpha) for all weight updates, and the learning rate does not change during training. A learning rate is maintained for each network weight (parameter) and adapted individually as learning progresses. The Adam optimizer combines the benefits of two other extensions to stochastic gradient descent: the adaptive gradient algorithm (AdaGrad), which maintains a per-parameter learning rate that improves performance for sparse gradient problems (e.g., natural language and computer vision problems), and root-mean-square propagation (RMSProp), which also maintains a per-parameter learning rate that is adapted based on the average of the recent magnitude of the weight gradient (e.g., how quickly it is changing). This means the algorithm works well online and for non-stationary problems. Adam recognizes the benefits of both AdaGrad and RMSProp. Instead of adapting the parameter learning rate based on the mean first moment (average) as in RMSProp, Adam also utilizes the average of the second moments of the gradients (uncentered variance). Specifically, the algorithm computes exponential moving averages of the gradients and squared gradients, and the parameters beta1 and beta2 control the decay rates of these moving averages. The initial values ​​of the moving averages and values ​​of beta1 and beta2 close to 1.0 (recommended) bias the moment estimates towards zero. This bias is overcome by first computing biased estimates and then bias-corrected estimates.

[0115] "Hyperbolic tangent function" in this context refers to a function of the form tanh(x) = sinh(x) / cosh(x). The tanh function is a common activation function in artificial neural networks. Like the sigmoid, the tanh function is also sigmoidal ("s" shaped), but instead outputs values ​​in the range (-1, 1). Thus, strongly negative inputs to tanh are mapped to negative outputs. Additionally, only zero-valued inputs are mapped to near-zero outputs. These properties make it less likely that a network will become "stuck" during training.

[0116] "Relative fluorescence units" in this context refers to measurements in electrophoretic methods, such as capillary electrophoresis for DNA sequence analysis. "Relative fluorescence units" are the units of measurement used in analyses that use fluorescence detection.

[0117] In this context, "CTC loss function" refers to a connectionist temporal classification, neural network output type, and associated scoring function for training recurrent neural networks (RNNs) such as LSTM networks, temporal convolutional networks (TCNs), and extended causal or acausal convolutional networks to address the problem of variable-timing sequences. CTC networks have continuous outputs (e.g., Softmax) and are adapted through training to model the probability of labels. CTC does not attempt to learn boundaries and timing. Label sequences are considered equivalent if they differ only in their placement, ignoring whitespace. Equivalent label sequences can arise in a variety of ways, making scoring a nontrivial task. Fortunately, scoring equivalent label sequences can be completed using a forward-backward algorithm. CTC scores can then be used with a backpropagation algorithm to update the neural network weights. Alternative approaches to CTC-adapted neural networks include hidden Markov models (HMMs).

[0118] "Polymerase" in this context refers to an enzyme that catalyzes polymerization. DNA and RNA polymerases build single-stranded DNA or RNA (respectively) from free nucleotides using another single-stranded DNA or RNA as a template.

[0119] "Sample data" in this context refers to the output of a single lane or capillary of a sequencing instrument. Sample data can be input into Sequencing Analysis, SeqScape, and other sequence analysis software manufactured by Applied Biosystems, Inc. and other manufacturers.

[0120] In this context, "plasmid" refers to a genetic structure within a cell that can replicate independently of chromosomes, typically a small circular DNA molecule in the cytoplasm of a bacterium or protozoan. Plasmids are commonly used for genetic manipulation in the laboratory.

[0121] "Beam search" in this context refers to a heuristic search algorithm that explores a graph by expanding the most promising nodes from a limited set. Beam search is an optimization of best-first search that reduces memory requirements. Best-first search is a graph search that orders all partial solutions (states) according to some heuristic. However, in beam search, only a predetermined number of the best partial solutions are retained as candidates. It is therefore a greedy method. Beam search uses breadth-first search to build a search tree. At each level of the tree, all successors of the states at the current level are generated and sorted in ascending order of heuristic cost. However, only a predetermined number K of the best states at each level are stored (called the "beamwidth"). Then, only those states are expanded. The larger the beamwidth, the fewer states are pruned. If the beamwidth is infinite, no states are pruned and beam search is the same as breadth-first search. The beamwidth limits the memory required to perform the search. Since the goal state can potentially be pruned, beam search sacrifices completeness (guarantee that the algorithm will terminate with a solution if one exists). Beam search is not optimal (i.e., there is no guarantee that the best solution will be found). Generally, beam search returns the first solution found. Beam search in machine translation is a different case. Once a configured maximum search depth (i.e., translation length) is reached, the algorithm evaluates solutions found during the search at various depths and returns the best one (the one with the highest probability). The beam width can be either fixed or variable. One approach using a variable beam width is to start with the smallest width. If no solution is found, the beam is widened and the procedure is repeated.

[0122] "Sanger sequencing" in this context refers to the DNA sequencing process that utilizes the ability of DNA polymerase to incorporate 2',3'-dideoxynucleotides—nucleotide base analogs lacking the 3'-hydroxyl group essential for phosphodiester bond formation. As originally designed, Sanger dideoxy sequencing required a DNA template, sequencing primers, DNA polymerase, deoxynucleotides (dNTPs), dideoxynucleotides (ddNTPs), and a reaction buffer. Four separate reactions were set up, each containing radiolabeled nucleotides: ddA, ddC, ddG, or ddT. The annealing, labeling, and termination steps were performed in separate heat blocks. DNA synthesis was performed at 37°C, the temperature at which DNA polymerase has optimal enzymatic activity. DNA polymerase adds either deoxynucleotides or the corresponding 2',3'-dideoxynucleotides at each step of chain elongation. Whether deoxynucleotides or dideoxynucleotides are added depends on the relative concentrations of both molecules. Addition of a deoxynucleotide (A, C, G, or T) to the 3' end allows chain elongation to continue. However, addition of a dideoxynucleotide (ddA, ddC, ddG, or ddT) to the 3' end terminates chain elongation. Sanger dideoxy sequencing results in the formation of extension products of various lengths that terminate at the 3' end with a dideoxynucleotide.

[0123] A "single nucleotide polymorphism" in this context refers to a variation in a single base pair in a DNA sequence.

[0124] "Mixed base" in this context refers to a base position that contains 2, 3, or 4 bases. These bases are assigned the appropriate IUB code.

[0125] A "softmax function" in this context refers to a function of the form f(xi) = exp(xi) / sum(exp(x)), where the sum is taken over the set of x. Softmax is used in different layers of an artificial neural network (often the output layer) to predict the classification of the input to those layers. The softmax function calculates the probability distribution of an event xi over "n" different events. In a general sense, this function calculates the probability of each target class over all possible target classes. The calculated probability helps predict which target class is represented in the input. The main advantage of using softmax is the range of the output probability. It ranges from 0 to 1, and the sum of all probabilities equals 1. For softmax functions used in multiclassification models, the probability of each class is returned, with the probability of the target class being higher. This formula calculates the exponent (e-power) of a given input value and the sum of the exponent values ​​for all values ​​in the input. The output of the softmax function is then the ratio of the exponent of the input value to the sum of the exponent values.

[0126] "Noise" in this context refers to the average background fluorescence intensity of each dye.

[0127] "Backpropagation" refers to an algorithm used in artificial neural networks to calculate the gradients needed to calculate the weights used in the network. It is commonly used to train deep neural networks, a term that refers to neural networks with two or more hidden layers. In backpropagation, the loss function calculates the difference between the network output and its expected output after a case has propagated through the network.

[0128] "Dequeue max finder" in this context refers to an algorithm that utilizes a double-ended queue to determine the maximum value.

[0129] "Pure base" in this context refers to a single base position that contains only one base or nucleotide (A, C, G, and T). These bases are assigned the appropriate IUPAC-IUB code.

[0130] "Primer" in this context refers to a short single strand of DNA that serves as a priming site for DNA polymerase in a PCR reaction.

[0131] A "loss function" in this context (sometimes called a cost function or error function) is a function that maps the values ​​of one or more variables to real numbers that intuitively represent some "cost" associated with those values.

Claims

1. 1. A method for automatically sequencing one or more deoxyribonucleic acid (DNA) molecules of a biological sample, comprising: measuring the biological sample using a capillary electrophoresis (CE) genetic analyzer to obtain an input trace comprising digital data corresponding to fluorescence values ​​comprising multiple scans of the biological sample; b. generating scan label probabilities for the plurality of scans using a trained artificial neural network including a plurality of layers, including a convolutional layer; c. determining a base call sequence comprising a plurality of base calls for the one or more DNA molecules based on the scan labeling probabilities for each of the plurality of scans; d. determining a scan number position for each of the plurality of base calls; e. Displaying the base call sequences on an electronic display; f. displaying on the electronic display a base call position representation for each of the plurality of base calls using the scan number position, the base call position representation visually indicating the relative spacing between adjacent base calls in the base call sequence.

2. 2. The method of claim 1, further comprising displaying the input trace on the electronic display such that an axis of the input trace corresponding to a relative scan number position of the fluorescence value of the input trace is aligned with an axis for displaying the base call position representation corresponding to a relative scan number position.

3. 3. The method of claim 2, wherein the axis of the input trace is the same axis as the axis for displaying the base call position representation.

4. The method of claim 1 , wherein the plurality of layers comprises a plurality of residual blocks, each residual block of the plurality of residual blocks comprising one or more non-causal convolutional layers.

5. The method of claim 4 , wherein a residual block of the plurality of residual blocks further comprises a skip connection.

6. The method of claim 4 , wherein a residual block of the plurality of residual blocks further comprises at least one spatial dropout layer following a non-causal convolutional layer.

7. 6. The method of claim 5, wherein a residual block of the plurality of residual blocks further comprises a 1x1 convolutional layer between an input and an output of the skip connection.

8. The method of claim 4 , wherein a residual block of the plurality of residual blocks further comprises at least one normalization layer following a non-causal convolutional layer.

9. The method of claim 4 , wherein a residual block of the plurality of residual blocks further comprises at least one modified linear activation function layer following a non-causal convolutional layer.

10. 5. The method of claim 4, wherein the plurality of residual blocks includes at least a first residual block including one or more acausal convolutional layers having a first expansion coefficient and a second residual block including one or more acausal convolutional layers having a second expansion coefficient different from the first expansion coefficient.

11. 2. The method of claim 1, wherein the trained artificial neural network is trained to minimize the loss between the scan labeling probabilities and the target sequence of bases using a connectionist temporal classification (CTC) loss function.

12. 10. The method of claim 1, wherein the trained artificial neural network further comprises a 1x1 convolutional reduction layer for reducing the number of extracted features to match the number of output labels.

13. The method of claim 1 , wherein the trained artificial neural network further comprises a softmax layer to obtain the scan label probabilities.

14. 2. The method of claim 1, wherein determining the base call sequence further comprises decoding the scan label probabilities using a prefix beam search.

15. decoding the scan signature probabilities using the prefix beam search; a. initializing an empty base call sequence as a prefix; b. In each scan t of the plurality of scans, i. extending the prefix with each of a plurality of extension indicators; ii. Scoring each prefix by incorporating the scan label probability of the extended label in the scan t; iii. Saving an expanded candidate subset containing the K highest scoring prefixes, wherein the subset does not exceed a beamwidth of size K; iv. if the prefix is ​​different from the highest-scoring prefix in the previous scan t-1, saving the highest-scoring prefix in the scan in a candidate subset; v. if the prefix is ​​different from the highest-scoring prefix in the previous scan t-1, assigning the scan to the highest-scoring prefix; vi. Returning the highest scoring prefix in the last scan as the final base call sequence; vii. returning the candidate subset of top candidates saved for each scan during the prefix beam search.

16. 16. The method of claim 15, wherein the plurality of extended labels comprises pure base labels, mixed base labels, and blank labels.

17. 16. The method of claim 15, further comprising finding a scan range for each base call and then using the scan range to find a scan position within the scan range that has a peak labeling probability.

18. using the peak labeling probability within the scan range of each base call to find the scan range and the scan position; a. starting from the first base call of the final base call sequence y; b. each base call y of the plurality of base calls in the final base call sequence y i In i. Within the candidate subset, the first i base calls of the base call sequence y define a base call subsequence y 1..i Search for and ii. The base call y is assigned to the found candidate. i setting a start scan of the scan range; iii. Next base call y i+1 The base call y is obtained by extending the starting scan using a prefix scan number until the starting scan i setting an end scan of the scan range; iv. Using the peak labeling probability, the base call y is calculated between the start scan and the end scan. i selecting a scan position for c) returning the start scan and the end scan and the scan positions of all base calls in the final base call sequence.

19. 2. The method of claim 1, further comprising: for a base call among the plurality of base calls, determining a quality value for each of the plurality of base calls by using a plurality of feature values ​​derived from scan labeling probabilities corresponding to scans within a scan range that includes the scan position of the base call, wherein the plurality of feature values ​​include a peak scan labeling probability, a signal-to-noise ratio, a base call spacing ratio, and a resolution value for the base call.

20. The method of claim 19 , further comprising using a machine learning algorithm to obtain the quality value using the plurality of feature values.

21. determining a quality value for each base call, wherein predicting the quality value further comprises: determining a feature vector for the base call, the feature vector comprising a plurality of feature values ​​comprising a scan label probability of the base call at a base call scan position, a signal to noise ratio, a base call spacing ratio, and a resolution value; b. finding a row with a minimum cut index in a quality value lookup table comprising a plurality of rows, each row having (1) a feature vector assigned to a cut comprising a plurality of base calls, and (2) a quality value corresponding to an empirical error rate for the cut; traversing the quality value lookup table to assign to the base call a quality value corresponding to the row with the minimum cut index, wherein the row with the minimum cut index comprises rows having feature vectors with all feature values ​​greater than or equal to the feature vector of the base call, or assigning a quality value of zero if no row with the minimum cut index is found.

22. The quality value lookup table is a. Initializing a quality value lookup table; b. Computing a feature vector for each base call in a quality value training dataset comprising a plurality of samples; c. grouping the base calls into cuts, each cut equalizing a histogram of the feature vectors, until all remaining cuts have been added to the quality value lookup table; d. Calculating an empirical error rate for each of the one or more cuts; e. adding the cut with the lowest empirical error rate to the quality value lookup table as a next new row containing the feature vector assigned to the cut and a quality value corresponding to the empirical error rate of the cut; f. removing the cut that was added to the quality value lookup table from the plurality of cuts; g. removing from the remaining cut all base calls of the cut that were added to the quality value lookup table; h) The method of claim 21, comprising repeating steps (c) through (g) until no cuts remain.

23. 22. The method of claim 19 or 21, wherein the noise-to-signal ratio comprises the ratio of (1) the maximum scan label probability from one or more uncalled bases, or noise scan label probability within a local scan window of the base call, to (2) the scan label probability of a called base at the scan position of the base call.

24. 22. The method of claim 19 or 21, wherein the base call spacing ratio comprises a ratio of a first base spacing value between the base call and a first adjacent base call to a second base spacing value between the base call and a second adjacent base call.

25. 22. The method of claim 19 or 21, wherein the resolution value comprises a ratio of a local base spacing value to a width value of a scan label probability peak of the base call.

26. 10. The method of claim 1, further comprising displaying the base call sequence and input analysis trace as an electropherogram on a display of a computing device.

27. 4. The method of claim 1, further comprising displaying on the electronic display a visual representation of a quality value for each of the plurality of base calls.

28. 4. The method of claim 1, further comprising displaying a visual representation of a quality value for each of the plurality of base calls, wherein the visual representation of the quality values ​​is also used as the base call position representation by placing the visual representation of the quality value in a position on the electronic display that represents the position of the base call.

29. When executed by one or more processors of at least one computing device, acquiring an input trace comprising digital data corresponding to fluorescence values ​​in multiple scans of a biological sample performed by a capillary electrophoresis genetic analyzer; b. generating scan label probabilities for the plurality of scans using a trained artificial neural network including a plurality of layers, including a convolutional layer; c. determining a base call sequence comprising a plurality of base calls for the one or more deoxyribonucleic acid (DNA) molecules based on the scan labeling probabilities for each of the plurality of scans; d. determining a scan number position for each of the plurality of base calls; e. Displaying the base call sequences on an electronic display; f. displaying on the electronic display a base call position representation of each of the plurality of base calls using the scan number positions to visually indicate the relative spacing between adjacent base calls in the base call sequence.

30. 1. A method for determining scan positions of base calls in a base call sequence obtained using scan labeling probabilities for each of a plurality of scans performed by a capillary electrophoresis device on one or more deoxyribonucleic acid (DNA) molecules of a biological sample, comprising: a. determining a scan range for each base call scan range probability, the scan range including scans ranging from the first scan corresponding to the base call to the last scan corresponding to the base call; b. searching within said scan range to determine peak scan signature probabilities within said scan range; c) using the scan positions of said peak scan labeling probabilities to display an indication of the positions of said base calls on an electronic display.

31. Determining the scan range and the scan position having the peak scan labeling probability within the scan range for each base call comprises: a. starting from the first base call of the final base call sequence y; b. Each base call y in the final base call sequence y i In i. Within the candidate subset, the first i base calls of the base call sequence y define a base call subsequence y 1..i Search for and ii. The base call y is assigned to the found candidate. i setting a start scan of the scan range; iii. Next base call y i+1 The base call y is obtained by extending the starting scan using a prefix scan number until the starting scan i setting an end scan of the scan range; iv. Using the peak scan labeling probability, the base call y i selecting a scan position for c. returning the start scan and the end scan and the scan positions of all base calls in the final base call sequence.

Citation Information

Patent Citations

  • Automated quality control and spectral error correction for sample analysis instruments

    JP2020510822A

  • Methods and systems for de novo peptide sequencing using deep learning

    US20190018019A1

  • Machine learning enabled pulse and base calling for sequencing devices

    US20190237160A1

  • Machine learning enabled biological polymer assembly

    US20190348152A1

  • Image driven quality control for array-based PCR

    US20200074303A1