Base sequence determining apparatus, capillary array electrophoresis apparatus and method
By performing mobility correction and peak extraction in a capillary array electrophoresis apparatus, combined with peak spacing constrained deconvolution processing, the problem of inaccurate base sequence determination caused by low peak separation was solved, and high-precision base sequence determination was achieved.
Patent Information
- Application Number
- CN201680079412.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2016-01-28
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2036-01-28
AI Technical Summary
Existing technologies struggle to accurately determine base sequences when peak separation is low, especially when base types overlap, making it difficult to accurately add peak types.
A deconvolution part is used for mobility correction and peak extraction. The base sequence is determined by deconvolution processing. The mobility correction signal is then corrected and peaks are extracted using the deconvolution part. Combined with peak spacing constraint deconvolution processing, candidate solutions are reduced to avoid local solutions and improve accuracy.
Even with low peak resolution, it can determine base sequences with high accuracy, making it particularly suitable for portable sequencers or those with short sequencing times.
Smart Images

Figure CN108473925B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technique for automatically determining base sequences. Background Technology
[0002] The device used to determine the base sequence that makes up nucleic acids is generally called a DNA (deoxyribonucleic acid) sequencer. Various detection technologies exist in DNA sequencers; the following describes capillary electrophoresis sequencers. Capillary electrophoresis sequencers determine the base sequence by electrophoresis of nucleic acid samples. Using a temporal input signal, the sequencer observes the spectra emitted by pigments corresponding to the four bases (A (adenine), G (guanine), C (cytosine), and T (thymine)) to determine the base sequence. This process of determining the base sequence based on a temporal input signal is called a "temporal base call."
[0003] In sequential base response analysis, peaks corresponding to each base are detected from the signal after color transformation and mobility correction of the input signal, and the sequence is determined based on the order of these peak positions. However, steep peaks are not always observable; there are cases where peak widths extend to the point of overlap, and multiple overlapping peaks may appear as a single peak on the input signal. In such cases, a single peak on the input signal must be decomposed into its original multiple peaks to determine their correct positions; otherwise, the analysis will be inaccurate.
[0004] Patent Document 1 describes a method that can determine the sequence with high accuracy even in the presence of overlapping peaks. The abstract of Patent Document 1 describes a method that includes: "(A) a base peak extraction step that extracts base peaks from electrophoretic data containing peaks of four base types obtained by electrophoresis separation of nucleic acids in a sample; (B) a condition setting step that sets a starting base peak and a peak interval reference value in the time series data composed of the extracted base peaks; (C) in the aforementioned time series data, starting from the starting base peak, scanning sequentially forward and backward between adjacent base peaks, comparing the interval between base peaks with the aforementioned peak interval reference value, and adding interpolated peaks in peak-deficient intervals, thereby determining the base sequence."
[0005] Existing technical documents
[0006] Patent documents
[0007] Patent Document 1: International Publication No. 2008 / 050426 Summary of the Invention
[0008] The problem that the invention aims to solve
[0009] As described above, Patent Document 1 discloses a technique of "adding interpolated peaks to the peak-deficient interval by comparing the interval between base peaks with the aforementioned peak interval reference value." However, even if the interval in which the peak should be added can be determined using the method of adding peaks to the deficient interval, it is impossible to determine which type of base peak should be added. For example, consider... Figure 1 As shown, the waveforms overlap in time between base type A and base type G. If a small peak P1 exists in base type A but cannot be detected due to noise or other reasons, it is impossible to determine at that moment whether peak P1 is added in base type A or base type G.
[0010] Therefore, the present invention provides a technique for determining base sequences with high accuracy even when peak separation is low.
[0011] Methods for solving problems
[0012] To solve the above problems, the present invention employs, for example, the structure described within the scope of the claims. This specification includes several ways to solve the above problems, one example of which is "a base sequence determination apparatus, comprising (1) a mobility correction unit that outputs a mobility correction signal after performing mobility correction on a time-series signal of a spectrum corresponding to each base; (2) a deconvolution unit that performs the following processes: processing to calculate the deconvolutioned signal of the mobility correction signal for multiple parameter candidates of a point spread function, processing to calculate the peak interval discretization of the calculated deconvolutioned signal, processing to determine the parameters of the point spread function using the calculated discretization, and processing to output the deconvolutioned signal corresponding to the point spread function having the determined parameters as an updated deconvolutioned signal; (3) a peak extraction unit that extracts peak waveforms from the updated deconvolutioned signal and outputs an updated peak extraction signal; and (4) a sequence determination unit that inputs the updated peak extraction signal and determines the base sequence."
[0013] The effects of the invention
[0014] According to the present invention, base sequences can be determined with high precision even when peak resolution is low. Other problems, structures, and effects beyond those described above will become clear from the following description of embodiments. Attached Figure Description
[0015] Figure 1 This graph illustrates the case of low peak separation.
[0016] Figure 2 This is a diagram showing the overall structure of the capillary array electrophoresis apparatus of Embodiment 1.
[0017] Figure 3 This is a diagram illustrating the structure of the fluorescence detection device in Example 1.
[0018] Figure 4 This is a diagram illustrating the functional structure of the signal processing unit.
[0019] Figure 5 This is a diagram illustrating the recalculation function (inter-block recalculation function) that uses the standard deviation obtained from calculations on multiple blocks.
[0020] Figure 6 This is a diagram showing an example of a user interface screen.
[0021] Figure 7 This is a flowchart illustrating the processing operations of the peak spacing constrained deconvolution part.
[0022] Figure 8 This is a diagram showing other structural examples of a capillary array electrophoresis apparatus. Detailed Implementation
[0023] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. Furthermore, the embodiments of the present invention are not limited to the examples described below, but various modifications can be made within the scope of its technical concept.
[0024] (1) Example 1
[0025] (1-1) Overall Structure
[0026] Figure 2 This illustrates a structural example of a capillary array electrophoresis apparatus 10. The capillary array electrophoresis apparatus 10 includes: a sample tray 12 containing multiple sample containers 11 (each sample container 11 contains a different sample) of a sample for which a fluorescently labeled sample (hereinafter referred to as a "sample") has been added to the DNA of the analyte; a conveyor 20 for transporting the sample tray 12; a capillary array 30 forming the electrophoresis path of the sample within the sample containers 11; a pump unit 42 for injecting electrophoresis medium 41 into the capillary array 30; a high-voltage power supply 21 applying a high voltage to both ends of the capillary array 30; a thermostat 31 maintaining a fixed temperature within the capillary array 30; detection positions 32 positioned along the electrophoresis path of the sample; a fluorescence detection device 50; and a control substrate 51.
[0027] The capillary array 30 is an assembly of multiple capillaries 33. Each capillary 33 is hollow. In this embodiment, the length of the capillary 33 is 30 cm or less. However, the length of the capillary 33 can also be 30 cm or more (e.g., 36 cm or more). Furthermore, one end of each capillary 33 is inserted into the sample in the sample container 11. The sample moves from the sample container 11 into the capillary 33 and then undergoes electrophoresis within the capillary. The pump unit 42 injects an electrophoresis medium 41 (e.g., a polymer) into each capillary 33. Thus, each capillary 33 is filled with the electrophoresis medium 41.
[0028] A high-voltage power supply 21 applies a high voltage (e.g., a maximum voltage of 20 kV) to both ends of each capillary 33 filled with electrophoresis medium 41. With the application of this high voltage, each sample passes through its respective capillary 33 and is electrophoresed from the sample container 11 to the detection position 32. After passing through the detection position 32, each sample is electrophoresed to the discharge position 34. The samples exhibit different electrophoretic speeds depending on their base lengths, thus arriving at the detection position 32 sequentially from the DNA with the shortest base lengths.
[0029] The fluorescence detection device 50 sequentially irradiates the sample from the point of arrival at the detection position 32 with excitation light, detecting the intensity and wavelength of the fluorescence signal emitted by the fluorescent label. In this embodiment, the fluorescence detection device 50 determines the DNA base sequence based on the detected fluorescence signal wavelength and outputs it to the control substrate 51. Furthermore, the fluorescence detection device 50 displays the determined base sequence and other information on the screen of the display device 52. The control substrate 51 transmits the base sequence of the fluorescence signal analyzed by the fluorescence detection device 50 to an external terminal (not shown). This external terminal preferably has a display device for confirming the analysis results.
[0030] Figure 3 This section illustrates an example of the internal structure of the fluorescence detection device 50. The fluorescence detection device 50 includes an excitation light source 61, a shutter 62, an excitation lens 63, an optical filter 66, a fluorescence lens 67, a diffraction grating 68, a CCD (Charge Coupled Device) 69, a CCD control unit 70, an ADC (Analog to Digital Converter) 71, a signal processing unit 72, and a memory 73.
[0031] Excitation source 61 is a light source that continuously illuminates excitation light 64. All capillaries 33 corresponding to detection position 32 are illuminated by excitation light 64. Shutter 62 repeatedly opens and closes at predetermined intervals. When shutter 62 is open, it allows excitation light 64 illuminating from excitation source 61 to pass through; when it is closed, it blocks the light. Excitation lens 63 focuses the excitation light 64 passing through shutter 62. The excitation light 64 focused by excitation lens 63 illuminates detection position 32. Due to electrophoresis, the fluorescent label added to the DNA at detection position 32 passing through each capillary 33 is excited by the illumination of excitation light 64, emitting a fluorescent signal 65.
[0032] An optical filter 66 (e.g., a color filter) filters out light other than the fluorescent signal 65 emitted from the fluorescent marker. A fluorescent lens 67 focuses the fluorescent signal 65 passing through the optical filter 66. The fluorescent signal 65 focused by the fluorescent lens 67 is split into wavelengths at a diffraction grating 68 and irradiates the light-receiving surface of the CCD 69, which serves as the light-receiving element (light-receiving section). When the fluorescent signal 65, split into wavelengths, is irradiated, a signal charge is generated on the light-receiving surface of the CCD 69. The conversion circuit of the CCD 69 converts the signal charge into a voltage (analog) output to the ADC 71. The CCD 69 can be of any type: frame transmission type, full frame transmission type, interline transmission type, or inter-frame transmission type.
[0033] The ADC71 converts the analog signal output from the CCD69 into a digital signal. The ADC71 outputs the converted digital signal to the signal processing unit 72. The signal processing unit 72 obtains the fluorescence signal intensity (approximately equivalent to the fluorescence signal intensity of the light-receiving surface bound as one pixel) from the output digital signal. In this embodiment, the signal processing unit 72 performs signal processing to determine the DNA base sequence based on the wavelength of the fluorescence signal obtained from the digital signal. Furthermore, the signal processing unit 72 stores the fluorescence signal intensity and wavelength obtained from the output digital signal in the memory 73.
[0034] (1-2) Functional structure of signal processing unit 72
[0035] Figure 4 This describes the functional structure of the signal processing unit 72. In this embodiment, the function of the signal processing unit 72 is implemented by a program executed by a computer. The signal input unit 101 receives an input signal as a digital signal from the ADC 71. As described above, the input signal is a temporal sequence of the spectra emitted from pigments corresponding to the four base types: A (adenine), G (guanine), C (cytosine), and T (thymine). The spectrum at each moment is a sequence of real values after discretizing each wavelength, and the sequence length N is, for example, an integer of 4 or more, such as 10 or 20. Therefore, the input signal is an N-channel real-valued signal.
[0036] The baseline removal unit 102 performs baseline removal processing, obtaining the baseline-removed signal from the input signal (N-channel real-value signal). The baseline removal processing can be performed using known methods such as subtracting the minimum value of the nearest interval at each time point. The color conversion unit 103 performs color conversion processing, calculating the four-channel signals corresponding to each of the four bases, i.e., the color-converted signal, based on the baseline-removed signal (N-channel real-value signal). The color conversion processing can be performed using known methods such as multiplication of the pseudo-inverse matrix of the pre-calibrated color conversion matrix.
[0037] The block segmentation unit 104 performs block segmentation processing, dividing the signals of each channel of the color-transformed signal into small intervals (hereinafter also referred to as "blocks") with the time when the intensity is sufficiently low as the boundary, and outputs the signal corresponding to each block as the segmented signal. Figure 5 This represents an example of block segmentation. For example... Figure 5 As shown, a segmented signal consists of a set of waveforms. By dividing it into blocks, the computational load of subsequent deconvolution processing and peak-space-constrained deconvolution processing can be reduced. This results in shorter processing time and reduced storage consumption.
[0038] The deconvolution unit 105 performs deconvolution processing on the cut signal to calculate the initial deconvolutioned signal. Generally, the difficulty of deconvolution processing differs depending on whether the point spread function (PSF) is known or unknown. When the PSF is known, deconvolution processing uses the input signal (cut signal) and the PSF to calculate the deconvolutioned signal. In this case, deconvolution can be achieved with relatively high accuracy by using methods such as Wiener filtering.
[0039] However, in the capillary array electrophoresis apparatus 10, the PSF sometimes changes. Therefore, the deconvolution unit 105 needs to perform deconvolution under the condition that the PSF is unknown. This process, which calculates both the PSF and the deconvolutioned signal using the input signal, is called "blind deconvolution". For example, such deconvolution with simultaneous optimization can be performed based on Nonnegative Matrix Factorization (NMF). However, there are generally countless candidate pairs of PSF and deconvolutioned signal, and countless local solutions. Therefore, it is generally difficult to obtain a high-precision solution. In NMF, although there are schemes that avoid local solutions and obtain global solutions by sparse regularization such as L1 regularization and L1 / L2 regularization, it is difficult to obtain sufficient accuracy even when used in the capillary array electrophoresis apparatus 10.
[0040] Therefore, in the deconvolution unit 105 of this embodiment, candidate solutions are reduced using the method described below to avoid local solutions. First, the deconvolution unit 105 limits the PSF to a Gaussian function, and limits its unknown parameter only to the standard deviation σ. By limiting the unknown parameter of the PSF in this way, local solutions can be avoided. The deconvolution unit 105 performs deconvolution processing on the corresponding Gaussian function as the PSF for each discrete finite number of standard deviations σ, and obtains the deconvolutioned signal for each standard deviation σ.
[0041] Next, the deconvolution unit 105 focuses the candidate solutions only on the group of standard deviations σ and deconvolutioned signals that satisfy the conditions (1), (2) and (3) shown below.
[0042] (1) The following calculated e(σ) is below a certain threshold.
[0043] First, the deconvolution unit 105 performs convolution processing using a Gaussian function corresponding to the standard deviation σ and the deconvolutioned signal using this function, calculating the convolutioned signal. Next, the deconvolution unit 105 calculates the error e(σ) between the convolutioned signal and the input signal for each standard deviation σ, based on the mean square error and the Kullback-Leibler divergence. Then, it selects only standard deviations σ below a certain threshold for the error e(σ). Based on this condition, it is possible to concentrate only the solutions obtained according to the convolution generation model on the parameter candidates.
[0044] (2) Regarding the intervals of each peak extracted from the deconvolutioned signal, the minimum value of the peak interval within a block is above a certain threshold.
[0045] According to this condition, when candidate solutions with excessively short peak intervals are excluded, local solutions can be avoided.
[0046] (3) σ is near the inflection point σ_r of e(σ).
[0047] The deconvolution unit 105, for example, uses the positive peak of the second-order differential value e""σ" related to the error e(σ) as the inflection point detection, and substitutes the position of this peak into the inflection point σ_r. According to this condition, solutions with an excessive number of peaks compared to the original solution can be eliminated. This is based on the following reasoning: e(σ) corresponding to solutions with an excessive number of peaks compared to the original solution becomes a value that is as small as e(σ) corresponding to the original solution. Furthermore, e(σ) corresponding to solutions with a small number of peaks compared to the original solution becomes a value that is larger than e(σ) corresponding to the original solution. Therefore, it can be said that σ, where e(σ) changes drastically, i.e., the solution corresponding to the inflection point σ_r, is close to the original solution.
[0048] Through the above processing, the deconvolution unit 105 calculates the standard deviation σ and the deconvolutioned signal for each block. Here, as... Figure 5 As shown, the standard deviation of the i-th block is σ_i. It can be assumed that the standard deviation σ will not change significantly in a short period of time. Under this assumption, the deconvolution unit 105 also includes a function to further improve the accuracy of deconvolution (inter-block recalculation function).
[0049] When this function is executed, the deconvolution unit 105 determines whether the difference between the standard deviation σ of a certain block and the average of the standard deviations of the blocks that appear before and after it exceeds the threshold T.
[0050]
[0051] If Equation 1 holds, the deconvolution unit 105 performs a second deconvolution using the PSF obtained through Equation 2 as the standard deviation, and calculates a new deconvolutioned signal. Here, K is a constant of some natural number (1, 2, 3...).
[0052]
[0053] For example in Figure 5 In the case of the i-th block, the standard deviation σ i The value is "1.2", which differs significantly from the average standard deviation of the blocks before and after it, "3.15". Therefore, the deconvolution unit 105 recalculates the signal after deconvolving the i-th block. Furthermore, considering that the recalculation function for these multiple blocks can be performed automatically by the deconvolution unit 105, or, as described later, the operator can specify the execution or non-execution of the function on the user interface screen. Figure 5 The displayed screen can also serve as a prompt to the operator.
[0054] return Figure 4 The peak extraction unit 106 performs peak extraction processing to obtain the peak-extracted signal from the initial deconvolutioned signal. In the peak extraction processing, known methods such as the zero-crossing calculation method for calculating the first-order differential can be used.
[0055] The mobility correction unit 107 performs mobility correction processing, calculating the mobility-corrected signal based on the peak-extracted signal and the color-transformed signal. The mobility correction processing is performed in the following order.
[0056] (1) The mobility correction unit 107 substitutes the peak-extracted signal into the "mobility correction mid-peak signal" and the color-transformed signal into the "mobility correction mid-way signal".
[0057] (2) For the mobility correction mid-peak signal, the mobility correction unit 107 calculates the average interval d(G) between peak PG and other peaks adjacent to it on the time axis, corresponding to bases other than G (guanine) in peak PG of G (guanine). Similarly, for peak PA of A (adenine) and other peaks adjacent to it on the time axis, the mobility correction unit 107 calculates the average interval d(A) between peak PA and other peaks adjacent to it on the time axis. Similarly, for peak PT of T (thymine) and other peaks adjacent to it on the time axis, the mobility correction unit 107 calculates the average interval d(T) between peak PT and other peaks adjacent to it on the time axis. Similarly, the mobility correction unit 107 calculates the average interval d(C) between the peak PC of C (cytosine) and the peaks of other bases that are adjacent to it on the time axis.
[0058] (3) The mobility correction unit 107 calculates the average value d_mean of d(G), d(A), d(T), and d(C).
[0059] (4) The mobility correction unit 107 calculates d'(G), d'(A), d'(T), and d'(C) obtained by the following formula.
[0060] d'(G)=d(G)-d_mean
[0061] d'(A)=d(A)-d_mean
[0062] d'(T)=d(T)-d_mean
[0063] d'(C)=d(C)-d_mean
[0064] (5) The mobility correction unit 107 moves the channel corresponding to G (guanine) of the mobility correction mid-peak signal by d'(G) on the time axis, moves the channel of A (adenine) on the time axis by d'(A), moves the channel of T (thymine) on the time axis by d'(T), and moves the channel of C (cytosine) on the time axis by d'(C), so that the moved signal covers the mobility correction mid-peak signal.
[0065] (6) The mobility correction unit 107 moves the channel corresponding to G (guanine) of the mobility correction intermediate signal by d'(G) on the time axis, moves the channel corresponding to A (adenine) on the time axis by d'(A), moves the channel corresponding to T (thymine) on the time axis by d'(T), and moves the channel corresponding to C (cytosine) on the time axis by d'(C), so that the moved signal covers the mobility correction intermediate signal.
[0066] (7) If the mobility correction unit 107 is small enough (smaller than each threshold), it will substitute the intermediate mobility correction signal after coverage into the mobility correction signal and end the process; otherwise, it will return to the process in (2).
[0067] The peak interval tolerance input unit 109 processes the input of a discrete tolerance α for the peak interval, as indicated by the operator, through an interface screen displayed on the display device 52. α refers to the degree of discrete tolerance of the peak interval. By allowing the operator to set α via the peak interval tolerance input unit 109, sequential base response can be performed with high precision using an optimal α that depends on the magnitude of noise.
[0068] Figure 6 This is an example of the interface screen 200 displayed by the peak interval tolerance input unit 109. Figure 6 The interface screen 200 shown displays the waveforms of the mobility-corrected signal and the updated deconvolutioned signal obtained by processing the temporal sequence of the spectra emitted from pigments corresponding to the bases A (adenine), G (guanine), C (cytosine), and T (thymine).
[0069] The interface screen 200 has an input field 201 for a discrete tolerance α with peak spacing. The interface screen 200 displays the updated convolutional signal calculated using the input tolerance α. Therefore, when the value of the tolerance α is changed, the waveform of the displayed updated convolutional signal also changes. By checking the updated convolutional signal calculated using the tolerance α, the operator can easily determine whether the input tolerance α is an appropriate value.
[0070] Furthermore, the peak spacing tolerance input unit 109 also displays the peak position 202 extracted from the updated convolutional signal, calculated using the input tolerance α. Figure 6In the diagram, peak positions 202 are represented by dashed lines. The operator can easily determine whether the input tolerance value α is appropriate by checking whether the interval of peak positions 202 is fixed. The closer the interval of peak positions 202 is to being fixed, the closer the tolerance value α is to being appropriate. Furthermore, by comparing the mobility-corrected signal displayed on the interface screen 200 with the updated convolution signal, the operator can easily determine whether the deconvolution process has been performed appropriately. That is, it is easy to determine whether the input tolerance value α is appropriate. Additionally, the interface screen 200 also includes a checkbox 203 for selecting whether the inter-block recalculation function of the deconvolution unit 105 is executed, a checkbox 204 for indicating the execution of computation processing in cases of high noise, and a checkbox 205 for indicating the execution of computation processing in cases of low noise. When checkbox 203 is selected, the processing content has already been used. Figure 5 Explanation has been provided. The handling of cases where checkboxes 204 or 205 are selected will be described later. Additionally, Figure 6 Here is an example of interface screen 200. Input bar 201 and check bar 203-205 can be either not displayed or only a portion of them can be displayed.
[0071] return Figure 4 The peak spacing constrained deconvolution unit 108 performs deconvolution processing on the mobility-corrected signal of the 4 channels to calculate the updated deconvolutioned signal of the 4 channels. As mentioned above, it is difficult to obtain sufficient accuracy through deconvolution processing when the PSF is unknown. Therefore, in this embodiment, the peak spacing constrained deconvolution processing unit 108 avoids local solutions by reducing the number of candidate solutions, as described below.
[0072] Figure 7 This describes the processing operation of the peak-spaced constrained deconvolution processing unit 108. First, the peak-spaced constrained deconvolution processing unit 108 limits the PSF to a Gaussian function, restricting its unknown parameters only to the standard deviation σ. By limiting the unknown parameters of the PSF in this way, local solutions can be avoided. Next, the peak-spaced constrained deconvolution processing unit 108 enumerates a discrete finite number of standard deviations σ as parameter candidates (step S301), and performs the following processing on each parameter candidate (step S302). First, the peak-spaced constrained deconvolution processing unit 108 performs deconvolution processing using the PSF as a Gaussian function corresponding to the standard deviation of the processing object, obtaining a 4-channel deconvolutioned signal (step S303).
[0073] Next, the peak spacing constrained deconvolution processing unit 108 calculates e(σ) and v(σ) for the parameter candidates (i.e., standard deviation σ) of the processing object in the order shown in (1) and (2) below.
[0074] (1) Similar to the above processing, the peak-spaced constrained deconvolution processing unit 108 performs convolution processing using a Gaussian function corresponding to the standard deviation σ and the deconvolutioned signal using the function, and calculates the 4-channel convolutioned signal. Then, the peak-spaced constrained deconvolution processing unit 108 calculates the error e(σ) between the calculated convolutioned signal and the input signal based on the mean square error and KL divergence.
[0075] (2) The peak spacing constrained deconvolution processing unit 108 calculates the spacing between each peak extracted from the calculated 4-channel convolution signal and adjacent peaks determined regardless of channel differences. Details have been provided. Figure 6 This has been explained. Next, the peak spacing constrained deconvolution processing unit 108 calculates discrete values v(σ) for all peaks corresponding to these spacings. However, v(σ) may not be a discrete value in the literal sense, but rather the standard deviation, the maximum value of the deviation from the mean, or the Median Absolute Deviation (MAD). In this specification, the term "discrete" is used in the sense that they are used collectively.
[0076] Next, the peak spacing constrained deconvolution processing unit 108 calculates the evaluation value c according to the following formula (step S304).
[0077] c=e(σ)+α×v(σ)……Equation 3
[0078] As described above, α is the tolerance for the discreteness of the peak spacing. For example, in the case of high noise, the peak spacing-constrained deconvolution processing unit 108 emphasizes the discreteness v(σ) of the peak spacing by setting the tolerance α to a large value. As a result, an acceptable solution can be obtained even if the convolution error e(σ) is large. On the other hand, in the case of low noise, the peak spacing-constrained deconvolution processing unit 108 emphasizes the convolution error e(σ) by setting the tolerance α to a small value, and an acceptable solution can be obtained even if the discreteness v(σ) of the peak spacing is large.
[0079] Therefore, the determination of whether the noise level is high can also be automatically executed by the peak-space-constrained deconvolution processing unit 108. Alternatively, it can also be achieved using... Figure 6 As shown, the peak spacing constrained deconvolution processing unit 108 detects which of the checkboxes 204 and 205 on the interface screen 200 the operator has selected, and changes the tolerance α according to the selected content.
[0080] The peak-space-constrained deconvolution processing unit 108 searches for the minimum standard deviation of the evaluation value c (step S305), and outputs the 4-channel convolutioned signal corresponding to the determined standard deviation as the 4-channel updated deconvolutioned signal (step S306). The updated deconvolutioned signal obtained in this way satisfies the constraint that the peak-space is approximately fixed. This solution is highly likely to be the original solution, thus resulting in a deconvolution result with sufficient accuracy. Alternatively, the standard deviation that takes the maximum value can also be searched based on the formula used in calculating the evaluation value c.
[0081] The above processing is performed in each block (see reference). Figure 5 The peak-interval constrained deconvolution unit 108 processes data similarly to the deconvolution unit 105 described above, and can further improve accuracy by using results calculated over multiple blocks. Here, the standard deviation σ of the i-th block is also set to σ_i. It is also assumed here that the standard deviation σ does not change significantly in a short period of time. The peak-interval constrained deconvolution unit 108 determines whether the difference between the standard deviation σ of a certain block and the average of the standard deviations of the blocks preceding and following it exceeds a threshold T.
[0082]
[0083] When Equation 4 holds, the peak-spacing constrained deconvolution unit 108 performs a second deconvolution using a PSF with a standard deviation of the value obtained through Equation 5, and calculates a new deconvolutioned signal. Here, K is a constant of some natural number (1, 2, 3...).
[0084]
[0085] return Figure 4 The peak extraction unit 110 performs peak extraction waveform processing, calculating the updated peak-extracted signal for the four channels from the updated deconvolutioned signal. In the peak extraction processing, methods such as calculating the zero-intersection of the first derivative of the updated deconvolutioned signal and replacing the signal value of the updated peak-extracted signal at the moment the zero-intersection exists with the signal value of the updated deconvolutioned signal are used. In calculating the first derivative of the updated deconvolutioned signal, spline interpolation is performed on the updated deconvolutioned signal, and the first derivative value of the spline-interpolated signal is calculated. Therefore, robust peak extraction of noise is possible.
[0086] In the peak extraction unit 110, in order to further improve accuracy, a judgment process based on the following judgment rule can be performed on all peaks, and peaks whose judgment result is "delete" can be deleted.
[0087] Judgment rule: If d < 0, then delete.
[0088] Where d is a discriminant function that depends on x_1 and x_2.
[0089] ·x_1 = Intensity of the most notable peak - Average intensity of nearby peaks
[0090] x_2 = (Intensity of the channel at the moment of the focal peak of the signal after mobility correction) / (Average intensity of the entire channel at the moment of the focal peak of the signal after mobility correction)
[0091] For example, in d, you can use a linear discriminant function d = w_1*x_1 + w_2*x_2 + w_3, or you can use SVM (Support Vector Machine), or you can use a decision tree.
[0092] The sequence determination unit 111 performs sequence determination processing, extracting the result sequence from the updated peak signals of the 4 channels. In the sequence determination processing, a tag column is generated by arranging the base types G, A, T, and C corresponding to each peak in the extracted updated peak signals in chronological order. This tag column is designated as the result sequence. The sequence output unit 112 outputs the result sequence to a storage device such as a hard disk or SSD (Solid State Drive), or displays it using the display device 52.
[0093] (1-3) Summary
[0094] As explained above, the capillary array electrophoresis apparatus 10 in this embodiment is equipped with a function that performs deconvolution processing based on limiting the peak spacing. Therefore, even when peak separation is low (especially when waveforms of different base types overlap in time), the base sequence can be determined with high accuracy. In particular, considering that peak separation tends to decrease when using portable (small) sequencers or sequencers with short measurement times, the method of this embodiment is effective for these sequencers.
[0095] (2) Example 2
[0096] In this embodiment, the function of employing various parameters of machine learning in the capillary array electrophoresis apparatus 10 will be described. Furthermore, the basic apparatus of the capillary array electrophoresis apparatus 10 in this embodiment is the same as that in Embodiment 1. In the following description, multiple input signals obtained by measuring samples with known base sequences under the same measurement conditions are input to the signal input unit 101.
[0097] First, as learning step 1, a method for automatically adjusting the discrete tolerance α of the peak interval will be described. The signal processing unit 72 changes the tolerance α at a certain time Δα within a certain range [α_min, α_max], and performs a series of processes described in Example 1 using each α to determine the base sequence. Then, the signal processing unit 72 calculates the matching rate p(α) between the base sequence determined using each tolerance α and the correct sequence, and sets the tolerance α in a way that maximizes the matching rate p(α). By executing this learning step, the input of the tolerance α using the user interface screen 200 can be omitted.
[0098] Next, in learning step 2, the method for adjusting parameters other than the tolerance α will be explained. For example, the signal processing unit 72 adjusts the parameters w_1, w_2, and w_3 of the deletion determination rule of the peak extraction unit 110 described in Example 1. Here, since the base sequence determination is known for the sample, it is possible to determine whether each peak was correctly extracted, incorrectly extracted (e.g., peaks that were not originally extracted), incorrectly failed to exclude (e.g., peaks that were originally extracted), or correctly excluded (e.g., peaks that were originally excluded). Using this correctness / incorrectness as a teacher signal, the signal processing unit 72 performs teacher-guided learning by comparing it with the output signal d.
[0099] When the output signal d is a linear discriminant function, the signal processing unit 72 can use algorithms such as error correction learning. When the output signal d is an SVM (Support Vector Machine), the signal processing unit 72 can use the SMO (Sequential Minimal Optimization) algorithm for teacher-guided learning. When the output signal d is a decision tree, the signal processing unit 72 can use algorithms such as ID3 or C4.5 for teacher-guided learning.
[0100] The results of learning process 1 and learning process 2 are interdependent. Therefore, the signal processing unit 72 performs the optimal parameter adjustment as a whole by alternately performing learning process 1 and learning process 2.
[0101] (3) Example 3
[0102] Figure 8 Other structural examples of the capillary array electrophoresis apparatus 10A are shown. Figure 8 In the middle, to and Figure 2 Corresponding parts are labeled with the same reference numerals. In the case of Embodiment 1, the base sequence determination process is performed by the signal processing unit 72 of the fluorescence detection device 50. In this embodiment, the fluorescence detection device 50A does not have a base sequence analysis function in its signal processing unit 72; this is performed by an external device 53, such as a computer, connected to the capillary array electrophoresis apparatus 10A. Therefore, in Figure 8 The capillary array electrophoresis apparatus 10A shown has an external device 53 connected to the control board 51. The external device 53 has sufficient computing resources to perform signal processing executed by the signal processing unit 72. Furthermore, the external device 53 is equipped with a display device for confirming processing results, etc. Communication between the control board 51 and the external device 53 can be either wired or wireless. For example, the external device 53 can be a desktop PC, a laptop PC, a smartphone, or a portable information terminal.
[0103] (4) Other embodiments
[0104] This invention is not limited to the embodiments described above, but includes various modifications. For example, the embodiments described above have been given in detail to make the invention easy to understand, but are not necessarily limited to all the structures described. Furthermore, a part of the structure of one embodiment can be replaced with the structure of another embodiment, and a structure of another embodiment can be added to the structure of one embodiment. In addition, other structures can be added, deleted, or replaced to a part of the structure of each embodiment.
[0105] Furthermore, the aforementioned structures, functions, processing units, and processing modules can be partially or entirely implemented in hardware, for example, using integrated circuits. Additionally, the aforementioned structures and functions can be implemented in software by interpreting and executing programs that implement the functions of the information processor. The programs, diagrams, folders, and other information implementing these functions can be stored in recording devices such as memory, hard disks, SSDs (Solid State Drives), or recording media such as IC cards, SD cards, and DVDs.
[0106] Furthermore, for control lines and information lines, only the parts deemed necessary in the description are shown; not all control lines and information lines on the product are necessarily shown. In fact, almost all structures can be considered interconnected.
[0107] Explanation of reference numerals in the attached figures
[0108] 10 capillary array electrophoresis device
[0109] 11 Sample Container
[0110] 12 sample trays
[0111] 20 conveyors
[0112] 21 High Voltage Power Supply
[0113] 30 capillary array
[0114] 31. Thermostatic bath
[0115] 32 Detection Locations
[0116] 33 Capillary
[0117] 34 Discharge location
[0118] 41 Electrophoretic Media
[0119] 42 Pump Unit
[0120] 50 Fluorescence Detection Device
[0121] 51 Control board
[0122] 52 Display devices
[0123] 72 Signal Processing Department
[0124] 101 Signal Input Section
[0125] 102 Baseline Removal Section
[0126] 103 Color Transformation Section
[0127] Block 104 Cutting Section
[0128] 105 Deconvolution part
[0129] 106 Peak Extraction Section
[0130] 107 Mobility Correction Department
[0131] 108 peak spacing constrained deconvolution part
[0132] 109 Peak Interval Tolerance Input Section
[0133] 110 Peak Extraction Section
[0134] 111 Sequence Determination Unit
[0135] 112 Sequence Output Section
[0136] 200 Interface Screen
[0137] 201. Input field for tolerance α
[0138] 202 Peak Position
[0139] 203 Checkbox
[0140] 204 Checkbox
[0141] 205 Checkbox.
Claims
1. A base sequence determining apparatus characterized by comprising: including: a signal input section that inputs a time-series signal of a spectrum detected by a capillary electrophoresis type sequencer from a pigment corresponding to each base; a mobility correction section that calculates an average interval of adjacent peaks of each base with respect to the time-series signal of the spectrum corresponding to each base, calculates a difference between the average interval of each base and an average value of the average intervals of the bases, moves a channel corresponding to each base to a rear direction on a time axis based on the difference, and outputs a mobility-corrected signal using a mobility-corrected interim signal after the movement in a case where the difference of each base is smaller than a threshold value set in advance; a deconvolution section that performs the following processing: limits a point spread function to a Gaussian function, and limits only a standard deviation to an unknown parameter in the point spread function; lists a discrete finite number of the standard deviations as parameter candidates; performs the following processing with respect to each of a plurality of parameter candidates of the point spread function, that is, processing of calculating a deconvolution-after signal of the mobility-corrected signal, and processing of calculating a convolution error and a dispersion of peak intervals with respect to each of the deconvolution-after signals calculated with respect to the plurality of parameter candidates, and calculates an evaluation value in accordance with c = e(σ) + α × v(σ), where c is the evaluation value, e(σ) is the convolution error, v(σ) is a size of the dispersion, and α is an allowance of the dispersion of the peak intervals; determines a parameter of the point spread function by searching for a candidate having the smallest evaluation value among all parameter candidates of the standard deviation; and outputs the deconvolution-after signal corresponding to the point spread function having the determined parameter as an updated deconvolution-after signal. a peak extraction section that extracts a peak waveform from the updated deconvolution-after signal and outputs an updated peak-extraction-after signal; and a sequence determination section that inputs the updated peak-extraction-after signal and determines a base sequence, the base sequence determination apparatus further has an input section that obtains the allowance by accepting an input to an interface screen of the input section, the dispersion of the peak intervals refers to a standard deviation of peak intervals, a maximum value of a deviation from an average value, or an absolute median deviation.
2. The base sequence determination apparatus according to claim 1, wherein: a waveform of the updated deconvolution-after signal displayed on the interface screen changes in correspondence with a change in an input value of the allowance.
3. The base sequence determination apparatus according to claim 1, further comprising: an input section that accepts whether or not to use a calculation result of the dispersion calculated with respect to a plurality of time intervals in the calculation of the dispersion of the peak intervals of the deconvolution-after signal performed by the deconvolution section, by presence or absence of a check in a check column displayed on an interface screen.
4. The base sequence determination apparatus according to claim 1, wherein: the deconvolution section corrects a calculation result of the dispersion of the peak intervals of the deconvolution-after signal using a calculation result of the dispersion calculated with respect to a plurality of time intervals other than a time interval corresponding to the calculation result, and determines a parameter of the point spread function using the corrected dispersion. 5. The base sequence determination apparatus according to claim 1, wherein: the deconvolution section extracts only a parameter candidate capable of obtaining a deconvoluted signal satisfying the following (1) to (3) before performing processing of parameters of the point spread function used in calculation of the updated deconvoluted signal, (1) a convolution error between the deconvoluted signal and the mobility-corrected signal is below a threshold value set in advance, (2) a minimum value of intervals of peaks extracted from the deconvoluted signal is above a threshold value, (3) the parameter candidate of the point spread function is in a vicinity of an inflection point of the convolution error. comprising:
6. A capillary array electrophoresis device characterized by, an array of capillaries for electrophoresis of a test sample; a high-voltage power supply for applying an electrophoresis voltage to the array of capillaries; a light-receiving section for detecting fluorescence from the array of capillaries; and a base sequence determination apparatus according to claim 1 for processing a signal from the light-receiving section to determine a base sequence of the test sample.
7. A base sequence determination method executed by a base sequence determination apparatus having a signal input section, a signal processing section, and a memory, the base sequence determination method comprising: the signal input section inputting a time-series signal of a spectrum emitted from a pigment corresponding to each base detected by a capillary electrophoresis type sequencer, the signal processing section calculating an average interval of adjacent peaks of each base with respect to the time-series signal of the spectrum corresponding to each base, and calculating a difference between the average interval of each base and an average value of the average intervals of the bases, moving a channel corresponding to each base to a rear direction on a time axis based on the difference, and outputting a mobility-corrected signal after the movement as a mobility-corrected signal, the signal processing section performing processing of limiting a point spread function to a Gaussian function, and limiting only a parameter unknown in the point spread function to a standard deviation, the signal processing section enumerating a discrete finite number of the standard deviations as parameter candidates, the signal processing section performing, for each of a plurality of parameter candidates of the point spread function, processing of calculating a deconvoluted signal of the mobility-corrected signal, and for each of the deconvoluted signals calculated with respect to the plurality of parameter candidates, calculating a convolution error and a dispersion of peak intervals, and calculating an evaluation value in accordance with c = e (σ) + α x v (σ), where c is the evaluation value, e (σ) is the convolution error, v (σ) is a size of the dispersion of peak intervals, and α is an allowable degree of the dispersion of peak intervals, the signal processing section determining parameters of the point spread function by searching for a candidate having the smallest evaluation value among all parameter candidates of the standard deviation, the signal processing section outputting, as an updated deconvoluted signal, the deconvoluted signal corresponding to the point spread function having the determined parameters, the signal processing section extracting a peak waveform from the updated deconvoluted signal, and outputting an updated peak-extracted signal, and the signal processing section inputting the updated peak-extracted signal, and determining a base sequence. The signal processing section obtains the tolerance by accepting input to an input field of the tolerance displayed on the interface screen, The dispersion of the peak interval refers to a standard deviation of the peak interval, a maximum value of a deviation from an average value, or an absolute median deviation.
8. The base sequence determination method according to Claim 7, wherein: The signal processing section changes the waveform of the updated deconvoluted signal displayed on the interface screen in accordance with a change in the input value of the tolerance.
9. The base sequence determination method according to Claim 7, wherein: The signal processing section displays a check field on the interface screen, the check field indicating whether or not to use the calculation result of the dispersion calculated for a plurality of time intervals in the calculation of the dispersion of the peak interval of the deconvoluted signal.
10. The base sequence determination method according to Claim 7, wherein: The signal processing section corrects the calculation result of the dispersion of the peak interval of the deconvoluted signal using the calculation result of the dispersion calculated for a plurality of time intervals other than the time interval corresponding to the calculation result, and determines the parameter of the point spread function using the corrected dispersion.
11. The base sequence determination method according to Claim 7, wherein: The signal processing section extracts only a parameter candidate capable of obtaining a deconvoluted signal satisfying (1) to (3) below before executing the process of determining the parameter of the point spread function used in the calculation of the updated deconvoluted signal, (1) a convolution error between the deconvoluted signal and the mobility-corrected signal is equal to or less than a threshold value set in advance, (2) a minimum value of the interval of peaks extracted from the deconvoluted signal is equal to or greater than a threshold value, (3) the parameter candidate of the point spread function is in the vicinity of an inflection point of the convolution error.
Citation Information
Patent Citations
Method of determining base sequence of nucleic acid
WO2008050426A1
Determination of nucleic acid base sequence
CN1363688A
Method of determining base sequence of nucleic acid
US20100089771A1
Signal processing by iterative deconvolution of time series data
US20100266177A1
Device for genotypic analysis and method for genotypic analysis
CN104870980A