Capillary electrophoresis device

By using chromatographic separation and comparison of model fluorescence signals, the problem of detecting multiple fluorophores under spectral and spatiotemporal overlap was solved, and accurate base response results were achieved.

CN115808461BActive Publication Date: 2025-11-04HITACHI HIGH TECH CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211503360.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-02-20
Publication Date
2025-11-04
Estimated Expiration
2037-02-20

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately detect the emission of multiple fluorophores under conditions of spectral and spatiotemporal overlap, especially when the number of fluorophore types exceeds the number of detection wavelengths, leading to erroneous base response results.

Method used

By separating the fluorescence signals of multiple components by chromatography and combining them with the time-series data of the model fluorescence signals, a computer is used to compare the data and determine the fluorophore label of each component, thus achieving N-color detection in N wavelength bands.

Benefits of technology

Even under conditions of spectral and spatiotemporal overlap, it can accurately detect M components, improving the accuracy and reliability of base response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115808461B_ABST
    Figure CN115808461B_ABST
Patent Text Reader

Abstract

A capillary electrophoresis device detects M kinds of components by N-color detection in M wavelength bands (M>N) in a state where emission fluorescence from M kinds of phosphors has spectral overlap and spatiotemporal overlap. The capillary electrophoresis device includes a sample containing four or more kinds of phosphors, a capillary for performing electrophoretic analysis of the sample, a light source that irradiates a laser beam to the capillary, an optical system that condenses fluorescence emitted from a light emission point of the capillary due to irradiation of the laser beam, and a sensor that measures an image of the light emission point generated by the optical system, wherein the sensor is an RGB color sensor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application entitled "Analysis system and analysis method", filed on February 20, 2017, with the Patent Office, application number 201780085297.8. TECHNICAL FIELD

[0002] The present application relates to an analysis system and an analysis method. BACKGROUND

[0003] An analysis method is widely used in which a plurality of components (however, not necessarily one-to-one) contained in a sample represented by DNA, protein, cells, various biological samples are labeled with a plurality of fluorescent substances, and detection is performed while identifying the luminescent fluorescence from the plurality of fluorescent substances, whereby the plurality of components are analyzed. As such analysis methods, there are, for example, chromatograms, DNA sequencers, DNA fragment analysis, flow cytometers, PCR, HPLC, Western / Northern / Southern blotting, microscopic observation, and the like.

[0004] Generally, the fluorescence spectra of the plurality of fluorescent substances used have overlaps with each other (hereinafter, referred to as spectral overlap). Furthermore, the luminescent fluorescence of the plurality of fluorescent substances has overlaps with each other in time or space (hereinafter, referred to as temporal-spatial overlap). In the presence of the above-described overlaps, a technique for performing detection while identifying the luminescent fluorescence from the plurality of fluorescent substances is required.

[0005] Next, the technique will be described taking a DNA sequencer using electrophoresis as an example. The DNA sequencer using electrophoresis has changed from slab electrophoresis in the 1980s to a method of capillary electrophoresis after the 1990s, but the technique for solving the above-described problems has not changed substantially. The method of slab electrophoresis described in Non-Patent Literature 1 is as follows. Figure 3 is a method of slab electrophoresis. Furthermore, the present technique is also used in capillary electrophoresis. The basic steps of the present technique consist of (1) to (4) below under the condition that M ≤ N.

[0006] (1) A sample containing M kinds of DNA fragments labeled with M kinds of fluorescent substances is separated by electrophoresis while irradiating a laser beam, whereby the fluorescent substances luminesce, N-color detection is performed in N wavelength bands, and time-series data of N-color fluorescence intensity are acquired.

[0007] (2) Color conversion is performed for each time of the time-series data (1), and time-series data of the concentrations of the M kinds of fluorescent substances, i.e., the M kinds of DNA fragments, are acquired.

[0008] (3) For each time of the time-series data (2), different correction based on the degree of movement of the M kinds of fluorescent substances (hereinafter, referred to as degree-of-movement correction) is performed, and time-series data of the concentrations of the M kinds of fluorescent substances, i.e., the M kinds of DNA fragments, after the above-described degree-of-movement correction are acquired.

[0009] (4) Based on the time-series data (3), a base response is performed.

[0010] The above (1) to (4) will be described in detail with reference to Non-Patent Literature 1. Figure 3 Here, M = 4, and N = 4. Hereinafter, a case in which one sample is analyzed using one electrophoresis path will be described. In a case in which a plurality of samples is analyzed using a plurality of electrophoresis paths, the following steps are performed in parallel.

[0011] First, (1) will be described. When copies of DNA fragments of various lengths corresponding to a template DNA are produced by a Sanger reaction, the DNA fragments are labeled with four kinds of fluorescent substances (hereinafter, these fluorescent substances will be referred to as C, A, G, and T, respectively, for simplicity) corresponding to the types of bases C, A, G, and T at the ends thereof. While the DNA fragments labeled with the fluorescent substances are separated by electrophoresis in terms of length, laser beams are sequentially irradiated to them, and the fluorescent substances are caused to emit light. The fluorescent light emitted is detected in four wavelength bands b, g, y, and r (each wavelength band is preferably identical to the maximum light emission wavelength of C, A, G, and T) (hereinafter, this will be referred to as 4-color detection). Thereby, time-series data of the fluorescent intensities I(b), I(g), I(y), and I(r) of the four colors are obtained. The above (1) to (4) of Non-Patent Literature 1 will be described in detail with reference to FIG. 1. Figure 3 A is time-series data (also referred to as raw data) of I(b), I(g), I(y), and I(r) represented by blue, green, black, and red, respectively.

[0012] Next, (2) will be described. As in Equation (1), the concentrations D(C), D(A), D(G), and D(T) of the four kinds of fluorescent substances C, A, G, and T at this time point represent the fluorescent intensities of the four colors at this time point.

[0013] [Math. 1]

[0014]

[0015] Here, the element w(XY) of the 4-row and 4-column matrix W represents the intensity proportion of the fluorescent substance type Y (C, A, G, or T) detected in the wavelength band X (b, g, y, or r) by spectral overlap. w(XY) is a fixed value determined only by the characteristics of the fluorescent substance type Y (C, A, G, or T) and the wavelength band X (b, g, y, or r) and does not change during electrophoresis.

[0016] Therefore, as in Equation (2), the concentrations of the four kinds of fluorescent substances at this time point are calculated from the fluorescent intensities of the four colors at this time point.

[0017] [Math. 2]

[0018]

[0019] Thus, by multiplying the inverse matrix W -1 by the fluorescence intensities of the 4 colors, the spectral overlap is eliminated (hereinafter, this step is referred to as color conversion). By this, the time series data of the concentrations of the 4 kinds of fluorescent substances, C, A, G, and T, i.e., the concentrations of the DNA fragments whose ends are the 4 kinds of bases, are obtained.

[0020] The Figure 3 B is the time series data. The color conversion can be performed regardless of the presence or absence of the spatiotemporal overlap. As seen from Figure 3 B, the presence of the concentrations of the multiple fluorescent substances at the same time indicates the presence of the spatiotemporal overlap.

[0021] The Figure 3 C is a step that does not correspond to (1) to (4) and is a step that is not always necessary. By deconvolution, the time series data of Figure 3 B are separated into individual peaks, i.e., signals indicating the concentrations of the single-length DNA fragments labeled with any one of the 4 kinds of fluorescent substances, C, A, G, and T, thereby eliminating the spatiotemporal overlap.

[0022] Finally, (3) and (4) are described. Generally, by electrophoresis, the DNA fragments of various lengths each having a different base length are separated at substantially equal intervals. However, due to the influence of the fluorescent substance labeled to the DNA fragments, the mobility (electrophoretic speed) varies, and the above equal intervals are sometimes broken. Therefore, the magnitude relationship of the mobility due to the kind of the labeled fluorescent substance is investigated in advance, and based on this information, the time series data of Figure 3 B or Figure 3 C of Non-Patent Literature 1 are corrected. By this, the time series data in which the DNA fragments of various lengths each having a different base length are arranged at substantially equal intervals are obtained. The Figure 3 D is the time series data. Each peak of blue, green, black, and red indicates the concentration of the single-length DNA fragments whose end base kind is C, A, G, and T, and they are arranged in order of each base length. Therefore, as shown in Figure 3 D, by sequentially recording the above end base kind, the result of the base response can be obtained.

[0023] In addition, in each of the above steps, or before and after each step, a process such as smoothing, noise filtering, baseline removal, and the like of the time series data is sometimes appropriately performed.

[0024] In the color conversion step of (2), as shown in Equation (1), for each time, the four kinds of fluorescent substance concentrations D(C), D(A), D(G), and D(T) are solved by solving a four simultaneous equations composed of the known numbers of the four kinds of fluorescent detection intensities I(b), I(g), I(y), and I(r) of b, g, y, and r. Generally, corresponding to solving M unknown numbers by N simultaneous equations, therefore as described above, the condition of M ≤ N is required. Assuming that if M > N, the solution cannot be uniquely solved (i.e., multiple solutions can exist), therefore the color conversion as Equation (2) cannot be performed.

[0025] However, in Non-Patent Literature 2, under the condition of M > N of M = 4 and N = 3, the DNA sequence generated by electrophoresis is acted upon. If fluorescent light emission is detected in three kinds of wavelength bands b, g, and r (hereinafter referred to as 3-color detection), the four-color fluorescent intensity of each time of Equation (1) is replaced by the three-color fluorescent intensity of each time of Equation (3).

[0026] [Mathematical Expression 3]

[0027]

[0028] At this time, the matrix W is 3 rows and 4 columns, and there is no inverse matrix, therefore as Equation (2), the concentrations of the four kinds of fluorescent substances cannot be uniquely solved. As described above, generally under the condition of M > N, the solution cannot be solved, but by adding a precondition as follows, the solution can be solved.

[0029] First, in the first step, it is assumed that there is no spatiotemporal overlap of multiple fluorescent substances, i.e., only one fluorescent substance emits light at a time. At this time, at each time, the Y (C, A, G, or T) whose ratio of the three-color fluorescent intensity (I(b) I(g) I(r)) T is closest to the ratio of the four columns (w(bY) w(gY) w(rY)) T of the matrix W can be selected. According to the other expression, in Equation (3), one kind of fluorescent substance is selected, and the concentrations of the remaining three kinds of fluorescent substances are set to 0. That is, when (D(C) D(A) D(G) D(T)) T is set to (D(C) 00 0) T , (0 D(A) 0 0) T , (0 0 D(G) 0) T , (0 0 0 D(T)) T , D(Y) (Y is C, A, G, or T) whose left and right sides are the smallest is solved, respectively. The Y (C, A, G, or T) whose left and right sides are the smallest at this time can be selected. Here, in the case where the left and right sides are not sufficiently small in any case, proceed to the second step below.

[0030] In the second step, it is assumed that only two phosphors emit light at a time. This time, in formula (3), two phosphors are selected, and the concentrations of the remaining two phosphors are set to 0. That is, when (D(C)D(A)D(G)D(T)) is used... T Let it be (D(C)D(A)0 0) T 、(D(C)0D(G)0) T 、(D(C)0 0D(T)) T 、(0D(A)D(G)0) T 、(0D(A)0D(T)) T 、or (0 0D(G)D(T)) T When the difference between the left and right sides is minimized, find the two D(Y) (Y can be C, A, G, or T). Then, select the two Y(C, A, G, or T) that minimize the difference between the left and right sides.

[0031] In this way, similar to step (2) in Non-Patent Literature 1, time-series data of the concentrations of four fluorophores, i.e., four DNA fragments, are obtained. Then, by performing the same steps as steps (3) and (4) in Non-Patent Literature 1, the results of base response can be obtained.

[0032] For the conditions for implementing the first and second steps in Non-Patent Document 2 to hold true, namely the assumption that only one or two phosphors emit light at a time, the spatiotemporal overlap must be small, i.e., the following two conditions must hold true.

[0033] (a) In the time series data of the concentrations of the four DNA fragments, the multiple peaks generated by each single-length DNA fragment are arranged at approximately equal intervals.

[0034] (b) In the time series data of the concentrations of the four DNA fragments, the two adjacent peaks generated by DNA fragments of different base lengths were well separated.

[0035] In Non-Patent Literature 2, four primers used in the Sanger reaction were labeled with four different fluorophores (more precisely, the labeling methods of three fluorophores were changed), and the difference in mobility caused by the influence of the fluorophores on the DNA fragments was made sufficiently small. Furthermore, conditions were set for sufficiently high electrophoretic separation performance. The result was the acquisition of the three-color fluorescence intensities (I(b)I(g)I(r)) established in (a) and (b) above. T Time series data (non-patent document 2) Figure 2 (Above paragraph). Furthermore, under these conditions, the first and second steps described above are performed. This yields time-series data on the concentrations of four fluorophores, i.e., four DNA fragments, and provides results on the base response (Non-Patent Document 2). Figure 2 next paragraph).

[0036] Prior Art Documents

[0037] Non-Patent Documents

[0038] Non-Patent Document 1: Genome Res. 1998 Jun; 8(6): 644-65

[0039] Non-Patent Document 2: Electrophoresis. 1998 Jun; 19(8-9): 1403-14 SUMMARY

[0040] Problems to be Solved by the Invention

[0041] The RGB color sensor is a two-dimensional sensor in which three kinds of pixels that detect red (R), green (G), and blue (B) wavelength bands corresponding to 3 primary colors that can be recognized by human eyes are arranged. The RGB color sensor is not only used in single-lens reflex digital cameras and compact digital cameras, but also in digital cameras mounted in smartphones, and has been rapidly spread worldwide in recent years. Therefore, the performance of the RGB color sensor has been significantly improved, and the price thereof has been significantly reduced. Thus, it is very useful to apply the RGB color sensor to an analysis method that detects a plurality of components while recognizing luminescent light from a plurality of phosphors. However, the RGB color sensor can detect only 3 colors, and thus it is difficult to recognize luminescent light from 4 or more kinds of phosphors. As described above, in order to recognize luminescent light from M kinds of phosphors, in a case where the M kinds of phosphors have spectral overlap and spatial overlap, N-color detection must be performed in N kinds of bands under the condition of M ≤ N.

[0042] In Non-Patent Document 2, in relation to a DNA sequencer that uses electrophoresis, the above-described difficulty is solved by setting a precondition for spatial overlap in a case where luminescent light from M = 4 kinds of phosphors is detected for N = 3 colors (that is, in a case where M > N). The condition is that only one or two kinds of phosphors perform luminescent light emission at a time, that is, the above-described (a) and (b) are satisfied.

[0043] However, the condition of the above-described (a) and (b) is not satisfied in many cases. Even in the case of Non-Patent Document 2, in the time-series data of the intensities of the 3 colors (upper row) and the time-series data of the concentrations of the 4 kinds of phosphors (lower row) in Figure 2 the region indicated by the asterisk, 3 or more peaks due to each single-length DNA fragment exist densely due to a phenomenon called compression, regardless of the high separation performance of the electrophoresis. Thus, the above-described condition (a) is broken, and a correct base response is not generated. In addition, in the case of Non-Patent Document 2, in the time-series data of the intensities of the 3 colors (upper row) and the time-series data of the concentrations of the 4 kinds of phosphors (lower row) in Figure 2In the latter stage of the electrophoresis, the separation performance of the electrophoresis decreases. Therefore, the separation of two adjacent peaks generated by DNA fragments differing by one base length is insufficient, the condition (b) is broken, and a correct base response is not generated.

[0044] In Non-Patent Literature 2, the condition (a) is achieved by using a primer labeling method in which four kinds of fluorescent bodies are respectively labeled to four kinds of primers used in the Sanger reaction. However, in recent years, a terminator labeling method in which four kinds of fluorescent bodies are respectively labeled to four kinds of terminators used in the Sanger reaction is mainly used instead of the primer labeling method. In the primer labeling method, the Sanger reaction using four kinds of primers must be performed respectively, that is, in four different sample tubes, and, in contrast, in the terminator labeling method, the Sanger reaction using four kinds of terminators can be performed together, that is, in one sample tube. Therefore, the terminator labeling method can greatly simplify the Sanger reaction.

[0045] However, in the primer labeling method, the difference in the moving degree of the DNA fragments labeled with four kinds of fluorescent bodies is small, and, in contrast, in the terminator labeling method, the difference in the moving degree of the DNA fragments labeled with four kinds of fluorescent bodies is large, and therefore the condition (a) is necessarily not established. That is, not limited to only one or two kinds of fluorescent bodies emit fluorescence, there are cases where three or more kinds of fluorescent bodies emit fluorescence at a time. Therefore, at least in the case of using the terminator labeling method, it is difficult to perform DNA sequencing by the method of Non-Patent Literature 2.

[0046] Therefore, the present application provides an analysis technique that detects M kinds of components by N-color detection in N wavelength bands from the emission fluorescence of M kinds of fluorescent bodies in a state where the emission fluorescence has spectral overlap and spatial overlap.

[0047] Means for solving the problem

[0048] For example, in order to solve the above problem, the structure described in the claimed scope is adopted. The present application has a plurality of means for solving the above problem, but if one example is cited, an analysis system is provided, which has: an analysis device that separates a sample containing a plurality of components labeled with any one of M kinds of fluorescent bodies by chromatography, and acquires first time series data of fluorescence signals detected in N wavelength bands in a state where at least a part of the plurality of components is not completely separated, where M>N; a storage unit that stores second time series data of model fluorescence signals of the plurality of components respectively; and a computer that determines which component of the plurality of components is labeled with which fluorescent body of the M kinds by comparing the first time series data with the second time series data.

[0049] According to another example, there is provided an analysis method including: separating a sample containing a plurality of components labeled with any one of M types of fluorescent substances by chromatography, acquiring first time-series data of fluorescent signals detected in N types of wavelength bands in a state in which at least a part of the plurality of components is not completely separated, where M>N; and determining each of the plurality of components as a component labeled with which of the M types of fluorescent substances by comparing the first time-series data with second time-series data of model fluorescent signals of the plurality of components.

[0050] Effects of Invention

[0051] According to the present application, even when M>N, M types of components can be detected in a state in which fluorescent light from the M types of fluorescent substances has spectral overlap and spatiotemporal overlap. Furthermore, more features associated with the present application can be understood from the description and drawings of the present specification. In addition, problems, structures, and effects other than those described above can be understood from the description of the following examples. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 is a diagram showing a DNA sequencing method of Non-Patent Literature 1 using model data.

[0053] Figure 2 is a diagram showing a DNA sequencing method of Non-Patent Literature 2 using model data.

[0054] Figure 3 is a diagram showing an example of a DNA sequencing method using model data of Example 1.

[0055] Figure 4 is a diagram showing another example of a DNA sequencing method using model data of Example 1.

[0056] Figure 5 is a diagram showing a processing procedure and system structure of Example 1.

[0057] Figure 6 is a diagram showing a processing procedure and system structure of Example 1.

[0058] Figure 7 is a diagram showing a processing procedure and system structure of Example 2.

[0059] Figure 8 is a diagram showing a processing procedure and system structure of Example 3 (M=4, N=3).

[0060] Figure 9 is a diagram showing a processing procedure and system structure of Example 4 (M=4, N=2).

[0061] Figure 10 is a block diagram of a capillary electrophoresis device.

[0062] Figure 11 is a block diagram of a multi-color detection device of a multi-capillary electrophoresis device.

[0063] Figure 12 is a block diagram of a computer.

[0064] Figure 13 represents an example of DNA sequencing using Figures 9-12 the processing procedure and system configuration shown in

[0065] Figure 14 is a graph showing time-series data of 2-color fluorescence intensities representing model peaks of single-length DNA fragments labeled with 4 kinds of fluorescent substances in Example 5.

[0066] Figure 15 is a graph showing a procedure of DNA sequencing in Example 5.

[0067] Figure 16 is a graph showing a procedure of DNA sequencing in Example 5.

[0068] Figure 17 is a graph summarizing end base types, fitting accuracy, and QV (quality value) in order of electrophoresis time for each model peak of (5) of Figure 15

[0069] is a graph summarizing end base types, fitting accuracy, and QV in order of corrected electrophoresis time for each model peak of (6) of Figure 18 DETAILED DESCRIPTION Figure 15 Embodiments of the present application will be described below with reference to the accompanying drawings. Note that the drawings used in the following description are schematic and for the purpose of explanation only. They are not intended to define the limits of the present application. In describing the embodiments of the present application, the following terms are used.

[0070] The following embodiment relates to a device that detects a sample containing a plurality of components labeled with a plurality of fluorescent substances while identifying each fluorescence and analyzes each component. The following embodiment can be applied to, for example, the fields of chromatography, a DNA sequencer, DNA fragment analysis, a flow cytometer, PCR, HPLC, Western / Northern / Southern Blot, microscopic observation, and the like.

[0071] For the case of performing electrophoresis-based DNA sequencing, the device of

[0072] Figures 1-4 ​​The contents of Non-Patent Documents 1 and 2 are explained in more detail.

[0073] Figure 1 The method of non-patent document 1. Figure 1 (1) represents the time-series data of the four-color fluorescence intensities I(b), I(g), I(y), and I(r) obtained by performing four-color detection of the emission fluorescence of four phosphors C, A, G, and T in four wavelength bands b, g, y, and r. The horizontal axis of the figure is time, and the vertical axis is fluorescence intensity.

[0074] Figure 1 Let (W) represent the 4x4 element w(XY) of matrix W. For example, the four black bars, from left to right, represent w(bC), w(gC), w(yC), and w(rC). Similarly, horizontal bars represent w(XA) (where X is b, g, y, and r), and diagonal bars represent w(XG) (where X is b, g, y, and r). Grid bars represent w(XT) (where X is b, g, y, and r). Furthermore, w(bY), w(gY), w(yY), and w(rY) are standardized such that w(bY) + w(gY) + w(yY) + w(rY) = 1 (where Y is C, A, G, or T).

[0075] Matrix W and its inverse matrix W -1 The details are as follows.

[0076] [Mathematical Expression 4]

[0077]

[0078]

[0079] By using the inverse matrix of matrix W -1 Multiply Figure 1 The values ​​of I(b), I(g), I(y), and I(r) at each time step (i.e., color transformation) in (1) are as follows: Figure 1 As shown in (2), the time-series data of the concentrations D(C), D(A), D(G), and D(T) of four fluorescent cells of C, A, G, and T, i.e., four DNA fragments with base-terminated ends C, A, G, and T, were obtained.

[0080] exist Figure 1In (2), four peaks were obtained: C, A, G, and T. The height of each peak, i.e., the concentration (in arbitrary units), is D(C) = 100, D(A) = 80, D(G) = 90, and D(T) = 80. In addition, the time of each peak (in arbitrary units) is 30, 55, 60, and 75. The peaks of A and G have a large spatiotemporal overlap, but the color transformation can still correctly determine the concentrations under such circumstances.

[0081] Figure 1 (3) is the difference in mobility of marker-based fluorophores based on a prior survey. Figure 1 The time-series data of the concentrations D(C), D(A), D(G), and D(T) of the four fluorophores or DNA fragments obtained by applying mobility correction to the results of (2). Specifically, it is known that the mobility of the DNA fragment labeled with fluorophore A is only delayed by a time width (arbitrary unit) of 10 relative to the mobility of the DNA fragment labeled with other fluorophores. Therefore, a mobility correction is performed to advance the peak detection time (arbitrary unit) of the DNA fragment labeled with fluorophore A by 10. That is, the time-series data of the concentrations D(C), D(A), D(G), and D(T) of the four fluorophores or DNA fragments obtained by applying mobility correction to the results of (2). Figure 1 The detection time (in any unit) of the peak of A in (2) was corrected from 55 to 45. As a result of this mobility correction, time-series data were obtained in which DNA fragments of various lengths with different base lengths were arranged at approximately equal intervals.

[0082] Figure 1 (4) is based on Figure 1 The results of the base call (3) are obtained. The labeling fluorescent species or terminal base species of the DNA fragment to which each peak belongs can be read in chronological order.

[0083] Figure 2 The method of representing non-patent document 2. Figure 2 (1) represents the time-series data of the three-color fluorescence intensities I(b), I(g), I(y), and I(r) obtained by three-color detection of the emission fluorescence of four phosphors C, A, G, and T through three wavelength bands b, g, and r.

[0084] Figure 2 (1) is from Figure 1 In (1), only the time series data of I(y) is excluded. The time series data of I(b), I(g) and I(r) are the same in the two figures.

[0085] Figure 2W (X, Y) indicates the elements w(XY) of the matrix W of 3 rows and 4 columns. For example, the 4 black bar patterns sequentially indicate w(bC), w(gC), and w(rC) from the left. Similarly, the horizontal bar patterns indicate w(XA) (X is b, g, and r), the diagonal bar patterns indicate w(XG) (X is b, g, and r), and the square bar patterns indicate w(XT) (X is b, g, and r). In addition, w(bY), w(gY), and w(rY) are normalized so that w(bY) + w(gY) + w(rY) = 1 (Y is C, A, G, or T), respectively.

[0086] Specifically, the matrix W is as follows.

[0087] [Equation 5]

[0088]

[0089] As a first step, a case in which only one phosphor emits light at a time is considered using Figure 2 W (X, Y) in (1) of Figure 2 . In (1) of Figure 2 , for the 2 peaks observed on the left and right sides, the ratios of the 3-color fluorescence intensities (I(b) I(g) I(r)) T are close to those of (0.63 0.31 0.06) T and (0.13 0.25 0.63) T , respectively, and thus it can be judged that they are the individual peaks of C and T, respectively. The heights of the peaks, i.e., the concentrations (arbitrary units) of C and T are D(C) = 100 and D(T) = 80, and the times (arbitrary units) of the peaks are 30 and 75, respectively.

[0090] On the other hand, for the peak observed in the center of (1) of Figure 2 , the ratios of the 3-color fluorescence intensities (I(b) I(g) I(r)) T are close to those of any one of (w(bY) w(gY) w(rY)) T (Y is C, A, G, or T). Thus, as a second step, a case in which only 2 phosphors emit light at a time is considered using Figure 2 W (X, Y). Here, a solution in which the peaks of A and G are detected at the same time can be derived. Specifically, when the heights of the peaks, i.e., the concentrations (arbitrary units) of A and G are D(A) = 80 and D(G) = 90, and the times (arbitrary units) of the peaks are 55 and 55, respectively, the difference between the left and right sides of Equation (3) becomes small, and it can be said that the peak observed in the center of (1) of Figure 2 . According to the above results, the following is obtained. Figure 2the time series data of the concentrations D(C), D(A), D(G), D(T) of the four kinds of fluorophores or DNA fragments whose ends are the base types C, A, G, and T.

[0091] With Figure 1 (3) of (2), Figure 2 (3) of (2) is the time series data of the concentrations D(C), D(A), D(G), D(T) of the four kinds of fluorophores or DNA fragments whose ends are the base types C, A, G, and T, obtained by performing the degree of movement correction. The peak detection time (arbitrary units) of A was made 10 earlier, from 55 to 45. Also, in order to unify the peak intervals, the peak detection time (arbitrary units) of G was made 5 later, from 55 to 60. As a result, Figure 2 (3) of (2) becomes the same as Figure 1 (3) of (1). Figure 2 (4) of (2) is the result of the base response, performed in the same way as Figure 1 (4) of (1) is the same as Figure 1 (4) of (2).

[0092] On the other hand, Figure 2 indicates Figure 2 (1) ~ (4) of (2). Figure 2 (2) of (2) is not uniformly determined. The first step is the same as above, but in the second step, another solution is derived, instead of Figure 2 (2) of (1) to obtain Figure 2 (2)'. That is, a solution in which the peaks of A and T are detected at the same time can be obtained. Specifically, when the height of the peak apex, that is, the concentration (arbitrary units) is D(A) = 107, D(G) = 34, and the time (arbitrary units) of the apex of each peak is 55 and 55, the difference between the left and right sides of Equation (3) becomes small, and it can be said that the peak observed in the center of Figure 2 (1) is the peak of A and T. According to the above results, the time series data of the concentrations D(C), D(A), D(G), D(T) of the four kinds of fluorophores or DNA fragments whose ends are the base types C, A, G, and T, of Figure 2 (2)' is obtained.

[0093] However, in Figure 2 (2)', G is not detected, and D(G) = 0. Figure 2 (3)' of (2)' is, in the same way as Figure 1 (3) of (1), the time series data of the concentrations D(C), D(A), D(G), D(T) of the four kinds of fluorophores or DNA fragments whose ends are the base types C, A, G, and T, obtained by performing the degree of movement correction. The peak detection time (arbitrary units) of A was made 10 earlier, from 55 to 45. Figure 2 (4)' of (2)' is obtained according to Figure 2the result of the base call made by (3)'. The same Figure 3 (1) and Figure 3 Independently of (W), Figure 2 (2) of (1), Figure 3 (3) of (1), Figure 3 (4) of (1) and Figure 3 (2)' of (1), Figure 3 (3)' of (1), Figure 3 (4)' of (1) are completely different, and derive different results of the base call. In the model data used here, Figure 3 the result of the base call shown by (4) of (1) is the correct answer, and thus Figure 3 (4) of (1) is the correct result of the base call, Figure 3 (4)' of (1) is the incorrect result of the base call.

[0094] Thus, in the method of Non-Patent Literature 2, it is possible to derive multiple solutions from the same measurement result, the incorrect result of the base call, and there is a risk that an incorrect analysis result can be generally produced.

[0095] [Example 1]

[0096] Figure 3 is a graph showing the DNA sequencing method using the model data of Example 1. Figure 3 (1) is time-series data of 3-color fluorescence intensities I(b), I(g), I(r) obtained by 3-color detection of the luminescence fluorescence of 4 kinds of fluorophores C, A, G, and T in 3 kinds of wavelength bands b, g, and r, and is the same as Figure 3 (1) of (1).

[0097] Figure 3 (5) is time-series data of 3-color fluorescence intensities I(b), I(g), I(r) of the model peak when a single length of the DNA fragment labeled with the fluorophore C is detected 3-color. Figure 3 The vertical axis of (5) of (1) is the fluorescence intensity, and the horizontal axis is the time. Figure 3 The data of (5) of (1) is data showing the time change of the fluorescence intensity ratio between the model peaks of 3-color fluorescence (b, g, and r) when a single length of the DNA fragment labeled with the fluorophore C is detected. Here, the 3-color fluorescence intensity ratio at each time is the matrix W of formula (6) (w(bC) w(gC) w(rC)) T = (0.63 0.31 0.06) T . Also, Figure 3 (6), (7), and (8) of (1) are time-series data of 3-color fluorescence intensities I(b), I(g), I(r) of the model peak when a single length of the DNA fragment labeled with the fluorophores A, G, T is detected 3-color. Figure 3The 3-color fluorescence intensity ratio in (6), (7), and (8) is (0.33 0.56 0.11) T (0.27 0.40 0.33) T (0.13 0.25 0.63) T

[0098] Here, the shape of each model data is a Gaussian distribution, and the variance thereof is consistent with the variance of a single length of DNA fragment observed through an experiment. Furthermore, the shape of each model data is not limited to this example, and can be other structures. Here, the fitting process is performed on the time series data of (1) of FIG. 5, (6), (7), and (8) of FIG. 6, only changing the height and timing of the vertex of the model peak. Figure 3 Figure 3 One example of the fitting process will be described using (1) of FIG. 5. For example, the fitting process is performed from the left end of the data of (1) of FIG. 5. For example, the height (fluorescence intensity) of (5) of FIG. 5 and the central value (electrophoresis time) of the Gaussian distribution are varied, and when the error with the data of (1) of FIG. 5 is less than a predetermined error, it is determined that the fitting is performed. That is, here, in the data of (1) of FIG. 5, the place where the fluorescence intensity and the electrophoresis time are most consistent is searched for. Furthermore, the fitting can be performed while varying the width of the Gaussian distribution. In the case where the fitting is not performed in (5) of FIG. 5, it is determined whether the fitting is performed based on other data (6), (7), and (8) of FIG. 6, or a combination thereof. In this way, the fitting is performed on the data of (1) of FIG. 5 based on any one of (5), (6), (7), and (8) of FIG. 6, or a combination thereof. For example, the error of the fitting process can be calculated based on the difference between the shape of the peak of (1) of FIG. 5 and the shape of the model peak. Various publicly known methods can be applied to the calculation of the error. Furthermore, from the viewpoint of efficiency, it is desirable to sequentially fit from one end of the data of (1) of FIG. 5. This is because the lower side of the adjacent peak intrudes into the peak of a certain fluorescence, and therefore if the fitting is performed from one end of the data, the fitting can be performed also considering the intrusion of the adjacent peak, and is efficient. Figure 3 Figure 3 Figure 3 Figure 3 Figure 3 Figure 1 Figure 3 Figure 1 Figure 3 Figure 3 Figure 3

[0099] Figure 2 ​​​​​​​​​​​​​(2) indicates the result of the fitting process. At this time, the peak shape of C is time series data obtained by making the height of the time series data of I(b) 1 / w(bC) = 1 / 0.63 = 1.59 times. The peak shape of A is time series data obtained by making the height of the time series data of I(g) 1 / w(gA) = 1 / 0.56 = 1.79 times. The peak shape of G is time series data obtained by making the height of the time series data of I(g) 1 / w(gG) = 1 / 0.40 = 2.50 times. The peak shape of T is time series data obtained by making the height of the time series data of I(r) 1 / w(rT) = 1 / 0.63 = 1.59 times.

[0100] These results, Figure 2 (2) becomes the same as Figure 2 (2) of Non-Patent Literature 2. The method of obtaining Figure 2 (3) and (4) after that is the same as the method of obtaining Figure 3 (3) and (4). According to the present embodiment, the step of deriving Figure 3 (2) from Figure 3 (1) is uniquely determined, and therefore, as shown in Figure 4 (4), a correct base response result can be obtained.

[0101] In the method of Non-Patent Literature 2, Figure 3 (1), only the 3-color fluorescence detection intensity ratio of the luminescent fluorescence of each fluorescent substance, that is, the matrix W shown in formula (6), is used when deriving Figure 4 (2) from Figure 3 (1) or Figure 4 (2)'. In the present embodiment shown in Figure 3 , when deriving Figure 4 (2) from Figure 3 (1), in addition to the 3-color fluorescence detection intensity ratio of the luminescent fluorescence of each fluorescent substance, that is, the matrix W shown in formula (6), the peak shape of the luminescent fluorescence of each fluorescent substance, that is, the time change information, is also used. These differences result in a difference in whether the solution is uniquely derived, that is, whether a correct base response result can be obtained.

[0102] Figure 4 indicates an example of a case in which 3-color detection implemented in Figure 4 is further combined with 2-color detection. Figure 4 (1) is time series data of 2-color fluorescence intensities I(b) and I(r) obtained by performing 2-color detection on the luminescent fluorescence of four fluorescent substances C, A, G, and T in two wavelength bands b and r. The time series data of I(g) is excluded from Figure 4 (1).

[0103] Figure 1(5) is the time-series data of the model peak fluorescence intensity I(b) and I(r) when detecting a single-length DNA fragment labeled with fluorophore C using 2-color imaging. Figure 4 The time series data of I(g) were excluded in (5). Similarly, Figure 1 (6), (7), and (8) are time-series data of the model peak fluorescence intensities I(b) and I(r) when detecting single-length DNA fragments labeled with fluorophores A, G, and T using a two-color method. Figure 4 (6), (7), and (8) respectively exclude the time series data of I(g).

[0104] Here, use Figure 4 (5), (6), (7), and (8) only change the height and time of the peak of the model, and perform targeted analysis. Figure 4 Fitting of time series data (1). Figure 5 (2) represents the result of the fitting process. At this point, the peak shape of C is the time series data obtained by making the height of the time series data of I(b) 1 / w(bC) = 1 / 0.63 = 1.59 times. The peak shape of A is the time series data obtained by making the height of the time series data of I(g) 1 / w(gA) = 1 / 0.33 = 3.03 times. The peak shape of G is the time series data obtained by making the height of the time series data of I(r) 1 / w(rG) = 1 / 0.33 = 3.03 times. The peak shape of T is the time series data obtained by making the height of the time series data of I(r) 1 / w(rT) = 1 / 0.63 = 1.59 times.

[0105] These results, Figure 3 (2) becomes with Figure 3 (2) is the same. Subsequent acquisitions Figure 6 Methods (3) and (4) and obtaining Figure 5 The methods in (3) and (4) are the same. According to this embodiment, the unique determination is made from... Figure 7 (1) Export Figure 3 Step (2), therefore, as Figure 8 As shown in (4), the correct base response results can be obtained.

[0106] Figure 8The processing steps and system structure of the present embodiment are shown. The system of Embodiment 1 has an analysis device 510, a computer 520, and a display device 530. The analysis device 510 is, for example, a liquid chromatograph. The computer 520 can be implemented using a general-purpose computer, for example. The processing section of the computer 520 can be implemented as a function of a program executed on the computer. The computer has at least a processor such as a CPU (Central Processing Unit) and a storage section such as a memory. Program codes corresponding to each processing can be stored in the memory, and each program code can be executed by the processor, whereby the processing of the computer 520 is implemented.

[0107] In this structure, the analysis device 510 separates a sample containing a plurality of components labeled with any one of M types of fluorescent substances by chromatography, and acquires first time-series data of fluorescent signals detected at N wavelength bands (M > N) in a state where at least a part of the plurality of components is not completely separated. The first time-series data of fluorescent signals corresponds to N-color detection time-series data 513 of M-color labeled samples described below. The computer 520 has a storage section (e.g., a memory, an HDD, or the like) that stores second time-series data of model fluorescent signals of the plurality of components in advance. The second time-series data of model fluorescent signals of the plurality of components corresponds to N-color detection time-series data 541 of a single peak of each M-color label described below. The computer 520 determines which component of the plurality of components is labeled with which fluorescent substance of the M types by comparing the first time-series data and the second time-series data. The display device 530 displays third time-series data of concentrations of the M types of fluorescent substances that act on the fluorescent signals. The third time-series data of concentrations of fluorescent substances corresponds to M-color labeling time-series data 523 described below. Hereinafter, the above processing is described in more detail.

[0108] First, an M-color labeled sample 501 containing a plurality of components labeled with M types of fluorescent substances is fed to the analysis device 510. Next, in the analysis device 510, separation analysis processing 511 of the plurality of components contained in the sample 501 is performed. The analysis device 510 detects (N-color detection) the luminescent fluorescence of the M types of fluorescent substances in N wavelength bands (M > N), and acquires N-color detection time-series data (fluorescent detection time-series data) 513 of the M-color labeled sample. Here, the plurality of components are not necessarily all well separated. That is, a part of different components labeled with different fluorescent substances is detected by fluorescence in a state of spatiotemporal overlap. The analysis device 510 outputs the N-color detection time-series data 513 to the computer 520.

[0109] Next, as input information, the computer 520 acquires the N-color detection time series data 513, and the N-color detection time series data of the single component labeled with any one of the M kinds of fluorescent substances, i.e., the N-color detection time series data of the single peak of each M-color label 541. The N-color detection time series data of the single peak of each M-color label 541 is, for example, data corresponding to (5), (6), (7), and (8) of Figure 9 . Further, the N-color detection time series data of the single peak of each M-color label 541 is stored in the first database 540 in advance.

[0110] Next, the computer 520 performs the comparison analysis process 521 of the N-color detection time series data 513 and the N-color detection time series data of the single peak of each M-color label 541. As a result, the computer 520 acquires the time series data of the concentration of the M kinds of fluorescent substances detected, i.e., the concentration of the component labeled with the M kinds of fluorescent substances, i.e., the M-color label time series data 523. The M-color label time series data 523 is, for example, data corresponding to (2) of Figure 8 . Finally, the display device 530 performs the display process 531 of the M-color label time series data 523.

[0111] Figure 8 is a diagram specifically showing the comparison analysis of the computer 520 Figure 10 . As the comparison analysis, the computer 520 performs the fitting analysis process 522 of the N-color detection time series data 513 using the N-color detection time series data of the single peak of each M-color label 541. As a result, the computer 520 acquires the M-color label time series data 523, and the difference between the N-color detection time series data 541 and the fitting result thereof, i.e., the fitting error data (or fitting accuracy data) 524. The display device 530 performs the display process 531 of either or both of the M-color label time series data 523 and the fitting error data 524.

[0112] [Example 2]

[0113] Figure 10 shows the processing steps and system configuration in the case where the present application is applied to the electrophoresis analysis of DNA fragments. The plurality of components to be analyzed are nucleic acid fragments of different lengths or different components, and the chromatography can be electrophoresis.

[0114] The analysis device 510 is an electrophoresis device. First, an M-color labeled DNA sample 502 containing a plurality of DNA fragments labeled with M kinds of fluorescent substances is put into the analysis device 510. Next, in the analysis device 510, an electrophoretic separation analysis process 512 of the plurality of DNA fragments contained in the DNA sample is performed. The analysis device 510 detects (N-color detection) the luminescent fluorescence of the M kinds of fluorescent substances in N kinds (M > N) of wavelength bands, and acquires N-color detection time-series data 513. Here, the plurality of DNA fragments are not necessarily all well separated. That is, a part of the different kinds of DNA fragments labeled with different fluorescent substances are detected by fluorescence in a state of spatiotemporal overlap. The analysis device 510 outputs the N-color detection time-series data 513 to the computer 520.

[0115] Next, as input information, the computer 520 acquires the N-color detection time-series data 513, and N-color detection time-series data of a single kind of DNA fragment labeled with any one of the M kinds of fluorescent substances, that is, N-color detection time-series data of a single peak of each M-color label 541. Next, the computer 520 performs a comparison analysis of the N-color detection time-series data 513 and the N-color detection time-series data of a single peak of each M-color label 541. Specifically, the computer 520 performs a fitting analysis process 522 of the N-color detection time-series data 513 using the N-color detection time-series data of a single peak of each M-color label 541. As a result, the computer 520 acquires time-series data of the concentrations of the detected M kinds of fluorescent substances, that is, M-color labeling time-series data 523 of the concentrations of the DNA fragments labeled with the M kinds of fluorescent substances. At the same time, the computer 520 acquires fitting error data (or fitting accuracy data) 524.

[0116] Here, the mobility of electrophoresis of the DNA fragments is affected by the labeled M kinds of fluorescent substances. Therefore, the computer 520 performs a process using different mobility difference data 551 of each M-color label representing the mobility of the labeled M kinds of fluorescent substances in order to reduce the effect thereof. The mobility difference data 551 of each M-color label is stored in advance in a second database 550. The computer 520 uses the mobility difference data 551 of each M-color label to perform a mobility correction process 525 of the M-color labeling time-series data 523, and acquires corrected data (hereinafter referred to as correction data) 526 of the M-color labeling time-series data. The correction data 526 is, for example, data corresponding to (3) of Figure 11 .

[0117] Finally, the display device 530 performs a display process 531 of a part or all of the M-color labeling time-series data 523, the fitting error data (or fitting accuracy data) 524, and the correction data 526.

[0118] [Example 3]

[0119] Figure 11The processing steps and system configuration when the present application is applied to electrophoresis-based DNA sequencing are shown. In Figure 11 N = 3, M = 4. The analysis device 510 is a DNA sequencer.

[0120] First, a 4-color labeled DNA sequencing sample 503 containing 4 kinds of DNA fragments labeled with 4 kinds of fluorescent substances corresponding to 4 kinds of terminal base species by Sanger method modulation using a target DNA as a template is prepared. Then, the 4-color labeled DNA sequencing sample 503 is fed to the analysis device 510. Next, in the analysis device 510, electrophoretic separation analysis processing 512 of the 4 kinds of DNA fragments contained in the DNA sequencing sample is performed. The analysis device 510 detects the luminescent fluorescence of the 4 kinds of fluorescent substances in 3 wavelength bands (3-color detection) and acquires 3-color detection time series data 513. Here, the 4 kinds of DNA fragments are not necessarily all well separated. That is, a part of the different lengths of DNA fragments labeled with different fluorescent substances are detected in a state of spatiotemporal overlap. The analysis device 510 outputs the 3-color detection time series data 513 to the computer 520.

[0121] Next, as input information, the computer 520 acquires the 3-color detection time series data 513, 3-color detection time series data of a single length of DNA fragment labeled with any one of the 4 kinds of fluorescent substances, that is, 3-color detection time series data of each 4-color labeled single peak 541. The 3-color detection time series data of each 4-color labeled single peak 541 is stored in the first database 540 in advance. The computer 520 performs comparison analysis of the 3-color detection time series data 513 and the 3-color detection time series data of each 4-color labeled single peak 541. Specifically, the computer 520 performs fitting analysis processing 522 of the 3-color detection time series data 513 of the 4-color labeled DNA sequencing sample using the 3-color detection time series data of each 4-color labeled single peak 541. As a result, the computer 520 acquires time series data of the concentrations of the detected 4 kinds of fluorescent substances, that is, the concentrations of the 4 kinds of DNA fragments different in terminal base species labeled with the 4 kinds of fluorescent substances, that is, 4-color labeling time series data 523. At the same time, the computer 520 acquires fitting error data (or fitting accuracy data).

[0122] Here, the mobility of the electrophoresis of the DNA fragments is affected by the 4 kinds of labeled fluorescent substances. Therefore, in order to reduce the effect thereof, processing is performed using different 4-color labeling mobility difference data 551 representing the mobility of the 4 kinds of labeled fluorescent substances. The 4-color labeling mobility difference data 551 is stored in the second database 550 in advance. The computer 520 uses the 4-color labeling mobility difference data 551 to perform mobility correction processing 525 on the 4-color labeling time series data 523 and acquires correction data (hereinafter referred to as correction data) 526 of the 4-color labeling time series data.

[0123] And, the computer 520 performs the DNA base sequence determination process 528 using the correction data 526. On the other hand, the computer 520 uses the fitting error data (or fitting accuracy data) 524 to obtain the base sequence determination error data (or base sequence determination accuracy data) 527 for each base for which the DNA base sequence determination is performed.

[0124] Finally, the display device 530 performs the display process 531 of some or all of the 4-color marker timing data 523, the fitting error data (or fitting accuracy data) 524, the correction data 526, the DNA base sequence determination result, and the base sequence determination error data (or base sequence determination accuracy data) 527.

[0125] [Example 4]

[0126] Figure 12 shows the processing steps and system configuration in the case where the structure of Figures 5-9 is replaced with the condition of N = 2, M = 4. The processing steps and system configuration are the same as Figures 5-9 , and thus the explanation is omitted.

[0127] Figures 5-9 is a structural diagram of a capillary electrophoresis device which is an example of the analysis device 510. The capillary electrophoresis device 100 is used as a DNA sequencer, a DNA fragment analysis device, or the like. The sample injection end 2 and the sample elution end 3 of the capillary 1 in which an electrophoretic separation medium containing an electrolyte is filled inside are immersed in the cathode side electrolyte solution 4 and the anode side electrolyte solution 5, respectively. In addition, the cathode electrode 6 is immersed in the cathode side electrolyte solution 4, and the anode electrode 7 is immersed in the anode side electrolyte solution 5. By applying a high voltage between the cathode electrode 6 and the anode electrode 7 through the high voltage power supply 8, electrophoresis is performed.

[0128] The sample injection into the capillary 1 is performed by immersing the sample injection end 2 and the cathode electrode 6 in a sample solution 9, and applying a high voltage between the cathode electrode 6 and the anode electrode 7 for a short time through the high voltage power supply 8. The sample solution 9 contains a plurality of components labeled with a plurality of fluorescent substances. After the sample is injected, the sample injection end 2 and the cathode electrode 6 are immersed in the cathode side electrolyte solution 4 again, and a high voltage is applied between the cathode electrode 6 and the anode electrode 7 to perform electrophoresis.

[0129] Negatively charged components in the sample, such as DNA fragments, are electrophoretically transferred from the sample injection end 2 to the sample dissolution end 3 in the capillary 1, in the electrophoresis direction 10 indicated by the arrow. Due to differences in mobility caused by electrophoresis, the various components contained in the sample solution 9 gradually separate. A laser beam 12 emitted from the laser source 11 is irradiated at a position (laser beam irradiation position 15) where the components have electrophoresed a certain distance in the capillary 1. This causes the components that sequentially pass through the laser beam irradiation position 15 to emit fluorescence 13 caused by various labeled phosphors. The fluorescence 13, which changes over time accompanying electrophoresis, is measured by a multicolor detection device 14 that detects light in multiple wavelength bands. Figures 13-18 In this paper, only one capillary 1 is depicted, but a multi-capillary electrophoresis apparatus that uses multiple capillary 1s for parallel electrophoretic analysis can also be used.

[0130] Figures 9-12 This is an example of a multicolor detection device representing a multicapillary electrophoresis apparatus. Multiple capillaries 1 are arranged at equal intervals on the same plane, with laser beam irradiation positions 15. Figure 13 The left figure is a cross-sectional view perpendicular to the major axis of the multiple capillaries 1. Figure 4 The right figure is a cross-sectional view parallel to the major axis of any capillary 1.

[0131] A laser beam 12 is irradiated along the plane of the arrangement of multiple capillaries 1. Thus, the laser beam 12 simultaneously irradiates multiple capillaries 1. The fluorescence 13 emitted by each capillary 1 is focused in parallel by a separate lens 16. Each focused beam is directly incident on a two-dimensional color sensor 17. The two-dimensional color sensor 17 is an RGB color sensor capable of detecting three colors in three wavelength bands. The fluorescence 13 emitted from each capillary 1 forms light spots at different positions on the two-dimensional color sensor 17, thus enabling independent three-color detection of each.

[0132] Figure 13 An example illustrating the structure of a computer 520. For example... Figure 14 As shown, computer 520 is connected to analysis device 510. Computer 520 does not only perform… Figure 14 The data analysis described can also be used to control the analysis device 510. Figure 4 In this context, the display device 530 and databases 540 and 550 are depicted outside the computer 520, but they may also be contained within the computer 520.

[0133] Computer 520 includes a CPU (processor) 1201, memory 1202, display unit 1203, HDD 1204, input unit 1205, and network interface (NIF) 1206. The display unit 1203 is, for example, a monitor, but can also be used as a display device 530. The input unit 1205 is, for example, an input device such as a keyboard or mouse. The user can set data parsing conditions and control conditions for the analysis device 510 via the input unit 1205. The N-color detection timing data 51 output from the analysis device 510 is sequentially stored in memory 1202.

[0134] HDD1204 may also include databases 540 and 550. Additionally, HDD1204 may include programs for performing fitting and analysis processing, mobility correction processing, and DNA base sequence determination processing by computer 520. The program code corresponding to each processing step can be stored in memory 1202, and CPU 1201 executes the program code to implement the processing by computer 520.

[0135] For example, the N-color detection timing data 541 of the single peak of each M-color marker stored in HDD 1204 is stored in memory 1202. CPU 1201 uses N-color detection timing data 513 and N-color detection timing data 541 of the single peak of each M-color marker to perform comparison and analysis processing. Display unit 1203 displays the analysis result. Alternatively, the analysis result can be compared with information on the network via NIF 1206.

[0136] [Example 5]

[0137] Figure 14 Indicates the use of Figure 4 The illustrated processing steps and system structure are examples of DNA sequencing.

[0138] As a 4-color labeled DNA sequencing sample, the results obtained from 3500 / 3500xLS Sequencing Standards, BigDye Terminator v3.1 (Thermo Fisher Scientific) were dissolved in 300 μL of formamide. The sample contained four DNA fragments labeled with terminal bases C, A, G, and T using four different fluorophores: dROX (maximum emission wavelength 618 nm), dR6G (maximum emission wavelength 568 nm), dR110 (maximum emission wavelength 541 nm), and dTAMRA (maximum emission wavelength 595 nm).

[0139] There are four capillary tubes 1, with an outer diameter of 360 μm, an inner diameter of 50 μm, a total length of 56 cm, and an effective length of 36 cm. POP-7 (Thermo Fisher Scientific) was used as the polymer solution for electrophoresis separation. During electrophoresis, capillary tube 1 was adjusted to 60°C and an electric field strength of 182 V / cm. Sample injection was performed by injecting an electric field at a strength of 27 V / cm for 8 seconds. The laser beam 12 had a wavelength of 505 nm and an output of 20 mW. A long-pass filter was used between lens 16 and the two-dimensional color sensor to shield the laser beam 12.

[0140] Figure 14 (1) and Figure 4 (1) corresponds to the two-color detection timing data obtained by detecting the fluorescence 13 emitted from one of the four capillaries 1 via an RGB color sensor during electrophoresis. In this embodiment, the RGB color sensor can perform three-color detection in three wavelength bands r, g, and b corresponding to R (red), G (green), and B (blue). However, the fluorescence emitted by the four phosphors is almost undetectable in the b wavelength band, and is only detected in the r and g wavelength bands. Therefore, Figure 14 Figure (1) shows the temporal sequence of the fluorescence intensities I(r) and I(g) of these two colors. The horizontal axis represents the elapsed time (electrophoresis time) from the start of electrophoresis in seconds, and the vertical axis represents the fluorescence intensity in arbitrary units.

[0141] Figure 13 Figure (1) corresponds to Figure 4 (5), and represents the time-series data of the model peak fluorescence intensities I(g) and I(r) for two-color detection of a single-length DNA fragment with terminal base type C labeled with the fluorophore dROX. Here, the ratio of the two-color fluorescence intensities at each time point is (w(gC)w(rC)). T =(0.02 0.98) T .same, Figure 13 (2) and Figure 13 (6) corresponds to the time-series data of the model peak fluorescence intensities I(g) and I(r) of a single-length DNA fragment of type A terminal base labeled with fluorescein dR6G, when subjected to two-color detection. The ratio of the two-color fluorescence intensities is (w(gA)w(rA)). T =(0.50 0.50) T . Figure 14 (3) and Figure 13(7) corresponds to the time-series data of the model peak fluorescence intensities I(g) and I(r) of a single-length DNA fragment with terminal base type G labeled with fluorophore dR110 during two-color detection. The ratio of the two-color fluorescence intensities is (w(gG)w(rG)). T =(0.71 0.29) T . Figure 13 (4) and Figure 14 (8) corresponds to the time-series data of the model peak fluorescence intensities I(g) and I(r) of a single-length DNA fragment with terminal base type T labeled with dTAMRA. The ratio of the two fluorescence intensities is (w(gT)w(rT)). T =(0.16 0.84) T The peak shapes of each model follow a Gaussian distribution with an offset of 0, and the standard deviation is set to 1 second. This corresponds to the shape of the peaks of a single-length DNA fragment measured under these electrophoresis conditions.

[0142] sequentially Figure 13 The two-color detection time series data of the peak values ​​of the four models shown in (1) to (4) are compared with those obtained by... Figure 13 The two-color detection time series data obtained from the electrophoretic analysis of (1) were fitted. In the fitting, while keeping the Gaussian distribution of the two-color fluorescence intensity and standard deviation, only the center value (electrophoresis time) and height (fluorescence intensity) of the Gaussian distribution were changed, and the center value and height were determined to minimize the error with the two-color detection time series data.

[0143] Figure 14 (2) is for Figure 13 The peak value of g observed on the far left of (1) is fitted with high precision. Figure 13 The result of the peak value of the model G(dR110) of (3). Figure 13 (3) is for Figure 13 The peak value of r, which is the right neighbor of the peak value of g in (1), is fitted with high precision. Figure 13 The result of the peak value of the model T(dTAMRA) of (4). Figure 13 (4) is for Figure 13 The peak value of the right neighbor of the peak value of r in (1) above is fitted with high precision. Figure 13 The results of the model peak of C(dROX) of (4). Figure 14 (5) is Figure 13 The overlapping 2-color detection time series data of (2), (3), and (4) are about to be... Figure 13 The time series data obtained by combining I(g) and I(r) of (2), (3), and (4). Figure 13 (5) can faithfully reproduce Figure 15 The corresponding part of (1).

[0144] The fitting error and accuracy are evaluated as follows: The fitting error is calculated by dividing the standard deviation of the difference between the fitted model peak and the measured two-color detection time series data within a 2-second wide interval (twice the standard deviation of the Gaussian distribution) centered on the peak time of the fitted model by the maximum measured two-color detection fluorescence intensity at the peak time of the model peak. The fitting accuracy is calculated by subtracting the fitting error from 1. The fitting accuracy is 100% when the fitted data perfectly matches the measured data, decreasing as the deviation increases. It is 0% if the deviation is equal to or greater than the maximum measured two-color detection fluorescence intensity. In this embodiment, the fitting error and accuracy are defined as described above, but other definitions are also possible.

[0145] Figure 13 The individual fitting accuracy of the T model peak in (3) is only 43.41%, but makes Figure 15 When the model peaks of G, T, and C in (5) overlap, the fitting accuracy of the model peak of T reaches 96.56%. This indicates that in the case of spatiotemporal overlap, it is important to fit together with the coexisting model peaks before and after the model peaks, rather than fitting with only a single model peak.

[0146] Similarly, for Figure 13 The peak values ​​of all g and r in (1) are sequentially... Figure 15 The two-color detection time series data of the peak values ​​of the four models shown in (1) to (4) were fitted. Figure 13 (6) is the two-color detection time series data obtained by making the peak values ​​of each of the above fitted models coincide. Figure 15 (6) can faithfully reproduce Figure 13 The entirety of (1).

[0147] Figure 17 (1) is a Figure 15 The combined results of I(g) and I(r) of (2) show the time-series data of the concentration of fluorophore dR110, i.e. the concentration of DNA fragments with terminal base type G. Figure 17 (2) is for Figure 17 The combined results of I(g) and I(r) of (3) show the time-series data of the concentration of the fluorophore dTAMRA, i.e. the concentration of DNA fragments with terminal base type T. Figure 15 (3) is for Figure 15 The combined results of I(g) and I(r) from (4) show the time-series data of the concentration of the fluorophore dROX, i.e., the concentration of DNA fragments with terminal base type C. Similarly, Figure 15 (5) indicates that with Figure 15The time-series data of the concentrations of DNA fragments with peak-related terminal base types C, A, G, and T in each fitted model (6).

[0148] Figure 15 against Figure 15 The peak values ​​of each model in (5) were summarized according to the electrophoresis time sequence, including the types of terminal bases, fitting accuracy, and QV (quality value). Here, when the fitting accuracy is set to S, QV is calculated according to QV = -10*Log(1-S). During the 200-second period from 1100 seconds to 1300 seconds of electrophoresis, 76 DNA fragments of different lengths for each base were measured, and the types of terminal bases could be determined. The fitting accuracy was 94.72% on average, and the QV was 13.95 on average, achieving high-precision fitting. In addition, it can also be calculated by computer 520. Figure 18 Data displayed by display device 530 Figure 15 The data.

[0149] Figure 18 (6) is based on the difference in mobility caused by the different four types of phosphors mentioned above. Figure 17 (5) Corrected time-series data of the concentrations of DNA fragments with terminal base types C, A, G, and T obtained by implementing mobility correction. Specifically, making Figure 18 The electrophoresis time of the model peak of the terminal base type G of (5) is delayed by 1.6 seconds, making Figure 18 The electrophoresis time for the median value of the model peak for terminal base type T in (5) is advanced by 1.1 seconds, and no correction is made for the model peaks for terminal base types C and A. The above results show the peaks of DNA fragments of different lengths arranged approximately at equal intervals from shortest to longest. The characteristic is that: in Figure 18 In (5), due to the difference in mobility caused by different fluorophores, DNA fragments with terminal base types G and T were measured in reverse (i.e., the longer DNA fragments exceeded the shorter DNA fragments), but... Figure 16 In (6), it was corrected very well.

[0150] Figure 16 against Figure 15 The peak values ​​of each model in (6) were summarized according to the corrected electrophoresis time order, including the types of terminal bases, fitting accuracy, and QV. Figure 4 right Figure 16 The data was sorted, and the same values ​​were used. Figure 15The arrangement of the terminal bases gives the DNA sequencing result (base response result). The fitting accuracy and QV of each base are indicators related to the accuracy of base arrangement determination, but they are not the same. Generally, the accuracy of base arrangement determination is expected to be greater than the fitting accuracy. In practice, this DNA sequencing result is 100% correct. Alternatively, it can be calculated by a computer (e.g., 520). Figure 4 The data is displayed by display device 530. Figure 16 The data.

[0151] Figure 15 The DNA sequencing steps in this embodiment are summarized. Figure 4 (1) and Figure 16 (1) is the same as the two-color detection time series data obtained by electrophoretic analysis, and is the same as Figure 16 (1) corresponds to. Figure 4 (2) and Figure 16 (5) is the same as the temporal data of the concentration of DNA fragments with terminal base types C, A, G, and T, and is consistent with ​ (2) Corresponding. ​ (3) and ​ Similar to (6), this is the corrected data obtained by applying mobility correction to the time-series data of DNA fragment concentrations of terminal base types C, A, G, and T, and is consistent with... ​ (3) corresponds. Finally, ​ (4) is based on ​ The results of (3) were used to perform base response, and the results were compared with those of (3). ​ (4) corresponds to. ​ (4) The base response results are consistent with the base arrangement of the target DNA.

[0152] In summary, the methods and effects of Examples 1 to 5 above are described. Based on the above examples, an analytical method can be provided as follows: when the fluorescence emitted from M fluorophores has spectral and spatiotemporal overlap, the M components can be identified and detected by N-color detection in the wavelength bands of the M types (M>N). Hereinafter, according to Non-Patent Literature 2, a unit for N=3-color detection of the fluorescence emitted by M=4 fluorophores using an electrophoresis DNA sequencer will be described.

[0153] The analysis device 510 performs 3-color detection of the luminescence fluorescence of the 4 types of fluorescent substances C, A, G, and T in the 3 wavelength bands b, g, and r. Until the 3-color fluorescence intensities of formula (3) are obtained for each time, it is the same as Non-Patent Literature 2. Here, the HDD (storage section) 1204 of the computer 520 stores 4 types of time-series data, i.e., 4 types of model peak data, when 3-color detection is performed on a single length of a DNA fragment labeled with any one of the 4 types of fluorescent substances C, A, G, and T. The 3-color fluorescence intensity ratio of the model peak data of the DNA fragment labeled with the fluorescent substance Y (C, A, G, or T) is (w(bY) w(gY) w(rY)) T That is, the model peak data contains information equivalent to the matrix W. In addition to this, the model peak data contains the shape of the peak, i.e., the time change information.

[0154] In Non-Patent Literature 2, for each time, one or two types of fluorescent substances that emit luminescence fluorescence are selected using the matrix W, and their concentrations are calculated. In the above-described embodiment, the computer 520 performs a fitting analysis process on the time-series data of the 3-color fluorescence intensities shown in formula (3) using the 4 types of model peak data. In a case where the luminescence fluorescence from the 4 types of fluorescent substances has spectral overlap and spatiotemporal overlap, the computer 520 can also perform the fitting analysis process. For example, it can also be a case where 3 or more types of fluorescent substances emit luminescence fluorescence at once. The set of the model peak data of C, A, G, and T that constitutes the fitting result represents the time-series data of the concentrations D(C), D(A), D(G), and D(T) of C, A, G, and T. That is, regardless of whether color conversion is performed, the time-series data of the 3-color fluorescence intensities can be used to obtain the time-series data of the 4 types of fluorescent substance concentrations, i.e., the 4 types of base species concentrations, which corresponds to formula (2) in Non-Patent Literature 1. The above-described embodiment is different from Non-Patent Literature 2 in that it flexibly uses information on the time change of the 4 types of fluorescent substance concentrations. Then, the computer 520 performs steps equivalent to steps (3) and (4) in Non-Patent Literature 1, and can obtain the result of the base response.

[0155] According to the above-described embodiment, in a state where the luminescence fluorescence from M types of fluorescent substances has spectral overlap and spatiotemporal overlap, the time-series data of N-color fluorescence intensities obtained by N-color detection in N wavelength bands (M > N) is fitted using model peak data of the M types of fluorescent substances, whereby the time-series data of the concentrations of the M types of fluorescent substances, i.e., the M types of components, can be analyzed.

[0156] In order to analyze by recognizing and detecting M kinds of components by N color detection of N wavelength bands in a state where the luminescence fluorescence from M kinds of phosphors has spectral overlap and spatiotemporal overlap, it has been necessary in the past that M ≤ N. According to the above-described embodiment, even if M > N, it is possible to recognize and detect M kinds of components as well. That is, it has the effect that the same analysis can be achieved by a more simple, small, and low-cost device structure. For example, it is possible to perform analysis by recognizing and detecting M = 4 or more kinds of components labeled with M = 4 or more kinds of phosphors by N = 3 color detection using an RGB color sensor that has rapidly developed by high performance and low cost. As a result of the above, it is possible to perform high-precision and low-cost analysis based on multi-color detection. For example, it is possible to perform N = 3 color detection using a low-cost RGB color sensor while performing electrophoretic separation of M = 4 kinds of DNA fragments labeled with M = 4 kinds of phosphors by a Sanger reaction, whereby even if different lengths of DNA fragments labeled with M = 4 kinds of phosphors are interlaced with each other, it is possible to acquire time-series data of the concentrations of the M = 4 kinds of DNA fragments, and it is possible to perform DNA sequencing well.

[0157] The present application is not limited to the above-described embodiments, and includes various modifications. The above-described embodiments are described in detail in order to easily understand the present application, and are not limited to necessarily having all of the structures described. Also, a part of the structure of an embodiment can be replaced with the structure of another embodiment. Also, the structure of an embodiment can be added with the structure of another embodiment. Also, other structures can be added, deleted, or replaced with respect to a part of the structure of each embodiment.

[0158] Also, each structure, function, processing section, processing unit, and the like of the above-described computer 520 can be implemented by hardware, for example, by designing with an integrated circuit. Also, each structure, function, and the like can be implemented by software by a processor interpreting and executing a program that realizes each function. A program, table, file, and the like that realize each function can be stored in a non-transitory computer readable medium. For example, a floppy disk, CD-ROM, DVD-ROM, hard disk, optical disk, magneto-optical disk, CD-R, magnetic tape, non-volatile memory card, ROM, and the like can be used as the non-transitory computer readable medium.

[0159] In the above-described embodiments, control lines and information lines that are considered necessary for explanation are shown, and are not limited to necessarily showing all of the control lines and information lines on the product. All of the structures can be connected to each other.

[0160] Symbol explanation:

[0161] c: fluorescent signal detected in wavelength band c; g: fluorescent signal detected in wavelength band g; y: fluorescent signal detected in wavelength band y; r: fluorescent signal detected in wavelength band r; C: fluorophore C or base type C; A: fluorophore A or base type A; G: fluorophore G or base type G; T: fluorophore T or base type T; 1: capillary; 2: sample injection end; 3: sample elution end; 4: cathode-side electrolyte solution; 5: anode-side electrolyte solution; 6: cathode electrode; 7: anode electrode; 8: high-voltage power supply; 9: sample solution; 10: electrophoresis direction; 11: laser light source; 12: laser beam; 13: fluorescence; 14: multicolor detection device; 15: laser beam irradiation position; 16: lens; 17: two-dimensional color sensor; 510: analysis device; 520: computer; 530: display device; 540, 550: database.

Claims

1. A capillary electrophoresis apparatus, comprising: The sample contains more than four different fluorescent agents; A capillary tube, used for electrophoretic analysis of the sample; A light source that irradiates a laser beam onto the capillary; An optical system that focuses the fluorescence emitted from the luminous point of the capillary due to irradiation by the laser beam; and A sensor that measures the image of the light-emitting point generated by the optical system. Its features are, The sensor is an RGB color sensor.

2. The capillary electrophoresis apparatus according to claim 1, characterized in that, The DNA of the sample was sequenced using the electrophoretic analysis.

3. The capillary electrophoresis apparatus according to claim 1, characterized in that, The DNA fragments of the sample are analyzed by the electrophoresis analysis.

4. The capillary electrophoresis apparatus according to any one of claims 1 to 3, characterized in that, The following steps are performed: the time-series data of the RGB color sensor signal obtained during the electrophoretic analysis of the sample is compared with the time-series data of the RGB color sensor signal obtained during the electrophoretic analysis of each monomer of the various phosphors.

5. The capillary electrophoresis apparatus according to any one of claims 1 to 3, characterized in that, The following steps are performed: the time-series data of the RGB color sensor signal obtained during the electrophoretic analysis of the sample are represented by combining the time-series data of the RGB color sensor signal obtained during the electrophoretic analysis of the individual monomers of the various phosphors.

6. A capillary electrophoresis apparatus, comprising: Multiple samples, each containing more than four different fluorescent agents; Multiple capillaries are used for electrophoretic analysis of the various samples. A capillary array that arranges the measurable portions of the plurality of capillaries on the same plane; A light source that irradiates a laser beam onto the capillary array; An optical system that focuses the fluorescence emitted from the light-emitting points of the plurality of capillaries due to irradiation by the laser beam; and A region sensor that measures an image of fluorescence from multiple emission points generated by the optical system. Its features are, The area sensor is an RGB color sensor.

7. The capillary electrophoresis apparatus according to claim 6, characterized in that, The DNA of the various samples was sequenced using the electrophoretic analysis.

8. The capillary electrophoresis apparatus according to claim 6, characterized in that, The DNA fragment analysis of the various samples was performed using the electrophoretic analysis.

9. The capillary electrophoresis apparatus according to any one of claims 6 to 8, characterized in that, The following steps are performed: the time-series data of the signals of the RGB color sensors obtained during the electrophoretic analysis of the various samples are compared with the time-series data of the signals of the RGB color sensors obtained during the electrophoretic analysis of the individual monomers of the various phosphors.

10. The capillary electrophoresis apparatus according to any one of claims 6 to 8, characterized in that, The following steps are performed: the time-series data of the signals of the RGB color sensors obtained during the electrophoretic analysis of the various phosphors are represented by combining the time-series data of the signals of the RGB color sensors obtained during the electrophoretic analysis of the various samples.

11. The capillary electrophoresis apparatus according to claim 6, characterized in that, The optical system includes multiple lenses that individually focus the fluorescence emitted from the light-emitting points of the multiple capillaries. The multiple fluorescent rays from the individually focused light are directly incident on the RGB color sensor.

Citation Information

Patent Citations

  • Chip component take-in apparatus

    US20020011398A1