Genetic analysis device and genetic analysis method

The gene analysis apparatus and method enhance the accuracy of signal detection in electrophoresis data by segmenting and analyzing fluorescent intensity patterns, effectively addressing errors in distinguishing signal and non-signal sections.

GB2641677APending Publication Date: 2025-12-10HITACHI HIGH TECH CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
GB2025012760
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-12
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Existing methods for determining base sequences from electrophoresis data face challenges in accurately distinguishing signal and non-signal sections, particularly with low-intensity signals, leading to erroneous determinations due to high-intensity signals or insufficient changes in local signal intensity.

Method used

A gene analysis apparatus and method that divides time-series electrophoresis data into sections, generating features based on the appearance frequency of maximum, minimum, and flat portions of fluorescent intensity data, and uses these features to detect signal areas with high accuracy by determining section characteristics.

Benefits of technology

The method enables precise detection of signal sections, even with low-intensity signals and varying intensities, improving detection accuracy and reducing errors in base sequence analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This genetic analysis device comprises an acquisition unit that acquires time series data indicating the result of electrophoresis of a sample; and an analysis unit that analyzes a base sequence of the sample from the time series data. The time series data includes multiple sets of fluorescence intensity data corresponding to multiple bases. The analysis unit divides the time series data into multiple sections, generates, for each of the multiple sets of fluorescence intensity data, a feature amount indicating an emergence frequency of at least one of a local maximum portion, a local minimum portion, and a flat portion of the fluorescence intensity data in each of the sections, determines, from among the multiple feature amounts generated for the multiple sets of fluorescence intensity data, a section feature amount on the basis of the magnitude relationship of the feature amounts, and detects, by using the section feature amount, a signal region which is an analysis target region of the base sequence in the time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Title of Invention: GENETIC ANALYSIS DEVICE AND GENETIC ANALYSIS METHOD Technical Field

[0001] The present invention relates to a gene analysis apparatus and a gene analysis method. Background Art

[0002] In the related art, in order to determine a base sequence, there is a technique described in WO2008 / 050426 (PTL 1) . PTL 1 describes that "a base sequence can be accurately analyzed even from migration data including a degraded part" and that "the base sequence of a nucleic acid is determined including the following steps (A) to (C) in that order. (A) a basic peak extracting step of extracting basic peaks from electrophoretic data including respective peaks of four kinds of bases obtained by electrophoresing a sample nucleic acid; (B) a condition setting step of setting a basic peak at search starting point from which a search is started and a standard peak interval based on time-series data including the extracted basic peaks, and (C) a step of determining a base sequence by starting a search from the basic peak at search starting point to sequentially scan intervals between adjacent basic peaks in temporal forward and backward directions in the time-series data and comparing a distance of each interval between adjacent basic peaks with a standard peak interval to add an interpolation peak to a peak-missing section". Citation List Patent Literature

[0003] PTL 1: WO2008 / 050426 Summary of Invention Technical Problem

[0004] In the related art, there is a problem in accuracy of identifying a signal section (signal area) and a non-signal section (non-signal area). In particular, detection accuracy of the signal area including a low-intensity signal is not sufficient. For example, in a method of obtaining a threshold for distinguishing between a signal and a non-signal based on a distribution of overall signal intensity, a threshold may increase due to an influence of a high-intensity signal, and a section including the low-intensity signal may be erroneously determined as the non-signal section. In a method of finding a change (rising or falling) in local signal intensity and distinguishing between the signal and the non-signal, it is difficult to distinguish the low-intensity signal since a change in intensity is relatively small. In a method of determining the signal section by extracting a periodic component of the signal, an erroneous determination occurs when periodicity close to that of the signal is observed in the non-signal area.

[0005] Therefore, an object of the invention is to detect a signal section (signal area) with high accuracy from timeseries data indicating an electrophoresis result. Solution to Problem

[0006] In order to achieve the above-described object, a representative gene analysis apparatus of the invention includes: an acquisition unit configured to acquire timeseries data indicating an electrophoresis result of a sample; and an analysis unit configured to analyze a base sequence of the sample from the time-series data, in which the time-series data includes a plurality of pieces of fluorescent intensity data corresponding to a plurality of bases, and the analysis unit divides the time-series data into a plurality of sections, generates, for each of the plurality of pieces of fluorescent intensity data, a feature indicating an appearance frequency of at least one of a maximum portion, a minimum portion, and a flat portion of the fluorescent intensity data in each section, determines, based on a magnitude relation of the plurality of features generated for the plurality of pieces of fluorescent intensity data, a section feature from the plurality of features, and detects a signal area that is an analysis target area of the base sequence in the time-series data by using the section feature. A representative gene analysis method of the invention includes: acquiring time-series data indicating an electrophoresis result of a sample and including a plurality of pieces of fluorescent intensity data corresponding to a plurality of bases; dividing the time-series data into a plurality of sections; generating, for each of the plurality of pieces of fluorescent intensity data, a non-signal feature of the fluorescent intensity data in each section based on an appearance frequency of at least one of a maximum portion, a minimum portion, and a flat portion of the fluorescent intensity data in each section; determining, based on a magnitude relation of the plurality of features generated for the plurality of pieces of fluorescent intensity data, a section feature from the plurality of features; and detecting a signal area that is an analysis target area of the base sequence in the time-series data by using the section feature. Advantageous Effects of Invention

[0007] According to the invention, a signal section (signal area) can be detected with high accuracy from time-series data indicating an electrophoresis result. Problems, configurations, and effects other than those described above will become apparent by the following description of an embodiment. Brief Description of Drawings

[0008] [FIG. 1] FIG. 1 is a configuration example of a gene analysis apparatus according to Embodiment 1. [FIG. 2] FIG. 2 is a configuration example of an electrophoresis device according to Embodiment 1. [FIG. 3] FIG. 3 is a flowchart illustrating the summary of a process that is executed by the gene analysis apparatus according to Embodiment 1. [FIG. 4] FIG. 4 is a flow of an electrophoresis process of an actual sample. [FIG. 5] FIG. 5 is a flow of base calling. [FIG. 6] FIG. 6 is a flow of signal section detection. [FIG. 7] FIG. 7 is a diagram illustrating characteristics of a non-signal section. [FIG. 8] FIG. 8 is a diagram illustrating characteristics of a signal section. [FIG. 9] FIG. 9 is a diagram illustrating generation of non-signal features based on shape patterns. [FIG. 10] FIG. 10 is a diagram illustrating non-signal feature determination in a section. [FIG. 11] FIG. 11 is a diagram illustrating signal boundary determination (part 1). [FIG. 12] FIG. 12 is a diagram illustrating the signal boundary determination (part 2) . [FIG. 13] FIG. 13 is a diagram illustrating a case where a threshold is determined based on a distribution of non-signal features in a section. [FIG. 14] FIG. 14 is a diagram illustrating a case where non-signal features are used in combination with other features (part 1). [FIG. 15] FIG. 15 is a diagram illustrating a case where non-signal features are used in combination with other features (part 2). [FIG. 16] FIG. 14 is a diagram illustrating a case where non-signal features are used in combination with other features (part 3) . [FIG. 17] FIG. 17 is a diagram illustrating a case where non-signal features are used in combination with other features (part 4). [FIG. 18] FIG. 18 is a diagram illustrating a case where non-signal features are used in combination with other features (part 5). [FIG. 19] FIG. 19 is a diagram illustrating a case where an analysis section is corrected by editing fluorescent intensity data. Description of Embodiments

[0009] Hereinafter, an embodiment will be described with reference to the drawings. [Embodiment 1]

[0010] FIG. 1 is a diagram illustrating a configuration example of a gene analysis apparatus 101 according to Embodiment 1. The gene analysis apparatus 101 includes an electrophoresis device 105 and a data analysis device 112. The electrophoresis device 105 and the data analysis device 112 are communicably connected using a communication cable.

[0011] The data analysis device 112 includes a central control unit 102, a storage unit 104, and a user interface unit 103. The central control unit 102 executes control and data processing of the electrophoresis device 105. The central control unit 102 is, for example, a central processing unit (CPU) and a graphics processing unit (GPU). The storage unit 104 stores a program to be executed by the central control unit 102, setting information of the electrophoresis device 105, information used for various processes, and the like. The storage unit 104 is, for example, a memory. The user interface unit 103 is an interface that connects an input device and an output device, or an interface that is connected to an external device via a network. The data analysis device 112 presents information to a user or receives information input by the user via the user interface unit 103.

[0012] By executing the program stored in the storage unit 104, the central control unit 102 operates as a sample information setting unit 106, an electrophoresis device control unit 108, a fluorescent intensity calculation unit 110, and a base calling unit 107. In the following description, when a functional unit is used as a subject to describe a process, it can be considered that the central control unit 102 executes the program.

[0013] The sample information setting unit 106 is a setting unit that sets information related to a sample. The electrophoresis device control unit 108 is a control unit that controls electrophoresis of the sample performed by the electrophoresis device 105. The fluorescent intensity calculation unit 110 is an acquisition unit that acquires, from the electrophoresis device 105, time-series data indicating an electrophoresis result. The time-series data includes a plurality of pieces of fluorescent intensity data corresponding to a plurality of bases. The base calling unit 107 is an analysis unit that analyzes a base sequence of the sample from time-series data. The base calling unit 107 includes an analysis section detection unit 109.

[0014] The analysis section detection unit 109 divides the time-series data into a plurality of sections, and generates, for each piece of fluorescent intensity data, a non-signal feature indicating a non-signal based on an appearance frequency of maxima, minima, and flat segments in the fluorescent intensity data in each section. This feature is set to a large value, resembling that of the non-signal. Alternatively, a value that increases as the appearance frequency decreases, such as a reciprocal of the appearance frequency or a value obtained by subtracting the appearance frequency from a fixed value, may be generated as a signal feature. In this case, the signal feature is set to a large value, resembling that of a signal. Hereinafter, an embodiment using the above-described non-signal feature will be described, but the present embodiment is similarly applicable to a case where the signal feature is used. The analysis section detection unit 109 determines, as the nonsignal feature of the section, a smallest value among a plurality of non-signal features generated for the plurality of pieces of fluorescent intensity data. When the signal feature is used, a maximum signal feature is determined as the signal feature of the section. Then, a signal section of the time-series data is detected using the signal feature of the section. The signal section (signal area) is a section in the time-series data in which there is a change in fluorescent intensity due to presence of a base. The non-signal section (non-signal area) is a section in the time-series data in which there is no change in the fluorescent intensity due to the presence of the base.

[0015] The electrophoresis device 105 electrophoreses a sample (DNA fragment) to acquire migration data. The migration data is time-series data of a luminance of the DNA fragment labeled with a fluorescent dye.

[0016] Here, a configuration of the electrophoresis device 105 will be described. FIG. 2 is a diagram illustrating a configuration example of the electrophoresis device 105 according to Embodiment 1.

[0017] The electrophoresis device 105 includes a detection unit 216, a thermostatic bath 218, a transport device 225, a high-voltage power supply 204, a first ammeter 205, an anode-side electrode 211, a second ammeter 212, a capillary array 217, and a pump mechanism 203.

[0018] The capillary array 217 is a replacement member including a plurality of (for example, eight) capillaries 202, and includes a load header 229, the detection unit 216, and a capillary head 233. Along with breakage or deterioration of quality of the capillaries 202, the capillary array 217 can be replaced with a new one.

[0019] The capillary 202 includes a glass tube having an inner diameter of several tens to several hundreds of microns and having an outer diameter of several hundreds of microns, and a surface thereof is coated with polyimide to improve strength. However, a light irradiation unit that emits a laser beam has a structure where a polyimide coating is removed such that the emitted light easily leaks from the inside to the outside. The capillary 202 is filled with a separation medium for applying a difference in migration speed during electrophoresis. As the separation medium, both a flowable medium and a non-flowable medium can be used but in Embodiment 1, a flowable polymer is used.

[0020] The high-voltage power supply 204 applies a high voltage to the capillary 202. The first ammeter 205 detects a current generated from the high-voltage power supply 204. The second ammeter 212 detects a current flowing through the anode-side electrode 211.

[0021] An optical detection unit that detects information light acquired from the sample includes a light source 214 that emits excitation light to the detection unit 216, an optical detector 215 that detects the emitted light in the detection unit 216, and a diffraction grating 232. The detection unit 216 is a member that acquires information depending on the sample.

[0022] When the sample in the capillary 202 separated by electrophoresis is detected, by emitting the excitation light from the light source 214 to the detection unit 216, fluorescence having a wavelength depending on the sample is generated as the information light. Further, the diffraction grating 232 disperses the information light in a wavelength direction, and the optical detector 215 detects the dispersed information light to analyze the sample.

[0023] Each capillary cathode end 227 is fixed through a metallic hollow electrode 226, and a tip of the capillary 202 protrudes from the hollow electrode 226 by about 0.5 mm. All of the hollow electrodes 226 provided in the capillaries 202 are integrated and mounted on the load header 229. All of the hollow electrodes 226 are electrically connected to the high-voltage power supply 204 mounted on a device main body and, when it is necessary to apply a voltage, for example, during electrophoresis or sample introduction, functions as a cathode electrode.

[0024] Capillary end portions (other end portion) opposite to the capillary cathode ends 227 are bundled into one by the capillary head 233. The capillary head 233 can be connected to a block 207 in a pressure-resistant confidential manner. A high voltage generated by the high-voltage power supply 204 is applied between the load header 229 and the capillary head 233. The capillaries 202 are filled with a new polymer from the other end portions by a syringe 206. The refill of the polymer in the capillaries 202 is executed per measurement to improve the performance of the measurement.

[0025] The pump mechanism 203 includes the syringe 206 and a mechanical system for pressurizing the syringe 206, and injects the polymer into the capillaries 202.

[0026] The block 207 is a connection portion for connecting the syringe 206, the capillary array 217, an anode buffer container 210, and a polymer container 209 to communicate with each other.

[0027] To keep the capillaries 202 in the thermostatic bath 218 at a constant temperature, the thermostatic bath 218 is covered with a heat insulating material, and the temperature is controlled by a heating and cooling mechanism 220. A fan 219 circulates and stirs air in the thermostatic bath 218, and a temperature of the capillary array 217 is kept positionally uniform and constant.

[0028] The transport device 225 transports various containers to the capillary cathode ends 227. The transport device 225 includes three electric motors and a linear actuator, and is movable in three axis directions including vertical, horizontal, and depth directions. At least one container can be placed on a moving stage 230 of the transport device 225. The moving stage 230 includes an electric grip 231 and can grip and release each of the containers. Therefore, a buffer container 221, a cleaning container 222, a waste solution container 223, and a sample plate 224 can be transported to the capillary cathode ends 227 as necessary. Unnecessary containers are stored in a predetermined storage in the electrophoresis device 105.

[0029] By using the data analysis device 112, the user can control various functions of the electrophoresis device 105 and acquire the migration data detected by the optical detection unit.

[0030] The electrophoresis device 105 may include a sensor for acquiring information regarding an observation environment that affects electrophoresis (observation environment information). The electrophoresis device 105 in FIG. 2 includes an internal sensor 240, a polymer sensor 241, and a buffer solution sensor 242.

[0031] The internal sensor 240 is a sensor for acquiring information regarding an internal environment of the electrophoresis device 105, and measures a temperature sensor, a humidity sensor, an air pressure sensor, and the like in the electrophoresis device 105.

[0032] The polymer sensor 241 is a sensor for acquiring information regarding the quality of the polymer and is, for example, a PH sensor or an electrical conductivity sensor. The polymer sensor 241 is installed in the polymer container 209 in FIG. 2, but an installation position is not limited thereto.

[0033] The buffer solution sensor 242 is a sensor for acquiring information regarding the quality of a buffer solution and is, for example, a temperature sensor. The buffer solution sensor 242 is installed in the anode buffer container 210 in FIG. 2, but an installation position is not limited thereto. For example, the buffer solution sensor 242 may be set in the buffer container 221.

[0034] FIG. 3 is a flowchart illustrating the summary of a process that is executed by the gene analysis apparatus 101 according to Embodiment 1.

[0035] The electrophoresis device 105 of the gene analysis apparatus 101 executes an electrophoresis process on a sample to be analyzed (step S301) . Details of the electrophoresis process will be described with reference to FIG. 4.

[0036] Next, the data analysis device 112 of the gene analysis apparatus 101 executes spectrum correction for correcting wavelength characteristics of the device (step S302), and executes a fluorescent intensity calculation process using the migration data (step S303). Specifically, the fluorescent intensity calculation unit 110 calculates time-series data of a fluorescent intensity of a fluorescent dye based on the migration data, and detects a center position, a height, a width, and the like of a peak based on the time-series data of the fluorescent intensity.

[0037] Next, the data analysis device 112 of the gene analysis apparatus 101 executes a mobility correction process on the time-series data of the fluorescent intensity (step S304) . Next, the data analysis device 112 of the gene analysis apparatus 101 executes base calling using the corrected time-series data of the fluorescent intensity based on a result of the mobility correction process (step S305) . Specifically, the base calling unit 107 identifies a base sequence of the sample using the corrected time series data of the fluorescent intensity.

[0038] FIG. 4 illustrates a flow of the electrophoresis process of an actual sample in S301. A basic procedure of electrophoresis can be roughly divided into sample preparation (S401), analysis start event (S402), migration medium filling(S403), preliminary migration (S404), sample introduction (S405), migration analysis (S406), and end of migration analysis (S407).

[0039] An operator of the present apparatus sets the sample or a reagent in the apparatus as the sample preparation (S401) before starting analysis. More specifically, first, the buffer container 221 and the anode buffer container 210 are filled with a buffer solution that forms a part of a conduction path. The buffer solution is, for example, an electrolytic solution that is commercially available from each manufacturer for electrophoresis. The sample to be analyzed is dispensed into a well of the sample plate 224. The sample is, for example, a PCR product of DNA. Further, a cleaning solution for cleaning the capillary cathode ends 227 is dispensed into the cleaning container 222. The cleaning solution is, for example, pure water. A migration medium for electrophoresing the sample is injected into the syringe 206. The migration medium is, for example, a polyacrylamide separation gel or a polymer that is commercially available from each manufacturer for electrophoresis. When deterioration of the capillaries 202 is expected or when lengths of the capillaries 202 are changed, the capillary array 217 is replaced.

[0040] At this time, as the samples to be set to the sample plate 224, in addition to the actual sample of DNA to be analyzed, a positive control, a negative control, and an allelic ladder can be set, and the samples are electrophoresed in different capillaries. The positive control is, for example, a PCR product including known DNA, and is a sample for a control experiment verifying that DNA is correctly amplified by PCR. The negative control is a PCR product not including DNA, experiment for verifying that of the operator and dust amplification product of PCR. and is a sample for a control a contamination such as DNA is not generated in an

[0041] The allelic ladder a large number of alleles that may be included generally in a DNA marker, and is generally provided from a reagent manufacturer as a reagent kit for a DNA test. The allelic ladder is used to finely adjust a correspondence between a DNA fragment length and the allele of the DNA marker.

[0042] A known DNA fragment called a size standard that is labeled with a specific fluorescent dye is mixed with all of the above-described actual sample, the positive control, the negative control, and the allelic ladder. A type of the fluorescent dye allotted to the size standard varies depending on the reagent kit to be used.

[0043] The operator designates a type of the allelic ladder, a type of the size standard, a type of the fluorescent reagent, a type of a sample set to the well on the sample plate 224 corresponding to the capillaries, and the like. In the present embodiment, as a type of sample, any one of the actual sample, the positive control, the negative control, and the allelic ladder is designated. These pieces of information are set by the sample information setting unit 106 via the user interface unit 103 of the data analysis device 112.

[0044] After the above-described sample preparation (S401) is completed, the operator operates the user interface unit 103 on the data analysis device 112 to instruct the start of the analysis. The instruction to start the analysis is passed to the electrophoresis device control unit 108. The electrophoresis device control unit 108 transmits an analysis start signal to the electrophoresis device 105 to start the analysis (S402).

[0045] Next, in the electrophoresis device 105, the migration medium filling (S403) is started. This step may be automatically executed after the start of the analysis or may be sequentially executed by transmitting a control signal from the electrophoresis device control unit 108. The migration medium filling is a procedure of filling the capillary 202 with a new migration medium to form a migration path.

[0046] In the migration medium filling (S403) in the present embodiment, first, the waste solution container 223 is transported to a position immediately below the load header 229 by the transport device 225, and an electromagnetic valve 213 is closed, and the used migration medium discharged from the capillary cathode end 227 is received. Then, the syringe 206 is driven to fill the capillary 202 with the new migration medium, and the used migration medium is discarded. Finally, the capillary cathode end 227 is immersed in the cleaning solution in the cleaning container 222 to clean the capillary cathode end 227 contaminated by the migration medium.

[0047] Next, the preliminary migration (S404) is executed. This step may be automatically executed or may be sequentially executed by transmitting the control signal from the electrophoresis device control unit 108. The preliminary migration is a procedure of applying a predetermined voltage to the migration medium to adjust the migration medium to a state suitable for electrophoresis. In the preliminary migration (S404) in the present embodiment, first, the capillary cathode end 227 is immersed in the buffer solution in the buffer container 221 by the transport device 225 to form a conduction path. Then, a voltage of about several kilovolts to several tens of kilovolts is applied to the migration medium for several minutes to several tens of minutes by the high-voltage power supply 204 to adjust the migration medium to a state suitable for electrophoresis. Finally, the capillary cathode end 227 is immersed in the cleaning solution in the cleaning container 222 to clean the contaminated capillary cathode end 227 with the buffer solution.

[0048] Next, the sample introduction (S405) is executed. This step may be automatically executed or may be sequentially executed by transmitting the control signal from the electrophoresis device control unit 108. In the sample introduction (S405), sample components are introduced into the migration path. In the sample introduction (S405) in the present embodiment, first, the capillary cathode end 227 is immersed in the sample held in the well of the sample plate 224 by the transport device 225, and then the electromagnetic valve 213 is opened. Accordingly, a conduction path is formed, and a state where the sample components are introduced into the migration path is established. Then, a pulse voltage is applied to the conduction path by the high-voltage power supply 204, and the sample components are introduced into the migration path Finally, the capillary cathode end 227 is immersed in the cleaning solution in the cleaning container 222 to clean the contaminated capillary cathode end 227 with the sample.

[0049] Next, the migration analysis (S406) is executed. This step may be automatically executed or may be sequentially executed by transmitting the control signal from the electrophoresis device control unit 108. In the migration analysis (S406), the sample components contained in the sample are separated and analyzed by electrophoresis. In the migration analysis (S406) in the present embodiment, first, the capillary cathode end 227 is immersed in the buffer solution in the buffer container 221 by the transport device 225 to form a conduction path. Next, a high voltage of about 15 kV is applied to the conduction path by the high-voltage power supply 204 to generate an electric field in the migration path. By the generated electric field, each of the sample components in the migration path moves to the detection unit 216 at a speed depending on characteristics of each of the sample components. That is, the sample components are separated by a difference in moving speed. Then, the sample components arrived at the detection unit 216 are detected in order of arrival. For example, when the sample includes a large number of DNAs having different numbers of bases, a difference in moving speed is generated depending on the number of bases, and the DNAs arrive at the detection unit 216 in order starting from a DNA having a shortest base length. Each DNA is attached with a fluorescent dye depending on its terminal base sequence. When the detection unit 216 is irradiated with the excitation light from the light source 214, information light, that is, fluorescence having a wavelength depending on the sample is generated from the sample and emitted to the outside. The information light is detected by the optical detector 215. During the migration analysis, the optical detector 215 detects the information light at a constant time interval and transmits image data to the data analysis device 112. Alternatively, in order to reduce the amount of information to be transmitted, not only the image data but also a luminance of only a partial region in the image data may be transmitted. For example, for each of the capillaries, a luminance that is sampled from only a wavelength position at a constant interval may be transmitted. Luminance data represents a spectrum waveform of each of the capillaries. The spectrum waveform is stored in the storage unit 104.

[0050] Finally, when scheduled image data is acquired, voltage application is stopped, and the migration analysis ends (S407) . The above is an example of the electrophoresis process (S301) in FIG. 4.

[0051] FIG. 5 illustrates a flow of the base calling in S305. First, the analysis section detection unit 109 of the base calling unit 107 detects a signal section from the corrected time-series data of the fluorescent intensity (step S501) . The base calling unit 107 analyzes the detected signal section and specifies a base sequence of the sample (step S502).

[0052] FIG. 6 illustrates a flow of signal section detection in S501. The signal section detection in S501 includes steps S601 to S604. In step S601, the analysis section detection unit 109 divides the entire time-series data into a plurality of small sections. Thereafter, the signal section detection proceeds to step S602. In step S602, the analysis section detection unit 109 selects one of the small sections, and generates non-signal features for each signal included in this section. Each signal is four pieces of fluorescent intensity data corresponding to four bases. The analysis section detection unit 109 generates a non-signal feature for each of the four pieces of fluorescent intensity data of the selected small section. Thereafter, the signal section detection proceeds to step S603.

[0053] In step S603, the analysis section detection unit 109 determines the non-signal feature of the selected small section. Specifically, the analysis section detection unit 109 sets, as the non-signal feature of the small section, a smallest feature among the four non-signal features obtained from the four pieces of fluorescent intensity data. After step S603, if a small section in which the non-signal feature is not determined remains, the analysis section detection returns to step S602. If the non-signal feature is determined for each of the small sections, the signal section detection proceeds to step S604. As described above when signal features are used instead of the non-signal features, a maximum signal feature is set as the signal feature of the small section, and the processing is executed in the same flow as described above.

[0054] In step S604, the analysis section detection unit 109 determines a boundary between the non-signal section and the signal section using the non-signal feature determined for each of the small sections, and ends the processing.

[0055] Here, the non-signal feature will be described. FIG. 7 is a diagram illustrating characteristics of a non-signal section. FIG. 8 is a diagram illustrating characteristics of a signal section. In FIGS. 7 and 8, Dyel to Dye4 represent four fluorescent dyes corresponding to four bases. In FIGS. 7 and 8, each of horizontal axes represents time, and each of vertical axes represents fluorescent intensity.

[0056] Comparing fluorescent intensity data in FIG. 7 and fluorescent intensity data in FIG. 8, there are many undulations and flat segments in the non-signal section, and there are few undulations and flat segments in the signal section. In particular, in the signal section, since transition of the fluorescent intensity is gentle, the flat portion is significantly reduced.

[0057] Therefore, the analysis section detection unit 109 generates the non-signal feature based on the number of appearances of the three shape patterns (maxima, minima, and flat segment) of the fluorescent intensity data. The generation of the non-signal feature based on the shape patterns corresponds to step S602.

[0058] FIG. 9 is a diagram illustrating the generation of non-signal features based on the shape patterns. In the shape pattern "flat segment", an intensity difference between adjacent points is ±hl. That is, the following formula is satisfied. Here, the point is a sample value of an electrophoresis signal, and is determined by the above-described time interval at which the optical detector 215 acquires data or a sampling rate. This time interval is predetermined as a default value of the apparatus by the user. -hl <(y[k + 1] - y[k]) <= hl The shape pattern "maxima" satisfies the following formula . y[k] - y[k - 1] >h2 &&y[k] - y[k + 1] >h2 The shape pattern "minima" satisfies the following f o rmu1a . y[k - 1] - y[k] >h3 &&y[k + 1] - y[k] >h3 The above-described hl, h2, and h3 may be predetermined values according to the sampling rate and an electrophoresis voltage.

[0059] The analysis section detection unit 109 sets the number of appearances of the three patterns as the nonsignal feature of the fluorescent intensity data in the section. The appearance frequency of the three patterns may be normalized by a section length and used as the nonsignal feature.

[0060] Among the three patterns, the shape pattern "flat segment" is less likely to appear in the signal section and has a high importance as a feature of the non-signal section Therefore, the shape pattern "flat segment" may be weighted more heavily than other shape patterns to generate the nonsignal feature.

[0061] FIG. 10 is a diagram illustrating non-signal feature determination in a section in S603. In FIG. 10, a minimum value of the non-signal features of the fluorescent intensity data in the section is set as the non-signal feature of this section. A graph illustrated in FIG. 10 is the fluorescent intensity data of Dyel to Dye4. F (Dye 1) to F (Dye 4) are non-signal features generated from the fluorescent intensity data of Dye 1 to Dye 4, respectively.

[0062] The analysis section detection unit 109 obtains a nonsignal feature Fq of a section q by Min (F(Dye 1), F(Dye 2), F(Dye 3) , F(Dye 4) ) . In FIG. 10, since the fluorescent intensity data of Dye 1 is the gentlest, F (Dye 1) is smaller than F (Dye 2) to F (Dye 4) . Therefore, Fq = F(Dyel) is obtained.

[0063] FIGS. 11 and 12 are each a diagram illustrating signal boundary determination in S604. As illustrated in FIG. 11, the analysis section detection unit 109 plots the non-signal feature of each section, and executes smoothing and interpolation. Accordingly, it is possible to prevent an influence of a fine fluctuation of the non-signal feature. The analysis section detection unit 109 sets, as a signal section boundary, a time when the smoothed and interpolated non-signal feature exceeds a threshold.

[0064] As illustrated in FIG. 12, the analysis section detection unit 109 may set, as a boundary, a time when a section in which the non-signal feature continuously exceeds the threshold is equal to or larger than a certain margin. Accordingly, it is possible to be robust against an influence of a fluctuation near the boundary.

[0065] FIG. 13 is a diagram illustrating a case where a threshold is determined based on a distribution of nonsignal features in a section. The distribution of the nonsignal features is assumed to be bimodal. A signal portion is low, and a non-signal portion is high. The analysis section detection unit 109 can determine the threshold based on the distribution. For example, X% of a peak Fp of a higher mountain (non-signal) can be used as the threshold. Further, a value at which a slope of the mountain becomes flat to some extent may be used as the threshold.

[0066] However, since the acquired electrophoresis data may not include the non-signal section depending on setting executed by the user, a predetermined fixed value may be always used. Alternatively, it may be possible to determine whether a distribution of features is bimodal and determine whether to dynamically determine a threshold or to obtain a fixed value .

[0067] FIGS. 14 to 18 are each a diagram illustrating a case where non-signal features are used in combination with other features . In FIG. 14, a threshold is changed according to a signal level. That is, the threshold is determined according to the following (1) and (2). (1) If a signal intensity is equal to or larger than a certain level, it is always determined as a signal section (2) If the signal intensity is equal to or less than a certain level, the threshold is lowered according to the signal intensity. Here, in a range of (2) , the lower the signal intensity, the easier it is to determine that it is a non-signal section In addition, a differential signal may be used. By lowering the threshold as the difference increases, rising and falling of the signal section can be detected.

[0068] In FIG. 15, a feature vector including a non-signal feature and other features is given to a signal section discriminator 121 as an input, and a signal section is obtained as an output. Other features may be any feature such as signal intensity or differential signal. Any model such as a deep neural network (DNN) , a support-vector machine (SVM), or a random forest can be used as the signal section discriminator 121. The output may be a determination result as to whether the section is the signal section or the non-signal section, a probability (likelihood) of the section being the signal section, or the like.

[0069] FIG. 16 illustrates ensemble learning in which a plurality of discriminators are combined. In an example illustrated in FIG. 16, a first feature vector is given to a first signal section discriminator 122 as an input, and a second feature vector is given to a second signal section discriminator 123 as an input. Then, outputs of the signal section discriminators 122 and 123 are used as inputs to the identifier 124 to finally obtain an output.

[0070] The first feature vector includes, for example, a nonsignal feature and a signal intensity, and the second feature vector includes, for example, a differential signal. The identifier 124 determines the output by majority decision or the like. Other methods such as bagging, boosting, and stacking may also be used.

[0071] FIG. 17 is a diagram illustrating a configuration for executing learning on the apparatus. A signal section information storage unit 125 is a storage unit provided in a gene analysis apparatus. The signal section information storage unit 125 stores signal section information (a label indicating whether the section is a signal section) and the feature vector in association with each other every time the user performs a measurement. The analysis section detection unit 109 can read out the feature vector and the label from the signal section information storage unit 125 and provide the feature vector and the label to a signal section discriminator to execute supervised learning.

[0072] FIG. 18 is a diagram illustrating learning reflecting an adjustment result executed by the user. When the user adjusts an analysis section and executes reanalysis, an operation and a result are stored in a predetermined information storage unit and reflected in the next learning. Accordingly, the detection accuracy of the signal section can be improved. FIG. 18 illustrates an example in which the user operates the boundary between the signal section and the non-signal section, but adjustment results of other parameters may be used. For example, if the user adjusts a non-signal feature parameter (a condition of the flat segment, conditions of a maximum point and a minimum point), threshold setting, and a parameter related to signal boundary determination, a result of the operation may be stored and learned.

[0073] The base calling unit 107 in the gene analysis apparatus 101 may reanalyze fluorescent intensity data other than fluorescent intensity data generated from an electrophoresis result in the fluorescent intensity calculation unit 110. As an example of this case, the fluorescent intensity data may be stored in the storage unit 104 or may be transmitted through a communication cable. In this case, the user can adjust the analysis section by editing the fluorescent intensity data. FIG. 19 is a diagram illustrating a case where the analysis section is corrected by editing the fluorescent intensity data. With respect to the fluorescent intensity data in an upper part in FIG. 19, by correcting the data to include many flat portions with respect to the fluorescent intensity data near the beginning as illustrated in a lower part in FIG. 19, the non-signal feature increases, and a signal start position can be moved. When a fluorescent intensity near the beginning is set to a zero value, a base calling result near the beginning may change from that before the correction due to a large change in the fluorescent intensity. For this reason, in the lower part in FIG. 19, only the signal start position is changed by increasing the flat portion such that the signal intensity deviates from the signal intensity before correction within a certain range (a range of gray in the lower part in FIG. 19) . In addition to the flat portion, the maximum point or the minimum point may be added. Such editing of the fluorescent intensity data may be executed using an external tool. Alternatively, the gene analysis apparatus 101 may have such a function of editing the fluorescent intensity data.

[0074] As described above, the disclosed gene analysis apparatus 101 includes the fluorescent intensity calculation unit 110 serving as the acquisition unit that acquires the time-series data indicating the electrophoresis result of the sample, and the base calling unit 107 serving as the analysis unit that analyzes the base sequence of the sample from the time-series data. The timeseries data includes a plurality of pieces of fluorescent intensity data corresponding to a plurality of bases, and the analysis unit divides the time-series data into a plurality of sections, generates, for each of the plurality of pieces of fluorescent intensity data, a feature indicating an appearance frequency of at least one of a maximum portion, a minimum portion, and a flat portion of the fluorescent intensity data in each section, determines, based on a magnitude relation of the plurality of features generated for the plurality of pieces of fluorescent intensity data, a section feature from the plurality of features, and detects a signal area that is an analysis target area of the base sequence in the time-series data by using the section feature. According to this configuration, the signal section (signal area) can be detected with high accuracy from the time-series data indicating the electrophoresis result. For example, a signal section including a low-intensity signal can also be accurately detected. In addition, the signal section can be accurately detected even when there is a large variation in signal intensity. Specifically, the detection accuracy of the signal section is improved for data including an unintentionally and sporadically high-intensity signal (Dye Blob) caused by sample preprocessing, data including a high-intensity signal appearing at the end of a PCR reaction, and data in which the signal intensity is attenuated due to a special sample or preprocessing.

[0075] In addition, the analysis unit determines a smallest feature among the plurality of features of the fluorescent intensity data as the section feature. As described above, the signal section can be detected by specifying the non-signal section using the shape patterns that characteristically appear not in the signal section but in the non-signal section. In addition, since only one base is present at the same position in a plurality of pieces of fluorescent intensity data due to characteristics of electrophoresis, the signal section can be detected with high accuracy using the characteristics of electrophoresis by selecting, as a representative feature of the section, a feature that is most likely to be a signal from the plurality of pieces of fluorescent intensity data.

[0076] The analysis unit compares the section feature with a threshold to determine whether a corresponding section is the non-signal section. As the threshold, a predetermined fixed value or a value obtained from a distribution of features of the section can be used. In this configuration, the signal section can be accurately detected in consideration of the distribution of the non-signal features.

[0077] When a certain number of continuous sections each have the section feature larger than the threshold, the analysis unit sets a boundary between the continuous sections and the signal area adjacent to the continuous sections as a boundary between the signal area and a non-signal area. In this configuration, it is possible to prevent an influence of a fluctuation near the boundary of the signal section.

[0078] The analysis unit detects the signal area by using a discrimination model in which the section feature is one of inputs . In this configuration, the signal section can be flexibly detected based on the non-signal features and other features .

[0079] In addition, the analysis unit generates the feature for each piece of fluorescent intensity data by using a weight for the flat portion of the fluorescent intensity data that is larger than those for the maximum portion and the minimum portion of the fluorescent intensity data. As described above, the signal section can be accurately detected by emphasizing the characteristic shape patterns of the non-signal section.

[0080] In addition, the analysis unit may detect a second signal section different from the fluorescent intensity data by using second fluorescent intensity data obtained by editing the fluorescent intensity data. Here, the second fluorescent intensity data is a deviation amount within a certain range from an intensity of the fluorescent intensity data . When the second signal section is detected using such second fluorescent intensity data, the analysis section can be changed while preventing an influence on an analysis result.

[0081] The invention is not limited to the above-described embodiment, and includes various modifications. For example the above-described embodiment has been described in detail to facilitate understanding of the invention, and the invention is not necessarily limited to those including all the configurations described above. In addition to deletion of such a configuration, it is also possible to replace or add a configuration. For example, in the above-described embodiment, a configuration in which a device that executes learning based on data and a device that updates the signal section discriminator are integrated is exemplified, but a configuration in which learning and updating of the signal section discriminator are executed by different devices may be adopted. Reference Signs List

[0082] 101: gene analysis apparatus 102: central control unit 103: user interface unit 104: storage unit 105: electrophoresis device 106: sample information setting unit 107: base calling unit 108: electrophoresis device control unit 109: analysis section detection unit 110: fluorescent intensity calculation unit 112: data analysis device 121 to 123: signal section discriminator 124: identifier 125: signal section information storage unit

Claims

1. A gene analysis apparatus comprising:an acquisition unit configured to acquire time-series data indicating an electrophoresis result of a sample; andan analysis unit configured to analyze a base sequence of the sample from the time-series data, whereinthe time-series data includes a plurality of pieces of fluorescent intensity data corresponding to a plurality of bases, andthe analysis unitdivides the time-series data into a plurality of sections,generates, for each of the plurality of pieces of fluorescent intensity data, a feature indicating an appearance frequency of at least one of a maximum portion, a minimum portion, and a flat portion of the fluorescent intensity data in each section,determines, based on a magnitude relation of the plurality of features generated for the plurality of pieces of fluorescent intensity data, a section feature from the plurality of features, anddetects a signal area that is an analysis target area of the base sequence in the time-series data by usingthe section feature.

2. The gene analysis apparatus according to claim 1, whereinthe analysis unit determines a smallest feature among the plurality of features as the section feature.

3. The gene analysis apparatus according to claim 1, whereinthe analysis unit compares the section feature with a threshold to determine whether a corresponding section is the signal area, andthe threshold is a predetermined fixed value or a value obtained from a distribution of the section feature.

4. The gene analysis apparatus according to claim 3, whereinwhen a certain number of continuous sections each have the section feature larger than the threshold, the analysis unit sets a boundary between the continuous sections and the signal area adjacent to the continuous sections as a boundary between the signal area and a non-signal area.

5. The gene analysis apparatus according to claim 1, whereinthe analysis unit detects the signal area by using a discrimination model in which the section feature is one of inputs .

6. The gene analysis apparatus according to claim 1, whereinthe analysis unit generates the feature by using a weight for the flat portion of the fluorescent intensity data that is larger than those for the maximum portion and the minimum portion of the fluorescent intensity data.

7. The gene analysis apparatus according to claim 1, whereinthe analysis unit detect a second signal section different from the fluorescent intensity data by using second fluorescent intensity data obtained by editing the fluorescent intensity data, andthe second fluorescent intensity data is a deviation amount within a certain range from an intensity of thefluorescent intensity data.

8. A gene analysis method comprising:acquiring time-series data indicating an electrophoresis result of a sample and including a plurality of pieces of fluorescent intensity data corresponding to a plurality of bases;dividing the time-series data into a plurality of sections;generating, for each of the plurality of pieces of fluorescent intensity data, a feature indicating a nonsignal degree of the fluorescent intensity data in each section based on an appearance frequency of at least one of a maximum portion, a minimum portion, and a flat portion of the fluorescent intensity data in each section;determining, based on a magnitude relation of the plurality of features generated for the plurality of pieces of fluorescent intensity data, a section feature from the plurality of features; anddetecting a signal area that is an analysis target area of the base sequence in the time-series data by using the section feature.

Citation Information

Patent Citations

  • Signal processing for determining base sequence nucleic acid

    JP1987225956A

  • Data processor

    JP1988210769A

  • Method for analyzing cataphoresis pattern of nucleic acid piece

    JP1999118760A

  • Data processing device, data processing method, and data processing program

    JP2012177568A

  • Living organism information processing device and method as well as program

    JP2018042560A