Spontaneous mutation detection method, device, program, and storage medium
The mutation detection method addresses the challenge of low-frequency mutation detection by creating a noise model from reference DNA samples to align and exclude noise peaks, improving accuracy and sensitivity.
Patent Information
- Application Number
- PCT/JP2025/017897
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-16
- Filing Date
- 2025-05-16
- Publication Date
- 2025-11-20
AI Technical Summary
Existing methods struggle to accurately detect low-frequency mutations in DNA strands due to irreproducible noise sources during DNA sample preparation, leading to false positives and negatives, and require additional costly methods like qPCR for noise removal.
A mutation detection method involving waveform data acquisition, noise model creation using multiple reference DNA samples, alignment with test data, and exclusion of noise model-matched peaks to identify true mutations.
Accurately detects low-frequency mutations by eliminating noise specific to DNA sample preparation, reducing false positives and negatives, and enhancing detection sensitivity.
Smart Images

Figure JP2025017897_20112025_PF_FP_ABST
Abstract
Description
Mutation detection method, device, program, and storage medium
[0001] The present invention relates to a mutation detection method, an apparatus, a program, and a storage medium.
[0002] Background art of the present invention includes, for example, Patent Document 1. The invention described in Patent Document 1 aims to provide a method that can be implemented using software for automatically detecting DNA mutations from sequence trace data. To achieve this objective, the invention described in Patent Document 1 includes the following steps: The invention described in Patent Document 1 aligns waveform data of a sample DNA sequence with waveform data of a reference DNA sequence to generate an aligned sample DNA sequence. The invention described in Patent Document 1 also calculates a correlation coefficient between the waveform data of bases of both the reference DNA sequence and the aligned sample DNA sequence for a specific frame number. If the DNA base waveform data of the aligned sample DNA sequence is a mutation compared to the reference DNA sequence, a correlation coefficient exceeding a minimum value is generated for the corresponding frame number. Furthermore, non-patent documents 1 and 2, for example, disclose DNA sequence alignment. Specifically, non-patent document 1 discloses the Smith-Waterman algorithm. Furthermore, non-patent document 2 discloses an algorithm based on this algorithm that is widely used for homology searches such as BLAST.
[0003] U.S. Patent Application Publication No. 2001 / 000952
[0004] Smith, Temple F. & Waterman, Michael S. (1981). "Identification of Common Molecular Subsequences". Journal of Molecular Biology. 147 (1): 195-197. Zheng Zhang, Scott Schwartz, Lukas Wagner, and Webb Miller (2000), "A greedy algorithm for aligning DNA sequences", J Comput Biol 2000; 7(1-2):203-14.
[0005] Generally, methods for detecting low-frequency variants (nucleic acid species having mutations) from wild-type nucleic acid species that are predominant in terms of molecular number have a detection sensitivity or limit. Mutations here include alleles, polymorphisms, etc. For example, a detection limit of 5% makes it impossible to identify mutations with frequencies lower than that value. The detection limit arises because the detection signal intensity of low-frequency variants is so low that it cannot be distinguished from noise generated in DNA by the Sanger method. The invention described in Patent Document 1 provides a method for distinguishing between noise generated under specific conditions and the detection signal of a variant. However, the invention described in Patent Document 1 does not provide a method for distinguishing between noise generated under general conditions and the detection signal of a variant.
[0006] Furthermore, when using a capillary electrophoresis device, there is noise that does not reproduce between the reference and the test sample whose DNA base sequence is being investigated. There are four sources of noise: (1) errors (base substitution, degradation) that occur during the preservation (formalin fixation) process of biological samples; (2) errors that occur during molecular amplification (PCR), a process specific to DNA samples; (3) equipment and consumables; and (4) contaminants in the preparation of sequencing samples.
[0007] The aforementioned irreproducible noise is often caused by (1) and (2). These noises are contained in the samples used for DNA sequencing and are generated during the process specific to DNA sample preparation. The appearance pattern of these noises varies depending on various factors, such as DNA purity and sample handling. Reproducible noise can cause false positives for mutations. Furthermore, setting a high filter to remove noise can bury positives and cause false negatives. To resolve these issues, users must use other methods (e.g., qPCR, protein staining), which incurs additional costs.
[0008] The invention described in Patent Document 1 can identify noise caused by consumables and equipment with high accuracy. However, the invention described in Patent Document 1 cannot remove noise that does not recur between the sample and the reference. For the reasons described above, there is a concern that the invention described in Patent Document 1 may have difficulty in accurately detecting low-frequency base mutations on DNA strands. This concern cannot be alleviated even when Non-Patent Documents 1 and 2 are taken into consideration, which merely explain DNA sequence alignment.
[0009] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a mutation detection method, device, program, and storage medium that can accurately detect low-frequency mutations in bases on a DNA strand.
[0010] The mutation detection method of the present invention, which solves the above-mentioned problems, is a method for detecting mutations in bases on a DNA strand, and includes: a waveform data acquisition step of acquiring waveform data including at least base sequence information and information on the appearance position of peak waveforms using a plurality of reference DNA samples that do not contain mutations; a noise model data creation step of creating noise model data using a collection of peak waveforms of the waveform data of the plurality of reference DNA samples; a test data acquisition step of acquiring test data, which is waveform data of a test sample; an alignment step of comparing the base sequences of the noise model data and the test data, and aligning the test data with the noise model data so that the same bases pair; and an exclusion step of excluding peak waveforms from the peak waveforms of the test data that appear in the same positions as the noise model data.
[0011] The mutation detection method, device, program, and storage medium according to the present invention can accurately detect low-frequency mutations in bases on DNA strands. Problems, configurations, and advantages other than those described above will become apparent from the following description of the embodiments.
[0012] 1 is a flowchart illustrating the contents of a mutation detection method according to one embodiment of the present invention. FIG. 1 is an explanatory diagram illustrating a comparison of peak waveforms of a test sample and a reference DNA sample. FIG. 2 is an explanatory diagram illustrating an example of a false positive occurring in a test sample. FIG. 2 is a flowchart illustrating a specific example of the contents of a mutation detection method according to one embodiment of the present invention. FIG. 3 is an explanatory diagram illustrating an example of a correspondence relationship between main peak positions between two samples. FIG. 4 is an explanatory diagram illustrating a method for determining a correspondence relationship between positions between two samples from peak waveform data. FIG. 5 is an explanatory diagram illustrating the concept of a noise model creation method according to a first aspect. FIG. 6 is an explanatory diagram illustrating identification of a mutation peak using peak information. FIG. 7 is an explanatory diagram illustrating an example of a peak information integration process in the noise model creation method according to the first aspect. FIG. 8 is an explanatory diagram illustrating another example of a peak information integration process in the noise model creation method according to the first aspect. FIG. 9 is an explanatory diagram illustrating the concept of a noise model creation method according to a second aspect. FIG. 10 is an explanatory diagram illustrating the concept of a noise model creation method according to the second aspect. FIG. 11 is a flowchart illustrating the contents of the noise model creation method according to the second aspect. FIG. 12 is an explanatory diagram illustrating cluster integration in the noise model creation method according to the second aspect. FIG. 13 is an explanatory diagram illustrating an example of a mutation peak detection method using a noise model according to the second aspect. 18 is an explanatory diagram illustrating another example of a method for detecting a mutation peak using a noise model of a second aspect. FIG. 19 is an explanatory diagram illustrating a case where a peak is unclear in the noise model of a second aspect. FIG. 20 is an explanatory diagram illustrating the concept of a noise model creation method of a third aspect. FIG. 21 is an explanatory diagram illustrating an example of creation of an integrated signal in the noise model creation method of a third aspect. FIG. 22 is an explanatory diagram illustrating an example of a method for detecting a mutation peak using a noise model of a third aspect. FIG. 23 is a block diagram illustrating an example of the configuration of a mutation detection device according to an embodiment of the present invention. FIG. 24 is a block diagram illustrating an example of the secondary analysis software shown in FIG. 17. FIG. 25 is a block diagram illustrating an example of the configuration of a mutation detection storage medium according to an embodiment of the present invention. FIG. 26 is a table showing information on primers used in the examples. FIG. 27 is a plot diagram illustrating the results of the examples. FIG. 28 is an explanatory diagram showing an example of an electrophoresis waveform in STR analysis. FIG. 29 is an explanatory diagram illustrating size calling in STR analysis.FIG. 1 is an explanatory diagram showing how to determine the relational expression "y=f(t)" between DNA electrophoresis time (t) and DNA base length (y). FIG. 2 is an explanatory diagram showing an example of an electrophoresis waveform of an allelic ladder in STR analysis. FIG. 3 is an explanatory diagram explaining the concept of basic information of an allelic ladder. FIG. 4 is an explanatory diagram explaining the concept of detecting a mutation peak in STR analysis. FIG. 5 is an explanatory diagram explaining the concept of detecting a mutation peak in STR analysis for the purpose of performing allele calling.
[0013] Hereinafter, a mutation detection method, an apparatus, a program, and a storage medium according to an embodiment of the present invention will be described with reference to the accompanying drawings. Note that common components in the following description and drawings may be designated by the same reference numerals, and redundant description may be omitted.
[0014] [First embodiment] (Mutation detection method) First, a mutation detection method according to one embodiment of the present invention (hereinafter, sometimes simply referred to as "this method") will be described. FIG. 1 is a flowchart explaining the details of this method. FIG. 2 is an explanatory diagram explaining a comparison of peak waveforms of a test sample and a reference DNA sample. FIG. 3 is an explanatory diagram showing an example of a false positive occurring in a test sample. FIG. 4 is a flowchart explaining a specific example of the details of this method.
[0015] This method detects mutations in bases on a DNA strand, i.e., mutations in DNA base sequences in sequence data. To detect mutations in DNA base sequences, this method includes a waveform data acquisition step S101, a noise model data creation step S102, a test data acquisition step S103, an alignment step S104, and an exclusion step S105, as shown in FIG.
[0016] In the waveform data acquisition step S101, waveform data including at least base sequence information and peak waveform appearance position information is acquired using multiple reference DNA samples that do not contain mutations or whose mutation locations are known. In the noise model data creation step S102, noise model data is created using a collection of peak waveforms from the waveform data of the multiple reference DNA samples. In the test data acquisition step S103, test data, which is waveform data of the test sample, is acquired. In the alignment step S104, the base sequences of the noise model data and the test data are compared, and the test data is aligned with the noise model data so that the same bases are paired. In the exclusion step S105, peak waveforms that appear in the same positions as the noise model data are excluded from the peak waveforms of the test data. The remaining peak waveforms of the test data represent mutations of bases on the DNA strand. This technique enables the method to accurately detect mutations that occur at low frequencies. The details of this method are described in more detail below.
[0017] In this method, mutation detection is performed by comparing sequence data 201 of a test sample (the test data described above) with sequence data 202 of a reference DNA sample (the noise model data described above), as shown in Figure 2. The sequence data 201 of the test sample is composed of a peak waveform 203 derived from wild-type nucleic acid species that account for the majority of molecules in the DNA solution, noise 204 having a lower signal intensity than the peak waveform 203, and a peak waveform 205 derived from a mutation. The sequence data 202 of the reference DNA sample is composed of the peak waveform 203 of the wild-type nucleic acid species and noise 204.
[0018] The reference DNA sample is DNA known to be mutation-free or DNA with known mutation locations. More specifically, the reference DNA sample may be wild-type genomic DNA, artificially synthesized DNA, or amplified DNA thereof. The test sample is a sample used to confirm the presence or absence of mutations (mutations) in bases on a DNA strand, and is typified by a patient sample. The test sample may be genomic DNA extracted from cells or cDNA obtained by reverse transcription of RNA. In this case, the position of noise may differ between the sequence data 201 of the test sample and the sequence data 202 of the reference DNA sample. Figure 3 shows an example.
[0019] Here, the test sample shown in FIG. 3 is known to be mutation-free. As shown in FIG. 3 , comparing sequence data 301 of the test sample with sequence data 302 of the reference DNA sample reveals a peak waveform 303 specific to the test sample, resulting in a false positive for a mutation. Such false positives are difficult to eliminate because they arise from multiple factors, such as reaction efficiency during sample preparation and DNA purity. In this embodiment, after obtaining each sequence data using the method described below, false positives are eliminated to accurately detect low-frequency base mutations in DNA strands. This will be described below with reference to FIG. 4 as appropriate. The sequence data is obtained by performing DNA preparation, PCR reaction, cycle sequencing reaction, purification, and base sequence determination.
[0020] (DNA Preparation) First, the preparation of DNA will be described. DNA can be prepared by any method. A commercially available purification kit may be used for DNA preparation. Next, the region of the DNA to be analyzed is amplified by a PCR reaction. If a sufficient number of molecules for detection can be obtained without amplification, this reaction does not need to be performed.
[0021] (PCR Reaction) Next, the PCR reaction will be described. The PCR reaction requires at least the following reagents. Specifically, the reagents are a test sample and a reference DNA sample, deoxyribonucleoside triphosphates (dNTPs) for each of the four bases (adenine (A), guanine (G), cytosine (C), and thymine (T)), a pH buffer to maintain the pH of the reaction solution, ions necessary for the activity of the thermostable DNA polymerase, the provided thermostable DNA polymerase, and a primer set that pairs with the base sequences at both ends of the DNA region to be analyzed. The test sample and the reference DNA sample are amplified in separate reaction tubes. Next, the reaction product is purified. Purification can be performed by conventional phenol treatment and ethanol precipitation. A commercially available DNA purification kit can also be used for purification.
[0022] (Cycle Sequencing Reaction) Next, the cycle sequencing reaction will be explained. The cycle sequencing reaction is performed using the purified reaction product (PCR product). The cycle sequencing reaction is a process in which a specific fluorescent dye is added to each base. The cycle sequencing reaction requires at least the following reagents: PCR product, dNTP, dideoxyribonucleoside triphosphates (ddNTPs) for each of the four bases, a thermostable DNA polymerase, and a sequencing primer. The ddNTPs are labeled with specific fluorescent dyes for each base.
[0023] (Purification Treatment) Next, the purification treatment will be described. After the cycle sequencing reaction, a purification treatment is performed. The purpose of the purification treatment is to exchange the solvent for one suitable for electrophoresis and to remove unreacted dNTPs, ddNTPs, and primers. Purification methods include ethanol precipitation and gel filtration, and the operator can select an appropriate method to use. Furthermore, a commercially available DNA purification kit can be used as the purification method.
[0024] (Determination of Base Sequence) Next, the determination of base sequence will be explained. The purified reaction products are separated and detected by electrophoresis. DNA molecules are separated by the molecular sieving effect of the separation medium, and fluorescent signals from the labeled substances derived from the labeled ddNTPs are detected. The base sequence is then determined based on the detected signals. The type of electrophoresis method is not particularly limited. Generally, denaturing polyacrylamide gel electrophoresis or capillary electrophoresis is performed. Here, a commercially available capillary electrophoresis device is used. Waveform data and base sequence are obtained by electrophoresis. In other words, each sequence data is obtained.
[0025] Here, waveform data will be described. Waveform data is a diagram in which the horizontal axis represents the signal detection position, usually the number of data scans, and the vertical axis represents the fluorescence intensity. Therefore, waveform data includes at least base sequence information and information on the position at which peak waveforms appear. The migration speed of DNA during capillary electrophoresis is affected not only by the base length of the DNA but also by the type of fluorescent dye bound to the DNA. Therefore, commercially available capillary electrophoresis devices correct the migration speed for each fluorescent dye and output data before and after correction. The corrected data is used in this embodiment. Peaks in waveform data to which base types have been assigned by the device are called main peaks. Main peaks have a higher signal intensity than peaks representing noise or low-frequency mutations.
[0026] The reference DNA sample is the same genetic region as the test sample. Unlike the test sample, the reference DNA sample does not contain mutations or has known mutation locations. Here, it is assumed that the reference DNA sample does not contain mutations. Therefore, all peaks other than the main peak detected in the reference DNA sample, such as peaks with low signal intensity, are noise. The reference DNA sample serves as a representative sample of noise generated in the sample during PCR, cycle sequencing, and purification steps. The location and signal intensity of these noises vary depending on the purity and handling of the DNA. Therefore, it is desirable to prepare multiple reference DNA samples prepared under various conditions. In other words, it is desirable for the multiple reference DNA samples to be multiple samples of different quality. Specifically, multiple types of reference DNA samples are prepared using template DNA in PCR, cycle sequencing, or purification steps to obtain waveform data (this corresponds to step S401 in FIG. 4). Furthermore, the reference DNA samples may be subjected to different sample injection and electrophoresis conditions during electrophoresis. In other words, waveform data for multiple reference DNA samples may be obtained using multiple devices and multiple different preparation reagent lots. On the other hand, if the preparation method of the test sample is predetermined, it is desirable to prepare the reference DNA sample using the same method as the test sample. In this way, by using multiple reference DNA samples, noise that does not recur between the test sample and the reference DNA sample is eliminated. Therefore, this method can accurately detect mutations that appear at low frequencies.
[0027] (Peak Detection) First, the position of a peak in the waveform data for each fluorescent dye is detected for each reference DNA sample. As an example, the time t at which the signal intensity of three consecutive points P[t-1], P[t], and P[t+1] is highest at P[t] may be searched for. Alternatively, the waveform data may be smoothed in advance to remove minute irregularities. If the peak detected from the waveform data is determined to be clearly identical to the main peak of the base sequence, the peak position may be replaced with the main peak position of the base sequence.
[0028] (Waveform Data Alignment) Next, data processing for removing noise using multiple reference DNA samples will be described. First, the positions of peaks in the waveform data are corrected. Waveform data obtained by capillary electrophoresis exhibits variations in the peak detection positions for each data. Therefore, it is necessary to first correct the deviation in the detection position. Base sequences are used to correct the deviation in the detection position. Specifically, one reference reference DNA sample is selected from multiple reference DNAs to serve as a position reference. The base sequences are then aligned so that as many identical bases as possible are paired between the reference reference DNA sample and the other reference DNA samples. One example is alignment so that the Levenshtein distance between the two base sequences is minimized. The Levenshtein distance is a type of distance that indicates the degree of difference between two character strings (in this method, the DNA base sequence of the reference reference DNA sample and the DNA base sequences of the other reference DNA samples). Other base sequence alignment methods that may be used include the Smith-Waterman algorithm (Non-Patent Document 1) and algorithms based on this algorithm that are widely used in homology searches, such as BLAST (Non-Patent Document 2). Furthermore, due to the characteristics of electrophoresis, the reliability of the beginning and end of data reading is low. Therefore, these regions may be omitted from the analysis. For example, the analysis target is limited to 40 to 500 bases from the beginning of reading. This process eliminates noise inherent to electrophoresis that occurs at the beginning and end of data reading from the analysis, resulting in more accurate results.
[0029] FIG. 5 shows an example of the correspondence relationship between the main peak positions of two samples obtained by the above-described base sequence alignment. FIG. 5 is an explanatory diagram illustrating an example of the correspondence relationship between the main peak positions of two samples. The figure shows a correspondence relationship 503 between the base sequence positions (i.e., the main peak positions) of a reference DNA sample 502 and another reference DNA sample 501. Because the peak appearance positions differ between different samples, the correspondence relationship 503 shown in the figure is used to convert the main peak positions of all reference DNA samples 501 to the position coordinate system of the reference DNA sample 502 (hereinafter referred to as the reference position coordinate system). Conversion of any position of a peak other than the main peak can be achieved by linear interpolation using the correspondence relationship 503 of the main peaks located on both sides of the arbitrary position. Therefore, any peak position detected from a waveform, even for peaks other than the main peak, can be converted to the reference position coordinate system.
[0030] (Alignment of Waveform Data Without Using Base Sequence) In the above example, the waveform data was aligned using the base sequence. However, the waveform data may also be aligned based only on the peak waveform data, without using the base sequence. An example of a method for aligning peak waveform data based only on the peak waveform data will be described with reference to FIG. 6 . FIG. 6 is an explanatory diagram illustrating a method for determining the positional correspondence between two samples from the peak waveform data. The lower part of the figure shows a peak waveform 601 of a reference reference DNA sample, and the upper part shows a peak waveform 602 of another reference DNA sample to be aligned with it. As shown in the figure, a peak waveform with a window width W is extracted at a certain position of the reference DNA sample, and a waveform most similar to the extracted peak waveform is searched for among the peak waveform 601 of the reference reference DNA sample. Known methods such as cosine similarity and cross-correlation can be used to calculate the waveform similarity. In this way, the position of the reference reference DNA sample that is most similar to the window at a certain position of the reference DNA sample is used as the transformed position, and the window is shifted by a fixed width across the entire waveform while calculation is performed at any position, thereby obtaining the correspondence between the two samples. To convert any position other than the position of the shifted window, linear interpolation can be performed using the correspondence between the windows located on both sides of that position, just as in the case of conversion using base sequences. In this way, waveform data can be aligned without using base sequences. In this case, it is possible to align data for which it is difficult to obtain base sequences, such as data with insufficient peak separation.
[0031] (Conversion to Relative Signal Intensity) As mentioned above, it is desirable to prepare multiple reference DNA samples under various conditions, but this results in greater variability in signal intensity. Therefore, to correct for variability in signal intensity of waveform data, a relative signal intensity is calculated for each fluorescent dye from the original signal intensity for each aligned waveform data. Any relative signal intensity can be used as long as it corrects for variability in signal intensity between electrophoresis runs. For example, the ratio of each fluorescent dye to its maximum value can be suitably used as the relative signal intensity. The ratio of each fluorescent dye to its average value, or the ratio of the signal intensity at the main peak position of each fluorescent dye to its average value, can also be used as the relative signal intensity. Either of these methods can correct for variability in signal intensity of waveform data. Specifically, it can reduce variability in waveform data peak height between electrophoresis runs. Therefore, subsequent analysis can be performed without being affected by variations in DNA quantity, handling, or sample injection efficiency into the capillary. The information on each peak obtained here (including at least the peak position in the reference position coordinate system, relative signal intensity, and signal intensity) is conveniently referred to as peak information.
[0032] (Identification of Mutant Peaks) Next, identification of mutant peaks using peak information will be described. As described above, to identify mutant peaks, a noise model is created from peak information of multiple reference DNA samples. The method for creating the noise model will be described in detail below.
[0033] (Creating a Noise Model) Even if the peaks of all reference DNA samples are converted into the reference position coordinate system by aligning the waveform data as described above, a shift in the position of the same peak between samples may occur. If the positions of peaks that should be treated as the same remain shifted, it becomes difficult to identify them as mutant peaks by comparing them with the peaks of the test sample described below. Therefore, if a shift in the position of the same peak occurs between samples, a noise model can be created by referencing the peak information of all samples for each fluorescent dye and integrating peaks that are close to each other as the same peak. In the following description, peaks that are close to each other and integrated as the same peak may be referred to as an integrated peak.
[0034] (Noise Model Creation Example 1) Next, as an example of noise model creation of the first embodiment, a method of dividing waveform data into windows of a fixed width and integrating peak information for each window will be described. FIG. 7 is an explanatory diagram illustrating the concept of the noise model creation method of the first embodiment. FIG. 7 shows the concept of integrating peak information for a reference DNA sample using a window width of one base length unit. The horizontal axis of the waveform in FIG. 7 represents the reference position coordinate, and the data is divided into windows with the midpoints between the main peaks at both ends. In other words, the windows are set by dividing the data to have a window width of one base length unit. At the bottom of the figure, if a peak with a relative signal intensity greater than 0 is detected within each window for each of the G, A, T, and C signals, that window is shown in black. In the example shown in FIG. 7, a small C peak (waveform) and a C signal (window: black) can be seen at the position of the second A peak from the left in the reference position coordinate. Furthermore, in the example shown in FIG. 7, a small G peak (waveform) and a G signal (window: black) can be seen at the position of the fourth T peak from the left in the reference position coordinate. The window width here may be any length. It may be shorter than one base length to increase the position resolution of the integrated peak, or it may be longer than one base length to absorb the shift in peak position.
[0035] FIG. 8 is an explanatory diagram illustrating the identification of mutation peaks using peak information. FIG. 8 shows an example in which the above-mentioned window-based peak information integration process is applied to four types of reference DNA samples. For convenience, the figure shows only the A signal out of the four bases A, C, G, and T. In other words, the vertical axis indicates the presence or absence of an A signal peak for each one-base-wide window (the A signal is shown in a black window, and the C, G, and T signals are shown in a white window). FIG. 8 shows peak information integrated in window units for reference DNA samples 1 to 4 (hereinafter referred to as window-based peak information 801).
[0036] Next, the creation of a noise model will be described. The noise model is created using the above-mentioned window-based peak information 801 (this corresponds to step S402 in FIG. 4). Any noise model can be used as long as it can integrate the window-based peak information 801 of multiple reference DNA samples. Here, as a simple example, a union 802 of the window-based peak information 801 of the reference DNA samples is used. However, other integration processes may also be included. Below, examples of other integration processes are shown in FIG. 9.
[0037] 9A is an explanatory diagram illustrating an example of the integration process of window-based peak information in the noise model creation method of the first aspect. In FIG. 9A, peak positions and peak relative signal intensities that are similar across all reference DNA samples are considered to be the same peak. In the noise model shown in FIG. 9A, window-based peak information 801 for four types of reference DNA samples is integrated into one representative peak position (union 802). The representative peak information may store the peak positions and peak relative signal intensities as average values (however, in FIG. 9A, the peak positions are rounded to the nearest base) or medians, respectively.
[0038] 9B is an explanatory diagram illustrating another example of the integration process of window-based peak information in the noise model creation method of the first embodiment. In FIG. 9B, low-frequency peaks 801a that can only be confirmed in a small number of reference DNA samples are removed in the noise model to obtain a union 802. The conditions for excluding low-frequency peaks 801a are that they can only be confirmed below a predetermined threshold (frequency for an arbitrary number of samples or for all samples), and that, as shown in FIG. 9A, no peaks that appear to be integrable exist near the low-frequency peaks 801a. By removing such noise with low reproducibility between samples in advance, the probability of preventing the detection of a mutation peak, as described below, can be reduced.
[0039] (Mutant Peak Detection Method 1) Next, a mutant peak detection method will be described. First, test data, which is waveform data of a test sample, is obtained (this corresponds to step S403 in FIG. 4). Next, as described above, the base sequences of the reference DNA sample data and the test data are compared, and the waveform data is aligned so that the same bases are paired (this corresponds to step S404 in FIG. 4). Next, the union 802 ( FIG. 8 ) of the window-unit peak information 801 of the noise model is compared with the window-unit peak information 803 ( FIG. 8 ) of the test sample, base by base. Specifically, as described above, the aligned waveform data of the reference DNA sample and the waveform data of the test DNA sample are compared (step S405 in FIG. 4 ). Peaks detected at the same positions are then excluded from the peak information of the test sample. The differential peaks that remain (are recognized) specific to the test data are mutant peaks 804 ( FIG. 8 ) (this corresponds to step S406 in FIG. 4 ).
[0040] In this embodiment, the window width is set to one base unit as described above, but the window width may be any length. That is, in this embodiment, the window width may be set to be shorter than one base length to increase the position resolution of the integrated peak, or may be set to be longer than one base length to absorb the shift in peak position. By widening the window width, it is expected that mutation peaks can be detected even if the accuracy of correction of the peak top detection position is low.
[0041] In the above embodiment, the ratio to the maximum value of each fluorescent dye is used as the relative signal intensity, but the ratio to the main peak at the same position may also be used. In this embodiment, the frequency of mutations can be quantified.
[0042] Furthermore, while all peaks in the peak information are used in the above embodiment, it is also possible to use only peaks located close to the main peak. Among peaks with low signal intensity, those located far from the main peak, i.e., peaks whose positions are shifted, are likely to be noise. Therefore, this embodiment improves the accuracy of noise removal.
[0043] In the above-described embodiment, all peaks in the peak information are used, but a threshold value may be set for the signal intensity, and only peaks above the threshold value may be used. The threshold value can be set for the relative signal intensity or the fluorescence intensity. In this embodiment, minute noise generated by the device or consumables can be removed.
[0044] (Noise Model Creation Example 2) In the above-described embodiment, waveform data was divided into windows of a fixed width and peak information was integrated for each window, but the resolution of peak positions for each window decreases. Therefore, in a second noise model creation embodiment, noise model data may be obtained by clustering peak information for all reference sample DNAs based on peak positions and relative signal intensities, without providing windows.
[0045] The concept of clustering will be explained with reference to FIG. 10. Both FIG. 10A and FIG. 10B are explanatory diagrams illustrating the concept of the noise model creation method of the second aspect. In FIG. 10A, peaks detected in all reference sample DNAs are plotted based on peak positions (reference positions) (reference position coordinate system) and relative signal intensities. In the figure, peaks with similar peak positions and relative signal intensities are considered to be the same and set to the same cluster. In the figure, the clusters are clustered into clusters A, B, C, D, and E. The clustering method will be described later. Each cluster has information on the number of peaks it contains (referred to as the cluster size) and the representative peak of that cluster. Typical examples of the representative peak may be the positions of the peaks contained in the cluster, and the average or median of the relative signal intensities.
[0046] FIG. 10B shows the concept of the representative peak of each cluster obtained in this manner. Note that in FIG. 10B, cluster A, which has a small cluster size, has been removed. This is for the same reason as the process used to remove low-frequency peak 801a, which can only be confirmed in a low number of reference DNA samples, in the noise model creation example 1 described above. By removing noise with low reproducibility between multiple reference DNA samples in advance, it is possible to reduce the number of undetected mutation peaks, as described below. Clustering is performed as shown in FIG. 10B, and the data obtained by calculating the representative peak of each cluster is used as noise model 1001.
[0047] (Example of Clustering) Next, an example of the clustering process will be shown with reference to Fig. 11 and Fig. 12. Fig. 11 is a flowchart illustrating the details of the noise model creation method of the second aspect. Fig. 12A and Fig. 12B are both explanatory diagrams illustrating cluster integration in the noise model creation method of the second aspect.
[0048] First, a different cluster number is assigned to each peak detected in all reference DNA samples. That is, in the initial cluster setting, all peaks are set as individual clusters. In this case, cluster numbers are assigned in ascending (or descending) order of peak position, and within the same peak position, in descending order of relative signal intensity (step S1101). Next, a counter i for the cluster number to be processed is set to 0 (step S1102). Next, the current cluster is set to cluster (i) and the next cluster is set to cluster (i+1) (step S1103), and it is determined whether or not to merge these clusters (step S1104).
[0049] The concept of cluster integration and its integration criteria will now be described with reference to FIGS. 12A and 12B . In the left diagram of FIG. 12A , peaks belonging to the current cluster are indicated by a "+" and peaks belonging to the next cluster are indicated by a "○." It is determined whether these peaks should be integrated into a single cluster (the next cluster indicated by a "○" in the right diagram of FIG. 12A ) (step S1104, as described above). FIG. 12B shows an example of the integration criteria. In FIG. 12B , if the maximum distance Dp between the positions (reference positions) of all peaks belonging to the current cluster and the next cluster (adjacent clusters) and the maximum difference Ds in relative signal strength are less than thresholds THp and THs (Dp<THp and Ds<THs), respectively, and the integration criteria are satisfied (Yes in step S1105 of FIG. 11 ), the peaks are integrated into a single cluster (the next cluster indicated by a "○" in the right diagram of FIG. 12A ) (S1106). The thresholds THp and THs are preset. By appropriately setting each threshold value depending on the electrophoresis conditions and experimental conditions, subsequent analysis can be performed without being affected by variations in peak position or relative signal intensity. For example, the smaller the relative signal intensity, the smaller the resolution of the relative signal intensity of the cluster, so THs may be set according to the relative signal intensity. During cluster integration, as shown in the right diagram of Figure 12A, by assigning the number of the next cluster to the peaks belonging to the current cluster, all peaks are considered to belong to the next cluster. Then, the representative peak information of the next cluster after integration is updated.
[0050] After steps S1101 to S1106 are performed, or if the integration condition Dp<THp and Ds<THs is not satisfied (No in step S1105), it is checked whether the last cluster has been reached (step S1107). If the last cluster has not been reached (No in step S1107), the cluster number counter i is updated to i+1 (step S1108), and the process returns to step S1103 to perform the above-described processes again. On the other hand, if the last cluster has been reached (Yes in step S1107), the process ends (end). Note that the above is an example of clustering processing, and is not necessarily limited to this. Various conditions may be applied to determine whether peaks are the same.
[0051] (Mutation Peak Detection Method 2) Next, a mutation peak detection method based on the noise model creation method of the second aspect described above will be described. First, test data, which is waveform data of a test sample, is acquired (this corresponds to step S403 in FIG. 4). Next, peak information (representative peaks) of the noise model 1001 (FIG. 10B) is compared with peak information of the test sample (this corresponds to step S405 in FIG. 4). Then, a peak whose peak position and relative signal intensity in both cases satisfy the mutation peak determination conditions described below is determined as a mutation peak 804 (this corresponds to step S406 in FIG. 4).
[0052] (Mutation Peak Determination Conditions) Next, examples of conditions for determining a mutation peak will be shown with reference to Fig. 13 . In this embodiment, determination is made under the following two conditions. Fig. 13A is an explanatory diagram illustrating an example of a mutation peak detection method using a noise model of the second embodiment. Fig. 13B is an explanatory diagram illustrating another example of a mutation peak detection method using a noise model of the second embodiment.
[0053] (1) Peak Height and Position in Test Sample Using FIG. 13A , an example of the conditions under which peak 1302 in a test sample becomes a mutation candidate will be described. (Condition 1-1) A main peak 1301 of another base must be present within ±D1 of the peak position Xs of peak 1302. (Condition 1-2) When the height of peak 1302 is Hts and the height of main peak 1301 is Hm, Hts / Hm must exceed a predetermined threshold. Note that D1 is the allowable range for the deviation between the mutation peak position and the main peak position. The predetermined threshold for Hts / Hm can be set arbitrarily. When both of the above conditions are met, peak 1302 becomes a mutation candidate because it has a signal intensity above a certain level relative to the signal intensity of main peak 1301, which is within a predetermined range. In other words, when the ratio of the signal intensity of the peak waveform of the test data to the signal intensity of the representative peak (main peak 1301) of multiple reference DNA samples (corresponding to the aforementioned Hts / Hm) is below a threshold value, the peak waveform of the test data is excluded from the mutation candidates.
[0054] Among peaks with low signal intensity, mutation peaks are likely to be located close to the main peak, while noise peaks are likely to be located farther from the main peak. In other words, noise peaks are likely not to exist within ±D1 of the peak position Xs of peak 1302. Therefore, condition 1-1 improves the accuracy of noise removal. Condition 1-2 sets a lower limit on the frequency of mutations to be detected, thereby making it possible to remove noise peaks with low signal intensity.
[0055] (2) Peak Height and Position in Reference DNA Sample Using FIG. 13B, an example of the conditions under which a peak 1302 in a test sample that satisfies conditions 1-1 and 1-2 is excluded from mutation candidates is described. (Condition 2-1) A representative peak 1303 of the same base in the reference DNA sample must be present within ±D2 of the peak position Xs of peak 1302. (Condition 2-2) When the height of the representative peak 1303 that satisfies condition 2-1 is Hns and the height of peak 1302 is Hs, Hns / Hs must exceed a predetermined threshold. Note that D2 is the allowable range for deviation of noise peaks that are considered to be identical. The predetermined threshold for Hns / Hs can be set arbitrarily.
[0056] The above conditions are used to determine whether the condition that a mutation peak is present specifically only in the test sample is satisfied. That is, if a representative peak 1303 that satisfies conditions 2-1 and 2-2 for peak 1302 in the test sample is present in the reference DNA sample, peak 1302 is excluded from the mutation candidates. This means that a peak corresponding to peak 1302 is also present in the reference DNA sample, and therefore it is determined that the peak is not derived from a mutation. Note that the above conditions for determining a mutation peak are not necessarily limited to these, and various conditions may be applied to more accurately distinguish between a mutation peak and other noise.
[0057] (Noise Model Creation Example 3) The above-mentioned conditions 2-1 and 2-2 are intended to detect noise peaks in a reference DNA sample, but noise signals typically have low signal intensity and may not form clear peaks. Here, FIG. 14A is an explanatory diagram illustrating a case where a peak is unclear in the noise model of the second aspect. FIG. 14A shows an example in which, in a certain reference DNA sample, a noise signal 1404 with significant signal intensity is present near the peak position Xs of peak 1402 in the test sample, but the detected peak 1403 (two peaks 1403a and 1403b are present in the figure) is not present at a position that satisfies condition 2-1. In such a case, peak 1402 is present specifically only in the test sample, and is therefore determined to be a mutation peak. However, if noise signal 1404 with significant signal intensity is present at the same position with high reproducibility in multiple reference DNA samples, it is more appropriate to determine that peak 1402 is highly reproducible noise and not a peak derived from a mutation.
[0058] Since the noise models of the first and second aspects described above both retain only peak information, they cannot detect noise from noise signals with unclear peaks, as described above, and may erroneously detect mutation peaks. Therefore, in the example of noise model creation of the third aspect described below, the noise model data is created by creating an integrated signal that represents the signals of multiple reference DNA samples, and this is added to the noise model 1001 of the second aspect ( FIG. 10B ), resulting in the noise model 1401 of the third aspect ( FIG. 14B ). Note that FIG. 14B is an explanatory diagram illustrating the concept of the noise model creation method of the third aspect.
[0059] (Creating an Integrated Signal) A method for generating an integrated signal of multiple reference DNA samples will be described using FIG. 15 . FIG. 15 is an explanatory diagram showing an example of generating an integrated signal in the noise model creation method of the third aspect. The upper diagram in the figure shows an enlarged view of the signal points from reference positions P-2 to P+2, with the signals of multiple reference DNA samples superimposed. As an example of a method for generating an integrated signal, a signal point corresponding to a predetermined percentile value is selected from the signal points of multiple samples at each reference position as a representative value for each reference position. As a typical example, the 50th percentile value (median) is selected as the representative value for each reference position, and a representative value is calculated for the positions of all points that make up the signal, and this is used as the integrated signal (lower diagram in FIG. 15 ). Using percentile values enables subsequent analysis without being affected by variations in signal strength or outliers.
[0060] (Mutation Peak Determination Condition Using Integrated Signal) (3) Height of Integrated Signal in Reference DNA Sample In addition to the two mutation peak determination conditions (1) and (2) in the noise model of the second aspect described above, the following mutation peak determination condition 3 using an integrated signal will be described with reference to Fig. 16. Fig. 16 is an explanatory diagram illustrating an example of a mutation peak detection method using the noise model of the third aspect.
[0061] (Condition 3) The integrated signal 1605 of the reference DNA sample has a point where Hns / Hs exceeds a predetermined threshold within ±D3 of the peak position Xs of the peak 1602 in the test sample. Note that in FIG. 16, D3 is the range of the integrated signal considered to be near the peak. Hs is the height of the peak 1602. Hns is the lower limit of the relative signal intensity of the integrated signal considered to be a significant peak. The predetermined threshold for Hns / Hs can be set arbitrarily.
[0062] Condition 3 is used to determine whether the condition that a mutation peak exists specifically only in the test sample is satisfied. That is, as described above, even if a noise peak satisfying conditions 2-1 and 2-2 is not detected, if integrated signal 1605 satisfies condition 3, peak 1602 in the test sample is excluded from the mutation candidates. That is, a peak (peak 1602) of the same base located near integrated signal 1605 (representative peak of the reference DNA sample) that is considered to be a significant peak is excluded from the mutation candidates. In other words, this means that integrated signal 1605 corresponding to peak 1602 is also present in the reference DNA sample, and therefore peak 1602 is determined to be a highly reproducible noise signal and not a peak derived from a mutation. Note that the conditions for determining a mutation peak based on integrated signal 1605 described above are not necessarily limited to these, and various conditions may be applied to more accurately distinguish between mutation peaks and other noise.
[0063] (Mutation Detection Apparatus) Next, a mutation detection apparatus according to one embodiment of the present invention (hereinafter, sometimes simply referred to as "the apparatus") will be described with reference to Figures 17 and 18. Figure 17 is a block diagram illustrating an example of the configuration of the apparatus. Figure 18 is a block diagram illustrating an example of the secondary analysis software shown in Figure 17.
[0064] The apparatus 1700 shown in Fig. 17 detects mutations in bases on DNA strands, i.e., mutations in DNA base sequences in sequence data. As shown in Fig. 17, the apparatus 1700 includes a capillary electrophoresis apparatus 1701 and a data analysis computer 1709. The capillary electrophoresis apparatus 1701 includes an apparatus control unit 1702, a storage device 1703, primary analysis software 1704, a communication device 1705, a signal detection unit 1706, an electrophoresis unit 1707, and a sample presentation unit 1708.
[0065] The device control unit 1702 controls the operation of each component, function, and processing unit included in the capillary electrophoresis device 1701. The device control unit 1702 is composed of, for example, a central processing unit (CPU) that executes a program that controls the operation of each component, function, and processing unit.
[0066] The storage device 1703 stores primary analysis software 1704. The storage device 1703 also stores the programs and various data such as signals detected by the signal detection unit 1706 and sequence data. Examples of the storage device 1703 include random access memory (RAM), a hard disk drive (HDD), and a solid state drive (SSD).
[0067] The primary analysis software 1704 creates waveform data based on the signal detected by the signal detection unit 1706, and assigns data on the bases corresponding to the waveform data (each peak), specifically data on the type of A, C, G, or T, and the letters A, C, G, or T corresponding to the bases.
[0068] In this apparatus 1700, primary analysis software 1704 causes apparatus control unit 1702 (i.e., a computer) to function as waveform data acquisition means (not shown). By functioning as waveform data acquisition means, this apparatus 1700 acquires waveform data including at least base sequence information and information on the position of peak waveform appearances using multiple reference DNA samples that do not contain mutations or whose mutation positions are known. The waveform data acquisition means corresponds to the waveform data acquisition step S101 described above.
[0069] Furthermore, in this apparatus 1700, the primary analysis software 1704 causes the apparatus control unit 1702 (i.e., the computer) to function as test data acquisition means (not shown). By functioning as the test data acquisition means, this apparatus 1700 acquires test data, which is waveform data of the test sample. The test data acquisition means corresponds to the test data acquisition step S103 described above.
[0070] The communication device 1705 is connected to the communication device 1710 of the data analysis computer 1709 via a wired or wireless connection, and transmits and receives data and signals to and from the communication device 1705. The data includes sequence data such as waveform data including base sequence information and information on the position at which peak waveforms appear. The communication device 1705 may also be connected to another computer via a wired or wireless internet connection. The communication device 1705 may also be connected to a peripheral device such as a printer via a wired or wireless connection.
[0071] The signal detection unit 1706 detects fluorescent signals from labeled substances derived from labeled ddNTPs in the reaction products after the cycle sequencing reaction and purification process, which have been separated into DNA molecule groups by the electrophoresis unit 1707. The electrophoresis unit 1707 separates the reaction products after the cycle sequencing reaction and purification process into DNA molecule groups by the molecular sieving effect of the separation medium. As described above, the electrophoresis unit 1707 generally performs denaturing polyacrylamide gel electrophoresis, capillary electrophoresis, or the like. The sample presentation unit 1708 is a section where the reaction products after the cycle sequencing reaction and purification process are temporarily stored in order to be provided to the electrophoresis unit 1707.
[0072] The data analysis computer 1709 includes a communication device 1710 , a storage device 1711 , secondary analysis software 1712 , a printer 1713 , a RAM 1714 , an image output unit 1715 , an input interface 1716 , an input device 1717 , and an image display device 1718 .
[0073] The communication device 1710 is connected to the communication device 1705 of the capillary electrophoresis device 1701 by wire or wirelessly, and transmits and receives data and signals to and from the communication device 1705. The communication device 1710 may also be connected to another computer by wire or wirelessly via the Internet.
[0074] The storage device 1711 stores secondary analysis software 1712. The storage device 1711 also stores, for example, programs that control the operation of each configuration, function, and processing unit, and stores sequence data received via the communication device 1710. Examples of the storage device 1711 include an HDD and an SSD.
[0075] The secondary analysis software 1712 causes a computer, specifically a central processing unit (CPU) 1719, to read the contents of the software and cause it to function as noise model data creating means, aligning means, and excluding means.
[0076] As an example, in this apparatus 1700, the secondary analysis software 1712 causes the CPU 1719 (i.e., the computer) to function as noise model data creation means (not shown). By functioning as noise model data creation means, this apparatus 1700 creates noise model data using a set of peak waveforms of waveform data from multiple reference DNA samples. The noise model data creation means corresponds to the noise model data creation step S102 described above.
[0077] Next, in the present apparatus 1700, the secondary analysis software 1712 causes the CPU 1719 (i.e., the computer) to function as alignment means (not shown). By functioning as alignment means, the present apparatus 1700 compares the base sequences of the noise model data and the test data, and aligns the test data with the noise model data so that the same bases are paired. The alignment means corresponds to the alignment step S104 described above.
[0078] Next, in the present apparatus 1700, the secondary analysis software 1712 causes the CPU 1719 (i.e., the computer) to function as an exclusion means (not shown). By functioning as the exclusion means, the present apparatus 1700 excludes peak waveforms from the test data that appear at the same positions as those of the noise model data. The exclusion means corresponds to the above-mentioned exclusion step S105.
[0079] The printer 1713 prints out on paper the analyzed DNA base sequence, waveform data, information on detected low-frequency base mutations on the DNA strand, and the like. The RAM 1714 is used when the CPU 1719 performs some processing or displays some data on the screen. The image output unit 1715 is a device (terminal) that enables connection to an image display device 1718. Examples of the image output unit 1715 include a VGA terminal, a DVI terminal, an HDMI (registered trademark) terminal, and a DisplayPort. The input interface 1716 is a device that enables connection to an input device 1717. Examples of the input interface 1716 include a connection device conforming to the Universal Serial Bus (USB) standard. The input device 1717 is generally a keyboard, a mouse, or the like. The image display device 1718 is generally a monitor, or the like.
[0080] (Mutation Detection Program) Next, a mutation detection program according to one embodiment of the present invention (hereinafter, sometimes simply referred to as "this program") will be described with reference to Fig. 18. As mentioned above, Fig. 18 is a block diagram illustrating an example of the secondary analysis software shown in Fig. 17, but it also represents a block diagram illustrating an example configuration of this program.
[0081] This program detects mutations in bases on DNA strands, i.e., mutations in DNA base sequences in sequence data. As shown in Fig. 18, this program is stored in a storage device 1711 as secondary analysis software 1712. This program, i.e., secondary analysis software 1712, has a noise model data creation step S102, an alignment step S104, and an exclusion step S105.
[0082] 18 , this program causes the RAM 1714 to acquire waveform data including at least base sequence information and peak waveform appearance position information from a waveform data acquisition step S101 using a plurality of reference DNA samples that do not contain mutations or whose mutation positions are known, and then causes the computer (CPU 1719) to execute a noise model data creation step S102. The noise model data creation step S102 creates noise model data using a set of peak waveforms from the waveform data of the plurality of reference DNA samples. The noise model data creation step S102 of this program can be executed in the same manner as the noise model data creation step S102 of the method described above. This embodies the noise model data creation means described above.
[0083] Next, after the RAM 1714 acquires test data, which is waveform data of the test sample, from the test data acquisition step S103, the program causes the computer (CPU 1719) to execute the alignment step S104 and the exclusion step S105. The alignment step S104 compares the base sequences of the noise model data and the test data, and aligns the test data with the noise model data so that the same bases are paired. Furthermore, the exclusion step S105 excludes peak waveforms from the test data that appear at the same positions as those in the noise model data. Note that the alignment step S104 and exclusion step S105 of the program can be executed in the same manner as the alignment step S104 and exclusion step S105 of the method described above. This embodies the alignment means and exclusion means described above.
[0084] (Mutation Detection Storage Medium) Next, a mutation detection storage medium according to one embodiment of the present invention (hereinafter, simply referred to as "this storage medium") will be described with reference to Fig. 19. Fig. 19 is a block diagram illustrating an example of the configuration of this storage medium.
[0085] This storage medium detects mutations in bases on DNA strands, i.e., mutations in DNA base sequences in sequence data. As shown in Fig. 19, this storage medium 1901 stores secondary analysis software 1712. The secondary analysis software 1712 stored in this storage medium 1901 includes a noise model data creation step S102, an alignment step S104, and an exclusion step S105.
[0086] 19 , in this storage medium 1901, after the RAM 1714 acquires waveform data including at least base sequence information and peak waveform appearance position information from a plurality of reference DNA samples that do not contain mutations or whose mutation positions are known, from waveform data acquisition step S101, the storage medium 1901 causes the computer (CPU 1719) to execute noise model data creation step S102. The noise model data creation step S102 creates noise model data using a set of peak waveforms from the waveform data of the plurality of reference DNA samples. Note that the noise model data creation step S102 of this storage medium 1901 can be executed in the same manner as the noise model data creation step S102 of the present method described above. This embodies the noise model data creation means described above.
[0087] Next, in this storage medium 1901, after the RAM 1714 acquires test data, which is waveform data of the test sample, from the test data acquisition step S103, the computer (CPU 1719) executes the alignment step S104 and the exclusion step S105. The alignment step S104 compares the base sequences of the noise model data and the test data, and aligns the test data with the noise model data so that the same bases are paired. Furthermore, the exclusion step S105 excludes peak waveforms of the test data that appear at the same positions as those of the noise model data. Note that the alignment step S104 and exclusion step S105 of this storage medium 1901 can be executed in the same manner as the alignment step S104 and exclusion step S105 of the present method described above. This embodies the alignment means and exclusion means described above.
[0088] As described above, the mutation detection method, device, program, and storage medium according to one embodiment of the present invention have the above-described configuration. In particular, this embodiment includes a noise model data creation step S102 (noise model data creation means), an alignment step S104 (alignment means), and an exclusion step S105 (exclusion means). The peak remaining after these steps (means) is the mutation peak. Therefore, this embodiment can accurately detect low-frequency mutations of bases on DNA strands.
[0089] (Comparison with the techniques described in documents 1 and 2) Here, techniques using the term "noise" are described in, for example, document 1 (JP 2000-131284 A) and document 2 (JP 2016-168707 A). Therefore, this embodiment will be compared with the techniques described in documents 1 and 2 (hereinafter referred to as "document technique 1" and "document technique 2", respectively), and the differences between this embodiment and document techniques 1 and 2, the advantages of this embodiment, and the like will be described.
[0090] The documented technique 1 aims to eliminate the influence of background noise in a chromatography mass spectrometer and obtain an accurate chromatogram. To achieve this, the documented technique 1 includes: a) a storage means for storing, as noise data, data obtained by introducing only a solvent containing no sample components and performing chromatography mass analysis under predetermined measurement conditions prior to measuring a target sample; b) a calculation means for reading noise data at the same retention time and mass number from the storage means and subtracting the data obtained by introducing the target sample and performing chromatography mass analysis under the predetermined measurement conditions; and c) a data processing means for creating a mass spectrum or chromatogram based on the data subtracted by the calculation means. In other words, the documented technique 1 includes: a) a solvent containing no sample components is measured under the same measurement conditions prior to measuring a target sample; and the noise data obtained thereby is compared with the measurement data of the target sample.
[0091] In the literature technique 1, noise data is generated using only a solvent that does not contain sample components. In contrast, in the present embodiment, noise data is generated using a sample solution that contains sample components (corresponding to a reference DNA sample). This difference arises from whether or not issues specific to DNA sequencers, such as capillary electrophoresis, are resolved. One issue specific to DNA sequencers is that the DNA sample used for analysis itself contains noise. This noise is caused by the contamination of the starting sample with external samples and side reactions of enzymes used during sample preparation. Both of these issues inevitably arise during processes unique to DNA analysis, such as the difficulty of obtaining highly pure samples from living organisms and the replication of DNA molecules over 100 million times during sample preparation. The noise inherent to DNA sequencers that is targeted for removal in this embodiment cannot be predicted from literature technique 1. Furthermore, the noise targeted for removal in this embodiment cannot be detected using only a solvent that does not contain sample components.
[0092] Furthermore, Document Technology 2 aims to provide an image forming apparatus having a voice recognition function with a function for performing good voice recognition based on voice received during job execution. To solve this problem, Document Technology 2 includes a generation unit that generates noise data by simulating the operating sounds of moving parts that operate during job execution, a voice input unit that inputs voice, and a voice data generation unit that receives the voice input to the voice input unit and generates voice data. Document Technology 2 also includes a removal unit that removes the noise data from the voice data generated by the voice data generation unit during job execution, and a recognition unit that performs voice recognition based on the voice data from which the noise data has been removed by the removal unit. As described above, Document Technology 2 is characterized by generating noise data by simulating the operating sounds of moving parts that operate during job execution.
[0093] In Document Technology 2, noise data is created using noise caused by the device. In contrast, in the present embodiment, noise data is created using data caused by wild-type DNA, i.e., a reference DNA sample. This difference, as above, arises from whether or not issues specific to DNA sequencers, such as capillary electrophoresis, are resolved. Therefore, for the same reasons as above, the noise generated specifically by DNA sequencers, which is the target of removal in this embodiment, cannot be predicted from Document Technology 2. Note that the noise to be removed in this embodiment is not detected from the operating sounds of the device.
[0094] As described above, the mutation detection method, device, program, and storage medium according to the present embodiment, which can accurately detect low-frequency mutations in bases on DNA strands, cannot be conceived based on literature techniques 1 and 2, etc., which make no mention whatsoever of the issues inherent to DNA sequencers, such as capillary electrophoresis.
[0095] The effect of this embodiment was confirmed using the BRAF gene as a target. Genomic DNA from HeLa cells was used as a reference DNA sample. Genomic DNA (HD238, Horizon discovery) containing a single base substitution in the BRAF gene was used as a test sample.
[0096] First, the BRAF region was amplified by PCR. Primer 2001 (SEQ ID NO: 1 in the Sequence Listing) and primer 2002 (SEQ ID NO: 2 in the Sequence Listing) shown in Figure 20 were used for amplification. Figure 20 is a table showing information on the primers used in the examples. Four types of reference DNA samples were prepared. Specifically, 10 ng of each reference DNA sample was dispensed into four tubes and amplified individually. Next, the PCR product was purified according to standard methods, and the concentration of the obtained PCR product was quantified.
[0097] Next, 0.95 ng of PCR product from the reference DNA sample was mixed with 0.05 ng of PCR product from the mutant sample to create a sample simulating a test sample with a 5% mutation rate. The test sample and four reference DNA samples were subjected to cycle sequencing reactions. The reactions were performed using sequence primer 2003 (Figure 20) (SEQ ID NO: 3 in the Sequence Listing) and commercially available sequencing reaction reagents. The reaction products were purified by ethanol precipitation and dissolved in 10 μL of formamide.
[0098] The resulting DNA was subjected to a capillary electrophoresis apparatus to obtain waveform data. Electrophoresis was performed for 900 seconds. Fluorescence signals were acquired every 120 milliseconds, yielding a total of 7,500 frames of data, equivalent to approximately 470 bases.
[0099] Next, the obtained waveform data was processed. Due to the characteristics of electrophoresis, highly reliable data cannot be obtained from the beginning and end of the data. Therefore, the parts of the waveform data from 1 to 47 bases and from 423 to 470 bases were deleted.
[0100] Next, peak information was obtained from each waveform data. Furthermore, the base sequences of the four reference DNA samples and the test sample were compared, and the waveform data was aligned so that the same bases were aligned. Since the reference DNA sample did not contain any mutations, the peak information of the reference DNA sample obtained here was the main peak and noise. Next, each waveform data was divided by base, and the signal intensity was corrected, with the maximum value set to 1. Hereinafter, the corrected signal intensity will be referred to as the relative signal intensity.
[0101] Next, a noise model of the first aspect described above was created with a window width of one base. Here, the noise model was the union of the peak information of four reference DNA samples. Next, the noise model data was compared with the test sample data, and peak information (peak waveforms) of the test sample detected at the same position as the noise model were excluded. A control experiment was also conducted in parallel using a system that did not use the noise model. In the control experiment, peak information obtained from one specific reference DNA sample out of the four reference DNA samples was used. Furthermore, to improve the accuracy of noise detection, a threshold value of 0.04 arb. U relative signal intensity was set. In this system, only peaks with a relative signal intensity greater than the threshold were considered mutant peaks. Figure 21 shows the results obtained.
[0102] FIG. 21 is a plot diagram showing the results of the example. In FIG. 21, the horizontal axis represents base length and the vertical axis represents relative signal intensity (arb. U), and all peak information that was not excluded in comparison with the reference DNA sample is plotted. Of these, peak information derived from mutant bases is indicated by cross plot 2101. In both plot 2102 (without model) of the control experiment and plot 2103 (with model) of this example, multiple peak information (noise) was observed in addition to the mutation peak information (cross plot 2101). However, in plot 2103 of this example, the only peak larger than threshold 2104 was the mutation peak. From the above, in this example, mutation peaks that appear at low frequencies were correctly identified. On the other hand, multiple noises larger than threshold 2104 were observed in plot 2102 of the control experiment, so mutation peaks could not be correctly identified.
[0103] [Second Embodiment] (Application to STR Analysis) In the above-described embodiment, a method for distinguishing between noise peaks and low-frequency mutation peaks was described, in which a noise model was created using multiple reference DNA samples, and the noise model was compared with that of a test sample, targeting mutations in DNA base sequences. This method can be applied not only to mutations in DNA base sequences, but also to distinguishing low-frequency mutations in the form of DNA polymorphisms such as short tandem repeats (STRs) or microsatellites. Below, an embodiment of the present invention for STR analysis will be described.
[0104] (Overview of STR analysis) STR is a characteristic sequence pattern in which a short sequence of about 2 to 7 bases in length is repeated several to several tens of times, and it is known that the number of repeats varies from individual to individual. Analyzing the combination of the number of STR repeats at the locus of a specific gene is called STR analysis.
[0105] In DNA testing for purposes such as criminal investigations, STR analysis is used, taking advantage of the property that the combination of STR repeat numbers varies between individuals. The FBI (Federal Bureau of Investigation) and the International Criminal Police Organization (ICP) define 10 to several dozen STR loci (locuses) used in DNA testing as DNA markers, and analyze the repeat number patterns of these STR sequences. Because differences in the number of STR repeats occur due to differences in alleles (alleles), hereafter, the number of STR repeats in each DNA marker will be referred to as an allele. In STR analysis, DNA fragments at the target locus are amplified by PCR, and the DNA fragments are subjected to electrophoresis to measure the DNA fragment length. The repeat number (allele) is identified from the obtained DNA fragment length. The methods used for determining the base sequence described above can be applied to this electrophoresis.
[0106] FIG. 22 shows an example of a fluorescence intensity waveform of a sample obtained by electrophoresis. FIG. 22 is an explanatory diagram showing an example of an electrophoresis waveform of STR analysis. The figure shows the intensity waveforms of five types of fluorescent dyes (6FAM, VIC, NED, PET, and LIZ) as an example. The horizontal axis of the figure represents electrophoresis time, which is related to DNA fragment length as described below. The time at which each fluorescence intensity peak occurs corresponds to the length of the DNA fragment labeled with each fluorescent dye, and differences in this length correspond to differences in alleles. The fluorescence intensity waveform in FIG. 22 contains one or two peaks for each DNA marker. When there is one peak, the fluorescence intensity of that peak is higher than the fluorescence intensity of a marker with two peaks. When there is one peak, it indicates homozygosity (the paternal allele and the maternal allele are the same), and when there are two peaks, it indicates heterozygosity (the paternal allele and the maternal allele are different). If mutations occur in the DNA of a sample, there may be three or more peaks for one DNA marker depending on the number of mutations.
[0107] (Size Call) In Figure 22, loci are assigned to the four fluorescent dyes 6FAM, VIC, NED, and PET, and the DNA fragment length is calculated from the detection time of the peak at each locus. This process is called size call. Size call is a process that correlates the time required for a DNA fragment to be detected as a peak by electrophoresis with the base length of the DNA fragment (hereinafter referred to as DNA base length). Specifically, electrophoresis is performed on a reagent called a size standard, which contains DNA fragments of known length and is labeled with a specific fluorescent dye. In Figure 22, the size standard is labeled with the fluorescent dye LIZ.
[0108] The concept of size calling using size standards is explained using Figure 23. Figure 23A is an explanatory diagram illustrating size calling in STR analysis. The size standard reagent illustrated in Figure 23A shows an electrophoretic waveform of known DNA fragments between 80 bp and 480 bp in length, labeled with the fluorescent dye LIZ. The center positions of peaks detected in the electrophoretic waveform, i.e., peak times, are associated with known DNA fragment lengths. This association is achieved using well-known dynamic programming techniques. A correspondence equation between electrophoretic time and DNA base length can be obtained by combining these peak times with known DNA base lengths. Figure 23B is a diagram illustrating how to determine the relationship between DNA electrophoresis time (t) and DNA base length (y), "y = f(t)." The known DNA base lengths of the size standards and their corresponding peak times are plotted, and the relationship y = f(t) that best approximates this plot is determined. A quadratic or cubic polynomial, etc., can be used as f(t), and an approximation that minimizes the squared error can be performed. The relational expression "y = f(t)" between DNA migration time (t) and DNA base length (y) obtained in this way is calculated for all capillaries and stored. Using this relational expression, the DNA base length corresponding to any peak time of the fluorescence intensity waveform measured in each capillary can be calculated.
[0109] (Allelic Ladder) It is known that the migration speed of DNA fragments varies depending on the environment, such as the migration medium, reagent performance, device temperature, and migration voltage. Changes in migration speed result in different measured DNA fragment lengths, making it impossible to accurately identify alleles. Therefore, a standard reagent called an allelic ladder is commonly used to accurately identify alleles in response to fluctuations in migration speed. As described below, an allelic ladder is an artificial sample containing all alleles that may generally be contained in a DNA marker. It can absorb fluctuations in migration speed and fine-tune the correspondence between alleles and DNA fragment lengths. Figure 24 shows an example of a fluorescence intensity waveform obtained by electrophoresis of an allelic ladder. Figure 24 is an explanatory diagram showing an example of the migration waveform of an allelic ladder in STR analysis. In this waveform, all alleles of DNA markers (e.g., D10S1248) for each fluorescent dye appear as peaks. The base lengths of individual alleles can be obtained by performing the size call described above on these peaks.
[0110] (Allele Calling) Allele calling is a process of identifying alleles from the DNA base lengths of individual peaks obtained by size calling from the electrophoretic waveform of the sample shown in Figure 22. The allelic ladder described above is used for allele calling. The basic information of the allelic ladder includes the locus name (Locus) labeled by each fluorescent dye (Dye), the allele name (Allele) contained in that locus, the DNA base length (Length) corresponding to that allele, and the allowable base length range (Min / Max) from the center position of each allele. Figure 25 shows an example of part of the basic information of the allelic ladder. Note that Figure 25 is an explanatory diagram explaining the concept of the basic information of the allelic ladder. In this figure, the DNA marker (locus) D10S1248 is labeled with 6FAM, and its alleles include 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, and 18, with standard DNA base lengths (units: bp) of 77, 81, 85, 89, 93, 97, 101, 105, 109, 113, and 117, respectively. All alleles have a tolerance of plus or minus 0.5 bp. By referencing the basic information of the allelic ladder, it is possible to identify which allele corresponds to each peak detected with each fluorescent dye in the sample. This allelic ladder is typically provided by reagent manufacturers as part of a DNA identification kit. Because fluctuations in electrophoretic speed due to environmental changes accumulate over time, it is recommended that allelic ladders be used in STR analysis with a regular frequency and that the basic information be updated. Depending on the purpose of STR analysis, there may be cases where allele identification is not necessary and only information on the distribution of base lengths is required. In such cases, there is no need to use an allelic ladder.
[0111] (Creating a Noise Model in STR Analysis) As described above, in STR analysis, as in the aforementioned base sequence analysis, the position of the peak (corresponding to the base length) and the relative signal intensity (corresponding to the fluorescence intensity) detected in the electrophoretic waveform of the sample contain information about the genotype. If a low-frequency mutation occurs in the sample, in addition to the peak occurring in the wild-type, a peak of low intensity appears at the position corresponding to the mutant type. However, when a noise peak occurs in the electrophoretic waveform, there is a problem in that it is difficult to distinguish between the low-intensity peak caused by the mutation and the noise peak. In this embodiment, as in the first embodiment described above, the noise peak is reduced by creating a noise model using multiple reference DNA samples, and further, by comparing the noise model with that of the test sample, it is possible to distinguish between the noise peak and the low-frequency mutation peak.
[0112] That is, to create the above-mentioned noise model from multiple DNA samples for STR analysis, as described in the creation of the noise model of the first embodiment, peaks with similar peak base lengths and relative signal intensities may be integrated into one window and low-frequency peaks may be excluded (see FIG. 9 ). In this case, the window width is preferably a small base length, at least 1 bp or less, that allows alleles to be identified without any problems. Alternatively, as described in the creation of the noise model of the second embodiment, clustering may be performed based on the peak base length and relative signal intensity without providing a window. Alternatively, a noise model may be created by creating an integrated signal from electrophoresis waveforms obtained from multiple DNA samples using the method described in the creation of the noise model of the third embodiment. However, in this case, it is desirable to convert the horizontal axis of each electrophoresis waveform to DNA base length using the relationship "y = f(t)" between DNA electrophoresis time (t) and DNA base length (y) obtained by size calling. This allows the difference in electrophoresis speed of each electrophoresis waveform to be unified into base length to generate an integrated signal. Furthermore, when the application involves allele calling by referring to an allelic ladder, multiple peaks identified as the same allele may be clustered as in Example 2 of noise model creation.
[0113] (Mutation Peak Determination in STR Analysis) Mutation peaks in test samples can be detected by applying to STR analysis methods similar to those described for the noise models of the first, second, and third aspects. However, in STR analysis, mutation peaks and peaks without mutations exist at positions with different base lengths, making it difficult to associate mutation peaks with main peaks as in the above-mentioned aspects. For this reason, in STR analysis, peaks with significant intensity that do not exist in the noise model can be calculated as candidates for mutation peaks.
[0114] FIG. 26 shows an example of mutation peak determination in STR analysis. FIG. 26 is an explanatory diagram illustrating the concept of mutation peak detection in STR analysis. The figure shows a created noise model and a group of peaks detected in a certain base length region for a certain fluorescent dye in a test sample. Of these, peaks 2601 and 2603 in the noise model and peaks 2606 and 2608 in the test sample are both located at approximately the same base length and have high relative signal intensities, and are therefore considered to be main peaks. The test sample also contains peaks 2605, 2607, and 2609, which have lower relative signal intensities than the main peaks. Of these, peaks 2605 and 2607 can be determined to be possible mutations because no peaks exist in the noise model at base length positions close to these peaks. The criteria for determining the level of signal intensity to be considered a mutation peak can be determined in advance. One example of such a criterion is the ratio of the signal intensity to that of nearby main peaks.
[0115] If the STR analysis is intended for allele calling and it is assumed that the mutation to be detected can be called as an allele, only peaks present only in the test sample that meet the allele conditions may be considered as mutant peaks. An example of a mutant peak in such a case is shown in FIG. 27. FIG. 27 is an explanatory diagram illustrating the concept of mutant peak detection in STR analysis intended for allele calling. This figure shows an example in which a base length region similar to that of FIG. 26 corresponds to an allele (AL) (8-18) of a specific locus (D10S1248 in this figure). Of these, peaks 2701 and 2703 in the noise model and peaks 2706 and 2708 in the test sample are both present at approximately the same base length, have high relative signal intensities, and are close in base length to allele 10 and allele 14 in the allelic ladder, respectively, and are therefore considered to be main peaks that can be called as alleles 10 and 14, respectively. In the test sample, peaks 2705, 2707, and 2709 are present, which have lower relative signal intensities than the main peaks. Peak 2705 is highly likely to be a noise peak because there is no peak with a base length close to the noise model and no corresponding allele exists (marked as Unknown). Peak 2707 is highly likely to be a noise peak because there is no peak with a base length close to the noise model and it can be called allele 12. Peak 2709 can be called allele 17, but since peak 2704 with the same base length also exists in the noise model, it is not considered a mutation peak.
[0116] As described above, in STR analysis, as in sequence analysis for identifying base sequences, noise models are created from multiple reference DNA samples to remove noise with low reproducibility and improve the accuracy of detecting mutations in test samples. In other words, in this embodiment, low-frequency mutations in bases on DNA strands can also be detected with high accuracy.
[0117] The above describes in detail the mutation detection method, device, program, and storage medium according to one embodiment of the present invention. However, the scope of the present invention is not limited to this and includes various modifications. For example, the above-described embodiment has been described in detail to clearly explain the present invention and is not necessarily limited to including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations. Furthermore, the above-described configurations, functions, processing units, processing means, control means, etc. may be implemented in hardware, in part or in whole, by, for example, designing an integrated circuit. Furthermore, the above-described configurations, functions, etc. may be implemented in software by a processor interpreting and executing a program that realizes each function. Information such as programs, tables, and files that realize each function can be stored in memory, storage devices such as HDDs and SSDs, or storage media such as IC cards, SD cards, and DVDs. Furthermore, the control lines and information lines shown are those considered necessary for explanation, and not all control lines and information lines are necessarily shown in the product. In reality, it is safe to assume that almost all components are interconnected.
[0118] S101 Waveform data acquisition step S102 Noise model data creation step S103 Test data acquisition step S104 Alignment step S105 Exclusion step 201 Sequence data (sequence data of test sample) 202 Sequence data (sequence data of reference DNA sample) 203 Peak waveform 204 Noise 205 Peak waveform 301 Sequence data (sequence data of test sample) 302 Sequence data (sequence data of reference DNA sample) 303 Peak waveform 501 Reference DNA sample 502 Reference reference DNA sample 503 Correspondence relationship 601 Peak waveform (peak waveform of reference reference DNA sample) 602 Peak waveform (peak waveform of other reference DNA samples) 801 Window-based peak information (window-based peak information of noise model) 801a Low-frequency peak 802 Union 803 Window-based peak information (window-based peak information of test sample) 804 Mutation peak 1001 noise model 1301 main peak 1302 peak 1303 representative peak 1401 noise model 1402 peak 1403 peak 1403a peak 1404 noise signal 1602 peak 1605 integrated signal 1700 present apparatus (mutation detection apparatus) 1701 capillary electrophoresis apparatus 1702 apparatus control unit 1703 storage device 1704 primary analysis software 1705 communication device 1706 signal detection unit 1707 electrophoresis unit 1708 sample presentation unit 1709 data analysis computer 1710 communication device 1711 storage device 1712 secondary analysis software 1713 printer 1714 RAM 1715 image output unit 1716 input interface 1717 input device 1718 Image display device 1719 CPU 1901 Main storage medium (mutation detection storage medium) 2001 Primer 2002 Primer 2003 Sequence primer 2101 Plot 2104 Threshold
Claims
1. A method for detecting mutations in bases on a DNA strand, comprising: a waveform data acquisition step of acquiring waveform data including at least base sequence information and information on the position of peak waveforms using a plurality of reference DNA samples that do not contain mutations or whose mutation positions are known; a noise model data creation step of creating noise model data using a collection of peak waveforms from the waveform data of the plurality of reference DNA samples; a test data acquisition step of acquiring test data, which is waveform data of a test sample; an alignment step of comparing the base sequences of the noise model data and the test data, and aligning the test data with the noise model data so that the same bases are paired; and an exclusion step of excluding peak waveforms from the peak waveforms of the test data that appear in the same positions as the noise model data.
2. The mutation detection method according to claim 1, wherein the plurality of reference DNA samples are of different qualities.
3. The mutation detection method according to claim 1, wherein the waveform data of the plurality of reference DNA samples are obtained using a plurality of devices and a plurality of different preparation reagent lots.
4. The mutation detection method described in claim 1, characterized in that in the alignment step, the transformation of any position in a peak other than the main peak is performed by linear interpolation using the correspondence between the main peaks located on both sides of the arbitrary position to obtain the transformed position.
5. A mutation detection method according to claim 1, wherein in the alignment step, the waveform data is aligned only from peak waveform data without using base sequences.
6. A mutation detection method according to claim 1, wherein the relative signal intensity of the waveform data is processed to reduce variations in pulse height that occur between electrophoresis runs.
7. A mutation detection method according to claim 1, wherein the relative signal intensity of the waveform data is a ratio to the maximum value of each fluorescent dye.
8. The mutation detection method according to claim 1, wherein the relative signal intensity of the waveform data is a ratio to a main peak at the same position.
9. The mutation detection method according to claim 1, wherein the waveform data excludes peaks located far from the main peak.
10. A mutation detection method according to claim 1, wherein the waveform data is generated based on a window width obtained by dividing the horizontal axis into sections of an arbitrary length.
11. A mutation detection method according to claim 1, wherein the waveform data is generated by removing low-frequency peaks that can only be confirmed in a small number of reference DNA samples.
12. The mutation detection method according to claim 1, wherein a threshold value is set for the signal intensity of the waveform data, and only peaks above the threshold value are used.
13. The mutation detection method according to claim 1, wherein the noise model data is obtained by clustering peak information of all reference DNA samples using peak positions and relative signal intensities.
14. The mutation detection method according to claim 13, wherein the clusters created by the clustering have information on the number of peaks contained in the cluster and the representative peak.
15. The mutation detection method according to claim 13, wherein the clustering is performed using the maximum distance between the positions of all peaks contained in adjacent clusters and the maximum difference in relative signal intensity.
16. The mutation detection method of claim 1, wherein the noise model data is generated from an integrated signal representative of signals from a plurality of reference DNA samples.
17. A mutation detection method as described in claim 16, characterized in that the integrated signal is a signal point corresponding to a predetermined percentile value from the signal points of multiple samples at each reference position, which is taken as the representative value for each reference position.
18. The mutation detection method according to claim 1, wherein in the excluding step, peaks having a signal intensity equal to or greater than a certain level relative to the signal intensity of the main peak are selected as mutation candidates.
19. The mutation detection method according to claim 1, wherein in the excluding step, peaks of the same base located near the representative peak of the reference DNA sample are excluded from mutation candidates.
20. A mutation detection method as described in claim 1, characterized in that in the exclusion step, when the ratio of the signal intensity of the peak waveform of the test data to the signal intensity of the representative peak of the multiple reference DNA samples is below a threshold, the peak waveform of the test data is excluded from mutation candidates.
21. A device for detecting mutations in bases on a DNA strand, comprising: waveform data acquisition means for acquiring waveform data including at least base sequence information and information on the position at which peak waveforms appear, using a plurality of reference DNA samples that do not contain mutations or whose mutation positions are known; noise model data creation means for creating noise model data using a collection of peak waveforms from the waveform data of the plurality of reference DNA samples; test data acquisition means for acquiring test data, which is waveform data of a test sample; alignment means for comparing the base sequences of the noise model data and the test data, and aligning the test data with the noise model data so that the same bases are paired; and exclusion means for excluding peak waveforms from the peak waveforms of the test data that appear in the same positions as the noise model data.
22. A mutation detection program for detecting mutations in bases on a DNA strand, comprising: a waveform data acquisition step in which waveform data including at least base sequence information and information on the position at which peak waveforms appear is obtained using a plurality of reference DNA samples that do not contain mutations or whose mutation positions are known; a noise model data creation step in which noise model data is created using a set of peak waveforms of the waveform data of the plurality of reference DNA samples; a test data acquisition step in which test data, which is waveform data of a test sample, is obtained; an alignment step in which the noise model data and the test data are compared in base sequence, and the test data is aligned with the noise model data so that the same bases are paired; and an exclusion step in which peak waveforms of the test data that appear in the same positions as the noise model data are excluded.
23. A storage medium for detecting mutations in bases on a DNA strand, comprising: a waveform data acquisition step in which waveform data including at least base sequence information and information on the position at which peak waveforms appear is obtained using multiple reference DNA samples that do not contain mutations or whose mutation positions are known; a noise model data creation step in which noise model data is created using a set of peak waveforms from the waveform data of the multiple reference DNA samples; and a test data acquisition step in which test data, which is waveform data of a test sample, is obtained; an alignment step in which the noise model data and the test data are compared in base sequence, and the test data is aligned with the noise model data so that the same bases are paired; and an exclusion step in which peak waveforms appearing in the same positions as the noise model data are excluded from the peak waveforms of the test data.
Citation Information
Patent Citations
Multi-positioning double tag adapter set for detecting gene mutations, and its preparation method and application
JP2019523638A
Nucleic acid analyzer and nucleic acid analysis method using same
WO2014188887A1
Biopolymer analysis method and biopolymer analysis device
WO2020026418A1
Electrophoresis device and analysis method
WO2021229700A1