Method of designing primer for amplicon methylation sequence analysis, production method, designing device, designing program and recording medium
Patent Information
- Application Number
- JP2024543796
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-06-06
- Filing Date
- 2023-06-06
- Publication Date
- 2025-05-12
AI Technical Summary
Current primer design software is inadequate for bisulfite-treated DNA and multiplex PCR, leading to low primer design success rates and increased primer dimer formation, making it difficult to efficiently analyze DNA methylation patterns.
A method and device for designing primers that involve selecting primer candidate sequences based on specific conditions such as Tm values, YG/CR sequence counts, and junctions with external sequences, using local alignment scores to optimize primer pairs for bisulfite-treated DNA, enhancing primer design success rates and reducing dimer formation.
The method significantly improves primer design success rates and reduces primer dimer formation, enabling more effective analysis of DNA methylation by allowing for the simultaneous amplification of multiple target sites with high accuracy.
Abstract
Description
Design method, manufacturing method, design device, design program and recording medium for primers for amplicon methylation sequence analysis
[0001] The present invention relates to a method, a manufacturing method, a design device, a design program, and a recording medium for designing primers for amplicon methylation sequence analysis, and more particularly to a method, a manufacturing method, a design device, a design program, and a recording medium for designing primers for simultaneously amplifying multiple amplification target regions, each containing multiple target sites, in bisulfite-treated or enzyme-treated DNA (deoxyribonucleic acid) by multiplex PCR (polymerase chain reaction).
[0002] DNA methylation is known as one of the epigenetic mechanisms that regulates gene expression without altering DNA base sequences. DNA methylation in mammals occurs primarily at the 5th carbon atom of cytosine (C) in CG sequences in DNA. Regions called CpG islands, where CG sequences frequently occur, are often found in gene promoter regions. Initially, many of the CG sequences in these regions are unmethylated, but they undergo methylation in association with disease, development, differentiation, inflammation, aging, and other conditions, resulting in the suppression of gene expression. For example, in cancer cells, increased methylation of CpG islands in gene promoter regions is known to inactivate many cancer-suppressing genes.
[0003] As described above, DNA methylation is closely related to the regulation of gene expression, and its information is useful for elucidating the mechanisms of diseases such as cancer, evaluating the differentiation state of various cells, etc., and has therefore attracted attention in various fields, such as diagnosis, treatment, drug discovery, and regenerative medicine, and active research and development is being conducted on this subject. For example, by measuring and analyzing the state of DNA methylation in specific regions, attempts are being made to examine the presence or absence of drug resistance for each type of cell when developing drugs, to evaluate the presence or absence and malignancy (progression) of cancer cells from the ratio of normal cells to abnormal cells, and to evaluate the differentiation state of stem cells and use it for quality control.
[0004] One method for analyzing DNA methylation status is to use the bisulfite reaction. For example, a cytosine (C) in a CG sequence related to a certain disease is selected and used as the target site (measurement site). In Figure 13A, [1] to [4] are methylation sites, and [2] and [4] are set as target sites A and B (Figure 13A shows only one strand).
[0005] Next, the template DNA is treated with bisulfite (a salt of hydrogen sulfite). If the cytosine (C) in the CG sequence on the template DNA is methylated, it remains as cytosine (C) after this treatment (see methylation sites [3] and [4] in Figure 13A). On the other hand, if the cytosine (C) in the CG sequence on the template DNA is not methylated, it is deaminated and converted to uracil (U) (see methylation sites [1] and [2] in Figure 13A). Recently, instead of bisulfite treatment, a method of base conversion similar to the above reaction using enzymes such as the NEB Next Enzymatic Methyl-seq Kit manufactured by New England Biolabs has also been used.
[0006] Next, the bisulfite-treated DNA is amplified using PCR (polymerase chain reaction) for sequence analysis. The amplified DNA, i.e., PCR amplification product, is then sequenced using a capillary sequencer or a next-generation sequencer (NGS). When the bisulfite-treated DNA is amplified using PCR, cytosine (C) remains intact (see methylation sites [3] and [4] in Figure 13A), but uracil (U) is amplified and replaced with thymine (T) (see methylation sites [1] and [2] in Figure 13A). By utilizing the difference between cytosine (C) and thymine (T) in the sequence of this PCR amplification product, for example, it is possible to detect the methylation state of a specific target site in DNA (template DNA) before bisulfite treatment, i.e., whether the DNA of a specific target site selected from a single cell is methylated. More specifically, whether the cytosine (C) at a predetermined target site in the template DNA is methylated or unmethylated can be determined depending on whether the base at the predetermined target site in the PCR amplification product is cytosine (C) or thymine (T). Referring to Figure 13A, since the base at target site A in the PCR amplification product is thymine (T), it can be seen that the cytosine (C) at target site A in the template DNA was unmethylated. On the other hand, since the base at target site B in the PCR amplification product is cytosine (C), it can be seen that the cytosine (C) at target site B in the template DNA was methylated.
[0007] Furthermore, by utilizing the difference between cytosine (C) and thymine (T) occurring in the sequence of this PCR amplification product, the methylation state (frequency) of DNA at specific target sites derived from multiple cells in DNA (template DNA) before bisulfite treatment, i.e., whether or not DNA at specific target sites derived from multiple cells is methylated, can be detected, and the proportion of cells in which DNA at specific target sites is methylated can be determined based on the detection results. When there are multiple specific target sites, whether or not DNA at each specific target site is methylated can be detected, and the proportion of cells in which DNA is methylated can be determined for each specific target site based on the detection results. More specifically, the methylation state (frequency) of DNA at specific target sites derived from multiple cells can be determined based on whether the base at the specific target site occurring in the sequence of the PCR amplification product is cytosine (C) or thymine (T). The DNA methylation state (frequency) of a specific target site can be obtained by measuring the methylation degree = C / (C + T) from the number of cytosine (C) and thymine (T) occurring at each target site (measurement site).If there are multiple specific target sites, the proportion of cells with methylated DNA can be determined for each specific target site.
[0008] For example, as shown in FIG. 13B , when multiple cells (cells C1 to C3 in the figure) are used to evaluate the methylation status (frequency) of target sites (measurement sites) A and B derived from multiple cells, the number of cytosines (C) occurring at target site A is two and the number of thymines (T) is one, so the calculated methylation degree is 2 / (2 + 1) = 0.67. Therefore, the DNA methylation status (frequency) at target site A in FIG. 13B is a methylation degree of 0.67 derived from three cells, which allows us to grasp the proportion of cells with methylated DNA. Meanwhile, the number of cytosines (C) occurring at target site B is three and the number of thymines (T) is zero, so the calculated methylation degree is 3 / (3 + 0) = 1. Therefore, the DNA methylation status (frequency) at target site B in FIG. 13B is a methylation degree of 1 derived from three cells, which allows us to grasp the proportion of cells with methylated DNA. Similarly, the methylation state (frequency) of target site A shown in Figure 13A can be detected as a methylation degree of 0 derived from one cell, and the methylation state (frequency) of target site B can be detected as a methylation degree of 1 derived from one cell.
[0009] Amplification of DNA after bisulfite treatment may involve multiplex PCR, which can simultaneously amplify two or more target regions on DNA in the same reaction. To understand the DNA methylation state of a specific target site or the DNA methylation state (frequency) of specific target sites derived from multiple cells using multiplex PCR, primer pairs (forward and reverse primers) are required to amplify one or more target regions each containing two or more target sites, as shown in Figure 13C (only one strand is shown in Figure 13C). As illustrated in Figure 13A, a primer pair is required to amplify the target region (amplified region) containing target site A, and a primer pair is required to amplify the target region (amplified region) containing target site B.
[0010] When designing primers for bisulfite-treated DNA, in addition to the conditions considered in designing ordinary primers (i.e., designing primers for non-bisulfite-treated DNA), the following conditions must also be considered. First, unlike the base sequence, the presence or absence of DNA methylation cannot be known in advance. In other words, there are bases whose fate after bisulfite treatment is uncertain, i.e., whether they will become thymine (T) or cytosine (C). Therefore, when designing primers for analyzing the methylation status of DNA, it is necessary to minimize the presence of CG sequences at the site where the primer joins, or, if present, to limit the position of the CG sequence in the primer to minimize its influence, so that the amplification efficiency of the primer does not change depending on the methylation status around the target site.
[0011] Furthermore, when double-stranded DNA is subjected to bisulfite treatment, many cytosines (C) in the DNA are converted to thymine (T), and therefore, after bisulfite treatment, the DNA sequence of each strand increases in regions consisting of three bases other than cytosine (C). Therefore, it is also necessary to consider the need to design primers that can specifically bind to regions consisting of three bases. Furthermore, since many cytosines (C) in the DNA are converted to thymine (T), and the double-stranded DNA loses its complementarity, if it is necessary to amplify and analyze both strands of double-stranded DNA, it is necessary to design primer pairs (forward primer and reverse primer) that respectively amplify one or more amplification target regions containing the target site in each strand, i.e., two sets of primer pairs.
[0012] Therefore, designing a primer for bisulfite-treated DNA, which has these unique circumstances, requires different design conditions and is more difficult than designing a normal primer. Although there are many primer design software programs available, most of them are designed for designing normal primers, such as Primer-BLAST, and therefore cannot set conditions that take into account cytosine, the base of which is converted by bisulfite treatment. In other words, normal primer design software does not take into account the unique circumstances involved in designing a primer for bisulfite-treated DNA, as described above, and therefore such software cannot design a primer for bisulfite-treated DNA.
[0013] Furthermore, when multiplex PCR is used to amplify the bisulfite-treated DNA, multiple amplification target regions each containing a target site for methylation analysis are amplified simultaneously, so it is also necessary to consider designing primers that suppress the formation of primer-dimers.
[0014] Therefore, when the bisulfite reaction and multiplex PCR are used to measure the degree of methylation of DNA at a predetermined site, there is a problem in that the work of designing primers for the multiplex PCR used in the analysis (i.e., primers for bisulfite amplicon sequence analysis) is more complicated and time-consuming than the design of primers targeted at bisulfite-treated DNA.
[0015] As mentioned above, most primer design software is related to ordinary primer design software, and there is little software related to the design of primers for bisulfite-treated DNA. Furthermore, there is even less primer design software that can design primers for amplifying bisulfite-treated DNA by multiplex PCR (i.e., primers for bisulfite amplicon sequence analysis). One example of the few available software is that described in Patent Document 1, proposed by the present inventors.
[0016] International Publication No. 2022 / 113835
[0017] In bisulfite amplicon sequencing, 5 to 1,000 target sites are generally preset as measurement targets, and it is desirable to output primer sequences for as many target sites as possible. In other words, a high primer design success rate (number of target sites for which primers could be designed / number of all target sites [%]) is required.
[0018] Although the software described in Patent Document 1 can improve the success rate of primer design compared to conventional software for designing primers for bisulfite-treated DNA, further improvement in the success rate of primer design is desired. Furthermore, even if the success rate of primer design could be improved, there is a possibility that the probability of primer dimer formation will increase, resulting in a problem of reduced primer accuracy.
[0019] Furthermore, when designing primers, users do not necessarily want designs with a high success rate because they select designs with a high success rate depending on the conditions of individual DNA samples and the content of their research. However, there is a problem in that designing multiple primers according to the success rate requires time, effort, and cost.
[0020] The present invention has been made to solve these problems, and aims to provide a design method, production method, design device, design program, and recording medium for primers for bisulfite amplicon sequencing analysis (more specifically, primers for amplicon methylation sequencing analysis) that can further improve the success rate of primer design. It also aims to provide a design method, production method, design device, design program, and recording medium for primers for bisulfite amplicon sequencing analysis (more specifically, primers for amplicon methylation sequencing analysis) that enable users to easily design primers that meet their desired success rate of design.
[0021] [1] The method for designing primers for amplicon methylation sequence analysis according to the present invention is a method for designing primers for amplicon methylation sequence analysis used to simultaneously amplify multiple regions each containing two or more target sites for which the methylation degrees are to be measured, using a bisulfite reaction or an enzymatic reaction and multiplex PCR to measure the methylation degree of at least one genomic double-stranded DNA, and includes the following steps: a complementary strand generation step of generating a strand complementary to the template strand of the DNA; a partial sequence excision step of selecting one of the two or more target sites and excising one or more partial sequences of a predetermined length from the base sequence located on the 5'-end side of the selected target site from each of the strands; a candidate primer sequence selection step of selecting the one or more excised partial sequences as one or more candidate primer sequences; and a primer sequence determination step of adopting and determining a forward primer sequence and a reverse primer sequence that amplify the region containing the selected predetermined target site from the one or more candidate primer sequences. and a repeating step of repeating the partial sequence excision step, the candidate primer sequence selection step, and the primer sequence determination step until all of the two or more target sites have been selected in the partial sequence excision step, (I) when one or more primer sequences for different target sites have not yet been determined, the primer sequence determination step: [1] selects one or more candidate primer sequence pairs spanning the predetermined target site from the one or more candidate primer sequences, [2] selects one pair from the one or more candidate primer sequence pairs for the predetermined target site and calculates a local alignment score between the sequences of the selected candidate primer sequence pair, [3] adopts and determines the candidate primer sequence pair for which the calculated local alignment score is lower than the predetermined threshold as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site, (II) when one or more primer sequences for different target sites have already been determined, the primer sequence determination step: [1] selects one or more candidate primer sequence pairs spanning the predetermined target site from the one or more candidate primer sequences,[2] selecting one pair from one or more candidate primer sequence pairs for the predetermined target site, and calculating local alignment scores between each candidate sequence of the selected candidate primer sequence pair and each primer sequence for a different target site that has already been determined, as well as local alignment scores between the sequences of the selected candidate sequence pair; [3] detecting the maximum value from all the calculated local alignment scores, and adopting and determining the candidate primer sequence pair for which the local alignment score is calculated, the value of which is equal to or less than a predetermined threshold, as the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site; if the candidate primer sequence pair is not adopted as the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site in step [3] of (I) and (II) above, selecting a different pair from the one or more candidate primer sequence pairs selected in [1] of (I) and (II) above, and repeating steps [2] and [3] above until at least one candidate primer sequence pair is adopted; The local alignment score is calculated as follows: (1) "X" for each position of a pair of complementary bases, (2) "Y" for each position of a non-complementary pair, and (3) "Z" for each position when there is an insertion or deletion between the candidate primer sequences; where "X" is 1, "Y" is -4 to -2, and "Z" is -6 to -3; and the predetermined threshold is 1 to 4.
[0022] [2] The primer sequence determination step includes the steps of: (I) above, in the case where there are two or more target sites and one or more primer sequences for different target sites have not yet been determined; in the step [2], selecting all pairs from one or more candidate primer sequence pairs for the predetermined target site and calculating a local alignment score between the sequences of the selected candidate primer sequence pair for each pair; and in the step [3], selecting one or more candidate primer sequence pairs for which the calculated local alignment score is equal to or less than the predetermined threshold; and further detecting, from all the selected pairs, the candidate primer sequence pair having the smallest maximum value of the local alignment score, and adopting and determining them as the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site; and (II) above, in the case where one or more primer sequences for different target sites have already been determined; the method for designing primers for amplicon methylation sequence analysis according to [1] above, wherein in step [2] all pairs are selected from one or more candidate primer sequence pairs for the predetermined target site, and for each pair, a local alignment score is calculated between each candidate sequence of the selected candidate primer sequence pair and each primer sequence for a different target site that has already been determined, as well as a local alignment score between the sequences of the selected candidate sequence pair; and in step [3], for each pair, the maximum value is detected from all the calculated local alignment scores, and the candidate primer sequence pair for which a local alignment score has been calculated that is equal to or less than a predetermined threshold is selected, and further, from all the selected pairs, the candidate primer sequence pair having the smallest maximum value of the local alignment score is detected, and these are adopted and determined as the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site.
[0023] [3] The method further comprises: a base sequence data acquisition step of acquiring base sequence data of the genomic double-stranded DNA; a target site information acquisition step of acquiring the two or more target sites and their positional information; and a base conversion step of converting "C" that can be methylated in the genomic double-stranded DNA to "Y" and converting other "C"s to "T" in the base sequence data; the complementary strand generation step of generating a complementary strand to each template strand of the genomic double-stranded DNA after the base conversion; the partial sequence excision step of selecting one from the two or more target sites and excising one or more partial sequences of a predetermined length from each strand from the base sequence located on the 5'-terminal side of the "Y" into which the selected target site is converted or the "R" complementary thereto, based on the positional information of the selected target site; the candidate primer sequence selection step of selecting, from the one or more partial sequences excised from each strand, those that satisfy predetermined selection conditions as candidate primer sequences; the "C" that can be methylated is "C" in a CG sequence; The method for designing a primer for amplicon methylation sequencing according to [1] or [2], wherein the predetermined selection conditions include: (1) the Tm value is within a predetermined range; (2) the number of YG sequences or CR sequences contained in the partial sequence is a predetermined number or less; and (3) the upper limit of the number of junctions with sequences outside the relevant region on the genomic double-stranded DNA after the base conversion is a predetermined number or more, where "C," "G," "Y," and "R" are base symbols defined by IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine.
[0024] [4] The primer design method according to [3] above, wherein the "C" that can be methylated further includes a "C" in a CHG sequence, and the predetermined selection conditions further include (4) the partial sequence containing no more than a predetermined number of YHG sequences or CDR sequences. [Here, "C," "G," "Y," "H," "R," and "D" are base symbols defined by IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine.] [5] The primer design method according to [3] or [4] above, wherein the "C" that can be methylated further includes a "C" in a CHH sequence, and the predetermined selection conditions further include (5) the partial sequence containing no more than a predetermined number of YHH sequences or DDR sequences. [Here, "Y", "H", "R", and "D" are base symbols defined by IUPAC, where "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine.]
[0025] [6] The step of selecting candidate primer sequences comprises: [7] The primer design method according to any one of [3] to [5] above, wherein the genomic double-stranded DNA after base conversion is used as a first template strand and a second template strand, the complementary strand of the first template strand is used as the first complementary strand, and the complementary strand of the second template strand is used as the second complementary strand, and one or more partial sequences excised from the first template strand that satisfy a predetermined selection condition are selected as candidate forward primer sequences for the first template strand, one or more partial sequences excised from the first complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the first template strand, one or more partial sequences excised from the second template strand that satisfy the predetermined selection condition are selected as candidate forward primer sequences for the second template strand, and one or more partial sequences excised from the second complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the second template strand. The primer sequence determination step calculates the lengths of PCR amplification products predicted to be amplified by PCR for all combinations of the candidate forward primer sequences of the one or more first template strands selected in the candidate primer sequence selection step and the candidate reverse primer sequences of the one or more first template strands selected in the candidate primer sequence selection step, and employs combinations of candidate primer sequences whose calculated PCR amplification product lengths are within a predetermined range as forward primer sequences and reverse primer sequences of the first template strands that amplify the region including the target site selected in the partial sequence excision step. [6] The primer design method according to any one of [3] to [6] above, wherein the method is a step of calculating the lengths of PCR amplification products predicted to be amplified by PCR for all combinations of candidate forward primer sequences of the selected second template strand and candidate reverse primer sequences of the selected second template strand, and determining combinations of candidate primer sequences for which the calculated PCR amplification product lengths are within the predetermined range as forward primer sequences and reverse primer sequences of the second template strand that amplify the region including the target site selected in the partial sequence excision step.
[0026] [8] The primer design method according to any one of [1] to [7] above, wherein a correspondence relationship between at least the number of target sites, the predetermined threshold, and the primer design success rate is measured in advance using the primer design method for amplicon methylation sequence analysis according to any one of [1] to [7] above, and the correspondence relationship is stored in a storage unit; and when a user sets at least the primer design success rate and the number of target sites desired by the user via an input unit and issues an instruction to execute primer design, the method reads out, from the correspondence relationships stored in the storage unit, the predetermined threshold corresponding to the primer design success rate and the number of target sites that are equal to or greater than the set values for the primer design success rate and the number of target sites and have a small difference therebetween, and selects and determines, from the one or more candidate primer sequences, a primer sequence that amplifies a region including the predetermined target site based on the read predetermined threshold.
[0027] [9] The method for producing a primer for amplicon methylation sequence analysis of the present invention comprises a primer design step according to any one of [1] to [8] above, and a synthesis step of synthesizing a primer based on the primer sequence designed in the primer design step, wherein the primer design step is carried out by the above-mentioned method for designing a primer for amplicon methylation sequence analysis.
[0028]
[10] The device for designing primers for amplicon methylation sequence analysis of the present invention is a device for designing primers for amplicon methylation sequence analysis used to simultaneously amplify multiple regions each containing two or more target sites for which the methylation degrees are to be measured, using a bisulfite reaction or an enzymatic reaction and multiplex PCR to measure the methylation degree of at least one double-stranded DNA, and includes: a complementary strand generator that generates a strand complementary to a template strand of the DNA; a partial sequence excision unit that selects one of the two or more target sites and excises one or more partial sequences of a predetermined length from the base sequence located on the 5'-end side of the selected target site from each of the strands; a candidate primer sequence selection unit that selects the one or more excised partial sequences as one or more candidate primer sequences; and a primer sequence determination unit that adopts and determines a forward primer sequence and a reverse primer sequence that amplify a region containing the selected predetermined target site from the one or more candidate primer sequences. and a control unit that controls the partial sequence excision unit to repeat the processes of the partial sequence excision unit, the candidate primer sequence selection unit, and the primer sequence determination unit until all of the two or more target sites are selected in the partial sequence excision unit, (I) when one or more primer sequences for different target sites have not yet been determined, the primer sequence determination step: [1] selects one or more candidate primer sequence pairs spanning the predetermined target site from the one or more candidate primer sequences, [2] selects one pair from the one or more candidate primer sequence pairs for the predetermined target site and calculates a local alignment score between the sequences of the selected candidate primer sequence pair, and [3] adopts and determines the candidate primer sequence pair for which the calculated local alignment score is equal to or less than the predetermined threshold as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site, (II) when one or more primer sequences for different target sites have already been determined, the primer sequence determination step: [1] selects one or more candidate primer sequence pairs spanning the predetermined target site from the one or more candidate primer sequences,[2] selecting one pair from one or more candidate primer sequence pairs for the predetermined target site, and for each pair, calculating a local alignment score between each candidate sequence of the selected candidate primer sequence pair and each primer sequence for a different target site that has already been determined, and a local alignment score between the sequences of the selected candidate sequence pair; [3] detecting the maximum value from all the calculated local alignment scores, and adopting and determining the candidate primer sequence pair for which the local alignment score is calculated, the value of which is equal to or less than a predetermined threshold, as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site; if the candidate primer sequence pair is not adopted as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site in step [3] of (I) and (II) above, selecting a different pair from the one or more candidate primer sequence pairs selected in [1] of (I) and (II) above, and repeating steps [2] and [3] above until at least one candidate primer sequence pair is adopted; The local alignment score is calculated as follows: (1) "X" for each position of a pair of complementary bases, (2) "Y" for each position of a pair of non-complementary bases, and (3) "Z" for each position when there is an insertion or deletion between the candidate primer sequences; where "X" is 1, "Y" is -4 to -2, and "Z" is -6 to -3; and the predetermined threshold is 1 to 4.
[0029]
[11] The primer sequence determination unit: (I) when one or more primer sequences for different target sites have not yet been determined; in step [2], selects all pairs from one or more candidate primer sequence pairs for the predetermined target site, and calculates a local alignment score between the sequences of the selected candidate primer sequence pair for each pair; in step [3], selects one or more candidate primer sequence pairs for which the calculated local alignment score is equal to or less than a predetermined threshold, and further detects, from all the selected pairs, the candidate primer sequence pair having the smallest maximum value of the local alignment score, and adopts and determines them as the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site; and (II) when one or more primer sequences for different target sites have already been determined; the step [2] selects all pairs from one or more candidate primer sequence pairs for the predetermined target site, and for each pair, calculates a local alignment score between each candidate sequence of the selected candidate primer sequence pair and each primer sequence for a different target site that has already been determined, and a local alignment score between the sequences of the selected candidate sequence pair; and the step [3] detects the maximum value from all the calculated local alignment scores for each pair, and selects the candidate primer sequence pair for which a local alignment score has been calculated that is equal to or less than a predetermined threshold value, and further detects the candidate primer sequence pair with the smallest maximum value of the local alignment score from all the selected pairs, and adopts and determines them as the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site.
[0030]
[12] The DNA sequencing system further comprises: a base sequence data acquisition unit that acquires base sequence data of the genomic double-stranded DNA; a target site information acquisition unit that acquires the two or more target sites and their positional information; and a base conversion unit that converts "C" that can be methylated in the genomic double-stranded DNA to "Y" and converts other "C" to "T" in the base sequence data; the complementary strand generation unit generates a complementary strand for each template strand of the genomic double-stranded DNA after the base conversion; the partial sequence excision unit selects one from the two or more target sites and, based on positional information of the selected target site, excises one or more partial sequences of a predetermined length from the base sequence located on the 5'-terminal side of the "Y" into which the selected target site is converted or the "R" complementary thereto; the candidate primer sequence selection unit selects, from the one or more partial sequences excised from each strand, one that satisfies predetermined selection conditions as a candidate primer sequence; the "C" that can be methylated is "C" in a CG sequence; The device for designing primers for amplicon methylation sequencing analysis according to
[10] or
[11] above, wherein the predetermined selection conditions include: (1) Tm being within a predetermined range; (2) the number of YG sequences or CR sequences contained in the partial sequence being a predetermined number or less; and (3) the upper limit of the number of junctions with sequences outside the relevant region on the genomic double-stranded DNA after the base conversion is a predetermined number or more, where "C," "G," "Y," and "R" are base symbols defined by IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine.
[0031]
[13] The primer design device according to
[12] above, wherein the "C" that can be methylated further includes "C" in a CHG sequence, and the predetermined selection conditions further include (4) that the partial sequence contains no more than a predetermined number of YHG sequences or CDR sequences [wherein "C," "G," "Y," "H," "R," and "D" are base symbols defined by IUPAC, "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine].
[14] The primer design device according to
[12] or
[13] above, wherein the "C" that can be methylated further includes "C" in a CHH sequence, and the predetermined selection conditions further include (5) that the number of YHH sequences or DDR sequences contained in the partial sequence is not more than a predetermined number [wherein "Y", "H", "R", and "D" are base symbols defined by IUPAC, "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine].
[0032]
[15] The candidate primer sequence selection unit uses the genomic double-stranded DNA after the base conversion as a first template strand and a second template strand, a complementary strand of the first template strand as a first complementary strand, and a complementary strand of the second template strand as a second complementary strand,
[16] The primer design device according to any one of
[12] to
[14] above, wherein one or more partial sequences excised from the first template strand that satisfy a predetermined selection condition are selected as candidate forward primer sequences for the first template strand, one or more partial sequences excised from the first complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the first template strand, one or more partial sequences excised from the second template strand that satisfy the predetermined selection condition are selected as candidate forward primer sequences for the second template strand, and one or more partial sequences excised from the second complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the second template strand. The primer sequence determination unit calculates, in the candidate primer sequence selection unit, the lengths of PCR amplification products expected to be amplified by PCR for all combinations of the candidate forward primer sequences of the one or more selected first template strands and the candidate reverse primer sequences of the one or more selected first template strands, and adopts combinations of candidate primer sequences for which the calculated PCR amplification product lengths are within a predetermined range as forward primer sequences and reverse primer sequences of the first template strand that amplify the region containing the target site selected in the partial sequence excision unit; calculates the lengths of PCR amplification products expected to be amplified by PCR for all combinations of the candidate forward primer sequences of the selected second template strand and the candidate reverse primer sequences of the selected second template strand, and adopts combinations of candidate primer sequences for which the calculated PCR amplification product lengths are within the predetermined range as forward primer sequences and reverse primer sequences of the second template strand that amplify the region containing the target site selected in the partial sequence excision unit. The primer design device described in
[15] above.
[0033]
[17] The primer design device for amplicon methylation sequence analysis according to
[10] above, further comprising: a storage unit that measures and stores in advance, using the primer design device according to
[10] above, at least a correspondence relationship between the number of target sites, the predetermined threshold, and the primer design success rate; and an input unit for a user to input instructions, wherein when a user sets at least a primer design success rate and the number of target sites desired by the user via the input unit and issues an instruction to execute primer design, the primer sequence determination unit reads out, from the correspondence relationships stored in the storage unit, the predetermined threshold corresponding to the primer design success rate and the number of target sites that are equal to or greater than the set values for the primer design success rate and the number of target sites and have a small difference therebetween, and selects and determines, from the one or more candidate primer sequences, a primer sequence that amplifies a region including the predetermined target site, based on the read predetermined threshold.
[0034] The primer design device for amplicon methylation sequence analysis according to any one of
[12] to
[17] above further includes a communication interface, and can be connected to a server via a communication network external to the device via the communication interface, and can execute at least one of the group consisting of the base sequence data acquisition unit, the target site information acquisition unit, the base conversion unit, the complementary strand generation unit, the partial sequence excision unit, the primer candidate sequence selection unit, and the primer sequence determination unit using a program in the server.
[0035] The program for designing primers for amplicon methylation sequence analysis according to any one of [1] to [8] above of the present invention is capable of executing the above-mentioned primer design method on a computer. The computer-readable recording medium according to
[19] above of the present invention is a recording medium on which the program for designing primers for amplicon methylation sequence analysis described above is recorded.
[0036] According to the present invention, not only can the design success rate of primers for bisulfite amplicon sequencing analysis (more specifically, primers for amplicon methylation sequencing analysis) be further improved compared to conventional methods, but the probability of primer dimer formation can also be reduced. Furthermore, primers based on the designs of the present invention can be obtained. As a result, many target sites can be amplified and measured. Furthermore, according to the present invention, primers for bisulfite amplicon sequencing analysis (more specifically, primers for amplicon methylation sequencing analysis) can be easily and quickly designed according to the user's desired design success rate. Furthermore, primers based on such designs can be obtained.
[0037] FIG. 1 is a block diagram conceptually showing an example of the configuration of a primer design device according to embodiment 1 of the present invention. FIG. 2 is a flowchart showing an example of a primer design method according to embodiment 1 carried out by the primer design device shown in FIG. 1. FIG. 3 is a schematic diagram illustrating a base sequence data acquisition step in the primer design method shown in FIG. 2. FIG. 4 is a schematic diagram illustrating a base conversion step in the primer design method shown in FIG. 2. FIG. 5 is a schematic diagram illustrating a complementary strand generation step in the primer design method shown in FIG. 2. FIG. 6 is a flowchart illustrating an example of the operation of a partial sequence excision unit 28, a candidate primer sequence selection unit 30, and a primer sequence determination unit 32. FIG. 5A is a diagram for explaining condition (3) that "the upper limit of the number of junctions with a sequence outside the relevant region on the genomic double-stranded DNA after base conversion is equal to or less than a predetermined number greater than or equal to 1." FIG. 5B is a diagram for explaining condition (3) that "the upper limit of the number of junctions with a sequence outside the relevant region on the genomic double-stranded DNA after base conversion is equal to or less than a predetermined number greater than or equal to 1." FIG. 6A is a diagram for explaining combinations of sequence comparisons for calculating a local alignment score. FIG. 6B is a diagram for explaining combinations of sequence comparisons for calculating a local alignment score. FIG. 7A is a diagram for explaining combinations of sequence comparisons for calculating a local alignment score. FIG. 7B is a diagram for explaining combinations of sequence comparisons for calculating a local alignment score. FIG. 8 is a diagram for explaining combinations of sequence comparisons for calculating a local alignment score. FIG. 9 is a diagram illustrating a method for calculating an alignment score and a method for determining a result based on a threshold value. FIG. 9 is a diagram illustrating the correspondence between the number of target sites, a threshold value, and a primer design success rate, which are stored in the storage unit of a primer design device according to a second modification of the first embodiment of the present invention. FIG. 10 is a block diagram conceptually illustrating an example of the configuration of a primer design device according to a second embodiment of the present invention. FIG. 11 is a block diagram conceptually illustrating an example of a connection between the primer design device according to the second embodiment of the present invention and an external server. FIG. 13A is a diagram illustrating the primer design success rates for Examples 1 to 4 and Comparative Examples 2 to 4. FIG. 13B is a schematic diagram illustrating an example of a method for analyzing the methylation status of DNA using a bisulfite reaction.Fig. 13B is a schematic diagram illustrating an example of a method for analyzing the methylation status (frequency) of DNA using the bisulfite reaction. Fig. 13C is a diagram illustrating a target site (measurement site) and a region to be amplified.
[0038] The following provides a detailed description of a method for designing, a manufacturing method, a design device, a design program, and a recording medium for a primer for bisulfite amplicon sequencing (a primer for amplicon methylation sequence analysis) of the present invention, based on the official embodiments shown in the accompanying drawings.
[0039] (Explanation of Terms) As used herein, "primers for bisulfite amplicon sequence analysis" refers to analysis primers for simultaneously amplifying multiple amplification target regions, each containing multiple target sites, in bisulfite-treated DNA by multiplex PCR. "Primers for amplicon methylation sequence analysis" refers to analysis primers for simultaneously amplifying multiple amplification target regions, each containing multiple target sites, in bisulfite-treated or enzyme-treated DNA by multiplex PCR. "Amplified target region" refers to a region amplified by a primer pair. "Methylation site" refers to a site that can be methylated. "Target site" refers to a "methylation site" and refers to a site (measurement site) where the degree of methylation is measured. "Candidate primer sequence" refers to either a forward candidate primer sequence or a reverse candidate primer sequence, unless otherwise specified. "Candidate primer sequence pair" refers to a combination of a forward candidate primer sequence and a reverse candidate primer sequence. "Primer sequence" refers to either a forward primer sequence or a reverse primer sequence, unless otherwise specified. A "primer sequence pair" refers to a combination of a forward primer sequence and a reverse primer sequence. Base sequences such as "GC sequence" and "YG sequence" refer to sequences read from the 5' end. Ranges expressed using "~" include both sides of "~". For example, a range expressed as "A to B" includes A and B.
[0040] [Embodiment 1] Fig. 1 is a block diagram conceptually showing an example of a primer design device according to embodiment 1 of the present invention. Fig. 2 conceptually shows a flowchart of an example of a primer design method carried out by the primer design device shown in Fig. 1. Figs. 3A to 3D are schematic diagrams illustrating each step of the primer design method.
[0041] 1, the primer design device 10 includes an input unit 12, a storage unit 14, an output unit 16, and a primer design processing unit 18. The input unit 12, the storage unit 14, the output unit 16, and the primer design processing unit 18 are connected to each other.
[0042] The input unit 12 acquires information entered by the user, as well as various setting instructions, selection instructions, input instructions, creation instructions, etc., and is composed of input devices such as a keyboard and a mouse. The storage unit 14 stores the operating program of the primer design device and can also temporarily store information and data necessary for executing the primer design process. The storage unit 14 can be, for example, a hard disk drive (HDD), a solid state drive (SSD), a flexible disk (FD), a magneto-optical disk (MO disk), a magnetic tape (MT), a random access memory (RAM), a compact disk (CD), a digital versatile disk (DVD), a secure digital card (SD card), a universal serial bus memory (USB memory), or other storage media. The output unit 16 outputs the DNA base sequence information, instructions, design conditions, and primer sequence information designed by the primer design processing unit 18 input from the input unit 12, and is composed of, for example, a display unit such as a liquid crystal display (LCD), an organic light emitting diode (OLED), a flat panel display, a solid-state display, or a cathode ray tube (CRT), as well as various types of printers.
[0043] The primer design processing unit 18 performs a series of processes for primer design. The primer design processing unit 18 includes a base sequence data acquisition unit 20, a target site information acquisition unit 22, a base conversion unit 24, a complementary strand generation unit 26, a partial sequence excision unit 28, a candidate primer sequence selection unit 30, a primer sequence determination unit 32, and a control unit 34. The primer design processing unit 18 can be configured by a processor including a central processing unit (CPU), a computer, etc. As shown in FIG. 2 , the primer design method includes a base sequence data acquisition step S10, a target site information acquisition step S12, a base conversion step S14, a complementary strand generation step S16, a partial sequence extraction step S18, a candidate primer sequence selection step S20, a primer sequence determination step S22, and a determination step S24, in which it is determined whether all target sites have been selected, and the partial sequence extraction step S18, the candidate primer sequence selection step S20, and the primer sequence determination step S22 are repeated until all target sites are detected.
[0044] (Base Sequence Data Acquisition Unit) The base sequence data acquisition unit 20 shown in FIG. 1 is a unit that performs the base sequence data acquisition step S10 shown in FIG. 2 and acquires double-stranded DNA sequence (reference sequence) data of the genome of the biological species for which primer design is being performed via the input unit 12. If reference sequence data is pre-stored in the storage unit 14, the data may be acquired from the storage unit 14. Here, the acquired genome double-stranded DNA sequence data is preferably data of the entire genome sequence of the biological species for which primer design is being performed. For the purpose of explaining the primer design method of this embodiment, the double-stranded DNA of the double-stranded DNA sequence data acquired in this step will be referred to as template DNA, and will be referred to as strand A and strand B, respectively (see FIG. 3A). The base sequence data acquisition unit 20 is configured by a computer and performs the function of acquiring the above-mentioned genome double-stranded DNA sequence data.
[0045] (Target Site Information Acquisition Unit) The target site information acquisition unit 22 shown in FIG. 1 is a unit that performs the target site information acquisition step S12 shown in FIG. 2 and can acquire, via the input unit 12, one or more target sites contained in the genomic double-stranded DNA acquired by the base sequence data acquisition unit 20, and their location information. If the target sites and their location information are stored in the storage unit 14 in advance, they may be acquired from the storage unit 14. Here, a "target site" refers to a site related to a predetermined biological phenomenon, a cytosine (C) in a CG sequence that can be methylated, and a site for measuring the degree of methylation. The number of target sites selected is not particularly limited as long as it is two or more, but it is preferable to select 5 to 1,000 sites from the perspective of significantly achieving the desired effects of the present invention. The position of each target site can be indicated by chromosome and genome coordinates, etc. The target site information acquisition unit 22 is configured by a computer and performs the function of acquiring two or more target sites contained in the above-mentioned genomic double-stranded DNA, and their location information.
[0046] (Base Conversion Unit) The base conversion unit 24 is a unit that performs the base conversion step S14 shown in FIG. 2. As shown in FIGS. 3A and 3B, the base conversion unit 24 converts cytosine (C) in the CG sequence on the template DNA acquired from the base sequence data acquisition unit 20 to "Y" (see the bases indicated by the arrows in FIGS. 3A and 3B), and converts cytosine (C) in other sequences to thymine (T). Because cytosine (C) in the CG sequence of DNA may be methylated or unmethylated, it is converted to "Y," which includes both the possibility of being converted to thymine (T) and the possibility of remaining as cytosine (C). This conversion process simulates the generation of DNA amplified using PCR after bisulfite treatment on a computer. The base conversion unit 24 is configured by a computer and performs the function of converting cytosine (C) in the CG sequence on the template DNA to "Y" and converting cytosine (C) in other sequences to thymine (T).
[0047] As mentioned above, bisulfite treatment causes double-stranded DNA to lose its complementarity. This is because bisulfite treatment converts the cytosine (C) in the complementary CG base pair to thymine (T), resulting in the loss of base pair complementarity (see the bold bases in Figures 3A and 3B). The target amplification region on the DNA after bisulfite treatment, which has lost its complementarity, cannot be amplified in both strands in the same way with a single pair of primers. Therefore, when analyzing the methylation status of double-stranded DNA, it is necessary to prepare primer pairs (forward and reverse primers) that amplify the target amplification region containing each target site in each strand for each target site. In other words, it is necessary to design a primer pair for the target amplification region containing the target site in strand A after base conversion in Figure 3B, and a primer pair for the target amplification region containing the target site in strand B after base sequencing. However, as will be explained in Modification Example 5 below, when it is desired to analyze only the A chain, or only the B chain, or when it is sufficient to analyze either the A chain or the B chain, it is not necessarily necessary to design two sets of primer pairs.
[0048] (Complementary Chain Generation Unit) The complementary chain generation unit 26 is the part that performs the complementary chain generation step S16 shown in FIG. 2 and generates complementary chains for each of the double-stranded DNA strands after the base conversion process. For the purpose of explaining the primer design method of this embodiment, the post-base conversion A chain and the post-base conversion B chain are referred to as the first template strand (A+ strand) and the second template strand (B+ strand), respectively, and the complementary chain of the first template strand is referred to as the first complementary chain (A- strand), and the complementary chain of the second template strand is referred to as the second complementary chain (B- strand) (see FIG. 3C). As shown in FIG. 3C, a complementary chain A- is generated by generating a sequence complementary to the base sequence of the A+ strand, and a complementary chain B- is generated by generating a sequence complementary to the base sequence of the B+ strand. The base complementary to "Y" is designated "R," which may be either adenine (A) or guanine (G). The complementary strand generating unit 26 is configured by a computer and performs the function of generating the above-mentioned complementary strand for each of the DNA double strands after the base conversion process.
[0049] As a result, the first template strand (A+ strand) is composed of three bases, thymine (T), adenine (A), and guanine (G), excluding "Y" (i.e., methylation site), and the first complementary strand (A- strand) is composed of three bases, thymine (T), adenine (A), and cytosine (C), excluding "R" (methylation site), but the first template strand (A+ strand) and the first complementary strand (A- strand) can be complementary. Similarly, the second template strand (B+ strand) is composed of three bases, thymine (T), adenine (A), and guanine (G), excluding "Y" (methylation site), and the second complementary strand (B- strand) is composed of three bases, thymine (T), adenine (A), and cytosine (C), excluding "R" (methylation site), but the second template strand (B+ strand) and the second complementary strand (B- strand) can be complementary.
[0050] (Partial Sequence Extraction Unit) The partial sequence extraction unit 28 is a unit that performs the partial sequence extraction step S18 shown in Figure 2. As shown in the flowchart of Figure 4, the partial sequence extraction unit 28 selects one target site from two or more target sites acquired by the target site information acquisition unit 22 (step S280), detects "Y" or its complementary "R" (i.e., a base located in the target site and at a methylation site) of the selected target site from the DNA sequence of each strand based on the position information of the selected target site, and extracts as many partial sequences as possible from the base sequences located on the 5'-terminal side of the detected "Y" and "R" ((1) to (4) in Figure 3D) from partial sequences of a predetermined length (step S282), thereby obtaining one or more partial sequences. Note that Figure 4 is a flowchart showing an example of the operations of the partial sequence extraction unit 28, the candidate primer sequence selection unit 30, and the primer sequence determination unit 32. The partial sequence extraction unit 28 is configured by a computer, and performs the function of extracting as many partial sequences of a predetermined length from the selected target site "Y" or its complementary "R" from the DNA sequence of each strand based on the position information of the selected target site described above, thereby obtaining one or more partial sequences.
[0051] Here, the length of one or more partial sequences to be excised is not particularly limited, but from the viewpoint of processing efficiency and achieving the desired effects of the present invention, it is preferable that the length be the difference between the maximum length of the PCR amplification product desired by the user and the minimum length of the primer, minus the length of the target site (one base). The length of the PCR amplification product is not particularly limited as long as it is within a known range, i.e., 70 to several kilobp. It is preferable to take into consideration the PCR success rate and the sequence decoding ability of the DNA sequencer. The length of the primer is not particularly limited as long as it is within a known range, i.e., 15 to 45 bases. It is preferable to take into consideration the specificity of the primer and the likelihood of primer-dimer formation.
[0052] For example, if the maximum length of the PCR product set by the user is 300 bases and the minimum length of the primer is 20 bases, the predetermined length to be excised is calculated as x = 300 - 20 - 1 (length of the target site) = 279, and first, 279 bases on the 5'-end side of each target site are excised. As shown in Figure 3D, 279 bases on the 5'-end side of the target site of each strand (i.e., "Y" on the A+ strand, "R" on the A- strand, "Y" on the B+ strand, and "R" on the B- strand) ((1) to (4) in Figure 3D)) are excised from each strand. Next, one or more partial sequences can be obtained by excising as many partial sequences as possible from the 279 bases within the length of the primer (20 bases or more and the predetermined length or less).
[0053] The values or ranges of values for the length of the PCR amplification product and the length of the primers are set by the user via the input unit 12. If these conditions are stored in advance in the storage unit 14, they can be obtained from the storage unit 14 and set.
[0054] (Candidate Primer Sequence Selection Section) The candidate primer sequence selection section 30 is a section that carries out the candidate primer sequence selection step S20 shown in Figure 2, and selects as candidate primer sequences those that satisfy all of the predetermined selection conditions (1) to (3) from one or more partial sequences of each strand excised by the partial sequence excision section 28. Specifically, one or more partial sequences excised from the first template strand (A+ strand) (i.e., one or more partial sequences excised from (1) in Figure 3D) that satisfy the predetermined selection conditions are selected as candidate forward primer sequences for the first template strand (A+ strand), and one or more partial sequences excised from the first complementary strand (A- strand) (i.e., one or more partial sequences excised from (2) in Figure 3D) that satisfy the predetermined selection conditions are selected as candidate reverse primer sequences for the first template strand (A+ strand). If one or more partial sequences excised from the second template strand (B+ strand) (i.e., one or more partial sequences excised from (3) in FIG. 3D) satisfy a predetermined selection condition, they are selected as candidate forward primer sequences for the second template strand (B+ strand), and if one or more partial sequences excised from the second complementary strand (B- strand) (i.e., one or more partial sequences excised from (4) in FIG. 3D) satisfy a predetermined selection condition, they are selected as candidate reverse primer sequences for the second template strand (B+ strand). The candidate primer sequence selection unit 30 is configured by a computer and performs the function of selecting candidate primer sequences from one or more partial sequences of each strand described above that satisfy all of the predetermined selection conditions (1) to (3).
[0055] The "predetermined selection conditions" for candidate primer sequences are the following items (1) to (3). The numerical values and numerical ranges of the predetermined selection conditions can be set in advance by the user via the input unit 12. (1) The Tm value is within a predetermined range, (2) The number of YG sequences or CR sequences contained in the partial sequence is a predetermined number or less, and (3) The upper limit of the number of junctions with base sequences outside the relevant region on the template strand DNA (genomic double-stranded DNA) after base conversion is 1 or more and less than a predetermined number.
[0056] The range of the "Tm value" according to the above condition (1) is not particularly limited as long as it is within a known numerical range, i.e., 45 to 70°C. It is preferable to take into consideration the thermal cycle conditions of PCR, the ease of PCR amplification (the temperature range in which amplification easily proceeds depending on the PCR enzyme used), and the specificity of PCR amplification. The Tm value can be calculated, for example, by the nearest neighbor base pair method.
[0057] The number of "YG sequences or CR sequences contained in the partial sequence" according to the above condition (2) is not particularly limited, but from the viewpoint of significantly obtaining the desired effect of the present invention, it is preferably 2 or less, more preferably 1 or less, and particularly preferably 0. By satisfying this condition, it is possible to reduce the influence of the joining of the primer with cytosine (C) of the CG sequence at the primer joining site.
[0058] In the above condition (3), the "sequence outside the relevant region on the template strand DNA (genomic double-stranded DNA) after base conversion" refers to a base sequence excluding the sequence at the position on the template strand DNA after base conversion that corresponds to the position of the partial sequence, or a base sequence complementary to the sequence excluding the partial sequence (template strand DNA sequence after base conversion). The "upper limit of the number of junctions with sequences outside the relevant region on the template strand DNA after base conversion" is not particularly limited, but from the viewpoint of significantly achieving the desired effect of the present invention, it is preferably 5 or less, and particularly preferably 2 or less. By satisfying this condition, the effect of primers junctioning with sequences outside the relevant region on the DNA after bisulfite treatment can be reduced.
[0059] As shown in FIG. 5A, when the number of heating cycles in PCR is n, and a primer pair (forward primer and reverse primer) is ligated to DNA, 2 nHowever, as shown in Figure 5B, when either one of the forward primer or the reverse primer is ligated to DNA, PCR amplification products are generated on the order of 2n (Figure 5B shows the case where only the forward primer is ligated). Therefore, when PCR is performed with a typical number of heating cycles (n is approximately 20 to 40), if the primer pair is ligated to DNA at a DNA sequence outside the target region for amplification, a large amount of non-specific products is generated, which is problematic. However, when either one of the forward primer or the reverse primer is ligated to a DNA sequence outside the relevant region, non-specific products are not generated in significant quantities, which is not particularly problematic. For this reason, the problem of non-specific products being generated when either one of the forward primer or the reverse primer is ligated to a DNA sequence outside the relevant region has not been particularly considered in the past. Note that (1) in Figure 5A shows the DNA sequence of the target region for amplification, and (2) shows the DNA sequence outside the target region for amplification. Furthermore, (3) in Figure 5B shows the DNA sequence of the relevant region of the partial sequence, and (4) shows the DNA sequence outside the relevant region. In this way, by adding the above condition (3) to the conventional conditions used in primer design, which allows each primer to bind to DNA outside the target region within a specified range, the success rate of primer design can be increased.
[0060] Here, the process of selecting, from one or more partial sequences excised from each strand, those that satisfy predetermined selection conditions as candidate primer sequences will be described with reference to the flowchart of FIG.
[0061] The candidate primer sequence selection unit 30 first obtains one of one or more partial sequences excised from the first template strand (A+ strand) (step S300) and determines whether the Tm value of that partial sequence is within a predetermined range (step S302). If the Tm value is not within the predetermined range, another partial sequence is obtained (step S300). If the Tm value is within the predetermined range, it determines whether the number of YG or CR sequences contained in the partial sequence is a predetermined number or less (step S304). If the number of YG or CR sequences contained in the partial sequence is not a predetermined number or less, another partial sequence is obtained (step S300). If the number of YG or CR sequences contained in the partial sequence is not a predetermined number or less, it determines whether the upper limit of the number of junctions with base sequences outside the relevant region on the template strand DNA after base conversion is not more than a predetermined number (step S306). If the upper limit of the number of junctions between the base sequence outside the relevant region on the template strand DNA after base conversion and the partial sequence is not equal to or less than a "predetermined number of 1 or more," another partial sequence is obtained (step S300). If the upper limit of the number of junctions between the base sequence outside the relevant region on the template strand DNA after base conversion and the partial sequence is equal to or less than a predetermined number of 1 or more, the partial sequence is selected as a candidate primer sequence (step S308). It is then determined whether or not all partial sequences excised from the first template strand (A+ strand) have been identified (step S310). If identification of all partial sequences excised from the first template strand (A+ strand) has not been identified, another partial sequence is obtained (step S300). If identification of all partial sequences has been identified, the one or more selected candidate primer sequences are determined as candidate forward primer sequences for the first template strand (A+ strand) (step S312).
[0062] Similar determinations are made for one or more partial sequences excised from the first complementary strand (A-strand), one or more partial sequences excised from the second template strand (B+ strand), and one or more partial sequences excised from the second complementary strand (B-strand) (steps S300 to S310), and candidate reverse primer sequences for the first template strand (A+ strand), candidate forward primer sequences for the second template strand (B+ strand), and candidate reverse primer sequences for the second template strand (B+ strand) are determined (step S312).
[0063] (Primer Sequencing Unit) The primer sequencing unit 32 is a unit that carries out the primer sequencing step S22 shown in FIG. 2 . The primer sequencing unit 32 creates predetermined combinations (pairs) of sequences from among the one or more candidate primer sequences determined in the candidate primer sequence selection unit 30, i.e., one or more candidate forward primer sequences for the first template strand (A+ strand), one or more candidate reverse primer sequences for the first template strand (A+ strand), one or more candidate forward primer sequences for the second template strand (B+ strand), and one or more candidate reverse primer sequences for the second template strand (B+ strand), dividing the cases into (I) cases where one or more primer sequences for different target sites have not yet been determined and (II) cases where one or more primer sequences for different target sites have already been determined, calculates a local alignment score between the sequences of each combination, and, based on whether the value exceeds a predetermined threshold, adopts and determines forward primer sequences and reverse primer sequences that will amplify a region in each strand (A+ strand or B+ strand) containing the predetermined target site selected in the partial sequence excision unit 28. Below we describe the primer sequencing method carried out on each strand.
[0064] (I) In the case where one or more primer sequences for different target sites have not yet been determined, [1] one or more candidate primer sequence pairs spanning a predetermined target site are selected from one or more candidate primer sequences of the first template strand (A+ strand); [2] one pair is selected from one or more candidate primer sequence pairs for the predetermined target site, and a local alignment score is calculated between the sequences of the selected candidate primer sequence pair; and [3] the candidate primer sequence pair for which a local alignment score is calculated that is equal to or less than the predetermined threshold is adopted and determined as a forward primer sequence and a reverse primer sequence for amplifying a region in the first template strand (A+ strand) including the predetermined target site.
[0065] Here, if the score of the candidate primer sequence pair selected in [2] above is higher than the threshold and a primer sequence pair (forward primer sequence and reverse primer sequence) cannot be determined, a different pair is selected from the candidate primer sequence pairs selected in [1] above, and steps [2] and [3] are performed. These steps are repeated until at least one primer sequence pair is determined. If at least one primer sequence pair can be determined, it is not necessary to perform steps such as calculating the scores of all candidate primer sequence pairs selected in [1] above. Instead, the process may return to the partial sequence extraction step S18, select another target site (step S280 in FIG. 4), and determine a primer sequence for another target site. This has the effect of reducing calculation costs and saving effort and time. Furthermore, if the scores of all candidate primer sequence pairs selected in [1] above are higher than the threshold and the pair cannot be adopted and determined as a primer sequence pair, the process may return to the partial sequence extraction step S18, select another target site (step S280 in FIG. 4), and determine a primer sequence for another target site.
[0066] (II) In the case where one or more primer sequences for different target sites have already been determined, [1] one or more candidate primer sequence pairs spanning the predetermined target site are selected from the one or more candidate primer sequences of the first template strand (A+ strand); [2] one pair is selected from the one or more candidate primer sequence pairs for the predetermined target site, and local alignment scores are calculated between each candidate sequence of the selected candidate primer sequence pair and each primer sequence for the different target site that has already been determined, as well as local alignment scores between the sequences of the selected candidate sequence pair; and [3] the maximum value of all the calculated local alignment scores (i.e., the score of the pair most likely to form primer dimers) is detected, and the candidate primer sequence pair for which a local alignment score has been calculated that is equal to or less than a predetermined threshold is adopted and determined as the forward primer sequence and reverse primer sequence for amplifying the region containing the predetermined target site in the first template strand (A+ strand).
[0067] Here, if the maximum score calculated for the candidate primer sequence pair selected in [2] above is higher than the threshold value and a primer sequence pair (forward primer sequence and reverse primer sequence) cannot be determined, a different pair is selected from the candidate primer sequence pairs selected in [1] above, and steps [2] and [3] are performed. These steps are repeated until at least one primer sequence pair is determined. If at least one primer sequence pair can be determined, it is not necessary to perform steps such as calculating the scores for all candidate primer sequence pairs selected in [1] above. Instead, the process may return to the partial sequence extraction step S18, select another target site (step S280 in FIG. 4), and determine a primer sequence for another target site. This has the effect of reducing calculation costs and saving effort and time. Furthermore, if the maximum score calculated for all candidate primer sequence pairs selected in [1] above is higher than the threshold value and the pair cannot be adopted and determined as a primer sequence pair, the process may return to the partial sequence extraction step S18, select another target site (step S280 in FIG. 4), and determine a primer sequence for another target site.
[0068] Here, the local alignment score is calculated as follows: <1> "X" is assigned to each position for pairs of complementary bases, <2> "Y" is assigned to each position for pairs of non-complementary bases, and <3> "Z" is assigned to each position for insertions or deletions between candidate primer sequences, or between candidate primer sequences and an already determined primer sequence. The "X" is 1, the "Y" is −4 to −2, and the "Z" is −6 to −3. The predetermined threshold is 1 to 4.
[0069] The present inventors have focused on the parameters (e.g., complementary score: 1, non-complementary score: -1, gap / deletion score: -2), threshold (0), and sequence comparison method (e.g., brute-force search of candidate sequences) typically used in score calculations. They have extensively investigated previously unexplored methods for calculating local alignment scores, the sequence comparison combinations and order involved in score calculation, and predetermined thresholds for selecting primer sequences. They have discovered that the above method can achieve a high success rate in primer design while minimizing the rate of primer-dimer formation at low computational cost. In particular, the greater the number of target sites, specifically, when designing primers with 50 or more target sites, the more pronounced the desired effects of the present invention can be achieved. Conventional methods for achieving the above effects require computationally expensive methods, such as chemical energy calculations and deep learning, to improve score calculations. However, when there are many target sites, the calculations cannot be completed within a realistic time frame. However, the method of the present application can achieve the above-mentioned effect within the "scope of score calculation by simple addition," and therefore has the effect of enabling design to be performed in a realistic time (about a few days on a general computer) with low computational costs if the number of target sites is in the order of several thousand.
[0070] The primer sequence determination unit 32 is configured by a computer and performs the function of selecting and determining a forward primer sequence and a reverse primer sequence from the one or more candidate primer sequences described above.
[0071] Here, the sequence comparison (sequence combination) method and its order for calculating the local alignment score will be described in more detail with reference to FIGS.
[0072] First, the above (I) primer sequence determination method when one or more primer sequences for different target sites have not yet been determined will be described. In the above step [1], first, all primer pairs (combinations of forward primer and reverse primer) that can be prepared from one or more candidate forward primer sequences of the first template strand (A+ strand) and one or more candidate reverse primer sequences of the first template strand (A+ strand) are obtained, and the length of the PCR amplification product expected to be amplified by PCR for each primer pair is calculated. Next, the calculated length of the PCR amplification product is determined to be within a predetermined range. If the calculated length of the PCR amplification product is within the predetermined range, the primer pair (i.e., the combination of the candidate forward primer sequence of the first template strand and the candidate reverse primer sequence of the first template strand) for which the length of the PCR amplification product was calculated is adopted as one or more candidate primer sequence pairs for amplifying the region containing the target site selected in the partial sequence extraction unit 28 (partial sequence extraction step), i.e., one or more candidate forward primer sequence and reverse primer sequence pairs of the first template strand (step S320 in FIG. 4). Here, the "predetermined range" used to determine the calculated length of the PCR amplification product is a range that includes the length of the PCR amplification product desired by the user, and as mentioned above, is not particularly limited as long as it is within a known range, i.e., 70 to several kilobp. It is preferable to take into account the success rate of the PCR and the sequencing capability of the DNA sequencer, etc. 6A shows candidate primer sequences (three candidate forward primer sequences and two candidate reverse primer sequences) selected by candidate primer sequence selection unit 30 (candidate primer sequence selection step). FIG. 6B shows one or more candidate primer sequence pairs spanning a predetermined target site selected in step [1] above (i.e., in this case, the lengths of the PCR amplification products predicted to be amplified by PCR for all pairs of the three candidate forward primer sequences and two candidate reverse primer sequences were determined to be within a predetermined range).
[0073] Next, in step [2] above, from the candidate primer sequence pairs (6 pairs) shown in Figure 6B, a "forward candidate sequence FC1" and a "reverse candidate sequence RC1" are selected as one pair, and a local alignment score between the sequences of that pair is calculated. Next, in step [3] above, if the calculated local alignment score value is equal to or less than a predetermined threshold, the pair of "forward candidate sequence FC1" and "reverse candidate sequence RC1" selected in step [2] above are adopted and determined as the forward primer sequence and reverse primer sequence for amplifying the region containing the selected predetermined target site (steps S324 and S322 in Figure 4).
[0074] Next, we will explain the primer sequence determination method (II) when one or more primer sequences for different target sites have already been determined. In step [1] above, similar to step (I)-[1] above, one or more candidate forward primer and candidate reverse primer sequence pairs for the first template strand are first selected to amplify the region containing the target site selected in the partial sequence extraction unit 28 (partial sequence extraction step) (step S320 in Figure 4). Figure 6A shows the candidate primer sequences (three candidate forward primer sequences and two candidate reverse primer sequences) selected in the candidate primer selection step. Figure 6B shows one or more candidate primer sequence pairs for the six predetermined target sites selected in [1] above (i.e., in this case, the lengths of the PCR amplification products predicted to be amplified by PCR for all pairs of the three candidate forward primer sequences and two candidate reverse primer sequences were determined to be within a predetermined range). Figure 7A shows the primer sequence pairs for different target sites P1 and P2 that have already been determined.
[0075] Next, in step [2] above, a "forward candidate sequence FC1" and a "reverse candidate sequence RC1" are selected as one pair from the candidate primer sequence pairs shown in Figure 6B, and local alignment scores are calculated between each candidate sequence and each primer sequence for a different target site that has already been determined, as well as between the candidate primer sequence paired with the selected candidate sequence. That is, as shown in Figure 7B, local alignment scores are calculated between the "forward candidate sequence FC1" and the "forward sequence of target site P1," the "reverse sequence of target site P1," the "forward sequence of target site P2," or the "reverse sequence of target site P2," as well as between the "reverse candidate sequence RC1" and the "forward sequence of target site P1," the "reverse sequence of target site P1," the "forward sequence of target site P2," or the "reverse sequence of target site P2," and between the pair of the "forward candidate sequence FC1" and the "reverse candidate sequence RC1." Next, the maximum value among the nine calculated local alignment scores in the above step [3] is detected, and the candidate primer sequence pair for which the calculated local alignment score is equal to or less than a predetermined threshold value is adopted and determined as the forward primer sequence and reverse primer sequence for amplifying the region containing the selected predetermined target site (steps S324 and S322 in FIG. 4).
[0076] Here, with reference to an example of a local alignment shown in FIG. 8 , a method for calculating a local alignment score and a method for determining whether or not a candidate primer sequence is to be used as a primer sequence will be described in more detail. FIG. 8 shows, from top to bottom, local alignments of (1) sequence [I] and sequence [II], (2) sequence [I] and sequence [III], and (3) sequence [I] and sequence [IV], which are used to determine whether or not a candidate primer sequence should be used as a primer sequence. In the figure, a "|" is added when bases between the sequences form a complementary pair, a ":" is added when a non-complementary pair is formed, a "-" is added to a gap, and nothing is added to a deletion. In calculating the local alignment score, the following conditions were set between the sequences: <1> "X" = 1 per base pair; <2> "Y" = -3 per base pair; and <3> "Z" = -6 per base insertion or deletion. The threshold was set to 4.
[0077] Between sequence [I] and sequence [II] in (1) above, there are five complementary pairs, so the score is 1×5−3×0−6×0=5 (top diagram of Figure 8). However, this score exceeds the threshold of 4, so candidate primer sequence [I] is not adopted. Between sequence [I] and sequence [III] in (2) above, there are four complementary pairs and one non-complementary pair, so the score is 1×4−3×1−6×0=1. This score is below the threshold of 4, so candidate primer sequence [I] can be adopted. Between sequence [I] and sequence [IV] in (3) above, there are nine complementary pairs and one deletion, so the score is 1×9−3×0−6×1=3. This score is below the threshold of 4, so candidate primer sequence [I] can be adopted.
[0078] In step [2] above, all pairs (6 pairs) are selected from the candidate primer sequence pairs shown in Figure 6B, and it is assumed that the maximum local alignment score calculated for each pair is as shown in Table 1. Here, if the predetermined threshold is set to 3, only four of the six candidate sequence primer pairs with a maximum value of 3 or less will be adopted and determined as the forward primer sequence and reverse primer sequence that amplify the region containing the predetermined target site. In other words, these four candidate primer sequence pairs are determined by adopting the forward sequence and reverse sequence of the first template strand (A+ strand) as the primer sequences.
[0079]
[0080] Similarly, first, from among one or more candidate forward primer sequences for the second template strand (B+ strand) and one or more candidate reverse primer sequences for the second template strand (B+ strand), predetermined combinations (pairs) of sequences are created, divided into cases where (I) one or more primer sequences for different target sites have not yet been determined and (II) one or more primer sequences for different target sites have already been determined. A local alignment score is calculated between the sequences in each combination, and based on whether or not this value exceeds a predetermined threshold, forward primer sequences and reverse primer sequences that will amplify a region in the B+ strand containing the predetermined target site selected by partial sequence excision unit 28 are adopted and determined (step S322).
[0081] Once it has been determined whether the length of the PCR amplification product is within a predetermined range for all primer pairs, the partial sequence extraction unit 28 (partial sequence extraction step) determines whether all target sites have been selected (step S24). If all target sites have not been selected, the process returns to the partial sequence extraction step S18 to select another target site (step S280). If all target sites have been selected, the process ends.
[0082] (Controller) The controller 34 is connected directly or indirectly to the input unit 12, storage unit 14, and output unit 16 as well as to each unit within the primer design processor 18, and controls each unit of the primer design device 10 to design primers based on user instructions from the input unit 12 or based on a predetermined operating program stored in the storage unit 14, and is configured, for example, with a CPU (Central Processing Unit) of a computer or the like. The controller 34 controls the candidate primer sequence selection unit 30 to repeat the determination process (steps S300 to S308) until it has been determined whether all partial sequences satisfy all of the predetermined selection criteria (step S310).
[0083] The control unit 34 controls the primer sequence determination unit 32 to repeat the determination operation (step 320) until it has been determined whether the length of the PCR amplification product is within a predetermined range for all prepared primer pairs. The control unit 34 controls the primer sequence determination unit 32 to select a different pair from the candidate primer sequence pairs selected in [1] of (I) and (II) and repeat the steps [2] and [3] of (I) and (II) until at least one primer sequence pair associated with a predetermined target site is determined. If a primer sequence pair associated with the predetermined target site has not been determined, the control unit 34 controls the primer sequence determination unit 32 to select a different target site in the partial sequence excision step (steps S18, S280 to S282) and perform the candidate primer sequence selection step (steps S20, S300 to S312) and the primer sequence determination step (steps S22, S320 to S322).
[0084] The control unit 34 controls the partial sequence excision unit 28, the candidate primer sequence selection unit 30, and the primer sequence determination unit 32 so that the partial sequence excision unit 28 repeats the partial sequence excision step (steps S18, S280 to S282), the candidate primer sequence selection step (steps S20, S300 to S312), and the primer sequence determination step (steps S22, S320 to S322) until all target sites acquired by the target site information acquisition unit 22 are detected (step S24).
[0085] The primer design device 10 according to the first embodiment of the present invention can design primers for amplicon methylation sequencing analysis with a high design success rate. Furthermore, primers based on the designs can be obtained. As a result, primers can be designed for a larger number of target sites and the degree of methylation can be measured.
[0086] [Modification 1] Next, a primer design device according to Modification 1 of Embodiment 1 of the present invention will be described. Note that in this primer design device according to Modification 1, explanations of processes similar to those in Embodiment 1 will be omitted. In Embodiment 1, in determining the primer sequence, the number of candidate primer sequence pairs for which local alignment scores are calculated and the number of forward primer sequences and reverse primer sequences that amplify a region including a predetermined target site are not particularly limited, but the present invention is not limited thereto, and it is also possible to calculate scores for all pairs and select only one primer sequence pair that amplifies a region including each target site.
[0087] In this first modification, the primer sequence determination unit 32 can also perform the following steps: (I) If one or more primer sequences for different target sites have not yet been determined, in step [2] above, select all pairs from one or more candidate primer sequence pairs for the predetermined target site and calculate a local alignment score between the sequences of the selected candidate primer sequence pairs for each pair, and in step [3] above, select one or more candidate primer sequence pairs whose calculated local alignment score is equal to or less than a predetermined threshold, and further detect the candidate primer sequence pair having the smallest maximum value of the local alignment score from all selected pairs, and adopt and determine them as the forward primer sequence and reverse primer sequence for amplifying a region including the predetermined target site.
[0088] (II) When one or more primer sequences for different target sites have already been determined, in step [2] above, all pairs are selected from one or more candidate primer sequence pairs for the predetermined target site, and for each pair, a local alignment score is calculated between each candidate sequence in the selected candidate primer sequence pair and each primer sequence for the different target site that has already been determined, as well as a local alignment score between the sequences of the selected candidate sequence pair; and in step [3] above, for each pair, the maximum value is detected from all the calculated local alignment scores, and the candidate primer sequence pair for which a local alignment score has been calculated that is equal to or less than a predetermined threshold is selected, and further, from all the selected pairs, the candidate primer sequence pair with the smallest maximum value of the local alignment score is detected, and these are adopted and determined as the forward primer sequence and reverse primer sequence for amplifying a region including the predetermined target site.
[0089] By carrying out such a process, it is possible to determine, from among one or more primer sequence pairs capable of amplifying a region including a predetermined target site, the pair with the lowest primer-dimer formation rate and the highest primer design success rate as the primer sequence.
[0090] For example, suppose that in step [2] above, all pairs are selected from the candidate primer sequence pairs shown in Figure 6B, and the maximum local alignment score calculated for each pair is as shown in Table 1. If the predetermined threshold is set to 3, pairs having a maximum score equal to or less than the predetermined threshold are selected, with four pairs having a maximum value of 3 or less. Furthermore, from all the selected pairs, the candidate primer sequence pair having the smallest maximum local alignment score is the candidate primer sequence pair of "forward candidate sequence FC2" and "reverse candidate sequence RC2," and therefore this pair is adopted and determined as the forward primer sequence and reverse primer sequence for amplifying a region including the predetermined target site.
[0091]
[0092] [Modification 2] Next, a primer design device according to Modification 2 of Embodiment 1 of the present invention will be described. Note that in this primer design device according to Modification 2, the same processes as those in Embodiment 1 will not be described again. In Embodiment 1, primers are designed without the user setting a primer design rate, but this is not limiting, and primer design can also be performed based on a preset primer design success rate desired by the user.
[0093] The correspondence relationships between at least a predetermined threshold, the number of target sites (measurement sites), and the primer design success rate are measured in advance using the primer design device (method) described in Embodiment 1 and each of the modified examples, and the correspondence relationships are stored in the storage unit 14. Here, the "predetermined threshold" is not particularly limited as long as it is 1 to 4. When X, Y, and Z are all integers, if the number of target sites is less than 1,000, it is preferable to create correspondence relationships for all of thresholds 1, 2, 3, and 4. If the number of target sites is 1,000 or more, it is preferable to create correspondence relationships for at least two or more thresholds. When X, Y, and Z include non-integer values, if the number of target sites is less than 1,000, it is preferable to create correspondence relationships for at least five or more thresholds. If the number of target sites is 1,000 or more, it is preferable to create correspondence relationships for at least two or more thresholds.
[0094] When the user sets at least the desired primer design success rate and number of target sites via the input unit 12 and inputs an instruction to execute primer design, the primer sequence determination unit 32 reads out the predetermined threshold value corresponding to the primer design success rate and number of target sites that is equal to or greater than the set values of the primer design success rate and number of target sites and whose difference is small from the corresponding relationships stored in the storage unit 14, and determines the primer sequence based on the predetermined threshold value.
[0095] With the primer design device of Modification 2, the user can easily and inexpensively design primer sequences according to the background and circumstances at the time of primer design, such as when there are few samples, when the user wants to attempt primer design with a desired primer design success rate or multiple primer design success rates, and can also obtain primers based on that design.
[0096] A method for selecting a threshold value used in determining primer sequences will be specifically described with reference to Figure 9. Figure 9 shows the primer design success rate measured in advance by designing primers for 100 target sites using the primer design device (method) described in embodiment 1 and each of the modifications, with the threshold value for determining the local alignment score of each pair being an integer between 1 and 4. This correspondence relationship is stored in the storage unit 17.
[0097] For example, if the user desires a primer design success rate of 30% or higher, the user sets at least the primer design success rate to 30% and the number of target sites to 100, and inputs an instruction to execute primer design via the input unit 12. From the correspondence relationships stored in the storage unit 14, the number of target sites that satisfy the conditions, 100, and the threshold value 2 corresponding to the primer design success rate of 31%, which is greater than and closest to the user's desired primer design success rate of 30%, are read out to the primer sequence determination unit 32, and a primer sequence determination step is executed based on the threshold value 2 in the correspondence relationship to obtain primer sequence pairs for 31 sites.
[0098] [Modification 3] Next, a primer design device according to Modification 3 of Embodiment 1 of the present invention will be described. Note that in this primer design device according to Modification 3, explanations of processes similar to those in Embodiment 1 will be omitted. In Embodiment 1, the cytosines (C) that can be methylated are limited to cytosines (C) in a CG sequence, and a cytosine (C) picked from among them is used as a target site, but this is not limited thereto, and the cytosines (C) that can be methylated may also include cytosines (C) in a CHG sequence, and a cytosine (C) picked from among them may also be used as a target site.
[0099] In this third modification, the target site information acquisition unit 22 further acquires, via the input unit 12, two or more target sites contained in the genomic double-stranded DNA acquired by the base sequence data acquisition unit 20, and their position information. The base conversion unit 24 further converts cytosine (C) in the CHG sequence on the template DNA acquired from the base sequence data acquisition unit 20 to "Y", and converts cytosine (C) in other sequences (i.e., sequences other than the CG sequence and the CHG sequence) to thymine (T).
[0100] The candidate primer sequence selection unit 30 selects, as candidate primer sequences, those that satisfy all of the predetermined selection conditions (1) to (4), including the following item (4), from one or more partial sequences of each strand excised by the partial sequence excision unit 28: (4) The partial sequence contains a predetermined number or less of YHG sequences or CDR sequences. Here, the number of "YHG sequences or CDR sequences contained in the partial sequence" according to the above condition (4) is not particularly limited, but from the viewpoint of significantly achieving the desired effect of the present invention, it is preferably 2 or less, more preferably 1 or less, and particularly preferably 0. By satisfying this condition, it is possible to reduce the influence of the junction between the primer and the cytosine (C) of the CHG sequence at the primer junction site.
[0101] The primer design device of Modification 3 of Embodiment 1 of the present invention allows for easy and rapid design of primers for amplicon methylation sequence analysis that also correspond to CHG sequences. Primers based on the design can also be obtained. As a result, analysis of these sequences becomes possible, allowing for more detailed analysis of the DNA methylation state (degree of methylation). Modification 3 can also be combined with Modifications 1 and 2.
[0102] [Modification 4] Next, a primer design device according to Modification 4 of Embodiment 1 of the present invention will be described. Note that in this primer design device according to Modification 4, explanations of processes similar to those in Embodiment 1 will be omitted. In Embodiment 1, the cytosines (C) that can be methylated are limited to cytosines (C) in a CG sequence, and a cytosine (C) picked from among them is used as a target site, but this is not limited thereto, and the cytosines (C) that can be methylated may also include cytosines (C) in a CHH sequence, and a cytosine (C) picked from among them may also be used as a target site.
[0103] In this fourth modification, the target site information acquisition unit 22 further acquires, via the input unit 12, two or more target sites contained in the genomic double-stranded DNA acquired by the base sequence data acquisition unit 20, and their position information. The base conversion unit 24 further converts cytosine (C) in the CHH sequence on the template DNA acquired from the base sequence data acquisition unit 20 to "Y", and converts cytosine (C) in other sequences (i.e., sequences other than the CG sequence and the CHH sequence) to thymine (T).
[0104] The candidate primer sequence selection unit 30 selects, as candidate primer sequences, those that satisfy all of the predetermined selection conditions (1) to (3) and (5), including the following item (5), from one or more partial sequences of each strand excised by the partial sequence excision unit 28: (5) The partial sequence contains a predetermined number or less of YHH sequences or DDR sequences. Here, the number of "YHH sequences or DDR sequences contained in the partial sequence" according to the above condition (5) is not particularly limited, but from the viewpoint of significantly achieving the desired effect of the present invention, it is preferably 2 or less, more preferably 1 or less, and particularly preferably 0. By satisfying this condition, it is possible to reduce the influence of the junction between the primer and the cytosine (C) of the CHH sequence at the primer junction site.
[0105] The primer design device according to the fourth modification of the first embodiment of the present invention allows for easy and rapid design of primers for amplicon methylation sequence analysis, even for CHH sequences. Furthermore, primers based on the design can be obtained. As a result, analysis of these sequences becomes possible, enabling more detailed analysis of the DNA methylation state (degree of methylation).
[0106] Variation 4 can also be combined with the previously described Variation 1 or 2. Variation 4 can also be combined with the previously described Variation 3. That is, the cytosines (C) that can be methylated may include both cytosines (C) in the CHG sequence and in the CHH sequence, and a cytosine (C) selected from among them may be used as the target site. In such cases, the candidate primer sequence selection unit 30 selects, as candidate primer sequences, those that satisfy all of the selection conditions (1) to (5) from one or more partial sequences of each strand excised by the partial sequence excision unit 28.
[0107] [Variation 5] Next, a primer design device according to Variation 5 of Embodiment 1 of the present invention will be described. Note that in this primer design device according to Variation 4, components similar to those in Embodiment 1 are assigned the same reference numerals, and descriptions of processes similar to those in Embodiment 1 will be omitted. In Embodiment 1, an apparatus and method for designing two pairs of primers to amplify and analyze both strands of double-stranded DNA were described. However, this is not limited to this. To analyze only one strand of double-stranded DNA, it is sufficient to design only one pair of primers. That is, while primers are designed based on strands A and B in FIG. 3B , primers may be designed based on strand A alone. Furthermore, if the DNA maintenance methylation mechanism is believed to be active, it is sufficient to design only one pair of primers. This is because, if a C in a CG sequence in one strand of DNA is methylated, it is highly likely that a C in a CG sequence in the other strand is also methylated, and, if a C in a CG sequence in one strand is unmethylated, it is highly likely that a C in a CG sequence in the other strand is also unmethylated. In such a case, if a pair of primers cannot be designed based on one strand, primers can be designed based on the other strand.
[0108] In this way, when only one set of primers is designed, only a complementary strand A- having a base sequence complementary to the base sequence of the A+ strand shown in Figure 3C is produced in the complementary strand generation unit 26. Next, the partial sequence excision unit 28 selects one target site from the two or more target sites acquired by the target site information acquisition unit 22 (step S280), and based on the position information of the selected target site, detects "Y" or its complementary "R" (i.e., a base that is in the target site and is at the methylation site) of the selected target site from the DNA sequences of the A+ strand and A- strand, and excises as many bases as possible from partial sequences of a predetermined length from the base sequences located on the 5'-terminal side of the detected "Y" and "R" ((1) and (2) in Figure 3D) to obtain one or more partial sequences (step S282).
[0109] The candidate primer sequence selection unit 30 is a unit that carries out the candidate primer sequence selection step S20 shown in Figure 2, and selects as candidate primer sequences those that satisfy all of the predetermined selection conditions (1) to (3) from one or more partial sequences of each strand excised by the partial sequence excision unit 28. One or more partial sequences excised from the first template strand (A+ strand) (i.e., one or more partial sequences excised from (1) in Figure 3D) that satisfy the predetermined selection conditions are selected as candidate forward primer sequences for the first template strand (A+ strand), and one or more partial sequences excised from the first complementary strand (A- strand) (i.e., one or more partial sequences excised from (2) in Figure 3D) that satisfy the predetermined selection conditions are selected as candidate reverse primer sequences for the first template strand (A+ strand).
[0110] The primer sequence determination unit 32 creates predetermined combinations (pairs) of sequences from among the one or more candidate forward primer sequences for the selected first template strand (A+ strand) and one or more candidate reverse primer sequences for the first template strand (A+ strand), dividing the sequences into (I) cases where one or more primer sequences for different target sites have not yet been determined and (II) cases where one or more primer sequences for different target sites have already been determined, calculates a local alignment score between the sequences of each combination, and, based on whether or not this value exceeds a predetermined threshold, adopts and determines the forward primer sequence and reverse primer sequence for amplifying the region containing the predetermined target site selected by the partial sequence excision unit 28 in the A+ strand. Note that Variation 5 can be combined with at least one of the previously described Variations 1 to 4.
[0111] [Embodiment 2] FIG. 10 is a block diagram conceptually illustrating an example of a primer design device according to Embodiment 2 of the present invention. The primer design device 10 of Embodiment 1 may also include a communication interface (communication device). The primer design device 10A of Embodiment 2 shown in FIG. 10 has the same configuration as the primer design device 10 of Embodiment 1 shown in FIG. 1, except for the inclusion of a communication interface 36. Therefore, the same components are designated by the same reference numerals and their description will be omitted. As shown in FIG. 11, the primer design device 10A can be connected to a search server 42 equipped with a public database installed outside the device via a communication network 38, such as the Internet. The device 10A of this embodiment can execute at least one of the base sequence data acquisition unit 20, target site information acquisition unit 22, base conversion unit 24, complementary strand generation unit 26, partial sequence excision unit 28, primer candidate sequence selection unit 30, and primer sequence determination unit 32 via the communication interface 36 using a program on the site of an external server 40. In such a case, the primer design device 10A of this embodiment does not need to include each means executed by a program on an external server.
[0112] For example, based on instructions from the control unit 34, the communication interface 36 can obtain DNA base sequences including genes and genomes from public databases via the communication network 38 and store the obtained sequences in the storage unit 14. Examples of public databases include GenBank from the National Center for Biotechnology Information (NCBI) in the United States, ENA from the European Molecular Biology Laboratory (EMBL), and DDBJ from the National Institute of Genetics. The base sequences obtained from the public databases may be partial base sequences of the genomic DNA of the organism species for which primers are designed, but preferably the entire sequence.
[0113] For example, the communication interface 36 can use a public search server 42 to perform a sequence homology search and a local alignment search of the primer sequence determination unit 32, based on instructions from the control unit 34, via a communication network 38. Here, an example of the public search server 42 is BLAST from the National Center for Biotechnology Information (NCBI) in the United States.
[0114] [Embodiment 3] Embodiment 3 is a method for producing a primer by synthesizing the primer based on the primer sequence designed by the primer design device and design method according to Embodiments 1 and 2. The primer design method is as shown in Embodiments 1 and 2. The primer can be synthesized by a known method, such as a method of chemically synthesizing the primer from a terminal base using a DNA synthesizer or an RNA synthesizer, using dNTP (Deoxyribonucleoside triphosphate) or the like as a material. A commercially available synthesizer can be used as the synthesizer.
[0115] The components of the device of the present invention may be configured as dedicated hardware, or may be configured as a programmed computer. The method of the present invention can be implemented, for example, by a program that causes a computer to execute each step. Also, a computer-readable recording medium on which this program is recorded can be provided.
[0116] Although the present invention has been described in detail above, the present invention is not limited to the above-described embodiments, and various improvements and modifications may be made without departing from the spirit and scope of the present invention.
[0117] [Example 1, Comparative Example 1] Using the primer design device of this embodiment 1, primers for multiplex PCR were designed, each having a PCR amplification product length of 70 bp to 120 bp, based on the base sequence data of reference genome GRCh37 (GenBank assembly accession: GCA_000001405.1, RefSeq assembly accession: GCF_000001405.13), 100 randomly selected measurement sites (target sites) shown in Table 1, and their location information. The primers were designed so that the length of the primers was 20 to 35 bases (mer), and the only C that could be methylated was C in the CG sequence. The conditions for determining the partial sequence were set as follows: Condition (1): The Tm value is 55°C to 65°C; Condition (2): The partial sequence contains zero YG or CR sequences; and Condition (3): The upper limit of the number of junctions with sequences outside the relevant region is 2.
[0118] In calculating the local alignment score, in Example 1, between the sequences, <1> for a pair of complementary bases, "X" = 1 per position, <2> for a non-complementary pair, "Y" = -3 per position, and <3> if there is an insertion or deletion, "Z" = -6 per position, and the threshold was set to 1. On the other hand, in Comparative Example 1, between the sequences, <1> for a pair of complementary bases, "X" = 1 per position, <2> for a non-complementary pair, "Y" = -1 per position, and <3> if there is an insertion or deletion, "Z" = -2 per position, and the threshold was set to 0.
[0119] Table 3 shows the success or failure of primer design at each measurement site in Example 1 and Comparative Example 1, and the primer design success rate calculated from the results of the success or failure of primer design. Table 4 shows the primers that could be designed in Example 1, and Table 5 shows the primers that could be designed in Comparative Example 1. For each primer pair, the first pair that resulted in a maximum local alignment score below the threshold was used.
[0120] As shown in Table 3, the success rate of each primer design was 62% in Example 1 and 4% in Comparative Example 1. These results confirmed that the success rate of primer design can be increased by setting the threshold for the maximum value of the local alignment score within a predetermined range in the primer sequencing step.
[0121]
[0122]
[0123]
[0124] [Examples 1 to 4, Comparative Examples 2 to 4] Using the primer design device of the present embodiment 1, primer sequences for multiplex PCR with PCR amplification product lengths of 70 bp to 120 bp were designed based on the base sequence data of reference genome GRCh37 (GenBank assembly accession: GCA_000001405.1, RefSeq assembly accession: GCF_000001405.13), 100 randomly selected measurement sites (target sites) shown in Table 1, and their location information. The primers were designed with a length of 20 to 35 bases (mer), and the only C that could be methylated was C in the CG sequence. The conditions for selecting the partial sequence were set as follows: Condition (1): Tm value is 55°C to 65°C Condition (2): The partial sequence contains zero YG or CR sequences Condition (3): The upper limit of the number of junctions with sequences outside the relevant region is 2
[0125] In calculating the local alignment score, between the sequences, <1> for a pair of complementary bases, "X" = 1 per position, <2> for a pair of non-complementary bases, "Y" = -3 per position, and <3> for an insertion or deletion, "Z" = -6 per position. The thresholds for Examples 1 to 4 and Comparative Examples 2 to 4 were set as shown in Table 6.
[0126] We also calculated the dimer formation rate of primers using the same local alignment score conditions (parameters and thresholds used for score calculation) used in each Example and Comparative Example. A set of primers with local alignment scores ranging from 0 to 6 was prepared to amplify 91 separately selected target sites (i.e., one pair of primers was designed for each target site, for a total of 182 primers). Bisulfite-treated sample DNA (Zymo Research Human WGA Methylated DNA) was amplified by multiplex PCR. The sequences of the resulting amplification products were obtained using a next-generation sequencer (Illumina MiSeq). The obtained sequences consisted of the desired amplification product containing the target site, primer dimers, and other nonspecific amplification products. All possible primer dimer sequences generated from the prepared primer sequences were generated in a computer, and the actual primer dimer sequences and their number were determined by comparing and summarizing them with the sequences obtained by the next-generation sequencer. All combinations of two sequences selected from the prepared primer sequences were divided into seven groups, from 0 to 6, based on the local alignment score. The percentage of two sequences in each group in which primer dimers were actually generated (10 or more sequences obtained by the next-generation sequencer) was calculated and used as the dimer formation rate.
[0127]
[0128] Table 7 shows the success or failure of primer design at each measurement site in Examples 1 to 4 and Comparative Examples 2 to 4, and the primer design success rate calculated from the results of the primer design success or failure. Tables 8 to 10 show the primers that could be designed in Examples 2 to 4, and Tables 11 to 13 show the primers that could be designed in Comparative Examples 2 to 4. Figure 12A shows the primer design success rate for each threshold value set during primer sequencing based on each Example and Comparative Example, and Figure 12B shows the dimer formation rate for each threshold value designed in each Example and Comparative Example.
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136] As shown in Table 7 and Figures 12A and 12B, the maximum local alignment score was determined using integer thresholds 1 to 4, and the adopted primer sequence pairs (Examples 1 to 4) achieved a high primer design success rate while suppressing dimer formation to an extremely low level of 2% or less. On the other hand, the maximum local alignment score was determined using thresholds 0, 5, and 6, and the adopted primer sequence pairs (Comparative Examples 2 to 4) achieved a low primer design success rate despite a low dimer formation rate, or a high design success rate despite a high dimer formation rate.
[0137] The primer design success rate in Comparative Example 3 (84%) was slightly higher than that in Example 4 (82%). However, in Example 4, i.e., when multiplex PCR was performed using primers designed and manufactured according to the present invention, the dimer formation rate was suppressed to 2% or less, whereas in Comparative Example 3, i.e., when the maximum local alignment score was determined at a threshold of 5 outside the numerical range according to the present invention, dimers were formed in approximately 20% of the primer sequence pairs used. Therefore, when the primer sequence pairs designed and manufactured in Comparative Example 3 are used in multiplex PCR, problems such as failure to amplify the desired target site or the generation of a large amount of primer dimers that inhibit the sequence of another amplified target site are likely to occur, resulting in failure.
[0138] 10, 10A Primer design device 12 Input unit 14 Storage unit 16 Output unit 18 Primer design processing unit 20 Base sequence data acquisition unit 22 Target site information acquisition unit 24 Base conversion unit 26 Complementary strand generation unit 28 Partial sequence excision unit 30 Primer candidate sequence selection unit 32 Primer sequence determination unit 34 Control unit 36 Communication interface 38 Communication line network 40 Server 42 Search server
[0139] The primers designed in the present invention can be used to measure the degree of DNA methylation in biological samples in the fields of drug discovery, diagnosis, and other bioindustry fields.
Claims
1. A method for designing primers for amplicon methylation sequence analysis used to simultaneously amplify multiple regions each including two or more target sites for measuring the methylation degree of at least one genomic double-stranded DNA, using a bisulfite reaction or an enzymatic reaction and multiplex PCR, the method comprising: a complementary strand generating step of generating a complementary strand to the template strand of the DNA; a partial sequence excision step of selecting one of the two or more target sites and excising one or more partial sequences of a predetermined length from a base sequence located on the 5'-terminal side of the selected target site from each of the strands; a step of selecting one or more candidate primer sequences from the one or more excised partial sequences as one or more candidate primer sequences; a primer sequence determination step of selecting and determining a forward primer sequence and a reverse primer sequence that amplify a region including the selected predetermined target site from among the one or more candidate primer sequences; a repeating step of repeating the partial sequence excision step, the candidate primer sequence selection step, and the primer sequence determination step until all of the two or more target sites are selected in the partial sequence excision step; having (I) If one or more primer sequences for different target sites have not yet been determined, The primer sequencing step comprises: [1] selecting one or more candidate primer sequence pairs spanning the predetermined target site from the one or more candidate primer sequences; [2] selecting one pair from one or more candidate primer sequence pairs for the predetermined target site, and calculating a local alignment score between the sequences of the selected candidate primer sequence pair; [3] determining a pair of candidate primer sequences for which the local alignment score is calculated to be equal to or less than the predetermined threshold value as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site; (II) When one or more primer sequences for different target sites have already been determined, The primer sequencing step comprises: [1] selecting one or more candidate primer sequence pairs spanning a predetermined target site from the one or more candidate primer sequences; [2] Selecting one pair from one or more candidate primer sequence pairs for the predetermined target site, and calculating a local alignment score between each candidate sequence of the selected candidate primer sequence pair and each primer sequence for a different target site already determined, and a local alignment score between the sequences of the selected candidate sequence pair; [3] A method for determining a candidate primer sequence pair for which a maximum value is detected from among all the calculated local alignment scores and the maximum value is equal to or less than a predetermined threshold value, the candidate primer sequence pair being selected as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site, In the step [3] of (I) and (II), if the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site are not adopted, a different pair is selected from the one or more candidate primer sequence pairs selected in [1] of (I) and (II), and the steps [2] and [3] are repeated until at least one candidate primer sequence pair is adopted; The local alignment score is calculated as follows: <1> each pair of complementary bases is represented by "X", <2> each pair of non-complementary bases is represented by "Y", and <3> each insertion or deletion is represented by "Z", where "X" is 1, "Y" is -4 to -2, and "Z" is -6 to -3; The predetermined threshold value is 1 to 4. Method for designing primers for amplicon methylation sequencing analysis.
2. The primer sequencing step comprises: In the case where the number of target sites is two or more and one or more primer sequences for the different target sites have not yet been determined, In the step [2], all pairs are selected from the one or more candidate primer sequence pairs for the predetermined target site, and a local alignment score between the sequences of the selected candidate primer sequence pair is calculated for each pair; In the step [3], one or more candidate primer sequence pairs having a calculated local alignment score equal to or less than the predetermined threshold value are selected, and further, from all the selected pairs, a candidate primer sequence pair having the smallest maximum value of the local alignment score is detected, and adopted and determined as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site; In the case where one or more primer sequences for different target sites have already been determined in the above (II), In the step [2], all pairs are selected from the one or more candidate primer sequence pairs for the predetermined target site, and for each pair, a local alignment score between each candidate sequence of the selected candidate primer sequence pair and each primer sequence of a different target site already determined, and a local alignment score between the sequences of the selected candidate sequence pair, In the step [3], for each pair, the maximum value is detected from all the calculated local alignment scores, and a candidate primer sequence pair for which a local alignment score is calculated with a value equal to or less than a predetermined threshold is selected, and further, from all the selected pairs, a candidate primer sequence pair having the smallest maximum value of the local alignment score is detected, and these are adopted and determined as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site. The method for designing a primer for amplicon methylation sequence analysis according to claim 1.
3. a base sequence data acquisition step of acquiring base sequence data of the genomic double-stranded DNA; a target site information acquisition step of acquiring the two or more target sites and their position information; a base conversion step of converting "C" that can be methylated in the genomic double-stranded DNA in the base sequence data to "Y" and converting other "C" to "T"; and The complementary strand generating step generates a complementary strand for each template strand of the genomic double-stranded DNA after the base conversion, The partial sequence excision step includes selecting one of the two or more target sites, and excising, from each strand based on positional information of the selected target site, one or more partial sequences of a predetermined length from a base sequence located on the 5'-terminal side of the "Y" into which the selected target site is converted or the "R" complementary thereto; The candidate primer sequence selection step includes selecting, as a candidate primer sequence, one or more partial sequences excised from each strand that satisfy a predetermined selection condition; The "C" that can be methylated is "C" in a CG sequence, The predetermined selection conditions are: (1) The Tm value is within a predetermined range; (2) the number of YG sequences or CR sequences contained in the partial sequence is a predetermined number or less; and (3) The upper limit of the number of junctions with a sequence outside the relevant region on the genomic double-stranded DNA after the base conversion is a predetermined number of 1 or more. The method for designing a primer for amplicon methylation sequence analysis according to claim 1 , comprising: [Note that "C", "G", "Y" and "R" are base symbols defined by IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine.]
4. The "C" that can be methylated further includes the "C" in the CHG sequence, The method for designing a primer according to claim 3, wherein the predetermined selection conditions further include (4) that the partial sequence contains no more than a predetermined number of YHG sequences or CDR sequences. [Note that "C", "G", "Y", "H", "R" and "D" are base symbols defined by IUPAC, "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, "H" represents adenine, cytosine or thymine, "D" represents thymine, guanine or adenine, and "R" represents adenine or guanine.]
5. The "C" that can be methylated further includes "C" in a CHH sequence, The method for designing a primer according to claim 3, wherein the predetermined selection conditions further include (5) that the partial sequence contains no more than a predetermined number of YHH sequences or DDR sequences. [Note that "Y", "H", "R" and "D" are base symbols defined by IUPAC, "Y" represents thymine or cytosine, "H" represents adenine, cytosine or thymine, "D" represents thymine, guanine or adenine, and "R" represents adenine or guanine.]
6. The step of selecting candidate primer sequences is as follows: the genomic double-stranded DNA after the base conversion is a first template strand and a second template strand, a complementary strand of the first template strand is a first complementary strand, and a complementary strand of the second template strand is a second complementary strand; 4. The primer design method according to claim 3, wherein the one or more partial sequences excised from the first template strand that satisfy a predetermined selection condition are selected as candidate forward primer sequences for the first template strand, the one or more partial sequences excised from the first complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the first template strand, the one or more partial sequences excised from the second template strand that satisfy the predetermined selection condition are selected as candidate forward primer sequences for the second template strand, and the one or more partial sequences excised from the second complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the second template strand.
7. The primer sequence determination step is a step of calculating the length of a PCR amplification product expected to be amplified by PCR for all combinations of the candidate forward primer sequences of the one or more selected first template strands and the candidate reverse primer sequences of the one or more selected first template strands in the candidate primer sequence selection step, adopting a combination of candidate primer sequences in which the calculated PCR amplification product length is within a predetermined range as the forward primer sequence and reverse primer sequence of the first template strand that amplify the region including the target site selected in the partial sequence excision step, and calculating the length of a PCR amplification product expected to be amplified by PCR for all combinations of the candidate forward primer sequences of the selected second template strand and the candidate reverse primer sequence of the selected second template strand, and adopting a combination of candidate primer sequences in which the calculated PCR amplification product length is within the predetermined range as the forward primer sequence and reverse primer sequence of the second template strand that amplify the region including the target site selected in the partial sequence excision step. The primer design method according to claim 3,
8. Calculating a correspondence relationship between at least the number of target sites, the predetermined threshold, and the primer design success rate in advance using the primer design method according to claim 1, and storing the correspondence relationship in a storage unit; When a user sets at least the primer design success rate and the number of target sites desired by the user via an input unit and instructs to execute primer design, the predetermined threshold value corresponding to the primer design success rate and the number of target sites that is equal to or greater than the set values of the primer design success rate and the number of target sites and has a small difference therebetween is read out from the correspondence relationship stored in the storage unit, The primer design method according to claim 1 , further comprising the step of selecting and determining a primer sequence that amplifies a region including the predetermined target site from among the one or more candidate primer sequences based on the read-out predetermined threshold value.
9. A primer design process; a synthesis step of synthesizing a primer based on the primer sequence designed in the primer design step; Equipped with A method for producing a primer, wherein the primer design step is carried out by the primer design method according to claim 1.
10. An apparatus for designing primers for amplicon methylation sequence analysis used to simultaneously amplify multiple regions each including two or more target sites for measuring the methylation degree of at least one double-stranded DNA, using a bisulfite reaction or an enzyme reaction and multiplex PCR, the apparatus comprising: a complementary strand generating unit that generates a complementary strand to the template strand of the DNA; a partial sequence excision unit that selects one of the two or more target sites and excises one or more partial sequences of a predetermined length from a base sequence located on the 5'-terminal side of the selected target site from each of the strands; a candidate primer sequence selection unit that selects one or more of the excised partial sequences as one or more candidate primer sequences; A primer sequence determination unit that adopts and determines a forward primer sequence and a reverse primer sequence that amplify a region including the selected predetermined target site from among the one or more candidate primer sequences; a control unit that controls the partial sequence excision unit to repeat the processes of the partial sequence excision unit, the candidate primer sequence selection unit, and the primer sequence determination unit until all of the two or more target sites are selected in the partial sequence excision unit; having (I) If one or more primer sequences for different target sites have not yet been determined, The primer sequence determination unit [1] selecting one or more candidate primer sequence pairs spanning the predetermined target site from the one or more candidate primer sequences; [2] selecting one pair from one or more candidate primer sequence pairs for the predetermined target site, and calculating a local alignment score between the sequences of the selected candidate primer sequence pair; [3] determining a pair of candidate primer sequences for which the local alignment score is calculated to be equal to or less than the predetermined threshold value as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site; (II) When one or more primer sequences for different target sites have already been determined, The primer sequence determination unit [1] selecting one or more candidate primer sequence pairs spanning a predetermined target site from the one or more candidate primer sequences; [2] Selecting one pair from one or more candidate primer sequence pairs for the predetermined target site, and for each pair, calculating a local alignment score between each candidate sequence of the selected candidate primer sequence pair and each primer sequence of a different target site already determined, and a local alignment score between the sequences of the selected candidate sequence pair; [3] A method for determining a candidate primer sequence pair for which a maximum value is detected from among all the calculated local alignment scores and the maximum value is equal to or less than a predetermined threshold value, the candidate primer sequence pair being selected as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site, In the step [3] of (I) and (II), if the forward primer sequence and reverse primer sequence for amplifying the region including the predetermined target site are not adopted, a different pair is selected from the one or more candidate primer sequence pairs selected in [1] of (I) and (II), and the steps [2] and [3] are repeated until at least one candidate primer sequence pair is adopted; The local alignment score is calculated as follows: <1> each pair of complementary bases is represented by "X", <2> each pair of non-complementary bases is represented by "Y", and <3> each insertion or deletion is represented by "Z", where "X" is 1, "Y" is -4 to -2, and "Z" is -6 to -3; The predetermined threshold value is 1 to 4. Primer design device for amplicon methylation sequencing analysis.
11. The primer sequence determination unit In the case where one or more primer sequences for different target sites in (I) have not yet been determined, In the step [2], all pairs are selected from the one or more candidate primer sequence pairs for the predetermined target site, and a local alignment score between the sequences of the selected candidate primer sequence pair is calculated for each pair; In the step [3], one or more candidate primer sequence pairs having a calculated local alignment score that is equal to or less than a predetermined threshold are selected, and further, from all the selected pairs, a candidate primer sequence pair having the smallest maximum value of the local alignment score is detected, and adopted and determined as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site; In the case where one or more primer sequences for different target sites have already been determined in the above (II), In the step [2], all pairs are selected from the one or more candidate primer sequence pairs for the predetermined target site, and for each pair, a local alignment score between each candidate sequence of the selected candidate primer sequence pair and each primer sequence of a different target site already determined, and a local alignment score between the sequences of the selected candidate sequence pair, In the step [3], for each pair, the maximum value is detected from all the calculated local alignment scores, and a candidate primer sequence pair for which a local alignment score is calculated with a value equal to or less than a predetermined threshold is selected, and further, from all the selected pairs, a candidate primer sequence pair having the smallest maximum value of the local alignment score is detected, and these are adopted and determined as a forward primer sequence and a reverse primer sequence for amplifying a region including the predetermined target site. The primer design device for amplicon methylation sequence analysis according to claim 10.
12. a base sequence data acquisition unit for acquiring base sequence data of the genomic double-stranded DNA; A target site information acquisition unit that acquires the two or more target sites and their position information; a base conversion unit that converts "C" that can be methylated in the genomic double-stranded DNA to "Y" and converts other "C" to "T" in the base sequence data; Further comprising: the complementary strand generating unit generates a complementary strand for each template strand of the genomic double-stranded DNA after the base conversion; the partial sequence excision unit selects one of the two or more target sites, and excises, from each of the strands based on positional information of the selected target site, one or more partial sequences of a predetermined length from the base sequence located on the 5'-terminal side of the "Y" into which the selected target site is converted or the "R" complementary thereto; the candidate primer sequence selection unit selects, from one or more partial sequences excised from each of the strands, a sequence that satisfies a predetermined selection condition as a candidate primer sequence; The "C" that can be methylated is "C" in a CG sequence, The predetermined selection conditions are: (1) Tm is within a predetermined range; (2) the number of YG sequences or CR sequences contained in the partial sequence is a predetermined number or less; and (3) The upper limit of the number of junctions with a sequence outside the relevant region on the genomic double-stranded DNA after the base conversion is a predetermined number of 1 or more. Including, The primer design device for amplicon methylation sequence analysis according to claim 10. [Note that "C", "G", "Y" and "R" are base symbols defined by IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine.]
13. The "C" that can be methylated further includes the "C" in the CHG sequence, The primer design device according to claim 12, wherein the predetermined selection conditions further include (4) that the partial sequence contains no more than a predetermined number of YHG sequences or CDR sequences. [Note that "C", "G", "Y", "H", "R" and "D" are base symbols defined by IUPAC, "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, "H" represents adenine, cytosine or thymine, "D" represents thymine, guanine or adenine, and "R" represents adenine or guanine.]
14. Contains the "C" in the CHH sequence, The device for designing a primer according to claim 12, wherein the predetermined selection condition further includes (5) that the partial sequence contains no more than a predetermined number of YHH sequences or DDR sequences. [Note that "Y", "H", "R" and "D" are base symbols defined by IUPAC, "Y" represents thymine or cytosine, "H" represents adenine, cytosine or thymine, "D" represents thymine, guanine or adenine, and "R" represents adenine or guanine.]
15. The primer candidate sequence selection section: the genomic double-stranded DNA after the base conversion is a first template strand and a second template strand, a complementary strand of the first template strand is a first complementary strand, and a complementary strand of the second template strand is a second complementary strand; 13. The primer design device according to claim 12, wherein one or more partial sequences excised from the first template strand that satisfy a predetermined selection condition are selected as candidate forward primer sequences for the first template strand, one or more partial sequences excised from the first complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the first template strand, one or more partial sequences excised from the second template strand that satisfy the predetermined selection condition are selected as candidate forward primer sequences for the second template strand, and one or more partial sequences excised from the second complementary strand that satisfy the predetermined selection condition are selected as candidate reverse primer sequences for the second template strand.
16. 16. The primer design device according to claim 15, wherein the primer sequence determination unit calculates, in the candidate primer sequence selection unit, lengths of PCR amplification products expected to be amplified by PCR for all combinations of the candidate forward primer sequences of the one or more selected first template strands and the candidate reverse primer sequences of the one or more selected first template strands, adopts combinations of candidate primer sequences for which the calculated PCR amplification product lengths are within a predetermined range as forward primer sequences and reverse primer sequences of the first template strand that amplify the region including the target site selected in the partial sequence excision unit, and calculates lengths of PCR amplification products expected to be amplified by PCR for all combinations of the candidate forward primer sequences of the selected second template strand and the candidate reverse primer sequences of the selected second template strand, and adopts combinations of candidate primer sequences for which the calculated PCR amplification product lengths are within the predetermined range as forward primer sequences and reverse primer sequences of the second template strand that amplify the region including the target site selected in the partial sequence excision unit.
17. a storage unit that measures and stores in advance the correspondence relationship between at least the number of target sites, the predetermined threshold value, and the primer design success rate using the primer design device according to claim 10; an input unit for a user to input instructions; Further comprising:
11. The primer design device for amplicon methylation sequence analysis according to claim 10, wherein, when a user sets at least the primer design success rate and the number of target sites desired by the user via an input unit and instructs to execute primer design, the primer sequence determination unit reads out the predetermined threshold value corresponding to the primer design success rate and the number of target sites that is equal to or greater than the set values of the primer design success rate and the number of target sites and has a small difference from the corresponding relationships stored in the storage unit, and adopts and determines a primer sequence that amplifies a region including the predetermined target site from among the one or more candidate primer sequences based on the read predetermined threshold value.
18. Further, the device is provided with a communication interface, 13. The primer design device according to claim 12, wherein the communication interface enables connection to a server via a communication network external to the device, and a program in the server enables execution of at least one of the group consisting of the base sequence data acquisition unit, the target site information acquisition unit, the base conversion unit, the complementary strand generation unit, the partial sequence excision unit, the primer candidate sequence selection unit, and the primer sequence determination unit.
19. A primer design program, which executes the primer design method according to claim 1 on a computer.
20. 20. A computer-readable recording medium having recorded thereon the primer design program according to claim 19.