Method for designing primers for amplicon methylation sequence analysis, manufacturing method, design apparatus, design program, and recording medium

The method and device enhance primer design for bisulfite-treated DNA and multiplex PCR by addressing low success rates and sequence-specific challenges, enabling efficient and detailed DNA methylation analysis in CG, CHG, and CHH sequences.

JP7715736B2Active Publication Date: 2025-07-30FUJIFILM CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022565257
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-26
Filing Date
2021-11-17
Publication Date
2025-07-30
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

Existing primer design software for bisulfite-treated DNA and multiplex PCR has low success rates and lacks functionality for analyzing DNA methylation in CG, CHG, and CHH sequences, requiring manual labor and time for primer design.

Method used

A method and device for designing primers that utilize nucleotide sequence data, target site information, and base conversion to enhance primer design success, considering specific conditions for bisulfite-treated DNA, including CG, CHG, and CHH sequences, using multiplex PCR to amplify multiple target sites efficiently.

Benefits of technology

Improves primer design success rate and enables detailed analysis of DNA methylation states, particularly in CG, CHG, and CHH sequences, facilitating high-throughput amplicon methylation sequence analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007715736000004
    Figure 0007715736000004
  • Figure 0007715736000005
    Figure 0007715736000005
  • Figure 0007715736000006
    Figure 0007715736000006
Patent Text Reader

Abstract

The purpose of the present invention is to provide a method of designing a primer for amplicon methylation sequence analysis by which primer design success rate can be improved, and a production method, a designing device, a designing program and a recording medium therefor. The present invention pertains to a method of designing a primer for amplicon methylation sequence analysis, said method comprising a base conversion step for, in genomic double-stranded DNA, converting "C", which is possibly methylated, into "Y" and converting other "C" into "T", and a primer candidate sequence selection step for selecting a sequence that satisfies preset selection conditions as a primer candidate sequence, wherein the C that is possibly methylated is C in a CG sequence, and the preset selection conditions include (1) having a Tm value within a definite range, (2) the number of YG sequences or CR sequences contained on the partial sequence being not larger than a definite number, and (3) after the base conversion, the upper limit of the number of junctions with sequences outside the relevant region on the genomic double-stranded DNA being not larger than a definite number that is 1 or more.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for designing primers for amplicon methylation sequence analysis, a manufacturing method, a design apparatus, a design program, and a recording medium. In particular, it relates to a method for designing primers, a manufacturing method, a design apparatus, a design program, and a recording medium for simultaneously amplifying a plurality of amplification target regions each containing a plurality of target sites in DNA (deoxyribonucleic acid) that has been bisulfite-treated or enzymatically treated by multiplex PCR (polymerase chain reaction).

Background Art

[0002] DNA methylation is known as one of the epigenetic mechanisms, which is a gene expression control mechanism without changes in the DNA base sequence. Mammalian DNA methylation mainly occurs at the 5th carbon atom of cytosine (C) in the CG sequence on DNA. Regions called CpG islands where CG sequences appear frequently are abundant in gene promoter regions. Initially, most of the CG sequences in these regions are not methylated, but they become methylated with the progression of diseases, development, differentiation, inflammation, or aging, etc., and it is known that gene expression is suppressed. For example, in cancer cells, it is known that the hypermethylation of CpG islands in the gene promoter region occurs, resulting in the inactivation of many genes that suppress cancer.

[0003] Thus, since DNA methylation is greatly involved in gene expression regulation, its information is attracting attention in various fields such as diagnosis, treatment, drug discovery, and regenerative medicine, and is being actively researched and developed as being useful for elucidating the mechanisms of diseases such as cancer and evaluating the differentiation states of various cells. For example, by measuring and analyzing the state of DNA methylation in a specific region, when developing a drug, attempts are made to examine the presence or absence of drug resistance for each cell type, and to evaluate the presence or absence and malignancy (progression degree) of cancer cells from the ratio of normal cells to abnormal cells. Attempts are also being made to evaluate the differentiation state of stem cells and use it for quality control.

[0004] As one of the methods for analyzing the state of DNA methylation, there is a method that utilizes the bisulfite reaction. For example, cytosine (C) in the CG sequence related to a certain disease is picked up and used as the target site (measurement site). In Fig. 12A, [1] to [4] are methylated sites, and from among them, [2] and [4] are set as target sites A and B (Fig. 12A shows only a single strand).

[0005] Subsequently, the template DNA is treated with bisulfite (hydrogen sulfite). When cytosine (C) in the CG sequence on the template DNA is methylated, it remains as cytosine (C) as it is after this treatment (see methylated sites [3] and [4] in Fig. 12A). On the other hand, when cytosine (C) in the CG sequence on the template DNA is not methylated, it is deaminated and converted to uracil (U) (see methylated sites [1] and [2] in Fig. 12A). Recently, instead of bisulfite treatment, for example, methods that perform the same base conversion as the above reaction using enzymes such as the NEB Next Enzymatic Methyl-seq Kit manufactured by New England Biolabs are also being used.

[0006] Subsequently, the DNA after bisulfite treatment is amplified using PCR (polymerase chain reaction) for sequence analysis. The amplified DNA, i.e., the PCR amplification product, is subjected to sequence analysis using a capillary sequencer or an NGS (Next Generation Sequencer). When the DNA after bisulfite treatment is amplified using PCR, cytosine (C) remains as it is (see methylation sites [3] and [4] in Fig. 12A), while uracil (U) is replaced by thymine (T) and amplified (see methylation sites [1] and [2] in Fig. 12A). By utilizing the difference between cytosine (C) and thymine (T) that occurs in the sequence of this PCR amplification product, for example, in the DNA (template DNA) before bisulfite treatment, it is possible to detect the methylation state of a predetermined target site, i.e., whether the DNA of a predetermined target site selected from one cell is methylated or not. More specifically, depending on whether the base at a predetermined target site of the PCR amplification product is cytosine (C) or thymine (T), it is possible to determine whether the cytosine (C) at the predetermined target site of the template DNA was methylated or not. Explained with reference to Fig. 12A, since the base at target site A of the PCR amplification product is thymine (T), it can be seen that the cytosine (C) at target site A of the template DNA was not methylated. On the other hand, since the base of the PCR amplification product at target site B is cytosine (C), it can be seen that the cytosine (C) at target site B of the template DNA was methylated.

[0007] In addition, by utilizing the difference between cytosine (C) and thymine (T) that occurs in the sequence of this PCR amplification product, in the DNA (template DNA) before bisulfite treatment, the methylation state (frequency) of the DNA of a specific target site derived from multiple cells, that is, whether the DNA of a specific target site derived from multiple cells is methylated or not can be detected, and based on the detection result, the proportion of cells in which the DNA of a specific target site is methylated can also be grasped. When there are multiple specific target sites, for each specific target site, it is also possible to detect whether the DNA of that site is methylated or not, and based on the detection result, detect the proportion of cells in which the DNA is methylated for each specific target site. More specifically, the methylation state (frequency) of the DNA of a specific target site can be grasped according to whether the base at a specific target site is cytosine (C) or thymine (T). The methylation state (frequency) of the DNA of a specific target site can be obtained by measuring the methylation degree = C / (C + T) from the number of cytosine (C) and thymine (T) that occur at each target site (measurement site). When there are multiple specific target sites, the proportion of cells in which the DNA is methylated can be grasped for each specific target site.

[0008] For example, as shown in FIG. 12B, when evaluating the methylation state (frequency) of target sites (measurement sites) A and B derived from multiple cells (in the figure, cells 1 to 3), the number of cytosine (C) that occurs at target site A is 2, and the number of thymine (T) is 1. Therefore, when calculating the methylation degree, it is 2 / (2 + 1) = 0.67. Thus, the methylation state (frequency) of the DNA at target site A in FIG. 12B can be regarded as a methylation degree of 0.67 derived from 3 cells, and the proportion of cells in which the DNA is methylated can be grasped. On the other hand, the number of cytosine (C) that occurs at target site B is 3, and the number of thymine (T) is 0. Therefore, when calculating the methylation degree, it is 3 / (3 + 0) = 1. Thus, the methylation state (frequency) of the DNA at target site A in FIG. 12B can be regarded as a methylation degree of 1 derived from 3 cells, and the proportion of cells in which the DNA is methylated can be grasped. Incidentally, similarly, the methylation state (frequency) of target site A shown in Fig. 12A can be detected as the methylation degree of 0 derived from one cell, and the methylation state (frequency) of target site B can be detected as the methylation degree of 1 derived from one cell.

[0009] For the amplification of DNA after bisulfite treatment, multiplex PCR that can amplify two or more amplification target regions on DNA at once may be used in the same reaction. In order to grasp the DNA methylation state of a predetermined target site or the DNA methylation state (frequency) of a specific target site derived from multiple cells using multiplex PCR, as shown in Fig. 12C (Fig. 12C shows only single strands), primer pairs (forward primer and reverse primer) that amplify one or more amplification target regions each containing one or more target sites are required. Explaining with Fig. 12A, a primer pair for amplifying the amplification target region (amplification region) containing target site A and a primer pair for amplifying the amplification target region (amplification region) containing target site B are required.

[0010] The design of primers for DNA subjected to bisulfite treatment requires consideration of the following conditions in addition to the conditions considered in the design of normal primers (i.e., the design of primers for DNA not subjected to bisulfite treatment). First, as a premise, the presence or absence of DNA methylation cannot be known in advance, unlike the base sequence. That is, there are bases whose identity as thymine (T) or cytosine (C) remains undetermined after bisulfite treatment. Therefore, in primer design for the purpose of analyzing the DNA methylation state, in order to prevent the amplification efficiency of the primer from changing depending on the methylation state around the target site, the primer should contain as few CG sequences as possible at the site where it binds, or if it contains CG sequences, the positions of the CG sequences in the primer should be limited to minimize the impact.

[0011] In addition, since many cytosines (C) on DNA are converted to thymines (T) by bisulfite treatment, the DNA sequence of each strand has an increased region composed of three bases other than cytosine (C) after bisulfite treatment. Therefore, it is also necessary to consider that primers specifically binding to the region composed of three bases must be designed. In addition, since many cytosines (C) on DNA are converted to thymines (T) and the double-stranded DNA loses its complementarity, when it is necessary to amplify and analyze both strands of the DNA double-strand, primer pairs (forward primers and reverse primers) that amplify one or more amplification target regions each containing the target site of each strand, that is, two sets of primer pairs need to be designed.

[0012] Therefore, the design of primers for DNA subjected to bisulfite treatment with such specific circumstances has different design conditions and higher difficulty compared to the design of ordinary primers. There are many primer design softwares, but most of them are for designing ordinary primers like Primer-BLAST, so conditions considering cytosine whose bases are converted by bisulfite treatment cannot be set. That is, in ordinary primer design softwares, the specific circumstances related to the design of primers for DNA subjected to bisulfite treatment as described above are not considered at all, so there is a situation where primers for DNA subjected to bisulfite treatment cannot be designed with those softwares.

[0013] In addition, when multiplex PCR is used for the amplification of DNA after bisulfite treatment, in order to simultaneously amplify a plurality of amplification target regions each containing the target site related to the analysis of the methylation degree, it is also necessary to consider designing primers that suppress the formation of primer-dimers.

[0014] Therefore, when bisulfite reaction and multiplex PCR are used to measure the methylation level of DNA at a specific site, the task of designing primers for multiplex PCR (i.e., primers for bisulfite amplicon sequence analysis) used in the analysis is more complicated and time-consuming than designing primers for DNA subjected to bisulfite treatment.

[0015] As described above, most of the primer design software is related to ordinary primer design software, and there is little software related to the design of primers for DNA subjected to bisulfite treatment. In addition, there is even less primer design software corresponding to the design of primers for amplifying DNA subjected to bisulfite treatment by multiplex PCR (i.e., primers for bisulfite amplicon sequence analysis). As one of the few available software, the software described in Non-Patent Document 1 can be mentioned.

Prior Art Documents

Non-Patent Documents

[0016]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0017] In bisulfite amplicon consequence analysis, generally, 5 to 1000 target sites are preset as the measurement targets, but it is desirable to output primer sequences for as many target sites as possible. That is, a high primer design success rate (the number of target sites for which primers can be designed / the number of all target sites [%]) is required.

[0018] However, generally, it is known that the primer design success rate by primer design software for DNA subjected to conventional bisulfite treatment, including the software described in Non-Patent Document 1, is low. Therefore, there is a need for primer design software that can increase the primer design success rate for bisulfite sequencing analysis as much as possible and enable more analysis of the methylation state of DNA at a predetermined site (i.e., measurement of the methylation degree).

[0019] In addition, in the genome of plants, it is known that DNA methylation can occur not only in cytosine (C) of the CG sequence, but also in cytosine (C) of the CHG sequence and the CHH sequence. However, there is no software for multiplex PCR corresponding to these sequences. Therefore, users who also desire analysis of these sequences have to design primers independently while considering all the special circumstances unique to the design of primers for bisulfite sequencing analysis as described above, which requires a great deal of labor and time.

[0020] The present invention has been made to solve such problems, and an object thereof is to provide a method for designing primers for bisulfite amplicon consequence analysis (more specifically, primers for amplicon methylation sequence analysis), a manufacturing method, a design device, a design program, and a recording medium that can improve the primer design success rate. Also, as cytosine (C) that can be methylated, provided are a method for designing primers for bisulfite amplicon sequence analysis corresponding to cytosine (C) in CHG sequences and CHH sequences (more specifically, primers for amplicon methylation sequence analysis), a manufacturing method, a design device, a design program, and a recording medium.

Means for Solving the Problem

[0021] The method for designing primers for amplicon methylation sequence analysis according to the present invention is a method for designing primers used to simultaneously amplify a plurality of regions each containing one or more target sites for measuring the methylation degree by using a bisulfite reaction or an enzymatic reaction and multiplex PCR for measuring the methylation degree of genomic double-stranded DNA at a predetermined site related to a predetermined biological phenomenon, comprising: a nucleotide sequence data acquisition step of acquiring nucleotide sequence data of genomic double-stranded DNA; a target site information acquisition step of acquiring one or more target sites and their position information; a base conversion step of converting "C" that can be methylated in genomic double-stranded DNA to "Y" and converting other "C" to "T" in the nucleotide sequence data; a complementary strand generation step of generating a complementary strand for each template strand of the genomic double-stranded DNA after base conversion; selecting one from among one or more target sites, and based on the position information of the selected target site, cutting out one or more partial sequences of a predetermined length from the nucleotide sequences located on the 5'-terminal side of the converted "Y" of the selected target site or "R" complementary thereto from each strand (partial sequence cutting-out step); a primer candidate sequence selection step of selecting, as primer candidate sequences, those that satisfy predetermined selection conditions from among one or more partial sequences cut out from each strand; a primer sequence determination step of adopting and determining a forward primer sequence and a reverse primer sequence that amplify a region containing the selected target site and are cut out from each template strand from among the one or more selected primer candidate sequences; In the partial array cutting-out step, a repeating step of repeating the partial array cutting-out step, the primer candidate array selection step, and the primer array determination step until all one or more target sites are selected, and having The "C" that can be methylated is the "C" in the CG sequence, The predetermined selection conditions are (1) The Tm value is within a predetermined range, (2) The number of YG sequences or CR sequences contained in the partial array is equal to or less than a predetermined number, and (3) The upper limit of the number of junctions with sequences outside the related region on the genomic double-stranded DNA after base conversion is equal to or less than a predetermined number of 1 or more including [wherein, "C", "G", "Y", and "R" are base notations defined by IUPAC, "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine].

[0022] Here, the "C" that can be methylated further includes the "C" in the CHG sequence, and the predetermined selection conditions can further include (4) the number of YHG sequences or CDR sequences contained in the partial array is equal to or less than a predetermined number [wherein, "C", "G", "Y", "H", "R", and "D" are base notations defined by IUPAC, "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine]. Further, the "C" that can be methylated further includes the "C" in the CHH sequence, and the predetermined selection conditions can further include (5) the number of YHH sequences or DDR sequences contained in the partial array is equal to or less than a predetermined number [wherein, "Y", "H", "R", and "D" are base notations defined by IUPAC, "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine].

[0023] The predetermined selection conditions preferably further include that (6) the three bases from the 3'-end of the partial sequence are not complementary to the three bases from the 3'-end of another partial sequence. The predetermined selection conditions preferably further include that (7) the junction with the genomic double-stranded DNA before base conversion is equal to or less than a predetermined number.

[0024] The predetermined selection conditions preferably further include that (8) in the condition of (2) above, when the predetermined number of YG sequences or CR sequences included in the partial sequence is set to 1 or more, the position range of the YG sequence or CR sequence in the partial sequence is also specified, and within the specified position range, the number of YG sequences or CR sequences included is preferably equal to or less than the predetermined number. The predetermined selection conditions preferably further include that (9) in the condition of (4) above, when the predetermined number of YHG sequences or CDR sequences included in the partial sequence is set to 1 or more, the position range of the YHG sequence or CDR sequence in the partial sequence is also specified, and within the specified position range, the number of YHG sequences or CDR sequences included is preferably equal to or less than the predetermined number. The predetermined selection conditions preferably further include that (10) in the condition of (5) above, when the predetermined number of YHH sequences or DDR sequences included in the partial sequence is set to 1 or more, the position range of the YHH sequence or DDR sequence in the partial sequence is also specified, and within the specified position range, the number of YHH sequences or DDR sequences included is preferably equal to or less than the predetermined number.

[0025] The primer candidate sequence selection step is Using the genomic double-stranded DNA after the base conversion as the first template strand and the second template strand, taking the complementary strand of the first template strand as the first complementary strand and the complementary strand of the second template strand as the second complementary strand, if one or more partial sequences excised from the first template strand meet the predetermined selection criteria, they are selected as the forward primer candidate sequences for the first template strand; if one or more partial sequences excised from the first complementary strand meet the predetermined selection criteria, they are selected as the reverse primer candidate sequences for the first template strand; if one or more partial sequences excised from the second template strand meet the predetermined selection criteria, they are selected as the forward primer candidate sequences for the second template strand; and if one or more partial sequences excised from the second complementary strand meet the predetermined selection criteria, they are selected as the reverse primer candidate sequences for the second template strand. In the primer sequence determination step, in the primer candidate sequence selection step, for all combinations of the selected one or more forward primer candidate sequences of the first template strand and the selected one or more reverse primer candidate sequences of the first template strand, calculate the length of the PCR amplification product expected to be amplified by PCR. Combinations of primer candidate sequences for which the calculated length of the PCR amplification product is within a predetermined range are adopted as the forward primer sequence and the reverse primer sequence of the first template strand for amplifying the region containing the target site selected in the partial sequence excision step. For all combinations of the selected forward primer candidate sequences of the second template strand and the selected reverse primer candidate sequences of the second template strand, calculate the length of the PCR amplification product expected to be amplified by PCR. Combinations of primer candidate sequences for which the calculated length of the PCR amplification product is within the predetermined range are adopted as the forward primer sequence and the reverse primer sequence of the second template strand for amplifying the region containing the target site selected in the partial sequence excision step to determine.

[0026] For all target sites, after adopting the forward primer sequence and the reverse primer sequence, the primer sequence determination step further preferably calculates the local alignment score for all combinations of the adopted primer sequences, and adopts and determines those with a calculated local alignment score lower than a predetermined threshold as the primer sequences. Under the conditions of (3) above, the upper limit of the number of junctions with the partial sequence is preferably 1 or 2.

[0027] The method for manufacturing primers for amplicon methylation sequence analysis in the present invention includes a primer design step and a synthesis step of synthesizing primers based on the primer sequences designed in the primer design step, and the primer design step is implemented by the above-described primer design method for amplicon methylation sequence analysis.

[0028] The primer design device for amplicon methylation sequence analysis in the present invention is a device for designing primers used to simultaneously amplify a plurality of regions each containing one or more target sites for measuring the methylation degree of genomic double-stranded DNA at a predetermined site related to a predetermined biological phenomenon, using bisulfite reaction or enzymatic reaction, and multiplex PCR, and a base sequence data acquisition unit for acquiring the base sequence data of the genomic double-stranded DNA, a target site information acquisition unit for acquiring the one or more target sites and their position information, a base conversion unit that converts the "C" that can be methylated in the genomic double-stranded DNA to "Y" and the other "C" to "T" in the base sequence data, a complementary strand generation unit for generating a complementary strand for each template strand of the genomic double-stranded DNA after the base conversion, Select one from among the one or more target sites, and based on the position information of the selected target site, from each of the strands, cut out one or more partial sequences of a predetermined length from among the base sequences located on the 5'-end side of the "Y" into which the selected target site has been converted or the complementary "R" thereof by a partial sequence cutting unit. A primer candidate sequence selection unit that selects, as primer candidate sequences, those that satisfy a predetermined selection condition from among the one or more partial sequences cut out from each of the strands. A primer sequence determination unit that adopts and determines, from among the one or more selected primer candidate sequences, a forward primer sequence and a reverse primer sequence that are cut out from each template strand and amplify the region including the selected target site. A control unit that controls the partial sequence cutting unit, the primer candidate sequence selection unit, and the primer sequence determination unit to repeat each process until all of the one or more target sites are selected in the partial sequence cutting unit. having The "C" that can be methylated is the "C" in the CG sequence. The predetermined selection conditions are (1) The Tm is within a predetermined range. (2) The number of YG sequences or CR sequences included in the partial sequence is equal to or less than a predetermined number, and (3) The upper limit of the number of junctions with sequences outside the relevant region on the genomic double-stranded DNA after the base conversion is equal to or less than one or more predetermined numbers. including [wherein, "C", "G", "Y", and "R" are base notations defined by IUPAC, "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine].

[0029] Here, the "C" that can be methylated further includes the "C" in the CHG sequence, and the predetermined selection conditions can further include that (4) the number of YHG sequences or CDR sequences contained on the partial sequence is equal to or less than a predetermined number [wherein, "C", "G", "Y", "H", "R", and "D" are base notations defined by the IUPAC, "C" is cytosine, "G" is guanine, "Y" is thymine or cytosine, "H" is adenine, cytosine, or thymine, "D" is thymine, guanine, or adenine, and "R" represents adenine or guanine]. In addition, the "C" that can be methylated further includes the "C" in the CHH sequence, and the predetermined selection conditions can further include that (5) the number of YHH sequences or DDR sequences contained on the partial sequence is equal to or less than a predetermined number [wherein, "Y", "H", "R", and "D" are base notations defined by the IUPAC, "Y" is thymine or cytosine, "H" is adenine, cytosine, or thymine, "D" is thymine, guanine, or adenine, and "R" represents adenine or guanine].

[0030] The predetermined selection conditions preferably further include that (6) the three bases from the 3'-end of the partial sequence are not complementary to the three bases from the 3'-end of another partial sequence. The predetermined selection conditions preferably further include that (7) the junction with the genomic double-stranded DNA before base conversion is equal to or less than a predetermined number.

[0031] The predetermined selection conditions preferably further include that (8) in the condition of (2) above, when the predetermined number of YG sequences or CR sequences contained on the partial sequence is set to 1 or more, the position range of the YG sequence or CR sequence in the partial sequence is also specified, and within the specified position range, the number of YG sequences or CR sequences is equal to or less than a predetermined number. The predetermined selection conditions preferably further include that (9) in the condition of (4) above, when the predetermined number of YHG sequences or CDR sequences contained on the partial sequence is set to 1 or more, the position range of the YHG sequence or CDR sequence in the partial sequence is also specified, and within the specified position range, the number of YHG sequences or CDR sequences is equal to or less than a predetermined number. The predetermined selection conditions further include the following: (10) When, in the condition of (5) above, a predetermined number of YHH sequences or DDR sequences included in the partial sequence is set to 1 or more, the position range of the YHH sequence or DDR sequence in the partial sequence is also specified, and it is preferable that the number of YHH sequences or DDR sequences included within the specified position range is equal to or less than the predetermined number.

[0032] The primer candidate sequence selection step is using the double-stranded genomic DNA after the base conversion as the first template strand and the second template strand, the complementary strand of the first template strand as the first complementary strand, and the complementary strand of the second template strand as the second complementary strand. For one or more partial sequences excised from the first template strand that satisfy the predetermined selection conditions, they are selected as forward primer candidate sequences for the first template strand; for one or more partial sequences excised from the first complementary strand that satisfy the predetermined selection conditions, they are selected as reverse primer candidate sequences for the first template strand; for one or more partial sequences excised from the second template strand that satisfy the predetermined selection conditions, they are selected as forward primer candidate sequences for the second template strand; and for one or more partial sequences excised from the second complementary strand that satisfy the predetermined selection conditions, they are selected as reverse primer candidate sequences for the second template strand. In the primer sequence determination step, in the primer candidate sequence selection step, for all combinations of the selected forward primer candidate sequences of the one or more first template strands and the selected reverse primer candidate sequences of the one or more first template strands, the length of the PCR amplification product expected to be amplified by PCR is calculated, and a combination of primer candidate sequences for which the calculated length of the PCR amplification product is within a predetermined range is used as the forward primer sequence and the reverse primer sequence of the first template strand that amplifies the region including the target site selected in the partial sequence excision step. For all combinations of the selected forward primer candidate sequences of the second template strand and the selected reverse primer candidate sequences of the second template strand, the length of the PCR amplification product expected to be amplified by PCR is calculated, and a combination of primer candidate sequences for which the calculated length of the PCR amplification product is within the predetermined range is used as the forward primer sequence and the reverse primer sequence of the second template strand that amplifies the region including the target site selected in the partial sequence excision step to determine.

[0033] For all target sites, after adopting the forward primer sequence and the reverse primer sequence, in the primer sequence determination step, further, for all combinations of the adopted primer sequences, a local alignment score is calculated, and it is preferable to adopt and determine as the primer sequence those for which the calculated local alignment score is lower than a predetermined threshold. Under the conditions of (3) above, the upper limit of the number of junctions with the partial sequence is preferably 1 or 2.

[0034] The primer design device for amplicon methylation sequence analysis further includes a communication interface, through which it can be connected to a server via a communication line network outside the device, and a program in the server can execute at least one of the group consisting of the nucleotide sequence data acquisition unit, the target site information acquisition unit, the nucleotide conversion unit, the complementary strand generation unit, the partial sequence extraction unit, the primer candidate sequence selection unit, and the primer sequence determination unit.

[0035] The primer design program for amplicon methylation sequence analysis in the present invention can execute the above-described primer design method on a computer. The computer-readable recording medium in the present invention is one in which the above-described primer design program for amplicon methylation sequence analysis is recorded.

Advantages of the Invention

[0036] According to the present invention, it is possible to improve the design success rate of primers for bisulfite amplicon sequence analysis (more specifically, primers for amplicon methylation sequence analysis). In addition, primers based on the design of the present invention can be obtained. As a result, many target sites can be amplified and measured. According to the present invention, furthermore, it is possible to easily and in a short time design primers for bisulfite amplicon sequence analysis (more specifically, primers for amplicon methylation sequence analysis) that also correspond to CHG sequences and CHH sequences. In addition, primers based on the design can be obtained. As a result, analysis related to these sequences becomes possible, and thus the methylation state (degree of methylation) of DNA can be analyzed in more detail.

Brief Description of the Drawings

[0037]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 12C

Mode for Carrying Out the Invention

[0038] Hereinafter, based on the public embodiments shown in the accompanying drawings, a method for designing a primer for bisulfite amplicon sequencing (primer for amplicon methylation sequence analysis) of the present invention, a manufacturing method, a design device, a design program, and a recording medium will be described in detail.

[0039] (Explanation of Terms) In this specification, the "primer for bisulfite amplicon sequencing analysis" means an analysis primer for simultaneously amplifying a plurality of amplification target regions each containing a plurality of target sites in DNA subjected to bisulfite treatment by multiplex PCR. The "primer for amplicon methylation sequence analysis" means an analysis primer for simultaneously amplifying a plurality of amplification target regions each containing a plurality of target sites in DNA subjected to bisulfite treatment or enzyme treatment by multiplex PCR. The "amplification target region" means a region amplified by a primer pair. The "methylation site" means a site that can be methylated. The "target site" is a "methylation site" and means a site (measurement site) for measuring the degree of methylation. Base sequences such as the "GC sequence" and the "YG sequence" all mean sequences read from the 5'-end side. The range represented using "~" shall include both sides of "~". For example, the range represented as "A~B" includes A and B.

[0040] [Embodiment 1] FIG. 1 is a block diagram conceptually showing an example of a primer design device according to Embodiment 1 of the present invention. Further, FIG. 2 conceptually shows a flowchart of an example of a primer design method implemented by the primer design device shown in FIG. 1. Further, FIGS. 3A to 3D show schematic diagrams for explaining each step of the primer design method.

[0041] As shown in FIG. 1, the primer design device 10 includes an input unit 12, a storage unit 14, an output unit 16, and a primer design processing unit 18. The input unit 12, the storage unit 14, the output unit 16, and the primer design processing unit 18 are connected to each other.

[0042] The input unit 12 acquires information input by the user, various setting instructions, selection instructions, input instructions, creation instructions, etc., and is constituted by, for example, input devices such as a keyboard and a mouse. The storage unit 14 stores the operation program of the primer design device, and can also temporarily store information and data necessary for executing the primer design process. As the storage unit 14, for example, recording media such as HDD (Hard Disc Drive), SSD (Solid State Drive), FD (Flexible Disc), MO disc (Magneto-Optical disc), MT (Magnetic Tape), RAM (Random Access Memory), CD (Compact Disc), DVD (Digital Versatile Disc), SD card (Secure Digital card), USB memory (Universal Serial Bus memory), etc. can be used. The output unit 16 outputs the DNA base sequence information, instructions, design conditions, primer sequence information designed by the primer design processing unit 18, etc. input from the input unit 12, and is composed of, for example, a display unit such as a liquid crystal display (LCD), an organic light emitting diode (OLED), a flat panel display, an individual display, and a cathode ray tube (CRT), and various forms of printers, etc.

[0043] The primer design processing unit 18 performs a series of processes for primer design. The primer design processing unit 18 includes a base sequence data acquisition unit 20, a target site information acquisition unit 22, a base conversion unit 24, a complementary strand generation unit 26, a partial sequence cutting unit 28, a primer candidate sequence selection unit 30, a primer sequence determination unit 32, and a control unit 34. The primer design processing unit 18 can be configured by a processor including a central processing unit (CPU), etc., and a computer, etc. Also, as shown in FIG. 2, the primer design method 10 includes a base sequence data acquisition step S10, a target site information acquisition step S12, a base conversion step S14, a complementary strand generation step S16, a partial sequence extraction step S18, a primer candidate sequence selection step S20, a primer sequence determination step S22, and a determination step S24 that determines whether all target sites have been selected, and includes a repetition step of repeating the partial sequence extraction step S18, the primer candidate sequence selection step S20, and the primer sequence determination step S22 until all target sites are detected.

[0044] (Base sequence data acquisition unit) The base sequence data acquisition unit 20 shown in FIG. 1 is a part that performs the base sequence data acquisition step S10 shown in FIG. 2, and acquires data of the double-stranded DNA sequence (reference sequence) of the genome of the species for which primer design is to be performed via the input unit 12. If data of the reference sequence is stored in the storage unit 14 in advance, it may be acquired from the storage unit 14. Here, the data of the genomic double-stranded DNA sequence to be acquired is preferably data of the entire sequence of the genome of the species for which primer design is to be performed. For the purpose of explaining the primer design method of the present embodiment, the double-stranded DNA of the double-stranded DNA sequence data acquired in this step is referred to as template DNA, and is respectively referred to as strand A and strand B (see FIG. 3A). The base sequence data acquisition unit 20 is configured by a computer and performs the function of acquiring the data of the double-stranded DNA sequence of the genome described above.

[0045] (Target site information acquisition unit) The target site information acquisition unit 22 shown in FIG. 1 is a part that performs the target site information acquisition step S12 shown in FIG. 2, and can acquire one or more target sites included in the genomic double-stranded DNA acquired by the base sequence data acquisition unit 20 and their position information via the input unit 12. If the target sites and their position information are stored in the storage unit 14 in advance, it may be acquired from the storage unit 14. Here, the "target site" is a site related to a predetermined biological phenomenon, which is the cytosine (C) of the CG sequence that can be methylated and is the site for measuring the degree of methylation. The number of selected target sites is not particularly limited, but from the viewpoint of significantly obtaining the desired effects of the present invention, it is preferable to select 5 to 1000 sites. The position of each target site can be indicated by a chromosome, genomic coordinates, etc. The target site information acquisition unit 22 is configured by a computer and functions to acquire one or more target sites included in the above-described genomic double-stranded DNA and their position information.

[0046] (Base conversion unit) The base conversion unit 24 is the part that performs the base conversion step S14 shown in FIG. 2. As shown in FIGS. 3A and 3B, the cytosine (C) of the CG sequence on the template DNA acquired from the base sequence data acquisition unit 20 is converted to "Y" (refer to the bases of the arrows in FIGS. 3A and 3B), and the cytosine (C) of other sequences is converted to thymine (T). Since the cytosine (C) of the CG sequence of DNA may or may not be methylated, it is converted to "Y" including both the possibility of being converted to thymine (T) and the possibility of remaining as cytosine (C). Note that this conversion process is a process that virtually reproduces, on a computer, the generation of DNA amplified using PCR after bisulfite treatment. The base conversion unit 24 is configured by a computer and functions to convert the cytosine (C) of the CG sequence on the above-described template DNA to "Y" and the cytosine (C) of other sequences to thymine (T).

[0047] As described above, through bisulfite treatment, the DNA double strand loses its complementarity. This is because through bisulfite treatment, cytosine (C) in complementary CG base pairs is converted to thymine (T), resulting in the loss of base pair complementarity (refer to the bold bases in FIGS. 3A and 3B). For the amplification target region on the DNA after bisulfite treatment where the complementarity is lost in this way, both strands cannot be amplified in the same manner with a single set of primers. Therefore, when analyzing the methylation state of double-stranded DNA, it is necessary to prepare primer pairs (forward primers and reverse primers) that amplify the amplification target regions each containing each target site of each strand for each target site. That is, it is necessary to design primer pairs for the amplification target region containing the target site of the A strand after base conversion in FIG. 3B and primer pairs for the amplification target region containing the target site of the B strand after the base sequence respectively. However, as will be described in Modification Example 7 below, when only the A strand or only the B strand is desired to be analyzed, or when it is sufficient to analyze either the A strand or the B strand, it is not always necessary to design two sets of primer pairs.

[0048] (Complementary strand generation unit) The complementary strand generation unit 26 is the part that performs the complementary strand generation step S16 shown in FIG. 2, and generates complementary strands for each of the DNA double strands after the base conversion treatment. Here, for the purpose of explaining the primer design method of the present embodiment, the A strand after base conversion and the B strand after base conversion are referred to as the first template strand (A+ strand) and the second template strand (B+ strand), and the complementary strand of the first template strand is referred to as the first complementary strand (A- strand), and the complementary strand of the second template strand is referred to as the second complementary strand (B- strand) (refer to FIG. 3C). As shown in FIG. 3C, the complementary strand A- is prepared by generating a sequence complementary to the base sequence of the A+ strand, and the complementary strand B- is prepared by generating a sequence complementary to the base sequence of the B+ strand. Note that the base complementary to "Y" is "R" which includes both the possibility of adenine (A) and guanine (G). The complementary strand generation unit 26 is configured by a computer and functions to generate the above-described complementary strands for each of the DNA double strands after the base conversion process.

[0049] As a result, the first template strand (A+ strand) is composed of three bases, thymine (T), adenine (A), and guanine (G), excluding "Y" (i.e., the methylated site), and the first complementary strand (A- strand) is composed of three bases, thymine (T), adenine (A), and cytosine (C), excluding "R" (the methylated site). However, the first template strand (A+ strand) and the first complementary strand (A- strand) can have complementarity. Similarly, the second template strand (B+ strand) is also composed of three bases, thymine (T), adenine (A), and guanine (G), excluding "Y" (the methylated site), and the second complementary strand (B- strand) is composed of three bases, thymine (T), adenine (A), and cytosine (C), excluding "R" (the methylated site). However, the second template strand (B+ strand) and the second complementary strand (B- strand) can have complementarity.

[0050] (Subsequence excision unit) The subsequence excision unit 28 is the part that performs the subsequence excision step S18 shown in FIG. 2. As shown in the flowchart of FIG. 4, one target site is selected from the one or more target sites acquired by the target site information acquisition unit 22 (step S280). Based on the position information of the selected target site, from the DNA sequences of each strand, "Y" of the selected target site or "R" complementary thereto (i.e., the base at the target site and in the methylated site) is detected, and as much as possible is excised from the base sequences located on the 5'-end side of the detected "Y" and "R" ( (1) to (4) in FIG. 3D) to obtain one or more subsequences. Note that FIG. 4 is a flowchart showing an example of the operations of the subsequence excision unit 28, the primer candidate sequence selection unit 30, and the primer sequence determination unit 32. The partial sequence extraction unit 28 is configured by a computer and, based on the position information of the selected target site described above, cuts out as much as possible a partial sequence of a predetermined length from the "Y" of the selected target site or "R" complementary thereto from the DNA sequences of each strand, and functions to obtain one or more partial sequences.

[0051] Here, the length of the one or more partial sequences to be cut out is not particularly limited, but from the viewpoint of processing efficiency and significantly obtaining the desired effects of the present invention, it is preferably the length obtained by subtracting the length of the target site (one base) from the difference between the maximum length of the PCR amplification product desired by the user and the minimum length of the primer. The length of the PCR amplification product is not particularly limited as long as it is within a known range, that is, 70 to several kilobases. It is preferable to consider the success rate of PCR and the sequence decoding ability of the DNA sequencer, etc. The length of the primer is not particularly limited as long as it is within a known range, that is, 15 to 45 bases. It is preferable to consider the specificity of the primer and the formation of primer-dimers.

[0052] For example, when the maximum length of the PCR product set by the user is 300 bases and the minimum length of the primer is 20 bases, the predetermined length x to be cut out is calculated as x = 300 - 20 - 1 (the length of the target site) = 279, and first, 279 bases on the 5'-end side of each target site are cut out. As shown in FIG. 3D, 279 bases ( (1) to (4) in FIG. 3D) on the 5'-end side of the target sites of each strand, that is, "Y" of the A+ strand, "R" of the A- strand, "Y" of the B+ strand, and "R" of the B- strand, are cut out from each strand, respectively. Subsequently, one or more partial sequences can be obtained by cutting out as much as possible partial sequences with the length of the primer (20 bases or more and a predetermined length or less) from the 279 bases.

[0053] Note that the numerical values or numerical ranges of the length of the PCR amplification product and the length of the primer are set by the user via the input unit 12. If these conditions are stored in the storage unit 14 in advance, they can also be obtained from the storage unit 14 and set.

[0054] (Primer candidate sequence selection unit) The primer candidate sequence selection unit 30 is a part that performs the primer candidate sequence selection step S20 shown in FIG. 2, and selects, as primer candidate sequences, those that satisfy all of the predetermined selection conditions (1) to (3) from one or more partial sequences of each strand cut out by the partial sequence cutting unit 28. Specifically, for one or more partial sequences cut out from the first template strand (A+ strand) (i.e., one or more partial sequences cut out from (1) in FIG. 3D) that satisfy the predetermined selection conditions, they are selected as forward primer candidate sequences for the first template strand (A+ strand). For one or more partial sequences cut out from the first complementary strand (A- strand) (i.e., one or more partial sequences cut out from (2) in FIG. 3D) that satisfy the predetermined selection conditions, they are selected as reverse primer candidate sequences for the first template strand (A+ strand). For one or more partial sequences cut out from the second template strand (B+ strand) (i.e., one or more partial sequences cut out from (3) in FIG. 3D) that satisfy the predetermined selection conditions, they are selected as forward primer candidate sequences for the second template strand (B+ strand). For one or more partial sequences cut out from the second complementary strand (B- strand) (i.e., one or more partial sequences cut out from (4) in FIG. 3D) that satisfy the predetermined selection conditions, they are selected as reverse primer candidate sequences for the second template strand (B+ strand). The primer candidate sequence selection unit 30 is configured by a computer and functions to select, as primer candidate sequences, those that satisfy all of the predetermined selection conditions (1) to (3) from one or more partial sequences of each of the above-mentioned strands.

[0055] The "predetermined selection conditions" for the primer candidate sequences are the following items (1) to (3). The numerical values and numerical ranges of the predetermined selection conditions can be preset by the user via the input unit 12 in advance. (1) The Tm value is within a predetermined range. (2) The number of YG sequences or CR sequences contained in the partial sequence is equal to or less than a predetermined number. (3) The upper limit of the number of junctions with the base sequence outside the related region on the template strand DNA (genomic double-stranded DNA) after base conversion is a predetermined number of 1 or more and a certain number or less.

[0056] Regarding the condition of (1) above, the range of the "Tm value" is not particularly limited as long as it is within a known numerical range, that is, 45 to 70°C. It is preferable to consider the thermal cycle conditions of PCR, the ease of PCR amplification (the temperature range in which amplification proceeds easily depending on the PCR enzyme used), and the specificity of PCR amplification. The Tm value can be calculated, for example, by the nearest neighbor base pair method.

[0057] Regarding the condition of (2) above, the number of "YG sequences or CR sequences contained on the partial sequence" is not particularly limited, but from the viewpoint of significantly obtaining the desired effect of the present invention, it is preferably 2 or less, more preferably 1 or less, and particularly preferably 0. By satisfying this condition, the influence of the primer on the junction with cytosine (C) of the CG sequence at the primer junction site can be reduced.

[0058] Regarding the "sequence outside the related region on the template strand DNA (genomic double-stranded DNA) after base conversion" according to the condition of (3) above, it refers to the base sequence excluding the sequence at the position on the template strand DNA after base conversion corresponding to the position of the partial sequence, and the base sequence complementary to the sequence excluding the partial sequence (the template strand DNA sequence after base conversion). The "upper limit of the number of junctions with the sequence outside the related region on the template strand DNA after base conversion" is not particularly limited, but from the viewpoint of significantly obtaining the desired effect of the present invention, it is preferably 5 or less, and particularly preferably 2 or less. By satisfying this condition, the influence of the primer on the junction outside the related region on the DNA after bisulfite treatment can be reduced.

[0059] Regarding PCR, when the number of heating cycles is n, as shown in FIG. 5A, when a primer pair (forward primer and reverse primer) is joined to DNA, 2

[0059] , , , nPCR amplification products are generated in the order of n, but as shown in Fig. 5B, when either one of the forward primer and the reverse primer binds to DNA, PCR amplification products are generated in the order of 2n (Fig. 5B shows the case where the forward primer binds). Therefore, when PCR is performed with a general number of heating cycles (n is about 20 to 40), if the primer pair binds to a DNA sequence outside the amplification target region in the DNA, a large amount of non-specific products will be generated, which becomes a problem. However, when either one of the forward primer and the reverse primer binds to a DNA sequence outside the relevant region, not so many non-specific products are generated and it is not particularly problematic. Therefore, conventionally, the problem of non-specific product generation when either one of the forward primer and the reverse primer binds to a DNA sequence outside the relevant region has not been particularly considered. Note that (1) in Fig. 5A shows the DNA sequence of the amplification target region, and (2) shows the DNA sequence outside the amplification target region. Also, (3) in Fig. 5B shows the DNA sequence of the relevant region of the partial sequence, and (4) shows the DNA sequence outside the relevant region. Therefore, as a result of intensive research this time, the inventors have found that by adding the condition of (3) above, which allows the binding of each primer to DNA outside the target region within a predetermined range to the conventional conditions determined in primer design, the primer design success rate can be increased.

[0060] Here, the process of selecting, as primer candidate sequences, those that satisfy predetermined selection conditions from among one or more partial sequences excised from each strand will be described using the flowchart of Fig. 4.

[0061] The primer candidate sequence selection unit 30 first obtains one (step S300) from among one or more partial sequences excised from the first template strand (A+ strand), and determines whether the Tm value of the partial sequence is within a predetermined range (step S302). If the Tm value is not within a predetermined range, another partial sequence is obtained (step S300). If the Tm value is within the predetermined range, it is determined whether the number of YG sequences or CR sequences included in the partial sequence is equal to or less than a predetermined number (step S304). If the number of YG sequences or CR sequences included in the partial sequence is not equal to or less than a predetermined number, another partial sequence is obtained (step S300). If the number of YG sequences or CR sequences included in the partial sequence is equal to or less than a predetermined number, it is determined whether the upper limit of the number of junctions with the base sequences outside the relevant region on the template strand DNA after base conversion is equal to or less than a predetermined number of 1 or more (step S306). If the upper limit of the number of junctions between the base sequence outside the relevant region on the template strand DNA after base conversion and the partial sequence is not equal to or less than "a predetermined number of 1 or more", another partial sequence is obtained (step S300). If the upper limit of the number of junctions with the sequence outside the relevant region on the template strand DNA after base conversion is equal to or less than a predetermined number of 1 or more, it is selected as a primer candidate sequence (step S308), and it is determined whether the determination of all partial sequences cut out from the first template strand (A+ strand) has been completed (step S310). If the determination of all partial sequences cut out from the first template strand (A+ strand) has not been completed, another partial sequence is obtained (step S300). If the determination of all partial sequences has been completed, the selected one or more primer candidate sequences are determined as the forward primer candidate sequences of the first template strand (A+ strand) (step S312).

[0062] The same determination is performed for one or more partial sequences cut out from the first complementary strand (A- strand), one or more partial sequences cut out from the second template strand (B+ strand), and one or more partial sequences cut out from the second complementary strand (B- strand) (steps S300 to S310), and the reverse primer candidate sequences of the first template strand (A+ strand), the forward primer candidate sequences of the second template strand (B+ strand), and the reverse primer candidate sequences of the second template strand (B+ strand) are determined (step S312).

[0063] (Primer sequence determination unit) The primer sequence determination unit 32 is a part that performs the primer sequence determination step S22 shown in FIG. 2, and adopts and determines a forward primer sequence and a reverse primer sequence that amplify a region including a target site selected by the partial sequence cutting unit 28 from one or more primer candidate sequences determined by the primer candidate sequence selection unit 30, that is, one or more forward primer candidate sequences of the first template strand (A+ strand), one or more reverse primer candidate sequences of the first template strand (A+ strand), one or more forward primer candidate sequences of the second template strand (B+ strand), and one or more reverse primer candidate sequences of the second template strand (B+ strand). The primer sequence determination unit 32 is configured by a computer and functions to adopt and determine a forward primer sequence and a reverse primer sequence from the one or more primer candidate sequences described above.

[0064] Specifically, first, all primer pairs (combinations of a forward primer and a reverse primer) that can be prepared from one or more forward primer candidate sequences of the first template strand and one or more reverse primer candidate sequences of the first template strand are obtained, and for each primer pair, the length of the PCR amplification product expected to be amplified by PCR is calculated. Next, it is determined whether the calculated length of the PCR amplification product is within a predetermined numerical range. If the calculated length of the PCR amplification product is within the predetermined numerical range, the primer pair for which the length of the PCR amplification product is calculated (that is, the combination of the forward primer candidate sequence of the first template strand and the reverse primer candidate sequence of the first template strand) is adopted (step S320) as the forward primer sequence and the reverse primer sequence of the first template strand that amplify the region including the target site selected by the partial sequence cutting unit 28 (partial sequence cutting step S18), and determined (step S322). Similarly, first, all primer pairs (combinations of a forward primer and a reverse primer) that can be prepared from one or more forward primer candidate sequences of the second template strand and one or more reverse primer candidate sequences of the second template strand are obtained, and for each primer pair, the length of the PCR amplification product expected to be amplified by PCR is calculated. Next, it is determined whether the length of the calculated PCR amplification product is within a predetermined numerical range (step S320). If the length of the calculated PCR amplification product is within the predetermined numerical range, the primer pair for which the length of the PCR amplification product has been calculated (that is, the combination of the forward primer candidate sequence of the second template strand and the reverse primer candidate sequence of the second template strand) is adopted and determined as the forward primer sequence and the reverse primer sequence of the second template strand that amplify the region including the target site selected in the subsequence cutting unit 28 (subsequence cutting step S18) (step S322). Here, the "predetermined numerical range" for determining the length of the calculated PCR amplification product is a range including the length of the PCR amplification product desired by the user, and as described above, it is not particularly limited as long as it is a known range, that is, 70 to several kilobases. It is preferable to consider the success rate of PCR and the sequence decoding ability of the DNA sequencer, etc.

[0065] After determining whether the length of the PCR amplification product is within the predetermined range for all primer pairs, it is determined in the subsequence cutting unit 28 (subsequence cutting step S18) whether all target sites have been selected (S24). If not all target sites have been selected, the process returns to the subsequence cutting step S18 to select another target site (step S280), and if all target sites have been selected, the process ends.

[0066] (Control unit) The control unit 34 is directly or indirectly connected not only to each part within the primer design processing unit 18, but also to the input unit 12, the storage unit 14, and the output unit 16. Based on the user's instructions from the input unit 12 or based on a predetermined operation program stored in the storage unit 14, etc., it controls each part of the primer design apparatus 10 to design a primer, and is composed of, for example, a CPU (Central Processing Unit) such as a computer. The control unit 34 controls the primer candidate array selection unit 30 to repeat the determination operation (steps S300 to S308) until the determination as to whether all partial arrays satisfy all predetermined selection criteria is completed (step S310). The control unit 34 controls the primer array determination unit 32 to repeat the determination operation (step 320) until the determination as to whether the length of the PCR amplification product is within a predetermined range is completed for all the produced primer pairs. The control unit 34 controls the partial sequence cutting-out unit 26, the primer candidate array selection unit 30, and the primer array determination unit 32 so that the partial sequence cutting-out process (steps S18, S280 to S282), the primer candidate array selection process (steps S20, S300 to S312), and the primer array determination process (steps S22, S320 to S322) are repeated until all the target sites acquired by the target site information acquisition unit 22 are detected (step S24).

[0067] According to the primer design apparatus 10 of Embodiment 1 of the present invention, a primer for amplicon methylation sequence analysis can be designed with an excellent design success rate. Also, a primer based on the design can be obtained. As a result, primers can be designed for more target sites and the degree of methylation can be measured.

[0068] [Modification Example 1] Next, a primer design apparatus according to Modification Example 1 of Embodiment 1 of the present invention will be described. Note that, in the primer design apparatus according to this Modification Example 1, the description of the same processing as in Embodiment 1 will be omitted. In Embodiment 1, cytosine (C) that can be methylated was limited to only cytosine (C) in the CG sequence, and the cytosine (C) picked up from among them was used as the target site. However, the present invention is not limited to this. Further, cytosine (C) in the CHG sequence may be included in cytosine (C) that can be methylated, and cytosine (C) picked up from among them may be used as the target site.

[0069] In this Modification 1, the target site information acquisition unit 20 further acquires, via the input unit 12, one or more target sites included in the genomic double-stranded DNA acquired by the base sequence data acquisition unit 20 and their position information. The base conversion unit 24 further converts cytosine (C) in the CHG sequence on the template DNA acquired from the base sequence data acquisition unit 20 to "Y", and cytosine (C) in other sequences (that is, sequences other than the CG sequence and the CHG sequence) is converted to thymine (T).

[0070] The primer candidate sequence selection unit 30 selects, from one or more partial sequences of each strand cut out by the partial sequence cutting unit 28, those that satisfy all of the following predetermined selection conditions (1) to (4) including the following item (4) as primer candidate sequences. (4) The number of YHG sequences or CDR sequences contained in the partial sequence is equal to or less than a predetermined number Here, the number of "YHG sequences or CDR sequences contained in the partial sequence" according to the condition of (4) above is not particularly limited. However, from the viewpoint of significantly obtaining the desired effects of the present invention, it is preferably 2 or less, more preferably 1 or less, and particularly preferably 0. By satisfying this condition, the influence of the primer on the junction with cytosine (C) in the CHG sequence at the primer junction site can be reduced.

[0071] With the primer design device of Modification Example 1 of Embodiment 1 of the present invention, primers for amplicon methylation sequence analysis corresponding to the CHG sequence can be easily designed in a short time. In addition, primers based on the design can be obtained. As a result, analysis related to these sequences becomes possible, so that the methylation state (degree of methylation) of DNA can be analyzed in more detail.

[0072] [Modification Example 2] Next, a primer design device according to Modification Example 2 of Embodiment 1 of the present invention will be described. In the primer design device according to this Modification Example 2, the description of the same processing as in Embodiment 1 will be omitted. In Embodiment 1, cytosine (C) that can be methylated is only cytosine (C) in the CG sequence, and the cytosine (C) picked up from among them is used as the target site, but the present invention is not limited to this. Furthermore, cytosine (C) in the CHH sequence may be included in cytosine (C) that can be methylated, and cytosine (C) picked up from among them may be used as the target site.

[0073] In this Modification Example 2, the target site information acquisition unit 20 further acquires one or more target sites included in the genomic double-stranded DNA acquired by the base sequence data acquisition unit 20 and their position information via the input unit 12. The base conversion unit 24 further converts cytosine (C) in the CHH sequence on the template DNA acquired from the base sequence data acquisition unit 20 to "Y", and cytosine (C) in other sequences (that is, sequences other than the CG sequence and the CHH sequence) is converted to thymine (T).

[0074] The primer candidate sequence selection unit 30 selects, as primer candidate sequences, those that satisfy all of the predetermined selection conditions (1) to (3) and (5) including the following item (5) from one or more partial sequences cut out by the partial sequence cutting unit 28. (5) The number of YHH sequences or DDR sequences included on the partial sequence is equal to or less than a predetermined number Here, the number of "YHH sequences or DDR sequences contained in the partial sequence" that satisfies the above condition (5) is not particularly limited. However, from the viewpoint of significantly obtaining the desired effects of the present invention, it is preferably 2 or less, more preferably 1 or less, and particularly preferably 0. By satisfying this condition, the influence of the primer on the binding with cytosine (C) of the CHH sequence at the primer binding site can be reduced.

[0075] With the primer design device of Modification 2 of Embodiment 1 of the present invention, primers for amplicon methylation sequence analysis corresponding to the CHH sequence can be easily designed in a short time. Further, primers based on the design can be obtained. As a result, analysis related to these sequences becomes possible, and the methylation state (methylation degree) of DNA can be analyzed in more detail.

[0076] Note that Modification 2 can be combined with the previously described Modification 1. That is, it may contain both cytosine (C) in the CHG sequence and cytosine (C) in the CHH sequence in cytosine (C) that can be methylated, and cytosine (C) picked up therefrom may be used as a target site. In such a case, the primer candidate sequence selection unit 30 further selects, as primer candidate sequences, those that satisfy all of the selection conditions (1) to (5) from one or more partial sequences of each strand cut out by the partial sequence cutting unit 28. Here, the process of selecting, as primer candidate sequences, those that satisfy the selection conditions (1) to (5) from one or more partial sequences cut out from each strand will be described using the flowchart of FIG. 6. Note that FIG. 6 is a flowchart showing a modification of condition A in FIG. 4, and the description of the same processes as those in Embodiment 1 will be omitted.

[0077] The primer candidate sequence selection unit 30 first obtains one (step S300) from one or more partial sequences cut out from the first template strand (A+ strand), and determines whether the Tm value of the partial sequence is within a predetermined range (step S302). If the Tm value is not within the predetermined range, another subarray is obtained (step S300). If the Tm value is within the predetermined range, it is determined whether the number of YG arrays or CR arrays included in the subarray is equal to or less than a predetermined number (step S304). If the number of YG arrays or CR arrays included in the subarray is not equal to or less than the predetermined number, another subarray is obtained (step S300). If the number of YG arrays or CR arrays included in the subarray is equal to or less than the predetermined number, it is determined whether the number of YHG arrays or CHR arrays included in the subarray is equal to or less than a predetermined number (step S314). If the number of YHG arrays or CHR arrays included in the subarray is not equal to or less than the predetermined number, another subarray is obtained (step S300). If the number of YHG arrays or CHR arrays included in the subarray is equal to or less than the predetermined number, it is determined whether the number of YHH arrays or DDR arrays included in the subarray is equal to or less than a predetermined number (step S316). If the number of YHH arrays or DDR arrays included in the subarray is not equal to or less than the predetermined number, another subarray is obtained (step S300). If the number of YHH arrays or DDR arrays included in the subarray is equal to or less than the predetermined number, it is determined whether the upper limit of the number of junctions with the sequences outside the relevant region on the template strand DNA after base conversion is equal to or less than a predetermined number greater than or equal to 1 (step S306). If the upper limit of the number of junctions with the sequences outside the relevant region on the template strand DNA after base conversion is not equal to or less than the predetermined number greater than or equal to 1, another subarray is obtained (step S300). If the upper limit of the number of junctions with the sequences outside the relevant region on the template strand DNA after base conversion is equal to or less than the predetermined number greater than or equal to 1, it is selected as a primer candidate sequence (step S308), and it is determined whether the determination of all subarrays excised from the first template strand (A+ strand) has been completed (step S310). If the determination of all subarrays excised from the first template strand (A+ strand) has not been completed, another subarray is obtained (step S300). If the determination of all subarrays has been completed, one or more selected primer candidate sequences are determined as the forward primer candidate sequences of the first template strand (A+ strand) (step S312).

[0078] By satisfying this condition, the influence of the primer and the binding of the cytosine (C) in the CHG sequence and the binding of the cytosine (C) in the CHH sequence at the primer binding site can be eliminated.

[0079] [Modification Example 3] Next, the primer design apparatus according to Modification Example 3 of Embodiment 1 of the present invention will be described. In the primer design apparatus according to this Modification Example 3, the description of the same processing as in Embodiment 1 will be omitted. In the present embodiment, in the primer candidate sequence selection unit 30, although items (1) to (3) are set as predetermined selection conditions used for selecting primer candidate sequences, it is preferable to further include the following item (6). (6) The three bases from the 3'-end of the partial sequence are not complementary to the three bases from the 3'-end of the other partial sequence The fact that the three bases from the 3'-end of the partial sequence are not complementary to each other means that for all combinations of primers obtained by taking out two primers from n primers, the three bases from the 3'-end of the primers are not complementary. For example, as shown in FIG. 7, for the combination of the primers of SEQ ID NO: 1 and SEQ ID NO: 40, the three bases from the 3'-end are (underlined part), AG pair or GG pair, and they are not complementary. By satisfying this condition, the binding of primers to each other on the 3'-end side can be avoided. Note that Modification Example 3 can be combined with at least one of the previously described Modification Examples 1 and 2. Further, according to Modification Example 3, the above-described effects can be additionally obtained.

[0080] [Modification Example 4] Next, the primer design apparatus according to Modification Example 4 of Embodiment 1 of the present invention will be described. In the primer design apparatus according to this Modification Example 4, the description of the same processing as in Embodiment 1 will be omitted. In the present embodiment, although items (1) to (3) are set as predetermined selection conditions used for selecting primer candidate sequences in the primer candidate sequence selection unit 30, it is preferable to further include the following item (7). (7) The number of junctions with the genomic double-stranded DNA before base conversion is equal to or less than a predetermined number Here, the upper limit of the number of junctions between the genomic double-stranded DNA before base conversion and the partial sequence is not particularly limited. However, from the viewpoint of significantly obtaining the desired effects of the present invention, it is preferably 5 or less, and particularly preferably 2 or less. Note that Modification Example 4 can be combined with at least one of the existing Modification Examples 1 to 3.

[0081] [Modification Example 5] Next, a primer design device according to Modification Example 5 of Embodiment 1 of the present invention will be described. In the primer design device according to this Modification Example 5, descriptions of the same processes as those in Embodiment 1 are omitted. In Modification Example 1 of Embodiment 1, although items (1) to (3) are set as predetermined selection conditions used for selecting primer candidate sequences in the primer candidate sequence selection unit 30, it is preferable to further include the following item (8). (8) In the condition of (2) above, when a predetermined number of YG sequences or CR sequences included in the partial sequence is set to 1 or more, the position range of the YG sequence or CR sequence in the partial sequence is also specified, and it is preferable that the number of YG sequences or CR sequences within the specified position range is equal to or less than a predetermined number. Here, the "position range of the YG sequence or CR sequence in the partial sequence" indicates the position where the YG sequence or CR sequence exists within the partial sequence. For example, a predetermined position range including the 5'-end of the partial sequence or a predetermined position range including the 3'-end can be specified. The specified "position range" is not particularly limited, but it is preferable to specify the position range on the 5'-end side of the partial sequence. This is because the influence of cytosine (C) methylation can be reduced more than when specifying the position range on the 3'-end side of the partial sequence. More specifically, a base mismatch (non-complementary pair) between the primer and the template DNA generally has a greater influence and more frequently inhibits ligation when it is on the 3'-end side than when it is on the 5'-end side. In the present invention, when there is a site with base indeterminacy in the region where the primer ligates, it is preferable to arrange it on the 5'-end side so that the ligation of the primer is not inhibited by its presence.

[0082] Similarly, in Modification 2 of Embodiment 1, in the primer candidate sequence selection unit 30, items (1) to (4) were set as predetermined selection conditions for selecting the primer candidate sequence, but it is further preferable to include the following item (9). (9) In the condition of (4) above, when the predetermined number of YHG sequences or CDR sequences included in the partial sequence is set to 1 or more, the position range of the YHG sequence or CDR sequence in the partial sequence is also specified, and the YHG sequence or CDR sequence is included in a predetermined number or less within the specified position range. Similar to condition (8), it is preferable to specify the position range on the 5'-end side of the partial sequence. As described above, this is because the influence of cytosine (C) methylation can be reduced more than when specifying the position range on the 3'-end side of the partial sequence.

[0083] Similarly, in Modification 3 of Embodiment 1, in the primer candidate sequence selection unit 30, items (1) to (3) and (5) were set as predetermined selection conditions for selecting the primer candidate sequence, but it is further preferable to include the following item (10). (10) In the condition of (5) above, when the predetermined number of YHH sequences or DDR sequences included in the partial sequence is set to 1 or more, it is preferable to also specify the position range of the YHH sequence or DDR sequence in the partial sequence and include the YHH sequence or DDR sequence in a predetermined number or less within the specified position range. Similar to condition (8), it is preferable to specify the "position range" on the 5'-end side of the subarray. This is because, as described above, the influence of cytosine (C) methylation can be reduced more than when specifying the position range on the 3'-end side of the subarray.

[0084] In addition, in the primer candidate sequence selection unit 30, as predetermined selection conditions used to select primer candidate sequences, items (1) to (5) are set, and further, the conditions of items (8) to (10) may be included. Note that Modification Example 5 can be combined with at least one of the existing Modification Examples 3 and 4. Further, according to Modification Example 5, the above-described effects can be additionally obtained. [[ID=⑧]]

[0085] [[ID=⑨]] [[ID=⑩]][Modification Example 6][[ID=⑪]] [[ID=⑫]]Next, a primer design apparatus according to Modification Example 6 of Embodiment 1 of the present invention will be described. In the primer design apparatus according to this Modification Example 6, the same components as those in Embodiment 1 are denoted by the same reference numerals, and the description of the same processes as those in Embodiment 1 is omitted. [[ID=⑬]] [[ID=⑭]]FIG. 8 is a flowchart showing a modification of Condition B in FIG. 4, and the description of the same processes as those in Embodiment 1 is omitted. [[ID=⑮]] [[ID=⑯]]

[0086] [[ID=⑰]] In Embodiment 1, the primer sequence determination unit 32 selects, based on the lengths of the PCR amplification products calculated respectively, a forward primer sequence and a reverse primer sequence that amplify the region including the target site selected by the partial sequence extraction unit 28 from all combinations of one or more primer candidate sequences determined by the primer candidate sequence selection unit 30 (step S312), that is, all combinations of one or more forward primer candidate sequences of the first template strand (A+ strand) and one or more reverse primer candidate sequences of the first template strand (A+ strand), all combinations of one or more forward primer candidate sequences of the second template strand (B+ strand), and all combinations of one or more reverse primer candidate sequences of the second template strand (B+ strand), and determines them (step S320). However, the present invention is not limited to this. As shown in FIG. 8, after adopting the forward primer sequence and the reverse primer sequence of the first template strand or the second template strand (step S320), further, for all combinations of the adopted primer sequences, a local alignment score is calculated, and it is determined whether the calculated local alignment score is lower than a predetermined threshold value. Those with a calculated local alignment score lower than the predetermined threshold value are adopted as primer sequences (step S324) and can be determined (step S322). Note that the above-mentioned "all combinations of the adopted primer sequences" is not limited to the combination of a forward primer and a reverse primer, but includes combinations of a forward primer and a forward primer, combinations of a reverse primer and a reverse primer, and is not limited to the forward primer sequence of the first template strand (A+ strand) and the reverse primer sequence of the first template strand (A+ strand), but includes combinations such as the combination of the forward primer sequence of the first template strand (A+ strand) and the reverse primer sequence of the second template strand (B+ strand).

[0087] Under this condition, the calculated local alignment score is based on a score matrix in which the score is set high when the bases are complementary and low when they are not complementary. For example, a method for calculating the local alignment score for the combination of array number 1 and array number 40 shown in FIG. 8 will be described. In the figure, when the bases of array number 1 and array number 40 match (complementary pair / match string), "|" is attached, when they do not match (non-complementary pair / mismatch string), ":" is attached, and "-" is attached to gaps. In FIG. 9, since there are 7 complementary pairs, 15 non-complementary pairs, and 4 gaps, taking the complementary pairs as "+1", the non-complementary pairs as "0", and the gaps as "-1", the score of the pairwise alignment (local alignment score) for the combination of array number 1 and array number 40 is calculated as (+1×7)+(0×15)+(-1×4)=+3. By satisfying this condition, the influence caused by the joining of primers can be avoided. Note that Modification Example 6 can be combined with at least one of the existing Modification Examples 1 to 5. Further, according to Modification Example 6, the above-described effects can be additionally obtained.

[0088] [Modification Example 7] Next, a primer design apparatus according to Modification Example 7 of Embodiment 1 of the present invention will be described. In the primer design apparatus according to this Modification Example 7, the same components as those in Embodiment 1 are denoted by the same reference numerals, and the description of the same processes as those in Embodiment 1 is omitted. In Embodiment 1, an apparatus and method for designing two sets of primers for amplifying and analyzing both strands of a DNA double strand were shown, but the present invention is not limited thereto. When analyzing only one of the strands of a DNA double strand, only one set of primers needs to be designed. That is, in FIG. 3B, primers were designed based on strand A and strand B, but primers may be designed based on only strand A. In addition, when it is considered that the DNA maintenance methylation mechanism is functioning, only one set of primers needs to be designed. This is because when the C in the CG sequence of one strand of DNA is methylated, it is very likely that the C in the CG sequence of the other strand is also methylated, and when the C in the CG sequence of one strand is not methylated, it is very likely that the C in the CG sequence of the other strand is also not methylated. In such a case, when a set of primers cannot be designed based on one strand, primer design can be performed based on the other strand.

[0089] Thus, when only one set of primers is designed, in the complementary strand generation unit 26, only the complementary strand A- having a base sequence complementary to the base sequence of the A+ strand shown in FIG. 3C is produced. Next, the partial sequence excision unit 28 selects one target site from among the one or more target sites acquired by the target site information acquisition unit 22 (step S280), and based on the position information of the selected target site, from the DNA sequences of the A+ strand and the A- strand, the "Y" at the selected target site or its complementary "R" (that is, the base at the target site and at the methylation site) is detected, and as much as possible is excised from the base sequences located on the 5'-terminal side of the detected "Y" and "R" (FIG. 3D (1) and (2)) from among partial sequences of a predetermined length (step S282), and one or more partial sequences are acquired.

[0090] The primer candidate sequence selection unit 30 is the part that performs the primer candidate sequence selection step S20 shown in FIG. 2, and selects, as primer candidate sequences, those that satisfy all of the predetermined selection conditions (1) to (3) from among the one or more partial sequences of each strand excised by the partial sequence excision unit 28. If one or more partial sequences excised from the first template strand (A+ strand) (i.e., one or more partial sequences excised from (1) in FIG. 3D) satisfy a predetermined selection condition, they are selected as forward primer candidate sequences for the first template strand (A+ strand). If one or more partial sequences excised from the first complementary strand (A− strand) (i.e., one or more partial sequences excised from (2) in FIG. 3D) satisfy a predetermined selection condition, they are selected as reverse primer candidate sequences for the first template strand (A+ strand).

[0091] The primer sequence determination unit 32 acquires all primer pairs (combinations of a forward primer and a reverse primer) that can be prepared from one or more forward primer candidate sequences of the selected first template strand (A+ strand) and one or more reverse primer candidate sequences of the first template strand (A+ strand), and calculates the length of a PCR amplification product that is expected to be amplified by PCR for each primer pair. Next, it is determined whether the calculated length of the PCR amplification product is within a predetermined numerical range. If the calculated length of the PCR amplification product is within the predetermined numerical range, the primer pair for which the length of the PCR amplification product has been calculated (i.e., the combination of the forward primer candidate sequence and the reverse primer candidate sequence) is adopted and determined as the forward primer sequence and the reverse primer sequence of the first template strand that amplify the region including the target site selected in the partial sequence excision unit 28 (partial sequence excision step S18). Note that Modification 7 can be combined with at least one of the above-described Modifications 1 to 6.

[0092] [Embodiment 2] FIG. 10 is a block diagram conceptually showing an example of a primer design apparatus according to Embodiment 5 of the present invention. The apparatus 10 of Embodiment 1 can also include a communication interface (communication device). The primer design apparatus 10A of Embodiment 2 shown in FIG. 10 has the same configuration as the primer design apparatus 10 of Embodiment 1 shown in FIG. 1, except that it has a communication interface 36. Therefore, the same reference numerals are given to the same components, and the description thereof is omitted. As shown in FIG. 11, the primer design apparatus 10A can be connected to a server 42 having a public database installed outside the apparatus via a communication network 38 such as the Internet. The apparatus 10A of the present embodiment can execute at least one of a base sequence data acquisition unit 20, a target site information acquisition unit 22, a base conversion unit 24, a complementary strand generation unit 26, a partial sequence extraction unit 28, a primer candidate sequence selection unit 30, and a primer sequence determination unit 32 on a program located on the site of an external server 40 via a communication interface 36. In such a case, the data processing apparatus 10A of the present embodiment may not include each means executed by the program of the external server.

[0093] For example, based on an instruction from the control unit 34, the communication interface 36 can acquire a DNA base sequence including genes and genomes from a public database via the communication network 38 and store it in the storage unit 14. Here, examples of the public database include GenBank of NCBI (National Center for Biotechnology Information) in the United States, ENA of EMBL (European Molecular Biology Laboratory), and DDBJ of the National Institute of Genetics. The base sequence acquired from the public database may be a partial sequence of the base sequence of the genomic DNA of the species for which the primer is designed, but is preferably the entire sequence.

[0094] <( For example, based on an instruction from the control unit 34, the communication interface 36 can perform a sequence homology search using a public search server 40 via the communication network 38, and perform junction determination according to condition 8 of the primer candidate sequence selection unit 30, local alignment search of the primer sequence determination unit 32 in Modification 6, and the like. Here, examples of the public search server include BLAST of NCBI (National Center for Biotechnology Information) in the United States.

[0095] [Embodiment 3] Embodiment 3 is a method for manufacturing a primer by synthesizing a primer based on a primer sequence designed by the primer design apparatus and design method according to Embodiments 1 and 2. The primer design method is as shown in Embodiments 1 and 2. As the primer synthesis method, a known method can be used. For example, there are methods such as chemical synthesis from the terminal base using dNTP (Deoxyribonucleoside triphosphate) and the like as materials by a DNA synthesizer or an RNA synthesizer. As the synthesizer, a commercially available product can be used.

[0096] Each component included in the device of the present invention may be configured by dedicated hardware, or each component may be configured by a programmed computer. The method of the present invention can be implemented, for example, by a program for causing a computer to execute each of its steps. Also, a computer-readable recording medium on which this program is recorded can be provided.

[0097] Although the present invention has been described in detail above, the present invention is not limited to the above embodiments, and it goes without saying that various improvements and modifications may be made without departing from the gist of the present invention.

Example

[0098] [Example 1, Comparative Example 1] Using the primer design apparatus of the first embodiment, multiplex PCR primers with a length of 100 bp to 300 bp for PCR amplification products were designed based on the nucleotide sequence data of the reference genome GRCh37 (GenBank assembly accession: GCA_000001405.1, RefSeq assembly accession: GCF_000001405.13), 50 measurement sites (target sites) shown in Table 1, and their position information. Note that the length of the primer was 20 to 30 bases (mer), and C that could be methylated was only C in the CG sequence when designing the primer. Also, the conditions of Example 1 for determining the partial sequence were set as follows. Condition (1): The Tm value is between 55°C and 60°C. Condition (2): The YG sequence or CR sequence contained in the partial sequence is 0. Condition (3): The upper limit of the number of junctions with sequences outside the relevant region is 2.

[0099] On the other hand, the conditions of Comparative Example 1 for determining the partial sequence were set as follows. Condition (1): The Tm value is between 55°C and 60°C. Condition (2): The YG sequence or CR sequence contained in the partial sequence is 0. Condition (3): The number of junctions with sequences outside the relevant region is 0.

[0100] Table 1 shows the success or failure of primer design for each measurement site in Example 1 and Comparative Example 1, and the primer design success rate calculated from the results of the success or failure of primer design. Also, Table 2 shows the primers that could be designed in Example 1, and Table 3 shows the primers that could be designed in Comparative Example 1.

[0101] As shown in Table 1, in the example, primers related to the target sites of ID9, 19, 28, and 50 could be designed, but in Comparative Example 1, primers related to these target sites could not be designed. The success rate of each design was 78% in Example 1 and 70% in Comparative Example 1. It was confirmed that by performing the determination of Condition (3) for the partial sequence, the design success rate could be increased.

[0102]

Table 1

[0103]

Table 2

[0104]

Table 3

Explanation of Symbols

[0105] 10, 10A Primer Design Device 12 Input Unit 14 Storage Unit 16 Output Unit 18 Primer Design Processing Unit 20 Nucleotide Sequence Data Acquisition Unit 22 Target Site Information Acquisition Unit 24 Nucleotide Conversion Unit 26 Complementary Strand Generation Unit 28 Subsequence Extraction Unit 30 Primer Candidate Sequence Selection Unit 32 Primer Sequence Determination Unit 34 Control Unit 36 Communication Interface 38 Communication Network 40, 42 Server

Industrial Applicability

[0106] The primers designed by the present invention can be used for measuring the DNA methylation degree of biological samples in the fields of drug discovery, diagnosis, and other biotech industries.

Claims

1. A method for designing primers used to simultaneously amplify a plurality of regions each containing one or more target sites for measuring the methylation level of genomic double-stranded DNA at a predetermined site related to a predetermined biological phenomenon, using a bisulfite reaction or an enzymatic reaction and multiplex PCR, comprising: a nucleotide sequence data acquisition step of acquiring nucleotide sequence data of the genomic double-stranded DNA; a target site information acquisition step of acquiring the one or more target sites and their position information; a base conversion step of converting "C" that can be methylated in the genomic double-stranded DNA to "Y" and converting other "C" to "T" in the nucleotide sequence data; a complementary strand generation step of generating a complementary strand for each template strand of the genomic double-stranded DNA after the base conversion; a partial sequence excision step of selecting one from the one or more target sites and excising one or more partial sequences of a predetermined length from the base sequences located on the 5'-end side of the "Y" or its complementary "R" into which the selected target site has been converted from each strand based on the position information of the selected target site; a primer candidate sequence selection step of selecting, as primer candidate sequences, those that satisfy a predetermined selection condition from the one or more partial sequences excised from each strand; a primer sequence determination step of adopting and determining a forward primer sequence and a reverse primer sequence that amplify a region containing the selected target site and are excised from each template strand from the one or more selected primer candidate sequences; a repetition step of repeating the partial sequence excision step, the primer candidate sequence selection step, and the primer sequence determination step until all of the one or more target sites are selected in the partial sequence excision step; having, wherein the "C" that can be methylated is the "C" in the CG sequence, the predetermined selection conditions are: (1) the Tm value is within a predetermined range; (2) the number of YG sequences or CR sequences contained in the partial sequence is not more than a predetermined number; and (3) the upper limit of the number of junctions with sequences outside the relevant region on the genomic double-stranded DNA after the base conversion is not more than a predetermined number of 1 or more including, a primer design method for amplicon methylation sequence analysis. [However, "C", "G", "Y", and "R" are base notations defined by the IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine.]

2. The "C" that can be methylated further includes the "C" in the CHG sequence, The predetermined selection conditions further include (4) that the number of YHG sequences or CDR sequences contained in the partial sequence is equal to or less than a predetermined number, according to the primer design method described in claim 1. [However, "C", "G", "Y", "H", "R", and "D" are base notations defined by the IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine.]

3. The "C" that can be methylated further includes the "C" in the CHH sequence, The predetermined selection conditions further include (5) that the number of YHH sequences or DDR sequences contained in the partial sequence is equal to or less than a predetermined number, according to the primer design method described in claim 1 or 2. [However, "Y", "H", "R", and "D" are base notations defined by the IUPAC, where "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine.]

4. The predetermined selection conditions further include (6) that the three bases from the 3'-end of the partial sequence are not complementary to the three bases from the 3'-end of another partial sequence, according to the primer design method described in any one of claims 1 to 3.

5. The predetermined selection conditions further include (7) that the junction with the genomic double-stranded DNA before the base conversion is equal to or less than a predetermined number, according to the primer design method described in any one of claims 1 to 4.

6. The predetermined selection conditions further include (8) that in the condition of (2), when the predetermined number of YG sequences or CR sequences contained in the partial sequence is set to 1 or more, the position range of the YG sequence or CR sequence in the partial sequence is also specified, within the specified position range, the number of YG sequences or CR sequences contained in the partial sequence is equal to or less than the predetermined number, according to the primer design method described in any one of claims 1 to 5.

7. The predetermined selection conditions further include: (9) in the condition of (4) above, when the predetermined number of YHG sequences or CDR sequences included in the partial sequence is set to 1 or more, the position range of the YHG sequence or CDR sequence in the partial sequence is also specified, The primer design method according to any one of claims 2 to 6, wherein the YHG sequence or CDR sequence within the specified position range is included in a number less than or equal to the predetermined number.

8. The predetermined selection conditions further include: (10) in the condition of (5) above, when the predetermined number of YHH sequences or DDR sequences included in the partial sequence is set to 1 or more, the position range of the YHH sequence or DDR sequence in the partial sequence is also specified, The primer design method according to any one of claims 3 to 7, wherein the YHH sequence or DDR sequence within the specified position range is included in a number less than or equal to the predetermined number.

9. The primer candidate sequence selection step is Using the double-stranded genomic DNA after the base conversion as the first template strand and the second template strand, and using the complementary strand of the first template strand as the first complementary strand and the complementary strand of the second template strand as the second complementary strand, For one or more partial sequences excised from the first template strand that satisfy the predetermined selection conditions, they are selected as forward primer candidate sequences for the first template strand. For one or more partial sequences excised from the first complementary strand that satisfy the predetermined selection conditions, they are selected as reverse primer candidate sequences for the first template strand. For one or more partial sequences excised from the second template strand that satisfy the predetermined selection conditions, they are selected as forward primer candidate sequences for the second template strand. For one or more partial sequences excised from the second complementary strand that satisfy the predetermined selection conditions, they are selected as reverse primer candidate sequences for the second template strand. The primer design method according to any one of claims 1 to 8 is a step of selection.

10. The primer sequence determination step is as follows: in the primer candidate sequence selection step, for all combinations of the selected forward primer candidate sequences of one or more of the first template strands and the selected reverse primer candidate sequences of one or more of the first template strands, calculate the length of the PCR amplification product expected to be amplified by PCR, and use the combination of primer candidate sequences for which the calculated length of the PCR amplification product is within a predetermined range as the forward primer sequence and the reverse primer sequence of the first template strand that amplifies the region including the target site selected in the partial sequence excision step. For all combinations of the selected forward primer candidate sequences of the second template strand and the selected reverse primer candidate sequences of the second template strand, calculate the length of the PCR amplification product expected to be amplified by PCR, and use the combination of primer candidate sequences for which the calculated length of the PCR amplification product is within the predetermined range as the forward primer sequence and the reverse primer sequence of the second template strand that amplifies the region including the target site selected in the partial sequence excision step to determine. The primer design method according to claim 9.

11. After adopting the forward primer sequence and the reverse primer sequence for all target sites, The primer sequence determination step further calculates the local alignment score for all combinations of the adopted primer sequences, and adopts and determines as the primer sequence those for which the calculated local alignment score is lower than a predetermined threshold. The primer design method according to any one of claims 1 to 10.

12. Under the condition of (3) above, the upper limit of the number of junctions with the partial sequence is 1 or 2. The primer design method according to any one of claims 1 to 11.

13. A primer design step, A synthesis step of synthesizing primers based on the primer sequences designed in the primer design step, Comprising: The primer design step is carried out by the primer design method according to any one of claims 1 to 12. A method for manufacturing primers, characterized in that.

14. An apparatus for designing primers used to simultaneously amplify a plurality of regions each containing one or more target sites for measuring the methylation level of genomic double-stranded DNA of a predetermined site related to a predetermined biological phenomenon, by using a bisulfite reaction or an enzymatic reaction and multiplex PCR, a nucleotide sequence data acquisition unit that acquires nucleotide sequence data of the genomic double-stranded DNA; a target site information acquisition unit that acquires the one or more target sites and their position information; a base conversion unit that, in the nucleotide sequence data, converts "C" that can be methylated in the genomic double-stranded DNA to "Y", and converts other "C" to "T"; a complementary strand generation unit that generates a complementary strand for each template strand of the genomic double-stranded DNA after the base conversion; a partial sequence excision unit that selects one from the one or more target sites, and excises one or more partial sequences of a predetermined length from the nucleotide sequences located on the 5'-terminal side of the "Y" into which the selected target site has been converted or "R" complementary thereto, from each strand, based on the position information of the selected target site; a primer candidate sequence selection unit that selects, as primer candidate sequences, those that satisfy a predetermined selection condition from among the one or more partial sequences excised from each strand; a primer sequence determination unit that adopts and determines a forward primer sequence and a reverse primer sequence that amplify a region containing the selected target site, excised from each template strand, from among the one or more selected primer candidate sequences; a control unit that controls the partial sequence excision unit, the primer candidate sequence selection unit, and the primer sequence determination unit to repeat each process until all of the one or more target sites are selected; comprising wherein the "C" that can be methylated is the "C" in the CG sequence; the predetermined selection condition includes (1) that the Tm is within a predetermined range; (2) that the number of YG sequences or CR sequences contained in the partial sequence is equal to or less than a predetermined number; and (3) that the upper limit of the number of junctions with sequences outside the relevant region on the genomic double-stranded DNA after the base conversion is equal to or less than a predetermined number of 1 or more ; a primer design apparatus for amplicon methylation sequence analysis. [However, "C", "G", "Y", and "R" are base notations defined by IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, and "R" represents adenine or guanine.]

15. The "C" that can be methylated further includes "C" in the CHG sequence, The predetermined selection conditions further include (4) that the number of YHG sequences or CDR sequences included in the partial sequence is equal to or less than a predetermined number, for the primer design device according to claim 14. [However, "C", "G", "Y", "H", "R", and "D" are base notations defined by IUPAC, where "C" represents cytosine, "G" represents guanine, "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine.]

16. The "C" that can be methylated further includes "C" in the CHH sequence, The predetermined selection conditions further include (5) that the number of YHH sequences or DDR sequences included in the partial sequence is equal to or less than a predetermined number, for the primer design device according to claim 14 or 15. [However, "Y", "H", "R", and "D" are base notations defined by IUPAC, where "Y" represents thymine or cytosine, "H" represents adenine, cytosine, or thymine, "D" represents thymine, guanine, or adenine, and "R" represents adenine or guanine.]

17. The predetermined selection conditions further include (6) that the three bases from the 3'-end of the partial sequence are not complementary to the three bases from the 3'-end of another partial sequence, for the primer design device according to any one of claims 14 to 16.

18. The predetermined selection conditions further include (7) that the junction with the genomic double-stranded DNA before the base conversion is equal to or less than a predetermined number, for the primer design device according to any one of claims 14 to 17.

19. The predetermined selection conditions further include (8) that in the condition of (2), when the predetermined number of YG sequences or CR sequences included in the partial sequence is set to 1 or more, the position range of the YG sequence or CR sequence in the partial sequence is also specified, and within the specified position range, the number of YG sequences or CR sequences included in the partial sequence is equal to or less than the predetermined number, for the primer design device according to any one of claims 14 to 18.

20. The predetermined selection conditions further include: (9) in the condition of (4) above, when the predetermined number of YHG sequences or CDR sequences included in the partial sequence is set to 1 or more, the position range of the YHG sequence or CDR sequence in the partial sequence is also specified, The primer design device according to any one of claims 15 to 19, wherein within the specified position range, the number of the YHG sequence or CDR sequence is included below the predetermined number.

21. The predetermined selection conditions further include: (10) in the condition of (5) above, when the predetermined number of YHH sequences or DDR sequences included in the partial sequence is set to 1 or more, the position range of the YHH sequence or DDR sequence in the partial sequence is also specified, The primer design device according to any one of claims 16 to 20, wherein within the specified position range, the number of the YHH sequence or DDR sequence is included below the predetermined number.

22. The primer candidate sequence selection unit uses the double-stranded genomic DNA after base conversion as the first template strand and the second template strand, and uses the complementary strand of the first template strand as the first complementary strand and the complementary strand of the second template strand as the second complementary strand. For one or more partial sequences excised from the first template strand that satisfy the predetermined selection conditions, they are selected as the forward primer candidate sequences for the first template strand. For one or more partial sequences excised from the first complementary strand that satisfy the predetermined selection conditions, they are selected as the reverse primer candidate sequences for the first template strand. For one or more partial sequences excised from the second template strand that satisfy the predetermined selection conditions, they are selected as the forward primer candidate sequences for the second template strand. For one or more partial sequences excised from the second complementary strand that satisfy the predetermined selection conditions, they are selected as the reverse primer candidate sequences for the second template strand. The primer design device according to any one of claims 14 to 21.

23. The primer sequence determination unit, in the primer candidate sequence selection unit, calculates the lengths of PCR amplification products expected to be amplified by PCR for all combinations of the selected forward primer candidate sequences of the one or more first template strands and the selected reverse primer candidate sequences of the one or more first template strands, and selects combinations of primer candidate sequences for which the calculated lengths of the PCR amplification products are within a predetermined range as the forward primer sequence and the reverse primer sequence of the first template strand that amplifies the region including the target site selected in the partial sequence excision unit. For all combinations of the selected forward primer candidate sequences of the second template strand and the selected reverse primer candidate sequences of the second template strand, the length of the PCR amplification product expected to be amplified by PCR is calculated, and combinations of primer candidate sequences for which the calculated length of the PCR amplification product is within the predetermined range are used as the forward primer sequence and the reverse primer sequence of the second template strand that amplifies the region including the target site selected in the partial sequence excision unit to determine. The primer design device according to claim 22.

24. For all target sites, after adopting the forward primer sequence and the reverse primer sequence, the primer sequence determination unit further calculates a local alignment score for all combinations of the adopted primer sequences, and adopts and determines as the primer sequence those for which the calculated local alignment score is lower than a predetermined threshold. The primer design device according to any one of claims 14 to 23.

25. Under the condition of (3) above, the upper limit of the number of junctions with the partial sequence is 1 or 2. The primer design device according to any one of claims 14 to 24.

26. Furthermore, it is provided with a communication interface, The communication interface can be used to connect to a server via a communication line network outside the device, and a program in the server can execute at least one of the group consisting of the nucleotide sequence data acquisition unit, the target site information acquisition unit, the nucleotide conversion unit, the complementary strand generation unit, the partial sequence extraction unit, the primer candidate sequence selection unit, and the primer sequence determination unit. The primer design device according to any one of claims 14 to 25.

27. A program for designing primers, characterized in that the primer design method according to any one of claims 1 to 13 is executed on a computer.

28. A computer-readable recording medium, characterized in that the program for designing primers according to claim 27 is recorded thereon.

Citation Information

Patent Citations

  • Multi-methylation specific PCR primer design method and system

    CN111653311A

  • Pipeline method for CpG island prediction and primer production

    KR1020180130376A

  • Conducting a fluid in a flow vessel

    WO2016156132A1