Design method for primer set, primer set, base sequence determination method for total length of mitochondrial DNA, design device for primer set, computer program, and storage medium
The method designs a primer set for mitochondrial DNA sequencing by aligning conserved regions and optimizing primer specificity, addressing the challenges of species commonality and fragmented DNA, achieving efficient and accurate sequencing across species.
Patent Information
- Application Number
- PCT/JP2024/042264
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-08
- Filing Date
- 2024-11-29
- Publication Date
- 2025-07-17
AI Technical Summary
Existing methods for determining the nucleotide sequence of mitochondrial DNA face challenges in achieving high commonality across species, handling fragmented DNA, and simplifying the process, particularly due to the need for numerous primer sequences and the amplification of non-mitochondrial DNA, leading to increased complexity and cost.
A method for designing a primer set that aligns mitochondrial DNA sequences from multiple species to identify conserved regions, sets primer candidates based on high conservation, and selects a set that amplifies the DNA into three to six overlapping fragments, suitable for long-read sequencing, using a computer program to optimize primer specificity and overlap length.
This approach allows for efficient and accurate determination of mitochondrial DNA sequences across various species with reduced complexity and cost, enabling high accuracy and simplification of the sequencing process even with fragmented DNA samples.
Smart Images

Figure JP2024042264_17072025_PF_FP_ABST
Abstract
Description
Method for designing a primer set, primer set, method for determining full-length nucleotide sequence of mitochondrial DNA, primer set design device, computer program, and storage medium
[0001] The present disclosure relates to a method for designing a primer set for determining the full-length base sequence of mitochondrial DNA, a primer set, a method for determining the full-length base sequence of mitochondrial DNA, a primer set design device, a computer program, and a storage medium.
[0002] Methods for investigating or monitoring biological species present in the environment include recovering biologically derived nucleic acids from the environment and analyzing the recovered nucleic acids to identify the biological species from which the nucleic acids originate. Biologically derived nucleic acids are nucleic acids that contain the genetic information of an organism and are released from an organism into the environment, and are included in nucleic acids present in the environment, such as environmental DNA. A known method for analyzing such environmental DNA involves using a known mitochondrial DNA base sequence as a reference sequence to identify the biological species from which mitochondrial DNA collected from the environment originates. To analyze environmental DNA using a mitochondrial DNA base sequence as a reference sequence, it is necessary to register mitochondrial DNA sequence information for various organisms in a database. In this case, to improve the accuracy of the analysis, it is desirable to accumulate the full mitochondrial DNA base sequences of a greater number of organisms in the database, thereby enriching the database.
[0003] In order to enrich the database and efficiently obtain information linking the entire mitochondrial DNA sequence with a wide range of biological species, it is desirable to satisfy, for example, the following conditions. That is, as a first condition, it is desirable to reduce the burden required for primer design by using primers that can be commonly applied to a fairly wide range of biological species as primers for obtaining mitochondrial DNA sequences. As a second condition, it is desirable to be able to analyze the mitochondrial DNA sequence even when the quality of environmental DNA containing mitochondrial DNA is relatively low and the mitochondrial DNA used is not circular but rather fragmented to some extent. As a third condition, it is desirable to be able to easily construct the entire mitochondrial DNA sequence with as few steps as possible.
[0004] As a method for obtaining the entire base sequence of mitochondrial DNA, for example, Patent Document 1 discloses a configuration in which a relatively short human mitochondrial genome fragment is amplified using 24 known primer pairs designed for humans to determine the entire length of mitochondrial DNA. Non-Patent Document 1 discloses a method in which the entire length of mitochondrial DNA is amplified as a single fragment using one pair of primers, and Non-Patent Documents 2 and 3 disclose methods in which the entire length of mitochondrial DNA is amplified as two fragments using two pairs of primers. Patent Document 2 also discloses a method for amplifying all DNA fragments of the entire genome, including mitochondrial DNA.
[0005] Special table 2004-508836 publication Special table 2002-524090 publication
[0006] Emser et al., "Extension of Mitogenome Enrichment Based on Single Long-Range PCR: mtENAs and Putative Mitochondrial-Derived Peptides of Five Rodent Hibernators", Front. Genet., 12 December 2021:685806Kneubehl et al., "Amplification and sequencing of entire tick mitochondrial genomes for a phylogenomic analysis", Scientific Reports, 12: 19310, 2022Karin et al., "Highly-multiplexed and efficient long-amplicon PacBio and Nanopore sequencing of hundreds of full mitochondrial genomes", BMC Genomics, 24:229, 2023
[0007] However, the method described in Patent Document 1 requires the design of a relatively large number of primer sequences corresponding to the number of mitochondrial genome fragments to be amplified, making it difficult to design primer sequences that are applicable to a wide range of biological species, and is therefore considered to be difficult to satisfy the first condition. Furthermore, the methods described in Non-Patent Documents 1 to 3 require the use of unfragmented, circular mitochondrial DNA samples to amplify relatively long DNA fragments, making it difficult to satisfy the second condition. Furthermore, the method described in Patent Document 2 amplifies mitochondrial DNA by amplifying all DNA fragments throughout the genome, but also amplifies DNA other than mitochondrial DNA, making the overall operation more complicated and increasing costs and the burden of data processing, making it difficult to satisfy the third condition. Therefore, there has been a need for a technology that uses primers that are highly universal and can be used across a fairly wide range of biological species, even in samples where mitochondrial DNA is somewhat fragmented, and that can simplify the process of determining the entire mitochondrial DNA base sequence.
[0008] The present disclosure can be realized in the following embodiments. (1) According to one embodiment of the present disclosure, a method for designing a primer set for determining the full-length base sequence of mitochondrial DNA is provided. This primer set design method involves obtaining known full-length mitochondrial DNA sequences from multiple biological species belonging to a specific biological taxonomy, aligning the obtained full-length mitochondrial DNA sequences, and extracting conserved regions from the full-length mitochondrial DNA sequences that meet predetermined criteria for the degree of base conservation. Candidate primer sequences are set using the base sequences of each of the extracted conserved regions. When the candidate primer sequences are combined and used as a primer set to amplify mitochondrial DNA, a combination of candidate primer sequences is identified as a candidate primer set, which combination results in three to six fragments that cover the full length of mitochondrial DNA and a desired length for the overlap region with adjacent fragments. At least one of the candidate primer sets is selected as the primer set. This primer set design method provides a primer set having a highly conserved base sequence in biological species belonging to a "specific biological taxonomy," which can obtain three to six fragments that cover the full length of mitochondrial DNA. This allows for highly universal primers that can be used across a fairly wide range of biological species. Furthermore, since mitochondrial DNA can be amplified as three to six fragments, even if the mitochondrial DNA used as a PCR template is not circular but partially degraded, it is possible to amplify the desired DNA fragment, and to obtain DNA fragments of a length suitable for base sequence analysis using a long-read sequencer. (2) In the above-described method for designing a primer set, the specific biological classification may be a class in the taxonomic hierarchy of organisms, or a taxonomic hierarchy lower than the class. With this configuration, it is possible to determine the full-length mitochondrial DNA sequence of various organisms belonging to a class or a taxonomic hierarchy lower than the class through a simple process using a common primer set.(3) In the method for designing a primer set according to the above aspect, the specific biological classification may be Mammalia or Aves. This configuration enables the determination of full-length mitochondrial DNA sequences for various organisms belonging to Mammalia or Aves through a simple process using a common primer set. (4) In the method for designing a primer set according to the above aspect, the candidate primer sets may be identified so that the length of the overlapping region is 500 bases or more and 4,000 bases or less. This configuration increases the likelihood of selecting all available primer sequences and identifying appropriate candidate primer sets. (5) In the method for designing a primer set according to the above aspect, the candidate primer sets may be identified so that the length of the overlapping region is 500 bases or more and 3,000 bases or less. This configuration enables the use of samples containing mitochondrial DNA from multiple different organisms belonging to a common specific biological classification as amplification templates, allowing the amplified DNA fragments to be distinguished by biological species based on sequence differences in the overlapping region, and the sequences of the amplified DNA fragments to be spliced together for each biological species. (6) In the method for designing a primer set of the above aspect, the primer set may be specified so that, when mitochondrial DNA is amplified using the primer set, four fragments covering the entire length of the mitochondrial DNA are obtained. This configuration can improve the accuracy of amplifying the desired DNA fragment and determining the full-length mitochondrial DNA sequence. (7) In the method for designing a primer set of the above aspect, the primer set may be selected from the candidate primer sets so that it is composed of a primer pair determined to have obtained the desired mitochondrial DNA fragment based on the amplification results, using each of the primer pairs comprising a pair of primers constituting the candidate primer set. This configuration can improve the accuracy of amplifying the desired DNA fragment and determining the full-length mitochondrial DNA sequence.(8) In the method for designing a primer set according to the above aspect, the candidate primer sequences may be set by identifying highly conserved sequences, which are base sequences at sites within each of the conserved regions where base conservation is relatively high, and by selecting as the candidate primer sequences sequences where a first parameter representing the degree of sequence specificity of the highly conserved sequence relative to mitochondrial DNA sequences of species belonging to the specific biological taxonomy is equal to or greater than a predetermined first reference value, and a second parameter representing the degree of sequence specificity of the highly conserved sequence relative to mitochondrial DNA sequences of species not belonging to the specific biological taxonomy is equal to or less than a predetermined second reference value that is smaller than the first reference value. This configuration can further enhance the specificity of the primer set for species belonging to the specific biological taxonomy. (9) In the method for designing a primer set according to the above aspect, the first parameter may be the ratio of the number of biological species having one or fewer mismatched bases when the highly conserved sequence is aligned with the full-length sequences of mitochondrial DNA of the plurality of biological species belonging to the specific biological taxonomy to the total number of the plurality of biological species belonging to the specific biological taxonomy, and the second parameter may be the ratio of the number of biological species having one or fewer mismatched bases when the highly conserved sequence is aligned with the full-length sequences of mitochondrial DNA of a plurality of biological species not belonging to the specific biological taxonomy to the total number of the plurality of biological species not belonging to the specific biological taxonomy, the first reference value being 85% and the second reference value being 15%. This configuration further enhances the specificity of the primer set for biological species belonging to the specific biological taxonomy. (10) According to another aspect of the present disclosure, a primer set for determining the full-length base sequence of mitochondrial DNA is provided.This primer set comprises a first primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 1 and 2 in the Sequence Listing at their 3'-terminus; a second primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 3 and 4 in the Sequence Listing at their 3'-terminus; a third primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 5 and 6 in the Sequence Listing at their 3'-terminus; and a fourth primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 7 and 8 in the Sequence Listing at their 3'-terminus, each of the primers having a length of 100 bases or less. This primer set allows the full-length nucleotide sequences of mitochondrial DNA of various species belonging to the class Mammalia to be determined by using a common primer set. (11) According to yet another aspect of the present disclosure, there is provided a primer set for determining the full-length nucleotide sequence of mitochondrial DNA. This primer set comprises a first primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 9 and 10 in the Sequence Listing at their 3'-terminus; a second primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 11 and 12 in the Sequence Listing at their 3'-terminus; a third primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 13 and 14 in the Sequence Listing at their 3'-terminus; and a fourth primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOS: 15 and 16 in the Sequence Listing at their 3'-terminus, each of the primers having a length of 100 bases or less. This primer set allows the full-length nucleotide sequences of mitochondrial DNA of various species belonging to the class Aves to be determined by using a common primer set. (12) According to yet another aspect of the present disclosure, there is provided a primer set for determining the full-length nucleotide sequence of mitochondrial DNA.This primer set comprises a first primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 130 and 131 in the Sequence Listing at their 3' ends; a second primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 132 and 133 in the Sequence Listing at their 3' ends; a third primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 134 and 135 in the Sequence Listing at their 3' ends; and a fourth primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 136 and 137 in the Sequence Listing at their 3' ends, each primer having a length of 100 bases or less. This primer set allows the full-length nucleotide sequences of mitochondrial DNA from various species belonging to the Mammalia to be determined by using a common primer set. (13) According to yet another aspect of the present disclosure, there is provided a primer set for determining the full-length nucleotide sequence of mitochondrial DNA. This primer set comprises a first primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 138 and 139 in the Sequence Listing at their 3' termini; a second primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 140 and 141 in the Sequence Listing at their 3' termini; a third primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 142 and 143 in the Sequence Listing at their 3' termini; and a fourth primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 144 and 145 in the Sequence Listing at their 3' termini, each primer having a length of 100 bases or less. This primer set allows the full-length nucleotide sequences of mitochondrial DNA of various species belonging to the class Aves to be determined by using a common primer set. (14) According to yet another aspect of the present disclosure, there is provided a method for determining the full-length nucleotide sequence of mitochondrial DNA.This method for determining the full-length nucleotide sequence of mitochondrial DNA includes preparing a DNA sample containing mitochondrial DNA, using the DNA sample as a template and a primer set designed by the method for designing a primer set described in any one of (1) to (9) to amplify DNA fragments, analyzing the nucleotide sequence of each of the amplified DNA fragments, and constructing a full-length mitochondrial DNA sequence using the nucleotide sequences of each of the DNA fragments. This method for determining the full-length nucleotide sequence of mitochondrial DNA enables efficient and easy deciphering of the full-length nucleotide sequence of mitochondrial DNA in various biological species whose full-length mitochondrial DNA sequences have not yet been deciphered. (15) According to yet another aspect of the present disclosure, there is provided a method for determining the full-length nucleotide sequence of mitochondrial DNA. This method for determining the full-length nucleotide sequence of mitochondrial DNA includes preparing a DNA sample containing mitochondrial DNA, using the DNA sample as a template and a primer set described in any one of (10) to (13) to amplify DNA fragments, analyzing the nucleotide sequence of each of the amplified DNA fragments, and constructing a full-length mitochondrial DNA sequence using the nucleotide sequences of each of the DNA fragments. This method for determining the full-length mitochondrial DNA sequence makes it possible to efficiently and easily determine the full-length mitochondrial DNA sequence of various biological species for which the full-length mitochondrial DNA sequence has not yet been determined. (16) In the method for determining the full-length mitochondrial DNA sequence of the above mode, the DNA sample may be prepared as a sample containing mitochondrial DNA derived from multiple types of organisms. This configuration simplifies the operation of determining the full-length mitochondrial DNA sequence of multiple types of organisms. (17) In the method for determining the full-length mitochondrial DNA sequence of the above mode, the DNA sample may be prepared as a sample containing genomic DNA other than mitochondrial DNA in addition to mitochondrial DNA.With this configuration, a wider variety of samples can be used, thereby increasing the degree of freedom in selecting samples containing template DNA for determining the full-length base sequence of mitochondrial DNA. The present disclosure can be realized in various forms other than those described above, such as an apparatus for designing a primer set for determining the full-length base sequence of mitochondrial DNA, a computer program for causing a computer to execute a method for designing a primer set for determining the full-length base sequence of mitochondrial DNA, a storage medium for storing the program, a database in which the full-length base sequences of mitochondrial DNA are registered, and a method for analyzing environmental DNA using the database.
[0009] 1 is a functional block diagram showing the configuration of a design device. A flowchart showing a method for designing a primer set. A schematic diagram showing an example of the configuration of a candidate primer set. A diagram showing an example of a primer set for Mammalia. A diagram showing an example of a primer set for Aves. A diagram showing an example of a primer set for Mammalia. A diagram showing an example of a primer set for Aves. A flowchart showing a method for determining the full-length base sequence of mitochondrial DNA. A diagram showing "candidate primer sequences" obtained for Mammalia. A diagram showing "candidate primer sequences" obtained for Aves. A diagram showing the results of an investigation into the sensitivity of matching of candidate primer sequences. A diagram showing the results of an investigation into the sensitivity of matching of candidate primer sequences. A diagram showing the results of amplification using a primer set for Mammalia. A diagram showing the results of amplification using a primer set for Aves. A diagram showing the results of searching a constructed sequence against a database. A diagram showing the results of searching a constructed sequence against a database. A diagram showing the results of searching a constructed sequence against a database. A diagram showing the "candidate primer set" for Aves, which has the highest degree of conservation. 19 is an explanatory diagram showing the results of amplification using the candidate primer sets in Figure 18. An explanatory diagram showing a primer set for Mammalia that yields two types of fragments. An explanatory diagram showing a primer set for Mammalia that yields three types of fragments. An explanatory diagram showing the results of amplification using a primer set for two types of fragments. An explanatory diagram showing the results of amplification using a primer set for three types of fragments. An explanatory diagram showing the process of narrowing down highly conserved sequences. An explanatory diagram showing the process of narrowing down highly conserved sequences. An explanatory diagram showing "candidate primer sequences" set for Mammalia. An explanatory diagram showing "candidate primer sequences" set for Aves. An explanatory diagram showing the sequences of each primer (Pr1 to Pr6). An explanatory diagram showing the results of amplification using a primer pair. An explanatory diagram showing the primer pair and the results of the amplification reaction together. An explanatory diagram showing the results of searching the constructed sequence against a database. An explanatory diagram showing the results of searching the constructed sequence against a database.
[0010] A. Primer Set Design Apparatus: FIG. 1 is a block diagram functionally illustrating the configuration of a primer set design apparatus 10 (hereinafter also referred to as "design apparatus 10") for determining mitochondrial DNA sequences according to an embodiment of the present disclosure. The design apparatus 10 includes a CPU 110, a storage unit 120, a RAM 130, an input interface 140, an output interface 150, and a communication interface 160. These components are interconnected via a bus. The CPU 110 controls the overall operation of the design apparatus 10 by loading a program 122 stored in the storage unit 120 into the RAM 130. An operation unit 170 such as a keyboard or mouse is connected to the input interface 140, and a display unit 180 such as a liquid crystal display is connected to the output interface 150.
[0011] The CPU 110 executes a program 122 stored in the storage unit 120 to realize the functions of an acquisition unit 111 , an extraction unit 112 , a setting unit 113 , an identification unit 114 , and a selection unit 115 .
[0012] The acquisition unit 111 acquires known full-length mitochondrial DNA sequences of multiple biological species belonging to a specific biological taxonomy. The extraction unit 112 aligns the full-length mitochondrial DNA sequences acquired by the acquisition unit 111 and extracts, from the full-length mitochondrial DNA sequences, conserved regions that satisfy predetermined criteria indicating the degree of base conservation. The setting unit 113 sets multiple candidate primer sequences using the base sequences of each conserved region extracted by the extraction unit 112. The identification unit 114 combines the candidate primer sequences set by the setting unit 113 to identify candidate primer sets. The selection unit 115 selects at least one of the candidate primer sets identified by the identification unit 114 as a primer set. The operation related to the design of a primer set will be described in detail later.
[0013] The storage unit 120 may be, for example, a hard disk, a storage medium, a nonvolatile memory, a storage device configured with a nonvolatile memory (SSD), etc. The storage unit 120 stores a program 122.
[0014] B. Primer Set Design Method: Figure 2 is a flowchart showing a primer set design method for determining a mitochondrial DNA sequence according to an embodiment of the present disclosure. The primer set design method shown in Figure 2 will be described below based on an embodiment in which the method is performed using the design device 10 shown in Figure 1. However, the primer set design method of this embodiment may also be performed using a device other than the design device 10. The primer set designed by the design method shown in Figure 2 is a primer set for determining the full-length base sequence of mitochondrial DNA (hereinafter also referred to as "full-length mitochondrial DNA sequence"), and is used to determine the full-length mitochondrial DNA sequence of a biological species belonging to a specific biological taxonomy.
[0015] When designing a primer set, the acquisition unit 111 first acquires known full-length mitochondrial DNA sequences of multiple biological species belonging to a specific biological taxonomy (step T100). Here, the "specific biological taxonomy" refers to a group of relatively closely related organisms for which a primer set designed by the primer set design method of this embodiment is commonly used as a universal primer to determine full-length mitochondrial DNA sequences. Specifically, the "specific biological taxonomy" can be, for example, a "class" of biological taxonomy, or a taxonomic level lower than the class (e.g., "order," "family," or "genus"). For example, in the classes Mammalia and Aves, full-length mitochondrial DNA sequences are relatively highly conserved between species. Therefore, by setting the "specific biological taxonomy" to a "class," it is possible to design a primer set as a universal primer that can be used in common across a wider range of biological species. Furthermore, for example, in the class Pisces, full-length mitochondrial DNA sequences are relatively less conserved between species. Therefore, by setting the "specific biological taxonomy" to a level lower than the class, it is possible to improve the accuracy of amplifying mitochondrial DNA fragments using the designed primer set. By defining the "specific biological classification" as a higher taxonomic level than "species," such as "class," "order," "family," or "genus," it becomes possible to design a primer set that can be used across different species.
[0016] The number of biological species from which full-length mitochondrial DNA sequences are acquired in step T100 is preferably, for example, 800 or more, more preferably 1000 or more, and even more preferably 1200 or more, from the viewpoint of increasing the accuracy of amplifying mitochondrial DNA fragments using the designed primer set. Alternatively, the number may be, for example, 3000 or less, or even 2000 or less, as long as the data processing load is tolerable. It is sufficient to acquire full-length mitochondrial DNA sequences from as many biological species as possible. To acquire full-length mitochondrial DNA sequences, a publicly available database containing information on full-length mitochondrial DNA sequences for various biological species may be appropriately used. For example, the acquisition unit 111 can access an external database 190 via the communication interface 160 and the Internet to perform the acquisition operation in step T100. Alternatively, if the database to be used is stored in the memory unit 120, the acquisition unit 111 can access the memory unit 120 to perform the acquisition operation in step T100.
[0017] Next, the extraction unit 112 aligns the full-length mitochondrial DNA sequences of multiple biological species belonging to the same specific biological taxonomy obtained in step T110, and extracts from the aligned full-length mitochondrial DNA sequences "conserved regions," which are regions that satisfy predetermined criteria indicating the degree of base conservation (step T110). The alignment of full-length mitochondrial DNA sequences in step T110 can be performed using conventionally known alignment algorithms, such as MAFFT, ClustalW, MUSCLE, ClustalOmega, and Kalign. The degree of conservation at each position (each base) in the aligned mitochondrial DNA can be evaluated, for example, by scoring the Shannon information as the degree of conservation. The formula for expressing the degree of base conservation using the Shannon information is shown below as equation (1).
[0018]
[0019] In formula (1), I represents the degree of base conservation, and pk represents the probability that each of the bases A, T, G, and C appears at each base position. The second term on the right side is the Shannon information at the base position of interest. That is, it represents the bias in the frequency of occurrence of each base, with a minimum value of 0 (when only one type of base appears) and a maximum value of 2 (when four types of bases appear with equal probability). For example, when only A appears at the base position of interest, that is, when the bases are completely conserved between sequences, the second term on the right side is 0. By subtracting the value of this second term on the right side from the maximum value of 2, the entire right side shows a higher value when bases are conserved more biasedly toward one type of base. That is, the value on the right side is an index indicating the degree of base conservation.
[0020] The "conserved region" extracted in step T110 can be, for example, a range in which the average degree of conservation of bases at each position of a "specific length sequence" having a specific length in a group of aligned mitochondrial DNA sequences continuously exceeds a predetermined threshold when the operation of calculating the average degree of conservation of bases at each position of the "specific length sequence" is performed while shifting the position of the "specific length sequence" by one base at a time. That is, the above-mentioned "region satisfying a predetermined criterion as a criterion for indicating the high degree of base conservation" can be a "region in which the average degree of conservation of a sequence having a specific length continuously exceeds a threshold." For example, consider a case in which the "specific length sequence" is a sequence of 20 consecutive bases, and the average degree of conservation of each base constituting the "specific length sequence" is sequentially calculated. After a "specific length sequence" whose average degree of conservation exceeds the threshold is found, the operation of calculating the average degree of conservation of the "specific length sequence" while shifting the position by one base is performed 10 times, but the average degree of conservation remains above the threshold, and the average degree of conservation of the "specific length sequence" falls below the threshold by the 11th operation of shifting by one base. In this case, a range of 30 bases in which the average value of the conservation degree of the "specific length sequence" continuously exceeds a threshold value is specified as a "conserved region."
[0021] In step T110, by appropriately setting a "predetermined criterion indicating the degree of base conservation," i.e., a "threshold value for the average degree of conservation in a sequence of a specific length," it is possible to identify a considerable number of "conserved regions" in which bases with a sufficiently high degree of conservation are consecutive. For example, the above-mentioned criterion (threshold value) may be set so that approximately 30 to 70 "conserved regions" are identified.
[0022] The extraction of the "conserved region" in step T110 may be performed by a method different from the above. For example, instead of the Shannon information described above, the proportion of the most frequent base may be used as an index of conservation. The "proportion of the most frequent base" is a value indicating what percentage of the total base is occupied by the most frequently occurring base at each position (each base) of the aligned mitochondrial DNA. For example, if the same base is present in all of the aligned mitochondrial DNA at a specific position, the "proportion of the most frequent base" at that position is 100%. In this case, the operation of calculating the average proportion of the most frequent base for each base in the "specific length sequence" is performed continuously while shifting the position of the "specific length sequence" by one base, and the range in which the average proportion of the most frequent base in the "specific length sequence" continuously exceeds a predetermined threshold may be defined as the "conserved region." Note that, unlike the method using the proportion of the most frequent base as an index of conservation, the method using the Shannon information as an index of conservation is preferable from the viewpoint of being able to evaluate the degree of conservation taking into account bases other than the most frequent base.
[0023] After extracting multiple conserved regions in step T110, the setting unit 113 sets "candidate primer sequences" using the base sequences of each "conserved region" (step T120). The "candidate primer sequences" in step T120 are set by identifying "highly conserved sequences," which are base sequences in each "conserved region" that have a relatively high degree of base conservation. A "highly conserved sequence" can be, for example, a "specific length sequence" (hereinafter referred to as the "highest conserved sequence") with the highest average degree of conservation within each "conserved region." For example, when a "conserved region" of 30 bases is extracted as described above, this "conserved region" contains 11 consecutive "specific length sequences" of 20 bases in length, each of which has an average degree of conservation exceeding a threshold, each of which is shifted by one base. In such a case, the "highest conserved sequence," which is the sequence with the highest average degree of conservation among these 11 "specific length sequences," can be designated as the "highly conserved sequence." The identified "highly conserved sequence" can then be designated as a "candidate primer sequence."
[0024] The identification of the "highly conserved sequence" for designing the "candidate primer sequence" in step T120 may be performed by a method different from the above. For example, it may be inappropriate to use the "most conserved sequence" as the "candidate primer sequence" as is, such as when the GC content of the "most conserved sequence" in the "conserved region" is excessively high or excessively low. In such cases, the "highly conserved sequence" can be obtained by adding approximately 2 to 5 bases adjacent to at least one of the 5' and 3' ends of the "most conserved sequence" in the conserved region. Alternatively, the "highly conserved sequence" can be obtained by deleting approximately 2 to 5 bases from at least one of the 5' and 3' ends of the "most conserved sequence." For example, if there are regions with a relatively high GC content on both the 5' and 3' ends of the "most conserved sequence," it is possible to add several bases of adjacent sequence contained in the "conserved region" to the 5' side of the "most conserved sequence" and delete several bases of sequence from the 3' end of the "most conserved sequence." When bases are added to one end of the "most conserved sequence" and bases are deleted from the other end, the length of the added bases and the length of the deleted bases may be the same or different.
[0025] The "candidate primer sequences" set in step T120 are candidate primers to be used when amplifying mitochondrial DNA fragments using the mitochondrial DNA of the aforementioned "organism belonging to a specific biological taxonomy" as a template. Therefore, they are required to sufficiently bind (anneal) to the mitochondrial DNA of the aforementioned "organism belonging to a specific biological taxonomy." From this perspective, it is desirable that the "candidate primer sequences" be able to bind to the mitochondrial DNA of a biological species belonging to the aforementioned "specific biological taxonomy" with no more than one mismatched base.
[0026] As described above, in order to ensure that the "candidate primer sequence" sufficiently binds to the mitochondrial DNA of a biological species belonging to a "specific biological taxonomy," the "candidate primer sequence" may be set by further screening the "highly conserved sequence" rather than using the "highly conserved sequence" as the "candidate primer sequence." Specifically, for example, the "candidate primer sequence" may be set by narrowing down the "highly conserved sequence" to those for which the "first parameter," which represents the level of sequence specificity for the mitochondrial DNA sequence of a biological species belonging to a "specific biological taxonomy," is equal to or greater than a predetermined "first reference value." The "candidate primer sequence" may then be set by further narrowing down the "highly conserved sequence" to those for which the "second parameter," which represents the level of sequence specificity for the mitochondrial DNA sequence of a biological species not belonging to a "specific biological taxonomy," is equal to or less than a predetermined "second reference value," which is smaller than the above-mentioned "first reference value."
[0027] The "first parameter" can be, for example, the ratio of the number of biological species in which the number of mismatched bases is one or less when a "highly conserved sequence" is aligned with the full-length sequences of mitochondrial DNA of multiple biological species belonging to a "specific biological taxonomy" to the total number of multiple biological species belonging to the "specific biological taxonomy." The "second parameter" can be the ratio of the number of biological species in which the number of mismatched bases is one or less when a "highly conserved sequence" is aligned with the full-length sequences of mitochondrial DNA of multiple biological species not belonging to the "specific biological taxonomy," to the total number of multiple biological species not belonging to the "specific biological taxonomy." The "first reference value" can be, for example, 85%. The "second reference value" can be, for example, 15%.
[0028] In the above, the "first parameter" and the "second parameter" are set based on the number of mismatches when the "highly conserved sequence" is aligned with the full-length sequence of mitochondrial DNA of a biological species that may or may not belong to a "specific biological taxonomy," but a different index may also be used. For example, the "first parameter" and the "second parameter" may be set based on the number of gaps when the "highly conserved sequence" is aligned with the full-length sequence of mitochondrial DNA of a biological species that may or may not belong to a "specific biological taxonomy," or the longest number of consecutive bases aligned without mismatches or gaps.
[0029] The length of the "candidate primer sequence" is not limited to 20 bases, and can be any length within the range of, for example, 17 to 25 bases. However, from the viewpoint of ensuring the accuracy of amplifying the desired fragment by PCR, it is necessary to ensure a sufficient primer length. Furthermore, from the viewpoint of making the primer set to be designed a universal primer that can be used in common among organisms belonging to the aforementioned "specific biological classification," it is desirable to keep the primer length small. Therefore, 20 bases can be mentioned as a particularly desirable length for the "candidate primer sequence."
[0030] After the "candidate primer sequences" are set in step T120, the identification unit 114 then combines these "candidate primer sequences" to identify "candidate primer sets" (step T130). The "candidate primer sets" are identified by combining multiple "candidate primer sequences" so that when mitochondrial DNA is amplified using the identified "candidate primer sets" as primer sets, three to six fragments covering the entire length of mitochondrial DNA are obtained and the length of the overlap region with adjacent fragments is the desired length.
[0031] FIG. 3 is an explanatory diagram schematically illustrating an example of the configuration of a "primer set candidate" identified in step T130. FIG. 3 shows an example of a "primer set candidate" that yields four fragments covering the entire length of mitochondrial DNA (mt). The primer set candidate shown in FIG. 3 is composed of eight primers, primers α1, α2, β1, β2, γ1, γ2, δ1, and δ2, selected from the multiple "primer candidate sequences" established in step T120. That is, the primer set candidate shown in FIG. 3 includes a first primer pair composed of a forward primer α1 and a reverse primer α2, a second primer pair composed of a forward primer β1 and a reverse primer β2, a third primer pair composed of a forward primer γ1 and a reverse primer γ2, and a fourth primer pair composed of a forward primer δ1 and a reverse primer δ2. A "fragment A" is obtained by performing PCR using a first primer pair consisting of primers α1 and α2, a "fragment B" is obtained by performing PCR using a second primer pair consisting of primers β1 and β2, a "fragment C" is obtained by performing PCR using a third primer pair consisting of primers γ1 and γ2, and a "fragment D" is obtained by performing PCR using a fourth primer pair consisting of primers δ1 and δ2. Note that while FIG. 3 schematically shows fragments A to D as being the same length and arranged at equal intervals, FIG. 3 does not accurately represent the dimensional ratios of each part.
[0032] In the fragments obtained when mitochondrial DNA is amplified using a candidate primer set, adjacent fragments have overlapping regions. Figure 3 shows how "Fragment A" and "Fragment B" form an "overlapping region AB," "Fragment B" and "Fragment C" form an "overlapping region BC," "Fragment C" and "Fragment D" form an "overlapping region CD," and "Fragment D" and "Fragment A" form an "overlapping region DA."
[0033] When identifying "primer set candidates" in step T130, as described above, "primer candidate sequences" are combined so as to obtain the desired length of the overlap region. The sequence of each overlap region is used to precisely join the sequences of each DNA fragment in the correct order when mitochondrial DNA is amplified using a primer set to obtain each fragment, and then the sequences of each DNA fragment are joined to obtain the full-length mitochondrial DNA sequence after sequencing each DNA fragment. In other words, by overlapping corresponding overlap regions, the full-length mitochondrial DNA sequence can be easily constructed.
[0034] Here, in each overlapping region, highly conserved "candidate primer sequences" are located at both ends, but in the middle region sandwiched between these "candidate primer sequences," there is a region (hereinafter also referred to as a "non-conserved region") that is less conserved than the "candidate primer sequences." That is, as described above, "candidate primer sequences" are set from each of the conserved regions, which are regions of consecutive highly conserved bases, and therefore, between the spaced apart "candidate primer sequences," there are regions of relatively low conserved bases where the consecutive highly conserved bases are interrupted. Such "non-conserved regions" can be said to be regions in which the base sequences differ to some extent between mitochondrial DNAs derived from different organisms, even if they belong to a "specific biological taxonomy."
[0035] Therefore, even when mitochondrial DNA fragments are amplified using the same primer set, fragments of the same species can be distinguished by species by comparing the sequences of the "non-conserved regions" within the overlapping regions of the same fragments, as in "Fragment A" shown in Figure 3. Therefore, even when a sample containing mitochondrial DNA from multiple organisms belonging to a common biological taxonomy but different from each other is used as an amplification template, fragments of the same species can be distinguished by species as described above, and fragments from the same organisms having overlapping regions with the same sequence can be joined together. This makes it possible to appropriately construct full-length mitochondrial DNA sequences for multiple organisms for each species. To distinguish the species from which template DNA is derived based on sequence differences in the overlapping regions, it is necessary to ensure that the overall length of the overlapping region is sufficient to include a sufficient length of non-conserved region in the overlapping region. From this perspective, it is desirable to specify "primer set candidates" so that the length of each overlapping region is 500 bases or more.
[0036] Furthermore, if the overlapping region is made excessively long, the length of each fragment obtained by PCR will consequently be longer. The longer the length of each fragment, the more likely it is that the accuracy of sequencing each fragment will decrease, and the sequencing operation will become more complicated. Furthermore, the longer the length of each fragment, the more difficult it becomes to obtain read data including the full length of the fragment, and the greater the data processing load. Therefore, the length of each overlapping region is preferably 4,000 bases or less from the viewpoint of ensuring the opportunity to select all available primer sequences and identify appropriate primer set candidates, and is preferably 3,000 bases or less from the viewpoint of suppressing the PCR fragment length, improving sequencing accuracy, and enhancing the effect of making it easier to obtain read data.
[0037] As described above, the longer the length of the fragment to be amplified, the more likely it is that the accuracy of amplification will decrease. Furthermore, if the mitochondrial DNA contained in the DNA sample used as the amplification template is not a complete circular DNA but contains fragments that are fragmented to some extent, the longer the length of the fragment to be amplified, the more likely it is that amplification will not be sufficient. Therefore, from the perspective of reducing the length of the fragments obtained by PCR and improving the amplification accuracy and amplification efficiency, in this embodiment, the number of types of fragments obtained by the "primer set candidate" is three or more, and preferably four or more.
[0038] In contrast, increasing the number of fragments obtained by amplification shortens the length of each fragment, thereby improving the accuracy of amplification and sequencing. However, excessively shortening the length of each fragment makes it difficult to ensure sufficient length for overlapping regions containing regions with relatively low conservation, which can make it difficult to construct a full-length mitochondrial DNA sequence from the amplified fragments. Furthermore, a full-length mitochondrial DNA sequence contains a certain number of consecutive regions with relatively low conservation. In such regions, it becomes difficult to set overlapping regions with candidate primer sequences at both ends, making it difficult to suppress the maximum fragment length, which may result in undesirable variations in fragment length. Furthermore, shortening the fragment length and increasing the number of fragment types increases the number of experimental steps required to obtain a full-length mitochondrial DNA sequence. Therefore, in this embodiment, the number of fragment types obtained by the "primer set candidate" is set to six or less, and preferably five or less. From the perspective of ensuring amplification accuracy and efficiency while reducing the number of experimental steps required to obtain a full-length mitochondrial DNA sequence, it is preferable that the number of fragment types obtained by the "primer set candidate" be four.
[0039] The length of each fragment obtained by amplification using the "primer set candidate" must be somewhat longer than the sum of the overlapping regions at both ends of the fragment. Therefore, even if the length of the overlapping region is, for example, about 500 bases, the length of the fragment is preferably 2000 bases or more, and more preferably 4000 bases or more in order to allow for longer overlapping regions. Furthermore, the length of each fragment is preferably 9000 bases or less, more preferably 8000 bases or less, in order to prevent a decrease in sequencing accuracy due to the length of each fragment, the complexity of sequencing operations, and an increase in data processing load.
[0040] In step T130, a full search is performed on the combinations of the multiple candidate primer sequences set in step T120 to obtain three to six fragments that cover the entire length of mitochondrial DNA, and primer combinations (eight primers: α1, α2, β1, β2, γ1, γ2, δ1, and δ2 in the example of FIG. 3 ) that are arranged to obtain a desired length as the overlap region with adjacent fragments are extracted and identified as candidate primer sets. The number of primers required to obtain three to six fragments is six, eight, ten, or twelve. In contrast, as described above, a relatively large number of sequences are set as candidate primer sequences. Therefore, by performing a full search on the combinations of the multiple candidate primer sequences, it is relatively easy to extract combinations of the desired number of candidate primer sequences that form overlap regions of the desired length.
[0041] After identifying the candidate primer sets in step T130, the selection unit 115 selects at least one of these candidate primer sets as a primer set (step T140). If only one candidate primer set is identified in step T130, the candidate primer set may be selected as the primer set. If multiple candidate primer sets are identified in step T130, all of these multiple candidate primer sets may be selected as primer sets, or a subset of the multiple candidate primer sets may be selected as primer sets. The selection operation in step T140 is performed according to predetermined criteria, such as selecting a primer set from among the multiple candidate primer sets by excluding a candidate primer set that exhibits particularly large variations in fragment length in a combination of three to six fragments obtained by amplification. The selection unit 115 can output the selected primer set to the display unit 180 via the output interface 150 for display.
[0042] Alternatively, the CPU 110 may not implement the function of the selection unit 115, and the primer set candidates identified by the identification unit 114 may be output from the design device 10 to the display unit 180 or the like, and the operation of selecting a primer set from the primer set candidates may be executed outside the design device 10. A method for selecting some of the multiple primer set candidates as primer sets outside the design device 10 includes, for example, a method in which a preliminary experiment is performed using each of the primer set candidates to perform PCR using mitochondrial DNA derived from an organism belonging to the aforementioned "specific biological classification" as a template, and the combination that produces the best results is selected as the primer set. Specifically, for each primer set candidate, PCR is performed using each primer pair constituting the primer set candidate, and the primer set candidate constituted by the primer pair that can obtain a fragment of the desired length is selected as the primer set.
[0043] Each primer constituting the primer set selected in step T140 may contain a different sequence in addition to the "candidate primer sequence" set based on the "highly conserved sequence" set within each "conserved region" described in step T120, i.e., a sequence that anneals to mitochondrial DNA, which serves as a template during PCR amplification. Specifically, a "5'-side additional sequence" may be added to the 5' end of the aforementioned "candidate primer sequence," which is a sequence of up to approximately 80 bases, preferably 60 bases or less, that is different from the adjacent sequence within the aforementioned "conserved region." When used as a primer in PCR, the "5'-side additional sequence" is not expected to anneal to mitochondrial DNA, but addition of such a sequence is permitted at the 5' end, which is different from the 3' end where the DNA elongation reaction proceeds. As described above, the "candidate primer sequence" is approximately 20 bases (e.g., approximately 16 to 27 bases, preferably approximately 18 to 25 bases), and therefore the length of each primer may be, for example, 100 bases or less, and preferably 80 bases or less.
[0044] In the embodiments of the primer set design method described above, some of the hardware-implemented components may be replaced with software, and conversely, some of the software-implemented components may be replaced with hardware. Furthermore, when some or all of the functions of the present disclosure are implemented by software, the software (computer program) may be provided in a form stored on a computer-readable storage medium. The term "computer-readable storage medium" is not limited to portable storage media such as floppy disks and CD-ROMs, but also includes internal storage devices within a computer, such as various RAMs and ROMs, and external storage devices fixed to a computer, such as a hard disk. In other words, the term "computer-readable storage medium" has a broad meaning, including any storage medium that can store data not temporarily but fixedly.
[0045] C. Primer Set: Figure 4 is an explanatory diagram showing, as an example of the primer set of this embodiment, an example of a primer set designed by the primer set design method described above and that can be used to determine the full-length sequence of mitochondrial DNA of Mammalia. Figure 5 is an explanatory diagram showing, as another example of the primer set of this embodiment, an example of a primer set designed by the primer set design method described above and that can be used to determine the full-length sequence of mitochondrial DNA of Aves. Figure 6 is an explanatory diagram showing, as yet another example of the primer set of this embodiment, an example of a primer set designed by the primer set design method described above and that can be used to determine the full-length sequence of mitochondrial DNA of Mammalia. Figure 7 is an explanatory diagram showing, as yet another example of the primer set of this embodiment, an example of a primer set designed by the primer set design method described above and that can be used to determine the full-length sequence of mitochondrial DNA of Aves.
[0046] 4 to 7 show primer sets composed of four primer pairs (primer pairs 1 to 4) that yield four fragments covering the entire length of mitochondrial DNA. In FIGS. 4 to 7, the forward primer of each primer pair is indicated by adding "for" to the primer name, and the reverse primer is indicated by adding "rev" to the primer name. The conditions for designing each primer set will be described in detail below. In the four fragments obtained by amplifying mitochondrial DNA from Mammalia using the primer set shown in FIG. 4, the lengths of the overlapping regions (overlapping regions AB to DA in FIG. 3) are approximately 1,300 bases, approximately 3,000 bases, approximately 1,100 bases, and approximately 1,200 bases, respectively. In the four fragments obtained by amplifying mitochondrial DNA from Aves using the primer set shown in FIG. 5, the lengths of the overlapping regions are approximately 1,600 bases, approximately 1,700 bases, approximately 3,000 bases, and approximately 2,000 bases, respectively. In the four fragments obtained by amplifying mitochondrial DNA of Mammalia using the primer set shown in Figure 6, the lengths of the overlapping regions (overlapping regions AB to DA in Figure 3) are approximately 3,100 bases, 3,000 bases, 1,300 bases, and 600 bases, respectively. In the four fragments obtained by amplifying mitochondrial DNA of Aves using the primer set shown in Figure 7, the lengths of the overlapping regions are approximately 600 bases, 3,000 bases, 2,500 bases, and 1,100 bases, respectively.
[0047] In the primer sets shown in Figures 4 to 7, at least one of the primers constituting each primer set may further have an additional sequence of about 60 bases or less in length at the 5' end.
[0048] D. Mitochondrial DNA Sequencing Method: Figure 8 is a flowchart showing the method for determining the full-length base sequence of mitochondrial DNA according to this embodiment. This method for determining the full-length sequence of mitochondrial DNA is carried out, for example, to determine the full-length mitochondrial DNA sequence of a species whose full-length mitochondrial DNA sequence has not yet been registered, in order to enrich a database in which nucleic acid sequence information from various species is registered.
[0049] When determining the full-length base sequence of mitochondrial DNA, first, a DNA sample containing mitochondrial DNA is prepared (step T200). The DNA sample prepared in step T200 may be a sample containing only mitochondrial DNA derived from one type of organism, or may be a sample containing mitochondrial DNA derived from multiple types of organisms. Furthermore, the DNA sample prepared in step T200 may be a sample containing only mitochondrial DNA as DNA, or may be a sample containing mitochondrial DNA as well as genomic DNA other than mitochondrial DNA. Furthermore, the mitochondrial DNA contained in the DNA sample prepared in step T200 may exist as circular DNA, or may be at least partially fragmented.
[0050] After preparing a DNA sample in step T200, PCR amplification of DNA fragments is performed using the prepared DNA sample as a template and a primer set designed by the primer set design method described with reference to FIG. 2 (step T210). That is, mitochondrial DNA fragments contained in the DNA sample are amplified. This results in the amplification of three to six types of fragments covering the entire length of the mitochondrial DNA of a species belonging to a specific biological taxonomy whose known full-length mitochondrial DNA sequence was obtained in step T100 of FIG. 2 . In step T210, PCR reactions for amplification can be performed separately for each primer pair constituting the primer set, from the viewpoint of efficiently amplifying the desired types of fragments. The number of cycles in the PCR performed in step T210 can be, for example, 25 to 35. The annealing temperature in the PCR can be, for example, 50°C to 60°C.
[0051] Examples of primer sets used in step T210 and designed by the primer set design method shown in Fig. 2 include the primer sets shown in Figs. 4 to 7. By using the primer set shown in Fig. 4 or 6, four types of fragments of mitochondrial DNA of the class Mammalia present in a DNA sample, which cover the entire length of the mitochondrial DNA, are amplified. By using the primer set shown in Fig. 5 or 7, four types of fragments of mitochondrial DNA of the class Aves present in a DNA sample, which cover the entire length of the mitochondrial DNA, are amplified.
[0052] After amplifying the mitochondrial DNA fragments in step T210, the base sequence of each amplified DNA fragment is analyzed (step T220). As previously described, the primer set of this embodiment can be designed so that the length of the mitochondrial DNA fragments obtained by amplification is approximately 2,000 to 8,000 bases. In this embodiment, the base sequence of each fragment is analyzed using a long-read sequencer (e.g., MinION by Oxford Nanopore Technologies or Sequel by PacBio; MinION and Sequel are registered trademarks). While conventional short-read sequencers, which have been widely used in the past, can read sequences of approximately 200 bases at a time, long-read sequencers are devices capable of reading DNA sequences of several thousand bases or even 10,000 bases in length in a continuous sequence. Therefore, by using a long-read sequencer, it is possible to analyze the base sequence of each fragment amplified in step T210 using longer reads rather than analyzing short, discrete fragments of about 300 bases, and the number of steps required for experimental operations to analyze the base sequence can be significantly reduced. Furthermore, by using a long-read sequencer, it is possible to analyze complex structures such as long repetitive sequences that are difficult to analyze using a short-read sequencer.
[0053] After analyzing the base sequence of each DNA fragment amplified in step T220, the full-length sequence of mitochondrial DNA is constructed using the base sequences of each DNA fragment (step T230). Specifically, the base sequence of each fragment is first determined from the shared sequence between the reads obtained by the long-read sequencer. Then, the corresponding overlapping regions at both ends of each fragment are overlapped to construct the full-length sequence of the circular mitochondrial DNA.
[0054] According to the primer set design method of the present embodiment configured as described above, it is possible to create a primer set having a highly conserved base sequence among biological species belonging to a "specific biological taxonomy," which can generate three to six fragments covering the entire length of mitochondrial DNA. Therefore, it is possible to obtain highly common primers that can be used across a fairly wide range of biological species. Furthermore, because mitochondrial DNA can be amplified as three to six fragments, even if the mitochondrial DNA used as a PCR template is not circular but is partially degraded, it is possible to amplify the desired DNA fragments and obtain DNA fragments of a length suitable for base sequence analysis using a long-read sequencer. Then, by analyzing the base sequence of each fragment obtained in this manner, the full-length sequence of mitochondrial DNA of various biological species belonging to a "specific biological taxonomy" can be determined in a simple process.
[0055] In this embodiment, when determining the full-length sequence of mitochondrial DNA from DNA fragments obtained by amplification using a primer set, the full-length sequence of mitochondrial DNA can be easily constructed by overlapping the overlapping regions at both ends of the DNA fragments. Even if the sample used as an amplification template contains mitochondrial DNA from multiple organisms, as described above, the sequences of the non-conserved regions in the overlapping regions can be compared to distinguish the biological species from which the DNA fragments originate, and the full-length sequences of mitochondrial DNA for each biological species can be constructed separately.
[0056] Furthermore, according to the method for designing a primer set of this embodiment, the number of types of fragments obtained can be adjusted within a range of 3 to 6, depending on the length distribution of DNA in a sample used as a PCR template. For example, if the DNA distribution in a sample is such that a high proportion of relatively short fragments is present, the number of types of fragments obtained can be set to a larger number. This makes it possible to optimize the balance between PCR amplification efficiency, amplification accuracy, and simplification of the sequencing process.
[0057] Furthermore, by using a primer set that can obtain 3 to 6 fragments that cover the entire length of mitochondrial DNA, the length of each DNA fragment obtained by amplification can be limited to, for example, approximately 8,000 bases or less. Therefore, even when a long-read sequencer is used, it becomes possible to use a DNA polymerase with higher reading accuracy, such as that used in general short-read sequencers, thereby improving the accuracy of determining the full-length sequence of mitochondrial DNA of a target organism.
[0058] As described above, the primer set designed by the primer set design method described above can be commonly used when determining the full-length sequences of mitochondrial DNA of various biological species belonging to the same "specific biological taxonomy." Therefore, when investigating the full-length sequences of mitochondrial DNA of various biological species, the cumbersome process of preparing different primer sets for each biological species can be avoided. Furthermore, since the number of DNA fragments amplified by the primer set is between three and six, the data processing load required to connect the base sequences of each fragment can be reduced. As a result, it becomes possible to efficiently and easily decode the full-length sequences of mitochondrial DNA of various biological species whose full-length sequences have not yet been decoded, thereby efficiently enriching the database. Note that when amplifying mitochondrial DNA fragments using a primer set designed by the primer set design method of this embodiment to derive the full-length sequence of mitochondrial DNA, the DNA fragments may be amplified by a method other than PCR, such as multiple displacement amplification (MDA).
[0059] Furthermore, in this embodiment, when setting "candidate primer sequences" from "highly conserved sequences," as described above, it is desirable to narrow down the "candidate primer sequences" to sequences for which the first parameter representing the level of sequence specificity of the "highly conserved sequence" relative to mitochondrial DNA sequences of biological species belonging to a "specific biological taxonomy" is equal to or greater than a predetermined first reference value, and the second parameter representing the level of sequence specificity of the "highly conserved sequence" relative to mitochondrial DNA sequences of biological species not belonging to a "specific biological taxonomy" is equal to or less than a predetermined second reference value that is smaller than the first reference value. This makes it possible to further increase the specificity of the primer set for biological species belonging to a "specific biological taxonomy."
[0060] <Design of First Primer Set> Primer sets for determining the full-length mitochondrial DNA sequence were designed for each of the "specific biological classifications" Mammalia and Aves by the primer set design method shown in FIG. 2 .
[0061] (Setting of candidate primer sequences) For each of the classes Mammalia and Aves, known full-length mitochondrial DNA sequences of multiple species were obtained (step T100). Specifically, full-length mitochondrial DNA sequence information for 1,379 species belonging to the class Mammalia and 990 species belonging to the class Aves was obtained from the RefSeq base sequence database of the National Center for Biotechnology Information. Subsequently, for each of the mammals and birds, alignment was performed using the sequence alignment tool MAFFT (Katoh et al., Mol Biol Evol, 30(4):772-780, 2013). The Shannon information for each base in the alignment was calculated and scored as the degree of conservation, and conserved regions meeting the criteria for high base conservation were extracted (step T110).
[0062] In step T110, the average degree of conservation was calculated for a 20-base region (specific length sequence) from one end of the aligned sequence group, and the calculated average value was designated as the "degree of conservation of the region." The same calculation was repeated from one end to the other, shifting the degree of conservation of the 20-base region by one base, and the calculation was compared with a predetermined threshold. The range in which the degree of conservation of the 20-base region continuously exceeded the threshold was identified as a "region that meets the predetermined criteria for indicating the high degree of base conservation," i.e., a "conserved region."
[0063] Then, "candidate primer sequences" were set from each of the identified "conserved regions" (step T120). Here, "highly conserved sequences" were identified using the consensus sequence (highly conserved sequence) of the 20-base region with the highest degree of conservation within each "conserved region," and "candidate primer sequences" were set. Forty "candidate primer sequences" were obtained for Mammalia and 55 for Aves. When setting "candidate primer sequences," as described above, if the 20-base "highly conserved sequence" had an excessively high GC content, bases were added or deleted at least one of the 5' and 3' ends of the "highly conserved sequence" to identify the "highly conserved sequence." Therefore, the length of the set "candidate primer sequences" ranged from 18 to 25 bases. Furthermore, when setting "candidate primer sequences," "highly conserved sequences" that did not meet the conditions generally considered desirable for primers were excluded from the candidates based on thermodynamic parameters. As a specific example, Primer3, a tool for designing PCR primers, was used to exclude those that did not satisfy the condition of "Tm value of 55°C or higher and 65°C or lower."
[0064] FIG. 9 is an explanatory diagram showing the "candidate primer sequences" obtained for Mammalia, and FIG. 10 is an explanatory diagram showing the "candidate primer sequences" obtained for Aves. In FIG. 9, the position of each "candidate primer sequence" in mitochondrial DNA is shown by representing the coordinates on the mitochondrial DNA sequence of Cavia porcellus (guinea pig) (RefSeq ID: NC_000884.1), with both ends of the sequence designated as "Start" and "End." In FIG. 10, the position of each "candidate primer sequence" in mitochondrial DNA is shown by representing the coordinates on the mitochondrial DNA sequence of Prodotiscus insignis (western olive honeyguide) (RefSeq ID: NC_039892.1), with both ends of the sequence designated as "Start" and "End." To determine how many mitochondrial DNA sequences these "primer candidate sequences" could anneal to in the same "specific biological taxonomy," we analyzed them using PrimerProspector (Walters et al., Bioinformatics, 27(8), 2011).
[0065] Fig. 11 is an explanatory bar graph showing the results of investigating the theoretical sensitivity of matching when 40 "candidate primer sequences" set in step T120 are combined (annealed) with each of the mitochondrial DNA sequences of 1,379 species of organisms belonging to the class Mammalia obtained in step T100. Fig. 12 is an explanatory bar graph showing the results of investigating the theoretical sensitivity of matching when 55 "candidate primer sequences" set in step T120 are combined (annealed) with each of the mitochondrial DNA sequences of 990 species of organisms belonging to the class Aves obtained in step T100. In Figs. 11 and 12, the horizontal axis represents the "candidate primer sequences," and the vertical axis represents the proportion of organisms belonging to the classes Mammalia and Aves that combine with the "candidate primer sequences" at a specific matching rate. 11 and 12, the black columns indicate the proportion of species having mitochondrial DNA sequences to which each "candidate primer sequence" binds with zero mismatches, and the gray columns indicate the proportion of species having mitochondrial DNA sequences to which each "candidate primer sequence" binds with one or fewer mismatches. In Figures 11 and 12, the dashed line indicates the point where the proportion of species shown on the vertical axis reaches 90%.
[0066] As shown in Figure 11, 35 of the 40 "candidate primer sequences" set for the class Mammalia were confirmed to bind to the mitochondrial DNA of more than 90% of the species belonging to the class Mammalia with no more than one mismatch. Also, as shown in Figure 12, 49 of the 55 "candidate primer sequences" set for the class Aves were confirmed to bind to the mitochondrial DNA of more than 90% of the species belonging to the class Aves with no more than one mismatch. As described above, it was demonstrated that the "candidate primer sequences" set in step T120 have high sensitivity for the base sequences of the mitochondrial DNA of the target species "belonging to a specific biological taxonomy."
[0067] (Primer Set Design) For the "candidate primer sequences" obtained above, all combinations of "candidate primer sequences" were enumerated, the number of which corresponded to the number of fragments obtained when mitochondrial DNA was amplified using the primer set. The lengths of the fragments produced when the "candidate primer sequences" were used for each combination, as well as the lengths of the overlapping regions with adjacent fragments, were then calculated. Combinations of "candidate primer sequences" that yielded the desired fragment lengths and overlapping region lengths were obtained and identified as "candidate primer sets" (step T130). Specifically, the number of fragments obtained was set to 4, and "candidate primer sets" were identified so that the length of the overlapping region was 500 to 3,000 bases. From these "candidate primer sets," "primer sets" were selected based on factors such as the conservation of the primers constituting the "candidate primer sets" (step T140). When selecting such "primer sets," "candidate primer sets" with highly conserved primers were further used to conduct experiments in which mitochondrial DNA fragments were actually amplified using the "candidate primer sets," and those that yielded favorable results were selected as "primer sets." The sequences of the primers constituting the "primer set" for Mammalia selected in this manner are shown in Figure 4, as described above. The sequences of the primers constituting the "primer set" for Aves selected in this manner are shown in Figure 5, as described above. The results of experiments in which mitochondrial DNA fragments were actually amplified using the "primer set candidates" will be described later.
[0068] <Evaluation of Primer Sets> (Evaluation Targeting Cattle, Pig, Sheep, and Chicken) DNA was extracted from meat sold at a general butcher shop for cattle, pigs, and sheep belonging to the class Mammalia, and chickens belonging to the class Aves, and DNA fragments were amplified using the primer sets shown in Figures 4 and 5. Note that the primer sets shown in Figures 4 and 5, the primer sets shown in Figures 6, 7, 18, 20, and 21 described below, and primer pairs set based on Figure 28 described below include primers whose sequences contain positions indicated by W, R, Y, M, N, etc., indicating that they correspond to multiple types of bases. When amplifying DNA fragments, primers containing a mixture of equal amounts of all possible sequences were used.
[0069] Specifically, genomic DNA was extracted from approximately 25 mg of beef, pork, sheep, and chicken meat using a NucleoSpin Tissue (Machley-Nagel) kit (Step T200). 2x KAPA Hifi HS Ready Mix (Nippon Genetics) and 0.4 μM primer DNA were added to 1 ng of the resulting genomic DNA, and a thermal cycling reaction was performed (Step T210). The thermal cycling conditions were: 1 cycle of 98°C for 3 minutes, 30 cycles of 98°C for 30 seconds, 55°C for 30 seconds, and 72°C for 3 minutes, and 1 cycle of 72°C for 5 minutes.
[0070] Figure 13 is an explanatory diagram showing the results of electrophoresis of the reaction product obtained by performing an amplification reaction using the primer set for Mammalia shown in Figure 4. Figure 14 is an explanatory diagram showing the results of electrophoresis of the reaction product obtained by performing an amplification reaction using the primer set for Aves shown in Figure 5. As shown in Figures 13 and 14, amplification of fragments of the expected size was confirmed for all four primer pairs for three species (bovine, porcine, and ovine) when the primer set for Mammalia was used, and for chicken when the primer set for Aves was used. It was also confirmed that fragments of the expected size were not amplified for biological species belonging to a "specific biological classification" different from the target of the primer set.
[0071] 13 and 14, for samples in which amplification of fragments of the expected size was confirmed (samples in which the "specific biological taxonomy" corresponded between the biological species from which the extracted DNA was derived and the target of the primer set), 25 fmol of each of the four fragments was mixed and sequenced using MinION (Oxford Nanopore Technologies). A sequencing library was constructed using the Ligation Sequencing kit (SQK-LSK109, Oxford Nanopore Technologies) and Native Barcoding Expansion 1-12 (EXP-NBD104, Oxford Nanopore Technologies), and sequencing was performed for 72 hours using an R9.4.1 flow cell (Oxford Nanopore Technologies) (step T220).
[0072] The reads obtained by sequencing contain the adapter sequences used during sequencing. These were removed from the resulting DNA sequences using Porechop (Wick et al., Microb. Genom. 2017; 3(10): e000132). The reads were then filtered by quality score using Chopper (De Coster & Rademakers, Bioinformatics, 2023; 39(5): btad311). The reads were then sorted by fragment type using Amplicon_sorter (Vierstraete & Braeckman, Ecol Evol, 2022). The entire fragment sequence was then determined using the consensus sequence between the reads. The sequences of each fragment were then combined at overlapping regions using Minimus2 (Sommer et al., BMC Bioinformatics, 2007; 8 64), and error correction was performed using Medaka (https: / / github.com / nanoporetech / medaka) to construct the full-length mitochondrial DNA sequence (step T230). To verify whether the full-length mitochondrial DNA sequences obtained for each of the bovine, porcine, ovine, and chicken species were correctly constructed, a search was performed against the National Center for Biotechnology Information sequence database using the search program Blastn (Altschul et al., J Mol Biol, 215(3), 1990).
[0073] Figure 15 is an explanatory diagram showing the results of searching a database for each full-length mitochondrial DNA sequence constructed as described above. In Figure 15, the species with the highest sequence identity (identity) with the constructed full-length mitochondrial DNA sequence is shown as the "Blastn top hit species." Also, Figure 15 shows the length of the constructed full-length mitochondrial DNA sequence as the "Contig Length." As shown in Figure 152, for all samples, the full-length mitochondrial DNA sequence of the same species as the analyzed DNA sample was hit, with a match rate of over 99%. This demonstrates that the use of the primer sets shown in Figures 4 and 5 achieves an error rate comparable to that of conventional base sequence determination methods (e.g., when using an Illumina sequencer or the Sanger method, the error rate is generally around 0.1%). Furthermore, because complex regions that are generally difficult to address using short-read sequencers could be constructed, the method using the primer sets described above demonstrates that it is an excellent method for constructing full-length mitochondrial DNA sequences, even if they contain complex regions.
[0074] (Evaluation of Other Organisms) To confirm that the primer set designed by the primer set design method according to the present disclosure can be widely applied to various biological species belonging to the target "specific biological classification (here, the class Mammalia or the class Aves)," the primer set was evaluated using DNA extracted from other biological species. Specifically, for the following biological species belonging to the class Mammalia: Sika deer, Brown bear, Steller sea lion, Racoon, Japanese badger, Wild boar, and European rabbit; and the following species belonging to the class Aves: Japanese green pigeon, Common kestrel, Wild duck, Green pheasant, and Japanese grossbeak, full-length mitochondrial DNA sequences were constructed using tissues obtained from meat or carcass samples. Extraction of genomic DNA from tissues, amplification of DNA fragments using the primer sets shown in Figures 4 and 5, sequencing, and construction of the full-length mitochondrial DNA sequence were performed in the same manner as in the evaluation of cattle, pigs, sheep, and chickens described above. To examine whether the obtained full-length mitochondrial DNA sequence was correctly constructed, a search was performed against the National Center for Biotechnology Information's base sequence database using the search program Blastn (Altschul et al., J Mol Biol, 215(3), 1990).
[0075] Figures 16 and 17 show the results of searching the database for each of the constructed full-length mitochondrial DNA sequences. Figure 16 shows the results for the mammalian species, and Figure 17 shows the results for the avian species. In Figures 16 and 17, the species with the highest sequence identity (identity) with the constructed full-length mitochondrial DNA sequence is shown as the "Blastn top hit species." In Figures 16 and 17, the length of the constructed full-length mitochondrial DNA sequence is shown as the "Contig Length." As shown in Figures 16 and 17, for most samples, the full-length mitochondrial DNA sequence of the same species as the analyzed DNA sample was found to match, with a match rate of over 99%. Even for the Japanese Green Pigeon, which showed a 97% match rate with the database, there were no significant structural differences, and the error was due to variations in the copy number of the repeat sequence. The above analysis demonstrated that by using the primer sets shown in Figures 4 and 5, it is possible to construct full-length mitochondrial DNA sequences with extremely high accuracy and without major structural errors for a wide range of biological species belonging to the target "specific biological classification (here, Mammalia or Aves)."
[0076] <Selection of Primer Set from Primer Set Candidates> The following describes the results of an experiment in which mitochondrial DNA fragments were actually amplified using the "primer set candidates" identified in step T130.
[0077] Figure 18 is an explanatory diagram showing the "primer set candidate" identified in step T130, which has the highest degree of conservation (average base conservation calculated using Shannon information) among the "primer set candidate" identified by combining the "primer set candidate sequences" related to the class Aves shown in Figure 10. Note that the "primer set" shown in Figure 5, selected from the "primer set candidate" in step T140, had the second highest degree of conservation (average base conservation calculated using Shannon information) among the "primer set candidate".
[0078] Figure 19 is an explanatory diagram showing the results of PCR performed using the primers constituting the candidate primer set shown in Figure 18 synthesized and genomic DNA extracted from chicken, Japanese green pigeon, and kestrel as templates. As shown in Figure 19, when primer pair numbers 3 and 4 were used, fragments of the desired length were amplified, but when primer pair numbers 1 and 2 were used, little or no amplification was observed. These results indicate that the usefulness of a "primer set" may be relatively significantly affected by factors that cannot be predicted solely by the degree of conservation of the constituent primers, depending on, for example, a "specific biological taxonomy." Therefore, when selecting a "primer set" from the "primer set candidates" in step T140, it is desirable to conduct a preliminary experiment in which the primer set is actually synthesized and amplified using mitochondrial DNA as a template. The "primer set" for Mammalia shown in Figure 4 had the highest degree of conservation of the constituent primers among the "primer set candidates" obtained by combining the "primer candidate sequences" shown in Figure 9.
[0079] <Number of fragments obtained by primer set> In the above example, an example was shown in which mitochondrial DNA was amplified using a primer set that obtained four types of fragments covering the entire length of mitochondrial DNA. Below, an example will be described in which a primer set was used that obtained fewer than four types of fragments (two or three types) that covered the entire length of mitochondrial DNA.
[0080] FIG. 20 is an explanatory diagram showing an example of a primer set for Mammalia designed to obtain two fragments spanning the entire length of mitochondrial DNA. FIG. 21 is an explanatory diagram showing an example of a primer set for Mammalia designed to obtain three fragments spanning the entire length of mitochondrial DNA. The primer sets for the two fragments shown in FIG. 20 were created by selecting two highly conserved primer pairs from the four sets of eight primers included in the primer sets for the four fragments shown in FIG. 4 . Both of these primer pairs were designed to obtain fragments of 10,000 bases or more in length. The primer sets for the three fragments shown in FIG. 21 were designed based on the "candidate primer sequences" for Mammalia shown in FIG. 9 so that the overlapping region was 500 to 3,000 bases long and three fragments of 4,000 to 9,000 bases long were obtained. Using these primer sets, amplification reactions of mitochondrial DNA fragments were performed under the same conditions as those for the primer sets shown in FIGS. 4 and 5 .
[0081] Fig. 22 is an explanatory diagram showing the results of amplification using the primer sets for two types of fragments shown in Fig. 20, and Fig. 23 is an explanatory diagram showing the results of amplification using the primer sets for three types of fragments shown in Fig. 21. Figs. 22 and 23 show the results of confirming, by electrophoresis, the amplification products obtained by amplification for each primer pair constituting each primer set.
[0082] As shown in Figure 22, when the primer set for two types of fragments was used, amplification of the DNA fragment was observed in cattle among mammalian species, but no amplification was observed at all in pigs or sheep. Thus, it was confirmed that the success rate of amplifying the desired DNA fragment was low when the primer set for two types of fragments was used.
[0083] In contrast, as shown in Figure 23, when primer sets for three types of fragments were used, amplification was observed for all primer pairs in all samples. However, when sheep-derived DNA was used as a template and primer pair No. 3 shown in Figure 21 was used, fragments shorter than the fragment length set during primer set design were mainly amplified, with a relatively low proportion of fragments of the desired length. Thus, although the desired DNA fragments can be amplified when using primer sets for three types of fragments, there was a possibility that undesired fragments may also be amplified, depending on the combination of the primer pair and the organism species from which the template DNA was derived. Therefore, from the perspective of suppressing the amplification of undesired fragments and increasing the accuracy of constructing a full-length mitochondrial DNA sequence, it was confirmed that it is desirable to design a primer set that will obtain four or more fragments that cover the entire length of mitochondrial DNA.
[0084] <Design of second primer set> Separately from the "design of the first primer set" described above, primer sets for determining the full-length mitochondrial DNA sequence were designed for each of the "specific biological classifications" of Mammalia and Aves, in a manner similar to the "design of the first primer set," by the primer set design method shown in Figure 2.
[0085] (Setting of candidate primer sequences) "Design of the second primer set" was performed using the same database as "design of the first primer set," but because it was performed after "design of the first primer set," the number of data registered in the database increased compared to when "design of the first primer set" was performed. Therefore, the number of biological species from which full-length mitochondrial DNA sequences were obtained from the database in step T100 was large for both the class Mammalia and the class Aves, and information on full-length mitochondrial DNA sequences was obtained for 1,620 species belonging to the class Mammalia and 1,081 species belonging to the class Aves.
[0086] In step T110, similar to the "design of the first primer set," full-length mitochondrial DNA sequences were aligned for each of the classes Mammalia and Aves to extract conserved regions. Then, in step T120, the "most conserved sequences" in each "conserved region" were used to identify 38 "highly conserved sequences" for Mammalia and 55 for Aves. Furthermore, among the "highly conserved sequences," sequences that did not meet the generally desirable conditions for primers were eliminated based on thermodynamic parameters. Specifically, Primer3, a PCR primer design tool, was used to eliminate sequences that did not meet the following conditions: "Tm value between 55°C and 65°C," "self-pairing score of 50 or less," "self-pairing score in the 3'-terminal region of 50 or less," "hairpin formation score of 50 or less," and "3'-end structural stability of 4 or less." As a result, the "highly conserved sequences" were narrowed down to 28 for Mammalia and 34 for Aves.
[0087] In step T120, the "highly conserved sequences" were further narrowed down using the level of sequence specificity for sequences of mitochondrial DNA of biological species belonging to a "specific biological taxonomy" as an index. That is, a "first parameter" representing the degree to which each "highly conserved sequence" can bind to the mitochondrial DNA of a biological species belonging to the corresponding "specific biological taxonomy" was calculated, and selection was performed to narrow down the "highly conserved sequences" to those for which the value of the obtained "first parameter" was equal to or greater than a predetermined "first reference value." Furthermore, the "highly conserved sequences" were narrowed down using the level of sequence specificity for sequences of mitochondrial DNA of biological species not belonging to the corresponding "specific biological taxonomy." That is, a "second parameter" representing the degree to which each "highly conserved sequence" can bind to the mitochondrial DNA of a biological species not belonging to the corresponding "specific biological taxonomy" was calculated, and selection was performed to narrow down the "highly conserved sequences" to those for which the value of the obtained "second parameter" was equal to or less than a predetermined "second reference value" that was smaller than the above-mentioned "first reference value."
[0088] Such analysis of sequence specificity was performed using PrimerProspector. Specifically, full-length mitochondrial DNA sequences obtainable from the database were first divided into "target organisms," sequences derived from organisms belonging to a "specific biological taxonomy," and "non-target organisms," sequences derived from organisms not belonging to a "specific biological taxonomy." That is, they were divided into sequences derived from the class Mammalia and sequences derived from non-Mammalia, or sequences derived from the class Aves and sequences derived from non-Aves. PrimerProspector analysis was then performed to examine sequence specificity between the "highly conserved sequences" described above.
[0089] The "first parameter" is the percentage of biological species among all target organisms belonging to a "specific biological taxonomy" whose "highly conserved sequences" align with mitochondrial DNA sequences within one mismatch, and the "first reference value" is set to 85%. The "second parameter" is the percentage of biological species among all non-target organisms not belonging to a "specific biological taxonomy" whose "highly conserved sequences" align with mitochondrial DNA sequences within one mismatch, and the "second reference value" is set to 15%.
[0090] 24 and 25 are explanatory diagrams showing two-dimensional plots of the results of examining the "first parameter" and the "second parameter" for each "highly conserved sequence." FIG. 24 shows the results for Mammalia, and FIG. 25 shows the results for Aves. In FIGS. 24 and 25, the horizontal axis represents the "first parameter" and the vertical axis represents the "second parameter." In FIGS. 24 and 25, the "first standard value" for the "first parameter" of 85% and the "second standard value" for the "second parameter" of 15% are indicated by dashed lines. In FIGS. 24 and 25, "highly conserved sequences" whose "first parameter" is equal to or greater than the "first standard value" and whose "second parameter" is equal to or less than the "second standard value" are indicated by black circles, and "highly conserved sequences" that do not satisfy the above two conditions are indicated by black triangles. In step T120, "highly conserved sequences" that satisfied the above two conditions were set as "candidate primer sequences." As a specific example, for Mammalia, 13 of the 28 "highly conserved sequences" were set as "candidate primer sequences," and for Aves, 20 of the 34 "highly conserved sequences" were set as "candidate primer sequences."
[0091] Fig. 26 is an explanatory diagram showing "candidate primer sequences" set for the class Mammalia, and Fig. 27 is an explanatory diagram showing "candidate primer sequences" set for the class Aves.
[0092] (Primer Set Design) For each of the Mammalia and Aves classes, the "candidate primer sequences" established as described above were used as PCR primers to enumerate combinations that yielded fragments of 4,000 to 8,000 bases in length, and PCR amplification experiments were performed on all of these combinations. For each primer combination, the length of the amplified product was confirmed by electrophoresis, and primer combinations (primer pairs) that yielded amplified products of the expected length were identified. For these primer pairs, primer sets that cover the entire length of mitochondrial DNA with four fragments and whose overlapping regions (overlapping regions AB to DA in Figure 3) were 500 to 4,000 bases in length were enumerated and identified as "candidate primer sets" (Step T130). Among these "candidate primer sets," the one with the shortest average length of the four amplified fragments was selected as the "primer set" (Step T140). The "primer sets" selected in this manner are the primer sets shown in Figures 6 and 7, as described above.
[0093] <Confirmation of the effect of narrowing down based on high sequence specificity> In the "design of the second primer set," as described above, when "candidate primer sequences" were set from the "highly conserved sequences" in step T120, narrowing down was performed based on sequence specificity for the "target organism" and "non-target organism." The effect of such narrowing down was verified as follows.
[0094] First, from among the 28 "highly conserved sequences" for the mammalian class shown in Figure 24, five primer pairs were set up, each combining a primer (Pr1) that satisfies the "sequence specificity condition" of "satisfying both the condition that a 'first parameter' is equal to or greater than a 'first reference value' and the condition that a 'second parameter' is equal to or less than a 'second reference value'," a primer (Pr2) that similarly satisfies the "sequence specificity condition," and primers (Pr3 to Pr6) that do not satisfy the "sequence specificity condition."
[0095] Figure 28 is an explanatory diagram showing the sequences of each of the primers (Pr1 to Pr6). In Figure 24, each of the primers (Pr1 to Pr6) is indicated by an arrow. Using these five primer pairs, PCR was performed on Mammalia (target organisms) and Aves (non-target organisms).
[0096] DNA was extracted from tissues obtained from meat or carcasses of mammals (cattle, pig, sheep, and grizzly bear), and birds (chicken, green pigeon, common kestrel, and duck), and the DNA fragments were amplified using the five primer pairs described above. The DNA extraction and PCR conditions were the same as those used in "Evaluation of primer sets" in "Design of first primer set."
[0097] Figure 29 is an explanatory diagram showing the results of electrophoresis of the reaction products obtained by performing amplification reactions using the above-mentioned five primer pairs and DNA from the above-mentioned four mammalian species and four avian species as templates. Figure 29 shows four lanes representing the results of PCR performed on mammalian species using each primer. These lanes, from left to right, show the results for cattle, pigs, sheep, and grizzly bears. Figure 29 also shows four lanes representing the results of PCR performed on avian species using each primer. These lanes, from left to right, show the results for chickens, green pigeons, common kestrels, and ducks.
[0098] Figure 30 is an explanatory diagram showing the breakdown of the five primer pairs used and the results of amplification reactions using each primer pair. For each primer, Figure 30 shows the "first parameter," i.e., the percentage (%) of the number of mammalian species for which the mitochondrial DNA sequence can be aligned with each primer within one mismatch relative to the total number of mammalian species for which sequences were obtained, as "Mammal specificity." The percentage (%) of the "second parameter" is also shown as "Non-mammal specificity."
[0099] As shown in Figures 29 and 30, when a primer pair (Pr1-Pr2) was used in which both primers in the primer pair satisfied the "sequence specificity condition," amplification was observed when the genomic DNA of any mammal was used as a template, but when the genomic DNA of Aves was used as a template, no amplification was observed in any organism. In contrast, when primer pairs combining primer (Pr1) with a primer that did not satisfy the "sequence specificity condition" were used, amplification was observed even with the genomic DNA of Aves (Pr1-Pr3, Pr1-Pr4, Pr1-Pr5), and amplification was weak even with the genomic DNA of Mammalia (Pr1-Pr3, Pr1-Pr6). These results suggest that narrowing down highly conserved sequences through analysis of sequence specificity can suppress amplification using DNA from organisms other than the target organism as a template, thereby improving the effectiveness of specifically amplifying mitochondrial DNA fragments from the target organism. In other words, narrowing down highly conserved sequences through analysis of sequence specificity can be said to be useful for designing primer pairs that are more suitable as universal primers. Although samples containing mitochondrial DNA that are prepared when determining the full-length mitochondrial DNA sequence are generally thought to contain almost no DNA from non-target organisms, it has been confirmed that narrowing down primers based on sequence specificity for "target organisms" and "non-target organisms" can have the effect of improving the performance of primers for determining the full-length mitochondrial DNA sequence.
[0100] (Evaluation by Sequencing) Using the primer set shown in Figure 6, which was the primer set selected in "Design of the Second Primer Set," the full-length sequences of mitochondrial DNA for various mammalian animals were determined. Specifically, PCR was performed using DNA from cattle, pigs, sheep, Hokkaido shika deer, grizzly bears, Steller sea lions, raccoons, Japanese badgers, wild boars, and rabbits as templates, and the full-length sequences of mitochondrial DNA for each animal were determined. Furthermore, the full-length sequences of mitochondrial DNA for various birds were determined using the primer set shown in Figure 7. Specifically, PCR was performed using DNA derived from chicken, green pigeon, common kestrel, duck, green pheasant, and Japanese grossbeak as templates, and the full-length sequence of mitochondrial DNA for each animal was determined.
[0101] The full-length mitochondrial DNA sequence was determined by PCR using each primer pair constituting each primer set, with DNA derived from Mammalia or Aves as a template, and recovering the resulting fragments. The specific method for determining the full-length mitochondrial DNA sequence was the same as that described in "Evaluation by Sequencing" under "Design of the First Primer Set." To determine whether the full-length mitochondrial DNA sequence obtained for each animal was correctly constructed, a search was performed against the National Center for Biotechnology Information's nucleotide sequence database using the search program Blastn (Altschul et al., J Mol Biol, 215(3), 1990).
[0102] Figures 31 and 32 show the results of searching the database for each of the constructed full-length mitochondrial DNA sequences. Figure 31 shows the results for the mammalian species, and Figure 17 shows the results for the avian species. In Figures 31 and 32, the species with the highest sequence identity (identity) with the constructed full-length mitochondrial DNA sequence is shown as the "Blastn top hit species." In Figures 31 and 32, the length of the constructed full-length mitochondrial DNA sequence is shown as the "Contig Length." As shown in Figures 31 and 32, all samples were matched with the full-length mitochondrial DNA sequence of the same species from which the analyzed DNA sample was derived, and most samples showed a match rate of 99% or higher. Even for the Green Pigeon, which showed a match rate of 97% with the database, there were no significant structural differences, and the error was due to variations in the copy number of the repeat sequence. The above analysis demonstrated that by using the primer sets shown in Figures 6 and 7, it is possible to construct full-length mitochondrial DNA sequences with extremely high accuracy and without major structural errors for a wide range of biological species belonging to the target "specific biological classification (here, Mammalia or Aves)."
[0103] The present disclosure is not limited to the above-described embodiments, and can be realized in various configurations without departing from the spirit thereof. For example, the technical features in the embodiments corresponding to the technical features in each aspect described in the Summary of the Invention section can be appropriately replaced or combined to solve some or all of the above-described problems or achieve some or all of the above-described effects. Furthermore, if a technical feature is not described as essential in this specification, it can be appropriately deleted.
[0104] The present disclosure can also be realized in the following forms. [Application Example 1] A method for designing a primer set for determining the full-length base sequence of mitochondrial DNA, comprising: acquiring known full-length mitochondrial DNA sequences of multiple biological species belonging to a specific biological taxonomy; aligning the acquired full-length mitochondrial DNA sequences to extract, from the full-length mitochondrial DNA sequences, conserved regions that meet predetermined criteria indicating a high degree of base conservation; designing candidate primer sequences using the base sequences of each of the extracted conserved regions; identifying, as candidate primer sets, combinations of candidate primer sequences that, when combined to form a primer set and used to amplify mitochondrial DNA, produce three to six fragments that cover the full length of mitochondrial DNA and achieve a desired length as the overlap region between adjacent fragments; and selecting at least one of the candidate primer sets as the primer set. [Application Example 2] The method for designing a primer set according to Application Example 1, wherein the specific biological taxonomy is a class in the taxonomic hierarchy of organisms or a taxonomic hierarchy lower than a class. [Application Example 3] The method for designing a primer set according to Application Example 2, wherein the specific biological classification is Mammalia or Aves. [Application Example 4] The method for designing a primer set according to any one of Application Examples 1 to 3, wherein the candidate primer sets are identified so that the length of the overlapping region is 500 bases or more and 4000 bases or less. [Application Example 5] The method for designing a primer set according to any one of Application Examples 1 to 4, wherein the candidate primer sets are identified so that the length of the overlapping region is 500 bases or more and 3000 bases or less.[Application Example 6] The method for designing a primer set according to any one of Application Examples 1 to 5, wherein the primer set is specified so that four fragments covering the entire length of mitochondrial DNA are obtained when mitochondrial DNA is amplified using the primer set. [Application Example 7] The method for designing a primer set according to any one of Application Examples 1 to 6, wherein a mitochondrial DNA fragment is experimentally amplified using DNA derived from an organism belonging to the specific biological taxonomy as a template, using each of a primer pair consisting of a pair of primers constituting the candidate primer set, and selecting the primer set from the candidate primer sets so that the primer set is composed of a primer pair that is determined to obtain the desired mitochondrial DNA fragment based on the amplification results. Application Example 8 The method for designing a primer set according to any one of Application Examples 1 to 7, wherein the candidate primer sequences are set by: identifying highly conserved sequences, which are base sequences at sites where the degree of base conservation is relatively high, within each of the conserved regions; and designating as the candidate primer sequences sequences where a first parameter representing the degree of sequence specificity of the highly conserved sequence to a sequence of mitochondrial DNA of a biological species belonging to the specific biological taxonomy is equal to or greater than a predetermined first reference value, and a second parameter representing the degree of sequence specificity of the highly conserved sequence to a sequence of mitochondrial DNA of a biological species not belonging to the specific biological taxonomy is equal to or less than a predetermined second reference value that is smaller than the first reference value.[Application Example 9] A method for designing a primer set according to Application Example 8, wherein the first parameter is the ratio of the number of biological species having one or less mismatched bases when the highly conserved sequence is aligned to the full-length sequences of mitochondrial DNA of the plurality of biological species belonging to the specific biological taxonomy to the total number of the plurality of biological species belonging to the specific biological taxonomy, and the second parameter is the ratio of the number of biological species having one or less mismatched bases when the highly conserved sequence is aligned to the full-length sequences of mitochondrial DNA of the plurality of biological species not belonging to the specific biological taxonomy to the total number of the plurality of biological species not belonging to the specific biological taxonomy, and the first reference value is 85%, and the second reference value is 15%. [Application Example 10] A primer set for determining the full-length base sequence of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOS: 1 and 2 in the Sequence Listing at their 3' ends; a second primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOS: 3 and 4 in the Sequence Listing at their 3' ends; a third primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOS: 5 and 6 in the Sequence Listing at their 3' ends; and a fourth primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOS: 7 and 8 in the Sequence Listing at their 3' ends, wherein each of the primers has a length of 100 bases or less. [Application Example 11] A primer set for determining the full-length base sequence of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOs: 9 and 10 in the Sequence Listing at their 3' ends; a second primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOs: 11 and 12 in the Sequence Listing at their 3' ends; a third primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOs: 13 and 14 in the Sequence Listing at their 3' ends; and a fourth primer pair consisting of a pair of primers containing the base sequences of SEQ ID NOs: 15 and 16 in the Sequence Listing at their 3' ends, wherein each of the primers has a length of 100 bases or less.[Application Example 12] A primer set for determining the full-length nucleotide sequence of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 130 and 131 of the Sequence Listing at their 3' ends; a second primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 132 and 133 of the Sequence Listing at their 3' ends; a third primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 134 and 135 of the Sequence Listing at their 3' ends; and a fourth primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 136 and 137 of the Sequence Listing at their 3' ends, wherein each of the primers has a length of 100 bases or less. [Application Example 13] A primer set for determining the full-length nucleotide sequence of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 138 and 139 of the Sequence Listing at their 3' ends; a second primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 140 and 141 of the Sequence Listing at their 3' ends; a third primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 142 and 143 of the Sequence Listing at their 3' ends; and a fourth primer pair consisting of a pair of primers containing the nucleotide sequences of SEQ ID NOs: 144 and 145 of the Sequence Listing at their 3' ends, wherein each of the primers has a length of 100 bases or less. [Application Example 14] A method for determining the full-length base sequence of mitochondrial DNA, comprising: preparing a DNA sample containing mitochondrial DNA; using the DNA sample as a template and a primer set designed by the method for designing a primer set described in any one of Application Examples 1 to 9, amplifying DNA fragments; analyzing the base sequence of each of the amplified DNA fragments; and constructing the full-length sequence of mitochondrial DNA using the base sequences of each of the DNA fragments.[Application Example 15] A method for determining the full-length base sequence of mitochondrial DNA, comprising: preparing a DNA sample containing mitochondrial DNA; using the DNA sample as a template and the primer set described in any one of Application Examples 10 to 13 to amplify DNA fragments; analyzing the base sequence of each of the amplified DNA fragments; and constructing the full-length sequence of mitochondrial DNA using the base sequences of each of the DNA fragments. [Application Example 16] A method for determining the full-length base sequence of mitochondrial DNA, according to Application Example 14 or 15, wherein samples containing mitochondrial DNA derived from multiple types of organisms are prepared as the DNA sample. [Application Example 17] A method for determining the full-length base sequence of mitochondrial DNA, according to any one of Application Examples 14 to 16, wherein a sample containing genomic DNA other than mitochondrial DNA in addition to mitochondrial DNA is prepared as the DNA sample. Application Example 18 An apparatus for designing a primer set for determining the full-length base sequence of mitochondrial DNA, comprising: an acquisition unit that acquires known full-length sequences of mitochondrial DNA of multiple biological species belonging to a specific biological taxonomy; an extraction unit that aligns the full-length sequences of mitochondrial DNA acquired by the acquisition unit and extracts, from the full-length sequences of mitochondrial DNA, conserved regions that satisfy predetermined criteria indicating a high degree of base conservation; a setting unit that sets candidate primer sequences using the base sequences of each of the conserved regions extracted by the extraction unit; and an identification unit that identifies, as candidate primer sets, combinations of candidate primer sequences that, when the candidate primer sequences are combined and used as a primer set to amplify mitochondrial DNA, result in three to six types of fragments that cover the full length of mitochondrial DNA and that have a desired length as the length of the overlap region with adjacent fragments.[Application Example 19] A computer program for designing a primer set for determining the full-length base sequence of mitochondrial DNA, the computer program causing a computer to execute the following functions: an acquisition function for acquiring known full-length mitochondrial DNA sequences of multiple biological species belonging to a specific biological taxonomy; an extraction function for aligning the full-length mitochondrial DNA sequences acquired by execution of the acquisition function and extracting, from the full-length mitochondrial DNA sequences, conserved regions that satisfy predetermined criteria indicating a high degree of base conservation; a function for setting candidate primer sequences using the base sequences of each of the conserved regions extracted by execution of the extraction function; and a function for identifying, as a candidate primer set, a combination of candidate primer sequences that, when the candidate primer sequences are combined and used as a primer set to amplify mitochondrial DNA, yields between three and six fragments that cover the full length of mitochondrial DNA and has a desired length as the length of overlap region with adjacent fragments. [Application Example 20] A storage medium storing the computer program of Application Example 19.
[0105] DESCRIPTION OF SYMBOLS 10: Design device 110: CPU 111: Acquisition unit 112: Extraction unit 113: Setting unit 114: Identification unit 115: Selection unit 120: Storage unit 122: Program 130: RAM 140: Input interface 150: Output interface 160: Communication interface 170: Operation unit 180: Display unit 190: External database
Claims
1. A method for designing a primer set for determining the nucleotide sequence of the entire length of mitochondrial DNA, comprising: obtaining the nucleotide sequences of the entire lengths of mitochondrial DNAs of a plurality of species belonging to a specific biological classification; aligning the obtained nucleotide sequences of the entire lengths of the mitochondrial DNAs, and extracting conserved regions, which are regions satisfying a criterion predetermined as a criterion indicating a high degree of nucleotide conservation, from the nucleotide sequences of the entire lengths of the mitochondrial DNAs; setting primer candidate sequences using the nucleotide sequences of the extracted conserved regions; identifying, as primer set candidates, combinations of the primer candidate sequences that, when used to amplify mitochondrial DNA as a primer set, result in obtaining three or more and six or fewer fragments that cover the entire length of the mitochondrial DNA and a desired length as the length of the overlapping region with an adjacent fragment; and selecting at least any one of the primer set candidates as the primer set.
2. The method for designing a primer set according to claim 1, wherein the specific biological classification is a class in the biological classification hierarchy or a classification hierarchy lower than the class.
3. The method for designing a primer set according to claim 2, wherein the specific biological classification is the class Mammalia or the class Aves.
4. The method for designing a primer set according to claim 1, wherein the primer set candidates are identified such that the length of the overlapping region is 500 bases or more and 4000 bases or less.
5. The method for designing a primer set according to claim 4, wherein the primer set candidates are identified such that the length of the overlapping region is 500 bases or more and 3000 bases or less.
6. The method for designing a primer set according to claim 1, wherein the primer set is identified such that, when mitochondrial DNA is amplified using the primer set, four fragments that cover the entire length of the mitochondrial DNA are obtained.
7. A method for designing a primer set according to claim 1, wherein each primer pair consisting of a pair of primers constituting the primer set candidate is used to experimentally amplify a mitochondrial DNA fragment using DNA derived from an organism belonging to the specific biological classification as a template, and based on the result of the amplification, the primer set is selected from the primer set candidate so as to be constituted by a primer pair that is determined to have obtained a desired mitochondrial DNA fragment. A method for designing a primer set.
8. A method for designing a primer set according to claim 1, wherein the setting of the primer candidate sequence is performed by specifying a highly conserved sequence that is a base sequence of a site with a relatively high degree of base conservation among each of the conserved regions, and a first parameter representing the degree of sequence specificity of the highly conserved sequence with respect to the sequence of mitochondrial DNA of a species belonging to the specific biological classification is equal to or higher than a predetermined first reference value, and a second parameter representing the degree of sequence specificity of the highly conserved sequence with respect to the sequence of mitochondrial DNA of a species not belonging to the specific biological classification is equal to or lower than a second reference value that is smaller than the first reference value. A method for designing a primer set is performed by using the sequence as the primer candidate sequence.
9. A method for designing a primer set according to claim 8, wherein the first parameter is the ratio of the number of species having 1 or less mismatched bases when the highly conserved sequence is aligned with the full-length sequences of mitochondrial DNA of the plurality of species belonging to the specific biological classification to the total number of the plurality of species belonging to the specific biological classification, and the second parameter is the ratio of the number of species having 1 or less mismatched bases when the highly conserved sequence is aligned with the full-length sequences of mitochondrial DNA of the plurality of species not belonging to the specific biological classification to the total number of the plurality of species not belonging to the specific biological classification, the first reference value is 85%, and the second reference value is 15%. A method for designing a primer set.
10. A primer set for determining the nucleotide sequence of the entire length of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 1 and SEQ ID NO: 2 in the sequence listing at the 3'-end side; a second primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 3 and SEQ ID NO: 4 in the sequence listing at the 3'-end side; a third primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 5 and SEQ ID NO: 6 in the sequence listing at the 3'-end side; a fourth primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 7 and SEQ ID NO: 8 in the sequence listing at the 3'-end side; and each of the primers has a length of 100 bases or less.
11. A primer set for determining the nucleotide sequence of the entire length of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 9 and SEQ ID NO: 10 in the sequence listing at the 3'-end side; a second primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 11 and SEQ ID NO: 12 in the sequence listing at the 3'-end side; a third primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 13 and SEQ ID NO: 14 in the sequence listing at the 3'-end side; a fourth primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 15 and SEQ ID NO: 16 in the sequence listing at the 3'-end side; and each of the primers has a length of 100 bases or less.
12. A primer set for determining the nucleotide sequence of the entire length of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 130 and SEQ ID NO: 131 in the sequence listing at the 3'-end side; a second primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 132 and SEQ ID NO: 133 in the sequence listing at the 3'-end side; a third primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 134 and SEQ ID NO: 135 in the sequence listing at the 3'-end side; a fourth primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 136 and SEQ ID NO: 137 in the sequence listing at the 3'-end side; and each of the primers has a length of 100 bases or less.
13. A primer set for determining the nucleotide sequence of the entire length of mitochondrial DNA, comprising: a first primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 138 and SEQ ID NO: 139 in the Sequence Listing at the 3'-terminal side; a second primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 140 and SEQ ID NO: 141 in the Sequence Listing at the 3'-terminal side; a third primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 142 and SEQ ID NO: 143 in the Sequence Listing at the 3'-terminal side; a fourth primer pair consisting of a pair of primers each containing the nucleotide sequence of SEQ ID NO: 144 and SEQ ID NO: 145 in the Sequence Listing at the 3'-terminal side; and each of the primers has a length of 100 bases or less. Primer set.
14. A method for determining the nucleotide sequence of the entire length of mitochondrial DNA, comprising: preparing a DNA sample containing mitochondrial DNA; using the DNA sample as a template and using a primer set designed by the primer set design method according to any one of claims 1 to 9 to amplify a DNA fragment; analyzing the nucleotide sequence of each of the amplified DNA fragments; and constructing the entire length sequence of mitochondrial DNA using the nucleotide sequence of each of the DNA fragments. Method for determining the nucleotide sequence of the entire length of mitochondrial DNA.
15. A method for determining the nucleotide sequence of the entire length of mitochondrial DNA, comprising: preparing a DNA sample containing mitochondrial DNA; using the DNA sample as a template and using the primer set according to any one of claims 10 to 13 to amplify a DNA fragment; analyzing the nucleotide sequence of each of the amplified DNA fragments; and constructing the entire length sequence of mitochondrial DNA using the nucleotide sequence of each of the DNA fragments. Method for determining the nucleotide sequence of the entire length of mitochondrial DNA.
16. The method for determining the nucleotide sequence of the entire length of mitochondrial DNA according to claim 14, wherein as the DNA sample, a sample containing mitochondrial DNA derived from a plurality of types of organisms is prepared. Method for determining the nucleotide sequence of the entire length of mitochondrial DNA.
17. A method for determining the full-length nucleotide sequence of mitochondrial DNA according to claim 15, wherein as the DNA sample, a sample containing mitochondrial DNA derived from a plurality of types of organisms is prepared. A method for determining the full-length nucleotide sequence of mitochondrial DNA.
18. A method for determining the full-length nucleotide sequence of mitochondrial DNA according to claim 14, wherein as the DNA sample, a sample containing genomic DNA different from mitochondrial DNA in addition to mitochondrial DNA is prepared. A method for determining the full-length nucleotide sequence of mitochondrial DNA.
19. A method for determining the full-length nucleotide sequence of mitochondrial DNA according to claim 15, wherein as the DNA sample, a sample containing genomic DNA different from mitochondrial DNA in addition to mitochondrial DNA is prepared. A method for determining the full-length nucleotide sequence of mitochondrial DNA.
20. A primer set design device for determining the full-length nucleotide sequence of mitochondrial DNA, comprising: an acquisition unit that acquires the full-length sequences of known mitochondrial DNAs of a plurality of species belonging to a specific biological classification; an extraction unit that aligns the full-length sequences of the mitochondrial DNA acquired by the acquisition unit and extracts a conserved region that is a region satisfying a predetermined criterion as a criterion indicating the degree of conservation of bases from the full-length sequence of the mitochondrial DNA; a setting unit that sets primer candidate sequences using the nucleotide sequences of the respective conserved regions extracted by the extraction unit; and a specifying unit that specifies a combination of primer candidate sequences as a primer set candidate, such that when mitochondrial DNA is amplified using the combination of the primer candidate sequences as a primer set, three or more and six or less fragments that cover the full length of the mitochondrial DNA are obtained, and a desired length is obtained as the length of the overlapping region with an adjacent fragment. A primer set design device.
21. A computer program for designing a primer set for determining the nucleotide sequence of the full length of mitochondrial DNA, comprising: an acquisition function for acquiring the full length sequences of known mitochondrial DNA of a plurality of species belonging to a specific biological classification; an extraction function for aligning the full length sequences of the mitochondrial DNA acquired by the execution of the acquisition function and extracting a conserved region, which is a region satisfying a predetermined criterion as a criterion indicating a high degree of nucleotide conservation, from the full length sequences of the mitochondrial DNA; a function for setting primer candidate sequences using the nucleotide sequences of the respective conserved regions extracted by the execution of the extraction function; and a function for specifying, as primer set candidates, combinations of primer candidate sequences that, when used to amplify mitochondrial DNA as a primer set, result in three or more and six or fewer fragments covering the full length of the mitochondrial DNA and a desired length as the length of the overlapping region with an adjacent fragment, and causing the computer to execute the functions.
22. A storage medium storing the computer program according to claim 21.
Citation Information
Patent Citations
Universal primer for amplification of whole genome of acarid mitochondria and amplification method
CN116121406A