Method for measuring germinal mutation rate of sexually reproducing animals and application thereof
By analyzing the frequency of variations in neutral regions of the genome and constructing phylogenetic trees, the problems of narrow applicability and high cost of phylogenetic mutation rate calculation have been solved, achieving low-cost and accurate phylogenetic mutation rate calculation applicable to all sexually reproducing animals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-03-24
AI Technical Summary
Existing methods for calculating phylogenetic mutation rates have narrow applicability, are difficult to obtain data for, and are costly to calculate. They cannot be widely applied to the animal kingdom, especially to wild or rare species with no clear kinship.
By analyzing the variation frequency of neutral regions in the genome, a phylogenetic tree is constructed. Germplasm mutation rate is calculated using nucleotide sequence alignment and divergence time of single-copy homologous genes. Conventional second-generation sequencing data is used, reducing computational resource requirements. This method is applicable to all sexually reproducing animals.
It enables widely applicable, low-cost, and accurate calculation of phylogenetic mutation rates, applicable to all sexually reproducing animals, reduces computational resource requirements, and is suitable for rapid clinical screening in non-bioinformatics analysis laboratories and general hospitals.
Smart Images

Figure CN121171342B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of bioengineering, and relates to a mutation detection method based on molecular biology, in particular to a method for measuring nucleotide germline mutation rate at the whole genome level of sexually reproducing animals and application thereof. BACKGROUND
[0002] Mutation is a core concept in biology. On the one hand, mutation is the direct cause of thousands of genetic diseases; on the other hand, mutation is the molecular mechanism of the source of biological diversity. Without mutation, there is no evolution. According to the cell type in which mutation occurs, mutation includes somatic mutation and germline mutation. Somatic mutation is a mutation that occurs in a single individual cell and is limited to the tissue produced by the mutant cell, and does not be passed on to offspring; while germline mutation is a mutation that occurs in germ cells (such as ova or sperm), which can be passed on to offspring, and each cell in the offspring will be affected by the mutation. Therefore, germline mutation is the ultimate source of all genetic variation, determines the speed of genome evolution, and provides raw materials for natural selection. Germline mutation may affect the adaptability of the entire population, such as the formation of sickle cell anemia disease in humans, which is closely related to germline mutation.
[0003] Germline mutation rate is one of the most basic parameters in evolutionary biology. Germline mutation rate is not only a key parameter for determining the age of speciation events, calculating population effective size, and reconstructing population historical dynamics. Germline mutation rate also has application prospects in biological research and biomedicine, such as: (1) providing information for protection strategies of endangered species (such as Chinese crane or Acrozooid), (2) guiding the management of invasive species (such as Pomacea or Tilapia), (3) the intrinsic germline mutation rate (iGMR) of a single family or population can be used as a new type of clinical biomarker. In summary, accurate calculation of germline mutation rate has important significance for both theoretical research and practical application; the core principle is to detect and count new generated and heritable variation sites under specific conditions.
[0004] Currently, the commonly used methods for estimating the germline mutation rate include the mutation accumulation method (MA method) and the pedigree sequencing method (PO method). Among them, the MA method is considered as the "gold standard" for estimating the germline mutation rate, and its core principle is: first, let the mutations accumulate in the environment without natural selection through laboratory culture for multiple generations (the culture process may take months or years); then, compare the genomes of the founder ancestor and the last generation of offspring, count all the mutation sites of the offspring, and calculate the germline mutation rate. Although this method can accurately calculate the germline mutation rate, the experimental period is long, the process is laborious and time-consuming, and it is only suitable for a few model organisms such as nematodes and fruit flies that have short generation times and are easy to culture in the laboratory, which greatly limits its application range. The PO method is an extension of the MA method, and its principle is to only sequence the whole genome of a specific individual and its parents at high depth, identify the mutation sites generated in a single generation, and then calculate the mutation rate. Since only one generation of family is needed to be determined, the PO method breaks through the limitation of the MA method that is only suitable for model species, and is also the most common method for estimating mutation rate at present. However, the PO method is only suitable for samples with exact kinship, and cannot be applied to most wild or rare species without clear kinship or difficult to obtain. In addition, since the number of mutations generated between one generation is between a few nucleotides and dozens of nucleotides (bp), in the face of a genome with a size of GB level (1GB = 1x10 9 bp), false positive or false negative mutation sites can greatly affect the results, leading to large variability in the PO method measurement results. Therefore, both the MA method and the PO method have the following shortcomings: (1) narrow applicability, cannot be widely applied to the animal kingdom; (2) need high-depth sequencing and chromosome-level genome assembly quality, high sequencing cost; (3) and the data analysis process of both methods is very complex, with high calculation cost.
[0005] In summary, it is an urgent technical problem to develop a universal, accurate and low-cost method for calculating the germline mutation rate. Currently, there is no related report. SUMMARY
[0006] In view of the problems of narrow applicability, difficulty in data acquisition and high calculation cost of the calculation method of germline mutation rate in the prior art, the present application provides a method for measuring the nucleotide germline mutation rate of the whole genome of a sexually reproducing animal. The specific principle is that neutral regions or neutral sites are widely distributed in the genome, and mutations at these positions are usually not affected by natural selection; by analyzing the variation frequency of these sites over time, the germline mutation rate can be calculated, and based on this principle, the present application is proposed. The method described in the present application is not only widely applicable to sexually reproducing animal species that have obtained genome or transcriptome data, but also accurate and efficient, filling the gap in the prior art, and having important significance for theoretical research and practical application of biomedical engineering.
[0007] The technical solution of the present application is as follows:
[0008] (All auxiliary analysis software and processes used in the present application have been checked by peer review in the academic field, and all original codes have been fully open sourced without any copyright requirements.)
[0009] A method for measuring the annual germline mutation rate of a sexually reproducing animal, comprising the following steps:
[0010] (1) Genome sequencing of the test species to obtain the sequence of all coding proteins (Coding sequence, CDS) of the test species. Specifically:
[0011] ① Whole genome sequencing (second or third generation technology) of the target species is performed to complete genome assembly and structural annotation;
[0012] ② The sequence of coding proteins of the test species is obtained using a genome structure annotation process (such as Maker Pipeline);
[0013] ③ The pseudogene sequence in the genome is obtained through a genome pseudogene identification process (such as PseudogenePipeline).
[0014] In the case of no reference genome for the target species: In the method described in the present application, routine second-generation sequencing (sequencing depth of 20x) of the genome can be used to calculate the nucleotide mutation rate of the species, and there is no special requirement for the assembly quality of the genome. Compared with the MA method and the PO method which require high-depth (sequencing depth of at least 35x) second-generation sequencing and high-quality genome assembly, the method described in the present application requires very simple data and has obvious advantages.
[0015] In the case of target species with reference genomes: the measurement method described in the present application can run in a conventional PC, only 8 threads, 16 GB running memory are required. While the MA method or the PO method needs to align the high-depth second-generation sequencing data back to the genome, and then perform single nucleotide polymorphism site identification steps, which require professional computing clusters with a minimum configuration of 256 threads and 500 GB of running memory. In contrast, the method described in the present application does not require a professional-level computing cluster, the required computing resources are significantly reduced, and a personal computer can complete all calculations, making it widely applicable, suitable for non-biological information analysis laboratories and research institutions, and also suitable for rapid screening in general hospitals.
[0016] In addition, the method described in the present application uses data for pseudogene sequence replacement sites in the genome or synonymous replacement sites in the coding protein sequence, and only the data of a single individual is required to calculate the germline mutation rate, which theoretically covers all sexually reproducing animals. While the commonly used MA method and PO method require all sites in the whole genome sequences of two or more generations, and the MA method is only suitable for laboratory-cultured model species, and the PO method is only suitable for families with clear kinship, both methods require sequencing data of multiple individuals, and the application range is limited.
[0017] (2) Constructing a phylogenetic tree of the species to be tested: select 4-8 close relatives of the species to be tested, obtain single-copy homologous genes common to the species to be tested and all close relatives; align the amino acid sequences of the single-copy homologous genes, and align the corresponding nucleotide sequences accordingly; then concatenate the nucleotide sequences of the single-copy homologous genes, and use the maximum likelihood method to construct a phylogenetic tree of the species to be tested and the close relatives. Unlike conventional analysis of directly aligning nucleotide sequences, the present application aligns nucleotide sequences based on amino acid sequences, which can preserve codon integrity and correct reading frames, reducing the probability of incorrect alignment of non-homologous sites. In addition, this alignment method is suitable for sequences of species with longer divergence times or distant genetic relationships in the present application, providing more accurate and biologically meaningful alignment results for downstream mutation rate calculation. The specific steps are as follows:
[0018] ① Select 4-8 close relatives with reference genomes, the species to be tested and the close relatives should belong to the same class, obtain the whole genome level CDS sequences of all species (the species to be tested and all close relatives), and then translate them into protein sequences (Peptide sequence, PEP).
[0019] ②According to the PEP sequences described above, use homologous gene identification tools (such as OrthoFinder tool) to identify single-copy homologous genes common to the selected species.
[0020] ③ Use a multiple sequence alignment tool (such as the MAFFT tool) to align all single-copy homologous PEP sequences, and then align the corresponding CDS sequences according to the aligned PEP sequences.
[0021] ④ Concatenate the CDS sequences based on codon alignment, set the first, second and third positions of the codon as three partitions, use the phylogenetic analysis software to automatically search for the most suitable nucleotide substitution model for the three partitions (for example, the " -m MFP + MERGE" parameter in the IQ-TREE software), and use the maximum likelihood method to construct a phylogenetic tree of the test species and close relatives.
[0022] (3) Obtain the divergence time T between the test species and close relatives divergence : Use the relaxed molecular clock method in the phylogenetic software to obtain the divergence time T between the target species and close relatives in the phylogenetic tree constructed in step (2) using the calibrated time node and the nucleotide substitution model divergence . The calibrated time node is specifically the fossil evidence time and the geological event time in the phylogenetic tree; the number of calibrated time nodes is not less than two. Specifically:
[0023] ① Use a divergence time inference tool based on the Markov chain Monte Carlo method (such as the MCMCTree module in the PAML software) to estimate the divergence time between the target species and close relatives in the above phylogenetic tree, and use the fossil evidence, geological event time node and calibration time as the calibrated time node. The fossil evidence time and the geological event time are obtained from public paleontology and geology documents, and the calibration time is obtained from the TimeTree 5 database.
[0024] ② Use the optimized parameters suitable for the present application, specifically as follows: split each single-copy homologous gene into 3 independent data according to the position of the codon to cope with different substitution rates at different sites. Select the independent rate molecular clock method among the relaxed molecular clock method to accurately simulate the real evolution process. Select the general time reversible (GTR) model as the nucleotide substitution model. (The above parameters are represented in MCMCTree as "ndata = 3, clock = 2, model = 7") Need to calculate at least 2 times independently to ensure that the calculation results converge and are consistent. Then summarize all the results to obtain the divergence time T divergence .
[0025] (4) Optimize the phylogenetic tree of the test species: use the four-fold degenerate sites in the single-copy homologous genes to optimize the phylogenetic tree constructed in step (2) using the maximum likelihood method to obtain the optimized phylogenetic tree; thereby obtaining the branch length D speciesand the branch length D of the closest relative species to the species to be tested phylogenetic The four-fold degenerate site is the third codon position, which encodes the same amino acid regardless of the base type, including sites such as GCN, CGN, GGN, CTN, CCN, TCN, ACN, GTN, wherein N represents any base type.
[0026] (5) Obtain the annual species mutation rate μ of the species to be tested year :
[0027] ① If the divergence time of the target species and the closest relative species selected is less than 5 Mya (million years ago), the annual species mutation rate μ of the target species is calculated by Eq. 1. year .
[0028] Eq. 1
[0029] wherein k is the divergence degree of the homologous pseudogene sequence, and θπ is the nucleotide polymorphism of the target species. The inventors believe that in the method described in the present application, when the divergence time of the target species and the closest relative species selected is less than 5 Mya, the difference between the "sequence divergence time" and the "species divergence time" needs to be considered, that is, the divergence time of the species is later than the divergence time of the sequence, so the mutation rate needs to be corrected by θπ. The divergence degree k of the homologous pseudogene sequence is obtained by the following method: first, obtain the pseudogene sequence of the closest relative species to the species to be tested in the phylogenetic tree constructed in step (2), then obtain the homologous pseudogene sequence by alignment, count all the alignment results, and calculate the k value by using the Kimura two-parameter model. The nucleotide polymorphism θπ of the target species is obtained by the following method: resequencing at least three individuals of the species to be tested to obtain the single nucleotide polymorphism sites at the population level, and then calculating the θπ value by using population genetics statistical tools.
[0030] ② If the divergence time of the target species and the closest relative species selected is between 5 and 20 Mya, the annual species mutation rate μ of the target species is calculated by Eq. 2. year .
[0031] Eq. 2
[0032] The inventors believe that in the method described in the present application, when the divergence time between the target species and the closest related species is within 5-20 Mya, the divergence time of the species can be considered to be approximately equal to the sequence divergence time, and the difference between the two can be ignored on a large time scale, without considering the influence of nucleotide polymorphism and incomplete lineage screening. The upper limit of the specified time is 20 Mya, because existing research shows that there is a "Twilight Zone" in the process of nucleotide sequence alignment, with a sequence similarity of 60-65%, that is, when the sequence similarity is less than 65%, the direct alignment result of the pseudogene nucleotide is unreliable. According to existing research on mutation rates of species, the inventors calculated that when the divergence time is greater than 20 Mya, the similarity of most homologous sequences will be less than 65%, and it is not possible to use the pseudogene region to estimate the mutation rate. Therefore, the present application uses the aforementioned simplified formula to estimate the phylogenetic mutation rate, which essentially uses four-fold degenerate sites instead of pseudogene sequences, thereby avoiding the risk of alignment in the Twilight Zone. D species The result has been corrected by multiple models and does not need to be corrected again. The final calculation result of the phylogenetic mutation rate has no significant difference compared with the MA method and the PO method.
[0033] ③When the divergence time between the target species and the closest related species is between 20-200 Mya, the annual phylogenetic mutation rate of the target species is calculated using the Eq.3 calculation formula,
[0034] Eq.3.
[0035] Wherein, Ks mid is the median number of synonymous substitutions of all homologous single-copy genes between the target species and the closest related species in the phylogenetic tree constructed in step (2); the calculation of Ks mid includes all degenerate codons.
[0036] The inventors believe that in the method described in the present application, when the divergence time between the target species and the closest related species is between 20-200 Mya, the sequence divergence time is too large, and the accuracy of sequence alignment needs to be considered. Therefore, the present application uses homologous protein sequences to guide the alignment of nucleotide sequences, and then includes all degenerate codons (two-fold degenerate sites, three-fold degenerate sites, four-fold degenerate sites) in the analysis, that is, calculates the number of synonymous substitutions of nucleotide sites (Ks), and takes the median (Ks mid ) to calculate the mutation rate of the species.
[0037] In addition, since the mutation rate of nucleotides is inconsistent between different species, the invention introduces a correction factor of germ line mutation rate according to the divergence time between the target species and its closest relative species, as shown in the right part of the multiplication sign in the formula.
[0038] In this step, different calculation paths are determined according to the divergence time between the target species and its closest relative species, ensuring that the most suitable data type and theoretical model are used for each evolutionary time scale, thereby excluding known confusion factors in phylogenetic analysis and making the calculation results more accurate. The inventors have derived germ line mutation rate calculation formulas applicable to different divergence times based on existing theories, and by using neutral regions matched with the divergence time, the germ line mutation rate can be calculated more accurately. The calculation strategy is very flexible and is one of the core innovative steps of the method.
[0039] Preferably, the method for measuring the annual germ line mutation rate of sexually reproducing animals further comprises step (5) verification: comparing the annual germ line mutation rate obtained in step (4) with existing methods to verify.
[0040] Preferably, the application of the annual germ line mutation rate obtained by the method as described above. It includes: applying the reliable germ line mutation rate of the target species to evolutionary simulation software such as PSMC, popsizeABC, Fastamical2, etc. to accurately simulate the population historical dynamic changes of the target species and conduct more in-depth research on the target species.
[0041] Preferably, a method for measuring the germ line mutation rate of sexually reproducing animals, the germ line mutation rate is obtained by using Eq. 4 calculation formula:
[0042] μ generation = μ year × T generation Eq. 4;
[0043] Wherein, μ year is the annual germ line mutation rate obtained by the method as described above, T generation is the generation time of the target species.
[0044] Application of the germ line mutation rate obtained by the method as described above.
[0045] The beneficial effects of the present application are:
[0046] (1) The present application innovatively proposes a method for measuring the germ line mutation rate of sexually reproducing animals, which requires simple data types, and conventional second-generation sequencing data can be used to calculate the nucleotide mutation rate of the species. The sequencing depth of 20x can meet the demand, significantly reducing the cost.
[0047] (2) The method for measuring the annual phylogenetic mutation rate of sexually reproducing animals described in this application only requires data from a single individual to complete the calculation of the phylogenetic mutation rate. It is applicable to all sexually reproducing animals and has a wide range of applications.
[0048] (3) The method for measuring the phylogenetic mutation rate of sexually reproducing animals described in this application adaptively determines the matching phylogenetic mutation rate calculation formula based on the estimated divergence time between the target species and its closest available relatives, thereby calculating the phylogenetic mutation rate more accurately.
[0049] Table 1 Comparison of Mutation Rate Estimation Methods
[0050] Attached Figure Description
[0051] Appendix Figure 1 This is a schematic diagram illustrating the steps of the method for measuring the nucleotide germline mutation rate at the whole genome level in sexually reproducing animals as described in this application.
[0052] Appendix Figure 2 This is a set of single-copy homologous PEP sequences aligned in Example 1.
[0053] Appendix Figure 3 This is a set of single-copy homologous CDS sequences aligned in Example 1.
[0054] Appendix Figure 4 This is the maximum likelihood tree re-optimized using four degenerate sites in Example 1.
[0055] Appendix Figure 5 This is an estimate of the divergence time between the scaly-foot snail and the shield snail in Example 1.
[0056] Appendix Figure 6 This is an example of an AXT format file constructed from some sequences in Example 1. Each pair of sequences in the AXT file contains three lines: one line for the sequence name and two lines for the sequence itself. Any pair of sequences is separated by a blank line.
[0057] Appendix Figure 7 This refers to the historical dynamics of the scaly-foot snail population described in Example 1.
[0058] Appendix Figure 8 This section shows the population history dynamics of the Argentine and Chinese populations of *Pomacea canaliculata* as described in Example 2. The green line represents the Chinese population, and the red line represents the Argentine population. Detailed Implementation
[0059] The present invention will be further described below with reference to the embodiments.
[0060] Example 1: Measurement of phylogenetic mutation rate of the endangered deep-sea mollusc snail *Leptochrocera*
[0061] Step (1) Data acquisition and sequence preparation:
[0062] The chromosomal level genome and annotation files of Chrysomallon squamiferum have been published, and the genome sequence, coding sequence (CDS), and peptide sequence (PEP) files can be directly obtained from public databases (such as NCBI GenBank and Dryad database). Therefore, in this embodiment, there is no need to re-sequence, assemble or annotate the target species, and directly use the published high-quality data, which embodies the low cost and high efficiency advantage of the method of the present application in data acquisition.
[0063] Step (2) Phylogenetic tree construction:
[0064] To construct the phylogenetic tree containing Chrysomallon squamiferum, 8 closely related species with existing reference genomes belonging to the same class Gastropoda were selected. The genomes and annotation files of these species were also obtained from public databases, and the details are shown in Table 2.
[0065] Table 2. Summary statistics of selected species for mutation rate calculation
[0066]
[0067] ① Single-copy ortholog identification: Obtain the CDS and PEP sequences of all 9 species including Chrysomallon squamiferum, a total of 197,701. Using the ortholog identification tool OrthoFinder v2.5.5, 3,112 groups of single-copy orthologs common to all species were identified.
[0068] ② Sequence alignment: The "amino acid sequence guided nucleotide sequence alignment" strategy described in the present application was used. First, use the multiple sequence alignment tool MAFFT v7.520 to align the 3,112 groups of single-copy ortholog PEP sequences respectively (see Appendix Figure 2 ). Then, using the PAL2NAL v14.1 tool, according to the aligned PEP sequences as a template, the corresponding CDS sequences were aligned based on the codon (see Appendix Figure 3 ). This strategy ensures the integrity of the codon reading frame, significantly reducing the problem of non-orthologous site misalignment that may occur in distant species comparison.
[0069] ③ Phylogenetic tree inference: Concatenate all 3,112 codon-aligned CDS sequences into one super-long sequence (total length: 1,822,122 bp). Divide the sequence into three partitions according to the 1st, 2nd, and 3rd sites of codons. Use phylogenetic analysis software IQ-TREE v2.2.2.7 with the parameters -m MFP+MERGE -B 1000, and the software automatically searches for the most suitable nucleotide substitution model for each partition. Finally, use Maximum Likelihood (ML) to construct a phylogenetic tree containing Lottia and its close relatives.
[0070] Step (3) Divergence time estimation
[0071] Using the phylogenetic tree constructed in step (2), estimate the divergence time between each species using relaxed molecular clock method combined with nucleotide substitution model.
[0072] ① Calibration time node setting: To calibrate the molecular clock, four calibration time nodes were obtained from published paleontology and geology literature. The specific settings are as follows: The maximum divergence time between African Lanistes nyassanus and South American Pomaceacanaliculata was set to 150 million years ago (Mya), corresponding to the geological event of the splitting of the South American and African continental plates; The minimum divergence time between the new Caenogastropoda and Heterobranchia was set to 390 Mya; The divergence time range between Aplysia californica and Radix auricularia was set to 168.6 Mya to 473.4 Mya; The divergence time range between A. californica and Lottia gigantea was set to 470.2 Mya to 531.5 Mya.
[0073] The referenced paleontology and geology literature is as follows:
[0074] 1. Hayes, K. A., Cowie, R. H., Jørgensen, A., Schultheiß, R., Albrecht, C., & Thiengo, S. C. (2009). Molluscan models in evolutionary biology: apple snails (Gastropoda: Ampullariidae) as a system for addressing fundamental questions. American Malacological Bulletin, 27(1 / 2), 47-58.
[0075] 2. Jörger, K. M., Stöger, I., Kano, Y., Fukuda, H., Knebelsberger, T., & Schrödl, M. (2010). On the origin of Acochlidia and other enigmatic euthyneuran gastropods, with implications for the systematics of Heterobranchia. BMC evolutionary biology, 10(1), 323.
[0076] 3. Benton, M. J., Donoghue, P. C., & Asher, R. J. (2009). Calibrating and constraining molecular clocks. The timetree of life, 35, 86.
[0077] ② Divergence time estimation: MCMCTree module in PAML v4.10.7 package was used for the calculation. The parameters optimized according to the present application were as follows: “ndata = 3” (corresponding to three codon position partitions), “clock = 2” (independent rate molecular clock model was selected to deal with the difference in evolutionary rate among different lineages), “RootAge < 550” (the maximum age of the root node was set according to the fossil record of Gastropoda), “model = 7” (the general time-reversible GTR model was selected). To ensure the convergence and stability of the results, the analysis was independently run three times, and then the posterior samples of the three runs were combined for summary. Finally, the phylogenetic tree with accurate divergence time was obtained (see FIG. 2).Figure 5 The results showed that the divergence time T between the scaly-foot snail and its closest relative, the cochlear snail (Scutellastra cochlear), was... divergence The value is 116.1 Mya.
[0078] Step (4) Phylogenetic tree optimization
[0079] To obtain the number of nucleotide substitutions at four-fold degenerate sites for all species in the phylogenetic tree, the branch lengths of the ML tree constructed in step 2 were optimized using four-fold degenerate sites in single-copy homologous genes. Four-fold degenerate sites are sites that ensure the amino acid remains unchanged regardless of substitutions at the third base of the codon encoding the same amino acid; mutations at these sites are considered approximately neutral. Branch lengths were re-estimated based on 352,890 four-fold degenerate sites using RAxML-NG v1.2.1 software. The optimized phylogenetic tree is attached. Figure 4 The diagram shows the branch length D of the tested species, *Spodoptera exigua*. species The branch length D of its closest relative, the shield snail, is 0.8407. phylogenetic It is 1.0862.
[0080] Step (5) Calculation of annual phylogenetic mutation rate
[0081] Based on the aforementioned steps, the divergence time T between *Symplocos leptospira* and its closest relative, *Syngonium scabra*, is determined. divergence The value was 116.1 Mya, which falls within the 20-200 Mya range defined in this invention. Therefore, Eq.3 was chosen for calculating the annual phylogenetic mutation rate.
[0082] Eq.3.
[0083] ① Calculate the median of synonym substitutions (Ks) mid ): 9,162 pairs of single-copy homologous genes shared by *Triplophysa* and *Triplophysa* were extracted. The protein sequence alignments were converted to codon alignments using ParaAT v2.0, and AXT format files were generated (see attached). Figure 6 Using KaKs_Calculator v3.0 software, the number of synonymous substitutions (Ks) and non-synonymous substitutions (Ka) for each homologous gene pair was calculated using the model mean method (-m MS). To ensure data quality, gene pairs with Ka values greater than 0.2, Fisher's exact test p-values greater than 0.01, or sequence lengths less than 300 bp were filtered out, ultimately retaining 6,025 high-quality homologous gene pairs. The median of these 6,025 Ks values was calculated to obtain Ks. mid .
[0084] (2) Calculate the annual germline mutation rate (μyear): Ks mid (1.88), T divergence (116.1 Mya), D species (0.8407), and D phylogenetic (1.0862) are substituted into Eq. 3, the annual germline mutation rate μ year of Onchidium struma is calculated to be 0.80 x 10 −8 / bp / year.
[0085] (3) Calculate the generation germline mutation rate (μ generation ): According to the reference, the generation time T generation of Onchidium struma is about 2 years. According to Eq. 4, the generation germline mutation rate μ generation of Onchidium struma is calculated to be 1.60 x 10 −8 / bp / generation.
[0086] μ generation = μ year × T generation Eq. 4
[0087] The calculation results of this example provide key parameters for subsequent population historical dynamic research of this endangered species, as shown in FIG. 1. Population history simulation using this mutation rate shows that the effective population size of Onchidium struma has been continuously decreasing in the past ten thousand years, indicating that it is experiencing a population bottleneck, which provides an important scientific basis for formulating its protection strategy. Figure 7 Example 2: Measurement and application of the germline mutation rate of the invasive species Pomacea canaliculata
[0088] This example takes the globally important invasive species Pomacea canaliculata as the research object, provides a method for calculating the germline mutation rate at a medium divergence time scale, and describes its application in invasion biology research.
[0089] Step (1) Data acquisition and sequence preparation
[0090] The high-quality reference genome of the test species Pomacea canaliculata has been published and can be obtained from the National Center for Biotechnology Information (NCBI) database, with the accession number GCF_003073045.1. All coding protein sequences (CDS) and their corresponding protein sequences (PEP) are extracted from the genome annotation file. This step can completely use the published data, and the low-cost characteristics of the method of the application are verified again.
[0091] Step (2) Phylogenetic tree construction
[0092] Step (2) Phylogenetic tree construction
[0093] The steps were basically the same as in Example 1, except that the closely related species Pomacea maculata (data from the NCBI database, accession number GCA_003920245.1) were added.
[0094] Step (3) Estimation of divergence time
[0095] As in Example 1, the selection of the calibration time point is the same and will not be repeated here. The calculation results show that the divergence time T between the golden apple snail (P. canaliculata) and its closest relative, the spotted apple snail (P. maculata), is... divergence It is approximately 12.5 Mya.
[0096] Step (4) Phylogenetic tree optimization
[0097] Similar to Example 1, after optimization, the branch length D of the golden apple snail was obtained. species It is 0.301.
[0098] Step (5) Calculation and application of annual phylogenetic mutation rate
[0099] The divergence time between *Pomacea canaliculata* and its closest relative, *Macrobrachium spp.*, was 12.5 Mya, which falls within the 5–20 Mya range defined in this invention. At this timescale, the difference between species divergence time and sequence divergence time is negligible. Therefore, Eq.2 was chosen for calculation.
[0100] Eq.2.
[0101] ① Calculate the annual phylogenetic mutation rate (μyear): D species (0.1152) and T divergence Substituting (12.5 Mya) into Eq.2, the annual phylogenetic mutation rate of the golden apple snail is calculated to be 2.40 × 10⁻⁶. -8 / bp / year.
[0102] ② Calculate the phylogenetic mutation rate (μ) of each generation. generation ): Generation time T of the golden apple snail in the invaded area generation Extremely short, 0.4 years. Based on Eq.4, the phylogenetic mutation rate was calculated to be 0.96 × 10⁻⁶. -8 / bp / generation.
[0103] application:
[0104] Golden apple snails are a global agricultural pest and a public health threat. Accurate phylogenetic mutation rates are key parameters for precisely reconstructing their invasion history and predicting their future dynamics using population genetics methods (such as PSMC and SMC++). (See attached...)Figure 8 As shown, the mutation rate calculated by this example can be used to simulate the population history dynamics of apple snails in their native Argentina and in the invaded China. The simulation results clearly show that after being introduced to China (about 40-200 years ago), the effective population size of the apple snails has undergone a sharp expansion, which is consistent with the known invasion history. This application demonstrates the important information provided by the present application in the management and prevention strategy of invasive species.
[0105] Example 3: Measure and verify the germline mutation rate of modern humans using ancient human genomes
[0106] This example provides a method for calculating the germline mutation rate of the whole genome of modern humans (Homo sapiens), which aims to demonstrate the application of the present application in primates. In particular, the sequenced ancient humans (such as Neanderthals) are used as outgroups, and compared with the recognized human mutation rate obtained by pedigree sequencing method (PO method) to verify the accuracy of the present application.
[0107] Step (1) Data acquisition and sequence preparation
[0108] The test species is modern humans, and the human reference genome GRCh38 / hg38 is used as a representative. High-quality Altai Neanderthal (Altai Neanderthal) genome and Denisovan genome are selected as outgroups. To ensure the accurate orientation of the phylogenetic tree and the robustness of the branch length estimation, chimpanzees (Pantroglodytes), bonobos (Pan paniscus), gorillas (Gorilla gorilla), orangutans (Pongo abelii) and gibbons (Nomascus leucogenys) are selected as outgroups. All genome data can be obtained from public databases such as UCSC Genome Browser or NCBI, as shown in Table 3.
[0109] Table 3 Species / Sample Information for Human Phylogenetic Analysis
[0110]
[0111] Step (2) Phylogenetic tree construction
[0112] Protein sequences for all species in Table 3 were analyzed using OrthoFinder v2.5.5 to identify single-copy orthologous genes common to all species. Subsequently, codon-based sequence alignments were performed using MAFFT and PAL2NAL, and maximum likelihood phylogenetic trees were constructed using IQ-TREE. The constructed tree is as follows, (( (Modern human, Neanderthal), Denisovan), ((Chimpanzee, Bonobo), Gorilla)), which is consistent with the accepted evolutionary history of humans and their close relatives.
[0113] Step (3) Divergence time estimation
[0114] Divergence time estimation was performed using the MCMCTree module of PAML.
[0115] 1. Calibration time node settings: widely accepted fossil calibration points in primatology research were used: the divergence time of the modern human and chimpanzee lineages. Genetic and fossil evidence suggests that this divergence occurred between approximately 6.5 and 9.0 Mya. The divergence time of the Hominidae and Cercopithecoidea, fossil evidence and molecular clock studies generally support the divergence occurring between 25 and 35 Mya. Referenced paleontological and geological literature as follows:
[0116] 1. Zalmout, I. S., Sanders, W. J., MacLatchy, L. M., Gunnell, G. F., Al-Mufarreh, Y. A., Ali, M. A.,... & Gingerich, P. D. (2010). New Oligocene primate from Saudi Arabia and the divergence of apes and Old World monkeys. Nature, 466(7304), 360-364.
[0117] 2. Patterson, N., Richter, D. J., Gnerre, S., Lander, E. S., & Reich, D. (2006). Genetic evidence for complex speciation of humans and chimpanzees. Nature, 441(7097), 1103-1108.
[0118] 2. The estimated divergence time T between modern humans and Neanderthals divergenceAbout 0.65 Mya (65 million years ago), which is consistent with the current scientific community based on a variety of methods estimated range.
[0119] Step (4) Phylogenetic tree optimization
[0120] This step is the same as the previous example, using four-fold degenerate sites to optimize branch length. Since the divergence time is very short, there is no need to perform this step.
[0121] Step (5) Calculation and verification of annual germline mutation rate
[0122] The divergence time between modern humans and the closest relative of ancient humans (Neanderthals) is 0.65 Mya, which is much smaller than the 5 Mya defined by the present invention. On this extremely short time scale, the influence of ancestral polymorphism (i.e. incomplete lineage sorting) cannot be ignored and must be corrected. Therefore, Eq. 1 is selected for calculation.
[0123] Eq. 1;
[0124] ① Obtain parameter k (homoeolog sequence divergence): Use the PseudogenePipeline to identify pseudogenes in the genomes of modern humans and Neanderthals. By whole-genome alignment, identify pairs of orthologous pseudogenes. Align these pseudogene sequences and calculate the average nucleotide divergence k using the Kimura 2-parameter model.
[0125] ② Obtain parameter θπ (nucleotide polymorphism): θπ reflects the effective population size and mutation rate of the ancestral population, and is an indicator of the level of genetic variation in the population. This value can be calculated from large-scale population resequencing data (such as 1000 Genomes Project, gnomAD). For human populations, the accepted value is about 1 × 10 -3 .
[0126] ③ Calculate annual germline mutation rate (μ year ): Substitute the calculated k, θπ and T divergence into Eq. 1 to calculate the annual germline mutation rate of modern humans μyear.
[0127] ④ Calculate the germline mutation rate (μ generation ) and verify: The average generation time of humans T generation is about 29 years. According to Eq. 4, the calculation result is about 1.20 × 10−8 / bp / generation.
[0128] Validation: Family sequencing approach (PO approach) by direct observation of de novo mutations through whole-genome sequencing of thousands of parent-offspring trios is considered the most direct method to estimate mutation rate in humans. Numerous studies have shown that the germline mutation rate in humans is in the range of 1.1 × 10 -8 to 1.3 × 10 -8 / bp / generation. Therefore, the result (1.20 × 10 −8 ) calculated by the method of the present application is highly consistent with the result of the PO approach, falling within its 95% confidence interval; this indicates the high accuracy and reliability of the method of the present application in recent divergent species.
[0129] Example 4: Measurement and verification of the germline mutation rate of Drosophila melanogaster
[0130] This example provides a method for calculating the whole-genome germline mutation rate of the model organism Drosophila melanogaster, aiming to demonstrate the application of the method of the present application in the classic model organism Drosophila melanogaster. By comparing with D. simulans and comparing with the "gold standard" mutation rate obtained by the mutation accumulation method (MA method), the accuracy of the present application is further verified.
[0131] Step (1) Data acquisition and sequence preparation
[0132] The test species is Drosophila melanogaster, and its reference genome and annotation file are obtained from the FlyBase database. According to user requirements, its close relative Drosophila simulans is selected as the main comparison object. In order to construct a robust phylogenetic tree, the genomes of Drosophila yakuba, Drosophila erecta, Drosophila sechellia and Drosophila grimshawi, which are unique to Hawaii, are also downloaded as outgroups. All data are publicly available resources, as shown in Table 4.
[0133] Table 4 Species information for phylogenetic analysis of Drosophila
[0134]
[0135] Step (2) Phylogenetic tree construction
[0136] The same pipeline (OrthoFinder, MAFFT, PAL2NAL, IQ-TREE) as the previous example was used to analyze all species in Table 5 to construct the maximum likelihood phylogenetic tree.
[0137] Table 5 Summary of KaKs value calculation results of fifteen pairs of homologous sequences
[0138]
[0139] Note: The meaning of the first column of this table is as follows, Sequence: the name of the homologous sequence; Ka: non-synonymous substitution rate; Ks: synonymous substitution rate; P-Value (Fisher): the value calculated by Fisher's exact test; Length: sequence length (after removing gaps and stop codons); S-Sites: synonymous sites; N-Sites: non-synonymous sites; Substitutions: substitutions between sequences; S-Substitutions: synonymous substitutions; N-Substitutions: non-synonymous substitutions; ML-Score: maximum likelihood score; AICc: AICc weight; Akaike-Weight: Akaike weight value for model selection; Model: selected model of MS method.
[0140] Step (3) divergence time estimation
[0141] The phylogeny and divergence time of Drosophila have been studied a lot. However, it can still be calculated again by the method of the present application, and the calculation results are consistent with the current research, as follows.
[0142] ① Calibration time node setting: Calibration point 1: divergence time of Drosophila subgenus and Sophophora subgenus (subgenus where D. melanogaster is located). Fossil and molecular evidence shows that the divergence time is about 40 Mya. The node representing the divergence of D. yakuba (Drosophila subgenus) and D. melanogaster / D. simulans (Sophophora subgenus) in the tree is set to 40 Mya. Calibration point 2: radiation and differentiation time of Hawaiian Drosophila. The formation of Hawaiian Islands is a "conveyor belt" type geological process, which provides an important biogeographical calibration point for the radiation and differentiation of Drosophila. The oldest island of Kauai in Hawaii was formed about 5.1 Mya, so the minimum age of the common ancestor of the Hawaiian Drosophila lineage can be set to 5.1 Mya.
[0143] The references of paleontology and geology are as follows:
[0144] 1. Tamura, K., Subramanian, S., & Kumar, S. (2004). Temporal patterns of fruit fly (Drosophila) evolution revealed by mutation clocks. Molecular biology and evolution, 21(1), 36-44.
[0145] 2. Price, J. P., & Clague, D. A. (2002). How old is the Hawaiian biota? Geology and phylogeny suggest recent divergence. Proceedings of the Royal Society of London. Series B: Biological Sciences, 269(1508), 2429-2435.
[0146] ② Calculation of divergence time: run the MCMCTree module of PAML. This method calculates T divergence to be 2.5 Mya.
[0147] Step (4) optimization of phylogenetic tree
[0148] As in the previous examples, the branch length is optimized using quadruple-degenerate sites.
[0149] Step (5) calculation and verification of annual phylogenetic mutation rate
[0150] The divergence time of D. melanogaster and D. simulans is 2.5 Mya, which is less than 5 Mya, so Eq. 1 is also selected for calculation.
[0151] ① Obtain parameters k and θπ: Obtain the orthologous pseudogene sequences of D. melanogaster and D. simulans by searching in the FlyBase database and related literature, and calculate their divergence k. Because the current fruit fly research is very sufficient, the fruit fly population genomic data provided by public databases such as Drosophila Genome Nexus can be directly used, and then the nucleotide polymorphism θπ of D. melanogaster is calculated, without the need for resequencing.
[0152] ② Calculation of annual phylogenetic mutation rate (μ year ): Substitute k, θπ and T divergence into Eq. 1 to calculate the annual phylogenetic mutation rate μyear 5.20 x 10 −8 / bp / year.
[0153] iii. Calculate the germline mutation rate (μ generation ) and validate: The generation time T generation of D. melanogaster is very short, and there are 10 generations per year in wild fruit fly populations. The calculated germline mutation rate μ generation is about 5.20 x 10 −9 / bp / generation.
[0154] Validate: The mutation accumulation method (MA method) is the traditional “gold standard” method for estimating mutation rates in model organisms such as fruit flies. This method directly counts the number of accumulated mutations by conducting hundreds of generations of inbreeding under no selective pressure. Published MA experimental results show that the mutation rate of D. melanogaster is in the range of 2.8 x 10 −9 to 8.4 x 10 −9 / bp / generation. The calculated result of this embodiment (5.20 x 10 −9 ) is completely within the 95% confidence interval determined by the MA experiment, and there is no significant difference from the “gold standard” method. This again strongly proves the accuracy of the method of the present application, and the present application does not require time-consuming laboratory culture for several years, greatly improving efficiency and broadening the application range.
[0155] Table 6: Summary of results of measuring mutation rates in Examples 1-4
[0156]
[0157] Note: The unit of mutation rate is “the number of mutations per generation at a single nucleotide site in the whole genome”; the P value is whether there is a significant difference between the nucleotide mutation rate calculated by the present application and other methods, and P > 0.05 indicates no significant difference. It should be noted that when calculating the mutation rate, all methods use thousands of homologous genes or homologous pseudogenes, and the median is finally displayed in the table.
[0158] Example 5: Application of evaluating mutation risks in different human family backgrounds
[0159] This example aims to demonstrate the application potential of the present application in the field of clinical genetics and precision medicine, by measuring and comparing the intrinsic germline mutation rates (iGMR) of different families, to provide a new dimension for evaluating genetic disease risks.
[0160] De novo mutations (DNMs) are new mutations that arise in a single generation. DNMs are the main cause of a range of severe, genetic diseases, especially in neurodevelopmental disorders, developmental retardation, and various congenital syndromes. These diseases often bring heavy psychological and economic burden to patients and their families. However, the biggest challenge for clinical genetic counseling comes from the “new” nature of DNMs. Since there is no precedent for these mutations in the family, traditional pedigree-based genetic risk analysis methods cannot predict such diseases. Generally, after the disease occurs, sequencing technology is used to find DNMs to determine the cause.
[0161] In current clinical practice, the main basis for assessing the risk of DNMs is the “Paternal Age Effect” (PAE). Father's age is the most significant factor affecting the number of DNMs in offspring. Spermatogonial stem cells continue to undergo cell division throughout life, and the increase in the number of cell divisions provides more opportunities for errors to occur during DNA replication. PAE is essentially a macroscopic statistical law, and for a biological individual or a family with a genetic history, this law lacks quantitative standards. The “germline mutation rate” calculated by the present application can be used as a clinical biomarker for individuals, which can quantify the mutation trend of DNM risk and complement the current clinical use of DNM.
[0162] Specifically, the standard method for directly detecting DNMs is the “Parent-Offspring” (PO) method, which is to sequence the genomes of parents and children, and identify new mutations in the child's genome by comparison. The essence of this method is to obtain a stochastic event result, similar to the specific score of a “dice roll”. It can only be used for retrospective analysis after the birth of the child and is used for disease diagnosis. In contrast, the method described in the present patent infers the mutation rate by analyzing the sequence differences accumulated in the evolutionary history of an individual over tens of thousands of years. Instead of the mutation quantity of a single generation, the method reflects the long-term, average mutation rate of a specific lineage. It is defined as the intrinsic germline mutation rate (iGMR). In short, the germline mutation rate obtained by the present application uses neutral mutations accumulated over evolutionary time; this long-term average process effectively filters out random fluctuation noise in single-generation inheritance, so that the final signal accurately reflects the basic truth of the DNA replication and repair machinery encoded in the individual's genome. iGMR is not the result of a stochastic event, but a stable biological trait, similar to the intrinsic property of describing whether a dice is “fair” or not.
[0163] While a count of DNM in a born child provides important diagnostic information, it is not very useful in predicting the risk of the next pregnancy, as the outcome of the next "roll of the dice" is still random. The iGMR, as a stable biological trait, can be a neonatal, predictive biomarker for assessing the internal consistency of a newborn. Table 7 summarizes and compares it with the existing clinical DNM analysis standard.
[0164] Table 7. Summary comparison of iGMR detection and trio sequencing
[0165]
[0166] The proposed embodiments are as follows:
[0167] Step (1) Data acquisition and sequence preparation
[0168] This study selected individual genomes from three representative families with different genetic backgrounds as test samples. These samples can be obtained by cooperating with clinical research institutions or from public biological sample banks (such as the NIGMS Human Genetic Cell Bank).
[0169] Family A: A family with a history of Lynch Syndrome. Lynch Syndrome is caused by germline mutations in mismatch repair (MMR) genes (such as MSH2, MLH1), which are known to cause genomic instability. A confirmed carrier in this family was selected as a sample.
[0170] Family B: A family with multiple children with unexplained, sporadic neurodevelopmental disorders (such as autism spectrum disorders). In such diseases, de novo mutations are considered an important causative factor. A healthy father in this family was selected as a sample to investigate whether his basal mutation rate was high.
[0171] Family C (control group): A healthy family with no major genetic disease history for at least three generations. A male individual from this family, matching the age of the father in Family B, was selected as a control sample.
[0172] The selection of outgroups and proximate species is the same as in Example 3, with Neanderthals as the main proximate species for comparison.
[0173] Steps (2)-(4) Phylogenetic tree construction, divergence time estimation, and optimization
[0174] Consistent with the above Example 3, no further description is provided.
[0175] Step (5) Calculation and comparison of iGMR in different human families
[0176] Consistent with Example 3, the iGMR of each family was calculated using the formula of Eq1 in the present application:
[0177] Step (6) Result analysis and application:
[0178] By comparing the iGMR values of the three families, the long-term and intrinsic mutation risk of each family can be evaluated.
[0179] Application 1 (clinical genetic warning): For couples preparing for pregnancy, especially from families with a history of genetic diseases or unexplained repeated occurrence of diseases related to de novo mutations, the present application can provide a more personalized DNM risk assessment beyond the age effect. A significantly high iGMR value can serve as a warning signal, prompting more in-depth genetic screening or considering assisted reproductive technology.
[0180] Application 2 (etiology research): For the high iGMR observed in Family B, the results of the present application can guide subsequent research. Researchers can focus on deep sequencing of genes related to DNA repair and cell cycle regulation in this family member to find the "causative" gene that leads to high mutation rate, which may discover new genetic disease mechanisms.
[0181] The method described in the present application can quantify and compare the long-term mutation rate in different genetic backgrounds, providing an innovative tool for risk stratification, etiology exploration, and personalized medicine of genetic diseases.
[0182] As can be seen from Examples 1-5, the present application provides a universal, accurate and low-cost germline mutation rate measurement method. The method described in the present application adaptively selects the optimal calculation model according to the divergence time between species, overcoming the limitations of narrow application range and high cost of existing methods. At the same time, the verification results of Examples 3 and 4 show that the calculation accuracy of the method of the present application is comparable to that of family sequencing and mutation accumulation system, while its application range can be extended to almost all sexually reproducing animals, and it has broad application prospects in the fields of population genetics and evolutionary medicine, filling an important technical gap in this field. Moreover, as can be seen from Example 5, the germline mutation rate measured by the germline mutation rate measurement method described in the present application provides a new dimension for evaluating genetic disease risk, and has application prospects in the fields of clinical genetics and precision medicine.
Claims
1. A method for measuring the annual phylogenetic mutation rate of sexually reproducing animals, characterized in that: Includes the following steps: (1) Perform genome sequencing on the species to be tested to obtain the sequences of all encoded proteins of the species to be tested; (2) Constructing a phylogenetic tree of the species to be tested: Select 4-8 closely related species of the species to be tested and obtain the single-copy homologous gene shared by the species to be tested and all closely related species; Align the amino acid sequence of the single-copy homologous gene and align its corresponding nucleotide sequence accordingly; Then, tandem the nucleotide sequences of the single-copy homologous gene and construct a phylogenetic tree of the species to be tested and closely related species using the maximum likelihood method. (3) Obtain the divergence time T between the species under test and closely related species. divergence Using the calibration time point, the relaxation molecular clock method in the phylogenetic software, combined with the nucleotide substitution model, is used to obtain the divergence time T between the target species and closely related species in the phylogenetic tree constructed in step (2). divergence ; (4) Optimize the phylogenetic tree of the species to be tested: Optimize the phylogenetic tree constructed in step (2) using the tetra-degenerate sites in the single-copy homologous genes to obtain the optimized phylogenetic tree; thereby obtain the branch length D of the species to be tested. species And the branch length D of the closest relative to the species being tested. phylogenetic ; (5) Obtain the annual phylogenetic mutation rate μ of the species to be tested. year : ① If the divergence time between the target species and the closest related species is less than 5 Mya, the annual phylogenetic mutation rate μ of the target species is [not specified]. year The calculation formula in Eq.1 is used to obtain Eq.1; Where k is the homologous pseudogene sequence divergence degree, and θπ is the nucleotide polymorphism of the target species; ② If the divergence time between the target species and the closest selected species is within 5-20 Mya, the annual phylogenetic mutation rate μ of the target species year The result is obtained using the formula Eq.2; Eq.2; ③ If the divergence time between the target species and its closest related species is between 20 and 200 Mya, the annual phylogenetic mutation rate of the target species is calculated using the formula in Eq.
3. Eq.3; Among them, Ks mid The median number of synonymous substitutions of all homologous single-copy genes between the species to be tested and the closest related species in the phylogenetic tree constructed in step (2).
2. The method for measuring the annual phylogenetic mutation rate of sexually reproducing animals according to claim 1, characterized in that: The homologous pseudogene sequence divergence degree k in step (3) ① is obtained by the following method: First, obtain the pseudogene sequence of the species to be tested and the closest related species in the phylogenetic tree constructed in step (2), then obtain the homologous pseudogene sequence by comparison, count all comparison results, and calculate the k value using the Kimura dual-parameter model; The nucleotide polymorphism θπ of the target species is obtained by the following method: resequencing individuals of at least three species to be tested to obtain single nucleotide polymorphism sites at the population level, and then calculating the θπ value using population genetic statistical tools.
3. The method for measuring the annual phylogenetic mutation rate of sexually reproducing animals according to claim 1, characterized in that: The Ks mentioned in step (3) ③ mid The calculation incorporates all degenerate codons.
4. The method for measuring the annual phylogenetic mutation rate of sexually reproducing animals according to any one of claims 1-3, characterized in that: The calibration time nodes mentioned in step (3) are the fossil evidence time and / or geological event time of the species to be tested. The number of calibration time nodes is not less than two, and the calibration time nodes are obtained from publicly available paleontological and geological literature or the TimeTree 5 database.
5. The method for measuring the annual phylogenetic mutation rate of sexually reproducing animals according to claim 4, characterized in that: Step (3) is as follows: First, the single-copy homologous gene is divided into 3 partitions according to the position of the codon, and the divergence time is calculated independently according to different partitions; the independent rate molecular clock model is selected for the relaxation molecular clock method; the universal time reversible (GTR) model is selected for the nucleotide substitution model; the maximum divergence time is set to 200 Mya according to the applicable scope of this patent.
6. The method for measuring the annual phylogenetic mutation rate of sexually reproducing animals according to claim 5, characterized in that: It also includes step (5) verification: the annual germplasm mutation rate obtained in step (4) is compared with existing methods for verification.
7. Application of the annual germline mutation rate obtained by the method described in any one of claims 1-6.
8. A method for measuring the phylogenetic mutation rate of sexually reproducing animals, characterized in that: The mutation rate of the lineage was obtained using the formula in Eq.4: μ generation =μ year × T generation Eq.4; Where, μ year T is the annual phylogenetic mutation rate obtained using the method described in any one of claims 1-5. generation To indicate the generation time of the target species.
9. Application of the line mutation rate obtained by the method described in claim 8.
Citation Information
Patent Citations
Analysis method for multi-sample-size comparative transcriptomes based on third-generation sequencing detection
CN111161797A
Methods and systems for differentiating somatic and germline variants
CN111357054A