Analysis method for revealing provenance and diffusion of tropical plants based on chloroplast genome
By optimizing the chloroplast genome data processing workflow and combining various software tools for phylogenetic and time-tree construction, the problem of analytical accuracy in high-similarity or low-coverage species populations has been solved by traditional methods. This has enabled precise analysis of plant origin and dispersal, and promoted biogeographical research.
Patent Information
- Application Number
- CN202511869892.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-01-09
AI Technical Summary
Existing chloroplast genome analysis methods struggle to provide accurate phylogenetic and temporal tree estimates when dealing with species populations with high similarity or low coverage. Traditional species regional role theories are limited to "source" and "sink," failing to effectively expand and differentiate the dualistic characteristics of biogeography.
The analysis method based on chloroplast genomes was adopted, including obtaining complete genome data from NCBI, optimizing assembly parameters, using MAFFT and RAxML software for sequence alignment and phylogenetic tree construction, combining PhyloSuite and treePL software for time tree construction, and using BioGeoBEARS for ancestral region reconstruction. Multiple biogeographical models were used for analysis.
It improved the accuracy and applicability of the analysis, quantified the role of different regions in diversity combinations, promoted the research progress of plant phylogenetics and biogeography, and provided a reproducible and scalable research paradigm.
Smart Images

Figure CN121306263A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plant analysis technology, and more specifically to an analytical method based on chloroplast genomes to reveal the origin and spread of tropical plants. Background Technology
[0002] Species radiation plays a significant role in evolutionary theory, particularly in island ecosystems. However, existing research largely relies on standard genome assembly and alignment methods, especially in chloroplast genome analysis, which has several limitations. Traditional methods such as MAFFT and RAxML often struggle to provide accurate phylogenetic and temporal tree estimates when dealing with highly similar or low-coverage species populations, resulting in low precision in the analytical results. Furthermore, traditional theories of species regionalization only focus on "source" and "sink," failing to recognize that migration and emigration are not opposing processes, thus expanding and distinguishing the binary characteristics widely used in biogeography. The complete chloroplast genome information of plants is invaluable for exploring phylogenetic relationships between species, inferring evolutionary history patterns, and for disseminating complete chloroplast genome data and evolutionary patterns. Summary of the Invention
[0003] To achieve the above objectives, this invention proposes an analytical method based on chloroplast genomes to reveal the origin and spread of tropical plants, including: Step S1: Obtain the complete chloroplast genome data of the plant to be analyzed that has not been uploaded from NCBI, in FASTQ or FASTA format; Step S2: In low-coverage or complex genome samples, preprocessing steps are used to improve assembly quality, and assembly parameters are optimized according to species characteristics; Step S3: Using the plant chloroplast genome of the plant to be analyzed that has been uploaded to NCBI as a reference, annotate the chloroplast genome sequence and check the annotation results; Step S4: Perform sequence alignment of the entire chloroplast genome to identify homologous regions and evolutionary variations; Step S5: Compare the sequences of multiple chloroplast genomes of the plant to be analyzed with the whole chloroplast genomes of two outgroup species to construct a phylogenetic tree; Step S6: Using the phylogenetic tree constructed in S5, construct a time tree using the ancient fossil nodes of the plant to be analyzed; Step S7: Use time trees to collect data on the modern distribution areas of the plants to be analyzed, and use different models to perform ancestral region reconstruction analysis.
[0004] Step S8: Based on the species diffusion areas and the number of species diffusions obtained in the ancestral region reconstruction, partition them to obtain a pattern for shaping overall diversity.
[0005] Further, step S5 includes: S51. Use MAFFT software to align chloroplast genome sequences, and select an alignment strategy based on the data size or optimize local alignment accuracy through custom parameters. S52. Use the trimAI function in PhyloSuite software to cut the blank area of the sequence after alignment in step S51 to obtain a complete aligned sequence. S53. Using outgroups, construct the maximum likelihood phylogenetic tree using the WorkflowonXSEDE model in RAxML software, and perform guided replication analysis to verify the reliability of the maximum likelihood phylogenetic tree.
[0006] Furthermore, in step S51, when performing sequence alignment using MAFFT software, the alignment strategy is adaptively selected based on the dataset size: For small datasets, use G-INS-i or L-INS-i high-precision strategies; For medium to large datasets, the auto strategy is adopted to automatically determine sequence differences and select the optimal alignment algorithm; When processing sequences with high similarity or containing repetitive regions, an enhanced local alignment strategy is adopted, which optimizes alignment accuracy by customizing gap opening and extending penalty parameters.
[0007] Further, step S6 includes: S61. Using the same phylogenetic tree file as in step S5, generate a constraint tree file suitable for time-based tree construction by setting the workflow, printing branch lengths, and introducing evolutionary relationship constraints. S62. Prepare the configuration file for the treePL software. The configuration file includes the data input directory, node time calibration points and their time ranges set based on fossil evidence, optimization parameters, random subsampling and cross-validation parameters, smoothing values, and result output directory. S63. First, run the run1 configuration file to obtain preliminary optimization parameters. Then, run the run2 configuration file based on the results of run1. Determine the optimal smoothing value through multiple iterations. Finally, run the run3 configuration file using the optimal smoothing value to generate a collection file containing 1000 time trees. S64. Use a tree merging tool to merge the 1000 time trees into a consensus time tree with a branch time median, which is used to represent the differentiation time of species.
[0008] Further, step S7 includes: S71. Collect the modern distribution of the plant species to be analyzed based on the global biodiversity information agency database, divide the modern distribution into multiple regions according to biogeographical zoning, encode the distribution of each species using a binary method, and generate a geographic data file containing the species and its distribution area codes. S72. In the R language environment, using the BioGeoBEARS package, input the time tree obtained in step S6 and the geographic data file generated in step S71 to perform ancestral region reconstruction analysis. S73. Six biogeographical models, DEC, DEC+J, DIVALIKE, DIVALIKE+J, BAYAREALIKE, and BAYAREALIKE+J, were used for analysis. By calculating and comparing the Akaike information content criterion values and their weights for each model, the optimal model was selected, and the ancestral distribution history of the plants to be analyzed was reconstructed based on the results of the optimal model.
[0009] Furthermore, the preprocessing step in step S2 includes: S21. Convert the downloaded SRA format data to FASTQ format; S22. Randomly select a portion of reads from the converted data; S23. Perform quality control on randomly selected reads and filter out low-quality sequences; S24. After activating the getorganelle assembly environment, de novo assembly of the chloroplast genome is performed using quality-controlled reads.
[0010] Further, step S3 includes: S31. Convert the FASTA format file to obtain the SEQ format file.
[0011] S32. Error correction is performed by adding or removing bases, including start codons, stop codons, and repeating sequences, and finally the correct FASTA and GB format files are output.
[0012] Further, step S1 includes: S11. Collect the SRA sequence numbers of the plant species to be analyzed; S12. Use the Prefetch tool to download SRA format data and convert it to FASTQ format.
[0013] Further, step S8 includes: S81. Based on the ancestral region reconstruction and the obtained species diffusion direction, the 7 regions are divided into immigrant species, emigrant species, and in-situ diversified species, and the quantity and proportion statistics are carried out.
[0014] S82. Use the ggplots package in R to perform regional analysis on the obtained proportional data.
[0015] The beneficial effects of this invention are: This invention constructs and integrates an efficient and comprehensive workflow, covering key steps from the assembly and annotation of the whole chloroplast genome of a species, to the inference of phylogenetic trees, the estimation of time trees, and the reconstruction of ancestral regions and the shaping of diversity by regional effects. This workflow efficiently links these steps into a unified process, forming a reproducible and scalable paradigm. This paradigm not only provides a clear execution path and reference standard for phylogenetic analysis based on chloroplast genomes, but also makes innovative improvements in algorithm optimization, parameter adjustment, and data integration, enhancing the accuracy and applicability of the analysis. It possesses high portability and reproducibility, and innovates the "source and sink" theory, quantifying the role of different regions in diversity combinations—the continental shelf is a radiator of diversity, archipelagos are diffusion corridors, and continental margins are accumulators of diversity. This workflow can serve as a precedent and template for similar studies, promoting research progress in plant phylogeny and biogeography. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0017] Figure 1 This is a phylogenetic diagram of chloroplasts in the genus *Syzygium* according to the present invention.
[0018] Figure 2 This is a tree diagram illustrating the origin timeline of the genus *Syzygium* in this invention.
[0019] Figure 3 This is a reconstruction map of the ancestral region of the genus Syzygium in this invention.
[0020] Figure 4 This is a diagram illustrating the diversity of roles within the Syzygium genus region in this invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0023] This invention provides an analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes. Specifically, the plant to be analyzed is the island species *Syzygium spp.*, and the analytical steps are described below: Step S1: Download complete chloroplast genome data of *Syzygium* species from the NCBI database, ensuring the data are complete and of acceptable quality, in FASTQ or FASTA format. Obtain SRA data using the Prefetch tool and convert it to the FASTQ format suitable for subsequent analysis.
[0024] NCBI has developed public databases such as Genbank and provides tools such as PubMed, BLAST, Entres, OMIM, Taxonomy, and Structure, which can be used to search and analyze international molecular databases and biomedical literature. It has also developed software tools for analyzing genomic data and disseminating biomedical information.
[0025] The required species information is manually retrieved and downloaded for subsequent processing and research. The specific steps are as follows: S11, SRA serial number of the collected species; S12. Download the specified sequence using prefetchSRRxxxxxxxx. After downloading, you will get an xx.sra file, which needs to be converted into a fastq file before it can be used.
[0026] Step S2: Chloroplast genome assembly. In low-coverage or complex genome samples, preprocessing steps are used to improve assembly quality. Assembly parameters are optimized according to species characteristics to ensure the integrity and accuracy of genome assembly.
[0027] All downloaded SRA data needs to be transformed, extracted, and analyzed before assembly. The specific steps are as follows: S21 and SRA format data are converted to fastq format for subsequent analysis of fastq files. S22. Randomly selecting a portion of reads can improve analysis efficiency, reduce time costs, and improve assembly quality. S23. The extracted reads are controlled to ensure the quality of the data in this invention; S24. The assembly of the chloroplast genome first requires activating the assembly environment, and then merging and assembling the prepared files.
[0028] Step S3: Chloroplast genome annotation. Using the *Syzygium* chloroplast genome uploaded to NCBI as a reference, annotate the sequences and check the annotation results. Specifically, this includes: S31. Convert the FASTA format file to obtain the SEQ format file.
[0029] S32. Use Sequin V16.0 software to correct errors by adding or removing bases, including start codons, stop codons, and repeating sequences, and finally output the correct FASTA and GB format files.
[0030] Step S4: Perform sequence alignment of the entire chloroplast genome. Integrate all Syzygium genus and outgroup data in FASTA format and perform MAFFT alignment to find homologous points and align evolutionary variations. Specific steps are as follows: S41. Standardize the file content format for the genus *Prunus* and its outgroups, including genus name and species name, and finally output a FASTA format file.
[0031] S42. Merge all species with unified names and perform MAFFT comparison.
[0032] Specifically, this involves using the MAFFT alignment program in PhyloSuite V1.2.3 software. This program is suitable for most situations, performs well on sequences with high similarity and moderate differences, and has a fast computation speed. The `--auto` parameter is often used to automatically determine sequence differences and select the optimal algorithm, or the MAFFT program can be downloaded, and the command (MAFFT input file .fasta > output file .fasta) can be entered to obtain a two-sided aligned sequence. The resulting MAFFT_results results are then processed using the `trimAI` function to remove missing parts, reducing inaccuracies caused by gaps in subsequent operations.
[0033] It should be noted that if most pairs have high similarity, too many overlapping regions, uniform length distribution, and obvious consistency in local regions, they are judged as having high similarity, and enhanced local alignment will be used in this case.
[0034] S43. Enhanced local alignment for repetitive regions.
[0035] Specifically, this includes adjusting global alignment parameters to enhance local alignment effects, slightly reducing the gap open parameter to allow areas with locally weakened gaps to be aligned more easily (the default value is usually 1, so it should be changed to between 0.5 and 1.0); and setting the gap extend parameter to a medium level to avoid highly expanding the gap in non-repeating areas (the value is usually set between 0.5 and 1.0, which is relatively low to allow short segments to be aligned in repeating areas without being overly penalized).
[0036] Step S5: Compare the chloroplast genomes of two hundred species in the genus Syzygium with the chloroplast genome sequence information of two outgroups (Syzygium spp. and Syzygium var. spp.), and construct a phylogenetic tree for the genus Syzygium using the RAxML-HPC2WorkflowonXSEDE model.
[0037] In steps S51 and S4, a FASTA format result file is obtained. The correct model is selected, meaning the mathematical formula that best fits the current sequence evolution pattern is found. Specific steps include: S511, Input File: Import the optimized alignment file into the model selection tool.
[0038] S512. Use the nucleotide substitution model selection tool, set the nucleotide model parameters, and the tool will calculate the most suitable substitution model for the dataset, i.e., the optimal model (such as GTR+G, HKY+I+G).
[0039] S52. Construct a phylogenetic tree using the maximum likelihood method.
[0040] S53. Select the RAxML-HPC2WorkflowonXSEDE model, input the alignment file of this invention, and set its parameters: Maximum Hours to Run: Initially set to a shorter time (less than 0.5) to verify the process's functionality. Once configured correctly, the run time can be increased. Select your workflow: ML + Thorough bootstrap; Outgroup: Enter the selected outgroup, separated by commas; otherwise, the program will fail.
[0041] Bootstrapiterations: Setting this to 1000 will affect the accuracy of the results. The other options are left as default. The final phylogenetic tree of the genus Syzygium is obtained. The Bootstrap value of the phylogenetic tree result is greater than 70% in all cases. This is used to verify and judge the branch stability of the phylogenetic tree of the present invention.
[0042] This invention first ensures that the analyzed sequences are homologous, then aligns homologous sites, uses a model to simulate real evolutionary patterns, uses an algorithm to transform sequence differences into a kinship tree, and finally verifies the reliability of the tree.
[0043] like Figure 1 The phylogenetic tree of the genus *Syzygium* is shown. It is a maximum likelihood phylogenetic tree based on the complete chloroplast genome sequences of 200 *Syzygium* species and 2 outgroups. The tree was constructed using RAxML and GTRGAMMA models with 1000 fast bootstrap replicates. Nodes with bootstrap support values greater than 70% are marked in the figure (see main text). Figure 1 .
[0044] Step S6: Using the phylogenetic tree constructed in S5, construct a time tree using ancient fossil nodes of the genus Syzygium.
[0045] First, a tree file suitable for treePL is constructed using CIPRESScienceGateway. Then, a timeline of the genus Syzygium is constructed using treePL, combining the phylogenetic tree with collected fossil nodes.
[0046] S61 uses the same model and operating steps as S5, but there are three differences in the parameter settings: Select your workflow: Bootstrap + Consensus. This model helps the invention filter out random errors and retain true evolutionary signals. Selecting the option to print branch lengths (-k) helps visualize the branch length values in this invention; Select Constraint(-g). This option requires the output result RAxML_bestTree.testB in step S5 to add a reliable evolutionary relationship as a priori, which helps the present invention solve the problem that the tree structure cannot reflect the true evolutionary history when there is insufficient data signal or conflict. The output result is RAxML_bipartitions.MLTB_output.
[0047] S62. Prepare the configuration files required by treePL, including the data input directory, molecular fossil setting time and range, optimal optimization parameters, random subsampling and repeated cross-validation, smoothing values, and output directory; S63, the configuration files include run1, run2, and run3, with the following differences in their contents: treefile= / home / train / zby_temp / besttree.tre#Input files containing the ML trees numsites=46051#General commands nthreads=20 thorough log_pen mrca=ARALIACEAE Schefflera_actinophylla Raukaua_simplex#Calibrations min=ARALIACEAE 38.3 max=ARALIACEAE 112 mrca=TORRICELLIACEAE Torricellia_tiliifolia Melanophylla_alnifolia min=TORRICELLIACEAE 48.5 max=TORRICELLIACEAE 122 mrca=APIALES Pennantia_cunninghamii Raukaua_simplex min=APIALES 69.9 max=APIALES 95.9 prime#Execute prime in the first round #opt=3#Best optimisation parameters #more detail #optad=3 #more detail ad #optcvad=5 #randomcv#Cross-validation analysis #cviter=5 #cvsimaniter=1000000000 #cvstart=100000 #cvstop=0.000000000001 #cvmultstep=0.1 #cvoutfile= / home / train / zby_temp / randomcv_PEN.txt #smooth=0.000001#Best smoothing value #outfile= / home / train / zby_temp / PEN_treePL_trials_dated.tre In the run1 file, the tree file to be input is the result of S5, RAxML_bestTree.testB. Fossil nodes need to be located between different species where they are located, including the maximum and minimum time points (the essence of fossil selection is to find a fossil with a specific survival age and map it to the branch node of the tree. The phylogenetic relevance between the fossil and the target branch, the reliability of the fossil time, and the time constraint of the fossil should be ensured). There should be at least one fossil node. In this invention, a fossil node of 20.9 - 22.1 Ma is selected to be added within Syzygium_luehmanniiSyzygium_hemilamprum.
[0048] By inputting the command (nohuptreePLSyzygium_treePLconfig_run1.txt>run1_1.log 2>&1&), the output result of run1 can obtain the three values required in prime; In the run2 file, the tree file to be input is the result of S5, RAxML_bestTree.testB. At this time, input the opt, optad, and potcvad values obtained from run1. By repeating the run2 step multiple times, multiple different output results will be obtained, and their contents will change multiple times. The value with the smallest and most occurrences among the final multiple results is the required result.
[0049] The number corresponding to the value with the smallest and most occurrences in the output result of inputting the command (nohuptreePLSyzygium_treePLconfig_run2.txt>run2_1.log 2>&1&) is the best smooth value; In the run3 file, the tree file to be input is the RAxML_bipartitions.MLTB_output tree file in S61, and input the best smooth value obtained from run2.
[0050] By inputting the command (nohuptreePLSyzygium_treePLconfig_run3.txt>run3.log 2>&1&) and running it, 1000 trees are obtained; Generate a file containing 1000 time trees (including species names and the time of their appearance). Use the treeannotator command (treeannotator-burnin0-heightmeanPEN_treePL_trials_dated.tretest.out.tre) to merge the 1000 trees, and finally obtain the time tree. The time tree can help this invention understand the origin time of this species as well as the time of its speciation and spread.
[0051] S64. Using Figtree visualization software, select the HPD option to obtain the confidence range of the true value of the parameter described by probability intervals. Table 1 is a list of confidence ranges of the true value of the parameter described by probability intervals.
[0052] Table 1: List of Confidential Ranges for Describing True Parameter Values Using Probability Intervals , , , , , .
[0053] like Figure 2 The timeline of the entire *Syzygium* genus is shown. Divergence times of the *Syzygium* genus and outgroups were estimated based on 202 complete chloroplast genomes. Time-corrected phylogenetics were generated using treePL. Horizontal bars represent the 95% highest posterior density (HPD) intervals for well-supported lineages.
[0054] Step S7: Use a time tree to search for the current distribution areas of the genus Syzygium and select different models to reconstruct and analyze its ancestral regions.
[0055] The current global distribution of *Syzygium* species was obtained using GBIF, an open-source platform providing global biodiversity data that stores extensive information on species distribution, classification, and specimens. Ancestor regions of the *Syzygium* genus were reconstructed using a combination of time-based trees and distribution locations. The specific steps are as follows: S71. Using GBIF, the pre-distribution locations of *Syzygium* species are collected. Based on the distribution patterns of ancient continents due to sea route isolation and differences in biological communities, and considering that the selected species are currently distributed worldwide, they are divided into seven regions: (A) Asia, (B) Sunda Shelf, (C) Wallace Region, (D) Sahual Continent, (E) Pacific Islands, (F) Africa, and (G) India. The distribution regions are encoded using a binary classification method (e.g., Asia is 1000000) and a data format file is generated, for example: 6 7(ABCDEFG) Syzygium_acre 0001000 Syzygium_acuminatissimum 1011010 Syzygium_adelphicum 1000000 Syzygium_alatum 1000000 Syzygium_amplifolium 0001000 Syzygium_amplum 1001000 S72. Using the BioGerBEARS package in R, input the time tree from S6 and the geographic information from S71, while eliminating outgroups, and perform operations according to the contents of the package.
[0056] S73. Using six models—DEC, DEC+J, DIVALIKE, DIVALIKE+J, BAYAREALIKE, and BAYAREALIKE+J—the AIC value was calculated using the AICtable function. The values with the lowest AIC (Akaike Information Criterion) and the highest AIC_wt (Weighted Akaike Information Criterion) were selected. The BAYAREALIKE+J model was chosen (see Table 2). The geographical distribution evolution of this group conforms to the principles of small-scale dispersal priority and founder events, indicating an island-diffusion group, which matches the differentiation of the *Syzygium* genus. The ancestral region reconstruction map was finally obtained.
[0057] Further instructions: Use the following command AICtable=calc_AIC_column(LnL_vals=restable$LnL,nparam_vals=restable$numparams); restable=cbind(restable,AICtable); restable_AIC_rellike=AkaikeWeights_on_summary_table(restable=restable,colname_to_use="AIC"); restable_AIC_rellike=put_jcol_after_ecol(restable_AIC_rellike); conditional_format_table(restable_AIC_rellike).
[0058] First, based on LnL and the number of parameters, the AIC is calculated and added to the data table restable. Then, the AIC values are used to calculate the Akaike weights and format the output to evaluate which model is better supported under the AIC framework. Table 2 shows the AIC and AIC_wt data lists for the six running models of BioGerBEARS.
[0059] Table 2: List of AIC and AIC_wt data for six BioGerBEARS operating models
[0060] like Figure 3 The ancestral regions of *Syzygium* species are shown in the diagram. A time-calibrated phylogenetic tree of 200 *Syzygium* species based on complete chloroplast genomes was constructed, and the ancestral distribution regions were reconstructed using the BAYAREALIKE+J model in BioGeoBEARS within a Bayesian framework.
[0061] Step S8: Based on the seven regions reconstructed from the ancestral region and the obtained species dispersal directions, species are categorized into invasive species, emigrant species, and in-situ diversified species. A diversity analysis of the *Syzygium* genus is conducted using quantitative and proportional statistics. The specific steps are as follows: S81, S81, Statistical analysis of species composition in ancestral region reconstruction. For example, for region A, first obtain the total number of all species in the region; based on the data in ancestral region reconstruction, obtain the number of species that migrated to region A from other regions; obtain the number of species that migrated out of region A; and the in-situ diversification is calculated by subtracting the number of species that migrated into region A from the total number of species in region A. Table 3 is a list of embedded migration data.
[0062] Table 3: List of Embedded Migration Data
[0063] S82. Input the statistical data according to the requirements of the ggplots package in R. The x-axis represents the species emigration rate and the y-axis represents the speciation rate. At the same time, set the radiation zone (S ≥ 0.5, E ≥ 0.5), hatching zone (S ≥ 0.5, E<0.5), corridor zone (S<0.5, E ≥ 0.5), and accumulation zone (S<0.5, E<0.5).
[0064] like Figure 4 The area shown represents the role of roles in shaping species diversity.
[0065] Thus, this invention covers key steps from the assembly and annotation of the whole chloroplast genome of a species, to the inference of phylogenetic trees, the estimation of time trees, and the reconstruction of ancestral regions. By combining these steps, it can simultaneously answer questions such as "who evolved into their current form, when, and from where, how they spread geographically, the influence of historical environment on this process, and the contribution of regional roles to shaping biodiversity," thereby better analyzing the origin of species, dispersal pathways, and the formation mechanisms of regional diversity.
[0066] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes, characterized in that, include: Step S1: Obtain the complete chloroplast genome data of the plant to be analyzed that has not been uploaded from NCBI, in FASTQ or FASTA format; Step S2: In low-coverage or complex genome samples, preprocessing steps are used to improve assembly quality, and assembly parameters are optimized according to species characteristics; Step S3: Using the plant chloroplast genome of the plant to be analyzed that has been uploaded to NCBI as a reference, annotate the chloroplast genome sequence and check the annotation results; Step S4: Perform sequence alignment of the entire chloroplast genome to identify homologous regions and evolutionary variations; Step S5: Compare the chloroplast genome of the plant to be analyzed with the whole chloroplast genome of outgroup species to construct a phylogenetic tree; Step S6: Using the phylogenetic tree constructed in S5, construct a time tree using the ancient fossil nodes of the plant to be analyzed; Step S7: Use time trees to collect data on the modern distribution areas of the plants to be analyzed, and use different models to reconstruct the ancestral regions. Step S8: Based on the combination of species immigration and emigration rates, innovatively construct new bioregion combination models.
2. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 1, characterized in that, Step S5 includes: S51. Use MAFFT software to align chloroplast genome sequences, and select an alignment strategy based on the data size or optimize local alignment accuracy through custom parameters. S52. Use the trimAI function in PhyloSuite software to cut the blank area of the sequence after alignment in step S51 to obtain a complete aligned sequence. S53. Using outgroups, construct the maximum likelihood phylogenetic tree using the WorkflowonXSEDE model in RAxML software, and perform guided replication analysis to verify the reliability of the maximum likelihood phylogenetic tree.
3. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 2, characterized in that, In step S51, when using MAFFT software for sequence alignment, the alignment strategy is adaptively selected based on the dataset size: For small datasets, use G-INS-i or L-INS-i high-precision strategies; For medium to large datasets, the auto strategy is adopted to automatically determine sequence differences and select the optimal alignment algorithm; When processing sequences with high similarity or containing repetitive regions, an enhanced local alignment strategy is adopted, which optimizes alignment accuracy by customizing gap opening and extending penalty parameters.
4. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 1, characterized in that, Step S6 includes: S61. Using the same phylogenetic tree file as in step S5, generate a constraint tree file suitable for time-based tree construction by setting the workflow, printing branch lengths, and introducing evolutionary relationship constraints. S62. Prepare the configuration file for the treePL software. The configuration file includes the data input directory, node time calibration points and their time ranges set based on fossil evidence, optimization parameters, random subsampling and cross-validation parameters, smoothing values, and result output directory. S63. First, run the run1 configuration file to obtain preliminary optimization parameters. Then, run the run2 configuration file based on the results of run1. Determine the optimal smoothing value through multiple iterations. Finally, run the run3 configuration file using the optimal smoothing value to generate a collection file containing 1000 time trees. S64. Use a tree merging tool to merge the 1000 time trees into a consensus time tree with a branch time median, which is used to represent the differentiation time of species.
5. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 1, characterized in that, Step S7 includes: S71. Collect the modern distribution of the plant species to be analyzed based on the global biodiversity information agency database, divide the modern distribution into multiple regions according to biogeographical zoning, encode the distribution of each species using a binary method, and generate a geographic data file containing the species and its distribution area codes. S72. In the R language environment, using the BioGeoBEARS package, input the time tree obtained in step S6 and the geographic data file generated in step S71 to perform ancestral region reconstruction analysis. S73. Six biogeographical models, DEC, DEC+J, DIVALIKE, DIVALIKE+J, BAYAREALIKE, and BAYAREALIKE+J, were used for analysis. By calculating and comparing the Akaike information content criterion values and their weights for each model, the optimal model was selected, and the ancestral distribution history of the plants to be analyzed was reconstructed based on the results of the optimal model.
6. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 1, characterized in that, The preprocessing step in step S2 includes: S21. Convert the downloaded SRA format data to FASTQ format; S22. Randomly select a portion of reads from the converted data; S23. Perform quality control on randomly selected reads and filter out low-quality sequences; S24. After activating the getorganelle assembly environment, de novo assembly of the chloroplast genome is performed using quality-controlled reads.
7. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 1, characterized in that, Step S3 includes: S31. Convert the FASTA format file to obtain the SEQ format file; S32. Error correction is performed by adding or removing bases, including start codons, stop codons, and repeating sequences, and finally the correct FASTA and GB format files are output.
8. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 1, characterized in that, Step S1 includes: S11. Collect the SRA sequence numbers of the plant species to be analyzed; S12. Use the Prefetch tool to download SRA format data and convert it to FASTQ format.
9. The analytical method for revealing the origin and dispersal of tropical plants based on chloroplast genomes according to claim 1, characterized in that, Step S8 includes: S81. Based on the ancestral region of S7, the number of species emigrating out, emigrating in, and the number and proportion of native species formed are obtained. S82. Use the ggplot2 package in R to analyze it.
Citation Information
Patent Citations
Erieaeeae plant analysis method based on chloroplast genome sequencing
CN113308525A