Method for developing shellfish germplasm resource evaluation molecular marker
By developing a whole-genome genetic variation site filtering strategy, molecular markers suitable for shellfish germplasm resource assessment were screened out, solving the problem of molecular marker screening at the whole-genome level, realizing efficient and low-cost large-scale shellfish germplasm resource assessment, and providing a scientific basis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- OCEAN UNIV OF CHINA
- Filing Date
- 2026-03-14
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies struggle to effectively screen for high-quality molecular markers at the whole-genome level, making it difficult to achieve a balance between the reliability and quantity of genetic variation sites in the assessment of shellfish germplasm resources. This affects the assessment of population structure, genetic diversity, and effective population size.
By developing a series of whole-genome genetic variation site filtering strategies and using high-throughput sequencing data analysis tools such as BCFTOOLS and VCFTOOLS, molecular markers suitable for different genetic assessments were screened, including parameters such as sequencing data alignment quality, genotyping quality, sequencing depth, genotype information missing ratio, and minimum allele frequency. A genetic variation site dataset was then constructed to assess population structure, genetic diversity, and effective population size.
It enables the acquisition of high-quality and sufficient molecular markers at the whole genome level, meeting the needs of different genetic analyses, providing reliable assessment of shellfish germplasm resources, and providing scientific basis for the protection of natural populations, improvement of breeding populations, and development of aquaculture populations.
Smart Images

Figure CN121862201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-quality whole-genome genetic variation site capture and shellfish germplasm resource assessment in shellfish, and specifically to a method for developing molecular markers for shellfish germplasm resource assessment. Background Technology
[0002] Shellfish are a key component of aquatic ecosystems and a major source of high-quality animal protein. On the one hand, shellfish help purify water, accelerate carbon, nitrogen, and phosphorus cycles, and provide diverse habitats for aquatic organisms, which is crucial for maintaining the function and diversity of aquatic ecosystems. On the other hand, shellfish are delicious, nutritious, and have extremely high economic value; shellfish farming has become an important way to obtain high-quality aquatic products. However, the extremely high ecological and economic value of shellfish has led to various challenges posed by human activities, including but not limited to global warming, habitat destruction, overfishing, high-density farming, and germplasm degradation, resulting in endangered natural populations and depletion of germplasm resources. Therefore, systematically assessing the population status of natural, breeding, and farmed shellfish populations from a genetic perspective is a prerequisite for the sustainable protection, management, development, and utilization of the ecological and economic value of shellfish. The rapid development of high-throughput sequencing technology has generated a large number of molecular markers, laying the foundation for efficient, low-cost, and large-scale assessment of shellfish germplasm resources. However, the genetic variation sites obtained through high-throughput sequencing contain a large number of potential errors. Eliminating these errors is a prerequisite for ensuring the reliability of genetic assessment of shellfish populations. Therefore, constructing an accurate and efficient whole-genome genetic variation site filtering strategy is an important prerequisite for obtaining high-quality molecular markers and conducting shellfish germplasm resource assessment.
[0003] Due to the biological and ecological characteristics of mollusks, capturing molecular markers for evaluating mollusk germplasm resources at the whole-genome level faces numerous challenges. First, the extremely high complexity of mollusk genomes leads to a significant reduction in the alignment quality of sequencing data, posing stricter requirements for obtaining reliable molecular markers through whole-genome genetic variation filtering. Second, the planktonic nature of mollusk larvae and the high genetic connectivity and low degree of genetic differentiation among different populations greatly increase the difficulty of population structure assessment, leading to uncertainties in genetic analyses that heavily rely on clear population structure. Therefore, obtaining molecular markers that can accurately resolve mollusk population structure is essential. Third, mollusks exhibit extremely high genetic diversity, the level of which is influenced by various factors. Using different molecular markers for genetic diversity assessment may yield diametrically opposed conclusions, making it crucial to identify molecular markers that accurately reflect the level of genetic diversity. Fourth, influenced by the natural environment and human activities, the effective population size of mollusks frequently experiences significant fluctuations. The quantity, quality, and distribution characteristics of molecular markers have a significant impact on estimating the dynamic changes in the effective population size of mollusks. However, currently, no whole-genome genetic variation filtering strategy suitable for mollusks has been established to address different genetic assessment needs. Lower filtering standards make it difficult to guarantee the reliability of genetic variation sites; while overly strict screening thresholds can improve data quality, they can significantly reduce the number of genetic variation sites and the number of samples, resulting in different degrees and patterns of bias in different downstream analyses.
[0004] Therefore, how to obtain high-quality, sufficient molecular markers that meet the needs of different genetic analyses at the whole genome level by balancing the relationship between the reliability of genetic variation sites, the number of genetic variation sites, and the number of samples is an urgent problem to be solved in the systematic evaluation of mollusk germplasm resources based on high-throughput sequencing. Summary of the Invention
[0005] The purpose of this invention is to establish a method for developing molecular markers for evaluating shellfish germplasm resources, in order to overcome the shortcomings of existing technologies.
[0006] This invention systematically evaluates the impact of a series of genome-wide genetic variation site filtering strategies on downstream genetic analysis, develops molecular markers suitable for different genetic assessments, and achieves efficient, low-cost, and large-scale assessment of shellfish germplasm resources. It provides scientific standards for the protection of natural populations, improvement of breeding populations, and development of aquaculture populations, and provides key information for the sustainable discovery and enhancement of the ecological and economic value of shellfish.
[0007] The assessment of shellfish germplasm resources can be carried out from three aspects: (1) population structure assessment, (2) genetic diversity assessment, and (3) effective population size assessment. High-throughput sequencing can generate a large number of genetic variation sites at the whole genome level. Developing personalized genetic variation site filtering strategies for different genetic analysis requirements can obtain molecular markers suitable for different genetic assessments.
[0008] A method for developing molecular markers for evaluating shellfish germplasm resources includes: collecting representative individuals from the population to be evaluated and performing high-throughput sequencing at the whole-genome level; obtaining genetic variation sites through high-throughput sequencing data analysis; formulating a series of whole-genome genetic variation site filtering strategies to obtain a series of genetic variation site datasets; evaluating population structure, genetic diversity, and effective population size based on each genetic variation site dataset; and determining high-quality molecular markers suitable for shellfish germplasm resource evaluation by evaluating the impact of different genetic variation site filtering strategies on different genetic evaluations.
[0009] Furthermore, the method for developing molecular markers for evaluating shellfish germplasm resources specifically includes the following steps: (1) Collect representative populations of the shellfish population to be evaluated; (2) Conduct high-throughput sequencing at the whole genome level; (3) Obtain genome-wide genetic variation sites using high-throughput sequencing data analysis tools; (4) Based on the actual sequencing situation and genetic evaluation requirements, formulate a series of whole-genome genetic variation site filtering strategies representing different screening criteria, and obtain the corresponding genetic variation site dataset based on the strategy; (5) Use datasets of different genetic variation sites to conduct population structure assessment, genetic diversity assessment and effective population size assessment; (6) By evaluating the performance of different genetic variation sites in different genetic assessments, multiple datasets were selected, which are molecular markers suitable for the assessment of shellfish germplasm resources.
[0010] Furthermore, in steps (1), (2), and (3), the whole genome genetic variation sites are obtained as follows: at least 30 representative individuals are randomly collected from each population to be evaluated, and high-throughput sequencing at the whole genome level with an average sequencing depth ≥10× is carried out on the representative individuals of the population to be evaluated. The whole genome genetic variation sites are obtained using high-throughput sequencing data analysis tools (e.g., BCFTOOLS, Genome Analysis Toolkit, etc.).
[0011] Furthermore, in step (4), the data acquisition method for the genetic variation sites representing different screening criteria is as follows: the whole-genome genetic variation site filtering strategy is formulated for five parameters: sequencing data alignment quality, genotyping quality, sequencing depth, genotype information missing ratio, and minimum allele frequency, using VCFTOOLS and BCFTOOLS. Specifically: (1) BCFTOOLS was used to filter the alignment quality of sequencing data. The highest requirement was to remove genetic variant sites with an RMS Mapping Quality Score of less than 30, and the lowest requirement was not to set an RMS Mapping Quality Score filtering requirement. (2) VCFTOOLS is used to filter the genotype quality. The highest requirement is to remove genotypes with a quality value of less than 20 and genetic variation sites with a quality value of less than 30. The lowest requirement is not to set any filtering requirements for genotype quality value and genetic variation site quality value. (3) Filtering for sequencing depth is performed using VCFTOOLS. The highest requirement is to retain only genetic variation sites whose sequencing depth is greater than or equal to the median sequencing depth of all genetic variation sites and less than or equal to 100. The lowest requirement is to retain only genetic variation sites whose sequencing depth is greater than or equal to the minimum sequencing depth of all genetic variation sites. (4) VCFTOOLS is used to filter the proportion of missing genotype information. Individuals with more than a certain proportion of genetic variation sites that lack genotype information need to be removed. Individuals with more than a certain proportion of genetic variation sites that lack genotype information need to be removed. (5) VCFTOOLS was used to filter the minimum allele frequency, and genetic variation sites with minimum allele frequencies less than 0.01, 0.02, 0.03, 0.04 and 0.05 were removed in different filtering strategies.
[0012] Furthermore, the filtering of the proportion of missing genotype information and the minimum allele frequency was carried out in two ways: (1) it was carried out on the entire set of the population to be evaluated, and the resulting genetic variation site dataset was used for population structure assessment; (2) it was carried out on each population to be evaluated, and the resulting genetic variation site dataset was used for genetic diversity and effective population size assessment. The screening criteria for sequencing data alignment quality, genotyping quality, sequencing depth, proportion of missing genotype information, and minimum allele frequency were cross-combined to construct a series of genetic variation site datasets representing different filtering criteria.
[0013] Furthermore, in step (5), for the assessment of population structure, ADMIXTURE ANALYSIS is used to estimate the ancestral components of the population to be assessed, principal component analysis is performed using the R software package HIERFSTAT, and genetic differentiation indices (Weir and Cockerham's) are calculated using VCFTOOLS. F ST ).
[0014] Furthermore, in step (5), for the assessment of genetic diversity, R software is used to calculate multilocus heterozygosity (MLH) and observed heterozygosity (H). obs ).
[0015] Furthermore, in step (5), for the assessment of effective population size, the GONE software is used to reconstruct the effective population size from hundreds of generations ago to the present, revealing the fluctuation pattern of effective population size and predicting the trend of change in effective population size.
[0016] The above methods were used to systematically evaluate the impact of different genetic variation site filtering criteria on different genetic analyses, determine the combination of genetic variation sites suitable for different genetic assessments, and provide reliable molecular markers for the assessment of shellfish germplasm resources.
[0017] Compared with the prior art, the present invention has the following advantages: This invention, by evaluating the performance of different genetic variation locus datasets in various genetic assessments, identifies key screening criteria that need to be considered when screening for genetic variation loci for different genetic assessments, providing a reference framework for developing genetic variation locus screening strategies. This invention constructs a series of genetic variation locus datasets covering different screening criteria by setting thresholds and cross-combining all screening criteria that affect the quality and quantity of genetic variation loci, which can provide a basis for parameter setting in the development of molecular markers for shellfish germplasm resource assessment. This invention, by evaluating the impact of different genetic variation locus screening criteria on different genetic analyses, determines genetic variation loci suitable for different genetic assessments, thereby developing molecular markers for shellfish germplasm resource assessment. Attached Figure Description
[0018] Figure 1This is a schematic diagram of ancestral component estimation and principal component analysis of *Scallop* populations under different genetic variation site filtering strategies in Example 1; where (a) is the population genetic structure analysis result based on dataset 1 with genetic variation sites having a minimum allele frequency greater than or equal to 0.05, (b) is the population genetic structure analysis result based on dataset 2 with genetic variation sites having a minimum allele frequency greater than or equal to 0.05, (c) is the population genetic structure analysis result based on dataset 3 with genetic variation sites having a minimum allele frequency greater than or equal to 0.05, and (d) is the population genetic structure analysis result based on dataset 3 with a minimum allele frequency greater than or equal to 0.05. (e) is the principal coordinate analysis result of the genetic structure of the population in dataset 4 with a minimum allele frequency of 0.05; (f) is the principal coordinate analysis result of dataset 1 with a minimum allele frequency of 0.05; (g) is the principal coordinate analysis result of dataset 2 with a minimum allele frequency of 0.05; (h) is the principal coordinate analysis result of dataset 3 with a minimum allele frequency of 0.05; and (h) is the principal coordinate analysis result of dataset 4 with a minimum allele frequency of 0.05.
[0019] Figure 2 The diagram shows the genetic differentiation index of the *Scallop* population under different genetic variation site filtering strategies in Example 1. (a) is the genetic differentiation index of dataset 1 / 2 / 3 / 4 based on genetic variation sites with a minimum allele frequency greater than or equal to 0.05; (b) is the genetic differentiation index of dataset 1 / 2 / 3 / 4 based on genetic variation sites with a minimum allele frequency greater than or equal to 0.05; (c) is the genetic differentiation index of dataset 1 / 2 / 3 / 4 based on genetic variation sites with a minimum allele frequency greater than or equal to 0.05; and (d) is the genetic differentiation index of dataset 1 / 2 / 3 / 4 based on genetic variation sites with a minimum allele frequency greater than or equal to 0.05.
[0020] Figure 3 This is a schematic diagram of the heterozygosity of the *Ctenopharynx* population under different genetic variation site filtering strategies in Example 2; where (a) is a graph showing the relationship between the average heterozygosity and the minimum allele frequency based on genetic variation site dataset 1, (b) is a graph showing the relationship between the average heterozygosity and the minimum allele frequency based on genetic variation site dataset 2, (c) is a graph showing the relationship between the average heterozygosity and the minimum allele frequency based on genetic variation site dataset 3, and (d) is a graph showing the relationship between the average heterozygosity and the minimum allele frequency based on genetic variation site dataset 4.
[0021] Figure 4This is a schematic diagram illustrating the observed heterozygosity of *Scallop* populations under different genetic variation site filtering strategies in Example 2; where (a) is the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 1 (which does not contain population-specific conserved sites), (b) is the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 2 (which does not contain population-specific conserved sites), (c) is the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 3 (which does not contain population-specific conserved sites), and (d) is the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 4 (which does not contain population-specific conserved sites). (e) is a graph showing the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 1 containing population-specific conserved loci; (f) is a graph showing the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 2 containing population-specific conserved loci; (g) is a graph showing the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 3 containing population-specific conserved loci; and (h) is a graph showing the relationship between the average observed heterozygosity and the minimum allele frequency based on dataset 4 containing population-specific conserved loci.
[0022] Figure 5 The diagrams illustrate the effective population size of *Scallop scallop* populations under different genetic variation site filtering strategies in Example 2. (a) shows the effective population size of population 1 based on a dataset of four genetic variation sites with a minimum allele frequency greater than or equal to 0.05 over time; (b) shows the effective population size of population 2 based on a dataset of four genetic variation sites with a minimum allele frequency greater than or equal to 0.05 over time; (c) shows the effective population size of population 3 based on a dataset of four genetic variation sites with a minimum allele frequency greater than or equal to 0.05 over time; (d) shows the effective population size of population 4 based on a dataset of four genetic variation sites with a minimum allele frequency greater than or equal to 0.05 over time; (e) shows the effective population size of population 5 based on a dataset of four genetic variation sites with a minimum allele frequency greater than or equal to 0.05 over time; and (f) shows the effective population size of population 6 based on a dataset of four genetic variation sites with a minimum allele frequency greater than or equal to 0.05 over time. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Example 1: A method for developing molecular markers for evaluating mollusk germplasm resources, specifically including the following steps: (1) For the six natural populations of scallops to be evaluated, 30 representative individuals were collected from each population and whole-genome sequencing with an average sequencing depth of 10× was carried out. Genome analysis toolkit was used to identify the genetic variation sites of the whole genome.
[0025] (2) Based on the actual sequencing situation and genetic evaluation requirements, a series of whole-genome genetic variation site filtering strategies representing different screening criteria were formulated, and a total of 20 genetic variation site datasets in 4 categories were obtained for population structure evaluation, as shown in Table 1. Among them, the filtering for sequencing data alignment quality was performed using BCFTOOLS, with the highest requirement being to remove genetic variation sites with an RMS Mapping Quality Score of less than 30, and the lowest requirement being not to set an RMS Mapping Quality Score filtering requirement; the filtering for genotype typing quality was performed using VCFTOOLS, with the highest requirement being to remove genotypes with a quality value of less than 20 and genetic variation sites with a quality value of less than 30, and the lowest requirement being not to set a genotype quality value and genetic variation site quality value filtering requirement; the filtering for sequencing depth was performed using VCFTOOLS, with the highest requirement being to retain only genetic variation sites with a sequencing depth greater than or equal to 10 and less than or equal to 100, and the lowest requirement being to retain only genetic variation sites with a sequencing depth greater than or equal to 3; the filtering for the proportion of missing genotype information was performed using VCFTOOLS. OOLS was performed on the aggregate of all six populations, mainly in two ways: the first method retained all individuals but only those with genotypic information in all individuals; the second method retained only those with genotypic information in more than 90% of the individuals but only those with genotypic information in more than 90% of the individuals. VCFTOOLS was used for filtering based on minimum allele frequency, removing genetic variants with minimum allele frequencies less than 0.01, 0.02, 0.03, 0.04, and 0.05 in different filtering strategies.
[0026] Table 1. Filtering strategies for genetic variation sites
[0027] (3) Population structure assessment was conducted using a dataset of 20 genetic variation sites representing different screening criteria, such as... Figure 1 , Figure 2As shown in the figure, the results when the minimum allele frequency screening criteria are 0.01, 0.02, 0.03, and 0.04 are consistent with the results when the minimum allele frequency screening criteria are 0.05. Only the results when the minimum allele frequency screening criteria are 0.05 are shown.
[0028] (4) By evaluating the population structure analysis results based on datasets of different genetic variation sites, it was found that, for example Figure 1 , Figure 2 As shown, the population structure under different genetic variation site screening strategies is basically consistent. ADMIXTURE ANALYSIS shows that when the assumed population size is set to 1 or 2, the CV error is minimized, indicating that the six populations are highly likely to be divided into 1 or 2 groups. Principal component analysis shows that the six populations cluster into two groups (populations 1, 2, and 3 form one cluster, and populations 4, 5, and 6 form another cluster), indicating that the six populations are highly likely to be divided into 2 groups. The degree of genetic differentiation between pairs of populations 1, 2, and 3 and between pairs of populations 4, 5, and 6 is extremely low; in contrast, the degree of genetic differentiation between any one of populations 1, 2, and 3 and any one of populations 4, 5, and 6 is relatively high, further confirming that the six populations are divided into 2 groups.
[0029] Based on the principles of reducing sample requirements to lower data collection costs and increasing the number of genetic variation sites to increase the amount of genetic information while ensuring the reliability of the assessment, filtering the genetic variation site dataset 4 with a minimum allele frequency of 0.01 can obtain reliable, sufficient molecular markers suitable for population structure assessment.
[0030] Example 2: A method for developing molecular markers for evaluating mollusk germplasm resources, specifically including the following steps: (1) For the six natural populations of scallops to be evaluated, 30 representative individuals were collected from each population and whole-genome sequencing with an average sequencing depth of 10× was carried out. Genome analysis toolkit was used to identify the genetic variation sites of the whole genome.
[0031] (2) Based on the actual sequencing situation and genetic evaluation requirements, a series of whole-genome genetic variation site filtering strategies representing different screening criteria were formulated, and a total of 120 genetic variation site datasets in 4 categories were obtained for genetic diversity and effective population size evaluation, as shown in Table 2. Among them, the filtering for sequencing data alignment quality was performed using BCFTOOLS, with the highest requirement being to remove genetic variation sites with an RMS Mapping Quality Score less than 30, and the lowest requirement being to not set an RMS Mapping Quality Score filtering requirement; the filtering for genotype typing quality was performed using VCFTOOLS, with the highest requirement being to remove genotypes with a quality value less than 20 and genetic variation sites with a quality value less than 30, and the lowest requirement being to not set a genotype quality value and genetic variation site quality value filtering requirement; the filtering for sequencing depth was performed using VCFTOOLS, with the highest requirement being to retain only genetic variation sites with a sequencing depth greater than or equal to 10 and less than or equal to 100, and the lowest requirement being to retain only genetic variation sites with a sequencing depth greater than or equal to 3; the filtering for the proportion of missing genotype information was performed using VCFTOOLS respectively. For each population, the genetic variation sites were analyzed using two main methods: the first method retained only those genetic variation sites for which all individuals possessed genotype information, while preserving all individuals; the second method retained only those genetic variation sites for which more than 90% of individuals possessed genotype information, while preserving only those genetic variation sites for which more than 90% of individuals possessed genotype information. Filtering for minimum allele frequency was performed using VCFTOOLS, and genetic variation sites with minimum allele frequencies less than 0.01, 0.02, 0.03, 0.04, and 0.05 were removed using different filtering strategies.
[0032] Table 2. Filtering strategies for genetic variation sites
[0033] (3) Genetic diversity was assessed using a dataset of 120 genetic variation sites representing different screening criteria from 6 populations, such as... Figure 3 , Figure 4As shown. With the removal of population-specific conserved loci, the mean multilocus heterozygosity of genetic variation loci datasets 1, 2, and 3 is higher in populations 1, 2, and 3, and the mean multilocus heterozygosity of genetic variation loci dataset 4 is higher in populations 4, 5, and 6. With the removal of population-specific conserved loci, the mean observed heterozygosity of genetic variation loci datasets 1, 2, and 3 is higher in populations 1, 2, and 3, and the mean observed heterozygosity of genetic variation loci dataset 4 is higher in populations 4, 5, and 6; with the retention of population-specific conserved loci, the mean observed heterozygosity of genetic variation loci datasets 1, 2, 3, and 4 is higher in populations 4, 5, and 6. Different genetic variation loci screening criteria can change the relationship between different populations in terms of genetic diversity. This phenomenon is caused by whether or not population-specific conserved loci are included in the analysis. Therefore, two sets of molecular markers should be developed for population genetic diversity assessment: (1) retaining population-specific conserved loci and (2) removing population-specific conserved loci. The two sets of molecular markers should be used for assessment and comparison, and the corresponding conclusions should be reported.
[0034] (4) The effective population size was evaluated using a dataset of 120 genetic variation sites representing different screening criteria from 6 populations, such as... Figure 5 As shown in Table 3, for population 1, the effective population size change trends based on genetic variation site datasets 1 and 2 are consistent, and the effective population size change trends based on genetic variation site datasets 3 and 4 are also consistent. For population 2, the effective population size change trends based on genetic variation site datasets 1, 2, 3, and 4 are all different. For population 3, the effective population size change trends based on genetic variation site datasets 1 and 2 are consistent, but the effective population size change trends based on other genetic variation site datasets are all different. For populations 4, 5, and 6, the effective population size change trends based on genetic variation site datasets are basically consistent. The assessment based on genetic variation site datasets 1 and 2 cannot capture some of the effective population size change trends in populations 1, 2, and 3. The effective population size change trend based on genetic variation site dataset 4 is basically consistent across the six populations. The number of genetic variation sites is a key factor affecting the accuracy of effective population assessment; a decrease in the number of genetic variation sites will lead to the inability to capture certain effective population size change trends, as shown in Table 3. Figure 5 As shown. When a population contains a large number of genetic variation sites (e.g., populations 4, 5, and 6), genetic variation site datasets 2, 3, and 4 can all be used as molecular markers for effective population size assessment; when a population contains a small number of genetic variation sites (e.g., populations 1, 2, and 3), only genetic variation site dataset 4 can be used as a molecular marker for effective population size assessment (e.g., populations 4, 5, and 6). Figure 5 (As shown).
[0035] The results shown when the minimum allele frequency screening criteria are 0.01, 0.02, 0.03, and 0.04 are consistent with the results shown when the minimum allele frequency screening criteria are 0.05. Only the results shown when the minimum allele frequency screening criteria are 0.05 are displayed.
[0036] Table 3. Number of samples and number of genetic variation sites in datasets with different genetic variation sites. .
[0037] (5) By evaluating the analysis results of genetic diversity and effective population size for different populations based on different genetic variation site datasets, it was found that the screening criteria for genetic variation sites that affect the assessment of genetic diversity and effective population size are different. Based on this, molecular markers suitable for the assessment of genetic diversity and effective population size were developed respectively.
[0038] This invention aims to address the bottleneck of lacking molecular markers suitable for systematic and comprehensive assessment of shellfish germplasm resources at the whole-genome level. By evaluating the impact of different whole-genome genetic variation site screening strategies on different genetic assessments, this invention develops reliable molecular markers for shellfish germplasm resource assessment to meet the requirements of different genetic assessments, providing a scientific basis for the protection and utilization of shellfish germplasm resources.
[0039] Finally, although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for developing molecular markers for evaluating shellfish germplasm resources, characterized in that, The method specifically includes the following steps: (1) Collect representative populations of the shellfish population to be evaluated; (2) Conduct high-throughput sequencing at the whole genome level; (3) Obtain genome-wide genetic variation sites using high-throughput sequencing data analysis tools; (4) Based on the actual sequencing situation and genetic evaluation requirements, formulate a series of whole-genome genetic variation site filtering strategies representing different screening criteria, and obtain the corresponding genetic variation site dataset based on the strategy; (5) Use datasets of different genetic variation sites to conduct population structure assessment, genetic diversity assessment and effective population size assessment; (6) By evaluating the performance of different genetic variation sites in different genetic assessments, multiple datasets were selected, which are molecular markers suitable for the assessment of shellfish germplasm resources.
2. The method for developing molecular markers for evaluating shellfish germplasm resources as described in claim 1, characterized in that, The whole-genome genetic variation sites are obtained as follows: multiple representative individuals are randomly collected from each population to be evaluated, and whole-genome high-throughput sequencing is performed on the representative individuals of the population to be evaluated. The whole-genome genetic variation sites are obtained using high-throughput sequencing data analysis tools.
3. The method for developing molecular markers for evaluating shellfish germplasm resources as described in claim 1, characterized in that, In step (4), the genetic variation site dataset representing different screening criteria is obtained as follows: the whole genome genetic variation site filtering strategy is formulated for 5 parameters: sequencing data alignment quality, genotyping quality, sequencing depth, genotype information missing ratio, and minimum allele frequency, using VCFTOOLS and BCFTOOLS.
4. The method for developing molecular markers for evaluating shellfish germplasm resources as described in claim 3, characterized in that, The specific steps are as follows: (1) BCFTOOLS was used to filter the alignment quality of sequencing data. The highest requirement was to remove genetic variant sites with an RMS Mapping Quality Score of less than 30, and the lowest requirement was not to set an RMS Mapping Quality Score filter requirement. (2) VCFTOOLS is used to filter the genotype quality. The highest requirement is to remove genotypes with a quality value of less than 20 and genetic variation sites with a quality value of less than 30. The lowest requirement is not to set any filtering requirements for genotype quality value and genetic variation site quality value. (3) Filtering for sequencing depth is performed using VCFTOOLS. The highest requirement is to retain only genetic variation sites whose sequencing depth is greater than or equal to the median sequencing depth of all genetic variation sites and less than or equal to 100. The lowest requirement is to retain only genetic variation sites whose sequencing depth is greater than or equal to the minimum sequencing depth of all genetic variation sites. (4) VCFTOOLS is used to filter the proportion of missing genotype information. Individuals with more than a certain proportion of genetic variation sites that lack genotype information need to be removed. (5) VCFTOOLS was used to filter the minimum allele frequency, and genetic variation sites with minimum allele frequencies less than 0.01, 0.02, 0.03, 0.04 and 0.05 were removed in different filtering strategies.
5. The method for developing molecular markers for evaluating shellfish germplasm resources as described in claim 3, characterized in that, The filtering of the proportion of missing genotype information and the minimum allele frequency was carried out in two ways: (1) it was carried out on the entire set of the population to be evaluated, and the resulting genetic variation site dataset was used for population structure assessment; (2) it was carried out on each population to be evaluated, and the resulting genetic variation site dataset was used for genetic diversity and effective population size assessment. The screening criteria of sequencing data alignment quality, genotyping quality, sequencing depth, proportion of missing genotype information, and minimum allele frequency were cross-combined to construct a series of genetic variation site datasets representing different filtering criteria.
6. The method for developing molecular markers for evaluating shellfish germplasm resources as described in claim 1, characterized in that, In step (5), for population structure assessment, ADMIXTURE ANALYSIS is used to estimate the ancestral components of the population to be assessed, the R software package HIERFSTAT is used for principal component analysis, and VCFTOOLS is used to calculate the genetic differentiation index between different populations.
7. The method for developing molecular markers for evaluating shellfish germplasm resources as described in claim 1, characterized in that, In step (5), for the assessment of genetic diversity, R software is used to calculate the multilocus heterozygosity (MLH) and the observed heterozygosity (H). obs .
8. The method for developing molecular markers for evaluating shellfish germplasm resources as described in claim 1, characterized in that, In step (5), for the assessment of effective population size, the GONE software is used to reconstruct the effective population size from hundreds of generations ago to the present, revealing the fluctuation pattern of effective population size and predicting the trend of change in effective population size.
Citation Information
Patent Citations
Method for evaluating breed conservation effect of preserved population based on low-depth whole genome re-sequencing technology
CN117542418A
Eggplant core SNP molecular marker set and application
CN120485410A