Core germplasm screening method based on genetic distance control

Through the core germplasm screening method based on genetic distance control, the problem of inefficiency in screening of germplasm resources in the prior art is solved, effective screening and diversity protection of germplasm resources are achieved, and the efficiency and resource utilization of forestry breeding process are promoted.

CN119979752APending Publication Date: 2025-05-13HEBEI ACAD OF FORESTRY SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062064.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively screen germplasm resources, resulting in resource redundancy, increasing resource utilization difficulty, and failing to effectively protect the gene bank integrity of species gene banks and ecosystems.

Method used

The core germplasm screening method based on genetic distance control was adopted. DNA samples of plant primitive germplasm populations were collected, PCR amplification was performed using SSR molecular markers, Nei’s genetic distance between individuals was calculated, the genetic distance matrix was constructed, and the screening range was formulated based on the genetic distance distribution, and multiple rounds of screening were performed to select the most suitable core germplasm population.

Benefits of technology

Effectively screen germplasm resources, ensure their diversity, protect the gene bank integrity of species gene banks and ecosystems, promote the effective utilization of germplasm resources, and improve the efficiency of forestry breeding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119979752A_ABST
    Figure CN119979752A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of germplasm diversity screening, and particularly provides a core germplasm screening method based on genetic distance control, which comprises the following steps of: collecting an original germplasm group tissue sample to extract DNA (Deoxyribonucleic Acid), and performing PCR (Polymerase Chain Reaction) amplification by utilizing an SSR (Simple Sequence Repeat) molecular marker to obtain site length polymorphism data; the Neii's genetic distance between individual sites is calculated to construct a genetic distance matrix, and multi-round screening is carried out according to specific conditions. Through multiple rounds of screening, comparing the genetic diversity preserving rates to determine a final alternative group, and verifying the effectiveness of the group through t test or a principal coordinate analysis method. According to the method, germplasm resources can be effectively screened, the diversity of the germplasm resources is guaranteed, and powerful support is provided for biodiversity protection and forestry production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of germplasm diversity screening, and in particular relates to a core germplasm screening method based on genetic distance control. Background Art

[0002] Conducting core germplasm screening on genetic resources with huge quantities of germplasm is of great significance. For biodiversity conservation, it helps to protect the gene pool of species and maintain the integrity of the gene pool of the entire ecosystem; it also helps to explore genetic laws and discover new genes and traits.

[0003] From the perspective of forestry production, excessively redundant germplasm resources not only fail to play a role in scientifically and rationally protecting resources, but also increase the difficulty of resource utilization. Screening germplasm resources can promote the effective utilization of germplasm resources, accelerate the breeding process, and provide scientific basis and material foundation for related genetic research. Therefore, how to effectively carry out core germplasm screening has become an urgent problem to be solved. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a core germplasm screening method based on genetic distance control, which can effectively screen germplasm resources and ensure their diversity.

[0005] The specific technical solution adopted by the present invention is:

[0006] A core germplasm screening method based on genetic distance control comprises the following steps:

[0007] S1. Collect tissue samples of the original germplasm population of the plant to be screened and extract DNA, perform PCR amplification using SSR molecular markers, and obtain the site length polymorphism data of all individuals in the original germplasm population;

[0008] S2. Based on the individual locus length polymorphism data, calculate the Nei's genetic distance between individual loci and construct the genetic distance matrix of the original germplasm population;

[0009] S3. According to the distribution of Nei's genetic distance between individual loci, a screening range is formulated, and a specific genetic distance value is used as the screening condition, and a specific, smaller genetic distance value is used as the interval, and multiple rounds of screening are performed in the order from small to large;

[0010] S4. Compare the genetic diversity preservation rates of candidate populations obtained under all screening conditions within the screening range. When there are multiple candidate populations obtained under different screening conditions with the same genetic diversity preservation rate, take the candidate population with the largest screening condition as the final candidate population;

[0011] S5. Compare the genetic diversity parameters of all final candidate populations with those of the original population, select the most suitable core germplasm population through t-test or principal coordinate analysis, and complete the verification of population validity at the same time.

[0012] Taking genetic distance as the primary screening condition, the screening condition in step S3 is to set the upper and lower thresholds [a, b] of the genetic distance value and the change gradient t. The genetic distance value starts from a and increases t each time until the genetic distance value reaches b, and multiple rounds of screening are performed.

[0013] Step S3 is processed before screening:

[0014] S301, according to the screening condition, find out all the values ​​greater than or equal to the screening condition from the genetic distance matrix row by row, column by column, to form a data set that meets the screening condition, and find out all the values ​​less than the screening condition row by column, column by row, to form a data set that does not meet the screening condition; wherein the bidirectional matrix does not consider the data above the diagonal line;

[0015] S302, using the sample number or column number of the column where the data is located as the first coordinate and the sample number or row number of the row where the data is located as the second coordinate, converting the data set that meets the screening condition and the data set that does not meet the screening condition into a coordinate set, i.e., a coordinate set that meets the screening condition and a coordinate set that does not meet the screening condition;

[0016] S303, starting from the first group of coordinates in the coordinate set that meets the screening conditions, screening and filling are performed in sequence; the first group of coordinates used for each screening is defined as the starting coordinate group, and then screening is performed according to the steps.

[0017] The screening steps in step S3 are:

[0018] S311, from the coordinate set that meets the screening condition, find all coordinate groups that use the first coordinate and the second coordinate of the starting coordinate group as the first coordinate, find all repeated coordinates from the second coordinates of these coordinate groups, and define these coordinates as candidate third coordinates;

[0019] S312, sequentially extracting 2 coordinates from n candidate third coordinates, n≥2, and combining them to obtain n*(n-1) / 2 coordinate groups; finding out from these coordinate groups the coordinate groups that belong to the coordinate set that does not meet the screening condition, extracting two coordinates from these coordinate groups and classifying them into the set to be deleted;

[0020] S313, find out the coordinates whose repetition times are greater than 2 from the set to be deleted, find out the coordinates whose repetition times are the largest and ranked at the top, delete the coordinates from the candidate third coordinates, and obtain a new set of candidate third coordinates;

[0021] S314, repeating steps S312 and S313 until no coordinate group that does not meet the filter condition appears, the candidate third coordinate obtained in the last step S313 is regarded as the selected third coordinate, and directly proceeding to step S316, or the number of occurrences of all coordinates in the set to be deleted is 1, and proceeding to step S315;

[0022] S315, from the a coordinate groups in the coordinate set that does not meet the screening condition obtained in the last step S312, a≥1, randomly select one coordinate in turn to form a deleted coordinate set. This will result in 2 a Different sets of deleted coordinates are deleted from the candidate third coordinates obtained in the last step 313, and the coordinates in these sets are deleted respectively, so as to obtain 2 a Different groups of selected third coordinates;

[0023] S316, add all the selected third coordinates to the starting coordinate group in sequence to obtain i final coordinate groups. If you go directly to step S316 from steps 1, 2, and 4, then i=1; if you go to step 6 from step S315, then i=2 a .

[0024] S317, if this screening is the first screening in this round, the final coordinate group is directly added to the solution set; if it is not the first screening, the final coordinate group is compared with the coordinate group in the solution set obtained by the previous screening, if the final coordinate group is a subset of a coordinate group in the solution set, the final coordinate group is deleted, otherwise the final coordinate group is added to the solution set;

[0025] S318. Convert the coordinates in all coordinate groups in the solution set into sample numbers, convert the coordinate groups into sample sets, and calculate the polymorphic site preservation rate, sample number and average genetic distance of all sample sets compared with the original population based on the polymorphism data matrix and the genetic distance matrix. Find the qualified sample population as the final candidate population for this round of screening in the order of maximum preservation rate, minimum sample number and maximum average genetic distance.

[0026] In the step S311, when there are less than two candidate third coordinates, all the candidate third coordinates are regarded as selected third coordinates, and the process directly goes to step S316.

[0027] In the step S312, if no coordinate group belonging to the coordinate set that does not meet the screening condition is found from these coordinate groups, all candidate third coordinates are regarded as selected third coordinates, and the process directly goes to step S316.

[0028] In step S313, if there are no coordinates in the set to be deleted whose number of repeated occurrences is greater than 2, the process directly proceeds to step S315.

[0029] The beneficial effects of the present invention are:

[0030] The core germplasm screening method based on genetic distance control provided by the present invention can effectively screen germplasm resources to ensure their diversity, thereby helping to protect the species gene pool and maintain the integrity of the gene pool of the entire ecosystem.

[0031] Using this screening method to screen germplasm resources in forestry breeding work can promote the effective utilization of germplasm resources and allow high-quality germplasm resources to better serve the forestry production link.

[0032] The present invention collects tissue samples of the original germplasm population of the plant to be screened and extracts DNA, and uses SSR molecular markers for PCR amplification to obtain the site length polymorphism data of all individuals in the original germplasm population, laying a data foundation for subsequent screening and analysis.

[0033] The present invention uses genetic distance as a control condition, adopts a method of formulating a screening range according to the distribution of genetic distances between individual loci, and performs multiple rounds of screening in ascending order according to the screening conditions. The final candidate population for each round of screening is determined in the priority order of maximum preservation rate, minimum number of samples, and maximum average genetic distance. This can effectively screen germplasm resources and retain the genetic diversity of germplasm to the greatest extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 The distribution of Nei's genetic distance (GD) among individuals of Ulmus germplasm resources; DETAILED DESCRIPTION

[0035] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments:

[0036] Specific embodiment 1, the present invention is a core germplasm screening method based on genetic distance control, comprising the following steps:

[0037] S1. Collect tissue samples of the original germplasm population of the plant to be screened and extract DNA, perform PCR amplification using SSR molecular markers, and obtain the site length polymorphism data of all individuals in the original germplasm population;

[0038] S2. Based on the individual locus length polymorphism data, calculate the Nei's genetic distance between individual loci and construct the genetic distance matrix of the original germplasm population;

[0039] S3. According to the distribution of Nei's genetic distance between individual loci, a screening range is formulated, and a specific genetic distance value is used as the screening condition, and a specific, smaller genetic distance value is used as the interval, and multiple rounds of screening are performed in the order from small to large;

[0040] S4. Compare the genetic diversity preservation rates of candidate populations obtained under all screening conditions within the screening range. When there are multiple candidate populations obtained under different screening conditions with the same genetic diversity preservation rate, take the candidate population with the largest screening condition as the final candidate population;

[0041] S5. Compare the genetic diversity parameters of all final candidate populations with those of the original population, select the most suitable core germplasm population through t-test or principal coordinate analysis, and complete the verification of population validity at the same time.

[0042] The present invention uses 184 germplasm resources of Ulmus genus as test materials and screens the core germplasms.

[0043] According to step S1 and step S2, the distribution of Nei's genetic distances between individual sites of 184 germplasm resources is obtained, and a genetic distance matrix as shown in Table 1 is established;

[0044] Table 1: Genetic distance matrix (pattern)

[0045]

[0046] Note: The horizontal and vertical axes in the table are the sample numbers of the 184 germplasm resources.

[0047] The screening conditions in step S3 are to set the upper and lower thresholds [a, b] of the genetic distance value and the change gradient t. The genetic distance value starts from a and increases t each time until the genetic distance value reaches b. Multiple rounds of screening are performed, and multiple candidate populations are determined according to step S4.

[0048] Specifically, the Nei's genetic distance (GD) between the 184 Ulmus germplasm resources ranged from 0.020 to 0.920. The GD data were divided into 9 groups with an interval of 0.1. The GD data of each group were counted and it was found that the GD data of the 5 groups between 0.22 and 0.72 were all greater than 2000 (see Appendix Figure 1 ). The upper and lower thresholds were set to [0.22, 0.72], and the gradient was 0.001. After multiple rounds of screening, it was found that when the screening condition was greater than 0.520, the number of individuals in the candidate population was less than 5% of the total number of individuals in the original population, and the screening was stopped. The genetic diversity preservation rates of the candidate populations obtained under all screening conditions were compared. When the genetic diversity preservation rates of the candidate populations obtained under multiple different screening conditions were the same, the candidate population with the largest screening condition was taken as the final candidate population, thereby determining 8 candidate core germplasm populations (Table 2).

[0049] Table 2 Information on candidate core germplasm populations

[0050]

[0051]

[0052] According to step S5, the most suitable core germplasm is screened out. Specifically, the N of the eight candidate populations is compared with that of the original population. a 、N e , H o , H e ,uH e , I and F (Table 3). According to the results of the t-test, Core group 4 retained 92.4% of the N a , is the candidate population with the minimum number of individuals that are not significantly different from the original population, and its N e ,I,H o , H e and uH e They were respectively increased to 135.4%, 118.8%, 109.1%, 113.6% and 115.3% of the original population, meeting the population validity verification and determining it as the core germplasm.

[0053] Table 3 Genetic diversity parameters of different populations ①

[0054]

[0055] ①* indicates significant difference at P≤0.05 level, ** indicates significant difference at P≤0.01 level.

Claims

1. A core germplasm screening method based on genetic distance control, characterized in that: The steps include: S1. Collect tissue samples of the original germplasm population of the plant to be screened and extract DNA, perform PCR amplification using SSR molecular markers, and obtain the site length polymorphism data of all individuals in the original germplasm population; S2. Based on the individual locus length polymorphism data, calculate the Nei's genetic distance between individual loci and construct the genetic distance matrix of the original germplasm population; S3. According to the distribution of Nei's genetic distance between individual loci, a screening range is formulated, and a specific genetic distance value is used as the screening condition, and a specific, smaller genetic distance value is used as the interval, and multiple rounds of screening are performed in the order from small to large; S4. Compare the genetic diversity preservation rates of candidate populations obtained under all screening conditions within the screening range. When there are multiple candidate populations obtained under different screening conditions with the same genetic diversity preservation rate, take the candidate population with the largest screening condition as the final candidate population; S5. Compare the genetic diversity parameters of all final candidate populations with those of the original population, select the most suitable core germplasm population through t-test or principal coordinate analysis, and complete the verification of population validity at the same time.

2. The core germplasm screening method based on genetic distance control according to claim 1, characterized in that: Taking genetic distance as the primary screening condition, the screening condition in step S3 is to set the upper and lower thresholds [a, b] of the genetic distance value and the change gradient t. The genetic distance value starts from a and increases t each time until the genetic distance value reaches b, and multiple rounds of screening are performed.

3. The core germplasm screening method based on genetic distance control according to claim 1, characterized in that: Step S3 is processed before screening: S301, according to the screening condition, find out all the values ​​greater than or equal to the screening condition from the genetic distance matrix row by row, column by column, to form a data set that meets the screening condition, and find out all the values ​​less than the screening condition row by column, column by row, to form a data set that does not meet the screening condition; wherein the bidirectional matrix does not consider the data above the diagonal line; S302, using the sample number or column number of the column where the data is located as the first coordinate and the sample number or row number of the row where the data is located as the second coordinate, converting the data set that meets the screening condition and the data set that does not meet the screening condition into a coordinate set, i.e., a coordinate set that meets the screening condition and a coordinate set that does not meet the screening condition; S303, starting from the first group of coordinates in the coordinate set that meets the screening conditions, screening and filling are performed in sequence; the first group of coordinates used for each screening is defined as the starting coordinate group, and then screening is performed according to the steps.

4. The core germplasm screening method based on genetic distance control according to claim 3, characterized in that: The screening steps in step S3 are: S311, from the coordinate set that meets the screening condition, find all coordinate groups that use the first coordinate and the second coordinate of the starting coordinate group as the first coordinate, find all repeated coordinates from the second coordinates of these coordinate groups, and define these coordinates as candidate third coordinates; S312, sequentially extracting 2 coordinates from n candidate third coordinates, n≥2, and combining them to obtain n*(n-1) / 2 coordinate groups; finding out from these coordinate groups the coordinate groups that belong to the coordinate set that does not meet the screening condition, extracting two coordinates from these coordinate groups and classifying them into the set to be deleted; S313, find out the coordinates whose repetition times are greater than 2 from the set to be deleted, find out the coordinates whose repetition times are the largest and ranked at the top, delete the coordinates from the candidate third coordinates, and obtain a new set of candidate third coordinates; S314, repeating steps S312 and S313 until no coordinate group that does not meet the filter condition appears, the candidate third coordinate obtained in the last step S313 is regarded as the selected third coordinate, and directly proceeding to step S316, or the number of occurrences of all coordinates in the set to be deleted is 1, and proceeding to step S315; S315, from the a coordinate groups in the coordinate set that does not meet the screening condition obtained in the last step S312, a≥1, randomly select one coordinate in turn to form a deleted coordinate set, and this will result in 2 a Different sets of deleted coordinates are deleted from the candidate third coordinates obtained in the last step 313, and the coordinates in these sets are deleted respectively, so as to obtain 2 a Different groups of selected third coordinates; S316, add all the selected third coordinates to the starting coordinate group in sequence to obtain i final coordinate groups. If you go directly to step S316 from steps 1, 2, and 4, then i=1; if you go to step 6 from step S315, then i=2 a ; S317, if this screening is the first screening in this round, the final coordinate group is directly added to the solution set; if it is not the first screening, the final coordinate group is compared with the coordinate group in the solution set obtained by the previous screening, if the final coordinate group is a subset of a coordinate group in the solution set, the final coordinate group is deleted, otherwise the final coordinate group is added to the solution set; S318. Convert the coordinates in all coordinate groups in the solution set into sample numbers, convert the coordinate groups into sample sets, and calculate the polymorphic site preservation rate, sample number and average genetic distance of all sample sets compared with the original population based on the polymorphism data matrix and the genetic distance matrix. Find the qualified sample population as the final candidate population for this round of screening in the order of maximum preservation rate, minimum sample number and maximum average genetic distance.

5. The core germplasm screening method based on genetic distance control according to claim 4, characterized in that: In the step S311, when there are less than two candidate third coordinates, all the candidate third coordinates are regarded as selected third coordinates, and the process directly goes to step S316.

6. The core germplasm screening method based on genetic distance control according to claim 4, characterized in that: In the step S312, if no coordinate group belonging to the coordinate set that does not meet the screening condition is found from these coordinate groups, all candidate third coordinates are regarded as selected third coordinates, and the process directly goes to step S316.

7. The core germplasm screening method based on genetic distance control according to claim 4, characterized in that: In step S313, if there are no coordinates in the set to be deleted that have a repetition frequency greater than 2, the process directly proceeds to step S315.