Peanut quality character selective breeding method based on whole genome SNP (Single Nucleotide Polymorphism) and application
By using whole-genome resequencing and a hybrid deep learning model, high-quality SNP loci were screened, and a hybrid deep learning model was constructed. This solved the problems of long cycle and low prediction accuracy in traditional peanut breeding methods, achieving efficient breeding and cultivating high-quality new peanut varieties.
Patent Information
- Application Number
- CN202511352921.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Traditional flowering breeding methods suffer from long cycles, low prediction accuracy, and significant susceptibility to environmental errors. Furthermore, existing whole-genome selection technologies have low coefficients of determination when predicting across populations, making it difficult to meet the needs of commercial breeding.
By integrating whole-genome resequencing, SNP optimized marker screening, and hybrid deep learning model construction, we can screen high-quality SNP loci, build a hybrid deep learning model, optimize hyperparameters, shorten the breeding cycle, and improve prediction accuracy.
It has enabled accurate prediction of peanut quality traits and evaluation of breeding values, shortened the breeding cycle, improved prediction accuracy, increased breeding efficiency, and cultivated new high-quality peanut varieties with high oleic acid and high protein.
Smart Images

Figure CN121171335A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of crop breeding, in particular to a peanut quality trait selection breeding method based on whole genome SNPs and application. BACKGROUND
[0002] As an important oil crop and economic crop in the world, the quality traits (such as oil content, protein content, fatty acid composition, etc.) of peanuts directly affect the economic benefits of the industry and the health needs of consumers. With the development of molecular breeding technology, whole genome selection (GS) technology has become a key means to break through the bottleneck of traditional phenotype selection because it can simultaneously analyze the synergistic regulation mechanism of multiple genes. However, the peanut genome is complex, and the quality traits are influenced by the interaction of multiple genes and the environment, which leads to problems such as long cycle and low prediction accuracy in traditional breeding methods. Therefore, it is necessary to develop a whole genome selection breeding technology that integrates high-precision genetic marker screening and model optimization strategies to achieve precise improvement of peanut quality traits.
[0003] However, traditional peanut breeding techniques mainly rely on direct phenotype selection or single-marker assisted selection, which has some limitations. First, phenotype selection is significantly affected by environmental errors. For example, the determination of oleic acid content requires gas chromatography, which is costly and has poor repeatability, resulting in significant bias in breeding value assessment. Second, single-marker assisted selection can only utilize a small number of known QTL sites, ignoring the cumulative effects of microgenes within the whole genome. For example, the quality trait of oleic acid / linoleic acid ratio is regulated by at least 12 gene loci, and traditional methods are limited in explaining the variation of the phenotype. Third, the breeding cycle is long, from hybridization to stable line breeding, which requires multiple generations of selfing. The separation of offspring traits leads to low selection efficiency. Although existing whole genome selection techniques attempt to introduce deep learning models, they do not integrate population structure correction and correlation contribution weighting mechanisms, resulting in generally lower determination coefficients when predicting across populations, which cannot meet the needs of commercial breeding.
[0004] Therefore, a peanut quality trait selection breeding method based on whole genome SNPs and application is developed. SUMMARY
[0005] The purpose of the present application is to make up for the shortcomings of the prior art, and provide a peanut quality trait selection breeding method based on whole genome SNP and application, and the present application realizes the prediction and breeding value evaluation of peanut quality traits by integrating whole genome resequencing, SNP optimized marker screening and mixed deep learning model construction, obtains an SNP variation map through whole genome resequencing and quality control process, covers ten key quality traits such as protein, oil content and oleic acid, and provides comprehensive genetic information for subsequent analysis, avoids the interference of low-quality sites by screening out sites to form a characteristic marker set in the SNP optimization and marker screening link, adjusts the training mixed deep learning model, optimizes the hyperparameters through ten-fold cross-validation, obtains a whole genome selection model, shortens the breeding cycle, reduces the prediction error rate, and provides technical support for the improvement of peanut quality traits.
[0006] To solve the above technical problems, the present application provides the following technical solutions: on the one hand, a peanut quality trait selection breeding method based on whole genome SNP, the specific steps of the method are:
[0007] Data acquisition: collect multiple peanut germplasms for planting, take 3 healthy plants at the mature stage of the germplasm, determine ten quality traits and take the average value to construct a phenotype matrix, extract the tender leaf DNA of the plant, and perform whole genome resequencing, compare the sequencing data with the peanut T2T reference genome after processing, screen SNP sites, and obtain an SNP variation map;
[0008] SNP optimization and marker screening: based on the SNP variation map, calculate the SNP site deletion rate and the minimum allele frequency, screen high-quality SNP sites and fill in the missing data, analyze the population structure, calculate the linkage disequilibrium, and carry out whole genome association analysis to screen key SNP sites to form a characteristic marker set;
[0009] Model construction: taking the characteristic marker set as input and the corresponding germplasm quality trait phenotype as output, a mixed deep learning model is constructed, the correlation contribution degree related index is introduced in the training process to adjust the model, the hyperparameters are optimized and selected through cross-validation, and a whole genome selection model is obtained;
[0010] New strain breeding: select the backbone parent according to the combination of the phenotype matrix, the SQI and the AC index and the genetic distance data of the germplasm, construct a breeding population, obtain the genotype of the offspring individuals, and obtain the original breeding value through the whole genome selection model, calculate the corrected breeding value of the offspring individuals and screen excellent single plants, and obtain high-quality peanut new strains through multi-generation aggregation selection.
[0011] Further, in the data acquisition, ten quality traits are determined, including protein, oil content, oleic acid, linoleic acid, oleic acid / linoleic acid ratio, stearic acid, palmitic acid, lysine, phenylalanine and proline; whole genome resequencing is performed by a sequencing platform to obtain short fragment base sequence data, the short fragment base sequence data is quality controlled to remove short fragment base sequence data containing adapter sequences, with a low-quality base proportion of >5% filtered, and repeated, then the short fragment base sequence data is aligned with a peanut T2T reference genome, GATK software is used for single nucleotide polymorphism detection of SNP sites and filtering of low-quality sites with a quality value <30 and a coverage depth <5, and a SNP variation map is obtained.
[0012] Further, in the SNP optimization and marker screening, based on the SNP variation map, Plink software is used to calculate the missing rate and minimum allele frequency of each SNP site, and a SNP site comprehensive quality index is calculated by a SNP site comprehensive quality screening index formula to screen high-quality SNP sites with an SQI of >0.6; for sites with a small amount of missing data after screening, Beagle software is used to complete missing data filling to obtain a complete SNP genotype data matrix; subsequently, Admixture software is used for population structure analysis, an optimal population structure grouping is determined by cross-validation error, a population structure matrix is output, and Plink software is used to calculate a linkage disequilibrium coefficient; then, TASSEL software is used to carry out whole genome association analysis, the P value of each SNP site associated with ten quality traits is calculated, and the association contribution of the SNP site to the target quality trait is calculated by a SNP-trait association contribution formula, and sites with an association contribution of >5 to the target quality trait are screened to form a characteristic marker set.
[0013] Further, in the SNP optimization and marker screening, a SNP site comprehensive quality index is calculated by a SNP site comprehensive quality screening index formula, and the SNP site comprehensive quality screening index formula is: SQI=0.6x(1-MR)+0.4xMAF, wherein SQI is the SNP site comprehensive quality index, the value range is 0-1, and the higher the value, the better the quality of the SNP site, MR is the SNP site missing rate, and MAF is the SNP site minimum allele frequency.
[0014] Further, in the SNP optimization and marker screening, the association contribution of the SNP site to the target quality trait is calculated by a SNP-trait association contribution formula, and the SNP-trait association contribution formula is: wherein AC is the association contribution of the SNP site to the target quality trait, P is the P value of the SNP site associated with ten quality traits, The linkage disequilibrium coefficient of the SNP site is used to measure the degree of correlation between the genotypes of two SNP sites. The closer to 1, the stronger the correlation between the two sites, the greater the interference; the closer to 0, the weaker the correlation, the smaller the interference.
[0015] Further, in the model construction, the feature marker set is input, and the phenotypes of ten quality traits of the corresponding germplasm are output, a hybrid deep learning model is constructed, including a VAE layer, a CNN layer and a Transformer layer; in the VAE layer training link, the average correlation contribution degree of the feature marker set, that is, the average value of the AC values of all key SNP sites, is introduced, and the VAE layer noise compression loss value is calculated through a noise compression loss formula; after the VAE layer output is extracted by the CNN layer Local linkage features and the Transformer layer captures the epistatic effect on the whole genome, the VAE latent dimension, the CNN convolution kernel parameter and the Transformer attention head number are iteratively optimized and adjusted, and at the same time, the hybrid deep learning model is evaluated through ten-fold cross-validation to determine the determination coefficient of each round of verification, and the average determination coefficient and the determination coefficient standard deviation are analyzed, and then the prediction accuracy correction coefficient is calculated through the prediction accuracy correction coefficient formula, when the prediction accuracy correction coefficient ≥0.8, the whole genome selection model is obtained.
[0016] Further, in the model construction, the VAE layer noise compression loss value is calculated through a noise compression loss formula, and the noise compression loss formula is: Wherein, L VAE is the noise compression loss value, is the mean square error of the genotype original data X and the VAE reconstructed data X is the genotype original data of the key SNP site screened, is the reconstructed output data, and AC avg is the average correlation contribution degree of the feature marker set.
[0017] The prediction accuracy correction coefficient is calculated through the prediction accuracy correction coefficient formula, and the prediction accuracy correction coefficient formula is: Wherein, C acc is the prediction accuracy correction coefficient, is the average determination coefficient, CV std is the R 2 value standard deviation.
[0018] Further, in the breeding of the new strain, the key parents with genetic complementation are selected according to the phenotypic matrix, SQI and AC index and germplasm genetic distance data, and the selection criteria are that the genetic distance between the parents is greater than the average genetic distance of the population, and the parents perform excellently in different quality traits; the breeding population is constructed by using the selected key parents, and after the offspring individuals grow, the DNA is extracted from the tender leaves to obtain the genotypes of the offspring individuals; the genotypes of the offspring individuals are matched with the characteristic marker set to determine the key SNP sites carried by the individuals, and the average association contribution of the key SNP sites carried by the offspring individuals is obtained by using the SNP-trait association contribution formula; the average phenotype values of the parents are obtained by extracting the corresponding quality trait phenotype data of the key parents in the phenotypic matrix and averaging; the genotypes of the offspring individuals are input into the whole genome selection model to obtain the original predicted breeding value, and the corrected breeding value of the offspring individuals is calculated by using the offspring individual correction breeding value formula, and the excellent single plants are screened according to the high breeding value standard that the corrected breeding value of the offspring individuals is greater than or equal to 8, and the excellent single plants are selfed or backcrossed to stabilize the excellent traits; the genotypes of the offspring individuals are obtained, the GBV cal calculation, excellent single plant screening and selfing / backcrossing operation are repeated, and a new high-quality peanut strain with increased oil content, protein and oleic acid is obtained after 3-5 generations of polymerization selection.
[0019] Further, in the breeding of the new strain, the key parents with genetic complementation are selected according to the phenotypic matrix, SQI and AC index and germplasm genetic distance data, and the selection criteria are that the genetic distance between the parents is greater than the average genetic distance of the population, and the parents perform excellently in different quality traits; the breeding population is constructed by using the selected key parents, and after the offspring individuals grow, the DNA is extracted from the tender leaves to obtain the genotypes of the offspring individuals; the genotypes of the offspring individuals are matched with the characteristic marker set to determine the key SNP sites carried by the individuals, and the average association contribution of the key SNP sites carried by the offspring individuals is obtained by using the SNP-trait association contribution formula; the average phenotype values of the parents are obtained by extracting the corresponding quality trait phenotype data of the key parents in the phenotypic matrix and averaging; the genotypes of the offspring individuals are input into the whole genome selection model to obtain the original predicted breeding value, and the corrected breeding value of the offspring individuals is calculated by using the offspring individual correction breeding value formula, and the excellent single plants are screened according to the high breeding value standard that the corrected breeding value of the offspring individuals is greater than or equal to 8, and the excellent single plants are selfed or backcrossed to stabilize the excellent traits; the genotypes of the offspring individuals are obtained, the GBV cal calculation, excellent single plant screening and selfing / backcrossing operation are repeated, and a new high-quality peanut strain with increased oil content, protein and oleic acid is obtained after 3-5 generations of polymerization selection. pred ind par cal pred ind par
[0020] On the other hand, the application of peanut quality trait whole genome selection breeding, the specific steps of the method are:
[0021] Data acquisition module: collect multiple peanut germplasms for planting, take 3 healthy plants at the mature stage, determine ten quality traits and take the average value to construct a phenotypic matrix, extract DNA from the tender leaves of the plants, and perform whole genome resequencing, align the sequencing data with the peanut T2T reference genome, screen SNP sites, and obtain an SNP variation map;
[0022] The SNP optimization and marker screening module: based on the SNP variation map, the missing rate of SNP sites and the minimum allele frequency are calculated, high-quality SNP sites are screened, missing data are filled, population structure is analyzed, linkage disequilibrium is calculated, and key SNP sites are screened to form a characteristic marker set through whole genome association analysis;
[0023] The model construction module: a mixed deep learning model is constructed with the characteristic marker set as input and the corresponding germplasm quality trait phenotype as output, the correlation contribution degree related index is introduced in the training process to adjust the model, the hyperparameters are optimized and screened through cross-validation, and a whole genome selection model is obtained;
[0024] The platform development module adopts the Vue.js+Django framework to develop a peanut intelligent breeding platform, embeds the whole genome selection model, integrates data retrieval, download, visualization and breeding value prediction, parent combination and custom training functions;
[0025] The new line breeding module: according to the combination of the phenotype matrix, the SQI and AC indexes and the genetic distance data of the germplasm, the backbone parents are selected to construct a breeding population, the genotypes of the offspring individuals are obtained and uploaded to the peanut intelligent breeding platform to obtain the predicted breeding value, the corrected breeding value of the offspring individuals is calculated and the excellent single plant is screened, and the high-quality peanut new line is obtained through multi-generation aggregation selection.
[0026] Compared with the prior art, the peanut quality trait selection breeding method based on whole genome SNP and application has the following beneficial effects:
[0027] 1. The present application realizes the prediction and breeding value evaluation of peanut quality traits by integrating whole genome resequencing, SNP optimization and marker screening and mixed deep learning model construction, obtains the SNP variation map through whole genome resequencing and quality control process, covers ten key quality traits such as protein, oil content and oleic acid, provides comprehensive genetic information for subsequent analysis, avoids the interference of low-quality sites by screening out characteristic marker sets through the SNP optimization and marker screening link, optimizes the hyperparameters through ten-fold cross-validation, obtains the whole genome selection model, shortens the breeding cycle, reduces the prediction error rate, and provides technical support for the improvement of peanut quality traits.
[0028] Secondly, the application constructs a backbone parent selection strategy by combining phenotype matrix, SQI / AC index and genetic distance data, ensures the initial genetic diversity of the breeding population through genetic complementarity screening, lays a foundation for multi-generation aggregation selection, integrates whole genome selection model prediction value, average association contribution of key SNP sites and average value of parent phenotype in the offspring evaluation stage, realizes dynamic correction of breeding value, and the accuracy of screening high breeding value individuals is obviously improved. In addition, the platform development module integrates data retrieval, visualization and breeding value prediction functions through the Vue.js+Django framework, can obtain parent mating suggestions in real time, solves the problems of data island and decision lag in traditional breeding, and after 3-5 generations of aggregation selection, the oil content, protein content and oleic acid / linoleic acid ratio of the new strain are obviously improved, the quality traits are improved, and the breeding paradigm for upgrading the rapeseed industry to high oleic acid and high protein is provided.
[0029] Other advantages, objects, and features of the application will be in part apparent and in part pointed out hereinafter in the specification, and in part will be observed by persons skilled in the art upon examination of the following specification, or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0031] Figure 1 Flow chart of the peanut quality trait selection breeding method based on whole genome SNP;
[0032] Figure 2 Framework diagram of the application of the peanut quality trait whole genome selection breeding. DETAILED DESCRIPTION
[0033] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined application purpose, the specific embodiments, structures, features and effects according to the present application will be described in detail below with reference to the drawings and preferred embodiments.
[0034] Example one:
[0035] Data acquisition: In the scenario of breeding new peanut lines with high oleic acid, 150 peanut germplasm resources from different regions were collected. The plants were planted in the test field according to the unified planting specifications. When the germplasm entered the mature stage (45-50 days after flowering), 3 healthy plants were randomly selected, and 10 quality traits were measured, including protein content, oil content, oleic acid content, linoleic acid content, oleic acid / linoleic acid ratio, stearic acid content, palmitic acid content, lysine content, phenylalanine content and proline content. Each plant was measured 3-4 times, and the average value was used to construct the phenotype matrix. At the same time, the tender leaves of the 3 plants were collected, and DNA was extracted. Whole genome resequencing was performed through the sequencing platform to obtain short fragment base sequence data. The short fragment base sequence data was quality controlled to remove short fragments containing adapter sequences, filter short fragments with low quality base ratio > 5%, and remove repeated short fragment base sequence data. The quality controlled short fragment base sequence data was aligned with the peanut T2T reference genome. GATK software was used for SNP site single nucleotide polymorphism detection. Low-quality sites with quality value < 30 and coverage depth < 5 were filtered. Finally, the SNP variation map was obtained.
[0036] SNP optimization and marker screening: Based on the obtained SNP variation map, Plink software was used to calculate the missing rate and minimum allele frequency of each SNP site. The comprehensive quality index of each SNP site was calculated by the SNP site comprehensive quality screening index formula: SQI = 0.6 x (1-MR) + 0.4 x MAF, where SQI is the SNP site comprehensive quality index, the value range is 0-1, the higher the value represents the better quality of the SNP site, MR is the SNP site missing rate, and MAF is the SNP site minimum allele frequency. High-quality SNP sites with SQI ≥ 0.6 were screened out. For sites with a small amount of missing data after screening, Beagle software was used to complete the missing data filling to obtain complete SNP genotype data matrix. Admixture software was used for population structure analysis, and the optimal population structure grouping was determined by cross-validation error. The population structure matrix was output. Plink software was used to calculate the linkage disequilibrium coefficient between the screened high-quality SNP sites. TASSEL software was used for whole genome association analysis to calculate the P value of each SNP site associated with ten quality traits. The SNP-trait association contribution formula was used to calculate the association contribution of each SNP site to the target quality trait (focus on oleic acid content): AC = -log10(P) x R2, where AC is the association contribution of the SNP site to the target quality trait, P is the P value of the SNP site associated with ten quality traits, and R2 is the linkage disequilibrium coefficient of the SNP site. The closer to 1, the stronger the correlation between the two sites, the greater the interference; the closer to 0, the weaker the correlation, the smaller the interference; screening AC≥5 sites to form a characteristic marker set for high-oleic peanut breeding, as shown in Figure 1 .
[0037] Model construction: taking the characteristic marker set screened as input and the phenotypes of ten quality traits of 150 accessions (focusing on the oleic acid content phenotype data) as output, a hybrid deep learning model including VAE layer, CNN layer and Transformer layer is constructed; in the VAE layer training link, first, the average correlation contribution of the characteristic marker set is calculated, that is, the average value of the AC value of the key SNP site, and the VAE layer noise compression loss value is calculated through the noise compression loss formula, so as to optimize the VAE layer parameters, realize the noise reduction and feature extraction of the genotype data, and its noise compression loss formula is: Wherein, L VAE is the noise compression loss value, is the mean square error of the genotype original data X and the VAE reconstruction data X is the original genotype data of the key SNP site screened, is the reconstructed output data, and AC avg is the average correlation contribution of the characteristic marker set;
[0038] The VAE layer output data is transmitted into the CNN layer, the CNN layer convolution kernel size is set to 3×3, the number is 64, and the step is 1, the local linkage features between SNP sites are extracted through convolution operation; then the CNN layer output is transmitted into the Transformer layer, the number of Transformer layer attention heads is set to 8, and the hidden layer dimension is set to 256, the epistatic effect of SNP sites in the whole genome range is captured, and the VAE latent dimension, CNN convolution kernel parameter and Transformer attention head number are iteratively optimized and adjusted, after each adjustment, the hybrid deep learning model is evaluated through ten-fold cross-validation, the determination coefficient of each validation is determined, and the average determination coefficient and the standard deviation of the determination coefficient are analyzed, and the prediction accuracy correction coefficient is calculated through the prediction accuracy correction coefficient formula, the prediction accuracy correction coefficient formula is: Wherein, C acc is the prediction accuracy correction coefficient, is the average determination coefficient, CV std is the standard deviation of R 2 value; when the iterative optimization is C acc ≥0.8, stop training, and obtain a whole genome selection model suitable for high-oleic peanut breeding.
[0039] New strain breeding: combine the phenotypic matrix of 150 accessions (focus on the oleic acid content phenotypic data), the SQI value of each SNP site and the AC value of the oleic acid related SNP site (focus on the AC value of the oleic acid related SNP site) and the genetic distance data between accessions, select complementary backbone parents, the selection criteria are: the genetic distance between parents is greater than the average genetic distance of 150 accessions, and the oleic acid content of the father is ≥ 78%, the protein content of the mother is ≥ 28% and the oil content is ≥ 52%;
[0040] Using the selected father and mother to cross, F1 generation breeding population is constructed, when the F1 generation plants grow to mature period, the DNA of each tender leaf is collected to obtain the F1 generation individual genotype data;
[0041] The F1 generation individual genotype data is matched with the feature marker set composed of the key SNP sites screened to determine the key SNP sites carried by each F1 generation individual, and the AC value of each carrying site is calculated by the SNP-trait association contribution formula, and the average association contribution of each F1 generation individual carrying the key SNP site is obtained, and the ten quality trait phenotypic data of the parents in the phenotypic matrix are extracted, and the average value is obtained to obtain the average phenotype value of the parents, the F1 generation individual genotype data is input into the constructed whole genome selection model to obtain the original predicted breeding value of each F1 generation individual, and then the corrected breeding value of each F1 generation individual is calculated by the offspring individual correction breeding value formula, the offspring individual correction breeding value formula is: GBV cal = GBV pred × AC ind + 0.1 × P par , wherein, GBV cal is the corrected offspring individual quality trait breeding value, GBV pred is the original breeding value predicted by the whole genome selection model for the offspring individual quality trait, AC ind is the average association contribution of the offspring individual carrying the key SNP site, P par is the average phenotype value of the backbone parent on the corresponding quality trait; according to the high breeding value standard of F1 generation individual corrected breeding value ≥ 8, the excellent single plant is screened, and the F1 generation excellent single plant is self-crossed to obtain the F2 generation population; repeat the F1 generation individual genotype acquisition, GBV cal calculation, excellent single plant screening and self-crossing operation, after 3 generations of polymerization selection, a new high-quality high-oleic acid peanut strain with oleic acid content ≥ 80%, protein content ≥ 25% and oil content ≥ 50% is obtained.
[0042] In summary, in the breeding scenario of high-oleic peanut new lines, first, collect peanut germplasm from different regions, standardize planting, and then determine ten quality traits to construct a phenotype matrix. At the same time, extract DNA for resequencing and process the data to obtain an SNP variation map. Then, optimize and screen SNP sites to obtain a characteristic marker set. Next, build a hybrid deep learning model, and through iterative optimization and cross-validation, obtain a whole-genome selection model that meets the requirements. Finally, select the backbone parents according to multi-dimensional data to construct a breeding population, obtain the genotype of the offspring, and calculate the corrected breeding value. After three generations of polymerization selection, a high-oleic peanut new line is successfully bred.
[0043] Example Two:
[0044] Data acquisition: In the breeding scenario of high-protein and high-oil peanut new lines, 180 peanut germplasm resources from different regions are collected. In the experimental field, the germplasm is planted according to the unified planting standard. At the mature stage (48-53 days after flowering), 3 healthy plants are randomly selected to determine ten quality traits (protein, oil content, oleic acid, linoleic acid, oleic acid / linoleic acid ratio, stearic acid, palmitic acid, lysine, phenylalanine, and proline content). Each plant is measured 3-4 times to obtain the average value, and a phenotype matrix is constructed. DNA is extracted from the tender leaves of the 3 plants, and whole-genome resequencing is performed with the help of a sequencing platform to obtain short fragment base sequence data. The short fragment base sequence data is quality controlled to remove short fragment base sequence data containing adapter sequences, low-quality base sequence data with a proportion of >5%, and repeated short fragment base sequence data. The quality-controlled short fragment base sequence data is then aligned with the peanut T2T reference genome. GATK software is used for SNP site single nucleotide polymorphism detection. Low-quality sites with a quality value <30 and a coverage depth <5 are filtered, and finally an SNP variation map is obtained.
[0045] SNP optimization and marker screening: Based on the SNP variation map, the Plink software is used to calculate the missing rate and minimum allele frequency of each SNP site. The comprehensive quality index of each SNP site is calculated by the SNP site comprehensive quality screening index formula: SQI = 0.6 x (1-MR) + 0.4 x MAF. High-quality SNP sites with SQI ≥ 0.6 are selected.
[0046] The sites with missing data after screening were filled with Beagle software to obtain complete SNP genotype data matrix, and then the population structure was analyzed by Admixture software, the optimal grouping was determined by cross-validation error, and the population structure matrix was output, the linkage disequilibrium coefficient between high-quality SNP sites was calculated by Plink software, and whole genome association analysis was carried out by TASSEL software, the P value of each SNP site associated with ten quality traits (focusing on protein and oil content) was calculated, and the association contribution of SNP site to protein and oil content was calculated by the SNP-trait association contribution formula, and the SNP-trait association contribution formula is: And the sites with AC≥5 were screened to form a characteristic marker set for high-protein and high-oil peanut breeding.
[0047] Model construction: taking the characteristic marker set of key SNP sites as input and the ten quality trait phenotypes (focusing on protein and oil content) of 180 accessions as output, a hybrid deep learning model containing VAE layer, CNN layer and Transformer layer was constructed.
[0048] In the VAE layer training link, the average association contribution of the characteristic marker set, that is, the average value of the AC value of the key SNP site, was calculated, the VAE layer noise compression loss value was calculated by the noise compression loss formula, and the VAE layer parameters were optimized, and the noise compression loss formula is:
[0049] The VAE layer output data is transmitted into the CNN layer (convolution kernel 3×3, number 80, step 1) to extract local linkage features, and then transmitted into the Transformer layer (attention head number 10, hidden layer dimension 320) to capture the epistasis on the whole genome, and the VAE latent dimension, CNN convolution kernel number and Transformer attention head number are iteratively optimized, after each optimization, the model is evaluated by ten-fold cross-validation, the determination coefficient of each validation is calculated, and the average determination coefficient and the standard deviation of the determination coefficient are analyzed, and then the prediction accuracy correction coefficient is calculated by the prediction accuracy correction coefficient formula, and the prediction accuracy correction coefficient formula is: When C acc ≥0.8, a whole genome selection model suitable for high-protein and high-oil peanut breeding is obtained, as shown in Figure 2 .
[0050] New strain breeding: combine the phenotype matrix (focus on protein and oil content) of 180 accessions, SQI value and AC value of SNP site and genetic distance, select backbone parents: the genetic distance between parents is greater than the average genetic distance, the protein content of the father is ≥30%, the oil content of the mother is ≥55%, the father and the mother are crossed to construct F1 generation population, the tender leaves of F1 generation are collected at mature stage to extract DNA, the genotype data is obtained, the genotype data of F1 generation individuals is matched with the feature marker set composed of the key SNP sites screened, the key SNP sites carried by each F1 generation individual is determined, and the AC value of each carrying site is calculated by the SNP-trait association contribution formula, the average association contribution of each F1 generation individual carrying the key SNP site is obtained by taking the average value, and the ten quality trait phenotype data of the parents in the phenotype matrix is extracted, and the average phenotype value of the parents is obtained after averaging, the genotype data of F1 generation individuals is input into the constructed whole genome selection model, and the original predicted breeding value of each F1 generation individual is obtained, and then the corrected breeding value of each F1 generation individual is calculated by the offspring individual correction breeding value formula, and the offspring individual correction breeding value formula is: GBV cal = GBV pred × AC ind + 0.1 × P par , GBV cal ≥ 8 is screened out, the F1 generation excellent single plants are self-crossed to obtain F2 generation population; the genotype of F1 generation individuals is obtained, the GBV cal is calculated, the excellent single plants are screened and self-crossed, and after 4 generations of polymerization selection, a high-quality peanut new strain with protein content ≥ 28% and oil content ≥ 54% is obtained.
[0051] In summary, in the breeding of high-protein and high-oil peanut new strains, 180 peanut accessions from different regions are collected for standard planting, quality traits are determined to construct a phenotype matrix, DNA is extracted for resequencing to obtain a SNP variation map, then SNP sites are optimized and screened to form a feature marker set, a mixed deep learning model is constructed, a qualified whole genome selection model is obtained after optimization and verification, then suitable backbone parents are selected for cross to construct a population, the genotype of offspring is obtained and the corrected breeding value is calculated, excellent single plants are screened and 4 generations of polymerization selection are carried out to breed high-protein and high-oil peanut new strains.
[0052] The above merely describes the preferred embodiments of the present application, and is not intended to limit the present application in any form. Although the present application has been disclosed with the preferred embodiments as above, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content to obtain equivalent embodiments with equivalent changes, as long as the changes or modifications do not deviate from the technical solution of the present application. Any modification, change, equivalent change and modification of the above embodiments made according to the technical essence of the present application still belong to the scope of the technical solution of the present application.
Claims
1. A peanut quality trait selection breeding method based on whole-genome SNPs, characterized in that, The specific steps of this method are as follows: Data acquisition: Multiple peanut germplasm samples were collected and planted. Three healthy plants were taken at the germplasm maturity period. Ten quality traits were measured and the average values were used to construct a phenotypic matrix. At the same time, DNA was extracted from the young leaves of the plants and whole-genome resequencing was performed. After processing the sequencing data, it was compared with the peanut T2T reference genome to screen SNP sites and obtain SNP variation maps. SNP optimization and marker screening: Based on the SNP variation map, calculate the SNP site deletion rate and minimum allele frequency, screen high-quality SNP sites and fill in missing data, analyze population structure, calculate linkage disequilibrium, and conduct genome-wide association analysis to screen key SNP sites to form a feature marker set; Model construction: Using the feature tag set as input and the corresponding germplasm quality phenotypic trait as output, a hybrid deep learning model is constructed. During training, correlation contribution-related indicators are introduced to adjust the model, hyperparameters are optimized, and cross-validation is used to screen and obtain a genome-wide selection model. New line breeding: Based on the combination of phenotypic matrix, SQI and AC indexes and germplasm genetic distance data, the backbone parents are selected to construct the breeding population, the genotypes of offspring individuals are obtained, the original breeding values are obtained through the whole genome selection model, the corrected offspring individual breeding values are calculated, and superior single plants are screened. Through multi-generation aggregation selection, high-quality new peanut lines are obtained.
2. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 1, characterized in that, In the data acquisition process, ten quality traits were measured, including protein, oil content, oleic acid, linoleic acid, oleic acid / linoleic acid ratio, stearic acid, palmitic acid, lysine, phenylalanine, and proline. Whole-genome resequencing was performed using a sequencing platform to obtain short-fragment base sequence data. The short-fragment base sequence data underwent quality control, removing short-fragment base sequence data containing adapter sequences, filtering out low-quality bases with a percentage greater than 5%, and eliminating repetitive short-fragment base sequence data. The short-fragment base sequence data was then compared with the peanut T2T reference genome. Single nucleotide polymorphisms (SNPs) at SNP sites were detected using GATK software, and low-quality sites with a quality value <30 and a coverage depth <5 were filtered out to obtain an SNP variation map.
3. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 1, characterized in that, In the SNP optimization and marker screening process, based on the SNP variant map, the deletion rate and minimum allele frequency of each SNP site are calculated using Plink software, and the comprehensive quality index of SNP sites is calculated using the SNP site comprehensive quality screening index formula to screen out high-quality SNP sites with SQI ≥ 0.
6. For loci with a small number of missing data after screening, Beagle software was used to fill in the missing data to obtain a complete SNP genotype data matrix. Then, Admixture software was used to analyze the population structure, and the optimal population structure grouping was determined by cross-validation error. The population structure matrix was output, and Plink software was used to calculate the linkage disequilibrium coefficient. Then, TASSEL software was used to conduct genome-wide association analysis, calculate the p-value of each SNP locus associated with ten quality traits, and calculate the association contribution of SNP loci to the target quality traits using the SNP-trait association contribution formula. Loci with an association contribution of ≥5 to the target quality traits were selected to form a feature marker set for use.
4. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 3, characterized in that, In the SNP optimization and marker screening, the SNP comprehensive quality index is calculated using the SNP comprehensive quality screening index formula: SQI = 0.6 × (1 - MR) + 0.4 × MAF, where SQI is the SNP comprehensive quality index, MR is the SNP deletion rate, and MAF is the minimum allele frequency of the SNP.
5. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 3, characterized in that, In the SNP optimization and labeling screening, the association contribution of SNP loci to the target quality trait is calculated using the SNP-trait association contribution formula, which is as follows: Where AC represents the association contribution of the SNP locus to the target quality trait, and P represents the P-value of the association between the SNP locus and the ten quality traits. This represents the linkage disequilibrium coefficient of the SNP site.
6. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 1, characterized in that, In the model construction, a hybrid deep learning model is constructed using a feature marker set as input and the corresponding ten quality traits of germplasm as output. This model includes a VAE layer, a CNN layer, and a Transformer layer. During the VAE layer training, the average association contribution of the feature marker set, i.e., the average AC value of all key SNP sites, is introduced. The noise compression loss value of the VAE layer is calculated using the noise compression loss formula. After the VAE layer output is processed by the CNN layer to extract local linkage features and the Transformer layer to capture genome-wide epistatic effects, the VAE latent dimension, CNN convolution kernel parameters, and Transformer attention head number are iteratively optimized and adjusted. At the same time, the hybrid deep learning model is evaluated through ten-fold cross-validation to determine the coefficient of determination for each round of validation. The average coefficient of determination and the standard deviation of the coefficient of determination are analyzed. The prediction accuracy correction coefficient is then calculated using the prediction accuracy correction coefficient formula. When the prediction accuracy correction coefficient is ≥0.8, the genome-wide selection model is obtained.
7. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 6, characterized in that, In the model construction, the noise compression loss value of the VAE layer is calculated using the noise compression loss formula, which is as follows: Among them, L VAE This represents the noise compression loss value. Reconstructing data from genotype raw data X and VAE The mean squared error X represents the original data of the genotypes of the key SNP loci obtained through screening. To reconstruct the output data, AC avg The average association contribution of the feature label set; The prediction accuracy correction factor is calculated using the following formula: Among them, C acc This is a correction factor for prediction accuracy. The average coefficient of determination, CV std For R 2 The standard deviation of the values.
8. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 1, characterized in that, In the cultivation of the new line, genetically complementary backbone parents were selected by combining phenotypic matrix, SQI and AC indices, and germplasm genetic distance data. The selection criteria were that the genetic distance between parents was greater than the population average genetic distance, and the parents showed excellent performance in different quality traits. A breeding population was constructed using the selected backbone parents. After the offspring individuals grew, DNA was extracted from young leaves to obtain the genotypes of the offspring individuals. The genotypes of the offspring individuals were matched with a set of characteristic markers to determine the key SNP loci carried by each individual. The AC value of the locus was obtained using the SNP-trait association contribution formula. The average association contribution of key SNP loci carried by offspring individuals is obtained by averaging. Simultaneously, the corresponding quality phenotypic data of the backbone parents are extracted from the phenotypic matrix, and the average phenotypic value of the parents is obtained by averaging. The genotypes of offspring individuals are input into a genome-wide selection model to obtain the original predicted breeding value. Then, the corrected breeding value of the offspring individuals is calculated using the formula for corrected breeding values. Superior individual plants are selected according to the high breeding value standard of ≥8 after correction. These superior individual plants are then self-crossed or backcrossed to stabilize the superior traits. The process of obtaining offspring individual genotypes and GBV is repeated. cal Through calculation, screening of superior single plants, and self-crossing / backcrossing operations, and after 3-5 generations of polymeric selection, high-quality new peanut lines with increased oil content, protein, and oleic acid were obtained.
9. The peanut quality trait selection breeding method based on whole-genome SNPs according to claim 1, characterized in that, In the breeding of the new strain, the corrected individual breeding value of the offspring is calculated using the formula for corrected breeding value of offspring individuals. The formula for corrected breeding value of offspring individuals is: GBV cal =GBV pred ×AC ind +0.1×P par Among them, GBV cal GBV represents the corrected breeding values for individual quality traits in offspring. pred AC represents the original breeding values obtained by predicting the quality traits of offspring individuals using a genome-wide selection model. ind P represents the average association contribution of offspring individuals carrying key SNP loci. par This represents the average phenotypic value of the core parent lines in the corresponding quality traits.
10. An application of the peanut quality trait selection breeding method based on whole-genome SNPs as described in any one of claims 1-9, characterized in that, The application includes: Data acquisition module: Collect multiple peanut germplasm samples for planting. At the germplasm maturity period, take 3 healthy plants, measure 10 quality traits and take the average value to construct a phenotypic matrix. At the same time, extract DNA from the young leaves of the plants and perform whole genome resequencing. After processing the sequencing data, compare it with the peanut T2T reference genome, screen SNP sites, and obtain SNP variation map. SNP optimization and marker screening module: Based on the SNP variation map, calculate the SNP site deletion rate and minimum allele frequency, screen high-quality SNP sites and fill in missing data, analyze population structure, calculate linkage disequilibrium, and conduct genome-wide association analysis to screen key SNP sites to form a feature marker set; Model building module: Using the feature tag set as input and the corresponding germplasm quality phenotypic trait as output, a hybrid deep learning model is constructed. During training, correlation contribution-related indicators are introduced to adjust the model, hyperparameters are optimized, and cross-validation is used to screen and obtain a genome-wide selection model. The platform development module uses the Vue.js+Django framework to develop a peanut smart breeding platform, embedding a whole genome selection model and integrating data retrieval, downloading, visualization, breeding value prediction, parent pairing, and custom training functions. New line breeding module: Based on the combination of phenotypic matrix, SQI and AC indexes and germplasm genetic distance data, the core parents are selected to construct the breeding population, the genotypes of offspring individuals are obtained and uploaded to the peanut smart breeding platform to obtain the predicted breeding value, and the corrected offspring individual breeding value is calculated and the superior single plants are screened. Through multi-generation aggregation selection, high-quality new peanut lines are obtained.
Citation Information
Patent Citations
Whole genome selection-based poplar growth trait optimal prediction system and construction method and application thereof
CN117594129A
Construction method and application of poplar property character optimal prediction system based on whole genome selection
CN117953974A
Geographic surveying and mapping data accurate processing method based on machine learning
CN119027793A
Peanut high-oil germplasm resource numerical prediction data model and construction method and application thereof
CN119314550A
Prediction method and system for new coronavirus variant with higher immune escape capability
CN120183485A
Cited By
Multi-character ginseng breeding material high-throughput intelligent screening method based on deep learning
CN122090927A
A High-Throughput Intelligent Screening Method for Multi-Track Ginseng Breeding Materials Based on Deep Learning
CN122090927B