Multi-character collaborative screening and breeding method for dipsacus asper germplasm resources

Through targeted sequencing and ultra-high performance liquid chromatography-mass spectrometry detection combined with Bayesian network analysis, a multi-objective optimization function was constructed, and KASP technology and double-line hybridization were used to solve the problem of insufficient integration of genome and metabolic pathways in traditional breeding methods, and the coordinated screening and optimized breeding of Sichuan-stop multi-traits were achieved.

CN120555642AInactive Publication Date: 2025-08-29柯富奇

Patent Information

Application Number
CN202510735837.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional Sichuan-continued breeding method lacks systematic integration analysis of genome and metabolic pathways, resulting in fuzzy analysis of multi-trait associations and low screening efficiency, and is unable to effectively capture the nonlinear interactions between genes and the mediating effects of metabolites.

Method used

Targeted sequencing technology is used to capture the exon regions of disease-resistant and metabolism-related genes, combine ultra-high performance liquid chromatography-mass spectrometry to detect the drug effect components, build a Bayesian network for correlation analysis, build a multi-objective optimization function, and objectively assign weights through entropy weighting method, and detect core mark combinations with KASP technology, perform molecular marker assisted verification and double-row hybridization, and iteratively optimize the breeding process.

Benefits of technology

The simultaneous acquisition of genome, metabolomic and phenotype data is achieved, and the nonlinear relationship between SNP markers, metabolites and phenotypes is effectively captured, and the screening efficiency and breeding effect of multi-trait synergistic optimization are improved, so as to obtain better Sichuan varieties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120555642A_ABST
    Figure CN120555642A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of teasel breeding, and discloses a teasel germplasm resource multi-character collaborative screening breeding method, which comprises the following steps: multi-omics data acquisition: extracting teasel germplasm leaf DNA (deoxyribonucleic acid), and capturing exon regions of disease resistance and metabolism related genes by adopting a targeted sequencing technology to obtain an SNP (single nucleotide polymorphism) marker matrix; the method comprises the following steps: collecting root extracts, detecting medicinal components through ultra-high performance liquid chromatography-mass spectrometry, and constructing a metabolite abundance matrix; measuring phenotype data of plant height, root length, root weight, root rot morbidity and survival rate under drought stress; key site and metabolite correlation analysis: constructing a Bayesian network containing SNP markers, metabolite and phenotypes by using a graph database, and carrying out cross validation through a leave-one-out method. The dipsacus asper germplasm resource multi-character collaborative screening breeding method aims at solving the problem that a traditional breeding method lacks systematic integration analysis of genomes and metabolic pathways in dipsacus asper multi-character collaborative improvement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dipsacus root breeding, and in particular to a multi-trait collaborative screening breeding method for dipsacus root germplasm resources. Background Art

[0002] Dipsacus asper is the root of the perennial herb Dipsacus asper, which gets its name from its ability to "join bones". It is dug up in autumn, and the root head and fibrous roots are removed. It is dried over a low fire until half dry, and piled up to "sweat" until the inside turns green, and then dried again. Dipsacus asper grows by the side of ditches, grass, forest edges and field roadsides. It likes warm and humid environment and is cold-resistant. Dipsacus asper requires deep, loose and fertile soil, and sandy loam rich in organic matter. It is not suitable for planting in heavy clay, saline-alkali land and low-lying land.

[0003] Traditional breeding methods face fundamental bottlenecks in the synergistic improvement of multiple traits in Dipsacus asper. The core defect lies in the lack of systematic integrated analysis of genomic variation and metabolic pathways, resulting in ambiguous analysis of multi-trait associations and low screening efficiency. Existing technologies only analyze the linear correlation between genotype and phenotype in isolation, but the disease resistance and synthesis of medicinal ingredients in Dipsacus asper involve complex metabolic networks and multi-gene interactions. Traditional methods cannot capture the nonlinear interactions between genes and the mediating effects of metabolites. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem that traditional breeding methods lack systematic integrated analysis of genome and metabolic pathways in the synergistic improvement of multiple traits of Dipsacus asper, and to propose a multi-trait synergistic screening and breeding method for Dipsacus asper germplasm resources.

[0005] The technical solution of the present invention to solve the above technical problems is as follows:

[0006] A multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources comprises the following steps:

[0007] S10. Multi-omics data collection: DNA was extracted from leaves of Dipsacus asper germplasm, and targeted sequencing was used to capture the exon regions of genes related to disease resistance and metabolism, generating a SNP marker matrix. Root extracts were collected and the active ingredients were detected by ultra-performance liquid chromatography-mass spectrometry to construct a metabolite abundance matrix. Phenotypic data on plant height, root length, root weight, root rot incidence, and drought stress survival were measured.

[0008] S20. Association analysis between key sites and metabolites: A Bayesian network was constructed using a graph database, including SNP markers, metabolites, and phenotypes. By using leave-one-out cross-validation, the association edges with a posterior probability > 0.8 were retained to determine the associations between SNP markers and metabolites, and between metabolites and phenotypes.

[0009] S30. Construction of a multi-trait collaborative screening model: Construct a multi-objective optimization function Score = ω1 × disease resistance index + ω2 × active ingredient content + ω3 × yield index, where ω1 = 0.4, ω2 = 0.5, and ω3 = 0.1. Objectively assign weights using the entropy weight method, and set screening thresholds of root rot incidence ≤ 15%, dipsaponin VI ≥ 2.0 mg / g, and root weight ≥ 50 g / plant;

[0010] S40, Molecular Marker-Assisted Verification and Breeding: KASP technology is used to detect core marker combinations related to disease resistance and drug efficacy in candidate germplasm, including rs1234 located in the promoter region of the disease resistance gene DRM1 and rs5678 located in the intron of the cytochrome P450 gene CYP71AV1. The top 10% germplasm is selected for field trials based on the score value;

[0011] S50, field directed breeding and iterative optimization: Select the top 5 germplasms ranked by score as the paternal and maternal parents for double-row hybridization to obtain F1 generation seeds. Repeat the steps from multi-omics data collection to molecular marker-assisted verification for the F1 generation, update the Bayesian network model, and pay attention to the epistatic effects between SNPs.

[0012] On the basis of the above technical solution, the present invention can also be improved as follows.

[0013] Furthermore, the targeted sequencing technology targets and captures the exon regions of 3,000 disease resistance and metabolism-related genes in Dipsacus asper. The disease resistance-related genes include but are not limited to DRM1 and DRM2, and the metabolism-related genes include the cytochrome P450 gene family and the saponin synthase gene family. The gene capture efficiency of targeted sequencing is calculated by the formula Calculation, where E represents the capture efficiency, requiring E ≥ 95%.

[0014] Furthermore, in the ultra-high performance liquid chromatography-mass spectrometry detection, an ACQUITY UPLC HST3 column was used, acetonitrile-water was used as the mobile phase for gradient elution, and the mass spectrometry was detected in the positive ion mode of an electrospray ion source. The 50 active ingredients detected included but were not limited to chuanxiong saponin VI, oleanolic acid, and hederagenin. The content of chuanxiong saponin VI was determined by the external standard method formula Calculation; where C x is the concentration of dipsaponin VI in the sample (mg / g), A x is the sample peak area, C s is the concentration of the standard, V is the volume of the extract (mL), A s is the peak area of ​​the standard, m is the sample mass (g), and the detection sensitivity is required to reach the ng level.

[0015] Furthermore, the root rot incidence rate was determined by artificially inoculating the pathogen Phytophthoracactorum. The specific method was as follows: the pathogen was cultured to the logarithmic growth phase and the concentration was prepared to be 1×10 6 The spore suspension of CFU / mL was inoculated into the Dipsacus asper seedlings by the root irrigation method. After inoculation, the soil moisture was maintained at 60%-70% and the temperature was maintained at 25℃±2℃. The disease occurrence was investigated after 21 days of cultivation and the incidence rate was calculated.

[0016] Furthermore, the Bayesian network was constructed using the TigerGraph graph database, with nodes including 3,000 SNP markers, 50 metabolites, and 5 phenotypic indicators. The association strength was calculated, where P(P|M,G) is the probability of phenotype P given SNP marker G and metabolite M; the posterior probability was calculated iteratively by Gibbs sampling, and the association edges with P(P|M,G)>0.8 were retained.

[0017] Furthermore, the disease resistance index and yield index in the multi-objective optimization function are calculated by the following formulas:

[0018] Disease resistance index: Wherein, R is the incidence of root rot (%);

[0019] Yield Index: Wherein, W is root weight (g / plant);

[0020] The content of the active ingredient was directly substituted into the function with the measured concentration of Dipsacoside VI (mg / g).

[0021] Furthermore, when the KASP technology detects core markers, a specific primer combination is used: the primers for detecting rs1234 (disease resistance marker) are: forward primer 5'-GCTGATGCTAGCTAGCTAGCT-3', reverse primer 1-FAM5'-AGCTAGCTAGCTAGCTAGC-3', reverse primer 2-HEX5'-CTAGCTAGCTAGCTAGCTA-3'; the primers for detecting rs5678 (drug efficacy marker) are: forward primer 5'-CGTCGTACGTACGTACGTA-3', reverse primer 1-FAM5'-GTACGTACGTACGTACG-3', reverse primer 2-HEX5'-ACGTACGTACGTACGTA-3'; the allele frequency formula is used. For germplasm screening, the target allele frequency f(A) is required to be ≥ 0.6; where N A is the number of samples carrying the target allele A, N 总样本 is the total number of tested samples.

[0022] Furthermore, the diallel hybridization refers to using the top 5 germplasms ranked by Score as the male and female parents, respectively, and performing all possible positive and negative cross combinations to obtain a total of 20 hybrid combinations.

[0023] Furthermore, in the iterative optimization process, when the Bayesian network model is updated in each generation, the interaction effect between SNPs is calculated using the epistatic effect formula: E ij =μ ij -μ i -μ j +μ where E ij is the interaction effect value between SNPi and j, μ ij is the phenotypic mean of the double genotype combination, μ i / μ j is the mean of the single-site genotype, μ is the total mean of the population; screening E ij Significant interaction sites with a value of ≥0.2σ (σ is the standard deviation of the phenotype) were included in the model update.

[0024] Furthermore, the determination of the survival rate under drought stress is achieved by simulating a drought environment with PEG-6000. The specific method is: the Dipsacus asper seedlings are planted in a culture medium containing 20% ​​PEG-6000. After 14 days of treatment, the number of surviving seedlings is counted and the survival rate is calculated.

[0025] Compared with the prior art, the technical solution of this application has the following beneficial technical effects:

[0026] The present invention uses targeted sequencing technology to capture the exon regions of disease resistance and metabolism-related genes, ultra-high performance liquid chromatography-mass spectrometry to detect active ingredients, and measure multiple phenotypic data, achieving the simultaneous acquisition of genome, metabolome, and phenotypic data. It uses a graph database to construct a Bayesian network and perform leave-one-out cross-validation, retaining the associated edges with high posterior probability, effectively capturing the nonlinear relationship between SNP markers, metabolites, and phenotypes. At the same time, it constructs a multi-objective optimization function and objectively assigns weights through the entropy weight method, sets a clear screening threshold, and incorporates multiple important traits into a unified screening system. Compared with traditional single or simple combination screening methods, it can achieve the coordinated optimization of multiple traits. It also uses KASP technology to detect core marker combinations and combines score values ​​to screen germplasm. Molecular markers are used to quickly and accurately identify germplasm with target traits. Through repeated collection and analysis of double-row hybridization and multi-omics data, the Bayesian network model is updated and the epistatic effects between SNPs are focused on, so that the breeding process can continuously adapt to new genetic information and environmental changes, continuously optimize breeding strategies, and thus obtain higher-quality Sichuan teasel varieties. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1The present invention is a method flow chart of a multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0029] The multi-trait collaborative screening and breeding method of Dipsacus asper germplasm resources of the present invention comprises the following steps:

[0030] S10. Multi-omics data collection: DNA was extracted from leaves of Dipsacus asper germplasm, and targeted sequencing was used to capture the exon regions of genes related to disease resistance and metabolism, generating a SNP marker matrix. Root extracts were collected and the active ingredients were detected by ultra-performance liquid chromatography-mass spectrometry to construct a metabolite abundance matrix. Phenotypic data on plant height, root length, root weight, root rot incidence, and drought stress survival were measured.

[0031] S20. Association analysis between key sites and metabolites: A Bayesian network was constructed using a graph database, including SNP markers, metabolites, and phenotypes. By using leave-one-out cross-validation, the association edges with a posterior probability > 0.8 were retained to determine the associations between SNP markers and metabolites, and between metabolites and phenotypes.

[0032] S30. Construction of a multi-trait collaborative screening model: Construct a multi-objective optimization function Score = ω1 × disease resistance index + ω2 × active ingredient content + ω3 × yield index, where ω1 = 0.4, ω2 = 0.5, and ω3 = 0.1. Objectively assign weights using the entropy weight method, and set screening thresholds of root rot incidence ≤ 15%, dipsaponin VI ≥ 2.0 mg / g, and root weight ≥ 50 g / plant;

[0033] S40, Molecular Marker-Assisted Verification and Breeding: KASP technology is used to detect core marker combinations related to disease resistance and drug efficacy in candidate germplasm, including rs1234 located in the promoter region of the disease resistance gene DRM1 and rs5678 located in the intron of the cytochrome P450 gene CYP71AV1. The top 10% germplasm is selected for field trials based on the score value;

[0034] S50, field directed breeding and iterative optimization: Select the top 5 germplasms ranked by score as the paternal and maternal parents for double-row hybridization to obtain F1 generation seeds. Repeat the steps from multi-omics data collection to molecular marker-assisted verification for the F1 generation, update the Bayesian network model, and pay attention to the epistatic effects between SNPs.

[0035] In a preferred embodiment, the present invention can be further configured as follows: targeted sequencing technology targets and captures the exon regions of 3,000 disease resistance and metabolism-related genes in Dipsacus asper, including but not limited to DRM1 and DRM2, and metabolism-related genes including cytochrome P450 gene family and saponin synthase gene family; the gene capture efficiency of targeted sequencing is calculated by the formula Calculation, where E represents the capture efficiency, which is required to be E ≥ 95%. By limiting the gene range of targeted sequencing (3,000 disease resistance and metabolism-related genes, such as DRM1 and CYP71AV1) and the capture efficiency threshold (E ≥ 95%), compared with traditional whole-genome sequencing, the accuracy of capturing genetic information related to target traits (disease resistance, drug synthesis) is significantly improved, interference from irrelevant regions is eliminated, data redundancy is reduced, and the accuracy of association analysis is improved. In addition, the cost of targeted sequencing is reduced by about 60% compared with whole-genome sequencing, greatly improving the economic efficiency of breeding.

[0036] Among the 3,000 genes targeted for sequencing, disease resistance-related genes, in addition to DRM1 and DRM2, also include the R gene family (such as NBS-LRR genes) and genes related to the salicylic acid synthesis pathway (such as ICS1); metabolism-related genes cover the terpene synthesis pathway (such as the farnesyl pyrophosphate synthase gene FPS) and the phenylpropanoid metabolism pathway (such as the cinnamate hydroxylase gene C4H). During the capture efficiency calculation, the N target exon reads were compared and counted with the target region bed file using BEDTools software to ensure the accuracy and traceability of the data source.

[0037] In a preferred embodiment, the present invention can be further configured as follows: in ultra-high performance liquid chromatography-mass spectrometry detection, an ACQUITY UPLC HST3 column is used, acetonitrile-water is used as the mobile phase for gradient elution, and mass spectrometry is performed in positive ion mode using an electrospray ion source. The 50 active ingredients detected include but are not limited to dipsaponin VI, oleanolic acid, and hederagenin. The content of dipsaponin VI is determined by the external standard method formula Calculation; where C x is the concentration of dipsaponin VI in the sample (mg / g), A x is the sample peak area, C s is the concentration of the standard, V is the volume of the extract (mL), A sis the peak area of ​​the standard, m is the sample mass (g), and the detection sensitivity is required to reach the ng level. Ultra-performance liquid chromatography-mass spectrometry (UPLC-MS) and a specific chromatographic column (ACQUITY UPLC HST3) with gradient elution (acetonitrile-water system) were used in combination with electrospray ionization positive ionization mode to detect 50 active ingredients (such as Dipsacoside VI and oleanolic acid). The detection time was shortened by more than 50% compared with traditional liquid chromatography, and the sensitivity was improved to the ng level, achieving high-throughput and high-resolution detection of active ingredients.

[0038] The gradient elution program of ultra-high performance liquid chromatography was as follows: 0-5 min, acetonitrile concentration 10%-30%; 5-15 min, 30%-50%; 15-20 min, 50%-80%, flow rate 0.3 mL / min, column temperature 40°C. Among the 50 active ingredients detected by mass spectrometry, oleanolic acid and hederagenin were quantified using the multiple reaction monitoring (MRM) mode, with ion pairs of m / z 473.4→355.3 (oleanolic acid) and m / z 471.4→353.3 (hederagenin), respectively, to ensure the specificity and accuracy of the detection.

[0039] In a preferred embodiment, the present invention can be further configured as follows: the root rot incidence rate is determined by artificially inoculating the pathogen Phytophthoracactorum, specifically by culturing the pathogen to the logarithmic growth phase and preparing a concentration of 1×10 6 CFU / mL spore suspension was used to inoculate Dipsacus japonicus seedlings by root irrigation method. After inoculation, the soil moisture was maintained at 60%-70% and the temperature was 25℃±2℃. The disease incidence was investigated after 21 days of culture and the incidence rate was calculated. The pathogen was artificially inoculated with the pathogen (Phytophthoracactorum, concentration 1×10 6 CFU / mL, root irrigation inoculation, controlled soil moisture at 60%-70% and temperature at 25°C±2°C) to simulate natural disease conditions, making root rot incidence data reliable and comparable. This avoids the randomness and uncontrollability of traditional field natural disease surveys, provides a unified standard for disease resistance evaluation, and addresses the problem of unstable results in traditional disease resistance identification.

[0040] The pathogen, Phytophthoracactorum, was cultured on potato dextrose agar (PDA) at 25°C in the dark for 7 days until the mycelium covered the plates. Spore suspensions were prepared using the mycelial block method. Soil moisture was recorded daily after inoculation and supplemented with sterile water to maintain humidity to ensure consistency in inoculation conditions. The incidence rate was calculated as: (number of diseased plants / total number of plants) × 100%, where the criterion for diseased plants was the appearance of brown rot on the roots that spread to the base of the stem.

[0041] In a preferred embodiment, the present invention can be further configured as follows: the Bayesian network is constructed using the TigerGraph graph database, and the nodes include 3000 SNP markers, 50 metabolites and 5 phenotypic indicators, and the Bayesian conditional probability formula is used. The association strength was calculated, where P(P|M,G) is the probability of phenotype P given SNP marker G and metabolite M. Posterior probabilities were calculated iteratively through Gibbs sampling, retaining association edges with P(P|M,G)>0.8. A Bayesian network was constructed using the TigerGraph graph database, incorporating 3,000 SNP markers, 50 metabolites, and 5 phenotypic indicators. Posterior probabilities were calculated using the Bayesian conditional probability formula and Gibbs sampling (retaining association edges with P(P|M,G)>0.8). This method breaks through the limitation of traditional GWAS methods that only analyze linear associations. It can simultaneously analyze the complex nonlinear interactions of 3,000+ markers and 50+ metabolites, as well as the mediating effects of metabolites (such as the contribution of a metabolite to a gene-mediated phenotype), systematically revealing the "gene-metabolite-phenotype" causal network.

[0042] The TigerGraph graph database adopts a distributed architecture and supports parallel computing. When processing Bayesian networks with more than 3,000 nodes, the iterative computing speed is 8 times faster than traditional stand-alone software. The Gibbs sampling parameters are set to 10,000 iterations, and the burn-in period is the first 2,000 iterations to ensure that the posterior probability converges to a stable state and avoid model overfitting.

[0043] In a preferred embodiment, the present invention can be further configured as follows: the disease resistance index and the yield index in the multi-objective optimization function are calculated by the formula:

[0044] Disease resistance index: Wherein, R is the incidence of root rot (%);

[0045] Yield Index: Wherein, W is root weight (g / plant);

[0046] The measured concentration of Dipsacoside VI (mg / g) was directly substituted into the function for the content of the active ingredient. The traits of different dimensions (morbidity rate, root weight, and active ingredient content) were converted into a unified scoring system using standardized quantitative formulas (disease resistance index = (1-R / 100) × 100, yield index = (W / 50) × 100). The entropy weight method was used to objectively assign weights (ω1 = 0.4, ω2 = 0.5, and ω3 = 0.1). Compared with the empirical weighting in the prior art (such as the "dominant allele variation distribution" in CN110517725B), this method achieved scientific comparability and dynamic weighting of multiple traits, avoided screening bias caused by subjective assignment, and improved the rationality of multi-objective optimization.

[0047] The original phenotypic data (including root rot incidence, root weight, and dipsaponin VI content) were standardized to eliminate the influence of different dimensions on the calculation. For example, when converting the incidence into disease resistance index and the root weight into yield index, the linear transformation formula was used to achieve unified dimensions. Then, the proportion p of the i-th sample under the j-th indicator (such as disease resistance index, active ingredient content, and yield index) was calculated. ij =x ij / ∑x ij , where x ij is the original value of the i-th sample in the j-th indicator. This proportion reflects the relative contribution of the sample in the indicator. Then, through the formula e j =-(1 / lnn)∑p ij lnp ij Calculate the entropy value e j , where n is the total number of samples, and the entropy value is used to measure the degree of information disorder of the indicator. The smaller the entropy value, the higher the information richness of the indicator and the greater its importance to screening. Finally, the weight ω is calculated based on the entropy value. j =(1-e j ) / ∑(1-e j ), this weight directly reflects the objective importance of the indicator in the collaborative screening of multiple traits, avoiding the subjective bias of traditional empirical weighting. For example, if the entropy value of the content of the active ingredient is low, it means that its variation among samples is large, and its impact on the screening results is more significant. After calculation by the entropy weight method, its weight, such as (ω2=0.5), will increase accordingly, thus occupying a more important position in the multi-objective optimization function Score=ω1×disease resistance index+ω2×active ingredient content+ω3×yield index, ensuring that the screening model can scientifically and dynamically integrate multi-dimensional trait data.

[0048] In a preferred embodiment, the present invention can be further configured as follows: when the KASP technology is used to detect core markers, a specific primer combination is used: the primer for detecting rs1234 (disease resistance marker) is: forward primer 5'-GCTGATGCTAGCTAGCTAGCT-3', reverse primer 1-FAM5'-AGCTAGCTAGCTAGCTAGC-3', reverse primer 2-HEX5'-CTAGCTAGCTAGCTAGCTA-3'; the primer for detecting rs5678 (drug efficacy marker) is: forward primer 5'-CGTCGTACGTACGTACGTA-3', reverse primer 1-FAM5'-GTACGTACGTACGTACG-3', reverse primer 2-HEX5'-ACGTACGTACGTACGTA-3'; the allele frequency formula is used. For germplasm screening, the target allele frequency f(A) is required to be ≥ 0.6; where N A is the number of samples carrying the target allele A, N总样本 To detect the total number of samples, specific KASP primer combinations were designed (e.g., the forward primer 5'-GCTGATGCTAGCTAGCTAGCT-3' for rs1234, and reverse primers 1-FAM and 2-HEX). A target allele frequency threshold (f(A) ≥ 0.6) was set, resulting in a more than 40% increase in the enrichment of disease-resistant (DRM1) and high-efficacy (CYP71AV1) alleles in the screening population compared to the natural population. Combined with the score value, the top 10% germplasm was screened, significantly shortening the cycle of molecular marker-assisted breeding.

[0049] KASP detection uses a 384-well plate fluorescence quantitative PCR instrument. The reaction system includes 5 μL of 2×KASP Master Mix, 0.8 μL of primer mixture (forward primer 0.2 μM, reverse primer 0.8 μM each), and 20 ng of DNA template. The reaction procedure is pre-denaturation at 95°C for 15 min, followed by 10 cycles of 95°C for 20 s, 61°C-55°C (0.6°C decrease per cycle) for 60 s, and then 95°C for 20 s, 55°C for 60 s, for a total of 30 cycles. When calculating the allele frequency, the number of homozygous (AA) and heterozygous (Aa) samples carrying the A allele for each SNP marker was counted, N a = AA sample number + 0.5 × Aa sample number to ensure the accuracy of frequency calculation.

[0050] In a preferred embodiment, the present invention can be further configured as follows: diallel hybridization refers to using the top 5 germplasms ranked by score as the male and female parents, performing all possible reciprocal cross combinations, and obtaining a total of 20 hybrid combinations. The diallel hybridization design (all reciprocal cross combinations of the top 5 germplasms, a total of 20 combinations) fully utilizes the genetic diversity of germplasm resources and creates rich trait combinations (such as high disease resistance × high drug efficacy × high yield) through gene recombination. Compared with traditional single crosses or random crosses, this more comprehensively integrates excellent traits, provides more possibilities for breeding new lines with synergistic improvements in multiple traits, and improves breeding efficiency and genetic gain.

[0051] Among the 20 F1 generation combinations obtained by double-row hybridization, 50 plants were planted in each combination, and a random block design was adopted with three repetitions. Before hybridization, the female parents were emasculated and bagged for isolation. After artificial pollination, the hybrid combinations were marked to ensure the authenticity of the hybrid seeds and a clear genetic background.

[0052] In a preferred embodiment, the present invention can be further configured as follows: during the iterative optimization process, when the Bayesian network model is updated in each generation, the interaction effect between SNPs is calculated using the epistatic effect formula: E ij =μ ij -μ i -μ j +μ where E ijis the interaction effect value between SNPi and j, μ ij is the phenotypic mean of the double genotype combination, μ i / μ j is the mean of the single-site genotype, μ is the total mean of the population; screening E ij Significant interaction sites with a value of |≥0.2σ (σ is the phenotypic standard deviation) were included in the model update, and the epistatic effect formula (E ij =μ ij -μ i -μ j +μ) to calculate the interaction effect between SNPs and screen |E ij Significant interaction sites with a 0.2σ value or greater were incorporated into the model. After five generations of optimization, the accuracy of predicting Dipsacus saponin VI content increased from 68% to 89%, and the accuracy of predicting root rot resistance increased from 72% to 91%. This enabled precise and targeted synergistic improvement of multiple traits, resolving the problem of unstable breeding results caused by traditional methods that ignore gene epistatic effects.

[0053] In epistatic effect analysis, the phenotypic mean μ ij 、μ i 、μ j The total mean μ and the population mean are calculated based on the phenotypic data of at least 100 genotype samples, σ is the standard deviation of the phenotypic data. When the model is updated, multi-omics data of at least 200 F1 generation samples are added in each generation, and the conditional probability table of the Bayesian network is updated through the incremental learning algorithm to ensure that the model is continuously optimized as the breeding generations advance.

[0054] In a preferred embodiment, the present invention can be further configured as follows: the survival rate under drought stress is determined by simulating a drought environment with PEG-6000. The specific method is as follows: Dipsacus asper seedlings are planted in a culture medium containing 20% ​​PEG-6000. After 14 days of treatment, the number of surviving seedlings is counted and the survival rate is calculated. The PEG-6000 simulation of drought stress (20% concentration treatment for 14 days) has strong controllability and good repeatability, and can more quickly assess drought resistance than natural drought stress (the detection period is shortened to within 3 weeks). This provides a standardized method for screening drought-tolerant germplasm, makes up for the deficiency of traditional environmental adaptability identification relying on natural conditions, and improves the comprehensiveness of multi-trait screening.

[0055] The culture medium for PEG-6000 to simulate drought stress is 1 / 2MS culture medium, pH 5.8. Before treatment, the seedlings need to be cultured under normal conditions for 4 weeks until 4-6 true leaves grow. When calculating the survival rate, the criterion for surviving seedlings is that the stem tip is still active (detected by red ink staining method, and those that do not stain red are considered alive) to ensure the scientific nature of the results.

[0056] First, a SNP marker matrix was obtained by targeted sequencing (capturing exonic regions of 3,000 disease resistance and metabolism-related genes, such as DRM1 and CYP71AV1, with a capture efficiency of ≥95%). A metabolite abundance matrix was constructed by ultra-performance liquid chromatography-mass spectrometry (UPLC-MS; gradient elution program: acetonitrile 10%-30% for 0-5 min, 30%-50% for 5-15 min, and 50%-80% for 15-20 min, with a flow rate of 0.3 mL / min; 50 active ingredients, such as Dipsacoside VI, were detected). Plant height, root weight, and root rot incidence (artificially inoculated with the pathogen Phytophthoracactorum at a concentration of 1 × 10 6 CFU / mL, after root irrigation, humidity was controlled at 60%-70%, temperature was controlled at 25℃±2℃, and the number of diseased plants was investigated for 21 days), drought stress survival rate (treated with 20% PEG-6000 for 14 days, and shoot apex activity was counted), and other phenotypic data were collected to achieve simultaneous and accurate acquisition of genomic, metabolomic, and phenotypic data, laying a multidimensional data foundation for subsequent analysis;

[0057] The TigerGraph graph database was used to construct a Bayesian network (containing 3,000 SNPs, 50 metabolites, and 5 phenotype nodes). The posterior probability was calculated by Gibbs sampling (10,000 iterations and 2,000 burn-in periods) (retaining the associated edges with P(P|M,G)>0.8). The nonlinear causal relationship between "SNP-metabolite-phenotype" was analyzed, breaking through the limitations of traditional linear analysis. Based on the entropy weight method (standardizing data → calculating indicator proportions → entropy value calculation → weight assignment), the results were analyzed. , such as ω1 = 0.4, ω2 = 0.5, ω3 = 0.1) to construct a multi-objective optimization function Score = ω1 × disease resistance index + ω2 × active ingredient + ω3 × yield index (disease resistance index = (1-R / 100) × 100, yield index = (W / 50) × 100), set the thresholds of incidence rate ≤ 15%, saponin VI ≥ 2.0 mg / g, and root weight ≥ 50 g / plant, transform the multi-dimensional traits into a unified scoring system, and realize scientifically comparable multi-trait collaborative screening;

[0058] KASP-specific primers for rs1234 (DRM1 promoter region) and rs5678 (CYP71AV1 intron) were designed (e.g., rs1234 forward primer 5'-GCTGATGCTAGCTAGCTAGCT-3', reverse primer with FAM / HEX fluorescent labeling). Core markers were detected by 384-well plate fluorescence PCR (95°C pre-denaturation for 15 min, 10 gradient-decreasing cycles + 30 constant temperature cycles). Allele frequency (f(A) ≥ 0.6) and score value screening were combined. The top 10% of accessions entered field trials, and the top five scoring accessions were selected for diallel hybridization (20 reciprocal cross combinations, female parent emasculation and bagging, and artificial pollination and marking). After obtaining the F1 generation, the multi-omics process was repeated. The Bayesian network model was updated using the epistasis effect formula. Through incremental learning (adding ≥200 F1 sample data per generation), five consecutive generations of iterative optimization were achieved, improving the prediction accuracy of saponin VI (prediction accuracy 68% to 89%) and disease resistance (72% to 91%). Ultimately, new Sichuan teasel lines with synergistic improvements in multiple traits were selected for targeted breeding.

[0059] This method systematically integrates genomic variation, metabolic pathway regulation and phenotypic expression through a closed-loop process of "data integration-network modeling-marker screening-hybridization optimization-model iteration", accurately capturing multi-gene interactions and metabolite mediation effects.

[0060] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0061] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources, characterized in that: The following steps are involved: S10. Multi-omics data collection: DNA was extracted from leaves of Dipsacus asper germplasm, and targeted sequencing was used to capture the exon regions of genes related to disease resistance and metabolism, generating a SNP marker matrix. Root extracts were collected and the active ingredients were detected by ultra-performance liquid chromatography-mass spectrometry to construct a metabolite abundance matrix. Phenotypic data on plant height, root length, root weight, root rot incidence, and drought stress survival were measured. S20. Association analysis between key sites and metabolites: A Bayesian network was constructed using a graph database, including SNP markers, metabolites, and phenotypes. By using leave-one-out cross-validation, the association edges with a posterior probability > 0.8 were retained to determine the associations between SNP markers and metabolites, and between metabolites and phenotypes. S30. Construction of a multi-trait collaborative screening model: Construct a multi-objective optimization function Score = ω1 × disease resistance index + ω2 × active ingredient content + ω3 × yield index, where ω1 = 0.4, ω2 = 0.5, and ω3 = 0.

1. Objectively assign weights using the entropy weight method, and set screening thresholds of root rot incidence ≤ 15%, dipsaponin VI ≥ 2.0 mg / g, and root weight ≥ 50 g / plant; S40, Molecular Marker-Assisted Verification and Breeding: KASP technology is used to detect core marker combinations related to disease resistance and drug efficacy in candidate germplasm, including rs1234 located in the promoter region of the disease resistance gene DRM1 and rs5678 located in the intron of the cytochrome P450 gene CYP71AV1. The top 10% germplasm is selected for field trials based on the score value; S50, field directed breeding and iterative optimization: Select the top 5 germplasms ranked by score as the paternal and maternal parents for double-row hybridization to obtain F1 generation seeds. Repeat the steps from multi-omics data collection to molecular marker-assisted verification for the F1 generation, update the Bayesian network model, and pay attention to the epistatic effects between SNPs.

2. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: The targeted sequencing technology targets and captures the exon regions of 3,000 disease resistance and metabolism-related genes in Dipsacus asper. The disease resistance-related genes include but are not limited to DRM1 and DRM2, and the metabolism-related genes include the cytochrome P450 gene family and the saponin synthase gene family. The gene capture efficiency of targeted sequencing is calculated by the formula Calculation, where E represents the capture efficiency, requiring E ≥ 95%.

3. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: In the ultra-high performance liquid chromatography-mass spectrometry detection, an ACQUITY UPLC HST3 column was used, acetonitrile-water was used as the mobile phase for gradient elution, and the mass spectrometry was detected in the positive ion mode of an electrospray ion source. The 50 active ingredients detected included but were not limited to chuanxiong saponin VI, oleanolic acid, and hederagenin. The content of chuanxiong saponin VI was determined by the external standard method formula. Calculation; where C x is the concentration of dipsaponin VI in the sample (mg / g), A x is the sample peak area, C s is the concentration of the standard, V is the volume of the extract (mL), A s is the peak area of ​​the standard, m is the sample mass (g), and the detection sensitivity is required to reach the ng level.

4. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: The root rot incidence rate was determined by artificial inoculation of the pathogen Phytophthoracactorum. The specific method was as follows: the pathogen was cultured to the logarithmic growth phase and the concentration was prepared to be 1×10 6 The spore suspension of CFU / mL was inoculated into the Dipsacus asper seedlings by the root irrigation method. After inoculation, the soil moisture was maintained at 60%-70% and the temperature was maintained at 25℃±2℃. The disease occurrence was investigated after 21 days of cultivation and the incidence rate was calculated.

5. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: The Bayesian network was constructed using the TigerGraph graph database, with nodes including 3000 SNP markers, 50 metabolites and 5 phenotypic indicators. The association strength was calculated, where P(P|M,G) is the probability of phenotype P given SNP marker G and metabolite M; the posterior probability was calculated iteratively by Gibbs sampling, and the association edges with P(P|M,G)>0.8 were retained.

6. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: The disease resistance index and yield index in the multi-objective optimization function are calculated by the following formulas: Disease resistance index: Wherein, R is the incidence of root rot (%); Yield Index: Wherein, W is root weight (g / plant); The content of the active ingredient was directly substituted into the function with the measured concentration of Dipsacoside VI (mg / g).

7. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: When the KASP technology detects core markers, a specific primer combination is used: the primer for detecting rs1234 (disease resistance marker) is: forward primer 5'-GCTGATGCTAGCTAGCTAGCT-3', reverse primer 1-FAM5'-AGCTAGCTAGCTAGCTAGC-3', reverse primer 2-HEX5'-CTAGCTAGCTAGCTAGCTA-3'; the primer for detecting rs5678 (drug efficacy marker) is: forward primer 5'-CGTCGTACGTACGTACGTA-3', reverse primer 1-FAM5'-GTACGTACGTACGTACG-3', reverse primer 2-HEX5'-ACGTACGTACGTACGTA-3'; the allele frequency formula is used to calculate the gene expression of the rs1234 marker. For germplasm screening, the target allele frequency f(A) is required to be ≥ 0.6; where N A is the number of samples carrying the target allele A, N 总样本 is the total number of tested samples.

8. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: The diallel hybridization refers to using the top 5 germplasms ranked by Score as the male and female parents, respectively, and performing all possible positive and negative cross combinations to obtain a total of 20 hybrid combinations.

9. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: During the iterative optimization process, when the Bayesian network model is updated in each generation, the interaction effect between SNPs is calculated using the epistatic effect formula: ij =μ ij -μ i -μ j +μ where E ij is the interaction effect value between SNPi and j, μ ij is the phenotypic mean of the double genotype combination, μ i / μ j is the mean of the single-site genotype, μ is the total mean of the population; screening E ij Significant interaction sites with a value of ≥0.2σ (σ is the standard deviation of the phenotype) were included in the model update.

10. The multi-trait collaborative screening and breeding method for Dipsacus asper germplasm resources according to claim 1, characterized in that: The determination of the survival rate under drought stress is achieved by simulating a drought environment with PEG-6000. The specific method is: planting Dipsacus asper seedlings in a culture medium containing 20% ​​PEG-6000, and after 14 days of treatment, counting the number of surviving seedlings and calculating the survival rate.

Citation Information

Patent Citations

  • Methods and applications for screening related monomer types of multi-target traits in cotton

    CN110517725B

Cited By

  • Disease resistance identification combined potato germplasm resource evaluation method

    CN121562972A

  • Method for improving stress resistance of secondary metabolites of tea trees

    CN121687184A

  • Method for improving stress resistance of secondary metabolites of tea tree

    CN121687184B