Molecular marker related to oil extraction rate of carya illinoensis and application of molecular marker

By developing SNP molecular markers and specific primer sets related to the oil yield of thin-shelled pecans, the technical bottleneck of oil yield assessment in thin-shelled pecan breeding has been solved, enabling precise and efficient breeding at the seedling stage and improving the development quality of the thin-shelled pecan industry.

CN122012805AActive Publication Date: 2026-05-12INST OF BOTANY JIANGSU PROVINCE & CHINESE ACADEMY OF SCI
View PDF 4 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF BOTANY JIANGSU PROVINCE & CHINESE ACADEMY OF SCI
Filing Date
2026-04-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, the assessment of oil yield in thin-shelled pecans relies on physicochemical testing after fruit ripening, which results in long breeding cycles, waste of resources, and high testing costs. Furthermore, the lack of precise molecular markers for seedling selection makes it difficult to meet the breeding needs of high-oil varieties.

Method used

We developed SNP molecular markers related to the oil yield of thin-shelled pecans, identified SNP sites using high-throughput sequencing technology, designed specific primer sets for detection, and constructed MAS breeding and GS models to achieve precise seedling prediction and targeted breeding.

Benefits of technology

It enables precise seed selection during the seedling stage, reduces breeding costs, improves the efficiency of genetic improvement of varieties, and provides full-chain technical support from variety cultivation to commercial circulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122012805A_ABST
    Figure CN122012805A_ABST
Patent Text Reader

Abstract

The invention discloses a group of molecular markers related to the oil yield of carya illinoensis and application of the molecular markers, and relates to the technical field of biology. The nucleotide sequence of the SNP molecular marker is as shown in SEQ ID NO. 1-2. The molecular marker disclosed by the invention can be used for marker assisted selection (MAS) and genome selection (GS) of the carya illinoensis, so that the genetic improvement progress of the variety of the carya illinoensis can be accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to a molecular marker related to the oil yield of thin-shelled pecans and its application. Background Technology

[0002] Pecan (Carya illinoensis), a world-renowned nut and oilseed tree species, boasts a kernel oil content of 51%–69%. The oil is primarily composed of unsaturated fatty acids such as oleic acid and linoleic acid, with a total unsaturated fatty acid content exceeding 90%, making it a significant source of high-quality edible oil and possessing substantial economic value. Oil yield, a key economic trait of pecans, directly determines its oilseed development value and market competitiveness. Varieties with high oil yields dominate oil processing and enhance industrial economic benefits, representing a crucial requirement for high-quality industrial development.

[0003] However, there are prominent technical bottlenecks in the current field of thin-shelled pecan oil yield assessment and oil breeding: On the one hand, traditional oil yield assessment relies on physicochemical testing after fruit maturity, which can only be carried out after the plant has completed its full growth cycle to the fruiting stage. The genotype related to oil yield cannot be determined by phenotype during the seedling stage, which leads to the need to cultivate a large number of plants simultaneously during breeding and then select target individuals after fruiting, resulting in a serious waste of land, manpower and other resources; on the other hand, traditional testing methods are cumbersome, time-consuming and labor-intensive, costly, and easily affected by factors such as planting environment and maturity, resulting in poor result stability.

[0004] At the breeding level, traditional breeding relies on phenotypic selection. Breeding for high oil yield requires multiple generations of screening and verification, which is lengthy, inefficient, and fails to meet the industry's urgent need for high-oil varieties. Modern technologies such as marker-assisted breeding (MAS) and genome selection (GS) have become important means of crop genetic improvement. SNP molecular markers have significant advantages due to their wide distribution, high polymorphism, and accurate detection. Existing research has developed SNP markers for traits such as phenols and flowering type in thin-shelled pecans, and has also measured the oil content and fatty acid composition of some varieties. However, no specific SNP molecular markers closely related to oil yield have been found, and there is a lack of precise molecular markers and supporting detection systems that can be directly applied to breeding practices. Existing molecular markers cannot meet the needs of oil yield-oriented breeding, and GS models, lacking core trait marker support, have insufficient predictive accuracy, thus hindering the genetic improvement process of oil-producing thin-shelled pecans.

[0005] Furthermore, the lack of rapid and accurate oil yield assessment technologies in various industrial scenarios, such as breeding and seed selection, seedling identification, and resource evaluation, leads to low efficiency in selecting high-quality seedlings and affects the quality of industrial development. Therefore, developing SNP molecular markers related to the oil yield of thin-shelled pecans, constructing an efficient detection system, and applying them to marker-assisted selection (MAS) breeding and genomic selection (GS) model construction to achieve accurate seedling prediction and targeted breeding is of great significance for reducing breeding costs, accelerating the cultivation of high-oil varieties, and promoting high-quality industrial development. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention has developed a set of molecular markers for evaluating the oil yield of thin-shelled pecans. Specifically, SNP molecular markers for evaluating the oil yield of thin-shelled pecans were developed using high-throughput sequencing technology and validated in genetic and natural populations. These markers can be used for the commercial identification of commercially available thin-shelled pecan seedlings. Furthermore, their application in MAS breeding and GS model construction can accelerate the genetic improvement of thin-shelled pecan varieties.

[0007] The primary objective of this invention is to identify molecular markers associated with the oil yield trait of thin-shelled pecans. These markers are located on the pecan genome (sequencing and assembled by our research group), and the specific SNP information is as follows:

[0008] Table 1 SNP locus information

[0009]

[0010] The sequence of the molecular marker SNP1 is shown in SEQ ID NO.1. The polymorphic site of SNP1 is located at position 301 of the sequence in SEQ ID NO.1, and the mutation type is G>A, where AA or GA represents the genotype with a high oil yield. The sequence of the molecular marker SNP2 is shown in SEQ ID NO.2. The polymorphic site of SNP2 is located at position 301 of SEQ ID NO.2, and the mutation type is C>T, where CC or CT represents the genotype with a high oil yield. Each sequence is a single copy within the genome, and the specific sequences are as follows:

[0011] SEQ ID NO.1

[0012] AAAACGTGGACCTACAAACCAACACATGCTCACATGCAGAAAGCATTCGGACTCAATCATCAATCAGACTATGCATTGCAAAAAATAGATGTCCTATATTTCTTCGACCACTGCATTCACAGGGTCATCCTTATAATTTTTCTTATTTCGGCTACCTCTTTTGATGTGGCCTAGCTTGCCACAATTCCAGCAAGTAATCTACTATCCAGATCTTGACTTGCTCCTGCCTCTGTTCCTAGACCTATCATGTCCCATGCCTCTATTGTGAATGTTCAAGGTAGAACCCAAACCTGAGAACTC[G\A]CCAGAATCTCTTCTGCACACTTCCTTAATAAGAATTAAATCACGAATGTCATTGTATTTTAATTTCAGCTTACCAGCAAAATTACTTACCTCCATTCTCACAGCCTCCCAACTATTCGGCAAAGATGCCAACATAATCAATGCTCTAAGCTCATCATCAAACTCATATTCAACAAATGATGATTGATTTGTGATAGTGTTAAATTCATTCAAATGTTAGGCAACAAATGTACCTTCAATAATTCTTATATTGGATAGTTTTTTCATCAAATGCACATTGTTATTCACATACGGATTTTCA。

[0013] SEQ ID NO.2

[0014] ATGAGAATTTTAGGGTTTTGAGTTTATATACTCAAATAGGCTAGTGGTTAGGCTATCCACTTGCTAATATTTGAGCTTTATATATTGGAAGGGATAAAATATACCAAGTACATAGTGTTTTCCAAGCCCAGCCTAAGGTGTTAAGTCTTAT CAGGAGGGAAGAAACAGAGAGGAGGGGAAAAAATGTTGAGGAAGGATAAAGTGTCAGGAAGAAAATGGGAGATGTGACATAAATGAGAGGAGATGAGACATAAATGAACGAGGAACAAAAGATTAAAGGGAGGGAACCAGACGACTT[C\ T]AGGGACTTTTGTCCAGACTTCTCTTGAAGGGATTTGAGGAGACGACTTCTTGGACTTCCACTTTGGCTTCTTCAATCTCACAATACGAGTTCCAGCAGTGATATCTCTACCAGAAGATAAATTTTCAATATCAATTGTGTAGTGAGGAC AAGAGATTATCAGAAACTGTAAAACAAGCATATTATAATGGATTCTCCCTCTCAAACGACTTGTGGATATAGGCATTACGCTGAACCAAATAAATTTGTATGTCTCCCTCCTTTACTTTTCCGGTGTTTATTTCCTGTTTACTCCAATGG.

[0015] To achieve efficient and accurate detection of the above molecular markers, this invention provides a set of specific primers, which includes two primer pairs. Primer pair 1 is used to detect SNP1, and its nucleotide sequences are shown in SEQ ID NO.3 and SEQ ID NO.4, respectively. Primer pair 2 is used to detect SNP2, and its nucleotide sequences are shown in SEQ ID NO.5 and SEQ ID NO.6, respectively.

[0016] SEQ ID NO.3

[0017] TTGACTTGCTCCTGCCTCTG;

[0018] SEQ ID NO.4

[0019] GGGAGGCTGTGAGAATGGAG;

[0020] SEQ ID NO.5

[0021] GCCCAGCCTAAGGTGTTAAGT;

[0022] SEQ ID NO.6

[0023] AGTCCAAGAAGTCGTCTCCTC;

[0024] The primers mentioned above can be used for first-generation or second-generation sequencing.

[0025] The present invention also provides a kit for detecting the above-mentioned molecular markers, the core component of which is the above-mentioned specific primer set.

[0026] This invention also provides a method for genetically improving the oil yield of thin-shelled pecans, the method comprising the following steps:

[0027] 1) Determine the genotypes of molecular marker sites related to the oil yield of thin-shelled pecans in the pecan resource population, wherein the molecular marker sites are the aforementioned SNP sites.

[0028] 2) Make appropriate selections based on the genotype of the molecular marker and the breeding objectives.

[0029] Step 1) includes the following steps:

[0030] 1.1) Extract genomic DNA from the thin-shelled pecans to be tested;

[0031] 1.2) Genotyping of the thin-shelled pecans to be tested;

[0032] 1.3) Based on the detection results, determine the genotypes of the thin-shelled pecans to be tested at SNP1 and SNP2 loci on chromosome 4 of the reference genome;

[0033] Step 2) includes at least one of the following steps:

[0034] 2.1) Eliminate individuals with the GG type at the SNP1 locus of thin-shelled pecans in the aforementioned thin-shelled pecan resource population, so as to increase the frequency of the AA and / or GA genotypes at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0035] 2.2) In the thin-shelled pecan resource population, individuals with the TT type at the SNP2 locus of thin-shelled pecans were eliminated to increase the frequency of the CC or CT genotype at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0036] Preferably, in step 2.1), individuals with the GA and GG types at the SNP1 locus of thin-shelled pecans are eliminated to increase the frequency of the AA genotype at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0037] Preferably, in step 2.2), individuals with the CT and TT genotypes at the SNP2 locus of thin-shelled pecans are eliminated to increase the frequency of the CC genotype at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0038] Preferably, in order to perform genetic probability breeding more accurately, step 2) includes at least one of the following steps:

[0039] 2.1) In the thin-shelled pecan resource population, individuals with SNP2 and SNP1 combination genotypes of CC_AA, CC_GG, CT_AA, CT_GA and / or CC_GA were retained to improve the oil yield of offspring thin-shelled pecans.

[0040] Preferably, in order to perform genetic probability breeding more accurately, step 2) is as follows:

[0041] In the thin-shelled pecan resource population, individuals with the genotype CC_AA for the SNP2 and SNP1 combination on chromosome 4 of the thin-shelled pecan were retained, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0042] The accuracy of the results was verified to be high through population testing, providing a reliable basis for seedling transactions.

[0043] This invention further provides a method for evaluating the oil yield of thin-shelled pecans, wherein the oil yield is the oil yield of the pecan kernels, and the method includes the following steps:

[0044] 1) Determine the genotypes of molecular marker sites related to the oil yield of thin-shelled pecans in the pecan resource population, wherein the molecular marker sites are the aforementioned SNP sites.

[0045] 2) Determine the oil yield of thin-shelled pecan samples based on the genotype of the molecular markers.

[0046] Step 1) includes the following steps:

[0047] 1.1) Extract genomic DNA from the thin-shelled pecans to be tested;

[0048] 1.2) Genotyping of the thin-shelled pecans to be tested;

[0049] 1.3) Based on the detection results, determine the genotypes of the thin-shelled pecans to be tested at SNP1 and SNP2 loci on chromosome 4 of the reference genome;

[0050] Step 2) determining the oil yield of the thin-shelled pecan sample based on the genotype of the molecular marker includes at least one of the following steps:

[0051] 2.1) In the thin-shelled pecan population, if the genotype of the thin-shelled pecan sample at the SNP1 locus is AA or GA, then the sample is determined to be a sample with a high oil yield.

[0052] 2.2) In the thin-shelled pecan population, if the genotype of the thin-shelled pecan sample at the SNP2 locus is CC or CT, then the sample is determined to be a sample with a high oil yield.

[0053] Preferably, in step 2.1), if the genotype of the thin-shelled pecan sample at the SNP1 locus is AA, then the sample is determined to be a sample with a high oil yield.

[0054] Preferably, in step 2.2), if the genotype of the thin-shelled pecan sample at the SNP2 locus is CC, then the sample is determined to be a sample with a high oil yield.

[0055] Preferably, in order to more accurately assess the oil yield, step 2) includes at least one of the following steps:

[0056] 2.1) In the thin-shelled pecan population, if the combined genotype of the thin-shelled pecan sample at SNP2 and SNP1 loci is CC_AA, CC_GG, CT_AA, CT_GA and / or CC_GA, then the sample is determined to be a sample with a high oil yield.

[0057] Preferably, in order to more accurately assess the oil yield, step 2.1) is as follows:

[0058] In a population of thin-shelled pecans, if the combined genotype of the thin-shelled pecan sample at SNP2 and SNP1 loci is CC_AA, then the sample is determined to be a sample with a high oil yield.

[0059] The beneficial effects of this invention are as follows: The aforementioned molecular markers have significant application value in breeding related to the oil yield trait of thin-shelled pecans or in predicting the oil yield of thin-shelled pecans. Applying them to MAS breeding enables precise seedling selection, eliminating individuals with non-target genotypes, and reducing breeding costs. Integrating them into the GS model can improve the predictive accuracy of multi-trait aggregation breeding and accelerate the process of variety genetic improvement. Simultaneously, the primer set or kit provided by this invention also has significant application advantages in breeding related to the oil yield trait of thin-shelled pecans or in predicting the oil yield of thin-shelled pecans. It exhibits high specificity and stable detection results, and can be widely applied to early screening in breeding units, oil yield identification in seedling enterprises, and resource evaluation in research institutions, providing full-chain technical support for the thin-shelled pecan industry from variety cultivation to commercial circulation. Attached Figure Description

[0060] The method of the present invention and its beneficial effects will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] Figure 1 These are the statistical results of oil yield from different genotypes at the SNP1 locus. Different lowercase letters indicate significant differences, and different uppercase letters indicate extremely significant differences.

[0062] Figure 2 These are the statistical results of oil yield from different genotypes at the SNP2 locus. Different capital letters indicate highly significant differences.

[0063] Figure 3 This is a heatmap of the significance test of oil yield differences between different genotype combinations. The values ​​represent the p-values ​​of the significance test between genotypes on the horizontal and vertical axes, and the values ​​are rounded to two decimal places. Detailed Implementation

[0064] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0065] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0066] Example

[0067] Example 1: Development of SNPs and QTL Localization

[0068] 1. Experimental materials

[0069] 112 thin-shelled pecan germplasm resource samples were collected.

[0070] 2. Sample DNA extraction, library construction, and sequencing

[0071] First, the leaves of each sample were preserved using liquid nitrogen. Genomic DNA was extracted from samples 1 to 112 using a kit. A total of 994 Gb of raw data was obtained from sequencing, and the sequencing results were 150 bp of paired-end data.

[0072] Specific experimental steps:

[0073] Library construction was initiated with 1 μg of DNA;

[0074] High-quality genomic DNA was extracted using the CTAB method;

[0075] 0.75% agarose gel electrophoresis was used to detect DNA fragment size and the degree of DNA degradation.

[0076] The NanoDrop One spectrophotometer (Thermo Fisher Scientific) was used to detect DNA purity, with an OD260 / 280 ratio between 1.8 and 2.2, indicating no protein or visible contaminants.

[0077] The Qubit 3.0 fluorescence analyzer (Life Technologies, Carlsbad, CA, USA) was used to detect DNA concentrations greater than 50 ng / μl and total amounts greater than 2 μg.

[0078] After DNA was broken down by sonication with a Covaris M220, magnetic beads were used for fragment selection, resulting in sample bands concentrated between 200-400 bp.

[0079] The qualified libraries were then put into the sequencing machine for 2*150bp sequencing.

[0080] 3. Data quality control

[0081] (1) Remove the adapter sequence from the sequence;

[0082] (2) Remove polyG and polyX from the end of the reads (minimum length is 10bp);

[0083] (3) The average quality value of the bases in the window is calculated by using a sliding window method. Low-quality sliding windows are clipped out. Its function is similar to Trimmomatic.

[0084] (4) Remove N reads that are greater than 5;

[0085] (5) Remove reads with a base content of less than 15 that is higher than 40%;

[0086] (6) Remove reads with a length of less than 15bp after filtering.

[0087] 4. Data comparison and SNP development

[0088] In this invention, we use the genome of thin-shelled pecan as a reference genome, use BWA alignment software to align the sequencing fragments back to the reference genome, and then use Picard-tools to remove the sequencing fragments generated by PCR-duplication.

[0089] For each sample, a bwa (version: 0.7.17; parameter: mem) alignment analysis was performed. The filtered clean reads were aligned to the reference genome, and the alignment results were statistically analyzed. The specific analysis steps are as follows:

[0090] (1) Use bwa alignment software (parameter: mem -R, other parameters are the software default parameters) to align the clean reads of all samples with the reference genome;

[0091] (2) Use samtools (parameter: sort) to convert the alignment results from the sam (Sequence Alignment / MAP) file to a sorted bam (binary Alignment / Map) file;

[0092] (3) Use samtools (parameter: markdup -r) to remove duplicates from the sorted alignment results for subsequent analysis;

[0093] (4) Develop SNPs using GATK software.

[0094] 5. Genome-wide association analysis

[0095] The SNPs obtained in step 4 above were used to perform GWAS correlation analysis with the oil yield of each sample's kernels at maturity (i.e., the percentage of crude fat in the total dry weight of the dried thin-shelled pecan kernels; the method for detecting kernel oil yield refers to the standard GB-5009.6-2016). The GWAS analysis in this invention used the lm and lmm models of GEMMA software, the GLM and MLM models of rMVP software, and the FarmCPU model of rMVP software.

[0096] Ultimately, two loci on chromosome 4 were found to have a significant impact on oil yield, as shown in Table 2:

[0097] Table 2. Associated SNP sites

[0098]

[0099] Note: Columns 4-8 in Table 2 are the significance P-values ​​of the association results between SNP sites and oil yield traits of different models.

[0100] The molecular marker SNP1 G>A is located at position 301 of the sequence in SEQ ID NO.1, and the molecular marker SNP2 C>T is located at position 301 of SEQ ID NO.2. The sequences are single-copy sequences within the genome.

[0101] SEQ ID NO.1

[0102] AAAACGTGGACCTACAAACCAACACATGCTCACATGCAGAAAGCATTCGGACTCAATCATCAATCAGACTATGCATTGCAAAAAATAGATGTCCTATATTTCTTCGACCACTGCATTCACAGGGTCATCCTTATAATTTTTCTTATTTCGGCTACCTCTTTTGATGTGGCCTAGCTTGCCACAATTCCAGCAAGTAATCTACTATCCAGATCTTGACTTGCTCCTGCCTCTGTTCCTAGACCTATCATGTCCCATGCCTCTATTGTGAATGTTCAAGGTAGAACCCAAACCTGAGAACTC[G\A]CCAGAATCTCTTCTGCACACTTCCTTAATAAGAATTAAATCACGAATGTCATTGTATTTTAATTTCAGCTTACCAGCAAAATTACTTACCTCCATTCTCACAGCCTCCCAACTATTCGGCAAAGATGCCAACATAATCAATGCTCTAAGCTCATCATCAAACTCATATTCAACAAATGATGATTGATTTGTGATAGTGTTAAATTCATTCAAATGTTAGGCAACAAATGTACCTTCAATAATTCTTATATTGGATAGTTTTTTCATCAAATGCACATTGTTATTCACATACGGATTTTCA。

[0103] SEQ ID NO.2

[0104] ATGAGAATTTTAGGGTTTTGAGTTTATATACTCAAATAGGCTAGTGGTTAGGCTATCCACTTGCTAATATTTGAGCTTTATATATTGGAAGGGATAAAATATACCAAGTACATAGTGTTTTCCAAGCCCAGCCTAAGGTGTTAAGTCTTAT CAGGAGGGAAGAAACAGAGAGGAGGGGAAAAAATGTTGAGGAAGGATAAAGTGTCAGGAAGAAAATGGGAGATGTGACATAAATGAGAGGAGATGAGACATAAATGAACGAGGAACAAAAGATTAAAGGGAGGGAACCAGACGACTT[C\ T]AGGGACTTTTGTCCAGACTTCTCTTGAAGGGATTTGAGGAGACGACTTCTTGGACTTCCACTTTGGCTTCTTCAATCTCACAATACGAGTTCCAGCAGTGATATCTCTACCAGAAGATAAATTTTCAATATCAATTGTGTAGTGAGGAC AAGAGATTATCAGAAACTGTAAAACAAGCATATTATAATGGATTCTCCCTCTCAAACGACTTGTGGATATAGGCATTACGCTGAACCAAATAAATTTGTATGTCTCCCTCCTTTACTTTTCCGGTGTTTATTTCCTGTTTACTCCAATGG.

[0105] Example 2: Correlation analysis between phenotype and genotype in a genetic population

[0106] Two pairs of primers were designed and used for PCR amplification to construct a high-throughput sequencing library. The primer sequences are shown in SEQ ID NO.3~SEQ ID NO.6.

[0107] SEQ ID NO.3

[0108] TTGACTTGCTCCTGCCTCTG

[0109] SEQ ID NO.4

[0110] GGGAGGCTGTGAGAATGGAG

[0111] SEQ ID NO.5

[0112] GCCCAGCCTAAGGTGTTAAGT

[0113] SEQ ID NO.6

[0114] AGTCCAAGAAGTCGTCTCCTC

[0115] Sequencing was performed on the two loci mentioned above in 112 thin-shelled pecan samples for verification. The oil yield of individuals with different genotypes (detection method as in Example 1) was statistically analyzed and subjected to Tukey's test. The results are shown in Tables 3-5. Figures 1-3 As shown:

[0116] Table 3 Mutation sites and oil yield characteristics

[0117]

[0118] Note: Different lowercase letters indicate significant differences between groups (P≤0.05), different uppercase letters indicate highly significant differences between groups (P≤0.01), and the same uppercase or lowercase letters indicate no highly significant or no significant differences; the statistical results exclude individuals whose genotypes were not detected, and the oil yield is the average value.

[0119] Depend on Figures 1-2 As shown in Table 3, the three genotypes at SNP1 showed significant differences, with oil yields from highest to lowest being AA, GA, and GG genotypes. At SNP2, the TT genotype showed highly significant differences from the other two genotypes, with oil yields from highest to lowest being CC, CT, and TT genotypes.

[0120] Further statistical analysis and Tukey's test were performed on the oil yield of different genotype combinations at the three loci. The results are shown in Tables 4 and 5. Figure 3 As shown:

[0121] Table 4. Oil yield of different genotype combinations

[0122]

[0123] Table 5. Significance test of oil yield differences among different genotype combinations.

[0124]

[0125] Note: * indicates p <= 0.05, ** indicates p <= 0.01, *** indicates <= 0.001, \ indicates that there is no statistically significant difference between the two genotype groups, and ns indicates that there is no significant difference between the two groups. The oil yield is the average value.

[0126] From Tables 4-5 and Figure 3 It can be seen that the oil yield of TT_GG is significantly lower than that of other genotype combinations, indicating that the loci associated by the GWAS analysis in Example 1 are more accurate.

[0127] Example 3: Method for genetic improvement of oil yield in thin-shelled pecans

[0128] A method for genetically improving the oil yield of thin-shelled pecans, the method comprising the following steps:

[0129] 1) Determine the genotypes of molecular marker sites related to the oil yield of thin-shelled pecans in the pecan resource population, wherein the molecular marker sites are the sites in the aforementioned embodiments.

[0130] 2) Make appropriate selections based on the genotype of the molecular marker and the breeding objectives.

[0131] Step 1) includes the following steps:

[0132] 1.1) Extract genomic DNA from the thin-shelled pecans to be tested;

[0133] 1.2) Genotyping of the thin-shelled pecans to be tested;

[0134] 1.3) Based on the detection results, determine the genotypes of the thin-shelled pecans to be tested at SNP1 and SNP2 loci on chromosome 4 of the reference genome;

[0135] Step 2) includes at least one of the following steps:

[0136] 2.1) In the thin-shelled pecan resource population, individuals with the GG type at the SNP1 locus of thin-shelled pecans were eliminated to increase the frequency of the AA and / or GA genotypes at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0137] 2.2) In the thin-shelled pecan resource population, individuals with the TT type at the SNP2 locus of thin-shelled pecans were eliminated to increase the frequency of the CC or CT genotype at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0138] Preferably, in step 2.1), individuals with the GA and GG types at the SNP1 locus of thin-shelled pecans are eliminated to increase the frequency of the AA genotype at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0139] Preferably, in step 2.2), individuals with the CT and TT genotypes at the SNP2 locus of thin-shelled pecans are eliminated to increase the frequency of the CC genotype at this locus generation by generation, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0140] Preferably, in order to perform genetic probability breeding more accurately, step 2) includes at least one of the following steps:

[0141] 2.1) In the thin-shelled pecan resource population, individuals with SNP2 and SNP1 combination genotypes of CC_AA, CC_GG, CT_AA, CT_GA and / or CC_GA were retained to improve the oil yield of offspring thin-shelled pecans.

[0142] Preferably, in order to perform genetic probability breeding more accurately, step 2) is as follows:

[0143] In the thin-shelled pecan resource population, individuals with the genotype CC_AA for the SNP2 and SNP1 combination on chromosome 4 of the thin-shelled pecan were retained, thereby increasing the oil yield of the offspring thin-shelled pecans.

[0144] Example 4: Method for evaluating the oil yield of thin-shelled pecans

[0145] A method for evaluating the oil yield of thin-shelled pecans, wherein the oil yield is the oil yield of the pecan kernels, the method comprising the following steps:

[0146] 1) Determine the genotypes of molecular marker sites related to the oil yield of thin-shelled pecans in the pecan resource population, wherein the molecular marker sites are the sites in the aforementioned embodiments.

[0147] 2) Determine the oil yield of thin-shelled pecan samples based on the genotype of the molecular markers.

[0148] Step 1) includes the following steps:

[0149] 1.1) Extract genomic DNA from the thin-shelled pecans to be tested;

[0150] 1.2) Genotyping of the thin-shelled pecans to be tested;

[0151] 1.3) Based on the detection results, determine the genotypes of the thin-shelled pecans to be tested at SNP1 and SNP2 loci on chromosome 4 of the reference genome;

[0152] Step 2) determining the oil yield of the thin-shelled pecan sample based on the genotype of the molecular marker includes at least one of the following steps:

[0153] 2.1) In the thin-shelled pecan population, if the genotype of the thin-shelled pecan sample at the SNP1 locus is AA or GA, then the sample is determined to be a sample with a high oil yield.

[0154] 2.2) In the thin-shelled pecan population, if the genotype of the thin-shelled pecan sample at the SNP2 locus is CC or CT, then the sample is determined to be a sample with a high oil yield.

[0155] Preferably, in step 2.1), if the genotype of the thin-shelled pecan sample at the SNP1 locus is AA, then the sample is determined to be a sample with a high oil yield.

[0156] Preferably, in step 2.2), if the genotype of the thin-shelled pecan sample at the SNP2 locus is CC, then the sample is determined to be a sample with a high oil yield.

[0157] Preferably, in order to more accurately assess the oil yield, step 2) includes at least one of the following steps:

[0158] 2.1) In the thin-shelled pecan population, if the combined genotype of the thin-shelled pecan sample at SNP2 and SNP1 loci is CC_AA, CC_GG, CT_AA, CT_GA and / or CC_GA, then the sample is determined to be a sample with a high oil yield.

[0159] Preferably, in order to more accurately assess the oil yield, step 2.1) is as follows:

[0160] In a population of thin-shelled pecans, if the combined genotype of the thin-shelled pecan sample at SNP2 and SNP1 loci is CC_AA, then the sample is determined to be a sample with a high oil yield.

[0161] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to the above embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A molecular marker sequence associated with the oil yield of thin-shelled pecans, characterized in that, The molecular marker sequence includes at least one of SEQ ID NO.1 to SEQ ID NO.2, wherein the polymorphic site of SEQ ID NO.1 corresponds to the G>A mutation at position 301 of SEQ ID NO.1, where AA is the genotype with high oil yield and GG is the genotype with low oil yield; the polymorphic site of SEQ ID NO.2 corresponds to the C>T mutation at position 301 of SEQ ID NO.2, where CC and / or CT are the genotypes with high oil yield and TT is the genotype with low oil yield.

2. A primer set for detecting molecular markers, characterized in that, The nucleotide sequences of the primer set are shown in SEQ ID NO.3 to SEQ ID NO.6, wherein SEQ ID NO.3 to SEQ ID NO.4 are used to detect molecular marker SNP1, and SEQ ID NO.5 to SEQ ID NO.6 are used to detect molecular marker SNP2. Molecular marker SNP1 corresponds to the G>A mutation at position 301 of SEQ ID NO.1, and its nucleotide sequence is shown in SEQ ID NO.1, wherein AA is the genotype with high oil yield and GG is the genotype with low oil yield; SNP2 corresponds to the C>T mutation at position 301 of SEQ ID NO.2, and its nucleotide sequence is shown in SEQ ID NO.2, wherein CC and / or CT are the genotype with high oil yield and TT is the genotype with low oil yield.

3. A kit for detecting molecular markers, characterized in that, The kit comprises the primer set as described in claim 2.

4. A method for evaluating the oil yield of thin-shelled pecans, characterized in that, The method includes the following steps: (1) Extract genomic DNA from the thin-shelled pecan samples to be tested; (2) The genotype of the thin-shelled pecan to be tested is detected using the primer set described in claim 2 or the kit described in claim 3; (3) Based on the test results, determine the genotype of the thin-shelled pecan to be tested; (4) The oil yield of thin-shelled pecans is evaluated based on the genotype results obtained from the test. The evaluation criteria are that individuals with the genotype at any of the following loci are individuals with high oil yield: individuals with the genotype AA at SNP1 locus and individuals with the genotype CC or CT at SNP2 locus. SNP1 corresponds to the G>A mutation at position 301 of SEQ ID NO.1, and its nucleotide sequence is shown in SEQ ID NO.

1. SNP2 corresponds to the C>T mutation at position 301 of SEQ ID NO.2, and its nucleotide sequence is shown in SEQ ID NO.

2.

5. A genetic breeding method for improving the oil yield trait of thin-shelled pecans, characterized in that, The method includes the following steps: (1) Extract genomic DNA from the thin-shelled pecan resource population to be tested; (2) The genotype of the thin-shelled pecan to be tested is detected using the primer set described in claim 2 or the kit described in claim 3; (3) Based on the test results, determine the genotype of the thin-shelled pecan to be tested; (4) Select individuals with different genotypes as parents according to the breeding goal. When the breeding goal is to breed varieties with high oil yield, select individuals with the following genotypes as parents: Individuals with SNP1 of type AA and / or SNP2 of type CC or CT, wherein SNP1 corresponds to the G>A mutation at position 301 of SEQ ID NO.1, and its nucleotide sequence is shown in SEQ ID NO.1; and SNP2 corresponds to the C>T mutation at position 301 of SEQ ID NO.2, and its nucleotide sequence is shown in SEQ ID NO.

2.

6. The application of the primer set of claim 2 or the kit of claim 3 in breeding related to the oil yield of thin-shelled pecans or in predicting the oil yield of thin-shelled pecans.