A soybean oil component gene-based positioning method based on residual shrinkage network
By constructing mixed pools of high-oil and low-oil soybeans and combining multi-source data analysis, deep residual shrinkage networks were used to screen soybean oil genes, solving the problems of low efficiency and insufficient accuracy in existing technologies, and enabling the rapid and accurate acquisition of core genes that control soybean oil content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-04-07
AI Technical Summary
Existing methods for screening and identifying key genes for soybean oil content are inefficient and imprecise, making it difficult to improve soybean oil yield and processing efficiency.
A soybean oil content gene localization method based on residual shrinkage network was adopted. By constructing high-oil and low-oil mixed pools, and combining soil physicochemical data, microclimate data and transcriptome sequence data, gene feature extraction and screening were performed using deep residual shrinkage network to obtain the core genes controlling soybean oil content.
This technology enables the rapid and accurate acquisition of core genes that control soybean oil content, improving the reliability and screening efficiency of candidate genes.
Smart Images

Figure CN121438940B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of agricultural biotechnology, and relates to a soybean oil content gene positioning method based on a residual shrinkage network. BACKGROUND
[0002] Soybean is an important oil crop and plant protein source, and its oil content is a key trait determining quality and economic value. The higher the oil content, the more beneficial it is to improve oil yield and processing efficiency. In order to effectively improve the oil content of soybean, it is necessary to screen and identify key genes that regulate oil content, so as to cultivate new soybean varieties with high oil and high quality by using the key genes.
[0003] In related technologies, a large number of seeds are measured for oil content to select excellent plants and breed offspring through traditional cross breeding, and excellent seeds are further selected according to the oil content of the offspring. However, this method has problems such as long cycle, low efficiency and insufficient precision.
[0004] Therefore, how to quickly and accurately obtain the core gene controlling the oil content of soybean has become a problem to be solved. SUMMARY
[0005] Therefore, the embodiments of the present application provide a soybean oil content gene positioning method based on a residual shrinkage network, which at least solves the problem that related technologies cannot quickly and accurately obtain the core gene controlling the oil content of soybean.
[0006] According to a first aspect of the embodiments of the present application, a soybean oil content gene positioning method based on a residual shrinkage network is provided, which comprises:
[0007] After a plurality of seeds randomly obtained from each F1 generation plant after maturation are sowed in the same hole in a grid of a test field, a random initial F2 generation plant, corresponding soil physical and chemical data and microclimate data are obtained; and based on the sorting result of the initial oil content of the seeds harvested from the initial F2 generation plant, the remaining F2 generation plants are obtained;
[0008] Based on the remaining F2 generation plants, a high-oil pool and a low-oil pool are constructed; and the original sequencing data of the high-oil pool, the low-oil pool and the parent are obtained respectively;
[0009] Based on the original sequencing data and the soybean reference genome, alignment is performed to obtain target single nucleotide polymorphism sites that are consistent with the genotypes of the corresponding oil content parents and inconsistent between the parents in the high-oil pool and the low-oil pool, and a first candidate gene is obtained based on the soil physical and chemical data, the microclimate data and the target single nucleotide polymorphism sites;
[0010] Expression characteristics were obtained using transcriptome sequence data from the parents at different developmental stages; and genotypic and genetic effect characteristics of the first candidate gene at the target single nucleotide polymorphism site were extracted.
[0011] Target features are obtained based on genotype characteristics, expression level characteristics, genetic effect characteristics, and deep residual shrinkage networks;
[0012] The first candidate gene was screened using target features to obtain the second candidate gene; and the third candidate gene enriched in the lipid synthesis pathway was obtained based on the second candidate gene and the functional database.
[0013] The probability values of the second candidate gene and the total genes of the lipid synthesis pathway were obtained, and the core genes controlling the soybean oil content were obtained based on the probability values and the third candidate gene.
[0014] According to a second aspect of the present invention, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method described in the first aspect.
[0015] According to a third aspect of the present invention, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0016] According to the scheme provided in the embodiments of the present invention, multiple seeds randomly obtained from each mature F1 generation plant are sown in the same hole in a grid of the experimental field to obtain a random initial F2 generation plant, corresponding soil physicochemical data, and microclimate data; and the remaining F2 generation plants are obtained by sorting the seeds harvested from the initial F2 generation plants based on their initial oil content; high-oil mixing pools and low-oil mixing pools are constructed based on the remaining F2 generation plants; and the original sequencing data of the high-oil mixing pool, low-oil mixing pool, and parent plants are obtained respectively; the original sequencing data are compared with the soybean reference genome to obtain the target single nucleotide polymorphism (SNP) sites that are consistent with the corresponding oil content parent plants in the high-oil mixing pool and low-oil mixing pool, and where the genotypes of the parents are inconsistent; and the first candidate gene is obtained based on the soil physicochemical data, microclimate data, and the target SNP site; the expression level characteristics are obtained using the transcriptome sequence data of the parents at different developmental stages; and the genotype characteristics and genetic effect characteristics of the first candidate gene at the target SNP site are extracted.
[0017] Target features are obtained based on genotype characteristics, expression level characteristics, genetic effect characteristics, and a deep residual shrinking network. These target features are used to screen first candidate genes, resulting in second candidate genes. Third candidate genes enriched in the lipid synthesis pathway are then obtained based on the second candidate genes and a functional database. Probability values are obtained based on the second candidate genes and the total number of genes in the lipid synthesis pathway, and the core genes controlling soybean oil content are then identified based on these probability values and the third candidate genes. In this process, inputting three types of features into the deep residual shrinking network not only fully explores the complex correlations between multi-source data but also achieves interpretable screening of key features. Screening starts from the first candidate gene and continues until the core gene is obtained, greatly improving the reliability of the candidate genes. In summary, this method can quickly and accurately identify the core genes controlling soybean oil content. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:
[0019] Figure 1 A flowchart illustrating a method for locating soybean oil genes based on a residual shrinkage network, provided in an embodiment of the present invention;
[0020] Figure 2 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0023] It should be noted that the terms "first, second, and third" used in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.
[0024] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which these embodiments of the invention pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0025] Figure 1 This is a flowchart illustrating a method for locating soybean oil genes based on residual shrinkage networks, as provided in an embodiment of the present invention. This method for locating soybean oil genes based on residual shrinkage networks can be executed by an electronic device, such as a computer or server.
[0026] like Figure 1 As shown, the method for locating soybean oil genes based on residual shrinkage networks includes:
[0027] S101. Multiple seeds randomly obtained from each mature F1 generation plant are sown in the same hole in the grid of the experimental field to obtain a random initial F2 generation plant, the corresponding soil physicochemical data and microclimate data; and the remaining F2 generation plants are obtained by sorting the seeds based on the initial oil content of the seeds harvested from the initial F2 generation plants.
[0028] In embodiments of the present invention, from a large number of soybean germplasm resources, based on known oil content or preliminary determination, a known high-oil variety and a known low-oil variety are pre-selected as parents. The low-oil variety with white flowers is used as the female parent, and the high-oil variety with purple flowers is used as the male parent. The hybridization process can involve manually emasculating the unopened flower buds of the female parent one day before flowering, collecting fresh pollen from the male parent within two hours of flowering the next day, pollinating the emasculated female parent, and labeling it to obtain multiple F0 generation seeds. The F0 generation seeds can be used for outdoor generation at a pre-designated seed propagation base. After the F0 generation seeds germinate and grow into plants, they are self-pollinated and their flower color is identified. Successfully hybridized purple-flowered F0 plants are retained. Multiple seeds (e.g., 2 to 3 seeds) are randomly selected from each mature F1 generation plant and sown in the same hole. After F1 generation seedlings emerge, only one seedling from the same hole is randomly retained to obtain F2 generation seeds. F2 generation seeds were sown in the same hole according to a 3m×3m grid. After single-seed propagation, one F2 generation plant (initial F2 generation plant) was obtained, and the harvested seeds were used as F3 generation seeds (seeds harvested from the initial F2 generation plant). Simultaneously, soil sensors were deployed on the grid to collect real-time soil physicochemical and microclimate data for the 0-20cm soil layer of the grid containing the initial F2 generation plant. The seeds harvested from the initial F2 generation plant were naturally air-dried until the moisture content met the preset requirements, then packed into kraft paper bags labeled with plant numbers and stored at room temperature. When using a near-infrared spectroscopy grain analyzer to determine the oil content of seeds harvested from initial F2 generation plants stored at room temperature in kraft bags, 50 seeds can be weighed each time for measurement, repeated 3 times, and the average value is taken as the initial oil content of the initial F2 generation plants. Then, the initial oil content is ranked, and based on the ranking, the initial F2 generation plants are removed. Finally, a predetermined number of plants with high and low oil content from the two extremes are selected as the remaining F2 plants. For example, 30 plants from each extreme oil content are selected as the remaining F2 plants.
[0029] Soil physicochemical data, including available nitrogen / phosphorus / potassium, organic matter, pH value, and volumetric water content, can be measured using various sensors deployed on a grid. Microclimate data, including air temperature and humidity, photosynthetically active radiation, precipitation, and wind speed recorded every 30 minutes during the growing season, can be obtained using small weather stations deployed at the center and corners of the field. After acquiring the soil physicochemical and microclimate data, data cleaning, data processing, and standardization can be performed.
[0030] After obtaining the initial F2 generation plants, 0.5g of seedling leaves were collected from the corresponding plant position of the initial F2 generation plants. After rinsing twice with deionized water and drying with filter paper, the leaves were quickly placed into 2mL centrifuge tubes and liquid nitrogen was added. The tubes were then flash-frozen at -196℃ for 10 minutes and subsequently transferred to an ultra-low temperature freezer at -80℃ for storage.
[0031] S102. Construct high-oil-mixed pools and low-oil-mixed pools based on the remaining F2 generation plants; and obtain the original sequencing data of the high-oil-mixed pools, low-oil-mixed pools, and parents, respectively.
[0032] In an embodiment of the present invention, high-oil-mixing pools and low-oil-mixing pools were constructed based on the genomic DNA of the remaining F2 generation plants. Finally, whole-genome resequencing technology was used to re-sequencing the genomic DNA of the high-oil-mixing pools, low-oil-mixing pools, and the parent plants, obtaining their respective raw sequencing data containing sequence and quality information. This data included deoxyribonucleic acid (DNA) and ribonucleic acid (RNA).
[0033] S103. Based on the comparison of the original sequencing data and the soybean reference genome, target single nucleotide polymorphism sites were obtained that were consistent with the corresponding parents in the high-oil-content pool and the low-oil-content pool, and the genotypes of the parents were inconsistent. The first candidate gene was obtained based on soil physicochemical data, microclimate data and target single nucleotide polymorphism sites.
[0034] In an embodiment of the present invention, the raw sequencing data of the high-oil mixed pool, the low-oil mixed pool, and the two parents were compared with the soybean reference genome to obtain the target single nucleotide polymorphism (SNP) site. Specifically, in the high-oil mixed pool, the genotype of the target SNP site was consistent with that of the high-oil paternal parent; in the low-oil mixed pool, the genotype of the target SNP site was consistent with that of the low-oil maternal parent; and there were genotypic differences between the high-oil paternal and low-oil maternal parents at the target SNP site itself. Then, by combining soil physicochemical data, microclimate data, and the target SNP site, stable and enriched genomic regions significantly associated with oil content in the high-oil and low-oil populations were located, and first candidate genes related to oil content were screened from these regions.
[0035] S104. Use transcriptome sequence data from the parents at different developmental stages to obtain expression level characteristics; and extract genotypic characteristics and genetic effect characteristics of the first candidate gene at the target single nucleotide polymorphism site.
[0036] In embodiments of the present invention, known high-oil and low-oil soybean varieties were used as materials. Unopened flower buds at the middle nodes of the plants were marked at the same node. Seeds were collected at different times after pollination of the unopened flower buds, such as 0, 10, 20, 30, 40, and 50 days. Three biological replicates were set for each period. A predetermined volume of soybean tissue could be collected after flowering, such as 0-20 days, and then all seeds from a predetermined number of pods could be collected starting from 30 days. The soybean tissue and all seeds were used as samples and flash-frozen in liquid nitrogen and stored at -80°C. Next, total RNA was extracted from the stored samples using the Trizol method. The extracted total RNA was purified and quality tested. After passing the test, a cDNA library was constructed using the extracted total RNA, and the cDNA library was subjected to next-generation sequencing to obtain transcriptome data. After a series of processing steps, including washing and alignment, the FPKM value (i.e., normalized expression level) of the gene was obtained and used as the expression level characteristic. Simultaneously, the specific genotypes at the target single nucleotide polymorphism sites within the first candidate gene region are extracted as key genotype features.
[0037] Among them, genetic effect characteristics are quantitative indicators or attributes of the influence of single nucleotide polymorphism sites on specific phenotypes, such as disease susceptibility, height, and metabolic level.
[0038] S105. Target features are obtained based on genotype characteristics, expression level characteristics, genetic effect characteristics, and deep residual shrinkage networks.
[0039] In embodiments of the present invention, genotype features, expression level features, and genetic effect features are concatenated to form fused features. These fused features are input into a deep residual shrinking network, and the attention mechanism within the deep residual shrinking network or a post-attribution method, such as SHAP, is used to calculate the contribution of each original feature dimension to the prediction of oil content traits. Based on a preset first screening threshold, features with a contribution higher than the first screening threshold are selected as target features related to oil content traits.
[0040] The core structure of the Deep Residual Shrinkage Network (DRSN) includes an input layer, a residual shrinkage module, residual connections, and fully connected layers. The residual shrinkage module contains two convolutional layers, a batch normalization layer, a ReLU activation function, an embedded soft thresholding function, and attention.
[0041] S106. Screening the first candidate gene using the target characteristics to obtain the second candidate gene; and obtaining the third candidate gene enriched in the lipid synthesis pathway based on the second candidate gene and the functional database.
[0042] In an embodiment of the present invention, the identifier of the target feature is matched with the identifiers of multiple first candidate genes to obtain first candidate genes that fail to match. The first candidate genes that fail to match are then deleted to obtain second candidate genes.
[0043] For example, the target feature is identified as Gene_XYZ_FPKM. If a first candidate gene is identified as Gene_XYZ, it proves that the first candidate gene Gene_XYZ is associated with the target feature Gene_XYZ_FPKM. Therefore, the first candidate gene Gene_XYZ is retained. Following this method, the first candidate gene is screened to obtain the second candidate gene. Next, the second candidate gene is compared with genes in a preset functional database to obtain the gene function corresponding to the second candidate gene. Based on the gene function, genes enriched in the lipid synthesis pathway are found in the second candidate gene and used as the third candidate gene.
[0044] S107. Based on the probability values of the second candidate gene and the total genes of the oil synthesis pathway, the core gene controlling the soybean oil content was obtained based on the probability values and the third candidate gene.
[0045] In embodiments of the present invention, the total annotated genes of the soybean reference genome are used as the background set. The overlap significance between the second candidate gene and the total genes of the lipid synthesis pathway is evaluated using the hypergeometric distribution test to obtain a probability value proving its significant enrichment. Further, a threshold, such as 0.05, is set. When the probability value is less than 0.05, the genetic location of the third candidate gene is compared with the reported quantitative trait loci (QTL) intervals related to soybean oil content to obtain a fourth candidate gene located within a significant QTL interval with reliable genetic evidence. The fourth candidate genes are sorted according to their corresponding contribution. Based on the sorting results and the threshold, a fifth candidate gene is obtained. Further, based on the functional annotation of the fifth candidate gene, a sixth candidate gene with a clear function that directly participates in biological processes such as lipid synthesis, lipid metabolism, or storage is selected. Finally, the sixth candidate gene is selected as the core gene controlling soybean oil content and supported by multidimensional evidence.
[0046] In cases where the probability value is not less than 0.05, or where the genetic location of the third candidate gene is compared with known quantitative trait loci (QTL) intervals related to soybean oil content and the comparison fails, all target features contained in the second candidate gene are obtained, and the sum or average of all target features is calculated as the contribution of the second candidate gene. Compared with the first screening threshold set in the specific process of S105, a larger screening threshold is set this time. This screening threshold and the contribution of the second candidate gene are used to screen for genes. The screened genes are then compared with the functional database to obtain genes that are clearly related to oil synthesis. These genes are then used as the core genes controlling soybean oil content.
[0047] It is understood that, in the embodiments of the present invention, multiple seeds randomly obtained from each mature F1 generation plant are sown in the same hole in a grid of the experimental field to obtain a random initial F2 generation plant, corresponding soil physicochemical data, and microclimate data; and the remaining F2 generation plants are obtained by sorting the seeds harvested from the initial F2 generation plants based on their initial oil content; high-oil mixing pools and low-oil mixing pools are constructed based on the remaining F2 generation plants; and the original sequencing data of the high-oil mixing pool, low-oil mixing pool, and parent plants are obtained respectively; based on the original sequencing data and the soybean reference genome, the target single nucleotide polymorphism sites that are consistent with the corresponding oil content parent plants in the high-oil mixing pool and low-oil mixing pool, and whose genotypes are inconsistent among the parents are obtained, are then identified. First candidate genes were obtained from soil physicochemical data, microclimate data, and target single nucleotide polymorphism (SNP) sites. Expression characteristics were obtained using transcriptome sequence data from parents at different developmental stages. Genotype and genetic effect characteristics of the first candidate genes at the target SNP sites were extracted. Target features were obtained based on genotype, expression, and genetic effect characteristics, using a deep residual shrinking network. The first candidate genes were then screened using these target features to obtain second candidate genes. Third candidate genes enriched in the lipid synthesis pathway were obtained based on the second candidate genes and a functional database. Probability values were obtained based on the second candidate genes and the total number of genes in the lipid synthesis pathway, and the core gene controlling soybean oil content was identified based on these probability values and the third candidate gene. In this process, inputting three types of features into a deep residual shrinking network not only fully explored the complex relationships between multi-source data but also enabled interpretable screening of key features. Screening from the first candidate gene to the core gene significantly improved the reliability of the candidate genes. In summary, this method can rapidly and accurately identify the core gene controlling soybean oil content.
[0048] In some embodiments of the present invention, the construction of high-oil-mixing pools and low-oil-mixing pools based on the remaining F2 generation plants in S102 can be achieved through S1021 to S1025, as described in the following steps.
[0049] S1021. The oil content of the seeds of the remaining F2 generation plants was retested using gas chromatography.
[0050] S1022. When the initial oil content and the retested oil content of the seeds of the remaining F2 generation plants are inconsistent, the inconsistent plants are removed from the remaining F2 generation plants to obtain the removed F2 generation plants; the initial oil content of the seeds of the remaining F2 generation plants is initially determined by a near-infrared spectroscopy grain analyzer.
[0051] In some embodiments of the present invention, the oil content of the seeds of the remaining F2 generation plants is re-determined using gas chromatography to obtain the re-measured oil content. The initial oil content of the seeds of the initial F2 generation plants is initially determined by a near-infrared spectroscopy grain analyzer. The initial oil content of the seeds of the remaining F2 generation plants is obtained from the initial oil content of the seeds of the initial F2 generation plants. Then, the initial oil content and the re-measured oil content are compared. Plants with inconsistent oil content are removed from the remaining F2 generation plants to obtain the removed F2 generation plants.
[0052] S1023. Based on the retesting of the oil content of the seeds of the removed F2 generation plants, the plants are screened to obtain the screened F2 generation plants.
[0053] In some embodiments of the present invention, in order to improve the accuracy of the data, the F2 generation plants are re-sorted by retesting the oil content of the seeds of the F2 generation plants after removal. At both ends of the sorting result, a predetermined number of extreme individuals, such as the top 10% of high-oil extreme individuals and the bottom 10% of low-oil extreme individuals, are used to obtain the F2 generation plants after initial screening. Plants with abnormal phenotypes such as seed malformation and disease are removed from the F2 generation plants after initial screening to obtain the final F2 generation plants after screening.
[0054] S1024. Genomic DNA was extracted from the seedling leaves of the selected F2 generation plants to obtain their respective genomic DNA.
[0055] In some embodiments of the present invention, seedling leaves corresponding to the screened F2 generation plants are obtained in an ultra-low temperature freezer, and then genomic DNA is extracted from the seedling leaves to obtain their respective genomic DNA.
[0056] For example, when extracting genomic DNA from seedling leaves, approximately 500 mg of seedling leaf tissue can be taken and placed in a 2 mL centrifuge tube. Two sterile steel balls are added to the tube, and after flash freezing in liquid nitrogen, the tissue is ground into a fine powder using a ball mill through the steel balls. 800 μL of CTAB extraction buffer is added to the powder, and the centrifuge tube is placed in a 65°C water bath for 1 hour. During this 1-hour period, the tube is gently inverted several times every 15 minutes to mix thoroughly. After the liquid in the centrifuge tube cools to room temperature, 5 μL of RNase is added again, and the tube is incubated at 37°C for 15 minutes. An equal volume of a mixture consisting of phenol, chloroform, and isoamyl alcohol in a 25:24:1 ratio is then added to the centrifuge tube. The tube is thoroughly inverted to mix, and the mixture is centrifuged at 12,000 rpm for 10 minutes. Finally, the supernatant is transferred to a new 2 mL centrifuge tube. Next, add an equal volume of pre-cooled isopropanol to a new centrifuge tube, mix gently, and precipitate the DNA in the new centrifuge tube at -20°C for at least 30 minutes. Then, centrifuge at 12,000 rpm for 10 minutes, remove the supernatant from the new centrifuge tube, wash the DNA precipitate in the new centrifuge tube twice with 75% ethanol, dry the washed DNA precipitate in a vacuum desiccator for about 40 minutes, dissolve it with an appropriate amount of sterile water, and finally obtain high-quality genomic DNA.
[0057] S1025. Remove plants with unqualified genomic DNA from the screened F2 generation plants to obtain the target F2 generation plants; and construct high-oil mixing pools and low-oil mixing pools using the genomic DNA corresponding to the target F2 generation plants.
[0058] In some embodiments of the present invention, genomic DNA samples can be screened using 1% agarose gel electrophoresis to obtain initial genomic DNA; then, secondary screening is performed based on the concentration and purity of the initial genomic DNA as determined by Nanodrop 2000 to obtain high-quality genomic DNA. Among the screened F2 generation plants, those with high-quality genomic DNA are retained, while those with substandard genomic DNA are removed to obtain the target F2 generation plants. The target F2 generation plants include high-oil target F2 generation plants and low-oil target F2 generation plants. A high-oil mixed pool is constructed using the genomic DNA of the high-oil target F2 generation plants, and a low-oil mixed pool is constructed using the genomic DNA of the low-oil target F2 generation plants.
[0059] In some embodiments of the present invention, the target single nucleotide polymorphism site obtained in S103, which is consistent with the parent with the corresponding oil content in the high-oil mixed pool and the low-oil mixed pool, and whose genotype is inconsistent between the parents, can be achieved through S1031 to S1034, as described in the following steps.
[0060] S1031. Remove unsuitable sequence data from the original sequencing data to obtain the first sequencing data; and match the first sequencing data with the soybean reference genome to obtain matching information.
[0061] In some embodiments of the present invention, during the generation of raw sequencing data, low-quality sequences, adapter contamination sequences, and polymerase chain reaction (PCR) repetitive sequences may be generated simultaneously due to issues such as instrument signal attenuation, library construction adapter ligation, and library amplification. Therefore, the raw sequencing data is filtered to remove low-quality sequences, adapter contamination sequences, and PCR repetitive sequences, resulting in first sequencing data. The first sequencing data is then compared with the soybean reference genome for DNA sequence fragment alignment. Based on the alignment results, the matching information of the first sequencing data in the soybean reference genome is obtained. This matching information may include the matching location, matching method, alignment quality, and the original base sequence of the first sequencing data.
[0062] S1032. Generate an initial comparison file based on the matching information and the first sequencing data; and sort the alignment information of the initial comparison file according to the genomic position to obtain the first comparison file.
[0063] In some embodiments of the present invention, a high-throughput sequence alignment tool is used to perform comparisons based on the first sequencing data and matching information, and then an initial comparison file is generated. Next, the alignment records in the initial comparison file are sorted by chromosome number and genomic physical location to generate a sorted first comparison file.
[0064] The initial comparison file can be a SAM format file.
[0065] S1033. Remove duplicate reads generated during the amplification process from the first comparison file to obtain the target comparison file. Perform variant detection on the target comparison file to obtain the initial single nucleotide polymorphism sites where the genotypes of the parents are inconsistent.
[0066] In some embodiments of the present invention, a single nucleotide polymorphism (SNP) site refers to a difference of a single base at the same location in the genome between different individuals. In the first comparison file, a deduplication tool is used to identify and remove optical repeats and sequence repeats introduced by the PCR (polymerase chain reaction) amplification process to obtain the target comparison file. Then, existing variant detection software can be used to identify SNP sites in the target comparison file, and by comparing the variant information of the two parents, initial SNP sites with inconsistent genotypes between the two parents can be screened out.
[0067] S1034. Obtain the target single nucleotide polymorphism (SNP) site from the initial SNP site. The target SNP site is the site that is consistent with the genotype of the high-oil parent in the high-oil mixed pool and consistent with the genotype of the low-oil parent in the low-oil mixed pool.
[0068] In some embodiments of the present invention, among multiple initial single nucleotide polymorphism sites, a site is found that is consistent with the genotype of the high-oil parent in the high-oil mixture and with the genotype of the low-oil parent in the low-oil mixture, and this site is used as the target single nucleotide polymorphism site.
[0069] For example, if a single nucleotide polymorphism (SNP) site exists, and the high-oil parent has the GG allele at this SNP site while the low-oil parent has the AA allele at this SNP site, then this SNP site is taken as the initial SNP site. Further, in a high-oil mixed pool, if an initial SNP site exists, and the frequency of the GG allele at this initial SNP site is higher than the frequency of the AA allele, while in a low-oil mixed pool, the frequency of the AA allele at the initial SNP site is higher than the frequency of the GG allele, then this initial SNP site is taken as the target SNP site.
[0070] In some embodiments of the present invention, obtaining the first candidate gene based on soil physicochemical data, microclimate data and target single nucleotide polymorphism sites in S103 can be achieved through S103A to S103F, as described in the following steps.
[0071] S103A: The proportion of alleles from the parental source at the target single nucleotide polymorphism site was statistically analyzed in both the high-oil-mixed pool and the low-oil-mixed pool.
[0072] S103B, the proportions are used as the single nucleotide polymorphism indices for each target single nucleotide polymorphism site in the high-oil-mixed pool and the low-oil-mixed pool, respectively.
[0073] In some embodiments of the present invention, the number of alleles carried by sequencing reads of each target single nucleotide polymorphism (SNP) site in the high-oil mixed pool that are identical to those of the high-oil paternal parent and the number of alleles carried by each read that are identical to those of the low-oil maternal parent are counted. The proportions are then calculated based on these counts. These proportions are used as the SNP indices for each target SNP site in the high-oil and low-oil mixed pools, respectively.
[0074] S103C: The difference between the single nucleotide polymorphism indices corresponding to the high-oil-mixing pool and the low-oil-mixing pool is calculated to obtain the single nucleotide polymorphism index difference; and the predicted oil content is obtained based on the oil content prediction model, soil physicochemical data and microclimate data.
[0075] S103D: The residual is calculated based on the predicted oil content and the initial oil content of the seeds of the initial F2 generation plants.
[0076] In some embodiments of the present invention, using the alleles of the high-oil parent as a benchmark, the difference between the single nucleotide polymorphism (SNP) index in the high-oil mixed pool and the corresponding SNP in the low-oil mixed pool is calculated to obtain the SNP difference value. That is, the difference in SNP index is calculated at the same physical location of the SNP site in the genomes of the high-oil mixed pool and the low-oil mixed pool. The average values of soil physicochemical data for each grid during the growth period are calculated, and the daily average temperature, cumulative radiation, and other indicators are statistically analyzed in segments according to the seedling stage, flowering stage, and grain-filling stage. The average values and indicators during the growth period are input into the oil content prediction model to obtain the predicted oil content corresponding to the seeds of the initial F2 generation plants. The initial oil content and predicted oil content of the seeds of the initial F2 generation plants are then compared to obtain the residual.
[0077] S103E: The residual is used to correct the single nucleotide polymorphism index difference, and the corrected target single nucleotide polymorphism index difference is obtained.
[0078] S103F, the first candidate gene was obtained based on the difference in the corrected target single nucleotide polymorphism index.
[0079] In some embodiments of the present invention, residuals are used to correct the single nucleotide polymorphism index difference to reduce the influence of environmental bias, thereby obtaining the corrected target single nucleotide polymorphism index difference. Then, based on the corrected target single nucleotide polymorphism index difference, regions with extremely large and stable differences in allele frequencies between high-oil plants and low-oil plants are obtained, and genes in these regions are selected as first candidate genes.
[0080] In some embodiments of the present invention, S103E can be implemented by S201 to S202, as described in the following steps.
[0081] S201. A linear regression model is constructed based on the single nucleotide polymorphism index difference and residuals; and the regression coefficients of the linear regression model are extracted.
[0082] S202. Using the regression coefficient as a weighting factor, and multiplying the single nucleotide polymorphism index difference by the weighting factor, the corrected target single nucleotide polymorphism index difference is obtained.
[0083] In some embodiments of the present invention, a linear regression model is constructed with the residual as the dependent variable and the single nucleotide polymorphism index difference as the independent variable. The regression coefficients of the linear regression model are obtained, and then the regression coefficients are used as weighting factors. The weighting factors are multiplied by the single nucleotide polymorphism index difference to obtain the corrected target single nucleotide polymorphism index difference.
[0084] In some embodiments of the present invention, S103F can be implemented by S301 to S302, as described in the following steps.
[0085] S301. Statistical tests are performed based on the difference in the corrected target single nucleotide polymorphism index to obtain the test probability value. Based on the test probability value, the target single nucleotide polymorphism sites are screened to obtain single nucleotide polymorphism sites that meet the conditions.
[0086] S302. In the soybean reference genome, obtain the physical location of single nucleotide polymorphism sites that meet the conditions, and select genes in the genomic region where the density of the physical location is greater than the threshold as the first candidate genes.
[0087] In some embodiments of the present invention, a statistical test is performed on the difference in the corrected target single nucleotide polymorphism (SNP) index to calculate the test probability value (P-value) for each target SNP site. This test probability value can assess the significance of the association between the target SNP site and soybean oil content. A threshold is set, and the test probability value of each target SNP site is compared with the threshold. When it is less than the threshold, it is considered a significantly associated SNP site. When it is not less than the threshold, it is eliminated. Finally, the physical location of these significantly associated target SNP sites in the soybean reference genome is located, such as chromosomal coordinates. A density threshold is set, and the genomic regions where the target SNP sites are densely distributed are obtained by combining the density threshold. All genes in the region are selected as first candidate genes.
[0088] Reference Figure 2 The diagram shows a structural schematic of an electronic device according to an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.
[0089] like Figure 2As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.
[0090] in:
[0091] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.
[0092] Communication interface 504 is used to communicate with other electronic devices or servers.
[0093] The processor 502 is used to execute program 510, specifically the relevant steps in the above method embodiments.
[0094] Specifically, program 510 may include program code that includes computer operation instructions.
[0095] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The smart device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.
[0096] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0097] Specifically, program 510 can be used to cause processor 502 to perform the operations corresponding to the methods described in the above method embodiments.
[0098] The specific implementation of each step in program 510 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0099] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of the present invention can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.
[0100] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0101] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present invention.
[0102] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.
Claims
1. A method for locating soybean oil genes based on residual shrinkage networks, characterized in that, include: Multiple seeds randomly obtained from each mature F1 generation plant were sown in the same hole in a grid in the experimental field to obtain a random initial F2 generation plant, corresponding soil physicochemical data and microclimate data; and the remaining F2 generation plants were obtained by sorting the seeds harvested from the initial F2 generation plants based on their initial oil content. High-oil-mixing ponds and low-oil-mixing ponds were constructed based on the remaining F2 generation plants; Raw sequencing data of high-oil mixed cells, low-oil mixed cells, and the parent lines were obtained separately. Based on the comparison of the original sequencing data and the soybean reference genome, the target single nucleotide polymorphism sites that are consistent with the corresponding oil content parents in the high oil mixed pool and low oil mixed pool, and whose genotypes are inconsistent between the parents, were obtained. The first candidate gene was obtained based on soil physicochemical data, microclimate data and target single nucleotide polymorphism sites. Expression level characteristics were obtained using transcriptome sequence data from each parent at different developmental stages; And extract the genotypic characteristics and genetic effect characteristics of the first candidate gene at the target single nucleotide polymorphism site; Target features are obtained based on genotype characteristics, expression level characteristics, genetic effect characteristics, and deep residual shrinkage networks; The first candidate gene was screened using target features to obtain the second candidate gene; and the third candidate gene enriched in the lipid synthesis pathway was obtained based on the second candidate gene and the functional database. The probability values of the second candidate gene and the total genes of the lipid synthesis pathway were obtained, and the core genes controlling the soybean oil content were obtained based on the probability values and the third candidate gene.
2. The method according to claim 1, characterized in that, The construction of high-oil-mixing and low-oil-mixing ponds based on the remaining F2 generation plants includes: The oil content of the seeds of the remaining F2 generation plants was retested using gas chromatography. When the initial oil content and the retested oil content of the seeds of the remaining F2 generation plants are inconsistent, the inconsistent plants are removed from the remaining F2 generation plants to obtain the removed F2 generation plants; the initial oil content of the seeds of the remaining F2 generation plants is initially determined by a near-infrared spectroscopy grain analyzer. Plant selection was carried out by retesting the oil content of the seeds of the removed F2 generation plants to obtain the selected F2 generation plants. Genomic DNA was extracted from the leaves of the F2 generation plants after screening to obtain their respective genomic DNA. Plants with unqualified genomic DNA were removed from the selected F2 generation plants to obtain the target F2 generation plants; and high-oil mixing pools and low-oil mixing pools were constructed using the genomic DNA corresponding to the target F2 generation plants.
3. The method according to claim 1, characterized in that, The target single nucleotide polymorphism (SNP) sites, obtained by comparing raw sequencing data with the soybean reference genome and identifying genotypes consistent with the corresponding oil-content parents in both high-oil-content and low-oil-content mixed pools, but with genotype inconsistencies between parents, include: The first sequencing data was obtained by removing unsuitable sequence data from the original sequencing data; and the first sequencing data was then matched with the soybean reference genome to obtain matching information. An initial alignment file is generated based on the matching information and the first sequencing data; and the alignment information of the initial alignment file is sorted according to the genomic position to obtain the first alignment file. Repeated reads generated during the amplification process are removed from the first comparison file to obtain the target comparison file; and variant detection is performed on the target comparison file to obtain the initial single nucleotide polymorphism sites where the genotypes of the parents are inconsistent. Target single nucleotide polymorphisms (SNPs) were obtained from the initial SNP sites. The target SNP sites were those that were identical to the genotype of the high-oil parent in the high-oil mixed pool and identical to the genotype of the low-oil parent in the low-oil mixed pool.
4. The method according to claim 1, characterized in that, The process of obtaining the first candidate gene based on soil physicochemical data, microclimate data, and target single nucleotide polymorphism sites includes: The proportion of alleles from the parental source at the target single nucleotide polymorphism site was statistically analyzed in both high-oil-mixed and low-oil-mixed pools. The proportions were used as the single nucleotide polymorphism indices for each target single nucleotide polymorphism site in the high-oil-mixed pool and the low-oil-mixed pool, respectively. The difference between the single nucleotide polymorphism indices in the high-oil-mixing pool and the low-oil-mixing pool was calculated to obtain the single nucleotide polymorphism index difference; and the predicted oil content was obtained based on the oil content prediction model, soil physicochemical data and microclimate data. The residuals were calculated based on the predicted oil content and the initial oil content of the seeds of the initial F2 generation plants. The residual is used to correct the single nucleotide polymorphism index difference, and the corrected target single nucleotide polymorphism index difference is obtained. The first candidate gene is obtained based on the difference in the corrected target single nucleotide polymorphism index.
5. The method according to claim 4, characterized in that, The step of correcting the single nucleotide polymorphism index difference using residuals to obtain the corrected target single nucleotide polymorphism index difference includes: A linear regression model was constructed based on the single nucleotide polymorphism index difference and residuals; and the regression coefficients of the linear regression model were extracted. The regression coefficients are used as weighting factors, and the single nucleotide polymorphism index difference is multiplied by the weighting factors to obtain the corrected target single nucleotide polymorphism index difference.
6. The method according to claim 4, characterized in that, The process of obtaining the first candidate gene based on the corrected target single nucleotide polymorphism index difference includes: Statistical tests were performed based on the difference in the corrected target single nucleotide polymorphism index to obtain the test probability value. Based on the test probability value, the target single nucleotide polymorphism sites were screened to obtain single nucleotide polymorphism sites that meet the conditions. In the soybean reference genome, the physical locations of eligible single nucleotide polymorphism sites are obtained, and genes in the genomic regions where the density of physical locations is greater than a threshold are selected as first candidate genes.
Citation Information
Patent Citations
Multi-omics and phenotype association prediction method based on deep residual shrinkage network
CN115132279A
Rice disease resistance character whole genome association analysis method based on SNP (Single Nucleotide Polymorphism) marker
CN120564824A