Method for efficiently excavating long-chain non-coding RNA (Ribonucleic Acid) for regulating and controlling soybean grain weight based on multiple omics and application technology thereof
Through multiomics methods, screening and analyzing lncRNAs related to grain weight in soybeans has solved the problem of failing to effectively explore lncRNAs that regulate soybean grain weight in the prior art, and the analysis of their genetic regulation mechanism is realized, providing important theoretical and application value for crop breeding.
Patent Information
- Application Number
- CN202510245422.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-27
AI Technical Summary
The existing technology has not yet effectively discovered long-chain non-coding RNA (lncRNA) that can regulate the weight of soybean grains, and there are few studies that analyze their genetic regulatory mechanisms.
Multiomics methods, including DNA resequencing and RNA transcriptome sequencing, combined with Pearson correlation or Spearman correlation analysis, were used to screen significantly associated lncRNA sites, and their genetic regulatory mechanisms were analyzed by mixed linear models in the GCTA software package.
It has achieved rapid screening of lncRNAs significantly related to soybean grain weight, and analyzed its genetic regulation mechanism, providing important genetic resources and theoretical value for crop breeding.
Smart Images

Figure CN120048350A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of plant biotechnology, and is related to the application of plant non-coding genes. In particular, it relates to a method for efficiently mining long non-coding RNAs that regulate soybean seed weight based on multi-omics and its application. Background Art
[0002] In addition to a part of coding genes in the eukaryotic genome, there are also a large number of non-coding regions. Such regions may be specifically transcribed during the growth and development stages of plants. Long non-coding RNA (lncRNA) is a type of non-coding RNA with a length greater than 200 nt. This type of RNA can transcribe the intergenic regions of plant genomes, the antisense strands of coding genes, and intron regions. So far, a large number of lncRNAs have been identified in various plants. Among them, a part of lncRNAs have been proven to participate in regulating various pathways such as plant flowering and seed development.
[0003] The components of soybean yield are complex, and seed weight is an important component factor affecting yield. Seed weight is a quantitative trait controlled by multiple genes. So far, soybean biological researchers have cloned multiple QTL loci related to seed weight and multiple coding genes involved in seed weight formation through various methods. For example, Phosphatase 2C-1 ( GmPP2C-1 ), BIG SEEDS1 ( GmBS1 ) and KINASE-INDUCIBLE DOMAIN INTERACTING 8 ( GmKIX8-1 ). However, lncRNAs that can regulate seed weight have not yet been mined. In addition, although some lncRNAs have been reported to have biological functions, their genetic regulatory mechanisms have rarely been analyzed. Summary of the Invention
[0004] Based on the above background, the technical problem to be solved by the present invention is a method for efficiently mining long non-coding RNAs that regulate soybean seed weight based on multi-omics and its application. The method can quickly screen out lncRNAs significantly related to seed weight and has important theoretical and practical application values.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: A method for efficiently mining long non-coding RNAs that regulate soybean seed weight based on multi-omics and its application, comprising the following steps: 1) Collect leaves of all individuals of germplasm resources for DNA extraction; collect seeds at different stages of seed development of the above all individuals for RNA extraction, and harvest seeds at the mature stage for investigation of seed weight; 2) Construct a resequencing library using the genomic DNA of the leaf samples described in step 1); construct a 3' RNA transcriptome sequencing library using the RNA of the seed samples described in step 1). 3) Identify the expression of genome-wide lncRNAs from the transcriptome sequencing data described in 2). 4) Use Pearson correlation or Spearman correlation to perform a correlation analysis on the sample phenotypic trait data detected in step 1) and the lncRNA expression data obtained in step 3), and screen for significantly associated lncRNA loci. 5) Identify the genetic variant loci SNP within the whole genome from the DNA resequencing data; use the mixed linear model in the GCTA software package to associate the SNP data with the lncRNA expression data. Analyze the genetic regulatory mechanism of the significantly associated lncRNA loci.
[0006] Preferably, the threshold for screening significantly associated lncRNA loci in step 4) is FDR < 0.05, and Benjamini-Hochberg (FDR) correction is used. Preferably, the software used to identify SNP and Indel loci within the whole genome in step 3) is the GATK software package, and the software package used to calculate lncRNA expression is edgeR.
[0007] Preferably, the DNA and RNA sequencing in step 2) is paired-end sequencing; the read length of the sequencing is 150 bp, the depth of the sequencing is 10X; 3' RNA technology is used for RNA sequencing. The sequencing is performed using the Illumina platform.
[0008] Preferably, the sample is a leguminous crop.
[0009] Preferably, the phenotypic data includes 100-seed weight, seed length, seed width, etc.
[0010] The present invention provides the application of the described method in plant molecular breeding.
[0011] Beneficial effects of the present invention: The method provided by the present invention screens for lncRNAs related to 100-seed weight and analyzes its genetic regulatory mechanism, providing important gene resources and theoretical value for crop breeding. Description of the Drawings
[0012] Figure 1 . It is the result of the transcriptome association analysis of 100-seed weight and lncRNA expression in Example 1.
[0013] Figure 2. Results of the genome-wide eQTL association analysis of lncRNAs in Example 2. Detailed implementation manners
[0014] The present invention provides a method for efficiently mining long non-coding RNAs that regulate soybean seed weight based on multi-omics, comprising the following steps: 1) Collect leaves of all individuals of germplasm resources for DNA extraction; collect seeds at different stages of seed development of the above all individuals for RNA extraction, and harvest seeds at the mature stage for investigation of seed weight; 2) Construct a re-sequencing library using the genomic DNA of the samples described in step 1); construct a transcriptome sequencing library using the genomic RNA of the samples described in step 1); 3) Identify genetic variation sites SNP within the whole genome from the DNA re-sequencing data; identify the expression of genome-wide lncRNAs from the transcriptome sequencing data; 4) Perform correlation analysis on the sample phenotypic trait data detected in step 1) and the lncRNA expression data obtained in step 3) using Pearson correlation or Spearman correlation, and screen for significantly associated lncRNA loci; 5) Use the mixed linear model in the GCTA software package to analyze the genetic regulatory mechanism of the significantly associated lncRNA loci.
[0015] In the present invention, when collecting the tissue of the library construction sample, it is preferably carried out at the same time as much as possible to avoid environmental errors. As much as possible, immediately after the tissue leaves the plant body, it is quickly frozen in liquid nitrogen. To obtain sample phenotypic trait data, the present invention has no special limitation on the species of the sample, preferably a plant, more preferably an annual leguminous plant; in the specific implementation process of the present invention, the sample is preferably soybean. In the present invention, the tissue is preferably leaves and seeds. The sampling time preferably remains consistent every day in the present invention, and the growth state of the growing individuals is preferably kept consistent as much as possible, so as to avoid the influence of environmental factors and plant growth state on subsequent library construction and sequencing. Thereby, stable and reliable SNP loci and the expression levels of lncRNAs are identified, so as to realize the analysis of the genetic regulatory mechanism of lncRNA expression. The present invention has no special limitation on the phenotypic traits, and conventional phenotypic traits of the sample are acceptable, preferably phenotypic traits with practical application significance. In the specific implementation process of the present invention, the phenotypic traits are preferably the 100-seed weight, seed length and seed width of the seeds. The present invention has no special limitation on the detection method of the phenotypic traits, and conventional phenotypic trait detection methods in the art can be used.
[0016] The present invention extracts genomic DNA from the sample to obtain the genomic DNA of the sample. The present invention has no special limitation on the method for extracting the genomic DNA, and any conventional genomic DNA extraction method in the art can be used; preferably, the CTAB method is used to extract plant genomic DNA. Specifically, a tissue grinder can be used to grind the leaf tissue sufficiently; subsequently, a water bath at 65 °C is carried out using CTAB; DNA is extracted by chloroform; finally, DNA is precipitated using isopropanol. After extracting the genomic DNA of the sample, the present invention also detects the purity and integrity of the sample DNA. For purity detection, a UV spectrophotometer is used to detect the OD260 / OD280 ratio of the DNA and determine the purity of the DNA sample through the curve; in addition, integrity detection is carried out using the agarose gel electrophoresis method.
[0017] The present invention extracts total RNA from the sample to obtain the total RNA of the sample. The present invention has no special limitation on the method for extracting the RNA, and any conventional RNA extraction method in the art can be used; preferably, the TRIzol method is used to extract RNA. Specifically, a tissue grinder can be used to grind the seed tissue sufficiently; subsequently, TRIzol reagent is added and left standing at room temperature for 20 minutes; RNA is extracted by chloroform; finally, RNA is precipitated using isopropanol, and the RNA is washed with 70% alcohol. After extracting the RNA of the sample, the present invention also detects the purity and integrity of the sample RNA. For purity detection, a UV spectrophotometer is used to detect the OD260 / OD280 ratio and determine the purity of the RNA sample; in addition, integrity detection is carried out using the agarose gel electrophoresis method. If all three RNA bands are complete and clear, it indicates that the extraction quality is good and can be used for subsequent library construction.
[0018] After obtaining the DNA re-sequencing and RNA transcriptome libraries, the present invention performs sequencing to obtain DNA re-sequencing data and transcriptome data. In the present invention, the sequencing data is preferably paired-end sequencing, the read length of the sequencing is preferably 150 bp, and the depth of the sequencing is preferably 10X; the sequencing is preferably carried out using the Illumina platform. In the specific implementation process of the present invention, the sequencing is preferably entrusted to Beijing Novogene Bioinformatics Technology Co., Ltd.
[0019] After obtaining the transcriptome sequencing data, the present invention calculates the expression levels of lncRNAs across the whole genome from the transcriptome data. In the present invention, preferably, the featureCounts software is used to calculate the counts of each lncRNA, and the R software package edgeR is used to calculate the expression values of each lncRNA. The lncRNA expression data obtained by the present invention can be used for lncRNA expression-phenotype association analysis, thereby providing support for subsequent exploration of the functions of lncRNAs.
[0020] After obtaining the DNA sequencing data, the present invention uses the GATK software package to identify genome-wide SNPs. After filtering the alignment quality and heterozygous sites, the remaining data are used as the final genetic variation data to support subsequent analysis.
[0021] After obtaining the lncRNA sites with significant association, the present invention uses the mixed linear model in the GCTA software package to analyze the genetic regulatory sites of genome-wide lncRNAs.
[0022] The present invention also provides the application of the above method in plant molecular breeding, and preferably the application in plant molecular breeding. The present invention has no special limitation on the specific method of the application.
[0023] The technical solutions provided by the present invention are described in detail below in conjunction with the embodiments, but they should not be construed as limiting the protection scope of the present invention. Embodiment
[0024] The specific operation steps are as follows: Step 1) Seeds of each individual are taken from the single plants of 238 germplasm resources planted in the Liuhe Experimental Base in Nanjing, Jiangsu Province as the transcriptome library construction samples of the association population. It is carried out at 10:00-11:00 in the morning, and fresh seeds 14 days and 21 days after flowering are taken. To prevent rapid degradation of RNA, after wrapping the samples with tin foil, they are quickly placed in liquid nitrogen for storage.
[0025] Step 2) TRIzol method is used to extract RNA from the seeds.
[0026] After completing the above steps, the quality of the obtained total RNA can be further detected. Specifically, a UV spectrophotometer can be used to detect the OD260 / OD280 ratio of the sample to determine the purity of RNA and DNA. Agarose gel electrophoresis is used to judge the integrity of RNA.
[0027] Then step 3) is implemented to construct a transcriptome sequencing library for the extracted RNA. The sequencing work is completed by Beijing Novogene Bioinformatics Technology Co., Ltd.
[0028] Step 4) is implemented to calculate the expression of lncRNA using the transcriptome sequencing data. The transcriptome sequencing data are aligned to the reference genome of the soybean cultivar Wm82 using the Hisat2 software. The featureCounts software is used to calculate the counts value of each lncRNA. Finally, the R software package edgeR is used to calculate the expression value of each lncRNA.
[0029] Step 5) Transcriptome-phenotype association. Calculate the correlation between the expression of each lncRNA and the phenotype using Pearson or Spearman correlation. Subsequently, correct the P-values of each association result using Benjamini-Hochberg (FDR). Finally, take the lncRNAs associated with an FDR less than 0.05 as candidate lncRNAs. Use Manhattan Figure 1 to display the association results. The specific association results are shown in Table 1.
[0030] Table 1 Results of transcriptome association analysis Example
[0031] The specific operation steps are as follows: Step 1) Collect leaves from 238 germplasm resources planted in the greenhouse to extract DNA. Conduct this at 9:00 - 11:00 in the morning. After wrapping the samples with tin foil, quickly store them in liquid nitrogen. For the DNA used in DNA resequencing library construction, take the leaves at the V2 stage of soybean development from the samples planted in the greenhouse.
[0032] Step 2) Extract DNA using the CTAB method.
[0033] After completing the above steps, further perform quality detection on the obtained genomic DNA. Specifically, use a UV spectrophotometer to detect the OD260 / OD280 ratio and the purity of the DNA of the sample. Use agarose gel electrophoresis to judge the integrity of the DNA.
[0034] Then implement Step 3) Construct resequencing libraries for the extracted DNA respectively. Hand over the sequencing work to Beijing Novogene Bioinformatics Technology Co., Ltd. to complete.
[0035] Implement Step 4) Identification of genome-wide SNPs. Use the bowtie2 software to align the resequencing data to the reference genome of the soybean cultivar Wm82. Use the GATK software for genome-wide SNP identification. After quality filtering, retain the SNPs with reliable quality for subsequent analysis.
[0036] Implement Step 5) Genome-wide eQTL association analysis. Use the mixed linear model in the GCTA software package to associate the expression value of each lncRNA as the molecular trait with SNPs. Screen for significantly associated eQTL loci, with the threshold of P < 1 / N (N is the number of SNPs). Then merge the adjacent eQTL loci and use the most significant SNP of each eQTL locus as a representative for subsequent analysis. The association results are as Figure 2 . The specific association results are shown in Table 2.
[0037] Table 2 Results of eQTL association analysis of TWAS-significant lncRNAs (only partial association results are shown due to the long table) From the above experimental data, it can be seen that the present invention has the following advantages: The method first uses transcriptome-wide association analysis to analyze lncRNAs related to grain weight. At the same time, multiple omics are used to analyze the genetic regulatory mechanisms of these lncRNAs that regulate grain weight. Therefore, the invention has important application value for cultivating high-yield soybeans and also provides theoretical guidance for the molecular mechanism of grain weight formation.
Claims
1. A method for efficiently mining long non-coding RNA that regulates soybean grain weight based on multi-omics and its application technology, comprising the following steps: 1) Collect leaves from all individuals in the germplasm resource sample for DNA extraction; collect seeds from all individuals at different developmental stages for RNA extraction; harvest seeds at maturity for grain weight inspection; 2) constructing a resequencing library using the sample genomic DNA described in step 1); Constructing a 3' RNA transcriptome sequencing library using the sample genomic RNA in step 1); 3) identifying genome-wide genetic variation sites, including SNP sites, from the DNA resequencing data; identifying genome-wide lncRNA expression from the transcriptome sequencing data; 4) Using Pearson correlation or Spearman correlation, the sample phenotypic trait data detected in step 1) and the lncRNA expression data obtained in step 3) are analyzed for correlation, and significantly associated lncRNA sites are screened; 5) The mixed linear model in the GCTA software package was used to analyze the eQTL sites significantly associated with the lncRNA sites.
2. The method according to claim 1, characterized in that The threshold for screening significantly associated lncRNA sites in step 4) was FDR<0.05, and the P value was corrected using Benjamini-Hochberg (FDR).
3. The method according to claim 1, characterized in that: The software used in step 3) for identifying SNP sites in the whole genome is the GATK software package, and the software package used for calculating lncRNA expression is edgeR.
4. The method according to claim 1, characterized in that: The DNA and RNA sequencing in step 2) is double-end sequencing, the read length of the sequencing is 150 bp, and the depth of the sequencing is 10X; the sequencing is performed using the Illumina platform.
5. The method according to any one of claims 1 to 4, characterized in that: The sample is a leguminous crop.
6. The method according to any one of claims 1 to 4, characterized in that: The phenotypic data include grain weight, grain length and grain width, etc.
7. Use of the method according to any one of claims 1 to 6 in plant molecular breeding.