A fine QTL mapping method for important traits in mulberry trees based on chromosome segment substitution lines
By constructing a BC4F2 backcross population in mulberry trees and performing targeted sequencing and recombination hotspot analysis, combined with multi-environment experiments and co-expression network models, the problems of incomplete fragment coverage and environmental sensitivity in mulberry QTL mapping were solved, achieving fine mapping of important traits in mulberry trees and high-precision QTL interval compression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional QTL mapping methods in mulberry trees suffer from low mapping accuracy, complex genetic background, and limited recombination events. Furthermore, the long generation cycle of mulberry trees, being perennial woody plants, makes it difficult to effectively isolate and verify target QTLs. In addition, the lack of efficient molecular marker-assisted selection strategies leads to problems such as incomplete fragment coverage, low background recovery rate, and self-inbreeding decay in the construction of CSSL libraries.
By constructing a BC4F2 backcross population through high heterozygous parental hybridization and multiple generations of backcrossing, and combining targeted sequencing and recombination hotspot analysis, a ladder-covered CSSL library was established. Phenotypic data were collected in multi-environment repeated experiments, a hybrid model was constructed, and key genes were screened by combining a co-expression network model to achieve fine localization of important traits of mulberry trees.
It effectively compresses the QTL interval, ensures complete coverage of the CSSL library and a pure genetic background, reduces environmental noise interference, accurately screens out key regulatory genes, avoids false positive results, and improves the accuracy of QTL localization.
Smart Images

Figure CN121191572B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular breeding technology, specifically to a method for fine mapping of important traits in mulberry trees based on chromosome segment substitution lines. Background Technology
[0002] Mulberry is an important economic crop, widely used in silk production, and also has significant value in ecological restoration and medicinal resource development. With the development of the mulberry industry, the need for genetic improvement of important agronomic traits such as yield, quality, and stress resistance is becoming increasingly urgent. Quantitative trait locus (QTL) mapping, as an important tool for elucidating the genetic mechanisms of complex traits, has been widely applied in molecular breeding research on various crops.
[0003] Traditional QTL mapping methods mainly rely on mapping populations such as F2, RIL (Recombinant Inbred Lines), or DH (Doubled Haploid). While these methods can achieve preliminary QTL mapping to some extent, their complex genetic background and limited recombination events result in low mapping accuracy and make it difficult to effectively isolate and validate target QTLs. Furthermore, most mulberry species are perennial woody plants with long generation cycles, making hybridization difficult and further limiting the efficiency of traditional QTL analysis methods.
[0004] Chinese invention patent CN107365870A discloses a method for rice gene mapping and molecular breeding, relating to the field of rice gene mapping and breeding technology. The method includes the following steps: (1) selecting parental lines for hybridization and self-pollination to construct an F2 population for mapping; (2) examining the phenotypic traits and genotypes of individual plants; (3) performing conventional linkage analysis; (4) selecting heterozygous individual plants and screening exchange plants; (5) performing iterative mapping to locate the target gene within 5 cM; (6) performing fine mapping or molecular breeding, using heterozygous individual plants in the target region to continue self-pollination to obtain a continuous F2 population of 2000 plants for mapping, developing fine mapping markers to locate the target gene within 100 kb; (7) using molecular marker-assisted selection breeding with molecular markers linked to the target gene as the goal of molecular breeding. This invention is time-saving and efficient, addressing the urgent need for rapid gene mapping in current rice functional gene research and molecular breeding.
[0005] Because the mulberry genome is highly heterozygous and has a complex polyploid background, and is sensitive to environmental responses, it is difficult to effectively resolve complex traits (such as leaf size / thickness) that are tightly linked or have epistatic interactions when using F2 or RILs populations for QTL mapping. Although fine mapping can be performed using chromosome segment replacement lines, the lack of efficient marker-assisted selection strategies in mulberry leads to problems such as incomplete segment coverage, low background recovery rate, and self-crossing decay in CSSL library construction, thus limiting the accurate compression of QTL intervals. Summary of the Invention
[0006] The purpose of this invention is to provide a method for fine mapping of important traits in mulberry trees based on chromosome segment substitution lines, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for fine mapping of QTLs of important traits in mulberry trees based on chromosome segment substitution lines, comprising:
[0008] S1: Constructing a genotype-phenotype association matrix: The backcross population was screened through high heterozygous parental hybridization and multiple generations of backcrossing, and the genomes of the screened backcross population were marked through targeted sequencing and recombination hotspots;
[0009] S2: Phenotypic modeling: Based on the experimental plants, samples were collected, and a mixture model was obtained based on the environmental effects of the experimental plants and the CSSL covering line.
[0010] S3: Gene Screening: Hub gene screening is performed using constructed co-expression network models and hybrid models, including:
[0011] S3.1: Constructing a co-expression network model: Based on the denoised phenotypic data and multiple preset soft threshold parameters, a co-expression network model is constructed.
[0012] S3.2: Phenotypic Association: Based on characteristic genes and environmental effects, determine the Pearson correlation coefficient, and compare the Pearson correlation coefficient with a preset correlation coefficient threshold to filter the Pearson correlation coefficients, specifically as follows:
[0013] If the Pearson correlation coefficient is greater than a preset correlation coefficient threshold, the Pearson correlation coefficient is retained; otherwise, the Pearson correlation coefficient is removed.
[0014] S3.3: Gene Screening: Based on the retained Pearson correlation coefficients, individual gene expression levels, and target phenotypic data, the sum of connectivity and association strength are obtained, and a comprehensive score for hub genes is determined. Simultaneously, the comprehensive score for hub genes is compared with a preset score threshold to determine the selected hub genes. Specifically:
[0015] When the comprehensive score of the hub gene is greater than the preset score threshold, the hub gene corresponding to the comprehensive score is the selected hub gene; otherwise, the hub gene corresponding to the comprehensive score is not the selected hub gene.
[0016] Furthermore, the genomes of the selected backcross populations are labeled, including:
[0017] S1.1: Constructing the BC4F2 backcross population: Using the cultivated species as the female parent and the wild species as the male parent, initial hybridization is performed to obtain the F1 hybrid. At the same time, the F1 hybrid is used as the female parent and the cultivated species as the male parent, and three generations of backcrossing are performed to obtain the BC4F2 backcross population.
[0018] S1.2: Enriched Sequencing: Using a set liquid-phase capture probe, targeted sequencing is performed on the BC4F2 backcross population, and non-synonymous SNPs in the coding region are screened to obtain sequencing data enriched in the target region.
[0019] S1.3: Dynamic labeling: The genetic distance of the sequencing data is obtained by the maximum likelihood method, and a genetic map is drawn to identify the recombination hotspots of the BC4F2 backcross population;
[0020] S1.4: Constructing the CSSL library: Using the genetic map, the target QTL fragment intervals in the BC4F2 backcross population are divided into modules, and the divided modules are covered in a stepwise manner using the CSSL line.
[0021] Furthermore, the BC4F2 backcross population was obtained, including:
[0022] S1.1.1: Constructing F1 generation hybrids: Using SNP genotyping, the proportion of heterozygous sites in the F1 hybrids is obtained. Simultaneously, F1 hybrids with a proportion of heterozygous sites greater than a preset threshold are retained to obtain the initial F1 hybrids. The comprehensive score of the initial F1 hybrids is then obtained. Based on the comprehensive score and a preset score threshold, the final F1 hybrids are determined. Specifically:
[0023] If the overall score is less than a preset score threshold, or if any of the pollen viability, seed setting rate, or seed quality scores are less than a preset single-item score threshold, then the corresponding F1 hybrid after initial screening will be removed; otherwise, the corresponding F1 hybrid after initial screening will be retained.
[0024] S1.1.2: Backcrossing and screening: The final F1 hybrid is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC1 population. At the same time, the BC1 population is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC2 population. The BC2 population is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC3 population. The BC3 population is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC4 population.
[0025] S1.1.3: BC4 Individual Screening: Based on the genetic map of the BC3 population, the major-effect QTL target interval is determined. Simultaneously, based on the genotype data marked in the major-effect QTL target interval and the whole-genome SNP data of the BC4 population, preliminary BC4 individuals are obtained. The donor residual rate in the non-target interval of the preliminary BC4 individuals is compared with a preset donor residual threshold. Based on the comparison results, candidate BC4 individuals are determined, specifically as follows:
[0026] When the donor residual rate in the non-target interval is greater than the preset donor residual threshold, the corresponding BC4 individuals in the initial screening are removed; otherwise, the corresponding BC4 individuals in the initial screening are retained.
[0027] S1.1.4: Introgression hybridization: Based on the QTL gene fragments carried, the candidate BC4 individuals are divided, and the divided candidate BC4 individuals are subjected to reciprocal crosses to obtain F2 hybrids. Based on the background recovery rate and non-target region donor residual rate of the F2 hybrids, the F2 hybrids are screened to obtain the BC4F2 backcross population.
[0028] Furthermore, during the backcross screening process, the backcross individuals are screened based on the background recovery rate and non-target region donor residual rate of the backcross individuals in each backcross population, and the screened backcross individuals are used as the parent stock for the next round of backcrossing, specifically as follows:
[0029] When the background recovery rate is not less than a preset recovery threshold and the donor residual rate in the target area is not greater than a preset residual threshold, the corresponding backcross is retained; otherwise, the corresponding backcross is removed.
[0030] Furthermore, sequencing data enriched in the target region were obtained, including:
[0031] S1.2.1: Probe design: Based on the target region of the major QTL and the preset extension range, the length of the liquid phase capture probe is set. At the same time, based on all coding region sequences in the reference genome, the location of the liquid phase capture probe is set. And based on the annotated non-synonymous SNPs screened from the public database, the probe library is obtained.
[0032] S1.2.2: Liquid phase hybridization: DNA was extracted from the leaf tissue of the BC4F2 backcross population using the CTAB method. Simultaneously, the DNA was disrupted by ultrasound and gel electrophoresis, and target-sized fragments were screened out. The probe library and the target-sized fragments were hybridized using streptavidin magnetic beads to obtain probe-target DNA complexes.
[0033] Furthermore, the CSSL system provides tiered coverage of the divided modules, including:
[0034] S1.4.1: Modular Segmentation: Based on the genotype data of the BC4F2 backcross population, the target QTL fragment interval is determined, and the recombination rate of the target QTL fragment interval is compared with a preset recombination threshold range. Based on the comparison result, modular segmentation is performed, specifically as follows:
[0035] When the recombination rate is less than the lower limit of the preset recombination threshold, the target QTL fragment interval is divided into segments according to the preset length. When the recombination rate is within the range of the preset recombination threshold, the target QTL fragment interval is used as a segmentation module. When the recombination rate is greater than the upper limit of the preset recombination threshold, the target QTL fragment interval is further divided into segments according to the determined low recombination sub-intervals.
[0036] S1.4.2: CSSL line coverage: Based on the divided modules, the BC4F2 backcross population is screened, and at least three coverage lines are set in each module to perform stepwise coverage. Targeted sequencing is performed on the candidate lines to determine the actual coverage range of each coverage line.
[0037] Furthermore, a hybrid model is obtained, including:
[0038] S2.1: Phenotypic Collection: Based on differences in climate type, soil type gradient and average annual temperature, representative ecological zones are set up, and repeated planting experiments are conducted in the representative ecological zones for no less than 2 years to obtain experimental plants;
[0039] S2.2: Single-cell transcription library construction: Sample slices were obtained from the experimental plants, and the sample slices were analyzed for library construction using photosynthetic gene probes;
[0040] S2.3: Modeling and Analysis: Based on the environmental effects of the experimental plants and the CSSL line, a hybrid model was constructed, specifically as follows:
[0041]
[0042] in: For observed trait data, The baseline level for all observations, Phenotypic variations caused by different CSSL lineages. Phenotypic differences are caused by different environments. This is random error.
[0043] Furthermore, library construction analysis is performed on the sample slices, including:
[0044] S2.2.1: Sample collection: Based on the three-dimensional model of the test plant, the target leaf area of the test plant is located and selected, and then cut by single-pulse laser to obtain test sample slices;
[0045] S2.2.2: Library Construction Analysis: The experimental sample slices were placed in pretreatment lysis buffer to obtain a processed cell suspension. Photosynthetic gene probes were added to the cell suspension. Based on the number of cells in the cell suspension and the target sequencing depth, the total sequencing throughput was determined, specifically as follows:
[0046]
[0047] in: Total data volume For cell number, For the target sequencing depth, This is the redundancy coefficient.
[0048] Furthermore, a co-expression network model is constructed, including:
[0049] S3.1.1: Constructing the intergene correlation matrix: Using a preset soft threshold parameter and denoised phenotypic data, the network characteristics under the soft threshold parameter are tested to obtain the scale-free topology fit index corresponding to the soft threshold parameter. Soft threshold parameters greater than the preset index threshold are retained. Simultaneously, the minimum soft threshold parameter is determined from the retained soft threshold parameters, and the intergene correlation matrix is constructed, specifically as follows:
[0050]
[0051] in: Let be the adjacency weights of gene i and gene j in the co-expression network. For gene i, For gene j, The Pearson correlation coefficient is... This is a soft threshold parameter;
[0052] S3.1.2: Module division: Based on the correlation matrix between genes, the topological overlap matrix is obtained, the distance between different genes is determined, the branch height of the tree diagram is set, and each gene is set as a leaf node of the tree diagram to construct the tree diagram. The tree diagram is then divided according to the preset height to obtain tree modules.
[0053] S3.1.3: Module Merging: Based on the feature genes of each tree module, obtain the Pearson correlation coefficients between different feature genes, compare the Pearson correlation coefficients with a preset correlation threshold, and merge the tree modules according to the comparison results, specifically as follows:
[0054] When the Pearson correlation coefficient is not greater than a preset correlation threshold, the tree module corresponding to the Pearson correlation coefficient is not merged; otherwise, the tree module corresponding to the Pearson correlation coefficient is merged, and step S3.1.3 is repeated to obtain the Pearson correlation coefficient again until the Pearson correlation coefficient is not greater than the preset correlation threshold.
[0055] Compared with the prior art, the beneficial effects of the present invention are:
[0056] Firstly, this invention uses multi-generational backcrossing, targeted sequencing, and recombination hotspot analysis to dynamically label the genome and construct a CSSL library with tiered coverage. This effectively compresses QTL intervals, overcoming the problem of ambiguous localization caused by limited recombination events in traditional F2 or RIL populations. At the same time, by modularly segmenting the target QTL intervals and screening the background recovery rate, it can ensure complete coverage of the CSSL library and a pure genetic background, overcoming the technical difficulties of self-crossing degradation and fragment loss during mulberry backcrossing.
[0057] Secondly, this invention separates genotype and environmental effects through repeated planting experiments in multiple environments and a hybrid model, thereby significantly reducing the interference of environmental noise on phenotypic data and solving the problem of poor QTL stability caused by the sensitivity of mulberry trees to the environment.
[0058] Thirdly, this invention combines gene connectivity with phenotypic association strength through a co-expression network model and hub gene scoring, enabling precise screening of key regulatory genes and thus avoiding false positive results caused by epistatic interactions or tight linkages in traditional methods. Attached Figure Description
[0059] Figure 1This is a flowchart illustrating the QTL fine mapping method for important traits of mulberry trees in this invention.
[0060] Figure 2 This is a graph showing the background recovery rate of the backcross population in this invention;
[0061] Figure 3 This is a diagram illustrating the compression effect of targeted sequencing on QTL regions in this invention.
[0062] Figure 4 This is a graph showing the association between gene expression and environmental covariates in this invention. Detailed Implementation
[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0064] Due to the highly heterozygous and complex polyploid background of the mulberry genome, coupled with its sensitivity to environmental changes, QTL mapping using F2 or RILs populations is challenging for effectively resolving complex traits (such as leaf size / thickness) with tight linkages or epistatic interactions involving minor QTLs. While fine mapping can be achieved using chromosome fragment replacement lines, the lack of efficient marker-assisted selection strategies in mulberry leads to incomplete fragment coverage, low background recovery rates, and self-crossing degradation in multiple cycles of inbreeding, thus limiting the precise compression of QTL intervals. This application's technical solution constructs a BC4F2 backcross population through high-heterozygous parental hybridization and multiple generations of backcrossing. Targeted sequencing and recombination hotspot analysis are used for dynamic genome marker application to establish a tiered-coverage CSSL library. Phenotypic data are collected in multi-environmental replicate experiments, and a hybrid model is constructed combining environmental effects and CSSL lines. A co-expression network model is then built based on the denoised phenotypic data, and pivot genes significantly associated with the target trait are screened using topological overlap matrix analysis. This solves the problems of incomplete segment coverage and environmental sensitivity in the localization of complex mulberry tree features, and significantly improves the accuracy of QTL localization.
[0065] Example 1
[0066] refer to Figure 1 This embodiment provides a fine mapping method for QTLs of important traits in mulberry trees based on chromosome segment substitution lines. The fine mapping method for QTLs of important traits in mulberry trees specifically includes the following steps:
[0067] Step S1: Constructing the genotype-phenotype association matrix. This involves screening the backcross population through high-heterozygous parental hybridization and multiple generations of backcrossing, and then labeling the genomes of the selected backcross populations using targeted sequencing and recombination hotspots. Simultaneously, the genomes of the selected backcross populations are segmented and matched with the labeled genomes to construct the genotype-phenotype association matrix. Details are as follows:
[0068] Step S1.1: Construct the BC4F2 backcross population. This involves selecting suitable cultivated varieties based on differences in genomic heterozygosity, and then hybridizing these cultivated varieties with wild species to obtain the BC4F2 backcross population. Specifically, varieties with excellent agronomic traits (e.g., high yield, disease resistance) but weaker target traits (e.g., stress resistance) are used as cultivated varieties, serving as recipients. Wild relatives with prominent target traits (e.g., drought tolerance, insect resistance) and a genomic heterozygosity difference of at least 40% with the selected cultivated varieties are used as wild species, serving as donors. Simultaneously, the cultivated and wild species are initially hybridized to obtain F1 hybrids. These F1 hybrids or backcross progeny are used as the female parent, and the recurrent parent (i.e., the cultivated variety) as the male parent, for three generations of backcrossing to obtain the corresponding BC4F2 backcross population.
[0069] Step S1.2: Enrichment Sequencing. This involves using a pre-set liquid-phase capture probe to perform targeted sequencing on the obtained BC4F2 backcross population, and screening for non-synonymous SNPs in the coding region to obtain sequencing data enriched with the target region. Details are as follows:
[0070] Step S1.2.1: Design the probe. Based on the QTL intervals set during the BC4F2 backcross population screening process in Step S1.1, these intervals are converted into physical distances using a genetic map, and this distance is used as the physical range of the target interval. Simultaneously, the length of the liquid-phase capture probe is set according to the physical range of the target interval and the preset extension range. It is worth noting that in this embodiment, the length of the liquid-phase capture probe is set to 80-120 bp.
[0071] Furthermore, based on the latest assembled version of the mulberry genome, a reference genome is set up. Each exon is covered by at least one liquid-phase capture probe according to all coding region sequences in the reference genome; that is, the location of the liquid-phase capture probe is set according to all coding region sequences in the reference genome. Simultaneously, annotated non-synonymous SNPs are selected from public databases (such as NCBI dbSNP), and liquid-phase capture probes are set up based on these selected non-synonymous SNPs, i.e., the liquid-phase capture probes cover the selected non-synonymous SNPs, thus obtaining the corresponding probe library.
[0072] Step S1.2.2: Liquid phase hybridization. DNA is extracted from the leaf tissue of the BC4F2 backcross population in step S1.1 using the CTAB method. The extracted DNA is then broken into 300–500 bp fragments by sonication, and the corresponding target size fragments are screened out by gel electrophoresis.
[0073] Furthermore, the probe library obtained in step S1.2.1 is hybridized with the screened target-size fragments, and a probe-target DNA complex is obtained using streptavidin magnetic beads. Simultaneously, the obtained probe-target DNA complex is eluted to remove non-binding DNA, thereby enriching the target region fragment and obtaining sequencing data enriched in the target region.
[0074] Step S1.3: Dynamic Marking. Based on the sequencing data enriched with the target intervals obtained in Step S1.2.2, the genetic distances between markers are obtained using the maximum likelihood method. A genetic map is then drawn using JoinMap or R / qtl tools. Further, based on the drawn genetic map, recombination hotspots in the BC4F2 backcross population are identified; intervals with recombination rates greater than 5 cM / Mb are designated as recombination hotspots.
[0075] In this embodiment, the marker density and marker spacing of the recombination hotspots are increased based on their physical coordinates. Specifically, amplicon primers are designed to cover the hotspot regions, and new SNP markers are added with a spacing of no more than 10 kb. For example, if the original marker spacing is 20 kb, 10 new markers are added. Notably, sequencing is performed after multiplex PCR amplification to ensure that the coverage depth of each newly added marker is no less than 30x.
[0076] Step S1.4: Construct the CSSL library. This involves using the genetic map obtained in step S1.3 to divide the target QTL fragment regions in the BC4F2 backcross population into modules, and then using CSSL lines to perform a stepwise coverage of these modules. Specifically:
[0077] Step S1.4.1: Modular Segmentation. Based on the genetic map obtained in Step S1.3, the genetic and physical locations of each SNP / InDel marker are determined. Based on the genotype data of the BC4F2 backcross population obtained in Step S1.1, the target QTL fragment interval is determined. Further, based on the genetic-physical distance correspondence of the species, the target QTL fragment interval is converted into physical distance.
[0078] In this embodiment, based on the recombination hotspots in each target QTL fragment interval, the recombination rate corresponding to each target QTL fragment interval is determined. The obtained recombination rate is then compared with a preset recombination threshold range. Based on the comparison result, modules are divided to obtain the genome map after module division. Specifically:
[0079] When the obtained recombination rate is less than the lower limit of the preset recombination threshold, the target QTL fragment interval corresponding to the recombination rate is divided into segments according to the preset length, and the corresponding segmentation modules are obtained.
[0080] When the obtained recombination rate is within the preset recombination threshold range, the target QTL fragment interval corresponding to the recombination rate is used as the partitioning module.
[0081] When the obtained recombination rate is greater than the upper limit of the preset recombination threshold, the low recombination interval is determined in the recombination hotspot based on genome annotation (such as gene density and conserved sequences). The target QTL fragment interval corresponding to the recombination rate is divided into intervals based on the determined low recombination intervals as the boundary, and the corresponding partitioning module is obtained.
[0082] Step S1.4.2: CSSL lineage coverage. Based on the segment intervals defined in step S1.4.1, BC4F2 individuals carrying those segment intervals are selected from the BC4F2 backcross population. These selected BC4F2 individuals are then pre-screened using module boundary markers to obtain the pre-screened BC4F2 individuals.
[0083] Furthermore, within each partitioning module, at least three coverage lines are set up for tiered coverage, and targeted sequencing is performed on candidate lines to determine the actual coverage range of each coverage line. It is worth noting that during the tiered coverage process, for target QTL fragment intervals where the recombination rate exceeds the upper limit of the preset recombination threshold, the coverage coefficient is increased according to a preset step size. The tiered coverage modules within the target QTL fragment intervals are then covered using the increased coverage lines and the original coverage lines.
[0084] Step S2: Phenotypic Determination and Modeling. This involves collecting samples from the obtained experimental plants, determining the corresponding environmental effects from these plants, and combining these environmental effects with the CSSL line covered in step S1.4.2 to obtain a hybrid model. Details are as follows:
[0085] Step S2.1: Phenotypic Data Collection. This involves establishing four representative ecological zones based on differences in climate type (e.g., arid, humid, high-altitude, plains), soil type gradients, and average annual temperature differences. These zones are: the Yangtze River Basin (humid), the Loess Plateau (arid), the Yunnan-Guizhou Plateau (high altitude), and the Northeast Plain (low temperature). Specifically, repeated planting experiments are conducted for at least two years in each of the four representative ecological zones to obtain the corresponding experimental plants.
[0086] Furthermore, using a portable laser scanner, the experimental plants were scanned in three dimensions every two weeks to obtain time-series data on the three-dimensional morphology of the leaves, namely leaf tilt angle, leaf area index, and daily plant height growth. At the same time, the core parameters of the experimental plants (such as ΦPSII (photosystem II quantum efficiency) and NPQ (non-photochemical quenching)) and light response curves were obtained using a pulse-modulated chlorophyll fluorometer.
[0087] Step S2.2: Single-cell transcription library construction. This involves obtaining sample sections from the experimental plants acquired in step S2.1, and then performing library construction analysis on these sample sections using photosynthetic gene probes. Details are as follows:
[0088] Step S2.2.1: Sample Collection. This involves wrapping the target leaf area of the test plant with aluminum foil pre-cooled by liquid nitrogen, simultaneously cutting the sample with a scalpel, and storing the sample at -80℃. Further, 5% PVP is added to the cryogenically stored sample to prevent plant cell collapse, and polylysine slides are used to enhance adhesion. The sample sections are then obtained using a cryostat microtome.
[0089] Furthermore, based on the three-dimensional model of the experimental plant, areas with leaf tilt angles greater than 45° were located, and the outlines of palisade cells in the located areas were delineated under a 40x microscope, avoiding vascular bundles during the delineation process. Simultaneously, the delineated palisade cell outlines were cut using a single-pulse laser to obtain experimental sample slices.
[0090] Step S2.2.2: Library construction and analysis. At least three consecutive sample slices located in the middle are laser-microdissected, and the dissected sample slices are placed in EP tubes containing 2 μL of RLT Plus buffer. After being flash-frozen in liquid nitrogen, they are stored at -80°C, fixed with 4% glutaraldehyde, and then embedded in resin and OTC.
[0091] Furthermore, cellulase at a final concentration of 0.5 U / μL and pectinase at a final concentration of 0.2 U / μL were both added to the basal lysis buffer, and the mixture was lysed at 4°C for 15 minutes to obtain the corresponding pretreated lysis buffer. Simultaneously, the pretreated sample slices were placed in the pretreated lysis buffer, gently pipetted, and then filtered through a 40 μm filter to remove cell debris, obtaining the treated cell suspension.
[0092] In this embodiment, photosynthetic gene probes are added to the obtained processed cell suspension, and the total sequencing throughput is determined based on the number of cells in the processed cell suspension and the target sequencing depth, specifically as follows:
[0093]
[0094] in: Total data volume For cell number, For the target sequencing depth, This is the redundancy coefficient.
[0095] Specifically, the concentration of cell suspensions in each treatment group was calibrated to 1000 cells / μL, and BioLegend Hashtag antibody was added to the cell suspensions in each calibrated treatment group. The treatment groups were divided into a control group, a drought treatment group, and a salt stress treatment group, and each group was fluorescently labeled with FITC, PE, and APC, respectively. Specifically, FITC was labeled green, PE was labeled orange-yellow, and APC was labeled red.
[0096] Step S2.3: Modeling and Analysis. Based on the environmental effects (e.g., drought and salinity) obtained in Step S2.2.2 and the genotypic effects (e.g., CSSL line) obtained in Step S1.4.2, a hybrid model is constructed, specifically as follows:
[0097]
[0098] in: For observed trait data, The baseline level for all observations, Phenotypic variations caused by different CSSL lineages. Phenotypic differences are caused by different environments. This is random error.
[0099] refer to Figure 4 , Figure 4 This is a correlation diagram between gene expression and environmental covariates in this embodiment. Figure 4It can be seen that under arid conditions in the Yangtze River Basin and the Loess Plateau, most genes (such as Gene_1 to Gene_10) show negative correlations (blue blocks), with Gene_3 and Gene_7 being the most prominent (correlation coefficient approximately -0.9). Therefore, these genes may be suppressed by drought stress. Meanwhile, some genes (such as Gene_5) maintain positive correlations under drought conditions (red blocks), suggesting that they may participate in drought resistance responses.
[0100] Furthermore, under the low-temperature conditions of the Northeast Plain, gene expression generally shows a weak negative correlation (for example, Gene_4 is about -0.3), but Gene_9 and Gene_10 show neutral (close to 0), which may indicate that they are not sensitive to low temperatures.
[0101] Step S3: Gene Screening. This involves denoising the original phenotypic data using a convolutional neural network model and constructing a co-expression network model. Then, hub genes are screened using this co-expression network model and the hybrid model constructed in step S2.3. Details are as follows:
[0102] Step S3.1: Construct the co-expression network model. This involves constructing a co-expression network model based on the denoised phenotypic data and several preset soft threshold parameters. Specifically:
[0103] Step S3.1.1: Construct the gene correlation matrix. This involves using multi-environmental phenotypic data (such as time-series data of leaf area and plant height) and environmental covariates (such as temperature and humidity) as input to the convolutional neural network model, and outputting denoised phenotypic data.
[0104] Furthermore, by using preset soft threshold parameters and denoised phenotypic data, the network characteristics under different soft threshold parameters are tested, and the scale-free topology fitting index corresponding to each soft threshold parameter is obtained. Simultaneously, the obtained scale-free topology fitting index is compared with a preset index threshold, and based on the comparison results, the preset soft threshold parameters are filtered. Specifically:
[0105] When the obtained scale-free topology fitting index is greater than the preset index threshold, the preset soft threshold parameter corresponding to the scale-free topology fitting index is retained. Conversely, when the obtained scale-free topology fitting index is not greater than the preset index threshold, the preset soft threshold parameter corresponding to the scale-free topology fitting index is deleted.
[0106] Specifically, based on the retained preset soft threshold parameters, a minimum soft threshold parameter is determined, and this minimum soft threshold parameter is the determined soft threshold parameter. Simultaneously, based on the determined soft threshold parameter, a gene correlation matrix is constructed, as follows:
[0107]
[0108] in: Let be the adjacency weights of gene i and gene j in the co-expression network. For gene i, For gene j, The Pearson correlation coefficient is... This is the soft threshold parameter.
[0109] Step S3.1.2: Module Partitioning. This involves obtaining the corresponding topological overlap matrix based on the gene correlation matrix constructed in step S1.1.1, specifically as follows:
[0110]
[0111] in: The topological overlap measure between gene i and gene j. Let be the adjacency weights of gene i and gene j in the co-expression network. Let be the adjacency weights of gene i and gene k in the co-expression network. Let K and J be the adjacency weights of gene k and gene j in the co-expression network. , , For gene indexing.
[0112] Furthermore, based on the obtained topological overlap matrix, the spacing between different genes is determined, and the branch height of the dendrogram is set according to the determined spacing. Simultaneously, each gene is set as a leaf node in the dendrogram, thus constructing the corresponding dendrogram. At the same time, the constructed dendrogram is divided according to the preset height to obtain the corresponding tree modules.
[0113] Step S3.1.3: Module Merging. Based on the tree-like modules obtained in Step S3.1.2, the expression matrix (i.e., sample * gene) of all genes within each tree-like module is extracted, and PCA analysis is performed on the extracted expression matrices, with PC1 used as the feature gene. Further, based on the obtained feature genes, the Pearson correlation coefficients between different feature genes are obtained, and these Pearson correlation coefficients are compared with a preset correlation threshold. Based on the comparison results, the tree-like modules are merged, specifically as follows:
[0114] If the obtained Pearson correlation coefficient is not greater than the preset correlation threshold, the tree module corresponding to the Pearson correlation coefficient will not be merged. Conversely, if the obtained Pearson correlation coefficient is greater than the preset correlation threshold, the tree module corresponding to the Pearson correlation coefficient will be merged, and step S3.1.3 will be repeated to obtain the corresponding Pearson correlation coefficient again until the obtained Pearson correlation coefficient is not greater than the preset correlation threshold.
[0115] Step S3.2: Phenotypic Association. Based on the dendritic modules obtained in Step S3.1.3, the first principal component PC1 of each dendritic module is obtained and used as the characteristic gene of each dendritic module. Simultaneously, based on the environmental effects (e.g., drought and salinity) obtained in Step S2.2.2 and the characteristic genes of each dendritic module, the correlation coefficient between the characteristic genes and the phenotype is determined, specifically as follows:
[0116]
[0117] in: The Pearson correlation coefficient is... Let m be the expression value of the module feature gene for the m-th sample. The mean expression of module-specific genes across all samples. Let m be the phenotypic value of the m-th sample. This represents the phenotypic mean for all samples.
[0118] Furthermore, the obtained Pearson correlation coefficients are compared with preset correlation coefficient thresholds, and the Pearson correlation coefficients are filtered based on the comparison. Specifically:
[0119] If the obtained Pearson correlation coefficient is greater than the preset correlation coefficient threshold, the obtained Pearson correlation coefficient is retained. Conversely, if the obtained Pearson correlation coefficient is not greater than the preset correlation coefficient threshold, the obtained Pearson correlation coefficient is removed.
[0120] Step S3.3: Gene Screening. Based on the Pearson correlation coefficients retained in Step S3.2, the topological overlap matrix corresponding to the feature genes is determined. Simultaneously, based on the determined topological overlap matrix, the sum of the corresponding connectivity degrees is obtained and scaled to the range of 0-1. Further, based on the expression level of a single gene in all samples (e.g., TPM value) and target phenotypic data (e.g., leaf area and drought resistance index), the Pearson correlation coefficient corresponding to each gene is determined, and the association strength between the corresponding gene expression and the target phenotypic is obtained by calculating the absolute value.
[0121] In this embodiment, the corresponding hub gene comprehensive score is obtained based on the determined sum of connectivity and association strength, specifically as follows:
[0122]
[0123] in: For the comprehensive score of pivot genes, For connectivity weights, For the degree of connectivity within a module, For association weights, This indicates gene significance.
[0124] Furthermore, the obtained comprehensive score of the hub gene is compared with a preset scoring threshold, and the selected hub genes are determined based on the comparison results, specifically as follows:
[0125] When the obtained comprehensive score of a hub gene is greater than the preset score threshold, the hub gene corresponding to that comprehensive score is the selected hub gene. Conversely, when the obtained comprehensive score of a hub gene is not greater than the preset score threshold, the hub gene corresponding to that comprehensive score is not the selected hub gene.
[0126] Example 2
[0127] This embodiment provides a method for fine mapping of QTLs for important traits in mulberry trees based on chromosome segment substitution lines. The specific implementation method is the same as in Embodiment 1. In step S1.1, during the backcrossing of F1 hybrids or backcross progeny with cultivated varieties, genotyping is performed to obtain the background reversion rate for each individual. Based on the obtained background reversion rate, the backcross population is screened. Simultaneously, target region screening is performed on the screened backcross population to obtain a highly purified BC4F2 backcross population. The invention will be illustrated below with specific examples of this embodiment.
[0128] In this embodiment, a highly purified BC4F2 backcross population was obtained, as detailed below:
[0129] Step S1.1.1: Constructing F1 generation hybrids. This involves artificially pollinating the cultivated species as the female parent and the wild species as the male parent to obtain F1 hybrids. Simultaneously, SNP genotyping is performed on each F1 hybrid to obtain the corresponding heterozygous locus ratio. The obtained heterozygous locus ratio is then compared with a preset ratio threshold. Based on the comparison results, the obtained F1 hybrids are initially screened to obtain the pre-screened F1 hybrids. Specifically:
[0130] When the proportion of heterozygous sites obtained is greater than a preset threshold, the F1 hybrids corresponding to that proportion are retained. Conversely, when the proportion of heterozygous sites obtained is not greater than the preset threshold, the F1 hybrids corresponding to that proportion are removed. In other words, all the retained F1 hybrids are the F1 hybrids after the initial screening.
[0131] Furthermore, based on the F1 hybrids after initial screening, the corresponding pollen viability, seed set rate, and seed quality are obtained, and the corresponding comprehensive score is determined, specifically as follows:
[0132]
[0133] in: For the overall score, As a weighting of pollen viability, For pollen vitality, As the weight of the fruit set rate, For the fruit setting rate, As a weight for seed quality, Seed quality.
[0134] In this embodiment, the obtained comprehensive score is compared with a preset score threshold, and pollen viability, seed set rate, and seed quality are also compared with preset individual score thresholds. Based on the comparison, the final F1 hybrid is determined. Specifically:
[0135] If the overall score obtained is less than a preset score threshold, or if any of the scores for pollen viability, seed set rate, or seed quality is less than a preset individual score threshold, the corresponding F1 hybrid after initial screening will be removed. Conversely, if the scores are higher than the preset threshold, the corresponding F1 hybrid after initial screening will be retained. In other words, all the retained F1 hybrids after initial screening are the final F1 hybrids.
[0136] Step S1.1.2: Backcrossing and Screening. The final F1 hybrid obtained in step S1.1.1 is used as the female parent, and the cultivated variety is used as the male parent for artificial pollination to obtain the BC1 population. Simultaneously, based on the obtained BC1 population, the background recovery rate and non-target area donor residual rate of each BC1 individual are determined. Further, the obtained background recovery rate is compared with a preset recovery threshold, and the obtained non-target area donor residual rate is compared with a preset residual threshold. Based on the comparison results, the selected BC1 individuals are determined. Specifically:
[0137] If the obtained background response rate is not less than the preset response threshold, and the obtained non-target area donor residual rate is not greater than the preset residual threshold, then the corresponding BC1 individuals are retained. Otherwise, the corresponding BC1 individuals are removed. In other words, all the retained BC1 individuals are the selected BC1 individuals.
[0138] In this embodiment, the selected BC1 individuals are used as the female parent, and the cultivated variety is used as the male parent for artificial pollination to obtain the BC2 population. Then, based on the selection criteria of the BC1 population, each BC2 individual in the BC2 population is selected to obtain the selected BC2 individuals.
[0139] Furthermore, the selected BC2 individuals were used as the female parent, and the cultivated variety was used as the male parent for artificial pollination to obtain the BC3 population. Then, based on the selection criteria of the BC1 population, each BC3 individual in the BC3 population was selected to obtain the selected BC3 individuals.
[0140] Furthermore, the selected BC3 individuals were used as the female parent, and the cultivated variety was used as the male parent for artificial pollination to obtain the BC4 population. Then, based on the selection criteria of the BC1 population, each BC4 individual in the BC4 population was selected to obtain the selected BC4 individuals.
[0141] Step S1.1.3: Screening of BC4 individuals. This involves obtaining the contribution rate of the QTL from the genetic map of the BC3 individuals screened in Step S1.1.2, comparing the obtained contribution rate with a preset contribution threshold, and determining the major-effect QTL based on the comparison results. Specifically:
[0142] When the obtained contribution rate is greater than the preset contribution threshold, the QTL corresponding to that contribution rate is the major effect QTL. Conversely, when the obtained contribution rate is not greater than the preset contribution threshold, the QTL corresponding to that contribution rate is the minor effect QTL.
[0143] Furthermore, based on the identified major QTL, the major QTL interval is set as the target interval, and the target interval is marked by SNP markers to cover the gene coding region and non-coding regulatory region.
[0144] In this embodiment, based on the genotype data marked in the target interval and the whole-genome SNP data of the BC4 individuals screened in step S1.1.2, BC4 individuals retaining donor fragments within the target interval are screened to obtain preliminary BC4 individuals. Simultaneously, candidate BC4 individuals are determined based on the donor residual rate in the non-target interval of the preliminary BC4 individuals. Specifically, the donor residual rate in the non-target interval of the preliminary BC4 individuals is compared with a preset donor residual threshold, and candidate BC4 individuals are determined based on the comparison result.
[0145] If the obtained donor residual rate in the non-target interval is greater than the preset donor residual threshold, the BC4 individuals corresponding to that non-target interval donor residual rate are removed. Conversely, if the obtained donor residual rate in the non-target interval is not greater than the preset donor residual threshold, the BC4 individuals corresponding to that non-target interval donor residual rate are retained. In other words, the retained BC4 individuals are the candidate BC4 individuals.
[0146] Step S1.1.4: Introgression hybridization. Based on the different QTL gene fragments carried, the candidate BC4 individuals obtained in Step S1.1.3 are divided to obtain BC4 individuals with at least two different parents. Specifically, the BC4 individuals with two different parents are crossed reciprocally to obtain the corresponding F2 hybrids. Further, according to the screening conditions of the BC1 population in Step S1.1.2, each F2 individual in the F2 hybrids is screened to obtain the selected F2 individuals, which constitute the highly purified BC4F2 backcross population.
[0147] refer to Figure 2 , Figure 2 This is a graph showing the background response rate of the backcross population in this embodiment. Figure 2 It can be seen that in all backcross generations (F11 hybrid-BC4 population), the background recovery rate of this scheme is always higher than that of the traditional method, and a 90% recovery rate can be achieved in the BC3 population, reaching the threshold one generation earlier than the traditional method.
[0148] refer to Figure 3 , Figure 3 This is a diagram illustrating the compression effect of targeted sequencing on QTL regions in this embodiment. Figure 3 It can be seen that the QTL interval size of the traditional method is 3000kb, while the QTL interval size of this scheme is 80kb, which is reduced by 37.5 times.
[0149] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended embodiments and their equivalents.
Claims
1. A method for fine mapping of QTLs of important traits in mulberry trees based on chromosome segment substitution lines, characterized in that, Including: S1: Constructing a genotype-phenotype association matrix: The backcross population was screened through high-heterozygous parental hybridization and multiple generations of backcrossing. The genomes of the selected backcross population were then labeled using targeted sequencing and recombination hotspots, including: S1.1: Constructing the BC4F2 backcross population: Using the cultivated species as the female parent and the wild species as the male parent, initial hybridization is performed to obtain the F1 hybrid. At the same time, the F1 hybrid is used as the female parent and the cultivated species as the male parent, and three generations of backcrossing are performed to obtain the BC4F2 backcross population. S1.2: Enriched Sequencing: Using a set liquid-phase capture probe, targeted sequencing is performed on the BC4F2 backcross population, and non-synonymous SNPs in the coding region are screened to obtain sequencing data enriched in the target region. S1.3: Dynamic labeling: The genetic distance of the sequencing data is obtained by the maximum likelihood method, and a genetic map is drawn to identify the recombination hotspots of the BC4F2 backcross population; S1.4: Constructing the CSSL library: Using the genetic map, the target QTL fragment intervals in the BC4F2 backcross population are divided into modules, and the divided modules are covered in a stepwise manner using the CSSL line; S2: Phenotypic modeling: Based on the experimental plants, samples were collected, and a mixture model was obtained based on the environmental effects of the experimental plants and the CSSL covering line. S3: Gene Screening: Hub gene screening is performed using constructed co-expression network models and hybrid models, including: S3.1: Constructing a co-expression network model: Based on the denoised phenotypic data and multiple preset soft threshold parameters, a co-expression network model is constructed. S3.2: Phenotypic Association: Based on characteristic genes and environmental effects, determine the Pearson correlation coefficient, and compare the Pearson correlation coefficient with a preset correlation coefficient threshold to filter the Pearson correlation coefficients, specifically as follows: If the Pearson correlation coefficient is greater than a preset correlation coefficient threshold, the Pearson correlation coefficient is retained; otherwise, the Pearson correlation coefficient is removed. S3.3: Gene Screening: Based on the retained Pearson correlation coefficients, individual gene expression levels, and target phenotypic data, the sum of connectivity and association strength are obtained, and a comprehensive score for hub genes is determined. Simultaneously, the comprehensive score for hub genes is compared with a preset score threshold to determine the selected hub genes. Specifically: When the comprehensive score of the hub gene is greater than the preset score threshold, the hub gene corresponding to the comprehensive score is the selected hub gene; otherwise, the hub gene corresponding to the comprehensive score is not the selected hub gene.
2. The method for fine mapping of important mulberry traits based on chromosome segment substitution lines according to claim 1, characterized in that, The BC4F2 backcross population obtained includes: S1.1.1: Constructing F1 generation hybrids: Using SNP genotyping, the proportion of heterozygous sites in the F1 hybrids is obtained. Simultaneously, F1 hybrids with a proportion of heterozygous sites greater than a preset threshold are retained to obtain the initial F1 hybrids. The comprehensive score of the initial F1 hybrids is then obtained. Based on the comprehensive score and a preset score threshold, the final F1 hybrids are determined. Specifically: If the overall score is less than a preset score threshold, or if any of the pollen viability, seed setting rate, or seed quality scores are less than a preset single-item score threshold, then the corresponding F1 hybrid after initial screening will be removed; otherwise, the corresponding F1 hybrid after initial screening will be retained. S1.1.2: Backcrossing and screening: The final F1 hybrid is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC1 population. At the same time, the BC1 population is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC2 population. The BC2 population is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC3 population. The BC3 population is used as the female parent and the cultivated variety as the male parent for backcrossing to obtain the BC4 population. S1.1.3: BC4 Individual Screening: Based on the genetic map of the BC3 population, the major-effect QTL target interval is determined. Simultaneously, based on the genotype data marked in the major-effect QTL target interval and the whole-genome SNP data of the BC4 population, preliminary BC4 individuals are obtained. The donor residual rate in the non-target interval of the preliminary BC4 individuals is compared with a preset donor residual threshold. Based on the comparison results, candidate BC4 individuals are determined, specifically as follows: When the donor residual rate in the non-target interval is greater than the preset donor residual threshold, the corresponding BC4 individuals in the initial screening are removed; otherwise, the corresponding BC4 individuals in the initial screening are retained. S1.1.4: Introgression hybridization: Based on the QTL gene fragments carried, the candidate BC4 individuals are divided, and the divided candidate BC4 individuals are subjected to reciprocal crosses to obtain F2 hybrids. Based on the background recovery rate and non-target region donor residual rate of the F2 hybrids, the F2 hybrids are screened to obtain the BC4F2 backcross population.
3. The method for fine mapping of important mulberry traits based on chromosome segment substitution lines according to claim 2, characterized in that, During the backcross screening process, the backcross individuals are screened based on the background recovery rate and non-target region donor residual rate of the backcross individuals in each backcross population, and the screened backcross individuals are used as the parent lines for the next round of backcrossing. Specifically: When the background recovery rate is not less than the preset recovery threshold and the donor residual rate in the target area is not greater than the preset residual threshold, the corresponding backcrossing body will be retained. Conversely, the corresponding backcross individuals are removed.
4. The method for fine mapping of important mulberry traits based on chromosome segment substitution lines according to claim 1, characterized in that, The sequencing data enriched in the target region were obtained, including: S1.2.1: Probe design: Based on the target region of the major QTL and the preset extension range, the length of the liquid phase capture probe is set. At the same time, based on all coding region sequences in the reference genome, the location of the liquid phase capture probe is set. And based on the annotated non-synonymous SNPs screened from the public database, the probe library is obtained. S1.2.2: Liquid phase hybridization: DNA was extracted from the leaf tissue of the BC4F2 backcross population using the CTAB method. Simultaneously, the DNA was disrupted by ultrasound and gel electrophoresis, and target-sized fragments were screened out. The probe library and the target-sized fragments were hybridized using streptavidin magnetic beads to obtain probe-target DNA complexes.
5. The method for fine mapping of important mulberry traits based on chromosome segment substitution lines according to claim 1, characterized in that, The CSSL system is used to perform tiered coverage of the divided modules, including: S1.4.1: Modular Segmentation: Based on the genotype data of the BC4F2 backcross population, the target QTL fragment interval is determined, and the recombination rate of the target QTL fragment interval is compared with a preset recombination threshold range. Based on the comparison result, modular segmentation is performed, specifically as follows: When the recombination rate is less than the lower limit of the preset recombination threshold, the target QTL fragment interval is divided into segments according to the preset length. When the recombination rate is within the range of the preset recombination threshold, the target QTL fragment interval is used as a segmentation module. When the recombination rate is greater than the upper limit of the preset recombination threshold, the target QTL fragment interval is further divided into segments according to the determined low recombination sub-intervals. S1.4.2: CSSL line coverage: Based on the divided modules, the BC4F2 backcross population is screened, and at least three coverage lines are set in each module to perform stepwise coverage. Targeted sequencing is performed on the candidate lines to determine the actual coverage range of each coverage line.
6. The method for fine mapping of important mulberry traits based on chromosome segment substitution lines according to claim 1, characterized in that, The resulting hybrid model includes: S2.1: Phenotypic Collection: Based on differences in climate type, soil type gradient and average annual temperature, representative ecological zones are set up, and repeated planting experiments are conducted in the representative ecological zones for no less than 2 years to obtain experimental plants; S2.2: Single-cell transcription library construction: Sample slices were obtained from the experimental plants, and the sample slices were analyzed for library construction using photosynthetic gene probes; S2.3: Modeling and Analysis: Based on the environmental effects of the experimental plants and the CSSL line, a hybrid model was constructed, specifically as follows: ; in: For observed trait data, The baseline level for all observations, Phenotypic variations caused by different CSSL lineages. Phenotypic differences are caused by different environments. This is random error.
7. The method for fine mapping of important mulberry traits based on chromosome segment substitution lines according to claim 6, characterized in that, Library construction analysis of the sample slices includes: S2.2.1: Sample collection: Based on the three-dimensional model of the test plant, the target leaf area of the test plant is located and selected, and then cut by single-pulse laser to obtain test sample slices; S2.2.2: Library Construction Analysis: The experimental sample slices were placed in pretreatment lysis buffer to obtain a processed cell suspension. Photosynthetic gene probes were added to the cell suspension. Based on the number of cells in the cell suspension and the target sequencing depth, the total sequencing throughput was determined, specifically as follows: ; in: Total data volume For cell number, For the target sequencing depth, This is the redundancy coefficient.
8. The method for fine mapping of important mulberry traits based on chromosome segment substitution lines according to claim 1, characterized in that, The co-expression network model is constructed, including: S3.1.1: Constructing the intergene correlation matrix: Using a preset soft threshold parameter and denoised phenotypic data, the network characteristics under the soft threshold parameter are tested to obtain the scale-free topology fit index corresponding to the soft threshold parameter. Soft threshold parameters greater than the preset index threshold are retained. Simultaneously, the minimum soft threshold parameter is determined from the retained soft threshold parameters, and the intergene correlation matrix is constructed, specifically as follows: ; in: Let be the adjacency weights of gene i and gene j in the co-expression network. For gene i, For gene j, The Pearson correlation coefficient is... This is a soft threshold parameter; S3.1.2: Module division: Based on the correlation matrix between genes, the topological overlap matrix is obtained, the distance between different genes is determined, the branch height of the tree diagram is set, and each gene is set as a leaf node of the tree diagram to construct the tree diagram. The tree diagram is then divided according to the preset height to obtain tree modules. S3.1.3: Module Merging: Based on the feature genes of each tree module, obtain the Pearson correlation coefficients between different feature genes, compare the Pearson correlation coefficients with a preset correlation threshold, and merge the tree modules according to the comparison results, specifically as follows: When the Pearson correlation coefficient is not greater than a preset correlation threshold, the tree module corresponding to the Pearson correlation coefficient is not merged; otherwise, the tree module corresponding to the Pearson correlation coefficient is merged, and step S3.1.3 is repeated to obtain the Pearson correlation coefficient again until the Pearson correlation coefficient is not greater than the preset correlation threshold.
Citation Information
Patent Citations
Method for rice gene location and molecular breeding
CN107365870A
Method by adopting molecular marker-assisted backcross to improve gibberellic disease expansion resistance of wheat
CN102599047A
Method for exploring salt-tolerant hub genes of salix matsudana
CN110564884A