Soybean ideal plant type directional improvement method based on br synthetase gene editing

By using a multi-omics data-driven target prediction model and a high-throughput phenotypic monitoring platform, combined with CRISPR-Cas12a gene editing, the problems of long cycle, high cost, and large design blindness in soybean breeding have been solved, realizing precise and targeted improvement of ideal soybean plant type and efficient breeding.

CN122177229APending Publication Date: 2026-06-09NANYANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANYANG NORMAL UNIV
Filing Date
2026-02-08
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing soybean gene editing breeding methods suffer from problems such as long cycles, high costs, and significant design blindness. It is difficult to accurately predict the impact of gene editing on plant dynamics and yield composition in the early stages, especially given the inconsistent performance in different ecological regions.

Method used

A multi-omics data-driven target prediction model was constructed. Combined with the CRISPR-Cas12a gene editing vector, high-throughput non-destructive phenotypic monitoring was carried out in a controlled environment growth chamber using Agrobacterium-mediated cotyledon transformation. Positive single plants that meet the preset ideal plant type were screened through multispectral imaging and lidar scanning. The prediction accuracy was improved through a closed-loop feedback optimization mechanism.

Benefits of technology

It has achieved precise and targeted improvement of the ideal plant type of soybean, shortened the breeding cycle, reduced labor, land and time costs, obtained new germplasm with moderate plant height, thick stems, compact branching and high yield, and provided genetically stable homozygous lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177229A_ABST
    Figure CN122177229A_ABST
Patent Text Reader

Abstract

This invention relates to the field of biotechnology and discloses a method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing. The method includes: integrating multi-omics data to construct a BR synthase gene regulatory network prediction model to accurately identify key functional sites affecting traits such as plant height and internode length; designing a CRISPR-Cas12a vector to edit target sites; achieving high-throughput, non-destructive phenotypic monitoring of T0 generation plants through a controlled-environment growth chamber combined with multispectral and lidar technologies; associating genotype and phenotypic data to screen for ideal plant architecture individuals, and optimizing the prediction model through data feedback to form a closed-loop iterative system. This invention shortens the breeding cycle, improves the predictability of editing, and efficiently obtains new soybean germplasm with compact growth, lodging resistance, and a high harvest index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to a method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing. Background Technology

[0002] With increasing global pressure on food security and rising demands for sustainable agricultural development, crop plant architecture improvement has become one of the core directions of modern breeding. As an important dual-purpose crop for grain and oil, soybean plant architecture directly affects population photosynthetic efficiency, lodging resistance, and final yield. Traditional breeding relies on phenotypic selection and backcrossing, which are time-consuming and inefficient; while marker-assisted selection has improved accuracy, it is still limited by genetic linkage baggage and the complexity of multi-gene interactions.

[0003] In recent years, the development of gene editing technology, especially the CRISPR / Cas system, has provided a new pathway for targeted regulation of key developmental genes. Among them, the brassinolide (BR) signaling pathway has become an important target for ideal plant architecture design due to its decisive role in cell elongation, vascular differentiation and plant architecture.

[0004] Targeted editing of the BR synthase gene can precisely regulate soybean plant height and construct a compact plant architecture by modulating endogenous BR levels. This strategy theoretically avoids the environmental residues and cost issues associated with exogenous hormone application and possesses genetic stability. However, BR signaling exhibits high dose sensitivity and tissue specificity; even minor changes in expression can lead to excessive dwarfing, impaired reproductive development, or decreased photosynthetic efficiency. This presents significant challenges in selecting the editing site, editing type (knockout, base substitution, promoter modification), and editing intensity.

[0005] In current technologies, the design of BR synthase gene editing protocols and their validation in various field environments typically require 3 to 5 growing seasons, involving significant investment of manpower, land, and time. Furthermore, due to the lack of prior prediction of genotype-phenotype-environment interactions, editing strategies often exhibit a degree of uncertainty. Even with the integration of high-throughput sequencing and transcriptome analysis, it remains difficult to accurately predict the impact of specific editing events on plant architecture dynamics and yield composition throughout the entire growth cycle. Especially under different ecological zones, densities, and water and fertilizer conditions, the same edited line may exhibit drastically different agronomical behaviors, further amplifying the uncertainty and resource consumption of field validation. Therefore, a forward-looking design method integrating high-quality genomic information and digital simulation technology is urgently needed to achieve virtual screening and optimization of BR editing protocols, fundamentally shortening the breeding cycle and improving the accuracy of targeted improvement. Summary of the Invention

[0006] This invention provides a method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing, comprising: A multi-source heterogeneous dataset of soybean throughout its entire growth period was obtained. The multi-source heterogeneous dataset includes whole-genome resequencing data of soybean core germplasm resources, transcriptome data of different tissues and organs at key developmental stages, ATAC-seq data of open chromatin regions, and three-dimensional phenotypic data of corresponding plants under standardized conditions. Based on the aforementioned multi-source heterogeneous dataset, a prediction model of the brassinolide synthase gene regulatory network was constructed to identify key functional sites in the BR synthase encoding genes and their upstream cis-acting elements that have regulatory effects on plant height, internode length, stem strength and branching angle. Based on the output of the prediction model, a CRISPR-Cas12a gene editing vector targeting the key functional sites was designed and synthesized. The vector contains a pair of specific guide RNA sequences whose target sequences are located within conserved cis elements in the promoter region of the BR synthase gene or on key exons of the coding region. The gene editing vector was introduced into soybean recipient material using Agrobacterium-mediated cotyledon node transformation to obtain a T0 generation transgenic plant population; High-throughput non-destructive phenotypic monitoring was conducted on the T0 generation plant population in a controlled environment growth chamber. The monitoring was carried out by combining a multispectral imaging system with lidar scanning to obtain the plant height dynamic growth curve, stem diameter change rate, center of gravity shift angle and biomass spatial distribution map in real time. The high-throughput phenotypic data were correlated with the corresponding gene editing event information to screen out positive single plants whose plant phenotypic parameters met the preset ideal threshold range. These plants were then used as parents for self-pollination to obtain genetically stable T2 generation homozygous lines.

[0007] Preferably, the construction of the brassinosteroid synthase gene regulatory network prediction model includes: SNP and InDel variant detection were performed on whole-genome resequencing data, and genome-wide association analysis was conducted to locate genomic regions associated with plant height and internode length phenotypes. Differential expression analysis was performed on transcriptome data to screen for BR synthase-related genes whose expression levels differed by more than 2 fold between dwarf mutants and wild-type controls and whose corrected P-value was less than 0.01. By integrating ATAC-seq data, we identified chromatin regions that were specifically opened in the shoot apical meristem and internode elongation zone, and then performed co-localization analysis on these regions with regions located by genome-wide association analysis and promoter regions of differentially expressed genes. Based on the colocalization results, a multi-task regression model was trained using the gradient boosting decision tree algorithm. The input features of the model are the sequence context, chromatin accessibility score and linkage disequilibrium coefficient of the candidate site, and the output is the predicted effect value of the site on the four core plant type traits.

[0008] Preferably, the design and synthesis of the CRISPR-Cas12a gene editing vector targeting the key functional sites includes: Based on the effect values ​​output by the prediction model, the top 5 sites with the highest absolute effect values ​​were selected as editing targets. For each target, the PAM sequence preference TTTV of the Cas12a protein was used to search for guide RNA sequences with an off-target score of less than 0.3 within 200 bases upstream and downstream of the target. The two pairs of guide RNA sequences selected were cloned into binary vectors containing the soybean U6 promoter and the Cas12a expression cassette, respectively. The binary vectors also contain a red fluorescent protein reporter gene driven by an embryo-specific promoter.

[0009] Preferably, the high-throughput non-destructive phenotypic monitoring of the T0 generation plant population within a controlled environment growth chamber includes: The T0 generation plant population was planted at a specified density in a standardized nutrient solution cultivation system. From the date of seedling emergence, a fully automatic phenotypic acquisition process is initiated at regular intervals. The fully automatic phenotypic acquisition process first obtains a top view of the canopy using a high-resolution RGB camera installed on the top, and calculates the plant's projected area and branch angle. Subsequently, stalk texture images were acquired using a side-mounted near-infrared camera, and the stalk mechanical strength index was calculated using the image gray-level co-occurrence matrix algorithm. Finally, the plant was scanned 360 degrees using a rotating lidar to reconstruct a three-dimensional point cloud model, from which plant height, internode length, center of gravity height, and aboveground biomass volume were extracted.

[0010] Preferably, the step of performing correlation analysis between the high-throughput phenotypic data and the corresponding gene editing event information includes: For each T0 generation plant, the target site region was amplified by leaf PCR and Sanger sequencing was performed to determine its specific genotype, including insertion / deletion type and heterozygosity. The sequencing results were matched one-to-one with the final phenotypic data of the same plant collected in the growth chamber. Set a preset threshold range for the ideal plant shape; Plants whose phenotypic parameters all fall within the threshold range and whose genotype is the expected editing type are selected as candidate superior plants.

[0011] Preferably, the method further includes a closed-loop feedback optimization step: The phenotypic and genotypic data of the selected candidate superior plants are fed back into the brassinolide synthase gene regulatory network prediction model as new training samples. Online learning algorithms are used to incrementally update the model parameters in order to improve the model's accuracy in predicting future editing events; The updated model will guide the design of the next round of gene editing targets.

[0012] Preferably, during the training process of the multi-task regression model, five-fold cross-validation is used to evaluate the model performance, and models with R² within a specified range are retained for subsequent target prediction.

[0013] Preferably, the off-target scoring utilizes an off-target scoring algorithm that comprehensively considers the number of mismatches, position weights, and chromatin state, retaining only guide RNA sequences with off-target scores less than 0.3.

[0014] Preferably, after ICP registration and Poisson reconstruction, the plant height, internode length, center of gravity height, and aboveground biomass volume are extracted from the three-dimensional point cloud model.

[0015] Preferably, the sequencing peak diagram of the Sanger sequencing is analyzed using the TIDE algorithm to determine the insertion / deletion type, editing efficiency, and heterozygosity.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 0. This invention fundamentally changes the blind design mode of traditional gene editing that relies on literature reports or homologous gene speculation by constructing a multi-omics data-driven target prediction model, and realizes rational and precise screening of key functional sites of BR synthase gene. 2. By integrating a high-throughput non-destructive phenotypic monitoring platform with multispectral imaging and lidar in a controllable environment growth chamber, the field verification cycle of up to 6 months can be compressed to within 2 months, significantly reducing manpower, land and time costs. 3. By establishing a precise association analysis between genotype and phenotype and a closed-loop feedback optimization mechanism, we ensured that each editing experiment could provide effective learning signals for the model, continuously improving the success rate and predictability of subsequent editing.

[0017] 4. This invention can efficiently and stably obtain new soybean germplasm with moderate plant height, thick stems, compact branching and high harvest index, providing a solid genetic basis for high-yield, dense planting and mechanized harvesting of soybeans. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the multi-omics data-driven brassinolide synthase gene regulatory network prediction model in this invention. Figure 3This is a flowchart illustrating the logical process of high-throughput gene editing vector design and T0 generation plant transformation in this invention. Figure 4 This is a logical flowchart of the high-throughput non-destructive phenotypic monitoring and three-dimensional reconstruction in the controllable environment growth chamber of this invention. Figure 5 This is a flowchart illustrating the logical flow of genotype-phenotype association analysis and ideal plant type single-plant screening in this invention. Figure 6 This is a schematic diagram of the multi-level interaction relationship and data flow between the closed-loop feedback optimization mechanism and model iterative update in this invention. Detailed Implementation

[0019] Please refer to Figures 1 to 6 This invention provides a method for targeted improvement of ideal soybean plant architecture based on brassinosteroid synthase gene editing. Its core lies in constructing a systematic technical framework comprised of a multi-omics data-driven target prediction model, a high-throughput gene editing and rapid phenotypic screening platform, and a closed-loop feedback optimization mechanism. This method accurately identifies key functional sites in the brassinosteroid synthase encoding gene that regulate plant height, internode length, stem strength, and branching angle. Targeted editing is then performed using the CRISPR-Cas12a system. Under controlled environmental conditions, high-throughput, non-destructive phenotypic monitoring of T0 generation plants is achieved, ultimately leading to the selection of genetically stable homozygous lines that meet the preset ideal plant architecture threshold.

[0020] The method includes the following steps: S1, obtain a multi-source heterogeneous dataset of soybeans throughout their entire growth period; S2, Construct a predictive model of the brassinolide synthase gene regulatory network based on the multi-source heterogeneous dataset; S3. Based on the output of the prediction model, design and synthesize the CRISPR-Cas12a gene editing vector; S4. The gene editing vector was introduced into soybean recipient material using Agrobacterium-mediated cotyledon node transformation to obtain a T0 generation transgenic plant population. S5, High-throughput non-destructive phenotypic monitoring of the T0 generation plant population in a controlled environment growth chamber; S6. The high-throughput phenotypic data is correlated with the corresponding gene editing event information to screen out positive single plants that meet the preset ideal plant type threshold range, and these plants are used as parents for self-crossing to obtain genetically stable T2 generation homozygous lines. S7. The phenotypic and genotypic data of the selected candidate superior single plants are fed back into the prediction model, and the model parameters are incrementally updated using an online learning algorithm to form a closed-loop feedback optimization mechanism.

[0021] In step S1, obtaining the multi-source heterogeneous dataset of the entire soybean growth period specifically includes collecting whole-genome resequencing data of core soybean germplasm resources, transcriptome data of different tissues and organs at key developmental stages, ATAC-seq data of open chromatin regions, and three-dimensional phenotypic data of corresponding plants under standardized conditions.

[0022] Whole-genome resequencing data were obtained by deep sequencing of no less than 200 core soybean germplasm materials with broad genetic backgrounds, with a sequencing depth greater than 30-fold coverage. The original sequencing reads were aligned to the Williams82 reference genome after quality control. SNP and InDel variant detection were performed using the GATK workflow, and whole-genome association analysis was conducted to locate genomic regions associated with target traits such as plant height and internode length.

[0023] Transcriptome data were obtained from samples of key tissues such as shoot apical meristem, internode elongation zone, young leaves, and root tips at three developmental stages: seedling stage, branching stage, and flowering stage. Three biological replicates were set up for each sample. After total RNA extraction, strand-specific libraries were constructed, and sequencing was performed using the Illumina platform to obtain no less than 20 million valid reads. STAR software was used for alignment, HTSeq was used for quantification, and DESeq2 was used for differential expression analysis. BR synthase-related genes with a fold change greater than 2 between dwarf mutants and wild-type controls and a corrected P-value less than 0.01 were screened.

[0024] ATAC-seq data were obtained by isolating cell nuclei from the key tissues mentioned above, treating them with Tn5 transposase, constructing libraries, and sequencing to obtain open chromatin regions. After MACS2 peak identification, the coordinates of the open regions were co-localized with genome-wide association analysis (GWAS) regions and promoter regions of differentially expressed genes.

[0025] The three-dimensional phenotypic data were obtained by planting the above-mentioned germplasm materials in a standard greenhouse environment and using a fixed multispectral imaging system and a mobile lidar scanning device to periodically collect information on the plant canopy structure, stem morphology and biomass spatial distribution. The three-dimensional point cloud model of each plant was reconstructed, and quantitative indicators such as plant height, internode length, branch angle, center of gravity height and aboveground volume were extracted from it.

[0026] In step S2, the construction of the brassinolide synthase gene regulatory network prediction model specifically includes integrating the aforementioned multi-omics data and establishing a multi-task regression prediction framework.

[0027] First, all variant sites within the intervals where SNP sites obtained from genome-wide association analysis are located are used as a set of candidate regulatory sites; Secondly, the BR synthase genes downregulated or upregulated in the differential expression analysis and their promoter sequences within 2000 bases upstream were included in the candidate target gene set. Next, the chromatin regions specifically open in the shoot apical meristem and internode elongation zone identified by ATAC-seq were intersected with the aforementioned candidate regulatory sites and candidate target gene promoter regions. The regions shared by the three were retained as a set of high-confidence regulatory elements.

[0028] For each locus in this set, its sequence context features are extracted, including k-mer frequency, GC content, and conservation score; its chromatin accessibility score, i.e., the standardized value of ATAC-seq signal intensity, is extracted; and its linkage disequilibrium coefficient, i.e., the r² value with the nearest GWASSNP, is extracted. Using these feature vectors as input, and the effect values ​​of this locus on four core plant type traits (plant height, internode length, stem strength index, and branch angle) as output labels, a gradient boosting decision tree model is constructed. The effect values ​​are estimated using a linear mixture model, which includes genotype, environment, and genotype-environment interaction terms, with residuals following a normal distribution.

[0029] During training, five-fold cross-validation was used to evaluate model performance, and models with R² greater than 0.75 were retained for subsequent target prediction within a specified range. The model output is the predicted effect value for each candidate site on four traits, sorted by absolute effect value, and the top five are selected as priority editing targets.

[0030] In step S3, the design and synthesis of the CRISPR-Cas12a gene editing vector specifically includes identifying five priority editing target sites based on the effect value ranking results output by the prediction model. For each target site, potential cleavage sites satisfying the PAM sequence preference TTTV of the Cas12a protein are searched within a 200-base range upstream and downstream of the target site.

[0031] Off-target risk of each potential guide RNA sequence in the entire soybean genome was calculated using an off-target scoring algorithm that comprehensively considers mismatch number, position weight, and chromatin state, retaining only sequences with off-target scores less than 0.3. From each target region, a pair of guide RNA sequences with the lowest off-target risk and the highest predicted cleavage efficiency were selected. These two pairs of guide RNA sequences were cloned into the pCAMBIA1300 binary vector backbone containing a soybean U6 promoter-driven gRNA expression cassette and a constitutive promoter-driven Cas12a nuclease expression cassette, respectively. This binary vector also contains a DsRed reporter gene expression unit driven by the embryo-specific promoter GmLEC1, enabling rapid identification of transgenic events via fluorescence microscopy in early seed germination.

[0032] After the vector was constructed, it was amplified by E. coli DH5α, plasmid extracted, verified by enzyme digestion and sequencing, and then transformed into Agrobacterium strain AGL1 for subsequent transformation experiments.

[0033] In step S4, the gene editing vector is introduced into soybean recipient material using Agrobacterium-mediated cotyledon node transformation to obtain a T0 generation transgenic plant population.

[0034] The recipient material was the high-regeneration soybean variety Jack. Seeds were surface-sterilized and germinated in the dark for 48 hours. Cotyledonary explants were harvested and co-cultured for 45 minutes in an Agrobacterium suspension with an OD600 of 0.6, followed by transfer to a co-culture medium and incubation in the dark at 22°C for 3 days. The explants were then transferred to a selective medium containing cephalosporin and glufosinate, and cultured for 6 weeks at a 16-hour photoperiod and 26°C, with the medium changed every two weeks. When resistant shoots reached 2 cm in length, they were cut and transferred to a rooting medium to obtain fully regenerated plants.

[0035] All regenerated plants were transplanted into nutrient pots and grown in a greenhouse until they reached the three-leaf stage. Leaf samples were then taken for initial PCR screening to confirm T-DNA insertion. Positive plants were further cultured until maturity, and T0 generation seeds were harvested.

[0036] In step S5, the high-throughput non-destructive phenotypic monitoring of the T0 generation plant population in the controlled environment growth chamber specifically includes planting the T0 generation plants in the hydroponic system at a specified density of 16 plants per square meter, with the nutrient solution formula referring to the Hoagland standard and supplied in a daily cycle.

[0037] Environmental parameters were strictly controlled as follows: 14-hour photoperiod, 400 micromoles per square meter per second light intensity, 28 degrees Celsius daytime temperature, 22 degrees Celsius nighttime temperature, and 70% relative humidity. From the date of seedling emergence, a fully automated phenotypic collection process was initiated every 48 hours at specified intervals.

[0038] The process begins by using a 5-megapixel RGB camera mounted on top to acquire a top-view image of the canopy. The image is then segmented using the OpenCV library to calculate the plant's projected area and the angle between the main stem and the first-order branches. Subsequently, a side-mounted near-infrared camera acquires stalk texture images, and the gray-level co-occurrence matrix algorithm is used to calculate four texture features: contrast, correlation, energy, and homogeneity. These features are then weighted and synthesized into a stalk mechanical strength index. Finally, a rotating lidar scanned the plant 360 degrees at a frequency of 10 Hz per second, generating three-dimensional point cloud data containing no less than 500,000 points. After ICP registration and Poisson reconstruction, parameters such as plant height, internode length, center of gravity height, and aboveground biomass volume were extracted.

[0039] All phenotypic data are automatically stored in a central database by plant number and synchronized with the growth log.

[0040] In step S6, the association analysis between high-throughput phenotypic data and corresponding gene editing event information specifically includes collecting leaf samples from each T0 generation plant at the V4 growth stage, extracting genomic DNA, designing primers with 300 bases upstream and downstream of the target site for PCR amplification, and purifying the product for Sanger sequencing.

[0041] The sequencing peak diagram was analyzed using the TIDE algorithm to determine the insertion / deletion type, editing efficiency, and heterozygosity. The sequencing results were matched one-to-one with the final phenotypic data collected from the same plant in the growth chamber. Preset thresholds for ideal plant type were set as follows: plant height less than 70 cm, average internode length less than 5 cm, stem strength index greater than 0.8, and branching angle less than 45 degrees.

[0042] Plants with all phenotypic parameters falling within the threshold range and genotypes matching the expected editing type (i.e., non-synonymous mutations at the target site or disruption of cis-elements in the promoter region) were selected as candidate superior plants. These plants were self-pollinated to harvest T1 generation seeds. The T1 generation plants were planted under the same conditions and their phenotypes were identified. Lines with consistent phenotypes and no segregation were selected for further self-pollination to obtain homozygous T2 generation lines.

[0043] In step S7, the closed-loop feedback optimization step specifically includes using the complete genotype data (including target site mutation type and linkage SNP background) and high-precision phenotypic data (including dynamic growth curve and final trait value) of the selected candidate superior single plants as new training samples, and feeding them back into the brassinolide synthase gene regulatory network prediction model.

[0044] An online gradient boosting algorithm is used for incremental learning of the model. This algorithm retains historical knowledge while assigning higher initial weights to new samples, ensuring that the model can quickly adapt to new editing effect patterns. After the model is updated, the predicted effect values ​​of all candidate sites are recalculated, and a bias analysis is performed between the predicted and actual observed effect values ​​to identify systematic biases in the model within specific sequence contexts or chromatin environments.

[0045] Based on the analysis results, feature engineering strategies are adjusted, such as introducing local secondary structure prediction features or enhancing the spatiotemporal dynamic weights of chromatin states. The updated model is used to guide the design of the next round of gene editing targets, thus forming a complete technical loop from data-driven prediction, accurate editing, rapid validation to iterative model optimization.

[0046] The entire methodology, through the coordinated execution of the above seven steps, achieves rational, efficient, and predictable targeted improvement of the soybean BR synthase gene.

[0047] The deep integration of multi-omics data ensures the scientific nature of target selection and avoids the blindness of relying on literature speculation or random screening in traditional methods. High-throughput non-destructive phenotypic monitoring under controlled conditions reduces the field validation cycle from 6 months to less than 2 months, thus reducing labor, land and time costs. Precise association analysis between genotype and phenotype ensures the reliability of the screening results; The closed-loop feedback mechanism enables the system to continuously learn and optimize, with each experiment providing effective guidance for the next round of design. The resulting T2 generation homozygous lines exhibited ideal agronomic traits such as moderate plant height, robust stems, compact branching, strong lodging resistance, and a high harvest index, providing excellent genetic material for high-yield, high-density soybean cultivation and mechanized harvesting.

[0048] The technical details of this method cover the entire operational chain from data acquisition, model building, vector design, genetic transformation, phenotypic monitoring to data analysis and model updates. All steps adopt standardized protocols to ensure the reproducibility of experiments and the comparability of results.

[0049] For example, in multi-omics data acquisition, all samples are processed in the same growth batch and under the same environmental conditions; In model training, a unified feature standardization method and hyperparameter tuning strategy are adopted; In gene editing, validated high-fidelity Cas12a variants are used to reduce off-target risks; in phenotypic monitoring, all imaging devices are calibrated regularly to ensure data consistency.

[0050] This standardized and automated design throughout the entire process makes this method not only suitable for scientific research but also feasible for translation into breeding practice.

[0051] In summary, this embodiment describes in detail a method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing. It solves the key technical bottlenecks in existing soybean gene editing breeding, such as long cycle, high cost, and blind design, by organically integrating three core technical modules: multi-omics driven target prediction, high-throughput editing and phenotypic screening, and closed-loop feedback optimization. It provides a replicable and scalable technical paradigm for crop molecular design breeding.

[0052] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0053] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing, characterized in that, include: A multi-source heterogeneous dataset of soybean throughout its entire growth period was obtained. The multi-source heterogeneous dataset includes whole-genome resequencing data of soybean core germplasm resources, transcriptome data of different tissues and organs at key developmental stages, ATAC-seq data of open chromatin regions, and three-dimensional phenotypic data of corresponding plants under standardized conditions. Based on the aforementioned multi-source heterogeneous dataset, a prediction model of the brassinolide synthase gene regulatory network was constructed to identify key functional sites in the BR synthase encoding genes and their upstream cis-acting elements that have regulatory effects on plant height, internode length, stem strength and branching angle. Based on the output of the prediction model, a CRISPR-Cas12a gene editing vector targeting the key functional sites was designed and synthesized. The vector contains a pair of specific guide RNA sequences whose target sequences are located within conserved cis elements in the promoter region of the BR synthase gene or on key exons of the coding region. The gene editing vector was introduced into soybean recipient material using Agrobacterium-mediated cotyledon node transformation to obtain a T0 generation transgenic plant population; High-throughput non-destructive phenotypic monitoring was conducted on the T0 generation plant population in a controlled environment growth chamber. The monitoring was carried out by combining a multispectral imaging system with lidar scanning to obtain the plant height dynamic growth curve, stem diameter change rate, center of gravity shift angle and biomass spatial distribution map in real time. The high-throughput phenotypic data were correlated with the corresponding gene editing event information to screen out positive single plants whose plant phenotypic parameters met the preset ideal threshold range. These plants were then used as parents for self-pollination to obtain genetically stable T2 generation homozygous lines.

2. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 1, characterized in that, The constructed brassinolide synthase gene regulatory network prediction model includes: SNP and InDel variant detection were performed on whole-genome resequencing data, and genome-wide association analysis was conducted to locate genomic regions associated with plant height and internode length phenotypes. Differential expression analysis was performed on transcriptome data to screen for BR synthase-related genes whose expression levels differed by more than 2 fold between dwarf mutants and wild-type controls and whose corrected P-value was less than 0.

01. By integrating ATAC-seq data, we identified chromatin regions that were specifically opened in the shoot apical meristem and internode elongation zone, and then performed co-localization analysis on these regions with regions located by genome-wide association analysis and promoter regions of differentially expressed genes. Based on the colocalization results, a multi-task regression model was trained using the gradient boosting decision tree algorithm. The input features of the model are the sequence context, chromatin accessibility score and linkage disequilibrium coefficient of the candidate site, and the output is the predicted effect value of the site on the four core plant type traits.

3. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 2, characterized in that, The design and synthesis of the CRISPR-Cas12a gene editing vector targeting the key functional sites includes: Based on the effect values ​​output by the prediction model, the top five sites with the highest absolute effect values ​​were selected as editing targets. For each target, the PAM sequence preference TTTV of the Cas12a protein was used to search for guide RNA sequences with an off-target score of less than 0.3 within 200 bases upstream and downstream of the target. The two pairs of guide RNA sequences selected were cloned into binary vectors containing the soybean U6 promoter and the Cas12a expression cassette, respectively. The binary vectors also contain a red fluorescent protein reporter gene driven by an embryo-specific promoter.

4. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 3, characterized in that, The high-throughput non-destructive phenotypic monitoring of the T0 generation plant population within a controlled environment growth chamber includes: The T0 generation plant population was planted at a specified density in a standardized nutrient solution cultivation system. From the date of seedling emergence, a fully automatic phenotypic acquisition process is initiated at regular intervals. The fully automatic phenotypic acquisition process first obtains a top view of the canopy using a high-resolution RGB camera installed on the top, and calculates the plant's projected area and branch angle. Subsequently, stalk texture images were acquired using a side-mounted near-infrared camera, and the stalk mechanical strength index was calculated using the image gray-level co-occurrence matrix algorithm. Finally, the plant was scanned 360 degrees using a rotating lidar to reconstruct a three-dimensional point cloud model, from which plant height, internode length, center of gravity height, and aboveground biomass volume were extracted.

5. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 4, characterized in that, The step of correlation analysis between the high-throughput phenotypic data and the corresponding gene editing event information includes: For each T0 generation plant, the target site region was amplified by leaf PCR and Sanger sequencing was performed to determine its specific genotype, including insertion / deletion type and heterozygosity. The sequencing results were matched one-to-one with the final phenotypic data of the same plant collected in the growth chamber. Set a preset threshold range for the ideal plant shape; Plants whose phenotypic parameters all fall within the threshold range and whose genotype is the expected editing type are selected as candidate superior plants.

6. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 5, characterized in that, The method also includes a closed-loop feedback optimization step: The phenotypic and genotypic data of the selected candidate superior plants are fed back into the brassinolide synthase gene regulatory network prediction model as new training samples. Online learning algorithms are used to incrementally update the model parameters in order to improve the model's accuracy in predicting future editing events; The updated model will guide the design of the next round of gene editing targets.

7. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 6, characterized in that, The training process of the multi-task regression model uses five-fold cross-validation to evaluate the model performance, and models with R² within the specified range are retained for subsequent target prediction.

8. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 7, characterized in that, The off-target scoring uses an off-target scoring algorithm that comprehensively considers the number of mismatches, position weights, and chromatin state, retaining only guide RNA sequences with an off-target score of less than 0.

3.

9. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 8, characterized in that, After ICP registration and Poisson reconstruction, the plant height, internode length, center of gravity height, and aboveground biomass volume of the three-dimensional point cloud model were extracted.

10. The method for targeted improvement of ideal soybean plant architecture based on BR synthase gene editing according to claim 9, characterized in that, The sequencing peak diagrams from the Sanger sequencing were analyzed using the TIDE algorithm to determine the insertion / deletion type, editing efficiency, and heterozygosity.