Soybean lodging grade prediction method and system based on machine learning

By establishing a soybean lodging grade prediction model using machine learning algorithms, the problems of low efficiency and poor accuracy of traditional methods are solved, and a fast and accurate lodging grade prediction model is achieved, supporting breeding and cultivation management.

CN121834541APending Publication Date: 2026-04-10XINJIANG ACADEMY OF AGRI & RECLAMATION SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional methods for studying soybean lodging are inefficient and inaccurate, making it difficult to predict early lodging and guide breeding. They are also destructive to plants and cannot meet the breeding needs of high-density populations.

Method used

Machine learning algorithms were used to establish a lossless prediction model by collecting multi-dimensional feature data of soybean plants. Random forest, XGBoost and other algorithms were used to predict the lodging level. The model was optimized by combining five-fold cross-validation and weighted F1 score.

Benefits of technology

It enables rapid and accurate prediction of lodging severity, improves detection efficiency and accuracy, adapts to various environmental conditions, and supports lodging-resistant breeding and precision cultivation management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834541A_ABST
    Figure CN121834541A_ABST
Patent Text Reader

Abstract

The invention discloses a soybean lodging grade prediction method and system based on machine learning. The method comprises the following steps: collecting multi-dimensional features of soybean plants in a pod bearing period as feature data; carrying out standardization processing on the collected feature data; inputting the preprocessed data into a pre-trained machine learning model to obtain a lodging level prediction result; and outputting the predicted lodging level and confidence. The method has the remarkable effects that not only is the breeding process of lodging-resistant soybean varieties accelerated, but also decision support can be provided for precise cultivation management, and particularly, the method has important application value and popularization prospect in an under-mulch drip irrigation super-high-yield cultivation mode in Xinjiang.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural information technology, specifically to a method and system for predicting soybean lodging levels based on machine learning. Background Technology

[0002] Soybeans, as a globally important dual-purpose crop for both grain and oil, occupy a pivotal position in agricultural production and food security. They not only provide abundant plant protein but are also a crucial raw material for the oil and feed industries. However, lodging has consistently been a key factor restricting yield and quality improvement in high-yield soybean cultivation. Studies show that lodging during the pod-filling stage can lead to yield losses exceeding 50%, while also inducing the spread of pests and diseases, increasing seed mold rates, and significantly reducing protein content. Lodging plants, due to weakened stem support, cannot complete photosynthesis and nutrient transport normally, resulting in insufficient grain filling, reduced 100-grain weight, and severely impacting the final harvest. The lodging problem is particularly prominent in Xinjiang's ultra-high-yield cultivation model using drip irrigation under mulch. While the region has enormous production potential, the high-input, high-density cultivation environment leads to vigorous plant growth and dense canopies, increasing the risk of lodging. Therefore, while pursuing ultra-high yields, effectively predicting and controlling lodging has become a critical issue urgently needing to be addressed in soybean breeding and cultivation research in this region.

[0003] Traditional research on soybean lodging resistance relies primarily on field phenotypic selection and empirical judgment. Breeders subjectively score lodging resistance by observing plant growth in the field. This method is not only inefficient but also highly susceptible to subjective human factors, resulting in limited repeatability and accuracy. In terms of research methods, traditional techniques include marker-assisted selection, statistical analysis of agronomic traits, and morphogenesis simulation. For example, existing methods for predicting soybean lodging typically require testing indicators such as root length, root weight, and aboveground weight. These testing processes often involve destructive sampling of plants, making sample reuse impossible and hindering continuous monitoring. Furthermore, while marker-assisted selection can analyze traits from a genetic perspective with high accuracy, it is expensive and susceptible to gene-environment interactions. Statistical analysis of agronomic traits, while revealing relationships between traits and guiding breeding directions, struggles to establish practical predictive models. Morphogenesis simulation, while capable of quantifying growth and development processes and possessing strong mechanistic insights, lacks precision in lodging prediction. Furthermore, traditional statistical analyses typically employ correlation and path analyses to explore the relationship between traits and lodging, providing direction for breeding selection but failing to establish accurate lodging prediction models. Traditional morphological marker selection methods are time-consuming and inefficient, making them unsuitable for the needs of lodging-resistant breeding in modern high-density populations. Soybean lodging research also faces the bottleneck of insufficient phenotypic identification accuracy. Because lodging is influenced by multiple interacting factors and exhibits different characteristics at different growth and development stages of soybeans, existing technologies mostly only allow for observation and recording after lodging occurs, lacking pre-emptive prediction capabilities and failing to provide effective guidance for breeders' early selection. Overall, traditional research methods have significant shortcomings in predictive accuracy, operational efficiency, and application breadth, urgently requiring the development of novel, efficient, and accurate soybean lodging grade prediction models.

[0004] With the development of information technology, machine learning, as a highly efficient data mining and pattern recognition tool, has demonstrated strong application potential in various agricultural fields. Machine learning algorithms can extract inherent patterns from complex nonlinear data and build predictive models, providing new ideas for solving complex problems in agriculture. By using directly measurable phenotypic features such as stem tension, plant height, petiole length, and number of nodes, machine learning models can achieve accurate lodging grade prediction without damaging the plant. This not only preserves the integrity of the plant, making long-term tracking studies possible, but also significantly improves detection efficiency. In particular, ensemble learning algorithms such as XGBoost and random forests can significantly improve the accuracy and stability of the model by constructing multiple decision trees and integrating their prediction results. Furthermore, machine learning algorithms can effectively handle complex interactions between features, which is particularly important in predicting soybean lodging grades under the influence of multiple factors.

[0005] In response to the technical bottlenecks in current research on soybean lodging and the actual needs of ultra-high yield cultivation under drip irrigation under plastic film in Xinjiang, this invention aims to develop a method and system for predicting soybean lodging levels based on machine learning. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the purpose of this invention is to provide a machine learning-based method and system for predicting soybean lodging levels. This method and system can accurately, quickly, and non-destructively predict soybean lodging levels and establish a technical solution for rapidly, accurately, and non-destructively predicting soybean lodging characteristics. This provides a sustainable and efficient support tool for screening lodging-resistant breeding materials and high-yield cultivation management.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: Firstly, this invention proposes a machine learning-based method for predicting soybean lodging levels, the key of which includes the following steps: Step 1: Data Collection: Collect multi-dimensional characteristics of soybean plants during the pod-setting stage as feature data; Step 2, Data Preprocessing: Standardize the collected feature data; Step 3, Model Prediction: Input the preprocessed data into the pre-trained machine learning model to obtain the lodging level prediction results; The training process of the pre-trained machine learning model is as follows: Step 3.1: Simultaneously establish multiple types of machine learning models and determine the loss function and optimizer for each model; Step 3.2: Set the hyperparameters for various machine learning models; Step 3.3: Train various machine learning models using the input sample dataset, and use five-fold cross-validation to select the optimal model structure and hyperparameters; Step 3.4: Evaluate the merits of various machine learning models based on evaluation metrics, and select the optimal machine learning model as the pre-trained machine learning model. Step 4, Output Results: Output the predicted lodging level and confidence level.

[0008] Furthermore, the multi-dimensional features collected in step 1 consist of basic features and derived features.

[0009] Furthermore, the basic characteristics include: stem tensile strength, plant height, petiole length, and number of nodes; The derived characteristics include the plant height-to-tension ratio and the petiole segment ratio.

[0010] Furthermore, the plant height-to-strength ratio is the ratio of plant height to stem strength; The petiole segment ratio is the ratio of petiole length to the number of segments.

[0011] Furthermore, the various machine learning models established in step 3.1 include: Random Forest, XGBoost, LightGBM, Logistic Regression, Support Vector Machine, and K Nearest Neighbors; The loss functions and optimizers for each model are as follows: Logistic regression, support vector machine, and K-nearest neighbor models have no explicit loss function and are optimized through classification accuracy; random forest, XGBoost, and LightGBM models use classification cross-entropy as the loss function. The XGBoost and LightGBM models use the Adam optimizer, while the Random Forest model has no optimizer.

[0012] Furthermore, the hyperparameters of the various machine learning models in step 3.2 are as follows: KNN model: K value is 5-15; Random forest model: number of trees is 100-500, maximum depth is 5-15; XGBoost and LightGBM models: number of trees 100-300, learning rate 0.01-0.1, maximum depth 3-10; Support Vector Machine model: kernel function parameters are 0.001-0.1.

[0013] Furthermore, the evaluation index mentioned in step 3.4 adopts a weighted F1 score, the expression of which is: Where k is the lodging level number, For the first i Weights of class samples, For accuracy, This refers to the recall rate.

[0014] Secondly, this invention proposes a machine learning-based soybean lodging grade prediction system for implementing the method described in the first aspect, the key feature of which includes: The data acquisition module is used to collect multi-dimensional characteristics of soybean plants during the pod-setting stage as feature data. The data preprocessing module is used to standardize the collected feature data; The model prediction module is used to input the pre-processed data into a pre-trained machine learning model to obtain the lodging level prediction result. The training process of the pre-trained machine learning model is as follows: Simultaneously build multiple types of machine learning models and determine the loss function and optimizer for each model; Set the hyperparameters for various machine learning models; The input sample dataset is used to train various machine learning models, and five-fold cross-validation is used to select the optimal model structure and hyperparameters. The merits of various machine learning models are evaluated based on evaluation metrics, and the optimal machine learning model is selected as the pre-trained machine learning model. The results output module is used to output the predicted lodging level and confidence level.

[0015] Thirdly, the present invention provides a computer device comprising: at least one processor, and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, wherein when executed by the at least one processor, the instructions cause the at least one processor to perform the method as described in the first aspect.

[0016] Fourthly, the present invention provides a computer-readable storage medium on which a computer program is stored, which, when executed by a processor, implements the method described in the first aspect.

[0017] The significant effects of this invention are: This invention cleverly combines machine learning technology with crop breeding requirements to construct a method and system specifically for predicting soybean lodging severity. This technology not only helps accelerate the breeding process of lodging-resistant soybean varieties but also provides decision support for precision cultivation management, particularly demonstrating significant application value and promising prospects under the ultra-high-yield cultivation model of drip irrigation under mulch film in Xinjiang. With the acceleration of agricultural digital transformation, such cross-integrated technologies will provide more innovative solutions to the challenges facing soybean production, driving the soybean industry towards intelligent and precise development.

[0018] Compared with existing technologies, the outstanding advantages of this invention are reflected in several aspects. In terms of phenotypic data collection, traditional methods rely on manual evaluation and are highly subjective, while this invention uses standardized feature collection to achieve objective quantification. Regarding prediction models, traditional single algorithms have poor adaptability, while this invention selects the best algorithm through multi-model comparison. In feature selection, traditional methods require destructive sampling, while this invention innovatively integrates large-sample data with machine learning to predict plant lodging levels non-destructively. In terms of application specificity, traditional general-purpose methods have poor regional adaptability, while this invention is specifically optimized for the Xinjiang drip irrigation model.

[0019] Furthermore, the present invention also has the following advantages: High-precision prediction: The optimal model achieved a weighted F1 score of 0.9597 and an accuracy of 0.959, which is significantly better than the accuracy of the traditional linear model. Fast response: The single-sample prediction process is completed in milliseconds, meeting the requirements for real-time prediction; Strong environmental adaptability: The model training takes into account a variety of environmental conditions and has good generalization ability. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method described in this invention; Figure 2 This is a schematic diagram of the system described in this invention; Figure 3 This is a schematic diagram of the structure of the device described in this invention. Detailed Implementation

[0021] The specific embodiments and working principles of the present invention will be further described in detail below with reference to the accompanying drawings.

[0022] Example 1: like Figure 1 As shown in the figure, this embodiment provides a machine learning-based method for predicting soybean lodging levels. The specific steps are as follows: Step 1: Data Collection: Collect multi-dimensional characteristics of soybean plants during the pod-setting stage as feature data; In this embodiment, the multi-dimensional features collected during the specific implementation process consist of basic features and derived features.

[0023] It should be noted that, in addition to the basic and derived features mentioned above, the sample data also includes label features, which are used to characterize the lodging level of soybeans. These are categorized by lodging severity, such as levels 0-4, with level 0 being no lodging and level 4 being complete lodging.

[0024] Specifically, the basic features include: Mechanical characteristics: plant tensile strength (reflects the stem's resistance to bending and is directly related to lodging resistance); Morphological characteristics: plant height (cm, reflecting the overall height of the plant), petiole length (cm, affecting the distribution of the plant's center of gravity), number of nodes (number, reflecting the stem support structure); The derived features include: Plant height to pull force ratio = Plant height / Pull force; Petiole segment ratio = petiole length / number of segments.

[0025] As can be seen from the calculation formula of the derived characteristics, the plant height-tensile strength ratio is used to quantify the balance relationship between "height-tensile strength". The smaller the ratio, the stronger the lodging resistance potential. The petiole segment ratio reflects the length of a single petiole segment. If the ratio is too large, it will easily lead to the shift of the plant's center of gravity.

[0026] This embodiment collected a large amount of data on the stem tension, plant height, petiole length, number of nodes, and corresponding lodging levels of soybean plants during the pod-setting stage. In addition, to improve the differentiation of lodging levels, two derived features, "plant height-tension ratio" and "petiole-node ratio", were added. By combining multiple data, not only can the accuracy of prediction be improved, but the generalization ability of this method can also be ensured.

[0027] Step 2, Data Preprocessing: Standardize the collected feature data; In the specific implementation process, only standardization processing is performed on the 7 collected features (5 basic features + 2 derived features): The Z-score standardization method is used to transform each feature into a distribution with a mean of 0 and a standard deviation of 1. The formula is as follows:

[0028] in, μ The characteristic mean, σ The characteristic standard deviation is used to eliminate the interference of different dimensions such as plant height and stem tension on model training.

[0029] Step 3, Model Prediction: Input the preprocessed data into the pre-trained machine learning model to obtain the lodging level prediction results; The training process of the pre-trained machine learning model is as follows: Step 3.1: Stratify the collected feature data into a training set for cross-validation and a test set for performance evaluation according to an 8:2 ratio. Simultaneously build multiple machine learning models and determine the loss function and optimizer for each model. In this example, the various machine learning models established include: Random Forest (RF), XGBoost, LightGBM, Logistic Regression (LR), Support Vector Machine (SVM), and K Nearest Neighbors (KNN). The loss functions and optimizers for each model are as follows: Logistic regression, support vector machine, and K-nearest neighbor models have no explicit loss function and are optimized through classification accuracy; random forest, XGBoost, and LightGBM models use classification cross-entropy as the loss function. The XGBoost and LightGBM models use the Adam optimizer with a learning rate of 0.01-0.1, while the Random Forest model has no optimizer and adjusts the number of trees.

[0030] Step 3.2: Set the hyperparameters for various machine learning models; In practice, the hyperparameters of various machine learning models are as follows: KNN model: K value is 5-15; Random forest model: number of trees is 100-500, maximum depth is 5-15; XGBoost and LightGBM models: number of trees 100-300, learning rate 0.01-0.1, maximum depth 3-10; Support Vector Machine model: Kernel function parameters (gamma value of RBF kernel) are 0.001-0.1.

[0031] Step 3.3: Standardize the training and test sets, input the training and test sets to train various machine learning models, and use five-fold cross-validation to select the optimal model structure and hyperparameters. In this example, for each type of model, five subsets are randomly divided in the training set. Four subsets are used for training and one subset is used for validation each time, and this process is repeated five times. Record the validation performance (e.g., F1 score) of the model under different hyperparameter combinations, and select the hyperparameter combination with the best average performance in 5-fold validation as the final structure of the model.

[0032] Step 3.4: Evaluate the merits of various machine learning models based on evaluation metrics, and select the optimal machine learning model as the pre-trained machine learning model. In some specific implementations, because there may be class imbalance in the lodging level (e.g., a large number of non-lodging samples), the evaluation index uses the weighted F1 score (assigning higher weights to minority class samples) as the core criterion for model performance, and its expression is: Where k is the lodging level number, For the first i Weights of class samples, For accuracy, This refers to the recall rate.

[0033] The weighted F1 score of each of the six models is calculated on the test set. The model with the highest weighted F1 score is selected as the final prediction model, which is the pre-trained machine learning model (in practical applications, LightGBM is often preferred because it balances accuracy and speed).

[0034] Step 4, Result Output: Collect soybean feature data to be predicted, standardize it, and input it into the selected pre-trained machine learning model to output the predicted lodging level and confidence level.

[0035] Example 2: See appendix Figure 2 This invention provides a machine learning-based soybean lodging grade prediction system for implementing the method described in Embodiment 1, comprising: The data acquisition module is used to collect multi-dimensional characteristics of soybean plants during the pod-setting stage as feature data. The data preprocessing module is used to standardize the collected feature data; The model prediction module is used to input the pre-processed data into a pre-trained machine learning model to obtain the lodging level prediction result. The training process of the pre-trained machine learning model is as follows: Simultaneously build multiple types of machine learning models and determine the loss function and optimizer for each model; Set the hyperparameters for various machine learning models; The input sample dataset is used to train various machine learning models, and five-fold cross-validation is used to select the optimal model structure and hyperparameters. The merits of various machine learning models are evaluated based on evaluation metrics, and the optimal machine learning model is selected as the pre-trained machine learning model. The results output module is used to output the predicted lodging level and confidence level.

[0036] Example 3: See appendix Figure 3 This embodiment proposes a computer device, one embodiment of which includes: One or more central processing units, one or more power supplies, one or more operating systems, one or more computer programs, one or more databases, memory, one or more network interfaces and one or more input / output interfaces.

[0037] The central processing unit is capable of executing the steps described in the aforementioned embodiment 1, which will not be repeated here.

[0038] This power supply can meet the power requirements of computer equipment for normal operation or overclocking.

[0039] The operating system, for example When choosing an operating system, such as TM, MacOSXTM, UnixTM, LinuxTM, etc., pay attention to the compatibility between the version of the code being run and the operating system.

[0040] The memory stores one or more application programs or data, and can be volatile or persistent storage. The program stored in the memory comprises one or more modules, each capable of executing a series of instructions on the computer device. The central processing unit (CPU) can communicate with the memory and execute the series of instructions stored in the memory on the computer device.

[0041] Example 4: This embodiment proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the insulator image anomaly detection method as described in Embodiment 1.

[0042] It is evident that the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a scaled storage medium, which includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory, a random access memory, or an optical disc. It includes several instructions to cause a computer device, such as a personal computer, a server, or a network device, to execute all or part of the steps of the methods described in the various embodiments of this application.

[0043] Application Example 1: 1. Data Foundation Based on the ultra-high yield cultivation management model of drip irrigation under film in Xinjiang, 5361 mixed field samples of 382 soybean varieties (lines) were collected over two consecutive years from 2024 to 2025. The stem tension, plant height, petiole length, number of nodes and lodging grade (0-4) at harvest were recorded during the pod-setting stage.

[0044] 2. Modeling Process ① The training set and the test set are obtained by stratified sampling at a ratio of 8:2; ②Standardized processing; ③ Train six mainstream algorithms (LR, SVM, RF, LightGBM, XGBoost, KNN); ④ Five-fold cross-validation to optimize hyperparameters; ⑤ Evaluate using weighted F1 and accuracy metrics.

[0045] 3. Results XGBoost performed best: weighted F1=0.9597, accuracy=0.959, significantly better than the traditional linear model (Accuracy=0.7008).

[0046] 4. Key Hyperparameters learning_rate=0.05, max_depth=5, n_estimators=400, subsample=0.8.

[0047] Application Example 2: 1. System Deployment The optimal model of Example 1 is encapsulated as a REST API, which can be called by the front-end mini-program / webpage with one click.

[0048] 2. Application Examples ① User input: tension 14.8N, plant height 82.3cm, petiole length 11.5cm, number of nodes 9.

[0049] ②System response: Predicted lodging level = 1, confidence level 0.9932; ③ Output format JSON structure: {"Rank":1,"Result":[0.006,0.993,0.001,0.000],"Confidence":0.9932}.

[0050] Application Example 3: 1. Data 26 varieties / lines, 7 traits (newly added plant height-to-stretch ratio, petiole segment ratio).

[0051] 2. Verification Strategy We continued to use the XGBoost model with fixed hyperparameters from Example 1, making direct predictions with zero training, and examined its generalization ability.

[0052] 3. Results

[0053] Accuracy =96.1%, Macro-F1= 0.958; As can be seen from Table 3, the experimental results show that the method of predicting the lodging grade of soybeans using the model is consistent with the results of the field test method, indicating that the method of predicting the lodging grade of soybean varieties using the model in this invention has high accuracy.

[0054] Application Example 4: Based on the combined results of 5,335 big data samples and 26 small sample samples, the plant height-to-tension ratio (27.8%), tension (21.4%), and plant height (18.6%) are the core indicators for determining lodging risk. They can be directly used for the breeding of lodging-resistant materials for spring soybeans under the ultra-high-yield cultivation management model of drip irrigation under film in Xinjiang.

[0055] In summary, this invention collects characteristics of soybean plants such as tensile strength, plant height, petiole length, and number of nodes, and uses a standardized machine learning model to predict lodging levels. Six mainstream machine learning algorithms were compared, and the XGBoost model, with the highest accuracy and weighted F1 score, was selected as the primary prediction model. Experiments show that this invention has high prediction accuracy and fast response speed, and can be widely applied to the breeding of lodging-resistant materials for spring soybeans under the ultra-high-yield cultivation management model of drip irrigation under plastic film in Xinjiang. It not only helps accelerate the breeding process of lodging-resistant soybean varieties but also provides decision support for precision cultivation management, offering more innovative solutions to the challenges faced by soybean production and promoting the development of the soybean industry towards intelligence and precision.

[0056] The technical solution provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make several improvements and modifications to this invention without departing from the principles of this invention, and these improvements and modifications also fall within the protection scope of the claims of this invention.

Claims

1. A machine learning-based method for predicting soybean lodging grade, characterized in that, Includes the following steps: Step 1: Data Collection: Collect multi-dimensional characteristics of soybean plants during the pod-setting stage as feature data; Step 2, Data Preprocessing: Standardize the collected feature data; Step 3, Model Prediction: Input the preprocessed data into the pre-trained machine learning model to obtain the lodging level prediction results; The training process of the pre-trained machine learning model is as follows: Step 3.1: Simultaneously establish multiple types of machine learning models and determine the loss function and optimizer for each model; Step 3.2: Set the hyperparameters for various machine learning models; Step 3.3: Train various machine learning models using the input sample dataset, and use five-fold cross-validation to select the optimal model structure and hyperparameters; Step 3.4: Evaluate the merits of various machine learning models based on evaluation metrics, and select the optimal machine learning model as the pre-trained machine learning model. Step 4, Output Results: Output the predicted lodging level and confidence level.

2. The soybean lodging grade prediction method based on machine learning according to claim 1, characterized in that, The multi-dimensional features collected in step 1 consist of basic features and derived features.

3. The soybean lodging grade prediction method based on machine learning according to claim 2, characterized in that, The basic characteristics include: stem tensile strength, plant height, petiole length, and number of nodes; The derived characteristics include the plant height-to-tension ratio and the petiole segment ratio.

4. The soybean lodging grade prediction method based on machine learning according to claim 3, characterized in that, The plant height-to-tightness ratio is the ratio of plant height to stem tension. The petiole segment ratio is the ratio of petiole length to the number of segments.

5. The soybean lodging grade prediction method based on machine learning according to claim 1, characterized in that, The various machine learning models established in step 3.1 include: Random Forest, XGBoost, LightGBM, Logistic Regression, Support Vector Machine, and K Nearest Neighbors; The loss functions and optimizers for each model are as follows: Logistic regression, support vector machine, and K-nearest neighbor models have no explicit loss function and are optimized through classification accuracy; random forest, XGBoost, and LightGBM models use classification cross-entropy as the loss function. The XGBoost and LightGBM models use the Adam optimizer, while the Random Forest model has no optimizer.

6. The soybean lodging grade prediction method based on machine learning according to claim 1, characterized in that, The hyperparameters of the various machine learning models in step 3.2 are as follows: KNN model: K value is 5-15; Random forest model: number of trees is 100-500, maximum depth is 5-15; XGBoost and LightGBM models: number of trees 100-300, learning rate 0.01-0.1, maximum depth 3-10; Support Vector Machine model: kernel function parameters are 0.001-0.

1.

7. The soybean lodging grade prediction method based on machine learning according to claim 1, characterized in that, The evaluation index mentioned in step 3.4 uses a weighted F1 score, the expression of which is: Where k is the lodging level number, For the first i Weights of class samples, For accuracy, This refers to the recall rate.

8. A machine learning-based soybean lodging grade prediction system for implementing the method as described in any one of claims 1-7, characterized in that, include: The data acquisition module is used to collect multi-dimensional characteristics of soybean plants during the pod-setting stage as feature data. The data preprocessing module is used to standardize the collected feature data; The model prediction module is used to input the pre-processed data into a pre-trained machine learning model to obtain the lodging level prediction result. The training process of the pre-trained machine learning model is as follows: Simultaneously build multiple types of machine learning models and determine the loss function and optimizer for each model; Set the hyperparameters for various machine learning models; The input sample dataset is used to train various machine learning models, and five-fold cross-validation is used to select the optimal model structure and hyperparameters. The merits of various machine learning models are evaluated based on evaluation metrics, and the optimal machine learning model is selected as the pre-trained machine learning model. The results output module is used to output the predicted lodging level and confidence level.

9. A computer device, characterized in that, include: At least one processor, and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, cause the at least one processor to perform the method as described in any one of claims 1-8.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.