Gradient boosting tree-based pesticide-loaded nanofiber slow release rate prediction algorithm and burst release risk judgment method

By constructing a gradient boosting tree model to predict the sustained release rate of pesticide-loaded nanofibers and determine the risk of burst release, the problem of long cycle and high cost in traditional methods is solved. This achieves high-precision sustained release rate prediction and burst release risk determination, supporting rapid screening and adaptation to the needs of agricultural scenarios.

CN121768518APending Publication Date: 2026-03-31NORTHEAST AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies rely on laboratory experiments for evaluating the sustained-release performance of pesticide-loaded nanofibers. These experiments are time-consuming, costly, and lack high-precision prediction models and methods for assessing burst release risks. Consequently, they are insufficient to meet the needs of rapid screening of large quantities of carrier materials and agricultural applications.

Method used

A gradient boosting tree-based algorithm for predicting the sustained release rate of pesticide-loaded nanofibers was constructed. By building a dataset containing six key features such as fiber diameter and contact angle, the algorithm uses a gradient boosting tree model to predict the sustained release rate at 12h, 24h, and 48h. Based on the 12h sustained release rate, the algorithm determines the risk of burst release, achieving high-precision sustained release rate prediction and burst release risk assessment.

Benefits of technology

It significantly shortens the sustained-release performance evaluation cycle, reduces R&D costs, enables rapid screening of large quantities of carrier materials, achieves a prediction accuracy of 0.9976, can accurately distinguish sustained-release rate differences within 1%, and effectively determine burst release risks, adapting to the needs of different agricultural scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768518A_ABST
    Figure CN121768518A_ABST
Patent Text Reader

Abstract

The invention provides a pesticide-loaded nanofiber slow release rate prediction algorithm based on a gradient boosting tree and a burst release risk judgment method, relates to the field of agricultural nano material performance prediction and machine learning algorithm application, and solves the problems that a traditional slow release performance evaluation period is long, the cost is high, large-scale screening cannot be performed, and the efficiency is high. The problems that an existing prediction model is weak in generalization ability, unstable in output result and lack of an agricultural scene key burst release risk judgment function, and the research and development efficiency and the industrialization process of a nano pesticide carrier are restricted are solved. According to the method, 1000 groups of sample data sets containing 6 features are constructed, after standardization preprocessing, the 12h / 24h / 48h slow release rate is predicted through a gradient boosting tree model configured by optimal parameters, the burst release risk is judged based on the 12h slow release rate, and rapid screening of nanofiber carriers is achieved. The evaluation period of the nanofiber carrier is greatly shortened, the cost is reduced, the prediction precision is high, the burst release risk can be accurately judged, the requirements of different agricultural scenes are met, and cross-platform calling is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural nanomaterial performance prediction and machine learning algorithm application technology, and in particular to a pesticide loading nanofiber slow release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree. Background Technology

[0002] In agricultural production, pesticides play an irreplaceable role in ensuring food security and controlling pests and weeds. However, the low utilization rate, environmental pollution from residues, and ecological damage caused by traditional pesticides have long plagued the sustainable development of agriculture. To address this challenge, nanotechnology has been widely applied to pesticide formulation improvement. Among these technologies, nanofibers, with their high specific surface area, controllable pore structure, and good carrier compatibility, have become the preferred material for constructing pesticide slow-release systems. Through interactions with pesticide molecules, such as hydrogen bonding, hydrophobic interactions, or ion exchange, nanofibers can effectively regulate the pesticide release rate, prolong the duration of effectiveness, and reduce environmental risks.

[0003] Currently, related research has focused on the preparation, sustained-release mechanism, and performance optimization of functionalized nanomaterials. Regarding carrier materials, researchers have prepared various nanocomposite systems loaded with different types of pesticides (herbicides, insecticides, growth regulators, etc.) using techniques such as in-situ synthesis, surface modification, and layer-by-layer self-assembly. They are attempting to optimize pesticide loading and sustained-release effects by controlling parameters such as carrier particle size, surface functional groups, and microstructure. Meanwhile, some studies have focused on the impact of environmental factors on sustained-release performance, finding that soil pH, ionic strength, humus content, and microbial communities can all alter the release kinetics of pesticides, providing a reference for the environmentally adaptable design of sustained-release systems.

[0004] In the evaluation and prediction of pesticide sustained-release performance, traditional methods mainly rely on laboratory static release experiments and field trials, analyzing release patterns by fitting kinetic models. However, these methods suffer from drawbacks such as long cycles, complex operations, and high costs, making it difficult to meet the needs of rapid screening of large quantities of carrier materials. With the development of machine learning technology, some studies have attempted to apply it to related performance predictions, including constructing models such as gradient boosting trees and random forests. By extracting material physicochemical characteristics and process parameters as input variables, quantitative predictions of target performance can be achieved. However, existing prediction models have significant limitations: on the one hand, some models are developed for non-pesticide sustained-release scenarios such as drug relocation and load prediction, and their feature systems are incompatible with nanofiber drug-loaded systems. When directly transferred and used, the prediction error is large, making it difficult to distinguish subtle differences in sustained-release rates. On the other hand, existing models mostly focus on predicting the sustained-release amount at a single time point, without considering the burst release problem in the early stages of pesticide release—this problem can lead to rapid pesticide loss, wasted efficacy, and may also exacerbate environmental risks. Currently, there is a lack of specific methods for determining burst release risks for nanofiber drug-loaded systems.

[0005] Furthermore, the development of existing nanofiber sustained-release systems still relies on empirical experimental optimization, requiring repeated adjustments to multiple influencing factors such as fiber diameter, preparation process parameters, and pesticide physicochemical properties. This is not only time-consuming and labor-intensive but also makes it difficult to systematically reveal the intrinsic relationship between each factor and sustained-release performance. Although some studies have attempted to assist analysis through molecular simulation and kinetic fitting, an integrated tool with both high-precision predictive capabilities and practical judgment functions has not yet been developed, severely restricting the development efficiency and industrialization process of nanofiber pesticide sustained-release formulations. Therefore, developing a method that is compatible with nanofiber drug-loaded systems, has high predictive accuracy, and can rapidly screen carrier materials and determine burst release risk has become an urgent need in the field of agricultural nanotechnology. Summary of the Invention

[0006] This invention proposes a gradient boosting tree-based algorithm for predicting the sustained-release rate of pesticide-loaded nanofibers and a method for determining burst release risk. By constructing a dataset of 1000 samples containing six key features such as fiber diameter and contact angle, and after standardized preprocessing, the algorithm uses a gradient boosting tree model with optimal parameter configuration to quantitatively predict the sustained-release rate at 12h, 24h, and 48h. The burst release risk is determined based on the 12h sustained-release rate. This addresses the problems of long evaluation cycles, high costs, and inability to conduct large-scale screening in traditional sustained-release performance assessments, as well as the weak generalization ability and unstable output results of existing prediction models, and the lack of crucial burst release risk determination functions for agricultural scenarios, which restrict the efficiency of nanopesticide carrier research and development and industrialization.

[0007] An algorithm for predicting the sustained release rate of pesticide-loaded nanofibers and a method for determining burst release risk based on gradient boosting trees are proposed, comprising the following steps: S1. Data preparation: Construct a dataset containing nanofiber and pesticide-related characteristics and slow-release rates; S2. Preprocess the dataset to eliminate differences in feature dimensions; S3. Construct a gradient boosting tree regression model and configure the core parameters of the model; S4. Train the gradient boosting tree regression model using the preprocessed dataset, and evaluate the model's fit using performance metrics. S5. Load the trained gradient boosting tree regression model and preprocessing tools, input the features of the nanofibers to be screened, predict the sustained release rate at multiple time points, and determine the burst release risk based on the 12-hour sustained release rate.

[0008] Furthermore, in S1, the dataset is a nanofiber_release_data.csv dataset containing 1000 samples. This dataset is constructed by extracting 200 sets of experimental data from a professional nanofiber database and supplementing them with 800 sets of similar data. Each sample contains 6 input features and 1 output target. The input features are fiber diameter, contact angle, pesticide LogP, pesticide molecular weight, electrospinning voltage, and adsorption energy. The output target is the 12h / 24h / 48h release rate, and the release rate ranges from 0% to 100%. The dataset is read using the pd.read_csv function.

[0009] Furthermore, in S2, the preprocessing is a standardization process implemented using the fit and transform methods. The fit method calculates the mean and standard deviation of each feature using the following formula:

[0010] The transform method standardizes the features using the following formula:

[0011] Avoid division by zero by setting self.scale_[self.scale_==0]=1.0.

[0012] Furthermore, in S3, the core parameters of the gradient boosting tree regression model are configured as follows: number of decision trees n_estimators=200, maximum depth of decision trees max_depth=5, learning rate learning_rate=0.05, minimum number of leaf samples min_samples_leaf=3, and random seed random_state=42.

[0013] Furthermore, in S3, the mathematical model implementation of the gradient boosting tree regression model includes: Initial prediction: The mean of the target variable is used as the initial prediction value, and the formula is:

[0014] Residual calculation: Considering sample weights, the residual formula is:

[0015] Decision tree construction: By traversing all features and thresholds, the optimal split point is selected with the goal of minimizing the mean squared error (MSE). The MSE calculation formula is as follows:

[0016] The total error after splitting is the weighted sum of the errors of the left and right subtrees; Prediction Update: After training each tree, through... Update the predicted value, and finally constrain the result to 0%~100% by np.clip(y_pred,0,100) when making the prediction.

[0017] Furthermore, in S4, the model is trained by calling `model.fit(X_processed, y)`, and the coefficient of determination R is used. 2 To evaluate the fit, R 2 The calculation formula is as follows:

[0018] Requirement R 2 ≥0.99; The model release_regressor.pkl and the preprocessing tool release_reg_preprocessor.pkl are saved as binary files using pickle.dump.

[0019] Furthermore, in S5, the specific process for determining the burst release risk is as follows: input the six features of the nanofiber to be screened, after standardization by the preprocessing tool, call the trained model to output the 12h / 24h / 48h sustained release rate, preset the 12h sustained release rate threshold, if the output 12h sustained release rate is higher than the threshold, it is determined that there is a burst release risk, which is suitable for the need for rapid control in agricultural scenarios; if the output 12h sustained release rate is lower than or equal to the threshold, it is determined that there is no burst release risk, which is suitable for the need for long-term sustained release in agricultural scenarios.

[0020] A storage medium storing a computer program, which, when executed by a processor, implements the aforementioned algorithm for predicting the sustained release rate of pesticide-loaded nanofibers based on gradient boosting trees and the method for determining the risk of burst release.

[0021] A computer device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described algorithm for predicting the sustained release rate of pesticide-loaded nanofibers based on gradient boosting trees and the method for determining the risk of burst release.

[0022] Compared with existing technologies, this invention achieves significant beneficial effects through the above-mentioned technical solutions: The gradient boosting tree-based algorithm for predicting the sustained-release rate of pesticide-loaded nanofibers and the method for determining burst-release risk effectively solves the drawback of traditional pesticide-loaded nanofiber sustained-release performance evaluation relying on laboratory experiments, significantly shortening the evaluation cycle, reducing R&D costs, and enabling rapid screening of large quantities of carrier materials. The gradient boosting tree model constructed by this invention has extremely high prediction accuracy, with a determination coefficient of 0.9976 and a prediction error of no more than 2%. It can accurately distinguish sustained-release rate differences within 1%, exhibiting stronger stability and adaptability compared to existing prediction models, perfectly meeting the performance analysis needs of nanofiber-loaded pesticide systems. Furthermore, the burst-release risk determination function added to this invention can distinguish the burst-release attributes of carriers based on the 12-hour sustained-release rate. This avoids rapid pesticide loss and environmental pollution caused by burst release and meets the needs of different agricultural scenarios. For pest and disease control requiring rapid effectiveness, carriers with burst-release risk can be selected; for scenarios requiring sustained effectiveness and reduced application frequency, carriers without burst-release risk can be selected. Meanwhile, this method supports the commonly used Python programming environment. The trained model and preprocessing tools can be saved as binary files for direct use, making it easy to deploy in industrial equipment. This can improve the efficiency of research and development of nano-pesticide sustained-release formulations and accelerate their industrial application. Attached Figure Description

[0023] Figure 1 Flowchart of the prediction model; Figure 2 This is a graph illustrating the linear correlation between experimental and predicted values ​​in a machine learning model. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Reference Figures 1-2 As shown, an algorithm for predicting the sustained release rate of pesticide-loaded nanofibers based on gradient boosting trees and a method for determining the risk of burst release include the following steps: S1. Data preparation: Construct a dataset containing nanofiber and pesticide-related characteristics and slow-release rates; S2. Preprocess the dataset to eliminate differences in feature dimensions; S3. Construct a gradient boosting tree regression model and configure the core parameters of the model; S4. Train the gradient boosting tree regression model using the preprocessed dataset, and evaluate the model's fit using performance metrics. S5. Load the trained gradient boosting tree regression model and preprocessing tools, input the features of the nanofibers to be screened, predict the sustained release rate at multiple time points, and determine the burst release risk based on the 12-hour sustained release rate.

[0026] Specifically, this invention constructs a dataset containing nanofiber and pesticide-related characteristics and sustained-release rates. After preprocessing to eliminate dimensional differences, it utilizes a gradient boosting tree regression model to predict sustained-release rates at multiple time points and determines burst release risk based on the 12-hour sustained-release rate. This effectively solves the drawback of traditional pesticide-loaded nanofiber sustained-release performance evaluation relying on laboratory experiments, significantly shortening the evaluation cycle and reducing R&D costs, enabling rapid screening of large batches of carrier materials. The gradient boosting tree model constructed in this invention exhibits excellent prediction accuracy, with a determination coefficient of 0.9976 and a prediction error of no more than 2%. It can accurately distinguish sustained-release rate differences within 1%, demonstrating stronger stability and adaptability compared to existing prediction models, perfectly meeting the performance analysis needs of nanofiber-loaded pesticide systems. Furthermore, the added burst release risk assessment function avoids rapid pesticide runoff and environmental pollution caused by burst release, while also meeting the needs of different agricultural scenarios. For pest and disease control requiring rapid effectiveness, carriers with burst release risk can be selected; for scenarios requiring sustained effectiveness and reduced application frequency, carriers without burst release risk can be selected. In addition, this method supports the Python programming environment, and the trained model and preprocessing tools can be saved as binary files for direct use, making it easy to deploy in industrial equipment.

[0027] Furthermore, in S1, the dataset is a nanofiber_release_data.csv dataset containing 1000 samples. This dataset is constructed by extracting 200 sets of experimental data from a professional nanofiber database and supplementing them with 800 sets of similar data. Each sample contains 6 input features and 1 output target. The input features are fiber diameter, contact angle, pesticide LogP, pesticide molecular weight, electrospinning voltage, and adsorption energy. The output target is the 12h / 24h / 48h release rate, and the release rate ranges from 0% to 100%. The dataset is read using the pd.read_csv function.

[0028] Specifically, this invention constructs a nanofiber_release_data.csv dataset containing 1000 samples, based on 200 sets of real experimental data from a professional nanofiber database, supplemented by 800 similar sets of data. This ensures the authenticity and reliability of the dataset while significantly expanding the sample coverage, effectively solving the problem of weak generalization ability in existing prediction models due to insufficient sample size. It can support the prediction of the sustained-release performance of multiple types of nanofiber carriers. Simultaneously, it precisely selects six key features directly related to pesticide sustained-release performance—fiber diameter, contact angle, pesticide LogP, pesticide molecular weight, electrospinning voltage, and adsorption energy—as inputs. Each feature specifically reflects the influence of the structural characteristics of the nanofiber carrier, the physicochemical properties of the pesticide, or the preparation process parameters on the sustained-release effect. This avoids the problem of decreased prediction accuracy caused by redundant or irrelevant feature selection in existing technologies, enabling the model to focus on core related factors and improve prediction accuracy. The system explicitly outputs the sustained-release rates at three key time points: 12h, 24h, and 48h, with a limited range of 0% to 100%. This aligns with the actual needs of pesticide sustained-release rhythm in agricultural scenarios and is more practical than existing single-time-point predictions. The use of the pd.read_csv function to read the dataset ensures the stability and convenience of data loading, laying a reliable foundation for subsequent data preprocessing and model training.

[0029] Furthermore, in S2, the preprocessing is a standardization process implemented using the fit and transform methods. The fit method calculates the mean and standard deviation of each feature using the following formula:

[0030] The transform method standardizes the features using the following formula:

[0031] Avoid division by zero by setting self.scale_[self.scale_==0]=1.0.

[0032] Specifically, this invention achieves data standardization through the `fit` and `transform` methods. It precisely calculates the mean and standard deviation of each feature using explicit formulas, and then transforms features with different dimensions to a unified scale using standardization formulas. This effectively eliminates dimensional differences between different types of features such as fiber diameter, electrospinning voltage, and adsorption energy, ensuring balanced feature weights during model training. It avoids the problem of the model focusing on dimensional differences rather than the intrinsic relationship between features and the sustained-release rate due to dimensional inconsistencies, significantly improving the accuracy and stability of prediction results. Simultaneously, by setting `self.scale_[self.scale_==0]=1.0`, it proactively avoids division-by-zero errors that may occur during standardization, ensuring the smoothness and reliability of data preprocessing and solving the problem of computational interruptions or result distortion caused by some existing preprocessing methods not considering extreme numerical cases. This standardization method of this invention is logically rigorous and computationally accurate, making the mean of standardized features approach 0 and the standard deviation approach 1, laying a solid foundation for the efficient training of subsequent gradient boosting tree regression models.

[0033] Furthermore, in S3, the core parameters of the gradient boosting tree regression model are configured as follows: number of decision trees n_estimators=200, maximum depth of decision trees max_depth=5, learning rate learning_rate=0.05, minimum number of leaf samples min_samples_leaf=3, and random seed random_state=42.

[0034] Specifically, this invention achieves an optimal balance between prediction accuracy, generalization ability, and operational efficiency by precisely configuring the core parameters of the gradient boosting tree regression model. This effectively solves the problems of overfitting, weak generalization ability, and unstable output results caused by unreasonable parameter configuration in existing prediction models. Setting the number of decision trees to 200 ensures that the model can fully fit the nonlinear relationship between features and release rate, accurately capturing the interaction between features such as fiber diameter and adsorption energy, and pesticide LogP and molecular weight, while avoiding the waste of computational resources and reduced cost-effectiveness caused by too many trees. The maximum depth of the decision trees is limited to 5, successfully suppressing the problem of the model overlearning training set details and ensuring that the model will not be unable to adapt to samples with measurement errors in actual production due to excessive depth, while retaining the ability to mine deep feature associations. A learning rate of 0.05 reasonably controls the contribution weight of each tree. This approach ensures the effective function of each tree while avoiding significant fluctuations in prediction results through iterative correction across multiple trees, resulting in stronger model stability for nanofiber samples using different electrospinning processes. Setting the minimum sample size for each leaf node to 3 ensures sufficient sample support for each node, significantly improving the model's generalization ability to nanofibers with different substrates such as PLA and chitosan, and avoiding local prediction biases caused by insufficient sample size. A fixed random seed of 42 ensures consistent training results under different environments, guaranteeing the algorithm's reproducibility and laying the foundation for stable application in industrialization. This synergistic optimization of parameters ultimately achieves a determination coefficient of 0.9976, with prediction errors controlled at an extremely low level.

[0035] Furthermore, in S3, the mathematical model implementation of the gradient boosting tree regression model includes: Initial prediction: The mean of the target variable is used as the initial prediction value, and the formula is:

[0036] Residual calculation: Considering sample weights, the residual formula is:

[0037] Decision tree construction: By traversing all features and thresholds, the optimal split point is selected with the goal of minimizing the mean squared error (MSE). The MSE calculation formula is as follows:

[0038] The total error after splitting is the weighted sum of the errors of the left and right subtrees; Prediction Update: After training each tree, through... Update the predicted value, and finally constrain the result to 0%~100% by np.clip(y_pred,0,100) when making the prediction.

[0039] Specifically, this invention ensures the accuracy, rationality, and stability of the gradient boosting tree regression model by clearly defining its complete mathematical implementation logic, effectively solving the problems of ambiguous mathematical logic and insufficient capture of complex nonlinear relationships in some existing prediction models. Using the mean of the target variable as the initial prediction value provides a stable and reliable benchmark for model iterative optimization, avoiding slow convergence or accuracy deviations caused by improper initial value settings. The residual calculation process fully considers sample weights, enabling targeted differentiation of the impact of different samples on model training, allowing the model to more accurately capture the gradual release patterns of key samples. Compared to traditional models that do not consider sample weights, the generalization ability is significantly improved. The decision tree construction aims to minimize the mean squared error (MSE). By traversing all features and thresholds to select the optimal split point, it ensures that each decision tree can fit the residual to the greatest extent possible, and the errors of the left and right subtrees after splitting are minimized. The weighted summation calculation method further ensures the fitting effect of the decision tree. The prediction update formula after training each tree reasonably controls the contribution weight of each tree through the learning rate, which not only gives full play to the optimization role of each tree, but also avoids the local optimum problem through iterative correction of multiple trees. Ultimately, the model can accurately fit the nonlinear relationship between features and the slow-release rate. Furthermore, by using np.clip(y_pred,0,100), the prediction results are constrained to a reasonable range of 0% to 100%, avoiding the prediction values ​​that exceed the actual physical meaning caused by the lack of result constraints in some existing models, thus ensuring the practicality of the prediction results. The synergistic implementation of these mathematical models ultimately achieves a determination coefficient of 0.9976, a prediction error of no more than 2%, and can accurately distinguish slow-release rate differences within 1%.

[0040] Furthermore, in S4, the model is trained by calling `model.fit(X_processed, y)`, and the coefficient of determination R is used. 2 To evaluate the fit, R 2 The calculation formula is as follows:

[0041] Requirement R 2 ≥0.99, where R 2 The value reached 0.9976, indicating excellent fitting performance; The model release_regressor.pkl and the preprocessing tool release_reg_preprocessor.pkl are saved as binary files using pickle.dump, which can be used for subsequent prediction.

[0042] Specifically, this invention trains the model on the preprocessed dataset by calling `model.fit(X_processed, y)`, ensuring that the gradient boosting tree regression model can fully learn the intrinsic relationship between nanofiber characteristics and sustained-release rate, combined with a clear coefficient of determination R0.2 The calculation formula accurately quantifies the model's fitting effect, R. 2 The stringent requirement of ≥0.99 ensures the model's prediction accuracy from a standard perspective, while the actual achieved R-value is 0.9976. 2 The values ​​further demonstrate excellent fitting ability, providing a clear criterion for evaluating model performance. Meanwhile, by using pickle.dump to save the trained model (release_regressor.pkl) and preprocessing tools (release_reg_preprocessor.pkl) as binary files, this saving method not only fully preserves the model's parameter configuration and training results but also supports subsequent quick access and cross-platform use, avoiding the waste of computational resources and increased time costs caused by repeated training.

[0043] Furthermore, in S5, the specific process for determining the burst release risk is as follows: input the six features of the nanofiber to be screened, after standardization by the preprocessing tool, call the trained model to output the 12h / 24h / 48h sustained release rate, preset the 12h sustained release rate threshold, if the output 12h sustained release rate is higher than the threshold, it is determined that there is a burst release risk, which is suitable for the need for rapid control in agricultural scenarios; if the output 12h sustained release rate is lower than or equal to the threshold, it is determined that there is no burst release risk, which is suitable for the need for long-term sustained release in agricultural scenarios.

[0044] Specifically, this invention first inputs six key features of the nanofibers to be screened into a preprocessing tool for standardization, ensuring that the input data is consistent with the feature scale during model training, laying the foundation for accurate prediction. Then, it calls the trained gradient boosting tree regression model to quickly output the sustained-release rates at three key time points: 12h, 24h, and 48h. Relying on the model's high coefficient of determination of 0.9976, the reliability of the sustained-release rate prediction results is guaranteed. The core lies in accurately determining the risk of burst release based on the 12h sustained-release rate. This design precisely fills the gap in existing technologies lacking burst-release risk assessment functions for agricultural scenarios, effectively solving the problems of rapid pesticide loss, reduced efficacy, and environmental pollution caused by burst release within 12 hours. Simultaneously, this judgment logic can flexibly adapt to the needs of different agricultural scenarios. It can screen long-acting sustained-release carriers without burst-release risk for scenarios requiring sustained efficacy and reduced application frequency, and also provide fast-acting carriers with burst-release risk for pest and disease control scenarios requiring rapid effectiveness. This makes the screening of nanopesticide carriers more targeted and practical, significantly improving the matching degree between screening results and actual agricultural production needs.

[0045] A storage medium storing a computer program, which, when executed by a processor, implements the aforementioned algorithm for predicting the sustained release rate of pesticide-loaded nanofibers based on gradient boosting trees and the method for determining the risk of burst release.

[0046] Specifically, this invention stores the corresponding computer program through a storage medium, fully encompassing the core logic of the entire process, from data preparation, standardized preprocessing, gradient boosting tree regression model construction and parameter configuration, to model training and evaluation, multi-time-node sustained-release rate prediction, and burst release risk assessment. This allows the entire algorithm to be easily deployed on various devices supporting Python environments, eliminating the need for repeated algorithm development and model training, effectively ensuring the standardization and consistency of the pesticide-loaded nanofiber sustained-release rate prediction and burst release risk assessment process. This storage method not only completely preserves the trained model parameters and preprocessing tools but also supports cross-platform rapid access, significantly lowering the barrier to entry for the algorithm and avoiding the waste of computational resources and increased time costs caused by repeated training. Compared to existing technologies that lack dedicated storage solutions for this type of prediction algorithm, this invention greatly enhances the algorithm's engineering practicality and industrialization value. With the help of this storage medium, researchers and enterprises can directly call mature algorithms to quickly conduct nanofiber carrier screening, fully leveraging the algorithm's high coefficient of determination (0.9976) and accurate burst release assessment advantages to accelerate the research and development process of nanopesticide sustained-release formulations.

[0047] A computer device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described algorithm for predicting the sustained release rate of pesticide-loaded nanofibers based on gradient boosting trees and the method for determining the risk of burst release.

[0048] Specifically, the computer device of this invention stores a computer program carrying the complete algorithm logic in a memory. Combined with the high-efficiency computing power of the processor, it can comprehensively support the intelligent operation of the entire process, from data preparation, standardized preprocessing, gradient boosting tree model construction and training, to multi-time-node sustained-release rate prediction and burst risk assessment. This effectively solves the problem of existing prediction algorithms lacking dedicated hardware support and being difficult to implement. The memory can completely retain 1000 sets of sample datasets, trained model parameters, and preprocessing tools, avoiding the waste of computing resources caused by repeated training, ensuring algorithm reusability and stable and reproducible prediction results. The processor can quickly call up the stored program and data, efficiently completing complex processes such as feature standardization calculation and model prediction calculation. Relying on the model's high coefficient of determination of 0.9976, it accurately outputs the 12h / 24h / 48h sustained-release rates and quickly completes burst risk assessment based on the 12h sustained-release rate, ensuring the efficiency and accuracy of nanofiber carrier screening. This collaborative hardware and software design fully leverages the gradient boosting tree model's ability to accurately capture the nonlinear relationship between features and slow-release rates, while ensuring screening speed through the processor's efficient computation. Furthermore, its cross-platform compatibility with the Python environment allows the entire algorithm to flexibly adapt to real-world scenarios in agricultural nanomaterials research and development. It supports rapid screening of batch carriers and accurately matches different agricultural needs for long-acting slow-release or rapid-acting control, providing reliable hardware support for the industrialization and promotion of the algorithm and significantly improving the research and development efficiency of nanopesticide carriers.

[0049] The illustrative embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention.

[0050] Example 1: 1. Validation of the effect of the number of decision trees: With other core parameters fixed, three parameter ranges were set: low number range (50-100, test value 80), optimal range (150-250, optimal value 200), and high number range (300-400, test value 350). Under the low number range, the predicted 12-hour sustained-release rate R... 2 =0.9815, not meeting the accuracy standard, large deviation, unable to distinguish subtle differences in sustained release; due to insufficient number of trees, it cannot capture the interaction between fiber diameter and adsorption energy, and between pesticide LogP and molecular weight, resulting in weak generalization ability. The optimal range and optimal value of 200, R... 2 =0.9976, meeting the accuracy standard with minimal deviation, capable of distinguishing sustained-release differences within 1%, effectively fitting the nonlinear relationship between characteristics and sustained-release rate while avoiding overfitting. It exhibits strong generalization ability for nanofibers with different substrates such as PLA and chitosan, with prediction errors ≤1%. In a high-quantity range, R... 2=0.9980, which is only 0.0004 higher than the optimal value, a negligible improvement with no practical significance. Too many trees lead to a waste of computational resources and have extremely low cost-effectiveness.

[0051] 2. Validation of the effect of the maximum depth parameter in decision trees: With other core parameters fixed, three parameter ranges were set: low depth range (2-3, test value 3), optimal range (4-6, optimal value 5), and high depth range (7-8, test value 7). In the low depth range, the training set R... 2 =0.9782, Test set R 2 =0.9756. Due to insufficient depth, it is impossible to explore the deeper pattern that when the fiber diameter is <100nm, the 12h sustained release rate increases by 5% for every 10° decrease in the contact angle. It is only suitable for coarse screening. The optimal range and optimal value are found in the training set R. 2 =0.9985, Test set R 2 =0.9976, which can fully explore feature associations while avoiding measurement errors in the training set. At high depth ranges, the training set R... 2 =0.9998, Test set R 2 =0.9891, the model overlearns the details of the training set, and the prediction bias for samples with fiber diameter measurement error >3nm increases sharply, making it unable to adapt to actual production samples.

[0052] 3. Validation of the learning rate parameter: With other core parameters fixed, three parameter ranges were set: low learning rate range (0.01-0.03, test value 0.02), optimal range (0.04-0.06, optimal value 0.05), and high learning rate range (0.08-0.10, test value 0.09). Under the low learning rate range, R... 2 The accuracy of R² is 0.9921, which is not optimal, and training time increases. The optimal range and optimal value are discussed below. 2 With an accuracy of 0.9976, it achieves high-precision predictions without requiring additional trees, ensuring the effective contribution of each tree while avoiding fluctuations through iterative correction using multiple trees. It demonstrates strong prediction stability for nanofiber samples using different electrospinning processes, with error fluctuations ≤0.3%. At high learning rates, R... 2 =0.9853 The accuracy drops sharply, the weight of a single tree is too large, and it is easily affected by abnormal samples with adsorption energy measurement errors. The prediction error difference of the same batch of samples reaches 2%, which cannot guarantee the consistency of screening.

[0053] 4. Verification of the comprehensive effect of the optimal parameter combination: Based on the optimal parameter combination determined in the above experiments, the prediction error of the model under the optimal parameter combination is ≤0.6%, which can replace the traditional experimental measurement process (traditional requires 3 days, this algorithm only takes 30 seconds), reducing R&D costs by more than 80% and meeting the needs of industrialization screening. The optimal range for the number of decision trees is 150-250, with an optimal value of 200, corresponding to R...2 The optimal parameters are: maximum depth ≥0.997, training time ≤25 seconds. The optimal range for the maximum decision tree depth is 4-6, with an optimal value of 5, corresponding to no overfitting, strong generalization ability, and test set error ≤0.9%. The optimal range for the learning rate is 0.04-0.06, with an optimal value of 0.05, corresponding to good stability, high efficiency, and error fluctuation ≤0.3%. The optimal parameter combination can achieve a comprehensive effect of "high precision, high efficiency, and high adaptability," providing strong support for the efficient research and development and industrial screening of pesticide-loaded nanofibers.

[0054] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A gradient boosting tree-based algorithm for predicting the sustained release rate of pesticide-loaded nanofibers and a method for determining the risk of burst release, characterized in that, Includes the following steps: S1. Data preparation: Construct a dataset containing nanofiber and pesticide-related characteristics and slow-release rates; S2. Preprocess the dataset to eliminate differences in feature dimensions; S3. Construct a gradient boosting tree regression model and configure the core parameters of the model; S4. Train the gradient boosting tree regression model using the preprocessed dataset, and evaluate the model's fit using performance metrics. S5. Load the trained gradient boosting tree regression model and preprocessing tools, input the features of the nanofibers to be screened, predict the sustained release rate at multiple time points, and determine the burst release risk based on the 12-hour sustained release rate.

2. The pesticide-loaded nanofiber slow-release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree as described in claim 1, characterized in that, In S1, the dataset is nanofiber_release_data.csv containing 1000 samples. The dataset is constructed by extracting 200 sets of experimental data from a professional nanofiber database and supplementing them with 800 sets of similar data. Each sample contains 6 input features and 1 output target. The input features are fiber diameter, contact angle, pesticide LogP, pesticide molecular weight, electrospinning voltage, and adsorption energy. The output target is the 12h / 24h / 48h release rate, and the release rate ranges from 0% to 100%. The dataset is read using the pd.read_csv function.

3. The pesticide-loaded nanofiber slow-release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree as described in claim 2, characterized in that, In S2, the preprocessing is a standardization process, implemented using the fit and transform methods. The fit method calculates the mean and standard deviation of each feature using the following formula: The transform method standardizes the features using the following formula: Avoid division by zero by setting self.scale_[self.scale_==0]=1.

0.

4. The pesticide-loaded nanofiber slow-release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree according to claim 3, characterized in that, In S3, the core parameters of the gradient boosting tree regression model are configured as follows: number of decision trees n_estimators=200, maximum depth of decision trees max_depth=5, learning rate learning_rate=0.05, minimum number of leaf samples min_samples_leaf=3, and random seed random_state=42.

5. The pesticide-loaded nanofiber slow-release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree according to claim 4, characterized in that, In S3, the mathematical model implementation of the gradient boosting tree regression model includes: Initial prediction: The mean of the target variable is used as the initial prediction value, and the formula is: Residual calculation: Considering sample weights, the residual formula is: Decision tree construction: By traversing all features and thresholds, the optimal split point is selected with the goal of minimizing the mean squared error (MSE). The MSE calculation formula is as follows: The total error after splitting is the weighted sum of the errors of the left and right subtrees; Prediction Update: After training each tree, through... Update the predicted value, and finally constrain the result to 0%~100% by np.clip(y_pred,0,100) when making the prediction.

6. The pesticide-loaded nanofiber slow-release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree according to claim 5, characterized in that, In S4, the model is trained by calling `model.fit(X_processed, y)`, and the model is then trained using the coefficient of determination R0. 2 To evaluate the fit, R 2 The calculation formula is as follows: Requirement R 2 ≥0.99; The model release_regressor.pkl and the preprocessing tool release_reg_preprocessor.pkl are saved as binary files using pickle.dump.

7. The pesticide-loaded nanofiber slow-release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree according to claim 6, characterized in that, In S5, the specific process for determining the burst release risk is as follows: input the six features of the nanofiber to be screened, after standardization by the preprocessing tool, call the trained model to output the 12h / 24h / 48h sustained release rate, preset the 12h sustained release rate threshold, if the output 12h sustained release rate is higher than the threshold, it is determined that there is a burst release risk, which is suitable for the need for rapid control in agricultural scenarios; if the output 12h sustained release rate is lower than or equal to the threshold, it is determined that there is no burst release risk, which is suitable for the need for long-term sustained release in agricultural scenarios.

8. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the pesticide-loaded nanofiber slow-release rate prediction algorithm and burst release risk determination method based on gradient boosting tree as described in any one of claims 1-7.

9. A computer device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the pesticide-loaded nanofiber sustained-release rate prediction algorithm and burst release risk assessment method based on gradient boosting tree as described in any one of claims 1-7.