Prediction method for nutrient content of hermetia illucens body

By establishing a nutritional component prediction model for black soldier fly bodies based on machine learning, the time-consuming and labor-intensive problem of traditional methods is solved, efficient and accurate nutritional component prediction and process optimization are achieved, and intelligence and precision of multi-source waste treatment scenarios are supported.

CN120340653APending Publication Date: 2025-07-18TIANJIN AGRICULTURE COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510297520.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional methods are difficult to quickly and accurately predict the nutrients of black soldier fly bodies, cannot meet the needs of large-scale breeding, and the existing technology is time-consuming and labor-intensive.

Method used

By obtaining multi-dimensional data of black soldier flies and combining machine learning algorithms, random forest (RF), support vector machine (SVM), XGBoost and GBRT models were established to predict the nutritional content of black soldier flies, and the SHAP value was used to analyze and interpret the model.

Benefits of technology

The rapid and accurate prediction of the nutrient components of the black soldier fly body has been achieved, the efficiency has been increased by 800 times, the cost has been reduced to 1%, the model is highly interpretable, and it supports multi-source waste treatment scenarios and promotes the intelligence and precision of the resource utilization of agricultural waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340653A_ABST
    Figure CN120340653A_ABST
Patent Text Reader

Abstract

The invention relates to a method for predicting the nutrient content of a hermetia illucens body. The method comprises the following steps: S1, acquiring sample data of the nutrient content of the hermetia illucens body and preprocessing the sample data; s2, based on the sample data processed in the step S1, establishing a machine learning model for predicting the nutrient content of the hermetia illucens body, and optimizing the prediction model; and S3, verifying and explaining the trained machine learning model for predicting the nutrient content of the hermetia illucens body based on a random forest (RF), a support vector machine (SVM), an XGBoost model and a GBRT model. According to the method, the multi-dimensional data of the hermetia illucens bodies are acquired, and the prediction model is established in combination with the machine learning algorithm, so that the nutritional ingredients of the hermetia illucens bodies can be rapidly and accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of agriculture and biotechnology, and relates to a method for predicting the content of nutritional components, in particular to a method for predicting the content of nutritional components in the body of Hermetia illucens. Background Art

[0002] With the growth of the population and the acceleration of the urbanization process, the treatment of organic waste has become an increasingly serious environmental challenge. Organic wastes such as food waste and livestock manure not only occupy a large amount of land resources, but also may cause environmental pollution and disease transmission. At present, the conventional disposal methods of organic solid waste in China mainly include landfill and incineration. However, landfill occupies a large area and pollutes land resources, while incineration has high operating costs and is prone to generate toxic and harmful gases. Therefore, it is particularly important to explore efficient and sustainable ways for the treatment and resource utilization of organic waste.

[0003] Meanwhile, with the rapid development of the economic society, the demand for high-quality protein sources in the aquaculture industry is increasing, while animal protein sources, especially fish meal resources, are scarce and expensive. The research on substituting fish meal protein sources has become a research hotspot in animal nutrition. Although certain progress has been made in the research on substituting fish meal with plant protein sources, there are many limiting factors in plant protein sources, such as poor palatability, amino acid imbalance, and the presence of anti-nutritional factors, etc., which cannot be widely applied. Therefore, the research on animal protein sources substituting fish meal has attracted much attention.

[0004] Hermetia illucens, also known as black soldier fly, is an insect of the genus Hermetia in the family Stratiomyidae of the order Diptera. It originated from the grasslands of South America and is now widely distributed in tropical and subtropical regions around the world. Hermetia illucens has a wide range of food habits and is saprophagous, feeding on decaying organic matter in nature. Its high conversion efficiency can be used to degrade organic pollutants including food waste and animal manure, and it is a resourceful insect. Through continuous in-depth research, the use of Hermetia illucens to degrade solid organic pollutants has gradually come into people's view. The biodegradation of organic pollutants by Hermetia illucens realizes the resource recovery of waste from the source, and at the same time can produce livestock protein feed. Research shows that Hermetia illucens larvae can not only effectively reduce the volume and weight of organic waste, but also significantly reduce the harmful microorganisms and nitrogen content in the waste, which is of great significance to environmental protection.

[0005] The crude protein content of black soldier fly larvae can reach over 30%, and the crude fat content can reach over 20%. Moreover, it contains various essential amino acids and trace elements, having relatively high nutritional value, which makes black soldier fly an ideal source of protein and fat. In addition, black soldier fly is also rich in essential amino acids and minerals, such as lysine, methionine, and arginine; and the amino acid composition ratio of black soldier fly is balanced, especially the contents of components such as glutamic acid, lysine, and glycine are abundant, which provides strong support for adjusting the feed formula to meet the specific needs of different livestock and poultry.

[0006] In recent years, scholars at home and abroad have conducted extensive research on black soldier fly. Some research has pointed out that the content of nutritional components in the insect body during the breeding process of black soldier fly is affected by various factors, including substrate components, environmental conditions, breeding management, etc. Studying the effects of five organic feeds, namely food waste, wheat bran, alfalfa powder, corn flour, and chicken manure, on the production performance and nutritional components of black soldier fly larvae, the results show that the black soldier fly larvae fed with food waste have the highest contents of crude protein, crude fat, and salt, and the insect conversion rate is significantly higher than that of other groups under the same feeding amount, the insect body is large, the appearance is good, and the larval development time is short. When conducting feeding experiments using four organic wastes, namely manure, banana peel, beer waste, and food waste, it is found that the crude protein of the larvae fed with manure and beer waste is significantly higher than the other two groups.

[0007] Some research has evaluated the growth status of black soldier fly at temperatures of 27 - 36 °C and found that 27 °C is the optimal environmental temperature for the growth and development of black soldier fly and the temperature is proportional to the development rate. Breeding density, light conditions, and the type of food waste also have significant effects on the growth of black soldier fly. Excessive breeding density will lead to problems such as slow growth of black soldier fly larvae and reduced feed utilization rate. There is a close relationship between the composition and proportion of the nutritional components of black soldier fly and the nutritional composition of the food it eats. Although the insect body of the larvae fed with animal viscera can enrich more crude protein and linoleic acid, it has the disadvantages of slow growth and low pupation rate, while the larvae fed with food waste have the best growth and development, and also enrich more insect body protein and Omega-3 fatty acids. At the same time, light can also control the growth and development process of black soldier fly larvae, and the shading environment is more conducive to the growth of larvae and the accumulation of their body components.

[0008] The complexity and differences of these factors make it particularly difficult to predict and control the nutritional components of black soldier fly insect body.

[0009] Traditional methods rely on experience accumulation and experimental determination, which are not only time-consuming and laborious, but also difficult to meet the needs of large-scale breeding. Therefore, developing an efficient and accurate prediction model for the nutritional components of black soldier fly insect body is of great significance for optimizing the breeding process of black soldier fly, improving product quality, and enhancing market competitiveness.

[0010] Therefore, the present invention proposes a method for predicting the content of nutritional components of black soldier fly insect body. Summary of the Invention

[0011] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a method for predicting the nutrient component content of black soldier fly larvae. By obtaining multi-dimensional data of black soldier fly larvae and combining machine learning algorithms, a prediction model is established, so as to quickly and accurately predict the nutrient components of black soldier fly larvae, and has a wide range of application prospects.

[0012] The present invention solves its practical problems by adopting the following technical solutions:

[0013] A method for predicting the nutrient component content of black soldier fly larvae, comprising the following steps:

[0014] S1. Obtain sample data of the nutrient components of black soldier fly larvae and perform preprocessing;

[0015] S2. Based on the sample data processed in step 1, establish a machine learning model for predicting the nutrient component content of black soldier fly larvae and optimize the prediction model;

[0016] S3. Verify and interpret the machine learning model for predicting the nutrient component content of black soldier fly larvae trained in step S2 based on the random forest (RF), support vector machine (SVM), XGBoost and GBRT models, and use R 2 Evaluate the prediction effect of the prediction model, draw the feature importance based on the SHAP value and the SHAP summary graph, and output the prediction effect by inputting the characteristic parameters of the experiment;

[0017] Moreover, the specific steps of step 1 include:

[0018] (1) Collect valid sample data of the nutrient components of black soldier fly larvae from the database, including crude protein in larvae, crude fat in larvae, feeding substrate, types of exogenous additives, feeding environment temperature, humidity, initial larval age, transformation period, light cycle and initial larval weight;

[0019] (2) Perform preprocessing on the collected valid sample data of the nutrient components of black soldier fly larvae;

[0020] Moreover, the specific steps of step (2) of step 1 include:

[0021] According to the characteristics of the data set, some data need to be processed before being applied to the model as input features;

[0022] (1) Data encoding and enhancement processing

[0023] The feeding substrate is encoded using one-hot encoding to avoid the sequential misguidance caused by numerical labels;

[0024] The types of exogenous additives are encoded in segments: 0 indicates no additive, 1 - 3 indicate the low-dose group, and 4 - 6 indicate the high-dose group, enhancing the sensitivity of the model to the dose effect;

[0025] (2) Data cleaning and standardization

[0026] Outlier handling: Use SPSS to perform box plot analysis on continuous variables and delete samples beyond 1.5 times the IQR;

[0027] Missing value filling: Fill continuous variables with the median and discrete variables with the mode;

[0028] Standardization processing: Perform Z-score standardization on continuous features such as temperature and humidity. The formula is:

[0029]

[0030] where μ is the mean and σ is the standard deviation. The light cycle is normalized to [0, 1] through Min - Max normalization;

[0031] Moreover, the specific steps of step 2 include:

[0032] (1) Establish machine learning models for predicting the nutritional component content of black soldier fly larvae based on Random Forest (RF), Support Vector Machine (SVM), XGBoost, and GBRT models respectively and conduct training;

[0033] (2) Optimize the model parameters through grid search respectively and introduce dynamic error weights to adjust and strengthen the learning of difficult samples;

[0034] (3) Perform principal component analysis PCA dimensionality reduction on multi - collinear features respectively to generate trained machine learning models for predicting the nutritional component content of black soldier fly larvae based on Random Forest RF, Support Vector Machine SVM, XGBoost, and GBRT models;

[0035] Moreover, the specific steps of step S3 include:

[0036] (1) Optimize the model parameters according to the size of the model evaluation metrics. After completing data collection and processing, when selecting a machine learning model, it is necessary to adjust the model parameters and optimize the model parameters by observing the best results of each run. The CART optimization parameters of the XGBoost model include gamma, max_depth, lambda, subsample, comsample_bytree, min_child_weight, eta. The model optimization parameters of RF include n_estimators, random_state, max_depth, max_features, min_samples_leaf, min_samples_split. The parameter optimization of SVM includes gamma, kernel. The optimization parameters of the GBRT model are n_estimators, learning_rate, max_depth, random_state;

[0037] (2) For the evaluation metrics of the regression prediction model, due to the large difference in the eigenvalue of the black soldier fly breeding environment, use R 2 for model evaluation. The closer R 2 is to 1, the better the model prediction effect;

[0038]

[0039] Where: n is the number of experimental data used in training;

[0040] y pre 、y exp and are the predicted values of the nutrient content of the insect body obtained from the model, experimental data, and average experimental data respectively.

[0041] (3) Use SHAP analysis to calculate the marginal contribution of the model, perform global and local interpretations, and identify key features. The specific calculation is as follows:

[0042] 1) Conduct feature importance calculation and analysis: Calculate Gain, that is, the average information gain brought by a certain feature after splitting in all trees. The specific calculation steps are as follows. Obtain the feature importance through sorting by average information gain and draw a feature importance graph;

[0043]

[0044] Where: Ent(D) is the information entropy of set D;

[0045] θ is the number of subsets after partitioning;

[0046] is the proportion of the partitioned set in the original set;

[0047] 2) Perform SHAP model interpretation. By calculating the Shapley values of each feature, obtain the contribution of each feature to the model regression and show the positive or negative of this contribution. Increase the interpretability of the model and draw a feature summary diagram for intuitive analysis of model interpretation. The specific calculation steps are as follows;

[0048]

[0049] Where: is the contribution of the i-th feature;

[0050] S is a subset of the given predictive features;

[0051] A is the set of all features;

[0052] f x (S∪{i}), f x (S) include the model results with and without the i-th feature.

[0053] (4) Input the feature parameters in the actual experiment according to the code prompt and display the prediction results.

[0054] Advantages and beneficial effects of the present invention:

[0055] 1. High-efficiency and accurate prediction performance

[0056] Efficiency breakthrough: The traditional chemical analysis method takes more than 72 hours and the cost per single sample exceeds 200 yuan. The present invention can complete the prediction within 5 minutes, with the single cost less than 2 yuan (computer resources), the efficiency is increased by more than 800 times, and the cost is reduced to 1%.

[0057] Precision advantage: In crude protein prediction, R 2 = 0.89 ± 0.03, significantly better than the traditional model (R 2 ≤ 0.85).

[0058] 2. Model interpretability and decision guidance;

[0059] Black box transparency: Based on the global feature importance analysis and local sample contribution analysis of SHAP (Shapley value), clearly quantify the influence weights of key factors such as temperature and humidity on nutrient components, break through the limitations of the traditional machine learning "black box", and provide a quantifiable basis for process optimization.

[0060] 3. Technical universality

[0061] The present invention can be migrated to other organic solid waste (such as municipal sludge, kitchen waste) treatment scenarios and support the co-disposal of multi-source waste;

[0062] The present invention deeply integrates machine learning prediction with process control, not only achieving rapid and low-cost detection of nutrient components, but also revealing key regulatory factors through interpretability analysis, providing a technical closed-loop for the precision and intelligence of the black soldier fly breeding process, and promoting the transformation of agricultural waste resource utilization from experience-driven to data-driven. Brief Description of the Drawings

[0063] Figure 1 is the processing flow chart of the present invention;

[0064] Figure 2 is the comparison chart of the predicted values and the true values of the training set and the test set for predicting the crude protein content based on four machine learning models of the present invention;

[0065] Figure 3 is the comparison chart of the predicted values and the true values of the training set and the test set for predicting the crude fat content based on four machine learning models of the present invention;

[0066] Figure 4 is the feature importance chart of the present invention's features for predicting the crude protein content of the insect body;

[0067] Figure 5 is the feature importance chart of the present invention's features for predicting the crude fat content of the insect body;

[0068] Figure 6 is the influence of the present invention's features on the prediction result of the crude protein content model of the insect body, that is, the SHAP summary chart;

[0069] Figure 7 is the influence of the present invention's features on the prediction result of the crude fat content model of the insect body, that is, the SHAP summary chart;

[0070] Figure 8 is the output result chart of the present invention in practical applications. Detailed Embodiments

[0071] The following further details the embodiments of the present invention in conjunction with the drawings:

[0072] A method for predicting the nutrient component content of black soldier fly insect bodies, as Figure 1 shown, includes the following steps:

[0073] S1. Obtain sample data of the nutrient components of black soldier fly insect bodies and perform preprocessing;

[0074] The specific steps of step 1 include:

[0075] (1) Collect valid sample data of the nutrient components of black soldier fly insect bodies from the database, including crude protein in the insect body, crude fat in the insect body, feeding substrate, types of exogenous additives, feeding environmental temperature, humidity, initial larval age, transformation cycle, light cycle, and initial insect weight;

[0076] In this embodiment, origin2022 is used to extract features from the collected data. The data features include crude protein (CP) of insect bodies, crude fat (EE) of insect bodies, feeding substrate (BD), exogenous additives (ADD), feeding environment temperature (TEMP), feeding environment humidity (HUMIDITY), initial larval instar (INSTAR), transformation cycle (EXPT DAYS), light cycle (PP), and initial larval weight (WT).

[0077] (2) Preprocess the valid data of the collected samples of the nutritional components of black soldier fly larvae.

[0078] The specific steps of step (2) of S1 include:[[]]

[0079] According to the characteristics of the data set, some data need to be processed before being applied to the model as input features.

[0080] (1) Data encoding and augmentation processing

[0081] The feeding substrate uses one-hot encoding to avoid the sequential misguidance caused by numerical labels.

[0082] The types of exogenous additives (n types) are encoded in segments: 0 indicates no additives, 1 - 3 indicate the low-dose group (≤3 types), and 4 - 6 indicate the high-dose group (>3 types) to enhance the sensitivity of the model to the dose effect.

[0083] (2) Data cleaning and standardization

[0084] Outlier processing: Use SPSS to perform box plot analysis on continuous variables (such as temperature and humidity), and delete samples that exceed 1.5 times the IQR (a total of 42 groups are excluded).

[0085] Missing value filling: Continuous variables are filled with the median, and discrete variables are filled with the mode.

[0086] Standardization processing: Perform Z-score standardization on continuous features such as temperature and humidity. The formula is:

[0087]

[0088] where μ is the mean and σ is the standard deviation. The light cycle is normalized to [0, 1] through Min - Max normalization.

[0089] S2. Based on the sample data processed in step 1, establish a machine learning model for predicting the nutritional component content of black soldier fly larvae and optimize the prediction model:

[0090] The specific steps of step 2 include:

[0091] (1) Establish machine learning models for predicting the nutrient component content of black soldier fly larvae based on the random forest (RF), support vector machine (SVM), XGBoost, and GBRT models respectively, and train them.

[0092] In this embodiment, four machine learning models, namely the random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), and gradient boosting regression tree (GBRT), are used for comparative analysis. The same data is used to predict the crude protein and crude fat content of the larvae on the models.

[0093] (2) Optimize the model parameters through grid search respectively, and introduce dynamic error weights to adjust and strengthen the learning of difficult samples.

[0094] In this embodiment, error feedback iteration: introduce dynamic error weight adjustment during model training. For samples with a prediction residual (|y_pre - y_exp|) greater than 15%, increase their sample weights in the next round of training to strengthen the model's learning of difficult samples.

[0095] (3) Perform principal component analysis (PCA) dimensionality reduction on the multi - collinear features respectively to generate trained machine learning models for predicting the nutrient component content of black soldier fly larvae based on the random forest (RF), support vector machine (SVM), XGBoost, and GBRT models.

[0096] Multi - collinearity detection: Use variance inflation factor (VIF) analysis to identify high - collinearity feature groups with VIF > 10. Perform principal component analysis (PCA) on each group of high - collinearity features to generate orthogonal principal component features with an explained variance > 95%.

[0097] S3. Verify and interpret the machine learning models for predicting the nutrient component content of black soldier fly larvae based on the random forest (RF), support vector machine (SVM), XGBoost, and GBRT models trained in step S2. Use R 2 Evaluate the prediction effect of the prediction model, draw the feature importance based on SHAP values and the SHAP summary graph, and output the prediction effect through the input feature parameters of the experiment.

[0098] The specific steps of step S3 include:

[0099] (1) Optimize the model parameters according to the magnitudes of the model evaluation metrics. After completing data collection and processing, when selecting a machine learning model, it is necessary to adjust the model parameters and optimize the model parameters by observing the best results of each run. The CART optimization parameters of the XGBoost model include gamma, max_depth, lambda, subsample, comsample_bytree, min_child_weight, and eta. The model optimization parameters of RF include n_estimators, random_state, max_depth, max_features, min_samples_leaf, and min_samples_split. The parameter optimization of SVM includes gamma and kernel. The optimization parameters of the GBRT model are n_estimators, learning_rate, max_depth, and random_state;

[0100] (2) For the evaluation metrics of the regression prediction model, due to the large differences in the characteristic values of the black soldier fly breeding environment, use R 2 to evaluate the model. The closer R 2 is to 1, the better the model prediction effect;

[0101]

[0102] where: n is the number of experimental data used in training;

[0103] y pre 、y exp and are the predicted values of the nutritional component content of the insect body obtained from the model, experimental data, and average experimental data, respectively.

[0104] (3) Use SHAP analysis to calculate the marginal contribution of the model for global and local interpretations, and identify key features. The specific calculations are as follows:

[0105] 1) Calculate and analyze feature importance: Calculate Gain, which is the average information gain brought by a certain feature after splitting in all trees. The specific calculation steps are as follows. Obtain the feature importance by sorting the average information gain and draw a feature importance graph;

[0106]

[0107] where: Ent(D) is the information entropy of set D;

[0108] θ is the number of subsets after partitioning;

[0109] is the proportion of the partitioned set in the original set;

[0110] 2) Perform SHAP model interpretation. By calculating the Shapley values of each feature, obtain the contribution of each feature to the model regression and show the positive or negative of this contribution; increase the interpretability of the model and draw a feature summary graph for intuitive analysis of model interpretation. The specific calculation steps are as follows;

[0111]

[0112] Where: is the contribution of the i-th feature;

[0113] S is a subset of the given predictive features;

[0114] A is the set of all features;

[0115] f x (S ∪ {i}), f x (S) include the model results with and without the i-th feature.

[0116] (4) Input the feature parameters in the actual experiment according to the code prompt and display the prediction results;

[0117] The present invention will be further described below through specific examples:

[0118] A prediction method for the nutritional component content of black soldier fly larvae includes the following steps:

[0119] S1. Collect data and preprocess the data;

[0120] S1 includes the following specific steps:

[0121] S11. Data source and extraction

[0122] Collect black soldier fly breeding experiment data from public databases (CNKI, Web of Science) and experimental breeding, and 575 groups remain after removing outliers.

[0123] Use Origin2022 to extract the following features: crude protein (CP) in the larvae, crude fat (EE) in the larvae, feeding substrate (BD), exogenous additive (ADD), feeding environment temperature (TEMP), feeding environment humidity (HUMIDITY), initial larval instar (INSTAR), transformation cycle (EXPT DAYS), photoperiod (PP), and initial larval weight (WT).

[0124] S12. Process and calculate the extracted data.

[0125] S12 includes the following specific steps:

[0126] (1) Data Encoding and Enhancement Processing

[0127] The feeding substrate uses one-hot encoding to avoid the sequential misguidance caused by numerical labels;

[0128] The types of exogenous additives (n types) are encoded in segments: 0 indicates no additive, 1-3 indicates the low-dose group (≤3 types), and 4-6 indicates the high-dose group (>3 types) to enhance the sensitivity of the model to the dose effect.

[0129] (2) Data Cleaning and Standardization

[0130] Outlier handling: Use SPSS to perform box plot analysis on continuous variables (such as temperature and humidity), and delete samples exceeding 1.5 times the IQR (a total of 42 groups were excluded).

[0131] Missing value filling: Continuous variables are filled with the median, and discrete variables are filled with the mode.

[0132] Standardization processing: Perform Z-score standardization on continuous features such as temperature and humidity. The formula is:

[0133]

[0134] where μ is the mean and σ is the standard deviation. The light cycle is normalized to [0,1] through Min-Max normalization.

[0135] S2. Machine Learning Modeling:

[0136] (1) Model Selection;

[0137] (2) Model Training and Parameter Optimization;

[0138] (3) Model Evaluation;

[0139] S2 includes the following specific steps:

[0140] S21. Model Selection and Parameter Tuning

[0141] Select four models, namely Random Forest (RF), Support Vector Machine (SVM), XGBoost, and GBRT, for analysis and prediction.

[0142] Parameter Optimization Method:

[0143] For RF, XGBoost, and GBRT, the optimal parameters are determined through grid search. For SVM: Use the RBF kernel function, gamma = 0.1, C = 10000.

[0144] S23. Dynamic Weight Adjustment and Feature Orthogonalization

[0145] For samples with a predicted relative error (|(y_pre - y_exp) / y_exp|) > 15%, adjust their sample weights to twice the current value in the next round of training, and normalize the weights by dividing by the mean.

[0146] Two groups of highly collinear features (temperature and humidity, feeding substrate and initial insect age) were identified through VIF analysis. PCA dimensionality reduction was performed on each group of features to ensure that the cumulative explained variance of each group of principal components > 95%, and the generated orthogonalized principal components were combined with the remaining features into a new feature set.

[0147] S24. The evaluation metrics of the regression prediction model mainly include the coefficient of determination (R 2 ) and the root mean square error (RMSE). Compared with RMSE, which is sensitive to absolute errors, the R 2 index, which characterizes the relative explanatory ability, can more effectively evaluate the applicability of the model. Due to the large difference in eigenvalue of the black soldier fly breeding environment, R 2 is used for model evaluation.

[0148] For R 2 calculation, the calculation formula is as follows

[0149]

[0150] where n is the number of experimental data used in training, ypre, yexp, and are the predicted values obtained from the model, experimental data, and average experimental data respectively.

[0151] To better understand the internal operation of the model, the feature importance of the model was calculated.

[0152] S25. For four models, namely Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Regression Tree (GBRT), and Extreme Gradient Boosting (XGBoost), SHAP analysis was performed.

[0153] SHAP reveals the decision-making basis of the black-box model from two levels: the global feature importance ranking and the local individual sample interpretation by quantifying the marginal contribution of features to the model prediction. The calculation result graph based on Python analysis shows the importance of each feature. By calculating the importance degree of each feature, it further guides the parameter optimization and regulation of the black soldier fly composting process.

[0154] Based on the SHAP value, calculate the feature marginal contribution degree, generate the feature importance graph and summary graph. The main calculation methods are as follows:

[0155] 1) Conduct feature importance calculation and analysis: Calculate Gain, that is, the average information gain brought by a certain feature after splitting in all trees, and obtain the feature importance through sorting by the average information gain:

[0156]

[0157] Where: Ent(D) is the information entropy of set D;

[0158] θ is the number of sets after partitioning;

[0159] is the proportion of the partitioned set in the original set;

[0160] 2) Perform SHAP model interpretation. By calculating the Shapley values of each feature, obtain the contribution of each feature to the model regression, and show the positive and negative of this contribution, increasing the interpretability of the model. Use the feature summary graph of SHAP for intuitive analysis of model interpretation;

[0161]

[0162] Where: is the contribution of the i-th feature;

[0163] S is a subset of the given predictive features;

[0164] A is the set of all features;

[0165] f x (S ∪ {i}), f x (S) includes and does not include the model results of the i-th feature.

[0166] Figure 2 and Figure 3 are the comparison graphs of the predicted values and the true values of the training set and the test set for predicting the crude protein and crude fat contents of the insect body based on four machine learning models; Figure 4 and Figure 5 is the importance of the features of the present invention for predicting the crude protein and crude fat contents of the black soldier fly insect body; Figure 6 and Figure 7 is the influence of the features of the present invention on the model prediction results of the crude protein and crude fat contents, that is, the SHAP summary graph; Figure 8 is the output result of the present invention in practical applications;

[0167] The present invention selects a total of 8 features (feeding substrate, exogenous additive, feeding environment temperature, feeding environment humidity, initial larval age, conversion period, light cycle, and initial larval weight) to predict the nutritional component content of the black soldier fly insect body.

[0168] Select four models of random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), and gradient boosting regression tree (GBRT) for learning and calculation. Figure 2 (crude protein) and Figure 3(Crude fat) shows the comparison results of the true values and predicted values of the training set and test set of four machine learning models.

[0169] In some data cases, model adjustment can obtain better results than in general cases. By adjusting the parameters, the evaluation metrics of the displayed models are compared, and through comparison, the nutrient component content of black soldier fly larvae can be predicted better.

[0170] After four machine learning models are used to predict the crude protein content of black soldier fly larvae, the model evaluation metrics R of the model training set and test set 2 As shown in Table 1, for the model prediction of the crude fat content of larvae, the model evaluation metrics R of the training set and test set 2 As shown in Table 2.

[0171] Through the comprehensive evaluation of the models, it can be seen that when using machine learning models to predict the nutrient component content of black soldier fly larvae, the prediction effect of the support vector machine (SVM) model is better. And when predicting the crude protein content of larvae, the R values of the training set and test set 2 are 0.912 and 0.909 respectively, and the prediction error is relatively low, indicating that it has a high prediction accuracy.

[0172] Table 1 Comparison of model evaluation metrics for the prediction of crude protein content of larvae in the training set and test set

[0173]

[0174]

[0175] Table 2 Comparison of model evaluation metrics for the prediction of crude fat content of larvae in the training set and test set

[0176]

[0177] Importance of features for the prediction of nutrient component content of larvae and SHAP summary diagram

[0178] Shapley values are used to explain the influence degree of each feature in the model on the prediction result, and the support vector machine (SVM) model has a better prediction effect. The SHAP model explanations are carried out for the four models respectively. Specifically as follows:

[0179] 1. Feature importance

[0180] The idea behind SHAP feature importance is very simple. Features with larger absolute Shapley values are important. Since global importance is required, the average value of the absolute Shapley values of each feature in the data is taken:

[0181]

[0182] Feature importance can directly reflect the importance of features and see which features have a greater impact on the final model. According to the calculation results of Python analysis, the features are sorted in descending order of importance and plotted. Figure 4 (crude protein) and Figure 5 (Crude fat) shows the importance of each feature, among which the two features that have a particularly prominent effect on the crude protein content of the insect body are temperature (TEMP) and initial larval age (INSTAR). The four features that have a particularly prominent effect on the crude fat content of the insect body are temperature (TEMP), humidity (HUMIDITY), stocking density (BD) and initial larval weight (WT)

[0183] 2. SHAP Summary

[0184] The SHAP summary graph combines feature importance and feature influence. Each point on the summary graph is the Shapley value of a feature and an instance. The position on the y-axis is determined by the feature, and the position on the x-axis is determined by the Shapley value. The color represents the feature value from small to large, and the overlapping points jitter in the y-axis direction. The position of each feature in the graph can represent the contribution of the feature. Red represents a high value of the feature, and blue represents a low value of the feature. In the vertical direction, the higher the feature value, the greater the contribution of the feature. For example: Figure 6 (c) It can be seen that the larger the initial larval age (INSTAR) of the experiment, the higher the crude protein content of the insect body.

[0185] Expected effect comparison data

[0186]

[0187] The working principle of the present invention is:

[0188] In recent years, the rapid development of machine learning technology has provided a powerful tool for modeling and predicting complex systems. Machine learning algorithms can learn from large amounts of data and discover hidden patterns and laws, and then predict and classify unknown data. In the fields of agriculture, biotechnology, and environmental protection, machine learning has been successfully applied to crop growth prediction, pest and disease identification, environmental quality monitoring, and many other aspects. Therefore, applying machine learning technology to the prediction of insect nutritional components in the process of black soldier fly breeding is expected to achieve accurate prediction and dynamic regulation of nutritional components, providing strong support for the intelligent and efficient breeding of black soldier flies.

[0189] The innovation of the present invention lies in:

[0190] An innovative multi-dimensional data collection scheme has been designed and implemented, which not only includes biological characteristic data (such as initial insect weight, insect age, etc.), but also introduces environmental factor data (such as feeding temperature, humidity, feed type, etc.). Through data fusion technology, these data from different sources and different scales are effectively integrated, providing a rich and comprehensive input feature set for the machine learning model..

[0191] For the specific problem of predicting the nutritional components of black soldier flies, customized machine learning models, namely Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Regression Tree (GBRT) and Extreme Gradient Boosting (XGBoost), have been developed. These models can handle complex non-linear relationships and improve the prediction accuracy.

[0192] Feature selection and importance evaluation methods are introduced to automatically screen out the feature variables that have the most influence on the prediction results, optimize the model structure, reduce the risk of overfitting, and improve the computational efficiency at the same time.

[0193] In subsequent experiments, a strategy for dynamically adjusting the feeding conditions can be designed in combination with the prediction results, such as adjusting the feed ratio, feeding density or environmental conditions according to the predicted nutritional component content, so as to optimize the growth and nutritional accumulation of black soldier flies. Through the prediction model, the balance between the nutritional components and the ecological footprint of black soldier flies is explored, providing new ideas for realizing more environmentally friendly and efficient insect farming.

[0194] It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific implementation manners. Any other implementation manners obtained by those skilled in the art according to the technical solutions of the present invention also belong to the scope of protection of the present invention.

Claims

1. A prediction method for the nutrient component content of black soldier fly larvae, characterized in that: It includes the following steps: S1. Obtain the sample data of the nutritional components of Hermetia illucens larvae and perform preprocessing; S2. Based on the sample data processed in step S1, establish a machine learning model for predicting the nutritional component content of Hermetia illucens larvae and optimize the prediction model; S3. Validate and interpret the machine learning model for predicting the nutrient content of black soldier fly larvae trained in step S2 based on the random forest (RF), support vector machine (SVM), XGBoost, and GBRT models, using R 2 Evaluate the prediction effect of the prediction model, draw the feature importance and SHAP summary diagram based on the SHAP value, and output the prediction effect by inputting the characteristic parameters of the experiment.

2. The prediction method of the nutrient component content of the black soldier fly larvae according to claim 1, wherein: The specific steps of step S1 include: (1) Collect the valid sample data of the nutritional components of Hermetia illucens larvae from the database, including crude protein in the larvae, crude fat in the larvae, feeding substrate, types of exogenous additives, feeding environmental temperature, humidity, initial larval age, transformation period, light cycle, and initial larval weight; (2) Perform preprocessing on the collected valid sample data of the nutritional components of Hermetia illucens larvae.

3. The prediction method of the nutritional component content of black soldier fly larvae according to claim 2, wherein: The specific steps of step (2) of step S1 include: According to the characteristics of the data set, some data need to be processed before being applied to the model as input features; (1) Data encoding and enhancement processing The feeding substrate uses one-hot encoding to avoid the order misguidance caused by numerical labels; The types of exogenous additives use segmented encoding: 0 indicates no additive, 1 - 3 indicates the low-dose group, and 4 - 6 indicates the high-dose group to enhance the sensitivity of the model to the dose effect; (2) Data cleaning and standardization Outlier processing: Use SPSS to perform box plot analysis on continuous variables and delete samples exceeding 1.5 times the IQR; Missing value filling: Continuous variables are filled with the median, and discrete variables are filled with the mode; Standardization processing: Perform Z-score standardization on continuous features such as temperature and humidity, and the formula is: where μ is the mean and σ is the standard deviation; the light cycle is normalized to [0, 1] through Min - Max normalization.

4. The prediction method of the nutrient content of the black soldier fly larvae according to claim 1, wherein: The specific steps of step S2 include: (1) Respectively establish machine learning models for predicting the nutritional component content of Hermetia illucens larvae based on the random forest (RF), support vector machine (SVM), XGBoost, and GBRT models and perform training; (2) Respectively optimize the model parameters through grid search and introduce dynamic error weight adjustment to strengthen the learning of difficult samples; (3) Respectively perform principal component analysis PCA dimensionality reduction on the multicollinear features to generate the trained machine learning models for predicting the nutritional component content of Hermetia illucens larvae based on the random forest RF, support vector machine SVM, XGBoost, and GBRT models.

5. The prediction method of the nutrient component content of the black soldier fly insect body according to claim 1, wherein: The specific steps of step S3 include: (1) Optimize the model parameters according to the magnitudes of the model evaluation metrics. After completing data collection and processing, when selecting a machine learning model, it is necessary to adjust the model parameters and optimize the model parameters by observing the best results of each run. The CART optimization parameters of the XGBoost model include gamma, max_depth, lambda, subsample, comsample_bytree, min_child_weight, eta. The model optimization parameters of RF include n_estimators, random_state, max_depth, max_features, min_samples_leaf, min_samples_split. The parameter optimization of SVM includes gamma, kernel. The optimization parameters of the GBRT model are n_estimators, learning_rate, max_depth, random_state; (2) Evaluation indicators of the regression prediction model. Since there are significant differences in the characteristic values of the black soldier fly breeding environment, R 2 is used for model evaluation. The closer R 2 is to 1, the better the model prediction effect; where: n is the number of experimental data used in training; y pre 、y exp and are the predicted values of the nutrient content of the worm body obtained from the model, experimental data, and average experimental data, respectively; (3) Use SHAP analysis to calculate the marginal contribution of the model for global and local interpretations, and identify the key features. The specific calculations are as follows: 1) Conduct feature importance calculation and analysis: Calculate Gain, which is the average information gain brought by a certain feature after splitting in all trees. The specific calculation steps are as follows. Obtain the feature importance by sorting the average information gain and draw a feature importance graph; where: Ent(D) is the information entropy of set D; θ is the number of subsets after partitioning; is the proportion of the divided set in the original set; 2) Conduct SHAP model interpretation. By calculating the shapley value of each feature, obtain the contribution of each feature to the model regression and show the positive and negative of this contribution; increase the interpretability of the model and draw a feature summary graph for intuitive analysis of the model interpretation. The specific calculation steps are as follows; Wherein: is the contribution of the i-th feature; S is a subset of the given prediction features; A is the set of all features; f x (S ∪ {i}), f x (S) includes and does not include the model results of the i-th feature; (4) Input the feature parameters in the actual experiment according to the code prompts and display the prediction results.