A method and system for constructing a detection model for crude protein content in chicken manure
By combining visible-near-infrared spectroscopy with various preprocessing and feature selection-supervised learning models, a simple and efficient model for detecting crude protein content in chicken manure is constructed. This solves the problems of high cost, high pollution, and low efficiency in existing technologies, and achieves rapid and accurate detection results.
Patent Information
- Application Number
- CN202511439432.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing technologies for detecting crude protein content in chicken manure are costly, polluting, and inefficient, making it difficult to meet the demands of modern agriculture for rapid, efficient, and non-destructive testing.
Using visible-near-infrared spectroscopy, a simple and efficient model for detecting crude protein content in chicken manure was constructed by combining various preprocessing methods and a feature selection-supervised learning model. The optimal preprocessing and feature selection-supervised learning model were selected to extract feature bands that are highly correlated with the crude protein content in chicken manure.
It enables rapid and accurate detection of crude protein content in chicken manure, improves detection precision and model robustness, is suitable for portable on-site testing equipment, and has good prospects for widespread application.
Smart Images

Figure CN120913661B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, specifically to the field of fecal detection. Background Technology
[0002] Chicken manure is one of the main organic wastes generated during livestock and poultry farming. It is rich in organic matter and various nutrients essential for plant growth, particularly protein, nitrogen, phosphorus, and potassium. Therefore, it is widely used as an organic fertilizer in agricultural planting. Crude protein content, as an important parameter reflecting the fertilizer efficiency of chicken manure, is a crucial indicator for evaluating its nutrient content and application value. Understanding the crude protein content in chicken manure is of great significance for rational fertilization, improving fertilizer utilization, and ensuring the green development of agriculture.
[0003] Traditional methods for detecting crude protein content in chicken manure mainly rely on chemical analysis techniques such as the Kjeldahl method. While these methods offer high accuracy, they suffer from drawbacks such as cumbersome operation, long testing cycles, and environmental pollution from reagents, making them unsuitable for the practical needs of modern agriculture for rapid, efficient, and non-destructive testing. In particular, the lack of convenient and rapid methods for detecting crude protein content in agricultural production sites or during organic fertilizer processing can easily lead to inaccurate fertilizer effectiveness assessments, affecting the rational use of fertilizers and the yield and quality of crops.
[0004] Visible-near-infrared spectroscopy, as an efficient, rapid, and non-destructive detection method, has been widely used in agricultural resource detection and management in recent years. This technology enables rapid quantitative analysis of the internal components of a sample by analyzing its absorption, reflection, or transmission characteristics to various wavelengths of light. However, due to the high dimensionality and redundancy of visible-near-infrared spectroscopy data, direct modeling often faces problems of information interference and degraded model performance.
[0005] In summary, existing technologies for detecting crude protein content in chicken manure suffer from high costs, significant pollution, and low efficiency. Summary of the Invention
[0006] This invention solves the problems of high cost, significant pollution, and low efficiency in existing chicken manure crude protein content detection technologies, thereby improving the accuracy of chicken manure resource application. This invention provides the following solution:
[0007] Option 1: A method for constructing a model for detecting crude protein content in chicken manure, comprising the following steps:
[0008] Step 1: Obtain a sample set, which includes the visible-near infrared reflectance spectrum and crude protein content of chicken manure;
[0009] Step 2: The visible-near infrared reflectance spectrum is preprocessed using multiple preprocessing methods to obtain the preprocessed spectrum corresponding to each preprocessing method;
[0010] Step 3: Based on the crude protein content and the preprocessing spectra corresponding to the various preprocessing methods, establish and evaluate the partial least squares regression models corresponding to the various preprocessing methods, and obtain the evaluation results; based on the evaluation results, select the best preprocessing method and the corresponding preprocessing spectra.
[0011] Step 4: Use multiple feature selection methods to perform feature selection on the preprocessed spectrum to obtain the feature bands corresponding to each feature method;
[0012] Step 5: Based on the crude protein content and the feature bands corresponding to the various feature selection methods, construct various supervised learning models corresponding to the various feature selection methods, and select the best feature selection-supervised learning model by evaluating the various supervised learning models.
[0013] Step six: Combining the preprocessing methods selected in step three and the feature selection-supervised learning model selected in step five, the final chicken manure crude protein content detection model is obtained.
[0014] Furthermore, in one embodiment of the present invention, obtaining the visible-near-infrared reflectance spectrum of the sample set in step one includes the following steps:
[0015] Step 11: Collect chicken manure and pre-treat the collected chicken manure to obtain pre-treated chicken manure;
[0016] Step 12: Collect the visible-near infrared reflectance spectrum of the pretreated chicken manure.
[0017] Furthermore, in one embodiment of the present invention, step 11 is:
[0018] The collected chicken manure samples were divided into portions and sealed in plastic bags within each group; 10 mL of 10% H2SO4 was added to each 100 g of manure sample for nitrogen fixation treatment; after 4 days of treatment, each replicate of chicken manure samples was thoroughly mixed and dried at 65°C for 72 hours, followed by 24 hours of rehydration; the dried manure samples were then pulverized to 40 mesh to ensure homogenization and obtain pretreated chicken manure.
[0019] Furthermore, in one embodiment of the present invention, the various preprocessing methods described in step two specifically refer to SG smoothing, multivariate scattering correction MSC, standard normalization SNV, average normalization MN, baseline offset B, detrending D, first derivative FD, or second derivative SD preprocessing methods.
[0020] Furthermore, in one embodiment of the present invention, the establishment and evaluation of the partial least squares regression model corresponding to the various preprocessing methods in step three includes the following steps:
[0021] Step 31: Based on the crude protein content and the preprocessing spectrum corresponding to each preprocessing method, establish datasets respectively, and perform the following steps 32-35 on each dataset to obtain evaluation results;
[0022] Step 32: Randomly divide the dataset into a training set and a test set in a 4:1 ratio;
[0023] Step 33: Based on the training set, determine the number of principal components using leave-one-out cross-validation.
[0024] Step 34: Based on the number of principal components and the training set, train the partial least squares regression model to obtain the trained partial least squares regression model;
[0025] Step 35: Based on the test set, evaluate the partial least squares regression model using the coefficient of determination and root mean square error to obtain the evaluation results.
[0026] Furthermore, in one embodiment of the present invention, the multiple feature selection methods mentioned in step four specifically refer to the competitive adaptive reweighted sampling method CARS, the continuous projection method SPA, or the random frog jumping method RF.
[0027] Furthermore, in one embodiment of the present invention, the various supervised learning models mentioned in step five specifically refer to: Particle Swarm Optimization Support Vector Regression Machine (POS-SVR), Backpropagation Neural Network (BP-NN), or Extreme Learning Machine (ELM).
[0028] Option 2: A system for constructing a model for detecting crude protein content in chicken manure, comprising the following modules:
[0029] Module 1 is used to obtain a sample set, which includes the visible-near infrared reflectance spectrum and crude protein content of chicken manure;
[0030] Module 2 is used to preprocess the visible-near-infrared reflectance spectrum using multiple preprocessing methods to obtain the preprocessed spectrum corresponding to each preprocessing method;
[0031] Module 3 is used to establish and evaluate the partial least squares regression models corresponding to the various preprocessing methods based on the crude protein content and the preprocessing spectra corresponding to the various preprocessing methods, and obtain the evaluation results; based on the evaluation results, the optimal preprocessing method and the corresponding preprocessing spectra are selected.
[0032] Module 4 is used to perform feature selection on the preprocessed spectrum using multiple feature selection methods to obtain the feature bands corresponding to each feature method;
[0033] Module 5 is used to construct multiple supervised learning models corresponding to the multiple feature selection methods based on the crude protein content and the feature bands corresponding to the multiple feature selection methods, and to select the best feature selection-supervised learning model by evaluating the multiple supervised learning models.
[0034] Module 6 combines the preprocessing methods selected from Module 3 and the feature selection-supervised learning model selected from Module 5 to obtain the final chicken manure crude protein content detection model.
[0035] Option 3: A method for detecting the crude protein content in chicken manure, comprising the following steps:
[0036] Step B1: Collect and preprocess the chicken manure to be tested to obtain preprocessed chicken manure;
[0037] Step B2: Use a chicken manure spectral data acquisition device to acquire spectral data of the pretreated chicken manure and obtain the visible-near infrared reflectance spectrum of the pretreated chicken manure;
[0038] Step B3: Input the visible-near-infrared reflectance spectrum into the chicken manure crude protein content detection model for detection to obtain the chicken manure crude protein content;
[0039] The crude protein content detection model for chicken manure is any of the methods described above.
[0040] Option 4: An electronic device according to the present invention includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.
[0041] Memory, used to store computer programs;
[0042] When the processor executes the program stored in the memory, it implements any of the above-described methods for constructing a model for detecting crude protein content in chicken manure.
[0043] Option 5: A computer-readable storage medium according to the present invention, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any of the above-described methods for constructing a chicken manure crude protein content detection model.
[0044] This invention provides a method and system for constructing a model for detecting the crude protein content of chicken manure. It effectively solves the problems of high cost, significant pollution, and low efficiency in existing chicken manure crude protein content detection technologies, improving the accuracy of chicken manure resource application. It provides technical support for the precise application of chicken manure resources and green agricultural development, and also lays the foundation for the development of portable on-site testing equipment, demonstrating promising prospects and industrialization value.
[0045] 1. The method for constructing a crude protein content detection model in chicken manure described in this invention is used to build a concise and efficient detection model. This method evaluates partial least squares regression models established by various preprocessing methods to select the optimal preprocessing method; it combines multiple feature selection methods and multiple supervised learning models in pairs, and based on the selected preprocessing methods, evaluates the combined feature selection-supervised learning model to select the optimal one; combining the obtained preprocessing methods and feature selection-supervised learning models, the final crude protein content detection model in chicken manure is obtained. This method can extract feature bands highly correlated with the crude protein content in chicken manure from high-dimensional spectral data and construct a concise and efficient crude protein content detection model in chicken manure, achieving rapid and accurate detection of crude protein content in chicken manure.
[0046] 2. The difference between the optimal preprocessing method described in this invention and existing technologies lies in the fact that existing technologies, such as the literature "Study on Near-Infrared Spectroscopic Analysis of Main Nutrient Components in Chicken Manure Factory Compost," which, according to the abstract, establishes near-infrared models for total nitrogen, total phosphorus, and total potassium by combining multiple spectral data processing methods and partial least squares regression, lack efficient and quantitative evaluation criteria for preprocessing methods. Furthermore, factors such as particle size, shape, and component content of chicken manure significantly influence the dataset. Once a specific preprocessing method is selected as the final tool, it becomes impossible to find the most suitable preprocessing method for different datasets, thus reducing the simplicity and accuracy of the model. To address the aforementioned technical problems, this invention evaluates the partial least squares regression (PLSR) models established using various chicken manure visible-near-infrared reflectance spectroscopy preprocessing methods, selecting the optimal preprocessing method. On one hand, PLSR models have the advantage of handling high-dimensional datasets with strong multicollinearity, making them particularly suitable for modeling chicken manure spectral data. By constructing PLSR models from various preprocessed and raw spectral data respectively, the impact of each method on modeling performance can be systematically compared, thus scientifically and objectively selecting the most suitable preprocessing method for chicken manure visible-near-infrared reflectance spectroscopy. On the other hand, by objectively and quantitatively evaluating various preprocessing methods and combinations, the preprocessing effect is improved, leading to the construction of a concise and efficient detection model. This significantly enhances the detection accuracy, robustness, and practicality of the chicken manure crude protein content detection model.
[0047] 3. The optimal feature selection-supervised learning model described in this invention is used to improve the accuracy and stability of the detection model, reduce computational resource consumption, and enhance model interpretability. This invention combines multiple feature selection methods and multiple supervised learning models in pairs, and evaluates the combined feature selection-supervised learning models based on the preprocessing methods after selection. The optimal feature selection-supervised learning model is then selected to extract feature bands highly correlated with the crude protein content in chicken manure from high-dimensional spectral data, and the optimal match between the feature selection method and the supervised learning model is found, thus constructing a concise and efficient detection model.
[0048] The method described in this invention is applicable to the detection of crude protein content in chicken manure using manure resources. Furthermore, this strategy has better interpretability and portability, making it suitable for integration into low-cost multispectral equipment and showing promising prospects for widespread application. Attached Figure Description
[0049] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0050] Figure 1 This is a flowchart of the method for constructing a detection model for crude protein content in chicken manure as described in Implementation Method 1.
[0051] Figure 2 This is a schematic diagram of the chicken manure spectral data acquisition device described in Embodiment 1, wherein: 1. Computer, 2. White reference plate, 3. Sample, 4. Blackboard, 5. Reflection probe, 6. Light source, 7. Spectrometer.
[0052] Figure 3 This is a graph showing the principal component count of the partial least squares regression model corresponding to the various preprocessing methods described in Implementation Method 1.
[0053] Figure 4 This is a graph showing the evaluation results of the partial least squares regression models corresponding to the various preprocessing methods described in Implementation Method 1.
[0054] Figure 5 The above is an image showing the effect of MSC preprocessing CARS feature bands as described in Implementation Method 1. (a) is a diagram of the MSC preprocessing CARS feature band extraction process, showing, from top to bottom, the number of band variables, the RMSEcv change trajectory, and the wavelength regression coefficient trend. In the wavelength regression coefficient trend, different colored curves represent the path of regression coefficient changes for each band variable in each Monte Carlo sampling. The vertical line marked with an asterisk indicates the sampling number corresponding to when RMSEcv reaches its minimum, i.e., the 22nd sampling. (b) is a point map of CARS feature band extraction.
[0055] Figure 6The above are the effect diagrams of the SNV preprocessing CARS feature bands described in Implementation Method 1. Among them, (a) is a diagram of the SNV preprocessing CARS feature band extraction process, which, from top to bottom, shows the change in the number of band variables, the trajectory of RMSEcv change, and the trend of wavelength regression coefficients. In the trend of wavelength regression coefficients, the curves of different colors represent the change path of the regression coefficients of each band variable in each Monte Carlo sampling. The vertical line marked with an asterisk represents the sampling number corresponding to when RMSEcv reaches its minimum, i.e., the 24th sampling. (b) is a point diagram of the feature bands extracted by SNV preprocessing CARS.
[0056] Figure 7 The diagram shows the effect of MSC preprocessing SPA feature bands as described in Implementation Method 1, where (a) is a diagram of the MSC preprocessing SPA feature band extraction process, and (b) is a diagram of the SPA extracted feature bands.
[0057] Figure 8 The above are the effect diagrams of SNV preprocessing SPA feature bands as described in Implementation Method 1, wherein (a) is a diagram of the SNV preprocessing SPA feature band extraction process, and (b) is a diagram of the SNV preprocessing SPA feature band extraction point.
[0058] Figure 9 The diagram shows the effect of RF feature bands in MSC preprocessing as described in Implementation Method 1. (a) is a diagram of the RF feature band extraction process in MSC preprocessing, with the red dashed line indicating a selection threshold of 0.2. (b) is a point diagram of RF feature band extraction.
[0059] Figure 10 The above are the effect diagrams of SNV preprocessing CARS feature bands as described in Implementation Method 1. Among them, (a) is a diagram of the SNV preprocessing RF feature band extraction process, and the red dashed line indicates that the selection threshold is 0.2. (b) is a point diagram of SNV preprocessing RF feature band extraction.
[0060] Figure 11 This is an internal parameter diagram of the POS-SVR model described in Implementation Method 1.
[0061] Figure 12 This is a graph showing the evaluation results of the POS-SVR model described in Implementation Method 1.
[0062] Figure 13 This is a graph showing the evaluation results of the BP-NN model described in Implementation Method 1.
[0063] Figure 14 This is the internal parameter diagram of the ELM model described in Implementation Method 1.
[0064] Figure 15 This is a graph showing the evaluation results of the ELM model described in Implementation Method 1. Detailed Implementation
[0065] Various embodiments of the present invention will now be clearly and completely described with reference to the accompanying drawings. The embodiments described with reference to the drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0066] Implementation Method 1: The method for constructing a detection model for crude protein content in chicken manure described in this implementation method, such as... Figure 1 As shown, it includes the following steps:
[0067] Step 1: Obtain a sample set, which includes the visible-near infrared reflectance spectrum and crude protein content of chicken manure;
[0068] Step 2: The visible-near infrared reflectance spectrum is preprocessed using multiple preprocessing methods to obtain the preprocessed spectrum corresponding to each preprocessing method;
[0069] Step 3: Based on the crude protein content and the preprocessing spectra corresponding to the various preprocessing methods, establish and evaluate the partial least squares regression models corresponding to the various preprocessing methods, and obtain the evaluation results; based on the evaluation results, select the best preprocessing method and the corresponding preprocessing spectra.
[0070] Step 4: Use multiple feature selection methods to perform feature selection on the preprocessed spectrum to obtain the feature bands corresponding to each feature method;
[0071] Step 5: Based on the crude protein content and the feature bands corresponding to the various feature selection methods, construct various supervised learning models corresponding to the various feature selection methods, and select the best feature selection-supervised learning model by evaluating the various supervised learning models.
[0072] Step six: Combining the preprocessing methods selected in step three and the feature selection-supervised learning model selected in step five, the final chicken manure crude protein content detection model is obtained.
[0073] In this embodiment, the sample set obtained in step one is as follows:
[0074] Step 11: Collect chicken manure and pre-treat the collected chicken manure to obtain pre-treated chicken manure;
[0075] Step 12: Collect the visible-near infrared reflectance spectrum of the pretreated chicken manure.
[0076] In step 11, the collected chicken manure is pretreated, preferably by dividing the collected chicken manure samples into portions and sealing them in plastic bags within each group. 10 mL of 10% H₂SO₄ is added to each 100 g of manure sample for nitrogen fixation. After 4 days of treatment, each replicate chicken manure sample is thoroughly mixed and dried at 65°C for 72 hours, followed by rehydration for 24 hours. The dried manure sample is then pulverized to 40 mesh to ensure homogenization, yielding pretreated chicken manure.
[0077] The process of collecting the visible-near-infrared reflectance spectrum of the pretreated chicken manure in step 12 is as follows:
[0078] When acquiring visible-near-infrared reflectance spectra, the following methods are used: Figure 2 The apparatus shown is implemented by first connecting the optical fiber of the reflective probe 5 to the spectrometer 7 and the light source 6, and then fixing the reflective probe 5 with a reflective probe holder to ensure that the angle between the probe 5 and the surface of the sample 3 is 90°. After the light source is preheated for 8 minutes, whiteboard calibration is performed using the blackboard 4 and the white reference board 2, and the measurement is then performed. The measurement results are displayed on the computer 1. During the experiment, whiteboard calibration is performed every 30 minutes. 50g of dried chicken manure from each group is taken as a sample and placed in a 90mm diameter petri dish. The corresponding visible-near-infrared reflectance spectrum is collected based on each sample. This visible-near-infrared reflectance spectrum and the corresponding crude protein content of the sample are used as sample data in the sample set.
[0079] In this embodiment, spectral data acquisition can be achieved using existing spectral data acquisition software tools, such as using AvaSoft 8 software (Aventes, Netherlands) in conjunction with a spectrometer to acquire spectral data and export the acquired spectral data. Preprocessing is preferably performed using The Unscrambler X 10.4 software (CAMO, Norway). The detection model for crude protein content in chicken manure is preferably constructed using Matlab 2023b software (MathWorks, USA). Plotting is preferably performed using Origin2021 software (OriginLab, USA).
[0080] In this embodiment, the crude protein content of chicken manure in the sample set can be obtained by detecting each sample using existing technology, for example, the Kjeldahl nitrogen determination method can be used, based on:
[0081]
[0082] Obtain crude protein content , in %, of which , The volume of standard hydrochloric acid solution consumed for titrating the blank, in mL; The volume of standard hydrochloric acid solution consumed for titrating the sample, in mL; The concentration of the hydrochloric acid standard titration solution is expressed in mol / L. The mass of the sample is expressed in grams. This represents the total volume of the sample decomposition solution, in mL. 14 represents the volume of sample decomposition solution used for distillation, in mL; 14 represents the molar mass of nitrogen, in g / mol; 6.25 represents the average coefficient for converting nitrogen to crude protein.
[0083] In this embodiment, the specific process of establishing and evaluating the partial least squares regression models corresponding to the various preprocessing methods in step three is as follows:
[0084] Step 31: Based on the crude protein content and the preprocessing spectrum corresponding to each preprocessing method, establish datasets respectively, and perform the following steps 32-35 on each dataset to obtain evaluation results;
[0085] Step 32: Randomly divide the dataset into a training set and a test set in a 4:1 ratio;
[0086] Step 33: Based on the training set, determine the number of principal components using leave-one-out cross-validation.
[0087] Step 34: Based on the number of principal components and the training set, train the partial least squares regression model to obtain the trained partial least squares regression model;
[0088] Step 35: Based on the test set, evaluate the partial least squares regression model using the coefficient of determination and root mean square error to obtain the evaluation results;
[0089] The selection of the number of principal components in step 33 has a significant impact on model performance. Too few principal components may prevent the model from fully extracting information related to the target variable from the spectral data; while too many principal components may introduce noise and redundant information, increasing model complexity and thus reducing the model's robustness and generalization ability.
[0090] In this embodiment, step three involves selecting the optimal preprocessing method, typically resulting in one optimal method. However, in practical applications, multiple optimal preprocessing methods may be selected. In such cases, steps four and five are executed for each method to obtain the optimal feature selection-supervised learning model among these methods. By evaluating the optimal feature selection-supervised learning model across these methods, the optimal preprocessing-feature selection-supervised learning model is selected as the model for detecting crude protein content in chicken manure.
[0091] The method for constructing a detection model for crude protein content in chicken manure described in this embodiment is a rapid detection model construction method based on visible-near-infrared spectroscopy. This method obtains the optimal preprocessing-feature selection-supervised learning model by screening preprocessing methods, feature wavelength extraction, and modeling algorithms, thereby achieving efficient and accurate detection of crude protein content in chicken manure samples. This provides technical support for the precise application of chicken manure resources and green agricultural development, and also provides the basic conditions for the development of portable on-site detection equipment, with good prospects for promotion and industrialization value.
[0092] This implementation method provides an example, with the following specific steps:
[0093] Step S1: Obtain a sample set, which includes the visible-near infrared reflectance spectrum and crude protein content of chicken manure;
[0094] The sample set is obtained in the following way:
[0095] Step S11: Pre-treat the collected chicken manure to obtain pre-treated chicken manure;
[0096] Specifically, the collected chicken manure samples were aliquoted into individual plastic bags within each group. 10 mL of 10% H₂SO₄ was added to each 100 g of manure sample for nitrogen fixation treatment. After 4 days of treatment, each replicate of chicken manure samples was thoroughly mixed and dried at 65°C for 72 hours, followed by rehydration for 24 hours. The dried manure samples were then pulverized to 40 mesh to ensure homogenization, thus obtaining the chicken manure samples.
[0097] Step S12: Collect the visible-near infrared reflectance spectrum of the pretreated chicken manure;
[0098] Specifically, when acquiring visible-near-infrared reflectance spectra, the following methods are used: Figure 2 The apparatus shown is implemented by first connecting the optical fiber of the reflective probe 5 to the spectrometer 7 and the light source 6, and then fixing the reflective probe 5 with a reflective probe bracket to ensure that the angle between the probe 5 and the surface of the sample 3 is 90°. After the light source is preheated for 8 minutes, whiteboard calibration is performed using the blackboard 4 and the white reference board 2, and the measurement is displayed on the computer 1. During the experiment, whiteboard calibration is performed every 30 minutes. 50g of dried chicken manure from each group is taken as a sample and placed in a 90mm diameter petri dish. Spectral data corresponding to each sample are collected, and this spectral data and the corresponding crude protein content of the sample are used as sample data in the sample set. Eight spectral data points are collected for each sample, resulting in a total of 120 spectral data points. The spectral information has 873 data points, forming a high-dimensional matrix of 873×120.
[0099] The crude protein content of chicken manure was determined using the Kjeldahl method, based on:
[0100]
[0101] Obtain crude protein content , in %, of which , The volume of standard hydrochloric acid solution consumed for titrating the blank, in mL; The volume of standard hydrochloric acid solution consumed for titrating the sample, in mL; The concentration of the hydrochloric acid standard titration solution is expressed in mol / L. The mass of the sample is expressed in grams. This represents the total volume of the sample decomposition solution, in mL. 14 represents the volume of sample decomposition solution used for distillation, in mL; 14 represents the molar mass of nitrogen, in g / mol; 6.25 represents the average coefficient for converting nitrogen to crude protein.
[0102] Step S2: The visible-near infrared reflectance spectrum is preprocessed using multiple preprocessing methods to obtain the preprocessed spectrum corresponding to each preprocessing method;
[0103] The various preprocessing methods include SG smoothing, multivariate scattering correction (MSC), standard normalization (SNV), mean normalization (MN), baseline shift (B), detrending (D), first derivative (FD), or second derivative (SD) preprocessing methods.
[0104] Step S3: Based on the crude protein content and the preprocessing spectra corresponding to the various preprocessing methods, establish and evaluate the partial least squares regression models corresponding to the various preprocessing methods, and obtain the evaluation results; based on the evaluation results, select the best preprocessing method and the corresponding preprocessing spectra.
[0105] Among them, leave-one-out cross-validation is used, and the optimal number of principal components is determined based on minimizing RMSEcv. The number of principal components is as follows: Figure 3 The figure shows the principal component counts of the partial least squares regression models corresponding to various preprocessing methods. The partial least squares regression models corresponding to various preprocessing methods are evaluated using the coefficient of determination and root mean square error. The evaluation results are shown below. Figure 4 As shown, the model established by the multivariate scattering correction and standard normalization preprocessing methods exhibits the highest accuracy. Specifically, in the test set, the coefficient of determination for the multivariate scattering correction preprocessing method is 0.7984, and the root mean square error is 0.4473; while the coefficient of determination for the standard normalization preprocessing method is 0.7993, and the root mean square error is 0.4466. Therefore, further analysis of the spectral data after multivariate scattering correction and standard normalization preprocessing is required.
[0106] Step S4: Use multiple feature selection methods to perform feature selection on the preprocessed spectrum to obtain the feature bands corresponding to each feature method;
[0107] Among them, the various feature selection methods are the competitive adaptive reweighted sampling method CARS, the continuous projection method SPA, or the random frog jumping method RF.
[0108] This embodiment uses 50 Monte Carlo sampling iterations and employs 10-fold cross-validation to select the final variables. The smaller the root mean square error (RMSECV) of the cross-validation, the better the subset of feature bands associated with the sample.
[0109] A competitive adaptive reweighted sampling method was used to analyze the spectral data after multivariate scattering correction and standard normalization preprocessing. Specifically:
[0110] In this embodiment, the CARS algorithm was run 50 times. The RMSECV reached its minimum value of 0.1653 on the 22nd sampling run. Using the CARS algorithm, 74 characteristic bands were extracted from the preprocessed MSC spectrum for crude protein analysis, representing 8.48% of the total bands. Figure 5 As shown, Figure 5 The following are the results of MSC preprocessing CARS feature band extraction: (a) shows the process of MSC preprocessing CARS feature band extraction. The top graph shows the change in the number of band variables. As the number of iterations increases, the number of variables screened decreases, reflecting that a large amount of useless spectral information is removed, greatly improving the efficiency of variable screening. The middle graph shows the RMSEcv change trajectory, reflecting that as the number of samplings increases, the RMSECV value gradually decreases, indicating that a large amount of irrelevant or noise information in the spectrum is removed. When the number of samplings reaches 22, the RMSECV value increases, indicating that some important variables related to the detection of crude protein content in chicken manure are removed from the spectrum. The bottom graph shows the trend of wavelength regression coefficients, representing the change of regression coefficients of all variables in the middle graph in each sampling. Combining the analysis with the middle graph, it can be found that the PLS model established by the subset of variables obtained in the 22nd sampling has the smallest RMSECV value. (b) is a point graph of feature bands extracted by MSC preprocessing CARS. This subset contains 74 variables.
[0111] In this embodiment, the RMSECV reached its minimum value of 0.2203 when the sampling was run for the 24th time. Using the CARS algorithm, 57 characteristic band points were extracted from the SNV preprocessed spectrum for crude protein analysis, accounting for 6.53% of the total band. Figure 6As shown, the SNV preprocessing CARS feature band effect diagram is shown. (a) is the SNV preprocessing CARS feature band extraction process diagram. The top diagram shows the change in the number of band variables. As the number of iterations increases, the number of variables screened decreases, reflecting that a large amount of useless spectral information is removed, which greatly improves the efficiency of variable screening. The middle diagram shows the RMSEcv change trajectory, reflecting that as the number of samplings increases, the RMSECV value gradually decreases, indicating that a large amount of irrelevant or noise information in the spectrum is removed. When the number of samplings reaches 24, the RMSECV value increases, indicating that some important variables related to the detection of crude protein content in chicken manure are removed from the spectrum. The bottom diagram shows the wavelength regression coefficient trend, which shows the change of regression coefficients of all variables in the middle diagram in each sampling. Combining the middle diagram with the analysis, it can be found that the PLS model established by the variable subset obtained in the 24th sampling has the smallest RMSECV value. (b) is the SNV preprocessing CARS feature band point diagram, which contains 57 variables.
[0112] The spectral data after multivariate scattering correction and standard normalization preprocessing were analyzed using the successive projection method, specifically as follows:
[0113] When preprocessing is MSC, seven bands were selected as the number of crude protein band variables in chicken manure. The corresponding bands were: 466.0513, 585.6104, 819.8179, 828.8155, 841.169, 845.0952, and 860.2189 nm. The selected characteristic bands accounted for 0.802% of the total bands. Figure 7 The image shows the effect of MSC preprocessing SPA feature bands, where (a) is a diagram of the MSC preprocessing SPA feature band extraction process. Figure 7 (a) shows the change of RMSECV with the number of band points. According to the RMSE combined with the F test (α=0.25), the optimal number of characteristic bands is determined to be 7, indicating that a large amount of collinear information is removed; (b) is the feature band point map extracted by SPA, which shows the positions of the 7 selected bands on the full spectrum.
[0114] When preprocessing is performed using SNV (Special Navier-Video) mapping, seven bands are selected as the number of crude protein band variables in chicken manure. The corresponding bands are: 512.8915, 719.5516, 819.8179, 824.3181, 828.8155, 852.9412, and 860.2189 nm. The selected characteristic bands account for 0.802% of the total band. Figure 8 The image shows the effect of SNV preprocessing SPA feature bands, where (a) is a diagram of the SNV preprocessing SPA feature band extraction process. Figure 8(a) shows the change of RMSECV with the number of band points. According to the RMSE combined with the F test (α=0.25), the optimal number of characteristic bands is determined to be 7, indicating that a large amount of collinear information is removed; (b) is the feature band point map extracted by SPA preprocessing of SNV, which shows the positions of the 7 selected bands on the full spectrum.
[0115] The random frog-jump method was used to analyze the spectral data after multivariate scattering correction and standard normalization preprocessing, specifically as follows:
[0116] After wavelength selection using the RF algorithm on the preprocessed full-spectrum information from the MSC, the probability values of each variable being selected are as follows: Figure 7 As shown, a total of 50 corresponding band points were selected, accounting for 5.73% of the entire band. Figure 9 The image shows the RF feature band extraction results of MSC preprocessing, demonstrating the use of the RF method to screen key variables in the MSC preprocessed spectrum. The selection threshold for the RF method was set to 0.2. (a) shows the RF feature band extraction process during MSC preprocessing. Figure 9 (a) shows the probability distribution of selected bands. The bands above the red line are the selected bands, and a total of 50 bands were selected. (b) is the feature band point map extracted by RF, which shows the positions of the 50 selected bands on the full spectrum.
[0117] After the full-spectral information preprocessed by SNV was subjected to wavelength selection using the RF algorithm, a total of 46 corresponding band points were selected, accounting for 5.27% of the entire band. For example... Figure 10 The image shows the effect of SNV preprocessing on CARS characteristic bands. The RF method was used to screen key variables in the SNV preprocessed spectrum. The selection threshold for the RF method was set to 0.2. Figure 10 This is a diagram illustrating the RF feature band extraction process during SNV preprocessing. Figure 10 (a) shows the probability distribution of selected bands. The bands above the red line are the selected bands, and a total of 46 bands were selected. (b) is the feature band point map of SNV preprocessing RF extraction, which shows the positions of the 46 selected bands on the full spectrum.
[0118] Step S5: Based on the crude protein content and the feature bands corresponding to the various feature selection methods, construct multiple supervised learning models corresponding to the various feature selection methods. Evaluate these supervised learning models to select the optimal feature selection-supervised learning model. The multiple supervised learning models include the POS-SVR model, the BP-NN model, and the ELM model, specifically:
[0119] Spectral data preprocessed by MSC and SNV, and spectral data after CARS, SPA, and RF feature filtering, were used as input variables. Crude protein content in chicken manure was used as the output variable. These were combined with Particle Swarm Optimization Support Vector Regression Machine (POS-SVR), Backpropagation Neural Network (BP-NN), and Extreme Learning Machine (ELM) respectively to establish supervised learning models for crude protein in chicken manure. The internal parameter diagram of the established POS-SVR is shown below. Figure 11 As shown, the structure of the established ELM model is as follows: Figure 14 As shown in the figure. Furthermore, the chicken manure crude protein supervised learning model was evaluated using the coefficient of determination and root mean square error. The evaluation results of the POS-SVR model are shown in the figure. Figure 12 As shown, the evaluation results of the BP-NN neural network are as follows: Figure 13 As shown, the evaluation results of the established ELM model are as follows: Figure 15 As shown, overall, under both MSC and SNV preprocessing, the chicken manure crude protein content detection model established by combining CARS feature screening with POS-SVR has the best effect.
[0120] Step S6: Combining the preprocessing method selected in Step 4 and the feature selection-supervised learning model selected in Step 6, a model for detecting crude protein content in chicken manure is obtained.
[0121] Implementation Method 2: This implementation method further defines the method for constructing a detection model for crude protein content in chicken manure described in Implementation Method 1. In this implementation method, the various preprocessing methods described in step 2 are specifically: SG smoothing, multivariate scattering correction MSC, standard normalization SNV, mean normalization MN, baseline shift B, detrending D, first derivative FD, or second derivative SD preprocessing methods.
[0122] In this embodiment, it is preferred to use any one or more of the preprocessing methods in combination.
[0123] This embodiment further defines step two, illustrating the various preprocessing methods described in step two. Because chicken manure samples have high moisture content, complex composition, and unstable morphology, and the original spectral data is also affected by stray light, instrument noise, sample background, baseline drift, and other factors, all of these factors can impact the efficiency of the detection model for chicken manure spectral data. Therefore, preprocessing of the original spectral data of chicken manure is necessary. The various preprocessing methods described in this embodiment, such as SG smoothing, can effectively reduce noise and interference signals in the data; Standard Normalization (SNV) processing removes the scattering effects of optical path length differences and particle inhomogeneity by independently centering and standardizing each spectral sample; Multivariate Scattering Correction (MSC) processing reduces the impact of scattering on the spectral curve; MN processing helps achieve consistency in spectral intensity; First Derivative (FD) and Second Derivative (SD) processing techniques can solve the problem of overlapping peaks in the spectral curve and improve the discrimination between spectra; Baseline Shift (B) processing increases the signal resolution, making features and peaks more obvious; Detrending (D) processing removes trend signals, thereby improving the accuracy of data analysis and enhancing signal characteristics.
[0124] Implementation Method 3: This implementation method further defines the method for constructing a detection model for crude protein content in chicken manure described in Implementation Method 1. In this implementation method, the multiple feature selection methods mentioned in step 4 are specifically: competitive adaptive reweighted sampling method CARS, continuous projection method SPA, or random frog jumping method RF.
[0125] This implementation further defines step four, providing an example of the feature selection method used in step four. The aforementioned method, Competitive Adaptive Reweighted Sampling (CARS), is an efficient feature band selection method with outstanding global search capabilities. This method combines Monte Carlo sampling and partial least squares regression, using the absolute values of regression coefficients as weights. Through adaptive reweighted sampling, it continuously filters, prioritizing the retention of bands that contribute most to the model while eliminating redundant and irrelevant information. CARS uses PLS to perform Monte Carlo cross-validation modeling on multiple wavelength subsets, and selects the subset of variables with the smallest root mean square error (RMSE) as the optimal feature. CARS exhibits excellent stability and robustness, significantly improving model detection accuracy while effectively reducing data dimensionality and computational complexity.
[0126] The Continuous Projection Method (SPA) is a stepwise forward selection variable selection method. Its basic idea is to start with an initial wavelength and, in each iteration, calculate the projection of that wavelength onto the unselected wavelengths, adding the wavelength with the shortest projection vector length to the variable combination. In this way, each newly selected wavelength maintains the lowest possible linear correlation with the already selected wavelengths. Through this process, the SPA method can extract the few most critical feature variables from a large number of original variables, significantly reducing the number of variables required for modeling and simplifying the model structure. Furthermore, the SPA method can effectively eliminate redundant variables that may introduce noise or interfere with the model, thereby improving the model's detection accuracy, stability, and interpretability.
[0127] Random frog leaping (RF) is a feature selection method based on randomization and Monte Carlo sampling. Its core idea is to calculate the probability of a variable being selected through iterative sampling, and then use this probability to measure its importance; a higher probability indicates a stronger correlation between the variable and the target. This method combines global search and local optimization capabilities, effectively avoiding getting trapped in local optima and selecting the most representative feature bands from high-dimensional spectral data. Compared to traditional stepwise regression or exhaustive methods, RF, relying on random sampling and probability update mechanisms, significantly reduces computational complexity and improves the efficiency of band selection.
[0128] Implementation Method 4: This implementation method further defines the method for constructing a chicken manure crude protein content detection model as described in Implementation Method 1. In this implementation method, the various supervised learning models mentioned in step 5 are specifically: Particle Swarm Optimization Support Vector Regression Machine (POS-SVR), Backpropagation Neural Network (BP-NN), or Extreme Learning Machine (ELM).
[0129] This implementation further defines step five, providing examples of various supervised learning models used in step five. In this method, the Support Vector Regression (SVR) model can obtain the optimal solution using existing information under limited sample conditions, effectively avoiding local extrema and exhibiting good generalization ability. However, the model parameters of SVR are difficult to select, and the quality of the parameters significantly affects the detection accuracy. The Particle Swarm Optimization (PSO) algorithm requires no complex parameter tuning and has advantages such as simple implementation, fast convergence speed, and fewer parameter settings. Therefore, PSO is used to optimize SVR, constructing a PSO-SVR model.
[0130] The BP neural network model BP-NN has powerful nonlinear modeling capabilities and adaptive learning capabilities, which can effectively capture complex input-output relationships and exhibit good generalization performance.
[0131] Extreme Learning Machine (ELM) is a fast learning method based on an improvement of classical neural networks. This method uses randomized input layer weights and biases during the training phase, enabling it to generalize well at extremely high speeds. It features fewer parameters to choose from, good learning performance, and wide applicability.
Claims
1. A method for constructing a model for detecting crude protein content in chicken manure, characterized in that, Includes the following steps: Step 1: Obtain a sample set, which includes the visible-near infrared reflectance spectrum and crude protein content of chicken manure; Step 2: The visible-near infrared reflectance spectrum is preprocessed using multiple preprocessing methods to obtain the preprocessed spectrum corresponding to each preprocessing method; The various preprocessing methods mentioned above specifically refer to SG smoothing, multivariate scattering correction MSC, standard normalization SNV, average normalization MN, baseline shift B, detrending D, first derivative FD, or second derivative SD preprocessing methods. Step 3: Based on the crude protein content and the preprocessing spectra corresponding to the various preprocessing methods, establish and evaluate the partial least squares regression models corresponding to the various preprocessing methods, and obtain the evaluation results; Based on the evaluation results, the optimal preprocessing method and the corresponding preprocessing spectrum were selected. Step 4: Use multiple feature selection methods to perform feature selection on the preprocessed spectrum to obtain the feature bands corresponding to each feature method; The various feature selection methods mentioned specifically refer to the competitive adaptive reweighted sampling method CARS, the continuous projection method SPA, or the random frog jumping method RF. Step 5: Based on the crude protein content and the feature bands corresponding to the various feature selection methods, construct various supervised learning models corresponding to the various feature selection methods, and select the best feature selection-supervised learning model by evaluating the various supervised learning models. The various supervised learning models mentioned specifically refer to: Particle Swarm Optimization Support Vector Regression (POS-SVR) model, Backpropagation Neural Network (BP-NN) model, or Extreme Learning Machine (ELM) model; Step six: Combining the preprocessing methods selected in step three and the feature selection-supervised learning model selected in step five, the final chicken manure crude protein content detection model is obtained.
2. The method for constructing a detection model for crude protein content in chicken manure according to claim 1, characterized in that, The process of obtaining the visible-near-infrared reflectance spectrum of the sample set as described in step one includes the following steps: Step 11: Collect chicken manure and pre-treat the collected chicken manure to obtain pre-treated chicken manure; Step 12: Collect the visible-near infrared reflectance spectrum of the pretreated chicken manure.
3. The method for constructing a detection model for crude protein content in chicken manure according to claim 2, characterized in that, Step 11 is: The collected chicken manure samples were divided into portions and sealed in plastic bags within each group; 10 mL of 10% H2SO4 was added to each 100 g of manure sample for nitrogen fixation treatment; after 4 days of treatment, each replicate of chicken manure samples was thoroughly mixed and dried at 65°C for 72 hours, followed by 24 hours of rehydration; the dried manure samples were then pulverized to 40 mesh to ensure homogenization and obtain pretreated chicken manure.
4. The method for constructing a detection model for crude protein content in chicken manure according to claim 1, characterized in that, Step 3, which involves establishing and evaluating the partial least squares regression models corresponding to the various preprocessing methods, includes the following steps: Step 31: Based on the crude protein content and the preprocessing spectrum corresponding to each preprocessing method, establish datasets respectively, and perform the following steps 32-35 on each dataset to obtain evaluation results; Step 32: Randomly divide the dataset into a training set and a test set in a 4:1 ratio; Step 33: Based on the training set, determine the number of principal components using leave-one-out cross-validation. Step 34: Based on the number of principal components and the training set, train the partial least squares regression model to obtain the trained partial least squares regression model; Step 35: Based on the test set, evaluate the partial least squares regression model using the coefficient of determination and root mean square error to obtain the evaluation results.
5. A system for constructing a model for detecting crude protein content in chicken manure, characterized in that, Includes the following modules: Module 1 is used to obtain a sample set, which includes the visible-near infrared reflectance spectrum and crude protein content of chicken manure; Module 2 is used to preprocess the visible-near-infrared reflectance spectrum using multiple preprocessing methods to obtain the preprocessed spectrum corresponding to each preprocessing method; The various preprocessing methods mentioned above specifically refer to SG smoothing, multivariate scattering correction MSC, standard normalization SNV, average normalization MN, baseline shift B, detrending D, first derivative FD, or second derivative SD preprocessing methods. Module 3 is used to establish and evaluate the partial least squares regression models corresponding to the various preprocessing methods based on the crude protein content and the preprocessing spectra corresponding to the various preprocessing methods, and obtain the evaluation results; based on the evaluation results, the optimal preprocessing method and the corresponding preprocessing spectra are selected. Module 4 is used to perform feature selection on the preprocessed spectrum using multiple feature selection methods to obtain the feature bands corresponding to each feature method; The various feature selection methods mentioned specifically refer to the competitive adaptive reweighted sampling method CARS, the continuous projection method SPA, or the random frog jumping method RF. Module 5 is used to construct multiple supervised learning models corresponding to the multiple feature selection methods based on the crude protein content and the feature bands corresponding to the multiple feature selection methods, and to select the best feature selection-supervised learning model by evaluating the multiple supervised learning models. The various supervised learning models mentioned specifically refer to: Particle Swarm Optimization Support Vector Regression (POS-SVR) model, Backpropagation Neural Network (BP-NN) model, or Extreme Learning Machine (ELM) model; Module 6 combines the preprocessing methods selected from Module 3 and the feature selection-supervised learning model selected from Module 5 to obtain the final chicken manure crude protein content detection model.
6. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes the program stored in the memory, it implements the method for constructing the crude protein content detection model of chicken manure as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method for constructing a detection model for crude protein content in chicken manure according to any one of claims 1-4.
Citation Information
Patent Citations
Method for detecting nutrient content in excrement and application thereof
CN116660205A
Near infrared spectrum rapid evaluation method for brown rice protein content
CN119598298A