Method for detecting INDF content in TMR of lactating cattle based on near infrared spectrum
By constructing a model based on near-infrared spectroscopy and a handheld device, the problem of rapidly and accurately detecting iNDF content in TMR of lactating cows on farms has been solved, achieving simplified operation and efficient detection, and adapting to the local environment in China.
Patent Information
- Application Number
- CN202511705235.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies cannot quickly and accurately detect iNDF content in TMR of lactating cows on-site, and existing models lack local applicability, resulting in insufficient convenience and effectiveness of detection.
A model based on near-infrared spectroscopy was constructed, and partial least squares regression algorithm and cross-validation optimization were adopted. Combined with a handheld near-infrared device, iNDF content was detected directly on the ranch by simplifying sample preprocessing and spectral data processing.
It enables rapid and accurate on-site detection of iNDF content in dairy farms, reduces operational complexity and skill requirements, improves the convenience and accuracy of detection, and is adapted to the compositional characteristics of TMR in Chinese dairy cows.
Smart Images

Figure CN121565291A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of animal feed testing technology, and more specifically, to a method for detecting INDF content in TMR of lactating cows based on near-infrared spectroscopy. Background Technology
[0002] Total Mixed Ration (TMR) is a nutritionally balanced diet prepared by thoroughly mixing roughage, concentrate, and various additives in a TMR mixer to meet the nutritional needs of dairy cows at different growth and development stages and lactation stages. Indigestible neutral detergent fiber (iNDF), as a core indicator of feed or plant fiber digestibility, is crucial for assessing energy intake and feed consumption in ruminants and predicting organic matter digestibility. Accurately detecting the iNDF content in lactating cow TMR is a key factor in ensuring the profitability of dairy farming. Currently, most near-infrared (NIIR) prediction models used for detecting iNDF content in TMR of lactating cows are based on benchtop NIIR equipment used in laboratories. The application of these models is subject to strict limitations. Not only do they require a standardized laboratory environment and operation of the benchtop NIIR equipment by professional personnel, but the samples also need to undergo a lengthy pretreatment process at the farm before being mailed to a specialized laboratory for analysis. Furthermore, benchtop NIIR equipment is generally expensive and bulky. Combined with the cumbersome sample pretreatment, mailing, and laboratory testing process, this significantly restricts the convenience and effectiveness of these models, severely limiting their promotion and application on farms and making it difficult to meet the actual needs of farms for timely detection of iNDF content in TMR of lactating cows. More importantly, currently available iNDF detection models for TMR in lactating cows rarely utilize samples from Chinese farms. This results in significant shortcomings in adapting these models to the compositional characteristics and feeding environments of TMR in Chinese lactating cows, further reducing their accuracy and applicability in Chinese farms. In summary, the market currently lacks a handheld near-infrared model that can be used directly by farm workers on-site for rapid and accurate detection of iNDF content in TMR in lactating cows. This fails to effectively address the issues of convenience, effectiveness, and local applicability of existing technologies, urgently requiring a new technological solution to fill this gap. Summary of the Invention
[0003] In view of this, this application provides a method for detecting INDF content in TMR of lactating cows based on near-infrared spectroscopy.
[0004] A method for detecting INDF content in TMR of lactating cows based on near-infrared spectroscopy includes: A near-infrared spectral model for predicting INDF content is constructed. The sample set for constructing the near-infrared spectral model includes near-infrared spectral data and actual INDF values corresponding to TMR samples from lactating cows. The construction method includes extracting wavelength points with non-zero coefficients from the near-infrared spectral data and using these wavelength points as feature variables. A partial least squares regression algorithm is used to construct the near-infrared spectral model based on the feature variables and the corresponding actual INDF values. The model parameters are optimized using cross-validation, and the optimal combination of model parameters is determined by combining the root mean square error of cross-validation (RMSECV) and the root mean square error of external validation (RMSEP), resulting in the near-infrared spectral model for predicting INDF content. The near-infrared spectral data of the TMR sample from the lactating cow to be tested are input into the near-infrared spectral model to obtain the INDF content detection results.
[0005] One possible implementation method for acquiring near-infrared spectral data corresponding to lactating bovine TMR samples includes: After drying, the TMR sample from lactating cows was pulverized to a particle size of 1.0 mm. The pulverized sample was then re-loaded at least twice to obtain the near-infrared spectral data. The two re-loading operations included first mixing the pulverized sample evenly and pouring it into a scanning box for a first spectral scan; after the first spectral scan, the sample was removed from the scanning box and mixed evenly a second time, then poured back into the scanning box for a second spectral scan; the average of the spectral data obtained from the first and second spectral scans was taken as the near-infrared spectral data corresponding to the lactating cow TMR sample.
[0006] One possible implementation method for obtaining the actual INDF value of the TMR sample from lactating cows includes: Take 2 grams of TMR sample from lactating cows with known NDF content a, place it in a filter bag made of polyester material with a pore size of 12 micrometers, an effective surface area of 100-200 square centimeters and an open porosity of 7%, and incubate the filter bag in the rumen for 288 hours. After incubation, the filter bag is removed, cleaned, and dried. The sample inside the bag is taken out and its NDF content is measured and recorded as b. The actual NDF value corresponding to the lactating cow TMR sample is calculated according to the formula INDF=b / a.
[0007] In one possible implementation, after obtaining the near-infrared spectrum corresponding to the TMR sample of a lactating cow, the method further includes: The near-infrared spectral data is preprocessed, including performing normal variable transformation (SNV), multivariate scattering correction (MSC), detrending, smoothing with a Savitzky-Golay smoothing filter, and differentiation.
[0008] In one possible implementation, after constructing a sample set by obtaining the near-infrared spectral data and actual INDF values corresponding to the TMR samples of lactating cows, the method further includes: The nearest neighbor pairing sample selection algorithm SPXY is used to divide the sample set into a modeling set and a validation set based on spectral distance and INDF content distance; The spectral distance is expressed as:
[0009] In the formula, p,q ∈ [1,N], where N is the total number of samples. x p , x q The first p, q Spectral data of each sample; The INDF content distance is represented as follows:
[0010] In the formula, y p , y q The first p, q The actual INDF value of each sample; The overall distance between samples is obtained by summing the spectral distance and the INDF content distance after standardization, and the sample set is then partitioned based on this overall distance; the overall distance between samples is expressed as: .
[0011] One possible implementation involves partitioning the sample set based on the weighted distance between the samples, including: All samples are used as candidate samples for the training set. The training set is filtered based on the comprehensive distance between the samples, and the two candidate samples with the greatest Euclidean distance are selected and included in the training set. Calculate the weighted distance from each remaining candidate sample to the selected samples in the training set, and select the sample with the largest Euclidean distance among all the selected samples in the training set and include it in the training set. Repeat the above screening steps until the number of samples in the training set reaches the preset requirement.
[0012] One possible implementation involves extracting wavelength points with non-zero coefficients from the near-infrared spectrum and using these wavelength points as feature variables, including: Feature variables were selected from the near-infrared spectral data using a regularized regression method, specifically LASSO regression. The loss function for LASSO regression is expressed as follows:
[0013] In the formula, Represents the loss function. Indicates the number of samples. Indicates the first i The actual output of each sample Indicates the first i The input feature vector of each sample, This represents the vector of regression coefficients to be estimated. This represents the regularization parameter, which controls the strength of the penalty term. yes Norm, t denotes transpose; The optimal regression coefficient vector is obtained using the coordinate descent method. w We selected the wavelength points corresponding to non-zero regression coefficients as feature variables.
[0014] One possible implementation involves using a partial least squares regression algorithm to construct a near-infrared spectral model based on the feature variables and their corresponding actual INDF values, including: A near-infrared spectral model was established by performing coefficient regression on the spectrum and INDF using partial least squares (PLS). in, The regression formula is as follows:
[0015] In the formula, Indicates the INDF test value. This is a spectral matrix, representing the absorbance after selecting characteristic variables. Represents the regression coefficient matrix. Represents the error matrix; The X and Y matrices are decomposed according to the PLS algorithm to obtain the score matrix T, load matrix P, and residual matrix E of the X matrix, and the load matrix Q and residual matrix F of the Y matrix, as follows:
[0016]
[0017] Based on the score matrix T of the X matrix, the loading matrix P, and the loading matrix Q of the Y matrix, the regression coefficient matrix B is solved, and expressed as:
[0018] in, W* represents the load weight matrix of X, and t represents the transpose.
[0019] One possible implementation involves optimizing the model parameters using cross-validation, and determining the optimal combination of model parameters by combining the root mean square error of cross-validation (RMSECV) and the root mean square error of external validation (RMSEP), including: Using a 10-fold cross-validation method, the evaluation parameters RMSECV and RMSEP are calculated. When the evaluation parameters RMSECV and RMSEP are at their minimum, the factor number of the partial least squares (PLS) is determined, expressed as:
[0020] Select the parameters that result in the lowest BIC as the model parameters; in, n To model the number of samples, k The number of factors selected for PLS MSE The formula is expressed as:
[0021] In the formula, , Samples i The predicted value and the corresponding INDF detection value, where N represents the total number of samples.
[0022] Compared with the prior art, the technical solution provided in this application has the following beneficial effects: This application constructs a dedicated near-infrared spectral model, combining the portability and real-time capability of a handheld near-infrared device to complete sample scanning and INDF content prediction on-site at the farm. This eliminates the need to send samples to a laboratory for cumbersome wet chemical testing, significantly shortening the testing cycle and saving time and costs associated with sample transportation and prolonged rumen incubation. From an operational perspective, the device is simple to use; farm staff can operate it after training, eliminating the need for specialized technicians and reducing skill requirements. Sample pretreatment requires only simple steps such as drying and pulverizing, further simplifying the testing process. Combined with LASSO feature variable selection, PLS regression algorithm, and 10-fold cross-validation optimization, the model's prediction accuracy is reliable, effectively meeting the farm's need for rapid and accurate detection of INDF content in TMR of lactating cows. This provides crucial data support for optimizing lactating cow diet formulations and refining feeding management. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating a method for detecting INDF content in TMR of lactating cows based on near-infrared spectroscopy, as provided in Embodiment 1 of this application.
[0024] Figure 2The flowchart illustrates a method for detecting INDF content in TMR of lactating cows based on near-infrared spectroscopy, as provided in Example 2 of this application. Detailed Implementation
[0025] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0026] Example 1 See Figure 1 This is a flowchart illustrating a method for detecting INDF content in the TMR of lactating cows based on near-infrared spectroscopy, provided in Embodiment 1 of this application. Figure 1 As shown, the specific implementation steps of the above method include: Step 101: Collect TMR samples from lactating cows and obtain the near-infrared spectral data and actual INDF values corresponding to the above lactating cow TMR samples.
[0027] Step 102: Extract the wavelength points with non-zero coefficients from the above near-infrared spectral data, and use these wavelength points as feature variables.
[0028] Step 103: Using the partial least squares regression algorithm, a near-infrared spectral model that can be used for INDF content detection is constructed based on the above feature variables and the corresponding actual INDF values.
[0029] Step 104: Optimize the model parameters using cross-validation, and determine the optimal combination of model parameters by combining the root mean square error of cross-validation (RMSECV) and the root mean square error of external validation (RMSEP).
[0030] Step 105: Input the near-infrared spectral data of the TMR sample of the lactating cow to be tested into the above near-infrared spectral model to obtain the INDF content detection results.
[0031] Compared with the prior art, the technical solution provided in Embodiment 1 of this application has the following beneficial effects: This technical solution first collects TMR samples from lactating cows and simultaneously acquires near-infrared spectral data and actual INDF values, providing accurate and matching basic data support for model construction. By extracting wavelengths with non-zero coefficients from the near-infrared spectral data as feature variables, key information strongly correlated with INDF content is effectively screened out, redundant interference is eliminated, and the model's computational efficiency and specificity are improved. A partial least squares regression algorithm is used to construct the near-infrared spectral model, adapting to the complex correlation between spectral data and component content, ensuring the model's fitting ability and predictive foundation. Cross-validation is used to optimize model parameters, and the optimal parameter combination is determined by combining the root mean square error of cross-validation (RMSECV) and the root mean square error of external validation (RMSEP), significantly improving the model's stability, accuracy, and generalization ability. Finally, only the near-infrared spectral data of the sample to be tested needs to quickly output the iNDF content detection results. The entire process is simple to operate, highly efficient, requires no complex preprocessing or professional laboratory conditions, and can meet the rapid detection needs of farms, providing timely and reliable evidence for TMR nutritional regulation, effectively solving the problems of cumbersome, time-consuming, and inaccurate traditional detection methods.
[0032] Example 2 See Figure 2 This is a flowchart of a method for detecting INDF content in TMR of lactating cows based on near-infrared spectroscopy, provided in Embodiment 2 of this application. Figure 2 As shown, the specific implementation steps of the above method include: Step 201: Collect TMR samples from lactating cows of different sources and with different formulations, and dry and pulverize the above-mentioned TMR samples from lactating cows.
[0033] Specifically, this application involves drying and pulverizing a TMR sample from a lactating cow to a particle size of 1.0 mm.
[0034] Step 202: Perform spectral scanning on the processed TMR samples from lactating cows to obtain spectral data within the preset band.
[0035] The near-infrared spectral scanning was performed in the preset wavelength range of 950 nm to 1650 nm, with 351 spectral data points. The sample was reloaded twice during the scanning process. First, the pulverized sample was thoroughly mixed and poured into the scanning container for the first spectral scan. After the first spectral scan, the sample was removed from the scanning container, thoroughly mixed a second time, poured back into the scanning container, and then subjected to a second spectral scan. Finally, the average of the spectral data obtained from the first and second spectral scans was taken as the near-infrared spectral data corresponding to the lactating cow TMR sample.
[0036] Step 203: Determine the actual INDF value corresponding to the TMR sample from lactating cows.
[0037] Specifically, this application uses the NorFor INDF testing standard. A 2-gram sample (NDF content of neutral detergent fibers, denoted as 'a') is placed in a bag with a pore size of 12 micrometers and an effective surface area of 100-200 square centimeters. The bag is then incubated in the rumen for 288 hours (the bag uses polyester material with a pore size of 12 micrometers and an open porosity of 7%). After 288 hours, the bag is removed, washed, and dried. The sample is then removed from the bag, crushed, and its NDF content (denoted as 'b') is measured. The INDF content is calculated using 'a' and 'b', and expressed as:
[0038] Step 204: Perform preprocessing operations on the above near-infrared spectral data, including normal variable transformation (SNV), multivariate scattering correction (MSC), detrending processing, smoothing using a Savitzky-Golay smoothing filter, and differentiation processing.
[0039] Specifically, a normal variable transformation (SNV) is performed on the near-infrared spectral data to remove background data from the spectrum, as shown below:
[0040] in The absorbance of the original spectrum. This represents the average absorbance of a single sample spectrum. The standard deviation of the spectral absorbance of a single sample is expressed as:
[0041] Multivariate scattering correction (MSC) is performed on near-infrared spectral data to correct for scattering components in the spectrum and reduce the influence of non-detection absorption. Specifically, MSC completes the correction by calculating the average spectrum of the samples, establishing a linear regression equation between the spectrum of each sample and the average spectrum, and combining this with the multivariate scattering correction formula.
[0042] The average spectrum is expressed as:
[0043] Linear regression, expressed as:
[0044] Multivariate scattering correction, denoted as
[0045] In the formula, A i Let be the spectral vector of the i-th sample. For the average spectrum, m i b i Let be the regression slope and intercept of the i-th sample, respectively.
[0046] Detrend the near-infrared spectral data to eliminate trend variations in the spectrum. Specifically, perform a linear regression between absorbance and wavelength, expressed as:
[0047] in Let be the i-th wavelength, and a and b be univariate regression coefficients.
[0048] Based on formula Complete the detrending process.
[0049] Near-infrared spectral data is smoothed by employing a Savitzky-Golay (SG) smoothing filter, selecting an odd number of data points to form a smoothing window, establishing a polynomial smoothing curve, solving the Vandermonde matrix coefficients to achieve spectral data smoothing, and further optimizing the spectral data quality and reducing noise fluctuations by combining derivative processing.
[0050] Specifically, take m (an odd number) consecutive data points, with intervals h for x, and define the variable z:
[0051] Let m points be used for smoothing, and let the values of the m points be z, respectively.
[0052] A smooth curve with m points is represented as:
[0053] Find the interval coefficient 'a', which can be expressed as:
[0054] Where J is the Vandermonde matrix, and the i-th row of J is...
[0055] If the SG parameters use 5-point 3rd degree polynomial smoothing, i.e. m=5, k=3, then:
[0056] The absorbance data after smoothing and differentiation is as follows (taking a 5-point 3rd-order polynomial fitting as an example):
[0057] The order of the derivative should be selected based on the results of cross-validation, and this application does not impose a specific limitation.
[0058] Step 205: Combine the above near-infrared spectral data with the actual INDF values to construct a sample set. Using a partitioning method that takes into account both spectral characteristics and component content distribution, the sample set is divided into a modeling set and a validation set.
[0059] In this embodiment, the nearest neighbor pairing sample selection algorithm SPXY is used to divide the above-mentioned sample set into a modeling set and a validation set based on spectral distance and INDF content distance.
[0060] The spectral distance is expressed as:
[0061] In the formula, p,q ∈ [1,N], where N is the total number of samples. x p , x q The first p, q Spectral data of each sample; The INDF content distance is represented as follows:
[0062] In the formula, y p , y q The first p, q The actual INDF value of each sample; The overall distance between samples is obtained by summing the spectral distance and the INDF content distance after standardization. The sample set is then partitioned based on this overall distance. The overall distance between samples is expressed as: .
[0063] Regarding the above-mentioned method of partitioning the sample set based on the comprehensive distance between samples, as an feasible approach, this embodiment of the application treats all samples as candidate samples for the training set, and sequentially selects samples from them to enter the training set. Specifically, firstly, the two samples with the greatest Euclidean distance are selected to enter the training set. Then, by calculating the Euclidean distance from each remaining sample to each known sample in the training set, the candidate sample with the largest distance among the minimum distances is found and added to the training set, and so on, until the required number of samples is reached.
[0064] Step 206: Extract the wavelength points with non-zero coefficients in the processed near-infrared spectrum, and use these wavelength points as characteristic variables.
[0065] Specifically, in this embodiment, a regularized regression method is used to select feature variables from the near-infrared spectral data. The regularized regression method is LASSO regression, and the loss function of LASSO regression is expressed as:
[0066] In the formula, Represents the loss function. Indicates the number of samples. Indicates the first iThe actual output of each sample Indicates the first i The input feature vector of each sample, This represents the vector of regression coefficients to be estimated. This represents the regularization parameter, which controls the strength of the penalty term. yes Norm. t denotes transpose.
[0067] This application embodiment employs a regularized regression method to construct a loss function containing a regularization term, solves for the optimal regression coefficients, and selects characteristic wavelength points that significantly influence INDF content prediction as feature vectors. As one feasible approach, the coordinate descent method is used to solve for the optimal regression coefficient vector. w We selected the wavelength points corresponding to non-zero regression coefficients as feature variables.
[0068] Step 207: Using partial least squares regression algorithm, construct a near-infrared spectral model based on the above feature variables and the corresponding actual INDF values.
[0069] Specifically, partial least squares (PLS) was used to perform coefficient regression on the spectrum and INDF to establish a near-infrared spectral model. in, The regression formula is as follows:
[0070] In the formula, Indicates the INDF test value. This is a spectral matrix, representing the absorbance after selecting characteristic variables. Represents the regression coefficient matrix. Represents the error matrix; The X and Y matrices are decomposed according to the PLS algorithm to obtain the score matrix T, load matrix P, and residual matrix E of the X matrix, and the load matrix Q and residual matrix F of the Y matrix, as follows:
[0071]
[0072] Based on the score matrix T of the X matrix, the loading matrix P, and the loading matrix Q of the Y matrix, the regression coefficient matrix B is solved, and is expressed as:
[0073] in, W* represents the load weight matrix of X. t represents the transpose.
[0074] Step 208: Optimize the model parameters using cross-validation, and determine the optimal combination of model parameters by combining the root mean square error of cross-validation (RMSECV) and the root mean square error of external validation (RMSEP).
[0075] Specifically, the RMSECV and RMSEP evaluation parameters are calculated using a 10-fold cross-validation method. When the RMSECV and RMSEP evaluation parameters are at their minimum, the factor number of the partial least squares (PLS) is determined, expressed as:
[0076] Select the parameters that result in the lowest BIC as the model parameters; in, n To model the number of samples, k The number of factors selected for PLS MSE The formula is expressed as:
[0077] In the formula, , Samples i The predicted value and the corresponding INDF detection value.
[0078] The root mean square error of cross-validation, RMSECV, is expressed as:
[0079] The root mean square error (RMSEP) of external validation is expressed as:
[0080] In the formula, , , where i is the predicted value and the corresponding INDF detection value, respectively, and N represents the total number of samples.
[0081] Step 209: Input the near-infrared spectral data of the TMR sample of the lactating cow to be tested into the near-infrared spectral model to obtain the INDF content detection result.
[0082] Compared with the prior art, the technical solution provided in Embodiment 2 of this application has the following beneficial effects: By employing a scientific approach to sample collection design, multi-step spectral preprocessing, LASSO feature variable selection, and PLS regression modeling, the constructed INDF prediction model (RSQ=0.69, SEC=1.48) can accurately predict INDF content, effectively replacing traditional wet chemical detection methods and avoiding the latter's drawbacks of being cumbersome, time-consuming, and dependent on professional laboratory conditions.
[0083] Example 3 This application provides a near-infrared detection device for implementing the INDF content method as described in Examples 1 and 2, comprising a hardware unit and an embedded INDF prediction model. The hardware unit has near-infrared spectral acquisition capabilities, enabling it to acquire near-infrared spectral data of lactating bovine TMR samples. The embedded INDF prediction model is the near-infrared detection model constructed in Examples 1 and 2. The device can acquire spectral data of the sample to be tested through the hardware unit, calculate the INDF content using the embedded model, and output the detection result.
[0084] Compared with the prior art, the technical solution provided in Embodiment 3 of this application has the following beneficial effects: Paired with handheld near-infrared devices, the testing scenario extends from the laboratory to the farm field. Staff can operate the equipment after simple training, eliminating the need for complex sample transportation and providing real-time test results. This significantly shortens the testing cycle and reduces testing costs. The equipment is also highly portable and stable, adapting to the diverse on-site testing needs of farms. Furthermore, the entire technical process balances model versatility with ease of testing. It ensures model compatibility with TMR from lactating cows of different origins through wide-ranging, multi-formulation sample coverage, while its simplified operation and real-time testing capabilities provide efficient technical support for farms to precisely regulate lactating cow diets, ensuring balanced nutrition and optimal production performance.
[0085] Although embodiments of this application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for detecting INDF content in TMR of lactating cows based on near-infrared spectroscopy, characterized in that, include: A near-infrared spectral model for detecting INDF content is constructed. The sample set for constructing the near-infrared spectral model includes near-infrared spectral data and actual INDF values corresponding to TMR samples from lactating cows. The construction method includes extracting wavelength points with non-zero coefficients from the near-infrared spectral data and using these wavelength points as feature variables. A partial least squares regression algorithm is used to construct the near-infrared spectral model based on the feature variables and the corresponding actual INDF values. The model parameters are optimized using cross-validation, and the optimal combination of model parameters is determined by combining the root mean square error of cross-validation (RMSECV) and the root mean square error of external validation (RMSEP), thus obtaining the near-infrared spectral model for detecting INDF content. The near-infrared spectral data of the TMR sample from the lactating cow to be tested are input into the near-infrared spectral model to obtain the INDF content detection results.
2. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 1, characterized in that, Methods for obtaining near-infrared spectral data corresponding to TMR samples from lactating cows include: After drying, the TMR sample from lactating cows was pulverized to a particle size of 1.0 mm. The pulverized sample was then re-loaded at least twice to obtain the near-infrared spectral data. The two re-loading operations included first mixing the pulverized sample evenly and pouring it into a scanning box for a first spectral scan; after the first spectral scan, the sample was removed from the scanning box and mixed evenly a second time, then poured back into the scanning box for a second spectral scan; the average of the spectral data obtained from the first and second spectral scans was taken as the near-infrared spectral data corresponding to the lactating cow TMR sample.
3. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 1, characterized in that, Methods for obtaining the actual INDF values of TMR samples from lactating cows include: Take 2 grams of TMR sample from lactating cows with known NDF content a, place it in a filter bag made of polyester material with a pore size of 12 micrometers, an effective surface area of 100-200 square centimeters and an open porosity of 7%, and incubate the filter bag in the rumen for 288 hours. After incubation, the filter bag is removed, cleaned, and dried. The sample inside the bag is taken out and its NDF content is measured and recorded as b. The actual NDF value corresponding to the lactating cow TMR sample is calculated according to the formula INDF=b / a.
4. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 1, characterized in that, After obtaining the near-infrared spectrum corresponding to the TMR sample of a lactating cow, the method further includes: The near-infrared spectral data is preprocessed, including performing normal variable transformation (SNV), multivariate scattering correction (MSC), detrending, smoothing with a Savitzky-Golay smoothing filter, and differentiation.
5. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 1, characterized in that, After obtaining the near-infrared spectral data and actual INDF values of the TMR samples from lactating cows to construct the sample set, the method further includes: The nearest neighbor pairing sample selection algorithm SPXY is used to divide the sample set into a modeling set and a validation set based on spectral distance and INDF content distance; The spectral distance is expressed as: In the formula, p,q ∈ [1,N], where N is the total number of samples. x p , x q The first p, q Spectral data of each sample; The INDF content distance is represented as follows: In the formula, y p , y q The first p, q The actual INDF value of each sample; The overall distance between samples is obtained by summing the spectral distance and the INDF content distance after standardization, and the sample set is divided based on the overall distance between samples; the overall distance between samples is expressed as: 。 6. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 5, characterized in that, The sample set is partitioned based on the weighted distance between the samples, including: All samples are used as candidate samples for the training set. The training set is filtered based on the comprehensive distance between the samples, and the two candidate samples with the greatest Euclidean distance are selected and included in the training set. Calculate the weighted distance from each remaining candidate sample to the selected samples in the training set, and select the candidate sample with the largest distance among the minimum Euclidean distances to all selected samples in the training set to include in the training set; Repeat the above screening steps until the number of samples in the training set reaches the preset requirement.
7. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 1, characterized in that, Extracting wavelengths with non-zero coefficients from the near-infrared spectrum, and using these wavelengths as feature variables, includes: Feature variables were selected from the near-infrared spectral data using a regularized regression method, specifically LASSO regression. The loss function for LASSO regression is expressed as follows: In the formula, Represents the loss function. Indicates the number of samples. Indicates the first i The actual output of each sample Indicates the first i The input feature vector of each sample, This represents the vector of regression coefficients to be estimated. This represents the regularization parameter, which controls the strength of the penalty term. yes Norm, t denotes transpose; The optimal regression coefficient vector is obtained using the coordinate descent method. w We selected the wavelength points corresponding to non-zero regression coefficients as feature variables.
8. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 1, characterized in that, A near-infrared spectral model is constructed using a partial least squares regression algorithm based on the aforementioned feature variables and their corresponding actual INDF values, including: A near-infrared spectral model was established by performing coefficient regression on the spectrum and INDF using partial least squares (PLS). in, The regression formula is as follows: In the formula, Indicates the INDF test value. This is a spectral matrix, representing the absorbance after selecting characteristic variables. Represents the regression coefficient matrix. Represents the error matrix; The X and Y matrices are decomposed according to the PLS algorithm to obtain the score matrix T, load matrix P, and residual matrix E of the X matrix, and the load matrix Q and residual matrix F of the Y matrix, as follows: Based on the score matrix T of the X matrix, the loading matrix P, and the loading matrix Q of the Y matrix, the regression coefficient matrix B is solved, and expressed as: in, W* represents the load weight matrix of X, and t represents the transpose.
9. The method for detecting INDF content in TMR based on near-infrared spectroscopy according to claim 1, characterized in that, The model parameters are optimized using cross-validation. The optimal combination of model parameters is determined by combining the root mean square error of cross-validation (RMSECV) and the root mean square error of external validation (RMSEP), including: Using a 10-fold cross-validation method, the RMSECV and RMSEP evaluation parameters are calculated. When the RMSECV and RMSEP evaluation parameters are at their minimum, the factor number of the partial least squares (PLS) is determined, expressed as: Select the parameters that result in the lowest BIC as the model parameters; in, n To model the number of samples, k The number of factors selected for PLS MSE The formula is expressed as: In the formula, , Samples i The predicted value and the corresponding INDF detection value, where N represents the total number of samples.
Citation Information
Cited By
SERS (Surface Enhanced Raman Scattering) spectrum quantitative detection method and system based on interpretable stacked ensemble learning
CN121838947A