Rapid beef ingredient detection method based on near infrared spectrum and multivariable modeling

By employing near-infrared spectroscopy and multivariate modeling, the problem of spectral variation caused by breed, environment, and post-slaughter time in beef component detection was solved. A rapid detection model for beef components with broad applicability was established, enabling efficient and accurate beef component analysis.

CN121595489APending Publication Date: 2026-03-03INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511785378.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing near-infrared spectroscopy technology for beef component detection suffers from spectral variation interference caused by factors such as breed, environment, and post-slaughter time, resulting in poor model versatility and insufficient robustness, making it difficult to achieve rapid detection in diverse scenarios.

Method used

Near-infrared spectrometers were used to collect spectral data of beef samples. Outlier testing and preprocessing were performed. A prediction model for beef component content was established by using multivariate modeling methods such as partial least squares regression, combined with competitive adaptive reweighted sampling and cross-validation.

Benefits of technology

It significantly improves the model's generalization ability and robustness, enabling accurate, stable, and rapid detection of various mainstream beef cattle breeds, thus meeting the high-efficiency quality control needs of the modern meat industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121595489A_ABST
    Figure CN121595489A_ABST
Patent Text Reader

Abstract

The invention discloses a beef component rapid detection method based on near infrared spectroscopy and multivariable modeling, and the method comprises the following steps: collecting spectral data of a beef sample by using a near infrared spectrometer, and determining the content of each component in the beef sample; carrying out abnormal value testing on the acquired near infrared spectrum data and the content of each component, and removing abnormal points; preprocessing the original spectral data, and dividing residual samples into a training set and a test set; on the basis of the training set, screening characteristic variables related to a target component from a full wave band; establishing a beef component content prediction model by using the training set and the corresponding characteristic variables by using a partial least square regression method, and testing the model by using the test set; and inputting to-be-detected spectral data into the corrected prediction model to obtain the content of the components in the to-be-detected beef sample. According to the method, the generalization ability and robustness of the model are remarkably improved, and wide applicability and accurate and stable rapid detection covering various mainstream beef cattle varieties are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of meat testing technology, and in particular to a rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling. Background Technology

[0002] Beef, as a meat product with significant nutritional and economic value in the global consumer market, relies heavily on its moisture, intramuscular fat, and protein content as core indicators for evaluating tenderness, flavor characteristics, and processing suitability. These factors also serve as crucial bases for meat industry grading and pricing, and for ensuring consumers' right to know. Currently, traditional testing methods such as direct drying (moisture determination), acid hydrolysis (intramuscular fat determination), and the Kjeldahl method (protein determination), while providing accurate results, generally suffer from inherent limitations such as cumbersome procedures, long processing times (a single test can take several hours to tens of hours), the use of chemical reagents, and sample destruction. These limitations make them unsuitable for the modern meat industry's demands for multi-batch, high-efficiency quality control, and they cannot achieve rapid, dynamic on-site testing.

[0003] Near-infrared spectroscopy (NIRS) is a rapid and non-destructive detection method based on the overtone and combination frequency absorption of vibrations of hydrogen-containing groups (such as OH, CH, NH, etc.) in molecules for qualitative and quantitative analysis. It can simultaneously determine the content of multiple components, providing a new technical path for the aforementioned traditional detection methods, and has been widely used in the field of agricultural product and food quality testing. However, for beef samples, the near-infrared spectral detection of its components still faces challenges from several interfering factors: different breeds of beef (such as Simmental, Wagyu, Angus, etc.) have significant differences in intramuscular fat distribution due to differences in genetic background; different experimental environments (temperature and humidity differences) and instrument status parameters may affect the stability and repeatability of spectral acquisition; time differences after slaughter (such as 12 h, 24 h, 48 h post-slaughter) will lead to a series of changes during beef maturation, such as water migration, protein degradation, and fat oxidation, resulting in more complex spectral response characteristics and increased overlap. The aforementioned factors significantly increase the difficulty of accurate prediction from spectral information, resulting in existing near-infrared spectral models generally having poor versatility and insufficient robustness, making them difficult to widely apply to the rapid detection of beef components in real-world diverse scenarios.

[0004] Therefore, proposing a rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling to overcome the difficulties of existing technologies is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling, which effectively overcomes the interference of spectral variations caused by factors such as breed, environment and post-slaughter time, significantly improves the generalization ability and robustness of the model, and achieves wide applicability and accurate and stable rapid detection covering a variety of mainstream beef cattle breeds.

[0006] To achieve the above objectives, the present invention provides the following solution: A rapid method for detecting beef components based on near-infrared spectroscopy and multivariate modeling includes the following steps: S1. Near-infrared spectrometer was used to collect spectral data of beef samples. Each beef sample was scanned twice and the average value was taken as the near-infrared spectral data of the beef sample. S2. Determine the content of each component in the beef sample, including moisture, intramuscular fat and protein, according to standard testing methods; S3. Perform outlier testing on the collected near-infrared spectral data and the content of each component, and remove outliers. S4. Preprocess the spectral data after removing outliers, divide the preprocessed dataset into a training set and a test set, and based on the training set, screen feature variables related to the target components from the entire spectrum. S5. Using partial least squares regression, a beef component content prediction model is established using the training set and its corresponding feature variables, and the model is tested using the test set. S6. Input the near-infrared spectral data of the beef sample to be tested into the trained beef component content prediction model to obtain the moisture, intramuscular fat and protein content of the beef sample to be tested.

[0007] Preferably, in S1, the scanning type of the spectral data of the beef sample acquired using a near-infrared spectrometer is Fourier transform near-infrared, and the wavenumber range of the acquired data is 11520 cm⁻¹. -1 Up to 3952 cm -1 The background was scanned 32 times, and the sampling method was diffuse reflection.

[0008] Preferably, the beef sample in S1 is a frozen sample of the longissimus dorsi muscle. The beef sample is thawed 24 hours in advance at a temperature of 4°C and minced into a paste.

[0009] Preferably, the standard detection method used in S2 is: Referring to GB 5009.3-2016 "National Food Safety Standard - Determination of Moisture in Food", the moisture content in beef samples was determined by direct drying method; referring to GB 5009.6-2016 "National Food Safety Standard - Determination of Fat in Food", the intramuscular fat content in beef samples was determined by acid hydrolysis method; referring to GB 5009.5-2016 "National Food Safety Standard - Determination of Protein in Food", the protein content in beef samples was determined by Kjeldahl method.

[0010] Preferably, the outlier test in S3 uses Monte Carlo cross-validation.

[0011] Preferably, in S4, the preprocessing method used includes at least one of the following: standard normal transformation, multivariate scattering correction, maximum-minimum normalization, Savitzky-Golay first derivative, and Savitzky-Golay second derivative. The preprocessed near-infrared spectral data is selected from the entire band using a competitive adaptive reweighted sampling method to screen feature variables related to the target components. The dataset is partitioned using the SPXY algorithm, with a training set to test set ratio of 8:2.

[0012] Preferably, in S5, the partial least squares regression method is used to establish a prediction model for beef component content, including: The maximum number of principal components was set to 20. The optimal number of principal components was determined by 5-fold cross-validation, and quantitative models for water, intramuscular fat and protein content were established respectively.

[0013] Preferably, the rapid detection method for beef components further includes: using the coefficient of determination, root mean square error, and residual prediction bias to comprehensively evaluate the performance of the beef component content prediction model.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling as described above.

[0015] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: (1) The method of this invention integrates frozen beef samples from multiple sources, varieties, ages, and genders to construct a widely representative sample set, effectively overcoming the overfitting problem caused by the single sample in traditional models, and significantly improving the model's versatility and generalization ability. Secondly, by combining preprocessing methods such as standard normal transformation and multivariate scattering correction with a competitive adaptive reweighted sampling algorithm, spectral noise and scattering interference are effectively eliminated, and feature information closely related to the target components is screened out, thereby improving the model's prediction accuracy and robustness. Finally, the optimal number of principal components is determined through cross-validation, and a quantitative model is established based on partial least squares regression, providing a new approach for the rapid, non-destructive, and reliable detection of the content of major beef components.

[0016] (2) The method of the present invention effectively overcomes the interference of spectral variation caused by factors such as breed, environment and post-slaughter time through the design of large-scale, multi-sample training and test sets, significantly improves the generalization ability and robustness of the model, and achieves wide applicability and accurate, stable and rapid detection covering a variety of mainstream beef cattle breeds. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic flowchart illustrating the rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling provided by this invention; Figure 2 The original spectral data and the optimally preprocessed spectral image of the beef sample in Example 1 of this invention are shown. Among them, (a) is the original spectrum of the meat sample, and (b) is the spectrum after optimal pretreatment; Figure 3 This is a diagram showing the results of outlier removal using the Monte Carlo Cross-Validation (MCCV) method in Embodiment 1 of the present invention. Among them, (a) represents the removal of samples with abnormal water content, (b) represents the removal of samples with abnormal intramuscular fat content, and (c) represents the removal of samples with abnormal protein content. Figure 4 This is a distribution diagram of the true and predicted values ​​based on the partial least squares regression model in Embodiment 1 of the present invention; Among them, (a) is the distribution map of the actual and predicted values ​​of water content, (b) is the distribution map of the actual and predicted values ​​of intramuscular fat content, and (c) is the distribution map of the actual and predicted values ​​of protein content.

[0019] Figure 5 This is a diagram showing the relative error distribution of predictions based on the optimal model in Embodiment 2 of the present invention; Among them, (a) is the distribution map of relative error in water content prediction, (b) is the distribution map of relative error in intramuscular fat content prediction, and (c) is the distribution map of relative error in protein content prediction. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] Example 1 In this embodiment, a total of 5,946 beef samples were collected for model establishment and internal validation, including 14 local varieties, 6 cultivated varieties, 5 introduced varieties and 21 hybrid populations, covering multiple different regions.

[0023] like Figure 1 As shown, a rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling specifically includes the following steps: (1) Instrument preparation and parameter setting Under room temperature conditions, the near-infrared spectrometer was fully preheated and its optical performance was verified (OVP test). The spectral scan type was Fourier transform near-infrared, the wavenumber range was 11520 cm-1 to 3952 cm-1, the background scan was performed 32 times, and the sampling method was diffuse reflectance.

[0024] (2) Sample preparation and spectral acquisition Several frozen samples of the longissimus dorsi muscle from different beef varieties (stored at -20℃) were selected. The samples were thawed at 4℃ for 24 hours. The thawed beef should be bright red or dark red with a glossy sheen; the fat should be milky white or yellow; the muscle structure should be firm and compact; the muscle fibers should be resilient; and the beef should have a characteristic product aroma without any off-odors. After removing visible fascia, the beef was repeatedly ground until homogenized. Near-infrared spectra of the beef were collected in a minced state. Each sample was scanned twice. After meeting repeatability requirements, the average spectrum was used as the original spectrum. Figure 2 As shown in (a).

[0025] (3) Determination of nutrient content Moisture content detection The moisture content in beef samples was determined using the direct drying method in accordance with GB 5009.3-2016, "National Food Safety Standard - Determination of Moisture in Food".

[0026] Intramuscular fat content testing The intramuscular fat content in beef samples was determined using the acid hydrolysis method in accordance with GB 5009.6-2016, "National Food Safety Standard - Determination of Fat in Food".

[0027] Protein content detection The protein content in beef samples was determined using the Kjeldahl method in accordance with GB 5009.5-2016, "National Food Safety Standard - Determination of Protein in Food".

[0028] (4) Abnormal sample removal and spectral data preprocessing Monte Carlo cross-validation (MCCV) is used to identify and remove outlier samples based on the relationship between the original spectra and the content of each component. Figure 3 As shown in (a)-(c), after removing outlier samples, the original spectral data is preprocessed to reduce noise and scattering interference. The methods used include one of the following: Standard Normal Transform (SNV), Multivariate Scattering Correction (MSC), Max-Min Normalization, Savitzky-Golay First Derivative, and Savitzky-Golay Second Derivative.

[0029] (5) Dataset partitioning and feature variable selection The SPXY algorithm was used to divide the preprocessed sample set into a training set and a test set in an 8:2 ratio. The range, mean, and standard deviation distribution of water content, intramuscular fat, and protein content in the training and test sets are shown in Table 1. Further, based on the training set, a competitive adaptive reweighting algorithm (CARS) was used to screen the feature variables most relevant to the target components from the entire band. In this embodiment, the Monte Carlo iterations were set to 50, with 5-fold cross-validation used in each iteration.

[0030] Table 1

[0031] (6) Prediction model establishment and optimization Based on the selected feature variables, a partial least squares regression (PLSR) model was established to predict the content of beef components. The optimal number of factors was determined through cross-validation to optimize the model complexity and avoid overfitting.

[0032] The near-infrared spectrometer was used to scan the spectrum of the beef sample to be tested. After preprocessing the obtained spectrum, it was input into the prediction model to obtain the moisture, intramuscular fat and protein content of the beef sample to be tested.

[0033] (7) Model performance evaluation The performance of the models was comprehensively evaluated using the coefficient of determination (R²), root mean square error (RMSE), and residual prediction bias (RPD). Table 2 shows the performance comparison results of models established using different spectral preprocessing methods. The comprehensive comparison indicates that preprocessing using the Savitzky-Golay second derivative method (SG_2nd)... Figure 2 The PLSR model established after (b) showed excellent performance on all indicators. Scatter plots of the true and predicted values ​​of the test set for each quantitative model are shown below. Figure 4 The specific data is as follows: For moisture content, the training set determination coefficient (Rc²) of the SG_2nd preprocessed model is 0.956, the test set determination coefficient (Rp²) is 0.979; the training set root mean square error (RMSEC) is 1.052, the test set root mean square error (RMSEP) is 1.205, and the RPD value is 6.981. For intramuscular fat content, the Rc² of the SG_2nd preprocessed model was 0.971, Rp² was 0.984, RMSEC was 1.019, RMSEP was 1.285, and RPD was 7.891. For protein content, the Rc² of the SG_2nd preprocessed model was 0.855, Rp² was 0.922, RMSEC was 0.673, RMSEP was 0.723, and RPD was 3.591.

[0034] Although the training set determination coefficient of the SG_2nd preprocessing model was slightly lower than that of other preprocessing methods for protein indicators, its test set determination coefficient and RPD value were both higher, indicating that the model has better generalization ability and prediction robustness. In conclusion, for the quantitative analysis of the above three components, SG_2nd is the preferred spectral preprocessing method.

[0035] Table 2

[0036] Example 2 568 beef samples were collected for external model validation. The samples included 14 local varieties, 6 bred varieties, 4 introduced varieties and 16 hybrid populations.

[0037] For the beef samples to be tested, sample preprocessing and near-infrared spectra were collected according to step (2) of Example 1. After SG_2nd preprocessing, the samples were input into the established PLSR model to obtain the predicted values ​​of moisture, intramuscular fat, and protein content of the samples. Using the method for determining the content of each index of beef samples described in step (3) of Example 1, the measured data of moisture, intramuscular fat, and protein content of 568 samples were determined. The measured data were compared with the predicted data, and their absolute error and relative error were calculated. The results are as follows. Figure 5 (The results of measured values ​​and relative error distribution for unknown beef samples are shown). National standards require that the absolute difference between two measurements should not exceed 10% of the arithmetic mean; that is, a relative error within 10% is considered a high precision measurement method. In this study, moisture ( Figure 5 (a) and protein ( Figure 5 The predicted results of (c) are basically consistent with the measured values, with an average relative error of less than 10%, which fully meets the requirements of national standards, indicating that the moisture and protein prediction model has high reliability and accuracy.

[0038] For intramuscular fat, the predicted value generally follows the same trend as the measured value, but the relative error varies with the actual content. Figure 5 (b) When the measured intramuscular fat content is higher than 7g / 100g, the median relative error of the model prediction is 9.714%, indicating high prediction accuracy. When the measured intramuscular fat content is in the low range of 0~7g / 100g, the median relative error is 22.273%, indicating a decrease in prediction accuracy. This may be because the fat content in the samples is relatively low and unevenly distributed within the 0~7g / 100g range, resulting in weak spectral feature signals or large signal variation amplitudes, which to some extent affects the model's prediction accuracy. Overall, the model can effectively fit the changing trend of sample fat content. Further robustness can be improved through data cleaning (such as removing abnormal samples and supplementing effective data in the low content range) or feature optimization.

[0039] The measured values ​​and predicted values ​​of each indicator were imported into SPSS 22.0 software for paired-samples t-tests, and the results are shown in Table 3. No significant differences were found between the predicted and measured values ​​for water, intramuscular fat, and protein content (P>0.05).

[0040] In conclusion, the t-test and error analysis results together verify the feasibility and accuracy of this model in rapidly predicting the content of major components in beef.

[0041] Table 3

[0042] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling as described above.

[0043] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0044] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling, characterized in that, Includes the following steps: S1. Near-infrared spectrometer was used to collect spectral data of beef samples. Each beef sample was scanned twice and the average value was taken as the near-infrared spectral data of the beef sample. S2. Determine the content of each component in the beef sample, including moisture, intramuscular fat and protein, according to standard testing methods; S3. Perform outlier testing on the collected near-infrared spectral data and the content of each component, and remove outliers. S4. Preprocess the spectral data after removing outliers, divide the preprocessed dataset into a training set and a test set, and based on the training set, screen feature variables related to the target component from the entire band. S5. Using partial least squares regression, a beef component content prediction model is established using the training set and its corresponding feature variables, and the model is tested using the test set. S6. Input the near-infrared spectral data of the beef sample to be tested into the trained beef component content prediction model to obtain the moisture, intramuscular fat and protein content of the beef sample to be tested.

2. The rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling according to claim 1, characterized in that, In step S1, the spectral data of the beef sample acquired using a near-infrared spectrometer is scanned using Fourier transform near-infrared, with a wavenumber range of 11520 cm⁻¹. -1 Up to 3952 cm -1 The background was scanned 32 times, and the sampling method was diffuse reflection.

3. The rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling according to claim 1, characterized in that, The beef sample in S1 is specifically a frozen sample of the longissimus dorsi muscle. The beef sample is thawed 24 hours in advance at a temperature of 4°C and minced into a paste.

4. The rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling according to claim 1, characterized in that, The standard detection method used in S2 is as follows: Referring to GB 5009.3-2016 "National Food Safety Standard - Determination of Moisture in Food", the moisture content in beef samples was determined by direct drying method; referring to GB 5009.6-2016 "National Food Safety Standard - Determination of Fat in Food", the intramuscular fat content in beef samples was determined by acid hydrolysis method; referring to GB 5009.5-2016 "National Food Safety Standard - Determination of Protein in Food", the protein content in beef samples was determined by Kjeldahl method.

5. The rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling according to claim 1, characterized in that, The outlier test in S3 uses Monte Carlo cross-validation.

6. The rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling according to claim 1, characterized in that, In S4, the preprocessing methods employed include at least one of the following: standard normal transformation, multivariate scattering correction, maximum-minimum normalization, Savitzky-Golay first derivative, and Savitzky-Golay second derivative. A competitive adaptive reweighted sampling method is used to screen feature variables related to the target components from the entire band of the preprocessed near-infrared spectral data. The dataset is partitioned using the SPXY algorithm, with a training set to test set ratio of 8:

2.

7. The rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling according to claim 1, characterized in that, In S5, the partial least squares regression method is used to establish a prediction model for beef component content, including: The maximum number of principal components was set to 20. The optimal number of principal components was determined by 5-fold cross-validation, and quantitative models for water, intramuscular fat and protein content were established respectively.

8. The rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling according to claim 1, characterized in that, The rapid detection method for beef components also includes using the coefficient of determination, root mean square error, and residual prediction bias to comprehensively evaluate the performance of the beef component content prediction model.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a rapid detection method for beef components based on near-infrared spectroscopy and multivariate modeling as described in any one of claims 1 to 8.