Rapid detection method for brown rice protein content based on Fourier transform infrared spectroscopy
Through Fourier transform infrared spectroscopy combined with characteristic band screening and machine learning algorithms, an optical correlation system for brown rice protein content was constructed, which solved the problem of detecting brown rice protein content in the existing technology, and achieved a fast, accurate and low-cost detection effect.
Patent Information
- Application Number
- CN202411773527.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to quickly and accurately detect the protein content of brown rice. The traditional methods are very destructive, time-consuming and costly, affecting large-scale sample detection and breeding of high-quality rice varieties.
Fourier transform infrared spectroscopy technology is used to combine smooth preprocessing, continuous projection algorithm, competitive adaptive reweighting sampling algorithm, non-information variable elimination, interval combination optimization feature band screening algorithm, as well as support vector machine and partial least squares algorithm, to construct an optical correlation system for brown rice protein content.
It realizes rapid and accurate detection of brown rice protein content, simplifies detection steps, saves time, reduces labor, and reduces measurement costs, and supports food companies to achieve fast, efficient and stable product quality control.
Smart Images

Figure CN120177407A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of spectral analysis and relates to the detection of brown rice quality by applying Fourier transform infrared spectroscopy technology. Background Art
[0002] Rice, as the staple food for 60% of Chinese people, plays a dominant role in ensuring food security. With the improvement of residents' living standards, consumers' requirements for rice quality are constantly increasing. The quality of rice is a comprehensive trait, generally divided into milling quality, appearance quality, cooking and eating quality, and nutritional quality, etc. Protein is the second type of storage substance in rice, accounting for 8%-12% of the dry weight of rice, mainly existing in the aleurone layer and the area near the aleurone layer, and is closely related to the cooking and eating quality and nutritional quality of rice. The traditional determination of protein content in rice usually uses brown rice flour or polished rice flour for determination. This destructive determination method requires a large amount of manpower for grinding brown rice flour and is difficult to achieve rapid detection of large-scale samples. The destructive grinding of rice also makes the seeds unable to be further planted. On the contrary, spectroscopy technology allows rapid acquisition of spectral information of samples with little or no damage to the samples, and realizes the analysis of quality traits such as protein content, moisture, and amylose content in a non-destructive and rapid manner through computer modeling. The commonly used near-infrared grain analyzer on the market often has problems with inaccurate measurement due to the influence of the analysis model and spectral signal acquisition. In recent years, breeding new high-yield and high-quality rice varieties has become the main breeding goal. However, the cumbersome and destructive determination of rice quality traits limits the breeding of high-quality rice varieties.
[0003] Vibration spectroscopy technology is an important means for rapid identification of the quality of agricultural products and processed foods. At present, the common vibration spectroscopy technologies related to rice include near-infrared (NIR) spectroscopy, Fourier transform infrared (FTIR) spectroscopy, hyperspectral imaging (HIS), Raman spectroscopy, fluorescence spectroscopy (FS), terahertz spectroscopy, and laser-induced breakdown spectroscopy (LIBS). These spectroscopy technologies have realized the prediction of quality indicators such as moisture, starch, and freshness. However, the accurate detection of protein content in brown rice has not been effectively achieved based on Fourier transform infrared spectroscopy technology. The fundamental reason is that molecular spectra are mostly band spectra, and spectral lines are prone to overlap, resulting in relatively serious spectral interference, which needs to be improved by combining chemometrics. In some current studies, different chemometric methods have given unsatisfactory prediction results, which may be related to the too small number of samples and the small differences between samples. Different spectroscopy technologies have different principles, so they have different prediction potentials for the quality attributes of brown rice. The method based on characteristic band screening can greatly improve the accuracy of the model and help to enhance the generalization ability of the model. This will contribute to the development of the rice industry. Summary of the Invention
[0004] The object of the present invention is to provide a novel method for rapidly detecting the protein content of brown rice. By using Fourier transform infrared spectroscopy technology, the optical signal of brown rice is obtained and the protein content of brown rice is measured. After preprocessing by smoothing (Savitzky-Golay), the characteristic band screening algorithms of successive projections algorithm (SPA), competitive adaptive reweighted sampling algorithm (CARS), uninformative variable elimination (UVE), and interval combination optimization (ICO) are combined with partial least squares (PLS) and support vector machine (SVM) algorithms for analysis to construct an optical correlation system for the protein content of brown rice, so as to realize the detection of the protein content of brown rice from the optical characteristics of brown rice.
[0005] The beneficial effects of the present invention are as follows: (1) The change laws of the chemical quality and optical characteristics of brown rice are explored, and a high-throughput optical database of brown rice is constructed; (2) A prediction model for the protein content of brown rice by Fourier transform infrared spectroscopy is developed, and the protein content of brown rice is digitally evaluated by using the characteristic band screening method; (3) A method for rapidly predicting the protein content of brown rice is developed, which simplifies the detection steps of the protein content of brown rice. Compared with the traditional determination of the protein content of brown rice, this invention saves time, reduces labor, and significantly reduces the determination cost; (4) Only by rapidly characterizing the optical characteristics of brown rice, the rapid prediction of the protein content of brown rice is realized, which provides rapid, efficient, and stable product quality control for food enterprises. Brief Description of the Drawings
[0006] Figure 1 : Abstract drawing;
[0007] Figure 2 : Schematic diagram of brown rice varieties;
[0008] Figure 3 : Frequency distribution diagram of protein content;
[0009] Figure 4 : Results of different characteristic screening algorithms in the mid-infrared band (a) SPA characteristic wavelength extraction process; (b) Characteristic bands selected by the SPA algorithm; (c) CARS characteristic wavelength extraction process; (d) Characteristic bands extracted by the CARS algorithm;
[0010] Figure 5 : (a) UVE stability diagram; (b) Characteristic bands extracted by the UVE algorithm; (c) Box plot of the root mean square error of the ICO algorithm varying with the number of iterations; (d) Characteristic bands extracted by the ICO algorithm Detailed Embodiments
[0011] A rapid detection method for the protein content of brown rice based on Fourier transform infrared spectroscopy, and the specific embodiments are as follows:
[0012] 1. Experimental Materials and Methods
[0013] Preparation of brown rice samples: 155 rice varieties were sown in the middle and late May. Harvest of mature seeds was completed in batches from mid-September to early October, with 40 days after heading as the standard. Varieties with similar maturity periods were selected. After the collected seeds were naturally air-dried, they were placed indoors for after-ripening for one month, and then brown rice was obtained through hulling and grinding. The rice grains were ground into fine powder through a 100-mesh sieve using a vibrating ball mill GT300 for later use.
[0014] Determination of protein content: Using a rapid N Exceed nitrogen analyzer (elementar, Germany), weigh 250.0 mg each of brown rice powder and aspartic acid (Asp) standard sample with an electronic balance accurate to one ten-thousandth, wrap them well with tin foil and compact. After the nitrogen analyzer finishes self-checking and warming up, open the CO2 and O2 valves and start sample loading. First place three blank samples (Blank), set Method as BlankwithO2; then place 5 standard samples, set Method as 250mg Standard, and set the protein factor as 6.25. Place the brown rice samples of each variety in turn. After sample loading, the instrument automatically analyzes. After the operation ends, read and analyze the data using GraphPad Prism.
[0015] Fourier transform infrared spectroscopy and spectral data collection: Fourier transform attenuated total reflection infrared spectroscopy (ATR-FTIR) uses a Nicolet iS 10 spectrometer (Thermo Fisher Sceientific., USA). This instrument is equipped with an attenuated total reflection (ATR) accessory and supporting OMNIC software (Thermo fisher scientific inc., USA). The design mode of the software sampling workflow (work flow) facilitates the management and control of the spectrometer. During collection, take a small amount of brown rice powder and place it on the surface of the ATR crystal window of the sample disk. Rotate the pressure tower downward to tightly press the powder to make good contact with the ATR crystal. The number of scans is 32 times, and the wavenumber scanning range is 525 - 4000 cm -1 The resolution is 4 cm -1 , and the background spectrum is collected every 1 h.
[0016] Spectral data set: A total of 465 spectra (155 varieties × 3 repeated measurements).
[0017] 2. Spectral preprocessing and model construction
[0018] Use R2023a (The MathWorks, USA) was used to construct support vector machine (SVM) and partial least squares (PLS) models. Before building the models, 465 spectral data were preprocessed by smoothing (Savitzky-Golay). The calibration set and validation set were divided according to the random principle in a ratio of 3:1. For the calibration set of the model: 155 varieties of rice, including 349 mid-infrared spectra. For the validation set of the model: 155 varieties of rice, including 116 mid-infrared spectra. In addition, PLS and SVM regression models were constructed based on the results of feature band screening by successive projections algorithm (SPA), competitive adaptive reweighted sampling algorithm (CARS), uninformative variable elimination (UVE), and interval combination optimization (ICO). The prediction performance of the model was evaluated by the root mean square error of the validation set (RMSEV), the prediction accuracy of the calibration set (R c 2 ) and the coefficient of determination of the prediction accuracy of the validation set (R v 2 ).
[0019] 3. Analysis of the changes in the physicochemical quality characteristics of brown rice
[0020] As Figure 2 shown, there were no significant differences in the appearance of brown rice of all varieties, indicating that the appearance characteristics could not intuitively reflect the differences in protein content among varieties. However, there were obvious changes in the protein content in brown rice of different varieties (see Figure 3 ). According to the data in Table 1, the total protein content of brown rice ranged from 6.55% to 11.99%, indicating that the differences in protein content among these varieties were significant. Further analysis found that the distribution of these protein contents generally showed a normal distribution trend (see Figure 3 ), indicating that the protein content of most varieties was concentrated around the median value, and extremely low or high protein contents were relatively rare.
[0021] Table 1 Changes in the protein content of brown rice of different varieties
[0022]
[0023] 4. Comparative analysis of the results of predicting the quality characteristics of brown rice based on machine learning combined with feature band selection algorithms
[0024] The results of selecting mid-infrared spectral data based on different feature band selection algorithms are as Figure 4As shown, Figure a shows that as the number of variables increases in the SPA algorithm, the root mean square error of the model gradually decreases until the model error no longer decreases when the number of variables reaches 10. As shown in Figure b, a total of 10 characteristic wavelengths are selected. Figure c shows the process of feature selection by the CARS algorithm. The root mean square error of cross-validation first decreases and then increases with the increase of the number of sampling times and reaches the lowest value at about 450 times. Finally, as shown in Figure d, 119 characteristic bands are selected.
[0025] The results of selecting mid-infrared spectral data based on the UVE and ICO algorithms are as Figure 5 shown. The results of UVE are shown in Figures a and b. It can be seen from Figure a that the red part is random noise and the blue part is the actual variable. By comparing the t-values of the actual variable and the random variable, it is determined which variables make important contributions to the model. It can be seen from the figure that there is an obvious distinction between the actual variable (blue line) and the random variable (red line), indicating that there are many significant contribution bands for the actual variable in this figure. Finally, 976 characteristic bands are selected. The results of ICO are shown in Figures c and d. By the 5th iteration, the root mean square error of the cross-validation set begins to rise and is close to the number of variables finally retained. Finally, 1081 characteristic bands are selected.
[0026] The results of predicting the protein content of brown rice based on the combination of PLS and SVM algorithms and the feature band selection algorithm are shown in Table 2. R c 2 represents the prediction accuracy of the model for the calibration set. R v 2 represents the prediction accuracy of the model for the validation set. RMSEV represents the average prediction deviation of the model for the validation set, and RMSEC represents the average prediction deviation of the model for the calibration set. Among them, the larger R c 2 and R v 2 , and the smaller RMSE, the better the prediction effect of the model. In Table 2, the effect of predicting the protein content of brown rice based on the SVM model is significantly better than that based on PLS. The CARS-SVM model constructed based on the near-infrared (NIR) band can also better predict the protein content of brown rice (R v 2 = 0.7849, RMSEV = 0.5940) and only uses 119 characteristic bands, which greatly improves the detection efficiency. In addition, among all the models for predicting the protein content of brown rice, the model based on CARS-SVM performs the best. Specifically, the determination coefficient of this model on the validation set (R v 2The coefficient of determination ($R^2$) reached 0.9437, and the root mean square error of validation (RMSEV) was 0.8173, indicating extremely high prediction accuracy. It is worth noting that the model only used 119 characteristic bands, which not only significantly reduced the computational cost but also greatly improved the detection efficiency. Based on this, the superior performance of the CARS-SVM model in protein content prediction demonstrates its great potential in commercial applications. Through further optimization, this technology is expected to achieve fast, efficient, and low-cost quality control in actual detection, promoting the wide application of Fourier transform infrared spectroscopy technology in the field of grain quality detection.
[0027] Table 2 Prediction of brown rice protein content based on PLS and SVM algorithms combined with characteristic band selection algorithms
[0028]
[0029] In summary, the prediction results of brown rice protein content based on machine learning technology combined with characteristic band selection algorithms show that by selecting a small number of key bands for modeling, not only the complexity of the model is greatly simplified, but also the prediction accuracy is significantly improved. This result indicates that by adopting an optimized characteristic band screening strategy, the protein content of brown rice can be accurately and quickly predicted. This study provides important support for the application of Fourier transform infrared spectroscopy technology in the rapid and non-destructive detection of brown rice quality, demonstrates its great potential and broad application prospects in actual detection, and helps to improve the detection efficiency in grain production and quality control.
[0030] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by using equivalent replacements or equivalent transformations fall within the protection scope of the present invention.
Claims
1. A rapid detection method for brown rice protein content based on Fourier transform infrared spectroscopy, comprising the following steps: S1. 155 varieties of brown rice were sown in mid-to-late May. Mature seeds were harvested in batches from mid-September to early October, 40 days after heading. Varieties with similar maturity were selected, and the collected seeds were naturally air-dried and placed indoors for one month of post-ripening. They were then husked and ground to obtain brown rice and polished rice. The rice grains were ground into fine powder using a vibrating ball mill and passed through a 100-mesh sieve for later use. S2. During the Fourier transform attenuated total reflection infrared spectroscopy (ATR-FTIR) acquisition, a small amount of brown rice powder was placed on the surface of the ATR crystal window of the sample tray, and the pressure tower was rotated to press the powder downward to make it in good contact with the ATR crystal. The scanning number was 32 times, and the wave number scanning range was 525-4000cm -1 Resolution 4cm -1 , and background spectra were collected every 1 h. S3. Using a rapid N Exceed nitrogen analyzer (elementar, Germany), 250.0 mg of brown rice flour and aspartic acid (Asp) standard sample were weighed using a 1 / 10,000 electronic balance, and the samples were wrapped in tin foil and compacted. Turn on the nitrogen analyzer and wait for the self-test to finish heating. Then open the CO2 and O2 valves and start loading samples. First place three blank samples (Blank), set the Method to BlankwithO2; then place 5 standard samples, set the Method to 250mg Standard, and set the protein factor to 6.
25. Place the brown rice samples of each variety in turn. After loading, the instrument automatically analyzes. After the run is completed, read and use GraphPadPrism to analyze the data. S4. Construct support vector machine (SVM) and partial least squares (PLS) models. Before building the model, the 465 spectral data were smoothed (Savitzky-Golay) preprocessed. The calibration set and validation set were divided into a 3:1 ratio using the random principle. For the calibration set of the model: 155 varieties of rice, including 349 mid-infrared spectra. For the validation set of the model: 155 varieties of rice, including 116 mid-infrared spectra. The root mean square error (RMSEV) of the validation set and the prediction accuracy (R c 2 ) and the prediction accuracy of the validation set (R v 2 ) was used to evaluate the prediction performance of the model. S5. In addition, PLS and SVM regression models are constructed based on the results of continuous projection algorithm (SPA), competitive adaptive reweighted sampling algorithm (CARS), uninformative variable elimination (UVE), iterative information retention algorithm (IRIV), and interval combination optimization (ICO) feature band screening. The root mean square error (RMSEV) of the validation set and the prediction accuracy (R c 2 ) and the prediction accuracy of the validation set (R v 2 ) was used to evaluate the prediction performance of the model.