A method for quantitatively detecting adulteration of wheat flour with barley flour based on near infrared spectroscopy
Patent Information
- Application Number
- CN202610683931.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-05-18
AI Technical Summary
[0004]一方面,小麦粉中大麦粉掺杂在低浓度(如5%)时,大麦粉的特征光谱信号极其微弱,极易被小麦基质的强背景光谱所淹没,导致传统建模方法对低掺杂比例的检测灵敏度不足、定量误差较大;另一方面,样品研磨后的物理状态差异,如水分活度不均、静电积累等,会引入与化学掺杂无关的光谱变异,进一步干扰微量掺杂成分的准确识别,影响模型的稳健性与泛化能力
[0030] 1. First, the method of this invention can effectively improve the detection sensitivity and quantitative accuracy of low-concentration doping. Specifically, this invention systematically solves the problem of the strong background spectrum of wheat matrix masking the signal of low-concentration barley flour doping by a step-by-step focusing strategy of "differential spectral calculation - multidimensional scaling and dimensionality reduction - SE attention adaptive enhancement". Differential operation is performed with the average spectrum of pure wheat flour as a reference to systematically subtract the absorption background of common wheat components, so that the differential spectrum reflects the net spectral changes introduced by doping. Multidimensional scaling is used to reduce the dimensionality with the pairwise distance between samples as the optimization target, so that samples with different doping ratios present a structured distribution consistent with the concentration gradient in the low-dimensional feature space. On this basis, the SE attention mechanism adaptively amplifies the weights of key feature peaks directly related to barley flour components, such as 1438nm, 1880nm, and 1979nm, to suppress residual noise and redundant band interference. Finally, it achieves high-precision quantitative detection performance with a relative error of only 3.8%, an independent prediction set determination coefficient of 0.984, a prediction root mean square error as low as 0.71%, and a relative analysis error of 4.48 at a doping ratio as low as 5%.
Smart Images

Figure CN122217908B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of doping detection technology, specifically to a method for quantitative detection of barley flour doping in wheat flour based on near-infrared spectroscopy. Background Technology
[0002] Near-infrared spectroscopy has been widely used in grain quality evaluation and adulteration detection due to its advantages such as speed, non-destructive nature, and simultaneous multi-component analysis. Current techniques typically employ single or combined methods such as multivariate scattering correction, standard normal variable transformation, Savitzky-Golay smoothing, and derivative processing to preprocess near-infrared spectra. This is then combined with traditional chemometric methods such as principal component analysis and partial least squares discriminant analysis, or by introducing deep learning models such as one-dimensional convolutional neural networks, to classify and quantify adulterated samples.
[0003] However, the aforementioned existing technologies still have the following shortcomings in practical applications:
[0004] On the one hand, when barley flour is mixed into wheat flour at a low concentration (e.g., 5%), the characteristic spectral signal of barley flour is extremely weak and easily drowned out by the strong background spectrum of the wheat matrix. This results in insufficient sensitivity and large quantitative error in traditional modeling methods for detecting low doping ratios. On the other hand, differences in the physical state of the sample after grinding, such as uneven water activity and electrostatic accumulation, can introduce spectral variations unrelated to chemical doping, further interfering with the accurate identification of trace doping components and affecting the robustness and generalization ability of the model. Summary of the Invention
[0005] To address the problems in related technologies, this invention provides a method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy, comprising the following steps:
[0008] Step S1: Mix pure wheat flour matrix of different varieties with pure barley flour of different varieties according to multiple preset mass fraction gradients to prepare doped samples. At the same time, set up pure wheat flour control samples to obtain the average spectrum of pure wheat flour, and divide the doped samples into a calibration set and an independent prediction set.
[0009] Step S2: Collect near-infrared spectral data of the doped sample and the pure wheat flour control sample;
[0010] Step S3: The acquired near-infrared spectral data is processed using a combined preprocessing scheme, which includes multivariate scattering correction, Savitzky-Golay smoothing filtering, and first derivative processing to obtain a preprocessed spectrum with a flat baseline and clear characteristic peaks.
[0011] Step S4: Using the average spectrum of the pure wheat flour control sample as a reference, calculate the difference spectrum between the preprocessed spectrum of each doped sample and the reference; use a multidimensional scaling method to reduce the dimensionality of the difference spectrum, and extract the samples whose cumulative contribution rate meets a preset threshold. Each principal component is used as an eigenvector to focus on the spectral differences introduced by doping and to preserve the pairwise distance structure between data.
[0012] Step S5: Input the extracted feature vector into the constructed SE attention-enhanced one-dimensional convolutional neural network model to output the doping ratio of barley flour in the sample to be tested.
[0013] Optionally, in step S1, after all the original samples are crushed and sieved and before spectral acquisition, they are allowed to stand for at least 2 hours in an environment with a temperature of 20±2℃ and a relative humidity of 45%-55% to eliminate spectral variations unrelated to chemical doping caused by differences in water activity and electrostatic state between samples, so as to ensure that the weak signals detected later truly reflect the changes in composition.
[0014] Optionally, the multiple preset mass fraction gradients include 5%, 20%, and 50%, and at least one set of parallel samples is set for each doping ratio; the sample preparation also includes setting a pure barley flour control sample and a non-barley grain interference sample.
[0015] Optionally, in step S2, the environmental conditions for acquiring the spectrum include: a spectral scanning range of 800-2500 nm and a scanning resolution of [missing information]. The number of scans was 32-64 per sample, the ambient temperature was 20±2℃, and the relative humidity was 45%-55%; the spectrum of each sample was collected at least 3 times and the average spectrum was used as the raw data.
[0016] Optionally, in step S4, the dimensionality reduction of the difference spectrum using a multidimensional scaling method specifically includes:
[0017] Calculate the difference spectrum for every two samples and The original distance between The calculation formula is as follows:
[0018]
[0019] In the formula, and The first The first sample and the first Differential spectral data of each sample, , and These are the L2 norms of the corresponding subscript samples within the differential spectrum;
[0020] Based on the original distance, an inner product distance matrix is constructed through a bi-centering operation. And solve the inner product distance matrix. eigenvalues of the covariance matrix and eigenvectors ;
[0021] Retain those with a cumulative contribution rate ≥90% The eigenvectors corresponding to each eigenvalue are used as the extracted principal components to obtain the dimensionality-reduced differential spectral features. .
[0022] Optionally, the elements of the inner product distance matrix S The calculation formula is:
[0023]
[0024] In the formula, Let be the average of the squared near-infrared spectral distances between the i-th sample and all other samples. Let be the average of the squared near-infrared spectral distances between the j-th sample and all other samples. It is the total average of the squared near-infrared spectral distances between all pairs of samples.
[0025] Optionally, in step S3, the window size of the Savitzky-Golay smoothing filter is set to 9-13 points, and the polynomial order is 2-3. The combined preprocessing scheme uses signal-to-noise ratio, characteristic peak identification, and 1979nm peak intensity variation coefficient as evaluation indicators, and selects the optimal combination from a variety of combination schemes including MSC+SG, SNV+SG, MSC+SG+FD, and SNV+SG+SD. Here, MSC is multivariate scattering correction, SG is Savitzky-Golay smoothing filter, FD is first derivative, SNV is standard normal variable transformation, and SD is second derivative.
[0026] Optionally, the SE attention-enhanced one-dimensional convolutional neural network model uses a basic one-dimensional convolutional neural network as its backbone and introduces an SE attention mechanism module with a compression ratio of 16 after the convolutional layer. The SE attention mechanism module generates channel weights through global average pooling, fully connected layer compression and recovery, and Sigmoid activation function to adaptively enhance the weights of key feature peaks at 1438nm, 1880nm, and 1979nm, and suppress spectral overlap interference and noise.
[0027] Optionally, during the training process of the SE attention-enhanced one-dimensional convolutional neural network model, the learning rate search range is set to 0.001-0.01, the number of iterations is 80-120, and the batch size is uniformly 16-64. The training adopts ten-fold cross-validation, and the calibration set is randomly divided into 10 subsets. Multiple rounds of training are completed by taking turns using 9 subsets as training samples and 1 subset as validation samples to avoid data bias.
[0028] Optionally, the non-barley grain interference samples include buckwheat flour, oat flour, and wheat bran; after step S5, the method further includes: processing the near-infrared spectrum of the non-barley grain interference samples through steps S2 to S5, and verifying whether the predicted doping ratio of the interference samples by the SE attention-enhanced one-dimensional convolutional neural network model is lower than a preset threshold, so as to determine the model's specific identification ability of barley flour doping in wheat flour.
[0029] Beneficial effects:
[0030] 1. First, the method of this invention can effectively improve the detection sensitivity and quantitative accuracy of low-concentration doping. Specifically, this invention systematically solves the problem of the strong background spectrum of wheat matrix masking the signal of low-concentration barley flour doping by a step-by-step focusing strategy of "differential spectral calculation - multidimensional scaling and dimensionality reduction - SE attention adaptive enhancement". Differential operation is performed with the average spectrum of pure wheat flour as a reference to systematically subtract the absorption background of common wheat components, so that the differential spectrum reflects the net spectral changes introduced by doping. Multidimensional scaling is used to reduce the dimensionality with the pairwise distance between samples as the optimization target, so that samples with different doping ratios present a structured distribution consistent with the concentration gradient in the low-dimensional feature space. On this basis, the SE attention mechanism adaptively amplifies the weights of key feature peaks directly related to barley flour components, such as 1438nm, 1880nm, and 1979nm, to suppress residual noise and redundant band interference. Finally, it achieves high-precision quantitative detection performance with a relative error of only 3.8%, an independent prediction set determination coefficient of 0.984, a prediction root mean square error as low as 0.71%, and a relative analysis error of 4.48 at a doping ratio as low as 5%.
[0031] Secondly, the method of this invention can effectively eliminate spectral interference introduced by differences in the physical state of the samples. Specifically, during the sample preparation stage, the original sample is placed in an environment with a temperature of 20±2℃ and a relative humidity of 45%-55% and allowed to stand for at least 2 hours to equilibrate. This can suppress non-chemical spectral variations in powder samples caused by differences in water activity and electrostatic accumulation from the source. This pretreatment operation makes the intensity of the OH absorption peak of water more uniform and reduces the influence of differences in particle packing density caused by electrostatics on the light scattering path, ensuring that the weak signals retained in the subsequent differential spectra truly reflect the changes in chemical composition caused by barley flour doping, rather than physical artifacts accidentally introduced during the sample preparation process.
[0032] Third, the method of this invention can effectively enhance the model's generalization ability and variety adaptability. Specifically, in the sample preparation stage, multiple pure wheat flour matrices of different varieties are cross-paired and doped with multiple pure barley flour matrices of different varieties. This ensures that the calibration set has broad representativeness in terms of variety, avoiding overfitting caused by the model only learning spectral differences between specific varieties. The multidimensional scaling method preserves the pairwise distance structure between samples driven by the doping ratio, rather than relying on the absolute spectral characteristics of specific varieties. This ensures that the extracted feature vectors have consistent expressive power for wheat and barley samples from different origins and varieties. External validation showed that the model achieved a prediction coefficient of determination of 0.988, a root mean square error of 0.76%, and a relative analysis error of 7.21 for Canadian wheat flour mixed with French barley flour, which is highly consistent with the internal validation results, indicating that this scheme has good cross-variety and cross-origin generalization ability.
[0033] Fourth, the method of this invention can effectively ensure the high specificity of the detection results. Specifically, by pre-setting non-barley grain interference samples such as buckwheat flour, oat flour, and wheat bran during the sample preparation stage, and verifying that the predicted doping ratio of all interference samples by the model is less than 0.8%, a clear quantification threshold is used to prove that the model only produces a quantitative response to barley flour doping, while maintaining effective rejection ability for other closely related grains or homologous processing by-products. The SE attention mechanism adaptively enhances the unique spectral markers of barley flour at 1438nm, 1880nm, and 1979nm that distinguish barley flour from other grains. From a mechanistic perspective, this ensures that the model's specificity does not depend on a broad range of common grain responses, but rather focuses on the specific chemical composition information of barley flour, effectively avoiding false positives in actual detection.
[0034] Fifth, the method of this invention can effectively ensure the quality of spectral data and the stability of model training. Specifically, in the spectral preprocessing stage, the MSC+SG+FD combined preprocessing scheme is objectively selected from multiple candidate schemes by establishing three quantitative indicators: signal-to-noise ratio, characteristic peak identification, and coefficient of variation of 1979 nm peak intensity. This ensures that the preprocessing achieves the optimal balance in the three dimensions of noise suppression, feature preservation, and signal stability. In the spectral acquisition stage, parameters such as scanning range, resolution, number of scans, and ambient temperature and humidity are specifically limited to ensure the comparability and signal-to-noise ratio of the original spectral data. In the model training stage, the stability and reproducibility of the training process are ensured by limiting the learning rate range, number of iterations, batch size, and ten-fold cross-validation strategy. This effectively avoids data partitioning bias and overfitting risks, making the model performance evaluation results true and reliable.
[0035] 2. Other beneficial effects or advantages of the present invention will be described in detail in the specific embodiments. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] in:
[0038] Figure 1 This is a schematic flowchart of a method for quantitative detection of barley flour doping in wheat flour based on near-infrared spectroscopy, provided by an exemplary embodiment of the present invention.
[0039] Figure 2 This is an exemplary embodiment of the present invention providing PCA extraction from imported wheat flour, domestic wheat flour, imported barley flour, and domestic barley flour. Figure 2 A schematic diagram of the PLS-DA model analysis for b, where... Figure 2 a is a graph showing the PCA extraction of imported wheat flour, domestic wheat flour, imported barley flour, and domestic barley flour. Figure 2 b is a PLS-DA model analysis diagram of imported wheat flour, domestic wheat flour, imported barley flour, and domestic barley flour;
[0040] Figure 3 This is a schematic diagram of near-infrared spectroscopy using a combination of MSC+SG+FD for spectral preprocessing, provided by an exemplary embodiment of the present invention.
[0041] Figure 4This is a near-infrared spectrum of pure wheat flour using a combined MSC+SG+FD pretreatment scheme provided in an exemplary embodiment of the present invention;
[0042] Figure 5 Image a is a near-infrared spectral data image of a wheat sample adulterated with barley, obtained by using a near-infrared spectroscopy instrument, taking Zhengmai 1860 wheat adulterated with Yangnongpi 13 as an example, according to an exemplary embodiment of the present invention. Figure 5 b is a near-infrared spectral data image of a wheat sample doped with barley, taken as an example of Zhengmai 1860 doped with Minmai 2, provided by an exemplary embodiment of the present invention;
[0043] Figure 6 This is an exemplary embodiment of the present invention, using near-infrared spectroscopy to detect spectral data of a wheat sample adulterated with barley, taking Zhengmai 1860 (doped with Minmai 2) as an example; wherein, Figure 6 a is the original near-infrared absorbance spectrum. Figure 6 b is the near-infrared spectrum after preprocessing with the MSC+SG+FD combination. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0045] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0046] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. It should also be noted that in embodiments of this invention, the words "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in embodiments of this invention should not be construed as preferred or advantageous over other embodiments or designs. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0047] To facilitate a clearer and more accurate understanding of the technical solutions of this invention by those skilled in the art, the existing related technologies and their technical problems will be described in more detail below.
[0048] Near-infrared spectroscopy is a non-destructive testing method that utilizes the absorption signals of overtone and combination vibrations generated by hydrogen-containing groups (such as OH, NH, CH, etc.) in the near-infrared region to perform qualitative or quantitative analysis of the chemical composition of samples. In the field of grain quality testing, this technology has been widely used to determine the content of major components such as moisture, protein, and starch, and in recent years it has also been gradually expanded to scenarios such as variety identification and adulteration screening.
[0049] In practical detection processes, the preprocessing and modeling methods of spectral data have a decisive impact on the final analytical accuracy. Since near-infrared raw spectra inevitably carry various physical interference signals, existing techniques generally employ multivariate scattering correction or standard normal variable transformation to eliminate scattering effects caused by sample particle size differences. This is supplemented by Savitzky-Golay smoothing filtering to suppress random noise, and first- or second-derivative processing to enhance the resolution of spectral characteristic peaks and eliminate baseline drift. Based on this, modeling methods are generally divided into two categories: one is traditional chemometric methods, represented by partial least squares discriminant analysis, principal component analysis combined with Mahalanobis distance, whose advantages lie in strong model interpretability and low computational resource requirements; the other is deep learning methods, represented by one-dimensional convolutional neural networks, residual networks, and long short-term memory networks, whose advantages lie in their ability to automatically mine deep nonlinear features in spectral data and their greater potential for modeling complex mixed systems. Some studies have attempted to introduce attention mechanisms into convolutional neural networks, improving the model's ability to extract key spectral information by adaptively adjusting the feature channel weights.
[0050] In the feature dimensionality reduction stage, principal component analysis is the most commonly used unsupervised feature extraction method. By performing orthogonal transformation on the original high-dimensional spectrum, the top principal components with high cumulative variance contribution rates are extracted, thereby achieving data compression and redundancy removal.
[0051] Although the above-mentioned technical approaches have made some progress, when near-infrared spectroscopy is applied to the quantitative detection of barley flour adulteration in wheat flour, it still faces several technical bottlenecks caused by the characteristics of the detection object itself, which are difficult to be effectively solved by existing methods.
[0052] The primary problem lies in signal submersion and insufficient detection sensitivity under low-concentration doping conditions. Wheat flour matrix itself contains various chemical components with strong near-infrared absorption, and its spectral absorption signals constitute a high-intensity matrix background. When the doping ratio of barley flour is low (e.g., near the 5% mass fraction threshold, which is of significant practical importance in identifying the authenticity of imported and exported wheat), the characteristic spectral changes introduced by barley flour are extremely weak, and their signal intensity is often on the same order of magnitude as baseline noise, easily masked by the strong background spectrum of wheat. While existing preprocessing methods can improve the signal-to-noise ratio to some extent, they lack an effective targeted mechanism for this essential signal aliasing caused by "strong background masking weak differences." Traditional full-spectrum PCA dimensionality reduction methods focus on preserving the principal component directions with the largest overall variance in the data. However, these principal components are often dominated by the variation of the main components of the wheat matrix and cannot effectively focus on the weak spectral differences introduced by low-concentration doping. Conventional partial least squares discriminant analysis and basic deep learning models, without the enhancement of differential features, also struggle to continuously capture these weak but crucial doping indicator signals from the global spectrum, resulting in large quantitative errors for doping ratios of 5% and below.
[0053] Secondly, the spectral variations introduced by differences in the physical state of the samples constitute another hidden source of interference affecting detection accuracy. During sample preparation, after the grain particles are crushed and sieved, different varieties and batches of samples exhibit differences in powder particle size distribution, surface electrostatic accumulation, and water activity balance. Although these subtle differences in physical state are unrelated to changes in chemical composition, they can affect the scattering path and absorption efficiency of light in the powder sample, manifesting as baseline drift or overall absorbance fluctuations in the spectrum. These signal characteristics may be confused with the spectral response of low-concentration dopant, further increasing the difficulty of accurately identifying trace dopant components. Current technologies typically have a rather rudimentary control over the sample preparation environment, and have not yet explicitly incorporated pretreatment conditions such as temperature and humidity balance and electrostatic elimination into a systematic consideration of their impact on detection accuracy.
[0054] The combined existence of the above problems makes it difficult for existing near-infrared spectroscopy-based detection schemes to simultaneously meet the practical detection requirements of high sensitivity, high specificity, and good robustness when dealing with low proportions of barley flour adulteration in wheat flour. This limits the application effectiveness of this technology in the rapid screening of the authenticity of imported and exported wheat.
[0055] In view of this, the present invention provides a novel solution, namely, a method for quantitative detection of barley flour doping in wheat flour based on near-infrared spectroscopy. The technical concept of the present invention is to construct a step-by-step enhanced detection link of "physical state homogenization preprocessing - difference spectral focusing enhancement - attention mechanism adaptive mining", which systematically solves the bottleneck that the low concentration doping signal of barley flour in wheat flour is weak and easily masked. Specifically, firstly, by allowing the samples to fully settle and equilibrate under specific temperature and humidity conditions, physical spectral variations unrelated to chemical doping caused by differences in water activity and electrostatic state between powder samples are eliminated, ensuring the authenticity of weak signals from the source. Secondly, using the average spectrum of pure wheat flour as a reference, the differential spectra of each doped sample are calculated. Differential operations are used to remove the strong background spectrum of the wheat matrix, forcing subsequent multidimensional scaling and dimensionality reduction to focus on the net spectral differences introduced by the doping components, achieving selective enhancement of low-concentration doping signals during feature extraction. Finally, the SE attention mechanism is used to adaptively increase the weights of key feature peaks directly related to barley flour components, such as 1438nm, 1880nm, and 1979nm, suppressing residual noise and redundant band interference. This allows the model to accurately capture the spectral expression of doping ratios from the focused weak differences, significantly improving the quantitative detection accuracy and model robustness under low doping ratios.
[0056] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0057] like Figure 1 As shown in the figure, this embodiment provides a method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy, including the following steps:
[0058] Step S1: Mix pure wheat flour matrix of different varieties with pure barley flour of different varieties according to multiple preset mass fraction gradients to prepare doped samples. At the same time, set up pure wheat flour control samples to obtain the average spectrum of pure wheat flour, and divide the doped samples into a calibration set and an independent prediction set.
[0059] Step S2: Collect near-infrared spectral data of the doped sample and the pure wheat flour control sample;
[0060] Step S3: The acquired near-infrared spectral data is processed using a combined preprocessing scheme, which includes multivariate scattering correction, Savitzky-Golay smoothing filtering, and first derivative processing to obtain a preprocessed spectrum with a flat baseline and clear characteristic peaks.
[0061] Step S4: Using the average spectrum of the pure wheat flour control sample as a reference, calculate the difference spectrum between the preprocessed spectrum of each doped sample and the reference; use a multidimensional scaling method to reduce the dimensionality of the difference spectrum, and extract the samples whose cumulative contribution rate meets the preset threshold. Each principal component is used as an eigenvector to focus on the spectral differences introduced by doping and to preserve the pairwise distance structure between data.
[0062] Step S5: Input the extracted feature vector into the constructed SE attention-enhanced one-dimensional convolutional neural network model, and output the doping ratio of barley flour in the sample to be tested.
[0063] Through the above technical solution, firstly, this invention introduces multiple varieties of wheat flour and barley flour for cross-pairing, ensuring that the adulterated samples are broadly representative in terms of variety. This avoids overfitting caused by the model only learning the spectral differences between specific varieties, thus guaranteeing the generalization ability of the subsequently built model for wheat and barley samples from different origins and varieties from the outset. Simultaneously, pre-setting multiple mass fraction gradients allows the model to fully learn the continuous spectral response pattern from low to high adulteration ratios during the training phase, establishing a complete concentration-spectral correspondence for quantitative prediction. Furthermore, setting up a pure wheat flour control sample and obtaining its average spectrum provides a unified reference benchmark for subsequent differential spectral calculations, allowing the spectral changes of different adulterated samples to be compared under a common reference system, laying a data foundation for the targeted removal of adulteration information.
[0064] Second, the spectra of the adulterated sample and the pure wheat flour control sample were acquired simultaneously to ensure that the spectral data were obtained under the same instrument conditions, environmental conditions and acquisition parameters. This eliminated the systematic bias that may be introduced due to inconsistent acquisition conditions and helped to ensure the accuracy and comparability of subsequent differential spectral calculations.
[0065] Third, the combined preprocessing scheme used in step S3 of this invention, through the synergistic cooperation of three types of processing, can systematically eliminate various physical interference signals superimposed in the near-infrared raw spectrum. Specifically, multivariate scattering correction compensates for the differences in light scattering paths caused by the uneven particle size of the sample powder, restoring a linear additive relationship between spectral absorbance and chemical composition content; Savitzky-Golay smoothing filtering suppresses high-frequency random noise while preserving the characteristic peak shape of the spectrum, improving the spectral signal-to-noise ratio; and first-derivative processing effectively eliminates baseline tilt caused by instrument baseline drift or differences in overall sample concentration, while sharpening overlapping characteristic peaks, allowing components with similar chemical compositions but different structures to exhibit identifiable features in the spectrum. Thus, the combined use of these three methods results in a flat spectral baseline and clear characteristic peak outlines after preprocessing, providing higher-quality data input for accurately extracting subtle chemical absorption changes related to barley flour doping from the spectrum.
[0066] Fourth, the introduction of differential spectroscopy is the core method that distinguishes step S4 from existing conventional feature extraction. Specifically, by performing differential operations with the average spectrum of pure wheat flour as a reference, the strong absorption background of common components such as moisture, protein, and starch in the wheat matrix can be systematically subtracted, so that only the net spectral change signal caused by barley flour doping is retained in the differential spectrum. This allows for the targeted enhancement of doping information before dimensionality reduction, effectively alleviating the problem of detecting weak barley doping signals masked by strong wheat background signals.
[0067] Building upon this, employing a multidimensional scaling method instead of traditional principal component analysis to reduce the dimensionality of differential spectra yields two benefits. First, multidimensional scaling directly optimizes the pairwise spectral distances between samples, preserving as much as possible the relative proximity of doped samples in the original high-dimensional space due to different doping ratios. This distance-preserving mapping characteristic allows samples with different doping ratios to exhibit a structured distribution consistent with the concentration gradient in the feature space, exhibiting higher sensitivity to minute spectral shifts caused by low-concentration doping. Second, selecting principal components based on cumulative contribution rate thresholds effectively compresses the spectral dimension while retaining key doping information, removing redundant variables and reducing the input complexity and overfitting risk of subsequent deep learning models.
[0068] Fifth, the feature vector output in step S4, used as input to the model in step S5, already possesses the advantages of matrix background stripping and doping distance structure preservation. The SE attention-enhanced one-dimensional convolutional neural network model further leverages its deep feature mining and adaptive focusing capabilities. The one-dimensional convolutional neural network, by sliding the convolution kernel along the spectral dimension, can capture the combination patterns and nonlinear response relationships between local bands in the feature vector, compensating for the limitations of multidimensional scaling as a linear dimensionality reduction method in representing complex spectral interaction effects. Based on this, the SE attention mechanism, through global average pooling, fully connected layer compression and recovery, and the sigmoid activation function, automatically learns the importance weights of each feature channel. This allows the model to adaptively amplify the response of feature channels highly correlated with the doping ratio during forward inference, while suppressing interference from residual channels unrelated to doping. Thus, even when facing low doping ratios, if the doping signal only shows a slight amplitude difference in the feature vector, the model can effectively extract it through attention weight allocation and use it for quantitative regression, ultimately outputting an accurate barley flour doping ratio.
[0069] In one embodiment of the present invention, in step S1, after all the original samples are crushed and sieved and before spectral acquisition, they can be allowed to stand and equilibrate for at least 2 hours in an environment with a temperature of 20±2℃ and a relative humidity of 45%-55% to eliminate spectral variations unrelated to chemical doping caused by differences in water activity and electrostatic state between samples, and to ensure that the weak signals detected subsequently truly reflect the changes in composition.
[0070] In this embodiment, firstly, after the grain sample is crushed and sieved, the specific surface area of the powder increases dramatically, and the rate of moisture exchange with the surrounding air accelerates significantly. During the process of reaching moisture equilibrium with the environment, samples of different varieties and with different initial moisture contents will experience a shift in the intensity of the moisture absorption peak at approximately 1438 nm, which belongs to the first overtone of the OH stretching vibration, due to differences in water activity. If the samples are not allowed to fully equilibrate before spectral acquisition, the difference in moisture peak intensity will be superimposed on the spectral changes caused by doping, resulting in non-chemical fluctuations unrelated to doping in the difference spectrum calculated in subsequent step S4. This embodiment limits the humidity of the settling environment to 45%-55% and the temperature to 20±2℃. This ensures that all samples achieve moisture equilibrium under uniform and controllable temperature and humidity conditions, guaranteeing that the water activity of different samples tends to be consistent, thereby eliminating spectral variations introduced by fluctuations in the moisture absorption peak intensity. This ensures that the signal changes retained in the difference spectrum truly belong to the compositional differences between barley flour and wheat flour, rather than accidental deviations in the moisture state between samples.
[0071] Secondly, the mechanical friction generated during the pulverization process causes varying degrees of electrostatic charge to accumulate on the surface of powder particles. The presence of static electricity leads to repulsion or aggregation effects between powder particles, altering the packing density and particle stacking pattern after the sample is placed in the sample cell, thereby affecting the scattering path length of near-infrared light in the powder layer. Since the scattering effect in near-infrared spectroscopy manifests as baseline tilt and overall absorbance drift, differences in electrostatic states introduce spectral background fluctuations unrelated to chemical composition. For the detection of low-concentration doping, the characteristic signal intensity of the dopant component is extremely weak, and baseline fluctuations caused by static electricity are sufficient to mask or distort it. In this implementation method, the sample is equilibrated in a static environment for at least 2 hours. This allows some of the accumulated static charge to be naturally released through surface conductive pathways under controlled humidity (45%-55%) conditions, making the electrostatic state of each sample powder more consistent. This effectively reduces spectral fluctuations caused by differences in packing density and inconsistent scattering paths, ensuring that the spectral data acquired in step S2 stably reflects the chemical composition of the sample.
[0072] In one embodiment of the present invention, multiple preset mass fraction gradients may include 5%, 20% and 50%, and at least one set of parallel samples is set for each doping ratio; sample preparation also includes setting pure barley flour control samples and non-barley grain interference samples.
[0073] In this implementation, firstly, the doping mass fraction gradients are set to 5%, 20%, and 50%, with parallel samples for each gradient. 5% corresponds to a low-concentration doping range, a key threshold with practical regulatory significance in the authentication of imported and exported wheat; 20% corresponds to a medium doping ratio, forming a transition range from trace to large doping levels; and 50% approximates an equal mixture of pure barley flour and wheat flour, providing a high-concentration reference for the full expression of barley flour spectral characteristics in the mixture. These three gradients spanning the complete low, medium, and high concentration ranges allow the model to learn the continuous evolution of the spectral response as the doping ratio changes from trace to large levels during training, establishing a complete concentration-spectrum mapping curve for quantitative prediction. Simultaneously, setting parallel samples for each doping ratio provides repeatable observations at the same concentration point, effectively assessing sample preparation consistency and spectral acquisition repeatability, providing a stable statistical basis for subsequent model training, and avoiding systematic deviations in model predictions for specific concentration points due to accidental biases of a single sample.
[0074] Secondly, the inclusion of pure barley flour as a control sample allows for the independent acquisition of the intrinsic spectral characteristics of barley flour without interference from the wheat matrix. In subsequent analyses, the pure barley flour spectrum can serve as a reference for locating characteristic peaks of barley flour components, helping to confirm which bands in the differential spectra truly contribute to the changes from barley flour, rather than being affected by other factors. Simultaneously, the pure barley flour control also provides an endpoint reference for analyzing the trend of spectral characteristic peaks in doped samples as a function of doping ratio, enabling the model to more accurately establish a quantitative relationship between the spectral characteristics of barley flour and the doping ratio.
[0075] Third, in practical grain adulteration detection scenarios, impurities that may be mixed into wheat flour are not limited to barley flour; buckwheat flour, oat flour, wheat bran, and other grains may also be present. These interfering substances each have their own specific absorption characteristics in the near-infrared spectrum. This implementation method pre-includes non-barley grain interfering samples during the sample preparation stage, providing necessary test data for subsequent verification of the model's specific ability to identify barley flour adulteration. If the model has not been exposed to interfering samples and only learns the differences between barley flour and wheat flour, it may misinterpret the spectral differences of unknown interfering substances as barley flour adulteration signals, leading to false positive results. By including interfering samples in the sample preparation stage, it is possible to effectively verify whether the model can accurately distinguish barley flour adulteration from other grain interfering substances during the model evaluation stage, thereby ensuring the specific response capability of the detection method of this invention to the target adulterant in practical applications.
[0076] In one embodiment of the present invention, the environmental conditions for acquiring the spectrum in step S2 include: a spectral scanning range of 800-2500 nm and a scanning resolution of [missing information]. (Preferred) The number of scans is 32-64 per sample (preferably 32 per sample), the ambient temperature is 20±2℃, and the relative humidity is 45%-55%; the spectrum of each sample is collected at least 3 times and the average spectrum is taken as the raw data.
[0077] In this embodiment, firstly, regarding the spectral scanning range, the 800-2500 nm band completely covers the main absorption regions of the overtone and combination vibrations of hydrogen-containing groups in the near-infrared spectrum. Specifically, the area around 1438 nm corresponds to the first overtone of the OH stretching vibration of water, and the 1880-2180 nm range covers the combination frequency of the NH stretching vibration of proteins and amide groups, as well as the related absorptions of the CH and OH stretching vibrations of starch. The differences between barley flour and wheat flour in terms of moisture content, protein composition, and starch structure are all reflected in their spectral expressions within this band. Thus, limiting the scanning range to 800-2500 nm ensures that characteristic absorption information related to doping identification is not overlooked, while avoiding the introduction of redundant bands that do not contribute information, thus balancing the integrity of spectral information and data acquisition efficiency.
[0078] Second, regarding scanning resolution, spectral resolution determines the spacing between adjacent wavelengths, directly affecting the spectrum's ability to resolve narrowband absorption peaks. In this embodiment, The resolution effectively distinguishes the subtle frequency shift of the NH vibrations of wheat glutenin and barley gliadin around 1979 nm in the near-infrared region, as well as the gradual redshift of the protein amide II band from 1202 nm to 1198 nm with varying doping ratios. If the resolution is too low, these minute peak shifts caused by differences in protein composition will be smoothed out or masked, losing their potential as an indicator of doping ratios. In other words, The resolution is such that it can balance the signal-to-noise ratio and data acquisition time of a single scan while ensuring that the characteristic peaks are identifiable.
[0079] Third, increasing the number of scans and averaging them multiple times improves the spectral signal-to-noise ratio. According to the principle of signal accumulation, the signal-to-noise ratio is proportional to the square root of the number of scans. 32 accumulations can reduce random noise to about 17.7% of that of a single scan. For low-concentration doping detection, the characteristic signal intensity of the dopant component is only a few percent or even lower than that of the main component signal. After 32 accumulations and averaging, the noise floor is effectively suppressed, allowing weak spectral differences to emerge from the noise, providing raw data with a sufficient signal-to-noise ratio for the difference spectrum calculation in step S4.
[0080] Fourth, the intensity and position of near-infrared absorption peaks are sensitive to temperature changes. Temperature fluctuations can cause changes in hydrogen bond strength, leading to slight shifts or broadening of OH and NH-related absorption peaks. This implementation method controls the ambient temperature within a narrow range of ±2℃ (20±2℃), effectively suppressing spectral drift introduced by temperature effects. Simultaneously, controlling the relative humidity at 45%-55% prevents significant water vapor exchange between the sample and the environment during collection, avoiding dynamic spectral changes caused by moisture absorption or loss.
[0081] Fifth, the microscopic uniformity of particle packing in the same powder sample may vary in different areas after it is loaded into the sample cell. Performing at least three repeated samplings of the same sample and averaging the results can, to some extent, smooth out local scattering fluctuations caused by differences in packing micro-regions, making the average spectrum more representative. This operation (repeatedly sampling the spectrum of each sample at least three times and averaging the spectrum) can further reduce the impact of intra-sample inhomogeneity on the accuracy of subsequent difference spectral calculations, ensuring that changes in the difference spectrum are mainly attributed to differences in dopant components, rather than accidental packing inhomogeneities in a single sampling.
[0082] In one embodiment of the present invention, step S4, which uses a multidimensional scaling method to reduce the dimensionality of the difference spectrum, specifically includes:
[0083] Calculate the difference spectrum between each pair of samples and The original distance between The calculation formula is as follows:
[0084]
[0085] In the formula, and The first The first sample and the first Differential spectral data of each sample, , and These are the L2 norms of the corresponding subscript samples within the differential spectrum;
[0086] Based on the original distance, an inner product distance matrix is constructed through a bi-centering operation. And solve the inner product distance matrix. eigenvalues of the covariance matrix and eigenvectors ;
[0087] Retain those with a cumulative contribution rate ≥90% The eigenvectors corresponding to each eigenvalue are used as the extracted principal components to obtain the dimensionality-reduced differential spectral features. .
[0088] In this embodiment, the first... In the calculation formula, and The first The and the first Differential spectral data of each sample, , These are the L2 norms of the differential spectra of each sample, reflecting the spectral energy of the sample itself; The inner product of the difference spectra of the two samples reflects their correlation. Since the input data is the difference spectrum after the wheat matrix background has been removed in step S4, the distance calculated here... The distance measurement directly measures the relative magnitude of spectral differences caused by doping components between samples. Compared to directly calculating the distance from the original spectra, distance measurement based on differential spectra can exclude the spectral contribution of common components in the wheat matrix, making the variation in distance between samples mainly driven by the difference in doping ratio (samples with similar doping ratios have smaller distances, while samples with large differences in doping ratios have larger distances). This provides an accurate numerical basis for subsequent dimensionality reduction that preserves the structured distribution of doping gradients.
[0089] Second, the bicentering operation is a mathematical transformation of the distance matrix. Its core function is to transform the distance matrix, which only reflects the relative distances between samples, into an inner product matrix, allowing the sample points to be assigned coordinates in Euclidean space. Solving for the inner product distance matrix... eigenvalues of the covariance matrix and eigenvectors Essentially, this involves decomposing the distribution of differential spectral samples in the distance space into principal directions. Each eigenvector represents an independent direction of change for a sample in the feature space, and the magnitude of the corresponding eigenvalue reflects the significance of the differences between samples in that direction. Since the input distance matrix originates from the differential spectrum, the principal directions extracted by the eigenvalue decomposition directly correspond to the spectral difference patterns caused by changes in doping ratios, rather than the interference directions of fluctuations in the wheat matrix's own variety or composition.
[0090] Third, the cumulative contribution rate represents the percentage of retained contributions. Each principal component explains the total difference between samples. Setting the threshold to ≥90% means that the dimensionality-reduced feature space can cover most of the spectral variations caused by doping in the differential spectra, while discarding the remaining principal components with a cumulative contribution rate of less than 10% (these low contribution rate principal components usually correspond to non-doping factors such as noise fluctuations or small instrument drift). In this way, dimensionality compression can be achieved while preserving the integrity of key doping information, reducing the input redundancy of the SE attention-enhanced one-dimensional convolutional neural network model in subsequent step S5, which helps to suppress the risk of overfitting and improve the training efficiency and prediction stability of the model.
[0091] In one embodiment of the present invention, the elements of the inner product distance matrix S The calculation formula is:
[0092]
[0093] In the formula, Let be the average of the squared near-infrared spectral distances between the i-th sample and all other samples. Let be the average of the squared near-infrared spectral distances between the j-th sample and all other samples. It is the total average of the squared near-infrared spectral distances between all pairs of samples.
[0094] In this embodiment, firstly, the calculated original distance is... It can only reflect the distance between samples in the difference spectral space, but the distance matrix itself cannot directly assign Cartesian coordinates to sample points, and cannot directly perform eigenvalue decomposition to obtain principal components. This implementation method uses the formula... Mathematically, this process performs a bi-centering transformation from the squared distance matrix to the inner product matrix. This transformation is equivalent to successively subtracting the row mean effect, column mean effect, and global mean effect of the distance matrix, so that the transformed matrix S satisfies the mathematical properties of the inner product matrix (symmetric and positive semi-definite, thus allowing it to be processed by subsequent eigenvalue decomposition). Without this step, multidimensional scaling cannot recover the low-dimensional coordinates of the sample points from the distance information, and subsequent dimensionality reduction operations cannot be performed.
[0095] Second, in the above formula ( Subtract from ) and The terms correspond to eliminating the first one. row and number The effect of the squared mean distance of the column, plus The term compensates for the global mean that was repeatedly subtracted due to two subtractions. This operation of "subtracting the row mean, subtracting the column mean, and adding the overall mean" makes the inner product matrix... The representation in Euclidean space is translation invariant, meaning that the overall position of the sample point set and the reference frame shift of individual samples do not affect the translation. The matrix represents the relative relationships between samples. In this way, it can be ensured that no matter how the global baseline of the differential spectral data fluctuates, the constructed low-dimensional feature space can faithfully reflect the relative structure between samples caused by the difference in doping ratio, without being disturbed by the overall spectral intensity baseline shift.
[0096] Third, the inner product distance matrix after doubly centered The inner product relationship between samples in the difference spectral space is directly encoded. The principal component directions obtained from its eigenvalue decomposition correspond one-to-one with the coordinates of the samples in the low-dimensional space, and the Euclidean distance between samples in this low-dimensional space can approximate the sample distance in the original high-dimensional difference spectrum to the greatest extent. For the detection of barley flour doping in wheat flour, this distance-preserving characteristic means that the spectral distance difference caused by small changes in the doping ratio in the difference spectral space can be proportionally preserved in the dimensionality-reduced low-dimensional feature space. This allows samples with different doping gradients, such as 5% and 20%, to present a clearly distinguishable distance interval in the feature space, providing the model in step S5 with a feature expression that is highly sensitive to changes in the doping ratio and has good resolution.
[0097] In one embodiment of the present invention, in step S3, the window size of the Savitzky-Golay smoothing filter is set to 9-13 points (preferably 11 points), and the polynomial order is 2-3 (preferably 2nd order). The combined preprocessing scheme is evaluated based on the signal-to-noise ratio, characteristic peak identification, and 1979nm peak intensity variation coefficient. The optimal combination is selected from a variety of combination schemes including MSC+SG, SNV+SG, MSC+SG+FD, and SNV+SG+SD. Here, MSC is multivariate scattering correction, SG is Savitzky-Golay smoothing filter, FD is first derivative, SNV is standard normal variable transformation, and SD is second derivative.
[0098] In this embodiment, firstly, for Savitzky-Golay smoothing filtering, setting the window size to 9-13 points determines the number of adjacent wavelength points participating in the local polynomial fitting, and setting the polynomial order to 2-3 defines the functional form of the fitted curve. Thus, through this parameter combination, a balance can be achieved between effectively suppressing random noise and preserving the subtle spectral features to the maximum extent. For low-concentration doping scenarios of barley flour in wheat flour, the characteristic signals of the dopant components often manifest as weak and narrow-band absorption peak changes in the spectrum. If the window value is too large, the smoothing operation will smooth these narrow-band features along with the noise, causing irreversible loss of dopant information; if the window is too small or the order is too low, the noise suppression effect is insufficient, and the signal-to-noise ratio improvement is limited. It should be particularly noted that the preferred 11-point window, combined with 2nd-order polynomial fitting, can closely follow the true trend of the spectral curve within the local band range where key characteristic peaks such as 1979nm are located, eliminating only random high-frequency fluctuations without peak clipping or broadening the characteristic peak shape, thus completely preserving the weak peak intensity changes of the dopant components in the difference spectrum in step S4.
[0099] Second, the signal-to-noise ratio (SNR) reflects the ratio of the effective signal intensity to the noise level in the spectrum, directly quantifying the suppression effect of the preprocessing scheme on random noise. Using it as an evaluation metric ensures that the selected scheme is optimal in terms of noise suppression, providing a low-noise base for the reliable extraction of low-concentration doped signals. The characteristic peak discrimination index measures the clarity of distinction between characteristic absorption peaks and adjacent baselines or neighboring peaks in the spectrum; a value closer to 1 indicates a sharper peak shape and more thorough shoulder peak separation. Incorporating this into the evaluation system ensures that the selected scheme can clearly present the key characteristic peaks that distinguish barley flour from wheat flour, rather than submerging them in a broad peak envelope. The 1979nm peak intensity variation coefficient quantifies the repeatability stability of the peak intensity in different doped samples, targeting the core identification peak of the NH stretching vibration and amide group combination frequency of wheat glutenin. A lower variation coefficient indicates higher signal stability of the characteristic peak across multiple acquisitions or parallel samples. The 1979 nm peak, located in the protein amide II band, is one of the key wavelengths that significantly influences the response of barley flour and wheat flour due to differences in gluten composition. Its peak intensity increases linearly with the doping ratio, making it one of the most important quantitative indicator peaks for doping. Its coefficient of variation is listed separately as an evaluation index to ensure optimal stability of the selected preprocessing scheme for this core quantitative signal and to avoid quantitative errors introduced by excessive fluctuations in the characteristic peak itself.
[0100] Through the synergistic evaluation of the above three indicators, the signal-to-noise ratio ensures noise removal, the characteristic peak identification ensures that effective information is not mistakenly deleted, and the coefficient of variation of the 1979nm peak intensity ensures the repeatability reliability of the core quantitative signal. In this way, by selecting the optimal scheme from the four candidate schemes based on the index system, the optimal balance can be achieved in the three dimensions of noise suppression, feature preservation and signal stability, providing high-quality spectral input for the subsequent step S4.
[0101] In one embodiment of the present invention, the SE attention-enhanced one-dimensional convolutional neural network model uses a basic one-dimensional convolutional neural network as its backbone and introduces an SE attention mechanism module with a compression ratio of 16 after the convolutional layer. The SE attention mechanism module generates channel weights through global average pooling, fully connected layer compression and recovery, and Sigmoid activation function to adaptively enhance the weights of key feature peaks at 1438nm, 1880nm, and 1979nm, and suppress spectral overlap interference and noise.
[0102] In this embodiment, firstly, the one-dimensional convolutional neural network slides the convolution kernel along the spectral dimension to perform weighted combination and nonlinear activation of the spectral response within a local band, enabling it to capture the covariance patterns and interactions between adjacent wavelength points in the feature vector. This allows it to adapt to the sequential characteristics of near-infrared spectral data, significantly reducing the number of training parameters compared to fully connected networks, lowering the risk of overfitting, while maintaining the ability to extract the combined features of local absorption peaks in the spectrum, providing a suitable backbone network for the subsequent introduction of attention mechanisms.
[0103] Second, an SE attention mechanism module is embedded on the multi-channel feature map output by the convolutional layer. Its function is to recalibrate the weights of each feature channel extracted by the convolution, enabling the model to adaptively select key feature channels. A compression ratio of 16 means that the feature description vectors of the channel dimension are compressed to 1 / 16 of their original dimension through the fully connected layer before being restored. This parameter strikes a balance between the ability to model inter-channel dependencies and computational efficiency (too low a compression ratio makes it difficult to fully capture the non-linear interactions between channels, while too high a ratio may lead to excessive information loss during compression). Thus, after the compression-restoration operation, the model can learn the inter-channel dependencies within the global receptive field, rather than relying solely on local information for channel weight allocation.
[0104] Third, global average pooling compresses the spatial dimensionality of each feature channel into a scalar, serving as a statistical description of the channel's global response intensity. This allows subsequent weight generation to be based on a comprehensive understanding of the entire spectral dimensionality. The fully connected layer's compression and recovery constitutes a bottleneck structure, generating importance scores for each channel by learning the nonlinear interactions between channels. The Sigmoid activation function maps these scores to the 0-1 range, generating a normalized channel weight vector. This weight vector is applied channel-by-channel to the original feature map of the convolutional layer, amplifying feature channels highly correlated with doping ratios and suppressing channels unrelated to doping or containing noise. Since the weight allocation is automatically learned from the training data, the model can autonomously identify the spectral band combinations most strongly correlated with doping components during training, without requiring manual pre-setting of fixed weights for feature bands.
[0105] Fourth, the differences between barley flour and wheat flour in key components are concentrated in the absorption characteristics at specific wavelengths. Specifically, the difference in peak intensity of the first overtone of the OH stretching vibration of moisture at 1438 nm reflects the low moisture content of barley flour; the area around 1880 nm encompasses the CH and OH stretching vibrations of starch, reflecting the difference in carbohydrate composition; and the combination of the NH stretching vibration and amide group at 1979 nm reflects the difference in the ratio of glutenin to gliadin, which are the core spectral markers distinguishing the protein composition of barley flour and wheat flour. This implementation method uses the SE attention mechanism to adaptively enhance the weights of the feature channels containing the above three key characteristic peaks. This allows the model to automatically focus on these key band responses that are significant for doping identification after receiving the difference spectral features extracted in step S4. At the same time, it suppresses residual noise in non-characteristic bands and spectral overlap interference from other components, thereby accurately extracting sensitive spectral features that are quantitatively related to the doping ratio of barley flour from weak difference signals, achieving precise quantitative detection of low-concentration doping.
[0106] In one embodiment of the present invention, during the training process of the SE attention-enhanced one-dimensional convolutional neural network model, the learning rate search range is set to 0.001-0.01, the number of iterations is 80-100 (preferably 100), and the batch size is uniformly 16-64 (preferably 32). The training adopts ten-fold cross-validation, and the calibration set is randomly divided into 10 subsets. Multiple rounds of training are completed by taking turns using 9 subsets as training samples and 1 subset as validation samples to avoid data bias.
[0107] In this implementation, firstly, the learning rate is a core hyperparameter controlling the step size of parameter updates during deep learning model training. Its search range is limited to 0.001 to 0.01. On one hand, this search boundary is reasonable, avoiding the risk of low training efficiency or model non-convergence due to blindly searching within an excessively large range. On the other hand, this range is suitable for the structural scale and input feature dimension of the SE attention-enhanced one-dimensional convolutional neural network in this invention, enabling the model to gradually approach the optimal solution of the loss function with an appropriate step size during training. This avoids gradient updates jumping past the optimal point due to an excessively large learning rate, or convergence speed being too slow due to an excessively small learning rate. The optimal learning rate determined through searching within this range allows the model to achieve a good balance between training loss and generalization performance within a limited number of iterations.
[0108] Second, the number of iterations determines the complete number of scan rounds the model performs on the training set. Setting the number of iterations to 80-100, combined with the sample size of the calibration set in this invention, allows the model to have ample learning opportunities while setting a clear termination boundary. During 80-100 iterations, the model can adequately fit the nonlinear mapping relationship between spectral difference features and doping ratios; terminating training after 80-100 rounds prevents overfitting due to excessive training rounds (i.e., the model excessively memorizes the specific spectral fluctuations of individual samples in the calibration set, thus losing its generalization ability to the independent prediction set). In other words, this setting of the number of iterations achieves a balance between training sufficiency and generalization performance.
[0109] Third, the batch size determines the number of samples used for each model parameter update. Standardizing the batch size to 16-64 has two benefits: First, compared to stochastic gradient descent with updates per sample, gradient estimation with 16-64 samples per batch has statistical stability, allowing parameter updates to more accurately align with the global loss function's descent direction and reducing parameter oscillations during training. Second, compared to full-batch training, the smaller batch size of 16-64 preserves a degree of randomness in gradient estimation, helping the model escape local optima during optimization and improving final convergence quality. Furthermore, a standardized batch size of 16-64 helps eliminate differences in training conditions introduced by varying batch sample sizes, ensuring fair and consistent performance comparisons between different models or training epochs.
[0110] Fourth, the effectiveness of 10-fold cross-validation lies in two aspects: systematically avoiding data bias and reliably assessing the model's generalization ability. Randomly dividing the calibration set into 10 equal subsets and rotating them for training and validation means that every sample in the calibration set has the opportunity to participate in the internal evaluation of model performance as a validation sample. This eliminates the accidental impact on model evaluation caused by insufficient representativeness or distribution bias in a single fixed partition. Using 9 subsets as training samples and 1 subset as validation samples for 10 rounds effectively evaluates the model's performance stability under different data subset combinations. This ensures that parameter tuning during training is based on reliable internal validation feedback, avoiding the model being guided towards suboptimal parameter directions that only perform well on specific data subsets due to data partition bias. This cross-validation strategy is connected with the two-layer evaluation system in step S1, which divides the samples into a calibration set and an independent prediction set (cross-validation guides the selection of model tuning parameters within the calibration set, while the independent prediction set performs final performance verification after the model training is completed and does not participate in parameter tuning throughout the process). This ensures that the final output model performance index can truly reflect the actual application effect of the detection method of this invention on unknown samples.
[0111] In one embodiment of the present invention, the non-barley grain interference sample includes buckwheat flour, oat flour, and wheat bran; after step S5, the method further includes: processing the near-infrared spectrum of the non-barley grain interference sample through steps S2 to S5, and verifying whether the predicted doping ratio of the interference sample by the SE attention-enhanced one-dimensional convolutional neural network model is lower than a preset threshold, so as to determine the model's specific identification ability of barley flour doping in wheat flour.
[0112] In this embodiment, firstly, non-barley grain interference samples are set as buckwheat flour, oat flour, and wheat bran. This ensures that the type of interference targeted in the specificity verification has a definite direction. Although buckwheat and oats belong to the same family (Poaceae) as barley in botanical classification, they have identifiable chemical differences in protein composition, starch structure, and cellulose content. Their near-infrared spectra may overlap with barley flour to varying degrees in specific bands. Wheat bran, as a byproduct of wheat processing, has spectral characteristics related to the wheat matrix and may produce interference patterns similar to the wheat matrix background in doping detection. Including these three types of substances in the interference sample category for specificity verification ensures that the verification process covers representative interference types with actual confusion potential, ensuring that specificity verification is not targeted at a single interference, but rather examines the model's specificity in identifying the target dopant under a diverse interference background.
[0113] Second, this implementation adds a verification operation after step S5. Firstly, this allows the method of the present invention to not only predict the doping ratio of unknown samples but also to self-verify whether the prediction results are affected by non-target interfering substances. Thus, the method of the present invention can provide auxiliary judgment criteria for the reliability of the results while outputting the prediction results. Secondly, the verification step requires the interfering sample to undergo the exact same processing flow as the sample to be tested (i.e., spectral acquisition in step S2, combined preprocessing in step S3, difference spectral calculation and multidimensional scaling and dimensionality reduction in step S4, and SE attention-enhanced one-dimensional convolutional neural network model prediction in step S5). This ensures the consistency between the verification conditions and the actual measurement conditions in the data processing path, enabling the verification conclusion to truly reflect the model's behavior when facing non-target interfering substances in the actual detection scenario, rather than an extrapolation conclusion obtained under simplified or idealized conditions.
[0114] The technical solution of the present invention will be further described below with reference to an exemplary embodiment.
[0115] I. Experimental Materials and Instruments.
[0116] 1. Grain varieties
[0117] Imported categories: Wheat: Russia, Canada, Kazakhstan, France, USA, Australia; Barley: France, Ukraine, Australia, Russia, Kazakhstan; Buckwheat: Russia, Kazakhstan; Oats: Kazakhstan; Wheat bran: Kazakhstan
[0118] Domestic wheat varieties: Wheat flour: Xinmai 32, Zhengmai 9023, Anong 1124, Zhengmai 1860, Jimai 480, Jimai 32, Mianmai 51, Xinong 511, Zhoumai 24, Zhoumai 36, Zhenmai 12, Yangmai 15, Huaimai 330, Jimai 44, Aikang 58, Yili wheat; Barley: Yangsimai No. 3, Yangnongpi No. 13, Minmai No. 2, Yangsimai No. 5, Yangnongpi No. 7, Yangnongpi No. 5, Yangnongpi No. 9, Yangnongpi No. 14, Fudamai No. 1.
[0119] 2. Experimental Instruments and Equipment
[0120] The main instruments and equipment used in this embodiment are shown in Table 1 below. All instruments have been calibrated before use to ensure the accuracy and reliability of the test data.
[0121] Table 1 Experimental Instruments
[0122]
[0123] II. Experimental Methods.
[0124] 1. Sample preprocessing
[0125] All original samples were sealed in self-sealing bags and equilibrated at room temperature (20±2℃) for at least 2 hours to eliminate the influence of temperature differences on spectral acquisition. Each sample was reduced to approximately 150g using a quartering method, pulverized using a high-speed multi-functional pulverizer, and then sieved through a 60-mesh sieve to ensure uniform particle size. Finished wheat flour was directly sieved without further pulverization. For doping sample preparation, 22 high-quality wheat samples were used as the matrix, paired with 15 barley samples for doping. Three mass fraction gradients (5%, 20%, and 50%) were set up, with one set of parallel samples for each doping ratio. Three doping samples could be prepared from a single wheat matrix and a single barley flour, resulting in a total of 990 doping samples from the 22 wheat flour samples and 15 barley flour samples. To ensure sample representativeness while maintaining experimental efficiency and avoiding excessive sample size that would increase the modeling burden, 198 representative samples were selected as the core doping samples through stratified random sampling at a sampling ratio of 20%. The remaining 792 doping samples were used as backup samples for subsequent model stability verification, repeated experiments, and supplementation of outlier data, ensuring the reproducibility and reliability of the experiment.
[0126] A total of 239 samples were used in the experiment, comprising 22 pure wheat flour controls, 15 pure barley flour controls, and 4 interference samples. The SPXY sample set partitioning method was used, dividing the 239 core samples into a calibration set of 179 samples and an independent prediction set of 60 samples at a 3:1 ratio. The calibration set was used for model building, parameter tuning, and internal validation, while the independent prediction set was only used for final model performance evaluation and did not participate in the parameter tuning process. All experiments employed ten-fold cross-validation for model training and parameter optimization. The calibration set was randomly divided into 10 subsets: 9 subsets containing 18 samples each, and 1 subset containing 17 samples. Multiple rounds of training were conducted using 9 subsets as training samples and 1 subset as validation samples, effectively avoiding data bias.
[0127] 2. Near-infrared spectroscopy operation procedures and parameter settings
[0128] Spectral acquisition parameters were optimized through preliminary experiments to obtain high-quality spectral data in the spectral range of 800-2500 nm; scanning resolution... 32 scans per sample, integrating sphere acquisition mode, reference standard spectral purity. A white background was used, with ambient temperatures of 20±2℃ and relative humidity of 45%-55%. Each sample was placed in a standard sample cell, compacted and leveled to avoid light scattering differences caused by interparticle gaps. Each sample was collected three times, and the average spectrum was used as the raw data. This standardized acquisition process effectively ensured the stability and consistency of the spectral data.
[0129] 3. Spectral preprocessing
[0130] Near-infrared raw spectra are susceptible to interference from sample particle size, scattering effects, baseline drift, and environmental noise. This chapter employs a combination of preprocessing methods for optimization. Multivariate scattering correction (MSC) can eliminate scattering effects caused by differences in sample particle size and enhance the correlation between spectra and chemical composition. Standard normal variable transformation (SNV) standardizes spectral data, eliminating baseline drift and differences in instrument response. Savitzky-Golay (SG) smoothing, with a window size of 11 points and a polynomial order of 2, suppresses random noise. First derivative (FD) and second derivative (SD) enhance the resolution of spectral characteristic peaks and eliminate baseline drift and background interference. The combined preprocessing schemes are compared, with the processing effects of four combinations: MSC+SG, SNV+SG, MSC+SG+FD, and SNV+SG+SD. The optimal scheme is selected based on SNR and characteristic peak identification as evaluation indicators.
[0131] 4. Feature Extraction
[0132] In the preliminary experiment, 93 samples were first classified into four categories: imported wheat flour, domestic wheat flour, imported barley flour, and domestic barley flour. The PCA method of unsupervised pattern recognition was used for preliminary analysis to explore the overall clustering effect of the samples. On the other hand, a quantitative PLS-DA model was constructed to carry out preliminary modeling attempts.
[0133] Depend on Figure 2 As shown in Figure a, although the score points of the three sample classes all lie within the 95% confidence ellipse, they do not exhibit a clear separation trend, indicating that unsupervised learning using PCA alone is insufficient to effectively distinguish between different sample classes. PCA analysis results show that the variance contribution rates of the first principal component (PC1) and the second principal component (PC2) are 85.5% and 11.1%, respectively. Only the QC samples are closely distributed, reflecting the instrument's high accuracy and stability throughout the detection process. However, the preliminary model built based on these PCA features has a predicted value and variance explained rate both below 0.99, indicating poor model fit and inability to meet the needs of subsequent classification and quantitative analysis.
[0134] Ultimately, to address the high-dimensional challenge in spectral doping analysis, multidimensional scaling (MDS) is employed to extract key features and reduce dimensionality while preserving pairwise spectral distances. x represents the collected spectral data. The first step is to calculate the distance between each data point. This is achieved by constructing a distance matrix, which is derived using the following formula:
[0135]
[0136] in and They represent the first The first sample and the first Spectral data of each sample It is the first in the original spectrum The first sample and the first The distance between samples , , It is the L2 norm of the corresponding subscript sample within the characteristic spectrum.
[0137] The distance equation can be defined as follows:
[0138]
[0139]
[0140]
[0141] From the above three equations, we can derive:
[0142]
[0143]
[0144]
[0145] Further exportable:
[0146]
[0147]
[0148] Furthermore, through combination, we can conclude that:
[0149]
[0150] The distance matrix S can be obtained using the organizational principle given in the following formula:
[0151]
[0152] Then, based on the above formula, and Determine the covariance matrix of the distance matrix S The eigenvalues λ and eigenvectors σ; where E is a unit vector, according to the formula ( Calculate the contribution rate of each eigenvalue on the principal component, denoted by R(i);
[0153] When the necessary conditions are met At that time, the first t principal components are extracted. The spectrum after dimensionality reduction. As shown in the following formula:
[0154]
[0155] The contribution rate of each principal component is calculated, and the top t principal components with a cumulative contribution rate ≥ 90% are selected as feature vectors. The extracted feature vectors are then input into PLS-DA and 1D-CNN algorithms respectively to establish a hybrid hypothesis model. All the above preprocessing and feature extraction are implemented using Python 3.8 code.
[0156] III. Test Results
[0157] 1. The impact of preprocessing analysis on spectral quality
[0158] This experiment employed multiple preprocessing methods, including MSC, SNV, Savitzky-Golay (SG) smoothing, first derivative (FD), and second derivative (SD). Four combination schemes were designed: MSC+SG, SNV+SG, MSC+SG+FD, and SNV+SG+SD. The processing effects of different preprocessing schemes on the spectra of low-doped samples were compared. The signal-to-noise ratio (SNR), characteristic peak discrimination (PID), and coefficient of variation (CV%) of the 1979 nm peak intensity were used as the core evaluation indicators. Among them, SNR reflects the spectral noise suppression effect, with higher values indicating less noise; PID reflects the clarity of characteristic peaks, with values closer to 1 indicating higher discrimination; and CV% reflects spectral stability, with lower values indicating better data repeatability. The specific evaluation data for each scheme are shown in Table 2 below.
[0159] Table 2 Evaluation of the effects of different pretreatment schemes
[0160]
[0161] Comprehensive evaluation shows that the MSC+SG+FD combined preprocessing scheme has the best effect, with the highest SNR and PID, and the lowest coefficient of variation of characteristic peak intensity. Figure 3 As shown, by eliminating scattering effects through MSC, smoothing and suppressing random noise through SG, and enhancing the resolution of characteristic peaks and eliminating baseline drift through the first derivative, this approach effectively balances noise suppression, baseline correction, and feature preservation. Combined with the data in Table 2, it can be seen that this scheme improves SNR by 6.4% and PID by 2.2% compared to SNV + SG + SD. Compared to the other three schemes, it effectively overcomes the limitations of single preprocessing methods, solves the problem of incomplete baseline correction in MSC+SG, and avoids the shortcomings of insufficient scattering suppression in SNV+SG and feature loss in SNV+SG+SD. Furthermore, the slight noise amplification phenomenon can be further mitigated through subsequent characteristic wavelength selection and modeling algorithm optimization, and will not have a significant impact on subsequent research.
[0162] The processing flow of this scheme is well-suited to the characteristics of weak spectral signals and numerous interferences in samples with low doping ratios. It can preserve the effective spectral information of the samples to the greatest extent, highlight the characteristic peaks corresponding to low doping components, and provide stable and reliable spectral data for subsequent model construction.
[0163] 2. Quantitative analysis based on a single type of spectrum
[0164] After preprocessing with a combination of MSC+SG+FD, the near-infrared spectrum showed a flat baseline and clearly visible characteristic peaks. Figure 4 As shown, the significant characteristic peaks of the pure wheat sample at key wavelengths reflect its chemical composition of moisture, protein, and starch.
[0165] The characteristic peaks of pure wheat are shown in Table 3 below. The absorption peak at 1438 nm in pure wheat flour belongs to the first overtone of the -OH stretching vibration, directly indicating that the moisture content of the sample and the peak intensity are positively correlated with the moisture concentration. The characteristic peaks of proteins are distributed in the range of 1880-2180 nm. Among them, 1880 nm corresponds to the second overtone of the -O=H stretching vibration and the starch CH stretching vibration, 1979 nm is the combination frequency of the =NH stretching vibration and amide II, and 2180 nm is the combination frequency of amide I and the combination frequency of C=O and amide III. These multi-peak combined characteristics can be used to distinguish protein types.
[0166] Table 3 Comparison of characteristic peaks of pure wheat flour
[0167]
[0168] Near-infrared spectroscopy was used to detect barley flour adulteration in wheat flour samples. The absorbance variations of different wheat varieties in the 200-2500 nm wavelength range were analyzed at adulteration ratios of 5%, 20%, and 50%. Taking Zhengmai 1860 adulterated with Yangnongpi 13 as an example... Figure 5 As shown in Figure a, in the 200-400 nm ultraviolet-visible short-wavelength region, the absorbance of all samples is generally low, indicating that the molecules in this range have weak absorption of near-infrared light. However, when the wavelength reaches 400-2000 nm, the absorbance generally shows an upward trend, and characteristic absorption peaks appear near 1500 nm and 2000 nm. This may be related to the vibrational absorption of components such as carbohydrates, proteins, or fats in the samples.
[0169] It is worth noting that the sample adulterated with 5% inferior wheat exhibited high absorbance across the entire wavelength range, while samples with different adulteration ratios showed a certain correlation in absorbance fluctuation trends. Taking Zhengmai 1860 wheat flour adulterated with Minmai No. 2 barley flour as an example, the near-infrared preprocessed spectra of different adulteration ratios are as follows: Figure 5 As shown in b.
[0170] As the doping ratio increases, the spectrum shows a regular change. The intensity of the 960nm peak decreases linearly with the increase of the doping ratio, while the intensity of the 1480nm peak increases linearly, and the intensity of the 1976nm peak increases linearly. This is consistent with the characteristics of barley flour, which is low in moisture and high in cellulose. The protein amide II band at 1200nm gradually redshifts from 1202nm to 1198nm. The redshift amplitude is significantly positively correlated with the doping ratio, reflecting the gradual change in protein content and composition. This can be used as a preliminary quantitative indicator of the doping ratio, and is especially suitable for rapid on-site screening.
[0171] 3. Model Building and Optimization
[0172] Feature vectors were extracted using the same MDS method as described above, and five quantitative models were constructed: PLS-DA, basic 1D-CNN, ResNet-18, LSTM, and a self-developed SE attention-enhanced 1D-CNN. All deep learning models employed a unified optimization standard to ensure fairness in the comparison. Specifically, ResNet-18 was modified to a 1D residual network structure, and the 1D-CNN (SSA) was improved by introducing an SE attention mechanism with a compression ratio of 16, making it suitable for near-infrared spectral sequence data. The learning rate search range was uniformly set to 0.001-0.01, the number of iterations was 100, and the batch size was uniformly set to 32 to ensure fairness in the comparison.
[0173] 4. Using the same evaluation metrics as above, the performance comparability of models in different sections is ensured. The performance comparison results of each model are shown in Table 4 below. The self-developed SE attention-enhanced 1D-CNN shows the best performance. With a resolution of 0.990, RMSEp=0.71, and RPD=4.48, its performance surpasses that of basic 1D-CNN and LSTM models, reducing RMSEp to 1.05%. Its core advantage lies in the SE attention mechanism, which enhances the weight allocation of key feature peaks such as 1438nm and 1880nm, effectively suppressing spectral overlap interference. The convolutional and pooling layers collaboratively mine deep nonlinear features. With a 5% low doping ratio, the relative error is only 3.8±0.2%, fully meeting the high-precision detection requirements for barley flour doping in wheat flour. Furthermore, compared to the ResNet-18 model, it reduces RMSEp by 32.4% and improves RPD by 20.3%, demonstrating significant advantages.
[0174] Table 4 Comparison of Performance Evaluation Indicators for Different Models
[0175]
[0176] 5. Adulteration Model Validation
[0177] Using Zhengmai 1860 wheat flour as the matrix, 5%, 20%, and 50% of Yangnongpi 13 barley flour were added respectively to obtain near-infrared spectra, and the original near-infrared spectra were preprocessed using the MSC+SG+FD combination scheme.
[0178] Based on the preprocessed spectral data, specificity detection and cross-validation were performed using an improved 1D-CNN model. The results... Figure 6As shown, baseline drift and multivariate scattering effects were suppressed, noise was reduced, and the characteristic peak differences of samples with different doping ratios were clear. For example, the key features at 330nm, 1202nm, 1466nm, 1933nm, and 2102nm showed slightly improved identification. Based on the preprocessed spectral data, a modified 1D-CNN model with an SE attention mechanism was used to perform specificity detection and 10-fold cross-validation. The results showed that the model achieved an accuracy of 92.5% in identifying barley flour doping in wheat flour.
[0179] 6. Validation of model specificity and generalization ability
[0180] The spectra of interfering samples such as buckwheat flour, oat flour, and wheat bran were input into the improved 1D-CNN model, and the prediction results are shown in Table 5 below. The predicted doping ratio of all interfering samples was less than 0.8%, and the model's recognition accuracy reached 92.6%, indicating that the model can effectively distinguish barley flour from other grain components and has good specificity.
[0181] Table 5. Model specificity validation results
[0182]
[0183] The generalization ability of the model was tested using external validation samples, Canadian wheat flour mixed with French barley flour. The prediction performance of the improved 1D-CNN model was [performance value missing]. =0.988, RMSEP=0.76, RPD=7.21, which are close to the internal validation results. It can adapt to wheat and barley samples from different origins and varieties. The modeling data has broad representativeness. The deep learning capability of the improved 1D-CNN enhances the model's adaptability to different sample matrices.
[0184] IV. Summary
[0185] This implementation employs a combined MSC+SG+FD preprocessing scheme. The introduced SE attention mechanism adaptively adjusts feature weights to enhance attention to key feature peaks and suppress the influence of noise and redundant information. Simultaneously, it increases the number of convolutional kernels and adjusts the pooling window to optimize the network structure, thereby improving the model's ability to mine deep nonlinear spectral features. (Test set) =0.99, RMSEp=0.71, RPD=4.48.
[0186] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy, characterized in that, Includes the following steps: Step S1: Mix pure wheat flour matrix of different varieties with pure barley flour of different varieties according to multiple preset mass fraction gradients to prepare doped samples. At the same time, set up pure wheat flour control samples to obtain the average spectrum of pure wheat flour, and divide the doped samples into a calibration set and an independent prediction set. Step S2: Collect near-infrared spectral data of the doped sample and the pure wheat flour control sample; Step S3: The acquired near-infrared spectral data is processed using a combined preprocessing scheme, which includes multivariate scattering correction, Savitzky-Golay smoothing filtering, and first derivative processing to obtain a preprocessed spectrum with a flat baseline and clear characteristic peaks. Step S4: Using the average spectrum of the pure wheat flour control sample as a reference, calculate the difference spectrum between the pre-processed spectrum of each doped sample and the reference. The differential spectra are reduced in dimensionality using a multidimensional scaling method, and the top spectra whose cumulative contribution rate meets a preset threshold are extracted. Each principal component is used as an eigenvector to focus on the spectral differences introduced by doping and to preserve the pairwise distance structure between data. Step S5: Input the extracted feature vector into the constructed SE attention-enhanced one-dimensional convolutional neural network model to output the doping ratio of barley flour in the sample to be tested.
2. The method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy according to claim 1, characterized in that, In step S1, after all the original samples are crushed and sieved and before the spectral acquisition, they are allowed to stand for at least 2 hours in an environment with a temperature of 20±2℃ and a relative humidity of 45%-55% to eliminate spectral variations unrelated to chemical doping caused by differences in water activity and electrostatic state between samples, so as to ensure that the weak signals detected later truly reflect the changes in composition.
3. The method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy according to claim 1, characterized in that, The multiple preset mass fraction gradients include 5%, 20% and 50%, and at least one set of parallel samples is set for each doping ratio; the sample preparation also includes setting up a pure barley flour control sample and a non-barley grain interference sample.
4. The method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy according to claim 1, characterized in that, In step S2, the environmental conditions for acquiring the spectrum include: a spectral scanning range of 800-2500 nm and a scanning resolution of [missing information]. The number of scans was 32-64 per sample, the ambient temperature was 20±2℃, and the relative humidity was 45%-55%; the spectrum of each sample was collected at least 3 times and the average spectrum was used as the raw data.
5. The method for quantitatively detecting adulteration of wheat flour with barley flour based on near infrared spectroscopy according to claim 1, characterized by, In step S4, the dimensionality reduction of the difference spectrum using a multidimensional scaling method specifically includes: The raw distance between each two samples in the difference spectrum is calculated and The raw distance between each two samples in the difference spectrum is calculated The formula is: ; In the formula, and The first The first sample and the first Differential spectral data of each sample, , and These are the L2 norms of the corresponding subscript samples within the differential spectrum; Based on the original distances, construct an inner product distance matrix by a double centering operation and solve the eigenvalues and eigenvectors of the covariance matrix of the inner product distance matrix ; Retain those with a cumulative contribution rate ≥90% The eigenvectors corresponding to each eigenvalue are used as the extracted principal components to obtain the dimensionality-reduced differential spectral features. .
6. The method for quantitatively detecting adulteration of wheat flour with barley flour based on near infrared spectroscopy according to claim 5, characterized by The elements of the inner product distance matrix S The calculation formula is: ; In the formula, Let be the average of the squared near-infrared spectral distances between the i-th sample and all other samples. Let be the average of the squared near-infrared spectral distances between the j-th sample and all other samples. It is the total average of the squared near-infrared spectral distances between all pairs of samples.
7. The method for quantitative detection of barley flour adulteration in wheat flour based on near-infrared spectroscopy according to claim 1, characterized in that, In step S3, the window size of the Savitzky-Golay smoothing filter is set to 9-13 points, and the polynomial order is 2-3. The combined preprocessing scheme uses signal-to-noise ratio, characteristic peak identification, and 1979nm peak intensity variation coefficient as evaluation indicators, and selects the optimal combination from a variety of combination schemes including MSC+SG, SNV+SG, MSC+SG+FD, and SNV+SG+SD. Among them, MSC is multivariate scattering correction, SG is Savitzky-Golay smoothing filter, FD is first derivative, SNV is standard normal variable transformation, and SD is second derivative.
8. The method for quantitatively detecting adulteration of wheat flour with barley flour based on near infrared spectroscopy according to claim 1, characterized by, The SE attention-enhanced one-dimensional convolutional neural network model uses a basic one-dimensional convolutional neural network as its backbone. After the convolutional layer, an SE attention mechanism module with a compression ratio of 16 is introduced. The SE attention mechanism module generates channel weights through global average pooling, fully connected layer compression and recovery, and Sigmoid activation function to adaptively enhance the weights of key feature peaks at 1438nm, 1880nm, and 1979nm, and suppress spectral overlap interference and noise.
9. The method for quantitatively detecting adulteration of wheat flour with barley flour based on near infrared spectroscopy according to claim 8, characterized by, During the training process of the SE attention-enhanced one-dimensional convolutional neural network model, the learning rate search range is set to 0.001-0.01, the number of iterations is 80-120, and the batch size is uniformly 16-64. The training adopts ten-fold cross-validation, and the calibration set is randomly divided into 10 subsets. Multiple rounds of training are completed by taking turns using 9 subsets as training samples and 1 subset as validation samples to avoid data bias.
10. The method for quantitatively detecting adulteration of wheat flour with barley flour based on near infrared spectroscopy according to claim 1, characterized by, Non-barley grain interference samples included buckwheat flour, oat flour, and wheat bran; After step S5, the method further includes: processing the near-infrared spectrum of the non-barley grain interference sample through steps S2 to S5, and verifying whether the predicted doping ratio of the interference sample by the SE attention-enhanced one-dimensional convolutional neural network model is lower than a preset threshold, so as to determine the model's specific identification ability of barley flour doping in wheat flour.
Citation Information
Patent Citations
Wheat flour quality characteristic prediction method based on near infrared spectrum
CN117330535A
Bacterial Raman spectrum chemical component cross-domain analysis method based on deep transfer learning
CN120673875A