Sesame paste food ingredient content detection method based on spectral analysis
By collecting near-infrared spectral data under different temperature gradients, constructing a temperature-compensated spectral feature matrix, and combining it with a multi-task regression model and a standard substance spectral library, the problem of sesame paste component detection under the influence of temperature fluctuations was solved, achieving high-precision and efficient component analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG SHIJICHUN FOOD
- Filing Date
- 2026-04-16
- Publication Date
- 2026-07-10
AI Technical Summary
In existing technologies, the detection of sesame paste components is affected by temperature fluctuations, resulting in inaccurate spectral feature matrices. Conventional detection methods are cumbersome and have limited accuracy, failing to effectively combine standard material spectral libraries for error calibration, making it difficult to achieve accurate detection.
Near-infrared transmission spectral data of target sesame paste samples under different temperature gradients were collected, a temperature-compensated spectral feature matrix was constructed, and a multi-task partial least squares regression model was used to simultaneously predict the protein, fat and carbohydrate content. The model was then compared and fine-tuned in conjunction with a spectral library of sesame paste standard substances to generate a comprehensive component content estimation vector.
It effectively avoids the interference of temperature fluctuations on spectral characteristics, simplifies the detection process, improves detection accuracy and efficiency, and ensures that the detection results are more consistent with the actual product conditions.
Smart Images

Figure CN122361353A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of food component detection technology, and in particular to a method for detecting the content of components in sesame paste based on spectral analysis. Background Technology
[0002] Sesame paste, a common processed grain food, has protein, fat, and carbohydrate content that are important indicators of product quality. Currently, the industry widely uses near-infrared spectroscopy to detect its component content. Near-infrared spectroscopy is widely used in food component detection due to its advantages such as ease of operation, rapid detection, and non-destructive nature. Conventional detection methods typically collect near-infrared spectral data at a single temperature, construct a spectral feature matrix, and then use a single-component regression model to predict the content of each chemical component.
[0003] In conventional detection techniques, temperature fluctuations can cause shifts in the position and intensity of spectral absorption peaks in sesame paste samples, thus affecting the accuracy of the spectral feature matrix and leading to deviations in the detection results. Furthermore, single-component regression models require separate modeling and prediction for proteins, fats, and carbohydrates, making the detection process cumbersome and failing to consider the correlations between components, resulting in limited prediction accuracy. In addition, conventional methods rely solely on model predictions without incorporating a spectral library of sesame paste standard materials for error calibration or linking them to theoretical values from the product formula, leading to discrepancies between the final detection results and the actual component content, making it difficult to meet the needs of accurate detection. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a method for detecting the content of sesame paste food ingredients based on spectral analysis.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for detecting the content of sesame paste food components based on spectral analysis, comprising: Near-infrared transmission spectral data of the target sesame paste sample under different temperature gradients were collected, and a temperature-compensated spectral feature matrix was constructed based on the near-infrared transmission spectral data. The temperature-compensated spectral feature matrix is input into a multi-task partial least squares regression model, which simultaneously predicts the content of protein, fat and carbohydrates. The predicted content values of each chemical component are extracted from the prediction results of the multi-task partial least squares regression model and compared with the pre-established spectral library of sesame paste standard substances. Based on the comparison results, the latent variable space of the multi-task partial least squares regression model is fine-tuned so that the model's fitting error to the spectral library of sesame paste standard substances converges. Using the fine-tuned multi-task partial least squares regression model, the temperature-compensated spectral feature matrix is recalculated once to obtain a set of corrected predicted chemical composition values. The corrected predicted values of chemical component content are fused with the theoretical values of the sesame paste product formula to generate a comprehensive component content estimation vector. Based on the comprehensive component content estimation vector, a detection result record containing the specific values of each target component is generated.
[0006] As a further aspect of the present invention, a temperature-compensated spectral feature matrix is constructed based on near-infrared transmission spectral data, including: The near-infrared transmission spectral data covers multiple preset wavelength ranges; The near-infrared transmission spectral data are subjected to baseline drift correction and scattering noise suppression processing to eliminate optical interference caused by uneven particle distribution in the sample. The preprocessed near-infrared transmission spectrum data is converted into a three-dimensional spectral tensor, the three dimensions of which are wavelength, temperature and absorbance, respectively. Based on the three-dimensional spectral tensor, a temperature-compensated spectral feature matrix is constructed, which is used to eliminate the influence of temperature changes on the absorption peaks of specific chemical components. The near-infrared transmission spectral data undergoes baseline drift correction and scattering noise suppression processing, specifically including: For each wavelength point in the near-infrared transmission spectrum data, calculate its mean and variance of absorbance over the entire temperature gradient range; Using the average absorbance as a baseline, a smooth baseline curve is fitted using the asymmetric least squares method, and the baseline curve is subtracted from the original spectral data to obtain baseline-free spectral data. The baseline-de-baseline spectral data are transformed using a standard normal variable to eliminate light scattering differences caused by sample surface roughness. The first derivative operation is used to process the spectral data after standard normal variable transformation to enhance the characteristics of subtle absorption peaks caused by changes in component concentration; The processed spectral data is bound to the original temperature gradient labels to form a spectral correction dataset with physical meaning.
[0007] As a further aspect of the present invention, the step of converting the preprocessed near-infrared transmission spectral data into a three-dimensional spectral tensor includes: Determine the effective wavelength range corresponding to the near-infrared transmission spectral data, and uniformly divide the effective wavelength range into several continuous spectral sub-bands; The spectral subbands, the temperature gradient, and the corresponding absorbance values are respectively mapped to the three coordinate axes of the three-dimensional spectral tensor; For each spectral sub-band, within its corresponding wavelength range, the arithmetic mean of all absorbance data points within the wavelength range is taken to generate a representative feature value; The generated representative feature values are filled into the slice positions in the three-dimensional spectral tensor corresponding to the wavelength range; After processing all wavelength ranges, a complete three-dimensional spectral tensor is obtained that reflects the spectral characteristics as a function of temperature.
[0008] As a further aspect of the present invention, a temperature-compensated spectral feature matrix is constructed based on the three-dimensional spectral tensor, including: In the three-dimensional spectral tensor, a reference temperature point is selected as the reference temperature; Calculate the difference spectrum between the spectral data at each of the remaining temperature points and the spectral data at the reference temperature; Principal component analysis was performed on the differential spectrum to extract several principal component loading vectors that could best characterize the temperature change. A temperature correction operator is constructed using the extracted principal component load vector, and the three-dimensional spectral tensor is convolved with the temperature correction operator. The result of the convolution operation is rearranged into a two-dimensional matrix structure, which is the temperature-compensated spectral feature matrix that eliminates temperature dependence.
[0009] As a further aspect of the present invention, the temperature-compensated spectral feature matrix is input into a multi-task partial least squares regression model, which simultaneously predicts the content of protein, fat, and carbohydrates, including: The temperature-compensated spectral feature matrix is split into two parts: a training set and a validation set. In the training set, latent variable spaces are constructed for the three prediction tasks of protein, fat and carbohydrate, respectively, and regularization constraints for cross-tasks are introduced during the construction process. The optimal projection direction of the latent variables is solved iteratively, so that the spectral features have the maximum explanatory variance of the response variables of the three components in the latent variable space. Using the projection direction obtained by the solution, the spectral features of the validation set are mapped into the latent variable space, and the residual between the predicted value and the true value is calculated. When the decrease in the residual over three consecutive iterations is less than a set threshold, the iteration is stopped and the parameters of the multi-task partial least squares regression model are fixed.
[0010] As a further aspect of the present invention, the predicted content values of each chemical component are extracted from the prediction results of the multi-task partial least squares regression model, and compared with a pre-established spectral library of sesame paste standard substances, including: Read the three sets of prediction vectors output by the multi-task partial least squares regression model, which correspond to protein, fat and carbohydrate respectively; For each data point in the three sets of prediction vectors, calculate the cosine similarity with its corresponding standard spectral fingerprint in the sesame paste standard material spectral library. Select sample data points with a cosine similarity lower than a preset matching threshold and mark them as suspicious data points; The distribution of the suspicious data points in the overall dataset is statistically analyzed to determine whether they are concentrated in specific chemical components or specific temperature ranges. The analysis results are used as weighting factors for subsequent model fine-tuning, and are used to apply differentiated penalties to different types of errors during the fine-tuning process.
[0011] As a further aspect of the present invention, the step of fine-tuning the latent variable space of the multi-task partial least squares regression model based on the comparison results includes: Based on the distribution of suspicious data points in the comparison results, a dynamic weight matrix is calculated; The dynamic weight matrix is applied to the loss function of the multi-task partial least squares regression model to amplify the error contribution of the suspicious data points. In the latent variable space, a gradient ascent correction is performed on the projection matrix along the direction of increasing error to increase the model's ability to distinguish boundary samples. After the correction is completed, the covariance matrix between each latent variable in the latent variable space is recalculated to ensure that it still satisfies the orthogonality constraint. Repeat the correction and verification steps until the norm of the dynamic weight matrix converges to a stable value.
[0012] As a further aspect of the present invention, the step of recalculating the temperature-compensated spectral feature matrix using the fine-tuned multi-task partial least squares regression model includes: Load the already fine-tuned projection matrix and regression coefficient matrix; Each sample vector in the temperature-compensated spectral feature matrix is sequentially projected into the fine-tuned latent variable space; Perform linear regression in the latent variable space to obtain a set of linear predictions without inverse transformation; A nonlinear correction term derived from a standard substance spectral library is applied to the linear prediction value to compensate for the nonlinear deviation between the spectral response and the actual chemical content. The result after nonlinear correction is used as the predicted value of the corrected chemical component content.
[0013] As a further aspect of the present invention, the step of integrating the corrected predicted values of chemical component content with the theoretical values of the sesame paste product formula includes: The theoretical formula values corresponding to the batch of products to be tested are retrieved from the production database of sesame paste products. The theoretical formula values include the theoretical feed ratio of each raw material and its theoretical nutrient content. Calculate the difference vector between the corrected predicted chemical component content and the theoretical value of the formula; The difference vector is smoothed by a Kalman filter to filter out high-frequency noise introduced by fluctuations in a single measurement. The filtered difference vectors are weighted proportionally and superimposed back onto the theoretical formula values to generate the comprehensive component content estimation vector that integrates measured information and prior knowledge.
[0014] As a further aspect of the present invention, the step of generating a detection result record containing specific values of each target component based on the comprehensive component content estimation vector includes: Each component in the comprehensive component content estimation vector is formatted and converted according to national standard units of measurement; Each target component is assigned a data quality identifier, which is calculated based on its corresponding spectral signal-to-noise ratio and model confidence. According to the preset report template, fill in the formatted values and corresponding data quality identifiers into the specified table fields; Add a metadata description at the end of the report, recording the spectral wavelength range, temperature gradient settings, and model version number used in this test; All content is packaged into an immutable electronic document, which becomes the final record of the test results.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: Near-infrared transmission spectral data of the target sesame paste sample were collected under different temperature gradients, and a temperature-compensated spectral feature matrix was constructed based on the near-infrared transmission spectral data. By capturing the spectral variation patterns at different temperatures and performing temperature compensation correction on the spectral data, the interference of temperature fluctuations on spectral features can be effectively avoided, and problems such as spectral absorption peak shift and intensity distortion caused by temperature factors can be eliminated. This allows the spectral feature matrix to more accurately reflect the compositional information of the sesame paste sample, making the spectral data more stable and reliable. Compared with conventional spectral acquisition methods at a single temperature, this method can reduce detection bias caused by temperature.
[0016] A multi-task partial least squares regression model was used to simultaneously predict the content of protein, fat, and carbohydrates. The model's latent variable space was fine-tuned based on comparisons with a spectral library of sesame paste standard substances. The fine-tuned model was then used to recalculate the temperature-compensated spectral feature matrix to obtain corrected predicted chemical component content values. These corrected predicted values were then fused with the theoretical values of the sesame paste product formula to generate a comprehensive component content estimation vector. The multi-task regression model enables simultaneous detection of multiple components without the need for separate individual models, simplifying the detection process. Fine-tuning the model's latent variable space using a spectral library of standard substances reduces model fitting errors and improves the accuracy of prediction results. Fusing the theoretical values of the formula further calibrates the prediction results, overcoming the limitations of single-model prediction and standard library comparisons, making the component content estimation more closely reflect the actual product situation. Compared to conventional single-component detection and uncalibrated prediction methods, this approach optimizes detection efficiency and accuracy. Attached Figure Description
[0017] Figure 1 This is a flowchart of the method for detecting the content of sesame paste food ingredients based on spectral analysis according to the present invention; Figure 2 A flowchart for converting preprocessed spectral data into a three-dimensional spectral tensor; Figure 3 A flowchart for prediction and training of a multi-task partial least squares regression model; Figure 4 The convergence curve of the norm of the dynamic weight matrix; Figure 5 To compare the absorbance stability before and after temperature compensation. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0019] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0020] See Figure 1Near-infrared transmission spectral data of the target sesame paste sample were collected at different preset temperature gradients. A temperature-compensated spectral feature matrix was constructed based on this data to eliminate the influence of temperature fluctuations. The constructed temperature-compensated spectral feature matrix was input into a pre-trained multi-task partial least squares regression model, which simultaneously outputs preliminary predicted values for the content of three target components: protein, fat, and carbohydrates. These preliminary predicted values were compared and analyzed with a pre-established sesame paste standard material spectral library containing known standard material spectra and their precise component contents. Based on the data deviation characteristics found in the comparison, the latent variable spatial parameters within the multi-task partial least squares regression model were fine-tuned to bring the model's fitting error to the standard material spectral library to a convergent state. Using the fine-tuned and optimized model, the original temperature-compensated spectral feature matrix was recalculated to obtain a set of corrected and more accurate predicted values for chemical component content. These corrected predicted values were then fused with the theoretical formula values of the batch of sesame paste obtained from the production process to generate a comprehensive component content estimation vector. Based on this estimation vector, the final detection result record is generated, including the specific quantitative values and quality indicators of each target component.
[0021] In one embodiment of the present invention, in an exemplary scenario, for a sample of black sesame paste powder, five temperature gradient points are set, with temperatures of 20°C, 30°C, 40°C, 50°C, and 60°C. A near-infrared spectrometer equipped with a temperature-controlled sample cell is used to collect near-infrared transmission spectra at each temperature point within a wavelength range of 900 nm to 1700 nm. Each spectrum contains 1024 wavelength data points, forming an original near-infrared transmission spectrum dataset. Baseline drift correction and scattering noise suppression are performed on the near-infrared transmission spectrum dataset. In a specific implementation, for each wavelength point in the near-infrared transmission spectrum data, the mean and variance of absorbance are calculated across the entire five temperature gradient points. Using the calculated mean absorbance as a baseline, a smooth baseline curve is fitted using the asymmetric least squares method. The fitting process is achieved by minimizing a weighted sum of squared residuals, and its objective function is expressed as: , in: The original spectrum is represented in the first... The absorbance value at each wavelength point The baseline curve to be determined is in the th... The value at each wavelength point It is a second-order difference operator. The smoothing parameter controls the degree of smoothness of the baseline. It uses asymmetric weights. Calculate the weights. The rule is that when When, assign a smaller weight, when At that time, a larger weight is assigned, and the objective function is minimized through iterative solution to obtain the baseline curve. The raw spectral data Subtract baseline curve The baseline-removed spectral data is obtained. A standard normal transformation (SNR) is applied to the baseline-removed spectral data. The SNR operates on each sample spectral vector, calculating the mean and standard deviation of the absorbance at all wavelengths. Then, the absorbance value at each wavelength is subtracted from the mean and divided by the standard deviation. First-order derivative operations are used to process the SNR-Gore filter to enhance the characteristics of subtle absorption peaks caused by variations in component concentration. The processed spectral data is then bound to the original temperature gradient labels to form a physically meaningful spectral correction dataset.
[0022] In some embodiments, the spectral correction dataset is converted into a three-dimensional spectral tensor, whose three dimensions are defined as wavelength, temperature, and absorbance, respectively. In a specific implementation, the effective wavelength range corresponding to the near-infrared transmission spectral data is determined to be 950 nm to 1650 nm. The effective wavelength range of 700 wavelengths is uniformly divided into 70 continuous spectral sub-bands, each covering 10 nm. The 70 spectral sub-bands, 5 temperature gradient points, and the representative absorbance feature value calculated within each sub-band are mapped to the three coordinate axes of the three-dimensional spectral tensor. For each spectral sub-band, within its corresponding 10 nm wavelength interval, the arithmetic mean of all absorbance data points within that interval is taken to generate a representative feature value. The generated 70 representative feature values are filled into the slice positions in the three-dimensional spectral tensor corresponding to the 70 wavelength intervals, with the temperature dimension corresponding to 5 slices. After processing all wavelength intervals, a complete three-dimensional spectral tensor is obtained, with dimensions of 70 x 5 x 1. This three-dimensional spectral tensor reflects the variation of spectral characteristics with temperature.
[0023] A temperature-compensated spectral feature matrix is constructed based on a three-dimensional spectral tensor. In specific implementation, 20 degrees Celsius is selected as the reference temperature point in the three-dimensional spectral tensor. The difference spectra between the spectral data at the other four temperature points (30 degrees Celsius, 40 degrees Celsius, 50 degrees Celsius, and 60 degrees Celsius) and the spectral data at the 20-degree Celsius reference temperature are calculated, resulting in four sets of difference spectral vectors. Principal component analysis is performed on the four sets of difference spectral vectors to extract two principal component loading vectors that best characterize temperature changes. Using the two extracted principal component loading vectors, a temperature correction operator is constructed, which is a projection matrix. A convolution operation is performed between the 70x5x1 three-dimensional spectral tensor and the temperature correction operator, with the convolution operation unfolded along the temperature dimension. The result of the convolution operation is rearranged into a two-dimensional matrix structure, which is the temperature-compensated spectral feature matrix with eliminated temperature dependence and a size of 70x5. It is understandable that the temperature-compensated spectral feature matrix will be used as input for the subsequent multi-task regression model, with each row corresponding to different spectral sub-band features and each column corresponding to the corrected spectral response at different temperatures.
[0024] In one embodiment of the present invention, see [reference] Figure 2 The process of converting preprocessed near-infrared transmission spectral data into a three-dimensional spectral tensor involves explicit data mapping and reconstruction. Based on a spectral correction dataset of a black sesame paste sample collected at five temperature gradients and with baseline correction and scattering suppression completed, the effective wavelength range corresponding to the spectral correction dataset was determined to be 950 nm to 1650 nm. The 700 nm wide effective wavelength range was uniformly divided into 70 continuous spectral sub-bands, each covering a 10 nm wavelength interval. The 70 spectral sub-bands, the 5 temperature gradients, and the representative absorbance feature values calculated within each preprocessed sub-band were mapped to the three coordinate axes of the three-dimensional spectral tensor, where the first coordinate axis index corresponds to the spectral sub-band number, the second coordinate axis index corresponds to the temperature gradient number, and the third coordinate axis index stores the corresponding absorbance feature value. For each spectral sub-band, within its corresponding 10 nm wavelength interval, the arithmetic mean of all absorbance data points within this interval was taken to generate a representative feature value, calculated using the following formula: , in: Representing the Representative absorbance characteristic values of each spectral subband The representative fell into the first The number of wavelength points within each sub-band Representing the Within the first generation The absorbance values at each wavelength point are calculated. The generated 70 representative feature values are then sequentially filled into the slice positions corresponding to the 70 wavelength intervals in the three-dimensional spectral tensor, with 5 slices corresponding to the temperature dimension. After filling all wavelength intervals, a complete three-dimensional spectral tensor is obtained. The size of the three-dimensional spectral tensor is represented as 70 (number of spectral subbands) multiplied by 5 (number of temperature gradients) multiplied by 1 (feature depth). This three-dimensional spectral tensor structures the characteristic information of spectral changes with temperature.
[0025] A temperature-compensated spectral feature matrix is constructed based on a three-dimensional spectral tensor. In some embodiments, spectral data at 20 degrees Celsius is selected as the reference. In a specific implementation, the data slice corresponding to 20 degrees Celsius is selected as the reference temperature point in the three-dimensional spectral tensor. The difference spectra between the spectral data slices at the other four temperature points (30 degrees Celsius, 40 degrees Celsius, 50 degrees Celsius, and 60 degrees Celsius) and the 20-degree Celsius reference temperature data slice are calculated, and the difference spectral vectors are obtained. We obtain the following by subtracting element by element: , in: This represents the difference spectral vector at temperature t. This represents the spectral vector in the three-dimensional spectral tensor at temperature t. This represents the spectral vector at a reference temperature of 20 degrees Celsius, with t taking values of 30, 40, 50, and 60. Optionally, principal component analysis (PCA) is performed on the four sets of differential spectral vectors. PCA aims to extract loading vectors that best characterize the temperature change pattern. The covariance matrix of the matrix formed by the four sets of differential spectral vectors is calculated, and eigenvalue decomposition is performed on the covariance matrix. The two eigenvectors with the largest eigenvalues are extracted as principal component loading vectors. A temperature correction operator is constructed using the two extracted principal component loading vectors. The temperature correction operator is a projection matrix used to subtract the main effects of temperature change from the original spectral data. A 70x5x1 three-dimensional spectral tensor is convolved with the temperature correction operator. The convolution operation is performed along the temperature dimension (the second dimension) of the three-dimensional spectral tensor. The result of the convolution operation can be understood as a new three-dimensional data array. The result of the convolution operation is rearranged into a two-dimensional matrix structure. The rearrangement process involves expanding the temperature dimension and the spectral subband dimension. This two-dimensional matrix is the temperature-compensated spectral feature matrix with temperature dependence eliminated.
[0026] In one embodiment of the present invention, inputting the temperature-compensated spectral feature matrix into a multi-task partial least squares regression model and comparing the prediction results with a standard material spectral library constitutes a complete modeling and verification process. In an example scenario, the constructed temperature-compensated spectral feature matrix contains data from 120 sesame paste samples, each with a spectral feature dimension of 85, corresponding to the corrected absorbance of 85 spectral sub-bands. This matrix is used as the input to the multi-task partial least squares regression model. For specific implementation details, please refer to [reference needed]. Figure 3 The temperature-compensated spectral feature matrix was split into a training set and a validation set. A stratified random sampling method was used, allocating 96 samples to the training set and 24 samples to the validation set to ensure that the distribution ranges of protein, fat, and carbohydrate content were largely consistent between the two sets. In the training set, independent latent variable spaces were constructed for the three prediction tasks: protein, fat, and carbohydrates. Each latent variable space was initialized with 5 latent variables, and orthogonal regularization constraints based on cross-tasks were introduced during the construction process. The constraint term is as follows: , in: Represents the regularization penalty term. and Representing tasks respectively and tasks The latent variable projection direction matrix, Denotes the square of the Frobenius norm of the matrix. This is a hyperparameter controlling the regularization strength, set to 0.01. The optimal projection direction of the latent variables is solved iteratively using a nonlinear conjugate gradient method to maximize the explained variance of the spectral features in the latent variable space for the response variables of the three components. The projection direction matrices for all three tasks are updated simultaneously in each iteration. Using the solved projection direction matrices, the spectral features of the 24 samples in the validation set are mapped to the latent variable space, and the residuals between the predicted values and the true values measured by standard chemical methods are calculated. The residuals are measured as root mean square error (RMSE). A threshold of 0.5% for the decrease in residuals is set. When the decrease in RMSE is less than 0.5% in three consecutive iterations, the iteration stops, and the parameters of the multi-task partial least squares regression model, including the final projection direction matrix and regression coefficient matrix, are solidified.
[0027] In some embodiments, the predicted content values of each chemical component are extracted from the prediction results of a trained multi-task partial least squares regression model and compared with a pre-established spectral library of sesame paste standard substances. Specifically, the three sets of prediction vectors output by the multi-task partial least squares regression model for 24 samples in the validation set are read. These vectors correspond to the predicted protein content, fat content, and carbohydrate content, respectively, with each vector having a length of 24. For each data point in the three sets of prediction vectors, a cosine similarity calculation is performed with its corresponding standard spectral fingerprint in the sesame paste standard substance spectral library. The sesame paste standard substance spectral library contains the spectral characteristics of 50 different proportions of sesame paste standard substances and their component contents certified by authoritative institutions. When calculating the cosine similarity, the predicted content vector of the sample (containing the predicted values of protein, fat, and carbohydrates) is multiplied by the standard content vector of the corresponding substance in the standard substance spectral library, and then divided by the product of their moduli. A preset matching threshold is set to 0.92, and sample data points with a cosine similarity lower than 0.92 are filtered out and marked as suspicious data points. The distribution of suspicious data points across the 24 validation set samples was statistically analyzed to determine if they were concentrated in specific chemical compositions or temperature ranges. For example, it was found that 5 out of 8 suspicious data points corresponded to fat content predictions, and 4 of these 5 samples were collected at temperatures above 45 degrees Celsius. This analysis result was quantified into a weighted factor dictionary, which was used to apply greater penalty weights to the errors generated by the fat prediction task and samples from high-temperature ranges during subsequent model fine-tuning.
[0028] In one embodiment of the present invention, the latent variable space of the multi-task partial least squares regression model is fine-tuned based on the comparison results. The implementation process involves dynamic weight adjustment and parameter optimization based on the error distribution. In an example scenario, following the analysis results of the embodiment, the multi-task partial least squares regression model generated 8 suspicious data points with a cosine similarity lower than 0.92 on the validation set, as shown in Table 1.
[0029] Table 1: Example of Distribution of Suspicious Data Points and Dynamic Weight Calculation
[0030] In practice, a dynamic weight matrix is calculated based on the distribution of suspicious data points in the comparison results. Each element of the dynamic weight matrix corresponds to the error weight of a training sample for a prediction task. Its calculation is determined by whether the sample's component category and its collected temperature range fall within the "concentrated occurrence temperature range" in Table 1. Samples falling within this range are assigned the corresponding initial weight factor in Table 1 for their respective component task; otherwise, the weight factor is 1.0. The dynamic weight matrix is applied to the loss function of the multi-task partial least squares regression model. The original loss function is a simple sum of the mean squared errors of the predictions for each task. The loss function after introducing the dynamic weight matrix is modified as follows: , in: Represents the weighted total loss. Represents the total number of training samples. Representing the The sample at the th The actual content value of each component in the task. This represents the corresponding predicted value. It is the corresponding number in the dynamic weight matrix. The first sample and the first The weights of each task (protein, fat, and carbohydrate, corresponding to k=1, 2, and 3 respectively). For example, for a sample belonging to the fat category and at a temperature of 45 degrees Celsius, its weights on the fat task (k=2) are... The value is 1.8. This operation amplifies the contribution of errors from suspicious data points to the total loss. In the latent variable space, a gradient ascent correction is performed on the projection matrix along the direction of increasing error. The step size of the gradient ascent correction is set to 0.1 times the original learning rate to increase the model's ability to distinguish the aforementioned boundary samples. After the correction, the covariance matrix between each latent variable in the latent variable space is recalculated to ensure that it still satisfies the orthogonality constraint. Orthogonality is determined by verifying whether the dot product of any two latent variable score vectors is close to zero. The correction and verification steps are repeated. After each iteration, the norm of the dynamic weight matrix is re-evaluated until the fluctuation of the norm of the dynamic weight matrix in five consecutive iterations is less than one-thousandth. At this point, the norm of the dynamic weight matrix is considered to have converged to a stable value, and the fine-tuning process stops.
[0031] In some embodiments, the temperature-compensated spectral feature matrix is recalculated using a fine-tuned multi-task partial least squares regression model, a deterministic computational process. Specifically, the fine-tuned projection matrix and regression coefficient matrix are loaded from storage; the projection matrix has a dimension of 85 x 5, and the regression coefficient matrix has a dimension of 5 x 3. Each sample vector in the temperature-compensated spectral feature matrix is sequentially projected into the fine-tuned latent variable space. Projection is achieved through multiplication of the sample feature vector with the projection matrix. Linear regression is performed in the latent variable space, multiplying the latent variable score vector of the sample with the regression coefficient matrix to obtain a set of linear predicted values without inverse transformation. Each linear predicted value is a vector containing the content of the three components. A nonlinear correction term derived from a standard substance spectral library is applied to the linear predicted values. This nonlinear correction term is a mapping based on a radial basis function network, taking the original spectral feature vector of the sample as input and outputting the correction amount for the three components to compensate for the nonlinear deviation between the spectral response and the actual chemical content. The result after nonlinear correction, i.e., the sum of the linear prediction value and the nonlinear correction amount, is used as the final corrected prediction value of chemical composition content.
[0032] See Figure 4 In the fine-tuning iteration phase of a multi-task partial least squares regression model, the convergence behavior of the dynamic weight matrix norm intuitively reflects the model's optimization process for boundary samples. As shown in the figure, the solid line represents the trend of the dynamic weight matrix norm with the number of iterations, and the dashed line represents the preset convergence threshold. In the early iterations (1-6 iterations), the weight matrix norm rapidly decreases from its initial value of 1.80 to 1.60. This stage corresponds to the model's focused correction of suspicious data points, amplifying their error contribution to prompt rapid adjustment of the projection matrix in the latent variable space, thereby enhancing its ability to distinguish boundary samples. As the number of iterations increases (7-11 iterations), the rate of decrease in the norm slows significantly, gradually approaching the convergence threshold. This indicates that the optimization space of the model parameters gradually narrows, and the ability to distinguish suspicious data tends to stabilize. When the number of iterations exceeds 11, the weight matrix norm basically stabilizes near the convergence threshold, with a fluctuation range of less than one-thousandth, satisfying the preset convergence condition, marking the completion of the model fine-tuning process. This convergence process ensures that the dynamic weight matrix amplifies the contribution of errors in questionable data without introducing excessive parameter perturbations, thereby maintaining the orthogonality constraint of the latent variable space and ultimately improving the robustness and accuracy of the model in predicting the content of sesame paste food ingredients.
[0033] In one embodiment of the present invention, the corrected predicted values of chemical component content are fused with the theoretical values of the sesame paste product formula to generate a final test result record. The operation process is clear and repeatable. Continuing with the aforementioned embodiment, there is a batch of black sesame paste product with the serial number SH20260315A. The corrected predicted values of chemical component content are calculated by a finely tuned multi-task partial least squares regression model. For a specific sample in this batch, the predicted values are: protein content 12.8 g / 100 g, fat content 9.5 g / 100 g, and carbohydrate content 72.3 g / 100 g. In practice, the theoretical formula values for batch number SH20260315A of sesame paste were retrieved from the production database. The database records show that the theoretical ingredient ratio for this batch is 60% black sesame, 30% rice, and 10% white sugar. Based on the nutritional information table, the theoretical nutrient content is calculated to be: 13.2 g / 100g protein, 9.8 g / 100g fat, and 71.5 g / 100g carbohydrates. The difference vector between the corrected predicted chemical composition and the theoretical formula values is calculated as follows: Δ = [12.8 - 13.2, 9.5 - 9.8, 72.3 - 71.5] = [-0.4, -0.3, 0.8], in grams per 100g. The difference vector is smoothed by a Kalman filter, which filters out high-frequency noise introduced by fluctuations in a single measurement. Its state vector contains the differences in protein, fat, and carbohydrate content. The observation vector is the directly calculated difference vector. The filter is recursively updated based on the covariance between the prediction and observation. The state update formula is: , in: This represents the filtered state estimate at time k (i.e., the smoothed difference vector). This represents the prior state estimate at time k. The Kalman gain matrix at time k. This represents the observed value at time k (i.e., the difference vector obtained by direct calculation). This is the observation matrix, which is the identity matrix in this case.
[0034] See Figure 5In the experiment comparing absorbance stability before and after temperature compensation (wavelength band 1), near-infrared spectroscopy was used to evaluate the effect of temperature change on the absorbance of sesame paste samples and the compensation effect. Specifically, near-infrared transmission spectral data of the samples were collected under a continuous temperature gradient from 25℃ to 60℃. Wavelength band 1 was selected as the characteristic analysis interval, and the absorbance sequences before and after compensation were calculated. Before compensation (dashed line), the absorbance fluctuated significantly with increasing temperature, rapidly decreasing from 0.90 to 0.58 in the 25℃ to 45℃ range, and then rebounding to 0.75 in the 45℃ to 60℃ range, exhibiting a strong temperature dependence. After compensation (solid line), the absorbance stabilized within a narrow range of 0.79 to 0.82, with a significantly narrowed fluctuation range, effectively eliminating the interference of temperature changes on absorbance. The core technology of the experiment relies on the construction of a three-dimensional spectral tensor and a temperature correction operator. First, the preprocessed spectral data is converted into a three-dimensional tensor with wavelength, temperature, and absorbance as dimensions. A reference temperature point is selected to calculate the difference spectrum. Principal component analysis is used to extract the loading vector representing temperature changes. A temperature correction operator is then constructed and convolved with the tensor. Finally, the convolution result is rearranged into a two-dimensional feature matrix that eliminates temperature dependence. During parameter configuration, the cumulative variance contribution rate threshold for principal component analysis is set to 95%, and the convolution kernel size of the temperature correction operator is matched with the temperature gradient step size to ensure the stability and effectiveness of the compensation effect.
[0035] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for detecting the content of sesame paste food components based on spectral analysis, characterized in that, include: Near-infrared transmission spectral data of the target sesame paste sample under different temperature gradients were collected, and a temperature-compensated spectral feature matrix was constructed based on the near-infrared transmission spectral data. The temperature-compensated spectral feature matrix is input into a multi-task partial least squares regression model, which simultaneously predicts the content of protein, fat and carbohydrates. The predicted content values of each chemical component are extracted from the prediction results of the multi-task partial least squares regression model and compared with the pre-established spectral library of sesame paste standard substances. Based on the comparison results, the latent variable space of the multi-task partial least squares regression model is fine-tuned so that the fitting error of the model to the spectral library of sesame paste standard substances converges. Using the fine-tuned multi-task partial least squares regression model, the temperature-compensated spectral feature matrix is recalculated once to obtain a set of corrected predicted chemical composition values. The corrected predicted values of chemical component content are fused with the theoretical values of the sesame paste product formula to generate a comprehensive component content estimation vector. Based on the comprehensive component content estimation vector, a detection result record containing the specific values of each target component is generated.
2. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 1, characterized in that, Based on near-infrared transmission spectral data, a temperature-compensated spectral feature matrix is constructed, including: The near-infrared transmission spectral data covers multiple preset wavelength ranges; The near-infrared transmission spectral data are subjected to baseline drift correction and scattering noise suppression processing to eliminate optical interference caused by uneven particle distribution in the sample. The preprocessed near-infrared transmission spectrum data is converted into a three-dimensional spectral tensor, the three dimensions of which are wavelength, temperature and absorbance, respectively. Based on the three-dimensional spectral tensor, a temperature-compensated spectral feature matrix is constructed, which is used to eliminate the influence of temperature changes on the absorption peaks of specific chemical components. The near-infrared transmission spectral data undergoes baseline drift correction and scattering noise suppression processing, specifically including: For each wavelength point in the near-infrared transmission spectrum data, calculate its mean and variance of absorbance over the entire temperature gradient range; Using the average absorbance as a baseline, a smooth baseline curve is fitted using the asymmetric least squares method, and the baseline curve is subtracted from the original spectral data to obtain baseline-free spectral data. The baseline-de-baseline spectral data are transformed using a standard normal variable to eliminate light scattering differences caused by sample surface roughness. The first derivative operation is used to process the spectral data after standard normal variable transformation to enhance the characteristics of subtle absorption peaks caused by changes in component concentration; The processed spectral data is bound to the original temperature gradient labels to form a spectral correction dataset with physical meaning.
3. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 2, characterized in that, The step of converting the preprocessed near-infrared transmission spectral data into a three-dimensional spectral tensor includes: Determine the effective wavelength range corresponding to the near-infrared transmission spectral data, and uniformly divide the effective wavelength range into several continuous spectral sub-bands; The spectral subbands, the temperature gradient, and the corresponding absorbance values are respectively mapped to the three coordinate axes of the three-dimensional spectral tensor; For each spectral sub-band, within its corresponding wavelength range, the arithmetic mean of all absorbance data points within the wavelength range is taken to generate a representative feature value; The generated representative feature values are filled into the slice positions in the three-dimensional spectral tensor corresponding to the wavelength range; After processing all wavelength ranges, a complete three-dimensional spectral tensor is obtained that reflects the spectral characteristics as a function of temperature.
4. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 3, characterized in that, Based on the aforementioned three-dimensional spectral tensor, a temperature-compensated spectral feature matrix is constructed, including: In the three-dimensional spectral tensor, a reference temperature point is selected as the reference temperature; Calculate the difference spectrum between the spectral data at each of the remaining temperature points and the spectral data at the reference temperature; Principal component analysis was performed on the differential spectrum to extract several principal component loading vectors that could best characterize the temperature change. Using the extracted principal component load vector, a temperature correction operator is constructed, and the three-dimensional spectral tensor is convolved with the temperature correction operator. The result of the convolution operation is rearranged into a two-dimensional matrix structure, which is the temperature-compensated spectral feature matrix that eliminates temperature dependence.
5. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 4, characterized in that, The temperature-compensated spectral feature matrix is input into a multi-task partial least squares regression model, which simultaneously predicts the content of protein, fat, and carbohydrates, including: The temperature-compensated spectral feature matrix is split into two parts: a training set and a validation set. In the training set, latent variable spaces are constructed for the three prediction tasks of protein, fat and carbohydrate, respectively, and regularization constraints for cross-tasks are introduced during the construction process. The optimal projection direction of the latent variables is solved iteratively, so that the spectral features have the maximum explanatory variance of the response variables of the three components in the latent variable space. Using the projection direction obtained by the solution, the spectral features of the validation set are mapped into the latent variable space, and the residual between the predicted value and the true value is calculated. When the decrease in the residual over three consecutive iterations falls below a set threshold, the iteration is stopped and the parameters of the multi-task partial least squares regression model are fixed.
6. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 5, characterized in that, The predicted content values of each chemical component are extracted from the prediction results of the multi-task partial least squares regression model and compared with a pre-established spectral library of sesame paste standard substances, including: Read the three sets of prediction vectors output by the multi-task partial least squares regression model, which correspond to protein, fat and carbohydrate respectively; For each data point in the three sets of prediction vectors, calculate the cosine similarity with its corresponding standard spectral fingerprint in the sesame paste standard material spectral library. Select sample data points with a cosine similarity lower than a preset matching threshold and mark them as suspicious data points; The distribution of the suspicious data points in the overall dataset is statistically analyzed to determine whether they are concentrated in specific chemical components or specific temperature ranges. The analysis results are used as weighting factors for subsequent model fine-tuning, and are used to apply differentiated penalties to different types of errors during the fine-tuning process.
7. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 6, characterized in that, The step of fine-tuning the latent variable space of the multi-task partial least squares regression model based on the comparison results includes: Based on the distribution of suspicious data points in the comparison results, a dynamic weight matrix is calculated; The dynamic weight matrix is applied to the loss function of the multi-task partial least squares regression model to amplify the error contribution of the suspicious data points. In the latent variable space, a gradient ascent correction is performed on the projection matrix along the direction of increasing error to increase the model's ability to distinguish boundary samples. After the correction is completed, the covariance matrix between each latent variable in the latent variable space is recalculated to ensure that it still satisfies the orthogonality constraint. Repeat the correction and verification steps until the norm of the dynamic weight matrix converges to a stable value.
8. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 7, characterized in that, The step of recalculating the temperature-compensated spectral feature matrix using the fine-tuned multi-task partial least squares regression model includes: Load the already fine-tuned projection matrix and regression coefficient matrix; Each sample vector in the temperature-compensated spectral feature matrix is sequentially projected into the fine-tuned latent variable space; Perform linear regression in the latent variable space to obtain a set of linear predictions without inverse transformation; A nonlinear correction term derived from a standard substance spectral library is applied to the linear prediction value to compensate for the nonlinear deviation between the spectral response and the actual chemical content. The result after nonlinear correction is used as the predicted value of the corrected chemical component content.
9. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 8, characterized in that, The process of integrating the revised predicted values of chemical component content with the theoretical values of the sesame paste product formula includes: The theoretical formula values corresponding to the batch of products to be tested are retrieved from the production database of sesame paste products. The theoretical formula values include the theoretical feed ratio of each raw material and its theoretical nutrient content. Calculate the vector difference between the corrected predicted chemical component content and the theoretical value of the formula; The difference vector is smoothed by a Kalman filter to filter out high-frequency noise introduced by fluctuations in a single measurement. The filtered difference vectors are weighted proportionally and superimposed back onto the theoretical formula values to generate the comprehensive component content estimation vector that integrates measured information and prior knowledge.
10. The method for detecting the content of sesame paste food components based on spectral analysis as described in claim 9, characterized in that, Based on the comprehensive component content estimation vector, a detection result record containing the specific values of each target component is generated, including: Each component in the comprehensive component content estimation vector is formatted and converted according to national standard units of measurement; Each target component is assigned a data quality identifier, which is calculated based on its corresponding spectral signal-to-noise ratio and model confidence. According to the preset report template, fill in the formatted values and corresponding data quality identifiers into the specified table fields; Add a metadata description at the end of the report, recording the spectral wavelength range, temperature gradient settings, and model version number used in this test; All content is packaged into an immutable electronic document, which becomes the final record of the test results.