Food additive identification method based on sparse features of surface enhanced Raman spectroscopy

Through compression perception theory, the sparse characteristics of the surface-enhanced Raman spectrum are extracted and optimized, and the identification model is established in combination with multiple algorithms, which solves the problems of high spectral similarity and high feature overlap, and achieves high-precision recognition of food additives.

CN120340652APending Publication Date: 2025-07-18SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510266785.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing surface-enhanced Raman spectroscopy technology has problems with high spectral shape similarity and high feature overlap in food additive recognition, resulting in insufficient recognition accuracy. Common dimensionality reduction algorithms such as PCA and t-SNE have insufficient feature extraction and model stability.

Method used

Using compression perception theory, by designing appropriate transformation basis and sensing matrix, the sparse features of the surface-enhanced Raman spectrum are extracted and optimized, and recognition models are established using sparse features, and training is combined with proximity algorithms, support vector machines, random forests and convolutional neural networks.

Benefits of technology

It improves the accuracy of food additive identification and the operating efficiency of the model, and can accurately judge similar food additives and achieve high-precision judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340652A_ABST
    Figure CN120340652A_ABST
Patent Text Reader

Abstract

The invention provides a food additive identification method based on sparse features of a surface enhanced Raman spectrum. The method comprises the following steps: acquiring SERS spectrums of food additive samples of known types; according to the SERS spectrum, SERS spectrum sparse features are extracted, and the sparse features are optimized; dividing the optimized sparse features into a training set and a test set, establishing an identification model by using the training set, and verifying the identification accuracy of the identification model by using the test set; and performing type identification on the food additives in a to-be-detected sample by adopting the identification model. According to the method, the recognition model is established based on the SERS spectrum sparse features, the similar food additives can be accurately judged, the accuracy of additive type recognition is improved, and high-precision judgment of the food additives is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of food additive identification, and specifically, to a method for identifying food additives based on sparse features of surface-enhanced Raman spectroscopy. Background Art

[0002] Food additives can improve food quality, extend the shelf life, enhance taste and visual effects, etc., and are widely used in modern food processing industries. However, long-term or excessive intake of certain food additives can pose serious health risks to humans, such as allergic reactions, liver damage, nervous system damage, etc., and these additives are listed in the prohibited list. In view of this, accurate identification of food additive types is of great significance for ensuring food safety.

[0003] Currently, methods such as gas chromatography, high-performance liquid chromatography, and mass spectrometry have been used for the detection and identification of food additives. These traditional methods are accurate and sensitive, but require cumbersome pretreatment and rely on bench-top equipment. In addition, spectroscopic methods such as fluorescence spectroscopy and ultraviolet-visible spectrophotometry are also applied in the detection of food additives, with the characteristics of simple operation and high sensitivity, but their specificity and stability are limited. Surface-enhanced Raman spectroscopy (SERS) is a sensitive, specific, and reliable detection and analysis technique that can comprehensively display the molecular fingerprint information of the measured sample and has been widely used in the field of food safety. However, due to the similarity of molecular structures or the consistency of functional groups, the SERS spectra of the same type of food additives often show highly similar spectral shapes, and even a large number of overlapping features appear. To effectively improve the recognition accuracy of similar spectra, dimensionality reduction algorithms are used to extract key information from a large amount of spectral data, and these key information are used as feature variables to input into the model for training, so as to improve the model prediction performance. Commonly used dimensionality reduction algorithms include principal components analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), etc. Among them, the PCA dimensionality reduction method based on variance maximization may lose some key information in the case of low variance; while the t-SNE method embeds a random initialization mode, and its reproducibility and the stability of the downscaling results are poor, thus affecting the model recognition performance.

[0004] Compressed sensing (CS) is a powerful signal processing method that provides an effective solution for data redundancy. CS utilizes the sparsity of signals. By designing an appropriate measurement matrix, it projects high-dimensional signals into a low-dimensional space, and then reconstructs the original signals from a small number of non-linear measurements through an optimization algorithm. It is widely used in fields such as magnetic resonance imaging and single-pixel cameras. In recent years, researchers have achieved the reconstruction of Raman spectroscopy raw data using the sparse representation in the latent space, demonstrating the feasibility of CS in one-dimensional spectral data. However, how to optimize the sparse features of Raman spectroscopy and establish an identification model based on the optimized sparse features has not been mentioned. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method for identifying food additives based on the sparse features of surface-enhanced Raman spectroscopy.

[0006] The present invention provides a method for identifying food additives based on the sparse features of surface-enhanced Raman spectroscopy, including:

[0007] Obtaining the SERS spectra of food additive samples of known types;

[0008] According to the SERS spectra, extracting the sparse features of the SERS spectra and optimizing the sparse features;

[0009] Dividing the optimized sparse features into a training set and a test set, establishing an identification model using the training set, and verifying the identification accuracy of the identification model using the test set;

[0010] Using the identification model to identify the types of food additives in the sample to be tested.

[0011] Further, before extracting the sparse features of the SERS spectra according to the SERS spectra, it includes: preprocessing the SERS spectra.

[0012] Further, the extracting the sparse features of the SERS spectra and optimizing the sparse features according to the SERS spectra includes:

[0013] Given that the length of the SERS spectrum f n×1 is n, the projection representation of the SERS spectrum signal on the transform basis Ψ n×n is the sparse vector z n×1 , that is, f n×1 =Ψ n×n ·z n×1 ;

[0014] Using the sensing matrix Φ kn×n to perform sub-Nyquist sampling on the SERS spectrum to obtain the observation vector gkn×1 , i.e., g kn×1 = Φ kn×n ·f n×1 ;

[0015] Based on the sparse vector z n×1 and the observation vector g kn×1 , solve for the optimal solution , i.e., the optimized sparse vector.

[0016] Furthermore, the transform basis Ψ n×n adopts any one of the fast Fourier transform, discrete cosine transform, and discrete wavelet transform.

[0017] Furthermore, the determination method of the optimal transform basis Ψ n×n is as follows: the proportion of zero elements in the sparse vector z n×n corresponding to the optimal transform basis Ψ n×1 is the largest, and the sparsity is the highest.

[0018] Furthermore, the sensing matrix Φ kn×n adopts any one of the Gaussian random matrix, partial Hadamard matrix, and Bernoulli random matrix.

[0019] Furthermore, the determination method of the optimal sensing matrix Φ kn×n is as follows: the reconstructed SERS spectrum kn×n corresponding to the optimal sensing matrix Φ has the highest similarity with the SERS spectrum f n×1 , where the reconstructed SERS spectrum is obtained using the optimized sparse vector and the transform basis Ψ n×n .

[0020] Furthermore, the sensing matrix Φ kn×n satisfies: compression ratio k < 1.

[0021] Furthermore, the process of solving for the optimal solution n×1 i.e., the optimized sparse vector, based on the sparse vector z kn×1 and the observation vector g includes:

[0022] Based on the sparse vector z n×1 and the observation vector g kn×1 , obtain g kn×1 = Φ kn×n ·f n×1 = Φ kn×n ·Ψ n×n ·z n×1 ;

[0023] Solve it by using the L1 norm minimization method.

[0024] Further, establishing the recognition model by using the training set includes: training the training set by using any one of the nearest neighbor algorithm, support vector machine, random forest, and convolutional neural network to establish the recognition model.

[0025] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0026] The present invention establishes a recognition model based on the sparse features of SERS spectra, has higher recognition accuracy, and has superiority in terms of operating efficiency and training speed. The food additive recognition method based on the sparse features of SERS spectra provided by the present invention has a significant advantage in recognizing spectra with high spectral shape similarity and feature overlap degree, can accurately judge the same type of food additives, helps to improve the accuracy of additive type recognition, and realizes high-precision discrimination of food additives. Description of the Drawings

[0027] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objectives, and advantages of the present invention will become more obvious:

[0028] Figure 1 It is a flowchart of the extraction and optimization of SERS spectral sparse features in an embodiment of the present invention;

[0029] Figure 2 It is the SERS spectral data of six food additives in an embodiment of the present invention, where: (a) represents malachite green, (b) represents crystal violet, (c) represents brilliant blue, (d) represents rhodamine B, (e) represents rhodamine 6G, and (f) represents carmine;

[0030] Figure 3 It is the sparse representation of (a) the SERS spectral data after discrete cosine transform basis transformation in an embodiment of the present invention, and (b) the sparsity evaluation of the sparse representation;

[0031] Figure 4 It is the optimization result of (a) sparse representation of SERS spectral data by using a Gaussian random matrix as a sensing matrix in an embodiment of the present invention, and (b) the SERS spectrum reconstructed by optimizing the sparse representation;

[0032] Figure 5 It is the comparison of (a) the training parameters of the CNN model based on the sparse features of SERS spectra with (b) the training parameters of the CNN model based on PCA and (c) t-SNE. Detailed Embodiments

[0033] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.

[0034] An embodiment of the present invention provides a method for identifying food additives based on the sparse features of surface-enhanced Raman spectroscopy to accurately identify food additives. The method includes the following steps:

[0035] S1. Obtain the SERS spectra of food additive samples of known types;

[0036] S2. Extract the sparse features of the SERS spectra according to the SERS spectra and optimize the sparse features;

[0037] S3. Divide the optimized sparse features into a training set and a test set, establish an identification model using the training set, and verify the identification accuracy of the identification model using the test set;

[0038] S4. Use the trained identification model to identify the types of food additives in the sample to be tested.

[0039] The embodiment of the present invention constructs an identification model based on the sparse features of surface-enhanced Raman spectroscopy to achieve high-precision discrimination of food additives, showing superiority in the model operation efficiency and training speed, and helping to improve the accuracy of additive type identification.

[0040] To obtain the SERS spectra of known food additive samples, specifically, select food additive samples of known exact types, collect SERS spectral signals for each sample respectively, and the total number of groups of SERS spectral data of all collected additive samples is M. During the collection of SERS spectral data of all additive samples, ensure that the test environment such as laser power and integration time is consistent.

[0041] In order to obtain reliable and effective data information and improve the identification accuracy, in some embodiments, before extracting the sparse features of the SERS spectra according to the SERS spectra, it includes: preprocessing the collected original SERS spectra. Specifically, perform spectral data preprocessing on the measured original SERS spectra of food additives to reduce the noise signals and fluorescence background in the spectral data, that is, baseline correction, curve smoothing, normalization, and averaging processing to obtain the average SERS spectrum.

[0042] In some embodiments, refer to Figure 1, the extraction and optimization of SERS spectral sparse features include the sparse representation of SERS spectra, the observed representation of SERS spectra, and the optimization of SERS spectral sparse features. Specifically:

[0043] S21. Sparse representation of the preprocessed SERS spectrum: Given that the length of the SERS spectrum f n×1 is n, and the projection representation of the SERS spectral signal on the transform basis Ψ n×n is the sparse vector z n×1 , that is, f n×1 = Ψ n×n ·z n×1 ;

[0044] S22. Observed representation of the preprocessed SERS spectrum: Use the sensing matrix Φ kn×n to perform sub-Nyquist sampling on the SERS spectrum to obtain the observed vector g kn×1 , that is, g kn×1 = Φ kn×n ·f n×1 ;

[0045] S23. Optimize the sparse vector z n×1 : According to the sparse vector z n×1 and the observed vector g kn×1 , solve the optimal solution That is, the optimized sparse vector.

[0046] In some embodiments, the transform basis Ψ n×n adopts any one of orthogonal matrices such as the fast Fourier transform, discrete cosine transform, discrete wavelet transform, etc. Through the above orthogonal basis transformation, the sparse representation of SERS spectral data can be realized to remove redundant information in the data.

[0047] In some embodiments, the best transform basis can be confirmed by the sparsity of the sparse vector z n×1 . The determination method of the best transform basis Ψ n×n is as follows: The proportion of zero elements in the sparse vector z n×n corresponding to the best transform basis Ψ n×1 is the largest, and the sparsity is the highest.

[0048] In some embodiments, the sensing matrix Φ kn×n adopts any one of matrices that are not related to the transform basis, such as Gaussian random matrices, partial Hadamard matrices, Bernoulli random matrices, etc., and combines with the transform basis to realize the optimization of sparse representation.

[0049] In some embodiments, the best sensing matrix can be confirmed by the similarity between the reconstructed SERS spectrum and the preprocessed SERS spectrum f n×1 . The best sensing matrix Φkn×n is determined as follows: the reconstructed SERS spectrum corresponding to the optimal sensing matrix Φ kn×n has the highest similarity with the SERS spectrum f, where the reconstructed SERS spectrum is obtained by using the optimized sparse vector n×1 and the transform basis Ψ and the transform basis Ψ is obtained n×n .

[0050] In the above embodiments of the present invention, the sparse representation can be optimized through the sensing matrix and the transform basis, further removing the redundant information in the sparse representation. The optimized sparse features contain all the features in the SERS data and can provide comprehensive training features for the recognition model.

[0051] In some embodiments, the sensing matrix Φ kn×n satisfies: the compression ratio k < 1, and the data of length n is transformed into kn to achieve data compression.

[0052] In some embodiments, according to the sparse vector z n×1 and the observation vector g kn×1 , the optimal solution , that is, the optimized sparse vector, is solved, including:[[]]

[0053] According to the sparse vector z n×1 and the observation vector g kn×1 , we get:[[]]

[0054] g kn×1 = Φ kn×n · f n×1 = Φ kn×n · Ψ n×n · z n×1 ;

[0055] The L1 norm minimization method is used for solving, specifically:[[]]

[0056]

[0057] The method for extracting and optimizing the sparse features of the SERS spectrum proposed in the above embodiments of the present invention can effectively compress the SERS spectrum data information, remove the redundant part in the data and retain the key comprehensive features. As a spectrum feature extraction method, it can be combined with various algorithms to establish a recognition model, improving the accuracy and precision of food additive type recognition.

[0058] In some embodiments, to establish an identification model, P groups are randomly selected from the optimized sparse features of the measured M groups of SERS spectral data as the training set data, and the remaining Q = M - P groups of data are used as the test set data. Any one of the nearest neighbor algorithm, support vector machine, random forest, and convolutional neural network is used to train the training set to establish an identification model, and the identification model established based on the sparse features of the SERS spectrum is used to identify the types of food additive samples.

[0059] In the above embodiments of the present invention, a suitable transform basis and sensing matrix are designed using the compressive sensing theory, the sparse features of the SERS spectrum are extracted and optimized, and an identification model of food additives based on sparse features is established. The identification model of food additives based on the sparse features of the SERS spectrum has excellent discrimination accuracy and significant advantages in terms of model operation efficiency and training speed.

[0060] In a specific embodiment, a method for identifying food additives based on the sparse features of surface-enhanced Raman spectroscopy is provided, including the following steps:

[0061] (1) Acquisition of SERS spectra of food additive samples of known types

[0062] In this embodiment, the SERS spectra of six food additive solutions, namely malachite green, crystal violet, brilliant blue, rhodamine B, rhodamine 6G, and carmine, are collected. 210 groups of SERS spectral data are collected for each food additive, and a total of 1260 groups of SERS spectral data are collected.

[0063] (2) Data preprocessing

[0064] The collected SERS spectral data is preprocessed by baseline correction, curve smoothing, area normalization, etc. The preprocessed SERS spectra are as Figure 2 shown.

[0065] (3) Extraction and optimization of sparse features of SERS spectra

[0066] ① Sparse representation of SERS spectra:

[0067] Given the preprocessed SERS spectrum f n×1 of length n, the preprocessed SERS signal can be projected and represented as a sparse vector z n×n on the transform basis Ψ n×1 , that is

[0068] f n×1 = Ψ n×n · z n×1 (1)

[0069] ② Observation representation of SERS spectra:

[0070] Utilize the sensing matrix Φ kn×n Perform sub-Nyquist sampling on the preprocessed SERS spectrum to obtain the observation vector g kn×1 , that is

[0071] g kn×1 = Φ kn×n ·f n×1 (2)

[0072] ③ Optimization of SERS spectrum sparse features

[0073] Solve for the optimized sparse vector through equations (1) and (2) From equations (1) and (2), we can obtain

[0074] g kn×1 = Φ kn×n ·f n×1 = Φ kn×n ·Ψ n×n ·z n×1 (3)

[0075] Adopt the L1 norm minimization method to obtain the optimal solution of equation (3) Optimal solution That is, the optimized sparse vector, namely

[0076]

[0077] For example Figure 3 , in this embodiment, the discrete cosine transform is selected to sparsely represent the SERS spectrum, and the sparsity is 0.9539.

[0078] Considering efficiency and inversion accuracy, k = 0.5 is selected in this embodiment.

[0079] For example Figure 4 , in this embodiment, the Gaussian random matrix is selected as the sensing matrix to reconstruct the SERS spectrum The correlation with the preprocessed SERS spectrum f n×1 reaches 0.9929.

[0080] (4) Divide the optimized sparse features into a training set and a prediction set

[0081] In this embodiment, the five-fold cross-validation method is used to classify the 1260 groups of SERS spectrum sparse features into a training set and a prediction set. The recognition model is established using the training set data, and the recognition performance of the model is verified using the prediction set.

[0082] (5) Use the recognition model based on SERS spectrum sparse features to discriminate food additives

[0083] Using the training set data, recognition models of the K-nearest neighbor algorithm (KNN), support vector machine (SVM), random forest (RF), and convolutional neural network (CNN) were established respectively, and the prediction set data was predicted. After calculation, the average discrimination accuracies of the KNN, SVM, RF, and CNN models based on the sparse features of SERS spectra for food additive samples were 0.97, 0.94, 0.98, and 0.99 respectively, and it can be considered that their recognition of food additives is efficient.

[0084] Comparative example

[0085] Different from the method of extracting and optimizing the sparse features of SERS spectra in the above embodiments and using the sparse features as feature variables to input into the model for training and prediction, the comparative example adopted PCA and t-SNE methods to reduce the dimension of SERS spectra, and used the PCA and t-SNE features after dimension reduction as feature variables to input into the model for training and prediction, and compared with the results of the above embodiments. The discrimination results are shown in Table 1, which proves that the recognition method based on the sparse features of SERS spectra proposed in the above embodiments has higher accuracy for the identification of food additive types.

[0086] Table 1 Comparison of discrimination accuracy results between the embodiment and the comparative example

[0087]

[0088] As Figure 5 shown, in the CNN models of the embodiment and the comparative example, the accuracy and loss function curves of the training set and the prediction set in the embodiment tend to be stable after about 150 training cycles. However, the accuracy and loss function curves of the CNN model based on PCA and t-SNE in the comparative example gradually converge after about 350 training cycles. It proves that the embodiment not only helps to improve the convergence speed and stability of the CNN model, but also reduces the complexity and computational requirements of the model while maintaining the performance.

[0089] The food additive recognition method based on the sparse features of SERS spectra provided in the above embodiments of the present invention has excellent discrimination accuracy. The food additive recognition model based on the sparse features of SERS spectra has significant advantages in the model operation efficiency and training speed, can accurately judge the same type of food additives, helps to improve the accuracy of additive type recognition, and realizes high-precision discrimination of food additives.

[0090] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention. The above preferred features can be used in any combination without conflict.

Claims

1. A method for identifying food additives based on sparse features of surface-enhanced Raman spectroscopy, characterized in that Comprising: Obtaining the SERS spectra of samples of food additives of known types; Extracting the sparse features of the SERS spectra according to the SERS spectra and optimizing the sparse features; Dividing the optimized sparse features into a training set and a test set, establishing an identification model using the training set, and verifying the identification accuracy of the identification model using the test set; Using the identification model to identify the types of food additives in the sample to be tested.

2. The food additive identification method based on the sparse features of surface enhanced Raman spectroscopy according to claim 1, wherein Before extracting the sparse features of the SERS spectra according to the SERS spectra, it includes: preprocessing the SERS spectra.

3. The method for identifying food additives based on the sparse features of surface-enhanced Raman spectroscopy according to claim 1, wherein The extracting the sparse features of the SERS spectra according to the SERS spectra and optimizing the sparse features includes: Given that the length of the SERS spectrum f n×1 is n, and the projection of the SERS spectrum signal on the transformation basis Ψ n×n is represented as the sparse vector z n×1 , that is, f n×1 = Ψ n×n ·z n×1 ; Using the sensing matrix Φ kn×n Perform sub-Nyquist sampling on the SERS spectrum to obtain the observation vector g kn×1 , that is, g kn×1 = Φ kn×n ·f n×1 ; According to the sparse vector z n×1 and the observation vector g kn×1 , solve for the optimal solution which is the optimized sparse vector.

4. The food additive identification method based on the sparse features of surface-enhanced Raman spectroscopy according to claim 3, wherein The transform basis Ψ n×n Adopts any one of fast Fourier transform, discrete cosine transform, and discrete wavelet transform.

5. The food additive identification method based on the sparse features of surface-enhanced Raman spectroscopy according to claim 4, wherein Optimal transformation basis Ψ n×n is determined as follows: the sparse vector z n×n corresponding to the optimal transformation basis Ψ n×1 has the largest proportion of zero elements and the highest sparsity degree.

6. The method for identifying food additives based on sparse features of surface-enhanced Raman spectroscopy according to claim 3, wherein The sensing matrix Φ kn×n Adopts any one of a Gaussian random matrix, a partial Hadamard matrix, and a Bernoulli random matrix.

7. The method for identifying food additives based on sparse features of surface-enhanced Raman spectroscopy according to claim 6, characterized in that, Optimal sensing matrix Φ kn×n is determined as follows: The reconstructed SERS spectrum kn×n corresponding to the optimal sensing matrix Φ has the highest similarity with the SERS spectrum f n×1 , where the reconstructed SERS spectrum is obtained by using an optimized sparse vector and a transform basis Ψ n×n .

8. The method for identifying food additives based on the sparse features of surface-enhanced Raman spectroscopy according to claim 3, wherein The sensing matrix Φ kn×n satisfies: the compression ratio k < 1.

9. The food additive identification method based on the sparse features of surface-enhanced Raman spectroscopy according to claim 3, wherein According to the sparse vector z n×1 and the observation vector g kn×1 , solve the optimal solution That is, the optimized sparse vector, including: Based on the sparse vector z n×1 and the observation vector g kn×1 , we obtain g kn×1 = Φ kn×n ·f n×1 = Φ kn×n ·Ψ n×n ·z n×1 ; Solve by using the L1 norm minimization method, 10. The method for identifying food additives based on the sparse features of surface-enhanced Raman spectroscopy according to claim 1, wherein The establishing an identification model using the training set includes: training the training set using any one of the nearest neighbor algorithm, support vector machine, random forest, and convolutional neural network to establish an identification model.

Citation Information

Cited By

  • Food illegal additive detection method and system based on surface enhanced Raman spectroscopy

    CN122173905A

  • Method and system for detecting illegal additives in food based on surface-enhanced raman spectroscopy

    CN122173905B