Fish oil adulteration identification method by combining Raman spectrum with multivariate statistical analysis
By combining Raman spectroscopy with multivariate statistical analysis, the PCA_LDA algorithm was improved, solving the problems of complex pretreatment and environmental interference in fish oil adulteration detection. This enabled rapid and accurate identification of fish oil adulteration with high classification accuracy.
Patent Information
- Application Number
- CN202511661013.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies for detecting adulteration in fish oil suffer from problems such as complex pretreatment steps, long processing time, cumbersome labeling process, and environmental interference affecting Fourier transform infrared spectroscopy and near-infrared spectroscopy. Furthermore, Raman spectroscopy is rarely used in detecting adulteration in fish oil and has insufficient classification accuracy.
A method combining Raman spectroscopy and multivariate statistical analysis was adopted. By improving the PCA_LDA algorithm, feature vector sorting was performed, and a linear discriminant analysis model was established to achieve rapid, non-destructive, and pre-processing-free identification of fish oil adulteration.
It achieves rapid and accurate detection of fish oil adulteration, with a classification accuracy rate of over 95%, especially for fish oil with an adulteration rate of 5%, the accuracy rate is over 95%, and the sample preparation process is simplified.
Smart Images

Figure CN121595530A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for identifying adulterated fish oil, and more particularly to a rapid and highly accurate method for identifying adulterated fish oil using Raman spectroscopy combined with multivariate statistical analysis without pretreatment. Background Technology
[0002] Fish oil, a natural lipid supplement extracted from deep-sea fish, is valued primarily for its high content of eicosapentaenoic acid (EPA) and docosahexaenoic acid (DHA). -3 polyunsaturated fatty acids. These active ingredients have been confirmed by numerous studies to have multiple health benefits: not only can they reduce the risk of blood clots and maintain cardiovascular health, but they also have positive effects on brain nerve development, vision protection, mood regulation, and even cancer prevention. With the increasing health awareness of modern people, the global market demand for fish oil is showing a continuous growth trend. However, the production of high-quality fish oil faces many challenges: firstly, the raw material processing technology is complex; secondly, the extraction process requires precise processes such as molecular distillation, which makes compliant production costs high. Driven by high profits, some unscrupulous merchants use cheap vegetable oils for adulteration. These counterfeit products contain almost no EPA / DHA, which have health benefits, and their appearance is extremely similar to genuine products, making it difficult for ordinary consumers to distinguish them. More seriously, long-term consumption of inferior fish oil not only fails to achieve the expected health benefits but may also damage liver and kidney function, and even increase the risk of chronic diseases, endangering human safety and health. This industry chaos not only directly harms consumer rights but also weakens consumer confidence, triggering widespread public skepticism about the safety of health products.
[0003] In recent years, various methods have been used to detect adulteration in fish oil, including liquid chromatography, gas chromatography, isotope methods, Fourier transform infrared spectroscopy (FTIR), near-infrared spectroscopy (NIR), and Raman spectroscopy. However, chromatographic methods suffer from cumbersome sample pretreatment steps and long analysis times; isotope methods require sample labeling and cannot achieve non-destructive testing; and Fourier transform infrared spectroscopy and near-infrared spectroscopy are easily affected by moisture and sample cell interference, resulting in strong absorption background signals and temperature dependence.
[0004] Raman spectroscopy is simple to operate, has a short testing time, and is unaffected by sample size, signal intensity, and transparency. It requires very little sample volume, making it ideal for rapid and simple analysis of liquid samples. Furthermore, Raman spectroscopy can provide more precise detection of oils and fats, and is often considered a complementary technique to FTIR. Raman spectroscopy is widely used in the detection of adulteration in vegetable oils. For example, Li et al. used a portable Raman spectrometer to rapidly detect multiple adulterations in sesame oil, identifying 40 sesame oil samples containing four different adulterants, demonstrating that Raman spectroscopy is an effective method for detecting multiple adulterations in sesame oil. Teng et al. established a one-dimensional convolutional neural network model and used a portable Raman spectrometer to achieve oil type identification and quantitative analysis of multiple adulterants. By testing commercially available sesame oil products and comparing their results with gas chromatography and colorimetry, they verified the effectiveness of their model in achieving higher detection accuracy under low-concentration adulteration conditions. While Raman spectroscopy is widely used in detecting adulteration in vegetable oils, its application in fish oil detection is relatively limited. For example, Hall et al. demonstrated a technique combining Raman spectroscopy with partial least squares (PLS) chemometrics to identify genuine shark liver oil (SLO) by monitoring squalene concentration. Bekhit et al. combined FT-IR, NIR, and Raman spectroscopy with chemometrics, showing that by monitoring the concentrations of eicosapentaenoic acid (EPA), docosahexaenoic acid (DHA), and n-3 polyunsaturated fatty acids (omega-3), Raman spectroscopy combined with PLS chemometrics can be used to identify genuine fish oil.
[0005] Raman spectroscopy faces challenges from multiple interference factors in practical applications. Raman spectral data is typically high-dimensional, but dimensionality reduction techniques such as principal component analysis (PCA) can effectively extract key feature variables, improving the discrimination efficiency of classification models. Secondly, when faced with scenarios involving the mixing and adulteration of multiple substances, traditional single-spectral analysis methods often prove limited. In such cases, employing supervised learning methods to construct multi-classification models can significantly improve recognition accuracy through feature space mapping. Summary of the Invention
[0006] This invention solves the following problems: (1) In response to the problems of complex pretreatment steps, labeling, and long time consumption in traditional chromatographic and isotopic methods, the method proposed in this invention has the advantages of no pretreatment process, non-destructive labeling, and rapid detection. (2) In response to the strong absorption background signal and temperature dependence caused by the interference of moisture and sample cell in Fourier transform infrared spectroscopy and near-infrared spectroscopy, Raman spectroscopy is not limited by the environment such as moisture, sample cell, and temperature. (3) Although existing Raman spectroscopy has been reported in fish oil, it is limited to component analysis and has not been applied to the field of fish oil adulteration. This invention combines Raman spectroscopy with the PCA-LDA algorithm and improves the algorithm to improve the classification accuracy.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] A method for identifying adulteration in fish oil using Raman spectroscopy combined with multivariate statistical analysis, the method being as follows:
[0009] Step 1: Collect the Raman spectrum of the fish oil to be tested using a Raman spectrometer;
[0010] Step 2: Preprocess the acquired Raman spectra by denoising, baseline removal, and normalization;
[0011] Step 3: Divide the spectral data into a training set and a test set in a 7:3 ratio;
[0012] Step 4: Add sorting to the traditional PCA and calculate the intra-class scattering matrix. and inter-class scattering matrix And calculate the generalized Rayleigh quotient based on it, and sort them according to the size of the generalized Rayleigh quotient;
[0013] Step 5: Train the linear discriminant analysis model using the data from the training set, and use the results from the test set as the final classification result.
[0014] Further, step one specifically involves: placing the fish oil sample to be tested on a glass slide and acquiring the Raman scattering signal of the sample using a high numerical aperture objective microscope; selecting a visible light wavelength, such as 532 nm, as the excitation source; the Raman spectroscopy acquisition range being the Raman fingerprint region; all measurements being performed at room temperature; the acquisition time being more than 1 second; and acquiring four spectra. This step requires no complex pretreatment steps, the sample preparation method is simple to operate, label-free and non-destructive, and the detection process is short.
[0015] Further, step two specifically involves baseline correction and Savitzky-Golay convolution smoothing, which are used to eliminate baseline shift and high-frequency noise, respectively, thereby improving the signal-to-noise ratio of the Raman spectra. Simultaneously, the spectral data are normalized to reduce scattering effects and facilitate sample comparison. Data preprocessing was performed using Matlab.
[0016] Furthermore, step four specifically involves:
[0017] (1) Form a matrix from the training set , , The number of samples in the training set. The dimension of the sample;
[0018] (2) Using a matrix Composition of covariance matrix , ,in For the training set One sample, The average value of the training set samples is ; the total dimension of the samples is . This refers to a dimension ranging from 1 to 1. That is, 1, 2, 3, ... , It's just a symbolic representation, referring to From , Keep taking until ;
[0019] (3) Calculate the intra-class scattering matrix based on the traditional PCA algorithm. Inter-class scattering matrix , , ,in The number of categories in the sample. For the first Number of samples in each class For the first The first in the class One sample, For the first Class sample mean, The mean of the total sample is... This is a matrix transpose operation; according to Derivation of generalized Rayleigh quotient ,in The eigenvectors obtained after eigenvalue decomposition in (3) are sorted according to the size of the generalized Rayleigh quotient, and a new set of eigenvectors is obtained.
[0020] The advantages of this invention over the prior art are as follows:
[0021] (1) Combining Raman spectroscopy with machine learning results in less solution consumption, simpler sample preparation steps, no need for complex sample pretreatment, and shorter analysis time, enabling rapid detection and classification.
[0022] (2) The PCA_LDA algorithm applied to Raman spectroscopy is improved by sorting the feature vectors according to certain rules. Compared with the traditional PCA algorithm, the main feature information is better preserved while reducing the dimensionality, so that the LDA results after classifying the Raman spectra of different types of fish oil are more accurate.
[0023] (3) The analysis time is short and the classification accuracy is high. For fish oil with a 5% adulteration ratio, the classification accuracy rate reaches more than 95%. The data accuracy rate for fish oil adulterated with grape seed oil can reach 95.37%, and the data accuracy rate for fish oil adulterated with avocado is 96.29%. Attached Figure Description
[0024] Figure 1 This is the experimental optical path diagram;
[0025] Figure 2 The images show the spectra of oil samples collected. A represents fish oil and grape seed oil, and B represents fish oil and avocado oil.
[0026] Figure 3 Box plots of characteristic peaks in Raman spectra for different groups;
[0027] Figure 4 The image shows the traditional PCA_LDA classification results.
[0028] Figure 5 This is a flowchart of spectral data processing.
[0029] Figure 6 The improved PCA_LDA classification results are shown in the figure. Detailed Implementation
[0030] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments, but it is not limited thereto. Any modifications or equivalent substitutions to the technical solution of the present invention that do not depart from the spirit and scope of the technical solution of the present invention should be covered within the protection scope of the present invention.
[0031] Example 1
[0032] To achieve rapid detection of adulteration in fish oil, this invention combines Raman spectroscopy with multivariate statistics to establish a linear discriminant analysis model for adulteration identification and classification. The overall process is as follows: Figure 5 As shown, the specific implementation steps are as follows:
[0033] Step 1: Sample preparation: Grape seed oil and avocado oil were added to the pure fish oil sample and mixed evenly using a vortex mixer to prepare adulterated fish oil samples with adulteration volumes ranging from 5% to 50%. A total of 165 samples were prepared, including 15 samples of pure fish oil, 15 samples of pure grapeseed oil, 15 samples of pure avocado oil, and 120 samples of adulterated fish oil containing grapeseed oil and avocado oil respectively (15 samples of adulterated oil with 5%, 10%, 20%, and 50% volume of grapeseed oil; 15 samples of adulterated avocado oil with 5%, 10%, 20%, and 50% volume of avocado oil). All adulterated samples were prepared in real time for subsequent experiments.
[0034] Step Two: Spectral Acquisition: Raman spectra of pure and adulterated oils were acquired using an Andor spectrometer. The specific procedure was as follows: The oil sample was placed on a glass slide, and Raman scattering signals of each sample were acquired using a high numerical aperture objective microscope. A visible light wavelength, such as 532 nm, was selected as the excitation source, and the Raman spectroscopy acquisition range was the Raman spectral fingerprint region. All measurements were performed at room temperature, with an acquisition time of at least 1 second. Four spectra were acquired for each oil sample. The acquisition optical path is as follows: Figure 1 As shown, Laser is the laser source, providing a 532 nm wavelength laser light; HWP is a half-wave plate, which, together with the polarization beam splitter (PBS), forms an energy adjustment module to adjust the laser power; M1 and M6 are mirrors used to change the laser illumination direction to adapt to actual optical path requirements; L1 and L2 are two lenses that form a beam expander system, ensuring that the laser maintains sufficient beam diameter and beam quality during transmission; M2, M3, M4, and M5 are all mirrors, mainly used to adjust the beam pitch and height to collimate the beam with the objective lens; L3 lens is designed to eliminate the influence of the tube lens (TL) in the microscope, ensuring that the light entering the objective lens is a parallel beam; DM is a dichroic mirror, separating the laser according to wavelength; LPF is a long-pass filter used to filter out excess 532 nm light; and L4 lens is used to eliminate dispersion, focusing the Raman scattered light onto the spectrometer's slit.
[0035] Step 3: Spectral Preprocessing: Data preprocessing was performed using Matlab. Baseline correction and Savitzky-Golay convolution smoothing were used to eliminate baseline shift and high-frequency noise, respectively, thereby improving the signal-to-noise ratio of the Raman spectra. Simultaneously, the spectral data were normalized to reduce scattering effects and facilitate sample comparison. Figure 2 The average Raman spectra of fish oil and various vegetable oils after baseline correction are shown. Figure 2 A represents the spectra collected from fish oil, grape seed oil, and fish oil mixed with grape seed oil, respectively. Figure 2 B represents the spectra collected from fish oil, avocado oil, and the mixture of avocado and fish oil. For example... Figure 2 As shown, the Raman peaks are mainly concentrated between 800 and 1700 cm. -1 Within the range. 841cm -1 943cm -1 1053cm -1 The three Raman peaks are part of the CC framework at 800-1200cm. -1 Characteristic peak, 841 cm⁻¹ -1 With 1053cm -1 The peak at 943 cm⁻¹ is related to the stretching vibration of the methylene chain backbone, while the peak at 943 cm⁻¹ is related to the stretching vibration of the methylene chain backbone. -1 The peak at 1242 cm⁻¹ reflects the bending vibration of the trans-C=C bond. -1The characteristic peak at 1278 cm⁻¹ reflects the CH vibration of unsaturated olefins (=CH) and corresponds only to the cis-unsaturated structure. -1 The nearby band corresponds to the in-phase methylene bending deformation vibration (CH deformation of -CH2). The C / C stretching vibration of the alkene molecule is at 1625 cm⁻¹. -1 A strong Raman scattering signal is generated at that location.
[0036] And such as Figure 3 As shown, when the concentration of the added vegetable oil is gradually increased from 0% to 100%, at 1242 cm⁻¹... -1 With 1625cm -1 The peak intensity at certain locations shows a regular decrease. This may be due to the increased proportion of vegetable oil in adulterated fish oil, leading to a corresponding increase in the intensity of its characteristic peaks.
[0037] Step 4: Dimensionality Reduction and Feature Extraction: The Raman spectra of the collected oil samples are high-dimensional data, with 1024 dimensions. The spectral data contains a large amount of noise and useless information. Dimensionality reduction of the Raman spectral data can extract the main feature information and reduce the difficulty of analysis. The specific steps are as follows:
[0038] (1) First, divide the data into training set and test set in a 7:3 ratio, and then form a matrix from the training set. , , The number of samples in the training set. The dimension of the sample;
[0039] (2) Using a matrix Composition of covariance matrix , ,in For the training set One sample, This represents the average value of the training set samples;
[0040] (3) According to For matrix Perform eigenvalue decomposition to obtain a set of eigenvectors. , , ..., , and These are the eigenvalues and their corresponding eigenvectors; conventional PCA sorts the eigenvectors according to the magnitude of the eigenvalues, while this invention calculates the intra-class scattering matrix based on this. Inter-class scattering matrix , , ,in The number of categories in the sample. For the first Number of samples in each class For the first The first in the class One sample, For the first Class sample mean, The mean of the total sample is... This is a matrix transpose operation; according to Derivation of generalized Rayleigh quotient ,in The eigenvectors obtained after eigendecomposition in (3) are sorted according to the size of the generalized Rayleigh quotient, and then a new set of eigenvectors is obtained.
[0041] Step 4: Project the training and test samples onto the new feature vector.
[0042] Step 5: Results Analysis: Linear Discriminant Analysis (LDA) is used to perform a qualitative analysis of the adulteration situation. The LDA model is evaluated using accuracy, confusion matrix, and scatter plot; these parameters are used to measure the model's classification ability.
[0043] The dimensionality reduction of the spectral data was performed using the traditional PCA method, followed by classification analysis using the LDA model. The results are as follows: Figure 4 As shown, the classification accuracy rates for adulterated grape seed oil and adulterated avocado oil were 91.67% and 90.74%, respectively.
[0044] The classification results using the improved PCA_LDA model are as follows: Figure 6 As shown. In the confusion matrix, the data in each diagonal square represents the number of predicted values that match the true values. For adulterated samples containing grape seed oil, the data is... Figure 6 As shown in A, among the 108 samples in the test set, the pure oil group and the group with an adulteration concentration of 50% were correctly classified. Overall, the classification accuracy rate for adulterated fish oil in grape seed oil was 95.37%. Further... Figure 6 As shown in C, samples of the same class will cluster together and be distinguishable from samples of other classes. When a sample is assigned to another class, overlap will occur between samples of different colors, but this overlap will not occur when the test set samples are correctly predicted for classification. The classification results for samples mixed with avocado oil are shown below. Figure 6 As shown in B and D. From Figure 6 As shown in B, similarly, the groups of pure fish oil, pure avocado oil, and 50% adulterated oil were correctly classified, with a classification accuracy rate of 96.29%. The classification scatter plot is shown below. Figure 6 As shown in D.
Claims
1. A method for identifying adulterated fish oil using Raman spectroscopy combined with multivariate statistical analysis, characterized by: The method is as follows: Step 1: Collect the Raman spectrum of the fish oil to be tested using a Raman spectrometer; Step 2: Preprocess the acquired Raman spectra by denoising, baseline removal, and normalization; Step 3: Divide the spectral data into a training set and a test set in a 7:3 ratio; Step 4: Add sorting to the traditional PCA and calculate the intra-class scattering matrix. and inter-class scattering matrix And calculate the generalized Rayleigh quotient based on it, and sort them according to the size of the generalized Rayleigh quotient; Step 5: Train the linear discriminant analysis model using the data from the training set, and use the results from the test set as the final classification result.
2. The method for identifying adulterated fish oil using Raman spectroscopy combined with multivariate statistical analysis according to claim 1, characterized in that: Step one specifically involves: placing the fish oil to be tested on a glass slide and acquiring the Raman scattering signal of the sample using a high numerical aperture objective microscope; using a visible light wavelength, such as 532 nm, as the excitation source; the Raman spectrum acquisition range is the Raman fingerprint region; all measurements are performed at room temperature; the acquisition time is more than 1 second; and four spectra are acquired.
3. The method for identifying adulterated fish oil using Raman spectroscopy combined with multivariate statistical analysis according to claim 1, characterized in that: Step two specifically involves baseline correction and Savitzky-Golay convolutional smoothing, which are used to eliminate baseline shift and high-frequency noise, respectively, while also normalizing the spectral data.
4. The method for identifying adulterated fish oil using Raman spectroscopy combined with multivariate statistical analysis according to claim 1, characterized in that: Step four specifically involves: (1) Form a matrix from the training set , , The number of samples in the training set. The dimension of the sample; (2) Using a matrix Composition of covariance matrix , ,in For the training set One sample, This represents the average value of the training set samples; (3) Calculate the intraclass scattering matrix Inter-class scattering matrix , , ,in The number of categories in the sample. For the first Number of samples in each class For the first The first in the class One sample, For the first Class sample mean, The mean of the total sample is... This is a matrix transpose operation; according to Derivation of generalized Rayleigh quotient ,in The eigenvectors obtained after eigenvalue decomposition in (3) are sorted according to the size of the generalized Rayleigh quotient, and a new set of eigenvectors is obtained.