A rapid identification method for edible oil adulteration based on differential raman spectroscopy technology and multi-model machine learning

By combining differential Raman spectroscopy with multi-model machine learning, the problems of fluorescence interference and quantitative analysis in the detection of adulterated edible oils have been solved, enabling rapid and accurate identification of adulterated edible oils, which is suitable for on-site testing.

CN122631613APending Publication Date: 2026-08-25HUNAN POLICE ACAD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610053473.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In the detection of adulteration in edible oils, fluorescence interference leads to insufficient sensitivity of Raman spectroscopy, making it unsuitable for practical detection. Furthermore, existing methods struggle to achieve rapid and accurate qualitative and quantitative analysis.

Method used

A multi-model machine learning framework was constructed to rapidly identify adulterated edible oils by using differential Raman spectroscopy combined with a multi-model machine learning approach. This framework suppresses fluorescence background through dual-wavelength differential operations and combines data standardization and characteristic peak extraction.

Benefits of technology

It achieves the acquisition of high-quality spectral data, improves identification accuracy and reliability, and can complete qualitative and semi-quantitative analysis of edible oils within minutes, making it suitable for rapid on-site testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122631613A_ABST
    Figure CN122631613A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on differential Raman spectroscopy technology and multi-model machine learning's edible oil adulteration rapid identification method, the method includes the following steps: S1 carries out the preparation and pretreatment of adulteration sample;S2 utilize differential Raman spectrum acquisition high-quality spectral data;S3 carries out spectral data processing;S4 feature extraction processing is based on the structural characteristics of edible oil and adulterant, confirms and records the characteristic Raman peak position, intensity and area of edible oil and adulterant;S5 constructs multi-model machine learning framework, for different analysis target, select and train corresponding model;S6 model output, the performance optimal model obtained in the step S5 after strict evaluation, can carry out rapid identification to unknown edible oil sample, output whether it is adulterated qualitative conclusion and adulteration concentration semi-quantitative prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of food analysis, specifically relating to a rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning. Background Technology

[0002] As a daily necessity, the purity and safety of cooking oil directly impact consumers' health and the fair order of the market. Adulteration not only constitutes serious economic fraud and infringes upon consumer rights, but also introduces unknown health risks and may even trigger food safety incidents. Therefore, establishing rapid, accurate, and efficient technologies for identifying genuine and counterfeit cooking oil and detecting adulteration is an urgent need to ensure food safety and regulate market order, possessing significant social and economic value.

[0003] For a long time, the detection of adulteration in edible oils has mainly relied on chromatographic methods (such as gas chromatography (GC) and high-performance liquid chromatography (HPLC) and their coupled techniques. Although these methods are highly accurate, they generally have significant drawbacks. First, the pretreatment is complex and cumbersome, requiring complicated extraction, derivatization, and purification steps for the oil sample, which is time-consuming and labor-intensive, making it difficult to meet the needs of rapid on-site screening. Second, the equipment is expensive and the operation is specialized, relying on large precision instruments and professional operators, resulting in high detection costs and making it difficult to popularize in grassroots supervision and production enterprises. Third, it is a destructive detection method, which usually requires the consumption of samples and is not suitable for non-destructive analysis. In contrast, conventional spectroscopic methods (such as ordinary Raman spectroscopy) theoretically have the advantages of being fast, non-destructive, and simple in sample pretreatment. However, in the detection of edible oils, they face strong fluorescence background interference. The natural pigments, impurities, and some adulterants in edible oils themselves produce strong fluorescence signals, the intensity of which often far exceeds the weak Raman signal, causing the Raman characteristic peaks to be completely submerged, resulting in insufficient sensitivity and inability to effectively identify trace adulteration.

[0004] Differential Raman spectroscopy (DRaM) suppresses fluorescence background and preserves pure Raman characteristic peaks through dual-wavelength differential operations, fundamentally solving the core bottleneck of fluorescence interference in Raman detection of edible oils and providing a high-quality spectral data foundation for subsequent analysis. However, to achieve accurate identification of complex adulteration systems in edible oils (such as distinguishing oil types, identifying unknown adulterants, and especially quantifying adulteration ratios), DRaM alone is insufficient. Powerful information analysis tools are urgently needed to extract effective features from high-dimensional spectral data and establish accurate discrimination models. The introduction of machine learning (ML) algorithms has brought a qualitative leap to this stage. This invention does not simply combine Raman spectroscopy with machine learning, but rather constructs a complete technical chain of "differential fluorescence suppression - data standardization processing - feature-oriented extraction - multi-model task-oriented analysis." Using rapeseed oil adulterated with kerosene and soybean oil adulterated with diesel as case studies, the effectiveness and universality of this fusion method in qualitative and semi-quantitative analysis are verified. Summary of the Invention

[0005] Technical issues

[0006] This invention aims to solve two major bottlenecks in the detection of adulteration in edible oils: (1) Fluorescence interference problem: Pigments and impurities often contained in edible oils and their adulterants cause strong fluorescence background, which completely drowns out the weak Raman characteristic signals, making conventional Raman spectroscopy technology insufficient in sensitivity and unable to be applied to actual detection. (2) Accurate identification and quantification problem: After obtaining the effective spectrum, how to quickly and accurately achieve the qualitative judgment of "whether it is adulterated" and the quantitative (or semi-quantitative) analysis of "how much adulteration" from complex high-dimensional data is a requirement for rapid on-site detection that is difficult to reliably complete by existing chromatographic methods (complex and time-consuming pretreatment) and single algorithm models.

[0007] Technical solution

[0008] Purpose of the invention: To address the shortcomings of existing technologies, this application provides a rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning. The method includes the following steps:

[0009] S1 is used to prepare and pre-treat adulterated standard samples. The adulterant is mixed with pure edible oil according to a preset gradient to simulate samples with different degrees of adulteration. The mixed adulterated standard samples are defoamed to obtain the pre-treated adulterated standard sample solution.

[0010] S2 utilizes differential Raman spectroscopy to acquire high-quality spectral data. Using a portable differential Raman spectrometer, optimized dual-wavelength excitation, laser power, and integration time are set to perform differential Raman spectral scanning on adulterated samples and acquire high-quality spectral data.

[0011] S3 performs spectral data processing. After acquiring high-quality spectral data, it processes the spectral data, including data cleaning to remove low wavenumber bands with low signal-to-noise ratio, using algorithms to perform baseline correction to eliminate drift, and performing Z-score standardization on the entire spectrum to offset the influence of environmental factors, so as to obtain clean and consistent spectral data for model analysis.

[0012] S4 Characteristic peak extraction based on chemical knowledge: Based on the molecular structure characteristics of edible oils and adulterants, the position, intensity and area of ​​characteristic Raman peaks of edible oils and adulterants are identified and recorded, and a feature vector with physical meaning is constructed; Data-driven feature dimensionality reduction and extraction: Principal component analysis is performed on the spectral data processed in step S3, and the scores of the main principal components are extracted as features to achieve data dimensionality reduction and retain the main variation information in the spectrum.

[0013] S5 constructs a multi-model machine learning framework, selects and trains corresponding models for different analysis objectives, inputs the feature information extracted in step S4 into the model for training, thereby forming an analysis system, and selects the optimal model based on evaluation indicators.

[0014] The output of model S6, the optimal performance model obtained through rigorous evaluation in step S5, can quickly identify unknown edible oil samples, outputting a qualitative conclusion on whether they are adulterated and a semi-quantitative prediction of the adulteration concentration.

[0015] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0016] 1. It fundamentally solves the problem of fluorescence interference in Raman detection of edible oils, and realizes the acquisition of high-quality spectral data.

[0017] Traditional Raman spectroscopy is unusable for detecting edible oils because the strong fluorescence background completely drowns out the characteristic peaks. This invention creatively introduces differential Raman spectroscopy as a front-end detection method. Through dual-wavelength differential operations, it actively and in real-time suppresses and even subtracts fluorescence background (such as...) at the hardware level. Figure 1 , Figures 3 to 4 (The calibration process is shown in the diagram). This makes it possible to directly obtain "pure" Raman spectra with high signal-to-noise ratio and clear characteristic peaks, laying an irreplaceable and reliable data foundation for all subsequent quantitative and model-based analyses. This is the primary prerequisite and core advantage of this method.

[0018] 2. A multi-model intelligent decision-making system for multiple identification tasks was constructed, which significantly improved the identification accuracy and reliability.

[0019] This invention transcends the limitations of single-model applications, innovatively constructing a "task-adaptive" multi-model machine learning fusion framework. This framework intelligently selects and coordinates the optimal model combination based on different identification requirements, achieving multifunctional and high-precision analysis.

[0020] In qualitative identification (whether it is adulterated): using a random forest classifier as the core, an overall classification accuracy of 74.3% was achieved in the case of rapeseed oil being adulterated with kerosene.

[0021] In semi-quantitative analysis (adulteration concentration): using a random forest regressor as the core, the example demonstrates high prediction accuracy for medium to high concentrations of adulteration, as well as strong nonlinear fitting ability, clearly verifying that concentration information must be extracted by a supervised model, highlighting the scientific nature and necessity of the model framework design.

[0022] Regarding algorithm robustness: ensemble models such as random forests can effectively avoid overfitting and show stable prediction errors for edible oil samples from different origins and brands (such as rapeseed oil from five different origins), proving the universality of the method.

[0023] 3. Standardized preprocessing and innovative feature engineering ensured the accuracy and reproducibility of the analysis results.

[0024] This method establishes a complete and rigorous data preprocessing and feature extraction pipeline, the advantages of which are:

[0025] Advantages of data standardization: Through unified Z-score standardization, baseline correction, and invalid band removal, spectral differences caused by instrument status, laser power fluctuations, and environmental factors are completely eliminated. Reproducibility and uniformity experiments in the examples demonstrate that the relative standard deviation of the processed data is generally less than 10%, ensuring high consistency and comparability of data acquired at different times and by different operators.

[0026] Innovative Advantages of Feature Engineering: A two-way feature strategy combining "full-spectrum data-driven" and "key feature peak knowledge-driven" approaches was proposed. This not only preserves complete spectral information for the model to uncover deeper patterns but also programmatically extracts feature peak parameters with clear chemical significance, such as C=O and C=C bonds. This approach of injecting domain-prior knowledge into the model greatly enhances its interpretability and guides the algorithm to focus on the chemical signals most relevant to adulteration, avoiding the data blindness of "black box" models and improving the targeting and accuracy of detection.

[0027] 4. A complete solution for efficient, non-destructive testing suitable for rapid on-site inspection has been developed.

[0028] This invention integrates portable differential Raman hardware, standardized sample preparation procedures, and embedded intelligent algorithm software, forming a complete solution suitable for rapid on-site detection.

[0029] Highly efficient and fast: The entire identification process can be completed in minutes, from sample preparation to obtaining qualitative and semi-quantitative results, which is much faster than chromatographic methods that require complex pretreatment.

[0030] Non-destructive and environmentally friendly: The testing process does not require the consumption or destruction of samples, nor does it require the use of chemical reagents, thus achieving green testing.

[0031] Highly applicable to the field: From the use of portable differential Raman equipment to standardized preprocessing procedures, and the integration of trained machine learning models back into hardware via API, this invention provides a fast, non-destructive, and efficient end-to-end detection solution that can be applied to market supervision sites, from sample to result, and has extremely high industrial transformation value.

[0032] 5. The advantages of this invention are also reflected in the following multi-level integration and innovation:

[0033] Front-end technology innovation: In response to the inherent bottlenecks in Raman detection of edible oils, we creatively selected and adapted differential Raman spectroscopy as the only feasible front-end signal acquisition scheme, which solved the fundamental obstacle of fluorescence interference and laid an irreplaceable data foundation for all subsequent analyses.

[0034] Innovative Analytical Framework: A "task-adaptive" multi-model machine learning fusion decision framework was constructed. Based on different analytical objectives such as qualitative, quantitative, and exploratory analysis, this framework intelligently calls and integrates the optimal model combination (e.g., random forest as the primary model, supplemented by PCA and clustering), realizing a complete analytical chain from "whether there is fraud" to "how much fraud is involved," transcending the limitations of a single model.

[0035] Methodological innovation: A two-way feature engineering strategy of "deep integration of domain knowledge (characteristic peaks) and data-driven (full spectrum)" is proposed. This not only improves the interpretability of the model and its ability to capture core chemical information, but also effectively avoids the blindness of purely data-driven models, representing an organic combination of chemometrics and artificial intelligence.

[0036] System integration innovation: This resulted in an end-to-end rapid on-site detection solution integrating a portable differential Raman spectroscopy device, standardized sample preparation procedures, and embedded multi-model intelligent algorithms. This system simplifies complex laboratory analysis processes into rapid on-site operations, offering non-destructive, fast, and accurate results, and possesses significant value for market supervision and on-site screening applications. Attached Figure Description

[0037] Figure 1 The differential Raman spectrum of pure rapeseed oil after drift removal;

[0038] Figure 2 The differential Raman spectrum after Z-score normalization;

[0039] Figure 3 The raw Raman data were randomly selected from CH1-D1 with 2% doping.

[0040] Figure 4 Y-axis intensity correction processing for the original Raman data of CH1-D1 with 2% doping;

[0041] Figure 5 Flowchart for loading preprocessing data for raw Raman data;

[0042] Figure 6 This is an overall flowchart of the rapid identification method for adulterated edible oils described in this invention;

[0043] Figure 7 The confusion matrix for the random forest test set;

[0044] Figure 8 Elbow rule (optimal k=3);

[0045] Figure 9 Distribution diagram of 3D clustering results;

[0046] Figure 10 Clustering results (a) Concentration distribution (b) Comparison with oil type (c);

[0047] Figure 11 Principal component analysis of CH1 & D1~6; Crusher plots;

[0048] Figure 12 Eigenvalue summary chart;

[0049] Figure 13 Load the reference spectrum;

[0050] Figure 14 Score graph (2D);

[0051] Figure 15 Score chart (3D);

[0052] Figure 16 Random forest confusion matrix;

[0053] Figure 17 K-means clustering results;

[0054] Figure 18 Agglomerative hierarchical clustering results (3D). Detailed Implementation

[0055] The preferred embodiments of the present invention will now be described in detail with reference to specific examples. It should be understood that the following examples are given for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art can make various modifications and substitutions to the present invention without departing from its spirit and essence.

[0056] Unless otherwise specified, the experimental methods used in the following examples are conventional methods.

[0057] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.

[0058] One embodiment of this application provides a rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning. The method includes the following steps:

[0059] S1 is used to prepare and pre-treat adulterated standard samples. The adulterant is mixed with pure edible oil according to a preset gradient to simulate samples with different degrees of adulteration. The mixed adulterated standard samples are defoamed to obtain the pre-treated adulterated standard sample solution.

[0060] S2 utilizes differential Raman spectroscopy to acquire high-quality spectral data. Using a portable differential Raman spectrometer, optimized dual-wavelength excitation, laser power, and integration time are set to perform differential Raman spectral scanning on adulterated samples and acquire high-quality spectral data.

[0061] S3 performs spectral data processing. After acquiring high-quality spectral data, it processes the spectral data, including data cleaning to remove low wavenumber bands with low signal-to-noise ratio, using algorithms to perform baseline correction to eliminate drift, and performing Z-score standardization on the entire spectrum to offset the influence of environmental factors, so as to obtain clean and consistent spectral data for model analysis.

[0062] S4 Characteristic peak extraction based on chemical knowledge: Based on the molecular structure characteristics of edible oils and adulterants, the position, intensity and area of ​​characteristic Raman peaks of edible oils and adulterants are identified and recorded, and a feature vector with physical meaning is constructed; Data-driven feature dimensionality reduction and extraction: Principal component analysis is performed on the spectral data processed in step S3, and the scores of the main principal components are extracted as features to achieve data dimensionality reduction and retain the main variation information in the spectrum.

[0063] S5 constructs a multi-model machine learning framework, selects and trains corresponding models for different analysis objectives, inputs the feature information extracted in step S4 into the model for training, thereby forming an analysis system, and selects the optimal model based on evaluation indicators.

[0064] The output of model S6, the optimal performance model obtained through rigorous evaluation in step S5, can quickly identify unknown edible oil samples, outputting a qualitative conclusion on whether they are adulterated and a semi-quantitative prediction of the adulteration concentration.

[0065] In one embodiment, the defoaming treatment of the mixed adulterated standard involves centrifuging and shaking the mixed adulterated standard and allowing it to stand in the dark to eliminate air bubbles and ensure the uniformity and stability of the sample.

[0066] In one embodiment, the differential Raman spectroscopy is performed by automatically subtracting strong fluorescence background interference from the original Raman signal through dual-wavelength differential operations, thereby outputting a Raman spectrum with clear characteristic peaks and a high signal-to-noise ratio. The reproducibility and uniformity of the data are verified by repeated measurements of the same and different regions.

[0067] In one embodiment, in step S3, baseline correction is performed using an algorithm to eliminate drift, including baseline correction by polynomial fitting to eliminate drift.

[0068] In one embodiment, in step S4, the structural features of the edible oil and the adulterant include vibrational peaks of C=O and C=C bonds in the edible oil and vibrational peaks of aromatic CH bonds in the kerosene.

[0069] In one embodiment, in step S5, for different analysis objectives, the corresponding models are selected and trained, including using principal component analysis for dimensionality reduction and preliminary exploration; using random forest model for qualitative classification of whether there is adulteration and semi-quantitative regression prediction of the amount of adulteration; and supplemented by K-means clustering, artificial neural network and other models for comparative verification.

[0070] In one embodiment, the data is input into different machine learning models or combinations of models for analysis, depending on the different identification targets:

[0071] For qualitative identification, to determine whether adulteration has occurred: the first choice is a random forest classifier for fast and accurate discrimination; at the same time, principal component analysis can be used to reduce dimensionality and visualize, intuitively verifying the separability of adulterated samples from pure oil samples.

[0072] For semi-quantitative analysis, to estimate adulteration concentration, the first choice is a random forest regressor to establish a nonlinear mapping relationship from spectral features to adulteration concentration, thereby achieving accurate prediction of medium to high concentrations of adulteration. At the same time, artificial neural networks can be used as a comparative model to evaluate the performance of different algorithms.

[0073] For unsupervised exploration, the inherent structure of data can be discovered by using methods such as K-means clustering to explore the clustering patterns of the data itself without relying on sample labels. This can be used to verify the conclusions of supervised models or to discover new data insights.

[0074] In one embodiment, in step S5, the unknown edible oil sample is treated as follows before detection: similar to the adulterated standard in step S1, the oil sample to be tested is centrifuged and allowed to stand in the dark to eliminate air bubbles and ensure the uniformity and stability of the sample, thereby ensuring the consistency of its conditions in subsequent spectral acquisition and model analysis.

[0075] In one embodiment, the model's output includes validating the trained model using test set data and outputting the final discrimination result:

[0076] For qualitative models, the output is a judgment of whether the oil is pure or adulterated, and evaluation indicators such as accuracy and confusion matrix are given.

[0077] For quantitative models, the predicted values ​​of doping concentrations are output, and evaluation indicators such as the coefficient of determination R² and root mean square error RMSE are given to clarify the predictive performance of the model in different concentration ranges.

[0078] In one embodiment, the random forest model is used to process feature information extracted from the spectrum. By constructing and integrating multiple decision trees, it can respectively achieve qualitative classification of whether edible oil is adulterated and semi-quantitative regression prediction of adulteration concentration.

[0079] The K-clustering algorithm is used to explore differential Raman spectral data in groups under unsupervised conditions to discover whether there are potential cluster structures related to doping in the spectral data;

[0080] Artificial neural networks (ANNs) simulate nonlinear mapping relationships to perform deep learning and pattern recognition of spectral features, thereby assisting in classification or regression prediction tasks.

[0081] In one embodiment, the use of principal component analysis for dimensionality reduction and preliminary exploration includes PCA as a data statistical tool, which involves finding intrinsic features from multiple input data points and integrating them functionally to obtain several features ranked from high to low, with features appearing earlier being stronger. The PCA process includes:

[0082] (1) Data preparation: The preprocessed (e.g., baseline removal, standardization) differential Raman spectral data is used as input, and each spectrum is regarded as a high-dimensional vector.

[0083] (2) Covariance matrix calculation: Calculate the covariance matrix between all spectral data to measure the correlation between wavelength points.

[0084] (3) Eigenvalue and eigenvector extraction: Eigenvalues ​​and their corresponding eigenvectors (i.e., principal components) are obtained by performing eigenvalue decomposition on the covariance matrix. The magnitude of the eigenvalues ​​reflects the variance of the data that the principal component can explain.

[0085] (4) Dimensionality reduction and feature extraction: Sort the feature values ​​from largest to smallest, select the first k principal components (usually with a cumulative contribution rate of > 85%–95%), and project the original high-dimensional spectral data into a new low-dimensional space composed of these k principal components to achieve data dimensionality reduction.

[0086] (5) Preliminary exploration and visualization: Use the data after dimensionality reduction (such as the first 2-3 principal components) to draw a score map, which can intuitively show the distribution and separation of different types of samples (such as pure oil and adulterated oil) in the low-dimensional space, and help to judge the distinguishability of the data.

[0087] One embodiment of this application provides an end-to-end solution for "differential fluorescence suppression - data standardization processing - feature-oriented extraction - multi-model task-oriented analysis". Its core lies in the pioneering deep integration of differential Raman spectroscopy with a task-adaptive multi-model machine learning framework, constructing a complete, rapid, and non-destructive detection system from sample to intelligent identification result. Specifically, it consists of two main stages: an offline modeling stage and an online detection stage.

[0088] 1. Offline modeling stage: Building a standard database and training an intelligent analysis model

[0089] The goal of this stage is to establish a correlation model between the "spectral characteristics and adulteration information" of the target edible oil and adulterants.

[0090] Step S1: Construction of a standardized adulterated sample library

[0091] A series of samples were prepared, containing target pure edible oils (such as rapeseed oil and soybean oil) and typical adulterants (such as kerosene and diesel oil). These samples were precisely mixed according to a preset gradient (e.g., 0.5%, 1%, 2%, 5%, 10%) to simulate different levels of adulteration. All mixed samples underwent uniform physical pretreatment (e.g., centrifugation and light-protected settling) to eliminate air bubbles and ensure sample homogeneity and stability, forming a standard sample set with known concentration labels.

[0092] Step S2: High-quality differential Raman spectroscopy acquisition

[0093] Using a portable differential Raman spectrometer, all standard samples were scanned under optimized parameters (e.g., excitation wavelength around 785 nm, power 440 mW, integration time 3000 ms). The key innovation of this step lies in utilizing the dual-wavelength excitation and real-time differential operation of DRS to actively subtract the strong fluorescence background from the original signal, directly outputting a "pure" Raman spectrum with clear characteristic peaks and a high signal-to-noise ratio (e.g., ...). Figure 1 This fundamentally solves the data quality bottleneck in subsequent analysis.

[0094] Step S3: Spectral data normalization preprocessing and two-way feature engineering

[0095] The acquired raw differential spectra are standardized to provide consistent and clean data input for machine learning: Data cleaning: data intervals with low wavenumbers (e.g., <680 cm⁻¹) and excessively low signal-to-noise ratios or severe noise interference are removed; Baseline correction: algorithms such as polynomial fitting are used to remove baseline drift in the spectrum; Full spectrum standardization: Z-score standardization is performed on the spectrum to eliminate intensity differences caused by instrument fluctuations and environmental factors; Bidirectional feature extraction (one of the core innovations): Strategy A (data-driven): the preprocessed complete spectral data points are used as input, retaining all spectral morphology information; Strategy B (knowledge-driven): based on the molecular structure knowledge of edible oils and adulterants, key Raman peaks (e.g., the C=O peak of ester bonds in edible oil at 1745 cm⁻¹) are identified, located, and quantified programmatically. -1 The aromatic CH peak of kerosene is at 835 cm⁻¹. -1 The model incorporates parameters such as location, intensity, and area to construct feature vectors with clear physicochemical significance. This approach injects domain knowledge into the model, significantly enhancing its interpretability and ability to focus on key difference signals.

[0096] Step S4: Construction and Training of a Multi-Task-Oriented Multi-Model Machine Learning Framework

[0097] For different identification targets, the optimal combination of models is adaptively selected and trained to form a hierarchical intelligent analysis system.

[0098] (1) Qualitative identification task (whether it is adulterated):

[0099] Core Model: A random forest classifier is used, which leverages its powerful non-linear classification ability and anti-overfitting properties to train a model that can quickly distinguish between "pure oil" and "adulterated oil".

[0100] Auxiliary verification: Principal component analysis was used to reduce the dimensionality of the spectrum, and the natural separability of different categories of samples in the feature space was verified by visualizing the 2D / 3D score map, providing spatial distribution evidence for qualitative conclusions.

[0101] (2) Semi-quantitative analysis task (adulteration concentration):

[0102] Core Model: A random forest regressor is used to establish a nonlinear mapping model from spectral features (especially the feature peak parameters extracted by strategy B) to doping concentration.

[0103] Performance Comparison and Validation: Artificial neural networks can be used as comparison models, and the model with the best prediction performance can be selected through evaluation indicators (such as coefficient of determination R², root mean square error RMSE). Experiments show (see examples) that the random forest regression model can achieve a prediction accuracy of over 90% for medium to high concentrations of adulteration (>6%).

[0104] Unsupervised exploration and model cross-validation: Unsupervised learning methods such as K-means clustering are used to explore the inherent clustering patterns of the spectral data without using sample labels. This step serves to demonstrate the necessity of supervised models (e.g., the example shows that the clustering results are mainly dominated by the brand of cooking oil rather than concentration), highlighting the scientific nature of the "task-oriented" model selection in this invention.

[0105] 2. Online detection phase: rapid identification of unknown samples

[0106] The model trained and evaluated offline is integrated into the software system to enable rapid analysis of unknown edible oil samples.

[0107] Step S5: Sample pretreatment and spectral acquisition

[0108] The edible oil samples to be tested underwent the same standardized pretreatment (centrifugation, settling) as S1. Spectra were acquired using the same differential Raman spectrometer under the same optimized parameters.

[0109] Step S6: Real-time data processing and feature extraction

[0110] The system automatically performs the same preprocessing (cleaning, baseline correction, standardization) and bidirectional feature extraction on the collected unknown sample spectra as S3.

[0111] Step S7: Intelligent Model Judgment and Result Output

[0112] The extracted features are first input into the trained optimal qualitative model (random forest classifier), which immediately outputs a qualitative conclusion of "pure oil" or "adulterated oil", and can be accompanied by classification confidence.

[0113] If the qualitative result is "adulteration", the system can further input the features into the trained optimal quantitative model (random forest regressor) and output the predicted value of the adulteration concentration and its possible confidence interval.

[0114] The entire analysis process can be completed in minutes, ultimately providing a test report that includes qualitative conclusions, quantitative predictions, and key evaluation indicators.

[0115] Sample pretreatment

[0116] The edible oils (6 batches) in this study were sourced from different production areas, including newly packaged or currently used finished soybean oils in restaurants; No. 0 diesel oil was sourced from Sinopec gas stations; rapeseed oil (5 batches) and kerosene (3 batches) were purchased from the Taobao platform.

[0117] Sample preparation: Rapeseed oil system: 5 brands of rapeseed oil × 3 types of kerosene × 10 concentration gradients (0.5%~10%), totaling 150 samples. Soybean oil system: 6 brands of soybean oil × diesel (No. 0), prepared with gradients of 0.5%, 1%, 2%, 5%, and 10%, totaling 37 samples.

[0118] Pretreatment: Centrifuge for 10 minutes, then let stand in the dark for 6 hours to defoam.

[0119] Differential Raman Spectroscopy Acquisition Optimization

[0120] Preprocessing and differential spectral acquisition

[0121] All prepared oil samples underwent standardized pretreatment: centrifugation for 10 minutes, followed by standing in the dark for 6 hours to completely eliminate air bubbles and ensure sample homogeneity and stability. Detection was performed using a portable differential Raman spectrometer (such as those from Nanjing Jianzhi Instrument Equipment Co., Ltd.). The core technology utilizes its dual-wavelength differential technique, with optimized parameters set: excitation wavelengths of 784.5 nm and 785.5 nm, laser power of 440 mW, and integration time of 3000 ms. The instrument automatically performs differential calculations, directly outputting a pure Raman spectrum with a high signal-to-noise ratio and effectively suppressed fluorescence background. Figure 1 , Figure 2 As shown.

[0122] Spectral data preprocessing and two-way feature engineering

[0123] Standardized data preprocessing workflow

[0124] Data cleaning and invalid band removal: Due to low wavenumber bands (<680 cm⁻¹) -1 Low signal-to-noise ratio and severe noise interference mean that data in this band should be uniformly removed (e.g., Figure 1 (As shown), retain 680-2800 cm. -1 The effective spectral range.

[0125] Baseline correction: A polynomial fitting algorithm (e.g., power of 5) is used to remove the baseline from the spectrum and correct baseline drift (see the processing results). Figure 3 , Figure 4 ).

[0126] Standardization: Z-score standardization is performed on the full-spectrum data to eliminate intensity differences caused by laser power fluctuations, instrument conditions, etc., allowing the model to focus on spectral shape and relative intensity information (see processing results). Figure 2 ).

[0127] Instrument parameters

[0128] The experiment used a portable differential Raman spectrometer (Nanjing Jianzhi Instrument Equipment Co., Ltd.). The differential Raman spectrometer probe emits laser light focused at a single point. Based on the properties of the adulterated oil, parameters were continuously adjusted experimentally. It was found that when the excitation wavelength was 785 (±1.5) nm, the measurement power was 440 mW, and the integration time was 3000 ms, the spectral wavenumber range was 192-2800 cm⁻¹. -1 At that time, the obtained data was relatively stable.

[0129] Reproducibility and homogeneity experiments

[0130] To verify the homogeneity of the samples and the accuracy of the instruments, and to ensure the accuracy and reliability of the experimental data, one sample from each of the two oils was randomly selected for reproducibility and homogeneity tests. Five tests were performed at the same location on the selected samples, and one test was performed at each of five different locations. The relative standard deviation (RSD) was then calculated.

[0131] Table 1. Reproducibility test results (peak intensity) of randomly selected soybean oil CH1-D1 2% doped samples.

[0132] 1 851.06 1026.31 1895.47 5698.86 5072.08 2 898.68 1054.45 1868.09 5995.83 5332.84 3 1056.34 1030.06 1945.74 6050.50 5356.75 4 1007.37 1122.61 1949.85 5964.64 5517.11 5 993.10 1064.46 1934.39 5918.82 5234.05 6 1089.32 993.74 1845.85 5734.95 5245.81 7 982.33 1130.64 1856.73 5942.53 5329.63 8 902.63 981.52 1927.29 5588.21 5181.34 9 1071.90 970.48 1858.24 5547.63 5119.88 10 1063.74 922.06 1761.25 5382.20 5053.17 RSD 8.37% 6.42% 3.11% 3.90% 2.76%

[0133] Table 2. RSD results at the same location in rapeseed oil samples.

[0134] 1 1130 1685 2884 2 1160 1638 3020 3 1045 1545 2770 4 1114 1521 2753 5 1176 1677 3095 RSD value / % 4.53 4.70 5.20

[0135] Table 3 RSD results at different locations

[0136] Top left 1169 1694 3081 Bottom left 999 1467 2659 center 1124 1718 3451 Top right 940 1432 3061 Bottom right 1176 1608 3131 RSD value / % 9.83 8.21 9.16

[0137] The homogeneity of the sample and the accuracy of the instrument were verified by selecting the intensity of several characteristic peaks based on the frequency and confidence level of the Raman results. The relative standard deviation (RSD) of the characteristic peak intensities was calculated, and the results are shown in Tables 1 and 2. All RSD values ​​were less than 10%, confirming that the equipment was functioning well and that the experimental data had high repeatability and accuracy. The relative standard deviation (RSD) of the differential Raman spectra from five different locations was calculated, and the results are shown in Table 3. It can be seen that the peak shape, number, and position of the characteristic Raman peaks in the samples from different locations are basically consistent, and the relative standard deviation (RSD) is less than 10%, which confirms that the sample has high homogeneity and accuracy.

[0138] Data preprocessing workflow

[0139] Signal preprocessing is a crucial step in improving the accuracy of data analysis. Baseline removal and noise reduction are automatically performed before the differential Raman instrument performs differential calculations. The data processed by differential calculations already has the advantage of removing fluorescence background, making it the primary data source for doping detection models.

[0140] Spectral preprocessing of rapeseed oil samples

[0141] In the Raman spectral data of rapeseed oil, the small scattering cross-section corresponding to low-frequency vibrations in the low-shift range, coupled with potentially strong Rayleigh scattering noise (Rayleigh scattering noise, as a type of elastic scattering, mainly concentrated near the low-shift range), easily masks the effective signal, resulting in significant noise effects. Furthermore, due to instrument resolution limitations, adjacent peaks are prone to overlap or broadening, leading to signal fluctuations. These noise effects and signal fluctuations are particularly pronounced in rapeseed oils No. 2 and No. 5. Therefore, Raman data in the low-shift range (Raman shift < 680 cm⁻¹) lacks universality and representativeness and should be excluded. Figure 1 Finally, the measured Raman spectral data were Z-score normalized to further eliminate the influence of environmental factors such as instrumentation. Figure 2 ).

[0142] Spectral preprocessing of soybean oil samples

[0143] In the raw Raman data ( Figure 3 Especially in the low-frequency region (192~800cm⁻¹), due to differences in instrument response, the baseline drift is large, the peak shape is not significant, and it is greatly affected by fluorescence interference, requiring further processing. Baseline removal was achieved through a polynomial with a power of 5, and the Y-axis intensity in the low-frequency region was corrected to obtain the figure (Figure 1). Figure 4 The same operation resulted in the corrected data from the original Raman data of the entire sample set.

[0144] To implement machine learning, import the raw Raman data as follows (see...). Figure 5 The data loading process is implemented using Python 3.13.2. Smoothing refers to low-pass filtering of the spectral curve to reduce high-frequency noise, preserve low-frequency information, and improve the signal-to-noise ratio. Common methods include Savitzky-Golay (SG), wavelet transform (WT), and moving average smoothing. Scattering correction can eliminate the influence of scattered light on the detection process. Common methods include multivariate scattering correction (MSC) and standard normal variable transformation (SNV).

[11] The `Read_differential data` function is responsible for reading all data files, converting them into numerical arrays, and then performing filtering operations to prepare them for PCA (Principal Component Analysis) analysis together with the differential Raman data.

[0145] Multi-model fusion analysis framework (see flowchart for details) Figure 6 )

[0146] Principal Component Analysis (PCA)

[0147] Using principal component analysis (PCA) for linear dimensionality reduction can reduce computational costs and remove redundant information. By merging differential Raman spectral data and baseline-removed raw Raman spectral data, features such as peak position, intensity, and area can be extracted more effectively than from the raw spectral data.

[0148] Random Forest Analysis

[0149] Random Forest (RF) is a classic supervised learning algorithm. It is a nonlinear data processing algorithm based on autonomous resampling technology to construct regression tree classifiers. The random forest model randomly selects some variables in the original data to form several different regression tree classifiers. By building a decision tree model for each group of data, all decision trees are combined into a random forest model. The final result is determined based on the average value of these different classifier models [9].

[0150] K-means clustering analysis

[0151] K-means clustering is a classic unsupervised learning algorithm that divides n objects in a dataset into K clusters through an iterative approach

[10] . The goal is to make the data points within the same cluster as similar as possible, while the data points between different clusters are as different as possible. Each object is assigned to the nearest cluster center, with the aim of minimizing the sum of the distances from each object to the center of its cluster. Euclidean distance is often used as the relevant metric.

[0152] Artificial Neural Networks (ANN)

[0153] Artificial Neural Networks (ANNs) are a crucial algorithm in deep learning. They simulate the "processing-transmission-generation" model of neurons in the human brain to classify and predict large-scale data, with particular advantages in identifying the content of the input. We standardize the data and then build an artificial neural network to achieve basic modeling.

[0154] Innovative bidirectional feature extraction method

[0155] This invention employs a two-way feature strategy. First, the preprocessed complete spectral data points are used as input, retaining all original information. Then, key characteristic Raman peaks of edible oils and adulterants are programmatically identified, located, and quantified from the preprocessed spectrum.

[0156] (1) Characteristic peaks of rapeseed oil and kerosene

[0157] The differential Raman characteristic peaks of rapeseed oil are summarized in Table 4, mainly concentrated at 1745 cm⁻¹. -1 (Ester bond carbonyl stretching vibration peak), 1655 cm⁻¹ -1(Vibration peak of cis(C=C) group in unsaturated fatty acids), 1438 cm⁻¹ -1 (Methylene scissor vibration peak), 1301 cm⁻¹ -1 (Methylene curl vibration peak), 1265 cm -1 (Vibration peak of cis(=C—H) group in unsaturated fatty acids), 1190 cm⁻¹ -1 (Carbon-carbon single bond stretching vibration), 1078 cm -1 and 868 cm -1 (Methylene long chain skeleton stretching vibration peak) and other positions.

[0158] Table 4 Characteristic Raman peaks of rapeseed oil

[0159] 1745 Stretching vibration peak of carbonyl group (C=O) in ester bond The characteristic peaks of triglycerides reflect the vibrations of ester bonds formed between fatty acids and glycerol. 1655 Vibration peak of cis (C=C) group in unsaturated fatty acids The characteristic stretching vibration peaks of cis double bonds are commonly found in edible oils containing unsaturated fatty acids, such as rapeseed oil, reflecting the presence and configuration of the double bonds. 1438 <![CDATA[Scissoring vibration peak of methylene (—CH2—)]]> In-plane bending vibration peak of methylene groups in long-chain fatty acids 1301 <![CDATA[The curling vibration peak of methylene (—CH2—)]]> The out-of-plane bending vibrations of the methylene group in the molecular chain reflect conformational changes in the fatty acid chain (such as ordered arrangement or disordered coiling). 1265 Vibration peak of cis (=C—H) group in unsaturated fatty acids The out-of-plane bending vibration peak of the adjacent C-H bond in the cis double bond is conjugate with the vibration of the cis (C=C) bond. 1190 Stretching vibrations of carbon-carbon single bonds (C—C) Stretching vibrations of carbon-carbon single bonds in fatty acid chains 1078 <![CDATA[Stretching vibration peak of the methylene long-chain backbone (—CH2—)]]> n > Symmetric stretching vibrations of the carbon chain backbone of long-chain fatty acids 1008 <![CDATA[In-plane rocking vibration of methyl group (—CH3)]]> In-plane vibrations of the terminal methyl group of fatty acid chains 970 Bending vibration peak of trans(C=C) out-of-plane bending vibration peak of trans double bond 868 <![CDATA[Stretching vibration peak of the methylene long-chain skeleton (—CH2—) n > Asymmetric stretching vibrations of the carbon chain backbone of long-chain fatty acids

[0160] The differential Raman characteristic peaks of kerosene are summarized in Table 5, mainly concentrated at positions such as 835 cm⁻¹ (CH bond bending vibration in aromatics), 956-963 cm⁻¹ (C=C bond stretching vibration in alkenes), 1446 cm⁻¹ (CH bond bending vibration in alkanes), 730-759 cm⁻¹ (CH bond bending vibration in aromatics), and 1066-1072 cm⁻¹ (CO bond stretching vibration in alkanes). The characteristic peaks of kerosene primarily reflect the vibrational characteristics of the CH, C=C, and C=C bonds in hydrocarbons such as alkanes, aromatics, and alkenes.

[0161] Table 5 Characteristic Raman peaks of kerosene

[0162] 835 Aromatic CH bond Bending vibration characteristics of aromatic hydrocarbons (such as benzene and toluene) in kerosene 956-963 C=C bond in olefins Reflects the presence of unsaturated hydrocarbons in kerosene 1448 alkane C-bond Key characteristic peaks of alkanes in kerosene, reflecting their bending vibrations 730-759 Aromatic CH bond Supplementing characteristic information of aromatics in kerosene 1066-1072 alkane CO bond This indicates that the kerosene may contain a small amount of oxygen-containing compounds.

[0163] These extracted quantitative parameters are then used to construct a structured feature vector.

[0164] (2) Characteristic peaks of soybean oil and diesel oil

[0165] The characteristic peaks of the diesel sample are quite messy at low frequencies under the differential algorithm. After processing the original Raman data, they are summarized in Table 6.

[0166] Table 6. Characteristic peaks (characteristic peak intensities) of diesel sample No. 0

[0167] 1 1375.87 5466.14 411.55 2 1429.91 5530.70 359.71 3 1394.94 5461.16 421.42 4 1401.76 5450.33 360.20 5 1352.09 5481.84 419.44

[0168] The above is a list of characteristic peaks after differencing randomly selected pure CH1 samples (see Table 6), 1300cm -1 The nearby Raman peak corresponds to the deformation vibration of hydrocarbon groups; 1446 cm⁻¹ -1 The corresponding scissor vibration of the hydrocarbon C-H bond (CH) is 1660 cm⁻¹. -1 It corresponds to the stretching vibration of an unsaturated double bond (C=C).

[0169] The characteristic peak region of soybean oil is concentrated at 1263 cm⁻¹. -1 and 1653 cm -1 : Corresponds to the stretching vibration of C=C in unsaturated fatty acids, reflecting the content of polyunsaturated fatty acids such as oleic acid and linoleic acid in soybean oil; 1298 cm -1 and 1437 cm -1 It is related to the CC skeletal vibration of saturated fatty acids and the bending vibration of CH2 / CH3.

[0170] Similarly, these extracted quantitative parameters are used to construct structured feature vectors.

[0171] Qualitative identification (whether adulterated) implementation example: Core model: Random Forest classifier

[0172] Model training: Using the training set data, call the RandomForestClassifier from the Scikit-learn library. Optimize hyperparameters through pre-experiments and cross-validation, for example, setting the number of decision trees (nestimators) to 150 and the maximum depth of each tree (max_depth) to 10.

[0173] Validation and Output: Input the test set into the trained model to obtain the classification prediction result of "pure oil" or "adulterated oil". The output is as follows: Figure 7 The confusion matrix is ​​shown, and overall accuracy, precision, recall, and other metrics are calculated.

[0174] Auxiliary Validation: Visualization of Principal Component Analysis

[0175] Perform PCA: Perform principal component analysis to reduce the dimensionality of the training set spectral data.

[0176] Visual verification: Project the test set data onto the principal component space to generate, for example... Figure 14 , Figure 15 The 2D / 3D score plots shown provide intuitive and auxiliary spatial distribution evidence for qualitative judgments by observing the clustering and separation of samples of different categories in the reduced-dimensional space.

[0177] Semi-quantitative analysis (adulteration concentration) example:

[0178] Core Model: Random Forest Regressor

[0179] Model training: Using the training set data and the corresponding real concentration labels, the RandomForestRegressor in the Scikit-learn library is called to train the regression model to establish a nonlinear mapping from spectral features to continuous concentration values.

[0180] Validation and Output: Concentration prediction was performed on the test set, and key indicators such as the coefficient of determination (R²) and root mean square error (RMSE) were calculated. A key finding of this invention is that the model exhibits extremely high prediction accuracy (accuracy > 90%, R² > 0.85) for medium-to-high concentration adulteration (>6%), while its prediction accuracy for low-concentration samples indicates the need for further integration with feature enhancement algorithms.

[0181] Unsupervised exploratory analysis (used for model comparison and data insights):

[0182] Perform K-means clustering: using preprocessed full-spectrum data, through the "elbow rule" (see...) Figure 8 Determine the optimal number of clusters K (e.g., K=3).

[0183] In-depth results analysis: After running the K-means algorithm, the following results were obtained: Figure 9 , Figure 10 The clustering results are shown below. A key insight of this invention stems from this: through cross-analysis of the concentration distribution of each cluster (see Tables 7 and 8) and the distribution of oil types (see Tables 9 and 10), it was found that the unsupervised clustering results are mainly dominated by the brand / type of the edible oil itself, and have no clear correlation with the adulteration concentration. This finding strongly demonstrates that in quantitative detection tasks, supervised learning models such as random forest regression must be used, rather than relying on unsupervised methods, highlighting the scientific validity and necessity of the "task-oriented" model selection framework constructed in this invention.

[0184] Example of blending rapeseed oil with kerosene (150 samples)

[0185] Random Forest Analysis

[0186] In this experiment, 150 samples were divided into training and test sets in an 8:2 ratio. A random forest model was established, with 150 decision trees specified in the random forest and a maximum depth of 10 for each decision tree.

[0187] The random forest model was used for classification on the test set, achieving an accuracy of 74.29%. The accuracy in identifying doping in low-concentration samples was significantly lower than that in high-concentration samples, indicating its effectiveness in detecting adulterated oil samples. The false positive rate for low-concentration samples was significantly higher than that for high-concentration samples. When the concentration was below 6%, the model's predictions were insufficient to accurately predict doping, especially for 0.5% and 1% doped samples, where the accuracy was below 60%. When the doped sample concentration was above 6%, the model's prediction accuracy was above 90% (see [link to relevant documentation]). Figure 7 Further analysis of the regression model's performance metrics revealed that the coefficient of determination R for the test set... 2 =0.89, cross-validation result R 2=0.74, indicating that the model's prediction results on the test set are better than those on the validation set, but there may be slight overfitting on the training data. Further optimization of the model can be achieved by increasing the sample size to improve generalization ability. The root mean square error (RMSE) for the test set and cross-validation is 0.99 and 1.63, respectively, reflecting that the model has a low prediction error for medium-to-high concentration samples.

[0188] Separate analyses were conducted on rapeseed oil samples from different origins, and a random forest model was used for prediction to verify the model's generalizability. The results showed that the adulteration spectral characteristics of rapeseed oil from the five origins were consistent, and the model prediction errors were all less than 2.5%. This further confirms that the random forest regression model can stably and effectively detect adulteration in rapeseed oil from different origins, demonstrating strong generalizability and providing reliable technical support for rapeseed oil quality testing.

[0189] K-means clustering analysis

[0190] This experiment standardized the acquired differential Raman spectroscopy data and selected K=3 data points as initial cluster centers using the elbow rule. The distance from each data point to each cluster center was calculated, and the data point was assigned to the cluster corresponding to the nearest cluster center. The centroid of each cluster (i.e., the mean of all points in that cluster) was calculated, and the centroid of each cluster was used as the new cluster center. These steps were repeated until the cluster centers no longer changed or the preset maximum number of iterations was reached.

[0191] According to the elbow rule, K=3 is optimal (see...). Figure 8 The clustering results successfully divided the dataset into three categories (see...). Figure 9 The evaluation results of the trained model show that the silhouette coefficient is 0.4508, indicating that the cohesion within clusters and the separation between clusters are at a good level, and the clustering effect is acceptable. However, there are still problems such as some samples being misclassified or the cluster boundaries being blurred. The Calinski-Harabasz index (CHI) is 332.6614, indicating that the dispersion between clusters is high and the compactness within clusters is good.

[0192] Statistical analysis of the concentration distribution of various indicators for each cluster revealed that the concentration differences within clusters were relatively small, but the data volume was uneven in some clusters (see [link]). Figure 10The mean concentration of cluster 0 was significantly lower than that of clusters 1 and 2 (approximately 36% lower), possibly representing a low-concentration sample. The mean concentrations of clusters 1 and 2 were similar, indicating weaker distinguishability. The standard deviations of the concentrations of each cluster fluctuated within a similar range, suggesting that the contribution of concentration variables to clustering may be limited (see Table 7). Table 8 shows that the frequency of occurrence of each concentration in the three clusters was relatively uniform, and no linear or non-linear relationship was found.

[0193] Table 7. Concentration characteristics of each cluster

[0194] 0 84 0.028810 0.027937 1 259 0.045483 0.030245 2 312 0.042003 0.030496

[0195] Table 8. Ranking of Cluster Concentrations

[0196]

[0197] Distribution analysis of rapeseed oil types in each cluster revealed significant differences between clusters. Cluster 0 was highly concentrated in rapeseed oil types 2 (50 types) and 5 (31 types), while types 1 and 3 were absent, indicating that the samples in this cluster were dominated by specific rapeseed oil types. Cluster 1 had a relatively even distribution, with type 5 having the highest proportion (114 types), followed by type 2 (77 types), possibly representing a mixed type or no obvious preference. In contrast, cluster 2 was extremely concentrated in types 1 (154 types) and 3 (153 types), with types 2 and 5 being almost absent, forming a stark contrast with cluster 0. This suggests that rapeseed oil types 1 and 3 are the core characteristics of cluster 2 (Table 9).

[0198] Table 9 Distribution of rapeseed oil in each cluster

[0199] 1 0 3 154 2 50 77 1 3 0 2 153 4 3 63 4 5 31 114 0

[0200] Distribution analysis of kerosene type for each cluster revealed small differences between clusters, with overall balance. The sample sizes of kerosene type 1, 2, and 3 in each cluster showed no significant difference (Table 10, cluster 0 slightly biased towards type 3, cluster 1 slightly biased towards type 1, and cluster 2 evenly distributed). This indicates that kerosene type has a weak distinguishing effect on clustering, in contrast to the strong distinguishing effect of rapeseed oil type.

[0201] Table 10 Distribution of kerosene in each cluster

[0202] 1 18 90 107 2 28 89 102 3 38 80 103

[0203] In summary, the analysis suggests that the trained model based on k-means clustering can classify the samples well, dividing them into three clusters with high inter-cluster dispersion and good intra-cluster compactness.

[0204] Comparative analysis of different models

[0205] First, the preprocessed differential Raman data were plotted and compared visually. It was found that under low-concentration adulteration conditions, the differences in the differential Raman spectra were not obvious. Especially in the absence of a reference standard, it was even more difficult to distinguish whether rapeseed oil was adulterated.

[0206] The collected differential Raman data were then further analyzed using machine learning algorithms. A random forest model was used for supervised learning of the data, establishing classification and regression models respectively. The classification prediction accuracy reached 74.3%, enabling qualitative judgment of whether rapeseed oil was adulterated. The regression model achieved a prediction accuracy of over 90% for kerosene adulteration concentrations of 6% and above, making it suitable for semi-quantitative analysis and capable of effectively predicting and analyzing kerosene adulteration concentrations.

[0207] Examples of soybean oil blended with diesel oil (37 samples)

[0208] This experiment was designed to investigate the qualitative analysis of diesel oil adulteration in soybean oil, while also exploring the feasibility of quantitative analysis of adulteration levels. Therefore, determining the brand and type of oil from unknown samples was not the primary purpose of classification, but rather a factor influencing adulteration analysis. We used prior knowledge to screen for specific diesel characteristic peaks for comparative analysis.

[0209] PCA Algorithm Analysis

[0210] Thirty soybean oil adulteration samples were divided into training and test sets in a 7:3 ratio, with pure diesel oil and soybean oil spectral data used as controls. This resulted in 105 data points in the training set and 45 data points in the test set. Model adjustments were then made to support these adjustments. PCA analysis was performed using OriginPro 2024 (64-bit) SR1 10.1.0.178 software, and the extracted feature factors were incorporated into random forest, K-means clustering analysis, and artificial neural network analysis. Partial least squares regression (PLSR) was also attempted to be used for modeling to detect adulteration levels.

[0211] Principal component analysis was performed on the differential Raman spectral data of No. 0 diesel and various edible oils using the PCA for Spectroscopy app in Origin software to obtain the gravel plots (see [link]). Figure 11 The inflection points of the scree plot are the principal components.

[0212] Combined with the eigenvalue summary chart (see...) Figure 12 As can be seen, the first three components explain 94.24% of the total variance, proving the accuracy of the PCA algorithm.

[0213] Observe the loaded reference spectrum (see Figure 13The first and second principal component loading spectra are visible, with CH1#1 below representing the reference sample spectrum. The largest peak in the loading spectrum was identified using a rapid peak finding tool. With baseline Y=0, the most important peaks for PC1 were found to be at 1439 cm⁻¹. -1 1660cm -1 The important peak of PC2 is at 1453 cm⁻¹. -1 1660cm -1 .

[0214] Through the score graph (see) Figure 14 , 15 As can be seen, the sample data of CH1 (i.e., No. 0 diesel oil) deviates significantly from those of the other components (D1~6) soybean oil, and pure soybean oil also differs from other adulterated samples. This enables the differentiation between adulterated samples and pure diesel oil and pure soybean oil.

[0215] Random Forest Analysis

[0216] This study constructs a soybean oil adulteration classification model based on the random forest algorithm. By dividing the training set and the test set through stratified sampling, it classifies whether soybean oil is adulterated with diesel.

[0217] Classification model performance analysis

[0218] The classification report (see Table 11) shows that the model has a strong ability to identify doped samples, effectively detecting 88% of the actual doped samples, and correctly identifying 88% of the samples predicted as doped. However, all indicators for the pure product category are 0.66, mainly because the test set contains 6 × 10 (tests) pure product samples, resulting in general classification bias. The overall classification accuracy is 78%, reflecting the model's basic classification ability under the existing data distribution. Key classification features were studied (see Table 12), and the model's predictive ability was improved by weighting important doping variation peaks. Different n_components were tried, and the optimal parameters were selected through cross-validation in the model. Further improvements were made to the model performance, achieving R² = 0.63. The confusion matrix was visualized (see...). Figure 16 This further validated the model's ability to identify advantageous categories, providing a preliminary basis for practical doping detection.

[0219] Table 11 Random Forest Classification Report

[0220] Doping 0.88 0.88 0.88 8 Pure 0.66 0.66 0.66 6 accuracy -- -- 0.78 -- Macro average 0.44 0.44 0.44 9 weighted average 0.78 0.78 0.78 9

[0221] Table 12 Key Classification Features

[0222] 790 0.0229 Significant changes in Raman signal intensity after doping serve as a distinguishing feature for doped cores. 648 0.0126 Related to the bending vibration of alkanes in diesel fuel 1904 0.0123 Unsaturated bond (C=C) stretching vibrations as a characteristic of medium- to high-concentration doping

[0223] Regression Model Performance Analysis

[0224] For the doping ratio prediction task, the random forest regression model achieved a determination coefficient (R²) of 0.63 and a mean absolute error (MAE) of 12.4810%, indicating that the model can capture a partially linear relationship between the doping ratio and spectral features, but there is still room for improvement in prediction accuracy. The low R² value (0.63) may be related to the complex noise in the Raman spectral data, the nonlinear mapping between the doping ratio and features, and the uneven distribution of sample concentration. Nevertheless, the model demonstrates a good ability to fit continuous variables in the regression task, and can output prediction results consistent with the actual concentration trend based on spectral features. Combining the comprehensive evaluation of classification and regression models, the detection system constructed in this study has initially achieved the dual functions of "whether it is doped" and "doping ratio discrimination" for soybean oil doping, laying a methodological foundation for further improvements in detection accuracy by optimizing data preprocessing, increasing sample diversity, and adjusting model parameters. The results show that in practical applications, this model needs further data augmentation and hyperparameter tuning to overcome the challenges of sample size limitations and complex spectral features, and achieve more accurate doping detection and quantitative analysis.

[0225] 3.7.3 K-means clustering analysis and agglomerative hierarchical clustering

[0226] For cluster analysis, this experiment was designed to use K-means clustering and agglomerative clustering separately. The primary analytical method was determined by comparing the presence of hierarchical relationships between soybean oil types and adulteration concentrations. The five concentrations labeled on the right correspond to the five clusters on the left (see [link to experimental designation]). Figure 17 The results showed good fitting performance and large relative distances between dissimilar substances, achieving a good clustering effect.

[0227] To explore the hierarchical relationship of predicted doping modes in unknown samples, an agglomerative hierarchical clustering method is employed, which can quickly achieve a preliminary assessment of a large number of unknown samples (see [link to study]). Figure 18 The results of agglomerative hierarchical clustering are presented in a mapping table to facilitate the classification of "cluster-concentration". The results are shown in Table 13.

[0228] Table 13. Cluster-Single Concentration Value Mapping Table for Agglomerative Hierarchical Clustering Results

[0229] 0 1 3 5 7 9 11 13 15 17 19 20 2224 26 28 30 32 34 36 3839 41 4345 47 5777 79 81 83 85 89 '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%''1%' '1%' '1%' '1%''1%' '1%' '1%' '1%' '1%' '1%''2%' '2%' '2%' '2%''2%' '2%''10%' '10%' '10%' '10%' '10%' '10%' 1 58 60 62 64 66 68 70 72 74 76 '5%' '5%' '5%' '5%' '5%' '5%' '5%' '5%' '5%' '5%' 2 59 61 63 65 67 69 71 73 75 '5%' '5%' '5%' '5%' '5%' '5%' '5%' '5%' '5%' 3 1821 23 25 27 29 31 33 35 374951 53 5587 88 '0.5%''1%' '1%' '1%' '1%' '1%''1%' '1%' '1%' '1%' '2%' '2%' '2%' '2%''10%' '10%' 4 0 2 4 6 8 10 12 14 1640 42 4446 48 50 52 54 5678 80 82 84 86 '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%' '0.5%'0.5%''2%' '2%''2%' '2%' '2%' '2%' '2%' '2%' '2%''10%' '10%' '10%' '10%' '10%'

[0230] Based on the model clustering data, K-means clustering analysis has a stronger indicative role. Furthermore, since this experimental design uses doping concentration as the variable, and the number of clusters is known to be five concentrations, K-means clustering analysis is more consistent with prior inferences. This allows for a classification and discussion of the model's differences when different brands of soybean oil are adulterated with the same concentration of diesel oil.

[0231] In summary, by establishing a K-means clustering analysis model, typical characteristics of known samples at different concentrations were clustered. Due to the unsupervised nature of agglomerative hierarchical clustering, the clustering results are of good reference value for similar samples (such as cluster 1, cluster 2, and cluster 3), but the detection effect is poor for highly doped and low-doped samples with large ranges (cluster 0 and cluster 4), and clustering phenomenon exists between high and low concentrations.

[0232] Artificial Neural Networks (ANNs)

[0233] Based on the above data, from the perspective of training loss, the model for determining whether a component is adulterated showed a significant decrease in loss during training, from an initial 0.16136 to 0.00011 after 900 training rounds. This indicates that the model can quickly learn the feature patterns in the data, and its ability to determine whether a component is adulterated continuously improves (Table 14). The model for predicting the adulteration ratio had a loss of 14.64596 at the beginning of training, which stabilized at 14.27595 after 900 training rounds. Although the loss also decreased to some extent, the decrease was relatively small, indicating that the model faces some difficulty in learning the complex relationship between the adulteration ratio and features (Table 15).

[0234] Regarding model evaluation, the model accuracy for determining whether soybean oil is adulterated was 50.00%, indicating that the model's ability to distinguish between soybean oil adulterated with diesel needs improvement. This may be due to insufficiently clear data features or inadequate model complexity to capture the key information about adulteration. The mean square error of the model predicting the adulteration ratio was 25.1091. The high mean square error reflects that the model's prediction of the adulteration ratio is not accurate enough, which may be due to noise in the data or the model structure failing to fully fit the patterns in the data (Table 16).

[0235] Table 14 Training loss for determining whether a model is mixed with other components

[0236] 0 0.1613596000004735 100 0.002779799201795645 200 0.0008049247587731219 300 0.00044548796403600136 400 0.00030239347313443335 500 0.00022683531141295003 600 0.000180541936864678 700 0.00014943574826129197 800 0.0001271748928391147 900 0.00011049766222934941

[0237] Table 15 Training loss of the prediction doping ratio model

[0238] 0 14.64596251088087 100 14.350385286683688 200 14.276926328627143 300 14.276297443131453 400 14.27613118798747 500 14.27605511553155 600 14.276011739780143 700 14.275983812185036 800 14.275964378438761 900 14.27595010269957

[0239] Table 16 Model Evaluation Indicators

[0240] Accuracy in determining doping 50.00% Mean square error of predicted doping ratio 25.1091

[0241] The ANN model has demonstrated some learning ability in detecting soybean oil adulterated with diesel, but its current performance is still unsatisfactory. Further improvements can be made by optimizing data preprocessing methods, adjusting the model structure, and increasing the amount of training data to enhance the model's accuracy and stability, thereby enabling more accurate judgment of soybean oil adulteration and precise prediction of adulteration ratios.

[0242] Feasibility Example of Semi-Quantitative Detection of Doped Samples

[0243] Experimental results show that Principal Component Analysis (PCA) has significant advantages in classifying doping types, but requires a large number of samples to ensure stability. Further optimization is needed for inference on samples with single unknown samples. Overall, it can effectively address the doping issue. The ANN model demonstrated some learning ability in detecting soybean oil adulterated with diesel, but its current performance is still unsatisfactory, requiring further optimization through better data normalization to achieve higher accuracy. Based on the model evaluation metrics, since different methods have different evaluation approaches, the R² score formula and prediction accuracy were used for overall evaluation. Cluster analysis was not used as it cannot focus on a single concentration. The results are shown in Table 17.

[0244] Table 17 Evaluation Indicators for Each Classification Method

[0245]

[0246] *PCA analysis was determined based on the characteristics of the first three principal components.

[0247] **The detection limit of random forest is based on low doping level (1%).

[0248] Combining the above algorithms, PCA achieves good qualitative analysis results by reducing the dimensionality of Raman spectral data and combining the differencing data. K-means clustering and agglomerative clustering demonstrate the reliability of classifying known data samples and regressing unknown data samples, respectively, but their overall accuracy is not high at low concentrations of doping.

[0249] The "differential Raman spectroscopy + multi-model fusion" method established in this application, through optimization of models such as dual-wavelength differential fluorescence suppression, key feature peak extraction, and random forest, achieves rapid qualitative and semi-quantitative analysis of adulteration in edible oils. Case studies show that the detection accuracy for kerosene adulteration concentrations above 6% in rapeseed oil is >74.3%; random forest regression outperforms clustering / ANN models; the detection accuracy decreases for low-concentration adulteration (rapeseed oil < 6%, soybean oil < 1%), requiring the integration of feature enhancement algorithms for low-concentration detection. This method is fast and non-destructive, effectively expanding the qualitative and semi-quantitative testing methods for adulteration content in edible oils, and has certain guiding significance for practice. In the future, the trained random forest model can be integrated into an improved differential Raman spectrometer via an API interface to achieve rapid qualitative and semi-quantitative analysis of adulteration in edible oils, providing a possible analytical approach and favorable technical support for food safety supervision and combating food safety crimes.

[0250] The above are merely preferred embodiments of the present invention. It should be noted that, for those skilled in the art, numerous improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning, characterized in that, The method includes the following steps: S1 is used to prepare and pre-treat adulterated standard samples. The adulterant is mixed with pure edible oil according to a preset gradient to simulate samples with different degrees of adulteration. The mixed adulterated standard samples are defoamed to obtain the pre-treated adulterated standard sample solution. S2 utilizes differential Raman spectroscopy to acquire high-quality spectral data. Using a portable differential Raman spectrometer, optimized dual-wavelength excitation, laser power, and integration time are set to perform differential Raman spectral scanning on adulterated samples and acquire high-quality spectral data. S3 performs spectral data processing. After acquiring high-quality spectral data, it processes the spectral data, including data cleaning to remove low wavenumber bands with low signal-to-noise ratio, using algorithms to perform baseline correction to eliminate drift, and performing Z-score standardization on the entire spectrum to offset the influence of environmental factors, so as to obtain clean and consistent spectral data for model analysis. S4 Characteristic peak extraction based on chemical knowledge: Based on the molecular structure characteristics of edible oils and adulterants, the position, intensity and area of ​​characteristic Raman peaks of edible oils and adulterants are identified and recorded, and a feature vector with physical meaning is constructed; Data-driven feature dimensionality reduction and extraction: Principal component analysis is performed on the spectral data processed in step S3, and the scores of the main principal components are extracted as features to achieve data dimensionality reduction and retain the main variation information in the spectrum. S5 constructs a multi-model machine learning framework, selects and trains corresponding models for different analysis objectives, inputs the feature information extracted in step S4 into the model for training, thereby forming an analysis system, and selects the optimal model based on evaluation indicators. The output of model S6, the optimal performance model obtained through rigorous evaluation in step S5, can quickly identify unknown edible oil samples, outputting a qualitative conclusion on whether they are adulterated and a semi-quantitative prediction of the adulteration concentration.

2. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 1, characterized in that, In step S1, the defoaming treatment of the mixed adulterated standard involves centrifuging and shaking the mixed adulterated standard and then letting it stand in the dark to eliminate air bubbles and ensure the uniformity and stability of the sample.

3. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 1, characterized in that, In step S2, the differential Raman spectroscopy is performed by automatically subtracting strong fluorescence background interference from the original Raman signal through dual-wavelength differential operation, thereby outputting a Raman spectrum with clear characteristic peaks and high signal-to-noise ratio. The reproducibility and uniformity of the data are verified by repeated measurements of the same and different parts.

4. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 1, characterized in that, In step S3, an algorithm is used to perform baseline correction to eliminate drift, including baseline correction through polynomial fitting to eliminate drift.

5. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 1, characterized in that, In step S4, the structural characteristics of edible oil and adulterants include the vibrational peaks of C=O and C=C bonds in edible oil, and the vibrational peaks of CH bonds in aromatic hydrocarbons in kerosene.

6. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 1, characterized in that, In step S5, for different analysis objectives, appropriate models are selected and trained, including using principal component analysis for dimensionality reduction and preliminary exploration; using random forest model for qualitative classification of whether there is adulteration and semi-quantitative regression prediction of the amount of adulteration; and supplementing with K-means clustering, artificial neural network and other models for comparative verification.

7. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 6, characterized in that, Depending on the different identification targets, the data is input into different machine learning models or combinations of models for analysis. For qualitative identification, to determine whether adulteration has occurred: the first choice is a random forest classifier for fast and accurate discrimination; at the same time, principal component analysis can be used to reduce dimensionality and visualize, intuitively verifying the separability of adulterated samples from pure oil samples. For semi-quantitative analysis, to estimate adulteration concentration, the first choice is a random forest regressor to establish a nonlinear mapping relationship from spectral features to adulteration concentration, thereby achieving accurate prediction of medium to high concentrations of adulteration. At the same time, artificial neural networks can be used as a comparative model to evaluate the performance of different algorithms. For unsupervised exploration, the inherent structure of data can be discovered by using methods such as K-means clustering to explore the clustering patterns of the data itself without relying on sample labels. This can be used to verify the conclusions of supervised models or to discover new data insights.

8. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 1, characterized in that, In step S5, the unknown edible oil sample is treated as follows before testing: similar to the adulterated standard in step S1, the oil sample to be tested is centrifuged and allowed to stand in the dark to eliminate air bubbles and ensure the uniformity and stability of the sample, thereby ensuring the consistency of its conditions in subsequent spectral acquisition and model analysis.

9. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 1, characterized in that, The model's output includes validating the trained model using test set data and outputting the final discrimination result: For qualitative models, the output is a judgment of whether the oil is pure or adulterated, and evaluation indicators such as accuracy and confusion matrix are given. For quantitative models, the predicted values ​​of doping concentrations are output, and evaluation indicators such as the coefficient of determination R² and root mean square error RMSE are given to clarify the predictive performance of the model in different concentration ranges.

10. The rapid identification method for adulterated edible oils based on differential Raman spectroscopy and multi-model machine learning according to claim 7, characterized in that, The random forest model is used to process feature information extracted from the spectrum. By constructing and integrating multiple decision trees, it can achieve qualitative classification of whether edible oil is adulterated and semi-quantitative regression prediction of adulteration concentration. The K-clustering algorithm is used to explore differential Raman spectral data in groups under unsupervised conditions to discover whether there are potential cluster structures related to doping in the spectral data; Artificial neural networks simulate nonlinear mapping relationships to perform deep learning and pattern recognition of spectral features, thereby assisting in classification or regression prediction tasks.