Method for rapidly detecting plasticizer in cottonseed oil by combining machine learning with Raman spectrum

By combining surface-enhanced Raman spectroscopy and machine learning models, the problems of long detection time and low accuracy of plasticizers in cottonseed oil have been solved, achieving rapid and accurate detection of plasticizers, which is suitable for oil refining process optimization and product safety control.

CN121595532APending Publication Date: 2026-03-03SHIHEZI UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511737373.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing methods for detecting plasticizers in cottonseed oil are time-consuming and inaccurate, failing to meet the requirements for rapid, accurate, non-destructive, and highly sensitive detection.

Method used

A surface-enhanced Raman spectroscopy (SERS) model was used to collect Raman spectral signals of cottonseed oil samples by preparing a gold nanoarray SERS substrate. After preprocessing, a machine learning model was constructed, and the plasticizer content was predicted using recursive feature elimination and coupled ridge regression models.

Benefits of technology

It enables rapid, accurate, and non-destructive detection of plasticizers in cottonseed oil, improving detection efficiency by tens of times, and significantly enhancing model prediction accuracy and stability. It is suitable for online quality monitoring in laboratories and production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121595532A_ABST
    Figure CN121595532A_ABST
Patent Text Reader

Abstract

The invention discloses a method for rapidly detecting a plasticizer in cottonseed oil by combining machine learning with Raman spectroscopy, and relates to the technical field of edible oil component detection. Comprising the steps of sample collection, substrate preparation, spectrum collection, completion of machine learning model building, data collection by adopting the method again, substrate preparation, spectrum collection and actual detection by applying the machine learning model. The invention provides the method for rapidly detecting the plasticizer in the cottonseed oil by combining machine learning with the Raman spectrum, which is rapid in detection and high in precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of edible oil component detection technology, specifically to a method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy. Background Technology

[0002] Phthalate esters (PAEs) are a class of important industrial organic compounds that are colorless, low in volatility, and poorly water-soluble. They are commonly used as plasticizers in plastic products, imparting flexibility and plasticity. However, they lack chemical bonds to polymers, making them prone to leaching from macro- and micro-plastics and migrating into processed products. Common phthalate esters include dibutyl phthalate (DBP), di(2-ethylhexyl)phthalate (DEHP), butyl benzyl phthalate (BBP), dimethyl phthalate (DMP), diethyl phthalate (DEP), and di-n-octyl phthalate (DOP), which have garnered significant public attention due to their chronic toxicity, such as teratogenicity, neurotoxicity, reproductive toxicity, and carcinogenicity at high concentrations.

[0003] Trace analysis of PAEs is crucial in oil processing. It allows for accurate identification and control of contamination, ensuring the safety and quality of edible oils. Common methods for detecting PAEs in edible oils include GC-MS, liquid chromatography-mass spectrometry (LC-MS), high-performance liquid chromatography (HPLC), ultra-high-performance liquid chromatography (UPLC), and gas chromatography-flame ionization detection (GC-FID). Among these, GC-MS, which measures the mass-to-charge ratio (m / z) of the analyte, is the most widely used technique, essential for accurate identification and quantification. However, GC-MS, LC-MS, HPLC, UPLC, and GC-FID all require complex sample pretreatment, reliance on sophisticated equipment, and several hours or days of dedicated testing time by skilled technicians. There is an urgent need in the oil industry chain for a rapid, accurate, non-destructive, easy-to-operate, highly sensitive, and stable analytical technique to detect whether the PAE content meets standards during oil production. Surface-enhanced Raman spectroscopy (SERS) is a powerful fingerprint spectral analysis technique with advantages including high sensitivity, strong selectivity, non-destructive nature, rapid detection, and chemical specificity. However, due to the high dimensionality of spectral data, it is necessary to combine machine learning methods to quickly and accurately analyze the data and identify key information in the spectral data.

[0004] Therefore, a new technical solution is needed to address the aforementioned technical problem, particularly a method for rapid detection of plasticizers in cottonseed oil based on surface-enhanced Raman spectroscopy combined with a machine learning model. Summary of the Invention

[0005] The purpose of this invention is to solve the problems of long detection time and inaccurate detection of plasticizers in cottonseed oil in traditional methods, and to provide a method for rapid detection of plasticizers in cottonseed oil based on surface-enhanced Raman spectroscopy combined with a machine learning model.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy, comprising the following steps:

[0007] S1. Collect cottonseed oil samples from different refining stages, and construct a sample library by taking samples from six process nodes: leaching, degumming, deacidification, decolorization, deodorization, and dewaxing.

[0008] S2. Prepare a surface-enhanced Raman spectroscopy substrate with a gold nanoarray structure, and after drying, use it to obtain the Raman spectral signals of samples in the sample library;

[0009] S3. Immerse the surface-enhanced Raman spectroscopy substrate in the cottonseed oil sample, and use an Invia micro-laser confocal Raman spectrometer to collect the surface-enhanced Raman spectrum of the surface-enhanced Raman spectroscopy substrate after immersion.

[0010] S4. The content of plasticizer in the cottonseed oil sample was determined by ultra-high performance liquid chromatography-mass spectrometry;

[0011] S5. The collected surface-enhanced Raman spectra are preprocessed by removing substrate noise interference, baseline correction and first derivative processing, distinguishing overlapping peaks and enhancing the differences of characteristic peaks to improve the quality of spectral data.

[0012] S6. Construct a machine learning model for plasticizer detection.

[0013] S6.1 The plasticizer content data detected by ultra-high performance liquid chromatography-mass spectrometry in S4 is converted into matrix Y; then the spectral data obtained in S5 is converted into matrix X, and a one-to-one mapping relationship is established to build a spectral database.

[0014] S6.2 The spectral database is divided into a test set, a training set, and a prediction set according to a ratio of 16%:80%:4%;

[0015] S6.3 Select the characteristic peaks of plasticizers corresponding to the spectral database, and establish a ridge regression model to build a functional relationship between the intensity of the characteristic peaks and the content of plasticizers;

[0016] S6.4 Construct a machine learning model that combines recursive feature elimination with cross-validation and coupled ridge regression; train the machine learning model using the training set data and evaluate its performance;

[0017] S6.5 The machine learning model is used to predict the plasticizer content of two unknown samples in the test set and prediction set, and compared with the plasticizer content data detected by the ultra-high performance liquid chromatography-mass spectrometry to evaluate the performance of the constructed machine learning model.

[0018] S7. The sample to be tested is processed by the methods in steps S1, S2, S3 and S5, and the machine learning model built in S6 is used to perform actual testing to obtain the plasticizer content value.

[0019] Preferably, step S2 includes the following steps:

[0020] S2.1 A monolayer colloidal crystal of polystyrene latex microspheres with uniform surface arrangement and close packing was prepared by self-assembly at the gas-liquid interface on a clean rectangular (2.54×7.62cm) silicon wafer.

[0021] S2.2 The polystyrene latex microsphere monolayer colloidal crystal is kept in an oven at 120°C for 60 seconds to make the polystyrene latex microspheres form planar contact with the silicon wafer;

[0022] S2.3 uses sulfur hexafluoride to perform plasma etching in a reactive ion etching machine. After etching for 30 seconds, a well-aligned array of silicon nanocones is formed on the silicon wafer.

[0023] S2.4 The residual polystyrene latex microspheres at the top of the silicon nanocone array are removed by calcination in air at 400°C;

[0024] S2.5 A thin gold layer is deposited on the silicon nanocone array using a sputtering instrument to obtain a surface-enhanced Raman spectroscopy substrate.

[0025] Preferably, step S3 includes the following steps:

[0026] S3.1 Immerse the surface-enhanced Raman spectroscopy substrate in anhydrous ethanol, remove it after 5 minutes and allow it to evaporate and dry naturally in a dry environment;

[0027] S3.2 The dried surface-enhanced Raman spectroscopy substrate is immersed in the cottonseed oil sample and soaked for 10 min in a dry environment;

[0028] S3.3 Subsequently, the surface-enhanced Raman spectroscopy substrate is placed on the observation platform of a microscope to complete the focusing of the laser and the collection of surface-enhanced Raman spectra;

[0029] S3.4 The collected surface-enhanced Raman spectra are preprocessed and stored using a Raman spectrometer.

[0030] Preferably, step S5 includes the following steps:

[0031] S5.1 The range of 500-2200 nm-1 in the surface-enhanced Raman spectrum was selected for analysis;

[0032] S5.2 performs the following operations in sequence: removing the base noise interference, limit correction, smoothing, first derivative processing, and standardization.

[0033] Preferably, S6.3 specifically includes: using ridge regression as the benchmark model for evaluating feature importance, calling the recursive feature elimination algorithm to perform preliminary feature screening on matrix X in the spectral database established in S6.1, obtaining feature coefficients β, then sorting by absolute value of coefficients, removing the feature with the smallest weight, repeating until 80 important features are finally retained, screening out the Raman shifts of the characteristic peaks of the vibrational modes of plasticizer molecules, and selecting the feature wavelengths with the most information content in matrix X; based on the recursive feature elimination algorithm, using the cross-validation algorithm to perform feature screening, calculating an importance score |ωi| for each feature, removing the feature with the smallest importance score |ωi| (step=1), and re-establishing a feature subset, repeating the above recursive process to obtain the next feature subset, until the number of remaining features reaches a preset minimum value of 1, and calculating the average performance score under cross-validation (CV=5) for each candidate feature subset Sk in the recursive process, the formula is as follows: The final number of features is selected using a dual feature selection mechanism that yields the highest cross-validation score. The scoring formula is as follows: Select the optimal number of features from the spectral data.

[0034] Preferably, S6.4 specifically includes: constructing a machine learning model framework combining recursive feature elimination and coupled ridge regression models with cross-validation; adding an L2 regularization term to the multiple linear regression, as shown in the formula... A set of selectable regularization parameters alphas = [0.01, 0.1, 1.0, 10.0, 100.0] is provided. The RidgeCV algorithm is used for cross-validation to determine the optimal α value. For each α value, k-fold cross-validation is performed, and the cross-validation results are stored. The regularization parameter α value that minimizes the mean square error is selected. The machine learning model is trained using the training set data, and the performance of the model is comprehensively evaluated by calculating indicators such as the coefficient of determination, root mean square error, and performance-to-bias ratio.

[0035] The beneficial effects of this invention are as follows:

[0036] The innovation of this patent lies in the construction of a complete, efficient, and high-precision detection system.

[0037] 1. Methodological innovation: This paper proposes for the first time to apply surface-enhanced Raman spectroscopy (SERS) technology to the tracking and detection of plasticizers in the entire refining process of cottonseed oil. Using a gold nanoarray SERS substrate, trace plasticizer signals in the complex matrix of oil are enhanced, and plasticizer signals are screened from high-dimensional data through machine learning algorithms, achieving a breakthrough from "complex and time-consuming sample pretreatment" to "rapid, in-situ, and highly sensitive" detection.

[0038] 2. Model Originality: A machine learning model combining recursive feature elimination and coupled ridge regression (RFECV-Ridge) with cross-validation is creatively constructed. This model can not only automatically and accurately screen out the feature peaks most relevant to plasticizer content from complex spectral data ("dimensionality reduction and noise reduction"), but also effectively solve the problem of strong collinearity in spectral data, and establish a mapping relationship between spectral signals and quantitative results of chromatography-mass spectrometry. The prediction accuracy and reliability surpass traditional linear models.

[0039] 3. System Integration: The system integrates sample library construction, SERS signal acquisition, spectral preprocessing, model building, and validation into a standardized, closed-loop process. This system is not only suitable for laboratory research but also has great potential for transformation into online quality monitoring on production lines, providing a powerful technical tool for optimizing oil refining processes and controlling product safety.

[0040] S1. Construct a sample library covering the entire refining process to provide a broad and representative sample library for subsequent model training. This enables the model to adapt to the complex matrix of oils with different refining levels, improves the model's universality and prediction stability, and avoids model overfitting or application limitations caused by a single sample.

[0041] S2. Prepare a SERS substrate with a gold nanoarray structure. The nanoarray structure can generate a strong and consistent surface plasmon resonance effect, which greatly enhances the Raman signal. This substrate has the advantages of strong signal enhancement, good reproducibility and high batch stability. It can provide high-quality and quantifiable spectral data for the model and overcome the defects of unstable enhancement effect of traditional colloidal nanoparticles.

[0042] S3. Surface-enhanced Raman spectra are acquired using the immersion method. Sample pretreatment is extremely simple, requiring no complicated extraction and purification steps, thus enabling rapid and non-destructive in-situ detection.

[0043] S4. The actual content of plasticizer was quantitatively determined using UPLC-MS to provide a reliable standard answer for the model.

[0044] S5. Perform spectral preprocessing including denoising, baseline correction, and first derivative, which can effectively separate overlapping peaks, highlight subtle differences in characteristic peaks, and eliminate interference from non-target signals, providing high-quality, highly discriminative data for subsequent machine learning models.

[0045] S6. Construct a machine learning model for plasticizer detection, combining recursive feature elimination and cross-validation to automatically and objectively select the most relevant variables, avoiding the subjectivity of manual selection, effectively preventing overfitting, and enhancing the model's generalization ability. Coupled ridge regression addresses the problem of numerous spectral variables and strong collinearity, improving the model's stability and anti-interference ability. Train the model to achieve accurate and stable mapping from complex spectra to content values.

[0046] S7. The sample to be tested is processed by the methods in steps S1, S2, S3 and S5, and the machine learning model built in S6 is used to perform actual testing and obtain the plasticizer content value. All the aforementioned optimization steps are integrated into a systematic testing process, which enables non-professionals to complete the test in a few minutes. The testing efficiency is increased by more than 10 times compared with traditional methods, and the cost is low. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the operation method of the present invention. Figure 1 .

[0048] Figure 2 This is a flowchart illustrating the operation method of the present invention. Figure 2 .

[0049] Figure 3 This is a diagram showing the effect of the surface-enhanced Raman spectroscopy instrument condition adjustment part of the present invention.

[0050] Figure 4 The images show surface-enhanced Raman spectra of different concentrations of plasticizers detected in this invention.

[0051] Figure 5 This is a surface-enhanced Raman spectrum for machine learning in this invention.

[0052] Figure 6 This is a visualization of machine learning predictions from the present invention. Detailed Implementation

[0053] To make the above-mentioned objectives, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to specific examples.

[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0055] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0056] like Figures 1 to 6 As shown, a method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy includes the following steps:

[0057] S1. Collect cottonseed oil samples from different refining stages, and construct a sample library by taking samples from six process nodes: leaching, degumming, deacidification, decolorization, deodorization, and dewaxing.

[0058] To construct a widely representative sample library, we can provide comprehensive data on the dynamic changes of plasticizer content in edible oils (including cottonseed oil) throughout the refining process. This will avoid problems such as model overfitting or narrow applicability caused by a single sample. By detecting the plasticizer content in oils at different stages, we can identify the process steps in cottonseed oil processing that produce plasticizers, and then optimize the process steps that produce them.

[0059] S2. Prepare a surface-enhanced Raman spectroscopy substrate with a gold nanoarray structure, and after drying, use it to obtain the Raman spectral signals of samples in the sample library.

[0060] Step S2 includes the following steps:

[0061] S2.1 A monolayer colloidal crystal of polystyrene latex microspheres with uniform surface arrangement and close packing was prepared by self-assembly at the gas-liquid interface on a clean rectangular (2.54×7.62cm) silicon wafer.

[0062] S2.2 The polystyrene latex microsphere monolayer colloidal crystal is kept in an oven at 120°C for 60 seconds to make the polystyrene latex microspheres form planar contact with the silicon wafer;

[0063] S2.3 uses sulfur hexafluoride to perform plasma etching in a reactive ion etching machine. After etching for 30 seconds, a well-aligned array of silicon nanocones is formed on the silicon wafer.

[0064] S2.4 The residual polystyrene latex microspheres at the top of the silicon nanocone array are removed by calcination in air at 400°C;

[0065] S2.5 A thin gold layer is deposited on the silicon nanocone array using a sputtering instrument to obtain a surface-enhanced Raman spectroscopy substrate.

[0066] S3. Immerse the surface-enhanced Raman spectroscopy substrate in the cottonseed oil sample, and collect the surface-enhanced Raman spectrum of the immersed surface-enhanced Raman spectroscopy substrate using an Invia micro-laser confocal Raman spectrometer.

[0067] Step S3 includes the following steps:

[0068] S3.1 Immerse the surface-enhanced Raman spectroscopy substrate in anhydrous ethanol, remove it after 5 minutes and allow it to evaporate and dry naturally in a dry environment;

[0069] S3.2 The dried surface-enhanced Raman spectroscopy substrate is immersed in the cottonseed oil sample and soaked for 10 min in a dry environment;

[0070] S3.3 Subsequently, the surface-enhanced Raman spectroscopy substrate is placed on the observation platform of a microscope to complete the focusing of the laser and the collection of surface-enhanced Raman spectra;

[0071] S3.4 The collected surface-enhanced Raman spectra are preprocessed and stored using a Raman spectrometer.

[0072] S4. The content of plasticizer in the cottonseed oil sample was determined by ultra-high performance liquid chromatography-mass spectrometry;

[0073] S5. The collected surface-enhanced Raman spectra are preprocessed by removing substrate noise interference, baseline correction, and first derivative processing to distinguish overlapping peaks and enhance the differences in characteristic peaks, thereby improving the quality of spectral data.

[0074] Step S5 includes the following steps:

[0075] S5.1 The range of 500-2200 nm-1 in the surface-enhanced Raman spectrum was selected for analysis;

[0076] S5.2 performs the following operations in sequence: removing the base noise interference, limit correction, smoothing, first derivative processing, and standardization.

[0077] S6. Construct a machine learning model for plasticizer detection.

[0078] S6.1 converts the plasticizer content data detected by ultra-high performance liquid chromatography-mass spectrometry in S4 into matrix Y; then converts the surface-enhanced Raman spectrum obtained in S5 into matrix X. The pd.merge and merged_df functions are called to establish a one-to-one mapping between the row-level data of X and Y, thus building a spectral database.

[0079] S6.2 The spectral database is divided into a test set, a training set, and a prediction set in a ratio of 16%:80%:4%. All datasets are standardized, with the training set (80%) used for model training and hyperparameter tuning, the test set used for final performance evaluation, and the prediction set used for testing in simulated real-world environments.

[0080] S6.3 Select the characteristic peaks of plasticizers corresponding to the spectral database, and establish a ridge regression model to build a functional relationship between the intensity of the characteristic peaks and the content of plasticizers.

[0081] Ridge regression was used as the benchmark model for evaluating feature importance. A recursive feature elimination algorithm was invoked to perform preliminary feature screening on matrix X in the spectral database established in S6.1, obtaining feature coefficients β. These coefficients were then sorted by absolute value, and features with the lowest weight were removed. This process was repeated until 80 important features were retained, resulting in the Raman shifts of the characteristic peaks corresponding to the vibrational modes of plasticizer molecules: 653 cm⁻¹, 1043 cm⁻¹, 1156 cm⁻¹, 1167 cm⁻¹, 1284 cm⁻¹, 1292 cm⁻¹, 1585 cm⁻¹, 1605 cm⁻¹, 1764 cm⁻¹, 1796 cm⁻¹, etc. The most informative feature wavelengths are selected from matrix X. The optimal number of features is determined based on cross-validation recursive feature elimination. The key feature is that, based on the recursive feature elimination algorithm, a cross-validation algorithm is used for feature selection. For each feature, an importance score |ωi| is calculated. The feature with the smallest importance score |ωi| is removed (step = 1), and a new feature subset is established. This recursive process is repeated to obtain the next feature subset until the number of remaining features reaches a preset minimum of 1. For each candidate feature subset Sk in the recursive process, the average performance score under cross-validation (CV = 5) is calculated using the following formula:

[0082] Where CV(S) k ) represents the average cross-validation score of the model when using the k-th hyperparameter configuration; To find the average coefficient, where N cv The number of folds for cross-validation; For the summation formula, i starts from 1 and is added up to the Ncv term; For the term being summed, where Score is the evaluation function, For the model trained using the k-th set of hyperparameters Sk, This is the dataset reserved for testing in the i-th fold cross-validation.

[0083] The final number of features is selected using a dual feature selection mechanism that yields the highest cross-validation score. The scoring formula is as follows: (Where k* is the optimal value, argmax(k) is the input value k that makes the function CV(Sk) reach its maximum value, and CV(Sk) is the cross-validation score of the corresponding model when the parameter is k). The optimal number of features is selected from the spectral data. The feature importance is visualized and qualitatively verified. The feature weights of the selected optimal feature subset are calculated and visualized to intuitively show the contribution distribution of each feature wavelength to the identification of plasticizers. The selected features are visualized through feature ranking calculation to show the feature weight distribution. The rationality of feature selection is verified by comparing with the Raman spectral characteristic peaks of plasticizers.

[0084] S6.4 Construct a machine learning model combining recursive feature elimination with cross-validation and coupled ridge regression; train the machine learning model using the training set data and evaluate its performance. Add an L2 regularization term to the multiple linear regression, as shown in the formula:

[0085] in β is the model parameter vector that minimizes the loss function; Let Y be the sum of squared residuals, where Y is the true target value vector, X is the feature matrix, and Xβ is the residual; ||||2 is the L2 norm of the vector. α is the L2 regularization term, where α is the regularization parameter alphas and β is the weight coefficient of the model.

[0086] Given another set of selectable regularization parameters alphas = [0.01, 0.1, 1.0, 10.0, 100.0], the RidgeCV algorithm is used for cross-validation to determine the optimal α value. For each α value, k-fold cross-validation is performed, and the cross-validation results are stored. The regularization parameter α value that minimizes the mean square error is selected. The machine learning model is trained using the training set data, and the performance of the model is comprehensively evaluated by calculating indicators such as the coefficient of determination, root mean square error, and performance-to-bias ratio.

[0087] S6.5 The machine learning model is used to predict the plasticizer content of two unknown samples in the test set and prediction set, and compared with the plasticizer content data detected by the ultra-high performance liquid chromatography-mass spectrometry to evaluate the performance of the constructed machine learning model.

[0088] like Figures 1 to 6 As shown, and in conjunction with specific machine learning model experiments, the present invention provides the following embodiments:

[0089] like Figures 3 to 5 As shown, the Raman spectroscopy dataset constructed in this invention, combined with the RFEVC-Ridge machine learning quantitative prediction model, was used to collect surface-enhanced Raman spectra at the cottonseed oil processing plant of Xinjiang Qianhai Oils & Fats Technology Co., Ltd. The data included six refining stages: leaching, deacidification, degumming, decolorization, deodorization, and dewaxing. Three samples were taken at each refining stage, and each sample was measured ten times, resulting in a total of 180 surface-enhanced Raman data points. The parameters of the Invia micro-laser confocal Raman spectrometer were adjusted, and surface-enhanced Raman spectroscopy measurements were performed on different concentrations of plasticizer (DBP) to ensure the substrate and Raman spectrometer had good and stable functionality. The following steps were then performed:

[0090] Step 1: Use an Invia micro-laser confocal Raman spectrometer to collect surface-enhanced Raman spectral data of cottonseed oil with different refining levels, analyze the peak values, peak positions and other characteristic information in the data, and establish a Raman spectral dataset of different concentrations of plasticizers.

[0091] The preliminary experiment for quantitative prediction of plasticizers in this example included seven concentrations of plasticizer (DBP) solutions: 10, 20, 50, 100, 200, 500, and 1000 ng / mL.

[0092] Step 2: In order to improve the training efficiency and performance of machine learning models, this invention performs preprocessing on surface-enhanced Raman spectroscopy data, including noise removal, baseline correction, smoothing, and first-order derivative processing, and uses sklearn.preprocessing for standardization.

[0093] Step 3: The collected cottonseed oil samples were analyzed using ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS / MS) to measure the plasticizer content. The specific steps are as follows:

[0094] (1) Preparation of standard solutions: Prepare a series of standard control solutions with different concentrations of 0.1, 0.5, 1, 2, 5, 8, 10, 20, 50, 100, 200, 500, 800, 1000, and 2000 ng / mL for the plasticizer standard substance.

[0095] (2) Weigh 1g of biological sample, add 5mL of pure acetonitrile, and shake thoroughly to mix;

[0096] (3) Extract by ultrasonication at room temperature for 10 min, then place in a -20℃ refrigerator for 30 min;

[0097] (4) Centrifuge at 16000g, 4℃ for 10min, and collect the supernatant;

[0098] (5) Filter the sample with a microporous membrane (0.22 μm pore size) and store it in a sample vial for UPLC-MS / MS analysis.

[0099] Step 4: The surface-enhanced Raman spectroscopy data analysis model for quantitatively measuring the concentration of plasticizers in cottonseed oil mainly adopted the RFECV and Ridge regression algorithms. The specific steps are as follows:

[0100] (1) Divide the data into a training set, a test set, and a prediction set;

[0101] (2) Start training from the original feature set and calculate an importance score |ωi| for each feature. Remove the least important feature (step=1). This process will be repeated until the number of remaining features reaches the preset minimum value.

[0102] (3) For each candidate feature subset Sk (containing k features) in the recursive process, calculate the average performance score under cross-validation (CV=5);

[0103] (4) Select the number of features that corresponds to the highest cross-validation score as the optimal number of features;

[0104] (5) Calculate the optimal hyperparameter (α) by calling the RidgeCV algorithm;

[0105] (6) Establishing a Ridge Regression Model Perform model training;

[0106] Where J(β) is the total error of the model under the current parameter β; The mean squared error is multiplied by 1 / 2, where n is the sample size and y is the mean squared error. i Let i be the true values ​​of the i samples. Let be the predicted value for the i-th sample; Let be the ridge penalty term, where α is the regularization parameter alphas, and β is the regularization parameter. j Let be the weight coefficient of the j-th feature. It is the sum of squares of p parameters.

[0107] (7) Determine the model hyperparameters;

[0108] (8) Use machine learning models to predict the data and obtain the predicted value of plasticizer content in cottonseed oil;

[0109] Step 5: Plot the relationship between predicted and measured values, and evaluate the model's predictive performance on the training, test, and prediction sets using the coefficient of determination (R²), root mean square error (RMSE), and performance-to-bias ratio (RPD) of the calibration and prediction sets obtained from model testing.

[0110] like Figure 6 As shown, the Raman spectroscopy combined with the RFEVC-Ridge machine learning quantitative prediction model constructed in this invention has good performance.

[0111] This invention discovers that the optimized machine learning model exhibits superior prediction accuracy, generalization ability, and robustness compared to the traditional multiple linear regression model.

[0112]

[0113] On the training set, RFECV-Ridge's RMSE (1.4058) is significantly lower than MLR (4.1211), while its R² (0.9997) is higher than MLR (0.9974). This indicates that RFECV-Ridge has a stronger ability to learn and fit the training data. The test set is crucial for evaluating generalization ability. RFECV-Ridge's RMSE (2.6228) is significantly lower than MLR (5.4146), while its R² is higher. 2 (0.9989) is higher. This indicates that RFECV-Ridge has a much higher prediction accuracy than MLR when faced with new, unseen data. The prediction set represents completely independent external samples used to simulate real-world application scenarios. RFECV-Ridge also performs excellently (RMSE = 2.1132, R 2 =0.9991), and its error is much smaller than that of MLR (RMSE = 4.8083). This strongly demonstrates that RFECV-Ridge has higher practical application value. The RPD of the MLR model on the training set (19.62) is high, but it drops significantly on the test set (14.49) and prediction set (14.85). This performance degradation is a typical sign of slight overfitting, that is, the model fits too closely to the noise in the training data, resulting in poor performance on new data. The RPD value of the RFECV-Ridge model is very high and stable: training set (47.53), test set (29.91), prediction set (33.78). All values ​​are far above the 5.0 excellent line, which indicates that the model has extremely strong stability and reliability, and strong and stable predictive ability for new samples.

[0114] This invention provides a method for sampling cottonseed oil at different processing stages and then detecting plasticizers in it. This method involves sampling different samples, generating a machine learning model, training and testing the model, and adjusting the test error results using statistical methods to ultimately improve its accuracy.

[0115] The RMSE is used to measure the average deviation between the model's predicted values ​​and the actual values. Its advantage is that it can intuitively quantify the model's prediction error, quickly determine whether the prediction accuracy meets the standard when training the model's hyperparameters, and reduce the large prediction deviation caused by multicollinearity in ridge regression.

[0116] Using R 2 The advantage of ridge regression in measuring model fit and explanatory power is that it directly quantifies the proportion of the dependent variable variation explained by the model. During training, it can clearly define the goodness of fit of the model and avoid ineffective modeling that only reduces error without improving explanatory power.

[0117] Using RPD to measure the practical predictive ability of a model has the advantage of objectively evaluating model performance, using a dimensionless proportional relationship to measure predictive reliability, and having a clear threshold standard to quickly determine whether ridge regression meets practical application requirements. It also allows for direct sampling and detection during subsequent measurements, resulting in short detection time and high accuracy.

[0118] The above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the present invention, and the patent protection scope of the present invention should be defined by the claims.

Claims

1. A method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy, characterized in that, Includes the following steps: S1. Collect cottonseed oil samples from different refining stages, and construct a sample library by taking samples from six process nodes: leaching, degumming, deacidification, decolorization, deodorization, and dewaxing. S2. Prepare a surface-enhanced Raman spectroscopy substrate with a gold nanoarray structure, and after drying, use it to obtain the Raman spectral signals of samples in the sample library; S3. Immerse the surface-enhanced Raman spectroscopy substrate in the cottonseed oil sample, and use an Invia micro-laser confocal Raman spectrometer to collect the surface-enhanced Raman spectrum of the surface-enhanced Raman spectroscopy substrate after immersion. S4. The content of plasticizer in the cottonseed oil sample was determined by ultra-high performance liquid chromatography-mass spectrometry; S5. The collected surface-enhanced Raman spectra are preprocessed by removing substrate noise interference, baseline correction and first derivative processing, distinguishing overlapping peaks and enhancing the differences of characteristic peaks to improve the quality of spectral data. S6. Construct a machine learning model for plasticizer detection. S6.1 The plasticizer content data detected by ultra-high performance liquid chromatography-mass spectrometry in S4 is converted into matrix Y; then the surface-enhanced Raman spectrum data preprocessed in S5 is converted into matrix X, and a one-to-one mapping relationship is established to build a spectrum database. S6.2 The spectral database is divided into a test set, a training set, and a prediction set according to a ratio of 16%:80%:4%; S6.3 Select the characteristic peaks of plasticizers corresponding to the spectral database, and establish a ridge regression model to build a functional relationship between the intensity of the characteristic peaks and the content of plasticizers; S6.4 Construct a machine learning model that combines recursive feature elimination with cross-validation and coupled ridge regression; train the machine learning model using the training set data and evaluate its performance; S6.5 The machine learning model is used to predict the plasticizer content of two unknown samples in the test set and prediction set, and compared with the plasticizer content data detected by the ultra-high performance liquid chromatography-mass spectrometry to evaluate the performance of the constructed machine learning model. S7. The sample to be tested is processed by the methods in steps S1, S2, S3 and S5, and the machine learning model built in S6 is used to perform actual testing to obtain the plasticizer content value.

2. The method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy as described in claim 1, characterized in that, Step S2 includes the following steps: S2.1 A monolayer colloidal crystal of polystyrene latex microspheres with uniform surface arrangement and close packing was prepared by self-assembly at the gas-liquid interface on a clean rectangular (2.54×7.62cm) silicon wafer. S2.2 The polystyrene latex microsphere monolayer colloidal crystal is kept in an oven at 120°C for 60 seconds to make the polystyrene latex microspheres form planar contact with the silicon wafer; S2.3 uses sulfur hexafluoride to perform plasma etching in a reactive ion etching machine. After etching for 30 seconds, a well-aligned array of silicon nanocones is formed on the silicon wafer. S2.4 The residual polystyrene latex microspheres at the top of the silicon nanocone array are removed by calcination in air at 400°C; S2.5 A thin gold layer is deposited on the silicon nanocone array using a sputtering instrument to obtain a surface-enhanced Raman spectroscopy substrate.

3. The method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy as described in claim 1, characterized in that, Step S3 includes the following steps: S3.1 Immerse the surface-enhanced Raman spectroscopy substrate in anhydrous ethanol, remove it after 5 minutes and allow it to evaporate and dry naturally in a dry environment; S3.2 The dried surface-enhanced Raman spectroscopy substrate is immersed in the cottonseed oil sample and soaked for 10 min in a dry environment; S3.3 Subsequently, the surface-enhanced Raman spectroscopy substrate is placed on the observation platform of a microscope to complete the focusing of the laser and the collection of surface-enhanced Raman spectra; S3.4 The collected surface-enhanced Raman spectra are preprocessed and stored using a Raman spectrometer.

4. The method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy as described in claim 1, characterized in that, Step S5 includes the following steps: S5.1 The range of 500-2200 nm-1 in the surface-enhanced Raman spectrum was selected for analysis; S5.2 performs the following operations in sequence: removing the base noise interference, limit correction, smoothing, first derivative processing, and standardization.

5. The method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy as described in claim 1, characterized in that: S6.3 specifically includes: using ridge regression as the benchmark model for evaluating feature importance, calling the recursive feature elimination algorithm to perform preliminary feature screening on matrix X in the spectral database established in S6.1, obtaining feature coefficients β, then sorting by the absolute value of the coefficients, removing the feature with the smallest weight, repeating until 80 important features are finally retained, screening out the Raman shifts of the characteristic peaks of the vibrational modes of plasticizer molecules, and selecting the feature wavelengths with the most information content in matrix X; based on the recursive feature elimination algorithm, using the cross-validation algorithm to perform feature screening, calculating an importance score |ωi| for each feature, removing the feature with the smallest importance score |ωi| (step=1), and re-establishing a feature subset, repeating the above recursive process to obtain the next feature subset, until the number of remaining features reaches the preset minimum value of 1, and calculating the average performance score under cross-validation (CV=5) for each candidate feature subset Sk in the recursive process, the formula is as follows. The final number of features is selected using a dual feature selection mechanism that yields the highest cross-validation score. The scoring formula is as follows: Select the optimal number of features from the spectral data.

6. The method for rapid detection of plasticizers in cottonseed oil using machine learning combined with Raman spectroscopy as described in claim 1, characterized in that: S6.4 specifically includes: constructing a machine learning model framework that combines recursive feature elimination with coupled ridge regression model and cross-validation; adding an L2 regularization term to multiple linear regression, as shown in the formula... A set of selectable regularization parameters alphas = [0.01, 0.1, 1.0, 10.0, 100.0] is provided. The RidgeCV algorithm is used for cross-validation to determine the optimal α value. For each α value, k-fold cross-validation is performed, and the cross-validation results are stored. The regularization parameter α value that minimizes the mean square error is selected. The machine learning model is trained using the training set data, and the performance of the model is comprehensively evaluated by calculating indicators such as the coefficient of determination, root mean square error, and performance-to-bias ratio.