Method for rapidly detecting Sudan red in chili powder based on nanoparticles and machine learning
By combining magnetic molecularly imprinted polymers and a random forest model, the problems of complex preprocessing and insufficient accuracy in Sudan Red detection are solved, enabling rapid and accurate on-site detection.
Patent Information
- Application Number
- CN202511772810.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-27
AI Technical Summary
Existing Sudan Red detection technologies suffer from problems such as complex and time-consuming sample pretreatment, insufficient accuracy, and high cost, making it difficult to meet the needs of rapid on-site screening.
Sudan I was enriched using magnetic molecularly imprinted polymers, and its Fourier transform infrared spectrum was obtained by magnetic separation. Quantitative analysis was performed using a random forest model to construct dual recognition sites for Sudan I and capsaicin, reducing non-specific occupation and cross-reaction of interfering substances and improving selectivity and accuracy.
It reduces sample pretreatment time to within 10 minutes, improves detection efficiency, and reduces false positive and false negative rates. It is suitable for large-scale on-site screening and has high accuracy and low cost.
Smart Images

Figure CN121577568A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of food detection, in particular to a rapid detection method for Sudan red in chili powder based on nanoparticles and machine learning. BACKGROUND
[0002] The present application relates to the field of food safety detection technology, specifically, a rapid quantitative detection method for illegally added dyes in food matrix. Sudan I, as a synthetic azo dye, has been banned in most countries and regions around the world due to its potential carcinogenicity. However, it is still detected in condiments such as chili powder in order to enhance the color of the product. Therefore, it is of great significance to establish a rapid, accurate and suitable on-site screening method for Sudan I detection to protect public health and market supervision.
[0003] Currently, the detection techniques for Sudan I can be mainly divided into two categories: laboratory standard methods and on-site rapid detection techniques. The laboratory standard methods, represented by high performance liquid chromatography (HPLC) and its combined techniques, have high accuracy and sensitivity, and are the basis for legal confirmation. In addition, various rapid detection techniques have been developed, mainly including immunochromatography based on antigen-antibody specific binding, spectroscopic methods using surface enhanced Raman spectroscopy (SERS), and sensor technology based on electrochemical reaction, etc. These techniques meet the detection needs in different scenarios to some extent.
[0004] However, the existing Sudan red detection techniques have the following technical defects, which directly lead to low detection efficiency, high cost or insufficient accuracy. The pretreatment is complex and time-consuming, traditional solid phase extraction (SPE) relies on centrifugal / filtration separation, and the operation steps are complicated, resulting in a single sample processing time of more than 30 minutes, which cannot meet the demand of on-site rapid screening. The selectivity is insufficient, the antibody or C18 adsorbent cross-adsorbs Sudan red structural analogues (such as Sudan red III / IV), resulting in a high false positive rate (more than 15% in HPLC verification), which requires secondary confirmation and increases the cost. The instrument is highly dependent, HPLC / LC-MS requires laboratory environment and professional operation, which is difficult for grassroots market supervision to popularize, resulting in a detection period of 24-48 hours. Spectral interference is serious, the chili powder matrix (such as capsaicin, carotenoids) overlaps with Sudan red FTIR peak, resulting in an accuracy rate of less than 85% for traditional PLS algorithm, and a high miss rate for low concentration samples (less than 1mg / kg). SUMMARY
[0005] In view of the deficiencies of the prior art, the present application provides a rapid detection method for Sudan red in chili powder based on nanoparticles and machine learning, which solves the comprehensive technical problems of low detection efficiency caused by complex and time-consuming sample pretreatment process, and poor accuracy and reproducibility of quantitative analysis results caused by insufficient selectivity and anti-interference ability.
[0006] To achieve the above object, the first aspect of the present application provides a method for rapid detection of Sudan red in paprika powder based on nanoparticles and machine learning, comprising the following steps:
[0007] Step a, providing a to-be-tested extract obtained from a paprika powder sample;
[0008] Step b, using a magnetic molecularly imprinted polymer to enrich Sudan I in the to-be-tested extract, and obtaining the magnetic molecularly imprinted polymer loaded with Sudan I through magnetic separation;
[0009] Step c, obtaining a Fourier transform infrared spectrum of the magnetic molecularly imprinted polymer loaded with Sudan I;
[0010] Step d, extracting a multi-dimensional feature containing characteristic peak information of Sudan I and characteristic peak information of a magnetic core in the magnetic molecularly imprinted polymer from the Fourier transform infrared spectrum;
[0011] Step e, inputting the multi-dimensional feature into a pre-trained machine learning regression model, and outputting the concentration of Sudan I in the to-be-tested extract from the model.
[0012] The second aspect of the present application provides a magnetic molecularly imprinted polymer for selectively enriching Sudan I, which comprises a magnetic core and a molecularly imprinted polymer layer coated on the surface of the magnetic core, and the molecularly imprinted polymer layer contains double recognition sites for Sudan I as a target analyte and capsaicin as a representative interferent.
[0013] As a further optimization of the technical solution of the first aspect of the present application:
[0014] In a preferred embodiment, the magnetic molecularly imprinted polymer in step b contains double recognition sites for the target analyte Sudan I and the representative interferent capsaicin in the molecularly imprinted polymer layer. By pre-constructing specific recognition sites for common strong interferents in the matrix, they are preferentially captured during the enrichment process, thereby reducing non-specific occupation of the Sudan I recognition site and cross-reaction, improving the selectivity and purity of Sudan I enrichment, and providing data for subsequent high-precision spectral quantitative analysis.
[0015] In a preferred embodiment, the multi-dimensional feature extraction in step d specifically includes calculating the characteristic peak area of Sudan I and the magnetic core as an internal standard in the magnetic molecularly imprinted polymer (for example the ratio of the characteristic peak area of Sudan I to the characteristic peak area of the Fe3O4 magnetic core. By constructing such a proportional feature, the overall intensity fluctuations of the spectrum caused by sample preparation (such as uneven sample weight during KBr tablet pressing) and system errors during spectrum acquisition can be effectively calibrated and eliminated, thereby greatly improving the accuracy of the quantitative analysis method.
[0016] Further, the calculation of the peak area ratio can specifically be the calculation of the ratio of the integral area of the characteristic peak of Sudan I in one or more predetermined wave number ranges (for example, 1600-1610 , 1545-1555 , 1345-1355 ) to the integral area of the characteristic peak of the Fe3O4 magnetic core in one or more predetermined wave number ranges (for example, 580-600 and 400-450 ).
[0017] In a more specific embodiment, the multi-dimensional feature in step d constitutes a feature vector, which, in addition to including the aforementioned ratio of the characteristic peak area of Sudan I to the characteristic peak area of the magnetic core, can also include: the characteristic peak area of the magnetic core, the characteristic peak area of Sudan I, the characteristic peak height of the magnetic core, and the characteristic peak height of Sudan I. By constructing a multi-dimensional feature combination containing peak area, peak height, and peak area ratio, the spectrum information can be more comprehensively represented, and the machine learning model can be provided with more abundant discriminant basis.
[0018] In a preferred embodiment, the pre-trained machine learning regression model in step e is a random forest model. Through systematic comparison of various regression models, it is found that the tree-based ensemble learning model (such as random forest) can most effectively capture the complex nonlinear relationship between the multi-dimensional features and the concentration of Sudan I, and its prediction performance is significantly better than that of traditional linear models.
[0019] Further, the training process of the random forest model is based on a standard spectrum data set specially constructed to improve the anti-interference ability of the model. The construction of this data set includes: preparing a series of gradient standard samples containing different concentration gradients of Sudan I and different concentration gradients of capsaicin, and obtaining the corresponding Fourier transform infrared spectrum. By training with such a two-dimensional gradient standard data set, the model can learn to recognize the target substance in different interferent backgrounds, thereby having excellent generalization ability and stability in real sample detection.
[0020] In a preferred embodiment, before step d of the present application, a step of pre-processing the Fourier transform infrared spectrum obtained in step c is further included. The pre-processing is used to eliminate noise and baseline drift in the spectrum. To ensure the pre-processing is optimized and subjective, the parameters of the pre-processing (e.g. window size and polynomial order of Savitzky-Golay smoothing) are determined by an automated optimization process. The process aims to select the combination of parameters that can minimize the root mean square error (RMSE) of the subsequent quantitative model in cross-validation.
[0021]
[0022] wherein: is the root mean square error; is the total number of samples used for cross-validation; is the true concentration value of the i-th sample; is the predicted concentration value of the i-th sample. In a specific embodiment, the enrichment step of step b specifically comprises: adding the magnetic molecularly imprinted polymer to the sample extraction solution for oscillation adsorption, and then performing rapid magnetic separation by applying a magnetic field. In a specific embodiment, the step of obtaining the Fourier transform infrared spectrum of step c comprises: mixing the magnetic molecularly imprinted polymer loaded with Sudan I obtained in step b with potassium bromide (KBr) powder uniformly and pressing into a thin sheet, and then performing spectrum acquisition.
[0023] The present application provides a rapid detection method for Sudan red in chili powder based on nanoparticles and machine learning. The present application has the following beneficial effects:
[0024] 1. On the one hand, the present application constructs double recognition sites for Sudan I and representative interferents (such as capsaicin) in the magnetic molecularly imprinted polymer, actively captures and isolates the main interferents in the matrix, and ensures the purity of the enrichment of Sudan I; on the other hand, the machine learning model is trained on a data set containing different concentration gradients of Sudan I and capsaicin, so that it can accurately identify the characteristics of the target in a complex interference background, solve the problem of low accuracy and high false negative rate caused by overlapping matrix peaks in traditional spectral analysis methods, and simplify the detection steps.
[0025] 1. On the one hand, the present application constructs double recognition sites for Sudan I and representative interferents (such as capsaicin) in the magnetic molecularly imprinted polymer, actively captures and isolates the main interferents in the matrix, and ensures the purity of the enrichment of Sudan I; on the other hand, the machine learning model is trained on a data set containing different concentration gradients of Sudan I and capsaicin, so that it can accurately identify the characteristics of the target in a complex interference background, solve the problem of low accuracy and high false negative rate caused by overlapping matrix peaks in traditional spectral analysis methods, and simplify the detection steps.
[0026] 2、The application adopts magnetic molecularly imprinted polymers to selectively enrich target objects, and combines an external magnetic field to rapidly separate, so that complicated steps such as centrifugation and filtration are not needed, sample pretreatment time is shortened to within 10 minutes, rapid response from sample to result is realized, the problem of complicated pretreatment steps and too long time in the prior art is solved, and the method is especially suitable for the demand of on-site large-scale screening, and thus detection efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 It is a method flowchart of the application;
[0028] Figure 2 It is a comparison chart of random forest model predicted values and true values of the application;
[0029] Figure 3 It is a Fourier transform infrared spectrum of the magnetic molecularly imprinted polymer loaded with Sudan I. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the specification of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0031] In order to better understand the application, the above content will be described in detail below in combination with specific embodiments.
[0032] Please refer to the drawings in the specification of the application Figure 1 - the drawings in the specification of the application Figure 3 The embodiments of the application provide a rapid detection method of Sudan in chili powder based on nanoparticles and machine learning, which comprises the following steps:
[0033] Step a, providing a to-be-tested extract solution obtained from a chili powder sample;
[0034] Step b, using a magnetic molecularly imprinted polymer to enrich Sudan I in the to-be-tested extract solution, and obtaining the magnetic molecularly imprinted polymer loaded with Sudan I through magnetic separation;
[0035] Step c, acquiring a Fourier transform infrared spectrum of the magnetic molecularly imprinted polymer loaded with Sudan I;
[0036] Step d, extracting multi-dimensional features containing characteristic peak information of Sudan I and characteristic peak information of a magnetic core in the magnetic molecularly imprinted polymer from the Fourier transform infrared spectrum;
[0037] Step e: Input the multidimensional features into a pre-trained machine learning regression model, and have the model output the concentration of Sudan I in the extract to be tested.
[0038] In one specific embodiment, the present invention first synthesizes imprinted polymers (MIPs) possessing dual recognition sites for Sudan I and capsaicin, namely... The specific steps include:
[0039] (1) Synthesis of magnetic nanonuclei: A chemical coprecipitation method was employed. Under a nitrogen atmosphere, 9.4 g of... With 3.5g Dissolve in 140 ml of deionized water. Heat the mixture to 80°C while stirring vigorously, then add 20 ml of concentrated ammonia dropwise. The reaction was initiated. The resulting black precipitate was collected using an external magnetic field and washed sequentially with deionized water and ethanol. Finally, it was dried under vacuum at 60°C for 24 hours to obtain... Nanoparticles. For For the preparation of nanoparticles, those skilled in the art can use other well-known synthesis methods, and the specific synthesis processes are well-known technologies in the field, and will not be described in detail here.
[0040] (2) Synthesis of intermediate: 100 mg Nanoparticles were ultrasonically dispersed in a 3:1 mixture of ethanol and water in 160 ml of water. 3 ml of concentrated ammonia and 1 ml of tetraethyl orthosilicate (TEOS) were added to the suspension, and the mixture was stirred continuously at room temperature for 8 hours. The reaction product was magnetically separated, washed, and dried at 80 °C to obtain the nanoparticles used as a magnetic carrier. Core-shell structured nanoparticles;
[0041] (3) After obtaining the magnetic carrier, a molecularly imprinted layer is constructed on the surface of the carrier using a reverse emulsion polymerization method and a dual-template, bifunctional monomer strategy. To improve the specific recognition ability and anti-interference performance of the material in complex matrices, this invention employs a dual-template strategy to prepare a magnetic molecularly imprinted polymer (… This polymer-imprinted layer not only constructs specific recognition sites for the target analyte (Sudan I), but also corresponding recognition sites for a common, potentially cross-reactive, representative interfering substance in the sample matrix (capsaicin in this example). The aim is to pre-occupy a portion of the region that might non-specifically bind to Sudan I, thereby maximizing the reduction of the co-adsorption effect of strong interfering substances at the material level. This ensures high fidelity in the enrichment process of the target analyte, Sudan I, laying the foundation for subsequent high-precision spectroscopic quantitative analysis. The specific synthesis process is as follows:
[0042] (3.1) Formation of the prepolymerized complex: In a 5 ml mixture of ethanol and water (volume ratio 3:1), Sudan I (58 mg, 0.2 mmol) as the target template molecule, capsaicin as the interfering template molecule, methacrylic acid (MAA, 69 μL, 0.8 mmol) as the functional monomer reacting with Sudan I (i.e., the target functional monomer), 4-vinylpyridine as the functional monomer reacting with capsaicin (i.e., the interfering functional monomer), ethylene glycol dimethacrylate (EGDMA, 760 μL, 4 mmol) as the crosslinking agent, and 50 mg of the aforementioned prepolymerized complex were added sequentially. Particles. The mixture was ultrasonically treated for 10 minutes under light-protected conditions to allow the components to fully interact, forming a prepolymerized complex containing two template molecules and corresponding functional monomers, which was then adsorbed onto the surface of a magnetic carrier to obtain an aqueous dispersion.
[0043] (3.2) Emulsion polymerization: 0.3 g of surfactant Span-80 and 10 mg of initiator azobisisobutyronitrile (AIBN) were dissolved in 20 mL of toluene to prepare an oil phase solution. The aforementioned aqueous dispersion was added to the oil phase solution, and the mixture was stirred at 800 rpm for 10 minutes to form a stable reverse emulsion system. The system was purged with nitrogen for 15 minutes to remove oxygen, and then reacted at 70 °C under nitrogen protection and in the dark for 24 hours to complete the polymerization process.
[0044] (3.3) Dual-template elution: After the reaction, the generated polymer particles were collected using an external magnetic field. They were then thoroughly washed sequentially with toluene, a methanol / acetic acid mixture (9:1 volume ratio), and deionized water. The methanol / acetic acid mixture was used to simultaneously and efficiently elute Sudan I and capsaicin template molecules from the polymer network, leaving recognition holes in the imprinted layer that match the spatial configuration and chemical properties of the two molecules. The product was finally vacuum-dried at 60°C for 12 hours to obtain the final product. .
[0045] To conduct comparative verification, non-imprinted polymers were also prepared. The preparation methods of NIPs and The preparation methods are basically the same, the key difference being that no template molecules are added in the prepolymerization stage (i.e., no Sudan I and capsaicin are added), while the remaining components and reaction conditions remain the same.
[0046] To verify the physicochemical properties of the prepared material and confirm its successful synthesis, its morphology, structure, and magnetic properties were characterized. Transmission electron microscopy (TEM) was used for observation, confirming that the prepared material possesses a clear core-shell structure. Fourier transform infrared spectroscopy (FTIR) analysis was performed on the material, revealing that the spectral density is within the range of 590°. The absorption peaks appearing nearby are attributed to the vibrations of the Fe-O bond, 1072. The absorption peak at that point is attributed to the stretching vibration of the Si-O-Si bond, while... In the spectrum, in 1736 A new absorption peak appears at [location], which corresponds to the C=O carbonyl stretching vibration of methacrylic acid in the polymer layer. This result proves that [the following is true]. A molecularly imprinted polymer layer was successfully constructed on the surface of a magnetic carrier. The magnetic properties of the material were measured using a vibrating sample magnetometer (VSM), and the results showed… The material possesses sufficient saturation magnetization to ensure rapid magnetic separation in an aqueous system under the influence of an external magnetic field, with a separation time of less than 1 minute.
[0047] To demonstrate that the Fe3O4@SiO2@MIPs prepared in this invention possesses the ability to achieve efficient and selective enrichment, thereby providing high-quality samples for subsequent quantitative spectroscopic analysis, this section presents a series of quantitative adsorption experiments to assess the specific recognition ability and adsorption performance of this molecularly imprinted polymer. In the static adsorption experiment, 10 mg of MIPs or NIPs were added to a series of Sudan I solutions with initial concentrations ranging from 1 mg / L to 100 mg / L, and incubated with shaking for 60 minutes to reach adsorption equilibrium. Equilibrium adsorption capacity. The calculation formula is:
[0048] ;
[0049] The equilibrium adsorption capacity at different initial concentrations was calculated by measuring the remaining concentration of the solution after adsorption equilibrium. The specific experimental data are recorded as follows:
[0050] Table 1: Equilibrium adsorption data for Sudan I
[0051]
[0052] Table 2: Equilibrium adsorption data for Sudan I
[0053]
[0054] As can be seen from the above data, the equilibrium adsorption capacity of MIPs was significantly higher than that of NIPs at all tested concentrations. For example, at an initial concentration of 100 mg / L, the equilibrium adsorption capacity of MIPs reached 48.67 mg / g, while that of NIPs was only 14.10 mg / g, meaning that the adsorption capacity of MIPs was 3.5 times that of NIPs. This indicates that a large number of specific recognition sites with high affinity for Sudan I were successfully constructed on the material surface using molecular imprinting technology. To evaluate the imprinting effect, an imprinting factor (…) was defined. As a metric, its calculation formula is:
[0055] ;
[0056] The adsorption data were fitted using the Freundlich isotherm adsorption model, and its linear expression is as follows:
[0057] ;
[0058] The fitting results are shown in the table below:
[0059] Table 3: Fitting parameters of the Freundlich model
[0060]
[0061] The Freundlich model showed a good fit to the adsorption data of both materials. >0.97). Adsorption capacity constant of MIPs (11.35) is NIPs The value (1.79) is 6.3 times higher, quantitatively demonstrating that MIPs possess a higher adsorption capacity. Simultaneously, the adsorption strength constant of MIPs... (1.71) Greater than NIPs The value (1.48) indicates that its adsorption process is more favorable.
[0062] In adsorption kinetics experiments, it is also necessary to measure the instantaneous adsorption capacity of MIPs and NIPs at different time points. The calculation formula is as follows: The experimental data are recorded as follows:
[0063] Table 4: Adsorption kinetic data
[0064]
[0065] Table 5: Adsorption kinetic data
[0066]
[0067] Data show that the adsorption rate and equilibrium adsorption capacity of MIPs are much higher than those of NIPs. Pseudo-first-order and pseudo-second-order kinetic models were used to fit the adsorption kinetic data of MIPs. The expression for the pseudo-first-order kinetic model is:
[0068] ;
[0069] The expression for the pseudo-second-order dynamic model is:
[0070] ;
[0071] The expression for the pseudo-second-order dynamic model is:
[0072] The fitting results for MIPs show that the correlation coefficient of the pseudo-second-order kinetic model ( =0.999) is significantly higher than that of the pseudo-first-order dynamic model ( =0.971), and its calculated theoretical equilibrium adsorption capacity (82.64 mg / g) and experimentally determined value (79.86 mg / g) is in high agreement. This result indicates that the adsorption process of Sudan I on MIPs is mainly controlled by a chemisorption mechanism. The combined results of static adsorption and adsorption kinetics experiments fully demonstrate that the product prepared in this invention... It exhibits high adsorption capacity, high selectivity, and rapid adsorption kinetics for Sudan I, meeting the requirements for a high-efficiency magnetic solid-phase extraction material. This is the key material basis for the subsequent realization of a high-sensitivity and high-precision detection method in this invention.
[0073] In the above formula: To balance the adsorption capacity (mg / g); Equilibrium adsorption capacity (mg / g) of MIPs. Equilibrium adsorption capacity of NIPs (mg / g); In time Instantaneous adsorption amount (mg / g); The equilibrium adsorption capacity (mg / g) was measured experimentally. The theoretical equilibrium adsorption capacity (mg / g) calculated for the model. The initial concentration of the solution is (mg / L). The equilibrium concentration of the solution is (mg / L). In time Instantaneous concentration (mg / L); The volume of the solution is in L. The mass of the adsorbent is (g). Freundlich adsorption capacity constant; It is the Freundlich adsorption strength constant; The pseudo-first-order adsorption rate constant ( ); The pseudo-second-order adsorption rate constant ( ); The coefficient of determination;
[0074] After preparing magnetic molecularly imprinted polymers (MIPs) with logic recognition capabilities, the next step is to use this material to construct a standard spectral dataset for training and validating subsequent intelligent analysis models. The construction of this dataset is fundamental to ensuring the accuracy and robustness of the models.
[0075] To establish the intelligent detection model, a dataset containing standard samples of varying concentrations and their corresponding infrared spectra is first required. This dataset aims to comprehensively cover the response of the target analyte at different concentration levels and under different interference backgrounds, providing sufficient data support for subsequent feature extraction and model training.
[0076] The construction process of this standard spectral dataset specifically includes the preparation of gradient-spiked standard samples, and the standardized magnetic solid-phase extraction and Fourier transform infrared spectroscopy (FTIR) acquisition of each standard sample.
[0077] First, graded-spiked standard samples are prepared. This step aims to simulate the possible concentration combinations of target and interfering substances in real samples.
[0078] Chili powder that was confirmed to be free of Sudan I was selected as the blank matrix.
[0079] Prepare acetonitrile standard stock solutions of Sudan I and capsaicin, respectively.
[0080] A series of spiked samples with different concentration combinations were prepared by adding different volumes of Sudan I and / or capsaicin standard stock solutions to an accurately weighed blank matrix (e.g., 1.0 g). The samples were thoroughly mixed and equilibrated at room temperature for a period of time before being used for subsequent analysis.
[0081] To ensure the model can learn the spectral characteristics under different concentrations and interferences, the spiked sample concentration design covers multiple gradients from low to high, and includes cases with a single component and cases with multiple components coexisting. An exemplary spiked sample design scheme is shown in Table 6 below.
[0082] Table 6: Design Scheme for Gradient Spike Standard Samples
[0083]
[0084] After preparing the graded-spiked standard samples, a standardized magnetic solid-phase extraction (MSPE) and FTIR spectroscopy acquisition procedure was performed on each sample to obtain its corresponding infrared spectral data. This standardized procedure ensures that all spectral data within the dataset were obtained under consistent experimental conditions, thereby guaranteeing data comparability. The specific operational steps are as follows:
[0085] Sample extraction: Accurately weigh 1.0 g of spiked sample and place it in a centrifuge tube. Add 10 mL of acetonitrile as the extraction solvent, and then sonicate for 10 minutes in an ultrasonic cleaner. After extraction, centrifuge the tube at a certain speed and collect the supernatant for subsequent enrichment.
[0086] Magnetic solid phase extraction (MSPE): Take a certain volume of the aforementioned supernatant and dilute it with deionized water. Add 20 mg of the aforementioned prepared [prepared compound] to the diluted extract. The mixture was shaken on an oscillator for a certain period of time to reach adsorption equilibrium.
[0087] Separation and purification: After adsorption, a strong magnet is placed on the outer wall of the centrifuge tube. The MIPs carrying the target and interfering substances will then be rapidly adsorbed onto the tube wall. The supernatant is then poured off and discarded. A small amount of solvent can be used to quickly wash the MIPs to remove physically adsorbed impurities, followed by another magnetic separation.
[0088] FTIR Sample Preparation and Spectral Acquisition: The separated MIPs were dried under a nitrogen stream. An appropriate amount of the dried MIPs powder was mixed evenly with spectroscopically pure potassium bromide (KBr) powder in an agate mortar. The mixed powder was transferred to a tableting mold and pressed into a uniform, transparent sheet under standard pressure. This KBr sheet was placed in the sample chamber of the Fourier Transform Infrared Spectrometer and incubated at 4000-4000 nm. Infrared absorption spectra were acquired within the specified wavenumber range. The spectral acquisition parameters were set as follows: resolution 4... The number of scans was 32.
[0089] Repeat the above standardization process for all standard samples in Table 6 to obtain a spectral dataset consisting of sample number, known concentration labels (Sudan I concentration, capsaicin concentration), and corresponding raw FTIR spectral data. This dataset serves as the direct basis for subsequent feature extraction, model training, and validation.
[0090] After obtaining the dataset containing standard samples with gradient concentrations and their corresponding infrared spectra, the next step is data processing and feature extraction. The primary task of this stage is to preprocess the raw spectral data to eliminate or reduce non-chemical interferences introduced by factors such as the instrument, environment, and the physical state of the sample itself.
[0091] Raw Fourier transform infrared (FTIR) spectral data typically contains high-frequency random noise caused by instrument electronics or environmental factors, as well as baseline drift due to sample particle scattering or uneven pellet compression. These interferences can mask the true chemical information related to the concentration of the target analyte, thus affecting the performance of subsequent quantitative models. To eliminate these interferences and improve the accuracy and robustness of subsequent quantitative analysis, spectral data preprocessing is necessary.
[0092] This invention proposes an adaptive spectral preprocessing method that transforms the selection of preprocessing parameters from a subjective setting dependent on operator experience into a data-driven, reproducible, and automated optimization process oriented towards the final model performance. Specifically, this method includes Savitzky-Golay (SG) smoothing and baseline correction of the spectrum, and automatically determines the parameter combination that optimizes the performance of the subsequent quantitative model through a systematic optimization algorithm.
[0093] The specific implementation steps of this adaptive spectral preprocessing method are as follows:
[0094] 1. Define the parameter search space: Construct a discrete search space containing various combinations of preprocessing parameters. This space must include at least the window size for SG smoothing. and polynomial order For example, you can set the window size. The candidate set is {9, 13, 17, 21}, and the polynomial order is... The candidate set is {2,3}. Furthermore, if multiple baseline correction algorithms are used, such as polynomial fitting or asymmetric least squares (ALS), it can also be used as a dimension of the search space.
[0095] 2. Establishing a performance evaluation function and iterative optimization: This step uses iterative calculations to find the optimal combination of parameters.
[0096] Select a set of unevaluated parameter combinations from the parameter search space (e.g., , ).
[0097] This parameter combination is used to perform SG smoothing and baseline correction on the training set portion of the standard spectral dataset. The calculation process for SG smoothing is as follows:
[0098] ;
[0099] On the training set data preprocessed with this parameter combination, a computationally inexpensive surrogate regression model, such as ridge regression, is trained. The purpose of this surrogate model is to quickly evaluate the effectiveness of the current preprocessed parameter combination.
[0100] The performance of the surrogate model was then evaluated using the K-fold cross-validation method, and its mean square error (MSE) across all cross-validation folds was calculated. E) This serves as the performance evaluation value for the current parameter combination. The formula for calculating the root mean square error is as follows:
[0101] ;
[0102] Then record the current parameter combination and its corresponding performance evaluation value.
[0103] 3. Determine the optimal parameter combination: Traverse all parameter combinations within the parameter search space, then compare the performance evaluation values of all parameter combinations, and select the parameter combination that minimizes the root mean square error (RMSE) as the final optimal preprocessing scheme, denoted as ( ). , ).
[0104] 4. Apply the optimal solution: Apply the optimal parameter combination to the entire standard spectral dataset (including the training set and the subsequent test set for validation), performing uniform and standardized preprocessing operations on all spectral data. The spectral data processed in this step will be used for subsequent multidimensional feature extraction.
[0105] As for SG smoothing and baseline correction, their specific algorithms are well-known technologies in this field and will not be elaborated here.
[0106] In the above formula:
[0107] After SG smoothing, in the 1st The spectral absorbance value of the point; For the original spectrum in the 1st The absorbance value of the point; For the SG smoothing filter, the first One coefficient; This is the root mean square error; This represents the total number of samples used for cross-validation; For the first The true concentration value of each sample; For the first The predicted concentration value for each sample;
[0108] After adaptive preprocessing of the raw spectral data, the next step is to transform the optimized spectral data into a feature vector consisting of multiple quantitative indicators, and then train and evaluate a series of machine learning regression models based on this to determine the optimal model for the final quantitative analysis.
[0109] This process specifically includes the extraction of multidimensional features, standardization of feature data, parallel training of multiple regression models, and performance comparison and evaluation.
[0110] First, feature engineering is performed to extract a set of quantitative features from the preprocessed spectral data that can effectively characterize the concentration of the target analyte. This invention does not use the entire spectrum, but rather extracts features related to the target analyte (Sudan Red I) and the internal standard (…). Based on specific spectral region information, a feature set with 16 dimensions is constructed. This feature set is designed to simultaneously capture peak intensity, area, and relative proportion information, thereby comprehensively reflecting the sample's state. The specific feature extraction steps are as follows:
[0111] Peak area calculation: Integrate over a pre-defined characteristic wavenumber range in the spectrum to calculate the peak area.
[0112] The characteristic peak area of the internal standard was calculated separately for 580-600. and 400-450 The integral area of the two regions.
[0113] The characteristic peak areas of Sudan I target analyte were calculated separately for 1600-1610. 1545-1555 and 1345-1355 The integral area of the three regions.
[0114] Peak height calculation: within a small window near the center wavenumber of the characteristic peak (e.g., center wavenumber ± 10). The peak height is determined by taking the maximum absorbance value.
[0115] Internal standard characteristic peak height: at 590 and 425 Calculation at the location.
[0116] Sudan I target analyte characteristic peak height: at 1600 1550 and 1350 Calculation at point.
[0117] Peak area ratio calculation: To achieve internal standard calibration and eliminate systematic errors in sample preparation and measurement, the peak area ratio of each Sudan I characteristic peak is calculated relative to the peak area of each Sudan I characteristic peak. The ratio of the areas of the internal standard characteristic peaks. This is because there are 3 Sudan I characteristic peaks and 2... Characteristic peaks: This step generates a total of 3×2=6 peak area ratio characteristics.
[0118] In summary (refer to) Figure 3 Through the above steps, each spectral sample is converted into a vector consisting of 16 features (2... Peak area + 3 Sudan I peak areas + 2 Peak height + 3 Sudan I peak heights + 6 peak area ratios) are used for subsequent model training.
[0119] After feature extraction, to eliminate the impact of differences in units and numerical ranges between different features on model training, the generated feature dataset needs to be standardized. This embodiment uses robust scalars for data standardization. This method uses statistics that are insensitive to outliers (such as quantiles) for scaling, which, compared to conventional standardization methods using mean and standard deviation, can more effectively reduce the negative impact of potential outliers on model training.
[0120] Next, we move on to the model training and evaluation phase. The goal of this phase is to systematically compare and select the best-performing model for this specific analytical task from a candidate model library containing various heterogeneous algorithms. First, the complete dataset, containing 16-dimensional features and corresponding concentration labels, is randomly divided into training and test sets in an 80:20 ratio. The training set is used to train all candidate models, while the test set is fully retained for the final, unbiased performance evaluation of the trained models.
[0121] This embodiment tested 10 regression models with different working principles, specifically including:
[0122] Linear models: Partial Least Squares Regression (PLS), Ridge Regression, Lasso Regression, ElasticNet Regression;
[0123] Support Vector Machine: Support Vector Machine Regression using Radial Basis Function Kernel (SVM_RBF);
[0124] Ensemble learning models: Random-Forest, Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost);
[0125] Neural network model: Multilayer perceptron (MLP);
[0126] The 10 models were trained on the training set, and then applied to an independent test set for prediction. The coefficient of determination was calculated by comparing the model predictions with the actual concentration values. The root mean square error (RMSE) and root mean square error are used as core metrics for evaluating model performance. The performance of each model on the test set is shown in the table below.
[0127] Table 7: Performance Comparison of 10 Regression Models
[0128]
[0129] Based on the evaluation results in the table above, the optimal model was analyzed and determined. Tree-based ensemble learning methods, especially Random Forest, XGBoost, and GBM, significantly outperformed linear models, support vector machines, and multilayer perceptrons. Among them, the Random Forest model exhibited the best predictive performance, with a determination coefficient (...). The mean square error (MSE) reached 0.938, while the root mean square error (RMSE) was... The concentration of Sudan I was the lowest among all models, at only 7.459 mg / L. This indicates that the random forest model is the most effective at capturing the complex nonlinear relationship between the 16-dimensional spectral features and the Sudan I concentration. In contrast, the linear model and the SVM_RBF model performed poorly. The values are all significantly lower than 0.75, indicating that they are inadequate for handling the inherent nonlinearity of the data in this task.
[0130] To further verify the stability and generalization ability of the model, cross-validation was performed on the random forest model. The results show that its average performance in cross-validation is highly consistent with its performance on the independent test set, indicating that the model has low overfitting and good robustness.
[0131] In summary, through systematic comparison and evaluation, random forest was determined to be the optimal model for this analysis task. This trained and validated random forest model will be used for quantitative prediction when detecting unknown samples. A scatter plot of the model's predicted values versus actual values on the test set (see reference...) demonstrates this. Figure 2 The data points are observed to be closely distributed around the ideal 1:1 fitted line, which visually demonstrates its excellent predictive accuracy.
[0132] In one application embodiment, to further verify the accuracy and reliability of the optimal quantitative model (i.e., random forest model) determined by the present invention in actual detection, this embodiment simulates the detection process of samples with unknown concentrations.
[0133] First, a set of independent test samples for model validation was prepared. This set of samples was prepared by artificially adding (i.e., spiked) a known amount of Sudan I standard solution to a blank sample matrix confirmed to be free of Sudan I (commercially available chili powder in this example), thereby obtaining a series of simulated test samples with precise real concentrations. The concentrations of this set of validation samples covered low, medium, and high concentration levels, as shown in the table below:
[0134] Table 8: Preparation of validation set samples
[0135]
[0136] Next, the five validation set samples were all tested according to the complete analysis workflow established in this invention. The testing steps for each sample strictly followed the standardized procedure below:
[0137] Sample pretreatment: Accurately weigh a certain mass of the verification sample, extract it with a solvent, and then use the magnetic molecularly imprinted polymer prepared above. Selective enrichment is performed, and then the polymer adsorbed with the target substance is dried.
[0138] Spectral acquisition: The dried polymer was mixed with potassium bromide and compressed into tablets, and its infrared spectrum was acquired using a Fourier transform infrared spectrometer.
[0139] Data processing and feature extraction: The acquired raw spectra are preprocessed using the determined optimal parameters. Then, a 16-dimensional feature vector consisting of peak area, peak height, and peak area ratio is calculated and extracted according to a predetermined algorithm.
[0140] Concentration prediction: The 16-dimensional feature vector obtained for each sample is input into the trained and fixed random forest model, and the model directly outputs the predicted value of Sudan I concentration in the sample.
[0141] The model's predicted concentration values for five validation samples were compared with their known true concentration values, and the relative error was calculated to evaluate the model's prediction accuracy. The formula for calculating the relative error is:
[0142] ;
[0143] In the formula, This is relative error; This represents the true concentration value of the sample. This represents the predicted concentration value from the model.
[0144] The model's prediction results and accuracy evaluation are shown in Table 9:
[0145] Table 9: Prediction results of the random forest model for the validation set samples
[0146]
[0147] As shown in the table above, for all validation samples with different concentrations, the random forest model determined in this invention provides predictions that are highly close to the true values. The relative error (RE) for all samples is less than 4%, indicating that the model has good prediction accuracy and stability in low, medium, and high concentration ranges.
[0148] The results of this application example clearly demonstrate that the method constructed in this invention, based on magnetic molecular imprinted polymer enrichment-Fourier transform infrared spectroscopy detection and combined with a random forest regression model, can perform rapid and accurate quantitative analysis of Sudan I in samples, and has value and potential for application in practical detection.
Claims
1. A rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning, characterized in that, Includes the following steps: Step a: Provide the extract to be tested obtained from the chili powder sample; Step b: Enrich Sudan I in the extract to be tested using a magnetic molecularly imprinted polymer, and obtain the magnetic molecularly imprinted polymer loaded with Sudan I by magnetic separation. Step c: Obtain the Fourier transform infrared spectrum of the magnetic molecularly imprinted polymer loaded with Sudan I; Step d: Extract multidimensional features from the Fourier transform infrared spectrum that include Sudan I characteristic peak information and magnetic core characteristic peak information in the magnetic molecularly imprinted polymer; Step e: Input the multidimensional features into a pre-trained machine learning regression model, and have the model output the concentration of Sudan I in the extract to be tested.
2. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 1, characterized in that, In step b, the magnetically imprinted polymer contains dual recognition sites for Sudan I, the target analyte, and capsaicin, a representative interfering agent.
3. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 1, characterized in that, The multidimensional feature extraction in step d includes calculating the ratio of the characteristic peak area of Sudan I to the characteristic peak area of the Fe3O4 magnetic core serving as an internal standard in the magnetic molecularly imprinted polymer.
4. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 3, characterized in that, The calculation of the peak area ratio specifically involves calculating the peak area ratio of Sudan I in the range of 1600-1610. 1545-1555 1345-1355 The characteristic peak area within the wavenumber range, and The magnetic core is between 580-600. and 400-450 The ratio of the characteristic peak areas within the wavenumber range.
5. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 3, characterized in that, The multidimensional features in step d further include: The characteristic peak area of the magnetic core, the characteristic peak area of Sudan I, and the The characteristic peak height of the magnetic core, the characteristic peak height of Sudan I, and the characteristic peak area of Sudan I compared with the... The ratio of the area of the characteristic peak of the magnetic core.
6. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 1, characterized in that, The pre-trained machine learning regression model is a random forest model.
7. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 6, characterized in that, The random forest model is trained using a standard spectral dataset. The construction of the standard spectral dataset includes: preparing gradient-spiked standard samples containing different concentration gradients of Sudan I and different concentration gradients of capsaicin, and obtaining their corresponding Fourier transform infrared spectra.
8. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 1, characterized in that, Before step d, a preprocessing step is also included for the Fourier transform infrared spectrum. The parameters for the preprocessing are determined by an automated optimization process, which aims to select the parameter combination that minimizes the root mean square error of the subsequent quantitative model.
9. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 1, characterized in that, The enrichment step in step b specifically includes: adding the magnetic molecularly imprinted polymer to the sample extract for oscillation and adsorption, and then performing magnetic separation by applying an external magnetic field.
10. The rapid detection method for Sudan Red in chili powder based on nanoparticles and machine learning according to claim 1, characterized in that, The step of obtaining the Fourier transform infrared spectrum in step c includes: mixing the magnetic molecularly imprinted polymer loaded with Sudan I obtained in step b with potassium bromide and pressing it into a thin sheet, and then performing spectral acquisition.