Method for rapidly identifying sun-screening agent components based on machine learning algorithm in combination with ultraviolet spectrum characteristics and application thereof
By combining machine learning algorithms with ultraviolet spectral features, the problems of high dimensionality, strong noise interference, and severe feature overlap in sunscreen ingredient analysis have been solved, enabling rapid, accurate identification and non-destructive testing of sunscreen ingredients. This method is suitable for sunscreen formulation optimization and quality control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies for analyzing sunscreen components suffer from problems such as high dimensionality of ultraviolet spectra, strong noise interference, and severe feature overlap, resulting in traditional chemical analysis methods being complex, time-consuming, and unable to achieve non-destructive and rapid analysis.
We employ machine learning algorithms combined with ultraviolet spectral features, and through data preprocessing and feature engineering optimization, we utilize random forest, GBDT, multi-output linear regression, MLP, PLA, and CNN algorithms to identify sunscreen ingredients. We also extract feature wavelengths using EN and PLS methods to establish an ingredient content prediction model.
It enables rapid and accurate identification of sunscreen ingredients, with an accuracy rate of 91.67%, significantly improving identification efficiency and accuracy, shortening the analysis time to 30 minutes, achieving non-destructive testing, and is suitable for rapid optimization of sunscreen formulations and quality control.
Smart Images

Figure CN121954889A_ABST
Abstract
Description
A method for rapid identification of sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features and its application. Technical Field
[0001] This invention belongs to the field of sunscreen ingredient analysis technology, specifically, it relates to a method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features and its application. Background Technology
[0002] Near-ultraviolet (UV) spectroscopy, due to its advantages of speed, non-destructive nature, and ability to simultaneously analyze multiple components, has shown broad application prospects in the field of material composition detection. Its principle is based on the selective absorption of ultraviolet light with wavelengths of 200–400 nm by molecules. By measuring the UV absorption spectrum or transmittance (the ratio of UV light transmitted through sunscreen), and combining this with chemometric methods, a quantitative relationship between UV absorption spectrum and component content is established. However, in the application of sunscreen component analysis, this technology faces several challenges: on the one hand, UV spectra have high dimensionality and are subject to interference such as baseline drift and random noise; on the other hand, the absorption spectra of various sunscreens overlap significantly. For example, the absorption peaks of zinc oxide at 310–340 nm and titanium oxide at 330–380 nm partially overlap, leading to mutual interference of spectral characteristic information. Simply relying on the traditional least squares method cannot meet the requirements for accurate component content detection in actual production.
[0003] Currently, the analysis of sunscreen components mainly relies on traditional chemical analysis methods such as high-performance liquid chromatography (HPLC) and gas chromatography-mass spectrometry (GC-MS). While these methods offer high accuracy, they suffer from significant drawbacks: complex sample pretreatment requiring multiple steps such as extraction and purification, taking several hours; and the analysis process consumes large amounts of organic solvents, causing environmental pollution and damage, thus failing to achieve non-destructive and rapid sample analysis. With the increasing demand for sunscreen formulation optimization, a novel detection technology is urgently needed that simultaneously achieves efficient, non-destructive, and multi-component quantitative analysis. Machine learning algorithms, such as random forests and support vector machines, have significant advantages in processing high-dimensional data and establishing complex nonlinear mapping relationships. Targeted data preprocessing and feature engineering optimization can improve model performance and address the problems of high dimensionality, strong noise interference, and severe feature overlap in ultraviolet spectral data. Summary of the Invention
[0004] The purpose of this invention is to provide a method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features. This method achieves rapid identification of sunscreen ingredients by combining ultraviolet spectral features with machine learning algorithms, reducing the time required for traditional chemical analysis to 30 minutes. It is non-destructive and accurate, solves the shortcomings of traditional chemical analysis methods, and meets the high-efficiency detection requirements for sunscreen formulation optimization.
[0005] Another objective of this invention is to provide an application of the above-mentioned method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features.
[0006] The objective of this invention is achieved through the following technical solution: A method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral characteristics, comprising the following steps: S1. Mixing glycerin, propylene glycol, magnesium sulfate, and deionized water, heating to 78-85°C to obtain a phase B mixture, and mixing zinc oxide, titanium dioxide, octocrylene, diethylethoxyphenol methoxyphenyl triazine, and humosasulfate to obtain a phase C mixture. The phase C mixture and phase B mixture are sequentially added to a phase A mixture obtained by mixing sorbitan sesquioleate, white beeswax, methylparaben, white oil, magnesium stearate, and 200 / 350 cst polydimethylsiloxane, heating to 78-85°C to obtain a mixture. This mixture is then cooled to 40-48°C, and isomeric alcohol ethers are added and stirred. The mixture is then heated to emulsify and form a sunscreen composition; S2. Spectral data acquisition: The sunscreen composition is coated onto a substrate and dried. Ultraviolet light at wavelengths of 290-400 nm is scanned to obtain the weighted average transmittance; S3. Spectral preprocessing and dimensionality reduction: The preprocessed data of the weighted average transmittance obtained in step S2, which undergoes SG filtering and SNV transformation sequentially, is compressed using the central difference method to obtain dimensionality-reduced data; S4. Machine learning identification: Random forest, GBDT, multi-output linear regression, MLP, PLA, and CNN algorithms are used to identify the C-phase components of the dimensionality-reduced data obtained in step S3. Then, the EN method, PLS method, or EN+PLS combination method are used to extract feature wavelengths that are highly correlated with the C-phase mixture content of the sunscreen composition as the dataset; The above dataset is divided into training and test sets in a 4:1 ratio using the KS method. The training model on the training set is trained by the R-squared of the test set. 2 ≥0.8 recognition accuracy, select the optimal preprocessing combination; S5. Model establishment and optimization: Based on the characteristic wavelengths that are highly correlated with the C phase mixture content of the sunscreen composition in step S4, establish a prediction model for the C phase mixture content of the sunscreen composition, and by adjusting the parameters of the prediction model, the number of trees in the random forest is 50~120, and the number of hidden layer neurons in the MLP is 50~120, to achieve rapid identification of sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features.
[0007] Preferably, in step S1, the mass ratio of glycerin, propylene glycol, magnesium sulfate, and deionized water is 8:5:1:44.45; the mass ratio of zinc oxide, titanium dioxide, octocrylene, diethylethoxyphenol methoxyphenyl triazine, and humosasulfate is (0~5):(0~5):(0~5):(0~5):(0~5), and not all of them are 0; the mass ratio of dehydrated sorbitan sesquioleate, white beeswax, methylparaben, white oil, magnesium stearate, and 200 / 350cst polydimethylsiloxane is 3:3:0.15:12:0.6:3; the mass ratio of the mixture to the isomeric alcohol ether is 99.2:0.8, and the C-phase mixture is 15wt% of the sunscreen composition.
[0008] Preferably, the mass of the coating in step S2 is 1~3 mg / cm³. 2 The slide is a PMMA plate or a quartz glass slide.
[0009] Preferably, the parameters of the SG filter in step S3 are set as follows: window width of 5~11, polynomial order of 2~4, and high-frequency noise is eliminated by local polynomial fitting; the SNV transform is performed according to the formula... , where μ i and σ i These are the mean and standard deviation of the i-th wavelength point, respectively, used to eliminate sample matrix effects.
[0010] Furthermore, the specific operation of the dimensionality reduction process in step S3 is as follows: The central difference method is used to eliminate spectral baseline drift, according to the formula... ,in, Let y(i) be the rate of change of transmittance at the i-th wavelength point, y(i) be the transmittance of the ultraviolet spectrum at the i-th wavelength point, and Δy represent the interval between wavelength points, i.e., the wavelength difference between two adjacent spectral sampling points; then calculate the contrast of the enhanced characteristic absorption peak according to the formula. ,in, Let be the rate of change of transmittance at the i-th wavelength point. is the first derivative of the ultraviolet spectrum at the i-th wavelength point, and Δy represents the interval between wavelength points, that is, the wavelength difference between two adjacent spectral sampling points.
[0011] Preferably, in step S4, the regularization parameter α of the EN method is set to 0.1~1, and the L1 regularization ratio is 0.5~1; the sampling scale parameter of the PLS method is set to calculate the mean and standard deviation of the spectral transmittance of the sample, and then the data is standardized. The standardized value = original transmittance value - mean transmittance / standard deviation, and linear combination variables are extracted from the relevant characteristics of 290-400nm ultraviolet transmittance and the content of each component of the sunscreen.
[0012] A computer-readable storage medium stores a computer program that, when executed by a processor, enables the method for rapid identification.
[0013] The method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features is applied in the field of sunscreen spectral analysis.
[0014] The core objective of the L1 regularization ratio in this invention is feature selection and sparsification, avoiding the loss of useful features due to excessive sparsity. The scale parameter, related to data standardization, is not a single value but a set of data scaling rules. Standardization is achieved by calculating "standardized value = original transmittance value - mean transmittance / standard deviation," eliminating characteristic differences in transmittance at wavelengths from 290 to 400 nm. There is no fixed numerical range; it needs to be dynamically calculated based on the actual mean and standard deviation of the sample transmittance. SG filtering is responsible for eliminating high-frequency noise, SNV's core function is to eliminate sample matrix effects (not primarily baseline removal), and 1D processing specifically targets baseline drift and enhances characteristic peak contrast, thus having different dimensions of action.
[0015] Compared with existing technologies, this invention has the following beneficial effects: This invention focuses on the rapid identification of sunscreen ingredients. Through a specific preparation process, sunscreens containing different proportions of zinc oxide, titanium dioxide, octocrylene, diethylethoxyphenol methoxyphenyl triazine, and homosalate are obtained. The ultraviolet transmittance in the 290-400 nm wavelength range is measured using an integrating sphere detector. After data processing (SG and SNV preprocessing, 1D and 2D derivatives of the spectral data for dimensionality reduction to eliminate baseline drift), the training and test sets are divided in a 4:1 ratio. Six machine learning algorithms—random forest, GBDT, multi-output linear regression, MLP, PLA, and CNN—are used to model and identify the ingredients of the sunscreens. The R-value is validated through the test set. 2 The optimal preprocessed data set was selected. Furthermore, characteristic wavelengths were extracted using EN, PLS, and EN+PLS combination methods to reduce data dimensionality. Based on the selected characteristic wavelengths, a prediction model for sunscreen ingredient content was established, and its performance was optimized through parameter tuning.
[0016] The method of this invention enables rapid and accurate identification of sunscreen ingredients, with an accuracy rate of 91.67%. 2 With a value ≥0.94, the identification efficiency and accuracy are significantly improved, providing strong technical support for the research and development and quality control of sunscreen products. Compared with traditional methods, the identification efficiency and accuracy are significantly improved, and it has important application value in the field of spectral analysis of sunscreen agents.
[0017] The method of this invention is compatible with different algorithms. It can shorten the time required for chemical analysis to within 30 minutes and achieve non-destructive testing of samples, making it suitable for rapid optimization of sunscreen formulations and quality control. Attached Figure Description
[0018] Figure 1 is a flowchart of the overall system processing of the present invention; Figure 2 is the original spectral curve of the sunscreen composition of Example 1; Figure 3 is the characteristic curve of SG and SNV treatment of Example 2; Figure 4 is the characteristic curve of 1D and 2D treatment of Example 2; Figure 5 is the characteristic curve of 1D+SNV treatment and 1D+SG treatment of Example 2; Figure 6 is the characteristic curve of 2D+SNV treatment and 2D+SG treatment of Example 2; Figure 7 is a scatter plot of the predicted results of zinc oxide in Example 2; Figure 8 is a scatter plot of the predicted results of titanium oxide and homosalate in Example 2; Figure 9 is a scatter plot of the predicted results of the content of octocrylene and diethylethoxyphenol methoxyphenyl triazine in Example 2. Detailed Implementation
[0019] In this invention, unless otherwise specified, all raw material components are commercially available products well-known to those skilled in the art. In this invention, unless otherwise specified, the technical means used in the embodiments are conventional means well-known to those skilled in the art. The invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0020] Figure 1 is a flowchart of the overall system processing of the present invention. As shown in Figure 1, the rapid identification system for sunscreen ingredients of the present invention follows a complete process of "sample preparation → spectral acquisition → data processing → feature extraction → model training → ingredient identification", and uses six machine learning algorithms such as random forest to build a prediction model and adjust its parameters; ultimately achieving rapid and non-destructive identification of sunscreen ingredients.
[0021] Table 1 Formulation of sunscreen compositions
[0022] Table 1 shows the formulation of the sunscreen composition of the present invention. As can be seen from Table 1, the basic formulation and other ingredients are consistent. The C-phase mixture (sunscreen agent) is 15 wt% of the sunscreen composition. Only the proportion of the components of the C-phase mixture is changed. The purpose is to reduce variable interference, eliminate interference from differences in total concentration, and focus on the correlation between component proportions and the spectrum. Example 1
[0023] 3g of dried sorbitan sesquioleate, 3g of white beeswax, 0.15g of methylparaben, 12g of white oil, 0.6g of magnesium stearate, and 3g of 200cst polydimethylsiloxane were mixed and heated to 80°C until completely melted to obtain phase A. 8g of glycerin, 5g of propylene glycol, 1g of magnesium sulfate, and 44.45g of deionized water were mixed and heated to 80°C to form a homogeneous solution to obtain phase B. 1g of zinc oxide, 3g of titanium dioxide, 4g of octocrylene, 4g of diethylethoxyphenol methoxyphenyl triazine, and 3g of homosalate were mixed evenly to obtain phase C. Phase C was first added to phase A and stirred homogenized for 5 minutes. Then, phase B was slowly added and homogenized for another 10 minutes. After cooling to 45°C, 0.8g of isomeric alcohol ether was added and stirred for 30 minutes to emulsify and form the sunscreen composition. Phase C (sunscreen agent) constituted 15 wt% of the sunscreen composition.
[0024] Spectral data acquisition: 3M medical tape was adhered to a clean quartz glass slide (2 mm thick). The prepared sunscreen composition was drawn up with a syringe and evenly dropped onto the tape surface, spreading it evenly to achieve a sample coverage of 2 mg / cm³ on the tape surface. 2 Then, dry in the dark for 20 minutes, scan the wavelength range of 290~400nm using an ultraviolet integrating sphere detector, take points at 1nm intervals, repeat the test 4 times for each sample, measure the ultraviolet transmittance (T%), take the weighted average transmittance data, and obtain the real data of each formulation; preprocess the collected spectral data: in the Python environment 3.12.9, use sklearn and numpy packages, and use standard normal variable transformation (SNV) and Savitzky-Golay smoothing (SG) to preprocess the original data (real data of each formulation) in step 2 respectively. First, use SG filtering with a window width of 7 and a polynomial order of 3 to eliminate high-frequency noise, and then use SNV transformation formula (1) to eliminate sample matrix effect (non-primary baseline removal). (Equation 1), where μ i and σ i The mean and standard deviation of the i-th wavelength point are used to obtain two new preprocessed spectral datasets, namely the SNV-processed spectral dataset and the SG-processed spectral dataset.
[0025] Dimensionality reduction: According to formula (2), the first derivative is calculated using the central difference method to eliminate baseline drift in 1D dimensionality reduction. (Equation 2), where, Let y(i) be the rate of change of transmittance at the i-th wavelength point, y(i) be the transmittance of the ultraviolet spectrum at the i-th wavelength point, and Δy represent the interval between wavelength points, i.e., the wavelength difference between two adjacent spectral sampling points; then, according to formula (3), the second derivative is calculated in 2D dimensionality reduction to enhance the contrast of characteristic peaks. (Equation 3), where, Let be the rate of change of transmittance at the i-th wavelength point. Let be the first derivative of the ultraviolet spectrum at the i-th wavelength, and Δy represent the interval between wavelength points, i.e., the wavelength difference between two adjacent spectral sampling points. Then, the data processed by SNV and SG are subjected to 1D and 2D smoothing respectively to further reduce dimensionality and noise, resulting in two sets of dimensionality-reduced preprocessed spectral data. Finally, 1D feature parameters such as characteristic peak position, intensity, full width at half maximum (FWHM), and band integral area are extracted from the preprocessed spectral data to construct a 1D feature matrix. Based on the preprocessed spectral data, a synchronous spectral matrix is constructed, and the correlation coefficient between spectra is calculated to generate a 2D feature matrix.
[0026] When building the machine learning model, six machine learning algorithms were used: Random Forest, GBDT, Multi-Output Linear Regression, MLP, PLA, and CNN. These were combined with 1D and 2D feature matrices to establish a predictive model for the proportion of sunscreen ingredients. In a Python 3.12.9 environment using sklearn and numpy packages, the initial screening was performed using the EN algorithm to filter features across the entire wavelength range. Cross-validation was then used to optimize the regularization parameter α, eliminating most irrelevant features and reducing the number of features from the original dimensionality to approximately 30%. With regularization parameter α=0.5 and L1 regularization ratio of 0.8, 37 feature wavelengths were selected. Then, the PLS algorithm was used to progressively extract components, reducing feature dimensionality while retaining key information, further reducing the number of features from the original dimensionality to approximately 10-20%. Ten feature wavelengths were selected through PLS sampling. Grid search combined with a validation test set was used to optimize the key parameters of each model, such as the number of decision trees and maximum depth in the Random Forest, and the kernel function type and number of neurons in the MLP, using R... 2 Using the correlation coefficient as an indicator, the random forest model was validated through a test set with the following optimization parameters: 100 trees and a maximum depth of 50.
[0027] Model Validation and Optimization: Dataset Partitioning: The Kennard-Stone (KS) dataset partitioning method was adopted to divide the preprocessed spectral data into a 4:1 ratio: a training set (48 samples covering the largest range of the data space selected by Euclidean distance) and a test set (the remaining 12 samples). This ensured that the training set contained samples across the entire concentration range, the test set was highly representative, the partitioning results were stable, and the model was suitable for small-sample, high-dimensional data scenarios. During training, the model performance was evaluated by reading the test set, and the feature combinations and algorithm parameters were continuously adjusted. Random forest was used to predict the homosalate identification (Ri) on the test set. 2 The accuracy was 0.94, and the average accuracy of total component recognition reached 91.67%.
[0028] The model achieves a recognition accuracy of ≥85% on the test set, R 2The requirement is ≥0.9.
[0029] Figure 2 shows the original spectral curve of the sunscreen composition of Example 1. As can be seen from Figure 2, there are obvious characteristic absorption peaks in the wavelength range of 290–400 nm, some component absorption peaks overlap, and the spectrum is accompanied by slight baseline drift and high-frequency noise, confirming the necessity of subsequent SG filtering, SNV transformation, and 1D dimensionality reduction processing. Example 2
[0030] The difference from Example 1 is as follows: In step 1, the C-phase mixture contains 3g zinc oxide, 4g titanium oxide, 3g octocrylene, 3g diethylethoxyphenol methoxyphenyl triazine, and 2g homosalate (accounting for 15wt% of the total mass); in step 3, the SG filter window width is set to 9 and the polynomial order to 2 in the data preprocessing, and dimensionality reduction is only performed using 1D first derivative processing; in step 4, feature extraction uses the EN+PLS combination, with α=0.3 and L1 ratio set to 0.6, and PLS sampling retains 10 feature wavelengths. Six machine learning algorithms, namely random forest, GBDT, multi-output linear regression, MLP, PLA, and CNN, were used for modeling. The optimized parameters were verified on the random forest model test set: 120 trees and a maximum depth of 60. The octocrylene recognition score on the test set was 0.9268, and the average recognition accuracy of the total components was 91.67%.
[0031] Figure 3 shows the SG and SNV processing feature curves of Example 2. As can be seen from Figure 3, the spectral curves after SG filtering and SNV processing are smoother than the original spectra, with significantly reduced high-frequency fluctuations, clearer feature peak outlines, and no obvious baseline shift. This indicates that SG filtering can effectively eliminate high-frequency noise, and SNV transformation can remove sample matrix effects. Both preprocessing methods significantly improve the purity of spectral data, laying the foundation for subsequent dimensionality reduction and feature extraction. Figure 4 shows the 1D and 2D processing feature curves of Example 2. As can be seen from Figure 4, the 1D processing curve highlights key information such as the position and intensity of feature peaks, while the 2D processing curve presents a clear inter-spectral correlation structure. Both significantly compress the dimensionality of the original data. This indicates that 1D feature extraction focuses on the core features of a single spectrum, while 2D feature extraction strengthens the inter-spectral correlation information. Both dimensionality reduction methods can retain effective information related to sunscreen ingredients, achieving the structured transformation of high-dimensional data. Figure 5 shows the 1D+SNV processing feature curve and 1D+SG processing feature curve of Example 2. As shown in Figure 5, the curves after combined 1D+SNV and 1D+SG processing show further reduction in noise interference, significantly improved contrast of characteristic peaks, and more prominent characteristic wavelengths related to the C-phase mixture components. This indicates that the combination of preprocessing and 1D dimensionality reduction has a synergistic effect, accurately screening key information and eliminating redundancy. Compared with single processing, it is more conducive to subsequent models capturing the correlation between components and spectra. Figure 6 shows the characteristic curves of 2D+SNV and 2D+SG processing in Example 2. As shown in Figure 6, the curves after combined 2D+SNV and 2D+SG processing show a more regular structure of the correlation matrix between spectra, reduced invalid signal interference, and more concentrated effective correlation features. This indicates that the combined processing can enhance the effectiveness of the 2D feature matrix, further compress irrelevant dimensions, and provide better data support for EN / PLS characteristic wavelength screening. Figure 7 is a scatter plot of the zinc oxide prediction results in Example 2. As shown in Figure 7, the predicted and actual values of zinc oxide are closely distributed around the diagonal, with no obvious deviation. This demonstrates that the model of the present invention has extremely high accuracy in predicting the content of zinc oxide, with excellent fitting effect, and can accurately establish the quantitative relationship between its spectral characteristics and content. Figure 8 is a scatter plot of the prediction results for titanium dioxide and homosalate in Example 2. As shown in Figure 8, the predicted values of titanium dioxide and homosalate are closely distributed with the actual values, generally fitting the diagonal line with small deviation. This indicates that the model has strong stability in the quantitative prediction of two different sunscreen ingredients, titanium dioxide and homosalate, is not affected by the overlapping of component spectra, and has the ability to simultaneously and accurately identify multiple components. Figure 9 is a scatter plot of the prediction results for octocrylene and diethylethoxyphenol methoxyphenyl triazine in Example 2. As shown in Figure 9, the predicted values of octocrylene and diethylethoxyphenol methoxyphenyl triazine do not have a significant discrete distribution with the actual values, and the overall fitting degree is high. This indicates that the model has no significant deviation in the identification of the two components, octocrylene and diethylethoxyphenol methoxyphenyl triazine, verifying the adaptability of the method to different sunscreen ingredients, and demonstrating comprehensive and reliable overall prediction ability.
[0032] The difference between Comparative Example 1 and Example 1 is that: data preprocessing only performed SNV transformation, directly inputting full-band spectral data (290-400nm) into six machine learning models, without using traditional feature extraction methods or performing EN and PLS feature extraction on the spectral acquisition. On the test set, homosalate identification R... 2 The value is 0.5301, and the octocrylene recognition value is R. 2 The value was 0.6679, significantly lower than in Examples 1 and 2, indicating that feature extraction can improve the model's accuracy. Examples 1 and 2 used the EN+PLS combination to screen feature wavelengths, reducing data dimensionality while retaining key information, significantly improving model accuracy compared to the full-band analysis in Comparative Example 1. The six machine learning algorithms performed better in handling nonlinear relationships, demonstrating the compatibility of the method with different algorithms. This method reduces the time required for traditional chemical analysis to within 30 minutes and achieves non-destructive sample detection, making it suitable for rapid optimization of sunscreen formulations and quality control.
[0033] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features, characterized in that, Includes the following steps: S1. A mixture of glycerin, propylene glycol, magnesium sulfate, and deionized water is heated to 78-85°C to obtain a phase B mixture. A mixture of zinc oxide, titanium oxide, octocrylene, diethylethoxyphenol methoxyphenyl triazine, and humosasulfate is then added sequentially to a phase A mixture obtained by mixing sorbitan sesquioleate, white beeswax, methylparaben, white oil, magnesium stearate, and 200 / 350 cst polydimethylsiloxane and heating to 78-85°C to obtain a mixture. When the mixture is cooled to 40-48°C, isomeric alcohol ethers are added and stirred. The mixture is then heated to emulsify and form a sunscreen composition. S2. Spectral Data Acquisition: The sunscreen composition is coated onto a substrate and dried. Ultraviolet light at wavelengths of 290–400 nm is scanned to obtain the weighted average transmittance. S3. Spectral Preprocessing and Dimensionality Reduction: The weighted average transmittance obtained in step S2 is preprocessed using the central difference method, followed by SG filtering and SNV transformation to compress the dimensionality, resulting in dimensionality-reduced data. S4. Machine Learning Recognition: Random Forest, GBDT, Multi-Output Linear Regression, MLP, PLA, and CNN algorithms are used to identify the C-phase components of the dimensionality-reduced data obtained in step S3. Then, the EN method, PLS method, or a combination of EN and PLS methods are used to extract characteristic wavelengths highly correlated with the C-phase mixture content of the sunscreen composition as the dataset. The dataset is then divided into a training set and a test set at a 4:1 ratio using the KS method. The training model is trained on the training set and then tested using the R-squared algorithm on the test set. 2 ≥0.8 recognition accuracy, select the optimal preprocessing combination; S5. Model establishment and optimization: Based on the characteristic wavelengths that are highly correlated with the C phase mixture content of the sunscreen composition in step S4, establish a prediction model for the C phase mixture content of the sunscreen composition, and by adjusting the parameters of the prediction model, the number of trees in the random forest is 50~120, and the number of hidden layer neurons in the MLP is 50~120, to achieve rapid identification of sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features.
2. The method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features according to claim 1, characterized in that, In step S1, the mass ratio of glycerin, propylene glycol, magnesium sulfate, and deionized water is 8:5:1:44.45; the mass ratio of zinc oxide, titanium dioxide, octocrylene, diethylethoxyphenol methoxyphenyl triazine, and humosatriol is (0~5):(0~5):(0~5):(0~5):(0~5), and they are not all 0 at the same time; the mass ratio of dehydrated sorbitan sesquioleate, white beeswax, methylparaben, white oil, magnesium stearate, and 200 / 350cst polydimethylsiloxane is 3:3:0.15:12:0.6:3; the mass ratio of the mixture to the isomeric alcohol ether is 99.2:0.8, and the C-phase mixture is 15wt% of the sunscreen composition.
3. The method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features according to claim 1, characterized in that, The coating mass in step S2 is 1~3 mg / cm²; the slide is a PMMA plate or a quartz glass slide.
4. The method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features according to claim 1, characterized in that, The parameters for the SG filtering in step S3 are set as follows: window width of 5~11, polynomial order of 2~4, and high-frequency noise is eliminated through local polynomial fitting; the SNV transform is performed according to the formula... , where μ i and σ i These are the mean and standard deviation of the i-th wavelength point, respectively, used to eliminate sample matrix effects.
5. The method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features according to claim 1, characterized in that, The specific operation of the dimensionality reduction process in step S3 is as follows: The central difference method is used to eliminate spectral baseline drift, according to the formula... ,in, Let y(i) be the rate of change of transmittance at the i-th wavelength point, y(i) be the transmittance of the ultraviolet spectrum at the i-th wavelength point, and Δy represent the interval between wavelength points, i.e., the wavelength difference between two adjacent spectral sampling points; then calculate the contrast of the enhanced characteristic absorption peak according to the formula. ,in, Let be the rate of change of transmittance at the i-th wavelength point. is the first derivative of the ultraviolet spectrum at the i-th wavelength point, and Δy represents the interval between wavelength points, that is, the wavelength difference between two adjacent spectral sampling points.
6. The method for rapidly identifying sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features according to claim 1, characterized in that, In step S4, the regularization parameter α of the EN method is set to 0.1~1, and the L1 regularization ratio is 0.5~1; the sampling scale parameter of the PLS method is set to calculate the mean and standard deviation of the spectral transmittance of the sample, and then the data is standardized. The standardized value = original transmittance value - mean transmittance / standard deviation. Linear combination variables are extracted from the relevant characteristics of ultraviolet transmittance in the 290~400nm range and the content of each component of the sunscreen.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 6 for rapid identification.
8. The application of the method for rapid identification of sunscreen ingredients based on machine learning algorithms combined with ultraviolet spectral features according to any one of claims 1 to 6 in the field of sunscreen spectral analysis.