Cosmetic ingredient sensitization prediction method based on machine learning
By expanding the cosmetic ingredient dataset and combining it with multiple machine learning models, the data and model interpretability issues in cosmetic sensitization prediction were resolved, resulting in more accurate sensitization prediction and ensuring cosmetic safety.
Patent Information
- Application Number
- CN202510929244.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-31
AI Technical Summary
Existing cosmetic ingredient sensitization prediction models suffer from data discrepancies and insufficient model interpretability, leaving consumers still at risk of cosmetic sensitization.
The Local Lymph Node Experiment (LLNA) dataset was expanded to include data from human, animal, and non-animal sources. Morgan fingerprints, MACCS Keys, and ModRed molecular descriptors were used, combined with Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM) to construct a universal consensus prediction framework. The LazyPredict tool was used to select models, and a dynamic weight voting mechanism was designed for combined prediction.
It improves the accuracy and interpretability of predicting the sensitization effects of cosmetic ingredients, provides a more reliable reference for cosmetic safety, and reduces the risks to consumers due to sensitization.
Smart Images

Figure CN120877908A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of machine learning and cosmetic ingredient sensitization prediction technology, and particularly relates to a machine learning-based method for predicting cosmetic ingredient sensitization. Background Technology
[0002] The beauty product industry is gradually gaining a stronger position in the global landscape, characterized by the dominance of major global players, and its market prospects are broad. With high-quality economic development and steadily increasing consumer demands, the safety of cosmetics is one of the most pressing concerns for every consumer. Therefore, analyzing and testing allergenic ingredients in cosmetics is of great significance for quality and safety control. However, with the booming development of the cosmetic market, the complexity and diversity of its ingredients have also increased, posing potential health risks to some consumers, especially allergic reactions. The issue of allergenic cosmetic ingredients, as one of the key factors affecting consumer experience and satisfaction, is receiving increasing attention both within and outside the industry.
[0003] Allergic reactions are essentially an abnormal immune response. When an individual comes into contact with certain substances (i.e., allergens), the immune system mistakenly perceives them as a threat, triggering a series of inflammatory responses. Mild cases may result in skin redness, swelling, itching, and stinging, while severe cases can even lead to systemic allergic reactions and endanger life. Fragrances, preservatives, certain active ingredients, and surfactants in cosmetics can all be culprits in triggering allergies.
[0004] Relevant regulatory authorities have also taken certain measures to reduce the frequency of allergic reactions. The REACH regulation introduced by the European Chemicals Agency (ECHA) requires all chemicals sold or used in the EU market to be registered, evaluated, authorized, and restricted to ensure the safe use of chemicals and protect human health and the environment. In 2009, the EU adopted its first Cosmetics Regulation (EC) 1223 / 2009, aiming to regulate the entire cosmetics market, eliminate trade barriers, and promote product circulation. The European Scientific Committee on Consumer Safety (SCCS) has published opinions on several cosmetic ingredient usage requirements, including the maximum permitted levels of ingredients such as citral, silver, aluminum compounds, and benzophenone-4 in cosmetics.
[0005] Furthermore, this research typically involves multidisciplinary collaboration, including chemistry, skin physiology, immunology, and bioinformatics. In these studies, scientists first establish a cosmetic ingredient database containing information on widely used cosmetic raw materials and their chemical structures. Hoffmann et al. expanded the Cosmetics Europe skin sensitization database using new substances and PPRA data, while Del Bufalo et al. used related databases and NAM for quantitative skin sensitization assessment. Subsequently, high-throughput screening techniques and machine learning algorithms are used to analyze the characteristics of known sensitizing ingredients to identify potential sensitizing structural patterns or "alarm structures." These patterns help to quickly screen new compounds or unknown ingredients for potential sensitization. Simultaneously, in vitro testing methods based on human skin cells are widely used, such as using skin immune-related cells like keratinocytes and Langerhans cells for toxicity testing and sensitization assessment. These tests can simulate the immune response of human skin to cosmetic ingredients, providing direct evidence for predicting sensitization. As early as 2007, Dr. Patricia Engasser et al. emphasized the importance of cosmetic safety assessment in all major global markets. In 2012, Goebel et al. specified the cosmetic safety assessment criteria to the potential for skin sensitization. Building on the "gold standard" test method of LLNA (mice local lymph node assay), COLIPA (European Cosmetic Association) funded a broad program for skin sensitization research, method development, and method evaluation, and helped coordinate the early evaluation of three test methods that are currently undergoing pre-validation.
[0006] The field of cosmetic analysis research still faces numerous challenges due to the discrepancies, diversity, and variability of relevant data. Identifying the relevant mechanistic steps and developing the optimal toolset for reliably predicting skin sensitization efficacy are crucial for future research. In 2015, Baba et al. compiled a sample dataset consisting of 261 structurally different permeabilities and 31 solvents. They then applied Support Vector Regression (SVR) and Random Forest (RF) with greedy stepwise descriptor selection to the model to predict compounds not yet synthesized, providing an attractive alternative for permeability experiments in drug and cosmetic candidate screening and optimizing topical skin permeability formulations. In contrast, in 2021, Wilm et al. used a machine learning model trained on biologically meaningful descriptors to predict the skin sensitization potential of small molecules, further enhancing the safety aspects of intelligent cosmetic analysis. The interpretability of existing models is limited by molecular fingerprints and other non-intuitive descriptors. This experiment, combined with Skin Doctor CP, developed different CP variants to predict LLNA results based on the Random Forest (RF) model. An important observation to make is that models based on MACCS keys clearly require more data than those based on predicting biological activities. Only when all available LLNA data are used can models based on MACCS keys be synchronized with those based on predicting biological activities. This highlights the relevance of the proposed approach to developing strategies to address the scarcity of measurement data in biology, pharmacology, and toxicology. Consequently, researchers have made greater efforts in recent years to improve data quality, which is a key focus of this paper. In response, this year, Virk et al. critically discussed how artificial intelligence (AI) and machine learning (ML) can assist in the development of key ingredient materials for cosmetics and formulation products (including surfactants, polymers, fragrances, preservatives, and hydrogels), reviewing relevant research over the past four years. Among these, Wang et al. combined XGBoost (Extreme Gradient Boosting), CNN (Convolutional Neural Network), and MLP (Multilayer Perceptron) with the molecular ACCess system (MACCS) chemical fingerprinting to predict the viscosity of polymer melts and polymer aqueous solutions, potentially applicable to the cosmetics field; Martí et al. combined molecular dynamics models and machine learning models to efficiently evaluate 546 polymers, significantly saving time and cost; Kan et al. developed a QSAR model using SMOTE (Synthetic Minority Oversampling) and a random forest algorithm, achieving 87.7% accuracy in independent testing of 452 chemicals, and so on. Sharma et al. obtained a comprehensive database of chemical allergens and non-allergens from the Immune Epitope Database (IEDB), and used the calculated descriptors to train allergy prediction models derived from different ML algorithms, finding that the model derived from the random forest method outperformed other models.The developed AI model was then applied to assess the sensitization potential of components in FDA-approved drugs, revealing that the drugs caused allergic symptoms. Im et al. collected 482 sensitization data points and developed an RF model, which, compared to SVM, QSAR, and linear models, facilitated rapid screening of chemical sensitization hazards.
[0007] In summary, while existing research has made some progress using high-throughput screening, machine learning, and in vitro testing, challenges remain, including data gaps and insufficient model interpretability. This means consumers still face the risk of cosmetic allergies. To improve this situation, it is necessary to improve datasets, develop more reliable and interpretable predictive tools, and explore the application of AI and ML in cosmetic ingredient screening to more accurately predict allergens and ensure consumer safety. Summary of the Invention
[0008] This invention presents an improved method for predicting the sensitization of cosmetic ingredients, aiming to enhance the detection accuracy of potential sensitizing substances. The original Localized Lymph Node Neuralgia (LLNA) dataset is expanded to include data from human, animal, and non-animal sources, providing a more comprehensive foundation for sensitization prediction. Building upon previous work by Fuadah et al., a universal consensus prediction framework is constructed using three molecular descriptors: Morgan fingerprint, MACCS Keys, and ModRed, combined with Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM). Furthermore, a more flexible prediction model is designed. First, multiple benchmark models are selected using the LazyPredict tool, and then these models are combined through an ensemble method. This invention improves the accuracy of the new model in predicting the sensitization of cosmetic ingredients through comprehensive analysis of multiple data sources and models, providing a reference for cosmetic safety.
[0009] A machine learning-based method for predicting the sensitization potential of cosmetic ingredients includes a cosmetic ingredient data extraction module and a cosmetic ingredient sensitization identification and prediction module. The acquired cosmetic ingredient-related dataset is input into the cosmetic ingredient data extraction module, processed, and then output to the cosmetic ingredient sensitization identification and prediction module, which predicts the sensitization potential of the cosmetic ingredients. The processing steps of the cosmetic ingredient data extraction module include integrating cosmetic ingredient sensitization-related data, reading SDF data as input, extracting ingredient molecular features, and obtaining clean ingredient data. The cosmetic ingredient sensitization identification and prediction module also includes a cosmetic ingredient sensitization prediction model construction module. The processing steps of the cosmetic ingredient sensitization prediction model construction module include extracting the input data from the cosmetic ingredient data extraction module and outputting it to a prediction improvement model, and then predicting the sensitization potential of the cosmetic ingredients based on the prediction improvement model. The processing steps of the cosmetic ingredient sensitization prediction model construction module include obtaining clean ingredient data, inputting molecular features, and constructing the sensitization prediction model.
[0010] Furthermore, the data acquisition for the cosmetic ingredient data extraction module includes collecting skin sensitization data from humans, animals, and non-animals, integrating and processing the constantly updated COSING database, the LLNA database used in comparative studies, and NICEATM used to assess the skin sensitization potential of several cosmetic ingredients, including the human predictive patch test database as part of improving the existing dataset.
[0011] Furthermore, the cosmetic ingredient sensitization data extraction module integrates the chemical substance columns in the COSING database with relevant data in the LLNA database to improve the sensitization dataset.
[0012] Furthermore, the cosmetic ingredient sensitization identification and prediction module predicts the sensitization of input data by combining Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM) into a consensus prediction framework. It selects appropriate machine learning methods through the LazyPredict method to establish single and combined prediction models.
[0013] Furthermore, an improved dedicated molecular characterization system for sensitization prediction was developed, employing Mordred for the calculation of three-dimensional molecular descriptors, thereby enabling dual-channel fingerprint operation using Morgan fingerprints and MACCS keys.
[0014] Furthermore, the SMILES representations of 10 molecules were read using the RDKit library and visualized as images. The integrated data underwent molecular data processing to generate various types of chemical descriptors, which were then standardized and missing values were filled. The Morgan fingerprint for each molecule was generated using the `GetMorganFingerprintAsBitVect()` function from the `AllChem` library in the RDKit library and stored in a new column `Morgan_Descriptors` in the `moldf` dataframe. The `GenMACCSKeys()` function was used to generate MACCS keys for each molecule and stored in a new column `MACCS_Descriptors` in the `moldf` dataframe. A `Calculator` object was created to compute the Mordred descriptors. Next, all columns in the dataframe containing floating-point numbers were selected, and missing values were filled using the KNN method. Finally, the filled data was converted to NumPy arrays. The features were standardized using `StandardScaler()`. The standardized data was then stored in a new column `Modred_Descriptor` in the `moldf` dataframe.
[0015] Furthermore, the KNN algorithm is used to fill in missing values in the molecular descriptor to ensure the continuity of the chemical space. Simultaneously, Z-score normalization is applied to the high-dimensional descriptor to address the issue of dimensional differences.
[0016] Furthermore, based on the improved combined model of Fuadah et al. in the comparative experiment, a three-level model selection system was developed: a three-level prediction system based on the LazyPredict automated model selection framework.
[0017] Level 1: Quickly evaluate the initial performance of 30+ basic models, and use LazyClassifier to quickly build a benchmark model as a reference for subsequent work;
[0018] Level 2: The top 3 models (LightGBM, XGBoost, Random Forest) are selected for in-depth optimization. The optimal hyperparameter combination is found through grid search, and the StratifiedShuffleSplit method in cross-validation is used to ensure that the proportion of classes in each group is the same, and the Kappa coefficient is used as the scoring criterion.
[0019] Level 3: Construct a weighted voting integration model.
[0020] Single-model optimization: A three-level model optimization system was adopted, selecting XGBoost, LightGBM, and Random Forest as the base models. A phased parameter optimization scheme was designed to address the characteristics of cosmetic ingredient data. The StratifiedShuffleSplit method was used for data partitioning, with five stratified random splits to ensure that the proportion of samples in each category remained consistent with the original data during each validation round. The max_depth was set to a four-level gradient of [3, 5, 7, 10] as the parameters for the tree structure, preventing overfitting while ensuring feature interaction; num_leaves was set to [31, 50, 100] to correspond to different complex molecular structures to control leaf nodes; the learning rate was set to three levels of [0.01, 0.1, 0.2] to balance convergence speed and accuracy; finally, the number of trees n_estimators was set to a gradient of [100, 200, 300] to balance computational cost and performance.
[0021] The core of the improved model is the design of a dynamic weighted voting mechanism to construct a dynamically weighted multi-model combination prediction method. Based on the improved model, multi-dimensional predictions are performed on a reserved 20% test set. First, feature descriptors are predicted, resulting in three prediction probabilities: `morgan_probs`, `maccs_probs`, and `modred_probs`. Then, physicochemical properties are predicted, testing the sensitization potential of over 50 molecular features, including surface properties and hydrogen bond characteristics. Finally, the prediction probabilities of different models under different descriptors are obtained. Details are as follows:
[0022] ● Differentiated weight allocation mechanism: Weights are configured with precision down to the single digit ([210,209,1]): The optimal weight ratio is determined by grid search based on the Kappa coefficient performance of each model on the validation set. LightGBM and XGBoost have similar weights (210:209), reflecting their similar predictive capabilities; the weight of Random Forest is set to 1, serving only as an auxiliary correction.
[0023] ● Probability Fusion Innovation: A dual-channel probability calibration method was developed, which first fuses the predicted probabilities of RF and XGB by taking the arithmetic mean. Then, SVM probability is introduced as the basis for boundary sample discrimination. When |mean_probs-0.5|<0.2, SVM verification is enabled for secondary correction.
[0024] ●Dynamic Feature Selection: The Morgan fingerprint (1024-bit), MACCS key (166-bit), and Mordred descriptor (1613-dimensional) are processed using single and combined models respectively with the help of the descriptor-fingerprint collaborative system.
[0025] The algorithm is optimized based on the quadratic weighted Cohen's Kappa score of the validation set performance, which is used as a metric to measure the classifier's performance, and the model instance with the best_estimator_ optimal parameters is obtained.
[0026] Furthermore, the model evaluation module;
[0027] (1) Prediction accuracy assessment: accuracy, AUC, precision, recall, F1 score, specificity.
[0028] (2) Assessment of clinically relevant indicators: CCR (correct classification rate), PPV (positive predictive value), NPV (negative predictive value). Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the overall structure of this method.
[0030] Figure 2 This is a distribution chart of non-sensitizing and sensitizing compounds in the Outcome column.
[0031] Figure 3 This is an example diagram of SMILES, the molecules of a cosmetic ingredient.
[0032] Figure 4 This is an example graph comparing the evaluation parameters for establishing the optimal model.
[0033] Figure 5 This is an evaluation graph of the prediction model by Fuadah et al.
[0034] Figure 6 This is the evaluation diagram of the improved prediction model of this invention. Detailed Implementation
[0035] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0036] like Figure 1As shown, the key structure of this invention consists of two parts. The first part is a cosmetic ingredient sensitization data extraction module. To compensate for data deficiencies, this module integrates relevant data from the chemical substance column of the COSING database with the LLNA database ("LLNA" usually refers to the Local Lymph Node Assay, a method used to assess the skin allergenicity of chemicals. This method is mainly used in the safety testing of cosmetics and chemicals to determine whether they will cause allergic reactions in humans), thus improving the sensitization dataset. The second part is a cosmetic ingredient sensitization identification and prediction module. This module mainly predicts the sensitization of the input data. Comparing the research of Fuadah et al., which combined Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM) into a consensus prediction framework, this module uses the LazyPredict method to select appropriate machine learning methods, such as Random Forest, XGBoost, and SVM, to establish single and combined prediction models. This fully leverages the fundamental supporting role of data integrity in the model, improves the quality of sensitization prediction, and provides valuable cosmetic selection advice for consumers, regulators, and others.
[0037] 1. Data Input Module
[0038] (1) Multi-source data integration: This invention constructs a dedicated data integration system for cosmetic ingredient sensitization data, and achieves data standardization through data processing technology. Specifically, it includes using RDKit's SDMolSupplier to parse the ingredient sensitization database from the input SDF file, automatically extracting the SMILES string and molecular objects; for datasets from different sources, after retaining key tags such as ingredient names, CAS numbers, and sensitization labels, heterogeneous data are fused.
[0039] This invention retains LLNA data from the comparative study (Fuadah et al., 2024) and incorporates DPRA, h-CLAT, Human, KeratinoSens, and LLNA data from Pred-Skin v.3.0 to form a new ensemble dataset, such as... Figure 2First, the `LoadSDF()` function from the Pandas Tools library was used to import DPRA, h-CLAT, Human, KeratinoSens, and LLNA data in SDF (a standard file format for storing chemical structure information) format. The `Mol` column serves as metadata records for the relevant molecules, used to check for NAN values. The `notnull()` function from the pandas library was used to determine the NAN value and remove rows containing NAN values. Next, the `concat()` function from the pandas library was used to merge and join the above data to form a new DataFrame. The `drop_duplicates()` function from the pandas DataFrame library was then used to check for and remove duplicate data rows, indicating no duplicates. This process ensures the integrity of the data in the cosmetic sensitization prediction model, avoiding the possibility of missing values reducing the overall quality of the dataset, thus facilitating subsequent analysis and improving the accuracy of model predictions.
[0040] (2) Molecular characterization engineering: Improved the dedicated molecular characterization system for sensitization prediction, and used Mordred to calculate three-dimensional molecular descriptors, thereby realizing dual-channel fingerprint operation of Morgan fingerprint and MACCS key.
[0041] like Figure 3 As shown, the SMILES representations of 10 molecules were read using the RDKit library and visualized as images. The integrated data underwent molecular data processing to generate various types of chemical descriptors (fingerprints and descriptors), which were then standardized and missing values were filled. The Morgan fingerprint for each molecule was generated using the `GetMorganFingerprintAsBitVect()` function from the `AllChem` library in the RDKit library and stored in a new column `Morgan_Descriptors` in the `moldf` dataframe. The `GenMACCSKeys()` function was used to generate MACCS keys for each molecule and stored in a new column `MACCS_Descriptors` in the `moldf` dataframe. A `Calculator` object was created to compute the Mordred descriptors. Next, all columns in the dataframe containing floating-point numbers were selected, and missing values were filled using the KNN method. Finally, the filled data was converted to a NumPy array. The features were standardized using `StandardScaler()`. The standardized data was stored in a new column `Modred_Descriptor` in the `moldf` dataframe.
[0042] (3) Innovative data preprocessing: For missing values in molecular descriptors, the KNN algorithm is used to fill in the missing values to ensure the continuity of the chemical space. At the same time, Z-score standardization is performed on high-dimensional descriptors to solve the problem of dimensional differences.
[0043] 2. Model Building Module
[0044] (1) Three-level model selection system: A three-level prediction system was developed based on the LazyPredict automated model selection framework.
[0045] Level 1: Quickly evaluate the initial performance of 30+ base models (including SVM, Random Forest, etc.). This method reduces the initial workload, allowing us to focus more on model selection and tuning. In this invention, LazyClassifier is used to quickly build a benchmark model as a reference for subsequent work.
[0046] Level 2: The top 3 models (LightGBM, XGBoost, Random Forest) are selected for in-depth optimization. The optimal hyperparameter combination is found through grid search (GridSearchCV). The StratifiedShuffleSplit method in cross-validation is used to ensure that the proportion of classes in each group is the same, and the Kappa coefficient is used as the scoring criterion.
[0047] Level 3: Construct a weighted voting integration model.
[0048] (2) Single model optimization: This invention adopts a three-level model optimization system and selects three high-performance algorithms, namely XGBoost (XGBClassifier), LightGBM (LGBMClassifier) and RandomForestClassifier, as the basic model.
[0049] To address the characteristics of cosmetic ingredient data, a phased parameter optimization scheme was designed. The StratifiedShuffleSplit method was employed for data partitioning, with five stratified random splits (test_size = 0.2) to ensure that the proportion of samples in each category remained consistent with the original data during each validation round. This method effectively solved the problem of imbalance between sensitized and non-sensitized samples in cosmetic ingredient data. For the parameter grid design, we set max_depth to [3, 5, 7, 10].
[0050] The four-level gradient is used as a parameter for the tree structure to prevent overfitting and ensure feature interaction. At the same time, num_leaves is set to [31, 50, 100] to correspond to different complex molecular structures to control the leaf nodes. The learning rate is set to three levels [0.01, 0.1, 0.2] to balance convergence speed and accuracy. Finally, the number of trees n_estimators is set to a gradient of [100, 200, 300] to balance computational cost and performance.
[0051] (3) Combinatorial Model Analysis: A dynamic weighted voting mechanism was designed to construct a dynamically weighted multi-model combination prediction method. Based on the improved model, multi-dimensional predictions were performed on the reserved 20% test set. First, feature descriptors were predicted to obtain three prediction probabilities: morgan_probs, maccs_probs, and modified_probs. Then, physicochemical properties were predicted, and the sensitization potential of more than 50 molecular features, including surface properties and hydrogen bond characteristics, was tested. Finally, the prediction probabilities of different models under different descriptors were obtained. Specific improvements are as follows:
[0052] Differentiated weight allocation mechanism: Breaking through the limitations of traditional equal-weight voting, a weight configuration accurate to the single digit is adopted ([210,209,1]): Based on the Kappa coefficient performance of each model on the validation set, the optimal weight ratio is determined through grid search. LightGBM and XGBoost obtain similar weights (210:209), reflecting their similar predictive capabilities; the weight of Random Forest is set to 1, serving only as an auxiliary correction.
[0053] Probabilistic fusion innovation: A dual-channel probability calibration method was developed, which first fuses the predicted probabilities of RF and XGB by taking the arithmetic mean. Then, SVM probability is introduced as the basis for boundary sample discrimination. When |mean_probs-0.5|<0.2, SVM verification is enabled for secondary correction.
[0054] Dynamic feature selection: Morgan fingerprint (1024-bit), MACCS key (166-bit), and Mordred descriptor (1613-dimensional) are processed using single and combined models respectively with the help of the descriptor-fingerprint collaborative system.
[0055] Finally, the algorithm is optimized based on the quadratic weighted Cohen's Kappa score derived from the validation set performance. This metric measures the classifier's performance, resulting in a model instance with the best_estimator_optimal parameters, such as... Figure 4 .
[0056] Model evaluation module;
[0057] (1) Prediction accuracy assessment: accuracy, AUC, precision, recall, F1 score, specificity.
[0058] (2) Assessment of clinically relevant indicators: CCR (correct classification rate), PPV (positive predictive value), NPV (negative predictive value).
[0059] Application examples of the present invention
[0060] The dataset acquisition for the cosmetic ingredient data extraction module includes collecting skin sensitization data from humans, animals, and non-animals [25-26] (without duplication), integrating and processing the continuously updated COSING database, the LLNA database used in comparative studies, NICEATM (http: / / ntp.niehs.nih.gov / go / 40500, NICEATM LLNA database), etc. Pred-Skin v.3.0 (2020) was referenced to evaluate the skin sensitization potential of several cosmetic ingredients, including the human predictive patch assay database, as part of supplementing the existing dataset.
[0061] The database URL is:
[0062] https: / / ntp.niehs.nih.gov / whatwestudy / niceatm / test-method-evaluations / skin-sens / hppt
[0063] For model development, the mixtures from the above experiments must be included in the learning process, even though some ingredients are rarely used in cosmetics, they are actually listed as cosmetic ingredients in the COSING database. For data on chemical structures not found in the COSING database, ChemSpider was used with the Chemical Abstracts Service (CAS) registry number and chemical name.
[0064] Chemical structures were retrieved from the databases http: / / www.chemspider.com / or SciFinder (https: / / scifinder.cas.org). The data was cleaned and preprocessed to ensure accuracy and completeness.
[0065] Import molecular properties from rdkit_cdk for testing sensitization potential. Calculate the average of all model predicted probabilities and convert it to the final prediction result based on a threshold of 0.5, such as... Figure 5 Analyze multi-model fusion, such as Figure 6The confusion matrix and various evaluation metrics provide detailed model performance information, validating the improved combined model framework. The results are significantly better than the predictions of single models, which helps reduce the bias of individual models and improves the robustness and generalization ability of the model. High accuracy and high AUC scores indicate that the improved model has good classification ability. Sensitivity and specificity reflect the model's performance in identifying positive and negative samples, respectively. High sensitivity means the model can capture true positive examples well, while high specificity means the model can distinguish negative examples well. Metrics such as CCR, PPV, and NPV further refine the model's performance in different scenarios, helping to comprehensively understand the model's strengths and weaknesses. The AUC of this invention can be maintained at 0.86-0.91, demonstrating more stable and improved model performance compared to the original experiment.
[0066] Furthermore, we can consider applying the results of this invention to toxicity prediction in other fields, such as drug toxicity and environmental chemical toxicity prediction, thereby providing more support for the development of related fields. In summary, this study provides new ideas and methods for predicting the sensitization of cosmetic ingredients, reducing the risk of sensitization for consumers when choosing cosmetics, and has important theoretical significance and application value.
Claims
1. A machine learning-based method for predicting the sensitization effects of cosmetic ingredients, characterized in that, The method includes a cosmetic ingredient data extraction module and a cosmetic ingredient sensitization identification and prediction module. The acquired cosmetic ingredient-related dataset is input into the cosmetic ingredient data extraction module, processed, and then output to the cosmetic ingredient sensitization identification and prediction module, which predicts the sensitization potential of the cosmetic ingredients. The processing steps of the cosmetic ingredient data extraction module include integrating cosmetic ingredient sensitization-related data, reading and inputting SDF data, extracting ingredient molecular features, and collecting clean ingredient data. The cosmetic ingredient sensitization identification and prediction module also includes a cosmetic ingredient sensitization prediction model construction module. The processing steps of this module include extracting the input data from the cosmetic ingredient data extraction module and outputting it to a prediction improvement model, based on which the prediction improvement model predicts the sensitization potential of the cosmetic ingredients. The processing steps of the cosmetic ingredient sensitization prediction model construction module include collecting clean ingredient data, inputting molecular features, and constructing the sensitization prediction model.
2. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 1, characterized in that, The data acquisition for the cosmetic ingredient data extraction module includes collecting skin sensitization data from humans, animals, and non-animals, integrating and processing the constantly updated COSING database, the LLNA database used in comparative studies, and NICEATM used to assess the skin sensitization potential of several cosmetic ingredients, including the human predictive patch test database as part of improving the existing dataset.
3. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 1, characterized in that, The cosmetic ingredient sensitization data extraction module integrates the chemical substance columns in the COSING database with relevant data in the LLNA database to improve the sensitization dataset.
4. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 1, characterized in that, The cosmetic ingredient sensitization identification and prediction module predicts the sensitization of input data by combining Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM) into a consensus prediction framework. It selects appropriate machine learning methods through the LazyPredict method to establish single and combined prediction models.
5. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 1, characterized in that, An improved molecular characterization system for sensitization prediction was developed, using Mordred to calculate three-dimensional molecular descriptors, thereby enabling dual-channel fingerprint operation using Morgan fingerprints and MACCS keys.
6. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 5, characterized in that, The process involves using the RDKit library to read the SMILES representations of 10 molecules and visualizing them as images. The integrated data is then processed to generate various types of chemical descriptors, which are then standardized and missing values are filled. A Morgan fingerprint for each molecule is generated using the `GetMorganFingerprintAsBitVect()` function from the `AllChem` library in the RDKit library and stored in a new column `Morgan_Descriptors` in the `moldf` dataframe. MACCS keys for each molecule are generated using the `GenMACCSKeys()` function and stored in the same column. A `Calculator` object is created to compute the Mordred descriptors. Next, all columns in the dataframe containing floating-point numbers are selected, and missing values are filled using the KNN method. Finally, the filled data is converted to a NumPy array. Features are standardized using `StandardScaler()`. The standardized data is then stored in a new column `Modred_Descriptor` in the `moldf` dataframe.
7. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 6, characterized in that, For missing values in molecular descriptors, the KNN algorithm is used to fill them in, ensuring the continuity of the chemical space; at the same time, Z-score standardization is applied to high-dimensional descriptors to solve the problem of dimensional differences.
8. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 7, characterized in that, Develop a three-level prediction system based on the LazyPredict automated model selection framework; Level 1: Quickly evaluate the initial performance of 30+ basic models, and use LazyClassifier to quickly build a benchmark model as a reference for subsequent work; Level 2: The top 3 models are selected for in-depth optimization. The optimal combination of hyperparameters is found through grid search, and the StratifiedShuffleSplit method in cross-validation is used to ensure that the proportion of classes in each group is the same, and the Kappa coefficient is used as the scoring standard. Level 3: Constructing a weighted voting integration model; Single-model optimization: A three-level model optimization system was adopted, with XGBoost, LightGBM, and Random Forest selected as the three high-performance algorithms as the base models. A phased parameter optimization scheme was designed to address the characteristics of cosmetic ingredient data. The StratifiedShuffleSplit method was used for data partitioning, with five stratified random splits to ensure that the proportion of samples in each category remained consistent with the original data during each round of validation. The max_depth was set to a four-level gradient of [3, 5, 7, 10] as the parameter for the tree structure, preventing overfitting while ensuring feature interaction. Simultaneously, num_leaves was set to [31, 50, 100] to correspond to different complex molecular structures and control the leaf nodes. Three learning rates of [0.01, 0.1, 0.2] were used to balance convergence speed and accuracy. Finally, the number of trees, n_estimators, was set to a gradient of [100, 200, 300] to balance computational cost and performance.
9. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 1, characterized in that, A dynamic weighted voting mechanism was designed to construct a dynamically weighted multi-model combination prediction method. Based on the improved model, multi-dimensional predictions were performed on the reserved 20% test set. First, feature descriptors were predicted to obtain three prediction probabilities: morgan_probs, maccs_probs, and modified_probs. Then, physicochemical properties were predicted, including surface properties and hydrogen bond features, to test the sensitization potential. Finally, the prediction probabilities of different models under different descriptors were obtained.
10. The method for predicting the sensitization of cosmetic ingredients based on machine learning according to claim 9, characterized in that, Differentiated weight allocation mechanism: Weight configuration with precision down to the single digit is adopted: Based on the Kappa coefficient performance of each model on the validation set, the optimal weight ratio is determined by grid search; LightGBM and XGBoost are given similar weights to reflect their similar predictive capabilities; Random Forest is set to 1 as an auxiliary correction only. Probabilistic fusion innovation: A dual-channel probability calibration method was developed, which first fuses the predicted probabilities of RF and XGB by taking the arithmetic mean. Then, SVM probability is introduced as the basis for boundary sample discrimination. When |mean_probs-0.5|<0.2, SVM verification is enabled for secondary correction. Dynamic feature selection: Morgan fingerprint, MACCS key, and Mordred descriptor are processed using single and combined models respectively with the help of a descriptor-fingerprint collaborative system; The algorithm is optimized based on the quadratic weighted Cohen's Kappa score of the validation set performance, which is used as a metric to measure the classifier's performance, and the model instance with the best_estimator_ optimal parameters is obtained.