An acute toxicity prediction method based on data quality-aware machine learning
Patent Information
- Application Number
- CN202611033323.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,现有毒性预测方法在实际应用中仍面临若干关键挑战
本发明通过多源数据整合与规则化清洗,有效扩大了毒性数据覆盖的化学空间,同时利用结构标准化与聚合去重消除了冗余记录对模型训练的干扰。采用来源标识进行整体式数据拆分,避免了同源数据在训练集与测试集间的交叉污染,确保了模型泛化性能评估的客观性。折外预测残差筛查机制进一步剔除了实验值偏离折外估计的异常样本,极大提升了训练标签的可靠性,从数据源头抑制了“脏数据”对模型拟合的负面影响。最终,模型仅基于高质量、高一致性数据集进行训练,在独立测试集上展现出更稳健、更准确的急性毒性预测能力,相比传统方法显著降低了误报与漏报风险,为化合物早期风险评估提供了可靠的智能决策工具。
Smart Images

Figure CN122822402A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of chemical safety evaluation, artificial intelligence toxicology and machine learning prediction technology, and in particular relates to an acute toxicity prediction method based on data quality-aware machine learning. Background Technology
[0002] Traditional acute toxicity assessments primarily rely on in vivo animal experiments. However, these experiments suffer from inherent limitations such as lengthy processing times, high costs, stringent ethical constraints, and limited throughput, making them insufficient to meet the growing demands for chemical safety assessments. With the rapid development of computational toxicology and artificial intelligence technologies, machine learning methods have been widely used for compound toxicity prediction. These methods transform molecular structures into numerical representations such as circular fingerprints, two-dimensional molecular descriptors, or graph neural networks, training statistical models to learn structure-activity relationships, thereby enabling the prediction of toxicity for unknown compounds. Compared to traditional animal experiments, computational prediction methods offer significant advantages such as low cost, high throughput, and coverage of a large chemical space, making them an important auxiliary tool for initial toxicity screening and hazard classification of chemicals.
[0003] However, existing toxicity prediction methods still face several key challenges in practical applications. First, experimental records in public toxicity databases are typically compiled from different literature sources, experimental institutions, and data processing procedures. This can lead to inconsistencies in structural representation of the same compound across different sources (e.g., inconsistent salt / solvent treatment), different toxicity units (e.g., mixing mg / kg and mmol / kg), conflicting labels from duplicate records (LD50 values of the same compound differing by tens of times), and systematic biases in the source data. Directly using such heterogeneous, low-quality data for model training will severely impair the model's generalization ability and predictive reliability. Second, existing methods often incorporate independently sourced test data or external validation data into model-assisted filtering, feature preprocessing fitting, or hyperparameter optimization during data cleaning and anomaly sample screening. This introduces the risk of information leakage, leading to a systematic overestimation of model performance and an inability to objectively reflect its true predictive performance on data from unknown sources. Furthermore, most prediction models only output a single toxicity value, lacking quantitative assessments of the reliability of the prediction results (such as applicability domains) and molecular-level interpretability analysis. This makes it difficult for users to evaluate the credibility of the prediction results and limits the practical application value of the models in chemical safety decision-making. In summary, there is an urgent need for an acute toxicity prediction method that can implement strict quality control in multi-source toxicity data, avoid information leakage, and simultaneously output prediction reliability assessments and interpretable results. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention proposes an acute toxicity prediction method based on data quality-aware machine learning, thereby resolving the issues present in the prior art.
[0005] To achieve the above objectives, this invention provides an acute toxicity prediction method based on data quality-aware machine learning, comprising: Acquire acute toxicity data of multi-source compounds, wherein the acute toxicity data of multi-source compounds includes compound structure, experimental toxicity value and source identification; The acute toxicity data of the multi-source compounds were cleaned according to rules and standardized in structure, and repeated records were aggregated according to compound identifiers to obtain standardized toxicity labels; Based on the source identifier, compounds from independent data sources are reserved as a test set, while the remaining compounds are used as a development data pool. The residual screening of each compound in the development data pool is performed using the out-of-bounds prediction method. Records exceeding the preset threshold are removed based on the deviation between the experimental value and the out-of-bounds prediction value to obtain the final development dataset. Based on the final development dataset, molecular structural features are extracted and a prediction model is trained. The molecular structure characteristics of the compound to be tested are input into the prediction model, and the predicted value of acute toxicity is output.
[0006] Optionally, the process of performing regular cleaning and structural standardization on the acute toxicity data of the multi-source compounds, and aggregating repeated records according to compound identifiers to obtain standardized toxicity labels includes: Based on the acute toxicity data of the multi-source compounds, the experimental toxicity values were uniformly converted to the same unit of measurement, and the effectiveness of the structure of each compound was verified to obtain the toxicity data after initial screening. Based on the toxicity data after the initial screening, the parent fragments of the multi-fragment compounds are extracted and subjected to charge neutralization to generate standardized molecular identifiers for each compound, thereby obtaining the structure-standardized toxicity data. Based on the standardized toxicity data, compounds with the same standardized molecular identifier are grouped into the same compound group. The statistical median of the experimental toxicity values within the group is used as the representative toxicity label for the compound, thus obtaining the standardized toxicity label for each compound.
[0007] Optionally, the process of using an out-of-bounds prediction method to perform residual screening on each compound in the development data pool, and removing records exceeding a preset threshold based on the deviation between experimental values and out-of-bounds prediction values, to obtain the final development dataset includes: The development data pool is divided into multi-fold data subsets; A temporary prediction model is trained based on the data subsets of each fold other than the current fold, and the temporary prediction model is used to predict each compound in the current fold to obtain the out-of-fold prediction value of each compound. The fold deviation is calculated based on the experimental toxicity value of each compound and the corresponding out-of-range predicted value. Compounds with fold deviations exceeding a preset threshold are removed from the development data pool to obtain the final development dataset.
[0008] Optionally, the process of calculating the multiple deviation based on the experimental toxicity value and the corresponding out-of-division predicted value of each compound includes: calculating the out-of-division multiple deviation based on the multiple relationship between the experimental toxicity value and the corresponding out-of-division predicted value of each compound, obtaining the deviation metric value of each compound, and comparing the deviation metric value with a preset threshold.
[0009] Optionally, the process of extracting molecular structural features and training a prediction model based on the final development dataset includes: Based on the standardized molecular structures of each compound in the final development dataset, molecular fingerprinting algorithm and molecular descriptor calculation method are used to extract the molecular structure features of each compound. Missing descriptors are filled with statistical parameters from the final development dataset and zero variance features are filtered to obtain the molecular structure feature vector of each compound. An ensemble learning algorithm was used to train multiple sub-models under random seeds based on the molecular structure feature vectors of each compound and their corresponding standardized toxicity labels. The prediction outputs of each sub-model were then averaged to obtain an acute toxicity prediction model.
[0010] Optionally, after inputting the molecular structure features of the compound to be tested into the prediction model, the model further includes an applicable domain determination process: Based on the molecular structure of the compound to be tested, the structural fingerprint of the compound to be tested is extracted using the same molecular fingerprint calculation method as that used for each compound in the final development dataset; The structural similarity is calculated based on the structural fingerprint of the compound to be tested and the structural fingerprint of each compound in the final development dataset, and the similarity is compared with a preset threshold. When the similarity between the test compound and at least one compound in the final development dataset reaches a preset threshold, the test compound is determined to be within the applicable domain, and the acute toxicity prediction value and applicable domain status are output based on the prediction model.
[0011] The present invention also provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described thereon.
[0012] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.
[0013] Compared with the prior art, the present invention has the following advantages and technical effects: This invention effectively expands the chemical space covered by toxicity data through multi-source data integration and rule-based cleaning. Simultaneously, it eliminates the interference of redundant records on model training by utilizing structure standardization and aggregation deduplication. A holistic data splitting approach using source identification avoids cross-contamination between the training and test sets, ensuring the objectivity of model generalization performance evaluation. An outlier prediction residual screening mechanism further eliminates anomalous samples where experimental values deviate from the outlier estimate, greatly improving the reliability of training labels and suppressing the negative impact of "dirty data" on model fitting from the data source. Ultimately, the model is trained only on high-quality, highly consistent datasets, demonstrating more robust and accurate acute toxicity prediction capabilities on independent test sets. Compared to traditional methods, it significantly reduces the risk of false positives and false negatives, providing a reliable intelligent decision-making tool for early risk assessment of compounds. Attached Figure Description
[0014] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of the acute toxicity prediction method based on data quality-aware machine learning according to an embodiment of the present invention; Figure 2 This is a flowchart of the OOF residual screening process according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the applicable domain determination process according to an embodiment of the present invention. Detailed Implementation
[0015] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0016] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0017] Example 1 like Figure 1As shown in the figure, this embodiment provides an acute toxicity prediction method based on data quality-aware machine learning, including the following steps: acquiring acute toxicity data of multi-source compounds, which includes compound structure, experimental toxicity value, and source identifier; performing rule cleaning and structural standardization on the acute toxicity data of multi-source compounds, and aggregating duplicate records according to compound identifier to obtain standardized toxicity labels; based on the source identifier, reserving compounds from independent data sources as a test set, and using the remaining compounds as a development data pool; using an out-of-place prediction method to perform residual screening on each compound in the development data pool, and removing records exceeding a preset threshold based on the deviation between the experimental value and the out-of-place prediction value to obtain the final development dataset; extracting molecular structure features based on the final development dataset and training a prediction model; inputting the molecular structure features of the compound to be tested into the prediction model and outputting the acute toxicity prediction value.
[0018] S1. Acute toxicity data of multi-source compounds are obtained. This data includes the compound structure, experimental toxicity value, and source identification. In this embodiment, acute toxicity records of compounds are obtained from the TOXRIC and PubChem databases. Records are selected based on the mouse species, the route of administration (intraperitoneal injection), and the toxicity endpoint (LD50). Initial records include the compound structure, experimental LD50 value and its unit, experimental species, route of administration, toxicity endpoint, and data source information.
[0019] S2 performs regular cleaning and structural standardization on acute toxicity data of multi-source compounds.
[0020] Furthermore, the process of performing rule-based cleaning and structural standardization on the acute toxicity data of multi-source compounds, and aggregating repeated records according to compound identifiers to obtain standardized toxicity labels includes: based on the acute toxicity data of multi-source compounds, converting each experimental toxicity value to the same unit of measurement, and validating the structure of each compound to obtain the initial screening toxicity data; based on the initial screening toxicity data, extracting the parent fragment from the multi-fragment compound structure and performing charge neutralization processing to generate standardized molecular identifiers for each compound, thus obtaining structurally standardized toxicity data; based on the structurally standardized toxicity data, grouping compound records with the same standardized molecular identifier into the same compound group, and using the statistical median of each experimental toxicity value within the group as the representative toxicity label for that compound, thus obtaining the standardized toxicity label corresponding to each compound.
[0021] The LD50 values were standardized to mg / kg, and the molecular weights of the parent compounds were converted to mol / kg. The regression target value was then defined as the negative logarithmic LD50. Records containing inequality signs, missing values, non-numerical values, non-positive values, invalid smiles, inorganic substances, metallic structures, complex mixtures, and extreme LD50 values were deleted. For salts and multi-fragment structures, the main organic parent fragment was retained, and charge neutralization was performed to generate normalized smiles.
[0022] S3, Perform compound-level polymerization. Group the records according to the normalized parent SMILES. For multiple experimental measurements of the same compound, if the ratio of the maximum to the minimum value exceeds 5, the compound is considered to have a serious label conflict and is removed; if it does not exceed 5, the median is used as the representative LD50 value of the compound.
[0023] In one specific embodiment, after rule cleaning, structure standardization, and compound-level polymerization, 36,027 standardized compounds were obtained, including 31,944 TOXRIC-only compounds, 3,015 compounds overlapping between the two databases, and 1,068 PubChem-only compounds.
[0024] S4. Source-level segmentation. Before model-assisted screening, 1068 PubChem-only compounds were reserved as the source-level test set; TOXRIC-only compounds and overlapping compounds from the two databases were combined to form a model development data pool of 34959 compounds. The PubChem-only source-level test set was not used for outlier residual screening, feature preprocessing fitting, model parameter selection, or final model training.
[0025] S5, perform external residual screening.
[0026] Furthermore, the process of using an out-of-flush prediction method to perform residual screening on each compound in the development data pool, and removing records exceeding a preset threshold based on the deviation between experimental values and out-of-flush predicted values, to obtain the final development dataset includes: dividing the development data pool into multi-fold data subsets; training a temporary prediction model based on the remaining data subsets outside the current fold, and using the temporary prediction model to predict each compound in the current fold to obtain the out-of-flush predicted value of each compound; calculating the fold deviation based on the experimental toxicity value of each compound and the corresponding out-of-flush predicted value, and removing compounds with fold deviations exceeding a preset threshold from the development data pool to obtain the final development dataset.
[0027] like Figure 2As shown, the model development data pool is divided into five folds. A temporary prediction model is trained using four folds each time, and predictions are made for the remaining fold, thus obtaining an out-of-fold predicted value for each development compound. The out-of-fold fold bias is calculated based on the difference between the experimental regression target value and the out-of-fold predicted regression target value. When the out-of-fold fold bias is greater than or equal to 20, the corresponding compound is marked as a potentially low-reliability record and removed from the model development data pool. In this embodiment, a total of 448 compounds were removed, accounting for 1.28% of the model development data pool, resulting in a final model development dataset consisting of 34,511 compounds.
[0028] S6, Molecular Characterization. ECFP4 counting fingerprints and RDKit 2D molecular descriptors are calculated for standardized molecular structures. The ECFP4 counting fingerprints, with a radius of 2 and a length of 2048, are used to characterize the local molecular environment and its frequency of occurrence. The RDKit 2D molecular descriptors are used to characterize molecular size, lipophilicity, partial charge distribution, molecular surface area, molecular connectivity, ring structures, and topological complexity. Logarithmic transformation is performed on the counting fingerprints, missing descriptors are imputed using the median of the model development dataset, and zero-variance features are removed.
[0029] S7, train the acute toxicity prediction model.
[0030] Furthermore, the process of extracting molecular structure features and training prediction models based on the final development dataset includes: extracting molecular structure features of each compound based on the standardized molecular structure of each compound in the final development dataset using molecular fingerprinting and molecular descriptor calculation methods; filling in missing descriptors using statistical parameters from the final development dataset and filtering zero-variance features to obtain molecular structure feature vectors for each compound; training multiple sub-models under random seeds using an ensemble learning algorithm based on the molecular structure feature vectors of each compound and their corresponding standardized toxicity labels; and averaging the prediction outputs of each sub-model to obtain an acute toxicity prediction model.
[0031] This embodiment uses the XGBoost regression model as the base model. First, the final model development dataset is divided into an internal training set and a validation set to determine the appropriate number of iteration rounds. Then, the model is trained using the entire final model development dataset. To reduce prediction fluctuations caused by random sampling and feature sampling, multiple XGBoost sub-models under random seeds are trained, and the prediction results of multiple sub-models are averaged to obtain the final prediction value.
[0032] S8. Model evaluation and application. The compound to be predicted is input into the final acute toxicity prediction model, which outputs a negative logarithmic LD50 predicted value, which can be inversely transformed into an LD50 predicted value at the mg / kg scale. In this embodiment, the model achieved predictive performance of R² 0.7440, RMSE 0.4703, and MAE 0.2445 on a strictly reserved PubChem-only source-level test set.
[0033] S9 performs the domain of application determination. For example... Figure 3 As shown, the Tanimoto similarity between the compound to be predicted and the compounds in the final model development dataset is calculated. When the Tanimoto similarity between the compound to be predicted and at least one model development compound is not less than 0.70, it is determined to be a compound within the applicable domain. Prediction results corresponding to compounds within the applicable domain have higher reliability. In this embodiment, the R² of compounds within the applicable domain is 0.9136, the RMSE is 0.2700, and the MAE is 0.1199.
[0034] S10, Model Interpretation Analysis. The SHAP method was used to calculate the contribution of different molecular features to the prediction results, and the influence of ECFP4 counting fingerprints and RDKit two-dimensional descriptors on the prediction results was evaluated. Key influencing factors such as charge-related surface area, molecular connectivity, lipophilicity, and structural complexity were identified.
[0035] Through the steps described above, this embodiment can construct an acute toxicity prediction model with high reliability and interpretability even when multi-source toxicity data exhibits structural inconsistencies, label conflicts, and source noise. This method can be used for initial screening of chemical acute toxicity, hazard prioritization, and chemical safety assessment scenarios requiring reliability judgment based on the applicable domain status.
[0036] Example 2 To address the issues of inconsistent quality of multi-source acute toxicity data, susceptibility of model training to outlier labels, and insufficient evaluation of source-level generalization ability in existing technologies, this invention provides an acute toxicity prediction method based on data quality-aware machine learning. This method improves the reliability of LD50 prediction for intraperitoneal injection in mice through rule cleaning, source-level retention, outlier residual screening, joint molecular characterization, machine learning ensemble modeling, applicability domain determination, and model interpretation analysis.
[0037] This embodiment provides an acute toxicity prediction method based on data quality-aware machine learning, including: acquiring acute toxicity data of compounds from multiple sources; performing rule cleaning, unit standardization, molecular structure standardization, and compound-level aggregation on the data; reserving compound data from independent sources as a source-level test set before model-assisted screening, and using the remaining data as a model development data pool; performing out-of-bounds residual screening only within the model development data pool to remove potentially low-reliability toxicity records; constructing molecular characterization based on the final model development dataset and training an acute toxicity prediction model; inputting the compound to be predicted into the prediction model to obtain the acute toxicity prediction results for intraperitoneal injection in mice.
[0038] Furthermore, acute toxicity data for multi-source compounds include the compound's molecular structure, experimental LD50 value, experimental species, route of administration, toxicity endpoint, and data source identifier. After endpoint screening, only records corresponding to mouse, intraperitoneal administration routes, and LD50 endpoints are retained.
[0039] Furthermore, the LD50 data were standardized to uniformly represent LD50 values of different units to mg / kg. The molecular weight of the standardized parent compound was then converted to mol / kg, and finally converted to a negative logarithmic regression target value.
[0040] Furthermore, the rule cleaning includes deleting records containing inequality signs, deleting missing values and non-positive values, deleting invalid molecular structures, deleting inorganic substances, metal-containing structures, complex mixtures and extreme LD50 values, and extracting parent fragments, neutralizing charges and generating normalized SMILES for salt and multi-fragment structures.
[0041] Furthermore, compound-level aggregation is performed on multiple records corresponding to the same normalized parent SMILES. When the difference between multiple experimental values of the same compound exceeds a preset conflict threshold, the compound is removed; when the difference does not exceed the preset conflict threshold, the median of multiple experimental values is used as a representative toxicity label.
[0042] Furthermore, source-level retention includes reserving PubChem-only compounds as a source-level test set before model-assisted screening, and ensuring that this test set does not participate in extrapolation residual screening, feature preprocessing fitting, model parameter selection, or final model training.
[0043] Furthermore, the out-of-compensation residual screening is performed using a five-fold out-of-compensation prediction. Each development compound is predicted only by a temporary model that has not seen the compound before, resulting in an out-of-compensation predicted value. The out-of-compensation fold bias is calculated based on the difference between the experimental regression target value and the out-of-compensation predicted regression target value. Compounds with an out-of-compensation fold bias greater than or equal to 20 are identified as potentially low-reliability records and removed from the model development data pool.
[0044] Furthermore, molecular characterization includes ECFP4 counting fingerprints and RDKit two-dimensional molecular descriptors. ECFP4 counting fingerprints are used to characterize local molecular environments and their frequency of occurrence, while RDKit two-dimensional descriptors are used to characterize molecular size, lipophilicity, charge distribution, surface area, topological connectivity, and structural complexity.
[0045] Furthermore, the acute toxicity prediction model is a multi-random-seed XGBoost ensemble model. After determining the model parameters and number of iterations through internal training-validation splits or cross-validation, multiple XGBoost sub-models under random seeds are trained using the final model development dataset, and the average of the predictions from multiple sub-models is used as the final prediction result.
[0046] Furthermore, the method also includes application domain determination. Application domain determination is based on the Tanimoto similarity between the compound to be predicted and the compounds in the model development dataset. When the similarity between the compound to be predicted and at least one model development compound reaches a preset threshold, it is determined to be a compound within the application domain, and the application domain status and toxicity prediction value are output together.
[0047] Furthermore, the method also includes model interpretation analysis. The contributions of molecular structural fragments and two-dimensional molecular descriptors to the prediction results are calculated using SHAP values, and key molecular features affecting acute toxicity prediction results are output, providing a reference for understanding the prediction results and evaluating chemical safety.
[0048] This invention performs source-level retention before model-assisted screening to avoid source-level test data participating in model-assisted filtering and preprocessing fitting, thereby reducing the risk of information leakage. This invention identifies potentially low-reliability records in multi-source public toxicity data through extrapolation residual screening, reducing the impact of anomalous labels on model training. This invention combines ECFP counting fingerprints with two-dimensional molecular descriptors to simultaneously characterize local structural fragments and overall physicochemical and topological properties. This invention improves model prediction stability through multi-random seed ensemble. This invention further outputs the applicable domain state and SHAP interpretation results, which helps identify more reliable predictions and improve model interpretability.
[0049] The present invention also provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described thereon.
[0050] The present invention also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.
[0051] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An acute toxicity prediction method based on data quality-aware machine learning, characterized in that, Includes the following steps: Acquire acute toxicity data of multi-source compounds, wherein the acute toxicity data of multi-source compounds includes compound structure, experimental toxicity value and source identification; The acute toxicity data of the multi-source compounds were cleaned according to rules and standardized in structure, and repeated records were aggregated according to compound identifiers to obtain standardized toxicity labels; Based on the source identifier, compounds from independent data sources are reserved as a test set, while the remaining compounds are used as a development data pool. The residual screening of each compound in the development data pool is performed using the out-of-bounds prediction method. Records exceeding the preset threshold are removed based on the deviation between the experimental value and the out-of-bounds prediction value to obtain the final development dataset. Based on the final development dataset, molecular structural features are extracted and a prediction model is trained. The molecular structure characteristics of the compound to be tested are input into the prediction model, and the predicted value of acute toxicity is output.
2. The acute toxicity prediction method based on data quality-aware machine learning according to claim 1, characterized in that, The process of performing regular cleaning and structural standardization on the acute toxicity data of the multi-source compounds, and aggregating repeated records according to compound identifiers to obtain standardized toxicity labels includes: Based on the acute toxicity data of the multi-source compounds, the experimental toxicity values were uniformly converted to the same unit of measurement, and the effectiveness of the structure of each compound was verified to obtain the toxicity data after initial screening. Based on the toxicity data after the initial screening, the parent fragments of the multi-fragment compounds are extracted and subjected to charge neutralization to generate standardized molecular identifiers for each compound, thereby obtaining the structure-standardized toxicity data. Based on the standardized toxicity data, compounds with the same standardized molecular identifier are grouped into the same compound group. The statistical median of the experimental toxicity values within the group is used as the representative toxicity label for the compound, thus obtaining the standardized toxicity label for each compound.
3. The acute toxicity prediction method based on data quality-aware machine learning according to claim 1, characterized in that, The process of using an out-of-bounds prediction method to perform residual screening on each compound in the development data pool, and removing records exceeding a preset threshold based on the deviation between experimental values and out-of-bounds prediction values to obtain the final development dataset includes: The development data pool is divided into multi-fold data subsets; A temporary prediction model is trained based on the data subsets of each fold other than the current fold, and the temporary prediction model is used to predict each compound in the current fold to obtain the out-of-fold prediction value of each compound. The fold deviation is calculated based on the experimental toxicity value of each compound and the corresponding out-of-range predicted value. Compounds with fold deviations exceeding a preset threshold are removed from the development data pool to obtain the final development dataset.
4. The acute toxicity prediction method based on data quality-aware machine learning according to claim 3, characterized in that, The process of calculating the multiple deviation based on the experimental toxicity value and the corresponding out-of-division predicted value of each compound includes: calculating the out-of-division multiple deviation based on the multiple relationship between the experimental toxicity value and the corresponding out-of-division predicted value of each compound, obtaining the deviation metric value of each compound, and comparing the deviation metric value with a preset threshold.
5. The acute toxicity prediction method based on data quality-aware machine learning according to claim 1, characterized in that, The process of extracting molecular structural features and training a prediction model based on the final development dataset includes: Based on the standardized molecular structures of each compound in the final development dataset, molecular fingerprinting algorithm and molecular descriptor calculation method are used to extract the molecular structure features of each compound. Missing descriptors are filled with statistical parameters from the final development dataset and zero variance features are filtered to obtain the molecular structure feature vector of each compound. An ensemble learning algorithm was used to train multiple sub-models under random seeds based on the molecular structure feature vectors of each compound and their corresponding standardized toxicity labels. The prediction outputs of each sub-model were then averaged to obtain an acute toxicity prediction model.
6. The acute toxicity prediction method based on data quality-aware machine learning according to claim 1, characterized in that, After inputting the molecular structure features of the compound to be tested into the prediction model, the process also includes a domain of application determination: Based on the molecular structure of the compound to be tested, the structural fingerprint of the compound to be tested is extracted using the same molecular fingerprint calculation method as that used for each compound in the final development dataset; The structural similarity is calculated based on the structural fingerprint of the compound to be tested and the structural fingerprint of each compound in the final development dataset, and the similarity is compared with a preset threshold. When the similarity between the test compound and at least one compound in the final development dataset reaches a preset threshold, the test compound is determined to be within the applicable domain, and the acute toxicity prediction value and applicable domain status are output based on the prediction model.
7. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in claim 1.
8. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in claim 1.