Knowledge base-based chemical substance toxicity prediction service provision method
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GAZZILABS INC
- Filing Date
- 2025-11-29
- Publication Date
- 2026-08-06
Smart Images

Figure KR2025020201_06082026_PF_FP_ABST
Abstract
Description
Method for providing a knowledge base-based chemical toxicity prediction service
[0001] The present invention relates to a method for providing a knowledge base-based chemical toxicity prediction service, and more specifically, to a knowledge base-based chemical toxicity prediction service capable of predicting the risk of chemical substances.
[0002] Chemicals are essential elements of science and technology that provide convenience to modern people. Various new chemical substances are being rapidly developed and easily entering human life in the form of household chemical products, while the use of existing chemical substances is also continuously increasing. However, compared to the speed of demand and supply of chemical substances, there is a severe shortage of data on the toxicity and risk assessment of currently available chemical substances. Furthermore, since toxicity assessments require significant budget and time, it is practically difficult to use chemical substances safely.
[0003] Laboratory animals are used in various fields, such as medicine, pharmaceuticals, and food, to evaluate the risks and hazards of chemical substances themselves and products containing them. However, due to the recent spread of awareness regarding animal bioethics, high costs, and animal welfare, animal testing is currently being banned or restricted.
[0004] (Patent Document 1) KR10-2012-0109461 A
[0005] The problem that the present invention aims to solve is to provide a method for providing a knowledge base-based chemical toxicity prediction service capable of predicting toxic effects in vivo based on integrated data and a knowledge base.
[0006] The problems that the present invention aims to solve are not limited to those mentioned above, and other unmentioned problems will be clearly understood by a person skilled in the art from the description below.
[0007] A method for providing a knowledge base-based chemical toxicity prediction service according to one embodiment of the present invention may include: a process of collecting first chemical data related to a target chemical from a first chemical database; a process of collecting second chemical data related to the target chemical from a second chemical database different from the first chemical database; a process of collecting specialized knowledge related to the target chemical from a knowledge base storing chemical data; a process of generating fused data by fusing the first chemical data and the second chemical data; and a process of fusing the fused data and the specialized knowledge.
[0008] A method for providing a knowledge base-based chemical toxicity prediction service according to one embodiment of the present invention for solving the above-mentioned problem can be performed by a computer program stored on a recording medium readable by a computer, combined with a computer which is hardware.
[0009] Other specific details of the present invention are included in the detailed description and drawings.
[0010] According to the present invention, the toxicity of a chemical substance can be predicted using a heterogeneous algorithm without directly performing tests on the chemical substance. Furthermore, the reliability of the toxicity prediction results can be enhanced by providing scientific grounds for the toxicity prediction results.
[0011] Non-testing methods such as QSAR can help reduce the need for animal testing and alleviate ethical concerns. For example, under the EU's REACH law, companies are required to demonstrate the safety of chemicals without using animal testing whenever possible, and alternatives like QSAR can meet these requirements. By applying advancements in AI technology to the field of chemical management, predictions can be made faster and more cheaply than animal testing. Furthermore, if alternative methods like QSAR can specifically design models based on human physiology in the future, the results could be more relevant to the human body. Non-testing methods like QSAR are more cost-effective, and technological advancements enable the provision of more accurate results.
[0012] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below.
[0013] FIG. 1 is a block diagram conceptually showing a knowledge base-based chemical toxicity prediction service providing device according to an embodiment of the present invention.
[0014] FIG. 2 is a diagram showing, in sequence, a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention.
[0015] FIG. 3 is a diagram conceptually illustrating a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention.
[0016] FIG. 4 is a diagram showing the structure of an artificial intelligence model used in a knowledge base-based chemical toxicity prediction service provision method according to an embodiment of the present invention.
[0017] FIG. 5 is a diagram exemplarily showing the process of training artificial intelligence in a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention.
[0018] FIG. 6 is a diagram exemplarily showing the results of searching for toxicity labels of chemical substances using a knowledge base-based chemical substance toxicity prediction service provision method according to an embodiment of the present invention.
[0019] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the present invention, and the present invention is defined only by the scope of the claims.
[0020] The terms used in this specification are for describing embodiments and are not intended to limit the invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. The terms "comprises" and / or "comprising" used in this specification do not exclude the presence or addition of one or more other components in addition to the components mentioned. Throughout the specification, the same reference numerals refer to the same components, and "and / or" includes each of the mentioned components and all combinations of one or more. Although terms such as "first," "second," etc., are used to describe various components, these components are not limited by these terms. These terms are used merely to distinguish one component from another. Therefore, the first component mentioned below may be the second component within the technical scope of the invention.
[0021] Unless otherwise defined, all terms used herein (including technical and scientific terms) may be used in a meaning commonly understood by those skilled in the art to which the present invention pertains. Additionally, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.
[0022]
[0023] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0024] FIG. 1 is a block diagram conceptually showing a knowledge base-based chemical toxicity prediction service providing device according to an embodiment of the present invention.
[0025] Referring to FIG. 1, a knowledge base-based chemical toxicity prediction service providing device according to an embodiment of the present invention may include a data collection unit (110), a database unit (120), a preprocessing unit (130), and a data processing unit (140). Here, the components of the chemical toxicity prediction device according to an embodiment of the present invention are not limited in this way, and components may be added or omitted through design changes by the manager of the chemical toxicity prediction device; however, the present invention describes the chemical toxicity prediction device focusing on the components shown in FIG. 1.
[0026] The data collection unit (110) can access a chemical substance database (not shown) to receive chemical substance data containing the names of chemical substances from the chemical substance database and store it in the database unit (120). Here, the chemical substance database is not limited to any database that the data collection unit (110) can access to receive (or acquire) chemical substance data; examples include a paper server storing various papers that is a database of activities regarding chemical molecules and biology papers, an experiment report server storing experiment reports, data servers of domestic and foreign institutions, and Pubchempy. In addition, the chemical substance database may include various web pages or servers containing chemical substance data.
[0027] The data collection unit (110) can collect chemical data from various chemical databases through searching and web crawling. The data collection unit (110) may include a Gazzirawler, which is a specialized chemical collection tool for collecting basic information on chemical substances and test data of substances. The data collection unit (110) can perform additional data collection by utilizing Pubchempy or web crawling to supplement the basic information on substances in the data, and can additionally collect essential information such as CID and Smiles by designating them for connectivity between heterogeneous DBs and for data consistency and usability.
[0028] The database unit (120) can be configured to store, process, and search for data for the training and operation of an artificial intelligence model. For example, the database unit (120) may include MongoDB. MongoDB is a document-oriented database that uses a BSON format similar to JSON, and since its schema is not fixed, the data structure can be dynamically changed as needed. In addition, MongoDB enables fast and flexible operations because there is no need to modify the existing structure when adding or changing data with different data formats and structures. MongoDB can efficiently process large amounts of data by distributing data across multiple servers through horizontal scaling (Sharding) capabilities.
[0029] The preprocessing unit (130) can preprocess the collected chemical data to build integrated data using the chemical data collected by the data processing unit, thereby generating or building source data for building a GSEC base. The preprocessing unit (130) can process the collected chemical data through a series of processes such as purification, inspection, and labeling, and transmit it to the data processing unit (140). At this time, the preprocessing unit (130) can automatically select data that does not conform to purification guidelines written based on expert standards in the chemical field, and generate or build source data that includes only data for which quality inspection has been completed after purification.
[0030] The data collection unit (110) and the preprocessing unit (130) may use a tool capable of systematically managing a series of processes from data collection, refinement, inspection, and labeling. For example, the tool may include LabelStudio, an open-source data labeling tool.
[0031] The data processing unit (140) can construct integrated data for tests / experiments on similar chemical substances using preprocessed data, such as source data, and generate integrated data in a form that presents scientific grounds for the prediction results. Additionally, since the collected chemical substance data may yield different results depending on the chemical substance, the data processing unit (140) can generate and construct integrated data by integrating chemical substance data collected from multi-type and heterogeneous databases that allow for objective judgment regarding the prediction. The data processing unit (140) may include a Large Language Model (LLM) capable of automating and extending the inference of the integrated data.
[0032] The data processing unit (140) can present scientific grounds through inference using a knowledge base with a graph structure. At this time, the data processing unit (140) can provide explainability of the inference results through Symbolic AI, which is robust to traditional inference and advantageous for processing unstructured data, and Symbolic AI, which is robust to traditional inference. That is, the data processing unit (140) can be configured by integrating heterogeneous algorithms that provide scientific grounds based on the prediction of skin toxicity. The data processing unit (140) can predict skin toxicity using the deep learning algorithm HGP-SL and search for similar substances using the RDKit library, and present scientific grounds through inference using a knowledge base with a graph structure.
[0033] Such a prediction device can be implemented by hardware circuits (e.g., CMOS-based logic circuits), firmware, software, or a combination thereof. For example, it can be implemented using transistors, logic gates, and electronic circuits in the form of various electrical structures.
[0034] Hereinafter, a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention will be described.
[0035] FIG. 2 is a diagram showing, in sequence, a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention, FIG. 3 is a diagram conceptually showing a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention, FIG. 4 is a diagram showing the structure of an artificial intelligence model used in a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention, and FIG. 5 is a diagram exemplarily showing the process of training an artificial intelligence in a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention.
[0036] Referring to FIG. 2, a method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention may include: a process of collecting first chemical data related to a target chemical from a first chemical database; a process of collecting second chemical data related to the target chemical from a second chemical database different from the first chemical database; a process of collecting specialized knowledge related to the target chemical from a knowledge base storing chemical data; a process of generating fused data by fusing the first chemical data and the second chemical data; and a process of fusing the fused data and the specialized knowledge.
[0037] Table 1 below provides examples of chemical data collected from various chemical databases (Tox21, eChemPortal, OECD QSAR TOOLBOX, PuBChem_HSDB). Chemical data may include various information such as chemical information, biological levels of chemicals, test species or strains, exposure concentrations, exposure methods, test types, test results, and sources.
[0038] Data Classification Tox21 eChemPortal OECD QSAR TOOLBOX PubChem_HSDB Chemical Information SMILES / SAMPLE_NAME / PUBCHEM_CID / CASPubChemCID / SubstanceName / Nametype / Number CAS Number / SMILESTEXT RAW DATA Biological Level Protein Species Test type Test Species or Strain PROTOCOL_NAME Species Strain / Testorganisms(species) Exposure Concentration CONC Executive summary Exposure Method Executive summary S9 Mix Status Executive summary Metabolic activation Test Type Unknown Conclusions keyword Test type Test Method Guideline GLP Status GLP Compliance SCI / Agency Report Status Participant / Rationaleforreliabilityincl.deficiencies Test Results Unknown Conclusions keyword Test type Source TOX21(NIH) Title / ReferenceTypeYear / ExecutivesummaryYear / Author / Referencesource
[0039] The method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention can generate fused data by fusing heterogeneous chemical data collected from heterogeneous chemical databases, as exemplified in Table 1, and construct integrated data by combining the generated fused data with specialized knowledge. Furthermore, the method for providing a knowledge base-based chemical toxicity prediction service according to an embodiment of the present invention can construct a sustainable integration pipeline for the flexible expansion of the previously collected chemical database and additionally updated data from research institutions, papers, and international organizations in accordance with the normalization standards for heterogeneous chemical data. The process of collecting the first chemical data and the process of collecting the second chemical data can collect chemical data from different chemical databases as described in Table 1. At this time, the process of collecting at least the second chemical data may be performed by crawling web pages.
[0040] The process of fusing the first chemical data and the second chemical data may include a process of preprocessing the collected chemical data and a process of fusing the preprocessed chemical data to generate fused data.
[0041] The pretreatment process may include purifying, inspecting, and labeling the chemical substances according to their physicochemical properties, classifying them, and storing them as source data in the database section (120).
[0042] In addition, the preprocessing process can generate additional data, such as physicochemical properties, that complement the basic information about the target chemical through heterogeneous chemical data.
[0043] The process of generating fused data by fusing the first chemical substance data and the second chemical substance data may include the process of constructing data on the scientific basis of skin toxicity by performing unstructured data test analysis and chemical substance correlation analysis through recursive text segmentation using LLM, and storing or embedding the result data in the database unit (120).
[0044] Furthermore, integrated data can be generated by combining the created fused data with retrieved expertise. Expertise serves as data for scientifically verifying the fused data and can play a role in identifying the sources of the integrated data.
[0045] Meanwhile, Figure 4 illustrates the structure of HGP-SL (Hierarchical Graph Pooling with Structure Learning), a model capable of inputting the structural characteristics of chemical substances. HGP-SL is a hierarchical graph pooling operator utilizing structure learning, which can generate a hierarchical representation of a graph by combining graph pooling and structure learning into an integrated module. HGP-SL is composed of a neural network that adds an HGP-SL layer for extracting structural information to a Graph Convolutional Network (GCN) and Multi-Layer Perceptron (MLP) layer that collects information from nodes and edges within the graph.
[0046] HGP-SL is a model released in 2019, but it is a SOTA model that still ranks first in the Graph Classification on PROTEINS field with an accuracy of 84.91.
[0047] In order to utilize not only the structural characteristics of chemical substances but also the chemical characteristics of molecules as features, the chemical informatics library RDKit can be used to calculate characteristics such as molecular weight and LogP and use them as training data.
[0048] The training dataset can be converted into graph data using the PyG library. Table 3 below provides an example of the execution environment and training and validation conditions for the artificial intelligence model.
[0049] Execution Environment CPU 12th Gen Intel(R) Core(TM) i9-12900KS X 24 Memory 125.6 GB BG PU NVIDIA RTX A6000 Storage 10.0 GB BOSU Ubuntu 20.04 Training and Validation Conditions Development Language Python 3.8 Frames CUDA 11.6, PyTorch 1.13.1 Training Algorithm HGP-SL Training Conditions batch_size = 512 max_epoch = 500 lr = 0.001 weight_decay = 0.001 hidden_size = 128 pooling_ratio = 0.5 dropout_ratio = 0.0 early_stop = 50
[0050] FIG. 5 is a diagram exemplarily showing the results of searching for toxicity labels of chemical substances using a knowledge base-based chemical substance toxicity prediction service provision method according to an embodiment of the present invention, where FIG. 5 (a) is a model training log screen, FIG. 6 (b) is a model training graph, and FIG. 5 (c) is a model test log screen.
[0051] The AI model performs validation and logs at every epoch, and the performance can be verified by testing the AI model with a test dataset.
[0052] Table 3 below shows the test results as an example.
[0053] Endpoint Goal Performance Result Performance Skin Hypersensitivity 60 or higher (F1-score) 67.94 Skin Irritation 74.95 Skin Corrosiveness 63.72 Skin Irritation / Corrosiveness 68.18
[0054] FIG. 6 is a diagram exemplarily showing the results of searching for toxicity labels of chemical substances using a knowledge base-based chemical substance toxicity prediction service provision method according to an embodiment of the present invention, and can explore graphs using inference verification and data visualization using Neo4j, substance quantity information, and key information visualization queries.
[0055] The steps of the method or algorithm described in connection with embodiments of the present invention may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), Flash Memory, a hard disk, a removable disk, a CD-ROM, or any form of computer-readable recording medium well known in the art to which the present invention belongs.
[0056]
[0057] Although embodiments of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention may be implemented in other specific forms without altering its technical concept or essential features. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive.
Claims
1. As a method for providing a knowledge base-based chemical toxicity prediction service, The process of collecting first chemical data related to the target chemical from the first chemical database; A process of collecting second chemical data related to the target chemical from a second chemical database different from the first chemical database; A process of collecting specialized knowledge related to the target chemical substance from a knowledge base storing chemical substance data; A process of generating fused data by fusing the first chemical data and the second chemical data; and A method for providing a knowledge base-based chemical toxicity prediction service comprising a process of fusing the above-mentioned fusion data and the above-mentioned expertise.