Methods for establishing a multi-species aquatic organism acute toxicity prediction model and toxicity prediction methods
By constructing an acute toxicity prediction model for multiple aquatic organisms using deep learning methods, this approach addresses the issues of high costs associated with traditional experiments and the limited applicability of existing models. It enables efficient and broad-based toxicity prediction for multiple aquatic organisms, making it suitable for chemical environmental health risk assessment and environmental evaluation.
Patent Information
- Application Number
- CN202411981843.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Traditional in vivo toxicology experiments are costly, time-consuming, and rely on a large number of experimental animals. Existing machine learning algorithms for predicting the toxicity of aquatic organisms only target a single species, have a narrow scope of application, and limited predictive performance. Furthermore, the scarcity and uneven distribution of aquatic organism toxicity data hinders comprehensive toxicity prediction.
A multi-species acute toxicity prediction model for aquatic organisms was established. By acquiring molecular graphs and species information of compounds, deep learning methods were used to optimize network parameters and construct a multi-species acute toxicity prediction model for aquatic organisms. This model includes Cytb common subsequences and species evolutionary information. Graph neural networks were used to extract species features, reducing the requirements for computing resources.
It enables efficient, wide-ranging, and convenient toxicity prediction for a variety of compounds and aquatic organisms, reduces computational costs, expands the prediction scope, and improves prediction performance, making it suitable for chemical environmental health risk assessment and environmental evaluation.
Smart Images

Figure CN119851806B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of environmental science and ecotoxicology, and more specifically, to a method for establishing an acute toxicity prediction model for multi-species aquatic organisms and a toxicity prediction method. Background Technology
[0002] In the fields of environmental science and ecotoxicology, the health status of aquatic ecosystems is widely considered an important indicator of environmental quality and pollution levels. Aquatic organisms, in particular, due to their high sensitivity to environmental changes and unique ecological adaptability, have become key biological indicators for monitoring environmental disturbances, especially ecological disturbances caused by industrial and agricultural chemicals. In ecotoxicological risk assessment, acute toxicity testing is a core indicator for evaluating the biological effects of chemicals, typically involving assessments of impacts on mortality, reproductive capacity, mobility, or growth status of specific species.
[0003] With the rapid growth of chemical product types, more than 350,000 chemicals are currently registered and in use globally, posing a potential threat to human health and environmental safety. A comprehensive and systematic assessment of these potential environmental pollutants is urgently needed. However, traditional in vivo toxicology experiments are not only costly and time-consuming, but also rely on a large number of laboratory animals. Summary of the Invention
[0004] In view of this, this disclosure provides a method for establishing an acute toxicity prediction model for multi-species aquatic organisms and a toxicity prediction method.
[0005] One aspect of this disclosure provides a method for establishing an acute toxicity prediction model for multiple aquatic organisms, comprising: acquiring actual acute toxicity data of multiple compounds on multiple aquatic organisms, molecular diagrams of the multiple compounds, and species information of the multiple aquatic organisms, wherein the species information includes Cytb common subsequences and / or species evolutionary information of the multiple aquatic organisms; inputting the molecular diagrams of the multiple compounds and the species information of the multiple aquatic organisms into an initial prediction model to obtain initial acute toxicity prediction data of the multiple compounds on the multiple aquatic organisms; and optimizing the network parameters of the initial prediction model based on the initial acute toxicity prediction data and the actual acute toxicity data to obtain a trained acute toxicity prediction model for multiple aquatic organisms.
[0006] Another aspect of this disclosure provides an acute toxicity prediction method, comprising: acquiring a molecular map of a compound to be predicted and species information of an organism to be predicted, wherein the species information includes the Cytb common subsequence and / or species evolution information of the organism to be predicted; inputting the molecular map of the compound to be predicted and the species information of the organism to be predicted into a multi-species aquatic organism acute toxicity prediction model to obtain acute toxicity prediction data of the compound to be predicted for the organism to be predicted; wherein the multi-species aquatic organism acute toxicity prediction model is obtained by the above-described establishment method.
[0007] Another aspect of this disclosure provides an apparatus for establishing an acute toxicity prediction model for multiple aquatic organisms, comprising: a first acquisition module for acquiring actual acute toxicity data of multiple compounds on multiple aquatic organisms, molecular diagrams of the multiple compounds, and species information of the multiple aquatic organisms, wherein the species information includes Cytb common subsequences and / or species evolutionary information of the multiple aquatic organisms; a first obtaining module for inputting the molecular diagrams of the multiple compounds and the species information of the multiple aquatic organisms into an initial prediction model to obtain initial acute toxicity prediction data of the multiple compounds on the multiple aquatic organisms; and an optimization module for optimizing the network parameters of the initial prediction model based on the initial acute toxicity prediction data and the actual acute toxicity data to obtain a trained acute toxicity prediction model for multiple aquatic organisms.
[0008] Another aspect of this disclosure provides an acute toxicity prediction apparatus, comprising: a second acquisition module for acquiring a molecular map of a compound to be predicted and species information of an organism to be predicted, wherein the species information includes the Cytb common subsequence and / or species evolutionary information of the organism to be predicted; and a second acquisition module for inputting the molecular map of the compound to be predicted and the species information of the organism to be predicted into a multi-species aquatic organism acute toxicity prediction model to obtain acute toxicity prediction data of the compound to be predicted for the organism to be predicted; wherein the multi-species aquatic organism acute toxicity prediction model is obtained by the method described above.
[0009] Another aspect of this disclosure provides an electronic device comprising:
[0010] One or more processors;
[0011] Memory, used to store one or more programs.
[0012] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described above.
[0013] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the method described above.
[0014] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the method described above.
[0015] According to embodiments of this disclosure, molecular diagrams of multiple compounds and species information of multiple aquatic organisms are used as inputs to an initial prediction model. The actual acute toxicity data of known compounds to multiple aquatic organisms are used as labels to train the initial prediction model. This allows the model to learn the molecular characteristics of the acute toxicity effects of different compounds on different aquatic organism species, thereby establishing a multi-species aquatic organism acute toxicity prediction model that can target multiple compounds and multiple aquatic organisms. This model has a wider range of convenient applications and better prediction performance.
[0016] According to embodiments of this disclosure, species information, including Cytb common subsequences and / or species evolutionary information of multiple aquatic organisms, can be better extracted to identify information between species. Cytb common subsequences exhibit relatively stable characteristics, while species evolutionary information enables better clustering analysis of species.
[0017] The acute toxicity prediction method provided in this disclosure is applicable to the prediction of acute toxicity of large-scale compounds and multiple species of aquatic organisms. The prediction method is simple and efficient, improves prediction performance, expands the prediction range, and has broad application prospects in the fields of chemical environmental health risk assessment and environmental assessment. Attached Figure Description
[0018] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0019] Figure 1 This illustration schematically shows an exemplary system architecture of a method for establishing a multi-species aquatic organism acute toxicity prediction model and a toxicity prediction method and apparatus thereof, which can be applied according to embodiments of the present disclosure.
[0020] Figure 2 A flowchart illustrating a method for establishing a multi-species aquatic organism acute toxicity prediction model according to an embodiment of the present disclosure is shown.
[0021] Figure 3 The schematic diagram illustrates the structure of a multi-species aquatic organism acute toxicity prediction model according to an embodiment of the present disclosure;
[0022] Figure 4This illustration schematically shows a method for obtaining initial acute toxicity prediction data according to embodiments of the present disclosure;
[0023] Figure 5 A flowchart illustrating a method for establishing a multi-species aquatic organism acute toxicity prediction model according to another embodiment of the present disclosure is shown.
[0024] Figure 6 A flowchart illustrating a method for predicting acute toxicity of multi-species aquatic organisms according to embodiments of the present disclosure is shown schematically.
[0025] Figure 7 A block diagram of an apparatus for establishing a multi-species aquatic organism acute toxicity prediction model according to an embodiment of the present disclosure is shown schematically.
[0026] Figure 8 A block diagram schematically illustrates an apparatus for establishing a multi-species aquatic organism acute toxicity prediction model according to embodiments of the present disclosure; and
[0027] Figure 9 A block diagram of an electronic device suitable for establishing a method for predicting the acute toxicity of multi-species aquatic organisms, according to an embodiment of the present disclosure, is shown schematically. Detailed Implementation
[0028] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0031] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0032] In the embodiments of this disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to maintain user personal information security and network security. In the embodiments of this disclosure, user authorization or consent has been obtained before acquiring or collecting user personal information.
[0033] Traditional in vivo toxicology experiments are not only costly and time-consuming, but also rely on a large number of laboratory animals. To address these issues, in vitro computational methods can be considered to replace or reduce traditional in vivo experiments. These computational methods can more quickly and accurately predict the potential toxicity of chemicals to aquatic organisms, providing more effective risk assessment and management strategies for environmental protection and human health.
[0034] In related technologies, prediction models built using traditional machine learning algorithms require manual calculation of molecular quantification parameters as molecular descriptors using quantum chemical calculation methods. The preliminary data preparation consumes a lot of time and computing resources. Furthermore, prediction models built using traditional machine learning algorithms are only designed for a single species, resulting in a narrow application scope, limited prediction capabilities, and a need to improve prediction performance.
[0035] Quantitative structure-activity relationship (QSAR) links the biological effects of chemical substances to their molecular structure and properties, becoming a key method for predicting molecular properties. QSAR is based on the crucial assumption that molecules with similar structures typically exhibit similar functional properties. By analyzing the relationship between molecular structure and properties, QSAR models can quickly and efficiently predict molecular characteristics without the need for laboratory experiments.
[0036] A major challenge in building efficient predictive models is the scarcity and imbalance of aquatic organism toxicity data. Toxicity databases show that high-quality aquatic organism toxicity data is scarce and unevenly distributed, with data primarily concentrated on fish, while data on other species such as algae and crustaceans is very limited. This uneven data distribution hinders the development of comprehensive aquatic organism toxicity prediction models. Furthermore, the differences in toxicity testing conditions among different species make data utilization difficult. Faced with these challenges, there is an urgent need to develop more accurate and comprehensive aquatic toxicity prediction models to more accurately assess the potential toxicity of chemicals to various aquatic organisms. These models need to cover not only common aquatic species, such as water fleas, algae, and fish as ecological indicator species, but also a variety of other aquatic organisms. This approach can provide stronger scientific support for environmental protection and also provide crucial data support for the safe use of chemicals.
[0037] In view of this, embodiments of the present disclosure provide a method for establishing an acute toxicity prediction model for multiple aquatic organisms, comprising: acquiring actual acute toxicity data of multiple compounds on multiple aquatic organisms, molecular maps of multiple compounds, and species information of multiple aquatic organisms, wherein the species information includes Cytb common subsequences and / or species evolutionary information of multiple aquatic organisms; inputting the molecular maps of multiple compounds and the species information of multiple aquatic organisms into an initial prediction model to obtain initial acute toxicity prediction data of multiple compounds on multiple aquatic organisms; and optimizing the network parameters of the initial prediction model based on the initial acute toxicity prediction data and the actual acute toxicity data to obtain a trained acute toxicity prediction model for multiple aquatic organisms.
[0038] Figure 1 The illustration schematically depicts an exemplary system architecture 100 for establishing a multi-species aquatic organism acute toxicity prediction model and a toxicity prediction method and apparatus, according to embodiments of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0039] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0040] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).
[0041] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0042] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0043] It should be noted that the method for establishing a multi-species aquatic organism acute toxicity prediction model and the toxicity prediction method provided in this disclosure embodiment can generally be executed by server 105. Correspondingly, the device for establishing a multi-species aquatic organism acute toxicity prediction model and the toxicity prediction device provided in this disclosure embodiment can generally be located in server 105. The method for establishing a multi-species aquatic organism acute toxicity prediction model and the toxicity prediction method provided in this disclosure embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the device for establishing a multi-species aquatic organism acute toxicity prediction model and the toxicity prediction device provided in this disclosure embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, the method for establishing a multi-species aquatic organism acute toxicity prediction model and the toxicity prediction method provided in this disclosure embodiment can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the apparatus for establishing a multi-species aquatic organism acute toxicity prediction model and the toxicity prediction apparatus provided in this disclosure embodiment can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0044] For example, the actual acute toxicity data of multiple compounds to multiple aquatic species, the molecular maps of multiple compounds, and the species information of multiple aquatic species can be originally stored in any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 (e.g., the first terminal device 101, but not limited thereto), or stored on an external storage device and can be imported into the first terminal device 101. Then, the first terminal device 101 can locally execute the method for establishing the acute toxicity prediction model for multiple aquatic species and the toxicity prediction method provided in the embodiments of this disclosure, or send the actual acute toxicity data of multiple compounds to multiple aquatic species, the molecular maps of multiple compounds, and the species information of multiple aquatic species to other terminal devices, servers, or server clusters, and the other terminal devices, servers, or server clusters that receive the actual acute toxicity data of multiple compounds to multiple aquatic species, the molecular maps of multiple compounds, and the species information of multiple aquatic species can execute the method for establishing the acute toxicity prediction model for multiple aquatic species provided in the embodiments of this disclosure.
[0045] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0046] Figure 2 A flowchart illustrating a method for establishing a multi-species aquatic organism acute toxicity prediction model according to an embodiment of the present disclosure is shown.
[0047] like Figure 2 As shown, the method includes operations S100~S300.
[0048] In operation S100, actual acute toxicity data of multiple compounds to multiple aquatic species, molecular maps of multiple compounds, and species information of multiple aquatic species are obtained. Among them, the species information includes Cytb common subsequences and / or species evolutionary information of multiple aquatic species.
[0049] In operation S200, molecular diagrams of multiple compounds and species information of multiple aquatic organisms are input into the initial prediction model to obtain initial acute toxicity prediction data of multiple compounds on multiple aquatic organisms.
[0050] In operating the S300, the network parameters of the initial prediction model are optimized based on the initial acute toxicity prediction data and the actual acute toxicity data to obtain a trained multi-species aquatic organism acute toxicity prediction model.
[0051] According to embodiments of this disclosure, the various compounds can be any type of compound. For example, they can be compounds that are toxic to aquatic organisms. Specifically, the various compounds can be halogenated phenolic compounds, nitrobenzene compounds, alkylbenzene compounds, etc.
[0052] According to embodiments of this disclosure, multi-species aquatic organisms may include, for example, aquatic organisms of various species such as algae, fish, and crustaceans.
[0053] According to embodiments of this disclosure, the molecular graph of a compound can be an undirected graph containing node features and edge features, wherein node features are used to characterize atomic features and edge features are used to characterize bond features.
[0054] According to embodiments of this disclosure, species information may include Cytb common subsequences of multiple aquatic species; species information may include species evolutionary information of multiple aquatic species; species information may also include Cytb common subsequences and species evolutionary information of multiple aquatic species.
[0055] According to embodiments of this disclosure, Cytb (Cytochrome b gene) is the cytochrome b gene. Cytb sequence information from different aquatic organisms can be obtained, and MEGA11 software can be used for matching, alignment, and trimming to obtain common Cytb subsequences as species information for multiple aquatic organisms.
[0056] According to embodiments of this disclosure, species evolution information includes species classification information of multiple aquatic organisms. The species classification information of an organism may include information such as the kingdom, phylum, class, order, family, genus, and species to which the organism belongs.
[0057] According to embodiments of this disclosure, acute toxicity data of the compound to organisms can be obtained using LC-MS. 50 Measurement, LC 50 The median lethal concentration (LD50) is the concentration of a compound that would kill half the population of a species. The higher the concentration, the lower the toxicity.
[0058] According to embodiments of this disclosure, molecular diagrams of multiple compounds and species information of multiple aquatic organisms are used as inputs to an initial prediction model. The actual acute toxicity data of known compounds to multiple aquatic organisms are used as labels to train the initial prediction model. This allows the model to learn the molecular characteristics of the acute toxicity effects of different compounds on different aquatic organism species, thereby establishing a multi-species aquatic organism acute toxicity prediction model that can target multiple compounds and multiple aquatic organisms. This model has a wider range of convenient applications and better prediction performance.
[0059] According to embodiments of this disclosure, species information, including Cytb common subsequences and / or species evolutionary information of multiple aquatic species, can be better extracted to identify information between species. Cytb common subsequences exhibit relatively stable characteristics, while species evolutionary information enables better clustering analysis of species.
[0060] According to embodiments of this disclosure, using the molecular map of the compound as input to the initial prediction model eliminates the need to use molecular weight parameters calculated using quantum chemical calculation methods in traditional machine learning as molecular descriptors, saving time and computational resources for molecular description calculation and selection, and reducing the requirements for computational chemistry foundation in applications.
[0061] The following is for reference. Figures 3-5 In conjunction with specific embodiments, Figure 2 The method shown will be further explained.
[0062] Figure 3 The schematic diagram illustrates the structure of a multi-species aquatic organism acute toxicity prediction model according to an embodiment of the present disclosure.
[0063] According to embodiments of this disclosure, such as Figure 3As shown, the initial prediction model includes a feature extraction network 310, an interaction layer 320, and a prediction layer 330 connected in sequence. The feature extraction network 310 includes a species feature extraction network 311 and a molecular feature extraction network 312.
[0064] According to embodiments of this disclosure, deep learning is a type of machine learning algorithm that can use hierarchical recombination of features to extract relevant information and learn the pattern properties of data representations.
[0065] Figure 4 The illustration schematically depicts a method for obtaining initial acute toxicity prediction data according to an embodiment of the present disclosure.
[0066] like Figure 4 As shown, molecular diagrams of multiple compounds and species information of multiple aquatic organisms are input into the initial prediction model to obtain initial acute toxicity prediction data of multiple compounds on multiple aquatic organisms, including operations S210~S240.
[0067] In operation S210, the species information 301 of multiple aquatic organisms is input into the species feature extraction network 311 to extract the species feature tensor of the species information of multiple aquatic organisms.
[0068] In operation S220, the molecular graphs 302 of various compounds are input into the molecular feature extraction network 312 to extract the molecular feature tensors of the molecular graphs of various compounds.
[0069] In operation S230, the species feature tensor and the molecular feature tensor are input into the interaction layer 320 to obtain a combined feature tensor of the species feature tensor and the molecular feature tensor.
[0070] In operation S240, the combined feature tensor is input into prediction layer 330 to obtain initial acute toxicity prediction data.
[0071] According to embodiments of this disclosure, a feature extraction network can be used to extract species feature tensors of species information from multiple aquatic organisms; molecular feature tensors can be extracted using molecular graphs of multiple compounds; an interaction layer is used to combine the species feature tensors and molecular feature tensors to obtain a combined feature tensor; and a prediction layer is used to predict the combined feature tensor to obtain initial acute toxicity prediction data. Therefore, by sequentially extracting and combining features from the species information of multiple aquatic organisms and the molecular graphs of multiple compounds, and then performing initial acute toxicity prediction, the initial prediction model can better learn the acute toxicity effects of different compounds on different aquatic organism species, resulting in a multi-species aquatic organism acute toxicity prediction model with a wider prediction range and higher prediction efficiency.
[0072] According to embodiments of this disclosure, species information may include Cytb common subsequences of multiple aquatic organisms. Operation S210 may include: extracting features from the Cytb common subsequences to obtain a Cytb feature tensor, which is the species feature tensor.
[0073] According to embodiments of this disclosure, the species information may further include evolutionary information of multiple aquatic species. Operation S210 may include: extracting features from the evolutionary information to obtain a species evolutionary tensor, which is the species feature tensor.
[0074] According to embodiments of this disclosure, species information may further include Cytb common subsequences of multiple aquatic species and species evolutionary information.
[0075] According to an embodiment of this disclosure, operation S210, which involves inputting the species information 301 of multiple aquatic organisms into the species feature extraction network 311 to extract the species feature tensor of the species information of multiple aquatic organisms, may include operations S211 to S213.
[0076] In operation S211, features are extracted from the common subsequences of Cytb to obtain the Cytb feature tensor.
[0077] In operation S212, feature extraction is performed on the species evolution information to obtain the species evolution tensor.
[0078] In operation S213, the Cytb feature tensor and the species evolution tensor are concatenated to obtain the species feature tensor.
[0079] According to embodiments of this disclosure, species information, including Cytb common subsequences of multiple aquatic species and species evolutionary information, can be better extracted to reveal information between species. Cytb common subsequences are commonly used to construct phylogenetic trees and possess relatively stable characteristics; species evolutionary information can better perform cluster analysis on species, but it neglects the microscopic differences in species relationships. Simultaneously using species information, including Cytb common subsequences of multiple aquatic species and species evolutionary information, as complementary methods, can fully extract the species characteristic tensor of each species.
[0080] According to an embodiment of this disclosure, operation S211, namely, extracting features from the Cytb common subsequence to obtain the Cytb feature tensor, includes: performing one-hot encoding on the Cytb common subsequence of multiple aquatic organisms according to the base sequence to obtain the Cytb feature tensor; and using a convolutional kernel to extract features from the Cytb feature tensor to obtain the Cytb feature tensor.
[0081] According to embodiments of this disclosure, the ATCG characters in the Cytb sequence can be encoded according to [1,0,0,0], [0,1,0,0], [0,0,1,0], [0,0,01]. Then, the evolutionary classification information of different aquatic organisms is obtained, all of which are summarized and converted into a 01 string, and the value is assigned as 1 or 0 according to whether the species conforms to the corresponding evolutionary taxonomy.
[0082] According to embodiments of this disclosure, the convolution kernel can be a 4×n×1 convolution kernel, specifically, for example, a 4×3×1 convolution kernel.
[0083] According to an embodiment of this disclosure, operation S212, which is to extract features from species evolutionary information to obtain a species evolutionary tensor, includes: performing one-hot encoding on species classification information to obtain a species evolutionary tensor; and using a fully connected layer to extract features from the species evolutionary tensor to obtain a species evolutionary tensor.
[0084] According to embodiments of this disclosure, evolutionary populations of multiple aquatic organisms can be aggregated and represented by strings of equal length to their number. The seven evolutionary populations of the organisms—kingdom, phylum, class, order, family, genus, and species—are assigned a 1, while the remaining positions are assigned a 0.
[0085] According to embodiments of this disclosure, molecular graphs of various compounds include node features for characterizing atomic features and edge features for characterizing bond features.
[0086] According to embodiments of this disclosure, molecular graphs of multiple compounds are constructed by adding node features for characterizing atomic feature information and edge features for characterizing bond feature information to an undirected graph based on the SMILES encoding of multiple compounds, thereby obtaining molecular graphs of multiple compounds.
[0087] According to embodiments of this disclosure, atomic feature information includes one or more of the following: atomic symbol type, atomic implicit valence, number of free electrons in the atom, number of atomic connections, atomic formal charge, atomic hybridization mode, number of hydrogen atoms connected to the atom, atomic chirality, and FCFP fingerprint features.
[0088] According to embodiments of this disclosure, the key feature information includes one or more of the following: key type, whether the key is conjugated, whether the key is on a ring, and key stereo configuration information.
[0089] According to embodiments of this disclosure, SMILES encoding is an encoding method used to describe molecular structures. Molecules are represented by simple symbols of chemical structures, such as element symbols, plus and minus signs, parentheses, etc., which can quickly and directly represent the topological structure and connectivity of molecules and can be easily recognized and processed by computer systems. However, features learned directly from SMILES encoding lose a significant amount of information. This disclosure converts SMILES encoding into molecular graphs, which can better characterize molecular structures, thereby extracting key information more efficiently and reducing unnecessary repetitive steps in computation. The conversion between SMILES encoding and molecular graphs makes it easier to use initial prediction models to learn the molecular characteristics of the acute toxicity of compounds to multiple aquatic species, thus saving computation time.
[0090] Specifically, the molecular graph converted from the SMILES encoding of the compound is used as the input of the initial prediction model. The actual acute toxicity data of known compounds and multiple species of aquatic organisms are used as labels to train the initial prediction model and extract the molecular features of the acute toxicity effect of the compound. The resulting multi-species aquatic organism acute toxicity prediction model is applicable to the prediction of the acute toxicity of large-scale compounds and multiple species of aquatic organisms, reducing the requirements for computational chemistry foundation when applying the prediction model.
[0091] According to embodiments of this disclosure, a molecular graph is a set of nodes and edges, where nodes correspond to atoms in the molecule and edges correspond to bonds in the molecule. Each node has node features for characterizing atomic features and each edge has edge features for characterizing bond features.
[0092] According to embodiments of this disclosure, an undirected graph G(V,E) can be created using PyG (a graph neural network library, PyTorchGeometric), where V represents the set of atoms and E represents the set of bonds.
[0093] Specifically, molecular graphs for various compounds are constructed using the following method: First, the number of non-hydrogen atoms N in the SMILES encoding of volatile organic compounds can be obtained using the RDKit toolkit. N nodes are added to the undirected graph G. For each node's corresponding non-hydrogen atom, atomic feature information is obtained using the RDKit toolkit and encoded using one-hot encoding to generate a 1×L matrix. Second, the node feature matrices corresponding to all nodes are connected to generate an N×L matrix, which is stored as the node features of the molecular graph G, where L is set as needed. Third, all combinations of non-hydrogen atoms in the SMILES encoding of volatile organic compounds are traversed, adding edges to the non-hydrogen atoms corresponding to all nodes in the molecular graph G. For each edge representing a chemical bond, bond feature information is obtained using the RDKit toolkit and encoded using one-hot encoding to generate a 1×K matrix. Finally, the edge feature matrices corresponding to all edges are connected to generate a 2M×K matrix, which is stored as the edge features of the molecular graph G, where M is the number of chemical bonds and K can be set as needed.
[0094] To better illustrate the process of converting SMILES codes into molecular diagrams, we will take the volatile organic compound isoprene as an example. The SMILES code for isoprene is C=CC(=C)C.
[0095] For example, the atomic feature information obtained for isoprene includes: atomic symbol type (10 bits), atomic implicit valence (2 bits), number of free electrons (2 bits), number of atomic connections (7 bits), atomic formal charge (3 bits), atomic hybridization type (4 bits), functional fingerprint (FCFP) feature (6 bits), number of hydrogen atoms connected to the atom (5 bits), atomic chirality (2 bits), and whether the atom is chiral (1 bit). Then, the first carbon atom feature in the SMILES code of isoprene is [True, False, False, False, False, False, False, False, False, False, False, False, False, False, True, True, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, False, 0, 0, 0, 0, 0, 0, False, False, False, False, False, 0]. A matrix of size 1×L=1×42 [1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 1, 0] is generated using one-hot encoding. Similarly, the node feature matrices corresponding to all nodes are concatenated to generate an N×L=5×42 matrix, which is stored as the node features of the molecular graph G.
[0096] For example, the bond feature information obtained for isoprene includes: bond type (4 bits), whether the bond is conjugated (1 bit), whether the bond is on a ring (1 bit), and the bond stereoconfiguration information (4 bits). Therefore, the feature of the first edge added to the first carbon atom in the SMILES encoding of isoprene (i.e., the C=C bond) is [False, True, False, False, True, False, True, False, False, False]. A matrix of size 1×K=1×10 is generated using one-hot encoding [0, 1, 0, 0, 1, 0, 1, 0, 0, 0]. Connecting all the edge feature matrices corresponding to the edges generates a 2M×K=8×10 matrix, which is stored as the edge features of the molecular graph G.
[0097] According to embodiments of this disclosure, the molecular graph based on SMILES encoding conversion contains atomic feature information and bond feature information, which can better characterize the molecular structure of the compound, thereby facilitating the subsequent learning of the molecular features of the compound's acute toxicity to aquatic organisms by the graph neural network.
[0098] According to embodiments of this disclosure, in operation S220, the molecular feature extraction network is a graph isomorphic network. A graph isomorphic network is a general framework for supervised learning of graph structure data. By extracting the molecular features of a compound using a graph isomorphic network, the correlation between each atom and its neighboring atoms and / or bonds can be learned more effectively, thereby improving the accuracy of the compound's acute toxicity to aquatic organisms.
[0099] Graph isomorphic networks (GNNs) are a type of graph neural network. Graph neural networks are a deep learning method that learns information based on graph representations, making them particularly suitable for research areas such as protein-protein interaction networks and molecular properties, where information is expressed using graphs and networks. Deep learning methods require encoding molecular structures mathematically, including graph-based representations and molecular descriptor-based representations, corresponding to graph neural network models and traditional models, respectively. Benchmark tests on multiple datasets show that graph-based models outperform other methods on most datasets, demonstrating a significant advantage. Applying graph neural network-based deep learning methods to predict molecular physicochemical properties and the toxicity of molecular compounds holds great potential.
[0100] According to an embodiment of this disclosure, operation S220 specifically includes: updating the hidden state of each node feature for each node in the molecular graph of the compound to obtain the updated atomic features corresponding to each node; and merging the updated atomic features corresponding to all nodes in the molecular graphs of multiple compounds to obtain a molecular feature tensor.
[0101] Specifically, each node The initial features are set as node features. For each time step (or layer) Update the features of each node by aggregating neighbor nodes. :
[0102] (1)
[0103] The MLP is a fully connected layer used to perform nonlinear transformations on the aggregate features of nodes. These are learnable parameters used to adjust the influence of self-loops, ensuring that each node's features are not completely smoothed out during aggregation. The final features of the nodes after T time steps (or layers) are... It combines its own initial characteristics with information from all its neighbors. For each node... Initial features Added to the final hidden state, To obtain the final characteristics of the node. Features of each node The summation yields the feature tensor A of the entire molecule, which represents the molecular characteristics of the entire compound and can be used for prediction tasks.
[0104] According to an embodiment of this disclosure, operation S230, which involves inputting the species feature tensor and the molecular feature tensor into the interaction layer to obtain a combined feature tensor of the species feature tensor and the molecular feature tensor, specifically includes operations S231 and S232.
[0105] In operation S231, the species feature tensor and the molecular feature tensor are multiplied by a dot product to obtain a mixed feature tensor.
[0106] In operation S232, the species feature tensor, the mixed feature tensor, and the molecular feature tensor are concatenated to obtain the combined feature tensor.
[0107] According to embodiments of this disclosure, before performing dot product on the species feature tensor and the molecular feature tensor, the method further includes converting the species feature tensor and the molecular feature tensor into two feature tensors with the same shape through fully connected layers.
[0108] Specifically, the molecular feature tensor A can be transformed into a tensor A' of a specific length L through a fully connected layer, and the species feature tensor B can be transformed into a tensor B' of a specific length L through a fully connected layer. The dot product of A' and B' is used to obtain tensor C. A', C and B' are then connected sequentially to obtain the final combined feature tensor D. The combined feature tensor D is then input into the prediction layer to obtain the acute toxicity prediction result.
[0109] According to embodiments of this disclosure, the molecular feature tensor A and the species feature tensor B can be converted to have the same shape before being multiplied and concatenated. This allows the molecular features and species information of the compound to be combined, and subsequent feature extraction can be performed based on the combined features. Further, for ease of understanding, the above-described conversion, multiplication, and concatenation operations are illustrated with examples below:
[0110] The molecular characteristic tensor A is represented as ([[1., 1., 1.],
[0111] [0., 0., 0.]]).
[0112] The shape of the molecular characteristic tensor is (2,3).
[0113] First, the molecular feature tensor A is transformed into a one-dimensional vector X1 of length 3, and the species information feature tensor B is transformed into a one-dimensional vector Y1 of length 3.
[0114] At this point, the two feature tensors X1 and Y1 have the same first two dimensions and can be multiplied to obtain tensor Z. The feature tensor X1 is represented as [1,0,1], the feature tensor Y1 is represented as [3,4,0], the feature tensor Z is represented as [3,0,0], and the concatenated feature tensor D is represented as [1,0,1,3,0,0,3,4,0].
[0115] To better illustrate the composition of toxicity data, we take the acute toxicity data of the volatile organic compound isoprene to zebrafish as an example. The composition of the toxicity data is as follows: compound SMILES: C=CC(=C)C, zebrafish Cytb gene sequence: GCTTATCC (partial sequence), toxicity value: 12 mg / L, kingdom: animal kingdom, phylum: chordate, class: osteoid, order: cypriniformes, family: cyprinid family, genus: scabra, species: zebrafish.
[0116] For example, SMILES encoding is a unique string obtained by mapping each compound one-to-one, reflecting the structure of the compound. The common sequence obtained after matching and aligning the Cytb sequence is a string of length 1041 containing the four codons ATCG. According to A-[1,0,0,0], T-[0,1,0,0], C-[0,0,1,0], G-[0,0,0,1], the one-hot encoding of the Cytb sequence (partial) GCTTATCC is obtained, which results in the following matrix: [[0,0,0,1],[0,0,1,0],[0,1,0,0],[0,1,0,0],[1,0,0,0],[0,1,0,0],[0,0,1,0],[0,0,1,0]]. Collect all biological evolutionary groups involved in the species and represent them with strings of equal length to their number. Assign 1 to the 7 evolutionary groups to which zebrafish belongs, and 0 to the rest.
[0117] According to an embodiment of this disclosure, in operation S240, the prediction layer 330 may include a plurality of fully connected layers, each fully connected layer including a dropout layer to prevent the initial prediction model from overfitting.
[0118] According to embodiments of this disclosure, combined feature tensors can be sequentially input into multiple fully connected layers to obtain initial acute toxicity prediction data.
[0119] According to embodiments of this disclosure, the dropout layer, with a coefficient of P, randomly shuts down a portion of neurons during training and temporarily removes them from the network with a certain probability. This avoids the model from relying too much on certain neurons or features, reduces the interdependence between neurons, and improves the robustness of the model.
[0120] According to embodiments of this disclosure, each node in the current layer of the fully connected layer is connected to all nodes in the previous layer, and the number of nodes in each layer is fA, fB, 1. Except for the last layer, the output of each layer uses the ReLU activation function (Linear rectification function) to convert the linear features in the neural network into nonlinear features. The last layer of the fully connected layer is the output layer, which outputs the initial acute toxicity data of the corresponding compound to the aquatic organism of the species.
[0121] Figure 5 A flowchart illustrating a method for establishing a multi-species aquatic organism acute toxicity prediction model according to another embodiment of the present disclosure is shown.
[0122] like Figure 5 As shown, the method for establishing a multi-species aquatic organism acute toxicity prediction model includes operations S501~S507.
[0123] In operation S501, actual acute toxicity data of multiple compounds on multiple aquatic species are obtained, along with species information for multiple compounds and multiple aquatic species.
[0124] In operation S502, the molecular diagram of the compound is drawn based on SMILES. The specific drawing method is the same as in the above embodiment and will not be repeated here.
[0125] In step S503, the species information is encoded. The specific encoding method is the same as described in the above embodiment and will not be repeated here.
[0126] When operating S504, the dataset is divided into training set, validation set and test set.
[0127] When operating the S505, the initial prediction model is trained and its parameters are adjusted. Specifically, the network parameters of the initial prediction model are optimized using data from the training and validation sets.
[0128] According to embodiments of this disclosure, in operations S300 and S505, the network parameters of the initial prediction model may include model parameters and model hyperparameters. To determine the optimal prediction model, the actual acute toxicity data of multiple compounds to multiple aquatic species, the molecular graphs of multiple compounds, and the species information of multiple aquatic species are divided into a training set and a validation set. Operations S300 and S505 specifically include operations S310 to sub-steps S340.
[0129] In operation 310, based on preset values of a set of model hyperparameter combinations, the graph neural network model is iteratively trained using the reaction rate constant data and the reaction rate constant prediction data in the training set, and the model parameters are optimized to obtain the trained graph neural network model.
[0130] In operation 320, the trained graph neural network model is used to predict the validation set and obtain statistical parameters.
[0131] In operation 330, based on the automatic hyperparameter tuning algorithm, the preset values of the model hyperparameter combination are adjusted, and the training and prediction operations are repeated iteratively until the optimal statistical parameters are determined. The set of model hyperparameter combinations corresponding to the optimal statistical parameters is the optimal model hyperparameter.
[0132] In operation 340, the trained graph neural network model obtained based on the optimal model hyperparameters was determined as the multi-species acute toxicity prediction model for compounds on multi-species aquatic organisms.
[0133] According to embodiments of this disclosure, the model hyperparameters are combined as follows: ,in, The learning rate is used to control the pace of convergence to a local minimum. is the weight decay parameter, used to reduce model complexity; P is the random inactivation rate of neurons in the fully connected layer, used to prevent overfitting; S is the batch data size; T is the iteration step size in the graph isomorphic network.
[0134] According to embodiments of this disclosure, to improve model prediction performance, the automatic parameter tuning algorithm is preferably a Bayesian optimization algorithm; the statistical parameters include at least one of Mean Squared Error (MSE), Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Edtermination (R²). The model parameters corresponding to the lowest RMSE in the validation set results during the E-generation (epaoch) iterations are selected and saved.
[0135] In operation S506, the initial prediction model is evaluated. Specifically, the initial prediction model is evaluated using data from the test set.
[0136] According to embodiments of this disclosure, a test set is further defined based on the actual acute toxicity data of multiple compounds to multiple aquatic organisms, the molecular diagrams of the multiple compounds, and the species information of the multiple aquatic organisms. The method of establishing this disclosure further includes: evaluating the performance of the acute toxicity prediction model for multiple aquatic organisms using the test set.
[0137] By operating S507, a well-trained multi-species aquatic organism acute toxicity prediction model is obtained.
[0138] Figure 6 A flowchart illustrating an acute toxicity prediction method according to an embodiment of the present disclosure is shown schematically.
[0139] like Figure 6 As shown, the method includes operations S610~S620.
[0140] In operation S610, the molecular map of the compound to be predicted and the species information of the organism to be predicted are obtained, wherein the species information includes the Cytb common subsequence of the organism to be predicted and / or species evolution information.
[0141] In operation S620, the molecular diagram of the compound to be predicted and the species information of the organism to be predicted are input into the multi-species aquatic organism acute toxicity prediction model to obtain the acute toxicity prediction data of the compound to be predicted and the organism to be predicted; wherein, the multi-species aquatic organism acute toxicity prediction model is obtained by the method described above.
[0142] The acute toxicity prediction method provided in this disclosure is applicable to the prediction of acute toxicity of large-scale compounds and multiple species of aquatic organisms. The prediction method is simple and efficient, improves prediction performance, expands the prediction range, and has broad application prospects in the fields of chemical environmental health risk assessment and environmental assessment.
[0143] The acute toxicity prediction method provided in this disclosure can reliably predict the acute toxicity of the compound to be predicted to a variety of aquatic organisms.
[0144] Figure 7 A block diagram of an apparatus for establishing a multi-species aquatic organism acute toxicity prediction model according to an embodiment of the present disclosure is shown schematically.
[0145] like Figure 7 As shown, the device 700 for establishing a multi-species aquatic organism acute toxicity prediction model includes a first acquisition module 710, a first acquisition module 720, and an optimization module 730.
[0146] The first acquisition module 710 is used to acquire actual acute toxicity data of multiple compounds to multiple aquatic organisms, molecular diagrams of multiple compounds, and species information of multiple aquatic organisms, wherein the species information includes Cytb common subsequences and / or species evolutionary information of multiple aquatic organisms.
[0147] The first acquisition module 720 is used to input molecular diagrams of multiple compounds and species information of multiple aquatic organisms into the initial prediction model to obtain initial acute toxicity prediction data of multiple compounds on multiple aquatic organisms.
[0148] The optimization module 730 is used to optimize the network parameters of the initial prediction model based on the initial acute toxicity prediction data and the actual acute toxicity data, so as to obtain a trained multi-species aquatic organism acute toxicity prediction model.
[0149] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0150] For example, any plurality of the first acquisition module 710, the first acquisition module 720, and the optimization module 730 can be combined into one module / unit / subunit, or any one of these modules / units / subunits can be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits can be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the first acquisition module 710, the first acquisition module 720, and the optimization module 730 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first acquisition module 710, the first acquisition module 720, and the optimization module 730 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0151] Figure 8 A block diagram of an apparatus for establishing a multi-species aquatic organism acute toxicity prediction model according to an embodiment of the present disclosure is shown schematically.
[0152] like Figure 8 As shown, the device 700 for establishing a multi-species aquatic organism acute toxicity prediction model includes a second acquisition module 810 and a second acquisition module 820.
[0153] The second acquisition module 810 is used to acquire the molecular map of the compound to be predicted and the species information of the organism to be predicted, wherein the species information includes the Cytb common subsequence and / or species evolution information of the organism to be predicted.
[0154] The second acquisition module 820 is used to input the molecular diagram of the compound to be predicted and the species information of the organism to be predicted into the multi-species aquatic organism acute toxicity prediction model to obtain the acute toxicity prediction data of the compound to be predicted in the organism to be predicted; wherein, the multi-species aquatic organism acute toxicity prediction model is obtained by the method establishment method.
[0155] It should be noted that the data processing system part in the embodiments of this disclosure corresponds to the data processing method part in the embodiments of this disclosure. The specific description of the data processing system part is referred to in the data processing method part, and will not be repeated here.
[0156] Figure 9 A block diagram of an electronic device suitable for implementing the methods described above, according to embodiments of the present disclosure, is illustrated schematically. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0157] like Figure 9 As shown, an electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0158] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 902 and / or RAM 903. It should be noted that the programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0159] According to embodiments of this disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.
[0160] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 909, and / or installed from removable medium 911. When the computer program is executed by processor 901, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0161] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0162] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0163] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.
[0164] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the method for establishing a multi-species aquatic organism acute toxicity prediction model or the toxicity prediction method provided in the embodiments of this disclosure.
[0165] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0166] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0167] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0169] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A method for establishing a multi-species aquatic organism acute toxicity prediction model, comprising: Acquire actual acute toxicity data of multiple compounds to multiple aquatic organisms, molecular diagrams of the multiple compounds, and species information of the multiple aquatic organisms, wherein the species information includes Cytb common subsequences and / or species evolutionary information of the multiple aquatic organisms; The molecular diagrams of the various compounds and the species information of the various aquatic organisms are input into the initial prediction model to obtain the initial acute toxicity prediction data of the various compounds on the various aquatic organisms. Based on the initial acute toxicity prediction data and the actual acute toxicity data, the network parameters of the initial prediction model are optimized to obtain a trained multi-species aquatic organism acute toxicity prediction model. The initial prediction model includes a feature extraction network, an interaction layer, and a prediction layer connected in sequence. The feature extraction network includes a species feature extraction network and a molecular feature extraction network. The step of inputting the molecular diagrams of the various compounds and the species information of the various aquatic organisms into the initial prediction model to obtain the initial acute toxicity prediction data of the various compounds on the various aquatic organisms includes: The species information of the multiple aquatic organisms is input into the species feature extraction network to extract the species feature tensor of the species information of the multiple aquatic organisms. The molecular graphs of the various compounds are input into the molecular feature extraction network to extract the molecular feature tensors of the molecular graphs of the various compounds. The species feature tensor and the molecular feature tensor are input into the interaction layer to obtain a combined feature tensor of the species feature tensor and the molecular feature tensor; The combined feature tensor is input into the prediction layer to obtain the initial acute toxicity prediction data; The species information includes the Cytb common subsequence of the multi-species aquatic organisms and species evolution information; The species feature tensor used to extract species information from the multi-species aquatic organisms includes: Feature extraction is performed on the Cytb common subsequences to obtain the Cytb feature tensor; Feature extraction is performed on the species evolution information to obtain the species evolution tensor; The Cytb feature tensor and the species evolution tensor are concatenated to obtain the species feature tensor.
2. The method for establishing according to claim 1, wherein, The step of inputting the species feature tensor and the molecular feature tensor into the interaction layer to obtain a combined feature tensor of the species feature tensor and the molecular feature tensor includes: Multiply the species feature tensor and the molecular feature tensor by a dot product to obtain a mixed feature tensor; The species feature tensor, the mixed feature tensor, and the molecular feature tensor are concatenated to obtain the combined feature tensor.
3. The method for establishing according to claim 1, wherein, The feature extraction of the Cytb common subsequences to obtain the Cytb feature tensor includes: The Cytb common subsequences of the multi-species aquatic organisms are encoded one-hot according to the base sequence to obtain the Cytb feature tensor. The Cytb feature tensor is obtained by using a convolution kernel to extract features.
4. The method for establishing according to claim 1, wherein, The species evolution information includes the species classification information of the multiple aquatic organisms; The feature extraction of the species evolutionary information to obtain the species evolutionary tensor includes: The species classification information is one-hot encoded to obtain the species evolution tensor; The species evolution tensor is obtained by feature extraction using a fully connected layer.
5. The method for establishing according to claim 1, wherein, The molecular feature extraction network is a graph isomorphic network.
6. The method for establishing according to claim 1, wherein, The prediction layer includes multiple fully connected layers, each of which contains a dropout layer to prevent the initial prediction model from overfitting.
7. The method for establishing according to claim 1, wherein, The molecular graphs of the various compounds include node features for characterizing atomic features and edge features for characterizing bond features. The molecular maps of the various compounds were constructed using the following methods: Based on the SMILES encoding of the various compounds, node features for characterizing atomic features and edge features for characterizing bond features are added to the undirected graph to obtain the molecular graph of the various compounds. The atomic feature information includes one or more of the following: atomic symbol type, atomic implicit valence, number of free electrons, number of atomic connections, atomic formal charge, atomic hybridization mode, number of hydrogen atoms connected to the atom, atomic chirality, and FCFP fingerprint features. The key feature information includes one or more of the following: key type, whether the key is conjugated, whether the key is on a ring, and key stereo configuration information.
8. A method for predicting acute toxicity, comprising: Obtain the molecular map of the compound to be predicted and the species information of the organism to be predicted, wherein the species information includes the Cytb common subsequence and / or species evolution information of the organism to be predicted; The molecular diagram of the compound to be predicted and the species information of the organism to be predicted are input into a multi-species aquatic organism acute toxicity prediction model to obtain the acute toxicity prediction data of the compound to be predicted on the organism. The multi-species aquatic organism acute toxicity prediction model is obtained by the establishment method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Model-estimation-based biological toxicity estimation method
CN105205322A
Crispr effector system based diagnostics for virus detection
CN111108220A