Chemical toxicity evaluation model construction method, evaluation method and related device

CN116386762BActive Publication Date: 2026-08-21HANGZHOU INST FOR ADVANCED STUDY UCAS +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310247852.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2026-08-21
Estimated Expiration
2043-03-15

AI Technical Summary

Technical Problem

[0004]本发明的主要目的是提出一种化学品毒性评估模型构建方法、评估方法及相关装置,旨在解决目前能够获得的化学品毒性数据可用性较差的问题,以及快速准确的预测待评估化学品的毒性

Benefits of technology

[0031]本发明的技术方案,通过化学品的分子结构相似度,分别建立目标毒性回归模型和目标毒性分类模型,以及对应的应用域,对于满足应用域的化学品,其分子结构能够保持与构建目标毒性回归模型的化学品样本较高的分子结构相似度,因此利用目标毒性回归模型能够较为准确的得到毒性终点数据,而对于不满足应用域的化学品,表明其分子结构特征较为特殊,利用目标分类模型也能够大致确定其毒性终点数据所处的范围区间,对于化学品进行评估时毒性时较为方便、快速。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386762B_ABST
    Figure CN116386762B_ABST
Patent Text Reader

Abstract

The application discloses a chemical toxicity evaluation model construction method, an evaluation method and related devices, and the toxicity evaluation model construction method comprises the following steps: obtaining an original chemical sample set; determining the molecular structure similarity of each chemical relative to other chemicals; screening a plurality of first candidate chemical sample sets; constructing a plurality of first candidate toxicity regression models respectively; determining a first target chemical sample set, a target toxicity regression model, and an application domain of the target toxicity regression model based on the plurality of first candidate toxicity regression models; removing the first target chemical sample set from the original chemical sample set to obtain a second target chemical sample set, and constructing a target toxicity classification model. The target toxicity regression model, the target toxicity classification model and the application domain are respectively established based on the molecular structure similarity of the chemicals, the target toxicity regression model can accurately predict toxicity endpoint data, the target classification model can predict the range interval in which the toxicity endpoint data is located, and evaluation is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chemical toxicity assessment, and particularly to a method for constructing a chemical toxicity assessment model, an assessment method, and related apparatus. Background Technology

[0002] In recent decades, the amount of chemicals and their byproducts produced by human activities has increased rapidly. Currently, the number of chemical substances registered with CAS exceeds 193 million, and more than 395,000 chemicals are circulating in the global market. The large-scale production and use of chemicals inevitably leads to the release of large quantities of chemicals into the environment during their life cycle, which will have a significant negative impact on human health and environmental safety. Therefore, obtaining information on the toxicity of chemicals has become crucial for the control and regulation of chemicals.

[0003] Traditional in vivo animal testing methods are time-consuming and labor-intensive, making it difficult to quickly assess the toxicity of chemicals. Although quantitative structure-activity models based on machine learning are widely used to predict various toxicity endpoints of chemicals, the poor availability of currently available chemical toxicity data hinders its effectiveness in machine learning modeling and toxicity assessment. Summary of the Invention

[0004] The main objective of this invention is to propose a method for constructing a chemical toxicity assessment model, an assessment method, and related apparatus, aiming to solve the problem of poor availability of currently available chemical toxicity data, and to quickly and accurately predict the toxicity of the chemical to be assessed.

[0005] To achieve the above objectives, in a first aspect, the present invention proposes a method for constructing a chemical toxicity assessment model, comprising:

[0006] Obtain a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals;

[0007] Based on the molecular fingerprints of each chemical in the original chemical sample set, the molecular structural similarity of each chemical relative to other chemicals is determined, wherein the molecular fingerprints of each chemical are of the same type.

[0008] Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened.

[0009] Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively;

[0010] Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined.

[0011] The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set;

[0012] A target toxicity classification model is constructed based on the second target chemical sample set.

[0013] Secondly, the present invention also proposes a chemical toxicity assessment model construction device, comprising:

[0014] The first acquisition module is used to acquire a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals.

[0015] The first processing module is used to determine the molecular structural similarity of each chemical relative to other chemicals based on the molecular fingerprints of each chemical in the original chemical sample set, wherein the molecular fingerprints of each chemical are of the same type.

[0016] Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened.

[0017] Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively;

[0018] Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined.

[0019] The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set;

[0020] A target toxicity classification model is constructed based on the second target chemical sample set.

[0021] Thirdly, the present invention also proposes a chemical toxicity assessment method, which utilizes the toxicity assessment model constructed by the method described in the first aspect, the toxicity assessment method comprising:

[0022] Obtain the chemicals to be evaluated;

[0023] The molecular fingerprint of the chemical to be evaluated and the application domain are used to determine the applicable evaluation model for the chemical to be evaluated.

[0024] Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

[0025] Fourthly, the present invention also proposes a chemical toxicity assessment device, comprising a toxicity assessment model constructed using the method described in the first aspect, and

[0026] The second acquisition module is used to acquire the chemicals to be evaluated.

[0027] The second processing module is used to determine the appropriate evaluation model for the chemical to be evaluated based on the molecular fingerprint of the chemical to be evaluated and the application domain.

[0028] Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

[0029] Fifthly, the present invention also proposes a medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first or third aspect.

[0030] In a sixth aspect, the present invention also proposes a computing device comprising a processor for executing a computer program stored in a memory to implement the method described in the first or third aspect.

[0031] The technical solution of this invention establishes a target toxicity regression model and a target toxicity classification model based on the molecular structure similarity of chemicals, as well as corresponding application domains. For chemicals that meet the application domain, their molecular structures can maintain a high degree of molecular structure similarity with the chemical samples used to construct the target toxicity regression model. Therefore, the target toxicity regression model can obtain toxicity endpoint data more accurately. For chemicals that do not meet the application domain, it indicates that their molecular structure characteristics are more special. The target classification model can also roughly determine the range of their toxicity endpoint data. This makes it more convenient and faster to assess the toxicity of chemicals. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0033] Figure 1 This is a flowchart illustrating the steps of a chemical toxicity assessment model construction method in one embodiment of the present invention;

[0034] Figure 2 This is a flowchart illustrating the steps of a chemical toxicity assessment method in one embodiment of the present invention;

[0035] Figure 3 This is a schematic diagram of the chemical toxicity assessment model construction device in one embodiment of this application;

[0036] Figure 4 This is a schematic diagram of the chemical toxicity assessment device in one embodiment of this application;

[0037] Figure 5 This is a schematic diagram of the structure of a computer storage medium according to an embodiment of the present invention;

[0038] Figure 6 This is a schematic diagram of the structure of a computing device according to an embodiment of the present invention.

[0039] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0040] The principles and spirit of the invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement the invention, and are not intended to limit the scope of the invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0041] Those skilled in the art will recognize that embodiments of the present invention can be implemented as an apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0042] According to embodiments of the present invention, a method for constructing a chemical toxicity assessment model, an assessment method, and related apparatus are proposed.

[0043] Exemplary methods

[0044] In this exemplary embodiment, a method for constructing a chemical toxicity assessment model is proposed, such as... Figure 1 As shown, the evaluation model construction method includes the following steps S100-S700:

[0045] Step S100: Obtain a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals.

[0046] In this embodiment, the original chemical sample set can be the toxicity data of currently available chemicals, such as some open-source chemical toxicity datasets. The original chemical sample set includes data on the same type of toxicity endpoints for each chemical. Currently, toxicity endpoints are categorized into median lethal dose (LD50).50 ), teratogenicity index, median lethal concentration (LC50) 50 Maximum no-effect concentration (NOEC), half-maximal effective concentration (EC) 50 This application does not restrict the type of toxicity endpoint data to be used, as long as each chemical in the original chemical sample set includes toxicity endpoint data of the same type. For example, when using the median lethal dose (LD50), each chemical in the original chemical sample set includes the specific value of its corresponding LD50; or, when using the median lethal concentration (LD50), each chemical in the original chemical sample set includes the specific value of its corresponding LD50. This application's embodiments use LD50-type toxicity endpoint data as an example, meaning the original chemical sample set can contain the specific value of the LD50 for each chemical.

[0047] Step S200: Based on the molecular fingerprints of each chemical in the original chemical sample set, determine the molecular structural similarity of each chemical relative to other chemicals, wherein the molecular fingerprints of each chemical are of the same type.

[0048] When compounds in a chemical product have similar molecular structures, they tend to have similar chemical properties, such as roughly equivalent toxicity. Molecular fingerprints can visually represent the molecular structure of a compound. Therefore, the similarity of molecular structures between chemicals can be determined based on their molecular fingerprints. Furthermore, toxicity assessment models can be constructed based on the molecular fingerprints and toxicity endpoint data of each chemical in a chemical sample set. These models can learn the correlation between the molecular fingerprints and toxicity endpoints of chemicals, thereby enabling rapid and accurate assessment of chemical toxicity.

[0049] Commonly used molecular fingerprints include Morgan molecular fingerprints, Avalon molecular fingerprints, topological molecular fingerprints, and atom-pair molecular fingerprints. This application does not restrict the choice of molecular fingerprint; all chemicals can use the same molecular fingerprint, such as all using Morgan molecular fingerprints or all using topological molecular fingerprints, etc.

[0050] Taking Morgan molecular fingerprinting as an example, in this embodiment of the application, based on the original chemical sample set, the main compounds of each chemical can be identified. Therefore, the Morgan molecular fingerprint of each chemical can be calculated using SMILES codes (Simplified molecular input lineentry system). The calculation of the Morgan molecular fingerprint of each chemical can be performed according to the following parameters:

[0051] radius = 2;

[0052] length = 2048.

[0053] Wherein, radius represents the cut-off radius of the Morgan molecular fingerprint, and length represents the length of the Morgan molecular fingerprint. That is, for each chemical, its Morgan molecular fingerprint is calculated with a radius of 2 and a length of 2048.

[0054] After calculating the Morgan molecular fingerprint of each chemical in the original chemical sample set, the molecular structural similarity of each chemical relative to other chemicals can be determined through the following steps S210-S220:

[0055] Step S210: Using the leave-one-out method, calculate the root mean square error of the molecular fingerprint of any chemical in the original chemical sample set relative to the molecular fingerprints of all other chemicals.

[0056] In this embodiment of the application, it is assumed that the original sample set has S(AJ) containing ten chemicals (ten chemicals are only for illustrative purposes and do not mean that the original sample set has only ten samples). Then, the ten chemicals will have ten Morgan molecular fingerprints. Each time, the Morgan molecular fingerprint of a chemical is selected, and the root mean square error is calculated with the Morgan molecular fingerprints of all the remaining chemicals in the original sample set. For example, the root mean square error of the Morgan molecular fingerprint of chemical A relative to the other nine Morgan molecular fingerprints of BJ is calculated, and the root mean square error of the Morgan molecular fingerprint of chemical A relative to the Morgan molecular fingerprint of chemical BJ is obtained. After multiple calculations, the root mean square error of the Morgan molecular fingerprint of each chemical relative to the Morgan molecular fingerprints of other chemicals can be obtained.

[0057] Step S210: Based on the root mean square error of the molecular fingerprint of each chemical relative to the molecular fingerprints of all other chemicals, determine the molecular structural similarity of each chemical relative to the other chemicals.

[0058] In the embodiments of this application, the Morgan molecular fingerprint can reflect the molecular structure of the compounds that make up the chemical. Therefore, the root mean square error between the Morgan molecular fingerprint of any chemical and the Morgan molecular fingerprint of other chemicals can reflect the structural similarity of the chemical to other chemicals. That is, the root mean square error of the Morgan molecular fingerprint of each chemical relative to the Morgan molecular fingerprints of all other chemicals can be used as the molecular structural similarity of the chemical relative to other chemicals.

[0059] Step S300: Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, screen multiple first candidate chemical sample sets.

[0060] In this embodiment of the application, multiple first candidate chemical sample sets can be screened using the following method:

[0061] Based on each first molecular structure similarity threshold in the preset set of first molecular structure similarity thresholds, each first candidate chemical sample set is obtained from the original sample set. The chemical samples in each first candidate chemical sample set have feature similarity relative to all other chemicals, and the similarity is greater than the corresponding first molecular structure similarity threshold.

[0062] The preset set of first molecular structure similarity thresholds can be obtained empirically. Assuming the preset set of first molecular structure similarity thresholds is Z(z1, z2, z3), then z1, z2, and z3 can be used as the first molecular structure similarity thresholds respectively, so as to select from the original chemical sample set.

[0063] For example, firstly, using z1 as the first molecular structure similarity threshold, the similarity of each feature of chemical AJ is compared with z1, and chemicals with similarity greater than z1 are selected to form the first candidate chemical sample set. Assuming that after selection, three first candidate chemical sample sets are obtained by using z1, z2, and z3 as three thresholds, with the ratios S1, S2, and S3, where S1 = (A, B, C), S2 = (D, E, F), and S1 = (G, H, I, J), then a toxicity regression model can be constructed based on the three first candidate chemical sample sets, as shown in the following step S400: Based on the multiple first candidate chemical sample sets, multiple first candidate toxicity regression models are constructed respectively.

[0064] In constructing the candidate toxicity regression model, it can be built based on the Morgan molecular fingerprint of each chemical in each first candidate chemical sample set and the corresponding toxicity endpoint data, thereby establishing a mapping between the Morgan molecular fingerprint and the toxicity endpoint data.

[0065] Step S500: Based on multiple first candidate toxicity regression models, determine the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model.

[0066] In this embodiment of the application, the first target chemical sample set and the target toxicity regression model can be obtained based on the following steps S510-S520, as follows:

[0067] Step S511: Determine the prediction accuracy of each first candidate toxicity regression model based on each first candidate toxicity regression model and the corresponding first candidate chemical sample set.

[0068] In this embodiment, it is assumed that the three first candidate toxicity regression models constructed based on the above three first candidate chemical sample sets S1, S2, and S3 are M1, M2, and M3, respectively. Taking the first candidate regression model M1 as an example, when calculating the prediction accuracy of the first candidate regression model M1, the Morgan molecular fingerprints of the first candidate chemical sample set S1(A, B, C) corresponding to M1 are respectively input into M1 to obtain the predicted values ​​of each toxicity endpoint data. Then, the prediction accuracy of the first candidate regression model M1 can be calculated using the following formula (1):

[0069]

[0070] Where R represents the prediction accuracy of the corresponding first candidate regression model, y exp This represents the true toxicity endpoint data for each chemical in the first candidate chemical sample set corresponding to the first candidate regression model, y pred This represents the predicted toxicity endpoint data obtained through the first candidate chemical sample set corresponding to the first candidate regression model. This represents the average value of the true toxicity endpoint data for each chemical in the first candidate chemical sample set corresponding to the first candidate regression model.

[0071] Therefore, based on S1, S2, S3, and the corresponding M1, M2, M3, the prediction accuracy of the three first candidate regression models can be calculated respectively.

[0072] Step S512: Select the first candidate toxicity regression model with the highest prediction accuracy as the target toxicity regression model, and select the first candidate chemical sample set corresponding to the first candidate toxicity regression model with the highest prediction accuracy as the first target chemical sample set.

[0073] In this embodiment of the application, the one with the highest prediction accuracy among the three first candidate regression models is selected as the target candidate toxicity regression model. Assuming that M3 has the highest prediction accuracy, M3 is selected as the target candidate toxicity regression model, while S3 is selected as the first target chemical sample set.

[0074] Steps S511-S512 detailed the method for obtaining the target candidate toxicity regression model, i.e., the first target chemical sample set. The following steps S521-S525 will then be used to determine the application domain of the target toxicity regression model:

[0075] Step S521: Based on the molecular fingerprint of each chemical in the first target chemical sample set, calculate the feature similarity index of each chemical sample in the first target chemical sample set.

[0076] In this embodiment, firstly, based on the Morgan molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample relative to the other chemical samples can be calculated. Specifically, it can be calculated using the following formula (2).

[0077]

[0078] Among them, T C (X1,X2) represents the feature similarity index of chemical X1 relative to X2, a and b represent the number of structural features that chemical X1 and X2 each have in their Morgan molecular fingerprints, and c represents the number of the same structural features that chemical X1 and X2 have in their Morgan molecular fingerprints.

[0079] Therefore, for the first target chemical sample set S3(G, H, I, J), the following calculations are required:

[0080] The similarity index T between chemical G and H characteristics G1 The feature similarity index T relative to I G2 The characteristic index T relative to J G3 ;

[0081] The similarity index T of chemical H relative to G H1 The feature similarity index T relative to I H2 The characteristic index T relative to J H3 ;

[0082] Chemical I vs. G: Feature Similarity Index T I1 The feature similarity index T relative to H I2 The characteristic index T relative to J I3 ;

[0083] Chemical J has a similarity index T relative to G. J1 The feature similarity index T relative to H J2 The characteristic index T relative to I J3 .

[0084] Then, the largest feature similarity index of each chemical sample relative to all other chemical samples is selected as the feature similarity index of that chemical sample.

[0085] Assuming,

[0086] T G1 >T G2 >T G3 ,

[0087] T H1 >T H2 >TH3 ,

[0088] T I1 >T I2 >T I3 ,

[0089] T J1 >T J2 >T J3 ;

[0090] Therefore, the similarity index of chemical G is T. G1 The similarity index of chemical H characteristics is T. H1 The chemical I feature similarity index is T. I1 The chemical J feature similarity index is T. J1 .

[0091] In another embodiment, the average of the feature similarity indices of each chemical sample relative to the other chemical samples can also be selected as the feature similarity index of the chemical sample.

[0092] For example, for chemical G, T can be selected. G1 T G2 T G3 The average value is used as its feature similarity index;

[0093] For chemical H, T can be selected. H1 T H2 T H3 The average value is used as its feature similarity index;

[0094] For chemical I, T can be selected. I1 T I2 T I3 The average value is used as its feature similarity index;

[0095] For chemical J, T can be selected. J1 T J2 T J3 The average value is used as its feature similarity index.

[0096] Step S522: Based on the feature similarity index of each chemical sample in the first target chemical sample set, and according to the preset feature similarity index gradient, select multiple second candidate chemical sample sets from the first target chemical sample set respectively.

[0097] In this embodiment of the application, the preset feature similarity index gradient is, for example, 0.1. Then, according to the gradient of 0.1, it gradually increases from 0 to 1, and selects a plurality of second candidate chemical sample sets from the first target chemical sample set S3 respectively.

[0098] It should be noted that the preset feature similarity index gradient can be set according to the computational cost; a smaller gradient results in a higher computational cost, and vice versa. Each increment in the gradient may result in the same or different second candidate chemical sample set. For example, increasing the preset feature similarity index from 0.5 to 0.6 may result in the same or different second candidate chemical sample set selected from S2. For each preset feature similarity index, chemical samples with a feature similarity index greater than that preset index are selected from S2.

[0099] Based on each preset feature similarity index gradient, increasing from 0 to 1, multiple second candidate chemical sample sets can be obtained from S2.

[0100] Step S523: Based on the multiple second candidate chemical sample sets, construct multiple second candidate toxicity regression models respectively.

[0101] In this embodiment of the application, in step S522, multiple second candidate chemical sample sets can be obtained from S2. In step S523, each second candidate toxicity regression model can be constructed based on the multiple second candidate chemical sample sets. The specific construction process is described in step S400, and will not be repeated here.

[0102] Step S524: Calculate the prediction accuracy of the multiple second candidate toxicity regression models respectively.

[0103] After constructing multiple second candidate toxicity regression models, the prediction accuracy of each second candidate toxicity regression model can be calculated based on step 511, which will not be elaborated here.

[0104] Step S525: Select the feature similarity index corresponding to the second candidate toxicity regression model with the highest prediction accuracy as the application domain.

[0105] In the embodiments of this application, after obtaining the prediction accuracy of each second candidate toxicity regression model, the feature similarity index corresponding to the second candidate toxicity regression model with the highest prediction accuracy can be selected as the application domain.

[0106] For example, if the feature similarity index increases to 0.8, the second candidate toxicity regression model pair has the highest prediction accuracy, then the application domain is 0.8 to 1.

[0107] It should be noted that when two identical values ​​exist for the highest prediction accuracy, the one with the larger feature similarity index should be chosen as the application domain. For example, if the feature similarity index increases from 0.8 to 0.9 without changing the corresponding second candidate chemical sample set, then the two second candidate toxicity regression models corresponding to the two second candidate chemical sample sets are identical, and their prediction accuracies are also identical. When the prediction accuracies of these two second candidate toxicity regression models are both at their maximum values, the lower feature similarity index of 0.8 should be chosen as the application domain. When prediction accuracies are the same, choosing the lower feature similarity index as the application domain can broaden the scope of application, provided that prediction accuracy can be guaranteed.

[0108] Furthermore, in step S521, two methods for calculating the chemical similarity index are introduced: The largest or average feature similarity index among all chemical samples relative to other chemical samples is selected as the feature similarity index of that chemical sample. When the average value is selected as the application domain, the entire first target chemical sample set is considered to determine the application domain. When the maximum value is selected as the application domain, the application domain is selected based on the highest molecular structure similarity. Therefore, when the largest value is selected as the application domain, its applicability is smaller than when the average value is selected, but the prediction accuracy is higher. Conversely, when the average value is selected as the application domain, its applicability is relatively larger, but the prediction accuracy is lower than when the maximum value is selected.

[0109] Once the application domain is obtained, this application domain is the scope of application of the target toxicity regression model. That is, for any chemical, when its feature similarity index is greater than that of the application domain, the target toxicity regression model can be used to predict toxicity endpoint data.

[0110] Step S600: Remove the first target chemical sample set from the original chemical sample set to obtain the second target chemical sample set.

[0111] Through steps S100-S600, a target toxicity regression model and its application domain are constructed. The toxicity endpoint can be predicted for any chemical whose feature similarity index is greater than that of the application domain. For other chemicals whose feature similarity index does not meet the requirements of the application domain, the following classification model can be used for prediction.

[0112] First, the first target chemical sample set is removed from the original chemical sample set, that is, S3 is removed from S, to obtain the second target chemical sample set S4(AF).

[0113] Step S700: Construct a target toxicity classification model based on the second target chemical sample set.

[0114] In step S600, the second target chemical sample set is determined as S4(AF). In step S700, a target toxicity classification model can be constructed based on the second target chemical sample set as S4(AF).

[0115] Among them, the target toxicity classification model can establish a mapping between the Morgan molecular fingerprint of a chemical and the data range of toxic endpoints. For chemicals whose feature similarity index does not meet the application domain, after inputting their Morgan molecular fingerprint into the toxicity classification model, the range in which the toxic endpoint data of the chemical is located can be obtained.

[0116] In this embodiment, a target toxicity regression model and a target toxicity classification model are established based on the molecular structure similarity of chemicals, along with corresponding application domains. For chemicals that meet the application domain, their molecular structures maintain a high degree of similarity to the chemical samples used to construct the target toxicity regression model. Therefore, the target toxicity regression model can accurately obtain toxicity endpoint data. For chemicals that do not meet the application domain, it indicates that their molecular structure characteristics are quite unique. The target classification model can also roughly determine the range of their toxicity endpoint data. This makes it more convenient and faster to assess the toxicity of chemicals.

[0117] This application also proposes a chemical toxicity assessment method, utilizing the toxicity assessment model constructed using the methods described in all the above embodiments, such as... Figure 2 As shown, the toxicity assessment method includes the following steps:

[0118] Step S810: Obtain the chemical to be evaluated.

[0119] In the embodiments of this application, after obtaining the chemical to be evaluated, its molecular fingerprint can be calculated.

[0120] Step S820: Determine the appropriate evaluation model for the chemical to be evaluated based on the molecular fingerprint of the chemical to be evaluated and the application domain.

[0121] In this embodiment of the application, the feature similarity index is first determined based on the molecular fingerprint. Then, the feature similarity index of the chemical is calculated using the above formula (2). The feature similarity index of the chemical and each chemical in the first target chemical sample set (the first target chemical sample set S3 corresponding to the target toxicity regression model) is calculated respectively, and the largest one is selected as the feature similarity index of the chemical to be evaluated.

[0122] Step S830: Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, obtain the toxicity endpoint data of the chemical to be evaluated.

[0123] In step S820, after obtaining the feature similarity index of the chemical to be evaluated, it is compared with the application domain. Assuming that the application domain of the target toxicity regression model is 0.8-1, if it is within the application domain, the target toxicity regression model is selected to predict the toxicity endpoint data of the chemical to be evaluated; if it is not within the application domain, the target classification model is selected to predict the toxicity endpoint data of the chemical to be evaluated.

[0124] The chemical toxicity assessment method in this application embodiment utilizes all the methods for constructing toxicity assessment models described above. Therefore, it has at least all the beneficial effects of the methods for constructing toxicity assessment models described above. Thus, by using this toxicity assessment method, the toxicity endpoint data of the chemical to be assessed can be determined quickly and accurately.

[0125] Exemplary device

[0126] After introducing the method of exemplary embodiments of the present invention, the exemplary toxicity assessment model construction apparatus of the present invention will be described next, such as... Figure 3 As shown, the chemical toxicity assessment model construction device 100 includes:

[0127] The first acquisition module 110 is used to acquire a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals.

[0128] The first processing module 120 is used to determine the molecular structural similarity of each chemical relative to other chemicals based on the molecular fingerprints of each chemical in the original chemical sample set, wherein the molecular fingerprints of each chemical are of the same type.

[0129] Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened;

[0130] Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively;

[0131] Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined.

[0132] The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set;

[0133] A target toxicity classification model is constructed based on the second target chemical sample set.

[0134] In this embodiment of the application, the first processing module 120 is further configured to:

[0135] Using the leave-one-out method, the root mean square error of the molecular fingerprint of any chemical in the original chemical sample set relative to the molecular fingerprints of all other chemicals is calculated.

[0136] The molecular structural similarity of each chemical relative to the other chemicals is determined based on the root mean square error of the molecular fingerprint of each chemical relative to the molecular fingerprints of all other chemicals.

[0137] In this embodiment of the application, the first processing module 120 is further configured to:

[0138] Based on each first molecular structure similarity threshold in the preset first molecular structure similarity threshold set, each first candidate chemical sample set is obtained by screening from the original sample set, wherein the chemical samples in each first candidate chemical sample set are similar in features to all other chemicals and are all greater than the first molecular structure similarity threshold corresponding to each first candidate chemical sample set.

[0139] In this embodiment of the application, the first processing module 120 is further configured to:

[0140] Based on each first candidate toxicity regression model and the corresponding first candidate chemical sample set, the prediction accuracy of each first candidate toxicity regression model is determined.

[0141] The first candidate toxicity regression model with the highest prediction accuracy was selected as the target toxicity regression model, and the first candidate chemical sample set corresponding to the first candidate toxicity regression model with the highest prediction accuracy was selected as the first target chemical sample set.

[0142] In this embodiment of the application, the first processing module 120 is further configured to determine the application domain of the target toxicity regression model in the following manner:

[0143] Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample in the first target chemical sample set is calculated respectively.

[0144] Based on the feature similarity index of each chemical sample in the first target chemical sample set, and according to the preset feature similarity index gradient, multiple second candidate chemical sample sets are obtained from the first target chemical sample set respectively.

[0145] Based on the aforementioned sample sets of multiple second candidate chemicals, multiple second candidate toxicity regression models are constructed respectively.

[0146] Calculate the prediction accuracy of each of the multiple second candidate toxicity regression models;

[0147] The feature similarity index corresponding to the second candidate toxicity regression model with the highest prediction accuracy is selected as the application domain.

[0148] In this embodiment of the application, the step of calculating the feature similarity index of each chemical sample in the first target chemical sample set based on the molecular fingerprint of each chemical in the first target chemical sample set includes:

[0149] Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample relative to the other chemical samples is calculated respectively.

[0150] The highest feature similarity index among all chemical samples relative to other chemical samples is selected as the feature similarity index of that chemical sample; or

[0151] The average of the feature similarity indices of each chemical sample relative to all other chemical samples is selected as the feature similarity index of that chemical sample.

[0152] The chemical toxicity assessment model construction apparatus in this application embodiment utilizes all the methods for constructing toxicity assessment models described above, and therefore has at least all the beneficial effects of the methods for constructing toxicity assessment models described above, which will not be elaborated here.

[0153] This application also proposes a chemical toxicity assessment device 200, such as... Figure 4 As shown, the chemical toxicity assessment device 200 includes a toxicity assessment model constructed using the toxicity assessment modeling methods described in all the above embodiments, and

[0154] The second acquisition module 210 is used to acquire the chemical to be evaluated;

[0155] The second processing module 220 is used to determine the appropriate evaluation model for the chemical to be evaluated based on the molecular fingerprint of the chemical to be evaluated and the application domain.

[0156] Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

[0157] The chemical toxicity assessment device in this application embodiment utilizes the toxicity assessment model constructed by all the toxicity assessment model construction methods described above, and therefore has at least all the beneficial effects of the above-described methods for constructing toxicity assessment models, which will not be elaborated here.

[0158] Exemplary media

[0159] After introducing the methods and apparatus of exemplary embodiments of the present invention, the following references are made. Figure 5 A computer-readable storage medium according to an exemplary embodiment of the present invention will be described.

[0160] Please refer to Figure 5 The computer-readable storage medium shown is an optical disc 70, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it performs the steps described in the above method implementation, such as: obtaining a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals.

[0161] Based on the molecular fingerprints of each chemical in the original chemical sample set, the molecular structural similarity of each chemical relative to other chemicals is determined, wherein the molecular fingerprints of each chemical are of the same type.

[0162] Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened;

[0163] Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively;

[0164] Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined.

[0165] The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set;

[0166] A target toxicity classification model is constructed based on the second target chemical sample set.

[0167] or,

[0168] Obtain the chemicals to be evaluated;

[0169] The molecular fingerprint of the chemical to be evaluated and the application domain are used to determine the applicable evaluation model for the chemical to be evaluated.

[0170] Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

[0171] The specific implementation methods for each step will not be repeated here.

[0172] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.

[0173] Exemplary computing device

[0174] After introducing the methods, apparatus, and media of exemplary embodiments of the present invention, the following references are made. Figure 6 A computing device 80 according to an exemplary embodiment of the present invention will be described.

[0175] Figure 6 A block diagram is shown of an exemplary computing device 80 suitable for implementing embodiments of the present invention, which may be a computer system or a server. Figure 6 The computing device 80 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.

[0176] like Figure 6 As shown, the components of the computing device 80 may include, but are not limited to: one or more processors or processing units 801, system memory 802, and bus 803 connecting different system components (including system memory 802 and processing unit 801).

[0177] The computing device 80 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computing device 80, including volatile and non-volatile media, removable and non-removable media.

[0178] System memory 802 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 8021 and / or cache memory 8022. Computing device 70 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 8023 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 6 Not shown in the image (usually referred to as a "hard drive"). Although not shown in Figure 6The diagram illustrates that disk drives for reading and writing to removable non-volatile disks (e.g., "floppy disks") and optical disc drives for reading and writing to removable non-volatile optical discs (e.g., CD-ROMs, DVD-ROMs, or other optical media) can be provided. In these cases, each drive can be connected to bus 803 via one or more data media interfaces. System memory 802 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0179] A program / utility 8025 having a set (at least one) of program modules 8024 may be stored, for example, in system memory 802, and such program modules 8024 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment. Program modules 8024 typically perform the functions and / or methods described in the embodiments of the present invention.

[0180] The computing device 80 can also communicate with one or more external devices 804 (such as a keyboard, pointing device, display, etc.). This communication can be performed through an input / output (I / O) interface. Furthermore, the computing device 80 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 806. Figure 6 As shown, network adapter 806 communicates with other modules of computing device 80 (such as processing unit 801) via bus 803. It should be understood that, although... Figure 6 As not shown, it can be used in conjunction with computing device 80 with other hardware and / or software modules.

[0181] The processing unit 801 executes various functional applications and data processing by running programs stored in the system memory 802, such as acquiring a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals.

[0182] Based on the molecular fingerprints of each chemical in the original chemical sample set, the molecular structural similarity of each chemical relative to other chemicals is determined, wherein the molecular fingerprints of each chemical are of the same type.

[0183] Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened;

[0184] Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively;

[0185] Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined.

[0186] The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set;

[0187] A target toxicity classification model is constructed based on the second target chemical sample set.

[0188] or,

[0189] Obtain the chemicals to be evaluated;

[0190] The molecular fingerprint of the chemical to be evaluated and the application domain are used to determine the applicable evaluation model for the chemical to be evaluated.

[0191] Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

[0192] The specific implementation methods of each step will not be repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the GNSS-R-based soil moisture inversion device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0193] Furthermore, although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0194] While the spirit and principles of the invention have been described with reference to several specific embodiments, it should be understood that the invention is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for ease of description. The invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

[0195] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

[0196] Based on the above description, the embodiments of this application provide at least the following technical solutions, but are not limited thereto:

[0197] 1. A method for constructing a chemical toxicity assessment model, comprising:

[0198] Obtain a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals;

[0199] Based on the molecular fingerprints of each chemical in the original chemical sample set, the molecular structural similarity of each chemical relative to other chemicals is determined, wherein the molecular fingerprints of each chemical are of the same type.

[0200] Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened;

[0201] Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively;

[0202] Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined.

[0203] The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set;

[0204] A target toxicity classification model is constructed based on the second target chemical sample set.

[0205] 2. The chemical toxicity assessment model construction method as described in technical solution 1, wherein determining the molecular structural similarity of each chemical relative to other chemicals based on the molecular fingerprints of each chemical in the original chemical sample set includes:

[0206] Using the leave-one-out method, the root mean square error of the molecular fingerprint of any chemical in the original chemical sample set relative to the molecular fingerprints of all other chemicals is calculated.

[0207] The molecular structural similarity of each chemical relative to the other chemicals is determined based on the root mean square error of the molecular fingerprint of each chemical relative to the molecular fingerprints of all other chemicals.

[0208] 3. The chemical toxicity assessment model construction method as described in technical solution 1 or 2, wherein the step of screening multiple first candidate chemical sample sets based on the molecular structure similarity and a preset first molecular structure similarity threshold set includes:

[0209] Based on each first molecular structure similarity threshold in the preset first molecular structure similarity threshold set, each first candidate chemical sample set is obtained by screening from the original sample set, wherein the molecular structure similarity of the chemical samples in each first candidate chemical sample set relative to all other chemicals is greater than the first molecular structure similarity threshold corresponding to each first candidate chemical sample set.

[0210] 4. The method for constructing a chemical toxicity assessment model as described in any one of technical solutions 1-3, wherein determining the first target chemical sample set and the target toxicity regression model based on multiple first candidate toxicity regression models includes:

[0211] Based on each first candidate toxicity regression model and the corresponding first candidate chemical sample set, the prediction accuracy of each first candidate toxicity regression model is determined.

[0212] The first candidate toxicity regression model with the highest prediction accuracy was selected as the target toxicity regression model, and the first candidate chemical sample set corresponding to the first candidate toxicity regression model with the highest prediction accuracy was selected as the first target chemical sample set.

[0213] 5. The method for constructing a chemical toxicity assessment model as described in any one of technical solutions 1-4, wherein the application domain of the target toxicity regression model is determined in the following manner:

[0214] Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample in the first target chemical sample set is calculated respectively.

[0215] Based on the feature similarity index of each chemical sample in the first target chemical sample set, and according to the preset feature similarity index gradient, multiple second candidate chemical sample sets are obtained from the first target chemical sample set respectively;

[0216] Based on the aforementioned sample sets of multiple second candidate chemicals, multiple second candidate toxicity regression models are constructed respectively.

[0217] Calculate the prediction accuracy of each of the multiple second candidate toxicity regression models;

[0218] The feature similarity index corresponding to the second candidate toxicity regression model with the highest prediction accuracy is selected as the application domain.

[0219] 6. The chemical toxicity assessment model construction method as described in any one of technical solutions 1-5, wherein calculating the feature similarity index of each chemical sample in the first target chemical sample set based on the molecular fingerprint of each chemical in the first target chemical sample set includes:

[0220] Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample relative to the other chemical samples is calculated respectively.

[0221] The feature similarity index of each chemical sample relative to all other chemical samples is selected as the feature similarity index of that chemical sample; or,

[0222] The average of the feature similarity indices of each chemical sample relative to all other chemical samples is selected as the feature similarity index of that chemical sample.

[0223] 7. A device for constructing a chemical toxicity assessment model, comprising:

[0224] The first acquisition module is used to acquire a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals.

[0225] The first processing module is used to determine the molecular structural similarity of each chemical relative to other chemicals based on the molecular fingerprints of each chemical in the original chemical sample set, wherein the molecular fingerprints of each chemical are of the same type.

[0226] Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened;

[0227] Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively;

[0228] Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined.

[0229] The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set;

[0230] A target toxicity classification model is constructed based on the second target chemical sample set.

[0231] 8. The chemical toxicity assessment model construction apparatus as described in technical solution 7, wherein the first processing module is further configured to:

[0232] Using the leave-one-out method, the root mean square error of the molecular fingerprint of any chemical in the original chemical sample set relative to the molecular fingerprints of all other chemicals is calculated.

[0233] The molecular structural similarity of each chemical relative to the other chemicals is determined based on the root mean square error of the molecular fingerprint of each chemical relative to the molecular fingerprints of all other chemicals.

[0234] 9. The chemical toxicity assessment model construction apparatus as described in technical solution 7 or 8, wherein the first processing module is further configured to:

[0235] Based on each first molecular structure similarity threshold in the preset first molecular structure similarity threshold set, each first candidate chemical sample set is obtained by screening from the original sample set, wherein the hierarchical structural similarity of the chemical samples in each first candidate chemical sample set relative to all other chemicals is greater than the first molecular structure similarity threshold corresponding to each first candidate chemical sample set.

[0236] 10. The chemical toxicity assessment model construction apparatus as described in any one of technical solutions 7-9, wherein the first processing module is further configured to:

[0237] Based on each first candidate toxicity regression model and the corresponding first candidate chemical sample set, the prediction accuracy of each first candidate toxicity regression model is determined.

[0238] The first candidate toxicity regression model with the highest prediction accuracy was selected as the target toxicity regression model, and the first candidate chemical sample set corresponding to the first candidate toxicity regression model with the highest prediction accuracy was selected as the first target chemical sample set.

[0239] 11. The chemical toxicity assessment model construction apparatus as described in any one of technical solutions 7-10, wherein the first processing module is further configured to determine the application domain of the target toxicity regression model in the following manner:

[0240] Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample in the first target chemical sample set is calculated respectively.

[0241] Based on the feature similarity index of each chemical sample in the first target chemical sample set, and according to the preset feature similarity index gradient, multiple second candidate chemical sample sets are obtained from the first target chemical sample set respectively.

[0242] Based on the aforementioned sample sets of multiple second candidate chemicals, multiple second candidate toxicity regression models are constructed respectively.

[0243] Calculate the prediction accuracy of each of the multiple second candidate toxicity regression models;

[0244] The feature similarity index corresponding to the second candidate toxicity regression model with the highest prediction accuracy is selected as the application domain.

[0245] 12. The chemical toxicity assessment model construction apparatus as described in any one of technical solutions 7-11, wherein calculating the feature similarity index of each chemical sample in the first target chemical sample set based on the molecular fingerprint of each chemical in the first target chemical sample set includes:

[0246] Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample relative to the other chemical samples is calculated respectively.

[0247] The highest feature similarity index among all other chemical samples is selected as the feature similarity index for that chemical sample; or,

[0248] The average of the feature similarity indices of each chemical sample relative to all other chemical samples is selected as the feature similarity index of that chemical sample.

[0249] 13. A method for assessing the toxicity of a chemical, comprising using a toxicity assessment model constructed by the method described in any one of technical solutions 1-7, wherein the toxicity assessment method includes:

[0250] Obtain the chemicals to be evaluated;

[0251] The molecular fingerprint of the chemical to be evaluated and the application domain are used to determine the applicable evaluation model for the chemical to be evaluated.

[0252] Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

[0253] 14. A chemical toxicity assessment device, comprising a toxicity assessment model constructed using the method described in any one of technical solutions 1-7, and

[0254] The second acquisition module is used to acquire the chemicals to be evaluated;

[0255] The second processing module is used to determine the appropriate evaluation model for the chemical to be evaluated based on the molecular fingerprint of the chemical to be evaluated and the application domain.

[0256] Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

[0257] 15. A medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of technical solutions 1-6 or 13.

[0258] 16. A computing device comprising a processor for executing a computer program stored in a memory to implement the method as described in any one of claims 1-6 or claim 13.

Claims

1. A method for constructing a chemical toxicity assessment model, comprising: Obtain a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals; Based on the molecular fingerprints of each chemical in the original chemical sample set, the molecular structural similarity of each chemical relative to other chemicals is determined, wherein the molecular fingerprints of each chemical are of the same type. Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened; Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively; Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined. The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set; A target toxicity classification model is constructed based on the second target chemical sample set; The process of determining the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model based on multiple first candidate toxicity regression models includes: Based on each first candidate toxicity regression model and the corresponding first candidate chemical sample set, the prediction accuracy of each first candidate toxicity regression model is determined. The first candidate toxicity regression model with the highest prediction accuracy was selected as the target toxicity regression model, and the first candidate chemical sample set corresponding to the first candidate toxicity regression model with the highest prediction accuracy was selected as the first target chemical sample set. Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample in the first target chemical sample set is calculated respectively. Based on the feature similarity index of each chemical sample in the first target chemical sample set, and according to the preset feature similarity index gradient, multiple second candidate chemical sample sets are obtained from the first target chemical sample set respectively. Based on the aforementioned sample sets of multiple second candidate chemicals, multiple second candidate toxicity regression models are constructed respectively. Calculate the prediction accuracy of each of the multiple second candidate toxicity regression models; The feature similarity index corresponding to the second candidate toxicity regression model with the highest prediction accuracy is selected as the application domain; The prediction accuracy of both the first candidate toxicity regression model and the second candidate toxicity regression model is determined based on the following formula: Where R represents the prediction accuracy of the candidate regression model. This represents the true toxicity endpoint data for each chemical in the candidate chemical sample set corresponding to the candidate regression model. This represents the predicted toxicity endpoint data obtained through the candidate regression model for each chemical in the candidate chemical sample set corresponding to the candidate regression model. This represents the average value of the true toxicity endpoint data for each chemical in the candidate chemical sample set corresponding to the candidate regression model.

2. The method for constructing a chemical toxicity assessment model as described in claim 1, wherein determining the molecular structural similarity of each chemical relative to other chemicals based on the molecular fingerprints of each chemical in the original chemical sample set includes: Using the leave-one-out method, the root mean square error of the molecular fingerprint of any chemical in the original chemical sample set relative to the molecular fingerprints of all other chemicals is calculated. The molecular structural similarity of each chemical relative to the other chemicals is determined based on the root mean square error of the molecular fingerprint of each chemical relative to the molecular fingerprints of all other chemicals.

3. The chemical toxicity assessment model construction method as described in claim 1, wherein the step of screening multiple first candidate chemical sample sets based on the molecular structure similarity and a preset first molecular structure similarity threshold set includes: Based on each first molecular structure similarity threshold in the preset first molecular structure similarity threshold set, each first candidate chemical sample set is obtained by screening from the original chemical sample set, wherein the molecular structure similarity of the chemical samples in each first candidate chemical sample set relative to all other chemicals is greater than the first molecular structure similarity threshold corresponding to each first candidate chemical sample set.

4. A device for constructing a chemical toxicity assessment model, comprising: The first acquisition module is used to acquire a raw chemical sample set, which includes data on the same type of toxicity endpoints for multiple chemicals. The first processing module is used to determine the molecular structural similarity of each chemical relative to other chemicals based on the molecular fingerprints of each chemical in the original chemical sample set, wherein the molecular fingerprints of each chemical are of the same type. Based on the molecular structure similarity and the preset first molecular structure similarity threshold set, multiple first candidate chemical sample sets are screened; Based on the aforementioned sample sets of multiple first candidate chemicals, multiple first candidate toxicity regression models are constructed respectively; Based on multiple first candidate toxicity regression models, the first target chemical sample set, the target toxicity regression model, and the application domain of the target toxicity regression model are determined. The first target chemical sample set is removed from the original chemical sample set to obtain the second target chemical sample set; A target toxicity classification model is constructed based on the second target chemical sample set; The first processing module is further configured to: determine the prediction accuracy of each first candidate toxicity regression model based on each first candidate toxicity regression model and the corresponding first candidate chemical sample set; The first candidate toxicity regression model with the highest prediction accuracy was selected as the target toxicity regression model, and the first candidate chemical sample set corresponding to the first candidate toxicity regression model with the highest prediction accuracy was selected as the first target chemical sample set. Based on the molecular fingerprint of each chemical in the first target chemical sample set, the feature similarity index of each chemical sample in the first target chemical sample set is calculated respectively. Based on the feature similarity index of each chemical sample in the first target chemical sample set, and according to the preset feature similarity index gradient, multiple second candidate chemical sample sets are obtained from the first target chemical sample set respectively. Based on the aforementioned sample sets of multiple second candidate chemicals, multiple second candidate toxicity regression models are constructed respectively. Calculate the prediction accuracy of each of the multiple second candidate toxicity regression models; The feature similarity index corresponding to the second candidate toxicity regression model with the highest prediction accuracy is selected as the application domain; The first processing module is configured to determine the prediction accuracy of the first candidate toxicity regression model and the second candidate toxicity regression model based on the following formula: Where R represents the prediction accuracy of the candidate regression model. This represents the true toxicity endpoint data for each chemical in the candidate chemical sample set corresponding to the candidate regression model. This represents the predicted toxicity endpoint data obtained through the candidate regression model for each chemical in the candidate chemical sample set corresponding to the candidate regression model. This represents the average value of the true toxicity endpoint data for each chemical in the candidate chemical sample set corresponding to the candidate regression model.

5. The chemical toxicity assessment model construction apparatus as described in claim 4, wherein the first processing module is further configured to: Using the leave-one-out method, the root mean square error of the molecular fingerprint of any chemical in the original chemical sample set relative to the molecular fingerprints of all other chemicals is calculated. The molecular structural similarity of each chemical relative to the other chemicals is determined based on the root mean square error of the molecular fingerprint of each chemical relative to the molecular fingerprints of all other chemicals.

6. The chemical toxicity assessment model construction apparatus as described in claim 4, wherein the first processing module is further configured to: Based on each first molecular structure similarity threshold in the preset first molecular structure similarity threshold set, each first candidate chemical sample set is obtained by screening from the original sample set, wherein the hierarchical structural similarity of the chemical samples in each first candidate chemical sample set relative to all other chemicals is greater than the first molecular structure similarity threshold corresponding to each first candidate chemical sample set.

7. A method for assessing the toxicity of a chemical, comprising using a toxicity assessment model constructed according to any one of claims 1-3, the toxicity assessment method comprising: Obtain the chemicals to be evaluated; The molecular fingerprint of the chemical to be evaluated and the application domain are used to determine the applicable evaluation model for the chemical to be evaluated. Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

8. A chemical toxicity assessment device, comprising a toxicity assessment model constructed using the method described in any one of claims 1-3, and The second acquisition module is used to acquire the chemicals to be evaluated; The second processing module is used to determine the appropriate evaluation model for the chemical to be evaluated based on the molecular fingerprint of the chemical to be evaluated and the application domain. Based on the molecular fingerprint of the chemical to be evaluated and the evaluation model applicable to the chemical to be evaluated, the toxicity endpoint data of the chemical to be evaluated are obtained.

9. A medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-3 or claim 7.

10. A computing device, characterized in that, The computing device includes a processor that executes a computer program stored in a memory to implement the method as claimed in any one of claims 1-3 or 7.

Citation Information

Patent Citations

  • QSAR (Quantitative Structure-Activity Relationships) model constructed based on comprehensive toxicity action mode classification for predicting acute toxicity of organic compound to daphnia magna

    CN105005641A

  • Modeling method and device of compound toxicity prediction model and application of compound toxicity prediction model

    CN110890137A

  • Rapid assessment platform and assessment method for risk characteristics of chemicals with complex components

    CN114220498A

  • Integrated learning method for screening carcinogenic chemicals

    CN114743614A