Method for comparing toxicity of organic compounds

By employing dimensionless parameter scaling and linear regression analysis, the problem of multi-endpoint and cross-group comparisons of organic compound toxicity was solved, enabling efficient and accurate identification and control of new pollutants.

CN116246730BActive Publication Date: 2025-12-05NANJING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310250489.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2025-12-05
Estimated Expiration
2043-03-15

AI Technical Summary

Technical Problem

In existing technologies, quantitative comparisons of the influence of chemical structural characteristics on toxic effects are limited to a specific group and a specific toxic endpoint, making it difficult to achieve multi-endpoint and cross-group comparisons of the toxicity of organic compounds, resulting in inaccurate environmental risk assessments of new pollutants.

Method used

Dimensionless parameter scaling and linear regression analysis were used to classify and scale the different chemical structural characteristics of organic compounds, and a linear regression model was established to achieve multi-endpoint, cross-group toxicity comparison.

Benefits of technology

It enables multi-endpoint, cross-group quantitative comparison of the toxicity of organic compounds, eliminates interference from irrelevant chemical structural features, and accurately identifies and prioritizes the control of high-risk new pollutants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246730B_ABST
    Figure CN116246730B_ABST
Patent Text Reader

Abstract

The application discloses a method for comparing toxicities of organic compounds, and belongs to the technical field of environmental risk assessment. The method is based on different chemical structure characteristics of the organic compounds, and the toxicity values of different compounds are scaled by dimensionless parameters, so that a linear regression model is established, the influence trend and degree of the chemical structure characteristics on the toxicities of the organic compounds are realized by multi-endpoint and cross-group comparison, and an important way is provided for efficiently and accurately identifying new pollutants to be controlled in priority.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of environmental risk assessment, and more particularly relates to a method for comparing toxicities of organic compounds. BACKGROUND

[0002] With the implementation of the pollution prevention and control strategy in China, China's environmental pollution control has achieved preliminary results. For example, traditional pollutants such as nitrogen and phosphorus in water pollution control have been effectively controlled, and new pollutant control has gradually become an important task for extending the depth and broadening the scope of the pollution prevention and control campaign. New pollutants are of various types and complex in toxicity. With the development of detection and identification technology and toxicological analysis technology, hundreds of new pollutants with different chemical structures and multiple toxic effects have been further revealed. Clarifying the complex relationship between different chemical structures and multiple toxic effects is an important prerequisite for accurately identifying and prioritizing high-risk new pollutants.

[0003] Determining the influence rule and degree of chemical structure characteristics of new pollutants on toxic effects is an important basis for evaluating environmental risks and a key problem and challenge currently faced by new pollutant prevention and control. Although the chemical structure and multiple toxic effects of new pollutants have been explored, the quantitative comparison of the influence of specific chemical structures on toxicity is limited to a certain group or a certain toxic endpoint. In the face of hundreds of organic compounds, exploring the quantitative comparison of toxicity of multiple endpoints and across groups is an important way to efficiently and accurately identify and prioritize new pollutants.

[0004] Therefore, how to provide a method for comparing toxicities of organic compounds across multiple endpoints and groups is a problem that needs to be solved by those skilled in the art. SUMMARY

[0005] 1. Problem to be solved

[0006] In view of the problem that the research on the influence of chemical structure characteristics on toxic effects is limited in the prior art, and the environmental risk assessment of new pollutants is not accurate, the present application provides a method for comparing toxicities of organic compounds. The present application establishes a linear regression model based on different chemical structure characteristics of organic compounds and dimensionless parameter scaling of toxic values of different compounds, so as to realize the trend and degree of influence of chemical structure characteristics on toxicities of organic compounds across multiple endpoints and groups, and provide an important way for efficient and accurate identification and prioritization of new pollutants.

[0007] 2. Technical solution

[0008] In order to solve the above problems, the technical scheme adopted by the present application is as follows:

[0009] The method for comparing toxicities of organic compounds provided by the present application comprises the following steps:

[0010] S10, collecting toxicity values of organic compounds at different toxicity endpoints to establish a toxicity database;

[0011] S20, classifying the organic compounds in step S10 into a plurality of groups according to a primary chemical structure feature, and classifying the organic compounds in each group into a plurality of subgroups according to a secondary chemical structure feature, wherein the secondary chemical structure feature is different from the primary chemical structure feature;

[0012] S30, performing dimensionless parameter scaling and linear regression analysis on the toxicity indexes of all compounds in the same subgroup at the same toxicity endpoint to compare the influence of the secondary chemical structure feature on the toxicity of the organic compounds across groups;

[0013] S40, using the dimensionless parameters obtained from different subgroups at the same toxicity endpoint to calculate and perform linear regression analysis to compare the influence of the secondary chemical structure feature on the toxicity of the organic compounds across endpoints.

[0014] Preferably, in step S10, the number of toxicity endpoints is not less than 2.

[0015] Preferably, in step S20, the primary chemical structure feature or the secondary chemical structure feature includes, but is not limited to, one or more of the following: compound parent structure, number of halogens, type of halogen, isomeric configuration.

[0016] Preferably, in step S20, the toxicity indexes of all organic compounds in each subgroup of the plurality of subgroups come from the same toxicity experiment.

[0017] Preferably, in step S30, the dimensionless parameter scaling is the ratio of the toxicity index of other organic compounds to the toxicity index of the organic compound with a specific secondary chemical structure feature in each subgroup, D, and each substance in the same subgroup corresponds to a dimensionless parameter.

[0018] Preferably, in step S40, the calculation is taking the mean value of the dimensionless parameters of the compounds with the same secondary chemical structure feature in different subgroups at the same toxicity endpoint, Each secondary chemical structure feature corresponds to a mean value of the dimensionless parameters at the same toxicity endpoint.

[0019] Preferably, in step S30, the linear regression analysis of each subgroup corresponds to a linear regression model D = kn + b, where D is the dimensionless parameter scaling and n is the value of the secondary chemical structure feature.

[0020] Preferably, the slopes k in the linear regression models are compared to compare the influence of the secondary chemical structure feature on the toxicity of the compounds across groups.

[0021] Preferably, in step S40, each toxicity endpoint corresponds to a linear regression model wherein, is the dimensionless parameter scale mean, and n is the secondary chemical structure feature value.

[0022] Preferably, in step S40, the slope of the linear regression model is compared with the value of the slope of the linear regression model, and the influence of the secondary chemical structure feature on the toxicity of the compound is compared across endpoints.

[0023] 3. Advantages

[0024] Compared with the prior art, the advantages of the present application are:

[0025] (1) The organic compound toxicity comparison method of the present application is different from the traditional research on the influence of chemical structure features on toxicity under a certain specific group and a certain specific toxicity endpoint. The present application realizes multi-endpoint and cross-group comparison based on the dimensionless parameter scale method, and eliminates the interference of irrelevant chemical structure features.

[0026] (2) The organic compound toxicity comparison method of the present application is based on the linear regression analysis method of the dimensionless parameter scale. Quantitative comparison is realized through the same parameter value, positive or negative, and size, thereby realizing quantitative comparison of the toxicity of organic compounds across multiple endpoints and groups.

[0027] (3) The organic compound toxicity comparison method of the present application realizes quantitative comparison of the toxicity of organic compounds across multiple endpoints and groups, explores the relationship between different chemical structure features and multiple toxicity effects, and provides an important way for efficient and accurate identification and priority control of new pollutants. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 is a flowchart of the organic compound toxicity comparison method of the present application. DETAILED DESCRIPTION

[0029] The present application will be further described below in conjunction with specific embodiments. The essential features and significant effects of the present application can be embodied in the following examples. The described examples are part of the present application, but not all of the examples. Therefore, they do not limit the present application in any way. Those skilled in the art can make some non-essential improvements and adjustments based on the content of the present application, which are all within the protection scope of the present application.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terms used herein have their ordinary meaning to a person of ordinary skill in the art and are not to be limited to a specialized meaning unless specifically so defined herein. The following description is further only meant to illustrate the application and is not meant to limit the application in any way. Unless otherwise specifically explained, the reagents, methods, and equipment employed in the following examples are of a type commonly employed in the art.

[0031] As shown in the following table, the method for comparing toxicity of organic compounds of the present application specifically comprises the following steps: Figure 1

[0032] S10, obtaining toxicity values of organic compounds under different toxicity endpoints (not less than 2 toxicity endpoints) from published academic journals (exemplary journals can be JOURNAL OF ENVIRONMENTAL SCIENCE, ENVIRONMENTAL SCIENCE & TECHNOLOGY, WATER RESEARCH), and / or widely recognized public databases (exemplary databases can be PubChem), and / or standardized, scientific biological laboratories (exemplary laboratories can be State Key Laboratory of Pollution Control and Resource Reuse, Nanjing University), calculating toxicity indexes, and establishing a database; wherein, the "toxicity value" refers to the concentration of an organic compound that can cause a half maximum effect concentration EC 50 or a half death concentration LC 50 of a test organism (unit: μM), and the "toxicity index" refers to a unified standard method and scale for indicating and comparing the toxicity of a poison, which is generally the inverse of the toxicity value, and the calculation method is toxicity index TI (TI = (EC 50 or LC 50 ) -1 *10 6 ).

[0033] S20, classifying the organic compounds in step S10 into multiple groups according to the primary chemical structure characteristics, and classifying the organic compounds in each group into multiple subgroups according to the secondary chemical structure characteristics, wherein the secondary chemical structure characteristics are different from the primary chemical structure characteristics, and the primary chemical structure characteristics or the secondary chemical structure characteristics comprise one or more of the following: compound parent structure, number of halogens, type of halogens, isomeric configuration; it should be noted that the toxicity indexes of all the organic compounds in each subgroup in the multiple subgroups come from the same toxicity experiment.

[0034] ​S30, the toxicity indexes of all compounds in the same sub-group under the same toxicity endpoint are scaled by a dimensionless parameter and linear regression analysis is performed to compare the influence of secondary chemical structure characteristics on the toxicity of organic compounds across groups, wherein the dimensionless parameter scaling is the ratio of the toxicity index of other organic compounds in each sub-group to the toxicity index of the organic compound having a specific secondary chemical structure characteristic; it should be noted that the present application scales the dimensionless parameter by the ratio of the toxicity index of other compounds in each sub-group to the toxicity index of the compound having a specific secondary chemical structure characteristic, so as to exclude the influence of primary chemical structure characteristics on the toxicity index, thereby determining the influence trend and influence degree of secondary chemical structure characteristics on the toxicity of organic compounds.

[0035] The linear regression analysis obtains a linear regression model D=k n+b corresponding to each sub-group, wherein D is the dimensionless parameter scaling, and n is the secondary chemical structure characteristic value; it should be noted that the present application determines the quantitative relationship between the toxicity index of the dimensionless parameter scaling and the secondary chemical structure characteristics of the organic compound by the linear regression model, and compares the slope k value in the linear regression model to compare the influence trend and influence degree of secondary chemical structure characteristics on the toxicity of compounds across groups.

[0036] S40, the dimensionless parameters obtained by different sub-groups under the same toxicity endpoint are used to calculate and perform linear regression analysis to compare the influence of secondary chemical structure characteristics on the toxicity of organic compounds across endpoints, wherein the dimensionless parameters of the compounds having the same secondary chemical structure characteristics in different sub-groups under the same toxicity endpoint are averaged,

[0037] The linear regression analysis obtains a linear regression model D=k n+b corresponding to each sub-group, wherein D is the dimensionless parameter scaling, and n is the secondary chemical structure characteristic value; it should be noted that the present application determines the quantitative relationship between the toxicity index of the dimensionless parameter scaling and the secondary chemical structure characteristics of the organic compound by the linear regression model, and compares the slope k value in the linear regression model to compare the influence trend and influence degree of secondary chemical structure characteristics on the toxicity of compounds across groups. wherein, is the average of the dimensionless parameter scaling, and n is the secondary chemical structure characteristic value; by comparing the slope k value in the linear regression model, the influence trend and influence degree of secondary chemical structure characteristics on the toxicity of compounds across endpoints are compared.

[0038] Example 1

[0039] The organic compound toxicity comparison method of the present embodiment specifically comprises the following steps:

[0040] ​S10, obtain toxicity values of organic compounds from published academic journals (JOURNAL OF ENVIRONMENTAL SCIENCE and ENVIRONMENTAL SCIENCE & TECHNOLOGY) and standardized, scientific biological laboratories (State Key Laboratory of Pollution Control and Resource Reuse, Nanjing University), calculate toxicity indexes, and establish a database;

[0041] As shown in Table 1, the toxicity values of 12 organic compounds under 3 toxicity endpoints were selected for calculation and comparison. Each organic compound corresponded to a toxicity value (EC 50 or LC 50 , unit: μM), and 36 toxicity values and 36 toxicity indexes were obtained.

[0042] S20, classify the organic compounds in step S10 into multiple groups according to the primary chemical structure characteristics;

[0043] As shown in Table 1, the 12 compounds were divided into 3 groups according to the parent structure (haloacetic acid, haloacetonitrile, and halophenol).

[0044] Table 1 Toxicity database of target organic compounds in Example 1

[0045]

[0046] According to the secondary chemical structure characteristics, the organic compounds in each group in step S20 were classified into multiple subgroups;

[0047] As shown in Table 2, 3 compounds in each group were selected to form a subgroup according to the halogen substitution type. According to the database, each group can have one or more subgroups.

[0048] S30, perform dimensionless parameter scaling and linear regression analysis on the toxicity indexes of all compounds in the same subgroup under the same toxicity endpoint, to compare the influence trend and degree of the secondary chemical structure characteristics on the toxicity of the compounds across groups;

[0049] As shown in Table 2, the ratio of the toxicity index of the brominated or iodinated compound to that of the chlorinated compound in each subgroup was the dimensionless parameter of bromination or iodination. Then, the secondary chemical structure characteristics of chlorination, bromination, and iodination were assigned with characteristic values of 1, 2, and 3, respectively. Combined with the above corresponding dimensionless parameters, linear regression analysis was performed to obtain the linear regression model as shown in Table 2. Each subgroup corresponded to a linear regression model;

[0050] As shown in Table 3, the dimensionless parameters of all organic compounds with the same secondary chemical structure feature of chloro or bromo or iodo under each toxicity endpoint are averaged, and linear regression analysis is performed in combination with the characteristic value to obtain the linear regression model as shown in Table 3, and one linear regression model can be obtained for each toxicity endpoint;

[0051] Table 2: Cross-group comparison method of organic compounds in Example 1 using dimensionless parameter scale

[0052]

[0053] S40, using the dimensionless parameters obtained from different subgroups under the same toxicity endpoint in step S30, linear regression analysis is performed to obtain a linear regression model, and the influence trend and degree of secondary chemical structure features on the toxicity of compounds are compared across endpoints.

[0054] As shown in Table 3, the dimensionless parameters of all organic compounds with the same secondary chemical structure feature of chloro or bromo or iodo under each toxicity endpoint are averaged, and linear regression analysis is performed in combination with the characteristic value to obtain the linear regression model as shown in Table 3, and one linear regression model can be obtained for each toxicity endpoint;

[0055] The slope of the linear regression model under different toxicity endpoints is positive, indicating that the toxicity of the compounds also increases with the change of the chlorine, bromine, and iodine halogen substitution type under different toxicity endpoints; the absolute value of the slope of the linear regression model under different toxicity endpoints is in the order of CHO > Hep G2 > MVLN cell toxicity, indicating that CHO cell toxicity is most affected by halogen substitution type, this toxicity endpoint is the most sensitive, Hep G2 cell toxicity is less affected, and MVLN cell toxicity is the least affected.

[0056] Table 3: Multi-endpoint comparison method of organic compounds in Example 1 using dimensionless parameter scale

[0057]

[0058]

[0059] Example 2

[0060] The basic content of this example is the same as that of Example 1, except that the organic compound toxicity comparison method of this example specifically includes the following steps:

[0061] S10, obtaining toxicity values of organic compounds from published academic journals (Journal of Environmental Science) and standardized, scientific biological laboratories (State Key Laboratory of Pollution Control and Resource Reuse, Nanjing University), calculating toxicity indexes, and establishing a database;

[0062] As shown in Table 4, the toxicity values of 12 organic compounds at 2 toxicity endpoints were selected for calculation and comparison. Each organic compound corresponded to a toxicity value (EC 50 or LC 50 , unit: μM), and corresponded to 24 toxicity indexes.

[0063] S20, classifying the organic compounds in step S10 into multiple groups according to the primary chemical structure characteristics;

[0064] As shown in Table 4, the 12 compounds were divided into 4 groups according to the parent structure (haloacetic acid, haloacetonitrile, haloacetaldehyde, and haloacetamide).

[0065] Table 4 Toxicity database of target organic compounds in Example 2

[0066]

[0067] According to the secondary chemical structure characteristics, the organic compounds in each group in step S20 were classified into multiple subgroups;

[0068] As shown in Table 5, 3 compounds in each group were selected to form a subgroup according to the number of halogen substitutions. Each group can have one or more subgroups according to the database.

[0069] S30, performing dimensionless parameter scaling and linear regression analysis on the toxicity indexes of all compounds in the same subgroup at the same toxicity endpoint, to compare the influence trend and degree of the secondary chemical structure characteristics on the toxicity of the compounds across groups;

[0070] As shown in Table 5, the toxicity index of trihalogenated compounds in each subgroup was taken as a reference, and the ratio of the toxicity index of monohalogenated or dihalogenated compounds to that of trihalogenated compounds was the dimensionless parameter of monohalogenation or dihalogenation. Then, the secondary chemical structure characteristics of monohalogenation, dihalogenation, and trihalogenation were assigned characteristic values of 1, 2, and 3, respectively, and combined with the corresponding dimensionless parameters, linear regression analysis was performed to obtain the linear regression model shown in Table 5. Each subgroup corresponded to a linear regression model;

[0071] As shown in Table 6, the dimensionless parameters of all organic compounds with the same secondary chemical structure feature of monohalogenation or dihalogenation or trihalogenation under each toxicity endpoint are averaged, and linear regression analysis is performed in combination with the characteristic value to obtain the linear regression model of Table 6, and one linear regression model can be obtained for each toxicity endpoint;

[0072] Table 5 Cross-group comparison method of organic compounds with dimensionless parameter scale in Example 2

[0073]

[0074] S40, the dimensionless parameters obtained by the same toxicity endpoint under different subgroups in step S30 are calculated, and linear regression analysis is performed to obtain a linear regression model, so as to compare the influence trend and degree of secondary chemical structure features on the toxicity of compounds.

[0075] As shown in Table 6, the dimensionless parameters of all organic compounds with the same secondary chemical structure feature of monohalogenation or dihalogenation or trihalogenation under each toxicity endpoint are averaged, and linear regression analysis is performed in combination with the characteristic value to obtain the linear regression model of Table 6, and one linear regression model can be obtained for each toxicity endpoint;

[0076] The slope of the linear regression model under different toxicity endpoints is negative, indicating that the toxicity of the compounds changes with the change of the number of monohalogenation, dihalogenation or trihalogenation under different toxicity endpoints; the absolute value of the slope of the linear regression model under different toxicity endpoints is in the order of CHO > MVLN cell toxicity, indicating that the degree of influence of CHO cell toxicity on the number of halogen substitution is greater than that of MVLN cell toxicity.

[0077] Table 6 Multi-endpoint comparison method of organic compounds with dimensionless parameter scale in Example 2

[0078]

Claims

1. A method for comparing toxicity of organic compounds, characterized by: The method comprises the following steps: S10, collecting toxicity values of organic compounds at different toxicity endpoints to establish a toxicity database; S20, classifying the organic compounds in step S10 into multiple groups according to primary chemical structure characteristics, and classifying the organic compounds in each group into multiple subgroups according to secondary chemical structure characteristics, wherein the secondary chemical structure characteristics are different from the primary chemical structure characteristics, and the primary chemical structure characteristics or the secondary chemical structure characteristics comprise one or more of the following: compound parent structure, number of halogens, type of halogens, isomeric configuration; S30, performing dimensionless parameter scaling on toxicity indexes of all compounds in the same subgroup at the same toxicity endpoint and performing linear regression analysis, and comparing the slope k values in the linear regression models to compare the influence of the secondary chemical structure characteristics on the toxicity of the organic compounds across groups; the dimensionless parameter scaling is the ratio of the toxicity index of other organic compounds in each subgroup to the toxicity index of the organic compound with a specific secondary chemical structure characteristic, D, and each substance in the same subgroup corresponds to a dimensionless parameter; S40, using the dimensionless parameters of different subgroups under the same toxicity endpoint to calculate and conduct linear regression analysis, comparing the slope in the linear regression model to compare the influence of secondary chemical structure characteristics on the toxicity of organic compounds across endpoints; the calculation is to take the mean of the dimensionless parameters of compounds with the same secondary chemical structure characteristics in different subgroups under the same toxicity endpoint, a mean of dimensionless parameters corresponding to each secondary chemical structure characteristic under the same toxicity endpoint.

2. The method of toxicological comparison of organic compounds according to claim 1, characterized in that: In step S10, the number of toxicity endpoints is not less than 2.

3. The method of claim 1, wherein: In step S20, the toxicity indexes of all organic compounds in each subgroup in the multiple subgroups come from the same toxicity experiment.

4. The method of claim 1, wherein: In step S30, the linear regression analysis corresponds to a linear regression model D=kn+b for each subgroup, wherein D is the dimensionless parameter scaling, and n is the secondary chemical structure characteristic value.

5. The method of claim 1, wherein: In step S40, each toxicity endpoint corresponds to a linear regression model obtained by the linear regression analysis wherein, is the dimensionless parameter scale mean, and n is the secondary chemical structure characteristic value.

Citation Information

Patent Citations

  • Method for improving accuracy of predicting toxic effect endpoint value through pollutant QSAR model

    CN113793651A