Method and device for obtaining general chemical combination for toxicity assessment of mixture
By constructing chemical detection matrix and frequent item set mining technology, the chemical combinations that frequently appear in the target environment are screened out, and the problem of low accuracy of toxicity assessment of mixed exposure of multiple chemicals in complex environments is solved, achieving more efficient and accurate toxicity assessment.
Patent Information
- Application Number
- CN202411984147.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The prior art is difficult to effectively evaluate the toxicity of mixed exposures of multiple chemicals in complex environments, resulting in low accuracy in toxicity assessment.
By statistics and constructing chemical detection matrix, combining frequent item set mining technology, chemical combinations that frequently appear in the target environment are screened out, and high-order frequent item sets are obtained through optimization strategies to reduce the number of chemical combinations and improve evaluation efficiency.
Effectively screen out chemical combinations that are common in the actual environment, improve the accuracy and efficiency of environmental mixture toxicity assessment and reduce unnecessary toxicity assessment work.
Smart Images

Figure CN120048386A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental monitoring, and in particular to a method and a device for obtaining a universal chemical combination for mixture toxicity assessment. Background Art
[0002] Mixed exposure to pollutants (chemicals) in the actual environment is a universal law. It is necessary to conduct toxicity studies on mixed exposure to pollutants (chemical combinations) based on the pollutants detected in the actual environment to assess the joint risk. At present, research on the toxicity of mixed exposure to pollutants mainly focuses on the toxicity studies of binary mixtures consisting of two pollutants (binary chemical combinations) or ternary mixtures consisting of three pollutants (ternary chemical combinations). However, since the types of pollutants (chemicals) and exposure scenarios in the actual environment are often more complex, a comprehensive assessment of the mixed exposure toxicity of chemicals in the environment that relies solely on the toxicity research results of binary or ternary mixtures will deviate greatly from reality, resulting in a low accuracy in the assessment of the toxicity of the mixture in the target environment. For example, the number of possible combinations of a co-exposure mixture consisting of only 20 chemicals is theoretically more than one million (2 20 -1=1048575), and the number of mixture combinations doubles with each additional chemical. Although the exposure to chemicals in the actual environment, especially the mixed exposure to chemicals, is not completely random and is subject to various construction processes, the actual number of n mixed exposures experienced in the real environment will be much less than 2 n -1, but there is a big gap between the mixed exposure of chemicals based on binary or ternary chemical combinations and the mixed exposure of the combinations of chemicals in the actual environment. Therefore, how to screen and identify the chemical combinations existing in the actual environment from a large number of possible chemical mixed exposures, so as to reduce the number of chemical combinations for environmental mixture toxicity assessment and improve the efficiency and accuracy of environmental mixture toxicity assessment, is a problem that must be solved in environmental toxicity risk assessment research. Summary of the invention
[0003] In view of this, an object of the present invention is to provide a method and an apparatus for obtaining a common chemical combination for toxicity assessment, so as to improve the accuracy of toxicity assessment of environmental mixtures.
[0004] In a first aspect, an embodiment of the present invention provides a method for obtaining a universal chemical combination for toxicity assessment, comprising: Counting the total number of chemicals detected at each sampling and testing point in the target environment, for each sampling and testing point, constructing a chemical detection matrix for the sampling and testing point based on the chemical concentration detected at the sampling and testing point and the total number of chemicals, and based on the chemical detection matrices of each sampling and testing point, constructing a compound detection matrix for the target environment; Querying the mapping relationship between the preset chemicals and the analytical detection limits, and converting the compound detection matrix into a compound presence / absence matrix based on the target chemical concentration in the compound detection matrix and the analytical detection limits mapped to the target chemical; For each candidate chemical in the compound presence / absence matrix, constructing a first frequent item set of chemicals based on the occurrence frequency of the candidate chemical in the compound presence / absence matrix and a preset first frequency threshold; Based on a preset item set optimization strategy, the first frequent item set of the chemical is optimized from a low-order chemical combination to a high-order chemical combination, and high-order frequent item sets are obtained in sequence; Based on the first frequent item set of chemicals and the successively obtained high-order frequent item sets, a common chemical combination for performing mixture toxicity assessment is obtained.
[0005] In a second aspect, an embodiment of the present invention provides a device for obtaining a universal chemical combination for mixture toxicity assessment, comprising: A concentration matrix construction module is used to count the total number of chemicals detected at each sampling detection point in the target environment, and for each sampling detection point, construct a chemical detection matrix for the sampling detection point based on the chemical concentration detected at the sampling detection point and the total number of chemicals, and based on the chemical detection matrices of each sampling detection point, construct a compound detection matrix for the target environment; A presence / absence matrix conversion module is used to query the mapping relationship between the preset chemicals and the analytical detection limits, and convert the compound detection matrix into a compound presence / absence matrix based on the target chemical concentration in the compound detection matrix and the analytical detection limit mapped to the target chemical; A low-order frequent item set construction module is used to construct a first frequent item set of chemicals for each candidate chemical in the compound presence matrix based on the occurrence frequency of the candidate chemical in the compound presence matrix and a preset first frequency threshold; A high-order frequent item set construction module is used to optimize the first frequent item set of chemicals from low-order chemical combinations to high-order chemical combinations based on a preset item set optimization strategy, and obtain high-order frequent item sets in sequence; The chemical combination acquisition module is used to acquire a common chemical combination for mixture toxicity assessment based on the first frequent item set of the chemical and the successively acquired high-order frequent item sets.
[0006] According to a third aspect of the present invention, there is provided a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for obtaining a common chemical combination for mixture toxicity assessment in any possible implementation of the first aspect.
[0007] According to a fourth aspect of the present invention, there is provided an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for obtaining a common chemical combination for mixture toxicity assessment in any possible implementation of the first aspect are implemented.
[0008] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0010] Figure 1 A schematic flow chart of a method for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention is shown; Figure 2 Another schematic diagram of a method for obtaining a universal chemical combination for mixture toxicity assessment provided by an embodiment of the present invention is shown; Figure 3 Another schematic diagram of the method for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention is shown; Figure 4 A schematic diagram of constructing multi-level nodes in a method for obtaining a universal chemical combination for mixture toxicity assessment provided by an embodiment of the present invention is shown; Figure 5 A schematic diagram of the structure of a device for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention is shown; Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0011] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present invention.
[0012] In the related technologies, the methods for studying the mixed exposure toxicity of pollutants in the actual environment based on binary mixtures (binary chemical combinations) or ternary mixtures are difficult to comprehensively evaluate the toxicity of chemical combinations that may exist in the environment, and the assessment accuracy of environmental toxicity is low. Therefore, based on the mixed exposure of chemicals in the actual environment, how to screen and identify the chemical combinations that exist in the actual environment from a large number of possible chemical combinations is a problem that must be solved in environmental toxicity risk assessment research.
[0013] In the actual environment, any combination of chemicals detected may constitute a mixed exposure combination of chemicals. For example, if 20 antibiotics are detected in a certain watershed, the number of antibiotic combinations composed of these 20 antibiotics will theoretically exceed one million (2 20 -1=1048575), but most of the one million antibiotic combinations in this theoretical combination do not actually exist in the environment, or their existence is extremely rare, and their toxic effects on the environment can be ignored. That is, among the chemical co-exposure combinations (chemical combinations) that have been constructed in various ways, only a small number of chemical combinations are common or frequently present in the actual environment, while a large number of chemical combinations do not exist in the actual environment, or their toxic effects on the environment can be ignored. Therefore, it is necessary to analyze the antibiotic combinations (mixtures) that frequently appear in the basin from the 20 antibiotics, that is, to obtain the common chemical combinations in the basin for corresponding mixture toxicity assessment, thereby effectively reducing the workload of mixture toxicity assessment and improving the efficiency and accuracy of mixture toxicity assessment.
[0014] In this embodiment, a small amount of chemical co-exposure combinations that are common or frequently present in the actual environment and whose toxic effects on the environment cannot be ignored are called "common chemical combinations". The method of this embodiment, by screening and obtaining common chemical combinations, provides a reference for subsequent environmental toxicity risk assessment, effectively reduces the number of chemical combinations for environmental mixture toxicity assessment, and improves the efficiency of environmental mixture toxicity assessment. Specifically, based on the chemical detection concentration data, the frequent item set mining (FIM) technology is used. For example, 20 antibiotics obtained by testing 20 sampling points in the target environment are subjected to frequent item set mining (FIM) to obtain the actual or high probability antibiotic combinations composed of 20 antibiotics in the target environment, and then the common antibiotic combinations are screened out according to the minimum support threshold (common level, such as 50% probability). Based on the screened common antibiotic combinations, high-order common antibiotic combinations are further screened for mixture toxicity assessment. For example, the common antibiotic combinations selected according to the minimum support threshold include the antibiotic combinations {C1, C2}, {C1, C4} and the antibiotic degree combination {C1, C2, C3}, among which the antibiotic degree combination {C1, C2, C3} is a high-order combination (third order), and the antibiotic combinations {C1, C2} and {C1, C4} are low-order combinations (second order). Therefore, in the common antibiotic combinations selected according to the minimum support, the high-order combination may contain the low-order combination (the high-order combination {C1, C2, C3} contains the low-order combination {C1, C2}, but does not contain the low-order combination {C1, C4}). Since the mixture toxicity assessment study of the mixture gives priority to the study of more complex combinations (high-order combinations), the low-order combination {C1, C2} can be eliminated through the high-order combination screening to further improve the efficiency of the mixture toxicity assessment. In this way, from the large number of chemical combinations that may be formed in theory, the common chemical combinations that actually exist and have universal combinations can be screened out, thereby providing key reference information for the coexistence characteristics and ecological risk assessment of chemicals in the actual environment, and at the same time providing important material basis and methodological support for mixture design in mixture toxicity research.
[0015] In this embodiment, FIM technology is a big data mining technology, which was originally used for shopping basket analysis, that is, identifying a set of goods (frequent itemsets) that are often purchased together (i.e., frequently purchased) from the list of goods purchased in each consumer's shopping basket, where a single commodity is called an item, and a set of k commodities is called a k-itemset. In this embodiment, different environmental objects (such as collection and detection points) are defined as a transaction, each chemical is an item, and a set of k chemicals is called a k-itemset. The common chemical combinations in different environments are equivalent to the frequent itemsets in FIM. In this way, the screening and identification of common chemical combinations in different environments are converted into determining frequent itemsets in FIM.
[0016] In this embodiment, the frequent item set mining technology in data mining is used to screen and identify the common antibiotic co-exposure combinations in different environmental systems, thereby forming a method for identifying common chemical combinations using FIM. Specifically, based on the basic data set of chemical concentrations in different actual environments acquired in advance, a corresponding concentration threshold is set for each chemical, and the original data (basic data set) is processed according to the concentration threshold to construct a "presence-absence" data set, and the FIM technology is used to preliminarily mine frequently occurring chemical combinations to form common chemical combinations. Furthermore, for the highest element combination existing in each common chemical combination screened, a frequent item set technology is established to identify high-order common chemical combinations.
[0017] The embodiment of the present invention provides a method for obtaining a universal chemical combination for mixture toxicity assessment, which is described below through an embodiment.
[0018] Figure 1 The schematic diagram of the method for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention is shown. Figure 1 As shown, the process includes: S101, counting the total number of chemicals detected at each sampling detection point in the target environment, and for each sampling detection point, constructing a chemical detection matrix for the sampling detection point according to the chemical concentration detected at the sampling detection point and the total number of chemicals, and constructing a compound detection matrix for the target environment based on the chemical detection matrices of each sampling detection point; In this embodiment, as an optional embodiment, the chemical detection matrix of the sampling detection point is a row matrix, the number of columns contained in the row matrix is the total number of chemicals, and each column corresponds to a detected chemical.
[0019] In this embodiment, as an optional embodiment, taking the target environment as a certain river basin as an example, taking the collection of 8 water samples (sampling detection points, referred to as sampling points) including 5 chemical concentrations (ng / L) as an example, the compound detection matrix of the target environment constructed based on the collected raw data is shown in Table 1. In Table 1, NR means not detected.
[0020]
[0021] S102, querying a preset mapping relationship between chemicals and analytical detection limits, and converting the compound detection matrix into a compound presence / absence matrix based on the target chemical concentration in the compound detection matrix and the analytical detection limit mapped to the target chemical; In this embodiment, when FIM technology is applied to shopping basket analysis, the analysis is based on the "presence-absence" of the goods in the consumer's shopping list. Therefore, when FIM technology is applied to the analysis of chemical concentrations in the actual environment, a "presence-absence" or "1-0" data set of relevant chemicals in different environmental objects is established, so that the concentration data of each chemical is converted into a "presence-absence" data set of the chemical according to the analytical detection limit of the chemical, thereby filtering out chemicals whose impact on the environment can be ignored.
[0022] In this embodiment, as an optional embodiment, based on the target chemical concentration in the compound detection matrix and the analytical detection limit mapped to the target chemical, the compound detection matrix is converted into a compound presence matrix, including: The target chemical concentration corresponding to each chemical in the compound detection matrix is traversed, and the analytical detection limit mapped to the chemical is obtained from the mapping relationship. If the target chemical concentration is greater than or equal to the analytical detection limit, the target chemical concentration corresponding to the chemical is updated to 1; otherwise, the target chemical concentration corresponding to the chemical is updated to 0.
[0023] In this embodiment, in the compound detection matrix, the columns are the detected chemicals, the rows are the sampling detection points, and the row and column values are the detected chemical concentrations. As an optional embodiment, if the chemical concentration is greater than the analytical detection limit corresponding to the chemical, it indicates that the chemical exists or is "yes" or "1", otherwise, it is determined that the chemical does not exist or is "no" or "0".
[0024] In this embodiment, as an optional embodiment, the analytical detection limit of the chemical may be set differently based on the impact of the chemical on the environment, the instrument used to detect the chemical, and the detection method. For example, for detecting antibiotic A using a certain detection method of a certain instrument, the analytical detection limit set is 0.04ng / L, and for detecting antibiotic A using another instrument and the same detection method, the analytical detection limit set may also be 0.035ng / L. In this embodiment, taking the detection using a certain detection method of a certain instrument as an example, for each sampling detection point, if the detection result of antibiotic A is <0.04ng / L, it means that antibiotic A does not exist in the corresponding sampling point, that is, "none" or "0". As an optional embodiment, the chemical concentration data in Table 1 is converted into "presence-absence" data, and the compound presence-absence matrix obtained is shown in Table 2.
[0025]
[0026] In this embodiment, by processing the chemical concentration data set of each sampling and detection point in the target environment into a "presence-absence" data set (compound presence-absence matrix), it is convenient to subsequently process the "presence-absence" data set using FIM technology.
[0027] S103, for each candidate chemical in the compound presence matrix, constructing a first frequent item set of chemicals based on the occurrence frequency of the candidate chemical in the compound presence matrix and a preset first frequency threshold; In this embodiment, let I = {a1, a2, …, am} be a set of items, where a1, a2, …, am are m different items, each item represents a chemical, m is the total number of chemicals, and is the number of columns of the compound presence matrix. D = {T1, T2, …, Tn} is a transaction data set, and each transaction Ti (i∈{1, 2, …, n}) is a subset of I. Among them, each transaction is identified by a corresponding transaction identifier (TID, Task Identifier), which represents a sampling detection point, and n is the number of sampling detection points in the target environment, that is, the number of rows of the compound presence matrix.
[0028] In this embodiment, if a sub-item set A of I satisfies A T, then transaction T is said to contain sub-item set A. The support of a sub-item set (item set) is defined as: the proportion of the number of transactions containing the sub-item set in the total number of transactions. If the support of a certain item set is greater than the preset minimum support threshold, the item set is called a frequent item set or frequent pattern, and a frequent item set with a length of k is called a frequent k-item set. As an optional embodiment, since the number of chemicals in the compound presence matrix is fixed, the support can be characterized by the frequency of occurrence.
[0029] In this embodiment, the total number of chemicals contained in different sampling and detection points corresponds to items, the item set corresponds to the mixture composed of chemicals, different sampling and detection points correspond to transactions, the support is the proportion of the number of chemicals containing the same item set (composed of the same phase) in the total number of chemicals, and the mixture greater than the preset support threshold is a common chemical mixture.
[0030] In this embodiment, it is assumed that the number of sampling and testing points is 5, and the total number of chemicals detected by the 5 sampling and testing points is 5. Si represents the i-th sampling and testing point, Ci represents the i-th chemical, and ND represents not detected.
[0031] Among them, S1={C2,C3,C5}, which means that sampling and testing point 1 contains three chemicals: C2, C3, and C5; S2={C1,C3,C4,C5}; S3={C1,C2,C3,C4}; S4={C1,C3,C4}; S5={C1,C2,C3,C4}.
[0032] The constructed compound matrix is shown in Table 3.
[0033]
[0034] In this embodiment, assuming that the first frequency threshold is set to 3, based on Table 3, for chemical C1, the frequency of occurrence in the compound presence matrix is 4, which is greater than the first frequency threshold 3, so C1 is placed in the first frequent item set of chemicals. In this way, the first frequent item set of chemicals finally obtained includes: C1, C2, C3 and C4, that is, (C1, C2, C3, C4).
[0035] S104, based on a preset item set optimization strategy, optimizing the first frequent item set of chemicals from a low-order chemical combination to a high-order chemical combination, and sequentially obtaining high-order frequent item sets; In this embodiment, as an optional embodiment, based on a preset item set optimization strategy, the first frequent item set of chemicals is optimized from a low-order chemical combination to a high-order chemical combination, and high-order frequent item sets are obtained in sequence, including: A11, combining candidate chemicals in the first frequent item set of chemicals in pairs to obtain binary chemical combinations, obtaining the occurrence frequencies of the binary chemical combinations in the compound presence / absence matrix, and constructing a second frequent item set of chemicals based on binary chemical combinations whose occurrence frequencies are greater than or equal to a preset second frequency threshold; In this embodiment, taking the first frequent item set of chemicals including: C1, C2, C3 and C4 as an example, the binary chemical combinations in pairs include six combinations: (C1, C2), (C1, C3), (C1, C4), (C2, C3), (C2, C4) and (C3, C4).
[0036] In this embodiment, for the binary chemical combination (C1, C2), the frequency of occurrence (co-occurrence frequency) in the compound presence matrix is 2; the frequency of occurrence of (C1, C3) is 4, the frequency of occurrence of (C1, C4) is 4, the frequency of occurrence of (C2, C3) is 3, the frequency of occurrence of (C2, C4) is 2, and chemicals C3 and C4 exist in sampling points S2, S3, S4 and S5 at the same time, so the frequency of occurrence of (C3, C4) is 4. If the second frequency threshold is set to 3, the second frequent item set of chemicals constructed includes: {(C1, C3), (C1, C4), (C2, C3), (C3, C4)}.
[0037] A12, for the candidate chemicals in the second frequent item set of chemicals, any three candidate chemical combinations are performed to obtain a third-order chemical combination, and the occurrence frequencies of the third-order chemical combination in the compound presence / absence matrix are obtained in turn, and a third frequent item set of chemicals is constructed based on the third-order chemical combinations whose occurrence frequencies are greater than or equal to a preset third frequency threshold, until there is no higher-order chemical combination whose frequency is greater than or equal to a preset higher-order frequency threshold.
[0038] In this embodiment, as an optional embodiment, any three candidate chemical combinations are performed on the candidate chemicals in the second frequent item set of chemicals, including: Eliminating binary chemical combinations whose occurrence frequency is equal to a preset second frequency threshold from the second frequent item set of chemicals to obtain a filtered second frequent item set of chemicals; Three candidate chemicals are randomly selected from the second frequent item filtering set of chemicals to obtain the third-order chemical combination.
[0039] In this embodiment, as an optional embodiment, the monotonicity of the support in the FIM algorithm (i.e., if a k-item set or a k-ary chemical mixture is expanded (e.g., by adding one or more chemicals to it), the support of the k-item set or the k-ary chemical mixture will not increase) can be used to screen common chemical combinations. For example, taking the above-mentioned item set {C3, C4} as an example, the support of the item set {C3, C4} is 0.8, then after adding new chemicals to the item set, the support of the new item set obtained cannot exceed the support of the item set 0.8. For example, by adding a new chemical C2 to the item set {C3, C4}, the new item set {C2, C3, C4} obtained appears twice in 5 samples, and the corresponding support is 0.4. That is, if the minimum prevalence level is 50%, if the support of a combination (item set) is 0.5, then all other combinations obtained by adding one or more chemicals to the combination cannot have a corresponding support greater than 0.5, and therefore, no further screening is required.
[0040] Figure 2 Another schematic diagram of the method for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention is shown. Figure 2As shown, in this embodiment, the Apriori algorithm is used to obtain the general chemical combination. By scanning all the acquired data, as an optional embodiment, the item sets (compound presence / absence matrix) are scanned in sequence, for example, the support (occurrence frequency) of the 1-item set C1 is scanned, assuming that the minimum support is set to 2, and the support of the scanned item sets is pruned according to the minimum support (frequency threshold), assuming that the set L1 of frequent 1-item sets is obtained, including antibiotics ({1}, {2}, {3}, {5}); then, the combination of secondary frequent items in the set L1 of frequent 1-item sets is constructed, for example, the combination of secondary frequent items in the set L1 ({1}, {2}, {3}, {5}) includes: ({1 2}, {2 3}, {3 5}, {1 3}, {2 5}, {15}), thereby generating the set L2 of candidate 2-item sets. Next, the support of each chemical combination in the candidate 2-item set (2-item set) is scanned for the second time, and then based on the minimum support, the frequent 2-item set set L2 is obtained by pruning. Specifically, the support of the chemical combination {1 2} is 1, the support of {2 3} is 2, the support of {3 5} is 1, the support of {1 3} is 2, the support of {2 5} is 3, and the support of {1 5} is 2. Among them, the support of {1 2} and {3 5} is less than the minimum support 2, so they are removed. In this way, the frequent 2-item set set L2 obtained by pruning includes ({2 3}, {1 3}, {2 5}, {1 5}); then, the combination of third-order frequent items in L2 is constructed to generate the candidate 3-item set set C3. Since the frequent 2-item set set L2 only contains antibiotic 1, antibiotic 2, antibiotic 3 and antibiotic 5, the generated candidate 3-item set set C3 includes: ({1 2 3}, {1 3}, {2 5}, {1 5}). 3 5}, {23 5}), scan the support of the 3-item set, and obtain the set L3 of frequent 3-item sets by pruning, which is {2 3 5}. Since the frequent 3-item set only contains three antibiotics, there is no need to construct a fourth-order set, and the maximum frequent item set is the 3-item set L3.
[0041] In this embodiment, as another optional embodiment, taking the second frequent item set of chemicals including: (C1, C3), (C1, C4), (C2, C3) and (C3, C4) as an example, the number of candidate chemicals included is 4, and the number of high-order combinations is 3, then 4 third-order chemical combinations can be constructed, namely: (C1, C2, C3), (C1, C2, C4), (C2, C3, C4).
[0042] In this embodiment, for the third-order chemical combination (C1, C2, C3), the frequency of occurrence in Table 3 is 2; the frequency of occurrence of (C1, C2, C4) is 2; and the frequency of occurrence of (C2, C3, C4) is 2. If the preset third frequency threshold is 3, the processing of the third-order chemical combination is completed, and the common chemical combinations obtained for the mixture toxicity assessment are (C1, C2, C3), (C1, C2, C4), (C2, C3, C4); if the preset third frequency threshold is 2, since the frequency of occurrence of the third-order chemical combination: (C1, C2, C3), (C1, C2, C4), (C2, C3, C4) is equal to the third frequency threshold, the third-order chemical combination: (C1, C2, C3), (C1, C2, C4), (C2, C3, C4) is placed in the third frequent item set of chemicals, and a fourth-order combination is performed. As an optional embodiment, the number of chemicals included in the third frequent item set of chemicals is 4, and thus the fourth-order chemical combination is (C1, C2, C3, C4), which is a common chemical combination obtained for mixture toxicity assessment.
[0043] In this embodiment, as another optional embodiment, based on a preset item set optimization strategy, the first frequent item set of chemicals is optimized from a low-order chemical combination to a high-order chemical combination, and high-order frequent item sets are obtained in sequence, including: B11, build the root node; B12, in the compound presence / absence matrix corresponding to the first frequent item set of the chemical, extract the candidate chemicals in the first row, arrange them in descending order according to the frequency of occurrence of the candidate chemicals, and sequentially construct multiple levels of child nodes under the root node, wherein the level of the child nodes is the number of candidate chemicals contained in the first row, the candidate chemicals with high occurrence frequency are the parent nodes of the candidate chemicals with low occurrence frequency, and the counts of the candidate chemicals at each level of child nodes are all set to 1; B13, traverse the other rows in the compound presence / absence matrix, arrange the candidate chemicals in descending order according to the frequency of occurrence, traverse each candidate chemical in the sorted row, if the i-th candidate chemical in the row is the same as the candidate chemical in the i-th child node under the root node, add 1 to the count of the child node, if they are not the same, create a child node under the i-1-th child node for the i-th candidate chemical in the row from the i-1-th child node under the root node, set the count of the candidate chemical in the newly created child node to 1, and under the newly created child node, sequentially construct multiple levels of child nodes for the candidate chemicals after the i-th candidate chemical in the row; Figure 3 FIG. 2 shows another schematic diagram of a method for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention. Figure 3As shown, in this embodiment, the chemicals detected at each sampling detection point (TID: 100-500) in the target environment are obtained ( Figure 3 Left side), count the number of times each chemical detected in the target environment appears in each sampling point, find the chemicals with a support greater than the minimum (for example, the minimum support is 3) based on the number of times the chemical appears, and obtain frequent items. For each frequent item, re-sort it in order of support size ( Figure 3 Based on the reordered frequent items, for each sampling detection point, the sampling point is updated according to the sorted frequent items to obtain the sampling point item set obtained according to the sorted frequent items ( Figure 3 on the right).
[0044] Figure 4 FIG. 1 is a schematic diagram showing a method for constructing a multi-level node in a method for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention. Figure 4 As shown, based on Figure 3 Construct FP-Tree: It consists of three parts: frequent 1-item set (also called frequent item header table), root node (root) marked as null, and item prefix subtree (head of node-links). For example, the frequent item header table is f, c, a, b, m, p arranged from large to small according to support, and the item prefix subtree contains three levels of nodes. The process of constructing FP-Tree is to insert the items in the transaction into the item prefix subtree starting from the root node. If the same item is encountered, the child node count is +1, and if different items are encountered, a new child node is created.
[0045] In this embodiment, the occurrence frequencies of the chemicals counted from high to low are: chemical f: 4, chemical c: 4, chemical a: 3, chemical b: 3, chemical m: 3, chemical p: 3, and the occurrence frequencies of the remaining chemicals are all less than 3. In this embodiment, the threshold is set to 3; Assume that for sampling detection point 100, the detected chemicals include: f, a, c, d, g, i, m, p. After threshold screening and sorting, the first row in the compound presence matrix is: f, c, a, m, p. For sampling detection point 200, the detected chemicals include: a, b, c, f, l, m, o. After threshold screening and sorting, the second row in the compound presence matrix is obtained: f, c, a, b, m; For the sampling detection point 300, the detected chemicals include: b, f, h, j, o. After threshold screening and sorting, the third row in the compound presence matrix is obtained: f, b; For the sampling detection point 400, the detected chemicals include: b, c, k, s, p. After threshold screening and sorting, the fourth row in the compound presence matrix is obtained: c, b, p; For sampling detection point 500, the detected chemicals include: a, f, c, e, l, p, m, n. After threshold screening and sorting, the fifth row in the compound presence matrix is obtained as: f, c, a, m, p; For the f, c, a, m, p sorted in the first row; under the root node, build an f child node with a count of 1, under the f child node, build a c child node with a count of 1, under the c child node, build an a child node with a count of 1, under the a child node, build an m child node with a count of 1, and under the m child node, build a p child node with a count of 1; For the f, c, a, b, m sorted in the second row, there is an f child node under the root node, so the count of the f child node is updated to 2. Similarly, for the second chemical c and the third chemical a, there are corresponding second-level child nodes (c child nodes) and third-level child nodes (a child nodes), and the counts of the c child node and the a child node are updated to 2 respectively; for the fourth chemical b, the fourth-level child node (m child node) does not match chemical b, so a b child node with a count of 1 is constructed under the third-level child node (a child node) (a parallel node to the m child node), and a m child node with a count of 1 is constructed under the b child node accordingly; For f, b in the third row; under the root node, there is an f child node, so the count of the f child node is updated to 3. For the second chemical b, the second-level child node (c child node) does not match chemical b, so a b child node with a count of 1 is constructed under the first-level child node (f child node) (a parallel node to the c child node); For c, b, p in the 4th row, there is no c child node under the root node, so a c child node with a count of 1 is constructed under the root node (a parallel node to the f child node), and under the c child node, a b child node and a p child node with counts of 1 are constructed in sequence; For f, c, a, m, p in the 5th row, there is a child node f under the root node, so the count of the child node f is updated to 4. Similarly, for the 2nd chemical c, the 3rd chemical a, the 4th chemical m and the 5th chemical p, there are corresponding 2nd-level child nodes (c child nodes), 3rd-level child nodes (a child nodes), 4th-level child nodes (m child nodes) and 5th-level child nodes (p child nodes). Add 1 to the count value of each child node and finally obtain the multi-level child nodes of the candidate chemicals.
[0046] B14, select candidate frequent chemicals from the multi-level child nodes, locate the position of the candidate frequent chemicals in the multi-level child nodes, trace the root node for each located position, obtain the child node chain corresponding to each position, obtain the common chemicals contained in each child node chain, accumulate the counts of the common chemicals contained in each child node chain, and obtain a frequent item set containing the common chemical counts. If the common chemical count exceeds a preset count threshold, determine that the frequent item set corresponding to the common chemical count is a high-order frequent item set.
[0047] In this embodiment, mining high-order frequent item sets starts from a specific suffix, constructs a conditional pattern base of the FP-Tree, and determines whether it is frequent according to the updated count of the item on the conditional pattern base. For example, taking antibiotic p as the suffix, the two prefixes of antibiotic p form the conditional pattern base: {{p, m, a, c, f: 2}, {p, b, c: 1}}, and the generated conditional FP-Tree is:<p:3, c:3> , and the frequent itemset finally obtained is {p, c: 3}. Specifically, in the item prefix subtree, the conditional pattern bases formed by the two chemical p prefixes are: {{p, m, a, c, f: 2}, {p, b, c: 1}}, where the number is the count of chemical p. In the conditional pattern base formed by the two chemical p prefixes, there are two chemicals p and c in total. The counts of chemicals p and c are accumulated respectively, and the conditional FP-Tree is:<p:3, c:3> , therefore, the final frequent itemset is {p, c: 3}.
[0048] S105, based on the first frequent item set of chemicals and the successively acquired high-order frequent item sets, a common chemical combination for toxicity test evaluation is acquired.
[0049] In this embodiment, the toxicity test evaluation is a mixture toxicity evaluation. Using the FIM algorithm, all common chemical combinations containing two or more chemicals can be preliminarily screened out. However, in practical applications, the components of some low-order common chemical combinations may be repeated with the components of high-order common chemical combinations, indicating that these low-order combinations actually exist in a more complex form. Therefore, studying high-order common chemical combinations can not only avoid repeated analysis, but also more comprehensively reveal the complex correlations between chemicals, which is more practical. Therefore, in this embodiment, all frequent item sets screened out by the FIM algorithm are further screened. As an optional embodiment, based on the first frequent item set of the chemical and the successively obtained high-order frequent item sets, a common chemical combination for mixture toxicity evaluation is obtained, including: C11, from the high-order frequent item sets obtained in sequence, obtain a high-order frequent item set containing the maximum number of chemicals, construct a first set based on the high-order frequent item set containing the maximum number of chemicals, and construct a second set based on the remaining frequent item sets except the first set, wherein the maximum number of chemicals is n; In this embodiment, the high-order frequent item set with the number of chemicals being n (natural number) is taken as the first set, and the remaining frequent item sets (the number of chemicals being n-1, n-2, …, 3, 2) are taken as the second set, wherein the high-order frequent item set containing the largest number of chemicals is the highest-order frequent item set, which is the item set containing the largest number of chemicals in each frequent item set. The second set includes the high-order frequent item sets other than the highest-order frequent item set in the first frequent item set of chemicals and the high-order frequent item sets obtained in sequence.
[0050] C12, extract all candidate frequent item sets with the number of chemicals n-1 from the second set. For each candidate frequent item set, if the candidate frequent item set is a subset of any frequent item set in the first set, remove the candidate frequent item set from the second set. If not, merge the candidate frequent item set with the first set and update the first set. In this embodiment, for each high-order frequent item set with the number of chemicals n-1 in the second set, it is determined whether the high-order frequent item set is a subset of any item set in the first set, that is, whether the first set contains the high-order frequent item set. If so, the high-order frequent item set is removed from the second set, and the next high-order frequent item set is determined; if not, the high-order frequent item set is cut from the second set, and the cut high-order frequent item set is merged with the first set to form a new first set, and the next high-order frequent item set is determined. The merging is to place the cut high-order frequent item set in the first set, and the high-order frequent item set with the original chemical number n in the first set becomes two subsets of the first set.
[0051] C13, traverse all candidate frequent item sets in the second set whose number of chemicals is less than n-1. For each candidate frequent item set, if the candidate frequent item set is a subset of any frequent item set in the first set, remove the candidate frequent item set from the second set. If not, merge the candidate frequent item set with the updated first set. In this embodiment, according to the above method, the item sets whose number of chemicals is n-2, n-3, etc. in the second set are judged in turn until all the item sets whose number of chemicals is 2 in the second set are judged.
[0052] C14, obtaining the first set that has been last updated, and obtaining the universal chemical combination used for the mixture toxicity evaluation.
[0053] In this embodiment, the second set includes all chemical combinations with chemical numbers n-1, n-2, ..., 3, 2. The chemical combination with chemical number n-1 in the second set is judged with the first set. If it is a subset, it is removed. If it is not a subset, it is merged with the first set to form a new first set. Then the chemical combination with chemical number n-2 in the second set is judged with the new first set, until the chemical number in the second set is 2. The final first set is the highest-order frequent item set, which represents the most complex form of each common chemical combination.
[0054] The method of this embodiment can ensure that only the most complex common chemical combinations are retained, avoiding duplication and redundancy.
[0055] In this embodiment, it is assumed that the first set and the second set are as shown in Table 4.
[0056]
[0057] In this embodiment, in Table 4 above, {'C2', 'C3', 'C4'} is the chemical combination with the largest number of chemicals (3), which is the first set. In actual applications, if there are multiple chemical combinations with 3 chemicals, the multiple chemical combinations are respectively placed in the first set. The remaining chemical combinations {'C4', 'C3'}, {'C2', 'C4'}, {'C2', 'C3'}, and {'C1', 'C3'} are the second set, that is, the second set contains 4 chemical combinations. Next, it is determined whether the chemical combination with the number of chemicals (3-1=2) in the second set is a subset of the first set, wherein the chemical combinations {'C4', 'C3'}, {'C2', 'C4'}, and {'C2', 'C3'} are all subsets of the first set {'C2', 'C3', 'C4'}, and therefore, the three chemical combinations are removed from the second set. If the chemical combination {'C1', 'C3'} is not a subset of the first set, the first set is updated according to the chemical combination to obtain a new first set ({'C1', 'C3'}, {'C2', 'C3', 'C4'}).
[0058] In this embodiment, since the chemical combination with the chemical number of 2 is the lowest-order chemical combination, the new set ({'C1', 'C3'}, {'C2', 'C3', 'C4'}) is the final universal chemical combination, as shown in Table 5.
[0059]
[0060] In this embodiment, for the five chemicals detected, there are theoretically 31 chemical combinations, but there are actually only two common chemical combinations finally obtained by the method of this embodiment. These two common chemical combinations occur frequently in the actual environment and are universal. Therefore, the method of this embodiment can screen out a small number of common combinations that actually exist from a large number of chemical combinations that may be formed theoretically.
[0061] In this embodiment, by obtaining a small amount of common chemical combinations, they can be studied as priority mixtures later. As an optional embodiment, each target chemical contained in the common chemical combination is obtained; according to the concentration range of each target chemical in the target environment, for each chemical combination in the common chemical combination, according to the concentration range of the chemicals contained in the chemical combination, a reagent for conducting a mixture toxicity experiment evaluation is configured; based on each configured reagent, a mixture toxicity experiment is performed to obtain a mixture toxicity evaluation result. For example, the common chemical combination {'C2', 'C3', 'C4'} obtained by screening is selected for mixture toxicity research, and according to the concentration range of the three chemicals 'C2', 'C3', and 'C4' in the actual environment, a mixture design is performed for the common chemical combination, and toxicity evaluation and toxicity interaction analysis are performed on each mixture to reveal the toxicity change law of the common chemical combination, and provide practical methods and technologies for implementing mixture toxicity evaluation and risk assessment of mixed pollutants in the actual environment.
[0062] In the related art, in the mixture study, the selected mixture is not necessarily present or ubiquitous in the actual environment, which is out of practical environmental significance. The common chemical combination screened by the method of this embodiment occurs frequently in the actual environment and is universal, and the corresponding mixture toxicity study has more practical environmental significance.
[0063] Figure 5 FIG. 1 is a schematic diagram showing the structure of a device for obtaining a common chemical combination for mixture toxicity assessment provided by an embodiment of the present invention. Figure 5 As shown, the device comprises: The concentration matrix construction module 501 is used to count the total number of chemicals detected at each sampling detection point in the target environment, and for each sampling detection point, construct the chemical detection matrix of the sampling detection point according to the chemical concentration detected at the sampling detection point and the total number of chemicals, and based on the chemical detection matrices of each sampling detection point, construct the compound detection matrix of the target environment; In this embodiment, as an optional embodiment, the chemical detection matrix of the sampling detection point is a row matrix, the number of columns contained in the row matrix is the total number of chemicals, and each column corresponds to a detected chemical.
[0064] A presence / absence matrix conversion module 502 is used to query the mapping relationship between the preset chemicals and the analytical detection limits, and convert the compound detection matrix into a compound presence / absence matrix based on the target chemical concentration in the compound detection matrix and the analytical detection limit mapped to the target chemical; In this embodiment, as an optional embodiment, the matrix conversion module 502 is specifically used for: The target chemical concentration corresponding to each chemical in the compound detection matrix is traversed, and the analytical detection limit mapped to the chemical is obtained from the mapping relationship. If the target chemical concentration is greater than or equal to the analytical detection limit, the target chemical concentration corresponding to the chemical is updated to 1; otherwise, the target chemical concentration corresponding to the chemical is updated to 0.
[0065] A low-order frequent item set construction module 503 is used to construct a first frequent item set of chemicals for each candidate chemical in the compound presence matrix based on the occurrence frequency of the candidate chemical in the compound presence matrix and a preset first frequency threshold; A high-order frequent item set construction module 504 is used to optimize the first frequent item set of chemicals from low-order chemical combinations to high-order chemical combinations based on a preset item set optimization strategy, and sequentially obtain high-order frequent item sets; In this embodiment, as an optional embodiment, the high-order frequent itemset construction module 504 is specifically used for: Combining candidate chemicals in the first frequent item set of chemicals in pairs to obtain binary chemical combinations, obtaining the frequencies of occurrence of the binary chemical combinations in the compound presence / absence matrix, and constructing a second frequent item set of chemicals based on binary chemical combinations whose frequencies of occurrence are greater than or equal to a preset second frequency threshold; For the candidate chemicals in the second frequent item set of chemicals, any three candidate chemical combinations are performed to obtain a third-order chemical combination, and the occurrence frequencies of the third-order chemical combinations in the compound presence / absence matrix are obtained in turn. Based on the third-order chemical combinations whose occurrence frequencies are greater than or equal to a preset third frequency threshold, a third frequent item set of chemicals is constructed until there is no higher-order chemical combination whose frequency is greater than or equal to a preset higher-order frequency threshold.
[0066] In this embodiment, as an optional embodiment, the values of the first frequency threshold, the second frequency threshold and the third frequency threshold are the same. As another optional embodiment, the values of the first frequency threshold, the second frequency threshold and the third frequency threshold may be different from each other or partially the same.
[0067] In this embodiment, as an optional embodiment, any three candidate chemical combinations are performed on the candidate chemicals in the second frequent item set of chemicals, including: Eliminating binary chemical combinations whose occurrence frequency is equal to a preset second frequency threshold from the second frequent item set of chemicals to obtain a filtered second frequent item set of chemicals; Three candidate chemicals are randomly selected from the second frequent item filtering set of chemicals to obtain the third-order chemical combination.
[0068] In this embodiment, as another optional embodiment, the high-order frequent itemset construction module 504 is specifically used for: Construct the root node; In the compound presence / absence matrix corresponding to the first frequent item set of the chemical, the candidate chemicals in the first row are extracted, and the candidate chemicals are arranged in descending order according to the frequency of occurrence of the candidate chemicals, and multiple levels of child nodes are sequentially constructed under the root node, wherein the level of the child node is the number of candidate chemicals contained in the first row, the candidate chemical with a high frequency of occurrence is the parent node of the candidate chemical with a low frequency of occurrence, and the count of the candidate chemicals at each level of the child node is set to 1; Traverse the other rows in the compound presence / absence matrix, arrange the candidate chemicals in descending order according to the frequency of occurrence, traverse each candidate chemical in the sorted row, if the i-th candidate chemical in the row is the same as the candidate chemical in the i-th child node under the root node, add 1 to the count of the child node, if they are not the same, create a child node under the i-1-th child node for the i-th candidate chemical in the row from the i-1-th child node under the root node, set the count of the candidate chemical in the newly created child node to 1, and under the newly created child node, sequentially construct multiple levels of child nodes for the candidate chemicals after the i-th candidate chemical in the row; From the multi-level child nodes, candidate frequent chemicals are selected, and the positions of the candidate frequent chemicals in the multi-level child nodes are located. The root node is traced for each located position to obtain the child node chain corresponding to each position, and the common chemicals contained in each child node chain are obtained. The counts of the common chemicals contained in each child node chain are accumulated to obtain a frequent item set containing the counts of the common chemicals. If the counts of the common chemicals exceed a preset count threshold, the frequent item set corresponding to the counts of the common chemicals is determined to be a high-order frequent item set.
[0069] The chemical combination acquisition module 505 is used to acquire a common chemical combination for mixture toxicity assessment based on the first frequent item set of chemicals and the successively acquired high-order frequent item sets.
[0070] In this embodiment, as an optional embodiment, the chemical combination acquisition module 505 is specifically used to: From the high-order frequent item sets obtained in sequence, obtain the high-order frequent item set containing the maximum number of chemicals, construct a first set based on the high-order frequent item set containing the maximum number of chemicals, and construct a second set based on the remaining frequent item sets except the first set, wherein the maximum number of chemicals is n; Extract all candidate frequent item sets with the number of chemicals being n-1 from the second set. For each candidate frequent item set, if the candidate frequent item set is a subset of any frequent item set in the first set, remove the candidate frequent item set from the second set. If not, merge the candidate frequent item set with the first set and update the first set. Traverse all candidate frequent item sets in the second set whose number of chemicals is less than n-1. For each candidate frequent item set, if the candidate frequent item set is a subset of any frequent item set in the first set, remove the candidate frequent item set from the second set. If not, merge the candidate frequent item set with the updated first set. The first set that is last updated is obtained to obtain the universal chemical combination used for mixture toxicity evaluation.
[0071] In this embodiment, as an optional embodiment, the device further includes: A mixture toxicity assessment module (not shown in the figure), used to obtain each target chemical contained in the general chemical combination; According to the concentration range of each target chemical in the target environment, for each chemical combination in the general chemical combination, according to the concentration range of the chemicals included in the chemical combination, a reagent for conducting a mixture toxicity test is configured; A mixture toxicity test was performed based on each reagent in the configuration to obtain the mixture toxicity assessment results.
[0072] Based on the same inventive concept, an embodiment of the present invention further provides a storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method for obtaining a common chemical combination for mixture toxicity assessment in any possible implementation manner described above are implemented.
[0073] Alternatively, the storage medium may be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0074] Based on the same inventive concept, see Figure 6The embodiment of the present invention also provides an electronic device, including a memory 101 (such as a non-volatile memory), a processor 102, and a computer program stored in the memory 101 and executable on the processor 102. When the processor 102 executes the program, the steps of the method for obtaining a universal chemical combination for mixture toxicity assessment in any possible implementation manner described above can be equivalent to the device for obtaining a universal chemical combination for mixture toxicity assessment as described above. Of course, the processor can also be used to process other data or operations. The electronic device can be a PC, a server, a terminal, and other devices.
[0075] like Figure 6 As shown, the electronic device may also generally include: a memory 103, a network interface 104, and an internal bus 105. In addition to these components, other hardware may also be included, which will not be described in detail.
[0076] It should be pointed out that the above-mentioned device for obtaining common chemical combinations for mixture toxicity assessment can be implemented by software. As a device in a logical sense, it is formed by the processor 102 of the electronic device in which it is located reading the computer program instructions stored in the non-volatile memory into the memory 103 for execution.
[0077] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0078] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for obtaining a universal chemical combination for mixture toxicity assessment, characterized in that: include: Counting the total number of chemicals detected at each sampling and testing point in the target environment, for each sampling and testing point, constructing a chemical detection matrix for the sampling and testing point based on the chemical concentration detected at the sampling and testing point and the total number of chemicals, and based on the chemical detection matrices of each sampling and testing point, constructing a compound detection matrix for the target environment; Constructing a mapping relationship between chemicals and analytical detection limits, and converting the compound detection matrix into a compound presence / absence matrix based on the target chemical concentration in the compound detection matrix and the analytical detection limit mapped to the target chemical; For each candidate chemical in the compound presence / absence matrix, constructing a first frequent item set of chemicals based on the occurrence frequency of the candidate chemical in the compound presence / absence matrix and a preset first frequency threshold; Based on a preset item set optimization strategy, the first frequent item set of the chemical is optimized from a low-order chemical combination to a high-order chemical combination, and high-order frequent item sets are obtained in sequence; Based on the first frequent item set of the chemicals and the successively obtained high-order frequent item sets, a common chemical combination for mixture toxicity assessment and risk evaluation is obtained.
2. The method for obtaining a common chemical combination for mixture toxicity assessment according to claim 1, characterized in that: The method of optimizing the first frequent item set of chemicals from low-order chemical combinations to high-order chemical combinations based on the preset item set optimization strategy to sequentially obtain high-order frequent item sets includes: Combining candidate chemicals in the first frequent item set of chemicals in pairs to obtain binary chemical combinations, obtaining the occurrence frequencies of the binary chemical combinations in the compound presence / absence matrix, and constructing the second frequent item set of chemicals based on binary chemical combinations whose occurrence frequencies are greater than or equal to a preset second frequency threshold; For the candidate chemicals in the second frequent item set of chemicals, any three candidate chemical combinations are performed to obtain a third-order chemical combination, and the occurrence frequencies of the third-order chemical combinations in the compound presence / absence matrix are obtained in turn. Based on the third-order chemical combinations whose occurrence frequencies are greater than or equal to a preset third frequency threshold, a third frequent item set of chemicals is constructed until there is no higher-order chemical combination whose frequency is greater than or equal to a preset higher-order frequency threshold.
3. The method for obtaining a common chemical combination for mixture toxicity assessment according to claim 1 or 2, characterized in that: The step of combining any three candidate chemicals in the second frequent item set of chemicals includes: Eliminating binary chemical combinations whose occurrence frequency is equal to a preset second frequency threshold from the second frequent item set of chemicals to obtain a filtered second frequent item set of chemicals; Three candidate chemicals are randomly selected from the second frequent item filtering set of chemicals to obtain the third-order chemical combination.
4. The method for obtaining a common chemical combination for mixture toxicity assessment according to claim 1 or 2, characterized in that: The method of optimizing the first frequent item set of chemicals from low-order chemical combinations to high-order chemical combinations based on the preset item set optimization strategy to sequentially obtain high-order frequent item sets includes: Construct the root node; In the compound presence / absence matrix corresponding to the first frequent item set of the chemical, the candidate chemicals in the first row are extracted, and the candidate chemicals are arranged in descending order according to the frequency of occurrence of the candidate chemicals, and multiple levels of child nodes are sequentially constructed under the root node, wherein the level of the child node is the number of candidate chemicals contained in the first row, the candidate chemical with a high frequency of occurrence is the parent node of the candidate chemical with a low frequency of occurrence, and the count of the candidate chemicals at each level of the child node is set to 1; Traverse the other rows in the compound presence / absence matrix, arrange the candidate chemicals in descending order according to the frequency of occurrence, traverse each candidate chemical in the sorted row, if the i-th candidate chemical in the row is the same as the candidate chemical in the i-th child node under the root node, add 1 to the count of the child node, if they are not the same, create a child node under the i-1-th child node for the i-th candidate chemical in the row from the i-1-th child node under the root node, set the count of the candidate chemical in the newly created child node to 1, and under the newly created child node, sequentially construct multiple levels of child nodes for the candidate chemicals after the i-th candidate chemical in the row; From the multi-level child nodes, candidate frequent chemicals are selected, and the positions of the candidate frequent chemicals in the multi-level child nodes are located. The root node is traced for each located position to obtain the child node chain corresponding to each position, and the common chemicals contained in each child node chain are obtained. The counts of the common chemicals contained in each child node chain are accumulated to obtain a frequent item set containing the counts of the common chemicals. If the counts of the common chemicals exceed a preset count threshold, the frequent item set corresponding to the counts of the common chemicals is determined to be a high-order frequent item set.
5. The method for obtaining a common chemical combination for mixture toxicity assessment according to any one of claims 1 to 4, characterized in that: The method of obtaining a common chemical combination for mixture toxicity assessment based on the first frequent item set of the chemical and the sequentially obtained high-order frequent item sets includes: From the high-order frequent item sets obtained in sequence, obtain the high-order frequent item set containing the maximum number of chemicals, construct a first set based on the high-order frequent item set containing the maximum number of chemicals, and construct a second set based on the remaining frequent item sets except the first set, wherein the maximum number of chemicals is n; Extract all candidate frequent item sets with the number of chemicals being n-1 from the second set. For each candidate frequent item set, if the candidate frequent item set is a subset of any frequent item set in the first set, remove the candidate frequent item set from the second set. If not, merge the candidate frequent item set with the first set and update the first set. Traverse all candidate frequent item sets in the second set whose number of chemicals is less than n-1. For each candidate frequent item set, if the candidate frequent item set is a subset of any frequent item set in the first set, remove the candidate frequent item set from the second set. If not, merge the candidate frequent item set with the updated first set. The first set that is last updated is obtained to obtain the universal chemical combination used for mixture toxicity assessment.
6. The method for obtaining a common chemical combination for mixture toxicity assessment according to any one of claims 1 to 4, characterized in that: The step of converting the compound detection matrix into a compound presence / absence matrix based on the target chemical concentration in the compound detection matrix and the analytical detection limit mapped to the target chemical comprises: The target chemical concentration corresponding to each chemical in the compound detection matrix is traversed, and the analytical detection limit mapped to the chemical is obtained from the mapping relationship. If the target chemical concentration is greater than or equal to the analytical detection limit, the target chemical concentration corresponding to the chemical is updated to 1; otherwise, the target chemical concentration corresponding to the chemical is updated to 0.
7. The method for obtaining a common chemical combination for mixture toxicity assessment and risk evaluation according to any one of claims 1 to 4, characterized in that: The method comprises: Obtaining each target chemical contained in the universal chemical combination; According to the concentration range of each target chemical in the target environment, for each chemical combination in the general chemical combination, according to the concentration range of the chemicals included in the chemical combination, a reagent for performing a mixture toxicity assessment is configured; A mixture toxicity assessment is performed based on each reagent configured to obtain a mixture toxicity assessment result.
8. A device for obtaining a common chemical combination for mixture toxicity assessment, characterized in that: The device for obtaining a common chemical combination for mixture toxicity assessment comprises: A concentration matrix construction module is used to count the total number of chemicals detected at each sampling detection point in the target environment, and for each sampling detection point, construct a chemical detection matrix for the sampling detection point based on the chemical concentration detected at the sampling detection point and the total number of chemicals, and based on the chemical detection matrices of each sampling detection point, construct a compound detection matrix for the target environment; A presence / absence matrix conversion module is used to query the mapping relationship between the preset chemicals and the analytical detection limits, and convert the compound detection matrix into a compound presence / absence matrix based on the target chemical concentration in the compound detection matrix and the analytical detection limit mapped to the target chemical; A low-order frequent item set construction module is used to construct a first frequent item set of chemicals for each candidate chemical in the compound presence matrix based on the occurrence frequency of the candidate chemical in the compound presence matrix and a preset first frequency threshold; A high-order frequent item set construction module is used to optimize the first frequent item set of chemicals from low-order chemical combinations to high-order chemical combinations based on a preset item set optimization strategy, and obtain high-order frequent item sets in sequence; The chemical combination acquisition module is used to acquire a common chemical combination for mixture toxicity assessment based on the first frequent item set of the chemical and the successively acquired high-order frequent item sets.
9. A storage medium, characterized in that: The storage medium stores a program or an instruction, and when the program or the instruction is executed by a processor, the steps of the method for obtaining a common chemical combination for mixture toxicity assessment as claimed in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for obtaining a common chemical combination for mixture toxicity assessment according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method for measuring pollutant concentration in human exposed environment and related product
CN108956859A
An atmospheric pollution binary mixture health risk evaluation method
CN109583662A
Training method and prediction method of chemical genetic toxicity prediction model
CN114678083A
Toxicity prediction method based on compound secondary mass spectrum data
CN117409871A
Method for evaluating reproduction exposure risk of mixed pollutants
CN117423401A