Antibiotic mixture toxicity prediction method and device
By using the minimum redundancy maximum correlation algorithm and the genetic algorithm to determine the coupled mixture descriptor of antibiotic mixtures, and using the neural network model of the Transformer framework for toxicity prediction, the problem of low prediction accuracy in the prior art is solved, and accurate prediction of the toxicity of antibiotic mixtures and environmental protection are achieved.
Patent Information
- Application Number
- CN202511733170.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-03
AI Technical Summary
Existing methods for predicting the toxicity of antibiotic mixtures have limited applicability and low accuracy, making it impossible to accurately predict the toxicity of antibiotic mixtures and thus hindering effective environmental remediation and protection.
The minimum redundancy maximum correlation algorithm and genetic algorithm are used to analyze the component information of the target antibiotic mixture, determine the coupled mixture descriptor, and use a neural network prediction model based on the Transformer framework to predict toxicity. Accurate prediction is achieved through multi-head attention analysis.
It enables accurate prediction of the toxicity of antibiotic mixtures, and allows for environmental remediation based on the prediction results, thereby achieving the goal of environmental protection.
Smart Images

Figure CN121601075A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of environmental toxicology and ecological risk assessment technology, specifically relating to a method and apparatus for predicting the toxicity of antibiotic mixtures. Background Technology
[0002] Antibiotic mixtures are complex systems composed of two or more single antibiotics, and their components may originate from combination therapy design, environmental residues, or biological metabolic processes. Antibiotic mixtures are widely present in the ecological environment, and with the residues and accumulation of antibiotic mixtures in the environment, the individual components of the mixture significantly enhance their ecotoxicity through synergistic effects, thus posing a profound and irreversible threat to ecosystem stability and public health. Therefore, research on the toxicity and risks of antibiotic mixtures is necessary.
[0003] However, existing methods for predicting the toxicity of antibiotic mixtures have limitations in terms of applicability and accuracy. They cannot accurately predict the toxicity of antibiotic mixtures, and therefore cannot be used to carry out environmental remediation to achieve environmental protection based on the predicted toxicity of antibiotic mixtures. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a method and apparatus for predicting the toxicity of antibiotic mixtures. The method utilizes a minimum redundancy maximum correlation algorithm and a genetic algorithm to analyze and obtain a coupled mixture descriptor for the target antibiotic mixture. This coupled mixture descriptor sets corresponding mixing rules for each of the multiple common descriptors in the target antibiotic mixture. Then, a neural network prediction model based on the Transformer framework is used to analyze the coupled mixture descriptor to obtain the predicted toxicity value of the target antibiotic mixture. This process enables accurate prediction of the toxicity of antibiotic mixtures, allowing for environmental remediation based on the predicted toxicity results, thereby achieving environmental protection.
[0005] In a first aspect, the present invention provides a method for predicting the toxicity of antibiotic mixtures, comprising: acquiring component information of a target antibiotic mixture; the component information of the target antibiotic mixture includes the name, concentration, and proportion of each component among the multiple components contained in the target antibiotic mixture. Based on the component information of the target antibiotic mixture, a minimum redundancy maximum correlation algorithm, and a genetic algorithm, a coupled mixing descriptor for the target antibiotic mixture is determined; the coupled mixing descriptor includes multiple common descriptors of the multiple components in the target antibiotic mixture that best reflect its toxicity, and a mixing rule for each common descriptor; the common descriptors are physicochemical property descriptors, topological chemical descriptors, geometric descriptors, or electronic property descriptors; the mixing rules are molecular linear contribution mixing rules, molecular square contribution mixing rules, or molecular regular contribution mixing rules. Based on the coupled mixing descriptor and toxicity prediction model of the target antibiotic mixture, a predicted toxicity value for the target antibiotic mixture is determined; the toxicity prediction model is a neural network prediction model based on the Transformer framework.
[0006] The antibiotic mixture toxicity prediction method provided by this invention first employs a minimum redundancy maximum correlation algorithm and a genetic algorithm to analyze the component information of the target antibiotic mixture, obtaining a coupled mixture descriptor. This coupled mixture descriptor includes multiple common descriptors representing the various components that best reflect the toxicity of the target antibiotic mixture, as well as the mixing rules for each common descriptor. Then, a toxicity prediction model based on the Transformer framework is used to analyze and predict the coupled mixture descriptor, obtaining the predicted toxicity value of the target antibiotic mixture. In this process, not only are the mixing rules for each common descriptor in the target antibiotic mixture specifically provided, but also multi-head attention analysis is performed on the coupled mixture descriptor using the toxicity prediction model based on the Transformer framework. This achieves parallel processing of each common descriptor and its mixing rules, as well as analysis of global dependencies, enabling accurate prediction of antibiotic mixture toxicity. Therefore, environmental remediation can be carried out based on the predicted toxicity of the antibiotic mixture to achieve environmental protection.
[0007] In one implementation of the first aspect, a coupled mixture descriptor for the target antibiotic mixture is determined based on the component information, the minimum redundancy maximum relevance algorithm, and a genetic algorithm. This includes: calculating an initial descriptor set for the target antibiotic mixture using the SMILES expression; the initial descriptor set includes descriptors for each of the multiple components contained in the target antibiotic mixture. The descriptors in the initial descriptor set are mixed based on molecular linear contribution mixing rules, molecular square contribution mixing rules, and molecular regular contribution mixing rules, respectively, to obtain multiple candidate descriptor sets. The maximum relevance minimum redundancy algorithm is used to screen from the multiple candidate descriptor sets to obtain multiple common descriptors most relevant to the toxicity of the target antibiotic mixture. For each common descriptor, the optimal mixing rule for the common descriptor is adaptively optimized from the molecular linear contribution mixing rules, molecular square contribution mixing rules, and molecular regular contribution mixing rules using a genetic algorithm. The multiple common descriptors and the optimal mixing rule for each common descriptor are combined to obtain the coupled mixture descriptor for the target antibiotic mixture.
[0008] In one implementation of the first aspect, descriptors in the initial descriptor set are mixed based on molecular linear contribution mixing rules, molecular squared contribution mixing rules, and molecular regularized contribution mixing rules to obtain multiple candidate descriptor sets. This includes: mixing descriptors in the initial descriptor pool using molecular linear contribution mixing rules to obtain a first candidate descriptor set; mixing descriptors in the initial descriptor pool using molecular squared contribution mixing rules to obtain a second candidate descriptor set; and mixing descriptors in the initial descriptor pool using molecular regularized contribution mixing rules to obtain a third candidate descriptor set. The multiple candidate descriptor sets include the first candidate descriptor set, the second candidate descriptor set, and the third candidate descriptor set.
[0009] In one implementation of the first aspect, the toxicity prediction value of the target antibiotic mixture is determined based on the coupled mixture descriptor of the target antibiotic mixture and the toxicity prediction model, including: inputting the coupled mixture descriptor of the target antibiotic mixture into the toxicity prediction model, the toxicity prediction model performing multi-head attention analysis on the coupled mixture descriptor of the target antibiotic mixture, and outputting the toxicity prediction value of the target antibiotic mixture.
[0010] Secondly, this invention provides an antibiotic mixture toxicity prediction device, comprising an information acquisition module, a descriptor determination module, and a toxicity prediction module. The information acquisition module is used to acquire component information of the target antibiotic mixture; the component information of the target antibiotic mixture includes the name, concentration, and proportion of each component among the multiple components contained in the target antibiotic mixture. The descriptor determination module is used to determine a coupled mixing descriptor of the target antibiotic mixture based on the component information, a minimum redundancy maximum correlation algorithm, and a genetic algorithm; the coupled mixing descriptor includes multiple common descriptors of the multiple components in the target antibiotic mixture that best reflect its toxicity, and a mixing rule for each common descriptor; the common descriptors are physicochemical property descriptors, topological chemical descriptors, geometric descriptors, or electronic property descriptors; the mixing rules are molecular linear contribution mixing rules, molecular square contribution mixing rules, or molecular regular contribution mixing rules. The toxicity prediction module is used to determine the predicted toxicity value of the target antibiotic mixture based on the coupled mixing descriptor and the toxicity prediction model; the toxicity prediction model is a neural network prediction model based on the Transformer framework.
[0011] Thirdly, the present invention provides an electronic device including a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the method as described in the first aspect above or any implementation thereof.
[0012] Fourthly, the present invention provides a computer-readable storage medium including computer program instructions that, when executed by a computer, cause the computer to perform the method as described in the first aspect above or any implementation thereof.
[0013] Fifthly, the present invention provides a computer program product, including computer program instructions, which, when executed on a computer, cause the computer to perform the method as described in the first aspect above or any of its implementations.
[0014] The technical effects corresponding to the second to fifth aspects and their possible implementations can be referred to the above description of the technical effects of the first aspect and its possible implementations, and will not be repeated here. Attached Figure Description
[0015] Figure 1 This is one of the schematic diagrams of an antibiotic mixture toxicity prediction method provided in the embodiments of this application; Figure 2 This is a second schematic diagram of a method for predicting the toxicity of antibiotic mixtures provided in the embodiments of this application; Figure 3 This is a schematic diagram of an antibiotic mixture toxicity prediction device provided in an embodiment of this application. Detailed Implementation
[0016] In the specification and claims of this invention, the terms "first" and "second," etc., are used to distinguish different objects, rather than to describe a specific order of objects.
[0017] In the embodiments of this application, "and / or" indicates a relationship between objects. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist simultaneously.
[0018] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0019] In the description of this invention, unless otherwise stated, "a plurality of" means two or more. For example, a plurality of shared descriptors means two or more shared descriptors.
[0020] The method and apparatus provided in this application relate to the prediction of the toxicity of antibiotic mixtures, and can be used to predict the toxicity of a target antibiotic mixture based on the component information of the target antibiotic mixture.
[0021] To address the limitations of existing methods for predicting the toxicity of antibiotic mixtures, which suffer from limited applicability and low accuracy, thus hindering accurate prediction of antibiotic mixture toxicity and consequently preventing environmental remediation based on the predicted toxicity, this application provides a method and apparatus for predicting the toxicity of antibiotic mixtures. The method utilizes a minimum redundancy maximum correlation algorithm and a genetic algorithm to analyze and obtain a coupled mixture descriptor for the target antibiotic mixture. This coupled mixture descriptor sets corresponding mixing rules for each of the multiple common descriptors in the target antibiotic mixture. A neural network prediction model based on the Transformer framework is then used to analyze the coupled mixture descriptor to obtain the predicted toxicity value of the target antibiotic mixture. This process enables accurate prediction of antibiotic mixture toxicity, allowing for environmental remediation based on the predicted toxicity, thereby achieving environmental protection.
[0022] For example, the antibiotic mixture toxicity prediction method provided in this embodiment of the invention can be executed by an electronic device with processing capabilities, such as a computer or server. Taking a computer as an example, the hardware components of the computer may include: a processor, memory, a network interface, a user interface, a communication bus, etc.
[0023] The processor controls the electronic device to perform related processing and computation tasks, such as acquiring component information of a target antibiotic mixture, determining the coupled mixing descriptor of the target antibiotic mixture, and determining the toxicity prediction value of the target antibiotic mixture. The processor may include a central processing unit (CPU) or other processors, and may be single-core or multi-core; for example, the processor may include multiple CPUs.
[0024] Memory is used to store computer instructions and related data, such as component information of a target antibiotic mixture, coupled mixture descriptors, and toxicity prediction values. Memory can be random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, disk storage media, or other magnetic storage devices, or any other medium capable of storing program code or data accessible by a computer. Optionally, memory can be integrated into the processor, or it can be independent of the processor.
[0025] A network interface is used for communication between a computer and other devices or communication networks. A network interface can be a transceiver with transmit and receive capabilities. Optionally, a network interface may include standard wired interfaces or wireless interfaces (such as Wi-Fi interfaces, Bluetooth interfaces, and 5G interfaces).
[0026] The communication bus is used to enable communication between different components. For example, the processor, memory, network interface and user interface mentioned above can be interconnected through the communication bus.
[0027] The user interface may include a display screen and an input unit (such as a keyboard). Optionally, the user interface may also include a standard wired interface or a wireless interface.
[0028] Those skilled in the art will understand that the computer described above may include more or fewer components, or combine certain components, or have different component arrangements; the embodiments of this application do not limit this.
[0029] like Figure 1As shown in the embodiments of this application, a method for predicting the toxicity of antibiotic mixtures includes S101-S103.
[0030] S101. Obtain component information of the target antibiotic mixture.
[0031] For example, the component information of the target antibiotic mixture may include the name, concentration, and proportion of each component in the target antibiotic mixture. The multiple components in the target antibiotic mixture may include sulfonamide antibiotics, quinolone antibiotics, and tetracyclines. This application embodiment does not limit the specific types of the multiple components in the target antibiotic mixture.
[0032] S102. Based on the component information of the target antibiotic mixture, the minimum redundancy maximum correlation algorithm, and the genetic algorithm, determine the coupled mixture descriptor of the target antibiotic mixture.
[0033] In this embodiment, the coupled mixing descriptor includes multiple shared descriptors for various components of the target antibiotic mixture that best represent its toxicity, and a mixing rule for each shared descriptor. The shared descriptors are physicochemical property descriptors, topological chemical descriptors, geometric descriptors, or electronic property descriptors. The mixing rule is a molecular linear contribution mixing rule, a molecular square contribution mixing rule, or a molecular canonical contribution mixing rule.
[0034] In one implementation, combined with Figure 1 ,like Figure 2 As shown, S102 includes S1021-S1025.
[0035] S1021. The initial descriptor set of the target antibiotic mixture is calculated using the SMILES expression.
[0036] The aforementioned initial set of descriptors includes descriptors for each of the multiple components contained in the target antibiotic mixture. It is understood that a descriptor, as a fundamental concept commonly used in this technical field, is a numerical or quantitative representation used to describe a specific physical, chemical, or structural property of a molecule (or more broadly, a chemical structure).
[0037] In this embodiment, the Padelpy toolkit is used to calculate the descriptor for each component in the target antibiotic mixture. Specifically, Padelpy, a Python wrapper based on PaDEL-Descriptor, obtains the corresponding SMILES expression from the CAS number of the substance, and then uses Padelpy to calculate the physicochemical property descriptor, topological chemical descriptor, geometric descriptor, and electronic property descriptor for each substance based on the SMILES expression. Since the SMILES expression is common knowledge in this technical field, the specific content of the SMILES expression will not be described in detail here.
[0038] S1022. The descriptors in the initial descriptor set are mixed according to the molecular linear contribution mixing rule, the molecular square contribution mixing rule, and the molecular regular contribution mixing rule, respectively, to obtain multiple candidate descriptor sets.
[0039] Optionally, the aforementioned multiple candidate descriptor sets include a first candidate descriptor set, a second candidate descriptor set, and a third candidate descriptor set; S1022 includes steps 1 to 3.
[0040] Step 1: Mix the descriptors in the initial descriptor pool using the molecular linear contribution mixing rule to obtain the first candidate descriptor set.
[0041] The calculation formula for the above molecular linear contribution mixing rule is shown in the following formula (1).
[0042] in, Indicates the linear contribution value of the molecule. i This represents the i-th component in the target antibiotic mixture. This represents the concentration fraction of the i-th component in the target antibiotic mixture. The descriptor representing the i-th component. C This indicates component C in the target antibiotic mixture. A This indicates another component A in the target antibiotic mixture.
[0043] Step 2: Mix the descriptors in the initial descriptor pool using the numerator square contribution mixing rule to obtain the second candidate descriptor set.
[0044] The calculation formula for the above molecular square contribution mixing rule is shown in the following formula (2).
[0045] in, This represents the squared contribution value of the numerator.
[0046] Step 3: Mix the descriptors in the initial descriptor pool using the molecular regular contribution mixing rule to obtain the third candidate descriptor set.
[0047] The calculation formula for the above molecular canonical contribution mixing rule is shown in the following formula (3).
[0048] in, This represents the canonical contribution value of the molecule.
[0049] In one embodiment of S1022 above, a coupled mixing rule is constructed by selecting the optimal mixing rule for different types of mixture descriptors. Specifically, three common mixing rules—molecular linear contribution mixing rule, molecular square contribution mixing rule, and molecular regular contribution mixing rule—are first used to calculate the first candidate descriptor set, the second candidate descriptor set, and the third candidate descriptor set, respectively.
[0050] Specifically, the first set of candidate descriptors includes: ['RPSA', 'naaaC', 'minHBint5', 'MATS5c', 'AATSC6p', 'ETA_Epsilon_2', 'SHsOH', 'GATS4e', 'mindssC', 'minHsOH', 'MATS6c', 'GATS4c', 'TDB5s', 'ATSC4c', 'AATSC6i', 'AATSC0v', 'MATS5e', 'AATS5s', 'minsOH', 'minHBint2'].
[0051] The second set of candidate descriptors includes: ['RPSA', 'AATS5s', 'mindssC', 'MATS6p', 'THSA', 'MATS5c', 'nHBint2', 'E1m', 'TDB5s', 'minHsOH', 'AATS5e', 'ETA_Epsilon_2', 'minHssNH', 'minHCsats', 'GATS4c', 'FNSA-1', 'hmin', 'TPSA', 'minHBint3', 'GATS4e'].
[0052] The third set of candidate descriptors includes: ['RPSA', 'naaaC', 'MATS5c', 'minHBint5', 'AATSC6p', 'nHBint3', 'minHsOH', 'minwHBa', 'MATS5s', 'AATSC2m', 'minsssCH', 'ETA_Epsilon_2', 'GATS4e', 'ATSC4c', 'MATS6c', 'mindssC', 'TDB5s', 'GATS4c', 'MATS6p', 'VE1_Dzs'].
[0053] S1023. The maximum correlation minimum redundancy algorithm is used to select multiple common descriptors that are most relevant to the toxicity of the target antibiotic mixture from multiple candidate descriptor sets.
[0054] Understandably, this application employs the Maximum Relevance Minimum Redundancy (mRMR) algorithm to screen common descriptors, aiming to select features with maximum relevance and minimum redundancy, thereby improving model performance. During the MRMR feature selection process, the algorithm prioritizes features with strong correlation to the target variable and avoids selecting redundant features highly correlated with already selected features. When implementing the MRMR algorithm, the Mutual Information Quotient (MIR) is used as the selection criterion, and the number of common descriptors to be screened is set. m = 20. Since the mRMR algorithm is a commonly used technique in this field, the specific process of using the mRMR algorithm to select multiple common descriptors most relevant to the toxicity of the target antibiotic mixture from multiple candidate descriptor sets will not be described in detail in this embodiment.
[0055] Continuing with an example from S1022 above, the mRMR algorithm is applied to the first candidate descriptor set, the second candidate descriptor set, and the third candidate descriptor set for feature selection to identify the top 20 descriptors most relevant to the mixed toxicity. Among the mixed descriptors obtained from the above three sets of mixing rules, there are multiple common descriptors, namely [RPSA, MATS5c, ETA_Epsilon_2, ATS4e, mindssC, minHsOH, GATS4c, TDB5, naaaC, minHBint5, AATSC6p, MATS6c, ATSC4c, AATS5s, MATS6p]. This indicates that these multiple common descriptors all exhibit a high correlation with the mixed toxicity under different mixing rules, meaning that these multiple common descriptors best reflect the toxicity of the antibiotic mixture.
[0056] S1024. For each common descriptor among multiple common descriptors, the optimal mixing rule for the common descriptor is adaptively optimized from the molecular linear contribution mixing rule, the molecular square contribution mixing rule, and the molecular regular contribution mixing rule using a genetic algorithm.
[0057] S1025. Combine multiple shared descriptors and the optimal mixing rules of each shared descriptor to obtain a coupled mixing descriptor for the target antibiotic mixture.
[0058] Continuing with an example from S1022 above, in the multiple shared descriptors, each shared descriptor corresponds to three different mixture values, derived from the calculation results of three mixture rules. To further determine the optimal mixture calculation method for each shared descriptor, a genetic algorithm is used for adaptive optimization selection. In the genetic algorithm optimization process implemented in this invention, the following parameter configurations are used: the number of generations is set to 100, the population size to 30, and the mutation probability to 0.1; the selection mechanism adopts a roulette wheel selection method based on fitness ratio, and the crossover method adopts single-point crossover; the model evaluation uses a random forest regressor, with its decision tree number set to 100 and random state set to 42; the fitness evaluation index is negative mean squared error (-MSE). When the GA fitness is minimized, the 15 descriptors find their respective optimal mixture rules. The final results show that: ETA_Epsilon_2, GATS4e, GATS4c, TDB5s, naaaC, minHBint5, AATSC6p, ATSC4c, and AATS5s use the linear molecular contribution rule; RPSA, mindssC, minHsOH, MATS6c, and MATS6p use the square molecular contribution mixing rule, while MATS5c is suitable for the regular molecular contribution rule. The final result is a coupled mixing descriptor for the target antibiotic mixture, as shown in the following formula (4).
[0059] in, R k ( ) indicates that it is applied to the first k A mixing rule for each descriptor, which is selected from at least three predefined mixing rules. This represents the nth predefined mixing rule.
[0060] S103. Based on the coupled mixture descriptor and toxicity prediction model of the target antibiotic mixture, determine the toxicity prediction value of the target antibiotic mixture.
[0061] The above S103 includes the following process.
[0062] The coupled mixture descriptor of the target antibiotic mixture is input into the toxicity prediction model. The toxicity prediction model performs multi-head attention analysis on the coupled mixture descriptor of the target antibiotic mixture and outputs the toxicity prediction value of the target antibiotic mixture.
[0063] In one application scenario of S103 above, the component information, toxicity test information, and coupled mixture descriptor of the target antibiotic mixture are input into the toxicity prediction model using one-hot encoded data. The toxicity prediction model performs multi-head attention analysis on the component information, toxicity test information, and coupled mixture descriptor of the target antibiotic mixture and outputs the toxicity prediction value of the target antibiotic mixture.
[0064] For example, the predicted toxicity values for the aforementioned target antibiotic mixture can be the median effective concentration (EC50), median lethal concentration (LC50), no-observable-effect concentration (NOEC), and their logarithmic transformations. EC50 refers to the concentration of a chemical substance that, within a given exposure time, causes a specific effect (non-mortality) in 50% of the test organisms. LC50 refers to the concentration of a chemical substance that causes the death of 50% of the test organisms within a given exposure time. NOEC refers to the highest experimental concentration, within a specified exposure time, where, through statistical analysis, no significant adverse effects were observed compared to a blank control group. EC50 and LC50 are used to quantify the intensity of toxicity (especially acute toxicity), while NOEC is used to find the threshold of toxicity, providing a scientific basis for developing safety standards to protect the environment and organisms.
[0065] It should be noted that the aforementioned toxicity prediction model is a neural network prediction model based on the Transformer framework. This means that the toxicity prediction model refers to a neural network prediction model built using the Transformer framework, trained on a dataset composed of coupled hybrid descriptors. This toxicity prediction model may contain only an encoder, such as existing models like BERT, RoBERTa, ALBERT, or DeBERTa. Alternatively, it may include both an encoder and a decoder, such as existing models like T5 or BART. The aforementioned neural network prediction model built using the Transformer framework is a commonly used artificial intelligence model in this technical field, and the embodiments of this application do not modify the model structure. The specific hierarchical structure of the aforementioned model will not be elaborated upon in the embodiments of this application.
[0066] Optionally, the above-mentioned toxicity prediction model can also be a random forest model, an XGBOOST model, or a KNN model that can predict the toxicity of antibiotic mixtures. This application embodiment does not further limit the model on which the above-mentioned toxicity prediction model is based.
[0067] The training process of the above toxicity prediction model in an application scenario is given below.
[0068] Step 1: Construct the training dataset.
[0069] Step 1.1: Collect measured toxicity data of antibiotic-related mixtures from publicly published literature and public databases, and record the corresponding test species, exposure time and experimental conditions of toxicity endpoint.
[0070] Specifically, using "Mixture toxicity" and "sulfonamides," as well as "antibiotics" and "Mixture toxicity" as keywords, relevant literature published between 2000 and 2025 was retrieved from the ACS, ScienceDirect, and Web of Science databases, totaling 10,881 articles. From these, articles providing complete reports on the composition of the mixed system, CAS numbers of the substances, component concentration / proportion information, test species for toxicity testing, mixed toxicity endpoints, and toxicity values were selected and manually compiled to construct a preliminary dataset. Duplicate samples and samples with missing mixed toxicity information were then removed from the dataset, and the data was compiled according to different species and different toxicity test information.
[0071] Step 1.2: Calculate and process the single-substance descriptors and characteristics of the mixture: Calculate the single-substance descriptors of each component in the mixture. The descriptor pool includes the physicochemical property descriptors, topological chemical descriptors, geometric descriptors and electronic property descriptors of each substance; invalid descriptors and descriptors with small contributions are removed, and zero-value filling is performed on null descriptors.
[0072] Specifically, the Padelpy toolkit was used to calculate the molecular descriptors of each component in the mixture. Padelpy is a Python wrapper tool based on PaDEL-Descriptor. The SMILES expression corresponding to each substance was obtained through its CAS number, and with the help of Padelpy, physicochemical property descriptors, topological chemical descriptors, geometric descriptors, and electronic property descriptors for each substance were calculated based on the SMILES, resulting in a total of 1875 molecular descriptors.
[0073] Clean the descriptors: 1. Delete invalid descriptors with all values of 0 in all samples; 2. Use the following formula (5) to remove relevant features with a Pearson correlation coefficient greater than 0.95; 3. Use the following formula (6) to delete descriptors with a standard deviation less than 0.05. Thus, the component information of the mixture of various antibiotics is obtained.
[0074] in: The Pearson correlation coefficient is used. n For the sample size, These are the values of the i-th sample in descriptors X and Y. The mean of descriptor X, Let Y be the mean of the descriptor.
[0075] in: Standard deviation, n The number of samples.
[0076] Step 1.3: Construct a training dataset using component information of a mixture of various antibiotics.
[0077] After obtaining the component information of the multiple antibiotic mixtures, a coupled mixture descriptor for each antibiotic mixture in the multiple antibiotic mixtures is obtained by referring to the specific embodiments of S1021-S1025 above, and the coupled mixture descriptors for each antibiotic mixture in the multiple antibiotic mixtures are integrated to obtain a set of coupled mixture descriptors.
[0078] Specifically, the dataset is partitioned based on the coupled hybrid descriptor set. Each toxicity endpoint is divided into a training set and a test set in an 8:2 ratio. The training and test sets of each endpoint are then merged to form the training dataset. Each experimental endpoint is standardized using Z-score normalization.
[0079] Specifically, in the training dataset, the training data for each antibiotic mixture in the multiple antibiotic mixtures includes component information, toxicity test information, coupled mixture descriptors, and measured toxicity values. During the training of the toxicity prediction model, one-hot encoding is used to input the component information, toxicity test information, and coupled mixture descriptors of the antibiotic mixtures into the toxicity prediction model. The measured toxicity values corresponding to the antibiotic mixtures are compared with the toxicity prediction values output by the toxicity prediction model to determine the prediction accuracy of the toxicity prediction model.
[0080] The aforementioned component information includes the name, concentration, and proportion of each component in the antibiotic mixture, and the aforementioned toxicity test information includes the test species, toxicity endpoint, exposure time, and toxic effects of the antibiotic mixture.
[0081] Step 2: Construct a neural network prediction model based on the Transformer framework.
[0082] A neural network prediction model based on the Transformer framework is constructed, consisting of two main modules: a Transformer Encoder and a Neural Network (NN), implemented using the PyTorch 2.5.1 deep learning framework. First, the Transformer Encoder utilizes Multi-Head Self-Attention (MHSA) to model global dependencies between features, and combines this with a Feed-Forward Network (FFN) to perform nonlinear transformations on the information, thereby extracting semantically rich high-order feature representations. Subsequently, the extracted features are input into the neural network module for regression prediction. This module consists of multiple fully connected layers and employs the ReLU (Rectified Linear Unit) activation function to enhance the model's nonlinear expressive power and fitting performance. To improve the model's robustness and generalization ability, a 5-Fold Cross-Validation strategy is introduced during training. The training set was divided into five subsets. The model was trained on four of the subsets and validated on the remaining subset. This process was repeated five times, and the average results were taken. Finally, the model parameters with the best performance were selected for prediction on the test set.
[0083] In the neural network model construction of this application embodiment, each key parameter has a clear definition and function: (1) hidden_layer_sizes determines the depth structure and complexity of the network, and represents the feature abstraction ability of the model; (2) The choice of activation function directly affects the nonlinear expressive ability of the network and is related to the degree of model fitting to complex relationships; (3) learning_rate_init, as the initial step size parameter in the training process, controls the magnitude of weight updates and has a key impact on the convergence speed and stability of the model; (4) The selection of the solver optimization algorithm determines the search strategy of the parameter space, which directly affects the training efficiency and final performance; (5) max_iter sets the termination condition of the training process, and balances model accuracy and computational cost by controlling the number of iterations; (6) random_state ensures the reproducibility of the model initialization and training process, providing a guarantee for the consistency of experimental results. The coordinated configuration of these parameters jointly determines the model's architectural features, training dynamics, and final performance.
[0084] Step 3: Use the training dataset constructed in Step 1 to train the neural network prediction model constructed in Step 2 to obtain the toxicity prediction model.
[0085] Specifically, the coupled hybrid descriptor set and three traditional hybrid descriptor sets were used as inputs to the neural network model, and their predictive performance was evaluated on the test set. The evaluation metrics included Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and Coefficient of Determination (R²), calculated as shown in Equations (7), (8), and (9). The results show that this new coupled hybrid descriptor set performs best as model input. The model parameters and results based on the four hybrid rules are shown in Tables 1 and 2. In the test set, the coupled hybrid rule has higher accuracy and lower variance compared to the linear hybrid rule. Compared to using a single hybrid rule globally, the ensemble strategy has significant advantages. The newly developed ensemble strategy with coupled descriptor set shows a more significant improvement in prediction performance, increasing the contribution by 12.84% compared to the numerator squared contribution, resulting in more accurate model predictions.
[0086] in, and It is the first i The observed and predicted values of the mixture, respectively, It is the average of the observed values.
[0087] Table 1 Model running parameters based on each hybrid rule
[0088] Table 2. Test set model performance results based on each hybrid rule.
[0089] Regarding Table 2 above, the meanings of each indicator in the table are as follows.
[0090] (1) Mean absolute error (MAE) measures the average absolute difference between the model prediction and the actual value, and can intuitively reflect the average size of the prediction error.
[0091] (2) Root mean square error (RMSE) is obtained by calculating the arithmetic square root of the mean of the squares of the deviations between the predicted and actual values. It is more sensitive to larger errors in the prediction and is an important indicator for measuring the accuracy of the prediction.
[0092] (3) The coefficient of determination (R²) is used to characterize the proportion of the variability of the dependent variable that the model can explain. Its value is between 0 and 1. The higher the value, the better the model fits the data. Step 4: Verify the predictive ability of the toxicity prediction model.
[0093] Step 4.1: Use SHAP and local cumulative effects to jointly analyze key descriptors and the direction of influence.
[0094] In a further embodiment, interpretability analysis and application domain determination are performed. The SHAP (Shapley Additive exPlanations) and Accumulated Local Effects (ALE) methods are used to interpret features in the machine learning model that have a key impact on mixed toxicity prediction. SHAP, based on the Shapley value in game theory, measures the importance of a feature in the global prediction; while ALE reveals the direction of the feature's influence on the prediction result. ALE calculates the average prediction change of the model across multiple intervals by dividing the feature value into multiple intervals, and then sums these local effects to form an ALE curve reflecting the overall impact of the feature. The Euclidean distance (structural domain) and standardized residuals (response domain) of the training dataset are used to define the spatial definition of the application domain for accurate prediction at the confidence level. Two scientifically accepted confidence levels (i.e., 95% and 99%) are used to define the application domain.
[0095] According to the explanatory analysis method described above, the SHAP ranking of features by importance showed that RPSA, minHCsats, MATS6p, TDB5s, nHBint2, nHBint3, minHsOH, AATSC6p, ATSC4c, FNSA-1, species, and effect contributed most significantly to the mixed toxicity value. Local cumulative effect results indicated that minHCsats and FNSA-1 were negatively correlated with mixed toxicity, while RPSA, nHBint3, MATS6p, TDB5s, minHsOH, and ATSC4c had a positive effect on mixed toxicity. AATSC6p and nHBint2 had a positive effect on mixed toxicity within a certain range, but a negative effect beyond that range.
[0096] Step 4.2: Calculate the structural domain and response domain based on the training set to obtain the applicable domain of the model.
[0097] Based on the application domain, the applicable range of the model is determined to be [0.364, 7.584] and [-0.775, 8.722] for the Euclidean distance of the training set.
[0098] Step 4.3: Conduct a mixed risk assessment of the coexisting antibiotics measured in the environment based on the model prediction structure.
[0099] In a further embodiment, the above mixed toxicity prediction results are applied to risk assessment. The risk quotient (RQ) is used for potential risk assessment. RQ is obtained by calculating the ratio of the environmental concentration (MEC) to the predicted no-effect concentration (PNEC), as shown in Equation 10. The MEC in the mixture is obtained by adding the concentrations of each component in the sample, and the PNEC is obtained by dividing the predicted EC50 of the mixture by the assessment factor (AF), as shown in Equation 11. To consider the uncertainty in risk assessment and provide a more conservative assessment, AF is set to 1000. For risk level classification, the following risk assessment ranking criteria are used: (1) RQ < 0.01, no risk; (2) 0.01 ≤ RQ < 0.1, low risk; (3) 0.1 ≤ RQ < 1, medium risk; (4) RQ ≥ 1, high risk.
[0100] In summary, the antibiotic mixture toxicity prediction method provided in this application first analyzes the component information of the target antibiotic mixture using a minimum redundancy maximum correlation algorithm and a genetic algorithm to obtain a coupled mixture descriptor. This coupled mixture descriptor includes multiple common descriptors representing the various components that best reflect the toxicity of the target antibiotic mixture, as well as the mixing rules for each common descriptor. Then, a toxicity prediction model based on the Transformer framework analyzes and predicts the coupled mixture descriptor to obtain the predicted toxicity value of the target antibiotic mixture. In this process, not only are the mixing rules for each common descriptor in the target antibiotic mixture specifically provided, but also multi-head attention analysis is performed on the coupled mixture descriptor using a toxicity prediction model based on the Transformer framework. This achieves parallel processing of each common descriptor and its mixing rules, as well as analysis of global dependencies, enabling accurate prediction of the antibiotic mixture toxicity. Therefore, environmental remediation can be carried out based on the predicted toxicity of the antibiotic mixture to achieve environmental protection.
[0101] Accordingly, embodiments of this application provide a device for predicting the toxicity of antibiotic mixtures, such as... Figure 3 As shown, it includes an information acquisition module 501, a descriptor determination module 502, and a toxicity prediction module 503.
[0102] The information acquisition module 501 is used to acquire component information of the target antibiotic mixture; the component information of the target antibiotic mixture includes the name, concentration, and proportion of each component among the multiple components contained in the target antibiotic mixture. For example, the information acquisition module 501 is used to implement S101 of the above method.
[0103] The descriptor determination module 502 is used to determine the coupled mixing descriptor of the target antibiotic mixture based on the component information, the minimum redundancy maximum correlation algorithm, and the genetic algorithm. The coupled mixing descriptor includes multiple shared descriptors of various components in the target antibiotic mixture that best represent its toxicity, and a mixing rule for each shared descriptor. The shared descriptors can be physicochemical property descriptors, topological chemical descriptors, geometric descriptors, or electronic property descriptors. The mixing rule can be a molecular linear contribution mixing rule, a molecular squared contribution mixing rule, or a molecular regular contribution mixing rule. For example, the descriptor determination module 502 is used to implement S102 of the above method.
[0104] The toxicity prediction module 503 is used to determine the predicted toxicity value of the target antibiotic mixture based on a coupled mixture descriptor and a toxicity prediction model; the toxicity prediction model is a neural network prediction model based on the Transformer framework. For example, the toxicity prediction module 503 is used to implement S103 of the above method.
[0105] Optionally, the descriptor determination module 502 is specifically used to: calculate an initial descriptor set of the target antibiotic mixture using the SMILES expression; the initial descriptor set includes descriptors for each of the multiple components contained in the target antibiotic mixture. The descriptors in the initial descriptor set are mixed based on molecular linear contribution mixing rules, molecular square contribution mixing rules, and molecular regular contribution mixing rules, respectively, to obtain multiple candidate descriptor sets. A maximum correlation minimum redundancy algorithm is used to screen from the multiple candidate descriptor sets to obtain multiple common descriptors most relevant to the toxicity of the target antibiotic mixture. For each common descriptor, a genetic algorithm is used to adaptively optimize and select the best mixing rule for the common descriptor from the molecular linear contribution mixing rules, molecular square contribution mixing rules, and molecular regular contribution mixing rules. The multiple common descriptors and the best mixing rule for each common descriptor are combined to obtain a coupled mixed descriptor of the target antibiotic mixture. For example, the descriptor determination module 502 is specifically used to implement S1021-S1025 of the above method.
[0106] Each module of the above-mentioned antibiotic mixture toxicity prediction device can also be used to perform other steps in the above method embodiments. All relevant contents involved in the above method embodiments can be referred to in the functional description of the corresponding functional module, and will not be repeated here.
[0107] This application also provides an electronic device, including: a processor and a memory coupled to the processor; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to cause the electronic device to perform the methods described in the above embodiments. The processor can implement the information acquisition module 501, the descriptor determination module 502, and the toxicity prediction module 503; the memory can also be used to store component information of the target antibiotic mixture, coupled mixture descriptors, and toxicity prediction values, etc.
[0108] This application also provides a computer-readable storage medium including a computer program that, when run on a computer, performs the methods described in the above embodiments.
[0109] This application also provides a computer program product, which includes computer program instructions that, when run on a computer, execute the methods described in the above embodiments.
[0110] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for predicting the toxicity of antibiotic mixtures, characterized in that, include: Obtain component information of the target antibiotic mixture; The component information of the target antibiotic mixture includes the name, concentration, and proportion of each component among the multiple components contained in the target antibiotic mixture; Based on the component information of the target antibiotic mixture, the minimum redundancy maximum correlation algorithm, and the genetic algorithm, a coupled mixing descriptor for the target antibiotic mixture is determined; The coupled mixing descriptor includes multiple common descriptors of various components in the target antibiotic mixture that best reflect the toxicity of the target antibiotic mixture, and a mixing rule for each of the multiple common descriptors; the common descriptors are physicochemical property descriptors, topological chemical descriptors, geometric descriptors, or electronic property descriptors; the mixing rule is a molecular linear contribution mixing rule, a molecular square contribution mixing rule, or a molecular regular contribution mixing rule; Based on the coupled mixture descriptor and toxicity prediction model of the target antibiotic mixture, the toxicity prediction value of the target antibiotic mixture is determined; the toxicity prediction model is a neural network prediction model based on the Transformer framework.
2. The method as described in claim 1, characterized in that, The step of determining the coupled mixture descriptor of the target antibiotic mixture based on the component information, the minimum redundancy maximum correlation algorithm, and the genetic algorithm includes: An initial set of descriptors for the target antibiotic mixture is calculated using the SMILES expression; the initial set of descriptors includes descriptors for each of the multiple components contained in the target antibiotic mixture. The descriptors in the initial descriptor set are mixed based on the molecular linear contribution mixing rule, the molecular square contribution mixing rule, and the molecular regular contribution mixing rule, respectively, to obtain multiple candidate descriptor sets. The maximum relevance minimum redundancy algorithm is used to select the multiple common descriptors most relevant to the toxicity of the target antibiotic mixture from the multiple candidate descriptor sets; For each of the plurality of shared descriptors, the optimal mixing rule for the shared descriptor is adaptively optimized and selected from the molecular linear contribution mixing rule, the molecular square contribution mixing rule, and the molecular regular contribution mixing rule using a genetic algorithm. The coupled mixing descriptor of the target antibiotic mixture is obtained by combining the plurality of shared descriptors and the optimal mixing rule of each shared descriptor.
3. The method as described in claim 2, characterized in that, The descriptors in the initial descriptor set are mixed based on molecular linear contribution mixing rules, molecular squared contribution mixing rules, and molecular regularized contribution mixing rules, respectively, to obtain multiple candidate descriptor sets, including: The descriptors in the initial descriptor pool are mixed using the molecular linear contribution mixing rule to obtain the first candidate descriptor set; The descriptors in the initial descriptor pool are mixed using the numerator square contribution mixing rule to obtain a second candidate descriptor set; The descriptors in the initial descriptor pool are mixed using a molecular regularized contribution mixing rule to obtain a third candidate descriptor set; The plurality of candidate descriptor sets include the first candidate descriptor set, the second candidate descriptor set, and the third candidate descriptor set.
4. The method as described in claim 1 or 2, characterized in that, The determination of the predicted toxicity value of the target antibiotic mixture based on the coupled mixture descriptor and toxicity prediction model includes: The coupled mixture descriptor of the target antibiotic mixture is input into the toxicity prediction model, which performs multi-head attention analysis on the coupled mixture descriptor of the target antibiotic mixture and outputs the toxicity prediction value of the target antibiotic mixture.
5. A device for predicting the toxicity of antibiotic mixtures, characterized in that, It includes an information acquisition module, a descriptor determination module, and a toxicity prediction module; The information acquisition module is used to acquire component information of the target antibiotic mixture; the component information of the target antibiotic mixture includes the name, concentration, and proportion of each component among the multiple components contained in the target antibiotic mixture. The descriptor determination module is used to determine a coupled mixing descriptor for the target antibiotic mixture based on the component information, a minimum redundancy maximum correlation algorithm, and a genetic algorithm. The coupled mixing descriptor includes multiple shared descriptors for various components in the target antibiotic mixture that best represent its toxicity, and a mixing rule for each shared descriptor. The shared descriptors are physicochemical property descriptors, topological chemical descriptors, geometric descriptors, or electronic property descriptors. The mixing rule is a molecular linear contribution mixing rule, a molecular square contribution mixing rule, or a molecular regular contribution mixing rule. The toxicity prediction module is used to determine the toxicity prediction value of the target antibiotic mixture based on the coupled mixture descriptor and toxicity prediction model of the target antibiotic mixture; the toxicity prediction model is a neural network prediction model based on the Transformer framework.
6. An electronic device, characterized in that, The device includes a processor and a memory coupled to the processor; the memory is used to store computer instructions, which, when the electronic device is running, are executed by the processor to cause the electronic device to perform the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, It includes computer program instructions that, when executed by a computer, cause the computer to perform the method as described in any one of claims 1 to 4.
8. A computer program product, characterized in that, It includes computer program instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 4.