Artificial intelligence-based antibacterial drug molecule screening method and application thereof

By constructing a hierarchical virtual screening system and combining deep learning with traditional computing tools, the problems of low efficiency and insufficient accuracy of existing antibacterial drug molecule screening methods are solved. This enables efficient and accurate screening of antibacterial drug candidate molecules with novel structures and good drug-like properties, and is applicable to the screening of a variety of antibacterial drugs.

CN122135786APending Publication Date: 2026-06-02WUHAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2026-02-28
Publication Date
2026-06-02

Smart Images

  • Figure CN122135786A_ABST
    Figure CN122135786A_ABST
Patent Text Reader

Abstract

This invention provides an artificial intelligence-based method for screening antimicrobial drug molecules and its applications, specifically relating to the field of drug screening technology. The method constructs a hierarchical virtual screening system. First, it performs two-stage screening using an affinity prediction model and molecular docking tools, and then constructs the affinity prediction model through machine learning. Subsequently, it integrates drug-likeness assessments such as water solubility and toxicity, as well as molecular similarity calculations, to reduce the probability of obtaining low-drug-likeness drug molecules or potentially cross-resistant molecules, filtering candidate molecules hierarchically. Finally, it verifies antimicrobial activity. This invention solves the problems of low efficiency and system scarcity in screening hundreds of millions of molecules by integrating deep learning with traditional computational methods. While improving screening accuracy, it introduces drug resistance risk screening, effectively reducing research components and shortening the cycle, providing high-quality candidate molecules for the clinical translation of anti-drug-resistant drugs, and has significant clinical translational value and application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of drug screening technology, specifically relating to an artificial intelligence-based molecular screening method for antibacterial drugs and its application. Background Technology

[0002] The emergence and spread of bacterial resistance has impacted the effectiveness of antimicrobial treatment. In current clinical practice, infections caused by Gram-negative bacteria are frequent, and drug-resistant strains are constantly emerging. The therapeutic effects of traditional antimicrobial drugs are declining, while the discovery of new antimicrobial drugs lags far behind the evolution of drug-resistant bacteria.

[0003] Furthermore, existing antimicrobial molecule screening methods also have significant shortcomings. On the one hand, traditional screening relies on extensive chemical synthesis and in vitro experiments, which are cumbersome and costly, making it difficult to meet the screening needs of molecular libraries with hundreds of millions of molecules. On the other hand, single virtual screening tools have limitations. While deep learning models are fast, they lack accurate binding energy calculations, and while traditional molecular docking tools have high accuracy, they are inefficient. Moreover, neither fully integrates molecular druggability prediction and drug resistance risk exclusion mechanisms, resulting in screened molecules often having poor druggability and being prone to inducing drug resistance, which seriously affects the success rate and cycle of drug development.

[0004] Therefore, developing an efficient, precise, and comprehensive method for screening antimicrobial molecules to rapidly identify candidate molecules with novel structures, good drug-like properties, and low risk of drug resistance from a library of hundreds of millions of molecules is of great significance for promoting the development of new antimicrobial drugs and addressing the crisis of drug-resistant bacterial infections. Summary of the Invention

[0005] The purpose of this invention is to address the aforementioned shortcomings of existing technologies by providing an artificial intelligence-based method for screening antimicrobial drug molecules and its application. This method aims to solve the problems of low efficiency, insufficient accuracy, high cost, and long development cycle in existing antimicrobial molecule screening. By constructing a multi-dimensional integrated screening system, this invention achieves efficient and accurate screening of antimicrobial molecules, providing novel candidate molecules for the treatment of drug-resistant Gram-negative bacteria.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: The first aspect of this invention provides an artificial intelligence-based method for screening antimicrobial drug molecules, comprising the following steps: A hierarchical virtual screening system is constructed, including a first-level screening of a candidate compound library based on an affinity prediction model; a second-level screening of the results of the first-level screening based on a molecular docking tool; and inputting the results of the second-level screening as training data into a machine learning model to construct a molecule-target protein affinity prediction model, wherein the prediction model is used to predict the binding ability of the first candidate molecule to the target protein. Drugability assessment includes predicting key drugability attributes of the first candidate molecule based on its structural representation information using a water solubility prediction model and a toxicity prediction model, and screening out a second candidate molecule with antibacterial potential based on the predicted structure; the key drugability attributes include solubility, median lethal dose, ADMET parameter and / or antibacterial activity threshold. Antimicrobial resistance risk screening includes constructing a structural database of known antimicrobial drugs, performing molecular similarity calculations between the second candidate molecule and known antimicrobial drugs based on the structural database, and eliminating the second candidate molecule when the similarity calculation results meet a preset threshold condition to obtain a third candidate molecule. Activity verification includes verifying the antibacterial activity of the third candidate molecule to confirm its biological activity against the target.

[0007] Furthermore, the candidate compound library is used to obtain structural representation information of various small molecule compounds, and the structural representation information is represented by molecular structure codes generated by a simplified molecular linear input system.

[0008] Furthermore, the affinity prediction model takes the graph structure representation of the candidate molecule as input, outputs the corresponding molecule-target binding affinity value, and screens the candidate molecules based on a preset affinity threshold, wherein the affinity threshold is pKd ≥ 6.0.

[0009] Furthermore, the molecular docking tool takes the first-level screening results as input, performs molecular docking calculations with the target protein NusG, and uses a molecular docking program that supports parallel acceleration to perform batch docking evaluations on the first-level screening results, outputting docking evaluation results to characterize the binding ability of candidate molecules with the target protein NusG, wherein the binding energy of the candidate molecules is ≤-7.0 kcal / mol.

[0010] Furthermore, the structural database includes structural information of β-lactam, quinolone and / or aminoglycoside antibacterial drugs, and the Tanimoto similarity calculation method is used to calculate the molecular similarity between candidate molecules and known antibacterial drugs, with the preset threshold being ≥0.5.

[0011] A second objective of this invention is to provide an antibacterial drug molecular screening system, comprising: The candidate compound building module is configured to acquire molecular structure information of a variety of small molecule compounds; The first-level screening module is configured to perform first-level screening of the candidate compound library based on an affinity prediction model; The second-level screening module is configured to perform a second-level screening on the results of the first-level screening based on the molecular docking tool. The learning prediction module is configured to input the results of the second-level screening as training data into the machine learning model to construct a molecule-target protein affinity prediction model, wherein the prediction model is used to predict the binding ability of the first candidate molecule to the target protein. The drug development assessment module is configured to predict the key drug development properties of the first candidate molecule based on its structural representation information using a water solubility prediction model and a toxicity prediction model, and to screen out a second candidate molecule with antibacterial potential based on the predicted structure; the key drug development properties include solubility, median lethal dose, ADMET parameter and / or antibacterial activity threshold; The drug resistance risk screening module is configured to construct a structural database of known antimicrobial drugs, and based on the structural database, perform molecular similarity calculations between the second candidate molecule and known antimicrobial drugs. When the similarity calculation results meet a preset threshold condition, the second candidate molecule is eliminated to obtain a third candidate molecule. The activity verification module is configured to verify the antibacterial activity of candidate molecules that meet the screening criteria. The system for screening antimicrobial drug molecules is used to perform the steps in the above-described antimicrobial drug molecule screening method.

[0012] A third objective of the present invention is to provide a computer-readable storage medium storing a program that can be executed by one or more processors to implement the above-described method for screening antimicrobial drug molecules.

[0013] A fourth objective of this invention is to provide the application of the above-mentioned antimicrobial drug molecular screening method, antimicrobial drug molecular screening system, and computer-readable storage medium in screening drugs against Gram-negative bacteria.

[0014] Furthermore, the Gram-negative bacteria include at least one of Escherichia coli, Klebsiella pneumoniae, drug-resistant Klebsiella pneumoniae, acid-producing Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Shigella flexneri group B, and Salmonella typhimurium serotype.

[0015] Furthermore, the selected anti-Gram-negative bacterial drug compounds have any of the following structures: .

[0016] Compared with the prior art, the beneficial effects of the technical solution provided by the present invention are as follows: (1) This invention provides an artificial intelligence-based method for screening antimicrobial drug molecules. This method uses a hierarchical screening system to first quickly screen a library of hundreds of millions of molecules, and then accurately calculates the binding energy, which greatly shortens the screening time, solves the problem of low efficiency of traditional methods, and significantly improves the screening efficiency.

[0017] (2) Integrate the advantages of deep learning and traditional computing tools to construct a molecular-target protein affinity prediction model to predict candidate molecules, reduce unnecessary chemical synthesis and experimental verification steps, and reduce the time and financial costs of drug development.

[0018] (3) Introduce drugability assessment and drug resistance risk screening during the screening process to reduce the probability of obtaining low drugability drug molecules or potential cross-resistant molecules, reduce screening errors, and improve the overall quality of candidate molecules.

[0019] (4) By verifying the antibacterial activity of the screening results, it is ensured that the obtained candidate molecules have actual antibacterial potential, so that the virtual screening results can be more effectively transformed into reliable candidates for subsequent drug development stages.

[0020] (5) The antimicrobial drug molecular screening method provided by the present invention can adjust parameters and replace models according to different target proteins or different types of bacteria, and is applicable to screening scenarios of various antimicrobial drugs. Attached Figure Description

[0021] Figure 1 A schematic flowchart of an artificial intelligence-based antimicrobial drug molecular screening method provided in an embodiment of the present invention; Figures 2a-2b The screening process and validation diagram for compound 7 are shown below. Figure 2a It is a novel antimicrobial molecule discovery process based on deep learning; Figure 2b The t-SNE of all molecules from the training dataset (blue), Approved drugs (green), antibiotics (orange), and Compound7 (yellow, pubchem cid: 57703834) reveals the chemical relationships between these libraries; Figures 2c-2d The similarity between compound 7 (structural illustration) and 2520 approved drugs (DrugBank) is Tanimoto; the most similar Tanimoto to compound 7 is the anticancer drug acalabrutinib (score ~0.23), while thiazole sulfate is the closest antibacterial drug (score ~0.198). Figure 2e This is a framework diagram of the virtual screening model ASBOV of the present invention; Figure 3 To compare the molecular weight (MW) distribution of 47,084 predicted chemical molecules (A) with that of known antibiotics (B); Figure 4 To compare the water solubility of 47,084 predicted chemical molecules (A) with that of known antibiotics (B); Figure 5To compare the median lethal dose (LD50) of 47,084 predicted chemical molecules (A) with that of a known antibiotic (B); Figure 6 The graph shows the results of the in vitro antibacterial activity verification, where the antibacterial activity units are μg / mL and μM, respectively. Figure 7 This is a crystal structure diagram of the NusG protein; Figure 8 To illustrate the core binding sites of the broad-spectrum molecule and the NusG target protein, (B) in the figure is the structural diagram of the NusG protein-compound 7 complex, and (C) is the binding of compound 7 to the NusG protein. Figure 9 The figures show the MD simulation results after mutation of key binding sites. In the figure, (D) represents wild type, (E) represents phenylalanine at position 64 mutated to alanine, and (F) represents phenylalanine at position 65 mutated to alanine. Figure 10 The figure shows the SPR detection results of compound 7 and NusG protein. In the figure, (G) is the SPR sensor response map and (H) is the affinity fitting curve. Figure 11 The results of MST assay for compound 7 and NusG protein; Figure 12 The chemical structure diagrams of small molecule compounds of Formulas 1 to 9 clearly show the molecular structural features of each compound. Figure 13 The synthetic route for compound 7 is shown below. Figure 14 A schematic diagram of the system structure provided by the present invention is shown. Figure 15 A block diagram of an electronic device suitable for implementing an information acquisition method according to an embodiment of the present invention is shown schematically. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the specific embodiments and accompanying drawings are described in further detail below. Where specific techniques or conditions are not specified in the embodiments, they are performed in accordance with the techniques or conditions described in the literature in this field or according to the product manual.

[0023] Explanation of English abbreviations in this invention: The protein NusG (N-utilization substance G) is a uniquely conserved transcription factor found only in all life forms (bacteria, archaea, and eukaryotes). In bacteria, it is primarily responsible for regulating transcriptional elongation, termination, and the coupling of transcription and translation.

[0024] SMILES (Simplified Molecular Input Line Entry System) is a linear notation system used to describe the chemical structures of small molecules.

[0025] LGBM is an efficient distributed gradient boosting (GBDT) framework based on decision tree algorithms.

[0026] refer to Figure 1 This invention provides a flowchart of an artificial intelligence-based method for screening antimicrobial drug molecules. First, a rapid affinity prediction model is used to initially screen a library of hundreds of millions of molecules. Then, traditional molecular docking tools are used to accurately calculate binding energies. Simultaneously, a multimodal druggability prediction model is developed and combined with a known antimicrobial drug structure exclusion procedure to screen candidate molecules. Using NusG, a conserved essential protein of Gram-negative bacteria, as a target, screening was conducted. Several novel lead molecules were identified through chemical synthesis, in vitro activity verification, molecular docking, surface plasmon resonance, and micro-thermophoresis. The specific steps include: A hierarchical virtual screening system is constructed, including a first-level screening of a candidate compound library based on an affinity prediction model; a second-level screening of the results of the first-level screening based on a molecular docking tool; and inputting the results of the second-level screening as training data into a machine learning model to construct a molecule-target protein affinity prediction model, wherein the prediction model is used to predict the binding ability of the first candidate molecule to the target protein. Drugability assessment includes predicting key drugability attributes of the first candidate molecule based on its structural representation information using a water solubility prediction model and a toxicity prediction model, and screening out a second candidate molecule with antibacterial potential based on the predicted structure; the key drugability attributes include solubility, median lethal dose, ADMET parameter and / or antibacterial activity threshold. Antimicrobial resistance risk screening includes constructing a structural database of known antimicrobial drugs, performing molecular similarity calculations between the second candidate molecule and known antimicrobial drugs based on the structural database, and eliminating the second candidate molecule when the similarity calculation results meet a preset threshold condition to obtain a third candidate molecule. Activity verification includes verifying the antibacterial activity of the third candidate molecule to confirm its biological activity against the target.

[0027] In some implementations, the candidate compound library is used to obtain structural representation information of a variety of small molecule compounds, which is represented by molecular structure codes generated by a simplified molecular linear input system. The candidate compound library can be the PubChem library, which contains 116M compounds.

[0028] In some implementations, the affinity prediction model takes the graph structure representation of the candidate molecule as input, outputs the corresponding molecule-target binding affinity value, and screens the candidate molecules based on a preset affinity threshold, wherein the affinity threshold is pKd ≥ 6.0; the selectable affinity prediction model is KarmaDock.

[0029] In some implementations, the molecular docking tool takes the first-level screening results as input, performs molecular docking calculations with the target protein NusG, and uses a parallel-accelerated molecular docking program to perform batch docking evaluations on the first-level screening results. It outputs docking evaluation results characterizing the binding ability of candidate molecules to the target protein NusG, where the binding energy of the candidate molecules is ≤-7.0 kcal / mol. The molecular docking tool can be a traditional Vina-GPU molecular docking tool with a physical scoring function.

[0030] More specifically, by combining the bacterial NusG protein and using the ASBOV-DTA model, the binding energy of the potential molecules obtained from the initial screening was accurately calculated, and molecules with a binding energy ≤ -7.0 kcal / mol were screened out.

[0031] It should be noted that this invention is not limited to the use of a specific open-source model; any model capable of achieving the same technical functionality can be used in this invention.

[0032] The invention has now been generally described, and will be more readily understood by referring to the following embodiments, which are provided by way of example and not by way of limitation.

[0033] Example 1 This embodiment provides an artificial intelligence-based method for screening antimicrobial drug molecules, used to efficiently screen small molecule compounds with potential antimicrobial activity from a library of hundreds of millions of molecules. The screening process is as follows: Figure 2a As shown, the specific steps include: (1) Construction and model training of hierarchical virtual screening system (1.1) Construction of the initial molecular library: SMILES structural information of approximately 116 million small molecule compounds was obtained from the PubChem database, and a molecular structure database or training dataset for an artificial intelligence model was constructed based on this SMILES structural information; such as... Figure 2b As shown, t-SNEs of all molecules from the training dataset (blue), Approved drugs (green), antibiotics (orange), and Compound7 (yellow, pubchem cid: 57703834) reveal the chemical relationships between these libraries; (1.2) Structure preprocessing: The SMILES structure information of the small molecule compounds obtained above was normalized using the RDKit open-source cheminformatics tool. The specific process is as follows: Input: Original structural information of the small molecule compounds to be processed; Processing: The original SMILES structural information was subjected to a series of standardization processes: ① Desalting: Inorganic ions (such as sodium, chloride, and potassium ions) that are not covalently bound to the target active molecule were identified and removed from the molecular structure, retaining only the core framework structure of the target organic molecule; ② Solvent removal: Residual solvent molecules (such as methanol, ethanol, water, and dichloromethane) in the molecular structure were screened and removed to avoid interference from solvent components in subsequent molecular property calculations; ③ Aromaticity determination: Based on Hückel's rule (cyclic, planar, conjugated, π electron number 4n+2), the aromaticity determination module of RDKit was used to identify the cyclic structures in the molecule one by one, marking the aromatic and non-aromatic ring regions to clarify the characteristics of the molecular conjugated system; ④ Stereochemical normalization: The chiral center configuration (R / S configuration) and double bond cis / trans configuration (E / Z configuration) in the molecule were standardized and normalized, and the stereochemical labels were calibrated using the IUPAC standard nomenclature rules to ensure that different drawing forms of the same molecule have a unique stereochemical expression; Output: After normalization and conversion, there are no redundant ions or solvent residues, the aromaticity is clearly marked, and the standard SMILES structural information with consistent stereochemical expression is provided; Applications: The above-mentioned standard SMILES structural information will be used to construct a molecular structure database of small molecule compounds, and at the same time serve as a basic dataset for training artificial intelligence models, providing a consistent structural basis for subsequent analysis processes such as affinity calculations, physicochemical characterization and prediction.

[0034] (1.3) Initial Affinity Screening: The standardized SMILES structural information is converted into a molecular graph structure and input into the deep learning model KarmaDock to predict the binding affinity between the molecule and the target protein. The processing logic is as follows: Standardized SMILES structural information input - conversion into a molecular graph structure - input into the KarmaDock graph neural network model - protein pocket feature extraction and interaction modeling. The KarmaDock model structure and processing procedure are as follows: The KarmaDock model employs a three-stage network structure: "molecular graph feature extraction - protein pocket feature interaction - affinity regression prediction," specifically including: ① Input layer: Receives the three-dimensional structural vector of the target protein pocket and the graph vector information of the molecule to be predicted, respectively; ② Molecular graph feature extraction module: A multi-layer graph convolutional neural network (GCN) is used to initialize node features through attributes such as atom type, chemical bond type, and atomic charge. The adjacency matrix is ​​used to aggregate the feature information of adjacent atoms to generate a high-dimensional feature vector that can characterize the global chemical environment and topological structure of the molecule. ③ Protein-molecule interaction module: Introducing an attention mechanism to calculate the interaction weights between molecular feature vectors and key amino acid residue feature vectors in the protein pocket, focusing on the feature information of key binding sites that may form hydrogen bonds, hydrophobic interactions, π-π stacking, etc., and enhancing the expression of specific binding features; ④ Regression prediction layer: The fused features after interaction are input into the fully connected neural network, and after being mapped by the activation function, the final binding affinity prediction value (pKd) is output.

[0035] The specific prediction process is as follows: First, the preprocessed molecules are converted into a graph structure representation (atoms as nodes, chemical bonds as edges with bond types labeled), and the pocket structure data of the target protein is input. Then, feature learning and interaction modeling are completed through the molecular graph feature extraction module and protein-molecule interaction module of the KarmaDock model. Finally, the regression prediction layer outputs the binding affinity value (pKd) corresponding to each molecule. The affinity threshold is set to pKd ≥ 6.0, and approximately 22.9 million potential molecules are obtained.

[0036] (1.4) Precise docking calculation: The molecules obtained from the initial screening are docked with the target protein NusG, and the ASBOV-DTA (structure- or depth-based drug-target affinity prediction) model is trained based on the docking results. The model structure is as follows: Figure 2e As shown, the details are as follows: ① Target protein and docking parameter preparation: NusG is a conserved essential transcriptional regulator of Gram-negative bacteria. Its three-dimensional structure was obtained from the Alphafold2 database. After PyMOL pretreatment (removing redundant water molecules, adding amino acid side chains and minimizing energy), a cube with a side length of 40 Å centered on the active site was set as the docking region. ② Vina-GPU Batch Docking: Optimize the preferred conformation of potential molecules using RDKit and output it in PDBQT format. Simultaneously, preprocess the receptor protein NusG by removing water of crystallization, removing impurities, performing charge distribution, and defining the docking pocket of the active site. Then, using NusG as the receptor and the optimized molecules as ligands, Vina-GPU is used to calculate the molecule-protein free energy in parallel. Combining the free energy values ​​and the interaction characteristics between the molecule and the protein active site, the optimal docking conformation of each molecule is screened. Finally, a docking result dataset containing molecule identifiers, optimal conformation coordinates, binding free energy, and interaction details is output.

[0037] ③ASBOV-DTA Model Structure: This is a gradient boosting tree model based on LGBM, consisting of a feature input layer, an ensemble decision tree layer, and an output layer. The feature input layer includes docking-derived features and molecular physicochemical descriptors. Specific physicochemical parameters include molecular molar mass, lipid-water partition coefficient (LogP), topological polar surface area (TPSA), number of hydrogen bond donors (HBD), number of hydrogen bond acceptors (HBA), molecular refractive index, polarizability, hydrophobic surface area ratio, and atomic charge distribution characteristics. It also incorporates docking-derived binding free energy, number of hydrogen bond interactions, hydrophobic contact area, and steric hindrance parameters. The ensemble decision tree layer uses multiple CART decision trees for gradient boosting training, introducing L1 regularization (Lasso) and L2 regularization (Ridge) to suppress overfitting. Its logistic transparency is specifically reflected in the fact that each CART tree's splitting node is based on a single feature threshold (e.g., LogP ≥ 2.5, TPSA ≤ 120). The tree's hierarchical structure, feature splitting paths, and the importance weights of each feature are clearly defined (Ų), and the decision logic and feature contribution of each decision node during the model prediction process are traceable and interpretable, without relying on black-box computation, ensuring the logical verifiability of the model prediction results; the output layer is used to output the molecular-protein free energy prediction results.

[0038] ④ Model Training: Select the 1 million + 2200 docking results mentioned in step ② to construct a dataset (8:2 split training / test set). After feature standardization and key feature screening, use 5-fold cross-validation for training. Optimize hyperparameters through grid search. Evaluate model performance with RMSE, MAE, and R² to ensure prediction accuracy and generalization ability.

[0039] (1.5) Constructing the machine learning model ASBOV-DTA: Using the binding free energy output by Vina-GPU batch docking as the prediction label (denoted as y), a feature matrix is ​​constructed by combining molecular physicochemical descriptors and protein-ligand interaction features. (n is the number of samples, m is the feature dimension), an LGBM model based on a gradient boosting tree framework was trained to quantitatively predict the binding ability of molecules to the target protein NusG. The model achieved a Pearson correlation coefficient R≈0.66 on the independent test set, indicating good predictive accuracy. Specifically, its structure consists of three key parts: an input layer, a core ensemble tree layer, and an output layer. The input layer receives a standardized preprocessed fusion feature vector, which is composed of molecular descriptors (such as topological, electrical, and hydrophobic features) and protein-ligand interaction features (such as the number of hydrogen bonds, hydrophobic interaction area, and van der Waals interaction energy), effectively integrating key binding information between the molecule and the target protein.

[0040] (2) Drugability assessment and drug resistance risk screening (2.1) Constructing a multimodal drug-likeness prediction model: Based on the SMILES structure of the molecule, the following sub-models are constructed respectively: a. Water solubility prediction model: The water solubility of the compound is predicted by the WatLoS model, and the result is taken as LogS≤-3.0 (based on the average predicted value of LogS of marketed antibiotics -2.97, it is set to -3.0). WatLoS is a regression model based on the Uni-Mol large model. Uni-Mol is a 3D structure-aware pre-trained model. The core architecture is based on the SE(3) equivariant Transformer. The overall structure consists of an input layer, a core Transformer encoding layer, a pre-trained task head and a downstream task adaptation layer. Each module works together to achieve end-to-end learning from molecular 3D structure to feature representation. The input layer is responsible for digitizing the molecular physical structure. On the one hand, it converts chemical properties such as atomic number, hybridization mode and charge state into high-dimensional atomic feature vectors through the embedding layer. On the other hand, it constructs paired representations based on the invariant spatial position encoding calculated by the three-dimensional coordinates of the atoms. These two representations remain stable under global rotation and translation, laying the foundation for subsequent spatial information learning. Its prediction process follows the logic of "input preprocessing - feature encoding - representation adaptation - result output": First, the input molecule is preprocessed to determine the atom type and three-dimensional coordinates and complete the standardization; then, the preprocessed atomic features and spatial position codes are input into the core Transformer encoding layer, and a high-dimensional representation containing the global structure and local interactions of the molecule is generated through multi-layer self-attention calculation and feature fusion; then, this representation is passed into the downstream task adaptation layer, and undergoes targeted feature transformation and mapping; finally, the prediction result is output through the regression head (suitable for quantitative properties such as water solubility) or the classification head (suitable for qualitative judgment tasks). The whole process fully relies on 3D structure perception capabilities to ensure the accuracy and generalization of the prediction.

[0041] b. Toxicity prediction model: Predict the median lethal dose (LD50) of compounds, screened based on the prediction threshold of 3.36 for the top 95% of marketed antibiotics; (2.2) The model input features include: molecular descriptors (such as molecular weight, LogP, TPSA, number of hydrogen bond donors / acceptors), molecular fingerprints (Morgan fingerprint, MACCS bond) and three-dimensional conformation features (such as USRCAT, 3D Overlap). All sub-models are constructed using the LGBM algorithm and deep learning. (2.3) Optionally, predict the potential antibacterial activity of the compound and screen out molecules with good antibacterial potential; (2.4) Construct a database of known antimicrobial drug structures, covering representative drugs such as β-lactams, quinolones, and aminoglycosides. Using the Tanimoto similarity calculation method, a similarity threshold ≥0.5 is set to automatically remove molecules with structures highly similar to known antimicrobial drugs, thus avoiding cross-resistance. Figure 2c As shown in the figure, the cumulative probability distribution of Tanimoto similarity of the molecular dataset is displayed. The extremely steep curve shows that the similarity of the vast majority of molecules is less than 0.2, and the proportion of molecules with similarity greater than or equal to 0.3 is 0%. This strongly proves that the molecular library has extremely high structural diversity and novelty. (2.5) After the above screening, approximately nine candidate molecules with novel structures, good druggability, and low risk of drug resistance were retained for subsequent experimental verification. The structural formulas of these nine molecules are referenced. Figure 12 As shown.

[0042] Target candidate molecules (e.g.) Figure 2d The novelty of the structure shown in the figure was further confirmed by a comparative analysis with 2,520 approved drugs. Experimental results showed that the average similarity between the target molecule and the existing drug library was only 0.083, with a median of 0.082. Even compared with the top 10 drugs with the closest structures, the peak similarity was only about 0.2, far below the threshold (0.7-0.8) typically considered to indicate structural similarity.

[0043] refer to Figure 3 By comparing the molecular weight (MW) distribution of the predicted 47,084 chemical molecules (A) with that of known antibiotics (B), significant differences in their physicochemical properties were revealed: the compounds in the 477,084 compounds were generally heavier, with their molecular weights mainly concentrated between 700 and 800; while antibiotics were lighter and more streamlined, with more than half of the antibiotics having molecular weights between 300 and 500.

[0044] refer to Figure 4 The difference in water solubility distribution between 47,084 predicted chemical molecules (A) and known antibiotics (B) revealed that antibiotics exhibit significantly higher water solubility. Molecules in (A) are extremely hydrophobic, with the vast majority (approximately 90.8%) having extremely low logS values ​​between -9.7 and -8.0 (median -8.87). In contrast, antibiotics show a wider and significantly higher water solubility distribution, with a median of -2.90, and 26.7% of molecules falling within the range of -3.0 to -2.0. This difference suggests that drugs effectively combating pathogens often require better solubility to adapt to the complex biological fluid environment compared to conventional small-molecule screening libraries.

[0045] refer to Figure 5 The differences in the distribution of median lethal dose (LD50) of 47,084 predicted chemical molecules (A) and known antibiotics (B) reflect the uniqueness of antibiotics in terms of safety or toxicity profiles. The toxicity values ​​of compounds in (A) are highly concentrated, with approximately 82.4% of the molecules falling between 2.5 and 3.1, and a median of 2.81. In contrast, the distribution of antibiotics is broader and the overall values ​​are lower (meaning they generally have better safety or different toxicity characteristics), with the distribution center shifting to the left, the median dropping to 2.43, and 27.2% of antibiotics falling in the low value range of 1.6–2.0.

[0046] Through comprehensive analysis of three sets of data—molecular weight, water solubility, and toxicity (LD50)—antibiotics exhibit characteristics of being smaller, more soluble, and having higher biosafety: In terms of molecular weight, the median of antibiotics (435.9) is much lower than that of the universal library (712.9), with most concentrated in the simplified range of 300-500; in terms of solubility, antibiotics show a strong hydrophilic advantage, with the median (logS = -2.90) significantly higher than that of the universal library (-8.87), reflecting their need to adapt to the biological fluid environment for survival; and in terms of toxicological indicators, the LD50 distribution of antibiotics is more biased towards the lower value range (median 2.43 vs 2.81). This physicochemical combination of "low molecular weight, high solubility, and specific toxicity range" constitutes the core identification feature that distinguishes antibiotics from ordinary synthetic small molecules, and follows more stringent biosafety screening standards.

[0047] (3) Chemical synthesis and in vitro antibacterial activity verification The chemical synthesis and antibacterial activity verification of nine candidate molecules were carried out, specifically including the following steps: (3.1) Chemical synthesis: A modular synthesis strategy was adopted, giving priority to routes with readily available starting materials, simple synthesis steps and high yield. All target compounds were purified by HPLC with a purity ≥95%, and their structures were confirmed by 1H-NMR, LCMS, MS Spectrum and HPLC. (3.2) Antibacterial activity assay: The half-maximal inhibitory concentration (IC50) of the candidate molecules against the following eight Gram-negative bacteria and one Gram-positive bacteria was determined using the micro-broth dilution method (referring to CLSI standards): *Escherichia coli* (ATCC25922); *Klebsiella pneumoniae* (ATCC13883); drug-resistant *Klebsiella pneumoniae* (BAA-2146); *Pseudomonas aeruginosa* (ATCC27853); *Acinetobacter baumannii* (ATCC19606); *Salmonella typhimurium* (ATCC14028); *Shigella flexneri* (group B); *Klebsiella pneumoniae* (ATCC13182); and *Clostridium difficile* (ATCC43255). Specific results are as follows: Figure 6 As shown.

[0048] (3.3) The activity units are expressed in μg / mL and μM. The positive control drugs are ampicillin and halicin. (3.4) Screening criteria: Compounds with an IC50 ≤ 16 μg / mL against at least 5 bacterial strains are defined as "broad-spectrum antibacterial molecules"; (3.5) Finally, a lead compound (compound 7) with broad-spectrum antibacterial activity was obtained, which can effectively inhibit 9 types of bacteria. Its chemical structure is as follows: Figure 12 As shown in Formula 7. To clarify its potential against drug-resistant bacteria, experiments were conducted against a drug-resistant strain of Klebsiella pneumoniae (BAA-2146). The results showed that compound 7 had a significant inhibitory potential against this drug-resistant strain, with an IC50 value of 16 μg / mL.

[0049] This embodiment also provides the synthetic route for compound 7, see reference. Figure 13 Starting with 4-cyanophenyl bromide ketone, the target molecule compound 7 was synthesized through a nine-step chemical reaction. SPR experiments confirmed that the molecule has an affinity of about 9 μM for NusG. Molecular dynamics (MD) simulations were also used to explain the irreplaceable role of the key amino acid sites F64 / F65 in maintaining this binding stability at the atomic scale.

[0050] Example 2 Molecular mechanism verification.

[0051] This embodiment verifies the mechanism of action of the obtained broad-spectrum antibacterial molecule, specifically including the following steps: (1) Molecular docking analysis: Using AutoDock Vina software, compound 7 was molecularly docked with NusG protein to identify its binding sites and key interactions. The docking results showed that these molecules mainly bind to key amino acid residues of NusG, including LEU43, PRO113, PRO115, HIS117, HIS118, ARG122, and LEU125. The main interaction types included hydrogen bonds, van der Waals forces, hydrophobic interactions, π-Sigma bonds, and π-π stacking. The crystal structure of NusG protein is shown in Figure 1. Figure 7 As shown, the structure of the NusG protein-compound 7 complex is as follows: Figure 8 As shown in (B), the binding of compound 7 to the NusG protein is as follows. Figure 8 As shown in (C); Figure 9 This study demonstrates, through molecular dynamics simulations, the crucial role of key amino acid sites (F64 and F65) in maintaining the stability of the complex binding of compound CPD7 with protein NusG. Mutations in F64 and F65 cause NusG to change from a “compact stable state” to a “loose unstable state”.

[0052] (2) Surface plasmon resonance (SPR) experiment: Using the Biacore 8K system, NusG protein was immobilized on the surface of the CM5 sensor chip, and small molecule solutions of different concentrations were injected. The binding and dissociation curves were measured, and the binding kinetic parameters were calculated. K D (value), the result is as follows Figure 10 As shown in the figure, (G) is the SPR sensor response map, and (H) is the affinity fitting curve. Compound 7 can stably bind to the NusG protein, and their interaction conforms to a typical 1:1 binding mode. K D The value was 9.07±0.80μM, indicating good binding ability; (3) Micro-thermophoresis (MST) experiment: The NusG protein was fluorescently labeled using the NanoTemper Monolith X system. Gradual dilutions of small molecule solutions were added, and the changes in thermophoretic signals were measured. The results are as follows: Figure 11 As shown, compound CPD7 can bind to NusG in a concentration-dependent manner, with an affinity at the micromolar level, consistent with the SPR experiment, further verifying the specific binding ability between the molecule and the target protein. (4) Based on the above experimental results, it is confirmed that the broad-spectrum antibacterial molecular compound 7 can exert its antibacterial effect by specifically binding to NusG protein and inhibiting its transcriptional regulatory function.

[0053] Example 3 Structural novelty analysis and model performance evaluation.

[0054] This embodiment systematically evaluates the structural novelty and model performance of the screened molecules: (1) Structural novelty analysis: Taking compound 7 as an example, the Tanimoto similarity calculation method was used to compare its structure with all approved drugs in the DrugBank database. The results showed that its maximum similarity with existing drugs was 0.23 (similar to the anticancer drug acalabrutinib), and its highest similarity with existing antibacterial drugs was 0.198 (similar to thiazole sulfate), indicating that its structure has significant novelty; (2) Evaluation of model prediction performance: The Pearson correlation coefficient R between the predicted binding energy and the experimental value of the ASBOV-DTA model on the test set was 0.66, indicating that the model has good predictive ability; (3) Screening efficiency evaluation: Nine lead molecules with broad-spectrum antibacterial activity were screened from 116 million molecules. The entire screening process took about 6 weeks, which is significantly shorter than traditional high-throughput screening methods (which usually take 6 to 12 months). (4) Experimental hit rate: In the in vitro inhibition of 9 strains of Escherichia coli MC4100 and even drug-resistant Klebsiella pneumoniae (BAA-2146), all 9 molecules showed clear antibacterial activity, with an experimental hit rate of 100%, which is much higher than the traditional virtual screening method (usually <10%).

[0055] Example 4 This embodiment, based on the design of Embodiment 1, discloses an antibacterial drug molecular screening system, such as... Figure 14 As shown, the system includes: The candidate compound building module is configured to acquire molecular structure information of a variety of small molecule compounds; The first-level screening module is configured to perform first-level screening of the candidate compound library based on an affinity prediction model; The second-level screening module is configured to perform a second-level screening on the results of the first-level screening based on the molecular docking tool. The learning prediction module is configured to input the results of the second-level screening as training data into the machine learning model to construct a molecule-target protein affinity prediction model, wherein the prediction model is used to predict the binding ability of the first candidate molecule to the target protein. The drug development assessment module is configured to predict the key drug development properties of the first candidate molecule based on its structural representation information using a water solubility prediction model and a toxicity prediction model, and to screen out a second candidate molecule with antibacterial potential based on the predicted structure; the key drug development properties include solubility, median lethal dose, ADMET parameter and / or antibacterial activity threshold; The drug resistance risk screening module is configured to construct a structural database of known antimicrobial drugs, and based on the structural database, perform molecular similarity calculations between the second candidate molecule and known antimicrobial drugs. When the similarity calculation results meet a preset threshold condition, the second candidate molecule is eliminated to obtain a third candidate molecule. The activity verification module is configured to verify the antibacterial activity of candidate molecules that meet the screening criteria. This antimicrobial drug molecular screening system is used to perform the steps in the antimicrobial drug molecular screening method provided by the present invention.

[0056] Example 5 Based on the design of Embodiments 1 and 4, this embodiment discloses a computer-readable storage medium storing a program that can be executed by one or more processors to implement the antibacterial drug molecule screening method provided by the present invention.

[0057] Figure 15 A block diagram schematically illustrates an electronic device suitable for implementing an information acquisition method according to an embodiment of the present disclosure.

[0058] like Figure 15 As shown, an electronic device 100 according to an embodiment of the present disclosure includes a processor 101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 102 or a program loaded from a storage portion 108 into a random access memory (RAM) 103. The processor 101 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 101 may also include onboard memory for caching purposes. The processor 101 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0059] The RAM 103 stores various programs and data required for the operation of the electronic device 100. The processor 101, ROM 102, and RAM 103 are interconnected via a bus 104. The processor 101 executes various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 102 and / or RAM 103. It should be noted that programs may also be stored in one or more memories other than ROM 102 and RAM 103. The processor 101 may also execute various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.

[0060] According to embodiments of this disclosure, the electronic device 100 may further include an input / output (I / O) interface 105, which is also connected to a bus 104. The electronic device 100 may also include one or more of the following components connected to the I / O interface 105: an input section 106 including a keyboard, mouse, etc.; an output section 107 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 108 including a hard disk, etc.; and a communication section 109 including a network interface card such as a LAN card, modem, etc. The communication section 109 performs communication processing via a network such as the Internet. A drive 110 is also connected to the I / O interface 105 as needed. A removable medium 111, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 110 as needed so that computer programs read from it can be installed into the storage section 108 as needed.

[0061] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0062] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 102 and / or RAM 103 and / or one or more memories other than ROM 102 and RAM 103 described above.

[0063] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the information acquisition method provided in the embodiments of this disclosure.

[0064] When the computer program is executed by the processor 101, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0065] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via communication section 109, and / or installed from removable medium 111. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0066] In such an embodiment, the computer program can be downloaded and installed from a network via communication section 109, and / or installed from removable medium 111. When the computer program is executed by processor 101, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0067] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0069] Where there is no conflict, the above embodiments and features described herein can be combined with each other.

[0070] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for screening antimicrobial drug molecules based on artificial intelligence, characterized in that, Includes the following steps: A hierarchical virtual screening system was constructed, including a first-level screening of the candidate compound library based on an affinity prediction model; A second-level screening is performed based on the results of the first-level screening using molecular docking tools; The results of the second-level screening are used as training data to input into the machine learning model to construct a molecule-target protein affinity prediction model. The prediction model is used to predict the binding ability of the first candidate molecule to the target protein. Drugability assessment includes predicting key drugability attributes of the first candidate molecule based on its structural representation information using a water solubility prediction model and a toxicity prediction model, and screening out a second candidate molecule with antibacterial potential based on the predicted structure; the key drugability attributes include solubility, median lethal dose, ADMET parameter and / or antibacterial activity threshold. Antimicrobial resistance risk screening includes constructing a structural database of known antimicrobial drugs, performing molecular similarity calculations between the second candidate molecule and known antimicrobial drugs based on the structural database, and eliminating the second candidate molecule when the similarity calculation results meet a preset threshold condition to obtain a third candidate molecule. Activity verification includes verifying the antibacterial activity of the third candidate molecule to confirm its biological activity against the target.

2. The method for molecular screening of antibacterial drugs according to claim 1, characterized in that, The candidate compound library is used to obtain structural representation information of various small molecule compounds, and the structural representation information is represented by molecular structure codes generated by a simplified molecular linear input system.

3. The method for molecular screening of antibacterial drugs according to claim 1, characterized in that, The affinity prediction model takes the graph structure representation of the candidate molecule as input, outputs the corresponding molecule-target binding affinity value, and screens the candidate molecules based on a preset affinity threshold, wherein the affinity threshold is pKd ≥ 6.

0.

4. The method for screening antimicrobial drug molecules according to claim 2 or 3, characterized in that, The molecular docking tool takes the first-level screening results as input, performs molecular docking calculations with the target protein NusG, and uses a molecular docking program that supports parallel acceleration to perform batch docking evaluations on the first-level screening results. It outputs docking evaluation results to characterize the binding ability of candidate molecules to the target protein NusG, wherein the binding energy of the candidate molecules is ≤-7.0 kcal / mol.

5. The method for screening antimicrobial drug molecules according to claim 4, characterized in that, The Tanimoto similarity calculation method is used to calculate the molecular similarity between candidate molecules and known antibacterial drugs, with a preset threshold ≥ 0.

5.

6. An antibacterial drug molecular screening system, characterized in that, include: The candidate compound building module is configured to acquire molecular structure information of a variety of small molecule compounds; The first-level screening module is configured to perform first-level screening of the candidate compound library based on an affinity prediction model; The second-level screening module is configured to perform a second-level screening on the results of the first-level screening based on the molecular docking tool. The learning prediction module is configured to input the results of the second-level screening as training data into the machine learning model to construct a molecule-target protein affinity prediction model, wherein the prediction model is used to predict the binding ability of the first candidate molecule to the target protein. The drug development assessment module is configured to predict the key drug development properties of the first candidate molecule based on its structural representation information using a water solubility prediction model and a toxicity prediction model, and to screen out a second candidate molecule with antibacterial potential based on the predicted structure; the key drug development properties include solubility, median lethal dose, ADMET parameter and / or antibacterial activity threshold; The drug resistance risk screening module is configured to construct a structural database of known antimicrobial drugs, and based on the structural database, perform molecular similarity calculations between the second candidate molecule and known antimicrobial drugs. When the similarity calculation results meet a preset threshold condition, the second candidate molecule is eliminated to obtain a third candidate molecule. The activity verification module is configured to verify the antibacterial activity of candidate molecules that meet the screening criteria. The antimicrobial drug molecular screening system is used to perform the steps in the antimicrobial drug molecular screening method according to any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that can be executed by one or more processors to implement the antimicrobial drug molecule screening method according to any one of claims 1-5.

8. The application of the antimicrobial drug molecular screening method according to any one of claims 1-5, the antimicrobial drug molecular screening system according to claim 6, and the computer-readable storage medium according to claim 7 in screening drugs against Gram-negative bacteria.

9. The application according to claim 8, characterized in that, The Gram-negative bacteria include at least one of Escherichia coli, Klebsiella pneumoniae, drug-resistant Klebsiella pneumoniae, acid-producing Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Shigella flexneri group B, and Salmonella typhimurium serotype.

10. The application according to claim 8, characterized in that, The selected anti-Gram-negative bacterial drug compounds have any of the following structures: 。