A method, system, apparatus, and storage medium for computational design of broad-spectrum antibacterial compounds

By performing multi-step computational design on the compound library, including screening, skeleton deduplication, derivatization, and screening, the problem of cross-species activity design in the discovery of antimicrobial drugs in the prior art has been solved, and efficient and reproducible broad-spectrum antimicrobial compound design has been achieved.

CN120977436BActive Publication Date: 2026-04-21CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2025-09-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing computer-aided drug design methods face difficulties in the systematic transformation from narrow-spectrum to broad-spectrum antimicrobial drug discovery, lack model robustness, struggle with cross-species knowledge transfer, and lack of streamlined generation and derivation processes, leading to structural searches deviating from the membrane-penetrating feature space.

Method used

A first predictive model trained on Gram-positive bacteria bioactivity data was used for virtual screening to identify and remove duplicate chemical skeletons, select high-potential skeletons, generate a derivative library through virtual chemical derivation, and then screen the derivatives using a second predictive model trained on Gram-negative bacteria bioactivity data to form broad-spectrum antibacterial compounds.

Benefits of technology

This system enables the transformation from narrow-spectrum compounds to broad-spectrum antibacterial compounds, improving screening efficiency and accuracy, saving R&D resources, providing a scientific data-driven design path, and breaking through the throughput, time, and cost limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977436B_ABST
    Figure CN120977436B_ABST
Patent Text Reader

Abstract

This invention discloses a computational design method, system, device, and storage medium for broad-spectrum antibacterial compounds, relating to the fields of computational chemistry and drug discovery. The method includes: virtual screening of an initial compound library to identify a preliminary set of active compounds; extracting pre-defined chemical skeletons from the preliminary set of active compounds to identify a group of dominant chemical skeletons; filtering the identified dominant chemical skeletons based on pre-defined medicinal chemistry indicators to select a group of high-potential skeletons; selecting one or more representative molecules from the original cluster members corresponding to the selected high-potential skeletons, and generating a derivative library through virtual chemical derivatization; and virtual screening of the derivative library to identify a final set of candidate compounds. This invention can transform narrow-spectrum compounds that exhibit activity only against Gram-positive bacteria into broad-spectrum antibacterial compounds that are also active against Gram-negative bacteria through systematic computational design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computational chemistry and drug discovery, and in particular to a computational design method, system, device, and storage medium for a broad-spectrum antibacterial compound. Background Technology

[0002] With the widespread use of antimicrobial drugs, bacterial resistance has become a global public health challenge, particularly among Gram-negative bacteria. Gram-negative bacteria have a double-membrane structure consisting of a plasma membrane and an outer membrane. The outer membrane is rich in negatively charged lipopolysaccharide (LPS), forming a selective permeation barrier for small molecules. To effectively target Gram-negative bacteria, candidate molecules must cross the outer membrane and maintain a sufficient concentration within the bacterial cell. "Amphiphilic design" (achieving a balance between hydrophobic fragments and hydrophilic / ionizable fragments) is one of the core strategies.

[0003] In recent years, with the development of computing technology, computer-aided drug design (CADD) and artificial intelligence drug discovery (AIDD) have provided new avenues for antimicrobial drug discovery. Existing CADD / AIDD workflows mainly include target structure-based docking / virtual screening, ligand-based modeling, molecular generation and structure optimization, and ADMET / permeability prediction. While each has its value, they share common limitations in the system transformation from "narrow spectrum to broad spectrum": models are mostly trained with single-species data, resulting in insufficient cross-species knowledge transfer; the generation / derivation stage does not streamline key physicochemical targets such as "amphiphilicity" upfront, making structure searches prone to deviating from the "membrane-penetrating" feature space; decision-making relies on single models, lacking robustness; and the workflow is fragmented and difficult to standardize and reuse.

[0004] Therefore, proposing a computational design method, system, device, and storage medium for broad-spectrum antibacterial compounds to address the problems existing in the prior art is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a computational design method, system, device and storage medium for broad-spectrum antibacterial compounds, which can transform narrow-spectrum compounds that only show activity against Gram-positive bacteria into broad-spectrum antibacterial compounds that are also active against Gram-negative bacteria through systematic computational design.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A computational design method for broad-spectrum antibacterial compounds, comprising:

[0008] S1. Apply the first prediction model trained based on Gram-positive bacteria bioactivity data to virtually screen the initial compound library and identify a preliminary set of active compounds.

[0009] S2. Extract the preset chemical skeleton of each compound in the preliminarily identified set of active compounds to obtain an initial list containing duplicate skeletons. Perform a deduplication operation on the initial list containing duplicate skeletons to identify a set of dominant chemical skeletons.

[0010] S3. Based on preset medicinal chemistry indicators, the identified advantageous chemical skeletons are filtered, and a set of high-potential skeletons are selected according to the drug-likeness, synthetic feasibility and chemical stability of the skeletons.

[0011] S4. From the original cluster molecular members corresponding to the selected high-potential skeleton, select one or more representative molecules and generate a derivative library through virtual chemical derivatization.

[0012] S5. Apply the second prediction model trained based on Gram-negative bacterial bioactivity data to virtually screen the derivative library and identify a final set of candidate compounds.

[0013] Optionally, in the above method, in S1, the first prediction model includes at least two models based on different algorithm principles. These at least two models are used to screen the activity of compounds in the initial compound library to obtain a preliminary set of active compounds. The compounds included in this preliminary set of active compounds satisfy the preset activity threshold requirements of the prediction results of at least two models. Optionally, the Gram-positive bacteria can be Staphylococcus aureus.

[0014] In the above method, optionally, in S2, the chemical skeleton of the compound is a Bemis–Murcko skeleton.

[0015] In the above method, optionally, in S3, the preset medicinal chemical indicators include:

[0016] Physicochemical limitations: molecular weight of the backbone, lipophilicity, number of hydrogen bond donors and hydrogen bond acceptors;

[0017] Structural complexity constraints: number of rotatable keys and number of ring systems in the skeleton;

[0018] Limitations on synthetic feasibility: The framework is a non-chiral structure;

[0019] Chemical stability: Excludes chemical skeletons that are unstable under common chemical modification conditions.

[0020] Optionally, in the above method, in S4, the purpose-oriented virtual chemical derivatization is performed under the constraints of predefined reaction rules and structural unit library, including at least one or more reaction sequences such as amidation, etherification-aminelysis, etherification-aminelysis-guanidinolation, and combined extensions of linking arms and terminal groups.

[0021] Optionally, in S5 of the above method, the second prediction model includes at least two models based on different algorithm principles. The activity of compounds in the derivative library is screened using at least two models to obtain compounds that are jointly predicted by at least two models to meet a preset activity threshold, thus forming the final candidate compound set. Optionally, the Gram-negative bacteria can be *Escherichia coli*.

[0022] A computational design system for broad-spectrum antibacterial compounds, comprising executing the computational design method for a broad-spectrum antibacterial compound as described in any one of the preceding embodiments, including:

[0023] The preliminary screening module performs virtual screening on the initial compound library to identify a preliminary set of active compounds.

[0024] The chemical skeleton extraction module extracts the preset chemical skeleton of each compound in the preliminarily identified set of active compounds, and identifies a set of dominant chemical skeletons;

[0025] The chemical skeleton screening module filters the identified advantageous chemical skeletons and selects a set of high-potential skeletons.

[0026] The virtual chemical derivation module selects one or more representative molecules from the original cluster molecular members corresponding to the selected high-potential skeletons and generates a derivative library through virtual chemical derivation.

[0027] The final screening module performs virtual screening on the derivative library to identify a final set of candidate compounds.

[0028] A computational design device for broad-spectrum antibacterial compounds, comprising: a processor, a memory, and a communication interface;

[0029] The memory stores the processor's executable instructions;

[0030] The processor executes instructions to implement a computational design method for a broad-spectrum antibacterial compound as described above.

[0031] A computer-readable storage medium storing executable instructions that, when executed by a processor, implement a computational design method for a broad-spectrum antibacterial compound as described above.

[0032] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a computational design method, system, device, and storage medium for broad-spectrum antibacterial compounds, which has the following beneficial effects: The present invention, through deep integration of multiple computational technologies, constructs an automated process for designing broad-spectrum antibacterial compounds, solving the problem that traditional computational methods are difficult to guide cross-species activity design, and providing a systematic, efficient, and feasible technical path for discovering broad-spectrum antibacterial compounds; The present invention establishes a multi-dimensional quantitative evaluation method for the potential of "broad-spectrum" modification of chemical skeletons. In this way, a scientific and reproducible selection from a massive number of "highly active molecules" to a few "high-quality potential skeletons" is realized, providing a clear data-driven foundation for subsequent targeted optimization; The method proposed in the present invention is entirely based on computer simulation, and its computational properties bring significant scale and cost advantages, which can improve screening efficiency and perform comprehensive virtual screening and multi-round optimization on chemical databases (such as natural product libraries) containing hundreds of thousands or even millions of molecules, breaking through the physical limitations of traditional screening methods in terms of throughput, time, and cost, and greatly saving research and development resources. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0034] Figure 1 A flowchart illustrating a computational design method for a broad-spectrum antibacterial compound provided by this invention;

[0035] Figure 2 A schematic diagram of the structure of a computational design system for broad-spectrum antibacterial compounds provided by the present invention;

[0036] Figure 3 This is a data flow diagram illustrating a computational design method for a broad-spectrum antibacterial compound, provided as a specific embodiment of the present invention. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0039] Reference Figure 1 As shown, this invention discloses a computational design method for broad-spectrum antibacterial compounds, comprising:

[0040] S1. Apply a first prediction model trained on known bioactivity data of Gram-positive bacteria (e.g., Staphylococcus aureus) to virtually screen the initial compound library and identify a preliminary set of active compounds.

[0041] S2. Extract the pre-defined chemical skeleton (e.g., Bemis–Murcko skeleton) of each compound in the preliminarily identified active compound set to obtain an initial list containing duplicate skeletons. Perform a deduplication operation on the initial list containing duplicate skeletons to identify a set of dominant chemical skeletons.

[0042] S3. Based on preset medicinal chemistry indicators, the identified advantageous chemical skeletons are filtered, and a set of high-potential skeletons are selected according to the drug-likeness, synthetic feasibility and chemical stability of the skeletons.

[0043] S4. From the original cluster molecular members corresponding to the selected high-potential skeleton, select one or more representative molecules and generate a derivative library through virtual chemical derivatization.

[0044] S5. Apply a second prediction model trained on known bioactivity data of Gram-negative bacteria (e.g., Escherichia coli) to virtually screen the derivative library and identify a final set of candidate compounds.

[0045] Furthermore, in S1, to improve the accuracy and robustness of the screening, the first prediction model includes at least two models based on different algorithm principles (e.g., one is a graph neural network model, and the other is an ensemble learning model based on molecular descriptors). The activity of compounds in the initial compound library is screened using at least two models to obtain a preliminary set of active compounds. The compounds included in the preliminary set of active compounds satisfy the requirement that the prediction results of at least two models meet the preset activity threshold.

[0046] Furthermore, in S2, cluster analysis can be performed on the dominant chemical skeletons in this group to assess their structural diversity and classify skeletons with similar chemical structures into different families, providing a structural classification perspective for subsequent screening.

[0047] Furthermore, in S3, the pre-defined medicinal chemistry indicators aim to evaluate the drug-likeness, synthetic feasibility, and chemical stability of the matrix. These pre-defined medicinal chemistry indicators include:

[0048] Physicochemical property limitations: molecular weight of the backbone is less than 400 Da, lipophilicity (e.g., Crippen MolLogP) is less than 3.0, the number of hydrogen bond donors is no more than 3, and the number of hydrogen bond acceptors is no more than 4.

[0049] Structural complexity constraints: The number of rotatable keys in the skeleton is no more than 5, and the number of ring systems is no more than 3;

[0050] Limitations on synthetic feasibility: The framework is a non-chiral structure;

[0051] Chemical stability: Excludes chemical skeletons that are unstable under common chemical modification conditions.

[0052] Furthermore, in S4, the purpose-oriented virtual chemical derivatization is performed under the constraints of predefined reaction rules and a library of structural units, including at least one or more reaction sequences such as amidation, etherification-amine hydrolysis, and etherification-amine hydrolysis-guanidinolation, as well as combined extensions of linker arms and terminal groups, in order to achieve controllable regulation of amphiphilicity and related physicochemical characteristics to enhance the penetration ability against Gram-negative bacteria.

[0053] Furthermore, in S5, to improve the accuracy and robustness of the final screening, the second prediction model includes at least two models based on different algorithm principles (e.g., one is a graph neural network model, and the other is an ensemble learning model based on molecular descriptors). The activity of compounds in the derivative library is screened using at least two models to obtain compounds that are jointly predicted by at least two models to meet the preset activity threshold, thus forming the final candidate compound set.

[0054] Reference Figure 2As shown, a computational design system for broad-spectrum antibacterial compounds executes the computational design method for a broad-spectrum antibacterial compound described in any one of the above-mentioned methods, including:

[0055] The preliminary screening module performs virtual screening on the initial compound library to identify a preliminary set of active compounds.

[0056] The chemical skeleton extraction module extracts the preset chemical skeleton of each compound in the preliminarily identified set of active compounds, and identifies a set of dominant chemical skeletons;

[0057] The chemical skeleton screening module filters the identified advantageous chemical skeletons and selects a set of high-potential skeletons.

[0058] The virtual chemical derivation module selects one or more representative molecules from the original cluster molecular members corresponding to the selected high-potential skeletons and generates a derivative library through virtual chemical derivation.

[0059] The final screening module performs virtual screening on the derivative library to identify a final set of candidate compounds.

[0060] A computational design device for broad-spectrum antibacterial compounds, comprising: a processor, a memory, and a communication interface;

[0061] The memory stores the processor's executable instructions;

[0062] The processor executes instructions to implement a computational design method for a broad-spectrum antibacterial compound as described above.

[0063] A computer-readable storage medium storing executable instructions that, when executed by a processor, implement a computational design method for a broad-spectrum antibacterial compound as described above.

[0064] In one specific embodiment, refer to Figure 3 As shown, novel compounds predicted to have potential activity against Gram-negative bacteria (e.g., Escherichia coli) are screened and designed from a database containing natural products and commercial compounds through a computational process executed entirely in a computer.

[0065] In this embodiment, compound structures were extracted from the COCONUT Natural Products Database (downloaded in December 2024) and the biogenic subset of the ZINC15 Commercial Compound Database (downloaded in December 2024), and the two were merged and deduplicated using their canonical SMILES descriptors. The cheminformatics toolkit RDKit (version v2023.09.1) was used to preprocess each molecule in the collection. The steps included: filtering out compounds with a value greater than 500 Da from the COCONUT natural product database; using the Standardizer module to neutralize the charge of two subsets, remove salt ions and small molecular fragments; and constructing an initial compound library containing 614,412 compounds. Subsequently, a data-rich surrogate model trained on Gram-positive bacteria (Staphylococcus aureus) was used to perform preliminary screening of the compound library. The training dataset for the surrogate model consisted of the bioactivity data (MIC values) of compounds against Staphylococcus aureus strains ATCC29213 and ATCC25923, sourced from the ChEMBL database. The threshold for defining activity was set to MIC ≤ 32 μg / mL. After data processing and screening, 9,310 samples were obtained. To improve prediction accuracy and reduce the false positive rate, this embodiment employs a dual-model ensemble prediction strategy, which specifically includes the following two models for parallel prediction and cross-validation: a graph neural network model based on Chemprop, which uses a message-passing neural network architecture to learn features from the molecular graph structure. Its key architectural parameters are: a message-passing hidden layer dimension of 1800, a network depth of 4, and a Dropout rate of 0.1 for regularization. In terms of training configuration, the model uses a batch size of 64 for 50 training rounds, with the initial, maximum, and final learning rates set to 0.00169, 0.00530, and 0.00045, respectively, to ensure optimal convergence; and an ensemble learning model based on molecular descriptors (implemented in this embodiment using the AutoGluon framework, with input features being a 2048-bit Morgan fingerprint (radius 2) and 208 physicochemical property descriptors calculated by RDKit. During training, TabularPredictor is called and presets='best_quality' is set).

[0066] To summarize the structures of the 29,443 candidate compounds selected in the initial screening, this embodiment used the Bemis–Murcko method to extract the skeletons, obtaining an initial skeleton list equal to the number of molecules. Subsequently, the skeletons were normalized (chiral markers removed, tautomer normalization performed), and duplicates were removed using canonical smiles, ultimately identifying 340 unique skeletons (hereinafter referred to as "dominant chemical skeletons"). The skeleton properties of these 340 skeletons were calculated using the RDKit with a uniform caliber (calculated after adding implicit hydrogens to the skeleton), and the following preferred hard screening rules were applied to select skeletons with higher potential for broad-spectrum modification:

[0067] (a) Molecular weight: Skeletal molecular weight < 400 Da;

[0068] (b) Lipophilicity: Crippen MolLogP < 3.0;

[0069] (c) Hydrogen bonds: Number of hydrogen bond donors (HBD) ≤ 3, number of hydrogen bond acceptors (HBA) ≤ 4;

[0070] (d) Flexibility: Number of rotatable bonds (NumRotatableBonds) ≤ 5;

[0071] (e) Complexity: Number of cycles ≤ 3;

[0072] (f) Chirality: The skeleton does not contain chiral centers.

[0073] After screening using multiple rigorous criteria, 175 skeletons were selected from 340 promising skeletons to form a set of high-potential skeletons for subsequent derivation.

[0074] From the 175 high-potential skeletal clusters selected, a representative score was used... For each cluster, several representative molecules are selected for subsequent derivation; the system can calculate a representativeness score for each molecule within a cluster. The score can be determined using the following formula:

[0075]

[0076] This represents the activity probability value of the single molecule predicted by the dual-model ensemble (value range: 0~1). The chemical modifiability score is calculated based on its own structure (values ​​range from 0 to 1, based on the normalized result of the number / accessibility of reactive sites and functional group conflict penalty). As the weight of the modifiability score, The weighting of the activity score; in this embodiment Set to 0.7, The value is set to 0.3 to indicate the emphasis on chemical modifiability in the molecular selection strategy; based on the selected representative molecules, purpose-oriented virtual chemical derivatization is performed through the virtual chemical derivatization module to generate a set of virtual derivative libraries; subsequently, based on the selected representative molecules, molecular derivatization is performed through the virtual chemical derivatization module; this module contains a structural unit library and a reaction rule engine.

[0077] The structural unit library includes:

[0078] Connecting arm library: Contains a series of dibromo-n-alkanes with carbon chain lengths ranging from dibromoethane to dibromododecane (C2-C12) to provide a wide range of hydrophobicity tuning;

[0079] Terminal group library: contains 142 amine compounds, covering a variety of chemical types such as aliphatic amines, aromatic amines and cyclic secondary amines, and sets the ratio of primary amines to secondary amines to 1:1;

[0080] The engine uses reaction-based SMARTS for enumeration and constraints, including the following three types of reaction sequences:

[0081] (a) Amideation: Nucleophilic amines are coupled to carboxylic acid derivatives (acids / acyl chlorides / activated esters) to form amide bonds;

[0082] (b) Etherification-Amine hydrolysis: Nucleophiles are activated to form ether bonds with linkers, followed by substitution / amine hydrolysis with amines to generate the target substituted product;

[0083] (c) Etherification-Aminolysis-Guidinolation: Based on (b), the terminal primary amine is modified by guanidinolation (e.g., using conventional guanidinolation reagents).

[0084] The reaction rules are implemented using predefined reaction SMARTS, with site selection priority at primary alcohol / primary amine levels; if functional group incompatibility or protecting group conflict exists, the path is skipped. Molecules without compatible reaction sites do not generate derivatives. Derivatives are generated according to the above procedure, and a full library deduplication and statistical analysis are performed using canonical SMILES, ultimately yielding 90,244 candidate derivative records, of which 78,610 are unique chemical structures.

[0085] For the 78,610 derivatives generated, a dual-model ensemble strategy was used for final screening. The training dataset consisted of the bioactivity data (MIC values) of the compounds against strain ATCC25922, which were obtained from literature mining and the ChEMBL database. The threshold for defining activity was set to MIC ≤ 32 μg / mL. After data processing and screening, 8,557 samples were obtained. Two prediction models based on different algorithm principles and trained on Gram-negative bacteria (Escherichia coli) were applied: one was a graph neural network model based on Chemprop, specifically trained using E. coli activity data. The basic architecture and training process of this model were inherited from the aforementioned Staphylococcus aureus model. Its key architectural parameters were: the message passing hidden layer dimension was 1500, Dropout was 0.2, and its initial, maximum, and final learning rates were set to 3.8017e-05, 0.00267, and 0.00010, respectively; the other was an ensemble learning model based on the AutoGluon framework. The basic architecture and training process of this model were inherited from the aforementioned Staphylococcus aureus model, but the training data was E. coli ATCC25922 activity data.

[0086] The system first filters out compounds that are predicted by both models to meet the activity threshold (prediction score > 0.5), forming a high-confidence preliminary candidate pool. To rank the compounds in this candidate pool, this embodiment defines a weighted fusion score, calculated as follows: ;in and The activity probability values ​​output by the Chemprop model and the AutoGluon model are respectively set to (0~1).

[0087] The weighting is designed to favor the more stable Chemprop model. The system sorts all compounds in the candidate pool in descending order based on the calculated weighted fusion scores and selects the top 100 compounds. These 100 compounds are then subject to final human decision-making. Factors considered in the human decision-making include, but are not limited to: (a) feasibility assessment of synthetic routes based on existing literature; and (b) commercial availability and cost analysis of key intermediates and raw materials.

[0088] As illustrated in this embodiment, the method and system disclosed in this invention successfully designed 100 novel broad-spectrum candidate compounds predicted to have high activity against Gram-negative bacteria, starting from an initial compound library containing over 600,000 compounds. This was achieved through a series of systematic and automated computational steps, including: initial screening based on a dual-model ensemble, backbone identification based on cluster analysis, potential backbone evaluation based on weighted scoring, goal-oriented virtual derivatization, and final precise screening based on a dual-model ensemble. Because these compounds were systematically endowed with enhanced membrane-penetrating ability (amphiphilicity) and predicted to have high anti-Gram-negative bacterial activity during the design process, they are significantly more likely to exhibit effective activity against Gram-negative bacteria in subsequent in vitro bioactivity tests than molecules obtained through traditional screening methods. This verifies that this invention can produce the expected beneficial technical effects.

[0089] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0090] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A computational design method for broad-spectrum antibacterial compounds, characterized in that, include: S1. Apply the first prediction model trained based on Gram-positive bacteria bioactivity data to virtually screen the initial compound library and identify a preliminary set of active compounds. S2. Extract the preset chemical skeleton of each compound in the preliminarily identified set of active compounds to obtain an initial list containing duplicate skeletons. Perform a deduplication operation on the initial list containing duplicate skeletons to identify a set of dominant chemical skeletons. S3. Based on preset medicinal chemistry indicators, the identified advantageous chemical skeletons are filtered, and a set of high-potential skeletons are selected according to the drug-likeness, synthetic feasibility and chemical stability of the skeletons. Pre-defined medicinal chemistry indicators include: Physicochemical limitations: molecular weight of the backbone, lipophilicity, number of hydrogen bond donors and hydrogen bond acceptors; Structural complexity constraints: number of rotatable keys and number of ring systems in the skeleton; Limitations on synthetic feasibility: The framework is a non-chiral structure; Chemical stability: Excluding chemical skeletons that are unstable under common chemical modification conditions; S4. From the original cluster molecular members corresponding to the selected high-potential skeleton, select one or more representative molecules and generate a derivative library through virtual chemical derivatization. S5. Apply the second prediction model trained based on Gram-negative bacterial bioactivity data to virtually screen the derivative library and identify a final set of candidate compounds.

2. The computational design method for a broad-spectrum antibacterial compound according to claim 1, characterized in that, In S1, the first prediction model includes at least two models based on different algorithm principles. The activity of compounds in the initial compound library is screened using at least two models to obtain a preliminary set of active compounds. The compounds included in the preliminary set of active compounds meet the preset activity threshold requirements as predicted by at least two models.

3. The computational design method for a broad-spectrum antibacterial compound according to claim 1, characterized in that, In S2, the pre-defined chemical skeleton of the compound is the Bemis–Murcko skeleton.

4. The computational design method for a broad-spectrum antibacterial compound according to claim 1, characterized in that, In S4, virtual chemical derivatization adopts a goal-oriented strategy. The goal-oriented virtual chemical derivatization is performed under the constraints of predefined reaction rules and structural unit library, including at least one or more reaction sequences such as amidation, etherification-aminelysis, etherification-aminelysis-guanidinolation, and combined extensions of linker arms and terminal groups.

5. The computational design method for a broad-spectrum antibacterial compound according to claim 1, characterized in that, In S5, the second prediction model includes at least two models based on different algorithm principles. The activity of compounds in the derivative library is screened using at least two models to obtain compounds that are jointly predicted by at least two models to meet the preset activity threshold, thus forming the final candidate compound set.

6. A computational design system for broad-spectrum antibacterial compounds, executing the computational design method for broad-spectrum antibacterial compounds as described in any one of claims 1-5, characterized in that, include: The preliminary screening module performs virtual screening on the initial compound library to identify a preliminary set of active compounds. The chemical skeleton extraction module extracts the preset chemical skeleton of each compound in the preliminarily identified set of active compounds, and identifies a set of dominant chemical skeletons; The chemical skeleton screening module filters the identified advantageous chemical skeletons and selects a set of high-potential skeletons. The virtual chemical derivation module selects one or more representative molecules from the original cluster molecular members corresponding to the selected high-potential skeletons and generates a derivative library through virtual chemical derivation. The final screening module performs virtual screening on the derivative library to identify a final set of candidate compounds.

7. A computational design apparatus for broad-spectrum antibacterial compounds, characterized in that, include: Processor, memory, and communication interface; The memory stores the processor's executable instructions; The processor executes instructions to implement a computational design method for a broad-spectrum antibacterial compound as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable instructions that, when executed by a processor, implement a computational design method for a broad-spectrum antibacterial compound as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Artificial intelligence engine for generating candidate medicaments

    CN114270376A

  • Organic molecule virtual screening library construction method, device, equipment and medium

    CN118412066A