Calculation design method, system and equipment for broad-spectrum antibacterial compound and storage medium
By performing multi-step computational design on the compound library, including screening, skeleton extraction and derivation, the problem of narrow-spectrum to broad-spectrum antimicrobial drug discovery in the prior art has been solved, and efficient design of broad-spectrum antimicrobial compounds has been achieved.
Patent Information
- Application Number
- CN202511393854.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing computer-aided drug design methods face difficulties in the systematic transformation from narrow-spectrum to broad-spectrum antimicrobial drug discovery, lack of model robustness, difficulty in cross-species knowledge transfer, and fragmented processes that are difficult to standardize and reuse.
A first prediction model trained on Gram-positive bacteria bioactivity data was used for virtual screening to identify a preliminary set of active compounds, extract chemical skeletons, select high-potential skeletons, generate a derivative library through virtual chemical derivation, and then a second prediction model trained on Gram-negative bacteria bioactivity data was used for screening to identify the final candidate compounds.
This has enabled a systematic transformation from narrow-spectrum compounds to broad-spectrum antibacterial compounds, improved the efficiency and accuracy of cross-species activity design, saved R&D resources, and provided a scientific data-driven approach.
Smart Images

Figure CN120977436A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computational chemistry and drug discovery, and in particular to a method, system, device and storage medium for the computational design of broad-spectrum antibacterial compounds. BACKGROUND
[0002] With the widespread use of antibacterial drugs, bacterial drug resistance has become one of the global public health challenges, and the problem of drug resistance of gram-negative bacteria is particularly serious. Gram-negative bacteria have a double-membrane structure composed of a cytoplasmic membrane and an outer membrane, and the outer membrane is rich in negatively charged lipopolysaccharides (LPS), which form a selective permeability barrier to small molecules. To achieve effective action on negative bacteria, candidate molecules must cross the outer membrane and maintain sufficient concentration in the bacterial intracellular. Among them, "amphiphilic design" (balance between hydrophobic fragments and hydrophilic / ionizable fragments) is one of the core strategies.
[0003] In recent years, with the development of computing technology, computer-aided drug design (CADD) and artificial intelligence drug discovery (AIDD) have provided new ways for antibacterial drug discovery. The existing CADD / AIDD process mainly includes docking / virtual screening based on target structure, ligand-based modeling, molecule generation and structure optimization, and ADME / T permeability prediction, etc., although each has its own value, but there are common limitations in the "narrow-spectrum to broad-spectrum" system transformation: the model is trained based on single bacterial data, and the cross-bacterial knowledge transfer is insufficient; the generation / derivation link does not process the key physicochemical targets such as "amphiphilicity" in advance, and the structure search is easy to deviate from the "membrane-penetrable" feature space; the decision relies on a single model, and the robustness is insufficient; the process is fragmented, and it is difficult to standardize and reuse.
[0004] Therefore, to solve the problems existing in the prior art, a method, system, device and storage medium for the computational design of broad-spectrum antibacterial compounds are proposed, which are urgently needed by those skilled in the art. SUMMARY
[0005] Therefore, to solve the problems existing in the prior art, a method, system, device and storage medium for the computational design of broad-spectrum antibacterial compounds are proposed, which are urgently needed by those skilled in the art.
[0006] To achieve the above purpose, the present application adopts the following technical solutions:
[0007] A method for the computational design of broad-spectrum antibacterial compounds, comprising:
[0008] S1, applying a first prediction model trained based on gram-positive bacterial biological activity data to perform virtual screening on an initial compound library to identify a preliminary active compound set;
[0009] S2, extract the preset chemical skeleton of each compound in the identified preliminary active compound set, obtain an initial list containing repeated skeletons, and perform a de-duplication operation on the initial list containing repeated skeletons to identify a set of dominant chemical skeletons;
[0010] S3, filtering the identified dominant chemical skeletons based on preset drug chemistry indicators, and selecting a set of high-potential skeletons according to the drug-likeness, synthetic feasibility, and chemical stability of the skeletons;
[0011] S4, selecting one or more representative molecules from the molecular members in the original cluster corresponding to the selected high-potential skeletons, and generating a derivative library through virtual chemical derivation;
[0012] S5, applying a second prediction model trained based on Gram-negative bacteria biological activity data to perform virtual screening on the derivative library to identify a final candidate compound set.
[0013] The above method, optionally, in S1, the first prediction model includes at least two models based on different algorithm principles, and the activity of the compounds in the initial compound library is screened using at least two models to obtain a preliminary active compound set, wherein the compounds included in the preliminary active compound set meet the requirement that the prediction results of at least two models meet the preset activity threshold value. Optionally, the Gram-positive bacteria can be Staphylococcus aureus.
[0014] The above method, optionally, in S2, the preset chemical skeleton of the compound is Bemis-Murcko skeleton.
[0015] The above method, optionally, in S3, the preset drug chemistry indicators include:
[0016] Physical and chemical property restrictions: molecular weight, lipophilicity, number of hydrogen bond donors, and number of hydrogen bond acceptors of the skeleton;
[0017] Structural complexity restrictions: number of rotatable bonds and number of ring systems of the skeleton;
[0018] Synthetic feasibility restrictions: the skeleton is a non-chiral structure;
[0019] Chemical stability: excluding chemical skeletons that are unstable under common chemical modification conditions.
[0020] The above method, optionally, in S4, the purpose-oriented virtual chemical derivation is performed under the constraints of predefined reaction rules and structure unit library, at least including one or more of the following reaction sequences: amidation, etherification-amination, etherification-amination-guanylation, and combination expansion of connecting arms and terminal groups.
[0021] The method, optionally, in S5, the second prediction model comprises at least two models based on different algorithm principles, the activity of the compounds in the derivative library is screened using at least two models, the compounds predicted by at least two models to meet the preset activity threshold are obtained, and a final candidate compound set is constituted. Optionally, the gram-negative bacteria can be escherichia coli.
[0022] A system for the computational design of broad-spectrum antibacterial compounds, which performs the method for the computational design of broad-spectrum antibacterial compounds according to any one of the preceding claims, comprising:
[0023] A preliminary screening module, which performs virtual screening on the initial compound library to identify a preliminary active compound set;
[0024] A chemical skeleton extraction module, which extracts a preset chemical skeleton of each compound in the identified preliminary active compound set to identify a group of dominant chemical skeletons;
[0025] A chemical skeleton screening module, which filters the identified dominant chemical skeletons to select a group of high-potential skeletons;
[0026] A virtual chemical derivation module, which selects one or more representative molecules from the molecular members in the original cluster corresponding to the selected high-potential skeletons, and generates a derivative library through virtual chemical derivation;
[0027] A final screening module, which performs virtual screening on the derivative library to identify a final candidate compound set.
[0028] A system for the computational design of broad-spectrum antibacterial compounds, comprising: a processor, a memory, and a communication interface;
[0029] The memory stores executable instructions of the processor;
[0030] The processor executes the instructions to implement the method for the computational design of broad-spectrum antibacterial compounds according to any one of the preceding claims.
[0031] A computer-readable storage medium, which stores executable instructions, and the instructions, when executed by a processor, implement the method for the computational design of broad-spectrum antibacterial compounds according to any one of the preceding claims.
[0032] Compared with the prior art, the application provides a method, system, device and storage medium for the computational design of broad-spectrum antibacterial compounds, and has the following beneficial effects: the application builds an automatic process for designing broad-spectrum antibacterial compounds by deeply integrating multiple computing technologies, solves the problem that traditional computing methods are difficult to guide cross-species activity design, and provides a systematic, efficient and implementable technical path for discovering broad-spectrum antibacterial compounds; the application establishes a multi-dimensional quantitative evaluation method for the "broad-spectrum" modification potential of chemical skeletons, and through this method, scientific and reproducible selection is realized from a large number of "high-activity molecules" to a few "high-quality potential skeletons", which provides a clear data-driven basis for subsequent targeted optimization; the method proposed by the application is completely based on computer simulation, and the computational properties bring significant scale and cost advantages, which can perform comprehensive virtual screening and multiple rounds of optimization on a chemical database (such as a natural product library) containing hundreds of thousands or even millions of molecules with an efficiency far exceeding traditional wet experiments by several orders of magnitude, breaking through the physical limitations of traditional screening methods in throughput, time and cost, and greatly saving research and development resources. BRIEF DESCRIPTION OF DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.
[0034] Figure 1 A flowchart of a method for the computational design of broad-spectrum antibacterial compounds provided by the present application;
[0035] Figure 2 A structural schematic diagram of a system for the computational design of broad-spectrum antibacterial compounds provided by the present application;
[0036] Figure 3 A data flow schematic diagram of a method for the computational design of broad-spectrum antibacterial compounds in a specific embodiment provided by the present application. DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0038] In this application, such terms as first and second, etc., are used to distinguish one entity or operation from another without necessarily requiring or implying any such actual relationship or order between such entities or operations, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by an expression "comprising a..." does not exclude the existence of additional identical elements in the process, method, article, or apparatus including the element.
[0039] Referring to Figure 1 The present application discloses a method for computationally designing broad-spectrum antibacterial compounds, comprising:
[0040] S1, applying a first prediction model trained based on known biological activity data of gram-positive bacteria (e.g., Staphylococcus aureus) to perform virtual screening on an initial compound library to identify a preliminary active compound set;
[0041] S2, extracting a pre-set chemical skeleton (e.g., Bemis-Murcko skeleton) of each compound in the identified preliminary active compound set to obtain an initial list containing repetitive skeletons, performing a deduplication operation on the initial list containing repetitive skeletons to identify a set of dominant chemical skeletons;
[0042] S3, filtering the identified dominant chemical skeletons based on pre-set drug chemistry indicators, and selecting a set of high-potential skeletons according to the drug-likeness, synthetic feasibility, and chemical stability of the skeletons;
[0043] S4, selecting one or more representative molecules from the molecular members in the original cluster corresponding to the selected high-potential skeletons, and generating a derivative library through virtual chemical derivatization;
[0044] S5, applying a second prediction model trained based on known biological activity data of gram-negative bacteria (e.g., Escherichia coli) to perform virtual screening on the derivative library to identify a final candidate compound set.
[0045] Further, in S1, to improve the accuracy and robustness of the screening, the first prediction model comprises at least two models based on different algorithm principles (for example, one is a graph neural network model, and the other is an ensemble learning model based on molecular descriptors), and at least two models are used to screen the activity of the compounds in the initial compound library to obtain a preliminary active compound set, wherein the compounds included in the preliminary active compound set meet the prediction results of at least two models that meet the preset activity threshold requirement.
[0046] Further, in S2, the group of dominant chemical scaffolds can be further subjected to cluster analysis to evaluate their structural diversity and divide the chemical structures similar to the scaffolds into different families, thereby providing a structural classification perspective for subsequent screening.
[0047] Further, in S3, the preset drug chemistry indicators aim to evaluate the drug-likeness, synthetic feasibility and chemical stability of the scaffold, and the preset drug chemistry indicators include:
[0048] Physical and chemical property restrictions: the molecular weight of the scaffold is less than 400 Da, the lipophilicity (such as Crippen MolLogP) is less than 3.0, the number of hydrogen bond donors is not more than 3, and the number of hydrogen bond acceptors is not more than 4;
[0049] Structural complexity restrictions: the number of rotatable bonds of the scaffold is not more than 5, and the number of ring systems is not more than 3;
[0050] Synthetic feasibility restrictions: the scaffold is a non-chiral structure;
[0051] Chemical stability: exclude chemical scaffolds that are unstable under common chemical modification conditions.
[0052] Further, in S4, the purpose-oriented virtual chemical derivation is performed under the constraints of predefined reaction rules and structure unit library, at least including one or more of the following reaction sequences: amidation, etherification-amination, etherification-amination-guanylation, and combination expansion of connecting arms and terminal groups, to achieve controllable adjustment of amphiphilicity and related physicochemical characteristics to enhance the penetration ability to gram-negative bacteria.
[0053] Further, in S5, to improve the accuracy and robustness of the final screening, the second prediction model comprises at least two models based on different algorithm principles (for example, one is a graph neural network model, and the other is an ensemble learning model based on molecular descriptors), and at least two models are used to screen the activity of the compounds in the derivative library to obtain compounds that are predicted by at least two models to meet the preset activity threshold, thereby constituting a final candidate compound set.
[0054] Reference Figure 2As shown, a system for computationally designing broad-spectrum antibacterial compounds, performing the method for computationally designing broad-spectrum antibacterial compounds of any of the preceding claims, comprising:
[0055] A preliminary screening module performs virtual screening on the initialized compound library to identify a preliminary active compound set;
[0056] A chemical skeleton extraction module extracts a pre-set chemical skeleton of each compound in the identified preliminary active compound set to identify a group of dominant chemical skeletons;
[0057] A chemical skeleton screening module filters the identified dominant chemical skeletons to select a group of high-potential skeletons;
[0058] A virtual chemical derivation module selects one or more representative molecules from the molecular members in the original cluster corresponding to the selected high-potential skeletons, and generates a derivative library through virtual chemical derivation;
[0059] A final screening module performs virtual screening on the derivative library to identify a final candidate compound set.
[0060] A system for computationally designing broad-spectrum antibacterial compounds, comprising: a processor, a memory, and a communication interface;
[0061] The memory stores executable instructions of the processor;
[0062] The processor executes the instructions to implement the method for computationally designing broad-spectrum antibacterial compounds of any of the preceding claims.
[0063] A computer-readable storage medium, the computer-readable storage medium storing executable instructions, the instructions being executed by a processor to implement the method for computationally designing broad-spectrum antibacterial compounds of any of the preceding claims.
[0064] In one specific embodiment, referring to Figure 3 As shown, from a database containing natural products and commercial compounds, a new compound with potential activity against gram-negative bacteria (e.g., E. coli) is screened and designed through a completely computer-executed computational process.
[0065] In this embodiment, the compound structures were extracted from the COCONUT natural product database (downloaded in December 2024) and the biogenic subset of the ZINC15 commercial compound database (downloaded in December 2024), combined and de-duplicated using their canonical SMILES descriptors. Each molecule in the set was pre-processed using the cheminformatics toolkit RDKit (version v2023.09.1), steps included: filtering out compounds above 500 Da from the COCONUT natural product database; using the Standardizer module to neutralize charges, remove salt ions and small molecular fragments for both subsets; an initial compound library containing 614412 compounds was constructed; a data-rich proxy model trained on Gram-positive bacteria (Staphylococcus aureus) was then used to perform preliminary screening on the compound library; the training data set of the proxy model was the biological activity data (MIC value) of the compound on Staphylococcus aureus ATCC29213 and ATCC25923 strains, which was obtained from the ChEMBL database, and the threshold for defining activity was set as MIC≤32μg / mL, and after data processing and screening, 9310 samples were obtained. In order to improve the prediction accuracy and reduce the false positive rate, this embodiment adopts a double-model integrated prediction strategy, which specifically includes the following two models for parallel prediction and cross-validation: a graph neural network model implemented based on Chemprop, which learns features from molecular graph structure using a message passing neural network architecture, with the following key architecture parameters: message passing hidden layer dimension 1800, network depth 4, and 0.1 Dropout rate for regularization. In terms of training configuration, the model uses a batch size of 64 for 50 rounds of training, with the initial, maximum and final learning rates set to 0.00169, 0.00530 and 0.00045 respectively to ensure optimal convergence of the model; an integrated learning model based on molecular descriptors (implemented in this embodiment based on the AutoGluon framework, input features are 2048-bit Morgan fingerprints (radius 2) and 208 physicochemical property descriptors calculated by RDKit. During training, TabularPredictor is called and presets='best_quality' is set.
[0066] To structure induce the 29443 candidate compounds after the primary screening, the Bemis-Murcko method is used to extract the skeleton in this embodiment, and an initial list of skeletons equivalent to the number of molecules is obtained. Then, the skeleton is standardized (remove chiral markers, and perform tautomer standardization), and the canonical SMILES of the skeleton is de-duplicated, and finally 340 unique skeletons (hereinafter referred to as “advantageous chemical skeletons”) are identified. The above 340 skeletons are calculated for their skeleton properties (after calculating the implicit hydrogen on the skeleton) according to the unified standard by RDKit, and the following preferred hard screening rules are applied to select skeletons with higher potential for broad-spectrum modification:
[0067] (a) molecular weight: skeleton molecular weight < 400 Da;
[0068] (b) lipophilicity: Crippen MolLogP < 3.0;
[0069] (c) hydrogen bond: number of hydrogen bond donors (HBD) ≤ 3, number of hydrogen bond acceptors (HBA) ≤ 4;
[0070] (d) flexibility: number of rotatable bonds (NumRotatableBonds) ≤ 5;
[0071] (e) complexity: ring system number ≤ 3;
[0072] (f) chirality: the skeleton does not contain chiral centers.
[0073] After multi-dimensional hard standard screening, 175 skeletons are finally selected from the 340 advantageous skeletons, constituting a high-potential skeleton set for subsequent derivation.
[0074] From the 175 high-potential skeleton clusters screened, a representative score Score repr is calculated for each cluster, and several representative molecules are selected for subsequent derivation; the system can calculate a representative score Score repr for each molecule in the cluster, which can be determined by the following formula:
[0075] Score repr = w site ·Score site +w activity ·Score activity
[0076] Score activity is the activity probability value (value 0-1) of the single molecule predicted by the double model integration; Score siteChemical modifiability score calculated based on its own structure, (value 0~1, based on the normalized result of reactive site number / accessibility and functional group conflict penalty); w site Weight of modifiability score, w activity Weight of activity score; in this embodiment w site Set as 0.7, w activity Set as 0.3, to represent the emphasis on chemical modifiability in the molecule selection strategy; based on the selected representative molecule, perform purpose-oriented virtual chemical derivation through the virtual chemical derivation module to generate a set of virtual derivative library; then, based on the selected representative molecule, perform molecule derivation through the virtual chemical derivation module; the module includes a structure unit library and a reaction rule engine.
[0077] The structure unit library includes:
[0078] Linker arm library: contains a series of dibromo-n-alkanes, with carbon chain length covering from dibromoethane to dibromododecane (C2-C12), to provide a wide range of hydrophobicity adjustment;
[0079] End group library: contains 142 amine compounds, covering various chemical types such as aliphatic amines, aromatic amines, and cyclic secondary amines, with the ratio of primary amines to secondary amines set at 1:1;
[0080] The engine uses reaction SMARTS-based enumeration and constraints, including the following three types of reaction sequences:
[0081] (a) Amidation: coupling of nucleophilic amine with carboxylic acid derivative (acid / acid chloride / activated ester) to form amide bond;
[0082] (b) Etherification-aminolysis: nucleophilic alcohol forms ether bond with activated linker, then substituted / aminolysis with amine to generate target substitution;
[0083] (c) Etherification-aminolysis-guanidination: guanidination modification of terminal primary amine based on (b) (e.g. using conventional guanidination reagent).
[0084] Reaction rules are implemented with predefined reaction SMARTS, with site selection priority as primary alcohol / primary amine; if there is functional group incompatibility or protecting group conflict, skip this path. No derivative is generated for molecules without compatible reaction sites. Derivatives are generated according to the above process, and canonical SMILES is used for full library deduplication and statistics, finally obtaining 90244 candidate derivative records, of which 78610 are unique chemical structures.
[0085] For the generated 78610 derivatives, a final screening was performed using a combination of double model integration strategy, the training data set was the biological activity data (MIC value) of the compounds on the ATCC25922 strain, the data was obtained from literature mining and ChEMBL database, and the threshold value for defining activity was set as MIC≤32μg / mL, after data processing and screening, 8557 samples were obtained. Two prediction models trained for gram-negative bacteria (E. coli) based on different algorithm principles were applied: one was a graph neural network model implemented by Chemprop, which was specially trained with E. coli activity data, the basic architecture and training process of the model inherited from the aforementioned S. aureus model, and the key architecture parameters were: the message passing hidden layer dimension was 1500, the Dropout was 0.2, and the initial, maximum and final learning rates were set as 3.8017e-05, 0.00267 and 0.00010 respectively; one was an ensemble learning model implemented based on the AutoGluon framework, the input feature basic architecture and training process of the model inherited from the aforementioned S. aureus model, but the training data was the E. coli ATCC25922 activity data.
[0086] The system first screened out the compounds that were simultaneously predicted by the above two models to meet the activity threshold (prediction score > 0.5), forming a high confidence preliminary candidate pool; in order to sort the compounds in the candidate pool, the embodiment defines a weighted fusion score (Weighted Score), the calculation formula of which is: WS = 0.6·P chemprop +0.4·P autogluon ; Wherein P chemprop and P autogluon are the activity probability values (0-1) output by the Chemprop model and the AutoGluon model respectively.
[0087] The weight setting aims to tilt towards the more stable Chemprop model, the system ranks all the compounds in the candidate pool in descending order based on the calculated weighted fusion score, and selects the top 100 compounds, these 100 compounds are then subjected to final manual decision. The factors considered in manual decision include but are not limited to: (a) feasibility evaluation of synthesis route based on existing literature; (b) commercial availability and cost analysis of key intermediates and raw materials.
[0088] As shown in the present embodiment, the disclosed method and system successfully started from an initial compound library containing more than 600,000 compounds, through a series of systematic and automated computational steps, including: initial screening based on double-model integration, skeleton identification based on cluster analysis, potential skeleton evaluation based on weighted scoring, goal-oriented virtual derivation, and finally double-model-based precise screening, ultimately designed 100 new broad-spectrum candidate compounds predicted to have high activity against gram-negative bacteria. These compounds, due to their enhanced membrane penetration ability (amphiphilicity) and high predicted activity against gram-negative bacteria during the design process, are therefore more likely to exhibit effective activity against gram-negative bacteria in subsequent in vitro biological activity tests than molecules obtained through traditional screening methods. This verifies that the present application can produce predictable beneficial technical effects.
[0089] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, it is described more simply, and the relevant part can be referred to the part of the method embodiment. The above-described system and system embodiment are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0090] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of computational design of broad-spectrum antibacterial compounds, characterized in that, The method comprises the following steps: S1, applying a first prediction model trained based on biological activity data of gram-positive bacteria to perform virtual screening on an initial compound library to identify a preliminary active compound set; S2, extracting a preset chemical skeleton of each compound in the identified preliminary active compound set to obtain an initial list containing repetitive skeletons, performing a de-duplication operation on the initial list containing repetitive skeletons to identify a group of dominant chemical skeletons; S3, filtering the identified dominant chemical skeletons based on preset drug chemistry indicators, and selecting a group of high-potential skeletons according to the drug-likeness, synthetic feasibility and chemical stability of the skeletons; S4, selecting one or more representative molecules from the original cluster corresponding to the selected high-potential skeletons, and generating a derivative library through virtual chemical derivation; S5, applying a second prediction model trained based on biological activity data of gram-negative bacteria to perform virtual screening on the derivative library to identify a final candidate compound set.
2. The method for computationally designing broad-spectrum antibacterial compounds according to claim 1, wherein in S1, the first prediction model comprises at least two models based on different algorithm principles, and the activity of the compounds in the initial compound library is screened using at least two models to obtain a preliminary active compound set, wherein the compounds included in the preliminary active compound set satisfy the prediction results of at least two models meeting the preset activity threshold requirement.
3. The method for computationally designing broad-spectrum antibacterial compounds according to claim 1, wherein in S2, the preset chemical skeleton of the compound is a Bemis-Murcko skeleton.
4. The method for computationally designing broad-spectrum antibacterial compounds according to claim 1, wherein in S3, the preset drug chemistry indicators include: Physical and chemical property restrictions: molecular weight, lipophilicity, number of hydrogen bond donors, and number of hydrogen bond acceptors of the skeleton; Structural complexity restrictions: number of rotatable bonds and number of ring systems of the skeleton; Synthetic feasibility restrictions: the skeleton is a non-chiral structure; Chemical stability: excluding chemical skeletons unstable under common chemical modification conditions.
5. The method for computationally designing broad-spectrum antibacterial compounds according to claim 1, wherein in S4, the virtual chemical derivation adopts a goal-oriented strategy, and the goal-oriented virtual chemical derivation is performed under the constraints of predefined reaction rules and structure unit library, at least including one or more of the following reaction sequences: amidation, etherification-amination, etherification-amination-guanylation, and combination expansion of connecting arms and terminal groups.
6. The method for computationally designing broad-spectrum antibacterial compounds according to claim 1, wherein in S5, the second prediction model comprises at least two models based on different algorithm principles, and the activity of the compounds in the derivative library is screened using at least two models to obtain compounds predicted by at least two models to meet the preset activity threshold, which constitutes the final candidate compound set. The method comprises the following steps: A preliminary screening module performs virtual screening on an initial compound library to identify a preliminary active compound set; 7. A system for the in silico design of broad-spectrum antibacterial compounds, which performs a method for the in silico design of broad-spectrum antibacterial compounds according to any one of claims 1 to 6, characterized in that, a chemical skeleton extraction module, which extracts a preset chemical skeleton of each of the identified preliminary active compound set, and identifies a group of dominant chemical skeletons; a chemical skeleton screening module, which filters the identified dominant chemical skeletons, and selects a group of high-potential skeletons; a virtual chemical derivation module, which selects one or more representative molecules from the molecular members in the original cluster corresponding to the selected high-potential skeleton, and generates a derivative library through virtual chemical derivation; a final screening module, which performs virtual screening on the derivative library, and identifies a final candidate compound set.
8. A device for the computational design of broad-spectrum antibacterial compounds, characterized by, comprising: a processor, a memory, and a communication interface; the memory stores executable instructions of the processor; the processor executes the instructions to implement the method for the computational design of broad-spectrum antibacterial compounds according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, a computer-readable storage medium stores executable instructions, and the instructions, when executed by a processor, implement the method for the computational design of broad-spectrum antibacterial compounds according to any one of claims 1-6.
Citation Information
Patent Citations
Artificial intelligence engine for generating candidate medicaments
CN114270376A
Organic molecule virtual screening library construction method, device, equipment and medium
CN118412066A
Method and apparatus for processing molecular scaffold transition, medium, electronic device, and computer program product
US20230083810A1
Automation-ready DNA cloning by bacterial natural transformation
WO2020201022A1