Automated molecular docking and screening analysis methods
Through automated molecular docking and screening analysis methods, the problems of low efficiency and insufficient screening in existing technologies are solved, and efficient and accurate molecular screening and enzyme catalysis optimization are achieved, which is suitable for drug design and enzyme catalysis optimization.
Patent Information
- Application Number
- CN202510854524.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing molecular docking methods lack automation and screening analysis capabilities, resulting in low efficiency and susceptibility to human factors. Relying solely on binding energy screening may be insufficient and unable to effectively consider the spatial relationship between the ligand and the catalytic site.
Automated molecular docking and screening analysis methods are used, including protein structure alignment, receptor processing, batch molecular docking, catalytic site distance calculation and multi-dimensional screening, combined with binding energy and spatial relationship for screening to generate visual results.
It achieves efficient and automated molecular docking and screening, improves the accuracy and efficiency of molecular screening, is suitable for drug design and enzyme catalysis optimization, and supports large-scale molecular screening tasks.
Smart Images

Figure CN120375912B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular docking, and in particular to an automated molecular docking and screening analysis method. Background Art
[0002] Molecular docking is a key technique in drug design, enzymology research, and other biomolecular disciplines, capable of simulating the binding patterns of small molecules (such as drug molecules) with target receptors (such as proteins). Traditional molecular docking methods typically rely on manual operations or semi-automated tools, such as the molecular docking method and system based on the AutoDock visualization platform disclosed in Publication No. CN115295091A. While this method addresses the operational difficulties associated with manual command-line script invocation in existing visualization platforms for molecular docking, it only performs single docking calculations and lacks subsequent screening and analysis capabilities, forcing researchers to manually organize large amounts of data, which is inefficient and susceptible to human error. Furthermore, relying solely on binding energy (affinity) as a screening criterion may be insufficient. For example, in enzyme catalysis optimization, the spatial relationship between the ligand and the catalytic site is crucial to catalytic efficiency.
[0003] Therefore, there is an urgent need for an automated, efficient molecular docking method with screening and analysis functions, which can not only perform batch docking, but also automatically screen and analyze complexes based on structural characteristics, thereby improving the accuracy of molecular screening and experimental efficiency. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide an automated molecular docking and screening analysis method.
[0005] The technical solution adopted by this application to solve the technical problem is: the automated molecular docking and screening analysis method includes the following steps:
[0006] S1: Protein structure alignment; batch align multiple protein structures to ensure their spatial consistency with the reference structure;
[0007] S2: Process receptor protein files; identify active sites, automatically generate docking grids, and optimize receptor structures;
[0008] S3: Batch molecular docking: perform batch docking, automatically extract ligand binding conformations, and automatically generate protein-ligand complexes to facilitate subsequent analysis;
[0009] S4: Calculate the distance between substrate and catalytic site; calculate the geometric center distance between ligand and key residues in catalytic site to screen the optimal binding mode;
[0010] S5: Combined screening of binding energy and spatial relationship; automatic screening based on docking affinity and catalytic site distance to remove unreasonable binding modes;
[0011] S6: Results summary and visualization; generate binding energy distribution diagram, distance distribution diagram, calculation time comparison diagram, and export the screened candidate molecule list.
[0012] The S1 includes the following sub-steps:
[0013] S1-1: Match key atom pairs; assume the reference molecular structure is The target molecular structure is , each molecule consists of several residues, each residue contains an α-carbon atom, and the coordinates of all corresponding α-carbon atoms in the two molecules are extracted by residue ID matching;
[0014] S1-2: Calculate the center of mass of the atomic coordinates of the reference molecular structure and the target molecular structure;
[0015] S1-3: Calculate the rotation matrix, perform rotation transformation on the target molecule and superimpose it on the reference molecule;
[0016] S1-4: Calculate alignment accuracy and evaluate alignment accuracy.
[0017] The corresponding α-carbon atom coordinates in S1-1 are shown as follows:
[0018] , ;
[0019] Where n is the number of matching residues, p i =(x i ,y i ,z i ) represents the three-dimensional coordinates of the i-th α-carbon atom, i = 1, 2…n;
[0020] The calculation formula in S1-2 is as follows:
[0021]
[0022] in, is the center of mass of the matching atomic coordinates in the reference molecular structure, is the mass center of the matched atomic coordinates in the target molecular structure, and N is the number of matched atoms;
[0023] By translating all coordinates to the origin, the influence of translation on the alignment result is eliminated. The calculation formula is:
[0024] ;
[0025] in, is the coordinate of the i-th matching atom relative to the mass center in the reference molecular structure, is the coordinate of the i-th matching atom in the target molecular structure relative to the center of mass.
[0026] The S1-3 includes the following sub-steps:
[0027] S1-3-1: Define the covariance matrix H, the formula is as follows:
[0028] ;
[0029] in, Represents the outer product operation of a matrix.
[0030] S1-3-2: Calculate the rotation matrix R through singular value decomposition SVD. The formula is as follows:
[0031] U,Σ,V T =SVD(H), R=VU T ;
[0032] If det(R)<0, adjust the sign of the last column of V to ensure the positive definiteness of the rotation matrix;
[0033] Among them, U is an orthogonal matrix composed of left singular vectors, V is an orthogonal matrix composed of right singular vectors, Σ is a diagonal matrix whose diagonal elements are singular values, V T is the transpose of the orthogonal matrix composed of right singular vectors, U T The transpose of the orthogonal matrix consisting of left singular vectors.
[0034] S1-3-3: Rotate the target molecule coordinates and superimpose them on the reference molecule:
[0035] ;
[0036] is the coordinate set of the target molecule after optimal rotation alignment.
[0037] S1-4: The alignment accuracy is evaluated by the root mean square deviation (RMSD) of the matched atoms, using the following formula:
[0038] ;
[0039] The RMSD reflects the overall deviation between the two groups of structures.
[0040] The S2 includes the following sub-steps:
[0041] S2-1: Binding molecular dynamics pretreatment receptor protein;
[0042] S2-2: Convert the aligned receptor protein structure from PDB format to the PDBQT format required for molecular docking. The generation process of PDBQT files introduces flexible processing of receptor proteins.
[0043] The S3 includes the following sub-steps:
[0044] S3-1: Generate a dynamic grid and calculate the center coordinates and size of the grid based on the structural characteristics of the receptor protein;
[0045] S3-2: Calculate the grid size. The grid size is calculated based on the distribution range of the active site atoms and a certain boundary extension value is added to ensure that the ligand can completely cover the active site.
[0046] S3-3: Parameters are templated through configuration files. This allows batch application to multiple docking tasks simply by modifying the template file.
[0047] S3-4: Set up multiple docking modes to explore various binding conformations of ligand and receptor;
[0048] S3-5: Set a fixed random seed to ensure that the results of each run are repeatable;
[0049] S3-6: Generate a separate working folder for each receptor and summarize the results into a unified file. Automatically analyze the docking results and generate a visual report.
[0050] The S4 includes the following sub-steps:
[0051] S4-1: Parse the PDB structure file, read the complex.pdb file, extract the atomic coordinates of the substrate, and select the key atoms in the substrate structure;
[0052] S4-2: Identify key residues in the catalytic site; set the catalytic active region of the enzyme and search for the specified residues. Suppose the catalytic site contains N atoms of key residues, whose coordinates are (x i ,y i , z i ), then the central coordinate of the catalytic site (X c , Y c , Z c ), the calculation formula is:
[0053] ;
[0054] S4-3: Calculate the distance between substrate and catalytic site, assuming the coordinates of the key atoms of the substrate are (x s ,y s , z s ), then the Euclidean distance d from the center of the catalytic site is calculated as follows:
[0055] .
[0056] The S5 comprises the following sub-steps:
[0057] S5-1: Binding free energy screening to select the top 20 protein sequences with the highest affinity;
[0058] S5-2: Substrate-catalytic site distance screening: further calculate the closest distance from the key atoms of the substrate to the catalytic site, and screen out the top 20 protein sequences with the highest distances;
[0059] S5-3: Based on the screening strategy, screen qualified ligand-receptor complexes.
[0060] The screening strategy described in S5-3 is as follows:
[0061] If the affinity of the protein in the ligand-receptor complex to the substrate is greater than that of the wild-type sequence or the distance from the substrate to the active site is less than that of the wild-type sequence, it is retained and the screening structure is output; otherwise, it is excluded.
[0062] Compared with the prior art, this application has the following beneficial effects:
[0063] This application is used to accelerate the molecular docking experimental process, optimize the molecular screening strategy, and improve the screening efficiency of receptor and ligand binding modes. It is suitable for drug screening, enzyme catalysis optimization and other fields.
[0064] This application integrates multiple open source tools (such as PyMOL, AutoDockVina, MDAnalysis, etc.) to batch process molecular docking tasks and further analyze and screen the docking results.
[0065] This application covers a fully automated process from protein structure alignment, receptor processing, molecular docking, complex generation, to catalytic site-ligand distance calculation and screening analysis. It has a high degree of automation, reduces manual operations, and improves research efficiency.
[0066] This application not only screens based on binding energy (Affinity), but also screens based on the distance between the ligand and the catalytic site, thereby improving the reliability of candidate molecules and achieving multi-dimensional screening.
[0067] This application supports the automated calculation and screening of multiple protein-ligand pairs at the same time, and outputs standardized results. It is suitable for large-scale molecular screening tasks and realizes efficient batch processing.
[0068] Based on this application, users can customize the docking parameters and screening criteria through the config.txt configuration file, such as the number of combination modes, search depth, screening distance threshold, etc., making the configuration more flexible.
[0069] This application automatically summarizes docking scores, catalytic site-ligand distances, and complex structures, supports data export and visual analysis, and has result visualization and data export functions.
[0070] The present invention can efficiently and accurately complete large-scale molecular docking tasks, providing strong technical support for fields such as drug design and enzymology research. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 Flowchart of the present invention;
[0072] Figure 2 This is a flowchart for step one;
[0073] Figure 3 This is the flowchart for step three;
[0074] Figure 4 Schematic diagram of the screening strategy. DETAILED DESCRIPTION
[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0076] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0077] In the present invention, unless otherwise specified or limited, the terms "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can mean fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0078] Reference Figure 1-Figure 4 , the automated molecular docking and screening analysis method comprises the following steps:
[0079] S1: Protein structure alignment: PyMOL was used to align protein structures to ensure that different receptors were docked in a unified coordinate system.
[0080] Furthermore, the alignment_pymol.py script is used to batch align multiple protein structures in PDB format to ensure their spatial consistency with the reference structure; the goal of this step is to align multiple protein structures on the same reference structure to facilitate subsequent molecular docking operations.
[0081] Combine Figure 2 , said S1 comprises the following sub-steps:
[0082] S1-1: Match key atom pairs; In protein structures, alpha-carbon is one of the main chain atoms of each amino acid, connecting the amino (-NH2) and carboxyl (-COOH) groups of the amino acid, and is also the connection point of the side chain groups of each amino acid. Assume that the reference molecular structure is The target molecular structure is , each molecule consists of several residues, each residue contains an α-carbon atom (CA), and the coordinates of all corresponding α-carbon atoms in the two molecules are extracted by residue ID matching;
[0083] The corresponding α-carbon atom coordinates in S1-1 are shown as follows:
[0084] , ;
[0085] Where n is the number of matching residues, p i =(x i ,y i ,z i ) represents the three-dimensional coordinates of the i-th α-carbon atom, i = 1, 2…n.
[0086] S1-2: Calculate the center of mass of the atomic coordinates of the reference molecular structure and the target molecular structure;
[0087] The calculation formula in S1-2 is as follows:
[0088]
[0089] in, is the center of mass of the matching atomic coordinates in the reference molecular structure, is the mass center of the matched atomic coordinates in the target molecular structure, and N is the number of matched atoms.
[0090] By translating all coordinates to the origin, the influence of translation on the alignment result is eliminated. The calculation formula is:
[0091] ;
[0092] in, is the coordinate of the i-th matching atom in the reference molecular structure relative to the center of mass (excluding the effect of overall translation), is the coordinate of the i-th matching atom in the target molecular structure relative to the center of mass.
[0093] This transformation translates all matched atomic coordinates to the origin so that they are calculated relative to their respective centers of mass, eliminating the effects of translational contrast alignment and retaining only rotation-related information.
[0094] S1-3: Calculate the rotation matrix, perform rotation transformation on the target molecule and superimpose it on the reference molecule; specifically, it includes the following sub-steps:
[0095] S1-3-1: Define the covariance matrix H, the formula is as follows:
[0096] ;
[0097] in, Represents the outer product operation of a matrix.
[0098] S1-3-2: Calculate the rotation matrix R through singular value decomposition (SVD), the formula is as follows:
[0099] U,Σ,V T =SVD(H), R=VU T ;
[0100] If det(R)<0, adjust the sign of the last column of V to ensure the positive definiteness of the rotation matrix.
[0101] U is an orthogonal matrix composed of left singular vectors, V is an orthogonal matrix composed of right singular vectors, Σ is a diagonal matrix, and the diagonal elements are singular values. T is the transpose of the orthogonal matrix composed of right singular vectors, U T is the transpose of the orthogonal matrix composed of left singular vectors.
[0102] S1-3-3: Rotate the target molecule coordinates and superimpose them on the reference molecule:
[0103] ;
[0104] is the coordinate set of the target molecule after optimal rotational alignment, that is, the atomic coordinates of the target molecule in the optimal alignment state.
[0105] S1-4: Calculate and evaluate the alignment accuracy. The alignment accuracy is evaluated by the root mean square deviation (RMSD) of the matched atoms, and the calculation formula is:
[0106] ;
[0107] RMSD reflects the overall deviation between two groups of structures. Users can set their own thresholds based on the similarity of the proteins, such as:
[0108] Strict alignment (high similarity proteins): RMSD < 1.0Å, suitable for homologous proteins and point mutation proteins (such as single point mutations and partial active site mutations);
[0109] Conventional alignment (functionally related proteins): RMSD < 2.0 Å, suitable for proteins with similar structures but different sequences, such as proteins in the same family;
[0110] Loose alignment (distantly homologous proteins): RMSD < 3.0~4.0Å, suitable for proteins with similar functions but large structural differences.
[0111] S2: Process the receptor protein file; identify active sites, automatically generate docking grids, and optimize the receptor structure. S2 includes the following sub-steps:
[0112] S2-1: Pre-process the receptor protein using molecular dynamics (MD). Receptor protein structures are often derived from X-ray, NMR, or AlphaFold predictions, but these structures may be conformationally unstable in solution. Therefore, before generating PDBQT, a short molecular dynamics (MD) pre-optimization, such as a 10-50 ns pre-equilibration using GROMACS, is performed to remove possible conformational stress and thus improve the biological plausibility of the receptor.
[0113] S2-2: Use the prepare_receptor.py script to convert the aligned receptor protein structure from PDB format to the PDBQT format required for molecular docking, namely:
[0114] PDBQT=PDB+Charge+Hydrogens;
[0115] Among them, the PDBQT file is a file format dedicated to the AutoDock series of tools (such as AutoDockVina), which adds the following information based on the PDB file:
[0116] Charge: Molecular docking requires calculating the electrostatic interactions between the ligand and receptor, so each atom must be assigned a charge. Charge information is typically calculated using a force field such as AMBER or GAFF.
[0117] Hydrogen atoms: Hydrogen atoms play an important role in molecular docking, especially in the calculation of hydrogen bonds and van der Waals forces. PDB files usually do not contain hydrogen atoms, so they need to be explicitly added when preparing the receptor.
[0118] During the generation of the PDBQT file (via the prepare_receptor4.py script), the receptor protein is treated flexibly. Traditional AutoDockVina is mainly used for rigid docking (RigidDocking), that is, the receptor protein is regarded as a static and unchanging structure. However, in real environments, proteins will undergo conformational changes during the molecular docking process. Therefore, flexible treatment of the receptor protein is introduced to allow limited rotation / conformational adjustments of the side chains of key residues or active pocket regions to more accurately simulate the binding mode in the biological environment. The scoring function of molecular docking is usually based on free energy calculations. Rigid docking ignores the conformational changes of the protein, while flexible docking allows specific residues to rotate, thereby being closer to the real binding mode.
[0119] S3: Batch molecular docking; use AutoDockVina for batch docking, automatically extract ligand binding conformations, and automatically generate protein-ligand complexes (complex.pdb) for subsequent analysis.
[0120] Furthermore, the docking.ipynb script was used to automate the molecular docking process, configure appropriate docking parameters, and run AutoDockVina to perform the docking calculations. This part automatically completed the docking calculations for multiple receptors and ligands and generated the docking result files.
[0121] Binding energy The calculation formula is as follows:
[0122] ;
[0123] in, 、 、 and represent the contributions of van der Waals forces, hydrogen bonds, electrostatic interactions, and desolvation, respectively.
[0124] Furthermore, combined Figure 3 , S3 includes the following sub-steps:
[0125] S3-1: Dynamic mesh generation: Based on the structural characteristics of the receptor protein (such as the size and shape of the active site), the center coordinates and size of the mesh are calculated. The center coordinates of the mesh are the geometric center of the active site atoms, and the calculation formula is as follows:
[0126] ;
[0127] ;
[0128] .
[0129] Where: x i ,y i , z i are the coordinates of the ith active site atom. N is the total number of active site atoms. 、 and are the center coordinates respectively.
[0130] S3-2: Calculate the grid size; the grid size is automatically calculated based on the distribution of the active site atoms, and a certain margin is added to ensure that the ligand can completely cover the active site. The calculation formula is as follows:
[0131] ;
[0132] ;
[0133] .
[0134] in: and are the maximum and minimum coordinates of the active site atoms in the x direction, respectively. Similarly, 、 、 、 are the maximum and minimum coordinates of the active site atoms in the y and z directions, respectively. margin is a user-defined boundary extension value, usually set to 5-10 Å to ensure that the grid can fully cover the active site. 、 、 Represents the size (length) of the grid in the x, y, and z directions, respectively, and is used to define the coverage of the grid to ensure that the active sites and ligands are fully included.
[0135] S3-3: Parameters are templated. This is achieved through a configuration file (e.g., config.txt). Users can apply this template to multiple docking tasks by simply modifying the template file. This allows batch modification of docking parameters, such as the number of binding modes and search depth.
[0136] S3-4: Supports setting multiple docking modes (num_modes) to explore various binding conformations of the ligand and receptor. By adjusting the exhaustiveness parameter, the search depth is improved, ensuring the optimal binding mode is found, and achieving high-precision search. This application supports flexible search modes and can achieve multi-mode docking.
[0137] S3-5: By setting a fixed random seed, ensure that the results of each run are repeatable, facilitating result verification and comparison. By adjusting the spacing parameters and setting a finer grid resolution, the grid resolution can be optimized to improve the accuracy of the docking results.
[0138] S3-6: Automatically generates a separate working folder for each receptor and aggregates the results into a unified file for subsequent analysis. Automatically analyzes the docking results (such as binding energy, conformational ranking, etc.) and generates a visual report.
[0139] S4: Calculate the distance between the substrate and the catalytic site; use MDAnalysis to calculate the geometric center distance between the ligand and the key residues of the catalytic site to screen the optimal binding mode. Furthermore, the distance between the substrate and the catalytic site directly affects the catalytic efficiency of the enzyme and determines whether the substrate is in the optimal reaction position. The binding energy score cannot guarantee the rationality of the substrate conformation, so calculating the distance between key atoms can be used to screen the optimal binding conformation and remove invalid docking results. This method can optimize mutant design, improve enzyme activity, and enhance the efficiency of rational enzyme engineering. MDAnalysis is used to parse the complex.pdb file obtained by molecular docking, and for the key residues of the catalytic active site (such as serine, tyrosine and lysine), the reactive atoms in the substrate are automatically identified and the Euclidean distance between them is calculated. Further, the following sub-steps are included:
[0140] S4-1: Parsing PDB structure files;
[0141] Read the complex.pdb file, extract the atomic coordinates of the substrate (UNL residue), and select the key atoms in the substrate structure (such as carbon skeleton C12).
[0142] S4-2: Identification of key residues in the catalytic site;
[0143] Set the catalytic active region of the enzyme and search for specified residues (such as THR, TYR, LYS). Assume that the catalytic site contains N atoms of key residues with coordinates (x i ,y i , z i ), then the central coordinate of the catalytic site (X c , Y c , Z c ), the calculation formula is:
[0144] .
[0145] S4-3: Calculation of substrate-catalytic site distance;
[0146] Assume that the coordinates of the key atom of the substrate (such as C12) are (x s ,ys , z s ), then the Euclidean distance d from the center of the catalytic site is calculated as follows:
[0147] .
[0148] S5: Combined screening of binding energy and spatial relationships; automatic screening based on docking affinity and catalytic site distance to remove unreasonable binding modes and improve screening accuracy. After completing molecular docking calculations and extracting key parameters of the ligand-receptor complex, screening is performed based on binding affinity and substrate-catalytic site distance to improve the accuracy of enzyme catalysis screening. Specifically, it includes the following sub-steps:
[0149] S5-1: Binding free energy screening: In the docking results, the affinity between the ligand and the receptor was calculated by AutoDockVina and used for preliminary screening. The top 20 protein sequences with the highest affinity were selected.
[0150] S5-2: Substrate-catalytic site distance screening; To ensure that the substrate is in a reasonable catalytic position, the closest distance from the key atom of the substrate (such as C12) to the catalytic site is further calculated, and the top 20 protein sequences with the closest distance (the closer the better) are screened.
[0151] S5-3: Based on the screening strategy, screen for qualified ligand-receptor complexes; refer to Figure 4 , the screening strategy is as follows:
[0152] If the affinity of the protein in the ligand-receptor complex to the substrate is greater than that of the wild-type sequence or the distance from the substrate to the active site is less than that of the wild-type sequence, it will be retained and the screening results will be automatically output for further molecular dynamics simulation or experimental verification; ligand-receptor complexes that do not meet the above conditions will be excluded.
[0153] The wild type refers to the natural protein sequence that has not been mutated, and the calculation basis of its structure and substrate interaction is determined by [molecular docking / experimental determination].
[0154] S6: Results summary and visualization; generate binding energy distribution diagram, distance distribution diagram, calculation time comparison diagram, and export the screened candidate molecule list.
[0155] The above descriptions are merely optional embodiments of the present invention and do not limit the patent scope of the present invention. All equivalent structural transformations made using the contents of the present invention specification under the concept of the present invention, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present invention.
Claims
1. An automated molecular docking and screening analysis method, characterized in that: The steps include: S1: Protein structure alignment; batch align multiple protein structures to ensure their spatial consistency with the reference structure; S2: Process the receptor protein file; identify active sites, automatically generate docking grids, and optimize the receptor structure; S2 includes the following sub-steps: S2-1: Binding molecular dynamics pretreatment receptor protein; S2-2: Convert the aligned receptor protein structure from PDB format to the PDBQT format required for molecular docking. The generation process of PDBQT files introduces flexible processing of receptor proteins. S3: Batch molecular docking: perform batch docking, automatically extract ligand binding conformations, and automatically generate protein-ligand complexes to facilitate subsequent analysis; S3 includes the following sub-steps: S3-1: Generate a dynamic grid and calculate the center coordinates and size of the grid based on the structural characteristics of the receptor protein; S3-2: Calculate the grid size. The grid size is calculated based on the distribution range of the active site atoms and a boundary extension value is added to ensure that the ligand can completely cover the active site. S3-3: Parameters are templated through configuration files. This allows batch application to multiple docking tasks simply by modifying the template file. S3-4: Set up multiple docking modes to explore various binding conformations of ligand and receptor; S3-5: Set a fixed random seed to ensure that the results of each run are repeatable; S3-6: Generate a separate working folder for each receptor and aggregate the results into a unified file. Automatically analyze the docking results and generate a visual report. S4: Calculate the substrate-catalytic site distance; Calculate the geometric center distance between the ligand and the key residues in the catalytic site to screen the optimal binding mode; S5: Combined screening of binding energy and spatial relationship; automatic screening based on docking affinity and catalytic site distance to remove unreasonable binding modes; S5 includes the following sub-steps: S5-1: Binding energy screening, screening out the top 20 protein sequences with the highest affinity; S5-2: Substrate-catalytic site distance screening: Calculate the shortest distance from the key atoms of the substrate to the catalytic site and select the top 20 protein sequences with the highest distances; S5-3: Screen qualified protein-ligand complexes based on the screening strategy; The screening strategy described in S5-3 is as follows: If the affinity between the protein and the substrate of the protein-ligand complex is greater than that of the wild-type sequence or the distance from the substrate to the active site is less than that of the wild-type sequence, it is retained and the screening structure is output; otherwise, it is excluded; S6: Results summary and visualization; generate binding energy distribution diagram, distance distribution diagram, calculation time comparison diagram, and export the screened candidate molecule list.
2. The automated molecular docking and screening analysis method according to claim 1, characterized in that: The S1 includes the following sub-steps: S1-1: Match key atom pairs; assume the reference molecular structure is The target molecular structure is , each molecule consists of several residues, each residue contains an α-carbon atom, and the coordinates of all corresponding α-carbon atoms in the two molecules are extracted by residue ID matching; S1-2: Calculate the center of mass of the atomic coordinates of the reference molecular structure and the target molecular structure; S1-3: Calculate the rotation matrix, perform rotation transformation on the target molecule and superimpose it on the reference molecule; S1-4: Calculate alignment accuracy and evaluate alignment accuracy.
3. The automated molecular docking and screening analysis method according to claim 2, characterized in that: The corresponding α-carbon atom coordinates in S1-1 are shown as follows: , ; Where n is the number of matching residues, p i =(x i ,y i ,z i ) represents the three-dimensional coordinates of the i-th α-carbon atom, i = 1, 2…n; The calculation formula in S1-2 is as follows: in, is the center of mass of the matching atomic coordinates in the reference molecular structure, is the mass center of the matched atomic coordinates in the target molecular structure, and N is the number of matched atoms; By translating all coordinates to the origin, the influence of translation on the alignment result is eliminated. The calculation formula is: ; in, is the coordinate of the i-th matching atom relative to the mass center in the reference molecular structure, is the coordinate of the i-th matching atom in the target molecular structure relative to the center of mass.
4. The automated molecular docking and screening analysis method according to claim 3, characterized in that: The S1-3 includes the following sub-steps: S1-3-1: Define the covariance matrix H, the formula is as follows: ; in, Represents the outer product operation of the matrix; S1-3-2: Calculate the rotation matrix R through singular value decomposition SVD. The formula is as follows: U, Σ, V T =SVD(H), R=VU T ; If det(R)<0, adjust the sign of the last column of V to ensure the positive definiteness of the rotation matrix; Among them, U is an orthogonal matrix composed of left singular vectors, V is an orthogonal matrix composed of right singular vectors, Σ is a diagonal matrix whose diagonal elements are singular values, V T is the transpose of the orthogonal matrix composed of right singular vectors, U T is the transpose of the orthogonal matrix composed of left singular vectors; S1-3-3: Rotate the target molecule coordinates and superimpose them on the reference molecule. The formula is as follows: ; is the coordinate set of the target molecule after optimal rotation alignment.
5. The automated molecular docking and screening analysis method according to claim 4, characterized in that: S1-4: The alignment accuracy is evaluated by the root mean square deviation (RMSD) of the matched atoms, using the following formula: ; The RMSD reflects the overall deviation between the two groups of structures.
6. The automated molecular docking and screening analysis method according to claim 1, characterized in that: The S4 includes the following sub-steps: S4-1: Parse the PDB structure file, read the complex.pdb file, extract the atomic coordinates of the substrate, and select the key atoms in the substrate structure; S4-2: Identify key residues in the catalytic site; set the catalytic active region of the enzyme and search for the specified residues. Suppose the catalytic site contains N atoms of key residues, whose coordinates are (x i ,y i , z i ), then the central coordinate of the catalytic site (X c , Y c , Z c ), the calculation formula is: ; S4-3: Calculate the distance between substrate and catalytic site, assuming the coordinates of the key atoms of the substrate are (x s ,y s , z s ), then the Euclidean distance d from the center of the catalytic site is calculated as follows: 。
Citation Information
Patent Citations
Molecular modification design method for improving catalytic efficiency of prolyl endopeptidase
CN111540404A
Molecular docking method and system based on AutoDock visual platform
CN115295091A