A protein pocket detection system, method, and storage medium
By constructing a protein pocket detection system, molecular dynamics simulation is used to generate protein conformational trajectories and determine dynamic pockets, solving the problem of difficulty in discovering dynamic protein targets in existing technologies. This achieves efficient and accurate target identification and intuitive display, and lowers the technical threshold.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies cannot efficiently and conveniently discover key drug targets generated by proteins in dynamic processes. Traditional methods rely on static structure detection, which carries the risk of missing important targets. Moreover, the complex and time-consuming manual operation is difficult for researchers without a computational background to master.
A protein pocket detection system is provided, including a conformation generation module, a pocket detection module, a task scheduling and management module, and a visualization module. It generates the conformational trajectory of the protein through molecular dynamics simulation, determines the pocket data in the dynamic process, and renders the output. It integrates the workflow of task scheduling and management, conformation generation, pocket detection, and visualization.
It enables efficient and accurate identification of stable pockets in the dynamic conformation of proteins, improves the efficiency and accuracy of drug target discovery, lowers the barrier to entry, provides intuitive data display, and significantly improves the efficiency and accuracy of drug target discovery.
Smart Images

Figure CN121302837B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of protein detection technology, and in particular to a protein pocket detection system, method, and storage medium. Background Technology
[0002] Proteins are the core executors of life activities, and their functions are closely related to their three-dimensional structure, especially the "pockets" or "binding sites" that bind to ligands. Traditional drug discovery and target identification heavily rely on static protein structures obtained through techniques such as X-ray crystallography and nuclear magnetic resonance. However, this approach has inherent limitations: proteins are dynamic in their physiological environment, and their structures are constantly in motion and undergoing conformational changes. Many crucial drug targets, such as allosteric pockets or cryptic pockets, are usually closed or invisible in static snapshots, only forming or opening under specific conformations during their dynamic changes. Therefore, traditional pocket detection methods that rely on a single static structure risk missing a large number of potentially important drug targets.
[0003] Currently, a complete dynamic pocket analysis workflow requires researchers to manually operate multiple independent software tools to sequentially complete simulation, trajectory processing, sampling, pocket calculation, and visualization analysis. This process is complex, time-consuming, and prone to errors. For researchers in biology or chemistry without a computational background, the technical challenges of configuring supercomputing environments, writing job scripts, and managing data flow are insurmountable.
[0004] In summary, existing technologies cannot efficiently and conveniently discover key drug targets generated by proteins during dynamic processes. Summary of the Invention
[0005] To address, or at least partially address, the aforementioned technical problems, this disclosure provides a protein pocket detection system, method, and storage medium.
[0006] This disclosure provides a protein pocket detection system, which includes: a conformation generation module, a pocket detection module, a task scheduling and management module, and a visualization module.
[0007] The task scheduling and management module acquires the protein data to be detected and coordinates and controls the execution of the conformation generation module, the pocket detection module and the visualization module.
[0008] The conformation generation module, in response to the scheduling of the task scheduling and management module, generates the conformation trajectory of the protein data based on the protein data obtained by the task scheduling and management module using molecular dynamics simulation; wherein, the conformation trajectory is used to characterize the structural change process of the protein data in the process of performing biological functions.
[0009] The pocket detection module, in response to the scheduling of the task scheduling and management module, determines pocket data based on the conformation trajectory generated by the conformation generation module; wherein, the pocket data is used to characterize the stability of the spatial position and geometric shape of the protein data during the change of the conformation trajectory;
[0010] The visualization module, in response to the scheduling of the task scheduling and management module, renders and outputs the image trajectory generated by the image generation module and the pocket data determined by the pocket detection module.
[0011] Optionally, the task scheduling and management module generates the conformational trajectory of the protein data based on the protein data obtained by the task scheduling and management module, including:
[0012] Determine the data type of the protein data;
[0013] Based on the aforementioned data type, a molecular dynamics simulation method is determined;
[0014] The conformational trajectory is generated using the aforementioned molecular dynamics simulation.
[0015] Optionally, determining the molecular dynamics simulation method based on the data type includes:
[0016] When the data type is a first data type, the molecular dynamics simulation method is determined to be a first molecular dynamics simulation method; wherein, the first data type is used to characterize protein sequence data;
[0017] When the data type is a second data type, the molecular dynamics simulation method is determined to be a second molecular dynamics simulation method; wherein, the second data type is used to characterize protein structure data, and the structures of the first molecular dynamics simulation method and the second molecular dynamics simulation method are different.
[0018] Optionally, the pocket detection module determines pocket data based on the conformation trajectory generated by the conformation generation module, including:
[0019] Multiple conformational samples are extracted from the conformational trajectory;
[0020] Simultaneously, pocket detection is performed on multiple conformational samples to obtain multiple pocket detection results;
[0021] The pocket detection results are clustered and / or filtered to determine the pocket data.
[0022] Optionally, the conformation generation module further includes:
[0023] Based on the protein data obtained by the task scheduling and management module, topological data of the protein data is generated.
[0024] Optionally, the visualization module further includes:
[0025] In response to a selection and / or query operation of the pocket data, the selected pocket data is extracted from the pocket data;
[0026] The selected pocket data is highlighted and output, and the attribute information of the selected pocket data is also output; wherein, the attribute information includes at least one of the following: average pocket volume, drug-likeness score, hydrophobicity parameter, and frequency of occurrence.
[0027] Optionally, the visualization module further includes:
[0028] During the process of outputting the conformation trajectory generated by the conformation generation module, the pocket data corresponding to the current output frame is identified, and the pocket data corresponding to the current output frame is synchronously highlighted and output.
[0029] This disclosure provides a protein pocket detection method, applied to a protein pocket detection system, the method comprising:
[0030] Obtain the protein data to be detected;
[0031] Based on the protein data, molecular dynamics simulations were used to generate the conformational trajectory of the protein data.
[0032] Based on the conformational trajectory, pocket data is determined; wherein, the pocket data is used to characterize the stability of the spatial position and geometric shape of the protein data during the change of the conformational trajectory;
[0033] The conformation trajectory and the pocket data are rendered and output.
[0034] This disclosure provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the protein pocket detection method.
[0035] This disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the protein pocket detection method.
[0036] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: the task scheduling and management module coordinates data flow, the conformation generation module uses molecular dynamics simulation to generate conformational trajectories that reflect structural changes in the biological function process of proteins, the pocket detection module determines the spatial location and geometrically stable pocket data from the conformational trajectories, and finally the visualization module realizes the rendering output of the conformational trajectories and pocket data. Overall, it can efficiently capture stable pockets under the dynamic conformation of proteins and intuitively present relevant data. It can efficiently and accurately identify stable pockets from the dynamic conformation of proteins and intuitively display them through interactive visualization, which significantly improves the efficiency and accuracy of drug target discovery, while greatly reducing the threshold for use. Attached Figure Description
[0037] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0038] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of the structure of the protein pocket detection system according to an embodiment of the present disclosure;
[0040] Figure 2 This is a schematic flowchart of the protein pocket detection method according to an embodiment of the present disclosure;
[0041] Figure 3 This is a schematic flowchart of the pocket detection method according to an embodiment of the present disclosure;
[0042] Figure 4 This is a schematic flowchart of the three-layer synergistic framework for protein pocket detection according to an embodiment of this disclosure. Detailed Implementation
[0043] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0044] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0045] Figure 1This is a schematic diagram of the structure of the protein pocket detection system according to an embodiment of the present disclosure, as shown below. Figure 1 As shown, the detection system 100 includes: a conformation generation module 101, a pocket detection module 102, a task scheduling and management module 103, and a visualization module 104.
[0046] The task scheduling and management module 103 acquires the protein data to be detected and coordinates the automated execution of the conformation generation module 101, the pocket detection module 102, and the visualization module 104.
[0047] Specifically, the protein data to be detected can cover two core types: one is protein sequence data (e.g., FASTA format data), which directly reflects the amino acid composition and sequence of the protein; the other is protein three-dimensional structure data (e.g., PDB format data), which presents the spatial configuration of the protein molecule, including details such as atomic positions and chemical bond connections. These two data types correspond to the selection criteria for different molecular dynamics simulation methods in the subsequent conformation generation module 101 of the system, providing initial data support for the subsequent targeted generation of conformational trajectories.
[0048] The task scheduling and management module 103, as the core control hub of the system, is responsible for initiating the entire protein pocket detection process. This module receives input protein data through a user interface, which can be protein sequence information or three-dimensional structure data.
[0049] After receiving the data, the task scheduling and management module 103 can perform the following key operations: First, it verifies the validity of the input protein data to ensure the standardization and integrity of the data format; then, it creates an independent task identifier and initializes the task status; next, it temporarily stores the verified data in the storage area specified by the system to prepare for subsequent calculation processes.
[0050] The task scheduling and management module 103 also undertakes the function of intelligent allocation of computing resources. Based on the characteristics of the protein data and the current system load, it dynamically plans the usage strategy of computing resources and generates corresponding computing job configurations. Through a connection established with the application programming interface of the high-performance computing platform, it submits the computing tasks and related parameters to the computing cluster.
[0051] During the task execution phase, the task scheduling and management module 103 continuously monitors the status changes of the computation job, including queuing status, running progress, and completion status. This status monitoring mechanism allows users to understand the task processing progress in real time, while also providing the system with the ability to promptly detect abnormal situations.
[0052] For example, the backend defines a RESTful API, and the frontend submits tasks via POST requests. The request body includes the input file and the selected molecular dynamics simulation method (bio-emu or AI). 2 The backend includes BMD (Browser Memory Management) parameters, and has built-in supercomputing job script templates for bio-emu and fpocket parallel computing. Upon receiving a task, the backend program fills the template with the task's specific parameters (e.g., input file path, CPU / GPU node requests) to generate an executable job script. Using the supercomputing's unified API, the backend submits the generated script to the computing cluster. After submission, the backend immediately returns a task ID to the frontend and begins periodically polling the task's status in the supercomputing queue. The task status is updated in real-time to the backend database, and the frontend can query the latest status using the task ID.
[0053] Once the computation task is completed, the task scheduling and management module 103 automatically retrieves the result data from the computing platform, performs standardization processing on it, and updates the task status to completion. Finally, the computation results are transmitted to the visualization module 104 through a standardized data interface, completing the entire data processing flow.
[0054] Through this series of coordinated management tasks, the task scheduling and management module 103 effectively decouples user interaction from complex backend calculations, providing users with a simple and clear operating experience, while at the same time achieving optimized utilization of computing resources and efficient management of task processes at the underlying level.
[0055] As the core computing unit of the system, the conformation generation module 101 responds to the scheduling of the task scheduling and management module 103. Based on the protein data obtained by the task scheduling and management module 103, it uses molecular dynamics simulation to generate the conformation trajectory of the protein data. The conformation trajectory is used to characterize the structural change process of the protein data in the process of performing biological functions.
[0056] Specifically, the conformation generation module 101 employs an intelligent model selection strategy. First, it identifies the type of the input protein data and automatically selects the most suitable molecular dynamics simulation method based on the data type characteristics. Specifically, when the protein data is of the first data type (protein sequence information), the system uses the first molecular dynamics simulation method for processing; when the protein data is of the second data type (protein three-dimensional structure information), it uses a second molecular dynamics simulation method with a different structure for analysis. This differentiated molecular dynamics simulation method selection mechanism ensures that all types of input data can achieve accurate and reliable conformation trajectory generation. The first data type can correspond to protein sequence data (e.g., FASTA format data reflecting the amino acid sequence), while the second data type can correspond to protein structure data (e.g., PDB format data presenting the molecular spatial configuration). The difference between the two data types directly determines the direction of subsequent molecular dynamics simulation selection.
[0057] During trajectory generation, this module executes a refined computational process: it performs standardized preprocessing on the input protein data, including necessary steps such as atomic nomenclature normalization and structural integrity checks; it performs depth sampling and generation calculations of the protein conformation space according to the selected molecular dynamics simulation parameters; and it optimizes the calculation results through specific algorithms to ensure that the generated conformational trajectories can fully reflect the natural dynamic characteristics of the protein while maintaining physical rationality.
[0058] For example, when the protein data is of the primary data type, i.e., protein sequence information, the system uses the bio-emu model for processing. This model is based on deep learning methods, employing a pre-trained AlphaFold2 sequence encoder to extract features from the input sequence and combining multiple sequence alignment (MSA) to generate characterization information. Subsequently, using a coarse-grained structure representation module, only the heavy atoms of the protein backbone are retained, and a local coordinate system is established through orthogonalization, thereby obtaining an efficient representation of the backbone framework. On this basis, bio-emu gradually removes noise through a diffusion condition generation model, generating protein conformations that conform to the target distribution. Simultaneously, a score model is introduced to comprehensively evaluate multi-dimensional information, ensuring the accuracy and stability of the predicted conformations. Finally, the system can output trajectory files (*.xtc) and topology files (*.pdb) containing multiple candidate conformations. This workflow avoids dependence on long-duration physical simulations and significantly shortens the computation cycle. bio-emu is mainly suitable for structural sampling of monomeric proteins; for polymers or complexes containing multiple chains, the system can use linkertrick or template-based methods for modeling.
[0059] For another example, when the protein data is of the second data type, namely the three-dimensional structure information of the protein, the system will prioritize calling AI.2 The BMD model first standardizes the input structure, including adding hydrogen atoms, capping the N-terminus and C-terminus (ACE and NME), and adjusting atom nomenclature using tools to ensure compatibility. Subsequently, the protein data undergoes solvation and energy minimization, and can be pre-equilibrated using either the FF19SB or AMOEBA force field scheme, depending on requirements, to obtain a reasonable initial configuration. In the core simulation phase, AI... 2 BMD breaks down proteins into fragmented units and uses a VisNet-based machine learning potential function to calculate the total energy and atomic forces of the protein data at each time step, thereby driving molecular dynamics evolution with DFT-level ab initio accuracy. In addition to the default fragmentation mode, AI... 2 BMD also supports the overall ViSNet model, enabling direct modeling of complete protein molecules, but requires pre-training by the user and provision of corresponding potential parameters. Ultimately, the process generates trajectory files (.traj) in ASE format, which can be converted to the commonly used .dcd format using provided tools for further analysis and visualization. Compared to traditional molecular mechanics methods, AI... 2 BMD demonstrates higher quantum chemical precision in energy and atomic force calculations, and maintains good consistency with experimental data at the kinetic and thermodynamic levels.
[0060] The conformation generation module 101 also includes: generating topological data of the protein data based on the protein data acquired by the task scheduling and management module 103; the topological data covers the basic structural information of the protein molecule, such as atomic connection mode, amino acid residue positional relationship, molecular skeleton features, etc., and can be used as supplementary data for conformational trajectory; on the one hand, the topological data can provide the basic framework of protein structure for the subsequent pocket detection module, assisting the module in analyzing the spatial position and geometric shape of the pocket; on the other hand, it can also provide basic information for the visualization module to render the protein structure, ensuring the complete presentation of dynamic trajectory and static structural features.
[0061] The pocket detection module 102, as a key analysis unit of the system, responds to the scheduling of the task scheduling and management module 103 and determines the pocket data based on the conformation trajectory generated by the conformation generation module 101. The pocket data is used to characterize the stable spatial position and geometric shape of protein data during the change of conformation trajectory.
[0062] Specifically, the pocket detection module 102 extracts multiple conformational samples from the conformational trajectory; simultaneously performs pocket detection on the multiple conformational samples to obtain multiple pocket detection results; and performs clustering and / or filtering processing on the pocket detection results to determine the pocket data.
[0063] The execution process of the pocket detection module 102 includes three core steps: First, it systematically extracts multiple representative conformational samples from the conformational trajectory. These samples are not randomly selected, but are uniformly sampled according to a preset time interval or frame number (e.g., one structure is selected every fixed simulation step size). This ensures that the selected samples can cover the main conformational change states of the protein, avoiding both biased analysis due to insufficient sample quantity and increased computational burden due to sample redundancy.
[0064] For the extracted multiple conformational samples, the pocket detection module 102 performs pocket detection simultaneously using parallel processing. This approach does not detect individual samples sequentially, but rather allocates the detection tasks for multiple conformational samples to different computing units, initiating detection operations simultaneously. Through parallel detection, tasks that would have required several hours of serial processing can be completed within minutes, significantly improving detection efficiency. Furthermore, the detection of each sample is based on a unified pocket recognition logic, ensuring the consistency and comparability of detection results for different samples, ultimately yielding the pocket detection result for each conformational sample.
[0065] Because of structural differences in samples with different conformations, the detected pockets may overlap or differ in location and morphology. To screen out pockets with stable spatial location and geometry during protein dynamic changes, the pocket detection module 102 further processes all pocket detection results: on the one hand, it uses a clustering algorithm to group pockets with similar spatial locations and geometric features into one category, which can be considered as the same pocket in different conformations; on the other hand, it combines indicators such as pocket occurrence frequency, volume stability, hydrophobicity parameters, and drug-likeness scores to filter out pocket categories with low occurrence frequency, large morphological fluctuations, or poor drug-likeness. Through the synergistic processing of clustering and filtering, the final pocket data can accurately characterize the pocket features that maintain stability during protein dynamic conformational changes, providing a reliable target reference for subsequent visualization analysis and drug design.
[0066] The visualization module 104, as an important interactive interface of the system, responds to the scheduling of the task scheduling and management module 103 and is responsible for rendering and outputting the image trajectory generated by the image generation module 101 and the pocket data determined by the pocket detection module 102.
[0067] The visualization module 104 first performs basic rendering processing on the conformational trajectories and pocket data. For conformational trajectories, the protein backbone structure is presented in a cartoon or ribbon mode, and the trajectory is played dynamically to fully demonstrate the structural changes of the protein during the performance of its biological functions, such as the movement of the molecular skeleton and the opening and closing of domains. For pocket data, it is superimposed on the protein structure in a semi-transparent surface mode, with different colors used to distinguish different pockets, allowing users to quickly locate the spatial distribution of pockets on the protein molecule and intuitively perceive the geometric shape of the pockets. Through this rendering method, the originally complex structural data is transformed into visualized images, lowering the user's understanding threshold.
[0068] The visualization module 104 further includes: in response to the selection and / or query operation of the pocket data, extracting the selected pocket data from the pocket data; highlighting the selected pocket data and outputting the attribute information of the selected pocket data; wherein the attribute information includes at least one of the following: average pocket volume, drugability score, hydrophobicity parameter, and frequency of occurrence.
[0069] Specifically, to meet users' targeted analysis needs, the visualization module 104 supports interactive operations on pocket data. When a user selects or queries specific pocket data through the interface, the visualization module 104 responds to these operations in real time: on the one hand, it extracts the selected pocket data from the overall pocket data and presents it in a highlighted form in the 3D view (e.g., increasing opacity, changing edge color) to clearly distinguish it from other pockets and help users focus on the target pocket; on the other hand, it synchronously outputs the attribute information of the pocket data, which includes the average pocket volume, drug-likeness score, hydrophobicity parameter, frequency of occurrence, etc., providing data support for users to deeply judge the research value of the pocket.
[0070] The visualization module 104 also includes: identifying the pocket data corresponding to the current output frame during the process of generating the conformation trajectory by the output conformation generation module 101, and synchronously highlighting the pocket data corresponding to the current output frame.
[0071] Specifically, considering the dynamic characteristics of conformational trajectories, the visualization module 104 also has the ability to synchronously display trajectory and pocket data. During the playback of the conformational trajectory, the visualization module 104 identifies the pocket data corresponding to the current output frame (i.e., the protein structure at a certain moment) in real time. For example, if a specific pocket is formed by the opening of the protein conformation in a certain frame, the visualization module 104 will immediately highlight that pocket synchronously in the 3D view, allowing users to intuitively observe the correlation between changes in protein structure and the appearance / disappearance / morphological adjustment of pockets. This synchronous highlighting design helps users clearly capture the patterns of pocket changes with the dynamic conformation of the protein, such as some pockets only appearing in specific conformational states or maintaining morphological stability during trajectory playback, providing a direct reference for understanding the dynamic mechanism of drug binding.
[0072] In this embodiment, after the task scheduling and management module 103 acquires protein data, it transmits the data to the conformation generation module 101. The conformation generation module 101 generates conformational trajectories and topological data, and transmits the conformational trajectories to the pocket detection module 102. The pocket detection module 102 determines the pocket data based on the conformational trajectories, and then the conformational trajectories and pocket data are transmitted together to the visualization module 104. The visualization module 104 renders and interactively displays the two types of data to achieve a comprehensive analysis and presentation of the dynamic pockets of the protein.
[0073] Based on the protein pocket detection system described above Figure 2 This is a schematic flowchart of the protein pocket detection method according to an embodiment of this disclosure. Applied to the aforementioned protein pocket detection system, such as... Figure 2 As shown, the protein pocket detection method includes the following steps:
[0074] S201. Obtain the protein data to be detected.
[0075] Specifically, the system receives protein information submitted by users for analysis through a task scheduling and management module. The system supports two main data input formats: one is protein sequence data, which records the amino acid composition and sequence of the protein, serving as the core basis for subsequent sequence-based prediction of the protein's dynamic structure; the other is protein structure data, which presents the three-dimensional spatial configuration of the protein molecule, including details such as atomic positions, chemical bond connections, and the spatial distribution of amino acid residues, providing a foundation for optimizing existing structures and simulating dynamic changes.
[0076] After receiving the data, the system will verify the integrity and format of the data, create an independent identifier and record for the analysis task, and include it in the task queue for unified management.
[0077] S202. Based on protein data, molecular dynamics simulation is used to generate conformational trajectories of protein data.
[0078] Specifically, the protein data acquired by S201 is type-identified to determine whether it belongs to protein sequence data or protein structure data. The structural characteristics of the two data types directly determine the choice of molecular dynamics simulation method.
[0079] Match the appropriate molecular dynamics simulation method to the data type. If the protein data is protein sequence data, select a molecular dynamics simulation method adapted to sequence analysis (e.g., the bio-emu model); if the data is protein structure data, select a molecular dynamics simulation method based on optimization of existing structures (e.g., AI). 2 The two types of molecular dynamics simulation methods (BMD model) have different structures and can achieve efficient conformation generation for data features respectively.
[0080] By using selected molecular dynamics simulations, conformational trajectories (e.g., xtc, traj format files) are generated, containing structural snapshots of proteins at different time points or functional states. These trajectories fully represent the dynamic structural changes of proteins, including main chain motion, domain opening and closing, and local conformational adjustments, providing comprehensive structural evidence for subsequent detection of dynamic pockets. Simultaneously, this step also generates topological data (e.g., .pdb format files containing atomic connection patterns and molecular skeleton features) based on the protein data, serving as supplementary information to the conformational trajectories and assisting in subsequent analysis.
[0081] S203. Determine pocket data based on conformational trajectory.
[0082] Among them, pocket data is used to characterize the spatial location and geometric stability of protein data during the process of conformational trajectory changes.
[0083] Specifically, from the conformational trajectories generated by S202, uniform sampling is performed at preset time intervals or frame numbers (e.g., extracting 100 conformations) to select multiple conformational samples that can cover the main conformational change states of the protein. This avoids both biased analysis due to insufficient sample size and increased computational burden due to sample redundancy.
[0084] By leveraging high-performance computing resources, the pocket detection tasks of multiple extracted conformational samples are allocated to different computing units, and pocket detection operations are initiated simultaneously (e.g., identifying pockets by analyzing the geometric features and hydrophobicity of protein surface depressions), significantly reducing detection time and obtaining pocket detection results for each conformational sample.
[0085] All pocket detection results are processed in multiple dimensions to determine the final pocket data. First, a clustering algorithm is used to group pocket detection results that are spatially close and have similar geometric features into one category, which is regarded as the same pocket under different conformations. Then, the pocket occurrence frequency, volume stability, hydrophobicity parameter, and drug-likeness score are combined to filter and remove pocket categories with low occurrence frequency, large morphological fluctuations, and poor drug-likeness. The final pocket data can accurately characterize the pocket features with stable spatial position and geometric shape of proteins during conformational trajectory changes.
[0086] For example, from the .xtc and .traj conformation trajectory files generated by the dynamic conformation generation module, uniform sampling is performed according to a preset time interval or number of frames (e.g., extracting 100 conformations) to obtain a series of representative protein three-dimensional structure snapshots; the pocket detection task of the N sampled structure snapshots is parallelized using a job scheduling system. Figure 3 This is a schematic flowchart of the pocket detection method according to an embodiment of the present disclosure, as shown below. Figure 3 As shown, the user provides protein data, the cloud server initiates a task to the supercomputing system, the supercomputing system predicts the dynamic structure of the protein after delivering the task, and then the pockets are detected in parallel by multiple Fpocket tools. Finally, the results are fed back to the user, forming a complete closed loop.
[0087] The specific implementation is as follows: the main job script submits N independent fpocket subtasks to the scheduling system, each subtask is assigned to a different computing core or node, and simultaneously performs calculations on a structural snapshot. This massively parallel processing method shortens the computation task that originally required several hours of serial processing to be completed within minutes. After all fpocket subtasks are completed, an aggregation script automatically collects the pocket detection results of all conformations. The algorithm performs spatial location clustering on pockets appearing in different frames, and ranks them comprehensively based on indicators such as frequency of occurrence, volume, hydrophobicity, and druggability score. Finally, a group of "Consensus Pockets" with the highest probability, strongest stability, and best druggability is selected, and each pocket is marked with the conformations in which it appears.
[0088] S204. Render and output the conformation trajectory and pocket data.
[0089] Specifically, a professional visualization engine is used to perform basic rendering of conformational trajectories and pocket data. For conformational trajectories, the protein main chain structure is displayed in a cartoon or ribbon mode, and the trajectory is played dynamically to fully present the structural change process of the protein. For pocket data, it is superimposed on the protein structure in a semi-transparent surface mode, and different colors are used to distinguish different pockets, helping users quickly locate the spatial distribution of pockets and intuitively perceive the geometric shape of the pockets.
[0090] For example, when a user requests to view the results of a completed task, the front end retrieves the URL of the result file from the back end. Mol* components can simultaneously load the protein topology (*.pdb), dynamic trajectories (*.xtc, *.traj), and integrated pocket data (e.g., generating a PDB file or geometric data containing all occurrences of each consensus pocket).
[0091] It also supports users to select or query pocket data. When a user triggers an operation, the selected pocket data is extracted from the pocket data in real time and presented in a highlighted form in the 3D view. At the same time, the attribute information of the pocket is output, providing data support for users to conduct in-depth analysis of the pocket research value.
[0092] During the playback of the conformational trajectory, the pocket data corresponding to the current output frame is identified in real time, and these pocket data are highlighted synchronously, allowing users to intuitively observe the correlation between protein structural changes and the appearance, disappearance, and morphological adjustment of pockets. This clearly captures the pattern of pocket changes with conformation, providing an intuitive reference for understanding drug binding mechanisms and target selection.
[0093] In the above scheme, diverse protein data covering sequences and structures are acquired to lay the foundation for analysis. Topological data and conformational trajectories reflecting structural changes in protein biological functions are generated by combining data type-adapted molecular dynamics simulation methods. After conformational sample extraction, parallel detection, and cluster screening, stable pocket data in terms of spatial location and geometry are determined. The conformational trajectories and pocket data are presented intuitively through basic rendering, interactive display, and dynamic synchronous highlighting. The overall process not only efficiently captures stable pockets in the dynamic conformation of proteins, but also significantly reduces the usage threshold for non-computational personnel. It can also provide clear and valuable references for drug target discovery, drug mechanism research, and drug design, significantly improving the efficiency and practicality of protein pocket analysis.
[0094] The present disclosure will be further illustrated below through specific embodiments.
[0095] Figure 4 This is a schematic flowchart of the three-layer synergistic architecture for protein pocket detection according to an embodiment of this disclosure, as shown below. Figure 4 As shown, the architecture includes an application layer, a scheduling layer, and a computing layer.
[0096] Specifically, the application layer provides a user interface, including task submission and preview, dynamic interaction of protein molecules, and multi-source data loading functions. It serves as the entry point for user interaction with the system, responsible for data import, task initiation, and result preview. It supports dynamic visualization analysis of protein molecules and meets the loading needs of various types of data.
[0097] The application layer is built on web technologies and can be accessed by users through a browser. Users can submit protein sequences (such as in FASTA format) or 3D structure files (such as in PDB format). After task submission, the page does not require waiting; users can view the real-time status of the task (queued, computing, completed) in the "Task Management" interface. For completed tasks, an integrated 3D preview window is provided, using the Mol* engine to load and render the results.
[0098] The scheduling layer is deployed on the cloud server and is responsible for submitting tasks, interacting with the front end, and managing the task queue. It is the task hub of the system, responsible for receiving application layer task requests, interacting with the front end interface, queuing and prioritizing tasks to ensure orderly execution; at the same time, it provides temporary storage services for user files.
[0099] The scheduling layer uses the Django REST Framework architecture, providing standard API interfaces for frontend calls. This layer does not perform computationally intensive tasks; it is only responsible for logic processing and task scheduling. Upon receiving a task, the backend dynamically generates a job script to be submitted to the supercomputing machine based on the input type (sequence or PDB) and parameters, and initiates the submission via the API. This layer periodically queries the task status and updates the database for frontend queries.
[0100] The computing layer consists of a supercomputing cluster. It possesses parallel computing and GPU-accelerated computing capabilities and is the core of the system's computing power. Parallel computing technology improves the efficiency of large-scale data processing, while GPU acceleration further enhances support for computationally intensive tasks such as protein dynamics simulations and pocket detection, including the use of bio-emu or AI. 2 BMD's dynamic configuration generation and parallel fpocket detection provide efficient computing power for upper-layer functions.
[0101] The computing layer pre-deploys bio-emu and AI on the supercomputing platform. 2 Core computing software includes BMD and fpocket. The scheduling layer initiates computation by submitting customized job scripts (such as Slurm scripts). Computational tasks are executed on the supercomputer's CPU or GPU nodes, fully utilizing their powerful parallel computing capabilities. After computation, result files (such as xtc, pdb, and pocket information files) are stored in the supercomputer's shared file system, awaiting access and retrieval by the scheduling layer via API.
[0102] The following example, analyzing the dynamic pocket of a kinase target, further illustrates this disclosure. The specific implementation process is as follows:
[0103] Users log in to the web platform and go to the task submission page. They select "Submit via Sequence," paste the FASTA sequence of the kinase into the input box, and click the "Submit" button.
[0104] The frontend sends ordinal data to the backend (scheduling layer). The backend validates the input, creates a new task, and submits a task based on bio-emu or AI via API. 2 BMD's calculation task.
[0105] The supercomputing (computing layer) receives jobs and runs bio-emu or AI on GPU nodes. 2 BMD, the dynamic trajectory of protein formation.
[0106] The task script then extracts 100 conformations from the trajectory and launches 100 parallel fpocket subtasks for pocket detection.
[0107] After all subtasks are completed, the aggregation script clusters and sorts the results, generating the final pocket data file and a summary report.
[0108] Users can see the task status change to "Completed" on the "Task Management" page and click "View Results".
[0109] The front-end loads the Mol* visualization component and retrieves the URL of the result file from the back-end.
[0110] Mol* simultaneously renders the dynamic structure of the kinase and the top 5 high-probability pockets detected (displayed on surfaces of different colors).
[0111] Users can click the play button to observe the opening and closing of the pockets during the movement of the kinase; click "Pocket 1" in the pocket list, and the corresponding pocket in the view will become opaque and highlighted, while the list will display information such as the average volume and drug-likeness score of the pocket; rotate and zoom the view to carefully observe the fine structure of the pocket and the surrounding amino acid residues from different angles.
[0112] In summary, this disclosure, through technological innovation and process integration, constructs a fully automated dynamic pocket analysis solution from input to visualization, significantly reducing the technical threshold, improving analysis efficiency, and providing a powerful computational tool for drug discovery.
[0113] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements steps such as the protein pocket detection method.
[0114] This embodiment also provides a computer-readable storage medium (including but not limited to disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.) storing computer program code. When the computer program code is run on a computer, the computer executes the above-mentioned related method steps to implement the protein pocket detection method provided in the above embodiment.
[0115] The beneficial effects of the above embodiments can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0116] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0117] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A protein pocket detection system characterized by, The detection system comprises a conformation generation module, a pocket detection module, a task scheduling and management module and a visualization module: The task scheduling and management module acquires protein data to be detected, and coordinates control of execution of the conformation generation module, the pocket detection module and the visualization module, wherein the protein data to be detected comprises protein sequence data and protein three-dimensional structure data, the protein sequence data reflects the amino acid composition and arrangement order of the protein, and the protein three-dimensional structure data presents the spatial configuration of the protein molecule; The conformation generation module generates a conformation trajectory of the protein data by using molecular dynamics simulation based on the protein data acquired by the task scheduling and management module in response to scheduling of the task scheduling and management module, wherein the conformation trajectory is used to represent the structural change process of the protein data in the execution of biological functions. The pocket detection module determines pocket data based on the conformation trajectory generated by the conformation generation module in response to scheduling of the task scheduling and management module, wherein the pocket data is used to represent data stable in spatial position and geometric shape in the change process of the conformation trajectory. The visualization module renders and outputs the conformation trajectory generated by the conformation generation module and the pocket data determined by the pocket detection module in response to scheduling of the task scheduling and management module. The conformation generation module is further configured to: determine a data type of the protein data; when the data type is a first data type, determine that the molecular dynamics simulation method is a first molecular dynamics simulation method, wherein the first data type is used to represent protein sequence data; when the data type is a second data type, determine that the molecular dynamics simulation method is a second molecular dynamics simulation method, wherein the second data type is used to represent protein structure data, and the first molecular dynamics simulation method and the second molecular dynamics simulation method are different in structure; and generate the conformation trajectory by using the molecular dynamics simulation.
2. The protein pocket detection system of claim 1, wherein, The pocket detection module determines pocket data based on the conformation trajectory generated by the conformation generation module, comprising: extracting a plurality of conformation samples from the conformation trajectory; simultaneously performing pocket detection on a plurality of the conformation samples to obtain a plurality of pocket detection results; performing clustering and / or screening processing on the pocket detection results to determine the pocket data.
3. The protein pocket detection system of claim 1, wherein, The conformation generation module further comprises: generating topological data of the protein data based on the protein data acquired by the task scheduling and management module.
4. The protein pocket detection system of claim 1, wherein, The visualization module further comprises: in response to selection and / or query operation on the pocket data, extracting selected pocket data from the pocket data; highlighting and outputting the selected pocket data, and outputting attribute information of the selected pocket data, wherein the attribute information comprises at least one of pocket average volume, drugability score, hydrophobicity parameter and appearance frequency.
5. The protein pocket detection system of claim 1, wherein, The visualization module further comprises: In the process of outputting the conformation trajectory generated by the conformation generation module, pocket data corresponding to the current output frame is identified, and the pocket data corresponding to the current output frame is synchronously highlighted and output.
6. A method of protein pocket detection, characterized by, The method is applied to the protein pocket detection system according to any one of claims 1-5, and the method comprises: obtaining protein data to be detected; generating a conformation trajectory of the protein data by using molecular dynamics simulation based on the protein data; determining pocket data based on the conformation trajectory, wherein the pocket data is used to represent data of stable spatial position and geometric shape of the protein data in the change process of the conformation trajectory; rendering and outputting the conformation trajectory and the pocket data.
7. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps of the protein pocket detection method according to claim 6.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the protein pocket detection method according to claim 6.
Citation Information
Patent Citations
Molecular generation and optimization method based on protein large language model
CN120877841A
Analysis method and system for revealing hidden binding pocket of drug target
CN121096423A